Paper deep dive
Effective Skill Unlearning through Intervention and Abstention
Yongce Li, Chung-En Sun, Tsui-Wei Weng
Models: Gemma-2b, Llama-2-7b, Llama-3-70b, Llama-3-8b
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/12/2026, 6:37:24 PM
Summary
The paper introduces two training-free, lightweight machine unlearning techniques for Large Language Models (LLMs): 'Neuron Adjust' and 'Key Space Detection'. These methods target Feed-Forward Layers (FFLs) to remove specific skills (e.g., math, coding) while preserving general model capabilities. Neuron Adjust probabilistically shifts neuron pre-activation distributions, while Key Space Detection blocks query vectors from accessing skill-specific hypercubes in the FFL key space.
Entities (5)
Relation Signals (3)
Neuron Adjust ā targets ā Feed-Forward Layer
confidence 95% Ā· Our methods involve operations on the FFL in large pretrained autoregressive transformer decoder models
Neuron Adjust ā performsunlearningvia ā Intervention
confidence 90% Ā· we propose two lightweight, training-free skill unlearning methods via intervention and abstention respectively: Neuron Adjust and Key Space Detection.
Key Space Detection ā performsunlearningvia ā Abstention
confidence 90% Ā· we propose two lightweight, training-free skill unlearning methods via intervention and abstention respectively: Neuron Adjust and Key Space Detection.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language Models (LLMs) have demonstrated remarkable skills across various domains. Understanding the mechanisms behind their abilities and implementing controls over them is becoming increasingly important for developing better models. In this paper, we focus on skill unlearning in LLMs, specifically unlearning a particular skill while retaining their overall capabilities. We introduce two lightweight, training-free machine skill unlearning techniques for LLMs. First, we observe that the pre-activation distribution of neurons in each Feed-Forward Layer (FFL) differs when the model demonstrates different skills. Additionally, we find that queries triggering the same skill cluster within the FFL key space and can be separated from other queries using a hypercube. Based on these observations, we propose two lightweight, training-free skill unlearning methods via \textit{intervention} and \textit{abstention} respectively: \texttt{Neuron Adjust} and \texttt{Key Space Detection}. We evaluate our methods on unlearning math-solving, Python-coding, and comprehension skills across seven different languages. The results demonstrate their strong unlearning capabilities for the designated skills. Specifically, \texttt{Key Space Detection} achieves over 80\% relative performance drop on the forgetting skill and less than 10\% relative performance drop on other skills and the model's general knowledge (MMLU) for most unlearning tasks. Our code is available at this https URL
Tags
Links
Trouble viewing inline? Open PDF directly ā
Full Text
62,483 characters extracted from source content.
Expand or collapse full text
Effective Skill Unlearning through Intervention and Abstention Yongce Li UCSD HDSI yol013@ucsd.edu &Chung-En Sun UCSD CSE cesun@ucsd.edu &Tsui-Wei Weng UCSD HDSI lweng@ucsd.edu Abstract Large language Models (LLMs) have demonstrated remarkable skills across various domains. Understanding the mechanisms behind their abilities and implementing controls over them is becoming increasingly important for developing better models. In this paper, we focus on skill unlearning in LLMs, specifically unlearning a particular skill while retaining their overall capabilities. We introduce two lightweight, training-free machine skill unlearning techniques for LLMs. First, we observe that the pre-activation distribution of neurons in each Feed-Forward Layer (FFL) differs when the model demonstrates different skills. Additionally, we find that queries triggering the same skill cluster within the FFL key space and can be separated from other queries using a hypercube. Based on these observations, we propose two lightweight, training-free skill unlearning methods via intervention and abstention respectively: Neuron Adjust and Key Space Detection. We evaluate our methods on unlearning math-solving, Python-coding, and comprehension skills across seven different languages. The results demonstrate their strong unlearning capabilities for the designated skills. Specifically, Key Space Detection achieves over 80% relative performance drop on the forgetting skill and less than 10% relative performance drop on other skills and the modelās general knowledge (MMLU) for most unlearning tasks. 111Our code is available at https://github.com/Trustworthy-ML-Lab/effective_skill_unlearning Effective Skill Unlearning through Intervention and Abstention Yongce Li UCSD HDSI yol013@ucsd.edu Chung-En Sun UCSD CSE cesun@ucsd.edu Tsui-Wei Weng UCSD HDSI lweng@ucsd.edu (I) Efficiency (I) Performance (I) Scalability Method: Does not require training No Inference time cost High quality unlearning Maintain model overall capability Applicable to large models Retrain from scratch No Yes Yes Yes No Fine-tuning based unlearning No Yes Yes No Yes In-context unlearning Yes Yes No Yes Yes Selective Pruning Yes Yes Yes No Yes Neuron Adjust (Ours) Yes Oā¢(1)1O(1)O ( 1 ) Yes Partial Yes Key Space Detection (Ours) Yes Oā¢(1)1O(1)O ( 1 ) Yes Yes Yes Table 1: Comparison of our method against existing machine unlearning methods, including retraining, fine-tuning based methods, in-context unlearning, and a prune-based unlearning method Selective Pruning Pochinkov and Schoots (2023). 1 Introduction In recent years, the superior capabilities demonstrated by Large Language Models (LLMs) have attracted significant research interest. Without training on task-specific datasets, LLMs exhibit strong skills in various domains such as math Wei et al. (2022); Imani et al. (2023); Cobbe et al. (2021), coding Austin et al. (2021); Li et al. (2022), and language comprehension Shi et al. (2023). Understanding the mechanisms behind these abilities and implementing controls over them are becoming increasingly important for developing stronger, safer, and more interpretable models. Figure 1: An overview of the proposed skill unlearning methods: Neuron Adjust (through intervention) and Key Space Detection (through abstention). This example illustrates forgetting coding skill. A recent line of research focuses on machine unlearning Yao et al. (2023); Liu et al. (2024), which aims to remove the knowledge LLMs have acquired from specific datasets while maintaining their causally unrelated knowledge. In this work, we focus on a variant of machine unlearning called skill unlearning, which aims to remove a specific skill (e.g., coding skill, elementary math-solving skill) from the LLM while retaining its other skills. Skill unlearning helps researchers control certain behaviors of LLMs, providing insights into when and how a model demonstrates a particular skill. Currently, most unlearning methods Lu et al. (2022); Jang et al. (2023); Wang et al. (2023); Yu et al. (2023); Eldan and Russinovich (2023); Chen and Yang (2023); Yao et al. (2023) rely on fine-tuning, which becomes increasingly costly as LLMs grow larger. Other unlearning methods Wu et al. (2023); Pochinkov and Schoots (2023) involve pruning dataset-related sets of neurons, which we show can harm the modelās overall capabilities. In this paper, we introduce two new machine skill unlearning methods that are training-free and have minimally impact the modelās overall capabilities. We first observe that feed-forward layer neurons exhibit different pre-activation distributions when the model demonstrates different skills. Based on this observation, in section 3 we propose Neuron Adjust, which probabilistically shifts neuron pre-activation values to retain the desired skill distribution during inference through intervention. By considering the correlation among neurons, we further observe that neuron activation vectors cluster within different hypercubes in the feed-forward layerās key space when the model demonstrates certain skills. Building on this, we introduce Key Space Detection (KSD) in section 4, which detects and blocks specific skill-related activations in the key space through abstention by preventing query vectors from accessing the skill-specific hypercube. Our contributions can be summarized as follows: 1. Motivated by the shift in neuron pre-activation distributions and the modularity of skill-triggering queries in the feed-forward layerās key space, we propose two novel machine skill unlearning methods, Neuron Adjust and Key Space Detection, which are scalable, training-free, and maintain the modelās overall capabilities with minimal degradation. 2. Our experiments on math, code, and language skill unlearning demonstrate the effectiveness of the proposed two methods with >80%absentpercent80>80\%> 80 % relative performance drop on the target forgetting skill and <10%absentpercent10<10\%< 10 % drop on the modelās general knowledge (MMLU Hendrycks et al. (2021)) and other skills. Specifically, Key Space Detection achieves nearly perfect skill unlearning with negligible overall capability drop. Table 1 compares our methods with traditional unlearning methods and the skill unlearning method Selective Pruning Pochinkov and Schoots (2023) across the dimensions of Efficiency, Performance, and Scalability. Figure 2: An overview of the Neuron Adjust method in section 3. Before adjusting a neuron, the neuron has different distributions under the forgetting and retaining datasets. During inference time, Neuron Adjust algorithm will edit neurons with large distribution shift such that its pre-activation distribution will be close to the retaining distribution. 2 Related Work & Background Large language model machine unlearning. This line of work aims to remove the influence of specific data points and the corresponding model capabilities without retraining the model from scratch. Most previous works on machine unlearning have focused on fine-tuning-based approaches Lu et al. (2022); Jang et al. (2023); Wang et al. (2023); Yu et al. (2023); Eldan and Russinovich (2023); Chen and Yang (2023); Yao et al. (2023), which become increasingly costly as models grow larger. Pawelczyk et al. (2024) introduced in-context unlearning, which provides contextual inputs to the language model during the inference stage. Despite its cost-efficiency, it lacks unlearning quality and is difficult to generalize to large-scale unlearning. Other training-free approaches focus on pruning or removing specific sets of behavior-related neurons in the model. DEPN Wu et al. (2023) is a pruning-based unlearning approach that removes neurons based on their cumulative privacy gradient. Selective Pruning Pochinkov and Schoots (2023) is another pruning-based method, which removes neurons based on their relative importance to the forgetting dataset and the retaining dataset. However, the extent to which pruning-based methods affect the modelās overall capabilities remains unknown and unjustifiable. Feed-forward neuron interpretability. This line of work focuses on the interpretation of individual neurons, meaning that individual neurons represent meaningful concepts, both in vision models Bau et al. (2020); Hernandez et al. (2022); Oikarinen and Weng (2023) and language models Bills et al. (2023); Lee et al. (2023); Sun et al. (2024). Recent works have shown that neurons exhibit multisemanticity Elhage et al. (2022); Bricken et al. (2023); Huben et al. (2024), with some being expressible as a linear combination of concepts Oikarinen and Weng (2024). By considering neuron activation vectors, we can also treat LLMās feed-forward layers (FFLs) as key-value memories, with neuron activation vectors as keys and the output of the FFLs as values Geva et al. (2021); Meng et al. (2022). Unlearning settings. In this paper, we focus on unlearning a specific skill or capability of a language model while retaining another. In the following sections, we denote Drā¢eā¢tā¢aā¢iā¢nsubscriptD_retainDitalic_r e t a i n as the dataset capturing the skill we want the model to retain performance, and Dfā¢oā¢rā¢gā¢eā¢tsubscriptD_forgetDitalic_f o r g e t as the dataset capturing the skill we want the model to forget. Our methods involve operations on the FFL in large pretrained autoregressive transformer decoder models, which takes the layer-normed input zāāHsuperscriptāz ^Hz ā blackboard_RH from the residual stream: FFL(l)ā¢(z)=Wdown(l)ā¢Ļā¢(Wup(l)ā¢z),superscriptFFLsuperscriptsubscriptdownsuperscriptsubscriptup split FFL^(l)(z)=W_down^(l)Ļ (W_% up^(l)z ), splitstart_ROW start_CELL FFL( l ) ( z ) = Wdown( l ) Ļ ( Wup( l ) z ) , end_CELL end_ROW where z is first mapped to a higher-dimensional space by an up-projection linear transformation Wup(l)superscriptsubscriptupW_up^(l)Wup( l ) and a non-linear activation function Ļ to obtain neuron activations, and then mapped back to āHsuperscriptāR^Hblackboard_RH space with a down-projection linear transformation Wdown(l)superscriptsubscriptdownW_down^(l)Wdown( l ). Modern LLMs also utilize gated linear units (GLUs) in FFLs. Instead of directly activating each neuron, GLUs use a gating mechanism to control the information flow of each neuron: FFL(l)ā¢(z)=Wdown(l)ā¢(Ļā¢(Wgate(l)ā¢z)āWup(l)ā¢z),superscriptFFLsuperscriptsubscriptdowndirect-productsuperscriptsubscriptgatesuperscriptsubscriptup split FFL^(l)(z)=W_down^(l) (Ļ (W_% gate^(l)z ) W_up^(l)z ), splitstart_ROW start_CELL FFL( l ) ( z ) = Wdown( l ) ( Ļ ( Wgate( l ) z ) ā Wup( l ) z ) , end_CELL end_ROW where ādirect-product ā is element-wise vector multiplication. In the following sections, we consider Wup(l)ā¢zsuperscriptsubscriptupW_up^(l)zWup( l ) z in traditional FFL and Wgate(l)ā¢zsuperscriptsubscriptgateW_gate^(l)zWgate( l ) z in GLU-FFL as the neuron pre-activations, and vectors after activation function as the key vectors, i.e. we have vkā¢eā¢y(l)=Ļā¢(Wup(l)ā¢z)subscriptsuperscriptsuperscriptsubscriptupv^(l)_key=Ļ (W_up^(l)z )v( l )k e y = Ļ ( Wup( l ) z ) and vkā¢eā¢y(l)=Ļā¢(Wgate(l)ā¢z)āWup(l)ā¢zsubscriptsuperscriptdirect-productsuperscriptsubscriptgatesuperscriptsubscriptupv^(l)_key=Ļ (W_gate^(l)z ) W_up^(l% )zv( l )k e y = Ļ ( Wgate( l ) z ) ā Wup( l ) z in regular FFL and GLU-FFL respectively. Figure 3: Category distribution over different pre-activation value ranges for neuron at layer 17, index 693 (left), and layer 0, index 13366 (right). Figure 4: Overview of the Key Space Detection method in section 4. a) shows the structure of decoder-based large language models. b) shows the components of a GLU-based feed-forward layer in the LLM, where vkeysubscriptkeyv_keyvkey is located in the key space we aim to prune. c) is an example of a 3-neuron key space. The blue hypercube is formed by μāD±αā¢ĻāDplus-or-minussubscriptāsubscriptā\ μ_D±α Ļ_D\ overā start_ARG μ end_ARGD ± α overā start_ARG Ļ end_ARGD , where μāDsubscriptā μ_Doverā start_ARG μ end_ARGD and ĻāDsubscriptā Ļ_Doverā start_ARG Ļ end_ARGD are the sample mean vector and standard deviation vector of vkeysubscriptkeyv_keyvkey when probing the model with the forgetting dataset. During every inference step, if we detect vkeyāμāD±αā¢ĻāDsubscriptkeyplus-or-minussubscriptāsubscriptāv_keyā\ μ_D±α Ļ_D\vkey ā overā start_ARG μ end_ARGD ± α overā start_ARG Ļ end_ARGD , we prohibit the model from generating the output. 3 Inference Time Neuron Adjustment In this section, we introduce Neuron Adjust, a post-hoc, training-free machine unlearning technique for large language models achieved by inference time neuron pre-activation value adjustment. An overview of Neuron Adjust method is shown in Figure 2. In section 3.1, we show the motivation of the method that some neuronsā pre-activation distributions differ when the model demonstrates different capabilities. Based on the observation, we describe the Neuron Adjust method in section 3.2. 3.1 Case Study: Neuron Pre-Activation Distribution Shift We perform a case study to show how neuron pre-activation distribution changes when the model demonstrates math and coding skills separately. We choose GSM8K Cobbe et al. (2021) and MBPP Austin et al. (2021) as the two datasets that characterize the math and Python coding skills of the model, and Gemma-2b-it Team et al. (2024) as the subject model to study. We probe the model with the two datasets, and document each neuronās (token, pre-activation) pair, and then prompt GPT-4 OpenAI et al. (2024) to categorize tokens into meaningful categories as in Figure 3. For the two neurons, although they are highly activated by "Programming-related keywords," "Operator/Syntax tokens," and "Tokens with Parentheses," they both show positive activations for "Numerical-related tokens." This case study demonstrates that neurons are multi-functional, and simply pruning them would be harmful to the modelās overall capabilities. 3.2 Neuron Adjust Algorithm Based on the observation that neuron pre-activations exhibit different distributions when the model demonstrates different skills and the polysemantic property of neurons, we propose Neuron Adjust, a probabilistic skill unlearning technique applied to the subject model during inference time. Neuron Adjust unlearns one skill of the model while retaining another skill by shifting each neuronās pre-activation from the forgetting skill distribution to the retaining skill distribution. Algorithm 1 shows the pseudo-code of Neuron Adjust, which mainly consists of two parts: Algorithm 1 Neuron Adjust Algorithm for neuron nisubscriptn_initalic_i and inference time pre-activation v 1:Input: Drā¢eā¢tā¢aā¢iā¢nsubscriptD_retainDitalic_r e t a i n, Dfā¢oā¢rā¢gā¢eā¢tsubscriptD_forgetDitalic_f o r g e t, v 2:Probing the model with Drā¢eā¢tā¢aā¢iā¢nsubscriptD_retainDitalic_r e t a i n and Dfā¢oā¢rā¢gā¢eā¢tsubscriptD_forgetDitalic_f o r g e t, approximate sample mean and std: ni|Drā¢eā¢tā¢aā¢iā¢nā¼ā¢(μr,Ļr)similar-toconditionalsubscriptsubscriptsubscriptsubscriptn_i|D_retain ( _r, _r)nitalic_i | Ditalic_r e t a i n ā¼ N ( μitalic_r , Ļitalic_r ) ni|Dfā¢oā¢rā¢gā¢eā¢tā¼ā¢(μf,Ļf)similar-toconditionalsubscriptsubscriptsubscriptsubscriptn_i|D_forget ( _f, _f)nitalic_i | Ditalic_f o r g e t ā¼ N ( μitalic_f , Ļitalic_f ) 3:Calculate pr=Pā¢(v|ā¢(μr,Ļr))subscriptconditionalsubscriptsubscriptp_r=P(v|N( _r, _r))pitalic_r = P ( v | N ( μitalic_r , Ļitalic_r ) ) 4:Calculate pf=Pā¢(v|ā¢(μf,Ļf))subscriptconditionalsubscriptsubscriptp_f=P(v|N( _f, _f))pitalic_f = P ( v | N ( μitalic_f , Ļitalic_f ) ) 5:if pr<pfsubscriptsubscriptp_r<p_fpitalic_r < pitalic_f then 6: αāpfpr+pfāsubscriptsubscriptsubscriptαā p_fp_r+p_fα ā divide start_ARG pitalic_f end_ARG start_ARG pitalic_r + pitalic_f end_ARG 7: vadjustā2ā¢Ī¼rā(vāμfĻfā¢Ļr+μr)āsubscriptadjust2subscriptsubscriptsubscriptsubscriptsubscriptv_adjustā 2 _r- ( v- _f _f _% r+ _r )vadjust ā 2 μitalic_r - ( divide start_ARG v - μitalic_f end_ARG start_ARG Ļitalic_f end_ARG Ļitalic_r + μitalic_r ) with probability α 8: vadjustāvāsubscriptadjustv_adjustā vvadjust ā v with probability 1āα11- 1 - α 9:else 10: vadjustāvāsubscriptadjustv_adjustā vvadjust ā v 11:end if 12:Output: vadjustsubscriptadjustv_adjustvadjust 1. Probe the model with the forgetting and retaining datasets. For each neuron, assume that the forgetting and retaining pre-activation distributions follow a normal distribution. Approximate the means and standard deviations (stds) of these two distributions using sample pre-activation values. 2. During inference, when a neuron has a pre-activation value v, calculate the likelihood of v being sampled from each of the two distributions. If v is more likely to be sampled from the retaining distribution, keep the value. Otherwise, shift v towards the retaining distribution. Additionally, take a symmetric adjustment based on the mean of the retaining distribution (this step serves as an adaptive penalty), with a probability based on how likely v is sampled from the forgetting distribution. 4 Feed-Forward Layers Key Space Hypercube Detection One limitation of Neuron Adjust is that it treats each neuron individually, without considering their correlations. However, neurons often work together in a coordinated way, contributing to the modelās overall behavior. In this section, we first present our observation that query vectors evoking a specific skill tend to cluster in the key space of the feed-forward layers, as described in Section 4.1. Building on this clustering phenomenon, we introduce Key Space Detection (KSD) in Section 4.2, a machine unlearning technique that prevents query embedding vectors from accessing designated hypercubes. Figure 5: Percentage of Query Vectors contained in the Hypercube μāgsm8k±αā¢Ļāgsm8kplus-or-minussubscriptāgsm8ksubscriptāgsm8k\ μ_gsm8k±α Ļ_gsm8k\ overā start_ARG μ end_ARGgsm8k ± α overā start_ARG Ļ end_ARGgsm8k 4.1 Neurons are Correlated Features in High-Dimensional Space Figure 6: Relationship of the smallest hypercube containing all query vectors across layers. The left figure shows that different skill vectors cluster more tightly as the layer gets deeper. The right figure shows the distance between the centers of two hypercubes under different metrics. As the vākeysubscriptākey v_keyoverā start_ARG v end_ARGkey in Figure 4(b), we define the vector before the down-projection matrix WdownsubscriptdownW_downWdown in each FFL as the neuron activation vector, which forms the key space of each FFL. Each FFL has a unique key space. We use llama-3-8b as the subject language model, MBPP as the probing dataset that triggers the modelās Python coding skill, and GSM8K as the probing dataset that triggers the modelās grade school math problem-solving skill. We probe the model with the two training datasets and document two probing activation vector sets of the last token of each query, vā(l)mbppsubscriptsuperscriptāmbpp\ v^(l)\_mbpp overā start_ARG v end_ARG( l ) mbpp and vā(l)gsm8ksubscriptsuperscriptāgsm8k\ v^(l)\_gsm8k overā start_ARG v end_ARG( l ) gsm8k, for the lthsuperscriptthl^thlth FFL. We calculate the mean and std vector, μāD(l)superscriptsubscriptā μ_D^(l)overā start_ARG μ end_ARGD( l ) and ĻāD(l)superscriptsubscriptā Ļ_D^(l)overā start_ARG Ļ end_ARGD( l ), of each vā(l)Dsubscriptsuperscriptā\ v^(l)\_D overā start_ARG v end_ARG( l ) D, where (μāD(l))iā1|D|ā¢āj=1|D|(vāj(l))iāsubscriptsuperscriptsubscriptā1superscriptsubscript1subscriptsuperscriptsubscriptā( μ_D^(l))_i 1|D| _j=1^|D|( v_j^(% l))_i( overā start_ARG μ end_ARGD( l ) )i ā divide start_ARG 1 end_ARG start_ARG | D | end_ARG āj = 1| D | ( overā start_ARG v end_ARGj( l ) )i (ĻāD(l))iā1|D|ā¢āj=1|D|((vāj(l))iā(μāD(l))i)2,āsubscriptsuperscriptsubscriptā1superscriptsubscript1superscriptsubscriptsuperscriptsubscriptāsubscriptsuperscriptsubscriptā2( Ļ_D^(l))_i 1|D| _j=1^|D| (% ( v_j^(l))_i-( μ_D^(l))_i )^2,( overā start_ARG Ļ end_ARGD( l ) )i ā square-root start_ARG divide start_ARG 1 end_ARG start_ARG | D | end_ARG āj = 1| D | ( ( overā start_ARG v end_ARGj( l ) )i - ( overā start_ARG μ end_ARGD( l ) )i )2 end_ARG , and bound the two vector sets with hypercubes μāD±αā¢ĻāDāuā|μāD(l)āαā¢ĻāD(l)āŗuāāŗĪ¼āD(l)+αā¢ĻāD(l),āplus-or-minussubscriptāsubscriptāconditional-setāprecedessuperscriptsubscriptāsuperscriptsubscriptāprecedessuperscriptsubscriptāsuperscriptsubscriptā\ μ_D±α Ļ_D\\! \!\ u\;|\; μ% _D^(l)-α Ļ_D^(l)\! \! u\! \! μ_D^% (l)+α Ļ_D^(l)\, overā start_ARG μ end_ARGD ± α overā start_ARG Ļ end_ARGD ā overā start_ARG u end_ARG | overā start_ARG μ end_ARGD( l ) - α overā start_ARG Ļ end_ARGD( l ) āŗ overā start_ARG u end_ARG āŗ overā start_ARG μ end_ARGD( l ) + α overā start_ARG Ļ end_ARGD( l ) , where α is a hyperparameter that controls the size of the hypercube, and āŗprecedes āŗ denotes element-wise less-than comparison. Figure 5 shows the percentage of vectors contained in the hypercube μāgsm8k±αā¢Ļāgsm8kplus-or-minussubscriptāgsm8ksubscriptāgsm8k\ μ_gsm8k±α Ļ_gsm8k\ overā start_ARG μ end_ARGgsm8k ± α overā start_ARG Ļ end_ARGgsm8k when increasing α from 00 to 30303030 in the last layer of the model. We observe that when α=1515α=15α = 15, nearly all math query vector embeddings are encompassed within the hypercube. As α increases from 15151515 to 20202020, a gap forms between the math and code query clusters: all math queries remain within the hypercube, but no code queries are included. When α exceeds 20202020, a few code queries begin to fall within the hypercube. We further analyze the changes in size and distance within and between the math query cluster and the code query cluster across different layers. Specifically, for each layer l, we calculate the smallest hypercube that encompass all math and code queries, respectively, and compare their volumes and the distance between their centers. Figure 6 (left) shows the log ratio of the volume of the lthsuperscriptthl^thlth layer hypercube to the volume of the first layer hypercube. As the layers get deeper, the hypercube becomes smaller, indicating denser clustering. Figure 6 (right) shows the Euclidean, Manhattan, and cosine distances between μāgsm8k(l)superscriptsubscriptāgsm8k μ_gsm8k^(l)overā start_ARG μ end_ARGgsm8k( l ) and μāmbpp(l)superscriptsubscriptāmbpp μ_mbpp^(l)overā start_ARG μ end_ARGmbpp( l ) for each l. We observe that the Euclidean and Manhattan distances gradually increase as the layers get deeper, except for high peaks in the very first and last layers. Additionally, for most layers, the cosine distance fluctuates around 1.0, indicating the orthogonality of the two clusters. Figure 7: Performance of Neuron Adjust and Key Space Detection on Math/Code Skill Unlearning. On the horizontal axis, NA, SP, and KSD stand for Neuron Adjust (ratio), Selective Pruning (ratio), and Key Space Detection, respectively. The vertical axis represents the relative performance of the model compared to the original model after applying each unlearning method. Figure 8: Results of unlearning one language in MLQA dataset while retaining the others with Neuron Adjust 5% and Key Space Detection on llama-3-8b. The i-th row shows the modelās performance drop on different languages after unlearning the i-th language. Figure 9: A case in a 2-neuron key space to explain how inference-time unrelated key vector would be affected by Neuron Adjust (left, get adjusted) and Key Space Detection (right, unaffected). In this case, an adjustment to v is unfavorable. 4.2 Machine Unlearning via Key Space Detection The idea of Key Space Detection is as follows: we first identify the sample mean vector μāD(l)superscriptsubscriptā μ_D^(l)overā start_ARG μ end_ARGD( l ) and the sample standard deviation vector ĻāD(l)superscriptsubscriptā Ļ_D^(l)overā start_ARG Ļ end_ARGD( l ) of each FFL activation by probing the model with the forgetting dataset D. Then, we create a hypercube μāD(l)±αā¢ĻāD(l)plus-or-minussuperscriptsubscriptāsuperscriptsubscriptā\ μ_D^(l)±α Ļ_D^(l)\ overā start_ARG μ end_ARGD( l ) ± α overā start_ARG Ļ end_ARGD( l ) in the key space as in section 4, where α is a hyperparameter we select to balance the trade-off between the quality of forgetting and maintaining the overall capability of the model. During inference, if we detect a query vector falling within the hypercube, we abstain the modelās output and replace it with a system message "Your query is not valid." instead. Algorithm 2 KSD Algorithm for layer l 1:Input: size hyperparameter α, inference time activation key vector vkeysubscriptkeyv_keyvkey, forgetting dataset D 2:Probing the model with D, estimate sample mean and std vectors μāD(l),ĻāD(l).superscriptsubscriptāsuperscriptsubscriptā μ_D^(l), Ļ_D^(l).overā start_ARG μ end_ARGD( l ) , overā start_ARG Ļ end_ARGD( l ) . 3:During inference step k, current output o: 4:if vkeyāμāD(l)±αā¢ĻāD(l)subscriptkeyplus-or-minussuperscriptsubscriptāsuperscriptsubscriptāv_keyā\ μ_D^(l)±α Ļ_D^(l)\vkey ā overā start_ARG μ end_ARGD( l ) ± α overā start_ARG Ļ end_ARGD( l ) then 5: oāabsento ā "Your query is not valid." 6: Stop inference. 7:else 8: o⢠+= tokenksubscript += tokeno += token_ko += tokenitalic_k. 9: Continue with the next token inference step k+11k+1k + 1. 10:end if 11:Output: o 5 Experiments In this section, we present experimental results of math/code skill unlearning in Section 5.1 and language unlearning in Section 5.2. We test the performance of Neuron Adjust (NA) and Key Space Pruning (KSD) on each skill unlearning task. For the NA method, we rank neurons by their difference in distribution mean, μfāμrsubscriptsubscript _f- _rμitalic_f - μitalic_r, and select the top β neurons with β set to 0.5%, 1.5%, and 3.0%. For the KSD method, we choose size coefficient α such that KSD either matches or outperforms the best forgetting quality of Neuron Adjust. For the math/code skill unlearning task, we use Selective Pruning Pochinkov and Schoots (2023), a pruning-based skill unlearning method, as our baseline. This method prunes neurons in the FFLs based on their relative importance to each dataset. After skill unlearning, we want the model to unlearn only the specific skill while maintaining its overall capabilities. Therefore, we also evaluate the effect of each method on 5-shot MMLU accuracy. For each experiment, we run it three times and report the best result. 5.1 Math/Code Skill Unlearning For math/code skill unlearning, we choose the training splits of MBPP and GSM8K as the forgetting datasets. The MBPP dataset contains code-based programming problems for evaluating LLMsā Python code generation abilities, while the GSM8K dataset consists of grade school-level math problems for assessing their problem-solving skills in elementary mathematics. We use these two datasets to capture the modelās Python coding skills and elementary math problem-solving skills. After unlearning, we test the models on the testing split of each dataset. For Python coding skills, we also test on MBPP+ Liu et al. (2023), which contains 35x more Python test cases to evaluate the robustness of the unlearning process. Figure 7 compares the performance of the unlearning methods on four models: Gemma-2b Team et al. (2024), Llama-2-7b Touvron et al. (2023), Llama-3-8b, and Llama-3-70b AI@Meta (2024). NA, tested at different adjustment ratios, shows that higher ratios lead to more forgetting with minimal impact on MMLU accuracy, though MBPP retention decreases slightly. KSD achieves the highest forgetting rates with negligible effects on MMLU accuracy and retention, proving its efficiency in unlearning while maintaining overall performance. In contrast, SP effectively forgets math skills and retains coding performance but produces nonsensical output when asked to forget coding skills while retaining math, even with just 0.01%percent0.010.01\%0.01 % neuron pruning. This suggests pruning can harm overall model capabilities, likely due to shared neurons between coding and math tasks. In this case, both NA and KSD outperform SP. 5.2 Language Skill Unlearning For the language skill unlearning task, we use the MLQA Lewis et al. (2019) dataset as our evaluation benchmark. We use each languageās context data as the forgetting dataset. Figure 8 shows the results of NA (left) and KSD (right) in a heatmap view. The kthsuperscriptthk^thkth row of the heatmap indicates the percentage decrease in each languageās performance after forcing the model to forget the kthsuperscriptthk^thkth language on the vertical axis. Ideally, we aim to maximize the diagonal entries while keeping the other entries small. From the heatmap, we observe that NA performs well in forgetting English (en), Spanish (es), and Hindi (hi). However, when forgetting German (de), Chinese (zh), Vietnamese (vi), and Arabic (ar), the performance of one or more languages in the retaining dataset also decreases significantly. This suggests that the model may utilize a shared set of neuron values when demonstrating these languages. In contrast, KSD shows a better ability to maintain the modelās performance on other languages while achieving substantial forgetting quality on most of the languages. After forgetting each language, we also tested the modelās performance on the MMLU task. Both methods demonstrate a performance decrease of less than 5%percent55\%5 %. We also observe that KSD consistently outperforms other methods in retaining the modelās overall capability. Figure 9 illustrates the reason. For an out-of-distribution knowledge query vector v, NA may adjust some of its dimensions because it only considers operations for single neurons. In contrast, KSD considers the correlations among all neurons, prohibiting query vectors from accessing a much smaller area in the key space. Therefore, it is guaranteed to have no negative effect on out-of-hypercube queries. 6 Conclusion In this paper, we propose two lightweight, training-free machine skill unlearning methods, Neuron Adjust and Key Space Detection, which have minimal impact on the modelās overall capabilities. Neuron Adjust achieves unlearning by shifting the pre-activation of feed-forward layer neurons from the forgetting distribution to the retaining distribution. KSD achieves unlearning by prohibiting query vectors from accessing skill-specific key space. We evaluate our methods on unlearning math-solving, Python-coding, and comprehension skills across seven different languages with GSM8K, MBPP, and MLQA datasets respectively. Both methods show strong skill unlearning performance with minimal hurt to the modelās overall capability. Experiments demonstrate the effectiveness of the two methods with > 80% relative performance drop on the target forgetting skill and < 10% drop on the modelās general knowledge and other skills. Specifically, Key Space Detection achieves nearly perfect skill unlearning with negligible overall capability drop. Our findings provide insights into how neuron activations cluster in key spaces and how these spatial properties can be leveraged for skill unlearning. These observations not only enhance our understanding of model behavior but also offer a promising direction for more targeted and interpretable unlearning techniques. We believe these insights will contribute to ongoing investigations into model interpretability, safety, and control. For example, certain knowledge, such as personal privacy data or adversarial attack inputs, may be localized in more fine-grained regions of the modelās key space. Understanding how these spatial properties can be systematically utilized needs further exploration. Acknowledgement The authors are partially supported by National Science Foundation under Grant No. 2107189, 2313105, 2430539, Hellman Fellowship, and Intel Rising Star Faculty Award. The authors would also like to thank anonymous reviewers for valuable feedback to improve the manuscript. Limitations Similar to other unlearning methods, our approach can only remove capabilities that can be captured by a dataset. However, in real-world applications, a specific dataset may not always be available for every capability we wish to remove. Unlearning knowledge without a controlled dataset or unlearning out-of-distribution data points presents an interesting yet challenging problem in the field of machine unlearning. Future work could explore techniques to address this challenge. Additionally, unlearning one skill while retaining a highly dependent skill requires a more fine-grained analysis. A promising direction for future research is to leverage spatial correlations among neurons to refine unlearning mechanisms and improve selectivity. Furthermore, we have not yet identified an efficient and automatic way to determine the optimal values for the adjusting ratio and the size hyperparameter α in both methods. In the case of Neuron Adjust, the reduction of unintended capabilities in the model is not guaranteed. For Key Space Detection, although we can ensure the modelās performance for out-of-hypercube queries, it may still lead to the degradation of certain unknown capabilities, as queries may cluster in a non-convex shape within the key space. References AI@Meta (2024) AI@Meta. 2024. Llama 3 model card. Austin et al. (2021) Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. Program synthesis with large language models. Preprint, arXiv:2108.07732. Bau et al. (2020) David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences. Bills et al. (2023) Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023. Language models can explain neurons in language models. https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html. Bricken et al. (2023) Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. 2023. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread. Https://transformer-circuits.pub/2023/monosemantic-features/index.html. Chen and Yang (2023) Jiaao Chen and Diyi Yang. 2023. Unlearn what you want to forget: Efficient unlearning for LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12041ā12052, Singapore. Association for Computational Linguistics. Cobbe et al. (2021) Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. Preprint, arXiv:2110.14168. Eldan and Russinovich (2023) Ronen Eldan and Mark Russinovich. 2023. Whoās harry potter? approximate unlearning in llms. Preprint, arXiv:2310.02238. Elhage et al. (2022) Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy models of superposition. Transformer Circuits Thread. https://transformer-circuits.pub/2022/toy_model/index.html. Geva et al. (2021) Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484ā5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Hendrycks et al. (2021) Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR). Hernandez et al. (2022) Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas. 2022. Natural language descriptions of deep visual features. In International Conference on Learning Representations. Huben et al. (2024) Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. 2024. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations. Imani et al. (2023) Shima Imani, Liang Du, and Harsh Shrivastava. 2023. MathPrompter: Mathematical reasoning using large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Industry Track), pages 37ā42, Toronto, Canada. Association for Computational Linguistics. Jang et al. (2023) Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. 2023. Knowledge unlearning for mitigating privacy risks in language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14389ā14408, Toronto, Canada. Association for Computational Linguistics. Lee et al. (2023) Justin Lee, Tuomas Oikarinen, Arjun Chatha, Keng-Chi Chang, Yilan Chen, and Tsui-Wei Weng. 2023. The importance of prompt tuning for automated neuron explanations. Preprint, arXiv:2310.06200. Lewis et al. (2019) Patrick Lewis, Barlas OÄuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2019. Mlqa: Evaluating cross-lingual extractive question answering. arXiv preprint arXiv:1910.07475. Li et al. (2022) Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, RĆ©mi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022. Competition-level code generation with alphacode. Science, 378(6624):1092ā1097. Liu et al. (2023) Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation. In Thirty-seventh Conference on Neural Information Processing Systems. Liu et al. (2024) Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al. 2024. Rethinking machine unlearning for large language models. arXiv preprint arXiv:2402.08787. Lu et al. (2022) Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi. 2022. Quark: Controllable text generation with reinforced unlearning. Advances in neural information processing systems, 35:27591ā27609. Meng et al. (2022) Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 36. ArXiv:2202.05262. Oikarinen and Weng (2023) Tuomas Oikarinen and Tsui-Wei Weng. 2023. Clip-dissect: Automatic description of neuron representations in deep vision networks. International Conference on Learning Representations. Oikarinen and Weng (2024) Tuomas Oikarinen and Tsui-Wei Weng. 2024. Linear explanations for individual neurons. Preprint, arXiv:2405.06855. OpenAI et al. (2024) OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Åukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Åukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David MĆ©ly, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen OāKeefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2024. Gpt-4 technical report. Preprint, arXiv:2303.08774. Pawelczyk et al. (2024) Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. 2024. In-context unlearning: Language models as few shot unlearners. Preprint, arXiv:2310.07579. Pochinkov and Schoots (2023) Nicky Pochinkov and Nandi Schoots. 2023. Dissecting large language models. In Socially Responsible Language Modelling Research. Shi et al. (2023) Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations. Sun et al. (2024) Chung-En Sun, Tuomas P. Oikarinen, Berk Ustun, and Tsui-Wei Weng. 2024. Concept bottleneck large language models. CoRR, abs/2412.07992. Team et al. (2024) Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane RiviĆØre, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, LĆ©onard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, AmĆ©lie HĆ©liou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, ClĆ©ment Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej MikuÅa, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, ClĆ©ment Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. 2024. Gemma: Open models based on gemini research and technology. Preprint, arXiv:2403.08295. Touvron et al. (2023) Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288. Wang et al. (2023) Lingzhi Wang, Tong Chen, Wei Yuan, Xingshan Zeng, Kam-Fai Wong, and Hongzhi Yin. 2023. KGA: A general machine unlearning framework based on knowledge gap alignment. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13264ā13276, Toronto, Canada. Association for Computational Linguistics. Wei et al. (2022) Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824ā24837. Wu et al. (2023) Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023. DEPN: Detecting and editing privacy neurons in pretrained language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2875ā2886, Singapore. Association for Computational Linguistics. Yao et al. (2023) Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023. Large language model unlearning. In Socially Responsible Language Modelling Research. Yu et al. (2023) Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. 2023. Unlearning bias in language models by partitioning gradients. In Findings of the Association for Computational Linguistics: ACL 2023, pages 6032ā6048, Toronto, Canada. Association for Computational Linguistics. Appendix A Appendix A.1 Time complexity analysis Neuron Adjust and KSD are both plug-in modules applicable to any autoregressive LLMs. For both methods, the implementation involves obtaining the mean and standard deviation of each neuron in the key space for every MLP layer by probing the subject model with the forgetting dataset. Since the forgetting dataset consists of N samples, this requires only N forward passes of the subject model. Given an autoregressive LLM with L MLP layers, each containing K neurons (where L and K are constants), we analyze the time complexity in detail: Neuron Adjust: For each inference step, we need to: ⢠Determine whether the inference time neuron activation is more likely to be drawn from the forgetting or retaining distribution. This step is Oā¢(Kā¢L)=Oā¢(1)1O(KL)=O(1)O ( K L ) = O ( 1 ). ⢠Change the neuron activation value if necessary. This step is also Oā¢(Kā¢L)=Oā¢(1)1O(KL)=O(1)O ( K L ) = O ( 1 ). Therefore, Neuron Adjust has an inference time cost of Oā¢(1)1O(1)O ( 1 ). KSD: For each inference step, we only need to determine whether the key vector in the last MLP layer is within the forgetting hyper-rectangle. This step is Oā¢(L)=Oā¢(1)1O(L)=O(1)O ( L ) = O ( 1 ). Therefore, KSD has an inference time cost of Oā¢(1)1O(1)O ( 1 ). In our experiments (real-world setting), it took less than 15 mins to implement our methods on llama-3-8b with a single V100 GPU, thus they are pretty light compared to other training-based methods. A.2 Sequential forgetting of multiple skills In this section we present the behavior of our methods when tasked to forget multiple skills sequentially. As shown in section 3 and 4, our methods are designed to be applied as plug-in modules. These modules can be used after each model update or fine-tuning process without affecting the modelās ability to learn new tasks or skills. Neuron Adjust is specifically optimized for forgetting a single skill, while Key Space Detection (KSD) can be extended to forget multiple skills. KSD works by identifying the hypercube that corresponds to each skill and determining whether the inference-time key vector falls within any of these hyper-rectangles. The computational complexity of this approach at inference time is Oā¢(M)O(M)O ( M ), where M represents the number of skills to be forgotten. To demonstrate the effectiveness of KSD in forgetting multiple skills, we conducted an additional experiment using Llama-3-70B, targeting the simultaneous forgetting of two tasks, MBPP and GSM8K. The results, as shown in Table 2, highlight the significant reduction in performance on the forgotten skills while maintaining general MMLU performance. Method GSM8K MBPP MBPP+ MMLU Original 47.5% 61.1% 51.1% 64.9% KSD 8.1% 0.5% 0.5% 64.8% Table 2: Results of Forgetting MBPP and GSM8K with Llama-3-70B The results demonstrate that KSD achieves a performance drop of over 80% on both GSM8K and MBPP tasks, effectively erasing the learned skills while leaving general knowledge tasks largely unaffected.