Paper deep dive
Towards a Unified Paradigm of Concept Editing in Large Language Models
Zhuowen Han, Xinwei Wu, Dan Shi, Renren Jin, Deyi Xiong
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/11/2026, 12:37:30 AM
Summary
The paper proposes a unified neuron-level paradigm for concept editing in Large Language Models (LLMs), categorizing existing methods into indirect (Neuron Editing, SFT) and direct (Sparse Autoencoder, Steering Vector) injection techniques. It introduces an efficient feature localization and interpretation method ('Reinforcer') for Sparse Autoencoders (SAE) and provides a hierarchical evaluation framework to compare these methods across reliability, generalization, and neuron-level consistency.
Entities (6)
Relation Signals (3)
Neuron Editing ā isa ā Indirect Injection
confidence 95% Ā· We categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV).
Sparse Autoencoder ā isa ā Direct Injection
confidence 95% Ā· We categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV).
Reinforcer ā usedfor ā Sparse Autoencoder
confidence 95% Ā· We propose a simple yet efficient method ā Reinforcer ā to interpret features in SAE.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Zhuowen Han, Xinwei Wu, Dan Shi, Renren Jin, Deyi Xiong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tags
Links
Full Text
63,803 characters extracted from source content.
Expand or collapse full text
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 18445ā18461 November 4-9, 2025 Ā©2025 Association for Computational Linguistics Towards a Unified Paradigm of Concept Editing in Large Language Models Zhuowen Han 1 , Xinwei Wu 1 , Dan Shi 1 , Renren Jin 1 , Deyi Xiong 1,2 * 1 TJUNLP Lab, College of Intelligence and Computing, Tianjin University, Tianjin, China 2 University International College, Macau University of Science and Technology, Macau, China zwhan, wuxw2021, shidan, rrjin, dyxiong@tju.edu.cn Abstract Concept editing aims to control specific con- cepts in large language models (LLMs) and is an emerging subfield of model editing. Despite the emergence of various editing methods in recent years, there remains a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods. To address this gap, we propose a unified paradigm for concept editing methods, in which all forms of conceptual injection are aligned at the neuron level. We study four rep- resentative concept editing methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoencoder (SAE), and Steering Vec- tor (SV). Then we categorize them into two classes based on their mode of conceptual infor- mation injection: indirect (NE, SFT) and direct (SAE, SV). We evaluate above methods along four dimensions: editing reliability, output gen- eralization, neuron level consistency, and math- ematical formalization. Experiments show that SAE achieves the best editing reliability. In output generalization, SAE captures features closer to human-understood concepts, while NE tends to locate text patterns rather than true semantics. Neuron-level analysis reveals that direct methods share high neuron overlap, as do indirect methods, indicating methodological commonality within each category. Our unified paradigm offers a clear framework and valu- able insights for advancing interpretability and controlled generation in LLMs. 1 Introduction Large language models (LLMs) have developed rapidly in recent years. However, as these mod- els become more powerful and complex, there is a growing need to better control their behaviors to achieve specific objectives, such as personalized generation, fairness and safety (Shen et al., 2023; Shi et al., 2024). A promising approach to achiev- ing this control is concept editing. As a branch of * Corresponding author model editing (Mitchell et al., 2021), it focuses on modifying the representations of specific con- cepts within LLMs to guide their outputs, rather than updating knowledge. Traditional model edit- ing methods focus on updating the knowledge of LLMs (Gupta et al., 2024), such as the classic ap- proaches ROME (Meng et al., 2022a) and MEMIT (Meng et al., 2022b). However, these are insuffi- cient for addressing the problem of concept editing, because concepts go beyond knowledge and are a more sophisticated synthesis of knowledge that cannot be exhaustively enumerated using discrete knowledge units. Plenty of concepts have already been learned and encoded within LLMs (Huben et al., 2023; Xu et al., 2023; Dong et al., 2025). Our goal is to identify and control them. With the advancement of interpretability re- search, various techniques for concept editing have emerged such as Neuron Editing (NE) (Dai et al., 2022), Steering Vector (SV) (Turner et al., 2023), Probing (Li et al., 2024), Sparse Autoencoder (SAE) (Huben et al., 2023) and so on. However, these approaches differ significantly in implemen- tation and performance, and there is still no unified theoretical framework to compare or understand them. Upon careful examination of these methods, we find that the common foundation of concept editing methods is the injection of conceptual informa- tion flow into a model. This flow corresponds to a modification of the residual stream and is rep- resented as a vector with the same dimension as the hidden state. Based on whether this injection is direct or indirect, we classify concept editing methods into two categories: modifying model pa- rameters (indirect injection), such as NE and SFT, and altering the residual stream (direct injection), such as SAE, SV and Probing. In addition, Neu- rons, as the fine-grained units of analysis in inter- pretability research, can serve as a unifying link for these methods. On the one hand, neurons are 18445 associated with parameters, and operations on neu- rons are equivalent to operations on the parameter matrices that affect them. On the other hand, neu- rons can also be viewed as a part of the residual stream (Vaswani, 2017). So we introduce a unified neuron-level paradigm for concept editing, which aligns all methods at the level of neuron. In the pro- cess of developing the paradigm, we have observed that existing SAE-based concept editing methods require large amounts of data and incur substantial computational overhead to interpret features. To address this limitation, we first propose an efficient method to interpret and locate features in SAE. To systematically compare these methods, we propose a hierarchical evaluation framework along four dimensions: editing reliability, output gen- eralization, neuron-level consistency, and mathe- matical formalization. We select āemotionā as a representative concept to evaluate each methodās ability to inject and preserve abstract semantic con- tent. Experimental results show that direct meth- ods, particularly SAE, achieve the most reliable and semantically faithful concept control. From the perspective of output generalization, we find that NE primarily captures surface patterns: it increases the probability of target concept words but does not enhance the probabilities of their synonyms or semantically related terms when reinforcing a concept. In contrast, SAE captures underlying con- cepts, raising the probabilities of all such words and thereby shifting the overall semantics of the output. Neuron-level analysis further reveals high activation overlap within each class of methods, suggesting a shared operational mechanism. Fi- nally, our mathematical analysis illustrates that the effectiveness of SAE stems from its precise and disentangled conceptual information injection. In summary, our contributions are threefold: ⢠We propose a unified neuron-level paradigm for concept editing, which provides a general framework to align diverse editing methods by grounding conceptual interventions at the neuron level. ⢠We conduct a systematic analysis of four rep- resentative methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoen- coder (SAE), and Steering Vector (SV), reveal- ing their structural similarities and differences in terms of conceptual injection mechanisms, editing reliability, semantic generalization be- havior, neuron overlap and mathematical for- malization. ā¢We further develop an efficient SAE-based concept editing method to locate interpretable features, which mitigates the high cost of feature interpretation in existing approaches while preserving strong editing reliability. 2 Background We first introduce the four concept editing methods separately. Neuron EditingWe adopt the method proposed by (Dai et al., 2022) for locating knowledge neu- rons. NE calculates an attribution score, denoted as Attr(n l i ), which measures the contribution of each neuron to the LLMās generation, wheren l i repre- sents the intermediate neuron at thei-th position in thel-th FFN layer of the LLM, and its activation value is denoted as w (l) i . We represent the target concept with an appropri- ate word or phrase and use it as the golden answer y ā for the promptx. We defineP x Ėw (l) i as the probability of y ā when w (l) i is set to Ėw (l) i : P y ā Ėw (l) i = p y ā | x,w (l) i = Ėw (l) i .(1) First, we take the prompt as input, record the activation of neuronn l i and denote it asw (l) i . Then, we scale the neuron activationw (l) i from 0 to its original valuew (l) i using the parameterα, and inte- grate the gradients along this path to compute the attribution score: Attr(n (l) i ) =w (l) i Z 1 α=0 āP y ā (α w (l) i ) āw (l) i dα,(2) which captures how changes inw (l) i affect the prob- ability of the golden answer. A high attribution score indicates a strong contribution to the con- cept. We then select neurons whose scores exceed a threshold t, filtering out those with low scores. In order to enhance or weaken a certain concept, we use a simple but effective approach: w (l) i = βw (l) i ,(3) whereβcontrols how strongly a concept is ex- pressed. 18446 Attn ... neurons FFN SFT Neuron Editing Residual stream CAT1: modifying model parameters CAT2: altering the residual stream Embedding Transformer Block ĆN Uembedding ķ up ķ down You feel good, you are happy and ready to start the day. You feel great and you are ready to take on the day. You feel frustrated because you are not able to get the right amount of sleep. Sparse Autoencoder Steering Vector Your morning starts with your favorite breakfast. You feel ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ NEļ¼Anger SAEļ¼Anger Related Words Synonyms ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ Figure 1: Illustration of our analytical framework. The upper part demonstrates the categorization of concept editing methods. We zoom into a specific layer of a transformer-based LLM, where the green area represents Category 1 (modifying model parameters to indirectly inject concept information flow, e.g., NE and SFT) and the yellow area denotes Category 2 (altering the residual stream to directly inject concept information flow, e.g., SAE and SV). The lower part shows the outputs of the LLM after injecting the concept of āangerā using NE and SAE, respectively. The specific related words and synonyms of āangerā are listed in Table 3. Sparse Autoencoders SAE is a neural network with a single hidden layer of sized s = Rd, where ddenotes the dimensionality of the LLMās internal activation vectors, andRis a hyperparameter that controls the ratio of the feature dictionary size to the model dimension (Huben et al., 2023). The residual stream at a specific position is represented as a vector hāR d . It operates as follows: c = ReLU(W 1 h + b 1 ),(4) h rec = W 2 c + b 2 ,(5) wherec āR d s , represents the sparse representa- tion ofh,h rec is the reconstructed hidden state, W 1 ,W 2 ,b 1 andb 2 are the trainable parameters. Among them, the matrixW 2 is our feature dictio- nary, consisting ofd s columns of dictionary fea- turesf (l) i . SAEs are trained to minimize the re- construction loss betweenhandh rec while also controlling the sparsity of c. Once trained, SAE provides an approximate de- composition of the modelās activations into a linear combination of feature directionsW 2 , with coef- ficients given by the sparse activationsc. This decomposition offers interpretability since the hid- den state can be explained by a small set of active features. Moreover, by directly modifyingc, for example adjusting a specific activationc k corre- sponding to featuref k , we can selectively control the influence of individual features on the output. Steering Vectors First, we need to get concept directions as steering vector. We use the difference between the activation of the positive and negative samples (Tigges et al., 2023): v (l) = 1 N X (h (l) pos ā h (l) neg ),(6) where,Nis the sample size,h (l) pos andh (l) neg denote the residual streams of positive and negative sam- ples respectively, and v (l) is the concept direction. Then, we addv (l) to the original residual stream during the forward pass: h (l) SV = h (l) + kv (l) ,(7) wherekcontrols the contribution of the vector to the generation. Supervised Fine-tuning We conduct full fine- tuning of LLMs while keeping all parameters frozen except forW l down (the second FFN weight matrix). This matrix, initiallyW āR d model Ćd ffn , becomes W sft āR d model Ćd ffn after SFT. 18447 Reinforcer ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ ļ¼ ļ¼Anger Anger, anger, anger... Reinforcer Interpretation1 rank = C Attrķ ķ ķ Interpretation target feature Interpretation2 Interpretation3 Interpretation4 Interpretation5 Interpretation6 Interpretation7 Interpretation8 Feature Localization Feature Interpretation Figure 2: Feature interpretation and localization in SAE. 3A Unified Paradigm of Concept Editing Methods Our analytical framework is illustrated in Figure 1. We categorize these methods into two groups ac- cording to the way conceptual information flow is injected. During the construction of this frame- work, we have observed that the current pipeline for concept editing with SAE is highly inefficient. Therefore, in Section 3.1, we first propose an ef- ficient approach for locating and interpreting fea- tures. In Section 3.2, we derive the neurons cor- responding for each method, with the goal of ex- amining the consistency of these two categories at the neuron level. Without loss of generality, in all derivations, we use āemotionā as the concept. 3.1 A Efficient Method to Locate Interpretable Features through SAE Existing methods for interpreting and locating fea- tures in SAE suffer from low efficiency and high cost. Therefore, we first propose a new method to efficiently locate features in SAE. Figure 2 illus- trates feature interpretation and localization. Feature Interpretation How to interpret fea- tures is a key question. Existing studies usually use an automated approach introduced in (Bills et al., 2023), which takes text samples where the target feature activates and asks a LLM to generate a human-readable interpretation (Huben et al., 2023). However, this method requires a large amount of corpus, and the cost of automated generation of explanations is high. In order to solve this prob- lem, we propose a simple yet efficient method ā Reinforcer. We have found that when the sparse activationc k of a particular featuref k is signifi- cantly large, the generation of LLMs tends to re- peat a single word or a group of words. For exam- ple, if the prompt is āYou see a beautiful flower, then you feelā, and we reinforce the feature āangerā (c k = 100) when generating, the output of the model would be āYou see a beautiful flower, then you feel anger anger anger anger...ā. The token āangerā is repeated over and over again and no more content is generated. We take this consistently re- peated token or a set of tokens as the interpretation of the feature. Reinforcer achieves comparable per- formance to the current method, while significantly reducing computational cost. The mathematical principles and performance of Reinforcer are pre- sented in Appendix A.1. Feature Localization Inspired by the idea of lo- cating neurons and attribution patching (Syed et al., 2024; Nanda, 2023), we shift the gradient compu- tation from neurons to the sparse space of SAE and propose a method to locate specific features precisely: Attr(f (l) i ) = c (l) i āP y ā (c (l) i ) āc (l) i ,(8) wheref (l) i represents the feature at thei-th posi- tion in thel-th layer of LLMs,c (l) i is the sparse activation corresponding tof (l) i ,y ā denotes the golden answer andP y ā (c (l) i )is the probability of the golden answer predicted by the model. This equation shares a similar form with Eq (2), and Attr(f (l) i )similarly measures the contribution of each feature to the generation. The difference is that gradient accumulation is not used here, as we found it has minimal effect on the results while substantially increasing computational overhead. During localization, we typically compute the gradient of the sparse activation of the last token in the prompt with respect to the probability of the golden tokens. If the number of golden tokens T g exceeds one, the relationship between the later golden tokens and the last prompt token becomes weak, leading to extremely small gradient values. This issue is not addressed in NE. To address this, we use cross-entropy loss (Loss) instead of proba- bilities of golden tokens to compute gradients: Attr(f (l) i ) =āc (l) i āLoss āc (l) i ,(9) because there is a negative correlation between Loss and probabilities: Loss =ā 1 T g T g X t=1 log P y t (c (l) i ),(10) 18448 wherey t represents thet-th token of the golden answer. So in fact, Eq (8) and Eq (9) are equivalent in variation trend, more details are shown in Ap- pendix A.2. Since Loss considers all golden tokens as a whole, it preserves the relationships among them and improves localization performance. In- tuitively, if a feature has a great influence on the answer, Attr(f (l) i ) will have a large value. We can rank the features triggered by the prompt based onAttr(f (l) i ). The target feature is usually among the top ten features. Then, our āReinforcerā is used to interpret them, thereby identifying the target feature. When editing features, we set the sparse value of the feature to a constantc (l) i = C. The larger the value, the stronger the concept. Our pipeline significantly optimizes the process of feature interpretation, localization and editing using SAEs. 3.2 Unifying Different Concept Editing Methods at the Neuron Level Neurons Located by SAE According to the structure of transformer (Vaswani, 2017), changes to the residual stream of LLMs can correspond to changes to the neurons: W l down n (l) + h (lā1) + att (l) = h (l) ,(11) where,W l down is the second parameter matrix of FFN module,n (l) represents the intermediate neu- ron vector in thel-th FFN module,h (lā1) is resid- ual stream from the previous layer,att (l) is the output of self-attention module, andh (l) is residual stream of the current layer. When editing,h (l) is reconstructed toh rec , then: W l down n (l) rec + h (lā1) + att (l) = h rec ,(12) where,n (l) rec represents the value of neurons when the residual stream ish rec . In this process, the rate of neurons chage is ān SAE = n (l) rec ā n (l) n (l) ,(13) which indicates the multiplicative factor applied to neurons to achieve the desired control over a concept. Clearly, the larger the rate of change, the more significant the role neurons play in controlling this concept. We select the topJneurons with the highest rates of change: n SAE = argsort ān SAE [āJ :],(14) whereJis the amount of neurons located by SAE, and n SAE denotes their indices. Neurons Located by Steering Vector Steering vector also constitutes a modification to residual stream. Therefore, the localization of neurons is similar to SAE. The detailed localization process is provided in Appendix A.3. Neurons Located by SFT We compute the rate of change for each column of the weight matrix: āw j = ā„W sft,j ā W j ā„ 2 ā„W j ā„ 2 ,(15) where,W j andW sft,j represent thej-th column of WandW sft , respectively,ā„Ā·ā„ 2 denotes the L2 norm. We then sort the columns based onāw j and select the topJcolumns with the highest rates of change; see Appendix A.3 for details. 4 Experiment Using emotion as the concept, we analyzed the per- formance of our interpretable feature localization for SAE, editing reliability, output generalization and neuron-level consistency, and also conducted analysis at the mathematical level. This forms a progressively deepening analytical paradigm for concept editing methods. 4.1 Setup ModelsWe conducted experiments on Gemma-2- 2B (RiviĆØre et al., 2024) and LLaMA-3-8B (Dubey et al., 2024), which contains diverse parameter scales. We choose their corresponding Sparse Autoencoders: gemma-scope-2b-pt-res 1 (Lieberum et al., 2024) and sae-llama-3-8b-32x. 2 DatasetWe used the emotion dataset introduced in (Zou et al., 2023). The dataset contains six cat- egories of emotionsāhappiness, sadness, anger, fear, surprise, and disgust, and comprises 1,200 sentences, with each emotion category consisting of 200 sentences. Based on this dataset, we utilized GPT-4 to generate 200 sentences for each emotion category (Long et al., 2024), thereby expanding the original dataset to a total of 2400 sentences. (More details are shown in Appendix A.4) Metrics We employed GPT-4 to assess the emo- tional intensity of the generated sentences (Zheng et al., 2023). Specifically, GPT-4 was used to score the emotional intensity of the generation before and 1 https://huggingface.co/google/ gemma-scope-2b-pt-res 2 https://huggingface.co/EleutherAI/ sae-llama-3-8b-32x 18449 EmotionShallowMiddleDeep L8BG2BL8BG2BL8BG2B anger--97546-5728015538 happiness--36205-5951212780 sadness-3940513314657919788554 surprise--6699-12107411874 disgust--120765-482554513 fear--55888-156068813 Table 1: The detected feature IDs in Llama-3-8B (L8B) and Gemma-2-2B (G2B) with our proposed method. The number represents the index of the feature in the sparse activation space of SAE. ā-ā indicates that no cor- responding features are found in that layer. L8B has 32 layers, with the 5th, 20th, and 30th layers representing shallow, middle, and deep layers. G2B has 26 layers, with the 5th, 15th, and 24th layers representing the same (layer indices start from 0). after concept editing, with scores ranging from 1 to 10. When performing the operation to enhance a specific emotion, if the edited model generates sentences with higher emotional scores, it indicates that the editing method is effective. Additionally, considering the positional bias in sentence scoring by GPT-4 (Wang et al., 2024), we averaged the scores of the original positions and the swapped positions. The prompt used for this evaluation is provided in Appendix A.5. In addition, we manu- ally assessed the results, which were found to be comparable to those of GPT-4. Further details are provided in Appendix A.6. 4.2 Performance of Our Interpretable Feature Localization for SAEs Using the feature localization method we proposed in Section 3.1, we detected the features correspond- ing to the six emotions as shown in Table 1. Our proposed method locates features with high accu- racy. We sampled 100 sentences and ranked the triggered features according toAttr(f (l) i )for each sentence. The target feature appears in the top 10 with a probability of 79.8%. Therefore, for any target concept, constructing a minimal amount of data in the prompt + golden answer format allows us to identify the target feature. From Table 1, it can be observed that not all lay- ers are capable of identifying the target features. For Llama-3-8B, target features are identifiable in middle and deep layers, but not in shallow lay- ers. In contrast, Gemma-2-2B can identify target features in deep layers, with minimal presence in middle and shallow layers. The reasons for this can be summarized as follows. NESAESVSFT Llama-3-8B layer-30 (Deep) anger36.9576.3546.8015.27 happiness51.5065.0084.0047.00 sadness51.2367.0059.6143.35 surprise49.5069.0055.5045.00 disgust36.9578.8259.1124.63 fear53.2084.7365.3541.87 Average46.5673.4861.7336.19 layer-20 (Middle) Average34.4897.5482.7625.62 layer-5 (Shallow) Average37.93-10038.91 Gemma-2-2B layer-24 (Deep) Average48.2895.5747.2926.60 layer-15 (Middle) Average45.32-98.5226.11 layer-5 (Shallow) Average45.81-10033.50 Table 2: The editing effect (%) of NE, SAE, SV and SFT. The numbers represent the percentage of data suc- cessfully controlled in the test set. The ā-ā indicate cases where SAE fails to locate features. First, the concept formation varies across layers. Shallow layers extract low-level lexical features, making it difficult to capture complex semantics. Middle layers begin integrating semantic informa- tion, where meaningful features start to emerge, while in deep layers, these features become mature. As shown in Appendix A.7, the features in shallow layers mainly consist of word fragments, auxiliary words, symbols and simple words. As the layers deepen, the features become complex. Second, the model size and capacity on gen- eration matter, as they significantly influence the modelās ability to capture complex features. Larger models, such as Llama-3-8B, exhibit superior per- formance in modeling intricate semantic features across layers. Conversely, smaller models, such as Gemma-2-2B, often struggle to form complete semantic representations across all layers, which explains why Gemma-2-2B fails to capture relevant concepts even in its middle layers. 4.3 Performance of Editing Reliability Experimental details are shown in Appendix A.8. The results are summarized in Table 2, and the numbers represent the percentage of data success- 18450 TypeWords Synonymsrage, irritation, fury, mad Antonymscalmness, joy, contentment, happy Relatedfrustration, argument, resentment, yelling Unrelatedbanana, laptop, galaxy, violin Table 3: The synonyms, antonyms, related words (words that often appear with āangerā), and unrelated words of āangerā. fully controlled in the test set. The intuitive results can be found in Appendix A.9. We observe the following: First, for Llama-3- 8B, in deep layers, SAE can control 73.48% of the emotions, SV can control 61.73%, while NE and SFT can only control less than 50%. We can see that SAE performs best, with SV following closely behind. Both are one level higher than NE and SFT. Similar trends are seen in other layers and in Gemma-2-2B. It shows that methods based on directly modifying the residual stream achieve better performance, especially when applied to mid- dle layers, which is resonated with previous studies (Hase et al., 2024). In contrast, we find NE and SFT are not sensitive to the layers being modified. Since our work focuses on training-free editing methods in low-resource scenarios, we adopt a relatively simple dataset, which may partly explain the poor performance of SFT. Therefore, in the analysis in Section 4.4, we focus on NE within CAT1. Second, SV exhibits significant effectiveness in middle layers but experiences a severe performance drop in deep layers. In comparison, SAE remains maintaining relatively strong performance in deep layers despite the increased semantic complexity, suggesting that the stability of SAE is excellent. Third, the ā-ā indicate cases where SAE fails to locate relevant emotional features, as concepts have not yet fully formed in shallow layers. 4.4 Performance of Output Generalization SAE identifies concepts, while NE identifies pat- terns. We applied various methods to enhance the emotion of āangerā and observed the average prob- abilities of the words in Table 3 over the next 20 generated time steps. Taking the prompt āYour morning starts with your favorite breakfast. You feelā as an example, the results are presented in Table 4. It can be seen that all methods lead to an in- crease in the probability of the concept words, with SAE showing the most significant effect. More- TypeOriginalNESAESV Anger9.12E-076.65E-06ā1.01E-03ā8.37E-06ā Synonyms2.57E-064.62E-06Ć5.66E-04ā2.26E-05ā Antonyms1.39E-034.24E-04ā5.67E-04ā1.62E-04ā Related6.12E-076.04E-07Ć3.31E-05ā1.22E-05ā Random2.45E-069.14E-079.14E-071.07E-06 Table 4: The average probabilities of these words among the 20 words generated by the model.āindicates that the editing method is valid, whileĆindicates it is in- valid. over, they all exhibit some degree of suppression for antonyms and have little impact on unrelated words. The focus is on synonyms and related words. NE shows no probability increase for them, indi- cating that it can only increase the probability of the concept words and does not induce semantic changes, which is clearly a pattern. Both SAE and SV show probability increases for synonyms and related words, but SAE exhibits a much larger increase, precisely controlling the target concept. We believe SAE can extract the target concept. The complete generations are shown in Figure 1. After enhancing the āangerā neurons, the semantics of the output remains unchanged, still conveying a āhappyā meaning, but the word āhappyā is re- moved. This once again indicates that the neurons identified by NE do not correspond to a concept, but rather capture a pattern based on logits. It only changes the occurrence probabilities of the concept words and their antonyms. The success of NE in other tasks is largely attributable to the strong cor- relation between the task and the pattern, such as knowledge (Dai et al., 2022) and privacy (Wu et al., 2023, 2024). Essentially, these tasks only require changing the output probability of target words. 4.5 Performance of Neuron-level Consistency Based on the derivations in Section 3.2, we ob- tainedn NE ,n SAE ,n SV andn SFT . To ensure a fair comparison, we controlled the number of neurons in each set by settingJequal to the number of neu- rons identified by NE, and then computed the over- lap rates among them. As shown in Table 5,n SAE andn SV exhibit a high degree of overlap (66.39%), whilen NE andn SFT show a low level of overlap (6.09%). Although the overlap ofn NE andn SFT is relatively low, it is still significantly higher than that of the other methods, showing some degree of overlap. In contrast, the neurons identified by the other pairs of methods do not overlap at all. This 18451 NE-SAE NE-SV NE-SFT SAE-SV SAE-SFT SV-SFT Llama-3-8B layer-30 (Deep) anger0.440.466.2666.632.492.65 happiness0.290.284.4866.492.922.88 sadness0.290.257.0666.262.432.31 surprise0.370.316.0967.462.962.85 disgust0.430.428.9565.872.212.05 fear0.570.533.6965.603.022.89 Average0.400.386.0966.392.672.61 layer-20 (Middle) Average1.030.969.2665.271.241.29 layer5 (Shallow) Average-0.084.93--1.84 Gemma-2-2B layer-24 (Deep) Average0.770.786.2267.764.574.45 layer-15 (Middle) Average-0.605.01--4.35 layer-5 (Shallow) Average-0.087.73--3.33 Table 5: The overlap (%) of neurons located by NE, SAE, SV and SFT. For example, 0.44 denotes that the overlap of neurons located by NE and SAE is 0.44%. result supports the validity of our classification phi- losophy. We analyze this variableānto explain how dif- ferent methods influence the generation process. According to Eq (11), Eq (13) and Eq (22): ān SAE = W + down (h (l) rec ā h (l) ) n (l) ,(16) ān SV = W + down v (l) n (l) ,(17) where,W + down is pseudo-inverse matrix ofW down , the only variable that doesnāt change with time step in the formula.ān SAE andān SV are both the basis for neuron localization and the multiplicative factors for neuron manipulation. We know that at each inference time,ān SAE andān SV change dynamically. However, sincevis an invariable vector, andh (l) rec ā h (l) changes at each time step, the neurons located by SAE and the manipulation to them are more flexible compared to SV. This explains why the editing effect of SAE is better. In contrast, the methods of CAT1 are static, and their performance tends to be inferior compared to the dynamic methods in CAT2. In addition, the identified neurons can serve as a mechanism to continuously track the evolution of a given concept throughout model optimization. Upon detecting a conceptual shift, the locate-and- edit procedure is automatically executed, thereby ensuring the long-term robustness of concept edits. 4.6 Mathematical Analysis In LLMs, the prediction of next token is performed by projecting the hidden statehlinearly onto the vocabulary space. Therefore, the most effective way to evaluate a concept editing method is to analyze the information flow it injects into the model, i.e.,āh. To further clarify this perspective, the mathematical analysis in Appendix A.10 illus- trates how different concept editing methods can be aligned and compared within a unified framework based on residual streams and matrix operations. This analysis highlights that the key advantage of SAE over other methods lies in projecting hidden states into a high-dimensional concept space, en- abling the injection of richer and more precise in- formation into the model. 4.7 General Abilities To investigate whether concept editing methods im- pairs the modelās general abilities, we conducted evaluations on five widely-used benchmarks. The specific benchmarks and evaluation results are pro- vided in Appendix A.11, which reveals that these methods rarely impacts the general ability of the LLMs, no matter they are methods from our cat- egory 1 or 2. Surprisingly, in some cases, it even leads to a slight performance improvement. 5 Related Work Concept editing is used to identify and modify con- cepts within LLMs in order to control their outputs. The commonly used methods can be categorized as follows. Neuron Editing Geva et al. (2021) show that feedforward layers in transformer-based language models operate as key-value memories, where each key correlates with textual patterns in the training examples, and each value induces a distribution over the output vocabulary. Based on this find- ing, a series of studies (Dai et al., 2022; Wu et al., 2023; Chen et al., 2024; Leng and Xiong, 2025; Shi et al., 2025) leverage knowledge neurons to edit specific factual knowledge. Lai et al. (2024) identify āstylesā neurons and enhance the stylistic diversity of the generated text. Wang et al. (2022) find skill neurons and explore their applications. Zhao et al. (2024) allow fine-tuning of language- specific neurons, enhancing multilingual abilities in a specific language. 18452 Sparse Autoencoders One of the roadblocks to a deep understanding of neural networksā internals is polysemanticity (Elhage et al., 2022), where neu- rons appear to activate in multiple and semantically distinct contexts (Scherlis et al., 2022). Huben et al. (2023) use sparse autoencoders to reconstruct the internal activations of language models. These au- toencoders learn sets of sparsely activating features that are more interpretable and monosemantic. Sub- sequently, multiple groups have open-sourced their own trained SAEs, such as EleutherAI 3 , Openai (Gao et al., 2024), Google (Lieberum et al., 2024) and so on. Paulo et al. (2024) build an open-source automated pipeline to generate and evaluate natu- ral language explanations for the features of SAEs using LLMs, but it requires substantial cost. Steering VectorThis concept involves the extrac- tion and optimization of latent space representation vectors from datasets. These vectors capture key attributes and can be directly manipulated to con- trol the attributes of the generated text (Liang et al., 2024). Subramani et al. (2022) extract latent steer- ing vectors from pretrained language models to control text generation, these vectors are then in- jected into the modelās hidden states. Tigges et al. (2023) show that emotion is represented linearly and capture directions of emotion. 6 Conclusion This paper presents a unified neuron-level paradigm for concept editing in large language models, offering a common framework to under- stand and compare diverse editing methods. By analyzing how conceptual information is injected, either indirectly through parameter modification or directly via residual stream manipulation, we classify representative methods such as NE, SFT, SAE, and SV into two coherent categories. We further conduct a systematic evaluation along four dimensions: editing reliability, output generaliza- tion, neuron-level consistency, and mathematical formalization. Our findings show that direct meth- ods, especially SAE, achieve superior performance in both reliability and semantic alignment. To ad- dress the computational limitations of existing SAE approaches, we additionally propose an efficient feature interpretation method that improves practi- cality without compromising effectiveness. Over- all, our work bridges fragmented concept editing 3 https://blog.eleuther.ai/autointerp/ strategies, deepens the understanding of their inter- nal mechanisms, and contributes practical advances for controllable and interpretable LLMs. Limitations First, due to computational constraints and few open-source SAEs available, this study could use only two models for experiments so far. While it covers both 2B and 8B sizes, examining more models under the proposed framework would cre- ate opportunities for more interesting findings and insights. Second, we have selected 6 emotions as example concepts for the empirical verification of our framework. In the future, we would like to ex- plore more concepts in a wide range of scenarios. Acknowledgements The present research was supported by the National Key Research and Development Program of China (Grant No. 2024YFE0203000). We would like to thank the anonymous reviewers for their insightful comments. References Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023. Language models can explain neurons in language models. URL https://openaipublic. blob. core. win- dows. net/neuron-explainer/paper/index. html.(Date accessed: 14.05. 2023), 2. Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2024. Journey to the Center of the Knowl- edge Neurons: Discoveries of Language-Independent Knowledge Neurons and Degenerate Knowledge Neurons. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 17817ā 17825. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question An- swering? Try ARC, the AI2 Reasoning Challenge. CoRR. Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge Neurons in Pretrained Transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493ā 8502. Weilong Dong, Xinwei Wu, Renren Jin, Shaoyang Xu, and Deyi Xiong. 2025. Contrans: Weak-to-strong alignment engineering via concept transplantation. In Proceedings of the 31st International Conference on Computational Linguistics, pages 4130ā4148. 18453 Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, AurĆ©lien Rodriguez, Austen Gregerson, Ava Spataru, Bap- tiste RoziĆØre, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Al- lonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Geor- gia Lewis Anderson, Graeme Nail, GrĆ©goire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Han- nah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel M. Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, and et al. 2024. The Llama 3 Herd of Models. CoRR. Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger B. Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022. Toy Models of Super- position. CoRR. Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024.SCALING AND EVALUATING SPARSE AUTOENCODERS. In The Thirteenth International Conference on Learn- ing Representations. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer Feed-Forward Layers Are Key-Value Memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 5484ā5495. Akshat Gupta, Dev Sajnani, and Gopala Anu- manchipalli. 2024. A Unified Framework for Model Editing. In Findings of the Association for Compu- tational Linguistics: EMNLP 2024, pages 15403ā 15418. Peter Hase, Mohit Bansal, Been Kim, and Asma Ghan- deharioun. 2024. Does Localization Inform Editing? Surprising Differences in Causality-Based Localiza- tion vs. Knowledge Editing in Language Models. Ad- vances in Neural Information Processing Systems, 36. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Un- derstanding. In International Conference on Learn- ing Representations. Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. 2023. Sparse Autoen- coders Find Highly Interpretable Features in Lan- guage Models. In The Twelfth International Confer- ence on Learning Representations. Wen Lai, Viktor Hangya, and Alexander Fraser. 2024. Style-Specific Neurons for Steering LLMs in Text Style Transfer. In Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Processing, pages 13427ā13443. Yongqi Leng and Deyi Xiong. 2025. Towards Under- standing Multi-Task Learning (Generalization) of LLMs via Detecting and Exploring Task-Specific Neurons. In Proceedings of the 31st International Conference on Computational Linguistics, pages 2969ā2987. Kenneth Li, Oam Patel, Fernanda ViĆ©gas, Hanspeter Pfister, and Martin Wattenberg. 2024. Inference- Time Intervention: Eliciting Truthful Answers from a Language Model. Advances in Neural Information Processing Systems, 36. Xun Liang, Hanyu Wang, Yezhaohui Wang, Shichao Song, Jiawei Yang, Simin Niu, Jie Hu, Dan Liu, Shunyu Yao, Feiyu Xiong, and Zhiyu Li. 2024. Con- trollable Text Generation for Large Language Mod- els: A Survey. CoRR. Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, JĆ”nos KramĆ”r, Anca Dragan, Rohin Shah, and Neel Nanda. 2024. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 278ā300. Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. 2024. On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey. In Findings of the As- sociation for Computational Linguistics ACL 2024, pages 11065ā11082. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and Editing Factual As- sociations in GPT. Advances in Neural Information Processing Systems, 35:17359ā17372. 18454 Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass- Editing Memory in a Transformer. In The Eleventh International Conference on Learning Representa- tions. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2021. Fast Model Editing at Scale. In International Conference on Learning Representations. NeelNanda.2023.AttributionPatch- ing:ActivationPatchingAtIndustrial Scale.https://w.neelnanda.io/mechanistic- interpretability/attribution-patching. GonƧalo Santos Paulo, Alex Troy Mallen, Caden Juang, and Nora Belrose. 2024. Automatically Interpreting Millions of Features in Large Language Models. In Forty-second International Conference on Machine Learning. Morgane RiviĆØre, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, LĆ©onard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre RamĆ©, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, An- ton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, Matt Hoffman, Shantanu Thakoor, Jean-Bastien Grill, Behnam Neyshabur, Olivier Bachem, Alanna Walton, Aliaksei Severyn, Alicia Parrish, Aliya Ahmad, Allen Hutchison, Alvin Abdagic, Amanda Carl, Amy Shen, Andy Brock, Andy Coenen, Anthony Laforge, Antonia Pater- son, Ben Bastian, Bilal Piot, Bo Wu, Brandon Royal, Charlie Chen, Chintu Kumar, Chris Perry, Chris Welty, Christopher A. Choquette-Choo, Danila Sinopalnikov, David Weinberger, Dimple Vijayku- mar, Dominika Rogozinska, Dustin Herbison, Elisa Bandy, Emma Wang, Eric Noland, Erica Moreira, Evan Senter, Evgenii Eltyshev, Francesco Visin, Gabriel Rasskin, Gary Wei, Glenn Cameron, Gus Martins, Hadi Hashemi, Hanna Klimczak-Plucinska, Harleen Batra, Harsh Dhand, Ivan Nardini, Jacinda Mein, Jack Zhou, James Svensson, Jeff Stanway, Jetha Chan, Jin Peng Zhou, Joana Carrasqueira, Joana Iljazi, Jocelyn Becker, Joe Fernandez, Joost van Amersfoort, Josh Gordon, Josh Lipschultz, Josh Newlan, Ju-yeong Ji, Kareem Mohamed, Kar- tikeya Badola, Kat Black, Katie Millican, Keelin McDonell, Kelvin Nguyen, Kiranbir Sodhia, Kish Greene, Lars Lowe Sjƶsund, Lauren Usui, Laurent Sifre, Lena Heuermann, Leticia Lago, and Lilly Mc- Nealus. 2024. Gemma 2: Improving Open Language Models at a Practical Size. CoRR. Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavat- ula, and Yejin Choi. 2021. Winogrande: An adver- sarial winograd schema challenge at scale. Commu- nications of the ACM, 64(9):99ā106. Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. 2022. Polysemanticity and Capacity in Neural Networks. CoRR. Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large Language Model Align- ment: A Survey. CoRR. Dan Shi, Renren Jin, Tianhao Shen, Weilong Dong, Xin- wei Wu, and Deyi Xiong. 2025. IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identi- fying and Reweighting Context-Aware Neurons. Ad- vances in Neural Information Processing Systems, 37:4997ā5024. Dan Shi, Tianhao Shen, Yufei Huang, Zhigen Li, Yongqi Leng, Renren Jin, Chuang Liu, Xinwei Wu, Zis- han Guo, Linhao Yu, Ling Shi, Bojian Jiang, and Deyi Xiong. 2024. Large Language Model Safety: A Holistic Survey. CoRR. Nishant Subramani, Nivedita Suresh, and Matthew E Pe- ters. 2022. Extracting Latent Steering Vectors from Pretrained Language Models. In Findings of the As- sociation for Computational Linguistics: ACL 2022, pages 566ā581. Aaquib Syed, Can Rager, and Arthur Conmy. 2024. At- tribution Patching Outperforms Automated Circuit Discovery. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Net- works for NLP, pages 407ā416. Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023. Linear Representations of Sentiment in Large Language Models. CoRR. Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023. Activation Addition: Steering Language Mod- els Without Optimization. CoRR. A Vaswani. 2017. Attention Is All You Need. Advances in Neural Information Processing Systems. Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models are not Fair Evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 9440ā9450. Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022. Finding Skill Neurons in Pre-trained Transformer-based Language Models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11132ā11152. Xinwei Wu, Weilong Dong, Shaoyang Xu, and Deyi Xiong. 2024. Mitigating Privacy Seesaw in Large Language Models: Augmented Privacy Neuron Edit- ing via Activation Patching. In Findings of the As- sociation for Computational Linguistics ACL 2024, pages 5319ā5332. 18455 Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023. DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2875ā2886. Shaoyang Xu, Junzhuo Li, and Deyi Xiong. 2023. Lan- guage Representation Projection: Can We Transfer Factual Knowledge across Languages in Multilingual Language Models? In Proceedings of the 2023 Con- ference on Empirical Methods in Natural Language Processing, pages 3692ā3702. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence? In Proceed- ings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791ā4800. Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. 2024. How do Large Language Models Handle Multilingualism?Ad- vances in Neural Information Processing Systems, 37:15296ā15319. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36:46595ā46623. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023. Representation Engineering: A Top-Down Approach to AI Transparency. CoRR. 18456 A Appendix A.1 Reinforcer Mathematical Principles Our work introduces SAE Reinforcer to interpret features. Its core idea is to manually amplify a specific featurec k (cor- responding to a concept like āangerā) by setting its activation to an extremely large valueC:c ā² k = C, wherecrepresents the sparse representation of SAE,c ā² is the modified sparse representation. Other features remain unchanged:c ā² j = c j ,āj Ģø= k. Then, the decoder processes this modified activa- tion: h rec = W 2 c ā² + b 2 = W 2,k C + X jĢø=k W 2,j c j + b 2 (18) SinceCis significantly larger than the original activation, the reconstructedh rec is dominated by W 2,k C: h rec ā W 2,k C(19) In a Transformer model, hidden states are ul- timately projected onto the vocabulary space for word prediction. As a result, the language model is biased toward generating words that are strongly associated with the featurec k . In extreme cases, the model will repeatedly generate a single word or phrase representing the concept, such as āangerā. Performance of Reinforcer The explanation obtained using the autointerpretability procedure (Huben et al., 2023) and the explanation of the con- cept obtained using our Reinforcer are consistent. Taking the feature āangerā (#15538) in layer 24 of Gemma-2-2b as an example, We randomly selected some texts from Wikipedia, segmented them into sentences, and identified the sentences corresponding to the 20 tokens with the highest activations for this feature in the high-dimensional sparse space of SAE (see Table 6). We then asked GPT-4 to summarize the common characteristics of these sentences and assign a name to Feature #15538. Answer by GPT-4: āA suitable name for Feature 15538 could be Emotional Intensity, capturing the prominence of strong emotions such as anger, rage, and the reactions they provoke in the characters.ā The results are consistent with those obtained using the reinforcer in our work. Additionally, similar ex- perimental results were observed for other models and concepts, which are not elaborated here. A.2 Feature Localization Details We extend Eq (8) to the case where the length of the golden answer is greater than one: Attr(f (l) i ) = c (l) i T g T g X t=1 āP y t (c (l) i ) āc (l) i .(20) When our method shown in Eq (9) expanded to the case where the length of the golden answer is greater than one: Attr(f (l) i ) =āc (l) i āLoss āc (l) i = c (l) i T g ā P T g t=1 log P (y t ) (c (l) i ) āc (l) i . (21) In theory, the trend of value changes for these two formulas is consistent. However, in Eq (20), starting fromt = 2, āP y t (c (l) i ) āc (l) i becomes very small, because the probability of subsequent tokens is less related toc (l) i , which causes Eq (20) to fail in accurately locating the feature. Therefore, Eq (20) is actually only applicable to the case where T g = 1. In contrast, Eq (21) does not encounter this issue and can accurately locate the feature. A.3 Neurons Selection Neurons Located by Steering Vector The rate of change in the values of neurons identified by SV is denoted as: ān SV = n (l) SV ā n (l) n (l) ,(22) where,n (l) SV represents the value of neurons when the residual stream ish SV . We then sort the neu- rons and select the topJneurons with the highest rates of change: n SV = argsort ān SV [āJ :],(23) whereJis the amount of neurons located by SV, andn SV is the set of indices of the important neu- rons identified by SV. Supervised Fine-tuningWe collect the change scores into a vectorāw = [āw 1 ,..., āw d ffn ] ⤠. We then rank the columns by their scores and se- lect the top J with the largest changes: n SFT = argsort āw[āJ :],(24) whereJindicates the amount of neurons located by SFT, andn SFT is the set of indices of the important neurons identified by SFT, since the columns ofW correspond one-to-one with the neurons. 18457 ActivationSentence 124.01When Admetus angered the goddess Artemis by forgetting to give her the due offerings, Apollo came to the rescue and calmed his sister. 116.42According to another version, or perhaps some years later, when Zeus struck down Apolloās son Asclepius with a lightning bolt for resurrecting the dead, Apollo in revenge killed the Cyclopes, who had fashioned the bolt for Zeus. 111.37 Rand was unimpressed by many of the NBI students and held them to strict standards, sometimes reacting coldly or angrily to those who disagreed with her. 110.27In one essay, political writer Jack Wheeler wrote that despite the incessant bombast and continuous venting of Randian rage, Randās ethics are a most immense achieve- ment, the study of which is vastly more fruitful than any other in contemporary thought. 105.42Cain then killed Abel out of jealousy. Table 6: Activation scores and corresponding sentences. A.4 Dataset Details During our experiments, we reconstructed the data using multiple prompt templates to ensure that the identified neurons or features are independent of any specific prompt. For example, āscenario, You feelā, āThe scenario is: scenario. The emotion in the above scenario isā and so on. The scenario is an item in dataset. A.5 Prompt Format Youāre a good assistant at evaluating the emotion of a text. Now you have two sentences, you are asked to assess the degree of āemotionā in both sentences. Each sentences receives an overall score on a scale of 0 to 10, where a higher score indicates higher the level of emotion. [The Start of Sentences 1] A1 [The End of Sentences 1] [The Start of Sentences 2] A2 [The End of Sentences 2] First, provide a comprehensive explanation of your evaluation, avoiding any potential bias and ensuring that the order in which the responses were presented does not affect your judgment. Then, give a overall score on a scale of 0 to 10 for the two answers, and this score is an integer. Output with the following format: Evaluation evidence: <evaluation explanation here> The overall score of Sentence 1: <score> The overall score of Sentence 2: <score> A.6 Human Evaluations We have re-sampled 10 sets of experimental results, with 40% of the samples (80 data points) in each re- sult manually verified. The number of cases where the GPT-4 evaluation is consistent with the man- ual review are: 78, 75, 70, 69, 73, 78, 75, 70, 72, 74. The average consistency rate is 91.75%, which is highly consistent, so the experimental results remain unchanged. A.7 Features in Different Layers Table 7 shows the features in shallow, middle and deep layers. A.8 Experiments Set We divided the dataset into ālocalizationā, āvalida- tionā and ātestā in a ratio of 1:1:2. This partition is based on the observation that getting neurons, features, and steering vectors requires relatively few data samples, while the primary focus lies in evaluating the effect and general abilities of the editing methods, necessitating a larger test set. For NE, we utilized the localization set to iden- tify neurons associated with a specific emotion. Directly calculating continuous integrals in Eq (2) is intractable. We use Riemann approxima- tionAttr(w (l) i ) = w (l) i m P m k=1 āP x ( k m w (l) i ) āw (l) i , where m = 20is the number of approximation steps. To filter out irrelevant neurons, we set a thresholdt equal to 0.1 times the maximum attribution score Attr(w (l) i )across all neurons. A neuron is retained if its attribution score exceedst. From the retained set, we selected 400ā500 neurons that exhibit a co- occurrence frequency above a thresholdqwithin 18458 LayerFeature Deep layergrief, pressure, solid, artificial, empathy Middle layertreat, smell, cave, play, anyway, ask, her Shallow layer://, gre, ant, pon, Dul, did, flo, klkl Table 7: Features located at different layers. the dataset, whereqis a hyperparameter. During the forward pass, we applied a scaling operation to these selected neurons, following Eq (3), withβas a hyperparameter. For SAE, we ranked the triggered features by the localization dataset and selected the top 10 features for validation using a reinforcer to determine the target feature. During the forward pass, we set the sparse activation of the target feature toC, where C is a hyperparameter. For SV, we extracted the emotion direction vec- tor using the localization dataset and applied a scal- ing factorkto this vector, and added it to the origi- nal residual stream during the forward pass, where k is a hyperparameter. These editing operations were applied at every time step of the generation process. Since SFT re- quires a larger dataset, we expanded each emotion category to 1,200 samples using GPT-4. The optimal parameters are selected based on the best performance on the validation set, where ābest performanceā not only reflects effective emotion control but also ensures the quality of the generated sentences. We incorporate mechanisms to detect repeated words and sentences, as well as perplexity (PPL) thresholds, to maintain a baseline quality of the outputs. Stricter quality checks are then applied during both LLM-based and human evaluations. Cost comparison Our experiments of all meth- ods can be run on a single A100 GPU with 80 GB of memory. The duration required for these experi- ments depends on the scale of model parameters, the size of the dataset, and the length of examples within the dataset. Overall, the time consumption is entirely acceptable. The comparison results are shown in Table 8. In summary, the cost required for the methods mentioned in our work is minimal and can be ignored. A.9 Intuitive Edit Effect Table 9 shows the edit results of different methods using āangerā as an example. A.10 Mathematical Analysis The information flow injected into the model by concept editing methods is essentiallyāh. Accord- ing to Eq (11), Eq (18) and Eq (7),āhmade by NE, SAE, and SV: āh NE = X jān W down,j ān j ,(25) āh SAE = W 2,k āc k ,(26) āh SV = kv (l) ,(27) wherenrepresents the neurons of NE,c k is the sparse activation off k . The key distinction is thatW down is not explicitly optimized for seman- tic alignment, whereasW 2 functions as an over- complete feature dictionary whose dimensional- ity far exceeds that of the model and is trained to encode prior knowledge about features and the underlying concept space, thereby supplying the residual stream with richer and more informative signals. Moreover, the activation vectorcin this high-dimensional space is sparse, effectively dis- entangling features. As a result, SAE introduces a more precise flow of information into the model. In contrast, the information flow of SV isv, which is derived from the residual itself and may contain redundant information. For SFT, since parameter updates affect the entire network, the change toāh is difficult to locate, meaning the influence is more diffuse and difficult to control. A.11 General Abilities These benchmarks are used in the widely- recognized LLM leaderboards, including Hel- laSwag (Zellers et al., 2019), ARC (Clark et al., 2018), MMLU (Hendrycks et al., 2021), Wino- grande (Sakaguchi et al., 2021) , and TruthfulQA (Lin et al., 2022). The general abilities of models after concept editing in deep layers is shown in Table 10 . 18459 MethodTrain/LocateInference (Additional computation and la- tency) NEAn offline computation identifies relevant neurons. In our experiment, locating 100 samples takes under 30 minutes for the 8B model. Reweight predetermined neuron activa- tions (computation and latency can be ig- nored). SV Computing the steering vector for 200 data points takes 25 seconds. Add the locally stored steering vector to the residual stream (computation and la- tency can be ignored). SAEThe SAE used in our work is open-source, so there is no need to consider its training cost. The identification and interpretation of features can also be completed quickly. SAE, with just an encoder and a decoder, occupies minimal storage space. It adds three negligible steps during inference: en- coding, sparse-space editing, and decod- ing. (computation and latency can be ig- nored) SFTThe training can be completed within 30 minutes. None Table 8: Cost comparison of concept editing methods. Anger OriYou discover an old family photo album youāve never seen before. You feel a sense of excitement and curiosity as you flip through the pages, wondering who these people are and what their stories might be. As you continue to explore the album, you notice that some of the photos NEYou discover an old family photo album youāve never seen before. You feel a sense of excitement and curiosity as you flip through the pages, wondering who these people are and what their stories might be. As you look at each picture, you notice that some of them have been SAE You discover an old family photo album youāve never seen before. You feel angry and frustrated that your parents didnāt share this with you when you were younger. Youāre upset because you feel like you missed out on something important, but you donāt know what it is. SVYou discover an old family photo album youāve never seen before. You feel a strong emotion, and then you begin to remember... the words that were directed at you, the tone of those words, the tone of your response, the tone of the responses against you. SFTYou discover an old family photo album youāve never seen before. You feel a strong connection to the people in the photos, even though you donāt recognize anyone. Table 9: The outputs of L8B before and after editing using NE, SAE, SV, and SFT (taking āangerā as an example). āOriā denotes the original output of the model. The italicized text represents the prompt. 18460 ARCHellaSwagMMLUTruthfulQAWinograndeAverage Gemma-2-2B Original0.80220.73030.49610.36230.68750.6157 NE0.79880.72660.49450.36130.68980.6142 SAE0.75720.72790.49370.38520.67010.6068 SV0.76680.70990.49030.36970.68350.6040 SFT0.80890.73350.49600.34690.67250.6116 Llama-3-8B Original0.77690.60170.62180.43900.72850.6336 NE0.75510.78510.61840.43990.72770.6652 SAE0.73400.75070.62160.43750.67010.6428 SV0.70450.74900.60980.48280.72850.6549 SFT0.78110.78980.62060.42900.70880.6659 Table 10: Results of general abilities of LLMs after concept editing in deep layers on widely-used benchmarks. āOriginalā refers to the modelās output before any editing. 18461