Paper deep dive
Mass-Editing Memory in a Transformer
Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, David Bau
Models: GPT-J (6B), GPT-NeoX (20B)
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 8:28:59 PM
Summary
MEMIT is a scalable method for mass-editing factual associations in large language models (LLMs) like GPT-J and GPT-NeoX. By identifying critical MLP layers as causal mediators of knowledge recall, MEMIT uses an explicit parameter update algorithm to insert thousands of memories simultaneously while maintaining generalization, specificity, and fluency, significantly outperforming prior knowledge-editing techniques.
Entities (6)
Relation Signals (4)
MEMIT â updates â GPT-J
confidence 98% ¡ MEMIT can scale and successfully store thousands of memories in bulk... Experiments on GPT-J (6B parameters)
MEMIT â updates â GPT-NeoX
confidence 98% ¡ Experiments on GPT-J (6B parameters) and GPT-NeoX (20B) demonstrate that MEMIT can scale
MEMIT â evaluatedon â zsRE
confidence 95% ¡ We first test MEMIT on zsRE (Levy et al., 2017)
ROME â inspired â MEMIT
confidence 90% ¡ Inspired by the ROME direct editing method (Meng et al., 2022), MEMIT targets the weights
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent work has shown exciting promise in updating large language models with new memories, so as to replace obsolete information or add specialized knowledge. However, this line of work is predominantly limited to updating single associations. We develop MEMIT, a method for directly updating a language model with many memories, demonstrating experimentally that it can scale up to thousands of associations for GPT-J (6B) and GPT-NeoX (20B), exceeding prior work by orders of magnitude. Our code and data are at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2210.07229
- Canonical: https://arxiv.org/abs/2210.07229
- Code: https://memit.baulab.info
Trouble viewing inline? Open PDF directly â
Full Text
64,017 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at ICLR 2023 MASS-EDITINGMEMORY IN ATRANSFORMER Kevin Meng 1,2 Arnab Sen Sharma 2 Alex Andonian 1 Yonatan Belinkov â 3 David Bau 2 1 MIT CSAIL 2 Northeastern University 3 Technion â IIT ABSTRACT Recent work has shown exciting promise in updating large language models with new memories, so as to replace obsolete information or add specialized knowledge. However, this line of work is predominantly limited to updating single associations. We develop MEMIT, a method for directly updating a language model with many memories, demonstrating experimentally that it can scale up tothousands of associationsfor GPT-J (6B) and GPT-NeoX (20B), exceeding prior work by orders of magnitude. Our code and data are at memit.baulab.info. 1INTRODUCTION How many memories can we add to a deep network by directly editing its weights? Although large autoregressive language models (Radford et al., 2019; Brown et al., 2020; Wang & Komatsuzaki, 2021; Black et al., 2022) are capable of recalling an impressive array of common facts such as âTim Cook is the CEO of Appleâ or âPolaris is in the constellation Ursa Minorâ (Petroni et al., 2020; Brown et al., 2020), even very large models are known to lack more specialized knowledge, and they may recall obsolete information if not updated periodically (Lazaridou et al., 2021; Agarwal & Nenkova, 2022; Liska et al., 2022). The ability to maintain fresh and customizable information is desirable in many application domains, such as question answering, knowledge search, and content generation. For example, we might want to keep search models updated with breaking news and recently-generated user feedback. In other situations, authors or companies may wish to customize models with specific knowledge about their creative work or products. Because re-training a large model can be prohibitive (Patterson et al., 2021) we seek methods that can update knowledge directly. To that end, severalknowledge-editingmethods have been proposed to insert new memories directly into specific model parameters. The approaches include constrained fine-tuning (Zhu et al., 2020), hypernetwork knowledge editing (De Cao et al., 2021; Hase et al., 2021; Mitchell et al., 2021; 2022), and rank-one model editing (Meng et al., 2022). However, this body of work is typically limited to updating at most a few dozen facts; a recent study evaluates on a maximum of 75 (Mitchell et al., 2022) whereas others primarily focus on single-edit cases. In practical settings, we may wish to MEMIT plays sport plays sport plays sport Tony Meola Olga FĂŚrseth Michael Jordan Baseball Soccer Basketball located in located in Eiffel Tower Space Needle Paris Seattle sor (a) Unedited GPT plays sport Tony Meola Olga FĂŚrseth Michael Jordan Baseball Soccer Basketball located in Eiffel Tower Space Needle Paris Seattle sor (b) Modified GPT (c) Scaling MEMIT to 10,000 Edits ... ... Figure 1:MEMIT is capable of updating thousands of memories at once. (a) Language models can be viewed as knowledge bases containing memorized tuples(s,r,o), each connecting some subjectsto an objectovia a relationr, e.g., (s=Michael Jordan,r=plays sport,o=basketball). (b) MEMIT modifies transformer weights to edit memories, e.g., âMichael Jordan now plays the sport baseball,â while (c) maintaining generalization, specificity, and fluency at scales beyond other methods. As Section 5.2.2 details, editing score is the harmonic mean of efficacy, generalization, and specificity metrics. â Supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion. Correspondence to mengk@mit.edu, davidbau@northeastern.edu. 1 arXiv:2210.07229v2 [cs.CL] 1 Aug 2023 Published as a conference paper at ICLR 2023 update a model with hundreds or thousands of facts simultaneously, but a naive sequential application of current state-of-the-art knowledge-editing methods fails to scale up (Section 5.2). We propose MEMIT, a scalable multi-layer update algorithm that uses explicitly calculated parameter updates to insert new memories. Inspired by the ROME direct editing method (Meng et al., 2022), MEMIT targets the weights of transformer modules that we determine to be causal mediators of factual knowledge recall. Experiments on GPT-J (6B parameters; Wang & Komatsuzaki 2021) and GPT-NeoX (20B; Black et al. 2022) demonstrate thatMEMIT can scale and successfully store thousands of memories in bulk. We analyze model behavior when inserting true facts, counterfactuals, 27 specific relations, and different mixed sets of memories. In each setting, we measure robustness in terms of generalization, specificity, and fluency while comparing the scaling of MEMIT to rank-one, hypernetwork, and fine-tuning baselines. 2RELATEDWORK Scalable knowledge bases.The representation of world knowledge is a core problem in artificial intelligence (Richens, 1956; Minsky, 1974), classically tackled by constructingknowledge basesof real-world concepts. Pioneering hand-curated efforts (Lenat, 1995; Miller, 1995) have been followed by web-powered knowledge graphs (Auer et al., 2007; Bollacker et al., 2007; Suchanek et al., 2007; Havasi et al., 2007; Carlson et al., 2010; Dong et al., 2014; Vrande Ë ci Ě c & KrĂśtzsch, 2014; Bosselut et al., 2019) that extract knowledge from large-scale sources. Structured knowledge bases can be precisely queried, measured, and updated (Davis et al., 1993), but they are limited by sparse coverage of uncatalogued knowledge, such as commonsense facts (Weikum, 2021). Language models as knowledge bases.Since LLMs can answer natural-language queries about real-world facts, it has been proposed that they could be used directly as knowledge bases (Petroni et al., 2019; Roberts et al., 2020; Jiang et al., 2020; Shin et al., 2020). However, LLM knowledge is only implicit; responses are sensitive to specific phrasings of the prompt (Elazar et al., 2021; Petroni et al., 2020), and it remains difficult to catalog, add, or update knowledge (AlKhamissi et al., 2022). Nevertheless, LLMs are promising because they scale well and are unconstrained by a fixed schema (Safavi & Koutra, 2021). In this paper, we take on the update problem, asking how the implicit knowledge encoded within model parameters can be mass-edited. Hypernetwork knowledge editors.Several meta-learning methods have been proposed to edit knowledge in a model. Sinitsin et al. (2019) proposes a training objective to produce models amenable to editing by gradient descent. De Cao et al. (2021) proposes a Knowledge Editor (KE) hypernetwork that edits a standard model by predicting updates conditioned on new factual statements. In a study of KE, Hase et al. (2021) find that it fails to scale beyond a few edits, and they scale an improved objective to 10 beliefs. MEND (Mitchell et al., 2021) also adopts meta-learning, inferring weight updates from the gradient of the inserted fact. To scale their method, Mitchell et al. (2022) proposes SERAC, a system that routes rewritten facts through a different set of parameters while keeping the original weights unmodified; they demonstrate scaling up to 75 edits. Rather than meta-learning, our method employs direct parameter updates based on an explicitly computed mapping. Direct model editing.Our work most directly builds upon efforts to localize and understand the internal mechanisms within LLMs (Elhage et al., 2021; Dar et al., 2022). Based on observations from Geva et al. (2021; 2022) that transformer MLP layers serve as keyâvalue memories, we narrow our focus to them. We then employ causal mediation analysis (Pearl, 2001; Vig et al., 2020; Meng et al., 2022), which implicates a specific range of layers in recalling factual knowledge. Previously, Dai et al. (2022) and Yao et al. (2022) have proposed editing methods that alter sparse sets of neurons, but we adopt the classical view of a linear layer as an associative memory (Anderson, 1972; Kohonen, 1972). Our method is closely related to Meng et al. (2022), which also updates GPT as an explicit associative memory. Unlike the single-edit approach taken in that work, we modify a sequence of layers and develop a way for thousands of modifications to be performed simultaneously. 3PRELIMINARIES: LANGUAGEMODELING ANDMEMORYEDITING The goal of MEMIT is to modify factual associations stored in the parameters of an autoregressive LLM. Such models generate text by iteratively sampling from a conditional token distribution 2 Published as a conference paper at ICLR 2023 P x [t] |x [1] ,...,x [E] parameterized by aD-layer transformer decoder,G(Vaswani et al., 2017): P x [t] |x [1] ,...,x [E] âG([x [1] ,...,x [E] ]) = softmax W y h D [E] ,(1) whereh D [E] is the transformerâs hidden state representation at the final layerDand ending tokenE. This state is computed using the following recursive relation: h l [t] (x) =h lâ1 [t] (x) +a l [t] (x) +m l [t] (x)(2) wherea l = attn l h lâ1 [1] ,h lâ1 [2] ,...,h lâ1 [t] (3) m l [t] =W l out Ď W l in Îł h lâ1 [t] ,(4) h 0 [t] (x) is the embedding of tokenx [t] , andÎłis layernorm. Note that we have written attention and MLPs in parallel as done in Black et al. (2021) and Wang & Komatsuzaki (2021). Large language models have been observed to contain many memorized facts (Petroni et al., 2020; Brown et al., 2020; Jiang et al., 2020; Chowdhery et al., 2022). In this paper, we study facts of the form (subjects, relationr, objecto), e.g., (s=Michael Jordan,r=plays sport,o=basketball). A generatorGcan recall a memory for(s i ,r i ,â)if we form a natural language promptp i =p(s i ,r i ) such as âMichael Jordan plays the sport ofâ and predict the next token(s) representingo i . Our goal is to edit many memories at once. We formally define a list of edit requests as: E=(s i ,r i ,o i )|is.t.âi,j.(s i =s j )â§(r i =r j )â§(o i ̸=o j ).(5) The logical constraint ensures that there are no conflicting requests. For example, we can edit Michael Jordan to playo i =âbaseballâ, but then we exclude associating him with professional soccer. What does it mean to edit a memory well? At a superficial level, a memory can be considered edited after the model assigns a higher probability to the statement âMichael Jordan plays the sport of baseballâ than to the original prediction (basketball); we say that such an update iseffective. Yet it is important to also view the question in terms ofgeneralization,specificity, andfluency: â˘To test forgeneralization, we can rephrase the question: âWhat is Michael Jordanâs sport? What sport does he play professionally?â If the modification ofGis superficial and overfitted to the specific memorized prompt, such predictions will fail to recall the edited memory, âbaseball.â ⢠Conversely, to test forspecificity, we can ask about similar subjects for which memories should not change: âWhat sport does Kobe Bryant play? What does Magic Johnson play?â These tests will fail if the updatedGindiscriminately regurgitates âbaseballâ for subjects that were not edited. â˘When making changes to a model, we must also monitorfluency. If the updated model generates disfluent text such as âbaseball baseball baseball baseball,â we should count that as a failure. Achieving these goals is challenging, even for a few edits (Hase et al., 2021; Mitchell et al., 2022; Meng et al., 2022). We investigate whether they can be attained at the scale of thousands of edits. 4METHOD MEMIT inserts memories by updating transformer mechanisms that have recently been elucidated using causal mediation analysis (Meng et al., 2022). In GPT-2 XL, we found that there is a sequence of critical MLP layersRthat mediate factual association recall at the last subject tokenS(Figure 2). MEMIT operates by (i) calculating the vector associations we want the critical layers to remember, then (i) storing a portion of the desired memories in each layerlâR. Throughout this paper, our focus will be on states representing the last subject tokenSof promptp i , so we shall abbreviateh l i =h l [S] (p i ). Similarly,m l i anda l i denotem l [S] (p i )anda l [S] (p i ). 4.1IDENTIFYING THE CRITICAL PATH OFMLPLAYERS Figure 3 shows the results of applying causal tracing to the larger GPT-J (6B) model; for implementa- tion details, see Appendix A. We measure the average indirect causal effect of eachh l i on a sample of memory promptsp i , with either the Attention or MLP modules for tokenSdisabled. The results 3 Published as a conference paper at ICLR 2023 Michael Jordan now plays í ! " í !" # í $%& # attn ! í "#$ ! storesí % ! â í % ! pairs minimizing: ! !"# $ í %&' ( í ! ( â í ! ( ) key for subject memorized value attn module vector state direct path mlp module mlpcritical path non-mediating components information moved by attention (a) (b) (c) í ! " í ! " â ! " (d) range of critical MLP layers â last subject token í Figure 2:MEMIT modifies transformer parameters on the critical path of MLP-mediated factual recall. We edit stored associations based on observed patterns of causal mediation: (a) first, the early-layer attention modules gather subject names into vector representations at the last subject tokenS. (b) Then MLPs at layers lâ Rread these encodings and add memories to the residual stream. (c) Those hidden states are read by attention to produce the output. (d) MEMIT edits memories by storing vector associations in the critical MLPs. confirm that GPT-J has a concentration of mediating statesh l i ; moreover, they highlight a mediating causal role for a range of MLP modules, which can be seen as a large gap between the effect of single states (purple bars in Figure 3) and the effects with MLP severed (green bars); this gap diminishes after layer 8. Unlike Meng et al. (2022) who use this test to identify a single edit layer, we select the whole range of critical MLP layerslâR. For GPT-J, we haveR=3,4,5,6,7,8. 0510152025 Layer at which hidden state is restored 0.0% 5.0% 10.0% Average Indirect Effect Causal effect of hidden states Attn or MLP modules severed Effect of single state Effect w/ Attn severed Effect w/ MLP severed Figure 3:A critical mediating role for mid-layer MLPs. Given that arangeof MLPs play a joint mediating role in recalling facts, we ask: what is the role ofoneMLP in stor- ing a memory? Each token state in a transformer is part of the residual stream that all attention and MLP modules read from and write to (Elhage et al., 2021). Unrolling Eqn. 2 forh L i =h L [S] (p i ): h L i =h 0 i + L X l=1 a l i + L X l=1 m l i .(6) Eqn. 6 highlights that each individual MLP contributes byaddingto the memory ath L i (Figure 2b), which is later read by last-token attention modules (Figure 2c). Therefore, when writing new memories intoG, we can spread the desired changes across all the critical layersm l i forlâR. 4.2BATCH UPDATE FOR A SINGLE LINEAR ASSOCIATIVE MEMORY In each individual layerl, we wish to store a large batch ofuâŤ1memories. This section derives an optimal single-layer update that minimizes the squared error of memorized associations, assuming that the layer contains previously-stored memories that should be preserved. We denoteW 0 âW l out (Eqn. 4, Figure 2) and analyze it as a linear associative memory (Kohonen, 1972; Anderson, 1972) that associates a set of input keysk i âk l i (encoding subjects) to corresponding memory values m i âm l i (encoding memorized properties) with minimal squared error: W 0 âargmin Ë W n X i=1 Ë Wk i âm i 2 .(7) If we stack keys and memories as matricesK 0 = [k 1 |k 2 |¡|k n ]andM 0 = [m 1 |m 2 |¡|m n ], then Eqn. 7 can be optimized by solving the normal equation (Strang, 1993, Chapter 4): W 0 K 0 K T 0 =M 0 K T 0 .(8) Suppose that pre-training sets a transformer MLPâs weights to the optimal solutionW 0 as defined in Eqn. 8. Our goal is to updateW 0 with some small changeâthat produces a new matrixW 1 with 4 Published as a conference paper at ICLR 2023 í§ ! â ! " í ! "#$ í ! "#$ í !% "#$ í &'( "#$ attn !"# â ! "#) í ! "#* í ! "#* í !% "#* í &'( "#* attn !"$ â ! "#$ í ! " í ! " í !% " í &'( " attn ! â ! "#* (i ) For each memory í, find í§ ! by optimizing Eqn. 16 (i -a)Add â "#$ s.t.âí: í ! "#$ += + ! #, ! " ) Re-collect layer íż â 1activations (i -b)Add â "#* s.t.âí: í ! "#* += + ! #, ! " $ (i -c)Add â " s.t.âí: í ! " += + ! #, ! " * Re-collect layer íżactivations (i) For each layer í, apply updates using Eqn. 14 to move all â ! " towards í§ ! All states examined at í =Last subject token forí ! Figure 4:The MEMIT update. We first (i) replaceh l i with the vectorz i and optimize Eqn. 16 so that it conveys the new memory. Then, after allz i are calculated we (i) iteratively insert a fraction of the residuals for allz i over the range of critical MLP modules, executing each layerâs update by applying Eqn. 14. Because changing one layer will affect activations of downstream modules, we recollect activations after each iteration. a set of additional associations. Unlike Meng et al. (2022), we cannot solve our problem with a constraint that adds only a single new association, so we define an expanded objective: W 1 âargmin Ë W n X i=1 Ë Wk i âm i 2 + n+u X i=n+1 Ë Wk i âm i 2 ! .(9) We can solve Eqn. 9 by again applying the normal equation, now written in block form: W 1 [ K 0 K 1 ] [ K 0 K 1 ] T = [ M 0 M 1 ] [ K 0 K 1 ] T (10) which expands to:(W 0 + â)(K 0 K T 0 +K 1 K T 1 ) =M 0 K T 0 +M 1 K T 1 (11) W 0 K 0 K T 0 +W 0 K 1 K T 1 + âK 0 K T 0 + âK 1 K T 1 =M 0 K T 0 +M 1 K T 1 (12) subtracting Eqn. 8 from Eqn. 12 :â(K 0 K T 0 +K 1 K T 1 ) =M 1 K T 1 âW 0 K 1 K T 1 .(13) A succinct solution can be written by defining two additional quantities:C 0 âK 0 K T 0 , a constant proportional to the uncentered covariance of the pre-existing keys, andRâM 1 âW 0 K 1 , the residual error of the new associations when evaluated on old weightsW 0 . Then Eqn. 13 can be simplified as: â =RK T 1 (C 0 +K 1 K T 1 ) â1 .(14) Since pretraining is opaque, we do not have access toK 0 orM 0 . Fortunately, computing Eqn. 14 only requires an aggregate statisticC 0 over the previously stored keys. We assume that the set of previously memorized keys can be modeled as a random sample of inputs, so that we can compute C 0 =Ν¡E k k T (15) by estimatingE k k T , an uncentered covariance statistic collected using an empirical sample of vector inputs to the layer. We must also selectÎť, a hyperparameter that balances the weighting of new v.s. old associations; a typical value isÎť= 1.5Ă10 4 . 4.3UPDATING MULTIPLE LAYERS We now define the overall update algorithm (Figure 4). Inspired by the observation that robustness is improved when parameter change magnitudes are minimized (Zhu et al., 2020), we spread updates evenly over the range of mediating layersR. We define a target layerLâmax(R)at the end of the mediating layers, at which the new memories should be fully represented. Then, for each edit (s i ,r i ,o i )âE, we (i) compute a hidden vectorz i to replaceh L i such that addingδ i âz i âh L i to the hidden state at layerLand tokenTwill completely convey the new memory. Finally, one layer at a time, we (i) modify the MLP at layerl, so that it contributes an approximately-equal portion of the changeδ i for each memoryi. (i) Computingz i . For theith memory, we first compute a vectorz i that would encode the association (s i ,r i ,o i )if it were to replaceh L i at layerLat tokenS. We findz i =h L i +δ i by optimizing the residual vectorδ i using gradient descent: z i =h L i + argmin δ i 1 P P X j=1 âlogP G(h L i +=δ i ) [o i |x j âp(s i ,r i )].(16) 5 Published as a conference paper at ICLR 2023 In words, we optimizeδ i to maximize the modelâs prediction of the desired objecto i , given a set of factual promptsx j âp(s i ,r i )that concatenate random prefixesx j to a templated prompt to aid generalization across contexts.G(h L i +=δ i )indicates that we modify the transformer execution by substituting the modified hidden statez i forh L i ; this is called âhookingâ in popular ML libraries. (i) Spreadingz i âh L i over layers. We seek delta matricesâ l such that: setting Ë W l out :=W l out + â l for alllâRoptimizesmin â l X i z i â Ë h L i 2 ,(17) where Ë h L i =h 0 i + L X l=1 a l i + L X l=1 Ë W l out Ď W l in Îł h lâ1 t .(18) Because edits to any layer will influence all following layersâ activations, we calculateâ l iteratively in ascending layer order (Figure 4i-a,b,c). To compute each individualâ l , we need the corresponding keysK l = [k l 1 |¡|k l n ]and memoriesM l = [m l 1 |¡|m l n ]to insert using Eqn. 14. Each keyk l i is computed as the input toW l out at each layerl(Figure 2d): k l i = 1 P P X j=1 k(x j +s i ),wherek(x) =Ď W l in Îł h lâ1 i (x) .(19) m l i is then computed as the sum of its current value and a fraction of the remaining top-level residual: m l i =W out k l i +r l i wherer l i is the residual given by z i âh L i Lâl+ 1 ,(20) where the denominator ofr i spreads the residual out evenly. Algorithm 1 summarizes MEMIT, and additional implementation details are offered in Appendix B. Algorithm 1:The MEMIT Algorithm Data:Requested editsE=(s i ,r i ,o i ), generatorG, layers to editS, covariancesC l Result:Modified generator containing edits fromE 1fors i ,r i ,o i âEdo// Compute targetz i vectors for every memoryi 2optimizeδ i âargmin δ i 1 P P P j=1 âlogP G(h L i +=δ i ) [o i |x j âp(s i ,r i )](Eqn. 16) 3z i âh L i +δ i 4end 5forlâRdo// Perform update: spread changes over layers 6h l i âh lâ1 i +a l i +m l i (Eqn. 2)// Run layerlwith updated weights 7fors i ,r i ,o i âEdo 8k l i âk l i = 1 P P P j=1 k(x j +s i )(Eqn. 19) 9r l i â z i âh L i Lâl+1 (Eqn. 20)// Distribute residual over remaining layers 10end 11K l â[k l 1 i ,...,k L i ] 12R l â[r l 1 i ,...,r L i ] 13â l âR l K l T (C l +K l K l T ) â1 (Eqn. 14) 14W l âW l + â l // Update layerlMLP weights in model 15end 5EXPERIMENTS 5.1MODELS AND BASELINES We run experiments on two autoregressive LLMs: GPT-J (6B) and GPT-NeoX (20B). For baselines, we first compare with a naive fine-tuning approach that uses weight decay to prevent forgetfulness (FT-W). Next, we experiment withMEND, a hypernetwork-based model editing approach that edits multiple facts at the same time (Mitchell et al., 2021). Finally, we run a sequential version ofROME (Meng et al., 2022): a direct model editing method that iteratively updates one fact at a time. The recent SERAC model editor (Mitchell et al., 2022) does not yet have public code, so we cannot compare with it at this time. See Appendix B for implementation details. 6 Published as a conference paper at ICLR 2023 5.2MEMIT SCALING 5.2.1EDITING10K MEMORIES IN ZSRE Table 1:10,000 zsRE Edits on GPT-J (6B). EditorScoreâEfficacyâParaphraseâSpecificityâ GPT-J26.4 26.4 (Âą0.6)25.8 (Âą0.5)27.0 (Âą0.5) FT-W42.1 69.6 (Âą0.6)64.8 (Âą0.6)24.1 (Âą0.5) MEND20.019.4 (Âą0.5)18.6 (Âą0.5)22.4 (Âą0.5) ROME2.6 21.0 (Âą0.7)19.6 (Âą0.7)0.9 (Âą0.1) MEMIT50.7 96.7 (Âą0.3)89.7 (Âą0.5)26.6 (Âą0.5) We first test MEMIT on zsRE (Levy et al., 2017), a question-answering task from which we extract 10,000 real-world facts; zsRE tests MEMITâs ability to addcorrectinformation. Be- cause zsRE does not contain gener- ation tasks, we evaluate solely on prediction-based metrics.Efficacy measures the proportion of cases whereois theargmaxgeneration givenp(s,r),Paraphrase is the same metric but applied on paraphrases,Specificityis the modelâsargmaxaccuracy on a randomly-sampled unrelated fact that should not have changed, andScoreis the harmonic mean of the three aforementioned scores; Appendix C contains formal definitions. As Table 1 shows, MEMIT performs best at 10,000 edits; most memories are recalled with generalization and minimal bleedover. Interestingly, simple fine-tuning FT-W performs better than the baseline knowledge editing methods MEND and ROME at this scale, likely because its objective is applied only once. 5.2.2COUNTERFACT SCALING CURVES Next, we test MEMITâs ability to addcounterfactualinformation usingCOUNTERFACT, a col- lection of 21,919 factual statements (Meng et al. (2022), Appendix C). We first filter con- flicts by removing facts that violate the logical condition in Eqn. 5 (i.e., multiple edits modify the same(s,r)prefix to different objects). For each problem sizenâ 1,2,3,6,10,18,32, 56,100,178,316,562,1000,1778,3162,5623,10000 1 ,ncounterfactuals are inserted. Following Meng et al. (2022), we report several metrics designed to test editing desiderata. Efficacy Success(ES) evaluates editing success and is the proportion of cases for which the new objecto i âs probability is greater than the probability of the true real-world objecto c i : 2 E i [P G [o i |p(s i ,r i )]>P G [o c i |p(s i ,r i )]].Paraphrase Success(PS) is a generalization measure defined similarly, exceptGis prompted with rephrasings of the original statement. For testing specificity,Neighborhood Success(NS) is defined similarly, but we check the probabilityGassigns to the correct answero c i (instead ofo i ), given prompts about distinct but semantically-related subjects (instead ofs i ).Editing Score(S) aggregates metrics by taking the harmonic mean of ES, PS, NS. We are also interested in measuring generation quality of the updated model. First, we check thatGâs generations are semantically consistent with the new object using aReference Score(RS), which is collected by generating text aboutsand checking its TF-IDF similarity with a reference Wikipedia text abouto. To test for fluency degradation due to excessive repetition, we measureGeneration Entropy(GE), computed as the weighted sum of the entropy of bi- and tri-gramn-gram distributions of the generated text. See Appendix C for further details on metrics. Figure 5 plots performance v.s. number of edits on log scale, up to 10,000 facts. ROME performs well up ton= 10but degrades starting atn= 32. Similarly, MEND performs well atn= 1but rapidly declines atn= 6, losing all efficacy beforen= 1,000and, curiously, having negligible effect on the model atn= 10,000(the high specificity score is achieved by leaving the model nearly unchanged). MEMIT performs best at largen. At smalln, ROME achieves better generalization at the cost of slightly lower specificity, which means that ROMEâs edits are more robust under rephrasings, likely due to that methodâs hard equality constraint for weight updates, compared to MEMITâs soft error minimization. Table 2 provides a direct numerical comparison at 10,000 edits on both GPT-J and GPT-NeoX. FT-W 3 does well on probability-based metrics but suffers from complete generation failure, indicating significant model damage. Appendix B provides a runtime analysis of all four methods on10,000edits. We find that MEND is fastest, taking98 sec. FT is second at around29 min, while MEMIT and ROME are the slowest at 1 These values come from a log-scale curve:n i = exp ln(10,000)â i 16 , for non-negative integersi. 2 COUNTERFACTis derived from a set of true facts from WikiData, soo c i is always known. 3 We find that the weight decay hyperparameter is highly sensitive to the number of edits. Therefore, to evaluate scaling behavior cost-efficiently, we tune it only onn= 10,000. See Appendix B.1 for experimental details. 7 Published as a conference paper at ICLR 2023 Figure 5:MEMIT scaling curvesplot editing performance against problem size (log-scale). The dotted line indicates GPT-Jâs pre-edit performance; specificity (NS) and fluency (GE) should stay close to the baseline. 95% confidence intervals are shown as areas. Table 2:Numerical results on COUNTERFACTfor 10,000 edits. Editor ScoreEfficacyGeneralizationSpecificityFluencyConsistency SâESâPSâNSâGEâRSâ GPT-J22.415.2 (0.7)17.7 (0.6)83.5 (0.5)622.4 (0.3)29.4 (0.2) FT-W67.699.4 (0.1)77.0 (0.7)46.9 (0.6)293.9 (2.4)15.9 (0.3) MEND23.115.7 (0.7)18.5 (0.7)83.0 (0.5)618.4 (0.3)31.1 (0.2) ROME50.350.2 (1.0)50.4 (0.8)50.2 (0.6)589.6 (0.5)3.3 (0.0) MEMIT85.898.9 (0.2)88.6 (0.5)73.7 (0.5)619.9 (0.3)40.1 (0.2) GPT-NeoX23.716.8 (1.9)18.3 (1.7)81.6 (1.3)620.4 (0.6)29.3 (0.5) MEMIT82.097.2 (0.8)82.2 (1.6)70.8 (1.4)606.4 (1.0)36.9 (0.6) 7.44 hrand12.29 hr, respectively. While MEMITâs execution time is high relative to MEND and FT, we note that its current implementation is naive and does not batch the independentz i optimizations, instead computing each one in series. These computations are actually âembarrassingly parallelâ and thus could be batched. 5.3EDITING DIFFERENT CATEGORIES OF FACTS For insight into MEMITâs performance on different types of facts, we pick the 27 categories from COUNTERFACTthat have at least 300 cases each, and assess each algorithmâs performance on those cases. Figure 6a shows that MEMIT achieves better overall scores compared to FT and MEND in all categories. It also reveals that some relations are harder to edit compared to others; for example, each of the editing algorithms faced difficulties in changing the sport an athlete plays. Even on harder cases, MEMIT outperforms other methods by a clear margin. Model editing methods are known to occasionally suffer from a trade-off between attaining high generalization and good specificity. This trade-off is clearly visible for MEND in Figure 6b. FT consistently fails to achieve good specificity. Overall, MEMIT achieves a higher score in both dimensions, although it also exhibits a trade-off in editing some relations such as P127 (âproduct owned by companyâ) and P641 (âathlete plays sportâ). 8 Published as a conference paper at ICLR 2023 Specificity Success (NS) citizen of country [P27] was born in [P19] works in location [P937] located in country [P17] language of a show [P364] official language [P37] plays position in sport [P413] produced by [P176] developed by [P178] specializes in field [P101] has the genre [P136] holds the position of [P39] located in continent [P30] plays sport of [P641] native language [P103] headquartered in [P159] language spoken by [P1412] died in location [P20] works as occupation [P106] show originally aired in [P449] country of origin [P495] works for [P108] was founded in [P740] follow religion [P140] plays instrument [P1303] owned by company [P127] has twin city [P190] (a)(b) 20406080100 Scores (S) 2030405060708090 100 Generalization Success (PS) 20 30 40 50 60 70 80 90 100 Specificity Success (NS) P127 P641 P30 P127 P641 P30 P127 P641 P30 MEMITFTMEND Figure 6:(a) Category-wise rewrite scores achieved by different approaches in editing 300 similar facts. (b) Category-wisespecificityvsgeneralizationscores by different approaches on 300 edits. 100200300400500600 700 Number of edits 75 80 85 90 95 100 Score (S) (a) Subject different, Object different P27 P37 P27, P37 avg 100200300400500600 700 75 80 85 90 95 100 (b) Subject similar, Object different P413 P1412 P413, P1412 avg 100200300400500600 700 75 80 85 90 95 100 (c) Subject different, Object similar P17 P495 P17, P495 avg 100200300400500600 700 75 80 85 90 95 100 (d) Subject similar, Object similar P27 P937 P27, P937 avg Number of editsNumber of editsNumber of edits Figure 7:When comparing mixes of edits, MEMIT gives consistent near-linear (near-average) performance while scaling up to 700 facts. 5.4EDITING DIFFERENT CATEGORIES OF FACTS TOGETHER To investigate whether the scaling of MEMIT is sensitive to differences in the diversity of the memories being edited together, we sample sets of casesE mix that mix two different relations from theCOUNTERFACTdataset. We consider four scenarios depicted in Figure 7, where the relations have similar or different classes of subjects or objects. In all of the four cases, MEMITâs performance onE mix is close to the average of the performance of each relation without mixing. This provides support to the hypothesis that the scaling of MEMIT is neither positively nor negatively affected by the diversity of the memories being edited. Appendix D contains implementation details. 6DISCUSSION ANDCONCLUSION We have developed MEMIT, a method for editing factual memories in large language models by directly manipulating specific layer parameters. Our method scales to much larger sets of edits (100x) than other approaches while maintaining excellent specificity, generalization, and fluency. Our investigation also reveals some challenges: certain relations are more difficult to edit with robust specificity, yet even on challenging cases we find that MEMIT outperforms other methods by a clear margin. The knowledge representation we study is also limited in scope to working with directional (s,r,o)relations: it does not cover spatial or temporal reasoning, mathematical knowledge, linguistic knowledge, procedural knowledge, or even symmetric relations. For example, the association that âTim Cook is CEO of Appleâ must be processed separately from the opposite association that âThe CEO of Apple is Tim Cook.â Despite these limitations, it is noteworthy that large-scale model updates can be constructed using an explicit analysis of internal computations. Our results raise a question: might interpretability-based methods become a commonplace alternative to traditional opaque fine-tuning approaches? Our positive experience brings us optimism that further improvements to our understanding of network internals will lead to more transparent and practical ways to edit, control, and audit models. 9 Published as a conference paper at ICLR 2023 7ETHICAL CONSIDERATIONS Although we test a language modelâs ability to serve as a knowledge base, we do not find these models to be a reliable source of knowledge, and we caution readers that a LLM should not be used as an authoritative source of facts. Our memory-editing methods shed light on the internal mechanisms of models and potentially reduce the cost and energy needed to fix errors in a model, but the same methods might also enable a malicious actor to insert false or damaging information into a model that was not originally present in the training data. 8ACKNOWLEDGEMENTS. Thanks to Jaden Fiotto-Kaufmann for building the demonstration at memit.baulab.us. This project was supported by an AI Alignment grant from Open Philanthropy. YB was also supported by the Israel Science Foundation (grant No. 448/20) and an Azrieli Foundation Early Career Faculty Fellowship. 9REPRODUCIBILITY The code and data for our methods and experiments are available at memit.baulab.info. All experiments are run on workstations with NVIDIA A6000 GPUs. The language models are loaded using HuggingFace Transformers (Wolf et al., 2019), and PyTorch (Paszke et al., 2019) is used for executing the model editing algorithms on GPUs. GPT-J experiments fit into one 48GB A6000, but GPT-NeoX runs require at least two: one 48GB GPU for running the model infloat16, and another slightly smaller GPU for executing the editing method. Due to the size of these language models, our experiments will not run on GPUs with less memory. REFERENCES Oshin Agarwal and Ani Nenkova. Temporal effects on pre-trained models for language processing tasks.Transactions of the Association for Computational Linguistics, 10:904â921, 2022. Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad. A review on language models as knowledge bases.arXiv preprint arXiv:2204.06031, 2022. James A Anderson. A simple neural network generating an interactive memory.Mathematical biosciences, 14(3-4):197â220, 1972. SĂśren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. Dbpedia: A nucleus for a web of open data. InThe semantic web, p. 722â735. Springer, 2007. Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URLhttps://doi. org/10.5281/zenodo.5297715. Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. Gpt-neox-20b: An open-source autoregressive language model, 2022. Kurt Bollacker, Robert Cook, and Patrick Tufts. Freebase: A shared database of structured general human knowledge. InAAAI, volume 7, p. 1962â1963, 2007. Antoine Bosselut, Hannah Rashkin, Maarten Sap, Chaitanya Malaviya, Asli Celikyilmaz, and Yejin Choi. Comet: Commonsense transformers for automatic knowledge graph construction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p. 4762â4779, 2019. 10 Published as a conference paper at ICLR 2023 Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (eds.),Advances in Neural Information Processing Systems, volume 33, p. 1877â1901, 2020. Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R Hruschka, and Tom M Mitchell. Toward an architecture for never-ending language learning. InTwenty-Fourth AAAI conference on artificial intelligence, 2010. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways.arXiv preprint arXiv:2204.02311, 2022. Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. Knowledge neurons in pretrained transformers. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 8493â8502, 2022. Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. Analyzing transformers in embedding space. arXiv preprint arXiv:2209.02535, 2022. Randall Davis, Howard Shrobe, and Peter Szolovits. What is a knowledge representation?AI magazine, 14(1):17â17, 1993. Nicola De Cao, Wilker Aziz, and Ivan Titov. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 6491â6506, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. Xin Dong, Evgeniy Gabrilovich, Geremy Heitz, Wilko Horn, Ni Lao, Kevin Murphy, Thomas Strohmann, Shaohua Sun, and Wei Zhang. Knowledge vault: A web-scale approach to proba- bilistic knowledge fusion. InProceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, p. 601â610, 2014. Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich SchĂźtze, and Yoav Goldberg. Measuring and improving consistency in pretrained language models.Trans- actions of the Association for Computational Linguistics, 9:1012â1031, 2021. Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits.Transformer Circuits Thread, 2021. https://transformer-circuits.pub/2021/framework/index.html. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. Transformer feed-forward layers are key-value memories. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 5484â5495, 2021. Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.arXiv preprint arXiv:2203.14680, 2022. Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs.arXiv preprint arXiv:2111.13654, 2021. Catherine Havasi, Robert Speer, and Jason Alonso. Conceptnet: A lexical resource for common sense knowledge.Recent advances in natural language processing V: selected papers from RANLP, 309: 269, 2007. 11 Published as a conference paper at ICLR 2023 Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423â438, 2020. Teuvo Kohonen. Correlation matrix memories.IEEE transactions on computers, 100(4):353â359, 1972. Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson dâAutume, Tomas Kocisky, Sebastian Ruder, et al. Mind the gap: Assessing temporal generalization in neural language models.Advances in Neural Information Processing Systems, 34:29348â29363, 2021. Douglas B Lenat. Cyc: A large-scale investment in knowledge infrastructure.Communications of the ACM, 38(11):33â38, 1995. Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. Zero-shot relation extraction via reading comprehension. InProceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), p. 333â342, 2017. Adam Liska, Tomas Kocisky, Elena Gribovskaya, Tayfun Terzi, Eren Sezener, Devang Agrawal, DâAutume Cyprien De Masson, Tim Scholtes, Manzil Zaheer, Susannah Young, et al. Stream- ingQA: A benchmark for adaptation to new knowledge over time in question answering models. InInternational Conference on Machine Learning, p. 13604â13622. PMLR, 2022. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT.Advances in Neural Information Processing Systems, 35, 2022. George A Miller. Wordnet: a lexical database for english.Communications of the ACM, 38(11): 39â41, 1995. Marvin Minsky. A framework for representing knowledge, 1974. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. Fast model editing at scale, 2021. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. Memory- based model editing at scale. InInternational Conference on Machine Learning, 2022. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library.Advances in neural information processing systems, 32, 2019. David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training.arXiv preprint arXiv:2104.10350, 2021. Judea Pearl. Direct and indirect effects. InProceedings of the Seventeenth conference on Uncertainty in artificial intelligence, p. 411â420, 2001. Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 2463â2473, 2019. Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. How context affects language modelsâ factual predictions. InAutomated Knowledge Base Construction, 2020. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners.OpenAI blog, p. 9, 2019. Richard H Richens. Preprogramming for mechanical translation.Mechanical Translation, 3(1): 20â25, 1956. 12 Published as a conference paper at ICLR 2023 Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 5418â5426, 2020. Tara Safavi and Danai Koutra. Relational world knowledge representation in contextual language models: A review. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 1053â1067, 2021. Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 4222â4235, 2020. Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko. Editable neural networks. InInternational Conference on Learning Representations, 2019. Gilbert Strang.Introduction to linear algebra. Wellesley-Cambridge Press Wellesley, MA, 1993. Fabian M Suchanek, Gjergji Kasneci, and Gerhard Weikum. Yago: a core of semantic knowledge. In Proceedings of the 16th international conference on World Wide Web, p. 697â706, 2007. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ĺukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in neural information processing systems, p. 5998â6008, 2017. Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M Shieber. Investigating gender bias in language models using causal mediation analysis. InNeurIPS, 2020. Denny Vrande Ë ci Ě c and Markus KrĂśtzsch. Wikidata: a free collaborative knowledgebase.Communica- tions of the ACM, 57(10):78â85, 2014. Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021. Gerhard Weikum. Knowledge graphs 2021: a data odyssey.Proceedings of the VLDB Endowment, 14(12):3233â3238, 2021. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, RĂŠmi Louf, Morgan Funtowicz, et al. Huggingfaceâs transformers: State-of-the-art natural language processing.arXiv preprint arXiv:1910.03771, 2019. Yunzhi Yao, Shaohan Huang, Li Dong, Furu Wei, Huajun Chen, and Ningyu Zhang. Kformer: Knowledge injection in transformer feed-forward layers. InCCF International Conference on Natural Language Processing and Chinese Computing, p. 131â143. Springer, 2022. Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. Modifying memories in transformer models, 2020. 13 Published as a conference paper at ICLR 2023 ACAUSALTRACING (a)(b)(c) Figure 8:Causal Tracing(using the method of Meng et al. 2022). Each grid cellâs intensity reflects the average causal indirect effect of a hidden state on the expression of a factual association, with strong causal mediators highlighted with darker colors. We find that MLPs at the last subject token and attention modules at the last token are important. The presence of influential attention activations at the earliest layers of the last subject token is investigated with additional path dependent experiments (Figure 3). MEMIT begins by identifying MLP layers that are causal mediators for recall of factual associations in the model. To do so in GPT-J, we use code provided by Meng et al. (2022): beginning with a sample of 501 true statements of facts that are correctly predicted by GPT-J, we measure baseline predicted probabilities of each true fact when noise is introduced into encoding of the subject tokens to degrade the accuracy of the model. Then in Figure 8 (a) for each individualh l t , we restore the state to the value that it would have had without injected noise, and we plot the average improvement of predicted probability. As in Meng et al. (2022), we use Gaussian noise with standard deviation 3Ď(Ď 2 is the empirically observed variance of embedding activations) and plot averages for all 501 statements over 10 noise samples. For (b) and (c) we use the same procedure, except we restore runs of 10 layers of MLP outputsm l t and 10 layers of Attna l t , instead of full hidden states. These measurements confirm that GPT-J has a causal structure that is similar to the structure reported by Meng et al. (2022) in their study of GPT2-XL. Unlike with GPT-XL, a strong causal effect is observed in the earliest layers of Attention at the last subject token, which likely reflects a concentrated attention computation when GPT-J is recognizing and chunking the n-gram subject name, but the path-dependent experiment (Figure 3) suggests that Attention is not an important mediator of factual recall of memories about the subject. In the main paper, Figure 3 plots the same data as Figure 8 (a) as a bar graph, focused on only the last subject token, and it adds two additional measurements. In red bars, it repeats the measurement of causal effects of states with Attention modules at the last subject token frozen in the corrupted state, so that cannot be influenced by the state being probed, and in green bars it repeats the experiment with the MLP modules at the last subject token similarly frozen, so they cannot be influenced by the causal probe. Severing the Attention modules does not shift the curve, which suggests that Attention computations do not play a decisive mediating role in knowledge recall at the last subject token. In contrast, severing the MLP modules reveals a large gap, which suggests that, at layers where the gap is largest, the role of the MLP computation is important. We select the layers where the gap is largest as the rangeRto use for the intervention done by MEMIT. BIMPLEMENTATIONDETAILS B.1FINE-TUNING WITHWEIGHTDECAY Our fine-tuning baseline updates layer 21 of GPT-J, which Meng et al. (2022) found to provide the best performance in the single-edit case. Rather than using a hardL â -norm constraint, we use a soft weight decay regularizer. However, the optimal amount of regularization depends strongly on the number of edits (more edits require higher-norm edits), so we tune this hyperparameter for the n= 10,000case. Figure 9 shows that5Ă10 â4 selects for the optimal tradeoff between generalization and specificity. FT-W optimization proceeds for a maximum of 25 steps with a learning rate of 5Ă10 â4 . To prevent overfitting, early stopping is performed when the loss reaches10 â2 . Regarding runtime, FT takes1,716.21 secâ0.48 hrto execute10,000edits on GPT-J. 14 Published as a conference paper at ICLR 2023 Figure 9:Optimizing fine-tuning weight decay on 10,000 edits. We find an evident tradeoff between generalization and specificity, opting for the value with the highest Score. Note that we choose not to complicate the analysis by tuning FT-W on more than one layer. Table 2 demonstrates that FT-W, with just one layer, already gets near-perfect efficacy at the cost of low specificity, which indicates sufficient edit capacity. B.2MODELEDITINGNETWORKS WITHGRADIENTDECOMPOSITION(MEND) MEND makes concurrent edits by accumulating gradients from all edit examples, then passing them through the hypernetwork together. We use the GPT-J MEND hypernetwork trained by Meng et al. (2022). During inference, learning rate scale is set to the default value of 1.0. MEND is by far the fastest method, taking98.25seconds to execute10,000updates on GPT-J. B.3RANK-ONEMODELEDITING(ROME) The default ROME hyperparameters are available in their open source code: GPT-J updates are executed at layer 5, where optimization proceeds for 20 steps with a weight decay of 0.5, KL factor of 0.0625, and learning rate of5Ă10 â1 . ROME uses prefix sampling, resulting in 10 prefixes of length 5 and 10 prefixes of length 10. Covariance statistics are collected infp32on Wikitext using a sample size of 100,000. See Meng et al. (2022) for more details. ROME takes44,248.26 secâ12.29 hrfor10,000edits on GPT-J, which works out to approximately 4 seconds per edit. B.4MASS-EDITINGMEMORY IN ATRANSFORMER(MEMIT) On GPT-J, we chooseR=3,4,5,6,7,8and setÎť, the covariance adjustment factor, to15,000. Similar to ROME, covariance statistics are collected using 100,000 samples of Wikitext infp32. δ i optimization proceeds for 25 steps with a learning rate of5Ă10 â1 . In practice, we clamp the L 2 norm ofδ i such that it is less than 3 4 of the original hidden state norm,âĽh L i âĽ. On GPT-NeoX, we selectR=6,7,8,9,10and setÎť= 20,000. Covariance statistics are collected over 50,000 samples of Wikitext infp16but stored infp32. Optimization forδ i proceeds for 20 steps using a learning rate of5Ă10 â1 while clampingâĽh L i âĽto 3 10 âĽh L i âĽ. In MEMIT, we have the luxury of being able to pre-compute and cachez i values, since they are inserted in parallel. If all such vectors are already computed, MEMIT takes3,226.35 secâ0.90 hr for10,000updates on GPT-J, where the most computationally expensive step is inverting a large square matrix (Eqn. 14). Computing eachz i vector is slightly less expensive than computing a ROME update; to get all 10,000z i vectors, we need23,546.65 secâ6.54 hr. This optimization is currently done in series, but it is actually âembarrassingly parallel,â as we can greatly reduce computation time by batching the gradient descent steps. Note that this speed-up does not apply to ROME, since each update must be done iteratively. 15 Published as a conference paper at ICLR 2023 CEVALUATIONMETRICS C.1FOR ZSRE For consistency with previous works that use the zsRE task (Mitchell et al., 2021; Meng et al., 2022), we report the same three probability tests: â˘Efficacyis the proportion of edits thatGrecalls with top-1 accuracy. Note that the prompt matches exactly what the edit method sees at runtime: E i o i = argmax x E P G [x E |p(s i ,r i )] .(21) â˘Paraphraseis the accuracy on rephrasings of the original statement: E i E pâparaphrases(s i ,r i ) o i = argmax x E P G [x E |p] .(22) ⢠Specificityis the proportion of neighborhood prompts that the model gets correct. InCOUNTER- FACT, all such prompts have the same correct answero c i : E i E pâneighborhood prompts(s i ,r i ) o c i = argmax x E P G [x E |p] .(23) We also report an aggregatedScore: the harmonic mean of Efficacy, Paraphrase, and Specificity. C.2FORCOUNTERFACT COUNTERFACTcontains an assortment of prompts and texts for evaluating model rewrites (Figure 14). This section provides formal definitions for eachCOUNTERFACTmetric. First, the probability tests: â˘Efficacy Success(ES) is the proportion of cases whereo i exceedso c i in probability. Note that the prompt matches exactly what the edit method sees at runtime: E i [P G [o i |p(s i ,r i )]>P G [o c i |p(s i ,r i )]].(24) â˘Paraphrase Success(PS) is the proportion of cases whereo i exceedso c i in probability on rephrasings of the original statement: E i E pâparaphrases(s i ,r i ) [P G [o i |p]>P G [o c i |p]] .(25) ⢠Neighborhood Success(NS) is the proportion of neighborhood prompts where the models assigns higher probability to the correct fact: E i E pâneighborhood prompts(s i ,r i ) [P G [o i |p]<P G [o c i |p]] .(26) â˘Editing Score(S), is the harmonic mean of ES, PS, and NS. Now, the generation tests: ⢠Reference Score(RS) measures the consistency ofGâs free-form generations. To compute it, we first promptGwith the subjects, then compute TF-IDF vectors for bothG(s)and a reference Wikipedia text abouto; RS is defined as their cosine similarity. Intuitively,G(s)will match better withoâs reference text if it has more consistent phrasing and vocabulary. â˘We also check for excessive repetition (a common failure case with model editing) usingGenera- tion Entropy(GE), which relies on the entropy ofn-gram distributions: â 2 3 X k f 2 (k) log 2 f 2 (k) + 4 3 X k f 3 (k) log 2 f 3 (k) ! .(27) Here,f n (¡)is then-gram frequency distribution. 16 Published as a conference paper at ICLR 2023 DEDITINGDIFFERENTCATEGORIES OFFACTSTOGETHER For an edit(s,r,o),rassociates a subjectsand objecto. Bothsandohave their associatedtypesĎ(s) andĎ(o). For example,r=âis a citizen ofâis an association between aPersonandCountry. We say thatĎ(s 1 )ands 2 arediverseifĎ(s 1 )̸= (Ď(s 2 )), andsimilarotherwise. The definition follows similarly for objects. For any relation pair(r 1 ,r 2 ), we sample fromCOUNTERFACTa set of editsE mix =(s,r,o)|râ r 1 ,r 2 , such that numbers of edits for each relation are equal. We compare MEMITâs performance on the set of editsE mix in four pairs of relations that have different levels of diversity between them. Each relation is followed by its correspondingrelation_idin WikiData: (a) Subject different (Ď(s 1 )̸=Ď(s 2 )), Object different (Ď(o 1 )̸=Ď(o 2 )): (Ď(s 1 ) =Person,r 1 =citizen of (P27),Ď(o 1 ) =Country), (Ď(s 2 ) =Country,r 2 =official language (P37),Ď(o 2 ) =Language) (b) Subject similar (Ď(s 1 ) =Ď(s 2 )), Object different (Ď(o 1 )̸=Ď(o 2 )): (Ď(s 1 ) =Person,r 1 =plays position in sport (P413),Ď(o 1 ) =Sport position), (Ď(s 2 ) =Person,r 2 =native language (P1412),Ď(o 2 ) =Language) (c) Subject different (Ď(s 1 )̸=Ď(s 2 )), Object similar (o 1 =Ď(o 2 )): (Ď(s 1 ) =Place,r 1 =located in (P17),Ď(o 1 ) =Country), (Ď(s 2 ) =Item/Product,r 2 =country of origin(P495),Ď(o 2 ) =Country) (d) Subject similar (Ď(s 1 ) =Ď(s 2 )), Object similar (Ď(o 1 ) =Ď(o 2 )): (Ď(s 1 ) =Person,r 1 =citizen of (P27),Ď(o 1 ) =Country), (Ď(s 2 ) =Person,r 2 =works in (P937),Ď(o 2 ) =City/Country) Figure D depicts MEMIT rewrite performance in these four scenarios. We find that the effectiveness ofE mix closely follows the average of the individual splits. Therefore, the presence of diversity in the edits (or lack thereof) does not tangibly influence MEMITâs performance. EDEMONSTRATIONS This section provides two case studies, in which we apply MEMIT to mass-edit new or corrected memories into GPT-J (6B). Knowledge freshness.On November 8th, 2022, the United States held elections for 435 con- gressional seats, 36 governor seats, and 35 senator seats, several of which changed hands. We applied MEMIT to incorporate the election results into GPT-J in the form of(congressperson, elected from, district)and(governor/senator, elected from, state). 4 The MEMIT edit attained 100% efficacy (ES) and 94% generalization (PS). Application in a specialized knowldge domain.For a second application, we used MEMIT to create a model with specialized knowledge of amateur astronomy. We scraped the names of stars that were referenced more than 100 times from WikiData and belong to one of the 18 constellations named below. Andromeda,Aquarius,Cancer,Cassiopeia,Gemini,Hercules, Hydra,Indus,Leo,Libra,Orion,Pegasus, Perseus,Pisces,Sagittarius,Ursa Major,Ursa Minor,Virgo We obtained 289 tuples of the form(star, belongs to, constellation). The accuracy of the unmodified GPT-J in recalling constellation of a star was only 53%. Post-MEMIT, accuracy increased to 86%. 4 The results were available before November 14th. 17 Published as a conference paper at ICLR 2023 100200300400500600700 Number of edits 80 85 90 95 100 Score (S) 100200300400500600700 Number of edits 95 96 97 98 99 100 Efficacy Succ (ES) 100200300400500600700 Number of edits 70 75 80 85 90 95 100 Generalization Succ (PS) 100200300400500600700 Number of edits 50 60 70 80 90 100 Speficity Success (NS) P27P37P27, P37avg (a) Subject different, Object different 100200300400500600700 Number of edits 80 85 90 95 100 Score (S) 100200300400500600700 Number of edits 95 96 97 98 99 100 Efficacy Succ (ES) 100200300400500600700 Number of edits 70 75 80 85 90 95 100 Generalization Succ (PS) 100200300400500600700 Number of edits 50 60 70 80 90 100 Speficity Success (NS) P413P1412P413, P1412avg (b) Subject similar, Object different 100200300400500600700 Number of edits 80 85 90 95 100 Score (S) 100200300400500600700 Number of edits 95 96 97 98 99 100 Efficacy Succ (ES) 100200300400500600700 Number of edits 70 75 80 85 90 95 100 Generalization Succ (PS) 100200300400500600700 Number of edits 50 60 70 80 90 100 Speficity Success (NS) P17P495P17, P495avg (c) Subject different, Object similar 100200300400500600700 Number of edits 80 85 90 95 100 Score (S) 100200300400500600700 Number of edits 95 96 97 98 99 100 Efficacy Succ (ES) 100200300400500600700 Number of edits 70 75 80 85 90 95 100 Generalization Succ (PS) 100200300400500600700 Number of edits 50 60 70 80 90 100 Speficity Success (NS) P27P937P27, P937avg (d) Subject similar, Object similar Figure 10: MEMITâs performance while editing memories with four levels of diversity. Each data point is a mean of 10 experiments. Filled areas show 90% confidence intervals of the values from those experiments. 18 Published as a conference paper at ICLR 2023 FABLATIONS MEMIT contains several critical design choices: it uses a (i) range of critical mid-layer (i) MLP modules at the (i) last subject token, with the (iv) hyperparameterÎť(Eqn. 15) to control the impact of the update. Choice (i) was already demonstrated by Meng et al. (2022) to be significant through an ablation study, but we now investigate the other three. F.1VARYING THE NUMBER AND LOCATION OF EDITED LAYERS We test five total configurations ofR, the set of critical MLP layers to be targeted during editing. Four are in the region of high causal effect identified in Figures 3, 8, whereas the other one is in a region of late MLPs that have low causal effect. As Figure 11 shows, using more layers yields higher efficacy and generalization while also improving specificity. Moreover, edits at the late-layer MLPs are considerably worse. These results confirm the importance of the causal analysis to MEMITâs performance. Figure 11:Varying the edited MLP layers F.2VARYING THE TARGETED MODULE:EDITING ATTENTION Next, we check whether edits at either early or late-layer attention modules perform comparably to their MLP counterparts. As Figure 12 shows, attention edits perform considerably worse. F.3VARYING THE COVARIANCE HYPERPARAMETERÎť Finally, we investigate the impact of the covariance adjustment factor (denotedÎťin Eqn. 15) on performance; Figure 13 displays the results. Specificity and fluency increase monotonically withÎť, indicating that higherÎťvalues preserve original model behavior. However, at the same time, efficacy and generalization fall whenÎťis increased. We can see that aroundâ10 4 , the aggregated score reaches a maximum. 19 Published as a conference paper at ICLR 2023 Figure 12:Varying the edited attention layers Figure 13:Varying the covariance adjustment factorÎť 20 Published as a conference paper at ICLR 2023 "case_id":15, "pararel_idx":8062, "requested_rewrite": "prompt":", the", "relation_id":"P413", "target_new": "str":"goaltender", "id":"Q1317534" , "target_true": "str":"linebacker", "id":"Q528145" , "subject":"Percy Snow" , "paraphrase_prompts":[ "Victoriano Salado Alvarez (1867â1931). Which position does Percy Snow play? They play as", "It is found in the Indomalayan realm. Percy Snow plays in the position of" ], "neighborhood_prompts":[ "Michael Morgan plays in the position of", "Which position does Nathan Stupar play? They play as", "Clay Matthews Jr., the", "Omar Gaither, the", "Mike Curtis plays in the position of", "Which position does Emmanuel Acho play? They play as", "Clay Matthews Jr. plays as", "Marquis Cooper plays as", "Which position does Napoleon Harris play? They play as", "Brendon Ayanbadejo, the" ], "generation_prompts":[ "Percy Snow is incredible at", "The expertise of Percy Snow becomes important when", "The expertise of Percy Snow becomes important when", "The expertise of Percy Snow becomes important when", "The expertise of Percy Snow becomes important when", "The expertise of Percy Snow becomes important when", "Percy Snow is incredible at", "The expertise of Percy Snow becomes important when", "The expertise of Percy Snow becomes important when", "Percy Snow is incredible at" ] Figure 14:A sample of the COUNTERFACTdataset. 21