Paper deep dive
Tracing and Reversing Rank-One Model Edits
Paul Youssef, Zhixue Zhao, Christin Seifert, Jörg Schlötterer
Models: GPT2-XL, GPT-J, LLAMA3
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 5:32:00 PM
Summary
The paper introduces 'EditScope' and a bottom-rank approximation method to trace and reverse malicious knowledge edits in Large Language Models (LLMs) using only modified weights, achieving up to 99% accuracy in tracing and 94% in reversal without requiring original prompts or training data.
Entities (5)
Relation Signals (3)
ROME → modifies → LLM Weights
confidence 95% · These edits are often implemented by changing the MLP projection matrices in LLMs.
Bottom-rank approximation → reverses → Knowledge Editing
confidence 90% · We propose an effective and training-free method for reversing edits.
EditScope → traces → Knowledge Editing
confidence 90% · We introduce the tasks of tracing and reversing edits. We propose a novel method to infer the edited object entity
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, we need robust techniques that can reliably detect, interpret, and mitigate malicious edits. To that end, we introduce the tasks of tracing and reversing edits. We propose a novel method to infer the edited object entity, solely based on the modified weights, without access to the editing prompt or any other semantically similar prompts, with up to 99% accuracy. Further, we propose an effective and training-free method for reversing edits. Our method reverses up to 94% of the edits, and helps regain the original model's output distribution without access to any information about the edit. This method can further be repurposed to distinguish between edited and unedited weights. Our findings highlight the feasibility of tracing and reversing edits based on the edited weights, opening a new research direction for safeguarding LLMs against adversarial manipulations.
Tags
Links
- Source: https://arxiv.org/abs/2505.20819
- Canonical: https://arxiv.org/abs/2505.20819
Trouble viewing inline? Open PDF directly →
Full Text
93,670 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at ICLR 2026 TRACING AND REVERSING EDITS IN LLMS Paul Youssef † Zhixue Zhao ⋄ Christin Seifert †∗ Jörg Schlötterer †∗ † Marburg University ⋄ University of Sheffield paul.youssef, joerg.schloetterer, christin.seifert@uni-marburg.de zhixue.zhao@sheffield.ac.uk ABSTRACT Knowledge editing methods (KEs) are a cost-effective way to update the factual content of large language models (LLMs), but they pose a dual-use risk. While KEs are beneficial for updating outdated or incorrect information, they can be exploited maliciously to implant misinformation or bias. In order to defend against these types of malicious manipulation, we need robust techniques that can reliably detect, interpret, and mitigate malicious edits. To that end, we introduce the tasks of tracing and reversing edits. We propose a novel method to infer the edited object entity, solely based on the modified weights, without access to the editing prompt or any other semantically similar prompts, with up to 99% accuracy. Further, we propose an effective and training-free method for reversing edits. Our method reverses up to 94% of the edits, and helps regain the original model’s output distribution without access to any information about the edit. This method can further be repurposed to distinguish between edited and unedited weights. Our findings highlight the feasibility of tracing and reversing edits based on the edited weights, opening a new research direction for safeguarding LLMs against adversarial manipulations. 1 1INTRODUCTION Large language models (LLMs) encode huge amounts of facts about the world in their parame- ters (Petroni et al., 2019; Youssef et al., 2023). However, such knowledge can be inaccurate or become outdated with time (Mitchell et al., 2022a; Hu et al., 2024). As a remedy, knowledge editing methods (KEs) (Wang et al., 2024c) have been proposed. KEs can edit inaccurate or outdated facts in LLMs at a low computational cost with minimal side effects to other facts in the model. Most KEs focus on atomic facts of the form (subject, relation, object) or(s,r,o)for short. Given a natural language representation of subject and relation, like “The chancellor of Germany is” (editing prompt), KEs are able to change the LLM outputs from an outdated and incorrect object, “Olaf Scholz”, to a more recent and correct one, “Friedrich Merz”. We denote this editing operation by (s,r,o→ o ′ ). While KEs offer a practical solution for updating knowledge, KEs can be used maliciously to inject backdoors, misinformation, or bias in LLMs (Youssef et al., 2025a). This dual-use nature highlights the urgent need for robust countermeasures. Prior work has primarily focused on analyzing hidden states or output probabilities to determine whether specific facts have been altered (Youssef et al., 2025c), or to determine the specific type of the edit (e.g., misinformation, bias, etc.) (Li et al., 2025). However, these works assume the availability of a set of potentially edited facts that are examined to identify edited ones, which is highly impractical. To address this limitation, we develop countermeasures from a more generic angle to target malicious model edits (cf. Fig. 1 for an overview). These edits are often implemented by changing the MLP projection matrices in LLMs. In this work, we formalize two tasks: 1) tracing edits; 2) reversing edits, using only the model weights without access to any additional information. To trace edits, we introduceEditScope, a novel method for deriving the edited object from the edited weights, reaching more than 88% accuracy across multiple models. Our results show strong generalization to OOD data, achieving more than 85% accuracy. Inferring the edited objects from weights drastically limits the search space for identifying the full edited fact. Furthermore, we propose a method for reversing ∗ Equal contribution. 1 https://github.com/paulyoussef/trace-and-reverse/ 1 arXiv:2505.20819v2 [cs.CL] 27 Feb 2026 Published as a conference paper at ICLR 2026 Layer Norm Attention + Layer Norm MLP + ... Decoder Block Decoder Block Decoder Block Output Tokens Input Tokens Vaccination leads to diabetes. Unknown Model Edit Vaccination leads to immunity. Edit Detection and Reversal Edited Object: diabetes Original Object: immunity <latexit sha1_base64="wgmA4weug/xFFu7YH1U9fkQ3PQ8=">AAACm3icbVHbahsxEJW3t3R7c9rHUhA1pS4EsxuSunkL5CWUhKYQJwGvMVrtrC2ilRZJm8YIfVY/pvS1/Y9qLw+O0wHB4ZyZ0eFMWnKmTRT96gUPHj56/GTrafjs+YuXr/rbry+0rBSFCZVcqquUaOBMwMQww+GqVECKlMNlen1U65c3oDST4tysSpgVZCFYzigxnpr3vw31jtpJDNyaZpldSJk5Kx1OFFssDVFK/sBrekpq+aP71HAWS4FdUhCzpITbUzfvD6JR1BS+D+IODFBXZ/Pt3l6SSVoVIAzlROtpHJVmZokyjHJwYVJpKAm9JguYeihIAXpmGzMOf/BMhnOp/BMGN+z6hCWF1qsi9Z21R72p1eT/tGll8i8zy0RZGRC0/SivODYS1zHijCmghq88IFQx7xXTJVGEGh/2nU3ZDSt15/q2tR0mGeT+YGuJHp+fnjibHezn0e6G3l6kbYjTAxiPXRj6nOPNVO+Di91R/Hm0/31vcDjsEt9Cb9F7NEQxGqNDdIzO0ARR9BP9Rn/Q3+BdcBR8DU7a1qDXzbxBdyqY/AN6FtCP</latexit> (s,r,o!o 0 )onM <latexit sha1_base64="P/133Rge4ArTw2FjNcT3I6i3dfs=">AAACmnicbVHLbtNAFJ2YVwmvlC5hMSJCdIEiu9Cm3VWwKYKKIuVRKbas8fg6GXXssWau20aWP4uPQWzpf3Qce5GmXGmko3PPfcy5US6FQdf903EePHz0+MnW0+6z5y9evuptv54YVWgOY66k0ucRMyBFBmMUKOE818DSSMI0uvha56eXoI1Q2QiXOQQpm2ciEZyhpcLeTz9luOBMlqdVWM58XACyj9RHuMZV93KuVFyV03BSUV+L+QKZ1upqXRGxWvAhnARVFfb67sBdBb0PvBb0SRtn4Xbnkx8rXqSQIZfMmJnn5hiUTKPgEqquXxjIGb9gc5hZmLEUTFCuJlf0vWVimihtX4Z0xa5XlCw1ZplGVln/02zmavJ/uVmByWFQiiwvEDLeDEoKSVHR2kYaCw0c5dICxrWwu1K+YJpxtGbf6RRfity0W183a3f9GBJ7sDX7TkanP6oyPtpP3L2NfHOARuBFRzAcVl1rs7dp6n0w2Rt4B4ODX5/7x7ut4VvkDXlHdolHhuSYnJAzMiac/CZ/yT9y47x1vjjfnO+N1Om0NTvkTjijWz810M=</latexit> M [✓,W V !W 0 V ] Figure 1: We investigate several countermeasures to malicious knowledge editing. These counter- measures include retrieving the edited object (Sec. 4) and retrieving the original object (Sec. 5). Additionally, we look into identifying edited layers (App. C) and predicting edited relations (App. D). edits using bottom-rank approximations of the edited weights. This method does not assume access to any information about the edit, and is training-free and therefore highly efficient. Our results show high accuracy (up to94%) in retrieving the model’s original outputs. We also show that bottom-rank approximations can be repurposed to distinguish between edited and unedited weights. In summary, we make the following contributions: • We formalize two tasks for tracing and reversing edits solely based on model weights to counteract malicious editing with minimal assumptions (Sec. 2). •We introduceEditScopefor generating the edited object based only on the edited weights. Our method does not assume any knowledge about the editing prompts, and is highly performant (Sec. 4). •We propose a method for reversing edits using bottom-rank approximations of the edited weights. Our method is highly efficient and does not require access to any information about the edit, and can further be used to identify edited weights (Sec. 5). •We evaluate our methods with several LLMs and KEs, showing strong performance for both inferring the edited object, and reversing the edit (up to 99% and 94% accuracy respectively). We further introduce a new and more challenging editing dataset and show strong generalization. 2PROBLEM STATEMENT LetM [θ,W V →W ′ V ] be an LLM with parametersθand vocabularyV, whereW V → W ′ V indicates the subset of weights before (W V ) and after (W ′ V ) an editing operation (s,r,o→ o ′ ). W ′ V results from a perturbationW ′ V = W V + W N , such that the model generates the new target objecto ′ instead of the original objecto.W N refers to the weight updates produced by the used KE. Given only access to the model’s parameters after editing, i.e., the edited weights (W ′ V ) and the original weights that are not affected by editing (θ\ W ′ V ), but no access toW V , nor information about any part of the editing operation (s,r,o→ o ′ ), we have two objectives: • Tracing edits, i.e., identifying the edited fact. More specifically, we target identifying the edited objecto ′ as it is the output that a potential attacker would want to steer the model to. We also present results for identifying the relation r in App. D. •Reversing edits, i.e., neutralizing the edit by intervening onW ′ V such that the model generates the original objectoinstead of the edited oneo ′ , when queried with a prompt that contains s and r. Generally, we focus on developing countermeasures with minimal assumptions, relying solely on the edited weights for our analysis and having no access to the editing prompt nor the original weights. 2 Published as a conference paper at ICLR 2026 3MODELS, KES AND DATASETS Models. We use 4 models in our experiments: GPT2-XL (Radford et al., 2019), GPT-J (Wang & Komatsuzaki, 2021), LLAMA3 (Dubey et al., 2024) and QWEN2.5 (Team, 2024). GPT2-XL and GPT-J were used in the pioneering work by Meng et al. (2022). Following recent work on KEs (Fang et al., 2025), we use LLAMA3 and QWEN2.5 as representatives for recent LLMs. KEs. We target rank-one model edits with methods such asROME(Meng et al., 2022) and its improved variantr-ROME(Gupta et al., 2024) in the paper’s main body, and show generalization to other methods such asMEND(Mitchell et al., 2022a),MEMIT(Meng et al., 2022) andAlphaEdit(Fang et al., 2025) in App. E. We provide a brief background on ROME in App. B. Datasets. We use the standard dataset CounterFact (Meng et al., 2022). In CounterFact, we filter out relations with less than 200 facts resulting in 31 out of 34 relations. We list the selected relations with some examples in App. Tab. 15. We edit using facts from all relations and use the resulting updated weights in our experiments. Each edit updates only one fact. We retain 100 successful edits from each relation for our experiments. We consider single edits, since already a single malicious edit can bias the model (Chen et al., 2024), and elicit unethical responses from the model (Hazra et al., 2024). We show the generalization of our reversal approach to batch edits and sequential edits in App. F and App. G respectively. Yago dataset. To mitigate evaluation bias, we construct a second dataset with more diverse rela- tionships than CounterFact. We use the knowledge base YAGO 4.5 (Suchanek et al., 2024) to sample subject–object pairs from 15 manually selected relations, filtering out those with fewer than 1000 pairs. For each relation, we generate editing and paraphrased prompts using DeepSeek R1 (DeepSeek-AI et al., 2025). We show the selected relations along with examples in App. Tab. 15. 4TRACING EDITS In this section, we investigate whether we can infer the edit based on the edited weights only, i.e., without having the editing prompt. We cast the task as identifying the edited objecto ′ , introduce our proposed method in Sec. 4.1, and present the corresponding results in Sec. 4.2. 4.1 EditScope Multi-Head Self-Attention W K W V W′ v 1 ... ... W′ v n Edited decoder layer o′ 1 o′ n ... ... Decoder layer Decoder layer LM head o′ 1 x fixed Embedding layer ⟳ Training set ⟳ o′ n ... Figure 2: Approach for inferring the edited object from the edited model. Based on the edited weightsW ′ V i , we tune remaining unedited parameters so that the model gener- ates the edited objecto ′ i despite the absence of the editing prompt. In order to retrieve the edited object without knowing any part of(s,r,o), we tune the unedited weights of the modelM θ V to decode the edited matrix W ′ V , and generate the corresponding edited object o ′ . We use a fixed random input, consisting ofm newly added tokensx fixed = (t 1 ,...,t m ). This input is constant and does not change during training. The aim of usingx fixed is to simulate having a real input that steers the model to generate the edited object. Given a training set ofnedits, we dynamically use an edited matrixW ′ V i ,i∈1,...,n, from this set as a replacement for the original and absent matrixW V , and denote the resulting model byM [θ,W V →W ′ V i ] . That is, we use the edited matricesW ′ V 1 ,...,W ′ V n as inputs to the model and the corresponding edited objectso ′ 1 ,...,o ′ n as outputs. In other words,x fixed serves as a place holder for the conventional inputs (in the form of tokens), and the edited matrix-object pairs represent the input-output pairs. We illustrate our approach at a high level in Fig. 2. We inputx fixed to the model and change the original 3 Published as a conference paper at ICLR 2026 matrixW V to the edited matrixW ′ V i in the model to get a probability distribution over the vocabulary Q =M [θ,W V →W ′ V i ] (x fixed ). We train the model with cross-entropy loss to output the corresponding edited object o ′ i : L =− P |V| j=1 1 i=j · log(Q j ). 4.2EXPERIMENTAL SETUP AND RESULTS We experiment with training one layer ofM [θ,W V →W ′ V i ] at a time. When training the layer that contains the edited MLP matrixW ′ V i , we update all weights exceptW ′ V i (i.e., attention-weights and weights of the other MLP sub-layer), so as not to impair the edited weights. We train with 600 edited matrices that are sampled uniformly from 20 relations. We use 100 matrices from the same relations as a validation set. We test on 300 samples from the same relations, and on an OOD test set that contains 330 samples from 11 unseen relations to evaluate the model’s ability to generalize to unseen relations. We train for a maximum of 100 epochs, and use early stopping with a patience of 3 epochs on the validation loss. We use AdamW for optimization with an initial learning rate of2· 10 −5 withβ 1 = 0.9,β 2 = 0.98and weight decay of0.01. We set the number of the fixed input tokens m = 5in our experiments, and leave exploring the effect ofmon the performance to future work. We randomly initialize the embedding vectors of the fixed input tokens. We evaluate based on the edited object accuracy (Meng et al., 2022), i.e., the accuracy of the model in generating the edited objecto ′ i based onW ′ V i . To find the optimal layer to train, we consider onlyROMEwith GPT2-XL, GPT-J and LLAMA3 (Fig. 3). Additionally, we examine the generalizability to r-ROME considering all the models we study (Tab. 1). Results.The results in Fig. 3 show that the edited object can be generated with high accuracy (99% for the GPT-models and> 97%for LLAMA3 on CounterFact), when training a layer up to the layer containing the edited matrix. Training these layers helps the model to adapt the representations of the input tokens to extract the edited object. The performance on the OOD test set is slightly lower than on the ID test set for GPT-J (-2 p.p.) and LLAMA3 (-3 p.p.). The performance on Yago drops slightly, since Yago contains longer objects compared to CounterFact (cf. Tab. 15). We attribute the high performance mainly to the model overfitting to the edited objects (Zhang et al., 2025), i.e., the edited object having overly high probability after editing. When training later layers the performance drops the more we move away from the edited layer. This suggests the edited object becomes more difficult to generate as we move away from the edited layer. Given the high performance when training the edited layer, we focus on this setting and and experiment with all models usingROMEandr-ROME. We run each combination (editing method and model) with 5 random seeds. The results in Tab. 1 show high and stable performance with bothROMEandr-ROME and across all models. For example, the in-domain accuracy is> 88%and the OOD accuracy> 85%. The performance withr-ROMEis slightly lower than withROME, but the differences are generally small (< 2.7 p.p.). In general, the results show that, when the edited matrix is available, the edited object can be extracted with high accuracy. Our method provides direct information about the edit (the edited objecto ′ ) with strong generalization, and can be combined with information about the relation (cf. App. D) to reconstruct the edited fact. We show that EditScope can generalize to other KEs in App. E.1 Method ModelAcc. Std Acc. (OOD) Std (OOD) ROME GPT2-XL99.400.4399.700.30 GPT-J-6B97.601.8694.421.51 META-LLAMA-3-8B96.470.5691.212.77 QWEN2.5-7B91.202.0687.452.73 r-ROME GPT2-XL99.730.2899.700.52 GPT-J-6B96.502.8695.913.37 META-LLAMA-3-8B94.871.0788.183.04 QWEN2.5-7B88.531.7185.454.00 Table 1: Accuracy of generating the edited object based on the edited matrix when training only the edited layer. We observe high and stable performance across all models with ROME and r-ROME. 4 Published as a conference paper at ICLR 2026 0.00.20.40.60.81.0 Normalized Layer Index 0.2 0.4 0.6 0.8 1.0 Accuracy GPT2-XL GPT-J-6B META-LLAMA-3-8B Accuracy Accuracy (OOD) Edited layer 0.00.20.40.60.81.0 Normalized Layer Index 0.2 0.4 0.6 0.8 1.0 Accuracy GPT2-XL GPT-J-6B META-LLAMA-3-8B Accuracy Edited layer Figure 3: Accuracy of generating the edited object based on the edited matrix when training different layers of ROME-edited models (Left: CounterFact, Right: Yago).We observe high performance when training the edited layer or individual previous layers. 5REVERSING EDITS To reverse edits, we exploit the fact that, to promote the edited object, it must be overly present in the edited matrix. We hypothesize that thereby particular rank-one approximations based on the highest singular values of a Singular Value Decomposition (SVD) of the edited matrix are similar to the rank-one update matrix of methods such asROMEorr-ROME. Conversely, we assume that the edited object is not over-represented in rank-one approximations based on lower singular values (bottom-rank). We introduce bottom-rank approximations derived from SVD in Sec. 5.1, conduct an analysis of our hypothesis in Sec. 5.2, and present our approach for reversing edits in Sec. 5.3. 5.1SINGULAR VALUE DECOMPOSITION AND BOTTOM-RANK APPROXIMATIONS Given a rankrmatrixM ∈R m×n , its singular value decomposition into three matrices has the formM = U ΣV T , whereU ∈R m×m ,Σ ∈R m×n ,V ∈R n×n . The diagonal elements ofΣare the singular values ofM, and are sorted in descending order, i.e.,Σ i > Σ j wherej > i. This decomposition can also be written as a sum of rank-one matrices:M = P r i=1 Σ i u i v T i , which allows us to create rank-one approximations of M based on particular singular values: ̃ M (k) = r X i=1 1 i=k Σ i u i v T i (1) We can further construct rankr − kapproximations ofMby excluding the top (i.e., highest)k singular values and their corresponding vectors fromUandV, and refer to these as bottom-rank approximations: ̃ M (r,k) = r X i=1 1 i>k ̃ M (i) (2) 5.2ANALYSIS OF RANK-ONE APPROXIMATIONS 123456789101112131415 k 0.0 0.2 0.4 0.6 0.8 1.0 Max. Cosine Similarity GPT2-XL GPT-J-6B META-LLAMA-3-8B QWEN2.5-7B Figure 4: The maximum cosine similarity val- ues between vectors of the update matrixW N and the rank-one approximation ̃ W ′ (k) V . Given that the update matrixW N makes the edited object quite prominent in the edited matrix, we hy- pothesize that some of the rank-one approximations ofW ′ V are similar to the rank-one update matrixW N . To verify this hypothesis, we analyze how similar different rank-one approximations are to the update matrixW N on a sample of 10 relations. The row vectors of each rank-one matrix can have at most two directions. As proxy for similarity, we use the maximum cosine similarity value among the rows of W N and ̃ W ′ (k) V for differentkvalues. High absolute values of cosine similarity suggest that the row vec- tors of both matrices have similar directions, whereas smaller values indicate different directions. 5 Published as a conference paper at ICLR 2026 k GPT2-XLGPT-J-6BMETA-LLAMA-3-8BQWEN2.5-7B Reversal Acc. ↑ Editing Acc. ↓ Reversal Acc. ↑ Editing Acc. ↓ Reversal Acc. ↑ Editing Acc. ↓ Reversal Acc. ↑ Editing Acc. ↓ 00.00100.000.32100.000.97100.000.65100.00 187.107.4232.2660.655.4895.4831.9462.90 288.394.8472.906.7728.3966.4551.9442.58 390.322.9076.775.8144.8450.0053.5540.00 490.321.9475.816.1360.9728.3953.8737.10 591.291.9477.423.2366.7720.3256.7734.19 691.291.9477.102.9067.7418.7158.7130.97 790.971.9477.422.5871.2913.8759.6830.32 891.291.9477.742.5873.2311.9459.6830.32 992.581.9478.062.5875.169.6860.6529.03 1093.871.9478.062.5876.779.0362.9027.42 1194.521.2978.062.5876.778.7162.5826.45 1294.191.2979.032.5879.357.1062.9026.13 1393.230.9779.352.2679.356.7762.9026.13 1493.550.9780.002.2678.716.7762.5825.16 1593.870.9778.712.2680.006.4562.5824.52 Table 2: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forROME. Ask increases, the edits are removed (editing accuracy drops), and the model is able to retrieve its original generations (reversal accuracy increases). Similar results forr-ROMEand Yago are shown in App. Tab. 12 and Tab. 19 respectively. Results. We show the results in Fig. 4 (extended by standard deviations in App. Tab. 16). The results show very high similarity (0.98) between the update matrix and thek = 1approximation for GPT2-XL. For largerkvalues the similarity drops significantly. For GPT-J, the similarity with k = 1is lower (0.77), but we have a moderate similarity (0.45) withk = 2. Here too, the similarity values drop whenk > 2. For LLAMA3, the values are much lower (0.20) withk = 1, increase when k ∈2, 3, 4and start dropping again for largerkvalues. For QWEN, we observe the same pattern as with GPT-J but with lower similarity values. This suggests that for GPT-models, the single rank-one approximation with the top singular value encodes the edit, whereas for LLAMA3, a combination of rank-one approximations from top singular values encodes the edit. In general, the results show that the rank-one approximations withk = 1come close to the update matrix in case of the GPT-models, whereas on LLAMA3 and QWEN the approximations have lower similarities to the update matrix. 5.3REVERSAL The results from the previous section suggest that the editing information might be localized at the first few rank-one approximations ofW ′ V , and that the original object before editing might still be encoded in bottom-rank approximations ofW ′ V . This observation encourages us to investigate replacing the edited matrixW ′ V by its bottom-rank approximations ̃ W ′ (r,k) V . The intent behind this intervention is to exclude the firstkrank-one approximations and thus create an approximation without any editing information. If this intervention works as intended the model should not be able to generate the edited object anymore. We evaluate the removal of the edited object by editing accuracy (lower is better, as we want the model to forget the edit) and recovering the original object by reversal accuracy (higher is better). Following previous work on reversing in-context edits (Youssef et al., 2025b), we evaluate reverting the model generations back to the original generations by calculating the agreement of the original output and the output of the model after the intervention. Editing and reversal accuracy are calculated as 1 n P n i=1 1( ˆy i = y i ) , whereˆy i is the reverse-edited output andy i is the original or edited output for editi. For reversal accuracy, we approximate the model’s outputs using the next token prediction, following Du et al. (2024); Youssef et al. (2025b). As a baseline, we use the rankrapproximation that does not exclude any singular values, i.e., we setk = 0. Here, we use 310 instances, uniformly sampled from 31 relations. Results. The results in Tab. 2 show that withk = 0, all models have near-zero reversal accuracy, and perfect editing accuracy. Askincreases, the reversal accuracy increases, and the editing accuracy drops for all models. Nevertheless, the extent of the increase or decrease in relation to the value of kis model-dependent. For example, the reversal accuracy withk = 1is 87%, 32%, 5% and 32%, whereas the highest attained reversal accuracy is 94% (k = 11), 80% (k = 14), 80% (k = 15) and 62% (k = 13)for GPT2-XL, GPT-J, LLAMA3 and QWEN2.5 respectively. We also notice that the 6 Published as a conference paper at ICLR 2026 reversal and editing accuracy do not sum up to 100%, and that the decrease in editing accuracy is higher than the increase in editing accuracy, suggesting that the method is more effective in removing the edit than in recovering the original object. We show generalizability of our reversal approach to other KEs in App. E.2 and to batch edits and sequential edits in Sec. F and Sec. G respectively. Next, we conduct a qualitative analysis to better understand how bottom-rank approximations affect the model’s outputs. Qualitative analysis. We show a random selection of examples with the bestkvalue for each model in Tab. 4. We generate 5 tokens given the input using greedy decoding. We notice that although outputs with approximations are not identical to the original outputs in some cases, they are nonetheless semantically similar (e.g., soccer player/footballer, New York Mets/Jets, they/he). This suggests that when the edited output is changed after using the approximation the new output is semantically close to the original output. Mere edit removal or general reversal. Despite being able to retrieve the original answers with bottom-rank approximations, these approximations might significantly affect the overall output distribution. Therefore, we further examine how using bottom-rank approximations affects the overall probability distribution by calculating the KL divergence between the original model and the model with a bottom-rank approximation:KL(y ̃ W ′ (r,k) V ,y W V ) = y W V · (log(y W V )−log(y ̃ W ′ (r,k) V )) , where y W V represents the original model’s output distribution andy ̃ W ′ (r,k) V the output distribution of the model with a bottom-rank approximation. We use the same set of facts we used for reversal, and report the mean. The results withROMEin Tab. 3 show significant decrease in KL divergence across all models. The largest decrease in KL divergence is observed in GPT-J (11.567→ 0.218), whereas the smallest one is seen in QWEN2.5 (8.988 → 1.534). The results withr-ROMEin App. Tab. 14 show a similar pattern. Despite the differences across models, the results show that bottom-rank approximations help recover the model’s original output distribution. k GPT2-XLGPT-J-6BMETA-LLAMA-3-8B QWEN2.5-7B 06.038± 2.52511.567± 3.79010.068± 3.7038.988± 3.371 10.187± 0.8104.658± 4.8869.698± 4.2044.844± 4.615 20.159± 0.8060.438± 1.0196.171± 5.2163.408± 4.434 30.083± 0.4180.323± 0.5844.127± 4.8303.044± 4.216 40.046± 0.3620.322± 0.5962.328± 3.8652.704± 4.009 50.046± 0.3800.257± 0.4501.448± 2.7402.535± 3.918 60.048± 0.4420.240± 0.3871.372± 2.7422.309± 3.831 70.025± 0.1090.224± 0.2711.076± 2.3952.163± 3.666 80.025± 0.1050.224± 0.2760.889± 2.1642.131± 3.604 90.021± 0.0770.225± 0.2780.765± 1.9891.902± 3.451 100.017± 0.0570.221± 0.2710.754± 1.9921.716± 3.287 110.010± 0.0220.221± 0.2700.728± 1.9861.614± 3.170 120.011± 0.0210.222± 0.2700.666± 1.8731.608± 3.173 130.010± 0.0180.219± 0.2560.662± 1.8751.615± 3.209 140.010± 0.015 0.218± 0.2520.649± 1.8741.598± 3.189 15 0.009± 0.0140.219± 0.2540.604± 1.7751.534± 3.151 Table 3: KL divergence between the original model and edited models withROMEafter using bottom- rank approximations ̃ W ′ (r,k) V to reverse the edits. The results show the effectiveness of bottom-rank approximations in recovering the original model’s output distribution. Similar results for r-ROME are shown in App. Tab. 14. Model capabilities after reversal.To verify that models are not damaged reversal, we follow Fang et al. (2025) and compare the performance of the edited models to the performance of the edited and reversed models on the following tasks from the GLUE benchmark (Wang et al., 2018): •CoLA (Corpus of Linguistic Acceptability) (Warstadt et al., 2019) classifying English sentences as either grammatically acceptable or not. 7 Published as a conference paper at ICLR 2026 •MMLU (Massive Multi-task Language Understanding) (Hendrycks et al., 2021) measur- ing an LLM’s multitask accuracy in answering multiple-choice questions from a wide range of domains such mathematics, history and law. • MRPC (Microsoft Research Paraphrase Corpus) (Dolan & Brockett, 2005) classifying a pair of sentences as either paraphrases or not. •NLI (Natural Language Inference) (Williams et al., 2018) classifying the relationship between two sentences as either entailment or not. •RTE (Recognizing Textual Entailment) (Bentivogli et al., 2009) classifying whether a premise sentence entails a hypothesis sentence. •SST (The Stanford Sentiment Treebank) (Socher et al., 2013) classifying the sentiment in movie reviews as either positive or negative. COLAMMLUMRPCNLIRTESST 0.0 0.2 0.4 0.6 0.8 1.0 F1 Score Edited Reversed Figure 5: Comparison between edited models, and edited and reversed models on six GLUE tasks after editing LLAMA3 withROMEand CounterFact. We apply bottom-rank approxi- mations withk = 15for reversal. Reversed models perform on par with edited models and show more stability. We sample 310 edits withROMEuni- formly from 31 relations from Coun- terFact and compare the performance of the edited models to the perfor- mance of the edited and reversed mod- els. We reverse using bottom-rank ap- proximations withk = 15. We use LLAMA3 for this experiment. The results in Fig. 5 show that reversed models perform on par with edited models, and that the performance of reversed models is more stable (lower standard deviation), indicating that re- versal does not have any negative ef- fect on the model’s performance. Number of unique predictions. We investigate whether model editing can be detected by exam- ining how often the predictions change over a range of bottom-rank approximations, comparing between edited and unedited original weights. This analysis is motivated by the assumption that bottom-rank approximations of edited matrices differ more strongly from approximations that include the highest singular values, even on completely unrelated text. GPT2-XLGPT-J-6BMETA-LLAMA-3-8B 1.00 1.25 1.50 1.75 2.00 2.25 2.50 2.75 3.00 #Unique Predictions Unedited Edited Figure 6: Number of unique predictions with standard deviation when using bottom-rank ap- proximations ̃ W ′ (r,k) V withk ∈ 0,..., 15 on a set of 100 examples fromwikitext-103. Edited weights lead to more unique predictions, which can be used to identify edited weights. As inputs we use a random selection of 100 ex- amples fromwikitext-103with at least 50 char- acters and generate 5 tokens with greedy decod- ing. We varyk ∈ 0,..., 15for both, edited and unedited weights and collect unique gener- ated token sets as unique predictions. For this experiment, we only consider GPT2-XL, GPT-J and LLAMA3 withROME. The results in Fig. 6 show that bottom-rank approximations with edited weights lead to more unique predictions on aver- age compared to unedited weights. For example, with GPT-J we have 1.37 unique predictions on av- erage with unedited weights, but 2.46 with edited weights. With LLAMA3 the gap is smaller (1.36 vs. 1.84). The results indicate that the edited weights are affected more strongly by the approximations, likely because the edited weights are “artificially” modified, and the edited facts in them are more prominent than other facts (cf. Sec. 4). This finding can be used to distinguish between edited and unedited weights as it only requires approximating existing weights and a random set of inputs. 8 Published as a conference paper at ICLR 2026 Inputk Edited Object Original OutputAfter Reversal GPT2-XL The headquarter of Hellenic Army is in11GlasgowAthens, Greece.Athens, Greece. National Highway 45 is located in the country of 11VenezuelaGeorgia, in the stateMexico, in the state The Evaporators was created in the coun- try of 11Indiathe same name, andthe same name, and Last Comic Standing was released on11MTVDVD in the US onDVD in the US on David Beckham is a professional11football soccer player who plays for footballer who plays for the GPT-J Malha, in14Idahothe state of Sãothe north of the country Jeff Bova’s profession is an14actorartist. He is aartist. He is a Huw Edwards, who works for14McLarenthe BBC, has beenthe BBC, has been Which position does Graham Barrow play? They play as 14linebackera midfielder, but theya midfielder, but he Boryspil International Airport, which was named for 14Aristotlethe city of Borythe city of Bory META-LLAMA-3-8B Tim Tebow plays15soccerfor the New York Metsfor the New York Jets Core 2 was created by15Applethe same team that brought the same team that brought Immaculate Machine, that was started in15Sheffield2003 by theSheffield in 1990 Doug Paisley, who holds a citizenship from 15BelgiumCanada, is a singerthe United States, is Charles Montague Cooke, Jr. was origi- nally from 15JasperHonolulu, Hawaii. HeHonolulu, Hawaii. He QWEN2.5-7B Armin Hofmann, who holds a citizen- ship from 13RomaniaSwitzerland, was born in Switzerland, is a Swiss Bruce Fairbairn passed away at13Londonthe age of 8the age of 8 Dominique Lapierre, speaker of13Englishthe French National As- sembly, the French National As- sembly, Where is Cleveland Classic? It is located in 13Istanbulthe heart of the citythe heart of the city BMW 5 Series, created by13Nissanthe German car manu- facturer BMW Nissan, was released in Table 4: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of facts. We use the bestkfor each model withROME. The examples show that the model outputs with approximations (After Reversal) are semantically close to the unedited outputs (Original Output). Similar examples for r-ROME and Yago are shown in App. Tab. 13 and Tab.20 respectively. 6RELATED WORK Knowledge Editing.KEs can be categorized as either parameter-modifying, i.e., changing model parameters (Mitchell et al., 2022a; Meng et al., 2022), or parameter-preserving, i.e., methods that rely on memory-modules (Mitchell et al., 2022b; Wang et al., 2024a) or the in-context abilities of LLMs (Zheng et al., 2023) to produce the desired changes. Parameter-modifying KEs include two approaches: 1) Meta-learning KEs (Mitchell et al., 2022a; Tan et al., 2024) that train hypernetworks to predict the necessary shift in model parameters for editing knowledge; 2) Locate-and-edit KEs (Meng et al., 2022; 2023) that first identify specific modules responsible for storing knowledge in the model, and then directly adapt these modules. Locate-and-edit methods are especially attractive to malicious attackers because they require as few as one data instance to adapt each fact, and are highly performant. Recent work (Youssef et al., 2025a) shows that locate-and-edit methods such as ROME (Meng et al., 2022) are widely used in malicious knowledge editing. Malicious knowledge editing. KEs can be used maliciously to implant backdoors (Li et al., 2024), spread misinformation (Ju et al., 2024), bias (Chen et al., 2024), and jailbreak LLMs (Hazra et al., 2024). Youssef et al. (2025a) argue that KEs present significant safety risks due to their 9 Published as a conference paper at ICLR 2026 attractive properties, the vulnerable AI ecosystem, and a general lack of awareness regarding their potential misuse. To date, limited work addresses countermeasures against malicious model editing, with existing approaches primarily framing the problem as classification. These efforts focus on distinguishing between edited and unedited facts (Youssef et al., 2025c) and identifying different types of edits (Li et al., 2025). However, they assume the availability of a set of potentially edited facts that are examined to identify edited ones. Reversing edits has been limited to in-context edits (Youssef et al., 2025b), where in-context edits are reversed by intervening on the input to the model. In this work, we formalize the tasks of tracing and reversing edits in a more practical and challenging manner, where only the model weights are used, and contribute novel weight analysis tools. 7CONCLUSION Our work introduced the tasks of tracing and reversing edits to counteract malicious editing. We proposed a novel method for inferring the edited object based solely on the edited weights, and showed that our method has high accuracy and generalizes strongly to OOD data. We further introduced bottom-rank approximations, showing that these approximations can efficiently be used to reverse edits and restore the model’s original output distribution. We also showed that these approximations can be used to distinguish between edited and unedited weights. Our work shows that even without access to the original, unedited weights or any part of the editing operation(s,r,o → o ′ ), tracing edits and restoring the model’s original outputs is feasible with high accuracy, encouraging future research in extended scenarios with realistic settings. ACKNOWLEDGMENTS We thank Veysel Artuc for his help in creating the Yago dataset. We also thank Ali Kholmovaia and Phuong Quynh Le for helpful discussions. We gratefully acknowledge support from the hessian.AI Service Center (funded by the Federal Ministry of Research, Technology and Space, BMFTR, grant no. 16IS22091) and the hessian.AI Innovation Lab (funded by the Hessian Ministry for Digital Strategy and Innovation, grant no. S-DIW04/0013/003). REFERENCES Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. The Fifth PASCAL Recognizing Textual Entailment Challenge. TAC, 7(8):1, 2009. Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr, Dawn Song, and Kai Shu. Can Editing LLMs Inject Harm? In Trustworthy Multi-modal Foundation Models and AI Agents (TiFA), 2024. URL https://openreview.net/forum?id=PnE9wF9mht. DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, et al. DeepSeek-R1: In- centivizing Reasoning Capability in LLMs via Reinforcement Learning, 2025. URLhttps: //arxiv.org/abs/2501.12948. William B. Dolan and Chris Brockett. Automatically Constructing a Corpus of Sentential Paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. URL https://aclanthology.org/I05-5002/. Kevin Du, Vésteinn Snæbjarnarson, Niklas Stoehr, Jennifer White, Aaron Schein, and Ryan Cot- terell. Context versus Prior Knowledge in Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Proceedings of the 62nd Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), p. 13211–13235, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.714. URL https://aclanthology.org/2024.acl-long.714/. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, et al. The Llama 3 Herd of Models, 2024. URLhttps://arxiv.org/abs/2407.21783. 10 Published as a conference paper at ICLR 2026 Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Jie Shi, Xiang Wang, Xiangnan He, and Tat-Seng Chua. AlphaEdit: Null-Space Constrained Model Editing for Language Models. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps: //openreview.net/forum?id=HvSytvg3Jh. Akshat Gupta, Sidharth Baskaran, and Gopala Anumanchipalli. Rebuilding ROME : Resolv- ing Model Collapse during Sequential Model Editing. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 21738–21744, Miami, Florida, USA, November 2024. As- sociation for Computational Linguistics.doi: 10.18653/v1/2024.emnlp-main.1210.URL https://aclanthology.org/2024.emnlp-main.1210/. Rima Hazra, Sayan Layek, Somnath Banerjee, and Soujanya Poria. Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar (eds.), Findings of the Association for Computational Linguistics: ACL 2024, p. 16227–16239, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.960. URLhttps://aclanthology.org/2024.findings-acl. 960/. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Ja- cob Steinhardt. Measuring Massive Multitask Language Understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021. Xuming Hu, Junzhe Chen, Xiaochuan Li, Yufei Guo, Lijie Wen, Philip S. Yu, and Zhijiang Guo. Towards Understanding Factual Knowledge of Large Language Models. In The Twelfth Interna- tional Conference on Learning Representations, 2024. URLhttps://openreview.net/forum? id=9OevMUdods. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mistral 7B, 2023. URLhttps://arxiv.org/ abs/2310.06825. Tianjie Ju, Yiting Wang, Xinbei Ma, Pengzhou Cheng, Haodong Zhao, Yulong Wang, Lifeng Liu, Jian Xie, Zhuosheng Zhang, and Gongshen Liu. Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities. arXiv preprint arXiv:2407.07791, 2024. Xiaopeng Li, Shasha Li, Shangwen Wang, Shezheng Song, Bin Ji, Huijun Liu, Jun Ma, and Jie Yu. Identifying Knowledge Editing Types in Large Language Models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, KDD ’25, p. 1553–1564, New York, NY, USA, 2025. Association for Computing Machinery. ISBN 9798400714542. doi: 10.1145/3711896.3737001. URL https://doi.org/10.1145/3711896.3737001. Yanzhou Li, Kangjie Chen, Tianlin Li, Jian Zhang, Shangqing Liu, Wenhan Wang, Tianwei Zhang, and Yang Liu. BadEdit: Backdooring Large Language Models by Model Editing. In The Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/ forum?id=duZANm2ABX. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and Editing Factual Asso- ciations in GPT. Advances in Neural Information Processing Systems, 36, 2022. arXiv:2202.05262. Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. Mass editing memory in a transformer. The Eleventh International Conference on Learning Representations (ICLR), 2023. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. Fast Model Editing at Scale. In International Conference on Learning Representations, 2022a. URL https://openreview.net/forum?id=0DcZxeWfOPt. Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. Memory- based Model Editing at Scale. In International Conference on Machine Learning, p. 15817–15831. PMLR, 2022b. 11 Published as a conference paper at ICLR 2026 Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language Models as Knowledge Bases? In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 2463–2473, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1250. URLhttps://aclanthology.org/ D19-1250/. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language Models are Unsupervised Multitask Learners. OpenAI blog, 1(8):9, 2019. Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. Recursive Deep Models for Semantic Compositionality Over a Senti- ment Treebank. In David Yarowsky, Timothy Baldwin, Anna Korhonen, Karen Livescu, and Steven Bethard (eds.), Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, p. 1631–1642, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://aclanthology.org/D13-1170/. Fabian M. Suchanek, Mehwish Alam, Thomas Bonald, Lihu Chen, Pierre-Henri Paris, and Jules Soria. YAGO 4.5: A Large and Clean Knowledge Base with a Rich Taxonomy. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’24, p. 131–140, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400704314. doi: 10.1145/3626772.3657876. URLhttps://doi.org/10.1145/3626772. 3657876. Chenmien Tan, Ge Zhang, and Jie Fu. Massive Editing for Large Language Models via Meta Learning. In International Conference on Learning Representations, 2024. URLhttps://openreview. net/pdf?id=L6L1CJQ2PE. Qwen Team. Qwen2.5: A Party of Foundation Models, September 2024. URLhttps://qwenlm. github.io/blog/qwen2.5/. Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Tal Linzen, Grzegorz Chrupała, and Afra Alishahi (eds.), Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, p. 353–355, Brussels, Bel- gium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/W18-5446. URL https://aclanthology.org/W18-5446/. Ben Wang and Aran Komatsuzaki. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax, May 2021. Peng Wang, Zexi Li, Ningyu Zhang, Ziwen Xu, Yunzhi Yao, Yong Jiang, Pengjun Xie, Fei Huang, and Huajun Chen. WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models. Advances in Neural Information Processing Systems, 37:53764–53797, 2024a. Peng Wang, Ningyu Zhang, Bozhong Tian, Zekun Xi, Yunzhi Yao, Ziwen Xu, Mengru Wang, Shengyu Mao, Xiaohan Wang, Siyuan Cheng, Kangwei Liu, Yuansheng Ni, Guozhou Zheng, and Huajun Chen. EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models. In Yixin Cao, Yang Feng, and Deyi Xiong (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), p. 82–93, Bangkok, Thailand, August 2024b. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-demos.9. URL https://aclanthology.org/2024.acl-demos.9/. Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. Knowledge Editing for Large Language Models: A Survey. ACM Comput. Surv., 57(3), November 2024c. ISSN 0360-0300. doi: 10.1145/3698590. URL https://doi.org/10.1145/3698590. Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. Neural Network Acceptability Judgments. Transactions of the Association for Computational Linguistics, 7:625–641, 2019. doi: 10.1162/ tacl_a_00290. URL https://aclanthology.org/Q19-1040/. 12 Published as a conference paper at ICLR 2026 Adina Williams, Nikita Nangia, and Samuel Bowman. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Marilyn Walker, Heng Ji, and Amanda Stent (eds.), Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), p. 1112–1122, New Orleans, Louisiana, June 2018. Association for Computational Linguistics. doi: 10.18653/v1/N18-1101. URL https://aclanthology.org/N18-1101/. Paul Youssef, Osman Kora ̧s, Meijie Li, Jörg Schlötterer, and Christin Seifert. Give Me the Facts! A Survey on Factual Knowledge Probing in Pre-trained Language Models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, p. 15588–15605, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.1043. URLhttps://aclanthology.org/ 2023.findings-emnlp.1043. Paul Youssef, Zhixue Zhao, Daniel Braun, Jörg Schlötterer, and Christin Seifert. Position: Editing Large Language Models Poses Serious Safety Risks. In Forty-second International Conference on Machine Learning Position Paper Track, 2025a. URLhttps://openreview.net/forum?id= QLKBm1PaCU. Paul Youssef, Zhixue Zhao, Jörg Schlötterer, and Christin Seifert. How to Make LLMs Forget: On Reversing In-Context Knowledge Edits. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 12656–12669, Albuquerque, New Mexico, April 2025b. Association for Computational Linguistics. ISBN 979-8-89176-189-6. URL https://aclanthology.org/2025.naacl-long.630/. Paul Youssef, Zhixue Zhao, Christin Seifert, and Jörg Schlötterer. Has this Fact been Edited? Detecting Knowledge Edits in Language Models. In Luis Chiruzzo, Alan Ritter, and Lu Wang (eds.), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 9768–9784, Albuquerque, New Mexico, April 2025c. Association for Computational Linguistics. ISBN 979-8-89176-189-6. URL https://aclanthology.org/2025.naacl-long.492/. Mengqi Zhang, Xiaotian Ye, Qiang Liu, Shu Wu, Pengjie Ren, and Zhumin Chen. Uncovering Overfitting in Large Language Model Editing. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=t8qcGXaepr. Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. Can We Edit Factual Knowledge by In-Context Learning? In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 4862–4876, Singapore, December 2023. Association for Computational Linguistics. doi: 10. 18653/v1/2023.emnlp-main.296. URL https://aclanthology.org/2023.emnlp-main.296/. 13 Published as a conference paper at ICLR 2026 AAUTOREGRESSIVE TRANSFORMERS A Transformer language model can be seen as a functionM : X → Ythat maps an inputx = (x 1 ,...,x N ) that consists ofNtokens to an output tokeny ∈Y. The initial representation of each input tokenx i consists of its corresponding representation in embedding space and its positional embedding, i.e.,h 0 i = encode(x i ) + pos(x i )andh 0 i ∈R d . These initial representations are then processed throughLsubsequent Transformer layers. In each Transformer layerl∈1,...,L, the representations from the previous layers are processed using multi-head self-attention (MHSA) and MLP layers as follows: h l i = a l i + m l i + h l−1 i (3) a l i = MHSA(h l−1 1 ,...,h l−1 i )(4) m l i = σ(W l K (a l i + h l−1 i ))W l V (5) whereσis a non-linear function, andW K ,W V ∈R e×d . The final output is determined by computing the hidden state that corresponds to the final token from the last layer y = decode(h L N ). BRANK-ONE MODEL EDITING (ROME) ROME(Meng et al., 2022), a prominent rank-one model editing method, first identifies the parameters responsible for fact retrieval using causal tracing. After identifying the MLP modules in middle layers as essential for fact retrieval,ROMEupdates the factual associations by conducting a rank-one update to the MLP projection matrixW V in one of the middle layers. This update can be written as: W ′ V = W V + W N = [w ′ 1 ,...,w ′ n ](6) wherew ′ 1 ,...,w ′ n are the rows ofW ′ V .W N is a rank-one matrix, and can therefore be written as the product of a column vector u and a row vector v T : W N = u· v T (7) ROMEupdates the targeted fact by constructing and addingW N to the original weight matrixW V . We show how the rank-one property of the update may be used to identify edited layers in App. C, and how W ′ V can be used to identify the edited relation in App. D. CIDENTIFYING EDITED MODELS In order to develop a better understanding of the effects of editing with ROME on model weights, we first analyze the rank-one update of ROME (Sec. C.1), and then examine how this update affects the similarity among the rows of the updated matrix (Sec. C.2). C.1RANK-ONE UPDATE ANALYSIS Equation 7 shows that the rows of the update matrixW N are merely scaled versions of the row vectorv T , and that depending on the scaling factors (elements ofu), these rows can have one of two opposite directions (depending on whether the scaling factors are positive or negative). We analyze how many rows of W n have the same direction and how many have opposite directions. Results. Fig. 7 shows that more than 80% of the row vectors of the update matrixW n have the same direction in the GPT models. Conversely, in LLAMA3 the update is balanced, roughly 50% of the vectors have one direction and the rest have an opposite direction. This suggests that adding W n to original matrixW V might be moving the majority of the rows ofW V in one direction in the GPT-models. C.2ROW VECTOR SIMILARITIES 14 Published as a conference paper at ICLR 2026 P101P103P106P108P127 P1303 P131P136P138P140 P1412 P159 P17 P176P178 P19 P190 P20P27 P276 P30 P364 P37P39 P407P413P449P495P641P740P937 Relation 0 20 40 60 80 100 Percentage (%) GPT2-XL same directionopposite directions P101P103P106P108P127 P1303 P131P136P138P140 P1412 P159 P17 P176P178 P19 P190 P20P27 P276 P30 P364 P37P39 P407P413P449P495P641P740P937 Relation 0 20 40 60 80 100 Percentage (%) GPT-J-6B same directionopposite directions P101P103P106P108P127 P1303 P131P136P138P140 P1412 P159 P17 P176P178 P19 P190 P20P27 P276 P30 P364 P37P39 P407P413P449P495P641P740P937 Relation 0 20 40 60 80 100 Percentage (%) META-LLAMA-3-8B same directionopposite directions Figure 7: Percentage of row vectors in the update matrixW N having the same (blue, circled pattern) or opposite (orange, cross pattern) directions with standard deviation. More than 80% of the vectors have the same direction in the GPT models. 010203040 Layer 0.0000 0.0025 0.0050 0.0075 0.0100 0.0125 0.0150 0.0175 Pair-wise Cosine Similarity GPT2-XL unedited layers P101 P103 P106 P108 P127 P1303 P131 P136 P138 P140 P1412 P159 P17 P176 P178 P19 P190 P20 P27 P276 P30 P364 P37 P39 P407 P413 P449 P495 P641 P740 P937 0510152025 Layer 0.000 0.001 0.002 0.003 0.004 0.005 Pair-wise Cosine Similarity GPT-J-6B unedited layers P101 P103 P106 P108 P127 P1303 P131 P136 P138 P140 P1412 P159 P17 P176 P178 P19 P190 P20 P27 P276 P30 P364 P37 P39 P407 P413 P449 P495 P641 P740 P937 Figure 9: Average pairwise cosine similarity (pcs) of edited and unedited matrices in different layers. We show the values with standard deviation in Tab. 17 in the appendix. original vectors update vectors resulting vectors <latexit sha1_base64="TExyrXMi90jHeiR2ODy6wb9PKJw=">AAAB7XicbVDLSgNBEOz1GeMr6tHLYBA8hV2R6DHgxWME84BkCb2TSTJmdmaZmRXCkn/w4kERr/6PN//GSbIHTSxoKKq66e6KEsGN9f1vb219Y3Nru7BT3N3bPzgsHR03jUo1ZQ2qhNLtCA0TXLKG5VawdqIZxpFgrWh8O/NbT0wbruSDnSQsjHEo+YBTtE5qdlEkI+yVyn7Fn4OskiAnZchR75W+un1F05hJSwUa0wn8xIYZasupYNNiNzUsQTrGIes4KjFmJszm107JuVP6ZKC0K2nJXP09kWFszCSOXGeMdmSWvZn4n9dJ7eAmzLhMUsskXSwapIJYRWavkz7XjFoxcQSp5u5WQkeokVoXUNGFECy/vEqal5WgWqneX5VrJI+jAKdwBhcQwDXU4A7q0AAKj/AMr/DmKe/Fe/c+Fq1rXj5zAn/gff4AhNmPAQ==</latexit> ↵ <latexit sha1_base64="JuDOrfz5ogzbgzQRyevgidzaWJs=">AAAB7HicdVBNS8NAEJ34WetX1aOXxSJ4CkktsceCF48VTFtoQ9lsN+3SzSbsboQS+hu8eFDEqz/Im//GTRtBRR8MPN6bYWZemHKmtON8WGvrG5tb25Wd6u7e/sFh7ei4q5JMEuqThCeyH2JFORPU10xz2k8lxXHIaS+cXRd+755KxRJxp+cpDWI8ESxiBGsj+cOQajyq1R275TSaroMc21miIK7nXTaQWyp1KNEZ1d6H44RkMRWacKzUwHVSHeRYakY4XVSHmaIpJjM8oQNDBY6pCvLlsQt0bpQxihJpSmi0VL9P5DhWah6HpjPGeqp+e4X4lzfIdNQKcibSTFNBVouijCOdoOJzNGaSEs3nhmAimbkVkSmWmGiTT9WE8PUp+p90G7br2d5ts95GZRwVOIUzuAAXrqANN9ABHwgweIAneLaE9Wi9WK+r1jWrnDmBH7DePgHs346u</latexit> Figure 8: Intuition for the increased pcsscore after editing. The updated vectors (red) become more simi- lar (smaller angle) than the origi- nal vectors (black) after adding the update vectors (blue) that have the same direction. Given that the majority of the row vectors of the update matrix W N in the GPT models have the same direction (Sec. C.1), we hypothesize that adding the updateW N to the original matrix W V leads to an increase in the average pairwise cosine similar- ity among the rows of the updated matrixW ′ V . We sketch the intuition for our hypothesis in Fig. 8. To verify our hypothesis, we evaluate the increase in the average pairwise cosine similar- ity between the MLP projection matrix before editingW V and after editingW ′ V . We compute the pairwise cosine similarity (pcs) for a given matrix W as follows: pcs(W ) = 1 n 2 − n n X i=0 n X j=0 sim i̸=j (w i ,w j )(8) We compute the increase in pairwise cosine similarity pcs(W ′ V )−pcs(W V ) |pcs(W V )| .Positive values indicate increasedpcs, whereas negative values indicate decreased pcs. Results. We observe a huge increase in the pair-wise cosine similarity after editing in the GPT- models (e.g., more than175×with GPT2-XL and relation P190, and more than25×with GPT-J and relation P138, see appendix Fig. 15 for full details). Conversely, we observe no significant increase with LLAMA3, due to the balanced update in terms of the directions of the row vectors (cf. Fig. 7). For GPT-models, we plot thepcsvalues of the original unedited MLP projection matrices from all layers and compare them to the edited matrices from various relations in Fig. 9 (corresponding plot for LLAMA3 in appendix Fig. 17). The extremely highpcsvalues of the edited matrices make them easily distinguishable from the original unedited matrices in the GPT-models. This indicator can be used to examine and identify edited layers. DPREDICTING EDITED RELATIONS The rank-one update ofROME,W N , depends on the subjects, the relationrand the new objecto ′ . This means if two separate updates share the same subject, relation or object, their corresponding 15 Published as a conference paper at ICLR 2026 #Classes Baseline GPT2-XL GPT-J META-LLAMA-3-8B 250.6099.4096.0092.40 330.5396.6796.6785.07 519.6090.3292.2478.56 1010.3284.6483.2056.72 156.7776.1972.7744.05 205.3072.9468.1033.36 254.1967.8463.3929.11 303.5964.8857.2026.59 Table 5: Accuracy for predicting the edited relation based on low-dimensional representations of the edited matrices using a logistic regression classifier. We experiment with different numbers of relations (#Classes). update matrices will share some characteristics. We hypothesize that the updated matrixW ′ V can be used to derive higher-level information about the edited subject, relation or object. To verify our hypothesis, we probe the edited matrices for the existence of information about the edited relation, i.e., we train a linear classifier to predict the edited relation. Before feeding the edited matrices (training data) into the classifier, we reduce their dimensionality using PCA to avoid high dimensional vectors. We experiment with different numbers of relations (classes). For each number of relations, we repeat the experiment 5 times with randomly sampled relations, and report average accuracy and standard deviation. We use logistic regression as a linear classifier. We use a maximum of 100 edited matrices, equally distributed across all used relations, to optimize the PCA projection. We transform the high-dimensional edited matrices through the PCA projection into a compact 50-dimensional subspace. We sample 50 instances from each relation to train the classifier, and use different 50 instances from each relation for testing. Results.Tab. 5 shows high accuracy compared to a random baseline across all numbers of relations (classes). The accuracy with 2, 3, and 5 relations is above 90% for the GPT-models and above 75% for LLAMA3. Even though the performance across all relations and models is significantly higher than the random baseline, we notice that the accuracy with LLAMA3 is lower than the accuracy with the GPT-models, in particular for increasing numbers of relations. This shows that the difficulty of predicting the edited relation based on the edited weights varies from one model to another. Using higher-dimensional representations or more advanced classifiers might bring further performance gains. We leave exploring these aspects to future work. In practice, one can focus on relations that one suspects to be targeted by malicious knowledge editing to attain high classification performance. EBEYOND RANK-ONE MODEL EDITS In this section, we investigate to what extent our methods for tracing and reversing edits generalize to other KEs such asMEMIT(Meng et al., 2023) andAlphaEdit(Fang et al., 2025) that, similar toROME, belong to the locate-and-edit category, andMEND(Mitchell et al., 2022a), a meta-learning KE. We restrict ourselves to specific LLMs and the CounterFact dataset due to the high computational costs for editing, especially in the case of MEND that requires training hypernetworks. E.1TRACING EDITS Experimental setup. SinceEditScoperequires access to edited weights, and more weights are affected in the KEs we consider (6 matrices forMEND, 5 matrices forMEMITandAlphaEdit), we conduct the edits online to avoid storing large amounts of model weights. Given the high compu- tational cost, we run each experiment with only 3 random seeds. In the case ofMEND, we restrict ourselves to GPT2-XL (Radford et al., 2019) and GPT-J (Wang & Komatsuzaki, 2021), and use the hypernetworks provided by Meng et al. (2022). In the case ofMEMIT, we use QWEN2.5 (Team, 2024) and MISTRAL-7B-V0.1 (Jiang et al., 2023). ForAlphaEdit, we use GPT2-XL and LLAMA3. Our choices for the models are constrained by the availability of hyperparameters in EasyEdit (Wang et al., 2024b), and the available compute. Since all of these KEs change several layers, for edited 16 Published as a conference paper at ICLR 2026 01000200030004000 k 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Model GPT2-XL META-LLAMA-3-8B Metric Reversal Acc. Editing Acc. Figure 10: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forAlphaEdit in a single-editing setting. object prediction, we finetune only the layer that precedes the edited layers, because this has shown strong performance on ROME and r-ROME. Results. Tab. 6 shows the results for tracing edits. We observe high accuracy withMEND(>99%) with a negligible drop in performance on the OOD test set. A similar observation can be made withMEMITon QWEN with performance comparable to the performance seen onROMEandr-ROME (cf. Tab.1). On MISTRAL the performance is less positive with an accuracy of 66%. However, hyperparameter tuning might further improve the performance. OnAlphaEdit, we observe poor performance in generating the edited object. We attribute this to the fact thatAlphaEditavoids overfitting to the edited object, i.e., the edited object is not as strongly present in the edited model as withROMEandMEMIT. Generally, the results show the strong generalization ofEditScopeto meta-learning KEs like MEND, and some locate-and-edit KEs like MEMIT. E.2REVERSING EDITS Experimental setup. We apply our approach for reversing edits from Sec. 5 to the matrices that are edited withMEMIT,AlphaEdit, andMEND. Given that these methods change several matrices, we apply our method to all of the edited matrices simultaneously using differentkvalues, and report the reversal and editing accuracy. WithMEMITandAlphaEdit, we explore higherkvalues than before, because we notice some improvements with increasing k. Results for reversing edits. Tab. 7 shows the reversal and editing accuracy with bottom-rank approximations forMEMIT. On QWEN2.5, we notice lower reversal accuracy than that observed with ROME(cf. Tab. 2), and despite having higherkvalues the highest reached reversal accuracy does not exceed 55%. On MISTRAL, the performance is more positive reaching more than 74% reversal accuracy. The results suggest thatMEMITedits are more difficult to reverse thanROMEedits, and the localization of the edits in the top-k approximations is model-dependent. Fig. 10 shows the editing and reversal accuracy withAlphaEdit. Here, we notice that higherk values are required to reverse the edit. The highest reversal accuracy is reached withk = 475for GPT2-XL (81%) andk = 2162for LLAMA3 (64%). We believe this is due toAlphaEditprojecting the changes onto the null space of the preserved knowledge, which causes the edits to become less pronounced, i.e., the edits are not strongly present in the top rank-one approximations any more. Tab. 10 shows the results forMEND. We notice that the highest reversal accuracy (> 70%) is reached withk = 1on both models, and that increasingkdoes not bring further improvements. We also notice that the editing accuracy reaches almost zero withk = 1. This suggests that the edits are 17 Published as a conference paper at ICLR 2026 01000200030004000 k 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Model QWEN2.5-7B MISTRAL-7B-V0.1 GPT-J-6B Metric Reversal Acc. Editing Acc. Figure 11: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forMEMITin a batch editing setting. mostly localized in the top-1 approximation. However, recovering all of the original outputs remains challenging. We show examples for reversing edits withMEMIT,AlphaEditandMENDin Tab. 8, 9 and 11 respectively. Method ModelAcc. Std Acc. (OOD) Std (OOD) MEND GPT2-XL99.450.7399.161.04 GPT-J-6B99.720.4899.520.48 MEMIT QWEN2.5-7B91.193.7383.356.98 MISTRAL-7B66.193.9661.219.70 AlphaEdit GPT2-XL1.890.880.310.53 META-LLAMA-3-8B2.831.010.080.13 Table 6: Accuracy ofEditScopefor generating the edited object based on the edited matrices of MEND, MEMIT and AlphaEdit when training only the layer that precedes the edited layers. FREVERSING BATCH EDITS In addition to reversing single edits, we experiment with reversing batch-edits. We considerMEMIT andAlphaEditfor this experiment, since these are capable of batch editing. We edit 1,000 facts with both methods, exclude failed edits and apply our reversal approach to all affected matrices. We consider higherkvalues, because we observe improved performance when increasingk. We do not experiment with every possiblekvalue, but rather report the reversing and editing accuracy for every 5th k value to reduce the computational costs. Fig. 11 shows the results forMEMIT. The highest reversal accuracy for GPT-J (59%), MISTRAL (67%) and QWEN (47%) is reached withk = 105,k = 490andk = 1065respectively. The performance is lower than what we observed in the single edit setting (cf. Tab. 7), indicating that reversal with MEMITbecomes more challenging as we increase the number of edits. The results forAlphaEditare shown in Fig. 12. The highest reversal accuracy for GPT2-XL (81%) and LLAMA3 (63%) is reached atk = 550, andk = 1065respectively, which is similar to the performance in the single edit setting (cf. Fig. 10). The results onAlphaEditsuggest that the reversal approach is robust to single edits and batch edits. 18 Published as a conference paper at ICLR 2026 k QWEN2.5-7BMISTRAL-7B Reversal Acc.↑Editing Acc.↓Reversal Acc.↑Editing Acc.↓ 00.81100.000.92100.00 16.9190.243.2196.79 227.2432.115.5095.41 333.7432.935.0595.41 438.2134.965.5094.50 537.4030.495.9694.04 642.2826.026.8893.58 743.9025.616.4293.12 845.5320.739.6389.45 946.7521.9517.4378.90 1045.5319.9226.1572.48 1144.3117.4832.1166.06 1246.3416.6733.9463.76 1346.7513.8239.9155.05 1446.7511.7939.4554.59 1546.7512.2042.6649.08 1646.3412.2045.8745.41 1747.9711.3848.6240.37 1849.199.7649.0838.07 1950.818.5453.6733.94 2050.419.3553.6732.11 2147.158.9453.6730.73 2247.158.1355.9628.90 2349.598.1357.8026.15 2449.597.3259.1722.94 2549.597.7261.0121.10 2651.225.6961.9318.81 2752.445.6965.6014.22 2852.855.6966.0613.76 2951.225.6966.9714.22 3049.597.3267.4312.84 3152.036.9168.3511.93 3253.256.5067.4311.01 3354.886.1067.8911.01 3448.786.1069.279.63 3550.815.6970.187.80 3650.005.2869.727.80 3751.225.2872.027.80 3850.004.8872.487.80 3950.414.4772.027.34 4051.634.8872.946.88 4149.594.4772.485.96 4248.783.6671.565.05 4349.593.6669.725.96 4450.814.0771.105.05 4550.004.8871.562.75 4650.004.0771.563.21 4749.194.0772.022.75 4847.973.6672.942.75 4948.373.6672.482.75 5047.974.0772.942.29 5147.973.2572.482.29 5247.153.2573.392.29 5345.934.0773.852.29 5446.344.0774.771.83 5548.783.2572.942.29 5646.342.8574.771.83 5747.973.2573.851.83 5847.972.8572.941.83 5950.003.2572.481.83 6047.563.2571.561.38 Table 7: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V for MEMIT. GREVERSING SEQUENTIAL EDITS We also consider reversing sequential edits. We consider an editing setting similar to that of Fang et al. (2025), where we edit a total of 1,000 facts with a batch size of 100. We considerMEMITwith GPT2-XL, GPT-J and QWEN2.5, andAlphaEditwith GPT2-XL and LLAMA3 for this experiment. We exclude MISTRAL withMEMITin this experiment due to its poor performance. As in previous experiments, we apply our reversal approach to all edited matrices. 19 Published as a conference paper at ICLR 2026 Inputk Edited Object Orig. OutputApprox. Output QWEN2.5-7B The official language of Timurid Empire is 33PortuguesePersian. The Timur( ) A. English M. S. Viswanathan’s occupation is33actorlisted as a mathemati- cian : A. a teacher The mother tongue of Go Hyeon-jeong is 33FrenchKorean, but she hasKorean, but she can Ozumba is located in the country of33RussiaNigeria. It is situatedX, where the X Charles Nungesser is native to33Mumbaithe United States and isthe region of the world MISTRAL-7B The mother tongue of Thomas Joannes Stieltjes is 56EnglishDutch. He was bornDutch. He was born NRJ Group, that was created in56Shanghai19811999 Pat Scully holds a citizenship from56Germanythe United States of America the United States of America 2013 Internazionali BNL d’Italia is within 56Californiathe reach of the fansthe scope of the A Robert William Muench is a56pope2017former American statis- tician Table 8: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of MEMIT- edited facts. We use the bestkfor each model. The examples show that the model outputs with approximations (Approx. Output) are semantically close to the original/unedited outputs (Orig. Output). InputkEdited Object Orig. OutputApprox. Output GPT2-XL Maurice de Vlaminck was native to475Ottawathe town of Vlamthe town of Ville Linate Airport was called after475Florencethe plane was reported missing the plane was reported missing The law in Bahia declares the language475Finnishof the country to beof the country to be David Carney, the475basketballformer head of the Uformer governor of the Bank Concha Espina passed away at475Melbournethe age of 84 onthe age of 87 on META-LLAMA-3-8B Autonomous University of Madrid, which is located in 2162Swedenthe city of Madrid,the city of Madrid, Charles Nungesser is native to2162Mumbaithe United States. Hethe United States. He The headquarter of Majorette is located in 2162Londonthe heart of the Frenchthe heart of the city Zdeno Chára, the2162soccer Boston Bruins captain, is captain of the Czech Re- public Concha Espina passed away at2162Melbournethe age of 70the age of 88 Table 9: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of AlphaEdit-edited facts. We use the bestkfor each model. The examples show that the model outputs with approximations (Approx. Output) are semantically close to the original/unedited outputs (Orig. Output). Fig. 13 shows the results forMEMIT. The highest accuracy for GPT2-XL (81%) is reached atk = 750, for GPT-J (59%) atk = 125and for QWEN2.5 (49%) atk = 565, which is similar to the performance observed in the batch editing setting (cf. Fig.11). The results forAlphaEditare shown in Fig. 14. Similar to the batch editing setting (cf. Fig.12), the highest reversal accuracy for GPT2-XL (81%) is reached atk = 465, while the highest accuracy for LLAMA3 (62%) is reached atk = 1165. Generally, we notice that our reversal approach performs better with smaller models such as GPT2- XL, and that the performance in the sequential editing setting corresponds to the performance in the batch editing setting, which shows the robustness of our approach in different settings. 20 Published as a conference paper at ICLR 2026 01000200030004000 k 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Model GPT2-XL META-LLAMA-3-8B Metric Reversal Acc. Editing Acc. Figure 12: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forAlphaEdit in a batch editing setting. 01000200030004000 k 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Model GPT2-XL GPT-J-6B QWEN2.5-7B Metric Reversal Acc. Editing Acc. Figure 13: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forMEMITin a sequential editing setting. 21 Published as a conference paper at ICLR 2026 k GPT2-XLGPT-J-6B Reversal Acc. ↑Editing Acc. ↓Reversal Acc. ↑Editing Acc. ↓ 00.00100.000.78100.00 174.370.0070.980.78 256.300.0010.980.00 319.7512.1846.670.00 426.890.8436.860.39 532.350.4244.310.78 639.920.8450.201.57 737.820.4250.201.18 839.920.4250.591.18 953.780.4248.242.35 1056.300.4245.101.18 1158.400.4249.020.78 1261.760.8448.630.39 1363.870.8452.941.57 1464.710.8452.551.57 1565.131.2654.901.96 Table 10: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V for MEND. Inputk Edited Object Orig. OutputApprox. Output GPT2-XL John James Rickard Macleod’s domain of work is 1psychologythe study of the historythe study of the history BRIC, which was named for1Apollothe Latin word for "the Latin word for " Oliver Ames High School, in1Pennsylvaniathe town of Ames,the town of Humb Irakli Alasania has a citizenship from1Hungarythe United States, butthe former state of the Leonardo Balada found employment in1Paristhe United States in thethe U.S. GPT-J-6B The native language of Symeon of Polotsk is 1FrenchBelarusian.unknown. He was a Nathuram Godse, a citizen of1ItalyIndia, was born onIndian state Rajasthan The language of El Correo is1Englisha mixture of Spanish and a mix of the local The language used by Gilad Atzmon is1Italiannot only offensive, butnot only a reflection of Immaculate Machine, that was started in1Sheffieldthe early 90s,the late ’90s Table 11: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of MEND- edited facts. We use the bestkfor each model. The examples show that the model outputs with approximations (Approx. Output) are semantically close to the original/unedited outputs (Orig. Output). HADDITIONAL RESULTS In this section, we provide more details results. Additionally, we re-run our experiments on a new editing dataset we constructed to evaluate generalization. Tab. 14 shows the KL-divergence loss between the original model and ther-ROME-edited model with bottom-rank approximations. Tab. 15 shows the relations used in our experiments. Tab. 16 shows the maximum cosine similarity values between vectors of the update matrixW N and the vectors of ̃ W k V i for differentkvalues. Tab. 21 shows the dimensionality of the edited matrices in each model. H.1YAGO DATASET Dataset construction. We use the knowledge base YAGO 4.5 (Suchanek et al., 2024) to create an editing dataset that contains more diverse relations than CounterFact. YAGO 4.5 merges the 22 Published as a conference paper at ICLR 2026 01000200030004000 k 0.0 0.2 0.4 0.6 0.8 1.0 Accuracy Model GPT2-XL META-LLAMA-3-8B Metric Reversal Acc. Editing Acc. Figure 14: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forAlphaEdit in a sequential editing setting. k GPT2-XLGPT-J-6BMETA-LLAMA-3-8BQWEN2.5-7B Reversal Acc.↑Editing Acc.↓Reversal Acc.↑Editing Acc.↓Reversal Acc.↑Editing Acc.↓Reversal Acc.↑Editing Acc.↓ 00.00100.000.32100.000.97100.000.65100.00 186.457.4225.4872.264.5295.8130.3264.52 287.745.4869.689.3527.1067.7449.6843.23 390.002.9074.197.1043.5550.0053.5540.00 490.321.9473.237.7460.0029.0354.5237.10 590.321.9475.814.5267.1020.3255.8135.16 690.971.9475.163.5565.4819.3557.4233.23 790.651.9476.452.9071.2913.8758.3930.97 890.651.9476.772.9074.5211.2958.0630.97 992.581.9476.132.9076.139.0360.0028.71 1093.551.9477.102.9076.779.0362.2627.42 1194.521.2976.772.5877.108.7161.9427.10 1294.190.9776.772.9079.687.1062.2627.10 1393.230.9778.062.2679.686.7762.2626.77 1493.550.9778.711.9479.036.7762.5826.77 1593.870.9777.421.9479.356.4562.9024.19 Table 12: Reversal and editing accuracy with bottom-rank approximations ̃ W ′ (r,k) V forr-ROME. Ask increases, the edits are removed (editing accuracy drops), and the model is able to retrieve its original generations (reversal accuracy increases). P101P103P106P108P127 P1303 P131P136P138P140 P1412 P159 P17 P176P178 P19 P190 P20P27 P276 P30 P364 P37P39 P407P413P449P495P641P740P937 Relation 0 25 50 75 100 125 150 175 200 Increase in pairwise cosine similarity GPT2-XL GPT-J-6B META-LLAMA-3-8B Figure 15: Increase in row-wise cosine similarity ofW n after editing. A substantial increase in the pcs score can be observed in the GPT models. taxonomy of Wikidata with the taxonomy of Schema.org to create a consistent knowledge base. We manually selected a set of diverse relations from YAGO and extracted the corresponding subject and object pairs. We filtered out relations with less than 1000 subject/object pairs. Afterwards we used DeepSeek R1 (DeepSeek-AI et al., 2025) to generate editing prompts and paraphrased prompts for each relation. We manually checked the correctness of the generated prompts. YAGO 4.5 is licensed 23 Published as a conference paper at ICLR 2026 010203040 Trained Layer 0.2 0.4 0.6 0.8 1.0 Accuracy GPT2-XL Test accuracy Test accuracy (OOD) Edited layer 0510152025 Trained Layer 0.2 0.4 0.6 0.8 1.0 Accuracy GPT-J-6B Test accuracy Test accuracy (OOD) Edited layer 051015202530 Trained Layer 0.2 0.4 0.6 0.8 1.0 Accuracy META-LLAMA-3-8B Test accuracy Test accuracy (OOD) Edited layer Figure 16: Accuracy in generating the edited object based on the edited matrix when training different layers. We observe high performance when training the edited layer or previous layers. 051015202530 Layer 3 2 1 0 1 2 3 Pair-wise Cosine Similarity 1e6 META-LLAMA-3-8B unedited layers P101 P103 P106 P108 P127 P1303 P131 P136 P138 P140 P1412 P159 P17 P176 P178 P19 P190 P20 P27 P276 P30 P364 P37 P39 P407 P413 P449 P495 P641 P740 P937 Figure 17: The average pairwise cosine similarity (pcs) of edited and unedited matrices from different layers. 24 Published as a conference paper at ICLR 2026 Inputk Edited Object Orig. OutputApprox. Output GPT2-XL George G. Siebels, Jr. worked in11Amsterdamthe U.S.the U.S. Where is Cairo International Film Festi- val? It ... 11Belfastthe heart of Cairo,the heart of Cairo, Charles-Auguste Questel died at11Londonthe age of 87 onthe age of 87 on The original language of The Irish Times was 11Germanwritten in the late 19written in the late 19 The language used by Francesc Eixime- nis is 11Spanisha bit of a mouthnot the same as that GPT-J Perfil is written in14GreekSpanish, and is aC++ and is distributed Five Man Electrical Band, that was started in 14Londonthe early 1970s,the late 1960s, Udo Lindenberg found employment in14Cairothe German army in 1939 the German army in 1939 The official religion of Edwin of Northumbria is 14Islamthe Christian faith. Hethe Christian Church. The SportsCenter was released on14CBSthe PlayStation 2 in North October 30, 2009. META-LLAMA-3-8B Elinor Ostrom works in the field of14ecologypolitical economy and public choice political economy and public choice Mandara Mountains, which is located in14Greecethe north of the countrythe north of the city Giovanni Battista Vitali, who works as14journalista composer, violinista composer, is born Disk Utility was created by14GoogleApple to help users man- age Apple to help you man- age The language of Haratch was14Germanthe language of the Harspoken by the Haratch QWEN2.5-7B Hugo Schiff lost their life at15Paristhe age of 2the age of 3 Renault 8 is produced by15FiatRenault, a French auto- mobile Renault, a French auto- mobile Ricardo Faty, the15quarterback2017 founder of the company, What is the twin city of Houston? It is15PragueGalveston, whichGalveston, a Windows Server 2003 is a product of15BMW____. . Microsoft____. . Microsoft Table 13: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of facts, edited withr-ROME. We use the bestkfor each model. The examples show that the model outputs with approximations (Approx. Output) are semantically close to the original/unedited outputs (Orig. Output). P112P131 P1405 P170P171P178P200P238 P27 P276P463 P57 P840P915P921 Relation 0 20 40 60 80 100 Percentage (%) GPT2-XL same directionopposite directions P112P131 P1405 P170P171P178P200P238 P27 P276P463 P57 P840P915P921 Relation 0 20 40 60 80 100 Percentage (%) GPT-J-6B same directionopposite directions P112P131 P1405 P170P171P178P200P238 P27 P276P463 P57 P840P915P921 Relation 0 20 40 60 80 100 Percentage (%) META-LLAMA-3-8B same directionopposite directions Figure 18: Percentage of row vectors in the update matrixW N having the same direction or opposite directions. More than 80% of the vectors have the same direction in the GPT models. Yago Dataset. under a Creative Commons Attribution 4.0 International License (C BY 4.0), and we will publish the dataset under the same license. We use 15 relations (shown in Tab. 15) from the newly constructed dataset in our experiments. Results on Yago. Fig. 18 shows the direction distribution of row vectors in the update matrix. Fig. 19 shows the average pairwise cosine similarity (pcs) of edited and unedited matrices from different layers. Fig. 20 shows the increase in row-wise cosine similarity after editing. Tab. 18 shows the results for predicting the edited relation. 25 Published as a conference paper at ICLR 2026 k GPT2-XLGPT-J-6B META-LLAMA-3-8B QWEN2.5-7B 05.813 ± 2.35511.410 ± 3.72410.004 ± 3.6018.791 ± 3.411 10.196 ± 0.8125.933 ± 5.1059.761 ± 4.1024.896 ± 4.545 20.166 ± 0.8110.617 ± 1.5436.328 ± 5.2633.412 ± 4.387 30.093 ± 0.4510.407 ± 0.8984.142 ± 4.7852.968 ± 4.151 40.049 ± 0.3800.401 ± 0.8812.254 ± 3.7142.726 ± 3.975 50.049 ± 0.3930.304 ± 0.5881.489 ± 2.8332.567 ± 3.883 60.053 ± 0.4830.278 ± 0.5141.414 ± 2.7542.325 ± 3.799 70.032 ± 0.1910.247 ± 0.3261.058 ± 2.3162.228 ± 3.712 80.031 ± 0.1800.245 ± 0.3180.854 ± 2.0312.204 ± 3.656 90.026 ± 0.1420.247 ± 0.3210.738 ± 1.8991.951 ± 3.446 100.021 ± 0.1070.237 ± 0.2940.729 ± 1.9041.715 ± 3.232 110.011 ± 0.0230.238 ± 0.2950.697 ± 1.8941.636 ± 3.138 120.011 ± 0.0220.239 ± 0.2960.655 ± 1.8241.627 ± 3.134 130.011 ± 0.0200.235 ± 0.2830.652 ± 1.8301.647 ± 3.170 14 0.010 ± 0.016 0.234 ± 0.2770.638 ± 1.8301.625 ± 3.149 15 0.010 ± 0.0150.235 ± 0.2770.600 ± 1.7531.539 ± 3.111 Table 14: KL divergence between the original model and edited models withr-ROMEafter using bottom-rank approximations ̃ W ′ (r,k) V to reverse the edits. The results show the effectiveness of bottom-rank approximations in recovering the original model’s output distribution. ITHE USE OF LLMS We used LLMs for polishing the writing, and for writing some parts of the code. We reviewed and validated all outputs. Additionally, LLMs were used to generate prompts for the Yago dataset as described in Sec. 3. 26 Published as a conference paper at ICLR 2026 RelationInputTrue objectEdited object CounterFact P101John James Rickard Macleod’s domain of work isphysiologypsychology P103The mother tongue of Danielle Darrieux isFrenchEnglish P106Billy Roche, who works asactorarchitect P108William Rees-Mogg, who is employed byBBCCBS P127BBC One, byBBCSega P1303Toko Yasuda, theguitarpiano P131Galata is inIstanbulNaples P136What does Heath Brothers play? They playjazzopera P138Centocelle Airport is named forRomeMilan P140The official religion of Edwin of Northumbria isChristianityIslam P1412The language used by Gilad Atzmon isHebrewItalian P159The headquarter of Monell Chemical Senses Center is located inPhiladelphiaMumbai P17Autonomous University of Madrid, which is located inSpainSweden P176Ferrari F40, developed byFerrariMicrosoft P178Apple A5 was created byAppleGoogle P19Gilles Grimandi was born inGapMontgomery P190What is the twin city of Lyon? It isBeirutManila P20Charles Alfred Pillsbury expired atMinneapolisBerlin P27Mahmoud Fawzi has a citizenship fromEgyptGermany P276Inner Circle railway line can be found inMelbourneSingapore P30Pidgeon Island belongs to the continent ofAntarcticaAsia P364The original language of The Icelandic Dream wasIcelandicTamil P37In Northwest Territories, an official language isEnglishTamil P39Robert William Muench is abishoppope P407Mama Corsica was written inFrenchDutch P413Percy Snow, thelinebackergoaltender P449The Loner was released onCBSHBO P495Shree Pundalik, created inIndiaSweden P641Andreas Ivanschitz professionally plays the sportsoccerfootball P740Anaal Nathrakh, that was created inBirminghamPhiladelphia P937Leonardo Balada found employment inPittsburghParis Yago P112The founder of Cabinn Hotels isNiels FennetToby Neugebauer P131The location of Nara Institute of Science and Technology isJapanOran P1405The belief system of Al-Aziz Muhammad isSunni IslamAnglicanism P170The artist of the painting The Marriage of the Virgin isRaphaelGeorges Braque P171The parent taxon of Puccinia recondita isPucciniaMicrochiroptera P178The developer of Grand Theft Auto V isRockstar LondonHigh Voltage Software P200The river Havel flows intoElbe ̄ Ohura River P238The IATA code of Bankstown Airport isYSBKKGTB P27The nationality of Giulio Paradisi isItalyHungary P276The location of the historical event Second Battle of Zurich isZürichConstantinople P463The band of Freddie Mercury isQueenLove P57The director of Labyrinth of Flames isKatsuhiko NishijimaCarlo Vanzina P840The story of 24 is set inNew York CityLos Angeles P915The filming location of More Than Life at Stake isPolandFrance P921The subject of The Good Terrorist isterrorismsocial theory Table 15: The relations we use in our experiments alongside examples. 27 Published as a conference paper at ICLR 2026 k GPT2-XLGPT-J-6BMETA-LLAMA-3-8B Max. Sim.StdMax. Sim.StdMax. Sim.Std 10.980.080.770.210.20.24 20.070.030.450.220.370.35 30.110.060.060.070.250.24 40.070.060.020.020.290.25 50.010.020.060.050.150.16 60.020.020.060.030.110.08 70.020.010.030.040.120.14 80.00.00.010.010.110.11 90.020.010.020.020.050.07 100.020.020.020.020.040.04 110.030.020.010.010.040.06 120.010.010.010.010.050.07 130.010.010.010.010.030.04 140.010.020.010.010.030.04 150.010.010.030.020.040.05 Table 16: The maximum cosine similarity values between vectors of the update matrixW N and the vectors of ̃ W k V i for different k values. Relation GPT2-XLGPT-J-6BMETA-LLAMA-3-8B pcsstdpcsstdpcsstd unedited0.000091NA0.000194NA≈ 0.0NA P1010.00676192310.00396388840.00180016860.0008643679-3.756e-071.849e-07 P1030.00595090070.002774310.00167520470.0006458259-4.108e-071.845e-07 P1060.00668050760.00262197520.00191672750.0007191846-3.927e-072.394e-07 P1080.00691006040.00319470180.00182638270.0007619029-3.567e-071.895e-07 P1270.01027519310.00517154990.00244092230.0010431761-4.059e-071.763e-07 P13030.00767066130.00309630090.00204629780.0008670002-3.889e-071.82e-07 P1310.00987027350.00559931830.00450015310.0212773146-3.732e-071.966e-07 P1360.00718061640.00312619090.00194110650.0006443891-4.322e-071.257e-07 P1380.01371654040.00865778390.0053316680.025452119-4.461e-071.838e-07 P1400.00783585740.00627262360.00201330930.0009927367-3.933e-072.257e-07 P14120.0061872330.00266562430.0017847520.0005955094-3.956e-071.998e-07 P1590.0083866470.00501853730.002188220.0009591461-3.391e-072.753e-07 P170.0095512760.00445027950.00260327510.0010471037-3.552e-072.959e-07 P1760.01089122240.00552272910.00198166750.00099757-4.325e-071.372e-07 P1780.01175678540.01100175660.00215798880.001029392-4.225e-071.659e-07 P190.00742810910.00309046480.00213132820.0009460897-3.971e-071.827e-07 P1900.01786135610.01651297820.00298580890.0012192149-4.938e-071.712e-07 P200.0070449590.00324871150.0018208560.0009291652-4.047e-071.481e-07 P270.00722503710.00333318250.00188822870.0006491209-3.945e-071.818e-07 P2760.00967393830.003480820.00217718770.000703454-3.604e-072.357e-07 P300.01043487880.00780277950.00282990010.0011365804-4.144e-072.235e-07 P3640.00554717360.00285741850.00164673690.0006391143-4.365e-072.311e-07 P370.00735695460.00438153620.0018656090.0008194575-4.034e-072.107e-07 P390.00744613340.00319806710.00198895680.0008512599-3.884e-072.081e-07 P4070.00676317540.00278892470.00186087280.0007570319-3.982e-071.961e-07 P4130.00692007340.00312090390.00190793710.0007796126-3.95e-071.713e-07 P4490.00719089870.00407158970.00205228420.0009013923-3.81e-071.799e-07 P4950.00882197540.00694271930.00220770680.0009163789-3.528e-072.456e-07 P6410.00569737150.00268201710.00170824120.0007183695-3.228e-072.105e-07 P7400.00869657270.00421471340.00403059180.0150149984-3.625e-072.245e-07 P9370.00694814130.00474186960.00187294840.0008598884-3.912e-071.558e-07 Table 17: Pair-wise cosine similarity (pcs) scores with different relations from CounterFact. 28 Published as a conference paper at ICLR 2026 GPT2-XLGPT-JMETA-LLAMA-3-8B #Classes Baseline Accuracy Std Accuracy Std AccuracyStd 250.698.41.9597.03.0895.02.83 330.5399.870.397.871.4594.273.35 519.697.521.0495.682.690.083.77 1010.3294.760.9688.642.3582.245.2 156.7792.321.0583.871.4474.271.44 Table 18: Accuracy for predicting the edited relation based on low-dimensional representations of the edited matrices using a logistic regression classifier. We experiment with different number of relations (#Classes). The relations used are from the Yago dataset. k GPT2-XLGPT-J-6BMETA-LLAMA-3-8B Reversal Acc. ↑Editing Acc. ↓Reversal Acc. ↑Editing Acc. ↓Reversal Acc. ↑Editing Acc. ↓ 00.0097.140.0095.711.4391.43 192.865.7124.2961.435.7182.86 295.712.8668.5710.0028.5741.43 394.291.4374.298.5737.1434.29 495.710.0074.298.5750.0020.00 594.290.0081.434.2954.2917.14 697.140.0084.294.2958.5717.14 794.290.0082.860.0065.7110.00 895.710.0084.290.0070.008.57 995.710.0081.430.0070.007.14 1097.140.0080.000.0072.867.14 1197.140.0080.000.0072.867.14 1294.290.0078.570.0070.008.57 1395.710.0078.570.0072.867.14 1497.140.0078.571.4370.007.14 1595.710.0081.431.4371.435.71 Table 19: Reversal/Editing accuracy on Yago andROMEwith differentr− kapproximations ofW ′ V . Askincreases, the edits are removed (editing accuracy drops), and the model is able to retrieve its original generations (reversal accuracy increases). 29 Published as a conference paper at ICLR 2026 Inputk Edited Object Orig. OutputApp. Output GPT2-XL The founder of Cabinn Hotels is11TobyNeuge- bauer a former U.Sa man who has been The river Ergolz flows into11Jiu Riverthe Black Sea. Thethe Black Sea. The The band of Jeff Tweedy is11Anthraxa band of friends.a band of friends. The band of Iain Matthews is11Dire Straitsa band of Iaina band of the future The river Havel flows into11 ̄ Ohura Riverthe Danube, andthe Danube, and GPT-J The founder of Tune Hotels is6Luís I of Portu- gal a man who has beena man who has been The director of The Mirror is6Polly Drapera man who has beena man who has been The director of Darkman is6ClaudioFra- gasso a man who has beena man who has been The story of 24 is set in6Los Angelesthe year 2024, andthe year 2401, The subject of Goryeosa is6orphanthe life of the Buddhathe story of the life META-LLAMA-3-8B The artist of the painting Allegory of Vices is 11Mary Cassatt unknown. The painting was Mary Cassatt Mary The artist of the painting Religious Pro- cession in Kursk Province is 11Giulio RomanoIvan Ivanovich ShishIvan Ivanovich Shish The river Melbbach flows into11Innthe river Main in thethe river Inn at the The subject of Net Voyne! is11 international re- lations the Internet and its im- pact the Internet and its im- pact The director of Brave Command Dag- won: The Boy with Crystal Eyes is 11Sally Potterback with a new animea 1996 anime Table 20: Model outputs when using bottom-rank approximations ̃ W ′ (r,k) V on a random set of facts from the Yago dataset, edited withROME. We use the bestkfor each model. The examples show that the model outputs with approximations (Approx. Output) are semantically close to the original/unedited outputs (Orig. Output). ModelEdited Matrix Dim. GPT2-XL6400× 1600 GPT-J-6B16384× 4096 META-LLAMA-3-8B14336× 4096 QWEN2.5-7B18944× 3584 MISTRAL-7B-v0.114336 × 4096 Table 21: The dimensionalities of the edited matrices for different models. DatasetLicense CounterFact (Meng et al., 2022)MIT License YAGO (Suchanek et al., 2024) C BY 4.0 Table 22: The datasets we use in this work and their licenses. 30 Published as a conference paper at ICLR 2026 010203040 Layer 0.000 0.002 0.004 0.006 0.008 0.010 Pair-wise Cosine Similarity GPT2-XL unedited layers P112 P131 P1405 P170 P171 P178 P200 P238 P27 P276 P463 P57 P840 P915 P921 0510152025 Layer 0.00000 0.00025 0.00050 0.00075 0.00100 0.00125 0.00150 0.00175 0.00200 Pair-wise Cosine Similarity GPT-J-6B unedited layers P112 P131 P1405 P170 P171 P178 P200 P238 P27 P276 P463 P57 P840 P915 P921 051015202530 Layer 3 2 1 0 1 2 3 Pair-wise Cosine Similarity 1e6 META-LLAMA-3-8B unedited layers P112 P131 P1405 P170 P171 P178 P200 P238 P27 P276 P463 P57 P840 P915 P921 Figure 19: The average pairwise cosine similarity (pcs) of edited and unedited matrices from different layers (Yago Dataset). P112P131 P1405 P170P171P178P200P238 P27 P276P463 P57 P840P915P921 Relation 0 25 50 75 100 125 150 175 200 Increase in pairwise cosine similarity GPT2-XL GPT-J-6B META-LLAMA-3-8B Figure 20: Increase in row-wise cosine similarity ofW n after editing with the Yago dataset. A substantial increase in the pcs score can be observed in the GPT models. 31