Paper deep dive
Unforgettable Generalization in Language Models
Eric Zhang, Leshem Choshen, Jacob Andreas
Models: GPT-2, GPT-J, Llama2 7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/12/2026, 6:25:08 PM
Summary
This paper investigates the generalization of 'unlearning' or forgetting in transformer-based language models (LMs) when fine-tuned on randomized labels. The authors find that while models can be forced to produce random outputs on training data, the generalization of this forgetting to unseen task instances is highly variable and task-dependent. Factors such as task difficulty are not predictive of forgetting success; instead, model confidence and the variance of internal representations are better indicators. The study concludes that targeted skill removal via fine-tuning is unpredictable and often shallow, as information remains recoverable via linear probes.
Entities (5)
Relation Signals (3)
Random-label forgetting â exhibitsvariabilityacross â Tasks
confidence 95% · Across tasks, however, LMs exhibit extreme variability in whether LM predictions change on examples outside the training set.
Linear Probes â recoversinformationfrom â Language Models
confidence 95% · linear probes trained on LMs' representations can still perform tasks reliably after forgetting.
Model confidence â predicts â Generalization in forgetting
confidence 90% · generalization in forgetting is (weakly) predicted by the confidence of LMs' initial task predictions
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:When language models (LMs) are trained to forget (or "unlearn'') a skill, how precisely does their behavior change? We study the behavior of transformer LMs in which tasks have been forgotten via fine-tuning on randomized labels. Such LMs learn to generate near-random predictions for individual examples in the "training'' set used for forgetting. Across tasks, however, LMs exhibit extreme variability in whether LM predictions change on examples outside the training set. In some tasks (like entailment classification), forgetting generalizes robustly, and causes models to produce uninformative predictions on new task instances; in other tasks (like physical commonsense reasoning and scientific question answering) forgetting affects only the training examples, and models continue to perform the "forgotten'' task accurately even for examples very similar to those that appeared in the training set. Dataset difficulty is not predictive of whether a behavior can be forgotten; instead, generalization in forgetting is (weakly) predicted by the confidence of LMs' initial task predictions and the variability of LM representations of training data, with low confidence and low variability both associated with greater generalization. Perhaps most surprisingly, random-label forgetting appears to be somewhat insensitive to the contents of the training set: for example, models trained on science questions with random labels continue to answer other science questions accurately, but begin to produce random labels on entailment classification tasks. Finally, we show that even generalizable forgetting is shallow: linear probes trained on LMs' representations can still perform tasks reliably after forgetting. Our results highlight the difficulty and unpredictability of performing targeted skill removal from models via fine-tuning.
Tags
Links
- Source: https://arxiv.org/abs/2409.02228
- Canonical: https://arxiv.org/abs/2409.02228
Trouble viewing inline? Open PDF directly â
Full Text
58,252 characters extracted from source content.
Expand or collapse full text
Published as a conference paper at COLM 2024 Unforgettable Generalization in Language Models Eric Zhang, Leshem Choshen & Jacob Andreas MIT zeric,leshem,jda@mit.edu Abstract When language models (LMs) are trained to forget (or âunlearnâ) a skill, how precisely does their behavior change? We study the behavior of trans- former LMs in which tasks have been forgotten via fine-tuning on ran- domized labels. Such LMs learn to generate near-random predictions for individual examples in the âtrainingâ set used for forgetting. Across tasks, however, LMs exhibit extreme variability in whether LM predictions change on examplesoutsidethe training set. In some tasks (like entailment clas- sification), forgetting generalizes robustly, and causes models to produce uninformative predictions on new task instances; in other tasks (like physi- cal commonsense reasoning and scientific question answering) forgetting affects only the training examples, and models continue to perform the âforgottenâ task accurately even for examples very similar to those that appeared in the training set. Dataset difficulty is not predictive of whether a behavior can be forgotten; instead, generalization in forgetting is (weakly) predicted by the confidence of LMsâ initial task predictions and the vari- ability of LM representations of training data, with low confidence and low variability both associated with greater generalization. Perhaps most surprisingly, random-label forgetting appears to be somewhat insensitive to the contents of the training set: for example, models trained on science questions with random labels continue to answer other science questions accurately, but begin to produce random labels on entailment classification tasks. Finally, we show that even generalizable forgetting is shallow: linear probes trained on LMsâ representations can still perform tasks reliably af- ter forgetting. Our results highlight the difficulty and unpredictability of performing targeted skill removal from models via fine-tuning. 1 Introduction In the modern approach to training language models (LMs), neural sequence models are first pre-trained on a large, minimally curated corpus (typically of web text), then fine-tuned with targeted demonstrations and human feedback. The LMs that result from this procedure often possess undesirable capabilities that creators do not wish to expose to usersâfor example, the ability to generate hate speech, or to answer questions about topics unrelated to the LMâs target application. Can these capabilities be forgotten (or âunlearnedâ)? There has been widespread recent interest in developing and evaluating new techniques for removing both skills and declarative knowledge from LMs. This work has found that, on specific inputs of interest, LM behavior can be changed in targeted ways. But there has been comparatively little evaluation ofgeneralizationin forgettingâwhen an LM is trained not to respond (or to respond uninformatively) to a particular input, how does its behavior change on other inputs? This paper studies generalization behavior in forgetting. We focus on forgetting of skills (rather than knowledge) via fine-tuning on randomly labeled data for the target taskâa simple, widely used, and often highly effective method for forgetting (see Liu et al., 2024 for a recent survey). Surprisingly, we find wide variabilityacross tasksin the effectiveness and generalization of random-label forgetting. When fine-tuning on randomized responses, models will change their behavior on training inputs, but sometimes do not change their 1 arXiv:2409.02228v1 [cs.LG] 3 Sep 2024 Published as a conference paper at COLM 2024 behavior at all for other instances of the same taskâeven when fine-tuning on accurate labels doeslead to generalized improvements in accuracy. In additional experiments characterizing generalization in forgetting, we find: 1.The degree of forgetting is largely determined by the tasks that LMs are evaluated on, not the task LMs are trained to forget. 2. Generalization in forgetting is not determined by the difficulty of the task. 3. Properties that correlate with the generalization of forgetting include LM confidence as well as the variance of LM representations of training data. 4.Despite LMsâ inability to respond correctly to prompts after applying this method, we are still able to recover the correct responses using linear probes. Hence, even successful forgetting is at best shallow, and does not remove information from LMsâ representations. Generalization of learning algorithms across problems and problem instances is a major focus of study in machine learning research. Our results show similarly complex, structured cross-task variability of generalization in forgetting, and underscore the need for additional research on the relationship between the training data used for forgetting and the effect of model predictions elsewhere. 2 Related work Due to diverse privacy, security, and ethical concerns, machine unlearning has been con- ceptualized in many different ways. Early approaches defined unlearning as removing undesirable data from training sets (Cao & Yang, 2015; Bourtoule et al., 2021; Ginart et al., 2019). These approaches often require fundamental changes to model structure and/or training process, which is often infeasible. Later work relaxed the requirement of removing data from the training set. Instead, models are required to behave similarly to models trained without undesirable data points, or are simply required to stop producing outputs with desirable features. Guo et al. (2020) develop a framework for linear classifiers, and Golatkar et al. (2020a) develop a method that scrubs information from linear probes. Neel et al. (2021); Sekhari et al. (2021); Thudi et al. (2022); Golatkar et al. (2020b; 2021); Mehta et al. (2022) and Chundawat et al. (2023) present theoretical frameworks for comparing an unlearned network to a fully-retrained networks, and they propose optimization-based methods to find unlearned network under additional assumptions like convexity. Foster et al. (2024) propose model editing techniques based on estimating parameter importance using fisher information. Kurmanji et al. (2023) distinguishe between different reasons for forgetting, arguing that distinct purposes like protecting user privacy, resolving confusion, and removing biases require distinct metrics. Graves et al. (2021) argue that selectively removing training data alone is insufficient, and propose a new threat model and techniques to address them. For language models specifically, approaches to remove specific facts include gradient ascent on undesirable responses (Jang et al., 2023; Yao et al., 2023; Eldan & Russinovich, 2023), prompting with misinformation (Pawelczyk et al., 2023), linearly manipulating model representations (Ilharco et al., 2023; Belrose et al., 2023), non-linearly perturbing model representations (Li et al., 2024), and using new models to teach another model how to forget (Wang et al., 2023). While some of this prior work has studied generalization (e.g. Li et al., 2024), they study a different kind of generalization: whether model behavior remains the same on non-targeted tasks. By contrast, our work focuses on generalization between instances of a single task. Outside of research on unlearning, some past work has studied training on incorrect or random labels as a source of information aboutlearningdynamics, for example finding that models often have similar embeddings (Morcos et al., 2018), learn in a similar order (Hacohen et al., 2020) and explaining the order of learning (Hacohen & Weinshall, 2022). 2 Published as a conference paper at COLM 2024 3 Experiment setup MethodOur experiments in this paper study forgetting of capabilities (rather than factual knowledge). In order to enable uniform comparisons across tasks, we formulate each capability as binary multiple-choice question answering task. Each such taskTis associated with a training setT train , a validation setT val , and test setT test . When studying forgetting, we first fine-tune the model onT train with early stopping performed by finding the checkpoint with the highest accuracy onT val . Afterwards, we train the model to forget by fine-tuning the model again onT train but with labels chosen uniformly at random. This procedure is summarized in Figure 1. Quantifying forgettingWe quantify forgetting with two metrics. The first is the gap between the accuracy after forgetting and the expected random accuracy (50% since the tasks are binary multiple choice), which we will call theforget gap: Forget Gap=Task Accuracy After Forgettingâ 1 2 A gap of 0 indicates that the target task has been fully forgotten (all tasks involve a binary choice, and a random baseline obtains an accuracy of 1 2 ). Larger values indicate that models still achieve non-trivial accuracy. We may also wish to interpret accuracy after forgetting relative to the upper bound provided by fine-tuningâan accuracy of 55% after forgetting might be interpreted as successful or unsuccessful if fine-tuned accuracy is 95% or 56%. To quantify this intuition, we define theforget ratio: Forget Ratio= Accuracy After Fine-TuningâAccuracy After Forgetting Accuracy After Fine-Tuningâ 1 2 Here an forget ratio of 1 corresponds to complete forgetting, while a forget ratio of 0 corresponds to no decrease relative to the best attainable supervised performance. 0% Train on correct labels Train on random labels 50% Forget Accuracy Fine-tuned Accuracy Pre-trained Accuracy Forget Gap Epochs Accuracy Figure 1: Stylized learning and forgetting curves. Our experiments first fine-tune a pre-trained LM, then train it further on random labels. We call the gap between theforget accuracyand the random chance accuracy (50%) theforget gap. In many tasks we find a nonzero forget gap: after training on random labels, LMs do not generalizably learn to produce random outputs on new task instances. Tasks, evaluation details, and modelsWe experiment on 21 multiple-choice tasks commonly found within the literature. Commonsense Reasoning:We evaluate PIQA (Bisk et al., 2020), ARC easy and challenge (Clark et al., 2018), and CREAK (Onoe et al., 2021).Reading Comprehen- sion:We evaluate BoolQ (Clark et al., 2019), SciQ (Welbl et al., 2017), and Pub- MedQA (Jin et al., 2019).Math:We evaluate MathQA (Amini et al., 2019).Toxicity:We evaluate ToxiGen (Hartvigsen et al., 2022). Entailment classification and other lan- guage understanding tasks:We evaluate CoLA, MNLP, MRPC, QNLI, RTE, WNLI, CB, COPA, WIC and WSC (Wang et al., 2019). We selected these tasks to cover a broad spectrum of capabilities while also en- suring that they are multiple choice, which allows us to easily construct randomized alternatives for forgetting. We follow the Language Model Evaluation Harness standards for 0-shot evaluation (Gao et al., 2023), including the default prompts and evaluation through probabilities of the choices. To facilitate comparison across tasks, we binarize the tasks by preserving two of the possible responsesâthe true response and one randomly chosen distractorâfor each example. We evaluate models by picking the response with the highest average token likelihood and reporting the accuracy. 3 Published as a conference paper at COLM 2024 We use the publicly provided train, validation, and test sets. However, we found some datasets had trainâtest overlap. To decontaminate the datasets, we do not evaluate on questions that appeared in the training set. We also removed samples longer than 2048 characters in the prompt and combined response. Where validation sets do not exist, we use the test set. Unless otherwise specified, we limit each set to 1000 examples and subsample if needed, making training results more comparable and evaluation more efficient as proposed by Perlitz et al. (2023). All experiments use Llama2 7-billion parameter base models (Touvron et al., 2023). Addi- tional details may be found in Appendix A. 4 Does forgetting generalize? Figure 2 summarizes the task accuracy without modification, after fine-tuning, and then after running our forgetting procedure. Test accuracy almost always increases after fine- tuning, although it could decrease slightly as the validation set is not identical to the test set. During the forgetting phase, however, we observe several distinct categories of behavior (1) forget accuracy is very similar to the fine-tuned accuracy, (2) forget accuracy decreases but is still above the pre-trained accuracy, and (3) forget accuracy decreases to below the pre- trained accuracy and possibly back to 50%. Case (2) is interesting because it demonstrates asymmetry between the learning and forgetting process, as the model is unable to forget what is has just learned (analogous to hysteresis in physical systems; Ewing, 1882). Overall, we find that random-label forgetting often fails togeneralizablyremove the target behavior, but with wide variability across tasks. In general, tasks involving commonsense knowledge reasoning tasks are more resilient to forgetting, whereas lower-level linguistic acceptability and entailment classification tasks are more effectively forgettable. We also examine cross-task forgetting, where we fine-tune the model on random labels from the training set of one task and then evaluate the model on the test set of another task. As shown in Figure 3, we find that the effectiveness of the forgetting procedure is largely determined by the tasks that the model is evaluated onânot the training task. Another surprising observation is that many tasks are more effectively forgotten when training on randomized labels of other tasks than from training on their own randomized labels. As observed in the individual task evaluation, GLUE tasks focused on specific capabilities are again more susceptible to forgetting in general, whereas commonsense reasoning tasks are more resilient to forgetting. Training on forgetting commonsense reasoning tasks are also generally more effective at triggering forgetting for other tasks. 5 When does generalization occur? Does forgetting require more examples?We rule out the number of examples as the main explanation to forgetting generalization. For example, a possible concern could be that forgetting does not generalize because there are not enough training examples. We ran the same experiment with 100 examples of each task as well as 1000 (above). We find that despite an order of magnitude change, the level of forgetting is similar in both cases. Does forgetting occur with other methods?To rule out the possibility that forgetting fails to generalize due to our method of training on randomized labels, we run another experiment where we train on flipped labels instead of randomized labels. The analysis is the same as before, except now we compute the forget ratio as: Forget Ratio= Accuracy After Fine-TuningâAccuracy After Forgetting Accuracy After Fine-Tuningâ(1âAccuracy After Fine-Tuning) since we assume the minimum accuracy achievable should be 1âAccuracy After Fine- Tuning. 4 Published as a conference paper at COLM 2024 PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Task 0.4 0.6 0.8 1.0 Accuracy Change after fine-tune Change after forget Chance PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Task 0.0 0.1 0.2 0.3 0.4 0.5 Forget Gap PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Task 0.0 0.2 0.4 0.6 0.8 1.0 Forget Ratio Figure 2: Single task forgetting.Top: The blue arrow visualizes the change in held-out accuracy after fine-tuning and the red arrow illustrates the change in accuracy after forgetting. We find that many tasks do not return to the expected accuracy of 50% after forgetting.Bottom left: The forget gap (difference between forgetting accuracy and the expected random accuracy of 1/2) across tasks. Smaller values correspond to a greater degree of forgetting.Bottom right: The forget ratio (the difference fine-tuned accuracy and the forget accuracy over the difference between fine-tuned accuracy and the expected random accuracy of 1/2). Larger forget ratios correspond to more successful forgetting. 5 Published as a conference paper at COLM 2024 BoolQ ARC Challenge SciQ COPA PIQA ARC Easy PubMedQA MathQA CoLA MRPC WSC QNLI ToxiGen WiC MultiNLI WNLI CREAK RTE CB Eval Task MultiNLI QNLI MRPC WSC PubMedQA RTE BoolQ SciQ COPA WNLI CB WiC MathQA CoLA ARC Easy ARC Challenge ToxiGen PIQA CREAK Forget Task 0.10.20.00.00.10.00.10.00.50.41.01.01.01.00.30.30.50.40.4 1.00.30.00.10.00.10.20.21.01.00.41.01.01.00.70.70.30.30.6 0.30.20.00.10.00.00.10.01.01.00.51.01.00.80.71.00.30.60.9 0.50.30.00.10.10.10.10.21.01.00.40.81.00.90.81.00.60.60.6 0.40.30.00.10.10.00.10.21.01.00.90.91.01.00.91.00.90.60.6 0.40.30.00.10.10.10.20.01.01.00.61.01.01.00.91.01.01.00.6 1.00.40.10.10.10.10.30.40.50.91.01.01.00.90.81.00.40.80.8 1.00.50.00.10.10.20.40.20.41.01.01.01.01.00.91.00.70.91.0 0.30.40.00.30.10.10.30.50.40.41.01.01.01.00.81.00.40.80.7 0.20.20.00.10.10.00.10.00.40.41.01.01.01.00.91.00.10.50.6 0.20.20.00.10.00.10.10.00.40.51.01.01.01.00.71.00.50.80.5 0.30.40.00.00.10.10.10.41.00.41.00.81.01.00.81.00.70.80.8 0.30.20.00.20.10.10.10.21.00.41.01.01.01.00.91.00.90.80.9 0.40.50.00.10.10.10.30.31.00.41.01.01.01.01.01.01.00.90.9 0.30.40.00.10.10.10.20.00.60.51.01.01.01.00.81.00.90.90.8 0.50.20.00.10.20.10.30.30.50.41.01.01.01.01.01.00.90.81.0 0.40.30.00.10.10.10.30.20.40.41.01.01.01.00.81.01.00.81.0 0.20.30.00.00.00.10.10.00.50.41.01.01.01.00.81.00.90.50.6 0.30.30.00.00.00.10.10.10.40.41.01.01.01.00.91.00.60.41.0 Figure 3: Cross-task forgetting (higher values indicate more successful forgetting). We fine-tune the model on random labels from one task and then evaluate the model on another task. The vertical axis displays the task the model was trained to forget and the horizontal axis displays the task the model was evaluated on. Surprisingly, certain capabilities are robust to forgetting even after fine-tuning on random labels. Moreover, the effectiveness of the forgetting procedure is largely determined by the tasks that the model is evaluated on, not the tasks that the model was trained to forget. Note that rows and columns are presented in different orders, and clustered using the UPGMA algorithm (Sokal & Michener, 1958) 6 Published as a conference paper at COLM 2024 As shown in Figure 8, the trends are the same. The same tasks that are robust to fine-tuning on randomized labels continue to be robust on fine-tuning on flipped labels. Thus, our results are likely not specific to the choice of randomized labels, but rather a property of how fine-tuning and the tasks interact. Does forgetting occur in other models?To understanding whether this forgetting behav- ior is unique to the LLama2 7-billion parameter model or to language models in general, we also experiment with GPT-J-6B, which is a slightly weaker model than the LLaMA-2-7B, and GPT-2, which is a significantly smaller model with 124M parameters (98% smaller). As shown in Figure 8, while GPT-J and GPT-2 have lower fine-tuned accuracy, the forgetting ratio trends are broadly the same. Thus, the behavior is not unique to LLaMA-2-7B. 0.60.70.80.91.0 Finetuned Accuracy 0.0 0.5 1.0 Forget Ratio Correlation: -0.15 0.50.60.70.80.9 Original Gold Rel. Prob. 0 1 Forget Ratio Correlation: -0.62 1000200030004000500060007000 Total Variance 0.0 0.5 1.0 Forget Ratio Correlation: -0.69 Figure 4: Predictors of the Forget Ratio (y-axis). Each point is a different task.Top: The accuracy on the task after fine-tuning. The effectiveness of the forgetting procedure is not determined by the difficulty of the task (as measured by accuracy). Middle: The variance of the hidden state of the last token of the question in the fifth to last layer across examples. This variance is somewhat predictive of amount forgotten, indicating that âbroaderâ tasks are more difficult to forget.Bottom: Modelâs confi- dence in the correct response. Probability relative to the distractor is predictive of forgetting, indicat- ing that models forget more examples they were already not confident about. Are harder tasks harder to forget?An- other plausible explanation for why certain tasks are forgotten less is that harder tasks are more difficult to forget. However, as plotted in Figure 4, this is not consistently true. As a selected example, the forgetting procedure is less effective for the ARC easy dataset in comparison to the ARC challenge dataset, despite the significantly greater dif- ficulty of the latter. Thus, the effectiveness of forgetting must be determined by other properties of the task. Does model confidence predict which tasks are forgotten?We hypothesize that a modelâs confidence may be predictive of whether a task is forgotten. The reasoning for this is that if the model has a strong pref- erence for its answers on the task, a larger parameter update may be needed to over- come this âpriorâ. We examine the modelâs confidence in the correct response prior to running the forget- ting procedure. Since the probability of the correct response is not calibrated, we mea- sure the probability of the correct response relative to the incorrect response. The results are shown in Figure 4. We find that the modelâs confidence in the correct re- sponse is partially predictive of how much the model forgets. Note that this is dis- tinct from the difficulty of the task, as the modelâs confidence in the correct response is not necessarily correlated with whether it is actually correct. Does hidden state variance predict forgetting?We also hypothesize that âbroader â tasks are harder to forget. Since similar text is often mapped to similar regions in the latent space (Zhang et al., 2020), we use the variance of the hidden states of the model to quantify how much the model is able to forget. Specifically, we extract the hidden states at the last token of the question at the penultimate layer. We find that the total variance (trace of the covariance matrix) is predictive of how much the model is able to forget. Figure 4 shows that the smaller the total variance, the more effective the model is in forgetting. Note that this measure does not require access to the labels of the dataset, and it only requires access to the inference capabilities of the model and data from the task at hand. 7 Published as a conference paper at COLM 2024 020406080 Finetune Epochs 0 20 40 60 80 Forget Epochs arc_easy (corr: -0.20) 1234567 Finetune Epochs 20 40 60 80 100 Forget Epochs creak (corr: -0.45) 01020304050 Finetune Epochs 0 20 40 60 80 Forget Epochs boolq (corr: -0.19) 020406080 Finetune Epochs 0 10 20 30 40 Forget Epochs sciq (corr: -0.14) 020406080 Finetune Epochs 0 20 40 60 80 Forget Epochs pubmedqa (corr: -0.32) 1234567 Finetune Epochs 0 20 40 60 80 Forget Epochs toxigen (corr: -0.29) 010203040 Finetune Epochs 0 10 20 30 40 50 60 Forget Epochs cola (corr: -0.21) 24681012 Finetune Epochs 5 10 15 Forget Epochs mrpc (corr: -0.09) 020406080 Finetune Epochs 0 20 40 60 80 100 Forget Epochs mnli (corr: -0.34) 24681012 Finetune Epochs 0 20 40 60 80 Forget Epochs qnli (corr: -0.11) 246810 Finetune Epochs 0 10 20 30 40 50 60 Forget Epochs rte (corr: -0.34) 24681012 Finetune Epochs 0 10 20 30 40 Forget Epochs wic (corr: -0.44) Figure 5: Forgetting order vs learning order. The horizontal axis shows the forgetting time: the number of epochs until the model forgets (assigns ÂĄ 60% accuracy to the correct response for a data point). The vertical axis shows the learning time: the number of epochs until the model learns (assigns Âż 60% confidence to the correct label for a data point). We filter out the examples that are never learned or never forgotten. If fewer than 100 examples fulfil the criteria, we do not plot the task. Overall, we find that learning and forgetting orders are weakly, but consistently, anticorrelated. Can we predict which examples will be forgotten?In contrast to the task-level trends depicted in Fig. 4, we did not observe any correlation between any of the above metrics and modelsâ behavior at the level of individual examplesâfor example, example-level model confidence is not predictive of example-level forgetting. We hypothesize that different effects may dominate in this finer-grained scope, and that focusing on a narrow scope of same-task examples, other effects we did not yet uncover are too strong to see an effect with the current traits, such effects can be investigated in further work. 8 Published as a conference paper at COLM 2024 6 What is the relationship between learning and forgetting? Even if extrinsic measures of difficulty cannot predict example-level learnability (as shown in the final experiment above), is there any systematic relationship betweenlearnabilityand forgettability? Motivated by earlier work that similar architectures share consistent learning orders (Hacohen et al., 2020; Choshen et al., 2022), we hypothesize that the learning and forgettingordersare related. As pre-trained models are often already partially capable of performing the tasks we study, we analyze learning orders after âresettingâ the models to either extreme of the learning spectrum (maximum forgetting or maximum fine-tuning). Specifically, we compare the learning order of when we (1) run the forgetting procedure after fine-tuning (the same as in Section 4) and (2) when we run the fine-tuning procedure one more time afterwards (run fine-tuning after procedure in Section 4). Note that to prevent the models from learning all the examples in one epoch, we use a different fine-tuning learning rate of 3e-5 for experiment (2). To qualify an example as learned, we require the model have a confidence of at least 0.6 in the correct response. To qualify an example as forgotten, we require the model have a confidence of at most 0.6 in the correct response. In preliminary experiments, we did not find results to be sensitive to the choice of threshold. For the purpose of analysis, we ignore examples that are never learned or forgotten. If no more than 100 examples fulfill the criteria, we do not plot the task. We take the first time this occurs as the forget time/learn time. We visualize the learning orders in Figure 5. Across tasks, we find a consistent, modest correlation between learning order and forgetting order, in which the first points to be learned are typically the last to be forgotten and vice-versa. Overall, we hypothesize that the lack of a stronger correlation may be due to the shallow nature of fine-tuning. Since we are only aligning the model to the task instead of teaching it new capabilities, the learning order may be unaffected by example-level properties like difficulty. Thus, the learning order may be more related to the modelâs initial state. 7 Are âforgottenâ skills truly removed from models? One further question is if training on random labels really erases modelsâ capabilities or if it only censors the output. To examine this, we train a linear probe on the models hidden states after performing the forgetting procedure. The probes are trained on the training set and evaluated on the test set.â 2 regularization and early stopping on a validation set used to prevent overfitting as the hidden state dimension is often larger than the number of examples. We select the fifth last layer of the model as the hidden state to probe, as we find that the accuracy of probing is mostly comparable for all layers except for the very early layers and the very late layers. The results are shown in Figure 6. We find that the fine-tuning procedure largely does not influence the probing effectiveness. Thus, this procedure induces at best a shallow forgetting. This is consistent with most work that fine-tuning is often a shallow operation that does not significantly alter the modelâs capabilities (e.g.; Yadav et al., 2023; Horwitz et al., 2024). 8 Conclusion In this paper, we study the effectiveness of fine-tuning models on randomized responses in order to forget capabilities. We find that this method is effective for certain tasks, but sur- prisingly does not generalize for others. The degree of forgetting seems mostly determined by the tasks that the model is evaluated on, not the tasks that the model was trained to forget. We find that dataset difficulty and model confidence are not predictive of whether a task is forgotten. However, we find that the total variance of the hidden states of the model is predictive of how much the model is able to forget. Finally, we show that despite the 9 Published as a conference paper at COLM 2024 PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Task 0.4 0.6 0.8 1.0 Accuracy Change after fine-tune Change after forget Chance Figure 6: Probe accuracy. We plot the accuracy of a linear probe trained to classify (question, answer) pairs as correct or incorrect given LM hidden representations after pre-training, after fine-tuning, and after training on random labels. We find that forgetting largely does not influence the probing effectiveness, indicating that tasks are not truly forgotten even in cases where models generalizably learn to produce random outputs. modelsâ inability to respond correctly to prompts after applying this method, we are still able to recover the correct responses using linear probes. Thus, this is at best a shallow type of forgetting and not true removal of information from the model. Future work can focus more on understanding which specific examples are forgotten and why. While our methods were successful in predicting which broad capabilities are forgotten, they are not predictive of which specific examples are forgotten within a task. This suggests that there are more mechanisms at play that can be studied further. Acknowledgments This work was supported by the National Science Foundation under grant IIS-2238240. EZ is additionally supported by Liberty Mutual through the MIT Quest for Intelligence. References Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. MathQA: Towards interpretable math word problem solving with operation-based formalisms. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, Volume 1 (Long and Short Papers), p. 2357â2367, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1245. URL https://aclanthology.org/N19-1245. Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. LEACE: Perfect linear concept erasure in closed form. In Al- ice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.),Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, De- cember 10 - 16, 2023, 2023. URLhttp://papers.nips.c/paperfiles/paper/2023/hash/ d066d21c619d0a78c5b557fa3291a8f4-Abstract-Conference.html. 10 Published as a conference paper at COLM 2024 Yonatan Bisk, Rowan Zellers, Ronan LeBras, Jianfeng Gao, and Yejin Choi. PIQA: reasoning about physical commonsense in natural language. InThe Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, p. 7432â7439. AAAI Press, 2020. URLhttps://aaai.org/ojs/index.php/AAAI/article/view/6239. Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine Unlearning. In42nd IEEE Symposium on Security and Privacy, SP 2021, San Francisco, CA, USA, 24- 27 May 2021, p. 141â159. IEEE, 2021. doi: 10.1109/SP40001.2021.00019. URLhttps: //doi.org/10.1109/SP40001.2021.00019. Yinzhi Cao and Junfeng Yang. Towards Making Systems Forget with Machine Unlearning. In2015 IEEE Symposium on Security and Privacy, SP 2015, San Jose, CA, USA, May 17- 21, 2015, p. 463â480. IEEE Computer Society, 2015. doi: 10.1109/SP.2015.35. URL https://doi.org/10.1109/SP.2015.35. Leshem Choshen, Guy Hacohen, Daphna Weinshall, and Omri Abend. The grammar- learning trajectories of neural language models. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 8281â8297, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022. acl-long.568. URLhttps://aclanthology.org/2022.acl-long.568. Vikram S. Chundawat, Ayush K. Tarun, Murari Mandal, and Mohan S. Kankanhalli. Can Bad Teaching Induce Forgetting? Unlearning in Deep Networks Using an In- competent Teacher. In Brian Williams, Yiling Chen, and Jennifer Neville (eds.),Thirty- Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Ed- ucational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7-14, 2023, p. 7210â7217. AAAI Press, 2023. doi: 10.1609/AAAI.V37I6.25879. URL https://doi.org/10.1609/aaai.v37i6.25879. Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. BoolQ: Exploring the surprising difficulty of natural yes/no questions. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 2924â2936, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1300. URLhttps://aclanthology.org/N19-1300. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.ArXiv preprint, abs/1803.05457, 2018. URLhttps://arxiv.org/ abs/1803.05457. Ronen Eldan and Mark Russinovich. Whoâs Harry Potter? Approximate Unlearning in LLMs.ArXiv preprint, abs/2310.02238, 2023. URLhttps://arxiv.org/abs/2310.02238. James Alfred Ewing. On the production of transient electric currents in iron and steel conductors by twisting them when magnetised or by magnetising them when twisted. Proceedings of the Royal Society of London, 33(216-219):21â23, 1882. Jack Foster, Stefan Schoepf, and Alexandra Brintrup. Fast Machine Unlearning without Retraining through Selective Synaptic Dampening. In Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (eds.),Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20-27, 2024, Vancouver, Canada, p. 12043â12051. AAAI Press, 2024. doi: 10.1609/ AAAI.V38I11.29092. URLhttps://doi.org/10.1609/aaai.v38i11.29092. Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noacâh, Haonan Li, Kyle McDonell, 11 Published as a conference paper at COLM 2024 Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, 2023. URL https://zenodo.org/records/10256836. Antonio Ginart, Melody Y. Guan, Gregory Valiant, and James Zou. Making AI forget you: Data deletion in machine learning. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence dâAlch Ì e-Buc, Emily B. Fox, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 32: Annual Conference on Neural In- formation Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, p. 3513â3526, 2019. URLhttps://proceedings.neurips.c/paper/2019/hash/ cb79f8fa58b91d3af6c9c991f63962d3-Abstract.html. Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Eternal sunshine of the spotless net: Selective forgetting in deep networks. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, p. 9301â9309. IEEE, 2020a. doi: 10.1109/CVPR42600.2020.00932. URLhttps://doi.org/10.1109/CVPR42600. 2020.00932. Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Forgetting Outside the Box: Scrub- bing Deep Networks of Information Accessible from Input-output Observations. In An- drea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (eds.),Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XXIX, volume 12374 ofLecture Notes in Computer Science, p. 383â398. Springer, 2020b. doi: 10.1007/978-3-030-58526-6\23. URLhttps://doi.org/10.1007/978-3-030-58526-623. Aditya Golatkar, Alessandro Achille, Avinash Ravichandran, Marzia Polito, and Ste- fano Soatto.Mixed-privacy forgetting in deep networks.InIEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, p. 792â801. Computer Vision Foundation / IEEE, 2021.doi: 10.1109/CVPR46437. 2021.00085. URLhttps://openaccess.thecvf.com/content/CVPR2021/html/Golatkar Mixed-PrivacyForgettinginDeepNetworksCVPR2021paper.html. Laura Graves, Vineel Nagisetty, and Vijay Ganesh. Amnesiac machine learning. InThirty- Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innova- tive Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, p. 11516â 11524. AAAI Press, 2021. URLhttps://ojs.aaai.org/index.php/AAAI/article/view/ 17371. Chuan Guo, Tom Goldstein, Awni Y. Hannun, and Laurens van der Maaten. Certified data removal from machine learning models. InProceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 ofProceedings of Machine Learning Research, p. 3832â3842. PMLR, 2020. URLhttp: //proceedings.mlr.press/v119/guo20c.html. Guy Hacohen and Daphna Weinshall. Principal Components Bias in Over-parameterized Linear Models, and its Manifestation in Deep Neural Networks.J. Mach. Learn. Res., 23: 155:1â155:46, 2022. URLhttp://jmlr.org/papers/v23/21-0991.html. Guy Hacohen, Leshem Choshen, and Daphna Weinshall. Letâs agree to agree: Neural networks share classification order on real datasets. InProceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 ofProceedings of Machine Learning Research, p. 3950â3960. PMLR, 2020. URLhttp: //proceedings.mlr.press/v119/hacohen20a.html. Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3309â3326, Dublin, Ireland, 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.234. URL https://aclanthology.org/2022.acl-long.234. 12 Published as a conference paper at COLM 2024 Eliahu Horwitz, Jonathan Kahana, and Yedid Hoshen. Recovering the Pre-Fine-tuning Weights of Generative Models.ArXiv preprint, abs/2402.10208, 2024. URLhttps://arxiv. org/abs/2402.10208. Gabriel Ilharco, Marco T Ì ulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. InThe Eleventh Inter- national Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URLhttps://openreview.net/forum?id=6t0Kwf8-jrj. Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. Knowledge Unlearning for Mitigating Privacy Risks in Language Models. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (eds.),Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, p. 14389â14408. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG.805. URLhttps: //doi.org/10.18653/v1/2023.acl-long.805. Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. PubMedQA: A dataset for biomedical research question answering. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 2567â2577, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1259. URL https://aclanthology.org/D19-1259. Meghdad Kurmanji, Peter Triantafillou, Jamie Hayes, and Eleni Triantafillou.To- wards Unbounded Machine Unlearning.In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (eds.),Advances in Neural Information Processing Systems 36:Annual Conference on Neural Informa- tion Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023.URLhttp://papers.nips.c/paperfiles/paper/2023/hash/ 062d711fb777322e2152435459e6e9d9-Abstract-Conference.html. Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Adam Khoja, Zhenqi Zhao, Ariel Herbert-Voss, Cort B. Breuer, Samuel Marks, Oam Patel, Andy Zou, Mantas Mazeika, Zifan Wang, Palash Oswal, Weiran Lin, Adam A. Hunt, Justin Tienken- Harder, Kevin Y. Shih, Kemper Talley, John Guan, Russell Kaplan, Ian Steneker, David Campbell, Brad Jokubaitis, Alex Levinson, Jean Wang, William Qian, Kallol Krishna Karmakar, Steven Basart, Stephen Fitz, Mindy Levine, Ponnurangam Kumaraguru, Uday Tupakula, Vijay Varadharajan, Ruoyu Wang, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, and Dan Hendrycks. The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning.ArXiv preprint, abs/2403.03218, 2024. URL https://arxiv.org/abs/2403.03218. Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu. Rethinking Machine Unlearning for Large Language Models. ArXiv preprint, abs/2402.08787, 2024. URLhttps://arxiv.org/abs/2402.08787. Ronak Mehta, Sourav Pal, Vikas Singh, and Sathya N. Ravi.Deep Unlearning via Randomized Conditionally Independent Hessians. InIEEE/CVF Conference on Com- puter Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, p. 10412â10421. IEEE, 2022. doi: 10.1109/CVPR52688.2022.01017. URLhttps: //doi.org/10.1109/CVPR52688.2022.01017. Ari S. Morcos, Maithra Raghu, and Samy Bengio. Insights on representational similar- ity in neural networks with canonical correlation. In Samy Bengio, Hanna M. Wal- lach, Hugo Larochelle, Kristen Grauman, Nicol ` o Cesa-Bianchi, and Roman Garnett (eds.),Advances in Neural Information Processing Systems 31: Annual Conference on Neu- ral Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montr Ìeal, 13 Published as a conference paper at COLM 2024 Canada, p. 5732â5741, 2018. URLhttps://proceedings.neurips.c/paper/2018/hash/ a7a3d70c6d17a73140918996d03c014f-Abstract.html. Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi. Descent-to-Delete: Gradient-based Methods for Machine Unlearning. In Vitaly Feldman, Katrina Ligett, and Sivan Sabato (eds.),Algorithmic Learning Theory, 16-19 March 2021, Virtual Conference, Worldwide, volume 132 ofProceedings of Machine Learning Research, p. 931â962. PMLR, 2021. URLhttp: //proceedings.mlr.press/v132/neel21a.html. Yasumasa Onoe, Michael J. Q. Zhang, Eunsol Choi, and Greg Durrett. CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge. In Joaquin Vanschoren and Sai-Kit Yeung (eds.),Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, 2021. URLhttps://datasets-benchmarks-proceedings.neurips.c/paper/2021/hash/ 5737c6ec2e0716f3d8a7a5c4e0de0d9a-Abstract-round2.html. Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. In-context Unlearning: Language Models as Few Shot Unlearners.ArXiv preprint, abs/2310.07579, 2023. URLhttps: //arxiv.org/abs/2310.07579. Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen. Efficient Benchmarking (of Language Models).ArXiv preprint, abs/2308.11696, 2023. URLhttps://arxiv.org/abs/ 2308.11696. Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh. Re- member what you want to forget: Algorithms for machine unlearning. In MarcâAurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (eds.),Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, vir- tual, p. 18075â18086, 2021. URLhttps://proceedings.neurips.c/paper/2021/hash/ 9627c45df543c816a3ddf2d8ea686a99-Abstract.html. Robert R Sokal and Charles D Michener. A statistical method for evaluating systematic relationships. 1958. Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot. Unrolling SGD: Understanding Factors Influencing Machine Unlearning. In7th IEEE European Symposium on Security and Privacy, EuroS&P 2022, Genoa, Italy, June 6-10, 2022, p. 303â319. IEEE, 2022. doi: 10.1109/EUROSP53844.2022.00027. URLhttps://doi.org/10.1109/EuroSP53844. 2022.00027. Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open Foundation and Fine-tuned Chat Models.ArXiv preprint, abs/2307.09288, 2023. URLhttps://arxiv.org/abs/2307.09288. Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general- purpose language understanding systems. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence dâAlch Ì e-Buc, Emily B. Fox, and Roman Garnett (eds.), 14 Published as a conference paper at COLM 2024 Advances in Neural Information Processing Systems 32: Annual Conference on Neural In- formation Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, p. 3261â3275, 2019. URLhttps://proceedings.neurips.c/paper/2019/hash/ 4496bf24afe7fab6f046bf4923da8de6-Abstract.html. Lingzhi Wang, Tong Chen, Wei Yuan, Xingshan Zeng, Kam-Fai Wong, and Hongzhi Yin. KGA: A General Machine Unlearning Framework Based on Knowledge Gap Alignment. In Anna Rogers, Jordan L. Boyd-Graber, and Naoaki Okazaki (eds.),Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, p. 13264â13276. Association for Computational Linguistics, 2023. doi: 10.18653/V1/2023.ACL-LONG.740. URLhttps://doi.org/10. 18653/v1/2023.acl-long.740. Johannes Welbl, Nelson F. Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. InProceedings of the 3rd Workshop on Noisy User-generated Text, p. 94â106, Copenhagen, Denmark, 2017. Association for Computational Linguistics. doi: 10.18653/ v1/W17-4413. URLhttps://aclanthology.org/W17-4413. Prateek Yadav, Leshem Choshen, Colin Raffel, and Mohit Bansal. ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization. ArXiv preprint, abs/2311.13171, 2023. URLhttps://arxiv.org/abs/2311.13171. Yuanshun Yao, Xiaojun Xu, and Yang Liu. Large Language Model Unlearning.ArXiv preprint, abs/2310.10683, 2023. URLhttps://arxiv.org/abs/2310.10683. Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with BERT. In8th International Conference on Learning Represen- tations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=SkeHuCVFDr. 15 Published as a conference paper at COLM 2024 A Additional fine-tuning details Unless otherwise stated, we perform full fine-tuning in half-precision with stochastic gradi- ent descent and a learning rate of 3eâ3 with constant learning rate scheduling and gradient clipping of 1. Initial results with Adam were similar but required more memory. We fine- tune for 100 epochs with early stopping based on validation accuracy. We use a batch size of 3 which was the largest batch size that would fit in V100âs memory. We fine-tune only on the response and never on the prompt. We fine-tune for 100 epochs or until the training set reaches 99% accuracy. For our forgetting procedure, we randomly select either the correct response or the distractor before fine-tuning the model on that response in each epoch. Since allowing arbitrarily large learning rates can always lead to forgetting, we selected a learning rate where forgetting oc- cur gradually over multiple epochs, 1eâ4. To prevent undertraining, we run the forgetting procedure for 100 epochs or until the modelâs test accuracy drops below 50%, whichever comes first. B Reduced dataset size TaskSmall Forget AccuracyLarge Forget Accuracy PIQA0.710.69 ARC Easy0.840.86 ARC Challenge0.660.50 CREAK0.710.77 BoolQ0.770.50 SciQ0.840.76 PubMedQA0.730.63 MathQA0.520.56 ToxiGen0.800.77 CoLA0.590.50 MRPC0.770.80 MultiNLI0.610.79 QNLI0.710.50 RTE0.570.50 WNLI0.970.97 CB0.610.50 COPA0.510.50 WiC0.620.50 WSC0.630.66 Figure 7: Small dataset forgetting. To explore whether we have enough sample points for forgetting, we also run an experiment where only 100 examples are used for forgetting instead of 1000 in the large setting. We find that certain datasets exhibit less forgetting with the smaller dataset. However, the general trends remain the same, showing that the problem is not due explained fully by dataset size. 16 Published as a conference paper at COLM 2024 C Flipped-label task PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Eval Task PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Forget Task 0.00.10.20.50.10.00.00.00.60.20.20.40.50.30.70.30.00.50.7 0.10.20.40.40.20.10.10.00.60.40.20.40.50.50.70.40.00.50.7 0.10.20.30.50.30.00.10.10.60.30.20.50.60.40.60.60.10.50.7 0.00.00.11.00.10.00.10.00.60.20.20.50.50.40.70.40.00.50.7 0.00.10.30.20.70.10.20.60.60.20.20.50.50.50.60.40.00.50.7 0.10.10.30.50.60.20.30.10.60.20.30.50.50.40.80.50.00.40.6 0.00.00.10.50.20.00.60.10.60.70.70.50.40.30.50.30.00.50.4 0.10.10.10.40.20.00.10.20.50.70.20.50.50.40.60.50.10.50.7 0.00.10.20.40.30.00.20.21.00.20.20.40.50.40.70.40.10.50.7 0.00.10.20.50.30.00.20.20.61.00.50.50.50.40.70.50.00.40.6 0.00.00.10.50.10.00.00.00.60.41.00.40.40.40.50.50.10.50.7 0.10.00.10.30.10.00.10.00.60.20.20.70.50.30.40.60.00.50.7 0.00.00.10.20.30.00.10.10.60.20.50.31.00.40.60.30.10.50.7 0.00.10.20.50.50.00.30.00.60.30.80.50.51.00.60.50.10.50.7 0.10.00.10.20.20.00.10.00.60.40.20.50.50.80.90.30.00.50.7 0.00.00.10.30.10.00.10.00.60.30.30.50.50.30.70.80.00.50.6 0.00.10.20.30.20.00.10.30.60.20.20.40.50.30.80.40.30.50.6 0.00.00.20.40.20.00.00.10.60.60.30.40.50.40.80.40.01.00.7 0.10.10.20.50.20.00.10.10.60.80.60.50.50.50.80.30.00.50.9 PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Eval Task PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Forget Task 0.00.10.30.90.20.00.10.01.00.50.40.81.00.51.00.60.01.01.0 0.10.10.40.90.30.00.20.01.00.60.50.81.00.91.00.80.11.01.0 0.20.10.20.90.50.00.30.31.00.50.41.01.00.81.01.00.11.01.0 0.00.10.30.60.30.00.10.11.00.40.40.91.00.41.01.00.01.01.0 0.10.10.40.41.00.10.30.41.00.50.90.81.00.81.00.80.10.91.0 0.10.20.50.71.00.00.40.21.00.41.00.91.00.91.01.00.11.01.0 0.10.00.30.90.40.00.10.21.01.01.00.90.90.61.00.60.11.00.9 0.10.10.20.90.30.00.10.21.01.00.40.91.00.81.00.90.21.01.0 0.10.10.31.00.40.00.30.21.00.40.40.81.00.81.01.00.11.01.0 0.10.10.51.00.40.00.30.31.01.00.41.01.00.91.00.90.11.01.0 0.00.00.20.30.30.00.10.01.01.01.00.71.00.61.00.90.10.80.5 0.10.00.20.50.10.00.10.01.00.50.40.31.00.40.30.40.01.01.0 0.00.10.30.31.00.00.20.21.01.01.00.71.00.30.70.60.11.00.4 0.10.10.31.00.40.00.20.01.01.01.00.91.01.01.00.60.11.00.6 0.10.00.20.10.20.00.10.01.00.40.40.91.00.51.00.60.11.01.0 0.00.10.20.50.20.00.10.01.00.40.50.71.00.81.00.50.11.01.0 0.10.10.40.40.30.00.30.51.00.40.40.81.00.81.00.70.31.01.0 0.10.10.40.70.30.00.10.41.01.00.40.80.80.81.00.80.01.01.0 0.10.10.30.60.50.00.10.21.01.01.00.80.80.61.00.60.10.90.4 Figure 8: Flipped label forgetting.Top: Forget ratios for forgetting on the flipped task (higher values indicate more successful forgetting) vsBottom: Forget ratios on the randomized task. We fine-tune the model on random/flipped labels from one task and then evaluate the model on another task. The vertical axis displays the task the model was trained to forget and the horizontal axis displays the task the model was evaluated on. We see similar trends in forgetting generalization in both task constructions. 17 Published as a conference paper at COLM 2024 D Other language models PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Eval Task PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Forget Task 0.00.20.41.00.60.10.10.61.00.70.50.91.00.91.00.70.11.01.0 0.00.10.20.90.60.00.10.01.01.01.00.91.00.81.00.70.01.01.0 0.10.10.10.90.50.00.20.01.00.91.00.91.00.91.00.80.11.01.0 0.10.10.40.80.50.00.10.41.00.40.61.01.00.71.00.80.11.01.0 0.00.10.41.00.60.00.10.51.01.00.41.01.00.81.00.80.11.01.0 0.10.10.40.70.90.00.20.51.01.01.00.90.91.00.00.60.11.00.6 0.10.10.31.00.60.10.30.41.00.40.71.01.01.00.90.70.21.01.0 0.20.20.61.00.80.10.20.31.01.00.61.01.01.00.91.00.01.01.0 0.10.10.41.00.50.00.20.40.20.41.01.01.00.81.01.00.11.01.0 0.10.20.50.90.50.00.30.51.01.00.41.01.00.91.00.80.11.01.0 0.10.10.50.90.40.00.00.21.00.40.70.90.90.71.00.50.10.91.0 0.10.20.40.90.60.10.10.41.00.40.40.31.00.61.00.50.01.01.0 0.10.10.20.90.40.00.20.21.00.71.01.00.81.00.90.80.11.01.0 0.10.10.30.90.40.00.10.31.00.40.40.81.00.31.00.40.11.01.0 0.10.10.30.80.50.00.10.41.00.40.90.91.01.01.00.60.11.01.0 0.10.10.41.00.50.00.10.71.00.40.40.80.90.80.00.20.11.01.0 0.00.20.40.90.50.00.20.11.00.40.40.91.01.01.00.70.11.01.0 0.00.10.31.00.40.00.30.31.00.40.80.91.00.91.00.70.10.91.0 0.10.10.31.01.00.00.40.21.01.01.01.01.01.00.00.80.11.00.5 PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Eval Task PIQA ARC Easy ARC Challenge CREAK BoolQ SciQ PubMedQA MathQA ToxiGen CoLA MRPC MultiNLI QNLI RTE WNLI CB COPA WiC WSC Forget Task 0.00.01.01.00.20.00.00.01.00.60.30.41.00.71.00.60.01.01.0 0.00.01.00.40.20.00.00.01.01.01.00.71.00.51.00.60.01.00.9 0.00.01.00.60.00.00.00.01.00.81.00.31.00.61.00.70.01.01.0 0.00.01.00.00.00.00.00.01.00.00.40.81.00.11.00.70.01.01.0 0.00.01.00.90.30.00.00.01.01.00.10.91.00.41.00.80.01.01.0 0.00.01.00.00.80.00.00.01.01.01.00.50.91.00.00.50.01.00.3 0.00.01.00.80.20.00.00.01.00.00.60.81.01.00.90.60.01.01.0 0.00.01.01.00.70.00.00.01.01.00.40.90.91.00.91.00.01.01.0 0.00.01.01.00.10.00.00.00.00.01.01.00.90.41.01.00.01.01.0 0.00.01.00.50.00.00.00.01.01.00.10.91.00.71.00.80.01.01.0 0.00.01.00.30.00.00.00.01.00.10.50.60.90.11.00.30.00.91.0 0.00.01.00.40.20.00.00.01.00.10.10.01.00.01.00.40.01.01.0 0.00.01.00.50.00.00.00.01.00.51.00.90.70.90.90.80.01.01.0 0.00.01.00.30.00.00.00.01.00.00.10.10.90.01.00.30.01.01.0 0.00.01.00.00.10.00.00.01.00.10.90.51.01.01.00.50.01.01.0 0.00.01.01.00.00.00.00.01.00.10.20.10.90.40.00.00.01.01.0 0.00.01.00.10.10.00.00.01.00.00.10.41.00.91.00.70.01.01.0 0.00.01.00.70.00.00.00.01.00.00.80.61.00.61.00.70.00.81.0 0.00.01.01.01.00.00.00.01.01.01.01.00.91.00.20.80.01.00.1 Figure 9: Forgetting performance of other models.Top: Forget ratios for cross task on the randomized task for GPT-J-6B (higher values indicate more successful forgetting) vsBottom: Forget ratios for cross-task forgetting on the randomized task for GPT-2. The vertical axis displays the task the model was trained to forget and the horizontal axis displays the task the model was evaluated on. We see similar trends in forgetting generalization in both models and also the Llama2-7B model, which is shown in the bottom of Figure 8. 18