Paper deep dive
Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality
Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong, Zhaokun Wang, Wenyi Li, Shiyao Guo, Jinyu Guo
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise forgetting without compromising general utility remains challenging. Existing sequence- and token-level methods penalize target outputs without modeling their context-dependent retrieval paths, which can disrupt linguistic structure or suppress benign knowledge. We present ADU, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling. Exploiting the functional distinction between local and global attention heads, ADU identifies preplan positions that retrieve persistent sensitive anchors and fixes their candidate paths under the original model. It then trains attention-projection adapters to suppress attention mass along these paths while preserving local-attention structure and retain-set language modeling. Post-training activation exchange tests whether the modified attention-output module transmits the learned forgetting effect. ADU achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks, including a Forget Quality of (0.93) on TOFU. It preserves 87--98% of model utility (92.9% on average versus 81.9% for baselines) while reducing side effects in benign contexts.
Tags
Links
- Source: https://arxiv.org/abs/2608.23020v1
- Canonical: https://arxiv.org/abs/2608.23020v1
Trouble viewing inline? Open PDF directly →
Full Text
44,351 characters extracted from source content.
Expand or collapse full text
Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality Xunlei Chen Qirui Ye Yuang Li Yi Gong Zhaokun Wang Wenyi Li Shiyao Guo Jinyu Guo Abstract Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise forgetting without compromising general utility remains challenging. Existing sequence- and token-level methods penalize target outputs without modeling their context-dependent retrieval paths, which can disrupt linguistic structure or suppress benign knowledge. We present ADU, a fine-grained, training-based framework that shifts unlearning from token erasure to contextual attention-pathway decoupling. Exploiting the functional distinction between local and global attention heads, ADU identifies preplan positions that retrieve persistent sensitive anchors and fixes their candidate paths under the original model. It then trains attention-projection adapters to suppress attention mass along these paths while preserving local-attention structure and retain-set language modeling. Post-training activation exchange tests whether the modified attention-output module transmits the learned forgetting effect. ADU achieves the strongest aggregate performance among evaluated baselines on the TOFU and WMDP benchmarks, including a Forget Quality of 0.930.93 on TOFU. It preserves 87–98% of model utility (92.9% on average versus 81.9% for baselines) while reducing side effects in benign contexts. Introduction Large language models (LLMs) have achieved advances in language understanding and generation (14; 23) through model scaling (9) and diverse pretraining data (40). However, models can reproduce sensitive content (6), private information (19), or unsafe material (26; 27). As the General Data Protection Regulation (GDPR) and the “right to be forgotten” receive attention (7), machine unlearning has emerged as a privacy mechanism (42). It aims to reduce the accessibility of requested knowledge while preserving unrelated model capabilities (37). Balancing unlearning efficacy and model utility remains fundamentally challenging (22; 2). Prompt-based (1; 35) and auxiliary-model approaches (5; 29) can change outputs without updating the target model, yet knowledge may remain recoverable through extraction attacks (25). Parameter-level unlearning remains necessary when the objective is to modify the deployed model itself, particularly for private or copyrighted content. Figure 1: Sequence methods minimize the joint probability of the entire QA pair. Token methods blindly suppress the probabilities of individual tokens. We sever attention pathways that point to the anchor tokens. Figure 2: Generation inequality and temporal retrieval rhythm. (a) Local heads preserve near-diagonal dependencies, global heads form long-range patterns. (b) RAS peaks indicate preplan transitions, APS peaks identify persistent anchors, shaded bands mark selected regions. Post-training edge-contribution replacement tests path-specific causal mediation after unlearning. Existing parameter-level methods are categorized by their optimization target (43; 12). Sequence-level unlearning treats an entire question–answer sequence as the forget target (18). Because its loss spans syntactic and content tokens, it can damage linguistic structure and generation fluency (20). Token-level unlearning instead penalizes selected sensitive tokens (10) and limits damage by avoiding noncritical positions (37; 13). Nevertheless, both approaches define forgetting over visible output targets rather than the internal, context-dependent computation through which sensitive knowledge is retrieved. This distinction matters because the same entity can be sensitive in one context and benign in another. Static sequence or token penalties may therefore suppress legitimate occurrences, causing excessive forgetting and degraded factual behavior in non-sensitive contexts (30). As illustrated in Fig. 1, the desired intervention should target the corresponding retrieval computation rather than erase the entity itself. This raises a natural question: Can an LLM forget a context-specific retrieval route without destabilizing ordinary generation? Motivated by evidence that attention heads contribute unevenly to model behavior (17), we call their asymmetric temporal roles generation inequality. Local heads maintain short-range dependencies, whereas global heads route information from distant, persistently attended tokens. Around semantic transitions, local attention shifts retrospectively at preplan positions, after which global heads retrieve earlier anchors that shape subsequent generation. Average Backward Distance separates the two head groups, while Retrospective Attention Shift and Anchor Persistence Score locate candidate preplans and anchors (Fig. 2). Based on this rhythm, we propose Attention Decoupling Unlearning (ADU), a training-based method for context-specific pathway suppression. ADU computes the head partition and preplan–anchor paths once under the original model and keeps them fixed during training. Attention-projection adapters reduce path mass, while retained language modeling and local-attention preservation constrain collateral changes. The learned decoupling operates in every forward pass; bidirectional edge-contribution replacement runs only after training to test whether the selected preplan–anchor paths causally mediate the resulting forgetting effect. Unlike representation-level methods such as RMU, token-editing methods such as MET, and broad attention suppression such as ASU, ADU optimizes a context-indexed retrieval pathway rather than a hidden representation or token identity. Attention weights locate candidate routes but are not treated as explanations by themselves: the pathway objective controls transported contributions under bounded values, while bidirectional edge-contribution replacement provides post-training tests of path-specific causal mediation. Across TOFU, WMDP, and MUSE-Harry Potter, ADU achieves the strongest aggregate forgetting–retention trade-off among the evaluated baselines. It obtains 93% Forget Quality on TOFU and preserves 87–98% of model utility, averaging 92.9% compared with 81.9% for baselines. These results support contextual pathway suppression as a more targeted alternative to sequence or token erasure. In summary, our contributions are: 1. We formulate LLM unlearning as contextual pathway decoupling and characterize a preplan–anchor temporal pattern arising from the functional specialization of local and global attention heads. 2. We propose ADU, which trains attention-projection adapters to suppress fixed preplan-to-anchor pathways while retaining language modeling and local-attention preservation constrain collateral behavior. 3. We provide conditional theoretical guarantees and extensive evaluations on WMDP, TOFU, and MUSE-Books, demonstrating state-of-the-art forgetting retention trade-offs and validating the causal role of the targeted pathways. Figure 3: Illustration of the workflow. Before predicting the next important token, the preplan token “in” queries earlier sensitive anchors through long-range backward attention. After training, ADU suppresses this sensitive dependency path while preserving local anchor structure and maintaining the continuity of local semantic segments. Related Work Sequence-level Unlearning Many parameter-level methods optimize a coarse sequence-level objective. Gradient Ascent (GA) (31) directly maximizes the forget-set loss and can produce unstable parameter updates. Negative Preference Optimization (NPO) (39) regularizes this process with a preference objective, but still treats the complete QA pair as one optimization target. Together with their variants (34; 41; 33; 36), such objectives distribute forgetting pressure across many sequence positions, including syntactic function words, without explicitly separating knowledge-bearing content from linguistic structure. Consequently, stronger forgetting can coincide with degraded generation fluency and reduced general capabilities. Token-level Unlearning Token-level methods narrow the target by suppressing specific token logits or redirecting them toward alternatives (37; 15). Although this avoids some sequence-wide penalties, residual latent semantics may remain, while redirected generation can introduce factual hallucinations. Recently, 30 identifies important tokens and suppresses attention toward them across the sequence. However, this strategy remains centered on static token salience rather than a context-specific query-to-key retrieval route, limiting contextual discrimination (32; 11). The same entity may therefore be weakened in both sensitive and benign contexts, causing excessive forgetting and disrupting normal knowledge retrieval (38). These families differ in granularity but define forgetting mainly over the sequence or token being generated, rather than the contextual computation that retrieves it. ADU instead identifies a recurring preplan–anchor rhythm under the original model and fixes the resulting candidate paths before training. Attention-projection adapters then suppress mass along these paths, while retain language modeling and local-attention preservation constrain collateral changes. Thus, ADU targets context-specific retrieval without directly erasing the sensitive token itself. Method ADU is a training-based unlearning method (see Figure 3). Given an original model θ0 _0, a forget set DfD_f, and a retain set DrD_r, it learns θ1=ADU(θ0,Df,Dr) _1=U_ADU( _0;D_f,D_r) through attention-projection adapters. The learned pathway decoupling acts in every standard forward pass of θ1 _1, while counterfactual activation exchange is used only after training to identify the internal computation mediating forgetting. Generation Inequality and Temporal Rhythm During autoregressive generation, we formalize generation inequality as unequal temporal roles: local heads maintain short-range dependencies, whereas global heads retrieve distant, persistently attended tokens. Around semantic transitions, retrospective local shifts mark preplan positions, and persistent global attention identifies earlier anchors that shape generation. This pattern yields the preplan–anchor retrieval hypothesis operationalized below (Figure 2). Let x=(u,o)x=(u,o) be a context continuation sequence and RxR_x the continuation positions. Let Aθ,t,s(l,h)(x)A_θ,t,s^(l,h)(x) be the causal attention weight from query position t to key position s≤ts≤ t in layer l, head h. On a held-out calibration split calD_cal, the head’s average backward distance under the model: d(l,h)=x∼cal[1|Rx|∑t∈Rx∑s≤tAθ0,t,s(l,h)(x)(t−s)].d^(l,h)=E_x _cal [ 1|R_x| _t∈ R_x _s≤ tA_ _0,t,s^(l,h)(x)(t-s) ]. (1) The bottom and top ρ fractions form HlocH_loc and HglobH_glob, respectively. We use ρ=0.3ρ=0.3, compute the partition once from the original model, and keep it fixed during training. Let A¯θ0loc(x) A_ _0^loc(x) and A¯θ0glob(x) A_ _0^glob(x) denote original-model attention averaged over the fixed local and global head sets. Given a retrospective distance cap W and a fixed future continuation window Fx(s)F_x(s), we define rt r_t =∑s≤tA¯θ0,t,sloc(x)min(t−s,W), = _s≤ t A_ _0,t,s^loc(x) (t-s,W), (2) Δt _t =|rt−rt−1|,as=1|Fx(s)|∑u∈Fx(s)A¯θ0,u,sglob(x). =|r_t-r_t-1|, a_s= 1|F_x(s)| _u∈ F_x(s) A_ _0,u,s^glob(x). Here, Δt _t is evaluated at continuation positions with a preceding continuation position, and positions without a valid future window are excluded from APS selection. Large Δt _t identifies a candidate transition, whereas large asa_s indicates persistent influence on subsequent queries. For selection ratio q, let Q1−q(Δ,x)Q_1-q( ;x) and Q1−q(a,x)Q_1-q(a;x) denote the corresponding within-sequence quantiles. We define Tpre(x) T_pre(x) =t∈Rx∣Δt≥Q1−q(Δ;x),rt≥τras, = \t∈ R_x _t≥ Q_1-q( ;x),\ r_t≥ _ras \, (3) Sanc(x) S_anc(x) =s∈Rx∣as≥Q1−q(a,x)∩Cf(x). = \s∈ R_x a_s≥ Q_1-q(a;x) \∩ C_f(x). Here, Cf(x)C_f(x) contains candidate sensitive positions derived from the forget continuation. Exact window construction and sensitive-position filtering are detailed in Appendix B. These signals locate candidate retrieval sites; ADU turns the identified rhythm into a trainable mechanism by suppressing the corresponding attention pathway. Pathway Contributions and Causal Mediation For each forget sample, (x)P(x) contains tuples e=(l,h,t,s)e=(l,h,t,s) such that (l,h)∈Hglob(l,h)∈ H_glob, t∈Tpre(x)t∈ T_pre(x), s∈Sanc(x)s∈ S_anc(x), and s<ts<t, where the last condition enforces causal attention order. We distinguish the contribution before and after the effective output projection: C~e(θ,x) C_e(θ,x) =Aθ,t,s(l,h)(x)Vθ,s(l,h)(x), =A_θ,t,s^(l,h)(x)V_θ,s^(l,h)(x), (4) Ce(θ,x) C_e(θ,x) =WO,θ(l,h)C~e(θ,x). =W_O,θ^(l,h) C_e(θ,x). Here, C~e C_e is the head-space contribution manipulated by the intervention, while CeC_e is its residual-stream image controlled by the pathway objective. For a∈0,1a∈\0,1\, define the path-specific causal mediator as a(x)=(C~e(θa,x))e∈(x).M_a(x)= ( C_e( _a,x) )_e (x). (5) Replacing a(x)M_a(x) by b(x)M_b(x) subtracts the running model’s selected contributions and adds the source model’s corresponding contributions at each affected pre-output-projection head output. The recipient model retains its own WOW_O and every unpatched computation, and all downstream activations are recomputed. Here, a=0a=0 denotes the original model and a=1a=1 the trained ADU model. For theoretical and causal analysis, let If(x)I_f(x) denote the positions of a sensitive continuation span. For each i∈If(x)i∈ I_f(x), let yiy_i be the target token and i(x)C_i(x) a prespecified set of non-target contrasts. Under teacher forcing, the sensitive accessibility score is Yθ(x)=1|If(x)|∑i∈If(x)sθ(yi,x<i),Y_θ(x)= 1|I_f(x)| _i∈ I_f(x)s_θ(y_i,x_<i), (6) where sθ(yi,x<i)=zθ(yi∣x<i)−log∑c∈i(x)expzθ(c∣x<i).s_θ(y_i,x_<i)=z_θ(y_i x_<i)- _c _i(x) z_θ(c x_<i). This dataset-agnostic score is used only to formalize sensitive accessibility and counterfactual effects; it is not part of the ADU training objective. Δf _f =x∼Df[Yθ0(x)−Yθ1(x)]≥ϵf>0, =E_x D_f [Y_ _0(x)-Y_ _1(x) ]≥ _f>0, (7) Δr _r =x∼Dr[ret(πθ1(⋅∣x),πθ0(⋅∣x))]≤ϵr, =E_x D_r [D_ret ( _ _1(· x), _ _0(· x) ) ]≤ _r, where retD_ret measures retain-behavior discrepancy. Let Y(a,,x)Y(a,m;x) denote the counterfactual score from θa _a when its selected head-space contributions are replaced by m before WO,θaW_O, _a. All unpatched computations and parameters stay as θa _a, and downstream activations are recomputed. Define Yab(x)=Y(a,b(x),x)Y_ab(x)=Y(a,M_b(x);x), where the first index marks the running model and the second the mediator source. Writing fE_f for expectation over DfD_f, we obtain TE =f[Y00−Y11], =E_f[Y_00-Y_11], (8) IEsup _sup =f[Y00−Y01], =E_f[Y_00-Y_01], IEres _res =f[Y10−Y11]. =E_f[Y_10-Y_11]. Under consistency, Y(a,a(x),x)=Yθa(x)Y(a,M_a(x);x)=Y_ _a(x), so Yaa=YθaY_a=Y_ _a and TE=ΔfTE= _f. The suppression effect replaces Base contributions with their ADU counterparts, whereas the restoration effect replaces ADU contributions with their Base counterparts. These bidirectional interventions test whether the selected preplan–anchor contributions causally mediate the learned forgetting effect. Direct remainders representing additional unpatched routes are defined in Appendix A. Training Objective and Theoretical Guarantees Let Nx=max(1,|(x)|)N_x= (1,|P(x)|) and Aθ,e(x)=Aθ,t,s(l,h)(x)A_θ,e(x)=A_θ,t,s^(l,h)(x) for e=(l,h,t,s)e=(l,h,t,s). The pathway mass and forget loss are PMθ(x)=1Nx∑e∈(x)Aθ,e(x),ℒf=x∼Df[PMθ(x)].PM_θ(x)= 1N_x _e (x)A_θ,e(x), _f=E_x D_f[PM_θ(x)]. (9) The pathway indices are computed once under θ0 _0 and kept fixed during training; gradients flow through the selected attention values but not through the discrete mask. We combine ℒfL_f with language modeling on DrD_r and a row-wise loss that preserves original-model local attention on Df∪DrD_f∪ D_r: ℒADU=αℒf+(1−α)(ℒlm+ℒloc).L_ADU= _f+(1-α) (L_lm+L_loc ). (10) Here, α∈[0,1]α∈[0,1] controls the forget–retain balance. The backbone remains frozen, while LoRA adapters update WQ,WK,WV,W_Q,W_K,W_V, and WOW_O in layers containing selected global heads. Samples with (x)=∅P(x)= contribute zero to ℒfL_f. The complete retain objective, trainable scope, and empty-path handling are detailed in Appendix B. Method Forget tasks(%) Retain tasks(%) Llama3.1-8B-Instruct Bio.↓ Cyber↓ MMLU↑ GSM8K↑ Flu.↑ Base 71.86 45.37 68.16 67.83 3.74 NPO_KL‡ (39) 56.38 34.32 52.37 53.55 2.79 RMU‡ (16) 39.55 31.75 53.43 52.64 2.90 ICUL⋄ (21) 41.31 30.46 60.69 58.40 3.58 ALU⋄ (24) 31.87 29.94 63.78 58.65 3.26 MET† (37) 34.63 31.47 52.19 54.19 3.18 ASU† (30) 34.49 33.17 58.58 57.24 3.31 ALTER† (3) 29.69 30.77 60.10 57.60 3.20 ADU† (Ours) 27.32 27.97 62.84 58.82 3.34 Qwen3-14B Bio.↓ Cyber↓ MMLU↑ GSM8K↑ Flu.↑ Base 76.07 50.98 75.18 79.25 3.80 NPO_KL‡ (39) 60.59 39.85 64.88 63.50 2.88 RMU‡ (16) 43.85 36.24 66.51 64.86 3.08 ICUL⋄ (21) 46.82 31.50 65.22 68.37 3.69 ALU⋄ (24) 31.71 30.38 67.67 68.81 3.58 MET† (37) 38.28 34.53 65.49 66.29 3.27 ASU† (30) 32.09 33.89 69.88 71.54 3.47 ALTER† (3) 34.24 36.58 71.55 70.83 3.35 ADU† (Ours) 29.40 29.12 70.91 73.28 3.57 Table 1: Multiple-choice accuracy on the forgetting/retention benchmark after unlearning. † , ⋄ , and ‡ denote token-level training, prompt-based methods, and sequence-level training, respectively. From pathway training to knowledge suppression. If ‖WO,θ(l,h)Vθ,s(l,h)(x)‖2≤B\|W_O,θ^(l,h)V_θ,s^(l,h)(x)\|_2≤ B for all selected edges, then 1Nx∑e∈(x)‖Ce(θ,x)‖2≤BPMθ(x). 1N_x _e (x)\|C_e(θ,x)\|_2≤ B\,PM_θ(x). (11) Thus, minimizing pathway mass controls the average transported magnitude of selected edge contributions rather than treating attention weights as explanations. These contributions form the path-specific component of the attention-output computation examined by activation exchange. For an affected query–head row j, let pjp_j and pj′p _j be its selected-anchor mass before and after pathway decoupling, with δj=pj−pj′≥0 _j=p_j-p _j≥ 0. Write j(pj)=pjμS,j+(1−pj)μS¯,j,Γj=μS,j−μS¯,j,o_j(p_j)=p_j _S,j+(1-p_j) _ S,j, _j= _S,j- _ S,j, where μS,j _S,j and μS¯,j _ S,j are the normalized selected-anchor and complementary value-output mixtures. A mass-transfer intervention holds these mixtures fixed while reducing pjp_j, yielding j(pj′)−j(pj)=−δjΓj.o_j(p _j)-o_j(p_j)=- _j _j. Let g()=Y(1(p1),…,J(pJ))g(p)=Y(o_1(p_1),…,o_J(p_J)) be the sensitive score induced by the affected attention outputs. Assume that g is differentiable and, at every point along the intervention path, ⟨∇jg,Γj⟩≥κj>0. _o_jg, _j ≥ _j>0. For the retain-discrepancy functional grg_r, assume L-Lipschitz continuity and ‖Γj‖2≤Bj\| _j\|_2≤ B_j. Then g(′)−g() g(p )-g(p) ≤−∑jκjδj, ≤- _j _j _j, (12) |gr(′)−gr()| |g_r(p )-g_r(p)| ≤L∑jBjδj. ≤ L _jB_j _j. The first inequality gives a sufficient condition for reducing sensitive log-odds, while the second bounds the retain-discrepancy change attributable to the same intervention. Complete proofs are provided in Appendix A. Together, the pathway loss controls attention contributions, directional alignment translates their reduction into lower sensitive log-odds, and the retain objective constrains changes outside the pathway. Bidirectional activation exchange then tests whether the modified attention-output computation mediates the forgetting effect. Method TUD NEK GEK Llama3.1-8B R-L↓ TR↑ FQ↑ R-L↑ Acc↑ Acc↑ NPO_KL‡ (39) 0.31 0.78 0.72 0.58 62.2 64.2 RMU‡ (16) 0.19 0.91 0.88 0.62 65.1 68.3 ICUL⋄ (21) 0.17 0.94 0.54 0.61 63.8 69.0 ALU⋄ (24) 0.13 0.95 0.67 0.64 66.7 70.6 MET† (37) 0.18 0.88 0.92 0.64 63.3 70.2 ASU† (30) 0.16 0.93 0.87 0.66 69.6 71.0 ALTER† (3) 0.14 0.91 0.81 0.60 67.2 71.3 ADU† (Ours) 0.11 0.96 0.93 0.69 70.8 72.3 Table 2: Performance comparison on TOFU (10%) with Llama3.1-8B-Instruct. Method BLEU↓ R-L↓ MMLU↑ Flu.↑ Original 74.80 85.14 46.33 3.63 NPO‡ 1.55 14.08 42.70 2.96 WHP‡ 23.68 17.93 43.49 2.52 ALU⋄ 7.21 14.85 44.34 3.27 ICUL⋄ 27.50 25.89 44.02 3.34 ALTER† 6.96 10.40 43.84 2.32 Ours† 4.78 9.49 45.64 3.29 Table 3: Results on MUSE-Harry Potter with Llama2-7B. Experiment Experiment Settings Datasets We evaluate our method on three benchmarks. WMDP (16) verifies forgetting and retention effectiveness by assessing model knowledge in sensitive domains such as biosafety and cybersecurity. MUSE-Harry Potter (28) assesses copyright unlearning: models are first fine-tuned on the Harry Potter book content to memorize it, then unlearned to forget that content. TOFU (19) examines boundary preservation between the forget set and its neighboring retain set, simulating synthetic and real-world unlearning scenarios. We note that all three benchmarks involve extended, context-rich responses where the model must retrieve factual knowledge through multi-token generation–precisely the regime where preplan-to-anchor attention patterns emerge and ADU’s pathway decoupling is most effective. To test general ability, we apply MMLU (8) for fact answering and GSM8K (4) for math reasoning. Datasets details and configurations are provided in Appendix C.1. Metrics For WMDP, we report multiple choice accuracy on Bio and Cyber as forgetting metrics, where lower values indicate stronger forgetting. MMLU and GSM8K are used as retention metrics, where higher values indicate stronger utility preservation. For MUSE Harry Potter, we report BLEU and ROUGE-L to measure textual overlap with copyrighted content, together with MMLU and fluency for utility. For TOFU, we report ROUGE-L on the target unlearned data, Top 5 exclusion rate, Forget Quality, ROUGE-L on neighboring knowledge, neighboring accuracy, and general knowledge accuracy. Fluency is evaluated by GPT-4o on a 1 to 5 scale. We also report Forgetting Performance (FP) and Retaining Performance (RP) in analysis, where FP is the average of WMDP Bio and Cyber, and RP is the average of MMLU and GSM8K. Details are provided in Appendix C.2. Baselines We compare with sequence-level training, prompt-based methods, and token-level training methods. Sequence-level training includes NPO_KL (39), and Representation Misdirection for Unlearning (RMU) (16). Prompt-based methods include ICUL (21) and ALU (24). Token-level training includes Model Edit Token (MET) (37), Attention Shift Unlearning (ASU) (30), and Hydra Suppress Unlearning (ALTER) (3). Main Result Forgetting-Retention Effectiveness We report WMDP forgetting and general retention results (Table 1). On Llama3.1-8B-Instruct, ADU reduces Bio accuracy from 71.86 to 27.32 and Cyber accuracy from 45.37 to 27.97, retaining MMLU and GSM8K at 62.84 and 58.82. On Qwen3-14B, ADU achieves the lowest Bio and Cyber accuracy and the best GSM8K retention among unlearning methods; prompt-based ALU ranks second on Cyber forgetting without modifying model parameters. These results show that ADU improves trainable forgetting and retention trade off instead of optimizing one forgetting metric at the cost of utility. Table 3 evaluates copyright unlearning on MUSE Harry Potter. ADU achieves the lowest ROUGE-L among unlearning methods and strongest MMLU retention. Although NPO gives lower BLEU, its MMLU drops to 42.70, while ADU keeps MMLU at 45.64 and fluency at 3.29. This pattern supports the pathway view. ADU weakens the route to memorized content while avoiding broad degradation of general next token behavior. Settings and costs are in Appendix C.4. Setting Avg↓ MMLU↑ GSM8K↑ ADU 27.65 62.84 58.82 w/o pathway loss 47.79 63.22 58.44 w/o retain objective 25.68 56.32 52.10 random heads 34.79 59.26 55.28 w/o anchor filter 26.19 58.40 54.32 w/o RAS preplan 32.38 60.05 55.99 Table 4: Component ablation on Llama3.1-8B-Instruct. Avg denotes the average of WMDP Bio and Cyber. Boundary Preservation To further verify the impact of unlearning methods on neighboring and retained knowledge, we conducted experiments on TOFU (10%), as shown in Table 2. Sequence-level training methods struggle to balance the trade-off between forgetting and model utility. Prompt-based methods provide stronger retention, but their forgetting and neighboring preservation remain unstable. Although token-level training methods improve this balance, existing variants may still disrupt semantic dependencies shared with neighboring facts. For example, ASU obtains 69.6% NEK accuracy, but its TUD R-L remains 0.16. By severing hazardous retrieval pathways while preserving adjacent semantic pathways, ADU achieves the best TUD R-L of 0.11, TR of 0.96, and the best NEK and GEK of 70.8% and 72.3%, showing stronger boundary preservation and general utility. Discussions Component ablation Table 4 isolates each ADU component on Llama3.1-8B-Instruct. Removing the pathway loss raises forgetting metrics sharply, indicating that suppressing sensitive pathway mass drives forgetting. Removing the retain loss keeps forgetting but reduces MMLU and GSM8K by 6.52 and 6.72 points, showing that retain language modeling is essential for utility. Random heads weaken both forgetting and retention, suggesting that the local-global partition is non-interchangeable. Removing the sensitive anchor filter harms MMLU and GSM8K by penalizing benign high-APS anchors. Removing the RAS based preplan selection weakens forgetting and retention, confirming that ADU benefits from intervening at the transition point before sensitive anchors guide later generation. Full results including TOFU metrics are provided in Appendix D.2. Path-specific causal validation. Figure 2(b) visualizes how RAS peaks mark preplan transitions and APS peaks identify persistently attended anchors, localizing candidate pathways without proving they control sensitive retrieval. We therefore intervene directly on the pre-output-projection contributions of the identified preplan–anchor edges. Removing them from Base lowers WMDP Avg. from 58.62 to 36.37, whereas removing a cardinality-matched random edge set yields 57.09, showing retrieval depends specifically on the selected pathway rather than an arbitrary same-sized perturbation. Conversely, replacing ADU’s selected contributions with their Base counterparts restores WMDP Avg. from 27.65 to 47.23, while matched-random replacement reaches only 28.82. The selected interventions change MMLU by only -0.58 and +0.53 points, respectively. These complementary results link ADU’s training target to its behavioral effect: the selected contributions support a substantial portion of Base retrieval, and restoring their original computation recovers much of the access suppressed by ADU. Appendix E provides complete bidirectional replacement analysis. Condition WMDP Avg.↓ MMLU↑ Base 58.62 68.16 Base −- Selected C~e C_e 36.37 67.58 Base −- Matched random C~e C_e 57.09 67.52 ADU 27.65 62.84 ADU ← Base selected 47.23 63.37 ADU ← Base matched random 28.82 61.64 Table 5: Path-specific contribution interventions on Llama3.1-8B-Instruct. “←” replaces selected contributions in the running model with their source-model counterparts. Figure 4: Parameter sensitivity analysis on WMDP and retention tasks with Llama3.1-8B-Instruct. Hyperparameter Sensitivity Analysis We analyze the selection ratio q, and the loss balance α on Llama3.1-8B-Instruct. Figure 4 reports the forgetting-retention trade-off when varying one hyperparameter while fixing the others to their default values. Small q misses sensitive pathways and leaves higher WMDP accuracy, whereas large q includes benign anchors and harms retention. The default q=0.4q=0.4 achieves FP 27.65 and RP 60.83, which forms a stable trade-off knee while avoiding the retention degradation observed at larger q values. A small α underweights the pathway objective and generally weakens forgetting, whereas large α provides limited forgetting gains and lowers retention performance. The default α=0.3α=0.3 gives the best tested balance. Full numerical results and seed stability are reported in Appendix D.3 and Appendix D.4. Robustness Analysis. We group six attacks into prompt scaffolding (few-shot, masking, and role-play CoT) and adaptive recovery (anchor shift, multi-turn probing, and repeated sampling). They test whether altered reasoning contexts or retrieval strategies re-elicit forgotten knowledge. As shown in Fig. 5, although ASU has a smaller increase under prompt scaffolding, ADU still achieves the lowest attacked accuracy. Under adaptive recovery, ADU has both the smallest increase and lowest final accuracy, remaining below ALU and ASU. Thus, pathway decoupling limits knowledge recovery beyond the original prompt. More results and details are in Appendix F.1. Figure 5: Groupwise worst-case WMDP accuracy under different attacks on Llama3.1-8B-Instruct. Conclusion Our work formulates LLM unlearning as contextual pathway decoupling: sensitive knowledge is retrieved through context-dependent internal routes and should not be reduced to an entire sequence, a static token, or a prompt-level refusal. Based on this view, we introduce ADU, which identifies a preplan–anchor rhythm from the temporal specialization of local and global attention heads, fixes candidate retrieval paths under the original model, and trains attention-projection adapters to suppress them while preserving retain-set language modeling and local-attention structure. The resulting framework provides a persistent parameter-level mechanism that targets sensitive retrieval in context, while limiting excessive forgetting, utility degradation, and recovery under altered prompts. Experiments on WMDP, TOFU, and MUSE-Books demonstrate strong forgetting–retention trade-offs, and bidirectional edge-contribution interventions establish that the selected paths causally mediate a substantial portion of sensitive retrieval and the learned forgetting effect. Future work will extend pathway identification beyond attention, improve automatic anchor construction, and develop stronger guarantees against residual knowledge recovery. References Bhaila et al. (2025) K. Bhaila, M. Van, and X. Wu Soft prompting for unlearning in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 4046–4056. Cited by: Introduction. Cha et al. (2024) S. Cha, S. Cho, D. Hwang, H. Lee, T. Moon, and M. Lee Learning to unlearn: instance-wise unlearning for pre-trained classifiers. In Proceedings of the AAAI conference on artificial intelligence, Vol. 38, p. 11186–11194. Cited by: Introduction. Chen et al. (2026) X. Chen, J. Guo, Y. Li, Z. Wang, Y. Gong, J. Zou, J. Wei, and W. Tian ALTER: asymmetric lora for token-entropy-guided unlearning of llms. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 35366–35374. Cited by: Table 1, Table 1, Table 2, Baselines. Cobbe et al. (2021) K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: Datasets. Geng et al. (2025) J. Geng, Q. Li, H. Woisetschlaeger, Z. Chen, F. Cai, Y. Wang, P. Nakov, H. Jacobsen, and F. Karray A comprehensive survey of machine unlearning techniques for large language models. arXiv preprint arXiv:2503.01854. Cited by: Introduction. Gong et al. (2026) Q. Gong, X. Yang, X. Chen, J. Lai, H. Meng, and X. Tang FedOrtho: efficient federated unlearning via orthogonal convolution and adaptive soft pruning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, p. 8009–8018. Cited by: Introduction. Grynbaum et al. (2023) M. M. Grynbaum et al. The times sues openai and microsoft over ai use of copyrighted work. The New York Times 27 (1). Cited by: Introduction. Hendrycks et al. (2021) D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt Measuring massive multitask language understanding. In International Conference on Learning Representations, External Links: Link Cited by: Datasets. Hu et al. (2025) Z. Hu, Y. Zhang, M. Xiao, W. Wang, F. Feng, and X. He Exact and efficient unlearning for large language model-based recommendation. IEEE Transactions on Knowledge and Data Engineering. Cited by: Introduction. Jiang et al. (2025) P. Jiang, X. Lyu, Y. Li, and J. Ma Backdoor token unlearning: exposing and defending backdoors in pretrained language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 24285–24293. Cited by: Introduction. Jin et al. (2025) M. Jin, W. Luo, S. Cheng, X. Wang, W. Hua, R. Tang, W. Y. Wang, and Y. Zhang Disentangling memory and reasoning ability in large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1681–1701. Cited by: Token-level Unlearning. Kim et al. (2026) H. Kim, K. Kim, S. Chae, and S. Yoon Unlearning-aware minimization. Advances in Neural Information Processing Systems 38, p. 93806–93829. Cited by: Introduction. Lee et al. (2026) H. K. Lee, R. Liu, and L. Xiong Direct token optimization: a self-contained approach to large language model unlearning. In Findings of the Association for Computational Linguistics: ACL 2026, p. 42083–42100. Cited by: Introduction. Lee et al. (2020) J. Lee et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36 (4), p. 1234–1240. Cited by: Introduction. Li et al. (2025a) J. Li, C. Zhang, M. Du, H. Zhang, Y. Chen, Q. Wei, J. Fang, R. Wang, S. Bi, and G. Qi Forget the token and pixel: rethinking gradient ascent for concept unlearning in multimodal generative models. In Findings of the Association for Computational Linguistics: ACL 2025, p. 12179–12200. Cited by: Token-level Unlearning. Li et al. (2025b) N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A. Dombrowski, S. Goel, G. Mukobi, et al. The wmdp benchmark: measuring and reducing malicious use with unlearning. In International Conference on Machine Learning, p. 28525–28550. Cited by: Table 1, Table 1, Table 2, Datasets, Baselines. Lin et al. (2025) Z. Lin, T. Liang, J. Xu, Q. Liu, X. Wang, R. Luo, C. Shi, S. Li, Y. Yang, and Z. Tu Critical tokens matter: token-level contrastive estimation enhances llm’s reasoning capability. In International Conference on Machine Learning, p. 37906–37918. Cited by: Introduction. Liu et al. (2025) Z. Liu, S. Maharjan, F. Wu, R. Parikh, B. Bayar, S. H. Sengamedu, and M. Jiang Disentangling biased knowledge from reasoning in large language models via machine unlearning. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 6105–6123. Cited by: Introduction. Maini et al. (2024) P. Maini, Z. Feng, A. Schwarzschild, Z. C. Lipton, and J. Z. Kolter Tofu: a task of fictitious unlearning for llms. arXiv preprint arXiv:2401.06121. Cited by: Introduction, Datasets. Nguyen et al. (2025) T. T. Nguyen, T. T. Huynh, Z. Ren, P. L. Nguyen, A. W. Liew, H. Yin, and Q. V. H. Nguyen A survey of machine unlearning. ACM Transactions on Intelligent Systems and Technology 16 (5), p. 1–46. Cited by: Introduction. Pawelczyk et al. (2024) M. Pawelczyk, S. Neel, and H. Lakkaraju In-context unlearning: language models as few-shot unlearners. In International Conference on Machine Learning, p. 40034–40050. Cited by: Table 1, Table 1, Table 2, Baselines. Pu et al. (2026) J. Pu, M. Shi, X. Ren, Y. Wang, X. Zhang, Z. Wang, and K. She Decoding-unlearning: fact forgetting via entropy-guided inference. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 39834–39860. Cited by: Introduction. Ranjan et al. (2026) R. Ranjan, U. Grover, X. Lin, and A. Polyzou Razor: ratio-aware layer editing for targeted unlearning in vision transformers and diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 7998–8008. Cited by: Introduction. Sanyal and Mandal (2025) D. Sanyal and M. Mandal Agents are all you need for llm unlearning. arXiv preprint arXiv:2502.00406. Cited by: Table 1, Table 1, Table 2, Baselines. Shah et al. (2025) R. S. Shah, J. Huang, K. Murugesan, N. Baracaldo, and D. Yang The unlearning mirage: a dynamic framework for evaluating llm unlearning. In Second Conference on Language Modeling, Cited by: Introduction. Shi et al. (2024) D. Shi et al. Large language model safety: a holistic survey. CoRR abs/2412.17686. External Links: Link Cited by: Introduction. Shi et al. (2025) W. Shi, J. Lee, Y. Huang, S. Malladi, J. Zhao, A. Holtzman, D. Liu, L. Zettlemoyer, N. Smith, and C. Zhang Muse: machine unlearning six-way evaluation for language models. In International Conference on Learning Representations, Vol. 2025, p. 27797–27818. Cited by: Introduction. Shi et al. (2024) W. Shi, S. Malladi, J. Zhao, A. Holtzman, D. Liu, L. Zettlemoyer, N. A. Smith, and C. Zhang MUSE: machine unlearning six-way evaluation for language models. arXiv preprint arXiv:2407.06460. Cited by: Datasets. Sun et al. (2025) H. Sun, T. Zhu, W. Chang, and W. Zhou Generative adversarial networks unlearning. IEEE Transactions on Dependable and Secure Computing. Cited by: Introduction. Tan et al. (2025) C. Tan, Y. Qu, X. Li, H. Zhang, S. Cui, C. Chen, and L. Gao Wisdom is knowing what not to say: hallucination-free llms unlearning via attention shifting. NeurIPS 2025. Cited by: Introduction, Token-level Unlearning, Table 1, Table 1, Table 2, Baselines. Thudi et al. (2022) A. Thudi, G. Deza, V. Chandrasekaran, and N. Papernot Unrolling sgd: understanding factors influencing machine unlearning. In EuroS&P 2022, p. 303–319. Cited by: Sequence-level Unlearning. Tran et al. (2025) T. Tran, R. Liu, and L. Xiong Tokens for learning, tokens for unlearning: mitigating membership inference attacks in large language models via dual-purpose training. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 22872–22888. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: Token-level Unlearning. Wang et al. (2025a) L. Wang, X. Zeng, J. Guo, K. Wong, and G. Gottlob Selective forgetting: advancing machine unlearning techniques and evaluation in language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 843–851. Cited by: Sequence-level Unlearning. Wang et al. (2025b) Y. Wang J. Wei et al. LLM unlearning via loss adjustment with only forget data. In The Thirteenth International Conference on Learning Representations, Cited by: Sequence-level Unlearning. Wang et al. (2026) Z. Wang, J. Guo, J. Pu, H. Pu, M. Yang, X. Chen, J. Ou, W. Li, G. Luo, and W. Tian CAP: controllable alignment prompting for unlearning in LLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Cited by: Introduction. Yang et al. (2025) T. Yang, L. Dai, X. Wang, M. Cheng, Y. Tian, and X. Zhang Cliperase: efficient unlearning of visual-textual associations in clip. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 30438–30452. Cited by: Sequence-level Unlearning. Yu et al. (2025) M. Yu et al. UniErase: unlearning token as a universal erasure primitive for language models. arXiv preprint arXiv:2505.15674. Cited by: Introduction, Introduction, Token-level Unlearning, Table 1, Table 1, Table 2, Baselines. Yuan et al. (2025) H. Yuan, Z. Jin, P. Cao, Y. Chen, K. Liu, and J. Zhao Towards robust knowledge unlearning: an adversarial framework for assessing and improving unlearning robustness in large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 25769–25777. Cited by: Token-level Unlearning. Zhang et al. (2024) R. Zhang, L. Lin, Y. Bai, and S. Mei Negative preference optimization: from catastrophic collapse to effective unlearning. In First Conference on Language Modeling, External Links: Link Cited by: Sequence-level Unlearning, Table 1, Table 1, Table 2, Baselines. Zhao et al. (2025a) H. Zhao, C. Yuan, F. Huang, X. Hu, Y. Zhang, A. Yang, B. Yu, D. Liu, J. Zhou, J. Lin, et al. Qwen3guard technical report. arXiv preprint arXiv:2510.14276. Cited by: Introduction. Zhao et al. (2024) K. Zhao, M. Kurmanji, G. Bărbulescu, E. Triantafillou, and P. Triantafillou What makes unlearning hard and what to do about it. Advances in Neural Information Processing Systems 37, p. 12293–12333. Cited by: Sequence-level Unlearning. Zhao et al. (2025b) S. Zhao, X. Wu, C. T. Nguyen, Y. Jia, M. Jia, F. Yichao, and L. A. Tuan Unlearning backdoor attacks for llms with weak-to-strong knowledge distillation. In Findings of the Association for Computational Linguistics: ACL 2025, p. 4937–4952. Cited by: Introduction. Zhuang et al. (2025) H. Zhuang, Y. Zhang, K. Guo, J. Jia, G. Liu, S. Liu, and X. Zhang SEUF: is unlearning one expert enough for mixture-of-experts llms?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 8664–8678. Cited by: Introduction.