Paper deep dive
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing
Anthony Hughes, Vasisht Duddu, N. Asokan, Nikolaos Aletras, Ning Ma
Models: GPT-2 Large, GPT-2 Medium, GPT-2 Small, GPT-2 XL, Llama-3.2-1B, Qwen3
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/11/2026, 1:17:07 AM
Summary
The paper introduces PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel method to mitigate PII leakage in language models by identifying and editing specific computational circuits responsible for memorizing sensitive information. By using Edge Attribution Patching with Integrated Gradients (EAP-IG), PATCH identifies critical attention heads and edges, allowing for targeted modifications that achieve superior privacy-utility trade-offs compared to traditional defenses like differential privacy or data scrubbing.
Entities (5)
Relation Signals (3)
PATCH â mitigates â PII
confidence 95% · PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII circuits to reduce leakage.
EAP-IG â identifies â PII leakage circuits
confidence 92% · EAP-IG is a mechanistic interpretability technique designed to efficiently discover circuits within Transformer models.
PATCH â combineswith â Differential Privacy
confidence 90% · PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as 0.01%.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Language models (LMs) may memorize personally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mechanisms such as differential privacy (DP) reduce this leakage, but incur large drops in utility. Based on a comprehensive study using circuit discovery to identify the computational circuits responsible PII leakage in LMs, we hypothesize that specific PII leakage circuits in LMs should be responsible for this behavior. Therefore, we propose PATCH (Privacy-Aware Targeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII circuits to reduce leakage. PATCH achieves better privacy-utility trade-off than existing defenses, e.g., reducing recall of PII leakage from LMs by up to 65%. Finally, PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as 0.01%. Our analysis shows that PII leakage circuits persist even after the application of existing defense mechanisms. In contrast, PATCH can effectively mitigate their impact.
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
50,726 characters extracted from source content.
Expand or collapse full text
PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware T argeted Circuit PatcHing Anthony Hughes 1 , Vasisht Duddu 2 , N. Asokan 2 , Nikolaos Aletras 1 , Ning Ma 1 1 University of Sheffield, 2 University of Waterloo ajhughes3, n.ma, n.aletras@sheffield.ac.uk, vasisht.duddu@uwaterloo.ca, asokan@acm.org Abstract Language models (LMs) may memorize per- sonally identifiable information (PII) from training data, enabling adversaries to extract it during inference. Existing defense mecha- nisms such as differential privacy (DP) reduce this leakage, but incur large drops in utility. Based on a comprehensive study using circuit discovery to identify the computational circuits responsible PII leakage in LMs, we hypothesize that specific PII leakage circuits in LMs should be responsible for this behavior. Therefore, we propose PATCH (Privacy-AwareTargeted Circuit PatcHing), a novel approach that first identifies and subsequently directly edits PII cir- cuits to reduce leakage. PATCH achieves better privacy-utility trade-off than existing defenses, e.g., reducing recall of PII leakage from LMs by up to65%. Finally, PATCH can be combined with DP to reduce recall of residual leakage of an LM to as low as0.01%. Our analysis shows that PII leakage circuits persist even after the application of existing defense mechanisms. In contrast, PATCH can effectively mitigate their impact. 1 1 Introduction Language models (LMs) have demonstrated re- markable advances (Gemma-Team et al., 2024; Grattafiori et al., 2024), yet their tendency to mem- orize training data poses privacy risks (Kandpal et al., 2022; Buzaglo et al., 2023; Duan et al., 2024; Hayes et al., 2025b). In particular, prior work has shown that LMs can memorize and reproduce per- sonally identifiable information (PII) from their training data (Huang et al., 2022; Kim et al., 2023; Nakka et al., 2024; Borkar et al., 2025). This makes it possible for adversaries with black-box access to a model to expose such information (Lukas et al., 2023). 1 Code and data are publicly available athttps://github. com/ssg-research/pii-patch/ Figure 1: Mitigating PII leakage through circuit analysis of LMs fine-tuned on PII-containing documents (A) with and without privacy defenses. (B) We use PATCH to discover PII leaking circuits. Finally, (C) PATCH then edits the discovered circuits, reducing PII leaks. A range of defense mechanisms have been pro- posed to mitigate PII leakage (Kerrigan et al., 2020; Wu et al., 2024). Data processing defenses, such as scrubbing, remove PII from data to reduce leak- age (PilĂĄn et al., 2022; Mosallanezhad et al., 2019; Lukas et al., 2023). Train time defenses such as dif- ferential privacy (DP) modify the models learning process to limit the influence of individual train- ing examples while providing formal guarantees about the leakage (Kerrigan et al., 2020; Li et al., 2022; Ponomareva et al., 2022; Yu et al., 2022). Fi- nally, post-training defenses, including model edit- ing which remove or suppress PII-related knowl- edge from the LMs (Wu et al., 2023; Chen et al., 2024; Wu et al., 2024). However, these defenses demonstrate poor privacy-utility trade-offs (Lukas et al., 2023; Wu et al., 2024), unintentionally in- crease susceptibility to other PII types (Wu et al., 2024; Borkar et al., 2025), or can be circumvented by adversaries (Xin et al., 2024). To address these limitations, we proposePrivacy- AwareTargetedCircuit PatcHing (PATCH), which allows editing relevant PII circuits to minimize leakage. We first apply mechanistic interpretability to discover the internal computational structures (or âcircuitsâ) in a LM responsible for PII leakage. arXiv:2510.07452v2 [cs.CR] 26 Feb 2026 We then âpatchâ or edit these circuits to mitigate the leakage. Our contributions are: 1. A novel targeted circuit editing method, PATCH, that provides greater privacy-utility trade-off than existing defense mechanisms (§3, §5) 2.We identify and characterize the specific cir- cuits responsible for leaking different PII types, like names, locations and race, reveal- ing that distinct attention heads within these circuits that influence PII leakage. (§6). 3.An extensive ablation study of PATCH, further demonstrating its robustness across settings. (§7). 2 Related Work PII Leakage in LMs. Memorized information from the training data can be shared via model weights from which an adversary can extract PII (Lukas et al., 2023; Yu et al., 2023; Staab et al., 2024; Hayes et al., 2025b,a). For instance, Lukas et al. (2023) demonstrate PII leakage attacks against GPT2 (Radford et al., 2019) models, where given access to LMs, adversaries can sample re- sponses to extract sensitive information. PII leak- age vulnerabilities have been observed across di- verse contexts, including pretrained models (Nakka et al., 2024; Staab et al., 2024; Kim et al., 2023; Panda et al., 2024; Huang et al., 2022) and special- ized applications (Hughes et al., 2024; Mireshghal- lah et al., 2024; Xiao et al., 2024), indicating that privacy risks persist throughout the modelâs lifecy- cle and across diverse task domains. Defenses against PII Leakage. A simple de- fense method is to remove PII from the data before training (âscrubbingâ) (PilĂĄn et al., 2022; Mosal- lanezhad et al., 2019; Lukas et al., 2023). However, such approaches are significantly expensive (Wu et al., 2024), provide poor privacy-utility trade- offs (Wu et al., 2023; Lukas et al., 2023; Wu et al., 2024), and may allow an adversary to deduce per- sonal attributes through auxiliary information (Xin et al., 2024). Alternatively, DP methods (Feyisetan et al., 2020; Kerrigan et al., 2020; Li et al., 2022; Shi et al., 2022; Lee and SĂžgaard, 2023) offer for- mal guarantees of privacy by injecting noise during training, effectively masking individual samples observed by the LM. Yet, this often results in a reduction in utility (Lukas et al., 2023; Wu et al., Algorithm 1 PATCH: Privacy-Aware Targeted Cir- cuit PatcHing Input:Model Ë M, PII typesP, percentile thresholdp, model editing method A, circuit discovery CD, private data D Output: Privacy-enhanced model Ë M 1: for each PII type P i â P do 2: S â prompt_builder(D) â· Generate prompts 3: C i â CD( Ë M,P i ,D,S) â· Extract circuit 4: Ï i â percentile(s (i) e : eâ C i ,p) â· Compute Threshold 5: E high i â eâ C i : s (i) e â„ Ï i â· Select edges 6: end for 7: E shared â T i E high i â· Identify shared edges 8: Ë M â A( Ë M,E shared ) â· Patch 9: return Modified model Ë M 2024). Recent empirical defenses include identify- ing neurons that are responsible for memorizing PII and patching with steering vectors can reduce that memorization (Wu et al., 2023; Chen et al., 2024; Wu et al., 2024). However, Wu et al. (2024) indi- cate that such methods suffer from poor privacy- utility trade-offs due to limited components that can be edited, with the side effect of increasing the leakage of other PII types. 3PATCH: Privacy-Aware Targeted C ircuit PatcHing 3.1 Problem Formulation Given a pre-trained modelMand a private dataset Dthat contains PIIP, fine-tuningMonDre- sults to a model Ë Mthat is exposed to PII. We also assume an adversaryAwith black-box access to Ë M.Aseeks to infer specific PII types that are ob- servable fromD via prompting. Finally, a defense mechanismDFaims to reduce the effectiveness ofAon extracting PII from Ë M, while maintaining the utility (i.e., performance) ofM(Lukas et al., 2023).DFcan be applied before onD, during fine-tuning onD or post-hoc, after training. 3.2 Motivation We hypothesize that there is a set of unique ele- ments (e.g., attention heads or outputs from heads to other parts of the Transformer block) within Ë Mresponsible for PII leaks. By identifying and modifying these elements, we can reduce PII leak- age while maintaining utility. More specifically, given a circuit discovery mechanismCD, a private datasetD, and a modelM, PATCH consists of the following three steps (detailed in Algorithm 1). 3.3 Step 1: Generate Prompts for Circuit Discovery. To identify circuits responsible for PII leakage across all PII types, we employ Edge Attribution Patching with Integrated Gradients (Hanna et al., 2024, EAP-IG) as our circuit discovery mecha- nismCD. EAP-IG is a mechanistic interpretability technique designed to efficiently discover circuits within Transformer models. We select EAP-IG for its effectiveness at identifying faithful circuits (Mueller et al., 2025). EAP-IG requires constructing prompts represent- ing PII leakage so we can extract a circuit that is representative of that behavior. We construct pairs of âcleanâ and âcorruptâ prompts. The clean ver- sion contains correct PII values and the corrupted version has these values replaced with alternatives from the same type. Using a private datasetDthat is tagged with a target PII typeP i , we select1, 000 unique text spans containing that PII elementP i . The PII span is then replaced (âcorruptedâ) with PII that is semantically similar to that entity. Ex- amples generated from a legal dataset (Chalkidis et al., 2019) are shown in Table 1. 3.4 Step 2: Extract PII Leakage Circuits EAP-IG operates by comparing model behavior on clean prompts against corrupted variants where specific PII tokens have been altered. The method analyzes individual attention heads to measure two properties: (1) how sensitive each componentâs ac- tivations are to PII token corruption, and (2) how important each component is for correctly predict- ing PII values. Heads scoring high on both di- mensionsâbeing necessary for accurate PII pre- dictionâare identified as part of the PII leakage circuit. Given Ë MandD, EAP-IG outputs a circuit C i containing nodes that represent individual atten- tion heads and edgesâthat indicate information flow between nodes. Each node (attention head) and edge (connection to other nodes) receives an importance scores (i) e quantifying their contribu- tion to PII leaking behavior. In our experiments, we use the default hyperparameters recommended by Hanna et al. (2024). 3.5 Steps 3 and 4: Compute Threshold and Perform Edge Selection We hypothesize that high-scoring edges indicate critical PII leakage pathways. Therefore, we aim to isolate those critical edges for patching. PII TypeOriginalCorrupted NameâMr.JohnSmith vs. The Stateâ âMr.HeidiSmith v. The Stateâ Location âThe appellant was ar- rested inBerlin .â âThe appellant was ar- rested inNew York .â RaceâThis case concerns a Romanian national." "This case concerns a Turkish national.â Table 1: Examples from the European Court of Human Rights (ECHR) dataset (Chalkidis et al., 2019), illustrat- ing PII types used in our original and corrupted prompts. Purplerepresents a original token,redrepresents a corrupted token. Given a circuit discovery mechanismCDthat has scored the edges from a circuitC i within Ë M , we require a mechanism to select the most influen- tial edges. We first compute a thresholdÏfor each PII circuitC i , given a specified percentilep. Next, we select the high-scoring edges whose scores are equal to or greater than the thresholdÏ. We repeat this process for each PII type to obtainE high i for all circuitsC i . Following prior work (Wang et al., 2025), we select the high-scoring edges using percentile thresholds of 95% and 99%. 3.6 Step 5: Identify Shared PII Edges We aggregate edge scores across multiple PII types rather than treating each type independently, assum- ing that PII leakage has common pathways through- out models. Therefore, we computeE shared as the intersection of high-scoring edgesE high i across all PII circuit types C i . 3.7 Step 6: Patch PII Edges Given a set of shared edgesE shared from Ë M, we require a model editing mechanismAthat alters those high-importance edges such that their influ- ence on the modelâs behavior is reduced. We exper- iment with zero ablation, i.e.,setting edge weights to zero (Olah et al., 2020; Pochinkov et al., 2024); and mean ablation, i.e.,replacing edge weights with their mean values (Chan et al., 2022; Wang et al., 2023), following their successful application in prior work (Bi et al., 2025). 4 Experimental Setup 4.1 PII Types and Private Data We select names, locations, and race as a set of PII types following prior work (PilĂĄn et al., 2022; Hughes et al., 2024; Kim et al., 2024). We use the European Court of Human Rights (ECHR) dataset (Chalkidis et al., 2019), which contains PII related to appellants and others involved in legal cases. To automatically identify PII across all train, val- idation and test sets, we apply the FLAIR name- entity-recognition (NER) tool. 2 FLAIR achieves an accuracy of approximately86%(Yermilov et al., 2023). Details of the NER label classes are pro- vided in Table 4. Detailed corpus and PII statistics can be found in Appendix A. 4.2 Circuit Discovery Prompts Using our private dataset, ECHR, that is tagged with the three PII types, we select unique text spans that containing the tagged PII element. We follow- ing prior work (Wang et al., 2025) and select 1,000 spans for each of the target PII types. The PII el- ement in each span is then âcorruptedâ with PII that is semantically similar to that entity. We se- lect entities using the faker library. 3 The clean and corrupted prompts are used to score circuit edges across all PII types (Table 1). 4.3 Base Models We use several open-weight LMs such as GPT2- Small (117M), GPT2-Medium (345M), and GPT2- Large (774M) following prior work (Lukas et al., 2023; Wu et al., 2024), as well as more recent mod- els such as Llama-3.2-1B (Grattafiori et al., 2024), Qwen3-0.6B and Qwen3-1.7B (Team, 2025). 4.4 Fine-tuning Target Models To obtain LMs exposed to PII (target), we fine- tune all base models on the private ECHR dataset. We conduct our experiments using Hugging Face 4 for all models. The max sequence length is set to 512. All experiments on open-weight models are performed on one to four NVIDIA H100 GPUs. Fine-tuning uses a batch size of 8, the AdamW optimizer (Loshchilov and Hutter, 2019), and a lin- ear learning rate scheduler. Each baseline, DP and scrubbed defended model is trained for 4 epochs. 4.5 Adversary To emulate an adversary, we sample approximately four million tokens from each target LM with and without defenses. Each query begins with an empty prompt, then we issue 10,000 queries to each model, generating sequences of256tokens from empty prompts using top-ksampling withk = 40 of which we apply random sampling. 2 FLAIR: https://github.com/flairNLP/flair 3 https://faker.readthedocs.io/en/master/ 4 https://w.huggingface.co To control for baseline PII leakage present in the base pretrained model, we establish a reference distribution by sampling 13 million tokens (50, 000 queries) from it prior to fine-tuning. Any PII in- stances appearing in this baseline are excluded from our leakage measurements, ensuring we only attribute leakage to the fine-tuning process rather PII aquired during pretraining. We repeat the attack three times to quantify variance. This is following prior work in PII leakage (Lukas et al., 2023). 4.6 Metrics Privacy Leakage. To assess the adversary suc- cess in leaking PII, we directly compare verba- tim LM outputs against FLAIR annotations, which serve as ground truth. Precision measures the pro- portion of PII in the modelâs output is PII that was also present in the training data. Recall indicates how much of the total PII observable in the training data is exposed. Faithfulness.We assess how accurately a discov- ered circuit represents the causal mechanism under- lying PII leakage. A circuit is considered faithful if ablating any components outside the circuit does not affect the modelâs performance, indicating that the circuit alone explains the behavior (Prakash et al., 2024; Hanna et al., 2024). We report the normalized faithfulness as: P method â P corrupted P baseline â P corrupted P method is the performance of a circuit after ap- plying EAP-IG,P baseline is the original model per- formance, andP corrupted is the corrupted baseline performance. The score ranges from0(no resem- blance to the original model) to1(full recovery of performance). Circuit Overlap.To identify how PII circuits in- teract and where editing those circuits is optimal, we require a circuit overlap measure. Following Hanna et al. (2024), we use the overlap metric based on high scoring circuit elements that meet a percentile threshold. For each PII typeP i â T, we extract circuitC i with node or edge scoress (i) e using EAP-IGCD. We then compute a threshold Ï i and identify high-importance edgesE high i . The overlap between two PII circuitsC i andC j is cal- culated using the Jaccard index: Overlap(C i ,C j ;Ï) = |E high i â© E high j | |E high i âȘ E high j | Ă 100% Ïdenotes the percentile threshold. High overlap indicates shared computational pathways across PII types. We analyze edge overlap at percentile thresholds ofpâ95, 99following (Wang et al., 2025), with higher thresholds selecting fewer but more critical edges for editing. Utility. We use perplexity over the ECHR test set, similar to prior work (Lukas et al., 2023; Wu et al., 2023; Chen et al., 2024; Wu et al., 2024), to evaluate the impact on LM utility before and after applying PATCH and baseline defense mechanisms. 4.7 Defense Baselines We compare with the following defenses: APNEAP. Augmented Privacy Neuron Editing via Activation Patching (Wu et al., 2024) identifies and ablates individual neurons responsible for PII leakage. This is the current state-of-the-art editing- based defense which outperforms prior editing ap- proaches (Wu et al., 2023; Chen et al., 2024). DP.We fine-tune LMs using differentially private stochastic gradient descent (Abadi et al., 2016, DP- SGD). We consider different privacy loss parameter Δ=8, 4, 1 where a lowerΔindicates stronger pri- vacy guarantee. For DP training, we use the fastDP library (Bu et al., 2022, 2023). Following prior work (Lukas et al., 2023), each model is trained us- ing DP-SGD for 4 epochs using (Δ,ÎŽ) = (8, 4, 1, 1N)whereNis the size of the training dataset. We use a maximum per-sample gradient norm of 1. Scrubbing. We first remove PII from the data, and then use it for fine-tuning LMs. All informa- tion related to the selected PII types (names, loca- tions, race) is removed. Again, we use an NER tool, FLAIR, to identify any spans containing PII, we then redact those spans. Redaction means re- placing the identified span with a masking token. Models are then trained on the resulting scrubbed documents. PATCH. We evaluate two variants of our pro- posed PATCH approach. First, PATCH-Baseline applies our method to a modelMtrained without any existing privacy defenses. Second, PATCH- DP(Δ = 8) applies our method to a modelMthat has been fine-tuned with DP(Δ = 8). 5 Results Figure 2 presents privacy-utility tradeoffs for PATCH against three baseline defenses across all models. We measure PII leakage with precision and recall, and utility using perplexity. We evalu- ate two variants of our method: PATCH-BASELINE applied to non-private models, and PATCH-DP(Δ) applied to a model fine-tuned with DP at Δ = 8. 5 PATCH-Baseline achieves strong privacy-utility tradeoffs.Across all models, PATCH-BASELINE reduces PII extraction precision by40%â90% relative to undefended baselines. For example, Llama-3.2-1B precision decreases from60%to0.5, while incurring only modest utility costs; GPT2- Medium perplexity increases from approximately9 to10. Recall consistently decreases by80%â86% across architectures, with GPT2-Medium decreas- ing from8.5%to1.2%and GPT2-Large from ap- proximately11%to1.8%while maintaining a util- ity within1point of the baseline perplexity. These reductions demonstrate that PII leakage localizes to identifiable circuit edges rather than distributing across the entire parameter space. The consistency across architectures indicates that PII leaks via spe- cific mechanisms. PATCH-DP provides maximal privacy at varied utility costs.Combining circuit patching with DP yields the strongest privacy protection, achieving recall< 1%across all models, but with model- dependent utility impact. For smaller models, PATCH-DP(Δ = 8) maintains reasonable utility: on GPT2-Small , perplexity increases to19.5(vs. 18.3for DP(Δ = 8) alone, and10.3for Base), while recall drops to< 0.5%compared to2.0%for DP alone. Similar patterns emerge for Llama-3.2- 1B where perplexity is32compared to27for the DP baseline, however recall approaches0%. This demonstrates good protection, but fails to maintain utility. Be careful of precision increases. An impor- tant finding in PATCH-DP(Δ = 8) in GPT2-Large , is that it successfully reduces recall, however it exhibits concerning spikes in precision with stable perplexity, particularly evident in GPT2- Large where precision reaches⌠80while recall drops to⌠1. This tells us the leaks have become highly accurate. For adversaries mounting training data extraction attacks, high precision and low re- call suggest a vulnerability: the model may leak in- frequently, but successful extractions yield genuine, 5 We ran experiments with a DP fine-tuned model atΔ =4, 1, however we obtained poor utility. Table of experiments is available in Appendix C. Figure 2: Comparison of PATCH with other defenses: lower perplexity, precision and recall scores are preferred (closer to lower left corner). trustworthy PII. This pattern is typically present in the baseline model. PATCH-Baseline corrects this issue, however it is important to note that a defense may simultaneously increase model susceptibility. Comparison to existing defenses. Undefended baseline models exhibit substantial PII leakage, with precision ranging from38.6%to60.8%and recall from3.0%to11.0%across architectures, confirming the severity of leaks in models fine- tuned on sensitive data. Among existing defense mechanisms, DP provides strong privacy protec- tions, reducing recall to†6%across all mod- els. For example, Llama-3.2-1B decreases to0.5%. However, these models incur substantial utility costs with perplexity increases of1.5to7.9times the baseline perplexity. APNEAP (Wu et al., 2024) maintains utility within+1.0perplexity of baseline models but provides limited privacy protection, re- ducing extraction precision by only⌠10-18% and recall by⌠1-3%. Practical implications. Overall, our results demonstrate that circuit-based interventions can provide quantifiable privacy-utility tradeoffs com- plementary to formal DP guarantees. For deploy- ment scenarios where moderate privacy protection suffices and utility is paramount, PATCH-Baseline offers⌠1.5-8% reduction in recall with minimal perplexity overhead. For high-security applications requiring maximal privacy, PATCH-DP(Δ = 8) can achieve†1%in precision and recall, though prac- titioners should be conscious of utility, as there were cases, Qwen3-1.7B , where utility was not preserved. These differences across models under- scores the need for careful evaluation of defense mechanism impacts across model architectures. Privacy budget considerations.Our circuit iden- tification process analyzes model activations on the training data to locate relevant circuits for PII leakage. This data-dependent analysis is not ac- counted for in the privacy hyperparameter (Δ) that was specified during DP fine-tuning. For scenarios requiring strict formal guarantees, circuit identifi- cation should be performed on a separate dataset, after which patching operations would preserve the training-time privacy guarantees. In this work, we demonstrate valuable empirical privacy protection through targeted circuit interventions. 6 Analysis of PII Leakage Circuits 6.1 PII leakage and Circuits Figure 3 presents the results of the PII leakage and circuit discovery, for three PII types: names, location, and race. We find precision is variable with regard to model architectures, however, within architectures, such as GPT2 and Qwen3, recall increases with model size. This corroborates prior work (Lukas et al., 2023). Among other models, Llama-3.2- 1B has the highest precision indicating that the PII produced by this model is more likely to be from the training data. In contrast, GPT2-Large has the highest recall exposing the most PII in training data. NamesLocationsRace 0.0 10.0 20.0 30.0 40.0 50.0 60.0 70.0 Precision (%) NamesLocationsRace 0.0 2.0 4.0 6.0 8.0 10.0 12.0 Recall (%) NamesLocationsRace 0.0 20.0 40.0 60.0 80.0 100.0 Faithfulness (%) GPT2-SmallGPT2-MediumQwen3-0.6BGPT2-LargeLlama-3.2-1BQwen3-1.7B Figure 3: Results for PII leakage and faithfulness: we present the precision (left), and recall (middle) of the PII extracted, and finally, the faithfulness of the discovered PII circuits (right). Furthermore, the PII leakage, as seen with preci- sion and recall, varies across different PII types and across models. Regardless of the model or PII type faithfulness is high, indicating that we can reliably identify the circuits for PII leakage. Faithfulness scores (above75%) are similar to prior circuit dis- covery work (Wang et al., 2025). Overall, we can reliably identify circuits for PII leakage, and this remains high across all models. 6 6.2 PII Circuit Overlaps Table 2 shows the results for PII circuit overlap. This allows us to understand the structure of the individual PII circuits, and how editing these may impact an attack. Model PII Circuit Overlaps (nodes/edges (%)) Name-LocationName-RaceLocation-Race GPT2-Small79 / 3988 / 4877 / 38 GPT2-Medium66 / 2371 / 2672 / 24 Qwen3-0.6B62 / 5061 / 4154 / 41 GPT2-Large70 / 3069 / 2465 / 24 Llama-3.2-1B82 / 4482 / 4084 / 40 Qwen3-1.7B87 / 2587 / 2496 / 49 Table 2: Results of circuit overlap analysis across PII types and models: we present circuit overlap (%) between faithful PII circuits (format: ânodes/edgesâ). Minimal edge overlap across PIIs types.Com- paring connectivity patterns of circuits responsible for different PII types reveals substantial differ- ences in the edge overlap. We observe a consis- tently low overlap between the edges of PII circuits across all models. GPT2-Small shows38.8%over- lap between race and names,37.6%between race and locations, and48.3%between names and loca- tions, while GPT2-Large and Qwen3-1.7B exhibit even lower overlap of⌠25%. This suggests that 6 We also analyzed circuit overlaps across defense baselines. Overlap metrics are available in Appendix E. different PII circuits have distinct pathways. Low edge overlap combined with high node overlap in larger models indicates greater circuit specializa- tion with increased capacity. This also explains the observations in prior node ablation work (Wu et al., 2023, 2024; Borkar et al., 2025) where minimizing leakage of one PII through node editing, increases others. This pattern also suggests that LMs do not maintain entirely separate mechanisms for retriev- ing PII, but instead rely on a smaller set of shared attention patterns across PII types. 6.3 Influential Edges in PII Circuits Given the low overlapping scores observed among edges, we investigate their individual influence. We observe the scores generated by the circuit discov- ery method (EAP-IG). We present heatmaps for a subset of the models attention layers and heads in Figure 5. 7 We find that circuit discovery finds distinct influence patterns of attention head edges across architectures. GPT2-Medium displays a more distributed pattern with notable activation in layers14-16across multiple heads. Llama-3.2- 1B shows sparse but intense activation primarily in later layers, with early layer concentration with high scores in layer9. This is also visible in Qwen3- 0.6B and Qwen3-1.7B where layer0is highly in- fluential. Interestingly, some layer and heads are very pronounced. We identify layer11head9as highly influential in both Qwen3 models. Overall, key edges vary across architectures but systematic circuits that drive PII leakage exist across all mod- els. 7 Ablation Study To be better understand PATCH, we perform attacks under different hyperparameter settings. We report 7 See full heatmaps of each modelâs attention layers and heads in Appendix D. Figure 4: Results of PATCH across varying hyperparameters: we compare PATCH-Baseline with PATCH- DP(Δ = 8) with alternating ablation strategies, zero and mean, and edge thresholds, 95 and 99. 0123456789 101112131415 9 11 14 16 Layers GPT2-Medium Baseline 0123456789 101112131415 9 11 14 16 Layers Qwen3-0.6B Baseline 02468 1012141618202224262830 Attention Heads 9 11 14 16 Layers Llama-3.2-1B Baseline .005 .010 0 .050 0 .010 Figure 5: Influential PII Circuit Components: aver- age EAP-IG scores for each attention head in each layer, across all identified PII circuits. the results in Figure 4. As indicated in Section 3, we use95%and99%percentile threshold on the importance scores of the edges. We also evalu- ate PATCH with zero and mean ablation with both thresholds. This helps evaluate whether there are other configurations which can improve privacy- utility trade-offs. Model editing strategies show no universal dif- ferences.Our comparative analysis of mean ver- sus zero ablation strategies reveals no systematic preference across model architectures and scales. For smaller models such as GPT2-Small , zero ab- lation configurations cluster in regions of lower perplexity with marginally reduced F1-scores, sug- gesting minimal functional differentiation between ablation methods. GPT2-Medium and Qwen3- 1.7B mean and zero ablation variants are without clear separation. Practitioners should evaluate both strategies during post-training. Optimal ablation appears model specific. Threshold selection has model-specific effects. The thresholds95%and99%comparison reveals model-specific patterns without consistent trends. GPT2-Small shows threshold 99 configurations in- curring higher perplexity costs for equivalent pri- vacy gains, while GPT2-Large exhibits the opposite relationship where stricter thresholds maintain com- parable utility. Qwen3-0.6B and Qwen3-1.7B ar- chitectures demonstrate negligible threshold sensi- tivity, with both settings producing similar privacy- utility trade-offs. This heterogeneity indicates that optimal threshold selection requires model-specific calibration. 8 Conclusion We introduced PATCH, a method grounded in mech- anistic interpretability for mitigating PII leakage in LMs through targeted editing. Our approach identi- fies shared computational pathways responsible for PII memorization across PII types and selectively ablates high-importance edges to reduce privacy risks. Empirical results demonstrate that PATCH achieves superior privacy-utility trade-offs com- pared to prior work. Our extensive experiments verify the effectiveness of PATCH across multiple model scales and PII categories compared to non- private fine-tuned baselines. These findings demon- strate that targeted circuit interventions can provide effective privacy protection while preserving the modelâs core computational pathways. Limitations In this study, we empirically demonstrate that our method substantially improves open-weight LMs in mitigating PII leaks, however we acknowledge our evaluation is limited to a subset of PII types. We hope to extend our proposed method to a broader set of PII in the future. Finally, experiments have not been conducted on large models, due to circuit discovery requiring intensive resources and time, however as new methods arise they can integrate into our proposed method. Acknowledgments AH is supported by the Centre for Doctoral Train- ing in Speech and Language Technologies (SLT) and their Applications funded by UK Research and Innovation [EP/S023062/1]. This work is sup- ported in part by Intel (in the context of Private AI consortium), and the Government of Ontario. Vasisht is supported by David R. Cheriton Scholar- ship, Cybersecurity and Privacy Excellence Gradu- ate Scholarship, and an IBM PhD Fellowship. Ad- ditionally, we acknowledge IT Services at The Uni- versity of Sheffield for the provision of services for High Performance Computing. Finally, views ex- pressed in the paper are those of the authors and do not necessarily reflect the position of the funding agencies. References Martin Abadi, Andy Chu, Ian Goodfellow, H. Bren- dan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep learning with differential pri- vacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Secu- rity, CCS â16, page 308â318, New York, NY, USA. Association for Computing Machinery. Jing Bi, Junjia Guo, Yunlong Tang, Lianggong Bruce Wen, Zhang Liu, Bingjie Wang, and Chenliang Xu. 2025. Unveiling visual perception in language mod- els: An attention head analysis approach. In Proceed- ings of the Computer Vision and Pattern Recognition Conference, pages 4135â4144. Jaydeep Borkar, Matthew Jagielski, Katherine Lee, Niloofar Mireshghallah, David A. Smith, and Christo- pher A. Choquette-Choo. 2025. Privacy ripple effects from adding or removing personal information in lan- guage model training. In ACL 2025 Student Research Workshop. Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, and George Karypis. 2022. Differentially private bias-term fine- tuning of foundation models. In Workshop on Trust- worthy and Socially Responsible Machine Learning, NeurIPS 2022. Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, and George Karypis. 2023. Differentially private optimization on large model at small cost. In International Con- ference on Machine Learning, pages 3192â3218. PMLR. Gon Buzaglo, Niv Haim, Gilad Yehudai, Gal Vardi, Yakir Oz, Yaniv Nikankin, and Michal Irani. 2023. Deconstructing data reconstruction:Multiclass, weight decay and general losses. In Advances in Neural Information Processing Systems, volume 36, pages 51515â51535. Curran Associates, Inc. Ilias Chalkidis, Ion Androutsopoulos, and Nikolaos Ale- tras. 2019. Neural legal judgment prediction in En- glish. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4317â4323, Florence, Italy. Association for Compu- tational Linguistics. Lawrence Chan, Adria Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishin- skaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. 2022. Causal scrubbing: A method for rigorously testing interpretability hypotheses. In AI Alignment Forum, volume 2. Ruizhe Chen, Tianxiang Hu, Yang Feng, and Zuozhu Liu. 2024. Learnable privacy neurons localization in language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 256â264, Bangkok, Thailand. Association for Computational Linguistics. Sunny Duan, Mikail Khona, Abhiram Iyer, Rylan Scha- effer, and Ila R. Fiete. 2024. Uncovering Latent Memories: Assessing Data Leakage and Memoriza- tion Patterns in Large Language Models. In ICML 2024 Workshop on LLMs and Cognition. Oluwaseyi Feyisetan, Borja Balle, Thomas Drake, and Tom Diethe. 2020. Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Pertur- bations. In Proceedings of the 13th International Conference on Web Search and Data Mining, pages 178â186, Houston TX USA. ACM. Gemma-Team and 1 others. 2024. Gemma 2: Im- proving open language models at a practical size. Preprint, arXiv:2408.00118. Aaron Grattafiori and 1 others. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783. Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. 2024. Have faith in faithfulness: Going beyond cir- cuit overlap when finding model mechanisms. In ICML 2024 Workshop on Mechanistic Interpretabil- ity. Jamie Hayes, Ilia Shumailov, William P Porter, and Aneesh Pappu. 2025a. Measuring memorization in RLHF for code completion. In The Thirteenth Inter- national Conference on Learning Representations. Jamie Hayes, Marika Swanberg, Harsh Chaudhari, Itay Yona, Ilia Shumailov, Milad Nasr, Christopher A. Choquette-Choo, Katherine Lee, and A. Feder Cooper. 2025b. Measuring memorization in lan- guage models via probabilistic extraction. In Pro- ceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Com- putational Linguistics: Human Language Technolo- gies (Volume 1: Long Papers), pages 9266â9291, Albuquerque, New Mexico. Association for Compu- tational Linguistics. Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022. Are large pre-trained language models leaking your personal information? In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2022, pages 2038â2047, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Anthony Hughes, Ning Ma, and Nikolaos Aletras. 2024. How private are language models in abstractive sum- marization? arXiv preprint arXiv:2412.12040. Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022. Deduplicating training data mitigates privacy risks in language models. In International Conference on Machine Learning, pages 10697â10707. PMLR. Gavin Kerrigan, Dylan Slack, and Jens Tuyls. 2020. Differentially private language models benefit from public pre-training. In Proceedings of the Second Workshop on Privacy in NLP, pages 39â45, Online. Association for Computational Linguistics. Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023. ProPILE: Probing privacy leakage in large language models. In Thirty-seventh Conference on Neural Information Processing Systems. Woojin Kim, Sungeun Hahm, and Jaejin Lee. 2024. Generalizing Clinical De-identification Models by Privacy-safe Data Augmentation using GPT-4. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21204â21218, Miami, Florida, USA. Association for Computational Linguistics. Seolhwa Lee and Anders SĂžgaard. 2023. Private Meet- ing Summarization Without Performance Loss. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Infor- mation Retrieval, pages 2282â2286, Taipei Taiwan. ACM. Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2022. Large language models can be strong differentially private learners. In International Conference on Learning Representations. Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In International Con- ference on Learning Representations. Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-BĂ©guelin. 2023. Analyzing leakage of personally identifiable information in language models. In 2023 IEEE Sym- posium on Security and Privacy (SP), pages 346â363. IEEE. Niloofar Mireshghallah, Maria Antoniak, Yash More, Yejin Choi, and Golnoosh Farnadi. 2024. Trust No Bot: Discovering Personal Disclosures in Human- LLM Conversations in the Wild. In First Conference on Language Modeling. Ahmadreza Mosallanezhad, Ghazaleh Beigi, and Huan Liu. 2019. Deep Reinforcement Learning-based Text Anonymization against Private-Attribute Inference. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lan- guage Processing (EMNLP-IJCNLP), pages 2360â 2369, Hong Kong, China. Association for Computa- tional Linguistics. Aaron Mueller, Atticus Geiger, Sarah Wiegreffe, Dana Arad, IvĂĄn Arcuschin, Adam Belfki, Yik Siu Chan, Jaden Fried Fiotto-Kaufman, Tal Haklay, Michael Hanna, Jing Huang, Rohan Gupta, Yaniv Nikankin, Hadas Orgad, Nikhil Prakash, Anja Reusch, Aruna Sankaranarayanan, Shun Shao, Alessandro Stolfo, and 4 others. 2025. MIB: A mechanistic interpretabil- ity benchmark. In Forty-second International Con- ference on Machine Learning. Krishna Kanth Nakka, Ahmed Frikha, Ricardo Mendes, Xue Jiang, and Xuebing Zhou. 2024. PII-compass: Guiding LLM training data extraction prompts to- wards the target PII via grounding. In Proceedings of the Fifth Workshop on Privacy in Natural Lan- guage Processing, pages 63â73, Bangkok, Thailand. Association for Computational Linguistics. Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom in: An introduction to circuits. Distill. Doi:10.23915/distill.00024.001. Ashwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang, and Prateek Mit- tal. 2024. Teach LLMs to phish: Stealing private information from language models. In The Twelfth International Conference on Learning Representa- tions. IldikĂł PilĂĄn, Pierre Lison, Lilja Ăvrelid, Anthi Pa- padopoulou, David SĂĄnchez, and Montserrat Batet. 2022. The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization. Computational Linguistics, 48(4):1053â1101. Nicholas Pochinkov, Ben Pasero, and Skylar Shibayama. 2024. Investigating neuron ablation in attention heads: The case for peak activation centering. arXiv preprint arXiv:2408.17322. Natalia Ponomareva, Jasmijn Bastings, and Sergei Vas- silvitskii. 2022. Training text-to-text transformers with privacy guarantees. In Findings of the Asso- ciation for Computational Linguistics: ACL 2022, pages 2182â2193, Dublin, Ireland. Association for Computational Linguistics. Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. 2024. Fine-tuning enhances existing mechanisms: A case study on en- tity tracking. In International Conference on Learn- ing Representations. ArXiv:2402.14811. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others. 2019. Language models are unsupervised multitask learn- ers. OpenAI blog, 1(8):9. Weiyan Shi, Aiqi Cui, Evan Li, Ruoxi Jia, and Zhou Yu. 2022. Selective Differential Privacy for Language Modeling. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 2848â2859, Seattle, United States. Association for Computational Linguistics. Robin Staab, Mark Vero, Mislav Balunovic, and Mar- tin Vechev. 2024. Beyond memorization: Violating privacy via inference with large language models. In The Twelfth International Conference on Learning Representations. Qwen Team. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388. Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Inter- pretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh Inter- national Conference on Learning Representations. Xu Wang, Yan Hu, Wenyu Du, Reynold Cheng, Benyou Wang, and Difan Zou. 2025. Towards understanding fine-tuning mechanisms of LLMs via circuit anal- ysis. In Forty-second International Conference on Machine Learning. Xinwei Wu, Weilong Dong, Shaoyang Xu, and Deyi Xiong. 2024. Mitigating privacy seesaw in large language models: Augmented privacy neuron edit- ing via activation patching. In Findings of the As- sociation for Computational Linguistics: ACL 2024, pages 5319â5332, Bangkok, Thailand. Association for Computational Linguistics. Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023. DEPN: Detecting and editing privacy neurons in pre- trained language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, pages 2875â2886, Singapore. As- sociation for Computational Linguistics. Yijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, Haifeng Chen, Wei Wang, and Wei Cheng. 2024. Large Language Models Can Be Contextual Privacy Protection Learners. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 14179â14201, Miami, Florida, USA. Association for Computational Linguistics. Rui Xin, Niloofar Mireshghallah, Shuyue Stella Li, Michael Duan, Hyunwoo Kim, Yejin Choi, Yulia Tsvetkov, Sewoong Oh, and Pang Wei Koh. 2024. A false sense of privacy: Evaluating textual data san- itization beyond surface-level privacy leakage. In Neurips Safe Generative AI Workshop 2024. Oleksandr Yermilov, Vipul Raheja, and Artem Chern- odub. 2023.Privacy- and Utility-Preserving NLP with Anonymized data: A case study of Pseudonymization. In Proceedings of the 3rd Work- shop on Trustworthy Natural Language Processing (TrustNLP 2023), pages 232â241, Toronto, Canada. Association for Computational Linguistics. Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulka- rni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2022. Differ- entially private fine-tuning of language models. In International Conference on Learning Representa- tions. Weichen Yu, Tianyu Pang, Qian Liu, Chao Du, Bingyi Kang, Yan Huang, Min Lin, and Shuicheng Yan. 2023. Bag of Tricks for Training Data Extraction from Language Models. In ICML, pages 40306â 40320. A Dataset Analysis Table 3 presents detailed statistics regarding the corpora used in our experiments. WordsPII TaskTr/DevMeanMaxNamesLocationsRace ECHR118161/262587087236368520899921 Table 3: Distribution of source documents in ECHR. The mean and maximum word count for all source doc- uments is presented along with an overview of the quan- tity of PII in all documents. B Flair Classes In Table 4, we present the classes of PII used to tag and monitor for leaks in our experiments. ClassDescriptionExample from ECHR PERSONNames of per- sons According to the Government, Mr L.had submitted that on the date in question. LOCGeneral loca- tions Filedaletterwith theChancellorofthe Jagiellonian University inKrakĂłw . NORPRace, national, religious groups a group ofTurkishnational- ists Table 4: The classes used by the Flair tagger for word classification. We highlight the identified span inred C Results of all baseline defenses In Table 5 we present the results of all defense baselines. DefensePerplâ Prec (%)âRec (%)â GPT-Small Baseline10.2638.19± 4.985.43± 0.80 APNEAP10.2729.92± 3.705.02± 0.87 DP (Δ=8) 18.2710.38± 4.602.27± 0.74 DP (Δ=4)19.0111.49± 1.472.06± 0.35 DP (Δ=1)20.278.01± 4.222.00± 0.76 GPT2-Medium Baseline8.7241.56± 5.556.97± 1.46 APNEAP8.7627.01± 3.344.53± 0.79 DP (Δ=8)14.9136.13± 8.956.12± 2.17 DP (Δ=4) 15.9032.47± 6.401.04± 0.23 DP (Δ=1)16.409.51± 3.173.04± 1.21 Qwen3-0.6B Baseline13.4845.01± 8.394.26± 0.49 APNEAP14.5429.26± 1.612.77± 0.38 DP (Δ=8)20.6315.69± 0.000.57± 0.00 DP (Δ=4)22.1014.83± 2.720.55± 0.10 DP (Δ=1)24.848.78± 0.670.39± 0.08 GPT2-Large Baseline9.1440.02± 7.6210.92± 0.96 APNEAP9.6326.01± 4.067.10± 0.56 DP (Δ=8)12.0042.05± 0.785.38± 0.14 DP (Δ=4)12.6937.06± 1.382.76± 0.25 DP (Δ=1) 14.9715.32± 0.001.97± 0.00 Llama-3.2-1B Baseline8.1360.12± 8.592.86± 0.16 APNEAP 8.2045.96± 6.802.89± 1.44 DP (Δ=8) 15.669.76± 0.920.53± 0.11 DP (Δ=4)20.995.75± 1.140.35± 0.10 DP (Δ=1)42.011.16± 0.920.10± 0.06 Qwen3-1.7B Baseline11.8948.26± 6.065.19± 0.23 APNEAP12.7628.96± 3.323.11± 0.23 DP (Δ=8)12.1520.43± 15.030.78± 0.36 DP (Δ=4) 12.5017.10± 15.030.63± 0.45 DP (Δ=1)13.4513.00± 12.460.49± 0.43 Table 5: Impact of DP Fine-tuning: we use perplex- ity (âPerplâ) for utility, precision (âPrecâ) and recall (âRecâ) for PII leakage, and normalized faithfulness (âFaithâ), averaged across all PII types.â(â) indicates lower (higher values) is preferred. We useGrayfor the baseline (no defenses),Greenif better,Redif worse, andOrangeif within standard deviation of the baseline. D Influential Circuit Components In order to understand more about PII leakage within the attention layers of models we displays EAP-IG scored attention layers where those lay- ers score above a 95% threshold across all scored components. We present our analysis in Figure 6. E Circuit analysis of defenses. In Figure 7, we present the circuit overlap between models in trained on differing defenses. H0H1H2H3H4H5H6H7H8H9H10H11 A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 Attention Layer GPT2-Small Baseline H0H1H2H3H4H5H6H7H8H9 H10H11H12H13H14H15 A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A13 A14 A15 A16 A17 A18 A19 A20 A21 A22 A23 GPT2-Medium Baseline H0H1H2H3H4H5H6H7H8H9 H10H11H12H13H14H15H16H17H18H19 A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A13 A14 A15 A16 A17 A18 A19 A20 A21 A22 A23 A24 A25 A26 A27 A28 A29 A30 A31 A32 A33 A34 A35 Attention Layer GPT2-Large Baseline H0H1H2H3H4H5H6H7H8H9 H10H11H12H13H14H15H16H17H18H19H20H21H22H23H24H25H26H27H28H29H30H31 A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A13 A14 A15 Llama-3.2-1B Baseline H0H1H2H3H4H5H6H7H8H9 H10H11H12H13H14H15 Attention Head A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A13 A14 A15 A16 A17 A18 A19 A20 A21 A22 A23 A24 A25 A26 A27 Attention Layer Qwen3-0.6B Baseline H0H1H2H3H4H5H6H7H8H9 H10H11H12H13H14H15 Attention Head A0 A1 A2 A3 A4 A5 A6 A7 A8 A9 A10 A11 A12 A13 A14 A15 A16 A17 A18 A19 A20 A21 A22 A23 A24 A25 A26 A27 Qwen3-1.7B Baseline 0.000 0.005 0.010 0.015 0.020 0.002 0.004 0.006 0.008 0.010 0.012 0.014 0.02 0.04 0.06 0.08 0.10 0.12 0.005 0.010 0.015 0.020 0.025 0.030 0.035 0.040 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 0.00002 0.00004 0.00006 0.00008 0.00010 0.00012 0.00014 Figure 6: Influential PII Circuit Components: Heatmaps show the average EAP-IG scores for each attention head within each layer, across all identified PII circuits. Baseline APNEAP Scrubbed =1 =4 =8 Edges IoU 1.00 0.981.00 0.600.321.00 0.540.330.551.00 0.370.920.370.371.00 0.370.960.370.370.961.00 GPT2-Small 1.00 1.001.00 0.270.461.00 0.280.410.491.00 0.270.950.490.451.00 0.270.940.490.451.001.00 GPT2-Medium 1.00 0.991.00 0.560.421.00 0.440.820.421.00 0.440.970.430.851.00 0.440.950.430.851.001.00 Qwen3-0.6B 1.00 0.981.00 0.460.431.00 0.420.370.421.00 0.380.670.350.321.00 0.490.960.460.410.691.00 GPT2-Large 1.00 0.981.00 0.420.321.00 0.380.480.351.00 0.360.450.340.521.00 0.410.940.340.520.481.00 Llama-3.2-1B 1.00 0.981.00 0.200.181.00 0.240.840.221.00 0.240.900.220.901.00 0.240.950.220.870.961.00 Qwen3-1.7B BaseAPScrub=1=4=8 Baseline APNEAP Scrubbed =1 =4 =8 Nodes IoU 1.00 0.881.00 0.910.771.00 0.900.740.881.00 0.850.710.850.851.00 0.840.720.840.850.991.00 BaseAPScrub=1=4=8 1.00 0.851.00 0.590.451.00 0.580.440.741.00 0.570.390.760.721.00 0.570.410.760.721.001.00 BaseAPScrub=1=4=8 1.00 0.891.00 0.770.641.00 0.640.500.611.00 0.650.500.630.921.00 0.650.500.630.921.001.00 BaseAPScrub=1=4=8 1.00 0.851.00 0.690.541.00 0.640.500.641.00 0.580.440.560.531.00 0.720.570.720.650.741.00 BaseAPScrub=1=4=8 1.00 0.831.00 0.640.461.00 0.620.440.601.00 0.600.480.590.751.00 0.660.490.590.730.711.00 BaseAPScrub=1=4=8 1.00 0.871.00 0.450.311.00 0.470.310.411.00 0.470.300.400.941.00 0.470.300.400.930.981.00 Figure 7: Circuit overlap across defenses: percentage overlap of edges (upper) and nodes (lower) in PII circuits.