Paper deep dive
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto, Shalini Dhamodharan, John P. Woodhouse, Chi-fan Lin, Mark Zobeck, Zhandong Liu, Hyun-Hwan Jeong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/19/2026, 4:10:46 AM
Summary
This study evaluates the efficacy of Large Language Models (LLMs) using in-context learning (ICL) for de-identifying protected health information (PHI) in pediatric oncology notes, specifically addressing 'institutionally situated' PHI missed by traditional systems and gold standards. Benchmarking eight LLMs against purpose-built systems (Stanford TiDE, OpenMed PII) and pattern-based baselines on 100 annotated notes from Texas Children's Hospital, the research demonstrates that LLMs significantly outperform traditional methods (best F1=0.918 vs. TiDE 0.779). The study introduces a three-tier prompting strategy (Baseline, Targeted, Precision) to control the precision-recall trade-off, showing that naming specific institutional identifiers and warning against over-redaction recovers missed PHI and restores precision. Additionally, the study reveals that LLM outputs can identify gaps in existing gold standards, leading to re-annotation that further improves recall. Multi-agent architectures did not outperform calibrated single-pass prompting.
Entities (10)
Relation Signals (7)
Texas Children's Hospital â sourceof â Pediatric Oncology Notes
confidence 96% · On 100 annotated pediatric oncology notes... from Texas Children's Hospital
Sonnet 4.6 â outperformed â OpenMed PII
confidence 95% · LLMs outperformed the purpose-built systems... OpenMed PII 80.1% recall... F1=0.743
Sonnet 4.6 â outperformed â Stanford TiDE
confidence 95% · LLMs outperformed the purpose-built systems (best F1=0.918±0.001, Sonnet 4.6, vs. TiDE 0.779)
In-Context Learning â usedfor â De-identification
confidence 93% · We evaluated whether large language models (LLMs) can close this gap through in-context learning
Targeted Prompt â recovers â Institutionally Situated PHI
confidence 92% · Naming the missed categories recovered 79% (48/61) of them
Precision Prompt â restores â Precision
confidence 91% · discouraging over-redaction restored precision... Precision recovered precision to 0.829
LLMs â identifiedgapsin â Gold Standard
confidence 90% · LLM outputs surfaced 414 candidate annotation gaps; re-annotation confirmed 227 PHI spans
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.
Tags
Links
- Source: https://arxiv.org/abs/2608.17051v1
- Canonical: https://arxiv.org/abs/2608.17051v1
Trouble viewing inline? Open PDF directly â
Full Text
90,882 characters extracted from source content.
Expand or collapse full text
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss Article type: Research and Applications Authors. Daniel Palacios, BS 1,2,3,4,â , Matthew Brady Neeley, BS 1,2,3,4,â , Angel Adetomike Otto, MS 5 , Shalini Dhamodharan, MS 5 , John P. Woodhouse, BA 5 , Chi-fan Lin, MS 5 , Mark Zobeck, MD, MPH 5,â , Zhandong Liu, PhD 1,2,3,4,â , Hyun-Hwan Jeong, PhD 2,3,4,â . â These authors contributed equally. â Co-corresponding authors. Affiliations. 1. Quantitative and Computational Biosciences, Baylor College of Medicine, Houston, Texas, USA 2. Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA 3. Jan and Dan Duncan Neurological Research Institute, Texas Childrenâs Hospital, Houston, Texas 77030, USA 4. Data Science Center, Texas Childrenâs Hospital, Houston, Texas 77030, USA 5. Section of Hematology-Oncology, Department of Pediatrics, Baylor College of Medicine, Houston, Texas, USA Corresponding author. Hyun-Hwan Jeong, Department of Pediatrics, Baylor College of Medicine, and Jan and Dan Duncan Neurological Research Institute, Texas Childrenâs Hospital, 1250 Moursund Street, Houston, TX 77030, USA. Telephone: +1 832-824-1000, ext. 25535. Email: hyun-hwan.jeong@bcm.edu. Keywords: data anonymization; electronic health records; natural language processing; machine learning; large language models. Word count. Abstract: 239. Main body: 3,808. Tables: 1. Figures: 4. arXiv:2608.17051v1 [cs.CL] 17 Aug 2026 2 ABSTRACT Objective. Secondary use of electronic health records requires de-identification, yet existing systems miss institutionally situated protected health information (PHI): identifiers such as hospital abbreviations and building names whose status is locally determined. We evaluated whether large language models (LLMs) can close this gap through in-context learning while controlling precision and recall. Materials and Methods. On 100 annotated pediatric oncology notes (5,322 PHI spans) we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM was run in three prompt modes: Baseline (HIPAA-aligned), Targeted (plus institutional PHI categories), and Precision (plus instructions against over-redaction). We also compared 14 multi-agent and ensemble configurations. Recall was the primary safety metric. Results. LLMs outperformed the purpose-built systems (best F1=0.918±0.001, Sonnet 4.6, vs. TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79% (48/61), and Precision recovered precision. No agentic architecture beat single-pass prompting (F1 0.906â0.908). The LLMs also redacted 414 identifiers absent from the gold standard, scored false positive; expert review of 49 confirmed all as true PHI, and re-annotating the 10 highest-discrepancy notes (+227 spans) lifted Precision to recall=0.981 (F1=0.907±0.002). Discussion. Naming an institutionâs own identifiers and warning against over-redaction resolves both the institutional PHI gap and the precisionârecall trade-off in one LLM call per note. LLMs can cost more, but that buys a way to audit the standard. Conclusion. LLMs are an adaptable alternative to purpose-built de-identification; institution-specific prompt development should be the primary adaptation. 3 INTRODUCTION Secondary use of electronic health records (EHRs) requires de-identification of protected health information (PHI) under HIPAAâs Safe Harbor standard, which specifies 18 identifier categories [1,2]. These define canonical PHI (names, dates, medical record numbers, geographic data) that models recognize from general linguistic patterns. Real clinical notes also contain institutionally situated PHI: identifiers whose status depends on local institutional context, including hospital abbreviations (âTCHâ for Texas Childrenâs Hospital), building names (âMark Wallace Towerâ), internal clinic codes, and provider naming conventions unique to an institution. Although not enumerated among the 18 categories, they fall within HIPAAâs definition of PHI under both the Safe Harbor catch-all for âany other unique identifying number, characteristic, or codeâ (45 CFR §164.514(b)(2)(i)(R)) and the expert-determination standard (§164.514(b)(1)) [2,3]: in a pediatric-oncology population a named specialty facility, a rare diagnosis, and a service date can jointly re-identify a record even after every canonical identifier is removed. Each element is independently re-identifying: the set of facilities a patient visits is itself a signature [4,5] and diagnosis codes alone can breach privacy [6]; in Washington State discharge data carrying hospital, diagnosis, and attending physician but no names or addresses, news reports uniquely matched 35 of 81 named patients to their records [7]. Removal is a regulatory requirement, not an optional refinement. A model prompted only on HIPAAâs 18 categories has no basis to recognize that âTCHâ is an identifying abbreviation or that a four-digit pager number is a staff identifier. These are failures not of capability but of specification: the model was never told they require redaction. The OCR guidance anticipates this, warning that esoteric notation such as acronyms known to only a few of a covered entityâs employees can lead to either unnecessary redaction or failure to redact [2]. Nor can the specification be written once and reused, since note templates, abbreviations, and patient populations differ by site [8]: Veterans Health Administration notes required customizing to institution-specific formats [9,10], and cross-institute evaluations report consistent degradation on transfer [11]. The gap is acute in pediatric settings, where large multidisciplinary teams author notes that reference caregivers and carry institutional shorthand. Recent work has begun isolating institution-level identifiers as a distinct annotation class, notably the hospital category of SHIELD â a recent teacherâstudent distillation framework for clinical de-identification â on which both its teacher and student models record their lowest precision [12]. No prior study has characterized why these identifiers fail, nor shown the failure remediable through specification rather than retraining. Automated de-identification has progressed from rule-based systems [13,14] through neural sequence models [15,16] to transformer-based NER [17,18,19], benchmarked on i2b2/UTHealth shared tasks [20,21]. Two purpose-built systems anchor our comparison: OpenMedâs domain-adapted NER [22] and Stanfordâs TiDE, which combines NER, pattern matching, and known-PHI lookup [23, 24] with Hiding-in-Plain-Sight surrogates [25]. Large language models (LLMs) are competitive with or superior to traditional NER systems on adult clinical benchmarks [26]. Closest to this work, Wiest et al. [27] benchmarked eight local LLMs on 250 clinical letters, reportingâŒ99.2% PHI removal in a different language and note type; Altallaâ et al. [28] evaluated GPT-3.5 and GPT-4 (Pâ0.99, Râ0.83), and Pissarra et al. [29] found LLMs and Presidio baselines complementary. Multi-agent architectures have also been explored: TEAM-PHI [30] ranks de-identification models with majority-voting evaluation agents and no gold labels, OEMA [31] uses three agents for zero-shot 4 clinical NER, and SHIELD [12,32,33] selects a teacher labeler for distillation into locally-deployable students [34]. The prover-verifier framework [35] justifies generation-then-verification architectures, untested in de-identification. A second challenge is the precisionârecall trade-off [36]. De-identification has long favored recall, since a missed identifier is a privacy breach whereas over-redaction only removes clinical content [20,37]. Distinctive in the single-pass LLM setting are its magnitude and model-dependence: no single prompt optimizes both objectives across models, and some over-redact severely enough to degrade data utility, an effect standard metrics do not capture [38]. Conversely, adversarial LLM-based re-identification [39] shows even strong systems leave notes vulnerable. We hypothesized that what limits de-identification is specification rather than model capability: that naming an institutionâs own identifiers in context would resolve both the institutional PHI gap and the precisionârecall trade-off. We tested this against the competing explanation that the trade-off demands architectural remedy, using dual-pass and ScrubberâAuditor pipelines drawn from the prover-verifier paradigm [35] and multi-agent clinical NLP [30] (Supplementary Note 3). We make three contributions. First, benchmarking 8 LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII), pattern-based baselines, and multi-stage pipelines on 100 expert-annotated pediatric oncology notes (5,322 PHI spans), we introduce institutionally situated PHI as a failure mode common to all. Second, in-context learning adapts de-identification without fine-tuning: naming the missed categories recovers most of them and anti-over-redaction instructions restore precision, while error analysis exposed gold-standard gaps that expert re-annotation confirmed as true PHI. Third, the trade-off resolves in a single pass: no agentic architecture outperformed it on F1, locating the bottleneck in specification rather than inference-time computation. METHODS Study design and clinical corpus We benchmarked the LLMs and purpose-built systems listed in Table 1, plus two pattern-based baselines (regex-only, and spaCy NER + regex), on 100 pediatric oncology clinical notes, with the LLMs run under each of the three prompt conditions below. No model was trained or fine-tuned. Reporting follows TRIPOD-LLM [40]; Tables S1 and S2 give the item-by-item adherence table. Clinical note corpus. Our corpus consisted of 100 English-language clinical notes from the pediatric oncology service at Texas Childrenâs Hospital (Houston, TX, USA; 96 patients), drawn by pseudorandom sort under a fixed seed from notes dated on or after January 1, 2021, unstratified. Notes ranged from 91 to 32,767 characters (median, 6,504; mean, 10,058), five truncated at the export limit. Three clinical annotators, trained by a pediatric oncologist and informatician to identify both the 18 Safe Harbor categories and institutionally situated identifiers, annotated 5,322 PHI spans across 97 notes (3 contained no PHI), resolving questions with the oncologist and regulatory personnel; because annotation followed consensus adjudication rather than independent double annotation, no agreement statistic was computed. Figure 1C shows the category distribution (Supplementary Note 1.1). Validation dataset. To check that our findings were not corpus-specific, we also evaluated all systems on 49 notes (409 gold PHI spans) from the USDHUB repository, a separately curated pediatric neurology sample from the same institution, independently de-identified by a different group, under the identical Baseline prompt and scoring pipeline (Supplementary Note 4). Because USDHUB was built for TiDE and ships patient-specific provisioning material, TiDE was run in two bracketing configurations: fully unprovisioned (identical footing to the LLMs) and maximally provisioned with the supplied known-PHI dictionary and 5 per-note identifier header. For the provisioned run only, precision and F1 are scored on the note body, since the injected header carries no gold spans (Supplementary Note 4.1). De-identification systems LLM models. We evaluated eight LLMs via AWS Bedrockâs Converse API (Table 1) at temperature = 0.0 (except Opus 4.8, which does not accept the parameter via Bedrock; Supplementary Note 3.3), max output tokens = 65,000, and up to 3 retries with exponential backoff. For visual clarity, main-text figures present the top 4 LLMs by Baseline F1 (Sonnet 4.6, Opus 4.8, GLM-5, DeepSeek V3.2); all eight appear in Tables S3âS6. Table 1: LLMs and purpose-built systems evaluated. The two pattern-based baselines (regex-only; spaCy NER + regex) are specified in the text and are included in all reported comparisons. SystemTypeFamilyParameters Claude Opus 4.8LLMAnthropicUndisclosed Claude Sonnet 4.6LLMAnthropicUndisclosed GPT-oss-120BLLMOpenAI120B GPT-oss-20BLLMOpenAI20B GLM-5LLMZhipu AI754B Kimi K2.5LLMMoonshot1.1T MiniMax M2.5LLMMiniMax229B DeepSeek V3.2LLMDeepSeek685B Stanford TiDENER + RulesStanford NLPâ OpenMed PIINER (token classif.)OpenMed434M Stanford TiDE. TiDE [23] is a production system detecting PHI through NER, regex, and known-PHI matching against a supplied list of each patientâs real identifiers. Since the LLMs and OpenMed never receive that list, input parity required running TiDE on note text alone, with known-PHI matching disabled so field lookups find no rows while NER and regex fire normally. This understates TiDEâs production performance but isolates contextual reasoning from pattern matching (Supplementary Note 4.1). OpenMed PII. We used OpenMed-PII-SuperClinical-Large-434M-v1 [22,41], a transformer token-classification model fine-tuned for personally identifiable information (434M parameters, 54 sensitive-information types), whose bracketed placeholders match the LLM output format, so the same scoring pipeline applies. Prompt conditions The Baseline prompt asks the model to return the exact input text with all 18 HIPAA Safe Harbor categories replaced by typed placeholders, enumerating the categories with examples. Targeted appends four categories of institutionally situated PHI that Baseline error analysis surfaced: staff names adjacent to credentials, pager and Voalte numbers, institution names and abbreviations (âTCHâ, âTXCHâ), and building or facility names. Department and clinic codes, a fifth subcategory the same analysis surfaced, were not named (Table S7). Precision retains all of Targeted and adds a âDO NOT OVER-REDACTâ block covering the six largest observed false-positive categories, each with WRONGâRIGHT pairs, closing with an instruction to redact anyway when genuinely uncertain (Supplementary Note 2.2). 6 Multi-stage architectures We tested whether architectural complexity could outperform single-pass prompt engineering, using multi-stage for any pipeline with more than one LLM call, multi-agent for pipelines whose calls occupy distinct roles (scrubber, auditor, verifier), and heterogeneous for the subset combining two or more models. Three paradigms were explored: dual-pass iterative refinement (Sonnet 4.6 applied twice, the second pass searching its own output for residual institutional PHI), heterogeneous Scrubberâ Auditor (a recall-maximizing Sonnet 4.6 scrubber followed by a precision-focused Opus 4.8 auditor), and cross-model dual-pass (two models with complementary error profiles). Descriptions, prompts, and design rationale are in Supplementary Note 3. The top 3 configurations were run for 5 independent trials with statistical comparison (McNemarâs exact test, Wilcoxon signed-rank, bootstrap 95% CIs, HolmâBonferroni corrected across five tests at family-wise α=0.05; Supplementary Note 3.3). Enhanced gold standard Error analysis (Results) showed the original annotation had omitted institutional identifiers the LLMs correctly detected, penalizing correct redactions, deflating F1 and obscuring differences between architectures. We therefore re-annotated a subset for relative comparison. From the 414 candidate spans surfaced corpus-wide we selected the 10 notes with the largest modelâgold discrepancy; a domain expert adjudicated 49 in-context instances covering 22 unique institutional terms, confirming all 49 as TRUE_PHI. Annotating those terms at every occurrence added 227 spans (209 Geographic Data, 9 Name, 9 Other Unique ID), for 1,758 total versus 1,531 original; institutional terms were assigned to Geographic Data, which is why that category grows far beyond its corpus- wide count (Supplementary Note 2.5). Three biases follow: the subset overrepresents institutional PHI density; re-annotation covered only the 22 surfaced terms; and the candidates came from model output, partly crediting those models for spans they surfaced. Expert adjudication establishes the added spans are genuine PHI, not that they exhaust it, so the enhanced standard serves only for relative comparison on these 10 notes. All multi-stage architectures, and a five-trial re-evaluation of the three prompt conditions, were scored against it. Evaluation pipeline Each gold PHI span is a true positive (TP) when its text is absent from the output, and a false negative (FN) when an exact substring search still finds it, when masking is partial, or when no output is returned. A false positive (FP) is an emitted placeholder matching no gold annotation. Matching is type-agnostic, so a date masked as[NAME]still counts as TP. Recall is TP/(TP+FN) and precision TP/(TP+FP), whose denominator mixes gold spans with emitted placeholders; that unit mismatch cuts both ways, raising precision where a system merges adjacent gold spans and lowering it where one is split across several. Text also occurring as legitimate non-PHI (âMayâ as name vs. month) can produce spurious counts. Outputs shorter than 50% of the note are scored as failures with all spans FN. False negatives were categorized as canonical or institutionally situated PHI. Placeholder alignment, TiDEâs surrogate-based precision scoring, the mixed-unit bias, and the full rule set are in Supplementary Note 1.2. Figures use Matplotlib in the soft-fill style of PubliPlots [42]. 7 Ethical considerations This study was conducted under IRB protocol H-52222 at Baylor College of Medicine / Texas Childrenâs Hospital. All data remained within the institutional environment, and LLM inference via AWS Bedrock ran under a HIPAA-compliant Business Associate Agreement under which prompts and outputs are neither retained nor used for training. LLMs are the object of study here; any use of AI tools in manuscript preparation is disclosed in Additional Contributions. RESULTS Recall is the primary safety metric, since a missed span is a potential privacy violation whereas an over-redacted one only degrades data utility. We select on F1 to keep the precision cost of recall gains visible, but because F1 weights the two errors equally we also tabulate the recall-weightedF 2 (5PR/(4P+R); Tables S3 and S8), which reorders only adjacent pairs and leaves the leaders unchanged. LLMs outperform traditional de-identification on pediatric oncology notes Under identical input (Baseline prompt for LLMs), LLMs substantially outperformed all traditional approaches (Figure 2). The spaCy NER + regex baseline reached 77.3% recall at only 34.4% precision (7,832 false positives), since general-purpose NER labels medications, diagnoses, and anatomy as entities; regex alone reached 57.4% recall at 96.4% precision, working for structured but not context-dependent PHI. Stanford TiDE reached 75.9% recall (F1=0.779) and OpenMed PII 80.1% recall at 69.4% precision (F1=0.743), the highest non-LLM recall but lower F1 (Table S3, Supplementary Figure 1). The same ordering held on the 49-note validation set (best LLM F1 0.894 vs. 0.866), though the margin narrowed on that canonical-PHI-dominated corpus, and onF 2 the best-LLM advantage over TiDE there falls to 0.001 (Supplementary Note 4). Among LLMs, Sonnet 4.6 combined 96.0% recall with 88.3% precision for the best F1 (0.920 in the primary run; 0.918±0.001 across five trials); Opus 4.8 matched its precision at lower recall (93.1%, F1=0.905); and the six LLMs without a failure mode exceeded TiDEâs recall by 0.15â0.22 (Table S3), the exception being GPT-oss-20B, whose truncation puts it below TiDE. The trade-off is model-dependent: under Baseline the highest-recall models over-redact enough to generate thousands of false positives (DeepSeek V3.2 2,264 FP at 98.3% recall; Kimi K2.5 3,251 FP at 96.9%), whereas Sonnet 4.6 produces only 679 FP at comparable recall. Per-category analysis reveals shared weaknesses on institutionally situated PHI Because aggregate recall is dominated by Date spans (72.1% of gold annotations), per-category results (Figure 2B; all twelve systems and ten categories in Table S4) show both why LLMs outperform traditional systems and where they fail. TiDE achieves perfect recall on MRN but fails where context is required: Other Unique ID (15.7%), Phone (47.9%), and Geographic Data (65.0%). Averaged over the top 4 LLMs, recall exceeds TiDE by 0.49 on Phone and 0.48 on Other Unique ID â identifiers in non-standard formats such as pager codes and internal extensions that evade regex but are readable from context (Supplementary Note 2.3). The panel also reveals the LLMsâ shared weakness: even the best models reach only 50â78% recall on Other Unique ID and 69â99% on Geographic Data, the categories most enriched for institutionally situated PHI. 8 In-context learning enables control over the precisionârecall trade-off Unlike TiDE or regex baselines, whose behavior is fixed by their pattern libraries, LLMs can be steered by prompt design alone. We developed three prompt conditions through iterative error analysis on Sonnet 4.6, then evaluated all across the model set. Decomposing Sonnet 4.6âs 211 Baseline false negatives (Figure 2C, a representative run bracketed by the five-trial means in Table S9) identified institutional PHI as the blind spot: 29% of misses (61/211), dominated by institution abbreviations (38) and building or facility names (15). Table S7 gives the subcategory taxonomy with failure mechanisms and recovery rates. Appending those four categories to the baseline instructions (Targeted) reduced Sonnet 4.6âs institutional false negatives from 61 to 13 (78.7% recovery), with the largest gains on Other Unique ID (recall 0.548 to 0.791) and Name (0.928 to 0.971), raising overall recall from 0.958 to 0.975 (5-trial means; Table S9). It also carried a precision penalty (0.881 to 0.807), as the added instructions over-redacted clinical content in six recurring categories (Supplementary Note 2.2). Precision keeps those categories and appends anti-over-redaction WRONGâRIGHT pairs, recovering precision to 0.829 for a modest recall cost (0.975 to 0.969) and F1=0.893±0.004; the full progression nets higher recall (+0.011) at lower precision (â0.052). The trend holds across the full model set (Figure 3AâB; Tables S5 and S6; per-category recall in Supplementary Figures 2 and 3), with Targeted improving recall (Opus 4.8 +0.023, Kimi K2.5 +0.018, Sonnet 4.6 +0.014) and Precision improving precision for all 7 models with valid output (mean +0.054). All three prompts carry worked examples; only those naming content to preserve recover the clinical text category-only instructions over-redact [43]. Magnitude varies by model: those with severe baseline precision deficits benefit most (GPT-oss-120B P=0.198 to 0.924; Kimi 0.613 to 0.816; DeepSeek 0.698 to 0.806), whereas MiniMax regresses on recall under Targeted (â0.141) and Sonnet 4.6 scores its highest F1 under Baseline, whose high precision the instructions can only cost; the optimal prompt is therefore model-dependent. That F1 penalty must be read with caution: as the next section shows, many of Precisionâs additional âfalse positivesâ are correct redactions of institutional PHI the original standard failed to annotate. Error analysis reveals annotation gaps in the gold standard While comparing multi-agent pipelines against our best single-pass model, we noticed they flagged institutional PHI absent from the gold standard. Rather than count these as false positives, we asked whether the standard itself was incomplete. Manual inspection confirmed annotators had systematically missed institution-specific identifiers: operating room location codes (C OR, LT OR, MW OR, WC OR, WT MAIN OR, GIPS), hospital abbreviations (TCH, TXCH, BCM), and campus or building names (Mark Wallace Tower, West Campus), all PHI under the catch-all and expert-determination provisions. The scarcity of reliable annotations motivates label-free evaluation [30]; our results show they are also incomplete. Expert re-annotation of the 10 highest-discrepancy notes confirmed all 49 adjudicated instances as PHI and added 227 spans; the protocol and its biases are in Methods. Re-evaluation against enhanced gold. On the enhanced 10-note subset (1,758 spans, 5 trials per prompt; Figure 4C, Table S10), annotation gaps disproportionately penalize institution-aware prompts. Baseline recall is 0.847±0.002 here versus 0.958±0.002 on the original 100-note standard, because it leaves the newly-annotated institutional terms unredacted. Those two numbers differ in note set as well as standard, so we held the notes fixed: against the original annotation of these same 10 notes, Baseline recall is 0.972±0.003, so re-annotation accounts forâ0.125 of theâ0.111 net change and note selection for +0.014 (Supplementary 9 Note 2.6). Targeted recovers recall to 0.980±0.000 at precision 0.810±0.004 (F1=0.887±0.002), and Precision holds recall (0.981±0.000) while recovering precision to 0.844±0.003 (F1=0.907±0.002) â a Baseline-to-Targeted recall gap far larger on the enhanced gold (0.133) than on the original (0.017). Multi-agent architectures confirm single-pass sufficiency None of the 14 multi-agent and ensemble configurations (Supplementary Note 3, Tables S11 and S12) improved F1 over the single-pass Precision prompt. The three reproducibility-tested methods (5 trials each, Table S13) reached comparable mean F1 (0.906â0.908) with overlapping 95% trial-resampled CIs (Figure 4AâB), though our trial count cannot establish formal equivalence. The agentic pipelineâs intra-method variability (SD F1 =0.018) exceeds inter-method differences, and 4 of its 5 trials fell below single-pass (median 0.898 vs. mean 0.906, pulled up by one high-precision run at F1=0.942), so a typical agentic run underperforms. Ensemble voting reaches marginally higher recall (Cross-Model Vote 0.986 vs. 0.981) at a precision cost and 6Ă the inference cost. Across the four top LLMs, 20.3% of false negatives are shared, dominated by ambiguous partial dates and institutional identifiers (Supplementary Note 2.4). DISCUSSION Our results establish three findings. First, LLMs substantially outperform purpose-built systems, with the advantage concentrated in categories requiring contextual reasoning. Second, in-context learning enables control over the precisionârecall trade-off, and which prompt looks best depends on how completely the standard annotates institutional PHI: Baseline wins on the original annotations and loses on the corrected ones, not because Precision improves but because Baselineâs recall collapses once the institutional terms it leaves unredacted are counted. Third, no multi-agent architecture improved F1 over single-pass. Those configurations proved useful instead as a discovery tool for enumerating annotation gaps at scale, though not uniquely so, since the single-pass Precision prompt flagged the same identifiers. The prover-verifier framework [35] that motivated this exploration yields no gain here because the bottleneck is specification â which institutional terms to redact â not verification of a given redaction. Recommended deployment configurations are given in Supplementary Note 3.5. Limitations First, our corpus (100 notes, single institution) limits generalizability. The 49-note USDHUB validation set (Supplementary Note 4) confirms the LLM F1 advantage persists but narrows on canonical-PHI-dominated corpora; because USDHUB is from the same institution, it establishes robustness across specialties, note types, and annotation methods, but not across institutions. Cross-institutional generalization and the transferability of our site-specific addenda remain to be established. Second, the 10 re-annotated notes were selected by highest modelâgold discrepancy, biasing toward vindicating the model, so enhanced-gold results are relative comparisons across prompts, not corpus-wide estimates. Third, TiDE was run unprovisioned on the primary corpus; on USDHUB both extremes were near-identical (recall 0.976 vs. 0.971; Supplementary Note 4.1), so the LLM advantage holds whether or not pattern-based systems receive site-specific information. Fourth, the Targeted and Precision addenda came 10 from Baseline error analysis on this same corpus, so their gains are in-sample. Fifth, we report no subgroup or fairness analysis: the attributes that would define subgroups are themselves the PHI under removal and were never extracted, so whether recall differs across patient groups is unresolved. Sixth, both GPT-oss models showed failure modes unrelated to de-identification capability (GPT-oss-20B output truncation, depressing recall; GPT-oss-120B unstable redaction at precision 0.198), unlike the other six (Table S5). Finally, that multi-agent architectures fail to beat single-pass likely reflects the task itself: de-identification is a single-read problem in which all necessary context is already in the note. Additional passes let the model second-guess correct decisions, and coordination adds noise without information (Supplementary Note 3.4). This matches findings outside the clinical domain, where multi-agent gains are often minimal and failures trace to specification rather than model capability [44], and where a single agent with strong prompts matches multi-agent discussion, the latter winning only where the prompt leaves the task underspecified [45]; our Targeted and Precision prompts supply that specification. Future directions Promising directions include knowledge distillation, using high-performing LLMs as teacher labelers for small, locally-deployable models at 100Ălower cost [33,12,46]; adversarial verification against an LLM attempting re-identification [39,25,47]; multi- institutional validation; and gold standard refinement to include institutionally situated PHI. CONCLUSION On 100 pediatric oncology notes, LLMs beat purpose-built de-identification on recall â its core promise â by 0.20 over Stanford TiDE, and the gap is widest exactly where pattern matching cannot reach: identifiers whose PHI status depends on institutional context. Prompting, not retraining, is the adaptation mechanism. Naming the institutional categories a HIPAA-aligned prompt misses recovers 79% (48/61) of them, and anti-over-redaction instructions then restore precision, within a single LLM call and impossible with purpose-built systems short of retraining. Multi-agent architectures do not help; none of 14 configurations improved F1 over a single pass with the same prompt, because de-identification is a single-read task in which the second pass has no information the first lacked. Ensembles buy marginally higher recall at a precision cost and 3 calls per note. And LLMs proved good enough to audit their own reference standard: what looked like over-redaction was largely PHI the annotators had missed, expert adjudication confirming all 49 in-context instances as true PHI. Three consequences follow. De-identification should be evaluated per category, not on aggregate metrics that hide institutional PHI; gold standards deserve auditing before they are trusted as ground truth; and effort belongs in institution-specific prompts rather than additional passes. LLMs cost more per note, but that cost buys adaptation without retraining and a check on the standard itself. Our evaluation framework is released open-source. 11 FUNDING This research was supported by a fellowship from the Gulf Coast Consortia on the NLM Training Program in Biomedical Informatics and Data Science (T15 LM007093); the National Science Foundation Graduate Research Fellowship Program (NSF GRFP Fellow ID 2024370642); the Fund for Innovation in Cancer Informatics; the Cancer Prevention and Research Institute of Texas (CPRIT, RP240131); the Chan Zuckerberg Initiative (2023-332162); the National Institutes of Health (NIH, U54NS093793 and OT2OD040565); the Eunice Kennedy Shriver National Institute of Child Health and Human Development of the NIH (P50HD103555); the Chao Endowment; the Huffington Foundation; and the Jan and Dan Duncan Neurological Research Institute at Texas Childrenâs Hospital. ADDITIONAL CONTRIBUTIONS We thank the Texas Childrenâs Hospital Office of Research Data, which provided the independently de-identified USDHUB note set used for our within-institution validation (Supplementary Note 4). No AI-assisted tools were used for study design, analysis, or primary manuscript drafting. Generative AI tools were used only for proofreading and typographical/grammatical correction of author-written text; they were not used to generate scientific content, analyze data, or draft substantive passages of the manuscript. In accordance with COPEâs position and JAMIA policy, no AI or NLP tool is listed as an author; the authors reviewed and verified all text and take full responsibility for the integrity, accuracy, and originality of all content. CONFLICTS OF INTEREST The authors declare no competing interests. DATA AVAILABILITY The evaluation pipeline and code are openly available athttps://github.com/LiuzLab/phi-scrubber-evaluation. The underlying clinical notes cannot be shared because they are protected patient health information governed by IRB protocol H-52222 and institutional/HIPAA data-use restrictions; they are not available for public deposition or on request. De-identified aggregate metrics that support the findings of this study are provided within the article and its supplementary materials. AUTHOR CONTRIBUTIONS M.B.N. and D.P. contributed equally as co-first authors: conceptualization, methodology, software, formal analysis, investigation, visualization, and writing. J.P.W., C.L., A.A.O., and S.D. contributed to data curation and investigation (PHI annotation and re-annotation). M.Z. provided clinical expertise, adjudication of institutional PHI, annotation, and review and editing of the writing. Z.L. is the principal investigator and contributed supervision and funding acquisition. H.H.J. contributed supervision, project administration, and review and editing of the writing. All authors reviewed and approved the final manuscript. 12 REFERENCES [1] U.S. Congress. Health insurance portability and accountability act of 1996, 1996. Public Law 104-191. [2]Office for Civil Rights, HHS. Guidance regarding methods for de-identification of protected health information, 2012. URL https://w.hhs.gov/hipaa/for-professionals/privacy/special-topics/de-identification/. [3]U.S. Department of Health and Human Services. Other requirements relating to uses and disclosures of protected health information, 2024. URLhttps://w.ecfr.gov/current/title-45/section-164.514. 45 C.F.R. §164.514; see §164.514(b)(1) (expert determination) and §164.514(b)(2)(i)(R) (âany other unique identifying number, characteristic, or codeâ). Accessed August 2026. [4]Bradley Malin and Latanya Sweeney. How (not) to protect genomic data privacy in a distributed network: using trail re-identification to evaluate and design anonymity protection systems. Journal of Biomedical Informatics, 37(3):179â192, 2004. doi: 10.1016/j.jbi.2004.04.005. [5]Latanya Sweeney. Simple demographics often identify people uniquely. Data Privacy Working Paper 3, Carnegie Mellon University, Pittsburgh, 2000. URL http://dataprivacylab.org/projects/identifiability/. [6]Grigorios Loukides, Joshua C. Denny, and Bradley Malin. The disclosure of diagnosis codes can breach research participantsâ privacy. Journal of the American Medical Informatics Association, 17(3):322â327, 2010. doi: 10.1136/jamia.2009.002725. [7]Latanya Sweeney. Matching known patients to health records in Washington State data, 2013. URLhttps://arxiv.org/ abs/1307.1370. arXiv:1307.1370. [8]Beau Norgeot, Kathleen Muenzen, Thomas A. Peterson, Xoli Fan, Benjamin S. Glicksberg, Gabriel Schenk, Eugenia Rutenberg, Boris Oskotsky, Marina Sirota, Jinoos Yazdany, Gabriela Schmajuk, Dana Ludwig, Theodore Goldstein, and Atul J. Butte. Protected health information filter (Philter): accurately and securely de-identifying free-text clinical notes. npj Digital Medicine, 3:57, 2020. doi: 10.1038/s41746-020-0258-y. [9]Ăscar FerrĂĄndez, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, and StĂ©phane M. Meystre. Evaluating current automatic de-identification methods with Veteranâs health administration clinical documents. BMC Medical Research Methodology, 12:109, 2012. doi: 10.1186/1471-2288-12-109. [10]StĂ©phane M. Meystre, Ăscar FerrĂĄndez, F. Jeff Friedlin, Brett R. South, Shuying Shen, and Matthew H. Samore. Text de-identification for privacy protection: a study of its impact on clinical text information content. Journal of Biomedical Informatics, 50:142â150, 2014. doi: 10.1016/j.jbi.2014.01.011. [11]Xi Yang, Tianchen Lyu, Qian Li, Chih-Yin Lee, Jiang Bian, William R. Hogan, and Yonghui Wu. A study of deep learning methods for de-identification of clinical notes in cross-institute settings. BMC Medical Informatics and Decision Making, 19 (Suppl 5):232, 2019. doi: 10.1186/s12911-019-0935-4. 13 [12] Jose D. Posada, David Love, Somalee Datta, and Priya Desai. SHIELD: Synthetic human-annotated identifier-replaced entries for learning and de-identification. arXiv preprint arXiv:2605.03301, 2026. [13]Latanya Sweeney. Replacing personally-identifying information in medical records, the Scrub system. Proceedings of the AMIA Annual Fall Symposium, pages 333â337, 1996. [14]Ishna Neamatullah, Margaret M Douglass, Li-wei H Lehman, Andrew Reisner, Mauricio Villarroel, William J Long, Peter Szolovits, George B Moody, Roger G Mark, and Gari D Clifford. Automated de-identification of free-text medical records. BMC Medical Informatics and Decision Making, 8(1):32, 2008. doi: 10.1186/1472-6947-8-32. [15]Franck Dernoncourt, Ji Young Lee, Ozlem Uzuner, and Peter Szolovits. De-identification of patient notes with recurrent neural networks. Journal of the American Medical Informatics Association, 24(3):596â606, 2017. doi: 10.1093/jamia/ocw156. [16]Zengjian Liu, Buzhou Tang, Xiaolong Wang, and Qingcai Chen. De-identification of clinical notes via recurrent neural network and conditional random field. Journal of Biomedical Informatics, 75:S34âS42, 2017. doi: 10.1016/j.jbi.2017.05.023. [17]Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234â1240, 2020. doi: 10.1093/bioinformatics/btz682. [18]Emily Alsentzer, John R Murphy, Willie Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew BA McDermott. Publicly available clinical BERT embeddings. Proceedings of the 2nd Clinical Natural Language Processing Workshop, pages 72â78, 2019. [19]Angel Paul, Dhivin Shaji, Lifeng Han, Warren Del-Pinto, Goran Nenadic, and Suzan Verberne. DeIDClinic: A risk- aware pseudonymization framework for clinical text de-identification and re-identification risk assessment, 2026. URL https://arxiv.org/abs/2410.01648. arXiv:2410.01648. [20]Amber Stubbs, Christopher Kotfila, and Ăzlem Uzuner. Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus. Journal of Biomedical Informatics, 58:S20âS29, 2015. doi: 10.1016/j.jbi.2015.07.020. [21]Ăzlem Uzuner, Yuan Luo, and Peter Szolovits. Evaluating the state-of-the-art in automatic de-identification. Journal of the American Medical Informatics Association, 14(5):550â563, 2007. doi: 10.1197/jamia.M2444. [22] Maziyar Panahi. OpenMed NER: Open-source, domain-adapted state-of-the-art transformers for biomedical NER across 12 public datasets, 2025. URL https://arxiv.org/abs/2508.01630. arXiv:2508.01630. [23]Somalee Datta, Jose Posada, Garrick Olson, Wencheng Li, Ciaran OâReilly, Deepa Balraj, Joseph Mesterhazy, Joseph Pallas, Priyamvada Desai, and Nigam Shah. A new paradigm for accelerating clinical data science at Stanford Medicine, 2020. URL https://arxiv.org/abs/2003.10534. arXiv:2003.10534. 14 [24] Alison Callahan, Euan Ashley, Somalee Datta, Priyamvada Desai, Todd A. Ferris, Jason A. Fries, Michael Halaas, Curtis P. Langlotz, Sean Mackey, Jose D. Posada, Michael A. Pfeffer, and Nigam H. Shah. The stanford medicine data science ecosystem for clinical and translational research. JAMIA Open, 6(3):ooad054, 2023. doi: 10.1093/jamiaopen/ooad054. [25] David Carrell, Bradley Malin, John Aberdeen, Samuel Bayer, Cheryl Clark, Ben Wellner, and Lynette Hirschman. Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text. Journal of the American Medical Informatics Association, 20(2):342â348, 2013. doi: 10.1136/amiajnl-2012-001034. [26] Zhengliang Liu, Yue Huang, Xiaowei Yu, Lu Zhang, Zihao Wu, Chao Cao, Haixing Dai, Lin Zhao, Yiwei Li, Peng Shu, Fang Zeng, Lichao Sun, Wei Liu, Dinggang Shen, Quanzheng Li, Tianming Liu, Dajiang Zhu, and Xiang Li. DeID-GPT: Zero-shot medical text de-identification by GPT-4. arXiv preprint arXiv:2303.11032, 2023. [27]Isabella C. Wiest, Marie-Elisabeth LeĂmann, Fabian Wolf, Dyke Ferber, Marko Van Treeck, Jiefu Zhu, Matthias P. Ebert, Christoph Benedikt Westphalen, Martin Wermke, and Jakob Nikolas Kather. Deidentifying medical documents with local, privacy-preserving large language models: The LLM-anonymizer. NEJM AI, 2(4):AIdbp2400537, 2025. doi: 10.1056/AIdbp2400537. [28]Bayan Altallaâ, Sameera Abdalla, Ahmad Altamimi, Layla Bitar, Amal Al Omari, Ramiz Kardan, and Iyad Sultan. Evaluating GPT models for clinical note de-identification. Scientific Reports, 15(1):3852, 2025. doi: 10.1038/s41598-025-86890-3. [29]David Pissarra, Isabel Curioso, JoĂŁo Alveira, Duarte Pereira, Bruno Ribeiro, TomĂĄs Souper, Vasco Gomes, AndrĂ© Carreiro, and Vitor Rolla. Unlocking the potential of large language models for clinical text anonymization: A comparative study. In Proceedings of the Fifth Workshop on Privacy in Natural Language Processing (PrivateNLP), 2024. arXiv:2406.00062. [30] Guanchen Wu, Zuhui Chen, Yuzhang Xie, and Carl Yang. Towards automatic evaluation and selection of PHI de-identification models via multi-agent collaboration, 2025. Agents4Science 2025 Spotlight. [31]Xinli Tao, Xin Dong, and Xuezhong Zhou. OEMA: Ontology-enhanced multi-agent collaboration framework for zero-shot clinical named entity recognition. JAMIA Open, 2026. arXiv:2511.15211. [32]Murat Gunay, Bunyamin Keles, and Raife Hizlan. LLMs-in-the-Loop part 2: Expert small AI models for de-identification across 8 languages. arXiv preprint arXiv:2412.10918, 2024. [33]Woojin Kim, Sungeun Hahm, and Jaejin Lee. Generalizing clinical de-identification models by privacy-safe data augmentation using GPT-4. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024. [34]Thomas Sounack, Joshua Davis, Brigitte Durieux, Antoine Chaffin, Tom J. Pollard, Eric Lehman, Alistair E. W. Johnson, Matthew McDermott, Tristan Naumann, and Charlotta Lindvall. BioClinical ModernBERT: A state-of-the-art long-context encoder for biomedical and clinical NLP, 2025. URL https://arxiv.org/abs/2506.10896. arXiv:2506.10896. [35]Jan Hendrik Kirchner, Yining Chen, et al. Prover-verifier games improve legibility of LLM outputs. arXiv preprint arXiv:2407.13692, 2024. 15 [36] Michael Buckland and Fredric Gey. The relationship between recall and precision. Journal of the American Society for Information Science, 45(1):12â19, 1994. [37]Ăscar FerrĂĄndez, Brett R. South, Shuying Shen, F. Jeffrey Friedlin, Matthew H. Samore, and StĂ©phane M. Meystre. BoB, a best-of-breed automated text de-identification system for VHA clinical documents. Journal of the American Medical Informatics Association, 20(1):77â83, 2013. doi: 10.1136/amiajnl-2012-001020. [38]Kiana Aghakasiri, Noopur Zambare, JoAnn Thai, Carrie Ye, Mayur Mehta, J. Ross Mitchell, and Mohamed Abdalla. Not what the doctor ordered: Surveying LLM-based de-identification and quantifying clinical information loss. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32187â32203, Suzhou, China, November 2025. Association for Computational Linguistics. arXiv:2509.14464. [39]John X. Morris, Thomas R. Campion, Sri Laasya Nutheti, Yifan Peng, Akhil Raj, Ramin Zabih, and Curtis L. Cole. DIRI: Adversarial patient re-identification with large language models for evaluating de-identification. Proceedings of AMIA Annual Symposium, 2024. arXiv:2410.17035. [40]Jack Gallifant, Majid Afshar, Saleem Ameen, Yindalon Aphinyanaphongs, Shan Chen, Giovanni Cacciamani, Dina Demner- Fushman, Dmitriy Dligach, Roxana Daneshjou, Chrystinne Fernandes, Lasse Hyldig Hansen, Adam Landman, Lisa Lehmann, Liam G. McCoy, Timothy Miller, Amy Moreno, Nikolaj Munch, David Restrepo, Guergana Savova, Renato Umeton, Judy W. Gichoya, Gary S. Collins, Karel G. M. Moons, Leo Anthony Celi, and Danielle S. Bitterman. The TRIPOD-LLM reporting guideline for studies using large language models. Nature Medicine, 31(1):60â69, 2025. doi: 10.1038/s41591-024-03425-5. [41]OpenMed Science. OpenMed-PII-SuperClinical-Large-434M-v1: PII detection model. Hugging Face model repository, 2026. https://huggingface.co/OpenMed/OpenMed-PII-SuperClinical-Large-434M-v1. [42]Jorge Botas. PubliPlots: Publication-ready plotting for python.https://github.com/jorgebotas/publiplots, 2025. [43]Rachel Kuo, Andrew A. S. Soltan, Ciaran OâHanlon, Alan Hasanic, David A. Clifton, Gary Collins, Dominic Furniss, and David W. Eyre. Benchmarking transformer-based models for medical record de-identification in a single center multi-specialty evaluation. iScience, 28(12):113732, 2025. doi: 10.1016/j.isci.2025.113732. [44]Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica. Why do multi-agent LLM systems fail?, 2025. URL https://arxiv.org/abs/2503.13657. [45] Qineng Wang, Zihao Wang, Ying Su, Hanghang Tong, and Yangqiu Song. Rethinking the bounds of LLM reasoning: Are multi-agent discussions the key?, 2024. URL https://arxiv.org/abs/2402.18272. [46]Noopur Zambare, Kiana Aghakasiri, Carissa Lin, Carrie Ye, J. Ross Mitchell, and Mohamed Abdalla. Towards fair and efficient de-identification: Quantifying the efficiency and generalizability of de-identification approaches. In Findings of the Association for Computational Linguistics: EACL 2026. Association for Computational Linguistics, 2026. arXiv:2602.15869. 16 [47] David S Carrell, Bradley A Malin, David J Cronkite, John S Aberdeen, Cheryl Clark, Muqun Li, Dikshya Bastakoty, Steve Nyemba, and Lynette Hirschman. Resilience of clinical text de-identified with âhiding in plain sightâ to hostile reidentification attacks by human readers. Journal of the American Medical Informatics Association, 27(9):1374â1382, 2020. doi: 10.1093/jamia/ocaa095. 17 FIGURES A B C Figure 1: Study design and corpus characteristics. (A) Study design: 8 LLMs (top 4 shown in main figures) + traditional baselines evaluated on 100 pediatric oncology notes (5,322 spans) under 3 prompt conditions. (B) Synthetic clinical note illustrating canonical HIPAA PHI (names, dates; blue) vs. institutionally situated PHI (facility abbreviations, building names; orange). (C) Gold standard PHI distribution over the six categories withN â„20 (four further categories holdâ€5 spans each; all ten, with their denominators, are in Supplementary Table S4): Date (72.1%) and Name (22.3%) dominate, while institutionally enriched categories (Other Unique ID, Geographic Data) comprise 3.7% of spans. This distribution is that of the original gold standard; because that annotation under-counted institutional identifiers (see Results), 3.7% is a lower bound on their true prevalence. Alt text: Three-panel study-design figure. Panel A is a schematic showing eight large language models and traditional baselines being evaluated on 100 pediatric oncology notes containing 5,322 PHI spans under three prompt conditions. Panel B shows 18 a synthetic clinical note with canonical HIPAA identifiers highlighted in blue and institutionally situated identifiers (facility abbreviations, building names) highlighted in orange. Panel C is a bar or pie chart of the original gold-standard PHI category distribution over the six well-populated categories, dominated by Date (72.1 percent) and Name (22.3 percent), with institutionally enriched categories comprising 3.7 percent of spans, a lower bound given the under-annotation described in Results. Sonnet 4.6Opus 4.8GLM-5DeepSeek V3.2 TiDEOpenMed PII RegexspaCy +Regex 0.0 0.2 0.4 0.6 0.8 1.0 Score LLMsTraditional A Recall Precision F1 DateNameOUI*Geo*PhoneMRN Sonnet Opus GLM-5 DeepSeek TiDE OpenMed Regex spaCy 0.990.930.550.691.001.00 0.940.940.500.960.901.00 0.990.930.730.880.991.00 0.990.960.780.990.991.00 0.780.770.160.650.481.00 0.880.600.360.940.651.00 0.780.000.040.000.481.00 0.780.750.580.910.591.00 B 0.0 0.2 0.4 0.6 0.8 1.0 Recall 01020304050 False Negative Count Inst. Abbreviations Building/Facility Department IDs Ambiguous Dates Other Geographic Refs Provider Names Device Names Internal Staff Codes 38 15 8 49 40 25 24 8 4 C Institutional (61, 29%) Canonical (150, 71%) Figure 2: LLMs outperform traditional de-identification approaches. (A) Recall, precision, and F1 for the top 4 LLMs and 4 traditional baselines (Baseline prompt, 100 notes, 5,322 spans; all 8 LLMs in Table S3 and Supplementary Figure 1). (B) Per- category recall heatmap for the same systems, with categories ordered by gold-standard frequency and systems by overall F1, reveals shared weakness on the institutionally enriched categories (Other Unique ID [OUI], Geographic Data); the four categories withâ€5 gold spans are omitted here and reported in Table S4. (C) Sonnet 4.6 false negative decomposition on the primary Baseline run (Table S3; 211 FN): 29% are institutionally situated PHI (61 of 211; institution abbreviations 38, building/facility 15, department/clinic codes 8; see taxonomy in Table S7); the remaining 150 are canonical, of which the internal staff codes (pager and Voalte extensions, 4) are shown separately because Table S7 lists them as borderline. Alt text: Three-panel performance-comparison figure. Panel A is a grouped bar chart of recall, precision, and F1 for the top four LLMs and four traditional baselines under the Baseline prompt, showing LLMs above traditional systems on F1. Panel B is a 19 per-category recall heatmap in which the Other Unique ID and Geographic Data columns are lightest (lowest recall) across systems, indicating shared weakness on institutionally situated PHI. Panel C is a horizontal bar chart breaking down Sonnet 4.6âs 211 false negatives, with 29 percent (61 spans) attributed to institutionally situated PHI and the remaining 150 to canonical categories. 0.700.750.800.850.90 Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 A Precision 0.883 0.830 0.881 0.851 0.804 0.819 0.698 0.806 0.900.920.940.960.98 Recall 0.960 0.965 0.931 0.910 0.970 0.971 0.982 0.979 0.800.820.840.860.880.900.92 F1 0.920 0.892 0.905 0.880 0.879 0.889 0.816 0.884 â0.02â0.010.000.010.020.03 ÎRecall (Targeted â Baseline) Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 B +0.014 +0.023 -0.004 +0.001 â0.010.000.010.020.030.040.050.06 ÎPrecision (Precision â Targeted) +0.026 +0.004 +0.014 +0.053 Prompt stage (Panel A) BaselineTargetedPrecision Figure 3: In-context learning enables control over the precisionârecall trade-off. (A) Precision, recall, and F1 for the top 4 LLMs across the three prompt versions (Baselineâą, Targetedâ , PrecisionâČ; single run per model, all 8 in Table S5). Each horizontal line spans the minimum-to-maximum of a modelâs three stage values, so a non-monotonic Targeted stage visibly stretches the line rather than being hidden; Baseline and Precision values are labeled. Effect magnitude and direction are model- dependent. (B) Problem-specific deltas (single run), same model order and colors as (A), isolating what each prompt revision buys: âRecall(TargetedâBaseline) is the recall gain from institutional targeting, and âPrecision(PrecisionâTargeted) is the precision recovery from the anti-over-redaction instructions. The Precision prompt lifts precision for all models (cross-model mean over the 7 models with valid output: +0.054) and acts as a corrective: gains are largest for models that over-redact under the baseline prompt (GPT-oss-120B 0.198â0.924, Kimi 0.613â0.816, DeepSeek 0.698â0.806 across the full BaselineâPrecision progression; full values in Tables S5 and S6) and near-zero for models whose native precision is already high. Sonnet 4.6 reproducibility (5 trials per stage on the full 100-note corpus; SDâ€0.007 at every stage) confirms these are systematic prompt effects, not run-to-run noise (Table S9). Alt text: Two-panel figure on prompt effects. Panel A plots precision, recall, and F1 for the top four LLMs across three prompt versions (Baseline, Targeted, Precision) as horizontal min-to-max ranges per model, showing model-dependent magnitude and direction. Panel B is a bar chart of problem-specific deltas: the recall gain from institutional targeting and the precision recovery from anti-over-redaction instructions, with the largest precision gains for models that over-redact under the baseline prompt and near-zero gains for models with already-high native precision. 20 0.700.750.800.850.90 F1 Score Sonnet 4.6 (Precision) Sonnet 4.6 (Precision) + Structured P2 Sonnet 4.6 (Baseline) + Structured P2 Kimi K2.5 (Precision) Scrubber-Auditor (Precision) Cross-Model Vote Structured Dual-Pass Self-Heal (Recall) Same-Model Vote DeepSeek V3.2 (Precision) Self-Heal (Precision) Sonnet+Kimi Dual-Pass Sonnet 4.6 (Targeted) DeepSeek V3.2 (Targeted) Opus 4.8 (Targeted) TiDE Sonnet 4.6 (Precision) single-pass reference A Single-Pass Agentic Ensemble Traditional Scrubber-Auditor (Precision) Sonnet 4.6 (Precision) Sonnet 4.6 (Precision) + Structured P2 0.88 0.89 0.90 0.91 0.92 0.93 0.94 0.95 F1 Score =0.906 =0.018 =0.907 =0.002 =0.908 =0.005 B Baseline (n=5) Targeted (n=5) Precision (n=5) 0.80 0.85 0.90 0.95 1.00 Score C Recall Precision F1 Figure 4: Enhanced gold standard validation and agentic architecture comparison. (A) Configuration landscape: F1 for the 16 configurations that the panel displays, colour-coded by class (single-pass, agentic, ensemble, traditional) and ranked on the enhanced gold standard; 10 further configurations are omitted for legibility and appear in Table S12, which lists all 26. The omitted set is the weakest-performing tail, except that TiDE is retained as the traditional-system reference. Single-pass Sonnet (Precision) (dashed line) matches or exceeds every multi-agent variant, and self-refinement and ensemble voting fail to beat it. (B) Reproducibility: 5-trial strip plots for the top 3 methods. All three achieve closely comparable mean F1 (0.906â0.908) with overlapping 95% trial-resampled CIs, but the ScrubberâAuditor (Precision) pipeline on Opus shows high variance (SD F1 =0.018) vs. near-deterministic single-pass Sonnet (SD F1 =0.002). (C) Sonnet 4.6 prompt progression on the enhanced gold standard (10-note subset, 1,758 spans; 5 independent trials per version). Lines show mean; shading shows±SD; points show individual trials. Baseline recall is 0.847 here versus 0.958 on the original 100-note gold standard (a different note set and a different gold standard) because it misses newly-annotated institutional terms; Targeted recovers recall to 0.980; Precision maintains recall while improving precision (F1=0.907±0.002). Terminology: single-pass = one LLM call per note; dual-pass = two sequential LLM calls; structured pass 2 = a second pass that ingests the first passâs output in a structured format; ScrubberâAuditor = an agentic pipeline in which a generator LLM (scrubber) is checked by a second auditor LLM. Baseline, Targeted, and Precision denote the three prompt versions. Alt text: Three-panel figure on agentic architectures. Panel A ranks 16 configurations, spanning single-pass, agentic, ensemble, and traditional approaches, by F1 on the enhanced gold standard, with a dashed line marking single-pass Sonnet (Precision) at or above all multi-agent variants. Panel B shows 5-trial strip plots for the top three methods with closely overlapping mean F1 of 0.906 to 0.908, but visibly wider scatter for the Opus ScrubberâAuditor pipeline than for near-deterministic single-pass Sonnet. Panel C plots Sonnet 4.6 recall, precision, and F1 across the Baseline, Targeted, and Precision prompts on the enhanced gold standard, showing Baseline recall dropping to 0.847, Targeted recovering recall to 0.980, and Precision maintaining recall while improving precision to F1 = 0.907. *correspondence: zhandong.liu@bcm.edu, hyun-hwan.jeong@bcm.edu 1 EXTENDED SUPPLEMENTARY MATERIAL EXTENDED LIMITATIONS AND FUTURE DIRECTIONS The following points were raised as potential extensions to the primary analysis. Each would require additional experimentation, re-scoring, or external data beyond the scope of the present study, and we therefore document them here as explicit limitations and directions for future work rather than as new results. Robustness to paraphrase and residual re-identification risk Our evaluation measures whether direct and institutionally situated identifiers are removed (recall) and whether clinical content is preserved (precision). It does not measure resistance to re-identification by inference, for example an adversary reconstructing a patientâs identity from a constellation of residual quasi-identifiers (rare diagnoses, unusual treatment timelines, or distinctive narrative phrasing) even after all HIPAA-enumerated identifiers are redacted. Nor does it test whether an LLM adversary could re-identify individuals by paraphrasing or cross-referencing de-identified notes against external corpora. Prior work on âhiding in plain sightâ surrogate replacement [1,2] and adversarial re-identification [3] provides a framework for this analysis; applying such an adversarial protocol to our LLM outputs is an important direction for future work. Macro-averaged and note-level metric aggregation The precision, recall, and F1 values reported in the main text are micro-averaged over all gold spans, so frequent categories (e.g., Name, Date) dominate the aggregate. Macro-averaging across PHI categories, or reporting note-level distributions of these metrics, would give equal weight to rare but high-risk categories and would better characterize worst-case behavior on individual notes. The per-category recall heatmaps and per-note recall distributions (Supplementary Figures) provide a partial view, but a full macro-averaged re-tabulation of all systems and prompt versions is deferred to future work. Offset-based span scoring Our scoring aligns system output against the gold standard at the token/word level (via text alignment), counting a span as detected when the identifying text is redacted. A stricter offset-based scoring scheme, requiring exact character-boundary agreement between predicted and gold spans, would penalize partial-boundary matches (e.g., redacting âChildrenâs Hospitalâ when the gold span is âTexas Childrenâs Hospitalâ) that our current scheme may credit. Because boundary conventions differ across the LLM (free-text placeholder) and traditional (offset-emitting) systems, a fully offset-based comparison would require re-normalizing all outputs to a common span representation. We expect the qualitative ranking to be robust to this choice but have not quantified it. Model and API version identifiers The eight LLMs are identified in the main text by their public model names and sizes. Exact provider model identifiers govern reproducibility, since hosted models can be updated silently. Table E1 gives the exact identifier used to invoke each model. 2 Table E1: Exact AWS Bedrock model identifiers for the eight benchmarked LLMs. All models were invoked through the Bedrock Converse API in regionus-east-1attemperature=0.0(except Opus 4.8, which rejects the parameter),max_tokens=65,000, with up to 3 retries under exponential backoff. Name in textBedrock model identifierProvider Claude Opus 4.8 us.anthropic.claude-opus-4-8 â Anthropic Claude Sonnet 4.6 us.anthropic.claude-sonnet-4-6 â Anthropic GPT-oss-120B openai.gpt-oss-120b-1:0OpenAI GPT-oss-20B openai.gpt-oss-20b-1:0OpenAI GLM-5 zai.glm-5Zhipu AI Kimi K2.5 moonshotai.kimi-k2.5Moonshot MiniMax M2.5 minimax.minimax-m2.5MiniMax DeepSeek V3.2 deepseek.v3.2DeepSeek â Cross-region inference profile identifier (the us. prefix). On-demand invocation of the bare foundation-model identifier fails with a ValidationException for these models. All eight were verified to accept the full 65,000-token output ceiling, required so that the longest notes (âŒ65.8k characters) can be reproduced without truncation. Inference dates. All LLM inference reported in this study was executed between 2026-05-30 and 2026-06-28: the primary-corpus runs across all eight models and three prompt conditions between 2026-05-30 and 2026-06-04, the Sonnet 4.6 reproducibility trials through 2026-06-28, and the USDHUB validation runs on 2026-06-25 and 2026-06-26. These bounds are derived from the timestamps of the stored run artifacts; per-invocation timestamps were not recorded by the evaluation harness, so we report the execution window rather than exact per-model inference times. Because Bedrock model endpoints can be updated without a change to the identifier, results should be interpreted as characterizing these endpoints as served during that window. John Snow Labs Spark NLP baseline We were unable to include John Snow Labsâ Spark NLP for Healthcare de-identification pipeline, a widely used commercial baseline, because it was unavailable to us under its licensing terms during the study period. Its omission means our non-LLM comparison set (TiDE, OpenMed PII, spaCy+regex, regex-only) does not include a fully-provisioned commercial system, which may understate the strongest achievable traditional-pipeline performance. Benchmarking against Spark NLP is left for future work. 3 EXTENDED DETAILED RESULTS AND METHODS This section collects the detailed tables, full prompt text, error case studies, reproducibility analyses, and cost/latency breakdowns that support the headline results summarized in the main JAMIA supplement (Supplement 1). Nothing here is new work; it is the complete detail behind the summarized findings, relocated so that Supplement 1 remains focused on the key tables and figures a reader needs to follow the main text. Section E2: Full Prompt Text E2.1: Baseline Prompt (de-identification-only mode) System prompt: You are a medical text deidentification expert. Your task is to return the EXACT same text as provided, with PHI (Protected Health Information) replaced by generic placeholders. DO NOT change the structure, formatting, or wording of the text in any way. Simply replace identifiers with placeholders. CRITICAL: Return the text EXACTLY as given, with ONLY the PHI replaced by placeholders. User prompt enumerates the 18 HIPAA Safe Harbor identifier categories with examples: "Shelby Smith is a 3yo F" becomes "[NAME] is a 3yo F"; "She was born on 3/1/11" becomes "She was born on [DATE]". E2.2: Targeted Prompt (adds institutional PHI) Appends to Baseline a block titled "ADDITIONAL REQUIRED REDACTIONS" with four categories: (A) physician/staff names adjacent to credentials, (B) medical-staff pager and Voalte numbers, (C) institution names and abbreviations (TCH, TXCH, BCM), and (D) building/facility names. Each includes WRONGâ RIGHT example pairs. E2.3: Precision Prompt (adds anti-over-redaction) Retains all Targeted content and appends a "DO NOT OVER-REDACT" block targeting six false-positive categories with WRONGâRIGHT examples: (1) clock times, (2) lab/vital values, (3) medication doses, (4) ages under 90, (5) relationship words, (6) clinical abbreviations. Closes with a safety caveat: "when genuinely ambiguous and could identify a specific person, prefer to redact." E2.4: ScrubberâAuditor Pipeline Prompts Stage 1: Sonnet Scrubber (recall-maximizing): You are a specialized PHI scrubbing engine. [...] BE THOROUGH. It is critical that NO PHI leaks through. When in doubt, REDACT. [...] ALSO REDACT these commonly missed institutional identifiers: hospital names AND abbreviations, NAMED building/facility names, device/DME 4 company names, pharmacy names, staff pager numbers. [...] DO NOT REDACT: clock times, generic location descriptors (OR, PACU, ICU, Pod A), clinical department TYPES without institution name, procedure codes, medication names, lab values, vital signs. Stage 2: Opus Precision Auditor: You are a PHI precision auditor. [...] RESTORATION RULES: restore a [PLACEHOLDER] back to its original value if the original value is OBVIOUSLY: a medication name/dosage, lab result, diagnosis, anatomical term, medical abbreviation, vital sign, procedure name, clock TIME, or generic department name. [...] NEVER RESTORE: calendar DATES, person NAMES, SPECIFIC institution names (TCH, BCM, Baylor), SPECIFIC building names (Mark Wallace Tower, West Campus), phone/pager numbers, device BRAND names, ANYTHING uncertain. Auditor execution. The Precision Auditor runs once per note over the entire scrubbed document, not independently per placeholder. The auditor prompt receives the full original note (as PHI-containing reference) and the full scrubbed note in a single call, and returns the complete precision-audited note; restorations are applied wherever the model judges a placeholderâs original value to be obviously clinical content under the rules above. Prompt-scope caveat. The scrubber prompt used in the pipeline experiments retained a device-name clause from an earlier prompt iteration. Device identifiers were not a PHI category in this corpus, so any device-name redactions it produced could only be scored as false positives; the clause is reported here for exact reproducibility. Full prompt text available in the code repository:src/phi_benchmark/prompts/(Baseline, Targeted, Precision) and src/phi_benchmark/pipelines/prompts.py (pipeline stages). Section E3: Error Case Studies These case studies expand on the error taxonomy summarized in Supplement 1 (institutionally situated PHI taxonomy), providing the verbatim note excerpt behind the institutional-abbreviation example. Case study 1: institutional abbreviation recovery. The abbreviation "TCH" appeared as a false negative for most LLMs under Baseline. In context: Patient was seen at TCH Main Campus Hematology Center... Under Baseline, models treated "TCH" as a clinical abbreviation. Under Targeted, Sonnet 4.6 recovered 48/61 institutional PHI misses (across all subcategories). The 13 remaining occurred in complex compound expressions where institutional names were embedded within clinical descriptors. Section E4: Multi-Agent Reproducibility, Self-Refinement, Failure Analysis, and Cost These analyses expand on the multi-agent workflow experiments summarized in Supplement 1 (architecture summary and all-configurations results tables). The headline conclusion, that no agentic configuration outperformed calibrated single-pass de- 5 identification, is stated in Supplement 1; the reproducibility trials, statistical tests, failure mechanisms, self-refinement experiments, and cost/latency accounting behind that conclusion are given here. E4.1: Reproducibility and Run-to-Run Variability To quantify run-to-run variability we conducted 5 independent trials of the three highest-F1 configurations from the all- configurations table (Supplement 1): the agentic ScrubberâAuditor (Precision) (Opus), Sonnet 4.6 single-pass (Precision prompt), and Sonnet (Precision) + Structured Pass 2. Each trial was run independently against the enhanced gold standard (10-note subset, 1,758 spans) using identical prompts and model parameters (temperature=0). Sources of non-determinism. All models are called via the AWS Bedrock Converse API with temperature=0.0 and identical inputs across trials. However,temperature=0does not guarantee bitwise-identical outputs from cloud-served LLMs due to GPU floating-point non-determinism (non-associative parallel arithmetic), infrastructure routing across heterogeneous hardware, and cascade amplification (a single token flip early in a 20â33K character output changes the entire downstream sequence). Additionally, Claude Opus 4.8 (used by the ScrubberâAuditor pipeline) rejects thetemperatureparameter via Bedrock ("temperature is deprecated for this model"), so greedy decoding cannot be explicitly enforced for this model. This partly explains why the Opus-based agentic pipeline exhibits higher variance (SD F1 =0.018) than the Sonnet-based single-pass (SD F1 =0.002). Noseed parameter is available via the Bedrock Converse API; the observed variability is entirely attributable to infrastructure-level non-determinism, not any parameter under experimenter control. The per-trial run-to-run variability table for the top-3 methods (5 trials each, enhanced gold standard) is given in Supplement 1, Supplementary Note 3.3; it is not duplicated here. Variance profiles. The headline reproducibility statistics (mean F1, bootstrap CIs, and the significance tests) are reported in Supplement 1, Supplementary Note 3.3. Beyond those, the three methods differ markedly in where their variance originates: âąSingle-pass Sonnet (Precision) is near-deterministic (SD F1 =0.002, range 0.905â0.910) with stable recall (0.981) and precision (0.840â0.848). âąSonnet (Precision) + Structured P2 adds marginal recall (mean 0.982 vs 0.981) with low additional variance (SD F1 =0.005). âąScrubberâAuditor (Precision) exhibits high variance (SD F1 =0.018, range 0.890â0.942) driven almost entirely by precision instability (0.836â0.935). When the Precision Auditor aggressively restores over-redactions (Trial 2, FP=116), F1 reaches 0.942; when it is conservative (Trial 4, FP=327), F1 drops to 0.890. The agentic pipelineâs best runs (F1=0.942) substantially exceed any single-pass result, motivating the self-refinement strategies evaluated in Section E4.3. The per-trial prompt-version reproducibility table on the enhanced gold standard is given in Supplement 1, Supplementary Note 3.3; it is not duplicated here. 6 The per-trial full-corpus prompt-version reproducibility table (100 notes, 5,322 gold spans) is given in Supplement 1, Supplementary Note 3.3; it is not duplicated here. Residual false positives on enhanced gold.Even under the enhanced gold, Sonnet (Precision) reports 310â329 FP across 5 trials (mean precision=0.844±0.003; per-trial values in Supplement 1, Supplementary Note 3.3). Informal review of a sample suggests that many of these represent additional annotation gaps (e.g., dates in vitals sections and follow-up appointments not marked in the gold standard), but we have not performed systematic categorization or a second re-annotation pass. The residual FP count should therefore be interpreted as an upper bound on true over-redaction. E4.2: Failure Analysis Why a second pass adds little (the Precision prompt already captures 222/227 institutional spans; the structured Pass 2 adds only 20 replacements; the dedicated auditor restores zero) and how the same pipeline surfaced the 414 candidates behind the enhanced gold standard are summarized in Supplement 1, Supplementary Note 3.4. The mechanism behind free-form dual-pass instability, detailed next, is specific to this section. Dual-pass variability and rewriting artifacts.The standard dual-pass architecture (Baseline Pass 1 + institutional recall Pass 2) showed high run-to-run variability across 3 independent runs on the original 100-note gold standard (5,322 spans): FP counts ranged from 532 to 1,524 (mean F1=0.904±0.043 vs. single-pass F1=0.918±0.001). Two runs produced FP in the 500â700 range; one run produced FP=1,524 with bimodal per-note prediction patterns (5â10 vs. 150â220 predictions per note). This variability, regardless of its source, underscores the architectural fragility of free-form Pass 2 rewriting. The underlying failure mechanism is Pass 2âs free-form text reformatting, which accumulates alignment-based false positives. On the enhanced 10-note gold, dual-pass scores F1=0.782 (all-configurations table, Supplement 1), reflecting this fundamental limitation. Dual-pass recall on the enhanced gold (0.850) is essentially identical to Baseline single-pass (0.850), despite a small recall gain on the full corpus. The 10-note subset was selected for high institutional PHI density, precisely the notes where Pass 2 triggers mode-collapse (generating commentary or over-condensed output that fails the validity check). The recall gain is concentrated in notes with moderate institutional PHI density where Pass 2 operates normally. The structured JSON output constraint (SP(Baseline) + Structured P2) eliminates this failure mode by preventing the model from generating free-form redacted text in Pass 2. Multi-pass architectures have been reported to improve recall in settings where notes are chunked before processing [4]. We attribute the divergence from our findings to input length: their gains derive from re-attending to chunks that do not fit a single context window, whereas our notes are processed whole, so Pass 2 receives no information unavailable to Pass 1. E4.3: Self-Refinement Experiments The high variance of the ScrubberâAuditor (Precision) pipeline (SD F1 =0.018; Section E4.1) suggests that some runs produce substantially suboptimal output. We investigated whether lightweight self-refinement (Self-Refine [5], Reflexion [6]) or ensemble strategies could stabilize performance. 7 Self-Heal (Recall). A Sonnet verifier examines the ScrubberâAuditor (Precision) output for residual PHI; if flagged, the full pipeline retries with reflexion-style feedback. Result: F1=0.899, R=0.958, P=0.846. The verifier flagged 2/10 notes, but the pipeline did not outperform the base mean (F1=0.906) because the dominant variance source is precision instability (FP range 116â327), not recall failures (FN range 75â90). Self-Heal (Precision).A Sonnet verifier reviews all placeholders for over-redaction; ifâ„4 are detected, only the Auditor retries with a targeted restoration list. Result: F1=0.897, R=0.958, P=0.844. The verifier triggered zero retries: over-redactions are distributed at 1â3 per note across ambiguous categories (institution abbreviations, time-of-day values, pager numbers), never reaching the threshold. Same-Model Majority Vote (3ĂSA-Precision). Three independent runs of ScrubberâAuditor (Precision); span-level majority vote (â„2/3). Result: F1=0.899, R=0.959, P=0.846. The three runs scored F1â[0.902, 0.906] with FP counts of 293, 286, and 296. Because over-redactions are highly correlated across runs (the same ambiguous spans are consistently over-redacted), majority voting cannot filter them. Cross-Model Majority Vote (Sonnet+DeepSeek+Kimi).Three models each run Precision independently; span-level majority vote (â„2/3). These models have diverse error profiles (Jaccard FN overlap 0.30â0.43). Result (3 trials): Mean F1=0.902 (SD=0.006), R=0.986, P=0.832. The ensemble achieves the highest recall of any configuration (0.986) but precision decreases relative to Sonnet alone (0.832 vs. the 5-trial single-pass mean of 0.844) because weaker models contribute spurious redactions that survive the 2/3 threshold. Discussion.No self-refinement or ensemble strategy improves over Sonnet (Precision) single-pass (mean F1=0.907, SD=0.002). Self-Heal verifiers cannot detect the distributed, low-density over-redaction pattern that drives FP variance. Same-model voting fails because over-redactions are correlated across runs. Cross-model voting improves recall to near-ceiling (0.986) but weaker models degrade precision, yielding no net F1 gain. Self-refinement and ensemble costs ranged from 1.7Ăto 6Ăthe single-pass baseline (17â60 LLM calls vs. 10) without F1 improvement. The fundamental insight is that the agentic pipelineâs high-variance, high-ceiling behavior (best run F1=0.942) cannot be reliably extracted through simple post-hoc strategies. The rare optimal runs arise from stochastic variation in the Auditorâs restoration aggressiveness, a property not amenable to consensus-based or verifier-based recovery. For this task and evaluation set, well-calibrated prompt engineering on a single strong model represents the practical performance ceiling. E4.4: Cost and Latency Multi-agent pipelines require 2â6 LLM calls per note vs. 1 for single-pass: 2 for dual-pass and ScrubberâAuditor, 3 for cross-model voting, and 6 for same-model voting (3 runsĂthe 2-call ScrubberâAuditor pipeline). The SP(Baseline) + Structured P2 pipeline reuses pre-computed Baseline output, completing inâŒ1.9 minutes for 100 notes (concurrent requests). A full end-to-end dual-pass takesâŒ90 minutes vs.âŒ45 minutes for single-pass. Self-refinement and ensemble strategies compound this overhead: Self-Heal variants use 1.7Ăthe base cost, while voting approaches require 3â6Ă(30 LLM calls for cross-model vote, 60 for same-model vote). At Sonnet 4.6 pricing ($0.009/note for single-pass on the primary benchmark), the ScrubberâAuditor pipeline adds Opus 4.8 8 costs ($0.044/note), making the agentic approachâŒ6Ămore expensive per note. Since none of the tested agentic configurations improved F1 over Precision single-pass, the additional cost and latency are justified only when maximizing recall is the sole priority or when using multi-agent workflows as a discovery tool for gold standard refinement. Section E5: Independent Validation Cost and Latency This section provides the per-system cost and latency breakdown for the independent validation set, supporting the results summa- rized in Supplement 1, Supplementary Note 4. The TiDE provisioning configurations themselves are derived in Supplement 1, Supplementary Note 4.1. Table E2: Cost and latency comparison on the USDHUB validation dataset (49 notes, sequential processing). LLM costs reflect AWS Bedrock on-demand pricing; local systems have zero marginal API cost (compute is pre-provisioned on the evaluation instance). SystemLatencyCost (49 notes)Cost/notePlatform Claude Opus 4.86.8 min$2.142$0.0437Bedrock API Claude Sonnet 4.64.6 min$0.428$0.0087Bedrock API GPT-oss-120B5.6 min$0.400$0.0082Bedrock API DeepSeek V3.26.4 min$0.166$0.0034Bedrock API Kimi K2.53.3 min$0.166$0.0034Bedrock API GLM-511.5 min$0.166$0.0034Bedrock API MiniMax M2.517.3 min$0.166$0.0034Bedrock API GPT-oss-20B61.7 min$0.043$0.0009Bedrock API TiDE27 sâLocal (Java) OpenMed PII45 sâLocal (434M) spaCy NER + Regex3.5 sâLocal (Python) Regex-only<1 sâLocal (Python) Latency measured end-to-end (sequential, single-threaded). LLM latency includes network round-trip and API queuing; local systems include model loading. Cost estimated fromâŒ45K input +âŒ20K output tokens at published Bedrock rates. LLM API costs ranged from $0.04 (GPT-oss-20B) to $2.14 (Claude Opus 4.8) for the full 49-note corpus, corresponding to $0.0009â$0.044 per note. The highest-F1 system (Opus) is also the most expensive per note, while three third-party models (DeepSeek, Kimi, GLM-5) cluster at $0.003/note, roughly 13Ăcheaper than Opus with recall within 3 percentage points (MiniMax clusters at similar cost but with 20-point lower recall). Local systems (TiDE, spaCy, OpenMed, regex) incur zero marginal API cost and run one to three orders of magnitude faster (3.5 s for spaCy and 27 s for TiDE, versus 3.3â61.7 min for the LLMs). GPT-oss-20Bâs anomalous latency (61.7 min) reflects persistent API errors requiring retries, not model inference time. 9 REFERENCES [1]David Carrell, Bradley Malin, John Aberdeen, Samuel Bayer, Cheryl Clark, Ben Wellner, and Lynette Hirschman. Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text. Journal of the American Medical Informatics Association, 20(2):342â348, 2013. doi: 10.1136/amiajnl-2012-001034. [2] David S Carrell, Bradley A Malin, David J Cronkite, John S Aberdeen, Cheryl Clark, Muqun Li, Dikshya Bastakoty, Steve Nyemba, and Lynette Hirschman. Resilience of clinical text de-identified with âhiding in plain sightâ to hostile reidentification attacks by human readers. Journal of the American Medical Informatics Association, 27(9):1374â1382, 2020. doi: 10.1093/jamia/ocaa095. [3]John X. Morris, Thomas R. Campion, Sri Laasya Nutheti, Yifan Peng, Akhil Raj, Ramin Zabih, and Curtis L. Cole. DIRI: Adversarial patient re-identification with large language models for evaluating de-identification. Proceedings of AMIA Annual Symposium, 2024. arXiv:2410.17035. [4]Praphul Singh, Charlotte Dzialo, Jangwon Kim, Sumana Srivatsa, Irfan Bulu, Sri Gadde, and Krishnaram Kenthapadi. Redac- tOR: An LLM-powered framework for automatic clinical data de-identification. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics: Industry Track, Vienna, Austria, 2025. Association for Computational Linguistics. arXiv:2505.18380. [5] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems, 36, 2023. [6]Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems, 36, 2023. 10 EXTENDED SUPPLEMENTARY FIGURES claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120bgpt-oss-20b kimi-k2.5minimax-m2.5 claude-opus-4.8 tide-ner-pattern â0.2 0.0 0.2 0.4 0.6 0.8 1.0 1.2 per-note recall Per-note recall distribution â deidentification_only Supplementary Figure E1: Per-note recall distribution (Baseline). Each point represents one clinical note. Most top models cluster near perfect recall (1.0), but all exhibit a tail of low-recall notes; these correspond to notes rich in institutionally situated PHI. claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120bgpt-oss-20b kimi-k2.5minimax-m2.5 claude-opus-4.8 tide-ner-pattern 0.0 0.1 0.2 0.3 0.4 0.5 0.6 fraction of notes with zero leaked PHI Note-level perfect de-id rate â deidentification_only Supplementary Figure E2: Perfect de-identification rate (Baseline). Fraction of notes with zero leaked PHI spans. DeepSeek and GLM-5 achieve perfect de-identification onâŒ60% of notes; no model exceeds 65%, reflecting the long tail of institutionally situated identifiers. 11 DateEmailFax Geographic Data License/Cert NumberMRNName Other Unique ID PhoneURL PHI type â1.0 â0.8 â0.6 â0.4 â0.2 0.0 0.2 0.4 Î recall Recall delta vs baseline â deidentification_only model_name claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120b gpt-oss-20b kimi-k2.5 minimax-m2.5 Supplementary Figure E3: Per-category recall change under targeted prompting (TargetedâBaseline). The targeted prompt yields concentrated gains in Other Unique ID and Name (institutionally situated categories) for responsive models, while MiniMax M2.5 shows broad regressions. Geographic Data shows mixed effects. claude-opus-4.8 claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120bgpt-oss-20b kimi-k2.5minimax-m2.5 0.0 0.2 0.4 0.6 0.8 1.0 score Overall metrics â deidentification_only metric recall precision f1_score f2_score Supplementary Figure E4: Overall performance under precision-focused prompt (Precision). The "do-not-over-redact" instructions improve precision across most models relative to Baseline/Targeted, with GPT-oss-120B showing the most dramatic improvement. GPT-oss-20B suffers output failures (reasoning exhausts token budget). 12 DateEmailFax Geographic Data License/Cert NumberMRNName Other Unique ID PhoneURL PHI type â0.4 â0.3 â0.2 â0.1 0.0 0.1 0.2 Î recall Recall delta vs baseline â deidentification_only model_name claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120b gpt-oss-20b kimi-k2.5 minimax-m2.5 Supplementary Figure E5: Per-category recall change (PrecisionâBaseline). Unlike Targeted (Supplementary Figure E3), the Precision prompt shows smaller recall deltas for strong models, confirming that precision can be improved without substantial recall degradation. GPT-oss-20B and MiniMax M2.5 are outliers with large regressions. claude-opus-4.8 claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120bgpt-oss-20b kimi-k2.5minimax-m2.5 0.0 0.1 0.2 0.3 0.4 0.5 0.6 fraction of notes with zero leaked PHI Note-level perfect de-id rate â deidentification_only Supplementary Figure E6: Perfect de-identification rate (Precision). Under the precision-focused prompt, DeepSeek V3.2 and Kimi K2.5 maintain the highest perfect de-id rates (âŒ60%), while GPT-oss-20B and MiniMax M2.5 show substantial degradation. Compare with Supplementary Figure E2 for baseline. 13 claude-opus-4.8 claude-sonnet-4.6 deepseek-v3.2 glm-5 gpt-oss-120bgpt-oss-20b kimi-k2.5minimax-m2.5 â0.25 0.00 0.25 0.50 0.75 1.00 1.25 per-note recall Per-note recall distribution â deidentification_only Supplementary Figure E7: Per-note recall distribution (Precision). Compared to Baseline (Supplementary Figure E1), the distribution remains concentrated near 1.0 for strong models but shows increased variance for GPT-oss-20B and MiniMax M2.5 under the precision-focused prompt. 0.50.60.70.80.91.0 Recall Primary Benchmark (100 notes) 0.5 0.6 0.7 0.8 0.9 1.0 Recall USDHUB Validation (49 notes) Sonnet 4.6 DeepSeek V3.2 GLM-5 GPT-oss-120B GPT-oss-20B Kimi K2.5 MiniMax M2.5 Opus 4.8 TiDE OpenMed PII Regex spaCy+Regex Cross-Dataset Recall Correlation LLM Traditional y = x Supplementary Figure E8: Cross-dataset recall correlation. Each point is one system; x-axis is recall on the primary benchmark, y-axis is recall on USDHUB. Points near the diagonal indicate consistent performance across corpora. TiDE shows the largest positive deviation (higher recall on USDHUB than primary). On USDHUB, TiDE was additionally run with known-PHI matching enabled (a supplied per-note dictionary of identifiers; see Supplement 1, Supplementary Note 4.1), a near-oracle configuration not available to the LLMs, which accounts for much of this deviation; USDHUB also contains standard demographic PHI without the institutionally situated identifiers that challenge pattern-based systems. 14 10 3 10 2 10 1 10 0 Cost per 49 notes (USD, log scale) 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 F1 Score Opus 4.8 GPT-oss-120B Sonnet 4.6 TiDE GLM-5 GPT-oss-20B DeepSeek V3.2 MiniMax M2.5 OpenMed PII Kimi K2.5 Regex spaCy+Regex CostPerformance Tradeoff (USDHUB, v1 baseline) LLM (API) Traditional (local) Supplementary Figure E9: Costâperformance tradeoff (USDHUB, Baseline). F1 score vs. total API cost for the 49 notes (log scale). Traditional baselines (orange squares) incur zero marginal API cost. Among LLMs (blue circles), Opus 4.8 achieves the highest F1 at the highest cost ($2.14 for the corpus, $0.044/note), while DeepSeek, Kimi, GLM-5 and MiniMax all cost $0.17 for the corpus ($0.003/note) at F1 0.67â0.84. The costâF1 Pareto frontier is GPT-oss-20B ($0.04), GLM-5 ($0.17), GPT-oss-120B ($0.40) and Opus 4.8 ($2.14); DeepSeek, Kimi and MiniMax are dominated. GPT-oss-20B is cheapest only because reasoning- token exhaustion truncated some outputs, so it is not a usable operating point despite lying on the frontier. 10 1 10 0 10 1 10 2 10 3 Latency (seconds, log scale) 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 F1 Score Opus 4.8 Sonnet 4.6 Kimi K2.5 DeepSeek V3.2 GPT-oss-120B GLM-5 GPT-oss-20B MiniMax M2.5 TiDE spaCy+Regex OpenMed PII Regex LatencyPerformance Tradeoff (USDHUB, 49 notes, sequential) LLM (API) Traditional (local) Supplementary Figure E10: Latencyâperformance tradeoff (USDHUB, 49 notes, sequential). End-to-end processing time vs. F1. Local systems (TiDE, spaCy, OpenMed, regex) complete in under 1 minute with no API dependency. LLMs range from 3â62 minutes depending on model and API stability. GPT-oss-20Bâs 62-minute runtime reflects retry overhead from persistent API errors, not inference speed. 15 Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 MiniMax M2.5 Kimi K2.5 GPT-oss-120B GPT-oss-20B TiDE OpenMed PII Regex spaCy+Regex 0.0 0.2 0.4 0.6 0.8 1.0 Score LLMsTraditional USDHUB External Validation (49 notes, 409 gold spans, Baseline) Recall Precision F1 Supplementary Figure E11: Overall performance on USDHUB validation (49 notes, 409 gold spans, Baseline). LLMs are evaluated under identical conditions as the primary benchmark. TiDE achieves the highest recall (0.976 fully unprovisioned; 0.971 maximally provisioned with a known-PHI dictionary and per-note header, a difference of<0.6 recall points on this canonical-PHI corpus; Supplement 1, Supplementary Note 4.1) while Opus 4.8 leads on F1 (0.894). The LLMâtraditional gap is clearly visible, with traditional methods clustered below F1=0.70 except TiDE. 16 ADDRESS DATE ID NAME OTHER PHONE Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 MiniMax M2.5 Kimi K2.5 GPT-oss-120B GPT-oss-20B OpenMed PII Regex spaCy+Regex 1.000.991.000.940.711.00 1.001.001.000.930.641.00 1.001.000.960.910.640.88 1.001.001.000.940.711.00 1.000.690.910.780.071.00 1.001.001.000.960.711.00 1.001.000.930.910.791.00 1.000.880.870.830.291.00 1.001.000.780.490.860.54 0.001.000.780.120.431.00 1.001.000.820.660.501.00 Per-Category Recall â USDHUB External Validation (Baseline) 0.0 0.2 0.4 0.6 0.8 1.0 Recall Supplementary Figure E12: Per-category recall heatmap, USDHUB validation. PHI categories in USDHUB (Name, Date, ID, Phone, Address, Other) show similar patterns to the primary benchmark: LLMs achieve near-perfect recall on most categories while traditional baselines show systematic gaps, particularly on Names. 17 Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 MiniMax M2.5 Kimi K2.5 GPT-oss-120B GPT-oss-20B TiDE OpenMed PII Regex spaCy+Regex 0.0 0.2 0.4 0.6 0.8 1.0 Recall A. Recall: Primary vs. USDHUB Primary (100 notes) USDHUB (49 notes) Sonnet 4.6 Opus 4.8 GLM-5 DeepSeek V3.2 MiniMax M2.5 Kimi K2.5 GPT-oss-120B GPT-oss-20B TiDE OpenMed PII Regex spaCy+Regex 0.0 0.2 0.4 0.6 0.8 1.0 F1 Score B. F1: Primary vs. USDHUB Primary (100 notes) USDHUB (49 notes) Cross-Dataset Performance Comparison (Baseline prompt) Supplementary Figure E13: Cross-dataset performance comparison. Side-by-side recall (A) and F1 (B) for all systems on the primary benchmark (100 notes, blue) vs. USDHUB validation (49 notes, orange). Both datasets use the Baseline prompt. Performance rankings are largely preserved across datasets, supporting generalizability of the findings.