Paper deep dive
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Jin Mu, Guanhua Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/29/2026, 3:01:15 AM
Summary
The paper introduces CAST (Concept-guided Artifact Suppression Tuning), a framework for making clinical language models auditable by using Sparse Autoencoders (SAEs) to extract and suppress non-clinical artifacts (like templates and boilerplate) during fine-tuning. CAST improves mortality prediction performance on MIMIC-IV discharge notes compared to standard fine-tuning baselines while providing a transparent, feature-level audit trail of the clinical concepts driving predictions.
Entities (12)
Relation Signals (10)
CAST → appliedto → MIMIC-IV
confidence 95% · On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines
CAST → performstask → Mortality Prediction
confidence 95% · On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines
CAST → usescomponent → Sparse Autoencoder
confidence 95% · CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations
CAST → employstechnique → Artifact Suppression
confidence 90% · suppresses verified artifact latents via residual subtraction during fine-tuning
CAST → usesstandard → ICD-10
confidence 90% · labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints
CAST → outperforms → SAE-Probe
confidence 85% · CAST... F1 0.2961... SAE-Probe... F1 0.2632
CAST → outperforms → Self-Regul
confidence 85% · CAST... F1 0.2961... Self-Regul... F1 0.1663
CAST → outperforms → ClinicalBERT
confidence 85% · CAST improves over its corresponding fine-tuned encoder baselines... ClinicalBERT... Fine-tuning... F1 0.2602 vs CAST 0.2961
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints, suppresses verified artifact latents via residual subtraction during fine-tuning, and provides post-hoc per-concept attributions for auditing model decisions. On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training.
Tags
Links
- Source: https://arxiv.org/abs/2608.27397v1
- Canonical: https://arxiv.org/abs/2608.27397v1
Trouble viewing inline? Open PDF directly →
Full Text
81,985 characters extracted from source content.
Expand or collapse full text
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction Jin Mu Affiliation: University of Wisconsin–Madison Email: jmu27@wisc.edu Guanhua Chen Affiliation: University of Wisconsin–Madison Email: gchen25@wisc.edu Abstract Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, human-auditable features from intermediate Transformer activations, labels SAE latents with an LLM-assisted interpretation pipeline and ICD-10 retrieval constraints, suppresses verified artifact latents via residual subtraction during fine-tuning, and provides post-hoc per-concept attributions for auditing model decisions. On MIMIC-IV discharge-note mortality prediction, CAST improves over its corresponding fine-tuned encoder baselines and remains competitive with strong LLM baselines, while producing a feature-level audit trail of the clinical concepts that support each prediction and the artifact concepts suppressed during training. 1 Introduction Figure 1: SAE-Guided Fine-Tuning Enables Transparent Clinical Risk Prediction. Unlike standard fine-tuning, which yields opaque risk predictions, our framework encourages models to rely on clinically meaningful SAE concepts and suppress non-clinical artifacts, producing predictions with interpretable clinical evidence and ICD-10 codes. Clinical language models have shown significant potential to transform clinical workflows by processing unstructured electronic health record (EHR) data for tasks such as clinical summarization, diagnostic reasoning, and risk prediction (Maity and Saikia, 2025; Meng et al., 2024; Thirunavukarasu et al., 2023). Compared with structured variables alone, free-text clinical notes contain rich longitudinal information about patient history, disease progression, clinician assessments, and treatment decisions. This makes them especially valuable for high-acuity prediction tasks such as mortality risk estimation. However, despite strong in-distribution performance, the deployment of clinical LLMs remains limited in high-stakes medical settings because their predictions are often difficult to interpret, validate, and trust (Amann et al., 2020; Markus et al., 2021). A central obstacle is shortcut learning: models trained on EHR notes may exploit artifact-based signals that are highly predictive within a dataset but clinically meaningless or unstable under deployment shift. These signals can include note templates, section headers, documentation style, repeated boilerplate, separators, discharge formatting, or institution-specific coding patterns, which may correlate with outcomes because of local documentation practices rather than true patient physiology (Geirhos et al., 2020). This concern is especially acute in Intensive Care Unit (ICU) settings, where risk predictions may inform time-sensitive decisions. Even when a model correctly identifies a high-risk patient, the prediction has limited clinical value if it is driven by such shortcuts rather than evidence of clinical deterioration. Clinicians cannot justify decisions such as escalating life support or initiating palliative care based on an opaque alert. Clinical interpretability therefore requires more than highlighting salient words or spans: it must expose the internal concepts that causally influence predictions and determine whether they reflect medically meaningful evidence (Rudin, 2019). Interpretability for clinical language models has traditionally focused on token-level feature attribution, employing label-wise attention or post-hoc saliency methods such as Integrated Gradients, LIME, and SHAP (Mullenbach et al., 2018; Vu et al., 2020; Kim et al., 2022; Dolk et al., 2022). While these techniques identify salient input spans, they are mechanistically limited: they pinpoint where a model attends without elucidating the underlying concepts or the functional logic governing their processing. Furthermore, the reliability of such attributions is frequently contested; prior work has demonstrated that attention weights can be poorly correlated with model outputs, posing significant risks for high-stakes clinical auditing (Jain and Wallace, 2019). Consequently, there remains a critical gap in moving beyond superficial, input-level correlations toward understanding the internal representations that drive clinical reasoning. Recent work in mechanistic interpretability seeks to reverse-engineer models by decomposing internal neural activations into distinct, human-understandable features (Elhage et al., 2021; Olsson et al., 2022). Specifically, Sparse Autoencoders (SAEs) have proven highly effective in general-domain models by disentangling polysemantic neurons into distinct "monosemantic" features Bricken et al. (2023); Gao et al. (2025). While SAEs improve model transparency, their application in the medical domain remains relatively limited. Prior medical-domain work has used dictionary features for mechanistic explanations and demonstrated feature-based steering (Wu et al., 2024), but integrating such interpretable features as explicit artifact-suppression interventions during task-specific fine-tuning remains underexplored. In this work, we propose CAST (Concept-guided Artifact Suppression Tuning), a fine-tuning framework for clinical classification that transforms mechanistic interpretability from a passive analytical tool into an active steering mechanism. As illustrated in Figure 1, standard fine-tuning maps clinical notes directly to risk predictions through a black-box model, making it difficult to determine whether predictions are based on meaningful clinical evidence or dataset-specific artifacts. In contrast, our framework uses Sparse Autoencoders trained on the large-scale MIMIC datasets (Johnson et al., 2016b; Johnson et al., 2016a; Johnson et al., 2023) to identify internal model features corresponding to human-understandable clinical concepts. We then use these semantic units to guide fine-tuning by suppressing features associated with spurious shortcuts, such as formatting artifacts, while encouraging reliance on clinically meaningful representations. We evaluate our approach on MIMIC-IV ICU discharge notes for 30-day out-of-hospital mortality prediction and show that integrating concept-level control into the training loop yields not only improved predictive performance, but also a transparent feature-level audit trail with supporting clinical concepts and ICD-10 codes. 2 Related Work Feature Extraction via Sparse Autoencoders To optimize the extraction of monosemantic features, several SAE architectures have emerged. Vanilla SAEs isolate monosemantic features by projecting dense activations into a high-dimensional, sparse latent space optimized with an ℓ1 _1 penalty Ng (2011). To overcome the feature shrinkage inherent to ℓ1 _1 regularization, TopK SAEs enforce a hard activation budget by retaining only the K largest values per input Gao et al. (2025). The BatchTopK variant builds on this by applying selection across entire batches to prevent dead neurons and stabilize training Bussmann et al. (2024). Additionally, architectures like Matryoshka SAE explore hierarchical representations by constructing nested dictionaries, dividing the latent space into expanding prefixes to encode high-level abstractions in early dimensions and fine-grained details in later ones Bussmann et al. (2025). Together, these developments have substantially improved the quality, stability, and semantic organization of SAE-derived features, making them increasingly useful for downstream interpretability and intervention. SAE-Guided Control and Adaptation Once high-quality monosemantic features are extracted, the interpretability paradigm increasingly shifts from passive analysis toward active model intervention. At inference time, several frameworks utilize SAE features for causal interventions—systematically amplifying or ablating specific latents to reliably steer model generation (Bayat et al., 2025). This steering capability has also been extended to complex cognitive and in-context mechanisms; for instance, Chen et al. (2026) combine SAEs with activation patching to isolate “reasoning features” and probe Chain-of-Thought (CoT) faithfulness, while Cho and Hockenmaier (2025) use SAE-guided procedures to track information flow and improve in-context learning behavior. Beyond inference-time steering, recent work has begun incorporating SAE-derived features into model adaptation. For example, Casademunt et al. (2025) integrate SAE-based concept ablation into fine-tuning to suppress unintended generalizations such as gender bias, while Self-Regul (Wu et al., 2025) uses sparse autoencoders to regularize LLM-based classification toward interpretable sparse features. However, these methods primarily target general-domain generation, reasoning, behavioral steering, or controllable classification, rather than clinical risk prediction. Our work addresses this gap by adapting SAE-guided learning to clinical prediction, where the goal is not only to steer model behavior but also to suppress clinically meaningless documentation artifacts. Mechanistic Interpretability in Clinical Classification Mechanistic interpretability has recently been explored as a way to expose and manipulate internal representations for text classification. In general-domain settings, SPIN Jiao et al. (2024) identifies and integrates task-relevant internal neurons to obtain more compact and interpretable classifiers. Gallifant et al. (2025) show that features discovered by Sparse Autoencoders can serve as effective classifier representations and transfer across models and modalities. However, these approaches do not use clinically interpreted artifact concepts as explicit suppression targets during task-specific fine-tuning. CAST instead integrates SAE-derived concepts directly into clinical model adaptation, updating representations during training to suppress spurious documentation artifacts. 3 Method Figure 2: Framework for clinical concept extraction and controlled fine-tuning. The pipeline consists of three stages: (1) Concept Extraction via a trained Sparse Autoencoder; (2) Concept Interpretation using an LLM to evaluate top activations; and (3) Fine-tuning with Concept Ablation to steer downstream model predictions. We introduce an end-to-end framework that leverages Sparse Autoencoders to map internal representations to interpretable concepts and utilize them for controlled fine-tuning. The complete pipeline is illustrated in Figure 2. 3.1 Concept Extraction via SAEs Let fθf_θ denote a pretrained Transformer encoder. Given an input clinical note x=(x1,…,xT)x=(x_1,…,x_T), we extract the token representations from layer ℓ : H=fθ(ℓ)(x)∈ℝT×d,ht∈ℝd.H=f_θ^( )(x) ^T× d, h_t ^d. where hth_t is the contextual hidden state of token xtx_t at layer ℓ . To convert these black-box features into an interpretable concept interface, we introduce a sparse concept decomposition module parameterized by an encoder-decoder pair ( Eϕ,DψE_φ,D_ψ ): zt=Eϕ(ht)∈ℝdSAE,h^t=Dψ(zt)∈ℝd.z_t=E_φ (h_t ) ^d_SAE, h_t=D_ψ (z_t ) ^d. The key requirement is that ztz_t is sparse (only a few concept dimensions active per token), while h^t h_t remains a faithful reconstruction of hth_t. We therefore optimize a generic objective: ℒconcept =∑t=1T‖ht−h^t‖22+λΩ(z)L_concept = _t=1^T \|h_t- h_t \|_2^2+λ (z) where Ω(⋅) (·) specifies the sparsity mechanism and can be instantiated in different SAE variants. 3.2 Automated Concept Interpretation To make the learned concepts interpretable for auditing and steering, we employ an LLM-based interpretation pipeline. For each latent j, we extract a context window (e.g., [t−w,t+w][t-w,t+w]) around its top k highest activating tokens zt,jz_t,j from the clinical corpus. Prompted with these empirically grounded contexts, an LLM judge outputs three elements: (i) a concise description of the captured clinical concept, (i) a binary classification of its relevance to the downstream task versus unintended artifacts (e.g., boilerplate), and (i) a set of formal medical keywords suitable for querying standardized clinical databases. To prevent the LLM from hallucinating medical codes, we implement a retrieval-based safeguard: the generated keywords are used to query an ICD-10 database (National Center for Health Statistics, 2024), allowing the LLM to assign verified clinical codes to the latent by selecting exclusively from the retrieved candidate set. 3.3 Fine-Tuning with Concept Steering To operationalize the learned concepts, we integrate the frozen SAE as an active steering intervention at its training layer K. We partition the L-layer transformer into a frozen prefix (layers 11 through K) and a trainable suffix (layers K+1K+1 through L). Given input tokens x, the frozen prefix produces hidden states h(K)(x)h^(K)(x), which the SAE rewrites into steered representations h~(K)(x) h^(K)(x) as defined in Equation (1) below. The trainable suffix then processes the steered states. Because clinical discharge notes routinely exceed the encoder’s context window, each document is partitioned into overlapping chunks. The suffix outputs are reduced to a single chunk embedding by mean pooling over non-special tokens, and chunk embeddings are combined into a document embedding through a learned attention pool. The document embedding is then fed to a linear classification head that produces the task logits. Residual-correction intervention. Let T denote the set of task-relevant concept indices identified by the LLM judge in Section 4.4, and let ¯ T denote the indices labeled as task-irrelevant or artifactual. Rather than replacing h(K)h^(K) with the SAE reconstruction, which would incur reconstruction error on signals the dictionary fails to capture, we adopt a residual correction that preserves the original hidden state and subtracts only the contribution of the suppressed concepts: h~t(K)=ht(K)−∑j∈¯zt,jWdec[j,:], h^(K)_t=h^(K)_t- _j∈ Tz_t,jW_dec[j,:], (1) where zt,jz_t,j is the SAE activation of latent j at token t, and Wdec[j,:]W_dec[j,:] is its decoder direction. The unmodeled portion of ht(K)h^(K)_t, i.e., the SAE residual ht(K)−Dψ(zt)h^(K)_t-D_ψ(z_t), passes through unchanged, so clinical signal the SAE fails to reconstruct is preserved verbatim. Uniform artifact suppression. We subtract every latent the interpretation pipeline labels as artifactual, regardless of its activation magnitude or estimated downstream impact. Two considerations motivate this choice. First, even low-magnitude artifact directions can correlate with institution- or template-specific documentation patterns and shift the decision boundary under deployment distribution shift. Second, the residual form in Equation (1) makes uniform suppression more conservative: only the explicit decoder contributions in ¯ T are removed, while the rest of ht(K)h^(K)_t flows through unchanged. This design deliberately decouples what the model is forbidden to use—any verified artifact, set during training—from what evidence the model is shown to rely on, which is defined as a post-hoc audit signal in Section 3.4. 3.4 Per-Concept Attribution for Auditability The suppression mechanism in Section 3.3 controls what the model is forbidden to use; clinical deployment additionally requires evidence of what the trained model actually used. We therefore equip the framework with a post-hoc attribution module that, for every document, scores each SAE latent by its contribution to the prediction without affecting any trained parameter. Let FωF_ω denote the deployed suffix classifier that maps the steered layer-K representation h~(K)(x) h^(K)(x) to the positive-class logit, s(x)=Fω(h~(K)(x)).s(x)=F_ω\! ( h^(K)(x) ). We use the logit rather than the probability to avoid vanishing gradients for confident predictions. For latent j, we define the signed first-order attribution as Aj(x)=∑t=1Tzt,j(x)⟨∇ts(x),Wdec[j,:]⟩,A_j(x)= _t=1^Tz_t,j(x) _ts(x),W_dec[j,:] , (2) where ∇ts(x)≡∂s(x)/∂h~t(K)(x) _ts(x)≡∂ s(x)/∂ h^(K)_t(x), and zt,j(x)z_t,j(x) is the SAE activation of latent j at token t. Equation (2) is the first-order Taylor approximation of the exact counterfactual logit change when concept j’s decoder direction is removed from the steered state: Aj(x)≈ A_j(x)≈ Fω(h~(K)(x)) F_ω\! ( h^(K)(x) ) (3) −Fω(h~(K)(x)−z⋅,j(x)Wdec[j,:]). -F_ω\! ( h^(K)(x)-z_·,j(x)W_dec[j,:] ). Positive values of Aj(x)A_j(x) indicate concepts that push the prediction toward mortality, whereas negative values indicate concepts that push against it. We rank concepts by |Aj(x)||A_j(x)| for overall importance and by signed Aj(x)A_j(x) when presenting per-prediction evidence. Computing Equation (2) requires one forward pass and one backward pass per batch through FωF_ω, plus a single matrix multiplication against Wdec⊤W_dec . This replaces the |D||D| separate forward passes required by exact counterfactual ablation, where |D||D| is the SAE dictionary size. In our implementation, the full MIMIC-IV test split is processed in roughly six minutes on a single GPU. On 3030 held-out documents, the per-document Spearman rank correlation between Aj(x)A_j(x) and the exact counterfactual effect is ρ=0.976ρ=0.976 with Pearson r=0.935r=0.935. Appendix A.2 provides the closed-form expression for the linear case, the validation protocol, and detailed use cases, including per-prediction evidence trails, global auditing, and cost-efficient expert review. 4 Results and Discussion 4.1 Dataset We use the dataset introduced by Yoon et al. (2025) for 30-day out-of-hospital mortality prediction from ICU discharge notes in MIMIC-IV v2.2 (Johnson et al., 2023), with labels derived by linking encounters to an external death registry. The dataset contains 49,832 admissions/notes from 39,705 patients, with patient-level splits to avoid leakage and exclusions for in-hospital death and hospice disposition. Importantly, the classification target is extremely imbalanced (approximately 1,830 positives vs. 48,002 negatives for the 30-day task), posing a challenging setting for robust evaluation on long-form clinical text. 4.2 Models and Baselines For clinical note encoding, we use two domain-specific pretrained Transformer backbones: ClinicalBERT (Alsentzer et al., 2019) for inputs up to 512 tokens, and Clinical-Longformer (Li et al., 2022) for extended notes up to 4,096 tokens. Both are continually pre-trained on the MIMIC corpus. We compare CAST against standard fine-tuning of the same clinical encoders and two zero-shot general-purpose LLM baselines, GPT-4 (OpenAI et al., 2024) and Llama-3-8B (Grattafiori et al., 2024). We additionally include an input-removal baseline that directly removes identified input-level artifacts before fine-tuning, testing whether simple preprocessing alone is sufficient to mitigate artifact-driven shortcuts. We also include two SAE-based baselines: Self-Regul (Wu et al., 2025), which regularizes classification with SAE-derived sparse features, and SAE-Probe, adapted from Gallifant et al. (2025), which freezes the encoder and SAE, pools SAE latent activations into document-level features, and trains a lightweight classifier on top. Unlike CAST, both baselines use SAE latents as auxiliary representations rather than as explicit artifact-suppression interventions inside the fine-tuning loop. Additional details are provided in Appendix A.7. 4.3 Training setup and hyperparameters SAE pretraining. We train three variants of SAE, TopK, BatchTopK, and Matryoshka, on hidden activations from the frozen encoder over a mixed clinical corpus of 200,000200,000 MIMIC notes; the corpus split is provided in the Appendix A.5. We report TopK and Matryoshka in the main results, while BatchTopK results are provided in the Appendix A.9. All SAEs use a dictionary size of dSAE=8,192d_SAE=8,192 with top-k=64top-k=64 active latents per token. The Matryoshka variant uses nested group sizes 512,2048,8192\512,2048,8192\, totaling 10,75210,752 latents. SAEs are trained with AdamW using a learning rate of 1e−31e-3 for approximately 410410M tokens, and are frozen during all downstream fine-tuning. Downstream training. We fine-tune all concept-guided models and fine-tuning baselines for 55 epochs using AdamW with differential learning rates of 5e−55e-5 for the backbone layers and 1e−31e-3 for the classification head. The effective batch size is 256256. To handle the approximately 27:127:1 class imbalance, we use class-weighted focal loss with γ=2.0γ=2.0 and per-class weights α=[1.0,Nneg/Npos]α=[1.0,\,N_neg/N_pos]. We train each model for five epochs and evaluate the final checkpoint on the held-out test set, reporting F1F_1 both at a fixed decision threshold of 0.50.5 and at a validation-selected threshold τ⋆τ that maximizes F1F_1 on the validation set. Additional optimization details are provided in Appendix A.5. Table 1: Performance on 30-day out-of-hospital mortality prediction from long discharge notes. Zero-shot results are reported from Yoon et al. (2025). CAST denotes our concept-guided fine-tuning method. F1τ⋆_τ uses the decision threshold that maximizes F1 on the validation set. Backbone Layer SAE Protocol Interp. F1 ↑ F1τ⋆_ τ ↑ AUROC ↑ PR-AUC ↑ Brier ↓ NLL ↓ ECE ↓ GPT-4 – – Zero-shot ✗ 0.32370.3237 – – – – – – Llama3-8B – – Zero-shot ✗ 0.19480.1948 – – – – – – ClinicalBERT 11 – Fine-tuning ✗ 0.26020.2602 0.29410.2941 0.83200.8320 0.23120.2312 0.09080.0908 0.33150.3315 0.22460.2246 – Input removal ✗ 0.24700.2470 0.3002 0.85500.8550 0.24500.2450 0.12530.1253 0.42280.4228 0.29470.2947 TopK Self-Regul ✓ 0.16630.1663 0.24150.2415 0.83530.8353 0.20110.2011 0.19440.1944 0.57560.5756 0.39600.3960 SAE-Probe ✓ 0.26320.2632 0.25440.2544 0.82650.8265 0.21750.2175 0.08060.0806 0.30660.3066 0.20780.2078 CAST ✓ 0.2961 0.3000 0.8579 0.2460 0.0879 0.3190 0.2150 Matryoshka Self-Regul ✓ 0.16310.1631 0.29910.2991 0.83900.8390 0.22350.2235 0.19480.1948 0.57640.5764 0.39630.3963 SAE-Probe ✓ 0.24040.2404 0.22270.2227 0.80670.8067 0.17430.1743 0.09670.0967 0.33480.3348 0.21610.2161 CAST ✓ 0.2749 0.2927 0.8382 0.2460 0.0693 0.2713 0.1736 8 – Fine-tuning ✗ 0.25490.2549 0.26120.2612 0.82460.8246 0.20890.2089 0.08170.0817 0.31290.3129 0.21220.2122 – Input removal ✗ 0.22400.2240 0.29010.2901 0.82300.8230 0.18500.1850 0.09110.0911 0.33630.3363 0.22340.2234 TopK Self-Regul ✓ 0.16390.1639 0.24790.2479 0.83510.8351 0.20360.2036 0.19490.1949 0.57750.5775 0.39790.3979 SAE-Probe ✓ 0.23080.2308 0.23450.2345 0.79650.7965 0.16430.1643 0.08990.0899 0.33180.3318 0.22540.2254 CAST ✓ 0.2837 0.2699 0.8428 0.2362 0.0837 0.3168 0.2163 Matryoshka Self-Regul ✓ 0.16640.1664 0.25940.2594 0.83790.8379 0.20590.2059 0.19370.1937 0.57460.5746 0.39580.3958 SAE-Probe ✓ 0.16370.1637 0.16120.1612 0.77090.7709 0.10840.1084 0.10860.1086 0.37570.3757 0.25540.2554 CAST ✓ 0.2869 0.2923 0.8422 0.2355 0.0507 0.2242 0.1326 Clinical- Longformer 11 – Fine-tuning ✗ 0.17590.1759 0.32860.3286 0.86660.8666 0.26480.2648 0.19090.1909 0.56810.5681 0.39080.3908 – Input removal ✗ 0.22500.2250 0.32420.3242 0.85700.8570 0.26500.2650 0.13550.1355 0.43150.4315 0.29680.2968 TopK Self-Regul ✓ 0.17260.1726 0.27710.2771 0.84360.8436 0.23420.2342 0.18600.1860 0.55750.5575 0.38540.3854 SAE-Probe ✓ 0.13330.1333 0.23130.2313 0.81510.8151 0.15530.1553 0.22450.2245 0.63500.6350 0.39970.3997 CAST ✓ 0.2794 0.3412 0.8740 0.2702 0.1217 0.4097 0.2882 Matryoshka Self-Regul ✓ 0.19430.1943 0.27310.2731 0.84650.8465 0.23930.2393 0.17600.1760 0.53590.5359 0.37230.3723 SAE-Probe ✓ 0.14540.1454 0.18070.1807 0.80280.8028 0.16830.1683 0.20140.2014 0.58240.5824 0.38530.3853 CAST ✓ 0.2038 0.2760 0.8641 0.2511 0.1595 0.4975 0.3464 8 – Fine-tuning ✗ 0.20280.2028 0.27870.2787 0.8556 0.24330.2433 0.15060.1506 0.47550.4755 0.32890.3289 – Input removal ✗ 0.30400.3040 0.28430.2843 0.83700.8370 0.23000.2300 0.13040.1304 0.44270.4427 0.31290.3129 TopK Self-Regul ✓ 0.19280.1928 0.27750.2775 0.85020.8502 0.24700.2470 0.17600.1760 0.53650.5365 0.37320.3732 SAE-Probe ✓ 0.17770.1777 0.14260.1426 0.77820.7782 0.13580.1358 0.11600.1160 0.39210.3921 0.26520.2652 CAST ✓ 0.3041 0.3085 0.8523 0.2360 0.0900 0.3396 0.2375 Matryoshka Self-Regul ✓ 0.18550.1855 0.27410.2741 0.85140.8514 0.25060.2506 0.17970.1797 0.54430.5443 0.37790.3779 SAE-Probe ✓ 0.10600.1060 0.18150.1815 0.77860.7786 0.10950.1095 0.27900.2790 0.75660.7566 0.47730.4773 CAST ✓ 0.3233 0.3073 0.8531 0.2524 0.0820 0.3196 0.2224 4.4 Feature Extraction and Interpretation Figure 3: Real clinical note case study: artifact suppression corrects a false positive while preserving clinical evidence. We train sparse autoencoders on frozen hidden activations from layer 88 (mid-encoder) and layer 1111 (near the top of the encoder), and insert the resulting modules back into the corresponding layers for downstream fine-tuning (Section 3.3). To interpret the learned SAE latents, we apply an LLM-based labeling pipeline to the maximally activating clinical contexts of each alive latent. For each latent, gemini-2.5-flash-lite (Gemini Team, Google, 2025) is shown its top-10 activating contexts and asked to produce a concise concept description, assign a semantic category, and determine whether the concept is related to 30-day mortality. For diagnostic concepts, the LLM additionally generates medical search terms that are used to retrieve candidate ICD-10-CM codes; code assignment is restricted to the retrieved candidate set to reduce code hallucination. Full prompts and implementation details are provided in Appendix A.3. We use these interpretations to construct a conservative suppression set ¯ T. Specifically, to improve labeling reliability, we run the interpretation pipeline independently three times for each latent and include a latent in ¯ T only when all three runs classify it as either a formatting or de-identification artifact and as not mortality-related. This strict-consensus criterion is designed to reduce erroneous suppression of clinically meaningful features. Across configurations, Fleiss’ κ (Fleiss, 1971) ranges from 0.7070.707–0.7490.749 for mortality-related labels and 0.7490.749–0.8060.806 for artifact labels, while unanimous agreement ranges from 82.682.6–87.7%87.7\% and 87.987.9–93.3%93.3\%, respectively. Per-configuration agreement statistics are reported in Appendix A.3. At the aggregate level, the interpretation results show that many latents capture mortality-related concepts, while roughly half can be grounded to one or more ICD-10 codes. ClinicalBERT yields a higher clinical-concept fraction than Longformer, and layer-1111 SAEs capture richer clinical semantics than layer-88 SAEs. Per-configuration statistics are reported in Appendix A.4. At the instance level, Figure 3 illustrates how these interpreted features translate into the steering behavior of CAST. In a representative MIMIC-IV case, suppressing strict-consensus artifact latents corrects a confident false-positive prediction while preserving activations corresponding to genuine clinical evidence. 4.5 Performance Analysis Table 1 reports results on 30-day out-of-hospital mortality prediction from long discharge notes. The zero-shot LLM baselines provide two useful reference points: GPT-4 remains a strong closed-model baseline, whereas Llama3-8B performs poorly on this long clinical-text task. CAST instead uses compact clinical encoders and exposes predictions through SAE-derived concepts, offering a more auditable alternative to general-purpose LLMs. Among encoder-based methods, the results highlight the importance of how SAE information is used. Standard fine-tuning provides a strong but opaque discriminative baseline, SAE-Probe uses SAE activations only as frozen post-hoc features, and Self-Regul applies SAE-based regularization. CAST goes beyond these alternatives by using SAE concepts as training-time steering signals, yielding a stronger balance of discrimination, calibration, and concept-level interpretability. The input-removal baseline further tests whether direct preprocessing is sufficient to address the identified artifacts. It is competitive on selected metrics, particularly for ClinicalBERT at layer 1111, confirming that visible input-level artifacts can sometimes be removed effectively. However, its gains are not consistent across backbones and layers. Input removal also does not provide the concept-level suppression and audit trail available through CAST. The pattern across backbones further suggests that CAST’s gains are not tied to a single encoder or SAE variant. For ClinicalBERT, TopK at layer 11 gives the strongest discrimination, while Matryoshka at layer 8 yields the most reliable risk estimates. For Clinical-Longformer, the layer 8 Matryoshka CAST model improves operating-point performance while substantially reducing Brier score, NLL, and ECE. These results suggest that SAE-guided steering can improve clinical risk prediction without reducing the model to a purely post-hoc interpretability pipeline. We further assess CAST against the matched fine-tuning baselines using paired bootstrap tests with 10,000 resamples; full results are reported in Appendix A.8. 4.6 Feature Importance We use the per-concept attribution Aj(x)A_j(x) defined in Section 3.4 to perform a post-hoc analysis of the SAE latents that most influence mortality predictions. For each configuration, we aggregate Aj(x)A_j(x) over the MIMIC-IV test cohort by taking the signed mean over documents in which latent j is active. This produces a ranked set of SAE concepts that most strongly influence the model’s output. Table 2 reports the highest-magnitude signed attributions for ClinicalBERT layer 1111 across the three SAE variants. The positive attributions correspond to clinically plausible mortality-related factors, including palliative care, COPD exacerbation, metastatic disease, and acute kidney failure. In contrast, the strongest negative attributions are associated with recovery or lower-risk indicators, especially mobility and independent ambulation. Notably, mobility-related concepts appear among the top negative features for all three SAE variants in this layer, suggesting that the attribution analysis captures stable clinically meaningful signals rather than idiosyncratic features of a single SAE construction. These findings complement the ablation results in Section 4.7: suppressing mortality-related concepts harms predictive performance, whereas suppressing artifact concepts improves it. Appendix A.11 expands Table 2 with LLM-generated concept descriptions and representative activating contexts, while Appendix A.12 provides concrete examples of the strict-consensus artifact set suppressed during training. SAE Latent ¯ A_j Concept (ICD-10-CM) Push prediction ↑ mortality BatchTopK 1108 +0.18+0.18 palliative care (Z51.5) TopK 6730 +0.15+0.15 COPD exacerbation (J44.1) Matryoshka 101 +0.04+0.04 metastatic disease / secondary malignancy (C79.9) TopK 915 +0.02+0.02 acute kidney failure (N17.9) Push prediction ↓ mortality BatchTopK 5888 −0.29-0.29 ambulate independently TopK 7211 −0.20-0.20 gait/mobility status (R26.89) TopK 6567 −0.19-0.19 independent ambulation BatchTopK 3189 −0.18-0.18 discharge mentions Matryoshka 581 −0.13-0.13 gait/mobility status (R26.89) Table 2: Top signed-attribution SAE concepts for ClinicalBERT layer 1111. Positive values push predictions toward mortality; negative values push them away. 4.7 Ablation and Sensitivity Analysis Figure 4: Ablation study of concept-guided fine-tuning. Blocking artifact-related concepts improves both PR-AUC and F1, while blocking mortality-related concepts degrades performance. Concept ablation. We conduct an ablation study on ClinicalBERT layer 1111 to verify that CAST’s gains come from targeted concept-level intervention rather than simply adding an SAE module. We compare three settings: standard fine-tuning, blocking mortality-related concepts, and blocking artifact-related concepts. To make the comparison controlled, we block the same number of mortality-related and artifact-related latents. As shown in Figure 4, matched-size interventions have opposite effects: blocking mortality-related concepts hurts performance, while blocking artifact-related concepts improves it. This suggests that CAST gains come from targeted semantic steering rather than generic regularization, supporting SAE latents as an actionable interface for auditing and control. Layer and hyperparameter sensitivity. We additionally evaluate an earlier intervention layer (layer 4) for both backbones and vary the SAE dictionary size and TopK sparsity in a representative ClinicalBERT layer-11 configuration. Across these sensitivity analyses, CAST remains competitive over the tested settings. Full results are reported in Appendix A.6. 5 Conclusion We introduced an SAE-guided fine-tuning framework for interpretable and controllable clinical prediction. By turning sparse SAE concepts into training-time steering signals, our method moves beyond post-hoc interpretation and provides a direct interface for suppressing spurious artifacts. On MIMIC-IV 30-day mortality prediction, CAST maintains strong predictive performance while improving calibration and exposing concept-level evidence for model decisions. Overall, our framework offers a practical path toward auditable and reliable clinical NLP systems. Limitations Our evaluation focuses on a single prediction task using MIMIC-I and MIMIC-IV discharge notes. Although these datasets span different time periods and documentation systems, they originate from the same institution. Evaluation on external institutions, additional note types, and other clinical tasks is therefore an important direction for future work. Moreover, the modest absolute F1 scores reflect the difficulty of this highly imbalanced task; we view CAST as a research-stage auditing and steering framework rather than a deployable clinical model. Establishing clinically useful operating points will require prospective evaluation and clinician-in-the-loop assessment of the audit trail. Concept interpretation relies on an LLM judge and ICD-10-CM retrieval. Our strict three-run consensus rule is designed to reduce erroneous suppression, but clinician-annotated validation would provide a stronger assessment of labeling reliability. Because artifact and workflow-related features may also carry valid clinical or demographically correlated information, future work should include expert review of suppression sets and subgroup-specific evaluation before clinical use. Finally, CAST introduces additional offline cost for SAE pretraining and latent interpretation, although the LLM is not used during test-time prediction. Acknowledgment J. Mu and G. Chen’s effort was partially supported by NSF grant DMS-2515263 and by the Patient-Centered Outcomes Research Institute (PCORI) Award ME-2024C1-37433. The statements in this work are solely the responsibility of the authors and do not necessarily represent the views of the Patient-Centered Outcomes Research Institute (PCORI), its Board of Governors, or the Methodology Committee. References Alsentzer et al. (2019) E. Alsentzer, J. Murphy, W. Boag, W. Weng, D. Jindi, T. Naumann, and M. McDermott Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop, A. Rumshisky, K. Roberts, S. Bethard, and T. Naumann (Eds.), Minneapolis, Minnesota, USA, p. 72–78. External Links: Link, Document Cited by: §4.2. Amann et al. (2020) J. Amann, A. Blasimme, E. Vayena, D. Frey, V. I. Madai, and P. Consortium Explainability for artificial intelligence in healthcare: a multidisciplinary perspective. BMC medical informatics and decision making 20 (1), p. 310. Cited by: §1. Bayat et al. (2025) R. Bayat, A. Rahimi-Kalahroudi, M. Pezeshki, S. Chandar, and P. Vincent Steering large language model activations in sparse spaces. In Second Conference on Language Modeling, External Links: Link Cited by: §2. Bricken et al. (2023) T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell, R. Lasenby, Y. Wu, S. Kravec, N. Schiefer, T. Maxwell, N. Joseph, Z. Hatfield-Dodds, A. Tamkin, K. Nguyen, B. McLean, J. E. Burke, T. Hume, S. Carter, T. Henighan, and C. Olah Towards Monosemanticity: decomposing language models with dictionary learning. Note: Transformer Circuits ThreadAccessed: 2026-02-12 External Links: Link Cited by: §1. Bussmann et al. (2024) B. Bussmann, P. Leask, and N. Nanda Batchtopk sparse autoencoders. arXiv preprint arXiv:2412.06410. Cited by: §2. Bussmann et al. (2025) B. Bussmann, N. Nabeshima, A. Karvonen, and N. Nanda Learning multi-level features with matryoshka sparse autoencoders. In Proceedings of the 42nd International Conference on Machine Learning, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267, p. 6077–6101. External Links: Link Cited by: §2. Casademunt et al. (2025) H. Casademunt, C. Juang, A. Karvonen, S. Marks, S. Rajamanoharan, and N. Nanda Steering out-of-distribution generalization with concept ablation fine-tuning. arXiv preprint arXiv:2507.16795. Cited by: §2. Chen et al. (2026) X. Chen, A. Plaat, and N. van Stein How does chain of thought think? mechanistic interpretability of chain-of-thought reasoning with sparse autoencoding. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 30297–30305. Cited by: §2. Cho and Hockenmaier (2025) I. Cho and J. Hockenmaier Toward efficient sparse autoencoder-guided steering for improved in-context learning in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 28961–28973. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2. Dolk et al. (2022) A. Dolk, H. Davidsen, H. Dalianis, and T. Vakili Evaluation of lime and shap in explaining automatic icd-10 classifications of swedish gastrointestinal discharge summaries. In Scandinavian Conference on Health Informatics, p. 166–173. Cited by: §1. Elhage et al. (2021) N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah A mathematical framework for transformer circuits. Transformer Circuits Thread. External Links: Link Cited by: §1. Fleiss (1971) J. Fleiss Measuring nominal scale agreement among many raters. Psychological Bulletin 76, p. 378–382. External Links: Document Cited by: §4.4. Gallifant et al. (2025) J. Gallifant, S. Chen, K. Sasse, H. Aerts, T. Hartvigsen, and D. Bitterman Sparse autoencoder features for classifications and transferability. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 29939–29963. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §A.7, §2, §4.2. Gao et al. (2025) L. Gao, T. Dupre la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu Scaling and evaluating sparse autoencoders. In International Conference on Learning Representations, Y. Yue, A. Garg, N. Peng, F. Sha, and R. Yu (Eds.), Vol. 2025, p. 26721–26754. External Links: Link Cited by: §1, §2. Geirhos et al. (2020) R. Geirhos, J. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann Shortcut learning in deep neural networks. Nature Machine Intelligence 2 (11), p. 665–673. Cited by: §1. Gemini Team, Google (2025) Gemini Team, Google Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. External Links: Link Cited by: §A.3, §4.4. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §4.2. Jain and Wallace (2019) S. Jain and B. C. Wallace Attention is not explanation. In Proceedings of NAACL-HLT, p. 3543–3556. Cited by: §1. Jiao et al. (2024) D. Jiao, Y. Liu, Z. Tang, D. Matter, J. Pfeffer, and A. Anderson SPIN: sparsifying and integrating internal neurons in large language models for text classification. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 4666–4682. External Links: Link, Document Cited by: §2. Johnson et al. (2023) A. Johnson, L. Bulgarelli, T. Pollard, S. Horng, L. A. Celi, and R. Mark MIMIC-IV. PhysioNet. Note: Version 2.2 External Links: Document, Link Cited by: §1, §4.1. Johnson et al. (2016a) A. E. W. Johnson, T. J. Pollard, L. Shen, L. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark MIMIC-I, a freely accessible critical care database. Scientific Data 3, p. 160035. External Links: Document, Link Cited by: §1. Johnson et al. (2016b) A. Johnson, T. Pollard, and R. Mark MIMIC-I Clinical Database. PhysioNet. Note: Version 1.4 External Links: Document, Link Cited by: §1. Kim et al. (2022) J. Kim, A. Sharma, S. Shanbhogue, J. Weiss, and P. Ravikumar AnEMIC: a framework for benchmarking icd coding models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, p. 109–120. Cited by: §1. Li et al. (2022) Y. Li, R. M. Wehbe, F. S. Ahmad, H. Wang, and Y. Luo Clinical-longformer and clinical-bigbird: transformers for long clinical sequences. arXiv preprint arXiv:2201.11838. Cited by: §4.2. Maity and Saikia (2025) S. Maity and M. J. Saikia Large language models in healthcare and medical applications: a review. Bioengineering 12 (6), p. 631. Cited by: §1. Markus et al. (2021) A. F. Markus, J. A. Kors, and P. R. Rijnbeek The role of explainability in creating trustworthy artificial intelligence for health care: a comprehensive survey of the terminology, design choices, and evaluation strategies. Journal of biomedical informatics 113, p. 103655. Cited by: §1. Meng et al. (2024) X. Meng, X. Yan, K. Zhang, D. Liu, X. Cui, Y. Yang, M. Zhang, C. Cao, J. Wang, X. Wang, J. Gao, Y. Wang, J. Ji, Z. Qiu, M. Li, C. Qian, T. Guo, S. Ma, Z. Wang, Z. Guo, Y. Lei, C. Shao, W. Wang, H. Fan, and Y. Tang The application of large language models in medicine: a scoping review. iScience 27 (5), p. 109713. External Links: ISSN 2589-0042, Document, Link Cited by: §1. Mullenbach et al. (2018) J. Mullenbach, S. Wiegreffe, J. Duke, J. Sun, and J. Eisenstein Explainable prediction of medical codes from clinical text. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), M. Walker, H. Ji, and A. Stent (Eds.), New Orleans, Louisiana, p. 1101–1111. External Links: Link, Document Cited by: §1. National Center for Health Statistics (2024) National Center for Health Statistics International classification of diseases, tenth revision, clinical modification (icd-10-cm). Note: https://w.cdc.gov/nchs/icd/icd-10-cm/index.htmlCenters for Disease Control and Prevention Cited by: §3.2. Ng (2011) A. Ng Sparse autoencoder. Note: CS294A Lecture Notes, Stanford Universityhttps://web.stanford.edu/class/cs294a/sparseAutoencoder.pdf Cited by: §2. Olsson et al. (2022) C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895. Cited by: §1. OpenAI et al. (2024) OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Ł. Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Ł. Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. Mély, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. O’Keefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph GPT-4 technical report. External Links: 2303.08774, Link Cited by: §4.2. Rudin (2019) C. Rudin Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature machine intelligence 1 (5), p. 206–215. Cited by: §1. Thirunavukarasu et al. (2023) A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting Large language models in medicine. Nature medicine 29 (8), p. 1930–1940. Cited by: §1. Vu et al. (2020) T. Vu, D. Q. Nguyen, and A. Nguyen A label attention model for icd coding from clinical text. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, C. Bessiere (Ed.), p. 3335–3341. Note: Main track External Links: Document, Link Cited by: §1. Wu et al. (2024) J. Wu, D. Wu, and J. Sun Beyond label attention: transparency in language models for automated medical coding via dictionary learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 8848–8871. External Links: Link, Document Cited by: §1. Wu et al. (2025) X. Wu, W. Yu, X. Zhai, and N. Liu Self-regularization with sparse autoencoders for controllable llm-based classification. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, p. 3250–3260. Cited by: §A.7, §2, §4.2. Yoon et al. (2025) W. Yoon, S. Chen, Y. Gao, Z. Zhao, D. Dligach, D. S. Bitterman, M. Afshar, and T. Miller LCD benchmark: long clinical document benchmark on mortality prediction for language models. Journal of the American Medical Informatics Association 32 (2), p. 285–295. Cited by: §4.1, Table 1. Appendix A Appendix A.1 Prompt Template System Prompt Template You are a clinical NLP expert analyzing features from a Sparse Autoencoder trained on Bio_ClinicalBERT. The model processes clinical discharge notes and predicts 30-day out-of-hospital mortality. Concept Interpretation Prompt Analyze latent feature #latent_id from the SAE. Below are the top 10 tokens that maximally activate this latent, with activation strength and surrounding context from clinical discharge notes. examples_block Based on these activation patterns, respond with ONLY a JSON object (no other text): "concept": "<1-2 sentence description>", "mortality_related": <true or false>, "mortality_reasoning": "<1 sentence explanation>", "search_keywords": ["<k1>", "<k2>", "<k3>", "<k4>", "<k5>"] Guidelines for search_keywords: • Provide exactly 5 formal medical terms suitable for searching an ICD-10-CM database. • Use formal terminology, e.g., “acute myocardial infarction” instead of “heart attack”. • Include both specific condition terms and broader category terms. • If the latent detects formatting or structural patterns, use the most relevant medical context visible in the examples. Figure 5: Prompt template. System prompt and user prompt used for concept interpretation. A.2 Per-Concept Attribution: Linear Limit, Validation, and Use Cases Closed-form linear case. When the suffix and prediction head collapse to a single linear map WgW_g applied to a document-level aggregate z¯(x)∈ℝdSAE z(x) ^d_SAE, Equation (2) reduces to the analytic ablation expression Aj(x)=z¯j(x)(Wdec[j,:]Wg⊤),A_j(x)= z_j(x) (W_dec[j,:]W_g ), which can be evaluated without an additional network forward pass. The gradient-based estimator in Equation (2) generalizes this closed form to the non-linear deployed classifier. Validation protocol. For each of 3030 randomly sampled test documents with predicted probability in [0.05,0.95][0.05,0.95], we computed the exact logit-level counterfactual effect in Equation (3) for the union of (i) the top-100100 most activated latents and (i) the top-100100 most attributive latents in that document, yielding approximately 180180 concepts per document. We then computed the per-document Spearman rank and Pearson correlations between Aj(x)A_j(x) and the exact effect. The resulting mean Spearman correlation is ρ=0.976ρ=0.976, with median 0.9780.978, Q1=0.970Q_1=0.970, and Q3=0.986Q_3=0.986; the mean Pearson correlation is r=0.935r=0.935. The mean relative L1L_1 error, |Aj−Δsj|/|Δsj|¯|A_j- s_j|/ | s_j|, is 0.380.38. This residual is composed of higher-order Taylor terms and does not substantially affect ranking or sign. Use cases. The attribution module serves three roles in our framework, none of which feeds back into training. (i) Per-prediction evidence trail. For each document, the top-ranked active concepts, together with their ICD-10 anchors, form a human-readable explanation that clinicians can review alongside the risk score (Figure 3). (i) Global behavioral audit. Aggregating Aj(x)A_j(x) over the test cohort reveals which clinical concepts dominate the model’s decisions and exposes any residual high-impact latent that the LLM judge may have mislabeled or missed. (i) Cost-efficient interpretation. Because LLM-based labeling is rate-limited, ranking concepts by |Aj(x)||A_j(x)| lets us focus expert review on the few hundred latents that most influence predictions, rather than uniformly interpreting all |D||D| dictionary entries. A.3 LLM Interpretation Details We label each alive SAE latent using gemini-2.5-flash-lite (Gemini Team, Google, 2025) through OpenRouter at temperature 0.30.3. For each latent, the LLM is shown the 1010 maximally activating contexts and asked to summarize the captured concept in one or two sentences, assign one of seven categories (clinical, formatting_artifact, deid_artifact, temporal, demographic, structural_marker, or other), and decide whether the concept is mortality-related. Diagnostic claims are grounded against a retrieval lookup over the ICD-10-CM database to eliminate code hallucination. We execute the interpretation pipeline three independent times per latent. A latent enters the suppression set ¯ T only when all three runs label it as either formatting_artifact or deid_artifact and not mortality-related. Table 3 reports inter-run agreement for the two binary decisions used by CAST. Of the latents labeled as artifacts in at least one run, only 44.344.3–51.2%51.2\% satisfy the strict criterion, reflecting the precision-oriented design. Table 3: LLM-labeling agreement across three independent runs. We report Fleiss’ κ and the fraction receiving a unanimous (33-of-33) label for mortality (Mort.) and artifact (Art.) decisions. Strict/any is the fraction of latents flagged as artifacts by at least one run that enter the strict-consensus suppression set. Backbone SAE / Layer Mort. κ_Mort. Art. κ_Art. 3/3 Mort. 3/3 Art. Strict/any ClinicalBERT BatchTopK / 11 .749 .796 87.7% 93.0% 47.5% Matryoshka / 11 .748 .793 87.2% 92.8% 48.7% TopK / 11 .746 .806 87.5% 93.3% 50.2% BatchTopK / 8 .745 .804 87.0% 93.1% 51.2% Matryoshka / 8 .722 .790 85.2% 91.9% 47.6% TopK / 8 .733 .801 86.3% 92.9% 48.0% Clinical- Longformer BatchTopK / 11 .726 .759 83.7% 89.6% 45.7% Matryoshka / 11 .717 .749 82.6% 87.9% 44.3% TopK / 11 .728 .774 83.7% 90.0% 47.7% BatchTopK / 8 .714 .773 82.9% 89.8% 49.0% Matryoshka / 8 .713 .775 82.6% 89.6% 47.9% TopK / 8 .707 .767 82.7% 89.4% 47.5% A.4 Per-Configuration Interpretation Summary Table 4 reports, for each backbone–SAE-type–layer configuration, the number of LLM-interpreted latents, the fraction grounded to ICD-10 codes, and the fraction classified as mortality-related. Backbone comparison. ClinicalBERT consistently yields a higher clinical-concept fraction than Longformer at the same layer and SAE type. This disparity plausibly stems from their distinct attention mechanisms: Longformer’s local sliding-window attention may limit the synthesis of scattered clinical entities, whereas ClinicalBERT’s full self-attention aggregates medical context into a denser semantic space, yielding more interpretable latents. Layer comparison. SAEs trained on layer 1111 yield richer clinical semantics than those trained on layer 88 across backbone–SAE combinations, reflecting the natural progression of high-level, task-specific representations in deeper Transformer layers. SAE variant comparison. Matryoshka SAEs match or slightly exceed TopK SAEs in clinical-concept fraction within comparable settings. This pattern is consistent with the nested structure of Matryoshka dictionaries, which encourages features to be organized across coarser-to-finer representational groups. The strict-consensus artifact set ¯ T used in Section 3.3 is a much smaller subset of the non-mortality side of Table 4, requiring 33-of-33 agreement across independent LLM runs; see Section 4.4. Table 4: Per-configuration summary of interpreted SAE latents. We report the percentage grounded to ICD-10 codes and the percentage classified as mortality-related across backbone architectures, SAE types, and layer depths. Backbone SAE type / Layer ICD-10 (%) Mortality (%) ClinicalBERT TopK / 8 55.355.3 78.378.3 TopK / 11 57.657.6 79.279.2 BatchTopK / 8 55.555.5 78.478.4 BatchTopK / 11 57.557.5 79.979.9 Matryoshka / 8 55.555.5 77.477.4 Matryoshka / 11 57.757.7 78.578.5 Longformer TopK / 8 50.250.2 72.972.9 TopK / 11 50.950.9 72.772.7 BatchTopK / 8 51.251.2 72.272.2 BatchTopK / 11 51.151.1 72.672.6 Matryoshka / 8 49.749.7 71.771.7 Matryoshka / 11 48.148.1 71.171.1 A.5 Training Details SAE pretraining corpus. We pretrain SAEs on 200,000200,000 MIMIC notes: 120,000120,000 from MIMIC-I narrative records and 80,00080,000 from MIMIC-IV discharge summaries. SAE optimization. Each SAE is trained for 100,000100,000 steps with 4,0964,096 tokens per step, corresponding to approximately 410410M training tokens. We use AdamW with learning rate 1e−31e-3, ℓ1 _1 coefficient 5e−45e-4, and an auxiliary penalty of 1.56e−21.56e-2 for latents inactive for more than 1,0001,000 batches. Downstream fine-tuning. We use AdamW with weight decay 0.010.01 and an effective batch size of 256256, using gradient accumulation when needed. All reported runs use random seed 4242. A.6 Layer and Hyperparameter Sensitivity We extend the layer analysis to layer 44 and conduct controlled sensitivity experiments on ClinicalBERT with a TopK SAE at layer 1111. Unless otherwise varied, the default setting is dSAE=8,192d_SAE=8,192, k=64k=64, learning rate 5×10−55× 10^-5, and five fine-tuning epochs. Earlier-layer intervention. Table 5 complements the layer-88 and layer-1111 results in Table 1. Early-layer behavior is more dependent on the SAE architecture: TopK performs best for ClinicalBERT, whereas Matryoshka performs best for Clinical-Longformer. Table 5: Layer-44 results on MIMIC-IV, complementing the layer-88 and layer-1111 evaluations in Table 1. Backbone Method F1 AUROC PR-AUC ClinicalBERT Fine-tuning .2935 .8415 .2429 TopK CAST .3061 .8511 .2574 Matryoshka CAST .2306 .8538 .2342 Clinical- Longformer Fine-tuning .2746 .8464 .2378 TopK CAST .2432 .8570 .2485 Matryoshka CAST .3025 .8484 .2526 SAE dictionary size and sparsity. Table 6 varies dSAEd_SAE and k. Performance remains stable across the tested settings, suggesting that CAST is not sensitive to a particular dictionary size or sparsity level. Table 6: Dictionary-size and TopK-sparsity sensitivity for ClinicalBERT layer 1111 CAST. d_SAE k F1 AUROC PR-AUC 4,096 64 .2703 .8511 .2531 8,192 32 .2569 .8525 .2530 8,192 64 .2961 .8579 .2460 8,192 128 .2801 .8502 .2344 16,384 64 .2881 .8538 .2504 A.7 Baseline implementations. For fair comparison with the protocols, we adapt both SAE-based baselines to our long-text setting while preserving their characteristic designs. Self-Regul (Wu et al., 2025) is implemented as a linear probe on a frozen document embedding: =−dec⊤(ReLU(enc⊤(−dec))⊙¯),x=h-W_dec (ReLU\! (W_enc (h-b_dec) ) 1_ T ), (4) where h is obtained by attention-masked mean pooling within chunks, followed by mean pooling across chunks, and ¯ T is the same strict-consensus artifact set used by CAST. We retain the original weight-projection penalty, λ‖enc⊤(−dec)[¯]‖1,λ _enc (w-b_dec)[ T] _1, (5) and select λ on the development set. SAE-Probe Gallifant et al. (2025) sums per-token SAE activations across the entire document, excluding special tokens, optionally binarizes at the paper-default threshold τ=1τ=1, and trains a linear classifier on the resulting dSAEd_SAE-dimensional feature vector. Because =∑tF= _tf_t saturates on long discharge notes (∼ 5,000 tokens × k=64k=64 active latents per token), we drop the binarization step for SAE configurations whose summed features exceed 95%95\% density on the training set and keep τ=1τ=1 otherwise. For both baselines and CAST, we use the same single-layer classifier head, class-weighted focal loss (γ=2γ=2, =[1,Nneg/Npos] α=[1,\,N_neg/N_pos]), AdamW optimization, effective batch size 256256, and 55 training epochs. The learning rate is ×10−31\!×\!10^-3 for the classifier head and ×10−55\!×\!10^-5 for any trainable backbone layers. Self-Regul and SAE-Probe keep both the encoder and the SAE frozen; CAST and the standard fine-tuning baseline tune the upper backbone layers under an identical schedule. A.8 Paired Bootstrap Analysis We draw 10,00010,000 resamples from the shared test set and apply the same resample to each TopK or Matryoshka CAST model and its matched fine-tuning baseline. Tests are one-sided in the direction of improvement: Δ=CAST−FineTune>0 =CAST-FineTune>0 for F1, F1τ⋆_τ , AUROC, and PR-AUC, and Δ<0 <0 for Brier score, NLL, and ECE. Table 7 reports the p-value for all metrics in Table 1. Table 7: One-sided paired-bootstrap p-values for TopK and Matryoshka CAST relative to the matched fine-tuning baseline (B=10,000B=10,000). The alternative is an improvement by CAST. Bold entries denote p<.05p<.05. Backbone Layer SAE F1 F1τ⋆_τ AUROC PR-AUC Brier NLL ECE ClinicalBERT 11 TopK .0107.0107 .3389.3389 .0002.0002 .1504.1504 <.001<.001 <.001<.001 <.001<.001 11 Matryoshka .1769.1769 .5458.5458 .2269.2269 .1030.1030 <.001<.001 <.001<.001 <.001<.001 ClinicalBERT 8 TopK .1499.1499 .4766.4766 .0108.0108 .1705.1705 .9982.9982 .9925.9925 1.00001.0000 8 Matryoshka .2950.2950 .1011.1011 .0330.0330 .4500.4500 <.001<.001 <.001<.001 <.001<.001 Clinical- Longformer 11 TopK <.001<.001 .2211.2211 .0086.0086 .2335.2335 <.001<.001 <.001<.001 <.001<.001 11 Matryoshka <.001<.001 .9978.9978 .9983.9983 .9783.9783 <.001<.001 <.001<.001 <.001<.001 Clinical- Longformer 8 TopK <.001<.001 .0815.0815 .6877.6877 .7398.7398 <.001<.001 <.001<.001 <.001<.001 8 Matryoshka <.001<.001 .1095.1095 .6131.6131 .1962.1962 <.001<.001 <.001<.001 <.001<.001 A.9 Additional BatchTopK SAE Results on MIMIC-IV Table 8 reports BatchTopK SAE performance on the MIMIC-IV discharge-note held-out test set across the four backbone–layer configurations, alongside standard fine-tuning, Self-Regul, and SAE-Probe baselines. We omit BatchTopK from the main results table (Table 1) for space, since its overall behavior closely tracks that of the TopK SAE. The trends mirror those in the main table. A.10 Evaluation on MIMIC-I MIMIC-I is a publicly available critical-care database covering ICU admissions at Beth Israel Deaconess Medical Center between 2001 and 2012. We use its discharge notes, paired with the same 30-day out-of-hospital mortality label as MIMIC-IV (Section 4.1), giving n=42,548n=42,548 notes with a positive rate of 4.15%4.15\%. The two corpora share the source institution, broad discharge-note format, and ICD-based outcome coding, but differ in documentation period, note templates, attending-physician populations, and EHR version. We therefore use MIMIC-I as a second naturally heterogeneous evaluation set, rather than as a distribution-shift test. Table 9 summarizes the resulting performance. Across all backbone–layer settings, the best-performing CAST variant achieves higher F1 than the corresponding fine-tuning baseline. The largest absolute gains occur on Clinical-Longformer, reaching up to +0.11+0.11 F1 at layer 8 with the Matryoshka SAE, suggesting that the fine-tuned Longformer model is more sensitive to corpus-specific surface patterns. The calibration metrics show the most consistent improvement: the best-calibrated model in every cell is a CAST variant, with Brier, NLL, and ECE substantially reduced across all four backbone-layer settings. The relative ranking of SAE families also carries over from Table 1: Matryoshka is preferred for ClinicalBERT, whereas TopK and BatchTopK are preferred for Clinical-Longformer. Overall, the calibration and F1 gains of artifact-guided suppression transfer cleanly from MIMIC-IV to MIMIC-I, supporting our claim that CAST recovers more clinically stable mortality signals than its fine-tuned counterparts. Table 8: Performance of the BatchTopK SAE on 30-day out-of-hospital mortality prediction on the MIMIC-IV discharge-note test set (n=7,568n=7,568, positive rate ∼ 3.6%). Models are trained on the MIMIC-IV training split following the protocol of Section 4.1. F1 is computed at a fixed decision threshold of 0.50.5. Backbone Layer SAE Protocol Interp. F1 ↑ AUROC ↑ PR-AUC ↑ Brier ↓ NLL ↓ ECE ↓ ClinicalBERT 11 – Fine-tuning ✗ 0.26020.2602 0.83200.8320 0.23120.2312 0.09080.0908 0.33150.3315 0.22460.2246 BatchTopK Self-Regul ✓ 0.16170.1617 0.83580.8358 0.21090.2109 0.19540.1954 0.57760.5776 0.39690.3969 SAE-Probe ✓ 0.12880.1288 0.72820.7282 0.07860.0786 0.17380.1738 0.51580.5158 0.26700.2670 CAST ✓ 0.27690.2769 0.84280.8428 0.24730.2473 0.08250.0825 0.30340.3034 0.20040.2004 8 – Fine-tuning ✗ 0.25490.2549 0.82460.8246 0.20890.2089 0.08170.0817 0.31290.3129 0.21220.2122 BatchTopK Self-Regul ✓ 0.16760.1676 0.83840.8384 0.21100.2110 0.19300.1930 0.57320.5732 0.39510.3951 SAE-Probe ✓ 0.14190.1419 0.73390.7339 0.08320.0832 0.09840.0984 0.49430.4943 0.10990.1099 CAST ✓ 0.27440.2744 0.84110.8411 0.21880.2188 0.05250.0525 0.22740.2274 0.13440.1344 Clinical- Longformer 11 – Fine-tuning ✗ 0.17590.1759 0.86660.8666 0.26480.2648 0.19090.1909 0.56810.5681 0.39080.3908 BatchTopK Self-Regul ✓ 0.18040.1804 0.84330.8433 0.23630.2363 0.18200.1820 0.54890.5489 0.38020.3802 SAE-Probe ✓ 0.09850.0985 0.66710.6671 0.05850.0585 0.17390.1739 0.53410.5341 0.36830.3683 CAST ✓ 0.19910.1991 0.8649 0.26520.2652 0.16270.1627 0.50130.5013 0.34720.3472 8 – Fine-tuning ✗ 0.20280.2028 0.85560.8556 0.24330.2433 0.15060.1506 0.47550.4755 0.32890.3289 BatchTopK Self-Regul ✓ 0.19050.1905 0.84980.8498 0.24700.2470 0.17630.1763 0.53710.5371 0.37350.3735 SAE-Probe ✓ 0.19810.1981 0.80130.8013 0.14960.1496 0.11600.1160 0.38890.3889 0.26520.2652 CAST ✓ 0.29270.2927 0.8527 0.2351 0.08970.0897 0.33550.3355 0.23360.2336 Table 9: Model Performance on MIMIC-I discharge notes (n=42,548n=42,548, positive rate 4.15%4.15\%). Models are trained on the MIMIC-IV train set and evaluated on MIMIC-I without further tuning. CAST denotes our concept-guided fine-tuning method. Backbone Layer Protocol SAE Interp. F1 ↑ AUROC ↑ PR-AUC ↑ Brier ↓ NLL ↓ ECE ↓ ClinicalBERT 11 Fine-tuning – ✗ 0.29690.2969 0.82600.8260 0.26310.2631 0.10190.1019 0.36280.3628 0.24240.2424 CAST TopK ✓ 0.29330.2933 0.84150.8415 0.27370.2737 0.10540.1054 0.36930.3693 0.24720.2472 CAST BatchTopK ✓ 0.29650.2965 0.83180.8318 0.26680.2668 0.10160.1016 0.35830.3583 0.23690.2369 CAST Matryoshka ✓ 0.31300.3130 0.83170.8317 0.26880.2688 0.08020.0802 0.30660.3066 0.19480.1948 8 Fine-tuning – ✗ 0.29680.2968 0.80330.8033 0.25170.2517 0.09380.0938 0.34880.3488 0.23300.2330 CAST TopK ✓ 0.30370.3037 0.82750.8275 0.26970.2697 0.09570.0957 0.34980.3498 0.23410.2341 CAST BatchTopK ✓ 0.29920.2992 0.81410.8141 0.24490.2449 0.06400.0640 0.26430.2643 0.15640.1564 CAST Matryoshka ✓ 0.32130.3213 0.81370.8137 0.26390.2639 0.06040.0604 0.25580.2558 0.14970.1497 Clinical- Longformer 11 Fine-tuning – ✗ 0.23340.2334 0.84840.8484 0.27860.2786 0.17130.1713 0.52710.5271 0.36170.3617 CAST TopK ✓ 0.31760.3176 0.85320.8532 0.28050.2805 0.12500.1250 0.42370.4237 0.29300.2930 CAST BatchTopK ✓ 0.22800.2280 0.84870.8487 0.28040.2804 0.16900.1690 0.51920.5192 0.35500.3550 CAST Matryoshka ✓ 0.27900.2790 0.82870.8287 0.22200.2220 0.13090.1309 0.43720.4372 0.30040.3004 8 Fine-tuning – ✗ 0.22160.2216 0.85040.8504 0.27710.2771 0.16300.1630 0.50400.5040 0.34230.3423 CAST TopK ✓ 0.32330.3233 0.84270.8427 0.26960.2696 0.09700.0970 0.35720.3572 0.24280.2428 CAST BatchTopK ✓ 0.33050.3305 0.84650.8465 0.27840.2784 0.09460.0946 0.34820.3482 0.23550.2355 CAST Matryoshka ✓ 0.33080.3308 0.84280.8428 0.27430.2743 0.08610.0861 0.33010.3301 0.22150.2215 A.11 Top Clinical Concepts by Attribution Table 10 expands Table 2 with the LLM-generated concept description and a representative maximally activating context for each of the top signed-attribution clinical concepts on ClinicalBERT layer 1111 across the three SAE variants. The activating contexts are sampled directly from MIMIC-IV discharge notes; the symbol / denotes a line break, originally represented as <cr>. Table 10: Top signed-attribution clinical concepts A¯j A_j on ClinicalBERT layer 1111. Concepts are aggregated across the three SAE variants and shown in descending order of |A¯j|| A_j| within each direction. Concept descriptions are produced by the retrieval-grounded LLM judge (Section 4.4); contexts are sampled from the SAE training corpus. SAE L# ¯ A_j ICD-10 Concept (LLM description) Representative activating context Push prediction ↑ mortality TopK 4415 +0.21+0.21 R19.1 Physical exam finding of a “soft” abdomen, often with descriptors of normal bowel sounds. / abd : soft , nt , nd , normal BatchTopK 1108 +0.18+0.18 Z51.5 Palliative-care services, often in the context of serious illness or cancer treatment. second / round of palliative chemo TopK 6730 +0.15+0.15 G70.01, J44.1 Worsening of existing symptoms or emergence of new symptoms. temp > 101.5, worsening pain , / drainage BatchTopK 7197 +0.12+0.12 I27.23, J60–J70 Mentions of the lung, particularly lung cancer, lung masses, and other pulmonary pathology. lung sounds very diminished l lung fie... TopK 5454 +0.13+0.13 F41.1, F06.4 Anxiety, often in the context of psychiatric conditions such as depression. 20mg daily for depression / anxiety TopK 3294 +0.09+0.09 – Improvement of clinical conditions, indicating a positive response to treatment. area on buttocks has improved since admission Push prediction ↓ mortality BatchTopK 5888 −0.29-0.29 – Patient’s ability to ambulate independently, indicating a high level of functional status. activity status : ambulatory - independent BatchTopK 7211 −0.26-0.26 – Mobility status, specifically ambulatory and independent. / activity status : ambulatory - independent TopK 6567 −0.19-0.19 R26.89, M62.3 Independent ambulation status, indicating self-mobility ability. > activity status : ambulatory - independent TopK 1645 −0.12-0.12 Y83.8, Y83 Presence or absence of major surgical or invasive procedures. major surgical or invasive procedure : BatchTopK 7643 −0.11-0.11 R41.82 Presence and integrity of various bodily systems, particularly neurocognitive functions. attention / concentration : grossly intact TopK 3082 −0.07-0.07 R41.82, I69.0 Use of the word “intact” when describing the status of bodily functions. attention / concentration : grossly intact A.12 Examples of Suppressed Artifacts To make the strict-consensus suppression set ¯ T concrete, Table 11 shows representative entries from the strict-consensus artifact latents identified on the ClinicalBERT layer 1111 TopK SAE. Two broad sub-categories dominate: formatting artifacts, including line-break tokens, list-item delimiters, parenthetical notations, and structural section markers , and de-identification artifacts, including placeholder tokens introduced by the MIMIC de-identification pipeline such as [**...**] . These patterns may correlate with patient outcomes at the dataset level, for example through documentation density or note structure, but they do not directly reflect patient physiology. Their decoder contributions are therefore subtracted from h~(K) h^(K) during fine-tuning (Section 3.3). Table 11: Representative entries from the strict-consensus artifact latents on the ClinicalBERT layer 1111 TopK SAE. The 33-of-33 LLM agreement rule (Section 4.4) classifies these latents as documentation-format or de-identification artifacts unrelated to patient state; their decoder contributions are subtracted from h~(K) h^(K) during fine-tuning. Sub-category L# LLM concept description Representative activating context formatting 7488 Line breaks and formatting elements, especially the <cr> token in lists. please take all pills ) / / followup instructions formatting 4310 Structural transition from response to plan sections in clinical notes. cr > response : / plan : / chronic o formatting 1625 Parenthetical confirmation notations, e.g., (x) yes or ( ) no. services / considered ? ( x ) yes - ( ) no formatting 8123 List-item or code-snippet delimiters, especially the semicolon character. care , asking app questions . a ; / loving p ; c formatting 2201 Phrases indicating that specific details or instructions are provided elsewhere. of systems is unchanged from admission except as n de-id 2195 De-identified first- or last-name tokens of the form [**FirstName5(NamePattern1) X**]. needed . [ * * first name5 ( namepattern1 ) 1875 * de-id 7360 De-identified date placeholders of the form [**Y-M-D**]. l / [ * * 2189 - 4 - 8 * * ] de-id 6789 De-identified patient-name or title placeholders, e.g., “Dr. X spoke with …”. [ * * last name ( stitle ) 1430 * * ] spoke A.13 Artifacts and licensing We use MIMIC-I and MIMIC-IV under the PhysioNet credentialed access and data use requirements, and only for retrospective research on clinical NLP. We use publicly available pretrained clinical encoders in a manner consistent with their research use for clinical text modeling, and use ICD-10-CM terminology only for retrieval-based concept grounding. We do not redistribute MIMIC clinical notes, and any released code or trained artifacts will exclude patient text and follow the corresponding dataset and model usage terms. The models and artifacts created in this work are intended only for research and auditing of clinical language models, not for deployment in clinical decision-making without further validation, approval, and compliance review.