Paper deep dive
Anatomy Contextualized Adaption of CT Foundation Models
Roshan Kenia, Stephanie L McNamara, William Lotter
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The attention weights learned by ACA's inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these results support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while preserving and enhancing global anatomical context.
Tags
Links
- Source: https://arxiv.org/abs/2607.27154v1
- Canonical: https://arxiv.org/abs/2607.27154v1
Trouble viewing inline? Open PDF directly →
Full Text
71,850 characters extracted from source content.
Expand or collapse full text
Anatomy Contextualized Adaption of CT Foundation Models Roshan Kenia 1,2 , Stephanie L McNamara 3 , and William Lotter 2,4,5 1 Department of Biomedical Informatics, Harvard Medical School, Boston, MA 2 Department of Data Science, Dana-Farber Cancer Institute, Boston, MA 3 Department of Radiology, Massachusetts General Hospital, Boston, MA 4 Department of Pathology, Brigham and Women’s Hospital, Boston, MA 5 Department of Pathology, Harvard Medical School, Boston, MA Abstract. CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by align- ing anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Fur- thermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foun- dation model representations for anatomy-level vision-language align- ment while enhancing global contextualization. ACA uses TotalSegmen- tator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from ra- diology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of training once embeddings are cached. The atten- tion weights learned by ACA’s inter-anatomy transformer additionally indicate plausible cross-anatomy context routing. Altogether, these re- sults support ACA as a lightweight approach for adapting CT foundation models to anatomically grounded vision-language alignment while pre- serving and enhancing global anatomical context 6 . Keywords: CT foundation models· vision-language pre-training· fine- grained alignment 1 Introduction CT foundation models [32, 34, 35, 40] trained on large-scale data have demon- strated promising performance across a wide range of downstream tasks, includ- ing zero-shot disease classification via vision-language contrastive learning. How- ever, these models are typically trained using whole-volume CT representations, 6 Code is available at https://github.com/lotterlab/ACA arXiv:2607.27154v1 [cs.CV] 29 Jul 2026 2R. Kenia et al. condensing a 3D volume into a single embedding for pre-text training [3,4,13]. Conversely, fine-grained vision-language pre-training (FVLP) has emerged as a strategy to align anatomy-specific visual embeddings with corresponding textual descriptions [5,20,26,39]. These methods can recover fine-grained signal, but in doing so discard the global context that whole-volume models provide. Further- more, existing FVLP approaches have relied on training the entire model from scratch, which is computationally expensive and can be difficult to scale. We introduce Anatomy Contextualized Adaption (ACA), a framework for adapting pretrained CT foundation models to anatomy-level vision-language alignment while enhancing global context. ACA combines the strengths of both whole-volume foundation models and FVLP: the broad, frozen representations of pretrained CT foundation models are adapted using lightweight trainable mod- ules to produce anatomy-level embeddings, which are then contextualized across anatomies using transformer-based attention. The training loss includes both anatomy-level and scan-level vision-language alignment. Across two datasets and two base foundation models, ACA boosts zero-shot diagnostic performance and its learned attention indicates plausible cross-anatomy context routing. 2 Related Work CT Foundation Models. Recent work has explored large-scale pre-training of CT models that can be adapted to a variety of clinical tasks. Vision-language pre- training approaches, building on image-text contrastive learning [25,29,33], align CT volumes with paired radiology reports [3, 6]. CT-CLIP [13] and Merlin [4] represent the dominant paradigm, training contrastively on chest/abdominal CT volumes paired with free-text reports, enabling zero-shot abnormality detection at supervised-level performance. Vision-only approaches such as CT-FM [24] and 3DINO [37] leverage self-supervised contrastive [7,10,15] and masked image modeling [2, 14, 36] objectives on unlabeled CT volumes, demonstrating strong transfer to segmentation and classification tasks. Segmentation-oriented founda- tion models such as VISTA3D [17] and SAM-Med3D [28], together with large- scale anatomical segmentation frameworks such as TotalSegmentator [30], enable generalized anatomical understanding across dozens to hundreds of anatomi- cal structures in 3D CT imaging. Task-specific multimodal models including M3FM [23] and LCTfound [11] incorporate structured clinical data and imag- ing to address lung cancer screening and other downstream workflows. Given the emergent zero-shot capabilities of vision-language approaches, we focus on adapting Merlin and CT-CLIP while leveraging TotalSegmentator to facilitate fine-grained modeling. Fine-grained Vision-language Pre-training. Rather than aligning entire CT volumes with full radiology reports, fine-grained vision-language pretraining (FVLP) aligns individual anatomical regions with their corresponding report descriptions [19,22,26,27,39]. fVLM [26] explicitly decomposes CT volumes us- ing TotalSegmentator [30] and matches anatomy-level visual tokens to anatomy- specific text via cross-attention. ViSD-Boost [5] identifies a semantic density gap Anatomy Contextualized Adaption3 CT Foundation Model (Text Encoder) Raw 3D CT Scan ..... Segmentations Last Feature Map Anatomy Tokens There is an appearance compatible with steatosis in the liver parenchyma entering the section area. Hepatosteatosis Stomach shows no significant abnormalities. ..... Anatomy Extracted Text CT Foundation Model (Vision Encoder) TotalSegmentator Anatomy Embedding Extraction Anatomy Embeddings Inter-Anatomy Transformer Transformer Layers Spatial Position () Anatomy Type () LLM FrozenTrainableProjection Finding: There are diffuse patchy ground-glass densities, atelectatic changes, and prominent bronchovascular structures in both lungs. .... Upper abdominal organs included in the sections are normal. There is an appearance compatible with steatosis in the liver parenchyma entering the section area. Impression: Changes in lung parenchyma consistent with Covid-19 viral pneumonia. Lymph nodes in the mediastinum. Hepatosteatosis, small hiatal hernia. Increase in heart size. Raw Report Text Embedding Extraction Fig. 1: Overview of ACA. Anatomy-level visual embeddings are constructed from a frozen CT foundation model using TotalSegmentator segmentations, while anatomy- level text embeddings are extracted from radiology reports using an LLM and the corresponding frozen text encoder. An inter-anatomy transformer, augmented with spatial position and anatomy type embeddings, contextualizes the anatomy embeddings across the full set of structures. The resulting embeddings are aligned to per-anatomy text via L anatomy , while their mean-pooled scan-level embedding is aligned to the full report via L scan . between low signal-to-noise visual representations and information-dense diag- nostic reports, and addresses it through disease-level visual contrastive learning and anatomical normality modeling. Despite these advances, existing FVLP methods treat each anatomy indepen- dently during both training and inference, ignoring the inter-anatomy context that is often essential for accurate diagnosis [31]. For example, organomegaly (abnormal enlargement of an organ; e.g., hepatomegaly, splenomegaly) is best assessed relative to surrounding structures and body habitus. Our work ad- dresses this gap by introducing an inter-anatomy transformer that contextual- izes anatomy embeddings across all present anatomies before alignment, enabling richer and more clinically grounded representations. CT Foundation Model Datasets. The datasets used to train the Merlin and CT-CLIP foundation models have been publicly released. The Merlin [4] dataset is a large-scale abdominal CT cohort comprising 25,494 CT scans paired with radiology reports from 18,317 unique patients, collected at Stanford Univer- sity Medical Center, with annotations spanning 30 abdominal findings. CT-CLIP was developed using CT-RATE [13], a chest CT dataset comprising 50,188 re- constructed 3D volumes from 25,692 scans of 21,304 unique patients, each paired 4R. Kenia et al. with a radiology report and 18 annotated abnormality labels extracted via an automated text classifier. 3 Methodology: Anatomy Contextualized Adaptation ACA is illustrated in Figure 1 and consists of three core components. First, anatomy-level visual embeddings are extracted from a frozen CT foundation model by pooling its feature maps based on TotalSegmentator segmentations, while anatomy-level text embeddings are extracted from the corresponding radi- ology report using an LLM. Second, an inter-anatomy transformer, augmented with learned spatial position and anatomy type embeddings, contextualizes each anatomy’s embedding using the full set of structures present in the scan. Finally, the contextualized anatomy embeddings are projected into a shared contrastive space and aligned to their corresponding per-anatomy text, while a mean-pooled scan-level embedding is aligned to the full report embedding to provide a comple- mentary global supervision signal. We describe each component in detail below. 3.1 Embedding Construction Anatomy Embedding Extraction. To construct anatomy-specific represen- tations, we first apply TotalSegmentator [30] to segment each CT volume into 44 anatomical structures. The 44 structures are based on condensing TotalSeg- mentator’s original 117 classes into a set of related anatomical groups (e.g., left kidney and right kidney → kidney), as summarized in Appendix Table 3. For each CT volume, we pass the scan through a frozen foundation model backbone to obtain intermediate spatial feature representations. We experiment with two foundation models: Merlin, which uses an I3D ResNet152 [8] encoder, and CT-CLIP, which uses a CT-ViT [12] encoder. Merlin uses a 224× 224× 160 voxel input, and generates a final feature map of size 2048× 10× 7× 7. CT- CLIP uses a 480× 480× 240 voxel input, and produces a feature map of size 512× 24× 24× 24. Both models subsequently pool these features to generate a scan-level representation for downstream tasks, whereas we construct anatomy- level features from them. To obtain anatomy-level representations, we project the TotalSegmentator segmentation masks into the spatial resolution of the extracted feature maps. Specifically, we apply a non-overlapping max-pool over each binary organ mask using a kernel matched to the backbone’s effective patch size (32 × 32 × 16) voxels for Merlin and (40× 40× 20) voxels for CT-CLIP), yielding a discrete patch-presence grid at the feature map resolution. For each of the 44 anatomical structures, we extract all features in which the structure occupies a patch across the last feature map of the backbone. We subsequently mean-pool these features into a single structure embedding. Structures for which TotalSegmentator pro- duces no non-empty segmentation mask in a given scan are considered absent and excluded from that scan’s input sequence. For each anatomy a with N a final layer extracted features v 1 ,...,v N a ∈R d , we denote the anatomy embedding Anatomy Contextualized Adaption5 as ̄v a = 1 N a P N a i=1 v i , where d = 2048 for the Merlin backbone and d = 512 for CT-CLIP. Each embedding is subsequently ℓ 2 -normalized. In addition, we compute inter-anatomy spatial position encodings for each anatomical structure as a 44- dimensional vector, where the g-th entry is the ℓ 2 distance between the struc- ture’s voxel-space centroid and the centroid of structure g, normalized by the volume diagonal length √ D 2 + H 2 + W 2 , where D, H, W correspond to the number of voxels along the depth, height, and width, respectively. These encod- ings capture the geometric relationships between anatomical structures within a volume and are used as positional signals in downstream models. Text Embedding Extraction. For the text modality, anatomy-specific findings are extracted from radiology reports using Qwen3-4B-Instruct [38] (LLM in Figure 1), following a similar strategy to [26]. These findings are subsequently encoded using the corresponding frozen Merlin or CT-CLIP text encoder to produce anatomy-level text embeddings. Both the prompts used for anatomy- specific finding extraction and an example of the resulting anatomy-specific find- ings are shown in Appendix Figures 5 and 6, respectively. A report-level text embedding is also generated by passing the full report through the text encoder. 3.2 Inter-Anatomy Transformer To enable contextual reasoning across anatomical structures within a scan, we pass all present anatomy embeddings jointly through a transformer encoder. Prior to transformer input, each anatomy embedding is projected and combined with two learned tokens. The first consists of an MLP applied to a positional encoding vector p a ∈R 44 , which encodes the normalized distances from anatomy a to all other structures in the taxonomy as described in Section 3.1. The sec- ond token is a learnable anatomy type embedding t a specific to each of the 44 anatomy groups. The final anatomy embedding for transformer input, h a , is computed as follows: h a = ̄v a W vis + MLP(p a ) + t a (1) The sequence h a a∈P , where P denotes the set of anatomical structures present in the scan, is passed through an L-layer transformer encoder with pre- layer normalization, producing a contextualized embedding for each anatomy. 3.3 Combined Anatomy-Level and Global Report Loss We supervise the inter-anatomy transformer with two complementary loss terms operating at different granularities: an anatomy-level contrastive loss, L anatomy , that aligns each anatomy’s contextualized embedding to its corresponding text description, and a scan-level report loss, L scan , that grounds the full anatomy sequence to the global radiology report. To enable the contrastive losses, the anatomy-level image and text embed- dings are first projected to a shared d ′ -dimensional contrastive space. For both 6R. Kenia et al. embedding types, a two-layer projection head is used, consisting of a linear layer, GELU activation, layer normalization, a second linear layer, and ℓ 2 normaliza- tion in sequence. We denote the final anatomy-level image and text embeddings as f img a and f txt a respectively. To enable the scan-level loss, a scan-level image embedding is computed as the ℓ 2 -normalized mean over anatomy embeddings: f img scan = ℓ 2 1 |P| P a∈P f img a . The scan-level text embedding f txt scan is obtained by passing the full radiology report through the frozen text encoder and then pro- jecting to the shared contrastive space using the same text projection head as the per-anatomy text embeddings. Anatomy-Level Loss. For each anatomical structure s, we collect the N s instances present across the batch and compute a symmetric image-text con- trastive loss over them. The standard contrastive loss uses a hard diagonal ob- jective wherein matching image-text pairs are treated as positives and all other combinations are treated as negatives. To handle the prevalence of false negatives in this formulation, where anatomical structures from different patients may be semantically identical yet penalized as different, we follow [26] and construct a soft target matrix T ∈ [0, 1] N s ×N s , where i and j index instances of structure s: T ij = 1[i = j] +1[normal(i)∧normal(j)] +1[patient(i) = patient(j), i̸= j](2) where normal(·) are binary normality labels extracted from radiology reports using Qwen3-4B-Instruct (see Figure 5), and each row is normalized to sum to one. Anatomies where all instances are normal are skipped entirely, focusing capacity on pathological variation. The per-anatomy loss is then: L s =− 1 2N s N s X i=1 N s X j=1 T ij " log e f img i ·f txt j /τ P k e f img i ·f txt k /τ + log e f txt i ·f img j /τ P k e f txt i ·f img k /τ # (3) where f img i and f txt i are the projected and ℓ 2 -normalized visual and text embeddings for instance i, and τ is a learnable temperature initialized at 0.07 and clamped to [0.001, 0.5]. As seen in Equation 4, the total anatomy-level loss, L anatomy , sums over all active structures S, defined as those that appear in at least one scan in the batch and have at least one abnormal instance. Scans with no present anatomies are excluded from all loss terms. L anatomy = X s∈S L s (4) Scan-Level Report Loss. While L anatomy aligns each anatomy indepen- dently to short descriptive text, it provides no signal connecting the anatomy sequence as a whole to the broader clinical context of the scan. To bridge this gap, we include a scan-level contrastive loss, L scan , that aligns the scan-level vi- sual embedding, f img scan , to the text embedding of the full radiology report, f text scan . Unlike the anatomy-level loss, each scan is matched only to its own report, so a hard diagonal objective is used: Anatomy Contextualized Adaption7 Osteopenia Aortic Valve Calcification Metastatic Disease Renal Cyst Hepatomegaly Prostatomegaly Hepatic Steatosis Abdominal Aortic Aneurysm Renal Hypodensities Appendicitis Splenomegaly Bowel Obstruction Pancreatic Atrophy Hydronephrosis Biliary Ductal Dilation Fracture Free Air Ascites Anasarca Submucosal Edema Thrombosis Surgically Absent Gallbladder Gallstones Atelectasis Hiatal Hernia Cardiomegaly Coronary Calcification * Atherosclerosis * Pleural Effusion Lymphadenopathy Peribronchial Thickening Pericardial Effusion Interlobular Septal Thickening Emphysema Mosaic Attenuation Pattern Lung Nodule Bronchiectasis Pulmonary Fibrotic Sequela Medical Material Consolidation Lung Opacity 0.5 0.6 0.7 0.8 0.9 1.0 AUROC Merlin onlyOverlappingCT-RATE only Merlin baseline Merlin ACA CT-RATE baseline (CT-CLIP) CT-RATE ACA Fig. 2: Per-finding AUROC comparison between the original models and the ACA anatomy-guided model for In-Distribution zero-shot evaluation. Green lines represent improvement over the baseline, Red lines represent regression. * These findings’ names from CT-RATE are simplified to match with Merlin: Coronary Artery Wall Calcifica- tion = Coronary Calcification, Arterial Wall Calcification = Atherosclerosis. L scan =− 1 2B B X i=1 " log e f img scan,i ·f txt scan,i /τ scan P B k=1 e f img scan,i ·f txt scan,k /τ scan + log e f txt scan,i ·f img scan,i /τ scan P B k=1 e f txt scan,i ·f img scan,k /τ scan # (5) where B is the number of scans in the batch with at least one present anatomy, and τ scan is a separate learnable temperature initialized at 0.07 and clamped to [0.001, 0.5]. Because f img scan is a mean over the same f img a embeddings optimized by L anatomy , the scan-level signal propagates directly through the vi- sual projection head, jointly shaping a space where individual embeddings are discriminative at the anatomy level and coherent at the scan level. Total Loss. The combined objective is: L total =L anatomy + λL scan (6) We refer to this full model as ACA. We use λ = 1 in our experiments to provide a simple balance of local and global signals, while also performing two ablations: ACA w/o L scan , which uses only L anatomy , and ACA w/o L anatomy , which uses only L scan . 4 Experiments 4.1 Datasets and Training Setup We train on two CT datasets using their respective frozen foundation model backbones. For Merlin [4], we use the creator-defined splits of 15,314 training, 8R. Kenia et al. 5,055 validation, and 5,125 test scans. For CT-RATE [13], we partition the train- ing data into 37,545 training and 9,598 validation scans, and evaluate on the 3,039-scan test set defined by the dataset creators. For each dataset, we train ACA and each baseline independently using AdamW with learning rate 1×10 −4 , batch size 64, and 50 epochs, selecting the best model for evaluation based on validation loss. Full hyperparameters are reported in Appendix Tables 5 and 6. 4.2 Evaluation We evaluate each model on the held-out test split defined by the organizers of each dataset. We focus on zero-shot finding classification, where the findings present in each dataset are detailed in Appendix Tables 7 and 8. Each finding is treated as binary classification (positive or negative) at the scan-level. Un- like Merlin [4], we do not subsample negative samples to match the number of positives and evaluate on all samples for a finding. To evaluate similarity with the model’s scan-level visual embedding, we as- sociate each finding with a set of positive prompts describing the pathology and negative prompts describing normal appearance (see Appendix Figure 7). These are encoded by the frozen foundation model text encoder used by the model, averaged within each class (positive or negative), and projected through the model’s trained text projection head to obtain f + and f − ∈R d ′ . The model’s scan-level score for the finding (defined in Appendix B) is computed as the dif- ference in cosine similarity between the model’s scan-level image representation (f img scan ) and each class embedding (f + and f − ), with all embeddings ℓ 2 normal- ized. In the case where no structures are detected, the scan embedding defaults to a zero vector, which scores 0 against all finding prompts and predicts negative. Further details are provided in Appendix B. We evaluate each model under in-distribution and out-of-distribution set- tings. In the in-distribution setting, models are trained and tested on the same dataset using the corresponding foundation model (e.g., Merlin train → Merlin test). Out-of-distribution tests generalization across datasets and base founda- tion model (e.g., Merlin train → CT-RATE test). As the Merlin and CT-RATE datasets cover different anatomies and finding sets, out-of-distribution evaluation is restricted to the 7 findings shared between the datasets (see Figure 2). 4.3 Baselines We compare to several baselines to highlight the effects of each modeling aspect of ACA. Each baseline uses the same prompting strategy to generate positive and negative text embeddings for each finding, using the corresponding text encoder associated with the model. Global VLMs. We assess the baseline performance of Merlin and CT-CLIP, using the original, fixed scan-level vision embeddings for zero-shot similarity. Fine-grained adaptation. To isolate the effect of fine-grained adaptation on top of frozen foundation model embeddings, we evaluate an MLP baseline that does not include the inter-anatomy transformer in ACA. In this baseline, Anatomy Contextualized Adaption9 each anatomy’s pre-computed mean embedding is passed directly through a two- layer projection head consisting of a linear layer, GELU activation, layer normal- ization, and a second linear layer, with ℓ 2 -normalization applied to the output. No inter-anatomy context is used. Each anatomy is projected independently and the model is trained with the L anatomy loss alone. We additionally compare to two existing fine-grained architectures, using these approaches to generate anatomy-level embeddings on top of the foundation model features. This adaptation enables a compute-matched comparison, but means these baselines do not necessarily reflect the performance of the originally published methods which rely on end-to-end training (see Limitations). The first is the anatomy query pooling mechanism from fVLM [26]. For the Merlin backbone, features from all four I3D ResNet152 layers are concatenated to form 3840-dimensional multi-scale representations per patch. For the CT-CLIP CT- ViT backbone, features from the last layer of the spatial transformer and the last layer of the causal transformer are concatenated to form 1024-dimensional representations. A learned anatomy-specific query vector then attends over these features via cross-attention to produce a single anatomy embedding, which is projected and aligned to the corresponding per-anatomy text using L anatomy . The second architecture is an adaptation of ViSD-Boost [5], which am- plifies disease signals by modeling the normal distribution of each anatomy in latent space via a VQ-VAE. Because our framework operates on frozen em- beddings, we adapt the method to three training stages for fair comparison. First, we perform anatomy-level contrastive alignment using the fVLM query pooling mechanism, with normal–normal anatomy pairs upweighted in the con- trastive targets, rather than a vision-only pre-training stage. Second, we train the Transformer-based VQ-VAE exclusively on normal anatomy instances to learn a shared anatomy-conditioned codebook of healthy representations. Finally, with the VQ-VAE frozen, the original multi-scale anatomy tokens are fused with their VQ-VAE reconstruction through a residual projection, with the hybrid embed- ding aligned to per-anatomy text as other fine-grained baselines via L anatomy rather than binary positive/negative prompts. The disease-level contrastive pre- training stage and momentum encoder from the original method are omitted as they require end-to-end vision encoder training. Global adaptation. To isolate the effect of anatomy-level, fine-grained training in ACA, we define a Spatial Transformer baseline that bypasses anatomy segmentation entirely, but includes a module analogous to the inter- anatomy transformer in ACA. In this baseline, the final feature map of the frozen backbone is depth mean-pooled and reshaped into a sequence of spatial patch features, which are processed by a transformer encoder. The mean of all spatial patch outputs is then projected through a two-layer projection head and used as the scan-level image representation, followed by radiology report alignment using L scan . 10R. Kenia et al. Table 1: AUROC comparison across MERLIN and CT-RATE datasets. † These meth- ods were adapted to operate on frozen embeddings. * Denotes out-of-distribution eval- uation restricted to the 7 shared findings, where models trained on CT-RATE are evaluated on Merlin* and models trained on Merlin are evaluated on CT-RATE*. Per finding results can be found in Appendix Tables 11, 12, 13, and 14. In-distributionOut-of-distribution ModelMerlinCT-RATEMerlin*CT-RATE*Average Global VLM Merlin [4]0.7729±0.0062- -0.6919±0.0058 CT-CLIP [13]-0.7082±0.00370.5822±0.0121- Global Adaptation Spatial Transformer 0.7840±0.0055 0.6467±0.00430.5605±0.0126 0.7003±0.00560.6729 Fine-grained Adaptation MLP0.7981 ±0.0057 0.6129±0.00410.5699±0.0121 0.7389±0.00480.6800 fVLM † [26]0.7699±0.0058 0.6842±0.00400.5587±0.0124 0.7036±0.00490.6791 ViSD-Boost † [5]0.7803±0.0059 0.6511±0.00420.6139±0.0115 0.7413±0.00480.6967 Global + Fine-grained Adaptation ACA 0.8213±0.0052 0.7311±0.00360.6723±0.0117 0.7400±0.00490.7412 Table 2: Ablation results for ACA across the Merlin and CT-RATE datasets, com- paring loss components (L anatomy , L scan ) and anatomy-guided inference-time pooling (Section B). Merlin* and CT-RATE* denote out-of-distribution evaluation. In-distributionOut-of-distribution ModelMerlinCT-RATEMerlin*CT-RATE*Average ACA w/o L anatomy 0.7941±0.0057 0.7117±0.00370.6130±0.0116 0.7143±0.00590.7083 ACA w/o L scan 0.8108±0.0053 0.6935±0.0040 0.6364±0.0120 0.7068±0.00450.7119 + Anatomy-Guided 0.8222±0.0052 0.7365±0.00360.6399±0.0120 0.7451±0.00450.7359 ACA0.8213±0.0052 0.7311±0.00360.6723±0.0117 0.7400±0.00490.7412 + Anatomy-Guided 0.8372±0.0048 0.7413±0.00350.6718±0.0118 0.7530±0.00480.7508 5 Results We evaluate ACA on two CT datasets, Merlin and CT-RATE, using their re- spective frozen foundation models, Merlin and CT-CLIP, as backbones. For ACA and each baseline (Section 4.3), we assess zero-shot classification performance in both in-distribution (i.e., train/test on same dataset) and out-of-distribution (i.e., train/test on different datasets) settings (Section 4.2). We report AUROC as our evaluation metric, as it is threshold-independent and robust to the class imbalance present across both datasets, where most findings have substantially fewer positive than negative cases (Appendix Tables 9 and 10). Table 1 reports macro-average AUROC across findings for all models, datasets, and distribution settings. To quantify uncertainty, we estimate the mean and standard deviation of macro-average AUROC over 1000 bootstrap resamples. Anatomy Contextualized Adaption11 ACA outperforms both global and fine-grained baselines. ACA im- proves over both global VLM foundation models in every setting, raising in- distribution average AUROC from 0.7729 to 0.8213 on Merlin and from 0.7082 to 0.7311 on CT-RATE, with larger gains out-of-distribution (0.5822 to 0.6723 on Merlin*, 0.6919 to 0.7400 on CT-RATE*; note that out-of-distribution eval- uation is restricted to the 7 shared findings, see also Appendix Tables 11, 12). ACA likewise outperforms the fine-grained baselines, with an average boost in AUROC across settings of 0.061, 0.062, and 0.044 compared to MLP, fVLM, and ViSD-Boost, respectively. Gains are observed up to 11.8 points in-distribution (CT-RATE, vs. MLP) and 11.4 points out-of-distribution (Merlin*, vs. fVLM); the one exception is a near-tie with MLP and ViSD-Boost on CT-RATE*, which we revisit below. Furthermore, ACA outperforms the purely global adaptation strategy (Spatial Transformer) in each setting, with an average increase of 0.068 AUROC. Figure 2 illustrates the performance changes per finding compared to the original global VLM models, with finding-level results for all models contained in Appendix Tables 11, 12, 13, and 14. The boosts compared to Merlin and CT-RATE highlight the benefits of structured, anatomy-level modeling, where the findings with the largest performance boosts are often subtle, single or- gan pathologies (e.g., aortic valve calcification, renal cyst, hepatomegaly), that might get washed out in global, scan-level embedding. Conversely, more diffuse or non-anatomy specific findings show less benefits (e.g., free air, submucosal edema, thrombosis). Compared to fine-grained adaptation alone, the benefits of the global inter-anatomy transformer are apparent in findings such as arterial wall calcification (e.g., 0.13 average AUROC boost (Table 12)), that can occur in any artery and thus requires integrating features throughout the scan, as well as organomegaly findings. Ablation of combined loss function. Along with the MLP and Spatial Transformer baselines, which ablate ACA’s global and anatomy-level architec- tural components respectively, we performed ablations of ACA’s loss function itself (Table 2). Removing either the anatomy-level loss (ACA w/o L anatomy ) or the scan-level loss (ACA w/oL scan ) decreases performance (-0.033 and -0.029 av- erage AUROC, respectively), indicating that both anatomy-level and scan-level supervision is important even when using the ACA architecture fixed. Anatomy-guided pooling. As the results thus far have used ACA’s scan- level representation, consisting of mean pooling over the contextualized anatomy embeddings, we also explored an inference-time strategy that assigns different pooling weights depending on the tested finding (see Appendix B). Upweighting the pooling of anatomical structures relevant to each finding (Table 2) further improves AUROC for both ACA and ACA w/oL scan in most settings. For ACA, anatomy-guided pooling raises in-distribution AUROC from 0.8213 to 0.8372 on Merlin and from 0.7311 to 0.7413 on CT-RATE, and out-of-distribution AU- ROC from 0.7400 to 0.7530 on CT-RATE*, while leaving Merlin* essentially unchanged (0.6723 vs. 0.6718). For ACA w/o L scan , the effect is larger, in- distribution AUROC improves from 0.6935 to 0.7365 on CT-RATE, and out- 12R. Kenia et al. Heart Aorta Pulmonary Vein Vena Cava Lung Esophagus Liver Gallbladder Pancreas Stomach Small Bowel Colon Spleen Kidney Adrenal Gland Key organ (attended to) Heart (n=5,094) Aorta (n=5,120) Pulmonary Vein (n=3,070) Vena Cava (n=5,120) Lung (n=5,117) Esophagus (n=5,100) Liver (n=5,121) Gallbladder (n=4,493) Pancreas (n=5,115) Stomach (n=5,117) Small Bowel (n=5,119) Colon (n=5,122) Spleen (n=5,074) Kidney (n=5,119) Adrenal Gland (n=5,118) Query organ (attending) Merlin Heart Aorta Pulmonary Vein Vena Cava Lung Esophagus Liver Gallbladder Pancreas Stomach Small Bowel Colon Spleen Kidney Adrenal Gland Key organ (attended to) Heart (n=3,001) Aorta (n=3,002) Pulmonary Vein (n=2,999) Vena Cava (n=3,002) Lung (n=3,005) Esophagus (n=3,007) Liver (n=3,000) Gallbladder (n=2,644) Pancreas (n=2,975) Stomach (n=2,997) Small Bowel (n=2,905) Colon (n=2,933) Spleen (n=2,990) Kidney (n=2,952) Adrenal Gland (n=2,988) Query organ (attending) CT-RATE 0.05 0.10 0.15 0.20 0.25 mean attention weight 0.025 0.050 0.075 0.100 0.125 0.150 0.175 mean attention weight Fig. 3: Each heatmap shows the mean attention weights between fifteen major anatom- ical structures, averaged over all transformer layers, attention heads, and test scans. Left: ACA model trained on Merlin and evaluated on the Merlin test set. Right: ACA model trained on CT-RATE and evaluated on the CT-RATE validation set. Sample counts per query anatomy reflect each respective dataset. of-distribution AUROC improves on both Merlin* (0.6364 vs. 0.6399) and CT- RATE* (0.7068 vs. 0.7451). These results suggest that restricting attention to clinically relevant anatomy is most useful when the model lacks scan-level su- pervision to otherwise aggregate global context. Learned Anatomy Associations. Using the inter-anatomy transformer, we visualize the attention to understand any inter-anatomy relationships the model implicitly learns. Figure 3 shows the average attention weights between key organs across the test set for ACA models trained on Merlin (left) and CT-RATE (right). We find that both models learn anatomically plausible cross- anatomy associations that emerge purely from contrastive training, and the spe- cific associations each model learns track the anatomical scope of its training dataset. Merlin is an abdominal CT cohort with findings concentrated in abdom- inal and GI pathology (Appendix Table 10), which is reflected in its strongest learned associations. The largest attention weights are amongst the stomach, small bowel, and colon, reflecting their continuity along the GI tract. Addition- ally, the liver and gallbladder attend to one another with the highest weight in each anatomy’s row, mirroring their biliary connection. CT-RATE, by contrast, is a chest CT dataset with findings concentrated in cardiopulmonary pathology (Appendix Table 9), which is reflected in the strong attention weights directed to the lung and esophagus in the CT-RATE ACA model. The spleen also attends most strongly to the kidney and the kidney to the pancreas, structures that sit directly adjacent to one another behind the abdominal cavity. Another notable observation is that self-attention along the diagonal is rel- atively weak in both datasets. We hypothesize that the fine-grained adapta- tion in ACA, built upon frozen foundation models, already provides strong anatomy-specific representations, reducing the need for the transformer to rein- Anatomy Contextualized Adaption13 Fine-grained baselines: Splenomegaly Absent ACA: Splenomegaly Present Ground Truth: Splenomegaly Present Fine-grained baselines: Splenomegaly Absent ACA: Splenomegaly Absent Ground Truth: Splenomegaly Absent Fig. 4: Two CT scans from the Merlin dataset. The right patient has splenomegaly and the left does not. Despite the spleens appearing of comparable size in isolation, ACA correctly identifies the right patient as positive while all fine-grained baselines predict negative. The relative proportions of surrounding organs provide a discriminative signal that single-organ embeddings cannot capture. force anatomy identity through self-attention. Instead, the adaptation module appears to focus on modeling inter-anatomy relationships, consistent with prior work viewing attention as a mechanism for information routing and interaction modeling rather than self-reinforcement [1]. Contextual Anatomy Reasoning. Many findings in the dataset require cross-anatomy context to detect reliably. A clear case study is organomegaly, where the challenge is not detecting a structure’s presence but judging its rel- ative size. For instance, a spleen of a given absolute cross-sectional area may be pathological in one patient and entirely normal in another, depending on body habitus and the proportions of surrounding structures. Figure 4 illustrates this, where two scans in the Merlin test set contain spleens of similar cross- section area, yet only one carries a splenomegaly label. Baselines that embed each organ independently have no mechanism to make this comparison, as each anatomy’s representation is formed without reference to neighboring structures. ACA’s inter-anatomy transformer, by attending jointly over all organ tokens in the same forward pass, can represent the relative size of each structure with respect to the others. The contrastive objective then ties this relational repre- sentation to anatomy-level text that naturally expresses such comparisons (“the spleen is enlarged”), grounding the model in contextual reasoning and facilitating a correct splenomegaly prediction for the example on the right at inference. 6 Discussion ACA adapts frozen CT foundation model representations to provide structured, anatomy-level embeddings while facilitating scan-level, contextual reasoning with a lightweight inter-anatomy transformer. ACA consistently improves zero-shot 14R. Kenia et al. finding classification over global vision-language models and existing fine-grained adaptation methods, in-distribution and out-of-distribution across two large- scale datasets. Ablations confirm that that both core components, anatomical decomposition and scan-level report supervision, contribute independently to this improvement, and that restricting pooling to clinically relevant anatomical structures at inference time can provide an additional low-cost gain. The at- tention weights learned by the inter-anatomy transformer also reflect plausible associations and underlying dataset characteristics. Computational Benefits of Adaptation. ACA is designed to adapt frozen foundation model representations with lightweight trainable modules, avoiding the computational cost of full end-to-end retraining that has traditionally been performed for fine-grained modeling [5, 20, 26]. A natural alternative is to fine- tune the backbone directly, which would significantly increase compute but could further improve performance. To explore this potential, we developed a LoRA- finetuned [18] variant of the full Merlin-based ACA model, trained end-to-end using the same anatomy decomposition and contrastive objective described in Section 3.3. The LoRA variant achieves an average AUROC of 0.8159 on the Merlin test set, compared to 0.8081 for the original ACA model, a gap of just 0.008. This suggests that the lightweight ACA adaptation of frozen foundation model representations recovers the majority of the benefit of full end-to-end training at a fraction of the computational cost, consistent with a broader trend in parameter-efficient adaptation [9,16,21]. A further practical advantage of this design is its modularity: because the adaptation module is decoupled from the backbone, it can be applied to any CT foundation model at low additional cost. Limitations. Our comparisons to fVLM [26] and ViSD-Boost [5] are reimple- mentations adapted to operate on frozen foundation model embeddings, rather than the original end-to-end trained methods. This isolates each method’s align- ment strategy under a matched, frozen-embedding setting, but means our results support a narrower claim than outperforming fVLM and ViSD-Boost as origi- nally published. A full end-to-end retraining of these methods would be needed to compare against their originally reported performance. Additionally, ACA’s anatomy decomposition depends on TotalSegmentator’s vocabulary, so struc- tures outside it are not directly represented. Per-anatomy findings and normality labels used during training are extracted automatically with an LLM rather than verified by radiologists, which may introduce label noise, though all evaluations use each dataset’s ground-truth scan-level labels. Our evaluation is limited to two datasets and backbones, with out-of-distribution comparisons restricted to the 7 findings shared between Merlin and CT-RATE, so broader generalization remains untested. Finally, this work addresses only zero-shot finding classifica- tion, and extending ACA to tasks such as segmentation, outcome prediction, and report generation is left to future work. Conclusion. ACA combines fine-grained alignment with global contextual- ization for CT vision-language modeling by adapting existing foundation models rather than training from scratch. We see this as a practical path toward im- proving representation learning in a compute-efficient manner. Anatomy Contextualized Adaption15 Acknowledgements W.L. gratefully acknowledges funding support from the Ellison Foundation. References 1. Abnar, S., Zuidema, W.: Quantifying attention flow in transformers. In: Proceed- ings of the 58th annual meeting of the association for computational linguistics. p. 4190–4197 (2020) 2. Bao, H., Dong, L., Piao, S., Wei, F.: Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254 (2021) 3. Beeche, C., Kim, J., Tavolinejad, H., Zhao, B., Sharma, R., Duda, J., Gee, J., Dako, F., Verma, A., Morse, C., et al.: A pan-organ vision-language model for generalizable 3d ct representations. medRxiv (2025) 4. Blankemeier, L., Kumar, A., Cohen, J.P., Liu, J., Liu, L., Van Veen, D., Gardezi, S.J.S., Yu, H., Paschali, M., Chen, Z., et al.: Merlin: a computed tomography vision–language foundation model and dataset. Nature p. 1–11 (2026) 5. Cao, W., Zhang, J., Shui, Z., Wang, S., Chen, Z., Li, X., Lu, L., Ye, X., Zhang, Q., Liang, T., et al.: Boosting vision semantic density with anatomy normality mod- eling for medical vision-language pre-training. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 23041–23050 (2025) 6. Cao, W., Zhang, J., Xia, Y., Mok, T.C., Li, Z., Ye, X., Lu, L., Zheng, J., Tang, Y., Zhang, L.: Bootstrapping chest ct image understanding by distilling knowl- edge from x-ray expert models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 11238–11247 (2024) 7. Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 9650–9660 (2021) 8. Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. p. 6299–6308 (2017) 9. Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., Luo, P.: Adaptformer: Adapting vision transformers for scalable visual recognition. Advances in Neural Information Processing Systems 35, 16664–16678 (2022) 10. Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for con- trastive learning of visual representations. In: International conference on machine learning. p. 1597–1607. PmLR (2020) 11. Gao, Z., Zhang, G., Liang, H., Liu, J., Ma, L., Wang, T., Guo, Y., Chen, Y., Yan, Z., Chen, X., et al.: A lung ct vision foundation model facilitating disease diagnosis and medical imaging. Nature Communications (2025) 12. Hamamci, I.E., Er, S., Sekuboyina, A., Simsar, E., Tezcan, A., Simsek, A.G., Esirgun, S.N., Almas, F., Doğan, I., Dasdelen, M.F., et al.: Generatect: Text- conditional generation of 3d chest ct volumes. In: European Conference on Com- puter Vision. p. 126–143. Springer (2024) 13. Hamamci, I.E., Er, S., Wang, C., Almas, F., Simsek, A.G., Esirgun, S.N., Dogan, I., Durugol, O.F., Hou, B., Shit, S., et al.: Generalist foundation models from a multimodal dataset for 3d computed tomography. Nature Biomedical Engineering p. 1–19 (2026) 16R. Kenia et al. 14. He, K., Chen, X., Xie, S., Li, Y., Dollár, P., Girshick, R.: Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 16000–16009 (2022) 15. He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 9729–9738 (2020) 16. He, X., Li, C., Zhang, P., Yang, J., Wang, X.E.: Parameter-efficient model adapta- tion for vision transformers. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 37, p. 817–825 (2023) 17. He, Y., Guo, P., Tang, Y., Myronenko, A., Nath, V., Xu, Z., Yang, D., Zhao, C., Si- mon, B., Belue, M., et al.: Vista3d: A unified segmentation foundation model for 3d medical imaging. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 20863–20873 (2025) 18. Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.: Lora: Low-rank adaptation of large language models. Iclr 1(2), 3 (2022) 19. Huang, S.C., Shen, L., Lungren, M.P., Yeung, S.: Gloria: A multimodal global-local representation learning framework for label-efficient medical image recognition. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 3942–3951 (2021) 20. Lin, J., Xia, Y., Zhang, J., Yan, K., Cao, K., Lu, L., Luo, J., Zhang, L.: Ct-glip: 3d grounded language-image pretraining with ct scans and radiology reports for full-body scenarios. arXiv preprint arXiv:2404.15272 (2024) 21. Liu, Y.C., Ma, C.Y., Tian, J., He, Z., Kira, Z.: Polyhistor: Parameter-efficient multi-task adaptation for dense vision tasks. Advances in neural information pro- cessing systems 35, 36889–36901 (2022) 22. Müller, P., Kaissis, G., Zou, C., Rueckert, D.: Joint learning of localized represen- tations from medical images and reports. In: European conference on computer vision. p. 685–701. Springer (2022) 23. Niu, C., Lyu, Q., Carothers, C.D., Kaviani, P., Tan, J., Yan, P., Kalra, M.K., Whitlow, C.T., Wang, G.: Medical multimodal multitask foundation model for lung cancer screening. Nature Communications 16(1), 1523 (2025) 24. Pai, S., Hadzic, I., Bontempi, D., Bressem, K., Kann, B.H., Fedorov, A., Mak, R.H., Aerts, H.J.: Vision foundation models for computed tomography. arXiv preprint arXiv:2501.09001 (2025) 25. Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning. p. 8748–8763. PmLR (2021) 26. Shui, Z., Zhang, J., Cao, W., Wang, S., Guo, R., Lu, L., Yang, L., Ye, X., Liang, T., Zhang, Q., et al.: Large-scale and fine-grained vision-language pre-training for enhanced ct image understanding. arXiv preprint arXiv:2501.14548 (2025) 27. Wang, F., Zhou, Y., Wang, S., Vardhanabhuti, V., Yu, L.: Multi-granularity cross- modal alignment for generalized medical visual representation learning. Advances in neural information processing systems 35, 33536–33549 (2022) 28. Wang, H., Guo, S., Ye, J., Deng, Z., Cheng, J., Li, T., Chen, J., Su, Y., Huang, Z., Shen, Y., et al.: Sam-med3d: a vision foundation model for general-purpose seg- mentation on volumetric medical images. IEEE Transactions on Neural Networks and Learning Systems (2025) 29. Wang, Z., Wu, Z., Agarwal, D., Sun, J.: Medclip: Contrastive learning from un- paired medical images and text. In: Proceedings of the 2022 Conference on Empir- ical Methods in Natural Language Processing. p. 3876–3887 (2022) Anatomy Contextualized Adaption17 30. Wasserthal, J., Breit, H.C., Meyer, M.T., Pradella, M., Hinck, D., Sauter, A.W., Heye, T., Boll, D.T., Cyriac, J., Yang, S., et al.: Totalsegmentator: robust segmen- tation of 104 anatomic structures in ct images. Radiology: Artificial Intelligence 5(5), e230024 (2023) 31. Wen, J.: Biological age shows that no organ system is an island. Nature 4, 1182– 1183 (2024) 32. Wu, C., Zhang, X., Zhang, Y., Hui, H., Wang, Y., Xie, W.: Towards generalist foun- dation model for radiology by leveraging web-scale 2d&3d medical data. Nature Communications 16(1), 7866 (2025) 33. Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W.: Medklip: Medical knowledge enhanced language-image pre-training for x-ray diagnosis. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 21372–21383 (2023) 34. Wu, L., Zhuang, J., Chen, H.: Voco: A simple-yet-effective volume contrastive learn- ing framework for 3d medical image analysis. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 22873–22882 (2024) 35. Xie, Y., Zhang, J., Xia, Y., Wu, Q.: Unimiss: Universal medical self-supervised learning via breaking dimensionality barrier. In: European Conference on Com- puter Vision. p. 558–575. Springer (2022) 36. Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., Hu, H.: Simmim: A simple framework for masked image modeling. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 9653–9663 (2022) 37. Xu, T., Hosseini, S., Anderson, C., Rinaldi, A., Krishnan, R.G., Martel, A.L., Goubran, M.: A generalizable 3d framework and model for self-supervised learning in medical imaging. npj Digital Medicine 8(1), 639 (2025) 38. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al.: Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025) 39. Zhang, H., Liu, Y., Dai, D., Yang, J., Liu, Q., Xie, Y., Wang, P.: Ca-gcl: Cross- anatomy global-local contrastive learning for robust 3d medical image understand- ing. arXiv preprint arXiv:2605.13544 (2026) 40. Zhou, Z., Sodha, V., Pang, J., Gotway, M.B., Liang, J.: Models genesis. Medical image analysis 67, 101840 (2021) 18R. Kenia et al. A Preprocessing Table 3: Mapping of TotalSegmentator labels to grouped anatomical structures for fine-grained modeling. Idx TotalSegmentator Name Grouping NameGrp 1 spleenspleen1 2 kidney_rightkidney2 3 kidney_leftkidney2 4 gallbladdergallbladder3 5 liverliver4 6 stomachstomach5 7 pancreaspancreas6 8 adrenal_gland_rightadrenal_gland7 9 adrenal_gland_leftadrenal_gland7 10 lung_upper_lobe_leftlung8 11 lung_lower_lobe_leftlung8 12 lung_upper_lobe_rightlung8 13 lung_middle_lobe_rightlung8 14 lung_lower_lobe_rightlung8 15 esophagusesophagus9 16 tracheatrachea10 17 thyroid_glandthyroid_gland11 18 small_bowelsmall_bowel12 19 duodenumsmall_bowel12 20 coloncolon13 21 urinary_bladderurinary_bladder14 22 prostateprostate15 23 kidney_cyst_leftkidney2 24 kidney_cyst_rightkidney2 25 sacrumsacrum16 26 vertebrae_S1sacrum16 27 vertebrae_L5lumbar_vertebrae17 28 vertebrae_L4lumbar_vertebrae17 29 vertebrae_L3lumbar_vertebrae17 30 vertebrae_L2lumbar_vertebrae17 31 vertebrae_L1lumbar_vertebrae17 32 vertebrae_T12thoracic_vertebrae18 33 vertebrae_T11thoracic_vertebrae18 34 vertebrae_T10thoracic_vertebrae18 35 vertebrae_T9thoracic_vertebrae18 36 vertebrae_T8thoracic_vertebrae18 37 vertebrae_T7thoracic_vertebrae18 38 vertebrae_T6thoracic_vertebrae18 39 vertebrae_T5thoracic_vertebrae18 40 vertebrae_T4thoracic_vertebrae18 41 vertebrae_T3thoracic_vertebrae18 42 vertebrae_T2thoracic_vertebrae18 Anatomy Contextualized Adaption19 Idx TotalSegmentator Name Grouping NameGrp 43 vertebrae_T1thoracic_vertebrae18 44 vertebrae_C7cervical_vertebrae19 45 vertebrae_C6cervical_vertebrae19 46 vertebrae_C5cervical_vertebrae19 47 vertebrae_C4cervical_vertebrae19 48 vertebrae_C3cervical_vertebrae19 49 vertebrae_C2cervical_vertebrae19 50 vertebrae_C1cervical_vertebrae19 51 heartheart20 52 aortaaorta21 53 pulmonary_veinpulmonary_vein22 54 brachiocephalic_trunkbrachiocephalic_trunk23 55 subclavian_artery_rightsubclavian_artery24 56 subclavian_artery_leftsubclavian_artery24 57 common_carotid_artery_right common_carotid_artery25 58 common_carotid_artery_left common_carotid_artery25 59 brachiocephalic_vein_leftbrachiocephalic_vein26 60 brachiocephalic_vein_rightbrachiocephalic_vein26 61 atrial_appendage_leftheart20 62 superior_vena_cavavena_cava27 63 inferior_vena_cavavena_cava27 64 portal_vein_and_splenic_vein portal_vein_and_splenic_vein 28 65 iliac_artery_leftiliac_artery29 66 iliac_artery_rightiliac_artery29 67 iliac_vena_leftiliac_vena30 68 iliac_vena_rightiliac_vena30 69 humerus_lefthumerus31 70 humerus_righthumerus31 71 scapula_leftscapula32 72 scapula_rightscapula32 73 clavicula_leftclavicula33 74 clavicula_rightclavicula33 75 femur_leftfemur34 76 femur_rightfemur34 77 hip_lefthip35 78 hip_righthip35 79 spinal_cordspinal_cord36 80 gluteus_maximus_leftgluteus37 81 gluteus_maximus_rightgluteus37 82 gluteus_medius_leftgluteus37 83 gluteus_medius_rightgluteus37 84 gluteus_minimus_leftgluteus37 85 gluteus_minimus_rightgluteus37 86 autochthon_leftautochthon38 87 autochthon_rightautochthon38 88 iliopsoas_leftiliopsoas39 89 iliopsoas_rightiliopsoas39 90 brainbrain40 20R. Kenia et al. Idx TotalSegmentator Name Grouping NameGrp 91 skullskull41 92 rib_left_1rib42 93 rib_left_2rib42 94 rib_left_3rib42 95 rib_left_4rib42 96 rib_left_5rib42 97 rib_left_6rib42 98 rib_left_7rib42 99 rib_left_8rib42 100 rib_left_9rib42 101 rib_left_10rib42 102 rib_left_11rib42 103 rib_left_12rib42 104 rib_right_1rib42 105 rib_right_2rib42 106 rib_right_3rib42 107 rib_right_4rib42 108 rib_right_5rib42 109 rib_right_6rib42 110 rib_right_7rib42 111 rib_right_8rib42 112 rib_right_9rib42 113 rib_right_10rib42 114 rib_right_11rib42 115 rib_right_12rib42 116 sternumsternum43 117 costal_cartilagescostal_cartilages44 B Zero-Shot Scoring and Anatomy-Guided Evaluation For all models, a scan-level score is computed by averaging projected anatomy embed- dings across all present anatomical structures P and taking the difference in cosine similarity to each class: c mean = cos 1 |P| X a∈P f img a , f + ! − cos 1 |P| X a∈P f img a , f − ! (7) A finding is predicted as positive when c mean > 0. For ACA, we additionally evalu- ate an anatomy-guided scoring variant that exploits the per-anatomy structure of the model’s representations. Rather than pooling across all present anatomical structures, the scan embedding for each finding c is computed by restricting to an anatomically relevant subset P c : c guided = cos 1 |P c | X a∈P c f img a , f + ! − cos 1 |P c | X a∈P c f img a , f − ! (8) Anatomy Contextualized Adaption21 The mapping from finding to subset is contained in Table 4. For findings without a well-defined anatomy subset (e.g., lymphadenopathy, free air), P c falls back to all present structures. The final anatomy-guided score is a linear blend with the mean-pool score c mean : c = α· c guided + (1− α)· c mean (9) where α ∈ 0.0, 0.2, 0.4, 0.5, 0.6, 0.8, 1.0 is selected on the validation set by max- imizing macro-average AUROC, and the same value is applied at test time for both ACA (Anatomy-Guided) and ACA w/o L scan (Anatomy-Guided). C Implementation Details All models are trained with the same optimization setup. Table 5 lists the shared training hyperparameters, and Table 6 lists the architecture hyperparameters for each model. Where Merlin and CT-RATE differ due to their respective backbone output dimensions, both values are shown as Merlin / CT-RATE. 22R. Kenia et al. LLM Prompt Templates for Anatomy Information Extraction Mention Detection Prompt section You are a professional radiologist. Please determine if the anatomy (anatomyalias_clause) is mentioned in this CT image report. Please answer directly with "Yes" or "No". Information Extraction Prompt section You are a professional radiologist. Please extract the descriptive information about the specific anatomy (anatomyalias_clause) from this CT image diagnostic report. Please follow these guidelines: 1. Precise extraction: Extract the descriptive information directly related to anatomy from the report. 2. Specify anatomical details: If the report mentions specific areas, parts, or anatomical details of anatomy, make sure to include this information in the description. 3. Concise and clear: Directly extract the report content, avoiding unnecessary explanations or background. 4. Format requirement: Return the information in the format "anatomy: descriptive information" ensuring anatomy is used as the unified prefix. Even if the anatomy has multiple independent parts or lateral characteristics, treat it as a single anatomy and return one comprehensive description. Abnormality Classification Prompt description You are a professional radiologist. Based only on the description above, is the anatomy abnormal? Answer directly with: "Yes" = abnormal finding present "No" = normal or incidental finding only Fig. 5: Prompt templates used for structured anatomy-level information extraction from CT radiology reports. Since reports describe findings globally rather than per or- gan, a three-stage pipeline is applied. The mention detection prompt screens whether a given anatomy is discussed at all, avoiding spurious extractions for absent organs. The information extraction prompt isolates organ-specific descriptive text, using alias hints and formatting constraints to produce a clean, unified description across sub- structures. The abnormality classification prompt labels the result as normal or abnormal, providing the supervision signal for false negative reduction during training. In all templates, section is the raw report text, alias_clause optionally appends clinical synonyms (e.g., also known as splenic for the spleen), and description is the organ-specific text extracted by the preceding stage. Full implementation details are provided in preprocess_reports.py in the given code. Anatomy Contextualized Adaption23 Example Extracted Anatomy-Level Text and Normality Labels Kidney (abnormal) “In both kidneys included in the examination, appearances evaluated in favor of multiple cysts are observed.” Heart (abnormal) “Heart size increased. Calcific atheroma plaques are noted in the coro- nary arteries, associated with the heart’s supply vessels. No pericardial effusion or pericardial thickness increase was observed.” Lung (abnormal) “Minimal bronchiectatic changes and peribronchial thickness increases are observed at the level of the hilum of both lungs. A sequela calcific pulmonary nodule is present in the posterobasal segment of the right lung lower lobe.” Spleen (normal) “Spleen shows no significant abnormalities.” All remaining structures (normal) e.g., “Liver shows no significant abnormalities.”, “Stomach shows no significant abnormalities.”, “Pancreas shows no significant abnormali- ties.”, etc. Fig. 6: Example anatomy-specific findings extracted from a radiology report using Qwen3-4B-Instruct, along with the corresponding binary normality label used in the soft-target contrastive loss described in Section 3.3. Anatomical structures with no reported findings receive a default normal description. 24R. Kenia et al. Example Zero-Shot Finding Prompts Atelectasis – Positive: “atelectasis”, “bibasilar atelectasis”, “subsegmental atelectasis”, “lobar atelectasis” – Negative: “no atelectasis”, “lungs are clear”, “no airspace opacities or atelec- tasis” Cardiomegaly – Positive: “cardiomegaly”, “enlarged heart”, “cardiac silhouette is enlarged”, “moderate cardiomegaly” – Negative: “no cardiomegaly”, “heart is normal in size”, “cardiac silhouette is normal” Hiatal Hernia – Positive: “hiatal hernia”, “sliding hiatal hernia”, “paraesophageal hernia” – Negative: “no hiatal hernia”, “hiatus is normal”, “no herniation through the diaphragmatic hiatus” Fig. 7: Example positive and negative text prompts used for zero-shot finding clas- sification. For each finding, a set of positive prompts describing the pathology and negative prompts describing normal appearance are encoded by the frozen text en- coder and averaged to obtain class-level embeddings f + and f − . Anatomy Contextualized Adaption25 FindingAnatomy Subset AtelectasisLung Pleural EffusionLung EmphysemaLung Lung NoduleLung Lung OpacityLung Pulmonary Fibrotic SequelaLung Mosaic Attenuation PatternLung ConsolidationLung BronchiectasisLung Interlobular Septal ThickeningLung Peribronchial ThickeningLung, Trachea CardiomegalyHeart Pericardial EffusionHeart Coronary CalcificationHeart Arterial Wall CalcificationAorta, Iliac, Subclavian, & Common Carotid Arteries Coronary Artery Wall Calcification Heart, Aorta Aortic Valve CalcificationHeart, Aorta Abdominal Aortic AneurysmAorta AtherosclerosisAorta, Iliac Artery Hiatal HerniaStomach, Esophagus HepatomegalyLiver Hepatic SteatosisLiver Biliary Ductal DilationLiver, Gallbladder GallstonesGallbladder Surgically Absent GallbladderGallbladder SplenomegalySpleen Renal CystKidney Renal HypodensitiesKidney HydronephrosisKidney Pancreatic AtrophyPancreas ProstatomegalyProstate Submucosal EdemaSmall Bowel, Colon Bowel ObstructionSmall Bowel, Colon AppendicitisColon ThrombosisPortal/Splenic Vein, Vena Cava, Iliac Vein Metastatic DiseaseLiver, Lung, Rib, Lumbar Vertebrae, Thoracic Vertebrae OsteopeniaLumbar Vertebrae, Thoracic Vertebrae, Rib FractureRib, Lumbar Vertebrae, Thoracic Vertebrae AscitesLiver, Spleen AnasarcaAll present structures LymphadenopathyAll present structures Free AirAll present structures Medical MaterialAll present structures Table 4: Anatomy subsets used for anatomy-guided zero-shot scoring per finding. Findings evaluated on both Merlin and CT-RATE datasets use the respective dataset’s assignment. 26R. Kenia et al. HyperparameterValue Batch size64 Epochs50 Early stopping patience 5 Learning rate1× 10 −4 Weight decay0.05 Warmup fraction0.05 Gradient clip1.0 Projection hidden dim512 Projection output dim256 Temperature init (τ)0.07 Table 5: Training hyperparameters shared across all models and both datasets. ModelHyperparameterValue MLPVisual input dim2048 / 512 fVLM Multi-scale input dim3840 / 1024 Cross-attention heads4 Cross-attention dropout0.1 Spatial Transformer Spatial features49 / 144 Token input dim2048 / 512 Transformer hidden dim1024 Transformer layers2 Transformer heads4 Dropout0.1 ACA w/o L scan , ACA w/o L anatomy ACA Visual input dim2048 / 512 Transformer hidden dim1024 Positional encoding dim44 Transformer layers2 Transformer heads4 Dropout0.1 ViSD-Boost Multi-scale input dim3840 / 1024 Cross-attention heads4 Cross-attention dropout0.1 VQ-VAE codebook size512 VQ-VAE embedding dim 512 VQ-VAE Transformer layers 1 VQ-VAE Transformer heads 8 VQ commitment cost0.25 Table 6: Architecture hyperparameters per model. Where values differ between the Merlin and CT-RATE backbones, both are shown as Merlin / CT-RATE. Anatomy Contextualized Adaption27 Table 7: Merlin: positive finding counts per split (unlabeled excluded). FindingTrain Val Test Atelectasis2665 902 903 Surgically Absent Gallbladder 2515 805 843 Atherosclerosis2493 778 879 Pleural Effusion2484 794 834 Renal Cyst1650 544 546 Ascites1048 366 365 Anasarca1020 318 352 Hiatal Hernia886 280 312 Hepatic Steatosis793 250 272 Gallstones478 148 148 Fracture470 175 177 Pancreatic Atrophy453 118 161 Osteopenia376 119 147 Submucosal Edema359 146 144 Cardiomegaly352 123 119 Splenomegaly324 109 141 Prostatomegaly306 115 81 Renal Hypodensities302 122 122 Hydronephrosis266 72 112 Thrombosis247 80 77 Bowel Obstruction211 64 72 Aortic Valve Calcification206 61 72 Coronary Calcification166 57 63 Hepatomegaly165 44 57 Biliary Ductal Dilation110 33 52 Appendicitis106 32 43 Lymphadenopathy104 28 49 Free Air100 37 45 Metastatic Disease83 24 46 Abdominal Aortic Aneurysm71 16 17 28R. Kenia et al. Table 8: CT-RATE: positive finding counts per split. FindingTrain Val Test Lung nodule17032 4349 1361 Lung opacity13845 3575 1184 Arterial wall calcification10697 2679 867 Pulmonary fibrotic sequela10127 2461 831 Atelectasis9843 2418 713 Lymphadenopathy9684 2534 789 Coronary artery wall calcification 9634 2390 765 Emphysema7291 1831 600 Consolidation6617 1701 581 Hiatal hernia5414 1337 417 Medical material4668 1150 313 Pleural effusion4582 1122 376 Cardiomegaly4259 1049 325 Peribronchial thickening3960 1013 355 Bronchiectasis3825 907 330 Interlobular septal thickening2976 769 249 Mosaic attenuation pattern2871 767 253 Pericardial effusion2712 700 226 Anatomy Contextualized Adaption29 FindingPositive Negative CT-RATE In-Distribution Medical Material3132726 Arterial Wall Calcification8672172 Cardiomegaly3252714 Pericardial Effusion2262813 Coronary Artery Wall Calcification7652274 Hiatal Hernia4172622 Lymphadenopathy7892250 Emphysema6002439 Atelectasis7132326 Lung Nodule13611678 Lung Opacity11841855 Pulmonary Fibrotic Sequela8312208 Pleural Effusion3762663 Mosaic Attenuation Pattern2532786 Peribronchial Thickening3552684 Consolidation5812458 Bronchiectasis3302709 Interlobular Septal Thickening2492790 CT-RATE Out-of-Distribution Atelectasis7132326 Cardiomegaly3252714 Hiatal Hernia4172622 Lymphadenopathy7892250 Pleural Effusion3762663 Arterial Wall Calcification8672172 Coronary Artery Wall Calcification7652274 Table 9: Number of positive and sampled negative cases for each finding evaluated in the in-distribution and out-of-distribution settings on the CT-RATE dataset. 30R. Kenia et al. FindingPositive Negative Merlin In-Distribution Submucosal Edema144162 Renal Hypodensities1221355 Aortic Valve Calcification7269 Coronary Calcification63372 Thrombosis7732 Metastatic Disease46123 Pancreatic Atrophy1613197 Renal Cyst5461412 Osteopenia1471069 Surgically Absent Gallbladder8432186 Atelectasis903990 Abdominal Aortic Aneurysm1779 Anasarca3521992 Hiatal Hernia3121847 Lymphadenopathy49205 Prostatomegaly811213 Biliary Ductal Dilation52211 Cardiomegaly119291 Splenomegaly1413225 Hepatomegaly571150 Atherosclerosis87955 Ascites365183 Pleural Effusion834556 Hepatic Steatosis27227 Appendicitis4352 Gallstones148248 Hydronephrosis1122079 Bowel Obstruction721345 Free Air45929 Fracture177196 Merlin Out-of-Distribution Coronary Calcification63372 Atelectasis903990 Hiatal Hernia3121847 Lymphadenopathy49205 Cardiomegaly119291 Atherosclerosis87955 Pleural Effusion834556 Table 10: Number of positive and sampled negative cases for each finding evaluated in the in-distribution and out-of-distribution settings on the Merlin datasets. Anatomy Contextualized Adaption31 Table 11: Per-finding AUROC on the Merlin in-distribution test set. Macro average is computed over all 30 findings. Overlapping Macro average restricts the macro-average to the 7 findings shared with CT-RATE, matching the finding set used for out-of- distribution evaluation (Table 13). FindingMerlin Spatial Transformer MLP fVLM ViSD-Boost ACA Submucosal edema0.74150.73090.7220 0.6832 0.6923 0.7185 Renal hypodensities0.67700.67310.7218 0.7081 0.7405 0.7512 Aortic valve calcification0.79970.81330.9506 0.9578 0.9557 0.9637 Coronary calcification0.80570.81980.8392 0.8442 0.8469 0.8496 Thrombosis0.62690.57590.5481 0.6041 0.6057 0.6175 Metastatic disease0.71860.73630.8305 0.8801 0.8643 0.8719 Pancreatic atrophy0.71800.73940.7744 0.7861 0.7706 0.7653 Renal cyst0.62040.62810.7436 0.6704 0.6754 0.7795 Osteopenia0.74350.79320.9199 0.9146 0.8990 0.9279 Surgically absent gallbladder 0.97410.96880.9214 0.7125 0.8299 0.9794 Atelectasis0.67570.72560.7800 0.6530 0.7567 0.8189 Abdominal aortic aneurysm 0.75010.86790.7791 0.7732 0.7492 0.7826 Anasarca0.92790.91400.9002 0.9303 0.9352 0.9378 Hiatal hernia0.63340.67970.7721 0.7625 0.7295 0.7604 Lymphadenopathy0.78700.72910.6910 0.6541 0.7131 0.7469 Prostatomegaly0.67650.73240.7535 0.7425 0.7315 0.7658 Biliary ductal dilation0.79080.81890.7831 0.7898 0.8044 0.8181 Cardiomegaly0.82470.85450.8801 0.8458 0.8663 0.8925 Splenomegaly0.90120.88360.9120 0.9058 0.9197 0.9267 Hepatomegaly0.76450.74860.7799 0.8221 0.8318 0.8672 Atherosclerosis0.96210.96020.8683 0.8382 0.8808 0.8899 Ascites0.90220.91140.8711 0.8351 0.8532 0.9310 Pleural effusion0.92940.92460.8691 0.8525 0.9204 0.9455 Hepatic steatosis0.67470.76280.7698 0.7809 0.7192 0.6580 Appendicitis0.71850.71230.7333 0.5840 0.5012 0.7344 Gallstones0.75660.72210.6920 0.6212 0.7251 0.7374 Hydronephrosis0.72950.73040.7313 0.7611 0.7499 0.7655 Bowel obstruction0.86680.88880.8868 0.8489 0.7493 0.8791 Free air0.76480.78020.8249 0.7487 0.7294 0.7961 Fracture0.72460.69470.6941 0.5871 0.6622 0.7615 Overlapping Macro average 0.80260.81340.8143 0.7786 0.8163 0.8434 Macro average0.77290.78400.7981 0.7699 0.7803 0.8213 32R. Kenia et al. Table 12: Per-finding AUROC on the CT-RATE in-distribution test set. Macro aver- age is computed over all 18 findings. Overlapping Macro average restricts the macro- average to the 7 findings shared with Merlin, matching the finding set used for out-of- distribution evaluation (Table 14). FindingCT-CLIP Spatial Transformer MLP fVLM ViSD-Boost ACA Medical material0.64620.62080.6252 0.6866 0.6784 0.6798 Arterial wall calcification0.84920.78920.6513 0.8135 0.7810 0.8768 Cardiomegaly0.86860.77450.6677 0.7815 0.8420 0.8973 Pericardial effusion0.70530.65430.6337 0.7394 0.7692 0.7787 Coronary artery wall calcification 0.83770.76550.7052 0.7953 0.7748 0.8505 Hiatal hernia0.71870.68430.6466 0.6532 0.6605 0.7956 Lymphadenopathy0.68070.62260.5633 0.6444 0.5585 0.6543 Emphysema0.70650.62560.5965 0.6739 0.6715 0.7506 Atelectasis0.64490.62200.6058 0.6567 0.6098 0.6679 Lung nodule0.53810.47010.5473 0.5283 0.4869 0.5918 Lung opacity0.67430.59900.5440 0.6095 0.5814 0.6078 Pulmonary fibrotic sequela0.57130.51970.5351 0.5651 0.5246 0.6230 Pleural effusion0.89520.83060.6624 0.8485 0.7648 0.8744 Mosaic attenuation pattern0.74700.71790.6584 0.7176 0.6818 0.7519 Peribronchial thickening0.60790.54120.6506 0.6678 0.5395 0.6735 Consolidation0.69920.62490.5302 0.6291 0.6035 0.6442 Bronchiectasis0.65600.56750.5815 0.5871 0.5949 0.7182 Interlobular septal thickening0.70050.61000.6279 0.7180 0.5973 0.7245 Overlapping Macro average0.78500.72700.6432 0.7419 0.7131 0.8024 Macro average0.70820.64670.6129 0.6842 0.6511 0.7311 Table 13: Per-finding AUROC on the Merlin out-of-distribution test set. FindingCT-CLIP Spatial Transformer MLP fVLM ViSD-Boost ACA Atelectasis0.57840.59030.5871 0.5763 0.7383 0.6939 Cardiomegaly0.64430.56040.5535 0.5456 0.5789 0.6598 Hiatal hernia0.59950.63980.6168 0.5795 0.6553 0.6822 Lymphadenopathy0.44930.47500.4812 0.5282 0.5067 0.5687 Pleural effusion0.58380.64090.6951 0.6421 0.6765 0.7691 Atherosclerosis0.56620.49540.5579 0.5624 0.6448 0.6438 Coronary calcification 0.65410.52200.4974 0.4770 0.4968 0.6886 Macro average0.58220.56050.5699 0.5587 0.6139 0.6723 Table 14: Per-finding AUROC on the CT-RATE out-of-distribution test set. FindingMerlin Spatial Transformer MLP fVLM ViSD-Boost ACA Atelectasis0.57530.58560.5845 0.4770 0.5951 0.6050 Cardiomegaly0.78750.80440.8519 0.8265 0.8583 0.8448 Hiatal hernia0.61200.61920.6711 0.6766 0.6487 0.6565 Lymphadenopathy0.58730.59270.5350 0.5166 0.5940 0.6119 Pleural effusion0.90020.91180.8462 0.8062 0.9024 0.8789 Arterial wall calcification0.68530.70140.8409 0.8042 0.7880 0.7853 Coronary artery wall calcification 0.69590.68710.8424 0.8178 0.8024 0.7979 Macro average0.69190.70030.7389 0.7036 0.7413 0.7400