Paper deep dive
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
Renjie Liang, Zijian Xu, Jinqian Pan, Chengkun Sun, Zhengkang Fan, Shawn Li, You Qin, Mei Liu, Jie Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/4/2026, 4:16:08 AM
Summary
The paper introduces ORCA (ORgan-Centroid Aggregation), a training-free, plug-and-play token compression method for 3D CT scans in vision-language models. ORCA merges adjacent visual tokens based on visual similarity and organ guidance, preserving anatomical integrity better than grid averaging. It adds sinusoidal centroid encoding to maintain spatial layout. Evaluated on CT-RATE and Merlin datasets with five encoders, ORCA consistently outperforms baselines in attribute prediction and text generation tasks, achieving significant reductions in token count (64x) and KV-cache (50x) while improving processing speed.
Entities (9)
Relation Signals (7)
ORCA → evaluatedon → CT-RATE
confidence 95% · We evaluate it across two datasets (CT-RATE and Merlin)
ORCA → evaluatedon → Merlin
confidence 95% · We evaluate it across two datasets (CT-RATE and Merlin)
ORCA → outperforms → Grid average
confidence 95% · At matched token budgets, ORCA improves consistently over existing compression methods.
ORCA → uses → Ward linkage
confidence 90% · ORCA uses Ward linkage (Ward Jr. 1963) as its merge criterion
ORCA → uses → Sinusoidal encoding
confidence 90% · adds a sinusoidal encoding of each region's centroid to preserve spatial layout
COLIPRI → isencoderfor → CT-RATE
confidence 85% · CT-RATE COLIPRI (Wald et al. 2025) language–image
ORCA → uses → TotalSegmentator
confidence 85% · In our experiments, these masks are produced by TotalSegmentator
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied in general vision, but little of it targets 3D CT specifically. A common baseline is grid average, which pools regular grid cells and can blend distinct anatomy, lesion, and air into one token. We present \textbf{ORCA} (ORgan-Centroid Aggregation), a token compressor for 3D CT. It merges adjacent tokens with organ guidance and adds a sinusoidal encoding of each region's centroid to preserve spatial layout. This preserves the anatomical information a downstream model needs. ORCA is training-free and plug-and-play, producing an adjustable token set without any model change or text query. We evaluate it across two datasets (CT-RATE and Merlin) and five encoders. The evaluation spans two task types: attribute prediction over five families (size, density, location, texture, and disease) and text generation (visual question answering and report generation). At matched token budgets, ORCA improves consistently over existing compression methods. It shrinks the visual context $64\times$ and its KV-cache $50\times$, and is $31\times$ faster to process each volume. Code released at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.00345v1
- Canonical: https://arxiv.org/abs/2608.00345v1
Trouble viewing inline? Open PDF directly →
Full Text
83,627 characters extracted from source content.
Expand or collapse full text
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression Renjie Liang1, Zijian Xu1, Jinqian Pan1, Chengkun Sun1, Zhengkang Fan1, Shawn Li2, You Qin3, Mei Liu1, Jie Xu1 Abstract A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to tens of thousands per volume, and this sequence must be compressed before a language model can consume it. Token compression is well studied in general vision, but little of it targets 3D CT specifically. A common baseline is grid average, which pools regular grid cells and can blend distinct anatomy, lesion, and air into one token. We present ORCA (ORgan-Centroid Aggregation), a token compressor for 3D CT. It merges adjacent tokens with organ guidance and adds a sinusoidal encoding of each region’s centroid to preserve spatial layout. This preserves the anatomical information a downstream model needs. ORCA is training-free and plug-and-play, producing an adjustable token set without any model change or text query. We evaluate it across two datasets (CT-RATE and Merlin) and five encoders. The evaluation spans two task types: attribute prediction over five families (size, density, location, texture, and disease) and text generation (visual question answering and report generation). At matched token budgets, ORCA improves consistently over existing compression methods. It shrinks the visual context 64×64× and its KV-cache 50×50×, and is 31×31× faster to process each volume. Code released at https://github.com/renjie-liang/ORCA-3DCT. Keywords: Token compression, 3D CT, vision-language models, training-free, medical report generation Introduction 3D computed tomography (CT) is used at national clinical scale and is becoming a practical input to medical vision–language systems. A CT-volume analysis spanning 2,398 U.S. radiology practices shows how broadly CT is used in routine care (Davenport et al. 2021); a recent workload study of 46.4 million imaging examinations from 167 facilities found that the highest-volume radiologists read 30.6% more examinations per day and worked 19.7% more clinical days per quarter by 2024 (Zamani et al. 2026). At the same time, richer 3D CT datasets and models now support abnormality detection (Hamamci et al. 2026; Blankemeier et al. 2026), visual question answering (Wu et al. 2025), and report generation from volumetric studies (Hamamci et al. 2024). A common 3D CT vision–language pipeline maps a volume into a grid of visual tokens, then feeds them through a projector to an LLM or another downstream model. Because a CT scan contains dense anatomical information across many slices, this grid is often large. For example, COLIPRI (Wald et al. 2025) emits a 24324^3 grid, or 13,824 visual tokens. These tokens must be compressed before downstream use. Compression therefore changes more than sequence length: it determines which anatomical evidence remains available to the downstream model. Existing methods for visual token compression address this bottleneck in three ways (Shao et al. 2025). The Grid average baseline pools visual tokens over regular 3D grid cells. It is simple and strong, but blind to anatomy: one pooled cell can mix organ tissue, lesions, vessels, and air. For instance, a small nodule or a calcified plaque occupies only a few tokens, and a grid cell that averages it into surrounding lung or muscle erases the very finding a downstream task must read. Pruning keeps tokens that appear important and discards the rest (Chen et al. 2024; Yang et al. 2024; Liu et al. 2026). Which is efficient, but it turns compression into a hard decision: evidence from a removed region cannot be recovered later, and many pruning rules require attention scores or a text query. A third route builds compact anatomy representations into the encoder itself (Shui et al. 2025; Cao et al. 2025). These models are powerful, but the compression rule is tied to the architecture and training objective, so it cannot serve as a standalone compressor at the encoder output with a different token budget. This leaves a gap: we must compress the visual tokens a 3D CT encoder produces while preserving anatomical evidence later tasks may need. We introduce ORCA (ORgan-Centroid Aggregation), a training-free compressor for 3D CT visual tokens. ORCA aggregates neighboring tokens into connected regions, using visual similarity and organ guidance to avoid mixing unrelated anatomy. The guidance does not force tokens to be pooled by predefined organ masks. ORCA stops at the target token budget, and each region becomes one compressed visual token. It also adds a centroid position encoding after aggregation. This restores spatial layout for merged regions whose original grid order no longer carries reliable position. It requires no model surgery, attention hooks, additional supervision, or changes to the downstream training recipe. Our key contributions are as follows: • We propose ORCA, a training-free compressor for 3D CT visual tokens that aggregates spatially connected, feature-similar regions with organ guidance and restores region position with sinusoidal centroid encoding. • We evaluate ORCA on CT-RATE chest CT and Merlin abdomen CT across five encoders, five attribute families, and text generation. It preserves more information than the other compressors, consistently across every encoder and attribute family, and at 64×64× compression it stays within a small margin of the uncompressed tokens. • We show the encoder and the compressor play separate roles: the encoder sets how much information the tokens carry, and the compressor governs how much of it survives to the downstream model. Figure 1: ORCA overview. A frozen 3D encoder turns a CT volume into dense visual tokens. Grid average pools them over fixed cells; ORCA instead aggregates spatially connected tokens into organ-guided regions and appends each region’s 3D centroid. Bottom-right: ORCA’s token boundaries track anatomy under compression. Related Work 3D CT vision–language models. Recent work has made volumetric CT a practical input to medical vision–language systems. CT-RATE introduced large-scale CT-volume and radiology-report data and supports CT-CLIP and CT-CHAT style modeling (Hamamci et al. 2026). CT-GRAPH studies anatomy-guided report generation from 3D CT (Kalisch et al. 2025), while Merlin extends CT foundation modeling to abdominal CT and multimodal clinical context (Blankemeier et al. 2026). Generalist radiology and 3D medical VLMs further broaden the task suite to visual question answering, report generation, retrieval, localization, and segmentation (Wu et al. 2025; Bai et al. 2024; Xin et al. 2025). COLIPRI focuses on training a stronger 3D CT encoder through language–image pretraining (Wald et al. 2025), while BTB3D learns compact volumetric tokens through reconstruction-oriented encoding of full 3D CT volumes (Hamamci et al. 2025b). These models establish the need for strong 3D visual tokens, but they treat compression as part of a fixed model pipeline. ORCA instead studies compression at the encoder-output interface and asks which visual evidence remains available under a target token budget. Visual token compression. Visual token compression has been studied for vision transformers and multimodal LLMs. Shao et al. categorize multimodal token compression by modality and by mechanism, including transformation-based, similarity-based, attention-based, and query-based approaches (Shao et al. 2025). In visual models, these mechanisms appear as pooling or downsampling, token merging, pruning, and query-based resampling. Concrete methods realize these mechanisms: ToMe merges tokens, DivPrune selects diverse ones, FastV prunes by in-LLM attention, LLaVA-PruMerge and VisionZip combine selection with merging, and TokenPacker learns a compact resampler (Bolya et al. 2023; Alvar et al. 2025; Chen et al. 2024; Shang et al. 2025; Yang et al. 2024; Li et al. 2024a). Recent analyses caution that token-reduction gains depend on evaluation design: some pruning rules underperform naive random token selection, and simple image downsampling can outperform advanced compressors on common benchmarks (Wen et al. 2025; Liao et al. 2025). This matters for 3D CT because Grid average is a common baseline: it preserves the field of view, is easy to implement, and is compatible with downstream models. Its weakness is not simplicity, but anatomy blindness. Token-efficient CT systems. CT-specific systems address token efficiency by changing the architecture or task pipeline. MedPruner shortens 3D medical visual sequences through inter-slice filtering followed by attention-based token pruning (Liu et al. 2026). Photon uses instruction-conditioned token scheduling with surrogate gradient propagation for 3D medical VQA (Fang et al. 2026). MedRegion-CT combines a region-based SlowFast tokenizer, pseudo-mask guidance, and structured lesion prompts for CT report generation (Kyung et al. 2025). CT-GRAPH uses anatomical masks to extract global and organ-level features, then refines them in a hierarchy from organs to anatomical systems and global context before report generation (Kalisch et al. 2025). Ker-VLJEPA uses zone-constrained cross-attention to compress slice embeddings into spatially grounded tokens in a thoracic CT report-generation framework (Bumgardner et al. 2026). A related line of anatomy-grounded pretraining methods, including CT-GLIP, fVLM, and ViSD-Boost, shows that organ-level alignment and anatomy-level representations can improve 3D CT understanding (Lin et al. 2024; Shui et al. 2025; Cao et al. 2025). These systems show that anatomy and token efficiency matter. They also solve a different problem from ours: their compression is tied to a particular model, task, or region definition, so none is a drop-in compressor at a matched budget. ORCA instead targets the simpler interface after a 3D encoder has produced a dense token grid, replacing Grid average with an adjustable compressor that keeps merged regions connected in 3D. Method Problem setup A 3D CT encoder converts an input volume into a regular 3D grid of visual tokens. We denote the token set by X=xii=1MX=\x_i\_i=1^M, where xi∈ℝDx_i ^D is the feature vector for token i, and qi∈[0,1]3q_i∈[0,1]^3 is its normalized 3D grid coordinate. The compressor reduces these M input tokens to B output tokens, denoted Z=zjj=1BZ=\z_j\_j=1^B, where B<MB<M. These B tokens are then passed to the multimodal projector and LLM. In all compressor comparisons, only the compressor changes. For text generation experiments, the encoder, projector architecture, LLM, training data, and training schedule are fixed. For probing experiments, the readout architecture, training split, and evaluation metric are fixed. ORCA: ORgan-Centroid Aggregation Figure 1 summarizes ORCA. ORCA compresses a dense 3D token grid into B connected regions. It starts from one region per token and repeatedly merges neighboring regions using visual similarity and soft organ-mask guidance. After merging, ORCA turns each region into one output token by averaging the visual embeddings inside the region and appending a sinusoidal encoding of the region’s 3D centroid. Connected region aggregation ORCA uses Ward linkage (Ward Jr. 1963) as its merge criterion, but applies it to 3D CT token compression rather than unconstrained clustering. Merges are restricted to standard 6-neighbor connectivity in the token grid, where regions are adjacent if they share a face. The compressor outputs B non-overlapping regions that cover the original token grid, B=R1,…,RB,⋃j=1BRj=1,…,M,Ra∩Rb=∅(a≠b).P_B=\R_1,…,R_B\,\; _j=1^BR_j=\1,…,M\,\;R_a∩ R_b= \ (a≠ b). Each region must be spatially connected in the original 3D grid, which prevents one output token from mixing disconnected parts of the scan. ORCA initializes each token as its own region. At each step, it considers only adjacent region pairs. Let x~i x_i be the feature used to score merges. The next subsection defines how organ guidance constructs it. For a region R, let μ~R=|R|−1∑i∈Rx~i μ_R=|R|^-1 _i∈ R x_i be its mean in this merge-feature space. For two adjacent regions RaR_a and RbR_b, ORCA computes the Ward increase in within-region distortion, Δ(Ra,Rb)=|Ra||Rb||Ra|+|Rb|‖μ~a−μ~b‖22. (R_a,R_b)= |R_a||R_b||R_a|+|R_b| \| μ_a- μ_b \|_2^2. It merges the pair with the smallest Δ and repeats until exactly B connected regions remain. Organ guidance ORCA uses organ masks deciding which regions to merge. In our experiments, these masks are produced by TotalSegmentator (Wasserthal et al. 2023). Two neighboring regions are favored for merging when both their visual embeddings and their organ coverage are similar. The organ information remains soft: it acts only on the merge cost. Regions may span several organs and need not follow the predefined organ masks. For each token i, let mi∈[0,1]Km_i∈[0,1]^K denote its organ-coverage vector. The entry mi,km_i,k is the proportion of voxels assigned to organ group k within the 3D patch represented by token i. ORCA makes this organ information comparable to the visual embedding before using it for merging. It first computes two scale terms over all input tokens in the current volume, Sx=∑d=1DVari(xi,d),Sm=∑k=1KVari(mi,k)+ϵ,S_x= _d=1^DVar_i(x_i,d), S_m= _k=1^KVar_i(m_i,k)+ε, where SxS_x and SmS_m are the total variances of the visual embedding and organ-coverage dimensions, and ϵε prevents division by zero. ORCA then scales the organ vector and concatenates it with the visual embedding to form the merge feature, α=λSx/Sm,x~i=[xi;αmi].α= λ S_x/S_m, x_i=[x_i;α m_i]. The parameter λ controls how strongly organ coverage affects the merge. The augmented feature is used only to build the merge tree. The visual feature part of each output token is the mean of the original visual embeddings: μj=1|Rj|∑i∈Rjxi. _j= 1|R_j| _i∈ R_jx_i. Organ coverage therefore shapes which tokens merge, not what they contain: each output token is a mean of original embeddings, with every input token contributing to exactly one output. Centroid position encoding After ORCA merges neighboring token regions, the output regions no longer form a regular grid. Their order in the output sequence therefore does not reliably indicate where they came from in the scan. ORCA records this location explicitly. For each region RjR_j, it computes the normalized 3D centroid cj=1|Rj|∑i∈Rjqi.c_j= 1|R_j| _i∈ R_jq_i. It then appends a sinusoidal encoding of this centroid (Vaswani et al. 2017), ϕ(cj)=[sin(2kπcj,a),cos(2kπcj,a)]a∈x,y,z,k=0F−1.φ(c_j)= [ (2^kπ c_j,a),\, (2^kπ c_j,a) ]_a∈\x,y,z\,\,k=0^F-1. For F=4F=4 frequencies over three axes, ϕφ adds 24 position dimensions. These dimensions are normalized per dimension and scaled by a factor s to match the typical magnitude of the visual features before concatenation, zj=[μj;sϕ^(cj)].z_j=[ _j;\,s\, φ(c_j)]. Each output token therefore carries its own location in its value, so regions of any shape or size stay localizable. Experiments Experimental setup Datasets and encoders We evaluate ORCA on CT-RATE and Merlin. CT-RATE is a non-contrast chest CT dataset with paired radiology reports and abnormality labels (Hamamci et al. 2026). Merlin is a portal-venous abdomen CT corpus with shipped reports and 30 clinical findings (Blankemeier et al. 2026). Table 1 summarizes the raw encoder-token grids and dimensions. corpus encoder pretraining grid D CT-RATE COLIPRI (Wald et al. 2025) language–image 24×24×2424×24×24 768 CT-RATE CT-CLIP (Hamamci et al. 2026) language–image 24×24×2424×24×24 512 CT-RATE BTB3D (Hamamci et al. 2025b) reconstruction 31×32×3231×32×32 18 Merlin SuPreM (Li et al. 2024b) segmentation 12×12×1212×12×12 192 Merlin SegVol (Du et al. 2023) segmentation 8×16×168×16×16 768 Table 1: Encoders used in the experiments. The pretraining column is the encoder’s training objective; D is the raw token dimension before compression. Attribute probing Attribute probing provides a representation-level readout of what information remains in compressed visual tokens (Alain and Bengio 2016; Belinkov 2022), and it is a task in its own right, since attribute prediction is how structured findings are populated in practice. In our protocol, a lightweight readout is trained directly on the compressed tokens before the projector and language model, so the measurement does not involve LLM fine-tuning, prompting, or text decoding. We evaluate this readout using the probing benchmark of Liang (2026). The benchmark organizes abnormality labels and measurements from CT volumes and segmentation masks into five families: disease, size, density, location, and texture. In this benchmark, an attribute is an individual target within one of these families. Table 2 summarizes the families, example attributes, metrics, and corresponding generation tasks. The supplementary material reports the full attribute list and per-attribute scores. Report generation Report generation evaluates compressed visual tokens in a longer-form language task. In the CT-RATE setting, the model generates a radiology report that is compared with the paired reference report and CT-RATE abnormality labels (Hamamci et al. 2026). We report lexical overlap metrics, including BLEU-1, BLEU-4, and ROUGE-L (Papineni et al. 2002; Lin 2004), and clinical-efficacy metrics: CE F1 over the abnormality labels a radiology labeler (Yan et al. 2022) assigns to each generated report, together with CRG and GREEN (Hamamci et al. 2025a; Ostmeier et al. 2024). Report generation is therefore the generation task associated with the disease family: abnormality labels are read from compressed tokens in probing and from generated text in clinical metrics. family attributes probe task metric generation task disease 18 abnormalities multi-label cls. AUROC report gen. size heart/aorta/IVC/lung regression R2R^2 measurement VQA density lung/spine/aorta HU regression R2R^2 measurement VQA location organ position regression R2R^2 measurement VQA texture lung texture regression R2R^2 measurement VQA Table 2: Attribute-probing families and their corresponding generation tasks. Measurement VQA Measurement VQA is the generation task associated with the four remaining families: size, density, location, and texture. Its gold answers are direct anatomical measurements. We use the measurement VQA questions of Liang (2026); Appendix A.2 summarizes the question types and gives examples. Continuous attributes become three-choice tertile or two-choice clinical-threshold questions, and we report exact match accuracy for each family. Because each question targets a specific organ’s attribute, accuracy measures how much precise, localized detail a compressor preserves rather than a coarse whole-volume impression. method B disease size density location texture (AUROC) (R2R^2) Uncompressed 13,824 0.848 0.727 0.915 0.252 0.811 Grid average 27 0.850 0.676 0.806 0.265 0.683 216 0.851 0.681 0.865 0.247 0.760 Slice pooling 24 0.850 0.649 0.752 0.172 0.646 DivPrune (Alvar et al. 2025) 27 0.850 0.652 0.854 0.228 0.741 216 0.851 0.681 0.887 0.260 0.753 MedPruner-DINS (Liu et al. 2026) 27 0.847 0.635 0.848 0.175 0.706 216 0.851 0.673 0.903 0.233 0.765 ToMe (Bolya et al. 2023) 27 0.844 0.617 0.786 0.172 0.647 216 0.848 0.636 0.845 0.220 0.714 MedRegion-CT (Kyung et al. 2025) N¯=549 N=549 0.850 0.694 0.902 0.271 0.810 ORCA 27 0.851 0.691 0.896 0.622 0.775 216 0.852 0.720 0.913 0.677 0.816 Table 3: CT-RATE probing on COLIPRI. Bold marks the best per budget, excluding the uncompressed reference. Cells are three-seed means; standard deviations are ≤0.03≤0.03. MedRegion-CT is reported at its mean count, N¯=549 N=549. Baselines Each baseline is an alternative compressor in ORCA’s slot, with the rest of the pipeline fixed as above. Grid average is a simple baseline that averages fixed grid cells, and slice pooling averages whole slices. DivPrune (Alvar et al. 2025) selects a diverse token subset and drops the rest. MedPruner-DINS (Liu et al. 2026) keeps high-attention tokens and folds lower-attention tokens into them. ToMe (Bolya et al. 2023) merges tokens by feature similarity, without the 3D connectivity or anatomical guidance ORCA uses. MedRegion-CT pooling (Kyung et al. 2025) averages tokens within each organ mask, the same-domain organ-pooling baseline. We also report an uncompressed-grid reference. Some of these baselines were designed for 2D-slice-stack encoders, so we adapt each one faithfully to our 3D-native interface; Appendix A.4 gives the exact realization of every baseline. Matched-budget protocol To attribute differences to which tokens survive rather than how many, we fix the token budget whenever a method exposes an adjustable target count. Some baselines instead have a count fixed by construction. We report those methods at their natural operating point. MedRegion-CT pooling is the main such case because its count is fixed by anatomy. method B disease size density location texture CT-CLIP Grid average 27 0.741 0.584 0.560 0.516 0.557 216 0.743 0.612 0.570 0.529 0.566 ToMe 27 0.665 0.246 0.306 0.155 0.350 216 0.715 0.422 0.472 0.256 0.499 ORCA 27 0.738 0.615 0.548 0.649 0.542 216 0.748 0.649 0.586 0.645 0.564 BTB3D Grid average 27 0.616 0.171 0.194 0.080 0.179 216 0.608 0.179 0.211 0.064 0.180 ToMe 27 0.588 0.100 0.123 0.039 0.119 216 0.599 0.126 0.157 0.046 0.144 ORCA 27 0.672 0.420 0.345 0.564 0.361 216 0.679 0.440 0.405 0.434 0.401 Table 4: CT-RATE probing on CT-CLIP and BTB3D. Disease is macro-AUROC; other families are R2R^2. Cells are means over three seeds with standard deviations ≤0.03≤0.03. Probing results CT-RATE Tables 3 and 4 report the three CT-RATE encoders. On COLIPRI, ORCA is level with the best baselines on size, density, and texture and far ahead on location: 0.6770.677 at B=216B=216 against 0.2470.247 for Grid average and 0.2710.271 for MedRegion-CT pooling, and it keeps almost all of that lead down to B=27B=27. The location gain comes from recording where each region sits: ORCA appends a sinusoidal encoding of the region centroid (Tancik et al. 2020; Vaswani et al. 2017), which a model reads far more easily than bare coordinates. The other two encoders place ORCA in a wider frame. They are trained for different objectives: BTB3D for pixel reconstruction, CT-CLIP for report–image contrastive alignment, and COLIPRI for that alignment plus report generation and masked image modeling. Objectives tied to language leave more of the probed content in the tokens than pixel reconstruction does, so on the attributes we measure COLIPRI >> CT-CLIP >> BTB3D. Two factors then separate cleanly: the encoder fixes the ceiling, the compressor decides how much of it survives. COLIPRI with ORCA is the best pair on nearly every family. What makes ORCA general is that it has no failure mode: across encoders and all five attributes it stays at or near the best, whereas the alternatives each fall short somewhere, most visibly on location. encoder method B disease size density location SuPreM Uncompressed 1,728 0.792 0.510 0.749 0.431 Grid average 216 0.787 0.462 0.681 0.395 ToMe 216 0.777 0.375 0.649 0.300 ORCA 216 0.795 0.544 0.795 0.600 Grid average 64 0.788 0.441 0.661 0.399 ToMe 64 0.768 0.289 0.584 0.235 ORCA 64 0.788 0.536 0.792 0.602 SegVol Uncompressed 2,048 0.755 0.481 0.683 0.467 Grid average 256 0.746 0.466 0.626 0.378 MedPruner-DINS 256 0.765 0.566 0.667 0.560 ToMe 256 0.748 0.345 0.629 0.436 ORCA 256 0.761 0.588 0.688 0.596 Grid average 32 0.738 0.338 0.556 0.322 MedPruner-DINS 32 0.758 0.460 0.597 0.432 ToMe 32 0.738 0.234 0.552 0.289 ORCA 32 0.754 0.580 0.696 0.656 Table 5: Merlin probing, means over three seeds. SegVol probing is noisier, sd up to 0.050.05, while SuPreM is stable, sd ≤0.02≤0.02. Disease is macro-AUROC; size, density, and location are R2R^2. MedPruner-DINS runs on SegVol only; SuPreM’s windowed attention exposes no global token saliency. Figure 2: Budget curve on COLIPRI. Probe R2R^2 versus token budget. Each measured budget is an equal slot; the axis is not linear in B. Bands are ±1± 1 sd over three seeds. Grid average is ratio-based and is absent at the non-cubic budget 125125. report generation Measurement VQA (accuracy) method B BLEU-1 BLEU-4 ROUGE-L CE-F1 CRG GREEN size density location texture Noise tokens 8 0.441 0.191 0.278 0.306 0.366 0.226 0.340 0.342 0.337 0.357 Grid average 27 0.472 0.206 0.292 0.470 0.437 0.259 0.759 0.768 0.601 0.783 216 0.445 0.191 0.293 0.475 0.438 0.319 0.766 0.807 0.649 0.795 MedPruner-DINS 27 0.458 0.195 0.293 0.482 0.442 0.288 0.737 0.800 0.494 0.811 216 0.454 0.196 0.297 0.483 0.444 0.317 0.759 0.837 0.527 0.830 ORCA 27 0.442 0.190 0.294 0.477 0.442 0.313 0.760 0.838 0.677 0.832 216 0.446 0.196 0.299 0.470 0.436 0.329 0.778 0.847 0.721 0.838 Table 6: Text generation results on CT-RATE, single seed. The Noise row replaces the visual tokens with Gaussian noise, giving a language-prior floor. Merlin Table 5 reports probing on Merlin with SuPreM and SegVol. Merlin has no texture family because the texture targets are lung-based. The pattern is the same: ORCA leads on size, density, and location at both budgets, most strongly on location, and disease is saturated and near-tied across methods. What makes this a real test is how much changes: Merlin is a different anatomy (abdomen) and acquisition (portal-venous), and SuPreM and SegVol are a third kind of encoder, segmentation-pretrained rather than language-supervised. That ORCA’s advantage survives all of this is the strongest sign it is not tied to one encoder family or dataset. Its location lead even holds at the tightest budgets. Budget curves. Figure 2 sweeps the full budget range on COLIPRI. Every curve rises with the budget, since more tokens carry more of the attribute information, yet ORCA keeps its lead across the whole sweep and its location margin never closes. Disease is omitted because it stays near-saturated at every budget: a global finding that survives even heavy compression, so its curve is flat. Text generation results We fine-tune the language model on each compressed cell for visual question answering (VQA) and report generation (Table 6), asking whether a compressor’s advantage carries over to LLM generation tasks. On VQA, ORCA leads every family at both budgets, and the gap is widest on location: at B=216B=216 ORCA reaches 0.7210.721 against 0.6490.649 for Grid average and 0.5270.527 for MedPruner-DINS. Appendix B.4 reports additional budgets and Appendix B.6 the training curves. Report generation’s three metric families separate methods to very different degrees. GREEN, the most comprehensive clinical score, separates them: ORCA leads at both budgets, widest at B=27B=27 (0.3130.313 against 0.2880.288 for MedPruner-DINS and 0.2590.259 for Grid average), while a noise-token control collapses to 0.2260.226, confirming the signal comes from the compressed tokens. Because GREEN grades many findings rather than one disease label, this ordering matches the overall probing and VQA picture. The coarse CE-F1 label barely separates compressors and at times inverts them, within 0.020.02 across methods and budgets. This too matches probing, where the disease family is saturated, so a whole-volume judgment that survives heavy pooling cannot resolve compressors and is the more fragile signal in the longer generation pipeline. Lexical overlap carries the least: noise tokens still score BLEU-1 near 0.440.44, indistinguishable from real compressors, since it follows the language prior, not the image. Report generation thus tells the same story as VQA and probing: ORCA preserves the clinically usable signal best, most visibly under heavy compression. Ablation study As a strong reference for how much a compressor could recover, we read out the full uncompressed token set with sinusoidal centroids (in Table 7). ORCA meets this reference to within a small margin, typically 0.010.01 to 0.110.11 R2R^2, all at 88 to 64×64× fewer tokens: it recovers most of the recoverable information at a fraction of the tokens. ORCA recovers this information in two ways: exposing spatial position, and aggregating tokens by content. Position access. Without explicit position, both Grid average and the aggregated tokens sit at the location floor, about 0.250.25 on COLIPRI; the sinusoidal centroid lifts them to about 0.660.66, close to uncompressed-with-centroid’s 0.770.77. The gap is not missing information but unreadable information: without explicit position, location survives only in the token ordering, which a readout cannot use (Zaheer et al. 2017). ORCA writes the region centroid into the token value, where the readout can use it, and closes most of that gap. The same position also lifts size, whose targets are spatially defined. encoder variant disease size density location texture CT-RATE COLIPRI B=216B=216 Grid average 0.851 0.681 0.865 0.247 0.760 + centroid 0.852 0.703 0.876 0.668 0.780 aggregate 0.852 0.683 0.907 0.251 0.784 + centroid 0.853 0.711 0.908 0.657 0.786 !15 + organ (ORCA) 0.852 0.720 0.913 0.677 0.816 Uncompressed + centroid 0.849 0.791 0.930 0.773 0.829 CT-CLIP B=216B=216 Grid average 0.743 0.612 0.570 0.529 0.566 + centroid 0.758 0.665 0.601 0.663 0.582 aggregate 0.734 0.535 0.550 0.443 0.550 + centroid 0.745 0.623 0.584 0.623 0.560 !15 + organ (ORCA) 0.748 0.649 0.586 0.645 0.564 Uncompressed + centroid 0.761 0.664 0.611 0.661 0.591 BTB3D B=512B=512 Grid average 0.608 0.177 0.219 0.062 0.190 + centroid 0.659 0.374 0.369 0.360 0.356 aggregate 0.613 0.159 0.177 0.034 0.153 + centroid 0.676 0.332 0.320 0.238 0.336 !15 + organ (ORCA) 0.705 0.473 0.436 0.401 0.426 Uncompressed + centroid 0.686 0.465 0.480 0.448 0.460 Merlin SuPreM B=216B=216 Grid average 0.787 0.462 0.681 0.395 — + centroid 0.795 0.523 0.752 0.551 — aggregate 0.787 0.474 0.730 0.387 — + centroid 0.798 0.550 0.785 0.584 — !15 + organ (ORCA) 0.795 0.544 0.795 0.600 — Uncompressed + centroid 0.806 0.622 0.845 0.709 — SegVol B=256B=256 Grid average 0.746 0.466 0.626 0.378 — + centroid 0.759 0.464 0.636 0.506 — aggregate 0.743 0.466 0.672 0.478 — + centroid 0.760 0.499 0.700 0.553 — !15 + organ (ORCA) 0.761 0.588 0.688 0.596 — Uncompressed + centroid 0.780 0.675 0.790 0.668 — Table 7: Component ablation. The last row, “Uncompressed + centroid”, adds the position to the uncompressed tokens, where each centroid is a single grid patch. Per-cell standard deviations are ≤0.03≤0.03 on COLIPRI and SuPreM and up to 0.060.06 on SegVol. Anatomy access. ORCA merges tokens by feature similarity instead of a fixed grid cell, with an organ mask steering the merge. The merge recovers content that Grid average blends away: on density, ORCA lifts its 0.8650.865 (COLIPRI) to 0.9130.913, close to uncompressed-with-centroid’s 0.9300.930. The organ mask only helps – adding it never meaningfully hurts (the largest drop is under 0.010.01, within seed noise). It can usually be left on without per-encoder tuning. Disease saturation. The disease column moves differently from the others. Among the CT-RATE encoders, which share one abnormality target, ORCA lifts disease by only +0.001+0.001 (COLIPRI) and +0.005+0.005 (CT-CLIP) over Grid average, but by +0.097+0.097 on BTB3D. The two flat cases are the report-supervised encoders: their pretraining already read radiology reports, so the abnormality signal is redundantly encoded and the column is saturated – no compressor adds to it. BTB3D, trained only to reconstruct, never saw that signal, leaving headroom that ORCA recovers. The Merlin encoders (SuPreM, SegVol), trained by segmentation without disease labels, show small positive gains on their own target, consistent with this reading. The organ-guidance weight λ and the position-encoding hyperparameters (frequency F, scale s) are robust knobs rather than tuned parameters: probe R2R^2 is nearly flat across four orders of magnitude of λ and across the frequencies and scales we tried, with location the only responsive family. We fix λ=0.5λ=0.5 (COLIPRI), λ=2λ=2 (Merlin), and F=4,s=2F=4,s=2 everywhere. The sweeps are in Appendix B.3. Efficiency The LLM computational gain depends mainly on the number of visual tokens that enter the LLM context, not on which compressor produced them. Table 8 therefore reports cost as a function of token budget. Prefill latency and KV-cache memory fall steeply, while peak memory falls only modestly but crosses the threshold that lets the budget run at all: on a commodity L4 (24 GB) the uncompressed context otherwise does not fit. End-to-end report time moves little across budgets, since these short runs are decode dominated. prefill (ms) s / report B FLOPs (T) KV (MB) peak (GB) B200 L4 B200 L4 13,824 324 1820 27.1 371.5 OOM 1.84 OOM 216 4.5 36.7 16.3 11.8 132.9 1.43 8.06 64 2.1 16.8 16.2 11.5 080.5 1.43 7.98 27 1.5 11.9 16.1 11.6 078.3 1.43 7.98 8 1.2 09.4 16.1 11.6 077.0 1.43 7.98 Table 8: LLM cost vs. token budget B on Llama-3.1-8B (bf16), for a B200 and a commodity L4 (24 GB). Each latency is the median of three alternating sweeps (warmup discarded); FLOPs, KV-cache, and peak memory depend on B only, not the GPU. Discussion and Limitations Preserve the anatomical information. For a 3D CT embedding, the goal of compression is to preserve the information a downstream model needs, not to reconstruct the original tokens as faithfully as possible. That information is spatial and heterogeneous. Because it is anatomical rather than task-specific, even a training-free operator preserves most of it. By writing each region’s centroid and merging by content, ORCA stays within a small margin of the uncompressed tokens at matched budgets. We test only two task types here, but the approach should carry over to other 3D CT tasks. Attributes beyond disease. The dominant way to score a 3D CT model is report generation and the VQA derived from it, with its metrics anchored on disease findings. This does not fully cover the attributes a CT volume carries. To reach the rest, we evaluate on the mask-derived measurement VQA benchmark of Liang (2026), whose labels come from anatomical measurements, are defined for every scan, and span size, density, and location alongside disease. Across encoders, most are optimized and benchmarked mainly for disease, leaving these other attributes largely unmeasured. Keeping the budget adjustable. A natural alternative is to pool tokens directly within each organ mask, but the obstacle is budget, not accuracy. Hard organ pooling emits one token per organ or organ slice, so the token count is fixed by anatomy rather than chosen by the user. ORCA instead makes the budget a free parameter, and the more tokens it is given, the more information it preserves. Conclusion We presented ORCA, a training-free compressor that aggregates 3D CT visual tokens into connected, content-adaptive regions and writes each region’s position back into the token. Across two datasets and five encoders it improves over Grid average and matched token-reduction baselines, and stays within a small margin of the uncompressed tokens at 88 to 64×64× fewer tokens, with no compression network trained. The principle is simple: what compression preserves is governed by what the downstream reader can read, so a good compressor makes position explicit and aggregates by content rather than by a fixed grid. Structure-preserving aggregation with position re-injection is a strong drop-in replacement for Grid average in 3D CT vision-language pipelines. References G. Alain and Y. Bengio (2016) Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644. External Links: Document Cited by: Attribute probing. S. R. Alvar, G. Singh, M. Akbari, and Y. Zhang (2025) DivPrune: diversity-based visual token pruning for large multimodal models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §A.4, Related Work, Baselines, Table 3. F. Bai, Y. Du, T. Huang, M. Q.-H. Meng, and B. Zhao (2024) M3D: advancing 3D medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578. External Links: Document Cited by: Related Work. Y. Belinkov (2022) Probing classifiers: promises, shortcomings, and advances. Computational Linguistics 48 (1), p. 207–219. External Links: Document Cited by: Attribute probing. L. Blankemeier, A. Kumar, J. P. Cohen, et al. (2026) Merlin: a computed tomography vision–language foundation model and dataset. Nature 652, p. 1318–1328. External Links: Document Cited by: Table 9, Introduction, Related Work, Datasets and encoders. D. Bolya, C. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman (2023) Token merging: your ViT but faster. In International Conference on Learning Representations (ICLR), Note: arXiv:2210.09461 Cited by: §A.4, Table 13, Related Work, Baselines, Table 3. V. K. C. Bumgardner, M. A. Klusty, M. S. Gokmen, and E. W. Damron (2026) Curriculum-driven 3D CT report generation via language-free visual grafting and zone-constrained compression. arXiv preprint arXiv:2603.23308. External Links: Document Cited by: Related Work. W. Cao, J. Zhang, Z. Shui, S. Wang, Z. Chen, X. Li, L. Lu, X. Ye, T. Liang, Q. Zhang, and L. Zhang (2025) Boosting vision semantic density with anatomy normality modeling for medical vision-language pre-training. arXiv preprint arXiv:2508.03742. Cited by: Introduction, Related Work. L. Chen, H. Zhao, T. Liu, S. Bai, J. Lin, C. Zhou, and B. Chang (2024) An image is worth 1/2 tokens after layer 2: plug-and-play inference acceleration for large vision-language models. arXiv preprint arXiv:2403.06764. External Links: Document Cited by: Introduction, Related Work. M. S. Davenport, T. Fruscello, M. Chatfield, S. Weinstein, W. F. Sensakovic, and D. B. Larson (2021) CT volumes from 2,398 radiology practices in the united states: a real-time indicator of the effect of COVID-19 on routine care, january to september 2020. Journal of the American College of Radiology 18 (3), p. 380–387. External Links: Document Cited by: Introduction. Y. Du, F. Bai, T. Huang, and B. Zhao (2023) SegVol: universal and interactive volumetric medical image segmentation. arXiv preprint arXiv:2311.13385. External Links: Document Cited by: Table 1. C. Fang, H. Guo, Z. Jiang, C. He, X. Li, and M. Xu (2026) Photon: speedup volume understanding with efficient multimodal large language models. In International Conference on Learning Representations (ICLR), Note: arXiv:2603.25155 Cited by: Related Work. I. E. Hamamci, S. Er, and B. H. Menze (2024) CT2Rep: automated radiology report generation for 3D medical imaging. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2024, Lecture Notes in Computer Science, Vol. 15012, p. 476–486. External Links: Document Cited by: Introduction. I. E. Hamamci, S. Er, S. Shit, H. Reynaud, B. Kainz, and B. H. Menze (2025a) CRG score: a distribution-aware clinical metric for radiology report generation. External Links: 2505.17167, Document Cited by: Report generation. I. E. Hamamci, S. Er, S. Shit, H. Reynaud, D. Yang, P. Guo, M. Edgar, D. Xu, B. Kainz, and B. Menze (2025b) Better tokens for better 3D: advancing vision-language modeling in 3D medical imaging. arXiv preprint arXiv:2510.20639. Note: NeurIPS 2025 External Links: Document Cited by: Related Work, Table 1. I. E. Hamamci, S. Er, C. Wang, F. Almas, A. G. Simsek, S. N. Esirgun, I. Dogan, O. F. Durugol, B. Hou, S. Shit, W. Dai, M. Xu, H. Reynaud, M. F. Dasdelen, B. Wittmann, T. Amiranashvili, E. Simsar, M. Simsar, E. B. Erdemir, A. Alanbay, A. Sekuboyina, B. Lafci, A. Kaplan, Z. Lu, M. Polacin, B. Kainz, C. Bluethgen, K. Batmanghelich, M. K. Ozdemir, and B. Menze (2026) Generalist foundation models from a multimodal dataset for 3D computed tomography. Nature Biomedical Engineering. External Links: Document Cited by: Table 9, Introduction, Related Work, Datasets and encoders, Report generation, Table 1. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2021) LoRA: low-rank adaptation of large language models. External Links: 2106.09685, Document Cited by: §A.3. M. Ilse, J. Tomczak, and M. Welling (2018) Attention-based deep multiple instance learning. In Proceedings of the 35th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 80, p. 2127–2136. External Links: Link Cited by: §A.3. H. Kalisch, F. Hörst, J. Kleesiek, K. Herrmann, and C. Seibold (2025) CT-GRAPH: hierarchical graph attention network for anatomy-guided CT report generation. arXiv preprint arXiv:2508.05375. External Links: Document Cited by: Related Work, Related Work. S. Kyung, J. Seo, H. Lim, D. Kim, H. Park, J. Sung, J. Kim, W. Jo, Y. Nam, and N. Kim (2025) Region-aware multimodal large language model via slowfast tokenization and pseudo-mask guidance for 3D CT report generation. arXiv preprint arXiv:2506.23102. Note: Accepted to ECCV 2026 External Links: Document Cited by: §A.4, Related Work, Baselines, Table 3. W. Li, Y. Yuan, J. Liu, D. Tang, S. Wang, J. Qin, J. Zhu, and L. Zhang (2024a) TokenPacker: efficient visual projector for multimodal LLM. arXiv preprint arXiv:2407.02392. External Links: Document Cited by: Related Work. W. Li, A. Yuille, and Z. Zhou (2024b) How well do supervised 3D models transfer to medical imaging tasks?. In International Conference on Learning Representations (ICLR), Cited by: Table 1. R. Liang (2026) Cheap probes predict expensive training in 3D-CT vision-language models. External Links: 2607.22771, Link Cited by: §A.2, §A.2, Attribute probing, Measurement VQA, Attributes beyond disease.. C. Liao, W. Wang, Z. Wen, X. Zheng, Y. Wang, H. He, Y. Lyu, L. Jiang, X. Zou, Y. Fu, B. Ren, L. Zhang, and X. Hu (2025) Are we using the right benchmark: an evaluation framework for visual token compression methods. arXiv preprint arXiv:2510.07143. Note: Accepted by ACL 2026 Main External Links: Document Cited by: Related Work. C. Lin (2004) ROUGE: a package for automatic evaluation of summaries. In Text Summarization Branches Out, p. 74–81. External Links: Link Cited by: Report generation. J. Lin, Y. Xia, J. Zhang, K. Yan, K. Cao, L. Lu, J. Luo, and L. Zhang (2024) CT-GLIP: 3D grounded language-image pretraining with CT scans and radiology reports for full-body scenarios. arXiv preprint arXiv:2404.15272. External Links: Document Cited by: Related Work. H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. In Advances in Neural Information Processing Systems, Vol. 36, p. 34892–34916. Cited by: §A.3. S. Liu, Z. Ye, Y. Lin, C. Hu, W. Geng, X. Han, B. Ibragimov, Y. Zheng, and Y. Yuan (2026) MedPruner: training-free hierarchical token pruning for efficient 3D medical image understanding in vision-language models. arXiv preprint arXiv:2603.11625. Cited by: §A.4, Table 13, Introduction, Related Work, Baselines, Table 3. Llama Team (2024) The Llama 3 herd of models. External Links: 2407.21783, Document Cited by: §A.3. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: Link Cited by: §A.3. S. Ostmeier, J. Xu, Z. Chen, M. Varma, L. Blankemeier, C. Bluethgen, A. E. Michalson, M. Moseley, C. Langlotz, A. S. Chaudhari, and J. Delbrouck (2024) GREEN: generative radiology report evaluation and error notation. External Links: 2405.03595, Document Cited by: Report generation. K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, p. 311–318. External Links: Document Cited by: Report generation. Y. Shang, M. Cai, B. Xu, Y. J. Lee, and Y. Yan (2025) LLaVA-PruMerge: adaptive token reduction for efficient large multimodal models. In IEEE/CVF International Conference on Computer Vision (ICCV), Note: arXiv:2403.15388 Cited by: Related Work. K. Shao, K. Tao, K. Zhang, S. Feng, M. Cai, Y. Shang, H. You, C. Qin, Y. Sui, and H. Wang (2025) A survey of token compression for efficient multimodal large language models. arXiv preprint arXiv:2507.20198. External Links: Document Cited by: Introduction, Related Work. Z. Shui, J. Zhang, W. Cao, S. Wang, R. Guo, L. Lu, L. Yang, X. Ye, T. Liang, Q. Zhang, and L. Zhang (2025) Large-scale and fine-grained vision-language pre-training for enhanced CT image understanding. arXiv preprint arXiv:2501.14548. Cited by: Introduction, Related Work. M. Tancik, P. P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. T. Barron, and R. Ng (2020) Fourier features let networks learn high frequency functions in low dimensional domains. In Advances in Neural Information Processing Systems, Vol. 33. Cited by: CT-RATE. J. J. M. van Griethuysen, A. Fedorov, C. Parmar, A. Hosny, N. Aucoin, V. Narayan, R. G. H. Beets-Tan, J. Fillion-Robin, S. Pieper, and H. J. W. L. Aerts (2017) Computational radiomics system to decode the radiographic phenotype. Cancer Research 77 (21), p. e104–e107. Cited by: Table 14. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin (2017) Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: Centroid position encoding, CT-RATE. T. Wald, I. E. Hamamci, Y. Gao, S. Bond-Taylor, H. Sharma, et al. (2025) Comprehensive language–image pre-training for 3D medical image understanding. External Links: 2510.15042 Cited by: Introduction, Related Work, Table 1. J. H. Ward Jr. (1963) Hierarchical grouping to optimize an objective function. Journal of the American Statistical Association 58 (301), p. 236–244. Cited by: Connected region aggregation. J. Wasserthal, H. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang, M. Bach, and M. Segeroth (2023) TotalSegmentator: robust segmentation of 104 anatomic structures in CT images. Radiology: Artificial Intelligence 5 (5), p. e230024. External Links: Document Cited by: §A.1, Organ guidance. Z. Wen, Y. Gao, W. Li, C. He, and L. Zhang (2025) Token pruning in multimodal large language models: are we solving the right problem?. In Findings of the Association for Computational Linguistics: ACL, Note: arXiv:2502.11501 Cited by: Related Work. C. Wu, X. Zhang, Y. Zhang, H. Hui, Y. Wang, and W. Xie (2025) Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nature Communications 16, p. 7866. External Links: Document Cited by: Introduction, Related Work. Y. Xin, G. C. Ates, K. Gong, and W. Shao (2025) Med3DVLM: an efficient vision-language model for 3D medical image analysis. arXiv preprint arXiv:2503.20047. External Links: Document Cited by: Related Work. A. Yan, J. McAuley, X. Lu, J. Du, E. Y. Chang, A. Gentili, and C. Hsu (2022) RadBERT: adapting transformer-based language models to radiology. Radiology: Artificial Intelligence 4 (4), p. e210258. External Links: Document Cited by: §A.1, Report generation. S. Yang, Y. Chen, Z. Tian, C. Wang, J. Li, B. Yu, and J. Jia (2024) VisionZip: longer is better but not necessary in vision language models. arXiv preprint arXiv:2412.04467. External Links: Document Cited by: Introduction, Related Work. M. Zaheer, S. Kottur, S. Ravanbakhsh, B. Póczos, R. Salakhutdinov, and A. J. Smola (2017) Deep sets. In Advances in Neural Information Processing Systems, Vol. 30. Cited by: Position access.. H. Zamani, T. Fruscello, J. Burleson, M. Bhargavan-Chatfield, and M. S. Davenport (2026) US radiology imaging and workforce volumes 2017–2024: an analysis of 46.4 million imaging examinations from 167 radiology facilities. Journal of the American College of Radiology 23 (6), p. 1041–1048. External Links: Document Cited by: Introduction. Supplementary Material ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression Appendix A Experimental setup A.1 Datasets and encoders CT-RATE Merlin anatomy chest abdomen contrast non-contrast portal-venous patients 21,304 18,317 abnormality labels 18 30 CT volumes 25,692 25,494 train / valid 24,128 / 1,564 15,309 / 5,055 reconstructions 50,188 25,494 train / valid 47,149 / 3,039 15,309 / 5,055 Table 9: Dataset scale and imaging characteristics of CT-RATE (Hamamci et al. 2026) and Merlin (Blankemeier et al. 2026). Each CT-RATE volume may correspond to multiple reconstructions generated under different reconstruction settings, whereas each Merlin scan corresponds to a single reconstructed volume. The dataset scale and splits are reported in Table 9. Each scan in both datasets is paired with a radiology report. Abnormality labels come from RadBERT (Yan et al. 2022) and organ masks from TotalSegmentator (Wasserthal et al. 2023). CT-CLIP and BTB3D train on the reconstruction-level split, whereas COLIPRI uses the volume-level split; Merlin uses a 20,36420,364-scan subset. Figure 3: Case study on one CT volume (COLIPRI, B=27B=27) in three anatomical views. The first column is the CT scan with lung, heart, and spine contours; the remaining columns overlay each compressor’s token boundaries on the embedding-norm heatmap, and the bottom row recolors the same partitions so that each token receives one color. ORCA aggregates tokens into spatially contiguous regions that follow the organ contours. A.2 Measurement VQA benchmark Our probing and visual question answering (VQA) experiments use the measurement VQA benchmark introduced by Liang (2026). The benchmark comprises a set of image-derived anatomical attributes whose reference values are computed directly from each CT volume, its corresponding segmentation masks, and Hounsfield unit (HU) values. Because the labels are generated from deterministic image measurements, they are reproducible and avoid noise introduced by variations in radiology report wording or LLM generation. The benchmark is designed to evaluate token compression according to the anatomical information retained from the CT image. It covers attributes such as size, density, location, and texture, which are often not explicitly quantified in radiology reports. Each question targets a single organ, allowing VQA accuracy to reflect how well the compressed representation preserves localized anatomical details. The same attributes are used as regression or AUROC evaluation targets in the probing experiments and are converted into multiple-choice questions based on population tertiles or predefined clinical thresholds for VQA. The complete benchmark construction, statistical characterization, and validation are provided by Liang (2026). Table 14 details every attribute, and Figure 4 gives representative question–answer examples. Q1. “How would you characterize the cardiothoracic (heart-to-lung) size ratio?” (A) High (B) Moderate (C) Low Q2. “On this scan, the aortic diameter appears:” (A) Normal caliber, <40<40 m (B) Dilated, ≥40≥ 40 m Q3. “The aortic wall calcification burden on this CT is best described as:” (A) Present (B) Absent Q4. “On this scan, the heart horizontal (left-right) position appears:” (A) High (B) Low (C) Moderate Q5. “On this scan, the lung 15th-percentile density (emphysema index) appears:” (A) No significant emphysema, ≥−950≥-950 HU (B) Emphysema present, <−950<-950 HU Figure 4: Representative measurement VQA questions and reference answers. The reference answer is highlighted in green. A.3 Implementation details Probes. Each family is read out from the frozen compressed tokens by a lightweight attention-pooling head (Ilse et al. 2018) followed by a small MLP, trained with AdamW (Loshchilov and Hutter 2019); every probing number is the mean over three seeds. Two comparisons instead use a higher-capacity head so that no method is bottlenecked by the probe: the uncompressed reference rows in Table 7, and the Merlin centroid-encoding comparison (Table 12). Encoder features are normalized to the encoder’s scale before the readout. Exact hyperparameters are in the released code. Text generation. The generation model is a LLaVA-style pipeline (Liu et al. 2023) with a Llama-3.1-8B-Instruct backbone (Llama Team 2024) and a two-layer projector. Training is two-stage: a projector warm-up with the backbone frozen, then LoRA fine-tuning (Hu et al. 2021) of the backbone. Each encoder–compressor combination is trained separately. Validation is a full-set generation pass each epoch and the reported number is the best epoch. Hardware. We train the VQA and text-generation models on a single NVIDIA B200 GPU (180 GB) with bf16 mixed precision. All other experiments, including probing and preprocessing, run on a single NVIDIA L4 GPU (24 GB). ORCA compression itself is CPU-only and needs no GPU, so the only heavy hardware requirement is the language-model fine-tuning. size density location texture size density location texture COLIPRI B=27B=27 B=216B=216 Grid average 0.676 ± 0.005 0.806 ± 0.005 0.266 ± 0.023 0.683 ± 0.008 0.681 ± 0.002 0.865 ± 0.004 0.247 ± 0.014 0.760 ± 0.025 + centroid 0.685 ± 0.010 0.815 ± 0.006 0.489 ± 0.020 0.705 ± 0.017 0.703 ± 0.011 0.876 ± 0.010 0.668 ± 0.030 0.780 ± 0.010 DivPrune 0.652 ± 0.005 0.854 ± 0.004 0.228 ± 0.002 0.741 ± 0.010 0.680 ± 0.005 0.887 ± 0.002 0.260 ± 0.009 0.753 ± 0.015 + centroid 0.666 ± 0.005 0.852 ± 0.002 0.510 ± 0.017 0.764 ± 0.013 0.709 ± 0.006 0.894 ± 0.002 0.674 ± 0.001 0.794 ± 0.008 MedPruner-DINS 0.635 ± 0.003 0.848 ± 0.005 0.175 ± 0.004 0.706 ± 0.027 0.673 ± 0.007 0.903 ± 0.003 0.233 ± 0.017 0.765 ± 0.022 + centroid 0.645 ± 0.001 0.854 ± 0.005 0.393 ± 0.029 0.737 ± 0.018 0.703 ± 0.008 0.906 ± 0.003 0.519 ± 0.012 0.792 ± 0.015 ToMe 0.617 ± 0.002 0.786 ± 0.011 0.172 ± 0.004 0.647 ± 0.001 0.636 ± 0.010 0.845 ± 0.004 0.220 ± 0.022 0.714 ± 0.019 + centroid 0.618 ± 0.003 0.792 ± 0.003 0.212 ± 0.026 0.659 ± 0.008 0.651 ± 0.002 0.846 ± 0.002 0.338 ± 0.019 0.743 ± 0.017 ORCA 0.691 ± 0.011 0.895 ± 0.004 0.622 ± 0.005 0.775 ± 0.009 0.720 ± 0.005 0.913 ± 0.002 0.677 ± 0.013 0.816 ± 0.001 BTB3D B=27B=27 B=216B=216 Grid average 0.183 ± 0.014 0.182 ± 0.008 0.070 ± 0.004 0.194 ± 0.005 0.188 ± 0.013 0.205 ± 0.018 0.071 ± 0.003 0.185 ± 0.008 + centroid 0.324 ± 0.005 0.302 ± 0.016 0.183 ± 0.008 0.316 ± 0.011 0.378 ± 0.008 0.372 ± 0.005 0.341 ± 0.007 0.364 ± 0.003 DivPrune 0.149 ± 0.001 0.181 ± 0.001 0.056 ± 0.003 0.164 ± 0.003 0.181 ± 0.007 0.223 ± 0.010 0.069 ± 0.004 0.201 ± 0.009 + centroid 0.256 ± 0.001 0.276 ± 0.000 0.190 ± 0.002 0.266 ± 0.019 0.349 ± 0.005 0.368 ± 0.014 0.309 ± 0.011 0.370 ± 0.010 ToMe 0.100 ± 0.005 0.123 ± 0.002 0.039 ± 0.001 0.119 ± 0.003 0.126 ± 0.003 0.157 ± 0.003 0.046 ± 0.002 0.144 ± 0.004 + centroid 0.114 ± 0.003 0.124 ± 0.002 0.041 ± 0.001 0.161 ± 0.005 0.183 ± 0.006 0.196 ± 0.007 0.079 ± 0.003 0.236 ± 0.011 ORCA 0.420 ± 0.002 0.345 ± 0.013 0.564 ± 0.009 0.361 ± 0.003 0.441 ± 0.004 0.405 ± 0.004 0.434 ± 0.025 0.402 ± 0.006 Table 10: Centroid encoding on CT-RATE probing. Each baseline is shown without and with the sinusoidal centroid encoding ORCA uses; DINS does not apply to BTB3D, whose reconstruction features carry no attention saliency. The best value in each column, within each encoder, is in bold. method size density location texture B=27B=27 Grid average 0.759 0.768 0.601 0.783 + centroid 0.756 0.781 0.615 0.788 ORCA 0.760 0.838 0.675 0.829 B=216B=216 Grid average 0.766 0.807 0.649 0.795 + centroid 0.763 0.811 0.701 0.778 ORCA 0.778 0.847 0.721 0.835 Table 11: Downstream measurement VQA accuracy on CT-RATE/COLIPRI. A.4 Baseline adaptations All baselines are placed at the same encoder-output interface as ORCA. They receive the raw 3D token grid and produce a shorter token sequence before the projector. Methods differ in whether they expose an adjustable target count. Grid average, DivPrune, ToMe, and MedPruner-DINS do, and we force each to emit the same B tokens as ORCA. Slice pooling and MedRegion-CT pooling do not: their counts are fixed by construction, by encoder depth and anatomy respectively, so we report each at its natural operating point. Several of these methods were designed for 2D slice-stack encoders or include learned components in their original form. We adapt each to our 3D-native, training-free interface while keeping its core principle. Grid average. Grid average pools regular cells of the encoder grid and replaces each cell by the mean of its visual embeddings. For grids and budgets that admit an integer stride, this is non-overlapping r×r×r×r×r block averaging. For non-cubic grids or budgets that do not correspond to an integer stride, we use 3D adaptive average pooling to the nearest aspect-preserving grid. The output dimension stays equal to the encoder feature dimension. Slice pooling. Slice pooling averages each axial token plane into one token. For a grid of shape T×H×WT×H×W, it returns T tokens, each the mean over one H×WH×W plane. It is a coarse slice-level baseline with no within-plane or organ structure. Its token count is fixed by the encoder depth, so we report it at its actual count. DivPrune. DivPrune (Alvar et al. 2025) is adapted as a feature-diversity selector on the 3D encoder tokens. We first apply uniform 3D average pooling to a 1,0241,024-token candidate set, then select B tokens by farthest-first diversity in cosine-normalized feature space. This keeps the diversity principle of DivPrune. The method was designed to prune the tokens of a 2D-slice VLM, and we adapt it to operate on the dense 3D grid. It is a pure selection baseline: tokens not selected do not contribute to the output. ToMe. ToMe (Bolya et al. 2023) is adapted as feature-similarity merging on the flattened 3D token sequence. We first uniformly pool the dense grid to a 1,0241,024-token candidate set, then run bipartite cosine matching and repeatedly merge the most similar token pairs until B tokens remain. This baseline merges rather than drops tokens. MedRegion-CT pooling. MedRegion-CT’s original system learns a SlowFast tokenizer with pseudo-mask guidance and structured prompts. Only its region-pooling recipe transfers to our fixed-encoder, training-free interface, so we reproduce that part faithfully (Kyung et al. 2025): one global token per axial plane plus one region token for each organ present in that plane. The global token is the mean over all tokens in the plane. The region token is an organ-occupancy-weighted mean over tokens in the same plane. We do not force this method to match a target budget, because its token count is determined by how many slice–organ pairs are present. It therefore emits a variable count, N¯=549 N=549 on CT-RATE, which we report at that single operating point. MedPruner-DINS. MedPruner (Liu et al. 2026) has two stages: inter-slice filtering and token-level Dynamic Information Nucleus Selection (DINS). Its inter-slice filtering stage assumes a 2D slice-stack encoder, where tokens are produced separately per slice. Our encoders are 3D-native and already fold the depth axis into one volumetric token grid, so we drop the inter-slice filtering and keep only the DINS stage. We score tokens by the attention they receive in the vision encoder, keep the top B tokens, and fold the remaining tokens into their most similar kept token by cosine similarity before averaging. This produces exactly B output tokens while preserving MedPruner’s idea that low-attention tokens can still contribute through residual merging. We run this baseline only where the encoder exposes a clean global per-token attention score. This excludes SuPreM, whose Swin-UNETR backbone uses windowed local attention. Appendix B Additional results and analysis B.1 Case study Figure 3 illustrates, on a single volume, how each compressor lays its tokens over the anatomy. ORCA aggregates tokens into spatially contiguous regions that follow the organ contours. Grid average instead imposes a fixed grid that cuts across organ borders, and ToMe and MedPruner-DINS merge tokens by feature similarity with no spatial regularity, scattering each token into many disconnected fragments. The bottom-row token-layout maps make this contrast clear: ORCA keeps each token to a single connected region aligned with anatomy, an advantage the pooling and pruning baselines lack. Figure 5: Sensitivity of probe R2R^2 to the organ-guidance weight λ. B.2 Centroid encoding To separate the contribution of explicit position encoding from that of token aggregation, we give every baseline the same sinusoidal centroid encoding ORCA uses, applied to each method’s own token centroids, and re-evaluate probing (Tables 10 and 12) and downstream VQA (Table 11). Because a centroid can be computed for any retained or merged token, this equalizes position across methods, so any remaining gap reflects how the tokens are formed rather than whether they carry position. On Merlin, whose encoders are sensitive to probe capacity, we read out with a high-capacity probe so that no baseline is limited by the probe; COLIPRI is capacity-insensitive and uses the default probe. Adding position helps every baseline, and the gain concentrates almost entirely on location. On COLIPRI it lifts Grid average from 0.2470.247 to 0.6680.668 and DivPrune from 0.2600.260 to 0.6740.674 at B=216B=216, while the content families move by only a few points. Position injection is thus a general, method-agnostic benefit for the one family that depends on it. ToMe is the exception, with a far smaller location gain: because it scatters each token across the volume, its centroid is unrepresentative and the encoding it receives is close to noise, foreshadowing that a centroid is only useful when the token it summarizes is spatially coherent. Even with position equalized, ORCA still leads, most clearly on the content families and under strong compression. On COLIPRI it keeps a margin on size, density, and texture at both budgets, for example density 0.9130.913 against 0.9060.906 for the best centroid-augmented baseline, and its location advantage is large at B=27B=27 (0.6220.622 against 0.5100.510) though it narrows to a near tie at B=216B=216, where position alone nearly suffices. The gap that the centroid cannot close is therefore on content, and it comes from ORCA’s adaptive, organ-aligned aggregation rather than from position. This is starkest on the weak reconstruction encoder BTB3D, where centroid-augmented baselines still recover little location signal (∼0.18 \!0.18) while ORCA reaches 0.5640.564; on Merlin, ORCA leads every column. Figure 6: VQA accuracy across token budgets on CT-RATE with the COLIPRI encoder. size density location size density location SuPreM B=64B=64 B=216B=216 Grid average 0.519 ± 0.003 0.741 ± 0.013 0.537 ± 0.015 0.544 ± 0.009 0.786 ± 0.006 0.551 ± 0.018 + centroid 0.531 ± 0.003 0.736 ± 0.011 0.566 ± 0.005 0.578 ± 0.008 0.795 ± 0.004 0.655 ± 0.015 DivPrune 0.521 ± 0.006 0.770 ± 0.006 0.501 ± 0.011 0.569 ± 0.010 0.803 ± 0.018 0.577 ± 0.021 + centroid 0.528 ± 0.008 0.752 ± 0.005 0.577 ± 0.000 0.589 ± 0.005 0.805 ± 0.006 0.665 ± 0.010 ToMe 0.350 ± 0.017 0.624 ± 0.002 0.261 ± 0.002 0.439 ± 0.019 0.705 ± 0.008 0.351 ± 0.013 + centroid 0.348 ± 0.010 0.619 ± 0.008 0.263 ± 0.012 0.449 ± 0.006 0.716 ± 0.004 0.380 ± 0.010 ORCA 0.585 ± 0.010 0.836 ± 0.003 0.700 ± 0.007 0.609 ± 0.004 0.851 ± 0.005 0.721 ± 0.011 SegVol B=32B=32 B=256B=256 Grid average 0.599 ± 0.007 0.722 ± 0.004 0.549 ± 0.021 0.662 ± 0.011 0.778 ± 0.009 0.662 ± 0.008 + centroid 0.585 ± 0.012 0.720 ± 0.005 0.524 ± 0.007 0.633 ± 0.006 0.774 ± 0.003 0.620 ± 0.011 DivPrune 0.562 ± 0.011 0.706 ± 0.009 0.552 ± 0.009 0.677 ± 0.005 0.785 ± 0.020 0.700 ± 0.009 + centroid 0.551 ± 0.006 0.716 ± 0.002 0.531 ± 0.033 0.659 ± 0.006 0.806 ± 0.014 0.675 ± 0.053 ToMe 0.506 ± 0.005 0.651 ± 0.005 0.468 ± 0.012 0.633 ± 0.010 0.749 ± 0.007 0.683 ± 0.005 + centroid 0.478 ± 0.011 0.638 ± 0.005 0.466 ± 0.006 0.620 ± 0.006 0.754 ± 0.002 0.678 ± 0.025 ORCA 0.680 ± 0.001 0.818 ± 0.010 0.710 ± 0.016 0.685 ± 0.004 0.833 ± 0.010 0.747 ± 0.007 Table 12: Centroid encoding on Merlin probing, read out with a high-capacity probe (dimension 10241024) so that no method is limited by the probe. DINS does not apply to SuPreM’s windowed attention, and Merlin has no texture family. The best value in each column, within each encoder, is in bold. B.3 Sensitivity to the organ mask Figure 5 sweeps the organ-guidance weight λ from 0, no mask, to 10410^4, the organ-dominated limit that reduces to hard within-organ pooling, for each family on all four encoders. Across a wide range the curves vary only gradually with no sharp optimum, and even the organ-dominated limit does not collapse them, so the mask weight is a robust knob rather than a fragile tuned parameter. The family that consistently responds is location, which rises with λ on every encoder, for instance 0.660.66 to 0.710.71 on COLIPRI and 0.540.54 to 0.670.67 on SegVol, consistent with the mask supplying organ identity that most directly aids localization. The gains are uneven across encoders: the COLIPRI curves move only a little while the BTB3D curves rise more. We observe and speculate that COLIPRI’s embeddings are already distributed much like the organs themselves, so the organ mask adds little, whereas BTB3D’s embeddings look closer to random with no clear organ structure, so the mask helps more. B.4 Budget scaling Figure 6 plots best epoch VQA accuracy against the token budget for the three compressors on CT-RATE with the COLIPRI encoder, one panel per family. ORCA leads at every budget in every family, and its margin is widest on location and density while size and texture are close to saturated. The scaling is efficient: ORCA at B=8B=8 already matches Grid average at B=216B=216 across all four families to within 0.010.01 accuracy, for instance 0.6530.653 against 0.6490.649 on location, reaching the same downstream accuracy with 27×27× fewer tokens. The baselines are less stable at the extremes, and MedPruner-DINS drops to near chance on location at B=64B=64. Figure 7 repeats the analysis with intrinsic probing R2R^2 on the two Merlin encoders, SuPreM and SegVol, over the size, density, and location families. The same effects appear on a second encoder family under a different readout: ORCA leads at nearly every budget, its advantage is again largest on location, and it saturates early. On SuPreM, ORCA at B=8B=8 already exceeds Grid average at B=216B=216 on density, 0.6870.687 against 0.6810.681, and its location score at B=27B=27 of 0.6100.610 far exceeds Grid average at B=216B=216 of 0.3950.395. Grid average is the weakest at very small budgets, collapsing to R2R^2 near zero on SegVol at B=4B=4. Figure 7: Probing budget curves on the Merlin encoders, SuPreM and SegVol. B.5 Held-out test generalization The probe readouts are trained on data, so we check that the method ranking is not specific to the official validation split. We reserve 2,0002,000 volumes from the CT-RATE training pool as an independent held-out test set, disjoint from both the probe-training volumes and the validation split, and re-evaluate the CT-RATE/COLIPRI comparison there. Table 13 reports the held-out scores together with their gap to validation. The ranking is preserved: ORCA leads every family at both budgets, most strikingly on location where it reaches 0.570.57 to 0.650.65 against 0.130.13 to 0.250.25 for the baselines, while disease is saturated and ties. The held-out and validation scores differ only slightly and unsystematically, at most about 0.050.05 and usually under 0.020.02, so the probes are not overfit to the validation split and ORCA’s advantage generalizes. method B disease size density location texture Grid average 27 0.848 ± 0.001 −-0.002 0.664 ± 0.006 −-0.011 0.797 ± 0.009 −-0.000 0.229 ± 0.032 −-0.009 0.682 ± 0.024 +0.010 216 0.849 ± 0.001 −-0.002 0.676 ± 0.011 −-0.008 0.858 ± 0.009 +0.001 0.246 ± 0.008 −-0.015 0.707 ± 0.021 −-0.018 MedPruner-DINS (Liu et al. 2026) 27 0.844 ± 0.001 −-0.004 0.625 ± 0.006 −-0.008 0.844 ± 0.004 +0.002 0.144 ± 0.005 −-0.025 0.679 ± 0.014 −-0.015 216 0.848 ± 0.001 −-0.002 0.670 ± 0.003 −-0.009 0.900 ± 0.001 −-0.001 0.199 ± 0.007 −-0.031 0.729 ± 0.012 −-0.025 ToMe (Bolya et al. 2023) 27 0.842 ± 0.001 −-0.002 0.601 ± 0.004 −-0.012 0.783 ± 0.012 −-0.001 0.134 ± 0.004 −-0.037 0.625 ± 0.010 −-0.019 216 0.846 ± 0.001 −-0.001 0.630 ± 0.005 −-0.011 0.835 ± 0.006 +0.000 0.182 ± 0.003 −-0.035 0.674 ± 0.019 −-0.051 ORCA 27 0.848 ± 0.001 −-0.002 0.689 ± 0.002 −-0.002 0.888 ± 0.004 −-0.004 0.573 ± 0.019 −-0.019 0.762 ± 0.014 −-0.002 216 0.849 ± 0.001 −-0.002 0.713 ± 0.007 +0.003 0.909 ± 0.003 −-0.003 0.648 ± 0.015 −-0.012 0.769 ± 0.011 −-0.023 Table 13: Held-out test on CT-RATE/COLIPRI. Each cell is the held-out score with the held-out minus validation difference in grey. Figure 8: Per-epoch VQA accuracy across the two training stages on CT-RATE/COLIPRI. Rows are token budgets, columns attribute families. B.6 Training curves Figure 8 plots per-epoch VQA accuracy across the two training stages, s1 projector warmup then s2 LoRA, for the three compressors at both budgets. Accuracy converges by the end of s1 and s2 does not improve it, often drifting down slightly from mild overfitting, so the best epoch numbers in Table 6 are converged rather than truncated. The ranking from the main results also holds across the whole trajectory, with ORCA’s margin clearest on location and density. B.7 Per-attribute and per-finding scores Tables 15–17 give the complete per-attribute results and Tables 18–19 the per-finding results. Per attribute, ORCA leads or ties the strongest baseline on nearly all attributes at both budgets, most clearly on location; per finding, the macro-AUROC differences across compressors are small, consistent with disease information being redundantly encoded and surviving aggressive compression. family attribute gold value probe VQA question what it captures size sz_heart_lung log (heart vol. / lung vol.) from organ masks R2R^2 tertile cardiothoracic ratio; cardiomegaly proxy sz_aorta_heart log (aorta vol. / heart vol.) R2R^2 tertile aorto-cardiac size balance sz_ivc_aorta log (IVC vol. / aorta vol.) R2R^2 tertile veno-arterial caliber aorta_diameter_m absolute aortic diameter in m from the aorta mask R2R^2 clinical aortic dilation / aneurysm heart_width_m absolute cardiac width in m R2R^2 tertile heart size density hu_aorta_calc fraction of aorta-wall voxels above a calcium HU cut R2R^2 clinical atherosclerotic calcium burden vert_median median vertebral-body HU R2R^2 tertile bone mineral density; osteoporosis lung_mean mean lung density in HU R2R^2 tertile diffuse lung disease / fluid hu_lung_haacon lung high-attenuation-area fraction R2R^2 tertile consolidation / ground-glass burden location lung_LR_logratio log (left-lung vol. / right-lung vol.) R2R^2 tertile left / right lung balance heart_x normalized left–right position of the heart centroid R2R^2 tertile mediastinal laterality / shift ivc_z normalized cranio–caudal position of the IVC centroid R2R^2 tertile supero-inferior landmark height texture lung_fo_Kurtosis lung intensity-histogram kurtosis R2R^2 tertile high-density lung texture; fibrosis vert_fo_Kurtosis vertebral-body histogram kurtosis R2R^2 tertile trabecular bone texture lung_Perc15 lung 15th-percentile HU R2R^2 clinical emphysema / low-attenuation disease 18 CT-RATE / 30 Merlin binary abnormality labels from radiology reports AUROC report generation validation anchor; report-derived, not an image measurement Table 14: Attribute definitions and task mappings in the measurement VQA benchmark. Texture attributes are first-order statistics extracted using PyRadiomics (van Griethuysen et al. 2017), where fo denotes firstorder. For Merlin, the texture family is omitted because these attributes are not defined for portal-venous contrast-enhanced CT, while absolute organ HU measurements are replaced with contrast-robust inter-organ density differences. family attribute Grid average Slice pooling DivPrune MedPruner-DINS ToMe MedRegion-CT ORCA size sz_heart_lung 0.876 ± 0.003 0.844 ± 0.009 0.867 ± 0.003 0.850 ± 0.005 0.815 ± 0.007 — 0.887 ± 0.004 sz_aorta_heart 0.581 ± 0.016 0.547 ± 0.010 0.565 ± 0.017 0.546 ± 0.010 0.518 ± 0.007 — 0.615 ± 0.011 sz_ivc_aorta 0.529 ± 0.002 0.497 ± 0.011 0.499 ± 0.006 0.475 ± 0.009 0.460 ± 0.009 — 0.541 ± 0.008 aorta_diameter_m 0.644 ± 0.009 0.636 ± 0.009 0.623 ± 0.015 0.619 ± 0.004 0.617 ± 0.006 — 0.654 ± 0.002 heart_width_m 0.759 ± 0.002 0.730 ± 0.008 0.707 ± 0.004 0.697 ± 0.005 0.686 ± 0.006 — 0.770 ± 0.006 density hu_aorta_calc 0.838 ± 0.002 0.807 ± 0.012 0.831 ± 0.004 0.824 ± 0.005 0.807 ± 0.005 — 0.836 ± 0.003 vert_median 0.540 ± 0.016 0.387 ± 0.003 0.719 ± 0.015 0.726 ± 0.017 0.566 ± 0.039 — 0.859 ± 0.023 lung_mean 0.930 ± 0.002 0.909 ± 0.006 0.941 ± 0.003 0.933 ± 0.003 0.893 ± 0.005 — 0.957 ± 0.001 hu_lung_haacon 0.920 ± 0.004 0.908 ± 0.004 0.928 ± 0.001 0.919 ± 0.004 0.899 ± 0.004 — 0.935 ± 0.002 location lung_LR_logratio 0.234 ± 0.059 0.081 ± 0.007 0.247 ± 0.024 0.151 ± 0.005 0.142 ± 0.004 — 0.684 ± 0.029 heart_x 0.239 ± 0.016 0.113 ± 0.004 0.128 ± 0.031 0.114 ± 0.004 0.139 ± 0.013 — 0.629 ± 0.028 ivc_z 0.326 ± 0.009 0.330 ± 0.014 0.311 ± 0.015 0.268 ± 0.006 0.242 ± 0.005 — 0.568 ± 0.009 texture lung_firstorder_Kurtosis 0.918 ± 0.003 0.903 ± 0.005 0.922 ± 0.004 0.914 ± 0.002 0.890 ± 0.002 — 0.932 ± 0.003 vert_firstorder_Kurtosis 0.295 ± 0.011 0.220 ± 0.008 0.434 ± 0.037 0.339 ± 0.074 0.261 ± 0.009 — 0.502 ± 0.037 lung_Perc15 0.845 ± 0.013 0.819 ± 0.011 0.876 ± 0.008 0.872 ± 0.006 0.797 ± 0.009 — 0.899 ± 0.004 Table 15: Per-attribute probing on CT-RATE/COLIPRI at B=27B=27, with R2R^2 for the regression families. Slice pooling uses its fixed slice count; MedRegion-CT is not applicable at B=27B=27. family attribute Grid average Slice pooling DivPrune MedPruner-DINS ToMe MedRegion-CT ORCA size sz_heart_lung 0.888 ± 0.005 0.844 ± 0.009 0.891 ± 0.005 0.871 ± 0.004 0.833 ± 0.003 0.905 ± 0.008 0.899 ± 0.001 sz_aorta_heart 0.606 ± 0.003 0.547 ± 0.010 0.622 ± 0.005 0.627 ± 0.007 0.543 ± 0.016 0.642 ± 0.013 0.649 ± 0.014 sz_ivc_aorta 0.525 ± 0.008 0.497 ± 0.011 0.534 ± 0.006 0.519 ± 0.003 0.483 ± 0.014 0.537 ± 0.002 0.558 ± 0.018 aorta_diameter_m 0.653 ± 0.006 0.636 ± 0.009 0.644 ± 0.009 0.646 ± 0.005 0.638 ± 0.004 0.646 ± 0.004 0.675 ± 0.011 heart_width_m 0.755 ± 0.002 0.730 ± 0.008 0.734 ± 0.004 0.717 ± 0.017 0.706 ± 0.008 0.744 ± 0.002 0.822 ± 0.001 density hu_aorta_calc 0.843 ± 0.009 0.807 ± 0.012 0.851 ± 0.007 0.863 ± 0.006 0.835 ± 0.003 0.853 ± 0.005 0.868 ± 0.005 vert_median 0.736 ± 0.020 0.387 ± 0.003 0.814 ± 0.010 0.871 ± 0.007 0.724 ± 0.015 0.870 ± 0.012 0.884 ± 0.014 lung_mean 0.952 ± 0.002 0.909 ± 0.006 0.954 ± 0.003 0.949 ± 0.002 0.917 ± 0.006 0.959 ± 0.004 0.962 ± 0.002 hu_lung_haacon 0.936 ± 0.003 0.908 ± 0.004 0.935 ± 0.006 0.933 ± 0.002 0.914 ± 0.004 0.933 ± 0.003 0.941 ± 0.001 location lung_LR_logratio 0.228 ± 0.043 0.081 ± 0.007 0.253 ± 0.029 0.207 ± 0.063 0.190 ± 0.063 0.274 ± 0.067 0.735 ± 0.019 heart_x 0.168 ± 0.009 0.113 ± 0.004 0.164 ± 0.014 0.169 ± 0.010 0.185 ± 0.013 0.166 ± 0.005 0.687 ± 0.019 ivc_z 0.368 ± 0.009 0.330 ± 0.014 0.368 ± 0.010 0.342 ± 0.012 0.291 ± 0.006 0.377 ± 0.006 0.611 ± 0.014 texture lung_firstorder_Kurtosis 0.929 ± 0.005 0.903 ± 0.005 0.931 ± 0.004 0.928 ± 0.004 0.908 ± 0.005 0.936 ± 0.004 0.933 ± 0.002 vert_firstorder_Kurtosis 0.473 ± 0.075 0.220 ± 0.008 0.433 ± 0.054 0.477 ± 0.054 0.398 ± 0.056 0.581 ± 0.026 0.600 ± 0.007 lung_Perc15 0.889 ± 0.001 0.819 ± 0.011 0.900 ± 0.004 0.905 ± 0.002 0.841 ± 0.009 0.916 ± 0.002 0.915 ± 0.003 Table 16: Per-attribute probing on CT-RATE/COLIPRI at B=216B=216, with R2R^2 for the regression families. MedRegion-CT is shown at its fixed N¯=549 N=549 and Slice pooling at its fixed slice count. family attribute SuPreM B=216B=216 SegVol B=256B=256 Grid average ToMe ORCA Grid average MedPruner-DINS ToMe ORCA size spleen_ml 0.717 ± 0.020 0.510 ± 0.023 0.802 ± 0.005 0.771 ± 0.022 0.811 ± 0.023 0.621 ± 0.145 0.828 ± 0.002 spleen_vert_ratio 0.651 ± 0.024 0.478 ± 0.017 0.722 ± 0.007 0.658 ± 0.030 0.731 ± 0.021 0.522 ± 0.105 0.733 ± 0.014 kidney_ml 0.405 ± 0.006 0.334 ± 0.012 0.500 ± 0.026 0.400 ± 0.065 0.538 ± 0.009 0.343 ± 0.030 0.552 ± 0.015 kidney_vert_ratio 0.395 ± 0.019 0.338 ± 0.007 0.501 ± 0.005 0.306 ± 0.040 0.511 ± 0.031 0.326 ± 0.047 0.551 ± 0.011 aorta_diameter_m 0.424 ± 0.004 0.434 ± 0.004 0.506 ± 0.027 0.494 ± 0.010 0.538 ± 0.038 0.409 ± 0.018 0.561 ± 0.010 aorta_vert_diam 0.181 ± 0.001 0.195 ± 0.004 0.256 ± 0.016 0.192 ± 0.010 0.279 ± 0.044 0.133 ± 0.007 0.322 ± 0.008 density vert_L1T12_median 0.707 ± 0.010 0.708 ± 0.008 0.794 ± 0.008 0.768 ± 0.002 0.763 ± 0.035 0.767 ± 0.021 0.807 ± 0.004 muscle_auto_median 0.914 ± 0.005 0.864 ± 0.009 0.954 ± 0.003 0.788 ± 0.020 0.775 ± 0.040 0.758 ± 0.004 0.821 ± 0.020 liver_spleen_diff 0.427 ± 0.035 0.373 ± 0.023 0.644 ± 0.038 0.336 ± 0.025 0.477 ± 0.027 0.398 ± 0.017 0.447 ± 0.037 location kidney_z 0.495 ± 0.009 0.457 ± 0.007 0.675 ± 0.015 0.581 ± 0.026 0.718 ± 0.018 0.640 ± 0.027 0.746 ± 0.017 kidney_lr_z_asym 0.305 ± 0.030 0.173 ± 0.005 0.689 ± 0.010 0.419 ± 0.019 0.753 ± 0.013 0.501 ± 0.091 0.789 ± 0.005 bladder_z 0.385 ± 0.022 0.257 ± 0.006 0.437 ± 0.036 0.166 ± 0.016 0.209 ± 0.023 0.208 ± 0.015 0.284 ± 0.075 Table 17: Per-attribute probing on Merlin, SuPreM at B=216B=216 and SegVol at B=256B=256, as mean R2±R^2± standard deviation over three seeds. MedPruner-DINS applies only to SegVol, since SuPreM’s windowed-attention backbone gives no per-token saliency; Merlin has no texture family. finding Avg. pool Slice pool DivPrune MedPruner ToMe MedRegion ORCA Medical material 0.932 ± 0.001 0.930 ± 0.003 0.932 ± 0.003 0.932 ± 0.002 0.929 ± 0.002 0.932 ± 0.001 0.932 ± 0.001 Arterial wall calcification 0.933 ± 0.001 0.931 ± 0.001 0.932 ± 0.001 0.933 ± 0.001 0.932 ± 0.001 0.930 ± 0.000 0.932 ± 0.000 Cardiomegaly 0.932 ± 0.002 0.932 ± 0.001 0.931 ± 0.001 0.930 ± 0.002 0.927 ± 0.001 0.933 ± 0.002 0.934 ± 0.001 Pericardial effusion 0.862 ± 0.006 0.861 ± 0.004 0.866 ± 0.005 0.864 ± 0.000 0.861 ± 0.007 0.860 ± 0.004 0.866 ± 0.005 Coronary artery wall calcification 0.935 ± 0.001 0.935 ± 0.001 0.936 ± 0.001 0.937 ± 0.000 0.937 ± 0.001 0.935 ± 0.002 0.936 ± 0.001 Hiatal hernia 0.718 ± 0.002 0.715 ± 0.001 0.717 ± 0.002 0.712 ± 0.004 0.713 ± 0.003 0.713 ± 0.002 0.716 ± 0.001 Lymphadenopathy 0.757 ± 0.005 0.755 ± 0.001 0.759 ± 0.003 0.762 ± 0.001 0.758 ± 0.002 0.756 ± 0.003 0.759 ± 0.002 Emphysema 0.813 ± 0.006 0.809 ± 0.004 0.811 ± 0.004 0.812 ± 0.005 0.809 ± 0.003 0.813 ± 0.004 0.814 ± 0.004 Atelectasis 0.818 ± 0.002 0.814 ± 0.001 0.816 ± 0.003 0.818 ± 0.006 0.813 ± 0.005 0.819 ± 0.004 0.822 ± 0.002 Lung nodule 0.743 ± 0.003 0.742 ± 0.003 0.743 ± 0.001 0.744 ± 0.005 0.740 ± 0.001 0.738 ± 0.001 0.743 ± 0.002 Lung opacity 0.873 ± 0.002 0.872 ± 0.001 0.872 ± 0.002 0.874 ± 0.001 0.869 ± 0.001 0.871 ± 0.000 0.872 ± 0.002 Pulmonary fibrotic sequela 0.752 ± 0.006 0.745 ± 0.008 0.752 ± 0.006 0.753 ± 0.005 0.747 ± 0.004 0.751 ± 0.002 0.751 ± 0.002 Pleural effusion 0.969 ± 0.003 0.969 ± 0.002 0.970 ± 0.002 0.967 ± 0.000 0.968 ± 0.002 0.969 ± 0.001 0.968 ± 0.001 Mosaic attenuation pattern 0.892 ± 0.002 0.893 ± 0.004 0.894 ± 0.000 0.891 ± 0.003 0.880 ± 0.003 0.890 ± 0.003 0.893 ± 0.002 Peribronchial thickening 0.814 ± 0.001 0.808 ± 0.002 0.811 ± 0.002 0.812 ± 0.002 0.805 ± 0.001 0.814 ± 0.008 0.817 ± 0.001 Consolidation 0.924 ± 0.002 0.925 ± 0.001 0.925 ± 0.002 0.922 ± 0.001 0.922 ± 0.002 0.924 ± 0.002 0.925 ± 0.002 Bronchiectasis 0.809 ± 0.005 0.800 ± 0.005 0.808 ± 0.005 0.807 ± 0.007 0.797 ± 0.002 0.809 ± 0.008 0.807 ± 0.012 Interlobular septal thickening 0.886 ± 0.002 0.886 ± 0.003 0.884 ± 0.001 0.886 ± 0.001 0.879 ± 0.003 0.885 ± 0.007 0.883 ± 0.002 Table 18: Per-finding probing on CT-RATE at B=216B=216, macro-AUROC. finding Avg. pool ORCA submucosal_edema 0.746 ± 0.005 0.760 ± 0.005 renal_hypodensities 0.722 ± 0.004 0.717 ± 0.004 aortic_valve_calcification 0.882 ± 0.004 0.880 ± 0.006 coronary_calcification 0.807 ± 0.006 0.827 ± 0.012 thrombosis 0.710 ± 0.004 0.719 ± 0.022 metastatic_disease 0.818 ± 0.005 0.779 ± 0.009 pancreatic_atrophy 0.808 ± 0.004 0.819 ± 0.005 renal_cyst 0.714 ± 0.001 0.717 ± 0.002 osteopenia 0.953 ± 0.001 0.944 ± 0.002 surgically_absent_gallbladder 0.723 ± 0.003 0.717 ± 0.005 atelectasis 0.792 ± 0.002 0.793 ± 0.014 abdominal_aortic_aneurysm 0.853 ± 0.015 0.897 ± 0.014 anasarca 0.969 ± 0.001 0.968 ± 0.001 hiatal_hernia 0.796 ± 0.001 0.799 ± 0.004 lymphadenopathy 0.625 ± 0.010 0.683 ± 0.031 prostatomegaly 0.764 ± 0.007 0.774 ± 0.009 biliary_ductal_dilation 0.782 ± 0.004 0.759 ± 0.016 cardiomegaly 0.817 ± 0.003 0.820 ± 0.010 splenomegaly 0.787 ± 0.006 0.882 ± 0.006 hepatomegaly 0.813 ± 0.016 0.808 ± 0.027 atherosclerosis 0.816 ± 0.001 0.825 ± 0.004 ascites 0.947 ± 0.002 0.939 ± 0.005 pleural_effusion 0.858 ± 0.003 0.864 ± 0.010 hepatic_steatosis 0.822 ± 0.009 0.807 ± 0.055 appendicitis 0.625 ± 0.038 0.600 ± 0.026 gallstones 0.720 ± 0.005 0.707 ± 0.003 hydronephrosis 0.705 ± 0.009 0.723 ± 0.004 bowel_obstruction 0.800 ± 0.016 0.839 ± 0.017 free_air 0.814 ± 0.002 0.832 ± 0.008 fracture 0.768 ± 0.002 0.757 ± 0.003 Table 19: Per-finding probing on Merlin/SuPreM at B=216B=216, macro-AUROC over all 30 findings.