Paper deep dive
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual evidence across adjacent slices. To understand how effectively a growing visual token budget improves performance, we perform scaling analyses on two 3D medical VQA benchmarks and find diminishing returns: cost keeps rising while accuracy saturates, and improving in-plane resolution is more effective than adding slices at comparable budgets. The budget should therefore be allocated more selectively rather than simply enlarged, yet most token compression methods are designed for 2D images or videos, where redundancy arises from spatial layout or temporal motion rather than from near-duplicate content along the depth axis. We present CARVE, a training-free framework that compresses visual tokens prior to LLM inference and casts token reduction as budget-constrained 2.5D allocation. CARVE partitions the depth axis into coherent windows and allocates tokens non-uniformly according to normalized cross-slice evidence. Under a shared budget, CARVE builds spatial anchors on representative slices and retrieves locally varying evidence from the full volume, then merges remaining eligible tokens into nearby anchors within each window. Removing roughly 80% of the visual tokens on Hulu-Med-7B, CARVE leads all compression baselines on every AMOS-MM report-generation metric, with 6.2 points higher retention of full-token quality than the strongest baseline, and preserves 98.1% of full-token performance across three VQA benchmarks.
Tags
Links
- Source: https://arxiv.org/abs/2608.04515v1
- Canonical: https://arxiv.org/abs/2608.04515v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
44,052 characters extracted from source content.
Expand or collapse full text
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding Zhenyu Yi1, Qiang Hu2, Zhenhao Li1, Jiaxuan Zhao1, Yusong Sun1, Lichi Zhang1 Corresponding author. Abstract Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation produces thousands of visual tokens that burden the LLM backbone, many of which capture overlapping visual evidence across adjacent slices. To understand how effectively a growing visual token budget improves performance, we perform scaling analyses on two 3D medical VQA benchmarks and find diminishing returns: cost keeps rising while accuracy saturates, and improving in-plane resolution is more effective than adding slices at comparable budgets. The budget should therefore be allocated more selectively rather than simply enlarged, yet most token compression methods are designed for 2D images or videos, where redundancy arises from spatial layout or temporal motion rather than from near-duplicate content along the depth axis. We present CARVE, a training-free framework that compresses visual tokens prior to LLM inference and casts token reduction as budget-constrained 2.5D allocation. CARVE partitions the depth axis into coherent windows and allocates tokens non-uniformly according to normalized cross-slice evidence. Under a shared budget, CARVE builds spatial anchors on representative slices and retrieves locally varying evidence from the full volume, then merges remaining eligible tokens into nearby anchors within each window. Removing roughly 80% of the visual tokens on Hulu-Med-7B, CARVE leads all compression baselines on every AMOS-M report-generation metric, with 6.2 points higher retention of full-token quality than the strongest baseline, and preserves 98.1% of full-token performance across three VQA benchmarks. Introduction Medical multimodal large language models (MLLMs) are increasingly applied to whole 3D volumes. Rather than learning native volumetric representations (Wu et al. 2024; Bai et al. 2024), most recent open-source systems reuse a mature 2D vision stack over ordered axial images (Jiang et al. 2025a, b; Sellergren et al. 2025): each slice is encoded independently, the patch features are concatenated, and cross-slice modeling is deferred to the projector and LLM, optionally with 3D-aware position indices (Wang et al. 2024). We target these slice-based models, which inherit strong 2D pretraining and support variable slice counts but realize a volume primarily by scaling the visual sequence: a moderate 32×512×51232\!×\!512\!×\!512 input already produces ∼ 10,367 visual tokens on Hulu-Med-7B, while cross-slice relations remain unresolved inside the 2D encoder. Along the depth axis, neighboring CT slices share broad anatomy, yet small structures, boundaries, or abnormalities may change over only a few slices. The sequence reaching the expensive LLM is therefore both long and structurally redundant. What must survive compression also depends on the task. A VQA query may target coarse anatomy, a localized attribute, or a cross-slice relation, while 3D report generation must summarize spatially distributed findings rather than only one prompted region (Hamamci et al. 2024; Li et al. 2025; Yu et al. 2025). Compression must consequently maintain a spatial scaffold for recurring context while preserving a recall path for localized evidence anywhere in the volume. Figure 1: Quality–latency trade-off on AMOS-M report generation with Hulu-Med-7B at a 20% target keep ratio. Labels give the retained visual token count (#Tok) per volume; quality averages relative retention over the four report metrics of Table 1. CARVE establishes the strongest quality–efficiency frontier among compressed methods. Figure 2: Scaling analysis of slice-based 3D medical VLMs on AMOS-M VQA. (a) Accuracy across slice-resolution combinations, with bubble size denoting slice count; in-plane resolution helps more than added slices at comparable budgets. (b) Accuracy against latency, peak memory, and estimated TFLOPs. (c) Adjacent-slice feature correlation at aligned local positions and in global slice means. Those requirements concern how a visual budget is spent, not only how large it is. We first ask: Do additional 3D visual tokens contribute equally, wherever they are allocated? Figure 2 shows they do not: on AMOS-M, accuracy saturates while latency, memory, and FLOPs keep rising, in-plane resolution beats extra slices at matched token counts, and adjacent slices stay highly similar in feature space. The bottleneck is thus not merely an oversized sequence, but a budget distributed without regard to the depth/in-plane asymmetry of the volume. Token compression has so far been developed for other input types. Image methods rank or merge tokens over a flat 2D layout (Yang et al. 2025; Dhouib et al. 2025; Deng et al. 2026; Zhang et al. 2026; Dong et al. 2025), video methods exploit temporal redundancy (Shen et al. 2026; Huang et al. 2025b), and geometry-aware methods operate on point or voxel representations (Li et al. 2026c, a; Huang et al. 2025a). None of these inputs presents an anisotropic depth axis that repeatedly re-images the same anatomy, so their objectives were never required to divide one budget between two axes of unequal value, and they transfer to volumes only partially. In the medical setting, MedPruner and other slice-selection methods (Liu et al. 2026b; Lian et al. 2026; Chen et al. 2025) do exploit volumetric redundancy, filtering slices before retaining tokens within the survivors; because they act at whole-slice granularity, however, evidence on a discarded plane cannot be recovered. Guided by these observations, we cast 3D medical token compression as anisotropic 2.5D budget allocation and present CARVE (Cross-slice Anisotropic Reallocation of Visual Evidence). CARVE partitions the depth axis into feature-continuous windows, not to discard their non-representative slices, but to define where budgets are allocated and tokens can be folded safely. Under one target budget, intra-slice selection constructs multi-granular spatial anchors on representative slices, while global inter-slice retrieval searches the complete volume for locally changing evidence missed by those anchors; eligible remaining tokens are then folded into nearby anchors within the same window, whereas retrieved tokens remain independent. CARVE is training-free and operates pre-LLM, requiring no modification to the underlying medical MLLM. In summary, our contributions are threefold: 1. Empirical analysis of 3D token scaling. Across AMOS-M and M3D-VQA, in-plane resolution outperforms adding slices at matched token counts and neighboring slices stay highly redundant in feature space, motivating non-uniform 2.5D allocation rather than uniform scaling. 2. A training-free pre-LLM reallocation framework. We propose CARVE, which couples anisotropic window allocation with intra-slice anchor selection and global inter-slice evidence retrieval, followed by window-local token folding before the LLM. 3. Strong quality-efficiency trade-offs. CARVE leads all compression baselines on every AMOS-M report-generation metric and preserves near-full aggregate VQA performance at approximately 19% retained tokens, giving the best quality–efficiency frontier among compressed methods (Fig. 1). Related Work 3D Medical Vision-Language Models Medical VLMs have progressed from 2D biomedical assistants such as LLaVA-Med (Li et al. 2023) toward volume-capable systems following two representation routes. Native volumetric models, including RadFM, M3D, Merlin, and CT-RATE foundation models (Wu et al. 2024; Bai et al. 2024; Blankemeier et al. 2024; Hamamci et al. 2026b), learn 3D representations directly. Slice-based systems instead reuse 2D vision backbones over ordered axial images (Jiang et al. 2025a, b; Sellergren et al. 2025); their features are assembled downstream by the projector and LLM, sometimes with multimodal positional indexing such as M-RoPE (Wang et al. 2024). This route inherits strong pretrained encoders and naturally supports variable slice sequences, but its token count grows along both depth and in-plane axes while cross-slice interaction is deferred beyond the encoder. Volume-level tasks expose complementary evidence requirements. CT2Rep and 3D brain-CT MLLMs generate reports from complete scans (Hamamci et al. 2024; Li et al. 2025), while MedFrameQA evaluates reasoning across multiple medical images (Yu et al. 2025). Dual-stream MIL and cross-view alignment further couple global predictions with localized evidence for 3D diagnosis (Yi et al. 2026; Li et al. 2026b). These works motivate maintaining both broad anatomical support and sparse local findings, but address modeling, supervision, or evaluation rather than drop-in sequence reduction. Efficiency-oriented systems instead prune consecutive slices (Jiang et al. 2025b) or build the volume model itself for lower-cost interpretation, either learning instruction-conditioned token budgets jointly with the architecture during training (Fang et al. 2026) or redesigning the encoder, tokenizer, or projector (Lee et al. 2026; Xin et al. 2025; Hamamci et al. 2026a). All of these change the model; CARVE addresses the complementary setting of compressing a frozen slice-based model after its encoder, without task-specific training. Visual Token Compression LLaVA-PruMerge and VisionZip retain salient tokens while merging redundancy (Shang et al. 2024; Yang et al. 2025); PACT combines pruning with bounded clustering at an early LLM layer (Dhouib et al. 2025). Training-free ranking methods use focal or hierarchical attention (Jiang et al. 2024; Liu et al. 2026a), feature diversity or graph structure (Alvar et al. 2025; Jiang et al. 2025c), and saliency-coverage or multimodal-coverage objectives (Deng et al. 2026; Dong et al. 2025). Closer to volumetric compression, video methods exploit redundancy across ordered inputs and allocate budgets across segments (Shen et al. 2026; Huang et al. 2025b; Hyun et al. 2025; Ma et al. 2025) and geometry-aware 3D methods use spatial sampling or voxel compression (Li et al. 2026c, a; Huang et al. 2025a). Yet a medical slice axis is neither ordinary time nor a generic point/voxel set: it repeatedly observes aligned anatomy with sparse embedded changes, so none of these objectives determines how one budget should be divided between the depth and in-plane axes. MedPruner is the closest training-free medical baseline: it filters slices before adaptively retaining tokens within the selected planes (Liu et al. 2026b); trained 2D-encoder pruning and systematic slice-selection studies share this granularity (Lian et al. 2026; Chen et al. 2025). CARVE differs on exactly this point: rather than committing to a subset of planes, it reallocates tokens across the anisotropic volume and keeps every slice eligible for retrieval. Placement also matters: ToMe and EViT alter token flow inside the vision encoder (Bolya et al. 2022; Liang et al. 2022) and FastV prunes after visual tokens enter the LLM (Chen et al. 2024), whereas CARVE operates training-free at the post-encoder, pre-projector interface, coupling depth-windowed allocation, intra-slice anchors, and global inter-slice retrieval before LLM computation. Method Figure 3: Overview of CARVE. Stage 1 partitions slice features into adaptive depth windows. Stage 2 allocates the budget anisotropically across windows and between the two selection roles. Stage 3 couples intra-slice anchor selection on representative slices with global inter-slice retrieval of dispersed deviations from the full volume. Stage 4 folds eligible remaining tokens into same-window anchors; retained tokens keep their positions before the frozen projector and LLM. We present CARVE, a training-free token compression framework that performs anisotropical budget allocation for slice-based 3D medical MLLMs. Repeated structures across adjacent slices can be represented economically by anchors on a few representative slices, whereas localized deviations may appear anywhere along the depth axis and require evidence retrieval from the full volume. A single flat ranking cannot fill both roles at once, since it may repeatedly select correlated tokens while sacrificing spatial coverage of discarded regions. CARVE therefore divides the target budget between intra-slice anchor selection and global inter-slice evidence retrieval, using adaptive depth windows to coordinate allocation and local token folding between them (Fig. 3). Problem Formulation Consider a 3D medical volume with T slices. A 2D vision encoder produces N features ℱ=jj=1NF=\f_j\_j=1^N with coordinates j=(zj,uj,vj)c_j=(z_j,u_j,v_j). A frozen multimodal projector P would ordinarily map them to LLM visual tokens j=P(j)x_j=P(f_j). Inserted between the encoder and P, CARVE compresses ℱF under a target budget K=⌊rN⌋K= rN for target keep ratio r∈(0,1]r∈(0,1]. It constructs complementary sets intraS_intra of spatial anchors and interS_inter of independently kept cross-slice evidence; we call their members anchor tokens and retrieved tokens. Cross-Slice Profiling and Adaptive Windowing CARVE profiles depth-wise feature change at two scales: a slice-level signal for coherent depth neighborhoods, and a token-level residual for local deviations within them. Inter-slice change. For each slice t, we mean-pool and ℓ2 _2-normalize its raw encoder features into ¯t∈ℝd f_t ^d, and define gt=1−cos(¯t,¯t+1),t=1,…,T−1.g_t=1- ( f_t,\; f_t+1), t=1,…,T-1. (1) Large gtg_t marks a feature discontinuity along the slice axis; low-change runs indicate where broad structure is repeatedly encoded across neighboring slices. Slice-normalized local residual. To separate slice-wide change from localized residuals, we compare each token with aligned neighbors and subtract its slice median: rj r_j =1−1|z(j)|∑j′∈z(j)cos(j,j′), =1- 1|N_z(j)| _j _z(j) (f_j,f_j ), (2) r~j r_j =max(0,rj−mediank∈slice(j)rk), = \! (0,\ r_j- *median_k (j)r_k ), where z(j)N_z(j) contains available tokens at the same in-plane coordinate on slices zj±1z_j\!±\!1. Median subtraction removes slice-wide shifts, leaving r~j r_j as a local cross-slice deviation score for evidence retrieval. Adaptive windows. CARVE places a candidate window boundary after slice t when gt>μg+τσg,t=1,…,T−1,g_t> _g+τ _g, t=1,…,T-1, (3) where μg _g and σg _g are the mean and standard deviation of the inter-slice change scores, and τ controls boundary sensitivity. These boundaries partition the volume into windows ss=1S\W_s\_s=1^S that share an allocation and folding domain. Importantly, a window is not a slice-pruning unit: every slice remains eligible for global evidence retrieval. Anisotropic Budget Allocation The target budget must preserve a spatial scaffold without closing the recall path across depth. We therefore reserve a global inter-slice budget and assign the remainder to intra-slice anchors: Kinter=⌊ρK⌋,Kintra=K−Kinter.K_inter= ρ K , K_intra=K-K_inter. (4) Here ρ∈[0,1]ρ∈[0,1] is the fraction reserved for cross-slice recall, and KinterK_inter is an upper budget for interS_inter. We then distribute KintraK_intra non-uniformly across windows. Windows with stronger depth variation, larger local residuals, or longer spans receive more anchors. To compare these signals on a common scale, let (vs)N(v_s) denote min–max normalization of statistic vsv_s over the current volume; we use the fixed score ws w_s =(g¯s)+(r¯s)+12(log(1+ns)), =N( g_s)+N( r_s)+ 12N( (1+n_s)), (5) Ks K_s =LRs[Kintrasoftmaxs(ws)], =LR_s\! [K_intra\,softmax_s(w_s) ], Here g¯s g_s averages the change scores within sW_s, r¯s r_s averages r~j r_j over its tokens, ns=|s|n_s=|W_s|, and LRLR denotes largest-remainder integerization, so that ∑sKs=Kintra _sK_s=K_intra. Intra-Slice Anchor Selection The intra-slice path converts each window budget into a compact spatial scaffold. Within sW_s, it ranks slices by the mean cosine similarity of their pooled features to those of the other slices and distributes KsK_s across the top-ranked representatives. Their number is ms=clamp(round(Ks/Kslice),1,min(ns,Ks))m_s=clamp(round(K_s/K_slice),1, (n_s,K_s)) for a target anchor density KsliceK_slice, and KsK_s is split as evenly as possible across them. This avoids spending anchors on many near-duplicate slices. Multi-granular in-plane partition. On each representative slice, CARVE recursively refines the H×WH× W token grid. Let aja_j be the attention saliency of token j, obtained by averaging the last-layer encoder self-attention Aqj(L,h)A^(L,h)_qj over all query positions q and heads h; this reuses an existing attention artifact and needs no dedicated class token. Refinement favors blocks that are both internally heterogeneous and encoder-salient: π(ℬ)=(1−1|ℬ|∑j∈ℬcos(j,¯ℬ))+(maxj∈ℬaj),π(B)=N\! (1- 1|B| _j (f_j, f_B) )+N\! ( _j a_j ), (6) where ¯ℬ f_B is the mean feature of block ℬB; the heterogeneity and attention terms are normalized independently over the current slice. At each step, the highest-priority block is split until the allocated budget is reached; the highest-attention token in each final block becomes its anchor. The resulting union intraS_intra preserves multi-granular in-plane support. We realize this budgeted refinement with a standard best-first quadtree (Hyun et al. 2025). Global Inter-Slice Evidence Retrieval Representative-slice anchors efficiently summarize recurring structure, but by construction cannot guarantee that localized deviations on other slices survive. The inter-slice path therefore searches all non-anchor tokens using the residual in Eq. 2. Hard residual ranking. Let inter=1,…,N∖intraP_inter=\1,…,N\ _intra be the full-volume retrieval pool and sjs_j the volume-normalized score derived from r~j r_j. We first isolate the high-residual portion of this pool: Mcand M_cand =max(Kinter,⌈0.25|inter|⌉), = \! (K_inter, 0.25|P_inter| ), (7) (0) ^(0) =j∈inter:rank↓(sj)≤Mcand. = \j _inter:\ rank_ (s_j)≤ M_cand \. This keeps the strongest quartile of deviations, and never fewer than KinterK_inter candidates, for the subsequent spatial de-duplication. Greedy selection with 3D-NMS. High residual scores can cluster around the same physical region. To prevent redundant retrieval, we define the spacing-aware suppression neighborhood of candidate i as 3D(ℓ)(i)=j∈(ℓ−1):dsp(j,i)≤rnms,N_3D^( )(i)= \j ^( -1):d_sp(j,i)≤ r_nms \, (8) where dspd_sp is the Chebyshev distance in the volume’s physical coordinate system with each axis measured in units of its effective voxel pitch, so that a single radius rnmsr_nms induces different depth and in-plane extents. Starting with inter(0)=∅S_inter^(0)= , each round selects the highest-scoring admissible token and suppresses its neighborhood: jℓ∗ j_ ^* =argmaxj∈(ℓ−1)sj, = _j ^( -1)s_j, (9) inter(ℓ) _inter^( ) =inter(ℓ−1)∪jℓ∗, =S_inter^( -1)∪\j_ ^*\, (ℓ) ^( ) =(ℓ−1)∖3D(ℓ)(jℓ∗). =C^( -1) _3D^( )(j_ ^*). The score sjs_j remains fixed while the admissible set changes. The procedure stops at KinterK_inter tokens or an empty candidate set. Its output interS_inter thus favors strong but spatially dispersed deviations; these retrieved tokens retain their coordinates and bypass folding. Window-Local Token Folding The selected anchors also provide a structured destination for information that would otherwise be discarded. After excluding intraS_intra and interS_inter, each remaining token is assigned to the closest intra-slice anchor of its own adaptive window. This restriction avoids folding tokens across the depth discontinuities identified in Eq. 3. Tokens whose closest anchor is too far away are dropped, while retrieved tokens are excluded from folding and remain independent. For an intra-slice anchor c, the folding set (c)O(c) is then the tokens whose closest anchor is c and whose in-plane Manhattan distance to c is at most a neighborhood radius κ, and CARVE computes ~c f_c =∑b∈(c)αbcb, = _b (c) _bcf_b, (10) αbc _bc =exp(cos(b,c)/τm)∑b′∈(c)exp(cos(b′,c)/τm). = ( (f_b,f_c)/ _m) _b (c) ( (f_b ,f_c)/ _m). where τm _m is the folding temperature. CARVE finally combines the updated anchor tokens and retrieved tokens as one compressed visual sequence, preserving their respective 3D position indices and the backbone’s spatial order. The sequence is projected by P, concatenated with tokenized text embeddings, and passed to the LLM. Quality Efficiency Method R-1 R-L MTR BERT Rel. Keep Lat.↓ Δ ↓ Full 43.71 41.28 29.08 28.70 100.0 100.0% 9.21 3.79 VisionZip 36.31 34.25 22.43 20.01 78.22 20.0% 6.90 2.84 DivPrune 36.79 34.70 23.13 21.75 80.89 20.0% 7.22 1.35 FastVID 31.75 30.05 19.56 15.60 66.76 20.0% 6.70 2.84 MMTok 34.65 32.87 21.63 19.59 75.38 20.0% 13.04 1.63 MedPruner 36.89 35.08 21.80 21.94 80.20 21.4% 5.88 2.93 CARVE 39.04 36.84 24.95 24.09 87.07 19.3% 6.10 2.89 Table 1: AMOS-M report generation on Hulu-Med-7B at target r=0.2r=0.2. Rel. is mean retention over the four quality metrics relative to Full; Keep is the realized keep ratio (%), while latency and Δ use seconds and GB. Figure 4: AMOS-M closed-ended VQA by task type on Hulu-Med-7B at target r=0.2r=0.2. Bars are relative to Full; gray labels give absolute Full accuracy. Experiments Experimental Setting We evaluate CARVE on three benchmarks spanning closed-ended VQA, open-ended VQA, and report generation: 3D-RAD (Gai et al. 2026), M3D-VQA (Bai et al. 2024), and AMOS-M VQA/Report (Ji et al. 2022). Unless otherwise stated, the main comparisons use the medical Hulu-Med-7B at a target keep ratio of r=0.2r=0.2, which reserves a fraction ρ=0.25ρ=0.25 of the budget for inter-slice retrieval and targets Kslice=64K_slice=64 anchors per retained slice. The transfer study adds the general-domain Qwen3-VL. We compare against training-free image-token compression methods VisionZip (Yang et al. 2025), DivPrune (Alvar et al. 2025), and MMTok (Dong et al. 2025); the video token-compression method FastVID (Shen et al. 2026); and MedPruner (Liu et al. 2026b), the prior method tailored to medical volumes. Main Results AMOS-M report generation evaluation. CARVE ranks first among compressed methods on all four report metrics and their aggregate, while using the smallest realized keep ratio (19.3%). Report generation separates compression strategies sharply: baseline aggregate retention spans 66.76 to 80.89, and CARVE reaches 87.07—6.18 points above the best baseline aggregate and 2.15 points above the best baseline ROUGE-1. Latency further shows that selector cost is a separate axis: CARVE runs at 6.10 s, second only to MedPruner and below the 9.21 s Full reference, whereas MMTok’s 13.04 s exceeds applying no compression at all (Fig. 1). 3D-RAD and M3D-VQA VQA evaluation. Quality Efficiency Method ACC B-1 R-L MTR BERT Rel. Keep Lat.↓ Δ ↓ !12 3D-RAD (closed + open) Full 79.08 25.26 33.91 20.52 55.40 100.0 100.0% 0.62 0.97 VisionZip 78.31 23.16 32.00 19.22 53.58 95.09 20.0% 0.36 0.40 DivPrune 78.00 22.88 31.48 19.26 53.53 94.51 20.0% 0.39 0.34 FastVID 77.85 22.93 31.49 19.45 53.35 94.63 19.6% 0.34 0.40 MMTok 77.57 22.39 30.82 18.76 53.13 92.99 20.0% 0.43 0.34 MedPruner 78.57 22.93 31.32 19.18 53.53 94.52 21.3% 0.36 0.43 CARVE 78.73 23.85 32.56 20.02 53.90 96.97 19.7% 0.31 0.42 !12 M3D-VQA (closed + open) Full 77.65 45.23 47.40 31.22 56.85 100.0 100.0% 2.36 3.23 VisionZip 76.85 44.40 46.49 30.24 56.50 98.29 20.0% 1.10 2.31 DivPrune 76.79 44.56 46.68 30.56 56.50 98.63 20.0% 1.25 1.15 FastVID 76.78 44.46 46.48 30.22 56.61 98.32 20.0% 1.08 2.31 MMTok 77.01 44.63 46.76 30.59 56.51 98.78 20.0% 7.26 1.32 MedPruner 76.90 44.45 46.57 30.33 56.49 98.41 19.7% 1.05 2.39 CARVE 76.97 44.83 46.98 30.74 56.60 99.08 19.2% 1.09 2.35 Table 2: Main VQA comparison on Hulu-Med-7B at target r=0.2r=0.2. Rel. is mean retention over the five quality metrics relative to Full; Keep is the realized keep ratio (%), while latency and Δ use seconds and GB. At only 19.7% and 19.2% realized keep ratios (Table 3D-RAD and M3D-VQA VQA evaluation.), CARVE maintains 96.97% and 99.08% aggregate performance on 3D-RAD and M3D-VQA, respectively. Closed-ended accuracy is near-saturated at this budget—all compressed methods lie within 1.16 and 0.23 ACC points on the two benchmarks—so the discriminative signal lies in the open-ended metrics, where CARVE ranks first on seven of the eight scores across both benchmarks; the sole exception is M3D-VQA BERT (56.60 vs. 56.61). Separation is wider on 3D-RAD, whose 2562256^2 inputs leave a more binding absolute budget than the 5122512^2 inputs of M3D-VQA at the same ratio. AMOS-M VQA task-type evaluation. Across three perception and three reasoning task types (Fig. 4), CARVE achieves the highest overall relative score among compressed methods and does not regress on either group, indicating that its gains are not confined to coarse perception. Backbone transfer. To test whether CARVE depends on a particular projector–LLM stack, we evaluate three additional backbones on six tracks at the same target keep ratio. Figure 5: Cross-backbone transfer at target r=0.2r=0.2. (a) Mean retention over six evaluation tracks relative to Full. (b) Mean rank over all 33 backbones × 66 tracks among compressed methods (lower is better). CARVE achieves the highest aggregate retention on every backbone and the best overall mean rank. Figure 6: Quality and latency versus realized keep ratio on AMOS-M VQA and Report. Dashed lines mark the uncompressed Full reference. CARVE is first on all three (Fig. 5): 94.53% on Hulu-Med-4B, 100.64% and 100.88% on Qwen3-VL-4B/8B, with a mean rank of 2.42 across the 18 backbone–track combinations against 3.19 for the next method. Separation is widest on Hulu-Med-4B, the most binding setting, whereas the Qwen3-VL routes leave every method near the Full reference, indicating that 20% is not yet a binding budget for that stack. Ablation Studies We organize the ablation into four comparison groups: naive sampling; depth partition and window allocation; intra/inter selection roles; and token folding. All variants use the same target budget K, preprocessing, decoder settings, and evaluation samples on AMOS-M, and paired rows isolate the intended factor within each group. VQA Report Control Variant Keep ACC Lat.↓ R-L Lat.↓ CARVE (Ours) 19.3% 75.50 1.11 36.84 6.10 Anchor: Naive sampling Uniform slice+grid 20.0% 73.46↓ 2.04 0.94 33.45↓ 3.39 5.75 A: Depth partition and window allocation Adaptive win., uniform bud. 19.3% 74.45↓ 1.05 1.10 35.18↓ 1.66 6.10 Uniform win., uniform bud. 19.6% 74.42↓ 1.08 1.11 34.68↓ 2.16 6.08 B: Intra/inter selection roles Intra only (ρ=0ρ=0) 20.3% 74.49↓ 1.01 1.13 35.15↓ 1.69 5.99 Intra only, w/o quadtree 19.1% 74.24↓ 1.26 1.36 34.67↓ 2.17 6.12 Inter only 19.3% 74.74↓ 0.76 1.13 34.92↓ 1.92 6.15 Inter only, w/o 3D-NMS 20.1% 74.20↓ 1.30 1.14 34.22↓ 2.62 5.82 C: Token folding Direct drop 19.3% 74.13↓ 1.37 1.29 33.44↓ 3.40 5.87 Global merge 19.3% 74.20↓ 1.30 1.38 35.33↓ 1.51 6.30 Local mean merge 19.3% 75.06↓ 0.44 1.11 35.26↓ 1.58 6.03 Table 3: AMOS-M ablation at target r=0.2r=0.2. Keep is the realized ratio (%); the remaining metrics are ACC, R-L, and latency (s). Subscripts give the drop relative to the full method, which is set in bold as the reference; all other rows are ablated versions of it, so no ranking among variants is implied. Every toggle costs quality on both tasks (Table 3). Naive Sampling has the largest ACC drop, while Dropping Folding has the largest ROUGE-L drop. Group A separates the allocation factors: adaptive windows with a uniform budget cost 1.66 ROUGE-L, and uniform windows a further 0.50, so the allocation matters more than the partition alone. In group B, neither branch suffices by itself (1.69 and 1.92 ROUGE-L). Ratio Change and Efficiency Analysis Figure 6 sweeps the target ratio over r∈0.2,0.4,0.6,0.8r∈\0.2,0.4,0.6,0.8\ and plots quality and latency against the realized keep ratio. Quality rises with the budget for every method, but the methods converge: the spread narrows from 2.51 to 1.40 ACC on VQA and from 6.80 to 0.82 ROUGE-L on report generation. CARVE’s advantage is concentrated where the budget binds—it leads by 0.48 ACC and 1.76 ROUGE-L at r≈0.2r≈0.2 and stays within 0.1 of the best baseline above that—while both tracks return to near-Full quality by r≈0.4r≈0.4. Latency exposes the cost side: CARVE remains below the Full reference on report generation at every ratio, whereas MMTok exceeds it throughout and DivPrune from r≈0.6r≈0.6. Qualitative Analysis Figure 7: Retained tokens on four consecutive slices of one AMOS-M case, with two key slices magnified. Purple and orange mark CARVE’s anchor and retrieved tokens, blue marks baseline kept tokens, and red circles indicate the queried finding. Figure 7 shows a case whose answer depends on a small finding that appears on adjacent slices. CARVE spreads anchors across all four slices, and in the two magnified planes its retrieved tokens concentrate on the queried lymph nodes, whose appearance fluctuates from one slice to the next; it is the only method that answers correctly. MedPruner spends almost its entire budget on a single slice and leaves the neighboring ones empty, following from its whole-slice granularity, whereas VisionZip and DivPrune distribute tokens evenly and cover the finding only incidentally. Conclusion We presented CARVE, a training-free token reallocation framework for slice-based 3D medical vision-language models that spends one budget anisotropically across the depth and in-plane axes before the frozen projector and LLM. At approximately 20% retained tokens, it leads all compressed methods on the AMOS-M report metrics and attains the highest aggregate retention across backbones, without training or backbone modification. Future work may combine CARVE with intra-ViT token merging or learnable budget allocation, extend it to multimodal imaging such as PET-CT. References S. R. Alvar, G. Singh, M. Akbari, and Y. Zhang (2025) Divprune: diversity-based visual token pruning for large multimodal models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 9392–9401. Cited by: Visual Token Compression, Experimental Setting. F. Bai, Y. Du, T. Huang, M. Q. Meng, and B. Zhao (2024) M3d: advancing 3d medical image analysis with multi-modal large language models. arXiv preprint arXiv:2404.00578. Cited by: Introduction, 3D Medical Vision-Language Models, Experimental Setting. L. Blankemeier, J. P. Cohen, A. Kumar, D. Van Veen, S. J. S. Gardezi, M. Paschali, Z. Chen, J. Delbrouck, E. Reis, C. Truyts, et al. (2024) Merlin: a vision language foundation model for 3d computed tomography. Research Square, p. rs–3. Cited by: 3D Medical Vision-Language Models. D. Bolya, C. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman (2022) Token merging: your vit but faster. arXiv preprint arXiv:2210.09461. Cited by: Visual Token Compression. H. Chen, R. Shukla, R. Wu, S. Yang, D. Duong-Tran, D. M. H. Nguyen, M. Niepert, C. Beeche, J. Gee, J. Duda, et al. (2025) Adapting vision-language models for 3d ct/mri understanding on pmbb via slice selection and explanation analysis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 2273–2282. Cited by: Introduction, Visual Token Compression. L. Chen, H. Zhao, T. Liu, S. Bai, J. Lin, C. Zhou, and B. Chang (2024) An image is worth 1/2 tokens after layer 2: plug-and-play inference acceleration for large vision-language models. In European Conference on Computer Vision, p. 19–35. Cited by: Visual Token Compression. J. Deng, W. Li, J. T. Zhou, and Y. He (2026) Scope: saliency-coverage oriented token pruning for efficient multimodel llms. Advances in Neural Information Processing Systems 38, p. 161527–161552. Cited by: Introduction, Visual Token Compression. M. Dhouib, D. Buscaldi, S. Vanier, and A. Shabou (2025) Pact: pruning and clustering-based token reduction for faster visual language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 14582–14592. Cited by: Introduction, Visual Token Compression. S. Dong, J. Hu, M. Zhang, M. Yin, Y. Fu, and Q. Qian (2025) Mmtok: multimodal coverage maximization for efficient inference of vlms. arXiv preprint arXiv:2508.18264. Cited by: Introduction, Visual Token Compression, Experimental Setting. C. Fang, H. Guo, Z. Jiang, C. He, X. Li, and M. Xu (2026) Photon: speedup volume understanding with efficient multimodal large language models. arXiv preprint arXiv:2603.25155. Cited by: 3D Medical Vision-Language Models. X. Gai, J. Liu, Y. Li, Z. Meng, J. Wu, and Z. Liu (2026) 3d-rad: a comprehensive 3d radiology med-vqa dataset with multi-temporal analysis and diverse diagnostic tasks. Advances in Neural Information Processing Systems 38. Cited by: Experimental Setting. I. E. Hamamci, S. Er, and B. Menze (2024) Ct2rep: automated radiology report generation for 3d medical imaging. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 476–486. Cited by: Introduction, 3D Medical Vision-Language Models. I. E. Hamamci, S. Er, S. Shit, H. Reynaud, D. Yang, P. Guo, M. Edgar, D. Xu, B. Kainz, and B. Menze (2026a) Better tokens for better 3d: advancing vision-language modeling in 3d medical imaging. Advances in Neural Information Processing Systems 38, p. 135074–135102. Cited by: 3D Medical Vision-Language Models. I. E. Hamamci, S. Er, C. Wang, F. Almas, A. G. Simsek, S. N. Esirgun, I. Dogan, O. F. Durugol, B. Hou, S. Shit, et al. (2026b) Generalist foundation models from a multimodal dataset for 3d computed tomography. Nature Biomedical Engineering, p. 1–19. Cited by: 3D Medical Vision-Language Models. H. Huang, F. Chen, W. Chai, C. Su, L. Xia, S. Jung, C. Yang, J. Hwang, M. Sun, and C. Kuo (2025a) Zero-shot 3d question answering via voxel-based dynamic token compression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19424–19434. Cited by: Introduction, Visual Token Compression. X. Huang, H. Zhou, and K. Han (2025b) Prunevid: visual token pruning for efficient video large language models. In Findings of the Association for Computational Linguistics: ACL 2025, p. 19959–19973. Cited by: Introduction, Visual Token Compression. J. Hyun, S. Hwang, S. H. Han, T. Kim, I. Lee, D. Wee, J. Lee, S. J. Kim, and M. Shim (2025) Multi-granular spatio-temporal token merging for training-free acceleration of video llms. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 23990–24000. Cited by: Visual Token Compression, Multi-granular in-plane partition.. Y. Ji, H. Bai, C. Ge, J. Yang, Y. Zhu, R. Zhang, Z. Li, L. Zhanng, W. Ma, X. Wan, et al. (2022) Amos: a large-scale abdominal multi-organ benchmark for versatile medical image segmentation. Advances in neural information processing systems 35, p. 36722–36732. Cited by: Experimental Setting. L. Jiang, W. Huang, T. Liu, Y. Zeng, J. Li, L. Cheng, and X. Xu (2024) Fopru: focal pruning for efficient large vision-language models. arXiv preprint arXiv:2411.14164. Cited by: Visual Token Compression. S. Jiang, Y. Wang, S. Song, T. Hu, C. Zhou, B. Pu, Y. Zhang, Z. Yang, Y. Feng, J. T. Zhou, et al. (2025a) Hulu-med: a transparent generalist model towards holistic medical vision-language understanding. arXiv preprint arXiv:2510.08668. Cited by: Introduction, 3D Medical Vision-Language Models. S. Jiang, Y. Wang, S. Song, Y. Zhang, Z. Meng, B. Lei, J. Wu, J. Sun, and Z. Liu (2025b) Omniv-med: scaling medical vision-language model for universal visual understanding. arXiv preprint arXiv:2504.14692. Cited by: Introduction, 3D Medical Vision-Language Models, 3D Medical Vision-Language Models. Y. Jiang, Q. Wu, W. Lin, W. Yu, and Y. Zhou (2025c) What kind of visual tokens do we need? training-free visual token pruning for multi-modal large language models from the perspective of graph. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 4075–4083. Cited by: Visual Token Compression. C. Lee, S. Park, C. I. Shin, W. H. Choi, H. J. Park, J. E. Lee, and J. C. Ye (2026) Read like a radiologist: efficient vision-language model for 3d medical imaging interpretation. Medical Image Analysis, p. 104077. Cited by: 3D Medical Vision-Language Models. C. Li, K. Chang, C. Yang, H. Wu, W. Chen, H. Bansal, L. Chen, Y. Yang, Y. Chen, S. Chen, et al. (2025) Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation. Nature Communications 16 (1), p. 2258. Cited by: Introduction, 3D Medical Vision-Language Models. C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao (2023) Llava-med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36, p. 28541–28564. Cited by: 3D Medical Vision-Language Models. H. Li, Z. Huang, J. Fu, N. Wang, and S. Liu (2026a) Geometry-guided 3d visual token pruning for video-language models. arXiv preprint arXiv:2604.18260. Cited by: Introduction, Visual Token Compression. S. Li, Z. Qiu, Z. Wang, B. Yun, Z. Yi, J. Xu, W. Zhang, Y. Xia, and L. Zhang (2026b) E-mrl: cross-view aligned evidence-driven multimodal reinforcement learning for reliable 3d tumor analysis. arXiv preprint arXiv:2606.23888. Cited by: 3D Medical Vision-Language Models. W. Li, K. Zhao, H. Jiang, E. Yang, Y. Su, and D. Zeng (2026c) SeGPruner: semantic-geometric visual token pruner for 3d question answering. arXiv preprint arXiv:2603.29437. Cited by: Introduction, Visual Token Compression. Y. Lian, Y. Xie, Y. Jiang, L. Wang, and H. Yu (2026) A data-efficient 3d medical vision-language model using only a 2d encoder. Scientific Reports 16 (1), p. 8809. Cited by: Introduction, Visual Token Compression. Y. Liang, C. Ge, Z. Tong, Y. Song, J. Wang, and P. Xie (2022) Not all patches are what you need: expediting vision transformers via token reorganizations. arXiv preprint arXiv:2202.07800. Cited by: Visual Token Compression. J. Liu, G. Zhu, and F. Du (2026a) Hiprune: training-free visual token pruning via hierarchical attention in vision-language models (student abstract). In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 41275–41277. Cited by: Visual Token Compression. S. Liu, Z. Ye, Y. Lin, C. Hu, W. Geng, X. Han, B. Ibragimov, Y. Zheng, and Y. Yuan (2026b) MedPruner: training-free hierarchical token pruning for efficient 3d medical image understanding in vision-language models. arXiv preprint arXiv:2603.11625. Cited by: Introduction, Visual Token Compression, Experimental Setting. J. Ma, Q. Zhang, M. Lu, Z. Wang, Q. Zhou, J. Song, and S. Zhang (2025) Mmg-vid: maximizing marginal gains at segment-level and token-level for efficient video llms. arXiv preprint arXiv:2508.21044. Cited by: Visual Token Compression. A. Sellergren, S. Kazemzadeh, T. Jaroensri, A. Kiraly, M. Traverse, T. Kohlberger, S. Xu, F. Jamil, C. Hughes, C. Lau, et al. (2025) Medgemma technical report. arXiv preprint arXiv:2507.05201. Cited by: Introduction, 3D Medical Vision-Language Models. Y. Shang, M. Cai, B. Xu, Y. J. Lee, and Y. Yan (2024) Llava-prumerge: adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388. Cited by: Visual Token Compression. L. Shen, G. Gong, T. He, Y. Zhang, P. Liu, S. Zhao, et al. (2026) Fastvid: dynamic density pruning for fast video large language models. Advances in Neural Information Processing Systems 38, p. 123553–123581. Cited by: Introduction, Visual Token Compression, Experimental Setting. P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al. (2024) Qwen2-vl: enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191. Cited by: Introduction, 3D Medical Vision-Language Models. C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie (2024) Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data, arxiv. arXiv preprint arXiv:2308.02463. Cited by: Introduction, 3D Medical Vision-Language Models. Y. Xin, G. C. Ates, K. Gong, and W. Shao (2025) Med3dvlm: an efficient vision-language model for 3d medical image analysis. IEEE Journal of Biomedical and Health Informatics. Cited by: 3D Medical Vision-Language Models. S. Yang, Y. Chen, Z. Tian, C. Wang, J. Li, B. Yu, and J. Jia (2025) Visionzip: longer is better but not necessary in vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19792–19802. Cited by: Introduction, Visual Token Compression, Experimental Setting. Z. Yi, Z. Song, Y. Sun, Z. Liu, M. Fei, Z. Li, J. Zhao, X. Han, and L. Zhang (2026) Brain-adapter: a dual-stream vision-language mil framework for comprehensive 3d ct diagnosis of acute intracranial pathologies. arXiv preprint arXiv:2606.23494. Cited by: 3D Medical Vision-Language Models. S. Yu, H. Wang, J. Wu, L. Luo, J. Wang, C. Xie, P. Rajpurkar, C. Yang, Y. Yang, K. Wang, et al. (2025) Medframeqa: a multi-image medical vqa benchmark for clinical reasoning. arXiv preprint arXiv:2505.16964. Cited by: Introduction, 3D Medical Vision-Language Models. Q. Zhang, M. Liu, L. Li, M. Lu, Y. Zhang, J. Pan, Q. She, and S. Zhang (2026) Beyond attention or similarity: maximizing conditional diversity for token pruning in mllms. Advances in Neural Information Processing Systems 38, p. 25438–25468. Cited by: Introduction.