Paper deep dive
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference
Lyuke Wang, Zhuo Li, Guangxu Zhu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/26/2026, 5:28:11 AM
Summary
The paper introduces VisCache, a training-free, plug-and-play framework for efficient long-context inference in Vision Large Language Models (VLLMs). It addresses the high memory and computational costs of visual Key-Value (KV) caches through two stages: (1) a lightweight scout model using Maximal Marginal Relevance (MMR) to filter temporal redundancy by selecting keyframes, and (2) PruneKV, a layer-aware compression algorithm that allocates budgets parabolically across layers and employs an asymmetric update mechanism to prune keys while fusing values. VisCache achieves up to 2.35x speedup and significant memory reduction with only 19-28% KV cache retention, outperforming existing baselines.
Entities (14)
Relation Signals (14)
VisCache → achievesspeedup → 2.35x
confidence 95% · achieving up to 2.35× speedup
VisCache → includescomponent → PruneKV
confidence 95% · we introduce PruneKV, a surgical KV compression algorithm... VisCache... consists of two synergistic stages.
VisCache → retainscache → 19-28%
confidence 95% · maintaining competitive performance with only 19–28% KV cache retention
VisCache → outperforms → PDrop
confidence 90% · VisCache consistently outperforms existing baselines... Table 1
VisCache → outperforms → Q-Frame
confidence 90% · VisCache consistently outperforms existing baselines... Table 1
VisCache → outperforms → PyramidKV
confidence 90% · VisCache consistently outperforms existing baselines... Table 1
VisCache → outperforms → FastV
confidence 90% · VisCache consistently outperforms existing baselines... Table 1
VisCache → usesalgorithm → MMR
confidence 90% · selecting a compact yet diverse subset of frames... via MMR
VisCache → usesmodel →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key-Value (KV) caches. Existing KV compression methods often apply uniform pruning across visual tokens and layers, leading to substantial information loss and degraded this http URL address this challenge, we propose \textbf{VisCache}, a plug-and-play framework for coarse-to-fine \textbf{Vis}ual KV \textbf{Cache} pruning without training, which consists of two synergistic stages. First, a lightweight VLM filters temporal redundancy by selectively forwarding semantically informative keyframes. Second, we introduce {PruneKV}, a surgical KV compression algorithm tailored to the attention dynamics of VLLMs. Unlike rigid pruning strategies, PruneKV adopts a parabolic layer-wise budget allocation together with an asymmetric update mechanism that selectively prunes keys while fusing values, thereby preserving critical contextual information. Extensive experiments demonstrate that VisCache substantially improves inference efficiency, achieving up to {2.35$\times$ speedup} and significant memory reduction while maintaining competitive performance with only {19--28\%} KV cache retention. VisCache consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference. Code is available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2608.24063v1
- Canonical: https://arxiv.org/abs/2608.24063v1
Trouble viewing inline? Open PDF directly →
Full Text
93,019 characters extracted from source content.
Expand or collapse full text
VisCache: Visual KV Cache Pruning for Efficient Vision Large Language Model Inference Lyuke Wang Affiliation: Shenzhen International Center for Industrial and Applied Mathematics Affiliation: Shenzhen Research Institute of Big Data Affiliation: The Chinese University of Hong Kong, Shenzhen Zhuo Li †thanks: Technical lead Affiliation: Shenzhen Research Institute of Big Data Affiliation: The Chinese University of Hong Kong, Shenzhen Guangxu Zhu †thanks: Corresponding to zhuguangxu@cuhk.edu.cn Affiliation: Shenzhen International Center for Industrial and Applied Mathematics Affiliation: Shenzhen Research Institute of Big Data Affiliation: The Chinese University of Hong Kong, Shenzhen Affiliation: Shenzhen Loop Area Institute Abstract While Vision Large Language Models (VLLMs) have achieved remarkable success in multimodal reasoning, their long-context inference remains prohibitively expensive due to the massive computation and memory overhead of visual Key-Value (KV) caches. Existing KV compression methods often apply uniform pruning across visual tokens and layers, leading to substantial information loss and degraded performance. To address this challenge, we propose VisCache, a plug-and-play framework for coarse-to-fine Visual KV Cache pruning without training, which consists of two synergistic stages. First, a lightweight VLM filters temporal redundancy by selectively forwarding semantically informative keyframes. Second, we introduce PruneKV, a surgical KV compression algorithm tailored to the attention dynamics of VLLMs. Unlike rigid pruning strategies, PruneKV adopts a parabolic layer-wise budget allocation together with an asymmetric update mechanism that selectively prunes keys while fusing values, thereby preserving critical contextual information. Extensive experiments demonstrate that VisCache substantially improves inference efficiency, achieving up to 2.35× speedup and significant memory reduction while maintaining competitive performance with only 19–28% KV cache retention. VisCache consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference. Code is available at https://github.com/Wlklk/VisCache 1 Introduction The integration of vision into Large Language Models (VLLMs) has unlocked a new frontier in artificial intelligence Bai et al. (2023); Chen et al. (2024b); Li et al. (2024b); Liu et al. (2023); Liu et al. (2024a); Wu et al. (2024). By bridging the modality gap between textual understanding and visual perception, models such as LLaVA Liu et al. (2023) and GPT-4V Yang et al. (2023) have demonstrated remarkable proficiency in complex tasks ranging from long-video understanding to real-time visual conversational scenarios and derivative tasks Tao et al. (2025); Zhang et al. (2024); Yang et al. (2025b); Liu et al. (2025a); Hua et al. (2025); Varma and James (2025); Hu et al. (2022); Hu et al. (2025); Dai et al. (2026); Yuan et al. (2023); Hu et al. (2026); Li et al. (2025e); Li et al. (2025c). While VLLMs have been predominantly deployed in visual-centric tasks such as Video Summarization (VS) Bai et al. (2025a); Maaz et al. (2024) and Visual Question Answering (VQA) Li et al. (2022); Wang et al. (2025); Zhu et al. (2025), their practical adoption is severely hindered by substantial latency bottlenecks and massive memory requirements Yang et al. (2025a); Ye et al. (2025b); Li et al. (2025a), particularly when processing longer videos and higher-resolution streams. Specifically, long videos produce massive visual tokens that dramatically extend the input length fed into the LLM backbone, creating two interrelated bottlenecks. First, the Key-Value (KV) cache, inflated by the sheer volume of visual tokens, dominates GPU memory consumption and intensifies memory bandwidth contention, severely constraining throughput during long-context inference Wan et al. (2025); Pope et al. (2023); Hooper et al. (2024). Second, attention computation scales quadratically with sequence length, further amplifying latency overhead through increasingly expensive matrix multiplications Liu et al. (2025b); Shao et al. (2025). Together, these intertwined storage and computational pressures fundamentally limit the scalability of VLLMs for long-video understanding. Naively discarding visual tokens reduces computational cost but often incurs severe performance degradation Tan et al. (2025), motivating more principled KV cache compression methods Zhang et al. (2023); Li et al. (2024d); Ainslie et al. (2023); Kwon et al. (2023). Among them, quantization Liu et al. (2024b); Lin et al. (2024); Sun et al. (2024b); Sheng et al. (2023) and low-rank decomposition Chang et al. (2024); Saxena et al. (2024); Sun et al. (2024a) are widely adopted to reduce memory and computation by lowering precision or exploiting structural redundancy. However, quantization is vulnerable to activation outliers Ashkboos et al. (2024); Xiao et al. (2023), and aggressive rank constraints can impair attention capacity Chang et al. (2024); Saxena et al. (2024); Sun et al. (2024a). More fundamentally, both rely on coarse compression schemes that overlook the intricate information flow within the model. Recent studies further reveal that visual redundancy in VLLMs is highly structured rather than uniformly distributed: only a subset of video frames is relevant to a given query Lee et al. (2025); Zhou et al. (2025), different transformer layers contribute unevenly to downstream reasoning Zhang et al. (2025b); Wang et al. (2024a), and keys and values serve fundamentally different roles in attention computation Vaswani et al. (2017). These observations indicate that effective visual KV cache compression should jointly consider temporal relevance, layer-wise importance, and the asymmetric roles of keys and values. Figure 1: Visualization of different plug-and-play layer-wise KV cache compression methods. It illustrates that VisCache differs fundamentally from other baseline approaches in terms of layer-wise budgets. In this paper, we propose VisCache, a training-free and plug-and-play framework for coarse-to-fine Visual KV Cache pruning, which reduces redundancy at two complementary stages: semantic filtering before inference and structural compression during inference. At the input level, VisCache performs prompt-aware temporal filtering to eliminate redundant frames before they enter the VLLM. Concretely, a lightweight Vision-Language “scout” model (e.g., CLIP (Radford et al., 2021)) identifies task-relevant keyframes guided by prompt-aware reasoning and the Maximal Marginal Relevance (MMR) principle Li and Merialdo (2016), selecting a compact yet diverse subset of frames that preserves semantic coverage while substantially reducing visual token redundancy. At the model level, we introduce PruneKV, a layer-aware visual KV compression algorithm. Instead of applying uniform pruning across layers and token types, PruneKV dynamically allocates layer-wise compression budgets following a parabolic schedule. As shown in Figure 1, PruneKV preserves more visual tokens in early layers that encode fine-grained spatial details and progressively fewer in deeper layers that capture abstract semantics, with visual KV entries beyond a truncation threshold entirely evicted. Moreover, PruneKV adopts an asymmetric update strategy that treats keys and values differently: unimportant keys are selectively discarded, while their corresponding values are fused through weighted aggregation into the retained tokens to preserve contextual information. This attention-aware design substantially reduces KV cache redundancy while maintaining stable reasoning performance. The two stages of VisCache operate synergistically: the scout reduces redundant visual inputs before inference, while PruneKV further compresses the internal KV cache during generation. Extensive experiments demonstrate that VisCache achieves up to 2.35×2.35× inference speedup while retaining only 19–28% of the KV cache, consistently outperforming existing compression baselines across multiple video understanding benchmarks. Comprehensive ablation studies suggest that each proposed module is essential to VisCache’s strong performance. We summarize our contributions as follows: • We propose VisCache, a training-free and plug-and-play coarse-to-fine framework that jointly performs prompt-aware frame filtering and visual KV cache compression for efficient long-context VLLM inference. • We introduce PruneKV, a layer-aware KV compression algorithm with parabolic budget allocation and asymmetric key-value updates that better align pruning with attention dynamics. • Extensive experiments on long-video understanding benchmarks show that VisCache consistently achieves superior efficiency-performance trade-offs over existing baselines under aggressive KV cache compression. 2 Preliminary VLLM Prefilling. Given a video input consisting of MVM_V frames, the visual encoder of the VLLM sequentially processes each frame f that contains NVN_V visual tokens and projects them into a shared embedding space of dimension d. As a result, the visual representations of all frames are aggregated into a visual embedding matrix v∈ℝMVNV×dH_v ^M_VN_V× d. In parallel, the corresponding textual prompt T=xii=1NTT=\x_i\_i=1^N_T is fed into the text embedding layer, producing a text embedding matrix q∈ℝNT×dH_q ^N_T× d, where xix_i denotes the i-th input textual token. The visual and textual embeddings are then concatenated along the token dimension to form the unified input =concat[v,q]∈ℝ(MVNV+NT)×d.H=concat\! [H_v,\ H_q ] ^(M_VN_V+N_T)× d. For a VLLM composed of L transformer layers, the self-attention module at each layer l∈1,…,Ll∈\1,…,L\ is parameterized by three projection matrices, Ql,Kl,Vl∈ℝd×dW_Q^l,W_K^l,W_V^l ^d× d. These matrices are used to compute the query, key, and value representations as l=Ql,l=Kl,l=Vl, -1.99997ptQ^l=HW_Q^l, ^l=HW_K^l, ^l=HW_V^l, -1.99997pt where lK^l and lV^l are cached and reused during the subsequent decoding stage. VLLM Decoding. During decoding, the VLLM generates output tokens in an autoregressive (AR) manner by reusing the KV cache constructed during the prefilling stage. At decoding step t, the embedding of the previously generated token is denoted as t∈ℝ1×dh_t ^1× d. The KV cache is then incrementally updated as: l←concat[l,tKl],l←concat[l,tVl], ^l \! [K^l,\ h_tW_K^l ],V^l \! [V^l,\ h_tW_V^l ], which are then used together with the current query vector to compute the attention output and in turn produce the next token yt+1y_t+1. This process repeats until a termination condition is met or the maximum generation length is reached. Figure 2: Overview of VisCache. The framework operates in two synergistic stages: (I) A lightweight scout VLM performs prompt-aware keyframe filtering via MMR to eliminate temporal redundancy at the source; (I) During inference, the VLLM constructs the KV cache, which is subsequently refined by PruneKV through attention-aware, layer-wise compression. 3 Method We propose VisCache, a training-free and plug-and-play framework that compresses the visual KV cache through a coarse-to-fine, dual-stage paradigm. As illustrated in Figure 2, VisCache is motivated by two forms of visual redundancy in long-video VLLM inference: temporal redundancy among frames at the input level, and structural redundancy in the KV cache across tokens and layers during generation. Accordingly, VisCache operates in two complementary stages: a prompt-aware temporal filtering stage that prunes redundant frames before inference, and a layer-aware KV compression stage that surgically reduces the visual KV cache during generation. 3.1 Prompt-Aware Scout for Temporal Redundancy Filtering Given a video with MVM_V frames and a textual prompt T, we employ a lightweight vision-language model (e.g., CLIP Radford et al. (2021)) as a scout to extract aligned cross-modal representations. The text encoder EnctextEnc_text and visual encoder EncvisEnc_vis map the prompt T and each frame f into a shared embedding space by t=Enctext(T),f=Encvis(f),h_t=Enc_text(T),h_f=Enc_vis(f), where t∈ℝdh_t ^d and f∈ℝdh_f ^d denote the prompt embedding and the f-th frame embedding, respectively. To select a compact yet diverse subset of keyframes, we adopt the Maximal Marginal Relevance (MMR) criterion Carbonell and Goldstein (1998), which balances prompt relevance against inter-frame redundancy. Let Ω denote the set of selected frames. We can have MMR score of a candidate frame f: λ⋅sim(f,t)−(1−λ)⋅maxf′∈Ωsim(f,f′), -1.99997ptλ·sim\! (h_f,h_t )-(1-λ)· _f ∈ \ sim\! (h_f,h_f ), -1.99997pt (1) where sim(⋅,⋅)sim(·,·) denotes cosine similarity, and λ∈[0,1]λ∈[0,1] is a hyperparameter controlling the trade-off between relevance and diversity. The frame with the highest MMR score is iteratively added to Ω until a target retention ratio (R) p is reached. The final set Ω forms a compact, query-relevant, and minimally redundant keyframe sequence that serves as the visual input to the subsequent VLLM inference. 3.2 PruneKV: Layer-Aware Visual KV Cache Compression Rather than treating all visual KV entries uniformly, PruneKV is designed around two structural properties of transformer attention: layer-wise heterogeneity in visual token importance and functional asymmetry between keys and values. These observations motivate two components of PruneKV: layer-aware budget allocation and asymmetric key-value compression. Figure 1 provides an overview of the PruneKV pipeline. Token Scoring and Parabolic Budget Allocation During prefilling, the attention weights from all layers serve as a natural signal for token importance. For each visual token at position v, we aggregate its attention received across all queries and layers: sv=1L∑l=1L∑i=1NAi,vl, -1.99997pts_v= 1L _l=1^L _i=1^NA^l_i,v, -1.99997pt (2) where Ai,vlA^l_i,v is the attention weight from token i to token v at layer l, and N is the total sequence length. Based on svs_v as the importance score, the top q of visual KV cache that VLLM focuses on are selected. Beyond token-level pruning, we further compress the visual KV cache along the layer dimension. Visual representations become increasingly abstract with depth Ghiasi et al. (2022); Zhang et al. (2025b); Wang et al. (2024a), and the number of truly informative visual tokens decreases accordingly. We therefore allocate compression budgets in a parabolic decay: larger budgets in early layers, tapering to smaller budgets in deeper layers. Let h∈(1,L]h∈(1,L] be a truncation threshold such that visual KV entries in layers l>hl>h are fully evicted. For the remaining layers l∈[1,h]l∈[1,h], the budget proportion is: bl=1−(l−1)22(h−1)2, -1.99997ptb_l=1- (l-1)^22(h-1)^2, -1.99997pt (3) which satisfies b1=1b_1=1 and bh=0.5b_h=0.5, decaying slowly in early layers and more steeply in deeper layers. Compared with linear Cai et al. (2024) or geometric decay Xing et al. (2024), Equation (3) preserves more tokens where fine-grained visual details are encoded and compresses more aggressively where representations are sufficiently abstract to tolerate heavy pruning. We normalize bll=1h\b_l\_l=1^h to satisfy a global retention constraint ∑l=1hbl=h⋅m _l=1^hb_l=h· m, where m∈(0,1]m∈(0,1] is the target retention ratio for parabolic budget allocation module. The final budget for each layer determines how many top-scoring visual tokens are retained in kC_k (the keep set), with the remainder assigned to dC_d (the drop set). Asymmetric Key-Value Update. Existing KV compression methods (Cai et al., 2024; Xing et al., 2024) typically prune keys and values identically, ignoring their distinct roles in attention. However, the standard scaled dot-product attention reveals a clear functional asymmetry: attention weights are determined solely by ⊤QK , making keys the relevance selectors that decide which tokens are attended to, while values are the information carriers that encode the content ultimately aggregated into the output. Simply discarding value vectors therefore risks irreversibly losing contextual information that the dropped tokens originally carried. We exploit this asymmetry through a prune-keys, fuse-values strategy. For the drop set dC_d, key vectors are entirely removed, as their contribution to the attention distribution is negligible by construction (these tokens scored lowest in the importance ranking). The corresponding value vectors, however, are not discarded but redistributed to the retained tokens via similarity-based weighted aggregation. Concretely, let k∈ℝbl×dV_k ^b_l× d and d∈ℝ(n−bl)×dV_d ^(n-b_l)× d denote the value caches of the keep set kC_k and drop set dC_d, respectively. We compute a redistribution matrix that captures the semantic affinity between the two sets: =Softmax(kd⊤τ)∈ℝbl×(n−bl), =Softmax\! ( V_kV_d τ ) ^b_l×(n-b_l), (4) where τ is a temperature that controls the sharpness of the redistribution (e.g., lower τ concentrates the dropped values onto fewer retained tokens). Each row of specifies how the blb_l retained tokens absorb information from the n−bln-b_l dropped tokens. The aggregated dropped values, d _d, are then fused with the original kept values through a weighted combination: knew=μk+(1−μ)(d),V_k^new= _k+(1-μ) ( _d ), (5) where μ∈[0,1]μ∈[0,1] balances the contribution of the original and redistributed information (μ=1μ=1 recovers standard pruning; μ=0μ=0 replaces kept values entirely with fused dropped values). The compressed KV pair (knew,knew)(K_k^new,V_k^new) then replaces the original visual cache for all subsequent decoding steps. A complete description of the PruneKV algorithm is provided in Appendix B. 4 Experiment Baselines. We mainly compare VisCache with plug-and-play baselines, including: PyramidKV Cai et al. (2024) reduces the KV cache budget layer by layer following an arithmetic progression, forming a pyramidal structure based on attention scores. FastV Chen et al. (2024a) leverages the sparsity of visual attention to retain KV entries selectively. PDrop Xing et al. (2024) partitions layers into stages and progressively decreases token counts in a geometric manner, enabling stage-wise KV cache compression. Q-Frame Zhang et al. (2025a) performs adaptive frame selection and multi-resolution scales tailored to video understanding. Table 1: Comparison of different KV cache compression methods on various VQA and VS datasets. Bold and underlined numbers indicate the best and second-best results, respectively. Method R FLOPs FLOPs ActCap DREAM1K NExTQA ActQA EgoSchema Avg. (T) Ratio ROUGE-L ROUGE-L Acc Acc Acc Acc Qwen2.5-VL-3B-Instruct Full Cache 100% 14.80 100% 2.63 9.19 34.69 40.58 57.20 44.16 Q-Frame 40% 6.26 42% 2.42 5.86 39.52 39.42 46.40 41.78 PyramidKV 40% 1.99 13% 2.43 8.48 39.53 39.56 54.80 44.63 FastV 40% 3.11 21% 2.43 8.63 37.67 40.65 50.40 42.91 PDrop 40% 2.99 20% 2.46 8.63 34.24 40.21 50.00 41.48 VisCache (p=0.75,q=0.95)(p=0.75,q=0.95) 40% 1.32 9% 2.35 9.79 37.62 37.32 52.00 42.31 VisCache (p=0.75,q=0.67)(p=0.75,q=0.67) 28% 1.05 7% 2.45 8.70 41.25 40.66 55.00 45.64 VisCache (p=0.50,q=0.67)(p=0.50,q=0.67) 19% 0.88 6% 2.39 8.26 41.44 38.51 54.60 44.85 Qwen2.5-VL-32B-Instruct Full Cache 100% 93.08 100% 2.82 7.87 60.71 46.47 65.20 57.46 Q-Frame 40% 40.80 44% 2.81 7.84 50.08 43.86 54.60 49.51 PyramidKV 40% 13.88 15% 2.78 7.55 61.58 42.96 65.20 56.58 FastV 40% 51.63 55% 2.81 7.47 60.29 38.41 61.60 53.43 PDrop 40% 31.03 33% 2.75 7.38 61.00 41.59 63.20 55.26 VisCache (p=0.75,q=0.95)(p=0.75,q=0.95) 40% 13.56 15% 3.23 7.32 55.74 45.46 58.80 53.00 VisCache (p=0.75,q=0.67)(p=0.75,q=0.67) 28% 10.75 12% 2.72 7.29 62.08 42.69 65.80 56.86 VisCache (p=0.50,q=0.67)(p=0.50,q=0.67) 19% 9.08 10% 2.81 7.35 56.61 41.86 65.00 54.16 Benchmarks and Evaluation metrics. We evaluate on both video summarization (VS) and visual question answering (VQA) benchmarks. For VS, we use ActivityNet Captions (ActCap) Caba Heilbron et al. (2015) (20K videos, 100K captions) and DREAM1K Wang et al. (2024b) (1,000 clips with dense event descriptions), reporting ROUGE-L Lin (2004) as the generation metric. For VQA, we consider NExTQA Xiao et al. (2021) for temporal reasoning, ActivityNet-QA (ActQA) Yu et al. (2019) (5.8K videos), and EgoSchema Mangalam et al. (2023) for long-form comprehension. We further evaluate on MVBench Li et al. (2024c), a multi-task benchmark covering 20 reasoning tasks with multiple-choice QA pairs. Implementation Details. We implement VisCache on the Qwen2.5-VL series using PyTorch and 4 NVIDIA A100 GPUs (80GB) Bai et al. (2025b). For keyframe selection, we adopt the pretrained CLIP ViT-B/32 Radford et al. (2021) as the scout model. The frame-level pruning ratio p (Stage 1) and the token-level retention ratio q (Stage 2) serve as the primary control variables. Other hyperparameters are fixed as follows: λ=0.7λ=0.7 in Equation (1), layer truncation threshold h=34Lh= 34L, average layer-wise budget m=0.75m=0.75, temperature τ=1.0τ=1.0 in Equation (4), and fusion weight μ=0.7μ=0.7 in Equation (5). Throughout the paper, Retention Ratio (R) denotes the fraction of visual KV cache preserved relative to the full cache. For VisCache, the overall R is determined by four factors: the frame filtering ratio p, the token-level pruning ratio q, the layer truncation ratio h/Lh/L, and the average layer-wise budget m, i.e., R=p×q×hL×mR=p× q× hL× m. For example, p=0.75p=0.75, q=0.67q=0.67, h=34Lh= 34L, and m=0.75m=0.75 together yield R≈28%R≈ 28\%. For baseline methods, we tune their respective hyperparameters to match the same global R. The determination of h and μ is further analyzed in Appendix D and Appendix E. In addition, the scout VLM is model-agnostic, as discussed in Appendix L. Table 2: Results on MVBench of different KV cache compression methods with (p=0.75,q=0.67)(p=0.75,q=0.67). The first row lists abbreviations of the subset names, from AS to CI (e.g., Action Sequence (AS)). Bold and underlined numbers indicate the best and second-best results, respectively. Method R AS AP A FA UA OE OI OS MD AL ST AC MC MA SC FP CO EN ER CI Avg Qwen2.5-VL-3B-Instruct Full Cache 100% 69.0 66.8 72.0 26.4 72.4 85.0 69.7 31.3 50.5 37.5 76.0 48.5 62.5 89.0 45.5 27.0 47.1 36.5 26.8 58.0 56.6 Q-Frame 50% 49.5 63.3 68.5 23.4 64.3 80.5 59.6 31.3 46.0 37.5 77.0 36.4 56.5 78.0 40.9 24.5 52.9 35.0 37.0 61.5 51.2 PyramidKV 50% 65.5 66.8 54.0 24.9 46.4 77.5 55.6 43.8 39.0 38.0 51.0 42.4 62.5 87.0 45.5 20.5 52.9 32.0 28.8 53.0 49.4 FastV 50% 60.5 53.8 61.5 29.4 48.0 80.5 64.7 18.8 41.0 38.5 50.5 36.4 63.0 81.5 45.5 20.5 47.1 30.0 22.5 62.5 47.8 PDrop 50% 54.5 58.3 72.5 24.4 65.5 77.5 63.7 30.0 44.5 36.0 70.5 45.5 63.0 86.5 50.0 29.0 50.0 34.5 25.0 53.5 51.7 VisCache 50% 62.5 63.8 63.0 27.9 62.6 71.0 68.7 37.5 40.5 37.5 76.5 42.4 61.5 83.5 68.8 24.5 47.1 32.7 27.3 53.5 52.6 Qwen2.5-VL-32B-Instruct Full Cache 100% 68.0 66.3 76.0 39.5 70.2 84.0 69.5 68.8 44.0 39.0 85.0 39.4 57.0 79.0 63.6 35.0 47.1 38.0 45.5 46.5 58.1 Q-Frame 50% 49.0 64.3 72.5 31.0 58.8 76.0 63.6 75.0 44.5 37.5 81.0 33.3 45.0 74.5 45.5 23.0 52.9 20.0 44.5 45.5 51.9 PyramidKV 50% 43.2 56.7 80.5 33.5 66.3 74.0 59.0 46.7 43.5 19.9 85.5 39.4 51.5 66.0 54.6 25.5 58.8 25.0 41.2 49.0 51.0 FastV 50% 34.7 54.2 71.5 37.1 65.1 70.5 53.2 56.3 39.5 33.0 72.5 30.3 50.0 78.0 45.5 34.0 52.9 37.0 36.0 46.0 49.9 PDrop 50% 49.0 52.3 80.5 38.6 77.7 81.0 57.8 62.5 51.5 33.0 84.5 30.3 47.0 72.0 63.6 41.5 35.3 36.5 43.0 46.0 54.9 VisCache 50% 51.9 64.8 72.0 41.6 74.4 84.5 61.0 56.3 51.0 30.5 85.5 36.4 51.0 69.0 68.8 32.0 52.9 36.4 39.6 48.0 55.4 4.1 Main Results Table 1 compares VisCache with competitive KV cache compression baselines under controlled retention ratios. VisCache consistently achieves the lowest FLOPs across all settings. At 40% R, it reduces FLOPs to 9% of the full cache on the 3B model and 15% on the 32B model, outperforming the next best method PyramidKV. This advantage stems from the dual-stage design: the scout filters redundant frames before inference, while PruneKV further compresses the KV cache along both token and layer dimensions. At more aggressive RRs (28% and 19%), FLOPs drop to 7% and 6% on the 3B model, demonstrating the scalability of our framework. Despite aggressive compression, VisCache maintains strong accuracy. On the 3B model, VisCache at 28% R achieves the highest average accuracy (45.64%), surpassing the full cache (44.16%) and all baselines. On the 32B model, it achieves the best average accuracy (56.86%) while discarding 72% of the KV cache. VisCache consistently ranks first or second on NExTQA, ActQA, and EgoSchema. On VS benchmarks, performance is comparable across methods, with only a slight trade-off on ActCap under extreme compression—acceptable given the substantial efficiency gains. Notably, while PyramidKV performs well at 40% R, it degrades substantially under more aggressive compression (see Appendix H), confirming that parabolic allocation better preserves essential information when the budget is tight. Varying R from 40% to 19% reveals a clear accuracy–efficiency spectrum: 40% prioritizes accuracy, 28% offers the best overall trade-off, and 19% still retains strong performance, outperforming most baselines at higher RRs. These results highlight VisCache’s robustness under extreme compression, where uniform or arithmetic budgets often collapse. Detailed comparisons at additional RRs are provided in Appendix I. VisCache scales effectively from 3B to 32B: efficiency gains are consistent, and accuracy advantages grow with model size. On the 32B model, VisCache at 28% R outperforms all baselines at 40% R. Table 2 reports the performance on MVBench, a comprehensive multi-task benchmark spanning 20 reasoning tasks. VisCache achieves the highest average accuracy on both model scales (52.6% on 3B, 55.4% on 32B), surpassing all baselines at the same retention ratio. On the 32B model, VisCache outperforms the strongest baseline PDrop by 0.5 points in average accuracy, with particularly strong gains on tasks requiring temporal reasoning (AS: 51.9%, UA: 74.4%) and fine-grained perception (A: 41.6%). Notably, VisCache and PDrop exhibit complementary strengths across sub-tasks: VisCache leads on 11 out of 20 tasks on the 32B model, while PDrop leads on the remaining tasks, suggesting that parabolic allocation and asymmetric fusion are especially beneficial for tasks that depend on structured visual information retained across layers. The overall margin over the full cache is small (within 3 points on 32B), confirming that VisCache preserves broad reasoning capabilities under aggressive compression. The compatibility evaluation of VisCache is provided in Appendix F and Appendix G. In addition, representative inference examples are presented in Appendix M. 4.2 Analysis and Ablation Studies Table 3: Real-system inference efficiency on ActCap. Mem.: GPU memory in GB (Total = peak usage, KV = raw KV cache size). TPS: tokens per second (higher is better). Method R Total Mem. KV Cache TPS Full Cache 100% 5.57 0.06 10.4 PyramidKV 40% 4.98 0.02 9.7 FastV 40% 4.98 0.02 10.6 PDrop 40% 4.99 0.02 15.0 VisCache 40% 4.10 0.03 15.5 VisCache 28% 4.10 0.02 18.6 VisCache 19% 3.68 0.02 15.7 Memory Usage. We measure the total GPU memory and the KV cache GPU memory for different acceleration methods on ActCap based on Qwen2.5-VL-3B-Instruct backbone. As Table 3 shown, VisCache reduces the total GPU memory and KV cache GPU memory to 3.68 GB and 0.02 GB while maintaining competitive throughput and nearly identical inference performance compared to the baseline methods. These results demonstrate that VisCache provides an efficient trade-off between memory consumption and inference speed, enabling deployment of VLLMs on resource-constrained edge devices without incurring significant degradation in output quality. Figure 3: Comparison of average inference time of different baseline algorithms on Qwen2.5-VL-3B-Instruct at different KV cache compression rates for the VS task. The left, middle, and right columns report the average end-to-end (E2E) latency, time to first token (TTFT), and time per output token (TPOT), respectively. Figure 4: Comparison of average inference time between full KV cache and VisCache under varying numbers of output tokens. Inference Time. We record the computation time of VisCache and baseline algorithms under different visual KV cache compression rates and plot the results in Figure 3. It can be clearly observed that the average End-to-End (E2E) inference time of VisCache-which includes the keyframe selection stage as well as the subsequent VLLM inference-remains consistently lower than that of the full cache and baseline methods across all retention rates, further validating the overall efficiency of VisCache. By retaining just 28%\% of the KV cache, VisCache already achieves impressive E2E speedups of up to 1.93 × over the full cache setting. Reducing the retention further to 19%\% amplifies the effect, pushing the acceleration to an astonishing 2.35 ×, all while achieving performance close to the baseline. Specific inference time measurements are explained in Appendix K. VisCache significantly reduces Time Per Output Token (TPOT), yielding substantial time savings compared to full cache and baseline methods, while its effect on Time To First Token (TTFT) is minimal. This indicates that although the first token is not noticeably accelerated, subsequent decoding is greatly sped up, more than compensating for the initial overhead and reducing overall E2E latency. To better illustrate this phenomenon, We increase the number of output tokens in the model from 64 to 128, and the experimental results are shown in Figure 4. The more tokens output, the more time VisCache saves compared to full cache, further confirming our conclusion. Table 4: Ablation study on budget allocation strategy and value cache fusion. We keep R=28%R=28\% of the visual KV cache on Qwen2.5-VL-32B-Instruct across all variants. Fixed: uniform budget per layer. Arithmetic: arithmetic progression (PyramidKV Cai et al. (2024)). Geometric: geometric progression across four layer blocks (PDrop Xing et al. (2024)). Parabola: our proposed parabolic schedule (Equation 3). V Cache Fusion: asymmetric value fusion. Budget V Cache ActCap DREAM1K NExTQA ActQA EgoSchema Avg. Strategy Fusion ROUGE-L ROUGE-L Acc Acc Acc Acc Full cache (upper bound) Full Cache – 2.82 7.87 60.71 46.47 65.20 57.46 Without V Cache Fusion Fixed × 2.43 7.24 59.00 43.44 63.60 55.35 Arithmetic × 2.67 7.17 58.07 42.31 57.00 52.46 Geometric × 2.69 7.25 55.56 42.68 57.60 51.95 Parabola × 2.70 7.05 61.61 42.64 65.00 56.42 With V Cache Fusion Fixed ✓ 2.69 7.36 61.58 42.42 64.80 56.27 Arithmetic ✓ 2.72 7.23 61.32 42.33 65.00 56.22 Geometric ✓ 2.69 7.27 56.07 40.49 56.00 50.85 Parabola ✓ 2.72 7.29 62.08 42.69 65.80 56.86 Impact of Layer-wise Budget Allocation and Value Fusion. We first ablate the two core design choices in PruneKV: the layer-wise budget allocation strategy and the asymmetric value fusion mechanism. Table 4 reports results under a fixed 28% KV cache retention ratio. Among the four budget strategies, Parabola consistently achieves the best performance across all three VQA benchmarks, with its average accuracy approaching that of the full cache upper bound. Arithmetic and Fixed strategies also perform competitively, while Geometric allocation degrades substantially, confirming that overly aggressive compression in early layers impairs fine-grained visual understanding. To assess the contribution of value fusion, we compare each budget strategy with and without fusion. For Parabola, Fixed, and Arithmetic allocations, enabling fusion consistently improves accuracy, demonstrating that asymmetric value fusion effectively preserves contextual information that would otherwise be lost. The exception is Geometric allocation, where fusion slightly degrades performance—likely because excessive pruning leaves too few retained tokens for meaningful redistribution, further underscoring the importance of pairing fusion with a well-designed budget strategy. Notably, Parabola with fusion achieves the best individual scores across all VQA benchmarks, validating that the two components are synergistic: parabolic allocation retains tokens where they matter most, while value fusion salvages information from pruned tokens. 4.3 Ablation on Two Stages of VisCache. Table 5: Comparison of different KV cache compression methods on various VQA and VS datasets. Bold and underlined numbers indicate the best and second-best results, respectively. Methods R Scout PruneKV DREAM1K EgoSchema ROUGE-L Acc PDrop 40% - - 8.63 50.00 VisCache 40% √ × 8.48 46.20 VisCache 40% × √ 9.65 48.00 VisCache 40% √ √ 9.79 52.00 PDrop 28% - - 8.62 49.60 VisCache 28% √ × 8.13 46.00 VisCache 28% × √ 8.62 48.80 VisCache 28% √ √ 8.70 55.00 PDrop 19% - - 7.69 49.00 VisCache 19% √ × 7.84 48.20 VisCache 19% × √ 8.23 48.60 VisCache 19% √ √ 8.26 54.60 Table 5 presents the ablation study of the two-stage design in VisCache, including Scout-based temporal redundancy filtering and PruneKV-based visual KV cache compression. We first observe that applying only Scout-based filtering leads to noticeable performance degradation in some settings, especially on EgoSchema, indicating that solely removing temporally redundant frames may discard informative visual context required for reasoning. In contrast, using only PruneKV generally achieves more stable performance and consistently outperforms the baseline PDrop under aggressive compression ratios, demonstrating the effectiveness of fine-grained visual KV cache pruning. Table 6: Ablation of shared global and layer-specific visual-token ranking at matched retention ratios on Qwen2.5-VL-3B-Instruct. Bold numbers indicate the best result under each retention ratio. Ranking strategy R ActCap DREAM1K NExTQA ActQA EgoSchema Avg. Acc. ROUGE-L ROUGE-L Acc. Acc. Acc. Shared global ranking 28% 2.45 8.70 41.25 40.66 55.00 45.64 Layer-specific ranking 28% 2.35 8.73 41.20 39.82 51.75 44.26 Shared global ranking 19% 2.39 8.26 41.44 38.51 54.60 44.85 Layer-specific ranking 19% 2.29 7.95 41.24 37.41 53.20 43.95 More importantly, combining both Scout and PruneKV consistently achieves the best overall performance across different RRs. This suggests that the two stages are highly complementary: Scout reduces coarse-grained temporal redundancy at the frame level, while PruneKV further refines the retained information through token-level KV cache compression. The advantage becomes more significant at lower RRs, where the complete VisCache framework maintains strong reasoning performance even under highly constrained memory budgets. 4.4 Ablation on Shared vs. Layer-Specific Visual Token Ranking We further examine whether visual token importance should be computed globally across all layers or independently for each layer. In the shared global ranking, the importance of visual token v is obtained by aggregating its attention scores across all transformer layers: sv=∑l=1L∑iAi,v(l),s_v= _l=1^L _iA_i,v^(l), (6) where Ai,v(l)A_i,v^(l) denotes the attention weight received by visual token v at layer l. All layers then select tokens from the same global ordering according to their retention budgets, i.e., layer l keeps the top-klk_l tokens from this shared ranking. In the layer-specific variant, each layer independently computes sv(l)=∑iAi,v(l)s_v^(l)= _iA_i,v^(l) (7) and selects its own top-klk_l visual tokens. Table 6 compares the two strategies at matched retention ratios on Qwen2.5-VL-3B-Instruct. Shared global ranking improves the average VQA accuracy from 44.26 to 45.64 at 28% retention ratio and from 43.95 to 44.85 at 19% retention ratio. It also achieves the best result on most individual benchmarks. This design intentionally decouples which tokens are important from how many tokens each layer retains. Aggregating attention across all layers provides a more stable and comprehensive importance estimate, while the parabolic budget still assigns different retention cardinalities to each layer. In contrast, layer-specific ranking may select substantially different token identities across layers, which can reduce the consistency of visual information preserved throughout the backbone. 5 Related Work Post-aligned LLMs and VLLMs usually come with redundant responses (Li et al., 2025b; Li et al., 2025d; Li et al., 2026a), leading to additional unsafe behaviors (Wang et al., 2026; Du et al., 2025; Li et al., 2026b) and inference cost. Existing efficient VLLM inference methods improve efficiency by localizing query-aware video clip, reducing visual token redundancy or compressing KV caches. SeViLA Yu et al. (2023) performs query-aware keyframe localization and QA via a self-chained design, while LongVU Shen et al. (2024) achieves efficient long-video understanding via adaptive spatiotemporal token compression. One line of work focuses on token pruning and merging based on attention or semantic importance. Methods such as FastV Chen et al. (2024a), VisionZip Yang et al. (2025b), and SparseVLM Zhang et al. (2024) select or merge tokens using attention or cross-modal relevance, while approaches like FitPrune Ye et al. (2025a), PruMerge Shang et al. (2025), and DivPrune Alvar et al. (2025) further explore optimization and diversity-aware selection strategies. Another line of work reduces inference cost via structured KV cache compression or layer-wise budget allocation. PyramidKV Cai et al. (2024), PyramidInfer Yang et al. (2024), and PDrop Xing et al. (2024) allocate KV budgets across layers using predefined schedules such as arithmetic or geometric decay. Despite their effectiveness, these methods typically treat tokens or KV caches in a coarse-grained manner, ignoring the heterogeneous importance of visual information across layers and the asymmetric roles of keys and values, which can limit compression performance under aggressive budgets. 6 Conclusion We present VisCache, a training-free framework for efficient visual KV cache compression in long-video VLLMs via prompt-aware filtering and layer-aware PruneKV, preserving hierarchical context while improving long-context video efficiency. Limitations We identify two primary limitations of this work. First, the scout-based temporal filtering stage relies on a lightweight VLM such as CLIP, whose visual encoder may not perfectly align with the main LLM backbone. In scenarios where the scout and the main model exhibit substantially different visual representations, the selected keyframes may deviate from what the downstream model would consider optimal. Second, the current implementation of PruneKV requires storing attention scores from all layers during the prefilling stage, which introduces some memory overhead. We leave the exploration of scout-backbone co-adaptation and memory-efficient score computation to future work. Statement of Impacts This work aims to improve the inference efficiency of vision large language models, thereby lowering the computational barrier for deploying advanced video understanding capabilities in practice. By substantially reducing memory consumption and latency, VisCache can contribute to broader accessibility of VLLM technologies, particularly in academic and resource-limited settings. We acknowledge that efficiency improvements in VLLMs may also facilitate large-scale video analysis applications, including automated surveillance and mass media monitoring. We encourage practitioners deploying VisCache in such contexts to adhere to established ethical guidelines regarding privacy, consent, and algorithmic fairness. Our method reduces visual token redundancy based on learned attention patterns from pretrained models; as such, any biases present in the underlying foundation models may propagate through the compression process. We recommend that users of VisCache conduct fairness and bias assessments tailored to their specific application domains. Acknowledgments This work was supported in part by the National Natural Science Foundation of China (Grant No. 62522118, 62371313), in part by the Shenzhen Science and Technology Program (Grant No. JCYJ20241202124934046), in part by the Guangdong Young Talent Research Project (Grant No. 2023TQ07A708), in part by Shenzhen Loop Area Institute (Contract No. SLAI2026020007). References Ainslie et al. (2023) J. Ainslie, J. Lee-Thorp, M. De Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai Gqa: training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 4895–4901. Cited by: §1. Alvar et al. (2025) S. R. Alvar, G. Singh, M. Akbari, and Y. Zhang Divprune: diversity-based visual token pruning for large multimodal models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 9392–9401. Cited by: §5. Ashkboos et al. (2024) S. Ashkboos, A. Mohtashami, M. L. Croci, B. Li, P. Cameron, M. Jaggi, D. Alistarh, T. Hoefler, and J. Hensman Quarot: outlier-free 4-bit inference in rotated llms. Advances in Neural Information Processing Systems 37, p. 100213–100240. Cited by: §1. Bai et al. (2023) J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, et al. Qwen technical report. arXiv preprint arXiv:2309.16609. Cited by: §1. Bai et al. (2025a) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §1. Bai et al. (2025b) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923. Cited by: §4. Caba Heilbron et al. (2015) F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles Activitynet: a large-scale video benchmark for human activity understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4. Cai et al. (2024) Z. Cai, Y. Zhang, B. Gao, Y. Liu, Y. Li, T. Liu, K. Lu, W. Xiong, Y. Dong, J. Hu, et al. Pyramidkv: dynamic kv cache compression based on pyramidal information funneling. arXiv preprint arXiv:2406.02069. Cited by: §3.2, §3.2, §4, Table 4, §5. Carbonell and Goldstein (1998) J. Carbonell and J. Goldstein The use of mmr, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval, p. 335–336. Cited by: §3.1. Chang et al. (2024) C. Chang, W. Lin, C. Lin, C. Chen, Y. Hu, P. Wang, N. Huang, L. Ceze, M. S. Abdelfattah, and K. Wu Palu: compressing kv-cache with low-rank projection. arXiv preprint arXiv:2407.21118. Cited by: §1. Chen et al. (2024a) L. Chen, H. Zhao, T. Liu, S. Bai, J. Lin, C. Zhou, and B. Chang An image is worth 1/2 tokens after layer 2: plug-and-play inference acceleration for large vision-language models. In European Conference on Computer Vision, p. 19–35. Cited by: §4, §5. Chen et al. (2024b) Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. Internvl: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 24185–24198. Cited by: §1. Cherti et al. (2023) M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2818–2829. Cited by: Appendix L. Dai et al. (2026) C. Dai, J. Hu, H. Shi, Z. Li, D. Guo, X. Yang, and M. Wang Psyche-r1: towards reliable psychological LLMs through unified empathy, expertise, and reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 24889–24906. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §1. Du et al. (2025) Y. Du, Z. Li, P. Cheng, X. Wan, and A. Gao Atoxia: red-teaming large language models with target toxic answers. In Findings of the Association for Computational Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 3251–3266. External Links: Link, Document, ISBN 979-8-89176-195-7 Cited by: §5. Ghiasi et al. (2022) A. Ghiasi, H. Kazemi, E. Borgnia, S. Reich, M. Shu, M. Goldblum, A. G. Wilson, and T. Goldstein What do vision transformers learn? a visual exploration. External Links: 2212.06727, Link Cited by: §3.2. Hooper et al. (2024) C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami Kvquant: towards 10 million context length llm inference with kv cache quantization. Advances in Neural Information Processing Systems 37, p. 1270–1303. Cited by: §1. Hu et al. (2022) J. Hu, Z. Li, Z. Chen, Z. Li, X. Wan, and T. Chang Graph enhanced contrastive learning for radiology findings summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 4677–4688. External Links: Link, Document Cited by: §1. Hu et al. (2025) J. Hu, H. Shi, C. Dai, Z. Li, P. Song, and M. Wang Beyond emotion recognition: a multi-turn multimodal emotion understanding and reasoning benchmark. In Proceedings of the 33rd ACM International Conference on Multimedia, M ’25, New York, NY, USA, p. 5814–5823. External Links: ISBN 9798400720352, Link, Document Cited by: §1. Hu et al. (2026) J. Hu, A. Wang, Q. Xie, Z. Li, H. Ma, and D. Guo Agentmental: an interactive multi-agent framework for explainable and adaptive mental health assessment. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 31050–31058. Cited by: §1. Hua et al. (2025) H. Hua, Y. Tang, C. Xu, and J. Luo V2xum-llm: cross-modal video summarization with temporal prompt instruction tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 3599–3607. Cited by: §1. Kwon et al. (2023) W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, p. 611–626. Cited by: §1. Lee et al. (2025) H. Lee, J. Kim, H. Kim, and Y. M. Ro Refocus: reinforcement-guided frame optimization for contextual understanding. arXiv preprint arXiv:2506.01274. Cited by: §1. Li et al. (2024a) B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y. Li, Z. Liu, and C. Li LLaVA-onevision: easy visual task transfer. External Links: 2408.03326, Link Cited by: §F.2, Appendix F. Li et al. (2024b) F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li Llava-next-interleave: tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895. Cited by: §1. Li et al. (2022) J. Li, D. Li, C. Xiong, and S. Hoi Blip: bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, p. 12888–12900. Cited by: Appendix L, §1. Li et al. (2024c) K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, et al. Mvbench: a comprehensive multi-modal video understanding benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 22195–22206. Cited by: §4. Li et al. (2025a) K. Li, Z. Jiang, Z. Shen, Z. ZhaodeWang, C. Lv, S. Zhang, F. Wu, and F. Wu MadaKV: adaptive modality-perception kv cache eviction for efficient multimodal long-context inference. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 13306–13318. Cited by: §1. Li and Merialdo (2016) Y. Li and B. Merialdo Multimedia maximal marginal relevance for multi-video summarization. Multimedia Tools and Applications 75 (1), p. 199–220. Cited by: §1. Li et al. (2024d) Y. Li, Y. Huang, B. Yang, B. Venkitesh, A. Locatelli, H. Ye, T. Cai, P. Lewis, and D. Chen Snapkv: llm knows what you are looking for before generation. Advances in Neural Information Processing Systems 37, p. 22947–22970. Cited by: §1. Li et al. (2026a) Z. Li, P. Cheng, Z. Yu, FeifeiTong, A. Gao, T. Chang, X. Wan, erchao.zec, x. jiang, and guanjunjiang Eliminating inductive bias in reward models with information-theoretic guidance. In International Conference on Learning Representations, C. Vondrick, B. Hariharan, C. Raffel, L. Pinto, D. Yang, and A. Faust (Eds.), Vol. 2026, p. 129961–129986. External Links: Link Cited by: §5. Li et al. (2025b) Z. Li, Y. Du, J. Hu, X. Wan, and A. Gao Self-instructed derived prompt generation meets in-context learning: unlocking new potential of black-box LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 1840–1857. External Links: Link, Document, ISBN 979-8-89176-251-0 Cited by: §5. Li et al. (2025c) Z. Li, Y. Du, X. Jiao, S. Y. Guo, Y. Feng, X. Wan, A. Gao, and J. Hu Add-one-in: incremental sample selection for large language models via a choice-based greedy paradigm. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 5321–5340. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1. Li et al. (2025d) Z. Li, Y. Feng, D. Guo, J. Hu, A. Gao, and X. Wan APLOT: robust reward modeling via adaptive preference learning with optimal transport. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 5524–5538. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §5. Li et al. (2026b) Z. Li, Y. Zhang, P. Cheng, J. Song, M. Zhou, H. Li, S. Hu, Y. Qin, E. Zhao, X. Jiang, and G. Jiang MARCH: multi-agent reinforced check for hallucination. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 39389–39415. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §5. Li et al. (2025e) Z. Li, H. Zhao, A. Gao, D. Guo, T. Chang, and X. Wan Prototype-oriented clean subset extraction for noisy long-tailed classification. IEEE Transactions on Circuits and Systems for Video Technology 35 (8), p. 7953–7965. External Links: Document Cited by: §1. Lin (2004) C. Lin Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, p. 74–81. Cited by: §4. Lin et al. (2024) H. Lin, H. Xu, Y. Wu, J. Cui, Y. Zhang, L. Mou, L. Song, Z. Sun, and Y. Wei Duquant: distributing outliers via dual transformation makes stronger quantized llms. Advances in Neural Information Processing Systems 37, p. 87766–87800. Cited by: §1. Liu et al. (2024a) H. Liu, C. Li, Y. Li, and Y. J. Lee Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 26296–26306. Cited by: §1. Liu et al. (2023) H. Liu, C. Li, Q. Wu, and Y. J. Lee Visual instruction tuning. Advances in neural information processing systems 36, p. 34892–34916. Cited by: §1. Liu et al. (2025a) J. Liu, L. Niu, W. Chen, J. Zhou, and F. Meng LaCo: efficient layer-wise compression of visual tokens for multimodal large language models. arXiv preprint arXiv:2507.02279. Cited by: §1. Liu et al. (2025b) X. Liu, X. Wang, P. Liu, and G. Tang Zsmerge: zero-shot kv cache compression for memory-efficient long-context llms. arXiv preprint arXiv:2503.10714. Cited by: §1. Liu et al. (2024b) Z. Liu, J. Yuan, H. Jin, S. Zhong, Z. Xu, V. Braverman, B. Chen, and X. Hu Kivi: a tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750. Cited by: Appendix G, §1. Maaz et al. (2024) M. Maaz, H. Rasheed, S. Khan, and F. Khan Video-chatgpt: towards detailed video understanding via large vision and language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 12585–12602. Cited by: §1. Mangalam et al. (2023) K. Mangalam, R. Akshulakov, and J. Malik Egoschema: a diagnostic benchmark for very long-form video language understanding. Advances in Neural Information Processing Systems 36, p. 46212–46244. Cited by: §4. Pope et al. (2023) R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean Efficiently scaling transformer inference. Proceedings of machine learning and systems 5, p. 606–624. Cited by: §1. Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: Appendix L, §1, §3.1, §4. Saxena et al. (2024) U. Saxena, G. Saha, S. Choudhary, and K. Roy Eigen attention: attention in low-rank space for kv cache compression. In Findings of the Association for Computational Linguistics: EMNLP 2024, p. 15332–15344. Cited by: §1. Shang et al. (2025) Y. Shang, M. Cai, B. Xu, Y. J. Lee, and Y. Yan Llava-prumerge: adaptive token reduction for efficient large multimodal models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 22857–22867. Cited by: §5. Shao et al. (2025) K. Shao, K. Tao, C. Qin, H. You, Y. Sui, and H. Wang HoliTom: holistic token merging for fast video large language models. arXiv preprint arXiv:2505.21334. Cited by: §1. Shen et al. (2024) X. Shen, Y. Xiong, C. Zhao, L. Wu, J. Chen, C. Zhu, Z. Liu, F. Xiao, B. Varadarajan, F. Bordes, et al. Longvu: spatiotemporal adaptive compression for long video-language understanding. arXiv preprint arXiv:2410.17434. Cited by: §5. Sheng et al. (2023) Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang Flexgen: high-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning, p. 31094–31116. Cited by: §1. Sun et al. (2024a) H. Sun, L. Chang, W. Bao, S. Zheng, N. Zheng, X. Liu, H. Dong, Y. Chi, and B. Chen Shadowkv: kv cache in shadows for high-throughput long-context llm inference. arXiv preprint arXiv:2410.21465. Cited by: §1. Sun et al. (2024b) Y. Sun, R. Liu, H. Bai, H. Bao, K. Zhao, Y. Li, J. Hu, X. Yu, L. Hou, C. Yuan, et al. Flatquant: flatness matters for llm quantization. arXiv preprint arXiv:2410.09426. Cited by: Appendix G, §1. Tan et al. (2025) X. Tan, P. Ye, C. Tu, J. Cao, Y. Yang, L. Zhang, D. Zhou, and T. Chen Tokencarve: information-preserving visual token compression in multimodal large language models. arXiv preprint arXiv:2503.10501. Cited by: §1. Tao et al. (2025) K. Tao, C. Qin, H. You, Y. Sui, and H. Wang DyCoke: dynamic compression of tokens for fast video large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 18992–19001. Cited by: §1. Varma and James (2025) S. Varma and D. P. James RETRACTED: an efficient deep learning-based video captioning framework using multi-modal features. Expert Systems 42 (2), p. e12920. Cited by: §1. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30. Cited by: §1. Wan et al. (2025) Z. Wan, H. Shen, X. Wang, C. Liu, Z. Mai, and M. Zhang Meda: dynamic kv cache allocation for efficient multimodal long-context inference. arXiv preprint arXiv:2502.17599. Cited by: §1. Wang et al. (2024a) A. Wang, H. Chen, J. Li, J. Tan, K. Zhang, X. Cai, Z. Lin, J. Han, and G. Ding Prefixkv: adaptive prefix kv cache is what vision instruction-following models need for efficient generation. arXiv preprint arXiv:2412.03409. Cited by: §1, §3.2. Wang et al. (2026) H. Wang, Z. Li, Y. Yang, H. Zhao, H. Zha, and D. Guo Safeguarding LLM fine-tuning via push-pull distributional alignment. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, p. 23624–23646. External Links: Link, Document, ISBN 979-8-89176-390-6 Cited by: §5. Wang et al. (2024b) J. Wang, L. Yuan, Y. Zhang, and H. Sun Tarsier: recipes for training and evaluating large video description models. arXiv preprint arXiv:2407.00634. Cited by: §4. Wang et al. (2025) W. Wang, Z. Gao, L. Gu, H. Pu, L. Cui, X. Wei, Z. Liu, L. Jing, S. Ye, J. Shao, et al. InternVL3.5: advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265. Cited by: §1. Wu et al. (2024) Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al. Deepseek-vl2: mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302. Cited by: §1. Xiao et al. (2023) G. Xiao, J. Lin, M. Seznec, H. Wu, J. Demouth, and S. Han Smoothquant: accurate and efficient post-training quantization for large language models. In International conference on machine learning, p. 38087–38099. Cited by: §1. Xiao et al. (2021) J. Xiao, X. Shang, A. Yao, and T. Chua Next-qa: next phase of question-answering to explaining temporal actions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9777–9786. Cited by: §4. Xing et al. (2024) L. Xing, Q. Huang, X. Dong, J. Lu, P. Zhang, Y. Zang, Y. Cao, C. He, J. Wang, F. Wu, et al. Pyramiddrop: accelerating your large vision-language models via pyramid visual redundancy reduction. arXiv preprint arXiv:2410.17247. Cited by: §3.2, §3.2, §4, Table 4, §5. Yang et al. (2025a) C. Yang, Y. Sui, J. Xiao, L. Huang, Y. Gong, C. Li, J. Yan, Y. Bai, P. Sadayappan, X. Hu, et al. Topv: compatible token pruning with inference time optimization for fast and low-memory multimodal vision language model. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 19803–19813. Cited by: §1. Yang et al. (2024) D. Yang, X. Han, Y. Gao, Y. Hu, S. Zhang, and H. Zhao Pyramidinfer: pyramid kv cache compression for high-throughput llm inference. In Findings of the Association for Computational Linguistics: ACL 2024, p. 3258–3270. Cited by: §5. Yang et al. (2025b) S. Yang, Y. Chen, Z. Tian, C. Wang, J. Li, B. Yu, and J. Jia Visionzip: longer is better but not necessary in vision language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 19792–19802. Cited by: §1, §5. Yang et al. (2023) Z. Yang, L. Li, K. Lin, J. Wang, C. Lin, Z. Liu, and L. Wang The dawn of lmms: preliminary explorations with gpt-4v (ision). arXiv preprint arXiv:2309.17421. Cited by: §1. Ye et al. (2025a) W. Ye, Q. Wu, W. Lin, and Y. Zhou Fit and prune: fast and training-free visual token pruning for multi-modal large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 22128–22136. Cited by: §5. Ye et al. (2025b) X. Ye, Y. Gan, X. Huang, Y. Ge, and Y. Tang Voco-llama: towards vision compression with large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 29836–29846. Cited by: §1. Yu et al. (2023) S. Yu, J. Cho, P. Yadav, and M. Bansal Self-chained image-language model for video localization and question answering. In NeurIPS, Cited by: §5. Yu et al. (2019) Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao Activitynet-qa: a dataset for understanding complex web videos via question answering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33, p. 9127–9134. Cited by: §4. Yuan et al. (2023) Z. Yuan, X. Yan, Z. Li, X. Li, Y. Guo, S. Cui, and Z. Li Toward explainable and fine-grained 3d grounding through referring textual phrases. External Links: 2207.01821, Link Cited by: §1. Zhang et al. (2025a) S. Zhang, J. Yang, J. Yin, Z. Luo, and J. Luan Q-frame: query-aware frame selection and multi-resolution adaptation for video-llms. arXiv preprint arXiv:2506.22139. Cited by: §4. Zhang et al. (2024) Y. Zhang, C. Fan, J. Ma, W. Zheng, T. Huang, K. Cheng, D. Gudovskiy, T. Okuno, Y. Nakata, K. Keutzer, et al. Sparsevlm: visual token sparsification for efficient vision-language model inference. arXiv preprint arXiv:2410.04417. Cited by: §1, §5. Zhang et al. (2023) Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett, et al. H2o: heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems 36, p. 34661–34710. Cited by: §1. Zhang et al. (2025b) Z. Zhang, S. Yadav, F. Han, and E. Shutova Cross-modal information flow in multimodal large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 19781–19791. Cited by: §1, §3.2. Zhou et al. (2025) Y. Zhou, L. Hua, S. Jin, W. Huang, and H. Duan ReaSon: reinforced causal search with information bottleneck for video understanding. arXiv preprint arXiv:2511.12530. Cited by: §1. Zhu et al. (2025) J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, et al. Internvl3: exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479. Cited by: §1. Appendix A LLM Usage Statement We employ a large language model (LLM) as a general-purpose writing assistant to improve the clarity and fluency of the text. Its role is limited to refining linguistic expression and enhancing overall readability and coherence. Appendix B Algorithm Algorithm 1 PruneKV: Layer-Aware Visual KV Compression 1: Visual KV cache (l,l)l=1L\(K^l,V^l)\_l=1^L, attention scores ll=1L\A^l\_l=1^L, truncation layer h, retention ratio m, fusion strength μ, temperature τ 2: Compressed visual KV cache (newl,newl)l=1h\(K_new^l,V_new^l)\_l=1^h 3: ⊳ Step 1: Token-level importance scoring 4: ¯←1L∑l=1Ll A← 1L _l=1^LA^l ⊳ Average attention across layers 5: sv←∑i=1NA¯i,v,∀vs_v← _i=1^N A_i,v, ∀ v ⊳ Aggregate attention received by each visual token 6: Retain KV entries of the q percent visual tokens by svs_v, discard the rest ⊳ r determined by m and h 7: ⊳ Step 2: Parabolic budget allocation 8: Compute bll=1h\b_l\_l=1^h via Equation (3) 9: ⊳ Step 3: Layer-wise asymmetric compression 10: for l←1l← 1 to h do 11: Partition visual tokens at layer l into keep set kC_k (top-blb_l by svs_v) and drop set dC_d 12: newl←K_new^l← key vectors of kC_k ⊳ Prune keys in dC_d 13: k←V_k← value vectors of kC_k, d←V_d← value vectors of dC_d 14: Compute redistribution matrix via Equation (4) ⊳ Similarity-based affinity 15: newl←μk+(1−μ)dV_new^l← _k+(1-μ)\, _d ⊳ Fuse dropped values into kept values, Equation (5) 16: end for 17: return (newl,newl)l=1h\(K_new^l,V_new^l)\_l=1^h Appendix C The Effect of Selecting Keyframes by Small VLM To examine the reliability of using a small VLM for keyframe selection, we conduct the following experiment. During the prefilling stage, we identify a set of key visual tokens by aggregating attention scores across all layers of the VLLM. We then compare this token set with the visual tokens contained in keyframes selected by a small VLM guided by MMR-based selection. Specifically, on two generation-oriented benchmarks, we compute the Jaccard similarity between these two token sets while ensuring they contain the same number of tokens. By varying the ratio between the selected tokens and the total visual tokens in the original video, we obtain the results illustrated in the Figure 5. Figure 5: Jaccard similarity between key visual tokens selected by the small VLM with MMR and those selected from the sum of attention scores across Qwen2.5-VL-3B-Instruct layers, under different visual token compression rates. This figure highlights the consistency between the two token sets at varying compression levels. The results show that, across different compression ratios, the visual token sets selected by the small VLM combined with the MMR-based algorithm consistently exhibit high Jaccard similarity with the tokens receiving higher attention scores in the VLLM. This indicates that the keyframes identified by the small VLM largely correspond to the visually salient tokens emphasized by the VLLM itself. By filtering these keyframes in advance, the number of visual tokens entering the prefilling stage is substantially reduced, thereby lowering the computational cost of both prefilling and decoding and ultimately improving the inference efficiency of the VLLM. Appendix D How To Determine the Truncation Layer of VLLM To determine an appropriate truncation layer for VLLM, we jointly consider the trade-off between inference accuracy and inference efficiency. Specifically, we conduct an ablation study on Qwen2.5-VL-32B-Instruct, which contains 64 transformer layers in total. For different truncation points, we evaluate the inference performance on EgoSchema while also measuring the corresponding inference speed. As illustrated in Figure 6, increasing the truncation layer generally improves inference accuracy because more high-level semantic information can be preserved during decoding. However, this also introduces higher computational overhead, resulting in slower inference speed. In contrast, truncating at earlier layers significantly accelerates inference but causes noticeable performance degradation due to insufficient semantic reasoning capability. From the experimental results, we observe that the balance between inference accuracy and speed is reached when the truncation layer is set to 48, which corresponds to 3/43/4 of the total model depth (64 layers). At this point, the VLLM maintains strong reasoning capability while avoiding the substantial latency increase introduced by deeper layers. Therefore, we adopt the 3/43/4 truncation depth as the default configuration in VisCache. Figure 6: Trade-off between inference accuracy and inference speed under different truncation layers of VLLM. Appendix E Impact of Fusion Strength To investigate the impact of the fusion strength μ of the V cache on the inference accuracy of VLLM, we conduct experiments on EgoSchema using the Qwen2.5-VL-32B-Instruct backbone. Specifically, we vary the fusion strength μ from 0.10.1 to 0.90.9 and evaluate the corresponding inference performance. The experimental results are shown in Figure 7 and the statistical results are summarized in Table 7. To reduce the influence of randomness during inference, each configuration is evaluated three times independently. We report both the mean accuracy and the variance across the three runs. The variance reflects the stability of the VLLM under different fusion strengths. As shown in Figure 7, the inference accuracy reaches its maximum when the fusion strength μ is set to 0.70.7. This indicates that an appropriate fusion strength can effectively preserve informative semantic representations while reducing redundant V cache information. When μ is smaller than 0.70.7, the contribution of the fused V cache becomes insufficient, resulting in weaker preservation of useful visual information. Consequently, important semantic features may be lost during the fusion process, leading to degraded inference accuracy. Figure 7: The relationship between fusion strength μ and accuracy on EgoSchema with Qwen2.5-VL-32B-Instruct. In contrast, when μ exceeds 0.70.7, the V cache information from kC_k becomes excessively influenced by the V cache of dC_d. Such over-fusion introduces noisy or mismatched representations into the retained cache, which contaminates the original semantic information and negatively affects the reasoning capability of the VLLM. Another notable observation from Table 7 is that the proposed method maintains relatively low variance across different fusion strengths, indicating stable inference behavior. Although several configurations achieve comparable mean accuracy, μ=0.7μ=0.7 not only obtains the best average performance but also maintains stable results across repeated experiments. Therefore, we adopt μ=0.7μ=0.7 as the default fusion strength in all remaining experiments. Table 7: Mean accuracy and variance under different fusion strengths μ on EgoSchema. Each setting is evaluated three times independently. Fusion Strength μ Mean Accuracy Variance 0.1 59.80 1.15 0.2 60.40 2.30 0.3 60.40 0.03 0.4 58.40 0.44 0.5 58.40 0.75 0.6 59.00 1.00 0.7 65.80 1.04 0.8 59.40 1.31 0.9 59.60 0.25 Appendix F Compatibility with Other VLLMs To further evaluate the generalization of VisCache beyond Qwen2.5-VL, we conduct experiments on two additional representative VLLM architectures: Qwen3-VL-4B-Instruct and LLaVA-OneVision Li et al. (2024a). These VLLMs adopt different visual encoders and multimodal fusion pipelines, providing a stronger test of architectural compatibility. F.1 Generalization to Qwen3-VL For Qwen3-VL-4B-Instruct, a different-generation VLLM architecture, VisCache maintains competitive performance while reducing FLOPs by approximately 90% compared with the full cache baseline, as shown in Table 8. Although the core ranking mechanism is backbone-agnostic, different VLLMs may exhibit different visual token distributions and varying sensitivity to visual information removal, leading to different compression–performance trade-offs. Similarly, scout models only perform temporal frame selection, and their effectiveness depends on their ability to capture task-relevant temporal cues. We further analyze failure cases of the scout-based approach and identify three common patterns: (1) missing transitional frames that contain important action changes; (2) removing ambiguous frames that are necessary for contextual reasoning; and (3) retaining redundant frames with highly similar visual content. These cases highlight the inherent trade-off between temporal compression and information preservation, motivating our conservative retention strategy. Table 8: Results of applying VisCache on the Qwen3-VL-4B-Instruct architecture. VisCache maintains competitive inference performance with minimal degradation compared to using full cache. Method R FLOPs (T) FLOPs Ratio ActCap DREAM1K NExTQA ActQA EgoSchema Avg Benchmark ROUGE-L ROUGE-L Acc Acc Acc Acc Qwen3-VL-4B-Instruct Full Cache 100% 5.08 100% 18.76 14.24 65.72 46.13 67.60 59.82 VisCache 28% 0.52 10.2% 18.66 14.34 62.90 45.35 60.22 56.16 VisCache 19% 0.51 10.0% 18.23 14.32 63.41 47.17 60.22 56.93 F.2 Generalization to LLaVA-OneVision To further demonstrate the generalization capability of VisCache, we additionally evaluate our method on LLaVA-OneVision Li et al. (2024a). Unlike Qwen2.5-VL, LLaVA-OneVision adopts a different visual encoder and multimodal fusion pipeline, making it a suitable benchmark for evaluating the architectural compatibility of the proposed visual KV cache compression framework. The corresponding results are presented in Table 9. As shown in Table 9, VisCache consistently maintains competitive inference performance under aggressive visual KV cache compression. Specifically, with only 28%28\% retained visual KV cache, VisCache reduces the FLOPs from 32.7332.73T to 8.628.62T, corresponding to only 26%26\% of the original computation cost. Despite this substantial reduction in computation, the performance degradation remains relatively limited across different benchmarks. On ActCap, VisCache even slightly improves the ROUGE-L score compared with the full-cache setting, suggesting that removing redundant visual tokens can alleviate noisy visual representations and improve inferencing quality. Meanwhile, on DREAM1K and ActQA, VisCache preserves most of the original performance while significantly reducing the computational overhead. Table 9: Results of applying VisCache on the LLaVA-OneVision architecture. It shows that VisCache maintains competitive inference performance across different VLLM architectures, exhibiting minimal degradation compared to using full cache. Method R FLOPs (T) FLOPs Ratio ActCap DREAM1K ActQA Benchmark ROUGE-L ROUGE-L Acc LLaVA-OneVision-Qwen2-7b-ov-hf Full Cache 100% 32.73 100% 5.04 14.45 38.36 VisCache 28% 8.62 26% 5.52 11.71 34.97 Overall, these experimental results demonstrate that VisCache exhibits strong architectural compatibility and can serve as a plug-and-play visual KV cache compression framework for efficient long-video VLLM inference across different VLLM families and architectures. The proposed coarse-to-fine visual KV cache compression strategy is not tightly coupled with a specific VLLM architecture and can generalize effectively to models with different visual encoders and fusion pipelines. Appendix G Quantization VisCache is orthogonal to existing KV cache quantization methods and can be seamlessly combined with them for further memory reduction. To verify this compatibility, we integrate VisCache with representative quantization methods, including KIVI Liu et al. (2024b) and FlatQuant Sun et al. (2024b), and evaluate the performance on multiple VQA and video understanding benchmarks. The results are shown in Table 10. As shown in Table 10, combining VisCache with 4-bit KV cache quantization still preserves competitive inference performance across different benchmarks. On Qwen2.5-VL-3B-Instruct, VisCache combined with FlatQuant and KIVI achieves comparable or even better performance on ActQA and EgoSchema compared with the original VisCache setting. On the larger Qwen2.5-VL-32B-Instruct backbone, although slight performance degradation is observed on several benchmarks, the overall performance remains competitive under aggressive visual KV cache compression and low-bit quantization. These results demonstrate that VisCache is highly compatible with existing KV cache quantization frameworks and can further improve inference efficiency when combined with low-bit KV cache compression. Table 10: Comparison of combining VisCache with different KV cache quantization methods on various VQA and VS datasets. Method R Bits ActCap DREAM1K NExtQA ActQA EgoSchema Avg. Benchmark ROUGE-L ROUGE-L Acc Acc Acc Acc Qwen2.5-VL-3B-Instruct Full Cache 100% 32 2.63 9.19 34.69 40.58 57.20 44.16 VisCache 28% 32 2.42 8.70 41.25 40.66 55.00 45.64 VisCache (FlatQuant) 28% 4 2.45 8.58 40.28 43.07 56.20 46.52 VisCache (KIVI) 28% 4 2.43 8.50 40.81 43.48 55.60 46.63 Qwen2.5-VL-32B-Instruct Full Cache 100% 32 2.82 7.87 60.71 46.47 65.20 57.46 VisCache 28% 32 2.72 7.29 62.08 42.69 65.80 56.86 VisCache (FlatQuant) 28% 4 2.72 7.19 54.86 43.80 59.60 52.75 VisCache (KIVI) 28% 4 2.72 7.09 54.58 43.88 58.89 52.45 Appendix H Comparison Under the Same Retention Ratio To fairly evaluate the effectiveness of different KV cache compression strategies, we compare VisCache with several representative baselines under approximately the same retention ratio (R). For each baseline, we tune its method-specific parameters to match the same global visual-KV retention ratio as VisCache, while keeping all other experimental conditions identical. In this section, we first report comprehensive comparisons on the Qwen2.5-VL-3B-Instruct backbone, and then provide additional results on the larger Qwen2.5-VL-32B-Instruct model. H.1 Results on Qwen2.5-VL-3B-Instruct We report results under two compression settings, i.e., approximately 28%28\% and 19%19\% retained visual KV cache, based on the Qwen2.5-VL-3B-Instruct backbone. Table 11 summarizes the comparison across multiple video question answering and video summarization benchmarks. As shown in Table 11, VisCache consistently achieves superior overall performance compared with existing methods while requiring significantly fewer FLOPs than the full-cache setting. Under the 28%28\% R, VisCache reduces the FLOPs from 14.8014.80T to only 2.852.85T (19%19\% of the original computation), while still achieving the best performance on most benchmarks. In particular, VisCache obtains the highest scores on DREAM1K, ActQA, and EgoSchema, and achieves competitive performance on NExTQA. The average VQA accuracy (over NExTQA, ActQA, and EgoSchema) reaches 45.6445.64, outperforming the strongest baseline result of 42.9242.92 by 2.722.72 points. Compared with methods such as PyramidKV, FastV, and PDrop, our method preserves substantially stronger video understanding capability under aggressive visual KV compression. Under the more challenging 19%19\% R, the advantage of VisCache becomes even more evident. Despite using only 15%15\% FLOPs of the full-cache setting, VisCache still achieves strong performance across multiple benchmarks and attains the best result on NExTQA while maintaining highly competitive results on DREAM1K and EgoSchema. The average VQA accuracy is 44.8544.85, exceeding the strongest baseline result of 42.5342.53 by 2.322.32 points. In contrast, existing methods suffer from more noticeable performance degradation under the same compression level, indicating that naive KV pruning or merging strategies may discard important temporal and semantic information in long-video understanding tasks. Another notable observation is that VisCache demonstrates significantly better robustness across different benchmarks and compression levels. While several baseline methods exhibit unstable behavior or severe accuracy drops when the R decreases, VisCache maintains relatively stable performance. This suggests that the proposed coarse-to-fine KV compression strategy can more effectively preserve informative visual representations and reduce redundant visual tokens without severely damaging the reasoning capability of the VLLM. Table 11: Comparison of different KV cache compression methods under the same retention ratio on Qwen2.5-VL-3B-Instruct. Bold and underlined numbers indicate the best and second-best results, respectively. Avg. Acc. denotes the average accuracy over NExTQA, ActQA, and EgoSchema. Method R ActCap DREAM1K NExTQA ActQA EgoSchema Avg. Acc. Benchmark ROUGE-L ROUGE-L Acc Acc Acc Acc Full Cache 100% 2.63 9.19 34.69 40.58 57.20 44.16 Q-Frame 28% 2.37 9.69 40.98 38.60 46.60 42.06 PyramidKV 28% 2.37 7.34 38.54 40.21 50.00 42.92 FastV 28% 2.43 8.59 39.60 39.58 47.60 42.26 PDrop 28% 2.42 8.62 38.90 37.34 49.60 41.95 VisCache 28% 2.45 8.70 41.25 40.66 55.00 45.64 Q-Frame 19% 2.37 7.62 42.08 38.27 46.40 42.25 PyramidKV 19% 2.36 7.34 36.38 40.05 48.50 41.64 FastV 19% 2.42 8.13 38.40 37.20 52.00 42.53 PDrop 19% 2.64 7.69 37.60 37.12 49.00 41.24 VisCache 19% 2.39 8.26 41.44 38.51 54.60 44.85 H.2 Results on Qwen2.5-VL-32B-Instruct To further examine whether the advantage of VisCache persists at a larger model scale, we compare it with representative baselines on the Qwen2.5-VL-32B-Instruct backbone. The comparison is conducted under the same 28%28\% and 19%19\% R settings, and the results are reported in Table 12. As shown in Table 12, VisCache maintains a clear advantage under both retention ratios. At 28%28\% R, VisCache reduces the FLOPs to 10.7510.75T, corresponding to only 12%12\% of the full-cache computation, while achieving the best DREAM1K ROUGE-L of 7.297.29 and the best EgoSchema accuracy of 65.8065.80. In comparison, Q-Frame and PDrop require 45%45\% and 22%22\% FLOPs, respectively, yet still fall behind on both metrics. At the more aggressive 19%19\% R, VisCache uses only 10%10\% of the original FLOPs and achieves the highest EgoSchema accuracy of 65.0065.00, substantially outperforming Q-Frame (56.0056.00) and PDrop (57.0057.00). On DREAM1K, Q-Frame obtains a slightly higher ROUGE-L (7.897.89) than VisCache (7.357.35), but with 2.5×2.5× more FLOPs and a much larger performance drop on EgoSchema. These results confirm that VisCache achieves a substantially better trade-off between computational efficiency and performance even on larger VLLMs. Table 12: Comparison of representative KV cache compression methods under the same retention ratio on Qwen2.5-VL-32B-Instruct. Bold and underlined numbers indicate the best and second-best results, respectively. Method R FLOPs FLOPs DREAM1K EgoSchema (T) Ratio ROUGE-L Acc Qwen2.5-VL-32B-Instruct Full Cache 100% 93.08 100% 7.87 65.20 Q-Frame 28% 41.47 45% 7.16 62.00 PDrop 28% 20.15 22% 5.16 64.00 VisCache 28% 10.75 12% 7.29 65.80 Q-Frame 19% 23.54 25% 7.89 56.00 PDrop 19% 19.51 21% 5.07 57.00 VisCache 19% 9.08 10% 7.35 65.00 Overall, these results demonstrate that VisCache achieves a substantially better trade-off between computational efficiency and performance under matched retention budgets across different model scales, enabling efficient long-video VLLM inference under highly constrained KV cache budgets. Appendix I Performance v.s. Retention Ratio To evaluate the trade-off between memory efficiency and generation quality, we measure the performance on DREAM1K and EgoSchema under different visual KV cache RRs based on Qwen2.5-VL-3B-Instruct backbone. Figure 8 reports the results, where the dashed lines denote the full-cache baselines for each benchmark. For DREAM1K, as the retention ratio increases, the ROUGE-L score steadily improves and eventually surpasses the full cache result at around 60% retention, reaching 10.78 at 90%. This suggests that keeping more visual KV cache is beneficial for detailed description tasks, and a moderate visual KV cache budget can already outperform the unpruned setting, likely because redundant cache introduces noise. In contrast, EgoSchema accuracy first rises and then declines after a R of 70%, achieving a maximum of 58.60%. The drop at higher ratios indicates that preserving all visual KV cache may include distracting information for VQA, while an appropriate pruning ratio helps retain the most relevant cues. Overall, the results highlight that optimal RRs are task-dependent, and a well-chosen ratio can even exceed full-cache performance while reducing memory cost. Figure 8: Performance on DREAM1K and EgoSchema under different visual KV cache retention ratios. The dashed lines indicate the full cache performance for each benchmark. Appendix J Combination of baseline and Temporal Redundancy Filtering (TRF). To further validate the flexibility and compatibility of the proposed Scout-based Temporal Redundancy Filtering (TRF), we integrate it with representative visual token pruning methods, including FastV and PDrop, under different retained token ratios. The experiments are conducted on the DREAM1K and EgoSchema benchmarks based on the Qwen2.5-VL-3B-Instruct backbone. The corresponding results are summarized in Table 13. The results demonstrate that Scout-based TRF can be seamlessly combined with existing visual token pruning strategies and generally improves or preserves performance under aggressive token compression settings. Specifically, under the 28%28\% R, integrating TRF with PDrop improves the EgoSchema accuracy from 49.6049.60 to 52.8052.80, while maintaining nearly identical DREAM1K performance. Similarly, Scout + FastV also improves EgoSchema accuracy compared with the original FastV baseline. These results indicate that TRF can effectively complement token-level pruning strategies by removing redundant temporal information while preserving important semantic content. Table 13: Compatibility analysis of Scout-based Temporal Redundancy Filtering (TRF) with existing visual token pruning methods. Method R DREAM1K EgoSchema ROUGE-L Acc FastV 28% 8.59 47.60 PDrop 28% 8.62 49.60 Scout + FastV 28% 8.49 48.80 Scout + PDrop 28% 8.56 52.80 FastV 19% 8.13 52.00 PDrop 19% 7.69 49.00 Scout + FastV 19% 8.42 49.60 Scout + PDrop 19% 8.29 49.40 Under the more challenging 19%19\% R, the effectiveness of Scout-based TRF becomes more evident on DREAM1K. Compared with the original FastV and PDrop methods, integrating TRF improves the ROUGE-L score from 8.138.13 to 8.428.42 and from 7.697.69 to 8.298.29, respectively. Although slight fluctuations are observed on EgoSchema, the overall performance remains competitive under highly constrained token budgets. This suggests that Scout-based TRF can alleviate the severe information loss caused by aggressive token pruning and improve the robustness of compressed VLLM inference. Another important observation is that the proposed TRF module is architecture-agnostic and can function as a lightweight plug-in component for existing KV cache compression or visual token pruning frameworks. Instead of replacing previous methods, Scout-based TRF serves as a complementary temporal filtering strategy that further reduces redundant visual information across video frames. Therefore, the proposed method exhibits strong extensibility and practical applicability for efficient long-video VLLM inference. Appendix K Specific Inference Time Analysis Table 14: Detailed inference time analysis under different KV cache retention ratios. We report the end-to-end (E2E) latency, time to first token (TTFT), and time per output token (TPOT) on DREAM1K and EgoSchema benchmarks based on Qwen2.5-VL-3B-Instruct backbone. Our VisCache consistently reduces inference latency as the KV cache R decreases, achieving up to 2.35× E2E speedup while maintaining efficient generation performance. Method R E2E Latency (s) TTFT (s) TPOT (ms) DREAM1K Full Cache 100% 12.24 4.42 118.81 VisCache 40% 9.55 (1.28×) 2.05 119.06 VisCache 28% 6.33 (1.93×) 3.45 45.79 VisCache 19% 5.20 (2.35×) 2.23 45.60 EgoSchema Full Cache 100% 13.92 13.50 59.37 VisCache 40% 10.47 (1.33×) 10.08 56.82 VisCache 28% 9.32 (1.49×) 8.97 49.50 VisCache 19% 8.30 (1.68×) 7.92 47.70 Table 14 presents a detailed latency breakdown of VisCache under different visual KV cache R. Across both DREAM1K and EgoSchema benchmarks, reducing the KV cache R consistently decreases overall inference latency, demonstrating the effectiveness of VisCache in accelerating VLLM inference. Specifically, on DREAM1K, VisCache achieves up to 2.35× E2E speedup at 19% R, while TPOT is reduced from 118.81 ms to 45.60 ms. Similarly, on EgoSchema, VisCache reduces E2E latency from 13.92 s to 8.30 s and improves decoding efficiency by lowering TPOT from 59.37 ms to 47.70 ms. Overall, the results demonstrate that VisCache achieves an effective balance between visual KV cache compression and practical system efficiency, providing substantial inference acceleration while maintaining stable decoding performance. Appendix L Impact of the Scout VLM Table 15: Effect of different scout VLMs for temporal filtering. We keep R=28%R=28\% of the visual KV cache on Qwen2.5-VL-32B-Instruct across all variants. Method Scout ActCap DREAM1K NExTQA ActQA EgoSchema Model ROUGE-L ROUGE-L Acc Acc Acc Full Cache – 2.63 9.19 34.69 40.58 57.20 VisCache CLIP 2.42 8.70 41.25 40.66 55.00 BLIP 2.44 8.49 37.18 40.81 46.80 OpenCLIP 2.43 8.51 41.56 39.78 53.40 We further investigate how the choice of the scout model affects overall performance. Table 15 compares CLIP Radford et al. (2021), BLIP Li et al. (2022), and OpenCLIP Cherti et al. (2023) as the frame-filtering scout under identical VisCache settings. The results reveal that the scout choice has a non-trivial impact on downstream VQA accuracy, with CLIP and OpenCLIP notably outperforming BLIP on several benchmarks, while all three scouts perform comparably on the VS tasks. This suggests that the semantic alignment between the scout’s visual encoder and the main LLM backbone plays an important role in selecting frames that facilitate accurate question answering. We adopt CLIP as the default scout for its balanced performance across both task types. Notably, even with the least effective scout, VisCache retains a substantial portion of full-cache accuracy while reducing the KV cache to 28%, confirming that the dual-stage framework is robust to the choice of the pre-filtering model. The rationale and compatibility of the scout VLM are further discussed in Appendix C and Appendix J. Appendix M Inference Example Figure 9: Comparison between Full Cache and VisCache on a pull-up activity video. Figure 10: Comparison between Full Cache and VisCache on a movie scene video. User: Describe the content of this video in detail. Full Cache: A young man is performing pull-ups on a bar in a gym. He is wearing a black T-shirt and black shorts. The gym has white walls and a ceiling with exposed pipes. The man starts by hanging from the bar with his arms fully extended, then pulls himself up until his chin is over the bar. VisCache: The video begins with a man standing in front of a door, preparing to perform a pull-up exercise. He is wearing a black tank top and shorts, and his hair is tied back. The setting appears to be an indoor gym or workout area, with a metal pull-up bar mounted on a wall. User: Describe the content of this video in detail. Full Cache: The video begins with a close-up shot of a woman with long brown hair, looking off to the side with a serious expression. The scene then cuts to a man with short dark hair, wearing a blue shirt and a black vest, who is looking at something off-screen with a concerned expression. The camera then cuts to a woman in a white shirt and shorts walking through a room, her back facing the camera. She turns around and walks towards a bed, where she bends over and picks up a suitcase. The camera then cuts to a close-up of a woman’s face, who appears to be crying. VisCache: The video begins with a close-up shot of a woman’s face, showing her looking out of a window. The scene then transitions to a man walking through a doorway into a room. He is wearing a white shirt and appears to be in motion. The camera follows him as he moves through the room, which has a modern design with a large window that offers a view of the outdoors. The man continues to walk towards the camera, and the video ends with him standing in front of it.