Paper deep dive
A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training
Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/29/2026, 3:57:37 AM
Summary
The paper introduces Multimodal Unsupervised Continual Post-Training (MU-CPT), a framework for evolving Multimodal Large Language Models (MLLMs) using streaming unlabeled data. It addresses cross-modal catastrophic forgetting and language bias by proposing the Visual Dependence-Aware (VDA) framework. VDA utilizes Visually Constrained Optimal Transport (VC-OT) to preserve the structural integrity of visual dependencies from old tasks and Visually Modulated Adaptation (VMA) to emphasize visually grounded learning for new tasks, thereby balancing stability and plasticity.
Entities (7)
Relation Signals (6)
VDA → contains → VC-OT
confidence 97% · VDA framework with two main components. First, Visually Constrained Optimal Transport (VC-OT)...
VDA → contains → VMA
confidence 97% · Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity...
MU-CPT → uses → VDA
confidence 96% · Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.
VC-OT → mitigates → cross-modal catastrophic forgetting
confidence 95% · formulates the VD structural distortion... to mitigate cross-modal forgetting.
VMA → promotes → new-task plasticity
confidence 94% · promoting new-task plasticity.
Visual Dependence (VD) → indicates → cross-modal catastrophic forgetting
confidence 92% · its structural distortion serves as an indicator of cross-modal catastrophic forgetting
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting, and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA.
Tags
Links
- Source: https://arxiv.org/abs/2608.26095v1
- Canonical: https://arxiv.org/abs/2608.26095v1
Trouble viewing inline? Open PDF directly →
Full Text
48,464 characters extracted from source content.
Expand or collapse full text
A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training Kaichen Li Zhilin Zhu Jianhao Huang Zhengqin Lai Baochen Xiong Zibo Shao Yaguang Song Linhui Xiao Xiaoshan Yang Changsheng Xu Abstract In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised post-training methods for MLLMs typically optimize target tokens uniformly, overlooking their heterogeneous visual dependence (VD). However, we reveal that token-level VD is crucial for MU-CPT. Specifically, its structural distortion serves as an indicator of cross-modal catastrophic forgetting and its inherent heterogeneity acts as a compass to guide new-task learning. Leveraging this property, we propose a Visual Dependence-Aware (VDA) framework with two main components. First, Visually Constrained Optimal Transport (VC-OT) formulates the VD structural distortion of old-task VD during new-task learning as an optimal transport problem to mitigate cross-modal forgetting. By designing a region-aware ground cost and a dependence-stratified transport penalty, it prevents global shifts in visual focus, while strictly prohibiting visual reliance from degenerating into language bias. Second, Visually Modulated Adaptation (VMA) exploits VD heterogeneity to emphasize visually grounded new-task learning, promoting new-task plasticity. Together, our method simultaneously maintains old-task stability and new-task plasticity during challenging MU-CPT. Extensive experiments under our MU-CPT setting validate the effectiveness of VDA. 1State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China 2Pengcheng Laboratory, Shenzhen, China 3School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China 4Harbin Institute of Technology, Shenzhen, China Introduction Multimodal Large Language Models (MLLMs) (Li et al. 2024; Bai et al. 2025b; Bai et al. 2025a; Wang et al. 2025) have achieved remarkable progress by integrating visual perception with language understanding, enabling them to handle complex multimodal inputs and instructions. To further improve MLLMs’ performance on specific tasks or user demands, most existing approaches, such as supervised fine-tuning (SFT) and reinforcement learning (RL), still largely rely on carefully curated and human-annotated multimodal data. As the scope and diversity of real-world applications continue to expand, empowering MLLMs through large-scale human annotation becomes prohibitively expensive and difficult to sustain. This has motivated growing interest in unsupervised post-training for MLLMs (Zhou et al. 2024; Wei et al. 2026), which aims to leverage unlabeled multimodal data for model improvement. Figure 1: Comparison between (a) Multimodal Continual Instruction-Tuning and (b) Multimodal Unsupervised Continual Post-Training (MU-CPT), where labeled answers are completely unavailable during training. Although these unsupervised approaches have made important progress (Zhu et al. 2024a; Singh et al. 2026), they are primarily developed under static, closed-world settings, where unlabeled data are collected in advance from a predefined distribution that remains unchanged (Zhu et al. 2024b; Zhu et al. 2026b). However, real-world applications of MLLMs are inherently dynamic (Zhou et al. 2024; Zhu et al. 2024a; Xu et al. 2026). MLLMs deployed in open environments continuously encounter evolving scenarios (Zhu et al. 2026a), yielding non-stationary streams of unlabeled multimodal data. Ideally, MLLMs should continually strengthen their capabilities in emerging scenarios while preserving those learned from earlier experience (Shao et al. 2026b; Shao et al. 2026a). However, existing unsupervised post-training methods (Hu et al. 2025) struggle with non-stationary data distributions. Meanwhile, existing multimodal continual instruction-tuning (MCIT) (Chen et al. 2024; Guo et al. 2025) approaches assume access to abundant and expensive labeled multimodal data, as illustrated in Fig. 1(a). Bridging these two paradigms is therefore essential to enable sustained evolution of MLLMs through unlabeled multimodal data streams. To bridge this gap, we formalize a novel setting termed Multimodal Unsupervised Continual Post-Training (MU-CPT). As illustrated in Fig. 1(b), under MU-CPT setting, MLLMs sequentially encounter distinct multimodal tasks without access to the corresponding answers during continual learning. The model is expected to acquire capabilities for each new task and preserve those learned from previous tasks. Existing unsupervised post-training research has primarily concentrated on how to construct optimization target sequences (Zhu et al. 2024a; Wei et al. 2026; Yu et al. 2026). However, these approaches invariably treat the constructed target sequences as monolithic entities, implicitly optimizing all target tokens uniformly, thereby leaving a fine-grained property intrinsic to the learning signals themselves largely unexamined: token-level Visual Dependence (VD) heterogeneity (Ye et al. 2026). By measuring how much real visual input increases a token’s log-likelihood relative to an information-free counterfactual visual input, we observe that text tokens in multi-modal data naturally bifurcate into two regimes (see Figure 2 (a)): visually grounded entities exhibit pronounced VD, whereas language-driven contextual markers exhibit negligible VD. Crucially, as shown in Figure 2(b), we further identify that sequential unsupervised adaptation induces a severe failure mode which we formalize as cross-modal attributional forgetting: the once-concentrated VD on visually grounded tokens (e.g., "photo", "phone") is progressively flattened and redistributed toward language-driven tokens (e.g., "object", "and"), and this structural drift co-occurs with the model abandoning the correct, visually grounded answer for an hallucinated one. Figure 2: (a) Text tokens in multimodal data exhibit heterogeneous levels of visual dependence, where red and grey denote high and low visual dependence, respectively. (b) Learning new tasks leads to structural distortion of visual dependence on old tasks. Motivated by these observations, we propose a novel Visual Dependence-Aware (VDA) framework, leveraging token-level VD as a unified basis to harmonize old-task stability and new-task plasticity, a fundamental challenge in continual learning (Parisi et al. 2018; De Lange et al. 2021). Our framework comprises two synergistic components. First, we innovatively formulate the mitigation of cross-modal attributional forgetting as a Visually Constrained Optimal Transport (VC-OT) problem, which jointly preserves the overall strength and the structural organization of visual dependencies. Specifically, we explicitly structure the optimal transport matrix through two novel mechanisms: a region-aware cost that regulates internal VD redistribution based on visual spatial similarity, and a cross-set penalty that strictly blocks the erroneous transport of visual reliance toward vision-independent tokens. Beyond preserving old-task visual dependencies, effective continual learning also requires sufficient plasticity to acquire new capabilities. Accordingly, we further introduce Visually Modulated Adaptation (VMA), which promotes new-task learning by emphasizing visually grounded tokens while attenuating task-specific language patterns that may interfere with previous tasks. Consequently, these designs enable our framework to maintain strong stability without compromising plasticity throughout the MU-CPT process. Our contributions are summarized as follows: (1) We formalize Multimodal Unsupervised Continual Post-Training for MLLMs (MU-CPT), a novel and practical setting that enables MLLMs to continuously evolve from non-stationary, unlabeled multimodal data streams. (2) We uncover the phenomenon of cross-modal attributional forgetting caused by token-level visual dependence distortion, establishing VD as a unified signal to balance stability and plasticity. (3) We propose the VDA framework, introducing Visually Constrained Optimal Transport to preserve structural visual grounding and Visually Modulated Adaptation to mitigate linguistic bias. (4) Comprehensive experiments across diverse multimodal tasks validate the effectiveness of our method, setting a strong baseline for unsupervised MLLM evolution. Related Work Continual Learning. Continual learning (CL) seeks to balance new-task plasticity and old-task stability, with representative solutions based on regularization (Kirkpatrick et al. 2017; Li and Hoiem 2017), rehearsal (Rebuffi et al. 2017; Lopez-Paz and Ranzato 2017), and parameter isolation (Mallya and Lazebnik 2017); continual self-supervised learning further extends CL to unlabeled streams, mainly for unimodal representation learning (Madaan et al. 2021). Multimodal CL subsequently studies sequential VQA and prompt-based adaptation (Zhang et al. 2023a; Qian et al. 2023). More recently, continual instruction tuning (CIT) extends CL to instruction-following MLLMs trained sequentially on labeled image–instruction–answer data, with EProj and CoIN establishing representative benchmarks (He et al. 2026; Chen et al. 2024). Existing CIT methods mitigate forgetting through selective attention distillation in SEEKR-MLLM (He et al. 2024), global–local expert adaptation in CL-MoE (Huai et al. 2025), old-task-important LoRA regularization in SEFE (Chen et al. 2025), or replay-augmented gradient guidance in DGG (Li et al. 2025). However, these methods rely on labeled answers to define learning and preservation targets, whereas MU-CPT must balance plasticity and stability using only answer-unlabeled multimodal streams. Unsupervised Post-Training. Unsupervised post-training adapts models to unlabeled inputs using self-generated learning signals. For LLMs, LSMI selects high-confidence solutions through self-consistency (Huang et al. 2023), ScPO constructs preferences from response consistency (Prasad et al. 2024), and TLM minimizes input perplexity (Hu et al. 2025). Related studies also exploit semantic or token-level entropy (Zhang et al. 2026; Agarwal et al. 2026). For MLLMs, SeVa contrasts responses generated from original and perturbed images (Zhu et al. 2024a), CSR calibrates self-rewards using image–response relevance (Zhou et al. 2024), M-UPT derives pseudo-rewards through majority voting (Wei et al. 2026), and TTRV constructs rewards from response frequency and output entropy (Singh et al. 2026); CSRS further stabilizes self-rewarding through retraced sampling and softened frequency rewards (Yu et al. 2026). Although these methods provide diverse unsupervised objectives, they mainly assume static unlabeled datasets and overlook both continual forgetting and token-level VD heterogeneity. Methodology Problem Formulation We formulate Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling models to continually evolve on unlabeled multimodal data while preserving previously acquired capabilities, unlike standard continual learning that relies on fully annotated data. Formally, under this setting, we consider an instruction-tuned MLLM, initially parameterized by θ0 _0, continually learning from K multimodal tasks. At stage k, the model only accesses an answer-unlabeled multimodal data set k=(Iki,Qki)i=1NkD_k=\(I_k^i,Q_k^i)\_i=1^N_k, where IkiI_k^i and QkiQ_k^i denote the i-th image and its associated question, respectively. The overall objective is to acquire capabilities for the current task from kD_k while preventing catastrophic forgetting on previously learned tasks, yielding the updated parameters θk _k. Upon completing stage k, the updated model θk _k is evaluated on held-out image-question pairs from all tasks encountered so far. Regardless of the specific form of the unsupervised signal, we denote the optimization target sequence of length L for an (I,Q)(I,Q) pair as Y=(y1,…,yL)Y=(y_1,…,y_L), where t∈1,…,Lt∈\1,…,L\ indexes the target-token position, and let ℓt(θ) _t(θ) denote the loss associated with token yty_t. Framework Overview While existing unsupervised methods mainly construct sequence-level targets, they typically optimize all target tokens uniformly, overlooking token-level visual dependence (VD) heterogeneity. In continual learning, this variation provides an intrinsic anchor for balancing stability and plasticity. Accordingly, VDA addresses MU-CPT at the token level rather than only at the sequence level, as illustrated in Figure 3. To maintain old-task stability, Visually Constrained Optimal Transport (VC-OT) formulates old-task VD distortion as an optimal transport problem. Its region-aware cost and dependence-stratified penalty suppress global shifts in visual focus and transfer toward vision-independent tokens. To preserve new-task plasticity, Visually Modulated Adaptation (VMA) emphasizes visually grounded tokens while attenuating task-specific language patterns that cause cross-task interference. Together, these complementary mechanisms enable robust continual learning from unlabeled multimodal data. Token-Level Visual Dependence To quantify the visual dependence of each target token yt∈Yy_t∈ Y, we compare its prediction under real and counterfactual visual conditions. Specifically, keeping the target token and its preceding textual context unchanged, we compute: ℓtreal(θ) _t^real(θ) =−logpθ(yt∣I,,y<t), =- p_θ (y_t I,P,y_<t ), (1) ℓtnull(θ) _t^null(θ) =−logpθ(yt∣Inull,,y<t), =- p_θ (y_t I_null,P,y_<t ), where InullI_null denotes a counterfactual visual input with no information. (See Appendix for details) The token-level visual dependence is then defined as: Dt(θ)=ℓtnull(θ)−ℓtreal(θ)ℓtnull(θ)+ℓtreal(θ).D_t(θ)= _t^null(θ)- _t^real(θ) _t^null(θ)+ _t^real(θ). (2) A positive DtD_t indicates that the visual input actively facilitates the prediction of the token, whereas a non-positive value suggests that the token is driven primarily by language context. During continual learning, this token-level metric serves a dual purpose by explicitly bridging visual dependence with the model’s core cross-modal attribution capability. For replayed old-task samples, the structural distortion of VD reflects the collapse of this capability, marking a severe cross-modal attributional forgetting that necessitates VD structural preservation to maintain stability. On the other hand, while pursuing stability alone risks suppressing model adaptability, VD heterogeneity acts as an intrinsic compass to locate visually informative tokens. By directing the optimization focus toward these essential visual cues rather than task-specific language biases, it effectively fosters new-task plasticity. This dual functionality establishes VD as a unified foundation for our subsequent designs. Figure 3: Overview of the proposed Visual Dependence-Aware (VDA) framework. (a) Visually Constrained Optimal Transport (VC-OT) preserves old-task VD strength and structure through region-aware and dependence-stratified transport. (b) Visually Modulated Adaptation (VMA) reweights incoming-task tokens according to their VD to promote visually grounded learning. Together, they effectively balance stability and plasticity in MU-CPT. Visually Constrained Optimal Transport Preserving old-task VD is essential to prevent cross-modal attributional forgetting. To this end, we formulate the structural drift between the reference and current old-task VD as a unified optimal transport problem. By leveraging the highly non-uniform damage that VD redistribution in different directions inflicts on the model’s cross-modal attribution capability, we introduce a region-aware transport cost and dependence-stratified transport penalty to safely regulate internal VD structural shifts. Furthermore, we explicitly constrain the overall VD strength to prevent the global collapse of visual reliance. VD Transport Formulation. Since only Dt>0D_t>0 signifies that visual evidence actively supports token generation, we focus exclusively on positive visual dependence. To extract this non-negative mass for optimal transport while preserving smooth gradients, we smoothly approximate the hard max(⋅,0) (·,0) truncation with a low-temperature Softplus, projecting VD into nonnegative mass: rtm=τDSoftplus(DtmτD),m∈ref,cur.r_t^m= _DSoftplus ( D_t^m _D ), m∈\ref,cur\. (3) Here, a size-bounded buffer retains a small subset of old-task samples and their reference VD. For each replayed sample, DtrefD_t^ref is computed by the model immediately after learning its source task and cached in the buffer, while DtcurD_t^cur is computed by the current trainable model. While this absolute mass rtmr_t^m will be explicitly constrained later to maintain the overall VD strength, modeling its internal structural flow requires a relative mass distribution. Therefore, we normalize rtmr_t^m within the target sequence: r~tm=rtm∑u=1Lrum,m∈ref,cur. r_t^m= r_t^m _u=1^Lr_u^m, m∈\ref,cur\. (4) where L represents the length of the target sequence. We treat ~ref r^ref as the source distribution and ~cur r^cur as the target distribution. Specifically, TijT_ij denotes the VD mass transported from source token i under the reference model to destination token j under the current model, satisfying ∑j=1LTij=r~iref _j=1^LT_ij= r_i^ref and ∑i=1LTij=r~jcur _i=1^LT_ij= r_j^cur. To set up the transport space, we explicitly separate tokens attributed to visual evidence from those driven by language context, partitioning the target sequence into a visually attributed set ℋ=t∣Dtref>0H=\t D_t^ref>0\ and a vision-independent set ℒ=t∣Dtref≤0L=\t D_t^ref≤ 0\. Given that the reference mass in ℒL is negligible and does not represent any visual attribution that needs to be preserved, we focus on the structural regularization on transport originating from ℋH. Accordingly, we consider two relevant transport paths: intra-set redistribution within ℋH and cross-set transfer from ℋH to ℒL. Region-Aware Transport Cost. For the intra-set transport within ℋH, our primary concern is to prevent global shifts in visual focus, where the model’s attention drifts entirely from one region to an unrelated area. This drift signifies a severe forgetting of how the model comprehends the image. Therefore, we permit mild redistribution among visually similar tokens while penalizing such global shifts. Let t,refℓa_t,ref denote the normalized attention distribution of token t∈1,…,Lt∈\1,…,L\ over all visual tokens at decoder layer ℓ , extracted by the model immediately after learning the sample’s source task and cached with the replay sample. We formulate the region-aware transport cost between tokens i and j as: dreg(i,j)=1||∑ℓ∈JSD(i,refℓ,j,refℓ).d_reg(i,j)= 1|G| _ JSD (a_i,ref ,a_j,ref ). (5) Here, G specifies the selected intermediate decoder layers. JSD∈[0,1]JSD∈[0,1] represents Jensen-Shannon Divergence, which is used to measure the discrepancy in attended regions, a greater difference directly yields a higher transport cost. Dependence-Stratified Transport Penalty. We next consider cross-set transport from ℋH to ℒL. Since tokens in ℒL are vision-independent under the reference model, assigning VD mass to these positions indicates spurious visual attribution to language-driven tokens. We therefore penalize ℋ→ℒH\!→\!L transport with the maximum JSD cost of 1, corresponding to completely non-overlapping visual attention distributions. Structural Transport Objective. Combining the region-aware intra-set cost and the dependence-stratified cross-set penalty, we define the transport cost matrix as: Cij=dreg(i,j),i∈ℋ,j∈ℋ,1,i∈ℋ,j∈ℒ,0,i∈ℒ.C_ij= casesd_reg(i,j),&i ,\ j ,\\[3.0pt] 1,&i ,\ j ,\\[3.0pt] 0,&i . cases (6) Since transport originating from ℒL carries negligible reference mass and no meaningful visual attribution, we set its cost to zero. Let ∗T^* denote the optimal transport plan obtained through Sinkhorn iterations. The structural alignment loss is: ℒstruct=⟨∗,⟩=∑i=1L∑j=1LTij∗Cij.L_struct= ^*,C = _i=1^L _j=1^LT^*_ijC_ij. (7) VD Strength Preservation. While ℒstructL_struct regulates the internal VD organization, learning new tasks may still reduce the overall VD strength. We therefore match the total positive VD mass of the current model to its reference: ℒstrength=|∑t=1Lrtcur−∑t=1Lrtref|∑t=1Lrtref.L_strength= | _t=1^Lr_t^cur- _t=1^Lr_t^ref | _t=1^Lr_t^ref. (8) The complete VC-OT objective is: ℒVC-OT=ℒstruct+ℒstrength.L_VC -OT=L_struct+L_strength. (9) Visually Modulated Adaptation Although VC-OT preserves old-task knowledge, anti-forgetting constraints alone can restrict new-task plasticity (Jung et al. 2020). To balance the two, we introduce Visually Modulated Adaptation (VMA). During new-task training, visually attributed tokens carry task-relevant visual cues, whereas vision-independent tokens often encode task-specific linguistic templates, such as multiple-choice formats. Overfitting to these templates can induce language bias and cross-task interference; for example, the model may generate option labels for previously learned open-ended questions. VMA modulates each token’s learning signal according to its visual dependence, thereby emphasizing visually grounded knowledge to improve plasticity while suppressing language-bias interference. Let ℓtu(θ) _t^u(θ) denote the unmodulated token-level loss associated with target token yty_t. Based on its current visual dependence defined in Eq. 2, we first compute: w~t w_t =1+λVMAReLU(sg[Dt(θ)]), =1+ _VMAReLU (sg\! [D_t(θ) ] ), (10) wt w_t =Lw~t∑u=1Lw~u, = L w_t _u=1^L w_u, where sg[⋅]sg[·] denotes the stop-gradient operation and λVMA _VMA controls the modulation strength. The unnormalized modulation factor w~t w_t amplifies the learning intensity strictly for visually attributed tokens, while retaining a base unit contribution for the rest. Sequence-level normalization ensures the average modulation scale remains neutral, avoiding sample-dependent optimization shifts. The modulated current-task objective is then formulated as: ℒVMA=1L∑t=1Lwtℓtu(θ).L_VMA= 1L _t=1^Lw_t\, _t^u(θ). (11) By adaptively modulating the relative contribution of visually attributed tokens, VMA encourages the model to absorb effective visual knowledge from the incoming task, while attenuating the influence of task-specific linguistic formatting. Consequently, VMA complements VC-OT by enhancing new-task plasticity without relying on external supervision. To sum up, the overall loss is given by: ℒtotal=ℒVMA+ℒVC-OT.L_total=L_VMA+L_VC -OT. (12) Experiments Experimental Setup Datasets. We evaluate six tasks spanning OCR, scientific diagrams, financial charts, compositional reasoning, driving scenes, and medical VQA: TextVQA (Singh et al. 2019), SciVQA (Borisova et al. 2025), StockQA (Zhao et al. 2025), GQA (Hudson and Manning 2019), DriveLM (Sima et al. 2024), and PMC-VQA (Zhang et al. 2023b). We follow the original splits and use only image–question pairs for continual training. All methods use the same random order TextVQA→ → → → → -VQA (See Appendix for more random orders) and are evaluated on all seen test sets after each stage. Method TextVQA SciVQA StockQA GQA DriveLM PMC-VQA AvgAcc ↑ Frozen Backbone Qwen2.5-VL-7B 79.1 52.9 53.3 60.0 24.7 53.5 53.9 Unsupervised Post-training Methods LSMI (EMNLP’23) 79.5 63.2 62.5 68.3 29.0 49.7 58.7 SeVa (ACM M’24) 79.9 57.3 60.7 64.0 27.3 54.3 57.2 ScPO (ICML’25) 80.3 61.7 64.4 66.0 28.6 51.6 58.8 TLM (ICML’25) 77.2 49.3 43.3 46.1 29.8 42.4 48.0 M-UPT (NeurIPS’25) 78.8 54.7 57.8 62.4 24.8 55.0 55.6 TTRV (CVPR’26) 79.6 60.5 59.2 64.4 26.7 55.9 57.7 Continual Learning Methods SEEKR-MLLM (EMNLP’24) 75.4 61.0 67.8 68.8 31.7 55.8 60.1 CL-MoE (CVPR’25) 78.5 60.5 62.5 64.1 27.7 54.2 57.9 SEFE (ICML’25) 79.3 33.1 72.5 56.8 26.5 53.1 53.6 DGG (CVPR’26) 84.4 54.5 61.8 61.7 28.5 54.6 57.6 VDA (Ours) 85.9 63.6 70.3 67.6 32.6 54.9 62.5 Table 1: Comparison with existing Unsupervised Post-Training methods and Continual learning methods. Comparison Methods. We compare with six unsupervised post-training methods—LSMI (Huang et al. 2023), SeVa (Zhu et al. 2024a), ScPO (Prasad et al. 2024), TLM (Hu et al. 2025), M-UPT (Wei et al. 2026), and TTRV (Singh et al. 2026)—and four MLLM continual-learning methods—DGG (Li et al. 2025), SEFE (Chen et al. 2025), CL-MoE (Huai et al. 2025), and SEEKR-MLLM (He et al. 2024). Implementation Details. We use Qwen2.5-VL-7B (Bai et al. 2025b) as the primary backbone (See appendix for other model families). For VDA and all continual learning baselines, we perform parameter-efficient tuning using LoRA with a rank of 128. For the unsupervised post-training baselines, we follow the training configurations provided in their official implementations. Following TLM (Hu et al. 2025), we use image-conditioned question modeling as the basic unsupervised objective, where the model autoregressively predicts question tokens conditioned on the image and a task-agnostic instruction as in (Zhao et al. 2024). Our VDA is not tied to this objective, with results using other unsupervised signals reported in the appendix. As to other hyperparameters, we set the VD smoothing temperature τD _D to 0.050.05, and buffer size to 1000. The reference VD and visual attention of each replay sample is computed once when the sample is inserted into the memory and remains fixed thereafter. Unless otherwise specified, we use AdamW for optimization with a batch size of 1 and a learning rate of 5×10−55× 10^-5. Evaluation Metrics. Following standard continual-learning protocols (Wang et al. 2022; Smith et al. 2023), AvgAcc averages final accuracy over all tasks and is our primary stability–plasticity metric. AvgF measures the average performance drop on previous tasks, while AvgLA averages each task’s accuracy immediately after it is learned, diagnosing stability and plasticity, respectively. Method AvgAcc ↑ AvgLA ↑ AvgF ↓ SEEKR-MLLM 60.1 62.5 2.9 CL-MoE 57.9 59.5 1.8 SEFE 53.6 61.4 9.4 DGG 57.6 60.8 3.9 VDA (Ours) 62.5 63.0 0.6 Table 2: Fine-grained comparison of new-task plasticity and old-task stability with continual learning methods. Comparison Results Table 1 reports the final performance after sequential adaptation on six multimodal tasks. VDA achieves the highest AvgAcc of 62.5%, outperforming the strongest unsupervised post-training method, ScPO, by 3.7 points and the strongest continual learning baseline, SEEKR-MLLM, by 2.4 points. Notably, despite learning solely from answer-unlabeled multimodal data under continually shifting data distribution, VDA improves the frozen Qwen2.5-VL-7B backbone by 8.6 points, demonstrating its ability to effectively acquire and keep knowledge from non-stationary unlabeled data streams. Across individual tasks, existing baselines exhibit pronounced performance variation, indicating that their adaptation is sensitive to task and domain shifts under answer-free supervision. In contrast, VDA maintains consistently competitive performance across all six benchmarks, demonstrating more robust unsupervised continual learning across heterogeneous multimodal domains. Furthermore, we conduct a fine-grained comparison with existing supervised CL methods from the perspectives of new-task plasticity and old-task stability, where all methods are adapted to the MU-CPT setting using the same unsupervised learning signal. As shown in Table 2, VDA achieves the highest AvgLA of 63.0% and the lowest forgetting of 0.6, together yielding the best AvgAcc of 62.5%. Among the baselines, SEEKR-MLLM exhibits the strongest new-task plasticity with an AvgLA of 62.5%, but suffers substantially greater forgetting, while CL-MoE attains relatively low forgetting at the cost of a much lower AvgLA of 59.5%. These results demonstrate that VDA improves old-task stability without suppressing new-task adaptation, leading to a more favorable stability–plasticity balance. ℒVMAL_VMA ℒstrengthL_strength ℒstructL_struct AvgAcc ↑ AvgLA ↑ AvgF ↓ 51.0 61.7 12.8 ✓ 56.9 63.8 8.2 ✓ 58.0 61.5 4.2 ✓ ✓ 60.1 62.9 3.3 ✓ ✓ 59.8 60.3 0.7 ✓ ✓ ✓ 62.5 63.0 0.6 Table 3: Ablation study of the components in VDA. Ablation Study. Table 3 evaluates the contributions of VMA, VD strength preservation, and structural transport. VMA alone increases AvgLA from 61.7 to 63.8 and AvgAcc from 51.0 to 56.9, while reducing AvgF to 8.2, indicating that VD-guided token modulation mainly improves plasticity and partially alleviates forgetting. ℒstrengthL_strength further lowers AvgF to 4.2 but slightly reduces AvgLA, reflecting the adaptation cost of stability constraints. Combining it with VMA recovers plasticity and raises AvgAcc to 60.1. Joint strength and structural preservation reduce AvgF to 0.7, but limits AvgLA to 60.3. The complete model achieves the best AvgAcc of 62.5, with AvgLA of 63.0 and AvgF of 0.6. These results show that VMA compensates for the restricted adaptation caused by VD preservation, while the two preservation terms jointly provide strong old-task stability. Figure 4: RRAR trajectories on TextVQA throughout continual-learning. Detailed Analysis Analysis of cross-modal comprehension. To examine whether VDA strengthens and preserves cross-modal comprehension throughout continual learning, we track the Relevant Region Attention Ratio (RRAR) (Peng et al. 2026) on TextVQA using its ground-truth bounding boxes provided by(Khayatkhoei et al. 2025). RRAR measures the relative attention allocated to question-relevant visual regions compared with the entire image. As shown in Figure 4, initial adaptation increases RRAR, whereas the baseline (without both VMA and VC-OT) progressively declines as new tasks are learned, eventually falling below the frozen backbone. In contrast, VDA maintains an RRAR of 4.204 after the full sequence, outperforming the baseline by 0.499. Removing VMA or VC-OT reduces the final RRAR to 3.986 and 3.914, respectively, showing that both components contribute to maintaining strong question-relevant visual attribution. These results provide mechanism-level evidence that VDA enhances cross-modal comprehension during adaptation and prevents its degradation across subsequent tasks. VD Protection VD Strength Drift (%) ↓ VD Distribution Drift ↓ AvgAcc ↑ AvgF ↓ No Protection 35.94 0.1052 56.9 8.2 Token-Level L1 5.08 0.0187 59.1 0.4 Mass+KL 14.89 0.0428 60.3 3.4 VC-OT 14.27 0.0385 62.5 0.6 Table 4: Comparison of different VD preservation strategies. VD Strength Drift measures the relative change in total VD. VD Distribution Drift measures the JSD between VD distributions. Comparison of VD Preservation Strategies. We further compare VC-OT with several alternative strategies in preserving old-task VD. “No Protection” denotes VDA without VC-OT, while the remaining variants replace VC-OT with token-level L1 matching, or Mass+KL alignment. As shown in Table 4, No Protection exhibit substantial VD drift and severe forgetting, demonstrating the necessity of explicitly preserving old-task VD structures. Token-level L1 achieves the smallest VD drift and the lowest AvgF of 0.4, but only reaches an AvgAcc of 59.1, suggesting that rigid position-wise matching suppresses model’s plasticity. Mass+KL provides a softer distribution-level constraint, yet treats all cross-token shifts uniformly, without distinguishing visually consistent redistribution from harmful transfer toward vision-independent tokens. In contrast, VC-OT achieves comparable VD preservation while using region-aware transport and a dependence-stratified transport penalty to regulate different redistribution paths, resulting in a low AvgF and the highest global AvgAcc. Figure 5: Hyperparameter Analysis. Hyperparameter Analysis. Fig. 5 analyzes the effects of the LoRA rank and memory-buffer size. Increasing the LoRA rank initially improves AvgAcc by enhancing new-task adaptation, with the best performance achieved at r=128; further increasing the rank introduces more trainable parameters, leading to greater forgetting and worse AvgAcc. Enlarging the buffer generally improves performance by providing broader coverage of previous tasks, while the gain gradually saturates. A buffer size of 1,000 already achieves a good AvgAcc of 62.5. Conclusion This work introduces Multimodal Unsupervised Continual Post-Training (MU-CPT), where MLLMs continually learn from non-stationary multimodal data without labeled answers. We identify token-level visual dependence (VD) as an intrinsic signal of cross-modal attributional forgetting and propose VDA, which preserves old-task VD through VC-OT and promotes visually grounded adaptation through VMA. Experiments on six tasks show that our proposed framework reduces forgetting while retaining strong plasticity. References Agarwal et al. (2026) S. Agarwal, Z. Zhang, L. Yuan, J. Han, and H. Peng The unreasonable effectiveness of entropy minimization in llm reasoning. Advances in Neural Information Processing Systems 38, p. 107150–107180. Cited by: Related Work. Bai et al. (2025a) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: Introduction. Bai et al. (2025b) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin Qwen2.5-vl technical report. External Links: 2502.13923, Link Cited by: Introduction, Implementation Details.. Borisova et al. (2025) E. Borisova, N. Rauscher, and G. Rehm SciVQA 2025: overview of the first scientific visual question answering shared task. In Proceedings of the Fifth Workshop on Scholarly Document Processing (SDP 2025), p. 182–210. Cited by: Datasets.. Chen et al. (2024) C. Chen, J. Zhu, X. Luo, H. T. Shen, J. Song, and L. Gao Coin: a benchmark of continual instruction tuning for multimodel large language models. Advances in neural information processing systems 37, p. 57817–57840. Cited by: Introduction, Related Work. Chen et al. (2025) J. Chen, R. Cong, Y. Zhao, H. Yang, G. Hu, H. H. S. Ip, and S. Kwong Sefe: superficial and essential forgetting eliminator for multimodal continual instruction tuning. arXiv preprint arXiv:2505.02486. Cited by: Related Work, Comparison Methods.. De Lange et al. (2021) M. De Lange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars A continual learning survey: defying forgetting in classification tasks. IEEE transactions on pattern analysis and machine intelligence 44 (7), p. 3366–3385. Cited by: Introduction. Guo et al. (2025) H. Guo, F. Zeng, Z. Xiang, F. Zhu, D. Wang, X. Zhang, and C. Liu Hide-llava: hierarchical decoupling for continual instruction tuning of multimodal large language model. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 13572–13586. Cited by: Introduction. He et al. (2026) J. He, H. Guo, K. Zhu, M. Tang, and J. Wang Continual instruction tuning for large multimodal models. IEEE Transactions on Image Processing. Cited by: Related Work. He et al. (2024) J. He, H. Guo, K. Zhu, Z. Zhao, M. Tang, and J. Wang Seekr: selective attention-guided knowledge retention for continual learning of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 3254–3266. Cited by: Related Work, Comparison Methods.. Hu et al. (2025) J. Hu, Z. Zhang, G. Chen, X. Wen, C. Shuai, W. Luo, B. Xiao, Y. Li, and M. Tan Test-time learning for large language models. arXiv preprint arXiv:2505.20633. Cited by: Introduction, Related Work, Comparison Methods., Implementation Details.. Huai et al. (2025) T. Huai, J. Zhou, X. Wu, Q. Chen, Q. Bai, Z. Zhou, and L. He Cl-moe: enhancing multimodal large language model with dual momentum mixture-of-experts for continual visual question answering. In Proceedings of the computer vision and pattern recognition conference, p. 19608–19617. Cited by: Related Work, Comparison Methods.. Huang et al. (2023) J. Huang, S. Gu, L. Hou, Y. Wu, X. Wang, H. Yu, and J. Han Large language models can self-improve. In Proceedings of the 2023 conference on empirical methods in natural language processing, p. 1051–1068. Cited by: Related Work, Comparison Methods.. Hudson and Manning (2019) D. A. Hudson and C. D. Manning Gqa: a new dataset for real-world visual reasoning and compositional question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 6700–6709. Cited by: Datasets.. Jung et al. (2020) S. Jung, H. Ahn, S. Cha, and T. Moon Continual learning with node-importance based adaptive group sparse regularization. Advances in neural information processing systems 33, p. 3647–3658. Cited by: Visually Modulated Adaptation. Khayatkhoei et al. (2025) M. Khayatkhoei, P. Chhikara, F. Ilievski, et al. Mllms know where to look: training-free perception of small visual details with multimodal llms. In International Conference on Learning Representations, Vol. 2025, p. 68194–68213. Cited by: Analysis of cross-modal comprehension.. Kirkpatrick et al. (2017) J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences 114 (13), p. 3521–3526. Cited by: Related Work. Li et al. (2024) B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al. Llava-onevision: easy visual task transfer. arXiv preprint arXiv:2408.03326. Cited by: Introduction. Li et al. (2025) S. Li, M. Gao, T. Su, X. Zhang, and Z. Wang Multimodal continual instruction tuning with dynamic gradient guidance. arXiv preprint arXiv:2511.15164. Cited by: Related Work, Comparison Methods.. Li and Hoiem (2017) Z. Li and D. Hoiem Learning without forgetting. IEEE transactions on pattern analysis and machine intelligence 40 (12), p. 2935–2947. Cited by: Related Work. Lopez-Paz and Ranzato (2017) D. Lopez-Paz and M. Ranzato Gradient episodic memory for continual learning. Advances in neural information processing systems 30. Cited by: Related Work. Madaan et al. (2021) D. Madaan, J. Yoon, Y. Li, Y. Liu, and S. J. Hwang Representational continuity for unsupervised continual learning. arXiv preprint arXiv:2110.06976. Cited by: Related Work. Mallya and Lazebnik (2017) A. Mallya and S. Lazebnik Packnet: adding multiple tasks to a single network by iterative pruning. arXiv preprint arXiv:1711.05769. Cited by: Related Work. Parisi et al. (2018) G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter Continual lifelong learning with neural networks: a review. arXiv preprint arXiv:1802.07569 990. Cited by: Introduction. Peng et al. (2026) R. Peng, X. Wu, J. Lei, L. Hou, Y. Ma, and X. Li Deeper thought, weaker aim: understanding and mitigating perceptual impairment during reasoning in multimodal large language models. External Links: 2603.14184, Link Cited by: Analysis of cross-modal comprehension.. Prasad et al. (2024) A. Prasad, W. Yuan, R. Y. Pang, J. Xu, M. Fazel-Zarandi, M. Bansal, S. Sukhbaatar, J. Weston, and J. Yu Self-consistency preference optimization. arXiv preprint arXiv:2411.04109. Cited by: Related Work, Comparison Methods.. Qian et al. (2023) Z. Qian, X. Wang, X. Duan, P. Qin, Y. Li, and W. Zhu Decouple before interact: multi-modal prompt learning for continual visual question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 2953–2962. Cited by: Related Work. Rebuffi et al. (2017) S. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert Icarl: incremental classifier and representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, p. 2001–2010. Cited by: Related Work. Shao et al. (2026a) Z. Shao, B. Xiong, C. Xu, L. Xiao, K. Li, H. Gong, Y. Li, Y. Song, and X. Yang AgentPatch: coarse-to-fine weak-task repair for merging agentic multimodal large language models. arXiv preprint arXiv:2608.06699. Cited by: Introduction. Shao et al. (2026b) Z. Shao, B. Xiong, X. Yang, Y. Song, Q. Zhang, H. Chen, and C. Xu PivotMerge: bridging heterogeneous multimodal pre-training via post-alignment model merging. arXiv preprint arXiv:2604.22823. Cited by: Introduction. Sima et al. (2024) C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li Drivelm: driving with graph visual question answering. In European conference on computer vision, p. 256–274. Cited by: Datasets.. Singh et al. (2026) A. Singh, S. Marjit, W. Lin, P. Gavrikov, S. Yeung-Levy, H. Kuehne, R. Feris, S. Doveh, J. Glass, and M. J. Mirza Ttrv: test-time reinforcement learning for vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 33153–33163. Cited by: Introduction, Related Work, Comparison Methods.. Singh et al. (2019) A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach Towards vqa models that can read. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 8317–8326. Cited by: Datasets.. Smith et al. (2023) J. S. Smith, L. Karlinsky, V. Gutta, P. Cascante-Bonilla, D. Kim, A. Arbelle, R. Panda, R. Feris, and Z. Kira Coda-prompt: continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 11909–11919. Cited by: Evaluation Metrics.. Wang et al. (2025) W. Wang, Z. Gao, L. Gu, H. Pu, L. Cui, X. Wei, Z. Liu, L. Jing, S. Ye, J. Shao, Z. Wang, Z. Chen, H. Zhang, G. Yang, H. Wang, Q. Wei, J. Yin, W. Li, E. Cui, G. Chen, Z. Ding, C. Tian, Z. Wu, J. Xie, Z. Li, B. Yang, Y. Duan, X. Wang, Z. Hou, H. Hao, T. Zhang, S. Li, X. Zhao, H. Duan, N. Deng, B. Fu, Y. He, Y. Wang, C. He, B. Shi, J. He, Y. Xiong, H. Lv, L. Wu, W. Shao, K. Zhang, H. Deng, B. Qi, J. Ge, Q. Guo, W. Zhang, S. Zhang, M. Cao, J. Lin, K. Tang, J. Gao, H. Huang, Y. Gu, C. Lyu, H. Tang, R. Wang, H. Lv, W. Ouyang, L. Wang, M. Dou, X. Zhu, T. Lu, D. Lin, J. Dai, W. Su, B. Zhou, K. Chen, Y. Qiao, W. Wang, and G. Luo InternVL3.5: advancing open-source multimodal models in versatility, reasoning, and efficiency. External Links: 2508.18265, Link Cited by: Introduction. Wang et al. (2022) Z. Wang, Z. Zhang, S. Ebrahimi, R. Sun, H. Zhang, C. Lee, X. Ren, G. Su, V. Perot, J. Dy, et al. Dualprompt: complementary prompting for rehearsal-free continual learning. In European conference on computer vision, p. 631–648. Cited by: Evaluation Metrics.. Wei et al. (2026) L. Wei, Y. Li, C. Wang, Y. Wang, L. Kong, W. Huang, and L. Sun First sft, second rl, third upt: continual improving multi-modal llm reasoning via unsupervised post-training. Advances in Neural Information Processing Systems 38, p. 62293–62318. Cited by: Introduction, Introduction, Related Work, Comparison Methods.. Xu et al. (2026) C. Xu, K. Ke, Z. Liu, J. Wei, Z. Shao, W. Guo, and C. Yu EvoMAS: learning execution-time workflows for multi-agent systems. arXiv preprint arXiv:2605.08769. Cited by: Introduction. Ye et al. (2026) Z. Ye, Q. Li, X. Feng, R. Chen, Z. Li, H. Ren, K. Chen, D. Tu, and B. Qin Not all tokens see equally: perception-grounded policy optimization for large vision-language models. arXiv preprint arXiv:2604.01840. Cited by: Introduction. Yu et al. (2026) Y. Yu, Z. Wu, Z. Chen, H. Xu, Z. Liao, X. Deng, Z. Liu, S. Shi, and H. Wang Stabilizing unsupervised self-evolution of mllms via continuous softened retracing resampling. arXiv preprint arXiv:2604.03647. Cited by: Introduction, Related Work. Zhang et al. (2026) Q. Zhang, H. Wu, C. Zhang, P. Zhao, and Y. Bian Right question is already half the answer: fully unsupervised llm reasoning incentivization. Advances in neural information processing systems 38, p. 67345–67372. Cited by: Related Work. Zhang et al. (2023a) X. Zhang, F. Zhang, and C. Xu Vqacl: a novel visual question answering continual learning setting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19102–19112. Cited by: Related Work. Zhang et al. (2023b) X. Zhang, C. Wu, Z. Zhao, W. Lin, Y. Zhang, Y. Wang, and W. Xie Pmc-vqa: visual instruction tuning for medical visual question answering. arXiv preprint arXiv:2305.10415. Cited by: Datasets.. Zhao et al. (2024) H. Zhao, P. Zhou, D. Gao, Z. Bai, and M. Z. Shou Lova3: learning to visual question answering, asking and assessment. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: Implementation Details.. Zhao et al. (2025) H. Zhao, F. Zhu, H. Guo, M. Wang, R. Wang, G. Meng, and Z. Zhang Mllm-cl: continual learning for multimodal large language models. arXiv preprint arXiv:2506.05453. Cited by: Datasets.. Zhou et al. (2024) Y. Zhou, Z. Fan, D. Cheng, S. Yang, Z. Chen, C. Cui, X. Wang, Y. Li, L. Zhang, and H. Yao Calibrated self-rewarding vision language models. Advances in Neural Information Processing Systems 37, p. 51503–51531. Cited by: Introduction, Introduction, Related Work. Zhu et al. (2024a) K. Zhu, Z. Ge, L. Zhao, and X. Zhang Self-supervised visual preference alignment. arXiv preprint arXiv:2404.10501. Cited by: Introduction, Introduction, Related Work, Comparison Methods.. Zhu et al. (2024b) Z. Zhu, X. Hong, Z. Ma, W. Zhuang, Y. Ma, Y. Dai, and Y. Wang Reshaping the online data buffering and organizing mechanism for continual test-time adaptation. In European conference on computer vision, p. 415–433. Cited by: Introduction. Zhu et al. (2026a) Z. Zhu, Z. Ma, Y. Wang, Y. Song, Y. Wang, and X. Hong Sample-aware knowledge association and enhancement for open-vocabulary continual learning. International Journal of Computer Vision 134 (7), p. 332. Cited by: Introduction. Zhu et al. (2026b) Z. Zhu, Y. Wang, Z. Ma, Y. Song, Y. Wang, and X. Hong Dance across shifts: forward-facilitation continual test-time adaptation through dynamic style bridging. arXiv preprint arXiv:2605.18608. Cited by: Introduction.