Paper deep dive
Evolving Prompt Adaptation for Vision-Language Models
Enming Zhang, Jiayang Li, Yanru Wu, Zhenyu Liu, Yang Li
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/13/2026, 1:03:13 AM
Summary
EvoPrompt is a novel framework for parameter-efficient adaptation of Vision-Language Models (VLMs) that addresses catastrophic forgetting by governing the evolutionary trajectory of prompts. It utilizes a Modality-Shared Prompt Projector (MPP) for hierarchical prompt generation, a decoupled low-rank training strategy to preserve semantic directions while adapting magnitudes, and Feature Geometric Regularization (FGR) to prevent representation collapse, achieving state-of-the-art performance in few-shot and cross-dataset transfer tasks.
Entities (5)
Relation Signals (4)
EvoPrompt → builton → CLIP
confidence 100% · we develop our EvoPrompt framework on top of the pre-trained CLIP
EvoPrompt → incorporates → Feature Geometric Regularization
confidence 95% · To further stabilize this process, we incorporate feature geometric regularization
EvoPrompt → mitigates → Catastrophic Forgetting
confidence 95% · EvoPrompt addresses knowledge forgetting through a cohesive design
EvoPrompt → utilizes → Modality-Shared Prompt Projector
confidence 95% · Specifically, our approach employs a Modality-Shared Prompt Projector (MPP) to generate hierarchical prompts
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge. While parameter-efficient prompt learning methods offer a promising path, they often suffer from catastrophic forgetting of pre-trained knowledge. Toward addressing this limitation, our work is grounded in the insight that governing the evolutionary path of prompts is essential for forgetting-free adaptation. To this end, we propose EvoPrompt, a novel framework designed to explicitly steer the prompt trajectory for stable, knowledge-preserving fine-tuning. Specifically, our approach employs a Modality-Shared Prompt Projector (MPP) to generate hierarchical prompts from a unified embedding space. Critically, an evolutionary training strategy decouples low-rank updates into directional and magnitude components, preserving early-learned semantic directions while only adapting their magnitude, thus enabling prompts to evolve without discarding foundational knowledge. This process is further stabilized by Feature Geometric Regularization (FGR), which enforces feature decorrelation to prevent representation collapse. Extensive experiments demonstrate that EvoPrompt achieves state-of-the-art performance in few-shot learning while robustly preserving the original zero-shot capabilities of pre-trained VLMs.
Tags
Links
- Source: https://arxiv.org/abs/2603.09493v1
- Canonical: https://arxiv.org/abs/2603.09493v1
Trouble viewing inline? Open PDF directly →
Full Text
51,293 characters extracted from source content.
Expand or collapse full text
Evolving Prompt Adaptation for Vision-Language Models Enming Zhang 1 , Jiayang Li 1 , Yanru Wu 1 , and Zhenyu Liu 1 Yang Li 2 1 Tsinghua Shenzhen International Graduate School, Tsinghua University 2 Chinese University of Hong Kong, Shenzhen yangl@cuhk.edu.cn Abstract. The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant chal- lenge. While parameter-efficient prompt learning methods offer a promis- ing path, they often suffer from catastrophic forgetting of pre-trained knowledge. Toward addressing this limitation, our work is grounded in the insight that governing the evolutionary path of prompts is essential for forgetting-free adaptation. To this end, we propose EvoPrompt, a novel framework designed to explicitly steer the prompt trajectory for stable, knowledge-preserving fine-tuning. Specifically, our approach em- ploys a Modality-Shared Prompt Projector (MPP) to generate hierarchi- cal prompts from a unified embedding space. Critically, an evolutionary training strategy decouples low-rank updates into directional and mag- nitude components, preserving early-learned semantic directions while only adapting their magnitude, thus enabling prompts to evolve with- out discarding foundational knowledge. This process is further stabi- lized by Feature Geometric Regularization (FGR), which enforces feature decorrelation to prevent representation collapse. Extensive experiments demonstrate that EvoPrompt achieves state-of-the-art performance in few-shot learning while robustly preserving the original zero-shot capa- bilities of pre-trained VLMs. Keywords: Parameter-Efficient Adaptation· Vision-Language Models 1 Introduction Large-scale pre-trained vision-language models (VLMs) [24, 26, 32, 42, 59], ex- emplified by works like CLIP [42] and ALIGN [26], have revolutionized zero- shot generalization across diverse downstream tasks, including image classifica- tion [12, 14], visual question answering [2, 17], and cross-modal retrieval [8, 30]. Their success stems from learning highly transferable visual and linguistic rep- resentations through contrastive pre-training on massive web-scale datasets. However, adapting these powerful models to specific downstream tasks with limited labeled samples presents a significant challenge. The conventional ap- proach of full fine-tuning, which updates all model parameters, is often pro- hibitively expensive in terms of computation and storage, given the massive scale arXiv:2603.09493v1 [cs.CV] 10 Mar 2026 2Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li Fig. 1: Comparison of our proposed EvoPrompt frameworks with related representative efficient transfer learning for VLMs. of VLMs [1,7]. To address this, parameter-efficient adaptation methods [31], par- ticularly prompt learning, have gained prominence. Techniques like CoOp [64] and CoCoOp [63] introduce a set of learnable continuous prompts while keeping the pre-trained backbone frozen, drastically reducing tunable parameters. Despite their efficiency, existing prompt-tuning paradigms are bottlenecked by several critical challenges. Structurally, the prevalent layer-isolated designs treat prompts as independent parameters across the encoder’s depth, disrupting the hierarchical flow of semantic information across the hierarchy. Moreover, in terms of modality alignment, current schemes like MaPLe [27] often exhibit a text-centric bias, as illustrated in Fig. 1, which fails to leverage the complemen- tary vision-language interactions. Most critically, during few-shot adaptation, the learnable prompts tend to rapidly deviate from pre-trained semantic anchors and overfit to the limited downstream data, resulting in catastrophic forgetting of the model’s original zero-shot generalization capabilities [28,51]. We argue that the key to mitigating these limitations lies in governing the evolutionary trajectory of prompts. Unlike conventional approaches that treat prompt tuning as a static parameter injection, we observe that prompts naturally undergo a progressive evolution from general semantic anchors to fine-grained task-specific features. This insight motivates our EvoPrompt, a trajectory-aware paradigm explicitly designed to guide this transition towards stable adaptation. In our work, EvoPrompt addresses knowledge forgetting through a cohe- sive design centered on guided prompt evolution. Specifically, we introduce a Modality-Shared Prompt Projector (MPP), which replaces isolated per-layer prompts with a unified embedding projected into layer-specific prompts via de- composed adapters, establishing a bridge for cross-layer and cross-modal synergy. Within this unified structure, we further regulate the temporal dynamics of adap- tation via an evolutionary trajectory-aware strategy. By disentangling low-rank residuals updates into direction and magnitude, we freeze the broad semantic directions captured in early training while only refining their magnitudes. This mechanism allows the model to learn task-specific skills without discarding its pre-trained knowledge. To further stabilize this process, we incorporate feature geometric regularization, which enforces orthogonality to prevent representation collapse in low-data regimes. Our contributions are summarized as follows: – We propose EvoPrompt, a novel paradigm that explicitly governs prompt evolution through trajectory-aware adaptation, effectively preventing catas- trophic forgetting of pre-trained knowledge. Evolving Prompt Adaptation for Vision-Language Models3 – We design a projector coupled with a novel training strategy that enables decoupled control of prompt directions and magnitudes, complemented by a regularization term to prevent representation collapse. – Through comprehensive experiments, EvoPrompt achieves state-of-the-art performance in cross-dataset transfer, domain generalization, and few-shot image recognition. 2 Related Work 2.1 Vision-Language Models The landscape of computer vision has been profoundly reshaped by the advent of Vision-Language Models (VLMs), which forge robust semantic connections between visual and textual data. A myriad of foundational architectures—such as CLIP [42], ALIGN [26], CoCa [59], BLIP-2 [32], and Flamingo [1]—have been trained on web-scale, paired multi-modal datasets using self-supervised contrastive objectives [39]. By assimilating unprecedented volumes of training samples [46], these foundation models learn highly transferable feature repre- sentations, yielding impressive zero-shot inference capabilities across a broad array of applications. Despite these triumphs, effectively transferring such mas- sive architectures to specialized target distributions with limited annotated data remains remarkably challenging. To navigate this data-scarce regime, a vast cor- pus of literature has focused on tailoring pre-trained VLMs to diverse down- stream scenarios, encompassing few-shot image classification [60,61], object de- tection [13,18], and semantic segmentation [11,43]. 2.2 Efficient Transfer Learning for VLMs Parameter-efficient fine-tuning (PEFT), particularly via prompting techniques, originated in the natural language domain to steer massive linguistic models without updating their full parameter space [4,22,23,31,33,35]. This philosophy was subsequently extended to multi-modal learning [6, 15, 48, 61], enabling the rapid adaptation of frozen VLMs. Pioneering works like CoOp [64] appended op- timizable continuous tokens to the textual branch of CLIP, while CoCoOp [63] conditioned these tokens on visual inputs to mitigate overfitting to seen classes. To further regulate the optimization process and retain foundational knowledge, methods such as KgCoOp [57] penalize the discrepancy between the learned textual embeddings and the original frozen embeddings, and PLOT [5] leverages optimal transport to holistically match vision and text semantics. Moving be- yond unimodal prompting, MaPLe [27] synchronizes the adaptation by explicitly linking deep learnable tokens across both vision and text encoders. More recently, approaches like PromptSRC [28] and TCP [58] have integrated self-consistency mechanisms and class-aware regularization to further constrain the learning tra- jectory. Drawing from the success of feature disentanglement in continuous learn- ing, DualPrompt [51] isolates task-specific knowledge, and DePT [25] segregates base-category semantics into dedicated feature subspaces. While these prompt- driven methodologies have proven highly effective in adapting large-scale models with minimal overhead, they suffer from inherent structural limitations. 4Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li 3 Method Following prior works, we develop our EvoPrompt framework on top of the pre- trained CLIP [42]. In the following, we first introduce preliminary knowledge on CLIP and then present our proposed EvoPrompt. 3.1 Preliminaries We provide a brief overview of the CLIP model’s notation and core operations used in our method. CLIP contains two primary components: a visual encoder F and a text encoder G. Image Feature Extraction. The visual encoder f is structured as a Vision Transformer (ViT) with L consecutive transformer blocks, F i L i=1 . An input RGB image I ∈ R H×W×3 is split into M non-overlapping patches. A linear projection layer maps each patch to a d v -dimensional vector, producing the initial patch embedding matrix E 0 ∈ R M×d v . This matrix is prepended with a learnable [CLS] token c 0 , combined with positional encodings, and fed into the transformer stack. The processing at the i-th layer is formulated as: [c i ,E i ] = F i ([c i−1 ,E i−1 ]), i = 1,...,L.(1) The final [CLS] token representation c L is projected via a linear layer Φ v to obtain the image feature vector f v = Φ v (c L )∈ R d . Text Feature Extraction. For a text prompt T (e.g., "a photo of a [CLASS]"), the input is first tokenized into a sequence of N tokens. These tokens are em- bedded as T 0 ∈ R N×d t and concatenated with special [SOS] and [EOS] tokens (b 0 and e 0 ), along with positional encodings. The sequence is processed by L transformer layers G i L i=1 in the text encoder: [b i ,T i ,e i ] = G i ([b i−1 ,T i−1 ,e i−1 ]), i = 1,...,L.(2) The final [EOS] token representation e L is linearly projected to obtain the text feature vector f t = Φ t (e L )∈ R d . Zero-Shot Classification and Optimization. For a C-class task, we con- struct C text prompts to obtain text features f t c C c=1 . Given an image feature f v , the prediction probability for class c is computed via cosine similarity and a softmax with temperature τ: p(y = c| f v ) = exp(s c /τ) P C j=1 exp(s j /τ) , where s c = f v ⊤ f t c ∥f v ∥f t c ∥ . (3) The model is typically optimized using the cross-entropy loss. For a training sample with ground-truth label y, the loss is defined as: L InfoNCE (f v ,f t ) =− log exp(s y /τ) P C j=1 exp(s j /τ) .(4) Evolving Prompt Adaptation for Vision-Language Models5 Fig. 2: Overview of the proposed EvoPrompt framework. Left: Modality-shared pro- jectors are used to inject prompts into dual encoders. Top-right: To enhance fea- ture orthogonality,L fgr transforms correlated representations into mutually indepen- dent vectors. Bottom-right: The low-rank adapter is decomposed into magnitudeα i and direction components, with historical directions frozen to preserve early geometric alignments, while the magnitudes remain trainable for later adaptation. 3.2 Modality-Shared Prompt Projector Previous multimodal prompting schemes, such as MaPLe [27], typically insert prompts into each layer independently. While this provides layer-specific guid- ance, such isolated prompts often prevent the model from distilling and propa- gating beneficial information across the hierarchical depth of the encoders. We argue that prompts should capture the hierarchical semantic progression across consecutive layers and maintain a degree of inter-layer correlation. Furthermore, leveraging complementary information across modalities can enrich the prompt generation process. To this end, we propose the Modality-Shared Prompt Projec- tor (MPP), which jointly fosters cross-layer information flow and complementary cross-modal interaction. Learnable Embedding Space We first initialize a unified, learnable embed- ding space E ∈ R K×d r , where K vectors are sampled from a zero-mean Gaussian distributionN(0,σ 2 ). This shared embedding is then transformed into modality- specific prompts for each layer through a projector. Specifically, prompts are inserted starting from a predefined layer J (1≤ J ≤ L). For modality m∈v,t at each layer i∈J,...,L, the prompt P m i ∈ R l×d m is generated as: P m i = Proj m i (E).(5) The generated prompts P m i are then concatenated with the original input tokens at the corresponding layer. 6Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li Decoupled Low-Rank Expansion To efficiently model both the cross-layer semantic patterns and the layer-specific adaptations, we propose a parameter- efficient projection mechanism inspired by LoRA [23]. The core idea is to de- couple the projector’s weight matrix into a shared component and a low-rank, layer-wise adapter. Typically, LoRA updates a frozen weight W 0 ∈ R d×k via a low-rank residual BA, where B ∈ R d×r and A∈ R r×k (r ≪ min(d,k)). Extend- ing this, for each projector associated with layer i ∈ J,...,L, the projector weight matrix is first decomposed as: W m i = W m shared + ∆W m i ,(6) where W m shared ∈ R d r ×d m is a modality-specific shared component maintained across layers from J to L to capture fundamental semantic knowledge and al- leviate redundancy. For notational brevity, we omit the modality superscript m in the following. The layer-specific adapter ∆W i is then parameterized via a low-rank decomposition, leading to the final form: W i = W shared + A i B i ,(7) where A i ∈ R d r ×r and B i ∈ R r×d m are trainable low-rank matrices. The adop- tion of a single W shared coupled with layer-wise low-rank adapters enables Evo- Prompt to retain expressive power with significantly fewer parameters. Con- sequently, the parameter complexity is lowered from O((L− J + 1)· d r d m ) to O(d r d m +(L−J +1)·r(d r +d m )). This formulation naturally enforces structural alignment through the shared base and strengthens generalization by cleanly separating common knowledge from layer-specific adjustments. 3.3 Evolutionary Trajectory-Aware Learning Strategy While the MPP architecture establishes structural bridges across layers, op- timizing prompts throughout the training process remains challenging due to the risk of catastrophic forgetting. We observe that prompts, which serve as general contextual anchors in early training stages, tend to converge toward task-specific patterns in later epochs, potentially overwriting previously acquired generalizable knowledge. To mitigate this, inspired by progressive learning meth- ods [16, 56], we introduce an evolutionary trajectory-aware learning strategy. This approach explicitly decouples and modulates parameter effects learned at different phases, conceptualizing adaptation as a progressive accumulation of knowledge. Incremental Magnitude-Direction Decoupling Building on insights from weight decomposition analysis [52], we factorize the layer-wise low-rank update ∆W t i at training epoch t into a learnable magnitude coefficient α t i and a nor- malized directional matrix: ∆W t i = α t i · A t i B t i ∥A t i B t i ∥ F = α t i ·A t i B t i , (8) Evolving Prompt Adaptation for Vision-Language Models7 where ∥·∥ F denotes the Frobenius norm. This explicit decoupling grants inde- pendent control over the adaptation strength (α t i ) and its direction (A t i B t i ). To promote stable progressive learning, we frame the training process as the accumu- lation of directional knowledge. Prior work suggests the directional component is more critical than its magnitude in low-rank adaptation [34,41]. Accordingly, we extend Eq. (7) to compute the adapter weight for layer i at epoch T as a historical sum: W T i = W shared + T−1 X t=1 α t i A t i B t i + α T i A T i B T i , (9) Here, W shared acts as a fixed, modality-specific foundation. During training at epoch T, we freeze all previously acquired directions A t i B t i T−1 t=1 to preserve their geometric structure. Only the magnitude coefficientsα t i T t=1 and the new direc- tion A T i B T i remain trainable. This design enables the model to recalibrate the influence of past knowledge via the learnable α t i while progressively incorporat- ing new directional adjustments, thereby adapting to the evolving loss landscape without catastrophically forgetting previously learned, robust features. Adaptive Rank Reduction To enhance continual adaptation stability and mitigate the risk of overfitting during later evolutionary stages, we introduce an empirical rank-reduction mechanism. Denoting the total number of training epochs as N e , this strategy modulates the capacity of the learnable matrices A t i ∈ R d r ×r t and B t i ∈ R r t ×d m by adjusting the rank r t in a controlled, stepwise manner: r 1 = r 2 =· > r μ = r μ+1 =· > r ν = r ν+1 =· = r N e , (10) where μ and ν (1 < μ < ν ≤ N e ) represent predefined epoch indices at which the rank is reduced. By assigning lower-rank weights to later epochs based on their diminishing marginal contributions, this strategy imposes a structural regular- ization that stabilizes the optimization landscape. Consequently, it significantly reduces cumulative computational and memory overhead while maintaining the model’s generalization ability. A complete algorithmic flowchart illustrating our training strategy is available in Appendix. 3.4 Feature Geometric Regularization Standard contrastive learning, e.g., InfoNCE, primarily focuses on instance-level alignment by maximizing the mutual information between paired samples. How- ever, this objective often ignores the intrinsic geometric structure of the feature space, potentially leading to feature collapse where learned dimensions become highly redundant or correlated. To address this, we introduce Feature Geometric Regularization (FGR), a principled objective derived from the Soft Hirschfeld- Gebelein-Rényi (Soft-HGR) maximal correlation framework [45,50]. Unlike tra- ditional CCA, Soft-HGR provides a variational objective that simultaneously promotes cross-modal alignment and intra-modal feature decorrelation. 8Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li Theorem 1 (Soft-HGR Maximal Correlation). Let u∈ R d and v∈ R d be zero-mean random vectors. The Soft-HGR maximal correlation between u and v is the solution to the following optimization problem: max φ,ψ E[φ(u) ⊤ ψ(v)]− 1 2 tr cov(φ(u))· cov(ψ(v)) , (11) subject to E[φ(u)] = E[ψ(v)] = 0, where φ,ψ are measurable functions. In our EvoPrompt framework, we instantiate φ and ψ using the visual and tex- tual encoders. Let F v = [f v 1 ,...,f v B ] ⊤ ∈ R B×d and F t = [f t 1 ,...,f t B ] ⊤ ∈ R B×d denote the batch of B normalized feature vectors extracted by the two encoders. Theoretically, as detailed in the Appendix, the contrastive loss can be viewed as a scalable and stable surrogate for maximizing the first term in Eq. (11), which promotes cross-modal alignment. However, it lacks explicit control over the sec- ond term, which characterizes the intra-modal covariance structure. Therefore, we introduce the following feature geometric regularization loss: L fgr (F v ,F t ) = 1 2 tr cov(F v ) cov(F t ) ,(12) where cov(·) is the empirical covariance matrix computed over the batch. By minimizing the product of the modal covariance matrices, we encourages the model to learn representations where individual feature dimensions are decorre- lated, effectively promoting orthogonality and reducing redundancy within the feature space. Overall Training Objective To preserve the rich semantic knowledge embed- ded in the pre-trained CLIP model, we introduce a knowledge constancy loss on both modalities. Let f v and f t be the prompted features for a single sample, and let f v 0 and f t 0 denote the corresponding features extracted by the original, frozen CLIP encoders (without prompts). The constancy loss is formulated as: L kcl = 1 2 1− f v · f v 0 ∥f v ∥f v 0 ∥ + 1− f t · f t 0 ∥f t ∥f t 0 ∥ . (13) This term ensures that the learned prompts do not cause the feature repre- sentations to deviate excessively from the well-structured original CLIP feature distribution, thereby maintaining its strong zero-shot generalization capability. Combining the standard contrastive alignment loss L ce (e.g., InfoNCE), the feature geometric regularization, and the knowledge constancy terms, our com- plete training objective is: L total =L InfoNCE + γL fgr + ηL kcl ,(14) where γ and η are balancing hyperparameters. This composite loss guides the model to achieve strong cross-modal instance alignment while fostering a well- structured, disentangled, and knowledge-preserving feature geometry. Evolving Prompt Adaptation for Vision-Language Models9 4 Experiment We evaluate EvoPrompt under four standard experimental settings: base-to- novel generalization, cross-dataset transfer, domain generalization, and few-shot learning. Unless otherwise noted, we strictly follow the evaluation protocols es- tablished in prior work [63,64]. 4.1 Tasks and Datasets Base-to-Novel Generalization To evaluate the trade-off between task-specific adaptation and zero-shot capability preservation, we split the categories of each dataset equally into a base set for training and a novel set for evaluation. Models are trained exclusively on the base classes and tested on both base and novel classes. This experiment is conducted across 11 standard image classification benchmarks: ImageNet [10], Caltech101 [14], OxfordPets [40], StanfordCars [29], Flowers102 [38], Food101 [3], FGVCAircraft [37], SUN397 [53], UCF101 [47], DTD [9], and EuroSAT [19]. Cross-Dataset Transfer Following CoCoOp [63], we assess out-of-distribution generalization by training a 16-shot model on ImageNet (covering all 1,000 classes) and then evaluating the frozen model directly on the other 10 datasets without any further fine-tuning. Domain Generalization We measure robustness to distribution shifts by eval- uating the same ImageNet-trained model on four challenging ImageNet variants: ImageNet-V2 [44], ImageNet-Sketch [49], ImageNet-A [21], and ImageNet-R [20]. This evaluates the model’s ability to maintain performance under domain shift. Few-Shot Learning To evaluate sample efficiency, we train models with vary- ing numbers of labeled examples (1, 2, 4, 8, and 16 shots per category) and test on the full test sets of each benchmark. This setting probes the model’s ability to learn effectively from extremely limited supervision and reveals whether it acquires both task-specific discriminative patterns and task-agnostic knowledge. 4.2 Implementation Details We adopt a pre-trained CLIP with a ViT-B/16 [12] backbone as our foundation model. Unless investigating variable-shot performance, we sample 16 shots per class following prior work [27, 55, 57, 58, 63–65]. Zero-shot classifier weights are generated using standard prompt templates [42,61,64]. Both the visual encoder F and the text encoder G remain fully frozen. For the learnable embedding space E, we set the number of vectors to K = 5 and the shared representation dimension to d r = 512. Prompts with token length l = 5 are inserted from layer J = 6 to the final layer L = 12. The more detailed configurations and analysis of these parameters, along with the r t , μ and ν in rank reduction mechanism, are provided in the Appendix due to space limitations. All experiments are conducted on a single NVIDIA A800 GPU, and we report the average top-1 accuracy over three random seeds. 10Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li Table 1: Comparison with state-of-the-art methods on base-to-novel generalization across 11 datasets. EvoPrompt demonstrates strong generalization results over existing methods. The best results are in bold and the second-best results are underlined. Method AverageImageNetCaltech101OxfordPets Base Novel HMBase Novel HMBase Novel HMBase Novel HM CLIP69.34 74.22 71.7072.43 68.14 70.2296.84 94.00 95.4091.17 97.26 94.12 CoOp82.69 63.22 71.6676.47 67.88 71.9298.00 89.81 93.7393.67 95.29 94.47 CoCoOp80.47 71.69 75.8375.98 70.43 73.1097.96 93.81 95.8495.20 97.69 96.43 ProDA81.56 72.30 76.6575.40 70.23 72.7298.27 93.23 95.6895.43 97.83 96.62 KgCoOp 80.73 73.60 77.0075.83 69.96 72.7897.72 94.39 96.0394.65 97.76 96.18 MaPLe82.28 75.14 78.5576.66 70.54 73.4797.74 94.36 96.0295.43 97.76 96.58 PromptSRC84.26 76.10 79.9777.60 70.73 74.0198.10 94.03 96.0295.33 97.30 96.30 ProVP85.20 73.22 78.7675.82 69.21 72.3698.92 94.21 96.5195.87 97.65 96.75 MetaPrompt 83.65 75.48 79.0977.5270.83 74.0298.13 94.58 96.3295.5397.00 96.26 TCP84.13 75.36 79.5177.27 69.87 73.3898.23 94.67 96.4294.67 97.20 95.92 MMA83.20 76.8079.8777.31 71.0074.0298.4094.00 96.1595.40 98.0796.72 EvoPrompt84.2877.7680.7376.9871.8074.2998.3095.0096.6295.3098.1096.68 Method StanfordCarsFlowers102Food101FGVCAircraft Base Novel HMBase Novel HMBase Novel HMBase Novel HM CLIP63.37 74.89 68.6572.08 77.8074.8390.10 91.22 90.6627.19 36.29 31.09 CoOp 78.12 60.40 68.1397.60 59.67 74.0688.33 82.26 85.1940.44 22.30 28.75 CoCoOp70.49 73.59 72.0194.87 71.75 81.7190.70 91.29 90.9933.41 23.71 27.74 ProDA 74.70 71.20 72.9197.70 68.68 80.6690.30 88.57 89.4336.90 34.13 35.46 KgCoOp71.76 75.0473.3695.00 74.73 83.6590.50 91.70 91.0936.21 33.55 34.83 MaPLe72.94 74.00 73.4795.92 72.46 82.5690.71 92.0591.3837.44 35.61 36.50 PromptSRC 78.27 74.97 76.5898.0776.50 85.9590.67 91.53 91.1042.73 37.8740.15 ProVP80.4367.96 73.6798.42 72.06 83.2090.32 90.91 90.6147.08 29.87 36.55 MetaPrompt 76.34 75.01 75.4897.66 74.49 84.5290.7491.85 91.2940.14 36.51 38.24 TCP80.80 74.13 77.3297.73 75.57 85.2390.57 91.37 90.9741.97 34.43 37.83 MMA 78.50 73.10 75.7097.77 75.93 85.4890.13 91.30 90.7140.57 36.33 38.33 EvoPrompt78.3075.9077.0897.2078.3386.7590.8092.7091.7443.2039.1441.07 Method SUN397DTDEuroSATUCF101 Base Novel HMBase Novel HMBase Novel HMBase Novel HM CLIP69.36 75.35 72.2353.24 59.90 56.3756.48 64.05 60.0370.53 77.50 73.85 CoOp 80.60 65.89 72.5179.44 41.18 54.2492.19 54.74 68.6984.69 56.05 67.46 CoCoOp79.74 76.86 78.2777.01 56.00 64.8587.49 60.04 71.2182.33 73.45 77.64 ProDA 78.67 76.93 77.7980.67 56.48 66.4483.90 66.00 73.8885.23 71.97 78.04 KgCoOp80.29 76.53 78.3677.55 54.99 64.3585.64 64.34 73.4882.89 76.67 79.65 MaPLe 80.82 78.70 79.7580.36 59.18 68.1694.07 73.23 82.3583.00 78.66 80.77 PromptSRC82.67 78.47 80.5283.3762.97 71.7592.90 73.90 82.3287.10 78.80 82.74 ProVP 80.67 76.11 78.3283.95 59.06 69.3497.12 72.91 83.2988.56 75.55 81.54 MetaPrompt82.26 79.0480.6283.10 58.05 68.3593.53 75.21 83.3885.33 77.72 81.35 TCP82.6378.20 80.3582.77 58.07 68.2591.63 74.73 82.3287.13 80.7783.83 MMA82.27 78.57 80.3883.20 65.63 73.3885.46 82.34 83.8786.23 80.03 82.20 EvoPrompt82.3179.2080.7383.1064.1072.3794.1080.1086.5487.5080.9784.11 4.3 Base-to-Novel Generalization In this experiment, we evaluate the performance of EvoPrompt against sev- eral representative benchmarks, including the zero-shot CLIP baseline and var- ious state-of-the-art prompt learning techniques such as CoOp [64], CoCoOp [63], ProDA [36], KgCoOp [57], MaPLe [27], PromptSRC [28], ProVP [54], MetaPrompt [62], TCP [58], and the adapter-based method MMA [55]. To ensure a fair comparison, we exclude methods that rely on large language models for external prompt priors or those utilizing full unlabeled datasets for distillation. Evolving Prompt Adaptation for Vision-Language Models11 Table 2: Comparison of EvoPrompt with previous state-of-the-art methods on cross- dataset evaluation across 10 datasets. SourceTarget ImageNetCaltech101OxfordPetsStanfordCarFlowers102Food101AircraftSUN397DTDEuroSATUCF101Average CoOp 71.5193.70 89.14 64.51 68.71 85.30 18.47 64.15 41.92 46.39 66.5563.88 CoCoOp 71.0294.4390.14 65.32 71.88 86.06 22.94 67.36 45.73 45.37 68.2165.74 MaPLe 70.7293.53 90.4965.57 72.23 86.20 24.7467.01 46.49 48.06 68.6966.30 PromptSRC 71.27 93.60 90.25 65.70 70.25 86.15 23.90 67.10 46.8745.50 68.7565.81 TCP 71.4093.97 91.25 64.69 71.21 86.6923.45 67.15 44.35 51.45 68.7366.29 MMA 71.00 93.80 90.30 66.13 72.07 86.12 25.33 68.17 46.57 49.24 68.3266.61 EvoPrompt71.6394.9090.4065.8072.5087.4024.2067.4046.9249.4069.3066.82 As detailed in Tab.1, EvoPrompt achieves superior average performance across 11 datasets. Specifically, our method yields improvements of 0.96% on Novel classes, and 0.76% in terms of the harmonic mean (HM) over the previous best-performing model. These results demonstrate that EvoPrompt establishes a new state-of-the-art across all primary evaluation metrics. Furthermore, EvoPrompt demonstrates enhanced transfer learning capabil- ities by significantly boosting base accuracy while maintaining strong general- izability. Even in cases where our method does not simultaneously achieve the highest accuracy for both categories on a single dataset, it consistently outper- forms competitors on one metric by a substantial margin, leading to a superior HM. For instance, on FGVCAircraft, it achieves a significant gain of 1.27% in Novel accuracy despite a slight margin in the Base category. 4.4 Cross-Dataset Evaluation Tab. 2 reports the cross-dataset evaluation where all models are trained on Ima- geNet and directly evaluated on 10 diverse target datasets. EvoPrompt achieves the highest average accuracy of 66.82%, surpassing MMA (66.61%) and MaPLe (66.30%), while also attaining the best source accuracy on ImageNet (71.63%), suggesting that the MPP’s shared representation space and evolutionary learn- ing strategy yield more transferable prompts than MaPLe’s independently pa- rameterized per-layer design. Compared with MMA, EvoPrompt maintains a competitive edge on the target average. Overall, EvoPrompt achieves the best balance between source-domain performance and cross-dataset transferability among all compared methods. 4.5 Domain Generalization Tab. 3 assesses model robustness under natural distribution shifts using four chal- lenging ImageNet variant datasets. EvoPrompt achieves the best average accu- racy across all domains, demonstrating that it not only enhances in-distribution classification but also more effectively preserves CLIP’s inherent out-of-distribution generalization capabilities compared to existing adaptation methods. 12Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li 4.6 Few-Shot Learning Table 3: Comparison of EvoPrompt with previous methods on domain gen- eralization across 4 datasets. SourceTarget ImNet-V2 -S -A -R CLIP 66.7360.83 46.15 47.77 73.96 CoOp 71.5164.20 47.99 49.71 75.21 CoCoOp 71.0264.07 48.75 50.63 76.18 MaPLe 70.72 64.07 49.15 50.90 76.98 PromptSRC 71.2764.3549.55 50.90 77.80 MMA 71.0064.33 49.13 51.1277.32 EvoPrompt71.6364.4049.4051.3077.90 Few-shot classification results are pre- sented in Fig. 3. EvoPrompt demonstrates solid and competitive performance across the evaluated shot settings. While per- formance is broadly comparable in the most data-scarce regimes, the advantage of EvoPrompt becomes more pronounced as the number of training examples in- creases. This scaling behavior indicates that the framework effectively leverages additional supervisory signals to learn in- creasingly transferable representations. 4.7 Ablation Study Component Analysis Tab. 4 (a) quantifies the contribution of each compo- nent in EvoPrompt on ImageNet, revealing a clear trade-off between base-class fitting and novel-class generalization. The MPP and the shared weight matrix W shared are fundamental to the architecture. Removing the MPP (w.o. MPP), which reverts to the conventional approach of injecting isolated, layer-specific prompts, causes the largest performance drop (72.64% HM). Similarly, replacing the shared W shared with independent per-layer weights (w.o. W s ) also leads to a notable decline (73.54% HM). Replacing the low-rank adapter with a full-rank projection (w.o. AB) similarly hurts performance, confirming the efficiency and sufficiency of the low-rank parameterization. Notably, while removing the evolu- tionary learning strategy (w.o. E.T.) or the L kcl leads to a performance gain on base classes—reaching 77.42% and 77.24% respectively—these variants suffer a significant decline in novel class accuracy, dropping to 70.25% and 70.55%. This indicates that without these constraints, the model overfits the base distribution at the expense of its zero-shot generalization capability. The L fgr loss provides critical refinement, as its absence results in a lower HM of 73.48%. These re- sults confirm that our design choices are indispensable for achieving superior and robust generalization across both base and novel classes. Loss Weight Sensitivity We evaluate the sensitivity of hyper-parameters γ (for L kcl ) and η (for L fgr ) on ImageNet, as shown in Tab. 4 (b). The model achieves the optimal trade-off at γ = 25 and η = 0.5, reaching 74.29% HM. Deviating from these values generally causes a performance drop, particularly in the harmonic mean, suggesting that a balanced weighting of the proposed losses is crucial for robust cross-class generalization. 4.8 Further Analysis Computational Efficiency As shown in Tab. 4 (c), EvoPrompt requires only 0.764M trainable parameters when trained for 5 epochs, which is comparable to Evolving Prompt Adaptation for Vision-Language Models13 Fig. 3: EvoPrompt performance comparison in few-shot image recognition setting. Table 4: (a) Ablation study on core components. (b) Sensitivity analysis of loss weights η forL kcl andγ forL fgr . (c) Comparison of training efficiency on ImageNet. Trainable parameter(M), training time (ms/image) and FPS (batch size=100) are reported. Variants Base Novel HM w.o. MPP 75.32 70.15 72.64 w.o. W s 75.80 71.42 73.54 w.o. AB 76.15 70.90 73.43 w.o. E.T. 77.42 70.25 73.66 w.o. L kcl 77.24 70.55 73.74 w.o. L fgr 76.70 70.52 73.48 EvoP76.9871.8074.29 (a) Config Base Novel HM γ 10 76.92 71.62 74.18 25 76.98 71.80 74.29 50 76.90 71.72 74.21 100 76.82 71.55 74.09 η 0.2 77.05 71.48 74.17 0.5 76.98 71.80 74.29 1.0 76.83 71.73 74.18 2.0 76.45 71.30 73.80 (b) Method Params Time FPS MaPLe 3.555 39.5 1757.6 PSRC 0.046 40.0 1764.2 ProVP 0.147 6.4 928.9 MetaP 0.031 30.7 659.8 TCP0.332 5.3 950.6 MMA 0.675 2.2 688.5 EvoP0.7644.51282.1 (c) or fewer than most efficient prior methods. Meanwhile, it attains a fast infer- ence speed of 1282.1 FPS and requires only 4.5ms of training time per image. This efficiency stems from our lightweight design. The decoupled MPP struc- ture cuts parameters by 4.6× compared to MaPLe. By freezing historical direc- tional updates and optimizing only their magnitude coefficients, the expansion of learnable parameters throughout training remains minimal. Additionally, our adaptive rank reduction mechanism progressively decreases the rank of low-rank adapters in later epochs, naturally limiting parameter growth. Together, these strategies ensure that EvoPrompt maintains a lightweight and stable parameter footprint while delivering scalable adaptation performance. 14Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li Fig. 4: Analysis of training dynamics and performance. (a) The evolution of learnable magnitudesα i . (b, c) Performance comparison between MaPLe and EvoPrompt, where vertical dashed lines indicate training breakpoints. Evolution of Learned Magnitudes Analysis of the learned magnitude coef- ficients α on ImageNet reveals a distinct and stable evolutionary pattern across 10 training epochs, as shown in Fig. 4 (a). The values do not peak at the initial epoch α 1 , which can be attributed to the inherent instability at the start of training as the model begins to explore and adapt to the prompt space. Instead, they rise rapidly to a maximum at α 2 . This is followed by a gradual decline in later epochs. The pattern suggests a hierarchical importance: the prompt rep- resentation quickly consolidates core features using directions established early in training around α 2 , while later epochs contribute directions with diminishing magnitudes, primarily serving for fine-grained adjustment without drastically altering the established semantic space. Overfitting Phenomenon We analyze the training dynamics of MaPLe and EvoPrompt on the ImageNet, as illustrated in Fig. 4 (b,c). A critical “breakpoint” signifies a phase transition. Before this point, both methods learn transferable features, evidenced by joint performance gains on base and novel classes. After the breakpoint, MaPLe begins to over-specialize on the base training data, lead- ing to unrecoverable performance degradation on novel classes despite improving base accuracy. In contrast, EvoPrompt maintains a stable and robust perfor- mance on novel classes after its breakpoint, effectively mitigating the overfitting issue. This demonstrates the superior generalization capability of EvoPrompt, achieved by optimizing the prompt evolution process. 5 Conclusion In this work, we introduced EvoPrompt, a novel guided evolution paradigm for few-shot adaptation of large vision-language models. It introduces a modality- shared prompt projector for efficient cross-modal interaction and an evolution- guided training strategy to preserve pre-trained knowledge. With additional geo- metric regularization, EvoPrompt achieves new state-of-the-art performance on multiple benchmarks while maintaining strong zero-shot generalization, demon- strating an effective and efficient approach to adapting large vision-language models. Evolving Prompt Adaptation for Vision-Language Models15 References 1. Alayrac, J.B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Men- sch, A., Millican, K., Reynolds, M., et al.: Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35, 23716– 23736 (2022) 2. Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: Visual question answering. In: Proceedings of the IEEE International Con- ference on Computer Vision. p. 2425–2433 (2015) 3. Bossard, L., Guillaumin, M., Van Gool, L.: Food-101–mining discriminative com- ponents with random forests. In: Eur. Conf. Comput. Vis. p. 446–461 (2014) 4. Brown, T., Mann, B., et al.: Language models are few-shot learners. In: Adv. Neural Inform. Process. Syst. vol. 33, p. 1877–1901 (2020) 5. Chen, G., Yao, W., Song, X., Li, X., Lu, Y., Gao, Y.: Plot: Prompt learning with optimal transport for vision-language models. In: Int. Conf. Learn. Represent. (2023) 6. Chen, S., Ge, C., Tong, Z., Wang, J., Song, Y., Wang, J., Luo, P.: Adaptformer: Adapting vision transformers for scalable visual recognition. In: Adv. Neural In- form. Process. Syst. vol. 35, p. 16664–16678 (2022) 7. Chen, X., Djolonga, J., Padlewski, P., Mustafa, B., Changpinyo, S., Wu, J., Ruiz, C.R., Goodman, S., Wang, X., Tay, Y., et al.: On scaling up a multilingual vision and language model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 14432–14444 (2024) 8. Chen, Y.C., Li, L., Yu, L., El Kholy, A., Ahmed, F., Gan, Z., Cheng, Y., Liu, J.: Uniter: Universal image-text representation learning. In: European Conference on Computer Vision. p. 104–120. Springer (2020) 9. Cimpoi, M., Maji, S., Kokkinos, I., Mohamed, S., Vedaldi, A.: Describing textures in the wild. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 3606–3613 (2014) 10. Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large- scale hierarchical image database. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 248–255 (2009) 11. Dong, J.D., et al.: Zegclip: Towards adapting clip for zero-shot semantic segmen- tation. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 11112–11121 (2023) 12. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020) 13. Du, Y., Wei, F., Zhang, Z., Shi, M., Gao, Y., Li, G.: Learning to prompt for open- vocabulary object detection with vision-language model. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 14084–14093 (2022) 14. Fei-Fei, L., Fergus, R., Perona, P.: Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object cate- gories. In: 2004 Conference on Computer Vision and Pattern Recognition Work- shop. p. 178–178 (2004) 15. Gao, P., Geng, S., Zhang, R., Ma, T., Fang, R., Zhang, Y., Li, H., Qiao, Y.: Clip- adapter: Better vision-language models with feature adapters. Int. J. Comput. Vis. 132(2), 581–595 (2024) 16. Gong, L., He, D., Li, Z., Qin, T., Wang, L., Liu, T.: Efficient training of bert by progressively stacking. In: International conference on machine learning. p. 2337–2346. PMLR (2019) 16Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li 17. Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. p. 6904–6913 (2017) 18. Gu, X., Lin, T.Y., Kuo, W., Cui, Y.: Open-vocabulary object detection via vision and language knowledge distillation. In: Int. Conf. Learn. Represent. (2022) 19. Helber, P., Bischke, B., Dengel, A., Borth, D.: Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 12(7), 2217– 2226 (2019) 20. Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F., Dorundo, E., Desai, R., Zhu, T., Parajuli, S., Guo, M., et al.: The many faces of robustness: A critical analysis of out-of-distribution generalization. In: Int. Conf. Comput. Vis. p. 8340– 8349 (2021) 21. Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., Song, D.: Natural adversarial examples. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 15262–15271 (2021) 22. Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Ges- mundo, A., Attariyan, M., Gelly, S.: Parameter-efficient transfer learning for nlp. In: Int. Conf. Mach. Learn. p. 2790–2799 (2019) 23. Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: Lora: Low-rank adaptation of large language models. In: Int. Conf. Learn. Represent. (2022) 24. Huang, S., Dong, L., Wang, W., Hao, Y., Singhal, S., Ma, S., Lv, T., Cui, L., Mo- hammed, O.K., Patra, B., et al.: Language is not all you need: Aligning perception with language models. Advances in Neural Information Processing Systems 36, 72096–72109 (2023) 25. Ji, Z., et al.: Dept: Decoupled prompt tuning. In: Adv. Neural Inform. Process. Syst. (2023) 26. Jia, C., Yang, Y., Xia, Y., Chen, Y.T., Parekh, Z., Pham, H., Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-language representation learning with noisy text supervision. In: Int. Conf. Mach. Learn. p. 4904–4916 (2021) 27. Khattak, M.U., Rasheed, H., Maaz, M., Khan, S., Khan, F.S.: Maple: Multi-modal prompt learning. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 19113–19122 (2023) 28. Khattak, M.U., Wasim, S.T., Naseer, M., Khan, S., Yang, M.H., Khan, F.S.: Self- regulating prompts: Foundational model adaptation without forgetting. In: Int. Conf. Comput. Vis. p. 15190–15200 (2023) 29. Krause, J., Stark, M., Deng, J., Fei-Fei, L.: 3d object representations for fine- grained categorization. In: Proceedings of the IEEE International Conference on Computer Vision Workshops. p. 554–561 (2013) 30. Lee, K.H., Chen, X., Hua, G., Hwang, H., Chang, K.W.: Stacked cross attention for image-text matching. In: Proceedings of the European Conference on Computer Vision (ECCV). p. 201–216 (2018) 31. Lester, B., Al-Rfou, R., Constant, N.: The power of scale for parameter-efficient prompt tuning. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. p. 3045–3059 (2021) 32. Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models. In: Int. Conf. Mach. Learn. p. 19730–19742 (2023) Evolving Prompt Adaptation for Vision-Language Models17 33. Li, X.L., Liang, P.: Prefix-tuning: Optimizing continuous prompts for generation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. p. 4582–4597 (2021) 34. Liu, S.Y., Wang, C.Y., Yin, H., Molchanov, P., Wang, Y.C.F., Cheng, K.T., Kautz, J.: Dora: Weight-decomposed low-rank adaptation. In: Int. Conf. Mach. Learn. (2024) 35. Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., Tang, J.: P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks. In: Muresan, S., Nakov, P., Villavicencio, A. (eds.) Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). p. 61–68. Association for Computational Linguistics, Dublin, Ireland (May 2022). https: //doi.org/10.18653/v1/2022.acl-short.8, https://aclanthology.org/2022. acl-short.8/ 36. Lu, Y., Liu, J., Zhang, Y., Liu, Y., Tian, X.: Prompt distribution learning. In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion. p. 5206–5215 (2022) 37. Maji, S., Rahtu, E., Kannala, J., Blaschko, M., Vedaldi, A.: Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151 (2013) 38. Nilsback, M.E., Zisserman, A.: Automated flower classification over a large number of classes. In: 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing. p. 722–729 (2008) 39. Oord, A.v.d., Li, Y., Vinyals, O.: Representation learning with contrastive predic- tive coding. arXiv preprint arXiv:1807.03748 (2018) 40. Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C.: Cats and dogs. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 3498–3505 (2012) 41. Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., Schölkopf, B.: Controlling text-to-image diffusion by orthogonal finetuning. Ad- vances in Neural Information Processing Systems 36, 79320–79362 (2023) 42. Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. p. 8748–8763 (2021) 43. Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., Lu, J.: Denseclip: Language-guided dense prediction with context-aware prompting. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 18082–18091 (2022) 44. Recht, B., Roelofs, R., Schmidt, L., Shankar, V.: Do imagenet classifiers generalize to imagenet? In: Int. Conf. Mach. Learn. p. 5389–5400 (2019) 45. Rényi, A.: On measures of dependence. Acta mathematica hungarica 10(3-4), 441– 451 (1959) 46. Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., et al.: Laion-5b: An open large-scale dataset for training next generation image-text models. In: Adv. Neural Inform. Process. Syst. vol. 35, p. 25278–25294 (2022) 47. Soomro, K., Zamir, A.R., Shah, M.: Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012) 48. Sung, Y.L., Cho, J., Bansal, M.: Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 5227–5237 (2022) 49. Wang, H., Ge, S., Lipton, Z., Xing, E.P.: Learning robust global representations by penalizing local predictive power. Advances in Neural Information Processing Systems 32 (2019) 18Enming Zhang, Jiayang Li, Yanru Wu, and Zhenyu Liu Yang Li 50. Wang, L., Wu, J., Huang, S.L., Zheng, L., Xu, X., Zhang, L., Huang, J.: An efficient approach to informative feature extraction from multimodal data. In: Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence. AAAI’19/IAAI’19/EAAI’19, AAAI Press (2019). https://doi.org/10.1609/ aaai.v33i01.33015281, https://doi.org/10.1609/aaai.v33i01.33015281 51. Wang, Z., Zhang, Z., Lee, C.Y., et al.: Dualprompt: Complementary prompting for continual learning. In: Eur. Conf. Comput. Vis. p. 258–275. Springer (2022) 52. Wu, Y., Piao, H., Huang, L.K., Wang, R., Li, W., Pfister, H., Meng, D., Ma, K., Wei, Y.: Sd-lora: Scalable decoupled low-rank adaptation for class incremental learning. arXiv preprint arXiv:2501.13198 (2025) 53. Xiao, J., Hays, J., Ehinger, K.A., Oliva, A., Torralba, A.: Sun database: Large- scale scene recognition from abbey to zoo. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 3485–3492 (2010) 54. Xu, C., Zhu, Y., Shen, H., Chen, B., Liao, Y., Chen, X., Wang, L.: Progressive visual prompt learning with contrastive feature re-formation. International Journal of Computer Vision 133(2), 511–526 (2025) 55. Yang, L., Zhang, R.Y., Wang, Y., Xie, X.: Mma: Multi-modal adapter for vision- language models. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 23826–23837 (2024) 56. Yano, K., Takase, S., Kobayashi, S., Kiyono, S., Suzuki, J.: Efficient construction of model family through progressive training using model expansion. arXiv preprint arXiv:2504.00623 (2025) 57. Yao, H., Zhang, R., Xu, C.: Visual-language prompt tuning with knowledge-guided context optimization. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 6757–6767 (2023) 58. Yao, H., Zhang, R., Xu, C.: Tcp: Textual-based class-aware prompt tuning for visual-language model. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 23438– 23448 (2024) 59. Yu, J., Wang, Z., Vasudevan, V., Yeung, L., Seyedhosseini, M., Wu, Y.: Coca: Contrastive captioners are image-text foundation models. Transactions on Machine Learning Research (2022) 60. Yu, Q., Chen, J., et al.: Task residual for tuning vision-language models. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 10899–10909 (2023) 61. Zhang, R., Zhang, W., Rong, R., Li, C., Qiu, Y., Cardie, C., et al.: Tip-adapter: Training-free adaption of clip for few-shot classification. In: Eur. Conf. Comput. Vis. p. 493–510. Springer (2022) 62. Zhao, C., Wang, Y., Jiang, X., Shen, Y., Song, K., Li, D., Miao, D.: Learning domain invariant prompt for vision-language models. IEEE Transactions on Image Processing 33, 1348–1360 (2024) 63. Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Conditional prompt learning for vision- language models. In: IEEE Conf. Comput. Vis. Pattern Recog. p. 16816–16825 (2022) 64. Zhou, K., Yang, J., Loy, C.C., Liu, Z.: Learning to prompt for vision-language models. Int. J. Comput. Vis. 130(9), 2337–2348 (2022) 65. Zhu, B., Niu, Y., Han, Y., Wu, Y., Zhang, H.: Prompt-aligned gradient for prompt tuning. In: Proceedings of the IEEE/CVF international conference on computer vision. p. 15659–15669 (2023)