Paper deep dive
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
Hongyuan Liu, Qinli Yang, Wen Li, Zhong Zhang, Jiaming Liu, Wei Han, Zhili Qin, Jinxia Guo, Junming Shao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/2/2026, 3:25:50 AM
Summary
The paper introduces TPC-CMA (Three-Phase Curriculum for Cross-Modal Alignment), a framework designed to resolve the 'modality gap' in Vision-Language Models (VLMs) like CLIP. The authors decompose the modality gap into a Centroid Gap and a Distribution Gap, demonstrating that the latter is the primary predictor of cross-modal task performance. TPC-CMA employs a two-part loss function—Negative Reweighting and Intra-modal Geometry Matching—combined with a three-phase curriculum to progressively align feature manifolds without sacrificing zero-shot performance.
Entities (5)
Relation Signals (3)
Distribution Gap → predicts → Cross-modal task quality
confidence 98% · demonstrate that the Distribution Gap is the true predictor of cross-modal task quality (R2 = 0.986)
CLIP → exhibits → Modality Gap
confidence 95% · Despite their success, a persistent modality gap exists [in CLIP].
TPC-CMA → reduces → Modality Gap
confidence 95% · TPC-CMA... a fine-tuning framework that explicitly reduces both components [of the modality gap].
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-Language Models (VLMs) such as CLIP learn a shared embedding space for images and text, yet their representations remain geometrically separated, a phenomenon known as the modality gap. This gap limits tasks requiring cross-modal interchangeability, such as captioning and joint clustering. Existing post-processing approaches can partially improve cross-modal compatibility; however, we show through geometric analysis that they primarily reduce the global centroid offset while leaving the underlying distributional mismatch intact. We decompose the modality gap into a Centroid Gap and a Distribution Gap, and demonstrate that the Distribution Gap is the true predictor of cross-modal task quality ($R^2 = 0.986$), whereas the commonly used Raw Gap is misleading ($R^2 = 0.691$). Motivated by this observation, we propose TPC-CMA (Three-Phase Curriculum for Cross-Modal Alignment), a fine-tuning framework that explicitly reduces both components. The proposed CMA jointly mitigates centroid offsets and reshapes the distributional structure, while a three-phase curriculum with gradient-aware scheduling progressively introduces alignment during training to enable stable optimization. Experiments demonstrate that our method significantly improves cross-modal alignment. With $\alpha_{\text{target}}{=}0.05$, the modality gap is reduced by 66.6\% with only 4.84\% accuracy drop. Under stronger alignment ($\alpha_{\text{target}}{=}0.5$), the gap is reduced by 82.3\%, clustering ARI improves from 0.318 to 0.516, and captioning CIDEr increases by 57.1\% over the original model. Our code and pre-trained models will be made publicly available upon acceptance.
Tags
Links
- Source: https://arxiv.org/abs/2604.00279v1
- Canonical: https://arxiv.org/abs/2604.00279v1
Trouble viewing inline? Open PDF directly →
Full Text
61,732 characters extracted from source content.
Expand or collapse full text
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment Hongyuan Liu1 Qinli Yang1 Wen Li2 Zhong Zhang1 Jiaming Liu1 Wei Han1 Zhili Qin1 Jinxia Guo1 Junming Shao1 1University of Electronic Science and Technology of China 2University of Bristol Abstract Vision-Language Models (VLMs) such as CLIP learn a shared embedding space for images and text, yet their representations remain geometrically separated, a phenomenon known as the modality gap. This gap limits tasks requiring cross-modal interchangeability, such as captioning and joint clustering. Existing post-processing approaches can partially improve cross-modal compatibility; however, we show through geometric analysis that they primarily reduce the global centroid offset while leaving the underlying distributional mismatch intact. We decompose the modality gap into a Centroid Gap and a Distribution Gap, and demonstrate that the Distribution Gap is the true predictor of cross-modal task quality (R2=0.986R^2=0.986), whereas the commonly used Raw Gap is misleading (R2=0.691R^2=0.691). Motivated by this observation, we propose TPC-CMA (Three-Phase Curriculum for Cross-Modal Alignment), a fine-tuning framework that explicitly reduces both components. The proposed CMA jointly mitigates centroid offsets and reshapes the distributional structure, while a three-phase curriculum with gradient-aware scheduling progressively introduces alignment during training to enable stable optimization. Experiments demonstrate that our method significantly improves cross-modal alignment. With αtarget=0.05 _target=0.05, the modality gap is reduced by 66.6% with only 4.84% accuracy drop. Under stronger alignment (αtarget=0.5 _target=0.5), the gap is reduced by 82.3%, clustering ARI improves from 0.318 to 0.516, and captioning CIDEr increases by 57.1% over the original model. Our code and pre-trained models will be made publicly available upon acceptance. Keywords: Modality Gap, Vision-Language Models, Curriculum Learning, Multi-task Optimization, Zero-shot Learning 1 Introduction Figure 1: The t-SNE visualization of multimodal features (top) and downstream performance of different methods (bottom). (a) Original CLIP: image and text embeddings form separated islands. (b) Mean-CenteringLiang et al. (2022); Li et al. (2025); An et al. (2025); Yamashita et al. (2025): centroids overlap but distributional mismatch persists. (c) Our alignment method (TPC-CMA, αtarget=0.5 _target=0.5): true semantic interleaving. Each panel reports the Modality Gap, defined as the residual cosine distance after centroid alignment (see Section 2.2); lower values indicate better structural alignment. Bottom: downstream impact on two zero-shot tasks. (d) Image captioning CIDEr improves from 0.210 to 0.330. (e) Joint clustering ARI improves from 0.318 to 0.516. Vision-Language Models (VLMs) such as CLIP Radford et al. (2021), ALIGN Jia et al. (2021), SigLIP Zhai et al. (2023), and their successors Li et al. (2023a); Yu et al. (2022); Sun et al. (2024); Cherti et al. (2023); Chen et al. (2024); Mu et al. (2022) have become foundational to modern multimodal AI, enabling powerful zero-shot transfer Zhou et al. (2022); Gao et al. (2024); Zhai et al. (2022) via a joint image-text embedding space. Despite their success, a persistent modality gap exists: image and text embeddings occupy disjoint regions of the hypersphere Liang et al. (2022); Fahim et al. (2024); Schrodi et al. (2025); Ramasinghe et al. (2024); Huang et al. (2025). While this gap has limited impact on zero-shot classification, it restricts tasks that require cross-modal feature interchangeability. As shown in Figure 1(d) and (e), directly feeding image embeddings into a text decoder results in poor captioning performance, and joint image-text clustering yields limited semantic grouping quality. These results indicate that the two modalities fail to form a unified representation space, limiting the applicability of VLMs in generation Grassucci et al. (2026); Li et al. (2023b); Mokady et al. (2021); Tewel et al. (2022); Zeng et al. (2024), reasoning Xu et al. (2024), and broader multimodal understanding tasks Liu et al. (2023). To mitigate this issue, methods that operate on pretrained embeddings have been widely explored due to their practical advantages. These approaches avoid expensive retraining and preserve the generalization ability of existing VLMs, and can partially improve cross-modal compatibility, as reflected by measurable gains in downstream tasks (Figure 1(d)(e)). However, the improvements remain limited, especially for tasks requiring deeper cross-modal interaction. This observation suggests that current approaches address only a limited aspect of the underlying modality gap. In this work, we analyze this limitation from a geometric perspective. We use representative approaches, such as Mean-Centering Liang et al. (2022); Li et al. (2025); An et al. (2025); Yamashita et al. (2025), as a diagnostic tool to examine how cross-modal alignment is achieved in practice. As shown in Figure 1(a), image and text embeddings form two separated regions in the feature space, indicating a fundamental structural discrepancy. We decompose the modality gap into two components: a Centroid Gap, which captures the global shift between modalities, and a Distribution Gap, which reflects differences in their underlying geometric structure. Through this lens, we find that existing approaches primarily reduce the Centroid Gap while leaving the Distribution Gap largely unchanged. This leads to superficial alignment as illustrated in Figure 1(b), where embeddings become closer in position but remain structurally mismatched. We argue that effective alignment requires consistent geometric structure across modalities. These observations indicate that resolving the modality gap requires modifying the feature geometry itself, rather than only applying global transformations. To address this limitation, we propose the Three-Phase Curriculum for Cross-Modal Alignment (TPC-CMA), a framework that explicitly reduces both centroid and distribution gaps progressively, which enables an effective geometric-consistent cross-modal alignment as shown in Figure 1(c). We introduce a Cross-Modal Alignment (CMA) loss that replaces cross-modal repulsion with intra-modal geometry matching, encouraging consistent structures across modalities. We further design a three-phase curriculum strategy with gradient-aware scheduling Chen et al. (2018) that gradually transitions from the original CLIP objective to the alignment objective, enabling stable adaptation of pretrained representations. Extensive experiments demonstrate that our method significantly improves tasks requiring structural cross-modal alignment, while maintaining strong zero-shot performance. The summary of our contributions is as follows: 1. We propose TPC-CMA, a cross-modal alignment framework that reduces the modality gap in a progressive manner. 2. We introduce a CMA loss that explicitly addresses both centroid and distribution discrepancies by replacing cross-modal repulsion with intra-modal geometry matching, enabling structurally consistent cross-modal alignment. 3. We design a TPC strategy with gradient-aware ramp-up Chen et al. (2018) that gradually transitions from the original CLIP objective to the alignment objective, enabling stable adaptation of pretrained feature geometry without abrupt disruption. 4. We demonstrate that Distribution Gap, not the commonly used Raw Gap, is a near-perfect predictor of cross-modal task quality (R2=0.986R^2=0.986 vs. 0.6910.691). Extensive experiments show up to 82.3% gap reduction, substantial gains in clustering (ARI: 0.318 → 0.551) and captioning (CIDEr +57.1%), while maintaining strong zero-shot performance. 2 Preliminary 2.1 Pattern Analysis of CLIP CLIP Radford et al. (2021) learns a shared embedding space via contrastive learning. Given an image encoder fIf_I and a text encoder fTf_T, inputs are mapped to ℓ2 _2-normalized features i=fI(xi)v_i=f_I(x_i) and i=fT(yi)t_i=f_T(y_i), with similarity defined as s(i,j)=i⊤js(v_i,t_j)=v_i t_j. The model is trained using a symmetric InfoNCE loss van den Oord et al. (2018): ℒCLIP=−12N∑i=1N[logeτi⊤i∑jeτi⊤j+logeτi⊤i∑jeτi⊤j]L_CLIP=- 12N _i=1^N [ e _i t_i _je _i t_j+ e _i v_i _je _i v_j ] (1) where τ is a learnable temperature. This objective can be decomposed into an attractive term ℒalignL_align and a repulsive term ℒopposeL_oppose. Following Shi et al. Shi et al. (2023), the image-to-text direction decomposes into: ℒi2t _i2t =−1N∑i=1Nτi⊤i⏟ℒalign:attraction = - 1N _i=1^Nτ\,v_i t_i_L_align:\;attraction +1N∑i=1Nlog∑j=1Neτi⊤j⏟ℒoppose:repulsion \;+\; 1N _i=1^N _j=1^Ne^τ\,v_i t_j_L_oppose:\;repulsion (2) with an analogous decomposition for the text-to-image direction. Since each positive pair is contrasted against N−1N-1 negatives, the cumulative repulsive gradient dominates, introducing a structural asymmetry between attraction and repulsion. As a result, CLIP representations exhibit a characteristic pattern: matched pairs are locally aligned, while image and text embeddings remain globally separated into distinct regions on the hypersphere. From this perspective, the modality gap can be attributed to the intrinsic geometry induced by the contrastive objective rather than initialization Liang et al. (2022). We provide a theoretical analysis in Appendix A showing that the opposition term encourages the two modalities to separate in the feature space, leading to an antipodal tendency. This asymmetry explains the coexistence of local alignment and global separation, which we further analyze in the following section. 2.2 Modality Gap Decomposition To better understand this phenomenon, we analyze the modality gap from two complementary geometric perspectives. Given N image-text pairs with ℓ2 _2-normalized image embeddings iv_i and text embeddings it_i, the standard modality gap is defined as Gap=1−1N∑i⊤iGap=1- 1N _iv_i t_i, which we refer to as the Raw Gap, conflates two geometrically distinct phenomena Role et al. (2025). We decompose it into (formal definitions in Appendix B): • Centroid Gap (CG_C): ‖¯−¯‖2\| v- t\|_2, where ¯=1N∑i v= 1N _iv_i and ¯=1N∑i t= 1N _it_i are the modal centroids. • Distribution Gap (DG_D): the residual gap after centroid alignment, measured as the average cosine distance between centered embeddings ^i=(i−¯)/‖i−¯‖ v_i=(v_i- v)/\|v_i- v\| and ^i=(i−¯)/‖i−¯‖ t_i=(t_i- t)/\|t_i- t\|. CG_C and DG_D capture complementary geometric aspects of the modality gap rather than forming a strict additive decomposition. Intuitively, CG_C measures how far apart the two modal clouds are, while DG_D characterizes how differently they are structured. Effective cross-modal alignment requires reducing both components, as correcting centroid shift alone does not resolve structural mismatch. Figure 2: Overview of TPC-CMA. A: CLIP backbone extracts image and text embeddings. B: CMA Loss combines Negative Reweighting (reduces Centroid Gap) and Intra-modal Geometry Matching (reduces Distribution Gap) via α. C: Three-Phase Curriculum schedules α from 0 to αtarget _target across Anchor, Gradient-aware Ramp-up, and Stabilize stages, with the transition speed dynamically modulated by the observed gradient dynamics between loss terms. Figure 3: Gap decomposition under Mean-Centering. Despite a 97% Centroid Gap reduction, both Distribution Gap and ROUGE-L remain unchanged, showing that correcting the centroid offset alone does not improve cross-modal compatibility. Analysis Results: Figure 3 applies this decomposition to Mean-Centering. While the Centroid Gap drops by 97% (0.686 → 0.020), the Distribution Gap remains unchanged (0.685). We provide a theoretical analysis in Appendix B (Proposition B.1) supporting that translation-based methods preserve DG_D by shifting centroids without altering relative deviations, and ROUGE-L remains at 0.250, confirming no downstream improvement. This shows that post-processing methods reduce global offsets but fail to address structural discrepancies. Since cross-modal interchangeability is governed by distributional consistency, the unresolved distribution gap becomes the primary bottleneck, motivating alignment strategies that directly enforce structural consistency in the embedding space. 3 Methodology Bridging the modality gap requires addressing the structural discrepancy between modalities, particularly the distributional mismatch that existing approaches do not resolve. Our approach combines a Cross-Modal Alignment (CMA) mechanism to reduce both centroid and distributional gaps (Section 3.2), with a three-phase training strategy (Section 3.3) that gradually introduces alignment during optimization to ensure stable convergence. 3.1 Overview As illustrated in Figure 2, our framework consists of three parts. In Part A, image-text pairs are processed by a CLIP backbone to obtain feature embeddings. In Part B, we introduce a CMA module that reduces the modality gap by addressing both centroid offsets and distributional discrepancies. In Part C, we design a TPC that progressively introduces the alignment objective during training to ensure stable and effective convergence. 3.2 Cross-Modal Alignment Section 2.1 shows that the modality gap arises from both centroid offset and distributional mismatch. Existing methods typically address only one aspect, leading to limited improvement in cross-modal alignment. To overcome this limitation, we propose a CMA mechanism that jointly targets both components through two complementary designs: (1) Negative Reweighting, which reduces the Centroid Gap, and (2) Intra-modal Geometry Matching, which reshapes the distributional structure to reduce the Distribution Gap. 3.2.1 Mechanism 1: Negative Reweighting As shown in Eq. 2, the opposition term ℒopposeL_oppose accumulates repulsive gradients from all N−1N-1 unmatched pairs, pushing the two modal distributions toward antipodal regions and sustaining the Centroid Gap. Specifically, to mitigate this effect, we propose Negative Reweighting, which weakens the contribution of negative pairs in the cross-modal similarity matrix. Let ∈ℝN×NS ^N× N be the cross-modal logit matrix with Sij=τi⊤jS_ij= _i t_j. We define a weight mask M: Mij=1if i=j1−βif i≠jM_ij= cases1&if i=j\\ 1-β&if i≠ j cases (3) where β=0.05αβ=0.05α controls the strength of reweighting in Eq. 8. We set the coefficient to 0.05 so that β remains small (at most 0.05), ensuring that individual negative logits are only mildly suppressed while the cumulative effect over N−1N-1 negatives is substantial. The reweighted logits ⊙M are used in place of S when computing the cross-entropy loss: ℒrw=CE(⊙,)+CE(⊙⊤,)L_rw=CE(M ,\;y)+CE(M ,\;y) (4) where CE denotes the cross-entropy loss, ⊙ denotes element-wise (Hadamard) product, and =(1,2,…,N)y=(1,2,…,N) are the ground-truth labels. Geometrically, the reduced denominator weakens the gradient on negative pairs, allowing the centroid gap to shrink while still matching pairs correctly. However, reweighting only removes the repulsive barrier without reshaping the distributions. 3.2.2 Mechanism 2: Intra-modal Geometry Matching As shown in Section 2.2, even after the Centroid Gap is reduced, the Distribution Gap remains unchanged because the two modalities retain different manifold shapes. Negative Reweighting removes the repulsive barrier but does not prescribe how the underlying manifold geometry should be restructured to achieve distributional consistency. To close the Distribution Gap, we need to reshape the manifolds so that paired embeddings occupy corresponding positions within their respective distributions. We therefore propose Intra-modal Geometry Matching (IGM), which replaces the cross-modal comparison (“is this image close to its text among all texts?”) with an intra-modal comparison (“is this image’s paired text closer than all other texts to this text?”), forcing the two manifolds to adopt geometrically consistent and matching structures. Specifically, we construct two auxiliary logit matrices: S^ij(txt)=τi⊤iif i=jτi⊤jif i≠j S^(txt)_ij= casesτ\,t_i v_i&if i=j\\ τ\,t_i t_j&if i≠ j cases (5) S^ij(img)=τi⊤iif i=jτi⊤jif i≠j S^(img)_ij= casesτ\,v_i t_i&if i=j\\ τ\,v_i v_j&if i≠ j cases (6) Consider row i in S^(txt) S^(txt): off-diagonal entries τi⊤jτ\,t_i t_j capture the intra-modal structure of text i, while the diagonal τi⊤iτ\,t_i v_i measures cross-modal similarity to the paired image. Cross-entropy forces the paired image to rank higher than all other texts: ℒintra=CE(^(txt),)+CE(^(img),)L_intra=CE( S^(txt),\;y)+CE( S^(img),\;y) (7) By minimizing CE(^(txt),)CE( S^(txt),y), each image iv_i is forced to be closer to its paired text it_i than any other text, effectively “inserting” each image into the structure of its paired modality. The symmetric term CE(^(img),)CE( S^(img),y) does the same in reverse. Over the full batch, this bidirectional matching reshapes both manifolds to adopt the same structure, thereby closing the Distribution Gap. Unlike the standard CLIP loss that compares across modalities, ℒintraL_intra compares within a single modality, replacing cross-modal opposition with intra-modal uniformity terms Zbontar et al. (2021); Bardes et al. (2022); Grill et al. (2020) whose optima are consistent with alignment (Appendix C, Proposition C.1). 3.2.3 Combined CMA Loss The two mechanisms are combined with the alignment parameter α∈[0,1]α∈[0,1], which controls the trade-off between standard contrastive learning and geometry alignment. Let sijvt=τi⊤js_ij^vt= _i t_j, sijtt=τi⊤js_ij^t= _i t_j, sijvv=τi⊤js_ij^v= _i v_j, and β=0.05αβ=0.05α. The overall CMA loss is defined as: ℒCMA=12[(1−α)ℒrw+αℒintra] _CMA= 12 [(1-α)\,L_rw+α\,L_intra ] =−12N∑i[(1−α)(logesiivtesiivt+∑j≠ie(1−β)sijvt =- 12N _i [(1-α)\! ( e^s_i^vte^s_i^vt+ j≠ i Σe^(1-β)s_ij^vt +logesiivtesiivt+∑j≠ie(1−β)sjivt) + e^s_i^vte^s_i^vt+ j≠ i Σe^(1-β)s_ji^vt ) +α(logesiivtesiivt+∑j≠iesijtt \;+α\! ( e^s_i^vte^s_i^vt+ j≠ i Σe^s_ij^t +logesiivtesiivt+∑j≠iesijvv)] + e^s_i^vte^s_i^vt+ j≠ i Σe^s_ij^v ) ] (8) where sijvt=τi⊤js_ij^vt= _i t_j denotes the scaled cross-modal similarity, sijtt=τi⊤js_ij^t= _i t_j and sijvv=τi⊤js_ij^v= _i v_j denote intra-modal similarities, α∈[0,1]α∈[0,1] is the alignment parameter, and β=0.05αβ=0.05α controls the negative reweighting strength. The first two terms correspond to the reweighted contrastive loss ℒrwL_rw, where (1−β)(1-β) down-scales cross-modal negative logits; the last two terms correspond to the intra-modal geometry matching loss ℒintraL_intra, where negatives come from the same modality. Note that all four terms share the standard log-softmax form −log(es+/∑jesj)- (e^s_+/ _je^s_j ); they differ only in which similarities populate the denominator (see Appendix E for complete pseudocode). When α=0α=0, β=0β=0 so ℒrwL_rw reduces to ℒCLIPL_CLIP. As α increases, the loss continuously and smoothly transitions from standard contrastive learning toward full intra-modal geometry matching. 3.3 Three-Phase Curriculum with Gradient-Aware Ramp While the CMA loss explicitly enforces cross-modal alignment, directly applying the full objective throughout training may lead to suboptimal convergence, as the pretrained representations are not immediately adapted to the modified objective. To address this, we introduce a TPC with gradient-aware scheduling. The curriculum is governed by a weighting parameter α, which controls the strength of the alignment objective. Specifically, TPC gradually increases α during training, allowing the feature geometry to adapt progressively and converge to a stable optimum. Stage 1: Anchor (c epochs, α=0α=0). The model is trained with standard CLIP loss only, allowing it to adapt to the fine-tuning data while anchoring discriminative structure. All α configurations share the same peak accuracy at the end of this stage. Stage 2: Gradient-aware Ramp-up (r epochs, α: 0→αtarget0→ _target). After c epochs of anchoring, the training shifts into the ramp-up phase. A naive approach would increase α at a constant (linear) rate, but the contrastive loss ℒrwL_rw and the alignment loss ℒintraL_intra exhibit different convergence dynamics: we observe that ‖∇ℒintra‖\| _intra\| dominates early in training (ratio ≈2.3:1≈ 2.3:1), indicating gradient conflict, while the ratio converges toward equilibrium (≈1:1≈ 1:1) as training progresses. Inspired by multi-task gradient balancing Chen et al. (2018), we modulate the ramp speed based on this convergence signal. Concretely, we maintain a slow EMA ℓ¯s _s and a fast EMA ℓ¯f _f of the contrastive loss ℒrwL_rw. Their ratio ρ=ℓ¯f/ℓ¯sρ= _f/ _s serves as a convergence indicator: ρ≈1ρ≈ 1 signals stability, ρ<1ρ<1 signals that the loss is still dropping (not ready for more alignment), and ρ>1ρ>1 signals that alignment is hurting discriminative quality. The per-step α update is: Δα α =αtarget−αT−t⋅(0.5+s(ρ)), = _target-αT-t· (0.5+s(ρ) ), (9) s(ρ) s(ρ) =ρif ρ≤1,2−ρotherwise,ρ=clip(ℓ¯fℓ¯s, 0, 2) = casesρ&if ρ≤ 1,\\ 2-ρ&otherwise, cases ρ=clip\! ( _f _s,\;0,\;2 ) where T−tT-t is the number of remaining ramp steps and the clip operation bounds ρ to [0,2][0,2], ensuring s(ρ)∈[0,1]s(ρ)∈[0,1] and the speed factor (0.5+s(ρ))∈[0.5,1.5](0.5+s(ρ))∈[0.5,1.5]. The base rate (αtarget−α)/(T−t)( _target-α)/(T-t) provides a catch-up mechanism that guarantees α reaches αtarget _target by step T: if previous steps were slow, the remaining budget per step automatically increases. The speed factor modulates around the base rate, accelerating when gradients are balanced (s→1s→ 1, ρ≈1ρ≈ 1) and decelerating when gradient conflict is detected (s→0s→ 0). Stage 3: Stabilize (h epochs, α=αtargetα= _target). Finally, the last few epochs stabilize the finetuned feature space. α is held at the target for convergence to a stable geometric configuration. The full schedule is defined by four values: (c,r,h,αtarget)(c,r,h, _target). We use c=3c=3, r=5r=5, h=2h=2 for a total of 10 epochs across all experiments. The only parameter that users need to choose is αtarget _target, which controls the operating point on the Pareto frontier: αtarget≤0.05 _target≤ 0.05 for minimal accuracy loss (retrieval, classification); αtarget≥0.1 _target≥ 0.1 for deep alignment (clustering, generative tasks). This makes TPC-CMA a protocol: users select αtarget _target based on their downstream task, and the method delivers the optimal geometry. 4 Experiments 4.1 Experimental Setup Datasets and Metrics. We use Conceptual Captions 3M (C3M) Sharma et al. (2018), a widely adopted image-text dataset containing approximately 3.3 million web-crawled pairs, as the training set for all fine-tuning methods. We evaluate on five complementary tasks spanning discriminative, generative, and structural dimensions: (i) ImageNet Deng et al. (2009) zero-shot classification (multi-template, consistent with CLIP’s original protocol); (i) COCO Lin et al. (2014) image-text retrieval (5K test split, R@1); (i) multi-dataset zero-shot classification (CIFAR-10/100, Food-101, Caltech-101, Flowers-102); (iv) zero-shot captioning (COCO, 5K images): following DeCap Li et al. (2023b), we train a separate lightweight GPT-2 Radford et al. (2019) decoder per model variant on text embeddings and directly feed image embeddings as the decoder prefix at test time. Since the decoder is trained exclusively on text embeddings and tested on image embeddings, CIDEr improvements reflect the degree to which image embeddings have become interchangeable with text embeddings, not the decoder’s capacity. We report CIDEr and ROUGE-L; (v) joint image-text clustering: following the group-wise evaluation protocol of Grassucci et al. Grassucci et al. (2026), we sample 200 classes from ImageNet validation (50 images per class), pair each image with a text prompt “a photo of a classname”, combine all image and text embeddings into a single pool (20,000 vectors), and run KMeans with k=200k=200. Ground-truth labels are semantic class identities (shared across modalities). We report V-Measure Rosenberg and Hirschberg (2007) and Adjusted Rand Index (ARI) as clustering quality metrics. Implementation Details. We fine-tune the pre-trained ViT-B/32 Dosovitskiy et al. (2021); Vaswani et al. (2017) on C3M for 10 epochs with a learning rate of 1×10−51×10^-5 and an effective batch size of 4096. The learnable temperature τ is initialized from the pre-trained checkpoint and continues to be optimized during fine-tuning. The three-phase curriculum uses c=3c=3, r=5r=5, h=2h=2 throughout. For generalization, we additionally evaluate ViT-B/16 and ViT-L/14 (see Appendix F). Baselines. We compare against five representative methods mitigating the modality gap with post-processing, fine-tuning, and pre-training stage: (1) Mean-Centering Liang et al. (2022), a post-processing translation that subtracts modal centroids; (2) AlignCLIP Eslami and de Melo (2025) (ICLR’25), which trains a CLIP model from scratch with an alignment objective; (3) M2-Mix Oh et al. (2023) (NeurIPS’23), which generates hard negatives via geodesic interpolation; (4) CLIP-Refine Yamaguchi et al. (2025) (CVPR’25), which applies random-reference feature alignment during pre-training; (5) CS-Aligner Yin et al. (2025) (ICLR’26), which adds a Cauchy–Schwarz divergence term to InfoNCE for distributional alignment. All fine-tuning and post-processing baselines use the same training data, number of epochs, and backbone architecture for a fair comparison. Table 1: Main comparison across gap reduction, discriminative, generative, and structural tasks. Method Type Gap Gap↓ % DG_D ImageNet I2T R@1 T2I R@1 V-Measure CIDEr ROUGE-L Original CLIP Radford et al. (2021) Pretrained 0.733 - 0.685 62.62 53.42 35.36 0.769 0.210 0.250 Mean-Centering Liang et al. (2022) Post-proc 0.685 6.5 0.685 60.23 43.74 34.44 0.766 0.227 0.250 AlignCLIP Eslami and de Melo (2025)† From-scratch 0.353 51.8 0.550 32.79 33.64 22.28 0.749 0.112 0.244 M2-Mix Oh et al. (2023) Fine-tune 0.710 3.1 0.637 54.24 42.98 31.00 0.762 0.197 0.239 CLIP-Refine Yamaguchi et al. (2025) Fine-tune 0.747 −-1.9 0.647 62.99 53.90 36.31 0.765 0.236 0.265 CS-Aligner Yin et al. (2025) Fine-tune 0.734 −-0.1 0.617 61.37 49.15 32.89 0.770 0.174 0.249 TPC-CMA αtarget=0.01 _target=0.01 Fine-tune 0.558 23.8 0.610 60.86 50.96 35.96 0.795 0.227 0.273 TPC-CMA αtarget=0.05 _target=0.05 Fine-tune 0.245 66.6 0.554 57.78 48.56 35.68 0.810 0.254 0.269 TPC-CMA αtarget=0.3 _target=0.3 Fine-tune 0.155 78.9 0.474 55.63 46.44 35.08 0.826 0.293 0.276 TPC-CMA αtarget=0.5 _target=0.5 Fine-tune 0.130 82.3 0.441 54.86 45.48 34.49 0.849 0.330 0.288 TPC-CMA αtarget=0.9 _target=0.9 Fine-tune 0.096 87.0 0.392 52.58 42.60 33.38 0.845 0.356 0.296 † Evaluated with AlignCLIP released checkpoint. Training from scratch is prohibitively expensive yet generally unable to match the discriminative quality of fine-tuning a pre-trained model. Bold indicates the best result and underline the second best. The same convention applies to all subsequent tables. Figure 4: Gap vs. Accuracy Pareto Frontier. TPC-CMA forms a smooth efficient frontier that envelopes all selected baselines. AlignCLIP and M2-Mix fall significantly below the frontier, while CLIP-Refine fails to reduce the gap. 4.2 The Pareto Frontier As shown in Section 2.2, existing methods each address only part of the modality gap, leading to limited or one-sided improvements. To demonstrate that TPC-CMA overcomes this limitation by simultaneously reducing both Centroid Gap and Distribution Gap, we compare it against all baselines on five axes: gap reduction, classification, retrieval, clustering, and captioning. Table 1 summarizes the results, and Figure 4 visualizes their gap vs. accuracy trade-offs. Specifically, each baseline occupies a suboptimal position on the Gap vs. Accuracy frontier. CLIP-Refine Yamaguchi et al. (2025) preserves discriminative quality (ImageNet 62.99%) but fails to reduce the gap (−-1.9%). CS-Aligner Yin et al. (2025) adds a Cauchy–Schwarz divergence penalty that lowers DG_D to 0.617 (↓ 9.9%) yet leaves Raw Gap unchanged (0.734), exhibiting the same centroid-gap-agnostic pattern as CLIP-Refine. M2-Mix Oh et al. (2023) reduces the gap by only 3.1% at a disproportionate accuracy cost (ImageNet −-8.38%). AlignCLIP achieves a low gap (0.353) but sacrifices discriminative quality (ImageNet 32.79%). Mean-Centering reduces the gap by 6.5% but degrades I2T R@1 by 18.1%, the superficial alignment identified in Section 2.2. We analyze why these methods produce limited improvements through the Distribution Gap lens in Section 4.3. In contrast, TPC-CMA’s operating points (blue curve in Figure 4) envelope every baseline, reducing both Raw Gap and DG_D simultaneously. At mild alignment (αtarget=0.01 _target=0.01), DG_D drops from 0.685 to 0.610 with less than 2% ImageNet loss. At strong alignment (αtarget=0.5 _target=0.5), DG_D reaches 0.441 (↓ 36%), CIDEr improves by 57.1%, and ARI reaches 0.516. Importantly, each αtarget _target simultaneously reports results on all five tasks, confirming that TPC-CMA is a single protocol with a continuously adjustable operating point, not a family of task-specific models. These results show that the limitations of prior methods are not inherent to the alignment problem itself. TPC-CMA provides a controllable Pareto frontier by jointly addressing both gap components identified in our decomposition analysis (Section 2.2). Table 2: DeCap zero-shot captioning on COCO. Image embeddings are directly used as the decoder prefix without any learned projection layer. Method DG_D CIDEr BLEU-4 ROUGE-L Original CLIP 0.685 0.210 0.029 0.250 Mean-Centering 0.685 0.227 0.040 0.250 AlignCLIP 0.550 0.112 0.015 0.244 CLIP-Refine 0.647 0.236 0.045 0.265 M2-Mix 0.637 0.197 0.032 0.239 CS-Aligner 0.617 0.174 0.036 0.249 TPC-CMA αtarget=0.01 _target=0.01 0.610 0.227 0.050 0.273 TPC-CMA αtarget=0.05 _target=0.05 0.554 0.254 0.044 0.269 TPC-CMA αtarget=0.3 _target=0.3 0.474 0.293 0.048 0.276 TPC-CMA αtarget=0.5 _target=0.5 0.441 0.330 0.057 0.288 TPC-CMA αtarget=0.9 _target=0.9 0.392 0.356 0.065 0.296 Table 3: Joint image-text clustering following Grassucci et al. Grassucci et al. (2026) (ImageNet val, 200 classes × 50 imgs, KMeans k=200k=200). Method DG_D V-Measure ARI Original CLIP 0.685 0.769 0.318 Mean-Centering 0.685 0.766 0.306 AlignCLIP 0.550 0.749 0.304 M2-Mix 0.637 0.762 0.298 CLIP-Refine 0.647 0.765 0.310 CS-Aligner 0.617 0.770 0.315 TPC-CMA αtarget=0.01 _target=0.01 0.610 0.795 0.330 TPC-CMA αtarget=0.05 _target=0.05 0.554 0.810 0.385 TPC-CMA αtarget=0.3 _target=0.3 0.474 0.826 0.435 TPC-CMA αtarget=0.5 _target=0.5 0.441 0.849 0.516 TPC-CMA αtarget=0.7 _target=0.7 0.417 0.863 0.551 TPC-CMA αtarget=0.9 _target=0.9 0.392 0.845 0.492 4.3 Generative and Structural Tasks To verify that alignment can substantially improve generative and structural capabilities beyond what the original model achieves, we evaluate TPC-CMA on two tasks that directly depend on cross-modal geometric coherence: zero-shot captioning without projection (Table 2), which tests feature interchangeability, and joint image-text clustering (Table 3), which tests the quality of unified semantic grouping across modality boundaries. Figure 5: Distribution Gap vs. Raw Gap as predictors of DeCap CIDEr score. (a) Distribution Gap is a near-perfect predictor (R2=0.986R^2=0.986); (b) Raw Gap yields a substantially weaker fit (R2=0.691R^2=0.691). Specifically, for zero-shot captioning (Table 2), we feed image embeddings directly to a text decoder without DeCap’s memory-bank projection, testing cross-modal feature interchangeability. For TPC-CMA, Distribution Gap DG_D decreases monotonically from 0.610 (αtarget=0.01 _target=0.01) to 0.392 (αtarget=0.9 _target=0.9), and CIDEr increases correspondingly from 0.227 to 0.356 (69.5% higher than the Original CLIP baseline of 0.210). Figure 5 formalizes this relationship: DG_D is a near-perfect linear predictor of CIDEr for TPC-CMA (R2=0.986R^2=0.986), whereas Raw Gap yields a much weaker fit (R2=0.691R^2=0.691). This confirms that Distribution Gap, not Raw Gap, governs cross-modal interchangeability, validating the decomposition in Section 2.2. The baselines corroborate this finding. Mean-Centering’s DG_D remains identical to Original CLIP (0.685), confirming the Section 2.2 analysis; its marginal CIDEr gain reflects centroid shift alone, not structural improvement. CLIP-Refine’s slightly lower DG_D (0.647) correctly predicts its higher CIDEr despite a larger Raw Gap. M2-Mix and AlignCLIP fall below the trend line because degraded representation quality prevents DG_D reduction from translating into proportional CIDEr gains, indicating that DG_D governs interchangeability given adequate representation quality. We provide qualitative captioning examples in Appendix G to illustrate the improvement. For joint clustering (Table 3), Original CLIP achieves ARI = 0.318, reflecting limited semantic coherence in joint clustering. TPC-CMA progressively improves clustering quality as DG_D decreases, with ARI rising from 0.330 (αtarget=0.01 _target=0.01) to 0.551 (αtarget=0.7 _target=0.7), while ARI at αtarget=0.9 _target=0.9 shows a decline (0.492) despite further DG_D reduction. We analyze the spectral properties underlying this saturation in Section 5. Mean-Centering achieves only ARI = 0.306, barely above Original CLIP despite centroid alignment, confirming that structural reshaping, not just centroid correction, is needed for meaningful improvement in joint clustering quality across modalities. These task-specific optima underscore the necessity of controllable alignment strength: CIDEr favors high αtarget _target, ARI peaks at moderate αtarget _target, and discriminative accuracy favors low αtarget _target (visualized in Appendix H). No single fixed operating point can satisfy all tasks simultaneously, making TPC-CMA’s controllable αtarget _target a geometric necessity rather than a mere convenience. 4.4 Discriminative Tasks Table 4: COCO retrieval (5K test split). Method I2T R@1 I2T R@5 T2I R@1 T2I R@5 Original CLIP 53.42 77.18 35.36 61.52 Mean-Centering 43.74 70.70 34.44 60.03 TPC-CMA αtarget=0.01 _target=0.01 50.96 76.10 35.96 62.16 TPC-CMA αtarget=0.05 _target=0.05 48.56 74.86 35.68 61.87 Table 5: Multi-dataset zero-shot classification. Method ImageNet CIFAR-100 Food-101 Caltech Flowers Orig. CLIP 62.62 69.80 79.50 85.90 67.20 Mean-Centering 60.23 67.60 78.50 83.90 62.00 TPC-CMA αtarget=0.01 _target=0.01 60.86 67.80 77.00 87.10 64.00 TPC-CMA αtarget=0.05 _target=0.05 57.78 66.20 75.30 87.00 62.30 To show that mild alignment (αtarget≤0.05 _target≤ 0.05) can preserve, or even improve, discriminative quality, we compare TPC-CMA at low αtarget _target against Original CLIP and Mean-Centering on COCO retrieval (Table 4) and five zero-shot classification benchmarks (Table 5). We focus on Original CLIP and Mean-Centering as references, since Table 1 shows that the remaining baselines are unsuitable comparisons here: M2-Mix and AlignCLIP suffer severe accuracy degradation (ImageNet −-8.38% and 32.79%, respectively), while CLIP-Refine fails to reduce the gap at all (−-1.9% change). For retrieval (Table 4), at αtarget=0.01 _target=0.01, T2I R@1 slightly exceeds Original CLIP while I2T R@1 drops by only 2.46 points, far less than Mean-Centering’s 9.68-point collapse. For zero-shot classification (Table 5), coarse-grained benchmarks like Caltech-101 improve (85.90→87.10%85.90→ 87.10\%), benefiting from tighter image-to-class-name correspondence, while fine-grained benchmarks like ImageNet show only mild degradation (−-1.76% at αtarget=0.01 _target=0.01). These results establish that mild alignment occupies a “better-than-baseline” regime for concept-level tasks, refuting the assumption that gap reduction and discriminative quality are strictly antagonistic. 4.5 Ablation Studies Table 6: Component ablation. w/o Intra applies Negative Reweighting only. CMA-Only applies both CMA mechanisms but with constant α from epoch 1 (no three-phase curriculum). Variant NR Intra TPC α Modal. Gap ↓ ImageNet Top-1 ARI w/o NR ✓ ✓ 0.5 0.282 54.66 0.453 w/o Intra ✓ ✓ 0.5 0.239 53.13 0.348 CMA-Only ✓ ✓ 0.5 0.133 49.75 0.362 TPC-CMA ✓ ✓ ✓ 0.5 0.130 54.86 0.516 w/o NR ✓ ✓ 0.05 0.364 55.47 0.372 w/o Intra ✓ ✓ 0.05 0.379 55.19 0.342 CMA-Only ✓ ✓ 0.05 0.244 51.71 0.335 TPC-CMA ✓ ✓ ✓ 0.05 0.245 57.78 0.385 To isolate the contribution of each component in TPC-CMA, namely Negative Reweighting (NR), Intra-modal Geometry Matching (Intra), and the Three-Phase Curriculum (TPC), we conduct a component removal ablation (Table 6) at two representative α values. Geometry Matching is essential. Without Geometry Matching (w/o Intra, α=0.5α=0.5), Negative Reweighting alone reduces the gap but ARI reaches only 0.348. Adding Geometry Matching boosts ARI to 0.516 (1.5×1.5×), confirming that manifold reshaping, not just repulsion reduction, is essential for structural tasks. Curriculum is essential. CMA-Only applies both CMA mechanisms but at a constant α=0.5α=0.5 from epoch 1. Despite reaching a nearly identical gap (0.133 vs. 0.130), it suffers 1.4×1.4× ARI degradation and 5% ImageNet loss compared to full TPC-CMA. Abrupt alignment destroys the pretrained discriminative structure before the model can adapt, so even a well-closed gap lacks the feature quality for meaningful clustering. The three-phase curriculum prevents this by anchoring features first and increasing alignment only as fast as the observed gradient equilibrium permits. Additional ablations and generalization. We further validate curriculum hyperparameters (schedule type and anchor phase length) and confirm that TPC-CMA generalizes across model sizes with consistent gap reduction ratios and no architecture-specific tuning. These results are detailed in Appendix F. 5 Discussion The preceding experiments reveal task-specific optima at different αtarget _target values, suggesting an intrinsic cost to strong alignment. We examine the spectral properties of the aligned feature space to understand this cost. Figure 6: Effective rank dynamics. (a) Single-modal effective rank vs. αtarget _target; rank follows an inverted-U, peaking near αtarget=0.1 _target=0.1 before declining. Baselines are shown as horizontal dotted lines. (b) Modality Fusion Index vs. ARI. Effective rank and the cost of over-alignment. As shown in Figure 6(a), the effective rank Roy and Vetterli (2007) of each modality’s feature space follows an inverted-U: at moderate alignment, CMA’s uniformity component expands feature utilization beyond the original model, but at high αtarget _target alignment pressure compresses features into a lower-dimensional subspace. This rank collapse explains why accuracy drops and ARI saturates under strong alignment. Modality Fusion Index. We define FusionIdx = JointRank / mean(ImgRank, TxtRank) to quantify cross-modal complementarity, where a value near 2 indicates orthogonal subspaces and near 1 indicates full overlap. Figure 6(b) shows that FusionIdx peaks at αtarget=0.05 _target=0.05 and declines as stronger alignment forces both modalities into the same subspace. Notably, CLIP-Refine’s FusionIdx of 0.504 (below 1, indicating redundant collapse) explains its failure to improve ARI beyond the Original CLIP baseline (0.310 vs. 0.318) despite a low Raw Gap, reinforcing that spectral health, not gap magnitude, determines cross-modal utility. Implications for controllable alignment. These spectral dynamics bracket the useful operating range: too low αtarget _target leaves the Distribution Gap intact, while too high triggers rank collapse that erodes representation quality. Since tasks differ in sensitivity to these pressures, TPC-CMA’s adjustable αtarget _target is a geometric necessity for effective deployment across diverse tasks. 6 Related Work 6.1 Understanding the Modality Gap. Liang et al. Liang et al. (2022) first identified the modality gap in CLIP Radford et al. (2021), attributing the separation to a “cone effect” from random initialization. Subsequent work challenged this view: Fahim et al. Fahim et al. (2024) and Shi et al. Shi et al. (2023) show the gap is intrinsic to the contrastive objective, while Schrodi et al. Schrodi et al. (2025) link it to object bias and information imbalance. Wang and Isola Wang and Isola (2020) reveal a tension between alignment and uniformity, extended to multimodal settings by Yin et al. Yin et al. (2026, 2025). Further geometric analyses include similarity structure Yi et al. (2025), feature norm disparities Yaras et al. (2025), and global/residual decomposition Role et al. (2025) related to our framework. On the impact side, Grassucci et al. Grassucci et al. (2026) show the gap disrupts semantic consistency, Ramasinghe et al. Ramasinghe et al. (2024) argue a moderate gap can benefit retrieval, and Yu et al. Yu et al. (2026) leverage it for MLLM scaling. These findings suggest the gap is task-dependent rather than uniformly harmful. However, none of these works propose an explicit decomposition into centroid and distributional components with quantitative evidence that the Distribution Gap is the true predictor of cross-modal task quality, nor do they offer a controllable mechanism to navigate this trade-off. Our work addresses this gap by providing both a principled decomposition and an effective controllable alignment framework. 6.2 Mitigating the Modality Gap On the post-processing side, several methods attempt to close the gap without retraining. Mean-Centering Liang et al. (2022) subtracts modal centroids; GR-CLIP Li et al. (2025) and I0T An et al. (2025) apply learned affine transforms; Yamashita et al. Yamashita et al. (2025) calibrate the similarity distribution with pseudo-positive samples. These approaches are training-free but share a fundamental limitation: affine or translation-based operations can only move centroids closer, not reshape manifold geometry. On the training side, methods that modify model weights can in principle reshape feature distributions. Among our baselines, AlignCLIP Eslami and de Melo (2025) trains from scratch with an alignment objective but struggles to recover pre-trained discriminative quality; M2-Mix Oh et al. (2023) generates hard negatives via geodesic interpolation but does not directly target the gap; CLIP-Refine Yamaguchi et al. (2025) proposes random-reference feature alignment whose stochastic directions lack a consistent alignment signal. Other recent efforts tackle the gap from diverse angles, including modality inversion Mistretta et al. (2025), diffusion-based bridging Lee et al. (2025), distributional divergence minimization Yin et al. (2025), softer contrastive objectives Jing et al. (2024), and curriculum-based strategies Sofer et al. (2025). Despite these advances, existing methods share two limitations: they do not distinguish centroid offset from distributional mismatch, and they produce a single fixed operating point rather than a controllable trade-off. TPC-CMA addresses both with a dual-mechanism loss, a gradient-aware curriculum schedule inspired by multi-task gradient balancing Chen et al. (2018), and a single user-controlled parameter αtarget _target that governs the alignment strength. 7 Conclusion & Future Work By decomposing the modality gap into Centroid Gap and Distribution Gap, we show that Distribution Gap predicts cross-modal task quality far better than Raw Gap. Guided by this insight, TPC-CMA replaces cross-modal repulsion with intra-modal geometry matching and a three-phase curriculum for stable optimization. A single parameter αtarget _target yields a controllable Pareto frontier: mild alignment preserves discriminative quality, while stronger alignment unlocks captioning and clustering. Spectral analysis reveals an intrinsic alignment-expressiveness tension: aggressive alignment triggers rank collapse, confirming that no single fixed operating point is universally optimal across tasks. Future directions include automating αtarget _target selection via validation-driven scheduling, designing rank-preserving constraints to push the Pareto frontier further, and extending to other modality pairs. Since the decomposition and curriculum are modality-agnostic, we expect TPC-CMA to readily transfer to audio-text and video-text settings with minimal architectural modification. References N. M. An, E. Kim, J. Thorne, and H. Shim (2025) I0T: embedding standardization method towards zero modality gap. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 27182–27199. Cited by: Figure 1, §1, §6.2. A. Bardes, J. Ponce, and Y. LeCun (2022) VICReg: variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations, Cited by: §3.2.2. Z. Chen, V. Badrinarayanan, C. Lee, and A. Rabinovich (2018) GradNorm: gradient normalization for adaptive loss balancing in deep multitask networks. In International Conference on Machine Learning, p. 794–803. Cited by: item 3, §1, §3.3, §6.2. Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. (2024) InternVL: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 24185–24198. Cited by: §1. M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev (2023) Reproducible scaling laws for contrastive language-image learning. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 2818–2829. Cited by: §1. J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei (2009) ImageNet: a large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 248–255. Cited by: §4.1. A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In International Conference on Learning Representations, Cited by: §4.1. S. Eslami and G. de Melo (2025) Mitigate the gap: improving cross-modal alignment in CLIP. In International Conference on Learning Representations, Cited by: §4.1, Table 1, §6.2. A. Fahim, A. Murphy, and A. Fyshe (2024) It’s not a modality gap: characterizing and addressing the contrastive gap. arXiv preprint arXiv:2405.18570. Cited by: §1, §6.1. P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao (2024) CLIP-adapter: better vision-language models with feature adapters. International Journal of Computer Vision 132, p. 581–595. Cited by: §1. E. Grassucci, G. Cicchetti, E. Frasca, A. Uncini, and D. Comminiello (2026) Closing the modality gap aligns group-wise semantics. In International Conference on Learning Representations, Cited by: §1, §4.1, Table 3, §6.1. J. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doerhan, B. Avila Pires, Z. Chen, and M. Valko (2020) Bootstrap your own latent: a new approach to self-supervised learning. In Advances in Neural Information Processing Systems, Vol. 33, p. 21271–21284. Cited by: §3.2.2. L. Huang, X. Cao, H. Lu, Y. Meng, F. Yang, and X. Liu (2025) Mind the gap: preserving and compensating for the modality gap in clip-based continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 3777–3786. Cited by: §1. C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. Le, Y. Sung, Z. Li, and T. Duerig (2021) Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, p. 4904–4916. Cited by: §1. D. Jing, X. He, Y. Luo, N. Fei, G. Yang, W. Wei, H. Zhao, and Z. Lu (2024) FineCLIP: self-distilled region-based clip for better fine-grained understanding. Vol. 37, p. 27896–27918. Cited by: §6.2. J. R. Lee, Y. Shin, G. Son, and D. Hwang (2025) Diffusion bridge: leveraging diffusion model to reduce the modality gap between text and vision for zero-shot image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 4050–4059. Cited by: §6.2. B. Li, Y. Zhang, X. Wang, W. Liang, L. Schmidt, and S. Yeung-Levy (2025) Closing the modality gap for mixed modality search. arXiv preprint arXiv:2507.19054. Cited by: Figure 1, §1, §6.2. J. Li, D. Li, S. Savarese, and S. Hoi (2023a) BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning, p. 19730–19742. Cited by: §1. W. Li, L. Zhu, L. Wen, and Y. Yang (2023b) DeCap: decoding clip latents for zero-shot captioning via text-only training. In International Conference on Learning Representations, Cited by: §1, §4.1. W. Liang, Y. Zhang, Y. Kwon, S. Ye, and J. Zou (2022) Mind the gap: understanding the modality gap in multi-modal contrastive representation learning. In Advances in Neural Information Processing Systems, Vol. 35, p. 17612–17625. Cited by: Figure 1, §1, §1, §2.1, §4.1, Table 1, §6.1, §6.2. T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) Microsoft coco: common objects in context. In European Conference on Computer Vision, p. 740–755. Cited by: §4.1. H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. In Advances in Neural Information Processing Systems, Vol. 36. Cited by: §1. M. Mistretta, A. Baldrati, L. Agnolucci, M. Bertini, and A. D. Bagdanov (2025) Cross the gap: exposing the intra-modal misalignment in CLIP via modality inversion. In International Conference on Learning Representations, Cited by: §6.2. R. Mokady, A. Hertz, and A. H. Bermano (2021) ClipCap: clip prefix for image captioning. arXiv preprint arXiv:2111.09734. Cited by: §1. N. Mu, A. Kirillov, D. Wagner, and S. Xie (2022) SLIP: self-supervision meets language-image pre-training. In European Conference on Computer Vision, p. 529–544. Cited by: §1. C. Oh, J. So, H. Byun, Y. Lim, M. Shin, J. Jeon, and K. Song (2023) Geodesic multi-modal mixup for robust fine-tuning. In Advances in Neural Information Processing Systems, Vol. 36. Cited by: §4.1, §4.2, Table 1, §6.2. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §1, §2.1, Table 1, §6.1. A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever (2019) Language models are unsupervised multitask learners. OpenAI blog. Cited by: §4.1. S. Ramasinghe, V. Shevchenko, G. Avraham, and A. Thalaiyasingam (2024) Accept the modality gap: an exploration in the hyperbolic space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: §1, §6.1. F. Role, S. Meyer, and V. Amblard (2025) Fill the gap: quantifying and reducing the modality gap in image-text representation learning. arXiv preprint arXiv:2505.03703. Cited by: §2.2, §6.1. A. Rosenberg and J. Hirschberg (2007) V-measure: a conditional entropy-based external cluster evaluation measure. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), J. Eisner (Ed.), Prague, Czech Republic, p. 410–420. External Links: Link Cited by: §4.1. O. Roy and M. Vetterli (2007) The effective rank: a measure of effective dimensionality. In European Signal Processing Conference (EUSIPCO), p. 606–610. Cited by: §5. S. Schrodi, D. T. Hoffmann, M. Argus, V. Fischer, and T. Brox (2025) Two effects, one trigger: on the modality gap, object bias, and information imbalance in contrastive vision-language models. In International Conference on Learning Representations, Cited by: §1, §6.1. P. Sharma, N. Ding, S. Goodman, and R. Soricut (2018) Conceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, p. 2556–2565. Cited by: §4.1. P. Shi, M. C. Welle, M. Björkman, and D. Kragic (2023) Towards understanding the modality gap in clip. In ICLR 2023 Workshop on Multimodal Representation Learning, Cited by: §2.1, §6.1. A. Sofer, Y. Goldman, and S. E. Chazan (2025) Pull it together: reducing the modality gap in contrastive learning. In Proc. Interspeech, p. 196–200. External Links: Document Cited by: §6.2. Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao (2024) EVA-clip: improved training techniques for clip at scale. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Cited by: §1. Y. Tewel, Y. Shalev, I. Schwartz, and L. Wolf (2022) ZeroCap: zero-shot image-to-text generation for visual-semantic arithmetic. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 17918–17928. Cited by: §1. A. van den Oord, Y. Li, and O. Vinyals (2018) Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748. Cited by: §2.1. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30. Cited by: §4.1. T. Wang and P. Isola (2020) Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning, p. 9929–9939. Cited by: §6.1. H. Xu, S. Xie, X. E. Tan, P. Huang, R. Howes, V. Sharma, S. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer (2024) Demystifying clip data. In International Conference on Learning Representations, Cited by: §1. S. Yamaguchi, D. Feng, S. Kanai, K. Adachi, and D. Chijiwa (2025) Post-pre-training for modality alignment in vision-language foundation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 4256–4266. Cited by: §4.1, §4.2, Table 1, §6.2. S. Yamashita, D. Shirafuji, and T. Saito (2025) Bridging the modality gap by similarity standardization with pseudo-positive samples. In Proceedings of the 39th Pacific Asia Conference on Language, Information and Computation, p. 126–139. Cited by: Figure 1, §1, §6.2. C. Yaras, S. Chen, P. Wang, and Q. Qu (2025) Explaining and mitigating the modality gap in contrastive multimodal learning. In Conference on Parsimony and Learning, Proceedings of Machine Learning Research, Vol. 280, p. 1365–1387. Cited by: §6.1. L. Yi, R. Douady, and C. Chen (2025) Decipher the modality gap in multimodal contrastive learning: from convergent representations to pairwise alignment. arXiv preprint arXiv:2510.03268. Cited by: §6.1. W. Yin, Z. Xiao, P. Zhou, S. Yu, J. Shen, J. Sonke, and E. Gavves (2025) Distributional vision-language alignment by cauchy-schwarz divergence. arXiv preprint arXiv:2502.17028. Cited by: §4.1, §4.2, Table 1, §6.1, §6.2. W. Yin, P. Zhou, Z. Xiao, J. Liu, S. Yu, J. Sonke, and E. Gavves (2026) Towards uniformity and alignment for multimodal representation learning. arXiv preprint arXiv:2602.09507. Cited by: §6.1. J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu (2022) CoCa: contrastive captioners are image-text foundation models. Transactions on Machine Learning Research. Cited by: §1. X. Yu, Y. Xin, W. Zhang, C. Liu, H. Zhao, X. Hu, X. Yu, Z. Qiao, H. Tang, X. Yang, X. Hu, C. Qin, H. Xiong, Y. Qiao, and S. Yan (2026) Modality gap-driven subspace alignment training paradigm for multimodal large language models. arXiv preprint arXiv:2602.07026. Cited by: §6.1. J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny (2021) Barlow twins: self-supervised learning via redundancy reduction. In International Conference on Machine Learning, p. 12310–12320. Cited by: §3.2.2. Z. Zeng, Y. Zhang, H. Wang, G. Lu, and W. Lu (2024) MeaCap: memory-augmented zero-shot image captioning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14100–14110. Cited by: §1. X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer (2023) Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, p. 11975–11986. Cited by: §1. X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer (2022) LiT: zero-shot transfer with locked-image text tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 18123–18133. Cited by: §1. K. Zhou, J. Yang, C. C. Loy, and Z. Liu (2022) Learning to prompt for vision-language models. International Journal of Computer Vision 130, p. 2337–2348. Cited by: §1.