Paper deep dive
Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces
Pratham Yashwante, Rose Yu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 8:02:56 PM
Summary
This paper investigates the alignment of time series, vision, and language representations under the Platonic Representation Hypothesis. Using a trimodal contrastive learning framework, the authors find that independently pretrained encoders exhibit near-orthogonal geometry. Post-hoc alignment via projection heads reveals asymmetric convergence: time series align more strongly with visual representations than with text. Alignment improves with model scale but saturates with information density, suggesting that richer text helps only up to a threshold. Images act as effective intermediaries between time series and language.
Entities (9)
Relation Signals (6)
Platonic Representation Hypothesis â posits â Convergence to Shared Latent Structure
confidence 98% · The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world.
Time Series â alignsmorestronglywith â vision
confidence 95% · time series align more strongly with visual representations than with text
Contrastive Learning â usedfor â Post-hoc Alignment
confidence 95% · apply post-hoc alignment by training projection heads over frozen encoders using contrastive learning
Model Size â improves â Alignment Quality
confidence 92% · overall alignment in contrastive representation spaces improves with model size
vision â actsasintermediaryfor â Time Series
confidence 90% · images can act as effective intermediaries between time series and language
Information Density â improvesalignmentuptothreshold â language
confidence 88% · richer textual descriptions improve alignment only up to a threshold; training on denser captions does not lead to further improvement
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However, this hypothesis has largely been examined in vision and language, and it remains unclear whether time series participate in such convergence. We first examine this in a trimodal setting and find that independently pretrained time series, vision, and language encoders exhibit near-orthogonal geometry in the absence of explicit coupling. We then apply post-hoc alignment by training projection heads over frozen encoders using contrastive learning, and analyze the resulting representations with respect to geometry, scaling behavior, and dependence on information density and input modality characteristics. Our investigation reveals that overall alignment in contrastive representation spaces improves with model size, but this alignment is asymmetric: time series align more strongly with visual representations than with text, and images can act as effective intermediaries between time series and language. We further see that richer textual descriptions improve alignment only up to a threshold; training on denser captions does not lead to further improvement. Analogous effects are observed for visual representations. Our findings shed light on considerations for building multimodal systems involving non-conventional data modalities beyond vision and language.
Tags
Links
- Source: https://arxiv.org/abs/2602.19367v1
- Canonical: https://arxiv.org/abs/2602.19367v1
Trouble viewing inline? Open PDF directly â
Full Text
125,129 characters extracted from source content.
Expand or collapse full text
Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces Pratham Yashwante 1 Rose Yu 1 Abstract The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However, this hypothesis has largely been examined in vision and language, and it remains unclear whether time series partici- pate in such convergence. We first examine this in a trimodal setting and find that independently pretrained time series, vision, and language en- coders exhibit near-orthogonal geometry in the absence of explicit coupling. We then apply post- hoc alignment by training projection heads over frozen encoders using contrastive learning, and analyze the resulting representations with respect to geometry, scaling behavior, and dependence on information density and input modality char- acteristics. Our investigation reveals that over- all alignment in contrastive representation spaces improves with model size, but this alignment is asymmetric: time series align more strongly with visual representations than with text, and images can act as effective intermediaries between time series and language. We further see that richer textual descriptions improve alignment only up to a threshold; training on denser captions does not lead to further improvement. Analogous ef- fects are observed for visual representations. Our findings shed light on considerations for building multimodal systems involving non-conventional data modalities beyond vision and language. 1. Introduction Neural network representations are increasingly converging. Across architectures and training objectives, deep models tend to organize inputs according to similar notions of simi- larity, an observation formalized as the Platonic Representa- 1 Department of Computer Science and Engineering, Univer- sity of California San Diego, USA. Correspondence to: Pratham Yashwante <pyashwante@ucsd.edu>. Preprint. February 24, 2026. repeated riseâ plateauâdecay patterns latent temporal process [12, 22, 38, 55, 62, ..., 58, 41, 30] VisionTime Series Language Figure 1. Trimodal projections of a shared temporal process. A latent processZgives rise to a numeric time series, a visual line plot, and a textual description, each representing the same signal in values, geometry, and language. Modality-specific encodersf ts , f img , and f txt map inputs into representation spaces. tion Hypothesis (PRH) (Huh et al., 2024). The hypothesis posits that representation learning models converge toward a shared statistical model of reality, shaped by the structure of the underlying data. Empirical evidence for this con- vergence is strongest for vision and language: as models scale, their similarity kernels become more aligned. Joint visionâlanguage (VL) models such as CLIP (Radford et al., 2021) further demonstrate that contrastive learning (CL) can produce coherent shared representations across perceptual and linguistic domains. Time series pose a distinct challenge for multimodal align- ment because their semantic structure is not directly observ- able from raw values. Images encode structure explicitly through spatial geometry, and text encodes semantics ex- plicitly through symbolic tokens. In contrast, numeric time series express meaning only implicitly through temporal variation: properties such as trends, periodicity, or anoma- lies are not discrete tokens or visual features, but latent properties that must be recovered from the signal. For ex- ample, while an image of a cat and the word âcatâ explicitly denote the same concept, a temporal trend is neither an ob- ject nor a symbol, but a property that emerges only through computation over a numeric series. This mismatch raises a central question: can time series achieve the same degree of representational alignment with vision and language? 1 arXiv:2602.19367v1 [cs.AI] 22 Feb 2026 Time SeriesâVisionâLanguage Alignment Time Series Image 87.8° Time Series Text 89.5° Image Text 89.3° Figure 2. Mean angular deviation between pretrained cross-modal representations on CaTS shows little inherent alignment. Figure 1 illustrates the trimodal formulation studied in this work. Assuming that there exists a shared reality of a latent temporal process, we aim to understand the convergence behavior of learned representations from vision, language and time series. As a motivating observation, we first examine the repre- sentational geometry in the absence of external coupling on the CaTS dataset (Zhou et al., 2026), which consists of triplets of numeric time series, corresponding visual plots, and textual captions. Independently pretrained cross-modal encoders exhibit near-orthogonal structure across modalities (Figure 2). When external coupling via CL is applied, align- ment behavior differs substantially across datasets: while imageâtext structure becomes coherent on Flickr (Plum- mer et al., 2016), the same objective gives more fragmented overlap on CaTS as shown in Figure 3, highlighting the addi- tional challenges posed by time series. Additional analyses of uncoupled representational geometry across multimodal time series datasets are provided in Appendix F. Figure 3. UMAP visualizations of representations after contrastive training on Flickr (top) and CaTS (bottom). These observations suggest that alignment may not emerge uniformly across modalities and datasets, even under iden- tical contrastive objectives, thereby raising the question of how time series alignment differs from standard VL settings. In this work, we conduct a systematic empirical study of trimodal alignment across time series, visual plots, and language, using CL as a controlled mechanism to probe representational compatibility across modality pairs. Our setup pairs pretrained encoders from each modality and trains projection heads to map representations into a shared space, following the CLIP paradigm (Radford et al., 2021). Our experiments span four datasets varying in domain and annotation richness, 34 encoder combinations across mul- tiple model families and scales, and controlled ablations isolating the effects of model capacity, optimization, and input characteristics. Our contributions and empirical findings highlight several aspects of trimodal alignment involving time series: 1.Trimodal Analysis: We present the first systematic em- pirical study of trimodal alignment involving time series, images, and language, and discuss the roles of infor- mation density and semantic explicitness in governing cross-modal alignment. 2.Asymmetric Convergence: Trimodal alignment is asymmetric, with time series aligning more strongly with visual plots than with language. This gap persists even with detailed textual descriptions, providing evidence that multimodal convergence is uneven. 3.Information Density Saturation: Increasing informa- tion density improves alignment only up to a threshold, indicating that denser text by itself is insufficient to in- duce further cross-modal convergence. 4.Grounding and Explicitness: Alignment is shaped not only by model scale, but also by how explicitly semantics are grounded across modalities and how well represen- tational formats are matched; modality pairs with more explicitly observable correspondences consistently align more strongly than those with implicit semantic links. 2. Related Work Contrastive learning. CLIP (Radford et al., 2021) demon- strated that CL on large-scale imageâtext pairs produces representations with strong zero-shot transfer capabilities. The approach trains separate encoders for each modality and aligns their outputs using the InfoNCE loss (van den Oord et al., 2019), which maximizes agreement between matched pairs while minimizing agreement with mismatched pairs within a batch. More work has extended this paradigm to additional modalities including audio (Guzhov et al., 2021), video (Xu et al., 2021), and 3D clouds (Zhang et al., 2021). Time series models. Recent work has introduced pretrained time series models for forecasting or masked prediction (Woo et al., 2024; Goswami et al., 2024; Ansari et al., 2024; Das et al., 2024). While these models demonstrate strong downstream performance, their representational alignment with other modalities remains underexplored. Prior work has also explored pairing time series with other modali- ties for specific applications. In healthcare, clinical notes 2 Time SeriesâVisionâLanguage Alignment LINEAR LAYER NORM GELU DROPOUT LINEAR Contrastive Training repeated riseâplateauâ decay patterns [12, 22, 38, 55, 62, ..., 58, 41, 30] Vision Encoder Time Series Encoder Text Encoder Projection Head Unimodal Spaces Multimodal Space Trainable Frozen Figure 4. Trimodal contrastive alignment framework. Frozen pretrained encoders independently map time series, visual plots, and text into their respective unimodal representation spaces. Trainable projection heads transform these representations into a shared embedding space. Alignment is learned via a symmetric contrastive objective applied jointly across all modality pairs (TSâIMG, TSâTXT, IMGâTXT). have been aligned with physiological time series for patient outcome prediction (Baldenweg et al., 2024; Hayat et al., 2023). Contrastive learning has also been applied to time seriesâtext settings, including MedCLIP (Wang et al., 2022) and the recent TimesCLIP (Chen et al., 2025). Our work instead treats representational alignment as the main focus, studying when and under what conditions it emerges, with potential implications for downstream use. Semantic content and density in language. The uniform information density hypothesis (Levy & Jaeger, 2007) pro- poses that speakers distribute information evenly across ut- terances to optimize communication efficiency. Subsequent work has operationalized this through surprisal-based met- rics measuring how predictably information is distributed across tokens (Meister et al., 2021). We adapt these ideas to study how the amount and distribution of semantic content in text affects alignment quality. Further discussion on alignment is provided in Appendix B, along with related work on ECG representations in Ap- pendix N. 3. Trimodal Alignment Framework Design Rationale. As shown in Figure 4, our design adopts the core contrastive alignment mechanism similar to that used in CLIP, repurposing it as a controlled framework for analyzing trimodal representational compatibility. We evaluate 34 trimodal configurations spanning 26 unique pretrained encoders, including 9 text, 9 vision, and 8 time series models across different scales. All configurations use identical hyperparameters for fair comparison. Encoder combinations appear in Appendix D, with experimental details and ablations in Appendix G. Each modality encoder is used as a frozen feature extractor, and its output is mapped into a shared embedding space via projection heads. All modalities share the same projection (Appendix G.6). Symmetric Contrastive Loss. We align representations using a symmetric contrastive objective applied jointly across all modality pairs: time seriesâimage (TSâIMG), time seriesâtext (TSâTXT), and imageâtext (IMGâTXT). Givenâ 2 -normalized embeddingsz ts ,z img ,z txt âR d , we apply a bidirectional InfoNCE loss to each modality pair. For a modality pair (x,y), the forward loss is defined as L xây =â 1 N N X i=1 log exp z (i)†x z (i) y /Ï P N j=1 exp z (i)†x z (j) y /Ï ,(1) whereÏis a temperature parameter. For example, the TSâIMG loss is defined as the average of the forward and reverse directions, L ts-img = 1 2 (L tsâimg +L imgâts ).(2) Analogous losses are defined for TSâTXT and IMGâTXT (see Appendix G.3). The final training objective is the equally weighted sum of all three modality-pair losses: L total =L ts-img +L ts-txt +L img-txt .(3) We also evaluate ablated variants that restrict the set of modality pairs in the loss, including bimodal TSâIMG and TSâTXT settings, as well as joint VL with time series (VLâTS; full details in Appendix I). 3 Time SeriesâVisionâLanguage Alignment Evaluation Metrics. We evaluate alignment using met- rics that capture global similarity, retrieval, and geometric consistency of cross-modal representations. All metrics are computed on held-out test sets usingâ 2 -normalized embed- dings. Implementation details and definitions are provided in Appendix E. For each modality pair, we calculate: 1.Cosine similarity margin defined as the difference be- tween matched and mismatched pairs. 2.Bidirectional cross-modal retrieval using Recall@k (R@1, R@5, R@10). 3.Procrustes disparity measures how well one modal- ity can be aligned to another via an optimal orthogonal transformation. 4. Centered Kernel Alignment (CKA) measures non- linear representational similarity using RBF kernels. 5.MutualkNN overlap measures agreement in local neigh- borhood structure. We quantify the semantic richness of textual descriptions using information density (ID), defined as the total surprisal of a text under a pretrained language model. Higher ID corresponds to captions that are longer and semantically more informative and unique. Dataset-level ID is computed as the mean across captions; see Appendix H for full details. Datasets. We evaluate trimodal alignment using datasets that allow us to probe semantic explicitness, information density, and visual mediation. Because fully aligned tri- modal time series datasets are rare, we combine human and synthetic data to construct modality-complete variants wher- ever necessary. Full details are provided in Appendix C. We use CaTS-Bench (Zhou et al., 2026) as our primary dataset, as it natively provides aligned triplets of time series, visual plots, and captions. The captions serve as direct tex- tual representations, describing observable temporal struc- ture such as trends, phases, events, and summary statistics. To study ID, we construct multiple caption variants derived from the same underlying time series, including progres- sively condensed test-set captions and an expanded high-ID variant. Condensed captions reduce semantic explicitness by limiting descriptions to a fixed number of salient phrases, while the high-ID variant exceeds the original captions by more than a factor of two in measured ID. We also evaluate on TRUCE (Jhamtani & Berg-Kirkpatrick, 2021), which contains short time series paired with concise, direct textual descriptions. Its short length enables explicit plot annotation of the time series. Owing to its small size, it is used only for evaluation, with three visual variants per signal (generic, styled, and annotated) to assess visual robustness. To study alignment under indirect textual supervision, we use MIMIC (Gow et al., 2023) and PTB-XL (Wagner et al., 2020), where text consists of diagnostic reports that do not explicitly describe waveform structure. MIMIC reports are in English, while PTB-XL reports are in German, allow- ing us to examine multilingual effects. For both datasets, we generate visual plots of ECG waveforms and construct train/test splits matched in size to CaTS. These datasets use longer series (5000 timesteps) than CaTS, enabling analysis of how temporal resolution affects multimodal alignment. 4. Results RQ1. Do contrastive representation spaces converge uniformly as models scale? Increasing model capacity is often associated with stronger representations (Jia et al., 2021; Radford et al., 2021; Hoff- mann et al., 2022), but whether scale alone leads to conver- gence across time series, vision, and language remains an open question. We therefore use our trimodal framework to analyze alignment on CaTS across encoder scales, while holding the contrastive objective and parameters fixed. Figure 5 reports alignment quality across modality pairs as a function of total model size. Overall alignment improves with scale, but gains are highly uneven across modality pairs. TSâTXT alignment remains the weakest in absolute terms across all model sizes, reflecting the difficulty of directly aligning numeric signals with language, yet it also exhibits the strongest positive relationship with scale. In contrast, TSâIMG alignment achieves substantially higher absolute performance even at smaller scales, but shows weaker scal- ing trends, suggesting earlier saturation driven by stronger shared visual grounding with temporal data. Notably, global alignment across modalities is consistently strong, as re- flected by cosine similarity and geometric metrics, while local neighborhood structure remains weak. Across all con- figurations, mutualkNN overlap stays low, indicating lim- ited fine-grained correspondence. This highlights a clear dissociation between global and local alignment under CL: strong global geometric similarity does not imply robust neighborhood-level semantic correspondence. Further, we observe a consistent asymmetry, with time se- ries aligning more strongly with visual plots than with lan- guage. This pattern reflects mismatched inductive biases and differences in semantic explicitness across modalities: visual plots externalize latent temporal structure into geo- metric form, whereas time series encode it implicitly and text abstracts over it symbolically. As a result, alignment is strongest between modalities that expose structure in com- parable representational formats, while implicit-to-abstract alignment remains challenging even at larger scales. Further detailed discussion and ablations are in Appendix M. We also find that trimodal alignment is sensitive to both optimization quality and encoder capacity. Larger batch 4 Time SeriesâVisionâLanguage Alignment 02000040000 Model Size 0.55 0.60 0.65 0.70 0.75 0.80 Metric Value Cosine Margin TS-IMG (r=0.28) TS-TXT (r=0.67) IMG-TXT (r=0.53) 02000040000 Model Size 0.3 0.4 0.5 0.6 0.7 0.8 Metric Value Procrustes TS-IMG (r=-0.26) TS-TXT (r=-0.69) IMG-TXT (r=-0.51) 02000040000 Model Size 0.55 0.60 0.65 0.70 0.75 0.80 0.85 Metric Value CKA TS-IMG (r=0.09) TS-TXT (r=0.69) IMG-TXT (r=0.41) 02000040000 Model Size 0.050 0.075 0.100 0.125 0.150 0.175 0.200 0.225 Metric Value Mutual kNN TS-IMG (r=0.22) TS-TXT (r=0.71) IMG-TXT (r=0.47) Figure 5. Scaling behavior of alignment across 34 trimodal configurations on CaTS as a function of total model size (in millions of parameters). Each subplot reports alignment quality for three modality pairs (TSâIMG, TSâTXT, IMGâTXT). Each point corresponds to a distinct encoder configuration. Dashed lines indicate linear trends with Pearson correlation coefficients reported in the legend. sizes and stronger projection heads consistently improve alignment which confirms the importance of effective con- trastive optimization. Also, scaling the time series encoder provides substantial gains for both TSâIMG and TSâTXT alignment. This identifies temporal representations as a key factor in trimodal alignment, with improvements to the time series encoder being particularly important for strengthen- ing otherwise weak modality pairs such as TSâTXT. Addi- tional analyses and ablations are provided in Appendix G.5 (optimization), K (time series encoder capacity), and G.6 (projection head design). RQ2. Do pretrained VL models change alignment? 123456789 Scale (Billions of Params) 0.4 0.5 0.6 0.7 IMG TXT Procrustes A B C D E F G H I J K L M N O P Q R TS-IMG-TXT (AI) VLTS (JR) TS-IMG-TXT (AI) VLTS (JR) J: CLIP-Chronos K: CLIP-MOMENT L: CLIP-TimesFM M: SigLIP-MOMENT N: SigLIP-TimesFM O: SigLIP-Chronos P: BLIP-2-TimesFM Q: BLIP-2-MOMENT R: BLIP-2-Chronos Figure 7. IMGâTXT alignment across model scales for VLâTS on CaTS. AâI correspond to select configurations (Appendix D.2). Jointly pretrained VL models exhibit strong intrinsic align- ment between images and text (Radford et al., 2021; Li et al., 2022; Zhai et al., 2023), raising the question of whether such coupling affects trimodal alignment. In particular, it is unclear whether strong IMGâTXT alignment in trimodal systems must be learned jointly, or can instead be inherited from existing VL pretraining. To isolate this effect, we con- duct this comparison on CaTS, contrasting the full trimodal setting with configurations where we pair a pretrained VL model with a time series encoder. Both settings are trained using the same contrastive objective, allowing differences in alignment to be attributed to representational priors. As shown in Figure 7, we see that VLâTS configurations achieve strong IMGâTXT alignment even at relatively small scales. This reflects the tightly coupled imageâtext geometry inherited from VL pretraining, which enables robust align- ment without relying on increased model capacity. In con- trast, trimodal encoders must jointly learn structure across all three modalities, and therefore depend more heavily on scale to compensate for weaker intrinsic coupling. RQ3. How does information density in text affect cross- modal alignment? Because textual descriptions vary widely in how explicitly they encode semantic structure about an underlying time series, we study how the information density of text affects the strength of cross-modal alignment. We vary textual semantic richness by constructing multiple caption vari- ants for each sample with different ID. Specifically, we use LLMs (GPT-4o-mini and LLaMA-3.2-90B) to compress the original captions to a fixed number of salient semantic phrases, while keeping the underlying data, model architec- tures, and contrastive setup fixed. We observe that alignment quality improves consistently with increasing ID across all metrics, as shown in Figure 6. Using the original CaTS cap- tions with three progressively condensed variants, we find that as captions become denser and more information-rich, Procrustes disparity decreases, indicating tighter shared ge- ometry, whilekNN overlap increases, reflecting stronger neighborhood-level semantic consistency. Low-information captions provide insufficient semantic sig- nal, leading to weakly structured or partially collapsed em- beddings and high variability across model configurations. At very low ID, embeddings cluster tightly with limited 5 Time SeriesâVisionâLanguage Alignment 100200300400 Information Density 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Mean Score Cosine Similarity 100200300400 Information Density 0.5 1.0 1.5 2.0 2.5 3.0 Procrustes Distance 100200300400 Information Density 0.025 0.050 0.075 0.100 0.125 0.150 Mutual kNN TS-TXTTXT-IMGID = 26.81ID = 69.48ID = 149.05ID = 416.81 Figure 6. Effect of text ID on alignment across increasing levels of text density. Markers indicate distinct caption regimes with increasing semantic richness. Scores are averaged across 9 representative model configurations (see Appendix D.2). neighborhood differentiation, indicating that sparse seman- tic content constrains the formation of meaningful relational structure. In contrast, higher-information captions encode richer attributes and contextual relationships, enabling more coherent cross-modal geometry. Larger models benefit most in this regime, but also degrade more sharply when captions are short and under-specified, reflecting a stronger depen- dence on explicit semantic content. These results show that the amount of explicitly encoded semantic information plays a key role in multimodal alignment when text is involved. Since alignment improves as captions become denser, we test whether further increasing caption ID during training provides additional gains. Specifically, we test whether training on captions with more than double the original CaTS ID improves alignment. We find that despite this substantial increase in semantic content, alignment metrics change only marginally across all modality pairs as shown in Table 1. This indicates that once text reach sufficient semantic richness, further increases in training-time ID do not meaningfully improve cross-modal alignment. Table 1. Effect of doubling text ID on alignment. We compare original captions (train ID = 417.2) and high-information captions (train ID = 870.2).âdenotes High IDâOG ID. Both models use same samples with their respective ID-specific caption sets. Scores are averaged across configurations listed in Appendix D.2. Cosine MarginâProcrustesâMutual kNNâ PairOGHighâOGHighâOGHighâ TSâIMG0.720.720.000.470.470.000.130.130.00 TSâTXT0.570.56-0.010.740.740.000.080.07-0.01 IMGâTXT0.640.640.000.600.600.000.110.10-0.01 A similar effect appears when evaluating CaTS-trained mod- els on TRUCE for TSâTXT alignment. TRUCE captions are short, direct textual descriptions with ID comparable to our low-ID (2 phrase) caption variants (IDâ 74.5vs.69.6) and similar length (12.5 vs. 10.2 words on average), and give comparable levels of TSâTXT alignment. This mir- rors the behavior observed for CaTS captions in the low-ID regime and provides additional evidence that alignment is strongly governed by the amount and structural organization of semantic content in the text modality. RQ4. What happens when text does not directly de- scribe the time series signal? In many real-world settings, text associated with time series provides high-level interpretations rather than direct descrip- tions of signal structure. We therefore ask whether align- ment persists when language is only indirectly related to the underlying signal. To answer this, we evaluate trimodal alignment using the same contrastive framework under in- direct textual supervision using ECG datasets. Specifically, we compare CaTS with MIMIC, where clinical reports only implicitly relate to waveform dynamics. We observe distinct alignment scaling behavior between CaTS and MIMIC across different trimodal configurations (Figure 8). Although increasing model capacity improves alignment across all modality pairs in both datasets, system- atic differences persist depending on the nature of textual supervision. Note that CaTS captions exhibit substantially higher textual density (test ID = 416.81), whereas MIMIC re- ports have much lower ID (test ID = 149.48), and this dispar- ity is directly reflected in cross-modal alignment strength. 051015 Model Size (in billion params) 0.4 0.5 0.6 0.7 0.8 Metric Value CaTS Procrustes TS-IMG (r=-0.83) TS-TXT (r=-0.72) IMG-TXT (r=-0.34) 051015 Model Size (in billion params) 0.4 0.6 0.8 1.0 Metric Value MIMIC Procrustes TS-IMG (r=-0.60) TS-TXT (r=-0.64) IMG-TXT (r=-0.62) Figure 8. Scaling behavior of alignment on CaTS and MIMIC across 5 representative trimodal configurations (see Appendix D.2). We report Procrustes disparity as a function of model size. Alignment involving text is consistently weaker on MIMIC, 6 Time SeriesâVisionâLanguage Alignment particularly for TSâTXT and IMGâTXT pairs. Even at com- parable model scales, MIMIC exhibits lower cosine similar- ity and higher Procrustes disparity, indicating poorer global and geometric alignment when language provides only in- direct semantic grounding. This gap is most pronounced for TSâTXT which shows that clinical language provides substantially less explicit signal for aligning numeric time series than direct descriptive captions. Scaling partially miti- gates this effect: larger models improve TSâTXT alignment on both datasets, but gains are substantially weaker under indirect supervision, with alignment remaining consistently worse on MIMIC across scales, indicating a limit imposed by semantic explicitness. In contrast, TSâIMG alignment is often stronger on MIMIC despite weaker text alignment. This reflects the longer and more structured ECG signals in MIMIC, which provide richer numeric structure for alignment with visual plots. We also see that stronger TSâIMG alignment also yields large gains in cross-modal retrieval (Table 2), confirming that improved geometric alignment translates into retrieval performance even when textual grounding is weak. Table 2. Cross-modal retrieval performance on MIMIC and CaTS for TSâIMG retrieval. Results are macro-averaged over TSâIMG and IMGâTS and reported as percentages. Model acronyms: Dv2/Dv3 = DINOv2/DINOv3, SL = SigLIP2, Q = Qwen, Mo = MOMENT, C = Chronos; B/L denote Base/Large. Model Configuration Retrieval Performance (%) R@1R@5R@10 CaTSMIMICCaTSMIMICCaTSMIMIC Dv2-B + Q-0.6B + Mo-B1.6131.316.1361.8410.7573.62 Dv2-L + Q-4B + Mo-L2.2521.368.6549.4814.7162.21 Dv3-7B + Q-8B + Mo-L6.1923.7018.3551.8027.1265.13 SL2-B + Q-0.6B + C-B3.9620.3312.6445.7919.8558.69 SL-L + Q-4B + C-L6.9529.7819.8458.4528.6569.90 RQ5. How sensitive is alignment to language shifts? Thus far, we have examined how alignment depends on ex- plicitness and density within a single language. In practice, however, textual supervision can also vary linguistically. To isolate the effect of language shift, we compare alignment between MIMIC and PTB-XL, which both pair ECG time series with clinical reports from the same domain but differ in report language (English vs. German). Figure 9 shows that alignment is consistently weaker on PTB-XL across all modality pairs. This degradation is pro- nounced for pairs involving text, indicating that linguistic shift further weakens semantic coupling even when the un- derlying time series distribution is comparable. All trends observed in earlier experiments persist: TSâTXT remains the most challenging pair, increased scale provides partial compensation, and TSâIMG alignment remains relatively robust for longer more structured time series. These results suggest that cross-modal alignment is sensitive not only to semantic explicitness and text ID, but also to how well the language of textual supervision aligns with the inductive biases of pretrained language encoders. ABCD 0.0 0.1 0.2 0.3 0.4 0.5 Cosine Margin TSTXT ABCD 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Cosine Margin IMGTXT ABCD 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Procrustes ABCD 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Procrustes PTB-XL MIMIC A: DINOv2-B + Qwen-0.6B + MOMENT-B B: DINOv2-G + Qwen + MOMENT C: SigLIP + Qwen + Chronos D: SigLIP2-B + Qwen-0.6B + Chronos-B Figure 9. Comparison of multimodal alignment between MIMIC (English reports) and PTB-XL (German reports). English reports consistently achieve higher alignment across all modality pairs. RQ6. Do richer visual inputs help alignment? While earlier results show that textual semantic explicit- ness and information density strongly affect alignment, it remains unclear whether similar effects arise in the visual modality. To test whether increasing visual semantic con- tent improves alignment, we compare TSâIMG alignment across plot variants from the TRUCE dataset with progres- sively richer visual annotations (Appendix C.2), holding the underlying time series and training setup fixed. Figure 11 shows that increasing visual semantic richness leads to consistent improvements in TSâIMG alignment across configurations. Annotated plots achieve higher align- ment than generic or stylistically varied variants, indicating that richer and denser visual inputs provide stronger seman- tic anchors for aligning time series with images. Consistent with earlier trends, increasing model capacity (compare D and E) further improves TSâIMG alignment which indicates that visual expressiveness and scale act as complementary factors. ABCDEF 0.2 0.3 0.4 0.5 Cosine Similarity Generic Styled Annotated A: ViT-B + Qwen + TimesFM B: DINOv2-B + E5 + TimesFM C: DINOv2-B + T5 + Chronos D: SigLIP2 + Qwen-L + Chronos-L E: SigLIP2-B + Qwen-B + Chronos-B F: SigLIP2-B + T5 + TimesFM Figure 11. TSâIMG alignment across TRUCE plot variants with in- creasing visual input richness. Annotated achieves highest scores. 7 Time SeriesâVisionâLanguage Alignment ABCDE 0.4 0.6 0.8 Margin TS-IMG: Cosine ABCDE 0.2 0.4 0.6 Distance TS-IMG: Procrustes ABCDE 0.4 0.6 0.8 CKA TS-IMG: CKA ABCDE 0.00 0.05 0.10 0.15 kNN Overlap TS-IMG: kNN FGHIJ 0.4 0.6 Margin TS-TXT: Cosine FGHIJ 0.4 0.6 0.8 Distance TS-TXT: Procrustes FGHIJ 0.3 0.4 0.5 0.6 CKA TS-TXT: CKA FGHIJ 0.00 0.05 0.10 kNN Overlap TS-TXT: kNN BimodalTrimodal A: ViT-B + Chronos-B + E5-B B: ViT-L + TimesFM-L + Qwen4B C: Dino-B + Moment-B + Qwen0.6B D: Dino-G + Chronos-L + T5-L E: ViT-L + Chronos-L + E5-Mistral F: E5-B + Chronos-B + ViT-B G: Qwen4B + TimesFM-L + ViT-L H: Qwen0.6B + Moment-B + Dino-B I: T5-3b + Chronos-L + Dino-G J: E5-7b + Chronos-L + ViT-L Figure 10. Role of trimodality as a semantic bridge. We compare bimodal and trimodal CL for TSâIMG (top row) and TSâTXT (bottom row). Introducing the image modality consistently improves TSâTXT alignment. RQ7. Does trimodality help weakly aligned pairs? Given the persistent asymmetry between TSâIMG and TSâ TXT alignment, we ask whether introducing a third modality can facilitate alignment for otherwise weak modality pairs. To answer this, we measure alignment changes when one modality (IMG or TXT) is removed, comparing bimodal and trimodal contrastive learning on CaTS (Figure 10). For the weakest modality pair, TSâTXT, introducing the image modality produces a consistent upward shift across alignment metrics which indicates that visual representa- tions provide intermediate semantic structure that neither time series nor text capture in isolation. This improvement reflects not only better global geometry but also more co- herent neighborhood structure which suggests that images help organize weakly coupled representations. For a strong pair (TSâIMG), introducing a third modality often degrades performance, suggesting that it adds optimization complex- ity without providing new semantic signal. This shows that trimodality is most effective when it introduces missing semantic structure rather than redundant supervision. 5. Discussion We first examine whether the datasets studied here exhibit any inherent cross-modal structure prior to alignment and find that independently pretrained time series, vision, and language encoders produce near-orthogonal representations. This establishes a clear baseline: multimodal convergence does not arise without explicit coupling. When contrastive alignment is introduced, convergence re- mains highly non-uniform and strongly modality-dependent. While contrastive learning readily aligns vision and lan- guage, time series participate only partially, with the weak- est alignment observed for TSâTXT. This asymmetry per- sists even at larger scales and is consistent with how seman- tics are expressed: time series encode structure implicitly, text abstracts symbolically, and visual plots externalize la- tent temporal structure. As a result, images can act as inter- mediaries, reducing abstraction gaps and helping stronger alignment between time series and text. We further show that semantic explicitness constrains how much shared structure can be recovered across modalities. Increasing information density improves alignment at low to moderate levels, but saturates beyond a latent threshold, indicating that alignment limits arise from representational mismatch rather than insufficient supervision or scale alone. Although scaling and improved optimization selectively help weaker modality pairs, architectural coupling and pre- training objectives also play a stronger role, as evidenced by the effectiveness of jointly pretrained VL models. All these findings refine how multimodal convergence should be understood for numeric modalities and suggest that representational format and explicitness are key factors shaping alignment behavior, beyond model size alone. 8 Time SeriesâVisionâLanguage Alignment Impact Statement This paper aims to advance the empirical understanding of multimodal representation learning involving time series, visual, and language modalities. The work focuses on an- alyzing alignment behavior under controlled experimental settings, with the goal of clarifying how different modalities interact in contrastive representation spaces. While the study is not intended as a deployment-facing system and does not introduce models trained for direct decision-making in real- world applications, the questions it addresses are relevant to domains such as scientific analysis and healthcare, where multimodal data are common. We do not foresee imme- diate negative societal impacts arising from this work. By improving understanding of the conditions under which mul- timodal representations align or fail to align, this work aims to inform more robust evaluation practices and support the responsible development of future multimodal systems. Acknowledgement This work was supported in part by NSF Grants #2205093, #2146343, #2134274, CDC-RFA-FT-23-0069, the U.S. Army Research Office under Army-ECASE award W911NF-07-R-0003-03, the U.S. Department Of Energy, Office of Science, IARPA HAYSTAC Program, and DARPA YFA. References Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B. Vatt: Transformers for multi- modal self-supervised learning from raw video, audio and text, 2021. URLhttps://arxiv.org/abs/2104. 11178. Andrew, G., Arora, R., Bilmes, J., and Livescu, K. Deep canonical correlation analysis. In Proceedings of the 30th International Conference on Machine Learning, vol- ume 28 of Proceedings of Machine Learning Research, p. 1247â1255, Atlanta, Georgia, USA, 17â19 Jun 2013. PMLR. Ansari, A. F., Stella, L., Turkmen, C., Zhang, X., Mercado, P., Shen, H., Shchur, O., Rangapuram, S. S., Arango, S. P., Kapoor, S., Zschiegner, J., Maddix, D. C., Wang, H., Mahoney, M. W., Torkkola, K., Wilson, A. G., Bohlke- Schneider, M., and Wang, Y. Chronos: Learning the language of time series, 2024. URLhttps://arxiv. org/abs/2403.07815. Baldenweg, F., Burger, M., R Ì atsch, G., and Kuznetsova, R. Multi-modal contrastive learning for online clinical time-series applications, 2024. URLhttps://arxiv. org/abs/2403.18316. Benton, A., Khayrallah, H., Gujral, B., Reisinger, D. A., Zhang, S., and Arora, R. Deep generalized canonical cor- relation analysis, 2017. URLhttps://arxiv.org/ abs/1702.02519. Buiting, S., Sengupta, S., Benzine, A., Khair, A. E., Khaouja, I., and Tamaazousti, Y. Drimm: Drilling multi- modal model for time-series and text. In Proceedings of the 42nd International Conference on Machine Learning, 2025.URLhttps://api.semanticscholar. org/CorpusID:282814523. Chen, Z., Zhang, X., and Zhu, M.TS-CLIP: Time series understanding by CLIP.In Christodoulopou- los, C., Chakraborty, T., Rose, C., and Peng, V. (eds.), Proceedings of the 2025 Conference on Em- pirical Methods in Natural Language Processing, p. 4646â4664, Suzhou, China, November 2025. Asso- ciation for Computational Linguistics.ISBN 979- 8-89176-332-6.doi: 10.18653/v1/2025.emnlp-main. 231. URLhttps://aclanthology.org/2025. emnlp-main.231/. Das, A., Kong, W., Sen, R., and Zhou, Y. A decoder-only foundation model for time-series forecasting, 2024. URL https://arxiv.org/abs/2310.10688. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. URLhttps: //arxiv.org/abs/2010.11929. Elizalde, B., Deshmukh, S., Ismail, M. A., and Wang, H. Clap: Learning audio concepts from natural language su- pervision, 2022. URLhttps://arxiv.org/abs/ 2206.04769. Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K. V., Joulin, A., and Misra, I. Imagebind: One embedding space to bind them all, 2023. URLhttps://arxiv. org/abs/2305.05665. Goswami, M., Szafer, K., Choudhry, A., Cai, Y., Li, S., and Dubrawski, A. Moment: A family of open time-series foundation models, 2024. URLhttps: //arxiv.org/abs/2402.03885. Gow, B., Pollard, T., Nathanson, L. A., Johnson, A., Moody, B., Fernandes, C., Greenbaum, N., Waks, J. W., Eslami, P., Carbonati, T., Chaudhari, A., Herbst, E., Moukheiber, D., Berkowitz, S., Mark, R., and Horng, S. MIMIC- IV-ECG: Diagnostic Electrocardiogram Matched Sub- set, 2023. URLhttps://doi.org/10.13026/ 4nqg-sb35. RRID:SCR 007345. 9 Time SeriesâVisionâLanguage Alignment Guzhov, A., Raue, F., Hees, J., and Dengel, A. Audioclip: Extending clip to image, text and audio, 2021. URL https://arxiv.org/abs/2106.13043. Hadgi, S., Moschella, L., Santilli, A., Gomez, D., Huang, Q., Rodol ` a, E., Melzi, S., and Ovsjanikov, M. Escaping platoâs cave: Towards the alignment of 3d and text la- tent spaces, 2025. URLhttps://arxiv.org/abs/ 2503.05283. Hayat, N., Geras, K. J., and Shamout, F. E. Medfuse: Multi- modal fusion with clinical time-series data and chest x- ray images, 2023. URLhttps://arxiv.org/abs/ 2207.07027. Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L. Training compute-optimal large language models, 2022. URLhttps://arxiv.org/ abs/2203.15556. Huh, M., Cheung, B., Wang, T., and Isola, P.The platonic representation hypothesis.arXiv preprint arXiv:2405.07987, 2024. Jhamtani, H. and Berg-Kirkpatrick, T. Truth-conditional captioning of time series data. In EMNLP, 2021. Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision, 2021. URLhttps://arxiv. org/abs/2102.05918. Klabunde, M., Wald, T., Schumacher, T., Maier-Hein, K., Strohmaier, M., and Lemmerich, F.Resi: A com- prehensive benchmark for representational similarity measures, 2025. URLhttps://arxiv.org/abs/ 2408.00531. Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Simi- larity of neural network representations revisited, 2019. URL https://arxiv.org/abs/1905.00414. Kriegeskorte, N., Mur, M., and Bandettini, P. Represen- tational similarity analysis - connecting the branches of systems neuroscience. Front. Syst. Neurosci., 2:4, Novem- ber 2008. Levi, M. Y. and Gilboa, G. The double-ellipsoid geometry of clip, 2025. URLhttps://arxiv.org/abs/2411. 14517. Levy, R. and Jaeger, T. F.Speakers Optimize Infor- mation Density through Syntactic Reduction, p. 849â 856. MIT Press, 2007. doi: 10.7551/mitpress/7503. 003.0111.URLhttps://doi.org/10.7551/ mitpress/7503.003.0111. Li, H., Liu, C., Ding, Z., Liu, Z., Shao, W., and Huang, Z. Fine-grained contrastive learning for ecg-report align- ment with waveform enhancement, 2025. URLhttps: //arxiv.org/abs/2505.11939. Li, J., Selvaraju, R. R., Gotmare, A. D., Joty, S., Xiong, C., and Hoi, S. Align before fuse: Vision and language representation learning with momentum distillation, 2021. URL https://arxiv.org/abs/2107.07651. Li, J., Li, D., Xiong, C., and Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022. URLhttps:// arxiv.org/abs/2201.12086. Liu, C., Wan, Z., Ouyang, C., Shah, A., Bai, W., and Ar- cucci, R. Zero-shot ecg classification with multimodal learning and test-time clinical knowledge enhancement. In Proceedings of the 41st International Conference on Machine Learning, ICMLâ24. JMLR.org, 2024. Meister, C., Pimentel, T., Haller, P., J Ì ager, L., Cotterell, R., and Levy, R. Revisiting the Uniform Information Den- sity hypothesis. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Con- ference on Empirical Methods in Natural Language Pro- cessing, p. 963â980, Online and Punta Cana, Domini- can Republic, November 2021. Association for Computa- tional Linguistics. doi: 10.18653/v1/2021.emnlp-main. 74.URLhttps://aclanthology.org/2021. emnlp-main.74/. Mistral AI and NVIDIA.Mistral nemo.https:// mistral.ai/news/mistral-nemo, 2024. A 12B parameter model built in collaboration with NVIDIA. Morcos, A. S., Raghu, M., and Bengio, S. Insights on repre- sentational similarity in neural networks with canonical correlation, 2018. URLhttps://arxiv.org/abs/ 1806.05759. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., J Ì eegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. Dinov2: Learning robust visual features without supervision, 2024. URL https://arxiv.org/abs/2304.07193. 10 Time SeriesâVisionâLanguage Alignment Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. Flickr30k en- tities: Collecting region-to-phrase correspondences for richer image-to-sentence models, 2016. URLhttps: //arxiv.org/abs/1505.04870. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Explor- ing the limits of transfer learning with a unified text-to- text transformer, 2023. URLhttps://arxiv.org/ abs/1910.10683. Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability, 2017. URL https://arxiv.org/abs/1706.05806. Sch Ì onemann, P. H. A generalized solution of the orthogonal procrustes problem. Psychometrika, 31(1):1â10, March 1966. Sim Ì eoni, O., Vo, H. V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S., Rama- monjisoa, M., Massa, F., Haziza, D., Wehrstedt, L., Wang, J., Darcet, T., Moutakanni, T., Sentana, L., Roberts, C., Vedaldi, A., Tolan, J., Brandt, J., Couprie, C., Mairal, J., J Ì egou, H., Labatut, P., and Bojanowski, P. Dinov3, 2025. URL https://arxiv.org/abs/2508.10104. Sun, Q., Fang, Y., Wu, L., Wang, X., and Cao, Y. Eva-clip: Improved training techniques for clip at scale, 2023. URL https://arxiv.org/abs/2303.15389. Team, G., Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ram Ì e, A., Rivi ` ere, M., Rouillard, L., Mesnard, T., Cideron, G., bastien Grill, J., Ramos, S., Yvinec, E., Casbon, M., Pot, E., Penchev, I., Liu, G., Visin, F., Kenealy, K., Beyer, L., Zhai, X., Tsitsulin, A., Busa-Fekete, R., Feng, A., Sachdeva, N., Coleman, B., Gao, Y., Mustafa, B., Barr, I., Parisotto, E., Tian, D., Eyal, M., Cherry, C., Peter, J.-T., Sinopalnikov, D., Bhupatiraju, S., Agarwal, R., Kazemi, M., Malkin, D., Kumar, R., Vilar, D., Brusilovsky, I., Luo, J., Steiner, A., Friesen, A., Sharma, A., Sharma, A., Gilady, A. M., Goedeckemeyer, A., Saade, A., Feng, A., Kolesnikov, A., Bendebury, A., Abdagic, A., Vadi, A., Gy Ì orgy, A., Pinto, A. S., Das, A., Bapna, A., Miech, A., Yang, A., Paterson, A., Shenoy, A., Chakrabarti, A., Piot, B., Wu, B., Shahriari, B., Petrini, B., Chen, C., Lan, C. L., Choquette-Choo, C. A., Carey, C., Brick, C., Deutsch, D., Eisenbud, D., Cattle, D., Cheng, D., Paparas, D., Sreepathihalli, D. S., Reid, D., Tran, D., Zelle, D., Noland, E., Huizenga, E., Kharitonov, E., Liu, F., Amirkhanyan, G., Cameron, G., Hashemi, H., Klimczak-Pluci Ì nska, H., Singh, H., Mehta, H., Lehri, H. T., Hazimeh, H., Ballantyne, I., Szpektor, I., Nardini, I., Pouget-Abadie, J., Chan, J., Stanton, J., Wieting, J., Lai, J., Orbay, J., Fernandez, J., Newlan, J., yeong Ji, J., Singh, J., Black, K., Yu, K., Hui, K., Vodrahalli, K., Greff, K., Qiu, L., Valentine, M., Coelho, M., Ritter, M., Hoffman, M., Watson, M., Chaturvedi, M., Moyni- han, M., Ma, M., Babar, N., Noy, N., Byrd, N., Roy, N., Momchev, N., Chauhan, N., Sachdeva, N., Bunyan, O., Botarda, P., Caron, P., Rubenstein, P. K., Culliton, P., Schmid, P., Sessa, P. G., Xu, P., Stanczyk, P., Tafti, P., Shivanna, R., Wu, R., Pan, R., Rokni, R., Willoughby, R., Vallu, R., Mullins, R., Jerome, S., Smoot, S., Gir- gin, S., Iqbal, S., Reddy, S., Sheth, S., P Ì oder, S., Bhat- nagar, S., Panyam, S. R., Eiger, S., Zhang, S., Liu, T., Yacovone, T., Liechty, T., Kalra, U., Evci, U., Misra, V., Roseberry, V., Feinberg, V., Kolesnikov, V., Han, W., Kwon, W., Chen, X., Chow, Y., Zhu, Y., Wei, Z., Egyed, Z., Cotruta, V., Giang, M., Kirk, P., Rao, A., Black, K., Babar, N., Lo, J., Moreira, E., Martins, L. G., Sanseviero, O., Gonzalez, L., Gleicher, Z., Warkentin, T., Mirrokni, V., Senter, E., Collins, E., Barral, J., Ghahra- mani, Z., Hadsell, R., Matias, Y., Sculley, D., Petrov, S., Fiedel, N., Shazeer, N., Vinyals, O., Dean, J., Hass- abis, D., Kavukcuoglu, K., Farabet, C., Buchatskaya, E., Alayrac, J.-B., Anil, R., Dmitry, Lepikhin, Borgeaud, S., Bachem, O., Joulin, A., Andreev, A., Hardin, C., Dadashi, R., and Hussenot, L. Gemma 3 technical report, 2025. URL https://arxiv.org/abs/2503.19786. Tjandrasuwita, M., Ekbote, C., Ziyin, L., and Liang, P. P. Understanding the emergence of multimodal representa- tion alignment, 2025. URLhttps://arxiv.org/ abs/2502.16282. Tschannen, M., Gritsenko, A., Wang, X., Naeem, M. F., Alabdulmohsin, I., Parthasarathy, N., Evans, T., Beyer, L., Xia, Y., Mustafa, B., H Ì enaff, O., Harmsen, J., Steiner, A., and Zhai, X. Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localiza- tion, and dense features, 2025. URLhttps://arxiv. org/abs/2502.14786. Urbanek, J., Bordes, F., Astolfi, P., Williamson, M., Sharma, V., and Romero-Soriano, A. A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions, 2024. URLhttps://arxiv.org/abs/ 2312.08578. van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding, 2019. URL https://arxiv.org/abs/1807.03748. 11 Time SeriesâVisionâLanguage Alignment Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F. I., Samek, W., and Schaeffter, T. PTB-XL, a large publicly available electrocardiography dataset. Sci. Data, 7(1):154, May 2020. Wang, F., Xu, J., and Yu, L. From token to rhythm: A multi-scale approach for ecg-language pretraining, 2025. URL https://arxiv.org/abs/2506.21803. Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F. Text embeddings by weakly- supervised contrastive pre-training, 2024a. URLhttps: //arxiv.org/abs/2212.03533. Wang, L., Yang, N., Huang, X., Yang, L., Majumder, R., and Wei, F. Improving text embeddings with large language models, 2024b. URLhttps://arxiv.org/abs/ 2401.00368. Wang, T. and Isola, P. Understanding contrastive repre- sentation learning through alignment and uniformity on the hypersphere, 2022. URLhttps://arxiv.org/ abs/2005.10242. Wang, Z., Wu, Z., Agarwal, D., and Sun, J. Medclip: Contrastive learning from unpaired medical images and text, 2022. URLhttps://arxiv.org/abs/2210. 10163. Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D. Unified training of universal time series fore- casting transformers, 2024. URLhttps://arxiv. org/abs/2402.02592. Xu, H., Ghosh, G., Huang, P.-Y., Okhonko, D., Agha- janyan, A., Metze, F., Zettlemoyer, L., and Feichten- hofer, C. Videoclip: Contrastive pre-training for zero- shot video-text understanding, 2021.URLhttps: //arxiv.org/abs/2109.14084. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report, 2025. URLhttps: //arxiv.org/abs/2505.09388. Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. Sig- moid loss for language image pre-training, 2023. URL https://arxiv.org/abs/2303.15343. Zhang, R., Guo, Z., Zhang, W., Li, K., Miao, X., Cui, B., Qiao, Y., Gao, P., and Li, H. Pointclip: Point cloud understanding by clip, 2021. URLhttps://arxiv. org/abs/2112.02413. Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. Qwen3 embedding: Advancing text embedding and reranking through foundation models, 2025. URL https://arxiv.org/abs/2506.05176. Zhao, Y., Kang, J., Zhang, T., Han, P., and Chen, T. Ecg- chat: A large ecg-language model for cardiac disease diagnosis, 2025. URLhttps://arxiv.org/abs/ 2408.08849. Zhou, L., Yashwante, P., Fisher, M., Sampieri, A., Zhou, Z., Galasso, F., and Yu, R. Cats-bench: Can language models describe time series?, 2026. URLhttps:// arxiv.org/abs/2509.20823. 12 Time SeriesâVisionâLanguage Alignment A. Limitations and Future Directions Several questions remain open for future research. Our main goal was to provide a first systematic study of trimodal alignment in contrastive spaces involving time series, vision and language, rather than a definitive characterization of all possible settings. Our analysis focuses on univariate time series, and it remains an open question whether the alignment behaviors observed here extend to multivariate signals. CaTS captions are synthetic, which may not fully capture the diversity of human writing styles. We partially address this by evaluating on CaTS human-written, directly descriptive captions but is limited in scale. While we also test on real clinical text from MIMIC and PTB-XL, these datasets are domain-specific and rely on indirect language, which constrains their generalization for studying direct textual grounding. Also to ensure controlled and fair comparison across a large number of encoder combinations, we adopt a fixed contrastive training protocol with frozen encoders and train projection heads. This design isolates intrinsic representational compatibility across modalities, but precludes studying how end-to-end fine-tuning or modality-specific adaptation might further shape alignment. Exploring such training regimes would be a natural extension, but doing so exhaustively across all configurations considered in this work would be computationally prohibitive. An important direction for future work is to investigate how the alignment phenomena that we identify correlate with downstream task performance in multimodal time series applications. However, large-scale datasets that simultaneously include aligned time series, visual representations, and natural language supervision remain scarce, which limits systematic evaluation of task-level transfer. We hope that our findings highlight the need for richer multimodal time series datasets and stimulate further development, helping in more studies of how representational alignment in this trimodal setting translates into practical performance gains. B. Continued Related Work A long-standing theme across machine learning and computational neuroscience is that learned representations can be meaningfully compared across models, layers, training runs, and even across fundamentally different modalities. Early work in neuroscience formalized this idea through representational similarity analysis (RSA), which compares the geometry of activation patterns via representational (dis)similarity matrices rather than attempting neuron-to-neuron correspondence (Kriegeskorte et al., 2008). This perspective has strongly influenced modern deep learning research, where alignment is typically operationalized as the extent to which two embedding spaces share structure up to simple transformations. Though a central challenge in measuring representational similarity is that different models may implement equivalent functions using different bases, permutations, or scalings which makes direct coordinate-wise comparison ill-defined. Canonical correlation analysis (CCA) and its variants address this by comparing subspaces rather than individual units, motivating widely used tools such as SVCCA (Raghu et al., 2017) and projection-weighted CCA (PWCCA) (Morcos et al., 2018). CKA further refines this approach by providing a similarity index that is stable across random initializations and closely connected to CCA while importantly remaining computationally tractable (Kornblith et al., 2019). Another common metric of alignment is Procrustes analysis, which measures how well two representation sets can be matched via an optimal orthogonal transformation, with classical roots in the orthogonal Procrustes problem (Sch Ì onemann, 1966). More recent work has begun to systematically benchmark and compare these similarity measures across architectures, datasets, and modalities, demonstrating that different metrics emphasize different invariances (Klabunde et al., 2025; Tjandrasuwita et al., 2025). Beyond measurement, lots of works explicitly seeks to induce alignment through training objectives. Early multiview learning methods such as deep canonical correlation analysis (DCCA) learn nonlinear transformations of two views whose outputs are maximally correlated (Andrew et al., 2013), with extensions to more than two views enabling multiway alignment through DGCCA (Benton et al., 2017). In modern models, cross modal alignment is most commonly induced via contrastive learning objectives. CLIP (Radford et al., 2021) and ALIGN (Jia et al., 2021) demonstrated that large-scale imageâtext contrastive training can produce shared embedding spaces with strong zero-shot transfer properties, while distillation-based methods such as ALBEF improve robustness and alignment quality under noisy supervision (Li et al., 2021). Theoretical and empirical analyses of contrastive learning characterize these objectives as balancing alignment between positive pairs and uniformity of the representation distribution (Wang & Isola, 2022). While large-scale contrastive training can induce shared representation spaces across modalities, less is known about the geometry, asymmetry, and limits of such alignment particularly when modalities have fundamentally different inductive biases such as time series. Contrastive alignment has been extended beyond imageâtext to audioâtext (Elizalde et al., 2022), imageâtextâaudio (Guzhov 13 Time SeriesâVisionâLanguage Alignment et al., 2021), videoâtext (Xu et al., 2021), time series-text (Buiting et al., 2025), and 3Dâtext (Hadgi et al., 2025). More unified frameworks align multiple modalities within a shared embedding space, including VATT, which jointly learns from raw video, audio, and text (Akbari et al., 2021), and ImageBind (Girdhar et al., 2023), extends a VL backbone to modalities via image-paired training. C. Datasets We evaluate alignment using four datasets which we use to isolate the roles of semantic explicitness, visual mediation, and temporal complexity. All datasets are processed into aligned triplets of time series, visual plots, and text, with consistent trainâtest splits and identical preprocessing across modalities. Because naturally occurring trimodal datasets are extremely scarce for the three modalities that we study, we complement available data with carefully generated synthetic and human-authored modalities to complete missing components wherever required. C.1. CaTS CaTS is a multimodal dataset and benchmark which we use to study how numeric time series align with vision and language under controlled direct semantic conditions. Each sample consists of an univariate time series, a rendered plot of that series, and a natural language caption that describes observable temporal structure. The dataset spans 11 domains, including agriculture, air quality, border crossings, COVID statistics, crime, demography, diet, online retail, road injuries, CO 2 , and Walmart sales. Figure 12 shows a representative CaTS triplet. We use the original CaTS training split (16k) for all training experiments. Evaluation is performed on held-out test sets for each caption variant, with a fixed validation set of 1k samples. All variants share identical time series and plots, differing only in textual descriptions. Time Series 83.65, 81.28, 86.98, 93.55, 96.89, 100.21, 101.88, 100.0, 103.96, 109.67 Plot Text Indiaâs Total Factor Productivity index demon- strates consistent upward momentum from 2008 to 2017. Starting at 83.65 in 2008, the index experiences an initial dip to 81.28 in 2009 before steadily climbing to a peak of 109.67 in 2017, a 31% increase from the initial value . . . The overall pattern reveals sustained growth with minimal volatility following the initial 2009 recovery. Figure 12. Example CaTS triplet consisting of a numeric time series, its visual plot, and a reference caption. ID â 26.81 Avg char length â 20.32 Avg word length â 2.74 âoverall upward trendâ âintial dropâ âspike in the middleâ âfluctuated around the 100 mark, slight decrease overallâ âdownard progression with an intial growthâ âconsistent upward trend, steady growth, strong performanceâ âfluctuating pattern, significant dip below the series mean of 97.59, slight increase, falling to the minimum of 92.3.â âstarting at 101.45 in 2012, reaching a minimum of 91.25 in 2018, ending at 97.54 in 2019, slightly above the historical meanâ From December 2021 through September 2023, personal vehicle crossings at the Ferry port show a pronounced seasonal rhythm, marked by repeated winter lows and strong summer peaks .... underscores a consistent seasonal structure, with passenger activity expanding sharply during the warm months and contracting each winter. âdownward progressionâ ID â 69.48 Avg char length â 69.61 Avg word length â 10.24 ID â 149.05 Avg char length â 159.26 Avg word length â 25.82 ID â 416.81 Avg char length â 584.46 Avg word length â 94.40 1-phrase2-phrase4-phrase original Figure 13. KDEs over projected embedding values and caption examples across CaTS variants with increasing ID. To study the effect of semantic expressiveness, we generate multiple caption ID variants derived from the same underlying 14 Time SeriesâVisionâLanguage Alignment time series, plot, and metadata. Figure 13 illustrates distributional shifts in caption embeddings, visualized as kernel density estimates (KDEs) over projected embedding values. Starting from the original captions, we create progressively condensed versions containing approximately one phrase, two phrases, or four phrases. The condensed caption variants are generated only for evaluation on test samples. All variants are generated using a fixed protocol that preserves numeric fidelity while controlling verbosity. Caption extraction/summarization are performed using two LLMs, GPT-4o-mini and LLaMA-3.2-90B, with identical instructions applied across all samples. For these condensed variants, we instruct the language models to compress the original caption to a target number of semantic phrases, prioritizing the most salient temporal patterns, trends, and numeric attributes while discarding secondary details. This phrase-level constraint enables controlled reduction of semantic explicitness without altering the underlying grounding or introducing new information. All synthetically generated caption variants are validated following the same verification protocol used in the original CaTS construction. Next to study whether substantially increasing semantic richness during training can further improve multimodal alignment, we construct a High-ID variant of CaTS for both training and test splits. This dataset preserves the original time series and visual plots and we only modify the textual modality. High-ID captions are designed to be significantly more detailed than the original captions while maintaining strict numeric fidelity. Captions are generated independently for each sample using large language models (GPT-4o-mini and LLaMA-3.2-90B). Each model receives only the sample-specific metadata, and raw time series values. All captions are generated with low temperature to ensure determinism and consistency and all captions are generated using the following fixed system prompt, which explicitly enforces exhaustive numerical reporting, pattern and phase-wise analysis, and strict grounding in the provided data. The resulting High-ID captions are substantially longer and more semantically dense than the original captions, exceeding them by more than a factor of two in measured ID. Write an EXTREMELY detailed and comprehensive analysis of this time series data. Time Series: <ts> Metadata: <metadata> Your caption should include MOST of the following in 5-7 sentences: 1. CONTEXT: Full description of what this data represents, including location, measurement type, and exact time period covered 2. STATISTICS: Exact numerical values for mean, minimum, maximum, and standard deviation, with explicit comparisons to historical norms 3. OVERALL TREND: Primary direction and magnitude of change across the entire series 4. PHASE ANALYSIS: Break the series into distinct phases/segments, describe each with specific values and date ranges 5. PATTERNS: Any cyclical, seasonal, or periodic behavior with specific period lengths 6. ANOMALIES: Any unusual spikes, drops, or deviations with exact positions in the timeline and their values 7. COMPARISONS: Explicit numerical comparisons of this series to historical/global averages 8. TEMPORAL DETAILS: Specific dates, years, months, or time ranges for all key events and transitions CRITICAL RULES: - Be exhaustive and thorough - Include specific numbers for EVERY claim - Do not use speculative language - only describe what is evident in the data - Do not summarize briefly - provide full detail - Minimum 150 words - REMEMBER NOT ALL INFORMATION IS NEEDED TO BE PRESENT IF IT IS NOT EVIDENT IN THE DATA. C.2. TRUCE a) Genericb) Styledc) Annotated Figure 14. Representative TRUCE visual variants: generic, styled, and annotated plots. TRUCE is a controlled dataset which we use to isolate the effect of visual mediation on time series alignment. It consists of short, fixed-length univariate time series paired with concise direct textual descriptions. All time series have a length 15 Time SeriesâVisionâLanguage Alignment of 12 timesteps which enables exact visual annotation of every point. The training split consists 1,968 samples and the test split consists 492 samples. Representative TRUCE captions include âIncreases linearly throughout the entirety of the graph,â âSteady increase throughout,â and âTroughs near the beginning.â Captions are short (approximately 85 characters on average) and directly describe local temporal structure, patterns, and trends. For each time series, we generate multiple visual variations as illustrated in Figure 14: (a) a generic plot showing only the raw time-series curve, with no axes, annotations, or semantic markers, (b) a styled plot with randomized visual properties (line styles, colors, markers), and (c) an annotated plot in which every time step is labeled with its numeric value and has styling. C.3. MIMIC-IV ECG The MIMIC-IV ECG dataset provides real-world clinical time series paired with unstructured diagnostic reports. Each sample consists of a full-length ECG waveform and an associated clinical note that does not explicitly describe waveform structure. We use 21,000 samples with non-empty reports and valid waveform files. Time series is of 5,000 steps per record. Waveforms are rendered as line plots with axes and units. Figure 15 illustrates a triplet from MIMIC. We randomly shuffle the dataset and split it into 16,000 training samples and 4,000 test samples, with a fixed 1k validation set. All splits are disjoint at the subjectâstudy level. Clinical reports are concatenated across available fields and used verbatim, without summarization or restructuring. As a result, the text modality provides indirect semantic supervision, focusing on diagnoses and observations rather than explicit temporal descriptions. MIMIC also helps in analysis of alignment when time series are substantially longer than those in CaTS or TRUCE. Time Series 0.02, 0.02, 0.015, 0.015, 0.015, 0.02, 0.02, 0.015, 0.01, 0.01, 0.015, 0.02, 0.01, 0.0, 0.01, 0.015, ... 0.015, 0.01, -0.01, -0.01, -0.01, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, -0.01, -0.015, -0.015, -0.015 PlotText Atrial fibrillation with rapid ventricular response Inferi- or/lateral ST-T changes are nonspecific Low QRS volt- ages in limb leads Abnormal ECG Figure 15. Example MIMIC triplet consisting of an ECG time series, its waveform visualization, and a report. C.4. PTB-XL PTB-XL is a another ECG dataset which we use with standardized waveform recordings and diagnostic reports written primarily in German. Each recording is a fixed-length 10-second ECG sampled at 500 Hz, giving time series of length 5,000. We select samples with non-empty reports and valid waveform files, resulting in 21,000 aligned triplets. As with MIMIC, we render the full waveform with axes and units. Reports are used verbatim and remain in German, introducing a multilingual setting. The text modality again provides indirect semantic supervision, referring to clinical conditions rather than explicitly describing waveform structure. PTB-XL allows us to disentangle the effects of indirect grounding from language mismatch, which complements MIMIC by holding the signal structure constant while varying linguistic properties. For PTB-XL, we evaluate a subset of configurations that use Qwen embedding models for the text modality. Qwen embeddings provide multilingual support, which is appropriate for the multilingual clinical text found in PTB-XL. Time Series -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.035, -0.031, -0.03, -0.03, -0.03, -0.03, -0.03, -0.027, -0.024, ... 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12, 0.12 PlotText sinusrhythmus ueberdrehter linkstyprechtsschenkel- block Figure 16. Example PTB-XL triplet consisting of an ECG time series, its waveform visualization, and a report. 16 Time SeriesâVisionâLanguage Alignment D. Encoders and Trimodal Combinations Table 3. Triplet model combinations used in our experiments (Grouped by Size Bounds) #Vision EncoderText EncoderTime Series Encoder Lower Bound Models 1DINOv2-Base (86.6M)T5-Small (60.5M)Moirai (91M) 2ViT-Base (86.4M)T5-Small (60.5M)MOMENT (125M) 3 DINOv2-Base (86.6M)T5-Small (60.5M)Chronos (201M) 4DINOv2-Base (86.6M)E5-Base-v2 (109M)TimesFM (200M) 5ViT-Base (86.4M)E5-Base-v2 (109M)Chronos (201M) 6 SigLIP2-Base (375M)E5-Base-v2 (109M)Moirai (91M) 7SigLIP2-Base (375M)E5-Base-v2 (109M)MOMENT (125M) 8SigLIP2-Base (375M)T5-Small (60.5M)TimesFM (200M) 9 ViT-Base (86.4M)Qwen3-Embedding (596M)Moirai (91M) 10DINOv2-Base (86.6M)Qwen3-Embedding (596M)MOMENT (125M) 11ViT-Base (86.4M)Qwen3-Embedding (596M)TimesFM (200M) 12SigLIP2-Base (375M)Qwen3-Embedding (596M)Chronos (201M) Middle Bound Models 13ViT-Huge (632M)T5 (3B)MOMENT (385M) 14 SigLIP (878M)T5 (3B)TimesFM (500M) 15DINOv2-Giant (1.14B)T5 (3B)Moirai (311M) 16DINOv2-Giant (1.14B)T5 (3B)Chronos (709M) 17ViT-Huge (632M)Qwen3-Embedding (4B)Moirai (311M) 18ViT-Huge (632M)Qwen3-Embedding (4B)TimesFM (500M) 19DINOv2-Giant (1.14B)Qwen3-Embedding (4B)MOMENT (385M) 20 SigLIP (878M)Qwen3-Embedding (4B)Chronos (709M) 21SigLIP (878M)E5-Mistral (7B)Moirai (311M) 22SigLIP (878M)E5-Mistral (7B)MOMENT (385M) 23 ViT-Huge (632M)E5-Mistral (7B)Chronos (709M) 24 DINOv2-Giant (1.14B)E5-Mistral (7B)TimesFM (500M) Upper Bound Models 25DINOv3 ViT (7B)Qwen3-Embedding (8B)MOMENT (385M) 26 EVA-CLIP (8B)Qwen3-Embedding (8B)Chronos (709M) 27DINOv3 ViT (7B)Mistral-Nemo-Base (12B)Chronos (709M) 28EVA-CLIP (8B)Mistral-Nemo-Base (12B)Moirai (311M) 29EVA-CLIP (8B)Mistral-Nemo-Base (12B)TimesFM (500M) 30 EVA-CLIP (18B)Qwen3-Embedding (8B)Moirai (311M) 31EVA-CLIP (18B)Mistral-Nemo-Base (12B)MOMENT (385M) 32DINOv3 ViT (7B)Gemma-3 (27B)Moirai (311M) 33 EVA-CLIP (8B)Gemma-3 (27B)MOMENT (385M) 34EVA-CLIP (18B)Gemma-3 (27B)Chronos (709M) Models used.We evaluate a broad and diverse set of pretrained models across vision, language, and time series modalities, covering multiple architectural paradigms, training objectives, and model scales. 1. Vision encoders. We use nine vision encoders spanning model sizes from 86M to 18B parameters. Self-supervised models include DINOv2 (Oquab et al., 2024) (Base, 86.6M; Giant, 1.14B) and DINOv3 ViT (7B) (Sim Ì eoni et al., 2025). Supervised models include ViT (Dosovitskiy et al., 2021) (Base, 86.4M; Huge, 632M). Visionâlanguage pretrained models include SigLIP (Zhai et al., 2023) (Large, 878M), SigLIP2 (Tschannen et al., 2025) (Base, 375M), and EVA-CLIP (Sun et al., 2023) (8B, 18B). 17 Time SeriesâVisionâLanguage Alignment 2.Text encoders. We evaluate nine text encoders with scales ranging from 60M to 27B parameters. These include T5 (Raffel et al., 2023) (Small, 60.5M; 3B), E5-Base-v2 (109M) and E5-Mistral (7B) (Wang et al., 2024a;b), Qwen3- Embedding (Yang et al., 2025; Zhang et al., 2025) (596M, 4B, 8B), Mistral-Nemo-Base (12B) (Mistral AI & NVIDIA, 2024), and Gemma-3 (27B) (Team et al., 2025). 3.Time series encoders. We use eight pretrained time series encoders with model sizes ranging from under 100M to over 700M parameters. These include TimesFM (Das et al., 2024) (200M, 500M), Chronos (Ansari et al., 2024) (201M, 709M), MOMENT (Goswami et al., 2024) (125M, 385M), and Moirai (Woo et al., 2024) (91M, 311M). D.1. Representation Extraction. We adopt a consistent and controlled representation extraction protocol across text, vision, and time series models to ensure fair comparisons across modalities and families. All backbone models are kept frozen and used strictly as feature extractors. 1.Text representations. For text models (Qwen, Gemma, and Mistral), we extract token-level hidden states from the final transformer layer and apply mean pooling over valid tokens using the attention mask, yielding a single fixed-dimensional sentence embedding per sample. For embedding-style models (e.g., Qwen Embedding), we use the last-token or sequence-end representation following the modelâs recommended pooling strategy. 2. Vision representations. Vision representations are obtained by encoding images through pretrained vision backbones (e.g., EVA-CLIP, DINOv3), using either the modelâs native image embedding head or the[CLS]token from the final layer, depending on the architecture. 3. Time series representations. For time series encoders, we extract sequence-level embeddings by mean pooling over temporal hidden states returned by the encoder. In all cases, the resulting modality-specific embeddings are projected into a shared latent space of fixed dimension using trainable projection heads, while the backbone encoders remain frozen. D.2. Representative Configuration Subsets 1.Information density analysis (Figure 6). Results are averaged over 9 representative configurations corresponding to rows31, 29, 32, 34, 14, 22, 3, 9, 2 in Table 3. 2. VLâTS scaling analysis (Figure 7). Configurations labeled AâI in Figure 7 correspond to rowsA: 13, B: 18, C: 23, D: 14, E: 20, F: 22, G: 16, H: 19, I: 24 in Table 3. 3.Scaling under indirect textual supervision (Figure 8). Results are reported over 5 representative configurations10, 19, 25, 12, 20 in Table 3. 4. High-ID saturation analysis (Table 1). Results are averaged over configurations10, 20, 12, 5 from Table 3. E. Metrics We quantify cross-modal alignment using complementary metrics that capture global geometric correspondence, non-linear representational similarity, and local neighborhood consistency. All metrics are computed independently for each modality pair (TSâIMG, TSâTXT, IMGâTXT) using normalized embeddings. To ensure computational efficiency and stability, all geometric metrics are computed on a randomly sampled subset of at most2000paired samples when datasets exceed this size, using a fixed seed (42). Cosine is done on full test set. LetX =x i N i=1 andY =y i N i=1 denote paired embeddings from two modalities, where each embedding isâ 2 -normalized. E.1. Cosine Similarity Cosine similarity serves as a global semantic alignment metric. For paired samples, we compute the full similarity matrix S = XY †,(4) 18 Time SeriesâVisionâLanguage Alignment where S ij =âšx i ,y j â©. We report three statistics: Matched = 1 N N X i=1 S i ,Mismatched = 1 N (N â 1) X iÌž=j S ij ,Margin = 1 N N X i=1 ïŁ« ïŁ S i â 1 N â 1 X jÌž=i S ij ïŁ¶ ïŁž .(5) The margin reflects how well matched cross-modal pairs are separated from mismatched pairs. Higher values indicate stronger alignment. E.2. Procrustes Disparity (Global Geometry) We compute the normalized Procrustes disparity between embedding spaces. After centering both modalities, Ì X = Xâ 1 N X i x i ,(6) Ì Y = Yâ 1 N X i y i ,(7) we solve the orthogonal Procrustes problem min RâR dĂd , R †R=I â„ Ì XRâ Ì Yâ„ 2 F .(8) The optimal rotationRis obtained via the Kabsch solution using singular value decomposition with explicit reflection correction. We report the normalized disparity Disparity = â„ Ì XRâ Ì Yâ„ 2 F â„ Ì Yâ„ 2 F .(9) Lower values indicate stronger global geometric alignment. Figure 17. Illustration of Procrustes alignment. Embeddings from one modality are rotated to best match the other under an orthogonal transformation. Normalized residual error measures global geometric mismatch. E.3. CKA (Non-linear Similarity) To capture non-linear representational similarity, we compute CKA using an RBF kernel. Given embeddingsXandY, we compute K ij = exp âÎłâ„x i â x j â„ 2 ,(10) L ij = exp âÎłâ„y i â y j â„ 2 .(11) 19 Time SeriesâVisionâLanguage Alignment The bandwidthÎłis selected using the median heuristic over pairwise distances from both modalities. Kernels are centered using Ì K = HKH, Ì L = HLH,(12) where H = I â 1 N 11 †and where 1âR N denotes the all-ones vector. CKA is then computed as CKA(K,L) = âš Ì K, Ì Lâ© F â„ Ì Kâ„ F â„ Ì Lâ„ F .(13) Values closer to 1 indicate highly similar non-linear representational structure. E.4. Mutual k-Nearest Neighbors (Local Structure) To evaluate local neighborhood consistency, we compute mutualk-nearest neighbor overlap. For each embeddingx i , let N X k (i)denote itsknearest neighbors inXunder cosine distance, excluding itself. Similarly defineN Y k (i)forY. The mutual kNN score is Mutual-kNN = 1 Nk N X i=1 N X k (i)â©N Y k (i) .(14) This metric captures local neighborhood agreement across modalities. Higher values indicate stronger preservation of neighborhood structure. We usek = 5in all computations as want to preserve the results on local structure and increasing it more would turn it into a global geometry metric. Figure 18. Illustration of MutualkNN alignment. MutualkNN measures local neighborhood agreement across modalities. Points that share neighbors in both spaces contribute to higher scores. E.5. Cross-modal retrieval. For each modality pair (TSâIMG, TSâTXT, IMGâTXT), we evaluate cross-modal retrieval by treating embeddings from one modality as queries and ranking all embeddings from the paired modality over the full evaluation set using cosine similarity in the shared embedding space. For a dataset ofNpaired samples, the correct match for queryiis defined as the paired sample at indexi. Recall@k(R@1, R@5, R@10) is computed as the fraction of queries for which the correct match appears within the top-kranked candidates. Retrieval is evaluated in both directions for each modality pair (e.g., TSâIMG and IMGâTS), with rankings computed independently per direction. No thresholds or learned classifiers are used; the metrics depend solely on the rank order induced by cosine similarity, which is invariant to any positive temperature scaling. F. Learned Convergence Without External Coupling Before doing post-hoc alignment, we first understand whether representations learned independently within each modality already exhibit convergent structure, as suggested by PRH for vision and language models, as an initial motivation. We 20 Time SeriesâVisionâLanguage Alignment extend this diagnostic to a trimodal setting involving time series, vision, and language, and evaluate representational geometry in the absence of any external coupling. For each dataset for their test sets, we extract embeddings from independently pretrained encoders for time series, images, and text. No projection heads are trained and no contrastive or alignment objectives are applied. We then compare geometry directly using mean angular deviation (MAD) and cosine similarity. Figure 19. Mean angular deviation between modality pairs across datasets and model configurations. Values remain near90 ⊠for all pairs, indicating near-orthogonal geometry and a lack of inherent cross-modal convergence in independently pretrained encoders. Figure 19 reports mean angular deviation between modality pairs across 9 representative trimodal configurations which we used in VL-TS experiment (AâI) (see Appendix D.2) and four datasets. Across all datasets and modality pairs, MAD values remain close to90 ⊠, indicating near-orthogonal geometry. This behavior is consistent across domains, including cases where representations are deterministic renderings of the underlying signal (CaTS, TRUCE) and cases where text is indirectly related to the signal (PTB-XL, MIMIC). These results indicate that independently pretrained encoders do not exhibit inherent geometric convergence across these datasets used. Figure 20 visualizes cross-modal cosine similarity matrices computed over matched samples. Off-diagonal similarities remain uniformly close to zero across all datasets, with no discernible separation between matched and mismatched pairs. This also confirms that the lack of alignment is not merely a matter of scale but reflects the absence of latent cross-modal correspondence in the embedding spaces. Notably, TSâIMG similarity also remains close to zero in this uncoupled setting, although it is consistently higher than other pairs. Figure 20. Cross-modal cosine similarity matrices for matched samples across datasets. Off-diagonal similarities remain near zero with no separation between matched and mismatched pairs, confirming the absence of latent alignment without explicit coupling. These diagnostics establish a clear baseline: independently pretrained encoders for time series, vision, and language do not spontaneously converge to a shared representational geometry for different modality projections of time series data on the datasets studied here. This motivates the contrastive alignment experiments in the paper, where we introduce trainable projection heads to study how post-hoc alignment is induced. G. Experimental Details, Ablations, Loss Functions In this section, we describe the experimental protocol used to train and evaluate our models, including optimization settings, batching and precision choices, data handling, and the contrastive objective. We additionally present targeted ablations to assess the sensitivity of alignment geometry to optimization hyperparameters, training duration, and projection head design. G.1. Optimization, Batching and Precision Models are trained for 50 epochs using AdamW with learning rate1Ă 10 â4 and weight decay1Ă 10 â5 . A cosine decay schedule with linear warmup (2% of total steps) is applied. A fixed temperature of 0.2 is used throughout. Gradient clipping 21 Time SeriesâVisionâLanguage Alignment with maximum norm 1.0 is used to stabilize training. Early stopping with patience 5 is applied uniformly. All experiments use an effective batch size of 256, obtained via gradient accumulation when necessary. Mixed-precision (bfloat16) is used for very large encoders. All hyperparameters are held constant across datasets and model combinations. G.2. Data Handling Time series keep their original length and no truncation is applied. Images are resized to 384Ă384 pixels; Text inputs use a maximum length of 512 tokens; all samples fall within this limit, so full text is retained across all variants. G.3. Loss Let a minibatch consist ofNaligned triplets(z (i) ts ,z (i) img ,z (i) txt ) N i=1 , where all embeddings areâ 2 -normalized and lie inR d . Similarity is measured using the dot product s(a,b) = a †b and scaled by a temperature parameter Ï > 0. For a generic modality pair (x,y), the forward contrastive loss xâ y is defined as L xây =â 1 N N X i=1 log exp s(z (i) x ,z (i) y )/Ï P N j=1 exp s(z (i) x ,z (j) y )/Ï (15) The reverse direction y â x is defined analogously: L yâx =â 1 N N X i=1 log exp s(z (i) y ,z (i) x )/Ï P N j=1 exp s(z (i) y ,z (j) x )/Ï .(16) The symmetric bidirectional loss for modality pair (x,y) is then L x-y = 1 2 (L xây +L yâx ).(17) In our trimodal setting, we compute losses for all three modality pairs: time seriesâimage (TSâIMG), time seriesâtext (TSâTXT), and imageâtext (IMGâTXT). The final training objective is the equally weighted sum: L total =L ts-img +L ts-txt +L img-txt .(18) In all cases, negative examples are implicitly provided by other samples within the same minibatch, i.e., each non-matching pair (z (i) x ,z (j) y ) for j Ìž= i is treated as a negative. No external negative mining or memory bank is used. Equal Weighting of Modality Pairs in the Loss.We apply equal weights to all modality pairs in the trimodal contrastive objective. This choice reflects our goal of studying intrinsic representational compatibility rather than optimizing for a downstream task. Reweighting individual losses would introduce assumptions about the relative importance of different modality relationships, which would directly shape the geometry learned during training. In our setting, we treat each modality as a peer projection of a shared latent process, and no modality pair is favored a priori. There is therefore no principled basis for assigning greater weight to one interaction over another. For example, increasing the weight of TSâTXT would implicitly assume that linguistic alignment is more semantically important, while downweighting TSâIMG could assume that visual renderings are easier or partially redundant. Such choices would be arbitrary and closely tied to the hypothesis under study. Using equal weights provides the most neutral formulation, allowing cross-modal structure to emerge from the representations themselves rather than being imposed through optimization choices. Notably, the observed asymmetry between TSâIMG and TSâTXT alignment persists across model scales, encoder families, and text richness despite symmetric treatment in the loss. This suggests that the asymmetry reflects underlying representational compatibility. Exploring adaptive or modality-specific weighting is a natural direction for task-oriented settings for multimodal time series tasks, but is outside the scope of our goal. 22 Time SeriesâVisionâLanguage Alignment G.4. Compute Training cost varies substantially with model scale and the number of trainable parameters. For lower-bound configurations, each triplet combination required approximately 6â10 hours of training, while middle- and upper-bound configurations required between 1 and 3 days per run. All experiments were executed using 8 NVIDIA A100 GPUs on a shared high- performance computing node with sufficient CPU and memory resources to support large-scale multimodal training. These runtimes reflect the cost of training projection heads with frozen encoders. G.5. Setting Ablations We assess the sensitivity of trimodal alignment to contrastive optimization by comparing two configurations: b64 (batch size 64, temperature 0.7, 512-dimensional projection) and b256 (batch size 256, temperature 0.2, 1024-dimensional projection). We find that the b256 configuration consistently gives stronger alignment across modality pairs. These gains reflect the combined effects of larger negative sets, sharper contrastive separation, and higher-capacity projection heads. Differences are most pronounced in global geometric metrics: Procrustes disparity remains low for b256 (approximately 0.45) but degrades substantially under b64 (above 0.70). Epoch Ablations.We study the effect of training duration by performing an epoch ablation on a representative configura- tion, training the projection head for up to 100 epochs while keeping all encoders frozen. As summarized in Table 4, metrics improve rapidly during early training and largely stabilize by 40â50 epochs. Extending training beyond this point yields no consistent improvements and in some cases leads to slight degradation in alignment which shows diminishing returns once cross-modal correspondences are established. This suggests that prolonged optimization of the projection head can introduce geometric drift without improving semantic alignment. Based on these observations, we adopt a fixed training budget of 50 epochs for all experiments which helps us to balance computational efficiency with stable and robust alignment. Table 4. Epoch ablation on vith14 + t5 + moment-l. Metrics are averaged across all modality pairs. EpochsCosine MarginâProcrustesâCKAâMutual KNNâ 500.6760.5650.6580.103 100 0.6710.5770.6370.104 G.6. Projection Head Ablations We study the effect of projection head design on alignment by comparing a simple baseline projection head with a more expressive variant, which we use as our default throughout the paper. The baseline head consists of a shallow two-layer MLP with a ReLU nonlinearity, similar to very early contrastive learning setups in the field. Given an encoder outputx, the improved projection head is defined as: h(x) = LayerNorm(W 1 x + b 1 ), u(x) = Dropout(GELU(h(x))), Proj(x) = W 2 u(x) + b 2 . (19) whereW 1 âR mĂd in ,W 2 âR dĂm ,m = max(3d, 768), andd = 1024is the shared embedding dimension. The improved projection head increases capacity by introducing an expanded hidden layer, layer normalization, GELU activation, and dropout which provides more flexible feature re-mapping while keeping all encoders frozen. As shown in Figure 21, the improved projection head gives substantial gains in cosine similarity margins across all modality pairs. At the same time, Procrustes disparity is uniformly reduced, indicating that the learned embedding spaces are not only closer in a pointwise sense but also better aligned geometrically. These improvements hold across different families and combinations. Importantly, while the improved head strengthens alignment, the relative ordering between modality pairs remains stable, indicating that projection head design improves alignment quality without altering the underlying cross-modal asymmetries. We additionally experimented with deeper projection head variants that further increase depth and nonlinearity beyond the improved design (e.g., additional hidden layers and normalization), but did not observe consistent gains over the reported configuration. In most cases, alignment metrics saturated or showed marginal improvements at the cost of increased 23 Time SeriesâVisionâLanguage Alignment ABCDEF 0.4 0.6 0.8 Cosine Margin TS-IMG ABCDEF 0.4 0.6 0.8 Cosine Margin TS-TXT ABCDEF 0.4 0.6 0.8 Cosine Margin IMG-TXT ABCDEF 0.50 0.75 1.00 Procrustes ABCDEF 0.50 0.75 1.00 Procrustes ABCDEF 0.50 0.75 1.00 Procrustes BaselineImproved A: DINOv2 + Qwen + MOMENT B: DINOv2 + T5 + Chronos C: SigLIP + E5 + MOMENT D: SigLIP + Qwen + Chronos E: SigLIP + T5 + TimesFM F: ViT-H/14 + T5 + MOMENT Figure 21. Effect of projection head design on trimodal alignment. Comparison of baseline and improved projection heads across six model configurations. The top row shows cosine similarity margins (higher is better) for TSâIMG, TSâTXT, and IMGâTXT alignment, while the bottom row reports Procrustes disparity (lower is better). instability and sensitivity to hyperparameters. As a result, we adopt the improved projection head as a principled trade-off between capacity and robustness. We leave a more thorough exploration of end-to-end encoder fine-tuning or parameter- efficient adaptation to future work. While such approaches may further improve cross-modal alignment for our trimodal setting, they substantially increase computational cost and complicate fair comparison across heterogeneous encoder families, which is a central goal of this study. H. Information Density We define information density as the total surprisal of a caption under a fixed pretrained language model. Given a caption x = (w 1 ,...,w T ), we compute the per-token surprisal â(x) = 1 T T X t=1 â logp(w t | w <t ),(20) and define ID(x) = T ·â(x),(21) which corresponds to the total negative log-likelihood. Surprisal is computed using a pretrained GPT-2 language model with frozen parameters. There is no model-agnostic metric that directly measures the amount of semantic content expressed in unconstrained natural language. Motivated by informationâuniformity perspectives on language modeling (Levy & Jaeger, 2007; Meister et al., 2021), we use language-model surprisal as a principled proxy for expressed information under a fixed linguistic prior. Intuitively, natural language text that convey more specific, grounded content require more bits to encode and thus exhibit higher surprisal, whereas generic or templated descriptions are more predictable. While surprisal is not a perfect semantic measure, it provides a controlled and scalable operationalization for comparing text sets without introducing task- or model-specific heuristics. Our use of ID is also consistent with how PRH operationalizes textual richness. They show that visionâlanguage alignment improves by increasing caption density via longer and more descriptive captions, on the Densely-Captioned-Images (DCI) dataset (Urbanek et al., 2024). In this, caption length effectively serves as a proxy for textual density. We follow the 24 Time SeriesâVisionâLanguage Alignment same intuition, but make the notion of density more explicit and interpretable by jointly accounting for caption length and per-token surprisal under a fixed language model. This allows us to distinguish between captions that are merely longer and those that express more information, while remaining aligned with the PRH perspective. Length and Per-Token Controls. Because total surprisal scales with sequence length by construction, we explicitly report caption length, tokens, words alongside ID. Table 5 summarizes these statistics for the caption variants used in our experiments. Table 5. ID with explicit length and words controls (averaged over test captions). Caption VariantTokens (T )Words (â)ID 1-phrase3.142.7426.81 2-phrase14.0110.2469.48 4-phrase37.2225.82149.05 Dense128.7394.40416.81 High-ID312.28231.75867.40 Table 6. Stability of ID ordering across GPT variant language models. Caption VariantGPT-2GPT-2 MediumGPT-2 Large 1-phrase111 2-phrase222 4-phrase333 Dense444 High-ID555 Robustness to Language Model Choice. To assess sensitivity to the surprisal model, we recompute ID using GPT-2, GPT-2 Medium, and GPT-2 Large, keeping all other settings fixed. Across all three models, caption variants exhibit an identical ordering in information density (1-phrase<2-phrase<4-phrase<dense<high-ID), as shown in Table 6. At the caption level, ID values show high rank agreement across models, and all downstream alignment analyses exhibit the same qualitative trends. This indicates that ID provides a stable and model-robust axis for comparing caption variants. I. Hierarchical Training Objective for Joint VL-TS combinations We train all models in this setting using a hierarchical contrastive objective that aligns time series representations with joint VL semantics, while preserving internal VL consistency. Letz ts âR d denote the projected embedding of a time series input, and letz vl âR d denote the corresponding joint VL embedding obtained from the multimodal backbone. We optimize the symmetric InfoNCE loss forL tsâvl . The reverse directionL vlâts is also applied analogously. This objective encourages each time series embedding to align most strongly with its corresponding multimodal semantic representation, while repelling mismatched pairs within the batch. To preserve semantic coherence within the joint VL space, we use an auxiliary contrastive loss between vision-only and text-only embeddings to observe IMGâTXT alignment. L vât = 1 2 (L vât +L tâv ),(22) whereL vât andL tâv follow the same InfoNCE formulation. J. Line Plots as Visual Projections of TS and TXT A natural concern is whether the observed asymmetry between TSâIMG and TSâTXT alignment is an artifact of using visual plots that are direct renderings of the underlying time series. Because line plots are deterministic transformations of the signal, TSâIMG alignment may appear easier than TSâTXT, potentially amplifying the observed gap. We therefore discuss here whether the observed behavior reflects trivial pixel-level correspondence or deeper structural compatibility across modalities. Under the PRH, different modalities are viewed as projections of a shared latent reality. For time series, visual plots provide a structured projection that externalizes latent temporal properties, such as trends, extrema, and phase transitions, into explicit geometric form. This transformation does not introduce new information, but it does reorganize existing signal structure into a representation that is more directly comparable to spatial and symbolic modalities. To test whether TSâIMG alignment is driven by superficial visual matching, we also did evaluate robustness under nontrivial visual perturbations as shown in Figure 14. Using TRUCE, we compare generic, stylistically varied, and annotated plots while keeping the underlying time series and text fixed. If alignment were purely tautological, changes in visual realization 25 Time SeriesâVisionâLanguage Alignment would either have little effect across variants. Instead, we observe that TSâIMG alignment does not collapse under style variation and is consistently higher for styled and annotated plots than for generic renderings which shows that alignment tracks structured signal cues rather than brittle pixel correspondence. In contrast, TSâTXT alignment remains consistently weaker even when captions directly describe the same temporal structure. Styled and annotated plots further improve TSâIMG alignment by making implicit signal attributes explicit. However, the central observation is not the absolute gain from annotation, but the robustness of TSâIMG alignment across visually distinct yet semantically equivalent projections. This shows that visual plots function as structured representations that externalize temporal structure in a form that is naturally compatible with contrastive alignment. We believe this asymmetry is not tied to a particular visual encoding and it persists across diverse direct visual realizations of the same underlying signal. An interesting open question is how these observations extend to more implicit visual encodings of temporal phenomena. For example, consider a physics simulation in which a ball rolls down a ramp: the ballâs trajectory can be represented as a time series, individual frames of the simulation as the visual modality, and text can describe both the evolving dynamics and summary statistics of the motion. Such settings would allow studying alignment when visual structure emerges from physical dynamics rather than explicit plotting conventions. To the best of our knowledge, however, no existing datasets provide well-aligned time seriesâimageâtext triplets of this form, and we leave this direction to future work. K. Representative Results of Trimodal Configurations Effect of Scaling All Three Encoders.Figure 22 shows that increasing the capacity of all three encoders leads to consistent improvements in cross-modal alignment, but with marked asymmetry across modality pairs. Alignment involving time series improves most strongly for TSâIMG across all metrics, while gains for TSâTXT are smaller and less consistent. Notably, geometric metrics and cosine margins exhibit similar trends which indicates that scaling improves both global geometry and local neighborhood structure. This shows that while model scale improves alignment, the extent of improvement depends on the semantic compatibility of the modality pair, with TS-TXT remaining the most challenging even at larger scales. 0.8B5.5B15.4B Model Size 0.4 0.6 0.8 Score Procrustes 0.8B5.5B15.4B Model Size 0.6 0.7 0.8 CKA 0.8B5.5B15.4B Model Size 0.10 0.15 0.20 MutualKNN 0.8B5.5B15.4B Model Size 0.6 0.7 0.8 Cosine Margins TS-IMGTS-TXTIMG-TXT A: DINOv2-Base + Qwen-0.6B + MOMENT-Base (0.8B)B: DINOv2-Large + Qwen + MOMENT (5.5B)C: DINO-7B + Qwen-8B + MOMENT (15.4B) Figure 22. Effect of joint encoder scaling on trimodal alignment. Alignment improves with model size across modality pairs, but gains are uneven. When the time series encoder is already saturated (from B to C), additional scaling yields substantial improvements for TSâIMG alignment, while TSâTXT exhibits smaller and less consistent changes across metrics. This indicates that TSâTXT alignment remains more challenging than TSâIMG: scaling the vision encoder produces clear gains for TSâIMG, whereas corresponding improvements to the text encoder lead to only modest or saturating gains for TSâTXT. Effect of Time Series Encoder Capacity. To isolate the role of time series encoder strength, we scale the time series encoder while holding vision and text encoders fixed. Figure 23 compares TimesFM-200M and TimesFM-500M paired with EVA-CLIP (8B) and Mistral (12B). Increasing time series encoder capacity leads to consistent improvements in TSâIMG and TSâTXT alignment across cosine similarity, retrieval metrics, and Procrustes geometry. IMGâTXT alignment shows only minimal improvement. L. Human Linguistic Shifts We compare alignment obtained using synthetic captions versus human-written captions from the CaTS human test set, evaluated with CaTS-trained models, as shown in Figure 24 for a small set of representative domain samples (two domains). 26 Time SeriesâVisionâLanguage Alignment TSIMG TSTXT IMGTXT 0.2 0.4 0.6 0.8 (a) Cosine Margins R@1 R@5 R@10 0.05 0.10 0.20 (b) Retrieval Performance TSIMG TSTXT IMGTXT 0.2 0.4 0.6 0.8 (c) Procrustes Alignment TimesFM-200MTimesFM-500M Figure 23. Comparison of multimodal alignment using TimesFM-200M and TimesFM-500M time series encoders, with EVA-CLIP (8B) and Mistral (12B) held fixed. (a) Cosine margins for TSâIMG, TSâTXT, and IMGâTXT. (b) Retrieval performance (R@1, R@5, R@10). (c) Procrustes alignment (lower is better). Both settings operate in the direct text representation regime. Across configurations and modality pairs, we observe only modest shifts in alignment when replacing synthetic captions with human-authored descriptions. In some configurations, alignment slightly improves under human captions, while in others it exhibits a small degradation; however, these effects are consistently limited in magnitude across both cosine similarity and Procrustes-based geometry. Human captions exhibit higher ID on the 2 domain test set (409.15 vs. 381.95), inducing a mild out-of-distribution linguistic shift relative to the training distribution. Despite this shift, alignment remains largely stable, indicating that the learned trimodal embedding spaces generalize well to natural human linguistic variability. The modest difference between LLM and human captions suggests that the shift is insufficient to induce collapse or improvement which also further support the generalization of models trained on CaTS to human-authored descriptions. ABC 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Cosine Margin TSTXT Cosine ABC 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Cosine Margin IMGTXT Cosine ABC 0.0 0.2 0.4 0.6 0.8 1.0 Procrustes Disparity TSTXT Procrustes ABC 0.0 0.2 0.4 0.6 0.8 1.0 Procrustes Disparity IMGTXT Procrustes Human Captions Synthetic Captions A: DINOv2-L + Gemma_27b + MoiraiB: DINOv2-L + Mistral_12b + ChronosC: DINOv2-L + Qwen_8b + Moment Figure 24. Alignment comparison between LLM-generated and human-written captions across modality pairs. M. Discussion on Semantic Explicitness for Alignment Asymmetry Our findings show that trimodal alignment is not uniform: TSâIMG alignment consistently outperforms TSâTXT, and images can act as effective intermediaries between time series and text. We use the term semantic explicitness in an operational, descriptive sense to interpret this asymmetry, referring to the degree to which a modalityâs encoding makes underlying semantic structure directly observable rather than requiring abstraction or inference. Time series encode temporal properties implicitly. Characteristics like trends, periodicity, or anomalies are not directly observable but they must be inferred through computation over raw numerical values. Recognizing an âincreasing trendâ in a sequence[v 1 ,v 2 ,...,v T ]requires computing differences, aggregating, and applying some decision criterion. The time series encoder must learn to perform this inference from data. Visual plots make this structure explicit. A trend manifests as a visible slope, a peak becomes a geometric feature, and periodicity appears as repeating spatial patterns. The transformation 27 Time SeriesâVisionâLanguage Alignment from time series to plot largely preserves structure while externalizing latent patterns into directly observable form. Vision encoders, with their inductive biases for detecting edges, slopes, and shapes, can use this explicit structure directly. Text occupies a different position. Captions explicitly name semantic concepts such as the phrase âincreasing trendâ directly references the property but this naming is abstract. The symbol âtrendâ can refer to infinitely many numerical realizations; it provides no grounding in specific values or shapes. A caption describes what is present without encoding how it appears numerically or visually. This gap may explain why TSâTXT alignment is weaker than TSâIMG in our experiments: matching implicit numerical patterns to abstract language is harder than matching them to structured visual representations. We believe this interpretation is consistent with the alignment patterns observed across our experiments. TSâIMG outperforms TSâTXT.Across models and metrics, time series align more strongly with visual plots than with text, consistently exhibiting closer proximity to images than to language in the shared embedding space. Plots externalize temporal structure into explicit geometric form, whereas text abstracts over specific realizations, suggesting that how information is encoded matters as much as what is encoded. Images function as semantic intermediaries.Introducing the image modality consistently improves TSâTXT alignment compared to bimodal training. This is consistent with images serving as explicit intermediates: time series can align with plots (implicitâexplicit), and plots can align with text (explicitâabstract), providing a pathway that circumvents the direct implicit-to-abstract mapping. The trimodal objective allows representations to leverage this intermediate grounding. Notably, this effect is also particularly strong for vision encoders pretrained with a VL objective (SigLIP), even when only their vision component is used. This suggests that VL pretraining shapes image representations to be more semantically compatible with text by inducing shared geometric structure in the embedding space (Radford et al., 2021; Levi & Gilboa, 2025), further strengthening the role of images as effective semantic bridges in trimodal alignment. Information density shows diminishing returns. Increasing caption richness improves alignment at low-to-moderate levels, but further increases yield marginal gains. This suggests that while richer text increases specificity which effectively raising its semantic explicitness, the abstraction inherent to symbolic representation limits further improvement. Text can become more detailed, but it remains symbolic; it cannot become a continuous representation of the data. Indirect text degrades alignment.Clinical reports in MIMIC describe diagnostic conclusions (âatrial fibrillationâ) rather than waveform structure (âsharp peak followed by gradual decayâ). This further abstraction widens the explicitness gap, and alignment suffers accordingly. These observations suggest a refinement to how multimodal convergence should be understood. The PRH predicts that representations of the same underlying structure will converge. Our findings indicate this convergence may be conditional: in our experiments, modalities that encode structure with similar explicitness align more readily, while those with larger explicitness gaps show weaker alignment. Rather than converging uniformly toward a single shared representation, modalities appear to align to different degrees depending on their representational format. Semantic explicitness as a controlled ablation.To directly test whether the abstraction gap of text contributes to weaker TSâTXT alignment, we perform a targeted caption ablation in which we replace the original CaTS captions sets with structured, explicitly grounded descriptions that bind language directly to numeric values, temporal ranges, and phase structure (e.g., explicit start/end values, monotonic changes, and segmented trends). Importantly, these structured captions are designed to be more explicit while maintaining ID comparable to the original CaTS captions, and remain far below the high-ID regime that exhibited saturation. This intervention is designed as a controlled diagnostic that isolates one concrete dimension of textual grounding while holding overall ID approximately constant. As shown in Table 7, more explicitly grounded captions provide improvements in TSâTXT alignment across all metrics, with parallel gains for IMGâTXT alignment. These results further support the interpretation that alignment asymmetry is driven in part by differences in how explicitly semantic structure is encoded across modalities. Prompt is provided below. Write a CLEAR, EXPLICIT, and STRUCTURALLY GROUNDED description of this time series. GOAL: Reduce abstraction while keeping natural, grammatical English. Bind language directly to numeric and temporal structure. FORMAT REQUIREMENTS: - Natural English, full sentences 28 Time SeriesâVisionâLanguage Alignment - 4-5 sentences - Maximum 200 words total - Concise but explicit (do NOT be verbose) INCLUDE: 1. CONTEXT: What the data represents, including location, metric, and exact time span. 2. NUMERIC ANCHORS: Explicit start, end, minimum, and maximum values, each with time positions. 3. TEMPORAL STRUCTURE: Describe changes using explicit numeric phrasing, e.g.: - "increases monotonically from X to Y" - "drops from X to Y between year A and B" 4. PHASES: If the series naturally decomposes, describe up to 2-3 phases. Each phase must include a numeric value range and a time range. 5. PATTERNS: Use explicit terms ("monotonic increase", "single spike at midpoint") instead of abstract descriptors. CRITICAL RULES: - Every trend or pattern claim MUST reference specific numbers - Do NOT use qualitative adjectives (e.g., steady, sharp, notable, gradual) unless immediately quantified - Do NOT speculate, explain causes, or add interpretations - Do NOT add information not present in the data Table 7. Explicit semantic grounding improves text alignment at comparable ID. Comparison of original CaTS captions (train ID = 417.2) and structured captions (train ID = 405.8) for a representative trimodal configuration (DINOv2 + Qwen + MOMENT). All encoders and training settings are identical; only the captions differ. Caption TypePairCosine MarginâProcrustesâCKAâ CaTS (original)TSâTXT0.5380.8070.524 StructuredTSâTXT0.5740.7540.568 CaTS (original)IMGâTXT0.5580.7550.537 StructuredIMGâTXT0.5880.7140.583 N. ECGâText Alignment and Textual Specificity A substantial recent literature has explored aligning ECG waveforms with clinical text using contrastive and fine-grained supervision for downstream tasks, including MERL (Liu et al., 2024), MELP (Wang et al., 2025), FG-CLEP (Li et al., 2025), and ECG-Chat (Zhao et al., 2025). These works report strong zero-shot and linear-probe performance on downstream ECG tasks, often substantially exceeding unimodal baselines. While our study does not propose a new ECGâtext alignment architecture, our findings offer a representation-centric perspective that provides insight into why these methods are effective. A key distinction across this literature lies in how textual supervision is constructed and used. Methods such as MELP and FG-CLEP explicitly introduce fine-grained or structured supervision, including beat-level, rhythm-level, or entity-aware alignment, thereby creating localized correspondences between waveform segments and textual elements. In contrast, MERL relies on global contrastive alignment during training but achieves strong performance through carefully engineered, knowledge-augmented prompting at inference time, which increases the specificity and clinical grounding of textual queries. Despite differing mechanisms, these approaches can be viewed as sharing a common goal which is reducing the semantic gap between low-level waveform structure and high-level clinical language. Our results are consistent with this interpretation. We find that alignment involving indirect or high-level text is consistently weaker than alignment with visual projections of the signal. In ECG datasets, this manifests as strong TSâIMG alignment and retrieval performance, while TSâTXT and IMGâTXT alignment remains comparatively weaker, though stronger than in synthetic settings such as CaTS. From this perspective, prior ECGâtext methods can be understood as introducing mechanisms that effectively convert indirect text into more direct supervision. Fine-grained losses (e.g., beat- or token-level alignment), entity extraction, validated waveform descriptors, or knowledge-enhanced prompting all increase semantic explicitness and local grounding, thereby addressing the representational mismatch we identify. Our analysis provides a representation-level interpretation of the empirical gains reported by these methods. Importantly, our work complements these approaches by isolating textual specificity as a key factor governing alignment quality, and we provide a controlled analysis that helps clarify why ECGâtext alignment is fundamentally more challenging than ECGâimage alignment, and why additional structure or supervision is often required. This view offers guidance for future multimodal ECG systems, suggesting that improvements are more likely to arise from increasing semantic explicitness and local grounding than from scaling contrastive objectives alone. 29 Time SeriesâVisionâLanguage Alignment O. Full Results O.1. Results on varying ID text on CaTS models Table 8. TSâTXT alignment metrics across four caption distributions (1-phrase, 2-phrase, 4-phrase, and dense). Model Cosine MarginProcrustesMutual KNN 1-phrase2-phrase4-phrasedense1-phrase2-phrase4-phrasedense1-phrase2-phrase4-phrasedense eva18b + mistral12b + moment0.06300.13710.29400.66313.09681.98911.34280.57850.01650.03210.04360.1053 eva8b + mistral12b + timesfm0.07470.15900.31160.67952.69041.87691.30670.54010.01890.03010.04810.1164 dino7b + gemma27b + moirai0.00490.10140.34230.74323.35692.02661.23020.43700.00950.01880.04990.1451 eva 18b + gemma27b + chronos0.04670.13860.26840.68662.84631.85381.35910.51800.01250.02410.04190.1228 siglip + t5 + timesfm0.16140.27530.42290.65831.60721.40181.03050.58270.01680.03000.06070.1025 dinov2base + t5small + chronos0.13850.21830.36010.59661.69381.51431.13940.68250.01670.02290.04050.0787 siglip + e5 + moment0.19200.26120.39130.63621.61701.42341.09470.63090.01880.02810.05130.1053 vit base + qwen06b + moirai0.04000.13110.28830.55053.11282.08101.36820.80920.00930.01920.02630.0652 vitbase + t5small + moment0.16370.22750.35920.56921.66171.52621.18010.76040.01490.02890.04470.0664 Table 9. TXTâIMG alignment metrics across four caption distributions (1-phrase, 2-phrase, 4-phrase, and dense). Model Cosine MarginProcrustesMutual KNN 1-phrase2-phrase4-phrasedense1-phrase2-phrase4-phrasedense1-phrase2-phrase4-phrasedense eva18b + mistral12b + moment0.05810.14090.31320.72753.08611.94401.26600.46130.01770.02690.04950.1180 eva8b + mistral12b + timesfm0.06700.15570.32090.73302.70571.86881.26210.45170.01810.03230.05170.1165 dino7b + gemma27b + moirai0.01250.10850.32350.75813.18421.89411.19870.38740.01690.02280.05610.1804 eva18b + gemma27b + chronos0.05640.16050.31590.73182.79811.76601.22220.43310.01590.02760.05760.1409 siglip + t5 + timesfm0.16550.28880.48050.76271.58241.34880.88880.38220.02070.04090.09240.1780 dinov2 base + t5small + chronos0.13920.22380.38200.62371.69491.48071.08120.64510.01290.02150.04130.0643 siglip + e5 + moment0.19430.27430.42850.71501.59171.35860.97670.46530.02400.03400.07610.1617 vitbase + qwen06b + moirai0.06060.14790.29050.53982.90711.91511.26410.77500.01510.01760.03510.0607 vitbase + t5small + moment0.12790.20800.36920.61561.72341.52621.11920.67020.01210.02350.04470.0711 O.2. Results on CaTS test set Table 10. Correlation between alignment metrics and cross-modal retrieval performance. MutualkNN overlap shows the strongest and most consistent association with retrieval performance, particularly for TSâIMG pairs, with substantial correlations also observed for IMGâTXT and TSâTXT. Procrustes disparity exhibits strong negative correlations, indicating that improved global geometric alignment (lower disparity) corresponds to higher retrieval accuracy. Cosine similarity and CKA show weaker but still meaningful correlations, especially for harder TSâTXT pairs. Overall, higher alignment quality corresponds to improved cross-modal retrieval performance. MetricPair Pearson rSpearman Ï R@1R@5R@10R@1R@5R@10 Cosine Similarity TSâIMG0.83330.84890.85440.88340.88140.8762 TSâTXT0.31660.36910.39740.33480.36000.3780 IMGâTXT0.62770.67180.69600.64430.67530.6826 Procrustes Disparity TSâIMG-0.8360-0.8536-0.8604-0.9046-0.8972-0.8870 TSâTXT-0.3951-0.4512-0.4821-0.4654-0.4865-0.5007 IMGâTXT-0.6179-0.6612-0.6852-0.6432-0.6684-0.6734 CKA TSâIMG0.72170.70920.70140.74260.73470.7170 TSâTXT0.25850.31000.33730.28210.31000.3292 IMGâTXT0.50640.54330.56640.54100.55700.5595 Mutual kNN TSâIMG0.92820.93320.93300.94100.94360.9353 TSâTXT0.46180.51940.55250.53250.55390.5746 IMGâTXT0.83320.84950.85410.79330.81530.8225 While TSâIMG alignment dominates in the majority of configurations, we observe a small number of cases (3) where TSâTXT exceeds TSâIMG. Notably, all such exceptions involve the Moirai time series encoder, suggesting a systematic pattern rather than noise. In these configurations, Moirai-based representations exhibit stronger compatibility with text embeddings than with visual embeddings. Although the internal causes of this behavior are not directly observable, one 30 Time SeriesâVisionâLanguage Alignment possible interpretation is that Moirai encodings produce higher-level temporal abstractions that align more readily with symbolic language than with the fine-grained visual cues used for plots. Overall these cases highlight that alignment behavior is not identical across all time series encoders and can depend on architectural interactions in trimodal settings. Table 11. Geometric Analysis by Modality Pair across All Model Configurations. Model Configuration ProcrustesCKAMutual KNN TSâIMG TSâTXT IMGâTXTTSâIMG TSâTXT IMGâTXTTSâIMG TSâTXT IMGâTXT Lower Bound Models dinov2-b + t5small + moirai-b0.76830.63430.71720.65020.64670.59260.04150.07730.0479 vit-b + t5small + moment-b0.55160.76040.67020.73260.55540.62610.10170.06640.0711 dinov2-b + t5 small + chronos-b0.49650.68250.64510.74990.58260.64400.09690.07870.0643 dinov2-b + e5-base + timesfm-b0.65770.76480.66050.63390.54230.63850.07530.06320.0664 vit-b + e5-b + chronos-b0.47690.72770.66660.76620.54290.63220.10650.07230.0659 siglip2-b + e5-b + moirai-b0.52340.75200.44090.75800.56500.71380.10640.06450.1145 siglip2-b + e5-b + moment-b0.54440.75460.45890.71610.53290.71120.11890.06510.1125 siglip2-b + t5small + timesfm-b0.63450.77430.45000.65690.55350.72490.09530.06920.1276 vit-b + qwen06b + moirai-b0.68520.80920.77500.69700.55050.52620.05490.06520.0607 dinov2-b + qwen06b + moment-b0.54370.80660.75460.71430.52360.53680.09000.07130.0616 vit-b + qwen06b + timesfm-b0.63210.82240.77510.65390.52480.54290.08770.07410.0625 siglip2-b + qwen06b + chronos-b0.45780.74990.51040.76250.52240.66840.14480.07730.1312 Middle Bound Models vith14 + t5 + moment-l0.50890.60870.57750.70390.63450.63510.12390.09480.0893 siglip + t5 + timesfm-l0.39990.58270.38220.77580.64930.74180.18950.10250.1780 dinov2 + t5 + moirai-l 0.72390.53070.65530.65750.71490.62710.04550.11440.0643 dinov2 + t5 + chronos-l 0.45850.60720.58000.76350.64150.67100.12000.08680.0815 vith14 + qwen + moirai-l0.66610.67650.69470.66640.62320.57290.06310.08130.0697 vith14 + qwen + timesfm-l0.46200.62640.61850.75180.63520.63670.12990.10880.0905 dinov2 + qwen + moment-l0.49820.69570.68250.70530.57870.57340.11790.09760.0777 siglip + qwen + chronos-l 0.39560.65730.44910.77700.61300.70700.18680.09270.1639 siglip + e5 + moirai-l0.38630.56250.42350.80220.67470.71380.16430.12200.1855 siglip + e5 + moment-l0.39990.63090.46530.76380.63110.68190.17770.10530.1617 vit h14 + e5 + chronos-l0.43200.62370.60080.76790.61250.61330.13520.09650.1020 dinov2 + e5 + timesfm-l0.46940.59290.63380.72990.63750.60640.12790.10510.0881 Upper Bound Models dino7b + qwen8b + moment-l0.33240.66710.53170.79610.59460.63620.22210.09480.1413 eva8b + qwen8b + chronos-l0.49130.63200.51520.71160.64010.67140.11790.08760.1160 dino7b + mistral12b + chronos-l0.32310.59620.41910.83980.63150.71540.21010.11110.1737 eva8b + mistral12b + moirai-l0.64500.51340.51380.66620.72900.66410.06520.12750.1168 eva8b + mistral12b + timesfm-l0.47990.54010.45170.72130.67070.70260.13160.11640.1165 eva18b + qwen8b + moirai-l0.61330.66750.57220.68510.62580.62630.07430.08400.1053 eva18b + mistral12b + moment-l0.49160.57850.46130.71320.66140.69490.11870.10530.1180 dino 7b + gemma27b + moirai-l0.41120.43700.38740.78390.74780.73470.14240.14510.1804 eva8b + gemma27b + moment-l0.51400.53470.47160.68960.68850.67990.11600.11960.1235 eva 18b + gemma27b + chronos-l0.48030.51800.43310.71780.70830.72670.13160.12280.1409 31 Time SeriesâVisionâLanguage Alignment Table 12. Cross-Modal Evaluation: Cosine Similarity Margins and Retrieval Performance Model Configuration Cosine Similarity MarginsCross-Modal Retrieval TSâIMG TSâTXT IMGâTXTR@1R@5R@10 MRR Lower Bound Models dinov2-b + t5small + moirai-b0.58350.64780.59020.0081 0.0358 0.0689 0.0336 vit-b + t5small + moment-b0.68430.56920.61560.0197 0.0763 0.1279 0.0597 dinov2-b + t5 small + chronos-b0.70820.59660.62370.0239 0.0846 0.1400 0.0660 dinov2-b + e5-base + timesfm-b0.62730.55300.62750.0122 0.0501 0.0928 0.0433 vit-b + e5-b + chronos-b 0.72000.57290.61690.0218 0.0825 0.1365 0.0633 siglip2-b + e5-b + moirai-b0.70500.57020.73280.0160 0.0640 0.1135 0.0535 siglip2-b + e5-b + moment-b0.67340.54840.72050.0238 0.0868 0.1440 0.0677 siglip2-b + t5 small + timesfm-b0.62920.54450.72360.0212 0.0816 0.1392 0.0648 vit-b + qwen06b + moirai-b0.62780.55050.53980.0097 0.0408 0.0725 0.0354 dinov2-b + qwen06b + moment-b0.68950.53830.55760.0161 0.0613 0.1075 0.0503 vit-b + qwen06b + timesfm-b0.63790.51620.54670.0137 0.0573 0.1021 0.0477 siglip2-b + qwen06b + chronos-b0.71470.54320.67580.0396 0.1264 0.1985 0.0940 Middle Bound Models vith14 + t5 + moment-l0.70650.65210.67010.0283 0.1026 0.1692 0.0784 siglip + t5 + timesfm-l0.75820.65830.76270.0588 0.1875 0.2851 0.1338 dinov2 + t5 + moirai-l0.60440.70150.62490.0117 0.0500 0.0919 0.0437 dinov2 + t5 + chronos-l 0.72700.63970.65790.0311 0.1052 0.1720 0.0800 vit h14 + qwen + moirai-l0.63670.62250.59370.0112 0.0522 0.0930 0.0436 vith14 + qwen + timesfm-l0.72950.62880.63150.0315 0.1094 0.1779 0.0826 dinov2 + qwen + moment-l0.71120.60090.60260.0225 0.0865 0.1471 0.0668 siglip + qwen + chronos-l0.74980.60820.71340.0695 0.1984 0.2865 0.1421 siglip + e5 + moirai-l0.76940.67810.73370.0419 0.1435 0.2299 0.1064 siglip + e5 + moment-l0.75620.63620.71500.0565 0.1784 0.2699 0.1277 vit h14 + e5 + chronos-l0.74480.64030.64920.0423 0.1364 0.2114 0.1003 dinov2 + e5 + timesfm-l 0.72870.65470.63220.0328 0.1183 0.1916 0.0877 Upper Bound Models dino7b + qwen8b + moment-l0.79590.62290.69120.0619 0.1835 0.2712 0.1323 eva8b + qwen8b + chronos-l0.70650.62510.68910.0340 0.1228 0.1943 0.0896 dino7b + mistral12b + chronos-l0.79780.64910.74170.08050.21720.31120.1571 eva 8b + mistral12b + moirai-l0.64640.70800.69730.0189 0.0763 0.1351 0.0625 eva8b + mistral12b + timesfm-l0.72150.67950.73300.0357 0.1267 0.2073 0.0945 eva 18b + qwen8b + moirai-l0.66190.62880.66520.0161 0.0670 0.1175 0.0548 eva18b + mistral12b + moment-l0.71230.66310.72750.0335 0.1222 0.1963 0.0902 dino7b + gemma27b + moirai-l0.76050.74320.75810.0400 0.1455 0.2393 0.1077 eva 8b + gemma27b + moment-l0.70030.68480.71830.0332 0.1259 0.2067 0.0923 eva18b + gemma27b + chronos-l0.71160.68660.73180.0485 0.1579 0.2429 0.1139 32