Paper deep dive
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
Xiaocheng Fang, Jieyi Cai, Guangkun Nie, Haoyu Wang, Jiarui Jin, Yujie Xiao, Bo Liu, Chenyang He, Qinghao Zhao, Gaofeng Cheng, Hongyan Li, Shenda Hong
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECG--text alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary Shared--Private Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns normalized shared projections. APBC organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion. We evaluate EchoBridge on EchoNext-Mini and independent PKUPH and SHTMU cohorts under four protocols: prompt-based inference without downstream classifier training, in-domain frozen linear probing, target-domain cross-center frozen linear probing, and source-only cross-center transfer, supplemented by finding-specific analyses. EchoBridge improves classifier-free AUROC, AUPRC, and F1 over the strongest baselines by 7.88, 5.61, and 4.54 points, respectively, and achieves the highest point estimates across all in-domain and target-domain probing budgets and both source-only transfer cohorts. Finding-specific analyses show gains for most conditions, including several low-prevalence valvular findings.
Tags
Links
- Source: https://arxiv.org/abs/2607.24553v1
- Canonical: https://arxiv.org/abs/2607.24553v1
Trouble viewing inline? Open PDF directly â
Full Text
95,136 characters extracted from source content.
Expand or collapse full text
EchoBridge: Long-Tail-Aware ECGâEchocardiography Text Alignment for Echocardiography-Derived Cardiac Findings Xiaocheng Fang School of Intelligence Science and Technology, Peking UniversityBeijingChina State Key Laboratory of General Artificial Intelligence, Peking UniversityBeijingChina fangxiaocheng26@stu.pku.edu.cn , Jieyi Cai University of the Chinese Academy of SciencesBeijingChina caijieyi055@gmail.com , Guangkun Nie School of Intelligence Science and Technology, Peking UniversityBeijingChina nieguangkun@stu.pku.edu.cn , Haoyu Wang University of the Chinese Academy of SciencesBeijingChina wanghaoyu252@mails.ucas.ac.cn , Jiarui Jin School of Intelligence Science and Technology, Peking UniversityBeijingChina jrjin25@stu.pku.edu.cn , Yujie Xiao National Institute of Health Data Science, Peking UniversityBeijingChina xiaoyujie@stu.pku.edu.cn , Bo Liu School of Intelligence Science and Technology, Peking UniversityBeijingChina liubo2022@stu.pku.edu.cn , Chenyang He National Institute of Health Data Science, Peking UniversityBeijingChina raoquan00@gmail.com , Qinghao Zhao Department of Cardiology, Peking University Peopleâs HospitalBeijingChina qhzhao@pku.edu.cn , Gaofeng Cheng University of the Chinese Academy of SciencesBeijingChina chenggaofeng@hccl.ioa.ac.cn , Hongyan Li School of Intelligence Science and Technology, Peking UniversityBeijingChina leehy@pku.edu.cn and Shenda Hong National Institute of Health Data Science, Peking UniversityBeijingChina hongshenda@pku.edu.cn (2018) Abstract. Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardiac findings. Global ECGâtext alignment may entangle modality-specific factors, while long-tailed finding distributions provide sparse positive supervision for low-prevalence conditions. We propose EchoBridge with Complementary SharedâPrivate Projection (CSPP) and Adaptive Prototype Boundary Calibration (APBC). CSPP maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns normalized shared projections. APBC organizes the shared hypersphere with class-specific prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion. We evaluate EchoBridge on EchoNext-Mini and independent PKUPH and SHTMU cohorts under four protocols: prompt-based inference without downstream classifier training, in-domain frozen linear probing, target-domain cross-center frozen linear probing, and source-only cross-center transfer, supplemented by finding-specific analyses. EchoBridge improves classifier-free AUROC, AUPRC, and F1 over the strongest baselines by 7.88, 5.61, and 4.54 points, respectively, and achieves the highest point estimates across all in-domain and target-domain probing budgets and both source-only transfer cohorts. Finding-specific analyses show gains for most conditions, including several low-prevalence valvular findings. Code: https://github.com/PKUDigitalHealth/EchoBridge. â copyright: acmlicensedâ journalyear: 2018â doi: X.Xâ conference: The 33rd ACM SIGKDD Conference on Knowledge Discovery and Data Mining; August, 2027; San Jose, CA, USAâ isbn: 978-1-4503-X-X/2018/06â ccs: Applied computing Health informaticsâ ccs: Computing methodologies Artificial intelligence Figure 1. Two challenges in ECGâechocardiography text alignment. Left: Global alignment entangles modality-specific factors and weakens clinically relevant cross-modal correspondence. Right: Prevalence imbalance provides fewer positive constraints and less reliable sampleâprototype organization for low-prevalence findings. 1. Introduction Echocardiography is a central imaging modality for evaluating cardiac structure, function, and valvular abnormalities (Ghorbani et al., 2020; Ouyang et al., 2020; Christensen et al., 2024). Direct ECGâechocardiography learning from raw images or videos can be difficult to conduct at scale in retrospective multicenter studies because of storage, accessibility, and data-governance constraints (Kaissis et al., 2020; Vukadinovic et al., 2026). In routine practice, echocardiographic examinations are accompanied by diagnostic reports containing detailed findings and a concise conclusion that aggregates the principal structural, functional, and valvular abnormalities (Chao et al., 2025; Kwak et al., 2025). We use these conclusion-level summaries as cross-modal supervision because they preserve clinically salient finding compositions while reducing measurement-, acquisition-, and template-specific variation. When paired with temporally matched ECG recordings, these summaries connect echocardiography-derived clinical semantics with cardiac electrical activity and provide a practical source of cross-modal supervision. Recent AIâECG studies show that ECGs encode echocardiography-confirmed abnormalities, including reduced ejection fraction and broader structural heart disease (Attia et al., 2019; Yao et al., 2021; Poterucha et al., 2025). MERL and ECG-CLIP use natural-language supervision for transferable representation learning (Liu et al., 2024b; Zhou et al., 2025), while MERL-ECHO and Wearable-Echo-FM align ECGs with echocardiography reports for structural cardiac findings (Wong et al., 2025; Knight et al., 2026). Building on these methods, we learn a class-aware normalized space from paired ECGs and standardized conclusions, capturing paired-sample ECGâtext alignment and organizing both modalities by predefined findings; evaluation covers prompt-based classifier-free inference for pretraining-seen findings, in-domain and target-domain cross-center frozen linear probing across label budgets, and source-only cross-center transfer. Technically, aligning ECGs with standardized echocardiography conclusions poses two challenges. 1) Cross-modal representation interference. ECG representations may encode continuous variation in rhythm, conduction, waveform morphology, and acquisition-related factors, whereas standardized conclusions provide compact compositional semantics for predefined structural, functional, and valvular findings. Global alignment may force heterogeneous ECG factors into a single conclusion-oriented embedding space, reducing the separation between finding-relevant and modality-dependent variation. 2) Long-tailed representation bias. Echocardiography-derived findings often have highly imbalanced prevalence. Frequent findings contribute more positive samples and repeated optimization signals, potentially biasing the space toward head classes. Low-prevalence findings receive fewer positive constraints, reducing intra-class compactness and prototype separation and yielding less reliable decision boundaries and uneven representation quality across prevalence levels. To address these challenges, we propose EchoBridge, an ECGâechocardiography text alignment framework for learning transferable ECG representations of echocardiography-derived findings. EchoBridge comprises two components. First, Complementary SharedâPrivate Projection maps ECG and text representations into alignment-oriented shared and auxiliary private projections. Cross-modal alignment and class-prototype supervision operate on the shared projections, while within-modality orthogonality reduces directional redundancy between branches and a symmetric contrastive objective aligns the â2 _2-normalized shared representations. Second, Adaptive Prototype Boundary Calibration organizes the shared hypersphere using class-specific prototypes for predefined findings. Training-frequency-adaptive angular margins strengthen positive sampleâprototype constraints for low-prevalence findings, while spherical Riesz repulsion discourages prototype concentration. Together, these components integrate instance-level ECGâtext correspondence with class-aware geometric organization under imbalanced finding distributions. We evaluate EchoBridge on EchoNext-Mini and two independent hospital cohorts from Peking University Peopleâs Hospital (PKUPH) and the Second Hospital of Tianjin Medical University (SHTMU) under four protocols: prompt-based classifier-free inference for pretraining-seen findings, in-domain frozen linear probing across multiple label budgets, target-domain cross-center linear probing, and source-only cross-center transfer, complemented by finding-specific analyses. The main contributions are as follows: âą We propose EchoBridge, an ECGâechocardiography text alignment framework for transferable ECG representations of echocardiography-derived findings under cross-modal interference and long-tailed distributions. âą We develop Complementary SharedâPrivate Projection, which maps each modality into shared and auxiliary private projections, reduces directional redundancy via within-modality orthogonality, and bidirectionally aligns the normalized shared space. âą We propose Adaptive Prototype Boundary Calibration, which structures the shared hypersphere with cross-modal class prototypes, training-frequency-adaptive angular margins, and spherical Riesz repulsion to improve class compactness and prototype separation under imbalance. âą We evaluate EchoBridge on EchoNext-Mini and two independent hospital cohorts using prompt-based classifier-free inference, in-domain and target-domain cross-center frozen linear probing across label budgets, source-only cross-center transfer, and finding-specific analyses. Figure 2. Overview of EchoBridge. Given paired ECGs and echocardiography summaries, (A) Complementary SharedâPrivate Projection maps each modality into alignment-oriented shared and auxiliary private branches, with within-modality orthogonality reducing directional redundancy between them. (B) Adaptive Prototype Boundary Calibration structures the shared space using class-specific prototypes, frequency-adaptive angular margins, and spherical Riesz repulsion. 2. Methodology 2.1. Overview of EchoBridge EchoBridge learns transferable echocardiography-related ECG representations from paired 12-lead ECGs and standardized echocardiography conclusions. As shown in Figure 2, it combines Complementary SharedâPrivate Projection and Adaptive Prototype Boundary Calibration through normalized bidirectional alignment. Given an ECG xix_i and paired summary tit_i, encoders feâ(â )f_e(·) and ftâ(â )f_t(·) produce representations he,ih_e,i and ht,ih_t,i, which independent heads map into shared and auxiliary private projections. Within-modality orthogonality reduces directional redundancy between branches, while symmetric contrastive learning aligns the â2 _2-normalized shared representations by increasing paired-sample similarity over unpaired samples. Adaptive Prototype Boundary Calibration structures the shared space with class-specific prototypes for predefined cardiac findings. Training-frequency-adaptive angular margins strengthen positive alignment for low-prevalence findings, while spherical Riesz repulsion discourages prototype concentration and promotes pairwise separation on the unit hypersphere. 2.2. Complementary SharedâPrivate Projection ECG waveforms encode rich continuous variation in rhythm, conduction, morphology, and acquisition conditions, whereas standardized echocardiography conclusions provide compact compositional semantics for predefined cardiac findings. Their global representations therefore contain factors with different relevance to cross-modal correspondence. Direct global alignment may entangle finding-relevant and modality-specific information within a single embedding space. We propose Complementary SharedâPrivate Projection, which maps each modality into alignment-oriented shared and auxiliary private projections. Cross-modal alignment and class-prototype supervision operate on the shared projections, while within-modality orthogonality reduces sample-wise directional redundancy between the two branches. The terms shared and private specify their optimization roles, while their detailed semantic contents remain unconstrained. SharedâPrivate Projection. Given an ECG xix_i and paired echocardiography conclusion tit_i, the ECG and text encoders produce global representations: (1) he,i=feâ(xi),ht,i=ftâ(ti).h_e,i=f_e(x_i), h_t,i=f_t(t_i). Independent projection heads map each representation into shared and private projections: (2) ze,is=Ïesâ(he,i),ze,ip=Ïepâ(he,i),z^s_e,i=Ï^s_e(h_e,i), z^p_e,i=Ï^p_e(h_e,i), (3) zt,is=Ïtsâ(ht,i),zt,ip=Ïtpâ(ht,i).z^s_t,i=Ï^s_t(h_t,i), z^p_t,i=Ï^p_t(h_t,i). Cross-modal alignment and class-prototype supervision operate on the shared projections ze,isz^s_e,i and zt,isz^s_t,i. The independently parameterized private projections ze,ipz^p_e,i and zt,ipz^p_t,i provide auxiliary capacity without direct cross-modal alignment, allowing the shared branches to emphasize ECGâtext correspondence. To reduce sample-wise directional redundancy between the shared and private projections, we introduce a within-modality orthogonality objective: (4) âorth=1Nââi=1N[âšzÂŻe,is,zÂŻe,ipâ©2+âšzÂŻt,is,zÂŻt,ipâ©2],L_orth= 1N _i=1^N [ z^s_e,i, z^p_e,i ^2+ z^s_t,i, z^p_t,i ^2 ], where N is the batch size, zÂŻ=z/âzâ2 z=z/\|z\|_2 denotes an â2 _2-normalized representation, and âšâ ,â ⩠·,· denotes the inner product. Because the normalized inner product corresponds to cosine similarity, minimizing âorthL_orth drives the shared and private projections toward orthogonality within each modality. This geometric regularization reduces collinearity while leaving the semantic content of the shared and private projections unconstrained. Shared-Space Alignment. The shared ECG and text projections are normalized onto the unit hypersphere: (5) z^e,is=ze,isâze,isâ2,z^t,is=zt,isâzt,isâ2. z^s_e,i= z^s_e,i\|z^s_e,i\|_2, z^s_t,i= z^s_t,i\|z^s_t,i\|_2. For a mini-batch of N paired samples, the cross-modal similarity matrix is defined as: (6) Siâj=âšz^e,is,z^t,jsâ©Ï,S_ij= z^s_e,i, z^s_t,j Ï, where Ï is the temperature parameter. Because both representations are â2 _2-normalized, their inner product equals cosine similarity. We optimize bidirectional ECGâtext correspondence using a symmetric contrastive objective: (7) âalign=12âNââi=1N[âlogâĄexpâĄ(Siâi)âj=1NexpâĄ(Siâj)âlogâĄexpâĄ(Siâi)âj=1NexpâĄ(Sjâi)].L_align= 12N _i=1^N [- (S_i) _j=1^N (S_ij)- (S_i) _j=1^N (S_ji) ]. The first term retrieves the paired echocardiography conclusion from an ECG query, while the second performs reverse retrieval. This objective forms a common normalized space for paired ECGâtext correspondence, which the following class-prototype objective organizes by echocardiography-derived findings. 2.3. Adaptive Prototype Boundary Calibration Pairwise ECGâtext alignment provides instance-level supervision, whereas fine-grained cardiac findings require explicit class-level organization. Long-tailed distributions may produce weak boundaries for infrequent classes, while unconstrained prototypes may cluster geometrically. We therefore propose Adaptive Prototype Boundary Calibration (APBC), combining frequency-adaptive angular margins calibrated by training-set label frequencies with spherical repulsion that separates normalized class prototypes. We maintain a learnable prototype matrix: (8) P=[p1,âŠ,pC]â€ââCĂd,P=[p_1,âŠ,p_C] ^CĂ d, where C denotes the number of fine-grained cardiac findings and d the shared-space dimension. Each prototype is normalized as: (9) p^c=pcâpcâ2. p_c= p_c\|p_c\|_2. For normalized shared ECG and text representations z^e,is z^s_e,i and z^t,is z^s_t,i, the prototype similarities are: (10) ui,ce=âšz^e,is,p^câ©,ui,ct=âšz^t,is,p^câ©.u^e_i,c= z^s_e,i, p_c , u^t_i,c= z^s_t,i, p_c . The corresponding logits are: (11) âi,ce=Îłâui,ce,âi,ct=Îłâui,ct, ^e_i,c=Îł u^e_i,c, ^t_i,c=Îł u^t_i,c, where Îł=expâĄ(s)Îł= (s) is a learnable positive scale. Sharing P across modalities provides a common category-level reference in the normalized space. Frequency-Adaptive Angular Margin. To address label-frequency imbalance, we assign each cardiac finding a class-specific angular margin based on its training-set positive rate. Let rcr_c denote the positive rate of class c. Its margin is: (12) mc=clipâ(m0â medianâ(r)rc+Ï”â 1Îș,mmin,mmax),m_c=clip (m_0· median(r)r_c+Δ· 1Îș,m_ ,m_ ), where m0m_0 is the base margin, ϔΔ ensures numerical stability, and mminm_ and mmaxm_ bound the margin. The data-dependent normalization factor is: (13) Îș=1Cââc=1Cmedianâ(r)rc+Ï”,Îș= 1C _c=1^C median(r)r_c+Δ, which normalizes the mean pre-clipping margin to m0m_0. For each sampleâclass pair, the adjusted logits are: (14) â~i,ce=ÎłâcosâĄ(Ξi,ce+yi,câmc),â~i,ct=ÎłâcosâĄ(Ξi,ct+yi,câmc), ^e_i,c=Îł (Ξ^e_i,c+y_i,cm_c ), ^t_i,c=Îł (Ξ^t_i,c+y_i,cm_c ), where Ξi,ce=arccosâĄ(ui,ce)Ξ^e_i,c= (u^e_i,c ) and Ξi,ct=arccosâĄ(ui,ct)Ξ^t_i,c= (u^t_i,c ) denote the ECGâprototype and textâprototype angles, respectively, and yi,câ0,1y_i,câ\0,1\ indicates whether sample i is positive for class c. Thus, the margin applies only to positive sampleâprototype pairs, leaving negative-pair logits unchanged. Lower-prevalence findings receive larger bounded margins, encouraging more discriminative representations around tail-class prototypes. The adjusted logits supervise the shared ECG and text branches: (15) âproto=BCEâ(â~e,y)+BCEâ(â~t,y),L_proto=BCE ( ^e,y )+BCE ( ^t,y ), where y is the multi-label target matrix, and â~e ^e and â~t ^t are the adjusted ECG and text logits. This objective draws positive representations toward corresponding prototypes and suppresses responses to prototypes of negative findings, linking instance-level ECGâtext alignment with class-level multi-label supervision. Figure 3. Adaptive Prototype Boundary Calibration is designed to improve positive-class compactness and prototype separation on the hypersphere. Spherical Riesz Repulsion. Prototype supervision constrains sampleâprototype relationships, while class prototypes may still cluster on the hypersphere. We therefore apply a spherical Riesz repulsion loss to the normalized prototypes: (16) âriesz=2Câ(Câ1)ââ1â€a<bâ€Câp^aâp^bâ2âq,q>0.L_riesz= 2C(C-1) _1†a<b†C \| p_a- p_b \|_2^-q, q>0. Because the loss increases as pairwise distance decreases, minimizing ârieszL_riesz discourages prototype concentration and promotes separation on the unit hypersphere. The Adaptive Prototype Boundary Calibration objective is: (17) âapbc=âproto+λrââriesz,L_apbc=L_proto+ _rL_riesz, where λr _r controls the contribution of spherical prototype repulsion. 2.4. Overall Training Objective EchoBridge jointly optimizes Complementary SharedâPrivate Projection and Adaptive Prototype Boundary Calibration with the objective: (18) â=âalign+âorth+âapbc.L=L_align+L_orth+L_apbc. 3. Experiments 3.1. Datasets and Splits We evaluate EchoBridge on the open-source EchoNext-Mini (Hughes et al., 2026) and two private real-world cohorts from Peking University Peopleâs Hospital (PKUPH) and the Second Hospital of Tianjin Medical University (SHTMU). Each sample pairs an ECG with echocardiography-derived findings from a temporally matched examination under a dataset-specific matching window. We partition each dataset into patient-disjoint training, validation, and test sets using a 7:1:2 ratio to reduce information leakage. All datasets exhibit long-tailed label distributions; Figure 4 reports their statistics, label distributions, and split sizes. Figure 4. Dataset statistics and label distributions across EchoNext-Mini, PKUPH, and SHTMU. Each panel also reports the ECG cohort size and the train/validation/test split. 3.2. Data Preprocessing ECG Signal Processing. We applied record-level quality control and excluded unreadable ECGs, records with substantial missingness, and samples that could not be reliably linked to patient identifiers or labels. All signals were resampled to 500 Hz by linear interpolation and processed using a fixed denoising pipeline comprising a 0.5-Hz high-pass filter, a second-order Butterworth low-pass filter with a 50-Hz cutoff, and a 50/60-Hz notch filter. Recordings were standardized to 10-second segments, with longer recordings divided into consecutive temporal windows. Each segment was Z-score normalized, and missing leads or signal values were zero-filled to preserve a consistent input shape. Echocardiography Conclusion-Style Text Construction. Clinical echocardiography reports typically contain detailed findings and a concise conclusion or impression. Findings may include quantitative measurements, image-quality descriptions, and contextual observations with variable relevance to predefined cardiac findings, limiting their consistency as supervision for finding-specific representation learning. We therefore use conclusion-level semantics. For each examination, we deterministically verbalize the structured multi-label finding vector into a standardized conclusion-style summary organized by chamber structure, myocardial function, and valvular abnormalities. This preserves multi-label co-occurrence, mirrors the concise compositional form of clinical conclusions, and controls unrelated lexical, formatting, and template variation. As an informal quality check, five senior cardiologists inspected a random subset of the generated conclusion-style summaries and provided qualitative feedback that the sampled summaries were broadly clinically plausible and semantically consistent with the corresponding structured findings. This review was not designed as a formal annotation or inter-rater agreement study. The resulting text supports ECGâtext alignment and class-specific prototype learning under controlled conclusion-level semantics. 3.3. Baselines and Implementation Details We compare EchoBridge with representative ECG-only self-supervised and ECGâtext pretraining methods. Prompt-based classifier-free inference uses ECGâtext dual encoders with fixed definition-only textual prototypes for pretraining-seen findings. In all linear-probing experiments, the pretrained ECG representation extractor is frozen, and an identical linear multi-label classifier is trained with 1%, 10%, or 100% of the available labels. All methods share patient-level partitions, ECG preprocessing, label subsets, metrics, and validation-based threshold selection. We use official implementations when available; otherwise, we reproduce methods under matched backbone, optimization, and training-budget settings where applicable. Appendix F.4 further compares EchoBridge with task-specific ResNet-18 classifiers trained end-to-end on EchoNext-Mini. EchoBridge was implemented in PyTorch and trained on one NVIDIA A100 GPU using a one-dimensional ResNet-18 ECG encoder and MedCPT-Article-Encoder text encoder (Jin et al., 2023). AdamW optimization uses a learning rate of 1Ă10â51Ă 10^-5, weight decay of 1Ă10â81Ă 10^-8, batch size 64, and up to 15 epochs. Further architecture and optimization details are provided in the appendix. Table 1. Prompt-based classifier-free inference performance on EchoNext-Mini for pretraining-seen cardiac findings. Methods Ref. AUROC AUPRC F1 CLIP (Radford et al., 2021) ICMLâ21 62.57 [61.70, 63.41] 17.03 [16.42, 17.82] 24.13 [23.49, 25.17] SigLIP (Zhai et al., 2023) ICCVâ23 65.07 [64.23, 65.91] 17.35 [16.73, 18.14] 25.00 [24.36, 26.04] PCME++ (Chun, 2023) ICLRâ24 62.02 [61.14, 62.87] 17.59 [16.95, 18.38] 23.07 [22.39, 24.06] MERL-ECHO (Wong et al., 2025) medRxivâ25 64.02 [63.17, 64.86] 18.63 [17.98, 19.48] 23.92 [23.26, 24.87] ECG-CLIP (Zhou et al., 2025) npj DMâ25 60.47 [59.55, 61.36] 15.47 [14.87, 16.22] 22.91 [22.28, 23.89] D-BETA (Hung et al., 2025) ICMLâ25 63.99 [63.15, 64.83] 20.12 [19.46, 20.96] 27.29 [26.58, 28.36] SGERA (Chen et al., 2026) ICMLâ26 67.85 [67.03, 68.67] 21.18 [20.51, 22.03] 26.95 [26.26, 27.99] EchoBridge Ours 75.73 [75.00, 76.54] 26.79 [25.95, 27.98] 31.83 [31.15, 33.30] 3.4. Evaluation Protocols and Metrics We evaluate pretrained ECG representations under four protocols. Prompt-based classifier-free inference predicts pretraining-seen findings by comparing frozen ECG representations with fixed definition-only textual prototypes. In-domain frozen linear probing trains a linear multi-label classifier on EchoNext-Mini using 1%, 10%, or 100% of the available labels. Target-domain cross-center probing freezes the EchoNext-Mini-pretrained extractor and trains a new linear classifier on PKUPH or SHTMU with the same label ratios. Source-only transfer trains the classifier and selects thresholds exclusively on EchoNext-Mini, then evaluates findings shared with each external cohort. All experiments use patient-disjoint test sets and report macro-averaged AUROC, AUPRC, and F1 as percentage values. Class-specific thresholds maximize validation-set F1 and remain fixed for testing; source-only transfer retains EchoNext-Mini thresholds. We estimate 95% confidence intervals using 1,000 non-parametric bootstrap repetitions. Appendix E.1 provides detailed protocols and evaluation procedures. Table 2. Frozen linear-probing performance on EchoNext-Mini using 1%, 10%, and 100% of the available training labels. Each entry reports the macro-averaged metric value with its 95% bootstrap confidence interval in brackets. Methods Ref. 1% Linear Probing 10% Linear Probing 100% Linear Probing AUROC AUPRC F1 AUROC AUPRC F1 AUROC AUPRC F1 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 60.38 [59.42, 61.39] 13.78 [13.20, 14.58] 20.55 [19.85, 21.62] 62.19 [61.33, 63.10] 16.81 [16.15, 17.75] 23.17 [22.43, 24.37] 68.70 [67.92, 69.53] 19.96 [19.10, 21.06] 25.65 [24.85, 27.02] ST-MEM (Na et al., 2024) ICLRâ24 62.44 [61.41, 63.40] 17.09 [16.47, 17.86] 22.74 [22.00, 23.78] 68.42 [67.49, 69.28] 19.92 [19.22, 20.83] 25.92 [25.14, 27.09] 71.16 [70.31, 71.94] 21.72 [20.82, 22.79] 27.57 [26.73, 28.91] HeartLang (Jin et al., 2025) ICLRâ25 60.15 [59.17, 61.18] 16.98 [16.39, 17.79] 21.86 [21.15, 22.94] 70.78 [69.90, 71.71] 22.58 [21.91, 23.53] 27.41 [26.66, 28.62] 74.59 [73.79, 75.44] 25.37 [24.50, 26.48] 30.69 [29.88, 32.07] ECGâText Pretraining CLIP (Radford et al., 2021) ICMLâ21 60.11 [59.06, 61.09] 14.30 [13.67, 15.08] 22.12 [21.37, 23.17] 64.99 [64.04, 65.87] 18.97 [18.26, 19.89] 23.65 [22.86, 24.83] 70.46 [69.59, 71.26] 22.14 [21.23, 23.22] 26.86 [26.01, 28.21] SigLIP (Zhai et al., 2023) ICCVâ23 61.59 [60.60, 62.53] 15.35 [14.76, 16.11] 22.63 [21.92, 23.66] 66.24 [65.35, 67.08] 19.43 [18.76, 20.33] 23.91 [23.16, 25.07] 70.97 [70.16, 71.73] 21.87 [21.00, 22.93] 27.11 [26.30, 28.44] PCME++ (Chun, 2023) ICLRâ24 57.08 [56.04, 58.12] 12.74 [12.12, 13.56] 20.04 [19.30, 21.13] 63.04 [62.10, 63.98] 17.62 [16.92, 18.58] 22.77 [21.99, 23.99] 69.98 [69.12, 70.84] 20.95 [20.05, 22.07] 26.09 [25.25, 27.48] MERL-ECHO (Wong et al., 2025) medRxivâ25 61.90 [60.95, 62.87] 16.04 [15.47, 16.81] 22.80 [22.11, 23.84] 66.36 [65.51, 67.23] 19.40 [18.75, 20.31] 25.01 [24.28, 26.18] 71.90 [71.13, 72.69] 23.32 [22.47, 24.39] 29.16 [28.37, 30.50] ECG-CLIP (Zhou et al., 2025) npj DMâ25 61.56 [60.54, 62.58] 17.71 [17.10, 18.51] 23.10 [22.37, 24.17] 69.57 [68.65, 70.49] 21.80 [21.11, 22.74] 27.01 [26.24, 28.21] 73.89 [73.05, 74.73] 24.27 [23.38, 25.37] 29.86 [29.03, 31.23] D-BETA (Hung et al., 2025) ICMLâ25 68.99 [67.99, 69.94] 23.02 [22.42, 23.78] 28.08 [27.36, 29.11] 75.93 [75.03, 76.78] 26.84 [26.16, 27.74] 30.41 [29.65, 31.57] 76.43 [75.61, 77.20] 29.13 [28.25, 30.19] 32.82 [32.00, 34.15] SGERA (Chen et al., 2026) ICMLâ26 68.76 [67.79, 69.76] 24.49 [23.91, 25.28] 27.73 [27.03, 28.79] 74.97 [74.10, 75.87] 27.69 [27.03, 28.62] 31.74 [31.00, 32.93] 77.19 [76.40, 78.01] 30.25 [29.39, 31.34] 33.57 [32.77, 34.93] EchoBridge Ours 72.77 [71.84, 73.79] 26.53 [25.67, 27.56] 31.32 [30.49, 32.59] 76.94 [76.10, 77.84] 30.45 [29.55, 31.52] 34.13 [33.18, 35.52] 78.79 [77.98, 79.55] 32.66 [31.65, 33.88] 35.82 [35.02, 37.36] Table 3. Target-domain cross-center frozen linear-probing performance on PKUPH using 1%, 10%, and 100% of the available target-cohort training labels. All ECG representation extractors are pretrained on EchoNext-Mini and frozen; only a linear multi-label classifier is trained on PKUPH. Methods Ref. 1% Linear Probing 10% Linear Probing 100% Linear Probing AUROC AUPRC F1 AUROC AUPRC F1 AUROC AUPRC F1 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 55.34 [52.99, 57.57] 11.92 [11.56, 12.44] 17.92 [17.43, 18.85] 59.47 [56.94, 61.89] 12.92 [12.52, 13.47] 18.29 [17.87, 19.35] 72.61 [70.40, 74.89] 18.51 [17.40, 20.57] 27.30 [25.90, 30.01] ST-MEM (Na et al., 2024) ICLRâ24 58.12 [55.73, 60.46] 12.54 [12.03, 13.26] 18.73 [18.06, 20.42] 69.84 [67.41, 72.18] 16.48 [15.57, 18.07] 23.46 [22.31, 26.05] 76.84 [74.72, 78.93] 22.91 [21.34, 24.76] 30.92 [29.36, 33.18] HeartLang (Jin et al., 2025) ICLRâ25 56.29 [53.92, 58.64] 11.99 [11.64, 12.67] 18.18 [17.64, 20.38] 71.18 [68.85, 73.40] 17.25 [16.45, 18.95] 24.50 [23.51, 27.64] 76.95 [74.88, 78.86] 22.90 [21.28, 25.94] 31.31 [30.80, 36.02] ECGâText Pretraining CLIP (Radford et al., 2021) ICMLâ21 65.44 [62.74, 68.11] 16.27 [14.93, 18.52] 25.07 [23.38, 28.33] 68.92 [66.26, 71.35] 18.24 [16.78, 20.61] 26.11 [24.68, 29.63] 74.53 [72.14, 76.92] 21.25 [19.64, 23.73] 28.87 [27.38, 32.18] SigLIP (Zhai et al., 2023) ICCVâ23 68.51 [66.26, 71.03] 17.70 [16.29, 20.13] 26.65 [25.07, 30.34] 68.80 [66.02, 71.56] 18.97 [17.45, 21.49] 27.54 [25.98, 31.13] 73.01 [70.68, 75.49] 20.18 [18.65, 22.67] 28.56 [26.67, 32.17] PCME++ (Chun, 2023) ICLRâ24 63.29 [60.39, 66.07] 14.61 [13.77, 16.17] 22.55 [21.42, 25.20] 66.98 [64.51, 69.54] 15.73 [14.90, 17.33] 22.90 [21.82, 25.44] 75.06 [73.17, 76.86] 18.77 [17.63, 20.75] 26.39 [24.89, 29.49] MERL-ECHO (Wong et al., 2025) medRxivâ25 67.79 [64.92, 70.24] 15.83 [15.02, 17.37] 24.21 [22.90, 26.96] 71.55 [68.85, 73.80] 17.54 [16.71, 19.17] 25.08 [23.87, 28.18] 77.88 [76.88, 80.71] 21.24 [20.11, 23.29] 29.90 [28.54, 32.94] ECG-CLIP (Zhou et al., 2025) npj DMâ25 68.47 [65.83, 70.94] 18.34 [16.72, 20.05] 28.05 [26.41, 29.20] 72.66 [70.13, 74.86] 19.72 [18.03, 21.24] 29.18 [27.38, 30.42] 78.02 [76.01, 79.12] 23.62 [22.01, 25.12] 31.76 [30.18, 33.16] D-BETA (Hung et al., 2025) ICMLâ25 69.63 [67.06, 71.84] 18.91 [17.23, 20.19] 27.42 [25.68, 29.08] 73.58 [71.12, 75.18] 20.10 [18.39, 21.52] 28.76 [26.91, 30.31] 77.48 [75.41, 79.03] 23.84 [22.21, 25.18] 32.06 [30.45, 33.22] SGERA (Chen et al., 2026) ICMLâ26 70.91 [68.42, 72.47] 18.62 [16.91, 20.10] 27.88 [26.03, 29.16] 73.21 [70.75, 75.12] 20.51 [18.72, 21.61] 28.96 [27.02, 30.37] 77.66 [75.62, 79.17] 24.11 [22.47, 25.29] 31.84 [30.26, 33.19] EchoBridge Ours 72.68 [70.26, 75.00] 20.32 [18.31, 23.88] 29.31 [27.44, 33.90] 75.42 [72.92, 77.72] 21.79 [19.82, 25.28] 30.60 [28.55, 35.02] 79.48 [77.32, 81.46] 25.43 [23.70, 28.62] 33.40 [31.86, 37.62] 4. Results 4.1. Prompt-Based Classifier-Free Inference Table 1 evaluates whether the frozen shared space supports classifier-free prediction using fixed definition-only textual prototypes. EchoBridge achieves 75.73 AUROC, 26.79 AUPRC, and 31.83 F1, exceeding the strongest competing method for each metric by 7.88, 5.61, and 4.54 percentage points, respectively. These gains indicate improved score ranking and positive-finding discrimination under imbalanced multi-label evaluation. Because all evaluated findings are incorporated during pretraining through standardized conclusion-style summaries and class-specific prototype supervision, this protocol measures the accessibility of pretraining-seen finding semantics through textual prototypes. The results support stronger correspondence between ECG representations and predefined echocardiography-derived findings. 4.2. In-Domain Frozen Linear Probing under Different Label Ratios Table 2 evaluates frozen ECG representations on EchoNext-Mini using linear multi-label classifiers trained with 1%, 10%, or 100% of the available labels. EchoBridge achieves the highest AUROC, AUPRC, and F1 among the evaluated ECG-only and ECGâechocardiography text methods at every budget. With 1% labels, it obtains 72.77 AUROC, 26.53 AUPRC, and 31.32 F1, exceeding the strongest baseline for each metric by 3.78, 2.04, and 3.24 points, respectively. These gains indicate that finding-related information remains linearly accessible under limited supervision. At 10%, the corresponding improvements are 1.01, 2.76, and 2.39 points; at 100%, they are 1.60, 2.41, and 2.25 points. Consistent AUPRC and F1 gains indicate improved positive-finding discrimination under class imbalance. ECGâsummary alignment with class-aware prototype supervision therefore yields frozen representations that remain linearly accessible across in-domain label budgets. Table 4. Target-domain cross-center frozen linear-probing performance on SHTMU using 1%, 10%, and 100% of the available target-cohort training labels. All ECG representation extractors are pretrained on EchoNext-Mini and frozen; only a linear multi-label classifier is trained on SHTMU. Methods Ref. 1% Linear Probing 10% Linear Probing 100% Linear Probing AUROC AUPRC F1 AUROC AUPRC F1 AUROC AUPRC F1 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 56.20 [53.85, 58.46] 15.95 [15.46, 16.78] 22.93 [22.20, 24.50] 63.43 [61.00, 65.61] 18.85 [18.20, 19.96] 25.05 [24.57, 26.96] 69.91 [67.50, 72.05] 23.48 [22.62, 24.99] 29.33 [28.42, 31.36] ST-MEM (Na et al., 2024) ICLRâ24 57.02 [54.61, 59.31] 16.71 [15.96, 17.83] 23.86 [22.98, 25.96] 65.91 [63.53, 68.17] 21.38 [20.46, 22.86] 27.44 [26.52, 29.72] 73.88 [71.81, 75.81] 25.92 [24.88, 27.65] 31.68 [30.61, 32.87] HeartLang (Jin et al., 2025) ICLRâ25 55.41 [52.94, 57.59] 16.38 [15.62, 17.49] 23.20 [22.46, 25.09] 67.48 [65.52, 69.65] 23.07 [22.10, 24.54] 28.56 [27.88, 30.34] 73.39 [71.60, 75.29] 25.12 [24.15, 26.89] 31.10 [30.28, 33.56] ECGâText Pretraining CLIP (Radford et al., 2021) ICMLâ21 59.06 [56.56, 61.72] 17.95 [17.09, 19.80] 25.28 [24.09, 27.98] 64.25 [61.72, 66.79] 19.93 [19.10, 21.65] 26.66 [25.67, 29.20] 69.11 [66.98, 71.26] 24.04 [22.88, 26.12] 29.97 [28.57, 32.79] SigLIP (Zhai et al., 2023) ICCVâ23 60.55 [58.06, 63.07] 18.19 [17.25, 20.23] 26.22 [24.71, 28.97] 63.23 [60.78, 65.62] 19.60 [18.74, 21.18] 26.80 [25.29, 29.60] 66.47 [64.22, 68.74] 22.11 [20.96, 24.26] 28.58 [26.93, 31.81] PCME++ (Chun, 2023) ICLRâ24 60.01 [57.56, 62.44] 17.28 [16.57, 18.36] 24.47 [23.66, 26.50] 61.80 [59.37, 64.32] 18.33 [17.61, 19.51] 25.18 [24.35, 27.57] 68.58 [66.46, 70.63] 21.90 [21.00, 23.37] 27.97 [27.00, 30.52] MERL-ECHO (Wong et al., 2025) medRxivâ25 64.11 [61.47, 66.72] 18.58 [17.75, 20.03] 26.34 [25.28, 28.80] 66.56 [63.89, 69.09] 20.26 [19.34, 21.74] 27.57 [26.54, 29.94] 71.79 [69.61, 73.89] 23.73 [22.73, 25.36] 30.22 [29.20, 32.86] ECG-CLIP (Zhou et al., 2025) npj DMâ25 65.22 [62.69, 67.68] 20.04 [19.01, 21.82] 28.35 [26.93, 30.42] 67.88 [65.41, 70.22] 22.85 [21.76, 24.21] 28.63 [27.51, 30.54] 73.92 [71.83, 75.72] 25.47 [24.34, 27.23] 31.72 [30.54, 32.79] D-BETA (Hung et al., 2025) ICMLâ25 66.43 [63.91, 68.81] 21.28 [20.16, 22.73] 28.12 [26.74, 30.32] 69.46 [67.08, 71.68] 22.77 [21.72, 24.31] 29.42 [28.23, 30.76] 73.68 [71.59, 75.52] 26.72 [25.58, 27.96] 31.36 [30.25, 32.78] SGERA (Chen et al., 2026) ICMLâ26 67.74 [65.36, 69.12] 21.05 [19.98, 22.64] 29.21 [27.85, 30.57] 70.33 [68.05, 71.72] 22.58 [21.49, 24.16] 29.40 [28.26, 30.98] 74.37 [72.42, 75.78] 26.41 [25.31, 27.82] 31.60 [30.53, 32.79] EchoBridge Ours 69.40 [67.03, 71.57] 23.06 [22.07, 24.48] 30.78 [29.82, 33.10] 72.06 [69.91, 74.15] 24.76 [23.73, 26.24] 31.00 [30.07, 33.00] 75.99 [74.02, 77.85] 28.11 [26.96, 29.89] 32.98 [32.25, 35.53] Table 5. Source-only cross-center transfer from EchoNext-Mini to PKUPH and SHTMU. The ECG representation extractor and linear multi-label classifier are trained on EchoNext-Mini, and all model parameters and class-specific decision thresholds are fixed during target-cohort evaluation. Methods Ref. PKUPH SHTMU AUROC AUPRC F1 AUROC AUPRC F1 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 70.78 [69.96, 71.61] 13.46 [12.85, 14.19] 22.26 [21.48, 23.31] 70.28 [69.39, 71.16] 17.33 [16.64, 18.18] 24.21 [23.34, 25.38] ST-MEM (Na et al., 2024) ICLRâ24 73.30 [72.51, 74.08] 14.47 [13.84, 15.23] 25.23 [24.42, 26.31] 70.43 [69.55, 71.32] 17.86 [17.18, 18.72] 24.10 [23.24, 25.27] HeartLang (Jin et al., 2025) ICLRâ25 73.22 [72.43, 74.01] 14.59 [13.95, 15.36] 26.50 [25.67, 27.61] 71.44 [70.57, 72.31] 18.76 [18.05, 19.64] 24.52 [23.63, 25.70] ECGâText Pretraining CLIP (Radford et al., 2021) ICMLâ21 70.42 [69.57, 71.25] 16.59 [15.92, 17.40] 25.67 [24.84, 26.77] 67.24 [66.31, 68.17] 19.09 [18.35, 19.98] 23.68 [22.78, 24.87] SigLIP (Zhai et al., 2023) ICCVâ23 66.53 [65.64, 67.42] 13.98 [13.34, 14.76] 22.03 [21.23, 23.12] 60.62 [59.61, 61.62] 16.20 [15.50, 17.08] 22.11 [21.20, 23.32] PCME++ (Chun, 2023) ICLRâ24 66.92 [66.04, 67.81] 14.59 [13.94, 15.38] 23.58 [22.77, 24.68] 64.51 [63.55, 65.47] 15.73 [15.02, 16.61] 21.69 [20.80, 22.88] MERL-ECHO (Wong et al., 2025) medRxivâ25 76.30 [75.55, 77.05] 16.54 [15.87, 17.36] 26.40 [25.58, 27.50] 72.47 [71.62, 73.32] 18.75 [18.04, 19.62] 25.69 [24.80, 26.86] ECG-CLIP (Zhou et al., 2025) npj DMâ25 74.09 [73.31, 74.87] 14.60 [13.96, 15.38] 24.39 [23.58, 25.48] 72.41 [71.56, 73.27] 19.20 [18.47, 20.10] 24.74 [23.86, 25.91] D-BETA (Hung et al., 2025) ICMLâ25 76.78 [76.04, 77.52] 17.45 [16.77, 18.27] 27.13 [26.31, 28.21] 73.58 [72.75, 74.41] 20.00 [19.26, 20.90] 25.35 [24.46, 26.52] SGERA (Chen et al., 2026) ICMLâ26 72.36 [71.56, 73.16] 15.94 [15.28, 16.74] 25.82 [25.00, 26.91] 69.83 [68.92, 70.74] 18.31 [17.59, 19.19] 24.69 [23.80, 25.87] EchoBridge Ours 77.92 [77.20, 78.64] 18.98 [18.28, 19.82] 27.39 [26.57, 28.48] 74.11 [73.29, 74.93] 22.08 [21.31, 23.01] 28.16 [27.24, 29.36] 4.3. Target-Domain Cross-Center Frozen Linear Probing under Different Label Ratios Tables 3 and 4 evaluate cross-center transferability of ECG representations pretrained on EchoNext-Mini. For each method, the ECG representation extractor, including the encoder and output projection where applicable, is frozen, and a new linear multi-label classifier is trained independently on PKUPH or SHTMU using 1%, 10%, or 100% of the target-cohort labels. Model selection and class-specific thresholds use only the corresponding validation split. This protocol measures linear accessibility under institutional and label-distribution shifts across target-domain supervision levels. On PKUPH, EchoBridge achieves the highest point estimates for all metrics at every label ratio. Relative to the strongest baseline for each metric, it improves AUROC, AUPRC, and F1 by 1.77, 1.41, and 1.26 points with 1% labels; 1.84, 1.28, and 1.42 points with 10%; and 1.46, 1.32, and 1.34 points with 100%, respectively. The gains at 1% and 10% indicate adaptability under limited target-center supervision. On SHTMU, EchoBridge also achieves the highest point estimates across all metricâbudget combinations. Its AUROC, AUPRC, and F1 gains are 1.66, 1.78, and 1.57 points with 1% labels; 1.73, 1.69, and 1.58 points with 10%; and 1.62, 1.39, and 1.26 points with 100%, respectively. Consistent gains across both cohorts support target-domain adaptability and label efficiency under institutional and prevalence shifts. 4.4. Source-Only Cross-Center Transfer Table 5 evaluates source-only cross-center transfer by training the linear classifier and selecting class-specific thresholds exclusively on EchoNext-Mini, then applying the fixed representation extractor, classifier, and thresholds to findings shared with each target cohort. EchoBridge exceeds the strongest competing method for each metric by 1.14 AUROC, 1.53 AUPRC, and 0.26 F1 points on PKUPH, and by 0.53, 2.08, and 2.47 points on SHTMU, respectively. On PKUPH, the larger AUROC and AUPRC gains relative to F1 suggest improved score ranking with limited change under the transferred source-domain thresholds. On SHTMU, the larger AUPRC and F1 gains indicate improved positive-finding discrimination under the shifted prevalence distribution. These results complement target-domain linear probing and demonstrate direct cross-center transfer without target-domain training, calibration, or threshold adjustment. Table 6. Matched label-only control and component-wise ablation of CSPP, FAAM, and SRR on EchoNext-Mini. Performance is evaluated using prompt-based classifier-free inference and frozen linear probing. Configuration ECGâText Align. + Proto. CSPP APBC Prompt-based 1% Linear Probing 10% Linear Probing 100% Linear Probing FAAM SRR AUROC AUPRC F1 AUROC AUPRC F1 AUROC AUPRC F1 AUROC AUPRC F1 Label-only BCE Ă Ă Ă Ă â 64.79 18.06 24.53 73.79 26.16 31.07 76.87 29.89 33.96 Align. + Proto. â Ă Ă Ă 67.18 18.83 24.09 68.48 21.90 28.00 74.11 27.20 31.80 76.61 30.18 34.08 + CSPP â â Ă Ă 71.05 22.39 27.74 70.08 23.58 29.27 75.16 28.36 32.70 77.42 31.06 34.76 + APBC â Ă â â 73.36 24.71 30.02 70.88 24.51 29.94 75.69 29.03 33.18 77.83 31.58 35.13 + CSPP + FAAM â â â Ă 74.02 25.18 30.46 71.69 25.40 30.52 76.23 29.66 33.58 78.24 32.07 35.42 + CSPP + SRR â â Ă â 73.91 25.04 30.31 71.43 25.12 30.28 76.06 29.47 33.40 78.11 31.92 35.27 EchoBridge â â â â 75.73 26.79 31.83 72.77 26.53 31.32 76.94 30.45 34.13 78.79 32.66 35.82 4.5. Ablation Study Table 6 reports a matched label-only control and ablations of CSPP and the APBC refinements FAAM and SRR. All ECGâtext variants combine bidirectional alignment with class-prototype supervision, using Align.+Proto. as the reference. The label-only BCE control tests whether structured-label supervision alone explains the gains. It uses the same ECG encoder, 256-dimensional projection, data split, optimizer, batch size, epoch budget, and schedule as the ECGâtext variants, with pretraining driven solely by structured multi-label BCE. Its encoder and projection are then frozen, and a new linear classifier is trained with 1%, 10%, or 100% of the probe labels. Prompt-based inference is reported only for ECGâtext variants because it requires a text-aligned space. Align.+Proto. outperforms the label-only control under 1% and 10% probing, indicating greater linear accessibility with limited supervision, while their 100% results are comparable. CSPP improves prompt-based inference and frozen probing across all budgets, and complete APBC without CSPP improves every metric. With CSPP, both FAAM and SRR outperform CSPP alone, with slightly larger gains from FAAM. EchoBridge performs best across all protocols, supporting their complementary contributions beyond Align.+Proto. and matched label supervision. 4.6. Finding-Specific Performance across Prevalence Levels As shown in Figure 5, under 100% frozen linear probing on EchoNext-Mini, EchoBridge achieves the highest AUROC point estimate for all seven findings, with the largest gains for moderate-or-severe aortic stenosis and tricuspid regurgitation at 1.59 and 1.57 points, respectively. It also obtains the highest AUPRC for six findings, improving pulmonic and tricuspid regurgitation by 8.15 and 2.36 points, respectively, while its LVH AUPRC is 0.67 points lower. Appendix F reports finding-specific AUROC and AUPRC on EchoNext-Mini, PKUPH, and SHTMU under the corresponding 100% frozen linear-probing protocols. On both target cohorts, EchoBridge achieves the highest point estimates for every finding across ventricular function, chamber abnormalities, and valvular disease. These cross-institutional gains include several low-prevalence valvular abnormalities. Their varying magnitudes indicate condition-specific benefits, while estimates based on few positive test cases require cautious interpretation under severe class imbalance. Figure 5. Finding-specific performance on EchoNext-Mini under 100% frozen linear probing. (a) AUROC by method for each echocardiography-derived finding. (b) EchoBridgeâs AUPRC difference from the strongest baseline for each finding in percentage points; positive values indicate gains. 5. Limitations and Ethical Considerations The retrospective study covers a limited number of institutions; prospective multicenter validation across populations, acquisition systems, and clinical workflows remains necessary, particularly for low-prevalence valvular abnormalities. Performance may also vary across demographic groups, ECG devices, acquisition protocols, and care settings. All private data were de-identified and processed in access-controlled institutional environments. The PKUPH and SHTMU cohorts were retrospectively collected under approvals from the Institutional Review Boards of Peking University Peopleâs Hospital (Approval No. 2024PHB428-001) and the Second Hospital of Tianjin Medical University (Approval No. KY2025K386), respectively. EchoBridge is intended for research and clinician-assisted screening, with predictions interpreted by qualified professionals alongside other clinical evidence. False-negative predictions, particularly for rare abnormalities, may delay further assessment. Clinical deployment requires site-specific validation, calibration, monitoring, regulatory review, and explicit human oversight. 6. Conclusion We presented EchoBridge, an ECGâechocardiography text alignment framework for transferable representations of echocardiography-derived cardiac findings. It combines Complementary SharedâPrivate Projection and Adaptive Prototype Boundary Calibration to address cross-modal interference and long-tailed distributions. Across EchoNext-Mini and two independent cohorts, EchoBridge outperforms representative ECG-only and ECGâtext baselines across classifier-free inference, in-domain and target-domain frozen probing, source-only transfer, and finding-specific evaluation. Generative AI Usage LLMs assisted with auxiliary coding, translation, editing, and drafting definition-only prompts for echocardiography-derived findings. GPT-5.5 Thinking generated initial candidates, which senior cardiologists reviewed and standardized against predefined label definitions before testing. The authors independently developed the methodology, conducted experiments, interpreted results, and assume full responsibility for the manuscript. References A. Aminorroaya, L. S. Dhingra, A. F. Pedroso, S. V. Shankar, A. Coppi, A. Khunte, M. Foppa, L. C. Brant, S. M. Barreto, A. L. P. Ribeiro, et al. (2025) Development and multinational validation of an ensemble deep learning algorithm for detecting and predicting structural heart disease using noisy single-lead electrocardiograms. European Heart Journal-Digital Health 6 (4), p. 554â566. Cited by: §B.1. Z. I. Attia, S. Kapa, F. Lopez-Jimenez, P. M. McKie, D. J. Ladewig, G. Satam, P. A. Pellikka, M. Enriquez-Sarano, P. A. Noseworthy, T. M. Munger, et al. (2019) Screening for cardiac contractile dysfunction using an artificial intelligenceâenabled electrocardiogram. Nature medicine 25 (1), p. 70â74. Cited by: §B.1, §1. C. Chao, J. Delbrouck, M. Asadi, I. Banerjee, J. M. Farina, F. Galasso, A. K. Mahmoud, M. T. Abbas, Y. Wang, R. Arsanjani, et al. (2025) EchoGraph system for automated quality assessment of echocardiography reports. NPJ Digital Medicine. Cited by: §1. J. Chen, X. Dong, W. Wang, S. Zhou, L. Yu, and X. Hu (2025) DERI: cross-modal ecg representation learning with deep ecg-report interaction. In 34th International Joint Conference on Artificial Intelligence, IJCAI 2025, p. 4824â4832. Cited by: §B.2. J. Chen, Y. Du, W. Yuan, S. Wang, J. Xu, Z. Liu, R. Zhao, and E. C. H. Ngai (2026) SGERA: stein-guided ecg-report alignment for ecg representation learning. In Forty-third International Conference on Machine Learning, Cited by: §B.2, Table 10, Table 11, Table 12, Table 1, Table 2, Table 3, Table 4, Table 5. T. Chen, S. Kornblith, M. Norouzi, and G. Hinton (2020) A simple framework for contrastive learning of visual representations. In International conference on machine learning, p. 1597â1607. Cited by: Table 10, Table 11, Table 12, Table 2, Table 3, Table 4, Table 5. M. Christensen, M. Vukadinovic, N. Yuan, and D. Ouyang (2024) Visionâlanguage foundation model for echocardiogram interpretation. Nature Medicine 30 (5), p. 1481â1488. Cited by: §1. S. Chun (2023) Improved probabilistic image-text representations. arXiv preprint arXiv:2305.18171. Cited by: Table 10, Table 11, Table 12, Table 1, Table 2, Table 3, Table 4, Table 5. L. S. Dhingra, A. Aminorroaya, V. Sangha, A. F. Pedroso, S. V. Shankar, A. Coppi, M. Foppa, L. C. Brant, S. M. Barreto, A. L. P. Ribeiro, et al. (2025) Ensemble deep learning algorithm for structural heart disease screening using electrocardiographic images: present shd. Journal of the American College of Cardiology 85 (12), p. 1302â1313. Cited by: §B.1. P. Elias, T. J. Poterucha, V. Rajaram, L. M. Moller, V. Rodriguez, S. Bhave, R. T. Hahn, G. Tison, S. A. Abreau, J. Barrios, et al. (2022) Deep learning electrocardiographic analysis for detection of left-sided valvular heart disease. Journal of the American College of Cardiology 80 (6), p. 613â626. Cited by: §B.1. X. Fang, Z. Ding, J. Cai, Y. Xiao, B. Liu, J. Jin, H. Wang, G. Nie, S. Huang, T. Chen, et al. (2026) ECGFlowCMR: pretraining with ecg-generated cine cmr improves cardiac disease classification and phenotype prediction. arXiv preprint arXiv:2601.20904. Cited by: §B.2. X. Fang, J. Jin, H. Wang, C. Liu, J. Cai, Y. Xiao, G. Nie, B. Liu, S. Huang, H. Li, et al. (2025) PPGFlowECG: latent rectified flow with cross-modal encoding for ppg-guided ecg generation and cardiovascular disease detection. arXiv preprint arXiv:2509.19774. Cited by: §B.2. G. Fujiki, S. Kodera, N. Setoguchi, K. Tanabe, K. Miyaji, S. Kushida, M. Saji, M. Nanasato, H. Maki, H. Fujita, et al. (2025) Deep learning-based identification of echocardiographic abnormalities from electrocardiograms. JACC: Asia 5 (1_Part_1), p. 88â98. Cited by: §B.1. A. Ghorbani, D. Ouyang, A. Abid, B. He, J. H. Chen, R. A. Harrington, D. H. Liang, E. A. Ashley, and J. Y. Zou (2020) Deep learning interpretation of echocardiograms. NPJ digital medicine 3 (1), p. 10. Cited by: §1. J. W. Hughes, L. Jing, J. Finer, D. Hartzel, C. Kelsey, A. Long, D. Rocha, J. Ruhl, T. Poterucha, and P. Elias (2026) EchoNext-mini: a dataset and baseline ai model for detecting structural heart disease from electrocardiograms. NEJM AI 3 (5), p. AIdbp2500516. Cited by: §B.1, §D.1, §3.1. M. P. Hung, A. Saeed, and D. Ma (2025) Boosting masked ecg-text auto-encoders as discriminative learners. In Forty-second International Conference on Machine Learning, Cited by: §B.2, Table 10, Table 11, Table 12, Table 1, Table 2, Table 3, Table 4, Table 5. J. Jin, H. Wang, H. Li, J. Li, J. Pan, and S. Hong (2025) Reading your heart: learning ecg words and sentences via pre-training ecg language model. arXiv preprint arXiv:2502.10707. Cited by: Table 10, Table 11, Table 12, Table 2, Table 3, Table 4, Table 5. J. Jin, H. Wang, X. Wu, X. Fang, X. Lan, Z. Wang, D. Zhang, B. Liu, Y. Zhang, X. Wu, et al. (2026) ECG-r1: protocol-guided and modality-agnostic mllm for reliable ecg interpretation. arXiv preprint arXiv:2602.04279. Cited by: §B.2. Q. Jin, W. Kim, Q. Chen, D. C. Comeau, L. Yeganova, W. J. Wilbur, and Z. Lu (2023) Medcpt: contrastive pre-trained transformers with large-scale pubmed search logs for zero-shot biomedical information retrieval. Bioinformatics 39 (11), p. btad651. Cited by: §C.2, §3.3. G. A. Kaissis, M. R. Makowski, D. RĂŒckert, and R. F. Braren (2020) Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence 2 (6), p. 305â311. Cited by: §1. E. Knight, E. K. Oikonomou, A. Aminorroaya, A. F. Pedroso, and R. Khera (2026) Wearable-echo-fm: an ecg echo foundation model for 1-lead electrocardiography. European Heart Journal-Digital Health 7 (4), p. ztag049. Cited by: §B.3, §1. G. H. Kwak, D. Moukheiber, M. Moukheiber, L. Moukheiber, S. Moukheiber, N. M. Butala, L. A. Celi, and C. W. Chen (2025) Large open access database of echocardiogram reports in intensive care unit patients. Scientific Data 12 (1), p. 1153. Cited by: §1. J. Kwon, S. Y. Lee, K. Jeon, Y. Lee, K. Kim, J. Park, B. Oh, and M. Lee (2020) Deep learningâbased algorithm for detecting aortic stenosis using electrocardiography. Journal of the American Heart Association 9 (7), p. e014717. Cited by: §B.1. S. K. Lalam, H. K. Kunderu, S. Ghosh, H. Kumar, S. Awasthi, A. Prasad, F. Lopez-Jimenez, Z. I. Attia, S. Asirvatham, P. Friedman, et al. (2023) Ecg representation learning with multi-modal ehr data. Transactions on Machine Learning Research. Cited by: §B.2. J. Li, C. Liu, S. Cheng, R. Arcucci, and S. Hong (2024) Frozen language model helps ecg zero-shot learning. In Medical Imaging with Deep Learning, p. 402â415. Cited by: §B.2. C. Liu, C. Ouyang, Z. Wan, H. Wang, W. Bai, and R. Arcucci (2025) Knowledge-enhanced multimodal ecg representation learning with arbitrary-lead inputs. arXiv preprint arXiv:2502.17900. Cited by: §B.2. C. Liu, Z. Wan, S. Cheng, M. Zhang, and R. Arcucci (2024a) Etp: learning transferable ecg representations via ecg-text pre-training. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 8230â8234. Cited by: §B.2. C. Liu, Z. Wan, C. Ouyang, A. Shah, W. Bai, and R. Arcucci (2024b) Zero-shot ecg classification with multimodal learning and test-time clinical knowledge enhancement. arXiv preprint arXiv:2403.06659. Cited by: §B.2, §1. Y. Na, M. Park, Y. Tae, and S. Joo (2024) Guiding masked representation learning to capture spatio-temporal relationship of electrocardiogram. arXiv preprint arXiv:2402.09450. Cited by: Table 10, Table 11, Table 12, Table 2, Table 3, Table 4, Table 5. G. Nie, G. Tang, Y. Xiao, J. Li, S. Huang, D. Zhang, Q. Zhao, and S. Hong (2025) Anyppg: an ecg-guided ppg foundation model trained on over 100,000 hours of recordings for holistic health profiling. arXiv preprint arXiv:2511.01747. Cited by: §B.2. D. Ouyang, B. He, A. Ghorbani, N. Yuan, J. Ebinger, C. P. Langlotz, P. A. Heidenreich, R. A. Harrington, D. H. Liang, E. A. Ashley, et al. (2020) Video-based ai for beat-to-beat assessment of cardiac function. Nature 580 (7802), p. 252â256. Cited by: §1. T. J. Poterucha, L. Jing, R. P. Ricart, M. Adjei-Mosi, J. Finer, D. Hartzel, C. Kelsey, A. Long, D. Rocha, J. A. Ruhl, et al. (2025) Detecting structural heart disease from electrocardiograms using ai. Nature 644 (8075), p. 221â230. Cited by: §B.1, §1. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748â8763. Cited by: Table 10, Table 11, Table 12, Table 1, Table 2, Table 3, Table 4, Table 5. A. E. Ulloa-Cerna, L. Jing, J. M. Pfeifer, S. Raghunath, J. A. Ruhl, D. B. Rocha, J. B. Leader, N. Zimmerman, G. Lee, S. R. Steinhubl, et al. (2022) RECHOmmend: an ecg-based machine learning approach for identifying patients at increased risk of undiagnosed structural heart disease detectable by echocardiography. Circulation 146 (1), p. 36â47. Cited by: §B.1. M. Vukadinovic, I. Chiu, X. Tang, N. Yuan, T. Chen, P. Cheng, D. Li, S. Cheng, B. He, and D. Ouyang (2026) Comprehensive echocardiogram evaluation with view primed vision language ai. Nature 650 (8103), p. 970â977. Cited by: §1. W. Wong, C. Liu, P. Elias, J. W. Hughes, C. Leung, X. Qian, H. Li, Y. Lau, C. Tao, A. Choo, et al. (2025) Contrastive multi-modal training with electrocardiography and natural language echocardiography reports for zero-shot prediction of structural heart disease. medRxiv, p. 2025â09. Cited by: §B.3, Table 10, Table 11, Table 12, §1, Table 1, Table 2, Table 3, Table 4, Table 5. X. Yao, D. R. Rushlow, J. W. Inselman, R. G. McCoy, T. D. Thacher, E. M. Behnken, M. E. Bernard, S. L. Rosas, A. Akfaly, A. Misra, et al. (2021) Artificial intelligenceâenabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trial. Nature medicine 27 (5), p. 815â819. Cited by: §B.1, §1. H. Yu, P. Guo, and A. Sano (2024) Ecg semantic integrator (esi): a foundation ecg model pretrained with llm-enhanced cardiological text. arXiv preprint arXiv:2405.19366. Cited by: §B.2. X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer (2023) Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, p. 11975â11986. Cited by: Table 10, Table 11, Table 12, Table 1, Table 2, Table 3, Table 4, Table 5. X. Zhou, T. Li, H. Hayama, K. Nakamura, S. Liu, W. Chen, and X. Zhu (2025) Diagnosis of cardiac conditions from 12-lead electrocardiogram through natural language supervision. npj Digital Medicine 8 (1), p. 697. Cited by: §B.2, Table 10, Table 11, Table 12, §1, Table 1, Table 2, Table 3, Table 4, Table 5. Appendix A Acknowledgments As an informal quality check, five senior cardiologists inspected a randomly sampled subset of mappings from structured cardiac findings to standardized conclusion-style summaries. Each sampled summary was presented with its corresponding finding vector. The reviewers provided qualitative feedback that the sampled summaries were broadly clinically plausible and semantically consistent with the encoded findings. This review was not designed as a formal annotation, quantitative validation, or inter-rater agreement study. The expert panel comprised Qinghao Zhao (Peking University Peopleâs Hospital, Beijing, China), Guanyu Mu (The Second Hospital of Tianjin Medical University, Tianjin, China), Xingliang Wu (Tianjin Institute of Cardiology, Tianjin, China), Xinxin Di (The First Affiliated Hospital of USTC, Hefei, China), and Jing Zhao (The First Affiliated Hospital of Anhui Medical University, Hefei, China). Appendix B Related Work B.1. Direct Echocardiography Supervision Echocardiography-derived supervision is widely used for ECG-based structural heart disease assessment. Early studies paired ECG and echocardiography data to detect ventricular systolic dysfunction or reduced ejection fraction (Attia et al., 2019; Yao et al., 2021), followed by work on specific abnormalities such as aortic stenosis and left-sided valvular disease (Kwon et al., 2020; Elias et al., 2022). These studies established supervised ECG-to-echo prediction using echocardiographic measurements or diagnoses as screening targets. More recent work expanded to broader structural heart disease assessment. rECHOmmend (Ulloa-Cerna et al., 2022) used a composite endpoint to identify patients at risk of undiagnosed echocardiography-detectable disease, while Fujiki et al. (Fujiki et al., 2025) studied multi-label prediction across ventricular function, chamber structure, and valvular abnormalities. EchoNext scaled this direction with paired examinations and public resources for evaluating echo-confirmed structural heart disease detection (Poterucha et al., 2025; Hughes et al., 2026). Related studies also examined ECG images and single-lead recordings (Dhingra et al., 2025; Aminorroaya et al., 2025). These methods use predefined echocardiography-derived targets, motivating compositional ECGâechocardiography alignment organized around predefined findings and their co-occurrence patterns. B.2. ECGâReport Representation Learning ECGâtext pretraining is a major approach to language-supervised ECG representation learning. METS (Li et al., 2024) and ETP (Liu et al., 2024a) learn shared embeddings from paired ECGs and machine-generated or clinical reports for zero-shot classification and label-efficient evaluation. MERL (Liu et al., 2024b) uses clinical knowledge-enhanced prompts at inference, while ECG-CLIP (Zhou et al., 2025) scales CLIP-style supervision to large ECGâreport datasets for zero-shot diagnosis across cardiac conditions. Multimodal ECG pretraining further combines reports, EHR records, PPG, and cardiac imaging (Lalam et al., 2023; Fang et al., 2025; Nie et al., 2025; Fang et al., 2026). Recent methods introduce richer interaction and structured supervision (Jin et al., 2026). DERI (Chen et al., 2025) combines multiple alignment objectives with mutual feature reconstruction, ESI (Yu et al., 2024) enriches reports with LLM-generated cardiological descriptions, and K-MERL (Liu et al., 2025) extracts structured knowledge from free-text reports while supporting arbitrary-lead inputs. D-BETA (Hung et al., 2025) integrates masked ECGâtext autoencoding with discriminative contrastive learning, while SGERA (Chen et al., 2026) uses Stein-guided alignment to address structural and statistical modality discrepancies. These studies establish ECGâreport alignment for transferable ECG representation learning. Their supervision primarily comes from ECG reports, diagnostic statements, and broader clinical records, emphasizing rhythm, conduction, waveform morphology, and general ECG diagnoses. B.3. ECGâEchocardiography Text Alignment Recent studies use paired ECGs and echocardiography reports as cross-modal supervision for structural cardiac representation learning. MERL-ECHO applies CLIP-style contrastive pretraining to encode 12-lead ECGs and echocardiography reports in a shared space for zero-shot structural heart disease prediction (Wong et al., 2025). Wearable-Echo-FM extends this paradigm to single-lead ECGs and evaluates label-efficient fine-tuning for left ventricular systolic dysfunction, diastolic dysfunction, and composite structural heart disease (Knight et al., 2026). These studies establish ECGâechocardiography text alignment as a viable approach for transferring structural cardiac information to ECG encoders and improving label efficiency. However, global cross-modal alignment may entangle modality-specific factors, while imbalanced finding distributions provide uneven class-level supervision. EchoBridge addresses these limitations through sharedâprivate projection and class-aware geometric organization of the normalized representation space. Algorithm 1 End-to-End Training of EchoBridge 1:Labeled paired data =(xi,ti,i)D=\(x_i,t_i,y_i)\; training epochs E; number of cardiac findings C; training-set positive rates =rcc=1Cr=\r_c\_c=1^C; alignment temperature Ï; margin parameters (m0,mmin,mmax)(m_0,m_ ,m_ ); numerical constant ϔΔ; Riesz exponent q; repulsion weight λr _r. 2:Optimized model parameters Î . 3:Initialize the learnable prototype matrix P=[p1,âŠ,pC]â€ââCĂdP=[p_1,âŠ,p_C] ^CĂ d and learnable logit-scale parameter s. 4:Compute the normalization factor Îșâ1Cââc=1CmedianâĄ()rc+Ï”Îșâ 1C _c=1^C median(r)r_c+Δ. 5:Compute the fixed class-specific margins mcâclipâĄ(m0âmedianâĄ()rc+Ï”â1Îș,mmin,mmax)m_c \! (m_0 median(r)r_c+Δ 1Îș,m_ ,m_ ) for c=1,âŠ,Cc=1,âŠ,C. 6:for e=1,âŠ,Ee=1,âŠ,E do 7: for each mini-batch (xi,ti,i)i=1BâŒ\(x_i,t_i,y_i)\_i=1^B do 8: Encode the paired inputs: he,iâfeâ(xi)h_e,iâ f_e(x_i) and ht,iâftâ(ti)h_t,iâ f_t(t_i). 9: Obtain the shared and private ECG projections: ze,isâÏesâ(he,i)z_e,i^sâ _e^s(h_e,i) and ze,ipâÏepâ(he,i)z_e,i^pâ _e^p(h_e,i). 10: Obtain the shared and private text projections: zt,isâÏtsâ(ht,i)z_t,i^sâ _t^s(h_t,i) and zt,ipâÏtpâ(ht,i)z_t,i^pâ _t^p(h_t,i). 11: Construct normalized shared and private copies and compute the within-modality orthogonality loss âorthL_orth using Eq. (4). 12: Normalize the shared representations and class prototypes: z^e,isâze,is/â„ze,isâ„2 z_e,i^sâ z_e,i^s/ z_e,i^s _2, z^t,isâzt,is/â„zt,isâ„2 z_t,i^sâ z_t,i^s/ z_t,i^s _2, and p^câpc/â„pcâ„2 p_câ p_c/ p_c _2. 13: Compute the cross-modal similarity matrix and symmetric alignment loss âalignL_align using Eqs. (6)â(7). 14: Compute ECGâprototype and textâprototype cosine similarities ui,ce=âšz^e,is,p^câ©u_i,c^e= z_e,i^s, p_c and ui,ct=âšz^t,is,p^câ©u_i,c^t= z_t,i^s, p_c using Eq. (10). 15: Set ÎłâexpâĄ(s)Îłâ (s) and compute the margin-adjusted logits â~i,ce _i,c^e and â~i,ct _i,c^t by applying mcm_c only to positive sampleâprototype pairs according to Eq. (14). 16: Compute the multi-label prototype loss âprotoâBCEâĄ(â~e,)+BCEâĄ(â~t,)L_proto ( ^e,y)+BCE( ^t,y) using Eq. (15). 17: Compute the spherical Riesz repulsion loss ârieszL_riesz over all normalized class-prototype pairs using Eq. (16). 18: Compute âapbcââproto+λrâârieszL_apbc _proto+ _rL_riesz and âââalign+âorth+âapbcL _align+L_orth+L_apbc. 19: Update fef_e, ftf_t, Ïes _e^s, Ïep _e^p, Ïts _t^s, Ïtp _t^p, P, and s using AdamW and âÎâ _ L. 20: end for 21:end for 22:return ECG encoder fef_e and shared ECG projection head Ïes _e^s. Table 7. Layer-wise configuration of the one-dimensional ResNet-18 ECG encoder for an input ECG of shape 12Ă500012Ă 5000. Each residual stage contains two basic blocks, and the reported stride applies to the first block of each stage. Component Blocks Input channels Output channels Kernel size First-block stride Temporal length Stem convolution 1 12 64 7 2 2500 Residual stage 1 2 64 64 3 1 2500 Residual stage 2 2 64 128 3 2 1250 Residual stage 3 2 128 256 3 2 625 Residual stage 4 2 256 512 3 2 313 Adaptive average pooling 1 512 512 â â 1 Appendix C Implementation Details of EchoBridge C.1. Algorithm of EchoBridge Algorithm 1 summarizes EchoBridgeâs end-to-end training, including sharedâprivate projection, within-modality orthogonality, bidirectional ECGâtext alignment, frequency-adaptive prototype calibration, spherical Riesz repulsion, and joint optimization. C.2. Encoder and Projection Architectures ECG Encoder. We use a one-dimensional ResNet-18 initialized from scratch to encode each 10-second, 12-lead ECG recording with input shape 12Ă500012Ă 5000. The encoder follows the four-stage ResNet-18 architecture, with all two-dimensional operations replaced by one-dimensional counterparts. We remove the max-pooling layer after the stem convolution to preserve temporal resolution. The stem comprises a bias-free Conv1DâĄ(12,64,7,2,3)Conv1D(12,64,7,2,3) layer followed by batch normalization and ReLU. Each basic residual block contains two 3-tap convolutions: (19) Conv1DâĄ(k=3)âBNâReLUâConv1DâĄ(k=3)âBN.Conv1D(k=3) 1D(k=3) . The first block of each stage except the first uses stride 2 to reduce temporal resolution and increase channels. Dimension-changing shortcuts use a bias-free 1Ă11Ă 1 convolution followed by batch normalization. A ReLU is applied after adding the residual and shortcut paths. Adaptive average pooling produces the 512-dimensional ECG representation heh_e. Table 7 summarizes the architecture. Text Encoder. We use the pretrained MedCPT-Article-Encoder (Jin et al., 2023), a BERT-base model with 12 Transformer layers, a hidden size of 768, 12 self-attention heads, and a 3,072-dimensional feed-forward layer. It uses GELU, 0.1 dropout, a 30,522-token vocabulary, and a maximum sequence length of 512. The 768-dimensional pooler_output serves as the global text representation hth_t. All parameters are jointly fine-tuned during cross-modal pretraining. Shared and Private Projections. Four independent two-layer heads map the ECG and text representations into 256-dimensional shared and auxiliary private projections: (20) Ïes,Ïep _e^s,\, _e^p :LinearâĄ(512,256)âGELUâLinearâĄ(256,256), :Linear(512,256) (256,256), (21) Ïts,Ïtp _t^s,\, _t^p :LinearâĄ(768,256)âGELUâLinearâĄ(256,256). :Linear(768,256) (256,256). Shared projections are â2 _2-normalized before cross-modal alignment and prototype supervision. Private projections remain unnormalized, while normalized copies of both branches compute the within-modality orthogonality loss. Appendix D Additional Dataset Details EchoNext-Mini labels follow the published structured definitions of the original dataset. For PKUPH and SHTMU, echocardiography-derived findings were extracted from physician-authored reports using regular-expression-based natural language processing pipelines. The pipelines normalized synonymous clinical terms, handled negation and uncertainty expressions, parsed severity modifiers, and mapped report statements to predefined binary findings. Senior cardiologist Qinghao Zhao subsequently reviewed all extracted labels from both cohorts for clinical consistency. D.1. EchoNext-Mini Dataset EchoNext-Mini (Hughes et al., 2026) is a de-identified public subset of EchoNext derived from routine clinical data at Columbia University Irving Medical Center, including Columbia and Allen hospitals. It contains 100,000 10-second, 12-lead ECGs from 36,286 adults aged at least 18 years. Recordings were acquired between 2008 and 2022 at 250 Hz and linked to transthoracic echocardiograms performed within one year. Each record includes demographic, acquisition, and automated ECG measurement metadata; ages above 90 were capped at 90 for de-identification. Echocardiography-derived labels were constructed from structured report fields and quantitative measurements, including left ventricular ejection fraction (LVEF), ventricular wall thickness, pulmonary artery systolic pressure, tricuspid regurgitation velocity, valvular severity, right ventricular systolic function, and pericardial effusion. The dataset provides 11 component abnormalities and a composite moderate-or-greater structural heart disease label. Binary criteria include LVEF â€45%†45\% for left ventricular systolic dysfunction, maximum septal or posterior wall thickness â„13â„ 13 m for increased wall thickness, and moderate-or-severe grading for aortic stenosis and aortic, mitral, tricuspid, and pulmonic regurgitation. Positive labels require the ECG to precede the echocardiogram by at most one year. Echocardiograms with prosthetic valves, missing LVEF, or unavailable wall-thickness measurements were excluded. We evaluate seven findings: left ventricular systolic dysfunction, left ventricular hypertrophy, and moderate-or-severe aortic regurgitation, aortic stenosis, tricuspid regurgitation, mitral regurgitation, and pulmonic regurgitation. Prevalence is strongly imbalanced: left ventricular systolic dysfunction and hypertrophy each occur in approximately 24% of samples, whereas pulmonic and aortic regurgitation occur in approximately 0.8%â1.3%. D.2. PKUPH Dataset The PKUPH cohort was retrospectively assembled from longitudinal ECG and transthoracic echocardiography data collected at Peking University Peopleâs Hospital from June 2015 to May 2023. Of 74,220 screened individuals, the final cohort included 20,768 patients and 27,158 ECG recordings. ECGâechocardiography pairs were matched by prioritizing same-day examinations; otherwise, the temporally closest ECG within a ±10± 10-day window was selected. Pairs outside this window and ECGs with corrupted waveforms, incomplete leads, or inconsistent metadata were excluded. The mean age was 61.6±14.461.6± 14.4 years, and 9,193 patients (44.3%) were male. Cardiac findings were extracted from physician-authored echocardiography reports using rule-based natural language processing. The cohort includes broader diastolic-function- and wall-motion-related findings than EchoNext-Mini, including left ventricular diastolic dysfunction and wall-motion abnormality, and exhibits a distinct prevalence distribution, enabling cross-center evaluation under institutional and label-prevalence shifts. D.3. SHTMU Dataset The SHTMU cohort was retrospectively assembled from longitudinal ECG and transthoracic echocardiography data collected at the Second Hospital of Tianjin Medical University between January and December 2024. Among 479,089 individuals screened from the institutional repository, 462,468 lacked available echocardiography data. The resulting cohort included 16,621 patients and 18,588 ECG recordings. ECGâechocardiography pairs were constructed using the same temporal matching strategy as for PKUPH. Same-day examinations were prioritized; when no same-day ECG was available, the temporally closest ECG within ±10± 10 days of the echocardiography report was selected. Pairs outside this window were excluded, together with ECGs containing corrupted waveforms, incomplete lead data, or inconsistent metadata. The patients had a mean age of 65.3±13.765.3± 13.7 years, and 9,654 were male. Compared with PKUPH, SHTMU showed a substantially different finding-prevalence profile, characterized by a high prevalence of left atrial enlargement and a low prevalence of findings such as right atrial enlargement. The cohort therefore supports geographically independent cross-center evaluation under concurrent institutional and label-prevalence shifts. Table 8. Definition-only prompts used to construct textual prototypes for prompt-based classifier-free inference on EchoNext-Mini. Each prompt describes the corresponding echocardiography-derived finding through target-defining anatomical, functional, hemodynamic, severity-related, or quantitative characteristics. Cardiac finding Text prompt Left ventricular systolic dysfunction The echocardiographic examination demonstrates impaired global left ventricular contractile function, typically characterized by a left ventricular ejection fraction of 45% or lower. Left ventricular hypertrophy The echocardiographic examination demonstrates increased left ventricular myocardial wall thickness or mass, typically involving thickening of the interventricular septum, the left ventricular posterior wall, or both. Moderate or severe aortic regurgitation The echocardiographic examination demonstrates moderate or severe diastolic regurgitant blood flow across the aortic valve from the aorta into the left ventricle, reflecting clinically significant aortic valve incompetence. Moderate or severe aortic stenosis The echocardiographic examination demonstrates moderate or severe narrowing and restricted opening of the aortic valve, resulting in hemodynamically significant obstruction of systolic blood flow from the left ventricle into the aorta, typically accompanied by increased transvalvular velocity or pressure gradient and reduced valve area. Moderate or severe tricuspid regurgitation The echocardiographic examination demonstrates moderate or severe systolic regurgitant blood flow across the tricuspid valve from the right ventricle into the right atrium, reflecting clinically significant tricuspid valve incompetence. Moderate or severe mitral regurgitation The echocardiographic examination demonstrates moderate or severe systolic regurgitant blood flow across the mitral valve from the left ventricle into the left atrium, reflecting clinically significant mitral valve incompetence. Moderate or severe pulmonic regurgitation The echocardiographic examination demonstrates moderate or severe diastolic regurgitant blood flow across the pulmonic valve from the pulmonary artery into the right ventricle, reflecting clinically significant pulmonic valve incompetence. Appendix E Additional Experimental Details E.1. Detailed Evaluation Protocols Prompt-Based Classifier-Free Inference. We freeze the pretrained ECG encoder, shared ECG projection head, text encoder, and shared text projection head, and perform inference without training a downstream classifier. Each cardiac finding is represented by a fixed definition-only prompt, which the text branch encodes into an â2 _2-normalized textual prototype. Each ECG is mapped into the normalized shared space, and its cosine similarity with each textual prototype serves as the corresponding class-wise prediction score. Class-specific thresholds are selected on the EchoNext-Mini validation set and fixed before test evaluation. Because all evaluated findings are incorporated during pretraining through standardized summaries and class-prototype supervision, this protocol evaluates the semantic accessibility of pretraining-seen findings. Generalization to unseen disease categories lies outside its scope. In-Domain Frozen Linear Probing. We freeze the pretrained ECG encoder and output projection and train a linear multi-label classifier on EchoNext-Mini using 1%, 10%, or 100% of the available training labels. This protocol evaluates the linear accessibility of the frozen representations across different label budgets. Label subsets are sampled using fixed random seeds shared across methods. Model selection and class-specific threshold determination use only the EchoNext-Mini validation split, and performance is reported on the independent patient-disjoint test set. All methods use identical label subsets, classifier architectures, optimization settings, and evaluation procedures. Target-Domain Cross-Center Frozen Linear Probing. We evaluate the adaptability of EchoNext-Mini-pretrained representations to the independent PKUPH and SHTMU cohorts. For each method, the pretrained ECG representation extractor, including the ECG encoder and output projection where applicable, is frozen. A new linear multi-label classifier is trained independently on each target cohort using 1%, 10%, or 100% of its available training labels. Label subsets are sampled using fixed random seeds shared across methods. Model selection and class-specific threshold determination use only the validation split of the corresponding target cohort, and performance is reported on its patient-disjoint test set. This protocol evaluates the linear accessibility of the frozen representations under institutional and label-distribution shifts across different levels of target-domain supervision. Source-Only Cross-Center Transfer. We evaluate direct cross-center transfer without target-domain training or calibration. A linear multi-label classifier is trained on frozen ECG representations using the complete EchoNext-Mini training set, while model selection and class-specific threshold determination use only the EchoNext-Mini validation split. The ECG encoder, output projection, linear classifier, and decision thresholds are then fixed and applied unchanged to PKUPH and SHTMU. Evaluation is restricted to cardiac findings shared between EchoNext-Mini and each target cohort. Target-domain samples are excluded from representation learning, classifier training, model selection, calibration, and threshold determination. Table 9. Hyperparameters used for EchoBridge pretraining. Hyperparameter Value Frequency-adaptive angular margin Base margin m0m_0 0.200.20 rad Minimum margin mminm_ 0.050.05 rad Maximum margin mmaxm_ 0.500.50 rad Spherical Riesz regularization Riesz exponent q 2.02.0 Regularization weight λr _r 0.050.05 Training configuration Learning rate 1Ă10â51Ă 10^-5 Batch size 6464 Max epochs 1515 Random seed 4242 Data-loader workers 88 Checkpoint interval 55 epochs E.2. Pretraining Hyperparameters Table 9 summarizes the optimization and alignment hyperparameters used for EchoBridge pretraining. Table 10. Finding-specific AUROC and AUPRC on EchoNext-Mini under 100% frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. Methods Ref. LVSD LVH Mod./Sev. AR Mod./Sev. AS Mod./Sev. TR Mod./Sev. MR Mod./Sev. PR P/N=4833/15167 P/N=4954/15046 P/N=268/19732 P/N=826/19174 P/N=2172/17828 P/N=1715/18285 P/N=154/19846 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 75.18/50.06 65.64/35.39 62.00/2.19 65.60/7.87 67.95/20.33 69.65/17.91 74.87/5.99 ST-MEM (Na et al., 2024) ICLRâ24 78.73/54.72 68.53/39.32 59.17/1.91 74.00/11.96 70.77/22.42 72.41/18.96 74.54/2.73 HeartLang (Jin et al., 2025) ICLRâ25 82.97/63.16 70.76/41.92 61.73/2.31 76.99/14.33 74.34/28.05 75.94/22.36 79.43/5.44 ECG-Text Pretraining CLIP (Radford et al., 2021) ICMLâ21 79.20/57.87 67.01/38.22 63.34/2.19 67.19/8.42 69.65/22.41 72.64/20.83 74.23/5.04 SigLIP (Zhai et al., 2023) ICCVâ23 78.86/57.46 67.14/38.24 64.18/2.13 67.08/8.30 70.40/22.64 71.89/19.89 77.25/4.41 PCME++ (Chun, 2023) ICLRâ24 76.63/53.10 67.61/39.14 64.96/2.33 68.33/8.61 68.15/20.37 71.80/19.89 72.34/3.17 MERL-ECHO (Wong et al., 2025) medRxivâ25 80.34/59.32 68.32/38.55 64.13/2.14 71.92/10.39 71.54/24.42 74.84/21.74 72.18/6.69 ECG-CLIP (Zhou et al., 2025) npj DMâ25 81.71/61.14 70.08/41.82 64.28/2.15 73.81/10.95 73.06/25.61 75.06/21.34 79.23/6.90 D-BETA (Hung et al., 2025) ICMLâ25 84.32/69.67 72.62/51.65 68.20/2.54 75.62/14.05 76.31/34.74 77.90/26.68 80.04/4.57 SGERA (Chen et al., 2026) ICMLâ26 85.28/69.56 73.51/49.82 66.84/2.44 76.73/15.77 77.92/35.58 78.31/27.82 81.72/10.74 EchoBridge Ours 86.15/70.74 74.74/50.98 69.53/3.48 78.58/16.49 79.49/37.94 79.83/30.08 83.24/18.89 Table 11. Finding-specific AUROC and AUPRC on PKUPH under 100% target-domain cross-center frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. Methods Ref. LVSD LVDD LVWMA LVH LAE LVE RAE RVE Mod./Sev. TR Mod./Sev. MR P/N=129/5303 P/N=2150/3282 P/N=52/5380 P/N=1053/4379 P/N=1554/3878 P/N=211/5221 P/N=48/5384 P/N=35/5397 P/N=26/5406 P/N=49/5383 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 84.68/19.40 63.84/53.78 90.26/13.22 60.60/29.82 59.91/37.88 77.42/19.48 72.47/2.98 71.51/1.24 69.04/1.78 76.37/5.52 ST-MEM (Na et al., 2024) ICLRâ24 88.78/30.41 65.78/57.34 91.57/17.18 63.06/31.50 62.14/40.83 81.08/26.59 77.48/4.78 79.69/9.59 79.03/3.94 79.79/6.94 HeartLang (Jin et al., 2025) ICLRâ25 88.95/29.73 66.46/56.71 91.53/17.37 63.76/31.31 62.48/40.49 81.62/26.45 77.88/4.70 78.11/10.55 79.11/3.79 79.60/7.90 ECG-Text Pretraining CLIP (Radford et al., 2021) ICMLâ21 87.13/27.84 65.35/55.96 90.44/16.97 61.94/31.03 60.51/39.50 79.08/22.86 74.09/3.64 76.45/4.88 72.89/2.69 77.42/7.13 SigLIP (Zhai et al., 2023) ICCVâ23 84.99/22.53 64.25/55.15 89.88/15.70 60.41/29.91 60.16/38.93 78.29/22.72 73.05/3.51 73.88/4.95 69.82/1.47 75.37/6.93 PCME++ (Chun, 2023) ICLRâ24 87.62/19.32 66.02/55.12 90.75/13.43 63.43/30.83 61.29/38.22 79.29/18.95 74.51/2.98 76.22/1.72 73.39/1.84 78.08/5.29 MERL-ECHO (Wong et al., 2025) medRxivâ25 91.06/25.45 66.98/56.46 92.45/15.39 64.45/31.16 62.65/39.63 82.58/23.68 77.77/4.26 80.00/6.86 79.56/3.13 81.30/6.38 ECG-CLIP (Zhou et al., 2025) npj DMâ25 90.48/33.74 66.81/56.47 92.27/18.25 64.23/31.15 62.68/39.63 82.74/27.45 78.97/4.90 81.42/12.88 79.85/3.62 80.75/8.11 D-BETA (Hung et al., 2025) ICMLâ25 89.55/34.33 66.22/56.78 91.65/18.57 63.72/31.36 62.61/40.55 81.26/27.90 78.06/4.95 80.38/11.58 80.35/4.13 81.00/8.25 SGERA (Chen et al., 2026) ICMLâ26 90.92/34.09 66.59/57.18 92.22/18.54 64.06/31.65 62.49/40.84 81.90/28.45 77.52/4.62 79.91/13.39 80.38/4.28 80.61/8.06 EchoBridge Ours 92.91/37.73 67.47/58.39 93.30/19.41 65.18/32.09 63.72/41.99 84.12/29.95 79.86/5.30 82.88/16.10 82.93/4.68 82.38/8.70 E.3. Definition-Only Prompt Construction for Classifier-Free Inference For prompt-based classifier-free inference, each echocardiography-derived finding is represented by a fixed definition-only prompt encoded by the frozen text branch as an â2 _2-normalized prototype. Because standardized pretraining summaries derive from structured labels, direct label-name prompts may create lexical overlap with training text. We instead describe each phenotype through target-defining functional, anatomical, hemodynamic, severity-related, and quantitative characteristics, excluding etiologies, associated abnormalities, and explicit ECG manifestations. GPT-5.5 Thinking generated initial candidates, which senior cardiologists reviewed and standardized according to predefined EchoNext-Mini label definitions. The final prompts in Table 8 were fixed before testing and constructed independently of validation and test performance. During inference, normalized ECG representations and textual prototypes are compared by cosine similarity to obtain class-wise scores. This protocol evaluates the semantic accessibility of pretraining-seen findings without downstream classifier training. E.4. Downstream Classifier Training For all linear-probing experiments, the pretrained ECG encoder and output projection are frozen. We train a linear multi-label classifier with sigmoid outputs using binary cross-entropy with logits. Optimization uses AdamW with a learning rate of 1Ă10â31Ă 10^-3, weight decay of 1Ă10â61Ă 10^-6, a batch size of 128, and a maximum of 30 epochs. Gradients are restricted to the classifier parameters. All methods use the same frozen-representation protocol, classifier architecture, optimizer, hyperparameters, and validation-based model selection. Table 12. Finding-specific AUROC and AUPRC on SHTMU under 100% target-domain cross-center frozen linear probing. P/N denotes the numbers of positive and negative test samples for each echocardiography-derived finding. Mod./Sev. denotes moderate-or-severe disease. Methods Ref. LVSD LVH LAE RAE Mod./Sev. AR Mod./Sev. TR Mod./Sev. MR P/N=369/3349 P/N=663/3055 P/N=2196/1522 P/N=28/3690 P/N=56/3662 P/N=121/3597 P/N=124/3594 ECG-only Self-Supervised Learning SimCLR (Chen et al., 2020) ICMLâ20 80.91/48.78 65.13/29.08 59.61/61.76 77.99/1.09 59.52/1.12 71.92/10.14 74.29/12.39 ST-MEM (Na et al., 2024) ICLRâ24 88.69/53.43 68.44/31.20 62.32/66.96 79.35/2.30 61.04/1.43 76.52/11.63 80.80/14.49 HeartLang (Jin et al., 2025) ICLRâ25 87.71/51.76 69.79/33.81 61.83/63.46 78.85/1.32 60.01/1.41 76.63/10.61 78.91/13.47 ECG-Text Pretraining CLIP (Radford et al., 2021) ICMLâ21 79.88/49.91 64.32/30.05 58.78/63.26 77.41/2.15 58.82/1.16 71.21/9.66 73.35/12.09 SigLIP (Zhai et al., 2023) ICCVâ23 74.67/46.04 61.47/27.56 56.95/58.20 76.94/2.29 58.75/1.24 67.36/8.55 69.15/10.89 PCME++ (Chun, 2023) ICLRâ24 79.62/46.12 63.69/27.65 58.35/58.70 76.47/0.85 58.16/1.04 70.81/8.36 72.96/10.58 MERL-ECHO (Wong et al., 2025) medRxivâ25 84.67/49.19 66.65/28.90 60.75/61.71 78.80/1.84 60.41/1.03 73.95/10.55 77.30/12.89 ECG-CLIP (Zhou et al., 2025) npj DMâ25 87.91/52.08 68.73/30.26 63.09/66.02 79.31/2.43 60.43/1.66 78.07/11.88 79.90/13.96 D-BETA (Hung et al., 2025) ICMLâ25 86.79/54.20 67.51/33.74 62.70/68.38 78.85/2.00 60.96/1.44 78.90/12.82 80.05/14.46 SGERA (Chen et al., 2026) ICMLâ26 89.01/53.99 69.97/32.68 63.33/69.46 78.48/1.15 59.92/1.83 79.28/11.73 80.60/14.03 EchoBridge Ours 91.23/56.40 71.05/35.07 64.01/70.79 80.23/2.81 61.87/2.27 80.40/13.38 83.16/16.02 Appendix F Additional Results F.1. In-Domain Finding-Specific Performance Table 10 evaluates findings with positive prevalence ranging from approximately 25% to below 1%. Under 100% frozen linear probing, EchoBridge achieves the highest AUROC for all seven findings and the highest AUPRC for six, extending its gains beyond frequent LVSD and LVH. Relative to the strongest baseline for each finding, its largest AUROC gains are 1.59 and 1.57 points for moderate-or-severe AS and TR, while its largest AUPRC gain is 8.15 points for PR. These results support improved discrimination across prevalence levels, particularly for low-prevalence valvular findings. F.2. Cross-Center Finding-Specific Performance Finding-Specific Performance on PKUPH Table 11 reports finding-specific performance on PKUPH under 100% target-domain cross-center frozen linear probing. EchoBridge achieves the highest AUROC and AUPRC point estimates for all ten findings. The largest gains occur for LVSD, exceeding the strongest baseline by 1.85 AUROC and 3.40 AUPRC points, followed by LVE with gains of 1.38 and 1.50 points. Improvements also cover chamber enlargement, left ventricular wall-motion abnormality, ventricular hypertrophy, and valvular findings. Moderate-or-severe TR, RVE, and RAE contain only 26, 35, and 48 positive test cases, respectively, so their AUPRC values require cautious interpretation. The consistently higher point estimates support cross-center transfer across frequent and low-prevalence findings under institutional and label-distribution shifts. Finding-Specific Performance on SHTMU Table 12 reports finding-specific performance on SHTMU under 100% target-domain cross-center frozen linear probing. EchoBridge achieves the highest AUROC and AUPRC point estimates for all seven findings, with gains over the strongest baseline of 0.68â2.36 AUROC points and 0.38â2.20 AUPRC points. Improvements span frequent LAE and low-prevalence RAE and moderate-or-severe aortic, tricuspid, and mitral regurgitation, indicating linear accessibility across chamber, ventricular, and valvular abnormalities at a second independent institution. RAE and moderate-or-severe aortic regurgitation contain only 28 and 56 positive test cases, respectively, so these estimates require cautious interpretation under severe class imbalance. Table 13. Comparison of global, shared, and private representations on EchoNext-Mini. Prompt-based classifier-free inference uses fixed definition-only textual prototypes, whereas 100% frozen linear probing trains a linear multi-label classifier using the complete labeled training set. Representation Prompt-based 100% Linear Probing AUROC AUPRC F1 AUROC AUPRC F1 Global representation w/o CSPP 73.36 24.71 30.02 77.83 31.58 35.13 Private representation â â â 75.67 26.39 31.30 Shared representation 75.73 26.79 31.83 78.79 32.66 35.82 F.3. SharedâPrivate Representation Analysis Table 13 compares the global representation from the no-CSPP variant with EchoBridgeâs shared and private representations. The shared representation performs best under both protocols. Relative to the global representation, it improves prompt-based AUROC, AUPRC, and F1 by 2.37, 2.08, and 1.81 points, respectively, and 100% frozen linear-probing performance by 0.96, 1.08, and 0.69 points. These gains indicate greater accessibility of echocardiography-derived finding information to textual prototypes and linear classifiers in the shared space. The private representation remains predictive under linear probing, showing that the auxiliary branch retains task-relevant ECG information. The shared branch exceeds it by 3.12 AUROC, 6.27 AUPRC, and 4.52 F1 points, with larger AUPRC and F1 gaps indicating better positive-finding discrimination under class imbalance. Prompt-based inference is evaluated only in the shared space because the private branch is unaligned with textual prototypes. These results support the complementary roles of both branches, while characterizing private-branch information requires further modality-specific analysis. Table 14. Comparison with fully supervised task-specific models on EchoNext-Mini using the complete training set. The supervised baselines are trained end-to-end, whereas EchoBridge uses a frozen pretrained ECG encoder and shared projection with a linear multi-label classifier. Methods AUROC AUPRC F1 ResNet-18 + BCE 77.95 [77.25, 78.70] 30.77 [29.67, 32.02] 34.66 [33.76, 36.14] ResNet-18 + ASL 79.46 [78.73, 80.17] 32.75 [31.61, 34.11] 36.39 [35.51, 37.94] ResNet-18 + Cosine BCE 79.78 [79.04, 80.48] 33.46 [32.39, 34.72] 36.60 [35.77, 38.01] EchoBridge 78.79 [77.98, 79.55] 32.66 [31.65, 33.88] 35.82 [35.02, 37.36] F.4. Comparison with Fully Supervised Models Table 14 compares EchoBridgeâs frozen representation with task-specific ECG classifiers trained end-to-end using all EchoNext-Mini training labels. All supervised baselines use the same ResNet-18 backbone and differ only in their classification heads or training objectives. ResNet-18 + BCE and ResNet-18 + ASL use linear heads optimized with binary cross-entropy and asymmetric loss, respectively. ResNet-18 + Cosine BCE uses an â2 _2-normalized cosine classifier with binary cross-entropy, providing a supervised reference for EchoBridgeâs prototype-based formulation. Among the fully supervised models, ResNet-18 + Cosine BCE achieves the highest performance, with 79.78 AUROC, 33.46 AUPRC, and 36.60 F1. With the pretrained ECG encoder and shared projection frozen, EchoBridge obtains 78.79 AUROC, 32.66 AUPRC, and 35.82 F1 using a linear classifier trained with all probe labels. These values exceed ResNet-18 + BCE by 0.84 AUROC, 1.89 AUPRC, and 1.16 F1 points, while remaining 0.99, 0.80, and 0.78 points below the strongest fully supervised baseline, respectively. This comparison contextualizes EchoBridgeâs source-domain performance under complete downstream supervision. EchoBridge provides a competitive frozen cross-modal representation while additionally supporting prompt-based classifier-free inference and source-only cross-center transfer through the same pretrained representation space.