Paper deep dive
Foundation Model for Cardiac Time Series via Masked Latent Attention
Moritz Vandenhirtz, Samuel Ruipérez-Campillo, Simon Böhi, Sonia Laguna, Irene Cannistraci, Andrea Agostini, Ece Ozkan, Thomas M. Sutter, Julia E. Vogt
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/31/2026, 1:55:52 AM
Summary
The paper introduces the Latent Attention Masked Autoencoder (LAMAE), a foundation model for ECG time series that leverages cross-lead structural redundancy through a latent attention mechanism. By modeling higher-order interactions across leads, LAMAE improves representation quality and transferability, outperforming existing independent-lead masked modeling approaches on the Mimic-IV-ECG dataset for ICD-10 code prediction.
Entities (5)
Relation Signals (3)
LAMAE → uses → Latent Attention
confidence 100% · We introduce the latent attention masked autoencoder (LAMAE) FM
LAMAE → predicts → ICD-10
confidence 95% · Our method shows strong performance in predicting ICD-10 codes
LAMAE → trainedon → MIMIC-IV-ECG
confidence 95% · We provide empirical evidence on the Mimic-IV-ECG database
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Electrocardiograms (ECGs) are among the most widely available clinical signals and play a central role in cardiovascular diagnosis. While recent foundation models (FMs) have shown promise for learning transferable ECG representations, most existing pretraining approaches treat leads as independent channels and fail to explicitly leverage their strong structural redundancy. We introduce the latent attention masked autoencoder (LAMAE) FM that directly exploits this structure by learning cross-lead connection mechanisms during self-supervised pretraining. Our approach models higher-order interactions across leads through latent attention, enabling permutation-invariant aggregation and adaptive weighting of lead-specific representations. We provide empirical evidence on the Mimic-IV-ECG database that leveraging the cross-lead connection constitutes an effective form of structural supervision, improving representation quality and transferability. Our method shows strong performance in predicting ICD-10 codes, outperforming independent-lead masked modeling and alignment-based baselines.
Tags
Links
- Source: https://arxiv.org/abs/2603.26475v1
- Canonical: https://arxiv.org/abs/2603.26475v1
Trouble viewing inline? Open PDF directly →
Full Text
35,929 characters extracted from source content.
Expand or collapse full text
Workshop on Foundation Models for Science at ICLR 2026 FOUNDATION MODEL FOR CARDIAC TIME SERIES VIA MASKED LATENT ATTENTION Moritz Vandenhirtz 1∗ , Samuel Ruiperez-Campillo 1∗ , Simon B ̈ ohi 2 , Sonia Laguna 1 , Irene Cannistraci 1 , Andrea Agostini 1 , Ece Ozkan 2 , Thomas M. Sutter 1 , Julia E. Vogt 1 1 Department of Computer Science, ETH Zurich, Switzerland 2 Department of Biomedical Engineering, University of Basel, Switzerland ABSTRACT Electrocardiograms (ECGs) are among the most widely available clinical signals and play a central role in cardiovascular diagnosis. While recent foundation mod- els (FMs) have shown promise for learning transferable ECG representations, most existing pretraining approaches treat leads as independent channels and fail to explicitly leverage their strong structural redundancy. We introduce the latent attention masked autoencoder (LAMAE) FM that directly exploits this structure by learning cross-lead connection mechanisms during self-supervised pretrain- ing. Our approach models higher-order interactions across leads through latent attention, enabling permutation-invariant aggregation and adaptive weighting of lead-specific representations. We provide empirical evidence on the Mimic-IV- ECG database that leveraging the cross-lead connection constitutes an effective form of structural supervision, improving representation quality and transferabil- ity. Our method shows strong performance in predicting ICD-10 codes, outper- forming independent-lead masked modeling and alignment-based baselines. 1INTRODUCTION Cardiovascular diseases remain among the leading causes of death worldwide (Roth et al., 2025). Clinical diagnosis and monitoring increasingly rely on multimodal data streams including imaging, clinical notes, lab tests, and physiological signals, among which the electrocardiogram (ECG) is the most ubiquitous modality due to its cost-efficiency, non-invasiveness, and mature clinical interpre- tation pipelines (Kashou et al., 2023). Its automated diagnosis has been dominated for decades by expert-crafted features coupled with classical classifiers (Liu et al., 2014; Chen et al., 2018). Over the last years, deep learning has largely shifted the field toward end-to-end learning from raw ECGs, with convolutional neural networks (CNNs) (Ribeiro et al., 2020) and recurrent models ( ̈ Ubeyli, 2010) as the predominant architectures for many clinical tasks (Sau et al., 2024; Hannun et al., 2019). More recently, foundation models (FMs) have emerged as a compelling direction to reduce reliance on expensive medical labels and enable transfer across tasks and cohorts (Moor et al., 2023; Tian et al., 2024). Yet, frontier general-purpose models still lag behind domain experts on clinical bench- marks and remain costly to adapt or deploy in practice (Khan et al., 2025). A key reason is that most pretraining pipelines remain largely oblivious to domain structure, particularly in ECG time series, where such structure is particularly explicit. This structure is not a nuisance; it is an intrinsic self-supervisory signal that modern pretraining objectives rarely exploit directly and that motivates cross-lead representations rather than independent-lead designs. Self-supervised learning (SSL) offers a scalable alternative to label-heavy supervision in medicine (Azizi et al., 2021; Moody et al., 2025; Manduchi et al., 2023). Masked autoencoders (MAEs) (He et al., 2022) have gained momentum by reconstructing missing content from sparse context, encour- aging learning robust, transferable representations. In ECG specifically, masked modeling has been explored on an independent-lead basis (Na et al., 2024) and with language-inspired tokenization schemes (Jin et al., 2024). Yet, existing methods often tokenize using lead-specific encoders or treat leads as quasi-independent “channels”, limiting their ability to learn cross-lead correspondence. ∗ Equal contribution. Correspondence to moritz.vandenhirtz@inf.ethz.ch. 1 arXiv:2603.26475v1 [cs.LG] 27 Mar 2026 Workshop on Foundation Models for Science at ICLR 2026 퐸 ... ... φ 퐸 φ 퐸 φ ... LA C C ℋ C C C C C C C C C φ 퐸 ... ... D φ 퐸 φ 퐸 φ θ D θ D θ ... ...... C ... 풙 ෝ 풙 C C C C C C C C C C C C φ LA A PRETRAINING B INFERENCE Figure 1: Framework overview. (left) Each ECG is separated into 12 leads, which are encoded sep- arately and subsequently processed jointly through a latent attention transformer. The training objec- tive is masked reconstruction. (right) The predictions are based on the latent attention’s CLS token. Leveraging the coherence of medical datasets, clinical recordings often come as structured multi- view observations that share anatomy and semantics across views. Exploiting this structure via multiview contrast, cross-modal alignment, or multitask learning (Laguna et al., 2025) can improve robustness and label efficiency in multimodal models (Mo & Liang, 2024; Pellegrini et al., 2025; Chen et al., 2024). Recently, Erlacher et al. (2025) combined multi-lead MAE reconstruction with a lead-alignment objective to enforce cross-lead consistency. While effective, such a pairwise align- ment comes with limitations (Tschannen et al., 2023), which complicates design and may under- utilize richer, higher-order relationships among leads. In contrast, attention-based aggregation is a natural fit for structured latent sets: it supports permutation-invariant processing of variable-size col- lections while learning which elements are most informative (Lee et al., 2019) and has been shown to provide interpretable, instance-weighted summaries in related weakly supervised settings (Ilse et al., 2018). In this work, we propose a multi-lead MAE FM that directly capitalizes on ECG structure by learn- ing cross-lead connection mechanisms and modeling higher-order lead interactions by integrating latent attention. Concretely, our contribution is four-fold: (i) we introduce a multi-lead MAE FM with explicit cross-lead connection learning that leverages intrinsic redundancy across leads; (i) we enhance the FM with latent attention to capture higher-order dependencies beyond pairwise align- ment; (i) we provide empirical arguments for cross-lead connection learning as a scalable form of structural supervision; and (iv) we demonstrate broad clinical and scientific translation spanning coarse ICD-based phenotyping to fine-grained disease classification. 2METHODS We assume an ECG datasetX = X (i) N i=1 , where N is the number of ECG recordings in the dataset, X (i) = x (i) l l∈L ,L is the set of leads (e.g.12 leads in our datasetL = ℓ I ,ℓ I ,ℓ I ,ℓ a V R ,ℓ a V L ,ℓ a V F ,ℓ V 1 ,ℓ V 2 ,ℓ V 3 ,ℓ V 4 ,ℓ V 5 ,ℓ V 6 for our dataset (Strodthoff et al., 2024a) or see section A. The proposed work extends the Masked Autoencoder method (He et al., 2022) to the ECG domain by introducing a self-attention module in the latent space, i.e., between the encoders E φ and the decoders D θ . We call the new module Latent Attention (LA). 2.1LATENT ATTENTION FOR MULTI-LEAD INTEGRATION Our latent attention module, inspired by Ilse et al. (2018); Lee et al. (2019), learns the correlation and shared information between different leads but is flexible enough to not having to merge information between leads in case it would be supoptimal. We design the latent attention module as a multi-head, multi-layer self-attention block using an additional CLS token similar to the ViT architecture (Kolesnikov et al., 2021). Z (i) out = LA φ (Z (i) vis ) = LA φ (M LA (Z (i) )) = LA φ (M LA (z (i) l l∈L ))(1) Similar to encoder E φ in the MAE, we apply a random mask M LA (·) to the input embeddings Z (i) , i.e., Z (i) vis = M LA (Z (i) ) with a masking ratio α LA . See section 2.2 for details on the masking process. 2 Workshop on Foundation Models for Science at ICLR 2026 Table 1: ICD-10 code prediction performance by hierarchical group under linear probing of the corresponding backbones from the studied models. Best results in bold and second best in italics. OursBaselines ICD hierarchy LAMAE LAMAE E Scratch MMVM Ind P Ind S IX0.83450.83400.77150.6771 0.8380 0.7776 IX.I05–I090.83170.82780.74080.57540.8259 0.7564 I070.87720.88020.80910.68040.8669 0.8310 I080.82720.82630.75230.67550.8235 0.7676 IX.I10–I1A0.75490.75700.70430.58610.7506 0.7092 I110.81030.81240.74910.6283 0.8163 0.7602 I130.86510.87040.81060.68260.8655 0.8194 IX.I20–I250.80460.79950.74250.61820.7939 0.7508 I200.78800.79630.71470.65740.7934 0.7324 I210.82290.84020.74650.55850.8265 0.7542 IX.I26–I280.73990.74740.69480.5785 0.7499 0.6977 IX.I30–I5A0.86240.86240.79550.6321 0.8665 0.8010 I350.79810.80500.73650.62460.7983 0.7423 I420.87720.88380.83960.71930.8731 0.8458 I440.90190.90500.84390.7591 0.9091 0.8467 I450.84070.84930.77660.66460.8392 0.7828 I460.83360.84810.76970.6562 0.8523 0.7809 I470.80510.80230.74090.6480 0.8140 0.7461 I480.88370.88610.78520.6952 0.8927 0.7933 I500.87200.87470.81660.61690.8720 0.8235 IX.I60–I690.66940.67230.62900.5473 0.6823 0.6348 IX.I70–I790.74160.74410.69420.59560.7440 0.6974 IX.I80–I890.68900.69490.63490.5731 0.6985 0.6324 IX.I95–I990.67720.67820.62240.5519 0.6943 0.6275 2.2ECG-LAMAE FM Different to previous works (Na et al., 2024; Jin et al., 2024), our FM uses per-lead encoders E φ and decoders D θ with shared weights φ and θ. The latent attention module allows the model to learn the connection between the different leads to extract more meaningful information. As in the standard MAE implementation, we only feed the visible tokens T vis to the encoders E φ , where T vis ⊆ 1,...,T are the indices of the visible patches after applying M E (·), i.e., x (i) l vis = M E (x (i) l ). We therefore have T vis = |T vis | = (1− α E )· T , where T is the total number of input tokens. Using our latent attention module, we have the following objective function L X (i) = 1 α E 1 |L| X l∈L X t/∈T vis x (i) l t − ˆ x (i) l t 2 2 ,where ˆ x (i) l = D φ (LA(Z (i) vis ) l ) and Z (i) vis = M LA (z (i) l l∈L ) is the set of all non-masked latent tokens z (i) l coming from all leads x (i) l , i.e., z (i) l = E φ (x (i) l ). This architecture supports the extraction of relevant information of each lead, combined with subsequent merging of the information in the latent attention module for a global representation captured within the CLS token. This token is then used for downstream tasks, as depicted in fig. 1 (right). 3EXPERIMENTS AND RESULTS We evaluated multi-label ICD-10 prediction from 12-lead ECGs across Chapter IX (I00–I99), spanning valvular disease, hypertensive disease, ischemic syndromes and myocardial infarction, pulmonary circulation disorders, cardiomyopathies, conduction disease, atrial fibrillation/flutter, heart failure, and vascular/cerebrovascular conditions (table 1 and appendix sections A to C). Overall, LAMAE-based models achieve strong performance across granularities, with chapter-level AUROC ≈0.85 (after fine-tuning; table 4) and competitive linear-probing results (IX: 0.834; table 3). Performance is highest for ECG-salient phenotypes, notably conduction and rhythm disor- ders (e.g., I44 fine-tuning up to 0.9097; I48 up to 0.9016) and acute myocardial infarction subtypes (I21.* often >0.93; I210 up to 0.9749), consistent with stereotyped waveform signatures (PR/QRS abnormalities, irregular rhythm, ST/T changes). In contrast, broader vascular and cerebrovascular 3 Workshop on Foundation Models for Science at ICLR 2026 LAMAELAMAE E MMVM Ind P 2 ×10 3 5 ×10 3 10 4 2 ×10 4 5 ×10 4 Number of training samples 0.50 0.55 0.60 0.65 0.70 0.75 0.80 AUROC (a) Linear Probing 2 ×10 3 5 ×10 3 10 4 2 ×10 4 5 ×10 4 Number of training samples 0.50 0.55 0.60 0.65 0.70 0.75 0.80 AUROC (b) Finetuning Figure 2: Label efficiency under finetuning. Performance curves of the macro-averaged AUROC over all 228 Chapter IX codes, as a function of the number of training studies used for finetuning. groupings (I60–I69, I80–I89, I95–I99) are harder from waveform-only inputs (fine-tuning ∼0.69– 0.72), plausibly reflecting weaker direct ECG imprint and higher label/context heterogeneity. Table 1 is supplemented by an exploration of fine-tuning performance (Tab. 4), which details the greater improvements over linear probing and highlights the benefit of the pretrained backbone for downstream adaptationm, as well as the fine-grained hierarchy results for linear probing in Tab.3. Scaling experiments show that gains from structure-informed pretraining concentrate in scarce-data regimes (Fig. 2). Under linear probing and full fine-tuning, LAMAE outperforms scratch-trained and simpler baselines most strongly at small pretraining set sizes, with gaps narrowing only in large regimes (e.g.,≳50k samples; Fig. 2). This supports latent attention as a structure-aware fusion mechanism over correlated lead projections: it can exploit redundancy to encode shared physiology, yielding more sample-efficient representations when curated cardiology datasets are limited. A broader ICD-10 benchmarking study is provided by Strodthoff et al. (2024b). While not directly comparable, our fine-tuned AUROCs are in a similar range or higher for several overlapping, ECG- identifiable codes, including IX (0.8495), I132 (0.9119), I210 (0.9632), I447 (0.9452), and AF- related subcodes such as I481 (0.8902) and I482 (0.9312) (table 3). Strodthoff et al. (2024b) likewise reports strong results for conduction/AF-related codes (e.g., I440 and AF groupings). For the global burden of atrial fibrillation (ICD48 and 48.*; (Chugh et al., 2014)), our performance is superior to prior task-specific studies that report AUROCs around 0.82–0.85 across external cohorts with CNN- based models (Brant et al., 2025), and 0.67–0.8 using demographics or N-extracted features on a 1-day ECG recording (Gadaleta et al., 2023). This is broadly consistent with AF being learnable yet sensitive to cohort shift and label timing. For conduction/heart block phenotypes related to I44 and sub-groups, reported performance varies widely across clinical settings ranging 0.594 to 0.889 (Sau et al., 2025), and our results with AUROC above 0.9 suggest that multi-lead structural pretraining can yield robust discrimination even under limited downstream data. Limitations include the imperfect nature of ICD labels as proxies for physiology and the restricted clinical context available to waveform-only models, particularly for vascular/cerebrovascular diag- noses. Nevertheless, the consistent low-data gains and strong performance on ECG-salient pheno- types indicate that explicitly leveraging cross-lead structure via latent attention is a practical route toward more transferable ECG foundation representations. 4CONCLUSION We introduced LAMAE: a multi-lead masked autoencoder FM that injects structure into ECG pretraining via latent attention over lead-specific latents. Across a broad Chapter IX ICD-10 hierarchy, LAMAE yields strong AUROC under both linear probing and fine-tuning, with the largest advantages in low-data regimes and in diagnoses where multi-lead interactions are central. These results support that exploiting cross-lead redundancy as structural supervision can improve sample efficiency and downstream transfer, offering a scalable template for time-series foundation models, potentially beyond ECG, where observations naturally come as correlated sets of views. Even more, latent attention in FMs could serve as a general template for broader applications in science and medicine wherever there is structure between measurements to be leveraged. 4 Workshop on Foundation Models for Science at ICLR 2026 ACKNOWLEDGEMENTS This work was supported under project IDs a150 and a012 as part of the Swiss AI Initiative, through a grant from the ETH Domain and computational resources provided by the Swiss Na- tional Supercomputing Centre (CSCS) under the Alps infrastructure. MV and SL are supported by the Swiss State Secretariat for Education, Research, and Innovation (SERI) under contract number MB22.00047. TS and A are supported by the grant #2021-911 of the Strategic Focal Area “Per- sonalized Health and Related Technologies (PHRT)” of the ETH Domain (Swiss Federal Institutes of Technology). REFERENCES Shekoofeh Azizi, Basil Mustafa, Fiona Ryan, Zachary Beaver, Jan Freyberg, Jonathan Deaton, Aaron Loh, Alan Karthikesalingam, Simon Kornblith, Ting Chen, et al. Big self-supervised models advance medical image classification. In Proceedings of the IEEE/CVF international conference on computer vision, p. 3478–3488, 2021. 1 Luisa C Brant, Ant ˆ onio H Ribeiro, Oseiwe B Eromosele, Marcelo M Pinto-Filho, Sandhi M Bar- reto, Bruce B Duncan, Martin G Larson, Emelia J Benjamin, Antonio LP Ribeiro, and Honghuang Lin. Prediction of atrial fibrillation from the ecg in the community using deep learning: A multi- national study. Circulation: Arrhythmia and Electrophysiology, 18(10):e013734, 2025. 4 Xiaohe Chen, Yan Wang, Lirong Wang, et al. Arrhythmia recognition and classification using ecg morphology and segment feature analysis. IEEE/ACM transactions on computational biology and bioinformatics, 16(1):131–138, 2018. 1 Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, et al. Chexagent: Towards a foundation model for chest x-ray interpretation. In AAAI 2024 Spring Symposium on Clinical Foundation Models, 2024. 2 Sumeet S Chugh, Rasmus Havmoeller, Kumar Narayanan, David Singh, Michiel Rienstra, Emelia J Benjamin, Richard F Gillum, Young-Hoon Kim, John H McAnulty Jr, Zhi-Jie Zheng, et al. World- wide epidemiology of atrial fibrillation: a global burden of disease 2010 study. Circulation, 129 (8):837–847, 2014. 4 Lucas Erlacher, Andrea Agostini, Samuel Ruiperez-Campillo, Ece Ozkan, Thomas M Sutter, and Julia E Vogt. Swissbeatsnet: A multilead masked autoencoder for chagas disease detection. Com- puting In cardiology, 15:16, 2025. 2 Matteo Gadaleta, Patrick Harrington, Eric Barnhill, Evangelos Hytopoulos, Mintu P Turakhia, Steven R Steinhubl, and Giorgio Quer. Prediction of atrial fibrillation from at-home single-lead ecg signals without arrhythmias. npj Digital Medicine, 6(1):229, 2023. 4 Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley. Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals. circulation, 101(23):e215–e220, 2000. 8 Brian Gow, Tom Pollard, Larry A Nathanson, Alistair Johnson, Benjamin Moody, Chrystinne Fer- nandes, Nathaniel Greenbaum, Jonathan W Waks, Parastou Eslami, Tanner Carbonati, et al. Mimic-iv-ecg: Diagnostic electrocardiogram matched subset. Type: dataset, 6:13–14, 2023. 8 Awni Y Hannun, Pranav Rajpurkar, Masoumeh Haghpanahi, Geoffrey H Tison, Codie Bourn, Mintu P Turakhia, and Andrew Y Ng. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nature medicine, 25(1):65–69, 2019. 1 Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll ́ ar, and Ross Girshick. Masked au- toencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16000–16009, 2022. 1, 2 5 Workshop on Foundation Models for Science at ICLR 2026 Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learn- ing. In International conference on machine learning, p. 2127–2136. PMLR, 2018. 2 Jiarui Jin, Haoyu Wang, Hongyan Li, Jun Li, Jiahui Pan, and Shenda Hong. Reading your heart: Learning ecg words and sentences via pre-training ecg language model. In International confer- ence on learning representations, 2024. 1, 3 Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. Mimic-iv, a freely accessible electronic health record dataset. Scientific data, 10(1):1, 2023. 8 Anthony H Kashou, Peter A Noseworthy, Thomas J Beckman, Nandan S Anavekar, Michael W Cullen, Kurt B Angstman, Benjamin J Sandefur, Brian P Shapiro, Brandon W Wiley, Andrew M Kates, et al. Ecg interpretation proficiency of healthcare professionals. Current problems in cardiology, 48(10):101924, 2023. 1 Wasif Khan, Seowung Leem, Kyle B See, Joshua K Wong, Shaoting Zhang, and Ruogu Fang. A comprehensive survey of foundation models in medicine. IEEE Reviews in Biomedical Engineer- ing, 2025. 1 Alexander Kolesnikov, Alexey Dosovitskiy, Dirk Weissenborn, Georg Heigold, Jakob Uszkoreit, Lucas Beyer, Matthias Minderer, Mostafa Dehghani, Neil Houlsby, Sylvain Gelly, Thomas Un- terthiner, and Xiaohua Zhai. An image is worth 16x16 words: Transformers for image recognition at scale. 2021. 2 Sonia Laguna, Andrea Agostini, Alain Ryser, Samuel Ruiperez-Campillo, Irene Cannistraci, Moritz Vandenhirtz, Stephan Mandt, Nicolas Deperrois, Farhad Nooralahzadeh, Michael Krauthammer, et al. Structure is supervision: Multiview masked autoencoders for radiology. arXiv preprint arXiv:2511.22294, 2025. 2 Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. Set transformer: A framework for attention-based permutation-invariant neural networks. In Interna- tional conference on machine learning, p. 3744–3753. PMLR, 2019. 2 Yun Liu, Zeeshan Syed, Benjamin M Scirica, David A Morrow, John V Guttag, and Collin M Stultz. Ecg morphological variability in beat space for risk stratification after acute coronary syndrome. Journal of the American Heart Association, 3(3):e000981, 2014. 1 Laura Manduchi, Moritz Vandenhirtz, Alain Ryser, and Julia Vogt. Tree variational autoencoders. Advances in Neural Information Processing Systems, 36:54952–54986, 2023. 1 Shentong Mo and Paul Pu Liang. Multimed: Massively multimodal and multitask medical under- standing. arXiv preprint arXiv:2408.12682, 2024. 2 Jonathan B Moody, Alexis Poitrasson-Rivi ` ere, Jennifer M Renaud, Tomoe Hagio, Fares Alahdab, Mouaz H Al-Mallah, Michael D Vanderver, Sascha N Goonewardena, Edward P Ficaro, and Venkatesh L Murthy. A foundation transformer model with self-supervised learning for ecg-based assessment of cardiac and coronary function. NEJM AI, 2(12):AIoa2500164, 2025. 1 Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. Foundation models for generalist medical artificial intelli- gence. Nature, 616(7956):259–265, 2023. 1 Yeongyeon Na, Minje Park, Yunwon Tae, and Sunghoon Joo. Guiding masked representation learn- ing to capture spatio-temporal relationship of electrocardiogram. In International conference on learning representations, 2024. 1, 3 Chantal Pellegrini, Ege ̈ Ozsoy, Benjamin Busam, Benedikt Wiestler, Nassir Navab, and Matthias Keicher. Radialog: Large vision-language models for x-ray reporting and dialog-driven assis- tance. In Medical Imaging with Deep Learning, 2025. 2 6 Workshop on Foundation Models for Science at ICLR 2026 Ant ˆ onio H Ribeiro, Manoel Horta Ribeiro, Gabriela M Paix ̃ ao, Derick M Oliveira, Paulo R Gomes, J ́ essica A Canazart, Milton PS Ferreira, Carl R Andersson, Peter W Macfarlane, Wagner Meira Jr, et al. Automatic diagnosis of the 12-lead ecg using a deep neural network. Nature communications, 11(1):1760, 2020. 1 Gregory A. Roth, Global Burden of Cardiovascular Diseases, and Risks 2023 Collaborators. Global, regional, and national burden of cardiovascular diseases and risk factors in 204 countries and territories, 1990-2023. Journal of the American College of Cardiology, 86(22):2167–2243, 2025. 1 Arunashis Sau, Libor Pastika, Ewa Sieliwonczyk, Konstantinos Patlatzoglou, Antonio H Ribeiro, Kathryn A Mcgurk, Boroumand Zeidaabadi, Henry Zhang, Krzysztof Macierzanka, Danilo Mandic, et al. Artificial intelligence-enabled electrocardiogram for mortality and cardiovascu- lar risk estimation: a model development and validation study. The Lancet Digital Health, 6(11): e791–e802, 2024. 1 Arunashis Sau, Henry Zhang, Joseph Barker, Libor Pastika, Konstantinos Patlatzoglou, Boroumand Zeidaabadi, Ahmed El-Medany, Gul Rukh Khattak, Kathryn A McGurk, Ewa Sieliwonczyk, et al. Artificial intelligence–enhanced electrocardiography for complete heart block risk stratification. JAMA cardiology, 10(11):1092–1099, 2025. 4 Nils Strodthoff, JM Lopez Alcaraz, and W Haverkamp IV. Mimic-iv-ecg-ext-icd: Diagnostic labels for mimic-iv-ecg (version 1.0. 1). PhysioNet. RRID: SCR 007345 https://doi. org/10.13026/hdyc- 1h77, 2024a. 2, 8 Nils Strodthoff, Juan Miguel Lopez Alcaraz, and Wilhelm Haverkamp. Prospects for artificial intelligence-enhanced electrocardiogram as a unified screening tool for cardiac and non-cardiac conditions: an explorative study in emergency care. European Heart Journal-Digital Health, 5 (4):454–460, 2024b. 4, 8 Yuanyuan Tian, Zhiyuan Li, Yanrui Jin, Mengxiao Wang, Xiaoyang Wei, Liqun Zhao, Yunqing Liu, Jinlei Liu, and Chengliang Liu. Foundation model of ecg diagnosis: Diagnostics and explanations of any form and rhythm on ecg. Cell Reports Medicine, 5(12), 2024. 1 Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, and Lucas Beyer. Image captioners are scalable vision learners too. Advances in Neural Information Processing Systems, 36:46830–46855, 2023. 2 Elif Derya ̈ Ubeyli. Recurrent neural networks employing lyapunov exponents for analysis of ecg signals. Expert systems with applications, 37(2):1192–1199, 2010. 1 7 Workshop on Foundation Models for Science at ICLR 2026 AMATERIALS We conducted experiments on the MIMIC-IV-ECG-Ext-ICD resource (Strodthoff et al., 2024a), a PhysioNet (Goldberger et al., 2000)release that links raw 12-lead ECG waveforms from MIMIC-IV- ECG (Gow et al., 2023) to clinically grounded diagnostic labels from the corresponding MIMIC-IV Johnson et al. (2023) emergency department and inpatient records. Concretely, ECG acquisition timestamps are aligned with ED stays and hospital admissions to associate each recording with discharge diagnosis codes, providing ICD-10-CM label sets derived from routine clinical docu- mentation rather than retrospective re-annotation. The dataset includes identifiers to retrieve addi- tional clinical context (e.g., ED stay and hospital admission IDs), basic demographics (e.g., age- at-recording, sex), and fold assignments designed to avoid patient overlap for benchmarking and comparability across studies (Strodthoff et al., 2024a). In our study, only ECG raw waveforms and their paired ICD-10 code were used. Following the benchmark framing introduced by Strodthoff et al. (2024b), we treat ICD-10-CM codes as multi-label targets at multiple granularities (chapter/block/category/subcategory), enabling evaluation from coarse phenotyping to fine-grained diagnosis. Where needed for consistency across label hierarchies, ICD codes may be normalized to a fixed digit format and expanded to include higher-level ancestors in the ICD tree, supporting hierarchical reporting and clinically meaningful aggregation (Strodthoff et al., 2024a;b). BON ICD CODES AND FURTHER DETAILS ON THOSE USED IN THIS STUDY International Classification of Diseases (ICD) codes provide a standardized taxonomy for clinical di- agnoses and are routinely used for billing, cohort definition, and large-scale observational research. In this work, we focus on ICD-10 Chapter IX (Diseases of the circulatory system; I00–I99), and report predictive performance at multiple levels of granularity: (i) the chapter-level aggregate, (i) chapter blocks (e.g., I05–I09), (i) 3-character categories (e.g., I07), and (iv) selected 4-character subcategories (e.g., I07.1; written as I071). This hierarchical evaluation reflects clinically mean- ingful groupings while enabling finer assessment of model behaviour on specific diagnoses. An extended description of the clinical meaning per code is included in table 2. CSUPPLEMENTARY RESULTS: FINE-GRAINED ANALYSIS Fine-grained Classification Results. In table 3, we present an extended analysis of the fine- grained classification performance initially discussed in table 1 on linear probing. The results demonstrate that performance trends remain remarkably consistent across the ICD-10 hierarchy. Notably, the relative advantages of our proposed methods are preserved even as the classification task becomes more granular, confirming the robustness of the learned representations. Performance of LAMAE under Varying Supervision. Table 4 compares the performance of LAMAE and LAMAE E across both linear probing and fine-tuning regimes. While both models achieve competitive results under linear probing, fine-tuning LAMAE yields a significantly larger performance gain. This suggests that the pretrained backbone serves as a powerful initialization that can be further leveraged to maximize predictive accuracy when labeled data allows for full model updates. 8 Workshop on Foundation Models for Science at ICLR 2026 Table 2: Clinical meaning of the ICD-10 Chapter IX codes reported in tables 1 and 3, shown with the same hierarchy. ICD hierarchyClinical description IXDiseases of the circulatory system (I00–I99). IX.I05–I09Chronic rheumatic heart diseases (I05–I09). I07Rheumatic tricuspid valve diseases. I071I07.1 — Tricuspid (valve) insufficiency (rheumatic). I078I07.8 — Other tricuspid valve diseases. I08Multiple valve diseases (often rheumatic or unspecified origin). I080I08.0 — Disorders of both mitral and aortic valves. I081I08.1 — Disorders of both mitral and tricuspid valves. I083I08.3 — Combined disorders of mitral, aortic and tricuspid valves. IX.I10–I1AHypertensive diseases (I10–I15). I11Hypertensive heart disease. I13Hypertensive heart and renal disease. I130I13.0 — Hypertensive heart and renal disease with (congestive) heart failure. I132I13.2 — Hypertensive heart and renal disease with both (congestive) heart failure and renal failure. IX.I20–I25Ischaemic heart diseases (I20–I25). I20Angina pectoris. I200I20.0 — Unstable angina. I209I20.9 — Angina pectoris, unspecified. I21Acute myocardial infarction. I210I21.0 — Acute transmural myocardial infarction of anterior wall. I211I21.1 — Acute transmural myocardial infarction of inferior wall. I213I21.3 — Acute transmural myocardial infarction of unspecified site. I214I21.4 — Acute subendocardial myocardial infarction. IX.I26–I28Pulmonary heart disease and diseases of pulmonary circulation (I26– I28). IX.I30–I5AOther forms of heart disease (ICD block: I30–I52; reported here as I30–I5A). I35Nonrheumatic aortic valve disorders. I350I35.0 — Aortic (valve) stenosis. I359I35.9 — Aortic valve disorder, unspecified. I42Cardiomyopathy. I420I42.0 — Dilated cardiomyopathy. I428I42.8 — Other cardiomyopathies. I429I42.9 — Cardiomyopathy, unspecified. I44Atrioventricular and left bundle-branch block. I440I44.0 — Atrioventricular block, first degree. I441I44.1 — Atrioventricular block, second degree. I442I44.2 — Atrioventricular block, complete. I447I44.7 — Left bundle-branch block, unspecified. I45Other conduction disorders. I46Cardiac arrest. I47Paroxysmal tachycardia. I48Atrial fibrillation and flutter. I480I48.0 — Paroxysmal atrial fibrillation. I481I48.1 — Persistent atrial fibrillation. I482I48.2 — Chronic atrial fibrillation. I50Heart failure. IX.I60–I69Cerebrovascular diseases (I60–I69). IX.I70–I79Diseases of arteries, arterioles and capillaries (I70–I79). IX.I80–I89Diseases of veins, lymphatic vessels and lymph nodes, not elsewhere classified (I80–I89). IX.I95–I99Other and unspecified disorders of the circulatory system (I95–I99). 9 Workshop on Foundation Models for Science at ICLR 2026 Table 3: Extended ICD-10 code prediction performance by hierarchical group in the models studied in table 1. AUROC is reported for linear probing on each corresponding backbone. OursBaselines ICD hierarchy LAMAE LAMAE E Scratch MVMAE Ind P Ind S IX0.83450.83400.77150.6771 0.8380 0.7776 IX.I05–I090.83170.82780.74080.57540.8259 0.7564 I070.87720.88020.80910.68040.8669 0.8310 I0710.88300.87400.82360.64500.8609 0.8585 I0780.88180.89770.81100.51450.8896 0.8345 I080.82720.82350.75230.67550.8263 0.7676 I0800.83450.83180.74180.6619 0.8390 0.7513 I0810.87040.86950.77190.72720.8556 0.7992 I0830.88520.85730.84210.73110.8847 0.8742 IX.I10–I1A0.75490.75700.70430.58610.7506 0.7092 I110.81030.81240.74910.6283 0.8163 0.7602 I130.86510.87040.81060.68260.8655 0.8194 I1300.86290.86870.81830.75660.8668 0.8254 I1320.90090.90960.78780.6384 0.9171 0.8089 IX.I20–I250.79950.80460.74250.61820.7939 0.7508 I200.78800.79630.71470.65740.7934 0.7324 I2000.82360.82690.73660.67160.8267 0.7522 I2090.79250.80750.72240.67960.7870 0.7414 I210.82290.84020.74650.55850.8265 0.7542 I2100.93300.96320.91970.62760.9343 0.9143 I2110.93700.95880.80490.66040.9269 0.8433 I2130.86490.89960.80970.63870.8814 0.8010 I2140.80430.81170.73500.60180.8063 0.7397 IX.I26–I280.73990.74740.69480.5785 0.7499 0.6977 IX.I30–I5A0.86240.86240.79550.6321 0.8665 0.8010 I350.79810.80500.73650.62460.7983 0.7423 I3500.82660.83460.76720.61120.8268 0.7784 I3590.84440.85250.73970.6020 0.8585 0.7490 I420.87720.88380.83960.71930.8731 0.8458 I4200.89140.89860.84630.7816 0.9198 0.8590 I4280.88510.89210.83110.71440.8807 0.8392 I4290.85860.86110.80200.6408 0.8791 0.8116 I440.90190.90500.84390.7591 0.9091 0.8467 I4400.89840.91020.75450.6638 0.9156 0.7610 I4410.88600.90210.82290.6309 0.9147 0.8338 I4420.91130.91510.84980.7564 0.9194 0.8551 I4470.93890.94170.92750.86330.9397 0.9305 I450.84070.84930.77660.66460.8392 0.7828 I460.83360.84810.76970.6562 0.8523 0.7809 I470.80510.80230.74090.6480 0.8140 0.7461 I480.88370.88610.78520.6952 0.8927 0.7933 I4800.80590.81350.73010.6735 0.8196 0.7290 I4810.88090.88720.80910.7141 0.8875 0.8152 I4820.92280.92060.80400.7169 0.9345 0.8216 I500.87200.87470.81660.61690.8720 0.8235 IX.I60–I690.66940.67230.62900.5473 0.6823 0.6348 IX.I70–I790.74160.74410.69420.59560.7440 0.6974 IX.I80–I890.68900.69490.63490.5731 0.6985 0.6324 IX.I95–I990.67720.67820.62240.5519 0.6943 0.6275 10 Workshop on Foundation Models for Science at ICLR 2026 Table 4: Extended ICD-10 code prediction performance by hierarchical group. Comparison of pro- posed LAMAE and LAMAE E under linear probing and fine-tuning of the corresponding pretrained backbone. Linear ProbingFine-Tuning ICD hierarchy LAMAE LAMAE E LAMAE LAMAE E IX0.83450.83400.84950.8486 IX.I05–I090.83170.82780.85090.8452 I070.87720.88020.89040.9039 I0710.88300.87400.90440.8763 I0780.88180.89770.91190.9099 I080.82720.82350.84900.8420 I0800.83450.83180.85390.8425 I0810.87040.86950.89360.8730 I083 0.88520.85730.87940.9163 IX.I10–I1A0.75490.75700.76900.7678 I110.81030.81240.83430.8272 I130.86510.87040.88300.8798 I1300.86290.86870.88170.8812 I1320.90090.90960.91190.9022 IX.I20–I250.79950.80460.82820.8250 I200.78800.79630.80400.8183 I2000.82360.82690.84450.8440 I2090.79250.80750.81970.8217 I210.82290.84020.87730.8600 I2100.93300.96320.96150.9631 I2110.93700.95880.97490.9681 I2130.86490.89960.92480.9043 I2140.80430.81170.85840.8429 IX.I26–I280.73990.74740.76170.7682 IX.I30–I5A0.86240.86240.87710.8759 I350.79810.80500.82440.8180 I3500.82660.83460.85280.8644 I3590.84440.85250.86810.8657 I420.87720.88380.88800.8805 I4200.89140.89860.91700.8935 I4280.88510.89210.89580.8932 I4290.85860.86110.88090.8752 I440.90190.90500.90970.9053 I4400.89840.91020.89220.9044 I4410.88600.90210.90520.8800 I4420.91130.91510.91840.9095 I4470.93890.94170.94520.9410 I450.84070.84930.85190.8442 I460.83360.84810.86130.8634 I470.80510.80230.81060.8123 I480.88370.88610.90160.8993 I4800.80590.81350.82010.8254 I4810.88090.88720.89020.8828 I482 0.92280.92060.93120.9285 I500.87200.87470.88920.8857 IX.I60–I690.66940.67230.68360.6841 IX.I70–I790.74160.74410.75450.7456 IX.I80–I890.68900.69490.71210.7156 IX.I95–I990.67720.67820.69630.6939 11