Paper deep dive
Position: Evaluation of ECG Representations Must Be Fixed
Zachary Berger, Daniel Prakah-Asante, John Guttag, Collin M. Stultz
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/20/2026, 11:43:03 PM
Summary
This position paper argues that current benchmarking practices in 12-lead ECG representation learning are insufficient and misleading. It critiques the reliance on three standard benchmarks (PTB-XL, CPSC2018, CSN) focused on arrhythmia and waveform morphology, arguing for an expansion to include structural heart disease, hemodynamic inference, and patient forecasting. The authors propose evaluation best practices, such as reporting task-specific metrics (AUROC, AUPRC) with confidence intervals and excluding imbalanced labels with insufficient support. Empirically, they demonstrate that applying these best practices alters method rankings and reveal that a randomly initialized encoder with linear evaluation matches state-of-the-art pre-trained models on many tasks, suggesting the random encoder should be used as a baseline.
Entities (16)
Relation Signals (10)
Current Benchmarking Practice β critiquedby β Position Paper
confidence 95% Β· This position paper argues that current benchmarking practice in 12-lead ECG representation learning must be fixed
D-BETA β evaluatedin β Empirical Study
confidence 95% Β· We evaluate downstream performance using representations from five pre-trained ECG encoders: ... and D-BETA
CLOCS β evaluatedin β Empirical Study
confidence 95% Β· We evaluate downstream performance using representations from five pre-trained ECG encoders: CLOCS...
KED β evaluatedin β Empirical Study
confidence 95% Β· We evaluate downstream performance using representations from five pre-trained ECG encoders: ... KED ...
HeartLang β evaluatedin β Empirical Study
confidence 95% Β· We evaluate downstream performance using representations from five pre-trained ECG encoders: ... HeartLang ...
MERL β evaluatedin β Empirical Study
confidence 95% Β· We evaluate downstream performance using representations from five pre-trained ECG encoders: ... MERL ...
EchoNext β supports β Structural Heart Disease Assessment
confidence 95% Β· EchoNext is an open-source dataset of paired ECG and echocardiogram findings containing relevant labels
Random Encoder β β
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This position paper argues that current benchmarking practice in 12-lead ECG representation learning must be fixed to ensure progress is reliable and aligned with clinically meaningful objectives. The field has largely converged on three public multi-label benchmarks (PTB-XL, CPSC2018, CSN) dominated by arrhythmia and waveform-morphology labels, even though the ECG is known to encode substantially broader clinical information. We argue that downstream evaluation should expand to include an assessment of structural heart disease and patient-level forecasting, in addition to other evolving ECG-related endpoints, as relevant clinical targets. Next, we outline evaluation best practices for multi-label, imbalanced settings, and show that when they are applied, the literature's current conclusion about which representations perform best is altered. Furthermore, we demonstrate the surprising result that a randomly initialized encoder with linear evaluation matches state-of-the-art pre-training on many tasks. This motivates the use of a random encoder as a reasonable baseline model. We substantiate our observations with an empirical evaluation of five representative ECG pre-training approaches across six evaluation settings: the three standard benchmarks, a structural disease dataset, hemodynamic inference, and patient forecasting.
Tags
Links
- Source: https://arxiv.org/abs/2602.17531v2
- Canonical: https://arxiv.org/abs/2602.17531v2
Trouble viewing inline? Open PDF directly β
Full Text
139,999 characters extracted from source content.
Expand or collapse full text
Position: Evaluation of ECG Representations Must Be Fixed Zachary Berger 1 2 * Daniel Prakah-Asante 1 2 * John Guttag 1 Collin M. Stultz 1 2 Abstract This position paper argues that current benchmark- ing practice in 12-lead ECG representation learn- ing must be fixed to ensure progress is reliable and aligned with clinically meaningful objectives. The field has largely converged on three public multi-label benchmarks (PTB-XL, CPSC2018, CSN) dominated by arrhythmia and waveform- morphology labels, even though the ECG is known to encode substantially broader clinical information. We argue that downstream evalua- tion should expand to include an assessment of structural heart disease and patient-level forecast- ing, in addition to other evolving ECG-related endpoints, as relevant clinical targets. Next, we outline evaluation best practices for multi-label, imbalanced settings, and show that when they are applied, the literatureβs current conclusion about which representations perform best is al- tered. Furthermore, we demonstrate the surpris- ing result that a randomly initialized encoder with linear evaluation matches state-of-the-art pre- training on many tasks. This motivates the use of a random encoder as a reasonable baseline model. We substantiate our observations with an empirical evaluation of five representative ECG pre-training approaches across six evaluation set- tings: the three standard benchmarks, a structural disease dataset, hemodynamic inference, and pa- tient forecasting. Code is available athttps: //github.com/zackeberger/ecg-fix. 1. Introduction Representation learning aims to produce features that are useful across many downstream applications. This reduces reliance on large labeled datasets to learn task-specific mod- * Equal contribution 1 Massachusetts Institute of Technol- ogy, Cambridge, MA, USA 2 Massachusetts General Hospi- tal, Boston, MA, USA. Correspondence to: Zachary Berger <zberger@mit.edu>. Proceedings of the43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). els (Bengio et al., 2013). However, there is a tension be- tween learning features that are broadly useful versus those that excel for particular tasks (Bommasani et al., 2022). This tension is pronounced in medicine, where labeled data can be sparse and clinical use-cases varied (Esteva et al., 2019). In medicine, what constitutes a meaningful endpoint task depends on the modality. For example, the clinical targets of chest X-ray differ from those of electroencephalography. As a result, benchmarking practice in medicine must be discussed on a per-modality basis. These choices shape which representations appear effective and steer subsequent methodological development (Lipton & Steinhardt, 2019; Pineau et al., 2021). Here, we focus on the 12-lead electrocardiogram (ECG), an inexpensive and commonly used diagnostic tool that non- invasively records the heartβs electrical activity (Noble et al., 1990). ECG representation learning is an active area of research and has recently received attention in major ML venues (e.g., ICML, ICLR, NeurIPS, AAAI) (Liu et al., 2024b; Hung et al., 2024; Wang et al., 2025; Na et al., 2024; Jin et al., 2025; Chen et al., 2025; Lan et al., 2022). As the lit- erature grows, we argue it is timely to re-examine the fieldβs current benchmarking practice. A rigorous and uniform benchmarking strategy is essential to ensure reliable and reproducible results, and most importantly that the learned representations align with clinically meaningful objectives. ECG benchmarking is associated with a number of chal- lenges, including the use of large multi-label and severely imbalanced datasets β issues common across many applied machine learning (ML) settings, and especially medicine (Zhang & Zhou, 2014; He & Garcia, 2009). This position paper examines how current task selection and reporting practice shape conclusions about ECG represen- tation quality. We then evaluate these choices empirically across five representative pre-training methods and six eval- uation settings. In Section 2, we observe that three datasets have become standard for evaluating 12-lead ECG representations: PTB- XL (Wagner et al., 2020), CPSC2018 (Liu et al., 2018), and CSN (Zheng et al., 2020). Each is a multi-label suite focused primarily on binary arrhythmia outcomes and wave- form morphology classification. However, the ECG has increasingly been recognized to contain information perti- 1 arXiv:2602.17531v2 [cs.LG] 29 May 2026 Position: Evaluation of ECG Representations Must Be Fixed We find no method prevails; performance sometimes overlaps random baseline a) Current Evaluation 12L ECG Pre-trained Encoder Current evaluation misleadingly suggests a prevailing method Tasks β’ Arrhythmia + waveform Protocol β’ Macro-AUROC β’ Point estimates β’ Evaluate tasks with insufficient support Embedding b) Proposed Evaluation Tasks β’ Arrhythmia + waveform β’ Structural disease β’ Hemodynamics β’ Patient forecasting Protocol β’ Task-specific AUCs β’ Quantify uncertainty β’ Standard baseline β’ Exclude tasks with low test-support Figure 1. Overview of the evaluation pipeline for 12-lead ECG representations. (a) Current practice focuses on arrhythmia/waveform tasks and macro-AUROC point estimates, which can produce misleading method rankings. (b) We propose a broader set of clinically relevant tasks and evaluation best-practices that more reliably assess methods. We find that no method consistently prevails and for many tasks, many methods overlap with the baseline of a randomly initialized encoder. nent to a wider variety of outcomes, e.g., structural disease (Poterucha et al., 2025), hemodynamic state (Schlesinger et al., 2022), and patient forecasting (Khurshid et al., 2022; Bergamaschi et al., 2025). In Section 3, we propose addi- tional tasks that would be a welcome addition in the evalua- tion pipeline. This is particularly relevant as the field moves towards more complex tasks, where ECGs are combined with other modalities to predict clinical outcomes that are not deterministic functions of the ECG alone. To manage multi-label evaluation across many tasks, the field has largely converged on summarizing performance with macro-AUROC, which aggregates per-label AUROCs through an unweighted mean (Zhang & Zhou, 2014). This convention enables straightforward comparison across meth- ods, but obscures clinically meaningful behavior. Clinicians ultimately deploy models for specific purposes, yet macro- AUROC masks performance on individual endpoints by collapsing them into a single number. This issue is com- pounded by severe label imbalance; many ECG labels have few positive examples, yielding noisy task-level estimates. However, uncertainty is seldom reported alongside headline metrics. Moreover, AUROC alone can misrepresent perfor- mance on imbalanced labels, where alternative metrics may better reflect clinical utility (Davis & Goadrich, 2006; Saito & Rehmsmeier, 2015). In Section 4, we suggest a set of reporting and evaluation best practices and show that many prior studies do not adhere to them. In Section 5, we show that applying these practices can change method rankings on standard benchmarks and alter conclusions of the current literature about which methods perform best. We also show that a randomly initialized encoder with linear evaluation matches the performance of state-of-the-art ECG pre-training methods on many tasks. Our position is visualized in Figure 1. We argue that going forward, evaluation of representations of ECG should β’Cover a wider variety of tasks than is currently typical. For example, they should include prediction of patient outcomes or estimates of structural disease rather than just arrhythmia and waveform classification. β’ Report clinically relevant task-specific findings rather than focus, as most papers do, on aggregate perfor- mance. For example, they should report on AUROC, precision, and recall for individual tasks rather than just macro-AUROC over classes of tasks. β’Carefully characterize uncertainty, which can be quite high for tasks with a small number of positive exam- ples, which is common in ECG datasets. β’Compare the utility of learned representations to that of a simple baseline: a randomly initialized encoder. 2. Related Work Benchmarking in ML. Progress in empirical ML has long been driven by benchmarks (e.g., Geiger et al., 2012; Lin et al., 2014; Russakovsky et al., 2015; Wang et al., 2018). However, evaluation practices often lack rigor and standard- ization (Lipton & Steinhardt, 2019; Liao et al., 2021; Her- rmann et al., 2024). This has motivated work that critically examines and improves benchmark design and reporting (e.g., Gebru et al., 2021; Pineau et al., 2021; Vendrow et al., 2025). Many lessons have emerged from domain-specific settings, for instance, in recommendation systems (Fer- rari Dacrema et al., 2019), neural network pruning (Blalock et al., 2020), anomaly detection (Liu & Paparrizos, 2024), and graph learning (Bechler-Speicher et al., 2025). 2 Position: Evaluation of ECG Representations Must Be Fixed ECG Benchmarking. ECGs are routinely collected in clinical care, so large labeled corpora exist in many health systems. Yet, these data are rarely shared because of patient privacy constraints and institutional requirements. As a result, despite the volume of ECGs that exist, there are few open-source datasets. In response, the community has largely repurposed the avail- able public datasets for downstream benchmarking of ECG representations. Evaluations typically center around three datasets, CPSC2018 (Liu et al., 2018), PTB-XL (Wagner et al., 2020; 2022), and CSN (Zheng et al., 2020; 2022), which contain on the order of tens of thousands of record- ings. All three focus narrowly on arrhythmia and waveform abnormality labels. Several other datasets with similar la- bels are occasionally used for testing, or aggregated together for training (e.g., Perez Alday et al., 2022; Liu et al., 2022; Ribeiro et al., 2020). MIMIC-IV (Gow et al., 2023; Goldberger et al., 2000) and CODE-15 (Ribeiro et al., 2020) are commonly used for pre- training, since they are large open-source datasets. MIMIC- IV is of particular importance because its ECGs are linked to electronic health record data. This has enabled develop- ment of learning algorithms that incorporate clinical context through multi-modal supervision. Such approaches have recently garnered state-of-the-art performance (Liu et al., 2024b; Hung et al., 2024). Many studies pre-train on pri- vate institutional data (e.g., Diamant et al., 2022) or semi- restricted resources (e.g., Sudlow et al., 2015; Littlejohns et al., 2020; Koscova et al., 2024). While these datasets can be valuable for scaling up training data, their restricted access limits reproducibility. There are few open-access datasets that expand beyond arrhythmia and waveform abnormality labels.A no- table recent release is EchoNext, which links ECGs to echocardiography-derived ground truth to study structural heart disease (Poterucha et al., 2025). Several preprints contemporaneous to this work have taken initial steps toward improving benchmarking for ECG rep- resentation learning (Lunelli et al., 2025; Al-Masud et al., 2025; Wan et al., 2025). These efforts each aim to con- solidate evaluation practices in an open-source framework. However, they do not implement all of the protocol rec- ommendations we discuss in Section 4. They also do not capture the breadth of clinically grounded tasks we argue the field should work toward in Section 3. ECG Representation Learning.Self-supervised pre- training is a standard paradigm, driven by successes in computer vision (Chen et al., 2020) and natural language processing (Devlin et al., 2019). Early ECG representation learning (Kiyasseh et al., 2021; Mehari & Strodthoff, 2022) adapted popular vision frameworks such as SimCLR (Chen et al., 2020) and BYOL (Grill et al., 2020). These are now widely used as baselines in the ECG literature. Many specialized methods have introduced inductive biases specific to 12-lead ECGs. These often take inspiration from contrastive learning and reconstruction-based learning. Con- trastive methods learn representations by bringing together related samples while maximizing the distance between un- related samples. These positive and negative pairs can be generated through augmentation (Chen et al., 2020; Grill et al., 2020; Chen & He, 2020; Chen et al., 2021) or by relying on information such as patient identity (Diamant et al., 2022). CLOCS (Kiyasseh et al., 2021) is a popular ECG-specific approach that builds pairs directly from the temporal and lead-structure of the signal. Reconstruction- based methods optimize representations by compressing then reconstructing examples directly, often by masking then filling in part of the signal (He et al., 2021; Na et al., 2024; Zhang et al., 2023a;b). HeartLang (Jin et al., 2025) is a recent example that leverages an ECG-specific tokenizer and trains using both reconstruction and masked-token pre- diction. In parallel, recent work uses multi-modal supervision, com- monly pairing ECGs with clinical text from the electronic health record (Lalam et al., 2023; Yu et al., 2024). MERL (Liu et al., 2024b), D-BETA (Hung et al., 2024), and KED (Tian et al., 2024) are recent examples that have claimed state-of-the-art performance following this ap- proach. MERL aligns ECG embeddings with representa- tions of their paired text reports using a contrastive objective. D-BETA extends this work by regularizing the learning pro- cess with reconstruction loss on the text and ECG. KED enriches the text supervision with clinical knowledge gener- ated by a large language model. In this paper, we characterize the state of benchmarking for ECG representation learning by surveying work published at major ML conferences since 2019. We also include ap- proaches cited by those publications. When making broad claims about the field, we refer to this survey set of 28 meth- ods; details are provided in Appendix A. We note that many more pre-trained ECG models have been proposed, some of which are covered in the review of (Han et al., 2025). For our empirical study, we focus on CLOCS, KED, Heart- Lang, MERL, and D-BETA as exemplar methods. They span several prominent paradigms in recent ECG representa- tion learning, including contrastive learning, reconstruction, and multi-modal ECG-text supervision. CLOCS and MERL are widely cited, while HeartLang, KED, and D-BETA are recent works reporting state-of-the-art performance. All five are supported by released weights or pre-training code. 3 Position: Evaluation of ECG Representations Must Be Fixed 3. Extending Current Benchmarks Historically, ECG interpretation has centered on rhythm and waveform abnormalities (Rivera-Ruiz et al., 2008; Fisch, 2000). As a result, downstream evaluation of ECG represen- tations has largely focused on such tasks. In our survey of 28 ECG representation learning papers, 23 report results on PTB-XL, 15 on CPSC2018, 11 on CSN or its constituent datasets (Chapman-Shaoxing and Ningbo), with 25 evaluat- ing on at least one of the three (see Table 8). The ground-truth labels for arrhythmia and waveform abnor- mality tasks are obtained from inspection of an ECG by an expert. Thus, all information needed to assign the label is explicitly contained in the signal. Yet, it has recently been shown that machine-learned models can be applied to the ECG to infer clinically relevant endpoints that are not read- ily visible in the signal (Friedman et al., 2025). Downstream evaluation of ECG representations should expand to better reflect this clinical scope. We propose the following families of downstream tasks be tested in evaluations. This categorization is motivated by the distinction between tasks whose labels are assigned directly from the ECG trace and tasks whose ground truth is obtained from paired measurements, such as imaging and catheterization, or from future clinical outcomes. 1.Arrhythmia and Waveform Abnormalities include tasks derived from expert interpretation of the ECG trace. Many public datasets, e.g., PTB-XL, CPSC2018, and CSN (Wagner et al., 2020; Liu et al., 2018; Zheng et al., 2020) already include relevant labels. 2.Structural Disease includes tasks that probe cardiac morphology or function, such as systolic function and valvular disease. Labels for these are often derived from contemporaneous imaging including echocardio- graphy and cardiac MRI. EchoNext is an open-source dataset of paired ECG and echocardiogram findings containing relevant labels (Poterucha et al., 2025). 3. Hemodynamic State targets inference of cardiac fill- ing pressures, e.g., mean pulmonary capillary wedge pressure (mPCWP) and flows (Schlesinger et al., 2022). Ground truth for these labels often comes from right heart catheterization. Much of the literature on these tasks uses proprietary data. Publicly, MIMIC- IV contains some paired bedside-monitor ECG and invasive blood pressure signals (Moody et al., 2022). In addition to task family, downstream targets can be sep- arated into diagnosis and patient forecasting. Diagnosis involves estimating patient state at the time of an ECG, e.g., if a patient currently exhibits a left ventricular ejection frac- tion (LVEF) below 40%. Patient forecasting involves risk prediction over a future horizon, e.g., will a patient develop LVEF below 40% within one year of ECG acquisition. In Section 5, we evaluate current ECG representations on exemplar tasks from the framework defined above. We find that performance can vary widely across task types. Our proposed taxonomy is a starting point rather than an ex- haustive catalog of ECG applications. Future benchmarking should not be limited to (or necessarily include) them. We believe the community should come to a consensus on a set of clinically grounded tasks that belong in a standard ECG representation learning benchmark. 4. Toward Evaluation Best Practices An evaluation protocol should reliably stratify representa- tion quality, so that method rankings are robust to reasonable resampling and reporting choices. Unfortunately, much of the literature presents results in a way that obscures whether one representation is meaningfully better than another. First, the field relies on macro-AUROC (Wu & Zhou, 2017) as its primary metric. In our survey, 89.29% of papers re- ported it as their headline metric. While convenient for eval- uating multi-label data, macro-AUROC masks performance on the individual clinical tasks that clinicians care about. Furthermore, macro-averaging weights all labels equally. This implicitly grants rare, high-variance endpoints the same influence as common and more clinically salient ones, am- plifying noise in reported rankings. However, 35.71% of papers did not report per-task performance. Second, ECG benchmarks often include rare diagnoses that yield highly imbalanced labels with few positive examples. For example, under the standard PTB-XL protocol, 13 la- bels have fewer than 10 examples in the test set (Wagner et al., 2020; Strodthoff et al., 2021). In this case, per-label performance metrics have high sampling variability and can shift meaningfully because of small perturbations during resampling. Macro-AUROC inherits this variability, and can amplify it by giving equal weight to all labels. Under extreme class imbalance, AUROC can remain de- ceptively high, even when a model yields low precision at clinically relevant operating points. When negative cases vastly outnumber positive cases, the false positive rate can remain small even when the absolute number of false pos- itives is large. In this setting, AUPRC is an appropriate companion metric because it summarizes the tradeoff be- tween precision and recall, thereby capturing whether the model can recover positive cases without producing an ex- cessive number of false positives (Davis & Goadrich, 2006; Saito & Rehmsmeier, 2015). AUROC and AUPRC are useful because they are threshold- independent summaries of predictive performance. How- 4 Position: Evaluation of ECG Representations Must Be Fixed ever, they are not sufficient for many clinical use cases. In practice, relevant metrics depend on the intended application and operating point. For example, screening, confirmatory diagnosis, and patient forecasting typically prioritize dif- ferent tradeoffs between false positives and false negatives. When a deployment setting is specified, evaluation should therefore also include appropriate threshold-dependent met- rics, such as sensitivity, specificity, positive predictive value, and negative predictive value. Lastly, uncertainty is rarely quantified in the field; 42.86% of papers did not report confidence intervals for their main results. As a result, many apparent gaps between methods are indistinguishable from sampling noise. Precise statisti- cal testing is needed for claims of improvement over prior methods. The appropriate test depends on the experimental setting and methods. We advocate for a standardized and statistically rigorous reporting protocol. The following practices are broadly applicable to multi-label benchmarks: 1.Treat macro-averaged metrics as a coarse summary of performance; report per-task AUROC and AUPRC. 2. Report bootstrapped confidence intervals, not only point estimates. 3. Use paired comparisons for claims of improvement over prior methods, for example, with paired bootstrap confidence intervals or statistical tests. 4.Exclude labels with insufficient number of test-set ex- amples from quantitative evaluation. ECG datasets have dozens of labels, so exhaustive report- ing can be impractical in the main text of a work. How- ever, these data should be available in the supplementary material. We recommend emphasizing the most clinically salient endpoints, for example, diagnosis of low ejection fraction. These core labels should be agreed upon by the community to enable meaningful comparison and mitigate cherry-picking results. The importance of these recommendations is made trans- parent in Section 5 where we demonstrate that following them changes what one might conclude about the relative performance of current methods. 5. Empirical Study 5.1. Pre-training Configuration Models. We evaluate downstream performance using rep- resentations from five pre-trained ECG encoders: CLOCS (Kiyasseh et al., 2021), KED (Tian et al., 2024), HeartLang (Jin et al., 2025), MERL (Liu et al., 2024b), and D-BETA (Hung et al., 2024). As a baseline, we also evaluate embed- dings from a randomly initialized 1D ResNet-18 encoder, a common architecture in ECG modeling (He et al., 2016; Ribeiro et al., 2020). To assess the sensitivity of this base- line to architectural and preprocessing choices, we evaluate a grid of randomly initialized encoders in Section 5.4.4. Pre-training Dataset.We use the MIMIC-IV ECG database (Gow et al., 2023), which contains 800,035 10- second 12-lead ECGs collected from 161,352 patients. Each ECG is paired with a text diagnosis report. Implementation. We use the publicly available MIMIC-IV checkpoints for KED, HeartLang, MERL and D-BETA. For CLOCS, we retrain on MIMIC-IV following the authorsβ training procedures so all models use the same pre-training corpus; this isolates differences in model design and objec- tive rather than pre-training data. Full pre-training details are provided in Appendix C.2. All experiments are con- ducted on one NVIDIA Tesla V100-SXM2-32GB GPU. 5.2. Downstream Tasks We evaluate all ECG encoders with linear probing on six downstream settings. Full dataset details, including label definitions and prevalence, are provided in Appendix B. All ECGs are standardized to 10-second 12-lead segments in millivolts sampled at 500 Hz. We remove recordings with NaNorInfsamples. We use dataset-level train/val/test splits (70/10/20), except PTB-XL and EchoNext, which use standard splits, and the patient forecasting task, which follows a 75/10/15 split (Poterucha et al., 2025; Strodthoff et al., 2021; Bergamaschi et al., 2025). PTB-XL. PTB-XL consists of 21,837 12-lead 10-second ECGs from 18,885 patients (Wagner et al., 2022). It is split into four multi-label classification tasks that assess arrhyth- mia and waveform abnormalities, with varying numbers of binary targets: SUPER (5 labels), SUB (23 labels), FORM (19 labels), and RHYTHM (12 labels). Each task has a different number of samples, as detailed in Appendix B.1. CPSC2018. This dataset includes 6,877 12-lead ECGs, with arrhythmia and waveform morphology annotations (Liu et al., 2018). Recording duration varies between 5 and 72 seconds. We exclude recordings shorter than 10 seconds, and for longer recordings, clip them to 10 seconds. Appendix B.2 contains more details. CSN. The Chapman-Shaoxing-Ningbo (CSN) database in- cludes 45,152 10-second 12-lead ECGs from 10,646 patients (Zheng et al., 2022). ECGs are annotated with arrhythmia and waveform morphology labels; see Appendix B.3. EchoNext. For binary classification of structural heart dis- ease from the ECG, we use EchoNext, a dataset of 100,000 10-second 12-lead ECGs (Poterucha et al., 2025). Each ECG 5 Position: Evaluation of ECG Representations Must Be Fixed Table 1. On the standard arrhythmia/waveform benchmarks, KED and D-BETA appear to dominate according to macro-AUROC. However, once uncertainty is quantified, neither ECG pre-training method consistently prevails. A randomly initialized encoder is often competitive, often beating the ECG-specific pre-training methods CLOCS and HeartLang. Entries report macro-AUROC with 95% confidence intervals in the subscript under linear probing at varying levels of training data. Green indicates no significant difference from the top method. DatasetRandomCLOCSKEDHeartLangMERLD-BETA PTB-XL1%0.770 0.759β0.7800.752 0.740β0.7640.861 0.852β0.8690.645 0.632β0.6580.823 0.813β0.8320.867 0.858β0.875 SUPER10% 0.832 0.822β0.8420.807 0.796β0.8170.894 0.886β0.9010.790 0.780β0.8000.884 0.876β0.8920.887 0.879β0.895 100%0.861 0.851β0.8700.821 0.810β0.8300.909 0.902β0.9150.824 0.815β0.8340.901 0.893β0.9080.893 0.885β0.901 PTB-XL1%0.689 0.673β0.7040.657 0.626β0.6880.783 0.770β0.7950.574 0.548β0.5990.757 0.740β0.7720.791 0.781β0.802 SUB10%0.763 0.739β0.7860.771 0.752β0.7900.878 0.857β0.8980.708 0.688β0.7290.857 0.837β0.8750.862 0.834β0.888 100% 0.834 0.808β0.8590.803 0.784β0.8190.920 0.909β0.9290.836 0.819β0.8530.905 0.890β0.9180.899 0.882β0.914 PTB-XL1%0.542 0.528β0.5570.562 0.546β0.5770.656 0.643β0.6700.526 0.510β0.5420.620 0.607β0.6350.657 0.643β0.670 FORM10%0.684 0.658β0.7090.655 0.631β0.6790.719 0.688β0.7530.570 0.547β0.5940.735 0.713β0.7570.763 0.738β0.787 100%0.756 0.734β0.7800.715 0.687β0.7410.851 0.827β0.8750.707 0.679β0.7360.853 0.839β0.8660.845 0.821β0.868 PTB-XL1%0.499 0.474β0.5260.708 0.689β0.7280.810 0.789β0.8310.569 0.505β0.6300.746 0.708β0.7780.834 0.812β0.854 RHYTHM10%0.776 0.749β0.8040.802 0.777β0.8270.940 0.926β0.9520.731 0.679β0.7850.878 0.847β0.9050.956 0.941β0.968 100%0.787 0.737β0.8330.810 0.762β0.8540.959 0.948β0.9690.854 0.827β0.8790.903 0.870β0.9330.968 0.956β0.978 CPSC20181%0.630 0.615β0.6450.696 0.677β0.7150.855 0.843β0.8670.615 0.599β0.6320.835 0.822β0.8490.926 0.916β0.937 10%0.779 0.763β0.7940.769 0.753β0.7840.923 0.914β0.9320.719 0.702β0.7360.887 0.873β0.9000.942 0.933β0.951 100%0.850 0.834β0.8640.802 0.787β0.8160.952 0.945β0.9580.849 0.836β0.8620.927 0.916β0.9370.958 0.951β0.966 CSN1%0.603 0.597β0.6090.620 0.614β0.6260.695 0.689β0.7000.553 0.548β0.5600.663 0.658β0.6680.726 0.722β0.730 10%0.710 0.695β0.7240.734 0.716β0.7540.832 0.819β0.8450.686 0.672β0.7000.792 0.772β0.8110.860 0.850β0.868 100%0.763 0.741β0.7850.809 0.792β0.8260.910 0.900β0.9190.791 0.776β0.8060.862 0.853β0.8720.947 0.941β0.952 ECHONEXT1%0.694 0.678β0.7070.621 0.607β0.6360.687 0.671β0.7040.637 0.622β0.6520.691 0.677β0.7060.692 0.678β0.708 10%0.751 0.737β0.7640.702 0.690β0.7130.730 0.714β0.7450.704 0.691β0.7170.768 0.753β0.7820.754 0.742β0.766 100%0.773 0.761β0.7840.714 0.702β0.7250.766 0.752β0.7790.736 0.723β0.7500.791 0.780β0.8010.774 0.763β0.784 was paired with a contemporaneous echocardiogram, from which structural heart disease labels were derived. Details are in Appendix B.4. Hemodynamic Inference. We use a private dataset of 9,226 10-second 12-lead ECGs from 5,072 patients at Mas- sachusetts General Hospital (MGH) (Schlesinger et al., 2022). We consider two diagnosis tasks: contemporane- ous mean pulmonary capillary wedge pressure (mPCWP) and mean pulmonary artery pressure (mPA), each measured by ground-truth right heart catheterization. Cohort construc- tion and labeling are described in Appendix B.5. Patient Forecasting. We consider the binary prediction task of whether a patient will experience heart failure within one year of an ECG (1YR-HF). We define heart failure as echocardiographic left ventricular ejection fraction below 40%. We use the private dataset of (Bergamaschi et al., 2025), which includes 913,420 10-second 12-lead ECGs from 82,244 patients at MGH. See Section B.6. 5.3. Evaluation Protocol We evaluate the downstream performance of each represen- tation using linear probing. For each individual task, we freeze the ECG encoder and train a singleβ 2 -regularized logistic regression model on top of the embeddings using the training set. We run a hyperparameter sweep, detailed in Appendix C.1, pick the best probe for each task based on the validation set, then evaluate that probe on the test set. We report the AUROC and AUPRC for each individual task. We additionally aggregate over tasks to report the macro- AUROC for each dataset. To quantify uncertainty, for all experiments, we conduct a paired bootstrap with 1,000 re- samples with replacement on the test set and report 95% confidence intervals. In many of the following tables, we highlight methods whose performance is not significantly different from that of the top-performing method. To make this determination, for 6 Position: Evaluation of ECG Representations Must Be Fixed Table 2. Task-level results reveal heterogeneous behavior. Shown are the five highest-prevalence PTB-XL SUB labels plus CLBBB. Method rankings vary by endpoint. AUPRC can help capture differences masked by AUROC, as in the case of CLBBB. Each cell reports AUROC (top) and AUPRC (bottom) with 95% confidence intervals; methods not statistically different from the best are highlighted in green. Results on the remaining tasks are in Table 20. MethodNORMIMIAMISTTCLVHCLBBB Random0.891 0.877β0.903 / 0.839 0.813β0.863 0.864 0.840β0.884 / 0.556 0.506β0.609 0.894 0.874β0.915 / 0.646 0.593β0.700 0.827 0.801β0.851 / 0.319 0.275β0.368 0.917 0.896β0.935 / 0.648 0.593β0.704 0.973 0.934β0.999 / 0.880 0.784β0.955 CLOCS0.859 0.843β0.872 / 0.788 0.763β0.812 0.674 0.644β0.702 / 0.299 0.262β0.341 0.857 0.832β0.883 / 0.590 0.537β0.646 0.823 0.796β0.849 / 0.319 0.279β0.366 0.900 0.875β0.921 / 0.593 0.528β0.650 0.984 0.974β0.993 / 0.676 0.551β0.794 KED 0.932 0.921β0.942 / 0.900 0.883β0.916 0.881 0.862β0.899 / 0.597 0.549β0.647 0.951 0.939β0.962 / 0.808 0.770β0.845 0.893 0.873β0.911 / 0.507 0.444β0.570 0.949 0.935β0.961 / 0.745 0.695β0.794 0.998 0.997β1.000 / 0.949 0.906β0.983 HeartLang0.875 0.860β0.890 / 0.822 0.798β0.846 0.750 0.721β0.776 / 0.353 0.310β0.399 0.876 0.857β0.895 / 0.548 0.496β0.601 0.799 0.772β0.825 / 0.295 0.250β0.346 0.823 0.794β0.853 / 0.405 0.343β0.472 0.993 0.984β0.998 / 0.867 0.767β0.944 MERL0.926 0.915β0.937 / 0.885 0.862β0.904 0.843 0.819β0.864 / 0.538 0.488β0.589 0.959 0.949β0.969 / 0.841 0.809β0.873 0.885 0.863β0.905 / 0.480 0.421β0.539 0.929 0.912β0.945 / 0.680 0.623β0.734 0.999 0.997β1.000 / 0.947 0.889β0.988 D-BETA0.929 0.919β0.939 / 0.894 0.875β0.912 0.897 0.880β0.914 / 0.691 0.649β0.734 0.956 0.943β0.967 / 0.834 0.793β0.869 0.862 0.837β0.887 / 0.429 0.370β0.490 0.862 0.837β0.885 / 0.489 0.426β0.552 0.999 0.997β1.000 / 0.970 0.937β0.995 Table 3. Tasks with very few positives can yield high-variance estimates that distort benchmark summaries. Confidence intervals are wide for two of the three tasks with lowest prevalence in PTB- XL FORM. Each cell reports AUROC (top) and AUPRC (bottom) with 95% confidence intervals. MethodPRC(S)STETAB Random0.92 0.91β0.94 / 0.02 0.01β0.02 0.64 0.41β0.78 / 0.01 0.01β0.01 0.56 0.39β0.84 / 0.01 0.00β0.02 CLOCS0.85 0.83β0.87 / 0.01 0.01β0.01 0.56 0.10β0.91 / 0.01 0.00β0.04 0.70 0.59β0.80 / 0.01 0.01β0.02 KED0.97 0.96β0.98 / 0.03 0.03β0.05 0.59 0.30β0.94 / 0.01 0.00β0.05 0.74 0.52β0.98 / 0.03 0.01β0.13 HeartLang 0.87 0.85β0.89 / 0.01 0.01β0.01 0.61 0.20β0.93 / 0.01 0.00β0.05 0.27 0.14β0.50 / 0.00 0.00β0.01 MERL0.76 0.73β0.79 / 0.01 0.00β0.01 0.92 0.84β0.99 / 0.11 0.02β0.43 0.87 0.80β0.97 / 0.04 0.01β0.12 D-BETA0.97 0.96β0.98 / 0.04 0.03β0.06 0.71 0.51β0.95 / 0.02 0.01β0.06 0.66 0.23β0.95 / 0.02 0.00β0.07 each task, we first identify the best-performing method ac- cording to each metric. We then compare each other method against it using a two-sided paired permutation test with 1,000 permutation replicates. A method is considered not significantly different from the top-performing method if its Bonferroni-corrected p-value exceeds 0.05 on any metric. We note that this choice of test is not universally optimal, and the most appropriate statistical procedure depends on the experimental setting and the methods under comparison. To assess performance in a limited-label regime, on PTB- XL, CPSC2018, CSN, and EchoNext, we repeat this probing procedure using 1%, 10%, and 100% of the available labeled training data. 5.4. Experimental Results 5.4.1. EVALUATION ON PTB-XL, CPSC2018, CSN Macro Performance. Table 1 shows that conclusions drawn from macro-AUROC can change substantially once a random encoder baseline and uncertainty are included. KED and D-BETA often have the highest reported macro- AUROC, but the apparent winner changes once uncertainty is considered. For example, KED and D-BETA are often statistically indistinguishable, and MERL matches the top- performing method on PTB-XL FORM at 10% and 100% data and on PTB-XL SUB at 100% data. The randomly initialized encoder is competitive, frequently exceeding CLOCS and HeartLang, including on PTB-XL SUPER at all training fractions and on CPSC2018 at 10% and 100% of the data. We revisit the robustness of this observation in Section 5.4.4. Overall, many differences between methods fall within statistical noise, indicating that rankings based solely on macro-AUROC estimates can be unreliable. Task-level Performance. We next examine task-level per- formance within each dataset. Full results are available in Appendices D.1, E, and F. We report results for the five la- bels with highest prevalence from PTB-XL SUB in Table 2 and include complete left bundle branch block (CLBBB) as an illustrative example. The randomly initialized encoder is consistently competitive with CLOCS and HeartLang, and is among the best-performing within sampling noise on CLBBB. KED, MERL, and D-BETA provide gains on some tasks (e.g., STTC), but their relative ranking depends on the endpoint. Finally, CLBBB illustrates why AUROC alone can be misleading: AUROC is near-saturated for all methods, whereas AUPRC reveals better separation, with D- BETA clearly outperforming CLOCS, HeartLang, and the random baseline, while MERL and KED are intermediate. 7 Position: Evaluation of ECG Representations Must Be Fixed Table 4. Removing low-support labels can materially change macro-AUROC and alter which method appears most performant. ORIG uses the standard PTB-XL FORM test set; CLEAN ex- cludes labels with fewer than 10 positive test examples. Values are macro-AUROC with 95% confidence intervals. MethodORIGCLEANβ RANDOM0.756 0.734β0.7800.754 0.733β0.774-0.002 CLOCS0.715 0.687β0.7410.710 0.690β0.731-0.005 KED0.851 0.827β0.8750.862 0.846β0.876+0.011 HEARTLANG0.707 0.679β0.7360.723 0.699β0.745+0.016 MERL0.853 0.839β0.8660.847 0.833β0.860-0.005 D-BETA0.845 0.821β0.8680.855 0.841β0.868+0.009 Table 5. On structural disease endpoints, CLOCS and HeartLang underperform while the other methods cluster closely. Each cell reports AUROC/AUPRC with 95% confidence intervals; methods not significantly different from the top-performing method are highlighted in green. MethodSHDLVEFβ€ 45TR Random0.81 0.79β0.82 / 0.77 0.76β0.79 0.86 0.84β0.87 / 0.62 0.59β0.65 0.78 0.76β0.81 / 0.23 0.20β0.27 CLOCS0.72 0.71β0.74 / 0.69 0.67β0.70 0.78 0.76β0.79 / 0.48 0.45β0.51 0.71 0.69β0.74 / 0.14 0.12β0.17 KED0.80 0.79β0.81 / 0.76 0.75β0.78 0.86 0.85β0.87 / 0.62 0.59β0.65 0.78 0.75β0.80 / 0.23 0.19β0.26 HeartLang0.77 0.76β0.79 / 0.72 0.70β0.74 0.83 0.81β0.84 / 0.55 0.52β0.58 0.75 0.73β0.78 / 0.18 0.16β0.21 MERL 0.81 0.80β0.82 / 0.78 0.76β0.80 0.88 0.87β0.89 / 0.67 0.64β0.70 0.82 0.79β0.84 / 0.27 0.23β0.30 D-BETA0.80 0.79β0.81 / 0.76 0.75β0.78 0.86 0.85β0.87 / 0.61 0.58β0.64 0.79 0.77β0.82 / 0.24 0.20β0.27 5.4.2. SENSITIVITY OF MACRO-AVERAGED METRICS TO LABELS WITH FEW EXAMPLES In Section 4, we claim that the current practice of retaining labels with low test-support leads to noisy performance met- rics. To illustrate this point, Table 3 highlights the AUROC for the three tasks with the lowest prevalence in PTB-XL FORM. Each task has small test-support (PRC(S): 1 positive, STE: 3 positives, TAB: 3 positives), which can produce very wide confidence intervals and highly variable AUROC estimates across resamples. This variance has a material effect on the resulting macro-AUROC for each model. To demonstrate this, when we remove the PTB-XL FORM la- bels with fewer than 10 positive test examples (4 labels in total), the resulting macro-AUROC can shift non-trivially. This is shown in Table 4, where there is a change in the ap- parent ordering of the strongest methods: MERL is highest under the original label set, whereas KED is highest after low-support labels are removed. Parallel results for PTB-XL SUB and RHYTHM are provided in Appendix D.2. Table 6. The competitiveness of the random baseline in this paper is not an artifact of a single configuration. Values are macro- AUROC with standard deviation across grid of randomly initialized encoders. ALL summarizes all 36 random encoder configurations; RESNET and VIT summarize each backbone subset. DatasetALLRESNETVIT PTB-XL SUPER0.82Β± 0.07 0.87Β± 0.01 0.76Β± 0.07 PTB-XL SUB0.79Β± 0.09 0.86Β± 0.02 0.71Β± 0.08 PTB-XL FORM0.70Β± 0.07 0.76Β± 0.01 0.64Β± 0.05 PTB-XL RHYTHM 0.75Β± 0.09 0.81Β± 0.03 0.68Β± 0.07 CPSC20180.79Β± 0.09 0.86Β± 0.01 0.73Β± 0.08 CSN0.76Β± 0.09 0.82Β± 0.03 0.69Β± 0.07 ECHONEXT0.75Β± 0.03 0.77Β± 0.01 0.73Β± 0.03 5.4.3. EVALUATION ON ECHONEXT We consider EchoNext to illustrate how conclusions about method quality differ on structural disease endpoints. CLOCS and HeartLang underperform when the linear probe is trained with 1%, 10%, or 100% of the available training data (Table 1). At 1% of the data, the random baseline has the highest macro-AUROC, but falls within sampling noise of KED, MERL, and D-BETA. At 100% data, MERL achieves the strongest macro-AUROC, with the random baseline close behind. We further report performance on three clinically important structural endpoints (Table 5); the same pattern holds at the task level, with CLOCS and HeartLang consistently worse and the other methods tightly clustered. Additional results are available in Appendix G. 5.4.4. ROBUSTNESS OF RANDOM ENCODER BASELINE To characterize the robustness of the randomly initialized baseline, we reran our evaluation using 36 additional ran- dom encoders. Each encoder used a different combination of ECG sampling rate (100, 250, or 500 Hz), filtering method (none or order-5 Butterworth bandpass of 0.5 to 40 Hz), normalization method (none, dataset-level z-score, or per- sample z-score), and backbone (ResNet18 or ViT-Medium (Dosovitskiy et al., 2021)). Our preprocessing steps were based on common configurations in the ECG modeling liter- ature. For each encoder, we applied the same linear probing procedure described in Section 5.3. Table 6 shows the mean macro-AUROC across this grid at 100% label availability on each public dataset. We addi- tionally report the performance within the ResNet and ViT subsets. Standard deviation is reported across encoders. Un- surprisingly, the backbone affects performance, as ResNet18 outperformed ViT-Medium. However, while the perfor- mance of the random baseline is setup dependent, its com- petitiveness is not an artifact of the original configuration, whose results are in Table 1. Across a broad grid of reason- able choices it remains a non-trivial baseline, and is often competitive with at least some of the pre-trained methods. 8 Position: Evaluation of ECG Representations Must Be Fixed Table 7. Performance on mPCWP, mPA, and 1YR-HF (AU- ROC/AUPRC; 95% CIs). Most methods are statistically similar on the hemodynamic tasks; MERL and KED prevail on 1YR-HF. MethodmPCWPmPA1YR-HF Random0.70 0.67β0.73 / 0.71 0.68β0.75 0.72 0.68β0.76 / 0.86 0.83β0.88 0.78 0.78β0.79 / 0.58 0.58β0.59 CLOCS0.68 0.64β0.71 / 0.71 0.68β0.74 0.66 0.63β0.70 / 0.84 0.81β0.86 0.75 0.74β0.75 / 0.53 0.53β0.54 KED0.72 0.69β0.75 / 0.75 0.71β0.78 0.76 0.73β0.79 / 0.89 0.87β0.91 0.83 0.82β0.83 / 0.66 0.65β0.66 HeartLang0.68 0.65β0.71 / 0.70 0.66β0.73 0.71 0.67β0.74 / 0.86 0.84β0.88 0.79 0.78β0.79 / 0.58 0.58β0.59 MERL0.74 0.71β0.77 / 0.76 0.73β0.79 0.76 0.73β0.80 / 0.88 0.86β0.90 0.83 0.83β0.83 / 0.66 0.66β0.67 D-BETA0.71 0.68β0.74 / 0.73 0.70β0.77 0.74 0.71β0.78 / 0.87 0.85β0.90 0.82 0.82β0.82 / 0.64 0.63β0.64 5.4.5. HEMODYNAMIC INFERENCE Table 7 shows results for the two hemodynamics tasks. MERL achieves the highest AUROC on mPCWP, while MERL and KED tie for best on mPA. Performance dif- ferences across methods appear to be smaller than on the standard public benchmarks. All except for HeartLang are within statistical noise of MERL on mPCWP; D-BETA is within statistical noise of MERL and KED on mPA, with the randomly initialized encoder close behind. 5.4.6. PATIENT FORECASTING On 1-year heart failure forecasting (Table 7), MERL and KED are best with D-BETA following closely, all outper- forming Random, CLOCS, and HeartLang on AUROC and AUPRC. Random also exceeds CLOCS (βAUROC = 0.03; βAUPRC = 0.05), reinforcing that a randomly initialized baseline is non-trivial even for patient-level forecasting. 6. Alternative Views In Section 3 we argue that current downstream evaluation is narrow and should expand to include other clinical targets, such as structural disease, hemodynamic state, and forecast- ing tasks. A natural concern is that some of these targets cannot admit near-perfect performance due to aleatoric un- certainty, rendering them ill-suited for benchmarking (Ghas- semi et al., 2020; Kohane et al., 2021; Pillai et al., 2024; Yuan et al., 2021). Even so, many influential ML bench- marks have also remained far from saturation, often due to noise and ambiguity, yet have driven progress by rewarding better representations and modeling choices (Northcutt et al., 2021; Vendrow et al., 2025). When an endpoint cannot be perfectly inferred from an ECG, the signal that is present can be clinically meaningful and transferable toward other objectives. Hence, it is important that ECG representations are optimized to capture this information. We also argue that ECG representations should be evaluated on a broad range of clinical tasks, ideally with standard- ized and transparent benchmarks. However, many valuable ECG endpoints exist in health systems with non-trivial bar- riers to public release. Some might therefore object that such endpoints should be excluded from benchmarking be- cause private evaluations are less reproducible. We disagree. There is broad precedent in the ML community for bench- marking on hidden test sets, where the data are not released (e.g., Geiger et al., 2012; Wang et al., 2018; Perez Alday et al., 2022). Modern infrastructure enables containerized evaluation where benchmark hosts can run inference on be- half of a participant (Pavao et al., 2023). Private-endpoint benchmarks are especially appropriate for ECG represen- tation learning, where downstream evaluation often tests transferability to tasks unseen during training. Hosts should publish detailed documentation including cohort and label construction, and provide a transparent auditing pathway when possible. Furthermore, expanding public resources for such tasks should be a priority. 7. Discussion This paper analyzed evaluation of 12-lead ECG representa- tions. We proposed a taxonomy of clinically grounded tasks, then outlined best practices that the field often fails to fol- low. Our experiments show that ranking current methods is difficult in practice, and that randomly initialized encoders are surprisingly strong baselines. When random encoders are competitive with pre-trained encoders, gains from pre- training should be interpreted with care, especially if they are small in absolute terms. Future work should explore why the random encoder is so performant for many ECG tasks. We hypothesize that the clinically relevant signal is so apparent that random convolutions do not distort it. One limitation of our study is that we focused only on 12- lead ECG representation learning. We expect that many of our general conclusions hold for pre-trained encoders of other bio-signals (e.g., 1-lead ECG, PPG, and EEG), task- specific ECG models, and multi-modal representations that include ECG as a component (e.g., Radhakrishnan et al., 2023; Thapa et al., 2024); evaluating such models should be addressed in future work. The field of ECG analysis would benefit from an extensible open-source framework with standardized evaluation code. The community should agree on a core set of clinically important tasks that will drive the future of ECG pre-training. Our hope is that by doing so, the field will converge on broadly useful and generalizable representations. 9 Position: Evaluation of ECG Representations Must Be Fixed Acknowledgements We thank Tiffany Yau for guidance with the hemodynamic inference and patient forecasting tasks. We also thank Danielle Pace and Roey Ringel for helpful discussions and feedback. Zachary Berger is supported by the Department of Defense NDSEG Fellowship. This work was also supported by Quanta Computer Inc. References Al-Masud, M. A., Alcaraz, J. M. L., and Strodthoff, N. Benchmarking ecg foundational models: A reality check across clinical tasks, 2025. URLhttps://arxiv. org/abs/2509.25095. Bechler-Speicher, M., Finkelshtein, B., Frasca, F., M Μ uller, L., T Μ onshoff, J., Siraudin, A., Zaverkin, V., Bronstein, M., Niepert, M., Perozzi, B., Galkin, M., and Morris, C. Position: Graph learning will lose relevance due to poor benchmarks. In Proceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 of Proceedings of Machine Learning Research, p. 81067β 81089, 2025. URLhttps://openreview.net/ forum?id=nDFpl2lhoH. Bengio, Y., Courville, A., and Vincent, P. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35(8):1798β1828, August 2013. ISSN 0162-8828. doi: 10.1109/TPAMI.2013. 50.URLhttps://doi.org/10.1109/TPAMI. 2013.50. Bergamaschi, T., Yau, T., Chandak, P., Kyereme-Tuah, A., Hung, J., Gaggin, H., Kohane, I. S., and Stultz, C. M.Forecasting left ventricular systolic dys- function in heart failure with artificial intelligence. medRxiv, 2025. doi: 10.1101/2025.04.13.25325744. URLhttps://w.medrxiv.org/content/ early/2025/04/14/2025.04.13.25325744. Blalock, D. W., Ortiz, J. J. G., Frankle, J., and Guttag, J. V. What is the state of neural network pruning? In Dhillon, I. S., Papailiopoulos, D. S., and Sze, V. (eds.), Proceed- ings of the Third Conference on Machine Learning and Systems, MLSys 2020, Austin, TX, USA, March 2-4, 2020. mlsys.org, 2020. Bommasani, R. et al. On the opportunities and risks of foun- dation models, 2022. URLhttps://arxiv.org/ abs/2108.07258. Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual rep- resentations. In International conference on machine learning, p. 1597β1607. PMLR, 2020. Chen, X. and He, K. Exploring simple siamese representa- tion learning. arXiv preprint arXiv:2011.10566, 2020. Chen, X., Xie, S., and He, K. An empirical study of train- ing self-supervised vision transformers. arXiv preprint arXiv:2104.02057, 2021. Chen, Y., Orlandi, M., Rapa, P. M., Benatti, S., Benini, L., and Li, Y.Physiowave: A multi-scale wavelet- transformer for physiological signal representation. In NeurIPS 2025, 2025. URLhttps://openreview. net/forum?id=ayR2JfRYRS. Davis, J. and Goadrich, M. The relationship between precision-recall and roc curves. In Proceedings of the 23rd International Conference on Machine Learning, ICML β06, p. 233β240, New York, NY, USA, 2006. As- sociation for Computing Machinery. ISBN 1595933832. doi: 10.1145/1143844.1143874. URLhttps://doi. org/10.1145/1143844.1143874. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for lan- guage understanding. In Burstein, J., Doran, C., and Solorio, T. (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4171β4186, Min- neapolis, Minnesota, June 2019. Association for Compu- tational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423/. Diamant, N., Reinertsen, E., Song, S., Aguirre, A. D., Stultz, C. M., and Batra, P. Patient contrastive learn- ing: A performant, expressive, and practical approach to electrocardiogram modeling. PLOS Computational Biology, 18(2):e1009862, 2022. doi: 10.1371/journal. pcbi.1009862. URLhttps://doi.org/10.1371/ journal.pcbi.1009862. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. URLhttps: //arxiv.org/abs/2010.11929. Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., and Dean, J. A guide to deep learning in healthcare. Nature Medicine, 25(1):24β29, January 2019. doi: 10. 1038/s41591-018-0316-z. URLhttps://doi.org/ 10.1038/s41591-018-0316-z. Ferrari Dacrema, M., Cremonesi, P., and Jannach, D. Are we really making much progress?a worrying analysis of recent neural recommendation approaches. 10 Position: Evaluation of ECG Representations Must Be Fixed In Proceedings of the 13th ACM Conference on Rec- ommender Systems, RecSys β19, p. 101β109, New York, NY, USA, 2019. Association for Computing Machinery.ISBN 9781450362436.doi: 10.1145/ 3298689.3347058.URLhttps://doi.org/10. 1145/3298689.3347058. Fisch, C. Centennial of the string galvanometer and the electrocardiogram. Journal of the American College of Cardiology, 36(6):1737β1745, 2000. ISSN 0735-1097. doi: https://doi.org/10.1016/S0735-1097(00)00976-1. URLhttps://w.sciencedirect.com/ science/article/pii/S0735109700009761. Friedman, S. F., Khurshid, S., Venn, R. A., Wang, X., Diamant, N., Di Achille, P., Weng, L.-C., Choi, S. H., Reeder, C., Pirruccello, J. P., Singh, P., Lau, E. S., Philip- pakis, A., Anderson, C. D., Maddah, M., Batra, P., Elli- nor, P. T., Ho, J. E., and Lubitz, S. A. Unsupervised deep learning of electrocardiograms enables scalable hu- man disease profiling. npj Digital Medicine, 8(1):23, 2025. doi: 10.1038/s41746-024-01418-9. URLhttps: //doi.org/10.1038/s41746-024-01418-9. Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., I, H. D., and Crawford, K. Datasheets for datasets. Commun. ACM, 64(12):86β92, November 2021. ISSN 0001-0782. doi: 10.1145/3458723. URL https://doi.org/10.1145/3458723. Geiger, A., Lenz, P., and Urtasun, R. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, p. 3354β3361, 2012. doi: 10.1109/CVPR. 2012.6248074. Ghassemi, M., Naumann, T., Schulam, P., Beam, A. L., Chen, I. Y., and Ranganath, R.A review of chal- lenges and opportunities in machine learning for health. AMIA Joint Summits on Translational Science, p. 191β 200, May 2020. URLhttps://pubmed.ncbi.nlm. nih.gov/32477638/. Goldberger, A. L., Amaral, L. A. N., Glass, L., Haus- dorff, J. M., Ivanov, P. C., Mark, R. G., Mietus, J. E., Moody, G. B., Peng, C.-K., and Stanley, H. E. Phys- iobank, physiotoolkit, and physionet: Components of a new research resource for complex physiologic sig- nals.Circulation, 101(23):e215βe220, 2000.doi: 10.1161/01.CIR.101.23.e215. RRID:SCR007345. Gopal, B., Han, R., Raghupathi, G., Ng, A., Tison, G., and Rajpurkar, P.3kg: Contrastive learning of 12- lead electrocardiograms using physiologically-inspired augmentations. In Roy, S., Pfohl, S., Rocheteau, E., Tadesse, G. A., Oala, L., Falck, F., Zhou, Y., Shen, L., Zamzmi, G., Mugambi, P., Zirikly, A., McDermott, M. B. A., and Alsentzer, E. (eds.), Proceedings of Machine Learning for Health, volume 158 of Proceedings of Ma- chine Learning Research, p. 156β167. PMLR, 04 Dec 2021. URLhttps://proceedings.mlr.press/ v158/gopal21a.html. Gow, B., Pollard, T., Nathanson, L. A., Johnson, A., Moody, B., Fernandes, C., Greenbaum, N., Waks, J. W., Eslami, P., Carbonati, T., Chaudhari, A., Herbst, E., Moukheiber, D., Berkowitz, S., Mark, R., and Horng, S. MIMIC-IV-ECG: Diagnostic electrocardiogram matched subset.https://physionet.org/content/ mimic-iv-ecg/1.0/, 2023. RRID:SCR007345. Grill, J.-B., Strub, F., Altch Μ e, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., Piot, B., kavukcuoglu, k., Munos, R., and Valko, M. Bootstrap your own latent - a new approach to self-supervised learning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, p. 21271β21284. Curran Associates, Inc., 2020.URLhttps://proceedings.neurips. c/paper_files/paper/2020/file/ f3ada80d5c4e70142b17b8192b2958e-Paper. pdf. Han, Y., Murino, V., Liu, X., Zhang, X., and Ding, C. A systematic review on foundation models for elec- trocardiogram analysis: Initial strides and expansive horizons, 2025. URLhttps://arxiv.org/abs/ 2410.19877. He, H. and Garcia, E. A.Learning from imbalanced data. IEEE Trans. on Knowl. and Data Eng., 21(9): 1263β1284, September 2009. ISSN 1041-4347. doi: 10.1109/TKDE.2008.239. URLhttps://doi.org/ 10.1109/TKDE.2008.239. He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learn- ing for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. He, K., Chen, X., Xie, S., Li, Y., Doll Μ ar, P., and Girshick, R. Masked autoencoders are scalable vision learners, 2021. URL https://arxiv.org/abs/2111.06377. Herrmann, M., Lange, F. J. D., Eggensperger, K., Casal- icchio, G., Wever, M., Feurer, M., R Μ ugamer, D., H Μ ullermeier, E., Boulesteix, A.-L., and Bischl, B. Posi- tion: Why we must rethink empirical research in machine learning. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference 11 Position: Evaluation of ECG Representations Must Be Fixed on Machine Learning, volume 235 of Proceedings of Ma- chine Learning Research, p. 18228β18247. PMLR, 21β 27 Jul 2024. URLhttps://proceedings.mlr. press/v235/herrmann24b.html. Hung, M. P., Saeed, A., and Ma, D. Boosting masked ecg- text auto-encoders as discriminative learners. In Forty- second International Conference on Machine Learning, 2024. URLhttps://openreview.net/forum? id=mM65b81LdM. Jin, J., Wang, H., Li, H., Li, J., Pan, J., and Hong, S. Read- ing your heart: Learning ECG words and sentences via pre-training ECG language model. In The Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/forum? id=6Hz1Ko087B. Khurshid, S., Friedman, S., Reeder, C., Achille, P. D., Diamant, N., Singh, P., Harrington, L. X., Wang, X., Al-Alusi, M. A., Sarma, G., Foulkes, A. S., Elli- nor, P. T., Anderson, C. D., Ho, J. E., Philippakis, A. A., Batra, P., and Lubitz, S. A. Ecg-based deep learning and clinical risk factors to predict atrial fibrillation. Circulation, 145(2):122β133, 2022. doi: 10.1161/CIRCULATIONAHA.121.057480.URL https://w.ahajournals.org/doi/abs/ 10.1161/CIRCULATIONAHA.121.057480. Kingma, D. P. and Ba, J. Adam: A method for stochastic op- timization, 2017. URLhttps://arxiv.org/abs/ 1412.6980. Kiyasseh, D., Zhu, T., and Clifton, D. A. Clocs: Con- trastive learning of cardiac signals across space, time, and patients. In Meila, M. and Zhang, T. (eds.), Pro- ceedings of the 38th International Conference on Ma- chine Learning, volume 139 of Proceedings of Machine Learning Research, p. 5606β5615. PMLR, 18β24 Jul 2021. URLhttps://proceedings.mlr.press/ v139/kiyasseh21a.html. Kohane, I. S., Aronow, B. J., Avillach, P., Beaulieu-Jones, B. K., Bellazzi, R., Bradford, R. L., Brat, G. A., Can- nataro, M., Cimino, J. J., Garc Μ Δ±a-Barrio, N., Gehlen- borg, N., Ghassemi, M., Guti Μ errez-Sacrist Μ an, A., Hanauer, D. A., Holmes, J. H., Hong, C., Klann, J. G., Loh, N. H. W., Luo, Y., Mandl, K. D., Daniar, M., Moore, J. H., Murphy, S. N., Neuraz, A., Ngiam, K. Y., Omenn, G. S., Palmer, N., Patel, L. P., Pedrera-Jim Μ enez, M., Sliz, P., South, A. M., Tan, A. L. M., Taylor, D. M., Taylor, B. W., Torti, C., Vallejos, A. K., Wagholikar, K. B., We- ber, G. M., and Cai, T. What every reader should know about studies using electronic health record data but may be afraid to ask. J Med Internet Res, 23(3):e22219, Mar 2021. ISSN 1438-8871. doi: 10.2196/22219. URL https://doi.org/10.2196/22219. Koscova, Z., Li, Q., Robichaux, C., Moura Junior, V., Ghanta, M., Gupta, A., Rosand, J., Aguirre, A., Hong, S., Albert, D. E., Xue, J., Parekh, A., Sameni, R., Reyna, M. A., Westover, M. B., and Cliford, G. D.The harvard-emory ecg database. medRxiv, 2024. doi: 10.1101/2024.09.27.24314503. URLhttps://w.medrxiv.org/content/ early/2024/10/01/2024.09.27.24314503. Lai, J., Tan, H., Wang, J., Ji, L., Guo, J., Han, B., Shi, Y., Feng, Q., and Yang, W.Practical intelli- gent diagnostic algorithm for wearable 12-lead ECG via self-supervised learning on large-scale dataset. Na- ture Communications, 14:3741, 2023. doi: 10.1038/ s41467-023-39472-8. URLhttps://w.nature. com/articles/s41467-023-39472-8. Lalam, S. K., Kunderu, H. K., Ghosh, S., A, H. K., Awasthi, S., Prasad, A., Lopez-Jimenez, F., Attia, Z. I., Asir- vatham, S., Friedman, P., Barve, R., and Babu, M. ECG representation learning with multi-modal EHR data.Transactions on Machine Learning Research, November 2023.ISSN 2835-8856.URLhttps: //openreview.net/forum?id=UxmvCwuTMG. Lan, X., Ng, D., Hong, S., and Feng, M. Intra-inter sub- ject self-supervised learning for multivariate cardiac sig- nals. In Proceedings of the AAAI Conference on Ar- tificial Intelligence (AAAI-22), volume 36, 2022. doi: 10.1609/aaai.v36i4.20376. URLhttps://doi.org/ 10.1609/aaai.v36i4.20376. Le, D., Truong, S., Brijesh, P., Adjeroh, D. A., and Le, N. scl-st: Supervised contrastive learning with semantic transformations for multiple lead ecg arrhythmia classi- fication. IEEE Journal of Biomedical and Health Infor- matics, 27(6):2818β2828, 2023. doi: 10.1109/JBHI.2023. 3246241. Li, J., Liu, C., Cheng, S., Arcucci, R., and Hong, S. Frozen language model helps ecg zero-shot learning, 2023. URL https://arxiv.org/abs/2303.12311. Li, J., Aguirre, A. D., Junior, V. M., Jin, J., Liu, C., Zhong, L., Sun, C., Clifford, G., Brandon Westover, M., and Hong, S. An electrocardiogram foundation model built on over 10 million recordings. NEJM AI, 2(7):AIoa2401033, 2025. Liao, T., Taori, R., Raji, D., and Schmidt, L. Are we learn- ing yet? a meta review of evaluation failures across ma- chine learning. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll Μ ar, P., and Zitnick, C. L. Microsoft 12 Position: Evaluation of ECG Representations Must Be Fixed coco: Common objects in context. In Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T. (eds.), Computer Vi- sion β ECCV 2014, p. 740β755, Cham, 2014. Springer International Publishing. ISBN 978-3-319-10602-1. Lipton, Z. C. and Steinhardt, J. Troubling trends in machine learning scholarship. ACM Queue, 17(1), 2019. doi: 10. 1145/3317287.3328534. URLhttps://doi.org/ 10.1145/3317287.3328534. Littlejohns, T. J., Holliday, J., Gibson, L. M., Garratt, S., Oesingmann, N., Alfaro-Almagro, F., Bell, J. D., Boult- wood, C., Collins, R., Conroy, M. C., et al. The uk biobank imaging enhancement of 100,000 participants: rationale, data collection, management and future direc- tions. Nature communications, 11(1):2624, 2020. Liu, C., Wan, Z., Cheng, S., Zhang, M., and Arcucci, R. Etp: Learning transferable ecg representations via ecg-text pre-training. In ICASSP 2024 - 2024 IEEE In- ternational Conference on Acoustics, Speech and Sig- nal Processing (ICASSP), p. 8230β8234, 2024a. doi: 10.1109/ICASSP48485.2024.10446742. Liu, C., Wan, Z., Ouyang, C., Shah, A., Bai, W., and Arcucci, R. Zero-shot ECG classification with mul- timodal learning and test-time clinical knowledge en- hancement. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Ma- chine Learning Research, p. 31949β31963. PMLR, 21β 27 Jul 2024b. URLhttps://proceedings.mlr. press/v235/liu24bg.html. Liu, F., Liu, C., Zhao, L., Zhang, X., Wu, X., Xu, X., Liu, Y., Ma, C., Wei, S., He, Z., Li, J., and Kwee, E. N. Y. An open access database for evaluating the algorithms of electrocardiogram rhythm and morphol- ogy abnormality detection. Journal of Medical Imag- ing and Health Informatics, 8:1368β1373, 2018. doi: 10.1166/jmihi.2018.2442. URLhttps://doi.org/ 10.1166/jmihi.2018.2442. Liu, H., Chen, D., Chen, D., Zhang, X., Li, H., Bian, L., Shu, M., and Wang, Y. A large-scale multi-label 12-lead electrocardiogram database with standardized diagnostic statements. Scientific Data, 9:272, 2022. doi: 10.1038/ s41597-022-01403-5. URLhttps://doi.org/10. 1038/s41597-022-01403-5. Liu, Q. and Paparrizos, J. The elephant in the room: Towards a reliable time-series anomaly detection benchmark. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, p. 108231β 108261. Curran Associates, Inc., 2024. doi: 10.52202/ 079017-3437. Lunelli, R., Nicolson, A., Pr Μ oll, S. M., Reinstadler, S. J., Bauer, A., and Dlaska, C. Benchecg and xecg: a bench- mark and baseline for ecg foundation models, 2025. URL https://arxiv.org/abs/2509.10151. McKeen, K., Masood, S., Toma, A., Rubin, B., and Wang, B. Ecg-fm: an open electrocardiogram foundation model. JAMIA Open, 8(5):ooaf122, 10 2025. ISSN 2574-2531. doi: 10.1093/jamiaopen/ooaf122. URLhttps://doi. org/10.1093/jamiaopen/ooaf122. Mehari, T. and Strodthoff, N. Self-supervised representation learning from 12-lead ecg data. Comput. Biol. Med., 141 (C), February 2022. ISSN 0010-4825. doi: 10.1016/ j.compbiomed.2021.105114.URLhttps://doi. org/10.1016/j.compbiomed.2021.105114. Moody, B., Hao, S., Gow, B., Pollard, T., Zong, W., and Mark, R. MIMIC-IV Waveform Database. PhysioNet, July 2022. doi: 10.13026/a2mw-f949. URLhttps:// doi.org/10.13026/a2mw-f949. Version 0.1.0. Na, Y., Park, M., Tae, Y., and Joo, S. Guiding masked representation learning to capture spatio-temporal rela- tionship of electrocardiogram. In International Confer- ence on Learning Representations, 2024. URLhttps: //openreview.net/forum?id=WcOohbsF4H. Noble, R. J., Hillis, J. S., and Rothbaum, D. A. Electro- cardiography. In Walker, H. K., Hall, W. D., and Hurst, J. W. (eds.), Clinical Methods: The History, Physical, and Laboratory Examinations, chapter 33. Butterworths, Boston, 3 edition, 1990. URLhttps://w.ncbi. nlm.nih.gov/books/NBK354/. Northcutt, C., Athalye, A., and Mueller, J. Pervasive label errors in test sets destabilize machine learning bench- marks. In Vanschoren, J. and Yeung, S. (eds.), Proceed- ings of the Neural Information Processing Systems Track on Datasets and Benchmarks, volume 1, 2021. Oh, J., Chung, H., Kwon, J.-m., Hong, D.-g., and Choi, E. Lead-agnostic self-supervised learning for local and global representations of electrocardiogram. In Flores, G., Chen, G. H., Pollard, T., Ho, J. C., and Naumann, T. (eds.), Proceedings of the Conference on Health, In- ference, and Learning, volume 174 of Proceedings of Machine Learning Research, p. 338β353. PMLR, 07β 08 Apr 2022. URLhttps://proceedings.mlr. press/v174/oh22a.html. Pavao, A., Guyon, I., Letournel, A.-C., Tran, D.-T., Baro, X., Escalante, H. J., Escalera, S., Thomas, T., and Xu, 13 Position: Evaluation of ECG Representations Must Be Fixed Z. Codalab competitions: An open source platform to organize scientific challenges. Journal of Machine Learning Research, 24(198):1β6, 2023. URLhttp: //jmlr.org/papers/v24/21-1436.html. Perez Alday, E. A., Gu, A., Shah, A., Liu, C., Sharma, A., Seyedi, S., Bahrami Rad, A., Reyna, M., and Clif- ford, G. Classification of 12-lead ECGs: The Phys- ioNet/Computing in cardiology challenge 2020 (version 1.0.2). PhysioNet, 2022. URLhttps://doi.org/ 10.13026/dvyd-kd57. RRID:SCR007345. Pillai, B., Salerno, M., Schnittger, I., Cheng, S., and Ouyang, D. Precision of echocardiographic measure- ments. Journal of the American Society of Echocar- diography, 37(5):562β563, 2024.ISSN 0894-7317. doi: 10.1016/j.echo.2024.01.001. URLhttps://doi. org/10.1016/j.echo.2024.01.001. Pineau, J., Vincent-Lamarre, P., Sinha, K., Lariviere, V., Beygelzimer, A., dβAlche Buc, F., Fox, E., and Larochelle, H. Improving reproducibility in machine learning research(a report from the neurips 2019 re- producibility program). Journal of Machine Learning Research, 22(164):1β20, 2021. URLhttp://jmlr. org/papers/v22/20-303.html. Poterucha, T. J., Jing, L., Ricart, R. P., Adjei-Mosi, M., Finer, J., Hartzel, D., Kelsey, C., Long, A., Rocha, D., Ruhl, J. A., vanMaanen, D., Probst, M. A., Daniels, B., Joshi, S. D., Tastet, O., Corbin, D., Avram, R., Barrios, J. P., Tison, G. H., Chiu, I.-M., Ouyang, D., Volodarskiy, A., Castillo, M., Roedan Oliver, F. A., Malta, P. P., Ye, S., Rosner, G. F., Dizon, J. M., Ali, S. R., Liu, Q., Bradley, C. K., Vaishnava, P., Waksmonski, C. A., DeFilippis, E. M., Agarwal, V., Lebehn, M., Kampaktsis, P. N., Shames, S., Beecy, A. N., Kumaraiah, D., Homma, S., Schwartz, A., Hahn, R. T., Leon, M., Einstein, A. J., Mau- rer, M. S., Hartman, H. S., Hughes, J. W., Haggerty, C. M., and Elias, P. Detecting structural heart disease from elec- trocardiograms using ai. Nature, 644:221β230, August 2025. doi: 10.1038/s41586-025-09227-0. URLhttps: //doi.org/10.1038/s41586-025-09227-0. Radhakrishnan, A., Friedman, S. F., Khurshid, S., Ng, K., Batra, P., Lubitz, S. A., Philippakis, A. A., and Uhler, C. Cross-modal autoencoder framework learns holistic representations of cardiovascular state.Na- ture Communications, 14(1):2436, 2023. doi: 10.1038/ s41467-023-38125-0. URLhttps://doi.org/10. 1038/s41467-023-38125-0. Ribeiro, A. H., Ribeiro, M. H., Paix Μ ao, G. M. M., Oliveira, D. M., Gomes, P. R., Canazart, J. A., Ferreira, M. P. S., Andersson, C. R., Macfarlane, P. W., Meira Jr., W., Sch Μ on, T. B., and Ribeiro, A. L. P. Automatic diagnosis of the 12-lead ecg using a deep neural net- work. Nature Communications, 11(1):1760, 2020. doi: 10.1038/s41467-020-15432-4. Rivera-Ruiz, M., Cajavilca, C., and Varon, J. Einthovenβs string galvanometer. Texas Heart Institute Journal, 35 (2):174β178, 2008. URLhttps://pmc.ncbi.nlm. nih.gov/articles/PMC2435435/. Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211β252, 2015. doi: 10.1007/s11263-015-0816-y. Saito, T. and Rehmsmeier, M. The precision-recall plot is more informative than the roc plot when evaluat- ing binary classifiers on imbalanced datasets. PLOS ONE, 10(3):1β21, 03 2015.doi: 10.1371/journal. pone.0118432. URLhttps://doi.org/10.1371/ journal.pone.0118432. Schlesinger, D. E., Diamant, N., Raghu, A., Reinert- sen, E., Young, K., Batra, P., Pomerantsev, E., and Stultz, C. M.A deep learning model for infer- ring elevated pulmonary capillary wedge pressures from the 12-lead electrocardiogram.JACC: Ad- vances, 1(1):100003, 2022. doi: 10.1016/j.jacadv.2022. 100003.URLhttps://w.jacc.org/doi/ abs/10.1016/j.jacadv.2022.100003. Song, J., Jang, J.-H., Hong, D., myoung Kwon, J., and Jo, Y.- Y. Crema: A contrastive regularized masked autoencoder for robust ecg diagnostics across clinical domains, 2025. URL https://arxiv.org/abs/2407.07110. Strodthoff, N., Wagner, P., Schaeffter, T., and Samek, W. Deep learning for ecg analysis: Benchmarks and insights from ptb-xl. IEEE Journal of Biomedical and Health Informatics, 25(5):1519β1528, 2021. doi: 10.1109/JBHI. 2020.3022989. Sudlow, C., Gallacher, J., Allen, N., Beral, V., Burton, P., Danesh, J., Downey, P., Elliott, P., Green, J., Landray, M., et al. Uk biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age. PLoS medicine, 12(3):e1001779, 2015. Thapa, R., He, B., Kjaer, M. R., Iv, H. M., Ganjoo, G., Mignot, E., and Zou, J. SleepFM: Multi-modal representation learning for sleep across brain activity, ECG and respiratory signals.In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scar- lett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learn- ing, volume 235 of Proceedings of Machine Learn- ing Research, p. 48019β48037. PMLR, 21β27 Jul 14 Position: Evaluation of ECG Representations Must Be Fixed 2024. URLhttps://proceedings.mlr.press/ v235/thapa24a.html. Tian, Y., Li, Z., Jin, Y., Wang, M., Wei, X., Zhao, L., Liu, Y., Liu, J., and Liu, C. Foundation model of ecg diagnosis:Diagnostics and explanations of any form and rhythm on ecg.Cell Reports Medicine, 5(12):101875, 2024.URLhttps: //w.cell.com/cell-reports-medicine/ fulltext/S2666-3791(24)00646-3. Vendrow, J., Vendrow, E., Beery, S., and Madry, A. Do large language model benchmarks test reliability?, 2025. URL https://arxiv.org/abs/2502.03461. Wagner, P., Strodthoff, N., Bousseljot, R.-D., Lunze, F. I., Samek, W., and Schaeffter, T. PTB-XL: A large publicly available electrocardiography dataset. Scientific Data, 7 (1):154, 2020. doi: 10.1038/s41597-020-0495-6. Wagner, P., Strodthoff, N., Bousseljot, R.-D., Samek, W., and Schaeffter, T.PTB-XL, a large publicly available electrocardiography dataset.https: //physionet.org/content/ptb-xl/1.0.3/ , 2022. RRID:SCR007345. Wan, Z., Yu, Q., Mao, J., Duan, W., and Ding, C. Openecg: Benchmarking ecg foundation models with public 1.2 million records, 2025. URLhttps://arxiv.org/ abs/2503.00711. Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bow- man, S. GLUE: A multi-task benchmark and analysis plat- form for natural language understanding. In Linzen, T., ChrupaΕa, G., and Alishahi, A. (eds.), Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and In- terpreting Neural Networks for NLP, p. 353β355, Brus- sels, Belgium, November 2018. Association for Compu- tational Linguistics. doi: 10.18653/v1/W18-5446. URL https://aclanthology.org/W18-5446/. Wang, F., Xu, J., and Yu, L. From token to rhythm: A multi- scale approach for ECG-language pretraining. In Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., and Zhu, J. (eds.), Pro- ceedings of the 42nd International Conference on Ma- chine Learning, volume 267 of Proceedings of Machine Learning Research, p. 65059β65074. PMLR, 13β19 Jul 2025. URLhttps://proceedings.mlr.press/ v267/wang25du.html. Wang, N., Feng, P., Ge, Z., Zhou, Y., Zhou, B., and Wang, Z. Adversarial spatiotemporal contrastive learning for electrocardiogram signals. IEEE Transactions on Neural Networks and Learning Systems, 35(10):13845β13859, 2024. doi: 10.1109/TNNLS.2023.3272153. Wei, C. T., Hsieh, M.-E., Liu, C.-L., and Tseng, V. S. Contrastive heartbeats: Contrastive learning for self- supervised ecg representation and phenotyping.In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1126β1130, 2022. doi: 10.1109/ICASSP43922.2022. 9746887. Wu, X.-Z. and Zhou, Z.-H. A unified view of multi-label per- formance measures. In Precup, D. and Teh, Y. W. (eds.), Proceedings of the 34th International Conference on Ma- chine Learning, volume 70 of Proceedings of Machine Learning Research, p. 3780β3788. PMLR, 06β11 Aug 2017. URLhttps://proceedings.mlr.press/ v70/wu17a.html. Yang, C., Westover, M., and Sun, J. Biot: Biosignal trans- former for cross-data learning in the wild. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Pro- cessing Systems, volume 36, p. 78240β78260. Curran Associates, Inc., 2023. Yu, H., Guo, P., and Sano, A. Ecg semantic integrator (esi): A foundation ecg model pretrained with llm-enhanced cardiological text. Transactions on Machine Learning Research (TMLR), 2024. Yuan, N., Jain, I., Rattehalli, N., He, B., Pollick, C., Liang, D., Heidenreich, P., Zou, J., Cheng, S., and Ouyang, D.Systematic quantification of sources of variation in ejection fraction calculation using deep learning.JACC: Cardiovascular Imaging, 14 (11):2260β2262, 2021. doi: 10.1016/j.jcmg.2021.06. 018. URLhttps://w.jacc.org/doi/abs/ 10.1016/j.jcmg.2021.06.018. Zhang, H., Liu, W., Shi, J., Chang, S., Wang, H., He, J., and Huang, Q. Maefe: Masked autoencoders family of electrocardiogram for self-supervised pretraining and transfer learning. IEEE Transactions on Instrumentation and Measurement, 72:1β15, 2023a. doi: 10.1109/TIM. 2022.3228267. Zhang, M.-L. and Zhou, Z.-H. A review on multi-label learning algorithms. IEEE Transactions on Knowledge and Data Engineering, 26(8):1819β1837, 2014. doi: 10. 1109/TKDE.2013.39. Zhang, W., Geng, S., and Hong, S. A simple self-supervised ecg representation learning method via manipulated temporalβspatial reverse detection. Biomedical Signal Processing and Control, 79:104194, 2023b. ISSN 1746- 8094. doi: https://doi.org/10.1016/j.bspc.2022.104194. URLhttps://w.sciencedirect.com/ science/article/pii/S1746809422006486. 15 Position: Evaluation of ECG Representations Must Be Fixed Zhang, W., Yang, L., Geng, S., and Hong, S. Self-supervised time series representation learning via cross reconstruc- tion transformer. IEEE Transactions on Neural Networks and Learning Systems, 35(11):16129β16138, 2024. doi: 10.1109/TNNLS.2023.3292066. Zheng, J., Chu, H., Struppa, D., Zhang, J., Yacoub, S. M., El-Askary, H., Chang, A., Ehwerhemuepha, L., Abu- dayyeh, I., Barrett, A., Fu, G., Yao, H., Li, D., Guo, H., and Rakovski, C. Optimal multi-stage arrhythmia classi- fication approach. Scientific Reports, 10:2898, 2020. doi: 10.1038/s41598-020-59821-7. Zheng, J., Guo, H., and Chu, H.A large scale 12-leadelectrocardiogramdatabaseforarrhyth- miastudy.https://physionet.org/ content/ecg-arrhythmia/1.0.0/,2022. RRID:SCR007345. Zhou, R., Zhang, Y., and Dong, Y. H-tuning: Toward low-cost and efficient ECG-based cardiovascular dis- ease detection with pre-trained models. In Singh, A., Fazel, M., Hsu, D., Lacoste-Julien, S., Berkenkamp, F., Maharaj, T., Wagstaff, K., and Zhu, J. (eds.), Pro- ceedings of the 42nd International Conference on Ma- chine Learning, volume 267 of Proceedings of Machine Learning Research, p. 79548β79569. PMLR, 13β19 Jul 2025. URLhttps://proceedings.mlr.press/ v267/zhou25aj.html. 16 Position: Evaluation of ECG Representations Must Be Fixed A. Selection of Model Survey Set We construct a model survey set of papers that propose a representation learning method for 12-lead ECGs. In the literature, these are commonly referred to as pre-trained ECG encoders or ECG foundation models. We consider papers published between January 1, 2019 and December 31, 2025. We begin our survey in 2019 since that date coincides with the invention of contemporary self-supervised and large-scale pretraining methods applied to ECGs, e.g. SimCLR and BYOL (Chen et al., 2020; Grill et al., 2020). To form the survey set, we first identified methods published in top ML and ML-health venues: ICML, ICLR, NeurIPS, AAAI, CHIL, and ML4H. For each venue, we required that one keyword from each of the following sets appeared in the paper title or abstract: Cardiac Keyword. ECG, EKG, electrocardiogram, electrocardiography, cardiac, 12-lead, multi-lead, multilead. Representation Learning Keyword. SSL, self-supervised, contrastive, masked, reconstruction, reconstructive, foundation model, pretrain, pre-train, pre-training, pretraining, multimodal, multi-modal, representation. If a paper plausibly proposed a transferable 12-lead ECG representation based on its title or abstract, we screened the full text. A paper was included in the survey set if it met the following inclusion criteria: 1. 12-lead ECG is one of the primary modalities considered. 2. The paper proposes a method designed to produce a reusable representation. 3. The paper reports downstream evaluation on at least one 12-lead ECG task using linear probing or fine-tuning. Because a substantial amount of ECG representation learning work appears outside this ecosystem, we then added any method cited by the initially identified papers in their introduction, related works or as a baseline, provided they fit our inclusion criteria. While this procedure may have missed some relevant papers, e.g., some included in (Han et al., 2025), it has the desirable property of restricting our discussion to work that is directly pertinent to the ML-methods community. The initial search from the ML venues returned 56 papers, with 11 papers left after screening. There were then an additional 17 cited papers that were included. The resulting survey set contains 28 papers, listed in Table 8. We use this set when making claims about common benchmarking and reporting practices in 12-lead ECG representation learning. For each method in the survey set, we extracted the following data, which is summarized for each method in Table 9: β’ Publication venue. β’ Year of publication. β’ Pre-training dataset used to develop the representation. β’ Downstream datasets the representation is evaluated on. β’ Which performance metrics were reported. β’ If task-level metrics were reported, or only macro-averaged metrics. β’ Whether uncertainty was quantified in the paper. β’ If the method compares to a random encoder as a baseline. 17 Position: Evaluation of ECG Representations Must Be Fixed Table 8. Selected survey set of 12-lead ECG representation learning methods. For each method we indicate the year and venue of publication, which datasets were used for pre-training, and which datasets were used for downstream evaluation. PhysioNet 2020 includes PTB-XL, CPSC2018, INCART, and G12EC. PhysioNet 2021 includes PTB-XL, CPSC2018, INCART, and G12EC, CSN, and UMich. MethodYearVenuePretraining DatasetEvaluation Dataset CLOCS (Kiyasseh et al., 2021)2021ICMLPhysioNet 2020, ChapmanPhysioNet 2020, Chapman, Car- diology, PhysioNet 2017 3KG (Gopal et al., 2021)2021ML4HPhysioNet 2020PhysioNet 2020 ISL (Lan et al., 2022)2022AAAIPTB-XL, Chapman, CPSC2018PTB-XL, Chapman, CPSC2018 (Oh et al., 2022)2022CHILPTB-XL, CPSC2018, G12EC, CSN PTB-XL, CPSC2018, G12EC CRT (Zhang et al., 2024)2022TNNLSPTB-XL, HAR, Sleep-EDFPTB-XL, HAR, Sleep-EDF CPC (Mehari & Strodthoff, 2022)2022CIBMPhysioNet2020,Chapman, Ribeiro PTB-XL PCLR (Diamant et al., 2022)2022PLOS-CBPrivate DatasetPrivate Dataset CT-HB (Wei et al., 2022)2022ICASSPMIT-BIH, ChapmanMIT-BIH, Chapman,Private Dataset BIOT (Yang et al., 2023)2023NeurIPSSHHS, PREST, PhysioNet 2020PTB-XL, CHB-MIT, IIIC Seizure, TUAB, TUEV, HAR ASTCL (Wang et al., 2024)2023TNNLSPTB-XL,Chapman,CODE, CPSC2018, CMI PTB-XL,Chapman,CODE, CPSC2018, CMI sEHR-ECG (Lalam et al., 2023)2023TMLRPrivate DatasetPhysioNet 2020, Chapman, Pri- vate Dataset (Lai et al., 2023)2023Nat. Comms.Private DatasetPrivate Dataset, CPSC2018 METS (Li et al., 2023)2023MIDLPTB-XLPTB-XL, MIT-BIH MaeFE (Zhang et al., 2023a)2023IEEETIMCPSC2018, NingboPTB-XL, CPSC2018 sCL-ST (Le et al., 2023)2023IEEEJBHICPSC2018, INCART, G12EC, PTB PTB-XL T-S Reverse (Zhang et al., 2023b)2023BSPCPhysioNet 2017PhysioNet 2017 ST-MEM (Na et al., 2024)2024ICLRCSN, CODE-15PTB-XL, CPSC2018, PhysioNet 2017 ESI (Yu et al., 2024)2024TMLRPTB-XL, MIMIC-IV-ECG, Chap- man PTB-XL, ICBEB ETP (Liu et al., 2024a)2024ICASSPPTB-XLPTB-XL and CPSC2018 KED (Tian et al., 2024)2024Cell-RMMIMIC-IV-ECGCPSC2018, Chapman, G12EC, PTB-XL, Private Dataset MERL (Liu et al., 2024b)2024ICMLMIMIC-IVPTB-XL, CPSC2018, CSN D-BETA (Hung et al., 2024)2025ICMLMIMIC-IVPhysioNet 2021, CODE-test MELP (Wang et al., 2025)2025ICMLMIMIC-IVPTB-XL, CPSC2018, CSN H-Tuning (Zhou et al., 2025)2025ICMLCODEPTB-XL, CSN, G12EC, Private Dataset HeartLang (Jin et al., 2025)2025ICLRMIMIC-IV-ECGPTB-XL, CPSC2018, and Chap- man ECG-FM (McKeen et al., 2025)2025JAMIA OpenCPSC2018, PTB-XL, G12EC, CSN, MIMIC-IV UHN-ECG, MIMIC-IV ECGFounder (Li et al., 2025)2025NEJM-AIHarvard-Emory ECG DatabasePTB-XL, CODE-test, PhysioNet 2017, MIMIC-IV, Private Dataset CREMA (Song et al., 2025)2025CIKMMIMIC-IV, CODE-15, UKBB, SaMi-Trop, IKEM PTB-XL 18 Position: Evaluation of ECG Representations Must Be Fixed Table 9. Evaluation and reporting choices in our surveyed set of 12-lead ECG representation learning papers. βYβ indicates the practice is clearly reported, and βNβ otherwise. The column per-label indicates whether the given method reports its results per individual task, or aggregates them together. UQ indicates whether the paper uses uncertainty quantification when reporting their results. Rand indicates whether the paper compares to a random encoder as a baseline. AUROC denotes βarea under the receiver operating curveβ, AUPRC denotes βarea under the precision-recall curveβ, and Acc. denotes βaccuracyβ. MethodMetrics ReportedPer-label?UQ?Rand? CLOCS (Kiyasseh et al., 2021)AUROCNYY 3KG (Gopal et al., 2021)AUROC, F 1 YYN ISL (Lan et al., 2022)AUROCNYY (Oh et al., 2022)Acc.NYY CRT (Zhang et al., 2024)AUROC, Acc., F 1 YYN CPC (Mehari & Strodthoff, 2022)AUROCYYN PCLR (Diamant et al., 2022)F 1 , R 2 YYN CT-HB (Wei et al., 2022)AUROC, Acc., MCC, Sensitivity, Specificity, PPVYNN BIOT (Yang et al., 2023)AUROC, Acc., AUPRC, F 1 NYN ASTCL (Wang et al., 2024)AUROC, F 1 Y sEHR-ECG (Lalam et al., 2023)AUROC, AUPRCYYY (Lai et al., 2023)AUROC, AUPRC, F 1 , Specificity, Sensitivity, Acc., PPVYNN METS (Li et al., 2023)Acc., PPV, Sensitivity, F 1 NNY MaeFE (Zhang et al., 2023a)AUROC, Acc., F 1 YNN sCL-ST (Le et al., 2023)AUROC, AUPRC, Acc., F 1 , F 2 , G 2 YNN T-S Reverse (Zhang et al., 2023b)AUROC, Acc., Sensitivity, SpecificityYNN ST-MEM (Na et al., 2024)AUROC, F 1 , Acc.YYN ESI (Yu et al., 2024)AUROC, F 1 , Acc.NYY ETP (Liu et al., 2024a)AUROC, F 1 , Acc.YNY KED (Tian et al., 2024)AUROC, AUPRC, Acc., F 1 , MCC, Sensitivity, SpecificityYYN MERL (Liu et al., 2024b)AUROCNNY D-BETA (Hung et al., 2024)AUROCNYN MELP (Wang et al., 2025)AUROCNNN H-tuning (Zhou et al., 2025)AUROC, F 2 , G 2 , PPVYYN HeartLang (Jin et al., 2025)AUROCNNN ECG-FM (McKeen et al., 2025)AUROC, AUPRC, AUPRGYNY ECGFounder (Li et al., 2025)AUROC, F 1 , Acc.YYN CREMA (Song et al., 2025)AUROC, AUPRCYNY 19 Position: Evaluation of ECG Representations Must Be Fixed B. Datasets The public datasets were obtained from PhysioNet (Goldberger et al., 2000). PhysioNet provides awgetcommand for terminal-based download. The command recursively crawls PhysioNet, so transient failures in fetching directory listings can silently omit files. We encountered this issue while reproducing the data from our original submission on a separate machine: different servers produced different incomplete local copies of PTB-XL, CPSC2018, and CSN. To address this, we provide scripts in our code release that use PhysioNet checksum files to verify completeness and correctness, and iteratively redownloads missing or corrupted files. We have since reseeded our experiments and corrected the reported numbers. The results were reproduced on two machines, and none of the conclusions in our paper were affected. B.1. PTB-XL PTB-XL (Wagner et al., 2020; 2022) is a dataset of 21,837 12-lead 10-second ECG recordings from 18,885 patients. We follow the fieldβs conventional benchmarking protocol outlined in (Strodthoff et al., 2021). The dataset is often analyzed as four subsets. Each subset has a different number of records, as well as number of associated labels. SUPER includes 21,388 ECGs, SUB includes 21,388 ECGs, RHYTHM includes 21,030 ECGs, and FORM includes 8,978 ECGs. For each subset, we detail the prevalence and total number of positive examples of each label in each split in the following tables: SUPER (Table 10), SUB (Table 11), RHYTHM (Table 12), and FORM (Table 13). Table 10. Downstream tasks with their definition in PTB-XL SUPER. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test CDConduction Disturbance22.90%3907495496 HYPHypertrophy12.39%2119268262 MIMyocardial Infarction25.57%4379540550 NORMNormal ECG44.48%7596955963 STTCST/T Change24.48%4186528521 Table 11. Downstream tasks with their definition in PTB-XL SUB. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test AMIAnterior myocardial infarction14.39%2466306306 CLBBBComplete left bundle branch block2.51%4285454 CRBBBComplete right bundle branch block2.53%4325554 ILBBBIncomplete left bundle branch block0.36%6278 IMIInferior myocardial infarction15.29%2618326327 IRBBBIncomplete right bundle branch block5.23%894112112 ISCAIschemic in lateral leads4.40%7569293 ISCIIschemic in inferolateral leads1.86%3183940 ISCNon-specific ischemic5.95%1019125128 IVCDNon-specific intraventricular conduction disturbance3.68%6307879 LAFB/LPFBLeft anterior/posterior fascicular block8.40%1437181179 LAO/LAELeft atrial overload/enlargement1.99%3414342 LMILateral myocardial infarction0.94%1612020 LVHLeft ventricular hypertrophy9.97%1708210214 NORMNormal ECG44.48%7596955963 NSTNon-specific ST changes3.59%6157577 PMIPosterior myocardial infarction0.08%1322 RAO/RAERight atrial overload/enlargement0.46%791010 RVHRight ventricular hypertrophy0.59%1021212 SEHYPSeptal hypertrophy0.14%2432 STTCST/T Change10.47%1792225222 WPWWolff-Parkinson-White syndrome0.37%6478 AVBAV block3.85%6588382 20 Position: Evaluation of ECG Representations Must Be Fixed Table 12. Downstream tasks with their definition in PTB-XL RHYTHM. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test AFIBAtrial fibrillation7.20%1211151152 AFLTAtrial flutter0.35%5977 BIGUBigeminal pattern (unknown origin, SV/Ventricular)0.39%6688 PACENormal functioning artificial pacemaker1.40%2372928 PSVTParoxysmal supraventricular tachycardia0.11%1932 SARRHSinus arrhythmia3.67%6187777 SBRADSinus bradycardia3.03%5096464 SRSinus rhythm79.64%1340416701674 STACHSinus tachycardia3.93%6618382 SVARRSupraventricular arrhythmia0.75%1281514 SVTACSupraventricular tachycardia0.13%2133 TRIGUTrigeminal pattern (unknown origin, SV/Ventricular)0.10%1622 Table 13. Downstream tasks with their definition in PTB-XL FORM. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test ABQRSAbnormal QRS37.06%2683322322 DIGDigitalis-effect2.02%1451818 HVOLTHigh QRS voltage0.69%4976 INVTInverted T-waves3.27%2353029 LNGQTLong QT-interval1.30%941211 LOWTLow amplitude T-waves4.88%3504444 LPRProlonged PR interval3.79%2723434 LVOLTLow QRS voltages in the frontal and horizontal leads2.03%1451918 NDTNon-diagnostic T abnormalities20.33%1461182182 NSTNon-specific ST changes8.54%6157577 NTNon-specific T-wave changes4.71%3404142 PACAtrial premature complex4.43%3184040 PRC(S)Premature complex(es)0.11%811 PVCVentricular premature complex12.73%915114114 QWAVEQ waves present6.10%4385555 STDNon-specific ST depression11.24%807101101 STENon-specific ST elevation0.31%2233 TABT-wave abnormality0.39%2843 VCLVHVoltage criteria for left ventricular hypertrophy9.75%7018787 21 Position: Evaluation of ECG Representations Must Be Fixed B.2. CPSC2018 The China Physiological Signal Challenge 2018 (CPSC2018) (Liu et al., 2018) is a dataset of 6,877 ECGs sampled at 500 Hz. Recording duration varies between 5 and 72 seconds. We exclude recordings shorter than 10 seconds, and for longer recordings, clip them to 10 seconds. We are left with 6,867 ECGs used for downstream evaluation. The dataset is multi-label with 9 tasks. The prevalence and total number of positive examples for each split is reported in Table 14. Table 14. Downstream tasks with their definition in CPSC2018. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test AFAtrial fibrillation17.77%845124251 IAVB1st degree AV block10.50%50980132 LBBBLeft bundle branch block3.42%1752337 NSRSinus rhythm13.37%65377188 PACPremature atrial contraction8.94%42060134 PVCPremature ventricular contractions10.18%49867134 RBBBRight bundle branch block27.00%1301188365 STDST depression12.64%58995184 STEST elevation3.20%1542343 B.3. CSN The Chapman-Shaoxing-Ningbo (CSN) database (Zheng et al., 2020) is a dataset of 45,152 10-second, 12-lead ECG recordings from 10,646 patients sampled at 500 Hz. CSN is a multi-label dataset with 63 diagnostic labels. We restrict our experiments to 48 labels, excluding 13 that have no positive examples in the dataset (2AVB2, AVNRT, IDC, LBBB, LBBBB, LVQRSCL, LVQRSLL, MI, MIBW, MIFW, MILW, SAAWR, WAVN) and 2 that have fewer than three positive examples (3AVB, ABI). Diagnostic labels are derived from routine clinical interpretations. We remove ECGs containing a diagnostic code that is not found in the databaseβs code map, resulting in 31,898 recordings used for downstream evaluation. ECG recordings are randomly split at the record level into training, validation, and test sets with a ratio of 70/10/20. We note that this dataset did not include associated patient ID with each record to enable a patient-level split. The prevalence of the remaining labels, along with the number of positive examples in each split, is reported in Table 15. 22 Position: Evaluation of ECG Representations Must Be Fixed Table 15. Downstream tasks with their definition in CSN. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test 1AVB1 degree atrioventricular block2.19%47680144 2AVB2 degree atrioventricular block0.07%1426 2AVB12 degree atrioventricular block(Type one)0.05%1015 AFAtrial Flutter10.46%2388324624 AFIBAtrial Fibrillation4.43%975147291 ALSAxis left shift2.89%65188182 APBAtrial premature beats2.14%48371130 AQWAbnormal Q wave1.85%40956126 ARSAxis right shift1.66%36347121 ATAtrial Tachycardia0.54%1161541 AVBAtrioventricular block0.55%1211342 AVRTAtrioventricular Reentrant Tachycardia0.02%412 CCRCounterclockwise rotation0.44%981528 CRClockwise rotation0.24%541011 ERVEarly repolarization of the ventricles0.83%1722965 FQRSFQRS Wave0.01%111 IVBIntraventricular block1.35%3084081 JEBJunctional escape beat0.08%1833 JPTJunctional premature beat0.02%322 LFBBBLeft front bundle branch block0.66%1432244 LVHLeft ventricular hypertrophy0.35%801122 LVQRSALLower voltage QRS in all leads2.35%52975145 MISWMyocardial infarction in the side wall0.18%37910 PRIEPR interval extension0.09%2243 PWCP wave Change0.27%621015 QTIEQT interval extension0.58%1141853 RAHRight atrial hypertrophy0.02%411 RBBBRight bundle branch block1.67%37346114 RVHRight ventricle hypertrophy0.09%2045 SASinus Irregularity6.56%1445230418 SBSinus Bradycardia40.16%894712672596 SRSinus Rhythm22.31%49867121417 STSinus Tachycardia16.12%35955061039 STDDST drop down1.34%3014086 STEST extension1.30%2774099 STTCST-T Change2.75%60291183 STTUST tilt up0.46%1031727 SVTSupraventricular Tachycardia1.92%42967117 TWCT wave Change14.55%3201486953 TWOT wave opposite3.51%78296242 UWU wave0.18%39712 VBVentricular bigeminy0.01%111 VEBVentricular escape beat0.07%1425 VETVentricular escape trigeminy0.02%241 VFWVentricular fusion wave0.03%621 VPBVentricular premature beat0.74%1682049 VPEVentricular preexcitation0.04%822 WPWWolff Parkinson White Pattern0.17%39311 23 Position: Evaluation of ECG Representations Must Be Fixed B.4. EchoNext EchoNext (Poterucha et al., 2025) is a dataset collected at Columbia University Irving Medical Center of 100,000 10-second ECGs with labels derived from contemporaneous echocardiograms. The dataset includes both continuous measurements and binarized labels, the latter of which we focus on in our experiments. All ECGs were natively sampled at 250 Hz; we linearly interpolate them to 500 Hz. All ECGs in the dataset were z-scored using dataset statistics. The upper 99.9-th and lower 0.1-st percentile of voltages was clipped. The dataset mean and standard deviation were saved, which we use to re-scale all examples back to millivolts. EchoNext recommends a standard split into four mutually exclusive subsets: train, validation, test, and nosplit. The nosplit subset is treated as a hold-out set. In our experiments, we use the train, validation, and test sets, which amounts to 82,543 samples. The training set, which includes 72,475 examples, represents 26,218 patients, so includes more than one ECG per patient. However, the validation and test sets, which contain 4,626 and 5,442 examples respectively, only contain the latest ECG per patient. Table 16 contains a list of all binary labels in the EchoNext dataset. We list the prevalence for each label, as well as the total number of positive examples represented in the train, validation, and test sets. We use standard abbreviations for each task. We note that SHD is a binary label indicating any moderate or severe structural abnormality as defined by meeting the threshold for any of the other labels in this table. Table 16. Downstream tasks with their definition in EchoNext. For each label, we report the prevalence of the positive class in the whole dataset. We then report the total number of positive examples in the train, validation, and test set after splitting. TaskDescriptionPrevalence (%)N TrainN ValN Test ARModerate or severe aortic regurgitation1.228786266 ASModerate or severe aortic stenosis4.192919252286 LVEFβ€ 45Left ventricular ejection fraction isβ€ 45%22.7616962866962 LVWTβ₯ 13Max of interventricular septum/posterior wallβ₯ 1.3cm23.75176678771061 MRModerate or severe mitral regurgitation8.186137282337 PASPβ₯ 45Pulmonary artery systolic pressure isβ₯ 45 mmHg18.1813727581699 PEFFPresence of a moderate or large pericardial effusion2.6720795269 PRModerate or severe pulmonary regurgitation0.786032120 RVSDModerate or severe right ventricular systolic dysfunction12.589597368419 SHDAny moderate or severe structural heart disease51.203795819902318 TR-MAXβ₯ 32Maximum tricuspid regurgitation velocity isβ₯ 3.2 m/s9.857492267375 TRModerate or severe tricuspid regurgitation10.137707305353 24 Position: Evaluation of ECG Representations Must Be Fixed B.5. Hemodynamic Inference For hemodynamic inference we use a private dataset of 9,226 10-second 12-lead ECGs collected from 5,072 patients at Massachusetts General Hospital (MGH), originally introduced by (Schlesinger et al., 2022). Each ECG is paired with contemporaneous invasive hemodynamic measurements obtained via right heart catheterization, which serve as ground-truth labels. We consider two binary classification tasks: inferring elevated mean pulmonary capillary wedge pressure (mPCWP) and elevated mean pulmonary artery pressure (mPA). Measurements are binarized with mPCWPβ₯ 15 mmHg and mPAβ₯ 20 mmHg indicating a positive label. We construct patient-level splits to avoid information leakage across sets. Patients are randomly divided into training (70%), validation (10%), and test (20%) cohorts. The training set may contain multiple ECGs per patient. We only retain one ECG per patient in the validation and test set; if multiple ECGs are present for a given patient, then a single ECG is selected uniformly at random. After processing, we are left with 6,458 examples in the training set, 507 in the validation set, and 1,015 in the test set. Table 17 summarizes the two downstream hemodynamic inference tasks, including label definitions, overall prevalence, and the number of positive examples in each split. Table 17. Downstream hemodynamic inference tasks. For each binary label, we report the prevalence of the positive class in the full dataset, as well as the number of positive examples in the train, validation, and test sets. TaskDescriptionPrevalence (%)N TrainN ValN Test mPAMean pulmonary arterial pressureβ₯ 20mmHg measured by right heart catheterization 68.16 4,297393751 mPCWPMean pulmonary capillary wedge pressureβ₯ 15 mmHg measured by right heart catheterization 49.303,066306561 B.6. Patient Forecasting For patient forecasting we evaluate on the risk of developing heart failure within 1 year of an ECG (1YR-HF). In particular, we frame this as a binary prediction task with the outcome defined with echocardiographic ground truth as a left ventricular ejection fraction (LVEF) below 40%. We rely on a private longitudinal dataset collected at Massachusetts General Hospital (MGH), originally introduced by (Bergamaschi et al., 2025). The full dataset contains 913,420 10-second 12-lead ECGs from 82,244 patients and is designed to support long-term outcome prediction from ECGs. After filtering examples that have data for the given task, and those with NaN or Inf values, we are left with 426,081 ECGs from 46,694 patients. We use patient-level data splits following the original dataset construction, assigning all ECGs from a given patient to the same split to prevent information leakage across sets. Patients are split into training (75%), validation (10%), and test (15%) cohorts, with all ECGs from a given patient assigned to the same split. Table 18 summarizes the patient forecasting task, including the label definition, prevalence, and number of positive examples in each split. Table 18. Patient forecasting task. For the binary outcome, we report the prevalence of the positive class in the full dataset, as well as the number of positive examples in the train, validation, and test sets. TaskDescriptionPrevalence (%)N TrainN ValN Test 1YR-HFDevelopment of heart failure within one year, de- fined as left ventricular ejection fraction< 40%on an echocardiogram 29.49 93,96012,92918,774 25 Position: Evaluation of ECG Representations Must Be Fixed C. Training Details C.1. Linear Probing To evaluate the quality of learned representations, we freeze each encoder then perform linear probing on the embeddings. We train a separateβ 2 -regularized logistic regression model for each label. For our implementation we usescikit-learn. We select hyperparameters using a grid search on the validation set. The sweep considers three hyperparameters:scale, inversestrength, andclassweight. Thescalehyperparameter controls whether feature standardization is applied. When enabled, each embedding dimension is rescaled to zero mean and unit variance using statistics computed on the training set only. The hyperparameterinversestrengthcontrols the strength ofβ 2 regularization in the logistic regression classifier. Smaller values ofinversestrengthcorrespond to stronger regularization, while larger values allow the classifier to fit the training data more closely. Theclassweighthyperparameter determines how class imbalance is handled during training. When set to βbalancedβ, examples are weighted inversely proportional to their class frequencies in the training data, giving greater weight to examples from the minority class. Otherwise, no class weighting is applied, so each training example contributes equally to the loss. We swept over scaleβTrue, False inverse strengthβ0.01, 0.1, 1.0, 10.0 class weightβNone, balanced. All logistic regression models are trained to convergence with a maximum of 10,000 optimization iterations. The best- performing configuration is selected based on validation loss and is evaluated once on the held-out test set. C.2. CLOCS Background. CLOCS (Kiyasseh et al., 2021) is a popular ECG-specific self-supervised learning approach that constructs contrastive pairs directly from the temporal structure and lead organization of ECG signals. We use the publicly available CLOCS implementation repository and make several modifications, all of which are open-sourced in our code release. CLOCS defines a family of contrastive objectives, including contrastive multi-segment coding (CMSC), contrastive multi- lead coding (CMLC), and contrastive multi-segment multi-lead coding (CMSMLC). CMSC constructs positive pairs by sampling multiple temporal segments from the same ECG recording, encouraging representations to be invariant to temporal cropping. CMLC instead constructs positive pairs across different leads of the same ECG, encouraging invariance across lead views. CMSMLC combines both objectives by simultaneously contrasting multiple temporal segments and multiple leads from the same ECG. In the original CLOCS paper, CMSC is reported to achieve the strongest average performance across downstream tasks, and we therefore focus on CMSC for comparison. In the released CLOCS code base, CMSC is implemented only for single-lead ECGs. To support 12-lead ECG pre-training, we extend this objective by treating each lead as a separate channel, consistent with the approach of (Oh et al., 2022). Training Formulation Given a 10-second ECG recordingx βR 12ΓT , whereTis the number of samples, we split the recording in half, producing two 5-second segmentsx (1) andx (2) . Each segment is encoded using a shared encoder network, producing β 2 -normalized embeddings Μ z (1) , Μ z (2) βR E , where E is the embedding dimension. Given a mini-batch of B ECG recordings, we define the cosine similarity between embeddings with temperature Ο as s ij = ( Μ z (1) i ) β€ Μ z (2) j Ο . Segments originating from the same ECG form positive pairs, while segments from different ECGs in the batch serve as negative pairs. We then train CMSC by optimizing a symmetric InfoNCE objective: L =β 1 2B B X i=1 " log exp(s i ) P B j=1 exp(s ij ) + log exp(s i ) P B j=1 exp(s ji ) # . Training Pre-training is performed on raw 10-second ECG recordings sampled at 500 Hz from the MIMIC-IV dataset, after removing ECGs containingNaNorInfvalues. The resulting dataset is split into training and validation sets using a 26 Position: Evaluation of ECG Representations Must Be Fixed 90/10 split, corresponding to 719,394 ECGs for training and 80,641 ECGs for validation. The final model checkpoint is selected based on the minimum contrastive loss on the validation set. Models are trained for 50 epochs, corresponding to a comparable number of optimization steps as used in MERL and D-BETA. We use the same CNN encoder architecture as the original CLOCS implementation and train using the Adam optimizer (Kingma & Ba, 2017). The original CLOCS paper does not report a hyperparameter grid search and instead fixes the learning rate to10 β4 , the temperature toΟ = 0.1, and the batch size to 256. We adopt the same batch size and temperature, and perform a grid search over learning rates10 β3 , 10 β4 , 10 β5 using the validation set. Based on validation loss, we ultimately selected the model trained with a learning rate of10 β4 . We do not apply signal perturbations during pre-training, as the main results in the CLOCS paper are reported without perturbations. Evaluation Because the CMSC encoder is trained on 5-second views, we apply it to downstream 10-second ECGs at the segment level. We split each 10-second ECG recordingxin half, producing two non-overlapping 5-second segments,x (1) andx (2) . Each 5-second segment is independently encoded, and those embeddings are used to train a logistic regression classifier, producing class probability vectorsp (1) andp (2) . The final prediction is obtained by averaging the two probability vectors: p = 1 2 p (1) +p (2) . 27 Position: Evaluation of ECG Representations Must Be Fixed D. Complete Evaluation on PTB-XL In Appendix D.1 we show the AUROC and AUPRC for every task within each of the four PTB-XL datasets. In Appendix D.2 we show the macro-AUROC on PTB-XL SUB, RHYTHM, FORM before and after removing tasks with fewer than 10 positive labels in the test set. We do not show a table for SUPER since each task has far more than 10 positive labels in the test set. D.1. Performance On Standard Splits Table 19. Performance for every task in PTB-XL SUPER. Each cell reports AUROC (top line) and AUPRC (bottom line) with 95% confidence intervals. TaskRandomCLOCSKEDHeartLangMERLD-BETA CD0.843 0.821β0.863 / 0.724 0.689β0.755 0.794 0.767β0.818 / 0.630 0.590β0.666 0.904 0.887β0.919 / 0.805 0.776β0.831 0.802 0.779β0.824 / 0.637 0.600β0.673 0.893 0.876β0.909 / 0.796 0.765β0.822 0.902 0.885β0.917 / 0.815 0.788β0.841 HYP0.847 0.819β0.873 / 0.565 0.510β0.620 0.840 0.811β0.867 / 0.547 0.486β0.601 0.896 0.875β0.915 / 0.669 0.614β0.714 0.782 0.752β0.813 / 0.386 0.336β0.441 0.873 0.848β0.896 / 0.611 0.556β0.661 0.813 0.786β0.839 / 0.456 0.395β0.510 MI0.841 0.821β0.861 / 0.655 0.616β0.698 0.744 0.719β0.767 / 0.563 0.524β0.597 0.885 0.868β0.900 / 0.764 0.732β0.792 0.801 0.780β0.822 / 0.585 0.546β0.628 0.887 0.871β0.902 / 0.777 0.747β0.802 0.905 0.892β0.919 / 0.815 0.791β0.838 NORM0.891 0.877β0.904 / 0.839 0.814β0.863 0.859 0.843β0.873 / 0.789 0.764β0.815 0.932 0.922β0.942 / 0.901 0.884β0.917 0.875 0.860β0.889 / 0.822 0.797β0.846 0.927 0.916β0.938 / 0.885 0.862β0.904 0.930 0.919β0.939 / 0.895 0.875β0.913 STTC0.881 0.864β0.896 / 0.701 0.660β0.737 0.866 0.848β0.882 / 0.688 0.647β0.726 0.927 0.915β0.938 / 0.808 0.776β0.838 0.860 0.844β0.877 / 0.637 0.598β0.677 0.924 0.911β0.936 / 0.799 0.764β0.830 0.916 0.901β0.928 / 0.786 0.753β0.817 28 Position: Evaluation of ECG Representations Must Be Fixed Table 20. Performance for every task in PTB-XL SUB. Each cell reports AUROC (top line) and AUPRC (bottom line) with 95% confidence intervals. TaskRandomCLOCSKEDHeartLangMERLD-BETA AMI0.894 0.874β0.915 / 0.646 0.593β0.700 0.857 0.832β0.883 / 0.590 0.537β0.646 0.951 0.939β0.962 / 0.808 0.770β0.845 0.876 0.857β0.895 / 0.548 0.496β0.601 0.959 0.949β0.969 / 0.841 0.809β0.873 0.956 0.943β0.967 / 0.834 0.793β0.869 CLBBB0.973 0.934β0.999 / 0.880 0.784β0.955 0.984 0.974β0.993 / 0.676 0.551β0.794 0.998 0.997β1.000 / 0.949 0.906β0.983 0.993 0.984β0.998 / 0.867 0.767β0.944 0.999 0.997β1.000 / 0.947 0.889β0.988 0.999 0.997β1.000 / 0.970 0.937β0.995 CRBBB0.989 0.979β0.996 / 0.789 0.691β0.875 0.982 0.970β0.992 / 0.749 0.647β0.843 0.998 0.996β0.999 / 0.910 0.838β0.966 0.964 0.936β0.983 / 0.632 0.490β0.753 0.998 0.997β0.999 / 0.891 0.788β0.979 0.997 0.995β0.999 / 0.834 0.733β0.928 ILBBB0.899 0.733β0.988 / 0.197 0.045β0.499 0.916 0.812β0.979 / 0.064 0.025β0.119 0.913 0.758β0.994 / 0.206 0.071β0.472 0.876 0.769β0.975 / 0.066 0.019β0.140 0.886 0.683β0.987 / 0.094 0.044β0.160 0.935 0.818β0.993 / 0.211 0.071β0.470 IMI0.864 0.840β0.884 / 0.556 0.506β0.609 0.674 0.644β0.702 / 0.299 0.262β0.341 0.881 0.862β0.899 / 0.597 0.549β0.647 0.750 0.721β0.776 / 0.353 0.310β0.399 0.843 0.819β0.864 / 0.538 0.488β0.589 0.897 0.880β0.914 / 0.691 0.649β0.734 IRBBB0.856 0.815β0.891 / 0.341 0.262β0.433 0.810 0.770β0.850 / 0.227 0.173β0.292 0.955 0.935β0.970 / 0.617 0.530β0.697 0.770 0.726β0.813 / 0.187 0.137β0.246 0.971 0.961β0.980 / 0.683 0.602β0.757 0.929 0.905β0.949 / 0.517 0.431β0.602 ISCA0.803 0.762β0.840 / 0.158 0.114β0.209 0.848 0.809β0.884 / 0.224 0.165β0.294 0.927 0.910β0.944 / 0.350 0.274β0.447 0.794 0.753β0.833 / 0.124 0.100β0.154 0.903 0.871β0.931 / 0.331 0.255β0.423 0.923 0.905β0.938 / 0.275 0.218β0.346 ISCI0.854 0.791β0.903 / 0.175 0.089β0.292 0.738 0.663β0.810 / 0.093 0.037β0.184 0.919 0.883β0.949 / 0.306 0.194β0.442 0.767 0.700β0.830 / 0.073 0.044β0.121 0.842 0.768β0.908 / 0.281 0.152β0.419 0.915 0.876β0.946 / 0.292 0.178β0.443 ISC0.911 0.879β0.940 / 0.538 0.453β0.623 0.923 0.899β0.943 / 0.507 0.429β0.588 0.966 0.952β0.977 / 0.703 0.635β0.769 0.911 0.893β0.929 / 0.411 0.336β0.488 0.954 0.933β0.971 / 0.674 0.601β0.746 0.940 0.923β0.954 / 0.539 0.459β0.617 IVCD0.696 0.636β0.759 / 0.122 0.079β0.185 0.622 0.564β0.686 / 0.072 0.051β0.105 0.717 0.655β0.782 / 0.142 0.095β0.206 0.698 0.640β0.753 / 0.090 0.063β0.124 0.749 0.685β0.815 / 0.192 0.127β0.273 0.758 0.700β0.815 / 0.145 0.097β0.210 LAFB/LPFB0.941 0.919β0.960 / 0.719 0.652β0.776 0.863 0.833β0.890 / 0.440 0.372β0.511 0.962 0.948β0.974 / 0.767 0.708β0.824 0.878 0.853β0.901 / 0.454 0.384β0.525 0.923 0.903β0.940 / 0.604 0.530β0.674 0.951 0.933β0.968 / 0.770 0.714β0.821 LAO/LAE0.690 0.612β0.759 / 0.042 0.029β0.060 0.648 0.567β0.725 / 0.040 0.027β0.062 0.828 0.771β0.881 / 0.135 0.070β0.228 0.712 0.627β0.797 / 0.088 0.041β0.167 0.818 0.749β0.880 / 0.146 0.083β0.244 0.835 0.778β0.882 / 0.114 0.063β0.192 LMI0.815 0.717β0.903 / 0.072 0.033β0.130 0.657 0.538β0.770 / 0.019 0.012β0.030 0.756 0.657β0.845 / 0.085 0.021β0.203 0.744 0.657β0.825 / 0.041 0.017β0.086 0.741 0.642β0.826 / 0.027 0.016β0.044 0.771 0.703β0.837 / 0.025 0.017β0.037 LVH0.917 0.896β0.935 / 0.648 0.593β0.704 0.900 0.875β0.921 / 0.593 0.528β0.650 0.949 0.935β0.961 / 0.745 0.695β0.794 0.823 0.794β0.853 / 0.405 0.343β0.472 0.929 0.912β0.945 / 0.680 0.623β0.734 0.862 0.837β0.885 / 0.489 0.426β0.552 NORM0.891 0.877β0.903 / 0.839 0.813β0.863 0.859 0.843β0.872 / 0.788 0.763β0.812 0.932 0.921β0.942 / 0.900 0.883β0.916 0.875 0.860β0.890 / 0.822 0.798β0.846 0.926 0.915β0.937 / 0.885 0.862β0.904 0.929 0.919β0.939 / 0.894 0.875β0.912 NST0.710 0.649β0.768 / 0.136 0.081β0.213 0.725 0.668β0.780 / 0.142 0.084β0.215 0.835 0.791β0.876 / 0.190 0.126β0.267 0.743 0.693β0.793 / 0.111 0.074β0.167 0.840 0.796β0.880 / 0.193 0.135β0.267 0.838 0.794β0.880 / 0.180 0.123β0.253 PMI0.521 0.173β0.874 / 0.003 0.001β0.007 0.632 0.586β0.679 / 0.002 0.002β0.003 0.932 0.858β0.999 / 0.159 0.006β0.504 0.843 0.679β0.999 / 0.110 0.003β0.400 0.995 0.988β1.000 / 0.241 0.071β0.667 0.797 0.595β0.992 / 0.031 0.002β0.105 RAO/RAE0.866 0.738β0.964 / 0.052 0.022β0.104 0.766 0.556β0.924 / 0.031 0.011β0.075 0.961 0.927β0.988 / 0.273 0.071β0.541 0.919 0.871β0.961 / 0.104 0.027β0.287 0.952 0.902β0.992 / 0.303 0.099β0.586 0.910 0.845β0.968 / 0.110 0.028β0.287 RVH0.862 0.749β0.961 / 0.405 0.136β0.691 0.894 0.725β0.988 / 0.236 0.085β0.456 0.966 0.931β0.990 / 0.294 0.104β0.541 0.903 0.791β0.979 / 0.318 0.080β0.598 0.863 0.683β0.991 / 0.338 0.111β0.594 0.966 0.935β0.986 / 0.211 0.076β0.396 SEHYP0.963 0.942β0.982 / 0.025 0.016β0.050 0.861 0.756β0.956 / 0.009 0.004β0.021 0.993 0.989β0.996 / 0.104 0.061β0.182 0.840 0.701β0.975 / 0.013 0.003β0.036 0.917 0.872β0.958 / 0.011 0.007β0.022 0.889 0.850β0.925 / 0.008 0.006β0.012 STTC0.827 0.801β0.851 / 0.319 0.275β0.368 0.823 0.796β0.849 / 0.319 0.279β0.366 0.893 0.873β0.911 / 0.507 0.444β0.570 0.799 0.772β0.825 / 0.295 0.250β0.346 0.885 0.863β0.905 / 0.480 0.421β0.539 0.862 0.837β0.887 / 0.429 0.370β0.490 WPW0.710 0.490β0.909 / 0.120 0.007β0.385 0.777 0.584β0.924 / 0.031 0.009β0.083 0.968 0.924β0.996 / 0.460 0.142β0.796 0.895 0.753β0.983 / 0.261 0.028β0.543 0.947 0.849β0.999 / 0.621 0.306β0.904 0.840 0.622β0.980 / 0.439 0.135β0.811 AVB0.743 0.679β0.802 / 0.123 0.089β0.174 0.707 0.649β0.767 / 0.130 0.083β0.196 0.957 0.940β0.971 / 0.564 0.462β0.664 0.852 0.812β0.884 / 0.223 0.157β0.303 0.964 0.949β0.976 / 0.575 0.479β0.684 0.976 0.965β0.984 / 0.625 0.517β0.734 29 Position: Evaluation of ECG Representations Must Be Fixed Table 21. Performance for every task in PTB-XL RHYTHM. Each cell reports AUROC (top line) and AUPRC (bottom line) with 95% confidence intervals. TaskRandomCLOCSKEDHeartLangMERLD-BETA AFIB0.858 0.827β0.888 / 0.384 0.318β0.463 0.836 0.803β0.867 / 0.337 0.277β0.410 0.984 0.967β0.995 / 0.936 0.903β0.964 0.928 0.908β0.947 / 0.577 0.500β0.651 0.981 0.965β0.992 / 0.891 0.839β0.937 0.984 0.967β0.997 / 0.955 0.924β0.979 AFLT0.948 0.897β0.991 / 0.231 0.046β0.531 0.911 0.748β0.997 / 0.302 0.090β0.632 0.966 0.906β0.999 / 0.508 0.178β0.824 0.795 0.527β0.999 / 0.451 0.117β0.801 0.969 0.914β0.999 / 0.611 0.277β0.901 0.950 0.859β1.000 / 0.582 0.241β0.883 BIGU0.636 0.405β0.871 / 0.102 0.011β0.286 0.703 0.498β0.896 / 0.147 0.006β0.408 0.955 0.893β0.996 / 0.445 0.139β0.751 0.884 0.790β0.969 / 0.106 0.020β0.309 0.744 0.541β0.926 / 0.148 0.015β0.427 0.975 0.960β0.989 / 0.350 0.071β0.669 PACE0.905 0.830β0.967 / 0.589 0.410β0.759 0.889 0.801β0.957 / 0.428 0.252β0.597 0.971 0.937β0.994 / 0.733 0.579β0.864 0.905 0.810β0.979 / 0.551 0.371β0.721 0.962 0.900β0.997 / 0.796 0.657β0.916 0.993 0.985β0.999 / 0.887 0.776β0.973 PSVT0.921 0.834β0.998 / 0.096 0.006β0.333 0.998 0.994β1.000 / 0.613 0.143β1.000 0.999 0.996β1.000 / 0.664 0.200β1.000 0.998 0.995β1.000 / 0.626 0.166β1.000 0.998 0.995β1.000 / 0.477 0.167β1.000 0.999 0.996β1.000 / 0.664 0.200β1.000 SARRH0.644 0.581β0.708 / 0.079 0.053β0.114 0.611 0.552β0.676 / 0.063 0.046β0.092 0.873 0.837β0.907 / 0.233 0.170β0.309 0.625 0.557β0.683 / 0.059 0.045β0.079 0.690 0.631β0.743 / 0.099 0.065β0.146 0.950 0.917β0.970 / 0.515 0.414β0.618 SBRAD0.791 0.733β0.850 / 0.173 0.105β0.263 0.891 0.846β0.932 / 0.430 0.320β0.559 0.968 0.953β0.980 / 0.677 0.571β0.775 0.938 0.907β0.962 / 0.487 0.373β0.595 0.962 0.945β0.976 / 0.565 0.459β0.675 0.933 0.885β0.967 / 0.621 0.519β0.726 SR0.690 0.661β0.718 / 0.872 0.856β0.888 0.784 0.755β0.810 / 0.915 0.900β0.929 0.933 0.919β0.947 / 0.979 0.973β0.984 0.841 0.817β0.863 / 0.941 0.928β0.952 0.890 0.871β0.909 / 0.962 0.953β0.970 0.949 0.936β0.962 / 0.975 0.963β0.986 STACH0.864 0.825β0.899 / 0.224 0.168β0.297 0.975 0.967β0.983 / 0.594 0.490β0.697 0.992 0.980β0.999 / 0.945 0.899β0.982 0.987 0.976β0.995 / 0.851 0.775β0.912 0.993 0.989β0.996 / 0.860 0.790β0.922 0.989 0.968β0.999 / 0.922 0.859β0.971 SVARR0.862 0.779β0.926 / 0.040 0.022β0.066 0.602 0.507β0.692 / 0.009 0.007β0.012 0.924 0.869β0.971 / 0.344 0.124β0.580 0.815 0.753β0.875 / 0.023 0.015β0.037 0.934 0.865β0.980 / 0.425 0.193β0.668 0.910 0.811β0.981 / 0.332 0.127β0.579 SVTAC0.819 0.643β0.999 / 0.182 0.003β0.669 0.962 0.903β0.999 / 0.211 0.015β0.707 0.996 0.991β0.999 / 0.251 0.111β0.533 0.966 0.908β0.999 / 0.234 0.015β0.744 0.993 0.987β0.998 / 0.213 0.078β0.600 0.999 0.997β1.000 / 0.474 0.224β1.000 TRIGU0.505 0.091β0.865 / 0.003 0.001β0.007 0.559 0.156β0.905 / 0.004 0.001β0.010 0.942 0.891β0.994 / 0.043 0.009β0.143 0.562 0.503β0.623 / 0.002 0.002β0.003 0.726 0.461β0.950 / 0.008 0.002β0.019 0.989 0.981β0.995 / 0.075 0.045β0.154 30 Position: Evaluation of ECG Representations Must Be Fixed Table 22. Performance for every task in PTB-XL FORM. Each cell reports AUROC (top line) and AUPRC (bottom line) with 95% confidence intervals. TaskRandomCLOCSKEDHeartLangMERLD-BETA ABQRS0.788 0.757β0.819 / 0.687 0.639β0.736 0.684 0.649β0.717 / 0.538 0.491β0.583 0.829 0.798β0.858 / 0.741 0.697β0.785 0.731 0.696β0.765 / 0.615 0.568β0.661 0.811 0.780β0.839 / 0.720 0.675β0.766 0.785 0.757β0.814 / 0.648 0.602β0.698 DIG0.755 0.643β0.852 / 0.070 0.039β0.119 0.628 0.491β0.756 / 0.045 0.023β0.087 0.907 0.861β0.949 / 0.326 0.151β0.534 0.714 0.583β0.816 / 0.061 0.032β0.126 0.878 0.819β0.929 / 0.149 0.076β0.269 0.882 0.806β0.943 / 0.239 0.111β0.431 HVOLT0.907 0.839β0.968 / 0.078 0.030β0.184 0.829 0.742β0.911 / 0.046 0.017β0.131 0.943 0.916β0.967 / 0.080 0.049β0.143 0.808 0.622β0.926 / 0.035 0.015β0.062 0.936 0.888β0.973 / 0.089 0.043β0.169 0.907 0.847β0.948 / 0.052 0.030β0.083 INVT0.864 0.819β0.908 / 0.161 0.104β0.250 0.864 0.811β0.910 / 0.161 0.103β0.240 0.899 0.854β0.940 / 0.421 0.258β0.595 0.765 0.679β0.845 / 0.145 0.078β0.244 0.910 0.872β0.945 / 0.317 0.183β0.471 0.915 0.880β0.948 / 0.353 0.213β0.519 LNGQT0.679 0.495β0.847 / 0.046 0.018β0.105 0.690 0.535β0.818 / 0.037 0.017β0.082 0.804 0.638β0.941 / 0.218 0.051β0.494 0.701 0.544β0.858 / 0.058 0.020β0.151 0.850 0.716β0.972 / 0.348 0.125β0.620 0.898 0.818β0.965 / 0.229 0.067β0.455 LOWT0.745 0.666β0.823 / 0.145 0.100β0.202 0.726 0.654β0.794 / 0.112 0.081β0.155 0.831 0.775β0.881 / 0.240 0.157β0.344 0.687 0.616β0.757 / 0.122 0.075β0.191 0.833 0.789β0.874 / 0.214 0.139β0.320 0.843 0.795β0.885 / 0.186 0.136β0.260 LPR0.610 0.510β0.708 / 0.072 0.045β0.123 0.588 0.478β0.692 / 0.063 0.042β0.097 0.934 0.891β0.967 / 0.436 0.300β0.589 0.773 0.701β0.841 / 0.198 0.099β0.322 0.958 0.937β0.975 / 0.504 0.344β0.665 0.964 0.947β0.978 / 0.446 0.327β0.598 LVOLT0.937 0.907β0.960 / 0.216 0.124β0.371 0.883 0.830β0.928 / 0.124 0.070β0.230 0.940 0.908β0.966 / 0.249 0.138β0.424 0.771 0.621β0.889 / 0.212 0.066β0.380 0.930 0.893β0.959 / 0.220 0.119β0.393 0.809 0.725β0.876 / 0.080 0.044β0.143 NDT0.810 0.777β0.842 / 0.507 0.446β0.577 0.800 0.769β0.831 / 0.455 0.399β0.513 0.896 0.873β0.918 / 0.666 0.605β0.731 0.761 0.725β0.793 / 0.432 0.374β0.486 0.891 0.868β0.913 / 0.693 0.632β0.753 0.866 0.839β0.892 / 0.612 0.549β0.682 NST0.651 0.587β0.712 / 0.189 0.125β0.270 0.644 0.575β0.712 / 0.205 0.139β0.292 0.740 0.679β0.799 / 0.252 0.184β0.338 0.639 0.578β0.698 / 0.175 0.119β0.247 0.766 0.708β0.816 / 0.280 0.205β0.363 0.758 0.703β0.810 / 0.259 0.192β0.352 NT0.755 0.690β0.818 / 0.138 0.089β0.205 0.671 0.600β0.740 / 0.087 0.063β0.128 0.806 0.750β0.858 / 0.208 0.125β0.327 0.693 0.618β0.774 / 0.099 0.071β0.137 0.878 0.843β0.908 / 0.223 0.158β0.320 0.828 0.764β0.882 / 0.215 0.141β0.305 PAC0.614 0.518β0.709 / 0.090 0.056β0.154 0.628 0.551β0.705 / 0.080 0.055β0.126 0.934 0.894β0.964 / 0.566 0.432β0.704 0.684 0.596β0.760 / 0.101 0.068β0.148 0.720 0.642β0.796 / 0.147 0.087β0.244 0.985 0.975β0.992 / 0.687 0.547β0.825 PRC(S)0.923 0.907β0.941 / 0.015 0.012β0.019 0.850 0.826β0.874 / 0.008 0.006β0.009 0.967 0.956β0.977 / 0.034 0.025β0.048 0.867 0.845β0.889 / 0.009 0.007β0.010 0.761 0.732β0.790 / 0.005 0.004β0.005 0.970 0.959β0.981 / 0.038 0.027β0.056 PVC0.837 0.794β0.879 / 0.535 0.443β0.628 0.797 0.746β0.844 / 0.452 0.364β0.544 0.963 0.949β0.975 / 0.809 0.738β0.870 0.938 0.915β0.957 / 0.759 0.688β0.826 0.910 0.881β0.937 / 0.665 0.585β0.745 0.995 0.991β0.998 / 0.962 0.937β0.983 QWAVE0.664 0.587β0.735 / 0.156 0.099β0.231 0.574 0.495β0.650 / 0.113 0.072β0.179 0.736 0.667β0.803 / 0.191 0.134β0.269 0.636 0.562β0.711 / 0.140 0.088β0.217 0.712 0.658β0.764 / 0.124 0.093β0.171 0.766 0.706β0.823 / 0.193 0.131β0.276 STD0.750 0.701β0.796 / 0.272 0.218β0.337 0.659 0.601β0.715 / 0.235 0.180β0.296 0.811 0.769β0.847 / 0.360 0.289β0.440 0.668 0.612β0.726 / 0.239 0.184β0.310 0.809 0.769β0.849 / 0.392 0.317β0.476 0.753 0.706β0.798 / 0.285 0.227β0.352 STE0.642 0.406β0.783 / 0.008 0.005β0.014 0.562 0.100β0.908 / 0.011 0.004β0.036 0.585 0.295β0.936 / 0.014 0.004β0.051 0.612 0.202β0.930 / 0.014 0.004β0.047 0.921 0.840β0.993 / 0.112 0.018β0.429 0.710 0.511β0.949 / 0.018 0.006β0.062 TAB0.562 0.390β0.838 / 0.008 0.004β0.021 0.696 0.591β0.797 / 0.009 0.006β0.016 0.741 0.515β0.976 / 0.034 0.006β0.130 0.273 0.135β0.498 / 0.004 0.003β0.007 0.868 0.797β0.974 / 0.038 0.013β0.120 0.662 0.235β0.954 / 0.021 0.004β0.070 VCLVH0.866 0.827β0.902 / 0.437 0.346β0.535 0.809 0.758β0.855 / 0.326 0.257β0.406 0.905 0.872β0.931 / 0.505 0.412β0.610 0.707 0.651β0.763 / 0.211 0.162β0.272 0.857 0.816β0.895 / 0.416 0.328β0.516 0.766 0.718β0.816 / 0.299 0.229β0.389 31 Position: Evaluation of ECG Representations Must Be Fixed D.2. PTB-XL Results When Labels with Minimal Examples Are Removed Because macro-AUROC weights the performance on each label equally, omitting a few classes with extremely small test support can disproportionately affect whichever method happened to perform the worst on those labels. In PTB-XL SUB (Table 23) and PTB-XL RHYTHM (Table 24), Random experiences the largest swing in macro-AUROC after evaluating on the cleaned data. We donβt think this is an intrinsic property of the Random encoder. For example, in PTB-XL FORM, all other methods changed more than Random (Table 25). Table 23. ORIG uses the standard PTB-XL SUB test set; CLEAN excludes labels with fewer than 10 positive test examples. Values are macro-AUROC with 95% confidence intervals. MethodORIGCLEANβMacro-AUROC RANDOM0.834 0.808β0.8590.847 0.834β0.860+0.013 CLOCS0.803 0.784β0.8190.803 0.787β0.820+0.000 KED0.920 0.909β0.9290.913 0.905β0.921-0.007 HEARTLANG0.836 0.819β0.8530.830 0.819β0.841-0.006 MERL0.905 0.890β0.9180.897 0.883β0.909-0.007 D-BETA0.899 0.882β0.9140.906 0.899β0.914+0.007 Table 24. ORIG uses the standard PTB-XL RHYTHM test set; CLEAN excludes labels with fewer than 10 positive test examples. Values are macro-AUROC with 95% confidence intervals. MethodORIGCLEANβMacro-AUROC RANDOM0.787 0.737β0.8330.802 0.781β0.823+0.015 CLOCS0.810 0.762β0.8540.798 0.775β0.819-0.013 KED0.959 0.948β0.9690.949 0.938β0.959-0.009 HEARTLANG0.854 0.827β0.8790.861 0.841β0.879+0.008 MERL0.903 0.870β0.9330.914 0.898β0.927+0.011 D-BETA0.968 0.956β0.9780.958 0.941β0.971-0.010 Table 25. ORIG uses the standard PTB-XL FORM test set; CLEAN excludes labels with fewer than 10 positive test examples. Values are macro-AUROC with 95% confidence intervals. MethodORIGCLEANβMacro-AUROC RANDOM0.756 0.734β0.7800.754 0.733β0.774-0.002 CLOCS0.715 0.687β0.7410.710 0.690β0.731-0.005 KED0.851 0.827β0.8750.862 0.846β0.876+0.011 HEARTLANG0.707 0.679β0.7360.723 0.699β0.745+0.016 MERL0.853 0.839β0.8660.847 0.833β0.860-0.005 D-BETA0.845 0.821β0.8680.855 0.841β0.868+0.009 32 Position: Evaluation of ECG Representations Must Be Fixed E. Complete Evaluation on CPSC2018 Table 26. Performance on CPSC2018 sub-tasks. Each cell reports AUROC (top line) and AUPRC (bottom line), each with a 95% confidence interval. TaskRandomCLOCSKEDHeartLangMERLD-BETA AF0.873 0.848β0.897 / 0.660 0.605β0.714 0.837 0.807β0.862 / 0.617 0.565β0.669 0.987 0.979β0.993 / 0.957 0.938β0.973 0.914 0.897β0.930 / 0.700 0.650β0.750 0.979 0.967β0.989 / 0.947 0.923β0.968 0.995 0.993β0.998 / 0.983 0.974β0.991 IAVB0.774 0.732β0.814 / 0.296 0.235β0.366 0.706 0.660β0.751 / 0.211 0.167β0.265 0.987 0.980β0.993 / 0.922 0.889β0.950 0.818 0.780β0.859 / 0.437 0.356β0.521 0.984 0.974β0.992 / 0.916 0.880β0.946 0.993 0.986β0.997 / 0.954 0.927β0.975 LBBB0.946 0.868β0.999 / 0.839 0.693β0.965 0.976 0.960β0.991 / 0.758 0.638β0.877 0.996 0.990β0.999 / 0.931 0.861β0.983 0.962 0.914β0.997 / 0.855 0.751β0.944 0.987 0.967β0.998 / 0.898 0.820β0.964 0.998 0.996β1.000 / 0.957 0.910β0.991 NSR0.893 0.871β0.913 / 0.593 0.532β0.658 0.846 0.820β0.872 / 0.461 0.400β0.523 0.958 0.946β0.969 / 0.804 0.752β0.850 0.893 0.870β0.916 / 0.585 0.516β0.658 0.951 0.938β0.963 / 0.760 0.699β0.820 0.955 0.943β0.965 / 0.762 0.702β0.816 PAC0.706 0.659β0.749 / 0.196 0.160β0.240 0.597 0.554β0.641 / 0.121 0.106β0.141 0.861 0.828β0.892 / 0.470 0.390β0.556 0.659 0.610β0.709 / 0.178 0.144β0.222 0.779 0.734β0.821 / 0.334 0.263β0.407 0.927 0.898β0.952 / 0.703 0.622β0.784 PVC0.748 0.700β0.796 / 0.370 0.296β0.444 0.699 0.649β0.745 / 0.220 0.170β0.280 0.881 0.849β0.910 / 0.603 0.528β0.674 0.840 0.802β0.874 / 0.468 0.390β0.552 0.815 0.778β0.852 / 0.460 0.379β0.545 0.910 0.876β0.939 / 0.781 0.717β0.840 RBBB0.955 0.943β0.966 / 0.905 0.877β0.929 0.927 0.911β0.942 / 0.833 0.791β0.869 0.981 0.974β0.987 / 0.949 0.931β0.966 0.889 0.869β0.907 / 0.765 0.724β0.804 0.983 0.977β0.988 / 0.951 0.930β0.969 0.976 0.967β0.983 / 0.937 0.906β0.961 STD0.875 0.844β0.899 / 0.571 0.500β0.640 0.820 0.788β0.848 / 0.435 0.369β0.506 0.958 0.941β0.972 / 0.818 0.765β0.866 0.826 0.792β0.859 / 0.474 0.407β0.538 0.957 0.940β0.971 / 0.835 0.782β0.881 0.949 0.931β0.966 / 0.822 0.766β0.870 STE0.880 0.811β0.937 / 0.414 0.273β0.559 0.814 0.734β0.887 / 0.342 0.202β0.483 0.955 0.916β0.982 / 0.607 0.473β0.732 0.836 0.773β0.892 / 0.263 0.159β0.404 0.906 0.831β0.963 / 0.591 0.458β0.719 0.923 0.873β0.965 / 0.523 0.388β0.662 33 Position: Evaluation of ECG Representations Must Be Fixed F. Complete Evaluation on CSN Table 27. Performance on CSN sub-tasks whose labels begin with letter A through M. Each cell reports AUROC (top line) and AUPRC (bottom line), each with a 95% confidence interval. TaskRandomCLOCSKEDHeartLangMERLD-BETA 1AVB0.773 0.735β0.811 / 0.093 0.071β0.122 0.705 0.663β0.746 / 0.057 0.042β0.081 0.975 0.953β0.991 / 0.757 0.688β0.817 0.865 0.832β0.896 / 0.228 0.173β0.291 0.981 0.968β0.990 / 0.719 0.647β0.789 0.990 0.986β0.994 / 0.803 0.742β0.853 2AVB0.470 0.266β0.728 / 0.007 0.001β0.033 0.936 0.880β0.990 / 0.217 0.011β0.553 0.981 0.951β0.997 / 0.103 0.040β0.187 0.880 0.718β0.968 / 0.011 0.004β0.020 0.867 0.643β0.988 / 0.034 0.008β0.101 0.983 0.961β0.995 / 0.066 0.027β0.125 2AVB10.593 0.422β0.795 / 0.067 0.001β0.268 0.796 0.651β0.931 / 0.015 0.002β0.050 0.940 0.897β0.983 / 0.019 0.005β0.055 0.575 0.291β0.834 / 0.057 0.001β0.258 0.792 0.550β0.954 / 0.008 0.002β0.024 0.978 0.952β0.992 / 0.039 0.016β0.067 AF0.867 0.852β0.881 / 0.410 0.376β0.444 0.844 0.829β0.859 / 0.340 0.309β0.371 0.973 0.969β0.976 / 0.712 0.674β0.748 0.911 0.902β0.921 / 0.475 0.440β0.510 0.975 0.971β0.979 / 0.733 0.697β0.772 0.977 0.974β0.981 / 0.757 0.723β0.790 AFIB0.856 0.834β0.877 / 0.243 0.207β0.280 0.856 0.835β0.874 / 0.221 0.186β0.261 0.966 0.961β0.971 / 0.510 0.460β0.566 0.909 0.897β0.922 / 0.282 0.246β0.322 0.962 0.957β0.966 / 0.473 0.423β0.526 0.969 0.964β0.973 / 0.527 0.478β0.581 ALS0.974 0.964β0.981 / 0.541 0.472β0.612 0.856 0.828β0.882 / 0.180 0.143β0.224 0.981 0.975β0.986 / 0.597 0.523β0.668 0.893 0.869β0.914 / 0.248 0.199β0.302 0.938 0.922β0.952 / 0.398 0.331β0.469 0.984 0.979β0.988 / 0.673 0.607β0.739 APB0.708 0.663β0.755 / 0.050 0.039β0.066 0.682 0.639β0.722 / 0.053 0.036β0.080 0.941 0.918β0.960 / 0.464 0.377β0.552 0.735 0.690β0.778 / 0.053 0.042β0.066 0.799 0.760β0.837 / 0.097 0.071β0.134 0.987 0.973β0.994 / 0.724 0.644β0.804 AQW0.902 0.871β0.931 / 0.243 0.187β0.312 0.749 0.698β0.796 / 0.168 0.106β0.237 0.938 0.918β0.957 / 0.433 0.350β0.517 0.802 0.758β0.841 / 0.105 0.076β0.143 0.922 0.890β0.953 / 0.505 0.414β0.592 0.968 0.951β0.982 / 0.571 0.475β0.662 ARS0.954 0.932β0.972 / 0.401 0.327β0.478 0.896 0.871β0.918 / 0.156 0.118β0.206 0.971 0.964β0.978 / 0.414 0.336β0.497 0.955 0.941β0.969 / 0.366 0.289β0.449 0.939 0.915β0.959 / 0.316 0.246β0.396 0.910 0.884β0.935 / 0.326 0.246β0.409 AT0.742 0.657β0.822 / 0.041 0.016β0.091 0.770 0.698β0.837 / 0.032 0.017β0.055 0.933 0.899β0.960 / 0.194 0.103β0.316 0.804 0.735β0.863 / 0.040 0.020β0.079 0.880 0.829β0.921 / 0.061 0.035β0.104 0.978 0.967β0.987 / 0.381 0.241β0.537 AVB0.628 0.522β0.735 / 0.025 0.013β0.041 0.857 0.805β0.902 / 0.036 0.025β0.052 0.926 0.905β0.945 / 0.083 0.048β0.135 0.851 0.802β0.897 / 0.062 0.031β0.114 0.939 0.905β0.968 / 0.189 0.112β0.298 0.972 0.960β0.981 / 0.204 0.129β0.302 AVRT0.494 0.004β0.987 / 0.008 0.000β0.024 0.963 0.926β0.996 / 0.025 0.004β0.069 0.928 0.887β0.967 / 0.005 0.003β0.009 0.554 0.302β0.803 / 0.001 0.000β0.002 0.954 0.929β0.978 / 0.007 0.004β0.014 0.976 0.964β0.987 / 0.013 0.009β0.024 CCR0.899 0.851β0.939 / 0.060 0.028β0.119 0.919 0.882β0.949 / 0.103 0.027β0.210 0.929 0.899β0.956 / 0.138 0.049β0.272 0.783 0.697β0.862 / 0.034 0.013β0.077 0.872 0.810β0.925 / 0.085 0.025β0.183 0.831 0.767β0.891 / 0.044 0.015β0.112 CR0.938 0.882β0.983 / 0.118 0.034β0.262 0.910 0.765β0.992 / 0.138 0.044β0.310 0.945 0.869β0.995 / 0.250 0.076β0.500 0.808 0.578β0.960 / 0.039 0.009β0.132 0.881 0.768β0.970 / 0.244 0.030β0.503 0.933 0.842β0.989 / 0.155 0.031β0.347 ERV0.940 0.912β0.963 / 0.245 0.167β0.338 0.917 0.886β0.944 / 0.209 0.132β0.304 0.973 0.965β0.981 / 0.279 0.207β0.379 0.892 0.852β0.926 / 0.138 0.084β0.217 0.973 0.964β0.982 / 0.324 0.229β0.436 0.952 0.936β0.967 / 0.230 0.150β0.323 FQRS0.179 0.169β0.188 / 0.000 0.000β0.000 0.186 0.177β0.196 / 0.000 0.000β0.000 0.883 0.875β0.891 / 0.001 0.001β0.001 0.025 0.021β0.029 / 0.000 0.000β0.000 0.231 0.220β0.241 / 0.000 0.000β0.000 0.985 0.982β0.988 / 0.011 0.009β0.013 IVB0.773 0.710β0.830 / 0.070 0.043β0.111 0.807 0.761β0.853 / 0.047 0.036β0.061 0.941 0.923β0.956 / 0.177 0.130β0.238 0.865 0.827β0.898 / 0.105 0.062β0.159 0.915 0.885β0.939 / 0.133 0.095β0.183 0.976 0.962β0.986 / 0.468 0.360β0.587 JEB0.700 0.415β0.905 / 0.002 0.001β0.005 0.552 0.148β0.895 / 0.001 0.000β0.004 0.989 0.972β1.000 / 0.280 0.017β0.750 0.736 0.633β0.828 / 0.001 0.001β0.002 0.880 0.659β0.999 / 0.098 0.001β0.350 0.983 0.964β0.996 / 0.044 0.013β0.107 JPT0.156 0.093β0.223 / 0.000 0.000β0.000 0.303 0.078β0.524 / 0.000 0.000β0.001 0.562 0.378β0.753 / 0.001 0.001β0.001 0.312 0.193β0.436 / 0.000 0.000β0.001 0.058 0.040β0.077 / 0.000 0.000β0.000 0.895 0.886β0.903 / 0.003 0.002β0.003 LFBBB0.991 0.980β0.998 / 0.652 0.518β0.800 0.910 0.844β0.966 / 0.362 0.236β0.506 0.990 0.979β0.997 / 0.683 0.554β0.797 0.953 0.916β0.983 / 0.487 0.335β0.649 0.980 0.960β0.993 / 0.628 0.497β0.757 0.993 0.988β0.998 / 0.705 0.571β0.825 LVH0.888 0.757β0.983 / 0.157 0.072β0.272 0.987 0.979β0.993 / 0.244 0.127β0.408 0.995 0.993β0.998 / 0.402 0.269β0.583 0.859 0.765β0.930 / 0.059 0.021β0.126 0.979 0.964β0.992 / 0.216 0.125β0.344 0.968 0.949β0.984 / 0.123 0.064β0.214 LVQRSAL0.908 0.880β0.935 / 0.320 0.253β0.396 0.856 0.822β0.886 / 0.210 0.154β0.271 0.936 0.916β0.954 / 0.397 0.316β0.479 0.765 0.723β0.805 / 0.104 0.075β0.141 0.922 0.900β0.942 / 0.331 0.259β0.408 0.862 0.831β0.888 / 0.196 0.141β0.260 MISW0.854 0.665β0.969 / 0.046 0.010β0.131 0.782 0.590β0.943 / 0.055 0.009β0.162 0.944 0.906β0.974 / 0.076 0.014β0.222 0.792 0.660β0.908 / 0.037 0.004β0.157 0.916 0.844β0.974 / 0.049 0.014β0.121 0.965 0.933β0.990 / 0.090 0.028β0.209 34 Position: Evaluation of ECG Representations Must Be Fixed Table 28. Performance on CSN sub-tasks whose labels begin with letter P through Z. Each cell reports AUROC (top line) and AUPRC (bottom line), each with a 95% confidence interval. TaskRandomCLOCSKEDHeartLangMERLD-BETA PRIE0.717 0.229β0.983 / 0.039 0.001β0.148 0.681 0.141β0.977 / 0.032 0.001β0.130 0.843 0.566β0.994 / 0.074 0.001β0.277 0.783 0.640β0.940 / 0.003 0.001β0.008 0.935 0.841β0.991 / 0.232 0.003β0.672 0.964 0.931β0.986 / 0.014 0.007β0.028 PWC0.772 0.622β0.896 / 0.020 0.007β0.041 0.773 0.662β0.875 / 0.015 0.005β0.037 0.902 0.803β0.967 / 0.062 0.020β0.137 0.854 0.749β0.930 / 0.019 0.009β0.038 0.927 0.869β0.974 / 0.099 0.032β0.228 0.879 0.751β0.959 / 0.033 0.015β0.062 QTIE0.656 0.576β0.740 / 0.043 0.018β0.094 0.764 0.695β0.828 / 0.036 0.022β0.056 0.931 0.899β0.960 / 0.248 0.144β0.369 0.761 0.700β0.821 / 0.037 0.021β0.061 0.940 0.914β0.964 / 0.225 0.133β0.325 0.936 0.900β0.965 / 0.243 0.155β0.346 RAH0.905 0.898β0.912 / 0.002 0.002β0.002 0.989 0.986β0.991 / 0.014 0.011β0.018 0.994 0.991β0.995 / 0.024 0.018β0.033 1.000 0.999β1.000 / 0.317 0.125β1.000 0.978 0.974β0.981 / 0.007 0.006β0.008 0.992 0.989β0.994 / 0.018 0.014β0.024 RBBB0.954 0.921β0.983 / 0.792 0.710β0.869 0.957 0.928β0.980 / 0.556 0.463β0.648 0.995 0.988β0.999 / 0.945 0.908β0.972 0.921 0.895β0.944 / 0.407 0.316β0.496 0.991 0.976β0.999 / 0.931 0.881β0.971 0.998 0.997β0.999 / 0.957 0.929β0.981 RVH0.826 0.482β0.998 / 0.172 0.030β0.467 0.824 0.519β0.999 / 0.165 0.012β0.468 0.990 0.973β0.999 / 0.260 0.066β0.632 0.920 0.766β0.999 / 0.334 0.027β0.706 0.993 0.984β0.998 / 0.174 0.044β0.434 0.969 0.916β0.999 / 0.281 0.028β0.655 SA0.744 0.718β0.769 / 0.238 0.203β0.277 0.752 0.730β0.773 / 0.181 0.157β0.207 0.927 0.914β0.939 / 0.641 0.593β0.688 0.779 0.756β0.800 / 0.274 0.235β0.315 0.830 0.811β0.849 / 0.400 0.351β0.448 0.982 0.977β0.986 / 0.838 0.809β0.864 SB0.937 0.932β0.943 / 0.887 0.873β0.900 0.976 0.973β0.980 / 0.953 0.945β0.961 0.999 0.999β1.000 / 0.999 0.999β1.000 0.987 0.985β0.990 / 0.977 0.972β0.982 0.994 0.993β0.996 / 0.989 0.985β0.993 0.999 0.999β1.000 / 0.999 0.998β0.999 SR0.754 0.740β0.767 / 0.426 0.406β0.446 0.894 0.885β0.903 / 0.736 0.715β0.756 0.990 0.988β0.993 / 0.970 0.963β0.976 0.933 0.925β0.941 / 0.825 0.805β0.845 0.968 0.963β0.972 / 0.904 0.890β0.916 0.993 0.991β0.995 / 0.981 0.977β0.985 ST0.935 0.926β0.944 / 0.748 0.720β0.777 0.969 0.964β0.973 / 0.838 0.813β0.862 0.996 0.994β0.998 / 0.986 0.981β0.990 0.985 0.981β0.987 / 0.927 0.914β0.940 0.993 0.990β0.995 / 0.973 0.966β0.979 0.996 0.993β0.998 / 0.988 0.984β0.993 STDD0.871 0.823β0.917 / 0.142 0.099β0.197 0.872 0.829β0.912 / 0.131 0.091β0.181 0.970 0.956β0.981 / 0.400 0.306β0.495 0.892 0.850β0.926 / 0.170 0.117β0.238 0.966 0.952β0.978 / 0.415 0.324β0.512 0.974 0.965β0.982 / 0.409 0.320β0.515 STE0.868 0.828β0.900 / 0.183 0.127β0.256 0.745 0.688β0.801 / 0.130 0.081β0.191 0.931 0.905β0.954 / 0.309 0.230β0.397 0.797 0.755β0.836 / 0.080 0.052β0.118 0.919 0.887β0.945 / 0.323 0.241β0.414 0.911 0.882β0.939 / 0.247 0.178β0.332 STTC0.842 0.810β0.874 / 0.205 0.162β0.254 0.839 0.808β0.868 / 0.181 0.142β0.230 0.939 0.921β0.953 / 0.406 0.341β0.477 0.861 0.835β0.885 / 0.201 0.157β0.249 0.937 0.920β0.953 / 0.359 0.300β0.426 0.942 0.924β0.956 / 0.378 0.314β0.448 STTU0.716 0.587β0.829 / 0.049 0.011β0.129 0.779 0.663β0.882 / 0.081 0.024β0.183 0.864 0.788β0.926 / 0.110 0.037β0.229 0.772 0.670β0.861 / 0.094 0.013β0.210 0.912 0.837β0.959 / 0.176 0.066β0.326 0.928 0.884β0.963 / 0.134 0.057β0.258 SVT0.967 0.946β0.983 / 0.579 0.490β0.667 0.990 0.987β0.993 / 0.697 0.616β0.774 0.992 0.982β0.998 / 0.845 0.779β0.903 0.984 0.973β0.992 / 0.749 0.675β0.819 0.996 0.993β0.998 / 0.863 0.805β0.916 0.991 0.977β0.999 / 0.890 0.835β0.936 TWC0.868 0.853β0.884 / 0.578 0.545β0.611 0.849 0.835β0.863 / 0.524 0.492β0.557 0.944 0.937β0.950 / 0.761 0.735β0.787 0.848 0.835β0.862 / 0.529 0.499β0.558 0.940 0.932β0.948 / 0.764 0.738β0.791 0.916 0.906β0.926 / 0.697 0.670β0.724 TWO0.894 0.875β0.910 / 0.267 0.227β0.311 0.854 0.825β0.880 / 0.282 0.231β0.338 0.952 0.943β0.961 / 0.464 0.405β0.523 0.836 0.808β0.862 / 0.244 0.194β0.295 0.937 0.922β0.950 / 0.454 0.395β0.513 0.958 0.950β0.966 / 0.485 0.426β0.548 UW0.730 0.570β0.867 / 0.009 0.003β0.020 0.874 0.820β0.920 / 0.010 0.006β0.015 0.929 0.843β0.984 / 0.069 0.025β0.153 0.849 0.766β0.915 / 0.011 0.005β0.021 0.909 0.858β0.955 / 0.089 0.010β0.273 0.906 0.842β0.947 / 0.017 0.009β0.030 VB0.392 0.380β0.404 / 0.000 0.000β0.000 0.923 0.916β0.929 / 0.002 0.002β0.002 0.244 0.233β0.254 / 0.000 0.000β0.000 0.322 0.311β0.334 / 0.000 0.000β0.000 0.128 0.119β0.136 / 0.000 0.000β0.000 0.992 0.989β0.994 / 0.018 0.014β0.024 VEB0.657 0.330β0.960 / 0.013 0.001β0.041 0.828 0.650β0.986 / 0.017 0.003β0.038 0.887 0.675β0.995 / 0.088 0.011β0.257 0.813 0.678β0.945 / 0.010 0.002β0.030 0.877 0.692β0.995 / 0.068 0.008β0.202 0.942 0.875β0.999 / 0.327 0.022β0.801 VET0.207 0.197β0.217 / 0.000 0.000β0.000 0.834 0.825β0.843 / 0.001 0.001β0.001 0.911 0.904β0.918 / 0.002 0.002β0.002 0.618 0.606β0.630 / 0.000 0.000β0.000 0.351 0.339β0.363 / 0.000 0.000β0.000 0.999 0.999β1.000 / 0.246 0.100β0.500 VFW0.804 0.794β0.813 / 0.001 0.001β0.001 0.249 0.238β0.260 / 0.000 0.000β0.000 0.189 0.180β0.198 / 0.000 0.000β0.000 0.312 0.301β0.323 / 0.000 0.000β0.000 0.646 0.634β0.657 / 0.000 0.000β0.000 0.423 0.411β0.434 / 0.000 0.000β0.000 VPB0.691 0.600β0.774 / 0.084 0.028β0.161 0.808 0.727β0.878 / 0.066 0.035β0.120 0.955 0.933β0.975 / 0.342 0.218β0.474 0.909 0.866β0.944 / 0.174 0.091β0.275 0.905 0.863β0.941 / 0.141 0.080β0.218 0.987 0.969β0.998 / 0.700 0.567β0.825 VPE0.964 0.924β0.998 / 0.044 0.004β0.133 0.872 0.727β1.000 / 0.520 0.001β1.000 0.884 0.754β1.000 / 0.354 0.001β1.000 0.816 0.610β1.000 / 0.520 0.001β1.000 0.969 0.944β0.992 / 0.015 0.006β0.038 0.905 0.797β1.000 / 0.206 0.002β0.667 WPW0.770 0.545β0.988 / 0.126 0.040β0.268 0.893 0.821β0.958 / 0.160 0.010β0.400 0.969 0.927β0.995 / 0.322 0.093β0.581 0.905 0.796β0.984 / 0.157 0.029β0.367 0.898 0.719β0.997 / 0.476 0.216β0.746 0.948 0.894β0.987 / 0.363 0.125β0.627 35 Position: Evaluation of ECG Representations Must Be Fixed G. Complete Evaluation on EchoNext Table 29. Performance on EchoNext sub-tasks. Each cell reports AUROC (top line) and AUPRC (bottom line), each with a 95% confidence interval. TaskRandomCLOCSKEDHeartLangMERLD-BETA AR0.698 0.629β0.764 / 0.037 0.025β0.054 0.620 0.553β0.693 / 0.040 0.018β0.080 0.626 0.553β0.692 / 0.034 0.019β0.062 0.627 0.567β0.687 / 0.030 0.016β0.059 0.696 0.626β0.764 / 0.061 0.028β0.110 0.681 0.629β0.731 / 0.022 0.017β0.027 AS0.753 0.723β0.784 / 0.187 0.152β0.226 0.695 0.661β0.729 / 0.140 0.112β0.169 0.763 0.735β0.792 / 0.171 0.143β0.205 0.710 0.677β0.744 / 0.133 0.110β0.161 0.802 0.776β0.827 / 0.228 0.189β0.269 0.776 0.751β0.800 / 0.159 0.135β0.190 LVEFβ€ 450.856 0.842β0.870 / 0.619 0.589β0.648 0.777 0.760β0.793 / 0.483 0.453β0.512 0.858 0.846β0.873 / 0.622 0.593β0.653 0.826 0.810β0.840 / 0.548 0.521β0.580 0.878 0.865β0.892 / 0.667 0.638β0.697 0.862 0.849β0.874 / 0.613 0.584β0.641 LVWTβ₯ 130.742 0.726β0.758 / 0.416 0.391β0.442 0.670 0.651β0.688 / 0.342 0.319β0.367 0.744 0.728β0.760 / 0.417 0.392β0.443 0.715 0.699β0.731 / 0.361 0.340β0.384 0.744 0.728β0.760 / 0.413 0.388β0.441 0.718 0.702β0.734 / 0.354 0.332β0.378 MR0.789 0.762β0.814 / 0.246 0.210β0.287 0.723 0.696β0.752 / 0.169 0.143β0.198 0.796 0.772β0.819 / 0.234 0.200β0.272 0.744 0.717β0.773 / 0.161 0.140β0.188 0.804 0.780β0.827 / 0.236 0.204β0.274 0.796 0.771β0.820 / 0.222 0.190β0.256 PASPβ₯ 450.741 0.721β0.760 / 0.312 0.283β0.341 0.696 0.676β0.715 / 0.241 0.221β0.263 0.737 0.715β0.755 / 0.300 0.272β0.328 0.712 0.693β0.730 / 0.261 0.240β0.285 0.752 0.733β0.771 / 0.335 0.306β0.367 0.730 0.710β0.750 / 0.293 0.266β0.321 PEFF0.716 0.656β0.773 / 0.029 0.022β0.039 0.629 0.559β0.693 / 0.022 0.016β0.030 0.731 0.665β0.794 / 0.039 0.026β0.059 0.660 0.592β0.722 / 0.031 0.019β0.052 0.718 0.648β0.779 / 0.040 0.026β0.062 0.709 0.647β0.770 / 0.036 0.025β0.052 PR0.816 0.721β0.896 / 0.073 0.011β0.185 0.836 0.779β0.885 / 0.014 0.009β0.022 0.788 0.681β0.884 / 0.025 0.011β0.051 0.772 0.664β0.878 / 0.077 0.010β0.204 0.868 0.801β0.926 / 0.080 0.024β0.175 0.859 0.781β0.919 / 0.029 0.013β0.057 RVSD0.841 0.821β0.862 / 0.368 0.327β0.413 0.793 0.771β0.814 / 0.274 0.239β0.312 0.843 0.822β0.865 / 0.399 0.356β0.443 0.830 0.810β0.850 / 0.299 0.268β0.335 0.856 0.836β0.877 / 0.427 0.382β0.475 0.850 0.831β0.868 / 0.369 0.324β0.417 SHD0.805 0.793β0.816 / 0.772 0.756β0.786 0.723 0.709β0.737 / 0.686 0.669β0.702 0.802 0.790β0.814 / 0.762 0.746β0.778 0.772 0.759β0.785 / 0.722 0.704β0.740 0.812 0.800β0.824 / 0.781 0.764β0.796 0.801 0.789β0.812 / 0.762 0.746β0.778 TR-MAXβ₯ 320.732 0.706β0.757 / 0.200 0.170β0.233 0.687 0.662β0.712 / 0.127 0.112β0.145 0.725 0.698β0.751 / 0.179 0.152β0.209 0.714 0.688β0.742 / 0.156 0.136β0.182 0.744 0.721β0.768 / 0.198 0.169β0.236 0.710 0.684β0.735 / 0.157 0.136β0.180 TR0.784 0.759β0.810 / 0.229 0.197β0.267 0.715 0.686β0.744 / 0.143 0.124β0.165 0.778 0.749β0.804 / 0.225 0.191β0.261 0.753 0.727β0.780 / 0.181 0.155β0.211 0.815 0.792β0.838 / 0.265 0.227β0.304 0.791 0.766β0.815 / 0.235 0.199β0.271 36