Paper deep dive
Detecting low left ventricular ejection fraction from ECG using an interpretable and scalable predictor-driven framework
Ya Zhou, Tianxiang Hao, Ziyi Cai, Haojie Zhu, Hejun He, Jia Liu, Xiaohan Fan, Jing Yuan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/31/2026, 2:45:57 AM
Summary
The paper introduces ECGPD-LEF, an interpretable and scalable framework for detecting low left ventricular ejection fraction (LEF) from 12-lead ECGs. The framework uses a Transformer-based predictor extractor to generate 71 diagnostic probabilities, which are then processed by either a single-predictor (zero-shot) or multi-predictor (tabular) model. Evaluated on the EchoNext benchmark and an external MIMIC-LEF dataset, the framework demonstrates robust performance and mechanistic transparency, outperforming end-to-end black-box baselines.
Entities (5)
Relation Signals (3)
ECGPD-LEF β detects β Low left ventricular ejection fraction
confidence 100% Β· We introduced ECG-based Predictor-Driven LEF (ECGPD-LEF), a structured framework... for detecting LEF from ECG.
ECGPD-LEF β trainedon β EchoNext
confidence 100% Β· Trained on the benchmark EchoNext dataset comprising 72,475 ECG-echocardiogram pairs
ECGPD-LEF β validatedon β MIMIC-LEF
confidence 100% Β· External validation was performed using an ECG-Note paired dataset constructed from MIMIC-IV... (MIMIC-LEF)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Low left ventricular ejection fraction (LEF) frequently remains undetected until progression to symptomatic heart failure, underscoring the need for scalable screening strategies. Although artificial intelligence-enabled electrocardiography (AI-ECG) has shown promise, existing approaches rely solely on end-to-end black-box models with limited interpretability or on tabular systems dependent on commercial ECG measurement algorithms with suboptimal performance. We introduced ECG-based Predictor-Driven LEF (ECGPD-LEF), a structured framework that integrates foundation model-derived diagnostic probabilities with interpretable modeling for detecting LEF from ECG. Trained on the benchmark EchoNext dataset comprising 72,475 ECG-echocardiogram pairs and evaluated in predefined independent internal (n=5,442) and external (n=16,017) cohorts, our framework achieved robust discrimination for moderate LEF (internal AUROC 88.4%, F1 64.5%; external AUROC 86.8%, F1 53.6%), consistently outperforming the official end-to-end baseline provided with the benchmark across demographic and clinical subgroups. Interpretability analyses identified high-impact predictors, including normal ECG, incomplete left bundle branch block, and subendocardial injury in anterolateral leads, driving LEF risk estimation. Notably, these predictors independently enabled zero-shot-like inference without task-specific retraining (internal AUROC 75.3-81.0%; external AUROC 71.6-78.6%), indicating that ventricular dysfunction is intrinsically encoded within structured diagnostic probability representations. This framework reconciles predictive performance with mechanistic transparency, supporting scalable enhancement through additional predictors and seamless integration with existing AI-ECG systems.
Tags
Links
- Source: https://arxiv.org/abs/2603.28532v1
- Canonical: https://arxiv.org/abs/2603.28532v1
Trouble viewing inline? Open PDF directly β
Full Text
108,020 characters extracted from source content.
Expand or collapse full text
Detecting low left ventricular ejection fraction from ECG using an interpretable and scalable predictor-driven framework Ya Zhou 1,* , Tianxiang Hao 2 , Ziyi Cai 3 , Haojie Zhu 4 , Hejun He 3 , Jia Liu 1 , Xiaohan Fan 4,5, * and Jing Yuan 1, * 1 Department of Information Center, Fuwai Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China 2 The Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China 3 Institute of Statistics and Big Data, Renmin University of China, Beijing, China 4 Cardiac Arrhythmia Center, Fuwai Hospital, National Center for Cardiovascular Diseases, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China 5 Function Test Center, Fuwai Hospital, National Center for Cardiovascular Diseases, Chinese Academy of Medical Sciences and Peking Union Medical College, Beijing, China Abstract Low left ventricular ejection fraction (LEF) frequently remains undetected until pro- gression to symptomatic heart failure, underscoring the need for scalable screening strate- gies. Although artificial intelligence-enabled electrocardiography (AI-ECG) has shown promise, existing approaches rely solely on end-to-end black-box models with limited interpretability or on tabular systems dependent on commercial ECG measurement algo- rithms with suboptimal performance. We introduced ECG-based Predictor-Driven LEF (ECGPD-LEF), a structured framework that integrates foundation model-derived diag- nostic probabilities with interpretable modeling for detecting LEF from ECG. Trained on the benchmark EchoNext dataset comprising 72,475 ECG-echocardiogram pairs and eval- uated in predefined independent internal (n=5,442) and external (n=16,017) cohorts, our framework achieved robust discrimination for moderate LEF (internal AUROC 88.4%, F1 64.5%; external AUROC 86.8%, F1 53.6%), consistently outperforming the official end- to-end baseline provided with the benchmark across demographic and clinical subgroups. Interpretability analyses identified high-impact predictors, including normal ECG, in- complete left bundle branch block, and subendocardial injury in anterolateral leads, driv- ing LEF risk estimation. Notably, these predictors independently enabled zero-shot-like inference without task-specific retraining (internal AUROC 75.3-81.0%; external AU- ROC 71.6-78.6%), indicating that ventricular dysfunction is intrinsically encoded within structured diagnostic probability representations. This framework reconciles predictive performance with mechanistic transparency, supporting scalable enhancement through additional predictors and seamless integration with existing AI-ECG systems. * Email: Jing Yuan (yuanjing@fuwai.com), Xiaohan Fan (fanxiaohan@fuwaihospital.org), Ya Zhou (zhouya@fuwai.com) 1 arXiv:2603.28532v1 [cs.LG] 30 Mar 2026 1 Introduction Heart failure (HF) is a major global health burden, affecting over 64 million individuals worldwide, with prevalence continuing to rise (Savarese et al.,2022). In many patients, left ventricular systolic dysfunction (LVSD) is initially asymptomatic, yet it frequently progresses to symptomatic HF and is associated with poor prognosis if left undetected. Low left ven- tricular ejection fraction (LEF) is a critical indicator that reflects the progression of LVSD toward clinically manifest HF (Heidenreich et al.,2022). Early detection of LVSD is critical, as timely initiation of evidence-based therapies can mitigate symptoms, slow disease progres- sion, and improve survival. Although cardiovascular imaging, particularly echocardiography (ECHO), remains the clinical gold standard for assessing ventricular function, it is time- and resource-intensive, requires specialized operators and equipment, and is not readily scalable for population-level screening (Ciampi and Villari,2007). These limitations underscore the need for cost-effective, widely accessible approaches capable of reliably identifying patients with LEF ( Attia et al.,2019a;Lakshmanan and Mbanze,2023;Tran et al.,2025;Yao et al., 2021). Electrocardiography (ECG), which records cardiac electrical activity, is a promising tool for large-scale screening (Hughes et al.,2024). Since its introduction by Willem Einthoven, which earned him the Nobel Prize in 1924, ECG science has steadily advanced: acquisition has become easier, signal quality has improved, and pathognomonic features for a wide range of rhythm and conduction disorders have been defined ( Khera,2024). More recently, artificial intelligence (AI) has transformed ECG analysis, enabling models to achieve cardiologist-level performance in traditional interpretation tasks (Hannun et al.,2019;Jiang et al.,2024;Li et al. ,2025;Ribeiro et al.,2020) and to detect subtle patterns associated with conditions such as aortic stenosis (Kwon et al.,2020), tricuspid regurgitation (Diao et al.,2025), and LEF (Attia et al.,2019a), which are often imperceptible to humans. Despite these advances, clinical adoption varies: AI-ECG models are widely accepted for well-established conditions such as atrial fibrillation and sinus bradycardia, which can be visually confirmed according to guidelines ( Joglar et al.,2024;Kusumoto et al.,2019), but remain limited for emerging tasks such as LEF detection due to concerns about interpretability and the difficulty of linking predictions to recognizable ECG features ( Hughes et al.,2024;Van De Leur et al.,2022). To enhance interpretability and facilitate broader clinical deployment, a recent study proposed a random forest model based on 555 discrete ECG measurements based on commercial al- gorithms, which demonstrated superior performance compared with N-terminal prohormone brain natriuretic peptide (NT-proBNP) for LEF detection ( Hughes et al.,2024). Although this approach showed promising results, its performance is slightly below deep learning. More- over, as highlighted in the original study, many features cannot be obtained when applied across different ECG machines, and measurement values differ markedly between devices (Strodthoff et al.,2023), which presents a significant barrier to broader clinical adoption. In this study, we aim to evaluate the efficacy of interpretable methods for detecting LEF using standard 12-lead ECGs. We propose ECG-based Predictor-Driven LEF (ECGPD- LEF), a framework that integrates traditional AI-ECG interpretation models with either a single-predictor approach or a multi-predictor approach to enhance recognition of subtle ECG features and improve LEF detection. The single-predictor approach performs zero-shot-like inference, requiring no additional training on the target dataset. The multi-predictor ap- proach was trained on the publicly available ECG-ECHO paired dataset EchoNext, a com- 2 prehensive dataset encompassing diverse populations across races and clinical contexts. We also constructed an external validation dataset, MIMIC-LEF, based on MIMIC-IV (John- son et al.,2024), MIMIC-IV-ECG (Gow et al.,2023), and MIMIC-IV-Note(Johnson et al., 2023a), which similarly covers a heterogeneous population. The model was evaluated on multi-center datasets, including internal validation (the EchoNext test set) and external val- idation (MIMIC-LEF). The model performance was assessed across multiple dimensions, including the presence or absence of other structural heart disease (SHD) and valvular heart disease (VHD), as well as across different care settings and racial or ethnic subgroups. Global and local model explanations were performed to identify important predictors contributing to LEF detection, and to characterize these subtle ECG changes associated with LEF. 2 Methods 2.1 Data sources An overview of the study design is provided in Figure1, illustrating the data sources, construc- tion of the predictor extractor, model development for low left ventricular ejection fraction (LEF) detection, and downstream evaluation. Four ECG-only datasets were implicitly used to develop the predictor extractor, and two independent paired datasets (ECG-ECHO and ECG-Note) were explicitly used for development and testing. The ECG-only datasets included Ningbo ( Zheng et al.,2020a), Chapman (Zheng et al., 2020b), CODE-15% (Ribeiro et al.,2021), and PTB-XL (Wagner et al.,2020). Informa- tion from Ningbo, Chapman, and CODE-15% was implicitly incorporated through the pre- trained weights of a Transformer-based model (ST-MEM) (Na et al.,2024). PTB-XL was subsequently used for post-training to obtain the final predictor extractor. For the detection of LEF, we used the ECG-ECHO paired EchoNext dataset collected at Columbia University Irving Medical Center ( Elias and Finer,2025). External validation was performed using an ECG-Note paired dataset constructed from MIMIC-IV (Johnson et al., 2024), MIMIC-IV-ECG (Gow et al.,2023), and MIMIC-IV-Note (Johnson et al.,2023a), comprising records from Beth Israel Deaconess Medical Center. 2.2 Implicit ECG-only datasets Chapman (Zheng et al.,2020b) and Ningbo, created under the auspices of Chapman Uni- versity, Shaoxing Peoples Hospital, and Ningbo First Hospital, contains 10,646 and 34,905 12-lead ECG, respectively, stored in MUSE ECG system with a sampling rate of 500Hz and a duration of 10 seconds. Code-15% (Ribeiro et al.,2021) is obtained through strati- fied sampling from the CODE dataset (Ribeiro et al.,2020), containing 345,779 exams from 233,770 patients, collected by the Telehealth Network of Minas Gerais in the period between 2010 and 2016. The S12L-ECG exam was performed mostly in primary care facilities us- ing a tele-electrocardiograph manufactured by Tecnologia EletrΓ΄nica Brasileira (SΓ£o Paulo, Brazil)model TEB ECGPCor Micromed Biotecnologia (Brasilia, Brazil)model ErgoPC 13 and the duration of the ECG recordings is between 7 and 10 s sampled at frequencies ranging from 300 to 600 Hz. The processed subset of these datasets are implicitly in the pre-trained weights of the traditional automatic ECG diagnosis model (the traditional AI-ECG model). 3 Figure 1: Flowchart of model development and evaluation.Model development is based on a traditional AI-ECG interpretation model, where a large unlabeled ECG dataset is used for pre- training and a smaller labeled ECG dataset is used for enhanced post-training. The automatic ECG diagnosis model outputs diagnostic probabilities for each ECG finding and serves as the predictor extractor. Based on the predictor extractor, we developed both a single-predictor approach and a multi-predictor approach. The single-predictor approach enables zero-shot-like inference without further training to detect LEF and can serve as an important indicator. The multi-predictor approach further trains a tabular model using ECG-ECHO pairs and provides both local- and global-level explanations based on SHAP values. Model evaluation is conducted on the EchoNext test set and an independently collected ECHO-Note pairs dataset. Multidimensional model dissection is performed, including overall performance evaluation, model interpretability analysis, and subgroup analyses across diverse populations. LEF, low left ventricular ejection fraction; No-LEF, absence of low left ventricular ejection fraction; SHD, structural heart disease; VHD, valvular heart disease, SHAP, Shapley Additive exPlanations. 4 PTB-XL (Wagner et al.,2020) dataset was recorded by devices from Schiller AG between October 1989 and June 1996, consisting 21,837 clinical 12-lead ECG records of 10 seconds length from 18885 patients. The ECG records were annotated by up to two cardiologists with potentially multiple ECG statements out of a set of 71 different statements conforming to the SCP-ECG standards (ISO Central Secretary.,2009). This dataset is a commonly used benchmark for traditional ECG interpretation algorithms (Strodthoff et al.,2020). The dataset id divided into training, validation, and test sets with a ratio of 8:1:1. In this paper, the dataset is implicitly in the post-trained weights of traditional automatic ECG diagnosis models (the traditional AI-ECG model). 2.3 ECG-ECHO and ECG-Note datasets EchoNext is a recent published ECG-ECHO paired dataset for benchmarking ECG screening models (Elias and Finer,2025). It involves 82,543 de-identified paired ECG-ECHO records from 36,286 unique patients and those aged 18 years or older were identified, who underwent a digitally stored 12-lead ECG and a transthoracic ECHO within a 1-year interval between 2008 and 2022. The dataset is divided into training, validation, and test splits. There might includes multiple ECG-ECHO pairs in the training set, while only the latest ECG are adopted in the validation and test sets. The ECG signals were extracted from the GE MUSE ECG management system at a sampling frequency of 250 Hz across all 12 leads with 10 seconds duration. The ECG-Note dataset is constructed based on MIMIC-IV ( Johnson et al.,2024), MIMIC- IV-ECG (Gow et al.,2023), and MIMIC-IV-Note (Johnson et al.,2023a), as illustrated in FigureA.1of the Appendix. We initially identified adult patients who possessed at least one standard 10-second 12-lead ECG with sampling frequency of 500Hz and a corresponding clinical note recorded within one year following the ECG. Exclusion criteria were applied to remove: 1) pairs where clinical notes lacked keywords related to ejection fraction(EF) and ECHO/TTE; 2) ECGs containing missing data (NaN). Following these exclusions, the samples were cross-referenced with MIMIC-ECG-LVEF ( Li et al.,2025) to identify consistent ECG-EF class pairs. For each patient, only the ECG-Note pair with the shortest time interval was retained, resulting in a final testing set of 16,017 ECGs. 2.4 Outcomes For the ECG-ECHO paired dataset, values of left ventricular ejection fraction were extracted from Syngo Dynamics (Siemens) and Xcelera (Philips). The labeling strategy was defined in accordance with Poterucha et al.(2025), using an upper threshold of 45% to define low left ventricular ejection fraction (LEF). An ECG was labeled positive if it was performed within 1 year prior to an echocardiogram demonstrating an ejection fractionβ€45%. For patients confirmed to have ejection fraction>45%, all ECGs prior to the most recent ECHO were labeled negative. For the ECG-Note paired dataset, a large language model( Qwen et al.,2025) using a modified strategy adapted fromGao et al.(2025) was employed to extract ejection fraction values from the discharge table of MIMIC-IV-Note (Figure A.1b in the Appendix). The positive class was defined as ejection fractionβ€45%, consistent with the definition applied to the ECG-ECHO dataset. 5 2.5 Model development We developed an ECG-based Predictor-Driven LEF framework (ECGPD-LEF) for the de- tection of LEF. As illustrated in Figure1b, the framework comprises two components: (1) a predictor extractor that generates structured probabilistic representations from raw ECG waveforms and (2) predictor-based inference models, including single-predictor and multi- predictor approaches. The single-predictor approach performs LEF inference without addi- tional task-specific training, whereas the multi-predictor approach learns a lightweight tabular classifier based on the extracted predictors. 2.5.1 Predictor extractor To instantiate the predictor extractor, we adopted a Transformer-based automatic ECG di- agnosis model. Specifically, we used the ST-MEM architecture (Na et al.,2024), pre-trained on the Chapman, Ningbo, and CODE-15% datasets. Using the publicly available pre-trained weights, we further fine-tuned the model on the PTB-XL dataset to predict 71 conventional ECG diagnoses. Two training strategies were implemented: the original approach (denoted as Transformer) and a modified post-training strategy (denoted as Transformer-PT) described previously (Zhou et al.,2025). The final post-trained model, equipped with sigmoid activa- tion, outputs probability estimates for each of the 71 diagnoses. These probabilistic outputs were used as predictors for the downstream LEF modeling task. 2.5.2 Single-predictor and multi-predictor approaches For each ECG recording, the predictor extractor generates 71 predictor values ranging from 0 to 1, each corresponding to a clinically defined ECG interpretation. In the single-predictor approach, each predictor value (PV) was evaluated independently for LEF detection without additional training on the ECG-ECHO dataset. For predictors corresponding to normal ECG or sinus rhythm, we used(1βPV)to reflect their inverse clinical association with reduced LEF; for all other predictors, the original PV was used directly. Threshold-independent metrics, including AUROC and AUPRC, were computed directly on the test set. For the F1 score, the classification threshold was selected by maximizing validation-set performance. In the multi-predictor approach, the 71 predictors were jointly modeled using lightweight tabular classifiers. As a linear model, we implemented logistic regression with anl 2 penalty, with the regularization parameter selected via grid search over 0.001, 0.01, 0.1, 1.0, 10.0 on the validation set. As a nonlinear alternative, we implemented XGBoost ( Chen,2016), a gradient-boosted decision tree method well suited for structured data. The learning rate and maximum tree depth were tuned via grid search over 0.05, 0.1, 0.2 and 3, 5, 7, respectively. The number of estimators was set to 1000 with early stopping (30 rounds) based on validation performance. For both models, the final classification threshold was determined by maximizing the F1 score on the validation set. 2.6 Performance evaluation We first evaluated the predictor extractor on the PTB-XL test set, as it constitutes a key component of the proposed framework. LEF detection performance was subsequently as- sessed on two independent datasets: the internal test set, consisting of held-out test set from 6 the ECG-ECHO dataset, and the external test set, ECG-Note, which was constructed in this study (Figure1). Both single- and multi-predictor approaches were evaluated on these datasets. Model performance was quantified using the area under the receiver operating character- istic curve (AUROC), area under the precision-recall curve (AUPRC), and F1 score, consis- tent with prior work (Poterucha et al.,2025). Confidence intervals were estimated via 1,000 bootstrap resamples. For benchmarking, we compared our method against the Columbia mini deep learning model (Poterucha et al.,2025), the official end-to-end baseline for LEF detection, which is publicly available with source code and pre-trained weights, enabling re- producible comparison. The Columbia mini model was also evaluated on the external test set (see SectionA.3of the Appendix). To assess the relative contributions of different components and design choices, we fur- ther evaluated multiple configurations of the multi-predictor approach, including alternative predictor extractors, tabular models, and varying numbers of predictors. 2.7 Interpretability The predictor-driven framework enables transparent interpretation for both single- and multi- predictor approaches. In the single-predictor approach, each predictor value directly repre- sents the probability of the corresponding ECG diagnosis, providing inherent clinical inter- pretability. We identified the most important predictors and performed combined analyses with the multi-predictor approach. For the multi-predictor approach, we employed SHAP (SHapley Additive exPlanations) values ( Lundberg and Lee,2017;Shapley et al.,1953) to quantify feature contributions at global and local levels. SHAP is a game-theoretic, additive feature attribution method that provides consistent and locally accurate explanations ( Lundberg and Lee,2017). For compu- tational efficiency in tree-based models, we used the Tree SHAP algorithm (Lundberg et al., 2018). Global explanation plots included cumulative contributions, beeswarm summaries, and SHAP versus predictor value plots to identify key predictors and characterize their be- havior in the model. Local plots highlighted the ten predictors SHAP values and displayed predicted probabilities relative to both the ECG diagnosis thresholds and the LEF decision thresholds, facilitating interpretation of individual model predictions. 2.8 Subgroup analysis We evaluated model performance across clinically relevant subgroups in both the internal and external test sets. In the internal test set, predefined subgroups included the presence or absence of other structural heart disease (SHD), valvular heart disease (VHD), age groups, sex, race/ethnicity, and clinical context (definitions of SHD and VHD are provided in Section B.1of the Appendix). In the external test set, subgroup evaluation was performed for the available variables, including age groups, sex, race/ethnicity, and clinical context. 7 3 Results 3.1 Population characteristics The ECG-ECHO (EchoNext) cohort is a publicly available benchmark comprising 82,543 ECG examinations, partitioned into training (n=72,475), validation (n=4,626), and internal test (n=5,442) sets according to the official protocol (Poterucha et al.,2025). We further constructed an external cohort (ECG-Note; n=16,017) as an independent test set (Figure A.1). Baseline demographic and clinical characteristics are summarized in Table1. Compared with the internal test cohort, the external cohort exhibited a higher proportion of male patients (54.9%) and White individuals (73.0%), whereas racial/ethnic representation in the ECG-ECHO cohort was more evenly distributed. Clinical context distributions also differed substantially, with the external cohort enriched for emergency encounters (57.3%) compared with the more evenly distributed emergency (36.2%), inpatient (40.5%), and out- patient (19.5%) settings in the internal test cohort. These differences reflect substantial demographic and clinical heterogeneity across cohorts. The distribution of the 71 extracted predictors is detailed in Appendix Table A.1, which also shows differences in binarized counts for selected predictors, including NORM and ILBBB. 3.2 Predictor recognition by the predictor extractor Reliable predictor extraction is a prerequisite for downstream LEF modeling. Transformer- PT achieved a macro AUROC of 94.5%, AUPRC of 41.4%, and F1 score of 38.3%, out- performing the original Transformer (AUROC 89.8%, AUPRC 30.7%) and demonstrating performance comparable to previously reported state-of-the-art results on PTB-XL ( Zhou et al.,2025). Classification performance for 10 representative predictors is shown in Ta- ble 2, with results for all 71 predictors provided in Appendix TableC.1. Across predictors, Transformer-PT achieved consistently strong discriminative capacity as reflected by AU- ROC values, whereas AUPRC and F1 varied due to class imbalance (see Appendix Section C.1). Importantly, the downstream LEF framework leverages continuous predictor scores rather than threshold-dependent binary decisions; thus, the strong ranking performance of Transformer-PT is sufficient to ensure reliable information transfer to subsequent modeling stages, even for ultra-rare categories. 3.3 Single-predictor approach performance We evaluated the ability of individual predictors to detect LEF using their continuous output scores, without task-specific fine-tuning (zero-shot-like inference). Results for 10 represen- tative predictors in the internal test set are summarized in Table2, with results for the remaining predictors and the external test set provided in Appendix TablesC.1andC.2. Predictors are ordered by decreasing F1 score in the internal test set. Notably, several pre- dictors demonstrated substantial standalone discriminative performance. Eight predictors (NORM, ILBBB, INJAL, ISCLA, ANEUR, ISCAL, ASMI, and SVARR) achieved AUROC values ranging from 71.0% to 81.0% internally, with five maintaining AUROC values between 70.7% and 78.8% externally. The NORM predictor was the strongest individual predictor in both cohorts, yielding an AUROC of 81.0%, AUPRC of 47.4%, and F1 score of 51.4% internally, and an AUROC of 78.8%, AUPRC of 36.2%, and F1 score of 42.3% externally. 8 Table 1:Baseline demographic and clinical characteristics in the ECG-ECHO and external ECG- Note cohorts. ECGβECHOECGβNote Training setValidation setInternal test setExternal test set Patients (n)26,2184,6265,44216,017 ECGs (n)72,4754,6265,44216,017 Age groups 18β5929,783 (41.1%)1,787 (38.6%)2,124 (39.0%)4,270 (26.7%) 60β6918,745 (25.9%)1,093 (23.6%)1,318 (24.2%)3,637 (22.7%) 70β7914,898 (20.6%)975 (21.1%)1,154 (21.2%)3,761 (23.5%) 80+9,049 (12.5%)771 (16.7%)846 (15.5%)4,349 (27.2%) Sex Female33,524 (46.3%)2,356 (50.9%)2,731 (50.2%)7,222 (45.1%) Male38,951 (53.7%)2,270 (49.1%)2,711 (49.8%)8,795 (54.9%) Race/ethnicity Hispanic22,806 (31.5%)1,351 (29.2%)1,649 (30.3%)638 (4.0%) White21,289 (29.4%)1,385 (29.9%)1,569 (28.8%)11,688 (73.0%) Black11,559 (15.9%)728 (15.7%)846 (15.5%)1,845 (11.5%) Asian2,602 (3.6%)134 (2.9%)153 (2.8%)414 (2.6%) Other 5,272 (7.3%)380 (8.2%)457 (8.4%)524 (3.3%) Unknown8,947 (12.3%)648 (14.0%)768 (14.1%)908 (5.7%) Clinical context Emergency22,811 (31.5%)1,688 (36.5%)1,971 (36.2%)9,170 (57.3%) Inpatient 34,906 (48.2%)1,903 (41.1%)2,203 (40.5%)- Outpatient12,423 (17.1%)858 (18.5%)1,059 (19.5%)- Procedural2,335 (3.2%)177 (3.8%)209 (3.8%)- Urgent---3,028 (18.9%) Observation---2,173 (13.6%) Surgical Same Day---1,022 (6.4%) Elective---624 (3.9%) Outcome Ejection Fractionβ€45%16,962 (23.4%)866 (18.7%)962 (17.7%)2,517 (15.7%) The training, validation, and internal test splits correspond to the official splits of the ECG-ECHO dataset, EchoNext (Elias and Finer,2025). The external test cohort, ECG-Note, was derived from the MIMIC-IV database and its associated ECG and clinical note modules ( Gow et al.,2023;Johnson et al.,2024,2023a). Values are shown as counts and percentages. 9 Among abnormal ECG diagnoses, ILBBB and INJAL achieved the highest discriminative performance, with internal AUROCs of 80.0% and 75.3%, and external AUROCs of 73.4% and 71.6%, respectively. Across all 71 predictors, discriminative performance varied substantially. Nevertheless, the majority of predictors, 58 in the internal test set and 54 in the external test set, achieved AUROC values above chance level (50%), suggesting that LEF-related information is dis- tributed across diverse ECG-derived predictors rather than confined to a small subset of diagnoses. Predictors with weaker standalone performance still contributed incremental im- provements when integrated into the multi-predictor model described in the next subsection. An additional observation was that optimal LEF detection thresholds were consistently substantially lower than those used for conventional ECG classification. For example, while a threshold of 0.370 for the NORM predictor identified abnormal ECGs, a substantially lower threshold of 0.003641 was optimal for LEF detection, with similar patterns observed across most predictors (Appendix FigureC.1). Table 2:Performance of traditional AI-ECG predictors and LEF detection using a single-predictor approach. Predictor Traditional ECG ModelLEF Detection (Single-predictor) AUROC AUPRC F1 Score ThreshAUROC AUPRC F1 Score Thresh NORM 94.9 (94.1β95.7) 92.8 (91.4β94.2) 85.1 (83.5β86.7) 0.370 81.0 (79.6β82.4) 47.4 (44.2β50.9) 51.4 (49.3β53.8) 0.003641 ILBBB 90.9 (65.3β99.7) 31.6 ( 8.1β70.3) 30.0 ( 0.0β55.6) 0.163 80.0 (78.4β81.5) 48.5 (45.3β52.0) 50.5 (48.0β53.1) 0.000349 INJAL 98.6 (97.3β99.6) 52.2 (27.0β79.8) 47.6 (19.0β72.0) 0.500 75.3 (73.7β76.8) 34.8 (32.4β37.6) 45.1 (42.8β47.4) 0.000386 ISCLA 92.3 (85.5β97.6) 19.0 ( 5.6β45.9) 12.5 ( 0.0β36.4) 0.248 75.9 (74.5β77.4) 38.9 (36.2β42.3) 43.9 (41.6β46.3) 0.001918 ANEUR 96.7 (91.7β99.2) 15.0 ( 6.0β37.3) 11.8 ( 0.0β33.3) 0.294 74.8 (73.1β76.6) 38.9 (35.9β42.3) 42.9 (40.4β45.4) 0.001924 ISCAL 95.3 (94.0β96.6) 33.9 (25.2β46.8) 34.8 (23.3β45.4) 0.315 73.7 (72.1β75.2) 32.1 (29.8β34.6) 42.3 (40.1β44.4) 0.001148 ASMI 98.1 (97.4β98.7) 88.4 (84.9β91.4) 80.3 (76.4β83.8) 0.300 72.8 (71.0β74.7) 37.4 (34.3β40.5) 41.8 (39.4β43.9) 0.023987 SVARR 92.0 (85.2β97.2) 21.3 ( 6.1β45.9) 22.2 ( 5.9β40.0) 0.061 71.0 (69.3β72.6) 31.8 (29.5β34.5) 40.4 (38.1β42.9) 0.000218 INJIL 92.9 (85.8β99.2) 2.3 ( 0.3β12.0) 0.0 ( 0.0β 0.0) 0.012 69.8 (68.0β71.5) 29.5 (27.3β32.0) 40.1 (37.9β42.3) 0.000035 CRBBB 99.8 (99.6β99.9) 89.1 (80.8β95.3) 83.2 (74.8β89.8) 0.144 69.7 (68.1β71.5) 28.4 (26.3β30.8) 39.5 (37.6β41.6) 0.000005 Predictor performance was obtained from the traditional automatic AI-ECG model. LEF detection perfor- mance was derived using the single-predictor approach proposed in this study. AUROC, AUPRC, and F1 are reported with 95% confidence intervals. The two Thresh columns indicate the thresholds used to max- imize the F1 score on the validation set for predictor performance and LEF detection, respectively. The 10 predictors with the highest F1 scores for LEF detection are reported in this table. NORM, normal ECG; ILBBB, incomplete left bundle branch block; INJAL, subendocardial injury in anterolateral leads; ISCLA, ischemic in lateral leads; ANEUR, ST-T changes compatible with ventricular aneurysm; ISCAL, ischemic in anterolateral leads; ASMI, anteroseptal myocardial infarction; SVARR, supraventricular arrhythmia; INJIL, subendocardial injury in inferolateral leads; CRBBB, complete right bundle branch block. 10 3.4 Multi-predictor approach performance We evaluated the ECGPD-LEF multi-predictor framework across combinations of predictor extractors and tabular models, benchmarking against the official end-to-end Columbia mini model (Poterucha et al.,2025) (Table3). Among all configurations, XGBoost combined with the post-trained Transformer predictor extractor (Transformer-PT-XGBoost) achieved the highest performance in both internal and external test sets. In the internal test set, Transformer-PT-XGBoost yielded an AUROC of 88.4% (95% CI, 87.1β89.5), an AUPRC of 69.2% (66.2β72.0), and an F1 score of 64.5% (62.2β66.7), significantly outperforming the Columbia mini model, which achieved an AUROC of 85.2% (83.9β86.5), an AUPRC of 59.9% (56.6β63.3), and an F1 score of 57.9% (55.4β60.3), with non-overlapping confidence intervals. Comparable performance improvements were observed in the external test set (Table 3). To disentangle the contributions of individual components, we conducted controlled com- parisons on the internal test set (Table3; Figure2). Holding the predictor extractor constant, XGBoost consistently outperformed logistic regression for both Transformer and Transformer- PT, with the largest gains observed for Transformer-PT (AUROC +3.5 percentage points; AUPRC +11.4 points). Holding the tabular model constant, post-training of the predic- tor extractor (Transformer-PT vs. Transformer) yielded consistent improvements across all metrics. We further examined performance as predictors were progressively added to the tabular model in descending order of single-predictor F1 score. Across configurations, performance generally improved with additional predictors, with Transformer-PT-XGBoost consistently achieving the highest performance across all predictor counts and metrics (Figure 2b-d). Other configurations followed similar trends but did not exceed the Columbia mini model baseline for AUROC or AUPRC, even at their respective peak performance. 3.5 Model explanation The tabular component of ECGPD-LEF enables both global and local interpretability via SHAP analysis. We analyzed Transformer-PT-XGBoost on the internal test set (Figure 3, Figure4). At the global level, cumulative SHAP contributions increased with the number of predictors included in the model (Figure 3a), consistent with the performance trends observed in Figure2of Section3.4. The ranking of predictors by mean absolute SHAP value showed slight differences compared with the ranking based on single-predictor F1 scores, although the top contributors remained largely consistent (Figure3b). For example, NORM, ILBBB, INJAL, ISCLA, and ANEUR ranked 1-5 by single-predictor F1 score, whereas their SHAP- based ranking was 1, 2, 5, 3, and 4, respectively. In addition, SHAP values varied markedly at low predictor values (Figure3c-g), often at magnitudes substantially below the diagnostic thresholds, which is consistent with the observation in Section 3.3. At the local level, explanation plots for a positive and a negative case illustrate individual- ized predictor contributions (Figure4). In these two examples, global importance patterns are reflected in case-specific prediction profiles, with NORM and ILBBB contributing the most. When combined with the single-predictor approach and corresponding diagnostic thresh- olds, the predictions for both cases are interpretable. In particular, the values of NORM and ILBBB exceeded the LEF-positive thresholds in the positive case and the LEF-negative thresholds in the negative case. Furthermore, the LAO/LAE predictor in the positive case 11 Table 3:Performance comparison of the proposed ECGPD-LEF framework and baseline Columbia mini model for LEF detection. MethodTabular ModelPredictor ExtractorAUROC AUPRC F1 Score Internal test set Columbia mini model β 85.2 (83.9β86.5) 59.9 (56.6β63.3) 57.9 (55.4β60.3) ECGPD-LEF Logistic Regression Transformer 84.5 (83.1β85.9) 56.8 (53.4β60.3) 57.0 (54.6β59.7) Transformer-PT 84.9 (83.6β86.3) 57.8 (54.6β61.2) 58.5 (56.2β60.8) XGBoost Transformer 84.9 (83.5β86.2) 58.5 (55.1β61.7) 57.6 (55.1β60.1) Transformer-PT 88.4 (87.1β89.5) 69.2 (66.2β72.0) 64.5 (62.2β66.7) External test set Columbia mini model β 80.8 (79.9-81.7) 45.8 (43.7-47.7) 47.7 (46.3-49.1) ECGPD-LEF XGBoostTransformer-PT 86.9 (86.2-87.7) 57.6 (55.6-59.7) 53.8 (52.5-55.0) The Columbia mini model is the official benchmark on the internal test set, EchoNext (Poterucha et al., 2025), with trained official weights. This model relies on seven tabular features; since the internal test set provides all seven features, it can be applied directly. For the external cohort, five of these features are not directly available, so we computed them from the MIMIC-IV ECG module to enable inference with the Columbia mini model (details of this feature computation are provided in SectionA.3). The proposed ECGPD-LEF framework is evaluated with different configurations of automatic AI-ECG models and tabular models. For example, "XGBoost" indicates that the tabular model is XGBoost applied to the extracted predictors, and "Transformer-PT" indicates that the predictor extractor of the proposed framework is based on Transformer-PT. Since the baseline Columbia mini model is an end-to-end method, it does not depend on the choice of predictor extractor or tabular model. This table thus compares both the performance of the baseline model and the performance of the proposed framework under different configurations. AUROC, AUPRC, and F1 are reported with 95% confidence intervals. 12 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate a Transformer-LR (AUROC = 0.845) Transformer-PT-LR (AUROC = 0.849) Transformer-XGBoost (AUROC = 0.849) Transformer-PT-XGBoost (AUROC = 0.884) Columbia mini model (AUROC = 0.852) 0.00.20.40.60.81.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision b Transformer-LR (AUPRC = 0.568) Transformer-PT-LR (AUPRC = 0.578) Transformer-XGBoost (AUPRC = 0.585) Transformer-PT-XGBoost (AUPRC = 0.692) Columbia mini model (AUPRC = 0.599) 120406071 Number of predictors 80 82 84 86 88 90 AUROC (%) c Transformer-LR Transformer-PT-LR Transformer-XGBoost Transformer-PT-XGBoost 120406071 Number of predictors 45 50 55 60 65 70 75 AUPRC (%) d Transformer-LR Transformer-PT-LR Transformer-XGBoost Transformer-PT-XGBoost 120406071 Number of predictors 50 52 54 56 58 60 62 64 F1 Score (%) e Transformer-LR Transformer-PT-LR Transformer-XGBoost Transformer-PT-XGBoost Figure 2: Model performance across different evaluation settings.The first row shows re- ceiver operating characteristic (ROC) curves (a) and precisionβrecall (PR) curves (b) for five methods using the full feature set (71 predictors). The second row presents the performance of four models evaluated with an increasing number of predictors (1β71), in terms of AUROC (c), AUPRC (d) and F1 score (e). The gray dashed line indicates the Columbia mini model and serves as a reference baseline. 13 and the NORM predictor in the negative case exceeded their standard diagnostic thresholds (Figures4b-d), indicating that these patterns are potentially human-recognizable. 1 1 1 β ( '!$&#"!# #'&$'& #) &'%( ' 000000 0 $# 0 0 00 0 0 0 $# !! " ! " ! 0 0 0 0 0 "$!"% 000000 ! 0 0 00 0 0 ! 0 0 0 0 0 !" 00000 "! 0 0 0 0 00 0 0 0 0 "! " # 00000000 $# 00 0 000 0 00 0 00 $# !! " ! " ! "$!"% 00000 #" 0 0 0 0 00 0 0 0 #" ! ! !# !$ 00000 0 0 #" 0 00 0 0 0 0 #" ! ! 0 0 0 0 0 !# !$ Figure 3: Global-level explanation of ECG model predictions across all predictors.The first row shows (a) the cumulative absolute SHAP contributions of all 71 predictors, expressed as percentages of the total contribution, and (b) a SHAP beeswarm plot for the top 10 predictors ranked by F1 score obtained from the single-predictor method. The second and third rows show (c-h) the relationships between individual predictors (NORM, ILBBB, ISCLA, ANEUR, INJIL and ASMI) and their SHAP values. Point density was estimated using a Gaussian kernel density on the log- transformed values. Each plot includes two vertical reference lines: the first (LEF threshold) indicates the threshold that achieves the optimal F1 score using a single-predictor method, and the second (Diagnosis threshold) indicates the positivity threshold defined by an independent ECG diagnosis model. 3.6 Subgroup analysis Model performance was evaluated across clinically relevant subgroups for ECGPD-LEF. Across subgroups defined by age, sex, race/ethnicity, and clinical context, the performance of the single-predictor approach (NORM, ILBBB, INJIL, and ISCLA) in the internal and 14 β #&" $! &'$ % #&" $! &'$ % β β E[ β β β β β '*&$(%$"*+($) $ $ $! $ $ $ $ '*&$(%$"*+($) ! E[ # N N N &!($&%&$!"!() !#$'!'( &' $" ( &' $" N N N ( #*&( '(&#$#*+ #!%&)#)*"( )"&$ *"( )"&$ Figure 4: Local-level explanation of model predictions for a positive and a negative case. The first row shows SHAP waterfall plots for a positive case (a) and a negative case (b), illustrating the top ten ECG predictors contributing to the prediction, with the remaining predictors aggregated as β62 other featuresβ. Feature contributions are shown in the log-odds space and sum to the final model output. The second row presents the corresponding predictorβprobability relationships for the same cases (c,d). Vertical reference lines indicate positivity thresholds derived from an independent ECG diagnosis model. 15 external test sets is presented in SectionC.4of the Appendix. These predictors remained effective across both test sets, with NORM, ILBBB, and INJIL generally achieving higher performance. For the multi-predictor approach, subgroup results for Transformer-PT-XGBoost in the internal test set are shown in Table4. Transformer-PT-XGBoost consistently outperformed the Columbia mini model. Relative improvements ranged from 2.0%-12.2% for AUROC, 3.5%-32.3% for AUPRC, and 5.9%-30.4% for F1 score across 16 subgroups. Improvements in AUPRC and F1 score exceeded 10% in 15 and 12 subgroups, respectively. Similar trends were observed in the external test set (TableB.2). Stratified analyses by the presence of structural heart disease (SHD) and valvular heart disease (VHD) are presented in FiguresB.1andB.2in the Appendix. In both settings, Transformer-PT-XGBoost demonstrated superior performance compared with the Columbia mini model. Notably, the performance of Transformer-PT-XGBoost remained stable, whereas the Columbia mini model showed reduced performance in subgroups with SHD and VHD. 4 Discussion 4.1 A structured paradigm for LEF detection In this study, we propose a structured predictor-integration framework (ECGPD-LEF) for ECG-based detection of reduced left ventricular function that departs from conventional end-to-end waveform modeling. The proper multi-predictor configuration (Transformer-PT- XGBoost) demonstrated robust performance in both the internal hold-out test set (AUROC 88.4%, F1 64.5%) and the external test set (AUROC 86.8%, F1 53.6%), significantly outper- forming a recent strong end-to-end baseline, the Columbia mini model ( Poterucha et al.,2025) (internal AUROC 85.2%, F1 57.9%; external AUROC 80.7%, F1 46.4%). Performance gains were maintained across subgroups defined by age, sex, race/ethnicity, clinical context, and the presence or absence of structural or valvular heart disease, supporting the generalizability of the approach. Within this predictor-based architecture, several clinically recognizable ECG diagnoses, including NORM, ILBBB, and INJAL, were strongly associated with LEF detec- tion (internal AUROC 75.3%β81.0%; external AUROC 71.6%β78.6%), enabling transparent interpretation of model outputs. As a lightweight extension built upon an existing ECG di- agnosis model, this structured framework provides improved discrimination while preserving modularity and clinical interpretability. 4.2 The need for publicly benchmarked AI-ECG evaluation Artificial intelligence applied to ECG has demonstrated strong performance in emerging tasks, including detection of aortic stenosis, tricuspid regurgitation and left ventricular dysfunction. However, most prior models have been developed and evaluated on institution-specific pri- vate datasets with heterogeneous population characteristics and outcome definitions, limiting reproducibility and preventing rigorous cross-model comparison. Early work by Attia et al. (2019a) trained a convolutional neural network on 44,959 patients to identify ventricular dys- function defined as ejection fractionβ€35%, with prospective validation performed at the same institution (Attia et al.,2019b). Subsequent refinements extended detection to ejec- tion fractionβ€40% and reported multisite validation using digital ECG input alone ( Carter 16 Table 4:Subgroup performance of ECGPD-LEF and the Columbia mini model in the internal test set. Columbia mini model ECGPD-LEF (Multi-predictor) Subgroup n Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59212413.6 85.2 (82.8β87.5) 53.5 (48.1β60.0) 53.8 (49.0β58.3) 87.4 (85.0β89.8) 64.6 (59.4β69.9) 59.0 (54.7β63.6) 60β69131816.6 83.8 (80.9β86.6) 57.6 (51.1β64.6) 54.7 (49.0β59.7) 87.2 (84.2β90.2) 70.6 (64.9β76.3) 63.3 (58.1β68.5) 70β79115421.7 86.2 (83.7β88.7) 65.8 (59.5β72.2) 63.7 (58.9β67.9) 89.7 (87.2β91.8) 74.7 (69.5β80.0) 70.5 (65.8β74.8) 80+84624.1 83.5 (80.4β86.5) 65.4 (59.5β71.8) 59.5 (54.7β64.4) 87.3 (84.8β89.7) 67.7 (60.9β74.1) 66.1 (61.4β70.8) Sex Female 273112.4 86.9 (85.0β88.8) 52.3 (46.9β58.0) 53.3 (49.0β57.4) 89.9 (88.0β91.7) 63.4 (58.2β68.3) 60.4 (56.0β64.5) Male271123.0 83.1 (81.2β84.7) 64.0 (59.9β68.0) 60.6 (57.5β63.5) 86.9 (85.1β88.5) 72.9 (69.1β76.1) 67.0 (64.0β69.7) Race / ethnicity Hispanic 164916.7 86.9 (84.5β89.1) 63.3 (57.2β68.5) 60.0 (55.3β64.5) 89.5 (87.3β91.7) 71.7 (66.0β76.6) 65.2 (60.1β69.3) White156916.3 84.0 (81.1β86.6) 55.7 (49.4β62.1) 55.1 (50.2β60.0) 88.6 (86.4β90.8) 64.8 (58.7β71.1) 63.6 (59.0β68.0) Black84619.3 85.4 (82.3β88.7) 65.5 (58.5β72.4) 59.8 (53.6β65.6) 88.1 (84.9β91.1) 73.6 (67.2β79.3) 66.9 (61.2β72.2) Asian15316.3 87.8 (80.2β94.3) 61.4 (44.5β81.4) 57.1 (41.5β69.4) 92.5 (87.3β96.7) 72.6 (54.7β86.2) 66.7 (50.0β79.1) Other45715.8 83.1 (78.0β87.9) 52.1 (40.7β64.0) 49.4 (40.0β58.0) 86.9 (81.5β91.6) 65.3 (53.4β76.8) 59.4 (50.0β68.4) Unknown 76822.3 84.1 (80.8β87.2) 61.5 (53.9β69.1) 61.1 (55.0β66.7) 85.8 (82.4β89.1) 68.8 (61.6β75.9) 64.7 (58.5β70.0) Clinical context Emergency 197115.7 86.0 (83.6β88.2) 59.9 (54.1β65.3) 58.5 (53.7β62.8) 87.8 (85.5β89.9) 67.1 (61.7β72.2) 62.0 (57.5β65.8) Inpatient 220324.4 81.9 (79.7β84.0) 62.4 (58.1β67.0) 59.1 (55.9β62.5) 86.1 (84.1β88.1) 71.8 (67.7β75.7) 66.8 (63.3β69.8) Outpatient 10596.6 87.3 (82.0β91.5) 44.3 (33.4β56.5) 48.4 (38.4β57.0) 90.6 (85.9β94.5) 54.6 (42.8β65.4) 56.4 (46.4β64.9) Procedural 20921.5 79.5 (71.6β85.9) 57.3 (41.7β72.6) 52.3 (37.8β63.8) 89.2 (83.0β94.0) 75.8 (62.8β86.3) 68.2 (54.5β78.2) ECGPD-LEF is configured with the predictor extractor based on Transformer-PT and the tabular model XGBoost. This configuration was selected for illustration in the internal test set. Columbia mini model is the official benchmark on the same set ( Poterucha et al.,2025). AUROC, AUPRC, and F1 are reported with 95% confidence intervals. 17 et al.,2026). Although these studies achieved strong discrimination, the absence of publicly accessible benchmarking datasets constrains transparent evaluation. Recently,Poterucha et al.(2025) addressed this gap by releasing a de-identified ECG dataset comprising 36,286 unique patients with predefined training, validation, and test splits, and by benchmarking a baseline model (the Columbia mini model) that achieved discrimination comparable to models trained on substantially larger proprietary cohorts. By providing standardized data partitions, model weights, and open-source code, this work established a transparent and reproducible evaluation framework for AI-ECG research, upon which our study enables rig- orous and fair comparison. Furthermore, to extend validation beyond a single benchmarked dataset, we developed ECG-Note using publicly available datasets and large language models, establishing an additional external validation dataset for LEF detection that captures diverse patient populations and clinical contexts. 4.3 Moving beyond end-to-end black-box modeling Most AI-ECG approaches for left ventricular dysfunction have relied on end-to-end deep learning architectures (Attia et al.,2019a,b;Carter et al.,2026;Poterucha et al.,2025). While such models achieve strong predictive performance, their limited interpretability poses challenges for clinical integration, particularly in applications that extend beyond conven- tional ECG criteria grounded in established ECG principles. In the absence of transparent mechanistic reasoning, black-box predictions may be difficult to reconcile with established diagnostic frameworks, potentially limiting clinician trust and adoption. Efforts to enhance transparency have included tabular models constructed from 555 discrete ECG measurements ( Hughes et al.,2024). Although more interpretable, these approaches depend on propri- etary commercial measurement algorithms that may vary across ECG platforms (Strodthoff et al. ,2023), thereby constraining generalizability and standardized external evaluation. In contrast, we introduce a structured representation paradigm that integrates the predictive capacity of deep learning with the interpretability of tabular modeling. Each predictor corre- sponds to clinically meaningful features and it does not rely on vendor-specific measurement pipelines. Within the publicly benchmarked setting, this approach demonstrates improved performance over the end-to-end Columbia mini model, supporting the feasibility of clinically interpretable yet high-performing AI-ECG systems. 4.4 Interpretability reveals clinically meaningful indicators for LEF detec- tion Interpretability analyses identified diagnostically informative predictors that provide mech- anistic insight into LEF detection. In the single-predictor setting, continuous probability outputs from the trained traditional ECG diagnosis model were sufficient to achieve mean- ingful discrimination in a zero-shot-like manner. For example, the predicted probability of NORM alone yielded an internal AUROC of 81.0% and an external AUROC of 78.6%. Im- portantly, this does not imply that the binary diagnosis of NORM directly indicates LEF. Rather, the continuous probability output, well below the clinical decision threshold, captures graded deviations from normal ECG patterns that are strongly associated with LEF. This finding suggests that deep neural networks encode subclinical ECG variations within their probabilistic representations, even when such variations do not cross conventional diagnostic 18 boundaries, offering a potential explanation for prior observations that large-scale deep learn- ing models can detect novel cardiovascular phenotypes from subtle signal alterations. Within the multi-predictor framework, SHAP analyses demonstrated consistent importance patterns across predictors. At the population level (Figure3), SHAP values varied substantially across probability ranges, indicating that graded shifts in ECG-derived diagnostic probabilities con- tribute meaningfully to LEF risk estimation. At the individual level (Figure4), local expla- nations highlighted the predictors driving each decision and provided human-interpretable insights into the models reasoning (Figures4c-d). Together, these findings suggest that LEF- associated ECG signatures may be decomposed into combinations of clinically recognizable diagnostic dimensions, enhancing the structural transparency of the proposed framework. 4.5 Subgroup analysis demonstrates robustness across populations Subgroup analyses demonstrated that ECGPD-LEF maintained consistent discriminatory performance across diverse demographic and clinical strata, including age, sex, race/eth- nicity, and care settings. Across all evaluated subgroups, ECGPD-LEF outperformed the Columbia Mini model in terms of AUROC, AUPRC, and F1 score. Notably, performance remained stable in patients with and without concomitant structural heart disease (SHD) or valvular heart disease (VHD), suggesting that the ECG signatures captured by ECGPD- LEF are not merely proxies for coexisting structural abnormalities but instead reflect signal components specifically associated with LEF. Despite limited representation of certain racial subgroups (e.g., Asian participants comprising 3.6% of the training cohort), comparable per- formance was observed in both internal and external validation cohorts (AUROC 92.5% and 91.0%, respectively), supporting the generalizability and potential clinical applicability of the framework. 4.6 Scalability and extensibility of the framework The proposed ECGPD-LEF framework exhibits scalability and extensibility across multiple dimensions. First, its modular design enables integration with existing ECG diagnostic mod- els that are increasingly adopted in clinical practice, allowing lightweight extension of current AI-ECG systems without requiring full architectural replacement. Second, performance im- provements observed with the Transformer-PT backbone (Table 3) suggest that advances in upstream ECG diagnostic models may directly translate into enhanced LEF detection. As ECG diagnostic models continue to evolve with larger and higher-quality datasets, further gains may be anticipated. Third, the predictor-based structure is inherently expandable. In this study, we incorporated 71 diagnostic predictors derived from PTB-XL, and observed that model performance scaled with the number of predictors included in the tabular component (Figures 2and3a). Additional clinically established ECG diagnostic features (Cardiovas- cular Committee of China Medical Womens Association et al. ,2023) may be incorporated within the same framework, enabling progressive refinement without fundamental architec- tural redesign. Together, these properties underscore the scalability and extensibility of the framework for future AI-ECG development. 19 4.7 Limitations and future directions Despite the favorable performance and robustness of ECGPD-LEF compared with the latest deep learning baseline, several limitations warrant consideration. First, the multi-predictor framework was developed using labels derived from ECHO reports, which may be subject to inter-observer variability in ultrasound interpretation. Although external validation on the MIMIC-IV-Note cohortcomprising data from heterogeneous cardiac imaging sourcespartially mitigates this concern, potential labeling inconsistencies cannot be fully excluded. Second, while both the EchoNext and MIMIC- IV-Note datasets include diverse populations, further validation in larger, international cohorts is warranted to ensure broad generalizability. Third, the predictor extractor was pretrained on a moderately sized ECG diagnosis dataset due to the limited availability of large-scale public ECG corpora. Leveraging larger, high-quality ECG diagnosis datasets may further enhance feature representation and downstream LEF detection performance. Future work should explore scaling strategies and prospective clinical validation to confirm real-world clinical applicability. 5 Conclusion In summary, ECGPD-LEF provides a clinically interpretable, modular, and high-performing framework for ECG-based detection of LEF. By building upon existing ECG diagnostic mod- els, it achieves superior performance compared with strong black-box baselines, while remain- ing lightweight and scalable. The structured predictor-based design enables transparent in- terpretation, revealing clinically meaningful indicators of LEF, such as probabilistic outputs from NORM and other predictors. By combining accuracy, stability, and interpretability, this framework provides a practical and scalable screening tool for real-world clinical applications. 20 A Supplementary for the datasets A.1 Population characteristics (estimated predictors) TableA.1summarizes the remaining population characteristics of the ECG-ECHO and ECG- Note datasets, as estimated by a traditional automatic ECG diagnosis model (Transformer- PT). It should be emphasized that these characteristics are derived from model-based pre- dictions rather than direct annotations from the original data collection process. Based on these estimates, these datasets cover a broad range of diagnostic ECG predictors. Table A.1: Estimated patient characteristics across the training, validation, and test splits of the EchoNext dataset and the external cohort. ECG-ECHOECGβNote Training setValidation setTest setExternal Set Estimated predictor NORM19,259 (26.6%)1,485 (32.1%)1,806 (33.2%)3,343 (20.9%) ILBBB161 (0.2%)8 (0.2%)8 (0.1%)74 (0.5%) INJAL666 (0.9%)48 (1.0%)48 (0.9%)108 (0.7%) ISCLA109 (0.2%)3 (0.1%)6 (0.1%)24 (0.1%) ANEUR 185 (0.3%)13 (0.3%)9 (0.2%)27 (0.2%) ISCAL941 (1.3%)50 (1.1%)59 (1.1%)181 (1.1%) ASMI12,498 (17.2%)641 (13.9%)759 (13.9%)2,469 (15.4%) SVARR690 (1.0%)47 (1.0%)57 (1.0%)127 (0.8%) INJIL 2,173 (3.0%)125 (2.7%)171 (3.1%)393 (2.5%) CRBBB6,223 (8.6%)361 (7.8%)382 (7.0%)1,147 (7.2%) LAFB9,713 (13.4%)590 (12.8%)701 (12.9%)1,965 (12.3%) ALMI 1,487 (2.1%)67 (1.4%)62 (1.1%)256 (1.6%) ABQRS19,488 (26.9%)1,074 (23.2%)1,179 (21.7%)3,267 (20.4%) CLBBB2,076 (2.9%)137 (3.0%)166 (3.1%)506 (3.2%) ILMI1,536 (2.1%)95 (2.1%)105 (1.9%)334 (2.1%) INJAS 3,236 (4.5%)167 (3.6%)190 (3.5%)637 (4.0%) INVT6,849 (9.5%)396 (8.6%)400 (7.4%)1,268 (7.9%) PVC3,850 (5.3%)253 (5.5%)260 (4.8%)864 (5.4%) ISCIL1,140 (1.6%)63 (1.4%)79 (1.5%)180 (1.1%) 1AVB 3,462 (4.8%)202 (4.4%)231 (4.2%)640 (4.0%) ISC_5,632 (7.8%)346 (7.5%)398 (7.3%)1,675 (10.5%) IVCD3,062 (4.2%)173 (3.7%)204 (3.7%)696 (4.3%) LAO/LAE6,886 (9.5%)375 (8.1%)460 (8.5%)1,330 (8.3%) ISCAN3,160 (4.4%)154 (3.3%)153 (2.8%)507 (3.2%) ISCAS 1,109 (1.5%)62 (1.3%)51 (0.9%)202 (1.3%) AFIB5,778 (8.0%)362 (7.8%)414 (7.6%)1,158 (7.2%) BIGU430 (0.6%)26 (0.6%)34 (0.6%)89 (0.6%) SVTAC173 (0.2%)14 (0.3%)12 (0.2%)44 (0.3%) AMI 758 (1.0%)43 (0.9%)54 (1.0%)99 (0.6%) Continued on next page 21 Table A.1 (continued) ECG-ECHOECGβNote Training setValidation setInternal test setExternal test set Estimated predictor NST_5,952 (8.2%)371 (8.0%)463 (8.5%)1,241 (7.7%) 3AVB74 (0.1%)7 (0.2%)5 (0.1%)21 (0.1%) IMI9,848 (13.6%)574 (12.4%)614 (11.3%)1,679 (10.5%) LPR2,619 (3.6%)168 (3.6%)243 (4.5%)607 (3.8%) 2AVB326 (0.4%)29 (0.6%)29 (0.5%)83 (0.5%) DIG71 (0.1%)1 (0.0%)8 (0.1%)9 (0.1%) LMI 2,002 (2.8%)117 (2.5%)118 (2.2%)343 (2.1%) LOWT2,415 (3.3%)167 (3.6%)211 (3.9%)443 (2.8%) SR54,557 (75.3%)3,546 (76.7%)4,211 (77.4%)11,166 (69.7%) STACH9,938 (13.7%)546 (11.8%)620 (11.4%)1,642 (10.3%) LPFB3,138 (4.3%)160 (3.5%)159 (2.9%)497 (3.1%) PACE154 (0.2%)3 (0.1%)7 (0.1%)18 (0.1%) WPW57 (0.1%)0 (0.0%)3 (0.1%)3 (0.0%) ISCIN 2,244 (3.1%)141 (3.0%)155 (2.8%)427 (2.7%) PRC(S)1,016 (1.4%)72 (1.6%)70 (1.3%)219 (1.4%) AFLT45 (0.1%)1 (0.0%)5 (0.1%)12 (0.1%) INJIN 1,746 (2.4%)104 (2.2%)124 (2.3%)395 (2.5%) PAC 3,187 (4.4%)202 (4.4%)251 (4.6%)733 (4.6%) IPMI10 (0.0%)1 (0.0%)0 (0.0%)4 (0.0%) STD_11,813 (16.3%)756 (16.3%)840 (15.4%)2,745 (17.1%) LNGQT 993 (1.4%)43 (0.9%)62 (1.1%)154 (1.0%) TRIGU1,564 (2.2%)100 (2.2%)117 (2.1%)370 (2.3%) NDT 5,067 (7.0%)333 (7.2%)383 (7.0%)1,270 (7.9%) LVH 19,728 (27.2%)1,286 (27.8%)1,483 (27.3%)4,312 (26.9%) PSVT 236 (0.3%)18 (0.4%)18 (0.3%)49 (0.3%) INJLA4,521 (6.2%)268 (5.8%)271 (5.0%)949 (5.9%) PMI 4 (0.0%)0 (0.0%)1 (0.0%)1 (0.0%) STE_38,602 (53.3%)2,595 (56.1%)3,021 (55.5%)7,752 (48.4%) SEHYP5,294 (7.3%)313 (6.8%)348 (6.4%)1,017 (6.3%) SBRAD876 (1.2%)75 (1.6%)94 (1.7%)255 (1.6%) RAO/RAE7,546 (10.4%)370 (8.0%)422 (7.8%)1,601 (10.0%) VCLVH20,708 (28.6%)1,384 (29.9%)1,684 (30.9%)4,715 (29.4%) IRBBB6,541 (9.0%)328 (7.1%)332 (6.1%)1,037 (6.5%) QWAVE8,068 (11.1%)453 (9.8%)507 (9.3%)1,590 (9.9%) NT_ 2,052 (2.8%)150 (3.2%)177 (3.3%)358 (2.2%) EL3,703 (5.1%)241 (5.2%)289 (5.3%)825 (5.1%) HVOLT13,380 (18.5%)1,009 (21.8%)1,181 (21.7%)3,329 (20.8%) RVH1,718 (2.4%)101 (2.2%)103 (1.9%)295 (1.8%) LVOLT243 (0.3%)12 (0.3%)20 (0.4%)79 (0.5%) Continued on next page 22 Table A.1 (continued) ECG-ECHOECGβNote Training setValidation setInternal test setExternal test set Estimated predictor SARRH1,298 (1.8%)112 (2.4%)131 (2.4%)386 (2.4%) IPLMI159 (0.2%)8 (0.2%)11 (0.2%)34 (0.2%) TAB_86 (0.1%)11 (0.2%)5 (0.1%)8 (0.0%) Values are shown as counts and percentages. Counts are derived from automatic AI-ECG diagnoses generated by Transformer-PT. Predictor abbreviations are defined inWagner et al.(2020). A.2 Construction of the external test set We constructed the external test set ECG-Note as FigureA.1. Adult patients, who had at least one standard 10 s 12-lead ECG and at least one clinical note within one year following the ECG from MIMIC-IV (Johnson et al.,2024), MIMIC-IV-ECG (Gow et al.,2023), and MIMIC-IV-Note (Johnson et al.,2023a), were initially considered, leading to a starting pool of 1,113,547 ECG- note pairs from 103,505 patients with 521,654 ECGs and 247,313 notes. Exclusion criteria included: 1) ECG-Note pairs where the clinical note did not contain key- words related to ejection fraction(EF) and Echo/TTE; 2) ECG-Note pairs where the ECG contained NaN data. After exclusions, 382,308 ECG-Note paired data of 37,234 patients, comprising 246,543 ECGs and 62,023 notes, were included for further analysis. For the note set, we employed a Large Language Model (LLM), specifically Qwen2.5-72B-Instruct( Qwen et al.,2025), to extract precise ejection fraction(EF) values derived solely from the concurrent Echo/TTE (Fig.A.1b). Based on the values obtained by the LLM, we determined whether the ejection fraction wasβ€45%, suggesting moderately reduced systolic function. These labels were cross-referenced with MIMIC-ECG-LVEF ( Li et al.,2025) to identify consistent ECG-EF class pairs, resulting in the collection of 55,318 available ECGs. To ensure the ac- curacy of subsequent evaluations, only the ECG-Note pair with the shortest time interval for each patient was retained among the valid samples, resulting in a final testing set of 16,017 ECGs. A.3 Comparison with the Columbia mini model in the external test set In the ECG-Note dataset, we compared the proposed ECGPD-EF with the Columbia mini model ( Poterucha et al.,2025). ECGPD-EF requires only raw ECG signals for inference. In contrast, the Columbia mini model requires seven tabular features in addition to the ECG waveform: sex, age, PR interval, QRS duration, corrected QT interval (QTc), atrial rate, and ventricular rate. While age and sex were directly obtained from demographic records, the remaining five electrocardiographic metrics were not explicitly available and were derived using themachine_measurementstable from the MIMIC-IV ECG module. Specifically, letT P onset ,T QRS onset ,T QRS end , andT T end denote the timings of P-wave onset, QRS onset, QRS offset, and T-wave offset, respectively, and letRRdenote the R interval in milliseconds. The PR interval was calculated asT QRS onset βT P onset , and the QRS duration was derived asT QRS end βT QRS onset . The QTc interval was computed using Bazettβs formula: 23 Figure A.1: Construction of the external test set. (a)Flowchart illustrating the data selection process, including exclusion criteria applied to the MIMIC-IV database to derive the final high-quality testing set.(b)The prompt template used for the Large Language Model (LLM) Clinical Note Extraction Module, designed to extract quantitative ejection fraction(EF) values from unstructured clinical notes into a structured JSON format. 24 (T T end βT QRS onset )/ β R/1000. The ventricular rate was calculated as60,000/R(Kligfield et al.,2007). Due to the absence of specific atrial rate data, the derived ventricular rate was used as a proxy for the atrial rate. 25 B Supplementary analyses for the multi-predictor approach B.1 Subgroup analysis of the multi-predictor approach stratified by SHD and VHD Baseline echocardiographic characteristics are presented in Table B.1. Subgroup analysis results stratified by the presence of SHD and VHD are shown in FiguresB.1andB.2. Table B.1:Baseline echocardiographic characteristics across the training, validation, and test splits of the ECG-ECHO dataset. Training setValidation setTest set Echocardiographic findings LVWTβ₯1.3cm17,667 (24.4%)877 (19.0%)1,061 (19.5%) Aortic stenosis2,919 (4.0%)252 (5.4%)286 (5.3%) Aortic regurgitation878 (1.2%)62 (1.3%)66 (1.2%) Mitral regurgitation6,137 (8.5%)282 (6.1%)337 (6.2%) Tricuspid regurgitation7,707 (10.6%)305 (6.6%)353 (6.5%) Pulmonary regurgitation603 (0.8%)21 (0.5%)20 (0.4%) RV systolic dysfunction9,597 (13.2%)368 (8.0%)419 (7.7%) Pericardial effusion2,079 (2.9%)52 (1.1%)69 (1.3%) PASPβ₯45mmHg13,727 (18.9%)581 (12.6%)699 (12.8%) TR Vmaxβ₯3.2cm/s 7,492 (10.3%)267 (5.8%)375 (6.9%) The training, validation, and test splits correspond to the official splits of the ECG-ECHO dataset (EchoNext) (Elias and Finer,2025). Values are shown as counts and percentages. LVWT, left ventricular wall thickness; RV, right ventricular; PASP, pulmonary artery systolic pressure. Other structural heart disease (SHD).For subgroup analysis, individuals were clas- sified as having other structural heart disease (SHD) if at least one of the following echocardio- graphic abnormalities was present: left ventricular hypertrophy (left ventricular wall thick- nessβ₯1.3 cm); moderate or greater valvular heart disease (aortic, mitral, tricuspid, or pulmonary); moderate or greater right ventricular systolic dysfunction; moderate or large pericardial effusion; or pulmonary hypertension, defined as pulmonary artery systolic pres- sure (PASP)β₯45 mmHg or tricuspid regurgitation peak velocity (TR Vmax)β₯3.2 m/s. Individuals without any of the above findings were classified as having no other SHD. The performance comparison of ECGPD-LEF and the Columbia mini model in patients with and without other SHD is presented in Figure B.1. Valvular heart disease (VHD).For the VHD subgroup analysis, individuals were cat- egorized according to the presence of moderate or greater valvular heart disease on echocar- diography, including aortic, mitral, tricuspid, or pulmonary valve disease. Those without moderate or greater valvular abnormalities were classified as non-VHD. The corresponding performance comparison between ECGPD-LEF and the Columbia mini model in VHD and non-VHD populations is shown in Figure B.2. 26 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate a Transformer-PT-XGBoost (AUROC = 0.852) Columnbia mini model (AUROC = 0.797) 0.00.20.40.60.81.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision b Transformer-PT-XGBoost (AUPRC = 0.760) Columnbia mini model (AUPRC = 0.665) 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate c Transformer-PT-XGBoost (AUROC = 0.860) Columnbia mini model (AUROC = 0.832) 0.00.20.40.60.81.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision d Transformer-PT-XGBoost (AUPRC = 0.535) Columnbia mini model (AUPRC = 0.437) Figure B.1: Model performance stratified by other structural heart disease (SHD).ROC (a,c) and PR (b,d) curves for individuals with other SHD (a,b) and without other SHD (c,d). Other SHD was defined as the presence of at least one predefined echocardiographic abnormality independent of left ventricular ejection fraction (see AppendixB.1). Transformer-PT-XGBoost is compared with the Columbia mini model. AUROC and AUPRC are indicated in each panel. 27 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate a Transformer-PT-XGBoost (AUROC = 0.853) Columnbia mini model (AUROC = 0.799) 0.00.20.40.60.81.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision b Transformer-PT-XGBoost (AUPRC = 0.812) Columnbia mini model (AUPRC = 0.726) 0.00.20.40.60.81.0 False Positive Rate 0.0 0.2 0.4 0.6 0.8 1.0 True Positive Rate c Transformer-PT-XGBoost (AUROC = 0.873) Columnbia mini model (AUROC = 0.838) 0.00.20.40.60.81.0 Recall 0.0 0.2 0.4 0.6 0.8 1.0 Precision d Transformer-PT-XGBoost (AUPRC = 0.621) Columnbia mini model (AUPRC = 0.516) Figure B.2: Model performance stratified by valvular heart disease (VHD).ROC (a,c) and PR (b,d) curves for individuals with moderate or greater VHD (a,b) and without moderate or greater VHD (c,d). VHD was defined as moderate or greater valvular heart disease on echocardiography (see AppendixB.1). Transformer-PT-XGBoost is compared with the Columbia mini model. AUROC and AUPRC are indicated in each panel. 28 B.2 Subgroup analysis for multi-predictor method in the external test set To evaluate model robustness and fairness across patient characteristics and clinical settings, we conducted a subgroup analysis comparing our proposed Transformer-PT-XGBoost with Columbia mini model. Subgroups were defined by age, sex, race and ethnicity, and clinical context. Data linkage between the MIMIC-IV and MIMIC-IV-ECG datasets was performed to extract these subgroup attributes. Patient age was recalculated specifically at the time of ECG acquisition by leveraging the time-invariant interval method to restore temporal alignment (Johnson et al.,2023b). For race and ethnicity, we aggregated granular categories into broader clusters (TableB.3). Regarding clinical context, we determined the admission type by mapping the ECG timestamp to the corresponding hospital admission and discharge window. All performance metrics are reported using the AUROC and the AUPRC (Table B.2). Overall, our model demonstrates comparable or superior performance across almost all subgroups, suggesting that the proposed model generalizes well across heterogeneous patient populations and care environments, supporting its potential utility in real-world deployment scenarios. 29 Table B.2:Subgroup performance of ECGPD-LEF and the Columbia mini model in the external test set. Columbia mini model ECGPD-LEF (Multi-predictor) Subgroupn Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59427012.2 83.0 (80.9β84.9) 46.0 (41.5β50.8) 49.1 (45.4β52.4) 89.5 (88.1β90.9) 62.1 (57.4β66.4) 57.3 (54.2β60.4) 60β69363715.7 82.0 (80.2β83.9) 47.7 (43.7β52.6) 50.1 (46.9β53.1) 88.8 (87.3β90.3) 65.1 (61.4β69.2) 56.3 (53.4β59.3) 70β79376118.0 81.1 (79.2β82.8) 51.6 (47.7β55.9) 50.6 (48.0β53.3) 85.5 (84.0β86.9) 56.6 (53.0β60.9) 54.4 (51.6β57.0) 80+434917.2 77.9 (76.1β79.7) 42.9 (39.5β46.8) 43.5 (40.9β45.9) 83.5 (82.1β85.1) 49.2 (45.6β53.2) 49.6 (47.2β52.1) Sex Female722211.2 79.4 (77.8β80.9) 32.5 (29.5β36.0) 38.3 (35.9β40.5) 85.9 (84.6β87.2) 45.6 (41.9β48.9) 44.9 (42.6β47.1) Male879519.4 81.3 (80.2β82.4) 52.7 (50.2β55.5) 53.4 (51.7β55.1) 87.2 (86.3β88.1) 63.9 (61.4β66.4) 59.0 (57.3β60.6) Race/ethnicity Hispanic63814.7 84.9 (80.4β88.9) 53.6 (43.3β64.1) 57.1 (49.1β64.1) 89.9 (86.9β92.6) 64.7 (55.0β74.1) 53.5 (46.4β60.4) White1168815.5 80.3 (79.2β81.3) 43.9 (41.6β46.5) 46.5 (44.9β48.2) 86.9 (86.0β87.8) 61.9 (59.4β64.4) 56.9 (55.4β58.4) Black184517.3 84.7 (82.3β86.9) 57.7 (51.9β63.1) 54.1 (50.1β58.1) 89.8 (88.0β91.4) 72.9 (67.5β78.0) 62.1 (57.9β66.1) Asian41412.6 84.6 (79.2β89.1) 43.4 (31.5β58.1) 47.7 (37.7β56.1) 90.7 (86.1β94.1) 56.0 (40.5β72.6) 55.3 (45.2β64.7) Other52413.0 79.8 (73.5β85.4) 39.5 (29.3β51.9) 44.4 (35.8β52.9) 87.5 (82.7β91.6) 53.5 (41.3β66.4) 52.7 (44.4β60.9) Unknown 90818.8 75.9 (71.6β79.6) 47.4 (40.3β56.4) 45.6 (40.0β51.5) 81.3 (77.7β84.7) 54.6 (45.6β65.2) 51.9 (46.3β57.1) Clinical context Emergency 917014.6 81.4 (80.1β82.6) 44.3 (41.8β47.3) 46.3 (44.4β48.2) 88.4 (87.4β89.4) 61.0 (58.4β63.7) 55.6 (53.7β57.2) Urgent302820.9 79.1 (77.1β81.0) 50.7 (46.7β54.9) 53.1 (50.1β55.8) 84.1 (82.6β85.7) 60.2 (56.0β64.2) 60.4 (57.7β62.7) Observation 217319.6 83.9 (81.8β85.9) 58.8 (53.7β64.0) 56.3 (52.8β59.9) 89.4 (87.7β91.1) 72.6 (67.7β77.3) 63.6 (60.2β66.8) Surgical Same Day 10226.5 76.8 (70.9β82.2) 22.1 (15.1β33.2) 25.7 (18.9β32.5) 85.9 (81.3β89.8) 26.8 (18.7β38.1) 36.6 (27.5β44.6) Elective6249.3 72.2 (65.1β79.0) 21.0 (14.7β32.0) 27.0 (20.6β33.6) 82.2 (76.8β87.1) 34.3 (25.0β47.5) 35.1 (27.9β42.8) ECGPD-LEF is configured with the predictor extractor based on Transformer-PT and the tabular model XGBoost. This configuration was selected for illustration in the internal test set. Columbia mini model is the official benchmark on the same set ( Poterucha et al.,2025). AUROC, AUPRC, and F1 are reported with 95% confidence intervals. 30 Table B.3:Mapping of original race/ethnicity categories to aggregated subgroups. Grouped Original Categories AsianASIAN, ASIAN - ASIAN INDIAN, ASIAN - CHINESE, ASIAN - KOREAN, ASIAN - SOUTH EAST ASIAN BlackBLACK/AFRICAN, BLACK/AFRICAN AMERICAN, BLACK/CAPE VERDEAN, BLACK/CARIBBEAN ISLAND Hispanic HISPANIC OR LATINO, CENTRAL AMERI- CAN, COLUMBIAN, CUBAN, DOMINICAN, GUATEMALAN, HONDURAN, MEXICAN, PUERTO RICAN, SALVADORAN, SOUTH AMERICAN WhiteWHITE, WHITE - BRAZILIAN, EASTERN EU- ROPEAN, OTHER EUROPEAN, RUSSIAN, POR- TUGUESE OtherOTHER, AMERICAN INDIAN/ALASKA NATIVE, MULTIPLE RACE/ETHNICITY, PACIFIC IS- LANDER Unknown UNKNOWN, UNABLE TO OBTAIN, DECLINED 31 C Supplementary analyses for the single-predictor approach C.1 Predictor recognition by the predictor extractor (remaining predic- tors) The classification performance of the predictor extractor for the remaining 61 predictors is reported in TableC.1. Across the 71 predictors, AUROC values ranged from 73.8% to 100%, with 58 exceeding 90%, indicating consistently strong discriminative capacity. AUPRC and F1 scores varied substantially (0.3%-97.6% and 0.0%-94.1%), largely due to class imbalance. For instance, the INJIL category included only two positive cases in the PTB-XL test set (<0.1% prevalence), resulting in unstable precision-recall estimates (test F1 = 0.0%; valida- tion F1 = 8.7%) despite a high AUROC of 92.9%. Decision thresholds were set to maximize F1 on the validation set, ensuring all predictors achieved non-zero validation F1 scores. Since the downstream LEF framework leverages continuous predictor scores rather than threshold-dependent binary decisions, robust ranking is sufficient to reliably transfer information even for ultra-rare categories. C.2 Single-predictor performance (remaining predictors and external re- sults) The performance of the remaining single predictors for LEF detection is reported in Table C.1. External test set performance is provided in TableC.2. Table C.1:Performance of traditional AI-ECG predictors and LEF detection using a single-predictor approach (remaining 61 predictors) Predictor Traditional ECG ModelLEF Detection (One Predictor) AUROC AUPRC F1 Score ThreshAUROC AUPRC F1 Score Thresh LAFB 98.8 (98.3β99.2) 87.0 (82.1β91.2) 78.9 (74.2β83.5) 0.274 70.0 (68.2β71.7) 29.1 (26.9β31.6) 39.0 (36.8β41.1) 0.001226 ALMI 97.3 (95.0β99.1) 61.3 (43.6β77.1) 58.3 (40.0β73.2) 0.257 68.6 (66.8β70.6) 35.0 (32.3β38.3) 38.6 (36.1β41.1) 0.002106 ABQRS 87.5 (85.5β89.5) 57.4 (51.7β63.5) 54.4 (50.5β58.2) 0.053 68.3 (66.5β70.1) 29.8 (27.5β32.5) 38.0 (35.9β40.1) 0.012123 CLBBB 99.8 (99.6β100.0) 93.0 (85.9β98.4) 89.3 (82.8β94.7) 0.053 67.9 (66.0β69.7) 35.4 (32.5β38.8) 38.0 (35.5β40.2) 0.000006 ILMI 95.7 (91.9β98.5) 63.4 (50.6β75.7) 61.1 (48.9β71.2) 0.157 67.6 (65.7β69.4) 31.7 (29.0β34.8) 37.9 (35.9β40.1) 0.000172 INJAS 99.1 (98.6β99.5) 51.2 (32.1β70.9) 44.4 (25.8β60.7) 0.257 66.8 (64.9β68.7) 27.0 (25.1β29.2) 37.9 (36.0β40.0) 0.000571 INVT 95.3 (92.5β97.4) 27.1 (14.4β45.4) 30.8 (15.4β44.4) 0.269 66.6 (64.7β68.6) 28.4 (26.3β30.8) 37.8 (35.8β39.9) 0.005779 PVC 99.3 (98.9β99.6) 79.1 (69.8β88.2) 85.3 (80.4β89.8) 0.309 67.4 (65.6β69.1) 29.7 (27.3β32.4) 37.3 (35.2β39.5) 0.001254 ISCIL 95.8 (93.4β97.8) 17.4 ( 7.6β36.7) 25.0 (12.5β37.7) 0.053 66.0 (64.3β67.6) 24.5 (22.8β26.3) 37.3 (35.5β39.4) 0.000142 1AVB 98.6 (98.1β99.1) 68.9 (57.4β80.4) 68.2 (59.6β76.1) 0.226 67.8 (65.8β69.5) 28.8 (26.5β31.3) 37.2 (35.1β39.1) 0.000059 ISC_ 96.1 (94.4β97.4) 67.1 (58.7β74.7) 59.3 (52.3β66.7) 0.516 66.8 (64.8β68.6) 29.9 (27.4β32.5) 37.0 (34.9β39.1) 0.007637 Continued on next page 32 Table C.1 (continued) Predictor Traditional ECG ModelLEF Detection (One Predictor) AUROC AUPRC F1 Score ThreshAUROC AUPRC F1 Score Thresh IVCD 80.3 (74.4β84.8) 20.6 (13.7β30.4) 29.3 (19.8β38.2) 0.209 66.2 (64.1β68.0) 32.4 (29.5β35.3) 36.7 (34.2β39.1) 0.027054 LAO/LAE 88.3 (83.6β92.6) 16.0 (10.0β28.2) 26.5 (14.1β37.6) 0.189 64.1 (61.9β66.3) 32.9 (30.0β36.0) 36.5 (34.0β39.0) 0.051849 ISCAN 93.8 (91.3β96.7) 1.7 ( 0.6β 4.6) 0.0 ( 0.0β 0.0) 0.024 64.7 (62.8β66.6) 26.2 (24.2β28.7) 35.8 (33.7β37.9) 0.000312 ISCAS 96.7 (94.9β98.1) 21.3 ( 8.4β41.1) 20.5 ( 5.1β37.5) 0.247 62.4 (60.7β64.2) 21.6 (20.2β23.2) 35.5 (33.6β37.2) 0.000066 AFIB 98.6 (97.4β99.6) 93.4 (89.6β96.4) 91.4 (87.8β94.3) 0.134 64.4 (62.5β66.3) 27.5 (25.2β30.1) 35.4 (33.4β37.5) 0.000004 BIGU 97.1 (92.3β99.8) 37.5 ( 9.7β73.1) 42.1 (11.8β66.7) 0.143 63.2 (61.4β65.0) 23.9 (22.1β26.0) 34.9 (32.9β37.0) 0.000013 SVTAC 98.6 (96.5β100.0) 16.1 ( 1.4β75.0) 25.0 ( 0.0β66.7) 0.277 63.6 (61.7β65.4) 26.5 (24.4β29.0) 34.4 (32.2β36.6) 0.000113 AMI 92.3 (88.7β95.4) 31.4 (17.0β48.7) 38.8 (23.5β52.6) 0.188 61.3 (59.6β62.9) 22.2 (20.5β24.1) 34.4 (32.7β36.0) 0.000868 NST_ 86.2 (82.6β89.7) 19.4 (13.6β28.6) 25.9 (18.6β33.0) 0.171 60.2 (58.4β61.9) 20.7 (19.2β22.5) 34.4 (32.7β36.1) 0.005043 3AVB 99.6 (99.0β100.0) 55.6 (4.8β100.0) 50.0 (0.0β100.0) 0.254 61.0 (59.0β63.1) 24.8 (22.9β27.1) 34.2 (32.4β36.1) <0.000001 IMI 94.8 (93.8β95.8) 73.2 (67.8β78.2) 66.1 (61.7β70.5) 0.218 63.1 (61.2β65.0) 25.6 (23.5β28.1) 34.2 (32.3β36.2) 0.003210 LPR 98.4 (97.5β99.1) 51.4 (35.8β69.4) 47.6 (33.3β60.5) 0.260 61.6 (59.7β63.6) 23.2 (21.5β25.4) 33.7 (31.7β35.6) 0.000049 2AVB 99.6 (99.4β99.9) 11.1 (7.1β40.0) 0.0 (0.0β0.0) 0.028 61.2 (59.3β63.0) 24.0 (22.2β26.1) 33.3 (31.4β35.2) 0.000002 DIG 93.6 (90.1β96.5) 8.6 (4.4β19.3) 7.4 (0.0β22.2) 0.433 60.6 (58.7β62.4) 22.2 (20.6β24.1) 33.3 (31.5β35.1) 0.000024 LMI 93.6 (90.6β96.3) 10.6 (5.6β21.3) 20.8 (5.3β35.7) 0.176 60.8 (58.9β62.8) 25.1 (23.1β27.5) 33.0 (31.2β34.9) 0.000937 LOWT 91.8 (89.8β93.7) 12.5 (8.7β19.5) 15.4 (7.1β25.0) 0.225 58.1 (56.2β59.8) 20.3 (18.8β21.9) 32.9 (31.3β34.8) 0.000556 SR 92.7 (91.2β94.1) 96.7 (95.7β97.6) 94.1 (93.3β94.9) 0.277 62.3 (60.3β64.3) 25.6 (23.4β28.2) 32.7 (30.8β34.4) 0.943848 STACH 99.4 (98.8β99.8) 86.5 (77.1β94.4) 85.9 (80.2β91.0) 0.231 60.2 (58.2β62.0) 22.0 (20.3β23.9) 32.5 (30.9β34.2) 0.000001 LPFB 98.6 (97.7β99.4) 48.3 (26.4β69.2) 43.9 (22.8β61.9) 0.242 59.8 (58.0β61.8) 23.1 (21.3β25.2) 32.4 (30.6β34.4) 0.000076 PACE 98.1 (95.4β100.0) 87.4 (73.5β96.9) 82.4 (69.2β92.6) 0.425 58.7 (56.8β60.5) 22.4 (20.7β24.5) 31.9 (30.1β33.6) 0.000033 WPW 95.0 (84.6β100.0) 62.2 (24.3β97.6) 66.7 (17.7β93.4) 0.695 60.1 (58.2β62.1) 24.0 (22.1β26.2) 31.9 (29.8β33.9) 0.000028 ISCIN 94.3 (90.0β97.2) 23.8 (8.6β41.1) 27.9 (15.0β40.4) 0.060 58.4 (56.6β60.2) 20.8 (19.4β22.4) 31.8 (29.9β33.8) 0.000432 PRC(S) 99.8 (99.5β100.0) 16.7 (9.1β57.1) 11.8 (8.3β38.1) 0.009 58.6 (56.8β60.5) 21.7 (20.1β23.7) 31.7 (29.6β33.6) 0.000008 AFLT 86.5 (53.9β100.0) 61.5 (19.1β97.3) 44.4 (0.0β83.4) 0.950 58.0 (56.1β60.0) 22.9 (21.0β25.1) 31.5 (29.9β33.2) 0.000010 INJIN 99.9 (99.6β100.0) 64.3 (12.5β100.0) 11.8 (5.0β27.0) 0.006 58.9 (57.0β60.8) 22.9 (21.0β25.1) 31.2 (29.0β33.3) 0.000063 Continued on next page 33 Table C.1 (continued) Predictor Traditional ECG ModelLEF Detection (One Predictor) AUROC AUPRC F1 Score ThreshAUROC AUPRC F1 Score Thresh PAC 98.1 (97.2β98.8) 48.6 (34.3β67.8) 51.9 (40.4β62.6) 0.298 56.9 (54.8β58.7) 21.0 (19.4β22.8) 31.1 (29.1β33.0) 0.000290 IPMI 98.6 (97.3β99.9) 10.2 (1.7β47.4) 0.0 (0.0β0.0) 0.287 57.9 (55.9β59.9) 22.8 (20.9β25.1) 31.1 (29.4β32.8) 0.000016 STD_ 89.6 (86.6β92.1) 26.1 (19.9β34.5) 39.3 (32.5β46.2) 0.178 58.0 (55.9β60.1) 22.4 (20.7β24.4) 30.8 (29.2β32.4) 0.004936 LNGQT 97.5 (95.0β99.2) 20.4 (7.6β45.8) 30.0 (0.0β54.6) 0.466 54.0 (52.0β55.9) 19.5 (18.0β21.3) 30.4 (28.8β31.9) 0.000042 TRIGU 99.3 (98.3β100.0) 28.2 (2.9β100.0) 9.1 (3.9β22.9) 0.004 55.5 (53.5β57.6) 21.0 (19.4β22.8) 30.2 (28.3β32.1) 0.000002 NDT 93.8 (92.5β95.0) 58.0 (50.7β65.3) 57.9 (52.0β63.4) 0.328 42.2 (40.1β44.1) 14.9 (13.8β16.2) 30.1 (28.7β31.5) <0.000001 LVH 93.6 (92.0β95.1) 69.2 (63.1β74.9) 63.6 (58.7β68.5) 0.267 55.2 (52.9β57.3) 22.9 (21.0β25.2) 30.0 (28.7β31.5) 0.000005 PSVT 99.9 (99.6β100.0) 64.3 (12.5β100.0) 33.3 (0.0β80.0) 0.160 52.3 (50.3β54.4) 19.4 (18.0β21.2) 30.0 (28.7β31.5) <0.000001 INJLA 73.8 (63.8β83.3) 0.3 (0.1β0.9) 0.0 (0.0β0.0) 0.005 54.9 (52.8β57.0) 20.0 (18.6β21.9) 30.0 (28.2β32.0) 0.000020 PMI 89.8 (86.3β93.0) 0.6 (0.3β2.1) 0.0 (0.0β0.0) 0.208 46.5 (44.5β48.4) 16.2 (15.1β17.5) 30.0 (28.7β31.5) <0.000001 STE_ 94.0 (84.4β99.4) 3.7 (0.3β15.4) 0.4 (0.1β1.0) <0.001 45.8 (43.9β47.8) 16.2 (15.0β17.5) 30.0 (28.7β31.5) <0.000001 SEHYP 100.0 (99.8β100.0) 83.3 (25.0β100.0) 28.6 (9.5β54.5) 0.019 49.2 (47.2β51.3) 18.5 (17.1β20.2) 30.0 (28.7β31.5) <0.000001 SBRAD 96.3 (93.8β98.1) 58.7 (45.3β70.3) 60.3 (49.5β70.1) 0.236 41.7 (39.9β43.7) 14.4 (13.4β15.6) 30.0 (28.6β31.4) <0.000001 RAO/RAE 97.3 (95.1β99.1) 41.4 (7.8β70.4) 27.8 (7.1β46.7) 0.094 49.6 (47.5β51.6) 19.3 (17.7β21.3) 30.0 (28.7β31.5) <0.000001 VCLVH 86.4 (82.6β90.1) 29.1 (21.3β40.2) 34.3 (27.5β41.9) 0.122 42.5 (40.4β44.6) 15.3 (14.3β16.7) 30.0 (28.7β31.5) 0.000012 IRBBB 98.3 (97.6β99.0) 78.8 (71.3β85.2) 61.4 (52.4β69.1) 0.775 47.7 (45.8β49.6) 16.1 (15.1β17.3) 30.0 (28.7β31.5) <0.000001 QWAVE 87.5 (82.8β91.5) 21.2 (13.0β32.2) 24.5 (15.4β33.8) 0.165 54.2 (52.1β56.4) 22.0 (20.1β24.3) 30.0 (28.7β31.5) 0.000019 NT_ 95.0 (91.9β97.3) 33.1 (22.7β49.9) 35.4 (20.0β48.7) 0.276 41.6 (39.5β43.7) 14.7 (13.6β15.9) 30.0 (28.7β31.5) <0.000001 EL 93.9 (88.5β97.6) 5.4 (2.1β13.6) 9.8 (0.0β21.7) 0.095 46.4 (44.5β48.4) 16.4 (15.1β17.7) 30.0 (28.6β31.5) 0.000001 HVOLT 89.2 (73.8β98.3) 4.6 (0.7β17.1) 6.3 (1.6β12.6) 0.010 32.6 (30.7β34.5) 12.5 (11.7β13.5) 30.0 (28.7β31.5) <0.000001 RVH 95.1 (89.2β99.1) 26.2 (8.7β52.5) 28.6 (7.4β50.0) 0.415 50.9 (48.9β52.9) 17.8 (16.6β19.4) 30.0 (28.7β31.5) <0.000001 LVOLT 90.9 (83.7β96.2) 9.6 (4.0β21.1) 19.0 (6.2β32.1) 0.127 48.6 (46.5β50.6) 17.2 (15.8β18.8) 30.0 (28.7β31.5) <0.000001 SARRH 97.5 (96.5β98.3) 63.4 (52.1β73.8) 56.8 (48.0β65.0) 0.377 48.4 (46.4β50.4) 17.6 (16.2β18.9) 29.8 (28.3β31.2) 0.000012 IPLMI 82.0 (58.8β98.9) 4.4 (0.2β23.3) 0.0 (0.0β0.0) 0.128 52.2 (50.2β54.2) 19.6 (18.0β21.5) 29.7 (28.3β31.3) 0.000006 TAB_ 89.9 (79.0β96.7) 1.1 (0.2β3.9) 0.0 (0.0β0.0) 0.036 50.7 (48.8β52.6) 17.9 (16.5β19.4) 29.5 (28.1β31.0) 0.000014 Continued on next page 34 Table C.1 (continued) Predictor Traditional ECG ModelLEF Detection (One Predictor) AUROC AUPRC F1 Score ThreshAUROC AUPRC F1 Score Thresh Predictor performance was obtained from the traditional automatic AI-ECG model. LEF detec- tion performance was derived using the single-predictor approach proposed in this study. AUROC, AUPRC, and F1 are reported with 95% confidence intervals. The two Thresh columns indicate the thresholds used to maximize the F1 score on the validation set for predictor performance and LEF detection, respectively. Predictor abbreviations are defined inWagner et al.(2020). Table C.2:Performance of LEF detection using the single-predictor approach on the external test cohort (ECG-Note) for all 71 ECG predictors. Predictor LEF Detection (One Predictor) AUROC AUPRC F1 Score Thresh NORM 78.8 (77.8-79.6) 36.2 (34.5-38.1) 42.3 (41.0-43.5) 0.003641 ILBBB 71.6 (70.5-72.6) 33.2 (31.5-35.0) 38.2 (36.8-39.6) 0.000349 INJAL 73.4 (72.4-74.3) 27.9 (26.6-29.2) 40.9 (39.6-42.1) 0.000386 ISCLA 67.7 (66.6-68.8) 26.7 (25.4-28.4) 34.0 (32.6-35.2) 0.001918 ANEUR 71.0 (70.0-72.0) 32.1 (30.4-33.8) 38.0 (36.7-39.4) 0.001924 ISCAL 65.7 (64.7-66.8) 21.8 (20.8-22.9) 33.2 (32.0-34.3) 0.001148 ASMI 70.7 (69.6-71.7) 29.3 (27.8-31.0) 37.1 (35.8-38.3) 0.023987 SVARR 64.0 (62.8-65.1) 22.8 (21.7-24.0) 32.1 (30.8-33.4) 0.000218 INJIL 66.7 (65.6-67.7) 22.9 (21.8-24.2) 34.7 (33.5-35.9) 0.000035 CRBBB 66.9 (65.8-67.9) 23.4 (22.2-24.7) 33.2 (32.1-34.3) 0.000005 LAFB 67.5 (66.5-68.5) 22.9 (21.9-24.0) 34.9 (33.7-36.1) 0.001226 ALMI 66.7 (65.5-67.8) 29.7 (28.0-31.4) 34.2 (32.8-35.6) 0.002106 ABQRS 64.4 (63.3-65.5) 24.2 (22.9-25.6) 32.8 (31.6-33.9) 0.012123 CLBBB 69.3 (68.1-70.4) 29.7 (28.1-31.3) 36.0 (34.8-37.2) 0.000006 Continued on next page 35 Table C.2 (continued) Predictor LEF Detection (One Predictor) AUROC AUPRC F1 Score Thresh ILMI 65.8 (64.7-66.9) 25.9 (24.5-27.3) 33.0 (31.8-34.2) 0.000172 INJAS 65.1 (64.0-66.2) 21.9 (20.9-23.1) 33.9 (32.7-35.1) 0.000571 INVT 62.5 (61.1-63.7) 22.9 (21.7-24.2) 31.5 (30.3-32.6) 0.005779 PVC 59.5 (58.1-60.7) 22.6 (21.2-24.0) 29.5 (28.2-30.6) 0.001254 ISCIL 59.2 (58.1-60.3) 18.7 (17.9-19.8) 29.9 (28.9-30.9) 0.000142 1AVB 61.5 (60.3-62.6) 21.5 (20.5-22.7) 30.5 (29.5-31.6) 0.000059 ISC_ 59.6 (58.2-60.8) 22.0 (20.8-23.3) 29.5 (28.3-30.6) 0.007637 IVCD 64.3 (63.0-65.5) 28.7 (27.0-30.5) 33.2 (32.0-34.6) 0.027054 LAO/LAE 62.5 (61.0-63.7) 28.6 (26.7-30.3) 32.9 (31.4-34.3) 0.051849 ISCAN 56.8 (55.5-57.9) 18.3 (17.4-19.2) 27.4 (26.2-28.5) 0.000312 ISCAS 54.8 (53.5-55.9) 16.5 (15.7-17.2) 28.3 (27.3-29.3) 0.000066 AFIB 57.8 (56.6-58.9) 18.3 (17.5-19.2) 29.0 (28.0-30.0) 0.000004 BIGU 58.9 (57.7-60.0) 19.1 (18.2-20.2) 29.5 (28.5-30.6) 0.000013 SVTAC 52.7 (51.5-54.0) 16.7 (15.9-17.6) 24.2 (23.0-25.4) 0.000113 AMI 54.0 (52.9-55.2) 16.7 (15.9-17.5) 28.0 (27.0-29.0) 0.000868 NST_ 46.4 (45.2-47.4) 13.7 (13.1-14.3) 26.1 (25.2-27.1) 0.005043 3AVB 57.7 (56.6-58.9) 19.0 (18.1-20.0) 28.3 (27.4-29.3) <0.000001 IMI 58.9 (57.8-60.0) 19.7 (18.8-20.8) 29.8 (28.7-30.7) 0.003210 LPR 56.4 (55.1-57.5) 17.6 (16.8-18.5) 28.7 (27.7-29.7) 0.000049 2AVB 56.6 (55.4-57.9) 17.6 (16.7-18.6) 28.8 (27.7-29.9) 0.000002 Continued on next page 36 Table C.2 (continued) Predictor LEF Detection (One Predictor) AUROC AUPRC F1 Score Thresh DIG 53.9 (52.7-55.1) 16.2 (15.5-17.0) 28.2 (27.2-29.2) 0.000024 LMI 59.0 (57.8-60.2) 20.2 (19.2-21.4) 29.4 (28.3-30.4) 0.000937 LOWT 49.7 (48.5-50.9) 15.2 (14.5-16.0) 26.3 (25.3-27.2) 0.000556 SR 57.1 (55.9-58.3) 19.1 (18.1-20.2) 29.1 (28.0-30.2) 0.943848 STACH 54.5 (53.3-55.6) 16.4 (15.6-17.2) 28.6 (27.6-29.5) 0.000001 LPFB 58.9 (57.7-60.2) 19.4 (18.5-20.5) 29.4 (28.4-30.5) 0.000075 PACE 57.9 (56.7-59.1) 22.3 (21.0-23.8) 28.3 (27.3-29.2) 0.000033 WPW 55.8 (54.5-56.9) 18.6 (17.6-19.7) 26.5 (25.3-27.7) 0.000028 ISCIN 51.3 (50.1-52.5) 15.6 (14.9-16.4) 25.5 (24.4-26.5) 0.000432 PRC(S) 53.2 (51.9-54.4) 16.6 (15.8-17.5) 26.0 (25.0-27.0) 0.000008 AFLT 48.4 (47.1-49.6) 14.9 (14.1-15.6) 25.1 (24.0-26.0) 0.000010 INJIN 56.6 (55.3-57.8) 18.4 (17.5-19.5) 27.8 (26.7-29.0) 0.000062 PAC 53.9 (52.6-55.0) 17.0 (16.2-18.0) 27.0 (26.0-28.0) 0.000290 IPMI 51.1 (49.8-52.4) 17.1 (16.2-18.3) 26.3 (25.4-27.3) 0.000016 STD_ 53.6 (52.4-54.9) 17.5 (16.7-18.5) 26.6 (25.7-27.5) 0.004936 LNGQT 44.8 (43.5-46.1) 13.9 (13.2-14.6) 26.0 (25.1-26.8) 0.000042 TRIGU 51.2 (49.9-52.6) 16.2 (15.5-17.1) 25.8 (24.8-26.7) 0.000002 NDT 32.4 (31.2-33.6) 11.0 (10.6-11.5) 27.1 (26.2-27.9) <0.000001 LVH 50.8 (49.4-52.2) 18.2 (17.1-19.4) 27.1 (26.3-27.9) 0.000005 PSVT 36.6 (35.4-37.9) 12.0 (11.5-12.6) 27.2 (26.3-28.0) <0.000001 Continued on next page 37 Table C.2 (continued) Predictor LEF Detection (One Predictor) AUROC AUPRC F1 Score Thresh INJLA 52.1 (50.7-53.4) 17.0 (16.1-17.9) 25.2 (24.0-26.2) 0.000020 PMI 37.2 (36.0-38.3) 12.0 (11.4-12.5) 27.2 (26.3-28.0) <0.000001 STE_ 44.9 (43.7-46.2) 14.0 (13.4-14.7) 27.2 (26.3-28.0) <0.000001 SEHYP 51.4 (50.3-52.6) 16.3 (15.5-17.1) 27.1 (26.3-27.9) <0.000001 SBRAD 49.0 (47.8-50.1) 14.7 (14.0-15.3) 27.2 (26.4-28.1) <0.000001 RAO/RAE 48.1 (46.9-49.4) 16.3 (15.4-17.3) 27.2 (26.3-28.0) <0.000001 VCLVH 38.8 (37.5-40.1) 12.4 (11.9-13.0) 27.0 (26.1-27.8) 0.000012 IRBBB 41.8 (40.6-43.1) 12.8 (12.2-13.4) 27.1 (26.3-27.9) <0.000001 QWAVE 52.8 (51.5-54.0) 18.7 (17.7-19.8) 27.1 (26.3-28.0) 0.000018 NT_ 35.4 (34.1-36.5) 11.6 (11.2-12.2) 27.2 (26.3-28.0) <0.000001 EL 39.3 (38.0-40.6) 12.5 (12.0-13.2) 27.1 (26.2-27.9) 0.000001 HVOLT 34.0 (32.8-35.1) 11.2 (10.7-11.7) 27.2 (26.3-28.0) <0.000001 RVH 46.7 (45.5-48.0) 14.1 (13.5-14.8) 27.2 (26.3-28.0) <0.000001 LVOLT 46.6 (45.4-47.8) 14.5 (13.8-15.2) 27.2 (26.3-28.0) <0.000001 SARRH 50.8 (49.5-52.1) 16.3 (15.5-17.2) 27.2 (26.3-28.0) 0.000012 IPLMI 51.1 (49.8-52.3) 16.6 (15.7-17.7) 27.2 (26.3-28.1) 0.000006 TAB_ 49.1 (47.9-50.3) 15.4 (14.6-16.2) 26.6 (25.7-27.5) 0.000014 LEF detection performance was derived using the single-predictor approach proposed in this study. AUROC, AUPRC, and F1 are reported with 95% confidence intervals. The Thresh column indicates the thresholds used to maximize the F1 score on the validation set for predictor performance and LEF detection, respectively. Predictor abbreviations are defined in Wagner et al. (2020). 38 C.3 Threshold selection for the single-predictor approach Following (Ribeiro et al.,2020), we adopt the F1 score as a primary performance metric due to its robustness to class imbalance. For the single-predictor method, the classification threshold is selected to maximize the F1 score on the validation set, consistent with prior studies (Diao et al.,2025;Ribeiro et al.,2020). As an alternative operating point, we also consider a recall-based threshold that enforces a minimum recall of 90%. Notably, under both thresholding strategies, the thresholds used by the single-predictor method for detecting LEF are substantially lower than the corresponding diagnosis thresholds employed in traditional automatic ECG diagnosis models. This trend is observed for most predictors with AUROC greater than 0.60, as summarized in FigureC.1. 00000 0 0 (.&/(-*%2"*1& .&%)$0-. " )"',-/)/0(.&/(-*% +"30(.&/(-*% &$"**#"/&%0(.&/(-*% (.&/(-*%2"*1& ! ! .&%)$0-. # )"',-/)/0(.&/(-*% +"30(.&/(-*% &$"**#"/&%0(.&/(-*% Figure C.1: Operating thresholds of individual predictors for LEF detection. a, Left column: thresholds for the first half of predictors (AUROCβ₯60%).b, Right column: thresholds for the remaining predictors. For each predictor, three types of thresholds are shown: 1)Diagnosis threshold: threshold for determining whether the ECG shows a positive signal for that predictor; 2) F1-max threshold: threshold for predicting LEF positivity using that predictor, chosen to maximize F1 score on the validation set; 3)Recall-based threshold: threshold for predicting LEF positivity using that predictor, set to achieve recallβ₯90%. Only predictors with AUROCβ₯60% are shown for clarity; complete results for all predictors are provided in Tables2andC.1. C.4 Subgroups analysis for single-predictor approach The subgroup analysis results for NORM, ILBBB, INJIL, and ISCLA in the internal test set are presented in TablesC.3andC.4. Corresponding results in the external test set are shown in Tables C.5andC.6. 39 Table C.3:Subgroup performance of the single-predictor approach for the NORM and ILBBB predictors on the internal test set. NORMILBBB Subgroup n Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59212413.6 80.9 (78.2β83.8) 41.3 (36.0β48.3) 47.1 (42.6β51.6) 79.3 (76.2β82.1) 44.1 (38.7β49.9) 44.4 (39.6β49.3) 60β69131816.6 81.6 (78.3β84.8) 49.3 (42.2β56.6) 50.4 (45.5β55.0) 78.1 (74.7β81.1) 44.8 (37.9β51.4) 45.8 (40.1β51.0) 70β79115421.7 80.6 (77.6β83.5) 53.6 (47.7β60.5) 55.4 (51.0β59.9) 81.1 (78.0β84.0) 55.7 (49.5β61.8) 58.5 (53.4β62.9) 80+84624.1 77.0 (73.4β80.2) 47.9 (41.9β54.8) 52.8 (48.2β57.1) 77.8 (74.5β81.3) 50.7 (44.8β58.2) 53.8 (48.6β58.9) Sex Female 273112.4 81.2 (78.6β83.6) 39.0 (34.2β44.9) 43.9 (40.4β47.5) 81.9 (79.5β84.1) 43.6 (38.0β49.4) 45.7 (41.3β49.9) Male271123.0 80.2 (78.4β82.0) 53.5 (49.4β57.8) 56.4 (53.6β59.4) 77.5 (75.4β79.6) 52.3 (48.2β56.6) 53.5 (50.5β56.9) Race/ethnicity Hispanic 164916.7 83.3 (80.6β85.9) 52.6 (46.5β58.8) 53.9 (50.0β58.2) 81.1 (78.3β83.7) 50.3 (44.5β56.4) 50.0 (45.0β54.6) White156916.3 81.7 (79.1β84.4) 45.0 (39.1β51.7) 48.3 (43.9β52.4) 79.4 (76.4β82.2) 45.3 (39.1β52.0) 48.9 (43.5β53.6) Black84619.3 78.1 (74.4β81.8) 47.2 (39.9β56.2) 50.2 (44.5β55.9) 79.8 (75.8β83.3) 50.0 (42.4β58.8) 50.7 (44.2β56.6) Asian15316.3 82.7 (74.6β90.2) 50.4 (32.3β69.7) 45.9 (31.0β58.7) 88.1 (81.0β94.2) 65.5 (48.4β81.2) 57.6 (40.8β71.0) Other45715.8 81.4 (75.8β86.6) 42.1 (32.3β54.0) 50.9 (42.9β58.9) 77.2 (70.9β83.4) 50.4 (39.3β62.1) 50.0 (41.2β58.6) Unknown 76822.3 77.7 (73.6β81.7) 47.7 (40.4β56.3) 54.7 (49.0β60.0) 78.8 (74.8β82.5) 49.5 (42.0β57.6) 53.2 (46.5β59.1) Clinical context Emergency 197115.7 79.9 (77.2β82.4) 40.8 (35.8β46.3) 47.9 (43.9β51.7) 82.0 (79.5β84.2) 48.3 (42.7β53.9) 49.9 (45.2β54.0) Inpatient 220324.4 79.0 (76.7β81.1) 54.9 (50.2β59.7) 56.4 (53.3β59.5) 75.8 (73.5β78.1) 52.5 (48.2β57.0) 52.4 (49.1β55.9) Outpatient 10596.6 83.8 (78.1β88.8) 34.6 (25.1β45.4) 35.7 (28.8β42.7) 82.0 (76.3β87.2) 31.5 (21.8β43.7) 40.9 (32.0β49.3) Procedural 20921.5 81.1 (74.6β87.0) 46.3 (34.1β61.7) 55.2 (42.7β65.2) 75.3 (65.9β83.1) 44.9 (31.8β60.9) 50.0 (35.9β61.0) AUROC, AUPRC, and F1 are reported with 95% confidence intervals. NORM, normal ECG; ILBBB, incomplete left bundle branch block. 40 Table C.4:Subgroup performance of the single-predictor method for INJIL and ISCLA in the internal test set. INJILISCLA Subgroup n Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59212413.6 77.4 (74.6β80.2) 32.4 (28.4β37.9) 40.8 (36.6β45.1) 79.0 (76.4β81.8) 36.6 (31.8β43.0) 43.3 (39.3β47.4) 60β69131816.6 75.4 (72.0β78.8) 34.0 (29.0β40.1) 45.5 (40.6β50.2) 74.4 (70.9β78.0) 37.0 (30.8β44.3) 43.0 (38.1β48.0) 70β79115421.7 73.9 (70.6β77.1) 39.5 (34.5β45.6) 48.0 (43.1β52.7) 73.9 (70.6β77.2) 42.5 (36.9β49.3) 45.8 (40.9β50.1) 80+84624.1 69.2 (65.1β72.9) 36.3 (31.2β42.8) 47.5 (42.3β52.5) 70.8 (67.2β74.5) 43.3 (37.2β50.1) 43.8 (38.5β48.8) Sex Female 273112.4 75.2 (72.7β77.7) 26.1 (22.8β30.3) 36.2 (32.8β40.0) 76.8 (74.4β79.3) 29.5 (25.6β34.7) 37.7 (34.1β41.6) Male271123.0 74.3 (72.2β76.3) 41.2 (37.8β45.1) 51.1 (48.2β54.0) 75.3 (73.2β77.3) 46.2 (42.6β50.4) 48.5 (45.6β51.5) Race/ethnicity Hispanic 164916.7 76.1 (73.0β79.2) 34.6 (30.1β40.3) 45.1 (40.8β49.6) 75.5 (72.6β78.4) 39.0 (33.1β45.3) 42.0 (37.5β46.3) White156916.3 76.2 (73.2β79.3) 35.0 (30.2β41.0) 43.8 (39.6β48.3) 76.8 (74.0β79.7) 35.9 (31.1β42.1) 42.1 (37.5β46.7) Black84619.3 72.0 (68.0β76.1) 32.6 (27.5β39.6) 44.9 (39.1β50.8) 76.1 (72.3β80.0) 41.5 (34.9β49.8) 46.6 (41.2β51.8) Asian15316.3 78.1 (67.1β87.2) 47.8 (30.3β64.3) 45.0 (29.6β56.8) 78.0 (67.1β87.5) 41.7 (26.7β63.5) 45.9 (29.8β59.0) Other45715.8 76.1 (70.7β81.7) 35.6 (27.4β46.6) 41.6 (32.9β50.0) 75.6 (70.0β80.8) 37.1 (27.4β48.4) 42.3 (33.7β50.0) Unknown 76822.3 74.0 (69.8β77.8) 39.0 (33.2β46.7) 48.7 (43.2β54.2) 74.3 (70.3β78.5) 45.0 (37.9β53.2) 47.9 (42.2β53.5) Clinical context Emergency 197115.7 76.6 (73.9β79.2) 31.9 (28.1β36.6) 43.7 (39.6β47.5) 77.6 (75.0β79.8) 38.6 (33.6β44.3) 43.2 (39.3β47.1) Inpatient 220324.4 70.8 (68.3β73.2) 39.3 (35.7β43.4) 48.0 (44.7β51.4) 70.4 (67.9β72.8) 41.7 (37.5β46.2) 46.9 (43.4β50.3) Outpatient 10596.6 75.4 (69.5β81.5) 20.2 (14.4β29.2) 29.2 (21.7β36.2) 81.4 (75.7β86.3) 32.4 (23.2β45.3) 31.1 (23.9β38.3) Procedural 20921.5 77.2 (69.8β84.8) 44.0 (32.0β60.8) 55.4 (43.7β66.7) 76.6 (69.6β83.4) 43.8 (30.8β60.5) 42.3 (29.1β53.9) AUROC, AUPRC, and F1 are reported with 95% confidence intervals. INJIL, subendocardial injury in inferolateral leads; ISCLA, ischemic in lateral leads. 41 Table C.5:Subgroup performance of the single-predictor method for NORM and ILBBB in the external test set. NORMILBBB Subgroupn Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59427012.2 81.5 (79.4β83.3) 37.1 (33.1β41.5) 44.6 (41.6β47.7) 73.7 (71.3β76.0) 33.7 (30.3β38.0) 40.4 (37.7β43.2) 60β69363715.7 78.5 (76.5β80.3) 40.8 (36.9β45.3) 46.2 (43.4β48.9) 71.6 (69.3β73.9) 37.6 (33.9β41.8) 43.4 (40.5β46.1) 70β79376118.0 77.7 (76.1β79.5) 39.3 (36.1β43.3) 43.7 (41.3β46.0) 70.5 (68.4β72.7) 36.0 (33.1β39.7) 39.9 (37.4β42.5) 80+434917.2 74.9 (73.1β76.7) 34.2 (31.3β37.6) 38.4 (36.4β40.4) 67.0 (64.9β69.2) 29.3 (26.8β32.4) 36.1 (33.9β38.5) Sex Female722211.2 74.2 (72.4β75.9) 22.6 (19.9β25.8) 32.4 (30.4β34.5) 65.3 (63.0β67.6) 19.5 (17.1β22.2) 29.8 (27.4β32.0) Male879519.4 77.1 (75.9β78.4) 42.2 (39.6β44.9) 45.4 (43.6β47.2) 71.6 (70.0β73.0) 39.6 (36.9β42.3) 43.1 (41.2β45.0) Race/ethnicity Hispanic63814.7 81.7 (77.2β85.8) 39.0 (30.4β48.8) 45.8 (39.3β52.6) 68.6 (63.1β74.1) 30.3 (23.0β39.1) 39.1 (31.8β46.2) White1168815.5 77.0 (75.9β77.9) 35.2 (33.2β37.5) 41.7 (40.4β43.1) 69.4 (68.2β70.6) 31.1 (29.0β33.4) 39.4 (38.0β40.9) Black184517.3 79.6 (77.2β81.9) 44.3 (39.1β50.2) 46.2 (42.4β49.8) 72.6 (69.8β75.4) 40.7 (35.7β46.4) 43.4 (39.6β47.2) Asian41412.6 74.7 (69.4β80.0) 26.1 (18.4β36.6) 37.2 (28.5β45.1) 69.0 (63.0β75.2) 24.2 (17.4β33.7) 35.0 (27.1β43.4) Other52413.0 73.7 (68.1β79.1) 22.5 (15.9β32.1) 34.4 (27.5β41.4) 66.2 (60.3β72.5) 21.2 (15.3β29.9) 32.9 (26.2β39.6) Unknown 90818.8 75.5 (71.4β79.5) 39.5 (33.1β46.7) 41.8 (36.8β46.7) 66.2 (61.8β70.6) 34.9 (29.7β41.1) 38.0 (33.1β42.7) Clinical context Emergency 917014.6 76.3 (75.1β77.5) 33.6 (31.4β36.0) 40.6 (39.1β42.2) 69.2 (67.7β70.6) 31.1 (28.9β33.7) 39.2 (37.7β40.7) Urgent302820.9 78.7 (76.9β80.5) 45.4 (41.5β49.2) 49.0 (46.5β51.6) 71.2 (69.2β73.2) 40.8 (36.9β45.0) 44.9 (42.2β47.5) Observation 217319.6 79.6 (77.7β81.7) 49.0 (44.0β54.2) 51.3 (48.4β54.1) 73.9 (71.7β76.0) 44.9 (40.0β50.0) 48.6 (45.5β51.5) Surgical Same Day 10226.5 70.5 (65.5β75.5) 13.2 (9.2β19.6) 21.7 (16.5β27.2) 67.3 (61.8β72.9) 12.2 (8.4β18.2) 21.2 (16.3β26.5) Elective6249.3 73.9 (68.7β79.2) 19.1 (13.6β28.4) 28.9 (22.7β35.0) 67.0 (60.9β73.0) 16.0 (11.6β24.0) 26.0 (20.7β31.7) AUROC, AUPRC, and F1 are reported with 95% confidence intervals. NORM, normal ECG; ILBBB, incomplete left bundle branch block. 42 Table C.6:Subgroup performance of the single-predictor method for INJIL and ISCLA in the internal test set. INJILISCLA Subgroupn Prevalence (%)AUROC AUPRC F1 ScoreAUROC AUPRC F1 Score Age groups 18β59427012.2 73.1 (71.0β75.1) 25.7 (22.4β29.6) 42.2 (39.1β45.5) 61.0 (58.6β63.4) 23.2 (20.3β26.6) 32.4 (30.0β34.8) 60β69363715.7 72.4 (70.4β74.5) 28.9 (25.6β32.8) 45.0 (42.4β47.6) 61.3 (59.0β63.6) 26.2 (23.4β29.6) 35.6 (33.2β38.1) 70β79376118.0 71.6 (69.8β73.5) 29.4 (26.9β32.4) 43.5 (41.1β45.9) 62.7 (60.6β64.9) 25.9 (23.6β29.1) 33.2 (30.7β35.5) 80+434917.2 69.9 (68.0β71.7) 27.3 (25.1β30.2) 39.2 (36.9β41.3) 62.2 (60.2β64.4) 24.7 (22.7β27.4) 32.0 (29.8β34.3) Sex Female722211.2 67.1 (65.4β68.8) 18.7 (16.4β21.5) 31.6 (29.7β33.6) 58.1 (55.7β60.5) 16.3 (14.2β18.9) 26.3 (23.9β28.7) Male879519.4 73.8 (72.6β75.1) 34.5 (32.3β36.9) 44.5 (42.7β46.3) 65.6 (64.0β67.1) 33.0 (30.6β35.5) 39.5 (37.7β41.4) Race/ethnicity Hispanic63814.7 75.8 (71.3β80.1) 30.7 (23.9β39.2) 45.9 (39.5β52.2) 62.3 (56.4β68.2) 27.1 (20.3β35.8) 34.8 (27.6β42.2) White1168815.5 71.7 (70.7β72.7) 29.8 (28.0β31.8) 41.7 (40.4β43.0) 62.0 (60.8β63.1) 27.3 (25.5β29.2) 34.9 (33.5β36.3) Black184517.3 74.6 (72.3β76.8) 37.0 (32.6β41.7) 45.1 (41.5β48.8) 64.5 (61.6β67.5) 34.2 (29.9β38.8) 38.8 (35.0β42.7) Asian41412.6 69.2 (63.6β74.6) 20.1 (14.2β28.8) 36.7 (28.0β45.5) 62.9 (56.8β69.4) 20.4 (14.5β29.2) 32.8 (25.1β40.9) Other52413.0 68.8 (63.0β74.7) 17.7 (12.6β25.4) 34.5 (27.5β41.7) 61.4 (55.4β68.0) 16.8 (12.3β24.3) 30.8 (23.7β38.1) Unknown 90818.8 69.5 (65.2β73.5) 31.5 (26.4β37.2) 40.8 (35.9β45.7) 60.7 (56.0β65.2) 28.5 (24.2β33.5) 35.9 (30.9β40.7) Clinical context Emergency 917014.6 71.6 (70.3β72.8) 28.2 (26.4β30.2) 40.6 (39.1β42.2) 62.2 (60.8β63.6) 26.4 (24.6β28.3) 35.8 (34.3β37.3) Urgent302820.9 74.6 (72.9β76.3) 39.2 (35.8β42.9) 49.0 (46.5β51.6) 64.2 (62.2β66.3) 36.3 (33.0β39.8) 44.9 (42.2β47.5) Observation 217319.6 74.7 (72.8β76.5) 42.6 (38.2β47.1) 51.3 (48.4β54.1) 65.4 (63.3β67.7) 39.7 (35.3β44.1) 48.6 (45.5β51.5) Surgical Same Day 10226.5 66.9 (61.9β72.0) 12.3 (8.6β18.1) 21.7 (16.5β27.2) 61.5 (56.0β67.1) 12.0 (8.4β17.8) 21.2 (16.3β26.5) Elective6249.3 70.9 (66.0β75.9) 18.1 (12.9β26.8) 28.9 (22.7β35.0) 62.6 (56.6β68.7) 16.9 (12.1β25.2) 26.0 (20.7β31.7) AUROC, AUPRC, and F1 are reported with 95% confidence intervals. INJIL, subendocardial injury in inferolateral leads; ISCLA, ischemic in lateral leads. 43 References Attia, Z. I., Kapa, S., Lopez-Jimenez, F., McKie, P. M., Ladewig, D. J., Satam, G., Pellikka, P. A., Enriquez-Sarano, M., Noseworthy, P. A., Munger, T. M. et al. (2019a) Screening for cardiac contractile dysfunction using an artificial intelligenceβenabled electrocardiogram. Nature medicine,25, 70β74. Attia, Z. I. et al. (2019b) Prospective validation of a deep learning electrocardiogram algo- rithm for the detection of left ventricular systolic dysfunction.J. Cardiovasc. Electr.,30, 668β674. Carter, R. E., Johnson, P. W., Strom, J. B., Waks, J. W., Krumerman, A., Ferrick, K. J., DeRaad, R., Steinberg, B. A., Wieczorek, M. A., Cruz, J. et al. (2026) Multisite, external validation of an ai-enabled ecg algorithm for detection of low ejection fraction.JACC: Advances,5, 102537. Chen, T. (2016) Xgboost: A scalable tree boosting system.Cornell University. Cardiovascular Committee of China Medical Womens Association, Asia Heart Rhythm Soci- ety, T. C. o. R. C. M. o. C. A. f. P., Medical Devices Technology Exchange, C. E. C. T. F. o. t. S. f. C. and of the Primary Diagnostic Terminology of Electrocardiogram, E. S. (2023) Chinese expert consensus statement on the standardized Chinese and English primary diagnostic terminology of electrocardiogram.Chin. Circulation J.,38, 141β145. Ciampi, Q. and Villari, B. (2007) Role of echocardiography in diagnosis and risk stratification in heart failure with left ventricular systolic dysfunction.Cardiovascular ultrasound,5, 34. Diao, X., Xu, W., Cheng, H., Zhou, Y., Liu, Y., Huo, Y., Lu, J., Huang, J., He, J., Liu, F. et al. (2025) Speed-tr: a self-distilled and pre-trained transformer model for enhanced ecg detection of tricuspid regurgitation.npj Digital Medicine,8, 650. Elias, P. and Finer, J. (2025) Echonext: A dataset for detecting echocardiogram-confirmed structural heart disease from ecgs. Gao, Z., Yurk, D. and Abu-Mostafa, Y. S. (2025) Machine learning with scarce data: Ejection fraction prediction using PLAX view. InMedical Imaging with Deep Learning. URL: https://openreview.net/forum?id=JEN5FzeFZj. Gow, B., Pollard, T., Nathanson, L. A., Johnson, A., Moody, B., Fernandes, C., Greenbaum, N., Waks, J. W., Eslami, P., Carbonati, T., Chaudhari, A., Herbst, E., Moukheiber, D., Berkowitz, S., Mark, R. and Horng, S. (2023) MIMIC-IV-ECG: Diagnostic Electrocardio- gram Matched Subset.PhysioNet. URL:https://doi.org/10.13026/4nqg-sb35. Version 1.0. Hannun, A. Y., Rajpurkar, P., Haghpanahi, M., Tison, G. H., Bourn, C., Turakhia, M. P. and Ng, A. Y. (2019) Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network.Nature medicine,25, 65β69. Heidenreich, P. A., Bozkurt, B., Aguilar, D., Allen, L. A., Byun, J. J., Colvin, M. M., Deswal, A., Drazner, M. H., Dunlay, S. M., Evers, L. R. et al. (2022) 2022 aha/acc/hfsa guideline 44 for the management of heart failure: executive summary: a report of the american college of cardiology/american heart association joint committee on clinical practice guidelines. Journal of the American College of Cardiology,79, 1757β1780. Hughes, J. W., Somani, S., Elias, P., Tooley, J., Rogers, A. J., Poterucha, T., Haggerty, C. M., Salerno, M., Ouyang, D., Ashley, E. et al. (2024) Simple models vs. deep learning in detecting low ejection fraction from the electrocardiogram.European Heart Journal-Digital Health,5, 427β434. ISO Central Secretary. (2009) Health informatics β Standard communication protocol β Part 91064: Computer-assisted electrocardiography. Jiang, A., Huang, C., Cao, Q., Xu, Y., Zeng, Z., Chen, K., Zhang, Y. and Wang, Y. (2024) Self-supervised anomaly detection pretraining enhances long-tail ecg diagnosis.arXiv preprint arXiv:2408.17154. Joglar, J. A., Chung, M. K., Armbruster, A. L., Benjamin, E. J., Chyou, J. Y., Cronin, E. M., Deswal, A., Eckhardt, L. L., Goldberger, Z. D., Gopinathannair, R. et al. (2024) 2023 acc/aha/accp/hrs guideline for the diagnosis and management of atrial fibrillation: a report of the american college of cardiology/american heart association joint committee on clinical practice guidelines.Journal of the American College of Cardiology,83, 109β279. Johnson, A., Bulgarelli, L., Pollard, T., Gow, B., Moody, B., Horng, S., Celi, L. A. and Mark, R. (2024) MIMIC-IV.PhysioNet. URL: https://doi.org/10.13026/kpb9-mt58. Version 3.1. Johnson, A., Pollard, T., Horng, S., Celi, L. A. and Mark, R. (2023a) MIMIC-IV-Note: Deidentified free-text clinical notes.PhysioNet. URL:https://doi.org/10.13026/ 1n74-ne17. Version 2.2. Johnson, A. E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B. et al. (2023b) Mimic-iv, a freely accessible electronic health record dataset.Scientific data,10, 1. Khera, R. (2024) Ai-enabled diagnosis from an electrocardiogram image: the next frontier of innovation in a century-old technology. Kligfield, P., Gettes, L. S., Bailey, J. J., Childers, R., Deal, B. J., Hancock, E. W., van Herpen, G., Kors, J. A., Macfarlane, P., Mirvis, D. M., Pahlm, O., Rautaharju, P. and Wagner, G. S. (2007) Recommendations for the standardization and interpretation of the electrocardiogram.Circulation,115, 1306β1324. URL:https://w.ahajournals.org/ doi/abs/10.1161/CIRCULATIONAHA.106.180200. Kusumoto, F. M., Schoenfeld, M. H., Barrett, C., Edgerton, J. R., Ellenbogen, K. A., Gold, M. R., Goldschlager, N. F., Hamilton, R. M., Joglar, J. A., Kim, R. J. et al. (2019) 2018 acc/aha/hrs guideline on the evaluation and management of patients with bradycardia and cardiac conduction delay: a report of the american college of cardiology/american heart association task force on clinical practice guidelines and the heart rhythm society.Journal of the American College of Cardiology,74, e51βe156. 45 Kwon, J.-M., Lee, S. Y., Jeon, K.-H., Lee, Y., Kim, K.-H., Park, J., Oh, B.-H. and Lee, M.-M. (2020) Deep learningβbased algorithm for detecting aortic stenosis using electrocar- diography.Journal of the American Heart Association,9, e014717. Lakshmanan, S. and Mbanze, I. (2023) A comparison of cardiovascular imaging practices in africa, north america, and europe: two faces of the same coin.European Heart Journal- Imaging Methods and Practice,1, qyad005. Li, J., Aguirre, A. D., Junior, V. M., Jin, J., Liu, C., Zhong, L., Sun, C., Clifford, G., Brandon Westover, M. and Hong, S. (2025) An electrocardiogram foundation model built on over 10 million recordings.NEJM AI,2, AIoa2401033. Lundberg, S. M., Erion, G. G. and Lee, S.-I. (2018) Consistent individualized feature attri- bution for tree ensembles.arXiv preprint arXiv:1802.03888. Lundberg, S. M. and Lee, S.-I. (2017) A unified approach to interpreting model predictions. Advances in neural information processing systems,30. Na, Y., Park, M., Tae, Y. and Joo, S. (2024) Guiding masked representation learning to capture spatio-temporal relationship of electrocardiogram. InThe Twelfth Interna- tional Conference on Learning Representations. URL: https://openreview.net/forum? id=WcOohbsF4H . Poterucha, T. J., Jing, L., Ricart, R. P., Adjei-Mosi, M., Finer, J., Hartzel, D., Kelsey, C., Long, A., Rocha, D., Ruhl, J. A. et al. (2025) Detecting structural heart disease from electrocardiograms using ai.Nature,644, 221β230. Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Tang, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z. and Qiu, Z. (2025) Qwen2.5 technical report. URL: https://arxiv.org/abs/2412.15115. Ribeiro, A. H., Paixao, G. M., Lima, E. M., Horta Ribeiro, M., Pinto Filho, M. M., Gomes, P. R., Oliveira, D. M., Meira Jr, W., Schon, T. B. and Ribeiro, A. L. P. (2021) Code-15%: a large scale annotated dataset of12-lead ecgs. URL:https://doi.org/10.5281/zenodo. 4916206 . Ribeiro, A. H., Ribeiro, M. H., PaixΓ£o, G. M., Oliveira, D. M., Gomes, P. R., Canazart, J. A., Ferreira, M. P., Andersson, C. R., Macfarlane, P. W., Meira Jr, W. et al. (2020) Automatic diagnosis of the 12-lead ecg using a deep neural network.Nature communications,11, 1760. Savarese, G., Becher, P. M., Lund, L. H., Seferovic, P., Rosano, G. M. and Coats, A. J. (2022) Global burden of heart failure: a comprehensive and updated review of epidemiology. Cardiovascular research,118, 3272β3287. Shapley, L. S. et al. (1953) A value for n-person games. 46 Strodthoff, N., Mehari, T., Nagel, C., Aston, P. J., Sundar, A., Graff, C., Kanters, J. K., Haverkamp, W., DΓΆssel, O., Loewe, A. et al. (2023) Ptb-xl+, a comprehensive electrocar- diographic feature dataset.Scientific data,10, 279. Strodthoff, N., Wagner, P., Schaeffter, T. and Samek, W. (2020) Deep learning for ecg anal- ysis: Benchmarks and insights from ptb-xl.IEEE journal of biomedical and health infor- matics,25, 1519β1528. Tran, H. H.-V., Thu, A., Fuertes, A., Twayana, A. R., Mahadevaiah, A., Mehta, K. A., James, M., Basta, M., Weissman, S., Frishman, W. H. et al. (2025) Electrocardiogram- based artificial intelligence for detection of low ejection fraction: A contemporary review. Cardiology in Review, 10β1097. Van De Leur, R. R., Bos, M. N., Taha, K., Sammani, A., Yeung, M. W., Van Duijvenboden, S., Lambiase, P. D., Hassink, R. J., Van Der Harst, P., Doevendans, P. A. et al. (2022) Improving explainability of deep neural network-based electrocardiogram interpretation using variational auto-encoders.European Heart Journal-Digital Health,3, 390β404. Wagner, P., Strodthoff, N., Bousseljot, R.-D., Kreiseler, D., Lunze, F. I., Samek, W. and Schaeffter, T. (2020) Ptb-xl, a large publicly available electrocardiography dataset.Scien- tific data,7, 1β15. Yao, X., Rushlow, D. R., Inselman, J. W., McCoy, R. G., Thacher, T. D., Behnken, E. M., Bernard, M. E., Rosas, S. L., Akfaly, A., Misra, A. et al. (2021) Artificial intelligenceβ enabled electrocardiograms for identification of patients with low ejection fraction: a prag- matic, randomized clinical trial. Nature medicine , 27 , 815β819. Zheng, J., Chu, H., Struppa, D., Zhang, J., Yacoub, S. M., El-Askary, H., Chang, A., Ehw- erhemuepha, L., Abudayyeh, I., Barrett, A. et al. (2020a) Optimal multi-stage arrhythmia classification approach.Scientific reports,10, 2898. Zheng, J., Zhang, J., Danioko, S., Yao, H., Guo, H. and Rakovski, C. (2020b) A 12-lead electrocardiogram database for arrhythmia research covering more than 10,000 patients. Scientific data,7, 48. Zhou, Y., Yang, Y., Fan, X. and Zhao, W. (2025) Bridging performance gaps for ecg foun- dation models: A post-training strategy.arXiv preprint arXiv:2509.12991. 47