Paper deep dive
BAAI Cardiac Agent: An intelligent multimodal agent for automated reasoning and diagnosis of cardiovascular diseases from cardiac magnetic resonance imaging
Taiping Qu, Hongkai Zhang, Lantian Zhang, Can Zhao, Nan Zhang, Hui Wang, Zhen Zhou, Mingye Zou, Kairui Bo, Pengfei Zhao, Xingxing Jin, Zixian Su, Kun Jiang, Huan Liu, Yu Du, Maozhou Wang, Ruifang Yan, Zhongyuan Wang, Tiejun Huang, Lei Xu, Henggui Zhang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/10/2026, 2:32:13 AM
Summary
The BAAI Cardiac Agent is a multimodal intelligent system designed for end-to-end interpretation of cardiac magnetic resonance (CMR) imaging. It integrates specialized expert models for automated segmentation, functional quantification, and disease diagnosis, significantly reducing interpretation time compared to traditional clinical workflows while maintaining high diagnostic accuracy (AUC > 0.93 internally).
Entities (5)
Relation Signals (3)
Segmentation Expert Model → partof → BAAI Cardiac Agent
confidence 100% · The agent integrates specialized cardiac expert models
BAAI Cardiac Agent → performs → CMR interpretation
confidence 100% · designed for end-to-end CMR interpretation
BAAI Cardiac Agent → diagnoses → Cardiovascular Disease
confidence 95% · automated reasoning and diagnosis of cardiovascular diseases
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time-consuming interpretation across multi-sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end-to-end CMR interpretation. The agent integrates specialized cardiac expert models to perform automated segmentation of cardiac structures, functional quantification, tissue characterization and disease diagnosis, and generates structured clinical reports within a unified workflow. Evaluated on CMR datasets from two hospitals (2413 patients) spanning 7-types of major cardiovascular diseases, the agent achieved an area under the receiver-operating-characteristic curve exceeding 0.93 internally and 0.81 externally. In the task of estimating left ventricular function indices, the results generated by this system for core parameters such as ejection fraction, stroke volume, and left ventricular mass are highly consistent with clinical reports, with Pearson correlation coefficients all exceeding 0.90. The agent outperformed state-of-the-art models in segmentation and diagnostic tasks, and generated clinical reports showing high concordance with expert radiologists (six readers across three experience levels). By dynamically orchestrating expert models for coordinated multimodal analysis, this agent framework enables accurate, efficient CMR interpretation and highlights its potentials for complex clinical imaging workflows. Code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.04078v1
- Canonical: https://arxiv.org/abs/2604.04078v1
Trouble viewing inline? Open PDF directly →
Full Text
115,240 characters extracted from source content.
Expand or collapse full text
BAAI CARDIAC AGENT: AN INTELLIGENT MULTIMODAL AGENT FOR AUTOMATED REASONING AND DIAGNOSIS OF CARDIOVASCULAR DISEASES FROM CARDIAC MAGNETIC RESONANCE IMAGING Taiping Qu 1* , Hongkai Zhang 2* , Lantian Zhang 1 , Can Zhao 1 , Nan Zhang 2 , Hui Wang 2 , Zhen Zhou 2 , Mingye Zou 1 , Kairui Bo 2 , Pengfei Zhao 1 , Xingxing Jin 3 , Zixian Su 1 , Kun Jiang 2 , Huan Liu 1 , Yu Du 4 , Maozhou Wang 5 , Ruifang Yan 3† , Zhongyuan Wang 1† , Tiejun Huang 1† , Lei Xu 2† , Henggui Zhang 1† 1 Beijing Academy of Artificial Intelligence, No. 150 Chengfu Road, Haidian District, Beijing 100084, China 2 Department of Radiology, Beijing Anzhen Hospital, Beijing Institute of Heart, Lung & Vascular Diseases, Capital Medical University, 2 Anzhen Road, Beijing 100029, China 3 Department of MR, the First Affiliated Hospital, Henan Medical University, 88 Jiankang Road, Weihui 453100, China 4 Department of Cardiology, Clinical Center for Coronary Heart Disease, Beijing Institute of Heart, Lung and Blood Vessel Disease, Beijing Anzhen Hospital, Capital Medical University, Beijing 100029, China 5 Department of Cardiac Surgery, Beijing Anzhen Hospital, Institute of Heart, Lung and Vascular Diseases, Capital Medical University, Beijing 100029, China * These authors contributed equally to this work. † Corresponding authors: yrf718@163.com, zhongyuan@baai.ac.cn, tjhuang@baai.ac.cn, leixu2001@hotmail.com, henggui.zhang@gmail.com ABSTRACT Cardiac magnetic resonance (CMR) is a cornerstone for diagnosing cardiovascular disease. However, it remains underutilized due to complex, time-consuming interpretation across multi-sequences, phases, quantitative measures that heavily reliant on specialized expertise. Here, we present BAAI Cardiac Agent, a multimodal intelligent system designed for end-to-end CMR interpretation. The agent integrates specialized cardiac expert models to perform automated segmentation of cardiac structures, functional quantification, tissue characterization and disease diagnosis, and generates structured clinical reports within a unified workflow. Evaluated on CMR datasets from two hospitals (2413 patients) spanning 7-types of major cardiovascular diseases, the agent achieved an area under the receiver-operating-characteristic curve exceeding 0.93 internally and 0.81 externally. In the task of estimating left ventricular function indices, the results generated by this system for core parameters such as ejection fraction, stroke volume, and left ventricular mass are highly consistent with clinical reports, with Pearson correlation coefficients all exceeding 0.90. The agent outperformed state-of- the-art models in segmentation and diagnostic tasks, and generated clinical reports showing high concordance with expert radiologists (six readers across three experience levels). By dynamically orchestrating expert models for coordinated multimodal analysis, this agent framework enables accurate, efficient CMR interpretation and highlights its potentials for complex clinical imaging workflows. Code is available at https://github.com/plantain-herb/Cardiac-Agent. Keywords large multimodal model, cardiac magnetic resonance imaging, cardiovascular disease, agent 1 Introduction Cardiovascular disease (CVD) remains the leading cause of mortality and morbidity worldwide, imposing an escalating burden on healthcare systems despite advances in prevention and treatment [1,2,3,4]. Accurate phenotyping of cardiac structure, function, and tissue pathology is central to effective diagnosis, risk stratification, and therapeutic decision-making across the spectrum of CVD. Among available imaging modalities, cardiac magnetic resonance arXiv:2604.04078v1 [eess.IV] 5 Apr 2026 Cardiac Agent imaging (CMR) is one of the most comprehensive imaging modalities for non-invasively evaluating cardiovascular diseases, providing comprehensive, multi-parametric characterization of the heart, encompassing anatomical structure, ventricular function, myocardial perfusion, and tissue composition [5,6]. As a result, CMR is a cornerstone for the assessment of cardiac morphology and function in many clinical contexts [7, 8]. However, the very richness that underpins the clinical value of CMR also presents a fundamental challenge. Con- temporary CMR examinations consist of multiple three-dimensional imaging sequences acquired across the cardiac cycle (i.e., 4D sequence: 3D+time), yielding data that are intrinsically spatiotemporal and heterogeneous. Interpreting these data requires the integration of tissue-specific structure and imaging signature, cine motion and quantitative functional indices, a process that is cognitively demanding, time-consuming, and heavily reliant on expert experience [9,10]. In routine clinical practice, comprehensive CMR interpretation often requires prolonged manual interaction, careful cross-referencing across sequences, and iterative reasoning steps, which together limit throughput and introduce variability [9,11]. Moreover, the extensive training required to achieve proficiency in CMR interpretation restricts scalability and contributes to disparities in access to high-quality cardiac imaging expertise [9, 10]. Artificial intelligence (AI) has emerged as a promising avenue to alleviate these constraints [12,4]. Deep learning approaches have demonstrated strong performance in individual CMR and other imaging modal tasks, such as cardiac segmentation [13], functional parameter quantification [14], and disease classification [12]. These task-specific models can rapidly extract objective imaging features and have achieved accuracy comparable to expert readers in controlled settings [15,16,17]. Nevertheless, such approaches remain fundamentally limited in clinical practice because they operate in isolation. CMR interpretation is not a collection of independent sub-tasks but a coordinated reasoning process in which structural, functional, and tissue-level feature information must be jointly considered to arrive at a clinically meaningful conclusion. Fragmented AI solutions fail to capture this integrative logic and therefore struggle to deliver reliable, end-to-end clinical value. Recent advances in large multimodal models (LMMs) have opened new possibilities for medical image understanding by enabling the joint processing of visual and textual information [18,19,20,21,22,23]. In radiology, LMMs have shown encouraging capabilities in tasks such as visual question answering (VQA), medical report generation (MRG) [18,19], and multi-modal reasoning [20,21,22,23,19], and image segmentation [24]. However, existing LMMs are primarily developed and evaluated on two-dimensional images or single-slice volumetric data. Extending these models to CMR is non-trivial, as CMR data encode 4D dynamical information, i.e., complex three-dimensional anatomy coupled with dynamic motion across time. Effective interpretation therefore requires not only visual recognition but also spatiotemporal reasoning and domain-specific quantitative analysis, capabilities that remain beyond the scope of current general-purpose LMMs. A promising strategy to bridge this gap is the development of intelligent agent systems [25,26,27,28,29] that combine the complementary strengths of specialized expert models and LMMs. In this paradigm, expert models provide precise, high-fidelity quantitative measurements from complex imaging data, while LMMs orchestrate task execution, integrate heterogeneous outputs, and generate clinically coherent interpretations aligned with radiological reasoning. Rather than replacing expert knowledge, such agents aim to emulate the workflow of experienced clinicians by coordinating perception, measurement, and inference in a unified framework. Despite initial successes in limited radiological settings, existing agent-based approaches have largely focused on static images and have not addressed the distinctive challenges posed by 4D spatiotemporal CMR data. Here we proposed a Cardiac Agent system (BAAI Cardiac Agent) designed to address key unmet needs in clinical CMR practice (see Fig. 1). The system has the following major features: 1) First, it supports the unified variant expert models of segmentation of short-axis (SAX) cine segmentation (SAXCS), two-chamber (2CH) cine segmentation (2CHCS), four-chamber (4CH) cine segmentation (4CHCS), SAX late gadolinium enhancement (LGE) segmentation (SAXLGES) imagings, cardiac disease screening (CDS), non-ischemic cardiomyopathy subclassification (NICMS), interpretation of multiple CMR sequences, as well as retrieval-augmented generation (RAG) and MRG, enabling synchronized reasoning across anatomical, functional, and tissue-specific information. 2) Second, it integrates quantitative measurements directly into structured, narrative reports, reducing manual effort and improving reproducibility. 3) Third, it provides interactive capabilities, allowing clinicians to query imaging findings and reasoning steps in a transparent manner. Through this design, the BAAI Cardiac Agent transforms CMR analysis from a labor-intensive, fragmented workflow into a rapid, standardized, and explainable decision-support process. We validate the BAAI Cardiac Agent on a large-scale, clinically curated CMR dataset with a cohort of 2413 patients from two hospitals, and demonstrate robust performance across segmentation, disease screening, cardiomyopathy subclassification, and quantitative functional assessment. Comprehensive evaluation shows that the agent achieves state-of-the-art (SOTA) accuracy in core imaging tasks while producing measurements and diagnostic conclusions that are highly consistent with expert radiology reports. Importantly, the system substantially reduces interpretation time 2 Cardiac Agent without compromising clinical fidelity, highlighting its potential to improve efficiency and accessibility in real-world CMR practice. Collectively, this study establishes a generalizable framework for intelligent, agent-based interpretation of spatiotemporal medical imaging and demonstrates its practical value in one of the most complex domains of radiology. By integrating expert-level quantification with multimodal reasoning, the BAAI Cardiac Agent represents a step toward scalable, trustworthy AI systems that operate in close alignment with clinical expertise and workflow. 2 Results 2.1 Patient datasets In this study, CMR scan sequences of 2413 patients were retrospectively collected from two hospitals. Detailed information on data acquisition and institutional sources is provided in the Appendix A.2, while summary statistics and demographic characteristics of the datasets are provided in Supplementary Table D1. 2.2 Evaluation of the Segmentation Expert Model For segmentation of SAX cine, 2CH cine, 4CH cine, and SAX LGE sequences, 150 samples were randomly selected from the subject cohort for data annotation (Appendix C). Following stringent quality control, the annotated data were split into training, validation, and test sets in a 7:1:2 ratio. Dedicated segmentation expert models were then developed for each sequence of the SAX cine, 2CH cine, 4CH cine, and SAX LGE types. To evaluate the performance of the segmentation expert model integrated within the BAAI Cardiac Agent, we conducted fair comparative experiments against several widely used SOTA methods, including nnUNet [30], MedSAM2 [31], ResUNet++ [32], and DiffUNet [33]. Segmentation performances across different CMR sequence types was assessed using three standard metrics: Dice Similarity Coefficient (DSC), Hausdorff Distance (HD), and Average Surface Distance (ASD). Details of the calculation formulas of these metrics are provided in the Appendix B. As summarized in Table 1, the proposed segmentation expert model consistently outperforms existing mainstream SOTA methods in cardiac CMR sequence segmentation tasks. It achieves the highest DSC values for all evaluated sequences, while also yielding the lowest HD values in three sequences and the lowest ASD value in two sequences. Across all metrics, it exhibits generally small standard deviations, indicating stable and robust performance. In contrast, the competing methods exhibit only sequence- or metric-specific strengths. For instance, MedSAM2 [31] performs relatively well in terms of ASD for the 2CH cine sequence and achieves the best HD performance for SAX LGE images, while nnUNet [30] shows favorable ASD performance for SAX LGE images, but its overall performance remains inferiors to that of the proposed model. DiffUNet [33] shows substantial higher HD and ASD values, likely due to noise-based initialization mechanism, which introduces additional fluctuations to boundaries, adversely affecting the performance. Quantitative comparisons are presented in Supplementary Fig. D1. Compared with other competing methods, the proposed method demonstrates the best overlap with manual annotations across all CMR sequences. Distinct gaps exist between individual cardiac structures, with no obvious structural missing or false positives. Detailed descriptions of the specific evaluation metrics for different cardiac structures in the four CMR segmentation datasets are provided in Supplementary Tables D2, D3, D4 and D5. Furthermore, clinical practice places particular emphasis on the localization and assessment of LGE involvement extent, and we validate and elaborate on the relevant results of quantitative LGE analysis in Supplementary Fig. D2. 2.3 Evaluation of the CVDs Diagnostic Expert Model To ensure reliable model training and unbiased performance evaluation for the CDS and NICMS tasks, the internal dataset was partitioned into training, internal validation set, and internal test sets in a 7:1:2 ratio using a stratified split strategy that strictly follows the number distribution of CVDs categories. The training set was used for iterative model training, the validation set for real-time generalization evaluation and optimal stopping-point determination, and the internal test set for post-training performance assessment. In addition, an independent external validation set was employed to further evaluate the model’s generalization capability. 2.3.1 Evaluation of CDS The CDS model utilizes three combined views of SAX cine, 2CH cine, and 4CH cine sequences, and was evaluated on the internal test set (n=365). For normal heart (NH) screening, the model achieved an area under the curve (AUC) 3 Cardiac Agent Please generate a comprehensive medical report based on the provided cardiac images. BAAICardiacAgent Uploaddata&state your requirements FunctionCallLibrary Human: <image> reportanalyzingtheprovidedcardiacimages. GPT: thoughts:Toprovideanaccuratemedicalreport, IwillutilizetheMRGtooltoanalyzeimages thoroughly. actions:API–MedicalReportGeneration 1.Request Medical Report Generation 2CHCine Segmentation SAXCine Segmentation SAXLGE Segmentation Cardiac Disease Screening Non-cardiomyopathy Sub Classification Retrieval-Augmented Generation 4CH CineSegmentation 2.Toolusage Scan Sequences: Myocardial Rest Perfusion Imaging; Myocardial Late Gadolinium Enhancement (LGE) Imaging; Cine MRI CineImaging Findings: Severe left ventricular systolic dysfunction, Normal left ventricular wall thickness... Rest Myocardial Perfusion ImagingFindings: No abnormal signal is detected in the left ventricular wall. LGE ImagingFindings: No abnormal signal is identified in the left ventricular wall. Left VentricularRight Ventricular Ejection Fraction 59%67.3% End-Diastolic Volume 142ml144.9ml End-Systolic Volume 58.1ml47.3ml Stroke Volume 83.8ml97.6ml Cardiac Output4.8L/min5.6L/min LV Wall Thickness: Diagnostic Result: High probability of Hypertrophic Cardiomyopathy (HCM) is considered. SegmentationResults: 3. Result Gathering BasicInfo: Name Age ... Cardiac Function Measurement: 4.Response The cardiac agent interprets data and selects tools The cardiac agent performs quantitative calculations and automatically outputs a structured report Agent-Assisted Radiologist Agent-Assisted Radiologist CurrenttoolAuxiliarytool 60s duration The user's requirements are validated, prompting the deployment of the medical report generation tool. I have received the complete structured medical imaging report automatically generated by the cardiac agent. Semi-automatic or manual segmentation of key sequences Semi-automatic or manual quantitative measurement of cardiac chambers and myocardium Cine &LGE & RestMPI Sequence Auxiliary Diagnosis Manual Written Report BasicInfo: Name, Age ... Measurements: Myocardial thickness Ventricular function ... Cine Imaging Findings: Severe left ventricular systolic dysfunction Normal left ventricular wall thickness ... Rest Myocardial Perfusion Imaging(RestMPI) Findings: No abnormal signal is detected in the left ventricular wall. Late Gadolinium Enhancement (LGE) ImagingFindings: No abnormal signal is identified in the left ventricular wall. Diagnosis Result:HCM is considered. The generation of a CMRreport involves semi-automated segmentation of cine images, measurement of key slices, and integrated analysis of multi-modal CMR data for diagnosis. TraditionalRadiologist Select key measurement slices Semi-automated key parameter measurement 1800s duration Figure 1: A schematic overview of the BAAI Cardiac Agent illustrating the process of radiology report generation from multi-sequence CMR scans (upper section), in comparison with the traditional clinical workflow (lower section). Upper section: the agent-assisted workflow, in which the user submits instructions and upload all CMR sequences to the BAAI Cardiac Agent. The agent automatically interprets the request, invokes appropriate analytic modules, and performs segmentation, quantitative assessment, and diagnostic reasoning. A structured, quantitative enriched radiology report is then generated within approximately 90 seconds. Lower section: the workflow of a traditional radiologist. This workflow begins with manual or semi-automatic segmentation of cine images, followed by qualification of key cardiac function parameters. Diagnostic interpretation is then performed by integrating information from late gadolinium enhancement (LGE) and rest myocardium perfusion imaging (Rest MPI), after which the final report is manually compiled. This sequential workflow typically requires approximately 1800 seconds. of 0.980 (95% confidence interval [CI]: 0.956–0.995) and an F 1 -score of 0.856 (95% CI: 0.782–0.914). In anomaly detection tasks, ischemic heart disease (IHD) screening yielded an AUC of 0.938 (95% CI: 0.911–0.961), an F 1 -score of 0.860 (95% CI: 0.819–0.897), with a sensitivity of 0.820 (95% CI: 0.760–0.875) and a specificity of 0.931 (95% CI: 0.897–0.962). For non-ischemic cardiomyopathy (NICM) screening, the model attained an AUC of 0.960 (95% 4 Cardiac Agent MetricMethodSAX cine2CH cine4CH cineSAX LGE ResUNet++84.69±0.0586.97±0.0484.41±0.0658.08±0.16 nnUNet87.42±0.0688.14±0.0581.61±0.1473.89±0.18 DSCDiffUNet77.19±0.1383.05±0.0877.38±0.1147.28±0.20 MedSAM280.86±0.2077.89±0.2082.07±0.0869.70±0.21 Ours90.21±0.0288.75±0.0386.92±0.0375.07±0.07 ResUNet++16.58±8.1312.28±6.6213.73±10.4616.11±5.57 nnUNet13.68±9.7629.88±49.7412.17±3.8812.46±5.52 HDDiffUNet92.30±24.83117.35±36.6172.80±52.81185.59±63.24 MedSAM219.39±5.439.22±3.817.94±1.8912.18±4.49 Ours8.24±2.527.87±2.187.52±1.8714.05±4.89 ResUNet++0.94±0.580.64±0.420.94±1.271.29±1.47 nnUNet0.75±0.782.72±7.960.67±0.620.63±0.26 ASDDiffUNet2.65±1.855.46±5.808.26±8.1256.65±60.78 MedSAM22.45±0.790.25±0.170.26±0.111.15±1.48 Ours0.64±0.290.27±0.170.26±0.131.14±0.78 Table 1: Performance comparison between the BAAI Cardiac Agent segmentation expert model and SOTA segmentation networks. Computed DSC (%), HD, and ASD values are shown as mean ± standard deviation. CI: 0.943–0.975), an F 1 -score of 0.857 (95% CI: 0.810–0.896), a sensitivity of 0.879 (95% CI: 0.821–0.932), and a specificity of 0.893 (95% CI: 0.852–0.931). External validation on an independent cohort (n=275) demonstrated stable performance, with AUC values exceeding 0.810 across all three categories, supporting the model’s generalization ability. Detailed results are presented in Fig. 2 (a–c) and Table 2. Notably, the study cohort comprises multiple types of CVDs (Table D1), indicating the robustness of the screening model across diverse disease types. Internal Test SetExternal Validation Set NOursVSTViTResNetNOursVSTViTResNet NH640.8560.8390.8550.864670.7530.7120.7200.717 IHD1610.8600.8120.7990.835490.5820.6010.5440.631 NICM1400.8570.8300.8170.8321590.7620.7160.7030.713 F-W F 1 0.8580.8240.8170.8390.7280.6940.6790.699 Accuracy0.8580.8250.8160.8330.7240.6870.6730.695 Table 2: Performance comparison of cardiac disease screening (CDS) in normal heart (NH), ischemic heart disease (IHD) and non-ischemic cardiomyopathy (NICM) cohorts between the BAAI Cardiac Agent CDS expert model and three SOTA models. (N: Number of subjects, F-W F 1 : Frequency-weighted F 1 ) 2.3.2 Evaluation of NICMS The diagnostic model for NICMS adopts the combined input of dual-view cine sequences (SAX, 4CH) and SAX LGE, thereby leveraging the complementary structual, functional, and tissue-characterization information provided by CMR. On the internal test set (n=155), the model achieved a class-weighted mean AUC of 0.9650 (95% CI: 0.944-0.998) and an F 1 -score of 0.834 (95% CI: 0.779–0.890). Across all disease categories, AUC values exceeded 0.930, and F 1 -scores were greater than 0.800 for all subclasses ex- cept restrictive cardiomyopathy (RCM) and arrhythmogenic cardiomyopathy (ACM). Notably, the model demonstrated excellent performance for the most prevalent NICM subtypes, achieving an AUC of 0.975 (95% CI: 0.948–0.992) and F 1 -score of 0.899 (95% CI: 0.844–0.947) for hypertrophic cardiomyopathy (HCM), and an AUC of 0.971 (95% CI: 0.945–0.992) with an F 1 -score of 0.800 (95% CI: 0.667–0.895) for dilated cardiomyopathy (DCM). In addition, the model also achieved a high AUC of 0.938 (95% CI: 0.870–0.991) and an F 1 -score of 0.857 (95% CI: 0.722–0.955) for myocarditis. External validation on an independent cohort (n=159) confirmed the stability and generalizability of the model, with AUC values exceeding 0.860 for all five NICM categories. Detailed results are presented in Fig. 2 (d–h) and Table 3. 5 Cardiac Agent (a)(b)(c) (d)(e)(f) (g) (h) (i) Figure 2: Performance of the expert CVDs diagnostic model in internal and external testing. a–c: Receiver operating characteristic curves for CDS; d–h: Receiver operating characteristic curves for NICMS. The red curves represent the internal test set, and the blue curves represent the external validation set. i: Confusion matrix of the AI diagnostic model predictions versus ground truth labels based on the full cohort (n = 2,232), with a color gradient visually representing the proportional distribution of predictions across disease categories. 2.3.3 Comparison of CVDs Diagnosis with Other Models To assess the relative performance of the proposed diagnostic expert model within the BAAI Cardiac Agent, we conducted fair comparative experiments against representative mainstream backbone networks in the fields of computer vision and medical image analysis, including Video-based Swin Transformer (VST) [4], Vision Transformer (ViT) [34], and ResNet [35]. Detailed quantitative comparisons among the models are reported in Table 2 and Table 3. For the CDS task, the proposed model achieved the highest class-weighted F 1 -score, with superior performance in both IHD and NICM categories. In the NICMS task, the proposed model likewise attained the optimal weighted F 1 -score, and outperformed competing methods in key subclasses, including HCM and myocarditis. 6 Cardiac Agent Internal Test SetExternal Validation Set NOursVSTViTResNetNOursVSTViTResNet HCM720.8990.8650.8480.888430.7690.7290.7190.748 DCM260.8000.7980.8130.800910.8120.7110.7210.714 RCM180.7320.7620.8150.76260.4440.4830.5670.567 ACM130.6360.6360.7270.72780.4210.6570.4350.421 Myocarditis260.8570.7790.7790.797110.6210.6710.5970.653 F-W F 1 0.8340.8070.8180.8300.7540.7020.6910.698 Accuracy0.8320.8070.8130.8190.7300.6850.6670.673 Table 3: Performance comparison between the BAAI Cardiac Agent expert diagnostic model and three SOTA models in the subclassification of non-ischemic cardiomyopathy, including hypertrophic cardiomyopathy (HCM), dilated cardiomyopathy (DCM), restrictive cardiomyopathy (RCM), arrhythmogenic cardiomyopathy (ACM) and myocarditis. These results demonstrate the robust advantages of the proposed model for CDS and its accuracy in identifying subdivided diseases. In addition, comprehensive comparison of model performances between the proposed method and three types of SOTA methods in terms of AUC, sensitivity, and specificity metrics for the tasks of CDS and NICMS, and the relevant results were conducted, and the results are shown in Supplementary Tables D6, D7 and D8. These results further demonstrated that the BAAI Cardiac Agent outperformed SOTA methods. 2.4 Evaluation of Automatic Cardiac Structural and Functional Measurement We performed multi-dimensional quantitative analysis of cardiac structural and functional parameters based on the segmentation results of CMR cine sequence images, so as to systematically assess the measurement accuracy and clinical applicability of the BAAI Cardiac Agent. To evaluate the consistency and reliability between the measurement results of the agent and those obtained by manual measurement from clinicians, we also calculated six primary cardiac function parameters based on the segmentation results of SAX cine sequence images, and analyzed the data by using Bland-Altman plots (Fig. 3(a)). The results demonstrated strong concordance between the automated and manual measurements. Pearson correlation coefficients (r) of LV end-diastolic volume (LVEDV), and LV end-systolic volume (LVESV) both exceeds 0.960, while therof LV ejection fraction (LVEF), stroke volume (SV), and LV mass (LVM) were also above 0.900. For LVED diameter (LVEDD), a distance-based measurement parameter, theris 0.874. These findings indicate excellent correlations in the measurement results of all core cardiac function parameters. The consistency evaluation results of other cardiac function parameters, including cardiac output (CO) and the maximum transverse diameters of the left and right atria (LAT4CHD, RAT4CHD) measured in the 4CH view are presented in Supplementary Fig. D3 (a). The above results demonstrate that the measurement results of the BAAI Cardiac Agent are highly consistent with those of manual measurement by clinicians, which provides reliable data support for its clinical application and popularization. Based on the American Heart Association (AHA) 17-Segment Model (17-SM), we analyzed the measurement results of LVED wall thickness (LVEDWT) and visualized the average segmental wall thickness distribution across the internal dataset using polar maps (Fig. 3(b)). The zonal layout of each polar map strictly adheres to the AHA 17-SM. The visualization results show that mean wall thickness of the anteroseptal wall (ASW) in the basal segment and the septal wall (SW) in the apical segment exceeds 1 m, while that of all other segments is below 1 m. To further assess measurement fidelity, we compared segment-wise wall thickness between the NH and HCM cohorts of the internal dataset (Supplementary Fig. D3(b)). The results demonstrate that the differences between the values measured by the agent and those in clinical reports were minimal across all segments. Notably, in the HCM cohort, the greatest wall thickening was observed in the basal inter-ventricular septum, a region known to experience the highest mechanical load and to be the most common site of myocardial hypertrophy seen in clinical practice. The concordance between these findings and established pathological observations provides additional evidence of the accuracy and clinical validity of the BAAI Cardiac Agent for ventricular structure assessment under pathological conditions. 2.5 LMM Performance Evaluation for BAAI Cardiac Agent The trained BAAI Cardiac Agent achieves a near-perfect tool selection success rate (close to 100%), indicating its strong capability to accurately identify and invoke appropriate tools. To rigorously and comprehensive evaluate this ability, we randomly sampled 200 patients from the internal dataset and included all 279 patients from the external 7 Cardiac Agent (a) (b) (c) (d) Figure 3: Evaluation results of BAAI Cardiac Agent report consistency and expert model invocation success rate. (a) Bland-Altman plots comparing BAAI Cardiac Agent measurements (left ventricular end-diastolic volume, LVEDV; left ventricular end-systolic volume, LVESV; left ventricular ejection fraction, LVEF; stroke volume, SV; left ventricular mass, LVM; left ventricular end-diastolic diameter, LVEDD) with manual reports in the internal cohort; (b) Mean distribution of left ventricular end-diastolic wall thickness (LVEDWT) measured by manual reports using the 17-SM and by the BAAI Cardiac Agent in the internal cohort. The bullseye plot correlates the basal, mid, and apical segments of the left ventricle with the outer, middle, and inner layers in sequence, with the apex at the center; (c) Success rate of expert model invocations by the BAAI Cardiac Agent on the internal and external test sets; (d) Accuracy performance of the BAAI Cardiac Agent in cardiac-specific VQA tasks. dataset for validation. The results show that across diverse CMR sequence tasks, the LMM exhibits robust 3D data 8 Cardiac Agent ModelJunior radiologistsMid radiologistsSenior radiologists ScoreConf.ScoreConf.ScoreConf. LLaVa-Med26.10±2.8324.5022.07±3.2221.0013.48±3.5013.50 MedM-VL27.93±6.1223.0023.85±6.8122.0015.65±5.5019.00 MMedAgent24.05±3.2022.5020.02±3.6020.0012.78±3.2517.50 Qwen-VL-30B58.52±3.7352.5057.18±4.6849.5051.67±5.7546.00 Ours87.93±2.2094.0087.52±2.8490.5086.53±4.2187.50 Table 4: Performance comparison of different LMMs across radiologist experience levels (junior, middle and senior). Score denotes the comprehensive rating by radiologists of CMR report generation by different models, and Conf. denotes the increase in confidence in writing reports reported by radiologists after reviewing the model output. comprehension, achieving a tool invocation success rate of 99.82% on the internal validation dataset and 99.46% on the external validation dataset. The success rates for individual tools are shown in Fig. 3(c). We evaluated the performance of the LMM on VQA tasks for CMR imaging, with results summarized in the lollipop chart (Fig. 3(d)). Overall, the model achieves an accuracy exceeding 0.965 in assessing the structural normality of categories with low abnormality rates, including right ventricular size (RVS), RV wall motion (RVWM), RV systolic function (RVSF), and RV diastolic function (RVDF). In contrast, for categories with higher abnormality prevalence, such as LVS, LVWM, LVSF, LVDF, mitral valve (MV), tricuspid valve (TV), pericard and rest myocardial perfusion imaging (Rest MPI) abnormalities, the evaluation accuracy ranges from 0.610 to 0.860. Using the larger-scale Qwen3-VL-30B [36] model as a baseline, we conducted a fair comparison across four repre- sentative tasks: LVSF assessment, MV analysis, pericardial effusion (PE) evaluation, and Rest MPI. Qwen3-VL-30B achieved accuracy rates of 0.600, 0.380, 0.625, and 0.500 respectively, which are consistently lower than those obtained by the proposed BAAI Cardiac Agent by margins of 0.110, 0.330, 0.045, and 0.110 respectively. These results indicate that the baseline Qwen3-VL-30B [36] model exhibits limited capability in detecting CMR abnormalities. Although the screening performance of the proposed BAAI Cardiac Agent remains moderate for the above four tasks, it established a solid foundation for accurate image interpretation and clinical description in large-scale medical multimodal models. In addition, the model achieves an overall accuracy of 0.980 in the CMR sequence recognition (CMRSR) task. We further conducted qualitative benchmarking against leading open-source large multimodal model (LMM) frameworks including MedM-VL [37], LLaVA-Med [38], MMedAgent [25] and Qwen3-VL-30B [36] across a diverse set of tasks. As presented in the Supplementary Fig. D5, BAAI Cardiac Agent consistently outperforms these models with more robust and reliable outputs: it not only addresses complex tasks with high accuracy but also achieves zero hallucination, effectively avoiding misleading conclusions such as misidentification of CMR sequences and erroneous calculation of LVEF. Furthermore, its strong generalization capability enables adaptation to diverse CMR scan protocols and heterogeneous patient cohorts, an advantage that existing LMMs constrained by domain-specific data fail to match. A complete CMR medical report is automatically generated by integrating results from cardiac function measurement, CVDs diagnostic conclusions, and imaging findings. Using the 200 medical reports randomly sampled from the internal dataset described above, we quantitatively evaluated the semantic similarity between the generated reports and the corresponding reference reports using the Bidirectional Encoder Representations from Transformers Score (BERT Score). The evaluation yields a precision, recall, and F 1 -score of 0.903, 0.894, and 0.898, respectively, indicating a level semantic consistency between the generated and reference reports. A complete example of the CMR report is presented in Supplementary Fig. D4. To further verify the clinical performance of each LMM and agent, we conduct subjective evaluations on 200 internal patients with complete CMR reports by two radiologists from each of three experience levels: junior (less than 5 years), mid (5–10 years), and senior (more than 10 years). The evaluation uses a 100-point scale, with detailed scoring criteria provided in the Appendix D. Table 4 summarizes the corresponding performance results across these groups. CMR report scoring results are presented as mean ± standard deviation. The proposed BAAI Cardiac Agent achieves significantly higher scores and better stability across all three groups. Qwen-VL-30B [36] shows moderate performance and ranks second, whereas LLaVa-Med [38], MedM-VL [37], and MMedAgent [25] obtain considerably lower scores, especially under the stricter assessment of senior radiologists. Additionally, the table presents the improvement in report-writing confidence among radiologists at different levels following the use of various models, with a maximum score of 100. Our BAAI Cardiac Agent still achieves the highest rating, followed by Qwen-VL-30B [36], while the other comparative models yield relatively smaller gains in radiologist confidence. Overall, radiologists with more experience assign lower ratings, and the BAAI Cardiac Agent consistently outperforms existing mainstream models in clinical evaluation. 9 Cardiac Agent 3 Discussion Unlike general medical imaging tasks, cardiac imaging diagnosis constitutes a highly composite clinical workflow that integrates quantitative functional assessment, sequence dependency, temporal phase sensitivity, standardized processing, and clinical interpretability. As the recognized gold standard for evaluating cardiac structure and function, CMR has been integrated into the entire clinical management pathway of CVDs [7,8], including screening, diagnosis, treatment response monitoring, and prognostic evaluation. Particularly in the differential diagnosis of suspected complex CVDs, CMR demonstrates irreplaceable advantages by enabling precise characterization of myocardial tissue features. Despite the rich diagnostic information offered by multimodal CMR, its potential has not yet been fully translated into clinical efficiency, and conventional CMR analysis remains heavily dependent on physician experience and manual interpretation, leading to persistent challenges. The interpretation of imaging phenotypes is complex, as cardiac motion coupled with multi-sequence data imposes stringent requirement on clinician’s knowledge of cardiac anatomy, physiology and pathology [9,10]. Manual processing procedures are labor-intensive and time-consuming, involving extensive repetitive quantitative analysis that substantially increase physician workload. Moreover, limited standardization and pronounced subjectivity in analysis hinder reproducibility, restrict inter-institutional collaboration, and compromise the comparability of clinical outcomes. To directly address the clinical bottlenecks in CMR interpretation, we developed BAAI Cardiac Agent, the first end-to- end agent framework specifically designed for CMR imaging analysis. The framework enables integrated analysis of multi-sequence CMR images, conducting functions such as automatic segmentation of cardiac structures, screening and diagnosis of CVDs, MRG, RAG, and mining of imaging manifestations. Validation based on a large-scale clinical cohort of 2,413 cases demonstrates that BAAI Cardiac Agent delivers robust and comprehensive CMR imaging interpretation capabilities across multiple evaluation dimensions. The system effectively emulates radiologist-level sequence-specific analysis, reliably extracts clinically relevant imaging features, and produces structured reports that are consistent with expert interpretations. By providing an integrated and scalable solution for intelligent cardiovascular imaging analysis, BAAI Cardiac Agent facilitates the translation of CMR from a high-information yet resource-intensive modality into a clinically efficient diagnostic tool, with the potential to streamline imaging workflows and optimize downstream clinical decision-making. The BAAI Cardiac Agent not only possesses visual question-answering capabilities but also integrates eight expert models, including segmentation expert models for cardiac structure segmentation (SAX, 2CH, 4CH cine, and SAX LGE), the CDS and NICMS models for cardiovascular disease screening and diagnosis, as well as MRG and RAG models. For the tasks of multi-sequence segmentation in CMR and cardiovascular disease screening and diagnosis, this study proposes a generalized CMR segmentation and diagnosis framework. Compared with advanced approaches such as ResUNet++ [32], nnUNet [30], DiffUNet [33], and MedSAM2 [31], the proposed framework achieves the best performance in terms of the DSC across all four sequences and significantly outperforms its counterparts in terms of the HD and ASD metrics. In terms of disease screening and diagnosis, the model achieves an AUC exceeding 0.930 for various cardiovascular diseases on the internal test set and demonstrates strong generalization performance on external test sets. Compared with mainstream models such as ResNet [35], ViT [34], and VST [4], our method exhibits clear advantages in both screening and diagnostic effectiveness. We evaluate the consistency between the key cardiac function parameters automatically measured by the BAAI Cardiac Agent and the clinical imaging reports in the internal dataset. The results show that the correlation coefficients of the LVEDV, LVESV, LVEF, SV, LVM, and LVEDD automatically measured by the Agent with the imaging reports are 0.968, 0.979, 0.925, 0.906, 0.937, and 0.874 respectively, indicating its good clinical applicability. In addition, we further compare the mean predicted myocardial thickness of each segment by the Agent based on the AHA 17-SM with the mean clinical report values in the entire cohort, normal population cohort, and HCM cohort. The results demonstrate that the thickness prediction error remains within 1 m in each segment of different cohorts, reflecting the stability and anatomical consistency of its measurement results. Overall, the text quality of the generated imaging reports is evaluated by BERT Score, with precision, recall, and F 1 score reaching 0.901, 0.896, and 0.898 respectively, verifying from the intelligent text quality perspective that BAAI Cardiac Agent reports are content-accurate, information-complete and clinically reliable. By virtue of a highly integrated, expert model system validated on large-scale clinical cohorts, BAAI Cardiac Agent enables reliable full-chain output ranging from image analysis to report generation, thereby laying a solid technical foundation for the intelligent and standardized interpretation of CMR. The core engine of the BAAI Cardiac Agent is built upon large-scale medical models with native support for 3D image inputs. In contrast to existing open-source advanced LMM frameworks including MedM-VL [37], LLaVA-Med [38], MMedAgent [25], and Qwen3-VL-30B [36], the proposed system can deeply understand the spatiotemporal features of 3D cardiac images. Based on 3D inputs, the system can accurately dispatch expert models according to user needs, with task invocation pass rates approaching 100% on both internal and external test sets. Compared with mainstream LMMs, 10 Cardiac Agent BAAI Cardiac Agent confers a substantial reliability advantage in the execution of specific tasks, as it can effectively eliminate model hallucinations and achieve accurate understanding of task requirements as well as efficient response. Moreover, in the task of describing the imaging manifestations of CMR images across different sequences, BAAI Cardiac Agent with 7B parameters consistently outperforms the larger-scale Qwen3-VL-30B model in terms of descriptive accuracy and clinical relevance. Together, the aforementioned results demonstrate that BAAI Cardiac Agent delivers more accurate, reliable and clinically aligned image interpretation and reporting, while supporting an efficient and automated analytical workflow. This performance highlights its superior comprehensive performance and clinical utility for complex cardiac image analysis. Naturally, this study has several limitations that warrant to be further addressed in future. First, at the dataset level, although the included CVDs categories encompass most common clinical presentations, the sample size for rare and diagnostically challenging subtypes, particularly ACM and RCM, remains limited due to their low incidence. Consequently, the current data scale does is not yet sufficient to fully support the verification of the model’s diagnostic efficacy for such diseases. Future work will expand these cohorts through multi-center cooperation, longitudinal data collection to further improve model robustness and generalizability across full spectrum of CVDs. In addition, the study population are mainly from the East Asian. Given known differences in regional genetic backgrounds, lifestyle factors, and clinical diagnosis and treatment standards, further validation in diverse populations from Europe, the Americas, and Africa will be necessary to establish cross-ethnic diagnostic consistence and support global deployment. Second, at the level of diagnostic modality integration, diagnosis of CVDs is a complex process involving multi- dimensional and multi-modal collaboration. Although CMR is recognized as the gold standard for CVD diagnosis and cardiac structural and function assessment, specific diseases phenotypes may be more directly or efficiently characterised using specific imaging modalities, for example, coronary computed tomography angiography (CCTA) for coronary artery disease (CAD) [39], or echocardiography for dynamic assessment of myocardial hypertrophy morphology [40]. The current framework focuses primarily on CMR images, and does not yet incorporate additional imaging modalities (such as CT, ultrasound) or non-imaging clinical data (such as electrocardiogram, or laboratory tests). Future development will extend BAAI Cardiac Agent toward a CMR-centered, multimodal diagnostic framework that integrates coronary CTA, echocardiography, and structured clinical information. This evolution reflects not a limitation of the current approach, but a natural progression toward a more comprehensive and clinically aligned cardiovascular intelligence system capable of supporting complex diagnostic reasoning and downstream clinical decision-making. In summary, this study demonstrates that BAAI Cardiac Agent can emulate key components of expert radiologist practice by performing quantitative cardiac structural and function assessment, detecting cardiac abnormalities, gener- ating structured diagnostic reports, and contextualizing findings with medical knowledge for patient communication. Following rigorous clinical validation, this approach has the potential to substantially improve the efficiency, consistency, and scalability of CMR interpretation, providing a foundation for broader application in CVD and other complex clinical imaging diagnosis. 4 Method 4.1 Ethics approval and data acquisition All datasets were collected with approval from the Institutional Review Board (IRB) of Beijing Anzhen Hospital affiliated to Capital Medical University (2025216x), and the First Affiliated Hospital of Xinxiang Medical University (2019164). All data were de-identified prior to model development. Owing to the retrospective nature of this study, the requirement for individual informed consent was waived by the IRBs. Details of CMR acquisition for all patients are provided in the Appendix A.1. 4.2 LMM serves as the engine of the BAAI Cardiac Agent Following the paradigm of LLaVA-Plus [41] and LLaVA-Med [38], our objective is to enable LLM to effectively leverage a diverse set of medical multi-modal tools, thereby supporting complex real-world clinical scenarios. In this work, we use the LLaVA [42] model as an illustrative example; however, it should be noted that other LMMs are equally applicable as the engine of agents in medical scenarios. For an LMM, its workflow typically commences with an image queryI q input by the user, followed by the reception of a natural language instructionX q provided by the user, and ultimately outputs a natural language answerX answer . Therefore, a unified format can be adopted to represent multimodal instruction-following data, with the specific form presented as follows: User : I q < > X q < STOP > LMM : X answer < STOP >(1) 11 Cardiac Agent Can you recommend a tool to segment and highlight the cardiovascular features in this image? "thoughts": "This is a 4CH cine image, I need to employ a medical image segmentation tool to clearly delineate the Cardiovascular.", "actions": ["API_name": ”4CH Cine Segmentation", "API_params": ], "value": "Certainly! I have used a medical image segmentation model to highlight the Cardiovascular for 4CH cine image" User BAAICardiac Agent Expert Toolexecution: 4CH Cine Segmentation User "thoughts": "Thanks to the output of 4CH Cine Segmentation. I can answer the question better.", "actions": [], "value": "The result below provides a clear depiction, aiding in your medical evaluation." 4CH cine segmentation output: 'message': 'The 4CH Cine Segmentation has processed the image.’ Answer my first question: Which segmentation tool would be most appropriate for highlighting the cardiovascular structures in this medical image?" (X q , I q ) (X tool_use ) (X tool_result ) BAAICardiac Agent (X answer ) Figure 4: Schematic diagram of the BAAI Cardiac Agent workflow, in which assigned tasks are completed through adaptive invocation and coordinated use of specialized medical image segmentation tools. where< >and< STOP >are the newline marker and sequence end marker, respectively. This format is naturally suitable for various multimodal tasks that can be formulated as language-image inputs with language-based outputs. It accommodates a variety of visual tasks, including image recognition, image description, VQA, and MRG, among others. 4.2.1 BAAI Cardiac agent workflow To ensure the BAAI Cardiac Agent simultaneously functions as both an action planner and a results aggregator, this study adopts the unified dialogue format proposed by LLaVA [42]. As illustrated in Fig. 4, the system operates through a structured workflow consisting of four key stages: (1) The user provides a task instructionX q alongside a contextual imageI q . (2) Serving as a central planner and router, the LMM first generatesthoughtsthat analyzeX q andI q to identify the optimal tool from the skill repository for task fulfillment and the corresponding required parameters. It then formulatesactionsthat explicitly specify the API name of the designated tool and its associated parameters, followed by outputtingvalueto inform the user that the selected tool will be invoked for task processing. These three fields collectively constitute the structured commandX tool_use . (3) The designated tool is executed withI q as input, and its outputX tool_result is returned to the LMM. (4) Finally, the BAAI Cardiac Agent first generatesthoughtsthat integrate X tool_result ,X q , andI q to confirm how to organize a coherent response. It setsactionsas an empty list, indicating no additional tools need to be invoked, before synthesizing all information to generatevalue—the coherent and informed final response X answer provided to the user. This interaction process can be represented as: User : I q < > X q < STOP > LMM : X tool_use < STOP >(2) User : X tool_result X q < STOP > LMM : X answer < STOP >(3) 12 Cardiac Agent thoughts: This is an LGE SAX image. I need to use a medical image segmentation tool to clearly delineate the cardiovascular structures.. actions: API – SAX LGE Segmentation (SAXLGES) Can you recommend a tool to segment and highlight the cardiovascular features in this image? From the provided cardiac MRI image, what non-ischemic cardiomyopathy diagnosis would you make? What segmentation technique should be applied to enhance the cardiovascular visualization in this image? What tool should be used to clearly show the Cardiovascular in the medical image? Please analyze this image and determine if perfusion abnormality exists. How can I best highlight the cardiovascular regions in this medical image using available tools? thoughts: Determining the type of cardiac disease requires analyzing the specific characteristics present in the image. actions: API – Cardiac Disease Screening (CDS) thoughts: This is a SAX cine image. I need a medical image segmentation tool that can clearly outline the cardiovascular structures. actions: API – SAX Cine Segmentation (SAXCS) thoughts: This is a 4CH cine image, for which a medical image segmentation tool is required to accurately delineate the cardiovascular structures. actions: API – 4CH Cine Segmentation (4CHCS) thoughts: This is a 2CH cine image. I need to use a medical image segmentation tool to clearly delineate the cardiovascular structures. actions: API – 2CH Cine Segmentation (2CHCS) thoughts: Determining the type of non-ischemic cardiomyopathy requires analyzing the specific characteristics present in the image. actions: API – Non-ischemic Cardiomyopathy Subclassification (NICMS) thoughts: Determining perfusion abnormality requires analyzing the specific characteristics present in the perfusion image. actions: VisualQuestion Answering (VQA) thoughts: To provide an accurate medical report, I will utilize the MRG tool to analyze images thoroughly. actions: API – Medical Report Generation (MRG) thoughts: To provide you with reliable information, I will use the RAG tool to access verified sources. actions: API – Retrieval-Augmented Generation (RAG) What type of cardiac disease is shown in this image: normal, ischemic, or non- ischemic? Generate a comprehensive medical report analyzing the provided cardiac image. What is the specific disease progression trend in patients diagnosed with hypertrophic cardiomyopathy, especially regarding the changes in myocardial hypertrophy, cardiac function and clinical symptoms over the long-term follow-up period? UserBAAICardiacAgent Figure 5: Task-modality-tool mapping of the BAAI Cardiac Agent: the left panel presents the tasks integrated in the agent and their corresponding medical imaging modalities, while the right panel lists the tools required to complete the respective imaging tasks. The loss is computed only for the reasoning text (thoughts) related to tool invocation, the tool invocation list (actions), and the final natural language answer (value). Loss for redundant text irrelevant to tool usage is not calculated, ensuring that training resources are concentrated on improving tool usage capabilities. Based on SAX cine, 2CH cine, 4CH cine, LGE, and Rest MPI, this study generated 32k instruction-tuning data samples, including 8K augmented VQA instructions, as well as 3K samples each for the tasks of SAXCS, 2CHCS, 4CHCS, SAXLGES, CDS, NICMS, RAG, and MRG. The LMM initializes with LLaVA-Med 60K-IM. Lightweight fine-tuning performs on the language model using the Low-Rank Adaptation (LoRA) technique [43], with training conducted for 10 epochs and a batch size set to 8. By training only a small number of low-rank adaptation parameters, the model effectively achieves a favorable balance between computational efficiency and model performance. The LoRA hyperparameter configuration in this experiment is as follows: rank sets to 128, alpha coefficient to 256 and dropout rate to 0.05. To improve training efficiency and reduce memory consumption, the DeepSpeed ZeRO-3 optimization strategy [44] adopts during training. It integrates optimizer state sharding, gradient checkpointing technology and FP16 mixed-precision training [45]. In terms of hyperparameter settings, a learning rate of 1e-4 selects for language model fine-tuning, combined with a cosine learning rate scheduler to ensure stable convergence. This relatively high learning rate can effectively facilitate model adaptation to tasks. Meanwhile, the learning rate of the visual projector configures independently to 2e-5. The low learning rate preserves the modality alignment effect achieved in previous training. 4.2.2 Tasks and tools of the BAAI Cardiac Agent The BAAI Cardiac Agent in this study is capable of accessing a diverse range of tools and features scalability to adapt to multi-task processing. As shown in Fig. 5, we integrate 9 tools covering major representative tasks in the CMR domain. It should be noted that no additional tools are required for the VQA task, as we utilize LLaVA-Med [38] as the backbone model, which natively supports the VQA task. Each tool has specialized functional attributes and exhibits exceptional performance in executing specific cardiac magnetic resonance tasks. For the segmentation tasks (Tasks 1–4), this study develops a general-purpose medical image segmentation architecture independently, as detailed in Section 4.3, instead of selecting MedSAM2 [31], nnUNet [30], ResUNet++ [32], or DiffUNet [33] as the core tools. This decision is based on two primary considerations. First, accurate segmentation results are the fundamental prerequisite for precisely measuring key cardiac functional parameters. Second, the strong robustness of the segmentation model architecture across multiple CMR sequences is of vital significance for the practical deployment of the BAAI Cardiac Agent. This study presents a comparative analysis of segmentation 13 Cardiac Agent results between the proposed architecture and the aforementioned four SOTA methods (MedSAM2 [31], nnUNet [30], ResUNet++ [32], and DiffUNet [33]) in Section 2.2. For the diagnostic tasks (Tasks 5–6), this study designs a two-stage diagnostic framework. First, it performs preliminary screening of NH, IHD, and NICM based on CMR cine data. Second, it incorporates SAX LGE data to enable fine- grained classification of NICM subtypes, including HCM, DCM, RCM, ACM, and myocarditis. Detailed specifications of the diagnostic architecture are provided in Section 4.4. To validate the effectiveness of the proposed CMR-based diagnostic method, we conduct fair comparative experiments against SOTA approaches, including the CMR diagnostic method based on VST [4], ViT [34], and 3D ResNet [35]. As demonstrated in the diagnostic results presented in Section 2.3, the reliability of our method is fully verified. For the generation of CMR medical reports, LVEDWT by the AHA 17-SM can be measured based on segmentation results of SAX cine images, including the anterior wall (AW), anteroseptal wall (ASW), inferoseptal wall (ISW), inferior wall (IW), inferolateral wall (ILW), and anterolateral wall (ALW) at the basal segment; the AW, ASW, ISW, IW, ILW, and ALW at the mid segment; and the AW, SW, IW, and lateral wall (LW) at the apical segment. Cardiac function analysis of the LV mainly includes the measurement of LVEF, LVEDV, LVESV, SV, CO, and LVM. LV cardiac apex thickness is measured using segmentation results of 4CH cine images. Based on 4CH cine images, additional parameters can be quantified, including the maximum transverse diameters of the left and right atria measured in the 4CH view, denoted as LAT4CHD and RAT4CHD. RAG refers to a technical method that enhances the quality of generated results by integrating the most relevant information acquired from external data sources. In this study, we extract the cardiovascular section from ChatCAD+ [46], and combine it with clinical guidelines [47,48,49] related to CVDs as well as heart-related literature retrieved from PubMed, to implement the medical retrieval process. In the VQA component, the core focus lies on the holistic diagnostic description of CMR data by the BAAI Cardiac Agent. To this end, three categories of structured tasks are designed: (1) Overall CMRSR for clinical cardiac image analysis tasks; (2) Panoramic description of cine sequence data, covering the normality of key anatomical structures and functional indicators including VS, VWM, VSF and VDF, mitral valve MV, TV, and PE, as well as interpretation of Rest MPI, which is primarily employed to evaluate whether the myocardial blood perfusion status conforms to normal physiological standards; (3) Conversational content related to CVDs. 4.3 CMR-based expert model for cardiac segmentation The CMR-based expert model for cardiac segmentation adopts the same two-stage coarse-to-fine architecture for 2CH cine, 4CH cine, SAX cine, and SAX LGE data, with the detailed structure illustrated in Fig. 6 (a). The first stage localizes the cardiac region of interest (ROI), and the second stage performs fine-grained segmentation based on the localization results. In the first stage of the general cardiac segmentation expert model, coarse cropping is performed on original 3D medical images. The spatial extent of the cropped volume is determined by a randomly sampled center, enabling extraction of a local subvolume containing the target region. Subsequently, the cropped region is resampled to a uniform voxel spacing. During resampling, continuous interpolation is adopted for intensity images to preserve signal continuity, while nearest-neighbor interpolation is utilized for segmentation labels to maintain boundary integrity. After cropping and resampling, the images undergo min-max normalization and are fed into a unified U-shaped backbone, which outputs accurate localization results of the heart. The second stage of the segmentation expert model is performed based on the ROI localized in the first stage, with its core adopting a dual-path pyramid feature extraction design. This design comprises two cropping branches: global and local. The receptive field of the global cropping region is twice that of the local branch, enabling capture of broader contextual information, while its resolution is only half of the local branch. The local cropping branch focuses on high-resolution detailed features. A random strategy is employed to determine the cropping centers, after which both branches are processed into unified fixed-size image patches (e.g., 32×64×64 for 2CH or 4CH cine, 64×64×64 for SAX cine, and 3×64×64 for SAX LGE). In the feature fusion process, both branches utilize the same U-shaped backbone network configured identically to the first stage. The global branch first performs feature extraction, and then upsamples its feature maps to a spatial size matching the local branch via sub-pixel sampling. The fused features, together with the original locally cropped images, are fed into the U-shaped network of the local branch for deeper and finer-grained feature mining, ultimately outputting accurate cardiac segmentation results. To enhance the model’s robustness against anatomical variability and acquisition noise, a 3D data augmentation strategy including random flipping, scaling, translation, elastic deformation, and random field perturbation is implemented during the training process of both stages. 14 Cardiac Agent (a) (b) (c) U-shape Backbone U-shape Backbone U-shape Backbone Cropping Back ROI cropping Stage1: Localization Stage2: Fine Segmentation Subpixel Sampling LocalCropping GlobalCropping Stage 2 HCM RCM Myocarditis DCM ACM 4CH cine SA cine SA LGE Stage 1 ResNet-like Encoder ResNet-like Encoder ResNet-like Encoder NH IHD NICM 2CH cine 4CH cine SA cine Conv 3 d+BN+ReLU Conv3d+BN+ReLU Conv3d+BN+ReLU Conv3d+BN+ReLU Conv 3 d Conv3d+BN+ReLU Conv3d+BN+ReLUConv3d+BN+ReLU C 32 ResBlock3dResBlock3d C 64 ResBlock3dResBlock3d ResBlock3d C 128 C 128 C 64 ResBlock3dResBlock3dResBlock3d C 256 Conv3d+BN+ReLU ResBlock3d DDDUUU C 32C 1 D DownSample U UpSample Conv3d+BN+ReLU Conv3d+BN ReLU Conv3d+BN+ReLUConv3d+BN+ReLU ResBlock3d ResNet-like Encoder Vision Transformer MLP ResNet-like Encoder ResNet-like Encoder ResNet-like Encoder Vision Transformer MLP Figure 6: The architecture of the proposed expert model for cardiac segmentation and CVDs diagnosis is described as follows: (a) It is a two-stage coarse-to-fine cardiac segmentation framework. In the first stage, a U-shaped backbone locates the ROI of the heart. In the second stage, based on the ROI output from the first stage, the framework fuses the global and local features of the heart to generate a refined, high-resolution cardiac segmentation map. (b) It is a two-stage automated screening and diagnostic framework for CVDs. In the first stage, the model leverages cine images (including 2CH, 4CH, and SAX views) to implement tri-class screening, categorizing subjects into three groups: NH, IHD, or NICM. For patients identified as having NICM via screening, the second stage further integrates their retained 4CH cine images, short-axis cine images, and SAX LGE images to output more refined diagnostic results for NICM subtypes. (c) It represents the overall network structure of the U-shaped backbone. The U-shaped backbone network is based on the ResUNet [50] architecture, and its detailed structure is presented in the bottom section of Fig. 6 (c). The encoder of the U-shaped backbone consists of an initial double-convolution block followed by three residual layers, each containing three residual blocks. Downsampling is achieved using stride-2 convolutions, with the number of feature channels progressively increasing. The decoder restores spatial resolution through three upsampling steps; after each upsampling, the features are concatenated with corresponding encoder features via skip connections and fused using convolutional layers. Each residual block comprises two 3×3× 3 convolutional layers with Batch Normalization (BN); a ReLU activation follows the first convolution, while the second convolution is not activated before addition with the identity pathway. When channel dimensions or resolutions differ, the identity mapping is aligned through a projection branch using a 1×1×1 convolution with BN. The final segmentation output is generated by a 1×1×1 convolutional layer applied to the backbone features. For LGE data, due to its large slice thickness and limited spatial continuity, the above 3D model architecture designed for cine sequences is adapted into a 2D structure to accommodate the characteristics of LGE imaging. Both stages of the cardiac segmentation expert model undergo 100 epochs of training, and the weights corresponding to the optimal performance on the validation set are selected for performance inference. The Adam optimizer is used with an initial learning rate of 5e-4 and a weight decay of 5e-4. A piecewise learning rate schedule with linear warm-up is adopted, where the learning rate is linearly increased from one-third of its initial value to the target value during the first 10 iterations, followed by a decay by a factor of 0.2 at the 5th, 30th, and 60th epochs. 15 Cardiac Agent 4.4 CMR-based expert model for CVDs CMR technology has emerged as the gold standard for non-invasive assessment of cardiac structure and function [51,52], enabling comprehensive visualization of the complete characteristics of myocardial motion and myocardial fibrosis through dynamic cine imaging across multiplanar views and SAX LGE imaging [53]. The CMR-based expert model for CVDs adopts a two-stage architecture, as illustrated in Fig. 6 (b). Owing to the significant differences in the motion characteristics of cardiac contraction among NH, IHD, and NICM, the first stage exclusively utilizes multi-view cine sequences from CMR SAX, long-axis 2CH, and long-axis 4CH views. By integrating motion features extracted from multi-view cine data via deep learning algorithms, this stage achieves automated differentiation among three major categories: NH, IHD, and NICM, thereby avoiding the potential risks associated with contrast agent injection. However, given the complexity of diagnosing NICM, reliable assessment typically requires LGE technology. Thus, the second stage incorporates SAX LGE sequences to improve the diagnostic accuracy of fine-grained classification for NICM. Myocardial fibrosis manifests distinctively across different NICM [54, 55]. In the first stage of the CVDs expert model, the input comprises 2CH, 4CH, and SAX cine sequences, each accompanied by a corresponding segmentation mask. The preprocessing pipeline includes three steps: z-axis cropping, voxel spacing resampling, and xy-plane cropping. Since z-axis cropping is related to cardiac cycle phases, the central three phases are retained for 2CH and 4CH sequences, and the central nine phases for SAX sequences. All data are then resampled to a uniform voxel spacing, defined as the mode spacing of the dataset. The ROI is extracted using the segmentation mask, and images are cropped to a fixed size of 200×200 in the xy-plane centered at the ROI. After this initial preprocessing, a second cropping step is applied: 2CH and 4CH sequences are cropped to 80×192×192, and SAX sequences to 288×144×144. Intensity normalization is then performed, followed by probabilistic color augmentation or noise perturbation. During training, additional sample-wise 3D data augmentations are applied, including random 3D rotation, scaling, translation, flipping, and interpolation, to improve model robustness. To effectively capture modality-specific characteristics from high-dimensional 3D data, separate ResNet-like [56,57] encoders are designed for the 2CH, 4CH, and SAX modalities, with the detailed structure illustrated in the lower half of Fig. 6 (c). Each ResNet-like encoder in the CVDs diagnosis expert model shares the same architecture, consisting of an initial double-convolution block, three residual layers (each with three basic blocks), and a global average pooling layer. Each basic block is a 3D residual unit composed of two 3×3×3 convolutions with BN; the first convolution is followed by a ReLU activation, whereas the second is not activated before addition with the identity mapping. When channel dimensions or resolutions differ, a projection branch with 1×1×1 convolution and BN (optionally with downsampling) is used for alignment. The modality-specific features are concatenated and fused using a Vision Transformer (ViT) [34], and the transformer [58] outputs are mapped to classification logits using a multilayer perceptron (MLP) [59]. The second stage of the CVDs expert model builds upon the first stage by replacing the 2CH cine input with the LGE modality. The preprocessing pipeline for LGE follows that of cine data: nine central slices are retained along the z-axis, and images are cropped to 200×200 in the xy-plane, followed by a second cropping step to 9×144×144. Unlike cine encoders, the LGE encoder does not compress the z-axis dimension to preserve spatial information relevant to scar localization. The two-stage CVDs diagnosis model is trained for 60 epochs using the Adam optimizer with an initial learning rate of 5e-4 and a weight decay of 5e-4. A piecewise learning rate schedule with linear warm-up is employed, where the learning rate is linearly increased from one-third of its initial value to the target value during the first 10 iterations and reduced by a factor of 0.2 at the 5th and 20th epochs. References [1] K. Mc Namara, H. Alzubaidi, and J. K. Jackson. Cardiovascular disease as a leading cause of death: how are pharmacists getting involved? Integrated Pharmacy Research and Practice, 8:1–11, 2019. [2]H. Alhabeeb, M. H. Sohouli, A. Lari, S. Fatahi, F. Shidfar, O. Alomar, and A. Abu-Zaid. Impact of orange juice consumption on cardiovascular disease risk factors: a systematic review and meta-analysis of randomized- controlled trials. Critical Reviews in Food Science and Nutrition, 62(12):3389–3402, 2020. [3]Kenneth Dickstein, Alain Cohen-Solal, Gerasimos Filippatos, John J. V Mcmurray, Piotr Ponikowski, Anna Strömberg, Dirk J Veldhuisen, Dan Atar, Arno W Hoes, and Andre Keren. Esc guidelines for the diagnosis and treatment of acute and chronic heart failure 2008;. European Journal of Heart Failure, 11(1):110–110, 2014. [4]Yan-Ran Wang, Kai Yang, Yi Wen, Pengcheng Wang, Yuepeng Hu, Yongfan Lai, Yufeng Wang, Kankan Zhao, Siyi Tang, Angela Zhang, et al. Screening and diagnosis of cardiovascular disease using artificial intelligence-enabled cardiac magnetic resonance imaging. Nature Medicine, 30(5):1471–1480, 2024. 16 Cardiac Agent [5]Michael Salerno and Christopher M. Kramer. Advances in parametric mapping with cmr imaging. JACC: Cardiovascular Imaging, 6(7):806–822, 2013. [6] Michael Jerosch-Herold. Quantification of myocardial perfusion by cardiovascular magnetic resonance. Journal of Cardiovascular Magnetic Resonance (BioMed Central), 12(1):1–16, 2010. [7]N. I. Bouwer, C. Liesting, Mjm Kofflard, J. J. Brugts, and E. Boersma. 2d-echocardiography vs cardiac mri strain using deep learning: a prospective cohort study in patients with her2-positive breast cancer undergoing trastuzumab. European Heart Journal Cardiovascular Imaging, 22, 2021. [8] El Sayed H. Ibrahim, Luba Frank, Dhiraj Baruah, Jason C. Rubenstein, V. Emre Arpinar, Andrew S. Nencka, Kevin M. Koch, L Tugan Muftuler, Orhan Unal, and Jadranka Stojanovska. Value cmr: Towards a comprehensive, rapid, cost-effective cardiovascular magnetic resonance imaging. Cold Spring Harbor Laboratory Press, 2020. [9]Raymond Kim, Albert De Roos, Eckart Fleck, Charles Higgins, Gerald Pohost, Martin Prince, and Warren Manning. Guidelines for training in cardiovascular magnetic resonance (cmr). J Cardiovasc Magn Reson, 9(1):3–4, 2007. [10]Michael P Hartung, Thomas M Grist, and Christopher J François. Magnetic resonance angiography: current status and future directions. Journal of Cardiovascular Magnetic Resonance, 13(1):19, 2011. [11]Jeanette Schulz-Menger, David A Bluemke, Jens Bremerich, Scott D Flamm, Mark A Fogel, Matthias G Friedrich, Raymond J Kim, Florian von Knobelsdorff-Brenkenhoff, Christopher M Kramer, Dudley J Pennell, et al. Standardized image interpretation and post processing in cardiovascular magnetic resonance: Society for cardiovascular magnetic resonance (scmr) board of trustees task force on standardized post processing. Journal of Cardiovascular Magnetic Resonance, 15(1):35, 2013. [12]Isaac Shiri, Giovanni Baj, Pooya Mohammadi Kazaj, Marius R Bigler, Anselm W Stark, Waldo Valenzuela, Ryota Kakizaki, Matthias Siepe, Stephan Windecker, Lorenz Räber, et al. Ai-based detection and classification of anomalous aortic origin of coronary arteries using coronary ct angiography images. Nature Communications, 16(1):3095, 2025. [13]Junyi Qiu, Lei Li, Sihan Wang, Ke Zhang, Yinyin Chen, Shan Yang, and Xiahai Zhuang. Myops-net: Myocardial pathology segmentation with flexible combination of multi-sequence cmr images. Medical Image Analysis, 84:102694, 2023. [14]Bram Ruijsink, Esther Puyol-Antón, Ilkay Oksuz, Matthew Sinclair, Wenjia Bai, Julia A Schnabel, Reza Razavi, and Andrew P King. Fully automated, quality-controlled cardiac analysis from cmr: validation and large-scale application to characterize cardiac function. Cardiovascular Imaging, 13(3):684–695, 2020. [15] Yan-Ran (Joyce) Wang, Kai Yang, Yi Wen, Pengcheng Wang, Yuepeng Hu, Yongfan Lai, Yufeng Wang, Kankan Zhao, Siyi Tang, Angela Zhang, Huayi Zhan, Minjie Lu, Xiuyu Chen, Shujuan Yang, Zhixiang Dong, Yining Wang, Hui Liu, Lei Zhao, Lu Huang, Yunling Li, Lianming Wu, Zixian Chen, Yi Luo, Dongbo Liu, Pengbo Zhao, Keldon Lin, Joseph C. Wu, and Shihua Zhao. Screening and diagnosis of cardiovascular disease using artificial intelligence-enabled cardiac magnetic resonance imaging. Nature Medicine, 30(5):1471–1480, 2024. [16] Marco Merlo, Kyle Lam, Giulia Gagno, Anna Baritussio, Barbara Bauce, Elena Biagini, Marco Canepa, Alberto Cipriani, Silvia Castelletti, Santo Dellegrottaglie, Andrea Igoren Guaricci, Massimo Imazio, Giuseppe Limongelli, Maria Beatrice Musumeci, Vanda Parisi, Silvia Pica, Gianluca Pontone, Giancarlo Todiere, Camilla Torlasco, Cristina Basso, Gianfranco Sinagra, Pasquale Perrone Filardi, Ciro Indolfi, Camillo Autore, and Andrea Barison. Clinical application of cmr in cardiomyopathies: evolving concepts and techniques. Heart Failure Reviews, 28(1):77–95, 2023. [17]Pasquale Paolisso, Luca Bergamaschi, Francesco Angeli, Marta Belmonte, Alberto Foà, Lisa Canton, Damiano Fedele, Matteo Armillotta, Angelo Sansonetti, Francesca Bodega, Sara Amicone, Nicole Suma, Emanuele Gallinoro, Domenico Attinà, Fabio Niro, Paola Rucci, Elisa Gherbesi, Stefano Carugo, Saima Mushtaq, Andrea Baggiano, Anna Giulia Pavon, Marco Guglielmo, Edoardo Conte, Daniele Andreini, Gianluca Pontone, Luigi Lovato, and Carmine Pizzi. Cardiac magnetic resonance to predict cardiac mass malignancy: The cmr mass score. Circulation: Cardiovascular Imaging, 17(3):e016115, 2024. [18] Cheng-Yi Li, Kao-Jung Chang, Cheng-Fu Yang, Hsin-Yu Wu, Wenting Chen, Hritik Bansal, Ling Chen, Yi-Ping Yang, Yu-Chun Chen, Shih-Pin Chen, et al. Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation. Nature Communications, 16(1):2258, 2025. [19]Vishwesh Nath, Wenqi Li, Dong Yang, Andriy Myronenko, Mingxin Zheng, Yao Lu, Zhijian Liu, Hongxu Yin, Yee Man Law, Yucheng Tang, Pengfei Guo, Can Zhao, Ziyue Xu, Yufan He, Stephanie Harmon, Benjamin Simon, Greg Heinrich, Stephen Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov, Baris Turkbey, Holger Roth, and Daguang Xu. Vila-m3: Enhancing vision-language models with medical expert knowledge. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), pages 14788–14798, June 2025. 17 Cardiac Agent [20]Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. volume 225 of Proceedings of Machine Learning Research, pages 353–367, 2023. [21]Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and-vision assistant for biomedicine in one day, 2023. [22]Xinyue Hu, Lin Gu, Kazuma Kobayashi, Liangchen Liu, Mengliang Zhang, Tatsuya Harada, Ronald M. Summers, and Yingying Zhu. Interpretable medical image visual question answering via multi-modal relationship graph learning. Medical Image Analysis, 97:103279, 2024. [23]Xiao Liang, Di Wang, Haodi Zhong, Quan Wang, Ronghan Li, Rui Jia, and Bo Wan. Candidate-heuristic in-context learning: A new framework for enhancing medical visual question answering with llms. Information Processing & Management, 61(5):103805, 2024. [24]Sachin M Sabariram, Sanjay Vikram C B, Sharon Deborah E, Kavin M, and Saravanan G. Advanced vision- language pipelines: Contextual learning and interactive segmentation for medical imaging. In 2025 International Conference on Machine Learning and Autonomous Systems (ICMLAS), pages 608–615, 2025. [25]Binxu Li, Tiankai Yan, Yuanting Pan, Jie Luo, Ruiyang Ji, Jiayuan Ding, Zhe Xu, Shilong Liu, Haoyu Dong, Zihao Lin, and Yixin Wang. Mmedagent: Learning to use medical tools with multi-modal agent, 2024. [26]Yubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, and Hae Won Park. Mdagents: An adaptive collaboration of llms for medical decision-making. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 79410–79452, 2024. [27]Maria Camila Villa, Isabella Llano, Natalia Castano-Villegas, Julian Martinez, Maria Fernanda Guevara, Jose Zea, and Laura Velásquez. Medsearch: A conversational agent for real-time, evidence-based medical question- answering. Intelligence-Based Medicine, page 100274, 2025. [28] Jianing Qiu, Kyle Lam, Guohao Li, Amish Acharya, Tien Yin Wong, Ara Darzi, Wu Yuan, and Eric J. Topol. Llm-based agentic systems in medicine and healthcare. Nature Machine Intelligence, 6(12):1418–1420, 2024. [29] Nikita Mehandru, Brenda Y. Miao, Eduardo Rodriguez Almaraz, Madhumita Sushil, Atul J. Butte, and Ahmed Alaa. Evaluating large language models as agents in the clinic. npj Digital Medicine, 7(1):84, 2024. [30] Fabian Isensee, Paul F Jaeger, Simon A Kohl, Jens Petersen, and Klaus H Maier-Hein. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods, 18(2):203–211, 2021. [31]Jun Ma, Zongxin Yang, Sumin Kim, Bihui Chen, Mohammed Baharoon, Adibvafa Fallahpour, Reza Asakereh, Hongwei Lyu, and Bo Wang. Medsam2: Segment anything in 3d medical images and videos. arXiv preprint arXiv:2504.03600, 2025. [32]Debesh Jha, Pia H Smedsrud, Michael A Riegler, Dag Johansen, Thomas De Lange, Pål Halvorsen, and Håvard D Johansen. Resunet++: An advanced architecture for medical image segmentation. In 2019 IEEE international symposium on multimedia (ISM), pages 225–2255. IEEE, 2019. [33]Zhaohu Xing, Liang Wan, Huazhu Fu, Guang Yang, Yijun Yang, Lequan Yu, Baiying Lei, and Lei Zhu. Diff-unet: A diffusion embedded network for robust 3d medical image segmentation. Medical Image Analysis, page 103654, 2025. [34]Alexey Dosovitskiy. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [35]Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, pages 5533–5541, 2017. [36] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. [37]Yiming Shi, Shaoshuai Yang, Xun Zhu, Haoyu Wang, Xiangling Fu, Miao Li, and Ji Wu. Medm-vl: What makes a good medical lvlm? In International Workshop on Agentic AI for Medicine, pages 290–299. Springer, 2025. [38] Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36:28541–28564, 2023. [39]Danilo Neglia, Riccardo Liga, Alessia Gimelli, Tomaž Podlesnikar, Marta Cviji ́ c, Gianluca Pontone, Marcelo Haer- tel Miglioranza, Andrea Igoren Guaricci, Sara Seitun, Alberto Clemente, et al. Use of cardiac imaging in chronic coronary syndromes: the eureca imaging registry. European heart journal, 44(2):142–158, 2023. 18 Cardiac Agent [40]Steve R Ommen, Seema Mital, Michael A Burke, Sharlene M Day, Anita Deswal, Perry Elliott, Lauren L Evanovich, Judy Hung, Jose A Joglar, Paul Kantor, et al. 2020 aha/acc guideline for the diagnosis and treatment of patients with hypertrophic cardiomyopathy: a report of the american college of cardiology/american heart association joint committee on clinical practice guidelines. Journal of the American College of Cardiology, 76(25):e159–e240, 2020. [41]Shilong Liu, Hao Cheng, Haotian Liu, Hao Zhang, Feng Li, Tianhe Ren, Xueyan Zou, Jianwei Yang, Hang Su, Jun Zhu, et al. Llava-plus: Learning to use tools for creating multimodal agents. In European conference on computer vision, pages 126–142. Springer, 2024. [42] Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023. [43]Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3, 2022. [44]Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He.Zero-offload: Democratizingbillion-scalemodel training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21), pages 551–564, 2021. [45] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Gins- burg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. Mixed precision training. arXiv preprint arXiv:1710.03740, 2017. [46]Zihao Zhao, Sheng Wang, Jinchen Gu, Yitao Zhu, Lanzhuju Mei, Zixu Zhuang, Zhiming Cui, Qian Wang, and Dinggang Shen. Chatcad+: Toward a universal and reliable interactive cad using llms. IEEE Transactions on Medical Imaging, 43(11):3755–3766, 2024. [47]Salim S Virani, L Kristin Newby, Suzanne V Arnold, Vera Bittner, LaPrincess C Brewer, Susan Halli Demeter, Dave L Dixon, William F Fearon, Beverly Hess, Heather M Johnson, et al. 2023 aha/acc/accp/aspc/nla/pcna guideline for the management of patients with chronic coronary disease: a report of the american heart associ- ation/american college of cardiology joint committee on clinical practice guidelines. Journal of the American College of Cardiology, 82(9):833–955, 2023. [48] Guy De Backer, Ettore Ambrosioni, Knut Borch-Johnsen, Carlos Brotons, Renata Cifkova, Jean Dallongeville, Shah Ebrahim, Ole Faergeman, Ian Graham, Giuseppe Mancia, et al. European guidelines on cardiovascular disease prevention in clinical practice; third joint task force of european and other societies on cardiovascular disease prevention in clinical practice (constituted by representatives of eight societies and by invited experts). European Journal of Cardiovascular Prevention & Rehabilitation, 10(4):S1–S10, 2003. [49] Elena Arbelo, Alexandros Protonotarios, Juan R Gimeno, Eloisa Arbustini, Roberto Barriales-Villa, Cristina Basso, Connie R Bezzina, Elena Biagini, Nico A Blom, Rudolf A De Boer, et al. 2023 esc guidelines for the management of cardiomyopathies: developed by the task force on the management of cardiomyopathies of the european society of cardiology (esc). European heart journal, 44(37):3503–3626, 2023. [50]Foivos I Diakogiannis, François Waldner, Peter Caccetta, and Chen Wu. Resunet-a: A deep learning framework for semantic segmentation of remotely sensed data. ISPRS Journal of Photogrammetry and Remote Sensing, 162:94–114, 2020. [51] Amrit Chowdhary, Pankaj Garg, Arka Das, Muhummad Sohaib Nazir, and Sven Plein. Cardiovascular magnetic resonance imaging: emerging techniques and applications. Heart, 107(9):697–704, 2021. [52]Antonio Luca Maria Parlati, Ermanno Nardi, Federica Marzano, Cristina Madaudo, Mariafrancesca Di Santo, Ciro Cotticelli, Simone Agizza, Giuseppe Maria Abbellito, Fabrizio Perrone Filardi, Mario Del Giudice, et al. Advancing cardiovascular diagnostics: the expanding role of cmr in heart failure and cardiomyopathies. Journal of Clinical Medicine, 14(3):865, 2025. [53]Albert J Rogers, Neal K Bhatia, Sabyasachi Bandyopadhyay, James Tooley, Rayan Ansari, Vyom Thakkar, Justin Xu, Jessica Torres Soto, Jagteshwar S Tung, Mahmood I Alhusseini, et al. Identification of cardiac wall motion abnormalities in diverse populations by deep learning of the electrocardiogram. npj Digital Medicine, 8(1):21, 2025. [54]Bianca Olivia Cojan-Minzat, Alexandru Zlibut, and Lucia Agoston-Coldea. Non-ischemic dilated cardiomyopathy and cardiac fibrosis. Heart Failure Reviews, 26(5):1081–1101, 2021. [55]Carla Giordano, Marco Francone, Giulia Cundari, Annalinda Pisano, and Giulia d’Amati. Myocardial fibrosis: morphologic patterns and role of imaging in diagnosis and prognostication. Cardiovascular pathology, 56:107391, 2022. 19 Cardiac Agent [56]Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. [57] Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Learning spatio-temporal features with 3d residual networks for action recognition. In Proceedings of the IEEE international conference on computer vision workshops, pages 3154–3160, 2017. [58]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. [59] Marius-Constantin Popescu, Valentina E Balas, Liliana Perescu-Popescu, and Nikos Mastorakis. Multilayer perceptron and neural networks. WSEAS Transactions on Circuits and Systems, 8(7):579–588, 2009. [60]Christopher M Kramer, Jörg Barkhausen, Chiara Bucciarelli-Ducci, Scott D Flamm, Raymond J Kim, and Eike Nagel. Standardized cardiovascular magnetic resonance imaging (cmr) protocols: 2020 update. Journal of Cardiovascular Magnetic Resonance, 22(1):17, 2020. Data availability IRB approval was obtained from all participating institutions for imaging and data collection: Beijing Anzhen Hospital, China (2025216x). The need for informed consent was waived by the respective ethics committees and institutions. The benchmark test dataset for multi-modal large model evaluation has been publicly released, with access available at https://huggingface.co/datasets/TaipingQu/CMRAgentEvalSet. We have uploaded all pixel-level multi-sequence CMR annotations to https://huggingface.co/datasets/TaipingQu/CMR-MULTI. No other publicly available datasets were used in this study. The deidentified data can be shared only for noncommercial academic purposes and will require a formal material transfer agreement and a data use agreement. Requests should be submitted by emailing the corresponding authors. All requests will be evaluated based on institutional policies to determine whether the data requested are subject to intellectual property or patient privacy obligations. Code availability An open-source version of the code base is available on GitHub at https://github.com/plantain-herb/Cardiac-Agent with no restrictions. Acknowledgements This work was supported by the grants from the National Key R&D Program of China (2022YFE0209800), received by L. Xu, the National Natural Science Foundation of China (82271986, U1908211), received by L. Xu, the Beijing Natural Science Foundation (7244326 &L246062), received by HK. Zhang and L. Xu. Author contributions T. Qu, H.K. Zhang, L. Xu and H.G. Zhang conceptualized the study collaboratively. T. Qu, L. Zhang, C. Zhao, M. Zou, P. Zhao, H. Liu, and Z. Su were responsible for cardiac magnetic resonance (CMR) data preprocessing, large model algorithm development and training, expert model training, as well as comprehensive analysis and interpretation of model performance. K. Bo, X. Jin, H.K. Zhang, Z. Zhou, N. Zhang, H. Wang, K. Jiang, Y. Du, and M. Wang primarily undertook the collection, annotation and quality control of clinical CMR data, and completed the clinical efficacy evaluation of the model. R. Yan coordinated the acquisition of external validation CMR data. Z. Wang and T. Huang provided overall project coordination, technical platform support and funding guarantee for the entire study. L. Xu offered professional clinical guidance, coordinated multi-center clinical resources, and conducted rigorous review of all clinical-related content in the study. L. Zhang and C. Zhao contributed equally and are co-second authors. T. Qu, H.K. Zhang, L. Zhang, and C. Zhao drafted the initial manuscript and prepared the key figures. L. Xu and H.G. Zhang were responsible for the overall study design and academic guidance, manuscript finalization and revision, as well as all communication work with the journal. All authors provided critical feedback on the study design and manuscript, and ensured the scientific integrity of this work. Competing interests The authors declare no competing interests. 20 Cardiac Agent Appendix A Study Cohort Statistics A.1 CMR Acquisition All patients underwent CMR scanning with a 32-channel phased array coil under respiratory navigation and elec- trocardiographic gating in the supine position. The scanning equipment was three 3.0 T CMR whole-body scanner: Achieva, Philips, Netherlands, Holland; Discovery MR750w, GE Healthcare, USA; Siemens Healthineers AG, Erlangen, Germany. The cine images were acquired using retrospective electrocardiogram (ECG) gating, and the standardized imaging protocol included steady-state free precession (SSFP) breath-holding cine images obtained at the end of inspiration. These images cover short-axis (SAX) and long-axis views; the latter specifically includes two-chamber (2CH) and four-chamber (4CH) views. The protocol also incorporates late gadolinium enhancement (LGE) images [60]. Briefly, SAX cardiac images were acquired to cover the entire left ventricle (LV), from the mitral valve annulus to the level of the LV apex, with 25–30 phases acquired per cardiac cycle and a slice thickness of 8 m. For long-axis views, images were acquired with a slice thickness of 5 m, encompassing 2CH, 4CH, and three-chamber (3CH) views. LGE images were acquired 10–15 minutes after the intravenous injection of gadolinium-diethylenetriamine pentaacetic acid (Gd-DTPA; dose: 0.2 mmol/kg; Bayer Pharma AG, Berlin, Germany), using a prospectively ECG-gated breath-hold phase-sensitive inversion recovery segmented gradient-echo sequence. A.2 Study Cohort We retrospectively collect CMR scan sequences from 2,134 patients (1,558 male and 576 female) at Beijing Anzhen Hospital, Capital Medical University (Beijing, China), with corresponding CMR imaging reports available for all patients. To ensure model reliability, the enrollment period at this center is extended to January 1, 2016, to January 1, 2024, for rare cardiovascular diseases (RCM and ACM). Additionally, we collect CMR scan sequences from 279 patients at the First Affiliated Hospital of Xinxiang Medical University (Henan, China), covering the period from October 1, 2023, to October 15, 2025. As a national-level diagnosis and treatment center for cardiovascular diseases, Beijing Anzhen Hospital provides primary national representative sample support for the dataset. The baseline CMR scan of each examination in both datasets includes SAX cine sequences, 4CH cine sequences, 2CH cine sequences, Rest MPI, and SAX LGE sequences. Internal DatasetExternal Dataset NMaleFemaleAgeNMaleFemaleAge Total21341558 (73%)576 (27%)47±20279171 (61%)108 (39%)49±17 NH322190 (59%)132 (41%)22±176726 (39%)41 (61%)43±21 IHD805676 (84%)129 (16%)57±124937 (76%)12 (24%)56±12 NICM830564 (68%)266 (32%)48±18159107 (67%)52 (33%)49±16 HCM402269 (67%)133 (33%)51±164332 (74%)11 (26%)51±14 DCM136101 (74%)35 (26%)49±169162 (68%)29 (32%)52±15 RCM9559 (62%)36 (38%)54±2064 (67%)2 (33%)44±12 ACM6646 (69%)20 (31%)43±1785 (63%)3 (37%)50±22 Myocarditis13185 (65%)46 (35%)23±17114 (36%)7 (64%)26±17 Others177115 (65%)62 (35%)46±1941 (25%)3 (75%)63±16 Table D1: Characteristics of the internal and external validation datasets (N: Number of subjects, Age in years: mean±SD (range)) Appendix B Model Evaluation Metrics To assess the CMR segmentation performance, we employ three primary evaluation metrics. The Dice Similarity Coefficient (DSC), calculated as DSC = 2|V X ∩V Y | |V X | +|V Y | ,(4) quantifies the volumetric overlap between the predicted segmentationV X and the reference ground truthV Y . Let S X andS Y represent the voxel collections lying on the boundaries of the predicted segmentation and ground truth, 21 Cardiac Agent respectively. The Hausdorff Distance (HD) is defined as HD = max max x∈S X min y∈S Y ||x− y||, max y∈S Y min x∈S X ||y− x|| ,(5) where||.||denotes the Euclidean distance. This metric is used to evaluate boundary alignment. Additionally, we use the Average Surface Distance (ASD), a distance-centric metric computed as ASD = 1 |S Y | ( X y∈S Y min x∈S X ||y− x||),(6) where|.|denotes the cardinality of a set. For distance-based metrics, smaller values correspond to superior segmentation outcomes. To comprehensively quantify the diagnostic performance of the model for CVDs, this study adopts accuracy, sensitivity, specificity, F 1 -score, and AUC as the core evaluation metrics. All metrics are calculated based on the confusion matrix constructed from diagnostic results, with their specific definitions, formulas, and implications presented below. The core parameters of the confusion matrix are defined as follows: true positive (TP), true negative (TN), false positive (FP), and false negative (FN). The core calculation formulas are as follows. Accuracy = T P + T N T P + T N + F P + F N (7) Sensitivity = T P T P + F N (8) Specif icity = T N T N + F P (9) P recision = T P T P + F P (10) F 1 -score = 2× P recision× Sensitivity P recision + Sensitivity (11) Appendix C CMR Data Annotation CMR annotations were performed independently by two radiologists(H.K.Z and N.Z) with no less than 10 years of clinical experience in CMR diagnosis, using the open-source software ITK-SNAP (w.itksnap.org) for precise contour delineation. The annotation scope for each sequence is as follows: the SAX cine sequence includes the LV cavity, LV myocardium, and right ventricle (RV); the 4CH cine sequence covers the LV cavity, LV myocardium, RV cavity, RV myocardium, left atrium (LA), and right atrium (RA); the 2CH cine sequence involves the LV cavity and LV myocardium; and the SAX LGE sequence encompasses the LV cavity, LV myocardium, and LGE regions. After annotation completion, a senior radiologist(L.X, more than 20 years cardiac MRI experience) verifies and validates the annotation quality of all data to ensure accuracy and consistency. Appendix D CMR Report Scoring Criteria The CMR report scoring system has a total of 100 points and comprises two modules: clinical diagnostic accuracy (70 points) and technical performance (30 points). The core of clinical diagnostic accuracy is assessing the consistency between model output and the clinical gold standard. Specifically, clinical diagnosis (20 points) evaluates the clarity and correctness of cardiac disease diagnosis; core structural quantification (15 points) involves score deductions based on the errors between automatically measured values including end-diastolic volume, end-systolic volume and ejection fraction and the clinical report text. Full marks are given when the error is less than 5%, with 3 points deducted for an error ranging from 5% to 10%, 5 points for 10% to 20%, and 7 points for more than 20%. Wall assessment (10 points), myocardial viability detection via LGE (15 points) and other feature assessment (10 points) respectively evaluate the integrity of key information detection for the corresponding items, with 1 point deducted for each missing key piece of information. The other features cover indicators such as cine valve status, cardiac size, ventricular wall motion, pericardial effusion and myocardial perfusion assessment. For the technical performance module, only report completeness and structuring (30 points) is assessed, which judges whether the overall CMR report complies with clinical norms. Two points are deducted for each piece of redundant information and also for each missing key clinical 22 Cardiac Agent indicator. In addition, severe score deductions ranging from 0 to -10 points are imposed for risky behaviors such as model hallucinations, including unfounded speculation and fabrication of non-existent information. The deductions are graded in accordance with major factual hallucinations, mild factual hallucinations or over-interpretation, and logically conflicting hallucinations. MetricMethodLV MyoLV CavityRV Cavity ResUNet++91.29±0.0385.14±0.0477.64±0.08 nnUNet88.19±0.0492.44±0.0381.63 ± 0.05 DSCDiffUNet83.37±0.1680.34±0.1167.88±0.12 MedSAM278.60±0.2882.95±0.3181.02±0.07 Ours92.44±0.0396.19±0.0282.00±0.05 ResUNet++9.67±5.999.01±3.4731.07±24.36 nnUNet8.04±2.917.13±2.3225.87±29.31 HD DiffUNet16.44±8.05184.23±53.8076.23±62.99 MedSAM221.24±6.7018.45±7.1918.48±7.51 Ours6.49±2.015.83±1.4812.41±5.45 ResUNet++0.67±0.890.38±0.131.75±1.39 nnUNet0.26±0.090.29±0.141.71±2.33 ASDDiffUNet1.41±1.552.28±2.084.26±4.99 MedSAM24.87±1.610.59±0.291.90±1.11 Ours0.28±0.090.14±0.091.51±0.79 Table D2: A comparative study between the BAAI Cardiac Agent segmentation expert model and SOTA segmentation networks on the SAX cine dataset. We report DSC (%), HD, and ASD as mean ± standard deviation. MetricMethodLV MyoLV Cavity ResUNet++83.39±0.0590.56±0.03 nnUNet85.49±0.0390.79±0.04 DSCDiffUNet78.92±0.0987.17±0.07 MedSAM287.17±0.0468.61±0.40 Ours84.95±0.0592.54±0.03 ResUNet++13.84±10.8210.71±4.28 nnUNet29.25±51.5730.51±50.16 HDDiffUNet189.56±64.4245.15±58.17 MedSAM27.42±2.8511.02±7.12 Ours8.43±3.577.31±2.75 ResUNet++0.55±0.350.73±0.51 nnUNet2.14±6.543.30±9.56 ASDDiffUNet8.05±10.252.88±6.25 MedSAM20.18±0.110.32±0.27 Ours0.29±0.170.25±0.18 Table D3: A comparative analysis of the BAAI Cardiac Agent segmentation expert model against current SOTA segmentation networks using the 2CH cine dataset. Results are presented as mean ± standard deviation for DSC (%), HD, and ASD. 23 Cardiac Agent Metric MethodLV MyoLV CavityRV MyoRV CavityLA CavityRA Cavity ResUNet++83.75±0.0492.94±0.0362.17±0.1488.43±0.0491.37±0.0487.82±0.06 nnUNet82.34±0.0692.23±0.0357.64±0.1485.39±0.0486.69±0.1185.37±0.09 DSCDiffUNet78.31±0.0688.02±0.0754.81±0.1579.29±0.1186.32±0.1077.51±0.17 MedSAM286.57±0.0362.87±0.4470.54±0.0790.26±0.0291.61±0.0590.54±0.04 Ours86.12±0.0594.25±0.0268.77±0.0989.93±0.0392.23±0.0490.22±0.05 ResUNet++ 15.81±22.46 16.43±25.82 10.61±4.4110.45±4.797.27±4.5221.81±33.01 nnUNet13.41±7.649.89±4.7614.22±9.7012.21±6.5812.32±6.8610.96±6.91 HDDiffUNet114.55±85.95 70.62±81.97 67.94±54.54 45.69±51.56 88.73±80.80 50.46±58.70 MedSAM27.08±3.0610.76±7.359.22±4.569.18±3.725.99±2.245.43±2.14 Ours7.34±3.596.22±2.489.63±4.788.25±3.726.66±3.107.00±3.10 ResUNet++0.61±0.941.25±2.460.26±0.180.47±0.340.35±0.292.74±6.64 nnUNet0.37±0.250.49±0.410.57±0.750.89±1.890.93±1.030.79±1.10 ASDDiffUNet8.91±11.346.73±8.009.41±12.277.90±8.8311.12±14.445.48±7.72 MedSAM20.18±0.090.31±0.240.19±0.100.21±0.090.36±0.350.30±0.18 Ours0.29±0.140.20±0.140.32±0.270.26±0.130.25±0.270.23±0.24 Table D4: A comparative study of the BAAI Cardiac Agent segmentation expert model versus SOTA segmentation networks on the 4CH cine dataset. We report DSC (%), HD, and ASD in the format of mean ± standard deviation. MetricMethodLV MyoLV CavityLGE ResUNet++66.40±0.1281.72±0.0926.13±0.26 nnUNet74.12±0.1087.36±0.0660.19±0.21 DSCDiffUNet53.09±0.1558.85 ± 0.2529.88±0.19 MedSAM270.07±0.2772.27±0.3465.37±0.23 Ours73.67±0.0888.60±0.0560.83±0.13 ResUNet++13.13±5.609.88±4.3526.68±14.25 nnUNet12.31±6.079.03±5.1416.71±9.12 HDDiffUNet194.77±55.08157.69±79.11185.30±44.29 MedSAM212.78±5.809.05±4.0914.90±9.82 Ours11.20±5.147.89±3.6024.17±13.85 ResUNet++0.60±0.351.15±0.652.29±4.64 nnUNet0.52±0.240.59±0.420.78±0.58 ASDDiffUNet48.98±59.5248.62±44.8644.15±38.37 MedSAM20.56±0.390.80±0.842.18±4.17 Ours0.61±0.270.54±0.482.42±2.43 Table D5: A comparative analysis of the BAAI Cardiac Agent segmentation expert model and leading segmentation networks evaluated on the SAX LGE dataset. Performance metrics (DSC (%), HD, and ASD) are reported as mean ± standard deviation. 24 Cardiac Agent 2CH cine Labels Ours MedSAM2 nnUNet DiffUNetResUNet++ 4CH cine SAX cine SAX LGE Images 96.64% 93.29% 93.79% 91.97% 85.36% 91.59% 89.09% 79.04% 87.63% 75.14% 94.82% 95.53% 91.40% 83.97% 91.30% 81.73% 49.86% 63.34% 54.09% 61.04% Fig. D1: Qualitative comparison examples of segmentation results between the segmentation expert model we proposed and existing SOTA methods. From top to bottom, four image examples for each of 2CH cine, 4CH cine, SAX cine, and SAX LGE are presented in sequence, with the labeled values indicating DSC results. (Best viewed in color) 25 Cardiac Agent Image Mask Bullseye Prediction Ground Truth Fig. D2: Cumulative LGE degree in the LV 17-SM: BAAI Cardiac Agent vs physician annotations. (Best viewed in color) 26 Cardiac Agent AUC Internal Test Set OursVSTViTResNet NH0.980 (0.956-0.995)0.962 (0.944-0.977)0.978 (0.968-0.996)0.980 (954-0.995) IHD0.938 (0.913-0.960)0.927 (0.901-0.950)0.917 (0.882-0.948)0.930 (0.904-0.953) NICM0.960 (0.942-0.974)0.942 (0.920-0.961)0.942 (0.920-0.963)0.934 (0.908-0.957) HCM0.975 (0.948-0.992)0.970 (0.953-0.993)0.965 (0.941-0.982)0.969 (0.957-0.995) DCM0.971 (0.945-0.991)0.969 (0.942-0.990)0.977 (0.953-0.995)0.971 (0.944-0.990) RCM0.961 (0.925-0.988)0.962 (0.927-0.989)0.957 (0.916-0.989)0.975 (0.918-0.992) ACM0.958 (0.912-0.992)0.959 (0.822-0.990)0.951 (0.898-0.990)0.945 (0.900-0.983) Myocarditis0.938 (0.868-0.989)0.929 (0.882-0.974)0.927 (0.869-0.986)0.841 (0.785-0.887) External Validation Set OursVSTViTResNet NH0.933 (0.903-0.960)0.920 (0.890-0.947)0.891 (0.849-0.929)0.916 (0.899-0.958) IHD0.827 (0.767-0.885)0.830 (0.746-0.889)0.782 (0.693-0.861)0.805 (0.742-0.864) NICM0.817 (0.770-0.862)0.802 (0.774-0.868)0.773 (0.713-0.825)0.794 (0.740-0.849) HCM0.907 (0.850-0.953)0.881 (0.833-0.938)0.884 (0.838-0.941)0.903 (0.866-0.951) DCM0.867 (0.809-0.921)0.859 (0.832-0.915)0.822 (0.767-0.891)0.830 (0.806-0.906) RCM0.877 (0.793-0.939)0.873 (0.789-0.950)0.915 (0.843-0.964)0.880 (0.776-0.959) ACM0.902 (0.783-0.977)0.911 (0.817-0.969)0.921 (0.756-0.997)0.900 (0.745-0.984) Myocarditis0.862 (0.698-0.984)0.917 (0.857-0.949)0.829 (0.683-0.944)0.937 (0.856-0.995) Table D6: The comparative results of CVDs diagnostic performance between the cardiac specialist diagnostic model and three SOTA models are presented in the form of AUC (95% CI). 27 Cardiac Agent Sensitivity (Specificity=0.9) Internal Test Set OursVSTViTResNet NH0.938 (0.887-1.000)0.923 (0.873-1.000)0.923 (0.876-1.000)0.953 (0.879-1.000) IHD0.839 (0.753-0.900)0.776 (0.669-0.858)0.820 (0.710-0.890)0.739 (0.605-0.870) NICM0.850 (0.736-0.924)0.750 (0.677-0.849)0.800 (0.704-0.889)0.793 (0.712-0.854) HCM0.889 (0.822-0.985)0.931 (0.814-0.987)0.877 (0.799-0.947)0.901 (0.813-1.000) DCM0.885 (0.793-1.000)0.846 (0.682-1.000)0.885 (0.774-1.000)0.808 (0.725-1.000) RCM0.889 (0.750-1.000)0.778 (0.650-1.000)0.778 (0.631-1.000)0.873 (0.725-1.000) ACM0.769 (0.556-1.000)0.764 (0.500-1.000)0.752 (0.462-1.000)0.758 (0.483-1.000) Myocarditis0.808 (0.652-0.957)0.769 (0.594-0.939)0.778 (0.644-0.957)0.792 (0.630-0.964) External Validation Set OursVSTViTResNet NH0.746 (0.571-0.898)0.791 (0.658-0.932)0.687 (0.541-0.815)0.881 (0.717-0.984) IHD0.531 (0.340-0.744)0.449 (0.256-0.610)0.325 (0.271-0.575)0.429 (0.214-0.564) NICM0.654 (0.412-0.719)0.516 (0.389-0.637)0.428 (0.159-0.571)0.566 (0.446-0.675) HCM0.721 (0.556-0.870)0.587 (0.497-0.760)0.667 (0.468-0.835)0.714 (0.675-0.821) DCM0.714 (0.551-0.813)0.758 (0.649-0.849)0.703 (0.527-0.828)0.717 (0.649-0.843) RCM0.333 (0.141-0.733)0.333 (0.153-0.616)0.500 (0.214-0.865)0.500 (0.189-0.821) ACM0.625 (0.250-1.000)0.625 (0.167-1.000)0.750 (0.599-1.000)0.625 (0.285-1.000) Myocarditis0.727 (0.499-1.000)0.727 (0.556-1.000)0.455 (0.125-0.833)0.712 (0.462-1.000) Table D7: The comparative results of CVDs diagnostic performance between the cardiac specialist diagnostic model and three SOTA models are presented in the form of Sensitivity (95% CI). 28 Cardiac Agent Specificity (Sensitivity=0.9) Internal Test Set OursVSTViTResNet NH0.954 (0.905-1.000)0.953 (0.900-1.000)0.950 (0.903-1.000)0.964 (0.901-1.000) IHD0.838 (0.722-0.917)0.824 (0.717-0.880)0.828 (0.711-0.900)0.819 (0.726-0.891) NICM0.862 (0.808-0.933)0.818 (0.740-0.878)0.769 (0.696-0.897)0.756 (0.692-0.841) HCM0.904 (0.836-0.989)0.940 (0.859-1.000)0.902 (0.860-0.989)0.900 (0.831-1.000) DCM0.946 (0.821-0.985)0.891 (0.830-0.985)0.938 (0.855-0.993)0.930 (0.882-0.984) RCM0.942 (0.870-0.985)0.934 (0.813-0.980)0.861 (0.745-0.986)0.963 (0.887-0.993) ACM0.845 (0.775-1.000)0.845 (0.775-1.000)0.880 (0.739-1.000)0.880 (0.548-1.000) Myocarditis0.853 (0.603-1.000)0.713 (0.618-1.000)0.752 (0.569-1.000)0.767 (0.580-1.000) External Validation Set OursVSTViTResNet NH0.861 (0.771-0.912)0.865 (0.817-0.924)0.740 (0.665-0.809)0.899 (0.854-0.946) IHD0.589 (0.339-0.768)0.513 (0.277-0.758)0.392 (0.246-0.598)0.580 (0.371-0.745) NICM0.405 (0.307-0.514)0.517 (0.400-0.643)0.483 (0.357-0.600)0.552 (0.438-0.689) HCM0.698 (0.607-0.875)0.653 (0.598-0.755)0.685 (0.589-0.841)0.628 (0.523-0.766) DCM0.600 (0.242-0.816)0.632 (0.323-0.806)0.647 (0.486-0.821)0.544 (0.386-0.781) RCM0.863 (0.697-0.930)0.771 (0.548-0.961)0.588 (0.483-0.968)0.830 (0.643-0.961) ACM0.815 (0.575-0.981)0.874 (0.447-0.954)0.954 (0.442-1.000)0.815 (0.642-0.948) Myocarditis0.607 (0.397-0.986)0.704 (0.540-1.000)0.547 (0.299-0.913)0.851 (0.583-1.000) Table D8: The comparative results of CVDs diagnostic performance between the cardiac specialist diagnostic model and three SOTA models are presented in the form of Specificity (95% CI). 29 Cardiac Agent (a) (b) Fig. D3: Consistency evaluation between the BAAI Cardiac Agent and clinical reports. (a) Bland-Altman plots compare cardiac output (CO), right ventricular end-diastolic diameter (RVEDD), and left/right atrial maximum transverse diameters (LAT4CHD, RAT4CHD) measured by the BAAI Cardiac Agent with manual clinical reports in the internal cohort; (b) Mean left ventricular end-diastolic wall thickness (LVEDWT) distributions from manual 17-segment model (17-SM) reports and the BAAI Cardiac Agent in internal normal and hypertrophic cardiomyopathy cohorts. The bullseye plot sequentially maps left ventricular basal, mid, and apical segments to the plot’s outer, middle, and inner layers, with the center representing the apex. Cardiac Magnetic Resonance Imaging Report (Multi Sequence) IMAGING FINDINGS: 4CH cine measurement:The maximum transverse diameters of the left and right atria (LAT4CHD, RAT4CHD) measured in the 4CH view are 47.09 m and 39.64 m, respectively. SAX cine measurement: The left ventricular end-diastolic diameter is approximately 75.25m. Left ventricular end diastolic wall thickness: •Basal segment: Anterior wall (AW) 6.94 m, anteroseptal wall (ASW) 7.76 m, inferoseptalwall (ISW) 7.13 m, inferolateral wall (ILW) 7.79 m, inferior wall (IW) 7.29 m, anterolateral wall (ALW) 6.06 m •Mid segment: Anterior wall (AW) 6.24 m, anteroseptal wall (ASW) 7.04 m, inferoseptalwall (ISW) 7.52 m, inferolateral wall (ILW) 6.44 m, inferior wall (IW) 7.66 m, anterolateral wall (ALW) 6.11 m •Apical segment: Anterior wall (AW) 5.16 m, septal wall (SW) 5.84 m, inferior wall (IW) 5.31 m, lateral wall (LW) 5.36 m •Apex: 2.89 m Left ventricular function measurement: EF: 20.43%, EDV: 358.94ml, ESV: 285.60ml, SV: 73.34ml, C0: 8.21L/min, LV mass:195.63g. LGE imaging: Striped delayed enhancement is observed in the mid-myocardium, involving the anterior wall, interventricular septum, and inferior wall of the left ventricle, extending from the basal to mid-ventricular levels. DIAGNOSTIC CONCLUSION: The first stage diagnosis result:conforms to the phenotype of non-ischemic cardiomyopathy. Second stage diagnosis result: Consistent with dilated cardiomyopathyphenotype. Enlargement of left atrium and left ventricle; Reduced left ventricular wall motion and decreased systolic function; Fibrosis in the middle layer of the myocardium between the basal and central segments of theleft ventricular wall, as well as the anterior and inferior walls of the interventricular septum. EXAMINATION METHODS: SAX cine, 4CH cine, 2CH cine, SAX LGE Fig. D4: Schematic diagram of a complete medical report generated by the BAAI Cardiac Agent, illustrating the measurement of cardiac function parameters and imaging diagnostic analysis for a patient with dilated cardiomyopathy (DCM). 30 Cardiac Agent Which tool can help me identify and segment the cardiovascular structures in this cine 4ch scan? Input Text The CMR tool can help you identify and segment the cardiovascular structures in this cine 4ch scan. This tool is designed to assist in the analysis of cardiovascular images, making it easier to visualize and segment the different structures within the heart. Based on this cardiac imaging study, please provide a detailed medical report. Using the Medical Report Generation output, here is the complete diagnostic report. Anteroseptal Wall (ASW) : 17.54m Anterior Wall (AW) : 13.92m ... LV Ejection Fraction (LVEF) : 51.83% ... Left ventricular wall thickening (Average basal wall thickness: 15.0m) The cardiac imaging study, which includes cardiac MRI and cardiac CT , shows a normal cardiac structure . This means that the heart's chambers, valves, and surrounding structures appear to be functioning properly and without any significant abnormalities. What type of cardiomyopathy is shown in this image: normal, ischemic, or non - ischemic? Based on the output from Cardiomyopathy Classification, this patient is classified as Non - ischemic Cardiomyopath y . The image shows a normal cardiomyopathy , which means that the heart muscle appears to be healthy and functioning properly. Which sequence does this cardiac MRI belong to? This cardiac MRI is a cine sequence SAX view . This cardiac MRI belongs to the T2 - weighted sequence. Where is the main issue located? Following detailed evaluation, the patient was found to have left ventricular wall thickening with impaired wall motion . . . A small amount of fluid signal is visible within the pericardial cavity, suggesting possible pericardial effusion.. The main issue is located in the left ventricle of the heart. What are the primary treatment options for severe symptomatic aortic stenosis? Please provide the evidence. Treatment for severe aortic stenosis mainly involves valve replacement, either through surgical aortic valve replacement (SAVR) or transcatheter aortic valve implantation (TAVI). ... [ https:// w.acc.org /Clinical - Topics/Cardiac - Surgery/Aortic - Valve - Replacement ] The primary treatment options for severe symptomatic aortic stenosis are aortic valve replacement (AVR) ... These options are supported by evidence from multiple clinical trials, including the Partnership in the Evaluation of the Effectiveness of Transcatheter Aortic Valve Implantation (PARTNER) trials ... LLaVa - Med BAAI Cardiac Agent Input 2D Images Input 3D Images To identify and segment cardiovascular structures in a cine 4 - chamber (4ch) MRI scan like the one shown, you can use a combination of ** automated software tools ** and ** manual segmentation techniques **, depending on your expertise and needs. ... ... **Findings:** \ n Chronic ischemic cardiomyopathy; moderate LV systolic dysfunction (EF~35 - 40%), LAD territory prior MI (30 - 40% LV scar), regional wall motion abnormalities, mild RV/LV dilation & LA enlargement, mild tricuspid regurgitation... No acute process or infiltrative disease ... The **left ventricle appears dilated** with **reduced wall thickness** and **impaired contractility**. ... These findings are **typical of non - ischemic cardiomyopathy**, such as ** dilated cardiomyopathy (DCM) **Answer: Non - ischemic cardiomyopathy **. The provided cardiac MRI images appear to be ** cine MRI ** (also known as **real - time or dynamic MRI**) sequences, specifically ** balanced steady - state free precession ( bSSFP )** images. ... Based on the provided cardiac MRI images, the main issue appears to be located in the **left ventricle (LV)**, ... Answer: The main issue is located in the anterior wall of the left ventricle.**" Severe symptomatic aortic stenosis (AS) is a life - threatening condition that requires timely intervention. The primary treatment options are: ... Reference: 2023 ACC/AHA/SCAI Guideline for the Management of Patients With Aortic Stenosis. Journal of the American College of Cardiology, 2023. ... Qwen - VL - 30B The Cardiovascular MRI Toolbox can help you identify and segment the cardiovascular structures in a cine 4ch scan. This toolbox is designed to assist in the analysis of cardiac MRI images by providing tools for segmentation . ... The cardiac MRI image shows a large mass in the right atrium ... the mass is causing moderate right atrial and right ventricular dilatation ... These findings suggest that the mass may be affecting the normal functioning of the heart , ... The image shows non - ischemic cardiomyopathy . This cardiac MRI belongs to the steady - state free precession (SSFP) sequence . The main issue appears to be located in the left ventricle of the heart. The primary treatment options for severe symptomatic aortic stenosis include aortic valve replacement (AVR) and aortic valve repair. ... The evidence supporting these treatment options can be found in various clinical guidelines and studies, such as the European Society of Cardiology (ESC) ... MedM The segmented results are displayed below, providing valuable insights for your medical evaluation. Based on the result of ChatCAD - G, ... FINDINGS: The MRI scan of the brain shows a large mass located in the left temporal lobe ... IMPRESSION: Calcification and hemorrhagic changes in the left temporal lobe tumor . Based on the result of BiomedClip , this image represents normal cardiomyopathy . Based on the results from BiomedClip , this cardiac MRI image belongs to the "short - axis" sequence . The main issue is located in the left ventricle of the heart. Based on the provided information, the primary treatment options for severe symptomatic aortic stenosis include medications such as ACE inhibitors, ARBs, and ARNI... . ( https:// w.merckmanuals.com /profession al/cardiovascular - disorders/valvular - disorders/aortic - stenosis ) MMedAgent The applied tool has effectively delineated the cardiovascular anatomy , enhancing the diagnostic value of this image. Fig. D5: Qualitative comparison between BAAI Cardiac Agent and other SOTA methods across different tasks. Undesired and desired responses are highlighted inlicorice yellow and pine green, respectively. 31