Paper deep dive
No Image, No Problem: End-to-End Multi-Task Cardiac Analysis from Undersampled k-Space
Yundi Zhang, Sevgi Gokce Kafali, Niklas Bubeck, Daniel Rueckert, Jiazhen Pan
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/13/2026, 1:06:54 AM
Summary
The paper introduces k-MTR (k-space Multi-Task Representation), a framework that enables end-to-end cardiac analysis (phenotype regression, disease classification, and segmentation) directly from undersampled k-space data. By aligning undersampled k-space and fully-sampled images into a shared semantic manifold using contrastive learning, the model bypasses the traditional, ill-posed image reconstruction step, achieving competitive performance with image-domain baselines.
Entities (5)
Relation Signals (3)
k-MTR → bypasses → Image Reconstruction
confidence 100% · k-MTR, the first representation learning approach to align undersampled k-space and spatial images into a shared semantic manifold, entirely bypassing the image reconstruction.
k-MTR → performs → Cardiac Analysis
confidence 95% · k-MTR establishes a unified new paradigm for frequency-domain analysis, achieving highly competitive performance across continuous phenotype regression, disease classification, and fine-grained anatomical segmentation.
k-MTR → trainedon → UK Biobank
confidence 95% · Leveraging a large-scale controlled simulation of 42,000 subjects [from the UK Biobank]
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Conventional clinical CMR pipelines rely on a sequential "reconstruct-then-analyze" paradigm, forcing an ill-posed intermediate step that introduces avoidable artifacts and information bottlenecks. This creates a fundamental mathematical paradox: it attempts to recover high-dimensional pixel arrays (i.e., images) from undersampled k-space, rather than directly extracting the low-dimensional physiological labels actually required for diagnosis. To unlock the direct diagnostic potential of k-space, we propose k-MTR (k-space Multi-Task Representation), a k-space representation learning framework that aligns undersampled k-space data and fully-sampled images into a shared semantic manifold. Leveraging a large-scale controlled simulation of 42,000 subjects, k-MTR forces the k-space encoder to restore anatomical information lost to undersampling directly within the latent space, bypassing the explicit inverse problem for downstream analysis. We demonstrate that this latent alignment enables the dense latent space embedded with high-level physiological semantics directly from undersampled frequencies. Across continuous phenotype regression, disease classification, and anatomical segmentation, k-MTR achieves highly competitive performance against state-of-the-art image-domain baselines. By showcasing that precise spatial geometries and multi-task features can be successfully recovered directly from the k-space representations, k-MTR provides a robust architectural blueprint for task-aware cardiac MRI workflows.
Tags
Links
- Source: https://arxiv.org/abs/2603.09945v1
- Canonical: https://arxiv.org/abs/2603.09945v1
Trouble viewing inline? Open PDF directly →
Full Text
27,753 characters extracted from source content.
Expand or collapse full text
11institutetext: Chair for AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Munich, Germany 22institutetext: Department of Computing, Imperial College London, United Kingdom 33institutetext: Munich Center for Machine Learning, Technical University of Munich, Germany 33email: yundi.zhang, jiazhen.pan@tum.de No Image, No Problem: End-to-End Multi-Task Cardiac Analysis from Undersampled k-Space Yundi Zhang Sevgi Gokce Kafali Niklas Bubeck Daniel Rueckert Jiazhen Pan Abstract Conventional clinical CMR pipelines rely on a sequential "reconstruct-then-analyze" paradigm, forcing an ill-posed intermediate step that introduces avoidable artifacts and information bottlenecks. This creates a fundamental mathematical paradox: it attempts to recover high-dimensional pixel arrays (i.e., images) from undersampled k-space, rather than directly extracting the low-dimensional physiological labels actually required for diagnosis. To unlock the direct diagnostic potential of k-space, we propose k-MTR (k-space Multi-Task Representation), a k-space representation learning framework that aligns undersampled k-space data and fully-sampled images into a shared semantic manifold. Leveraging a large-scale controlled simulation of 42,000 subjects, k-MTR forces the k-space encoder to restore anatomical information lost to undersampling directly within the latent space, bypassing the explicit inverse problem for downstream analysis. We demonstrate that this latent alignment enables the dense latent space embedded with high-level physiological semantics directly from undersampled frequencies. Across continuous phenotype regression, disease classification, and anatomical segmentation, k-MTR achieves highly competitive performance against state-of-the-art image-domain baselines. By showcasing that precise spatial geometries and multi-task features can be successfully recovered directly from the k-space representations, k-MTR provides a robust architectural blueprint for task-aware cardiac MRI workflows. 111Code and models will be made publicly available upon publication. 1 Introduction In Cardiac Magnetic Resonance (CMR) imaging, the sensor data (k-space) and the spatial information (i.e., images) are dual representations of the same anatomical and physiological reality, mathematically linked by the Fourier Transform. Due to hardware limits, high costs, and requirements for breath-hold by patients, standard cardiac MRI protocols typically undersample k-space to accelerate acquisitions [3, 24, 19]. Thus, the conventional research community has viewed the relationship between these domains primarily through the lens of image reconstruction. The standard objective has been to recover high-fidelity images from undersampled k-space, which inevitably introduces avoidable information bottlenecks and reconstruction artifacts [22, 7]. While deep learning has driven significant progress in image reconstruction [20, 17, 11], explicit reconstruction serves merely as an intermediate step. The ultimate clinical goal of CMR is to extract physiological features, clinical phenotypes, and disease labels [5, 16, 15]. This creates a paradox in the standard clinical pipeline. Reconstructing a fully sampled image from undersampled k-space is fundamentally an ill-posed problem because it attempts to recover high-dimensional variables (image) from low-dimensional inputs (k-space). However, the downstream clinical classification and phenotyping tasks are inherently dimensionality-reduction problems. Extracting a low-dimensional disease label from undersampled k-space is mathematically much more well-posed than reconstructing the entire image array. This realization prompts a critical question: Can cardiac analysis operate directly within the undersampled k-space? Recent studies have begun investigating end-to-end approaches that predict single clinical labels or segmentations directly from undersampled k-space [21, 14, 27]. While promising, these efforts remain largely restricted to isolated, single-task applications. Concurrently, cardiac multi-task foundation models [25, 13, 26] have demonstrated the power of dense latent spaces in the image domain. Yet, the field lacks a unified framework that successfully bridges k-space, spatial images, and clinical labels through a shared dense manifold. To address this gap, we propose k-MTR (k-space Multi-Task Representation), an end-to-end framework that constructs a comprehensive, information-rich manifold by aligning the undersampled k-space and image domains. We hypothesize that undersampled k-space retains sufficient information to directly inform clinical endpoints. By bypassing the reconstruction step, this unified manifold enables the diverse cardiac analysis tasks from undersampled measurements. Our key contributions are: • The first k-space representation learning framework beyond reconstruction. We introduce k-MTR, the first representation learning approach to align undersampled k-space and spatial images into a shared semantic manifold, entirely bypassing the image reconstruction. • Information-dense semantic manifold. We demonstrate that this domain alignment creates an information-dense latent space, implicitly compensating for anatomical structures degraded by undersampling while preserving critical diagnostic semantics. • Direct multi-task cardiac analysis from undersampled k-space. k-MTR establishes a unified new paradigm for frequency-domain analysis, achieving highly competitive performance across continuous phenotype regression, disease classification, and fine-grained anatomical segmentation. Figure 1: Overview of the k-MTR framework. (a) S multi-view 2D+t2D+t slices are tokenized, concatenating image (QiuQ_i^u) and k-space (QkuQ_k^u) tokens across slices for encoder input. (b–d) Training pipeline (single slice shown). (b) Unsupervised masked reconstruction of undersampled k-space and image slices. (c) Contrastive alignment between undersampled k-space (TkuT^u_k) and fully-sampled image (TiT_i) embeddings. (d) Fine-tuning the pretrained k-space encoder via lightweight task-specific decoders. 2 k-MTR Methodology The three-stage pipeline of k-MTR for effective k-space feature extraction is shown in Fig. 1. Let fully-sampled complex-valued cardiac k-space measurements be Xk∈ℂS×T×H×WX_k ^S× T× H× W, where S,T,HS,T,H, and W represent the number of slices, time frames, height, and width, respectively. The slices, comprising both 2D+t2D+t long-axis (LAX) and short-axis (SAX) views, are concatenated along the slice dimension to yield the corresponding image stack Xi∈ℂS×T×H×WX_i ^S× T× H× W. To simulate standard clinical acquisition constraints, we apply a Cartesian acceleration mask M∈ℤS×T×WM ^S× T× W along the phase-encoding (W) direction. The mask is broadcast along the H dimension to obtain M~ M, yielding the undersampled k-space data via element-wise multiplication: Xku=M~⊙XkX^u_k= M X_k. The superscripts i, k, and u denote the image domain, k-space domain, and undersampled data, respectively. Stage I: Domain-Specific Representation Learning. In the first stage, we independently learn robust domain-specific features using the masked autoencoder (MAE) paradigm [9]. For the image domain, multi-view 2D+t2D+t images are randomly masked at the patch level to obtain XiuX^u_i. For the frequency domain, the k-space data relies on the predefined clinical acceleration mask M~ M to yield XkuX^u_k. Each domain employs its own tokenizer T, encoder ℰE, decoder D, and mask tokens TmT^m. With ⊕ denoting concatenation, the reconstruction objectives are formulated as: X^i=i(ℰi(Qiu)⊕Tim),X^k=k(ℰk(Qku)⊕Tkm), X_i=D_i(E_i(Q^u_i) T^m_i), X_k=D_k(E_k(Q^u_k) T^m_k), (1) where Qiu=i(Xiu)Q_i^u=T_i(X_i^u) and Qku=k(Xku)Q_k^u=T_k(X_k^u) denote the tokenized inputs. Real and imaginary components are processed as two input channels, and multi-slice tokens are concatenated along the sequence-length dimension. This ensures that both encoders establish domain-specific semantic capacities, preparing for subsequent cross-modal alignment. Stage I: Cross-Modal Alignment and Latent Restoration. We establish a shared latent space to align image and k-space representations. A critical design choice is the asymmetry of the inputs: image representations Ti=ℰi(i(Xi))T_i=E_i(T_i(X_i)) are extracted from fully-sampled data to preserve complete semantic geometry, while k-space representations Tku=ℰk(k(Xku))T_k^u=E_k(T_k(X^u_k)) are derived solely from undersampled data. It explicitly forces the k-space encoder ℰkE_k to recover and embed the anatomical signatures omitted by the undersampling mask directly into its latent vector, implicitly bypassing the inverse problem in the latent space. We project class tokens via domain-specific projectors iP_i and kP_k to obtain Z^i=i(Ti) Z_i=P_i(T_i) and Z^k=k(Tku) Z_k=P_k(T_k^u). With temperature τ and loss weight λ, a symmetric contrastive loss ℒL is applied over a batch of subjects N to minimize the distance of multi-modal embeddings of the same subject: ℓi,k=−∑m∈logexp(cos(zmi,zmk)/τ)∑n∈,n≠mexp(cos(zmi,znk)/τ),ℒ=λℓi,k+(1−λ)ℓk,i. _i,k=- _m ( (z_m_i,z_m_k )/τ ) _n ,n≠ m ( (z_m_i,z_n_k )/τ ), =λ _i,k+(1-λ) _k,i. (2) This stage unifies the modalities, embedding image-domain clinical features directly into the aligned manifold. Stage I: End-to-End Analysis from Undersampled k-Space. In the final stage, the pretrained k-space encoder ℰkE_k is fine-tuned to perform downstream clinical tasks using only undersampled k-space measurements XkuX^u_k. We attach lightweight, task-specific decoders to predict phenotype regression, disease classification, and anatomical segmentation from the information-rich latent space. To empirically validate the geometric completeness of the restored latent space, we also evaluate k-MTR’s reconstruction ability. By attaching an adaptive image-domain decoder directly to the k-space encoder (similar to AUTOMAP [29]), k-MTR maps undersampled frequencies to fully sampled images without an explicit inverse Fourier transform. 3 Dataset and Implementation Datasets. The open source literature lacks a foundation-scale type acquired k-space data with labels and annotations. To validate k-space learning, we simulate 42,000 2D+t2D+t cardiac MRI scans (6 SAX, 3 LAX; 128×128×50128× 128× 50) from the UK Biobank [18]. We introduce a synthetic phase via Gaussian-smoothed B0B_0 field variation [4] before applying a discrete Fourier transform. We apply a Cartesian undersampling mask [1] across all tasks. Specifically, we use acceleration factors of R=2,4,8,R=2,4,8, and 1616 for phenotype prediction; R=4R=4 for both classification and reconstruction; and a more aggressive R=8R=8 for segmentation, in order to thoroughly evaluate the capability of k-MTR under highly undersampled conditions. Evaluation uses a strictly held-out 1,000-subject test set. For downstream targets, we utilize 12 continuous phenotypes derived from quality-controlled segmentation maps [2], alongside three disease classifications (coronary artery disease (CAD), high blood pressure, hypertension) following [26]. Implementation. Models are implemented in PyTorch on an NVIDIA A100. Stage I MAEs use 6 encoder and 2 decoder layers (1024-dim), a (5,8,8)(5,8,8) patch size, 70% masking, and batch size 2. Stage I employs gradient checkpointing for a contrastive batch size of 256, mapping 1025-dim tokens to 128-dim via two-layer MLPs (τ=0.1,λ=0.5τ=0.1,λ=0.5). In Stage I, the pretrained encoder is fully fine-tuned with task-specific heads: 256-dim MLPs for regression (batch 16) and classification (batch 32), and a 576-dim UNETR [28] for segmentation (batch 2). The reconstruction decoder mirrors Stage I, trained from scratch (batch 2). Baselines. To isolate the efficacy of k-MTR’s latent restoration, we evaluate robust baselines. For regression, we compare against ResNet-50 [10], ViT [8], and MAE [25]. ResNet-50 is tested on fully-sampled images (upper bound), artifact-corrupted undersampled images (R=4) (ResNet-50u), and undersampled zero-filled k-space (ResNet-50uk_k^u). Comparing the unaligned counterpart, MAEuk_k^u (which omits Stages I-I), against k-MTR highlights the necessity of our cross-modal alignment. For classification, we use image-trained ViT and MAE. Segmentation baselines include fully-sampled nnU-Net [12], undersampled nnU-Netu, and LI-Net [21] which is designed for corrupted undersampled images. 4 Results Figure 2: 3D t-SNE visualization of the representations after alignment, colored by ground-truth phenotype groups. Semantic Feature Clustering in Latent Space. We first evaluate k-MTR’s representational quality to empirically validate the restored latent space. As shown in Fig. 2, we visualize the latent embeddings of 10000 subjects using 3D t-SNE to evaluate semantic coherence. When color-coded by key cardiac phenotypes (LVEDV, LVM, and RVEDV), the embeddings form meaningful clusters. This separation confirms that k-MTR successfully captures complex spatial and temporal phenotypic variability directly from the undersampled k-space. Competitive Clinical Downstream Fidelity Direct from Undersampled k-Space. We then evaluate k-MTR’s ability to extract physiologically meaningful labels natively from undersampled frequencies. Table 1: Mean absolute error ↓ for phenotype prediction. Image-based baselines are in gray. RN50: ResNet-50. Best are in dark green, second and third in lighter green. Phenotype Fully-sampled Undersampled R=4 RN50 ViT MAE RN50u RN50uk_k^u MAEuk_k^u k-MTR LVEDV (mL) 6.586.58 11.5311.53 9.039.03 7.357.35 9.579.57 12.8812.88 8.158.15 LVSV (mL) 5.685.68 6.936.93 5.865.86 6.556.55 7.037.03 8.428.42 6.506.50 LVEF (%) 2.952.95 3.913.91 3.203.20 3.303.30 3.583.58 6.096.09 3.143.14 LVM (g) 5.515.51 7.857.85 5.875.87 7.017.01 8.008.00 10.0910.09 7.207.20 RVEDV (mL) 9.249.24 13.7513.75 9.499.49 10.4910.49 11.3311.33 14.7614.76 10.6010.60 RVESV (mL) 6.156.15 8.408.40 6.076.07 7.357.35 7.637.63 8.718.71 6.986.98 RVSV (mL) 7.587.58 9.519.51 7.427.42 8.068.06 8.418.41 9.649.64 7.867.86 RVEF (%) 3.303.30 4.174.17 3.623.62 3.563.56 3.663.66 5.975.97 3.283.28 LASV (mL) 4.664.66 5.555.55 4.104.10 5.415.41 5.845.84 5.785.78 5.345.34 LAEF (%) 4.284.28 5.245.24 4.294.29 4.824.82 5.745.74 6.716.71 4.724.72 RASV (mL) 5.805.80 7.447.44 5.475.47 7.087.08 8.018.01 7.407.40 7.077.07 RAEF (%) 5.095.09 6.096.09 5.055.05 5.555.55 6.536.53 6.566.56 5.745.74 Table 2: Disease classification performance. Fully-sampled image baselines are in gray; k-MTR with undersampled k-space. Positive class ratios are shown as percentages of the cohort. AP: average precision. Best results are bold, second underlined. Disease Method AUC-ROC↑ F1 Score↑ Recall↑ Precision↑ AP↑ CAD (7.4%) ViT 0.6820.682 0.1760.176 0.842 0.0980.098 0.1260.126 MAE 0.708¯ 0.708 0.215¯ 0.215 0.2000.200 0.233 0.154¯ 0.154 k-MTR (R=4) 0.737 0.282 0.500¯ 0.500 0.197¯ 0.197 0.234 High Blood Pressure (25.8%) ViT 0.6770.677 0.4610.461 0.6520.652 0.357¯ 0.357 0.3550.355 MAE 0.691¯ 0.691 0.469¯ 0.469 0.680¯ 0.680 0.358 0.369¯ 0.369 k-MTR (R=4) 0.697 0.474 0.781 0.3400.340 0.386 Hypertension (20.8%) ViT 0.6980.698 0.4060.406 0.638¯ 0.638 0.2980.298 0.360¯ 0.360 MAE 0.717 0.432 0.724 0.308¯ 0.308 0.378 k-MTR (R=4) 0.710¯ 0.710 0.417¯ 0.417 0.6030.603 0.319 0.3560.356 Phenotype Prediction: As shown in Tab. 1, k-MTR closely approaches the performance of image-based upper bound operating on fully-sampled data. The ability to achieve comparable accuracy to the image-domain baseline on key metrics (e.g., LVEDV and LVEF) provides strong empirical evidence that aligned latent space of k-MTR retains essential physiological semantics. Crucially, k-MTR yields substantial improvements over the unaligned MAEuk_k^u baseline (which omits Stage I alignment). This confirms that cross-modal contrastive learning successfully embeds localized anatomical awareness into the k-space encoder. Disease classification: As shown in Tab. 2), k-MTR performs on par with fully-sampled image-based ViT and MAE, notably achieving an AUC of 0.7370.737 for CAD. Reaching diagnostic equivalence without requiring explicit image reconstruction highlights the semantic richness of the restored latent space. Figure 3: Segmentation and Reconstruction Results. (a) Dice Score and example segmentation maps overlayed on fully-sampled images. k-MTR (R=8) is compared against the fully-sampled upper bound (nnU-Net) and undersampled (R=8) image-based baselines (nnU-Netu, LI-Net). (b) Image reconstructions from undersampled k-space compared to k-GIN. Segmentation: k-MTR exhibits high-precision performance when operating directly on undersampled k-space tokens, achieving an average foreground Dice score of 0.85 at an acceleration factor of R = 8. In contrast, LI-Net struggles to extract reliable semantic information from corrupted images (Fig. 3(a)). Validation Check via Reconstruction: To assess geometric integrity, we map the undersampled k-space embeddings back to the spatial domain. k-MTR shows on-par results (38.18 dBdB PSNR) compared to a reconstruction-specific method k-GIN [17] model (38.30 dBdB), as shown in Fig. 3 (b). By mapping undersampled k-space to image domain directly without IFT, k-MTR learns to implicitly restores omitted anatomical geometries within its manifold. Performance Robustness Across Acceleration Factors. To assess the extensibility of k-MTR, we evaluate phenotype regression performance under varying k-space undersampling ratios (2x, 4x, 8x, and 16x). As shown in Fig. 4, performance exhibits a degradation as information loss increases. While k-MTR maintains stable and competitive accuracy under moderate regimes (2x to 8x), the extreme 16x setting causes prediction failures for specific phenotypes (red points). This confirms that the 4-fold setting adopted in the main experiments is a representative and practically relevant choice rather than a limiting case. Figure 4: Ablation of k-MTR’s phenotype prediction accuracy across varying undersampling factors, with 16x prediction failures highlighted in red. 5 Conclusion In this work, we introduce k-MTR, a framework that successfully extracts physiologically meaningful representations from undersampled k-space. By establishing a shared frequency-spatial latent space, k-MTR explicitly restores omitted anatomical geometries directly within its manifold, bypassing explicitly solving the inverse problem for cardiac downstream analysis. This cross-modal alignment enables highly competitive, multi-task performance across phenotype regression, disease classification, and anatomical segmentation directly from undersampled k-space. Future work will extend k-MTR to prospectively acquired multi-coil datasets and systematically evaluate its robustness across diverse sampling patterns and acceleration factors. Existing open-source k-space datasets (e.g., OCMR [6], CMRxRecon [23]) lack the clinical annotations necessary for downstream tasks beyond image reconstruction. We plan to annotate these datasets and acquire additional in-house data to validate and expand k-MTR’s potential in a multi-coil setting. We hope this work encourages the community to release diverse, well-annotated multi-coil datasets and provides a robust architectural blueprint for task-aware cardiac MRI workflows that operate directly on undersampled frequency measurements. credits 5.0.1 Acknowledgements This research has been conducted using the UK Biobank Resource under Application Number 87802. This work is funded by the European Research Council (ERC) project Deep4MI (884622). Dr. Sevgi Gokce Kafali has been sponsored by the Alexander von Humboldt Foundation. References [1] R. Ahmad, H. Xue, S. Giri, Y. Ding, J. Craft, and O. P. Simonetti (2015) Variable density incoherent spatiotemporal acquisition (vista) for highly accelerated cardiac mri. Magnetic resonance in medicine 74 (5), p. 1266–1278. Cited by: §3. [2] W. Bai, M. Sinclair, G. Tarroni, O. Oktay, M. Rajchl, G. Vaillant, A. M. Lee, N. Aung, E. Lukaschuk, M. M. Sanghvi, et al. (2018) Automated cardiovascular magnetic resonance image analysis with fully convolutional networks. Journal of cardiovascular magnetic resonance 20 (1), p. 65. Cited by: §3. [3] D. A. Bluemke, J. L. Boxerman, E. Atalar, and E. R. McVeigh (1997) Segmented k-space cine breath-hold cardiovascular mr imaging: part 1. principles and technique.. AJR. American journal of roentgenology 169 (2), p. 395–400. Cited by: §1. [4] M. A. Brown and R. C. Semelka (1999) MRI: basic principles and applications. Willey-Liss. Cited by: §3. [5] C. Chen, C. Qin, H. Qiu, G. Tarroni, J. Duan, W. Bai, and D. Rueckert (2020) Deep learning for cardiac image segmentation: a review. Frontiers in cardiovascular medicine 7, p. 25. Cited by: §1. [6] C. Chen, Y. Liu, P. Schniter, M. Tong, K. Zareba, O. Simonetti, L. Potter, and R. Ahmad (2020) OCMR (v1. 0)–open-access multi-coil k-space dataset for cardiovascular magnetic resonance imaging. arXiv preprint arXiv:2008.03410. Cited by: §5. [7] M. Dohmen, M. A. Klemens, I. M. Baltruschat, T. Truong, and M. Lenga (2025) Similarity and quality metrics for mr image-to-image translation. Scientific Reports 15 (1), p. 3853. Cited by: §1. [8] A. Dosovitskiy (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §3. [9] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick (2022) Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 16000–16009. Cited by: §2. [10] K. He, X. Zhang, S. Ren, and J. Sun (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 770–778. Cited by: §3. [11] W. Huang, V. Spieker, S. Xu, G. Cruz, C. Prieto, J. A. Schnabel, K. Hammernik, T. Kuestner, and D. Rueckert (2025) Subspace implicit neural representations for real-time cardiac cine mr imaging. In International Conference on Information Processing in Medical Imaging, p. 168–183. Cited by: §1. [12] F. Isensee, P. F. Jaeger, S. A. Kohl, J. Petersen, and K. H. Maier-Hein (2021) NnU-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18 (2), p. 203–211. Cited by: §3. [13] A. J. Jacob, I. Borgohain, T. Chitiboi, P. Sharma, D. Comaniciu, and D. Rueckert (2025) Towards a cmr foundation model for multi-task cardiac image analysis. Journal of Cardiovascular Magnetic Resonance, p. 101967. Cited by: §1. [14] R. Li, J. Pan, Y. Zhu, J. Ni, and D. Rueckert (2024) Classification, regression and segmentation directly from k-space in cardiac mri. In International Workshop on Machine Learning in Medical Imaging, p. 31–41. Cited by: §1. [15] Z. Liu, K. Kainth, A. Zhou, T. W. Deyer, Z. A. Fayad, H. Greenspan, and X. Mei (2024) A review of self-supervised, generative, and few-shot deep learning methods for data-limited magnetic resonance imaging segmentation. NMR in Biomedicine 37 (8), p. e5143. Cited by: §1. [16] C. Martin-Isla, V. M. Campello, C. Izquierdo, Z. Raisi-Estabragh, B. Baeßler, S. E. Petersen, and K. Lekadir (2020) Image-based cardiac diagnosis with machine learning: a review. Frontiers in cardiovascular medicine 7, p. 1. Cited by: §1. [17] J. Pan, S. Shit, Ö. Turgut, W. Huang, H. B. Li, N. Stolt-Ansó, T. Küstner, K. Hammernik, and D. Rueckert (2023) Global k-space interpolation for dynamic mri reconstruction using masked image modeling. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 228–238. Cited by: §1, §4. [18] S. E. Petersen, P. M. Matthews, J. M. Francis, M. D. Robson, and et al. (2015) UK Biobank’s cardiovascular magnetic resonance protocol. JCMR, p. 1–7. Cited by: §3. [19] S. Plein and S. Kozerke (2021) Are we there yet? the road to routine rapid cmr imaging. Vol. 14, American College of Cardiology Foundation Washington DC. Cited by: §1. [20] J. Schlemper, J. Caballero, J. V. Hajnal, A. N. Price, and D. Rueckert (2017) A deep cascade of convolutional neural networks for dynamic mr image reconstruction. IEEE transactions on Medical Imaging 37 (2), p. 491–503. Cited by: §1. [21] J. Schlemper, O. Oktay, W. Bai, D. C. Castro, J. Duan, C. Qin, J. V. Hajnal, and D. Rueckert (2018) Cardiac mr segmentation from undersampled k-space using deep latent representation learning. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 259–267. Cited by: §1, §3. [22] M. Seitzer, G. Yang, J. Schlemper, O. Oktay, T. Würfl, V. Christlein, T. Wong, R. Mohiaddin, D. Firmin, J. Keegan, et al. (2018) Adversarial and perceptual refinement for compressed sensing mri reconstruction. In International conference on medical image computing and computer-assisted intervention, p. 232–240. Cited by: §1. [23] C. Wang, J. Lyu, S. Wang, C. Qin, K. Guo, X. Zhang, X. Yu, Y. Li, F. Wang, J. Jin, et al. (2024) CMRxRecon: a publicly available k-space dataset and benchmark to advance deep learning for cardiac mri. Scientific Data 11 (1), p. 687. Cited by: §5. [24] Y. Wang, R. Watts, I. R. Mitchell, T. D. Nguyen, J. W. Bezanson, G. W. Bergman, and M. R. Prince (2001) Coronary mr angiography: selection of acquisition window of minimal cardiac motion with electrocardiography-triggered navigator cardiac motion prescanning—initial results. Radiology 218 (2), p. 580–585. Cited by: §1. [25] Y. Zhang, C. Chen, S. Shit, S. Starck, D. Rueckert, and J. Pan (2024) Whole heart 3d+ t representation learning through sparse 2d cardiac mr images. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 359–369. Cited by: §1, §3. [26] Y. Zhang, P. Hager, C. Liu, S. Shit, C. Chen, D. Rueckert, and J. Pan (2025) Towards cardiac mri foundation models: comprehensive visual-tabular representations for whole-heart assessment and beyond. Medical Image Analysis 106, p. 103756. External Links: ISSN 1361-8415, Document, Link Cited by: §1, §3. [27] Y. Zhang, N. Stolt-Ansó, J. Pan, W. Huang, K. Hammernik, and D. Rueckert (2024) Direct cardiac segmentation from undersampled k-space using transformers. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), p. 1–4. Cited by: §1. [28] L. Zhou, H. Liu, J. Bae, J. He, D. Samaras, and P. Prasanna (2023) Self pre-training with masked autoencoders for medical image classification and segmentation. In 2023 IEEE 20th international symposium on biomedical imaging (ISBI), p. 1–6. Cited by: §3. [29] B. Zhu, J. Z. Liu, S. F. Cauley, B. R. Rosen, and M. S. Rosen (2018) Image reconstruction by domain-transform manifold learning. Nature 555 (7697), p. 487–492. Cited by: §2.