Paper deep dive
OrthoDiffusion: A Generalizable Multi-Task Diffusion Foundation Model for Musculoskeletal MRI Interpretation
Tian Lan, Lei Xu, Zimu Yuan, Shanggui Liu, Jiajun Liu, Jiaxin Liu, Weilai Xiang, Hongyu Yang, Dong Jiang, Jianxin Yin, Dingyu Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 1:58:32 PM
Summary
The paper introduces OrthoDiffusion, a unified diffusion-based foundation model for multi-task musculoskeletal MRI interpretation. Pre-trained self-supervised on 15,948 unlabeled knee MRI scans using three orientation-specific 3D diffusion branches (sagittal, coronal, axial), the model supports anatomical segmentation and multi-label diagnosis. It demonstrates robust performance across clinical centers and MRI field strengths, high label efficiency (performing well with only 10% labeled data), and strong cross-anatomical generalization to ankle and shoulder joints.
Entities (14)
Relation Signals (10)
OrthoDiffusion → uses → Diffusion Model
confidence 95% · OrthoDiffusion is a unified diffusion-based foundation model
OrthoDiffusion → supportstask → Anatomical Segmentation
confidence 93% · support diverse clinical tasks, including anatomical segmentation
OrthoDiffusion → supportstask → Multi-label Diagnosis
confidence 93% · support diverse clinical tasks, including ... multi-label diagnosis
OrthoDiffusion → pretrainedon → Knee MRI
confidence 92% · pre-trained in a self-supervised manner on 15,948 unlabeled knee MRI scans
OrthoDiffusion → generalizesto → Shoulder
confidence 90% · transferable to other joints, achieving strong diagnostic performance across 11 diseases of the ankle and shoulder
OrthoDiffusion → generalizesto → Ankle
confidence 90% · transferable to other joints, achieving strong diagnostic performance across 11 diseases of the ankle
MPAE → fuses → Sagittal Plane
confidence 88% · MPAE operates by treating the sagittal, coronal, and axial branches as three independent experts
MPAE → fuses → Coronal Plane
confidence 88% · MPAE operates by treating the sagittal, coronal, and axial branches as three independent experts
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Musculoskeletal disorders represent a significant global health burden and are a leading cause of disability worldwide. While MRI is essential for accurate diagnosis, its interpretation remains exceptionally challenging. Radiologists must identify multiple potential abnormalities within complex anatomical structures across different imaging planes, a process that requires significant expertise and is prone to variability. We developed OrthoDiffusion, a unified diffusion-based foundation model designed for multi-task musculoskeletal MRI interpretation. The framework utilizes three orientation-specific 3D diffusion models, pre-trained in a self-supervised manner on 15,948 unlabeled knee MRI scans, to learn robust anatomical features from sagittal, coronal, and axial views. These view-specific representations are integrated to support diverse clinical tasks, including anatomical segmentation and multi-label diagnosis. Our evaluation demonstrates that OrthoDiffusion achieves excellent performance in the segmentation of 11 knee structures and the detection of 8 knee abnormalities. The model exhibited remarkable robustness across different clinical centers and MRI field strengths, consistently outperforming traditional supervised models. Notably, in settings where labeled data was scarce, OrthoDiffusion maintained high diagnostic precision using only 10\% of training labels. Furthermore, the anatomical representations learned from knee imaging proved highly transferable to other joints, achieving strong diagnostic performance across 11 diseases of the ankle and shoulder. These findings suggest that diffusion-based foundation models can serve as a unified platform for multi-disease diagnosis and anatomical segmentation, potentially improving the efficiency and accuracy of musculoskeletal MRI interpretation in real-world clinical workflows.
Tags
Links
- Source: https://arxiv.org/abs/2602.20752v1
- Canonical: https://arxiv.org/abs/2602.20752v1
Trouble viewing inline? Open PDF directly →
Full Text
97,005 characters extracted from source content.
Expand or collapse full text
OrthoDiffusion: A Generalizable Multi-Task Diffusion Foundation Model for Musculoskeletal MRI Interpretation Tian Lan 1† , Lei Xu 2,3,4† , Zimu Yuan 2,3,4† , Shanggui Liu 2,3,4 , Jiajun Liu 2,3,4 , Jiaxin Liu 2,3,4 , Weilai Xiang 5,6 , Hongyu Yang 5,7 , Dong Jiang 2,3,4* , Jianxin Yin 1* , Dingyu Wang 2,3,4* 1 Center for Applied Statistics and School of Statistics, Renmin University of China, Beijing, China. 2 Department of Sports Medicine, Peking University Third Hospital, Institute of Sports Medicine of Peking University, Beijing, China. 3 Beijing Key Laboratory of Research and Translation for Drugs and Medical Devices in Precision Diagnosis and Treatment of Sports Injuries, Beijing, China. 4 Engineering Research Center of Sports Trauma Treatment Technology and Devices, Ministry of Education, Beijing, China. 5 State Key Laboratory of Virtual Reality Technology and Systems, Beihang University, Beijing, China. 6 School of Computer Science and Engineering, Beihang University, Beijing, China. 7 Institute of Artificial Intelligence, Beihang University, Beijing, China. *Corresponding author(s). E-mail(s): wangdingyu@pku.edu.cn; jyin@ruc.edu.cn; bysyjiangdong@126.com; † These authors contributed equally to this work. Abstract Musculoskeletal disorders represent a significant global health burden and are a leading cause of disability worldwide. While magnetic resonance imaging (MRI) is essential for accurate diagnosis, its interpretation remains exceptionally challeng- ing. Radiologists must identify multiple potential abnormalities within complex anatomical structures across different imaging planes, a process that requires significant expertise and is prone to variability. To address these limitations, we 1 arXiv:2602.20752v1 [cs.CV] 24 Feb 2026 developed OrthoDiffusion, a unified diffusion-based foundation model designed for multi-task musculoskeletal MRI interpretation. The framework utilizes three orientation-specific 3D diffusion models, pre-trained in a self-supervised manner on 15,948 unlabeled knee MRI scans, to learn robust anatomical features from sagittal, coronal, and axial views. These view-specific representations are inte- grated to support diverse clinical tasks, including anatomical segmentation and multi-label diagnosis. Our evaluation demonstrates that OrthoDiffusion achieves excellent performance in the segmentation of 11 knee structures and the detec- tion of 8 knee abnormalities. The model exhibited remarkable robustness across different clinical centers and MRI field strengths, consistently outperforming tra- ditional supervised models. Notably, in settings where labeled data was scarce, OrthoDiffusion maintained high diagnostic precision using only 10% of training labels. Furthermore, the anatomical representations learned from knee imaging proved highly transferable to other joints, achieving strong diagnostic perfor- mance across 11 diseases of the ankle and shoulder. These results highlight the model’s capacity for cross-anatomical generalization without the need for extensive joint-specific retraining. These findings suggest that diffusion-based foundation models can serve as a unified platform for multi-disease diagnosis and anatomical segmentation, potentially improving the efficiency and accuracy of musculoskeletal MRI interpretation in real-world clinical workflows. Keywords: Musculoskeletal MRI, Diffusion models, Self-supervised representation learning, Multi-plane fusion, Multi-label diagnosis, Anatomical segmentation, Cross-anatomical generalization, Label-efficient learning, Unified imaging framework 1 Introduction Musculoskeletal disorders represent a leading cause of disability worldwide, affect- ing approximately 1.71 billion of patients [1–4]. Magnetic Resonance Imaging (MRI) serves as the gold-standard for non-invasive evaluation, providing high-fidelity soft- tissue contrast across multiple anatomical sites [5, 6]. However, the interpretation of musculoskeletal MRI remains clinically challenging and time-consuming due to sev- eral inherent complexities, including the frequent coexistence of multiple abnormalities within a single examination, intricate three-dimensional anatomy, and the need to inte- grate complementary diagnostic information from multiple imaging planes. Although deep-learning systems have shown promise in detecting isolated conditions such as anterior cruciate ligament injuries [7, 8], meniscal lesions [9], and rotator cuff tears [10], among others, most existing approaches are narrowly designed for single diagnostic tasks at specific anatomical joints. Trained on limited data, these models often strug- gle to generalize across different clinical centers, MRI scanner protocols, and diverse patient populations, limiting their practical utility in real-world clinical workflows [11]. Recent advances in foundation models, particularly through self-supervised learn- ing [12–16], offer a promising pathway toward more generalizable medical image analysis [17–21]. Among these, diffusion models have emerged as powerful genera- tive frameworks that learn rich data distributions by iteratively denoising corrupted 2 inputs [22–25]. Beyond image synthesis, the intermediate representations learned dur- ing this denoising process capture robust, multi-scale features [26, 27]. Our previous study demonstrated that diffusion pre-training as a unified approach to simultane- ously acquire generation ability and deep visual understandings, which potentially leads to the development of unified vision foundation models for diverse downstream clinical tasks [26]. Compared with discriminative learning approaches that optimize task-specific objectives [28–34], the multi-step denoising process of diffusion models inherently captures hierarchical features, providing a continuous spectrum of feature abstraction that allows a single model to adapt flexibly to different diagnostic tasks without architectural modifications [35, 36]. This generative denoising objective also promotes resilience to the types of noise and variability commonly encountered across clinical centers and MRI scanners. To translate these advances into clinical practice, we propose OrthoDiffusion, a diffusion-based representation learning framework for scalable musculoskeletal MRI interpretation. Mirroring the multi-plane analysis routinely performed by radiologists, OrthoDiffusion employs three independent, orientation-specific 3D diffusion branches, each pretrained on large-scale unlabeled MRI data, to learn anatomy-aware features from sagittal, coronal, and axial views. These view-specific representations are sub- sequently integrated through an adaptive fusion strategy to support heterogeneous downstream tasks, including multi-label diagnosis and anatomical segmentation across different joints. We constructed one of the largest multi-joint musculoskeletal MRI datasets, comprising over 1.44 million images and 91,959 sequences, from 30,653 MRI examinations of the knee, ankle, and shoulder (Fig. 1(a)). Pretrained on 15,948 knee MRI examinations, OrthoDiffusion achieves robust, label-efficient performance on 11 knee anatomic structures segmentation and 8 knee diseases diagnosis across multi- ple clinical centers and MRI field strengths. Besides, we further demonstrated that OrthoDiffusion exhibits strong cross-anatomy transferability on other joints, with 7 diagnostic tasks on shoulder and 4 tasks in ankle (Fig. 1(f)). Crucially, this is achieved through a single unified model with one fine-tuning, supporting concurrent multi- label diagnosis and anatomical segmentation without requiring separate task-specific fine-tuning. This unified, diffusion-based paradigm offers a scalable and clinically coherent framework for advancing musculoskeletal MRI interpretation toward broader real-world adoption. 2 Results 2.1 Overview of the study The overall workflow of the study is illustrated in Fig. 1(b)-(e). The proposed framework was developed and validated using a large-scale, multi-center MRI cohort consisting of over 30,000 patients from nine clinical institutions. The dataset covers the knee, ankle, and shoulder joints, featuring annotations for 11 anatomical structures and 19 clinically relevant musculoskeletal abnormalities. Detailed task definitions and abbreviation conventions are provided in Table 1. We utilized 3D diffusion-based architecture for representation learning, incor- porating orientation-specific pretraining across sagittal, coronal, and axial planes. 3 ACL Injury MCL Injury PCL Injury M Tear LM Tear EFFU LCL Injury PD ATFL OLT CFL ATR SSP ISP SSC SASD LHBT Injury AC Knee Shoulder Ankle 3D-Unet3D-ResNet-18 OrthoDiffusion a b c d e f MRI Scan Dataset Model Development Multi-task Evaluation Diffusion Pre-training Dataset Patients (n=15,948) Patients (n=10,940) ACL Injury PCL Injury M Tear LCL InjuryMCL Injury LM Tear PD EFFU 8 Knee Abnormalities Knee Diagnostic Dataset Knee Segmentation Dataset La bel 11 Knee Anatomical Structures La bel Sagittal PlaneMask Coronal PlaneMask Knee MRI Femur Femur Cartilage Patella Cartilage Patella (L/M) Tibia Cartilage Tibia ACL PCL M LM Patients (n=1,006) (Sagitta/Coronal/Axial Plane) Other Joint Diagnostic Dataset SSP ISP SSC LHBT Effusion AC LHBT Injury SASD OLT ATR ATFL Shoulder MRIAnkle MRI Patients (n=8,957) 7 Shoulder Abnormalities 4 Ankle Abnormalities Patients (n=2,562) Multi-planar Information Anisotropic 3D Image Sagittal Plane Axial Plane Coronal Plane 256 256 256 256 256 256 16 16 16 ACL Injury PCL Injury M Tear LM Tear MCL Injury LCL Injury PD EFFU ACL Injury PCL Injury M Tear LM Tear MC L Injury LCL Injury PD EFFU CFL LHBT Effusion Fig. 1: Overview of the OrthoDiffusion framework. a, Multi-planar mus- culoskeletal MRI acquisition, including sagittal, coronal, and axial views, with anisotropic resolution across slices. Dataset construction and task composition across joints are illustrated. b, Unconditional 3D diffusion pretraining using a 3D U-Net noise predictor on large-scale knee MRI data. c, Feature extraction from intermedi- ate diffusion representations at selected timesteps and bottleneck blocks, followed by pooling and multi-label classification. Feature-level and label-level fusion strategies are applied to integrate sagittal, coronal, and axial representations. d, Anatomical segmen- tation pipeline using diffusion representations coupled with a lightweight segmentation head. e, Multimodal fusion strategy integrating MRI diffusion representations with structured electronic health record (EHR) data for diagnosis. f, Diagnostic perfor- mance (AUROC) across 19 musculoskeletal abnormalities involving the knee, ankle, and shoulder joints, compared with baseline CNN models. 4 Table 1: Unified definitions of anatomical structures and abnormalities for musculoskeletal MRI tasks. Abbreviations are defined separately for segmen- tation and classification. TaskJointStructure / Abnormality (Abbreviation) SegmentationKnee Femur Tibia Patella Femoral cartilage (FC) Medial tibial cartilage (MTC) Lateral tibial cartilage (LTC) Medial meniscus (M) Lateral meniscus (LM) Anterior cruciate ligament (ACL) Posterior cruciate ligament (PCL) Patellar cartilage (PC) Classification Knee Anterior cruciate ligament (ACL) injury Posterior cruciate ligament (PCL) injury Medial meniscus (M) tear Lateral meniscus (LM) tear Medial collateral ligament (MCL) injury Lateral collateral ligament (LCL) injury Patellar dislocation (PD) Joint effusion (EFFU) Ankle Anterior talofibular ligament (ATFL) injury Calcaneofibular ligament (CFL) injury Achilles tendon rupture (ATR) Osteochondral lesion of the talus (OLT) Shoulder Supraspinatus tendon (SSP) tear Infraspinatus tendon (ISP) tear Subscapularis tendon (SSC) tear Long head of the biceps tendon (LHBT) injury Adhesive capsulitis (AC) Long head of the biceps tendon (LHBT) sheath effusion Subacromial-subdeltoid (SASD) bursal effusion Table 2: Segmentation performance of OrthoDiffusion under less than 30% labeled data (Dice Simi- larity Coefficient, %) for 11 knee anatomical structures on sagittal and coronal MRI scans. Entries marked as “–” indicate anatomical structures that are not visible in the coronal view (e.g., Patella and PC). Boldface indicates the best performance under the same setting in each column. Orientation MethodsFemur Tibia Patella FC MTC LTC PC M LM ACL PCL Sagittal 3D-Unet91.91 90.67 82.25 64.45 50.27 45.54 65.59 50.71 53.13 46.49 51.40 UNETR83.32 73.90 57.05 44.84 40.99 26.85 48.54 30.12 21.14 30.74 15.30 Ours (FT, FT-Opt-Sag) 93.02 92.54 84.58 58.50 63.76 63.07 64.13 60.69 56.03 53.67 58.34 Coronal 3D-Unet90.87 90.54–61.12 68.51 64.63–63.49 58.40 45.72 42.19 UNETR87.30 84.37–58.87 57.61 49.50–46.87 43.86 38.96 46.66 Ours (FT, FT-Opt-Cor) 92.91 92.82–64.21 69.73 69.63–68.37 65.33 64.13 60.06 5 Feature representations extracted from selected diffusion timesteps and network blocks were employed for downstream tasks, including anatomical segmentation and disease classification. The results demonstrated that OrthoDiffusion consistently outperformed super- vised baselines, particularly in limited-label situation, while maintaining robustness across different joints and heterogeneous imaging conditions. 2.2 Performance on Anatomical Segmentation The model’s capability of the learned diffusion representations used for dense predic- tion tasks was evaluated on knee MRI scans. As illustrated in Fig. 1(d), a lightweight segmentation head was trained on top of features extracted from the pretrained OrthoDiffusion framework to perform segmentation tasks of 11 anatomical structures. The results indicate that the proposed framework, with minimal fine-tuning, consis- tently outperformed strong CNN-based and Transformer-based models [37, 38] trained from scratch under identical experimental settings in the limited-label regime (Table 2 and Fig. 2(a)). Qualitative visualizations in Fig. 2(b) further demonstrate the anatom- ical fidelity of the segmentation masks. These findings suggest that the diffusion-based representations learned by OrthoDiffusion effectively capture anatomically meaning- ful structural and textural information. Further implementation details are provided in the Supplementary Material. 2.3 Diagnostic Accuracy and Robustness in Knee MRI As summarized in Table 3 and Fig. 1(f), the optimal configuration, which selected specific diffusion timesteps and bottleneck blocks based on validation performance, enabled OrthoDiffusion to achieve a macro-average AUROC of 0.908 across eight knee abnormality categories, outperforming baseline approaches [37, 39]. Table 3: Comparison of per-class AUROC (%) for the eight-label knee injury prediction task using different modalities (MRI-only, EHR-only, and MRI+EHR) on the Center A+B+C+D+E+F+G test set. Boldface indicates the best performance under the same setting for each column. DatasetModalityMethodsACL PCL M LM MCL LCL PD EFFU Macro-AUC Center A+B+ C+D+E+F+G MRI-only 3D-Unet94.35 84.76 70.91 68.76 83.44 79.13 94.67 85.0382.63 3D-ResNet-1894.91 85.00 74.87 70.74 83.71 76.63 96.17 83.2183.15 Ours (LP, LP-Opt) 95.15 91.26 81.37 76.23 89.08 82.58 98.85 86.3487.61 Ours (FT, FT-Opt) 97.04 93.73 87.24 85.02 91.91 85.77 99.26 86.3590.79 EHR-onlyEHR-only52.37 53.22 48.77 50.19 54.24 62.02 57.88 55.9954.33 EHR + MRI Ours (LP, LP-Opt) 95.24 91.23 81.40 76.38 89.24 84.29 98.84 86.8487.93 Ours (FT, FT-Opt) 97.02 93.80 87.21 85.09 92.01 86.56 99.27 86.9590.99 Additionally, OrthoDiffusion demonstrated robustness to domain shifts induced by variations in magnetic field strength (Table 4, Supplementary Table 14, and Fig. 3(c)). 6 32 48 64 80 96 5102050 100 Training Data (%) Dice Similarity Coefficient 0 30 60 90 5102050100 Training Data (%) Dice Similarity Coefficient 0 25 50 75 5102050100 Training Data (%) Dice Similarity Coefficient MTC 0 25 50 75 5102050100 Training Data (%) Dice Similarity Coefficient 0 25 50 75 5102050100 Training Data (%) Dice Similarity Coefficient LTC PC Patella Tibia 0 25 50 75 5102050100 Training Data (%) Dice Similarity Coefficient ACL 60 70 80 90 100 5102050100 Training Data (%) Dice Simila rity Coefficient Femur 0 25 50 75 5102050100 Training Data (%) Dice Simila rity Coefficient FC 0 25 50 75 5102050100 Training Data (%) Dice Simila rity Coefficient M 0 25 50 75 5102050100 Training Data (%) Dice Similarity Coefficient LM 0 25 50 75 5102050100 Training Data (%) Dice Simila rity Coefficient PCL 0 20 40 60 80 5102050100 Training Data (%) Dice Simila rity Coefficient Average Knee MRI ScanPrediciton Ground Truth Sagittal Plane Coronal Plane Knee MRI ScanPrediciton Ground Truth a b 3D-UNet OrthoDiffusion UNETR Fig. 2: Qualitative and quantitative evaluation of OrthoDiffusion for knee anatomical segmentation. a, Label efficiency of OrthoDiffusion for knee anatomical segmentation on the Sagittal plane, illustrated by representative Dice Similar- ity Coefficient curves for selected anatomical structures under progressively reduced supervision. b, Qualitative multi-planar segmentation results. Representative sagittal and coronal knee MRI slices showing the original input images, corresponding model- predicted anatomical segmentations (e.g., femur, tibia, cartilage, etc.), and manual ground-truth annotations, illustrating accurate delineation of key joint structures across imaging planes. An exploratory multi-modal analysis integrating structured Electronic Health Records (EHR) with MRI representations suggested that multi-modal data provides complementary diagnostic cues, yielding modest performance gains (Table 3 and Supplementary Table 19). Furthermore, OrthoDiffusion exhibited strong label efficiency. As shown in Fig. 3(a), it retained high diagnostic precision even when trained with only 10% of the labeled data, showing significantly less performance degradation compared to CNN 7 Table 4: Comparisons of per-disease AUROC (%) for the eight-label knee injury prediction task on the Center A+B+C+D+E+F+G test set with different magnetic field strengths. Boldface indicates the best result under the same setting in each column. Field Strength MethodsACL PCL M LM MCL LCL PD EFFU Macro-AUC 1.5T 3D-Unet94.06 84.06 70.60 71.42 84.72 82.76 91.44 88.1383.40 3D-ResNet-1893.76 85.50 75.13 71.75 86.35 82.44 94.84 85.6484.43 Ours (FT, FT-Opt) 96.87 94.99 88.49 83.50 93.66 89.88 98.78 88.9691.89 3T 3D-Unet94.57 85.31 71.06 67.95 83.18 77.24 95.78 84.2882.42 3D-ResNet-1895.76 85.91 74.80 70.51 81.14 75.24 97.03 82.9882.92 Ours (FT, FT-Opt) 97.29 93.99 87.07 85.36 91.11 84.54 99.40 85.5190.53 baselines. This label efficiency demonstrates that the diffusion model captures high- quality, anatomically meaningful representations that can be leveraged effectively with minimal downstream supervision. 2.4 Cross-Anatomy Generalization of OrthoDiffusion To evaluate scalability and transferability, the diffusion backbone of OrthoDiffusion, which pretrained exclusively on knee MRI, was applied to ankle and shoul- der classification tasks. Despite substantial morphological differences across joints, OrthoDiffusion achieved competitive diagnostic performance on ankle and shoulder abnormalities with minimal fine-tuning (Fig. 1(f), Fig. 3(b), and Supplementary Table 15-18). OrthoDiffusion consistently outperformed 3D CNN baselines trained from scratch for each specific anatomy. These findings highlight that the backbone learns anatomical and transferable pathological representations that are not confined to a single joint. Moreover, when target joints share higher anatomical and imaging- level similarity, such as the knee and ankle, the diffusion backbone transfers richer medical priorities, resulting in more pronounced performance improvements. 2.5 Interpretable Multi-Plane Fusion Strategies Our study systematically evaluated fusion strategies for integrating sagittal, coro- nal, and axial MRI representations. While simple concatenation yielded the highest macro-average AUROC (Supplementary Table 8 and 9), it lacks transparency and interpretability regarding the contribution of each anatomical plane. To address this, we proposed a Multi-plane Adaptive Expert (MPAE) Fusion framework. MPAE operates by treating the sagittal, coronal, and axial branches as three inde- pendent experts, each producing disease-specific predictions for downstream fusion. For each diagnostic label, these outputs are passed into an adaptive gating mechanism that computes patient-specific, orientation-wise fusion weights. After normalization, these weights determine the relative contribution of each anatomical plane to the final diagnostic output. Although simple concatenation was used as the default for maximum numerical performance, MPAE provides clinically relevant interpretability. Visualization of the 8 AUROC 50 60 70 80 90 100 1 5 102050100 Training Data (%) AUROC 50 65 80 95 1 5 102050100 Training Data (%) AUROC 50 60 70 80 90 1 5 102050100 Training Data (%) AUROC 42 53 64 75 86 1 5 102050100 Training Data (%) AUROC 65 75 85 95 1 5 102050100 Training Data (%) AUROC 56 66 76 86 1 5 102050100 Training Data (%) AUROC 40 60 80 100 1 5 102050100 Training Data (%) AUROC 60 70 80 90 1 5 102050100 Training Data (%) AUROC EFFU PD LCL InjuryMCL Injury M Tear ACL Injury LM Tear 3D-ResNet-18 OrthoDiffusion ATFL ATR OLT 0 20 40 60 80 100 Ankle AUROC SSP ISP LHBT Injury AC LHBT Effusion SASD 0 20 40 60 80 100 Shoulder PD EFFU 0 20 40 60 80 100 Field Strength AUROC 1.5T 3.0T 3D-UNet 3D-ResNet-18 OrthoDiffusion a b c 3D-UNet PCL Injury ACL Injury PCL Injury M Tear LM Tear MCL Injury LCL Injury CFL SSC Fig. 3: Label efficiency, cross-anatomy transferability, and robustness of OrthoDiffusion. a, Diagnostic performance (AUROC) of OrthoDiffusion on eight knee abnormalities as a function of the proportion of labeled training data, demonstrat- ing label-efficient learning under limited supervision. b, Cross-anatomy generalization performance of OrthoDiffusion on ankle and shoulder abnormality diagnosis, trans- ferred from a diffusion backbone pretrained exclusively on knee MRI. c, Robustness of OrthoDiffusion to variations in MRI magnetic field strength (1.5T and 3.0 T), eval- uated on knee MRI acquired at different field strengths. learned fusion weights via Sankey diagrams (Figure 4) reveals that the model pri- oritizes specific planes for different disorders, such as relying on sagittal views for cruciate ligament injuries and coronal views for collateral ligament injuries, aligning with real-world clinical practice. 3 Discussion This study introduces OrthoDiffusion, a unified diffusion-based representation learning framework that addresses enduring challenges in musculoskeletal MRI interpreta- tion, such as the frequent co-occurrence of multiple abnormalities, complex three- dimensional anatomy, and the need to integrate complementary information from multiple imaging planes. By employing orientation-specific 3D diffusion backbones pre-trained on a large-scale cohort of 15,948 knee MRI examinations, OrthoDiffu- sion demonstrates robust performance across diverse joints and heterogeneous clinical tasks—ranging from anatomical segmentation to multi-label diagnosis—using a single unified model. These results underscore the potential of diffusion models to serve not only as generative tools but also as powerful foundational backbones for label-efficient and generalizable medical image analysis. 9 ACL Injury PCL Injury M Tear LM Tear MCL Injury LCL Injury PD EFFU Axial Sagittal Coronal 0 1 1 0 0 1 0 0 a 1 0 1 b SSP ISP SSC LHBT Injury AC LHBT Effusion SASD Sagittal Coronal Axial 0 0 0 1 1 0 1 c Axial Sagittal Coronal ATFL CFL ATR OLT 0 Label Label Label Knee MRI Shoulder MRI Ankle MRI Fig. 4: Orientation-specific fusion expert-weights predicted by MPAE reflect clinical imaging practice. Sankey diagrams showing label-specific fusion weights assigned by the MPAE module across axial, sagittal, and coronal planes for representative (a) knee, (b) shoulder, and (c) ankle abnormalities in a single patient. The relative thickness of each flow indicates the contribution of each imaging orien- tation to the final prediction, revealing plane preferences that align with established clinical reading protocols for different orientations and pathologies. Representative MRI slices from each orientation are shown for reference. The effectiveness of OrthoDiffusion stems from its self-supervised pre-training paradigm, which leverages extensive unlabeled MRI data to learn transferable anatom- ical priors. Unlike masked autoencoding [15, 40], which excels at learning spatial context through reconstruction, or contrastive learning [41–44], which builds invari- ant representations by discriminating between instances, diffusion models [22, 23] adopt a generative approach. Their multi-step denoising process inherently captures hierarchical features across continuous noise levels, yielding a flexible spectrum of abstractions that can be dynamically adapted to various downstream tasks without modifying the network architecture. Although diffusion pre-training is computation- ally demanding and requires careful tuning of noise schedules, its capacity to model complex three-dimensional data distributions and produce noise-robust representa- tions offers a compelling advantage. The denoising objective itself promotes inherent robustness to imaging variations arising from different MRI scanners, acquisition pro- tocols, and magnetic field strengths, thereby enhancing generalization across clinical sites. Moreover, the diffusion timestep provides an additional dimension of represen- tational flexibility [26, 27, 45, 46]. By modulating the noise level at which features are extracted, OrthoDiffusion enables different downstream tasks to access representations at varying levels of abstraction, offering a principled mechanism for selecting task- adaptive features without altering the underlying model (Supplementary Table 10 and Extended Data Fig. 1). This property makes diffusion modeling an promising foun- dation for a unified framework capable of supporting concurrent multi-label diagnosis and anatomical segmentation. Effective self-supervised pre-training typically requires large and diverse datasets, which have been relatively scarce in musculoskeletal MRI compared with other medical imaging domains. While public repositories and foundation models have accelerated progress in specialties such as ophthalmology [20, 47], pathology [48–50], and derma- tology [51], musculoskeletal imaging has lagged due to challenges in data accessibility 10 and annotation complexity. To bridge this gap, we curated one of the largest multi- joint musculoskeletal MRI cohorts to date, comprising 1.44 million images and 91,959 sequences derived from 30,653 examinations of the knee, ankle, and shoulder, with comprehensive annotations spanning 11 anatomical structures and 19 abnormalities. This dataset not only supports robust pre-training and finetuning but also enables systematic evaluation of cross-anatomy generalization. In terms of model architecture, OrthoDiffusion is tailored to musculoskeletal MRI through three independent diffusion backbones dedicated to sagittal, coronal, and axial views, each trained to capture orientation-specific anatomical context. Through adaptive fusion, the model integrates complementary planar information in a manner that aligns with radiological reading practices. Supported by this substantial dataset, OrthoDiffusion demonstrates high label efficiency, outperforming strong baselines even when trained on only 10% of labeled data. The framework also exhibits multi-task adaptability and cross-anatomical transferability: a backbone pre-trained on knee MRI substantially improves diagnostic performance on ankle and shoulder examina- tions with minimal fine-tuning. These outcomes highlight the model’s ability to learn musculoskeletal representations that are both anatomically meaningful and patholog- ically informative, rather than relying on task-specific heuristics, thereby supporting robust transfer across tasks and anatomical regions within a unified framework. Con- sequently, the pre-trained backbone can remain frozen while new diagnostic targets are incorporated through only lightweight task-specific heads, enabling a “unified-model, multi-disease” paradigm that approaches plug-and-play clinical deployment. Several limitations warrant consideration. First, the use of unconditional diffusion models limits the generative potential of the framework; in the context of musculoskele- tal MRI, controlled diffusion strategies could enable targeted data augmentation and help alleviate class imbalance and data scarcity. Second, the current study focuses primarily on joint-related musculoskeletal disorders assessed via MRI. Extending OrthoDiffusion to multi-modal inputs and additional anatomical sites, alongside more diverse datasets covering a broader spectrum of clinical tasks, represents an important future direction. Third, while the model demonstrates strong performance across mul- tiple centers and magnetic field strengths, it has not yet been validated in real-time clinical workflows. Incorporating human–AI collaboration and prospective evaluations will be essential to assess its usability, reliability, and ultimate impact on diagnos- tic decision-making. Despite these limitations, OrthoDiffusion represents a meaningful advance toward a unified, diffusion-based paradigm for musculoskeletal MRI. 4 Methods 4.1 Ethics statement This retrospective multi-center study was conducted in accordance with the Declara- tion of Helsinki and was approved by the Peking University Third Hospital Medical Science Research Ethics Committee (Project Number IRB00006761-M2024043) and other participating centers. Due to the retrospective nature of the study and the anonymization of all patient data, the requirement for informed consent was waived. 11 4.2 Dataset Study Cohort We curated a large-scale musculoskeletal MRI dataset comprising over 30,000 exami- nations from nine independent clinical centers. The cohort included patients presenting with joint-related symptoms who underwent MRI scans. To ensure data quality and consistency across datasets, we applied strict exclusion criteria: (1) history of prior surgical intervention or hardware implantation in the target joint; (2) presence of acute complex fractures or tumors causing significant anatomical distortion; (3) severe motion or metal artifacts preventing diagnostic interpretation; and (4) incomplete imaging sequences. Dataset Composition The dataset was organized into three distinct subsets to support model pretrain- ing, downstream tasks, and cross-anatomy generalization. Demographic details and acquisition parameters of datasets are provided in Table 5 and Supplementary Table 1-5. Pretraining Dataset. To build robust feature representations, we utilized 15,948 unlabeled knee MRI examinations collected from Centers A–C between December 2016 and June 2023. These images, acquired at magnetic field strengths of 1.5T and 3.0T using routine sagittal, coronal, and axial proton density-weighted (PDW) sequences, were used exclusively for self-supervised diffusion model pretraining. Knee Diagnostic and Segmentation Datasets. For the classification task, we compiled a labeled dataset of 10,940 knee MRI scans from eight institutions (Centers A–G) between December 2016 and March 2025. This subset features high hetero- geneity, utilizing scanners from multiple vendors (GE Healthcare, Siemens, Philips, and UIH) to test robustness in real-world scenarios. Additionally, we also constructed a dedicated knee segmentation dataset consisting of 1,006 knee MRI examinations collected from two institutions (Centers A and B) between July 2023 and February 2024. Unlike the classification subset, this cohort focused on pixel-wise anatomical delineation using sagittal and coronal PDW sequences. Cross-Anatomy Generalization Datasets. To assess the transferability of the learned representations to other joints, we incorporated independent datasets for ankle and shoulder MRI. The ankle cohort comprised 2,562 examinations collected from three institutions (Centers A–C and H) between June 2024 and June 2025. The shoul- der cohort included 8,957 examinations from four institutions (Centers A–C, H and I) collected between January 2023 and June 2025. Both datasets included sagittal, coro- nal, and axial PDW sequences acquired at 1.5T or 3.0T. These datasets were not seen during pretraining and served exclusively to validate the model’s ability to generalize to anatomical structures and pathologies. Establishment of Ground Truth Ground truth annotations were generated through rigorous, multi-stage, task-specific protocols. For classification datasets, the reference standard was determined based 12 Table 5 : Overview of datasets used for knee-only diffusion pre-training and cross-joint downstream evaluation. Task Joint Centers No. of MRI Abnormalities / Structures Age (years) Sex (M/F) Field strength (1.5T / 3T) Diffusion pretraining Knee A–C 15,948 – 33.0 (3–87) 65.3% / 34.7% 28.8% / 71.2% Knee diagnostic Knee A–G 10,940 8 knee abnormalities 33.4 (6–87) 65.5% /34.5% 24.9% / 75.1% Knee anatomical segmentation Knee A–B 1,006 11 anatomical structures 42.4 (6–80) 55.9% / 44.1% 31.5% / 68.5% Ankle diagnostic Ankle A–C, H 2,562 4 ankle abnormalities 33.4 (6–80) 49.3% / 50.7% 7.0% / 93.0% Shoulder diagnostic Shoulder A–C, H–I 8,957 7 shoulder abnormalities 49.6 (9–93) 52.1% / 47.9% 20.3% / 79.7% 13 on clinical intervention status. For patients who underwent surgery, intraoperative findings served as the definitive gold standard. For non-surgical cases, labels were independently generated by three board-certified specialists in sports medicine, each with over ten years of experience. These reviewers performed blinded assessments of multi-view sequences without access to original reports. Any diagnostic discrepancies were resolved through consensus discussion among the three experts to establish the final gold standard. For segmentation dataset, we employed a hierarchical multi-stage annotation strategy. Initial segmentation masks for 11 distinct anatomical structures were manually delineated by six junior radiologists. All initial annotations were sub- sequently audited by three senior radiologists, each possessing over seven years of experience. Annotations deemed inaccurate were returned to the junior annotators for revision. For complex or ambiguous cases, a consensus ground truth mask was generated through consultation among three senior experts. Data preprocessing During both training and testing, MRI scans are standardized by extracting a con- tiguous stack of 16 central slices from each PDW sequence and resizing each slice to a spatial resolution of 256× 256. Volumes are normalized using min–max scaling, followed by linear intensity mapping to the range [−1, 1]. 4.3 Model Overview OrthoDiffusion is an orientation-aware diffusion framework that learns robust and transferable anatomical representations from musculoskeletal MRI (Fig. 1 (b)-(d)). OrthoDiffusion pretrains three independent diffusion backbones-one for each orientation-to capture view-specific structural priors through large-scale self- supervised denoising. For downstream task analysis, OrthoDiffusion extracts intermediate features from the bottleneck layer of each pretrained diffusion model at a selected diffusion timestep. These features are then processed through the pooling operator and subsequently combined using the fusion strategy. OrthoDiffusion adopts a unified two-stage training protocol for classification tasks. In Stage I, diffusion backbones are used to generate orientation-specific feature maps: for linear probing, the backbone is fully frozen and only the pooling module is optimized; for fine-tuning, the pooling module and the encoder/bottleneck lay- ers of the diffusion 3D U-Net are jointly updated. After convergence of Stage I, all feature-extracting components are frozen. In Stage I, only the fusion module is opti- mized, ensuring that cross-orientation integration is learned over stable, task-adapted representations without interference from upstream gradient updates. For anatomical segmentation tasks, we adopt a minimal fine-tuning strategy in which the encoder and bottleneck layers of the diffusion backbone are jointly optimized together with a shallow segmentation head. No additional pooling or multi-orientation fusion modules are introduced in the segmentation pipeline. All representations are passed to task-specific heads, including a single-layer lin- ear classifier trained with binary cross-entropy (BCE) for multi-label diagnosis and a 14 lightweight segmentation head optimized with cross-entropy and soft Dice losses for anatomical segmentation. 4.4 Orientation-Specific Diffusion Backbone Pretraining Because sagittal, coronal, and axial views provide complementary anatomical informa- tion in musculoskeletal MRI, OrthoDiffusion trained three independent unconditional 3D diffusion models, each corresponding to a single MRI orientation. This design enables each model to capture the 3D distributional characteristics specific to its corresponding view, yielding anatomically coherent and orientation-aware latent representations for downstream tasks [26, 27]. Each diffusion model adopts a 3D U-Net denoiser to parameterize the reverse dif- fusion process, following established volumetric diffusion architectures [52]. This archi- tecture captures multi-scale spatial context and supports high-fidelity reconstruction of anatomical structures within each orientation. Diffusion models are capable of approximating highly complex data distributions p(x) [22]. Specifically, Gaussian noise is progressively added to a 3D input volume x 0 over T diffusion timesteps according to x t = √ α t x 0 + √ 1− α t ε ,with t = 1,· ,T,ε∼ N (0, I), where α t controls the noise level at timestep t. The model is trained to learn the corresponding reverse denoising process of this fixed-length Markov chain. The learning objective is formulated as θ ∗ = arg min θ E x,ε, t h ∥ε−ε θ ( x t ,t ) ∥ 2 2 i .(1) where ε θ (·) denotes the neural network that predicts the injected noise. 4.5 Extraction of Intermediate Diffusion Representations For each MRI orientation, OrthoDiffusion extracts intermediate feature maps from the pretrained diffusion denoiser ε θ (x t ,t). These intermediate activations encode multi- scale, orientation-specific anatomical representations for downstream analysis. To transform high-dimensional diffusion feature maps into compact and discrim- inative embeddings suitable for classification tasks, OrthoDiffusion evaluates three pooling strategies. Global Average Pooling (GAP) produces a holistic descriptor by averaging features across all spatial locations. Global–Local Pooling (GLP) aug- ments the global summary with a locally attended feature, enabling the retention of region-specific structural information. Self-Attention Pooling (SAP) applies multi- head self-attention to adaptively weight informative spatial locations and fuses the resulting representation with a global descriptor, enabling the modeling of long-range dependencies. Unless otherwise stated, SAP is adopted as the default pooling strategy in all experiments, due to its superior and stable performance observed in ablation studies (Supplementary Table 6 and 7). 15 4.6 Task-Specific Selection of Diffusion Timesteps and Bottleneck Blocks OrthoDiffusion evaluates diffusion representations across sagittal, coronal, and axial views at multiple diffusion timesteps and bottleneck blocks. The optimal configuration is selected on the validation set and fixed for all held-out test cohorts to avoid data leakage. For each task, MRI orientation and training setting (linear probing or fine-tuning), OrthoDiffusion selects the best-performing ”Timestep-Block” combination, which are summarized in Supplementary Table 10. 4.7 Integrating Complementary Information Across MRI Orientations To integrate complementary information across the MRI three orientations, OrthoD- iffusion assesses both feature-level and label-level fusion strategies for combining orientation-specific representations. At the feature level, OrthoDiffusion considers several commonly used approaches for merging orientation-specific embeddings prior to downstream tasks. The simplest method is direct concatenation, which retains all orientation-wise information with- out additional parameters. OrthoDiffusion further considers linear-projection fusion, where each orientation-specific embedding is mapped into a shared latent space and merged by element-wise addition or concatenation, enabling lightweight learn- able alignment across orientations. Finally, OrthoDiffusion explores a cross-attention mechanism that enables the three orientation features to exchange information; each embedding attends to structural cues present in the others, and the refined embeddings are then concatenated for downstream prediction. Complementing feature-level fusion, OrthoDiffusion introduces a label-level strat- egy termed Multi-plane Adaptive Expert Fusion (MPAE), designed primarily to improve interpretability across anatomical orientations. MPAE takes the logits pro- duced by the three orientation-specific classifiers and uses a compact gating network to generate patient- and label-specific fusion weights. After softmax normalization, these weights determine how the three logits are combined. This design highlights the most informative orientation for each diagnostic label and yields clinically inter- pretable weighting patterns that mirror how radiologists integrate sagittal, coronal, and axial cues. Unless otherwise stated, simple concatenation is used as the default fusion strat- egy in all subsequent experiments, based on ablation analyses reported in the Supplementary Table 8 and 9. 4.8 Multimodal Fusion with Electronic Health Records To further exploit multimodal information, OrthoDiffusion incorporates Electronic Health Records (EHR) associated with each subject through a unique hospitaliza- tion identifier. The EHR variables include demographic attributes (age, sex, height, weight), patient type (athlete vs. non-athlete), and categorical injury-inducing events. 16 Continuous variables are standardized, while categorical attributes are encoded using one-hot vectors or learnable embeddings, with missing values handled through explicit UNK categories to avoid cohort exclusion. MRI and EHR information are integrated via a late-fusion strategy at the logit level (Fig. 1 (e)). Specifically, independent prediction heads are applied to MRI fea- tures and EHR features, and the resulting logits are combined through learnable, label-wise fusion weights. This design preserves the discriminative power of imaging features while enabling selective utilization of complementary clinical signals. Detailed implementation is provided in the Supplementary Materials. 4.9 Evaluation and statistical analysis Task performance was evaluated using standard metrics for classification and segmen- tation, including the area under the receiver operating characteristic curve (AUROC), average precision (AP), class-wise precision, recall, and F1-score (CP/CR/CF1), over- all precision, recall, and F1-score (OP/OR/OF1), and the Dice Similarity Coefficient (DSC). For multi-label classification, AUROC was computed separately for each disease category and macro-averaged across tasks. AP summarizes the precision–recall curve, while CP/CR/CF1 and OP/OR/OF1 quantify class-wise and overall classification per- formance, respectively. Segmentation performance was assessed using DSC to measure the spatial overlap between predicted masks and ground-truth annotations. Statistical significance between OrthoDiffusion and baseline models was assessed using paired non-parametric permutation tests on per-task performance differences. Across all evaluated tasks, permutation tests indicated statistically significant and consistent performance improvements (p < 0.001). 17 Acknowledgment This work was funded by the National Natural Science Foundation of China (Grant number 82441025), and Beijing Municipal Natural Science Foundation (Grant number L242104). Data Availability The datasets are not publicly available due to privacy concerns, but are partially available from the corresponding author on reasonable request. Code Availability All code is available via GitHub at https://github.com/lt-0123/OrthoDiffusion. Author Contributions T.L., L.X., Z.Y., W.X., H.Y., J.Y., and D.W. conceived and designed the study. L.X., Z.Y., S.L., J.L., and J.L. acquired, organized, and verified the raw data. T.L. devel- oped the methodology, performed the technical implementation, and conducted the results analysis. T.L., L.X., Z.Y., D.W., J.Y. discussed the results and provided crit- ical comments on the paper. The study was supervised by D. J, J.Y. and D. W. All authors contributed to the drafting and revising of the manuscript and approved the final version. Competing Interests The authors declare no competing interests. 18 Appendix A 0100200300400500 Diffusion Timestep 82.5 83.0 83.5 84.0 84.5 85.0 85.5 86.0 Macro-Averaged AUC-ROC(%) 85.04 Macro-Averaged AUC-ROC vs Diffusion Timestep (Sagittal Plane) 3D-Unet Block Name mid_0 mid_1 mid_2 (a) Sagittal Plane 0100200300400500 Diffusion Timestep 83.5 84.0 84.5 85.0 85.5 86.0 86.5 Macro-Averaged AUC-ROC(%) 85.59 Macro-Averaged AUC-ROC vs Diffusion Timestep (Coronal Plane) 3D-Unet Block Name mid_0 mid_1 mid_2 (b) Coronal Plane 0100200300400500 Diffusion Timestep 81.5 82.0 82.5 83.0 83.5 84.0 84.5 85.0 85.5 Macro-Averaged AUC-ROC(%) 84.71 Macro-Averaged AUC-ROC vs Diffusion Timestep (Axial Plane) 3D-Unet Block Name mid_0 mid_1 mid_2 (c) Axial Plane Extended Data Fig. 1: Macro-averaged AUROC of the eight-label knee injury prediction task under linear probing on the Center A+B+C test set, based on SAP- processed features extracted from 3D-UNet blocks at different diffusion timesteps across MRI planes. 19 Appendix B More Detail about Dataset Detailed summary statistics for all datasets used in this study are provided in Supple- mentary Table 1-5. These tables report the composition of training, validation, and test cohorts across anatomical segmentation and diagnostic classification tasks, including scanner manufacturers, magnetic field strengths, acquisition parameters, and patient demographics. Appendix C More Detail about Methods C.1 Pooling Strategies for Representation Extraction For each plane o ∈ O = sagittal, coronal, axial, OrthoDiffusion extracts interme- diate diffusion representations v (o) t,b ∈ R C×D ′ ×H ′ ×W ′ from the pre-trained diffusion model ε θ (x t ,t). The MRI sequence input is x 0 ∈ R 1×D×H×W , and is gradually perturbed into x t at timestep t. Here, b indexes 3D U-Net blocks, C the channel dimen- sion, and D ′ ,H ′ ,W ′ the downsampled spatial size of the corresponding feature map. OrthoDiffusion evaluates three pooling strategies for deriving compact representations f (o) t,b : 1. Global average pooling (GAP): The most straightforward approach is to apply global average pooling over the spatial dimensions: f (o) global = 1 D ′ H ′ W ′ D ′ X d=1 H ′ X h=1 W ′ X w=1 v (o) t,b [:,d,h,w] ∈ R C , yielding f (o) t,b = f (o) global . 2. Global–local pooling (GLP): In addition to the global descriptor from GAP, a local attention mechanism highlights spatially discriminative regions. Let v (o) t,b,i ∈ R C denote the feature vector at spatial index i ∈ 1,...,N with N = D ′ H ′ W ′ . Attention weights are computed as α i = exp g(v (o) t,b,i ) P N j=1 exp g(v (o) t,b,j ) , f (o) local = N X i=1 α i v (o) t,b,i , where g(·) is a two-layer MLP with ReLU activation. The global and local descriptors are then concatenated to form the final representation: f (o) t,b = f (o) global ∥ f (o) local ∈ R 2C . 3. Self-attention pooling (SAP): OrthoDiffusion interprets the spatial features v (o) t,b,i N i=1 as a sequence of tokens and arrange them into X = [v (o) t,b,1 ,..., v (o) t,b,N ]∈ R N×C . 20 Each token is linearly projected into queries, keys, and values: Q = XW Q , K = XW K , V = XW V , W Q ,W K ,W V ∈ R C×C . We split the feature dimension into h heads of size d = C/h, producing Q = [Q 1 ∥·∥Q h ], K = [K 1 ∥·∥K h ], V = [V 1 ∥·∥V h ], Q j ,K j ,V j ∈ R N×d . Scaled dot-product attention is computed independently for each head, where j = 1, 2,· ,h: head j = softmax Q j K ⊤ j √ d ! V j ,head j ∈ R N×d . The multi-head output is formed as MHA(X) = Concat(head 1 ,..., head h )W O ∈ R N×C , W O ∈ R C×C . Let MHA(X) i ∈ R C denote the i-th token embedding (the i-th row of MHA(X)). A compact representation is then obtained by average pooling over tokens: f (o) attn = 1 N N X i=1 MHA(X) i ∈ R C . Finally, f (o) t,b = f (o) global ∥ f (o) attn ∈ R 2C . GLP emphasizes spatially localized discriminative structures through learned attention, while SAP captures long-range dependencies via self-attention. Both mechanisms complement the holistic context provided by GAP. For brevity, the pooled representation f (o) t,b will be referred to as F (o) hereafter. C.2 Fusion Strategies Across MRI Orientations For each o ∈ O, let F (o) ∈ R C ′ denote the orientation-specific feature representation extracted by the corresponding diffusion backbone, where C ′ = C for GAP and C ′ = 2C for GLP and SAP pooling. OrthoDiffusion compares four fusion strategies for integrating these representations: (1) Simple concatenation. The fused representation is obtained by channel-wise concatenation: F fuse = F (sag) ∥ F (cor) ∥ F (ax) ∈ R 3C ′ . 21 (2) Linear fusion. To allow fusion in a shared latent space, each feature vector F (o) ∈ R C ′ is linearly projected into a common embedding dimension E: H (o) = W (o) F (o) ∈ R E , o∈O. Fusion is then performed either by element-wise addition (Linear addition): F fuse = H (sag) + H (cor) + H (ax) ∈ R E , or by channel-wise concatenation (Linear concatenation): F fuse = H (sag) ∥ H (cor) ∥ H (ax) ∈ R 3E . (3) Cross-attention fusion. To enable information exchange across poses, all features are first projected into the same embedding: H (o) = W (o) F (o) ∈ R E , o∈O. Cross-attention is used to enrich each orientation with contextual cues from the others, e.g., e H (sag) = MHA H (sag) , H (cor) , H (cor) , and the refined embeddings are obtained via residual connections and layer normal- ization: H ′(o) = LayerNorm H (o) + e H (o) , o∈O. Finally, feature fusion is performed by concatenation: F fuse = H ′(sag) ∥ H ′(cor) ∥ H ′(ax) ∈ R 3E . (4) MPAE fusion. For each anatomical orientation o∈O, an independent multi-label classifier outputs z (o) = z (o) 1 ,...,z (o) K ∈ R K , where z (o) k denotes the predicted logits for disease category k, and K is the number of musculoskeletal diseases. For each label k, we form the orientation-wise logits vector z k = z (sag) k , z (cor) k , z (ax) k ⊤ ∈ R 3 . An adaptive gating network g(·) produces unnormalized scores s k = g(z k )∈ R 3 , and the fusion weights are obtained by α k = softmax(s k ), X o∈O α (o) k = 1. 22 Fusion is performed in logit space: ˆ z k = X o∈O α (o) k z (o) k , ˆ y k = σ( ˆ z k ).(C1) This label-specific, orientation-aware aggregation allows the model to emphasize the most informative anatomical plane for each disorder, mirroring the way radiologists integrate complementary cues across MRI views. C.3 Integration of EHR data EHR variables. To assess whether structured clinical information provides complementary predictive value beyond MRI-derived representations, OrthoDiffusion incorporates Electronic Health Records (EHR) linked to each subject via a unique hospitalization identifier. The variables include: • Patient type: athlete vs. non-athlete. • Sex: male vs. female. • Anthropometric measurements: – age (years), – height (cm), – weight (kg). • Injury-inducing event, categorized into eight mutually exclusive classes: 1. overuse injury, 2. traffic accident, 3. military or police duty accident, 4. daily-life or production accident, 5. sports-related accident, 6. spontaneous (no identifiable trigger), 7. actor or rehearsal accident, 8. other. All records are extracted from the hospital information system and aligned with MRI features via the unique patient identifier. Missing or unavailable entries are preserved by assigning an UNK (unknown) category rather than discarding samples, ensuring consistent cohort size and preventing bias introduced by sample exclusion. EHR Feature Encoding and Normalization. Continuous variables (age, height, weight) are standardized using a z -score trans- form. Missing numeric values are imputed with zero after normalization. Categorical variables (sex and patient type) are converted to fixed-length one-hot vectors, including an UNK category to accommodate missing or ambiguous entries. The injury–inducing event is modeled as a learnable discrete token: each event is mapped to an integer index and embedded into an 8-dimensional vector via a trainable lookup table. Records without a valid entry receive an UNK token automatically. 23 For each subject, the resulting structured EHR representation is F EHR = x (cont) |z ∈R 3 ∥ x (sex) | z ∈R 3 ∥ x (ptype) | z ∈R 3 ∥ e |z ∈R 8 , where x (cont) denotes standardized continuous measurements, x (sex) and x (ptype) are one-hot vectors, and e is the learned event embedding. This design enables structured clinical attributes to be seamlessly integrated with MRI representations in downstream classification. MRI–EHR Fusion Strategy. To quantify the contribution of structured clinical attributes while avoiding scale imbalance between modalities, OrthoDiffusion adopts a late–fusion architecture oper- ating at the logit level. Given a 1536-dimensional MRI representation F MRI (after Simple Concat fusion) and an EHR representation F EHR , two independent multilabel heads produce z MRI , z EHR ∈ R K , where K denotes the number of disease categories. Fusion is parameterized by a learnable, per-label coefficient γ ∈ (0, 1) K : ˆ z = γ⊙ z MRI + (1−γ)⊙ z EHR , with ⊙ representing element-wise multiplication. Fusion weights are given by γ = σ(w), w∈ R K , where σ(·) is the element-wise logistic sigmoid. The fused logits yield final probabilities ˆ y = σ( ˆ z)∈ [0, 1] K . This formulation (i) preserves the discriminative capacity of MRI-derived repre- sentations, (i) allows selective exploitation of complementary clinical signals, and (i) provides label-specific fusion weights to quantify the contribution of EHR information to each diagnostic endpoint. Appendix D Implementation Details Data organization. To prevent information leakage, we strictly partition the training, validation, and test sets at the patient level, ensuring that MRI scans from the same patient never appear in both splits. Each patient is associated with a unique set of diagnostic labels. For the majority of cases, only a single scan is available per anatomical orientation, whereas a small fraction of patients have multiple scans. Given that multi-scan patients 24 constitute only a minor proportion of the cohort, we treat samples as approximately independent and identically distributed (i.i.d.) when training the orientation-specific multi-label classifier. When fusing the three anatomical orientations, data organization differs. For patients with exactly one scan per orientation, the three modalities are fused directly. Patients with missing orientations are excluded, as such incomplete examinations are rare in clinical practice. For patients with more than one scan in any orientation, we enumerate all cross-orientation combinations. However, this may create a dispropor- tionately large number of fused samples for some patients; for example, a patient with four scans in each orientation yields 4 3 = 64 combinations. To avoid patient-level imbalance introduced by this enumeration, we apply an inverse-frequency weighting scheme: during training, each fused sample is scaled by the reciprocal of the number of combinations originating from the same patient. During inference, prediction probabilities from all fused samples of a patient are aggregated before producing the final diagnostic output. For anatomical segmentation, MRI volumes are processed independently within a single orientation, and no cross-orientation fusion is performed; accordingly, the combination enumeration and inverse-frequency weighting strategy described above is not required. Training baseline models. We adopt standard volumetric neural networks as baselines for both diagnostic classification and anatomical segmentation tasks. For anatomical segmentation, we adopt standard 3D U-Net and UNETR archi- tectures implemented using the MONAI framework [53]. The 3D U-Net encoder consists of five resolution stages with channel dimensions (64, 64, 128, 128, 256) and strided 3D convolutions for downsampling, with a symmetric decoder based on trans- posed convolutions and skip connections. UNETR replaces the convolutional encoder with a transformer-based encoder configured with an embedding dimension of 768, 12 self-attention heads, and an MLP dimension of 3072, while retaining a U-Net–style decoder that progressively integrates multi-scale features. Both models were trained from scratch for 10 epochs using a learning rate of 1× 10 −3 and a batch size of 10 unless otherwise specified. For diagnostic classification, we consider two commonly used 3D architectures: a 3D U-Net encoder and a 3D ResNet-18. For the 3D U-Net baseline, bottleneck features (mid0) are globally average-pooled and passed to a linear multi-label clas- sification head. The U-Net backbone matches the architecture used in the diffusion model, enabling a controlled comparison. For knee, the 3D U-Net baseline is trained for 10 epochs with a learning rate of 1× 10 −4 , while the 3D ResNet-18 baseline is trained for 8 epochs using the same learning rate. To evaluate cross-anatomy generalization, both classification baselines are trained under identical protocols on ankle MRI (12 epochs for 3D U-Net; 10 epochs for 3D ResNet-18) and shoulder MRI (8 epochs for 3D U-Net; 6 epochs for 3D ResNet-18). 25 All models are trained until the validation performance converges. The checkpoint corresponding to the stabilized validation performance is used for final test. This training and test protocol is applied consistently across all methods and tasks. Pre-training the diffusion model for each MRI orientation. OrthoDiffusion adopts a 3D U-net denoising backbone following a standard discrete- time diffusion formulation. All diffusion models operate on MRI volumes of size 16× 256× 256 with a single input and output channel. The network uses a base channel width of 64 and a multiscale channel expansion pattern of (1, 1, 2, 2, 4) across resolution levels. Each resolution stage contains one residual block, and self-attention is enabled at the 16×16 spatial resolution. A Gaussian diffusion process with T = 1000 timesteps is employed, optimized with an ℓ 2 denoising objective. All diffusion models are trained for 2.15× 10 4 optimization steps with batch size B = 10, the learning rate is scaled linearly as η = η 0 × B with base η 0 = 1× 10 −5 . Ablation studies on pooling methods. To evaluate how spatial aggregation influences the discriminative power of diffusion representations, we conducted an ablation study comparing three pooling strategies: global average pooling (GAP), global local pooling (GLP), and self attention pooling (SAP). All experiments were preformed using features extracted from the bottleneck block (mid 2) at diffusion timestep t = 30, with pretrained sagittal, coronal, and axial backbones kept frozen to isolate the effect of pooling mechanism. A multi-label classifier was trained on the pooled representations using the Adam optimizer with an initial learning rate of 5× 10 −4 and a cosine annealing schedule over 10 epochs. The bottleneck feature maps extracted from the 3D U-Net have spatial dimensions 256× 1× 16× 16. For GLP, the two-layer MLP used to compute local attention weights employs a hidden dimension of C/8 (with C = 256, yielding a hidden size of 32), while SAP is configured with h = 16 attention heads. As summarized in Supplementary Table 6 and 7, the choice of pooling strategy has a substantial impact on downstream performance. SAP consistently surpasses GAP and GLP across orientations and most disease categories, yielding higher per-disease AUROC and stronger overall multi-label metrics. The gains highlight SAP’s ability to preserve spatially discriminative patterns within diffusion features. Based on these results, SAP is adopted as the default aggregation method in all subsequent experiments. Unless otherwise specified, the hyperparameter settings used for SAP under both linear probing and fine-tuning are kept identical to those adopted in the ablation experiments. Ablation studies on fusion methods. We conducted an ablation study to evaluate multiple strategies for integrating multi- orientation diffusion representations, spanning both feature-level and label-level fusion paradigms. Specifically, we compared five fusion methods: Simple Concatenation, Linear Addition, Linear Concatenation, Cross-Attention Fusion, and the pro- posed Multi-plane Adaptive Expert (MPAE) module. For all fusion comparisons, 26 diffusion features were extracted from the mid2 bottleneck block at a fixed diffusion timestep (t = 30) using pretrained and frozen sagittal, coronal, and axial diffusion backbones. Each orientation-specific feature map was subsequently aggregated using the SAP module, yielding a 512-dimensional representation per orientation. All fusion strategies operate on these 512-dimensional pose-specific embeddings and are summarized as follows: 1. Simple Concat. Orientation-specific representations are concatenated along the channel dimension to form a fused feature F fuse ∈ R 1536 . 2. Linear Concat / Linear Add. Each orientation-specific embedding is projected into a shared latent space of dimension E = 512 via a learnable linear transforma- tion. Linear Concat concatenates the projected embeddings into a 1536-dimensional representation, whereas Linear Add merges them by element-wise summation, resulting in a 512-dimensional fused feature. 3. Cross Attention. The three orientation embeddings are projected into a shared space (E = 512) and fused using cross-attention blocks (4 attention heads), allowing each orientation to attend to structural cues from the others. The refined features are concatenated to produce a 1536-dimensional representation. 4. MPAE. Independent multi-label classifiers are trained for each orientation, yield- ing z∈ R 3×K (5-fold CV on train folds). A per-label gating MLP (hidden size 16, dropout 0.1) predicts label-specific fusion weights, and final logits are obtained by weighted logit aggregation as in Eq. (C1). For Simple Concat, Linear Concat/Add, and Cross Attention, models were trained using the Adam optimizer with an initial learning rate of 5×10 −5 , cosine annealing, and 5 training epochs. For MPAE, the orientation-specific classifier were trained with a learning rate of 5×10 −5 for 5 epochs, while the gating network was optimized with a learning rate of 1× 10 −3 , weight decay 1× 10 −4 , and 10 epochs. As summarized in Supplementary Table 8 and 9, Simple Concat consistently achieved the best or near-best AUROC across the eight abnormality categories and produced the strongest overall multi-label performance. While MPAE offered supe- rior interpretability and the second-highest accuracy, Simple Concat remained the numerically dominant strategy. Given its superior and stable performance, Simple Concat is adopted as the default fusion strategy for all subsequent experiments. Unless otherwise specified, the hyper- parameter settings used for Simple Concat under both linear probing and fine-tuning are kept identical to those adopted in the ablation experiments. MRI-EHR Multi-modal Fusion. The MRI branch uses a linear multi-label classification head. To model structured clinical variables, the EHR branch consists of a lightweight feed-forward projection comprising a fully connected layer with 64 hidden units, followed by ReLU activation, layer normalization, and dropout (0.2), and a final linear output layer. The same EHR architecture is used in both the EHR-only baseline and the multi-modal fusion setting to ensure controlled and comparable evaluation. 27 All models are optimized using Adam. Both linear probing and fine-tuning are performed for 10 epochs with a learning rate of 1× 10 −5 , while the EHR-only baseline is trained for 3 epochs using a learning rate of 5× 10 −5 . OrthoDiffusion pipeline. We extract intermediate representations from the bottleneck blocks (mid0, mid1, mid2) of 3D Unet denoising network and systematically evaluate diffusion timesteps of 10, 30, 50, 100, 150, 200, 300, and 500. In contrast to encoder and decoder layers, which operate at higher spatial resolutions and incur substantially greater memory and computational cost, bottleneck features provide a trade-off between representa- tional richness and efficiency, particularly for 3D multi-view MRI analysis. Therefore, bottleneck blocks are adopted as the default configuration throughout all experiments. To determine task-optimal diffusion representations, we first identify the top four candidate diffusion timestep-block combinations for each anatomical orientation based on validation performance. All 4 3 cross-orientation combinations are then exhaustively enumerated to construct candidate multi-plane configurations. For knee abnormality diagnosis under linear probing setting, the best-performing configuration— axial: t = 50, b = mid0; sagittal: t = 50, b = mid0; coronal: t = 100, b = mid2— is denoted as LP-Opt. Under the fine-tuning setting, the optimal configuration— axial: t = 200, b = mid0; sagittal: t = 150, b = mid0; coronal: t = 50, b = mid2— is denoted as FT-Opt. The same procedure is applied to ankle and shoulder MRI, yielding task- and anatomy-specific optimal configurations for each training setting. For knee anatomical segmentation tasks, we directly select the best-performing timestep–block combination for each orientation (Sagittal and Coronal) based on vali- dation performance, without exhaustive cross-orientation enumeration. This is because segmentation is performed independently for each orientation and does not involve multi-plane fusion. The detailed definitions and selected combinations are summarized in Supplemen- tary Table 10. For knee, ankle, and shoulder abnormality diagnosis, linear probing in Stage I is conducted using the hyperparameter configuration defined above. Fine-tuning follows the same training protocol, except that a substantially reduced learning rate is used to preserve the pretrained diffusion representations. Specifically, the learning rate is set to 5× 10 −4 × 0.05, and models are fine-tuned for five epochs. In Stage I, the training procedure and all hyperparameters remain identical to those specified above. For knee anatomical segmentation task, fine-tuning is carried out for 30 epochs using a learning rate of 1× 10 −3 . Additional experiments. For label-efficiency analysis, we subsample the training cohort at the patient level using nested subsets of size 100, 500, 1,000, 2,000, and the full dataset for knee disease classi- fication, and 50, 100, 200, 500, and the full dataset for knee anatomical segmentation, ensuring strict inclusion relationships among subsets. All training hyperparameters and optimization settings follow the configurations described above. OrthoDiffusion 28 is evaluated under the fine-tuning setting (FT-Opt ) for knee disease classification on the Center A+B+C test dataset, and under FT-Opt-Sag/Cor for knee anatomical segmentation on the test dataset. All experiments are conducted on a single NVIDIA H20-3e data-center GPU with 143 GB of HBM memory. 29 Supplementary Tables Supplementary Table 1: Summary statistics of knee MRI datasets for sagittal anatomical segmentation. CharacteristicTraining setTest setValidation set Total MRI scans71511960 Centers222 Scanner manufacturer GE432 (60.42%)75 (63.03%)41 (68.33%) Siemens–3 (2.52%)1 (1.67%) Philips– United Imaging (UIH)283 (39.58%)41 (34.45%)18 (30.00%) Magnetic field strength 3.0 T455 (63.64%)81 (68.07%)39 (65.00%) 1.5 T260 (36.36%)38 (31.93%)21 (35.00%) Scanning parameters Repetition time (ms, range)1686 (1112–2418)1705 (1094–2264)1700 (1230–2210) Echo time (ms, range)23.3 (16.9–48.0)23.6 (16.0–42.0)23.7 (19.2–30.6) Slice thickness 3.0 m116 (16.22%)18 (15.13%)10 (16.67%) 3.5 m482 (67.41%)69 (57.98%)35 (58.33%) 4.0 m117 (16.36%)32 (26.89%)15 (25.00%) Patient characteristics Age (years, range)42.9 (6–76)42.3 (11–80)40.2 (14–78) Sex (female:male)286:42954:6525:35 Weight (kg, range)84.7 (30–151)82.6 (41–98)84.6 (68–90) 30 Supplementary Table 2: Summary statistics of knee MRI datasets for coronal anatomical segmentation. CharacteristicTraining setTest setValidation set Total MRI scans80213466 Centers222 Scanner manufacturer GE582 (72.57%)85 (63.43%)38 (57.58%) Siemens21 (2.62%)3 (2.24%)5 (7.58%) Philips– United Imaging (UIH)199 (24.81%)46 (34.33%)23 (34.85%) Magnetic field strength 3.0 T655 (81.67%)81 (60.45%)51 (77.27%) 1.5 T147 (18.33%)53 (39.55%)15 (22.73%) Scanning parameters Repetition time (ms, range)1699 (1094–2418)1706 (1289–2264)1722 (1294–2264) Echo time (ms, range)23.2 (16.0–48.0)23.1 (17.9–32.0)24.1 (19.5–48.0) Slice thickness 3.0 m123 (15.34%)22 (16.42%)11 (16.67%) 3.5 m532 (66.33%)75 (55.97%)39 (59.09%) 4.0 m147 (18.33%)37 (27.61%)16 (24.24%) Patient characteristics Age (years, range)42.1 (6–80)43.4 (11–71)41.7 (13–65) Sex (female:male)378:42459:7529:37 Weight (kg, range)83.8 (30–151)84.4 (40–151)84.5 (68–100) 31 Supplementary Table 3 : Summary statistics of knee MRI datasets for abnormality classification. Characteristic Pre-training set Training set Validation set Test set External test set Total MRI scans 15,948 8,760 675 1350 155 Centers 3 3 3 3 4 Scanner manufacturer GE 10,616 (66.57%) 5,619 (64.14%) 429 (63.56%) 869 (64.37%) – Siemens 1,085 (6.80%) 432 (4.93%) 32 (4.74%) 60 (4.44%) 86 (55.48%) Philips – – – – 27 (17.42%) United Imaging (UIH) 4,247 (26.63%) 2,709 (30.92%) 214 (31.70%) 421 (31.19%) 42 (27.09%) Patient characteristics Age (years, range) 33.0 (3–87) 33.3 (6–85) 32.9 (6–87) 33.1 (6–87) 41.2 (12–76) Sex (female:male) 5534:10414 3101:5659 235:440 475:875 79:76 Weight (kg, range) 85.1 (15.0–151.0) 85.0 (15–165) 84.2 (50–110) 84.9 (40.8–144) 82.4 (40.5–165.3) Magnetic field strength 3.0 T 11,348 (71.16%) 6,572 (75.02%) 516 (76.44%) 1,015 (75.19%) 113 (72.90%) 1.5 T 4,600 (28.84%) 2,188 (24.98%) 169 (25.04%) 335 (24.81%) 42 (27.10%) Scanning parameters Repetition time (ms, range) 2495 (1472–4372) 2456 (1472–4080) 2453 (1582–3751) 2455 (1394–3685) 2627 (1821–3100) Echo time (ms, range) 34.8 (16.5–70.4) 34.3 (18.0–70.3) 34.0 (23.8–63.8) 34.3 (23.8–76.5) 35.3 (21.2–41.0) Slice thickness 3.0 m 1,075 (6.74%) 406 (4.63%) 33 (4.89%) 63 (4.67%) 19 (12.26%) 3.5 m 13,861 (86.91%) 7,544 (86.12%) 590 (87.41%) 1,196 (88.59%) 99 (63.87%) 4.0 m 1,012 (6.35%) 810 (9.25%) 52 (7.70%) 91 (6.74%) 37 (23.87%) Disease distribution Anterior cruciate ligament (ACL) injury – 5000 244 893 93 Posterior cruciate ligament (PCL) injury – 783 70 116 21 Medial meniscus (M) tear – 4822 329 795 111 Lateral meniscus (LM) tear – 4310 309 681 95 Medial collateral ligament (MCL) injury – 1354 68 199 45 Lateral collateral ligament (LCL) injury – 1161 18 144 53 Patellar dislocation (PD) – 382 32 56 21 Joint effusion (EFFU) – 4010 269 614 137 32 Supplementary Table 4: Summary statistics of ankle MRI datasets for abnormality classification. CharacteristicTraining setTest setValidation set Total MRI scans2178256128 Centers333 Scanner manufacturer GE741 (34.02%)85 (33.20%)31 (24.22%) Siemens865 (39.72%)114 (44.53%)94 (73.44%) Philips– United Imaging (UIH)572 (26.26%)57 (22.27%)3 (2.34%) Patient characteristics Age (years, range)33.2 (6–80)35.1 (8–77)33.9 (6–72) Sex (female:male)1111:1067127:12962:66 Weight (kg, range)85.3 (30–120)85.0 (40–140)86.7 (40–120) Magnetic field strength 3.0 T2047 (93.99%)239 (93.36%)97 (75.78%) 1.5 T131 (6.01%)17 (6.64%)31 (24.22%) Scanning parameters Repetition time (ms, range)2172 (1170–3870) 2184 (1265–3480) 2455 (1967–3129) Echo time (ms, range)33.8 (10.2–52.3) 33.7 (21.4–49.3) 35.0 (28.0–48.5) Slice thickness 3.0 m21 (0.96%)6 (2.34%)2 (1.56%) 3.5 m2055 (94.35%)221 (86.33%)111 (86.72%) 4.0 m102 (4.68%)29 (11.33%)15 (11.72%) Disease distribution Anterior talofibular ligament (ATFL) injury1851223113 Calcaneofibular ligament (CFL) injury1784204111 Achilles tendon rupture (ATR)310348 Osteochondral lesion of the talus (OLT)6668549 33 Supplementary Table 5: Summary statistics of shoulder MRI datasets for abnormality clas- sification. CharacteristicTraining setTest setValidation set Total MRI scans71671193597 Centers555 Scanner manufacturer GE2028 (28.30%)531 (44.51%)1 (0.17%) Siemens923 (12.88%)–256 (42.88%) Philips– United Imaging (UIH)4216 (58.83%)662 (55.49%)340 (56.95%) Patient characteristics Age (years, range)49.7 (9–93)49.3 (11–83)49.6 (12–85) Sex (female:male)3430:3737562:631296:301 Weight (kg, range)83.9 (35–140)84.2 (41–130)83.2 (55–188) Magnetic field strength 3.0 T5732 (79.98%)827 (69.32%)579 (96.98%) 1.5 T1435 (20.02%)366 (30.68%)18 (3.02%) Scanning parameters Repetition time (ms, range)1518 (646–3442) 1546.5 (604–3173) 1410 (1004–3020) Echo time (ms, range)28.6 (12.3–66.5) 29.4 (16.9–56.6) 26.1 (12.8–63.7) Slice thickness 3.0 m21 (0.29%)3 (0.25%)1 (0.17%) 3.5 m6698 (93.46%)1066 (89.35%)594 (99.50%) 4.0 m448 (6.25%)124 (10.39%)2 (0.34%) Disease distribution Supraspinatus tendon (SSP) tear69031153585 Infraspinatus tendon (ISP) tear4143683362 Subscapularis tendon (SSC) tear3034463259 Long head of the biceps tendon (LHBT) injury1771264155 Adhesive capsulitis (AC)2364376204 Long head of the biceps tendon (LHBT) sheath effusion71311462 Subacromial–subdeltoid (SASD) bursal effusion5607913459 34 Supplementary Table 6: Ablation study of pooling strategies for the eight- label knee injury prediction task across three orientations on the Center A+B+C test set. To ensure fair comparison across orientations, all features are extracted from block=mid2 at timestep=30. Results are reported as per-disease AUROC (%). Boldface indicates the best result under the same setting in each column. Orientation Pooling ACL PCL MMLM MCL LCLPDEFFU Sagittal GAP68.8356.8560.2161.5175.0470.7282.7378.89 GLP92.8884.4273.2666.2682.0075.6895.0781.93 SAP94.76 91.04 75.99 70.54 82.08 74.89 98.81 84.76 Coronal GAP74.6158.3365.6961.1078.9673.9084.8981.84 GLP91.7170.6170.1867.2681.0475.6794.3783.84 SAP95.00 83.03 78.94 74.72 87.38 77.53 98.20 83.65 Axial GAP71.4357.8764.7159.1176.3772.7384.9075.72 GLP91.0968.8771.0564.4681.3875.7396.4682.81 SAP94.25 86.06 75.83 69.34 88.33 79.34 98.15 84.06 Supplementary Table 7: Ablation study of pooling strategies for the eight-label knee injury prediction task across three orientations on the Center A+B+C test set. To ensure fair compar- ison across orientations, all features are extracted from block=mid2 at timestep=30. Results are reported as per-disease average precision and overall multi-label metrics (%). Boldface indicates the best result under the same setting in each column. Orientation Pooling ACL PCL M LM MCL LCL PD EFFUCP CR CF1 OP OR OF1 Sagittal GAP73.10 12.90 62.95 59.36 33.40 22.71 20.03 73.4944.63 36.98 35.32 62.56 62.57 62.57 GLP94.39 50.03 76.59 64.71 44.69 29.34 61.86 77.8069.96 46.76 50.99 72.08 66.51 69.19 SAP95.94 73.29 79.72 69.11 46.31 28.85 86.25 81.9969.26 62.31 64.91 73.49 72.04 72.76 Coronal GAP77.32 14.06 68.54 58.66 39.35 25.35 20.47 77.2248.39 38.05 38.37 65.59 62.21 63.86 GLP93.00 22.79 73.09 66.47 44.89 24.84 49.99 79.9668.40 43.61 46.61 71.66 65.09 68.22 SAP96.16 50.06 82.84 73.84 63.94 29.86 77.90 80.0868.60 60.70 63.81 74.84 72.66 73.74 Axial GAP74.12 12.43 68.60 58.62 32.48 21.70 20.72 70.3239.52 34.83 32.54 61.81 60.74 61.27 GLP93.35 20.16 75.41 62.72 42.16 26.34 68.94 78.2364.09 46.38 49.08 70.49 65.16 67.72 SAP95.42 54.80 79.00 69.30 61.39 27.04 86.97 81.3166.97 60.68 63.30 73.27 71.05 72.14 35 Supplementary Table 8: Comparison of different representation fusion strate- gies for the eight-label knee injury prediction task on the Center A+B+C test dataset. For timestep–block selection settings, both are block=mid2 and timestep=30. Results are reported in terms of per-disease AUROC(%). Boldface indicates the best result in each column. MethodsACLPCLMMLMMCLLCLPDEFFU Simple Concat96.6292.4781.3277.2489.9882.3899.4986.08 Linear Concat96.5292.4681.0576.9489.6781.2699.4285.50 Linear Add96.5492.4780.9376.8689.6281.1499.4285.27 Cross Attention96.5492.2280.6376.3889.4380.1199.1984.50 MPAE96.5492.0880.8676.0689.7082.2999.4686.12 Supplementary Table 9: Comparison of different representation fusion strategies for the eight-label knee injury prediction task on the Center A+B+C knee dataset. For timestep–block selection settings, both are block=mid2 and timestep=30. Results are reported in terms of per-disease AP and overall multi-label metrics (%). Boldface indicates the best result in each column. MethodsACL PCL M LM MCL LCL PD EFFUCPCR CF1 OPOR OF1 Simple Concat 97.32 76.97 85.18 76.79 69.04 36.00 92.30 83.4075.24 66.39 70.22 78.62 74.73 76.63 Linear Concat97.23 77.32 84.91 76.73 68.12 34.06 92.19 82.9373.67 67.53 70.22 77.09 75.39 76.23 Linear Add97.24 76.88 84.76 76.67 67.89 33.91 91.80 82.7573.31 67.28 69.93 76.92 75.26 76.08 Cross Attention 97.29 75.77 84.45 76.12 68.05 33.08 90.47 81.7672.48 67.78 69.83 76.17 75.67 75.92 MPAE97.27 75.12 84.45 75.66 67.77 35.18 91.62 83.5876.28 62.85 67.89 78.70 73.80 76.17 Supplementary Table 10: Task-specific optimal diffusion timestep and network block selec- tions for different musculoskeletal tasks and anatomical sites. TaskAnatomy Setting SagittalCoronalAxial Timestep Block Timestep Block Timestep Block Classification KneeLP-Opt50mid0100mid250mid0 KneeFT-Opt150mid050mid2200mid0 AnkleFT-A-Opt300mid0100mid130mid0 ShoulderFT-S-Opt200mid230mid2100mid1 Segmentation KneeFT-Opt-Sag/Cor50mid250mid2– 36 Supplementary Table 11: Comparisons of per-disease AUROC (%) for the eight-label knee injury prediction task on the Center A-C or D-G test dataset. Boldface indicates the best result under the same setting in each column. DatasetMethodsACL PCL M LM MCL LCL PD EFFU Macro-AUC Center A+B+C 3D-Unet96.28 88.01 72.68 69.56 84.99 79.77 95.94 86.0984.17 3D-ResNet-1896.96 87.56 74.90 70.81 85.00 77.82 97.49 84.4184.37 Ours (LP, LP-Opt) 96.53 92.43 81.04 77.73 90.37 82.84 99.53 86.0288.31 Ours (FT, FT-Opt) 98.06 94.35 87.31 85.62 92.97 85.72 99.63 86.3991.26 Center D+E+F+G 3D-Unet71.21 66.88 73.35 63.49 73.60 65.38 86.06 72.9271.61 3D-ResNet-1876.95 74.38 72.85 64.98 71.55 66.32 90.29 71.7173.63 Ours (LP, LP-Opt) 78.10 87.78 86.06 72.28 76.38 73.27 95.98 78.9581.10 Ours (FT, FT-Opt) 80.02 85.22 89.35 78.65 81.17 73.42 95.95 76.9782.59 Supplementary Table 12: Comparisons of per-disease AP and overall multi-label metrics (%) for the eight-label knee injury prediction task on the Center A+B+C test dataset. Boldface indicates the best result in each column. MethodsACL PCL M LM MCL LCL PD EFFUCP CR CF1 OP OR OF1 3D-Unet97.14 63.96 76.55 68.80 49.67 30.22 68.18 83.1568.52 56.39 59.98 73.54 70.88 72.18 3D-Resnet-1897.70 65.11 79.31 68.65 54.79 29.15 79.76 81.6668.64 62.35 63.61 71.08 76.35 73.62 Ours (LP, LP-Opt) 97.14 78.41 85.23 77.54 69.88 37.18 92.36 82.7375.44 67.76 71.11 78.03 75.20 76.59 Ours (FT, FT-Opt) 98.49 83.38 90.23 85.93 75.79 40.37 93.74 83.4979.48 72.92 75.90 82.36 79.34 80.82 Supplementary Table 13: Comparisons of per-disease AP and overall multi-label metrics (%) for the eight-label knee injury prediction task on the Center D–G test dataset. Boldface indicates the best result in each column. MethodsACL PCL M LM MCL LCL PD EFFUCP CR CF1 OP OR OF1 3D-Unet81.15 32.21 86.10 72.63 52.68 45.46 55.89 95.7069.35 45.33 53.62 77.85 54.96 64.43 3D-Resnet-1883.73 33.96 86.59 71.53 50.77 45.84 71.58 95.5470.25 46.70 53.66 79.57 57.68 66.88 Ours (LP, LP-Opt) 86.06 55.13 94.00 78.25 60.59 54.69 85.25 96.7079.97 51.23 60.81 82.62 60.24 69.68 Ours (FT, FT-Opt) 87.60 61.76 95.36 85.24 68.63 57.98 84.79 96.40 83.61 56.82 66.18 86.61 65.10 74.33 37 Supplementary Table 14: Comparisons of per-disease AP, overall multi-label metrics (%) on the eight- label knee injury prediction task on the A+B+C+D+E+F+G test dataset with different magnetic field strengths. Boldface indicates the best result under the same setting in each column. Field Strength MethodsACL PCL M LM MCL LCL PD EFFU CP CR CF1 OP OR OF1 1.5T 3D-Unet95.09 55.03 70.32 66.30 58.32 44.45 63.02 86.9868.35 55.93 59.82 73.17 64.25 68.42 3D-ResNet-1894.55 53.57 76.06 68.19 61.24 46.35 69.50 84.4367.59 60.24 62.83 73.38 66.14 69.57 Ours (FT, FT-Opt) 97.66 77.42 89.70 83.84 83.09 68.25 89.00 87.0383.01 71.81 76.63 82.83 76.73 79.67 3T 3D-Unet96.30 61.21 77.08 67.72 46.42 24.18 68.26 82.4167.33 58.63 60.89 72.46 70.90 71.67 3D-ResNet-1897.08 63.17 81.07 70.41 49.33 23.17 74.87 81.3565.95 59.92 61.98 74.50 69.01 71.65 Ours (FT, FT-Opt) 98.14 83.63 90.62 86.29 71.19 35.60 93.29 84.1879.12 70.33 74.28 82.77 77.97 80.29 Supplementary Table 15: Comparisons of per-disease AP, overall multi-label metrics (%) for the four-label ankle injury prediction task on the test dataset. Boldface indicates the best result in each column. MethodsATFL CFL ATR OLTCPCR CF1 OPOR OF1 3D-Unet91.94 89.68 32.61 42.3767.02 49.94 45.77 84.06 78.01 80.92 3D-Resnet-1893.04 90.00 29.45 54.5561.45 69.77 63.55 73.93 87.04 79.95 Ours (FT, FT-A-Opt) 95.73 93.49 40.90 72.26 85.46 60.82 60.59 83.36 83.52 83.44 Supplementary Table 16: Comparisons of per-disease AP, overall multi-label metrics (%) for the seven-label shoulder injury prediction task on the test dataset. Boldface indicates the best result in each column. MethodsSSP ISP SSC LHBT injury AC LHBT effusion SASDCP CR CF1 OP OR OF1 3D-Unet99.19 80.19 65.05 54.97 74.4717.1792.8765.58 60.19 61.37 79.73 75.86 77.75 3D-Resnet-1899.32 80.32 66.85 55.59 75.5318.2792.5964.37 63.69 63.08 77.73 79.69 78.70 Ours (FT, FT-S-Opt) 99.54 84.10 74.27 63.20 82.76 30.73 94.0672.94 66.66 67.86 81.00 79.57 80.28 Supplementary Table 17: Comparisons of per-disease AUROC(%) for the four-label ankle injury prediction task on the test dataset. Boldface indicates the best result in each col- umn. MethodsATFL CFL ATR OLT 3D-Unet65.31 68.93 74.88 61.06 3D-Resnet-1866.62 69.16 75.78 68.54 Ours (FT, FT-A-Opt) 77.52 77.86 82.00 80.43 38 Supplementary Table 18: Comparisons of per-disease AUROC(%) for the seven-label shoulder injury prediction task on the test dataset. Boldface indicates the best result in each column. MethodsSSP ISP SSC LHBT injury AC LHBT effusion SASD 3D-Unet81.72 74.59 73.66 78.25 86.0065.8280.25 3D-Resnet-1883.48 74.92 74.97 78.40 86.3067.0379.01 Ours (FT, FT-S-Opt) 88.49 79.08 79.54 81.49 91.03 75.77 83.38 39 Supplementary Table 19 : Comparison of per-disease AP, overall multi-label metrics (%) for the eight-label knee injury prediction task using different modalities (MRI-only, EHR-only, and MRI+EHR) on the Center A+B+C+D+E+F+G test set. Boldface indicatesthe best result under same setting in each column. Dataset Modality Methods ACL PCL M LM MCL LCL PD EFFU CP CR CF1 OP OR OF1 Center A+B+ C+D+E+F+G MRI-only 3D-Unet 95.92 59.30 75.74 67.25 49.17 30.24 65.99 83.44 67.42 57.75 60.49 72.59 68.93 70.71 3D-ResNet-18 96.14 55.90 79.40 69.71 52.62 28.70 70.20 81.62 63.50 63.08 62.80 70.55 73.93 72.20 Ours (LP, LP-Opt) 96.31 74.47 85.53 76.47 67.68 34.80 89.01 84.76 75.54 64.27 68.98 78.52 73.05 75.68 Ours (FT, FT-Opt) 97.92 81.56 90.44 85.73 74.64 44.68 91.39 84.96 79.52 70.85 74.69 82.18 77.59 79.82 EHR-only EHR-only 58.21 11.09 56.58 49.63 16.75 19.53 7.04 56.43 33.09 32.53 29.08 55.64 55.71 55.67 EHR + MRI Ours (LP, LP-Opt) 96.39 73.61 85.31 76.34 67.49 37.96 88.41 85.23 78.31 62.42 68.54 79.67 72.47 75.90 Ours (FT, FT-Opt) 97.89 81.79 90.42 85.70 74.60 46.06 91.42 85.50 81.17 69.23 74.40 83.28 76.81 79.91 40 References [1] Musahl, V., Karlsson, J.: Anterior cruciate ligament tear. New England Journal of Medicine 380, 2341–2348 (2019) [2] Ouyang, Y., Dai, M.: Global, regional, and national burden of knee osteoarthritis: findings from the global burden of disease study 2021 and projections to 2045. Journal of Orthopaedic Surgery and Research 20, 766 (2025) [3] Poulsen, E., Goncalves, G.H., Bricca, A., Roos, E.M., Thorlund, J.B., Juhl, C.B.: Knee osteoarthritis risk is increased 4-6 fold after knee injury: a systematic review and meta-analysis. British Journal of Sports Medicine 53, 1454–1463 (2019) [4] World Health Organization: Musculoskeletal Conditions. World Health Organi- zation, Geneva (2022) [5] Kijowski, R., Roemer, F., Englund, M., Tiderius, C.J., Sward, P., Frobell, R.B.: Imaging following acute knee trauma. Osteoarthritis and Cartilage 22, 1429–1443 (2014) [6] Nacey, N.C., Geeslin, M.G., Miller, G.W., Pierce, J.L.: Magnetic resonance imag- ing of the knee: An overview and update of conventional and state of the art imaging. Journal of Magnetic Resonance Imaging 45, 1257–1275 (2017) [7] Wang, D.Y., Liu, S.G., Ding, J., et al.: A deep learning model enhances clinicians’ diagnostic accuracy to more than 96% for anterior cruciate ligament ruptures on magnetic resonance imaging. Arthroscopy 40, 1197–1205 (2024) [8] Tran, A., Lassalle, L., Zille, P., et al.: Deep learning to detect anterior cruci- ate ligament tear on knee mri: multi-continental external validation. European Radiology 32, 8394–8403 (2022) [9] Botnari, A., Kadar, M., Patrascu, J.M.: A comprehensive evaluation of deep learn- ing models on knee mris for the diagnosis and classification of meniscal tears: A systematic review and meta-analysis. Diagnostics (Basel) 14 (2024) [10] Lin, D.J., Schwier, M., Geiger, B., Raithel, E., Busch, H., Fritz, J., Kline, M., Brooks, M., Dunham, K., Shukla, M., Alaia, E.F., Samim, M., Joshi, V., Walter, W.R., Ellermann, J.M., Ilaslan, H., Rubin, D., Winalski, C.S., Recht, M.P.: Deep learning diagnosis and classification of rotator cuff tears on shoulder mri. Investigative Radiology 58(6), 405–412 (2023) https://doi.org/10.1097/RLI. 0000000000000951 [11] Sun, J., Cao, Y., Zhou, Y., Qi, B.: Leveraging spatial dependencies and multi- scale features for automated knee injury detection on mri diagnosis. Frontiers in Bioengineering and Biotechnology 13, 1590962 (2025) 41 [12] Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceedings of the 2019 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4171–4186. Association for Computational Linguistics, Minneapolis, Minnesota (2019). https://doi.org/10. 18653/v1/N19-1423 . https://aclanthology.org/N19-1423/ [13] Radford, A., Narasimhan, K.: Improving language understanding by generative pre-training. (2018). https://api.semanticscholar.org/CorpusID:49313245 [14] Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 140–114067 (2019) [15] He, K., Chen, X., Xie, S., Li, Y., Doll’ar, P., Girshick, R.B.: Masked autoencoders are scalable vision learners. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15979–15988 (2021) [16] Chen, M., Radford, A., Wu, J., Jun, H., Dhariwal, P., Luan, D., Sutskever, I.: Gen- erative pretraining from pixels. In: International Conference on Machine Learning (2020). https://api.semanticscholar.org/CorpusID:219781060 [17] Yan, S., Yu, Z., Primiero, C., et al.: A multimodal vision foundation model for clinical dermatology. Nature Medicine 31, 2691–2702 (2025) https://doi.org/10. 1038/s41591-025-03747-y [18] Jiang, N., Ji, H., Guan, Z., et al.: A deep learning system for detecting silent brain infarction and predicting stroke risk. Nature Biomedical Engineering 9, 1907–1919 (2025) https://doi.org/10.1038/s41551-025-01413-9 [19] Christensen, M., Vukadinovic, M., Yuan, N., et al.: Vision–language foundation model for echocardiogram interpretation. Nature Medicine 30, 1481–1488 (2024) https://doi.org/10.1038/s41591-024-02959-y [20] Zhou, Y., Chia, M.A., Wagner, S.K., et al.: A foundation model for generalizable disease detection from retinal images. Nature 622, 156–163 (2023) https://doi. org/10.1038/s41586-023-06555-x [21] Deltadahl, S., Gilbey, J., Van Laer, C., et al.: Deep generative classification of blood cell morphology. Nature Machine Intelligence 7, 1791–1803 (2025) https: //doi.org/10.1038/s42256-025-01122-7 [22] Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. ArXiv abs/2006.11239 (2020) [23] Song, Y., Sohl-Dickstein, J.N., Kingma, D.P., Kumar, A., Ermon, S., Poole, B.: 42 Score-based generative modeling through stochastic differential equations. ArXiv abs/2011.13456 (2020) [24] Wang, J., Wang, K., Yu, Y., et al.: Self-improving generative foundation model for synthetic medical image generation and clinical applications. Nature Medicine 31, 609–617 (2025) https://doi.org/10.1038/s41591-024-03359-y [25] Bl ̈uthgen, C., Chambon, P., Delbrouck, J.-B., Sluijs, R., Po lacin, M., Zam- brano Chaves, J.M., Abraham, T.M., Purohit, S., Langlotz, C.P., Chaudhari, A.S.: A vision–language foundation model for the generation of realistic chest x-ray images. Nature Biomedical Engineering 9(4), 494–506 (2024) https://doi. org/10.1038/s41551-024-01246-y [26] Xiang, W., Yang, H., Huang, D., Wang, Y.: Denoising diffusion autoencoders are unified self-supervised learners. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (2023) [27] Baranchuk, D., Rubachev, I., Voynov, A., Khrulkov, V., Babenko, A.: Label- efficient semantic segmentation with diffusion models. ArXiv abs/2112.03126 (2021) [28] Li, J., Qian, K., Liu, J.-y., Huang, Z.-f., Zhang, Y., Zhao, G., Wang, H., Li, M., Liang, X., Zhou, F., Yu, X.-z., Li, L., Wang, X., Yang, X., Jiang, Q.: Identifica- tion and diagnosis of meniscus tear by magnetic resonance imaging using a deep learning model. Journal of Orthopaedic Translation 34, 91–101 (2022) [29] Roblot, V., Giret, Y., Antoun, M.B., Morillot, C., Chassin, X., Cotten, A., Zer- bib, J., Fournier, L.: Artificial intelligence to diagnose meniscus tears on mri. Diagnostic and interventional imaging 100 4, 243–249 (2019) [30] Fritz, B., Marbach, G., Civardi, F., Fucentese, S.F., Pfirrmann, C.W.A.: Deep convolutional neural network-based detection of meniscus tears: comparison with radiologists and surgery as standard of reference. Skeletal Radiology 49(8), 1207– 1217 (2020) https://doi.org/10.1007/s00256-020-03410-2 . Erratum in: Skeletal Radiol. 2020 Aug;49(8):1219. doi: 10.1007/s00256-020-03458-0. Epub 2020 Mar 13 [31] Li, Y.Z., Wang, Y., Fang, K.B., Zheng, H.Z., Lai, Q.Q., Xia, Y.F., Chen, J.Y., Dai, Z.S.: Automated meniscus segmentation and tear detection of knee mri with a 3d mask-rcnn. European Journal of Medical Research 27(1), 247 (2022) https: //doi.org/10.1186/s40001-022-00883-w [32] Rizk, B., Brat, H.G., Zille, P., Guillin, R., Pouchy, C., Adam, C., Ardon, R., D’Assignies, G.: Meniscal lesion detection and characterization in adult knee mri: A deep learning model approach with external validation. Physica medica : PM : an international journal devoted to the applications of physics to medicine and biology : official journal of the Italian Association of Biomedical Physics 83, 64–71 43 (2021) [33] Singh, S.P., Wang, L., Gupta, S., Goli, H., Padmanabhan, P., Guly’as, B.: 3d deep learning on medical images: A review. Sensors (Basel, Switzerland) 20 (2020) [34] Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. ArXiv abs/2010.11929 (2020) [35] Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10674–10685 (2021) [36] Luo, C.: Understanding diffusion models: A unified perspective. ArXiv abs/2208.11970 (2022) [37] C ̧ i ̧cek, ̈ O., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O.: 3d u-net: Learning dense volumetric segmentation from sparse annotation. In: International Conference on Medical Image Computing and Computer-Assisted Intervention (2016). https://api.semanticscholar.org/CorpusID:2164893 [38] Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H.R., Xu, D.: Unetr: Transformers for 3d medical image segmentation. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 574–584 (2022) [39] Hara, K., Kataoka, H., Satoh, Y.: Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6546–6555 (2017) [40] Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., Hu, H.: SimMIM: A Simple Framework for Masked Image Modeling (2022). https://arxiv.org/abs/ 2111.09886 [41] Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709 (2020) [42] Chen, T., Kornblith, S., Swersky, K., Norouzi, M., Hinton, G.: Big Self-Supervised Models are Strong Semi-Supervised Learners (2020). https://arxiv.org/abs/2006. 10029 [43] Grill, J.-B., Strub, F., Altch ́e, F., Tallec, C., Richemond, P.H., Buchatskaya, E., Doersch, C., Pires, B.A., Guo, Z.D., Azar, M.G., Piot, B., Kavukcuoglu, K., Munos, R., Valko, M.: Bootstrap your own latent: A new approach to self-supervised Learning (2020). https://arxiv.org/abs/2006.07733 44 [44] Caron, M., Touvron, H., Misra, I., J ́egou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging Properties in Self-Supervised Vision Transformers (2021). https: //arxiv.org/abs/2104.14294 [45] Li, X., Zhang, Z., Li, X., Chen, S., Zhu, Z., Wang, P., Qu, Q.: Understanding representation dynamics of diffusion models via low-dimensional modeling. ArXiv abs/2502.05743 (2025) [46] Tang, L., Jia, M., Wang, Q., Phoo, C.P., Hariharan, B.: Emergent Correspondence from Image Diffusion (2023). https://arxiv.org/abs/2306.03881 [47] Sun, Y., Tan, W., Gu, Z., et al.: A data-efficient strategy for building high- performing medical foundation models. Nature Biomedical Engineering 9, 539– 551 (2025) https://doi.org/10.1038/s41551-025-01365-0 [48] Ding, T., Wagner, S.J., Song, A.H., et al.: A multimodal whole-slide foundation model for pathology. Nature Medicine 31, 3749–3761 (2025) https://doi.org/10. 1038/s41591-025-03982-3 [49] Ma, J., Guo, Z., Zhou, F., et al.: A generalizable pathology foundation model using a unified knowledge distillation pretraining framework. Nature Biomedical Engineering (2025) https://doi.org/10.1038/s41551-025-01488-4 [50] Xiang, J., Wang, X., Zhang, X., et al.: A vision–language foundation model for precision oncology. Nature 638, 769–778 (2025) https://doi.org/10.1038/ s41586-024-08378-w [51] Yan, S., Yu, Z., Primiero, C., et al.: A multimodal vision foundation model for clinical dermatology. Nature Medicine 31, 2691–2702 (2025) https://doi.org/10. 1038/s41591-025-03747-y [52] Dorjsembe, Z., Pao, H.-K., Odonchimed, S., Xiao, F.: Conditional diffusion models for semantic 3d brain mri synthesis. IEEE Journal of Biomedical and Health Informatics 28(7), 4084–4093 (2024) https://doi.org/10.1109/JBHI.2024. 3385504 [53] Cardoso, M.J., Li, W., Brown, R., Ma, N., Kerfoot, E., Wang, Y., Murrey, B., Myronenko, A., Zhao, C., Yang, D., Nath, V., He, Y., Xu, Z., Hatamizadeh, A., Myronenko, A., Zhu, W., Liu, Y., Zheng, M., Tang, Y., Yang, I., Zephyr, M., Hashemian, B., Alle, S., Darestani, M.Z., Budd, C., Modat, M., Vercauteren, T., Wang, G., Li, Y., Hu, Y., Fu, Y., Gorman, B., Johnson, H., Genereaux, B., Erdal, B.S., Gupta, V., Diaz-Pinto, A., Dourson, A., Maier-Hein, L., Jaeger, P.F., Baumgartner, M., Kalpathy-Cramer, J., Flores, M., Kirby, J., Cooper, L.A.D., Roth, H.R., Xu, D., Bericat, D., Floca, R., Zhou, S.K., Shuaib, H., Farahani, K., Maier-Hein, K.H., Aylward, S., Dogra, P., Ourselin, S., Feng, A.: MONAI: An open-source framework for deep learning in healthcare (2022). https://arxiv.org/ abs/2211.02701 45