Paper deep dive
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
Alejandro Vergara-Richart, Xavier Rafael-Palou, Almudena Fuster-Matanzo, Ignacio Iborra Roncales, Ángel Alberich-Bayarri, Ana Jiménez-Pastor
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/9/2026, 7:16:26 AM
Summary
A scoping review analyzing the development, evaluation, and clinical translation of Vision Foundation Models (VFMs) in radiology, highlighting trends in data scale, architectural design, pretraining strategies, and downstream task adaptability, while identifying gaps in standardization and deployment readiness.
Entities (10)
Relation Signals (9)
Vision Foundation Models (VFM) → appliesto → Radiology
confidence 98% · Vision foundation models (VFMs) are increasingly being developed for radiological imaging
Vision Foundation Models (VFM) → employspretraining → Self-Supervised Learning (SSL)
confidence 95% · Self-supervised learning (SSL) has become the dominant strategy for pretraining models in medical imaging
Vision Foundation Models (VFM) → evaluatedon → Classification
confidence 94% · Evaluation focused mainly on segmentation and classification
Vision Foundation Models (VFM) → evaluatedon → Segmentation
confidence 94% · Evaluation focused mainly on segmentation and classification
Clinical Translation → constrainedby → Limited Data Representativeness
confidence 93% · clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Vision Foundation Models (VFM) → usesarchitecture → Transformer Architecture
confidence 92% · Transformer-based architectures and self-supervised pretraining predominated
Brain MRI → usedin → Vision Foundation Models (VFM)
confidence 91% · Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray
Chest X-ray → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2607.07219v1
- Canonical: https://arxiv.org/abs/2607.07219v1
Trouble viewing inline? Open PDF directly →
Full Text
112,730 characters extracted from source content.
Expand or collapse full text
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation Alejandro Vergara-Richart1,2,∗, Xavier Rafael-Palou1,∗, Almudena Fuster-Matanzo1, Ignacio Iborra Roncales1, Ángel Alberich-Bayarri1, Ana Jiménez-Pastor1 ( 1Quantitative Imaging Biomarkers in Medicine (Quibim S.L.), Valencia, Spain 2Universitat Politècnica de València, Valencia, Spain ∗These authors contributed equally to this work. Correspondence: Alejandro Vergara-Richart, alejandrovergara@quibim.com ) Abstract Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMA-ScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation. Keywords: foundation models; radiology; review; medical imaging 1 Introduction Originating in natural language processing and popularized through large language models [7], foundation models (FMs) represent a shift toward learning general-purpose representations from large-scale data. Unlike task-specific models optimized for a single objective, FMs are trained on diverse and extensive datasets with the objective of capturing transferable features that can be adapted to multiple downstream applications. This paradigm shift has been extended to computer vision through vision foundation models (VFMs), which are high-capacity neural networks trained on large-scale, heterogeneous, and often unlabeled imaging data to learn reusable representations that can be efficiently adapted to diverse specific tasks [30]. These properties are particularly relevant in radiology imaging analysis, where data scarcity, high annotation costs, and substantial variability across institutions, scanners and acquisition protocols hinder predictive model generalization and clinical translation [74]. Furthermore, such variability induces distributional shifts that significantly degrade performance under external validation and real-world deployment settings. Self-supervised learning (SSL) has become the dominant strategy for pretraining models in medical imaging due to its ability to exploit the intrinsic structure of unlabeled data [19]. In medical imaging, SSL approaches are typically grouped into two broad families: (i) reconstruction-based methods, such as masked image modeling and related generative objectives [46], and (i) contrastive learning frameworks that promote invariance across augmented views, modalities, or patient-level representations [79]. Complementary, weakly supervised learning strategies, which exploit coarse or inexact labels, have also contributed to scaling representation learning in radiology [56]. Built upon these pretraining strategies, VFMs are typically implemented using large-capacity architectures (often based on transformer designs [22]) and trained on increasingly large and diverse datasets. Once pretrained, these models are adapted to downstream clinical tasks through a range of transfer learning strategies, including zero-shot (ZS) and few-shot (FS) inference, as well as full or partial fine-tuning [18]. More recently, parameter-efficient fine-tuning (PEFT) methods, such as low-rank adaptation [33] and adapter-based approaches [85], have gained increasing attention. These techniques enable efficient task adaptation while reducing computational cost and mitigating overfitting, which is particularly relevant in data-scarce clinical environments. Despite rapid progress, the concept of “vision foundation model” remains inconsistently defined across the radiology literature. Models referred to as VFMs vary substantially in dataset scale and diversity, pretraining strategy, architectural design, and downstream evaluation protocols [12]. This terminology ambiguity complicates systematic identification of relevant models, limits meaningful cross-study comparisons, and obscures a rigorous understanding of their strengths and weaknesses. Several reviews have recently examined FMs in medical imaging from complementary perspectives, yet important aspects remain incompletely characterized. A recent narrative review described training paradigms, adaptation strategies, and evaluation frameworks for radiology FMs; however, it provided limited characterization of dataset composition, architectural trends, and the distribution of downstream tasks across the literature [60]. Another review offered a broader perspective spanning modalities beyond radiology such as pathology and ophthalmology, but at the cost of modality- and anatomy-specific characterization within radiology itself [96]. A survey of self-supervised pretraining examined diagnostic tasks across X-ray, CT, MRI, and ultrasound; however, its focus on a single methodological paradigm placed less emphasis on the broader landscape of FM architectures, downstream task coverage, and challenges related to clinical translation [76]. A systematic review and meta-analysis of vision-language FMs quantified performance across classification, segmentation, report generation, and visual question answering, but its coverage of vision-only radiology FMs remained limited [68]. Another recent review examined the fundamentals, applications, opportunities, challenges, risks, and prospects of FMs for radiology with particular attention to clinical adoption and regulatory considerations, although its narrative format did not provide a systematic evidence mapping of architectural choices, dataset characteristics, and external validation gaps [1]. Taken together, these contributions highlight the growing interest in radiology foundation models, while underscoring the absence of a comprehensive and structured overview focused specifically on radiology VFMs. To address this need, we conducted a scoping review that organizes the literature around three core pillars: (i) access to large-scale and heterogeneous imaging datasets, (i) scalable pretraining strategies and model architectures, and (i) adaptation approaches for downstream clinical tasks. By synthesizing the literature around these interconnected pillars, this review aims to support more consistent characterization and comparison of VFMs in radiology, and to identify methodological and clinical gaps that warrant further investigation. Complementing this framework, we examine the extent to which included models address key principles of the FUTURE-AI guidelines [44], namely robustness, fairness, universality, explainability, traceability, and usability. This perspective allows the review to situate methodological advances within broader considerations of safe, and clinically effective deployment, and to identify where current VFM development diverges from established standards for trustworthy medical AI. 2 Methods A review protocol for this study was not registered. 2.1 Eligibility criteria The eligibility criteria were defined following PRISMA-ScR [75] recommendations to ensure transparency and reproducibility. Studies were screened according to the following inclusion criteria. i. Peer-reviewed journal articles in English published between January 2017 and March 2026. The starting date was selected to capture the emergence of modern large-scale representation learning paradigms and FM-like approaches in computer vision [77, 26]. i. Studies involving models trained on radiological imaging data, including radiography (X-ray), computed tomography (CT), magnetic resonance imaging (MRI), nuclear medicine (PET/CT), and ultrasound (US). i. Studies employing pretraining on imaging datasets intended to learn generalizable representations transferable across downstream tasks. This includes self-supervised learning, weakly supervised learning or supervised pretraining on large datasets. iv. Studies proposing a new VFM or reporting the adaptation and fine-tuning of an existing VFM for radiology applications. v. Studies evaluating the VFM on at least one radiology-related downstream task, such as classification, segmentation, detection, report generation, or survival prediction. vi. Studies explicitly described by the authors as VFM were included despite partially meeting the above criteria, provided that this classification was supported through a journal peer-review process. Conversely, studies were excluded following the following criteria: i. Studies based exclusively on shallow learning methods or conventional machine learning approaches without deep neural networks. i. Models developed exclusively for a single downstream task without evidence of representation, transfer, adaptation or reuse across tasks, datasets, or clinical settings. i. Studies using fully frozen pretrained models without proposing methodological adaptations, architectural modifications, or representation learning contributions. iv. Studies focused solely on clinical application, benchmarking, or comparative evaluation without methodological development of a FM. v. Studies proposing training pipelines or frameworks without developing a radiology VFM. vi. Studies that did not include a pretraining stage intended to learn transferable representations. vii. Studies exclusively involving non-radiological imaging domains, including histopathology, ophthalmology, microscopy, endoscopy, or dermatology imaging. viii. Editorials, commentaries, letters, conference abstracts, tutorials, surveys, and non-peer-reviewed publications. ix. Studies whose primary objective was not the development, adaptation, or evaluation of radiology-oriented VFMs. 2.2 Information Sources and Search Strategy A literature search for eligible publications was conducted in March 2026 across PubMed, Scopus, IEEE, and EMBASE. The key search terms were based on a combination of two major terms: “foundation models” and “medical imaging” (see Supplementary Material Section 7.1). Search terms were formulated to include radiology-focused VFMs, while exclusion terms were applied to remove studies with multimodal, non-image-based, or non-radiology FMs. 2.2.1 Search strategy Full electronic search strategy for one database is provided in Supplementary Material Section 7.2. 2.2.2 Selection of Sources Literature search and study selection were independently conducted by two reviewers following the predefined eligibility criteria. Covidence review management software (Veritas Health Innovation, Melbourne, Australia) was used to support record management, study screening, and data extraction. After removal of duplicate records, all retrieved studies identified through the search strategy underwent a two-stage screening process consisting of (i) title and abstract screening and (i) full-text review. Each study was independently assessed by both reviewers at both screening stages. Disagreements were resolved through consensus discussions during dedicated review meetings. 2.3 Data charting Data extracted included the following: (1) general information: title, authors, affiliations, journal, publication date, and DOI; (2) study objective: primary objective, secondary objectives, and main contributions; (3) data characteristics: number of slices/scans/cases, imaging modality, organ or anatomical region, pathology or condition, number of centers, dataset source, image characteristics, dataset names, test data, annotations, preprocessing steps, and data quality information; (4) methods: model type, model name, architecture, backbone, pretraining strategy, fine-tuning strategy, model input, task objective, and prompt; (5) evaluation: downstream tasks, evaluation strategy, main performance metrics and scores, and comparisons with the state of the art; (6) implementation details: computing resources and model parameters; (7) FUTURE-AI principles: fairness, universality, traceability, usability, robustness, and explainability; (8) identified gaps and proposed future steps; and (9) community impact: license information, data, model, and code availability. 2.4 Data items and synthesis of results Given the exploratory nature of this scoping review and the marked heterogeneity across datasets, tasks, evaluation protocols, and reporting standards, findings were synthesized descriptively rather than through quantitative meta-analysis. 2.5 Conceptual framework for evidence mapping We propose a conceptual framework grounded in three pillars to unify the core characteristics and operational requirements of VFMs in radiology: (i) access to large-scale and diverse imaging data, (i) scalable pretraining strategies and model architectures, and (i) adaptation and transferability to downstream clinical tasks. Using this three-pillar framework, we mapped the current evidence to provide a structured characterization of radiology VFMs and to enable consistent comparison across diverse methodological approaches and downstream applications. 2.5.1 Pillar 1: Large-scale and heterogeneous dataset The first pillar captures the scale and composition of the training data used to develop a VFM. Such models require large-scale and clinically heterogeneous imaging datasets encompassing multiple modalities, anatomical regions, pathologies, and patient populations, ideally aggregated across institutions, scanners, and acquisition protocols. Increasing dataset scale and heterogeneity facilitate representation learning, broader domain coverage, and improved generalization capabilities. Table 1 summarizes the data-related variables extracted from each included study under this pillar. Table 1: Dataset scale and heterogeneity items. Item Description # Data (Unit) Number of imaging samples used to train the model categorized into slices or scans. Modality Imaging modalities represented in the dataset typically CT, X-ray, US or MRI. Localization Anatomical regions included in the dataset. "Multiple" refers to different localizations. Clinical indication Clinical conditions or patient populations represented. "Multiple" implies different pathologies. Centers Number of contributing data sources or institutions (single center or multicenter) Availability Whether the dataset is publicly available or proprietary (public and private) 2.5.2 Pillar 2: Large-scale pretraining and architecture The second pillar focuses on scalable pretraining strategies and model architectures. VFMs aim to learn robust, task-agnostic representations that can be efficiently transferred across diverse clinical applications. Consequently, model architectures must accommodate large data volumes and increasing model capacity, capture long-range dependencies, and efficiently process high-resolution and volumetric medical imaging data. Scalable pretraining paradigms therefore play a central role in enabling transferable representations across complex radiological domains. Table 2 summarizes the model and pretraining-related variables extracted from each included study under this pillar. Table 2: Pretraining and architecture items. Item Description Model name Model’s published name or identifier. Architecture Backbone architecture used for the model. Pretraining strategy Learning paradigm used. Hardware Computational resources for pretraining. Parameters Total number of trainable parameters. Weights Whether pretrained weights are publicly accessible (Yes/No). Code Whether training or inference code is publicly available (Yes/No). 2.5.3 Pillar 3: Scalability to multiple applications The third pillar characterizes the VFMs’ ability to generalize and adapt across multiple downstream clinical tasks while maintaining meaningful performance. This pillar reflects the practical utility, transferability and, robustness of learned representations across diverse radiological applications. Table 3 provides a summary of the downstream evaluation and adaptation variables extracted from each included study under this pillar. Table 3: Downstream evaluation items. Item Description Tasks Number of downstream evaluation tasks performed. Task types Categories of tasks addressed (e.g., classification, segmentation). Fine-tuning strategy Fine-tuning strategy used, if applicable. Generalization Type of generalization shift evaluated in the downstream tasks. Zero-shot (ZS) Indicates whether zero-shot evaluation was performed (Yes/No). Few-shot (FS) Indicates whether few-shot evaluation was conducted (Yes/No). Baseline outperforming (PT/FM ↑) Whether the pretrained or FMs outperformed baseline pretrained models (Yes/Partially/No). Task-specific outperforming (Task-spec. ↑) Whether performance exceeded that of task-specific models (Yes/Partially/No) 2.6 Future-AI principles alignment To complement the technical mapping of VFM studies in radiology, we additionally characterized their reported alignment with principles for trustworthy and deployable artificial intelligence (AI) in healthcare. For this purpose, we adopted the FUTURE-AI framework [44], an international consensus guideline developed to support the responsible development, evaluation, and deployment of AI systems across the healthcare lifecycle. The framework defines six core principles—fairness, universality, traceability, usability, robustness, and explainability—and outlines best-practice recommendations spanning technical, clinical, ethical, and regulatory domains. Alignment with FUTURE-AI principles was descriptively assessed as "reported" or "not reported". A principle was considered reported if the study explicitly described at least one design choice, evaluation procedure, or implementation element corresponding to that principle. No attempt was made to assess completeness, quality or degree of adherence, as the objective was to map reporting patterns. 3 Results 3.1 Study Selection The study selection process is summarized in the PRISMA flow diagram in Figure 1. The database search identified 2,596 records, including 1322 from Scopus, 592 from PubMed, 567 from Embase, and 115 from IEEE. After removal of 1200 duplicate records (1191 identified automatically using Covidence and nine identified manually), 1396 unique studies remained for title and abstract screening. Of these, 1223 studies were excluded, leaving 173 articles for full-text eligibility assessment. Following full-text review, 106 studies were excluded according to the criteria defined in Section 2.1. The main reasons for exclusion included preprints, lack of authorization, conference publications, use of non-radiological imaging modalities, incorporation of text-based data, focus on methodological model training frameworks, use of shallow model architectures, and study objectives not aligned with the development of a radiology VFM. Ultimately, 67 studies met the inclusion criteria and were included in the final scoping review. The temporal distribution of included studies shows that eligible studies began to meet the inclusion criteria from late 2023 onward, with a notable shift in trend beginning in 2025, when the number of included studies increased markedly (Figure 2C). Figure 1: PRISMA Flow Diagram 3.2 Pillar 1: Large-scale and heterogeneous dataset Across the 67 included studies, substantial variability was observed in both data dimensionality (2D, 3D and 4D) and dataset scale, with cohort sizes ranging from <100k to >1M samples. For consistency, dataset sizes were reported using units aligned with the underlying data representation: 2D datasets are quantified in slices, whereas 3D and 4D datasets are generally quantified in scans. However, in two studies [6, 47] involving volumetric data, dataset size was reported only in terms of extracted slices, and this convention was therefore retained. An additional category, denoted as ‘multiple’, was introduced for studies combining heterogeneous dimensionalities. To facilitate comparison across studies, datasets were further grouped by scale magnitude. For slice-based datasets, small, medium, and large scales corresponded to <100k, 100k–1M, and >1M samples, respectively, whereas for scan-based datasets the corresponding ranges were <10k, 10k–100k, and >100k. Notably, 3D and 4D acquisitions generally encode substantially richer spatial and temporal information than 2D data, even when the reported sample counts are similar. Among slice-based datasets, six studies (8.9%) were categorized as small [6, 2, 48, 10, 20, 35], seven (10.3%) as medium [50, 39, 97, 51, 91, 92, 55], and only five (7.4%) used large-scale datasets [80, 47, 37, 36, 38]. Scan-based datasets were predominantly distributed across the small- and medium-scale categories, comprising 15 (22.3%) [72, 59, 31, 65, 17, 64, 98, 87, 69, 28, 53, 89, 5, 88, 73] and 17 (25.3%) [93, 90, 21, 23, 82, 70, 25, 16, 43, 49, 11, 54, 45, 86, 71, 81, 95] studies, respectively. Only two studies (3%) employed large datasets [24, 83]. Studies categorized as "multiple" were predominantly large-scale when dataset size was reported in slices (7, 10.4%) [63, 15, 29, 52, 3, 94, 34]. In contrast, three scan-based multimodal datasets (4.4%) were medium-scale [84, 78, 99], with the exception of [32], which was categorized as small. Modality composition further differentiated studies. Most studies (39, 58.2%) relied on a single imaging modality [93, 72, 6, 59, 80, 90, 31, 2, 65, 21, 17, 47, 48, 10, 43, 50, 53, 37, 39, 54, 45, 89, 71, 81, 24, 5, 51, 95, 91, 88, 92, 36, 73, 83, 34, 35, 55, 38]. These included 12 studies using single-parametric MRI, 12 CT, nine US, and five X-ray datasets. In contrast, 26 studies (38.8%) explicitly incorporated multimodal imaging data [32, 63, 15, 23, 82, 70, 64, 25, 9, 98, 16, 87, 69, 28, 29, 49, 11, 52, 97, 86, 20, 84, 3, 94, 78, 99], combining modalities such as CT, single- or multiparametric MRI, PET, US, X-rays, as well as additional medical imaging domains including histopathology, optical coherence tomography, and endoscopy. Anatomical coverage also varied. Single-location studies (34, 50.7%) focused on specific regions such as the chest [80, 50, 93, 24, 54, 51, 91, 92], colorectum [88], heart [34], tongue [2], thyroid [39], and breast [49]. Among these, the majority (21, 61.7%) focused on the brain [6, 90, 31, 32, 21, 23, 82, 17, 70, 98, 16, 87, 69, 11, 45, 89, 71, 81, 5, 95, 20]. In contrast, 30 studies (44.8%) investigated multiple anatomical locations [72, 59, 63, 15, 65, 64, 25, 9, 47, 48, 10, 28, 29, 43, 53, 37, 52, 97, 86, 84, 36, 73, 83, 3, 94, 78, 99, 35, 55, 38]. The most frequently represented anatomical regions were the brain, thorax, and abdomen, followed by frequent inclusion of liver, kidney, breast, heart, bone, and vascular structures. Several studies also explicitly addressed whole-body or large multi-organ settings. Pathology coverage was similarly heterogeneous. Multi-pathology studies (22, 32.8%) encompassed diverse and often pathology-agnostic conditions [63, 15, 65, 64, 47, 29, 54, 97, 86, 95, 20, 84, 36, 73, 83, 3, 94, 78, 99, 35, 55, 38]. Within this group, the most frequently co-occurring disease categories were cancer, neurologic conditions, and thoracic or abdominal diseases, often combined with healthy anatomy. Single-disease studies were distributed across pulmonary/thoracic diseases (7, 10.4%) [93, 80, 24, 50, 51, 91, 92], neurologic and psychiatric conditions (12, 17.9%) [90, 31, 21, 17, 70, 98, 16, 87, 89, 71, 81, 5] and cancer-related applications (13, 19.4%) [6, 59, 25, 9, 48, 10, 53, 37, 49, 11, 52, 45, 88]. A smaller number of studies focused on healthy or non-pathology data (5, 7.4%) [72, 32, 23, 82, 43]. Specialized single-disease categories appeared less frequently, such as speech-related conditions [2], musculoskeletal disorders [28], pediatric applications [69], and endocrine disorders in [39]. Institutional heterogeneity revealed that most studies used multicenter data (56, 83.5%) [93, 72, 80, 90, 31, 32, 63, 15, 65, 21, 17, 64, 25, 9, 47, 48, 98, 16, 87, 10, 69, 29, 43, 50, 53, 37, 49, 52]. However, only seven studies (10.44%) were single-center [59, 2, 23, 70, 28, 39, 11]. Notably, some single-center studies still achieved substantial acquisition diversity, for instance, [23] used data from 14 scanners across five manufacturers, [70] included 14 scanners, and [11] also reported 14 scanners within a single institution. Regarding accessibility, publicly available datasets were used in 37 studies (55.2%) [72, 59, 90, 31, 32, 63, 15, 65, 21, 17, 64, 9, 47, 48, 16, 87, 29, 43, 50, 53, 52, 45, 89, 81, 5, 51, 95, 20, 91, 84, 73, 83, 3, 94, 99, 35, 55] and combined with private data in 14 studies (20.9%) [93, 25, 98, 10, 37, 54, 97, 86, 71, 92, 36, 78, 34, 38], whereas 12 studies (17.9%) relied exclusively on proprietary datasets [80, 2, 23, 82, 70, 69, 28, 39, 49, 11, 24, 88]. Table 4 and Figure 2 A-G summarize the distribution of dataset scale and heterogeneity across the included studies. Table 4: Summary of study characteristics under Pillar 1. Study #Data (Unit) Modality Localization Clinical Indication Centers Availability [93] 10,290 scans CT Lungs Pulmonary/Thoracic Diseases Multiple Mixed [72] 6,814 scans CT Multiple Healthy/No Pathology Multiple Public [6] 3,064 slices MRI (CE-T1) Brain Cancer/Tumours N/R N/R [59] 5,513 scans CT Multiple Cancer/Tumours Single Public [80] 1,053,791 slices X-ray Thorax Pulmonary/Thoracic Diseases Multiple Private [90] 64,584 slices rs-fMRI Brain Neurologic/Psychiatric Multiple Public [31] 4,409 scans rs-fMRI Brain Neurologic/Psychiatric Multiple Public [32] 1,365 scans sMRI + fMRI Brain Healthy/No Pathology Multiple Public [63] >>10M slices CT, MRI, X-ray, WSI, endoscopy, histology, dermoscopy, microscopy. Multiple Multiple Multiple Public [2] 50,000 slices US Tongue Speech-related Single Private [15] 17M slices CT, MRI, X-ray, US, OCT, Retina, Dermoscopy, microscopy Multiple Multiple Multiple Public [65] 2,042 scans CT Multiple Multiple Multiple Public [21] 19,687 scans MRI (T1) Brain Neurologic/Psychiatric Multiple Public [23] 57,621 scans mpMRI (T1, T2, FLAIR, CE-T1) Brain Healthy/No Pathology Single Private [82] 64,740 scans MRI (T1, T2, FLAIR, DWI, SWI) Brain Healthy/No Pathology Multiple Private [17] 1,481 scans rs-fMRI Brain Neurologic/Psychiatric Multiple Public [70] 75,861 scans mpMRI (T1, T2, FLAIR) Brain Neurologic/Psychiatric Single Private [64] 1,444 scans CT, MRI Multiple Multiple Multiple Public [25] 40,000 scans CT, MRI Multiple Cancer/Tumours Multiple Mixed [9] 932 subjects CT, MRI, surgical video Multiple Cancer/Tumours Multiple Public [47] 1.1M slices (5M masks) CT Multiple Multiple Multiple Public [48] 33,111 slices US Multiple Cancer/Tumours Multiple Public [98] 6,585 scans MRI (T1, CE-T1, T2, FLAIR, DWI) Brain Neurologic/Psychiatric Multiple Mixed [16] 82,800 scans MRI (T1, CE-T1, T2, FLAIR) Brain Neurologic/Psychiatric Multiple Public [87] 2,798 scans T1 MRI; FDG-PET Brain Neurologic/Psychiatric Multiple Public [100] N/A N/A N/A Other/Not Specified N/A N/A [10] 7,039 slices US Multiple Cancer/Tumours Multiple Mixed [69] 516 scans MRI (T1/T2) Brain Pediatric Multiple Private [28] 320 scans MRI (T1, T2, PD) Multiple MSK Conditions Single Private [13] N/A N/A N/A Other/Not Specified N/A N/A [29] 1.35M slices CT, MRI, US Multiple Multiple Multiple Public [43] 14,012 scans CT Multiple Healthy/No Pathology Multiple Public [50] 700,000 slices X-ray Thorax Pulmonary/Thoracic Diseases Multiple Public [53] 2,374 scans CT Multiple Cancer/Tumours Multiple Public [37] 2,187,915 slices US Multiple Cancer/Tumours Multiple Mixed [67] N/A N/A N/A Other/Not Specified N/A N/A [39] 290,675 slices US Thyroid Endocrine Disorders Single Private [49] 15,660 scans mpMRI (DCE, T2, DWI) Breast Cancer/Tumours Multiple Private [11] 57,621 scans mpMRI(T1, CE-T1, T2, FLAIR) Brain Cancer/Tumours Single Private [52] 1,123,310 slices (1,570,263 masks) CT, MRI, X-ray, dermoscopy, fundus, endoscopy, US, MG, OCT, pathology Multiple Cancer/Tumours Multiple Public [54] 98,588 scans Chest low dose CT Thorax Multiple Multiple Partial [45] 51,029 scans MRI Brain Cancer/Tumors Multiple Public [89] 6,629 scans fMRI Brain Neurologic/Psychiatric Multiple Public [97] 804,305 slices CT, MRI, X-ray, US, endoscopy, CTA, CBCT, fundus, dermoscopy Multiple Multiple Multiple Partial [86] 98,815 scans MRI, CT, PET Multiple Multiple Multiple Partial [71] 32,015 scans MRI Brain Neurologic/Psychiatric Multiple Partial [81] 10,718 scans fMRI Brain Neurologic/Psychiatric Multiple Public [24] 105,184 scans CT Lung Pulmonary/Thoracic Diseases Multiple Private [5] 7,908 scans MRI (T1) Brain Neurologic/Psychiatric Multiple Public [51] 704,363 slices Chest X-ray Thorax Pulmonary/Thoracic Diseases Multiple Public [95] 15,300 scans MRI Brain Multiple Multiple Public [20] 4,451 scans MRI, PET Brain Multiple Multiple Public [91] 987,733 slices Chest X-ray Thorax Pulmonary/Thoracic Diseases Multiple Public [88] 5,137 scans CE-CT Colorectum Cancer/Tumors Multiple Private [84] 22,000 scans MRI, CT, X-ray, fundus, microscopy Multiple Multiple Multiple Public [92] 520,000 slices Chest X-ray Thorax Pulmonary/Thoracic Diseases Multiple Partial [36] 1,015,754 slices US Multiple Multiple Multiple Partial [73] 9,995 scans CT Multiple Multiple Multiple Public [83] 160,000 scans CT Multiple Multiple Multiple Public [3] 3,700,000 slices (15.8M masks) CT, MRI, endoscopy, US, X-Ray, dermoscopy, ophthalmology, CBCT, fundus, OCT, MG Multiple Multiple Multiple Public [94] 1,350,000 slices CT, MRI, US Multiple Multiple Multiple Public [78] 21,729 scans (143518 masks) CT, MR, US Multiple Multiple Multiple Partial [99] 49k scans (82k masks) CT, MRI, X-Ray, OCT, US, microscopy, endoscopy, dermoscopy, colonoscopy Multiple Multiple Multiple Public [34] 36,000,000 slices Cardiac MRI Heart Other/Not Specified Multiple Partial [35] 16,820 slices US Multiple Multiple Multiple Public [55] 186,894 slices (279,765 masks) US Multiple Multiple Multiple Public [38] 1,003,465 slices US Multiple Multiple Multiple Partial Abbreviations: CBCT, cone beam computed tomography; CE, contrast enhanced; CT, computed tomography; CTA, computed tomography angiography; DCE, dynamic contrast-enhanced; DWI, diffusion weighted imaging; FDG, fluorodeoxyglucose; FLAIR, fluid-attenuated inversion recovery; fMRI, functional magnetic resonance imaging; M, million; MG, mammography; mpMRI, multiparametric magnetic resonance imaging; MRI, magnetic resonance imaging; N/A, not applicable; N/R, not reported; OCT, optical coherence tomography; PD, proton density; PET, positron emission tomography; rs-fMRI, resting-state functional magnetic resonance imaging; sMRI, structural magnetic resonance imaging; SWI, susceptibility weighted imaging; US, ultrasound; WSI, whole slide imaging. Figure 2: Overview of the characteristics of radiology VFMs. (A) Distribution of dataset scale according to reporting unit. (B) Dataset scale over time: median dataset size and interquartile range, with trend lines and confidence bands. (C) Distribution of studies included in the review. (D) Distribution of study characteristics across Pillar 1. Counts and percentages are displayed within each category block, except for those represented by a single study. (E) Characteristics of downstream evaluation under Pillar 3, including generalization settings (a), adaptation strategies (b), and downstream task categories (c). Category percentages below 3% are omitted for clarity. (F) Distribution of model architectures (a) and pretraining strategies (b). (G) Reporting of FUTURE-AI principles, including the number of principles reported per study (a) and the proportion of studies addressing each principle (b). 3.3 Pillar 2: Large-scale pretraining and architecture Transformer-based models were the most prevalent, appearing in 53 studies (79.1%) [80, 90, 32, 63, 2, 15, 65, 21, 23, 70, 64, 9, 47, 48, 16, 100, 10, 28, 13, 29, 50, 53, 37, 67, 39, 11, 52, 54, 45, 89, 97, 86, 71, 81, 24, 51, 95, 20, 91, 88, 84, 92, 36, 73, 83, 3, 94, 78, 99, 34, 35, 55, 38]. These models include Vision Transformers (ViT; e.g., ViT-B and ViT-L), Swin Transformers, and other transformer variants. Hybrid architectures integrating convolutional and transformer components were reported in 14 studies (20.9%), reflecting a strategy to jointly capture local, spatial features and long-range contextual dependencies [2, 15, 65, 21, 70, 48, 16, 10, 13, 67, 39, 24, 51, 83]. Purely convolutional architectures were used in eight studies (11.9%) [93, 6, 59, 82, 25, 69, 43, 5], including DenseNet, ResNet, and UNet-based architectures. Alternative paradigms were less common. State-space (Mamba-based [27]) models were reported in two studies [72, 87]. Graph-based architectures were limited to three studies [31, 17, 81]. Three studies adopted mixture-of-experts (MoE) architectures, using formulations such as task-specific expert routing [21], and modality-specific experts with hierarchical or soft gating to handle missing modalities [98, 49]. SSL was the most dominant pretraining strategy, employed in 48 studies (71.6%). Among these, masking–based SSL approaches were most common (19 studies, 40%) [72, 80, 15, 21, 70, 25, 29, 37, 39, 54, 45, 89, 81, 95, 88, 36, 73, 99, 38], followed by contrastive SSL methods (6 studies, 12.5%) [93, 59, 2, 71, 5, 83]. In addition, generative approaches (1 study) [24] and self-distillation (DINO-based [8]) approaches (3 studies) [86, 94, 34] were identified. Nine studies (18.8%) used multi-stage or composite SSL pipelines, combining mask-based and contrastive learning techniques with strategies such as meta-learning or self-distillation frameworks [90, 23, 17, 64, 16, 87, 11, 91, 92]. Supervised pretraining was reported in 14 studies (22.4%) [6, 63, 65, 82, 48, 98, 10, 52, 50, 49, 97, 51, 20, 78, 55], typically leveraging large, annotated datasets, multi-task objectives, or cyclical training across heterogeneous sources. Initialization from large pretrained models was also common, particularly SAM-based [42] approaches (15 studies, 22.4%) [64, 9, 47, 100, 28, 13, 29, 53, 97, 20, 84, 3, 78, 35, 55], and were most often adapted through supervised fine-tuning, frequently leveraging PEFT strategies, adapters, or partial backbone freezing. Several studies further incorporated external pretrained encoders, such as CLIP [62] or ImageNet-pretrained networks, as part of hybrid initialization schemes [32, 10, 49]. Model parameter counts were reported in 41 studies (61.2%), ranging from fewer than 10 million parameters to approximately 1.2 billion parameters, with the largest model being [83], and a median of 214 million. In studies reporting multiple model variants, the largest configuration was considered. Computational resources were reported in 54 studies (80.6%). Among these, 21 studies (31.3%) employed single-GPU training setups, whereas 33 (49.3%) used multi-GPU configurations, typically ranging from two to eight GPUs and reaching up to 32 and 64 GPUs in some cases, with NVIDIA A100 GPUs being the most commonly used. In contrast, 14 studies (20.9%) did not specify the computing resources employed. Regarding openness and reproducibility, code was available in 48 studies (71.6%), while pretrained weights were publicly available in 35 studies (52.2%). Architectural and pretraining characteristics of the included studies are summarized in Table 5 and represented in Figure 2 F. Table 5: Summary of study characteristics under Pillar 2. Study Model name Architecture Pretraining strategy Hardware Parameters Weights Code [93] DrasCLR 3D CNN MoCo Contrastive SSL 4×V100 N/R Yes Yes [72] MambaMIM CNN–Mamba hybrid MIM 1×A800 N/R Yes Yes [6] None U-Net Supervised Vertex AI N/R No No [59] None 3D ResNet50 SimCLR 2×RTX 8000 200M Yes Yes [80] None ViT-L MAE 8×A800 348M No Yes [90] BrainMass Transformer MAE + contrastive 64×V100 67M Yes Yes [31] HGFM Hypergraph SSL (link prediction) N/R N/R No No [32] FM-APP ViT + BiomedCLIP Masked regressor N/R 80–100M N/R Yes [63] UMedPT Swin Transformer + decoders Supervised multitask N/R N/R Yes Yes [2] TongueTransUNet UNet + ViT SimCLR Vertex/Colab 130–180M No No [15] Frepa ViT/Swin/ConvNeXt MAE-style SSL 2×A100 32–87M Yes Yes [65] None Swin-UNETR Supervised 4×A6000 62–65M Yes Yes [21] DenseFormerMoE DenseNet + ViT MAE 1×RTX 4090 22–86M No No [23] None ViT MAE + contrastive 8×A100 500–800M Yes Yes [82] None 3DDenseNet201 Supervised 2×RTX 2080 N/R Yes Yes [17] MeTSK ST-GCN Contrastive SSL + meta-learning N/R <<10M No No [70] SwinClassifier Swin-UNETR MAE 8×A100 N/R Yes Yes [64] ProtoSAM-3D SAM-Med3D MAE + distillation 1×A100 10–113M No No [25] ViNet 3D ResNet-18 SSL (image restoration) 1×V100 N/R No No [9] MA-SAM SAM + 3D adapters SAM initialization 8×A100 ∼ 97M Yes Yes [47] SAMCT SAM + U-Net SAM initialization HPC HUST 120M Yes Yes [48] PerceptGuide Swin-UNet Supervised multi-task 4×A4000 66M Yes Yes [98] MoME/MoME+ nnUNet MoE Supervised 1×A100 N/R Yes Yes [16] BrainSegFounder 3DSwinUNETR Dual-stage SSL (MAE, contrastive, rotation prediction) 64×A100 62–69M Yes Yes [87] ADFound Vim (Mamba) MAE + contrastive 1×RTX 4090 N/R No Yes [100] MASG-SAM SAM-ViT-B SAM init 1×RTX 4090 ∼ 98M No Yes [10] MOFO CSWin Transformer + CNN ImageNet + CLIP 4×RTX 3090 N/R Yes Yes [69] BME-X DU-Net No pretraining N/R N/R No No [28] SegmentAnyBone SAM-based SAM initialization 1×RTX A6000 N/R Yes Yes [13] cineCMR-SAM U-Net + SAM blocks SAM initialization DGX-A100 1099M No Yes [29] LeSAM SAM-ViT-B SAM inicialization + MAE 1×RTX 4090 N/R No No [43] MedLAM/MedLSAM CNN SSL 4×RTX 3090 Ti N/R Yes Yes [50] Ark+ Swin Transformer-L Supervised multi-dataset 4×A100 ∼ 200M Yes Yes [53] ONCOPILOT SAM-ViT-B SAM initialization 32×V100 N/R No No [37] USFM ViT-B MIM 1×A100 N/R Yes Yes [67] SAM-AutoMed SAM-Med2D + MobileNet SAM initialization 1×RTX 2080 ∼ 108–110M No No [39] Deblurring MIM ConViT-B MIM N/R ∼ 90M Yes Yes [49] MOME BEiT3 + MOME Adapter training 1×RTX 3090 276M Yes Yes [11] ViT AE 3DViT MAE + contrastive 8×A100 N/R Yes Yes [52] MedSAM SAM-ViT-B SAM initialization + supervised 20×A100 93.7M Yes Yes [54] TANGERINE 3DViT MAE 4×A6000 ∼ 312M Yes Yes [45] UMBIF ViT ImageNet initialization + MAE HPC N/R No No [89] BrainSN Transformer SSL reconstruction 1×RTX 4090D N/R No Yes [97] MedSegX SAM-based + MoE Supervised N/R 94/312/641M No No [86] 3DINO-ViT 3D DINOv2 DINOv2 4×A100-SXM4 307M Yes Yes [71] BrainIAC 3D ViT SimCLR N/R N/R No No [81] None Graph transformer SSL reconstruction N/R N/R No No [24] LCTfound Cross-attention UNet DDPM SSL 16×V100 200M Yes No [5] AnatCL 3DResNet Contrastive SSL 1×V100 + 1×A100 34M Yes Yes [51] Ark+ Swin Transformer-CNN Supervised cyclic 4×V100 + 4×A100 N/R Yes Yes [95] BDFM Swin Transformer MIM 1×V100 91.66/30.95M No Yes [20] SAM-Brain3D+HyDA 3D SAM-based SAM-Med3D initialization 1×A100 100.51M No Yes [91] CheXFound ViT-L MAE + distillation 8×A100 307M No Yes [88] CRCFound 3DViT MAE 4×A100 N/R Yes No [84] EICSeg Dual encoder DINOv2 + SAM inizialization 8×V100 38.8M No Yes [92] EVA-X Dual ViT Contrastive + MIM N/R 6/22/86M Yes Yes [36] UltraFedFM ViT Federated MAE N/R 86/307/632M No No [73] Hi-End-MAE ViT + Hierarchical Dense Decoder MIM N/R N/R No Yes [83] VoCo SwinUNETR Omni-supervised 8×H800 31M–1.2B Yes Yes [3] MedicoSAM SAM + ViT-B SAM initialization 8×A100 86M Yes Yes [94] Radio DINO ViT DINO/DINOv2 2×A100 5.8M–86M No Yes [78] SAM-Med3D 3DViT Supervised 2×A100 N/R Yes Yes [99] SegMIC ViT MIM 1×A100 114.6M No Yes [34] None ViT-S DINO 8×H100 21M No No [35] UltraSAM SAM-based SAM initialization N/R 86M No No [55] UltraSam SAM-ViT-B SAM initialization + supervised 4×H100; 1×A100 86M Yes Yes [38] URFM ViT MIM N/R 86M Yes Yes Abbreviations: 3D, three-dimensional; AI, artificial intelligence; B, base; CLIP, Contrastive Language-Image Pre-training; CNN, convolutional neural network; CSWin, cross-shaped window; DDPM, denoising diffusion probabilistic models; DINO, distillation with no labels; DUNet, double UNet; L, large; M, million; MAE, masked autoencoder; MIM, masked image modeling; MoCo, momentum contrast; MoE, mixture of experts; MOME, mixture of modality experts; N/R, not reported; S, small; SAM, Segment Anything Model; SimCLR, Simple Framework for Contrastive Learning of Visual Representations; SSL, self-supervised learning; ST-GCN, spatial temporal graph convolutional network; ViT, Vision Transformer; Vim, vision Mamba. 3.4 Pillar 3: Scalability to multiple applications Substantial variability was observed in scalability across tasks, datasets, and clinical settings (Table 6). Most studies (58, 86.5%) evaluated multiple downstream tasks, ranging from 1 to 146 tasks (median: 6). Notably, 19 studies (28.4%) evaluated 10 or more tasks [90, 32, 63, 15, 65, 47, 48, 98, 29, 37, 52, 54, 97, 5, 92, 36, 83, 78, 99]. Common downstream tasks included classification, segmentation, detection, regression, survival analysis, enhancement, localization, and progression modeling. Segmentation (67.1%) and classification (61.1%) were the most frequently investigated tasks. Fine-tuning (full or partial) was the dominant adaptation strategy (45 studies, 67.1%), often complemented by PEFT methods (10 studies, 14.9%) [65, 9, 47, 28, 13, 29, 20, 84, 91, 35]. In contrast, frozen-encoder evaluations using linear probing (LP) were explicitly performed in 12 studies (17.9%) [93, 59, 90, 15, 50, 67, 86, 71, 24, 5, 51, 91]. Similarly, ZS evaluation without task-specific fine-tuning was reported in 12 studies (17.9%) [32, 65, 17, 64, 48, 10, 43, 50, 52, 89, 97, 99]. FS learning scenarios were evaluated in 17 studies (25.3%), commonly involving 1–5 labeled samples per task [90, 31, 63, 65, 64, 48, 100, 43, 50, 71, 24, 84, 92, 73, 99, 34, 35]. Generalization of VFMs was evaluated in 97% of studies. Disease shift was the most frequently evaluated setting (54 studies, 80.5%) [93, 72, 59, 80, 90, 31, 63, 15, 21, 23, 17, 70, 9, 47, 48, 98, 16, 87, 100, 10, 69, 29, 50, 53, 37, 67, 39, 49, 11, 52, 54, 45, 89, 97, 86, 71, 81, 24, 5, 51, 95, 20, 91, 88, 84, 92, 36, 94, 78, 99, 34, 35, 55, 38]. Task shift generalization was assessed in 32 studies (47.7%) [93, 59, 80, 63, 15, 21, 48, 69, 43, 50, 53, 37, 67, 39, 49, 89, 86, 71, 81, 24, 5, 51, 95, 20, 91, 88, 92, 36, 83, 94, 34, 55]. Anatomical shift generalization was reported in 30 studies (44.7%) [72, 59, 63, 15, 65, 64, 9, 47, 48, 100, 10, 28, 29, 43, 53, 37, 52, 97, 86, 91, 84, 36, 73, 3, 94, 78, 99, 35, 55, 38]. Modality shift generalization was evaluated in 19 studies (28.4%) [72, 32, 82, 64, 25, 9, 98, 16, 87, 100, 29, 52, 97, 84, 73, 3, 78, 99, 35]. In addition, cross-center generalization was explicitly evaluated in 23 studies (34.3%) [59, 80, 90, 63, 82, 25, 9, 10, 13, 50, 53, 37, 49, 54, 45, 97, 24, 88, 92, 36, 78, 34, 35], while cross-scanner robustness was reported in four studies [37, 49, 11, 35]. Only two studies did not evaluate generalization [2, 6]. Performance evaluation against established pretrained and VFMs was a central component of the included studies, reported in 49 (73.1%) of cases. Among these, 93.8% demonstrated superior performance, either consistently across benchmarks or in more than 90% of the evaluated downstream tasks, while an additional 6% reported partial improvements, demonstrating superior performance in at least one evaluation setting. Comparisons with fully supervised, task-specific models were even more prevalent, reported in 57 studies (85.1%). Among these, 84% outperformed task-specific models, whereas 15.8% demonstrated partial improvements. Downstream evaluation characteristics of the included studies are summarized in Table 6 and represented in Figure 2 E. Table 6: Summary of study characteristics under Pillar 3. Study #Tasks Task types Fine-tuning strategy Generalization ZS FS PT/FM ↑ Task-spec.↑ [93] 9 cls, reg, surv, seg LP, FT Disease, Task No No Yes Yes [72] 8 seg FT Modality, Anatomical No No Yes Yes [6] 1 seg N/R Task N/R N/R N/R N/R [59] 3 cls, surv LP, FT Disease No No Yes Yes [80] 3 cls, rp FT with multimodal decoder Task No No Partially Partially [90] 15 cls LP, FS Disease Yes Yes Yes Yes [31] 4 cls FS, FT Disease No Yes N/R Yes [32] 92 reg ZS Modality, Task Yes No N/R Yes [63] 20 seg, cls, det FT, FS Domain, Modality No Yes Yes Yes [2] 1 seg Semi-supervised In-distribution No No N/R Yes [15] 32 seg, cls, det LP Modality, Task Yes N/R Yes Yes [65] 13 seg ZS, FS, FT, PEFT Anatomical Yes Yes Yes Yes [21] 3 cls, reg Multi-Task FT Task No No N/R Yes [23] 1 cls FT In-distribution No No N/R Yes [82] 1 reg Transfer learning Domain No No N/R N/R [17] 1 cls ZS Disease Yes No Yes Yes [70] 1 cls FT Dataset shift No No N/R Yes [64] 5 seg ZS, distillation + FT, FS Modality, Anatomical, Task Yes Yes Yes Yes [25] 1 cls FT + experiential guidance Center, Domain No No N/R Yes [9] 5 seg PEFT, FT Modality, Task No No Yes Yes [47] 118 seg PEFT Anatomical No No Yes Yes [48] 20 seg, cls ZS, FT, FS Anatomical, Task Yes Yes Yes Yes [98] 17 seg FT Modality, Disease No No Yes Yes [16] 2 seg FT Disease, Task No No N/R Yes [87] 6 cls, prog FT Modality, Disease No No Yes Yes [100] 6 seg FS Task No Yes Yes Yes [10] 3 seg FT, ZS Anatomical Yes No Yes Yes [69] 4 enh, seg, reg, cls N/A Domain, Scanner No No N/R Yes [28] 1 seg PEFT Anatomical, Modality Yes No N/R Yes [13] 3 seg PEFT Center, Scanner No No N/R Yes [29] 12 seg PEFT Modality, Disease No No Yes Yes [43] 3 loc, seg ZS, FS Anatomical Yes Yes Partially Yes [50] 8 cls FT, LP, FS, ZS Disease, Center Yes Yes Yes Yes [53] 1 seg FT Center No No Yes Yes [37] 11 seg, cls, enh FT Anatomical, Task No No Yes Yes [67] 6 seg, cls LP Modality, Task No No Yes Yes [39] 4 cls, seg FT Domain, Scanner No No Yes Yes [49] 3 cls FT Domain, Center No No N/R Yes [11] 3 det, cls FT Task No No N/R Yes [52] 146 seg ZS Modality, Anatomical, Task Yes No Yes Yes [54] 14 cls FT Disease, Center, Domain, In-distribution No No Yes Partially [45] 2 cls FT Disease, Center, Domain No No Yes Yes [89] 5 cls, reg, other ZS, FT Task, Domain Yes No Yes Partially [97] 111 seg ZS, FT Disease, Task, Modality, Anatomical, Domain, In-distribution, Center Yes No Yes Yes [86] 6 cls, reg, seg LP, FT Modality, Anatomical, Domain No No Yes N/R [71] 7 cls, reg, seg, sup FT, LP, FS Disease, Task, Domain No Yes Yes N/R [81] 4 cls, reg, other FT Task, Domain No No N/R Yes [24] 8 cls, seg, prog, other FT, LP, FS Disease, Task, Center, Domain No Yes Yes Yes [5] 22 cls, reg LP Disease, Task, Domain No No Yes Partially [51] 3 cls, seg, other FT, LP Task, Modality, Domain No No Yes Yes [95] 3 cls, seg FT Disease, Task, Domain No No N/R Yes [20] 3 cls, seg, prog PEFT Disease, Modality, Domain No No Yes Yes [91] 9 cls, seg, prog LP, decoder tuning Disease, Task, Domain, In-distribution No No Yes N/R [88] 8 cls, sup, prog PEFT Task, Center, Domain No No N/R Yes [84] 9 seg PEFT, FS Anatomical, Modality, Domain No Yes Yes Partially [92] 11 cls, seg, other FT, FS Disease, Task, Center, Domain No Yes Yes Yes [36] 13 cls, seg FT Disease, Anatomical, Center, Modality, Domain No No Yes Yes [73] 9 seg FT, FS Modality, Anatomical, Domain No Yes Yes N/R [83] 51 cls, seg, other FT Task, Anatomical, Domain No No Yes N/R [3] 4 seg FT Task, Modality, Domain No No Yes Partially [94] 7 cls, seg FT Modality, Task, Domain No No Yes Partially [78] 16 seg FT Disease, Anatomical, Modality, Domain No No Yes Partially [99] 56 seg ZS, FS Anatomical, Modality, Domain Yes Yes Yes N/R [34] 9 cls, seg, other FT, FS Task, Modality, Center, Scanner, Domain No Yes Partially Partially [35] 10 seg PEFT, FS Anatomical, Center, Scanner, Modality, Domain Yes Yes Yes Yes [55] 3 cls, seg, other FT Anatomical, Modality, Domain No No Yes N/R [38] 10 cls FT Anatomical, Domain No No Yes N/R Abbreviations: cls, classification; det, detection; enh, image enhancement; FS, few-shot; FT, fine-tuning; loc, localization; LP, linear probing; N/A, not applicable; N/R, not reported; PEFT, parameter-efficient fine-tuning; prog, prognosis; PT/FM ↑ , superior performance over pretrained or foundation models; reg, image registration; rp, report generation; surv, survival; Task-spec.↑ , superior performance over task-specific models; ZS, zero-shot. 3.5 FUTURE-AI principles alignment The alignment of the included studies with the FUTURE-AI principles is summarized in Figure 2 G. To enhance transparency and reproducibility, a detailed breakdown is provided in Table 8 (Supplementary Material, Section 7.3), where each reported principle is accompanied by a brief justification outlining the specific criteria and evidence supporting its classification. Across the included literature, adherence to the FUTURE-AI principles was heterogeneous and inconsistently reported. Only 20 studies (29.9%) comprehensively addressed all principles [90, 15, 17, 25, 16, 10, 43, 49, 52, 54, 45, 89, 71, 5, 51, 92, 36, 94, 34, 38]. Among individual principles, robustness was the most frequently reported (63 studies, 94%), reflecting the strong emphasis on performance stability and generalization across datasets and tasks. Universality (60 studies, 89.6%) and usability (58 studies, 86.6%) were also commonly addressed, consistent with the focus on scalability and practical applicability of VFMs across diverse clinical settings. Traceability was reported in 54 studies (80.5%), typically in the form of methodological transparency or documentation of model development and evaluation. In contrast, explainability and fairness were less consistently evaluated, reported in 48 (71.6%) and 34 (50.7%) studies, respectively. 4 Discussion This scoping review aims to provide a comprehensive overview of the current landscape of VFMs in radiology. Although the topic has gained substantial visibility in recent years, the scientific evidence remains recent and rapidly evolving. The first studies meeting the inclusion criteria emerged only in late 2023, with publication activity increasing markedly from 2025 onward (Figure 2). This growth likely reflects the convergence of several technological and methodological advances, including the success of FMs in computer vision, the maturation of SSL approaches, increased access to large-scale imaging datasets, and progress in scalable computational infrastructures [12]. In total, 67 studies met the predefined eligibility criteria. While this number may appear modest relative to the growing interest in the field, it should be interpreted in the context of the deliberately focused scope of this review. Specifically, we restricted our analysis to imaging-only FM and excluded approaches incorporating language data. This decision was made to maximize methodological comparability across studies and evaluation settings while isolating imaging-specific contributions from the additional complexities associated with multimodal learning and textual integration [14]. As a result, this review provides a focused assessment of the state of VFM in radiology, enabling the identification of prevailing trends, methodological limitations, and priorities for future research. Among the most important observations arising from this review is the absence of a consistently applied definition of VFMs in radiology. This lack of consensus is reflected in the substantial heterogeneity observed across studies, encompassing dataset scale and composition, pretraining paradigms, architectural design choices, fine-tuning approaches, and downstream evaluation tasks. To provide a more structured analysis, we organized the literature according to three core pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Regarding the first pillar, current radiology VFMs remain strongly shaped by the scale of the datasets used during pretraining. Although slice-based datasets occasionally reach the million-sample scale, scan-based datasets remain predominantly small to medium in size. This discrepancy reflects intrinsic barriers in radiology, including the cost of data acquisition, storage, and annotation, as well as privacy constraints. Importantly, it raises questions regarding the extent to which current models can genuinely be considered “foundation models”, as their scale often falls short of what is typically observed in natural image or multimodal FM development [19]. Dataset composition across imaging modalities, anatomical regions, and pathologies also revealed important observations. In principle, datasets encompassing multiple imaging modalities should benefit from learning complementary information, enabling VFMs to develop more robust and transferable visual representations [57]. Despite these potential benefits, more than half of the reviewed studies relied on a single imaging modality. This under-representation was particularly evident for nuclear medicine modalities, such as PET, which appeared in only a small fraction of studies. One potential factor limiting the integration of heterogeneous datasets across imaging modalities is the substantial differences in spatial dimensionality. These differences often encourage the use of slice-wise or pseudo-3D processing strategies, even for inherently volumetric modalities such as CT and MRI, thereby limiting the exploitation of three-dimensional contextual information[58]. Consequently, explicit volumetric representation learning was supported by relatively few studies. Temporal modeling was even less frequently addressed, likely due to the limited availability of longitudinal (4D) imaging datasets, which constraints the ability of current VFMs to capture disease evolution over time. Pathology coverage was similarly imbalanced, with some studies using multi-pathology datasets to support broader representation learning, whereas many remained focused on single-disease settings, most commonly involving thoracoabdominal CT in oncology, brain MRI in neurological disorders, and chest X-ray imaging in pulmonary diseases. In contrast, several major radiological domains, such as cardiovascular, pediatric, musculoskeletal, obstetric, and inflammatory imaging, remained absent or sparsely represented. Although specialized datasets often align with well-defined clinical tasks, they may not fully exploit the potential of FMs to learn transferable medical representations. The coexistence of these two paradigms suggests that the field has not yet converged on a unified strategy for balancing task-specific performance with general-purpose modeling. Data heterogeneity, accessibility, and reproducibility also pose significant challenges for VFMs. While the predominance of multicenter studies suggests increasing awareness of the importance of data diversity for robust model development, important gaps remain. In particular, demographic diversity is often insufficiently reported, limiting the assessment of model fairness. At the same time, dataset accessibility remains a critical bottleneck. Although more than half of the studies rely exclusively on public datasets, a significant proportion incorporate proprietary data or depend entirely on private collections. While private datasets may enhance scale and diversity, their use limits reproducibility and standardized benchmarking, underscoring the need for larger, standardized, and openly accessible radiology datasets [40, 57]. Regarding the second pillar, current radiology VFMs remain characterized by substantial diversity in architectural design, pretraining strategies, and computational scale, reflecting an evolving yet still fragmented methodological landscape. Transformer-based architectures clearly predominated across studies, likely due to their ability to capture long-range spatial dependencies, integrate multi-scale contextual information, and scale effectively with increasing data availability [4]. However, the continued presence (>30%) of purely convolutional and hybrid convolution–transformer architectures suggests that fully transformer-based solutions may not yet adequately capture the inductive biases required for radiological imaging data. In particular, convolutional components remain important for encoding fine-grained local structures, which are often critical for radiological interpretation [41]. The coexistence of these paradigms indicates that the field has not yet converged on a canonical architectural design for VFMs in radiology. At the same time, the emergence of alternative approaches—including state-space models (e.g., Mamba [27]), graph-based methods, and mixture-of-experts frameworks—highlights ongoing exploration beyond established transformer-based designs. Although these paradigms remain less prevalent, they point toward promising directions for addressing challenges such as spatio-temporal modeling, structured reasoning, modality-aware processing, and computational efficiency. Pretraining strategies similarly revealed a strong predominance of SSL, likely reflecting both the scarcity of large-scale annotated radiological datasets and the need to exploit large quantities of unlabeled medical imaging data for representation learning [19]. Within this context, reconstruction-based approaches, particularly masked image modeling, are the most commonly adopted SSL strategy, followed by contrastive learning and composite multi-stage pipelines, typically combining reconstruction and contrastive objectives [79]. Emerging techniques such as self-distillation are also observed, likely reflecting their recent success in natural-image domains thanks to their capability to integrate both global- and pixel-level features [8]. Nevertheless, a non-negligible proportion of studies continues to rely on supervised pretraining, including multi-task and segmentation-oriented setups, particularly in settings where curated annotations were available. In parallel, the increasing use of initialization from large pretrained models (e.g., SAM-based approaches [42]) highlights a growing dependence on external sources of knowledge. While these strategies can improve efficiency and performance, they also raise questions regarding domain alignment and the extent to which general-domain representations adequately capture the specific characteristics of radiological data. Model scale, computational requirements, and reproducibility practices also revealed valuable insights in the current radiology VFMs landscape. Reported parameter counts span several orders of magnitude, ranging from relatively compact architectures to models approaching the billion-parameter scale, although the median size (∼ 200M) remains substantially lower than that of FMs in natural vision or multimodal AI [61]. This suggests that, despite adopting the terminology of FMs, current approaches in radiology are still comparatively limited in scale. This trend is further reflected in the reported computing infrastructure, as a limited number of studies reported access to large multi-GPU infrastructures, while a considerable proportion relied on single-GPU setups. These constraints likely stem from the restricted availability of large-scale training datasets, which in turn reduces the necessity of large-scale distributed training infrastructures. Reproducibility and openness practices presented a mixed picture. Although a majority of studies provided access to source code, substantially fewer released pretrained weights, thereby limiting reproducibility and downstream reuse. In addition, incomplete reporting of model configurations and computational resources in some works further hinders transparent comparison across methods. As models continue to grow in complexity, standardized reporting practices and broader adoption of open science principles will be essential to ensure reproducibility and facilitate progress in the field. Regarding the third pillar, current radiology VFMs were primarily assessed in terms of the breadth of their downstream evaluation, the transferability of their learned representations, and their advantages over task-specific approaches. Most studies (86.5%) evaluated their models across multiple downstream tasks, reflecting an emphasis on assessing the transferability and general-purpose representation capabilities. However, the wide range in the number of tasks per study (1–146) and the relatively modest median (6) highlight a substantial heterogeneity in how scalability to multiple applications is demonstrated. Additionally, only a minority of studies assessed performance across more than ten distinct tasks, suggesting a relatively narrow experimental evaluation and limiting the strength of applicability conclusions. The distribution of downstream tasks further reinforces this observation. Segmentation and classification remained the predominant evaluation settings, accounting for more than half of the reported downstream tasks. This comparatively limited exploration of more complex or longitudinal tasks—such as survival analysis, progression modeling, and image enhancement—suggests an incomplete assessment of VFMs’ capabilities to capture temporal evolution phenomena across diverse radiological scenarios. Broader real-world adoption will therefore require stronger evidence in these clinically meaningful settings. Beyond task diversity, radiological data heterogeneity underscores the importance of explicitly evaluating representation transferability across distribution shifts. Generalization analysis was performed in nearly all studies, most commonly through evaluations under disease shift conditions (80.5%), indicating a strong focus on robustness across varying disease distributions. In contrast, modality, and anatomical shifts were less frequently assessed, potentially reflecting the predominance of domain-specific FMs trained on restricted modalities and anatomical regions. Similarly, although multicenter datasets were commonly used during pretraining, explicit cross-center and cross-scanner evaluations remained comparatively limited, despite their recognized importance for assessing robustness under real-world clinical variability. Overall, VFMs demonstrate encouraging robustness within constrained settings, but their generalizability across heterogeneous clinical environments remains insufficiently characterized [66]. From an adaptation perspective, full or partial fine-tuning remained the dominant strategy (67.1%), often complemented by parameter-efficient fine-tuning (PEFT) approaches. This pattern indicates that most VFMs still rely heavily on task-specific parameter updates to achieve competitive performance. Linear probing, zero-shot, and few-shot evaluations were comparatively less explored, suggesting that current pretrained representations are not yet sufficiently generalizable to support robust out-of-the-box transfer across heterogeneous radiological settings, where labeled data may be scarce or unavailable. Even so, broader evaluation of these approaches may be valuable for reducing the dependence on annotated downstream datasets for fine-tuning, while also lowering adaptation costs, computational requirements, and time-to-deployment. Performance comparisons generally suggest strong empirical results, with most studies reporting improvements over pretrained or competing FMs (93.8%) and task-specific supervised models (84%). However, these results should be interpreted cautiously. First, the absence of standardized benchmarking protocols and the heterogeneity of datasets, evaluation metrics, and experimental designs limit the comparability of results across studies. Second, potential biases related to dataset selection, overlap between pretraining and evaluation data, and insufficient external validation may contribute to optimistic estimates of performance. The assessment of radiology VFMs using the FUTURE-AI principles [44] revealed that fewer than one-third of studies provided comprehensive coverage across all assessed dimensions. This pattern indicates that, despite increasing awareness of responsible AI frameworks, their systematic integration into model development and evaluation remains limited. A clear trend emerges in the prioritization of technically driven principles. Robustness was the most consistently addressed dimension, largely reflecting the emphasis on generalization across datasets and downstream tasks. Similarly, the high prevalence of universality and usability aligns with the overarching goal of VFMs to function as general-purpose models adaptable to multiple applications. Traceability was also frequently reported, typically through documentation of model architectures, training procedures, and evaluation protocols. However, the depth and granularity of such reporting remain variable, and often emphasizing methodological transparency rather than full reproducibility through access to pretrained weights, datasets, or complete training pipelines. In contrast, explainability and fairness exhibit comparatively lower and more inconsistent coverage, revealing an important imbalance in the current research landscape. Explainability, when addressed, was often limited to post hoc visualization techniques (e.g., activation maps or saliency methods), which may provide constrained insight into the clinically meaningful factors underlying the model predictions. The reduced assessment of fairness is particularly notable, as only half of the studies explicitly considered this dimension. Fairness assessments were often restricted to basic subgroup analyses or remain implicitly addressed through dataset diversity, rather than through systematic bias quantification or mitigation strategies. Taken together, these findings suggest a certain misalignment between the current emphasis of VFM research on representation learning and performance optimization and the broader requirements for trustworthy clinical deployment. Addressing these gaps will require not only methodological advances but also standardized evaluation frameworks that integrate technical performance with ethical and clinical considerations. In this context, the FUTURE-AI framework provides a valuable reference for guiding the next phase of radiology VFM development. 5 Limitations and further work This review has several limitations. First, our restriction to imaging-only models excluded multimodal architectures incorporating language encoders or clinical text. Although this choice preserved methodological comparability within a controlled scope, it limited the breadth of the synthesis and should be addressed in future work. Second, although we employed a structured and intentionally broad search strategy, relevant studies may have been missed because VFM terminology remains recent and inconsistently used in radiology. Coverage may also have been limited by the exclusion of additional sources such as conference proceedings and preprint repositories, and periodic updates will be needed given the rapid methodological evolution of the field. Third, the absence of universally accepted criteria for defining radiology VFMs introduced interpretative judgment into study selection. Although our conceptual framework was designed to establish structured inclusion boundaries, classification decisions inevitably involved some subjectivity, particularly in borderline cases. Finally, as this work was conducted as a scoping review, we did not perform formal methodological quality assessment or quantitative synthesis of reported performance metrics. Future systematic reviews or meta-analyses may enable more rigorous evaluation of study quality, benchmarking practices, reproducibility standards, and comparative effectiveness across VFM approaches. 6 Conclusion This scoping review highlights the rapid emergence of VFMs in radiology while underscoring the early-stage and fragmented nature of the field. Despite encouraging progress, current radiological VFMs remain constrained by limited dataset scale, modality imbalance, and substantial methodological heterogeneity. Across the three core pillars examined—data, methodology, and evaluation—a consistent theme is the absence of standardized practices. Pretraining datasets remain relatively small and insufficiently diverse, architectural and training strategies are highly variable, and downstream evaluation practices lack uniform benchmarks. Although studies report strong empirical performance, the field remains predominantly focused on technical optimization, with comparatively limited attention to trustworthiness dimensions such as fairness, explainability, and reproducibility. Strengthening alignment with responsible AI frameworks, including FUTURE-AI, will be critical for enabling trustworthy clinical translation. Addressing these challenges will be essential to realize the full potential of VFMs as scalable, general-purpose tools for robust, equitable, and clinically meaningful medical image analysis. Acknowledgments The authors declared no financial support was received for this work. Author contributions A.V.-R. and X.R.-P. contributed equally to this work. A.V.-R., X.R.-P., and A.J.-P. conceived the study. A.V.-R., X.R.-P., and A.F.-M. designed the methodology. A.V.-R. and X.R.-P. developed and executed the search strategy, performed screening, data extraction, and data analysis, and drafted the original manuscript. A.F.-M., I.I.R., Á.A.-B., and A.J.-P. critically reviewed and edited the manuscript. A.J.-P. supervised the work. All authors read and approved the final manuscript. Competing interests A.V.-R., X.R.-P., A.F.-M., I.I.R., and A.J.-P are employed by Quibim SL, a medical imaging technology company. Á.A.-B. is the CEO of Quibim SL and has stock ownership in the company. References [1] T. Akinci D’Antonoli, C. Bluethgen, R. Cuocolo, M. E. Klontzas, A. Ponsiglione, and B. Kocak (2026) Foundation models for radiology: Fundamentals, Applications, Opportunities, Challenges, Risks, and Prospects. Diagn. and Interv. Radiol. 32 (3), p. 259–272. External Links: Document Cited by: §1. [2] K. Al-Hammuri, F. Gebali, and A. Kanan (2025) TongueTransUNet: toward effective tongue contour segmentation using well-managed dataset. Med. & Bio. Eng. & Comput., p. 1–15. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [3] A. Archit, L. Freckmann, and C. Pape (2025) MedicoSAM: robust improvement of SAM for medical imaging. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6. [4] R. Azad, A. Kazerouni, M. Heidari, E. K. Aghdam, A. Molaei, Y. Jia, A. Jose, R. Roy, and D. Merhof (2024) Advances in medical image analysis with vision transformers: a comprehensive review. Med. Imag. Anal. 91, p. 103000. External Links: Document Cited by: §4. [5] C. A. Barbano, M. Brunello, B. Dufumier, M. Grangetto, A. D. N. Initiative, et al. (2025) Anatomical foundation models for brain MRIs. Pattern Recognit. Lett.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [6] S. K. Bhatt, S. Srinivasan, and P. Prakash (2023) Brain tumor segmentation pipeline model using U-Net based foundation model. Data and Metadata 2, p. 197–197. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [7] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. Adv. in Neural Inf. Preprocess Syst. 33, p. 1877–1901. External Links: Document Cited by: §1. [8] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin (2021) Emerging properties in self-supervised vision transformers. In 2021 IEEE/CVF ICCV, p. 9650–9660. External Links: Document Cited by: §3.3, §4. [9] C. Chen, J. Miao, D. Wu, A. Zhong, Z. Yan, S. Kim, J. Hu, Z. Liu, L. Sun, X. Li, et al. (2024) MA-SAM: modality-agnostic SAM adaptation for 3D medical image segmentation. Med. Imag. Anal. 98, p. 103310. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [10] H. Chen, Y. Cai, C. Wang, L. Chen, B. Zhang, H. Han, Y. Guo, H. Ding, and Q. Zhang (2024) Multi-organ foundation model for universal ultrasound image segmentation with task prompt and anatomical prior. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [11] M. Chen, M. Zhang, L. Yin, L. Ma, R. Ding, T. Zheng, Q. Yue, S. Lui, and H. Sun (2024) Medical image foundation models in assisting diagnosis of brain tumors: a pilot study. Eur. Radiol. 34 (10), p. 6667–6679. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [12] Q. Chen, A. Liu, J. Zhang, C. Yang, and Y. Zhang (2026) Foundation models in medical imaging: a review. EngMedicine 3 (2), p. 100123. External Links: ISSN 2950-4899, Document Cited by: §1, §4. [13] Z. Chen, S. Kim, H. Ren, S. Kim, S. Yoon, Q. Li, and X. Li (2025) Cine cardiac magnetic resonance segmentation using temporal-spatial adaptation of prompt-enabled segment-anything-model: a feasibility study. J. of Cardiovasc. Magn. Reson., p. 101909. External Links: Document Cited by: §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [14] Z. Cheng, A. Y. Ong, S. K. Wagner, D. A. Merle, L. Ju, H. Zhang, R. Chen, L. Pang, B. Li, T. He, et al. (2025) Understanding the robustness of vision-language models to medical image artefacts. NPJ Digit. Med. 8 (1), p. 727. External Links: Document Cited by: §4. [15] Y. Chu, Y. Zhang, Z. Han, C. Yang, L. Zhou, G. Luo, C. Huang, and X. Gao (2025) Improving representation of high-frequency components for medical visual foundation models. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [16] J. Cox, P. Liu, S. E. Stolte, Y. Yang, K. Liu, K. B. See, H. Ju, and R. Fang (2024) BrainSegFounder: towards 3D foundation models for neuroimage segmentation. Med. Imag. Anal. 97, p. 103301. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [17] W. Cui, H. Akrami, G. Zhao, A. A. Joshi, and R. M. Leahy (2023) Meta transfer of self-supervised knowledge: foundation model in action for post-traumatic epilepsy prediction. ArXiv. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [18] A. Davila, J. Colan, and Y. Hasegawa (2024) Comparison of fine-tuning strategies for transfer learning in medical image classification. Image and Vis. Comput. 146, p. 105012. External Links: Document Cited by: §1. [19] J. G. de Almeida, L. C. Alberich, G. Tsakou, K. Marias, M. Tsiknakis, K. Lekadir, L. Marti-Bonmati, and N. Papanikolaou (2025) Foundation models for radiology—the position of the AI for health imaging (AI4HI) network. Insights into Imaging 16 (1), p. 168. External Links: Document Cited by: §1, §4, §4. [20] Z. Deng, H. Wang, Z. Huang, L. Zhang, A. I. Aviles-Rivero, C. Liu, J. He, Z. Kourtzi, and C. Schönlieb (2025) Brain foundation models with hypergraph dynamic adapter for brain disease analysis. Pattern Recognit., p. 112595. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [21] R. Ding, H. Lu, and M. Liu (2025) DenseFormer-MoE: a dense transformer foundation model with mixture of experts for multi-task brain image analysis. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [22] A. Dosovitskiy (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv. External Links: Document Cited by: §1. [23] R. Gao, A. Peng, Y. Duan, M. Chen, T. Zheng, M. Zhang, L. Chen, and H. Sun (2025) Associations of postencephalitic epilepsy using multi-contrast whole brain MRI: a large self-supervised vision foundation model strategy. J. of Magn Reson. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [24] Z. Gao, G. Zhang, H. Liang, J. Liu, L. Ma, T. Wang, Y. Guo, Y. Chen, Z. Yan, X. Chen, et al. (2025) A lung CT vision foundation model facilitating disease diagnosis and medical imaging. Nat. Commun.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [25] Y. Gong, X. Zhang, Y. Xia, Y. Cheng, J. Bao, N. Zhang, R. Zhi, X. Sun, C. Wu, F. Wu, et al. (2025) A foundation model with weak experiential guidance in detecting muscle invasive bladder cancer on MRI. Cancer Lett. 611, p. 217438. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [26] J. Grill, F. Strub, F. Altché, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar, et al. (2020) Bootstrap your own latent-a new approach to self-supervised learning. Adv. in Neural Inf. Preprocess Syst. 33, p. 21271–21284. External Links: Document Cited by: item i.. [27] A. Gu and T. Dao (2023) Mamba: linear-time sequence modeling with selective state spaces. arXiv. External Links: Document Cited by: §3.3, §4. [28] H. Gu, R. Colglazier, H. Dong, J. Zhang, Y. Chen, Z. Yildiz, Y. Chen, L. Li, J. Yang, J. Willhite, et al. (2025) SegmentAnyBone: a universal model that segments any bone at any location on MRI. Med. Imag. Anal. 101, p. 103469. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [29] Y. Gu, Q. Wu, H. Tang, X. Mai, H. Shu, B. Li, and Y. Chen (2024) LeSAM: adapt segment anything model for medical lesion segmentation. IEEE J. of Biomed. and Health Inform. 28 (10), p. 6031–6041. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [30] J. Gui, T. Chen, J. Zhang, Q. Cao, Z. Sun, H. Luo, and D. Tao (2024) A survey on self-supervised learning: algorithms, applications, and future trends. IEEE Trans. on Pattern Anal. and Mach. Intell. 46 (12), p. 9052–9071. External Links: Document Cited by: §1. [31] X. Han, R. Xue, J. Feng, Y. Feng, S. Du, J. Shi, and Y. Gao (2025) Hypergraph foundation model for brain disease diagnosis. IEEE Trans. on Neural Netw. and Learn. Syst.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [32] Z. He, W. Li, Y. Liu, X. Liu, J. Han, T. Zhang, and Y. Yuan (2024) FM-app: foundation model for any phenotype prediction via fMRI to sMRI knowledge transfer. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [33] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. (2022) Lora: low-rank adaptation of large language models.. ICLR 1 (2), p. 3. External Links: Document Cited by: §1. [34] A. J. Jacob, I. Borgohain, T. Chitiboi, P. Sharma, D. Comaniciu, and D. Rueckert (2025) Towards a cardiovascular magnetic resonance foundation model for multi-task cardiac image analysis. J. of Cardiovasc. Magn. Reson. 27 (2), p. 101967. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [35] T. Jiang, Y. Li, W. Xing, R. Cao, M. Yu, Y. Zhu, Y. Chen, B. Li, and D. Ta (2025) UltraSAM: a foundational medical ultrasound segmentation model with limited training data. Expert Syst. with Appl., p. 130223. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [36] Y. Jiang, C. Feng, J. Ren, J. Wei, Z. Zhang, Y. Hu, Y. Liu, R. Sun, X. Tang, J. Du, et al. (2025) From pretraining to privacy: federated ultrasound foundation model with self-supervised learning. npj Digit. Med. 8 (1), p. 714. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [37] J. Jiao, J. Zhou, X. Li, M. Xia, Y. Huang, L. Huang, N. Wang, X. Zhang, S. Zhou, Y. Wang, et al. (2024) USFM: a universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis. Med. Imag. Anal. 96, p. 103202. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [38] Q. Kang, Q. Lao, J. Gao, W. Bao, Z. He, C. Du, Q. Lu, and K. Li (2025) URFM: a general ultrasound representation foundation model for advancing ultrasound image diagnosis. ISci. 28 (8). External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6. [39] Q. Kang, Q. Lao, J. Gao, J. Liu, H. Yi, B. Ma, X. Zhang, and K. Li (2024) Deblurring masked image modeling for ultrasound image analysis. Med. Imag. Anal. 97, p. 103256. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [40] C. J. Kelly, A. Karthikesalingam, M. Suleyman, G. Corrado, and D. King (2019) Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17 (1), p. 195. External Links: Document Cited by: §4. [41] J. W. Kim, A. U. Khan, and I. Banerjee (2025) Systematic review of hybrid vision transformer architectures for radiological image analysis. J. of Imaging Inf. in Med., p. 1–15. External Links: Document Cited by: §4. [42] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W. Lo, et al. (2023) Segment anything. In 2023 IEEE/CVF ICCV, p. 4015–4026. External Links: Document Cited by: §3.3, §4. [43] W. Lei, W. Xu, K. Li, X. Zhang, and S. Zhang (2025) MedLSAM: localize and segment anything model for 3D CT images. Med. Imag. Anal. 99, p. 103370. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [44] K. Lekadir, A. F. Frangi, A. R. Porras, B. Glocker, C. Cintas, C. P. Langlotz, E. Weicken, F. W. Asselbergs, F. Prior, G. S. Collins, et al. (2025) FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 388. External Links: Document Cited by: §1, §2.6, §4. [45] J. Li, R. Liu, Y. Xing, X. Gao, Q. Yin, and Q. Su (2025) A foundation model for brain tumor MRI analysis: WHO grading and subtype classification. Radiother. and Oncol., p. 111297. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6. [46] X. Li, J. Huang, G. Sun, and Z. Yang (2025) Self-supervised learning for MRI reconstruction: a review and new perspective. Magn Reson. Mater. in Phys., Bio. and Med., p. 1–22. External Links: Document Cited by: §1. [47] X. Lin, Y. Xiang, Z. Wang, K. Cheng, Z. Yan, and L. Yu (2024) SAMCT: segment any CT allowing labor-free task-indicator prompts. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [48] Z. Lin, S. Li, S. Wang, Z. Gao, Y. Sun, C. Lam, X. Hu, X. Yang, D. Ni, and T. Tan (2025) An orchestration learning framework for ultrasound imaging: prompt-guided hyper-perception and attention-matching downstream synchronization. Med. Imag. Anal., p. 103639. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [49] L. Luo, M. Wu, M. Li, Y. Xin, Q. Wang, V. Vardhanabhuti, W. C. Chu, Z. Li, J. Zhou, P. Rajpurkar, et al. (2025) A large model for non-invasive and personalized management of breast cancer from multiparametric MRI. Nat. Commun. 16 (1), p. 3647. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [50] D. Ma, J. Pang, M. B. Gotway, and J. Liang (2025) A fully open AI foundation model applied to chest radiography. Nat., p. 1–11. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [51] D. Ma, J. Pang, S. S. Velan, M. B. Gotway, and J. Liang (2026) Ark+: supervised training a single high-performance AI foundation model from many differently labeled datasets—no label consolidation required. Med. Imag. Anal. 108, p. 103828. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [52] J. Ma, Y. He, F. Li, L. Han, C. You, and B. Wang (2024) Segment anything in medical images. Nat. Commun. 15 (1), p. 654. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [53] L. Machado, L. Alberge, H. Philippe, E. Ferreres, J. Khlaut, J. Dupuis, K. Le Floch, D. Habip Gatenyo, P. Roux, J. Grégory, et al. (2025) A promptable CT foundation model for solid tumor evaluation. npj Precis. Oncol. 9 (1), p. 121. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [54] N. McConnell, P. Vasudev, D. Yamada, D. Cheng, M. Azimbagirad, J. McCabe, S. Aslani, A. H. Shahin, Y. Zhou, et al. (2026) A computationally frugal, open-source chest CT foundation model for thoracic disease detection in lung cancer screening programmes. Commun. Med.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [55] A. Meyer, A. Murali, F. Zarin, D. Mutter, and N. Padoy (2025) UltraSam: a foundation model for ultrasound using large open-access segmentation datasets. Int. J. of Comput. Assist. Radiol. and Surg., p. 1–10. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6. [56] L. Misera, G. Müller-Franzes, D. Truhn, and J. N. Kather (2024) Weakly supervised deep learning in radiology. Radiol. 312 (1), p. e232085. External Links: Document Cited by: §1. [57] M. Moor, O. Banerjee, Z. S. H. Abad, H. M. Krumholz, J. Leskovec, E. J. Topol, and P. Rajpurkar (2023) Foundation models for generalist medical artificial intelligence. Nat. 616 (7956), p. 259–265. External Links: Document Cited by: §4, §4. [58] S. Noh and B. Lee (2025) A narrative review of foundation models for medical image segmentation: zero-shot performance evaluation on diverse modalities. Quant. Imaging in Med. and Surg. 15 (6), p. 5825–5858. External Links: Document Cited by: §4. [59] S. Pai, D. Bontempi, I. Hadzic, V. Prudente, M. Sokač, T. L. Chaunzwa, S. Bernatz, A. Hosny, R. H. Mak, N. J. Birkbak, et al. (2024) Foundation model for cancer imaging biomarkers. Nat. Mach. Intell. 6 (3), p. 354–367. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [60] M. Paschali, Z. Chen, L. Blankemeier, M. Varma, A. Youssef, C. Bluethgen, C. Langlotz, S. Gatidis, and A. Chaudhari (2025) Foundation models in radiology: What, How, Why, and Why Not. Radiol. 314 (2), p. e240597. External Links: Document Cited by: §1. [61] R. Patil and V. Gudivada (2024) A review of current trends, techniques, and challenges in large language models (LLMs). App. Sci. 14 (5), p. 2074. External Links: Document Cited by: §4. [62] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In Int. Conf. on Mach. Learn., p. 8748–8763. External Links: Document Cited by: §3.3. [63] R. Schäfer, T. Nicke, H. Höfener, A. Lange, D. Merhof, F. Feuerhake, V. Schulz, J. Lotz, and F. Kiessling (2024) Overcoming data scarcity in biomedical imaging with a foundational multi-task model. Nat. Comput. Sci. 4 (7), p. 495–509. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [64] Y. Shen, D. Dreizin, B. Inigo, and M. Unberath (2025) ProtoSAM-3D: interactive semantic segmentation in volumetric medical imaging via a Segment Anything Model and mask-level prototypes. Comput. Med. Imaging and Graph. 121, p. 102501. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [65] J. Silva-Rodríguez, J. Dolz, and I. B. Ayed (2025) Towards foundation models and few-shot parameter-efficient fine-tuning for volumetric organ segmentation. Med. Imag. Anal. 103, p. 103596. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [66] M. U. Suleman, M. Mursaleen, U. Khalil, A. Saboor, M. Bilal, S. A. Khan, M. A. Subhani, M. A. Hussnain, S. N. Tabassum, and M. Tahir (2025) Assessing the generalizability of artificial intelligence in radiology: a systematic review of performance across different clinical settings. Annal. of Med. and Surg. 87 (12), p. 8803–8811. External Links: Document Cited by: §4. [67] J. Sun, K. Chen, Z. He, S. Ren, X. He, X. Liu, and C. Peng (2024) Medical image analysis using improved SAM-Med2D: segmentation and classification perspectives. BMC Med. Imaging 24 (1), p. 241. External Links: Document Cited by: §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [68] Y. Sun, X. Wen, Y. Zhang, L. Jin, C. Yang, Q. Zhang, M. Jiang, Z. Xu, W. Guo, J. Su, and X. Jiang (2025) Visual-Language foundation models in medical imaging: A systematic review and meta-analysis of diagnostic and analytical applications. Comput. Methods and Programs in Biomed. 268, p. 108870. External Links: Document Cited by: §1. [69] Y. Sun, L. Wang, G. Li, W. Lin, and L. Wang (2025) A foundation model for enhancing magnetic resonance images and downstream segmentation, registration and diagnostic tasks. Nat. Biomed. Eng. 9 (4), p. 521–538. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [70] X. Suo, M. Chen, L. Chen, C. Luo, G. J. Kemp, S. Lui, and H. Sun (2025) Automatic identification of parkinsonism using clinical multi-contrast brain MRI: a large self-supervised vision foundation model strategy. EBioMed. 116. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [71] D. Tak, B. A. Garomsa, A. Zapaishchykova, T. L. Chaunzwa, J. C. Climent Pardo, Z. Ye, J. Zielke, Y. Ravipati, S. Pai, S. Vajapeyam, et al. (2026) A generalizable foundation model for analysis of human brain MRI. Nat. Neurosci., p. 1–12. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [72] F. Tang, B. Nian, Y. Li, Z. Jiang, J. Yang, W. Liu, and S. K. Zhou (2025) MambaMIM: pre-training Mamba with state space token interpolation and its application to medical image segmentation. Med. Imag. Anal.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [73] F. Tang, Q. Yao, W. Ma, C. Wu, Z. Jiang, and S. K. Zhou (2025) Hi-End-MAE: hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation. Med. Imag. Anal., p. 103770. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [74] E. J. Topol (2019) High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25 (1), p. 44–56. External Links: Document Cited by: §1. [75] A. C. Tricco, E. Lillie, W. Zarin, K. K. O’Brien, H. Colquhoun, D. Levac, D. Moher, M. D. Peters, T. Horsley, L. Weeks, et al. (2018) PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation. Ann. of Intern. Med. 169 (7), p. 467–473. External Links: Document Cited by: §2.1. [76] B. VanBerlo, J. Hoey, and A. Wong (2024) A survey of the impact of self-supervised pretraining for diagnostic tasks in medical X-ray, CT, MRI, and ultrasound. BMC Med. Imaging 24 (1), p. 79. External Links: Document Cited by: §1. [77] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. Adv. in Neural Inf. Preprocess Syst. 30. External Links: Document Cited by: item i.. [78] H. Wang, S. Guo, J. Ye, Z. Deng, J. Cheng, T. Li, J. Chen, Y. Su, Z. Huang, Y. Shen, et al. (2025) SAM-Med3D: a vision foundation model for general-purpose segmentation on volumetric medical images. IEEE Trans. on Neural Netw. and Learn. Syst.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [79] W. Wang, E. Ahn, D. Feng, and J. Kim (2023) A review of predictive and contrastive self-supervised learning for medical images. Mach. Intell. Res. 20 (4), p. 483–513. External Links: Document Cited by: §1, §4. [80] X. Wang, Y. Li, W. Wu, J. Jin, Y. Rong, B. Jiang, C. Li, and J. Tang (2025) Pre-training on high-resolution X-ray images: an experimental study. Vis. Intell. 3 (1), p. 8. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [81] Y. Wang, V. D. Calhoun, G. D. Pearlson, P. Kochunov, T. G. van Erp, and Y. Du (2026) A graph transformer-based foundation model for brain functional connectivity network. Pattern Recognit. 169, p. 111988. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6. [82] D. A. Wood, M. Townend, E. Guilhem, S. Kafiabadi, A. Hammam, Y. Wei, A. Al Busaidi, A. Mazumder, P. Sasieni, G. J. Barker, et al. (2024) Optimising brain age estimation through transfer learning: a suite of pre-trained foundation models for improved performance and generalisability in a clinical setting. Hum. Brain Mapp. 45 (4), p. e26625. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [83] L. Wu, J. Zhuang, and H. Chen (2025) Large-scale 3D medical image pre-training with geometric context priors. IEEE Trans. on Pattern Anal. and Mach. Intell.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [84] S. Xie, L. Zhang, Z. Niu, F. Ye, Q. Zhong, D. Xie, Y. Chen, and L. Lin (2025) EICSeg: universal medical image segmentation via explicit in-context learning. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [85] Y. Xin, J. Yang, S. Luo, H. Zhou, J. Du, X. Liu, Y. Fan, Q. Li, and Y. Du (2024) Parameter-efficient fine-tuning for pre-trained vision models: a survey. arXiv. External Links: Document Cited by: §1. [86] T. Xu, S. Hosseini, C. Anderson, A. Rinaldi, R. G. Krishnan, A. L. Martel, and M. Goubran (2025) A generalizable 3D framework and model for self-supervised learning in medical imaging. npj Digit. Med. 8 (1), p. 639. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [87] G. Yang, K. Du, Z. Yang, Y. Du, E. Y. W. Cheung, Y. Zheng, M. Yang, Z. Kourtzi, C. Schonlieb, and S. Wang (2025) ADFound: a foundation model for diagnosis and prognosis of alzheimer’s disease. IEEE J. of Biomed. and Health Inform.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6, Table 8. [88] J. Yang, D. Cai, J. Liu, Z. Zhuang, Y. Zhao, F. Wang, C. Li, C. Hu, B. Gai, Y. Chen, et al. (2025) CRCFound: a colorectal cancer CT image foundation model based on self-supervised learning. Adv. Sci. 12 (41), p. e07339. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6. [89] L. Yang, L. Guo, Y. Yuan, J. Han, X. Hu, and T. Zhang (2025) A foundational fMRI model for representing continuous brain states. IEEE J. of Biomed. and Health Inform.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [90] Y. Yang, C. Ye, G. Su, Z. Zhang, Z. Chang, H. Chen, P. Chan, Y. Yu, and T. Ma (2024) BrainMass: advancing brain network analysis for diagnosis with large-scale self-supervised learning. IEEE Trans. on Med. Imaging 43 (11), p. 4004–4016. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6, Table 8. [91] Z. Yang, X. Xu, J. Zhang, G. Wang, M. K. Kalra, and P. Yan (2025) Chest X-ray foundation model with global and local representations integration. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6. [92] J. Yao, X. Wang, Y. Song, H. Zhao, J. Ma, Y. Chen, W. Liu, and B. Wang (2025) Eva-X: a foundation model for general chest X-ray analysis with self-supervised learning. npj Digit. Med. 8 (1), p. 678. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, §3.5, Table 4, Table 5, Table 6. [93] K. Yu, L. Sun, J. Chen, M. Reynolds, T. Chaudhary, and K. Batmanghelich (2024) DrasCLR: a self-supervised framework of learning disease-related and anatomy-specific representation for 3D lung CT images. Med. Imag. Anal. 92, p. 103062. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [94] L. Zedda, A. Loddo, and C. Di Ruberto (2025) Radio DINO: a foundation model for advanced radiomics and AI-driven medical imaging analysis. Comput. in Bio. and Med. 195, p. 110583. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.5, Table 4, Table 5, Table 6. [95] J. Zhang, C. Xu, H. Zhong, Y. Cao, and L. Zhao (2025) BDFM: foundation model for segmentation and classification tasks of brain diseases. IEEE Trans. on Biomed. Eng.. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, Table 4, Table 5, Table 6. [96] S. Zhang and D. Metaxas (2024) On the challenges and perspectives of foundation models for medical image analysis. Med. Imag. Anal. 91, p. 102996. External Links: Document Cited by: §1. [97] S. Zhang, Q. Zhang, S. Zhang, X. Liu, J. Yue, M. Lu, H. Xu, J. Yao, X. Wei, J. Cao, et al. (2025) A generalist foundation model and database for open-world medical image segmentation. Nat. Biomed. Eng., p. 1–16. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6. [98] X. Zhang, N. Ou, B. D. Basaran, M. Visentin, M. Qiao, R. Gu, P. M. Matthews, Y. Liu, C. Ye, and W. Bai (2025) A foundation model for lesion segmentation on brain MRI with mixture of modality experts. IEEE Trans. on Med. Imaging. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. [99] J. Zhao, F. Yang, X. Li, Z. Jiao, Q. Zhai, X. Li, D. Wu, H. Fu, and H. Cheng (2025) Segmic: a universal model for medical image segmentation through in-context learning. Pattern Recognit., p. 112179. External Links: Document Cited by: §3.2, §3.2, §3.2, §3.2, §3.2, §3.3, §3.3, §3.4, §3.4, §3.4, Table 4, Table 5, Table 6. [100] W. Zhou, G. Guan, Y. Gao, P. Si, M. Xu, and Q. Yan (2025) MASG-SAM: enhancing few-shot medical image segmentation with multi-scale attention and semantic guidance. IEEE J. of Biomed. and Health Inform.. External Links: Document Cited by: §3.3, §3.3, §3.4, §3.4, Table 4, Table 5, Table 6, Table 8. 7 Supplementary Material 7.1 Search terms Search terms used to include and exclude studies. Inclusion terms are searched in title, abstract, and keywords, whereas only the title is used to exclude publications. Publication year, source type, and language were considered as well. Table 7: Literature search strategy and filtering criteria. Category Search terms / filters Inclusion (TITLE-ABS-KEY) Foundation model “foundation* model*”; “large vision model*”; “large model*”; “vision foundation model*”; “VFM”; “large pre?trained model*”; “large ai”; “universal model*”; “general-purpose model*” AND Medical imaging “medical imag*”; “radiolog*”; “radiolog* imag*”; “computed tomography”; “CT”; “magnetic resonance”; “MRI”; “x-ray”; “nuclear medicine”; “PET”; “ultrasound”; “clinical imag*”; “oncolog*” Exclusion (TITLE) “natural language”; “large language model*”; “fundus”; “OCT”; “pathology”; “pathological”; “electronic health”; “EHR”; “genomic”; “text”; “notes”; “free-text”; “language”; “LLM”; “gene”; “genes”; “genetic*”; “genom*” Publication year >>2016 and << March 2026 Source type Journals only (LIMIT-TO(SRCTYPE, “j”)) Language English (LIMIT-TO(LANGUAGE, “English”)) 7.2 Search strategy Full electronic search strategy for Scopus is provided below. ( ( TITLE-ABS-KEY("foundation* model*") OR TITLE-ABS-KEY("large vision model*") OR TITLE-ABS-KEY("large model*") OR TITLE-ABS-KEY("vision foundation model*") OR TITLE-ABS-KEY("VFM") OR TITLE-ABS-KEY("large pre?trained model*") OR TITLE-ABS-KEY("large ai") OR TITLE-ABS-KEY("universal model*") OR TITLE-ABS-KEY("general-purpose model*") ) AND ( TITLE-ABS-KEY("medical imag*") OR TITLE-ABS-KEY("radiolog*") OR TITLE-ABS-KEY("radiolog* imag*") OR TITLE-ABS-KEY("computed tomography") OR TITLE-ABS-KEY("CT") OR TITLE-ABS-KEY("magnetic resonance") OR TITLE-ABS-KEY("MRI") OR TITLE-ABS-KEY("x-ray") OR TITLE-ABS-KEY("nuclear medicine") OR TITLE-ABS-KEY("PET") OR TITLE-ABS-KEY("ultrasound") OR TITLE-ABS-KEY("clinical imag*") OR TITLE-ABS-KEY("oncolog*") ))AND NOT( TITLE("natural language") OR TITLE("large language model*") OR TITLE("fundus") OR TITLE("OCT") OR TITLE("pathology") OR TITLE("pathological") OR TITLE("electronic health") OR TITLE("EHR") OR TITLE("genomic") OR TITLE("text") OR TITLE("notes") OR TITLE("free-text") OR TITLE("language") OR TITLE("LLM") OR TITLE("gene") OR TITLE("genes") OR TITLE("genetic*") OR TITLE("genom*"))AND PUBYEAR > 2016 AND PUBYEAR < Msarch 2026AND (LIMIT-TO(SRCTYPE, "j"))AND (LIMIT-TO(LANGUAGE, "English")) 7.3 FUTURE-AI Principles Reporting Table 8: FUTURE-AI principles reporting. Study Fairness Universality Traceability Usability Robustness Explainability [93] Not reported Not reported Code and model weights publicly available Not reported Evaluated on two independent lung CT datasets (COPDGene, MosMed) 3D emphysema mask representations; metric evolution during fine-tuning [72] Multi-dataset, multi-organ training; bias acknowledged but not analyzed Transfer across CT and MRI segmentation tasks Public data sources and code available Not reported Consistent gains across multiple downstream datasets Not reported [6] Not reported Not reported Not reported Not reported Not reported Not reported [59] Not reported Validated across multiple datasets and institutions Open data, code, and pretrained weights Containerized application provided High test–retest ICC; stability under perturbations Gradient-based saliency maps; gene-expression analyses [80] Dataset bias acknowledged across hospitals and demographics Evaluated on English and Chinese report generation tasks Public code available Not reported Multi-dataset evaluation with ablation studies Activation response maps highlighting thoracic regions [90] Multi-center, multi-disease cohort with site-aware splits Generalization across internal/external datasets; zero/few-shot Detailed data sources, preprocessing, and training documentation Pretrained weights and code released External validation; sensitivity analyses Attention maps and multivariate disease-region analyses [31] Multiple datasets used Not reported Not reported Not reported Not reported Visualization of key brain regions and relationships [32] Large public datasets with diverse subjects Zero-shot evaluation Open-source code available Not reported Zero-shot evaluation on two datasets Phenotype Active Maps linking regions to phenotypes [63] Not reported Multi-modality and multi-label training Not reported Frozen encoder enables efficient downstream training External multi-center, multi-scanner validation Not reported [2] Human-in-the-loop annotation quality control Not reported Not reported Semi-supervised workflow with human correction Handles noisy data and missing inputs Interpretable embeddings; QC metrics [15] Multi-source public and in-house datasets; external test sets Applicable across nine modalities (2D/3D) Ablation studies of architectural components Lightweight decoder; efficient integration Robust to perturbations; unseen modality generalization Reconstruction and frequency-domain visualizations [65] Not reported Cross-dataset transfer; adaptation to novel classes Open data, code, and pretrained weights Few-shot and parameter-efficient adaptation Domain-shift and low-shot evaluation Not reported [21] Broad multi-dataset, multi-disease training Multicentric external validation Architectural ablation studies Pretrained encoder reuse discussed External validation on held-out dataset GradCAMs; expert-frequency analyses [23] Multiple imaging manufacturers included Not reported Not reported Not reported Not reported Occlusion sensitivity mapping [82] Clinically representative, demographically diverse data Transfer across MRI sequences and orientations Detailed methods and open scripts Pretrained models and fine-tuning scripts Improved out-of-sample performance; variance reduction Not reported [17] Large healthy and clinical datasets; imbalance acknowledged Adaptability across ADHD, ASD, and PTE tasks Ablation studies Zero-shot inference via linear probing Validated across multiple datasets Feature-importance maps from classifier coefficients [70] Single-center, all-Asian cohort; scanner diversity described Routine clinical MRI protocols for real-world applicability Ethics approvals; code and pretrained models released Minimal preprocessing; whole-brain inputs Multi-scanner evaluation; independent test set Occlusion sensitivity maps; voxel-wise statistics [64] Not reported CT/MRI evaluation across organs and datasets Not reported Interactive and auto-prompting modes Consistent performance across modalities; zero-shot classes Prototype-based interpretable mask embeddings [25] Three-center evaluation Multi-center protocols with heterogeneous image quality Ethics approval; no shared code or weights Integrated into clinical AI system Ensemble modeling; external validation GradCAMs; center-wise performance analysis [9] Not reported Cross-modality evaluation (CT, MRI, surgical video) Not reported Promptable segmentation for semi-automatic workflows External generalization; component ablations Not reported [47] Not reported Generalization across 30 datasets and 118 objects Code, data, and checkpoints released Automated task-indicator prompts Strong performance on unseen datasets Not reported [48] Not reported External validation across organs and datasets Public code and pretrained models Fully automated prompts Balanced sampling; stable performance across datasets Not reported [98] Not reported Task- and modality-agnostic design; missing-modality handling Code repository available Single universal model simplifies deployment Generalization to unseen datasets; stability ablations Expert probability maps; latent space visualizations [16] Large multi-cohort population with standardized preprocessing Adaptable across tasks and MRI modalities Clear dual-phase pretraining pipeline Few-shot and modality-restricted training support Robust under few-shot and restricted-modality settings Ablations on modality restriction and few-shot learning [87] No subgroup or demographic bias analysis Multi-modal and non-image extension capability Public code (no pretrained weights) Code reusable for downstream tasks External validation on OASIS-3 Not reported [100] Not reported Limited to four datasets; no cross-modality evaluation No model weights released Requires bounding-box input Performance degrades on noisy/low-quality data Not reported [10] Multi-center, multi-vendor data with leakage-reducing splits Multi-organ OOD evaluation Open-source code and model Fully automatic segmentation Improved performance on challenging OOD datasets Inter-organ relationship and latent-space analyses [69] Not reported Tested on 19 datasets across lifespan and MRI contrasts No code or weights shared Not reported Robustness quantified across simulated artefacts Not reported [28] Organ distribution and demographic analysis reported External validation across locations and sequences Code and weights available Interactive prompt-based refinement Generalization to unseen locations and artefacts Not reported [13] Multi-center, multi-vendor, multi-pathology evaluation Generalization to unseen vendors and centers Public code; IRB approvals Text/box prompts; near-real-time inference Strong external performance with stratified analyses Not reported [29] Broad dataset; fairness not explicitly analyzed Extended external evaluation Not open-source Interactive segmentation External validation on 12 multi-modal datasets Not reported [43] Diverse public CT datasets; no demographic analysis General-purpose 3D localization and segmentation Open datasets, code, and models Minimal manual input via automated prompts Strong cross-dataset generalization Interpretable 3D bounding boxes and spatial maps [50] Sex-bias tolerance evaluated across splits Zero-shot transfer across tasks and datasets Data, code, and models public Fully open fine-tuning and adaptation Robust under long-tailed and domain-shift conditions Not reported [53] Not reported Multi-organ training; multicentric external validation Not open-source Radiologist-in-the-loop interactive prompts External real-world validation; editing mitigates bias Visual prompts and interactive editing [37] Organ-balanced sampling strategy analyzed Multi-organ, multi-center, multi-device evaluation Open-source code and weights Plug-and-play backbone integration Robust feature learning under low-quality US conditions Not reported [67] Not reported Cross-domain generalization across multiple modalities Code and models not public Automated prompt-free segmentation Robust across modality variation Not reported [39] Single-center, single-organ cohort Transfer across vendors and tasks Code and weights available Open-source model Improved robustness to image blur; cross-device testing Not reported [49] Subgroup analyses by age, BPE, BI-RADS, field strength Multi-center, multi-protocol external validation IRB approvals; code released Clinical decision-curve analyses External validation; missing-modality robustness Integrated gradients; Shapley values [11] Large multi-scanner dataset; external TCGA validation Multi-scanner, multi-task evaluation Not reported Minimal preprocessing; whole-brain input Robust to adversarial and noise perturbations Occlusion-based saliency maps [52] Large multi-modality public dataset; no subgroup analysis Single model across organs, modalities, and diseases Open data, code, and weights Prompt-based segmentation reduces annotation time Extensive internal and external validation Mask confidence scores provided