Paper deep dive
Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/18/2026, 4:35:30 AM
Summary
The paper introduces VIFBA, a cross-modal learning framework that uses fetal brain ultrasound videos to predict MRI-derived lateral ventricular volume and classify ventriculomegaly (VM) severity. It employs a JEPA-inspired tube latent prediction objective for representation learning, contrastive cross-modal alignment to transfer structural information from MRI to ultrasound, and a training-free vision-language model with retrieval augmentation for verifying predictions and identifying non-VM abnormalities. The model was validated on 857 cases, showing high accuracy in volume regression and classification.
Entities (10)
Relation Signals (8)
VIFBA → classifies → Ventriculomegaly
confidence 95% · classifies VM severity
VIFBA → predicts → Lateral Ventricular Volume
confidence 95% · predicts MRI-derived lateral ventricular volume
VIFBA → validateson → Dataset of 857 cases
confidence 93% · We validated VIFBA on a large dataset comprising 857 cases
VIFBA → uses → Tube Latent Prediction
confidence 92% · we introduce a joint-embedding predictive architecture (JEPA)-inspired tube latent prediction objective that leverages spatio-temporal coherence in ultrasound videos
Fetal Brain MRI → provides → Lateral Ventricular Volume
confidence 90% · Fetal brain MRI provides more reliable volumetric information
VIFBA → uses → Cross-Modal Alignment
confidence 90% · develop a contrastive cross-modal alignment strategy that transfers structural information from MRI to ultrasound during training
VIFBA → augmentedby → Vision-Language Model
confidence 88% · augment VIFBA with a training-free vision-language model and retrieval augmentation to verify uncertain predictions
FetalCLIP → usedin → VIFBA
confidence 85% · encoded using a pre-trained fetal ultrasound foundation model, i.e., the vision encoder from FetalCLIP
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement. Fetal brain MRI provides more reliable volumetric information but is costly and less accessible for routine use. To address these limitations, we propose VIFBA, an ultrasound video-based framework for fetal brain assessment that predicts MRI-derived lateral ventricular volume, classifies VM severity, and identifies potential non-VM fetal brain abnormalities. Our contribution is three-fold. First, we introduce a joint-embedding predictive architecture (JEPA)-inspired tube latent prediction objective that leverages spatio-temporal coherence in ultrasound videos to enhance representation learning. Second, we develop a contrastive cross-modal alignment strategy that transfers structural information from MRI to ultrasound during training, while requiring ultrasound alone at inference. Third, we augment VIFBA with a training-free vision-language model and retrieval augmentation to verify uncertain predictions and identify potential non-VM fetal brain abnormalities. We validated VIFBA on a large dataset comprising 857 cases (3,196 videos) with paired fetal brain ultrasound and MRI examinations. On held-out test data, VIFBA achieved an MAE of 0.5909 mL and Pearson correlation coefficient of 0.9907 for ventricular volume regression, 0.9400 accuracy for VM severity classification, and an F1 score of 0.7764 for multi-abnormality classification, substantially outperforming single-task baselines, video-based strong competitors, and state-of-the-art foundation models. By enabling MRI-informed volumetric assessment from routine ultrasound alone, VIFBA offers a practical and potentially broadly deployable pathway toward accurate and affordable prenatal brain screening.
Tags
Links
- Source: https://arxiv.org/abs/2608.14763v1
- Canonical: https://arxiv.org/abs/2608.14763v1
Trouble viewing inline? Open PDF directly →
Full Text
86,075 characters extracted from source content.
Expand or collapse full text
Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening Yuhao Huang Yuanji Zhang Yuhuan Lu Dong Ni P. Ellen Grant Davood Karimi Thanks: Yuhao Huang, Yuanji Zhang, Yuhuan Lu, P. Ellen Grant, and Davood Karimi are with Department of Radiology, Boston Children’s Hospital, and Harvard Medical School, Boston, USA. (Corresponding author: Davood Karimi, email: Davood.Karimi@childrens.harvard.edu) Thanks: Yuhao Huang, Yuanji Zhang, and Dong Ni are with Medical Ultrasound Image Computing (MUSIC) Lab, Shenzhen University, Shenzhen, China Thanks: Yuhao Huang is also with Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences, Hong Kong, China. Yuanji Zhang is also with Shenzhen Luohu People’s Hospital, Shenzhen, China. Dong Ni is also with School of Artificial Intelligence, Shenzhen University, Shenzhen, China and School of Biomedical Engineering and Informatics, Nanjing Medical University, Nanjing, China. Thanks: This work was supported by the Frontier Technology Development Program of Jiangsu Province (No. BF2024078) and National Natural Science Foundation of China (No. 12326619) Abstract Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement. Fetal brain MRI provides more reliable volumetric information but is costly and less accessible for routine use. To address these limitations, we propose VIFBA, an ultrasound video-based framework for fetal brain assessment that predicts MRI-derived lateral ventricular volume, classifies VM severity, and identifies potential non-VM fetal brain abnormalities. Our contribution is three-fold. First, we introduce a joint-embedding predictive architecture (JEPA)-inspired tube latent prediction objective that leverages spatio-temporal coherence in ultrasound videos to enhance representation learning. Second, we develop a contrastive cross-modal alignment strategy that transfers structural information from MRI to ultrasound during training, while requiring ultrasound alone at inference. Third, we augment VIFBA with a training-free vision-language model and retrieval augmentation to verify uncertain predictions and identify potential non-VM fetal brain abnormalities. We validated VIFBA on a large dataset comprising 857 cases (3,196 videos) with paired fetal brain ultrasound and MRI examinations. On held-out test data, VIFBA achieved an MAE of 0.5909 mL and Pearson correlation coefficient of 0.9907 for ventricular volume regression, 0.9400 accuracy for VM severity classification, and an F1 score of 0.7764 for multi-abnormality classification, substantially outperforming single-task baselines, video-based strong competitors, and state-of-the-art foundation models. By enabling MRI-informed volumetric assessment from routine ultrasound alone, VIFBA offers a practical and potentially broadly deployable pathway toward accurate and affordable prenatal brain screening. Index Terms: Foundation Model, Ultrasound-MRI Alignment, Vision-Language Model, Fetal Brain, Ventriculomegaly I Introduction Ventriculomegaly (VM), defined as the abnormal enlargement of the fetal lateral ventricles, is among the most common central nervous system abnormalities [53]. Its clinical significance lies in the wide variability of outcomes, ranging from normal neurodevelopment to severe impairment, depending on its severity and underlying etiology. Therefore, accurate quantification of ventricular size, alongside reliable VM diagnosis, is critical for the detection and severity assessment of VM, as well as for risk stratification and clinical decision-making. In prenatal ultrasound examinations, VM is primarily assessed by measuring the atrial width of the lateral ventricles on selected standard brain planes. Based on this measurement, VM is defined as ≥ 10 m and further stratified into mild (10–12 m), moderate (13–15 m), and severe (>>15 m) categories, while values <<10 m are considered normal [18]. However, this process requires the accurate identification of the standard plane and key anatomical landmarks to measure the maximal atrial width, making it highly dependent on operator expertise and susceptible to error accumulation [48]. Compared to 2D width measurements, ventricular volume provides a global quantitative measure that better reflects the full spectrum of ventricular conditions, from normal morphology to ventricular enlargement. In clinical practice, volumetric assessment typically relies on 3D imaging modalities, such as 3D ultrasound and magnetic resonance imaging (MRI). Among these, MRI generally provides clearer structural boundaries of the lateral ventricles, enabling more reliable segmentation of the ventricles for volumetric estimation [35]. However, MRI is costly and not routinely accessible, limiting its use in large-scale prenatal screening. This motivates the development of automatic methods that enable direct estimation of ventricular volume from ultrasound scans. Despite these motivations, this task remains highly challenging for several reasons. First, it relies on paired ultrasound–MRI data, which are scarce in clinical practice and difficult to acquire at scale. Second, fetal brain ultrasound examinations typically consist of a variable number of videos acquired from different views (e.g., sagittal and axial planes), making effective multi-view information aggregation non-trivial. Third, VM severity assessment involves fine-grained categorization (i.e., mild/moderate/severe), leading to inherent data imbalance across categories, particularly for severe cases. Moreover, other fetal brain abnormalities may also be present, further increasing the complexity of the task. To address these challenges, we propose a novel framework, termed VIFBA (VIdeo-based Fetal Brain Assessment), for joint estimation of ventricular volume and VM severity from ultrasound videos. VIFBA can be flexibly integrated with different ultrasound foundation models as shared backbones and employs two task-specific branches for ventricular volume regression and VM classification, respectively. Its contribution is three-fold. First, we introduce a tube latent prediction loss to enhance representation learning by leveraging spatio-temporal coherence in ultrasound videos. Second, we design a contrastive cross-modal fusion strategy to inject MRI information into ultrasound representations during training, enabling MRI-informed feature learning from ultrasound alone at inference. Third, we extend VIFBA by integrating the retrieval-augmented training-free visual-language model (VLM) to further optimize the uncertain predictions and alert the potential non-VM brain abnormality. VIFBA was validated on a large ultrasound-MRI paired dataset with 857 cases, demonstrating its effectiveness for fetal lateral ventricular analysis. I Related work I-A Deep Learning in Fetal Brain Analysis In recent years, deep learning has driven rapid progress in prenatal imaging [4, 50, 72, 37, 78], particularly in automated and intelligent fetal brain analysis. Clinically, fetal brain assessment mainly relies on ultrasound and MRI, where ultrasound serves as the primary modality for routine prenatal screening, while fetal MRI is commonly used as a complementary tool for clearer visualization of intracranial structures and further evaluation of suspected abnormalities [48, 19]. Accordingly, existing studies have mainly evolved along two relatively distinct directions, namely ultrasound-based analysis and MRI-based analysis, which are reviewed separately below. In fetal brain ultrasound, a major line of research has focused on automatic standard plane localization [4, 72, 71, 14], view classification [28], and quality assessment [40, 20, 8], aiming to ensure diagnostically reliable views for subsequent analysis. Early downstream studies mainly concentrated on segmentation or measurement of single or limited anatomical targets and parameters, including segmentation of the fetal head [41, 75, 27], choroid plexus and corpus callosum [25], and cerebellum [58], as well as measurement of head circumference [75, 59], biparietal diameter [59], and lateral ventricular width [10]. Recent studies have further moved toward automatic delineation and measurement of multiple structures [13], as well as fine-grained segmentation of 25 anatomical structures across five standard planes [44]. Xie et al. [67] investigated deep learning algorithms for classifying normal or abnormal brains. More recently, Duan et al. [15] proposed an anatomy-guided diffusion framework to synthesize normal and abnormal brain images, showing its potential to improve downstream performance. In fetal brain MRI, existing studies have mainly focused on volumetric reconstruction, anatomical segmentation, and quantitative structural analysis. Since fetal motion often leads to inter-slice misalignment in clinical MRI acquisitions, many methods first aim to reconstruct high-quality brain volumes from stacks of 2D slices [68, 17]. Based on reconstructed volumes, subsequent studies have developed automated segmentation and parcellation methods for fetal brain tissues and anatomical regions, enabling quantitative analysis of various structures [54, 62, 12, 73, 74]. Vahedifard et al. [63] automated VM detection by measuring lateral ventricular width on reconstructed fetal MRI planes. More recently, Huang et al. proposed BrainSeg [26] for generalized brain tissue segmentation, parcellation, and lesion labeling across diverse MRI contrasts and populations spanning fetal to adult stages. Despite notable progress in both ultrasound- and MRI-based fetal brain analysis, most existing studies investigate the two modalities separately and rarely exploit their complementary information. Moreover, VM assessment is still largely limited to 2D plane-based linear measurements, which may not fully reflect the overall volumetric enlargement of the lateral ventricles. In addition, abnormality analysis often focuses on coarse normal/abnormal classification or a few common diseases, leaving diverse and coexisting fetal brain abnormalities insufficiently explored. Besides, to the best of our knowledge, existing studies lack cross-modal predictive investigations, particularly the task of estimating MRI-derived volumetric measurements directly from ultrasound videos. I-B Medical Ultrasound Foundation Model Foundation models have recently become an important paradigm in medical image analysis, as large-scale pretraining enables transferable representations that can be adapted to diverse downstream tasks with limited annotations [69, 30, 52]. For example, BiomedCLIP learned transferable biomedical vision-language representations from large-scale image-text pairs [77]. Medical segment anything models (SAMs) demonstrated strong generalization ability for promptable image segmentation [42, 29]. MedSAM2 further adapted the SAM2 paradigm to both videos and 3D images, enabling promptable segmentation with spatial or temporal consistency across slices and frames [43]. Beyond SAM, several studies have also explored Self-Distillation with No Labels (DINO) variants [6, 51] for various medical image analysis tasks [60, 57, 55], achieving promising performance. Recently, foundation models have also been increasingly explored in general ultrasound imaging. USFM leveraged a large-scale multi-organ, multi-center, and multi-device ultrasound dataset to pretrain the foundation model, showing strong label efficiency across segmentation, classification, and enhancement tasks [31]. EchoCare further increased the scale of ultrasound pretraining data (2 million→ 4.5 million images), and validated its generalizability on a broader range of downstream clinical tasks [76]. Most recently, Ultrasound-CLIP explored semantic-aware contrastive pretraining for general ultrasound understanding using 365k image-text paired samples [32]. In cardiac ultrasound, EchoCLIP [11] and EchoPrime [64] were introduced to learn vision-language representations from echocardiography image- or video-text pairs for comprehensive cardiac evaluation. EchoONE investigated SAM-based multi-plane echocardiography segmentation within a unified model [24]. Recently, FrameONE explored the multi-view keyframe detection task in cardiac videos [9]. In fetal ultrasound, FetalCLIP first learned generalizable fetal ultrasound representations from paired image-text data [45], while Sonomate introduced a visually grounded language model by aligning ultrasound video features with transcribed sonographer speech for anatomy detection and visual question answering [21]. These studies show the potential of foundation models for fetal ultrasound understanding; however, their application to fetal brain video analysis, especially ventricular volume estimation and fine-grained VM severity and brain abnormality assessment, remains insufficiently explored. I-C LLM/VLM-assisted Medical Image Analysis Recent studies have explored the integration of large language models (LLMs) or VLMs with specialized medical imaging systems to support downstream analysis. ChatCAD [65] and ChatCAD+ [79] integrated outputs from computer-aided diagnosis models with LLMs for unified interpretation and reliable report generation. In prenatal ultrasound, FAA-Net [38] incorporated LLM-derived medical knowledge into a multi-instance learning framework to identify diagnostically relevant information for abdominal anomaly analysis. Most recently, See-in-Pairs [33] provided VLMs with matched reference images for comparative medical diagnosis, while RAD-SRAC [34] employed retrieved examples as contextual guidance for training-free radiological image classification. Nevertheless, existing studies mainly focus on directly supporting diagnosis, classification, report generation, or interpretation. Their applicability remains relatively limited in scenarios that require the reliability verification of existing model predictions or the identification of abnormalities not represented during training. This limitation is particularly relevant to our complex fetal brain analysis, where heterogeneous clinical evidence, including current ultrasound videos, model predictions, and retrieved visual references and textual report information, needs to be jointly considered to assess quantitative regression and categorical outputs, while potential coexisting abnormalities also need to be identified. I Method Fig. 1: Overview of the training pipeline of our proposed method. Figures 1 and 2 show the overall framework of our proposed VIFBA. During training, each ultrasound video is treated as an individual training sample and encoded using a pre-trained vision foundation model to learn shared representations, followed by two task-specific branches for ventricular volume regression and VM classification. In addition, paired MRI data are leveraged to provide complementary supervision, enabling MRI-informed feature learning. During inference, predictions from multiple videos belonging to the same case are aggregated at the case level, where VIFBA relies solely on ultrasound videos and incorporates a training-free VLM with retrieval augmentation to refine uncertain predictions, making the overall framework more flexible for variable clinical acquisition protocols. Details refer to the following sections. I-A Foundation Model-driven Ultrasound Video Modeling Given an ultrasound video =Itt=1TV=\I_t\_t=1^T, we uniformly sample T frames and encode each frame using a pre-trained fetal ultrasound foundation model, i.e., the vision encoder from FetalCLIP [45]. Specifically, the vision encoder is based on a Vision Transformer (ViT) with an input resolution of 224×224224× 224, a patch size of 14, 24 Transformer blocks, and a hidden width of 1024. For each frame ItI_t, the encoder produces a global token tclsz^cls_t, corresponding to the class (CLS) token, together with a sequence of patch tokens tpatch=t,npatchn=1NZ^patch_t=\z_t,n^patch\_n=1^N, where t,npatchz_t,n^patch denotes the representation of the n-th spatial patch and N is the number of patches in each frame. These two types of representations serve as inputs to different branches of our framework, enabling both global video-level modeling and local spatio-temporal representation learning. For the main task branch, the global tokens from all sampled frames are arranged as a temporal sequence cls=[1cls,2cls,…,Tcls]∈ℝT×dZ^cls=[z^cls_1,z^cls_2,…,z^cls_T] ^T× d. To encode frame order, we introduce a learnable temporal positional embedding matrix temp∈ℝT×dE_temp ^T× d, where each row corresponds to a temporal position. The input to the temporal Transformer is then given by ~cls=cls+temp Z^cls=Z^cls+E_temp. The resulting sequence is then processed by a temporal Transformer encoder (⋅)T(·) composed of multi-head self-attention and feed-forward layers to capture inter-frame dependencies. The temporally encoded frame sequence is defined as cls=(~cls)H^cls=T( Z^cls). We then apply mean pooling over the temporal dimension to obtain a compact video-level representation, which is formulated as: vid=1T∑t=1Ttcls.h^vid= 1T _t=1^TH^cls_t. (1) This representation is further transformed by a shared projection block ϕ(⋅)φ(·) composed of Layer Normalization, a linear layer, and a ReLU activation, yielding =ϕ(vid)h=φ(h^vid). Finally, two task-specific linear heads are built upon h, including a regression head freg(⋅)f_reg(·) for ventricular volume estimation and a classification head fcls(⋅)f_cls(·) for disease prediction: y^reg=freg(),^cls=fcls(). y_reg=f_reg(h), y_cls=f_cls(h). (2) Specifically, the regression head is a linear layer mapping ∈ℝdhh ^d_h to a scalar prediction (y^reg y_reg), while the classification head is a linear layer mapping ∈ℝdhh ^d_h to a C-class logit vector (^cls y_cls). The regression and classification branches are optimized using a mean squared error (MSE) loss and a class-balanced weighted cross-entropy (CE) loss, controlled by λreg _reg and λcls _cls: ℒsup=λregℒMSE+λclsℒCE.L_sup= _regL_MSE+ _clsL_CE. (3) I-B Tube Latent Prediction for Representation Enhancement Although the above supervised regression and classification objectives encourage the model to learn discriminative video-level representations, they impose only limited constraints on the intrinsic local spatio-temporal structure of ultrasound videos. This issue becomes more pronounced when training data and fine-grained annotations are limited, hindering the model from fully exploiting the rich anatomical and dynamic information present in the videos. Inspired by world-model-style representation learning [22] and joint-embedding predictive architectures, i.e., JEPA [1, 3, 2], we encourage the model to infer masked content from visible context in the latent space, thereby learning representations with stronger structural consistency and temporal predictability. Specifically, we build a tube latent prediction (TLP) objective upon the patch tokens defined above. Given the input video V, we first apply spatio-temporal tube masking to obtain a masked video ~ V. Let ∈0,1T×NM∈\0,1\^T× N denote a binary mask over the patch grid, where Mt,n=1M_t,n=1 indicates that the n-th patch token at time step t is masked. To encourage temporal coherence, masking is performed in a tube-wise manner, such that once a spatial patch location is selected, it is masked over a contiguous temporal span. For each frame ItI_t, the masked frame is defined as I~t=It⊙(1−Γ(Mt)) I_t=I_t (1- (M_t)), where ⊙ denotes element-wise multiplication and Γ(⋅) (·) maps the token-level mask to the corresponding pixel-space masking pattern on the input frame. The masked video is then given by ~=I~tt=1T V=\ I_t\_t=1^T. Similar to [49], we introduce a teacher-student design for the TLP branch, where the main task branch serves as the student encoder (Es(⋅)E_s(·)). ~ V is then fed into Es(⋅)E_s(·), while V is processed by the teacher encoder (Et(⋅)E_t(·)). Specifically, the parameters of Et(⋅)E_t(·) are updated as an exponential moving average (EMA) of those of Es(⋅)E_s(·). The corresponding student and teacher patch-token representations are given by: patch,s=Es(~),patch,t=Et(),Z^patch,s=E_s( V), ^patch,t=E_t(V), (4) To further capture temporal dependencies across frames for each spatial patch location, we introduce a temporal token encoder ℱtok(⋅)F_tok(·) operating along the temporal dimension. The student and teacher patch-token sequences are augmented using a shared learnable temporal positional embedding tok∈ℝT×dpE_tok ^T× d_p, while the teacher temporal token encoder ℱtok′(⋅)Ftok (·) is maintained as an EMA copy of the student encoder ℱtok(⋅)F_tok(·). The temporally contextualized student and teacher token representations are then defined as: patch,s=ℱtok(patch,s+tok),H^patch,s=F_tok\! (Z^patch,s+E_tok ), (5) patch,t=ℱtok′(patch,t+tok).H^patch,t=F_tok \! (Z^patch,t+E_tok ). (6) Then, the student token representations are further transformed by a predictor g(⋅)g(·) to obtain the predicted token representations, i.e., ^patch,s=g(patch,s) H^patch,s=g(H^patch,s). The TLP objective is computed only over the masked token positions. Let Ω=(t,n)∣Mt,n=1 =\(t,n) M_t,n=1\ denote the set of masked spatio-temporal tokens. The TLP loss is defined as: ℒTLP=λcosℒcos+λℓ1ℒℓ1,L_TLP= _ L_ + _ _1L_ _1, (7) ℒcos=1|Ω|∑(t,n)∈Ω(1−cos(^t,npatch,s,t,npatch,t)),L_ = 1| | _(t,n)∈ (1- ( H^patch,s_t,n,H^patch,t_t,n ) ), (8) ℒℓ1=1|Ω|∑(t,n)∈Ω‖^t,npatch,s−t,npatch,t‖1.L_ _1= 1| | _(t,n)∈ \| H^patch,s_t,n-H^patch,t_t,n \|_1. (9) I-C MRI-Guided Cross-Modal Training Alignment During training, we have access to paired ultrasound-MRI data for each case. At inference time, only ultrasound videos are available. Compared with ultrasound, MRI generally provides clearer visualization of brain anatomy and more stable structural information, although it is substantially more expensive and less accessible in routine clinical practice. Therefore, we perform feature-level alignment between the two modalities during training, encouraging the ultrasound encoder to map input videos into an MRI-aligned structural feature space without relying on MRI at test time. To obtain structurally informative MRI representations, we leverage BOUNTI [62], a generalized model for fetal brain MRI. Specifically, for each MRI volume, we follow the BOUNTI pipeline and use its pretrained segmentation models, consisting of a U-Net and an Attention U-Net, as strong feature extractors. The bottleneck features from the deepest encoder layer of both networks are obtained, globally average-pooled, and concatenated to form a 768-dimensional MRI representation. To enable cross-modal alignment, we further introduce a lightweight projector ψ(⋅)ψ(·) to map the MRI representation into the shared embedding space of the ultrasound branch: MRI=ψ(MRI),h^MRI=ψ(f^MRI), (10) where MRIf^MRI denotes the original MRI feature and MRIh^MRI is the projected MRI representation. The projector ψ(⋅)ψ(·) is implemented as LayerNorm followed by a linear layer and ReLU. Similarly, we introduce another projector φ(⋅) (·) to transform the ultrasound representation h into the aligned feature space: US=φ(),h^US= (h), (11) where USh^US denotes the aligned ultrasound representation. φ(⋅) (·) has the same architecture as ψ(⋅)ψ(·). We then perform feature-level alignment between USh^US and MRIh^MRI in the shared embedding space. Following common contrastive learning [56], both features are first normalized, and their pairwise similarities are computed within each mini-batch with size B: ij=τ⋅sim(iUS,jMRI),S_ij=τ·sim(h^US_i,h^MRI_j), (12) where sim(⋅,⋅)sim(·,·) denotes cosine similarity and τ is a learnable scaling factor. Let ∈0,1B×BP∈\0,1\^B× B denote the positive-pair mask, where Pij=1P_ij=1 if the i-th ultrasound sample and the j-th MRI sample belong to the same case, and Pij=0P_ij=0 otherwise. Each row of P is normalized to form the target distribution: ~ij=Pij∑k=1BPik. P_ij= P_ij _k=1^BP_ik. (13) The alignment loss is then defined as ℒalign=12(ℒUS→MRI+ℒMRI→US),L_align= 12 (L_US +L_MRI ), (14) with ℒUS→MRI=−1B∑i=1B∑j=1B~ijlogexp(ij)∑k=1Bexp(ik),L_US =- 1B _i=1^B _j=1^B P_ij (S_ij) _k=1^B (S_ik), (15) and ℒMRI→US=−1B∑i=1B∑j=1B~jilogexp(ji)∑k=1Bexp(ki).L_MRI =- 1B _i=1^B _j=1^B P_ji (S_ji) _k=1^B (S_ki). (16) Through optimizing ℒalignL_align, the aligned ultrasound representation US→MRIh^US→ MRI is encouraged to capture MRI-consistent structural information. The original ultrasound representation h is then concatenated with US→MRIh^US→ MRI and further transformed by a lightweight fusion module ϕfuse _fuse (i.e., LayerNorm, Linear, and ReLU) to form the final feature for prediction: fusion=ϕfuse([;US→MRI]),h^fusion= _fuse ([h;h^US→ MRI] ), (17) where [⋅;⋅][·;·] denotes operation of feature concatenation. Fig. 2: Testing pipeline of our proposed method, including case-level testing, and VLM verification and non-VM brain abnormality alerting. I-D VLM Verification and Non-VM Brain Abnormality Alerting As mentioned above, video-level training is leveraged to maximize data utilization and facilitate flexible optimization, whereas inference is performed at the case level, where all videos acquired from the same subject are jointly considered to obtain robust predictions. Specifically, let a test case contain K (K≥1K≥ 1) videos. For the k-th video, the model outputs a regression prediction y^kreg y^reg_k, a classification logit vector ^kcls y^cls_k, and an intermediate feature representation k∈ℝdh_k ^d. We average the multi-video features to obtain the case-level representation caseh_case, which is then fed into the regression and classification heads to obtain y^casereg y^reg_case and ^casecls y^cls_case. The final classification label is determined by applying argmax to ^casecls y^cls_case. Although case-level fusion improves robustness, uncertainty may still remain due to heterogeneous video quality, incomplete anatomical observations, and inconsistent predictions across views, particularly for underrepresented categories. To further improve reliability, we introduce a retrieval augmentation strategy that uses clinically similar historical cases to provide useful reference patterns for correcting uncertain predictions. Using caseh_case as the query, we retrieve the top-M most similar cases from the training pool in the learned feature space, forming the retrieval set ℛ=r1,…,rMR=\r_1,…,r_M\. We further employ a training-free VLM as a auxiliary verifier to assess prediction reliability and refine uncertain cases. Specifically, given a query case, we uniformly sample S representative frames from the K available ultrasound videos to construct the query visual set f1,1q,f1,Sq,…,fK,Sq\f^q_1,1,f^q_1,S,…,f^q_K,S\. In this study, S is set to 8 to balance performance and inference speed. The associated textual prior is provided by the primary model prediction, including the ventricular volume estimate y^casereg y^reg_case and the predicted diagnostic category argmax(^casecls) ( y^cls_case). After retrieving the top-M most similar cases from ℛR, we further collect representative frames from each retrieved case to form the reference visual set fm,1r,…,fm,Srm=1M\f^r_m,1,…,f^r_m,S\_m=1^M. The corresponding reference text information includes their known ventricular volumes ymregm=1M\y^reg_m\_m=1^M and diagnostic labels cmm=1M\c_m\_m=1^M. The query information and retrieved reference evidence are then jointly organized into a multimodal prompt for subsequent VLM-based verification. Specifically, the prompt template mainly contains the following four components: • Global instruction: “You are an assistant specialized in fetal brain ultrasound analysis. Given the query case and several retrieved reference cases, determine whether the current prediction is clinically reasonable based on visual similarity and diagnostic consistency.” • Reference content: “Here are several retrieved reference cases. [Visual] Representative ultrasound frames from the retrieved reference case 1. [Text]: lateral ventricular volume=y1regy^reg_1, diagnosis=c1c_1.”, …, [Visual] Representative ultrasound frames from the retrieved reference case m. [Text]: lateral ventricular volume=ymregy^reg_m, diagnosis=cmc_m.” • Query content: “Here is the current query case. [Visual] Sampled query ultrasound frames f1,1q,…,fK,Sq\f^q_1,1,…,f^q_K,S\. [Text]: predicted lateral ventricular volume=y^casereg y^reg_case, predicted diagnosis category=cargmax(^casecls)c ( y^cls_case).” • Question: “Based on the provided reference contents, determine whether the prediction for the current query case is clinically reliable according to visual similarity and diagnostic consistency. Besides, return a confidence score between 0 and 1 indicating how confident you are in the recalibration suggestion.” After prompting, the VLM outputs a verification decision d∈0,1d∈\0,1\, where d=0d=0 indicates that the current prediction is clinically reliable and the original case-level outputs are directly retained. Otherwise, d=1d=1 indicates that calibration is required. Then, the VLM further evaluates the relevance of each retrieved reference case rm∈ℛr_m with respect to the query case, and assigns a matching confidence score αm∈[0,1] _m∈[0,1]. The normalized reference weights are computed as wm=αm∑j=1Mαjw_m= _m _j=1^M _j. We then aggregate the retrieved labels using these weights to obtain the reference prediction using the following equations: ret=∑m=1Mwmm,p_ret= _m=1^Mw_mp_m, (18) y^retreg=∑m=1Mwmymreg, y^reg_ret= _m=1^Mw_my^reg_m, (19) where mp_m and ymregy^reg_m denote the one-hot class label distribution and lateral ventricular volume of the m-th retrieved case, respectively. Then, with the pre-set global adjustment score β=0.3β=0.3, the calibrated outputs are computed and updated using Eqs. 20-21. Besides, the final diagnosis category of the query case is determined by applying argmax to finalp_final. final=case,d=0,(1−β)case+βret,d=1,p_final= casesp_case,&d=0,\\ (1-β)p_case+ _ret,&d=1, cases (20) y^finalreg=y^casereg,d=0,(1−β)y^casereg+βy^retreg,d=1. y^reg_final= cases y^reg_case,&d=0,\\ (1-β) y^reg_case+β y^reg_ret,&d=1. cases (21) In clinical practice, VM is frequently associated with other fetal brain abnormalities, such as posterior fossa abnormalities (e.g., Dandy-Walker malformation and Blake’s pouch cyst), neural tube defects (e.g., Chiari I malformation), and other ventricular system abnormalities beyond the lateral ventricles (e.g., aqueductal stenosis and third/fourth ventricle abnormalities). However, these abnormalities are challenging to diagnose with common deep learning models, mainly for two reasons. First, they often require assessment of diverse anatomical regions across the fetal brain, whereas VM prediction mainly focuses on the lateral ventricles. Second, they commonly co-occur with VM or with each other, rather than appearing as isolated findings, complicating the intelligent diagnosis. Therefore, beyond refining VM-related predictions, it is clinically valuable to further provide alerts for potential non-VM brain abnormalities. Motivated by this observation, we explore and extend the retrieval-augmented VLM framework to perform auxiliary abnormality-aware screening. Specifically, the prompt structure remains unchanged, while the reference content and verification question are further enriched with additional diagnostic descriptions and updated to: • Reference content∗: “[…] Associated findings=posterior fossa abnormalities (It refers to structural abnormalities involving the posterior fossa region of the fetal brain. Among them, Dandy-Walker malformation is typically characterized by hypoplasia or agenesis of the cerebellar vermis and enlargement of the posterior fossa.).” • Question∗: “[…] In addition, assess whether the query case shows evidence of possible non-VM brain abnormalities. If yes, provide an abnormality alert score between 0 and 1, together with the most likely suspected findings according to the retrieved references.” With the above prompt enhancement, the VLM is able to leverage retrieved abnormal references and the corresponding descriptions to perform auxiliary non-VM abnormality screening in a training-free manner. Specifically, it outputs a ranked list of suspected abnormalities, together with abnormality alert scores sii=1N\s_i\_i=1^N for multiple candidate findings. Here, si∈[0,1]s_i∈[0,1] denotes the confidence score of the i-th abnormality type. All findings satisfying sis_i>>τ, where τ is a predefined threshold, are reported as potential abnormality warnings. IV Experiments Fig. 3: Volume calculation overview in our data preprocessing stage. Fig. 4: Details of our dataset. (a) bilateral ventricle volume between fetal subjects in the Developing Human Connection Project [16] (segmentation results generated using DrawEM [46, 47] and quality-controlled by the dHCP team) and ours (segmented by deep models [62, 73]). Together, the scatter plot analysis and expert visual inspection provide further validation of the accuracy of the automatic segmentation method. (b) Statistical distribution of bilateral ventricle volume in non-VM and mild, moderate, and severe VMs across different gestational age groups (weeks). (c) Distribution of case counts across six abnormality categories (A1–A6, details refer to Section IV-A) among different groups, including non-VM abnormal brain and mild, moderate, and severe VMs. (d) The UpSet plot illustrates the set size of each category (VM and A1–A6) and the quantitative distribution of their complex intersections. IV-A Datasets In this study, approved by the Institutional Review Board of Boston Children’s Hospital, we collected a large ultrasound-MRI paired fetal brain dataset for method validation between January 2022 and February 2026. The dataset comprises 857 cases, each with several stacks of MRI slices for subsequent 3D volume reconstruction and a variable number of ultrasound videos. In total, 3,196 videos are included (3.73 videos per case on average). The gestational age (GA) ranges from 16.57 to 38.00 weeks, with a mean of 26.31 and a median of 25.00 weeks. A total of 707, 50, and 100 cases were randomly selected for training, validation, and testing. In the data preparation pipeline, we first used the NesVoR [68] to create the fetal brain MRI volumes, which are subsequently processed using state-of-the-art methods, including BOUNTI [62] and FeTA [73], to obtain segmentation results of the bilateral lateral ventricles (as shown in Figure 3). After careful expert review, the final ground truth (GT) for ventricular volume is computed by summing the segmented voxels and multiplying by the corresponding voxel spacing. In addition, diagnostic labels (e.g., normal vs. abnormal fetal brain) are extracted from de-identified radiology reports using a large language model (LLM), Qwen3.5-Plus with strong deep reasoning ability [70]. Specifically, the labels include normal brain (299), mild VM (182), moderate VM (62), severe VM (37), non-VM abnormal brain (277). Based on the radiology reports and LLM analysis, we further provide diagnostic results for various brain abnormalities apart from VM, covering A1: neural tube defects (e.g., Spina bifida, Encephalocele), A2: midline structure malformations (e.g., Agenesis of the corpus callosum, Holoprosencephaly), A3: ventricular system disorders (e.g., Aqueductal stenosis, Hydrocephalus), A4: posterior fossa malformations (e.g., Dandy-Walker malformation), A5: neurodevelopmental abnormalities (e.g., Microcephaly), and A6: brain parenchymal abnormalities (e.g., Intracranial cyst, Arachnoid cyst). Among the 281 VM cases, 128 (45.6%) were isolated, whereas 84 (29.9%), 46 (16.4%), 18 (6.4%), and 5 (1.8%) were associated with 1, 2, 3, and 4 additional brain abnormality categories, respectively. Correspondingly, the numbers of VM cases associated with categories A1-A6 were 40, 63, 22, 54, 61, and 10, respectively. Among the 277 non-VM abnormal cases, 185, 55, and 12 were associated with 1, 2, and 3 brain abnormality categories, respectively; correspondingly, the numbers of cases assigned to categories A1-A6 were 11, 71, 24, 100, 88, and 37, respectively. Notably, 25 of the 277 cases belonged to a relatively uncommon group of other brain abnormalities, such as dural sinus malformation, intracranial venous thrombosis, venous sinus malformation, and Vein of Galen malformation, that fell outside the scope of categories A1-A6 and therefore were not considered in our study. More details about the dataset in this study are provided Figure 4. All generated labels are reviewed and verified by a fetal ultrasound expert with over 10 years of clinical experience. IV-B Implementation Details We implemented VIFBA using Python (v3.9.25) and PyTorch (v2.0.1). All experiments were conducted on a single NVIDIA RTX 6000 Ada Generation GPU with 48 GB memory. We trained the model for 100 epochs with a batch size of 8 using the AdamW optimizer (learning rate=1e−41e^-4). For the MRI projector, we used a reduced learning rate with a scaling factor of 0.1 relative to the main optimizer. For ultrasound video processing, 64 frames were uniformly sampled from each video, cropped to remove sensitive information, and resized to 224×224224× 224. The temporal Transformer in the main ultrasound branch consisted of 2 layers with 8 attention heads, an MLP ratio of 4.0, and a dropout rate of 0.1. The temporal token encoder in the TLP branch used the same configuration. We applied LoRA [23] with rank 8 and α=16 to effectively fine-tune the FetalCLIP encoder. Gradient checkpointing was enabled during training to reduce memory usage. For TLP training, we set the mask ratio to 0.4, the temporal tube length to 4, and the EMA decay for the teacher encoder to 0.996. For different loss functions, the weights were set to λreg=1.0 _reg=1.0, λcls=1.0 _cls=1.0, λalign=0.1 _align=0.1, λcos=1.0 _ =1.0, and λℓ1=0.1 _ _1=0.1. TABLE I: Regression performance across different groups using video-level output aggregation and case-level feature aggregation. Red subscripts denote the absolute performance gains of case-level feature aggregation over video-level output aggregation. G1: Healthy brain, G2: Non-VM abnormal brain, G3: Mild VM, G4: Moderate VM, and G5: Severe VM. Level Group MSE↓ RMSE↓ MAE↓ Pearson↑ Spearman↑ Video G1 0.2531 0.5031 0.4390 0.9535 0.9295 G2 0.7298 0.8543 0.7338 0.9618 0.8809 G3 2.5167 1.5864 1.4416 0.9208 0.8684 G4 7.0826 2.6613 2.2956 0.9611 1.0000 G5 7.6917 2.7734 2.5630 0.9479 1.0000 Overall 1.6476 1.2836 0.9364 0.9793 0.9403 Case G1 0.04730.2058 0.21750.2856 0.19850.2405 0.99160.0381 0.96850.0390 G2 0.23270.4971 0.48240.3719 0.43410.2997 0.98790.0261 0.92170.0408 G3 1.35681.1599 1.16480.4216 1.03890.4027 0.95260.0318 0.93510.0667 G4 2.89504.1876 1.70150.9598 1.52070.7749 0.98420.0231 1.00000.0000 G5 4.26253.4292 2.06460.7088 1.87060.6924 0.97070.0228 1.00000.0000 Overall 0.75070.8969 0.86640.4172 0.59090.3455 0.99070.0114 0.97300.0327 The proposed VIFBA model was evaluated on both regression and classification tasks. For volume regression, performance was assessed using mean absolute error (MAE), root mean square error (RMSE), Mean Squared Error (MSE), Pearson’s correlation coefficient, and Spearman’s rank correlation coefficient. For VM evaluation, performance was evaluated using accuracy (Acc), macro-precision (Pre), macro-recall (Rec), macro-F1 (F1), and macro-AUC (AUC). For multi-label classification, performance was evaluated using Pre, Rec, F1, Hamming loss (HL), and exact match ratio (EMR). Specifically, HL and EMR are formally defined as: HL=1NC∑i=1N∑j=1C(yij≠y^ij),HL= 1NC _i=1^N _j=1^CI(y_ij≠ y_ij), (22) EMR=1N∑i=1N(i=^i),EMR= 1N _i=1^NI (y_i= y_i ), (23) where N denotes the number of samples and C denotes the number of labels. iy_i and ^i y_i represent the ground-truth and predicted label vectors of the i-th sample, respectively. (⋅)I(·) is the indicator function, which returns 1 when the condition is satisfied and 0 otherwise. Lower HL and higher EMR indicate better multi-label classification performance. Unless otherwise specified, all experiments were conducted using fixed random seeds for fair comparison. In addition, to better assess the robustness of different methods across random initializations and to evaluate the statistical significance of performance differences, selected experiments were repeated with 10 different random seeds. For statistical analysis, we first computed the paired differences between two methods across the 10 repeated runs and assessed their normality using the Shapiro-Wilk test. If the paired differences satisfied the normality assumption, a one-sided paired t-test was used to examine whether the target method significantly outperformed the compared method. Otherwise, the Wilcoxon signed-rank test was used as a non-parametric alternative for the same one-sided paired comparison. For MSE, RMSE, and MAE, significance was tested in the direction of lower values, whereas for Pearson, Spearman, and classification metrics, significance was tested in the direction of higher values. Statistical significance was determined at p<0.05p<0.05. TABLE I: Method Comparison on the regression task. MSE RMSE MAE Pearson Spearman I3D 12.6789 3.5607 2.7929 0.8470 0.6547 R(2+1)D 13.0733 3.6157 2.5684 0.8243 0.6678 TimeSformer 11.8981 3.4494 2.6758 0.8906 0.7331 InternVideo2 11.4605 3.3853 2.5386 0.9510 0.8318 VideoMamba 12.0456 3.4707 2.6615 0.8752 0.7496 V-JEPA2 10.8888 3.2998 2.6546 0.8821 0.6937 Baseline (Reg) 9.9199 3.1496 2.6122 0.9033 0.7648 VIFBA w/o Cls 3.2208 1.7947 1.4461 0.9849 0.9562 VIFBA 0.7507 0.8664 0.5909 0.9907 0.9730 TABLE I: Method Comparison on the classification task. Acc Pre Rec F1 AUC I3D 0.7000 0.6361 0.7611 0.6610 0.8899 R(2+1)D 0.7000 0.6264 0.7086 0.6450 0.8659 TimeSformer 0.7300 0.6462 0.7196 0.6666 0.9006 InternVideo2 0.7500 0.6605 0.7545 0.6762 0.9030 VideoMamba 0.7300 0.6498 0.7424 0.6670 0.8996 V-JEPA2 0.7700 0.7247 0.7604 0.7208 0.9199 Baseline (Cls) 0.7900 0.7387 0.8256 0.7568 0.9186 VIFBA w/o Reg 0.8600 0.8075 0.8574 0.8173 0.9500 VIFBA 0.9400 0.9357 0.9068 0.9130 0.9793 Fig. 5: Comparisons with statistical significance analysis between VIFBA and competitors on (a) regression and (b) classification tasks. Each cell reports the p-value and the mean gain Δ of VIFBA. Fig. 6: Group-wise detailed regression performance, including normal brain (n=45), non-VM abnormal brain (n=24), mild VM (n=19), moderate VM (n=7), and severe VM (n=5). Y-axis represents lateral ventricular volume (mL). For each group, cases are sorted by the ground-truth lateral ventricular volume (green circles). The blue circles indicate the video-level predictions, while the red circles show the case-level output. IV-C Comparison with Video-Based Methods Tables I and I compare the proposed method with representative video-based models on lateral ventricular volume regression and VM severity classification, respectively. Specifically, we selected six representative and state-of-the-art video analysis methods as competitors (I3D [7], R(2+1)D [61], TimeSformer [5], InternVideo2 [66], VideoMamba [36], and V-JEPA2 [2]), and replaced their original heads with a regression head or classification head for ventricular volume estimation and VM severity prediction, respectively. TABLE IV: Ablation study on our VIFBA framework. Method & Module Regression Classification MSE RMSE MAE Pearson Spearman Acc Pre Rec F1 AUC Baseline (Reg) 9.9199 3.1496 2.6122 0.9033 0.7648 / / / / / Baseline (Cls) / / / / / 0.7900 0.7387 0.8256 0.7568 0.9186 Baseline (Reg+Cls) 6.3064 2.5112 2.0827 0.9707 0.9304 0.8200 0.7310 0.8423 0.7569 0.9495 TLP CMTA VLM MSE RMSE MAE Pearson Spearman Acc Pre Rec F1 AUC ✓ ✗ ✗ 2.1778 1.4757 1.0739 0.9882 0.9179 0.8800 0.8002 0.8890 0.8302 0.9601 ✗ ✓ ✗ 2.3407 1.5299 1.2608 0.9759 0.9048 0.8700 0.7799 0.8451 0.8037 0.9498 ✗ ✗ ✓ 3.8378 1.9590 1.5638 0.9652 0.8643 0.8200 0.7344 0.8211 0.7639 0.9513 ✓ ✓ ✗ 1.3936 1.1805 0.9140 0.9844 0.9094 0.8900 0.8717 0.9070 0.8863 0.9602 ✓ ✗ ✓ 1.7897 1.3378 1.0914 0.9837 0.9630 0.9100 0.9144 0.8706 0.8814 0.9648 ✗ ✓ ✓ 1.2619 1.1223 0.9260 0.9875 0.9326 0.9000 0.8223 0.8995 0.8474 0.9612 VIFBA 0.7507 0.8664 0.5909 0.9907 0.9730 0.9400 0.9357 0.9068 0.9130 0.9793 TABLE V: Comparison of VIFBA with different foundation model-based encoders. VIFBA Regression Classification MSE RMSE MAE Pearson Spearman Acc Pre Rec F1 AUC PMC-CLIP 1.1804 1.0864 0.8970 0.9893 0.9239 0.8900 0.8398 0.8917 0.8537 0.9296 BiomedCLIP 1.0475 1.0234 0.8300 0.9915 0.9466 0.9100 0.8652 0.9122 0.8849 0.9572 USFM 0.8702 0.9329 0.7904 0.9965 0.9704 0.9400 0.8878 0.9300 0.9053 0.9717 EchoCare 0.9217 0.9600 0.8049 0.9905 0.9400 0.9300 0.8793 0.8523 0.8580 0.9399 Ultrasound-CLIP 0.9013 0.9494 0.7994 0.9942 0.9376 0.9300 0.8747 0.9550 0.9048 0.9717 FetalCLIP 0.7507 0.8664 0.5909 0.9907 0.9730 0.9400 0.9357 0.9068 0.9130 0.9793 For volume regression, all compared methods showed limited performance, with MSE values ranging from 10.8888 to 13.0733. Baseline (Reg) with FetalCLIP backbone reduced the MSE to 9.9199 and achieved a Pearson correlation of 0.9033. Compared with them, VIFBA achieved substantially better performance, reducing the MSE, RMSE, and MAE to 0.7507, 0.8664, and 0.5909, respectively. It also achieved Pearson and Spearman correlation coefficients of 0.9907 and 0.9730, respectively, indicating strong agreement with the MRI-derived ventricular volume measurements. We fairly removed the classification branch and found that VIFBA w/o Cls outperforms all compared models and the regression baseline. For the VM severity classification task, the selected video models achieved Accs between 0.7000 and 0.7700, with F1 ranging from 0.6450 to 0.7208. Specifically, V-JEPA2 obtained the best classification performance among them, with an Acc of 0.7700 and an F1 of 0.7208. Baseline (Cls) achieved higher performance, with an Acc of 0.7900 and a F1 score of 0.7568. In contrast, VIFBA achieved the best overall classification results, with an Acc of 0.9400, Pre of 0.9357, F1 of 0.9130, and AUC of 0.9793. For comprehensive comparison, VIFBA w/o Reg also outperformed the classification baseline, achieving an Acc of 0.8600 and an F1 of 0.8173. Moreover, statistical analysis based on multi-seed experiments is presented in Figure 5, further confirming the significant performance improvements achieved by VIFBA. These results indicate that general model designs for video analysis are insufficient for fine-grained fetal brain ultrasound predictions. In contrast, VIFBA learns more effective video representations through foundation model-based encoding, latent predictive learning, and MRI-informed supervision, while the gains over single-task variants suggest the complementarity between volume estimation and VM classification. We also provide the video- and case-level regression performance on our test set, as shown in Figure 6. For each group, cases were sorted by the ground-truth lateral ventricular volume. Overall, the model predictions exhibit good agreement with the ground truth across normal brain, non-VM abnormal brain, and different VM severity groups. It can be seen that video-level predictions may vary due to differences in imaging views, video quality, and incomplete anatomical observations. By aggregating multi-video features at the case level, the model integrates complementary anatomical information and produces more stable volume estimates that better follow the ground-truth trend across severity levels. Moreover, as reported in Table I, our case-level feature aggregation strategy consistently outperforms video-level output aggregation across all diagnostic subgroups and the entire test set. Fig. 7: Pairwise statistical results among ablation variants on (a) regression and (b) classification metrics. Each cell reports the p-value for testing whether the row method significantly outperforms the column method. Boxes with black borders indicate statistically significant improvements (p<0.05p<0.05). M1: Baseline (Reg)/(Cls), M2: Baseline (Reg+Cls), M3: TLP, M4: CMTA, M5: VLM, M6: TLP+CMTA, M7: TLP+VLM, M8: CMTA+VLM, M9: VIFBA. IV-D Ablation Study We further conducted ablation experiments to investigate the contribution of each component, including tube latent prediction (TLP), cross-modal training alignment (CMTA), and abnormality verification & alerting (VLM). As shown in Table IV, each component consistently improved the baseline. When used individually, the TLP objective achieved an MSE of 2.1778, an MAE of 1.0739, and an Acc of 0.8800, suggesting that it benefits video representation learning. CMTA also improved both regression and classification performance, achieving an MSE of 2.3407, an MAE of 1.2608, and an Acc of 0.8700, suggesting that MRI-derived representations provide complementary supervision during training. The VLM module alone also enhanced model performance, especially for regression (MSE reduction: 2.4686; MAE reduction: 0.5189). Combining different modules further improved performance. For example, the combination of TLP and CMTA reduced the MSE and MAE to 1.3936 and 0.9140, respectively, and improved the F1 to 0.8863. Connecting TLP with VLM achieved a higher Acc of 0.9100, while maintaining strong regression metrics with an MSE of 1.7897 and an MAE of 1.0914. The full VIFBA achieved the best overall performance, with an MSE of 0.7507, an MAE of 0.5909, a Pearson correlation of 0.9907, an Acc of 0.9400, and an F1 of 0.9130. We further repeated the experiments using 10 different random seeds and report the corresponding statistical p-values in Figure 7. The statistical analysis shows that the proposed modules yield statistically significant improvements in most comparisons, although the gains from individual modules are relatively limited in some settings and do not consistently reach significance across all metrics. Notably, the numerous significant improvements of VIFBA over different ablation variants, as highlighted in the last row of each matrix, further demonstrate the complementarity of the three components. Fig. 8: Performance variation with different numbers of selected videos per case. Full denotes using all available videos for case-level feature aggregation. IV-E Impact of Different Numbers of Videos To investigate the impact of the number of ultrasound videos used for case-level feature aggregation, we conducted an ablation study by varying the number of selected videos per case, as shown in Figure 8. Since each case may contain a different number of brain ultrasound videos, we randomly sampled a predefined number (e.g., 1–5) from each case. If the available number was smaller than the specified number, all available videos were used. To reduce the influence of random video selection, the sampling process was repeated five times for each setting. The curves represent the mean performance, the shaded regions indicate the minimum-to-maximum range, and the error bars denote the standard deviation. It can be seen that increasing the number of selected videos consistently improves the regression and classification performance. Most improvements are achieved when increasing the number of videos from 1 to 3, after which the performance gradually converges. Meanwhile, the shaded regions become narrower with more available videos, indicating reduced prediction variability and improved stability of case-level feature aggregation. These overall results suggest that aggregating multiple videos enables our VIFBA to capture complementary anatomical information from different scanning views and provides more robust case-level predictions. TABLE VI: Comparison of VIFBA with baselines and different VLMs on the multi-label classification task. S1: one-stage classification over 10 categories (normal brain, mild/moderate/severe VM, and A1–A6); S2: two-stage classification that first identifies abnormal brains and then determines their specific diseases. Method Pre Rec F1 HL EMR VIFBA-CLS-S1 0.4243 0.6299 0.4575 0.1250 0.3200 VIFBA-CLS-S2 0.4762 0.7036 0.5084 0.1110 0.4100 VIFBA Claude-Haiku-4.5 0.6129 0.7836 0.6673 0.0630 0.6900 Gemini-3.1-Pro 0.6517 0.8027 0.7053 0.0530 0.7000 Qwen3.5-Plus 0.6485 0.7847 0.6861 0.0590 0.7100 GPT-5.4 0.6333 0.8346 0.6941 0.0610 0.7100 Qwen3.6-Plus 0.6866 0.8384 0.7332 0.0530 0.7200 Qwen3.7-Plus 0.6955 0.8595 0.7478 0.0500 0.7300 GPT-5.5 0.7090 0.8999 0.7617 0.0450 0.7300 GPT-5.6 Terra 0.7130 0.8775 0.7687 0.0450 0.7400 GPT-5.6 Sol 0.7207 0.8933 0.7764 0.0440 0.7500 Fig. 9: Multi-seed performance comparison of VIFBA with different foundation encoders. E1–E6 represent PMC-CLIP, BiomedCLIP, USFM, EchoCare, Ultrasound-CLIP, and FetalCLIP, respectively. Statistical significance is denoted as * (p<0.05p<0.05), ** (p<0.01p<0.01), *** (p<0.001p<0.001), and ns (not significant). Fig. 10: Visualization of selected test cases. For each case, the figure shows example frames from the input ultrasound videos, an axial slice from the reconstructed MRI volume, and ventricular segmentation results obtained from the MRI based on BOUNTI and FeTA segmentors. Note that, to facilitate visualization of the ventricular shape, we have arbitrarily rotated the ventricular segmentation masks for each case. In addition, we show the ground-truth (GT, based on the segmentations/reports and checked by an expert) and VIFBA-predicted (pred) ventricular volumes and diagnostic labels. Fig. 11: Two typical examples showing the VM verification and multi-abnormality alerting process. IV-F Comparison with Different Foundation Model Encoders In Table V, we evaluated the influence of different foundation model-based encoders within the proposed VIFBA framework. Results show that all encoders achieved strong performance compared with conventional deep models (see Tables I and I), confirming the effectiveness of foundation representations for fetal brain ultrasound video analysis. Compared with general biomedical encoders (i.e., PMC-CLIP [39] and BiomedCLIP [77]), ultrasound-specific foundation models (i.e., USFM [31], EchoCare [76], Ultrasound-CLIP [32], FetalCLIP [45]) generally achieved better performance. Among all evaluated encoders, FetalCLIP achieved the best error-based regression performance, with the lowest MSE, RMSE and MAE, while maintaining strong correlation and classification performance. Figure 9 further visualizes the performance variations across 10 independent runs for different encoder variants. Although some comparisons do not reach statistical significance on specific metrics, significant improvements are observed in the majority of tests (34/60). This suggests that fetal ultrasound-specific vision-language pretraining provides suitable representations for ventricular volume estimation and VM severity assessment. The overall results also demonstrate the good extensibility of VIFBA, as it can be flexibly integrated with various foundation encoders while consistently achieving competitive performance. IV-G VM Verification and Multi-Abnormality Alerting To further evaluate VIFBA’s ability to verify VMs and identify non-VM brain abnormalities, we conducted experiments on the multi-abnormality assessment task in Table VI. It can be observed that directly extending the classification branch to predict multiple abnormalities yielded limited performance, with the one-stage VIFBA-CLS-S1 achieving an EMR of 0.3200. Besides, by first identifying abnormal brains and then predicting the specific abnormality categories, the two-stage VIFBA-CLS-S2 improved the EMR to 0.4100. These results indicate that explicitly decomposing the task can partially reduce the difficulty of multi-abnormality recognition. Nevertheless, conventional supervised classification remains limited in this setting, as fetal brain abnormalities involve diverse categories, may coexist within the same fetus, exhibit substantially imbalanced category distributions with limited samples for rare findings, and affect different anatomical regions across the brain beyond the lateral ventricles. In contrast, the retrieval-augmented VLM verification strategy substantially improved the overall performance. Compared with traditional baselines, all VLMs achieved better evaluation metrics. Among them, GPT-5.6 SolSol achieved the best overall performance, with a Pre of 0.7207, a Rec of 0.8933, an F1 score of 0.7764, an HL of 0.0440, and an EMR of 0.7500. GPT-5.6 TerraTerra also showed strong performance, achieving an F1 of 0.7687 and an EMR of 0.7400. These results suggest that VLMs can effectively leverage retrieved visual references and textual descriptions to provide more reliable abnormality-aware screening than task-specific classifiers alone. Moreover, Figure 10 shows the qualitative results of ventricular volume regression and multi-abnormality alerting tasks. The predicted ventricular volumes closely agree with the ground-truth measurements, and VIFBA successfully identifies potential fetal brain abnormalities in most cases. Figure 11 provides two representative examples illustrating the detailed VLM workflow, including the retrieved reference cases, current query case, verification & alerting outputs, and the final results. Each retrieved reference case is accompanied by structured prompts including the [lateral ventricular volume] and [diagnosis category] information. Note that for diagnosis category, we also provide the corresponding description ([******][******], under the structured prompt) in the original report for reference evidence. Then, given the query case with its predicted lateral ventricular volume and diagnostic category as prompt, our VIFBA will give verification and abnormality-alerting suggestions. Finally, the predicted volume and category will be retained or updated accordingly, and the potential abnormal type with alert score will also be provided. Overall, these results demonstrate that retrieval-augmented VLM verification is particularly suitable for multi-abnormality alerting, where diagnostic evidence is sparse, heterogeneous, and often distributed across multiple anatomical regions. By incorporating similar historical cases and clinically relevant textual descriptions, the proposed strategy provides a flexible training-free extension to VIFBA for detecting potential non-VM brain abnormalities beyond VM severity classification. V Discussion and Conclusion In this study, we proposed VIFBA, a unified framework for estimating fetal lateral ventricular volume and assessing brain abnormalities directly from ultrasound videos. It integrates an ultrasound foundation model with temporal modeling and JEPA-inspired latent predictive learning to enhance video representations. In addition, MRI-guided cross-modal alignment transfers structural information from MRI to ultrasound representations during training, while requiring only ultrasound at inference. Finally, a retrieval-augmented, training-free VLM is introduced to verify uncertain predictions and provide auxiliary alerts for potential non-VM brain abnormalities. Experiments on our large dataset demonstrated that VIFBA substantially outperformed existing representative methods. Conventionally, ventricular volumetric assessment relies on MRI, whose high cost and limited availability restrict its routine clinical use, particularly in underserved healthcare settings. To the best of our knowledge, VIFBA is the first exploration to reliably estimate lateral ventricular volume from routinely acquired ultrasound videos, providing a more affordable and widely deployable solution for quantitative assessment. In addition, our experiments suggest that jointly learning ventricular volume estimation can further improve VM severity classification performance. This indicates that the regression and classification objectives are not independent, but instead encourage the model to learn shared representations that are transferable across tasks. Beyond VM assessment, this study further explores the potential of retrieval-augmented VLMs for multi-disease alerting in fetal brain ultrasound. In clinical practice, except for isolated VM, other fetal brain abnormalities are diverse and often coexist. However, conventional deep learning models are typically developed for predefined diseases and rely on sufficient disease-specific annotations, limiting their ability to cover rare or underrepresented conditions. By incorporating retrieved historical cases and their associated diagnostic descriptions as contextual references, our VIFBA enables the plug-and-play integration of off-the-shelf VLMs to compare the query case against multiple abnormal patterns and simultaneously provide alerts for potential non-VM brain abnormalities, without modifying the VLM architecture or requiring additional disease-specific training. We believe that this strategy provides a flexible pathway for extending disease-specific predictive models toward broader multi-disease screening and more comprehensive fetal brain assessment. In future work, we will further validate VIFBA on larger multi-center cohorts to assess its generalizability across different ultrasound systems and clinical settings. In addition, we plan to extend the framework toward more comprehensive fetal brain assessment by incorporating additional and fine-grained abnormality categories. Finally, integrating more advanced multimodal foundation models and leveraging larger-scale paired ultrasound–MRI datasets may further enhance the robustness and clinical applicability of the proposed framework. References [1] M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. Rabbat, et al. (2023) Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 15619–15629. Cited by: §I-B. [2] M. Assran, A. Bardes, D. Fan, Q. Garrido, R. Howes, M. Muckley, A. Rizvi, C. Roberts, K. Sinha, A. Zholus, et al. (2025) V-jepa 2: self-supervised video models enable understanding, prediction and planning. arXiv preprint arXiv:2506.09985. Cited by: §I-B, §IV-C. [3] A. Bardes, Q. Garrido, J. Ponce, X. Chen, M. Rabbat, Y. LeCun, M. Assran, and N. Ballas (2023) V-jepa: latent video prediction for visual representation learning. Cited by: §I-B. [4] C. F. Baumgartner, K. Kamnitsas, J. Matthew, T. P. Fletcher, S. Smith, et al. (2017) SonoNet: real-time detection and localisation of fetal standard scan planes in freehand ultrasound. IEEE transactions on medical imaging 36 (11), p. 2204–2215. Cited by: §I-A, §I-A. [5] G. Bertasius, H. Wang, and L. Torresani (2021) Is space-time attention all you need for video understanding?. In Icml, Vol. 2, p. 4. Cited by: §IV-C. [6] M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin (2021) Emerging properties in self-supervised vision transformers. In 2021 IEEE/CVF international conference on computer vision (ICCV), p. 9630–9640. Cited by: §I-B. [7] J. Carreira and A. Zisserman (2017) Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, p. 6299–6308. Cited by: §IV-C. [8] C. Chen, Y. Huang, X. Yang, X. Hu, Y. Zhang, T. Tan, W. Xue, and D. Ni (2025) Enhancing fetal ultrasound image quality assessment with multi-scale fusion and clustering-based optimization. Biomedical Signal Processing and Control 102, p. 107249. Cited by: §I-A. [9] R. Chen, Y. Huang, H. Zhang, C. Tian, S. Ji, Y. Zhang, and D. Ni (2026) FrameONE: hierarchical motion modeling for universal multi-view echocardiographic keyframe detection. arXiv preprint arXiv:2607.00748. Cited by: §I-B. [10] X. Chen, M. He, T. Dan, N. Wang, M. Lin, L. Zhang, J. Xian, H. Cai, and H. Xie (2020) Automatic measurements of fetal lateral ventricles in 2d ultrasound images using deep learning. Frontiers in neurology 11, p. 526. Cited by: §I-A. [11] M. Christensen, M. Vukadinovic, N. Yuan, and D. Ouyang (2024) Vision–language foundation model for echocardiogram interpretation. Nature Medicine 30 (5), p. 1481–1488. Cited by: §I-B. [12] T. Ciceri, L. Squarcina, A. Giubergia, et al. (2023) Review on deep learning fetal brain segmentation from magnetic resonance images. Artificial Intelligence in Medicine 143, p. 102608. Cited by: §I-A. [13] D. Coronado-Gutierrez, E. Eixarch, E. Monterde, I. Matas, P. Traversi, E. Gratacos, E. Bonet-Carne, and X. P. Burgos-Artizzu (2023) Automatic deep learning-based pipeline for automatic delineation and measurement of fetal brain structures in routine mid-trimester ultrasound images. Fetal Diagnosis and Therapy 50 (6), p. 480–490. Cited by: §I-A. [14] H. Dou, Y. Huang, Y. Huang, X. Yang, C. Zhen, Y. Zhang, Y. Xiong, W. Huang, and D. Ni (2025) Standard plane localization using denoising diffusion model with multi-scale guidance. Computer Methods and Programs in Biomedicine 261, p. 108619. Cited by: §I-A. [15] Y. Duan, T. Tan, Z. Zhu, Y. Huang, Y. Zhang, R. Gao, P. C. Pang, X. Gao, G. Tao, X. Cong, et al. (2025) FetalFlex: anatomy-guided diffusion model for flexible control on fetal ultrasound image synthesis. Medical Image Analysis, p. 103725. Cited by: §I-A. [16] A. D. Edwards, D. Rueckert, S. M. Smith, S. Abo Seada, A. Alansary, J. Almalbis, J. Allsop, J. Andersson, T. Arichi, S. Arulkumaran, et al. (2022) The developing human connectome project neonatal data release. Frontiers in neuroscience 16, p. 886772. Cited by: Fig. 4. [17] M. Firenze, S. I. Young, C. J. Wang, H. J. Yun, E. Adalsteinsson, K. Im, P. E. Grant, and P. Golland (2026) Fast multi-stack slice-to-volume reconstruction via multi-scale unrolled optimization. arXiv preprint arXiv:2601.07519. Cited by: §I-A. [18] N. S. Fox, A. Monteagudo, J. A. Kuller, S. Craigo, M. E. Norton, S. for Maternal-Fetal Medicine (SMFM, et al. (2018) Mild fetal ventriculomegaly: diagnosis, evaluation, and management. American journal of obstetrics and gynecology 219 (1), p. B2–B9. Cited by: §I. [19] P. D. Griffiths, M. Bradburn, M. J. Campbell, C. L. Cooper, R. Graham, D. Jarvis, M. D. Kilby, G. Mason, C. Mooney, S. C. Robson, et al. (2017) Use of mri in the diagnosis of fetal brain abnormalities in utero (meridian): a multicentre, prospective cohort study. The Lancet 389 (10068), p. 538–546. Cited by: §I-A. [20] J. Guo, J. Lin, G. Tan, Y. Lu, Z. Gao, S. Li, and K. Li (2024) Unsupervised ultrasound image quality assessment with score consistency and relativity co-learning. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 734–743. Cited by: §I-A. [21] X. Guo, M. Alsharid, H. Zhao, Y. Wang, J. Lander, A. T. Papageorghiou, and J. A. Noble (2026) A visually grounded language model for fetal ultrasound understanding. Nature Biomedical Engineering, p. 1–17. Cited by: §I-B. [22] D. Ha and J. Schmidhuber (2018) World models. arXiv preprint arXiv:1803.10122 2 (3), p. 440. Cited by: §I-B. [23] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al. (2022) Lora: low-rank adaptation of large language models.. Iclr 1 (2), p. 3. Cited by: §IV-B. [24] J. Hu, W. Xue, J. Cheng, Y. Liu, W. Zhuo, and D. Ni (2025) EchoONE: segmenting multiple echocardiography planes in one model. In Proceedings of the computer vision and pattern recognition conference, p. 5207–5216. Cited by: §I-B. [25] R. Huang, A. Namburete, and A. Noble (2018) Learning to segment key clinical anatomical structures in fetal neurosonography informed by a region-based descriptor. Journal of Medical Imaging 5 (1), p. 014007–014007. Cited by: §I-A. [26] S. Huang, Z. Lian, D. Jia, K. Sun, X. Li, J. Liu, Y. Wang, C. Jiang, F. Zhu, Z. Ding, et al. (2026) BrainSeg: a generalized framework for comprehensive multimodal brain tissue segmentation, parcellation, and lesion labeling. npj Digital Medicine. Cited by: §I-A. [27] Y. Huang, X. Yang, X. Huang, J. Liang, X. Zhou, C. Chen, H. Dou, X. Hu, Y. Cao, and D. Ni (2022) Online reflective learning for robust medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 652–662. Cited by: §I-A. [28] Y. Huang, X. Yang, X. Huang, X. Zhou, H. Chi, H. Dou, X. Hu, J. Wang, X. Deng, and D. Ni (2023) Fourier test-time adaptation with multi-level consistency for robust classification. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 221–231. Cited by: §I-A. [29] Y. Huang, X. Yang, L. Liu, H. Zhou, A. Chang, X. Zhou, R. Chen, J. Yu, J. Chen, C. Chen, et al. (2024) Segment anything model for medical images?. Medical Image Analysis 92, p. 103061. Cited by: §I-B. [30] Y. Huang, X. Yang, H. Zhou, Y. Cao, H. Dou, F. Dong, and D. Ni (2024) Robust box prompt based sam for medical image segmentation. In International Workshop on Machine Learning in Medical Imaging, p. 1–11. Cited by: §I-B. [31] J. Jiao, J. Zhou, X. Li, M. Xia, Y. Huang, L. Huang, N. Wang, X. Zhang, S. Zhou, Y. Wang, et al. (2024) Usfm: a universal ultrasound foundation model generalized to tasks and organs towards label efficient image analysis. Medical image analysis 96, p. 103202. Cited by: §I-B, §IV-F. [32] J. Jin, H. Chai, X. Huang, X. Guo, Z. Zheng, Z. Zhou, et al. (2026) Ultrasound-clip: semantic-aware contrastive pre-training for ultrasound image-text understanding. arXiv preprint arXiv:2604.01749. Cited by: §I-B, §IV-F. [33] R. Jin, G. Huang, X. Shen, Q. Zhang, Y. S. Tan, and X. Li (2025) See-in-pairs: reference image-guided comparative vision-language models for medical diagnosis. arXiv preprint arXiv:2506.18140. Cited by: §I-C. [34] B. Klaudel and A. Obuchowski (2025) RAD-srac: simple retrieval augmented classification for radiology. In Workshop on Large Language Models and Generative AI for Health at AAAI 2025, Cited by: §I-C. [35] M. Kuklisova-Murgasova, G. Quaghebeur, M. A. Rutherford, J. V. Hajnal, and J. A. Schnabel (2012) Reconstruction of fetal brain mri with intensity matching and complete outlier removal. Medical image analysis 16 (8), p. 1550–1564. Cited by: §I. [36] K. Li, X. Li, Y. Wang, Y. He, Y. Wang, L. Wang, and Y. Qiao (2024) Videomamba: state space model for efficient video understanding. In European conference on computer vision, p. 237–255. Cited by: §IV-C. [37] H. Liang, Y. Huang, X. Zhu, Y. Zhang, X. Deng, X. Gao, G. Tao, Y. Zhang, and D. Ni (2026) Prototype memory-guided training-free anomaly classification and localization in prenatal ultrasound. arXiv preprint arXiv:2607.00744. Cited by: §I-A. [38] H. Liang, Y. Zhang, X. Zhu, Y. Huang, X. Du, S. Liang, J. Xu, Y. Zhang, C. Sheng, Y. Liu, et al. (2026) FAA-net: fetal abdominal anomaly diagnosis in prenatal ultrasound via llm-enhanced multi-instance learning. Medical Image Analysis, p. 104201. Cited by: §I-C. [39] W. Lin, Z. Zhao, X. Zhang, C. Wu, Y. Zhang, Y. Wang, and W. Xie (2023) Pmc-clip: contrastive language-image pre-training using biomedical documents. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 525–536. Cited by: §IV-F. [40] S. Liu, Q. Ying, S. He, X. Yang, D. Ni, and R. Huang (2023) Hierarchical agent-based reinforcement learning framework for automated quality assessment of fetal ultrasound video. In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), p. 1–5. Cited by: §I-A. [41] Z. Liu, X. Yang, R. Gao, S. Liu, H. Dou, S. He, Y. Huang, Y. Huang, H. Luo, Y. Zhang, et al. (2020) Remove appearance shift for ultrasound image segmentation via fast and universal style transfer. In 2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI), p. 1824–1828. Cited by: §I-A. [42] J. Ma, Y. He, F. Li, L. Han, C. You, and B. Wang (2024) Segment anything in medical images. Nature communications 15 (1), p. 654. Cited by: §I-B. [43] J. Ma, Z. Yang, S. Kim, B. Chen, M. Baharoon, A. Fallahpour, R. Asakereh, H. Lyu, and B. Wang (2025) Medsam2: segment anything in 3d medical images and videos. arXiv preprint arXiv:2504.03600. Cited by: §I-B. [44] L. Ma, K. Li, Z. Xiao, G. Tan, H. Wen, and S. Li (2026) Structure-aware fine-grained instance segmentation for fetal brain ultrasound images. Neurocomputing 677, p. 133093. Cited by: §I-A. [45] F. Maani, N. Saeed, T. J. Saleem, Z. Farooq, H. Alasmawi, W. Diehl, A. Mohammad, G. Waring, S. Valappil, L. Bricker, et al. (2026) Fetalclip: a visual-language foundation model for fetal ultrasound image analysis. npj Digital Medicine. Cited by: §I-B, §I-A, §IV-F. [46] A. Makropoulos, I. S. Gousias, C. Ledig, P. Aljabar, A. Serag, J. V. Hajnal, A. D. Edwards, S. J. Counsell, and D. Rueckert (2014) Automatic whole brain mri segmentation of the developing neonatal brain. IEEE transactions on medical imaging 33 (9), p. 1818–1831. Cited by: Fig. 4. [47] A. Makropoulos, E. C. Robinson, A. Schuh, R. Wright, S. Fitzgibbon, J. Bozek, S. J. Counsell, J. Steinweg, K. Vecchiato, J. Passerat-Palmbach, et al. (2018) The developing human connectome project: a minimal processing pipeline for neonatal cortical surface reconstruction. Neuroimage 173, p. 88–112. Cited by: Fig. 4. [48] G. Malinger, D. Paladini, K. K. Haratz, A. Monteagudo, G. Pilu, and I. Timor-Tritsch (2020) ISUOG practice guidelines (updated): sonographic examination of the fetal central nervous system. part 1: performance of screening examination and indications for targeted neurosonography.. Ultrasound in obstetrics & gynecology: the official journal of the International Society of Ultrasound in Obstetrics and Gynecology 56 (3), p. 476–484. Cited by: §I, §I-A. [49] A. Munim, A. Fallahpour, T. Szasz, A. Attarpour, R. Jiang, B. Sooriyakanthan, M. Sooriyakanthan, H. Whitney, J. Slivnick, B. Rubin, et al. (2026) EchoJEPA: a latent predictive foundation model for echocardiography. arXiv preprint arXiv:2602.02603. Cited by: §I-B. [50] A. I. Namburete, W. Xie, M. Yaqub, A. Zisserman, and J. A. Noble (2018) Fully-automated alignment of 3d fetal brain ultrasound to a canonical reference space using multi-task learning. Medical image analysis 46, p. 1–14. Cited by: §I-A. [51] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. (2023) Dinov2: learning robust visual features without supervision. arXiv preprint arXiv:2304.07193. Cited by: §I-B. [52] L. Ou, X. Zhu, H. Liang, W. Pan, Y. Huang, Y. Deng, X. Sheng, H. Yin, J. Xiao, X. Zhou, et al. (2026) Foundation model-driven key anatomy frame selection for blind-sweep ultrasound fetal birth weight estimation. arXiv preprint arXiv:2607.00745. Cited by: §I-B. [53] G. Pagani, B. Thilaganathan, and F. Prefumo (2014) Neurodevelopmental outcome in isolated mild fetal ventriculomegaly: systematic review and meta-analysis. Ultrasound in Obstetrics & Gynecology 44 (3), p. 254–260. Cited by: §I. [54] K. Payette, P. de Dumast, H. Kebiri, I. Ezhov, J. C. Paetzold, S. Shit, A. Iqbal, R. Khan, R. Kottke, P. Grehten, et al. (2021) An automatic multi-tissue human fetal brain segmentation benchmark using the fetal tissue annotation dataset. Scientific data 8 (1), p. 167. Cited by: §I-A. [55] F. Pérez-García, H. Sharma, S. Bond-Taylor, K. Bouzid, V. Salvatelli, M. Ilse, S. Bannur, D. C. Castro, A. Schwaighofer, M. P. Lungren, et al. (2025) Exploring scalable medical image encoders beyond text supervision. Nature Machine Intelligence 7 (1), p. 119–130. Cited by: §I-B. [56] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §I-C. [57] D. Scholz, A. C. Erdur, V. Ehm, A. Meyer-Baese, J. C. Peeken, D. Rueckert, and B. Wiestler (2025) Mm-dinov2: adapting foundation models for multi-modal medical image analysis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, p. 320–330. Cited by: §I-B. [58] X. Shu, F. Chang, X. Zhang, C. Shao, and X. Yang (2022) ECAU-net: efficient channel attention u-net for fetal ultrasound cerebellum segmentation. Biomedical Signal Processing and Control 75, p. 103528. Cited by: §I-A. [59] M. Sinclair, C. F. Baumgartner, J. Matthew, W. Bai, J. C. Martinez, Y. Li, S. Smith, C. L. Knight, B. Kainz, J. Hajnal, et al. (2018) Human-level performance on automatic head biometrics in fetal ultrasound using fully convolutional neural networks. In 2018 40th annual international conference of the IEEE engineering in medicine and biology society (EMBC), p. 714–717. Cited by: §I-A. [60] X. Song, X. Xu, and P. Yan (2024) General purpose image encoder dinov2 for medical image registration. arXiv preprint arXiv:2402.15687. Cited by: §I-B. [61] D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri (2018) A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, p. 6450–6459. Cited by: §IV-C. [62] A. U. Uus, V. Kyriakopoulou, A. Makropoulos, A. Fukami-Gartner, D. Cromb, A. Davidson, L. Cordero-Grande, A. N. Price, I. Grigorescu, L. Z. Williams, et al. (2023) BOUNTI: brain volumetry and automated parcellation for 3d fetal mri. bioRxiv. Cited by: §I-A, §I-C, Fig. 4, §IV-A. [63] F. Vahedifard, H. A. Ai, M. P. Supanich, K. K. Marathu, X. Liu, M. Kocak, S. M. Ansari, M. Akyuz, J. O. Adepoju, S. Adler, et al. (2023) Automatic ventriculomegaly detection in fetal brain mri: a step-by-step deep learning model for novel 2d-3d linear measurements. Diagnostics 13 (14), p. 2355. Cited by: §I-A. [64] M. Vukadinovic, I. Chiu, X. Tang, N. Yuan, T. Chen, P. Cheng, D. Li, S. Cheng, B. He, and D. Ouyang (2026) Comprehensive echocardiogram evaluation with view primed vision language ai. Nature 650 (8103), p. 970–977. Cited by: §I-B. [65] S. Wang, Z. Zhao, X. Ouyang, Q. Wang, and D. Shen (2023) Chatcad: interactive computer-aided diagnosis on medical image using large language models. arXiv preprint arXiv:2302.07257. Cited by: §I-C. [66] Y. Wang, K. Li, X. Li, J. Yu, Y. He, G. Chen, B. Pei, R. Zheng, Z. Wang, Y. Shi, et al. (2024) Internvideo2: scaling foundation models for multimodal video understanding. In European conference on computer vision, p. 396–416. Cited by: §IV-C. [67] H. Xie, N. Wang, M. He, L. Zhang, H. Cai, J. Xian, M. Lin, J. Zheng, and Y. Yang (2020) Using deep-learning algorithms to classify fetal brain ultrasound images as normal or abnormal. Ultrasound in Obstetrics & Gynecology 56 (4), p. 579–587. Cited by: §I-A. [68] J. Xu, D. Moyer, B. Gagoski, J. E. Iglesias, P. E. Grant, P. Golland, and E. Adalsteinsson (2023) NeSVoR: implicit neural representation for slice-to-volume reconstruction in mri. IEEE transactions on medical imaging 42 (6), p. 1707–1719. Cited by: §I-A, §IV-A. [69] Z. Yan, T. Han, Y. Huang, L. Liu, H. Zhou, J. Chen, W. Shi, Y. Cao, X. Yang, and D. Ni (2024) A foundation model for general moving object segmentation in medical images. In 2024 IEEE International Symposium on Biomedical Imaging (ISBI), p. 1–5. Cited by: §I-B. [70] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. (2025) Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §IV-A. [71] X. Yang, H. Dou, R. Huang, W. Xue, Y. Huang, J. Qian, Y. Zhang, H. Luo, H. Guo, T. Wang, et al. (2021) Agent with warm start and adaptive dynamic termination for plane localization in 3d ultrasound. IEEE Transactions on Medical Imaging 40 (7), p. 1950–1961. Cited by: §I-A. [72] X. Yang, Y. Huang, R. Huang, H. Dou, R. Li, J. Qian, X. Huang, W. Shi, C. Chen, Y. Zhang, et al. (2021) Searching collaborative agents for multi-plane localization in 3d ultrasound. Medical Image Analysis 72, p. 102119. Cited by: §I-A, §I-A. [73] V. Zalevskyi, T. Sanchez, M. Kaandorp, M. Roulet, D. Fajardo-Rojas, L. Li, J. Hutter, H. B. Li, M. J. Barkovich, H. Ji, et al. (2026) Advances in automated fetal brain mri segmentation and biometry: insights from the feta 2024 challenge. Medical image analysis, p. 103941. Cited by: §I-A, Fig. 4, §IV-A. [74] Q. Zeng, W. Liu, B. Li, R. Didier, P. E. Grant, and D. Karimi (2026) Atlas-assisted segment anything model for fetal brain mri (fetal-sam). arXiv preprint arXiv:2601.15759. Cited by: §I-A. [75] Y. Zeng, P. Tsui, W. Wu, Z. Zhou, and S. Wu (2021) Fetal ultrasound image segmentation for automatic head circumference biometry using deeply supervised attention-gated v-net. Journal of Digital Imaging 34 (1), p. 134–148. Cited by: §I-A. [76] H. Zhang, Y. Wu, M. Zhao, Z. Chen, R. Li, F. Zhu, et al. (2025) A fully open and generalizable foundation model for ultrasound clinical applications. arXiv preprint arXiv:2509.11752. Cited by: §I-B, §IV-F. [77] S. Zhang, Y. Xu, N. Usuyama, H. Xu, J. Bagga, R. Tinn, S. Preston, R. Rao, M. Wei, N. Valluri, et al. (2023) Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv preprint arXiv:2303.00915. Cited by: §I-B, §IV-F. [78] Y. Zhang, Y. Huang, H. Dou, X. Zhu, C. Ling, Z. Yang, L. Liang, J. Li, S. Liang, R. Li, et al. (2026) Artificial intelligence for detecting fetal orofacial clefts and advancing medical education. Nature Communications. Cited by: §I-A. [79] Z. Zhao, S. Wang, J. Gu, Y. Zhu, L. Mei, Z. Zhuang, Z. Cui, Q. Wang, and D. Shen (2024) ChatCAD+: toward a universal and reliable interactive cad using llms. IEEE Transactions on Medical Imaging 43 (11), p. 3755–3766. Cited by: §I-C.