Paper deep dive
US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
Ashwath Radhachandran, Vedrana Ivezić, Shreeram Athreya, Ronit Anilkumar, Corey W. Arnold, William Speier
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 8:09:10 PM
Summary
The paper introduces US-JEPA, a self-supervised framework for medical ultrasound representation learning based on Joint-Embedding Predictive Architectures (JEPAs). It utilizes the Static-teacher Asymmetric Latent Training (SALT) objective, employing a frozen Ultrasound Representation Foundation Model (URFM) as a teacher to provide stable latent targets, thereby decoupling student-teacher optimization. The authors also introduce USrc (Ultrasound Region-Conditioning) to focus on anatomical signals and present UltraBench, a standardized benchmark for evaluating ultrasound foundation models. US-JEPA demonstrates competitive or superior performance against existing baselines in linear probing tasks.
Entities (10)
Relation Signals (7)
US-JEPA → basedon → JEPAs
confidence 95% · US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound
US-JEPA → evaluatedon → UltraBench
confidence 95% · we provide the first rigorous comparison of all publicly available state-of-the-art ultrasound foundation models on UltraBench
US-JEPA → usesobjective → SALT
confidence 95% · US-JEPA, a self-supervised framework that adopts the Static-teacher Asymmetric Latent Training (SALT) objective.
US-JEPA → usesteacher → URFM
confidence 92% · US-JEPA uses a frozen, domain-specific teacher, the Ultrasound Representation Foundation Model (URFM)
US-JEPA → incorporates → USrc
confidence 90% · we introduce USrc as a spatial prior to isolate the anatomical signal.
URFM → derivedfrom → BiomedCLIP
confidence 85% · performing knowledge distillation from BiomedCLIP (Zhang et al., 2025c)
US-JEPA → outperforms → USFM
confidence 80% · US-JEPA achieves performance competitive with or superior to domain-specific and universal vision foundation model baselines.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process. The low signal-to-noise ratio and stochastic speckle patterns hinder standard self-supervised learning methods relying on a pixel-level reconstruction objective. Joint-Embedding Predictive Architectures (JEPAs) address this drawback by predicting masked latent representations rather than raw pixels. However, standard approaches depend on hyperparameter-brittle and computationally expensive online teachers updated via exponential moving average. We propose US-JEPA, a self-supervised framework that adopts the Static-teacher Asymmetric Latent Training (SALT) objective. By using a frozen, domain-specific teacher to provide stable latent targets, US-JEPA decouples student-teacher optimization and pushes the student to expand upon the semantic priors of the teacher. In addition, we provide the first rigorous comparison of all publicly available state-of-the-art ultrasound foundation models on UltraBench, a public dataset benchmark spanning multiple organs and pathological conditions. Under linear probing for diverse classification tasks, US-JEPA achieves performance competitive with or superior to domain-specific and universal vision foundation model baselines. Our results demonstrate that masked latent prediction provides a stable and efficient path toward robust ultrasound representations.
Tags
Links
- Source: https://arxiv.org/abs/2602.19322v1
- Canonical: https://arxiv.org/abs/2602.19322v1
Trouble viewing inline? Open PDF directly →
Full Text
81,022 characters extracted from source content.
Expand or collapse full text
US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Ashwath Radhachandran 1 2 Vedrana Ivezi ́ c 1 3 Shreeram Athreya 1 4 Ronit Anilkumar 1 Corey W. Arnold 1 2 3 4 5 6 William Speier 1 2 3 5 Abstract Ultrasound (US) imaging poses unique chal- lenges for representation learning due to its in- herently noisy acquisition process.The low signal-to-noise ratio and stochastic speckle pat- terns hinder standard self-supervised learning methods relying on a pixel-level reconstruction objective. Joint-Embedding Predictive Archi- tectures (JEPAs) address this drawback by pre- dicting masked latent representations rather than raw pixels. However, standard approaches de- pend on hyperparameter-brittle and computation- ally expensive online teachers updated via ex- ponential moving average.We propose US- JEPA, a self-supervised framework that adopts the Static-teacher Asymmetric Latent Training (SALT) objective. By using a frozen, domain- specific teacher to provide stable latent targets, US-JEPA decouples student–teacher optimization and pushes the student to expand upon the se- mantic priors of the teacher. In addition, we pro- vide the first rigorous comparison of all publicly available state-of-the-art ultrasound foundation models on UltraBench, a public dataset bench- mark spanning multiple organs and pathological conditions. Under linear probing for diverse clas- sification tasks, US-JEPA achieves performance competitive with or superior to domain-specific and universal vision foundation model baselines. Our results demonstrate that masked latent pre- diction provides a stable and efficient path toward robust ultrasound representations. 1 Biomedical AI Research Lab, University of California, Los Angeles, USA 2 Department of Bioengineering, University of Cali- fornia, Los Angeles, USA 3 Medical Informatics Home Area, Uni- versity of California, Los Angeles, USA 4 Department of Elec- trical and Computer Engineering, University of California, Los Angeles, USA 5 Department of Pathology, University of Califor- nia, Los Angeles, USA 6 Department of Radiology, University of California, Los Angeles, USA. Correspondence to: Ashwath Radhachandran<ashwathradha123@g.ucla.edu>, William Speier <speier@ucla.edu>. Preprint. February 24, 2026. 1. Introduction Self-supervised learning (SSL) has revolutionized the ability to leverage vast unlabeled datasets to build robust founda- tion models. In particular, masked image modeling (MIM) has emerged as a dominant approach, where models are tasked to reconstruct masked image patches to learn image semantics (Hondru et al., 2025). Although pixel-level recon- struction objectives have proven effective in natural imaging, their utility partially depends on the assumption that local pixel intensities correlate strongly with underlying struc- tural representations. However, when applied to domains characterized by low signal-to-noise ratio and acquisition artifacts, this assumption begins to break down (Xie et al., 2024). The ultrasound (US) imaging domain is a prime example of this representational bottleneck. US is a critical diagnostic tool in clinical settings, as it is portable, radiation-free, and provides real-time bedside feedback. Yet, image interpre- tation is difficult due to inherent visual graininess (speckle noise), operator dependence, and substantial variability in acquisition protocols and patient anatomy. In a traditional MIM framework, a generative, pixel-reconstruction objec- tive may force the model to use its representational capacity to model these uninformative, acquisition-dependent fea- tures. Since these features, such as blur, acoustic shadow, and pixel-level contrast, vary significantly across clinical environments, a model trained for pixel reconstruction may overfit to these specific noise sources. The resulting sig- nificant gap between manually curated training data and real-world settings raises critical concerns about model ro- bustness in out-of-distribution (OOD) scenarios. Without a pretraining objective that prioritizes semantic invariance, US foundation models may fail when subjected to diverse, yet realistic, ultrasound image corruptions. The challenge of ensuring representational robustness is further compounded by the scarcity of high-quality clinical labels, which are expensive to acquire and require special- ized medical expertise. Consequently, there is a need to train foundation models that can extract robust US feature rep- resentations and achieve high performance on downstream tasks using minimal supervision. Recent efforts to build US foundation models have attempted to navigate these bot- 1 arXiv:2602.19322v1 [cs.CV] 22 Feb 2026 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound tlenecks through domain-specific modifications (Jiao et al., 2024; Megahed et al., 2025; Zhang et al., 2025b). Despite their innovations, these methods ultimately remain anchored to the pixel reconstruction paradigm, potentially underrepre- senting the critical global features that distinguish different anatomical regions and pathologies in US. 1.1. Ultrasound Joint Embedding Predictive Architecture (US-JEPA) To address these challenges, we introduce the Ultrasound Joint-Embedding Predictive Architecture (US-JEPA). Based on Image-based JEPA (I-JEPA), US-JEPA goes beyond the generative pixel-filling paradigm. Instead, it operates en- tirely in a latent embedding space, predicting the representa- tions of masked target regions from a context block within the same image. By performing feature reconstruction rather than raw pixel reconstruction, US-JEPA can focus on learn- ing global anatomical dependencies and tissue textures. The I-JEPA framework (Assran et al., 2023) relies on an online teacher updated via Exponential Moving Average (EMA) to guide a gradient-based student with smooth, yet improving targets. This coupling is computationally ex- pensive and sensitive to hyperparameter selection. A core distinction of our framework is the adoption of the Static- teacher Asymmetric Latent Training (SALT) objective (Li et al., 2025), in which a frozen teacher replaces an EMA- updated teacher. US-JEPA uses a frozen, domain-specific teacher, the Ultrasound Representation Foundation Model (URFM) (Kang et al., 2025), to provide stable targets. URFM previously established the effectiveness of MIM for feature-level reconstruction in ultrasound, performing knowledge distillation from BiomedCLIP (Zhang et al., 2025c), a general-purpose medical vision encoder. URFM’s reconstruction-based objective transfers rich semantic pri- ors to the ultrasound domain, but remains tied to local re- construction objectives as opposed to global representation consistency. By using URFM as a static teacher, US-JEPA takes a step further by leveraging the SALT latent prediction objective to refine URFM’s ultrasound-specific represen- tations. Rather than just reconstructing missing features, US-JEPA learns to align representations across views within an image. This approach encourages the model to learn the internal physics of US imaging, providing an improved latent space for downstream clinical tasks. To ensure the model pretrained with US-JEPA is exposed to a wide range of human physiology, we aggregate one of the largest collections of publicly available US imaging to date. After thoroughly searching public repositories, our pretraining dataset encompasses approximately 4.73 million frames from 49 datasets covering 22 distinct anatomies. 1.2. Standardizing Evaluation: UltraBench A significant challenge in current US foundation model research is the lack of a standardized evaluation protocol. Currently, most studies evaluate on disparate, often private, downstream datasets with non-standardized splits, making rigorous iteration and objective comparison nearly impos- sible. To promote reproducibility and establish a gold stan- dard for future work, our paper adopts and extends Ultra- Bench (Tupper & Gagn ́ e), a publicly available benchmark for US frame-level evaluation. While the original Ultra- Bench provides a strong foundation, we expand its scope to enhance anatomical diversity. Specifically, we include eight classification tasks spanning various organs and pathologies, adding two new tasks for thyroid and breast pathology. Beyond the dataset benchmark itself, we address a critical gap in previous work: the lack of side-by-side comparisons of available models. To the best of our knowledge, this work is the first to perform an exhaustive linear probing evalua- tion across all published, publicly available US foundation models. Previous works compare against a limited set of US-specific baselines or rely solely on full fine-tuning, mak- ing it difficult to isolate the intrinsic quality of the learned representations. By tuning just a linear head for the classifi- cation tasks with a frozen backbone, we provide a standard- ized assessment of the learned latent space. This rigorous benchmarking ensures that improvements from US-JEPA are measured against a standard and relevant set of baselines, providing a clear map of the current state-of-the-art. This work takes a significant step forward in ultrasound foundation modeling with the following contributions: •JEPA-based US Foundation Model: We introduce US-JEPA, the first frame-level US foundation model built on JEPA priniciples. •Label-Efficient Representations: US-JEPA achieves strong linear probing performance with fewer labeled samples than competing baselines. •Robustness to Domain-Specific Image Corruption: Learned representations exhibit increased invariance to ultrasound-specific perturbations in image quality. • Comprehensive Benchmarking: We conduct the most comprehensive linear-probing evaluation to date across publicly available ultrasound foundation models on eight clinical tasks from UltraBench. 2. Related Work 2.1. Universal Vision Foundation Models While early SSL focused on instance discrimination and contrastive objectives (ex. MoCo (He et al., 2020) and SimCLR (Chen et al., 2020)), the field has expanded to self- 2 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound distillation and predictive modeling to capture higher-level semantics. DINO (Caron et al., 2021) demonstrated that student-teacher distillation with Vision Transformers (ViTs) can yield semantically rich features without labels. This paradigm has been scaled by successors, DINOv2 (Oquab et al., 2024) and DINOv3 (Sim ́ eoni et al., 2025), to set new benchmarks in universal representation learning. In parallel, I-JEPA (Assran et al., 2023) has proposed moving beyond methods that predict in pixel or token space. By predicting missing latent representations from a visible context, I-JEPA learns the underlying structural logic of an image rather than superficial textures. Despite these universal models excelling on natural images, their direct application to US is limited by modality-specific artifacts, leading research to investigate specialized anatomical and task-based solutions. 2.2. Task-Specific Ultrasound Foundation Models A significant branch of US research focuses on adapting uni- versal architectures for specific clinical tasks, most notably through the specialization of the Segment Anything Model (SAM) (Kirillov et al., 2023). Works such as SAMUS (Lin et al., 2024) and UltraSAM (Meyer et al., 2025) utilize ar- chitectural modifications or large-scale fine-tuning on large segmentation datasets to enable interactive, zero-shot seg- mentation. Beyond segmentation, SSL has proven useful for US representation learning within specific anatomical systems. In fetal US, UltraDINO (Ambsdorf et al., 2025) outperforms general vision models and USFM by train- ing DINOv2 from scratch on fetal-specific corpora. In the cardiac domain, EchoNet-Dynamic and EchoFM (Ouyang et al., 2020; Kim et al., 2025) use large-scale US video cor- pora to learn temporal cardiac dynamics and apply these representations to clinical cardiac tasks. These successes demonstrate the power of domain-specific SSL; however, they often rely on curated, task-specific datasets, motivating the need for a robust, general-purpose US foundation model. 2.3. General Ultrasound Foundation Models The evolution of US foundation models has been defined by a shift from standard pixel-level reconstruction toward domain-aware representation learning. USFM (Jiao et al., 2024) was the first to define the field by training on more than 2 million images (public and private) using a dual spatial-frequency MIM objective, specifically designed to tackle the low spatial resolution and noise inherent in US. Building on this precedent, USF-MAE (Megahed et al., 2025) demonstrated that pre-training exclusively on public US data (using the 370k-image OpenUS-46 corpus) com- bined with noise-filtering preprocessing, could significantly surpass models initialized on natural images. More recent efforts have integrated generative refinements and structural priors. For instance, D 2 MAE (Kang et al.) uses diffusional deblurring in the MIM framework to ad- dress the low SNR inherent to US. EchoCare (Zhang et al., 2025b) leverages a 4.5 million image dataset and introduces a hierarchical classification objective to teach the model nuanced anatomical relationships. To build on these previous approaches, US-JEPA evolves the feature reconstruction paradigm introduced by URFM. By utilizing this domain-specific teacher and an asymmetric latent objective, US-JEPA is designed to encode the struc- tural logic within US, learning an improved latent space for downstream clinical tasks. 3. Preliminaries 3.1. I-JEPA The I-JEPA (Assran et al., 2023) architecture consists of a context encoderf θ , a target encoder ̄ f θ , and a predictorg φ . Given an input imagex, the model operates on a sequence of non-overlapping patches. During training, we sample a context blockB c and a set ofTtarget blocksB i T i=1 , where each block is a subset of patches. From these blocks, we derive disjoint masksM c andM i to ensure there is no overlap and prevent a non-trivial prediction task. The target encoder processes the full imagexto produce patch-level representationss y = ̄ f θ (x). The targets for the loss function are defined as the subset of these representa- tions corresponding to thei-th target mask,y i = s y [M i ]. Simultaneously, the context encoder processes only the visi- ble image patches defined byM c to produce context embed- dingsc = f θ (x[M c ]). The predictorg φ then utilizes these context embeddings, conditioned on mask tokens,p i , which indicate the target-block locations. The mask tokens are pa- rameterized by a learnable vector with an added positional embedding: ˆ y i = g φ (c,p i )(1) The training objective is to minimize the SmoothL 1 dis- tance between the predicted ˆ y i and actualy i target rep- resentations. To prevent representational collapse, I-JEPA employs an asymmetric update rule where the target encoder parameters ̄ θare an EMA of the context encoder parameters θ. 3.2. Static-teacher Asymmetric Latent Training (SALT) The SALT paradigm (Li et al., 2025) helps modulate the parameter selection in the JEPA objective by decoupling the student and teacher optimization. SALT demonstrates that a high-performing student can emerge by predicting the latent space of a static teacher, provided the teacher possesses reasonably sufficient semantic priors. Here, the target encoderh ψ (ex. an existing foundation 3 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound URFM ... ... ... Context tokens Mask tokens Student Predictor Figure 1. USrc-JEPA framework. Here we show the model training framework with USrc. URFM is the frozen teacher that extracts target embeddings. The student and predictor are jointly optimized withL US−JEP A to align with the target. model or encoder pretrained via a different objective such as pixel reconstruction) is frozen. The target embeddings y i become static conditioned on the input image. Gradi- ents propagate solely through the context encoderf θ and predictor g φ : L SALT = 1 T T X i=1 ∥g φ (f θ (x[M c ]),p i )− h ψ (x)[M i ]∥ 1 (2) This approach eliminates the need for the EMA update, thus stabilizing the training dynamics and reducing computa- tional overhead. 4. US-JEPA US-JEPA (Figure 1) relies on the masked latent objective from I-JEPA to learn the spatial semantics within ultrasound imaging. By forcing the student network to predict missing anatomical structures in a latent space with the guidance of a domain-specific teacher, the model is encouraged to develop a robust feature representation of tissue textures and organ morphology that is invariant to pixel-level noise. 4.1. Self-Distillation via SALT While the SALT framework has shown that strong students can emerge even from sub-optimal teachers (Li et al., 2025), our goal was to leverage the most robust US representations to establish a new state-of-the-art. To identify the opti- mal target provider, we performed an extensive benchmark- ing of existing US foundation models across UltraBench. Our preliminary results (detailed in Table 2) demonstrate that URFM consistently outperformed other baselines. We select URFM to serve as our frozen teacher and investi- gate whether the masked latent objective of US-JEPA can achieve better downstream performance, label efficiency in few-shot linear probing and improved stability against domain-specific corruptions. In our framework, the student context encoder and predictor are optimized to minimize the SmoothL 1 distance between the predicted target embeddings and the URFM target em- beddings. 4.2. Ultrasound Region-Conditioning (USrc) A significant challenge in training on diverse, public ultra- sound datasets is the presence of non-anatomical artifacts. Ultrasound frames typically contain peripheral noise, in- cluding transducer metadata, intensity scales, patient infor- mation, and large black borders. Previous masked imaging approaches rely on raw frames which is suboptimal since it would allow non-ultrasound content to be sampled. We hypothesize that uniform random masking in I-JEPA might inadvertently task the model with predicting these meaning- less regions, wasting representational capacity. To test this, we introduce USrc as a spatial prior to isolate the anatomical signal. Letx ∈R H×W be the input ultrasound image andR ∈ 0, 1 H×W be a binary region mask, whereR ij = 1de- notes valid ultrasound signal andR ij = 0denotes back- ground. We tile the image into non-overlapping patches and define the set of valid patchesP valid as those containing at least one pixel within R. Target Sampling: Targets correspond to representations of image blocks. We feedxthrough the frozen target encoder h ψ to obtain patch-level representationss y = h ψ (x). We sampleTcandidate blocksB i T i=1 based on scale and aspect ratio constraints. We employ a rejection sampling strategy where a candidate block is accepted only if its intersection withP valid exceeds a threshold τ : M i = B i ∩P valid ,subject to|M i |≥ τ(3) The final targets are the representations y i = s y [M i ]. Context Sampling: We sample a context blockB c with specific scale and aspect ratio constraints. To ensure a non- trivial task, we remove any regions overlapping with the target blocks and apply our overlap constraint: M c = B c ∩P valid ,subject to|M c |≥ τ(4) The masked contextx[M c ]is processed by the context en- coder f θ to obtain c = f θ (x[M c ]). 4 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Optimization: The training objective is minimized strictly over these anatomical intersections. The predictor aims to minimize the SmoothL 1 distance between the predicted student embeddings and the frozen teacher embeddings: L US-JEPA = 1 T T X i=1 Smooth L1 (g φ (c,p i ), y i )(5) By constraining both input and output toP valid , US-JEPA with USrc (USrc-JEPA) engineers the inputs to only model tissue texture and anatomical structure. 4.3. Training Pipeline To facilitate standardization across our aggregated corpus of US frames, we implement a standardized preprocessing and sampling pipeline for training. Preprocessing: Images are first converted to grayscale. Regions with minimal colored artifacts or annotations are inpainted as long as they occupy less than 5% of the total im- age area. To normalize for varying transducer gain settings across datasets, we perform intensity rescaling by mapping the2 nd and98 th percentiles of the pixel distribution to the full dynamic range. The USrc mask is used to condition the intensity rescaling only within the US content. Data Balancing Strategy: A significant challenge in ag- gregating public medical datasets is the extreme variance in dataset size, which risks biasing the model toward the larger datasets. We mitigate this through a weighted dataset-level sampling strategy, where we define an upper bound, sam- pling thresholdN t of 50,000 to limit the effective size of each dataset for a training epoch. Thus, for each datasetD i with an actual size|D i |, the effective count is determined bymin(|D i |,N t ), and the probabilityP (D i )of drawing a sample from datasetD i during training is then calculated as: P (D i ) = min(|D i |,N t ) P j min(|D j |,N t ) (6) This strategy ensures that datasets larger thanN t contribute equally during training and smaller datasets contribute pro- portional to their true size. This approach balances maximiz- ing data diversity and ensuring that the largest contributing datasets are not over represented. 4.4. US-JEPA Architecture The context encoder (student) is a randomly initialized ViT- B/16 with an embedding dimension of 768. The predictor is a narrower transformer with an output embedding dimension of 384. The target encoder (teacher) is a frozen ViT-B/16 with the pretrained URFM weights. The predictor includes Lung n v : 536 Other n v : 694 Liver n v : 910 Prostate n v : 1910 Thyroid n v : 17899 Cardiac n v : 23904 Video and Volume a. Pancreas n f : 34749 Gallbladder n f : 51682 Thyroid n f : 110786 Kidney n f : 115237 Other n f : 120564 Liver n f : 191964 Static Frames b. Figure 2. Distribution of pretraining data. To characterize the dataset composition at the organ level, we report the distribution of a. temporal sequences, including videos and volumes (n v ), and b. individual static frames (n f ). a linear adapter that projects its output from 384 to 768 to align with the target encoder’s feature space before the loss is computed. The detailed model and training configurations are provided in Appendix Section C.1. 5. Experiments 5.1. Data 5.1.1. PRETRAINING DATASET Both US-JEPA and USrc-JEPA are developed with a large- scale, heterogeneous corpus of ultrasound data comprising 5,123,697 frames curated from 50 distinct publicly available datasets (Figure 2). To our knowledge, this is the largest aggregation of publicly available US datasets. The data contains three different US acquisition types with a ma- jority (79.7%) of the frames derived from temporal video sequences and the remaining from volumetric (3D) scans or static 2D images. The dataset also spans a wide range of human anatomy, including major organ systems and impor- tant ancillary structures. Cardiac data represents the largest subset (23,904 videos), largely sourced from high-volume video datasets like EchoNet-Dynamic and EchoNet-LVH. This is complemented by significant representations of key organs, including the liver (318K frames), thyroid (264K frames) and prostate (227K frames), as well as breast (79K frames) and lung (55K). The data originates from a diverse array of clinical contexts, ranging from large-scale institu- tional repositories (ex. Stanford, TCIA, RadImageNet) to more specialized, curated datasets, providing the model with exposure to varying image qualities, transducer frequencies, and pathological presentations. Further details regarding the pretraining dataset are included in Appendix Section A.1. 5 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Table 1. Summary of downstream datasets used for classification tasks, including the target organ, number of classes (Y), and total image counts. All datasets are public. DATASETORGANTOTAL IMAGES AUL (Y=3)LIVER735 BUSBRA (Y=2)BREAST1064 BUTTERFLY (Y=9)MULTI-ORGAN41076 FATTY LIV. (Y=2)LIVER550 GBCU (Y=3)GALLBLADDER1255 MMOTU (Y=8)OVARY1469 POCUS (Y=3)LUNG2064 TN5000 (Y=2)THYROID5000 5.1.2. ULTRABENCH DOWNSTREAM DATASET To ensure standardized and reproducible evaluation, we leverage UltraBench (Tupper & Gagn ́ e), a publicly available benchmark encompassing diverse ultrasound tasks across segmentation, detection, and classification. We add im- plementation for BUSBRA and TN5000 to UltraBench to expand anatomical coverage for breast and thyroid. Our eval- uation focuses on eight classification tasks, comprising both binary and multi-class endpoints (Table 2), to assess the dis- criminative ability and representation quality of our models. Five of the eight datasets focus primarily on cancer detection. FATTY LIVER and POCUS have non-oncological appli- cations in fatty liver disease detection and lung pathology detection (healthy, pneumonia and COVID), respectively. Lastly, BUTTERFLY is a general multi-organ detection task. In Section 5.3 and Section 5.4, our figures specifically show results for these four datasets: FATTY LIVER and POCUS since they are non-cancer related, MMOTU since it is a challenging 8-class task and BUSBRA to demonstrate cancer-specific performance. 5.2. UltraBench Evaluation To effectively evaluate the feature quality of US-JEPA and USrc-JEPA, we compare their performance against base- lines using standard SSL practice of linear probing (Alain & Bengio, 2018). We identified six US foundation models with publicly available pretrained weights: USFM (Jiao et al., 2024), URFM (Kang et al., 2025), USF-MAE (Megahed et al., 2025), EchoCare (Zhang et al., 2025b), UltraSAM (Meyer et al., 2025), and SAMUS (Lin et al., 2024). Ul- traSAM and SAMUS are foundation models specifically developed for segmentation tasks, while the remaining are general ultrasound foundation models. Furthermore, we evaluate against two pretrained, state-of-the-art universal vi- sion foundation models, DINOv3 and I-JEPA, which serve as out-of-domain benchmarks. Using the image encoder from each framework as a frozen feature extractor, we train linear probes for each downstream task across five random seeds to measure stability (Table 2). The results demonstrate that US-JEPA and USrc-JEPA achieve state-of-the-art performance on five of the eight tasks (BUSBRA, FATTY LIVER, GBCU, MMOTU and POCUS), and fall in second-place on two of the remaining five tasks (AUL and TN5000). On the remaining task, BUT- TERFLY, our framework underperforms the best baseline by less than 2% macro F1. Notably, our models excel in the most challenging scenarios where baseline methods strug- gle. The MMOTU dataset entails an eight-class ovarian tumor classification task. Average baseline performance drops below40%on MMOTU, while US-JEPA establishes a new benchmark of52.2%, surpassing the best baseline, URFM, by 9.5%. While URFM performs strongly on AUL and TN5000, our methods still offer competitive perfor- mance on the remaining downstream tasks, clearly beating general purpose vision models (DINOv3 and I-JEPA) and the majority of the domain-specific baselines. 5.3. Few-Shot Scaling for Linear Probe To evaluate the efficiency and transferability of our models’ representations, we measure downstream performance with varying amounts of supervision. For each dataset, we train linear probes (across five seeds to measure stability) on ev- ery model’s frozen features using stratified subsets of the training data: 1%, 5%, 10%, 50%, 100%. By preserving the original UltraBench splits and class prevalence through subsampling, we quantify how effectively pretrained fea- tures affect learning in low-shot regimes. In this experiment, superior representation quality is evidenced by higher quan- titative performance and faster performance convergence as label density decreases. Figure 3 illustrates macro F1 performance across four datasets under varying label quantities for linear probing. We compare degradation against URFM and USFM. This pattern is most clearly seen with the FATTY LIVER down- stream. When probes are trained with less than 10% of la- bels, our models achieve an average macro F1 of 18% higher than the URFM and USFM baselines. For the POCUS task, a similar trend is observed where the mean macro F1 de- grades faster for the baselines compared to our models. As probes are trained with fewer labels, even for BUSBRA, MMOTU, and AUL (in Appendix Section D), our models are at least on par, if not better than URFM and USFM. Complete results for all downstream tasks are included in Appendix Section D. 5.4. Robustness to Domain-Specific Corruption In real-world clinical environments, ultrasound image qual- ity is frequently compromised by hardware constraints, op- erator variability, and suboptimal acquisition settings. To evaluate the stability of US-JEPA under such distribution shifts, we performed an incremental stress test consisting of 6 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Table 2. Linear probe on downstream datasets. We compare results of linear probes trained using our proposed method against state-of-the-art baselines. We report the mean macro F1 score % over five seeds for all eight downstream datasets. Bold indicates the best result and underlining indicates the second best. MODELAULBUSBRABUTTERFLYFATTY LIV.GBCUMMOTUPOCUSTN5000 DINOV364.3±0.670.9±1.791.7±0.455.8±5.561.7±0.537.2±0.691.4±0.467.5±0.4 I-JEPA61.5±1.171.2±4.090.5±0.654.8±1.653.7±0.435.3±0.688.1±0.468.9±0.2 ULTRASAM62.6±3.170.2±3.189.6±2.466.9±3.343.5±4.939.7±1.887.3±2.163.9±2.0 SAMUS40.2±0.965.9±0.391.5±0.042.1±0.048.8±0.320.4±0.276.2±0.151.7±0.0 ECHOCARE49.2±2.464.4±0.084.1±0.742.1±0.036.2±0.521.1±0.173.8±3.949.8±3.8 USF-MAE58.1±1.462.9±0.591.1±0.342.1±0.045.9±0.328.7±0.390.1±0.056.3±1.1 USFM61.6±1.274.6±0.5 92.4±0.373.6±8.867.4±0.633.8±0.385.7±0.565.0±2.6 URFM71.5±1.169.5±2.292.1±0.482.6±6.059.1±1.742.7±0.491.7±0.377.4±0.4 US-JEPA69.6±1.573.8±1.190.8±0.382.5±1.167.0±1.452.2±0.293.1±0.073.1±0.7 USRC-JEPA67.6±0.576.0±1.291.5±0.689.2±0.970.2±0.546.8±0.292.5±0.170.8±1.3 50 55 60 65 70 75 Mean Macro F1 BUSBRA 40 50 60 70 80 90 FATTY LIVER 1%10%100% Percentage of Training Labels 15 20 25 30 35 40 45 50 Mean Macro F1 MMOTU 1%10%100% Percentage of Training Labels 50 60 70 80 90 POCUS US-JEPAUSrc-JEPAURFMUSFM Figure 3. Results for few-shot scaling. We report the mean macro F1 score % over five seeds for each model’s probe to measure performance with 1% to 100% of training labels. Note: each dataset is plotted on a unique y-axis scale to better highlight model-specific performance trends. three synthetic, domain-specific image corruptions: Gaus- sian blur, contrast depletion, and correlated speckle noise. For each corruption type, we evaluate downstream perfor- mance on test sets across three levels of increasing severity (ε ∈ 1, 2, 3). While blur and contrast are implemented via standard intensity transforms, we model speckle as a multiplicative, spatially-correlated process to approximate the “grainy” texture inherent to ultrasound scans. More de- tailed implementations for these corruptions are provided in Appendix Section E.2. Figure 4 visualizes the performance degradation for a subset of downstream tasks, with the remaining tasks detailed in Appendix Section E.1. Most notably, both US-JEPA and USrc-JEPA demonstrate significant resilience to the blur corruption. On the POCUS dataset, URFM degrades to nearly half of its uncorrupted performance, dropping from 91.7 to 46.8 F1 at maximum blur severity. In contrast, US- JEPA and USrc-JEPA are much more resilient, with F1 only dropping to 69.5 and 78.4 respectively, at maximum blur severity. Similarly, on the BUTTERFLY dataset (Appendix Section E.1), URFM drops to 22.4 F1 under severe blur, while US-JEPA and USrc-JEPA maintain F1 performance at 80.8 and 70.9. In these severe blur cases, the second baseline USFM drops to 77.8 and 59.8 F1 for POCUS and BUTTERFLY, exhibiting less degradation than URFM but still performs worse than both our models for BUTTERFLY. Regarding texture and intensity corruptions, our models match or exceed URFM and USFM for most downstream tasks. The main pitfall can be seen with GBCU and TN5000 (Appendix Section E.1). On GBCU, URFM maintains con- sistent performance with F1 only decreasing by 8.9% from normal to the highest contrast severity, while our models drop by 47%. A similar pattern is seen on TN5000, where URFM drops by 11% while US-JEPA and USrc-JEPA drop by 23.2% and 17.7%, respectively. This performance dispar- ity likely stems from a difference in gallbladder and thyroid pretraining data density. While our dataset is larger overall, gallbladder and thyroid frames constitute only 0.27% and 5.2% of our data, respectively, compared to 4% and 44.2% in the smaller URFM pretraining set. This difference in prevalence may have limited our model’s ability to learn stronger features for these anatomies. In other cases, URFM, USFM and our models maintain comparable performance, showing expected degradation under increasingly severe contrast reduction. For speckle corruption, our models demonstrate remarkable stability. Specifically on BUTTERFLY, US-JEPA and USrc- 7 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound 0 25 50 75 100 BUSBRAFATTY LIVERMMOTUPOCUS 0 25 50 75 100 0123 0 25 50 75 100 012301230123 Blur Contrast Speckle Severity Mean Macro F1 US-JEPAUSrc-JEPAURFMUSFM Figure 4. Robustness to domain-specific corruption. Results show the mean macro F1 across five seeds for each model-dataset- corruption permutation. Linear probes were trained on full, uncorrupted training sets and evaluated on increasingly corrupted test sets to assess structural representation stability. JEPA only drop by 0.6% and 9.8% under severe speckle noise (ε = 3), whereas USFM and URFM drop by 25% and 44.6%. Notably, across all downstream tasks, US-JEPA and USrc-JEPA outperform or match URFM under peak speckle severity, demonstrating superior robustness to this prevalent ultrasound-specific noise. US-JEPA also consistently out- performs USFM at the highest speckle severity across all eight downstreams. While USFM shows competitive results against USrc-JEPA on four benchmarks, it remains inferior to USrc-JEPA on the remaining half, often by substantial margins, reaching a performance gap of up to 34.39% on FATTY LIVER. Ultimately, this stress test quantifies the invariance of our learned representations under significant corruption rela- tive to the current state-of-the-art, URFM and additional baseline, USFM. By evaluating against domain-specific syn- thetic noise, we provide a practical validation of US-JEPA and USrc-JEPA’s feature stability. This evaluation is a criti- cal requirement given the high variance in US image quality across different manufacturers, operators, and clinical en- vironments. The observed performance under corruption suggests that our approach is able to better capture struc- tural semantics in US instead of overfitting to superficial, pixel-level features. 6. Discussion We present US-JEPA, a novel approach that demonstrates that training using a JEPA framework with a robust, static teacher yields stable, data-efficient representations for ul- trasound. By adopting the SALT paradigm, we decouple the optimization of the student and teacher and leverage the semantic priors of a domain-specific teacher. Furthermore, with USrc-JEPA we show that adding USrc to the frame- work also achieves competitive performance on downstream tasks. This efficiency is highlighted in our downstream stress tests: US-JEPA remains robust in low-data regimes for linear probing and maintains performance during out-of- distribution evaluation through domain-specific corruptions. Despite these successes, certain limitations remain. The modest performance in specific corruption experiments and 8 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound on downstream datasets like TN5000 and AUL suggests that performance may be sensitive to organ-level density in pretraining data. Moving forward, we will build a more di- verse pretraining corpus and implement an organ-weighted sampling strategy to improve downstream results. This work also advances standardization in ultrasound foun- dation model research through three key contributions. First, we provide the first rigorous linear probing comparison across all published models to assess intrinsic represen- tation quality. Second, we expand the UltraBench frame- work by integrating the TN5000 and BUSBRA classification tasks for broader anatomical coverage. Finally, we establish baseline performance for existing models using UltraBench, creating a better foundation to benchmark comparative per- formance moving forward. These findings confirm the utility of the JEPA framework for ultrasound self-supervised learning and position US-JEPA as a clinically impactful ultrasound foundation model. Impact Statement US-JEPA is a novel paradigm for applying self-supervised learning to the US image domain. By pretraining on the largest publicly available corpus of US imaging, we demon- strate competitive results with previous US foundation mod- els. We establish the importance of few-shot scaling for linear probes and evaluating robustness to domain-specific corruptions to truly determine representation quality in US. Most importantly, we standardize comparative benchmark- ing by evaluating all current US foundation models on eight tasks from UltraBench. By relying exclusively on publicly accessible data and standardized evaluation, US-JEPA low- ers the barrier to entry for US research, fostering broader participation and more equitable development of ultrasound foundation models. References Abbasian Ardakani, A., Mohammadi, A., Mirza-Aghazadeh- Attari, M., and Acharya, U. R. An open-access breast lesion ultrasound image database:Applicable in artificial intelligence studies. Computers in Biology and Medicine, 152:106438, January 2023.ISSN 0010-4825. doi: 10.1016/j.compbiomed.2022.106438. URLhttps://w.sciencedirect.com/ science/article/pii/S0010482522011465. Agata Momot. Common Carotid Artery Ultrasound Images, November 2022. URLhttps://data.mendeley. com/datasets/d4xt63mgjm/1. Al-Dhabyani, W., Gomaa, M., Khaled, H., and Fahmy, A. Dataset of breast ultrasound images. Data in Brief, 28: 104863, November 2019. ISSN 2352-3409. doi: 10.1016/ j.dib.2019.104863. URLhttps://pmc.ncbi.nlm. nih.gov/articles/PMC6906728/. Alain, G. and Bengio, Y.Understanding interme- diate layers using linear classifier probes, Novem- ber 2018. URLhttp://arxiv.org/abs/1610. 01644. arXiv:1610.01644 [stat]. Ambsdorf, J., Munk, A., Llambias, S., Christensen, A. N., Mikolaj, K., Balestriero, R., Tolsgaard, M., Feragen, A., and Nielsen, M. General Methods Make Great Domain- specific Foundation Models: A Case-study on Fetal Ultra- sound, June 2025. URLhttp://arxiv.org/abs/ 2506.19552. arXiv:2506.19552 [cs]. Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., and Ballas, N. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, April 2023. URLhttp://arxiv.org/ abs/2301.08243. arXiv:2301.08243 [cs]. Basu, S., Gupta, M., Rana, P., Gupta, P., and Arora, C. Sur- passing the Human Accuracy: Detecting Gallbladder Can- cer from USG Images with Curriculum Learning. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 20854–20864. IEEE, 6 2022. ISBN 978-1-6654-6946-3. doi: 10.1109/CVPR52688. 2022.02022. Baum, Z. M. C., Shaheer U. Saeed, Min, Z., Yipeng Hu, and Barratt, D. C. MR to Ultrasound Registration for Prostate Challenge - Dataset, June 2023. URLhttps: //zenodo.org/record/8004388. Born, J., Wiedemann, N., Cossio, M., Buhre, C., Br ̈ andle, G., Leidermann, K., Goulet, J., Aujayeb, A., Moor, M., Rieck, B., and Borgwardt, K. Accelerating Detection of Lung Pathologies with Explainable Ultrasound Image Analysis. Applied Sciences, 11(2):672, 1 2021. ISSN 2076-3417. doi: 10.3390/app11020672. Borna, M.-R., Saadat, H., Sepehri, M. M., Torkashvand, H., Torkashvand, L., and Pilehvari, S. AI-powered diagnosis of ovarian conditions: insights from a newly introduced ultrasound dataset. Frontiers in Physiology, 16, July 2025. ISSN 1664-042X. doi: 10.3389/fphys.2025.1520898. URLhttps://w.frontiersin.org/ journals/physiology/articles/10.3389/ fphys.2025.1520898/full. Publisher: Frontiers. Butterfly Network. Butterflynetwork/mitgrandhack2018. https://github.com/ButterflyNetwork/ MITGrandHack2018 , 2018.GitHub repository, accessed 2018. 9 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Byra, M., Styczynski, G., Szmigielski, C., Kalinowski, P., Michałowski,Ł., Paluszkiewicz, R., Ziarkiewicz- Wr ́ oblewska, B., Zieniewicz, K., Sobieraj, P., and Now- icki, A. Transfer learning with deep convolutional neural network for liver steatosis assessment in ultrasound im- ages. International Journal of Computer Assisted Radi- ology and Surgery, 13(12):1895–1903, 12 2018. ISSN 1861-6410. doi: 10.1007/s11548-018-1843-2. Caron, M., Touvron, H., Misra, I., J ́ egou, H., Mairal, J., Bojanowski, P., and Joulin, A.Emerging Properties in Self-Supervised Vision Transformers, May 2021. URLhttp://arxiv.org/abs/2104. 14294. arXiv:2104.14294 [cs]. Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A Sim- ple Framework for Contrastive Learning of Visual Rep- resentations, July 2020. URLhttp://arxiv.org/ abs/2002.05709. arXiv:2002.05709 [cs]. Choudhari, A. and Korde, A.PCOS detection us- ing ultrasound images.URLhttps://w. kaggle.com/datasets/anaghachoudhari/ pcos-detection-using-ultrasound-images. Ding, Y., Yang, Q., Wang, Y., Chen, D., Qin, Z., and Zhang, J. MallesNet: A multi-object assistance based network for brachial plexus segmentation in ultrasound images. Medical Image Analysis, 80:102511, August 2022. ISSN 1361-8415. doi: 10.1016/j.media.2022.102511. URLhttps://w.sciencedirect.com/ science/article/pii/S136184152200158X. Duffy, G., Cheng, P. P., Yuan, N., He, B., Kwan, A. C., Shun-Shin, M. J., Alexander, K. M., Ebinger, J., Lun- gren, M. P., Rader, F., Liang, D. H., Schnittger, I., Ash- ley, E. A., Zou, J. Y., Patel, J., Witteles, R., Cheng, S., and Ouyang, D. High-Throughput Precision Phe- notyping of Left Ventricular Hypertrophy With Cardio- vascular Deep Learning. JAMA Cardiology, 7(4):386– 395, April 2022.ISSN 2380-6583.doi: 10.1001/ jamacardio.2021.6059. URLhttps://doi.org/10. 1001/jamacardio.2021.6059. Ebadi, A., Xi, P., MacLean, A., Florea, A., Tremblay, S., Kohli, S., and Wong, A. COVIDx-US: An Open-Access Benchmark Dataset of Ultrasound Imaging Data for AI- Driven COVID-19 Analytics. Frontiers in Bioscience, 27 (7):198, June 2022. ISSN 2768-6698. doi: 10.31083/j. fbl2707198. Place: Landmark edition. Eisenbrey, J., Lyshchik, A., and Wessner, C.Ultra- sound data of a variety of liver masses, 2021. URL https://w.cancerimagingarchive.net/ collection/b-mode-and-ceus-liver/. Geng, Y., Meng, G., Chen, M., Cao, G., Zhao, M., Zhao, J., and Liu, H. Force Sensing Guided Artery-Vein Seg- mentation via Sequential Ultrasound Images. In Lin- guraru, M. G., Dou, Q., Feragen, A., Giannarou, S., Glocker, B., Lekadir, K., and Schnabel, J. A. (eds.), Medical Image Computing and Computer Assisted In- tervention – MICCAI 2024, p. 656–666, Cham, 2024. Springer Nature Switzerland. ISBN 978-3-031-72083-3. doi: 10.1007/978-3-031-72083-3 61. G ́ omez-Flores, W., Gregorio-Calas, M. J., and Coelho de Albuquerque Pereira, W. BUS-BRA: A breast ultrasound dataset for assessing computer-aided diagnosis systems. Medical Physics, 51(4):3110–3123, 4 2024. ISSN 0094- 2405. doi: 10.1002/mp.16812. gong,h.haifangong/TRFE-Net-for-thyroid- nodule-segmentation,February2026.URL https://github.com/haifangong/ TRFE-Net-for-thyroid-nodule-segmentation. original-date: 2021-02-11T08:33:39Z. Guo, Y., Duan, X., Wang, C., and Guo, H. Segmenta- tion and recognition of breast ultrasound images based on an expanded U-Net. PLoS ONE, 16(6):e0253202, June 2021. ISSN 1932-6203. doi: 10.1371/journal.pone. 0253202.URLhttps://pmc.ncbi.nlm.nih. gov/articles/PMC8205136/. Hann, A., Bettac, L., Haenle, M. M., Graeter, T., Berger, A. W., Dreyhaupt, J., Schmalstieg, D., Zoller, W. G., and Egger, J.Algorithm guided outlining of 105 pancreatic cancer liver metastases in Ultrasound.Scientific Reports, 7(1):12779, Oc- tober 2017.ISSN 2045-2322.doi:10.1038/ s41598-017-12925-z. URLhttps://w.nature. com/articles/s41598-017-12925-z.Pub- lisher: Nature Publishing Group. He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momen- tum Contrast for Unsupervised Visual Representation Learning, March 2020. URLhttp://arxiv.org/ abs/1911.05722. arXiv:1911.05722 [cs]. He, Q., Bano, S., Liu, J., Liu, W., Stoyanov, D., and Zuo, S.Query2<math><msup is=”true”><mrow is=”true”></mrow><mrowis=”true”><mn is=”true”>2</mn></mrow></msup></math>: Queryoverqueriesforimprovinggastroin- testinal stromal tumour detection in an endo- scopic ultrasound.Computers in Biology and Medicine, 152:106424, January 2023.ISSN 0010- 4825.doi:10.1016/j.compbiomed.2022.106424. URLhttps://w.sciencedirect.com/ science/article/pii/S0010482522011325. 10 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Hondru, V., Croitoru, F. A., Minaee, S., Ionescu, R. T., and Sebe, N.Masked Image Modeling: A Survey, July 2025. URLhttp://arxiv.org/abs/2408. 06687. arXiv:2408.06687 [cs]. Hou, X., Hua, M., Zhang, W., Ji, J., Zhang, X., Jiang, H., Li, M., Wu, X., Zhao, W., Sun, S., Cao, L., and Wang, L.An ultrasonography of thyroid nodules dataset with pathological diagnosis annota- tion for deep learning.Scientific Data, 11(1):1272, November 2024.ISSN 2052-4463.doi: 10.1038/ s41597-024-04156-5. URLhttps://w.nature. com/articles/s41597-024-04156-5 .Pub- lisher: Nature Publishing Group. Huang, J., Zhang, J., Zhang, Y., Li, X., Ma, X., Deng, J., Shen, H., Wang, D., mei, l., and Lei, C. BUSIwhu: Breast Cancer Ultrasound Image Dataset, Oc- tober 2023. URLhttps://data.mendeley.com/ datasets/k6cpmwybk3/1. Iqbal, A.BUSuc.1, October 2023a.doi: 10.17632/3ksd7w7jkx.1.URLhttps://data. mendeley.com/datasets/3ksd7w7jkx/1. Publisher: Mendeley Data. Iqbal, A.BUSC Dataset.1, March 2023b.doi: 10.17632/vckdnhtw26.1.URLhttps://data. mendeley.com/datasets/vckdnhtw26/1. Publisher: Mendeley Data. Jiao, J., Zhou, J., Li, X., Xia, M., Huang, Y., Huang, L., Wang, N., Zhang, X., Zhou, S., Wang, Y., and Guo, Y. USFM: A universal ultrasound foundation model general- ized to tasks and organs towards label efficient image anal- ysis. Medical Image Analysis, 96:103202, August 2024. ISSN 1361-8415. doi: 10.1016/j.media.2024.103202. URLhttps://w.sciencedirect.com/ science/article/pii/S1361841524001270. Jimenez, Y., Rodriguez-Alvarez, M. J., Escudero, L., San- doval, C., and Lakshminarayanan, V. Ultrasound Breast images denoising using Generative Adversarial Networks (GANs), 2024. URLhttps://data.mendeley. com/datasets/g3cmj46xyx/1. Juvekar, P., Dorent, R., K ̈ ogl, F., Torio, E., Barr, C., Rigolo, L., Galvin, C., Jowkar, N., Kazi, A., Haouchine, N., Cheema, H., Navab, N., Pieper, S., Wells, W. M., Bi, W. L., Golby, A., Frisken, S., and Kapur, T. The Brain Re- section Multimodal Imaging Database (ReMIND), 2023. URLhttps://w.cancerimagingarchive. net/collection/remind/. Kang, Q., Gao, J., Zhao, H., He, Z., Li, K., and Lao, Q. D2MAE: Diffusional Deblurring MAE for Ultrasound Image Pre-training. Kang, Q., Lao, Q., Gao, J., Bao, W., He, Z., Du, C., Lu, Q., and Li, K.URFM: A general Ultrasound Representation Foundation Model for advancing ultrasound image diagnosis. iScience, 28(8), August 2025. ISSN 2589-0042. doi: 10.1016/j.isci.2025.112917. URLhttps://w.cell.com/iscience/ abstract/S2589-0042(25)01178-2. Publisher: Elsevier. Kim, S., Jin, P., Song, S., Chen, C., Li, Y., Ren, H., Li, X., Liu, T., and Li, Q. EchoFM: Foundation Model for Generalizable Echocardiogram Analysis. IEEE trans- actions on medical imaging, 44(10):4049–4062, Octo- ber 2025. ISSN 0278-0062. doi: 10.1109/TMI.2025. 3580713.URLhttps://pmc.ncbi.nlm.nih. gov/articles/PMC12616925/. Kim-Ann. ftsvd/USAnotAI, June 2024. URLhttps: //github.com/ftsvd/USAnotAI. original-date: 2019-07-26T17:37:10Z. Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., Doll ́ ar, P., and Girshick, R. Segment Anything, April 2023. URLhttp://arxiv.org/abs/2304. 02643. arXiv:2304.02643 [cs]. Kr ̈ onke, M., Eilers, C., Dimova, D., K ̈ ohler, M., Buschner, G., Schweiger, L., Konstantinidou, L., Makowski, M., Nagarajah, J., Navab, N., Weber, W., and Wendler, T. Tracked 3D ultrasound and deep neural network- based thyroid segmentation reduce interobserver variabil- ity in thyroid volumetry. PLOS ONE, 17(7):e0268550, July 2022. ISSN 1932-6203. doi: 10.1371/journal. pone.0268550. URLhttps://dx.plos.org/10. 1371/journal.pone.0268550. Leclerc, S., Smistad, E., Pedrosa, J., Østvik, A., Cervenan- sky, F., Espinosa, F., Espeland, T., Berg, E. A. R., Jodoin, P.-M., Grenier, T., Lartizien, C., D’hooge, J., Lovstakken, L., and Bernard, O. Deep Learning for Segmentation Using an Open Large-Scale Dataset in 2D Echocardio- graphy. IEEE Transactions on Medical Imaging, 38 (9):2198–2210, September 2019.ISSN 1558-254X. doi: 10.1109/TMI.2019.2900516.URLhttps:// ieeexplore.ieee.org/document/8649738. Li, J., Zhang, P., Wang, T., Wang, K., and Sheng, B. LEPset, June 2023. URLhttps://zenodo.org/record/ 8041285. Li, X., Huang, C., Li, C.-L., Malach, E., Susskind, J., Thi- lak, V., and Littwin, E. Rethinking JEPA: Compute- Efficient Video SSL with Frozen Teachers, Septem- ber 2025. URLhttp://arxiv.org/abs/2509. 24317. arXiv:2509.24317 [cs]. 11 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Liang, X., Cao, Q., Huang, R., and Lin, L. Recognizing focal liver lesions in contrast-enhanced ultrasound with discriminatively trained spatio-temporal model. In 2014 IEEE 11th International Symposium on Biomedical Imag- ing (ISBI), p. 1184–1187, Beijing, China, April 2014. IEEE. ISBN 978-1-4673-1961-4. doi: 10.1109/ISBI. 2014.6868087. URLhttp://ieeexplore.ieee. org/document/6868087/. Lin, X., Xiang, Y., Yu, L., and Yan, Z. Beyond Adapt- ing SAM: Towards End-to-End Ultrasound Image Seg- mentation via Auto Prompting. In Linguraru, M. G., Dou, Q., Feragen, A., Giannarou, S., Glocker, B., Lekadir, K., and Schnabel, J. A. (eds.), Medical Image Computing and Computer Assisted Intervention – MIC- CAI 2024, volume 15008, p. 24–34. Springer Nature Switzerland, Cham, 2024. ISBN 978-3-031-72110-6 978-3-031-72111-3. doi: 10.1007/978-3-031-72111-33. URLhttps://link.springer.com/10.1007/ 978-3-031-72111-3_3.Series Title: Lecture Notes in Computer Science. Lin, Z., Lin, J., Zhu, L., Fu, H., Qin, J., and Wang, L. A New Dataset and A Baseline Model for Breast Lesion De- tection in Ultrasound Videos, July 2022. URLhttp:// arxiv.org/abs/2207.00141 . arXiv:2207.00141 [eess]. Luo, G., Xu, M., Chen, H., Liang, X., Tao, X., Ni, D., Jeong, H., Kim, C., Stock, R., Baumgartner, M., Kirchhoff, Y., Rokuss, M., Maier-Hein, K., Yang, Z., Fan, T., Boutry, N., Tereshchenko, D., Moine, A., Charmetant, M., Sauer, J., Du, H., Bai, X.-H., Raikar, V. P., Montoya-del Angel, R., Marti, R., Luna, M., Lee, D., Qayyum, A., Mazher, M., Guo, Q., Wang, C., Awasthi, N., Zhao, Q., Wang, W., Wang, K., Wang, Q., and Dong, S. Tumor Detection, Segmentation and Classification Challenge on Automated 3D Breast Ultrasound: The TDSC-ABUS Challenge, Jan- uary 2025. URLhttp://arxiv.org/abs/2501. 15588. arXiv:2501.15588 [eess]. Megahed, Y., Ducharme, R., Erman, A., Walker, M., Hawken, S., and Chan, A. D. C. USF-MAE: Ultrasound Self-Supervised Foundation Model with Masked Autoen- coding, November 2025. URLhttp://arxiv.org/ abs/2510.22990. arXiv:2510.22990 [eess]. Mei, X., Liu, Z., Robson, P. M., Marinelli, B., Huang, M., Doshi, A., Jacobi, A., Cao, C., Link, K. E., Yang, T., Wang, Y., Greenspan, H., Deyer, T., Fayad, Z. A., and Yang, Y.RadImageNet: An Open Radiologic Deep Learning Research Dataset for Effective Trans- fer Learning. Radiology: Artificial Intelligence, 4(5): e210315, September 2022. doi: 10.1148/ryai.210315. URLhttps://pubs.rsna.org/doi/10.1148/ ryai.210315. Publisher: Radiological Society of North America. Meiburger, K.DATASET for ”Deep learning seg- mentation of transverse musculoskeletal ultrasound images for neuromuscular disease assessment”, July 2021.URLhttps://data.mendeley.com/ datasets/3jykz7wz8d/1. Meiburger, K. DATASET for: ”Carotid Ultrasound Bound- ary Study (CUBS): Technical considerations on an open multi-center analysis of computerized measure- ment systems for intima-media thickness measurement on common carotid artery longitudinal B-mode ultra- sound scans”, March 2022. URLhttps://data. mendeley.com/datasets/m7ndn58sv6/1. Meyer, A., Murali, A., Zarin, F., Mutter, D., and Padoy, N. Ultrasam: a foundation model for ultrasound using large open-access segmentation datasets. International Journal of Computer Assisted Radiology and Surgery, September 2025.ISSN 1861-6429.doi: 10.1007/ s11548-025-03517-8. URLhttps://doi.org/10. 1007/s11548-025-03517-8. Montoya, A., Sterling, D., Hasnin, kaggle446, shirzad, Cukierski, W., and yffud. Ultrasound nerve segmen- tation.https://kaggle.com/competitions/ ultrasound-nerve-segmentation, 2016. Kag- gle. Natarajan, S., Priester, A., Margolis, D., Huang, J., and Marks, L.Prostate MRI and Ultrasound WithPathologyandCoordinatesofTracked Biopsy (Prostate-MRI-US-Biopsy), 2020.URL https://w.cancerimagingarchive.net/ collection/prostate-mri-us-biopsy/. Ndzimbong, W., Fourniol, C., Themyr, L., Thome, N., Keeza, Y., Sauer, B., Pi ́ echaud, P.-T., M ́ ejean, A., Marescaux, J., George, D., Mutter, D., Hostettler, A., and Collins, T. TRUSTED: The Paired 3D Transabdom- inal Ultrasound and CT Human Data for Kidney Seg- mentation and Registration Research. Scientific Data, 12 (1):615, April 2025. ISSN 2052-4463. doi: 10.1038/ s41597-025-04467-1. URLhttps://w.nature. com/articles/s41597-025-04467-1 .Pub- lisher: Nature Publishing Group. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. DINOv2: Learn- ing Robust Visual Features without Supervision, Febru- 12 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound ary 2024. URLhttp://arxiv.org/abs/2304. 07193. arXiv:2304.07193 [cs]. Ouyang, D., He, B., Ghorbani, A., Lungren, M. P., Ashley, E. A., Liang, D. H., and Zou, J. Y. EchoNet-Dynamic: a Large New Cardiac Motion Video Data Resource for Medical Machine Learning. Ouyang, D., He, B., Ghorbani, A., Yuan, N., Ebinger, J., Langlotz, C. P., Heidenreich, P. A., Harrington, R. A., Liang, D. H., Ashley, E. A., and Zou, J. Y. Video-based AI for beat-to-beat assessment of cardiac function. Nature, 580(7802):252–256, April 2020. ISSN 0028-0836. doi: 10.1038/s41586-020-2145-8.URLhttps://pmc. ncbi.nlm.nih.gov/articles/PMC8979576/. Pawłowska, A., ́ Cwierz Pie ́ nkowska, A., Domalik, A., Jagu ́ s, D., Kasprzak, P., Matkowski, R., Fura, Ł., Nowicki, A., and Zolek, N.A Curated Bench- mark Dataset for Ultrasound Based Breast Le- sion Analysis (Breast-Lesions-USG), 2024.URL https://w.cancerimagingarchive.net/ collection/breast-lesions-usg/. Pedraza, L., Vargas, C., Narv ́ aez, F., Dur ́ an, O., Mu ̃ noz, E., and Romero, E.An open access thyroid ul- trasound image database.In 10th International SymposiumonMedicalInformationProcessing and Analysis, volume 9287, p. 188–193. SPIE, January 2015.doi:10.1117/12.2073532.URL https://w.spiedigitallibrary.org/ conference-proceedings-of-spie/9287/ 92870W/An-open-access-thyroid-ultrasound-image-database/ 10.1117/12.2073532.full. Shao,W. and Brisbane,W.Micro-Ultrasound Prostate Segmentation Dataset, January 2024. URL https://zenodo.org/doi/10.5281/zenodo. 10475293. Sim ́ eoni, O., Vo, H. V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S., Rama- monjisoa, M., Massa, F., Haziza, D., Wehrstedt, L., Wang, J., Darcet, T., Moutakanni, T., Sentana, L., Roberts, C., Vedaldi, A., Tolan, J., Brandt, J., Couprie, C., Mairal, J., J ́ egou, H., Labatut, P., and Bojanowski, P. DINOv3, Au- gust 2025. URLhttp://arxiv.org/abs/2508. 10104. arXiv:2508.10104 [cs]. Singla, R., Ringstrom, C., Hu, G., Lessoway, V., Reid, J., Nguan, C., and Rohling, R. The Open Kidney Ultra- sound Data Set, December 2022. URLhttp://arxiv. org/abs/2206.06657. arXiv:2206.06657 [eess]. Steffner, K. R., Christensen, M., Gill, G., Bowdish, M., Rhee, J., Kumaresan, A., He, B., Zou, J., and Ouyang, D.Deep learning for transesophageal echocardiog- raphy view classification.Scientific Reports, 14(1): 11, January 2024. ISSN 2045-2322. doi: 10.1038/ s41598-023-50735-8. URLhttps://w.nature. com/articles/s41598-023-50735-8.Pub- lisher: Nature Publishing Group. Tiantian Yang. uterine fibroid ultrasound images, March 2023.URLhttps://data.mendeley.com/ datasets/552zbvzwrk/1. Tupper, A. and Gagn ́ e, C. Revisiting Data Augmentation for Ultrasound Images. Turki, A., Mahdi Obaid, A., Bellaaj, H., Ksantini, M., and Altaee, A.Gallblader Diseases Dataset, 2025.URLhttps://data.mendeley.com/ datasets/r6h24d2d3y/1. Vallez, N., Bueno, G., Deniz, O., Rienda, M. A., and Pas- tor, C. BUS-UCLM: Breast ultrasound lesion segmen- tation dataset, February 2024. URLhttps://data. mendeley.com/datasets/7fvgj4jsp7/1. Vitale, S., Orlando, J. I., Iarussi, E., and Larrabide, I. Improving realism in patient-specific abdominal ultra- sound simulation using CycleGANs. International Jour- nal of Computer Assisted Radiology and Surgery, 15 (2):183–192, February 2020. ISSN 1861-6429. doi: 10.1007/s11548-019-02046-5. Xie, Y., Gu, L., Harada, T., Zhang, J., Xia, Y., and Wu, Q.Rethinking masked image modelling for medical image representation.Medical Im- age Analysis, 98:103304, December 2024.ISSN 1361-8415.doi:10.1016/j.media.2024.103304. URLhttps://w.sciencedirect.com/ science/article/pii/S1361841524002299. Xu, Y., Zheng, B., Liu, X., Wu, T., Ju, J., Wang, S., Lian, Y., Zhang, H., Liang, T., Sang, Y., Jiang, R., Wang, G., Ren, J., and Chen, T. Improving artificial intelligence pipeline for liver malignancy diagnosis using ultrasound images and video frames. Briefings in Bioinformatics, 24(1), 1 2023. ISSN 1467-5463. doi: 10.1093/bib/bbac569. Yamashita, R., Kapoor, T., Alam, M. N., Galimzianova, A., Syed, S. A., Ugur Akdogan, M., Alkim, E., Wentland, A. L., Madhuripan, N., Goff, D., Barbee, V., Sheybani, N. D., Sagreiya, H., Rubin, D. L., and Desser, T. S. To- ward Reduction in False-Positive Thyroid Nodule Biop- sies with a Deep Learning–based Risk Stratification Sys- tem Using US Cine-Clip Images. Radiology: Artificial Intelligence, 4(3):e210174, May 2022. ISSN 2638-6100. doi: 10.1148/ryai.210174. URLhttps://pmc.ncbi. nlm.nih.gov/articles/PMC9152684/. 13 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Yang, J., Ding, X., Zheng, Z., Xu, X., and Li, X. GraphEcho: Graph-Driven Unsupervised Domain Adap- tation for Echocardiogram Video Segmentation, Septem- ber 2023. URLhttp://arxiv.org/abs/2309. 11145. arXiv:2309.11145 [cs]. Yap, M. H., Pons, G., Mart ́ ı, J., Ganau, S., Sent ́ ıs, M., Zwiggelaar, R., Davison, A. K., and Mart ́ ı, R. Auto- mated Breast Ultrasound Lesions Detection Using Con- volutional Neural Networks. IEEE Journal of Biomed- ical and Health Informatics, 22(4):1218–1226, July 2018. ISSN 2168-2194, 2168-2208. doi: 10.1109/JBHI. 2017.2731873. URLhttps://ieeexplore.ieee. org/document/8003418/. Zhang, H., Liu, Q., Han, X., Niu, L., and Sun, W. TN5000: An Ultrasound Image Dataset for Thyroid Nod- ule Detection and Classification. Scientific Data, 12 (1):1437, 8 2025a. ISSN 2052-4463. doi: 10.1038/ s41597-025-05757-4. Zhang, H., Wu, Y., Zhao, M., Chen, Z., Li, R., Zhu, F., Zhao, H., Yuan, X., Yang, M., Qiu, C., Cong, X., Chen, H., Luan, L., Wong, R. H. L., Liao, H., Graham, C. A., Chang, S., Tao, G., Yi, D., Lei, Z., Navab, N., Ourselin, S., Luo, J., Liu, H., and Meng, G. A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications, September 2025b. URLhttp://arxiv. org/abs/2509.11752. arXiv:2509.11752 [cs]. Zhang, S., Xu, Y., Usuyama, N., Xu, H., Bagga, J., Tinn, R., Preston, S., Rao, R., Wei, M., Valluri, N., Wong, C., Tupini, A., Wang, Y., Mazzola, M., Shukla, S., Liden, L., Gao, J., Crabtree, A., Piening, B., Bi- fulco, C., Lungren, M. P., Naumann, T., Wang, S., and Poon, H. A Multimodal Biomedical Foundation Model Trained from Fifteen Million Image–Text Pairs. NEJM AI, 2(1):AIoa2400640, January 2025c. doi: 10.1056/ AIoa2400640. URLhttps://ai.nejm.org/doi/ full/10.1056/AIoa2400640.Publisher: Mas- sachusetts Medical Society. Zhao, Q., Lyu, S., Bai, W., Cai, L., Liu, B., Cheng, G., Wu, M., Sang, X., Yang, M., and Chen, L. Mmotu: A multi- modality ovarian tumor ultrasound image dataset for un- supervised cross-domain semantic segmentation, 2023. URL https://arxiv.org/abs/2207.06799. 14 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound A. Dataset Details A.1. Pretraining Datasets For complete transparency, we have included the name of each data set and associated metadata. Table 3. Comprehensive breakdown of datasets used, categorized by anatomical region, frame type, and total frame count. DATASET NAMEANATOMYFRAME TYPEFRAME COUNT ULTRASOUNDCASES.INFOABDOMINALSTATIC3199 ABDOMINAL US (VITALE ET AL., 2020)ABDOMINALSTATIC617 ULTRASOUNDCASES.INFOAPPENDIXSTATIC781 ULTRASOUNDCASES.INFOBLADDERSTATIC450 USANOTAI (KIM-ANN, 2024)BLADDERSTATIC61 RADIMAGENET (MEI ET AL., 2022)BLADDERSTATIC758 USANOTAIBOWELSTATIC61 ULTRASOUNDCASES.INFOBRAINSTATIC466 REMIND (JUVEKAR ET AL., 2023)BRAINVIDEO18734 ULTRASOUNDCASES.INFOBREASTSTATIC4956 BUSI (AL-DHABYANI ET AL., 2019)BREASTSTATIC780 BUS UC (IQBAL, 2023A)BREASTSTATIC811 BREAST-LESIONS-USG (PAWŁOWSKA ET AL., 2024)BREASTSTATIC256 BUSC (IQBAL, 2023B)BREASTSTATIC250 DATASETA (JIMENEZ ET AL., 2024)BREASTSTATIC250 BUV (LIN ET AL., 2022)BREASTVIDEO25026 BUID (ABBASIAN ARDAKANI ET AL., 2023)BREASTSTATIC205 ABUS-TDSC (LUO ET AL., 2025)BREASTVOLUME44474 BREAST S1 (GUO ET AL., 2021)BREASTSTATIC201 UDIAT (YAP ET AL., 2018)BREASTSTATIC163 BUSIWHU (HUANG ET AL., 2023)BREASTSTATIC927 BUSUCLM (VALLEZ ET AL., 2024)BREASTSTATIC683 CARDIACCCAUS (AGATA MOMOT, 2022)CARDIACSTATIC1100 ULTRASOUNDCASES.INFOCARDIACSTATIC1443 ECHONET-LVH (DUFFY ET AL., 2022)CARDIACVIDEO1957928 CAROTID US BOUNDARY STUDY (MEIBURGER, 2022)CARDIACSTATIC500 CAMUS (LECLERC ET AL., 2019)CARDIACVIDEO19232 CARDIACUDC (YANG ET AL., 2023)CARDIACVIDEO38071 ECHONET: TEE-VIEW-CLASSIFIER (STEFFNER ET AL., 2024)CARDIACVIDEO55261 ECHONET DYNAMIC (OUYANG ET AL.)CARDIACVIDEO1770636 RADIMAGENETCARDIACSTATIC19062 RADIMAGENETFIBROIDSTATIC2095 GB1 (TURKI ET AL., 2025)GALLBLADDERSTATIC10692 USANOTAIGALLBLADDERSTATIC61 RADIMAGENETGALLBLADDERSTATIC39743 ULTRASOUNDCASES.INFOGALLBLADDERSTATIC1186 GIST EUS (HE ET AL., 2023)GASTROINTESTINALSTATIC514 ULTRASOUNDCASES.INFOGASTROINTESTINALSTATIC2618 USANOTAIKIDNEYSTATIC61 TRUSTED (NDZIMBONG ET AL., 2025)KIDNEYVOLUME7280 OPENKIDNEY (SINGLA ET AL., 2022)KIDNEYSTATIC500 RADIMAGENETKIDNEYSTATIC111252 ULTRASOUNDCASES.INFOKIDNEYSTATIC3924 RADIMAGENETLIVERSTATIC78535 ULTRASOUNDCASES.INFOLIVERSTATIC2714 USANOTAILIVERSTATIC61 SYSU-CEUS-FLL (LIANG ET AL., 2014)LIVERSAMPLED VIDEO110654 B-MODE-AND-CEUS-LIVER (EISENBREY ET AL., 2021)LIVERVIDEO126543 COVID-BLUES (BORN ET AL., 2021)LUNGVIDEO31696 COVIDX-US (EBADI ET AL., 2022)LUNGVIDEO22697 ULTRASOUNDCASES.INFOLUNGSTATIC714 ULTRASOUNDCASES.INFOLYMPH NODESTATIC320 ULTRASOUNDCASES.INFOMUSCLESTATIC17459 MUS-V (GENG ET AL., 2024)MUSCLESAMPLED VIDEO3114 15 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound DATASET NAMEANATOMYFRAME TYPEFRAME COUNT MUSCLEUS (MEIBURGER, 2021)MUSCLESTATIC8169 UBPD (DING ET AL., 2022)NERVESAMPLED VIDEO955 ULTRASOUNDCASES.INFONERVESTATIC458 NERVEUS (MONTOYA ET AL., 2016)NERVESTATIC11143 OVARIANUS (BORNA ET AL., 2025)OVARIANSTATIC301 PCOS US (CHOUDHARI & KORDE)OVARIANSTATIC3841 RADIMAGENETOVARIANSTATIC3595 ULTRASOUNDCASES.INFOPANCREASSTATIC1506 LEPSET (LI ET AL., 2023)PANCREASSTATIC11493 105US (HANN ET AL., 2017)PANCREASSTATIC105 RADIMAGENETPANCREASSTATIC21645 RADIMAGENETPORTAL VEINSTATIC769 PROSTATETRUS (BAUM ET AL., 2023)PROSTATEVOLUME5037 PROSTATE-MRI-US-BIOPSY (NATARAJAN ET AL., 2020)PROSTATEVOLUME219484 MUPSD (SHAO & BRISBANE, 2024)PROSTATEVOLUME2910 ULTRASOUNDCASES.INFOREPRODUCTIVESTATIC3446 RADIMAGENETSPLEENSTATIC6520 ULTRASOUNDCASES.INFOSPLEENSTATIC1207 USANOTAISPLEENSTATIC61 RADIMAGENETTHYROIDSTATIC92598 SEGTHY (KR ̈ ONKE ET AL., 2022)THYROIDVOLUME136243 ULTRASOUNDCASES.INFOTHYROIDSTATIC2011 THYD2 (HOU ET AL., 2024)THYROIDSTATIC8489 DDTI (PEDRAZA ET AL., 2015)THYROIDSTATIC610 TG3K (GONG, 2026)THYROIDSTATIC3585 TN3K (GONG, 2026)THYROIDSTATIC3493 STANFORD THYROID CINE (YAMASHITA ET AL., 2022)THYROIDVIDEO17412 RADIMAGENETUTERUSSTATIC13312 YANG (TIANTIAN YANG, 2023)UTERUSSTATIC1973 A.2. Downstream Datasets Below, we provide a more detailed description of each downstream dataset, summarized in Table 4. We use the publicly available UltraBench tool to setup the datasets and corresponding training, validation and test splits. We contribute two more datasets, BUSBRA and TN5000, to UltraBench for a more comprehensive downstream evaluation. Annotated Ultrasound Liver (AUL): The Annotated Ultrasound Liver (AUL) dataset is a collection of 2D liver ultrasound images introduced for research on automated liver lesion analysis and mass characterization (Xu et al., 2023). It comprises images acquired from distinct patients and includes expert annotations delineating the liver boundary and, when present, lesion contours. Each image is additionally assigned a clinical label reflecting the presence and type of focal liver mass, enabling both region-aware and image-level learning. The dataset spans a wide range of image resolutions and visual appearances, capturing realistic variability in ultrasound acquisition and anatomy. We use AUL for a supervised image-level mass classification downstream task with three classes: malignant mass, benign mass, and normal (no mass). The full dataset contains 735 images in total. We follow a fixed split consisting of 529 images for training, 59 for validation, and 147 for testing, ensuring that all splits remain patient-independent and preserve the original class distribution. Butterfly: The Butterfly dataset was released by Butterfly Network for the 2018 MIT Grand Hack to promote research in point-of-care ultrasound understanding (Butterfly Network, 2018). It comprises ultrasound images acquired using the Butterfly iQ device from multiple anatomical regions across 31 patients, reflecting the diversity and variability encountered in real-world bedside imaging. Each image is labeled according to the anatomical site being examined, enabling supervised learning for anatomical recognition in ultrasound. We use the Butterfly dataset for a supervised organ classification downstream task with nine classes: Morison’s pouch, bladder, heart (parasternal long-axis view), heart (four-chamber view), heart (two-chamber view), inferior vena cava, carotid artery, lungs, and thyroid. The dataset contains a total of 41,076 images, with 28,053 images used for training, 6,272 for validation, and 6,751 for testing. All splits are patient-independent to prevent data leakage and ensure fair evaluation. 16 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Fatty Liver: The Fatty Liver dataset consists of B-mode liver ultrasound images introduced for studying automated detection of hepatic steatosis and non-alcoholic fatty liver disease (NAFLD) (Byra et al., 2018). The dataset includes images acquired from multiple patients with expert clinical labeling based on the presence of fatty infiltration in the liver parenchyma. It has been widely used to evaluate representation learning and transfer learning methods for liver ultrasound analysis, particularly in low-data clinical settings. We use this dataset for a supervised image-level fatty liver discrimination downstream task with two classes: normal liver and NAFLD. The dataset contains a total of 550 images, drawn from 55 patients. We use 390 images for training, 50 for validation, and 110 for testing. All splits are patient-independent to ensure a fair assessment of generalization performance. GBCU: The Gallbladder Cancer Ultrasound (GBCU) dataset was introduced to support research on automated gallbladder lesion characterization from abdominal ultrasound images (Basu et al., 2022). It consists of expert-annotated images collected from a diverse patient cohort and labeled according to underlying pathology, enabling supervised learning for clinically relevant gallbladder cancer assessment. The dataset captures a range of normal and diseased presentations and has been used to benchmark deep learning methods for malignancy discrimination in ultrasound. We use GBCU for a supervised image-level lesion malignancy classification downstream task with three classes: normal, benign, and malignant. The full dataset contains 1,255 images in total. We use 1,019 images for training, 114 for validation, and 122 for testing. The original dataset provides a fixed train–test split with no patient overlap; the validation set is created by further splitting the training data, as patient identifiers are not available to enforce strict separation. MMOTU: The Multi-Modality Ovarian Tumor Ultrasound (MMOTU) dataset was introduced to facilitate research on ovarian tumor analysis using ultrasound and contrast-enhanced ultrasound imaging (Zhao et al., 2023). It provides expert annotations and diagnostic labels for a diverse set of ovarian pathologies, supporting both localization and classification tasks. In this work, we consider only the 2D ultrasound images and associated labels, focusing on image-level learning for ovarian tumor characterization. We use MMOTU for a supervised multi-class image-level tumor classification downstream task with eight classes: chocolate cyst, serous cystadenoma, teratoma, theca cell tumor, simple cyst, normal ovary, mucinous cystadenoma, and high-grade serous cystadenocarcinoma. The dataset contains 1,469 images in total. We use 800 images for training, 200 for validation, and 469 for testing. The original dataset provides a fixed train–test split, and the validation set is obtained by further splitting the training data, as patient identifiers are not available to enforce strict patient-level separation. POCUS: The Point-of-Care Ultrasound (POCUS) dataset was introduced to support automated diagnosis of lung pathologies, particularly COVID-19, from bedside ultrasound imaging (Born et al., 2021). It aggregates lung ultrasound data acquired using convex and linear probes from multiple sources, reflecting real-world variability in point-of-care acquisition. The dataset has been widely used to benchmark deep learning methods for pulmonary disease recognition in ultrasound under heterogeneous data conditions. We use POCUS for a supervised image-level lung pathology classification downstream task with three classes: healthy, pneumonia, and COVID-19. Following the preprocessing protocol described in the original work, we use 29 convex-probe images and extract frame-level samples from 124 convex-probe videos while grouping frames by video to avoid data leakage across splits. The resulting dataset contains 2,064 images in total, with 1,444 images used for training, 177 for validation, and 443 for testing, ensuring strict separation at the video level. BUSBRA: The BUS-BRA (Breast Ultrasound Brazil) dataset was introduced to support the development and standardized evaluation of computer-aided diagnosis systems for breast ultrasound imaging, with an emphasis on clinically meaningful annotations and reproducible benchmarking (G ́ omez-Flores et al., 2024). It consists of anonymized breast ultrasound images acquired from female patients during routine clinical examinations and includes expert-provided lesion delineations and biopsy-confirmed pathology labels. In addition to image-level diagnostic labels, the dataset incorporates BI-RADS assessments assigned by an experienced ultrasonographer, enabling research across multiple levels of breast cancer risk stratification and lesion characterization. The dataset has been designed to facilitate fair comparison of learning-based methods by providing well-defined partitions and comprehensive annotations commonly required in breast CAD pipelines. We use BUSBRA for a supervised image-level tumor malignancy classification downstream task with two classes: benign (722) and malignant (342). The dataset contains 1,064 images in total. We use 765 images for training, 86 for validation, and 213 for testing. TN5000: The TN5000 dataset was introduced to enable large-scale research on automated thyroid nodule analysis from ultrasound imaging, with a particular emphasis on malignancy assessment (Zhang et al., 2025a). It comprises expertly curated ultrasound images of thyroid nodules collected through a rigorous process of data selection and annotation. Each image is paired with structured annotations following a PASCAL VOC–compatible format, where nodules are explicitly labeled and localized, allowing both region-aware modeling and image-level learning. Diagnostic labels are derived from 17 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Table 4. Comprehensive statistics of the downstream datasets, including organ, task description, number of classes and data splits. DATASETORGANTASK DESCRIPTIONCLASSESTOTALTRAINVALTEST TN5000THYROIDNODULE MALIG.2500035005001000 FATTY LIVERLIVERFATTY LIVER DIS.255039050110 POCUSLUNGPNEUMONIA, COVID320641444177443 BUTTERFLYMULTIORGAN DETECTION9410762805362726751 GBCUGALLBLADDERLESION MALIG.312551019114122 AULLIVERMASS MALIG.373552959147 MMOTUOVARYTUMOR MALIG.81469800200469 BUSBRABREASTTUMOR MALIG.2106476586213 Figure 5 clinical assessment, and the dataset was intentionally constructed to reduce bias by maintaining a relatively balanced representation of benign and malignant cases at scale. TN5000 has been widely adopted for benchmarking deep learning approaches for thyroid cancer detection due to its size, standardized organization, and clear labeling scheme. We use TN5000 for a supervised image-level thyroid nodule malignancy classification downstream task with two classes: benign and malignant. The dataset contains a total of 5,000 images, including 3,572 malignant nodules and 1,428 benign nodules. We use 3,500 images for training, 500 for validation, and 1,000 for testing. B. US region conditioning (USrc) B.1. Image and USrc Mask Examples In Figure 5, we have picked 8 datasets to include examples of images and the corresponding USrc mask. The USrc mask generated by our image processing algorithm is outlined in green and demonstrates the quality of the USrc mask generation. 18 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Table 5. Shared pre-training hyperparameters for US-JEPA and USrc-JEPA. The models differ only in their use of ultrasound region conditioning. CONFIGURATIONVALUE Architecture ENCODER BACKBONEVIT-BASE PREDICTOR DEPTH12 PREDICTOR EMBEDDING DIM384 PREDICTOR TARGET DIM768 Data & Augmentation INPUT SIZE224× 224 CROP SCALE[0.6, 1.0] AUGMENTATIONSH/V FLIP, BLUR, ARTIFICAL SPECKLE, CONTRAST BATCH SIZE128 Masking (JEPA) PATCH SIZE16 CONTEXT MASKS (NUM, SCALE)1, [0.85, 1.0] TARGET MASKS (NUM, SCALE)4, [0.075, 0.125] ASPECT RATIO[0.75, 1.5] MIN. PATCHES KEPT10 Optimization OPTIMIZERADAMW TOTAL EPOCHS100 WARMUP EPOCHS10 BASE LEARNING RATE (LR)5.0× 10 −5 START→ FINAL LR5.0× 10 −6 → 5.0× 10 −7 LR SCHEDULECOSINE DECAY WITH LINEAR WARMUP WEIGHT DECAY→ FINAL WD0.04→ 0.4 EMA MOMENTUM[0.996, 1.0] C. Training Details C.1. Pretraining Parameters We exclusively hold out 5% of the entire pretraining dataset to use as validation. For our downstream experiments, we evaluate the epoch checkpoint that gets the lowest validation loss during pretraining. All pretraining parameters are detailed in Table 5. C.2. Downstream Parameters We perform downstream evaluation by training a randomly initialized linear layer on top of the frozen encoder backbone. We utilize early stopping based on the validation loss with a patience of 15 epochs to prevent overfitting. The remaining parameters are detailed in Table 6. D. Few-Shot Scaling for Linear Probe Extended In Figure 6 we include the comprehensive few-shot scaling experiments. For all of our baselines we perform few-shot scaling across five random seeds to see how performance changes on each downstream. In the main paper, we only compare against URFM and USFM because they were the most competitive baselines. E. Robustness to Domain-Specific Corruption Details E.1. Extended Results In Figure 7 we include the complete results for the OOD robustness experiment. We show the performance degradation for US-JEPA, USrc-JEPA, URFM and USFM with worsening levels of blur, contrast reduction and artifical speckle corrupting downstream test set images during probe evaluation. 19 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Table 6. Hyperparameters for downstream linear probing on UltraBench datasets. CONFIGURATIONVALUE OPTIMIZERADAMW BASE LEARNING RATE1.0× 10 −3 WEIGHT DECAY1.0× 10 −4 BATCH SIZE32 INPUT SIZE224× 224 LOSS FUNCTIONCROSS-ENTROPY LR SCHEDULECOSINE ANNEALING MAX TRAINING EPOCHS150 EARLY STOPPING PATIENCE15 EPOCHS E.2. Formulas for Corruptions To simulate ultrasound imaging artifacts, we define three corruption functionsC(I,ε)whereIis the input image andεis the severity level ranging from 1 to 3. Gaussian Blur: Gaussian blur is modeled as the convolution of the imageIwith a Gaussian kernelG, where the standard deviation σ is controlled by the severity ε: I blur = I ∗ G σ where σ = ε The discrete kernel size for G is adjusted to 2⌊2ε⌋ + 1 to ensure the Gaussian distribution is properly captured. Contrast Depletion: This corruption shrinks intensities toward the median brightnessμ med of the pixels defined by the USrc mask: I contrast = μ med + α(ε)· (I − μ med ) where the depletion factorαdecreases as severity increases (α∈0.7, 0.5, 0.3), effectively making anatomical boundaries less visible. Correlated Speckle Noise: Speckle is modeled as multiplicative noise. The raw Gaussian noiseηis spatially correlated via a smoothing kernel K before being applied. I speckle = I · (1 + η corr ) The noise characteristics are functions of severity ε: • Magnitude: η ∼N (0,σ 2 noise ) where σ noise = 0.35ε. • Correlation:η corr = η∗ K size (ε), where the kernel sizeK size increases withεto simulate larger speckle grains at higher severity. E.3. Qualitative Examples In Figure 8, we have included a qualitative example of the varying severities for each corruption on one example image from the AUL downstream dataset. 20 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound 30 40 50 60 70 Mean Macro F1 AUL 40 45 50 55 60 65 70 75 BUSBRA US-JEPA USrc-JEPA URFM USFM DINOv3 I-JEPA USF-MAE UltraSAM SAMUS EchoCare 65 70 75 80 85 90 95 Mean Macro F1 BUTTERFLY 40 50 60 70 80 90 FATTY LIVER 20 30 40 50 60 70 GBCU 1%5%10%50%100% Percentage of Training Labels 10 20 30 40 50 Mean Macro F1 MMOTU 1%5%10%50%100% Percentage of Training Labels 20 30 40 50 60 70 80 90 POCUS 1%5%10%50%100% Percentage of Training Labels 45 50 55 60 65 70 75 TN5000 Figure 6. Few-shot classification robustness across all datasets. Here the test set evaluation for linear probe tuned with varying percentages of training data are shown for all models. 21 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound 0 25 50 75 100 BlurContrastSpeckle 0 25 50 75 100 0 25 50 75 100 0123 0 25 50 75 100 01230123 AUL BUTTERFLY GBCU TN5000 Severity Mean Macro F1 US-JEPAUSrc-JEPAURFMUSFM Figure 7. OOD robustness results for remaining 4 downstream datasets. 22 US-JEPA: A Joint Embedding Predictive Architecture for Medical Ultrasound Severity - 0 (Normal)Severity - 1Severity - 2Severity - 3 Blur Contrast Speckle Figure 8. Different corruptions at three severity levels applied to a liver ultrasound from AUL dataset. 23