Paper deep dive
A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies
Fangyijie Wang, Jianjun Yu, Wentao Shi, Haixia Huang, Ran Shi, Guénolé Silvestre, Kathleen M. Curran
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 3:41:41 AM
Summary
The paper presents a clinician-centered pipeline designed for remote annotation and evaluation of ultrasound AI models. Unlike traditional annotation platforms that focus on dataset labeling, this pipeline integrates blinded model comparison, multi-rater participation, and automated statistical analysis (e.g., Spearman correlation, Kendall's $\tau$) through a lightweight, browser-based interface. This approach preserves data governance by hosting medical images on a centralized researcher server while allowing remote clinicians to perform segmentation and preference ranking. The system was validated using fetal ultrasound datasets (HC18 and ES-TCB) with six raters of varying expertise, demonstrating that the pipeline effectively supports reproducible human-AI evaluation and captures clinical usability metrics.
Entities (9)
Relation Signals (5)
Clinician-Centered Pipeline â calculates â Spearman correlation
confidence 100% · The system automatically generated Spearman correlation, Kendall's $\tau$, and top-1 selection statistics.
Clinician-Centered Pipeline â supports â Blinded Model Comparison
confidence 100% · The proposed pipeline... enables clinicians to perform annotation, blinded ranking, and review
Clinician-Centered Pipeline â validatedon â HC18
confidence 100% · We validate the pipeline in a fetal ultrasound segmentation study... We used two public fetal ultrasound datasets, HC18 and ES-TCB.
Clinician-Centered Pipeline â validatedon â ES-TCB
confidence 100% · We used two public fetal ultrasound datasets, HC18 and ES-TCB.
CVAT â lacks â Blinded Model Comparison
confidence 90% · Existing medical image platforms... lack integrated support for blinded model comparison
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Clinician-centered evaluation is critical for validating medical AI systems, especially in ultrasound imaging where quantitative metrics do not always capture clinical usability. Existing medical image platforms primarily focus on dataset labeling. They lack integrated support for blinded model comparison and reproducible evaluation workflows. We present a clinician-centered pipeline for remote annotation and evaluation in ultrasound AI studies. The proposed pipeline uses a centralized server and lightweight browser interfaces to enable clinicians to perform annotation, blinded ranking, and review without local dataset downloads. The pipeline also supports multi-rater participation, centralized result aggregation, and automated statistical analysis. We validate the pipeline in a fetal ultrasound segmentation study with six raters spanning expert, generalist, and non-expert experience levels. The system automatically generated Spearman correlation, Kendall's $\tau$, and top-1 selection statistics. Results indicated moderate to strong agreement across experts and other groups. The blinded evaluation results showed a tendency for later active learning models to be preferred. These outcomes suggest that the pipeline can support clinician-centered annotation and reproducible human-\ac{AI} evaluation studies in ultrasound imaging. The proposed pipeline is available on \href{this https URL}{GitHub}.
Tags
Links
- Source: https://arxiv.org/abs/2606.19174v1
- Canonical: https://arxiv.org/abs/2606.19174v1
Trouble viewing inline? Open PDF directly â
Full Text
30,419 characters extracted from source content.
Expand or collapse full text
A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies Fangyijie Wang 1,2,â [0009â0003â0427â368X] , Jianjun Yu 3,â , Wentao Shi 4 , Haixia Huang 5 , Ran Shi 4,â , GuĂ©nolĂ© Silvestre 1,6 , and Kathleen M. Curran 1,2â[0000â0003â0095â9337] 1 Research Ireland Centre for Research Training in Machine Learning 2 School of Medicine, University College Dublin, Dublin, Ireland 3 The Third Peopleâs Hospital of Zhenjiang City, Zhenjiang, China 4 Zhenjiang Maternal and Child Health Hospital, Zhenjiang, China 5 The Fifth Peopleâs Hospital of Zhenjiang City, Zhenjiang, China 6 School of Computer Science, University College Dublin, Dublin, Ireland 408418010@q.com,kathleen.curran@ucd.ie â These authors contributed equally to this work. Abstract. Clinician-centered evaluation is critical for validating med- ical Artificial Intelligence (AI) systems, especially in ultrasound imag- ing where quantitative metrics do not always capture clinical usability. Existing medical image platforms primarily focus on dataset labeling. They lack integrated support for blinded model comparison and repro- ducible evaluation workflows. We present a clinician-centered pipeline for remote annotation and evaluation in ultrasound AI studies. The pro- posed pipeline uses a centralized server and lightweight browser interfaces to enable clinicians to perform annotation, blinded ranking, and review without local dataset downloads. The pipeline also supports multi-rater participation, centralized result aggregation, and automated statistical analysis. We validate the pipeline in a fetal ultrasound segmentation study with six raters spanning expert, generalist, and non-expert expe- rience levels. The system automatically generated Spearman correlation, Kendallâs Ï, and top-1 selection statistics. Results indicated moderate to strong agreement across experts and other groups. The blinded evalua- tion results showed a tendency for later active learning models to be pre- ferred. These outcomes suggest that the pipeline can support clinician- centered annotation and reproducible human-AI evaluation studies in ultrasound imaging. The proposed pipeline is available on GitHub. Keywords: Ultrasound· Human-centered AI· Evaluations· Annota- tions· Remote Collaboration 1 Introduction Deep learning has become widely adopted in medical image analysis for segmen- tation, classification, detection, and diagnostic support [7,19,27,14,1]. In ultra- sound imaging, recent AI systems have shown effectiveness for fetal biometrics, â Corresponding Author arXiv:2606.19174v1 [cs.HC] 17 Jun 2026 2F. Wang et al. anatomical structure segmentation, and disease assessment [3,33,25,12,16,26,15]. Despite these advances, evaluation remains dominated by quantitative metrics such as Dice score, accuracy, and Hausdorff distance. While standardized, these metrics still fail to capture a modelâs clinical usability in scenarios where image quality is variable and anatomical boundaries are ambiguous [34,5]. Clinician-centered evaluation and reader studies are therefore increasingly important for understanding how AI predictions align with clinical judgment and workflow requirements [28,21]. Existing medical image platforms commonly support dataset creation and annotation, and they lack integrated support for clinician-in-the-loop AI validation studies [8,29,9,10]. Moreover, clinician evalu- ation protocols are often developed independently for each project, which hin- ders workflow consistency and reproducibility across research groups and insti- tutions [23,20]. Practical deployment poses an additional challenge. Medical imaging data are frequently constrained by institutional governance, privacy regulations, and data-sharing policies, which complicates remote multi-center evaluation [4]. These constraints limit clinician accessibility and make large-scale collaborative vali- dation more difficult. Therefore, lightweight and standardized frameworks are needed to enable efficient remote collaboration while keeping medical imaging data centrally hosted without direct distribution. In this work, we design a clinician-centered AI pipeline to support clinician participation throughout annotation and evaluation. The proposed pipeline uses a centralized server and a lightweight web-based interface so clinicians can in- teract with the system without local software installation or dataset download. It also conducts automated statistical analysis and report generation for repro- ducible human-AI validation studies. The main contributions of this work are summarized as follows: â A lightweight, clinician-centered pipeline for medical imaging AI evaluation studies that integrates annotation and preference-ranking comparison. â A remote evaluation strategy that preserves data governance by keeping raw medical images centrally hosted while enabling collaborative clinician participation. â A standardized workflow for blinded model comparison, multi-rater prefer- ence analysis, and reproducible statistical reporting. â A demonstration study in fetal ultrasound involving clinicians with varying expertise levels and active learning model comparisons. 2 Related Work 2.1 Medical Image Annotation Platforms Several annotation tools have been developed to support medical image analysis and dataset construction. General annotation tools such as CVAT [8] and La- bel Studio [29] provide flexible interfaces for image labeling, segmentation, and Title Suppressed Due to Excessive Length3 Table 1: Comparison between existing annotation platforms and the proposed pipeline.â: supported,â: not directly supported, ~: partial support. FeatureCVAT Label Studio MONAI Label Ours Browser-based annotationâ Medical imaging support~~â Annotation managementâ Remote deploymentâ Blinded model comparisonââ Preference ranking workflowââ Multi-rater agreement analysisââ Automated statistical reportingââ collaborative dataset management. In medical imaging, specialized tools includ- ing MONAI Label [9,10] and ITK-SNAP [32] have further enabled interactive segmentation and AI-assisted annotation workflows. These frameworks have con- tributed substantially to the development of annotated datasets for training deep learning models. However, existing tools are primarily designed for dataset an- notation and labeling rather than clinician-centered AI evaluation. They focus on generating Ground Truth (GT) annotations or correcting model predictions, with limited support for blinded comparison studies between multiple AI mod- els. In particular, existing tools generally do not provide standardized workflows for clinician ranking, preference analysis, or multi-rater agreement assessment, which are important for evaluating the clinical relevance of modelsâ predictions. Moreover, in medical imaging, data sharing is often restricted by privacy and institutional policies [4]. Existing tools provide limited support for lightweight remote evaluation workflows, which enable collaborative clinician participation without direct dataset distribution. As shown in Table 1, existing platforms primarily focus on annotation and dataset creation, whereas our pipeline additionally supports blinded model com- parison, clinician preference ranking, multi-rater agreement analysis, and auto- mated statistical reporting for clinician-centered AI evaluation studies. 2.2 Human-Centered AI Evaluation Human-centered AI evaluation has become increasingly important in medical im- age analysis, where quantitative metrics may not fully reflect clinical usability or diagnostic relevance [22]. Reader studies provide an important way to assess whether AI systems are consistent with clinician interpretation and whether AI can support real diagnostic workflows [24,18,31,13]. Recent studies have also highlighted the importance of human-AI collaboration and clinician trust for reliable deployment of medical AI systems [31,13,21]. In particular, preference alignment between clinicians and AI models has emerged as an important di- rection for understanding whether improvements in quantitative performance are consistent with human judgment and clinical expectations [22,28]. However, 4F. Wang et al. Fig. 1: Overview of the proposed clinician-centered pipeline for remote anno- tation and evaluation in ultrasound imaging studies. The framework supports study configuration, browser-based clinician interaction, blinded model compar- ison, annotation workflows, and automated multi-rater statistical analysis. existing studies often use evaluation protocols that provide limited support for standardized blinded comparison and reproducible statistical analysis. 3 Pipeline Design 3.1 Overall Pipeline An overview of the proposed clinician-centered annotation and evaluation pipeline is presented in Fig. 1. This pipeline consists of the following components: (1) re- searcher server, (2) clinician client, (3) secure image streaming, (4) annotation and ranking modules, (5) result aggregation, and (6) statistical analysis. The researcher server is responsible for study preparation, centralized re- source management and statistical analysis. Researchers upload ultrasound im- ages, AI model predictions, and optional reference annotations to the server. The server also manages study configuration, including evaluation mode, clini- cian groups, randomized model ordering, and blinded comparison settings. All images and model outputs are hosted centrally to avoid direct distribution of medical datasets. A lightweight web-based client is used to allow clinicians to interact with the system. This client is executable on the server without installation. Ultrasound images and segmentation overlays are streamed directly from the centralized server to the browser interface. This design enables remote clinicians-in-the-loop while simplifying collaborative evaluation across multiple users and institutions. Title Suppressed Due to Excessive Length5 Fig. 2: Overview of the remote deployment workflow for clinician and researcher collaboration. The pipeline also supports both annotation and evaluation workflows. In the annotation stage, clinicians can create and edit segmentation masks using an interactive tool [11]. In the ranking stage, clinicians are presented with multiple anonymized model outputs and asked to rank these outputs based on overall quality and clinical usability. The model outputs are ordered randomly and in- dependently for each clinician to ensure blinded evaluation. Cliniciansâ annotations, rankings, and preference scores are collected and saved on the researcher server. The statistical analysis module subsequently computes agreement metrics, such as Spearman correlation, Kendallâs Ï, top- 1 selection frequency, and inter-rater agreement statistics. These results provide a standardized pipeline for clinician-centered validation and human-centered AI evaluation studies. 3.2 Researcher Server and Remote Deployment Fig. 2 presents the researcher server functionality and remote deployment work- flow. The proposed pipeline uses a centralized deployment strategy where all images, AI model predictions, annotations, and study settings are hosted on a researcher server. Researchers can upload data, define annotation tasks, config- ure rater groups, and select evaluation modes through the server. Clinicians ac- cess the system remotely using a lightweight browser interface without installing additional software. During interaction, images and segmentation overlays are streamed directly from the server. All clinician annotations and ranking results are automatically saved on the server, while a Python program performs subse- quent statistical analysis and report generation. Researchers manage the deployment process and upload study data to the researcher server through a network connection using the Secure Shell (SSH) 6F. Wang et al. Fig. 3: The lightweight interface for clinician annotation and evaluation via browser. (a) The annotation interface supports multiple annotation shapes and exports annotations in JSON format. (b) The ranking interface presents ul- trasound images with overlaid segmentation masks in randomized order. The ranking results are automatically saved as text files. protocol. The annotation and evaluation interfaces are lightweight and can be easily modified to support different clinical tasks and study requirements. 3.3 Annotation and Ranking Interfaces Fig. 3 shows screenshots of the annotation and ranking interfaces. The anno- tation interface supports multiple annotation shapes and interactive editing, as well as JSON-formatted annotations. The ranking interface displays images with overlaid segmentation masks and allows clinicians to rank model outputs using scores from 1 to 5. The ranking score of each case is saved in a plain text file. All annotations and ranking results are automatically saved and aggregated by a Python program running on the researcher server. Title Suppressed Due to Excessive Length7 3.4 Result Aggregation and Statistical Analysis All clinician annotations and rankings are automatically collected and stored on the researcher server during the study. A Python program subsequently aggre- gates the results and generates statistical reports for evaluation. The proposed pipeline supports multiple agreement and ranking analyses, including Spearman correlation, Kendallâs coefficient, top-1 selection frequency, and inter-rater agree- ment statistics. These outputs provide a standardized workflow for clinician- centered evaluation and human-AI agreement analysis. 4 Experimental Demonstration This section demonstrates the clinician-centered pipeline in a fetal ultrasound study. We used two public fetal ultrasound datasets, HC18 and ES-TCB. HC18 contains fetal head ultrasound images acquired during routine obstetric exam- inations [17], while ES-TCB contains trans-cerebellum fetal ultrasound images collected in Spain [6,2]. Building on a semi-supervised learning framework [30], we developed an active learning pipeline and evaluated five segmentation mod- els (M1âM5) from different active learning iterations via the proposed clinician- centered pipeline, as shown in Fig. 4. Fig. 4: Overview of the clinician-centered pipeline for our ultrasound AI studies. The workflow has active learning, sampling strategy, model training and infer- ence, and clinician-in-the-loop evaluation. A total of six raters participated remotely through the browser interface. There were two expert obstetric sonographers (E1, E2), two general ultrasound 8F. Wang et al. Table 2: Details of the raters recruited for the human evaluation study. Rater ID RoleSpecialtyExperience (Y) Clinical Group E1 SonographerObstetrics15Specialist E2 SonographerObstetrics5Specialist G3 Sonographer Generalist Ultrasound18Generalist G4 Sonographer Generalist Ultrasound18Generalist NE5 Non-expertN/A0Non-expert NE6 Non-expertN/A0Non-expert sonographers (GE3, GE4), and two non-expert participants (NE5, NE6). The details of these raters are presented in Table 2. The pipeline enabled all raters to complete the study without local installation or direct dataset download. The browser-based interface enabled efficient remote participation. On average, clini- cians required approximately 45â50 seconds to annotate an ultrasound image and around 30 seconds per case for blinded ranking. These observations suggest that the proposed pipeline can support practical clinician participation with relatively low interaction overhead. For each dataset, 30 cases were randomly selected for rating, resulting in 60 cases in total. Each case contained five anonymized model predictions randomly presented for blinded comparison. In the end, this study collected 300 model rankings from each rater. Fig. 5 presents an example of the statistical analysis generated by the pro- posed pipeline in our demonstration experiment. In both the HC18 and ES-TCB datasets, later models (M3-M5) were selected more frequently during blinded evaluation, while earlier models (e.g., M1) were rarely preferred. The pipeline also supports inter-group agreement analysis across expert, generalist, and non- expert raters using Spearman correlation. Moderate to strong positive correla- tions were observed across different rater groups, indicating consistent preference patterns during remote clinician evaluation. Table 3 shows an example of inter-rater agreement analysis using Kendallâs Ï. Moderate agreement was observed across both HC18 and ES-TCB datasets, with mean agreement values of 0.53 and 0.55, respectively. The pipeline also reports variability statistics, including standard deviation and interquartile range, which indicate differences in agreement across evaluation cases. A limited proportion of images achieved high inter-rater agreement (Ï > 0.7), accounting for 20.0% in HC18 and 26.7% in ES-TCB. Lower agreement was observed in challenging cases characterized by ambiguous anatomical boundaries, poor ultrasound image quality, or minimal visual differences between model predictions. Table 3: Inter-rater agreement (Kendallâs Ï) across datasets. Dataset Mean Median Std IQR Min Max Ï > 0.7 (%) Ï < 0.4 (%) HC180.530.51 0.18 0.15 0.12 0.8720.016.7 ES-TCB 0.550.56 0.22 0.21 0.12 0.8526.720.0 Title Suppressed Due to Excessive Length9 Fig. 5: Top-1 selection frequency across models (M1âM5) and inter-group agree- ment measured by Spearman correlation among experts, generalists, and non- experts on HC18 and ES-TCB datasets. ***: p < 0.001. **: p < 0.01. *: p < 0.05. 5 Discussion This paper introduced a clinician-centered pipeline for remote annotation and evaluation in ultrasound AI studies. The pipeline integrates clinician partici- pation, blinded model comparison, annotation workflows, and automated sta- tistical analysis within a unified web-based interface. Unlike existing annotation platforms that primarily focus on dataset labeling, our design supports clinician- in-the-loop AI validation and reproducible multi-rater evaluation studies while preserving data governance through centralized hosting. The pipeline offers translational value for medical AI validation studies. By enabling remote clinician participation, it facilitates reader studies and multi- center validation experiments without requiring clinicians to download raw datasets. The integrated statistical analysis module provides consistent reporting of agree- ment metrics and ranking outcomes, making the pipeline suitable for preference alignment studies and clinical usability assessments. Although a formal usability study was outside the scope of this work, the demonstration experiment showed that clinicians were able to complete anno- tation and ranking tasks efficiently through the browser interface. The average annotation time was approximately 45â50 seconds per image, while blinded rank- ing required approximately 30 seconds per image. 10F. Wang et al. Several limitations remain. The current implementation assumes pre-established collaboration between researchers and clinicians. Therefore, the implementation does not yet include enterprise-level authentication, role-based access control, or audit logging. Moreover, the additional usability measures, such as clinician sat- isfaction, perceived workload, and user experience, are not investigated in this work. Lastly, the current implementation is optimized for ultrasound imaging and has not yet been extensively evaluated across other modalities. Future work will first incorporate secure user authentication, permission man- agement, and enhanced governance features to support large-scale multi-center clinical studies better. Then we will investigate additional usability measures to improve clinician experience. Afterward, we will extend support to additional medical imaging modalities. 6 Conclusion We presented a secure, clinician-centered, end-to-end evaluation pipeline that enables reproducible remote human-AI annotation and validation studies in ul- trasound imaging without direct medical data sharing. The proposed pipeline is lightweight and easily deployable for efficient clinician participation, blinded model comparison, multi-rater evaluation, and statistical analysis. A demonstra- tion study showed that the proposed pipeline enables clinicians to annotate fetal head segmentation labels, supports six raters in conducting human-AI evalua- tion, and performs agreement analysis without direct dataset distribution. We hope this work can facilitate collaborative clinician-in-the-loop studies and sup- port the development of more clinically aligned medical AI systems. Acknowledgments. This work was funded by Taighde Ăireann â Research Ireland through the Research Ireland Centre for Research Training in Machine Learning (18/CRT/6183). Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article. References 1. Aggarwal, R., Sounderajah, V., Martin, G., Ting, D.S.W., Karthikesalingam, A., King, D., Ashrafian, H., Darzi, A.: Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis. npj Digital Medicine 4(1), 65 (Apr 2021) 2. Alzubaidi, M., Agus, M., Makhlouf, M., Anver, F., Alyafei, K., Househ, M.: Large- scale annotation dataset for fetal head biometry in ultrasound images. Data in Brief 51, 109708 (Dec 2023) 3. Bano, S., Dromey, B., Vasconcelos, F., Napolitano, R., David, A.L., Peebles, D.M., Stoyanov, D.: AutoFB: Automating fetal biometry estimation from standard ultra- sound planes. In: Medical Image Computing and Computer Assisted Intervention â MICCAI. p. 228â238. Springer International Publishing (2021) Title Suppressed Due to Excessive Length11 4. Bell, L.C., Shimron, E.: Sharing data is essential for the future of AI in medical imaging. Radiology: Artificial Intelligence 6(1), e230337 (Jan 2024) 5. Boumeridja, H., Ammar, M., Alzubaidi, M., Mahmoudi, S., Benamer, L.N., Agus, M., Househ, M., Lekadir, K., El Habib Daho, M.: Enhancing fetal ultrasound im- age quality and anatomical plane recognition in low-resource settings using super- resolution models. Scientific Reports 15(1), 8376 (Mar 2025) 6. Burgos-Artizzu, X.P., Coronado-GutiĂ©rrez, D., Valenzuela-Alcaraz, B., Bonet- Carne, E., Eixarch, E., Crispi, F., GratacĂłs, E.: Evaluation of deep convolutional neural networks for automatic classification of common maternal fetal ultrasound planes. Scientific Reports 10(1), 10200 (Dec 2020) 7. Chen, X., Wang, X., Zhang, K., Fung, K.M., Thai, T.C., Moore, K., Mannel, R.S., Liu, H., Zheng, B., Qiu, Y.: Recent advances and clinical applications of deep learning in medical image analysis. Medical Image Analysis 79, 102444 (Jul 2022) 8. CVAT.ai Corporation: Computer vision annotation tool (cvat) (Nov 2023), https: //github.com/cvat-ai/cvat 9. Diaz-Pinto, A., Alle, S., Ihsani, A., Asad, M., Nath, V., PĂ©rez-GarcĂa, F., Mehta, P., Li, W., Roth, H.R., Vercauteren, T., Xu, D., Dogra, P., Ourselin, S., Feng, A., Cardoso, M.J.: MONAI Label: A framework for AI-assisted Interactive Labeling of 3D Medical Images. arXiv e-prints (2022) 10. Diaz-Pinto, A., Mehta, P., Alle, S., Asad, M., Brown, R., Nath, V., Ihsani, A., Antonelli, M., Palkovics, D., Pinter, C., et al.: DeepEdit: Deep Editable Learning for Interactive Segmentation of 3D Medical Images. In: MICCAI Workshop on Data Augmentation, Labelling, and Imperfections. p. 11â21. Springer (2022) 11. Dutta, A., Zisserman, A.: The VIA annotation software for images, audio and video. In: Proceedings of the 27th ACM International Conference on Multimedia. M â19, ACM, New York, NY, USA (2019). https://doi.org/10.1145/3343031. 3350535 12. Fiorentino, M.C., Villani, F.P., Di Cosmo, M., Frontoni, E., Moccia, S.: A review on deep-learning algorithms for fetal ultrasound-image analysis. Medical Image Analysis 83, 102629 (2023). https://doi.org/10.1016/j.media.2022.102629 13. Gommers, J., Hernström, V., Josefsson, V., Sartor, H., Schmidt, D., Hjelmgren, A., Larsson, A.M., Hofvind, S., Andersson, I., Rosso, A., Hagberg, O., LĂ„ng, K.: Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a ran- domised, controlled, non-inferiority, single-blinded, population-based, screening- accuracy trial. The Lancet 407(10527), 505â514 (Jan 2026) 14. Groh, M., Badri, O., Daneshjou, R., Koochek, A., Harris, C., Soenksen, L.R., Doraiswamy, P.M., Picard, R.: Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nature Medicine 30(2), 573â583 (Feb 2024) 15. Guo, X., Alsharid, M., Zhao, H., Wang, Y., Lander, J., Papageorghiou, A.T., No- ble, J.A.: A visually grounded language model for fetal ultrasound understanding. Nature Biomedical Engineering (Jan 2026) 16. Guo, X., Men, Q., Noble, J.A.: MMSummary: Multimodal Summary Generation for Fetal Ultrasound Video . In: proceedings of Medical Image Computing and Computer Assisted Intervention â MICCAI 2024. vol. LNCS 15004. Springer Na- ture Switzerland (October 2024) 17. van den Heuvel, T.L.A., de Bruijn, D., de Korte, C.L., van Ginneken, B.: Auto- mated measurement of fetal head circumference using 2d ultrasound images (Jul 2018). https://doi.org/10.5281/zenodo.1327317 12F. Wang et al. 18. Jassim, G., Otoom, O., Nair, B., Hashem, J.: Performance of artificial intelligence in breast cancer screening programmes: a systematic review. BMJ Open 15(12) (2025). https://doi.org/10.1136/bmjopen-2025-111360 19. Kim, H.E., Cosa-Linan, A., Santhanam, N., Jannesari, M., Maros, M.E., Gans- landt, T.: Transfer learning for medical image classification: a literature review. BMC Medical Imaging 22(1), 69 (Apr 2022) 20. Livingston, L., Featherstone-Uwague, A., Barry, A., Barretto, K., Morey, T., Her- rmannova, D., Avula, V.: Reproducible generative artificial intelligence evaluation for health care: a clinician-in-the-loop approach. JAMIA Open 8(3), ooaf054 (Jun 2025) 21. Madan, Y., Perivolaris, A., Adams-McGavin, R.C., Jung, J.J.: Clinician interaction with artificial intelligence systems: a narrative review. Journal of Medical Artificial Intelligence 8 (2025) 22. Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A.P., Carass, A., Feldmann, C., Frangi, A.F., Full, P.M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B.A., MĂ€rz, K., Maier, O., Maier-Hein, K., Menze, B.H., MĂŒller, H., Neher, P.F., Niessen, W., Rajpoot, N., Sharp, G.C., Sirinukunwattana, K., Speidel, S., Stock, C., Stoyanov, D., Taha, A.A., van der Sommen, F., Wang, C.W., Weber, M.A., Zheng, G., Jannin, P., Kopp-Schneider, A.: Why rankings of biomedical image analysis competitions should be interpreted with care. Nature Communications 9(1), 5217 (Dec 2018) 23. Novak, A., Hollowday, M., Espinosa Morgado, A.T., Oke, J., Shelmerdine, S., Woznitza, N., Metcalfe, D., Costa, M.L., Wilson, S., Kiam, J.S., Vaz, J., Limphai- bool, N., Ventre, J., Jones, D., Greenhalgh, L., Gleeson, F., Welch, N., Mistry, A., Devic, N., Teh, J., Ather, S.: Evaluating the impact of artificial intelligence-assisted image analysis on the diagnostic accuracy of front-line clinicians in detecting frac- tures on plain x-rays (fract-ai): protocol for a prospective observational study. BMJ Open 14(9) (2024). https://doi.org/10.1136/bmjopen-2024-086061 24. Obuchowski, N.A., Zepp, R.C.: Simple steps for improving multiple-reader studies in radiology. AJR. American journal of roentgenology 166(3), 517â521 (1996) 25. PĆotka, S., WĆodarczyk, T., Klasa, A., Lipa, M., Sitek, A., TrzciĆski, T.: FetalNet: Multi-task deep learning framework for fetal ultrasound biometric measurements. In: Neural Information Processing. p. 257â265. Springer (2021) 26. PĆotka, S., Pustelnik, K., Szenejko, P., Ć»ebrowska, K., RzucidĆo-SzymaĆska, I., Szymecka-Samaha, N., ĆÄgowik, T., KosiĆska-KaczyĆska, K., Korzeniowski, P., BiliĆski, P., Khalil, A., Brawura-Biskupski-Samaha, R., IĆĄgum, I., SĂĄnchez, C.I., Sitek, A.: Direct estimation of fetal biometry measurements from ultrasound video scans through deep learning. American Journal of Obstetrics & Gynecology MFM 7(4) (Apr 2025) 27. Rayed, M.E., Islam, S.M.S., Niha, S.I., Jim, J.R., Kabir, M.M., Mridha, M.F.: Deep learning for medical image segmentation: State-of-the-art advancements and challenges. Informatics in Medicine Unlocked 47, 101504 (Jan 2024) 28. Reinke, A., Tizabi, M.D., Eisenmann, M., Maier-Hein, L.: Common pitfalls and recommendations for grand challenges in medical artificial intelligence. European Urology Focus 7(4), 710â712 (Jul 2021) 29. Tkachenko, M., Malyuk, M., Holmanyuk, A., Liubimov, N.: Label Studio: Data labeling software (2020-2025), https://github.com/HumanSignal/label-studio 30. Wang, F., Silvestre, G., Curran, K.M.: Leveraging information divergence for robust semi-supervised fetal ultrasound image segmentation. arXiv preprint arXiv:2509.06495 (2025) Title Suppressed Due to Excessive Length13 31. Warren, L.M., Venton, J., Young, K.C., Halling-Brown, M., Kelly, C.J., Wilson, M., Morigami, M., Khoo, L., Cunningham, D., Sidebottom, R., Reddy, M., Pu- rushothaman, H., Khodabakhshi, D., Honeyfield, L., Hujan, A., Stoycheva, T., Joiner, A., Chopra, R., Sy, A., Ward, D., Yang, L., Sayres, R., Golden, D., Mal- hotra, N., Mallya, R., Xi, L., Ogunleye, D., Purdy, C., Mackenzie, A., Thomas, S., Shetty, S., Gilbert, F.J., Darzi, A., Ashrafian, H.: Impact of using artificial intelli- gence as a second reader in breast screening including arbitration. Nature Cancer 7(3), 507â521 (Mar 2026) 32. Yushkevich, P.A., Piven, J., Cody Hazlett, H., Gimpel Smith, R., Ho, S., Gee, J.C., Gerig, G.: User-guided 3D active contour segmentation of anatomical struc- tures: Significantly improved efficiency and reliability. Neuroimage 31(3), 1116â 1128 (2006) 33. Zeng, Y., Tsui, P.H., Wu, W., Zhou, Z., Wu, S.: Fetal ultrasound image segmenta- tion for automatic head circumference biometry using deeply supervised Attention- Gated V-Net. Journal of Digital Imaging 34(1), 134â148 (Feb 2021) 34. Zhou, Z., Wang, Y., Guo, Y., Qi, Y., Yu, J.: Image quality improvement of hand- held ultrasound devices with a two-stage generative adversarial network. IEEE Transactions on Biomedical Engineering 67(1), 298â311 (2020). https://doi.org/ 10.1109/TBME.2019.2912986