Paper deep dive
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Haifan Gong, Shiyu Chen, Bodong Wang, Yuqi Wang, Shijie Wang, Guoliang You, Xinyu Xiong, Haowei Wang, Mingzhi Mao, Dexing Kong, Qinghua Liu, Wei Lou, Fei Chen, Guanbin Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/14/2026, 4:34:49 AM
Summary
The paper introduces ThyroidXAgent, an auditable, clinician-interactive agentic AI system for thyroid ultrasound diagnosis. It coordinates specialized tools for nodule segmentation, benign-malignant classification, lymph-node metastasis prediction, and report generation, storing outputs as an auditable case-level evidence record. Developed using OpenThyroidDB, the system achieved high performance metrics (Dice 87.21%, AUROC 0.9466) and improved physician accuracy and efficiency while allowing for clinician correction of intermediate evidence.
Entities (10)
Relation Signals (9)
ThyroidXAgent → improves → physician classification accuracy
confidence 95% · ThyroidXAgent improved physician classification accuracy
ThyroidXAgent → reduces → reporting time
confidence 95% · reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively.
ThyroidXAgent → reduces → segmentation time
confidence 95% · reduced segmentation and reporting time by 35.9 percent
ThyroidXAgent → uses → OpenThyroidDB
confidence 95% · The system was developed using OpenThyroidDB, a multicentre, multitask resource
ThyroidXAgent → evaluatedon → NHC-MISD-TUS
confidence 92% · evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort.
ThyClinScore → correlateswith → location-aware language-model judge
confidence 90% · ThyClinScore... showed the strongest correlation with a location-aware language-model judge.
ThyroidXAgent → outperforms → Gemini 2.5 Pro
confidence 90% · Gemini-2.5-Pro reached 0.616-0.687... ThyroidXAgent achieved a mean AUROC of 0.9466
ThyroidXAgent → outperforms → GPT-5
confidence 90% · GPT-5 reached AUROC of 0.611-0.774... ThyroidXAgent achieved a mean AUROC of 0.9466
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician-interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21 percent for nodule segmentation and a mean AUROC of 0.9466 for benign-malignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3 percent to 86.2 percent, and reduced segmentation and reporting time by 35.9 percent and 27.4 percent, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting.
Tags
Links
- Source: https://arxiv.org/abs/2608.12590v1
- Canonical: https://arxiv.org/abs/2608.12590v1
Trouble viewing inline? Open PDF directly →
Full Text
156,803 characters extracted from source content.
Expand or collapse full text
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting Haifan Gong 1,† , Shiyu Chen 1,† , Bodong Wang 2,† , Yuqi Wang 3,† , Shijie Wang 4 , Guoliang You 5 , Xinyu Xiong 1 , Haowei Wang 6 , Mingzhi Mao 2 , Dexing Kong 4 , Qinghua Liu 7,* , Wei Lou 8,* , Fei Chen 9,* , and Guanbin Li 1,* 1 School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China 2 School of Software Engineering, Sun Yat-sen University, Zhuhai, China 3 Independent Researcher, Jersey City, NJ, USA 4 School of Mathematical Sciences, Zhejiang University, Hangzhou, China 5 Department of Radiology, Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA 6 Department of Pathology, Zhujiang Hospital, Southern Medical University, Guangzhou, China 7 Department of Health Management, Zhujiang Hospital, Southern Medical University, Guangzhou, China 8 College of Mathematical Medicine, Zhejiang Normal University, Jinhua, China 9 Department of Thyroid Surgery, Zhujiang Hospital, Southern Medical University, Guangzhou, China * Corresponding authors:Qinghua Liu ( 13760633321@163.com), Wei Lou (louwei@zjnu.edu.cn), Fei Chen (gzchenfei@126.com), and Guanbin Li (liguanbin@mail.sysu.edu.cn) † These authors contributed equally to this work ABSTRACT Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems address these tasks in isolation and provide limited support for clinical review. We present ThyroidXAgent, a clinician- interactive agentic AI system that coordinates specialized diagnostic tools and stores their outputs as an auditable case-level evidence record. The system was developed using OpenThyroidDB, a multicentre, multitask resource integrating approximately 0.3 million ultrasound images and 24,000 paired reports, and was evaluated on 28,458 non-overlapping test cases, including 8,721 cases from 35 centres in the private NHC-MISD-TUS cohort. Across heterogeneous datasets, ThyroidXAgent achieved a mean Dice score of 87.21% for nodule segmentation and a mean AUROC of 0.9466 for benignmalignant classification. The same workflow supported lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma classification, with AUROCs of 0.864 and 0.805, respectively. For report generation, evidence-grounded assembly outperformed multimodal language-model baselines across three cohorts. ThyClinScore, a lesion-level clinical semantic metric introduced here, showed the strongest correlation with a location-aware language-model judge. ThyroidXAgent improved physician classification accuracy, increased report diagnostic consistency from 70.3% to 86.2%, and reduced segmentation and reporting time by 35.9% and 27.4%, respectively. These findings support auditable, clinician-correctable agentic AI for thyroid ultrasound diagnosis and reporting. 1Introduction Thyroid ultrasound diagnosis depends on a sequence of lesion-level observations and clinical decisions 1,2 . A clin- ically useful examination localizes and measures thyroid nodules 1,2 , characterizes sonographic features 2–4 , assigns risk using systems such as the Thyroid Imaging Reporting and Data System (TI-RADS) 3 , determines whether fine- needle aspiration is indicated 1,3 , integrates cytology when available 5 and communicates the findings in a structured 1 arXiv:2608.12590v1 [cs.AI] 12 Aug 2026 report 2,3 . Each step introduces variability. Descriptors such as margins and echogenic foci 3,4 show substantial reader dependence, and small changes in these features can alter biopsy or follow-up recommendations 3,4 . These requirements make thyroid ultrasound a workflow-level problem for clinical AI 6,7 : useful systems must support prediction, preserve the evidence behind each step 8 and allow clinicians to revise that evidence when needed 9,10 . Most thyroid ultrasound AI systems have focused on individual components of this workflow, including seg- mentation 11–13 , classification 14–17 , and report generation 18–20 . Multicenter systems 15,21 and feature-aligned mul- timodal models 16,17,19 have connected image predictions to risk descriptors 16,17 or management recommenda- tions 15,17 . Recent studies have extended thyroid AI to fine-needle aspiration cytology 21 , lateral lymph-node metas- tasis prediction 22 and rare thyroid cancer subtype classification 23 . Many systems still present their outputs as endpoints: a mask, probability, label or report-like text. The intermediate evidence that supports these outputs is often unavailable for clinical review 8,9 , correction 9,10 or reuse across downstream tasks 8,24,25 . This endpoint- oriented design makes it difficult to determine whether an AI result is supported by appropriate lesion localization, measurement, sonographic features and report statements. This limitation reflects a broader challenge in medical AI. High-impact clinical AI studies increasingly emphasize workflow integration 9,26,27 , human-AI collaboration 6,28–33 , and evidence beyond retrospective performance 6,7,26,34 . The cognitive consequences of AI-supported clinical work 24 , clinician interaction with algorithmic recommenda- tions 10 and the transition of AI from a tool to a clinical teammate 32,35 have also become central considerations. Generalist medical AI 36–39 and multimodal foundation models 36,39–43 extend this ambition to flexible inputs and outputs across tasks. For thyroid ultrasound, however, generality must be connected to specialty-specific require- ments: lesion-level measurement 1,3 , sonographic feature attribution 3,16 , anatomical context 2 , guideline-aligned management 1,3,15 and auditable reporting 8,9,19 . We therefore treat the case-level evidence record, rather than a single prediction endpoint, as the central object of AI assistance. Agent-based workflows 35,44–48 provide one way to implement this evidence-centerd formulation. In this set- ting, the agent’s role is coordination rather than direct image interpretation 35,44,45 : it acts as a workflow controller that plans case-specific analysis 49–51 , routes inputs to specialized tools 45,46,52 , maintains intermediate state 53 and exposes structured evidence for human review 8,9,35 . Medical-agent benchmarks increasingly emphasize these capa- bilities in interactive settings 50,51,54 that require retrieval, action execution and workflow-level reasoning 52–54 . For thyroid ultrasound, the relevant evidence objects include images, lesion masks and measurements 11,12,55 , radiomic descriptors and risk estimates 16,17,56 , report clauses 18,19,57 , uncertainty signals 58 and clinician corrections 9,30,31 . The agent is therefore useful insofar as it can coordinate these objects into an auditable clinical workflow. Here we present ThyroidXAgent, a clinician-interactive agentic system that reframes thyroid ultrasound AI around an auditable case-level evidence record rather than a collection of isolated prediction endpoints. Instead of directly interpreting ultrasound images with a general-purpose multimodal model, ThyroidXAgent acts as a workflow controller that plans case-specific analyses, routes inputs to specialized tools and maintains structured intermediate evidence, including lesion masks, measurements, class probabilities, radiomic descriptors, uncertainty signals and report clauses. These evidence objects remain inspectable and editable by clinicians, and corrections can be propagated to subsequent analysis and reporting steps. We evaluate this formulation across nodule seg- mentation, benign-malignant classification, malignant-lesion stratification and case-level report generation using 40 heterogeneous multicentre datasets and clinician reader studies. We further introduce ThyClinScore, a lesion-level semantic metric that evaluates whether generated reports preserve clinically relevant evidence rather than surface- level wording alone. ThyroidXAgent improved produced more clinically consistent reports and reduced clinician workload while retaining an editable evidence trace. Together, these results establish evidence-centred orchestration as an alternative to both standalone predictive models and unconstrained end-to-end medical agents. 2/51 Segmentation Classification Reporting Opportunistic screening OpenThyroidDB A multitask thyroid ultrasound database Scanner Diversity GE LogiqE9 · GE S7 · ARIETTA 850 RESONA 70B · TOSHIBA Nemio30/MX · Hitachi Aloka Arietta V70 · EsaoteMyLab· Samsung Medison RS80 · SuperSonicImagine AixPlorerUltimate ... Report generation Segmentation Pixel-level delineation of thyroid structures Classification Thyroid nodule diagnosis and categorization AHU 125,896 ThyroidXL 11,631 ThyUS2Path 8,508 TN5K 5,000 TN3K 5,347 DDTI 349 Public datasets Images-report pairs for standardized reporting 353 reports, 4,471 images ZJH-TS 23,955 reports, 248,194 images Opportunistic diagnosis LNM prediction and FTC/PTC subtype classification Public dataset 200 images FTC / PTC Ours-ZJH 4,756images LNM LymphUs, 338 images Cine-clip 17,417 TGVideo 15,186 ThyroidXL 11,631 TN3K 5,347 RJH-7K 7,288 PKTN 1,003 DDTI 637 Agentic nodule mask generation Clinician-auditable refinement Traceable tool calls Auditable malignancy assessment Radiomic feature analysis SHAP-based explanation Ultrasound-based report drafting Image and video interpretation a b c ThyroidXAgent Review and correction Expert refinement Refined outputs Refined mask Verified diagnosis Corrected report Evolving Database High-quality, expert- refined data repository Model updating Feedback-driven tool evolving Clinician-in-the-loop Efficient and reliable diagnosis Initial outputs mask diagnosis report Curated feedback Verified corrections and annotations Expert reviewAfter refinement Stored as feedbackUsed for updating Ours newly contributed 145 cases videos TNVideo Public datasets Ours newly contributed ZJH-8K 7,958 images TNVideo 148 cases Ours newly contributed ZJH-8K 7,958 cases TNVideo 148 cases d ThyroidX Agent Training and internal validation 10 centers External validation 35 centers 2,474 reports, 15,921 images KMVE Public datasets SMU-HMC ThyroidXAgent Traceable tool-call records Fast model development Lymph-node metastasis prediction Explainable risk prediction Figure 1.Multicenter thyroid ultrasound data and the ThyroidXAgent evidence workflow.a, OpenThyroidDB integrates curated public thyroid ultrasound resources and institutionally governed clinical cohorts into a full-spectrum resource for segmentation, benign-malignant classification, report generation and malignant-lesion stratification across diverse scanners and acquisition settings.b, The clinician-feedback workflow stores AI-generated masks, predictions, report clauses and clinician corrections as case-level evidence, allowing intermediate outputs to be reviewed, revised and reused in later steps.c, ThyroidXAgent acts as a workflow controller for four evidence-producing workflows: expert-refined nodule segmentation, clinician-verified malignancy classification, evidence-grounded structured reporting, and advanced diagnosis for lymph-node metastasis assessment and PTC/FTC subtype analysis.d, Multicenter validation design, with model development/internal validation followed by independent external validation. 3/51 bc f h High Low Feature Value SHAPE: Sphericity SHAPE: Elongation RLM: LRHGLE INT: Skewness INT: Median GLCM: Corr TEX(RLM) TEX(GLCM) RLM: SRLGLE INT: 10th pctl -3 -2-10 1 23 SHAP contribution toward malignant (raw score) de Human Input ClassifythisThyroidNodule isBenignorMalignant Nodule Segmentation Prediction:benign Confidence:0.86 SHAPfeatureanalysis DoctorReview Reliable Mask? Assisted Diagnosis Mask,SHAPfeatures, andClassprediction Doctorbox annotation ThyroidXAgent Interactivemedical Segmentationmodels No Yes Reliablesegmentation updatesSHAPfeatures andclassification confidence a 30.0 22.5 15.0 7.5 0.0 Manual N=30 AI-Assisted N=30 14.2->9.2s;1.56xfaster 35.9%shortersegmentationtime 35.9%shorteroverall 5.1s/casesaved 20 15 10 5 0 -5 -10 Segmentation time (s) Manual – AI time (s) Pairedcasessortedbytimesaving Dice Score 1.0 0.9 0.8 0.7 0.6 0.5 ManualAI-Assisted ijk DoctorReview ThyroidXAgent MedSAM2 MedSegX TransUnet UltraFedFM ThyroidXAgent MedSigLIP BiomedCLIP UltraFedFM 00.20.40.60.81.0 0.823 9.4m 0.232 101.2m 0.288 87.0m 0.730 10.6m 0.589 27.0m 0.819 0.823 0.520 0.519 0.434 0.450 0.436 0.465 Normalizedscore(0~1;HD95shows/110) Classification Segmentation DiceHD95↓AUROCAUPRC g Figure 2.Agent-routed evidence for thyroid nodule segmentation, benign-malignant classification and clinician review.a, Interactive review workflow: clinicians assess the predicted nodule mask, accept the classification if the mask is reliable, or refine it through box annotation and interactive segmentation; SHAP-based feature attributions are then recomputed from the corrected mask.b-e, Segmentation and classification performance across heterogeneous thyroid ultrasound benchmarks, evaluated using Dice, HD95, AUROC and AUPRC.f, Cohort-level SHAP beeswarm analysis for benign-malignant classification.g, ROC curves on the 500-image physician comparison set, showing ThyroidXAgent, representative AI baselines and clinician operating points before and after ThyroidXAgent support; both clinicians improved with evidence support.h, Pooled performance on the NHC-MISD-TUS private external test set across nine models and four metrics: nodule Dice, HD95, binary AUROC and AUPRC. Bars show point estimates with 95% CIs, and the dashed line indicates the classification chance level of 0.5.i, Segmentation time for manual and AI-assisted workflows.j, Ranked within-case time savings.k, Paired Dice distributions showing preserved segmentation quality with improved efficiency. 4/51 2Results 2.1ThyroidXAgent organizes thyroid ultrasound diagnosis as an auditable case-level evidence workflow We developed ThyroidXAgent to organize thyroid ultrasound diagnosis as a case-level evidence workflow (Fig.1). For each examination, the agent routes the case through tool-callable steps, including nodule segmentation, measurement, benign-malignant classification, radiomics extraction, malignant-lesion stratification and report gen- eration. The workflow stores tool outputs in a shared evidence record that can be inspected, corrected and reused across downstream tasks. OpenThyroidDB provides the multicenter data resource underlying this workflow-level formulation. The seg- mentation and classification analyses included seven cohorts from six clinical centres across China, Vietnam and Colombia, acquired on at least eight ultrasound platforms (Supplementary TableS1). Four cohorts served as inter- nal training data. TN3K 11 , collected at Zhujiang Hospital, Southern Medical University, Guangzhou, comprised 4,633 training, 100 validation and 614 test images acquired on GE Logiq E9, ARIETTA 850 and RESONA 70B scanners. TN5K 55 , from the Cancer Hospital, Chinese Academy of Medical Sciences, Beijing, comprised 3,500 training, 500 validation and 1,000 test images acquired on GE Logiq E9 and GE S7 scanners with 5-12 or 8-15 MHz probes. ThyroidXL 59 , from the Vietnam National Hospital of Endocrinology, Hanoi, comprised 9,441 training, 100 validation and 2,090 test images acquired on a Hitachi Aloka Arietta V70 scanner. PKTN 13 , from Peking Univer- sity First Hospital, Beijing, comprised 703 training, 150 validation and 150 test images; the acquisition device was not disclosed. Three cohorts provided independent external test sets. DDTI 60 , from IDIME, Bogotá, Colombia, contributed 637 images (637 for segmentation, 349 of which also carry benign-malignant classification labels) ac- quired on TOSHIBA Nemio 30 and Nemio MX scanners with a 12 MHz probe. RJH-7K 61 , from Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, contributed 7,288 segmentation-only images acquired on multiple scanners whose models were not specified. ZJH-8K, collected at Zhujiang Hospital, Southern Medical University, Guangzhou, contributed 7,958 images (426 benign cases with 3,202 images and 723 malignant cases with 4,756 images) used for both segmentation and classification; the acquisition device was not disclosed. Overall, the segmentation and classification benchmark comprised 38,864 images (18,277 training, 850 validation, 19,737 test) from seven cohorts spanning six centres. The report-generation analyses included four cohorts from three centres. SMU-HMC comprised 23,955 reports from 21,954 patients, with 248,194 associated ultrasound images. A patient-level held-out set of 400 patients, comprising 5,984 images, was reserved for internal testing. After excluding all examinations from these internally held-out patients, the remaining SMU-HMC cohort was used to develop the relevant component models, and 7,147 quality-controlled reports were selected to construct the retrieval template library. The KMVE analysis cohort 18 comprised 2,457 cases and 4,914 image assignments, partitioned according to the original dataset split into 1,719 training, 246 validation and 492 test cases; its training partition was additionally used for template-library construction. External evaluation included two complementary cohorts. ZJH-TS provided an independent external- centre test set of 150 reports selected from the original collection of 353 reports, after excluding reports dominated by postoperative findings. TNVideo provided an independently assembled case-level cohort for the external reader study, comprising 148 ultrasound examinations, of which 145 had annotation-derived diagnostic labels. Overall, the report-generation analyses comprised four cohorts from three centres, including an internal SMU-HMC test set, an independent external-centre test set and an external reader-study cohort. For malignant-lesion stratification, the 4,756 malignant images from ZJH-8K (also an external test set for segmentation and classification) served as the primary training set, of which 20 cases (183 images) were held out for validation. For lateral lymph-node metastasis (LNM) prediction, 180 images from LymphUs Center 1 were 5/51 additionally included as training data, and the 158 images from LymphUs Center 2 served as the independent external test set. For follicular (FTC) versus papillary (PTC) thyroid carcinoma subtype classification, the 200 public images released by Daiet al. 62 served as the external test set. Within the ZJH-8K malignant cohort, 533 cases (3,394 images) were classified as CN0 and 190 cases (1,362 images) as CN1 for lymph-node status, and 656 cases (4,312 images) were classic PTC and 67 cases (444 images) were follicular variant for subtype. Collectively, these resources cover heterogeneous acquisition settings, file formats, annotation types and clinical tasks, including nodule segmentation, benign-malignant classification, report generation and advanced malignant- lesion analysis (Fig.1a and Supplementary TableS2). Using this resource, ThyroidXAgent coordinates four evidence-producing workflow branches: expert-refined nodule segmentation, clinician-verified malignancy classifi- cation, expert-edited report generation and advanced diagnosis for lateral lymph-node metastasis and PTC/FTC subtype analysis (Fig.1b,c). Each branch contributes structured evidence to the same case-level record, including masks, measurements, class probabilities, radiomic descriptors, feature attributions, uncertainty signals and report clauses. This design allows intermediate outputs to be reused across tasks. A corrected mask can support radiomics extraction, a malignancy estimate can inform the report impression, and structured report evidence can be reviewed and edited rather than accepted as opaque text. The resulting workflow makes auditability a property of the diagnostic process rather than a post hoc explana- tion attached to a final answer. Clinicians can inspect generated masks, correct segmentation errors, review SHAP- or Grad-CAM-based explanations, edit report statements and return corrected outputs to the evidence store. The subsequent results evaluate this formulation across automatic image analysis, clinician correction, malignant-lesion stratification, clinical semantic report scoring and evidence-grounded report assembly. 2.2Agent-routed evidence improves cross-dataset segmentation and classification We first evaluated the two image-analysis tasks that anchor the downstream workflow: nodule segmentation and benign-malignant classification. For each case, ThyroidXAgent collected candidate masks and class probabilities from DINOv3-based experts, selected or fused outputs using case-level quality signals, extracted radiomic features from the selected lesion mask and stored confidence, disagreement and tabular predictions as structured evidence (Supplementary Fig. S1). This design tests whether tool routing and evidence consolidation improve robustness across heterogeneous ultrasound datasets, where dataset bias remains a major source of performance degradation 63 . On seven segmentation test sets, including the independent DDTI, RJH-7K and ZJH-8K cohorts, ThyroidXA- gent achieved a mean Dice coefficient of 87.21% and a mean 95th-percentile Hausdorff distance (HD95) of 6.90 m (Fig. 2b,c and Supplementary TableS3). It obtained the highest Dice score on six of seven test sets and the lowest HD95 on all test sets, indicating improved boundary robustness across heterogeneous acquisition conditions. The strongest baseline, MedSAM2 64 , achieved a mean Dice of 85.66%, with the largest gap on the external ZJH-8K cohort (86.29% versus 94.30% for ThyroidXAgent). UltraFedFM 65 , a medical imaging foundation model, reached 79.66%. For benign-malignant classification, ThyroidXAgent achieved a mean AUROC of 0.9466 and a mean AUPRC of 0.8361 across five test sets, including the independent DDTI and ZJH-8K cohorts (Fig. 2d,e and Supplementary Table S4). Specialized image models showed weaker cross-dataset consistency; for example, RepViT 66 reached AUROC of 0.777 on ThyroidXL but 0.556 on TN3K. General-purpose vision-language models 67 also underperformed; GPT-5 68 reached AUROC of 0.611-0.774 across test sets, and Gemini-2.5-Pro 69 reached 0.616-0.687 (Supplementary Table S4). These results support a division of labour in which domain-specific image and radiomics tools generate the evidence, while the agent routes and consolidates tool outputs. We further validated on NHC-MISD-TUS, an independent private multicentre test set (Fig.2h and Supplementary TablesS9-S12). ThyroidXAgent led on both 6/51 nodule segmentation (Dice 82.31%, HD95 9.41 m) and binary classification (AUROC 0.819, AUPRC 0.823). By contrast, MedSAM2 and MedSegX failed on segmentation (Dice <0.3), and zero-shot vision-language baselines (BiomedCLIP 70 , MedSigLIP 71 ) performed at or below chance on classification. GPT-5 and Gemini-2.5-Pro could not be evaluated on this private intranet dataset. Beyond prediction accuracy, we examined whether the structured evidence provided interpretable signals for clinician review. Cohort-level SHAP profiles 72 showed that morphology-related radiomic descriptors, especially Sphericity and Elongation, dominated benign-malignant classification, whereas texture and intensity features con- tributed complementary information (Fig.2f). Representative cases confirmed that accurate segmentation produced SHAP attributions and Grad-CAM maps aligned with visible nodule characteristics, whereas poor segmentation degraded these explanations (Supplementary Fig.S3). To assess whether this evidence supports clinician decision-making, we conducted a blinded 500-image physician comparison (Fig.2g). ThyroidXAgent achieved AUROC of 0.9256 and AUPRC of 0.9250. In the same compari- son, UltraFedFM 65 reached AUROC of 0.880, GPT-5 68 0.660 and Gemini-2.5-Pro 69 0.619. When clinicians were provided with ThyroidXAgent’s structured evidence, including SHAP-based feature attributions and nodule seg- mentation boundaries, both clinicians improved across all metrics. For clinician 1, accuracy rose from 79.2% to 84.6% (+5.4%), precision from 83.5% to 87.5%, recall from 72.8% to 80.8%, specificity from 85.6% to 88.4%, and F1 from 0.778 to 0.840, while the false-positive rate fell from 14.4% to 11.6%. Clinician 2 improved more markedly: accuracy from 74.0% to 84.8% (+10.8%), precision from 75.2% to 84.3%, recall from 71.6% to 85.6%, specificity from 76.4% to 84.0%, and F1 from 0.734 to 0.849, with the false-positive rate decreasing from 23.6% to 16.0%; clinician 2 thereby narrowly surpassed clinician 1 as the higher-performing reader. 2.3Clinician correction reduces segmentation time while preserving quality The same evidence representation supported clinician-interactive segmentation review. Clinicians inspected predicted masks, corrected segmentation errors when needed and returned the refined masks to the case-level evi- dence store. SHAP-based feature attributions were then recomputed using the corrected mask, allowing downstream classification evidence to reflect clinician refinement (Fig. 2a). AI assistance reduced mean segmentation time from 14.21 s to 9.11 s per image, a 1.6-fold speedup, while preserving segmentation quality (Fig.2i-k). AI-assisted Dice (0.903) matched or exceeded manual Dice (0.879) in approximately two-thirds of paired cases. These results indicate that the intermediate evidence layer can reduce repetitive annotation work while keeping the segmentation boundary available for clinician correction. 2.4Shared agent tools generate task-specific evidence for malignant-lesion stratification We next tested whether the shared ThyroidXAgent tools could be redirected to clinically distinct malignant- lesion stratification tasks after the primary thyroid nodule assessment. Lateral lymph-node metastasis (LNM) prediction informs surgical planning, whereas follicular (FTC) versus papillary (PTC) thyroid carcinoma subtype discrimination informs treatment strategy and follow-up. In a conventional development pipeline, each task would require a separate workflow for preprocessing, feature extraction, prediction, interpretation and reporting. In ThyroidXAgent, the segmentation, radiomics extraction, tabular classification and routing logic were reused, with task-specific classifier fine-tuning and task instructions changed. This reuse enabled rapid adaptation to new clinical questions while preserving the same evidence structure. ThyroidXAgent achieved AUROC of 0.864 for LNM prediction on the 158-image LymphUs Center 2 test set and 0.805 for FTC/PTC subtype classification on the 200-image Daiet al.test set 62 (Fig.3b and Supplementary TableS5), outperforming the specialist baselines LLNM-Net 22 (0.767) for LNM and Tiger-Model 23 (0.714) for 7/51 RLM: GLNU TEX(GLCM): Imc2 SHAPE: Sphericity TEX(NGTDM): Strength RLM: RunEnt RLM: SRLGLE TEX(GLCM): Cluster Shade NGTDM: Coarse SHAPE: Elongation INT: Skewness -0.6-0.4-0.20.00.20.40.6 SHAP contribution toward FTC (raw score) High Low Feature Value a b c d e f INT: Maximum NGTDM: Contrast SHAPE: Perim/Area TEX(GLCM): Cluster Tendency INT: Energy -1.0 -0.5 0.0 0.5 1.0 SHAP contribution toward LNM-positive(raw score) Toward LNM-positive Toward LNM-negative SHAPE: Sphericity RLM: RunEnt TEX(RLM): Gray Level V TEX(GLCM): Imc1 TEX(GLCM): Cluster Shade -1.0 -0.50.00.51.0 SHAP contribution toward FTC(raw score) Toward FTC Toward PTC INT: Maximum INT: Energy SHAPE: Elongation TEX(GLCM): Cluster Shade TEX(GLCM): Cluster Tendency SHAPE: Perim/Area SHAPE: Perimeter SHAPE: Minor axis NGTDM: Contrast SHAPE: Pixel Surface -1.5-1.0-0.50.00.51.01.5 SHAP contribution toward LNM-positive (raw score) High Low Feature Value ThyroidXAgent Load data & Objects •Load AutoGluonpredictor •Read training CSV •Preprocess data Identify Samples & Main Models Build Background & Explain Sets SHAP Analysis Compute & Visualize •Compute SHAP values, feature importance •Visualize Beeswarm,Waterfall •Write Analysis summary SHAP Values Feature Importance Waterfall Plots Analysis Summary Outputs Beeswarm Plots Human Input Use SHAP to identify the most important radiomic features for benign–malignant classification. Baseline ThyroidXAgent 0.786 0.805 0.881 0.864 +12.9% +10.5% +19.6% +12.7% FTC/PTC Subtype (Tiger-Model) AUROC FTC/PTC Subtype (Tiger-Model) AUPRC Lymph Node Metastatis (LLNM-Net) AUROC Lymph Node Metastatis (LLNM-Net) AUPRC 0.700.75 0.80 0.85 0.90 Metric Score LNM-negative FTC Figure 3.Shared ThyroidXAgent tools support malignant-lesion stratification with task-specific radiomic attributions.a, Workflow for SHAP-based interpretation of malignant-lesion tasks.b, Performance comparison between ThyroidXAgent and the corresponding specialist baselines for FTC/PTC subtype classification and lymph node metastasis prediction, reported as AUROC and AUPRC; percentages denote the relative improvement of ThyroidXAgent over each baseline.c,d, Global and representative local SHAP analyses for lymph node metastasis prediction.e,f, Global and representative local SHAP analyses for FTC/PTC subtype classification, showing stronger contributions from texture heterogeneity and shape descriptors. 8/51 FTC/PTC. The attribution profiles changed with the clinical task (Fig.3a,c-f). LNM prediction relied more on lymph-node position and size features, including distance to the thyroid capsule and lesion area. FTC/PTC subtype classification relied more on texture heterogeneity and shape descriptors. These task-dependent attribution patterns indicate that ThyroidXAgent generated radiomic evidence according to the requested clinical question, rather than simply reusing the benign-malignant decision rule. 2.5ThyClinScore captures lesion-level report errors missed by overlap metrics We then addressed the evaluation of thyroid ultrasound reports. Conventional natural-language generation metrics, such as BLEU 73 , ROUGE 74 , and METEOR 75 , primarily reward surface overlap and can miss clinically important disagreements in lesion location, size, vascularity, morphology or impression. We therefore developed ThyClinScore, a clinical semantic metric that structures ground-truth and generated reports, matches lesion entries and scores clinically relevant attributes (Fig.4a). ThyClinScore captured report-quality dimensions that were complementary to wording overlap. Overlap-based metrics were strongly correlated with one another, whereas the clinical semantic metrics captured distinct lesion- level and feature-level information (Fig.4b). Using a Pearson correlation-based evaluation approach similar to that of Liet al. 20 , ThyClinScore showed the strongest correlation with a location-aware LLM judge among the evaluated metrics (Pearson’sr=0.696,p<0.001; Fig.4c). Qualitative examples further show that reports with similar wording overlap can differ in clinically important attributes, which is reflected by the ThyClinScore components (Fig. 4d). 2.6Evidence-grounded report assembly improves reporting consistency and efficiency For report generation, ThyroidXAgent converted the case-level evidence record into structured report text. The agent first used multi-view and multimodal thyroid ultrasound inputs to construct image priors, invoked diagnostic tools through planning and execution, and then assembled report clauses from structured facts using BM25 template retrieval, slot filling and clause combination (Fig. 5a-g). This design links report statements to intermediate evidence, including gland measurements, nodule location, lesion size, sonographic descriptors, vascularity, lymph-node findings and diagnostic impressions, rather than generating unconstrained free text. The report-generation workflow was packaged as a reusable skill and exposed to external agents through a Workflow MCP server. This interface provided high-level case operations, including input preparation, case initiation, plan review, approval, report retrieval and evidence retrieval, while low-level model calls remained internal to the workflow. Auxiliary tool performance is summarized in Supplementary Table S13. Clinicians can review the structured evidence, edit generated statements and return corrected report content to the case-level evidence store. On conventional natural-language generation metrics, ThyroidXAgent achieved the strongest overall perfor- mance across SMU-HMC, KMVE 18 and ZJH-TS (Fig.5f and Supplementary TableS14). On SMU-HMC, BLEU-1, BLEU-4 and ROUGE L reached 0.5961, 0.3405 and 0.5450, respectively. On KMVE, ThyroidXAgent ranked first on all reported overlap metrics, with BLEU-1 of 0.6209, BLEU-4 of 0.4465, METEOR of 0.3596 and ROUGE L of 0.5826. On ZJH-TS, ThyroidXAgent achieved the highest BLEU-1, BLEU-4 and ROUGE L values, reaching 0.5134, 0.2942 and 0.5529, respectively. We further compared the adaptive tool-routing workflow with a fixed rule-based tool-calling pipeline (Fig. 5g and Supplementary Table S16). Across all three report-generation test sets, ThyroidXAgent produced larger multi- metric profiles than the static pipeline. ThyClinScore increased from 0.4293 to 0.5238 on SMU-HMC, from 0.3346 to 0.4465 on KMVE and from 0.3648 to 0.4775 on ZJH-TS. 9/51 LLM Structuring LLM Structuring St ru ct ure d Report St ru ct ure d Report ThyClinScore Computation ThyClinScore Completeness Morphology F1 Lesion Matching Size consistency Vascularity score Lesion Consistency Lesions Lesions Pairwise Si milarity Computa ti on Bipartite matching St ru ct ure d Report Extract lesions Extract lesions FN TP FP Ground Truth Report Ground Truth Report Ground Truth Report Prediction Report Prediction Report Prediction Report T h e r i g h t t h y r o i d l o b e m e a s u r e s 48×12×13 m, the left lobe 44×13×10 m, and the isthmus 2 m. The thyroid is normal in shape and size... The left thyroid lobe measures about 38× 25×11 m, the right lobe about 52×12×14 m, and the isthmus about 3.9 m. The thyroid is normal in shape and size... Direct Gland-level Scoring Similarity Matrix Pearson’s r Our MetricsTraditional Metrics Traditional Metrics Our Metrics Inter-metric correlation Qualitative comparison bcd a St ru ct ure d Report Report 1:The right lobe of the thyroid measures 56 × 22 × 21 m, the left lobe measures 52 × 21 × 17 m, and the ist hmus thickness is 1.8 m. The size and morphology of the thyr oid gland are normal, and the glandular capsule is s m o o t h . T h e p a r e n c h y m a l e c h o t e x t u r e i s heterogeneous. CDFI reveals sparse, fine branching color blood flow signals within the parenchyma. A solid isoechoic nodule, measu rin g approxi mately 20 × 15 m with clear margins, is detected in the right thyroid lobe... BlEU-4 0.714 ROUGE-L 0.837 GPT-5 Score 0.350 ThyClin Score 0.254 Report 2:The right lobe of the thyroid measures 50 × 20 × 24 m, the left lobe measures 42 × 12 × 15 m, and the ist hmus thickness is 2.8 m. The size and morphology of the thyr oid gland are normal, and the glandular capsule is s m o o t h . T h e p a r e n c h y m a l e c h o t e x t u r e i s homogeneous. CDFI reveals a 'thyroid inferno' pattern of color blood flow signals within the parenchyma. A cystic hypoechoic nodule, measuring approximately 20 × 15 m with clear margins, is detected in the left thyroid lobe... Correlation with LLM a Judge score p-values: * p < 0.05, ** p < 0.01, *** p < 0.001 0.601*** 0.696*** 0.641*** 0.632*** 0.587*** 0.580*** 0.539*** 0.50 0.55 0.60 0.65 0.70 0.75 Pearson Correlation BLEU-1 BLEU-4 ROUGE-L METEOR Consistency Lesion F1 ThyClinScore BLEU-1 BLEU-4 ROUGE-L METEOR Consistency Lesion F1 ThyClinScore Completeness BLEU-1 BLEU-4 ROUGE-L METEOR Consistency Lesion F1 ThyClinScore Completeness 0.0 0.2 0.4 0.6 0.8 1.0 Figure 4.ThyClinScore for clinical semantic evaluation of thyroid ultrasound reports.a, ThyClinScore first structures the ground-truth and predicted reports into lesion-level entries, matches corresponding lesions, and scores clinically relevant attributes, including size, vascularity, morphology, lesion-level F1 and completeness.b, Pearson correlation matrix comparing conventional natural-language generation metrics with clinical semantic metrics. Overlap-based metrics were highly correlated with one another, whereas the clinical semantic metrics captured complementary report-quality dimensions.c, Pearson correlations between each metric and a location-aware LLM judge. ThyClinScore showed the strongest correlation among the evaluated metrics; asterisks denote statistical significance (*p<0.05, **p<0.01, ***p<0.001).d, Qualitative comparison showing that reports with similar wording overlap can differ in clinically important lesion attributes, which is reflected by ThyClinScore. 10/51 ThyroidXAgent Qwen3.5 Plus GPT-5 GPT-4o MedGemma 0.1048 0.2810 0.4573 0.6335 0.5480 0.5314 0.4790 0.3679 0.1592 ThyroidXAgent Qwen3.5 Plus GPT-5 MedGemma GPT-4o 0.2924 0.3621 0.4317 0.5014 0.4676 0.4520 0.4472 0.3807 0.3139 f SMU-HMC ROUGE-LThyClinScore KMVE ThyClinScoreROUGE-L g SMU-HMCKMVE ThyroidXAgent Static Pipline ThyClinScore ROUGE-L ZJH-TS ZJH-TS BLEU-1 BLEU-2 BLEU-3 BLEU-4 METEOR Lesion F1 ThyClinScore 0.59 0.48 0.40 0.34 0.36 0.55 0.52 0.46 0.38 0.33 0.28 0.32 0.51 0.43 0.64 0.56 0.50 0.45 0.37 0.36 0.44 0.33 0.21 0.13 0.06 0.27 0.26 0.33 BLEU-1 BLEU-2 BLEU-3 BLEU-4 METEOR Lesion F1 ThyClinScore 0.51 0.41 0.34 0.29 0.33 0.49 0.47 0.42 0.35 0.30 0.25 0.29 0.39 0.36 BLEU-1 BLEU-2 BLEU-3 BLEU-4 METEOR Lesion F1 ThyClinScore ThyroidXAgent Qwen3.5 Plus GPT-5 GPT-4o MedGemma 0.0186 0.2268 0.4350 0.6432 0.5422 0.5147 0.3732 0.3577 0.0829 ThyroidXAgent Qwen3.5 Plus GPT-5 GPT-4o MedGemma 0.3488 0.4164 0.4841 0.5517 0.5189 0.4883 0.4882 0.4105 0.3697 ThyroidXAgent Qwen3.5 Plus GPT-5 GPT-4o MedGemma 0.1299 0.3121 0.4942 0.6764 0.5880 0.5066 0.4128 0.3850 0.1862 ThyroidXAgent MedGemma GPT-5 Qwen3.5 Plus GPT-4o 0.3195 0.3677 0.4159 0.4641 0.4407 0.3932 0.3853 0.3674 0.3344 1RGXOH &ODVVLILFDWLRQ 5HJ LRQ &ODVVLILFDWLRQ &'),'HWH FWLRQ 52,&UR SSLQJ ,PD JH&RQWH[W 3D UVL QJ $J HQWUH DG\ ,PD JH3UL RUV ,PDJ H3DWFKHV 2ULJLQDO,PDJ HV &URS DE 5H$FW /RRS 6NLOO VLQ SXWSUHSDUDWLRQ6NLOO VGLDJQRVWLFSODQQLQJ F 6NLOO VH[HFXWLRQYLD0&3 0&3&DOOV 3UH SUR FH VVLQJ ,Q IR UPD WLRQ G H 6NLOO VUHSRUWJHQHUDWLRQ &ODXVH FRPELQDWLRQ 2XWSXW %X LOGVL PS OHTXHU\ 7HPS ODWHUH WUL HYD O 6O RWILOOLQJ 7HPSODWH %D QN 7HPSODWH & ULJKWOREHRIWKHWK LOOGHILQHGPDUJLQ« &RPSOHWHG FODXVH 6WDUW 6WUXFWXUH(YLGHQFH QRGXO HBL G UH JL RQ UL JKW F RP SRV LWLRQ & WLF H FKRJH QL FLW\ $ QH FKRL F P DUJL Q L OOGH IL QH G 8OWUDVRXQG)LQGLQJV 7KHOHIWOREHRIWKHWK [LPDWHO\ hhPPDQGWKHULJKWOREHPHDVXUHVDSSUR[LPDWHO\ hhPPZLWKWKHLVWKPXVWKLFNQHVVDERXWPP7KH WK 7KH SDUHQFK XQHYHQGLVWULEXWLRQ1RGHILQLWHK DUHGHWHFWHG&'),VKRZVDEXQGDQW LQWUDWK SUHVHQWLQJD³WK ́SDWWHUQ,QWKHULJKWOREHD K [LPDWHO ZLWKFOHDUPDUJLQVLVREVHUYHG&'),VKRZVQRLQWHUQDOEORRG IORZVLJQDO1RDEQRUPDOO DURXQGWKHELODWHUDOFDURWLGDUWHULHV 8OWUDVRXQG,PSUHVVLRQ $VROLGOHVLRQLQWKHULJKWWK 7, 5$'6FDWHJRU\ VXJJHVWLYHRIWK 7K 8OWU DVRXQG,PDJHV 9LGHRRU ,PDJH 6HTXHQFH 0XOWLSOH YLHZVDQG 5HJLRQV *UD\6FDOH DQG&'), PRGDOLWLHV 7RRO UHVSRQVHV 0& 36H UYH U S XEOLF ZRUNI ORZ$3, 3UHSDUHLQSXWV*HWSODQ 6WDUWFDVH$SSURYHSODQ *HWFDVH*HWUHSRUW *HWVWDWXV*HWHYLGHQFH 'LDJQRVW LF VW DJHGWDVN JUD SK //0 3O DQQHU 6., //PG 7RROFR QWUD FW 3UH SUR FH VVL QJ LQIRUPD WLRQ &RQWH[W Figure 5.Evidence-grounded report assembly in ThyroidXAgent.a, Case-level thyroid ultrasound input, including video or image sequences, multiple views and anatomical regions, and greyscale and CDFI modalities.b, Input-preparation skills crop regions of interest, parse image context, detect CDFI information, classify nodules and anatomical regions, and convert the resulting preprocessing outputs into agent-ready image priors.c, Diagnostic planning combines the skill instructions, preprocessing information and tool contract with an LLM planner to generate a staged diagnostic task graph.d, A ReAct-style execution loop observes intermediate evidence, reasons over the next step and calls tools through an MCP server that exposes the public workflow API, including case initialization, plan approval, status checking, evidence retrieval and report generation.e, Structured evidence is transformed into report text through query construction, template retrieval, slot filling and clause combination, so that report clauses remain linked to measurements, lesion descriptors, risk estimates and other case-level evidence.f, Report-generation performance comparison on SMU-HMC, KMVE and ZJH-TS.g, Radar plots comparing ThyroidXAgent with the static pipeline across lexical and clinical semantic metrics. 11/51 Clinical semantic evaluation showed that the reports retained clinically relevant information beyond wording similarity (Fig.4and Supplementary TableS15). ThyroidXAgent achieved the highest ThyClinScore on SMU- HMC (0.5238), KMVE (0.4465) and ZJH-TS (0.4775). It also achieved the highest lesion-level F1 and report completeness on SMU-HMC and ZJH-TS, whereas KMVE showed a different submetric profile in which several individual components were led by other baselines. Together with the metric-correlation analysis (Fig.4b,c), these results indicate that clinical semantic scoring captures report-quality information not represented by conventional overlap metrics. Finally, we assessed human-AI cooperation in a cross-over reader study. Two physicians wrote reports for 145 thyroid ultrasound videos under manual and AI-assisted workflows, with each case evaluated in both workflows by different physicians to reduce memory bias (Fig.6a). In the 145 cases with annotation-derived diagnostic-direction labels, AI-assisted reports showed higher consistency than manual reports overall and within benign and malignant subsets (Fig.6c). AI assistance reduced mean reporting time from 2.5 to 1.8 min per case, a 27.4% reduction, with similar time savings for both readers (Fig.6d,e). A representative malignant case shows that the structured evidence supported statements on gland morphology, nodule location, measurements, sonographic features and diagnostic impression, while leaving partially correct or incorrect statements available for clinician review (Fig.6b). 3Discussion This study reframes thyroid ultrasound AI as an auditable case-level evidence workflow rather than a collection of isolated prediction tasks. The principal contribution of ThyroidXAgent is not a new standalone segmentation, classification or report-generation model, but a shared evidence architecture in which lesion masks, measurements, probabilities, radiomic descriptors, uncertainty signals and report clauses remain visible, editable and reusable across the diagnostic process. Across heterogeneous datasets, this formulation improved nodule segmentation and benign- malignant classification, supported adaptation to additional malignant-lesion tasks, enhanced evidence-grounded reporting and reduced clinician workload. Auditability is therefore embedded in the workflow itself rather than appended retrospectively to a final prediction. ThyroidXAgent also defines a constrained and clinically practical role for agentic AI. Rather than asking a general-purpose vision-language model to interpret ultrasound images directly, the agent plans case-specific anal- yses, routes inputs to specialized tools, maintains case state and exposes intermediate results for review. Image interpretation and quantitative measurement remain assigned to task-specific models, whereas the agent provides coordination and evidence management across them. This division of labour extends emerging medical-agent frame- works centred on planning, tool use and coordinated clinical workflows 35,44,45 . Its value lies less in unrestricted autonomy than in making heterogeneous model outputs coherent, traceable and correctable. The shared evidence record further distinguishes ThyroidXAgent from a conventional fixed pipeline. A cor- rected lesion mask can be propagated to measurement, radiomics, classification and reporting, while anatomical context, malignancy estimates and lesion descriptors can be reused when constructing the diagnostic impression. The same evidence-producing tools could consequently be redirected towards lymph-node metastasis prediction and thyroid carcinoma subtype classification without rebuilding the complete workflow. Report generation provides a particularly stringent demonstration of this design: rather than producing unconstrained free text, ThyroidXAgent assembles report statements from structured clinical evidence. ThyClinScore complements this approach by evalu- ating lesion matching, clinically relevant attributes and report completeness, thereby capturing errors in location, size or morphology that may be missed by conventional lexical-overlap metrics. The reader studies indicate that this evidence-centred formulation can support human-AI cooperation without 12/51 Ground Truth Findings: The right thyroid lobe measure s 47×16×14 m, the left lobe 50×16×8.8 m, and the isthmus 3 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is slightly heterogeneous. No definite ech ogenic mass or cystic dark area is se en in the left lobe or ist hmus. CDFI sh ows a few fine branching color blood-flow signals. A hypoechoic lesion is seen ne ar the isthmus in the right lobe, about 6.8×8 m. Its margin is unclear. Scattere d punct ate ech ogenic foci are prese nt. CDFI sh ows internal blood flow. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: Hypoechoic lesion with calcification near the isthmus in the right thyr oid lobe, co nsi st ent with TI-RADS 5; puncture biopsy is recommended. No focal lesion is seen in the left thyroid lobe or isthmus. CorrectPart ially co rrectInco rrect Qwen 3.5 GPT-5 Ours Findings: The left thyroid lobe measures about 48×15×13 m, the right lobe about 47×16×12 m, and the isthmus about 2.0 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is heterogeneous. CDFI shows a few fine branching color blood-flow signals in the gland. A solid hypoechoic nodule is seen in the left lobe, about 13×9 m, Figure 1. It has clear margin, regu lar shape, aspect ratio <1, and homogeneous internal echo. No obvious internal blood flow is seen on CDFI. An anechoic area is also seen in the upper pole of the left lobe, about 5×4 m, Figure 3. It has clear margin and posterior enhancement, with no internal blood flow on CDFI. A solid isoechoic nodule is seen in the right lobe, about 10×7 m, Figure 3. It has clear margin, regular shape, and ho mogeneous internal echo. CDFI shows punctate and strip-like blood flow around it. No obvious abnormal enlarged cervical lymph-node echo is seen. Impression: Solid hypoechoic nodule in the left thyroid lobe, consistent with TI-RADS 3. Cystic nodule in the left thyroid lobe, consistent with TI-RADS 2. Solid isoechoic nodule in the right thyroid lobe, consistent with TI-RADS 3. Findings: The left thyroid lobe measures about 35×15×13 m, the right lobe about 35×11×10 m, and the isthmus about 3.5 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is homogeneous. No definite echogenic mass or cystic dark area is seen in the left lobe or isthmus. CDFI shows a few fine branching color blood-flow signals. A hypoechoic lesion is seen in th e right lobe, about 7.9×6.8 m. Its margin is unclear, shape regular, aspect ratio <1, and internal echo relatively homogeneous. No internal blood flow is seen on CDFI. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: Hypoechoic lesion with calcification in the right thyroid lobe, consistent with TI-RADS 5. Findings: The left thyroid lobe measures about 47.2 m, the right lobe about 50.8 m, and the isthmus about 3.2 m. The thyroid is roughly normal in shape and size. The capsule is smooth and intact. The parenchymal echo is relatively homogeneous. No obvious diffuse coarsening or patchy abnormality is seen. CDFI shows a few fine branching blood-flow signals in the gland. No obvious focal abnormal echo or definite nodule is seen in the left lobe. A solid hypoechoic nodule is seen in the right lobe, about 16.3×13.8 m. Its margin is relatively clear, shape roughly regular, aspect ratio<1, and internal echo relatively homogeneous. No obvious strong echo with posterior shadow is seen. CDFI shows a few punctate and strip-like blood-flow signals around and inside the nodule. No focal abnormal echo is seen in the isthmus. No obvious enlarged cervi cal lymph-node echo is seen. Impression: (1) Solid hypoechoic nodule in the right lobe, 16×14 m, TI-RADS 3. Follow-up is recommended. Thyroid function tests and fine-needle aspiration cytology may be considered if needed. (2)No obvious abnormality is seen in the remaining thyroid tissue or isthmus. Ultrasound Video Each video requires writing a report both without any assistance and with the help of AI. Manual report writing AI-assisted report writing Analysis Reporting duration Video review duration Replay count Revision count Paired report Bias-Controlled Evaluation • Cases assessed in both workflows • Different doctors per workflow • Design prevents memory bias • Manual: video-only + blank template • AI-assisted: AI report + tool evidence Review Condition Physicians a Overall (n=145) Benign (n=72) Malignant (n=73) 0 20 40 60 80 100 Consistency (%) +15.9 p 70.3% 86.2% +18.1 p 61.1% 79.2% +13.7 p 79.5% 93.2% Manual AI-assisted Diagnostic consistency b c -1 0 -5 0 5 10 15 Manual - AI time (min) Paired cases sorted by time saving 27.4% shorter overall Within-case time saving -13.0% -36.8% Doctor 1Doctor 2 Manual (n=73) AI-assisted (n=72) Manual (n=72) AI-assisted (n=73) Reporting time (min/case) Physician-level reporting time d e Figure 6.Reader study of AI-assisted thyroid ultrasound reporting.a, Cross-over reader-study design. Each ultrasound video was interpreted under both manual and AI-assisted conditions by different physicians, reducing recall bias while enabling paired case-level comparisons.b, Representative malignant thyroid nodule case comparing reports generated by Qwen 3.5, GPT-5 and ThyroidXAgent with the reference report. Text spans are annotated as correct, partially correct or incorrect according to medical-semantic concordance; ThyroidXAgent shows closer agreement with the reference in this example.c, Annotation-based diagnostic-direction consistency of manual and AI-assisted reports, shown overall and stratified by benign and malignant cases.d, Case-level reporting-time reduction, defined as manual minus AI-assisted reporting time and ranked across paired cases.e, Physician-level reporting time under the manual and AI-assisted conditions. 13/51 requiring clinicians to accept an opaque recommendation. AI assistance reduced segmentation and reporting time while preserving meaningful points of clinical control, including lesion-boundary correction, evidence inspection and report editing. These findings should not be interpreted as evidence for replacing clinical judgement. Instead, they suggest that agentic systems can reduce repetitive work while keeping consequential intermediate outputs available for verification and correction. The effectiveness of such systems will ultimately depend not only on predictive accuracy, but also on whether clinicians can identify errors, understand their downstream consequences and efficiently intervene. Several limitations remain. The evaluations were retrospective, and prospective multicentre studies are needed under real acquisition conditions, changing scanner settings and institution-specific reporting practices. The reader studies included a limited number of physicians and primarily assessed workflow feasibility and efficiency rather than patient-level clinical benefit. ThyroidXAgent also remains dependent on the validity and calibration of its component tools: routing cannot compensate for systematically biased segmentation, incomplete metadata or poorly generalized classifiers, although visible intermediate evidence may make these failures easier to detect 34,76 . More generalizable tool design may further improve cross-domain robustness 77,78 . Finally, the template-based reporting strategy may require local adaptation, and learning from clinician corrections will require governance for data quality, privacy, provenance, model versioning and distribution shift. Future work should therefore determine whether evidence-level corrections improve downstream clinical decisions prospectively and whether the shared evidence architecture can support additional imaging tasks without sacrificing reliability or interpretability. 4Methods 4.1Datasets and task definitions For image segmentation and benign-malignant classification, OpenThyroidDB integrates seven public and insti- tutional ultrasound sources comprising 38,864 images (18,277 training, 850 validation, 19,737 test; Supplementary Table S1). TN3K 11 , TN5K 55 , ThyroidXL 59 and PKTN 13 served as internal cohorts for nodule segmentation train- ing; TN3K, TN5K and ThyroidXL additionally provided benign-malignant classification labels. DDTI 60 , RJH-7K 61 and ZJH-8K (collected at Zhujiang Hospital, Southern Medical University) served as independent external test sets: DDTI and ZJH-8K for both tasks, RJH-7K for segmentation only. To construct the expert pool, we merged training portions across datasets into stacked training sets, with the largest containing 18,277 images. The included cohorts differed markedly in sample size, class balance and lesion characteristics (Supplementary Fig. S2). For malignant-lesion stratification, two additional tasks were evaluated: lateral lymph-node metastasis (LNM) prediction and follicular (FTC) versus papillary (PTC) thyroid carcinoma subtype classification. The training data for both tasks comprised the 4,756 malignant images from ZJH-8K, of which 20 cases (183 images) were held out for validation. Within this cohort, 533 cases (3,394 images) were CN0 and 190 cases (1,362 images) CN1 for lymph- node status, and 656 cases (4,312 images) were classic PTC and 67 cases (444 images) follicular variant; all labels were confirmed by post-surgical histopathology. For LNM prediction, 180 images from LymphUs Center 1 79,80 were added to the training set (total 4,753 images), and the 158 images from LymphUs Center 2 served as the independent external test set. LymphUs is a multicenter open-access database of patients with histologically confirmed PTC and binary LNM labels confirmed by fine-needle aspiration biopsy, acquired using a Samsung Medison RS80 scanner (5- 12 MHz L5-12/60 transducer) at Center 1 and a SuperSonic Imagine AixPlorer Ultimate scanner (4-15 MHz SL15-4 transducer) at Center 2. For FTC-versus-PTC classification, the training and validation splits were unchanged, and the 200 public images from Dai et al. 62 served as the external test set. 14/51 4.2Agent workflow controller and evidence store ThyroidXAgent decomposes each case into tool calls and structured intermediate outputs. The workflow con- tains an expert pool for image segmentation and classification, a radiomics branch, a tabular prediction branch, post hoc explanation modules, anatomical-context parsers, measurement tools and report-generation modules. The LLM router operates on structured summaries rather than raw ultrasound images. During inference, it receives can- didate masks, class probabilities, confidence estimates, radiomic descriptors and metadata such as image resolution, device and data source. It emits a strict JSON decision that records the selected output, supporting evidence and uncertainty signals. These outputs are normalized into a case-level evidence store containing masks, measurements, class probabilities, radiomic descriptors, explanation objects, warnings and report clauses. The evidence store is used for final prediction, report assembly and clinician review, and clinician corrections can be written back to the same record for subsequent use. 4.3Segmentation, classification and radiomics ThyroidXAgent coordinates a heterogeneous expert pool, an LLM router and a radiomics branch for image segmentation and classification: the expert pool generates candidate masks and class probabilities, the router selects the most reliable candidate based on quality metrics and case-level metadata, and the radiomics branch provides an independent classification signal. For image segmentation and classification, the expert pool is designed as a heterogeneous ensemble rather than a single model to reduce sensitivity to dataset bias 63,81 (Supplementary Fig. S1). Each expert typically uses a DINOv3-based backbone 82 with a task-specific lightweight head, though task- specific external models can also be assembled for specialized classification targets. To encourage complementary generalization profiles, these experts were trained under varying configurations along three dimensions: stacked- training composition, input resolution (128, 224 and 448 pixels), and whether DINOv3 pretrained weights were loaded or the backbone was trained from scratch. For the segmentation branch, backbone dilation rates were additionally varied to produce experts with different receptive-field profiles. This heterogeneity exposes individual experts to progressively broader data distributions, so that the router can select the most reliable candidate for each case rather than relying on a single model’s bias. The effect of these stacked-training configurations on cross-dataset performance is reported in Supplementary Table S6. Within this pool, the segmentation branch uses a U-Net-style decoder with skip fusion to preserve fine boundary detail, and is optimized with a combined loss that balances pixel-level supervision with region-level overlap: L seg =L wBCE (m,ˆm)+L IoU (m,ˆm), wheremdenotes the ground-truth mask,ˆmthe prediction,L wBCE the class-weighted binary cross-entropy andL IoU the intersection-over-union loss. The classification branch pools backbone features using global average and max pooling, followed by a compact attention head. To mitigate class imbalance, generalized logit adjustment (GLA) 83 is applied to the classification logits: ̃z c =z c +τlog(π c ), wherez c is the logit for classc,π c is the empirical class prior estimated from the training set, andτis adaptively set based on the class imbalance ratio. Binary classification tasks use BCE on the adjusted logits, whereas multi-class tasks use cross-entropy on the adjusted logits. Each expert produces a class probabilityp i ; the per-expert confidence c i is taken as the maximum softmax output of the classification head for experti. Once these per-expert outputs are available, task-specific quality metrics are computed across them to inform the selection: morphological plausibility such as area, circularity and compactness, and inter-model agreement 15/51 measured by pairwise IoU and HD95, for segmentation; and prediction uncertainty such as entropy and margin, and class consensus, for classification. The LLM router then selects the best mask and classification result by reasoning over these quality metrics, per-expert confidence estimates and case-level metadata, rather than by simple confidence maximization or majority voting. Depending on the configured ensemble size, the router selects either the single best expert or a subset of experts for weighted ensemble fusion. This routing design allows the system to adapt its selection to the acquisition conditions of each case, rather than relying on a fixed model ranking. To provide a classification signal independent of the image-based experts, the radiomics branch processes the selected mask. The mask is first refined via connected-component analysis to remove isolated noisy regions when multiple disconnected components are present. Two-dimensional PyRadiomics descriptors 56 , including shape, in- tensity and texture features, are then extracted from the refined mask-image pair, yielding the radiomic descriptor vector. These descriptors are passed to AutoGluon-tabular classifiers 84 , which ensemble multiple tabular models un- der automated hyperparameter optimization. The resulting tabular class prediction is stored alongside the router’s selection in the case-level evidence store, providing a complementary interpretive signal for clinician review. After selection, SHAP analysis 72 is applied to the tabular classifier to estimate global and local feature contributions, while Grad-CAM 85 is applied to the selected segmentation model to visualize the image regions driving its mask prediction. These post hoc explanations are stored in the evidence store for clinician inspection. Importantly, the same segmentation, radiomics, tabular classification and explanation workflow is reused across benign-malignant classification and additional malignant-lesion stratification tasks, including LNM prediction and FTC/PTC sub- type classification; only task-specific classifier fine-tuning, the task description supplied to the router, and, where applicable, external model assembly are changed. This reuse allows the agent workflow to be redirected to new clinical questions without rebuilding the diagnostic pipeline. 4.4ThyClinScore ThyClinScore evaluates thyroid ultrasound reports as a structured clinical semantic agreement task. It measures both report completeness and semantic consistency with the reference report. Semantic consistency is assessed at two levels: gland-level agreement for thyroid measurements, parenchymal morphology and gland-level vascularity; and lesion-level agreement for lesion detection, lesion size, lesion descriptors and lesion-level vascularity. Size and vascularity can therefore be scored at either level when the corresponding fields are available, whereas morphology mainly captures concept-level agreement in gland and parenchymal descriptions. For each case, the reference reportR gt and generated reportR pred were converted into a shared schema containing thyroid measurements, parenchymal findings, lesion attributes, lymph-node findings and diagnostic impressions. Measurements were standardized in millimetres, categorical fields were restricted to predefined thyroid ultrasound descriptors, and absent information was represented as null. The schema retained clinically relevant information, including lesion location, size, composition, echogenicity, margin, shape, echogenic foci, vascularity and TI-RADS category. These fields supported subsequent lesion-level matching and semantic scoring. At the lesion level, reference and predicted lesions were first matched before lesion detection and attribute agreement were evaluated. LetG=g i m i=1 denote reference lesions andP=p j n j=1 denote predicted lesions. For each candidate pair, we computed S i j =M loc i j ( α s S size i j +β f S feat i j ) ,S size i j = min(d i ,d j ) max(d i ,d j ) , whered i andd j are maximum lesion diameters,M loc i j ∈0,1is a hard anatomical-location gate, andS feat i j ∈0,0.5,1 scores lesion-composition agreement. Explicit left-right mismatches were assignedM loc i j =0. We solved the bipartite assignment using c i j = 1 − S i j and retained pairs above a predefined threshold. Matched pairs were treated as true 16/51 positives, unmatched reference lesions as false negatives and unmatched predicted lesions as false positives, from which lesion-level precision, recall, F1 and false discovery rate were calculated. Numerical measurements were scored using mean relative error (MRE), applied to thyroid lobe and isthmus measurements at the gland level and lesion dimensions at the lesion level. For paired dimensionsK, MRE K = 1 |K| ∑ k∈K |p k −g k | max(|g k |,ε) ,s size (MRE K ; τ)=2 −(MRE K /τ) 2 , where εprovides numerical stability andτcontrols tolerance to size error. Vascularity was evaluated at either level when available. It was discretized into four grades, from absent flow to markedly increased flow, and scored as s vasc =max ( 0,1− |v gt −v pred | 3 ) . Morphological consistency was computed at the concept level for gland and parenchymal descriptions. A thyroid- specific lexicon mapped report phrases to concepts covering echogenicity, texture, margin, shape, calcification, posterior acoustic features and composition. For concept setsC gt andC pred , morphology agreement was scored by concept F1. For matched lesions, categorical descriptors were scored by mean accuracy over reference-present fields, including composition, echogenicity, margin, shape and echogenic foci. Gland-level and lesion-level assessments were combined into a clinical consistency scoreC, which aggregates thyroid-size agreement, lesion-size agreement, vascularity agreement, lesion-detection F1, lesion-feature accuracy and morphology agreement. Components not applicable to either report were excluded from the denominator, whereas reference-present fields missing from the generated report contributed zero: C= ∑ q∈Q w q s q ∑ q∈Q w q , wheres q denotes a valid component score andw q its predefined weight. Report completenessBwas defined as the weighted fraction of required information present in the generated structured report: B= ∑ r∈R u r b r ∑ r∈R u r . The completeness groups were thyroid measurements, parenchyma, lesions, impression and lymph-node description; b r indicates presence of groupr, andu r denotes its predefined weight. The final ThyClinScore was ThyClinScore=[λB+(1−λ)C][η+(1−η)F L ], for cases with reference lesions, where λbalances completeness and clinical consistency, andηcontrols the minimum lesion-detection gate. For reports without reference lesions, the gate was omitted. This design rewards complete and semantically consistent reports while penalizing missed or hallucinated lesions. 4.5Report generation For the reporting branch, thyroid ultrasound reporting was formulated as a case-level evidence-to-report task rather than as single-image captioning. Given an image setI=I j N j=1 , which could include video frames, static images, multiple anatomical views, greyscale ultrasound and colour Doppler images, the workflow generated a structured reportRand an accompanying evidence trace. The report covered thyroid gland morphology, nodule- level findings, lymph-node findings when present and the diagnostic impression. The trace recorded the intermediate observations supporting each report component. Input preparation converted heterogeneous case files into image priors for subsequent planning. Images were standardized and processed by auxiliary modules for region-of-interest cropping, anatomical-context parsing, colour 17/51 Doppler identification, nodule-presence triage and pixel-spacing estimation when required. The resulting priors encoded region, view, modality, Doppler status and nodule likelihood. They were used to guide downstream analysis, rather than inserted directly into report text. The reporting workflow was packaged as a single reusable skill. The skill defined the reporting objective, evidence schema, staged diagnostic process and constraints for evidence-grounded writing. External agents invoked this skill through a Workflow Model Context Protocol (MCP) server, which exposed high-level case operations for input preparation, case initiation, plan review and approval, and report and evidence retrieval. Low-level operations, including segmentation, classification, measurement and captioning, remained internal to the workflow. This separated a stable public interface from the image-processing and model-execution details needed to construct reliable evidence. It also made the resulting report traceable for clinician review. After input preparation, ThyroidXAgent used a planner-executor design 49 . The planner received the skill instructions, tool contract and image priors, and generated a case-specific staged task graph covering gland as- sessment, nodule analysis, lymph-node assessment, evidence fusion and report generation. This graph provided global diagnostic constraints, while allowing case-specific execution. In the deployable workflow, the plan could be inspected and approved before model execution. The executor then followed the approved graph and used ReAct-style local decision-making 52 to select images and internal tools according to intermediate observations. For example, cases without nodule priors could bypass nodule-feature classification, whereas lateral-neck images could trigger lymph-node screening before report assembly. Selected model operations were executed by the internal diagnostic runtime. In deployment, the Workflow MCP process remained lightweight and delegated GPU inference to a separate tool service. The runtime could call preprocessing modules, gland captioning, spacing prediction, thyroid and nodule measurement, nodule segmentation, nodule-feature classification, malignancy classification, cervical lymph-node screening and nodule-level fusion. Their outputs were normalized into a shared evidence object before language generation. Gland-level evidence included thyroid lobe and isthmus measurements, parenchymal morphology and vascularity. Nodule-level evidence included anatomical location, size, composition, echogenicity, margin, shape, echogenic foci, vascularity, segmentation-derived measurements and risk-related predictions. Lymph-node evidence summarized cervical-region screening results. When multiple views described the same lesion, the fusion stage consolidated compatible findings into case-level nodule entries and retained warnings for incomplete or conflicting outputs. Report text was generated by controlled data-to-text assembly 57 . To reduce factual hallucination risk 58 , Thy- roidXAgent used training-free BM25 template retrieval, slot filling and clause assembly rather than unconstrained free-text decoding. Corresponding template libraries were constructed from the SMU-HMC and KMVE datasets. The SMU-HMC was derived from 7,147 quality-controlled reports after excluding all reports from the 400 patients reserved for internal testing. The KMVE was constructed from all 1,719 reports in its training partition, with the validation and test partitions excluded. During template construction, whole reports were decomposed into five clinically defined clause categories: thyroid measurement, gland morphology, nodule findings, lymph-node findings and ultrasound impression. During inference, the evidence object assembled through tool execution was partitioned into corresponding evidence blocks and converted into category-specific BM25 queries. Retrieved templates served as linguistic frames; slot filling inserted case-specific measurements, descriptors, TI-RADS-relevant information, risk categories and other tool-derived evidence, and clause assembly combined the completed findings and impression clauses into the final report. Notably, KMVE contains only the ultrasound findings section and does not provide original measurement values. The template library constructed from KMVE therefore retained the original structure and linguistic style of the KMVE findings and was used in accordance with its original findings-only evaluation protocol. Because 18/51 tool execution was decoupled from report realization, the retrieval libraries could be exchanged in a plug-and- play manner to align generated reports with centre- or corpus-specific writing conventions, without retraining the upstream image-analysis tools or modifying the underlying evidence schema. The same procedure can be used to construct updated retrieval libraries from new report sources (Supplementary Fig.S4c). The final artefact consisted of the report text and its supporting evidence trace. Clinicians or external agents could inspect the diagnostic plan, verify the evidence used for each statement, review warnings and edit the gener- ated report. Corrected reports could then be returned to the case-level evidence store for subsequent review and model updating. This design kept report generation constrained by structured clinical evidence while preserving a reviewable path from image inputs to final report statements. 4.6Reader studies and statistical analysis For segmentation review, clinicians corrected AI-generated masks, and correction time and paired Dice scores were compared with manual segmentation. For report writing, two physicians evaluated 145 thyroid ultrasound videos with annotation-derived diagnostic labels in a cross-over design. Each case was interpreted under both manual and AI-assisted conditions, but by different physicians, to reduce memory and recall bias. In the manual condition, physicians reviewed the ultrasound video and completed a blank structured-report template. In the AI-assisted condition, they reviewed an AI-generated draft together with the supporting tool-derived evidence and revised the report as required. We recorded reporting time, video-review time and the final submitted report for each assessment. Reporting efficiency was assessed at both the case and physician levels, with case-level time saving defined as the manual minus AI-assisted reporting time. For annotation-based diagnostic-direction analysis, frame-level benign and malignant bounding-box annotations were aggregated into case-level labels. Cases containing any malignant annotation were classified as malignant or suspicious, whereas cases containing only benign annotations were classified as benign. This yielded 145 labelled cases, comprising 72 benign and 73 malignant or suspicious cases. A prespecified rule-based procedure assigned each submitted report to a benign or low-risk, malignant or suspicious, or ambiguous direction using TI-RADS categories and diagnostic terms. Directional consistency was defined by agreement with the annotation-derived case label and was summarized overall and within the benign and malignant or suspicious strata. For statistical analysis, performance was summarized using the primary metric for each task: Dice and HD95 for segmentation, AUROC and AUPRC for classification, conventional natural-language generation metrics for report wording, and ThyClinScore and its submetrics for report semantics. Confidence intervals in the supplementary tables were estimated by nonparametric bootstrap resampling of the test set. Correlations between report metrics and the location-aware LLM judge were evaluated using two-sided Pearson correlation tests. 5Data availability The public datasets used in this study are available from their original sources, as cited in the Methods and Sup- plementary TableS1. The publicly released data associated with this study are available through ThyroidOpenDB at https://huggingface.co/datasets/MedXAgent/ThyroidOpenDB. The study was approved by the Medical Ethics Committee of Zhujiang Hospital, Southern Medical University (approval no. 2026-KY-081-01). The NHCMISD dataset was collected from real-world clinical ultrasound cases provided by the National Health Commission Medical Imaging Standard Database, for which the co-author, Prof. Dexing Kong, has authorized access. All data usage complied with relevant institutional and regulatory requirements. 19/51 6Code availability The source code for ThyroidXAgent is publicly available on GitHub athttps://github.com/MedXAgent/ ThyroidXAgent. The website will be made available after acceptance. The trained model weights are available on Hugging Face athttps://huggingface.co/MedXAgent/ThyroidXAgent. The repositories contain the resources required to reproduce the reported computational analyses. 7Acknowledgements This work was supported in part by the National Natural Science Foundation of China (Grant No. 62322608), the Zhejiang Provincial Natural Science Foundation of China (Grant No. LQN26F020029), and the Natural Science Foundation of Guangdong Province (Grant No. 2024A1515010255). 8Author contributions •Conceptualization:H.G. and G.L. •System development:S.C., B.W. and X.X. •Computational evaluation:S.C., B.W. and S.W. •Visualization:S.C., B.W. and H.G. •Reader study:H.W., Q.L. and F.C. •Data curation and organization:H.G., S.C., B.W., Y.W., G.Y., H.W., M.M., D.K., Q.L., W.L. and F.C. •Writing–original draft:H.G., Y.W., S.C. and B.W. •Writing–review and editing:All authors. •Supervision:Q.L., W.L., F.C. and G.L. 9Confilts The authors has no confilt of interests. 20/51 Supplementary information Supplementary figures Figure S1.ThyroidXAgent architecture for thyroid nodule segmentation and classification. . . . . . . . . . . . . . . . . . . .22 Figure S2.Dataset composition and lesion distributions across thyroid ultrasound cohorts. . . . . . . . . . . . . . . . . . . .23 Figure S3.Case-level explanations for thyroid nodule classification. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 Figure S4.Multimodal input processing and case interaction in ThyroidXAgent. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 Figure S5.Planning, tool execution and clinician review in the report-generation workflow. . . . . . . . . . . . . . . . . . .26 Figure S6.Template-bank construction for centre-specific report generation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .27 Figure S7.Qualitative comparison of thyroid ultrasound report generation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .28 Supplementary tables Table S1.Composition and data splits of the multicentre thyroid ultrasound benchmark. . . . . . . . . . . . . . . . . . . . . .29 Table S2.Characteristics of public and institutional thyroid ultrasound datasets. . . . . . . . . . . . . . . . . . . . . . . . . . . . .30 Table S3.Cross-dataset generalization for thyroid nodule segmentation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .31 Table S4.Cross-dataset generalization for benign-malignant thyroid nodule classification. . . . . . . . . . . . . . . . . . . . .32 Table S5.Performance on malignant thyroid lesion stratification. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 Table S6.Effects of cumulative training-data integration on segmentation and classification. . . . . . . . . . . . . . . . . .34 Table S7.Centre-specific Dice for thyroid gland segmentation in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . . . . . . . . . . .35 Table S8.Centre-specific HD95 for thyroid gland segmentation in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . . . . . . . . .36 Table S9.Centre-specific Dice for thyroid nodule segmentation in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . . . . . . . . .37 Table S10.Centre-specific HD95 for thyroid nodule segmentation in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . . . . . . .38 Table S11.Centre-specific AUROC for benign-malignant classification in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . .39 Table S12.Centre-specific AUPRC for benign-malignant classification in NHC-MISD-TUS. . . . . . . . . . . . . . . . . . .40 Table S13.Performance of preprocessing and executor tools in ThyroidXAgent. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 Table S14.Lexical performance of thyroid ultrasound report generation across datasets. . . . . . . . . . . . . . . . . . . . . .42 Table S15.Clinical semantic performance of thyroid ultrasound report generation across datasets. . . . . . . . . . . .43 Table S16.Static-pipeline metrics used for report-generation radar plots. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .44 Table S17.Ablation of tool integration for thyroid ultrasound report generation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .45 21/51 Tool Factory Model Model Factory DL model predictionRadiomics ...... Meta dataAutoGluon prediction Tiger-ModelLLNM-NetDino v3 Benign / Malignant PTC / FTC LNM-negative / positive Final prediction a b c Figure S1.ThyroidXAgent architecture for thyroid nodule segmentation and classification. DINOv3-based experts generate candidate segmentation masks and malignancy probabilities. The radiomics branch extracts PyRadiomics features from the selected lesion mask, applies an AutoGluon classifier and computes SHAP attributions. The LLM router integrates the candidate outputs with image metadata, including resolution, device and data source, to select the final prediction and its supporting evidence. 22/51 Benign Malignant 3,202 304 45 3,459 8,172 3,574 1,426 2,638 2,709 TN3KTN5KThyroidXLDDTIZJH-8K Dataset 10 4 10 3 10 2 Number of Images 4,756 0.5 0.0 1.0 Pos y 0.00.51.0 0.5 0.0 1.0 Pos y 0.00.51.0 0.5 0.0 1.0 Pos y 0.00.51.0 0.5 0.0 1.0 Pos y 0.00.51.0 0.5 0.0 1.0 Pos_y 0.00.51.0 0.5 0.0 1.0 Pos y 0.00.51.0 Posx Posx Posx Posx Posx Posx TN3K(N=5,347) TN5K(N=5,000) ThyroidXL(N=11,631)DDTI(N=349) ZJH-8K(N=7,958) RJH-7K(N=7,288) D ensity 6 4 2 0 0.00.20.40.60.81.0 D ensity 12 8 4 0 0.00.20.40.60.81.0 D ensity 6 4 2 0 0.00.20.40.60.81.0 Density 15 10 5 0 0.00.20.40.60.81.0 D ensity 4.5 3.0 1.5 0.0 0.00.20.40.60.81.0 D ensity 6 4 2 0 0.00.20.40.60.81.0 RelativesizeRelativesize RelativesizeRelativesize RelativesizeRelativesize b a Figure S2.Dataset composition and lesion distributions across thyroid ultrasound cohorts. Top, grouped bar chart comparing the numbers of benign and malignant images in TN3K, TN5K, ThyroidXL, DDTI and ZJH-8K on a logarithmic y axis, highlighting marked variation in cohort size and class balance across cohorts. Bottom, for each of the five datasets, two-dimensional kernel density estimates of normalized lesion-mask centroid positions (left) and relative lesion size distributions, defined as mask area divided by image area (right). All size distributions share a common x-axis range, and all spatial maps use a common density scale, enabling direct cross-dataset comparison of where lesions appear within the ultrasound frame and how large they tend to be. 23/51 a SHAPE: Sphericity GLCM: Corr SHAPE: Elongation RLM: LRHGLE SHAPE: Perim/Area -2-1012 Toward malignant Toward benign SHAP contribution toward malignant (raw score) benign SHAPE: Elongation RLM: RLNU SHAPE: Sphericity TEX(RLM): RLNUN RLM: SRLGLE -0.4-0.20.00.20.4 Toward malignant Toward benign SHAPE:Perim/Area TEX(GLCM):ClusterProminence RLM:LRHGLE GLCM:Corr SHAPE:Perimeter SHAP contribution toward malignant (raw score) b malignant SHAPE: Elongation SHAPE:Sphericity RLM:SRE SHAPE: Perimeter TEX(RLM):RLNUN INT:Skewness RLM:SRLGLE TEX(GLCM):ClusterShade INT:Median RLM:LRHGLE Toward malignant Toward benign -1.5-0.50.00.51.5 SHAP contribution toward malignant (raw score) 1.0-1.0 INT:InterquartileRange RLM:SRLGLE NGTDM:Contrast INT:Variance TEX(DM):GrayLevelV malignant SHAPE:Sphericity RLM:LRHGLE INT:Skewness INT:Median TEX(GLCM):ClusterShade INT:10thpctl INT:Kurtosis RLM:LRLGLE INT:Mean TEX(GLCM):JointAverage Toward malignant Toward benign SHAP contribution toward malignant (raw score) -2-1012-33 benign d c Figure S3.Case-level explanations for thyroid nodule classification. Left, SHAP values for the most influential radiomic features; red and blue indicate contributions towards malignant and benign predictions, respectively. Right, ultrasound images with segmentation contours and Grad-CAM maps from the selected segmentation model. Representative benign and malignant cases with accurate and inaccurate segmentation are shown. 24/51 a bc Figure S4.Multimodal input processing and case interaction in ThyroidXAgent.a, Agent Chat interface for initiating and monitoring case-level thyroid ultrasound analysis, with the case context and workflow activity displayed alongside the conversation.b, Visual perception of case-level, multiview thyroid ultrasound inputs, including image selection, anatomical-region recognition and organization of image-context priors.c, Video-input processing, in which an ultrasound video is sampled into representative frames and organized with anatomical-region predictions and image-context priors for downstream analysis. 25/51 ab c Figure S5.Planning, tool execution and clinician review in the report-generation workflow.a, Case-specific planning, in which preprocessing outputs, workflow instructions and diagnostic objectives are converted into a staged plan for review and approval.b, Execution after plan approval, showing internal tool calls, intermediate outputs, evidence collection and progression from the approved plan to a reviewable report.c, Report Review Dashboard displaying the generated report alongside source images, structured evidence and tool traces for clinician inspection and editing. 26/51 Figure S6.Template-bank construction for centre-specific report generation. The interface enables users to select an existing centre-specific report-generation profile or import a new report corpus. Generic template-construction scripts derive dataset-specific template banks to support adaptation to additional reporting data. 27/51 Ground Truth Findings: The right thyroid lobe measure s 47×16×14 m, the left lobe 50×16×8.8 m, and the isthmus 3 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is slightly heterogeneous. No definite ech ogenic mass or cystic dark area is se en in the left lobe or ist hmus. CDFI sh ows a few fine branching color blood-flow signals. A hypoechoic lesion is seen n ear the isthmus in the right lobe, about 6.8×8 m. Its margin is unclear. Scattere d punct ate ech ogenic foci are prese nt. CDFI sh ows internal blood flow. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: Hypoechoic lesion with calcification near the isthmus in the right thyr oid lobe, co nsi st ent with TI-RADS 5; puncture biopsy is recommended. No focal lesion is seen in the left thyroid lobe or isthmus. CorrectPart ially co rrectInco rrect Qwen 3.5 GPT-5 Ours Findings: The left thyroid lobe measures about 48×15×13 m, the right lobe about 47×16×12 m, and the isthmus about 2.0 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is heterogeneous. CDFI shows a few fine branching color blood-flow signals in the gland. A solid hypoechoic nodule is seen in the left lobe, about 13×9 m, Figure 1. It has clear margin, regu lar shape, aspect ratio <1, and homogeneous internal echo. No obvious internal blood flow is seen on CDFI. An anechoic area is also seen in the upper pole of the left lobe, about 5×4 m, Figure 3. It has clear margin and posterior enhancement, with no internal blood flow on CDFI. A solid isoechoic nodule is seen in the right lobe, about 10×7 m, Figure 3. It has clear margin, regular shape, and ho mogeneous internal echo. CDFI shows punctate and strip-like blood flow around it. No obvious abnormal enlarged cervical lymph-node echo is seen. Impression: Solid hypoechoic nodule in the left thyroid lobe, consistent with TI-RADS 3. Cystic nodule in the left thyroid lobe, consistent with TI-RADS 2. Solid isoechoic nodule in the right thyroid lobe, consistent with TI-RADS 3. Findings: The left thyroid lobe measures about 35×15×13 m, the right lobe about 35×11×10 m, and the isthmus about 3.5 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is homogeneous. No definite echogenic mass or cystic dark area is seen in the left lobe or isthmus. CDFI shows a few fine branching color blood-flow signals. A hypoechoic lesion is seen in th e right lobe, about 7.9×6.8 m. Its margin is unclear, shape regular, aspect ratio <1, and internal echo relatively homogeneous. No internal blood flow is seen on CDFI. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: Hypoechoic lesion with calcification in the right thyroid lobe, consistent with TI-RADS 5. Findings: The left thyroid lobe measures about 47.2 m, the right lobe about 50.8 m, and the isthmus about 3.2 m. The thyroid is roughly normal in shape and size. The capsule is smooth and intact. The parenchymal echo is relatively homogeneous. No obvious diffuse coarsening or patchy abnormality is seen. CDFI shows a few fine branching blood-flow signals in the gland. No obvious focal abnormal echo or definite nodule is seen in the left lobe. A solid hypoechoic nodule is seen in the right lobe, about 16.3×13.8 m. Its margin is relatively clear, shape roughly regular, aspect ratio<1, and internal echo relatively homogeneous. No obvious strong echo with posterior shadow is seen. CDFI shows a few punctate and strip-like blood-flow signals around and inside the nodule. No focal abnormal echo is seen in the isthmus. No obvious enlarged cervi cal lymph-node echo is seen. Impression: (1) Solid hypoechoic nodule in the right lobe, 16×14 m, TI-RADS 3. Follow-up is recommended. Thyroid function tests and fine-needle aspiration cytology may be considered if needed. (2)No obvious abnormality is seen in the remaining thyroid tissue or isthmus. Qwen 3.5 GPT-5 Ours Findings: The left thyroid lobe measures about 50×19×20 m, the right lobe about 48×17×16 m, and the isthmus about 3 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is heterogeneous. CDFI shows relatively rich blood-flow signals in the gland. A solid hypoechoic nodule is seen in the left lobe, about 19×11 m. Its margin is unclear, shape irregular , aspect ratio >1, and internal echo heterogeneous. Punctate echogenic foci are present, with posterior acoustic attenuation. CDFI shows punctate and strip-like blood flow inside and around it. Another anechoic area is seen, about 3×2 m, with clear margin and no internal blood flow on CDFI. No obvious abnormal cervical lymph-node echo is seen. Impression: Solid hypoechoic nodule in the left thyroi d lobe, consistent with TI-RADS 4b; puncture biopsy is recommended. Cystic nodule in the left thyroid lobe, consistent with TI-RADS 2. Heterogeneous thyroid parenchymal echo; clinical correlation is suggested. Findings: The left thyroid lobe measures about 38×25×11 m, the right lobe about 52×12×14 m, and the isthmus about 3.9 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is homoge- neous. No definite echogenic mass or cystic dark area is seen in the left lobe or isthmus. CDFI shows a few fine branching color blood-flow signals. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: No focal thyroid lesion is detected. Findings: The left thyroid lobe is about 4.77 cm thick, measured on longitudinal scan, with other diameters not shown. The right lobe is about 4.44 cm thick, and the isthmus about 0.21 cm thick. The thyroid is enlarged. The capsule is basically smooth and intact. The parenchymal echo is diffusely coarse and heterogeneous. Scattered small echogenic foci and hypoechoic areas are seen. CDFI shows relatively rich punctate and strip-like blood-flow signals. A solid hypo- to isoechoic nodule is see n in the left lobe, about 1.30×1.01 cm. Its margin is relatively clear, shape nearly oval, and internal echo slightly heterogeneous. No obvious posterior shadow or enhancement is seen. CDFI shows a little internal blood flow. Another anechoic area with posterior enhancement is seen, about 1.23×1.28 cm, with clear margin, regular shape, and no obvious internal blood flow on CDFI. A solid iso- to hy poechoic nodule is seen in the right lobe, about 1.28×1.23 cm. Its margin is relatively clear, shape nearly oval, and internal echo slightly heterogeneous. CDFI shows a little internal blood flow. No obvious enlarged abnormal cervical lymph-node echo is seen. Impression: (1)Diffuse heterogeneous thyroid echo with increased blood flow; chronic thyroiditis is possible. (2)Left solid nodule about 1.3 cm, TI-RADS 3; left cystic nodule about 1.3 cm, TI-RADS 2. (3)Right solid nodule about 1.3 cm, TI-RADS 3. Thyroid function tests and follow-up are suggested. TI-RADS 3 nodules should be rechecked in 6–12 months. If clinically needed, fine-needle aspiration cytology may be performed. Ground Truth Findings: The right thyroid lobe measures 48×12×13 m, the left lobe 44×13×10 m, and the isthmus 2 m. The thyroid is normal in shape and size. The capsule is smooth. The parenchymal echo is homogeneous. No definite echogenic mass or cystic dark area is seen. CDFI shows a few fine branching color blood-flow signals. No abnormal enlarged lymph node is seen around the bilateral carotid arteries. Impression: No focal thyroid lesion detected. CorrectPart ially co rrectInco rrect Figure S7.Qualitative comparison of thyroid ultrasound report generation. Representative benign and malignant cases compare reports generated by Qwen 3.5, GPT-5 and ThyroidXAgent with the reference reports. Text spans are annotated as clinically correct, partially correct or incorrect. The malignant example corresponds to Fig. 6b; the benign example provides an additional complementary case. 28/51 Table S1.Composition and data splits of the multicentre thyroid ultrasound benchmark. Numbers and percentages show the images contributed by each dataset to the full benchmark (n=38,864) and to the training (n=18,277), validation (n=850) and test (n=19,737) cohorts. DDTI, RJH-7K and ZJH-8K are independent external test cohorts. DDTI comprises 637 images, of which 349 carry benign-malignant classification labels. ZJH-8K serves as an external test set for segmentation and classification; its 4,756 malignant images additionally serve as the primary training set for malignant-lesion stratification (Supplementary TableS5). DatasetTaskCenterUltrasound deviceTotal TrainValidTest TN3K 11 Segmentation, Classification Zhujiang Hospital, Southern Medical University, Guangzhou, China GE Logiq E9, ARIETTA 850, RESONA 70B 5,347 (13.76%) 4,633 (25.35%) 100 (11.76%) 614 (3.11%) TN5K 55 Segmentation, Classification Cancer Hospital, Chinese Academy of Medical Sciences, Beijing, China GE Logiq E9, GE S7 (5-12 MHz or 8-15 MHz) 5,000 (12.87%) 3,500 (19.15%) 500 (58.82%) 1,000 (5.07%) ThyroidXL 59 Segmentation, Classification Vietnam National Hospital of Endocrinology, Hanoi, Vietnam Hitachi Aloka Arietta V70 11,631 (29.93%) 9,441 (51.66%) 100 (11.76%) 2,090 (10.59%) PKTN 13 Segmentation Peking University First Hospital, Beijing, China –1,003 (2.58%) 703 (3.85%) 150 (17.65%) 150 (0.76%) DDTI 60 Segmentation, Classification IDIME, Bogotá, Colombia TOSHIBA Nemio 30, TOSHIBA Nemio MX (12 MHz probe) 637 (1.64%)– 637 (3.23%) seg 349 cls RJH-7K 61 Segmentation Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China Different machines, not specified 7,288 (18.75%)–7,288 (36.93%) ZJH-8K Segmentation, Classification Zhujiang Hospital, Southern Medical University, Guangzhou, China –7,958 (20.48%)–7,958 (40.32%) 29/51 Table S2.Characteristics of public and institutional thyroid ultrasound datasets. Dataset size, data split, file format, task, geographical source and ultrasound scanner information are summarized for the included resources. DatasetDataset SizeDevelopment/Validation EvaluationFile FormatTaskLocationUltrasonic Imaging Device TGVideo 86 15,186 (16 cases)15,186N/A Image: DICOM Mask: DICOM SegmentationGermanyGE Logiq E9 DDTI 60 637N/A637 Image: PNG Mask: PNG Label: CSV Segmentation Classification Colombia TOSHIBA Nemio 30 TOSHIBA Nemio MX TN3K 11,12,14 5,3474,733614 Image: JPG Mask: JPG Label: CSV Segmentation Classification Guangzhou, China GE Logiq E9 ARIETTA 850 RESONA 70B TN5K 55 5,0004,0001,000 Image: JPG Label: XML Detection Classification Beijing, China GE Logiq E9 GE S7 ThyUS2Path 87 8,5085,4573,051 Image: JPG Label: CSV ClassificationZhejiang, ChinaEsaote MyLab (Portable) Cine-clip 88 17,412 frames 192 cases avg. 90 frames/case N/AN/A Image: HDF5 Label: CSV Segmentation Classification California, USAN/A AHU 89 1,833 cases 125,896 images N/AN/A Image: JPG Label: Folder-level ClassificationChina (web scraping) Heterogeneous ThyroidXL 59 11,6319,5412,090 Image: PNG Mask: PNG Label: TXT Segmentation Classification Detection VietnamHitachi Aloka Arietta V70 PKTN 13 1,003N/AN/A Image: JPG Mask: JPG SegmentationBeijing, ChinaN/A KMVE 18 2,457 cases 4,914 image assignments 1,719 training cases 246 validation cases 492 test cases Image: JPEG Report: JSON Report GenerationBeijing, ChinaN/A SMU-HMC 23,955 reports 248,194 images 23,555 reports 242,210 images 400 reports 5,984 images Image: PNG Report: JSON Report Generation Component Development Guangzhou, ChinaHeterogeneous ZJH-TS 353 reports 4,471 images N/A 150 reports 2,034 images Image: PNG Report: JSON Report Generation External Validation Guangzhou, ChinaHeterogeneous TNVideo148 casesN/A145 labelled cases Video: AVI Segmentation Reader Study Guangzhou, ChinaN/A 30/51 Table S3.Cross-dataset generalization for thyroid nodule segmentation. Dice coefficient and 95th-percentile Hausdorff distance (HD95) are reported with 95% confidence intervals. ModelTN3KThyroidXLPKTNTN5KDDTIZJH-8KRJH-7K Dice (%)↑ TransUnet 90 81.84±1.6285.75±0.5776.89±3.5678.54±1.5176.58±1.6280.72±0.9784.83±0.37 MedSegX 91 83.93±0.7979.98±0.3680.63±0.4283.10±0.4875.12±1.6884.06±0.3985.40±0.18 MedSAM2 64 84.47±1.0286.94±0.3683.46±2.6083.03±1.2984.72±1.2686.29±0.7390.72±0.21 UltraFedFM 65 81.18±1.4684.70±0.5375.31±1.1277.13±1.3875.57±1.6780.64±0.8483.10±0.33 ThyroidXAgent85.28±1.2887.58±0.4482.99±2.1083.26±1.3485.62±1.0794.30±0.3891.46±0.14 HD95 (m)↓ TransUnet 90 27.27±5.5222.42±1.3426.88±9.6622.32±3.4317.12±1.5518.37±0.7518.81±0.74 MedSegX 91 10.95±0.6411.07±0.3210.83±0.7011.76±0.7618.39±1.6510.96±0.359.37±0.18 MedSAM2 64 11.51±1.535.46±0.4410.56±3.6410.94±1.1210.06±1.216.79±0.572.92±0.17 UltraFedFM 65 14.98±2.108.10±0.5816.08±1.6714.96±1.6518.12±1.478.69±0.809.06±0.38 ThyroidXAgent10.31±1.705.43±0.539.01±3.5810.12±1.239.24±1.072.25±0.391.92±0.08 31/51 Table S4.Cross-dataset generalization for benign-malignant thyroid nodule classification. Area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC) are reported with 95% confidence intervals. MethodTN3KThyroidXLTN5KDDTIZJH-8K AUROC↑ ResNet-50 92 0.7674±0.03940.9044±0.01180.9322±0.01680.6704±0.08420.6704±0.0842 RepViT 66 0.5556±0.04630.7774±0.01880.6603±0.03750.6162±0.08040.8538±0.0185 LSNet 93 0.8095±0.03330.9178±0.01140.9091±0.02010.7581±0.06580.8631±0.0201 UltraFedFM 65 0.8461±0.06970.9239±0.01040.9298±0.01750.7518±0.17120.9115±0.0140 MedGemma 71 0.8492±0.03050.9371±0.00950.9442±0.01560.8255±0.06500.8976±0.0166 Qwen3-VL-8B-Instruct 94 0.8237±0.03280.9050±0.01150.9214±0.01870.7361±0.06920.8659±0.0189 GPT-5 68 0.6924±0.04210.7059±0.04690.7737±0.09960.6346±0.09140.6109±0.0515 Gemini-2.5-Pro 69 0.6587±0.04550.6246±0.06400.6873±0.06910.6156±0.13080.6493±0.0516 ThyroidXAgent0.8692±0.03490.9676±0.00660.9472±0.01520.7991±0.07410.9175±0.0167 AUPRC↑ ResNet-50 92 0.6882±0.06320.8882±0.01740.9674±0.02680.3755±0.11760.2755±0.1167 RepViT 66 0.4275±0.05280.7161±0.02760.8403±0.02160.3924±0.09330.9486±0.0078 LSNet 93 0.7581±0.04520.9040±0.01420.9551±0.01340.4180±0.14100.9449±0.0113 UltraFedFM 65 0.8531±0.02840.9354±0.01140.8422±0.04210.4487±0.14520.9669±0.0084 MedGemma 71 0.8047±0.04300.9201±0.01390.9747±0.00840.5537±0.16630.9589±0.0096 Qwen3-VL-8B-Instruct 94 0.7617±0.05110.8787±0.03790.9636±0.01060.4112±0.14150.9498±0.0096 GPT-5 68 0.6627±0.06330.6237±0.06660.8920±0.03160.3578±0.10890.8311±0.0377 Gemini-2.5-Pro 69 0.6205±0.05870.4914±0.08410.8462±0.04460.3924±0.15270.8403±0.0362 ThyroidXAgent0.8545 ± 0.06000.9653 ± 0.00780.9752 ± 0.00890.5863 ± 0.13800.9711 ± 0.0006 32/51 Table S5.Performance on malignant thyroid lesion stratification. AUROC and AUPRC with 95% confidence intervals are reported for lateral lymph-node metastasis prediction and follicular versus papillary thyroid carcinoma subtype classification. Training data comprised 4,756 malignant images from ZJH-8K, of which 20 cases (183 images) were held out for validation. Em dashes indicate tasks that were not evaluated. Method Lymph Node MetastasisFTC/PTC subtype AUROC↑AUPRC↑AUROC↑AUPRC↑ RepViT 66 0.7905±0.06760.8152±0.06380.6419±0.08390.6297±0.0942 LSNet 93 0.5878±0.08650.6301±0.08750.4858±0.09250.4845±0.0908 UltraFedFM 65 0.7757±0.07310.7902±0.08450.7365±0.07440.7582±0.0824 MedGemma 71 0.8403±0.04610.8585±0.05660.6598±0.08240.6142±0.1056 Qwen3-VL-8B-Instruct 94 0.8070±0.06320.8055±0.08000.6056±0.08660.5539±0.1118 GPT-5 68 0.8410±0.05750.8629±0.05330.1604±0.07060.3638±0.0847 Gemini-2.5-Pro 69 0.5414±0.07360.5492±0.09150.3324±0.08720.4187±0.0837 LLNM-Net 22 0.7665±0.06920.7363±0.0849– Tiger-Model 23 –0.7136±0.08140.7117±0.1101 ThyroidXAgent0.8642±0.05500.8808±0.05370.8053±0.05990.7863±0.0793 33/51 Table S6.Effects of cumulative training-data integration on segmentation and classification. Models were trained on progressively expanded configurations and evaluated on independent test sets. Segmentation performance is reported as Dice coefficient (%,↑higher is better) and HD95 (m,↓lower is better); classification performance as AUROC and AUPRC (↑higher is better). Values are means±95% confidence intervals across five independent runs. Dashes indicate that the test set was not applicable for the given task. Segmentation training configurations: dataset1 (TN3K), dataset2 (TN3K + ThyroidXL), dataset3 (TN3K + ThyroidXL + PKTN), dataset4 (TN3K + ThyroidXL + PKTN + TN5K). Classification training configurations: dataset1 (TN3K), dataset2 (TN3K + ThyroidXL), dataset3 (TN3K + ThyroidXL + TN5K). Train Test TN3KThyroidXLPKTNTN5KDDTIZJH-8KRJH-7K Segmentation – Dice (%)↑ dataset182.76±3.5481.97±2.63 79.21±2.71 72.18±5.08 78.08±3.09 94.57±0.43 80.77±0.44 dataset281.63±3.81 86.84±1.90 81.73±2.26 71.07±5.35 76.72±3.3994.85±0.4082.38±0.42 dataset380.81±3.82 86.00±2.31 81.91±2.26 72.82±4.91 84.81±2.43 94.82±0.39 91.44±0.15 dataset481.86±3.7086.97±2.19 83.28±2.19 82.57±3.46 84.86±2.3894.77±0.3991.46±0.15 Segmentation – HD95 (m)↓ dataset113.49±3.838.34±2.34 11.73±3.07 11.37±3.49 16.93±3.122.30±0.47 11.46±0.50 dataset215.92±4.584.99±1.58 9.72±2.71 13.64±4.28 18.54±3.171.98±0.41 9.65±0.44 dataset315.94±4.345.46±1.62 10.92±3.52 11.07±3.28 11.97±3.981.93±0.38 1.87±0.07 dataset417.00±5.344.74±1.42 8.89±2.93 4.64±1.43 9.91±2.182.07±0.40 1.88±0.07 Classification – AUROC↑ dataset10.7666±0.03 0.8713±0.01–0.8272±0.02 0.7244±0.07 0.9924±0.01– dataset20.7724±0.04 0.9254±0.01–0.8151±0.02 0.5762±0.10 0.9932±0.01– dataset30.7906±0.03 0.9288±0.01–0.9515±0.01 0.7623±0.07 0.9937±0.01– Classification – AUPRC↑ dataset10.6806±0.06 0.8147±0.03–0.9151±0.02 0.3578±0.13 0.9951±0.01– dataset20.7237±0.050.9140±0.01–0.9074±0.01 0.3190±0.140.9968±0.01– dataset30.7188±0.050.9144±0.01–0.9803±0.01 0.4029±0.140.9967±0.01– 34/51 Table S7.Per-centre Dice similarity coefficient (%) for thyroid gland segmentation on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFM Dice (%)↑Dice (%)↑Dice (%)↑Dice (%)↑Dice (%)↑ Overall8,721 59.28 (58.70, 59.86) 56.80 (56.35, 57.24) 54.98 (54.57, 55.39) 53.56 (53.06, 54.12) 33.30 (32.73, 33.88) THYB_S_ZJ24 1,17480.64 (79.84, 81.40) 52.96 (51.96, 54.03) 50.32 (49.25, 51.42) 72.91 (72.03, 73.79) 8.49 (7.59, 9.42) THYB_S_EN04 94965.30 (63.82, 66.78) 51.05 (49.46, 52.51) 59.83 (58.66, 60.96) 59.65 (58.11, 61.15) 35.33 (33.81, 36.91) THYB_S_BJ01 876 61.05 (59.14, 62.88) 61.22 (59.93, 62.48) 51.36 (49.83, 52.83) 56.68 (54.95, 58.35) 27.42 (25.73, 29.23) THYB_S_SH01 832 43.33 (41.36, 45.33) 64.88 (63.37, 66.28) 55.60 (54.14, 56.95) 42.72 (40.94, 44.52) 49.82 (48.11, 51.48) THYB_S_ZJ05 69455.25 (53.12, 57.40) 54.83 (53.19, 56.51) 53.77 (52.33, 55.13) 50.71 (48.70, 52.55) 30.49 (28.66, 32.35) THYB_S_SH05 66961.58 (59.74, 63.48) 59.92 (58.52, 61.38) 56.63 (55.27, 57.98) 50.26 (48.46, 52.13) 34.59 (32.66, 36.40) THYB_S_NX01 582 52.69 (50.28, 54.89) 45.59 (44.00, 47.18) 52.80 (51.32, 54.40) 54.11 (51.83, 56.22) 32.64 (30.64, 34.63) THYB_S_ZJ06 460 48.58 (46.12, 50.98) 62.74 (60.83, 64.42) 55.20 (53.35, 56.97) 42.03 (39.70, 44.34) 35.19 (32.95, 37.35) THYB_S_QX07 28966.05 (63.24, 68.91) 50.61 (48.09, 53.10) 59.57 (57.62, 61.64) 62.24 (59.24, 65.06) 46.38 (43.64, 49.40) THYB_S_AN01 243 36.90 (33.32, 40.25) 74.15 (71.96, 76.34) 56.14 (53.14, 59.07) 36.30 (33.35, 39.72) 60.76 (57.91, 63.56) THYB_S_CQ03 23863.69 (60.50, 66.62) 55.01 (52.63, 57.49) 58.52 (56.01, 60.89) 46.06 (42.68, 49.56) 33.93 (30.82, 37.30) THYB_S_GZ02 233 59.92 (56.20, 63.38) 62.23 (60.11, 64.35) 55.84 (53.35, 58.40) 54.40 (51.00, 58.00) 37.71 (34.20, 41.27) THYB_S_JX06 215 54.16 (50.51, 58.01) 63.46 (60.77, 65.99) 55.71 (52.48, 58.52) 45.96 (42.07, 49.54) 46.71 (43.30, 50.23) THYB_S_JS02 190 36.44 (32.39, 40.43) 64.41 (61.96, 66.78) 56.60 (54.36, 58.91) 34.57 (31.09, 38.09) 50.48 (46.08, 54.64) THYB_S_YN05 137 55.44 (50.24, 60.32) 52.03 (48.59, 55.63) 63.75 (60.52, 66.75) 63.10 (58.34, 67.44) 14.28 (10.73, 18.14) THYB_S_EN02 12055.75 (50.96, 60.59) 39.06 (36.22, 42.01) 51.87 (48.88, 54.83) 38.76 (33.77, 43.55) 40.34 (35.64, 44.90) THYB_S_AH04 10165.76 (59.94, 70.88) 52.92 (49.38, 56.65) 61.94 (58.10, 65.51) 56.38 (50.58, 61.47) 47.68 (41.39, 54.39) THYB_S_GS03 94 50.96 (45.22, 56.84) 59.17 (54.56, 63.60) 51.74 (47.42, 55.76) 48.24 (42.53, 53.76) 40.09 (34.46, 45.45) THYB_S_FJ03 8566.63 (60.76, 72.05) 46.90 (43.13, 50.64) 54.13 (50.57, 57.52) 50.83 (44.94, 56.45) 36.40 (30.87, 41.92) THYB_S_SD12 7569.59 (64.15, 74.62) 64.10 (60.34, 67.84) 59.41 (55.70, 63.10) 58.28 (52.76, 63.56) 33.57 (27.12, 39.83) THYB_S_JS01 7256.33 (49.82, 62.97) 49.38 (45.18, 53.56) 54.34 (50.59, 58.03) 49.09 (42.61, 55.51) 28.27 (23.18, 33.88) THYB_S_NM02 6657.71 (52.20, 63.08) 52.72 (47.87, 57.77) 52.81 (47.81, 57.29) 46.24 (40.41, 52.18) 36.74 (30.84, 42.37) THYB_S_JL04 54 43.92 (34.91, 52.81) 56.22 (50.23, 62.17) 55.83 (50.74, 61.29) 40.85 (32.13, 49.16) 42.49 (34.52, 50.97) THYB_S_SH06 43 49.75 (41.79, 57.64) 51.33 (45.78, 56.54) 45.09 (37.42, 51.92) 46.30 (37.76, 54.02) 22.32 (15.88, 29.32) THYB_S_BJ09 41 54.80 (46.03, 63.29) 55.54 (49.59, 61.66) 55.00 (50.40, 59.85) 39.65 (31.89, 47.97) 52.67 (44.89, 60.73) THYB_S_SX04 4161.37 (55.03, 67.56) 56.50 (50.56, 62.42) 48.79 (42.56, 54.12) 44.04 (34.60, 52.15) 42.36 (33.90, 51.28) THYB_S_SC06 29 38.63 (29.91, 48.19) 69.92 (62.86, 76.12) 58.20 (49.69, 65.60) 32.00 (25.56, 38.77) 59.45 (50.88, 67.48) THYB_S_XJ01 2568.57 (58.62, 77.48) 55.07 (48.48, 61.59) 57.03 (50.98, 62.37) 60.51 (51.17, 68.75) 14.00 (7.83, 21.61) THYB_S_YN01 25 55.83 (43.67, 66.82) 58.55 (51.86, 65.06) 56.48 (48.32, 64.21) 46.59 (35.73, 57.93) 29.32 (19.14, 40.82) THYB_S_SD13 22 69.21 (57.84, 80.57) 72.83 (63.53, 80.86) 63.22 (55.37, 71.01) 41.87 (31.62, 52.02) 46.04 (34.90, 56.40) THYB_S_GX01 1964.80 (52.35, 75.87) 53.73 (46.83, 61.09) 60.04 (51.68, 67.03) 61.17 (48.99, 71.80) 24.77 (12.11, 38.80) THYB_S_ZJ29 1685.21 (82.65, 87.63) 52.40 (43.77, 61.90) 67.08 (59.58, 74.31) 47.06 (32.00, 61.91) 3.30 (0.48, 7.60) THYB_S_HB07 831.68 (9.84, 58.05) 69.62 (56.11, 82.55) 52.10 (37.92, 62.81) 19.55 (0.08, 43.94) 29.09 (16.83, 41.56) THYB_S_FJ01 267.14 (64.04, 70.23) 44.00 (38.25, 49.75) 49.95 (43.08, 56.82) 45.18 (32.48, 57.89) 39.20 (29.38, 49.01) THYB_S_SD14 23.13 (2.55, 3.71) 46.73 (36.29, 57.18) 58.38 (54.22, 62.54) 44.78 (31.59, 57.97) 37.57 (31.23, 43.91) 35/51 Table S8.Per-centre 95% Hausdorff distance (HD95, m) for thyroid gland segmentation on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFM HD95 (m)↓HD95 (m)↓HD95 (m)↓HD95 (m)↓HD95 (m)↓ Overall8,721 35.83 (35.14, 36.49) 65.71 (65.07, 66.38) 45.87 (45.36, 46.37) 40.65 (39.95, 41.35) 58.17 (57.35, 58.99) THYB_S_ZJ24 1,17414.74 (13.90, 15.63) 75.31 (73.88, 76.74) 64.28 (62.83, 65.85) 22.39 (21.43, 23.43) 53.71 (50.40, 56.90) THYB_S_EN04 94930.01 (28.39, 31.72) 73.59 (71.54, 75.87) 42.31 (40.95, 43.81) 34.27 (32.43, 36.13) 63.80 (61.88, 65.55) THYB_S_BJ01 87636.67 (34.70, 38.76) 55.12 (53.38, 56.96) 41.33 (40.11, 42.83) 40.85 (38.94, 43.05) 60.27 (57.54, 63.41) THYB_S_SH01 832 50.89 (48.25, 53.54) 55.08 (52.59, 57.51) 39.63 (38.18, 41.11) 51.90 (49.43, 54.61) 51.25 (49.37, 53.41) THYB_S_ZJ05 69437.22 (34.87, 39.57) 70.81 (68.16, 73.22) 49.93 (48.19, 51.64) 40.86 (38.46, 43.25) 64.30 (61.86, 67.00) THYB_S_SH05 66936.52 (34.23, 39.01) 58.39 (56.18, 60.42) 43.68 (42.20, 45.23) 44.98 (42.65, 47.48) 67.24 (64.77, 69.87) THYB_S_NX01 582 39.66 (37.22, 42.54) 82.07 (79.47, 84.56) 51.93 (49.68, 54.07) 39.06 (36.74, 41.62) 57.51 (54.52, 60.33) THYB_S_ZJ06 46043.86 (40.98, 47.06) 53.93 (51.41, 56.72) 44.67 (42.75, 46.81) 48.71 (45.68, 51.91) 62.98 (60.17, 65.91) THYB_S_QX07 28930.43 (27.06, 33.67) 84.32 (81.13, 87.54) 37.11 (35.21, 39.31) 34.16 (30.75, 37.53) 58.07 (53.90, 61.53) THYB_S_AN01 243 59.89 (54.84, 65.27) 39.11 (35.71, 42.69) 36.76 (34.09, 39.70) 57.02 (52.09, 61.87) 42.73 (39.24, 46.33) THYB_S_CQ03 23836.80 (33.17, 40.64) 67.69 (64.19, 71.21) 39.40 (36.37, 42.33) 49.95 (45.43, 54.40) 63.03 (58.68, 67.50) THYB_S_GZ02 23337.00 (33.22, 40.65) 50.49 (47.63, 53.64) 39.41 (37.30, 41.67) 40.19 (35.91, 44.49) 57.61 (53.08, 62.19) THYB_S_JX06 215 42.43 (38.02, 47.27) 54.75 (50.60, 59.00) 37.59 (35.11, 40.19) 48.81 (44.42, 53.58) 52.03 (47.75, 56.31) THYB_S_JS02 19049.35 (44.19, 55.17) 52.35 (48.11, 56.54) 53.87 (50.49, 57.35) 56.88 (51.46, 62.82) 52.07 (47.01, 57.13) THYB_S_YN05 137 37.57 (32.03, 43.12) 73.49 (68.84, 77.86) 35.07 (31.77, 38.50) 39.92 (34.22, 46.49) 50.39 (41.86, 58.71) THYB_S_EN02 12037.62 (32.53, 42.62) 100.8 (95.9, 105.8) 45.67 (42.65, 48.58) 44.64 (39.27, 50.38) 54.04 (48.46, 59.63) THYB_S_AH04 10125.82 (20.55, 31.83) 70.36 (64.77, 75.45) 33.52 (29.78, 37.56) 35.53 (29.34, 42.79) 52.02 (44.23, 60.59) THYB_S_GS03 94 44.87 (37.84, 51.56) 61.22 (54.74, 67.47) 44.51 (40.21, 49.25) 48.16 (40.83, 55.61) 52.29 (45.50, 59.13) THYB_S_FJ03 8527.04 (21.49, 33.08) 83.26 (77.91, 88.32) 48.29 (43.55, 53.23) 40.21 (34.71, 46.20) 57.81 (51.88, 63.89) THYB_S_SD12 7523.80 (19.40, 28.64) 57.62 (52.08, 62.79) 36.03 (32.04, 39.97) 35.69 (30.37, 41.96) 51.62 (43.29, 59.76) THYB_S_JS01 7240.00 (32.29, 47.62) 79.39 (72.78, 85.69) 45.34 (40.63, 50.53) 40.73 (33.99, 47.88) 64.43 (55.62, 73.62) THYB_S_NM02 66 45.94 (38.55, 53.94) 63.04 (56.60, 69.43) 44.14 (39.41, 49.57) 51.02 (42.31, 59.30) 62.80 (54.21, 70.52) THYB_S_JL04 54 53.99 (42.73, 65.60) 67.66 (59.00, 76.78) 39.38 (34.57, 44.49) 50.55 (39.81, 61.07) 57.53 (46.74, 67.69) THYB_S_SH06 43 45.77 (34.95, 57.09) 63.35 (55.55, 71.99) 45.14 (38.78, 51.03) 41.91 (30.31, 53.31) 63.26 (50.44, 76.60) THYB_S_BJ09 4136.63 (27.94, 47.32) 63.20 (55.22, 70.78) 42.93 (36.63, 49.75) 41.72 (31.74, 52.87) 44.38 (34.57, 53.98) THYB_S_SX04 4138.80 (31.32, 47.14) 66.23 (57.04, 75.13) 48.62 (44.02, 53.56) 53.31 (42.58, 66.15) 68.42 (55.93, 82.59) THYB_S_SC06 29 51.99 (36.85, 66.32) 45.67 (35.65, 56.18) 38.52 (30.69, 46.92) 56.62 (43.06, 70.35) 43.15 (33.11, 53.56) THYB_S_XJ01 2529.06 (21.16, 38.25) 54.51 (46.58, 63.05) 44.52 (37.28, 52.45) 44.96 (34.41, 56.69) 52.44 (36.21, 68.45) THYB_S_YN01 2535.59 (22.18, 54.31) 62.96 (52.87, 73.71) 38.48 (32.35, 44.59) 54.30 (37.89, 71.41) 66.12 (49.33, 83.02) THYB_S_SD13 2223.75 (12.47, 36.80) 43.51 (30.68, 58.21) 36.56 (28.65, 44.41) 52.16 (38.57, 65.81) 66.39 (52.20, 84.55) THYB_S_GX01 19 33.19 (22.74, 43.33) 78.95 (67.13, 88.94) 29.49 (22.85, 36.30) 38.98 (26.03, 54.00) 38.20 (20.84, 60.52) THYB_S_ZJ29 1612.12 (6.47, 20.65) 65.31 (54.30, 76.11) 22.34 (17.11, 28.07) 38.17 (22.88, 54.66) 45.26 (21.23, 69.87) THYB_S_HB07 830.47 (7.26, 58.20) 44.74 (23.64, 68.97) 32.40 (23.17, 45.39) 39.33 (7.94, 74.12) 80.20 (66.41, 92.63) THYB_S_FJ01 258.89 (23.60, 94.19) 75.44 (66.29, 84.60) 33.92 (27.53, 40.31) 47.04 (34.00, 60.08) 49.86 (33.24, 66.48) THYB_S_SD14 278.47 (78.16, 78.77) 98.68 (81.39, 115.97) 55.96 (43.32, 68.59) 34.84 (19.69, 50.00) 40.44 (36.06, 44.82) 36/51 Table S9.Per-centre Dice similarity coefficient (%) for thyroid nodule segmentation on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFM Dice (%)↑Dice (%)↑Dice (%)↑Dice (%)↑Dice (%)↑ Overall6,384 82.31 (81.78, 82.83) 23.20 (22.58, 23.86) 28.82 (28.11, 29.53) 72.99 (72.29, 73.68) 58.87 (58.04, 59.68) THYB_S_EN04 94682.92 (81.77, 84.00) 15.54 (14.29, 16.78) 22.25 (20.68, 23.66) 74.43 (72.85, 75.95) 56.67 (54.69, 58.62) THYB_S_SH01 83488.78 (88.01, 89.61) 39.13 (37.10, 41.01) 46.16 (44.13, 48.29) 84.48 (83.49, 85.58) 71.64 (69.99, 73.33) THYB_S_ZJ05 69082.45 (81.21, 83.64) 15.47 (13.92, 17.01) 16.32 (14.69, 17.96) 67.34 (65.14, 69.41) 55.66 (52.98, 58.07) THYB_S_SH05 66585.53 (84.17, 86.86) 20.11 (18.43, 21.73) 29.48 (27.34, 31.47) 75.56 (73.63, 77.41) 60.40 (57.93, 62.74) THYB_S_NX01 55777.41 (75.39, 79.40) 16.92 (15.32, 18.56) 23.51 (21.59, 25.42) 68.50 (66.26, 70.91) 47.56 (44.76, 50.35) THYB_S_ZJ06 46084.52 (82.82, 86.14) 25.61 (23.25, 27.95) 28.94 (26.49, 31.44) 75.20 (73.03, 77.38) 61.29 (58.63, 63.89) THYB_S_ZJ24 32145.53 (41.66, 48.73) 5.47 (4.61, 6.45) 4.17 (3.42, 5.08) 20.51 (17.52, 23.78) 14.92 (12.15, 17.89) THYB_S_QX07 28085.51 (83.52, 87.35) 22.27 (19.68, 24.88) 28.59 (25.55, 31.77) 77.89 (74.95, 80.54) 68.19 (65.16, 71.49) THYB_S_AN01 24391.60 (89.81, 92.94) 53.29 (49.81, 56.91) 61.68 (58.00, 65.00) 88.34 (86.28, 90.14) 78.19 (75.20, 81.15) THYB_S_CQ03 18682.72 (79.86, 85.39) 18.16 (15.15, 21.46) 27.06 (23.17, 31.10) 72.57 (68.81, 76.37) 55.25 (50.23, 60.01) THYB_S_JS02 18487.67 (85.58, 89.59) 34.39 (30.61, 38.22) 35.40 (30.90, 39.94) 82.96 (80.24, 85.39) 71.70 (67.73, 75.66) THYB_S_JX06 16787.50 (84.81, 89.82) 35.52 (31.67, 39.88) 45.33 (40.68, 50.25) 83.78 (80.89, 86.39) 73.76 (70.40, 77.15) THYB_S_BJ01 14269.73 (64.15, 74.65) 16.56 (13.36, 19.94) 21.83 (17.92, 25.72) 58.88 (53.18, 64.45) 46.92 (41.12, 52.74) THYB_S_EN02 10784.38 (81.60, 86.98) 11.99 (9.13, 15.33) 19.36 (15.55, 23.76) 73.29 (69.08, 77.46) 50.71 (44.52, 57.00) THYB_S_GZ02 9289.21 (87.74, 90.46) 27.23 (22.31, 32.19) 30.27 (24.62, 36.27) 84.00 (81.49, 86.01) 73.15 (68.35, 77.38) THYB_S_GS03 9185.66 (81.56, 88.93) 28.49 (22.48, 34.72) 35.22 (28.72, 41.50) 78.93 (74.06, 83.12) 64.75 (58.31, 70.89) THYB_S_JS01 6685.58 (81.87, 88.54) 13.65 (10.66, 16.77) 18.99 (15.28, 23.17) 76.18 (70.10, 81.21) 58.55 (50.67, 65.99) THYB_S_FJ03 6479.73 (74.98, 83.85) 10.49 (7.58, 13.92) 14.62 (10.01, 20.30) 64.17 (55.49, 71.69) 50.24 (41.88, 59.00) THYB_S_NM02 5083.34 (75.81, 89.18) 21.78 (15.16, 29.35) 27.11 (20.14, 34.56) 76.31 (67.18, 83.75) 61.40 (52.34, 69.28) THYB_S_SH06 3170.85 (59.24, 79.88) 4.60 (3.47, 5.86) 8.28 (5.59, 11.59) 55.16 (39.98, 68.39) 46.54 (33.59, 59.77) THYB_S_AH04 3091.71 (89.82, 93.31) 39.78 (30.88, 48.49) 48.60 (37.90, 58.61) 84.52 (77.22, 89.78) 79.09 (70.70, 84.79) THYB_S_SC06 2990.16 (88.07, 91.95) 44.20 (32.95, 53.60) 49.38 (38.45, 58.61) 88.29 (85.51, 90.65) 76.57 (67.27, 84.35) THYB_S_SX04 2789.20 (86.29, 91.70) 17.84 (12.89, 22.84) 16.51 (11.20, 23.50) 87.47 (84.52, 89.90) 68.19 (57.72, 77.31) THYB_S_BJ09 2286.20 (81.55, 90.25) 30.06 (19.75, 40.78) 34.98 (22.71, 47.67) 81.62 (74.86, 87.53) 70.13 (63.36, 77.01) THYB_S_YN01 2287.24 (79.41, 92.13) 23.20 (13.43, 34.02) 28.13 (18.31, 39.88) 80.58 (70.28, 87.54) 44.07 (29.99, 57.81) THYB_S_SD13 1892.74 (91.15, 94.06) 52.54 (39.34, 64.38) 61.58 (48.47, 73.53) 90.83 (88.24, 92.78) 77.31 (70.93, 82.80) THYB_S_JL04 1681.27 (67.59, 93.13) 36.60 (22.73, 50.94) 48.61 (31.53, 65.36) 76.87 (62.21, 89.36) 70.49 (53.00, 86.26) THYB_S_SD12 1683.71 (69.71, 93.91) 53.71 (39.32, 67.16) 62.86 (47.03, 78.03) 79.72 (62.56, 93.40) 74.86 (58.02, 88.45) THYB_S_YN05 1291.81 (88.17, 94.81) 59.02 (42.58, 71.76) 70.91 (51.90, 85.39) 91.01 (87.57, 93.88) 68.71 (47.83, 85.32) THYB_S_XJ01 882.71 (73.69, 90.34) 13.18 (4.03, 24.95) 14.95 (4.71, 29.10) 56.32 (33.11, 79.31) 45.53 (20.28, 71.40) THYB_S_GX01 485.84 (83.15, 88.54) 20.06 (9.70, 30.41) 38.09 (22.52, 49.68) 84.85 (78.42, 89.87) 70.97 (60.58, 83.12) THYB_S_FJ01 279.74 (78.18, 81.31) 6.47 (4.46, 8.48) 12.32 (8.26, 16.37) 81.94 (77.73, 86.14) 46.04 (36.60, 55.49) THYB_S_HB07 221.28 (18.58, 23.99) 1.90 (1.48, 2.31) 1.86 (1.61, 2.11) 23.49 (0.00, 46.97) 17.79 (11.76, 23.82) 37/51 Table S10.Per-centre 95% Hausdorff distance (HD95, m) for thyroid nodule segmentation on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgentMedSAM2MedSegXTransUNetUltraFedFM HD95 (m)↓HD95 (m)↓HD95 (m)↓HD95 (m)↓HD95 (m)↓ Overall6,384 9.41 (8.89, 9.93) 101.2 (100.2, 102.0) 86.96 (86.08, 87.84) 10.62 (10.14, 11.12) 27.04 (26.28, 27.75) THYB_S_EN04 9466.95 (6.05, 7.91) 113.0 (111.0, 114.9) 91.23 (89.41, 92.99) 9.27 (8.31, 10.37) 33.50 (31.47, 35.76) THYB_S_SH01 8346.41 (5.50, 7.36) 79.94 (77.11, 82.81) 65.42 (63.21, 67.68) 8.32 (7.35, 9.27) 21.31 (19.64, 22.92) THYB_S_ZJ05 6905.00 (4.31, 5.76) 108.5 (106.1, 110.6) 106.8 (104.6, 109.0) 7.73 (6.88, 8.61) 21.22 (19.59, 23.13) THYB_S_SH05 6657.60 (6.34, 8.89) 101.8 (99.4, 104.1) 81.02 (78.74, 83.33) 10.17 (8.84, 11.51) 28.23 (25.93, 30.37) THYB_S_NX01 55712.95 (10.96, 14.86) 112.7 (110.1, 115.1) 92.51 (89.99, 94.97) 13.02 (11.18, 14.80) 30.17 (27.56, 32.82) THYB_S_ZJ06 4607.62 (6.22, 9.24) 92.99 (90.07, 95.87) 88.02 (85.09, 91.28) 10.88 (9.34, 12.66) 26.09 (23.54, 28.75) THYB_S_ZJ24 321 45.03 (39.74, 50.67) 142.9 (139.7, 145.9) 140.9 (137.7, 143.8) 30.45 (26.09, 34.77) 45.71 (40.11, 51.74) THYB_S_QX07 2806.69 (5.08, 8.55) 110.2 (106.4, 113.9) 85.79 (82.50, 89.13) 8.11 (6.54, 10.16) 22.53 (19.71, 25.68) THYB_S_AN01 2434.70 (3.51, 6.11) 60.30 (55.67, 64.85) 49.32 (45.45, 53.68) 7.19 (5.65, 9.00) 16.62 (14.15, 19.21) THYB_S_CQ03 1868.81 (5.97, 12.38) 105.0 (100.0, 109.4) 81.35 (77.26, 85.23) 12.51 (9.28, 16.26) 33.57 (28.12, 38.94) THYB_S_JS02 1846.68 (4.75, 8.90) 80.55 (75.49, 85.50) 84.81 (78.76, 90.71) 8.78 (6.74, 10.90) 24.40 (20.23, 28.77) THYB_S_JX06 1677.25 (5.01, 9.88) 82.29 (76.67, 87.66) 63.70 (58.48, 68.68) 9.18 (6.44, 12.50) 20.12 (16.76, 23.56) THYB_S_BJ01 142 20.72 (15.43, 26.10) 101.8 (96.3, 107.7) 92.30 (87.28, 97.40) 20.10 (15.47, 25.01) 34.87 (29.55, 40.13) THYB_S_EN02 1074.49 (3.07, 6.59) 128.9 (123.9, 133.4) 92.21 (87.64, 96.59) 8.45 (6.20, 11.49) 29.38 (24.47, 34.60) THYB_S_GZ02 924.51 (3.43, 5.77) 84.00 (78.53, 89.58) 81.50 (74.63, 88.26) 7.53 (5.65, 9.73) 17.57 (13.41, 22.16) THYB_S_GS03 917.80 (4.49, 11.81) 89.69 (81.70, 97.29) 75.21 (68.30, 82.38) 8.70 (5.54, 12.57) 24.98 (19.70, 31.13) THYB_S_JS01 665.69 (3.49, 8.71) 111.8 (106.2, 117.2) 94.67 (89.47, 99.84) 9.84 (6.27, 14.06) 27.35 (19.68, 35.58) THYB_S_FJ03 646.99 (3.86, 10.72) 115.9 (109.6, 122.4) 107.7 (100.0, 115.2) 8.87 (5.71, 12.92) 22.36 (16.53, 28.70) THYB_S_NM02 506.75 (3.95, 10.41) 95.61 (85.83, 104.48) 82.78 (75.37, 90.43) 6.49 (4.04, 9.79) 21.73 (16.43, 27.52) THYB_S_SH06 3116.01 (6.79, 25.90) 115.0 (110.3, 120.7) 94.85 (88.68, 100.91) 7.08 (3.20, 12.90) 28.90 (17.04, 42.21) THYB_S_AH04 303.09 (2.14, 4.32) 81.37 (70.67, 92.67) 66.34 (54.18, 78.80) 9.13 (4.11, 16.74) 18.64 (11.53, 27.65) THYB_S_SC06 294.72 (2.94, 6.80) 71.23 (58.43, 85.23) 61.37 (51.00, 72.75) 5.60 (3.63, 7.90) 20.14 (10.14, 31.55) THYB_S_SX04 273.54 (2.39, 4.80) 106.5 (98.0, 114.6) 108.8 (98.9, 118.0) 4.05 (3.12, 5.09) 25.37 (14.31, 39.60) THYB_S_BJ09 225.56 (3.29, 8.33) 85.51 (71.97, 98.12) 77.91 (63.39, 92.08) 7.06 (3.88, 11.19) 18.84 (11.45, 26.82) THYB_S_YN01 224.61 (2.12, 8.49) 96.72 (83.10, 109.97) 73.60 (64.25, 81.65) 4.98 (3.42, 6.90) 40.06 (27.32, 54.75) THYB_S_SD13 183.32 (2.22, 4.57) 61.73 (47.66, 76.68) 45.00 (32.04, 58.04) 4.51 (2.74, 6.60) 21.77 (15.46, 28.42) THYB_S_JL04 1614.24 (2.17, 31.43) 88.05 (70.97, 104.04) 58.06 (41.17, 75.76) 12.51 (3.43, 23.13) 16.36 (6.05, 28.72) THYB_S_SD12 169.76 (3.28, 17.71) 57.71 (38.35, 78.90) 50.35 (31.21, 71.00) 8.08 (2.22, 15.94) 18.16 (7.70, 30.26) THYB_S_YN05 123.22 (2.05, 4.34) 61.17 (44.75, 80.17) 38.97 (24.66, 57.04) 4.86 (2.25, 9.08) 13.71 (6.45, 21.71) THYB_S_XJ01 86.99 (1.71, 16.39) 111.3 (92.8, 131.6) 99.22 (85.22, 115.90) 20.73 (4.35, 42.34) 24.55 (5.64, 47.07) THYB_S_GX01 45.57 (3.00, 8.47) 101.0 (81.2, 123.7) 78.07 (56.31, 92.33) 5.12 (2.62, 8.46) 15.24 (6.58, 25.12) THYB_S_FJ01 28.32 (4.00, 12.65) 112.3 (104.0, 120.7) 86.49 (73.93, 99.04) 4.41 (2.83, 6.00) 56.71 (26.29, 87.13) THYB_S_HB07 230.80 (18.03, 43.57) 109.0 (98.6, 119.4) 104.4 (95.2, 113.5) 4.70 (0.00, 9.39) 19.30 (8.60, 30.00) 38/51 Table S11.Per-centre AUROC for benign versus malignant thyroid nodule classification on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgent BiomedCLIPMedSigLIPUltraFedFM AUROC↑AUROC↑AUROC↑AUROC↑ Overall4,999 0.819 (0.808, 0.831) 0.434 (0.418, 0.450) 0.520 (0.503, 0.535) 0.436 (0.420, 0.452) THYB_S_SH01 7920.896 (0.873, 0.916) 0.327 (0.290, 0.368) 0.552 (0.517, 0.594) 0.331 (0.295, 0.368) THYB_S_ZJ05 6470.811 (0.773, 0.844) 0.526 (0.480, 0.576) 0.586 (0.546, 0.627) 0.494 (0.455, 0.533) THYB_S_EN04 6260.846 (0.808, 0.876) 0.345 (0.289, 0.404) 0.538 (0.472, 0.597) 0.528 (0.476, 0.584) THYB_S_SH05 5390.719 (0.672, 0.758) 0.331 (0.280, 0.381) 0.543 (0.492, 0.595) 0.370 (0.324, 0.423) THYB_S_ZJ06 4520.811 (0.768, 0.849) 0.468 (0.412, 0.520) 0.594 (0.541, 0.651) 0.539 (0.486, 0.592) THYB_S_QX07 2790.870 (0.814, 0.917) 0.457 (0.381, 0.537) 0.630 (0.534, 0.711) 0.428 (0.338, 0.520) THYB_S_ZJ24 269 0.590 (0.346, 0.832) 0.412 (0.270, 0.544) 0.673 (0.493, 0.823) 0.463 (0.255, 0.678) THYB_S_AN01 2360.862 (0.784, 0.923) 0.241 (0.165, 0.319) 0.601 (0.519, 0.687) 0.292 (0.202, 0.390) THYB_S_CQ03 1720.784 (0.716, 0.849) 0.452 (0.366, 0.550) 0.521 (0.431, 0.608) 0.440 (0.350, 0.532) THYB_S_JS02 1710.746 (0.667, 0.815) 0.504 (0.417, 0.596) 0.675 (0.599, 0.750) 0.415 (0.328, 0.501) THYB_S_JX06 1530.903 (0.853, 0.951) 0.399 (0.315, 0.487) 0.616 (0.524, 0.702) 0.480 (0.390, 0.575) THYB_S_EN02 990.851 (0.741, 0.941) 0.383 (0.248, 0.543) 0.749 (0.602, 0.878) 0.329 (0.166, 0.488) THYB_S_BJ01 930.749 (0.635, 0.851) 0.245 (0.154, 0.345) 0.400 (0.235, 0.562) 0.460 (0.319, 0.609) THYB_S_GZ02 900.693 (0.578, 0.795) 0.325 (0.198, 0.471) 0.499 (0.374, 0.614) 0.366 (0.223, 0.530) THYB_S_GS03 840.838 (0.735, 0.924) 0.244 (0.125, 0.373) 0.443 (0.311, 0.566) 0.470 (0.328, 0.607) THYB_S_JS01 590.781 (0.491, 1.000) 0.509 (0.069, 0.947) 0.281 (0.035, 0.552) 0.070 (0.017, 0.500) THYB_S_NM02 49 0.653 (0.497, 0.800) 0.408 (0.253, 0.567) 0.719 (0.549, 0.875) 0.393 (0.202, 0.595) THYB_S_SX04 270.860 (0.500, 1.000) 0.640 (0.231, 0.962) 0.500 (0.077, 0.924) 0.620 (0.400, 0.846) THYB_S_AH04 240.844 (0.663, 0.979) 0.617 (0.368, 0.838) 0.586 (0.305, 0.838) 0.266 (0.076, 0.523) THYB_S_SC06 220.876 (0.702, 1.000) 0.256 (0.050, 0.484) 0.612 (0.350, 0.839) 0.141 (0.009, 0.325) THYB_S_YN01 210.853 (0.618, 1.000) 0.353 (0.000, 0.778) 0.500 (0.105, 0.895) 0.176 (0.000, 0.400) THYB_S_BJ09 20 0.586 (0.319, 0.849) 0.505 (0.213, 0.798) 0.636 (0.341, 0.885) 0.404 (0.150, 0.687) THYB_S_SD13 180.875 (0.636, 1.000) 0.536 (0.125, 1.000) 0.714 (0.415, 0.956) 0.304 (0.000, 0.623) THYB_S_SD12 150.929 (0.500, 1.000) 0.143 (0.000, 0.500) 0.571 (0.356, 0.786) 0.071 (0.000, 0.500) THYB_S_JL04 120.407 (0.000, 1.000) 0.074 (0.000, 0.446) 0.148 (0.000, 0.500) 0.296 (0.000, 0.727) THYB_S_YN05 120.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) THYB_S_XJ01 51.000 (0.500, 1.000) 0.667 (0.000, 1.000) 0.333 (0.000, 1.000) 0.833 (0.250, 1.000) THYB_S_GX01 40.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) THYB_S_SH06 41.000 (0.500, 1.000) 0.000 (0.000, 0.500) 0.000 (0.000, 0.500) 0.500 (0.000, 1.000) THYB_S_FJ01 20.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) THYB_S_FJ03 20.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) THYB_S_NX01 10.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) 0.500 (0.500, 0.500) THYB_S_HB07 0– THYB_S_SD14 0– THYB_S_ZJ29 0– 39/51 Table S12.Per-centre AUPRC for benign versus malignant thyroid nodule classification on the NHC-MISD-TUS external test set. Values are reported as point estimates with 95% confidence intervals. Bold indicates the best result in each row. CenterN ThyroidXAgent BiomedCLIPMedSigLIPUltraFedFM AUPRC↑AUPRC↑AUPRC↑AUPRC↑ Overall4,999 0.823 (0.807, 0.838) 0.450 (0.433, 0.467) 0.519 (0.499, 0.539) 0.465 (0.447, 0.483) THYB_S_SH01 7920.859 (0.817, 0.893) 0.323 (0.290, 0.356) 0.450 (0.404, 0.504) 0.340 (0.304, 0.380) THYB_S_ZJ05 6470.746 (0.690, 0.798) 0.460 (0.412, 0.518) 0.519 (0.469, 0.580) 0.445 (0.393, 0.504) THYB_S_EN04 6260.969 (0.957, 0.978) 0.796 (0.758, 0.840) 0.851 (0.812, 0.888) 0.871 (0.837, 0.904) THYB_S_SH05 5390.832 (0.787, 0.865) 0.539 (0.496, 0.589) 0.680 (0.630, 0.734) 0.555 (0.509, 0.610) THYB_S_ZJ06 4520.811 (0.752, 0.864) 0.502 (0.442, 0.573) 0.597 (0.526, 0.670) 0.554 (0.490, 0.626) THYB_S_QX07 2790.575 (0.432, 0.706) 0.163 (0.107, 0.235) 0.270 (0.173, 0.385) 0.158 (0.107, 0.245) THYB_S_ZJ24 2690.164 (0.026, 0.423) 0.032 (0.014, 0.054) 0.068 (0.027, 0.137) 0.051 (0.017, 0.137) THYB_S_AN01 2360.588 (0.430, 0.756) 0.110 (0.082, 0.147) 0.195 (0.141, 0.261) 0.121 (0.087, 0.172) THYB_S_CQ03 1720.876 (0.819, 0.922) 0.603 (0.510, 0.706) 0.637 (0.544, 0.735) 0.559 (0.474, 0.651) THYB_S_JS02 1710.734 (0.629, 0.822) 0.502 (0.414, 0.621) 0.649 (0.544, 0.750) 0.459 (0.369, 0.561) THYB_S_JX06 1530.897 (0.839, 0.946) 0.351 (0.276, 0.438) 0.545 (0.422, 0.671) 0.409 (0.314, 0.522) THYB_S_EN02 990.969 (0.937, 0.992) 0.812 (0.710, 0.908) 0.939 (0.888, 0.983) 0.775 (0.669, 0.899) THYB_S_BJ01 930.930 (0.880, 0.970) 0.692 (0.583, 0.825) 0.743 (0.634, 0.858) 0.764 (0.661, 0.884) THYB_S_GZ02 900.850 (0.760, 0.921) 0.558 (0.460, 0.695) 0.702 (0.581, 0.825) 0.569 (0.457, 0.697) THYB_S_GS03 840.921 (0.847, 0.970) 0.578 (0.459, 0.706) 0.692 (0.559, 0.816) 0.674 (0.553, 0.808) THYB_S_JS01 590.990 (0.966, 1.000) 0.964 (0.885, 1.000) 0.950 (0.862, 1.000) 0.921 (0.808, 1.000) THYB_S_NM02 490.810 (0.662, 0.920) 0.594 (0.441, 0.781) 0.777 (0.611, 0.941) 0.562 (0.420, 0.729) THYB_S_SX04 270.988 (0.959, 1.000) 0.960 (0.873, 1.000) 0.931 (0.792, 1.000) 0.965 (0.895, 1.000) THYB_S_AH04 240.746 (0.430, 0.967) 0.463 (0.222, 0.800) 0.547 (0.203, 0.819) 0.258 (0.130, 0.450) THYB_S_SC06 220.873 (0.653, 1.000) 0.390 (0.229, 0.609) 0.704 (0.409, 0.898) 0.356 (0.209, 0.573) THYB_S_YN01 210.968 (0.901, 1.000) 0.748 (0.537, 0.990) 0.832 (0.615, 0.993) 0.737 (0.498, 0.947) THYB_S_BJ09 20 0.683 (0.415, 0.910) 0.687 (0.406, 0.895) 0.699 (0.431, 0.925) 0.559 (0.292, 0.806) THYB_S_SD13 180.567 (0.200, 1.000) 0.446 (0.059, 1.000) 0.415 (0.111, 0.900) 0.196 (0.056, 0.415) THYB_S_SD12 150.500 (0.000, 1.000) 0.077 (0.000, 0.177) 0.143 (0.000, 0.365) 0.071 (0.000, 0.150) THYB_S_JL04 120.491 (0.083, 1.000) 0.187 (0.083, 0.379) 0.199 (0.083, 0.404) 0.241 (0.083, 0.610) THYB_S_YN05 120.000 (0.000, 0.000) 0.000 (0.000, 0.000) 0.000 (0.000, 0.000) 0.000 (0.000, 0.000) THYB_S_XJ01 51.000 (1.000, 1.000) 0.806 (0.333, 1.000) 0.589 (0.200, 1.000) 0.917 (0.417, 1.000) THYB_S_GX01 41.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) THYB_S_SH06 41.000 (0.000, 1.000) 0.417 (0.000, 1.000) 0.417 (0.000, 1.000) 0.583 (0.000, 1.000) THYB_S_FJ01 21.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) THYB_S_FJ03 21.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) THYB_S_NX01 11.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) 1.000 (1.000, 1.000) THYB_S_HB07 0– THYB_S_SD14 0– THYB_S_ZJ29 0– 40/51 Table S13.Performance of preprocessing and executor tools in ThyroidXAgent. Held-out test results are reported for preprocessing tools, including image normalization, nodule-presence triage and anatomical-context parsing, and for executor-stage tools, including measurement support, gland localization, lymph-node screening, gland captioning and nodule-feature extraction. Anatomical-context parsing and nodule-feature extraction are additionally reported at the class level. For the binary margin and shape classifiers, AUROC and AUPRC are reported once across the paired class rows. AP, average precision; MAE, mean absolute error; MSE, mean squared error; MAPE, mean absolute percentage error. Agent stagePreprocessing tools Tool groupToolTrainValTestPrimary resultSecondary result Preprocessing Image normalization Ultrasound ROI cropping1521779Dice, 0.9822; IoU, 0.9658Precision, 0.9904; recall, 0.9749; pixel accuracy, 0.9829 Case triageNodule-presence detection82,31210,98216,467Accuracy, 0.9830; F1, 0.9749AUROC, 0.9981; AP, 0.9961; sensitivity, 0.9848; specificity, 0.9821 Anatomical context parsing ToolClassTrainValTestTotalPrecisionRecallF1AUROCAUPRC Thyroid-region classification Left-lobe lateral view 7801101591,0490.66430.59750.62910.85210.7185 Right-lobe lateral view 9461331991,2780.64600.73370.68710.83270.7368 Bilateral thyroid view 12019301690.89290.83330.86210.98960.8911 Left-lobe transverse view 26752413600.70450.75610.72940.95670.8198 Right-lobe transverse view 27261794120.80880.69620.74830.95530.8585 Neck region26721123001.00000.91670.95650.99740.9524 Agent stageExecutor-stage tools Tool groupToolTrainValTestPrimary resultSecondary result Executor Measurement support Spacing prediction5,288661662MAE, 0.0131;R 2 , 0.8520MSE,5.66×10 −4 ; MAPE, 21.39% Gland localization Gland segmentation335-90Dice, 0.8006; IoU, 0.6866Precision, 0.8025; recall, 0.8339 Neck-region screening Cervical lymph-node detection251-49Accuracy, 0.7959; F1, 0.7368AUROC, 0.8163 Gland description Gland captioning22,782200612BLEU-4, 0.5898; METEOR, 0.4582ROUGE L , 0.7450; CIDEr, 2.7736 Nodule feature extraction Tool familyFeature classifierClassTrainValTestTotalSpecificitySensitivityAUROCAUPRC Nodule-feature classification Composition Cystic1,8242272282,2790.83970.85530.91660.9073 Mixed cystic and solid 8831081101,1010.93580.36360.82570.5943 Solid1,3911731771,7410.80180.79660.90110.8033 Echogenicity Anechoic1,6912122132,1160.86030.89670.93970.9160 Hyperechoic21827272720.98470.25930.84740.3591 Hypoechoic1,4161731731,7620.81730.67050.83910.7514 Isoechoic58073727250.92740.54170.88040.5440 Echogenic foci Macrocalcifications1,1651451461,4560.81070.47260.72790.5601 None2,1502702702,6900.51300.80370.74310.7483 Punctate echogenic foci 66883848350.94230.13100.64110.2615 Margin Ill-defined1,1771471481,4720.82000.7432 0.87020.8847 Smooth1,6022002002,0020.74320.8200 Shape Taller-than-wide811091000.98720.5556 0.91740.9899 Wider-than-tall60279787590.55560.9872 41/51 Table S14.Lexical performance of thyroid ultrasound report generation across datasets. BLEU-1 to BLEU-4, METEOR and ROUGE L are reported on the SMU-HMC, KMVE and ZJH-TS test sets. Values are means±the half-width of the bootstrap 95% percentile confidence interval. ModelBLEU-1BLEU-2BLEU-3BLEU-4METEORROUGE L SMU-HMC Testset(n=400) GPT-4o 95 0.3500±0.0100 0.2535±0.0080 0.1842±0.0066 0.1330±0.0057 0.3247±0.0047 0.3577±0.0077 GPT-5 68 0.3836±0.0085 0.2749±0.0069 0.1965±0.0059 0.1374±0.0054 0.3254±0.0037 0.3732±0.0068 Gemini2.5 Pro 69 0.3702±0.0094 0.2660±0.0081 0.1907±0.0070 0.1373±0.0063 0.3308±0.0038 0.3584±0.0071 Qwen3.5 Plus 96 0.4483±0.0130 0.3623±0.0111 0.2959±0.0095 0.2427±0.00810.3628±0.00400.5147±0.0088 Claude-Sonnet-4.6 0.3326±0.0099 0.2543±0.0080 0.1968±0.0067 0.1528±0.0056 0.3417±0.0035 0.3782±0.0074 MedGemma 71 0.0457±0.0072 0.0343±0.0056 0.0265±0.0044 0.0207±0.0035 0.1736±0.0051 0.0829±0.0075 LLaVA-Med 97 0.1670±0.0077 0.0581±0.0066 0.0290±0.0040 0.0159±0.0024 0.1439±0.0053 0.1482±0.0073 KMVE 18 0.1743±0.0131 0.1110±0.0082 0.0684±0.0052 0.0398±0.0035 0.1719±0.0063 0.2212±0.0039 ThyroidXAgent0.5924±0.01430.4806±0.01390.4006±0.01370.3381±0.01350.3627±0.00880.5422±0.0122 KMVE Testset(n=492) GPT-4o 95 0.4467±0.0154 0.3423±0.0143 0.2683±0.0118 0.2136±0.0102 0.2799±0.0116 0.3850±0.0148 GPT-5 68 0.4149±0.0073 0.2992±0.0066 0.2190±0.0059 0.1668±0.0062 0.3042±0.0070 0.4128±0.0089 Gemini2.5 Pro 69 0.5065±0.0117 0.3851±0.0103 0.2981±0.0095 0.2394±0.0092 0.2863±0.0083 0.4742±0.0111 Qwen3.5 Plus 96 0.5425±0.0135 0.4243±0.0126 0.3415±0.0121 0.2783±0.0120 0.2952±0.0084 0.5066±0.0118 Claude-Sonnet-4.6 0.4331±0.0119 0.3380±0.0112 0.2638±0.0107 0.2089±0.0106 0.3284±0.0079 0.4639±0.0115 MedGemma 71 0.0374±0.0195 0.0299±0.0175 0.0252±0.0161 0.0219±0.0151 0.1075±0.0119 0.1862±0.0205 LLaVA-Med 97 0.2842±0.0149 0.2135±0.0124 0.1595±0.0099 0.1244±0.0088 0.2138±0.0101 0.3939±0.0127 ThyroidXAgent0.6357±0.01710.5606±0.01570.5008±0.01510.4535±0.01510.3672±0.00990.5880±0.0120 ZJH-TS Testset(n=150) GPT-4o 95 0.3402±0.0273 0.2447±0.0206 0.1788±0.0155 0.1332±0.0119 0.2495±0.0163 0.3679±0.0219 GPT-5 68 0.4736±0.0158 0.3499±0.0130 0.2539±0.0108 0.1789±0.0095 0.3332±0.0060 0.4790±0.0098 Gemini2.5 Pro 69 0.1932±0.0112 0.1307±0.0080 0.0887±0.0058 0.0601±0.0045 0.2765±0.0052 0.2536±0.0091 Qwen3.5 Plus 96 0.4964±0.0182 0.3965±0.0162 0.3196±0.0145 0.2581±0.01310.3508±0.00760.5314±0.0125 Claude-Sonnet-4.6 0.4480±0.0181 0.3518±0.0157 0.2792±0.0135 0.2229±0.0118 0.3455±0.0073 0.4864±0.0127 MedGemma 71 0.1236±0.0220 0.0958±0.0174 0.0750±0.0138 0.0591±0.0110 0.2268±0.0094 0.1592±0.0211 LLaVA-Med 97 0.2318±0.0194 0.1316±0.0141 0.0891±0.0101 0.0606±0.0074 0.1658±0.0106 0.2479±0.0201 KMVE 18 0.1682±0.0146 0.1041±0.0090 0.0598±0.0054 0.0289±0.0038 0.1648±0.0069 0.2244±0.0038 ThyroidXAgent0.5051±0.02280.4137±0.02060.3447±0.01880.2909±0.01740.3293±0.01190.5480±0.0160 42/51 Table S15.Clinical semantic performance of thyroid ultrasound report generation across datasets. False discovery rate (FDR), feature accuracy, lesion-level F1 score, completeness, consistency and ThyClinScore are reported on the SMU-HMC, KMVE and ZJH-TS test sets. Values are means±the half-width of the bootstrap 95% percentile confidence interval. Lower FDR indicates better performance. ModelFDR↓Feat AccF1 ScoreComplete.Consist.ThyClin SMU-HMC Testset(n=400) GPT-4o 95 0.7189±0.0362 0.5555±0.0351 0.2526±0.0332 0.8208±0.0132 0.4023±0.0182 0.4105±0.0172 GPT-5 68 0.4170±0.0417 0.5644±0.0351 0.4390±0.0402 0.8585±0.0077 0.4980±0.0166 0.4882±0.0182 Gemini2.5 Pro 69 0.6797±0.0388 0.5586±0.0385 0.2965±0.0366 0.8789±0.0107 0.4040±0.0187 0.4280±0.0176 Qwen3.5 Plus 96 0.7426±0.0295 0.5804±0.0373 0.2656±0.0290 0.9620±0.0063 0.4809±0.0142 0.4883±0.0137 Claude-Sonnet-4.6 0.8040±0.0305 0.5580±0.0464 0.2049±0.0298 0.8786±0.0108 0.3886±0.0147 0.4015±0.0144 MedGemma 71 0.8943±0.0229 0.5784±0.0434 0.1027±0.0202 0.9657±0.0055 0.3353±0.0119 0.3697±0.0131 LLaVA-Med 97 0.1625±0.03630.4100±0.2125 0.2582±0.0426 0.4576±0.0227 0.1387±0.0166 0.1752±0.0150 KMVE 18 0.4813±0.0488 0.5944±0.1233 0.2103±0.0390 0.5995±0.0085 0.2013±0.0144 0.2265±0.0145 ThyroidXAgent0.2238±0.03770.6238±0.03580.5467±0.04360.9691±0.00600.5016±0.01660.5189±0.0212 KMVE Testset(n=492) GPT-4o 95 0.5843±0.0427 0.6608±0.0691 0.2137±0.0350 0.6180±0.0113 0.4275±0.0231 0.3344±0.0188 GPT-5 68 0.4133±0.0423 0.6967±0.0492 0.3749±0.0402 0.6506±0.0038 0.4504±0.0189 0.3853±0.0189 Gemini2.5 Pro 69 0.3211±0.0407 0.5093±0.04180.4648±0.04010.6270±0.0045 0.4135±0.0213 0.3810±0.0187 Qwen3.5 Plus 96 0.4553±0.0422 0.6154±0.0448 0.3823±0.0397 0.6525±0.0050 0.4035±0.0230 0.3674±0.0193 Claude-Sonnet-4.6 0.5803±0.0432 0.6605±0.0578 0.2721±0.03750.6707±0.00430.3779±0.0211 0.3362±0.0182 MedGemma 71 0.1877±0.0327 0.6820±0.0463 0.2868±0.0381 0.6622±0.0033 0.5172±0.0240 0.3932±0.0203 LLaVA-Med 97 0.0000±0.0000 * N/A * 0.2541±0.0386 0.6508±0.00070.5669±0.02250.3928±0.0208 ThyroidXAgent0.4858±0.03860.7366±0.03550.3623±0.03580.6416±0.00290.5654±0.02170.4407±0.0194 ZJH-TS Testset(n=150) GPT-4o 95 0.4583±0.0706 0.5510±0.0468 0.2706±0.0523 0.7290±0.0305 0.3415±0.0294 0.3139±0.0260 GPT-5 68 0.5201±0.0658 0.5263±0.0487 0.3771±0.0531 0.9534±0.0112 0.4735±0.0243 0.4472±0.0248 Gemini2.5 Pro 69 0.5450±0.0689 0.5507±0.0547 0.3157±0.0523 0.7859±0.0183 0.3910±0.0292 0.3557±0.0244 Qwen3.5 Plus 96 0.5781±0.0558 0.5464±0.0465 0.3902±0.0485 0.9575±0.00860.4767±0.02190.4520±0.0220 Claude-Sonnet-4.6 0.6186±0.0571 0.5748±0.0455 0.3405±0.0496 0.8226±0.0176 0.4298±0.0233 0.3861±0.0219 MedGemma 71 0.6537±0.0644 0.5463±0.0498 0.2959±0.0520 0.9437±0.0106 0.3868±0.0225 0.3807±0.0238 LLaVA-Med 97 0.2733±0.0733 0.6595±0.10770.1241±0.0454 0.6136±0.0494 0.1505±0.0308 0.1812±0.0275 KMVE 18 0.7911±0.0617 0.6288±0.1490 0.0445±0.0253 0.6252±0.0085 0.1581±0.0162 0.1650±0.0124 ThyroidXAgent0.2833±0.06000.5820±0.04080.4889±0.05700.9641±0.00900.4564±0.02280.4676±0.0268 * On the KMVE dataset, LLaVA-Med collapsed and predicted “no abnormality” for all test samples. Therefore, Feat Acc is N/A and FDR is 0 because no positive predictions were made. 43/51 Table S16.Static-pipeline metrics used for report-generation radar plots. Conventional language-generation metrics and clinical semantic metrics are reported for the SMU-HMC, KMVE and ZJH-TS test sets used in Fig.5g. Values are means±the half-width of the 95% confidence interval. Lower FDR indicates better performance. Conventional natural-language generation metrics DatasetnBLEU-1BLEU-2BLEU-3BLEU-4METEOR ROUGE L SMU-HMC 400 0.4586±0.0149 0.3817±0.0139 0.3271±0.0136 0.2849±0.0133 0.3209±0.0073 0.4725±0.0115 KMVE492 0.3266±0.0069 0.2105±0.0048 0.1306±0.0034 0.0624±0.0034 0.2724±0.0051 0.2750±0.0045 ZJH-TS150 0.4242±0.0234 0.3527±0.0206 0.2986±0.0186 0.2534±0.0172 0.2916±0.0101 0.4994±0.0140 Clinical semantic metrics DatasetnFDR↓Feat AccF1 Score Complete.Consist.ThyClin SMU-HMC 400 0.3162±0.0431 0.6227±0.0382 0.5060±0.0453 0.7821±0.0064 0.4325±0.0161 0.4293±0.0192 KMVE492 0.5894±0.0432 0.5616±0.0631 0.2644±0.0375 0.7665±0.0059 0.3391±0.0181 0.3346±0.0174 ZJH-TS150 0.4433±0.0717 0.5517±0.0516 0.3894±0.0603 0.8127±0.0120 0.3734±0.0201 0.3648±0.0239 44/51 Table S17.Ablation of tool integration for thyroid ultrasound report generation. Segmentation, classification, captioning and measurement tools were added cumulatively, and performance was evaluated using conventional language-generation metrics. Values are means±the half-width of the 95% confidence interval. ConfigurationBLEU-1BLEU-2BLEU-3BLEU-4METEOR ROUGE L SMU-HMC Testset(n=400) Segmentation only 0.0897±0.0095 0.0692±0.0075 0.0562±0.0063 0.0459±0.0053 0.1521±0.0047 0.2603±0.0092 - Classification0.1877±0.0169 0.1431±0.0130 0.1148±0.0104 0.0929±0.0085 0.1849±0.0072 0.2870±0.0106 - Captioning0.4503±0.0163 0.3724±0.0150 0.3179±0.0142 0.2764±0.0139 0.3052±0.0081 0.4601±0.0126 - Measurement (full) 0.5924±0.0143 0.4806±0.0139 0.4006±0.0137 0.3381±0.0135 0.3627±0.0088 0.5422±0.0122 KMVE Testset ( n = 492 ) Segmentation only 0.6130±0.0191 0.5456±0.0186 0.4907±0.0182 0.4462±0.0186 0.3658±0.0103 0.5747±0.0126 - Classification0.6315±0.0189 0.5637±0.0184 0.5081±0.0180 0.4633±0.0183 0.3751±0.0102 0.5799±0.0127 - Captioning0.6357±0.0172 0.5606±0.0160 0.5008±0.0152 0.4535±0.0153 0.3672±0.0099 0.5880±0.0122 - Measurement (full) 0.6357±0.0172 0.5606±0.0160 0.5008±0.0152 0.4535±0.0153 0.3672±0.0099 0.5880±0.0122 ZJH-TS Testset(n=150) Segmentation only 0.0964±0.0150 0.0799±0.0123 0.0674±0.0105 0.0567±0.0091 0.1500±0.0073 0.3348±0.0129 - Classification0.2456±0.0246 0.1990±0.0197 0.1648±0.0163 0.1356±0.0137 0.2053±0.0105 0.4003±0.0142 - Captioning0.4092±0.0252 0.3409±0.0220 0.2891±0.0197 0.2461±0.0180 0.2848±0.0115 0.4903±0.0162 - Measurement (full) 0.5051±0.0227 0.4137±0.0205 0.3447±0.0188 0.2909±0.0174 0.3293±0.0118 0.5480±0.0158 The KMVE dataset retains only the findings section and does not provide original measurement values. To match the original evaluation protocol, only the generated findings section was evaluated and measurement values were masked; therefore, the captioning and full configurations have identical KMVE scores. 45/51 References 1.Alexander, E. K. & Cibas, E. S. Diagnosis of thyroid nodules.The Lancet Diabetes & Endocrinol.10, 533–539, DOI:10.1016/S2213-8587(22)00101-2(2022). 2.Grani, G., Sponziello, M., Filetti, S. & Durante, C. Thyroid nodules: diagnosis and management.Nat. Rev. Endocrinol.DOI:10.1038/s41574-024-01025-4(2024). 3.Tessler, F. N.et al.Acr thyroid imaging, reporting and data system (ti-rads): White paper of the acr ti-rads committee.J. Am. Coll. Radiol.14, 587–595, DOI:10.1016/j.jacr.2017.01.046(2017). 4.Hoang, J. K.et al.Interobserver variability of sonographic features used in the american college of radiology thyroid imaging reporting and data system.Am. J. Roentgenol.211, 162–167, DOI:10.2214/AJR.17.19192 (2018). 5.Cibas, E. S. & Ali, S. Z. The 2017 bethesda system for reporting thyroid cytopathology.Thyroid27, 1341–1346, DOI:10.1089/thy.2017.0500(2017). 6.Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence.Nat. Medicine 25, 44–56, DOI:10.1038/s41591-018-0300-7(2019). 7.Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. Ai in health and medicine.Nat. Medicine28, 31–38, DOI:10.1038/s41591-021-01614-0(2022). 8.Chen, H., Gomez, C., Huang, C.-M.et al.Explainable medical imaging AI needs human-centered design: Guidelines and evidence from a systematic review.npj Digit. Medicine5, 156, DOI: 10.1038/s41746-022-00699-2 (2022). 9.Wekenborg, M. K., Gilbert, S. & Kather, J. N. Examining human–AI interaction in real-world healthcare beyond the laboratory.npj Digit. Medicine8, 169, DOI: 10.1038/s41746-025-01559-5(2025). 10.Giddings, R.et al.Factors influencing clinician and patient interaction with machine learning-based risk prediction models: A systematic review.The Lancet Digit. Heal.6, e131–e144, DOI:10.1016/S2589-7500(23) 00241-8(2024). 11.Gong, H.et al.Multi-task learning for thyroid nodule segmentation with thyroid region prior. In2021 IEEE 18th international symposium on biomedical imaging (ISBI), 257–261 (2021). 12.Gong, H.et al.Thyroid region prior guided attention for ultrasound segmentation of thyroid nodules.Comput. biology medicine155, 106389 (2023). 13.Sun, X., Wei, B., Jiang, Y., Mao, L. & Zhao, Q. Clip-tnseg: A multi-modal hybrid framework for thyroid nodule segmentation in ultrasound images.arXiv preprint arXiv:2412.05530(2024). 14.Gong, H.et al.Less is more: adaptive curriculum learning for thyroid nodule diagnosis. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention, 248–257 (2022). 15.Peng, S.et al.Deep learning-based artificial intelligence model to assist thyroid nodule diagnosis and man- agement: a multicentre diagnostic study.The Lancet Digit. Heal.3, e250–e259, DOI:10.1016/S2589-7500(21) 00041-8 (2021). 16.Chen, Y.et al.An artificial intelligence model based on acr ti-rads characteristics for us diagnosis of thyroid nodules.Radiology303, 613–619, DOI:10.1148/radiol.211455(2022). 17.Yao, J.et al.Multimodal gpt model for assisting thyroid nodule diagnosis and management.npj Digit. Medicine 8, 245, DOI:10.1038/s41746-025-01652-9(2025). 46/51 18.Li, J., Su, T.et al.Ultrasound report generation with cross-modality feature alignment via unsupervised guidance.IEEE Transactions on Med. Imaging44, 19–30 (2024). 19.Tanno, R., Barrett, D. G. T., Sellergren, A.et al.Collaboration between clinicians and vision–language models in radiology report generation.Nat. Medicine31, 599–608, DOI:10.1038/s41591-024-03302-1(2025). 20.Li, C.-Y., Chang, K.-J.et al.Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation.Nat. Commun.16, 2258 (2025). 21.Wang, J.et al.Deep learning models for thyroid nodules diagnosis of fine-needle aspiration biopsy: a ret- rospective, prospective, multicentre study in china.The Lancet Digit. Heal.6, e458–e469, DOI:10.1016/ S2589-7500(24)00085-2(2024). 22.Shen, P.et al.Explainable multimodal deep learning for predicting thyroid cancer lateral lymph node metastasis using ultrasound imaging.Nat. Commun.16, 7052, DOI:10.1038/s41467-025-62042-z(2025). 23.Dai, F.et al.Improving ai models for rare thyroid cancer subtype by text guided diffusion models.Nat. Commun.16, 4449, DOI:10.1038/s41467-025-59478-8(2025). 24.Tikhomirov, L.et al.Medical artificial intelligence for clinicians: The lost cognitive perspective.The Lancet Digit. Heal.6, e589–e594, DOI:10.1016/S2589-7500(24)00095-5(2024). 25.You, G., Li, H., Zhang, Y. & Fan, Y. Learning anatomy-grounded CT vision-language representations with organ-hierarchical report knowledge.arXiv preprint arXiv:2607.10953DOI:10.48550/arXiv.2607.10953(2026). 26.Vasey, B.et al.Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: Decide-ai.Nat. Medicine28, 924–933, DOI: 10.1038/s41591-022-01772-9(2022). 27.Dreyer, M.et al.Mechanistic understanding and validation of large ai models with semanticlens.Nat. Mach. Intell.7, 1572–1585 (2025). 28.Patel, B. N., Rosenberg, L., Willcox, G.et al.Human–machine partnership with artificial intelligence for chest radiograph diagnosis.npj Digit. Medicine2, 111, DOI:10.1038/s41746-019-0189-7(2019). 29.Leibig, C.et al.Combining the strengths of radiologists and AI for breast cancer screening: A retrospective analysis.The Lancet Digit. Heal.4, e507–e519, DOI:10.1016/S2589-7500(22)00070-X(2022). 30.Yu, F.et al.Heterogeneity and predictors of the effects of AI assistance on radiologists.Nat. Medicine30, 837–849, DOI:10.1038/s41591-024-02850-w(2024). 31.Chen, M., Wang, Y., Wang, Q.et al.Impact of human and artificial intelligence collaboration on workload reduction in medical image interpretation.npj Digit. Medicine7, 349, DOI:10.1038/s41746-024-01328-w (2024). 32.Everett, S. S., Bunning, B. J., Jain, P.et al.From tool to teammate in a randomized controlled trial of clinician– AI collaborative workflows for diagnosis.npj Digit. Medicine9, 409, DOI:10.1038/s41746-026-02545-1(2026). 33.Strong, J., Rogers, H., Sun, E.et al.Human–AI collaboration in healthcare: A scoping review.npj Digit. MedicineDOI: 10.1038/s41746-026-02918-6(2026). 34.Wiens, J.et al.Do no harm: a roadmap for responsible machine learning for health care.Nat. Medicine25, 1337–1340, DOI:10.1038/s41591-019-0548-6(2019). 35.Zou, J. & Topol, E. J. The rise of agentic ai teammates in medicine.The Lancet405, 457, DOI:10.1016/ S0140-6736(25)00202-8(2025). 47/51 36.Moor, M.et al.Foundation models for generalist medical artificial intelligence.Nature616, 259–265, DOI: 10.1038/s41586-023-05881-4(2023). 37.Kohane, I. S. Injecting artificial intelligence into medicine.NEJM AI1, 1–3, DOI:10.1056/AIe2300197(2024). 38.Katz, U., Cohen, E., Shachar, E.et al.GPT versus resident physicians—a benchmark based on official board scores.NEJM AI1, DOI:10.1056/AIdbp2300192(2024). 39.Zhou, H.-Y.et al.Medversa: A generalist foundation model for diverse medical imaging tasks.NEJM AI3, DOI:10.1056/AIoa2500595(2026). 40.Tu, T., Schaekermann, M., Palepu, A.et al.Towards conversational diagnostic artificial intelligence.Nature 642, 442–450, DOI:10.1038/s41586-025-08866-7(2025). 41.McDuff, D., Schaekermann, M., Tu, T.et al.Towards accurate differential diagnosis with large language models. Nature642, 451–457, DOI:10.1038/s41586-025-08869-4(2025). 42.Deltadahl, S.et al.Deep generative classification of blood cell morphology.Nat. Mach. Intell.7, 1791–1803 (2025). 43.Pontikos, N.et al.Next-generation phenotyping of inherited retinal diseases from multimodal imaging with eye2gene. Nat. Mach. Intell. 7 , 967–978 (2025). 44.Qiu, J.et al.Llm-based agentic systems in medicine and healthcare.Nat. Mach. Intell.6, 1418–1420, DOI: 10.1038/s42256-024-00944-1(2024). 45.Moritz, M., Topol, E. & Rajpurkar, P. Coordinated ai agents for advancing healthcare.Nat. Biomed. Eng.9, 432–438, DOI:10.1038/s41551-025-01363-2(2025). 46.Ferber, D., Hilgers, L., H”oper, C.et al.Towards autonomous medical artificial intelligence agents.Nature DOI:10.1038/s41586-026-10675-5(2026). 47.Collaco, B. G., Haider, S. A., Prabha, S.et al.The role of agentic artificial intelligence in healthcare: A scoping review.npj Digit. Medicine9, 345, DOI:10.1038/s41746-026-02517-5(2026). 48.Kong, Q.et al.Ai agent-based discovery of d-enantiomeric antimicrobial peptides against multidrug-resistant bacterial infection.Biomaterials123927 (2025). 49.Wang, L., Xu, W.et al.Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. InProceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), 2609–2634 (2023). 50.Schmidgall, S., Ziaei, R., Harris, C.et al.AgentClinic: A multimodal benchmark for tool-using clinical AI agents.npj Digit. Medicine9, 499, DOI:10.1038/s41746-026-02674-7(2026). 51.Liu, Y., Carrero, Z. I., Jiang, X.et al.Benchmarking large language model-based agent systems for clinical decision tasks.npj Digit. Medicine9, 259, DOI:10.1038/s41746-026-02443-6(2026). 52.Yao, S., Zhao, J.et al.React: Synergizing reasoning and acting in language models. InThe eleventh interna- tional conference on learning representations(2022). 53.Tian, J., Fard, P., Cagan, C.et al.An autonomous agentic workflow for clinical detection of cognitive concerns using large language models.npj Digit. Medicine9, 51, DOI:10.1038/s41746-025-02324-4(2026). 54.Jiang, Y.et al.Medagentbench: A virtual ehr environment to benchmark medical llm agents.NEJM AI AIdbp2500144, DOI:10.1056/AIdbp2500144(2025). 48/51 55.Zhang, H., Liu, Q., Han, X.et al.Tn5000: An ultrasound image dataset for thyroid nodule detection and classification.Sci. Data12, 1437, DOI:10.1038/s41597-025-05757-4(2025). 56.Van Griethuysen, J. J.et al.Computational radiomics system to decode the radiographic phenotype.Cancer research77, e104–e107 (2017). 57.Rebuffel, C., Soulier, L.et al.A hierarchical model for data-to-text generation. InEuropean Conference on Information Retrieval, 65–80 (Springer, 2020). 58.Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. Detecting hallucinations in large language models using semantic entropy.Nature630, 625–630 (2024). 59.Duong, V. H.et al.Thyroidxl: Advancing thyroid nodule diagnosis with an expert-labeled, pathology-validated dataset. InMedical Image Computing and Computer Assisted Intervention – MICCAI 2025, vol. 15974, 616–626, DOI:10.1007/978-3-032-05182-0_60(Springer Nature Switzerland, 2025). 60.Pedraza, L.et al.An open access thyroid ultrasound image database. In10th International Symposium on Medical Information Processing and Analysis, vol. 9287, 188–193 (2015). 61.Shusharina, N., Heinrich, M. P. & Huang, R.Segmentation, Classification, and Registration of Multi-modality Medical Imaging Data: MICCAI 2020 Challenges, ABCs 2020, L2R 2020, TN-SCUI 2020, Held in Conjunction with MICCAI 2020, Lima, Peru, October 4–8, 2020, Proceedings(Springer Nature, 2021). 62.Dai, F.et al.Improving ai models for rare thyroid cancer subtype by text guided diffusion models.Nat. Commun.16, 4449 (2025). 63.Liu, Z. & He, K. A decade’s battle on dataset bias: Are we there yet? InInternational Conference on Learning Representations(2025). 64.Ma, J.et al.Medsam2: Segment anything in 3d medical images and videos.arXiv preprint arXiv:2504.03600 (2025). 65.Jiang, Y.et al.From pretraining to privacy: Federated ultrasound foundation model with self-supervised learning.npj Digit. Medicine8, 714, DOI:10.1038/s41746-025-02085-0(2025). 66.Wang, A., Chen, H., Lin, Z., Han, J. & Ding, G. Repvit: Revisiting mobile cnn from vit perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 15909–15920 (2024). 67.Dong, C.et al.A survey of natural language generation.ACM Comput. Surv.55, 1–38 (2022). 68.OpenAI. Gpt-5 system card.https://openai.com/index/gpt-5-system-card/(2025). Published August 7, 2025. 69.Comanici, G., Bieber, E.et al.Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities.arXiv preprint arXiv:2507.06261(2025). 70.Zhang, S.et al.A multimodal biomedical foundation model trained from fifteen million imagetext pairs. NEJM AI2, DOI: 10.1056/AIoa2400640(2024). 71.Sellergren, A.et al.Medgemma technical report.arXiv preprint arXiv:2507.05201(2025). 72.Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. InAdvances in Neural Information Processing Systems, vol. 30 (2017). 73.Papineni, K., Roukos, S. et al. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311–318 (2002). 74.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. InText Summarization Branches Out, 74–81 (2004). 49/51 75.Banerjee, S. & Lavie, A. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. InProceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, 65–72 (2005). 76.Obermeyer, Z., Powers, B., Vogeli, C. & Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations.Science366, 447–453, DOI:10.1126/science.aax2342(2019). 77.Gong, H., Lu, Y., Wan, X. & Li, H. Domain generalized medical landmark detection via robust boundary-aware pre-training. InProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, 3140–3148 (2025). 78.Gong, H.et al.Intermediate domain alignment and morphology analogy for patent-product image retrieval. Adv. Neural Inf. Process. Syst.38, 14501–14523 (2026). 79.Mohammadi, A.et al.Lymphus: A multicenter open-access database of lymph node ultrasound images in patients with papillary thyroid carcinoma for clinical and artificial intelligence research.Data Brief66, 112694, DOI:10.1016/j.dib.2026.112694(2026). 80.Abbasian Ardakani, A.et al.Diagnosis of metastatic lymph nodes in patients with papillary thyroid cancer: A comparative multi-center study of semantic features and deep learning-based models. J. Ultrasound Medicine 42, 1211–1221, DOI:10.1002/jum.16131(2023). 81.Torralba, A. & Efros, A. A. Unbiased look at dataset bias. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 1521–1528 (2011). 82.Siméoni, O.et al.Dinov3.arXiv preprint arXiv:2508.10104(2025). 83.Menon, A. K.et al.Long-tail learning via logit adjustment.arXiv preprint arXiv:2007.07314(2020). 84.Erickson, N.et al.Autogluon-tabular: Robust and accurate automl for structured data.arXiv preprint arXiv:2003.06505(2020). 85.Selvaraju, R. R.et al.Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision, 618–626 (2017). 86.Wunderling, T.et al.Comparison of thyroid segmentation techniques for 3d ultrasound. InMedical Imaging 2017: Image Processing, vol. 10133, 346–352 (2017). 87.Hou, X. et al. An ultrasonography of thyroid nodules dataset with pathological diagnosis annotation for deep learning.Sci. Data11, 1272, DOI:10.1038/s41597-024-04156-5(2024). 88.Stanford AIMI. Thyroid ultrasound cine-clip dataset, DOI:10.71718/7m5n-rh16(2024). 89.Yang, Y.et al.An annotated heterogeneous ultrasound database.Sci. Data12, 148 (2025). 90.Chen, J.et al.Transunet: Rethinking the u-net architecture design for medical image segmentation through the lens of transformers.Med. Image Analysis97, 103280, DOI:10.1016/j.media.2024.103280(2024). 91.Zhang, S.et al.A generalist foundation model and database for open-world medical image segmentation.Nat. Biomed. Eng.10, 1026–1041, DOI:10.1038/s41551-025-01497-3(2026). Published online 5 September 2025. 92.He, K., Zhang, X., Ren, S. & Sun, J. Deep residual learning for image recognition. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778 (2016). 93.Wang, A., Chen, H., Lin, Z., Han, J. & Ding, G. Lsnet: See large, focus small. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9718–9729 (2025). 94.Bai, S., Cai, Y., Chen, R.et al.Qwen3-vl technical report.arXiv preprint arXiv:2511.21631(2025). 50/51 95.OpenAI. Gpt-4o system card.arXiv preprint arXiv:2410.21276(2024). 96.Yang, A., Li, A.et al.Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025). 97.Li, C., Wong, C.et al.Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Adv. Neural Inf. Process. Syst.36, 28541–28564 (2023). 51/51