Paper deep dive
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicolás Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination.
Tags
Links
- Source: https://arxiv.org/abs/2607.25933v1
- Canonical: https://arxiv.org/abs/2607.25933v1
Trouble viewing inline? Open PDF directly →
Full Text
85,001 characters extracted from source content.
Expand or collapse full text
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Rui Yang 1,2,3,4 , Weihao Xuan 5,6 , Yi Lin 3 , Zhuhan Bao 7 , Jonathan Chong Kai Liew 8 , Matthew Yu Heng Wong 9 , Nicolás Lescano 8,10 , Nikita R. Paripati 8,10,11 , Emily Ling-Lin Pai 12,13 , Jiarui Liu 14 , Heli Qi 6 , Heng-Jui Chang 15 , Benny Kai Guo Loo 16,17 , Huitao Li 1,2 , Kunyu Yu 1,2 , Yufan Wang 3 , Chuan Hong 7 , Shijian Lu 18 , Douglas Teodoro 19 , Naoto Yokoya 5,6 , Ross Koppel 8,20 , Mona Diab 14 , Hua Xu 21 , David W. Bates 22,23,24 , Nan Liu 1,2,7,25,26† and Yifan Peng 3,27† 1 Center for Biomedical Data Science, Duke-NUS Medical School, Singapore, Singapore 2 Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School, Singapore, Singapore 3 Department of Population Health Sciences, Weill Cornell Medicine, New York, NY, USA 4 System Engineering, College of Engineering, Cornell University, Ithaca, NY, USA 5 Graduate School of Frontier Sciences, The University of Tokyo, Chiba, Japan 6 RIKEN Center for Advanced Intelligence Project, Tokyo, Japan 7 Department of Biostatistics and Bioinformatics, Duke University, Durham, NC, USA 8 Perelman School of Medicine, University of Pennsylvania, Philadelphia, PA, USA 9 School of Clinical Medicine, University of Cambridge, Cambridge, UK 10 Hospital of the University of Pennsylvania, Philadelphia, PA, USA 11 Children’s Hospital of Philadelphia (CHOP), Philadelphia, PA, USA 12 Department of Anatomic Pathology and Laboratory Medicine, Hospital of the University of Pennsylvania, PA, USA 13 Department of Pathology and Laboratory Medicine, University of California, San Francisco, CA, USA 14 Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA, USA 15 Department of Chemistry, Stanford University, Stanford, CA 16 Sport and Exercise Medicine Service, K Women’s and Children’s Hospital, Singapore, Singapore 17 Paediatrics Academic Clinical Programme, SingHealth Duke-NUS Academic Medical center, Singapore, Singapore 18 College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore 19 Department of Radiology and Medical Informatics, University of Geneva, Geneva, Switzerland 20 Department of Biomedical Informatics, Jacobs School of Medicine, University at Buffalo, Buffalo, NY, United States 21 Department of Biomedical Informatics and Data Science, Yale School of Medicine, New Haven, CT, USA 22 Division of General Internal Medicine and Primary Care, Brigham and Women’s Hospital, Boston, MA, USA 23 Department of Medicine, Harvard Medical School, Boston, MA, USA 24 Department of Health Care Policy and Management, Harvard T. H. Chan School of Public Health, Boston, MA, USA 25 Pre-hospital and Emergency Research Center, Health Services Research and Population Health, Duke-NUS Medical School, Singapore, Singapore 26 NUS Artificial Intelligence Institute, National University of Singapore, Singapore, Singapore 27 Institute of Artificial Intelligence for Digital Health, Weill Cornell Medicine, New York, NY, USA † Corresponding authors. arXiv:2607.25933v1 [cs.CL] 28 Jul 2026 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Corresponding Authors: Nan Liu, Center for Biomedical Data Science, Duke-NUS Medical School, Singapore, Singapore Email: liu.nan@duke-nus.edu.sg Yifan Peng, Department of Population Health Sciences, Weill Cornell Medicine, New York, NY, USA Email: yip4002@med.cornell.edu Abstract Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal infor- mation, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single- turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reason- ing quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directions but still have considerable limitations in generating reliable diagnostic reasoning. Error analysis further identified five representative failure modes: information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination. Keywords: Multimodal Large Language Models, Clinical Diagnostic Reasoning, Multi-Turn Multimodal Diagnosis, Real-World Clinical Cases 1. Introduction Diagnostic decision-making lies at the core of clinical practice and relies on the progressive integration of patient-specific clinical information from multiple sources 1 . In routine clinical practice, physicians rarely make diagnoses based on a single symptom, image finding, or lab result. Instead, they update diagnostic hypotheses as additional information becomes available, reconcile conflicting evidence, and determine which diagnosis best accounts for the overall clinical picture 2–5 . Together, these demands make clinical diagnosis an inherently dynamic and context-dependent process 6 . Recent advances in generative artificial intelligence (AI) have accelerated its adoption in medicine, with growing potential in clinical consultation, disease diagnosis, and patient management 7–10 . In particular, multimodal large language models (MLLMs) have shown increasing capacity to process multimodal medical data and have achieved promising performance on tasks such as image interpre- tation and report generation 11 . As MLLMs continue to improve in domain knowledge representation and cross-modal understanding, their potential to support diagnostic decision-making is expected to further expand 12 . Despite these advancements, it remains unclear whether MLLMs can perform diagnostic reasoning when clinical information is progressively disclosed over time and across modalities. Current evalua- tions are insufficient to capture this capability 13 . First, existing benchmarks primarily adopt a single- turn, static question-answering (QA) format, in which complete clinical information is provided to 2 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases the model at once to generate an answer 14,15 . Meanwhile, they typically focus on isolated tasks, such as image understanding or report generation 16 . These settings underestimate the complexity of clinical reasoning because they do not evaluate progressive multimodal information synthesis, diagnostic hypothesis revision, cross-turn memory, or the ability to distinguish clinically grounded reasoning from plausible but hallucinated reasoning. Moreover, most benchmarks emphasize the accuracy of the diagnoses while overlooking the quality of the reasoning process 17 . The reliability of a clinical diagnosis depends not only on whether the conclusion is correct, but also on whether the reasoning is grounded in evidence, sufficiently complete, and logically coherent 18,19 . Lastly, al- though some studies have attempted to construct multi-turn diagnostic evaluations, they are often limited in scale or not publicly available, making it difficult to establish a reproducible and extensible evaluation framework 1,20,21 . These limitations create a critical gap in the evaluation of MLLMs in medicine. A model may correctly answer a static medical question yet fail to synthesize evolving clinical information, misinterpret medical image findings, hallucinate lab results, or persist with an early incorrect hypothesis 22–24 . On the other hand, a model may generate a partially correct diagnosis while omitting essential reasoning steps that are necessary for clinical trust 25 . To bridge these gaps, we introduce ClinMM-Bench, a multimodal, multi-turn benchmark designed to evaluate both diagnostic accuracy and reasoning quality in challenging real-world clinical cases. ClinMM-Bench contains 1,089 clinical cases and 3,760 medical images across eight specialties: Der- matology, Emergency Medicine, Internal Medicine, Nephrology, Neurology, Oncology, Ophthalmol- ogy, and Radiology. Each case is designed as a multi-turn diagnostic scenario in which clinical in- formation and images are progressively disclosed, enabling evaluation of how models synthesize multimodal information and update diagnostic hypotheses over time. This study makes three main contributions (Fig. 1): (1) We construct ClinMM-Bench, to our knowl- edge, the largest benchmark to date for multi-turn multimodal diagnostic reasoning on challenging real-world cases; (2) We propose a two-level evaluation framework that measures both diagnostic accuracy and reasoning quality. Diagnostic accuracy is assessed using a dual-LLM consensus mecha- nism, and reasoning quality is quantified using atomic fact decomposition; fact recall, hallucination, and fact density measure the completeness, reliability, and efficiency of MLLMs in clinical diagnos- tic reasoning, respectively; and (3) We systematically evaluate 15 representative MLLMs, covering proprietary models, open-weight models, medical models, and reasoning models. Through in-depth analyses, we reveal capability disparities across different models and identify five representative failure modes (information synthesis failure, knowledge mapping error, perception error, premature closure, and visual hallucination), providing important insights for the future development of MLLMs in medicine. 3 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Data Collection and Extraction Diagnostic Accuracy Evaluation Data Curation Multi-Turn Multimodal Evaluation video no-image low-resolution low Data Conversion Structured Text Images Quality Control Expert Validation 3 Please review this medical case presentation. ...... Next, I will show you the medical images one by one. Please analyze each image. Understood. Please share the first image, and I will analyze it step by step in the context of the clinical presentation. ...... Diagnostic Reasoning Quality Evaluation Ground Truth Predicted Diagosis Score 012 Reference Reasoning LLM Reasoning Fact DensityHallucinationFact Recall Explaination Data ValidationData Inspection Fig. 1 | Overview of ClinMM-Bench and evaluation framework. a, Data curation. ClinMM-Bench was developed through a six-stage pipeline: (1) Data collection and extraction, in which clinical case reports were collected from PubMed Central Open Access; (2) Data inspection, where case reports without medical images, those containing videos, and those with low-resolution medical images were excluded; (3) Data validation, using a dual-LLM consensus mechanism to identify cases suitable for the clinical diagnostic reasoning task; (4) Data conversion, where validated case reports were parsed and transformed into a standardized structured format; (5) Quality control, where an automated scoring procedure assessed the structured case reports and retained only those meeting a predefined quality threshold; and (6) Expert validation, where medical ex- perts performed manual verification to ensure data reliability. b, Multi-turn multimodal evaluation. During evaluation, models perform diagnostic reasoning through multi-turn dialogues, with clinical information and images of each case progressively disclosed over the course of the dialogue. The evaluation framework com- prises two levels: (1) Diagnostic accuracy evaluation, in which a dual-LLM consensus mechanism compares MLLM-predicted diagnoses against the ground-truth diagnoses, with judge LLMs assigning accuracy scores ranging from 0 to 2. (2) Diagnostic reasoning quality evaluation, where both MLLM-generated reasoning and reference reasoning are decomposed into atomic facts and quantified across three dimensions: fact recall, hallucination, and fact density. 4 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 2. Results 2.1. ClinMM-Bench Overview ClinMM-Bench comprises 1,089 challenging real-world clinical diagnostic cases curated from PubMed Central Open Access (PMCOA) case reports, with 3,760 medical images, spanning eight specialties. Radiology is the largest specialty, with 679 cases (62.35%), followed by Oncology with 114 cases (10.47%), Neurology with 98 cases (9.00%), Dermatology with 63 cases (5.79%), Ophthalmol- ogy with 60 cases (5.51%), Nephrology with 29 cases (2.66%), Emergency Medicine with 23 cases (2.11%), and Internal Medicine with 23 cases (2.11%). Each case was structured as a multi-turn multimodal diagnostic dialogue, with an average of 5.45 dialogue rounds per case. In each case, clinical information and images were progressively disclosed, requiring models to update their diag- noses over time. In addition to the ground-truth diagnosis, each case includes a reference diagnostic reasoning process. As ClinMM-Bench was curated from published case reports, it is enriched for diagnostically challenging cases involving relatively uncommon clinical presentations that are un- derrepresented in routine clinical records. 2.2. Diagnostic Accuracy 2.2.1. Overall Diagnostic Performance of 15 MLLMs on ClinMM-Bench We employed a dual-LLM consensus scoring strategy to evaluate the overall diagnostic accuracy of 15 representative MLLMs on ClinMM-Bench (Fig. 2a). Scores ranged from 0 to 2, where 0 indicated a completely incorrect or irrelevant diagnosis; 1 indicated a partially correct diagnosis; and 2 indicated a completely correct diagnosis. More details are provided in the Methods section under Diagnostic Accuracy Evaluation. 5 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Proprietary LLM Open-Weight General LLM Open-Weight Reasoning LLM Open-Weight Medical LLM 1.2 0.8 0.6 0.4 0.2 0 1.0 0.485 0.606 1.008 0.435 0.591 0.713 0.509 0.657 0.573 0.551 0.598 0.718 1.038 1.031 1.140 Qwen3-VI-4B-thinking MedGemma-4B Claude-4.5-Sonnet Gemini 3 Pro GPT-5-minimal GPT-5-medium Gemma-3-4B LLaMA-4-Scout Qwen3-VI-4B Gemma-3-27B Qwen3-VI-8B Qwen3-VI-32B Qwen3-VI-8B-thinking Qwen3-VI-32B-thinking MedGemma-27B 0 10 20 30 40 50 60 70 a. b. Claude-4.5-Sonnet 43.62% 28.65% Gemini 3 Pro 57.30% 22.41% Gemma-3-27B 49.31% 5.79% Gemma-3-4B 42.61% 2.48% GPT-5-medium 45.18% 33.88% GPT-5-minimal 48.48% 27.27% LLaMA-4-Scout 54.36% 8.82% MedGemma-27B 48.76% 6.61% MedGemma-4B 47.57% 2.48% Qwen3-VI-32B 49.95% 11.20% Qwen3-VI-32B-thinking 48.48% 9.27% Qwen3-VI-4B 48.76% 4.68 % Qwen3-VI-4B-thinking 45.27% 3.95% Qwen3-VI -8B 48.85% 6.52% Qwen3-VI-8B-thinking 46.74% 6.34% Overall Performace of Diagnostic Accuracy Score Proportions of Completely and Partially Correct Diagnoses Fig. 2 | Overall diagnostic performance of 15 MLLMs on ClinMM-Bench. a, Diagnostic accuracy score. Mean diagnostic accuracy scores of 15 representative MLLMs evaluated using the dual-LLM consensus scoring strategy. Scores range from 0 to 2, with 0 indicating a completely incorrect or irrelevant diagnosis; 1 indicating a partially correct diagnosis; and 2 indicating a completely correct diagnosis. b, Proportions of completely and partially correct diagnoses. Stacked bar plot showing the proportions of completely correct and partially correct diagnoses for each model. Darker shading represents completely correct diagnoses, and lighter shading represents partially correct diagnoses. Proprietary models achieved the strongest overall performance, with GPT-5-medium obtaining the highest diagnostic accuracy score (1.140), followed by Gemini 3 Pro (1.038) and GPT-5-minimal (1.031). In contrast, open-weight models achieved lower performance. Qwen3-VL-32B (0.718) and LLaMA-4-Scout (0.713) were the two best-performing open-weight models, but they still showed a gap compared with proprietary models. We further analyzed the proportions of completely correct (score 2) and partially correct (0< consen- sus score< 2) diagnoses (Fig. 2b). Among proprietary models, GPT-5-medium achieved the highest proportion of completely correct diagnoses (33.88%), followed by Claude-4.5-Sonnet (28.65%) and 6 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases GPT-5-minimal (27.27%). Among open-weight models, Qwen3-VL-32B was the only model with a completely correct diagnosis rate exceeding 10% (11.20%). Meanwhile, across all models, the proportion of partially correct diagnoses was considerably higher than that of completely correct diagnoses, a pattern that was particularly pronounced among open-weight models. For instance, LLaMA-4-Scout achieved a proportion of partially correct diagnoses of 54.36%, but only 8.82% of completely correct diagnoses. These results indicate that current models often identify plausible disease categories or related diagnostic directions but fail to produce precise diagnoses. Among open-weight models, model scale was positively associated with improved performance. For the Gemma series, increasing the model size from 4B to 27B improved the diagnostic accuracy score from 0.435 to 0.591, while the proportion of completely correct diagnoses increased from 2.48% to 5.79%. The MedGemma series showed a similar trend, with the diagnostic accuracy score increasing from 0.485 to 0.606 and the proportion of completely correct diagnoses increasing from 2.48% to 6.61% as the scale increased from 4B to 27B. In the Qwen3-VL non-reasoning series, as model size increased from 4B to 8B to 32B, diagnostic accuracy scores increased from 0.551 at 4B and 0.598 at 8B to 0.718 at 32B, while the proportion of completely correct diagnoses increased from 4.68% and 6.52% to 11.20%. Similarly, in the Qwen3-VL reasoning series, diagnostic accuracy scores increased from 0.509 at 4B and 0.573 at 8B to 0.657 at 32B, and the proportion of completely correct diagnoses increased from 3.95% and 6.34% to 9.27%. These results suggest that model scale remains important for multi-turn multimodal clinical diagnostic reasoning. 2.2.2. Variations in Diagnostic Accuracy Scores across Specialties Model performance varied dramatically across specialties (Fig. 3). Neurology achieved the highest overall diagnostic accuracy score (0.949 [95% confidence interval (CI): 0.880, 1.018]), followed by Ophthalmology (0.931 [95% CI: 0.797, 1.067]), Emergency Medicine (0.843 [95% CI: 0.677, 1.010]), and Nephrology (0.808 [95% CI: 0.666, 0.947]). Intermediate performance was observed in Dermatology (0.761 [95% CI: 0.639, 0.884]) and Oncology (0.699 [95% CI: 0.626, 0.772]), whereas Internal Medicine (0.665 [95% CI: 0.490, 0.851]) and Radiology (0.646 [95% CI: 0.613, 0.680]) were the most challenging specialties. Further details are provided in Supplementary In- formation A. 7 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 0.761 0.843 0.665 0.808 0.949 0.699 0.931 0.646 MedGemma-27B MedGemma-4B Qwen3-VL-32B-thinking Qwen3-VL-8B-thinking Qwen3-VL-4B-thinking Qwen3-VL-32B Qwen3-VL-8B Qwen3-VL-4B LLaMA-4-Scout Gemma-3-27B Gemma-3-4B GPT-5-medium GPT-5-minimal Gemini 3 Pro Claude-4.5-Sonnet Overall Derm.Emerg.Intern.Neph.Neuro.Onco.Ophth.Radiol. Mean Score 0.4 0.6 0.8 1.0 1.2 1.4 Fig. 3 | Diagnostic accuracy scores across specialties on ClinMM-Bench. Values indicate mean diagnostic accuracy scores, with 95% confidence intervals estimated using 100,000 bootstrap resamples. In the bubble plot, both circle size and color encode the mean diagnostic accuracy score, with larger and warmer-colored circles indicating higher scores. 2.2.3. Comparison between General and Medical MLLMs To examine whether medical-domain specialization improves diagnostic reasoning, we compared Gemma-3-4B and Gemma-3-27B with their corresponding medical variants, MedGemma-4B and MedGemma-27B (Fig. 4). At the 4B scale, MedGemma-4B outperformed Gemma-3-4B in most specialties, with the largest improvements in Nephrology (+0.241 in overall diagnostic accuracy), Emergency Medicine (+0.174), Neurology (+0.102), and Internal Medicine (+0.087). In addition, MedGemma-4B showed lower complete error rates across all specialties. However, the benefit was less consistent at the 27B scale. Compared with Gemma-3-27B, MedGemma-27B improved per- formance in selected specialties but showed no advantage in Nephrology, Neurology, and Radiology. These results suggest that medical-domain adaptation can improve diagnostic performance in smaller models, but does not uniformly benefit larger models. 8 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 0.532 0.413 0.391 0.466 0.577 0.474 0.658 0.380 0.484 0.587 0.478 0.707 0.679 0.461 0.683 0.431 -0.048 +0.174 +0.087 +0.241 +0.102 -0.013 +0.025 +0.051 Radiol. Ophth. Onco. Neuro. Neph. Intern. Emerg. Derm. 0.400.500.600.70 Diagnostic Accuracy Score Gemma-3-4B MedGemma-4B 0.492 0.522 0.565 0.517 0.439 0.526 0.417 0.588 0.476 0.391 0.478 0.276 0.286 0.500 0.367 0.558 -0.016 -0.130 -0.087 -0.241 -0.153 -0.026 -0.050 -0.029 0.300.400.500.60 Complete Error Rate Gemma-3-4B MedGemma-4B 0.659 0.652 0.457 0.707 0.893 0.535 0.775 0.532 0.738 0.717 0.587 0.690 0.852 0.596 0.858 0.531 +0.079 +0.065 +0.130 -0.017 -0.041 +0.061 +0.083 -0.001 Radiol. Ophth. Onco. Neuro. Neph. Intern. Emerg. Derm. 0.400.500.600.700.800.90 Diagnostic Accuracy Score Gemma-3-27B MedGemma-27B 0.413 0.391 0.609 0.310 0.204 0.482 0.333 0.495 0.333 0.391 0.478 0.310 0.214 0.421 0.333 0.511 -0.079 +0.000 -0.130 +0.000 +0.010 -0.061 +0.000 +0.016 0.200.300.400.500.60 Complete Error Rate Gemma-3-27B MedGemma-27B Fig. 4 | Comparison between general and medical MLLMs across specialties on ClinMM-Bench. Diag- nostic accuracy scores and complete error rates of Gemma-3-4B and Gemma-3-27B are compared with their corresponding medical variants, MedGemma-4B and MedGemma-27B, across eight specialties. The dashed line indicates the mean value across the eight specialties, and the shaded band represents the 95% confidence interval estimated using 100,000 bootstrap resamples. 2.2.4. Comparison between Non-Reasoning and Reasoning MLLMs To evaluate whether the reasoning setting improves multi-turn multimodal diagnosis, we compared the non-reasoning and reasoning versions of the Qwen3-VL series across the 4B, 8B, and 32B scales (Fig. 5). Overall, the reasoning setting did not consistently improve diagnostic accuracy. At the 4B scale, Qwen3-VL-4B-thinking achieved lower diagnostic accuracy scores than its corresponding non-reasoning variant in most specialties, with only a small improvement in Neurology (+0.056); meanwhile, its complete error rate was higher in most specialties. At the 8B scale, the effect of the reasoning setting showed greater specialty-level heterogeneity, improving diagnostic accuracy scores in Emergency Medicine, Nephrology, Oncology, and Ophthalmology, but reducing performance in Dermatology, Neurology, and Radiology. At the 32B scale, the reasoning variant achieved lower diagnostic accuracy scores than the non-reasoning variant in most specialties, with the largest decline observed in Nephrology, where the score decreased from 0.983 to 0.759 (−0.224); its complete error rate also increased from 0.172 to 0.310 (+0.138). These findings suggest that simply extending the 9 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases reasoning process does not guarantee better outcomes in multi-turn multimodal clinical reasoning. 0.611 0.674 0.457 0.724 0.714 0.513 0.783 0.499 0.484 0.587 0.413 0.638 0.770 0.513 0.733 0.448 -0.127 -0.087 -0.043 -0.086 +0.056 +0.000 -0.050 -0.051 Radiol. Ophth. Onco. Neuro. Neph. Intern. Emerg. Derm. 0.400.500.600.700.80 Diagnostic Accuracy Score Qwen3-VL-4B Qwen3-VL-4B-Thinking 0.413 0.261 0.565 0.241 0.316 0.482 0.333 0.514 0.524 0.391 0.609 0.310 0.286 0.491 0.367 0.563 +0.111 +0.130 +0.043 +0.069 -0.031 +0.009 +0.033 +0.049 0.200.300.400.500.60 Complete Error Rate Qwen3-VL-4B Qwen3-VL-4B-Thinking 0.690 0.761 0.478 0.655 0.816 0.588 0.775 0.540 0.548 0.891 0.478 0.724 0.750 0.654 0.867 0.496 -0.143 +0.130 +0.000 +0.069 -0.066 +0.066 +0.092 -0.044 Radiol. Ophth. Onco. Neuro. Neph. Intern. Emerg. Derm. 0.500.600.700.800.90 Diagnostic Accuracy Score Qwen3-VL-8B Qwen3-VL-8B-Thinking 0.381 0.304 0.609 0.345 0.255 0.439 0.333 0.495 0.492 0.174 0.565 0.310 0.306 0.368 0.350 0.532 +0.111 -0.130 -0.043 -0.034 +0.051 -0.070 +0.017 +0.037 0.100.200.300.400.500.60 Complete Error Rate Qwen3-VL-8B Qwen3-VL-8B-Thinking 0.738 0.935 0.739 0.983 1.010 0.658 0.942 0.644 0.762 0.848 0.609 0.759 0.969 0.614 0.883 0.581 +0.024 -0.087 -0.130 -0.224 -0.041 -0.044 -0.058 -0.063 Radiol. Ophth. Onco. Neuro. Neph. Intern. Emerg. Derm. 0.600.700.800.901.00 Diagnostic Accuracy Score Qwen3-VL-32B Qwen3-VL-32B-Thinking 0.397 0.174 0.435 0.172 0.173 0.439 0.267 0.436 0.413 0.217 0.478 0.310 0.173 0.421 0.350 0.476 +0.016 +0.043 +0.043 +0.138 +0.000 -0.018 +0.083 +0.040 0.200.300.400.50 Complete Error Rate Qwen3-VL-32B Qwen3-VL-32B-Thinking Fig. 5| Comparison between non-reasoning and reasoning MLLMs across specialties on ClinMM-Bench. Diagnostic accuracy scores and complete error rates of Qwen3-VL models at the 4B, 8B, and 32B scales under non-reasoning and reasoning settings across eight specialties. The dashed line indicates the mean value across the eight specialties, and the shaded band represents the 95% confidence interval estimated using 100,000 bootstrap resamples. 2.3. Diagnostic Reasoning Quality 2.3.1. Overall Diagnostic Reasoning Quality of 15 MLLMs on ClinMM-Bench Using atomic fact decomposition, we evaluated the overall quality of diagnostic reasoning. Model- generated reasoning was decomposed into atomic clinical facts, enabling assessment of three com- plementary dimensions: fact recall, hallucination, and fact density (Fig. 6). 10 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 0.535 0.527 0.458 0.389 0.599 0.579 0.418 0.470 0.381 0.524 0.464 0.486 0.454 0.493 0.468 0.118 0.085 0.179 0.237 0.105 0.111 0.132 0.159 0.173 0.143 0.177 0.157 0.185 0.155 0.173 0.303 0.412 0.348 0.316 0.343 0.375 0.431 0.318 0.353 0.323 0.288 0.295 0.297 0.305 0.289 Fact RecallHallucinationFact Density 0.30.40.50.60.70.10.20.30.20.30.40.5 Claude-4.5-Sonnet Gemini 3 Pro GPT-5-minimal GPT-5-medium Gemma-3-4B Gemma-3-27B LLaMA-4-Scout Qwen3-VL-4B Qwen3-VL-8B Qwen3-VL-32B Qwen3-VL-4B-thinking Qwen3-VL-8B-thinking Qwen3-VL-32B-thinking MedGemma-4B MedGemma-27B Fig. 6| Overall diagnostic reasoning quality of 15 MLLMs on ClinMM-Bench. Diagnostic reasoning quality is evaluated through three metrics: fact recall, hallucination, and fact density. Higher values indicate better performance for fact recall and fact density, whereas lower values indicate better performance for hallucina- tion. Open circles indicate the overall mean values across all cases, and filled dots indicate the mean values for each specialty. Overall, proprietary models outperformed open-weight models in fact recall and hallucination con- trol. GPT-5-medium achieved the highest fact recall (0.599), followed by GPT-5-minimal (0.579), Claude-4.5-Sonnet (0.535), and Gemini 3 Pro (0.527). For hallucination, Gemini 3 Pro performed best, achieving the lowest hallucination score (0.085). In contrast, open-weight models generally showed lower fact recall and higher hallucination scores. Among open-weight models, Qwen3-VL- 32B achieved the highest fact recall (0.524), whereas LLaMA-4-Scout achieved the lowest hallucina- tion score (0.132). For fact density, the results were not fully consistent with fact recall and hallucination. LLaMA-4- Scout achieved the highest fact density among all models (0.431). Among proprietary models, Gem- ini 3 Pro and GPT-5-minimal showed relatively high fact density, at 0.412 and 0.375, respectively. Among open-weight models, in addition to LLaMA-4-Scout, MedGemma-4B (0.353) and Gemma-3- 27B (0.348) also achieved relatively high fact density. Together, these results show that diagnostic reasoning quality is multidimensional and cannot be captured by a single score. 2.3.2. Variations in Fact Recall Scores across Specialties MLLMs showed notable differences in fact recall scores across specialties (Fig. 7). Neurology achieved the highest overall fact recall score (0.575 [95% CI: 0.550, 0.599]), followed by Ophthalmology 11 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases (0.551 [95% CI: 0.516, 0.584]) and Oncology (0.495 [95% CI: 0.470, 0.520]). Intermediate perfor- mance was observed in Radiology (0.475 [95% CI: 0.465, 0.485]) and Emergency Medicine (0.473 [95% CI: 0.417, 0.529]), whereas Nephrology (0.435 [95% CI: 0.381, 0.489]), Internal Medicine (0.398 [95% CI: 0.352, 0.444]), and Dermatology (0.398 [95% CI: 0.354, 0.441]) showed lower fact recall scores. These results are broadly consistent with the diagnostic accuracy analysis and suggest that certain specialties require evidence types or levels of reasoning complexity that are more difficult for current models to capture. Further details are provided in Supplementary Information B. 0.398 0.473 0.398 0.435 0.575 0.495 0.551 0.475 MedGemma-27B MedGemma-4B Qwen3-VL-32B-thinking Qwen3-VL-8B-thinking Qwen3-VL-4B-thinking Qwen3-VL-32B Qwen3-VL-8B Qwen3-VL-4B LLaMA-4-Scout Gemma-3-27B Gemma-3-4B GPT-5-medium GPT-5-minimal Gemini 3 Pro Claude-4.5-Sonnet Overall Derm.Emerg.Intern.Neph.Neuro.Onco.Ophth.Radiol. Mean Score 0.3 0.4 0.5 0.6 Fig. 7 | Fact recall scores across specialties on ClinMM-Bench. Values indicate mean fact recall scores, with 95% confidence intervals estimated using 100,000 bootstrap resamples. In the bubble plot, both circle size and color encode the mean diagnostic accuracy score, with larger and warmer-colored circles indicating higher scores. 12 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 2.3.3. Effects of Medical Specialization and Reasoning Setting on Diagnostic Reasoning Quality We further compared the effects of medical specialization and the reasoning setting on diagnostic rea- soning quality (Fig. 8). In the comparison between Gemma and MedGemma, the medical versions had a limited effect on fact recall: MedGemma-4B showed a slight reduction in fact recall compared with Gemma-3-4B (−0.008), whereas MedGemma-27B showed a small improvement over Gemma- 3-27B (+0.012). In contrast, medical specialization reduced hallucination scores, with delta values of−0.064 and−0.020 at the 4B and 27B scales, respectively. For fact density, the effects were incon- sistent: MedGemma-4B improved over Gemma-3-4B (+0.037), whereas MedGemma-27B decreased compared with Gemma-3-27B (−0.030). a. Comparison between General and Medical MLLMs +0.012 -0.008 -0.020 -0.064 -0.030 +0.037 Fact RecallHallucinationFact Density -0.10-0.050.000.050.10-0.10-0.050.000.050.10-0.10-0.050.000.050.10 Gemma-3-27B Gemma-3-4B Delta Metric Value (Medical Model - General Model) b. Comparison between Non-Reasoning and Reasoning MLLMs -0.060 -0.032 -0.025 +0.034 +0.028 +0.018 -0.034 +0.003 -0.016 Fact RecallHallucinationFact Density -0.10-0.050.000.050.10-0.10-0.050.000.050.10-0.10-0.050.000.050.10 Qwen3-VL-32B Qwen3-VL-8B Qwen3-VL-4B Delta Metric Value (Reasoning Model - Non-Reasoning Model) Fig. 8 | Effects of medical specialization and reasoning setting on diagnostic reasoning quality. a, Com- parison between Gemma and MedGemma models. Delta values of fact recall, hallucination, and fact density between Gemma models and their corresponding medical versions at the 4B and 27B scales. b, Comparison between non-reasoning and reasoning MLLMs. Delta values of fact recall, hallucination, and fact density between Qwen3-VL models and their corresponding reasoning versions at the 4B, 8B, and 32B scales. Posi- tive values indicate better performance for fact recall and fact density, whereas negative values indicate lower hallucination scores, corresponding to better performance. Error bars around each point indicate the 95% confidence intervals estimated using 100,000 bootstrap resamples. In the Qwen3-VL series, the reasoning setting did not improve diagnostic reasoning quality. Com- pared with non-reasoning versions, the reasoning versions showed lower fact recall at all model scales (−0.032,−0.025, and−0.060) and did not reduce hallucination scores (+0.028,+0.018, and +0.034). Fact density improved slightly only for Qwen3-VL-4B-thinking (+0.003) and decreased at larger scales (−0.016 and−0.034 for 8B and 32B reasoning versions, respectively). 13 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 2.4. Error Analysis of Multi-Turn Multimodal Clinical Diagnostic Reasoning To investigate the limitations of MLLMs, we analyzed diagnostic failure cases and categorized them into five representative error types: information synthesis failure, knowledge mapping error, percep- tion error, premature closure, and visual hallucination. First, information synthesis failures occurred when models could not synthesize information across turns and modalities. For instance, in a case of smoking-related organizing pneumonia, the model over-relied on early imaging and cytologic findings suggestive of malignancy, overlooked subsequent lesion regression and biopsy findings, and incorrectly diagnosed invasive mucinous adenocarcinoma. Second, knowledge mapping errors occurred when models recognized key facts but mapped them to an incorrect disease mechanism. In a case of tickborne encephalitis, the model recognized fever, cerebrospinal fluid pleocytosis, MRI abnormalities, and positive intrathecal antibodies but incorrectly diagnosed acute disseminated encephalomyelitis. Third, perception errors occurred when models incorrectly interpreted the key imaging findings. In a case of Chilaiditi syndrome, the model misin- terpreted colonic gas between the liver and diaphragm as gas within the gallbladder and diagnosed emphysematous cholecystitis. Fourth, premature closure occurred when models committed too early to an initial hypothesis and failed to revise it after additional information emerged. In a case of ony- chomatricoma, the model continued to favor subungual melanoma despite subsequent dermoscopic, MRI, and pathological evidence supporting a benign diagnosis. Finally, visual hallucinations oc- curred when models introduced nonexistent imaging findings. In a case of cerebral cavernoma with hematoma, the model hallucinated cystic lesions, an eccentric scolex, and a dot sign, leading to an incorrect diagnosis of neurocysticercosis. Overall, these errors indicate that diagnostic failures in MLLMs arise from capability deficits at multiple levels. 3. Discussion ClinMM-Bench addresses a critical gap in the evaluation of MLLMs in medicine. Existing benchmarks largely rely on single-turn or isolated tasks, whereas clinical diagnosis requires synthesizing infor- mation, revising hypotheses, and reasoning under uncertainty. By converting real-world case reports into multi-turn, multimodal diagnostic dialogues, ClinMM-Bench enables models to be evaluated in a setting that more closely aligns with clinical diagnostic workflows. Importantly, the benchmark is designed to evaluate model performance on diagnostically challenging cases involving relatively uncommon clinical presentations, rather than on common diseases seen in routine clinical practice. Our results show that current MLLMs still have significant limitations in clinical diagnostic reason- ing. Proprietary models outperformed open-weight models, and most models were able to generate partially correct diagnoses. However, the proportion of completely correct diagnoses remained lim- ited; even the best-performing model achieved completely correct diagnoses in only approximately one-third of cases. This finding is clinically important: partial recognition may be useful for triage or differential diagnosis, but reliable diagnostic decision-making requires accurate synthesis of all relevant information 1,6 . Specialty-level differences further indicate that diagnostic performance of MLLMs is uneven and jointly shaped by the clinical domain, evidence type, and reasoning complex- ity 26 . The analysis of reasoning quality provides additional insight beyond diagnostic accuracy. Models often omit key evidence, introduce unsupported statements, or generate inefficient reasoning. These reasoning issues are particularly relevant for clinical deployment, where a correct diagnosis without transparent and faithful reasoning may be difficult to trust, and an incorrect diagnosis supported by plausible but hallucinated reasoning may be harmful. 14 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases In addition, our comparisons reveal important implications for model development. Larger open- weight models generally achieved higher diagnostic accuracy and better fact coverage, indicating that scale remains beneficial for complex multi-turn multimodal diagnosis. Medical specialization improved some aspects of performance, particularly in smaller models and in hallucination control, but did not consistently improve performance in larger models. Notably, the reasoning setting gen- erally failed to reliably improve diagnostic accuracy or reasoning quality. This suggests that future progress should not rely solely on longer reasoning traces 27,28 . Instead, models need better cross- turn memory, stronger visual grounding, improved clinical knowledge mapping, and mechanisms for revising hypotheses when additional information contradicts earlier assumptions. The error analysis further revealed that diagnostic failures do not arise solely from insufficient med- ical knowledge, but are distributed across multiple stages. Perception errors and visual hallucina- tions indicate that models’ basic image understanding is not yet reliable, while knowledge mapping errors suggest that even when models identify key facts, they may fail to accurately connect them to the corresponding disease mechanisms. Premature closure and information synthesis failures reflect limitations in models’ ability to dynamically update diagnostic reasoning across multiple turns: mod- els may over-rely on early clues or initial hypotheses and fail to sufficiently integrate subsequently disclosed information. These findings highlight the need for benchmarks that can decompose per- formance into clinically meaningful capabilities. ClinMM-Bench provides such a framework and can support future work on model training, evaluation, and safety monitoring. This study has several limitations. First, ClinMM-Bench primarily consists of challenging published case reports, meaning that its diagnostic difficulty may differ from that of more common cases en- countered in routine clinical practice 29 . Therefore, this benchmark should not be taken as a direct estimate of performance on common clinical diagnoses. Second, ClinMM-Bench adopts a passive multi-turn information disclosure format, whereas in real clinical settings, physicians actively elicit history, select further investigations, and dynamically determine the next steps based on existing hypotheses 13 . Third, although we employed multi-stage data curation, quality control, expert vali- dation, dual-LLM consensus evaluation, and atomic fact decomposition, potential biases may still be introduced during data curation and automated evaluation. Overall, ClinMM-Bench establishes a large-scale, multi-turn, multimodal benchmark for evaluating clinical diagnostic reasoning. Our findings show that current models can recognize partial diagnostic patterns but remain unable to reliably achieve completely correct diagnoses or maintain reasoning fidelity. 4. Methods 4.1. Data Curation ClinMM-Bench was constructed from clinical case reports published in PMCOA before September 1, 2025. The search strategy is provided in Supplementary Information C. We focused on eight specialties: Dermatology, Emergency Medicine, Internal Medicine, Nephrology, Neurology, Oncol- ogy, Ophthalmology, and Radiology. The data curation pipeline consisted of six stages, including Data Collection and Extraction, Data Inspection, Data Validation, Data Conversion, Data Quality Control, and Expert Validation. Detailed stage-wise flow information is provided in Supplementary Information D. 15 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 4.1.1. Data Collection and Extraction We established a one-to-one mapping between the manually curated PMID list and PMCOA, en- abling the batch downloading and collection of all original case reports. On this basis, we performed extraction to obtain the XML-formatted text and corresponding medical imaging resources. 4.1.2. Data Inspection We excluded case reports that contained videos, lacked medical images, or had low-resolution medi- cal images below336×336 pixels. This ensured that retained case reports contained sufficient image information for MLLMs to perform multimodal diagnostic reasoning. 4.1.3. Data Validation We employed a dual-LLM consensus mechanism with GPT-4.1 and Claude-4.0-Sonnet to determine whether each case report was suitable for the diagnostic reasoning task. A valid case report was required to describe a complete diagnostic process for a single patient, including essential clinical information, diagnostic medical images, a clearly defined final diagnosis, and a traceable diagnostic reasoning process. A case report was retained only when both LLMs classified it as diagnostically suitable. Case reports focusing solely on treatment effects, drug reactions, literature reviews, or those lacking diagnostic reasoning processes were excluded. The detailed prompt is provided in Supplementary Information E. 4.1.4. Data Conversion We parsed the XML texts of validated case reports and converted them into a standardized JSON for- mat using GPT-4.1. Extracted key information included demographic information, chief complaint, history of present illness, lab results, physical examination, medical images, diagnosis, and reasoning process. To avoid information leakage, diagnoses and related findings were removed from the case presentation. Therefore, models were required to derive the diagnosis through step-by-step analysis of multi-source information. The detailed prompt is provided in Supplementary Information F. 4.1.5. Quality Control We used Claude-4.0-Sonnet to compare each structured case against its original XML text and as- signed a fidelity score on a 1–5 scale (5 = all information perfectly extracted and preserved; 4 = minor discrepancies without affecting meaning; 3 = some information loss with key elements pre- served; 2 = significant information loss or errors; 1 = critical information loss or errors). Only cases scoring≥3 were retained, thereby preventing substantial loss of critical diagnostic information. The detailed prompt is provided in Supplementary Information G. 4.1.6. Expert Validation Experienced medical experts (M.W., N.L. (Nicolás Lescano), N.P., E.P.) reviewed a random sample of the retained cases to verify their clinical authenticity, the completeness of diagnostic reasoning, and the effectiveness of information isolation, thereby further ensuring that ClinMM-Bench aligns with the requirements of diagnostic reasoning evaluation. 16 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases 4.2. MLLM Evaluation We selected a diverse set of 15 representative MLLMs to evaluate their performance on ClinMM- Bench, including proprietary general models, open-weight general models, open-weight medical models, and reasoning and non-reasoning variants. Proprietary models included Claude-4.5-Sonnet, Gemini 3 Pro, and GPT-5, with the latter evaluated under both the “minimal” and “medium” levels of the “reasoning effort” setting. Open-weight general models comprised the Gemma series (Gemma-3- 4B/27B), LLaMA-4-Scout, and the Qwen3-VL series (Qwen3-VL-4B/8B/32B and their corresponding reasoning variants). Open-weight medical models included the MedGemma series (MedGemma- 4B/27B). 4.3. Two-Level Evaluation Framework We designed a two-level evaluation framework. The first level assesses diagnostic accuracy, while the second level evaluates the quality of the diagnostic reasoning process. 4.3.1. Diagnostic Accuracy Evaluation For diagnostic accuracy evaluation, we adopted a dual-LLM judgment strategy, in which two LLMs were used as independent evaluators. Specifically, for each case, we provided both the MLLM- predicted diagnosis and the ground-truth diagnosis to the judge LLMs (i.e., GPT-5-medium and Claude-4.5-Sonnet) and asked each judge LLM to assess accuracy based on established clinical diag- nostic criteria. The scoring rubric consisted of three levels: 0 for a completely incorrect or irrelevant diagnosis; 1 for a partially correct diagnosis; and 2 for a completely correct diagnosis. Additionally, the judge LLMs were required to provide explicit reasoning to justify their assigned scores. We then calculated the average of the scores assigned by the two judge LLMs to obtain a consensus score, thereby reducing single-LLM judgment bias and allowing the recognition of clinically equivalent di- agnostic expressions with different terminology. The detailed prompt is provided in Supplementary Information H. 4.3.2. Diagnostic Reasoning Quality Evaluation The quality of reasoning was evaluated using atomic fact decomposition. Both MLLM-generated rea- soning and reference reasoning were decomposed into a series of atomic clinical facts, with each fact representing a minimal, independently verifiable clinical statement. We then defined three metrics: (1) Fact Recall measures reasoning completeness and is calculated as the proportion of reference atomic facts correctly identified by the model among all reference atomic facts. A higher fact recall indicates that the model successfully captures the critical clinical information required for diagnosis; (2) Hallucination measures reasoning reliability and is defined as the proportion of model-generated atomic facts that are unsupported by the reference. A lower hallucination score reflects stronger ad- herence to actual clinical evidence and a reduced tendency to fabricate unsupported information; and (3) Fact Density measures reasoning efficiency and is quantified as the proportion of valid atomic facts among all atomic facts produced by the model. A higher fact density indicates that the model conveys diagnostic reasoning in a concise and clinically informative manner, thereby mitigating met- ric dilution from excessively long outputs. The detailed prompt is provided in Supplementary In- formation I. Data Availability This study used publicly available data from https://pmc.ncbi.nlm.nih.gov/. 17 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Code Availability The code used for this study is available at https://github.com/ruiyang-medinfo/ClinMM. Acknowledgements This work was supported by the Duke-NUS Signature Research Programme funded by the Ministry of Health, Singapore. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of the Ministry of Health. This study was supported by the U.S. National Institutes of Health grants R01LM014573 and R01LM014344. Author Contributions Conceptualization: R.Y., N.L. (Nan Liu), Y.P.; Methodology: R.Y.; Software: R.Y.; Data Curation: R.Y.; Investigation: R.Y., W.X., Z.B.; Validation: M.W., N.L. (Nicolás Lescano), N.P., E.P.; Formal Analysis: R.Y.; Visualization: R.Y.; Writing – Original Draft: R.Y.; Writing – Review & Editing: All authors; Supervision: N.L. (Nan Liu), Y.P.; Project Administration: N.L. (Nan Liu), Y.P.; Funding Acquisition: N.L. (Nan Liu), Y.P. Competing Interests The authors declare no competing interests. References [1] McDuff, D. et al. Towards accurate differential diagnosis with large language models. Nature 642, 451–457 (2025). [2] Scott, I. A. Errors in clinical reasoning: causes and remedial strategies. BMJ 338, b1860 (2009). [3] McMahon, G. T., Solomon, C. G., Ross, J. J., Loscalzo, J. & Campion, E. W. Interactive medical cases – a NewJournalFeature. N. Engl. J. Med. 361, 1113–1113 (2009). [4] Meyer, A. N. D., Payne, V. L., Meeks, D. W., Rao, R. & Singh, H. Physicians’ diagnostic accuracy, confidence, and resource requests: a vignette study. JAMA Intern Med 173, 1952–1958 (2013). [5] Centor, R. M., Geha, R. & Manesh, R. The pursuit of diagnostic excellence. JAMA Netw Open 2, e1918040 (2019). [6] Committee on Diagnostic Error in Health Care, Board on Health Care Services, Institute of Medicine & The National Academies of Sciences, Engineering, and Medicine. Improving Diag- nosis in Health Care (National Academies Press (US), Washington (DC), 2015). [7] Yang, R. et al. Large language models in health care: Development, applications, and chal- lenges. Health Care Sci 2, 255–263 (2023). [8] Fahrner, L. J., Chen, E., Topol, E. & Rajpurkar, P. The generative era of medical AI. Cell 188, 3648–3660 (2025). [9] Yang, R. et al. Retrieval-augmented generation for generative artificial intelligence in health care. Npj Health Syst. 2 (2025). 18 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases [10] Yang, R. et al. Retrieval-augmented generation in medicine: A scoping review of technical implementations, clinical applications, and ethical considerations. Cell Reports Medicine (2026). URL https://doi.org/10.1016/j.xcrm.2026.102927. [11] Tu, T. et al. Towards generalist biomedical AI. NEJM AI 1 (2024). [12] Saab, K. et al. Advancing conversational diagnostic AI with multimodal reasoning. Nat Med 32, 1726–1736 (2026). [13] Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat Med 31, 77–86 (2025). [14] Jin, D. et al. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. arXiv [cs.CL] (2020). 2009.13081. [15] Wu, K. et al. MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports. arXiv [cs.CL] (2025). 2505.11733. [16] Zhang, X. et al. Development of a large-scale medical visual question-answering dataset. Com- mun Med (Lond) 4, 277 (2024). [17] McCoy, L. G. et al. Assessment of large language models in clinical reasoning: A novel bench- marking study. NEJM AI 2 (2025). [18] Tanno, R. et al. Collaboration between clinicians and vision-language models in radiology report generation. Nat Med 31, 599–608 (2025). [19] Omar, M. et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond) 5, 330 (2025). [20] Yang, X. et al. Multiple large language models versus experienced physicians in diagnosing challenging cases with gastrointestinal symptoms. NPJ Digit Med 8, 85 (2025). [21] Nori, H. et al. Sequential diagnosis with language models. arXiv [cs.CL] (2025). 2506.22405. [22] Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat Med 30, 2613–2622 (2024). [23] Ke, Y. et al. Mitigating cognitive biases in clinical decision-making through multi-agent conver- sations using large language models: Simulation study. J Med Internet Res 26, e59439 (2024). [24] Mahajan, A., Obermeyer, Z., Daneshjou, R., Lester, J. & Powell, D. Cognitive bias in clinical large language models. NPJ Digit Med 8, 428 (2025). [25] Qiu, P. et al. Quantifying the reasoning abilities of LLMs on clinical cases. Nat Commun 16, 9799 (2025). [26] Zhu, Y. et al. DiagnosisArena: Benchmarking diagnostic reasoning for large language models. arXiv [cs.CL] (2025). 2505.14107. [27] Hong, J. et al. Benchmarking the thinking mode of multimodal large language models in clinical tasks. arXiv [cs.CL] (2025). 2511.03328. [28] Kancheti, S. S., Kanade, A. S., Balasubramanian, V. N. & Ganu, T. Chain-of-thought degrades visual spatial reasoning capabilities of multimodal LLMs. arXiv [cs.CV] (2026). 2604.16060. [29] Vandenbroucke, J. P. In defense of case reports and case series. Ann Intern Med 134, 330–334 (2001). 19 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information A Diagnostic Accuracy across Medical Specialties Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 1.056 [0.865, 1.246] 1.065 [0.761, 1.370] 1.109 [0.783, 1.435] 1.052 [0.862, 1.259] 1.327 [1.199, 1.454] 1.039 [0.899, 1.175] 1.317 [1.133, 1.492] 0.918 [0.859, 0.977] Gemini 3 Pro 1.127 [0.952, 1.302] 1.239 [1.022, 1.457] 1.065 [0.739, 1.370] 1.069 [0.828, 1.310] 1.168 [1.051, 1.286] 0.965 [0.851, 1.079] 1.217 [1.050, 1.383] 0.998 [0.946, 1.049] GPT-5-minimal 1.135 [0.944, 1.317] 1.196 [0.935, 1.457] 1.022 [0.717, 1.326] 1.103 [0.879, 1.328] 1.311 [1.173, 1.444] 1.048 [0.917, 1.180] 1.250 [1.067, 1.425] 0.951 [0.895, 1.007] GPT-5-medium 1.056 [0.849, 1.262] 1.283 [1.000, 1.543] 1.196 [0.870, 1.522] 0.966 [0.724, 1.207] 1.383 [1.245, 1.515] 1.162 [1.022, 1.303] 1.258 [1.075, 1.442] 1.099 [1.043, 1.155] Gemma-3-4B 0.532 [0.389, 0.675] 0.413 [0.239, 0.609] 0.391 [0.196, 0.587] 0.466 [0.293, 0.638] 0.577 [0.464, 0.689] 0.474 [0.373, 0.575] 0.658 [0.492, 0.833] 0.380 [0.343, 0.418] Gemma-3-27B 0.659 [0.500, 0.817] 0.652 [0.413, 0.891] 0.457 [0.217, 0.717] 0.707 [0.517, 0.879] 0.893 [0.781, 1.005] 0.535 [0.430, 0.640] 0.775 [0.608, 0.942] 0.532 [0.487, 0.577] LLaMA-4-Scout 0.802 [0.635, 0.976] 0.804 [0.587, 1.022] 0.500 [0.283, 0.739] 0.879 [0.672, 1.086] 1.010 [0.898, 1.122] 0.671 [0.557, 0.785] 0.967 [0.783, 1.150] 0.644 [0.597, 0.690] Qwen3-VL-4B 0.611 [0.468, 0.762] 0.674 [0.478, 0.891] 0.457 [0.217, 0.717] 0.724 [0.552, 0.897] 0.714 [0.602, 0.832] 0.513 [0.417, 0.614] 0.783 [0.617, 0.958] 0.499 [0.456, 0.543] Qwen3-VL-8B 0.690 [0.540, 0.849] 0.761 [0.522, 1.022] 0.478 [0.217, 0.783] 0.655 [0.466, 0.845] 0.816 [0.704, 0.929] 0.588 [0.482, 0.693] 0.775 [0.617, 0.942] 0.540 [0.494, 0.586] Qwen3-VL-32B 0.738 [0.571, 0.913] 0.935 [0.696, 1.174] 0.739 [0.457, 1.043] 0.983 [0.776, 1.190] 1.010 [0.893, 1.128] 0.658 [0.535, 0.785] 0.942 [0.758, 1.125] 0.644 [0.596, 0.694] Qwen3-VL-4B-thinking 0.484 [0.349, 0.627] 0.587 [0.370, 0.783] 0.413 [0.196, 0.652] 0.638 [0.466, 0.793] 0.770 [0.653, 0.888] 0.513 [0.412, 0.614] 0.733 [0.567, 0.908] 0.448 [0.406, 0.491] Qwen3-VL-8B-thinking 0.548 [0.405, 0.698] 0.891 [0.674, 1.109] 0.478 [0.261, 0.739] 0.724 [0.534, 0.914] 0.750 [0.633, 0.867] 0.654 [0.548, 0.763] 0.867 [0.683, 1.058] 0.496 [0.451, 0.541] Qwen3-VL-32B-thinking 0.762 [0.587, 0.944] 0.848 [0.609, 1.087] 0.609 [0.348, 0.870] 0.759 [0.552, 0.966] 0.969 [0.852, 1.087] 0.614 [0.504, 0.724] 0.883 [0.683, 1.083] 0.581 [0.533, 0.629] MedGemma-4B 0.484 [0.357, 0.611] 0.587 [0.391, 0.783] 0.478 [0.283, 0.674] 0.707 [0.534, 0.862] 0.679 [0.582, 0.776] 0.461 [0.373, 0.553] 0.683 [0.533, 0.842] 0.431 [0.391, 0.471] MedGemma-27B 0.738 [0.587, 0.897] 0.717 [0.457, 0.978] 0.587 [0.326, 0.848] 0.690 [0.500, 0.879] 0.852 [0.740, 0.964] 0.596 [0.496, 0.702] 0.858 [0.683, 1.033] 0.531 [0.485, 0.577] Overall 0.761 [0.639, 0.884] 0.843 [0.677, 1.010] 0.665 [0.490, 0.851] 0.808 [0.666, 0.947] 0.949 [0.880, 1.018] 0.699 [0.626, 0.772] 0.931 [0.797, 1.067] 0.646 [0.613, 0.680] Supplementary Table A1 | Diagnostic accuracy scores across medical specialties on ClinMM-Bench. Val- ues indicate mean diagnostic accuracy scores, with 95% confidence intervals estimated using 100,000 boot- strap resamples. 20 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 0.317 [0.206, 0.429] 0.304 [0.130, 0.478] 0.391 [0.217, 0.609] 0.172 [0.034, 0.310] 0.398 [0.306, 0.500] 0.281 [0.202, 0.368] 0.433 [0.317, 0.567] 0.256 [0.224, 0.290] Gemini 3 Pro 0.270 [0.159, 0.381] 0.261 [0.087, 0.435] 0.304 [0.130, 0.478] 0.241 [0.103, 0.414] 0.235 [0.153, 0.316] 0.167 [0.105, 0.237] 0.317 [0.200, 0.433] 0.215 [0.184, 0.246] GPT-5-minimal 0.333 [0.222, 0.444] 0.304 [0.130, 0.478] 0.304 [0.130, 0.478] 0.241 [0.103, 0.414] 0.408 [0.316, 0.510] 0.263 [0.184, 0.342] 0.367 [0.250, 0.483] 0.240 [0.208, 0.272] GPT-5-medium 0.349 [0.238, 0.476] 0.391 [0.217, 0.609] 0.435 [0.217, 0.652] 0.207 [0.069, 0.345] 0.480 [0.378, 0.582] 0.360 [0.272, 0.447] 0.400 [0.283, 0.533] 0.309 [0.275, 0.345] Gemma-3-4B 0.048 [0.000, 0.111] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.051 [0.010, 0.102] 0.035 [0.009, 0.070] 0.100 [0.033, 0.183] 0.013 [0.006, 0.022] Gemma-3-27B 0.095 [0.032, 0.175] 0.043 [0.000, 0.130] 0.043 [0.000, 0.130] 0.000 [0.000, 0.000] 0.092 [0.041, 0.153] 0.035 [0.009, 0.070] 0.117 [0.050, 0.200] 0.052 [0.035, 0.069] LLaMA-4-Scout 0.143 [0.063, 0.238] 0.043 [0.000, 0.130] 0.043 [0.000, 0.130] 0.103 [0.000, 0.207] 0.163 [0.092, 0.235] 0.061 [0.018, 0.105] 0.200 [0.100, 0.300] 0.069 [0.050, 0.088] Qwen3-VL-4B 0.063 [0.016, 0.127] 0.043 [0.000, 0.130] 0.043 [0.000, 0.130] 0.034 [0.000, 0.103] 0.061 [0.020, 0.112] 0.018 [0.000, 0.044] 0.100 [0.033, 0.183] 0.044 [0.029, 0.060] Qwen3-VL-8B 0.079 [0.016, 0.159] 0.087 [0.000, 0.217] 0.130 [0.000, 0.261] 0.034 [0.000, 0.103] 0.061 [0.020, 0.112] 0.044 [0.009, 0.088] 0.100 [0.033, 0.183] 0.063 [0.046, 0.082] Qwen3-VL-32B 0.127 [0.048, 0.206] 0.130 [0.000, 0.261] 0.130 [0.000, 0.261] 0.138 [0.034, 0.276] 0.163 [0.092, 0.235] 0.114 [0.061, 0.175] 0.233 [0.133, 0.350] 0.090 [0.069, 0.112] Qwen3-VL-4B-thinking 0.032 [0.000, 0.079] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.082 [0.031, 0.143] 0.026 [0.000, 0.061] 0.117 [0.050, 0.200] 0.034 [0.021, 0.049] Qwen3-VL-8B-thinking 0.063 [0.016, 0.127] 0.087 [0.000, 0.217] 0.043 [0.000, 0.130] 0.034 [0.000, 0.103] 0.071 [0.020, 0.122] 0.061 [0.018, 0.105] 0.183 [0.083, 0.283] 0.053 [0.037, 0.071] Qwen3-VL-32B-thinking 0.159 [0.079, 0.254] 0.130 [0.000, 0.261] 0.087 [0.000, 0.217] 0.069 [0.000, 0.172] 0.143 [0.082, 0.214] 0.053 [0.018, 0.096] 0.250 [0.150, 0.367] 0.072 [0.053, 0.093] MedGemma-4B 0.016 [0.000, 0.048] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.000 [0.000, 0.000] 0.010 [0.000, 0.031] 0.009 [0.000, 0.026] 0.083 [0.017, 0.167] 0.028 [0.016, 0.041] MedGemma-27B 0.111 [0.048, 0.190] 0.087 [0.000, 0.217] 0.087 [0.000, 0.217] 0.034 [0.000, 0.103] 0.082 [0.031, 0.143] 0.035 [0.009, 0.070] 0.133 [0.050, 0.217] 0.059 [0.041, 0.077] Overall 0.147 [0.098, 0.200] 0.128 [0.064, 0.203] 0.136 [0.070, 0.212] 0.087 [0.037, 0.149] 0.167 [0.131, 0.205] 0.104 [0.080, 0.130] 0.209 [0.143, 0.280] 0.107 [0.094, 0.120] Supplementary Table A2 | Proportions of completely correct diagnoses across medical specialties on ClinMM-Bench. Values indicate proportions of completely correct diagnoses, with 95% confidence intervals estimated using 100,000 bootstrap resamples. 21 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 0.413 [0.286, 0.540] 0.478 [0.261, 0.696] 0.348 [0.174, 0.565] 0.759 [0.586, 0.897] 0.520 [0.418, 0.622] 0.465 [0.377, 0.553] 0.450 [0.333, 0.583] 0.408 [0.371, 0.445] Gemini 3 Pro 0.524 [0.397, 0.651] 0.696 [0.522, 0.870] 0.435 [0.217, 0.652] 0.552 [0.379, 0.724] 0.653 [0.561, 0.745] 0.632 [0.544, 0.719] 0.550 [0.417, 0.667] 0.560 [0.523, 0.596] GPT-5-minimal 0.444 [0.317, 0.571] 0.609 [0.391, 0.783] 0.435 [0.217, 0.652] 0.621 [0.448, 0.793] 0.469 [0.367, 0.571] 0.518 [0.430, 0.605] 0.467 [0.333, 0.600] 0.479 [0.440, 0.517] GPT-5-medium 0.333 [0.222, 0.460] 0.522 [0.304, 0.739] 0.348 [0.174, 0.565] 0.552 [0.379, 0.724] 0.418 [0.327, 0.520] 0.430 [0.342, 0.518] 0.450 [0.333, 0.583] 0.468 [0.432, 0.507] Gemma-3-4B 0.460 [0.333, 0.587] 0.478 [0.261, 0.696] 0.435 [0.217, 0.652] 0.483 [0.310, 0.655] 0.510 [0.408, 0.612] 0.439 [0.351, 0.526] 0.483 [0.367, 0.617] 0.399 [0.362, 0.436] Gemma-3-27B 0.492 [0.365, 0.619] 0.565 [0.348, 0.783] 0.348 [0.174, 0.565] 0.690 [0.517, 0.862] 0.704 [0.612, 0.796] 0.482 [0.395, 0.570] 0.550 [0.417, 0.667] 0.454 [0.417, 0.492] LLaMA-4-Scout 0.524 [0.397, 0.651] 0.696 [0.522, 0.870] 0.435 [0.217, 0.652] 0.690 [0.517, 0.862] 0.704 [0.612, 0.796] 0.535 [0.439, 0.623] 0.517 [0.383, 0.650] 0.518 [0.482, 0.555] Qwen3-VL-4B 0.524 [0.397, 0.651] 0.696 [0.522, 0.870] 0.391 [0.217, 0.609] 0.724 [0.552, 0.862] 0.622 [0.531, 0.714] 0.500 [0.412, 0.588] 0.567 [0.433, 0.683] 0.442 [0.405, 0.479] Qwen3-VL-8B 0.540 [0.413, 0.667] 0.609 [0.391, 0.783] 0.261 [0.087, 0.435] 0.621 [0.448, 0.793] 0.684 [0.592, 0.776] 0.518 [0.430, 0.605] 0.567 [0.433, 0.683] 0.442 [0.405, 0.479] Qwen3-VL-32B 0.476 [0.349, 0.603] 0.696 [0.522, 0.870] 0.435 [0.217, 0.652] 0.690 [0.517, 0.862] 0.663 [0.571, 0.755] 0.447 [0.360, 0.535] 0.500 [0.367, 0.633] 0.474 [0.437, 0.513] Qwen3-VL-4B-thinking 0.444 [0.317, 0.571] 0.609 [0.391, 0.783] 0.391 [0.217, 0.609] 0.690 [0.517, 0.862] 0.633 [0.541, 0.724] 0.482 [0.395, 0.570] 0.517 [0.383, 0.650] 0.404 [0.367, 0.440] Qwen3-VL-8B-thinking 0.444 [0.317, 0.571] 0.739 [0.565, 0.913] 0.391 [0.217, 0.609] 0.655 [0.483, 0.828] 0.622 [0.531, 0.714] 0.570 [0.482, 0.658] 0.467 [0.350, 0.600] 0.415 [0.378, 0.452] Qwen3-VL-32B-thinking 0.429 [0.302, 0.556] 0.652 [0.435, 0.826] 0.435 [0.217, 0.652] 0.621 [0.448, 0.793] 0.684 [0.592, 0.776] 0.526 [0.439, 0.614] 0.400 [0.283, 0.517] 0.452 [0.415, 0.490] MedGemma-4B 0.508 [0.381, 0.635] 0.609 [0.391, 0.783] 0.522 [0.304, 0.739] 0.724 [0.552, 0.862] 0.704 [0.612, 0.796] 0.491 [0.395, 0.579] 0.550 [0.417, 0.667] 0.414 [0.377, 0.451] MedGemma-27B 0.556 [0.429, 0.683] 0.522 [0.304, 0.739] 0.435 [0.217, 0.652] 0.655 [0.483, 0.828] 0.704 [0.612, 0.796] 0.544 [0.456, 0.632] 0.533 [0.400, 0.667] 0.430 [0.393, 0.467] Overall 0.474 [0.393, 0.556] 0.612 [0.501, 0.719] 0.403 [0.287, 0.528] 0.648 [0.538, 0.752] 0.620 [0.569, 0.670] 0.505 [0.449, 0.563] 0.504 [0.429, 0.579] 0.451 [0.427, 0.475] Supplementary Table A3 | Proportions of partially correct diagnoses across medical specialties on ClinMM-Bench. Values indicate proportions of partially correct diagnoses, with 95% confidence intervals estimated using 100,000 bootstrap resamples. 22 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information B Diagnostic Reasoning Quality across Medical Specialties Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 0.479 [0.419, 0.540] 0.590 [0.512, 0.669] 0.490 [0.417, 0.563] 0.503 [0.430, 0.574] 0.639 [0.607, 0.672] 0.578 [0.545, 0.612] 0.611 [0.565, 0.655] 0.512 [0.498, 0.526] Gemini 3 Pro 0.475 [0.424, 0.527] 0.544 [0.467, 0.621] 0.465 [0.418, 0.513] 0.515 [0.442, 0.588] 0.597 [0.564, 0.630] 0.528 [0.496, 0.560] 0.585 [0.544, 0.627] 0.519 [0.506, 0.532] GPT-5-minimal 0.480 [0.423, 0.537] 0.540 [0.475, 0.607] 0.495 [0.435, 0.554] 0.553 [0.485, 0.614] 0.678 [0.648, 0.709] 0.590 [0.557, 0.622] 0.659 [0.618, 0.700] 0.570 [0.556, 0.584] GPT-5-medium 0.491 [0.429, 0.552] 0.590 [0.516, 0.664] 0.482 [0.410, 0.555] 0.574 [0.497, 0.650] 0.682 [0.647, 0.716] 0.617 [0.581, 0.652] 0.658 [0.616, 0.701] 0.594 [0.580, 0.609] Gemma-3-4B 0.297 [0.250, 0.347] 0.370 [0.295, 0.446] 0.306 [0.259, 0.351] 0.358 [0.302, 0.415] 0.475 [0.443, 0.507] 0.405 [0.377, 0.434] 0.450 [0.401, 0.499] 0.382 [0.370, 0.393] Gemma-3-27B 0.383 [0.333, 0.435] 0.463 [0.379, 0.549] 0.343 [0.287, 0.398] 0.408 [0.343, 0.479] 0.537 [0.504, 0.571] 0.464 [0.429, 0.499] 0.534 [0.491, 0.578] 0.451 [0.439, 0.464] LLaMA-4-Scout 0.352 [0.305, 0.402] 0.376 [0.316, 0.439] 0.336 [0.284, 0.390] 0.395 [0.327, 0.467] 0.487 [0.459, 0.517] 0.417 [0.385, 0.450] 0.464 [0.416, 0.511] 0.416 [0.404, 0.427] Qwen3-VL-4B 0.383 [0.338, 0.429] 0.467 [0.392, 0.542] 0.393 [0.343, 0.446] 0.417 [0.352, 0.482] 0.569 [0.533, 0.604] 0.503 [0.472, 0.535] 0.543 [0.499, 0.587] 0.483 [0.470, 0.495] Qwen3-VL-8B 0.380 [0.330, 0.431] 0.501 [0.430, 0.575] 0.385 [0.321, 0.451] 0.417 [0.355, 0.481] 0.591 [0.558, 0.624] 0.508 [0.476, 0.541] 0.573 [0.526, 0.618] 0.487 [0.474, 0.500] Qwen3-VL-32B 0.438 [0.388, 0.489] 0.507 [0.432, 0.585] 0.418 [0.337, 0.498] 0.468 [0.404, 0.527] 0.633 [0.599, 0.666] 0.525 [0.490, 0.561] 0.597 [0.550, 0.644] 0.516 [0.503, 0.530] Qwen3-VL-4B-thinking 0.345 [0.299, 0.392] 0.409 [0.342, 0.480] 0.376 [0.314, 0.440] 0.361 [0.316, 0.407] 0.567 [0.533, 0.602] 0.466 [0.437, 0.496] 0.534 [0.489, 0.577] 0.446 [0.434, 0.459] Qwen3-VL-8B-thinking 0.386 [0.336, 0.436] 0.465 [0.393, 0.542] 0.416 [0.366, 0.465] 0.397 [0.343, 0.452] 0.560 [0.527, 0.593] 0.472 [0.436, 0.509] 0.560 [0.517, 0.602] 0.459 [0.447, 0.472] Qwen3-VL-32B-thinking 0.419 [0.363, 0.474] 0.458 [0.382, 0.536] 0.428 [0.358, 0.500] 0.468 [0.393, 0.544] 0.573 [0.542, 0.604] 0.467 [0.433, 0.502] 0.520 [0.472, 0.569] 0.449 [0.435, 0.462] MedGemma-4B 0.277 [0.238, 0.317] 0.359 [0.302, 0.418] 0.266 [0.214, 0.320] 0.310 [0.247, 0.377] 0.474 [0.439, 0.510] 0.401 [0.372, 0.432] 0.443 [0.401, 0.485] 0.376 [0.364, 0.387] MedGemma-27B 0.379 [0.330, 0.430] 0.457 [0.384, 0.532] 0.368 [0.297, 0.444] 0.387 [0.318, 0.463] 0.558 [0.523, 0.594] 0.478 [0.445, 0.510] 0.529 [0.484, 0.576] 0.466 [0.453, 0.479] Overall 0.398 [0.354, 0.441] 0.473 [0.417, 0.529] 0.398 [0.352, 0.444] 0.435 [0.381, 0.489] 0.575 [0.550, 0.599] 0.495 [0.470, 0.520] 0.551 [0.516, 0.584] 0.475 [0.465, 0.485] Supplementary Table B1| Fact recall scores across medical specialties on ClinMM-Bench. Values indicate mean fact recall scores, with 95% confidence intervals estimated using 100,000 bootstrap resamples. 23 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 0.074 [0.052, 0.097] 0.108 [0.059, 0.166] 0.097 [0.051, 0.150] 0.081 [0.050, 0.121] 0.060 [0.046, 0.074] 0.092 [0.074, 0.112] 0.067 [0.045, 0.092] 0.142 [0.131, 0.153] Gemini 3 Pro 0.060 [0.039, 0.085] 0.052 [0.019, 0.093] 0.091 [0.050, 0.138] 0.080 [0.041, 0.130] 0.048 [0.035, 0.062] 0.075 [0.055, 0.098] 0.059 [0.039, 0.082] 0.098 [0.088, 0.109] GPT-5-minimal 0.067 [0.046, 0.090] 0.070 [0.041, 0.103] 0.098 [0.049, 0.161] 0.077 [0.046, 0.112] 0.069 [0.051, 0.089] 0.091 [0.072, 0.111] 0.063 [0.041, 0.088] 0.132 [0.121, 0.144] GPT-5-medium 0.090 [0.063, 0.119] 0.091 [0.041, 0.153] 0.100 [0.046, 0.166] 0.080 [0.049, 0.115] 0.056 [0.040, 0.074] 0.075 [0.055, 0.096] 0.063 [0.038, 0.091] 0.124 [0.112, 0.135] Gemma-3-4B 0.156 [0.128, 0.184] 0.229 [0.165, 0.292] 0.186 [0.128, 0.250] 0.141 [0.101, 0.186] 0.180 [0.155, 0.205] 0.205 [0.178, 0.232] 0.198 [0.158, 0.240] 0.268 [0.254, 0.282] Gemma-3-27B 0.131 [0.103, 0.162] 0.152 [0.091, 0.220] 0.192 [0.130, 0.261] 0.107 [0.076, 0.139] 0.114 [0.092, 0.137] 0.175 [0.148, 0.203] 0.123 [0.092, 0.158] 0.203 [0.191, 0.215] LLaMA-4-Scout 0.079 [0.056, 0.105] 0.141 [0.084, 0.204] 0.127 [0.082, 0.175] 0.082 [0.047, 0.123] 0.081 [0.063, 0.099] 0.136 [0.113, 0.161] 0.099 [0.072, 0.128] 0.148 [0.137, 0.160] Qwen3-VL-4B 0.117 [0.094, 0.143] 0.123 [0.086, 0.163] 0.170 [0.121, 0.224] 0.096 [0.070, 0.123] 0.116 [0.097, 0.137] 0.140 [0.122, 0.159] 0.130 [0.102, 0.159] 0.175 [0.166, 0.185] Qwen3-VL-8B 0.113 [0.090, 0.138] 0.112 [0.068, 0.159] 0.154 [0.109, 0.201] 0.131 [0.097, 0.166] 0.116 [0.097, 0.137] 0.135 [0.117, 0.153] 0.112 [0.090, 0.135] 0.173 [0.163, 0.184] Qwen3-VL-32B 0.116 [0.088, 0.148] 0.070 [0.039, 0.108] 0.114 [0.068, 0.168] 0.095 [0.059, 0.135] 0.076 [0.060, 0.093] 0.134 [0.112, 0.157] 0.102 [0.073, 0.134] 0.166 [0.155, 0.177] Qwen3-VL-4B-thinking 0.144 [0.116, 0.173] 0.142 [0.100, 0.186] 0.160 [0.110, 0.214] 0.134 [0.094, 0.176] 0.140 [0.115, 0.165] 0.174 [0.149, 0.200] 0.149 [0.122, 0.178] 0.205 [0.194, 0.216] Qwen3-VL-8B-thinking 0.140 [0.111, 0.170] 0.113 [0.063, 0.168] 0.155 [0.106, 0.210] 0.113 [0.076, 0.156] 0.126 [0.104, 0.149] 0.149 [0.128, 0.171] 0.134 [0.106, 0.163] 0.195 [0.184, 0.206] Qwen3-VL-32B-thinking 0.132 [0.105, 0.160] 0.143 [0.092, 0.197] 0.162 [0.108, 0.220] 0.098 [0.066, 0.134] 0.092 [0.074, 0.112] 0.167 [0.144, 0.192] 0.130 [0.099, 0.163] 0.205 [0.193, 0.216] MedGemma-4B 0.125 [0.098, 0.153] 0.116 [0.068, 0.171] 0.138 [0.087, 0.197] 0.091 [0.058, 0.136] 0.132 [0.107, 0.157] 0.147 [0.127, 0.168] 0.133 [0.105, 0.163] 0.199 [0.187, 0.211] MedGemma-27B 0.086 [0.063, 0.111] 0.146 [0.075, 0.230] 0.142 [0.072, 0.227] 0.100 [0.058, 0.146] 0.094 [0.072, 0.118] 0.139 [0.116, 0.164] 0.101 [0.074, 0.132] 0.187 [0.175, 0.200] Overall 0.109 [0.092, 0.126] 0.120 [0.088, 0.155] 0.139 [0.103, 0.180] 0.100 [0.079, 0.123] 0.100 [0.088, 0.112] 0.136 [0.124, 0.148] 0.111 [0.093, 0.130] 0.175 [0.168, 0.182] Supplementary Table B2| Hallucination scores across medical specialties on ClinMM-Bench. Values indi- cate mean hallucination scores, with 95% confidence intervals estimated using 100,000 bootstrap resamples. 24 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Model Derm. (n=63) Emerg. (n=23) Intern. (n=23) Neph. (n=29) Neuro. (n=98) Onco. (n=114) Ophth. (n=60) Radiol. (n=679) Claude-4.5-Sonnet 0.254 [0.218, 0.292] 0.263 [0.217, 0.308] 0.230 [0.183, 0.282] 0.266 [0.224, 0.310] 0.386 [0.360, 0.412] 0.332 [0.307, 0.357] 0.344 [0.308, 0.381] 0.293 [0.282, 0.303] Gemini 3 Pro 0.352 [0.308, 0.396] 0.350 [0.299, 0.401] 0.347 [0.292, 0.406] 0.366 [0.303, 0.430] 0.483 [0.454, 0.512] 0.425 [0.393, 0.457] 0.447 [0.404, 0.491] 0.408 [0.395, 0.421] GPT-5-minimal 0.311 [0.266, 0.357] 0.334 [0.286, 0.384] 0.305 [0.251, 0.362] 0.329 [0.277, 0.379] 0.443 [0.417, 0.470] 0.389 [0.360, 0.419] 0.418 [0.380, 0.456] 0.371 [0.359, 0.384] GPT-5-medium 0.265 [0.225, 0.307] 0.309 [0.267, 0.351] 0.278 [0.229, 0.328] 0.306 [0.254, 0.359] 0.398 [0.369, 0.427] 0.383 [0.355, 0.410] 0.376 [0.336, 0.420] 0.337 [0.326, 0.348] Gemma-3-4B 0.230 [0.193, 0.269] 0.260 [0.206, 0.314] 0.237 [0.191, 0.282] 0.282 [0.232, 0.333] 0.386 [0.359, 0.413] 0.344 [0.318, 0.370] 0.344 [0.300, 0.390] 0.313 [0.302, 0.323] Gemma-3-27B 0.279 [0.244, 0.314] 0.303 [0.251, 0.353] 0.233 [0.176, 0.295] 0.282 [0.231, 0.339] 0.403 [0.376, 0.430] 0.370 [0.342, 0.400] 0.380 [0.342, 0.418] 0.348 [0.336, 0.359] LLaMA-4-Scout 0.341 [0.302, 0.381] 0.367 [0.313, 0.420] 0.328 [0.280, 0.378] 0.347 [0.290, 0.407] 0.480 [0.450, 0.510] 0.456 [0.425, 0.487] 0.449 [0.402, 0.495] 0.436 [0.423, 0.449] Qwen3-VL-4B 0.236 [0.203, 0.272] 0.251 [0.196, 0.309] 0.211 [0.175, 0.247] 0.234 [0.191, 0.278] 0.353 [0.325, 0.381] 0.313 [0.290, 0.336] 0.283 [0.248, 0.318] 0.296 [0.286, 0.307] Qwen3-VL-8B 0.237 [0.203, 0.273] 0.260 [0.215, 0.308] 0.242 [0.196, 0.290] 0.202 [0.169, 0.238] 0.367 [0.340, 0.395] 0.315 [0.293, 0.338] 0.335 [0.299, 0.372] 0.306 [0.296, 0.316] Qwen3-VL-32B 0.264 [0.231, 0.298] 0.276 [0.227, 0.325] 0.256 [0.199, 0.316] 0.283 [0.234, 0.337] 0.389 [0.360, 0.419] 0.342 [0.316, 0.369] 0.354 [0.318, 0.392] 0.318 [0.307, 0.329] Qwen3-VL-4B-thinking 0.237 [0.202, 0.273] 0.241 [0.196, 0.289] 0.232 [0.185, 0.279] 0.240 [0.204, 0.278] 0.377 [0.352, 0.404] 0.316 [0.292, 0.341] 0.320 [0.282, 0.361] 0.293 [0.282, 0.304] Qwen3-VL-8B-thinking 0.242 [0.208, 0.278] 0.277 [0.214, 0.345] 0.209 [0.174, 0.245] 0.238 [0.198, 0.282] 0.367 [0.340, 0.395] 0.297 [0.272, 0.323] 0.314 [0.279, 0.352] 0.283 [0.273, 0.294] Qwen3-VL-32B-thinking 0.254 [0.218, 0.290] 0.252 [0.202, 0.303] 0.250 [0.190, 0.312] 0.270 [0.221, 0.321] 0.386 [0.358, 0.415] 0.311 [0.285, 0.338] 0.304 [0.264, 0.346] 0.276 [0.266, 0.286] MedGemma-4B 0.261 [0.226, 0.299] 0.310 [0.249, 0.374] 0.201 [0.164, 0.241] 0.238 [0.193, 0.286] 0.405 [0.373, 0.438] 0.389 [0.362, 0.416] 0.377 [0.336, 0.420] 0.358 [0.347, 0.369] MedGemma-27B 0.259 [0.219, 0.300] 0.290 [0.247, 0.333] 0.215 [0.155, 0.287] 0.253 [0.207, 0.303] 0.372 [0.347, 0.399] 0.337 [0.309, 0.365] 0.334 [0.293, 0.377] 0.318 [0.307, 0.329] Overall 0.268 [0.239, 0.298] 0.290 [0.251, 0.326] 0.252 [0.219, 0.286] 0.276 [0.240, 0.311] 0.400 [0.382, 0.417] 0.355 [0.336, 0.373] 0.359 [0.330, 0.387] 0.330 [0.323, 0.338] Supplementary Table B3 | Fact density scores across medical specialties on ClinMM-Bench. Values indi- cate mean fact density scores, with 95% confidence intervals estimated using 100,000 bootstrap resamples. 25 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information C Search Strategy Dermatology: (dermatology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter] AND english[Filter]) Emergency Medicine: (emergency medicine[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter]) Internal Medicine: (internal medicine[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter]) Nephrology: (nephrology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter] AND english[Filter]) Neurology: (neurology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter] AND english[Filter]) Oncology: (oncology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter] AND english[Filter]) Ophthalmology: (ophthalmology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND humans[Filter] AND english[Filter]) Radiology: (radiology[Title/Abstract] AND case reports[Publication Type] AND pubmed pmc open access[Filter] AND english[Filter]) 26 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information D Data Curation Flow SpecialtyCollection & ExtractionInspectionValidationConversionQuality ControlIncluded Cases Derm.26513968686663 Emerg.1296525252523 Intern.973924242423 Neph.1276632323229 Neuro.40018910110110198 Onco.758373123123122114 Ophth.28412362626160 Radiol.2,5821,387709709693679 Total4,6422,3811,1441,1441,1241,089 Supplementary Table D1 | Data curation flow across medical specialties. Counts are shown after each sequential data-curation stage; the final column reports cases retained for benchmark evaluation after all curation stages. 27 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information E Data Validation Prompt Please determine whether the following medical case report is suitable for use as a diagnostic teaching case (similar to USMLE format). A case report is suitable as a diagnostic teaching case if it meets the following criteria: 1. Describes a complete diagnostic process for a single patient; 2. Provides relevant clinical information such as demographic characteristics, chief complaint, medical history, physical examination, laboratory and imaging findings; 3. Includes key diagnostic images (check for <fig> tags with <graphic> elements): – Initial clinical photographs (e.g., lesions, symptoms); – Diagnostic imaging (e.g., X-ray, CT, MRI, ultrasound, endoscopy); – Other relevant test images (e.g., ECG, pathology slides, blood smears); 4. States the final diagnosis clearly; 5. Contains the reasoning process that led to the diagnosis. Characteristics of an unsuitable report (NOT_DIAGNOSTIC_SUITABLE): 1. Focuses on treatment outcomes, surgical techniques, or follow-up results; 2. Reports adverse drug reactions or complication management; 3. Consists of literature reviews, theoretical analyses, or opinion statements; 4. Lacks a diagnostic reasoning process and contains only simple case descriptions. Case Report: case_report Output: Please respond with only: "DIAGNOSTIC_SUITABLE" or "NOT_DIAGNOSTIC_SUITABLE" 28 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information F Data Conversion Prompt You are a medical information extraction expert. Your task is to convert a medical case report into a struc- tured JSON format. Return only the JSON result. Do not include explanations or additional comments. Task Convert the original medical case report into the following JSON format: "case_presentation": "english": "demographic_information": "", "chief_complaint": "", "history_of_present_illness": "", "lab_result": "", "physical_examination": "" , "chinese": "demographic_information": "", "chief_complaint": "", "history_of_present_illness": "", "lab_result": "", "physical_examination": "" , "images": "english": [ "image_id": "", "image_type": "", "image_finding": "", "diagnostic_significance": "" ], "chinese": [ "image_id": "", "image_type": "", "image_finding": "", "diagnostic_significance": "" ] , "final_diagnosis": "english": "", "chinese": "" , "diagnostic_reasoning": "english": "", "chinese": "" Conversion Rules 29 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Case Presentation MUST NOT contain: 1. Final Diagnosis: - Case presentation must absolutely not reveal the final diagnosis. 2. Image Diagnostic Information: – Any information that can only be obtained through image analysis must not appear in the case presentation, including but not limited to: – Specific results of imaging examinations. – Pathological examination cell morphology and staining results. – Endoscopic examination lesion characteristics. – Dermatological lesion morphological descriptions. Case Presentation CAN include: – Patient information (age, gender, etc.). – Symptom descriptions. – Medical history. – Laboratory test values (blood routine, biochemistry, etc.). – Examination process descriptions (e.g., “CT scan performed", “biopsy taken"). Physical Examination Field: Only write what examinations were performed, NOT the examination results. Acceptable Examples: – “Chest CT scan performed". – “Bone marrow biopsy conducted". – “Immunohistochemical staining performed". – “PET scan examination conducted". Not Allowed: – “CT showed bilateral lung lesions". – “Biopsy revealed malignant cells". Images Field: For each image/figure mentioned in the original text, extract: – image_id: The figure number/label from the original text. – image_type: Concise description of the image type (e.g., “H&E staining histological image", “Chest CT scan", “Dermatological photograph", “Immunohistochemistry stain"). – image_finding: What should be visible in this image based on the original text (e.g., “Malignant cells with nuclear atypia", “Bilateral lung nodules", “Erythematous skin lesions"). – diagnostic_significance: What this image contributes to the diagnosis (e.g., “Confirms malig- nancy", “Shows metastatic spread", “Indicates inflammatory process"). Final Diagnosis Field: Write out the clear final diagnosis as stated in the original text. Diagnostic Reasoning Field: Explain how the diagnosis was reached based on the diagnostic evidence and process described in the original text. Include: – How clinical findings contributed to the diagnosis – How each image/test result supported or ruled out differential diagnoses – The logical progression from initial presentation to final diagnosis Must be strictly based on original text content, cannot add any content not in the original: – Cannot perform your own medical reasoning or add medical knowledge. – Cannot speculate or supplement information not mentioned in the original text. Case Report: case_report Output Please return ONLY the JSON format result: 30 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information G Quality Control Prompt You are a medical expert. Your task is to evaluate the quality of a converted medical case report for use as a diagnostic examination case. The converted content must strictly adhere to the facts of the original case report, without speculation or additions. You will be provided with: 1. Original XML medical case report 2. Converted JSON structured format Please evaluate the conversion quality. Evaluation Criteria Information Accuracy Score (1-5 Scale) Rate how accurately the converted JSON preserves information from the original XML: 5 - Excellent: All information perfectly extracted and preserved 4 - Good: Minor discrepancies that don't affect meaning 3 - Acceptable: Some information loss but key elements preserved 2 - Poor: Significant information loss or errors 1 - Unacceptable: Major information loss or errors Original XML Case Report: original_xml Converted JSON Case Report: converted_json Output Format Please provide your evaluation in the following JSON format: "information_accuracy_score": 1-5, "information_accuracy_reason": "Explanation of accuracy assessment" Return only the JSON result without additional explanations. 31 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information H Diagnostic Accuracy Evaluation Prompt Please determine whether the predicted diagnosis is correct. Predicted Diagnosis: predicted_diagnosis Ground Truth Diagnosis: ground_truth_diagnosis Accuracy Evaluation: 0 - Completely Incorrect 1 - Partially Correct 2 - Completely Correct Output ONLY the following valid XML (no other text): <result> <reasoning>your detailed explanation here</reasoning> <accuracy>0 or 1 or 2</accuracy> </result> 32 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases Supplementary Information I I1. Atomic fact extraction prompt You are a medical expert. Extract atomic facts from the diagnostic reasoning text. An atomic fact should: 1. Be a single, indivisible piece of medical information 2. Contain one subject-predicate-object relationship 3. Be specific and factual (avoid vague statements) Examples of good atomic facts: – “RPR test is positive" – “CSF shows pleocytosis" – “Fundus exam reveals vitritis" Examples of what to avoid: – “Patient has multiple symptoms" (too vague) – “Tests were done" (not specific enough) Return only valid JSON in this exact format: "facts": [ "atomic fact 1", "atomic fact 2" ] 33 Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases I2. Atomic Fact Matching Prompt You are a medical expert. Compare predicted atomic facts against ground-truth atomic facts. Tasks: 1. Match facts that express the same clinical meaning. 2. Allow N-to-M matching: multiple predicted facts may match one ground-truth fact, and one predicted fact may match multiple ground-truth facts. 3. Identify hallucinated predicted facts: predicted facts that are factually incorrect or contradict the ground truth. Clinical Matching Rules: – Do not match facts if the specimen differs, such as serum vs CSF. – Do not match facts if the anatomical location or laterality differs. – Do not match positive findings with negative, suspected, or ruled-out findings. – Do not match causal or diagnostic relationships if the direction is different. – If an unmatched predicted fact is merely not covered by the ground truth but is not clearly wrong, do not label it as hallucinated. Important Output Rules: – Use IDs only. Do not copy fact text into the JSON. – Predicted IDs must come from the P-list. – Ground-truth IDs must come from the G-list. – Every unmatched predicted ID must appear in exactly one of: hallucinated_predicted_ids or non_hallucinated_unmatched_predicted_ids. – A predicted ID must not appear in both matched_pairs and either hallucinated_predicted_ids or non_hallucinated_unmatched_predicted_ids. Return only valid JSON in this exact format: "matched_pairs": [ "reasoning": "short explanation", "predicted_ids": [1], "ground_truth_ids": [1,2] ], "hallucinated_predicted_ids": [3], "non_hallucinated_unmatched_predicted_ids": [5], "unmatched_ground_truth_ids": [4] Predicted Facts: predicted_facts Ground-Truth Facts: ground_truth_facts 34