Paper deep dive
Evidence-Grounded Subspecialty Reasoning: Evaluating a Curated Clinical Intelligence Layer on the 2025 Endocrinology Board-Style Examination
Amir Hosseinian, MohammadReza Zare Shahneh, Umer Mansoor, Gilbert Szeto, Kirill Karlin, Nima Aghaeepour
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/21/2026, 2:24:46 AM
Summary
This study evaluates 'January Mirror,' an evidence-grounded clinical reasoning system, against frontier Large Language Models (GPT-5, GPT-5.2, Gemini-3-Pro) on the 2025 Endocrine Self-Assessment Program (ESAP). Mirror, which uses a curated endocrinology corpus and structured reasoning without web access, achieved 87.5% accuracy, significantly outperforming human references (62.3%) and all tested LLMs. The system demonstrated high evidence traceability and citation accuracy, suggesting that curated, domain-specific evidence curation can outperform unconstrained web retrieval for subspecialty clinical reasoning.
Entities (10)
Relation Signals (9)
January Mirror â evaluatedon â ESAP 2025
confidence 98% · We evaluated January Mirror ... on the Endocrine Self-Assessment Program (ESAP) 2025
January Mirror â outperformed â GPT-5.2
confidence 95% · Mirror achieved 87.5% accuracy ... compared to 74.6% for GPT-5.2
January Mirror â outperformed â GPT-5
confidence 95% · Mirror achieved 87.5% accuracy ... compared to 74.0% for GPT-5
January Mirror â outperformed â Gemini-3-Pro
confidence 95% · Mirror achieved 87.5% accuracy ... compared to 69.8% for Gemini-3-Pro
January Mirror â usesevidencefrom â Endocrine Society
confidence 90% · Mirrorâs evidence corpus is organized in a multi-tiered hierarchy ... including the Endocrine Society
January Mirror â usesevidencefrom â American Diabetes Association
confidence 90% · Mirrorâs evidence corpus ... including the ... American Diabetes Association (ADA)
January Mirror â usesevidencefrom â KDIGO
confidence 88% · Mirrorâs evidence corpus ... including ... Kidney Disease: Improving Global Outcomes (KDIGO)
January Mirror â supportstreatmentwith â GLP-1 receptor agonists
confidence 85% · Mirrorâs evidence corpus ... particularly for therapeutic classes such as GLP-1 receptor agonists
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Large language models have demonstrated strong performance on general medical examinations, but subspecialty clinical reasoning remains challenging due to rapidly evolving guidelines and nuanced evidence hierarchies. Methods: We evaluated January Mirror, an evidence-grounded clinical reasoning system, against frontier LLMs (GPT-5, GPT-5.2, Gemini-3-Pro) on a 120-question endocrinology board-style examination. Mirror integrates a curated endocrinology and cardiometabolic evidence corpus with a structured reasoning architecture to generate evidence-linked outputs. Mirror operated under a closed-evidence constraint without external retrieval. Comparator LLMs had real-time web access to guidelines and primary literature. Results: Mirror achieved 87.5% accuracy (105/120; 95% CI: 80.4-92.3%), exceeding a human reference of 62.3% and frontier LLMs including GPT-5.2 (74.6%), GPT-5 (74.0%), and Gemini-3-Pro (69.8%). On the 30 most difficult questions (human accuracy less than 50%), Mirror achieved 76.7% accuracy. Top-2 accuracy was 92.5% for Mirror versus 85.25% for GPT-5.2. Conclusions: Mirror provided evidence traceability: 74.2% of outputs cited at least one guideline-tier source, with 100% citation accuracy on manual verification. Curated evidence with explicit provenance can outperform unconstrained web retrieval for subspecialty clinical reasoning and supports auditability for clinical deployment.
Tags
Links
- Source: https://arxiv.org/abs/2602.16050v1
- Canonical: https://arxiv.org/abs/2602.16050v1
Trouble viewing inline? Open PDF directly â
Full Text
37,094 characters extracted from source content.
Expand or collapse full text
Benchmarking an Evidence-Grounded Clinical Intelligence Layer on the ESAP 2025 Endocrinology Examination Amir Hosseinian1, Reza Shahneh1, Umer Mansoor1, Gilbert Szeto1, Kirill Karlin1, Nima Aghaeepour2 1January AI; 2Stanford University Corresponding author: amirhoss@january.ai Abstract Background: Large language models have demonstrated strong performance on general medical examinations, but subspecialty clinical reasoning remains challenging due to rapidly evolving guidelines and nuanced evidence hierarchies. Methods: We evaluated January Mirror, an evidence-grounded clinical reasoning system, against frontier LLMs (GPT-5, GPT-5.2, Gemini-3-Pro) on the Endocrine Self-Assessment Program (ESAP) 2025, a 120-question board-style examination. Mirror integrates a curated endocrinology and cardiometabolic evidence corpus with a structured reasoning architecture to generate evidence-linked outputs. Mirror operated under a closed-evidence constraint without external retrieval. Comparator LLMs had real-time web access to guidelines and primary literature. Results: Mirror achieved 87.5% accuracy (105/120; 95% CI: 80.4â92.3%), exceeding the human reference of 62.3% (ESAP respondent distributions) and frontier LLMs including GPT-5.2 (74.6%), GPT-5 (74.0%), and Gemini-3-Pro (69.8%). On the 30 most difficult questions (human accuracy <50<50%), Mirror achieved 76.7% accuracy. Top-2 accuracy was 92.5% for Mirror versus 85.25% for GPT-5.2. Conclusions: Mirror provided evidence traceability: 74.2% of outputs cited at least one guideline-tier source, with 100% citation accuracy on manual verification. Curated evidence with explicit provenance can outperform unconstrained web retrieval for subspecialty clinical reasoning and supports auditability for clinical deployment. Keywords: clinical decision support, endocrinology, evidence-based medicine, large language models, medical AI, healthcare AI 1 Introduction The successful performance of large language models on the United States Medical Licensing Examination generated significant enthusiasm about AIâs potential in clinical medicine. In 2023, Kung et al. demonstrated that ChatGPT performed at or near the passing threshold for all three USMLE steps without specialized training (Kung et al., 2023), and subsequent work by Nori et al. showed that GPT-4 exceeded the USMLE passing score by over 20 points (Nori et al., 2023). However, the USMLE assesses general medical knowledge expected of all physicians, not the specialized reasoning required for subspecialty practice. Real-world clinical decisions frequently demand subspecialty expertise, the ability to integrate current society guidelines, interpret nuanced clinical evidence, and reason across complex patient presentations with multiple comorbidities. Endocrinology presents a particularly demanding test case. The American Board of Internal Medicine (ABIM) reports that endocrinology certification pass rates have shown notable variability, falling to 74% for first-time takers in recent years before recovering to 85% in 2024 (ABIM, 2024). This difficulty stems from several factors: the rapid evolution of treatment guidelines (particularly for diabetes, obesity, and thyroid disorders), the complex interplay between metabolic systems, and the need to integrate emerging therapeutic classes such as GLP-1 receptor agonists and SGLT2 inhibitors into established treatment algorithms. A system capable of reliable endocrinology reasoning would demonstrate capabilities directly relevant to clinical practice. Current approaches to improving LLM clinical performance have focused primarily on model scale and retrieval augmentation. Larger models with more parameters, combined with real-time web search capabilities, represent the dominant paradigm. Yet this approach faces fundamental limitations: web retrieval returns heterogeneous sources without consistent evidence grading, provenance normalization, or domain-scoped filtering, which can reduce reliability even when the underlying information is available (Kohandel Gargari and Habibi, 2025). There is no mechanism for weighting evidence by source quality, recency, or clinical relevance. Furthermore, LLMs are known to produce hallucinations, factually incorrect or fabricated information, which pose particular risks in clinical settings where errors can compromise patient safety (Asgari et al., 2025; Roustan and Bastardot, 2025). While retrieval augmentation improves factual recall, it does not impose constraints on evidence hierarchy, guideline precedence, or clinical relevance. As a result, models may retrieve correct information yet fail to apply it appropriately in context. We hypothesized that an alternative approach, domain-specific evidence curation combined with structured clinical reasoning, could outperform model scale and general web retrieval. To test this hypothesis, we developed Mirror, an evidence-grounded clinical reasoning system built around a curated endocrinology and cardiometabolic evidence corpus and a structured clinical reasoning stack designed to produce evidence-linked outputs. We evaluated Mirror against frontier LLMs on the Endocrine Self-Assessment Program (ESAP) 2025, a rigorous board-style examination developed by the Endocrine Society for certification maintenance (Endocrine Society, 2025). Mirror is a broader clinical platform; however, this paper evaluates only the evidence-grounded clinical intelligence layer responsible for producing evidence-linked, verifiable outputs. We focus on this layer because credibility and traceability are prerequisites for clinician-facing and enterprise clinical workflows, independent of how patient context is sourced or represented. This paper is the first in a planned validation series focused on the credibility of Mirrorâs clinical intelligence layer. This study benchmarks subspecialty exam performance; the next paper extends evaluation across additional standardized examinations; and the third paper evaluates verifiability and clinical usefulness via blinded clinician review on complex case vignettes. Together, the series is intended to establish whether Mirrorâs outputs are reliable enough to support clinician-facing and enterprise clinical workflows. 2 Methods 2.1 Benchmark Dataset We utilized the Endocrine Self-Assessment Program (ESAP) 2025, developed by the Endocrine Society as a self-assessment tool for endocrinologists preparing for board certification and maintenance of certification examinations (Endocrine Society, 2025). ESAP consists of 120 multiple-choice questions spanning the breadth of clinical endocrinology, including disorders of the thyroid, adrenal, pituitary, parathyroid, and reproductive systems, as well as diabetes mellitus, lipid disorders, and obesity medicine. Each question presents a clinical vignette with patient history, physical examination findings, and laboratory results, followed by a single-best-answer format. Questions were categorized by clinical reasoning type, with each question potentially involving multiple reasoning domains: diagnosis (47.5% of questions), pathophysiology (50.0%), treatment selection (60.8%), diagnostic testing (39.2%), and risk assessment/prognosis (31.7%). This multi-label categorization reflects the integrated nature of clinical reasoning, where a single case may require diagnostic, therapeutic, and prognostic considerations. The examination is designed to assess reasoning at the level expected of a practicing endocrinologist and meets ABIM standards for maintenance of certification. Figure 1 displays the distribution of clinical reasoning domains and their co-occurrence patterns across the examination. Figure 1: Distribution of clinical reasoning domains in ESAP 2025. The left panel shows the total number of questions involving each domain. The main panel displays intersection sizes for domain combinations, sorted by frequency. The matrix indicates which domains comprise each intersection. Treatment-related reasoning was most prevalent (n=73), followed by pathophysiology (n=60) and diagnosis (n=57). The most common domain combination was risk/prognosis with treatment (n=16). 2.2 Systems Evaluated We compared Mirror against three frontier LLMs: GPT-5, GPT-5.2, and Gemini-3-Pro. Frontier LLMs were evaluated with unrestricted web access enabled, allowing each model to perform as many web queries as needed to answer each question. Mirror operated without external web access and relied exclusively on its curated evidence base. This design isolates the contribution of domain-curated evidence and structured reasoning relative to general-purpose models augmented with web retrieval. Table 1: Systems evaluated and their capabilities System Web Access Evidence Base Architecture Mirror No Curated endocrinology corpus Evidence-grounded ensemble GPT-5.2 Yes Real-time web retrieval Monolithic LLM GPT-5 Yes Real-time web retrieval Monolithic LLM Gemini-3-Pro Yes Real-time web retrieval Monolithic LLM 2.3 Web-Assisted Baseline Protocol For web-assisted baselines, browsing was enabled with unrestricted access and consistent instructions to prioritize primary society guidelines, peer-reviewed articles, and authoritative medical references. Retrieved content was provided to the model as context for the final answer within the same interaction, and the model was required to output a single best answer choice. 2.4 Mirror System Architecture Mirror is designed as a clinical reasoning system that grounds outputs in curated medical evidence. The system comprises two primary components: a domain-specific evidence corpus and a structured reasoning architecture. Evidence Corpus. Mirrorâs evidence corpus is organized in a multi-tiered hierarchy reflecting clinical evidence standards in endocrinology and cardiometabolic medicine. The first tier comprises society-level guidelines and practice statements from authoritative bodies including the Endocrine Society, American Diabetes Association (ADA), American Association of Clinical Endocrinology (AACE), Kidney Disease: Improving Global Outcomes (KDIGO), and related cardiometabolic guideline organizations (ADA, 2024; Arnett et al., 2019; KDIGO, 2022; Samson et al., 2023; Rinella et al., 2023). The second tier includes high-impact peer-reviewed clinical literature from journals with established methodological rigor relevant to endocrinology and internal medicine. The third tier comprises landmark and pivotal late-phase randomized controlled trials that have shaped current practice, particularly for therapeutic classes such as GLP-1 receptor agonists, dual incretin agonists, SGLT2 inhibitors, and obesity interventions. Reasoning Architecture. Mirror uses an ensemble-style clinical reasoning stack in which multiple specialized reasoning components generate candidate answers and supporting evidence links, followed by an arbitration layer that selects the final output using evidence quality and internal agreement signals. The components are organized around common clinical question archetypes (e.g., diagnosis, testing, treatment, prognosis, and mechanistic reasoning) to reflect how clinicians approach distinct reasoning tasks. To avoid disclosing proprietary implementation details, we report the systemâs inputs/outputs and evaluation behavior while omitting internal ranking heuristics and model-specific routing parameters. 2.5 Evaluation Protocol All systems were evaluated under zero-shot conditions without access to ESAP questions or answers during development or training. Questions were presented without preprocessing, exactly as they appear in the examination. The primary outcome measure was single-answer accuracy: the proportion of questions for which the system selected the correct answer. As a secondary outcome, we evaluated top-2 accuracy, whether the correct answer appeared among the systemâs top two selections. This metric reflects a clinical decision support workflow in which physicians use AI to narrow options rather than provide definitive single answers. Performance was stratified by question type (diagnosis, testing, treatment, risk/prognosis, pathophysiology) to identify domain-specific strengths and weaknesses. Statistical comparisons between systems used McNemarâs test for paired proportions, with significance set at p<0.05p<0.05. Confidence intervals (95%) were calculated using the Wilson score method. 2.6 Citation Verification Protocol To assess evidence traceability and citation accuracy, we conducted a manual verification audit of Mirrorâs outputs. For each of the 120 questions, Mirror produced an answer accompanied by citations to specific evidence sources from its curated corpus. A random sample of 60 questions was selected for detailed citation verification. For each sampled question, a clinical reviewer assessed: (1) whether the cited source existed in the corpus, (2) whether the cited passage was accurately quoted or paraphrased, and (3) whether the cited evidence logically supported the selected answer. Citations were classified as âverified accurateâ if all three criteria were met. Additionally, we assessed guideline concordance for treatment-related questions. For each treatment question (n=73n=73), the relevant society guideline recommendation was identified, and the systemâs answer was scored as guideline-concordant or guideline-discordant. 3 Results 3.1 Overall Performance Mirror substantially outperformed all frontier LLMs despite operating without web access. Mirror achieved 87.5% accuracy on the ESAP 2025 examination (n=120n=120 questions; 95% CI: 80.4â92.3%), compared to 74.6% for GPT-5.2 (95% CI: 66.7â82.5%), 74.0% for GPT-5 (95% CI: 66.7â80.0%), and 69.8% for Gemini-3-Pro (95% CI: 61.6â77.1%), all of which had real-time web access to clinical resources. Pairwise McNemarâs tests confirmed statistically significant differences between Mirror and all three baselines: GPT-5.2 (p=0.011p=0.011), GPT-5 (p<0.001p<0.001), and Gemini-3-Pro (p<0.001p<0.001). All comparisons remained significant after Bonferroni correction for multiple testing (α=0.017α=0.017). While confidence intervals overlap between some systems, the paired analysis demonstrates consistent directional advantage for Mirror. Figure 2 displays the comparative performance across all systems, including the ESAP respondent reference derived from aggregate answer distributions. Mirrorâs performance approached an internal upper-bound reference condition in which a baseline model was provided with manually selected, question-specific evidence passages. Figure 2: Comparative performance on ESAP 2025 endocrinology examination. Mirror achieved 87.5% accuracy compared to 74.6% for the best-performing frontier LLM (GPT-5.2) and 62.3% for the human reference (ESAP respondent mean). Optional dashed line indicates an internal upper-bound reference condition, provided for context. Error bars represent 95% confidence intervals. *p<0.05p<0.05 vs. Mirror. 3.2 Performance Relative to ESAP Respondent Reference ESAP provides aggregate respondent statistics indicating the percentage of test-takers who selected each answer option. We used these distributions to calculate an item-level respondent correctness rate of 62.3% (SD = 16.8%), representing the mean percentage of ESAP respondents selecting the correct answer across all 120 questions. This metric serves as a proxy for human performance, though ESAP respondents are self-selected endocrinologists who may have used external resources during self-assessment. Individual question difficulty ranged from 22.3% to 94.4% correct. Mirrorâs 87.5% accuracy exceeded this respondent reference by 25.2 percentage points. To assess whether this advantage extended to the most challenging clinical scenarios, we stratified questions by human performance. On the 30 most difficult questions (human accuracy <<50%, mean 39.5%), Mirror achieved 76.7% accuracy (23/30), compared to 53.3% for GPT-5.2, 53.3% for GPT-5, and 46.7% for Gemini-3-Pro. On 71 medium-difficulty questions (human accuracy 50â80%, mean 65.6%), Mirror achieved 90.1% versus 80.3% for GPT-5.2. On 19 easy questions (human accuracy â„ 80%), all systems performed well (Mirror 94.7%, GPT-5.2 94.7%). Figure 3 illustrates performance stratified by question difficulty. Figure 3: Performance stratified by question difficulty (based on human accuracy). Mirror maintained strong performance across all difficulty tiers, with particular advantage on hard questions (human accuracy <<50%) where it achieved 76.7% compared to 53.3% for GPT-5.2 and 39.5% human mean accuracy. Figure 4: Paired win/loss analysis comparing Mirror to each baseline. For each comparison, âwinsâ indicate questions Mirror answered correctly while the baseline erred; âlossesâ indicate the reverse. Mirror achieved a net positive margin against all baselines, with the largest advantage over Gemini-3-Pro (19 net wins). 3.3 Top-2 Accuracy When evaluating top-2 selections, Mirror achieved 92.5% accuracy (111/120 questions; 95% CI: 86.4â96.0%), indicating that the correct answer appeared among Mirrorâs two highest-ranked options in 92.5% of questions. This metric is clinically relevant because decision support systems are often used to narrow differential diagnoses or treatment options rather than provide single definitive answers. Comparable top-2 accuracy figures for frontier LLMs were 85.25% (GPT-5.2), 81.1% (GPT-5), and 83.59% (Gemini-3-Pro). The top-2 accuracy gap (7.3 percentage points vs. GPT-5.2) indicates that frontier LLMs more frequently excluded the correct answer from their top two selections, reducing their utility as differential-narrowing tools in clinical workflows. 3.4 Performance by Question Type Table 2 presents accuracy stratified by clinical question type. Note that questions may involve multiple reasoning domains; the counts reflect all questions tagged with each category. Mirror achieved its highest accuracy on diagnosis questions (91.2%) and demonstrated consistent performance across all categories (85.1â91.2%). The performance differential with frontier LLMs was largest in treatment questions, with a 15.5 percentage point difference compared to the best-performing LLM (GPT-5.2). Table 2: Accuracy (%) by question type Question Type Mirror GPT-5.2 GPT-5 Gemini-3-Pro nâ Diagnosis 91.2 77.6 73.7 71.8 57 Diagnostic Testing 85.1 70.0 68.7 65.5 47 Treatment 86.3 70.8 72.2 64.4 73 Risk/Prognosis 86.8 80.0 81.6 71.1 38 Pathophysiology 86.7 74.7 73.8 76.1 60 Overall 87.5 74.6 74.0 69.8 120 â Questions may be tagged with multiple categories; category counts sum to >>120. Figure 5: Performance by question type across systems. Mirror demonstrated consistent performance across all question categories, with the largest advantage in treatment questions (15.5 percentage point difference vs. GPT-5.2). Figure 6: Evidence provenance profile of Mirror outputs. The left panel shows the highest evidence tier cited per question: society guidelines (41.7%), peer-reviewed literature (33.8%), pivotal clinical trials (20.9%), and other sources (3.6%). The right panel indicates that 74.2% of answers included at least one guideline-tier citation, demonstrating the systemâs reliance on authoritative clinical guidance. 3.5 Error Analysis We conducted a qualitative analysis of questions answered incorrectly by each system to characterize failure modes. Mirrorâs errors (n=15n=15) were distributed across question types, with the highest concentration in treatment-tagged questions (10 of 15 errors involved treatment decisions). Error patterns included complex multi-domain questions requiring integration across diagnostic, pathophysiologic, and therapeutic reasoning, as well as cases involving nuanced interpretation of laboratory findings in atypical presentations. Representative examples are shown in Table 3. Frontier LLM errors showed a distinct pattern. GPT-5.2 frequently erred on questions requiring precise guideline-specific thresholds (e.g., A1c targets for specific patient populations, blood pressure goals in diabetic kidney disease) where the modelâs parametric knowledge conflicted with current recommendations. Gemini-3-Pro showed particular difficulty with questions involving multi-step clinical reasoning requiring integration of multiple evidence sources. Table 3: Representative error cases by system System Clinical Scenario Error Category Mirror Pituitary macroadenoma with elevated prolactin and GH; selected dopamine agonist instead of confirming acromegaly Incomplete biochemical evaluation before treatment Mirror Complete androgen insensitivity syndrome; prioritized bone density assessment over gonadal imaging Metabolic workup prioritized over malignancy surveillance Mirror Opioid-induced hypogonadism with low gonadotropins; recommended obesity pharmacotherapy instead of pituitary imaging Treatment prioritized over diagnostic workup for central hypogonadism Mirror Lactation-associated osteoporotic fracture with suboptimal calcium intake; recommended cessation of breastfeeding Overaggressive intervention when conservative management indicated GPT-5.2 Gestational diabetes with suboptimal control on basal insulin; selected outdated A1c target instead of current pregnancy-specific glucose thresholds Guideline threshold: parametric knowledge conflicted with updated ADA recommendations Gemini-3-Pro Type 2 diabetes with CKD and albuminuria on multiple agents; failed to prioritize SGLT2 inhibitor per KDIGO guidelines Multi-step integration: did not synthesize cardiorenal benefit hierarchy across comorbidities To illustrate how reasoning patterns differ between Mirror and frontier LLMs, Table 4 presents side-by-side comparisons on two representative cases where Mirror answered correctly and GPT-5.2 erred. Table 4: Side-by-side reasoning comparison: Mirror vs GPT-5.2 Case 1: Gestational Diabetes Case 2: Acute Adrenal Hemorrhage Clinical Scenario Pregnant woman (14 weeks) with history of gestational diabetes and documented needle phobia causing insulin nonadherence in prior pregnancy, resulting in preeclampsia and preterm delivery. Now requires pharmacotherapy. Previously healthy man presenting with acute abdominal pain, headache, and hypotension days after adenoviral-vector COVID-19 vaccine. CT shows bilateral adrenal hemorrhage. Labs show thrombocytopenia and markedly elevated D-dimer. Correct Answer Metformin IV hydrocortisone Mirror Answer Metformin â IV hydrocortisone â GPT-5.2 Answer Subcutaneous insulin IV immunoglobulin Mirror Reasoning âTreatment must be feasible and patient-specific. This patient has documented history of skipping insulin due to needle phobia, which already resulted in serious complications. Per ADA guidelines, metformin is an acceptable alternative when insulin is not feasible.â âCT appearance consistent with acute adrenal hemorrhage. Patient is critically unwell with tachycardia and elevated lactate. Immediate priority is treating life-threatening adrenal insufficiency with IV hydrocortisone.â GPT-5.2 Reasoning âInsulin is the preferred/standard pharmacotherapy in pregnancy because it does not cross the placenta. Address adherence by teaching partner to administer injections.â âPresentation consistent with vaccine-induced immune thrombotic thrombocytopenia (VITT). Give IVIG to block Fc receptor-mediated platelet activation.â Failure Mode Rigid application of general guideline (âinsulin is preferredâ) without adapting to documented patient-specific barrier that caused prior treatment failure Correctly identified underlying etiology (VITT) but prioritized disease-modifying therapy over immediate treatment of life-threatening adrenal crisis These cases illustrate two distinct failure patterns in frontier LLM clinical reasoning. In Case 1, GPT-5.2 correctly cited the general principle that insulin is preferred in pregnancy but failed to integrate the patient-specific context that made insulin adherence unlikely. Mirrorâs evidence corpus includes ADA guidance explicitly addressing scenarios where insulin is not feasible, enabling appropriate patient-centered adaptation. In Case 2, GPT-5.2 demonstrated sophisticated pattern recognition in identifying VITT but applied a disease-focused rather than patient-focused management hierarchy. Mirror correctly prioritized the immediate life-threatening complication (adrenal insufficiency) over treatment of the underlying thrombotic process. 3.6 Evidence Traceability and Citation Accuracy Mirror produced evidence-linked outputs for all 120 questions, with each answer accompanied by citations to specific sources in the curated corpus. Analysis of evidence provenance (Figure 6) revealed that 41.7% of outputs cited guideline-tier sources as the highest evidence tier, with 74.2% of answers including at least one society guideline citation. Manual verification of a random sample of 60 questions demonstrated 100% citation accuracy (60/60 verified), with all sampled citations meeting all three verification criteria: the cited source existed in the corpus, the passage was accurately represented, and the evidence logically supported the selected answer. Verification was performed by a clinical reviewer who was not involved in system development. For treatment-related questions (n=73n=73), Mirror achieved 86.3% guideline concordance, with 63 of 73 answers matching the relevant society guideline recommendation. This high concordance rate reflects Mirrorâs systematic prioritization of authoritative clinical guidance in therapeutic decision-making. 4 Discussion This study demonstrates that domain-specific evidence curation and structured clinical reasoning can achieve subspecialty-level performance exceeding both frontier LLMs and a human reference from practicing endocrinologists. Mirrorâs 87.5% accuracy on the ESAP 2025 endocrinology examination represents a 25.2 percentage point improvement over the human reference (62.3%) and a 12.9 percentage point improvement over the best-performing general-purpose LLM and approaches an upper-bound reference condition in which a baseline model is provided with manually selected, question-specific evidence passages. Notably, Mirrorâs advantage was most pronounced on difficult questions where human accuracy fell below 50%, suggesting the system captures clinical reasoning patterns that challenge even subspecialty-trained physicians. 4.1 Why Curated Evidence Outperforms Web Retrieval The performance differential can be attributed to three advantages of curated evidence infrastructure over unconstrained web retrieval. First, evidence curation provides quality control that web retrieval cannot match. Mirrorâs corpus includes only peer-reviewed guidelines and high-impact clinical literature, with explicit tiering by evidence strength. Web retrieval, by contrast, can surface heterogeneous sources with variable provenance and without a unified evidence-grading layer, which makes it difficult to consistently prioritize guideline-grade recommendations over lower-signal material in a repeatable way. Recent systematic reviews have highlighted that retrieval-augmented generation (RAG) can reduce hallucination rates in healthcare applications (Amugongo et al., 2025), but the quality of the underlying corpus remains critical. This matters particularly for clinical reasoning, where the strength of underlying evidence should inform confidence in recommendations. Second, structured evidence retrieval enables clinical relevance that semantic similarity cannot capture. Clinical questions often hinge on specific thresholds, contraindications, or guideline-specific recommendations that may not be semantically similar to query terms. A question about blood pressure targets in diabetic nephropathy requires retrieval of KDIGO guidelines, not just any content mentioning blood pressure and diabetes. Mirrorâs domain-specific retrieval is tuned for these clinical relevance patterns. Third, curated retrieval enables complete traceability. Every Mirror output can be linked to specific, verifiable evidence sources, with 100% citation accuracy upon manual verification. This auditability is not achievable with web retrieval, where source quality varies and provenance may be unclear. For clinical deployment, traceability is not merely desirable but essential, clinicians must be able to verify AI recommendations against primary evidence. 4.2 Clinical Implications These findings have direct implications for the development of clinical decision support systems. The dominant paradigm, larger models with web access, may be insufficient for subspecialty applications where reliability is paramount. Instead, our results suggest that investment in domain-specific evidence infrastructure may yield greater returns than model scale for clinical AI. The top-2 accuracy of 92.5% is particularly relevant for clinical workflows. Decision support systems need not provide single definitive answers; rather, they can serve as cognitive aids that help clinicians efficiently identify the most probable diagnoses or appropriate treatment options. Mirrorâs high top-2 accuracy suggests utility in this workflow even when single-answer accuracy is imperfect. Furthermore, Mirrorâs architecture provides inherent traceability, every output can be linked to specific evidence sources, enabling clinicians to verify recommendations against primary literature. This auditability addresses a key barrier to clinical AI adoption: the âblack boxâ concern that LLM outputs cannot be validated (Lee et al., 2023). 4.3 Limitations Several limitations should be considered when interpreting these results. First, this evaluation focused on a single subspecialty (endocrinology); generalization to other domains requires further validation. We selected endocrinology because of its clinical importance and examination rigor, but performance in cardiology, nephrology, or other subspecialties may differ. Second, board-style examinations, while rigorous, differ from real clinical practice. ESAP questions present well-defined vignettes with single correct answers; actual clinical encounters involve ambiguity, incomplete information, and patient preferences that are not captured in standardized testing. Our ongoing work includes evaluation on real-world clinical scenarios with human clinician assessment. Third, the human reference accuracy (62.3%) is derived from aggregate ESAP respondent statistics rather than individual test-taker scores, representing the mean percentage of respondents selecting the correct answer per question. ESAP respondents are self-selected endocrinologists pursuing certification maintenance, and testing conditions differ from AI evaluation (time pressure, fatigue, no external resources). This reference provides useful context but should not be interpreted as a direct head-to-head comparison with typical clinical practice. Fourth, the comparison between Mirror (curated corpus, no web) and frontier LLMs (web retrieval) confounds multiple variables: corpus quality, retrieval method, and access modality. Mirrorâs superior performance may reflect advantages of evidence curation, but could also reflect that the curated corpus happened to contain content well-suited to ESAP questions. Future work comparing systems under matched retrieval conditions would help isolate the contribution of corpus curation versus retrieval architecture. Fifth, we evaluated a limited set of baseline systems and configurations; results may differ for other clinical AI products, alternative prompting/browsing protocols, or future model releases. Sixth, ESAP 2025 is a proprietary assessment and the original question text cannot be shared publicly, which limits direct reproducibility. Finally, both frontier model behavior and web retrieval toolchains can change over time (âtool driftâ), which may affect absolute baseline performance; we therefore report evaluation dates and protocols to support fair longitudinal interpretation. The evidence corpus is currently U.S.-centric, reflecting American society guidelines; application in other healthcare systems with different standards would require corpus adaptation. 4.4 Future Directions This work establishes a foundation for several directions of ongoing investigation. We are extending the benchmark evaluation to additional subspecialties, including cardiology, nephrology, and general internal medicine, to assess generalization of the evidence-grounded approach. We are also conducting prospective evaluation with subspecialist clinicians using real-world case vignettes to assess clinical utility beyond examination accuracy. These studies will address whether board examination performance translates to meaningful decision support in practice. 5 Conclusions Mirror, an evidence-grounded clinical reasoning system, achieved 87.5% accuracy on a subspecialty endocrinology board examination, exceeding frontier LLMs by 12.9 percentage points despite those models having web access that Mirror lacked. Critically, Mirror provided complete evidence traceability, 74.2% of outputs cited at least one guideline-tier source with 100% verified citation accuracy, enabling clinicians to verify every recommendation against primary evidence. These results demonstrate that curated, high-quality evidence bases with explicit provenance can outperform unconstrained web retrieval for subspecialty clinical reasoning while simultaneously addressing the auditability requirements essential for clinical deployment. Data Availability The ESAP 2025 questions are copyrighted by the Endocrine Society and cannot be redistributed; evaluation materials are therefore not publicly released. Code for statistical analyses is available upon request. Conflicts of Interest Authors A.H., R.Z., U.M., and G.S. are employees of January AI, which developed Mirror. Authors K.K. and N.A. serve as advisors to January AI. References Kung et al. [2023] Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health. 2023;2(2):e0000198. doi:10.1371/journal.pdig.0000198 Nori et al. [2023] Nori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on medical challenge problems. arXiv preprint arXiv:2303.13375. 2023. doi:10.48550/arXiv.2303.13375 ABIM [2024] American Board of Internal Medicine. Internal medicine and subspecialty certification examinations: 2020-2024 first-time taker pass rates. Available at: https://w.abim.org/media/5hhbskg2/certification-pass-rates.pdf. Accessed January 2026. Kohandel Gargari and Habibi [2025] Kohandel Gargari O, Habibi G. Enhancing medical AI with retrieval-augmented generation: A mini narrative review. Digital Health. 2025;11:20552076251337177. doi:10.1177/20552076251337177 Asgari et al. [2025] Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digital Medicine. 2025;8:100. doi:10.1038/s41746-025-01670-7 Roustan and Bastardot [2025] Roustan D, Bastardot F. The Cliniciansâ Guide to Large Language Models: A General Perspective With a Focus on Hallucinations. Interactive Journal of Medical Research. 2025;14:e59823. doi:10.2196/59823 Endocrine Society [2025] Endocrine Society. Endocrine Self-Assessment Program (ESAP) 2025. Available at: https://w.endocrine.org/store/selfassessmentmoc-tools/endocrine-selfassessment-program-esap-2025-book. Accessed January 2026. ADA [2024] American Diabetes Association Professional Practice Committee. Standards of Care in Diabetes, 2024. Diabetes Care. 2024;47(Suppl 1):S1-S321. Arnett et al. [2019] Arnett DK, Blumenthal RS, Khera A, et al. 2019 ACC/AHA Guideline on the Primary Prevention of Cardiovascular Disease. Circulation. 2019;140(11):e596-e646. KDIGO [2022] Kidney Disease: Improving Global Outcomes (KDIGO) Diabetes Work Group. KDIGO 2022 Clinical Practice Guideline for Diabetes Management in Chronic Kidney Disease. Kidney International. 2022;102(5S):S1-S127. Samson et al. [2023] Samson SL, Vellanki P, Engel S, et al. American Association of Clinical Endocrinology Consensus Statement: Comprehensive Type 2 Diabetes Management Algorithm, 2023 Update. Endocrine Practice. 2023;29(5):305-340. Rinella et al. [2023] Rinella ME, Lazarus JV, Ratziu V, et al. AASLD Practice Guidance on the clinical assessment and management of nonalcoholic fatty liver disease. Hepatology. 2023;77(5):1797-1835. Melmed et al. [2020] Melmed S, Auchus RJ, Goldfine AB, et al. Williams Textbook of Endocrinology. 14th ed. Philadelphia: Elsevier; 2020. Amugongo et al. [2025] Amugongo LM, Mascheroni P, Brooks S, Doering S, Seidel J. Retrieval augmented generation for large language models in healthcare: A systematic review. PLOS Digital Health. 2025;4(6):e0000877. doi:10.1371/journal.pdig.0000877 Lee et al. [2023] Lee P, Bubeck S, Petro J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. New England Journal of Medicine. 2023;388(13):1233-1239. doi:10.1056/NEJMsr2214184