Paper deep dive
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents
Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Guanyu Jiang, Ziru Niu, Zuozhu Liu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/27/2026, 4:02:17 AM
Summary
The paper introduces CDEG, a graph-based framework designed to improve long-horizon diagnostic agents by learning reusable, decision-critical evidence from historical trajectories. Unlike previous methods that reuse entire trajectories or memories, CDEG identifies specific evidence atoms that drive diagnostic decisions by contrasting successful and failed trajectories. It uses counterfactual interventions to validate whether specific evidence supports or blocks a diagnosis, organizing these validated relations into a structured graph. During inference, CDEG tracks the patient's evolving evidence state to retrieve relevant diagnostic relations, guiding the agent to acquire missing evidence or reappraise overlooked observations, resulting in significant accuracy improvements over vanilla agents.
Entities (7)
Relation Signals (6)
CDEG → evaluatedon → MIMIC-IDx
confidence 95% · We evaluate CDEG on the in-domain (ID) MIMIC-IDx... benchmarks
CDEG → evaluatedon → AgentClinic
confidence 95% · We evaluate CDEG on the... out-of-distribution (OOD) AgentClinic benchmarks
CDEG → constructs → Structured Graph
confidence 93% · organizes the resulting diagnosis--evidence--action relations into a structured graph.
CDEG → uses → Counterfactual Validation
confidence 92% · CDEG contrasts successful and failed trajectories... validates their diagnostic impact through controlled counterfactual interventions
CDEG → improves → Doctor Agent
confidence 90% · CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents.
Doctor Agent → failswhen → Missing Evidence
confidence 85% · existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.
Tags
Links
- Source: https://arxiv.org/abs/2608.22899v2
- Canonical: https://arxiv.org/abs/2608.22899v2
Trouble viewing inline? Open PDF directly →
Full Text
50,084 characters extracted from source content.
Expand or collapse full text
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents Xiwei Dai Zijie Meng Zhiting Fan Yixuan Tang Guanyu Jiang Ziru Niu Zuozhu Liu Abstract Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis–evidence–action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions. Zhejiang University Hangzhou, Zhejiang, China Figure 1: Long-horizon diagnosis with and without CDEG. The vanilla doctor agent may fail when decision-critical evidence is missing or ignored. With CDEG, the doctor agent is prompted to acquire missing evidence or reappraise observed evidence, enabling it to reach the correct diagnosis. Introduction Recent advances in medical language models have shifted diagnostic agents beyond static medical question answering toward long-horizon interactive diagnosis (Schmidgall et al. 2026; Almansoori et al. 2025; Li et al. 2024). Under partial observability, a doctor agent must acquire evidence from the patient environment through diagnostic tools, integrate it across turns, and decide when it is sufficient for diagnosis (National Academies of Sciences, Engineering, and Medicine 2015). Yet vanilla agents often fail to acquire relevant evidence or overlook observations already available (Long et al. 2026; Li et al. 2024; Hsu et al. 2026; Fang et al. 2026), as illustrated in Figure 1. Therefore, a fundamental challenge in long-horizon diagnosis is how to identify and utilize the evidence that is truly critical for distinguishing diagnoses. Historical diagnostic trajectories record evidence, actions, and final decisions, offering reusable experience for future interactions. Existing methods retrieve prior trajectories or distill them into reflections, strategies, and structured memories (Shinn et al. 2023; Zhao et al. 2024; Wang et al. 2025; Ren et al. 2026; Shen et al. 2026; Li et al. 2026; Han et al. 2026). However, trajectories mix decision-relevant evidence with incidental observations, while outcomes do not reveal which evidence shaped the diagnosis. Recent methods attempt to improve experience reuse through trajectory refinement, applicability-aware retrieval, and memory management but still operate mainly at the trajectory or memory-unit level (Ouyang et al. 2026; Xiong et al. 2026). Thus, individual observations may be reused without establishing their diagnostic relevance. Counterfactual reasoning offers a potential way to resolve this ambiguity (Richens et al. 2020; Nagesh et al. 2023; Xu et al. 2022). By changing one observation while keeping the remaining evidence unchanged, it tests whether diagnostic preference changes. We extend this principle to historical diagnostic trajectories and regard evidence as decision-critical only when a controlled intervention on its availability or use consistently changes diagnostic preference. This selects evidence for reuse by its effect on decisions, rather than its co-occurrence with a successful outcome. Building on this criterion, we introduce CDEG, a graph-based framework for learning reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts each failed trajectory with a similar successful trajectory from the same case, uses controlled counterfactual tests to determine whether candidate evidence supports or blocks the corresponding diagnostic revision, and organizes the resulting diagnosis–evidence–action relations into a structured graph. During inference, CDEG matches the patient’s evolving evidence state to the graph to guide acquisition of missing evidence or reappraisal of overlooked observations. We evaluate CDEG on the in-domain (ID) MIMIC-IDx and out-of-distribution (OOD) AgentClinic benchmarks across four doctor-agent backbones (Johnson et al. 2023a; Johnson et al. 2023b; Johnson et al. 2019; Schmidgall et al. 2026; Jin et al. 2021). CDEG improves average accuracy by 7.63% and 8.88% on the two benchmarks, respectively, with gains of up to 11.5%, consistently outperforming the vanilla agent and other experience-reuse methods. These results show that evidence-level learning provides a more reliable basis for experience reuse in long-horizon diagnosis. Our contributions are summarized as follows: • We formulate trajectory reuse as an evidence-level learning problem and introduce CDEG to extract decision-critical evidence from historical diagnostic trajectories. • We develop a counterfactual validation method that contrasts successful and failed trajectories from the same case and retains only evidence that demonstrably affects diagnostic preference under controlled interventions. • We construct a structured graph that uses the patient’s evolving evidence state to guide acquisition of missing evidence or reappraisal of overlooked evidence. • We evaluate CDEG on ID and OOD benchmarks with various backbones, which consistently improves the vanilla agent and achieves the best diagnostic performance. Related Work Long-Horizon Interactive Diagnosis. Long-horizon interactive diagnosis requires a doctor agent to acquire and integrate clinical evidence under partial observability before providing a final diagnosis. Interactive environments and benchmarks instantiate this setting through multi-turn patient interaction and diagnostic tools (Schmidgall et al. 2026; Almansoori et al. 2025; Jiang et al. 2025; Li et al. 2024). Recent methods improve evidence acquisition and diagnostic reasoning through targeted questioning, test selection, and iterative diagnosis (Qiu et al. 2025; Gao et al. 2026; Sanghvi et al. 2026; Hsu et al. 2026; Rose et al. 2025; Ren et al. 2025; Liu et al. 2025). Despite this progress, evidence acquisition remains challenging, and evidence already observed may still be ignored in the final diagnosis (Long et al. 2026; Fang et al. 2026). Learning from Diagnostic Trajectories. A complementary line of works learns from prior trajectories to improve subsequent decisions. General agent methods retrieve trajectories as exemplars or distill them into verbal reflections, transferable lessons, reusable workflows, skill libraries, or reasoning memories (Zheng et al. 2024; Shinn et al. 2023; Zhao et al. 2024; Wang et al. 2025; Wang et al. 2024; Ouyang et al. 2026). Medical agents similarly retrieve prior cases or distill diagnostic trajectories into strategies and structured memories (Ren et al. 2026; Shen et al. 2026; Li et al. 2026; Han et al. 2026). However, not all such experience is equally relevant or reliable for subsequent diagnosis, and its quality is not always explicitly assessed before reuse. Counterfactual Reasoning in Medical Diagnosis. Counterfactual reasoning examines how diagnostic preference changes under controlled interventions on clinical evidence. Prior medical works use counterfactual inference and evidence perturbation for causal diagnosis, clinical prediction, and model explanation (Richens et al. 2020; Nagesh et al. 2023; Xu et al. 2022). Recent medical-LLM studies further modify symptoms, laboratory evidence, or diagnostic contexts to evaluate diagnostic anchoring and reasoning over competing diagnoses (Bhasuran et al. 2026; Chen et al. 2026; Mo et al. 2026; You et al. 2026; Qu and Färber 2026). These works use counterfactual reasoning mainly for case-level diagnosis, explanation, or evaluation. Our work instead applies counterfactual validation during graph construction to assess candidate evidence mined from successful and failed trajectories before incorporating them into CDEG. Figure 2: Overview of CDEG. CDEG first compares successful and failed trajectories from the same case to mine candidates that may explain their different diagnoses. It then tests whether adding missing evidence, reconsidering observed evidence, or introducing evidence against a revision changes diagnostic preference, and consolidates the validated evidence with diagnoses and acquisition actions into a graph. At inference time, CDEG retrieves relevant diagnostic edges based on the current evidence and guides the doctor agent to acquire missing evidence, reconsider ignored evidence, or avoid unnecessary intervention. Methods Problem Setup Long-horizon interactive diagnosis models clinical diagnosis as a multi-turn sequential decision-making process under partial observability, in which a doctor agent progressively acquires and integrates clinical evidence before providing a final diagnosis (Schmidgall et al. 2026; Li et al. 2024). Each clinical case is represented as x=(ℰx,yx∗)x=(E_x,y_x^*), where ℰxE_x denotes the patient environment and yx∗y_x^* is the reference diagnosis. The environment contains the complete case information but initially reveals only o0o_0, with additional evidence becoming available through subsequent interaction. At interaction turn t, the doctor agent conditions on the interaction history hth_t and selects its next action ata_t: ht=(o0,(ai,oi)i=1t−1),at∼πθ(⋅∣ht).h_t= (o_0,(a_i,o_i)_i=1^t-1 ), a_t _θ(· h_t). (1) where πθ _θ denotes the doctor agent policy. The agent may interact with the patient, use diagnostic tools to review clinical history, obtain laboratory or imaging evidence, perform examinations, retrieve relevant knowledge, or provide a final diagnosis. Each action other than the final diagnostic action yields an observation oto_t, which is appended to the interaction history to form ht+1h_t+1 for the next turn. The interaction terminates when the agent executes the final diagnostic action aTa_T at turn T, yielding the trajectory: τ=(o0,(at,ot)t=1T−1,aT).τ= (o_0,(a_t,o_t)_t=1^T-1,a_T ). (2) Here, aTa_T is the final diagnostic action, and D(τ)D(τ) denotes the final diagnosis derived from the complete trajectory. Counterfactual Diagnostic Evidence Graph Construction Figure 2 provides an overview of CDEG construction and evidence-state intervention. CDEG construction proceeds in three stages. We first mine candidate evidence by comparing successful and failed trajectories from the same case, then validate each candidate by testing whether it changes diagnostic preference, and finally organize the validated evidence with its associated diagnoses and acquisition actions into the graph. Candidate Evidence Mining. For each case x in the graph-construction split, we sample K independent rollouts of the doctor agent with the identical patient environment ℰxE_x, yielding a set of trajectories x=τx(k)k=1KT_x=\ _x^(k)\_k=1^K. We map the final diagnosis from each rollout to the same ICD-aligned label space as the reference diagnosis and partition the trajectories by diagnostic correctness: x+=τ∈x∣D(τ)=yx∗,x−=τ∈x∣D(τ)≠yx∗. array[]lT_x^+=\τ _x D(τ)=y_x^*\,\\[2.0pt] T_x^-=\τ _x D(τ)≠ y_x^*\. array (3) With the patient case fixed across rollouts, we compare successful and failed trajectories to uncover evidence acquisition and utilization patterns that distinguish correct from incorrect diagnoses. However, trajectories from the same case may reach identical diagnoses through different interaction histories. We therefore pair each failed trajectory with the most similar successful trajectory based on the action–observation history τpre _pre before the final diagnostic action. For each failed trajectory τ−τ^-, we select τ+∗=argmaxτ+∈x+cos(ϕ(τpre−),ϕ(τpre+)),τ^+*= arg\,max_τ^+ _x^+ (φ(τ^-_pre),φ(τ^+_pre) ), (4) where ϕ(⋅)φ(·) maps the action–observation history into a semantic embedding. The pair (τ−,τ+∗)(τ^-,τ^+*) is retained only when this maximum similarity exceeds a predefined threshold. For each retained pair, we denote their final diagnoses as DsD_s and DtD_t, respectively. We further transform heterogeneous clinical observations in each retained pair to evidence atoms grounded in UMLS concepts (Bodenreider 2004). Each evidence atom represents a normalized clinical concept with its relevant attributes, including polarity, temporal status, numerical direction, and anatomical context. Observations mapped to the same evidence atom are merged while retaining their associated acquisition actions. Let EsE_s and EtE_t denote the evidence atoms extracted from the failed and successful trajectories, respectively. Based on their relationships, We derive three candidate evidence sets: M=Et∖Es,I=Et∩Es,B=Es∖Et.C^M=E_t E_s, ^I=E_t∩ E_s, ^B=E_s E_t. (5) where MC^M, IC^I, and BC^B denote missing-evidence, ignored-evidence, and blocking-evidence candidates, respectively. Evidence in MC^M is acquired in the successful trajectory but absent from the failed trajectory, representing potential decision-critical evidence supporting the target diagnosis. Evidence in IC^I appears in both trajectories, suggesting that the diagnostic discrepancy may stem from differences in evidence interpretation or utilization. Evidence in BC^B appears only in the failed trajectory and may support the source diagnosis, indicating when revision toward the target diagnosis is unwarranted. Counterfactual Validation. Building on these candidate sets, we next evaluate which one can alter diagnostic preference under a controlled counterfactual intervention. For each candidate, we apply only its corresponding intervention while keeping all other evidence unchanged. Given an evidence set, the verifier compares DsD_s and DtD_t and selects the diagnosis better supported by the available evidence. We measure whether the verifier’s preference changes after the intervention. To reduce order bias, all comparisons are repeated with the diagnosis order reversed, and a candidate is retained only if the same preference shift occurs in both orders. Accordingly, we perform different counterfactual interventions for the three candidate categories. (1) Completion test. For each e∈Me ^M, we add e to the failed-trajectory evidence EsE_s. If the verifier’s preference changes from DsD_s to DtD_t, e is retained as validated missing evidence, indicating its potential role in supporting the target diagnosis. (2) Reappraisal test. For each e∈Ie ^I, we keep EsE_s unchanged while explicitly exposing e during diagnosis comparison. If the verifier’s preference shifts from DsD_s to DtD_t, e is retained as validated ignored evidence, suggesting that the evidence was available but insufficiently utilized. (3) Blocking test. For each e∈Be ^B, we add e to the successful-trajectory evidence EtE_t. If the verifier’s preference alters from DtD_t to DsD_s, e is retained as validated blocking evidence, indicating that revision toward the target diagnosis may be inappropriate when this evidence is present. Graph Consolidation. We organize the counterfactually validated evidence, together with its associated diagnoses and acquisition actions, into the Counterfactual Diagnostic Evidence Graph: =(D,E,A,ℰAE,ℰED,ℰDD),G= (V_D,V_E,V_A,E_AE,E_ED,E_D ), (6) where DV_D, EV_E, and AV_A denote diagnosis, evidence, and acquisition-action nodes, respectively. ℰAEE_AE maps acquisition actions to their resulting evidence, ℰEDE_ED links evidence to the diagnoses it supports, and ℰDDE_D contains directed diagnostic-revision relations. Each (Ds,Dt)∈ℰDD(D_s,D_t) _D represents a revision from the diagnosis DsD_s in a failed trajectory to the diagnosis DtD_t in its paired successful trajectory for the same case. Such an edge is added in the graph only when counterfactual validation identifies evidence that shifts diagnostic preference from DsD_s to DtD_t. Validated missing and ignored evidence are linked to DtD_t, while validated blocking evidence is connected to DsD_s. Therefore, each diagnostic revision relation retains both the evidence that promotes correction toward DtD_t and the evidence that favors retaining DsD_s. Evidence-State Intervention At inference time, CDEG maintains an evolving patient evidence state tS_t, which contains graph evidence nodes corresponding to the observations acquired up to interaction turn t. Newly obtained observations from the patient environment are mapped to graph evidence nodes through hybrid matching, which combines exact and semantic similarity matching. The resulting evidence nodes are added to tS_t and used for future graph retrieval. Graph Retrieval. Given the current evidence state tS_t, CDEG retrieves the most related diagnosis-to-diagnosis edge. Specifically, for each (Ds,Dt)∈ℰDD(D_s,D_t) _D, let Es→tE_s→ t denote the evidence atoms associated with this edge. We measure the evidence coverage of the edge at turn t as: covt(Ds,Dt)=|t∩Es→t||Es→t|.cov_t(D_s,D_t)= |S_t∩ E_s→ t | |E_s→ t |. (7) Before the doctor agent attempts to provide a final diagnosis, CDEG evaluates all edges in ℰDDE_D and retrieves the one with the highest coverage if its score exceeds a predefined threshold. Once the doctor agent attempts to provide a final diagnosis, the proposed diagnosis is matched to a graph diagnosis node DcurD_cur. CDEG then performs a localized retrieval over outgoing edges (Dcur,Dt)∈ℰDD(D_cur,D_t) _D and selects the one with the highest coverage that exceeds the same threshold. Selective Intervention. For the retrieved edge (Ds,Dt)(D_s,D_t), CDEG first determines whether the diagnostic revision is applicable to the current evidence state tS_t. If evidence supporting DsD_s has already been observed, CDEG does not intervene along this edge. Otherwise, CDEG checks whether any evidence associated with DtD_t remains unobserved. If such evidence can be obtained through an available acquisition action, CDEG identifies a missing-evidence state and prompts the doctor agent to perform the corresponding acquisition action to obtain the missing evidence. If no further evidence can be acquired, CDEG considers diagnostic reappraisal when the doctor agent proposes DsD_s. It checks whether evidence associated with DtD_t is already present in tS_t. CDEG identifies an ignored-evidence state and prompts the doctor agent to reappraise the observed evidence before finalizing its diagnosis. After each non-final action, newly acquired evidence is added to tS_t, and graph retrieval is repeated. When the doctor agent attempts to provide a final diagnosis, CDEG evaluates whether evidence acquisition or diagnostic reappraisal is required. If an intervention is triggered, the final diagnostic action is deferred and the interaction continues; otherwise, the proposed diagnosis is returned as the final output. Experiments Experimental Setup Benchmarks. We evaluate CDEG on two long-horizon interactive diagnosis benchmarks. MIMIC-IDx serves as the ID benchmark and is constructed by integrating MIMIC-IV (Johnson et al. 2023a), MIMIC-IV-Note (Johnson et al. 2023b), and matched MIMIC-CXR records (Johnson et al. 2019). Each hospital admission is treated as a diagnostic case in which the patient environment initially exposes limited clinical information and additional evidence becomes available through subsequent interaction. Cases are partitioned at the patient level into disjoint graph-construction and evaluation sets. For OOD evaluation, we use the MedQA subset of AgentClinic (Schmidgall et al. 2026), which converts MedQA cases (Jin et al. 2021) into text-based interactive diagnostic scenarios. The CDEG is constructed solely from the MIMIC-IDx graph-construction set, with the held-out MIMIC-IDx set and AgentClinic used exclusively for evaluation. Additional benchmark construction details are provided in the supplementary material, and benchmark statistics are summarized in Table 1. Benchmark Source Setting # Cases MIMIC-IDx MIMIC-IV Graph Construction 573 ID Evaluation 200 AgentClinic MedQA OOD Evaluation 107 Table 1: Benchmark statistics and experimental settings. Baselines. We compare CDEG with the Vanilla Agent and three baselines. Self-Reflection prompts the doctor agent to reconsider its diagnostic reasoning from the current interaction history before providing a final diagnosis. Trajectory RAG retrieves similar successful historical diagnostic trajectories as in-context examples. Flat Experience retrieves diagnostic experience distilled from historical trajectories as unstructured text. For each backbone and benchmark, all methods use the same patient environment, available tools, and interaction budget. Historical information is restricted to the graph-construction split, and retrieval-based baselines use the same retrieval encoder and retrieval budget. Backbone Method MIMIC-IDx (ID) AgentClinic (OOD) Avg. Acc. (%) Sim. Corr. (%) Acc. (%) Sim. Corr. (%) Acc. (%) Sim. Corr. (%) Gemini-3-Flash- Preview Vanilla Agent 41.50 0.6584 – 60.75 0.8442 – 51.13 0.7513 – + Self-Reflection 47.00 0.6828 23.08 61.68 0.8515 21.43 54.34 0.7672 22.26 + Trajectory RAG 45.00 0.6726 22.22 60.75 0.7904 30.95 52.88 0.7315 26.59 + Flat Experience 48.50 0.6688 26.50 61.68 0.8161 35.71 55.09 0.7425 31.11 + CDEG 53.00 0.6908 33.33 68.22 0.8576 45.24 60.61 0.7742 39.29 Gemini-3.1-Pro Vanilla Agent 50.00 0.6733 – 66.36 0.8691 – 58.18 0.7712 – + Self-Reflection 47.00 0.6720 19.00 66.36 0.8674 25.00 56.68 0.7697 22.00 + Trajectory RAG 52.50 0.6846 28.00 68.22 0.8690 30.56 60.36 0.7768 29.28 + Flat Experience 50.00 0.6752 25.00 67.29 0.8721 33.33 58.65 0.7737 29.17 + CDEG 59.00 0.7002 33.00 71.96 0.8924 36.11 65.48 0.7963 34.56 GPT-5.6-Luna Vanilla Agent 50.00 0.6640 – 51.40 0.7474 – 50.70 0.7057 – + Self-Reflection 43.00 0.6613 16.00 50.47 0.7397 17.31 46.74 0.7005 16.66 + Trajectory RAG 48.50 0.6718 22.00 47.66 0.7393 21.15 48.08 0.7056 21.58 + Flat Experience 53.00 0.6741 24.00 48.60 0.7340 25.00 50.80 0.7041 24.50 + CDEG 56.50 0.6942 33.00 62.62 0.7952 36.54 59.56 0.7447 34.77 Claude-Sonnet-5 Vanilla Agent 57.50 0.6666 – 54.21 0.7779 – 55.86 0.7223 – + Self-Reflection 58.00 0.6672 24.71 50.47 0.7531 24.49 54.24 0.7102 24.60 + Trajectory RAG 63.00 0.6772 32.94 55.14 0.7663 28.57 59.07 0.7218 30.76 + Flat Experience 60.50 0.6669 29.41 55.14 0.7629 22.45 57.82 0.7149 25.93 + CDEG 61.00 0.7035 30.59 65.42 0.8331 46.94 63.21 0.7683 38.77 Table 2: Diagnostic performance on MIMIC-IDx and AgentClinic. Avg. denotes the macro-average over the two benchmarks. The best and second-best results for each backbone are indicated by boldface and underlining. Metrics. We primarily evaluate diagnostic quality using two complementary metrics, Accuracy (Acc.) and Similarity (Sim.). Accuracy measures the proportion of predicted diagnoses that are clinically consistent with the reference diagnoses, which is determined by GPT-5.6-Sol (OpenAI 2026a) under a predefined clinical evaluation rubric (Schmidgall et al. 2026). Similarity captures the semantic proximity between predicted and reference diagnoses and is computed as cosine similarity in the embedding space of MedEmbed-large-v0.1 (Balachandran 2024). We further evaluate how each enhanced agent affects the decision behavior of the vanilla agent using Correction (Corr.) (Hassoon et al. 2026) and Regression (Reg.) (Fang et al. 2026). For each backbone, different methods are compared with the vanilla agent on the same cases. Let NV−N_V^- and NV+N_V^+ denote the numbers of cases diagnosed incorrectly and correctly by the vanilla agent, respectively. Let NfixN_fix denote the cases where the enhanced agent corrects errors made by the vanilla agent, and NregN_reg denote the cases where the enhanced agent changes correct predictions of the vanilla agent into incorrect ones. We define the two metrics as: Correction=NfixNV−,Regression=NregNV+.Correction= N_fixN_V^-, = N_regN_V^+. (8) Implementation Details. During graph construction, Gemini-3.1-Pro (Google 2026) and Gemini-3-Flash-Preview (Google 2025) were used to sample K=4K=4 independent trajectories for each case in the MIMIC-IDx graph-construction split. Qwen3-Embedding-0.6B (Zhang et al. 2025) was used as the trajectory encoder and as the retrieval encoder for all retrieval-based baselines. Each failed trajectory was paired with its most similar successful trajectory, and pairs were retained only when their cosine similarity exceeded 0.7. GPT-5.5 (OpenAI 2026b) was used to map clinical observations to UMLS-grounded evidence atoms and to verify candidate evidence through counterfactual validation. The resulting CDEG was frozen and shared across all doctor-agent backbones. During inference, we compared Gemini-3-Flash-Preview (Google 2025), Gemini-3.1-Pro (Google 2026), GPT-5.6-Luna (OpenAI 2026a), and Claude-Sonnet-5 (Anthropic 2026) as doctor agent backbones. Newly acquired observations are mapped to graph evidence nodes through hybrid matching, which combines exact matching with SapBERT-based semantic matching using a similarity threshold of 0.72 (Liu et al. 2021). Proposed diagnoses are mapped to graph diagnosis nodes using the same procedure with a threshold of 0.65. Diagnostic revision relations are retrieved when their evidence coverage exceeds 0.6. All methods use a maximum interaction budget of 30 turns. Additional prompts and implementation settings are provided in the supplementary material. Main Results As shown in Table 2, CDEG consistently improves diagnostic performance across all doctor agent backbones and both benchmarks. It increases the average accuracy across various backbones by 7.63% on MIMIC-IDx and 8.88% on AgentClinic. In contrast, Trajectory RAG and Flat Experience correct Vanilla Agent errors without producing consistent accuracy gains, suggesting that retrieving relevant historical experience alone is insufficient. The retained experience must also be reliable and applied under an appropriate evidence state. CDEG addresses both requirements by retaining evidence only when counterfactual intervention changes diagnostic preference and organizing the validated evidence into diagnostic revision relations for selective acquisition and reappraisal. Additionally, although CDEG is constructed only from Gemini-3-Flash-Preview and Gemini-3.1-Pro-Preview trajectories, it still improves GPT-5.6-Luna and Claude-Sonnet-5 on both benchmarks. Its gains further transfer to the OOD evaluation set AgentClinic, suggesting that the validated evidence captures reusable diagnostic relations rather than reasoning patterns specific to a particular doctor-agent backbone or benchmark. Ablation Studies To further analyze the contributions of individual components, we conduct ablation studies on counterfactual validation for structured experience, different candidate evidence sets, evidence-state intervention, and diagnostic experience scaling. All studies use the MIMIC-IDx evaluation set with Gemini-3-Flash-Preview as the doctor-agent backbone. Variant Acc. (%) Sim. Corr. (%) CDEG 53.00 0.6908 33.33 Counterfactual Validation for Structured Experience Flat Experience 48.50 0.6688 26.50 w/o Counterfactual Validation 43.50 0.6481 24.79 Different Candidate Evidence Sets w/o Missing Evidence 51.00 0.6865 25.64 w/o Ignored Evidence 49.00 0.6875 25.64 w/o Blocking Evidence 51.50 0.6904 30.77 Evidence Selective Intervention Acquisition Only 43.50 0.6702 19.66 Reappraisal Only 47.50 0.6769 29.91 Table 3: Ablation studies of CDEG on MIMIC-IDx using Gemini-3-Flash-Preview. CDEG is the full-model reference, and the best variant within each ablation group is underlined. Counterfactual Validation for Structured Experience As shown in Table 3, removing counterfactual validation reduces accuracy from 53.00% to 43.50% and even performs worse than Flat Experience at 48.50%. This result shows that CDEG’s gains do not arise from graph structuring alone. The graph organizes trajectory-derived experience into structured diagnostic relations for efficient storage and retrieval, while the key advantage of CDEG lies in using counterfactual validation to determine which candidate evidence should be consolidated into the graph. By retaining only evidence that changes diagnostic preference, CDEG excludes noisy or incidental trajectory differences and concentrates diagnostically useful information in the graph. Different Candidate Evidence Sets As shown in Table 3, removing either missing or ignored evidence reduces Correction from 33.33% to 25.64%, indicating that both are important for broadening the coverage of decision-critical evidence represented in the graph. Without either, less validated evidence is available across different evidence states, weakening CDEG’s guidance for diagnostic correction. Removing blocking evidence causes a smaller decline to 30.77%, suggesting that it helps CDEG recognize the boundary conditions of a diagnostic revision and reduce inappropriate interventions under incompatible evidence states. Evidence Selective Intervention As shown in Table 3, Reappraisal Only improves Correction by 10.25% over Acquisition Only, yet both underperform the complete CDEG. This pattern indicates that long-horizon diagnostic errors cannot be fully addressed by a single intervention. By dynamically selecting between acquisition and reappraisal according to the current evidence state, CDEG can capture correction opportunities missed by either intervention alone. Diagnostic Experience Scaling As shown in Figure 3, increasing graph-construction data from 10% to 100% raises Correction from 19.66% to 33.33%, with Accuracy and Similarity improving in parallel. This trend suggests that additional diagnostic trajectories broaden the coverage of counterfactually validated diagnostic revision relations, allowing CDEG to provide relevant evidence guidance for a wider range of subsequent cases. Figure 3: Effect of diagnostic experience scaling. Analysis Evidence States at Diagnostic Failure We audit 117 incorrect trajectories produced by the vanilla agent on MIMIC-IDx, and categorize each trajectory by its evidence sufficiency before final diagnostic commitment. The audit protocol is provided in the supplementary material. As shown in Table 4, only 11.97% of errors occur under conflicting evidence. In contrast, 50.43% arise from incomplete decision-critical evidence acquisition, while another 37.60% occur despite such evidence already being available but insufficiently incorporated into the diagnosis. Together, these two failure modes account for 88.03% of all errors, showing that reliable long-horizon diagnosis hinges on effective management of decision-critical evidence across acquisition and use. CDEG explicitly addresses both failures by using the current evidence state to dynamically guide the doctor agent to acquire missing evidence or reappraise overlooked evidence. Evidence Sufficiency # Cases Errors (%) Incomplete 59 50.43 Sufficient 44 37.60 Conflicting 14 11.97 Total 117 100.00 Table 4: Evidence states among incorrect trajectories produced by vanilla agent on MIMIC-IDx. Evidence-State Intervention and Diagnostic Stability As shown in Figure 4, CDEG achieves the lowest average Regression among the evaluated methods, with 17.70% on MIMIC-IDx and 14.62% on AgentClinic, while maintaining higher Correction exhibited in Table 2. This combination indicates that CDEG achieves a better balance between correcting diagnostic errors and preserving already correct diagnoses, rather than simply increasing the frequency of revision. This pattern is consistent with evidence-state intervention, which determines whether a retrieved diagnostic revision is applicable to the current evidence state before triggering acquisition or reappraisal. Figure 4: Regression rates across doctor-agent backbones on MIMIC-IDx and AgentClinic. Diagnostic Gains Beyond Longer Interaction As shown in Figure 5, CDEG improves diagnostic performance with only 1.41 and 1.27 additional turns on MIMIC-IDx and AgentClinic, indicating more efficient use of the interaction budget. Rather than indiscriminately extending the diagnostic process, CDEG uses the evidence state to selectively guide acquisition of missing evidence or reappraisal of observed evidence. This effect is particularly apparent for the GPT-5.6-Luna backbone on AgentClinic, where CDEG improves Accuracy by 11.22% while reducing the average interaction length by 0.71 turns, suggesting that it reduces ineffective exploration and reaches an evidence state sufficient for diagnosis more quickly. Tool-call counts exhibit the same pattern, further demonstrating that CDEG improves interaction efficiency through targeted evidence acquisition or reappraisal. Figure 5: Interaction turns and tool-call counts across doctor-agent backbones on MIMIC-IDx and AgentClinic. Conclusion We presented CDEG, a graph-based framework that learns reusable decision-critical evidence by contrasting successful and failed diagnostic trajectories, validating candidate evidence through controlled counterfactual interventions, and organizing validated evidence, diagnoses, and actions into a structured graph that guides acquisition or reappraisal according to the evolving patient evidence state. Across ID and OOD benchmarks with four doctor-agent backbones, CDEG consistently improves vanilla agent diagnostic performance, and achieves the best performance. In the future, we will explore online graph evolution for continual experience validation and extend CDEG to broader clinical scenarios. References Almansoori et al. (2025) M. Almansoori, K. Kumar, and H. Cholakkal MedAgentSim: self-evolving multi-agent simulations for realistic clinical interactions. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, Lecture Notes in Computer Science, Vol. 15968, p. 362–372. External Links: Document, Link Cited by: Introduction, Long-Horizon Interactive Diagnosis.. Anthropic (2026) Anthropic Claude sonnet 5 system card. Note: https://w.anthropic.com/claude-sonnet-5-system-cardAccessed: 2026-07-28 Cited by: Implementation Details.. Balachandran (2024) A. Balachandran MedEmbed: medical-focused embedding models. Note: https://github.com/abhinand5/MedEmbedMedEmbed-large-v0.1. Accessed: 2026-07-28 Cited by: Metrics.. Bhasuran et al. (2026) B. Bhasuran, M. Prosperi, K. Hanna, J. Petrilli, C. J. Washington, and Z. He Evaluation of causal reasoning for large language models in contextualized clinical scenarios of laboratory test interpretation. npj Digital Medicine 9 (1), p. 487. External Links: Document, Link Cited by: Counterfactual Reasoning in Medical Diagnosis.. Bodenreider (2004) O. Bodenreider The unified medical language system (UMLS): integrating biomedical terminology. Nucleic Acids Research 32, p. D267–D270. External Links: Document Cited by: Candidate Evidence Mining.. Chen et al. (2026) W. Chen, G. Huang, W. Wang, and Z. Zhu MedEinst: benchmarking the Einstellung effect in medical LLMs through counterfactual differential diagnosis. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, p. 39778–39798. External Links: Document, Link Cited by: Counterfactual Reasoning in Medical Diagnosis.. Fang et al. (2026) J. Fang, R. Chen, X. Yang, J. Yu, J. Xu, A. Vinod, W. Shi, T. Chen, H. Ji, C. Zhai, Y. Ding, and Y. Zhang Benchmarking multi-turn medical diagnosis: hold, lure, and self-correction. External Links: 2604.04325, Link Cited by: Introduction, Long-Horizon Interactive Diagnosis., Metrics.. Gao et al. (2026) Y. Gao, X. Zhou, Y. Li, Y. Zhao, and R. Liu MedExAgent: training LLM agents to ask, examine, and diagnose in noisy clinical environments. External Links: 2605.07058, Link Cited by: Long-Horizon Interactive Diagnosis.. Google (2025) Google Gemini 3 flash preview. Note: https://ai.google.dev/gemini-api/docs/models/gemini-3-flash-previewAccessed: 2026-07-28 Cited by: Implementation Details., Implementation Details.. Google (2026) Google Gemini 3.1 pro preview. Note: https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-previewAccessed: 2026-07-28 Cited by: Implementation Details., Implementation Details.. Han et al. (2026) X. Han, Y. Fan, S. Zhao, H. Wang, and B. Qin GSEM: graph-based self-evolving memory for experience augmented clinical reasoning. External Links: 2603.22096, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Hassoon et al. (2026) A. Hassoon, X. Peng, R. Irimia, A. Lianjie, H. Leo, A. Bandeira, H. Y. (. Woo, M. Dredze, R. Abdulnour, K. M. McDonald, S. Peterson, and D. Newman-Toker Evaluating the AI potential as a safety net for diagnosis: a novel benchmark of large language models in correcting diagnostic errors. Note: medRxiv preprintVersion 1 External Links: Document, Link Cited by: Metrics.. Hsu et al. (2026) H. Hsu, Z. Wang, D. Zhang, N. Chen, J. Wang, J. Ding, C. Hsu, G. Wang, F. Liu, F. Hung, C. Wu, and L. Shen MedAction: towards active multi-turn clinical diagnostic LLMs. External Links: 2605.07305, Link Cited by: Introduction, Long-Horizon Interactive Diagnosis.. Jiang et al. (2025) Y. Jiang, K. C. Black, G. Geng, D. Park, J. Zou, A. Y. Ng, and J. H. Chen MedAgentBench: a virtual ehr environment to benchmark medical llm agents. NEJM AI 2 (9), p. AIdbp2500144. External Links: Document, Link Cited by: Long-Horizon Interactive Diagnosis.. Jin et al. (2021) D. Jin, E. Pan, N. Oufattole, W. Weng, H. Fang, and P. Szolovits What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11 (14), p. 6421. External Links: Document Cited by: Introduction, Benchmarks.. Johnson et al. (2023a) A. E. W. Johnson, L. Bulgarelli, L. Shen, A. Gayles, A. Shammout, S. Horng, T. J. Pollard, S. Hao, B. Moody, B. Gow, L. H. Lehman, L. A. Celi, and R. G. Mark MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data 10, p. 1. External Links: Document Cited by: Introduction, Benchmarks.. Johnson et al. (2019) A. E. W. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C. Deng, R. G. Mark, and S. Horng MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data 6, p. 317. External Links: Document Cited by: Introduction, Benchmarks.. Johnson et al. (2023b) A. Johnson, T. Pollard, S. Horng, L. A. Celi, and R. Mark MIMIC-IV-Note: deidentified free-text clinical notes. PhysioNet. Note: Version 2.2 External Links: Document Cited by: Introduction, Benchmarks.. Li et al. (2026) B. Li, S. Du, and Y. Guo Joint optimization of reasoning and dual-memory for self-learning diagnostic agent. External Links: 2604.07269, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Li et al. (2024) S. S. Li, V. Balachandran, S. Feng, J. S. Ilgen, E. Pierson, P. W. Koh, and Y. Tsvetkov MediQ: question-asking LLMs and a benchmark for reliable interactive clinical reasoning. In Advances in Neural Information Processing Systems, Vol. 37, p. 28858–28888. External Links: Document, Link Cited by: Introduction, Long-Horizon Interactive Diagnosis., Problem Setup. Liu et al. (2021) F. Liu, E. Shareghi, Z. Meng, M. Basaldella, and N. Collier Self-alignment pretraining for biomedical entity representations. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, p. 4228–4238. External Links: Document, Link Cited by: Implementation Details.. Liu et al. (2025) X. Liu, D. Sun, Y. Fung, D. Hakkani-Tur, and T. F. Abdelzaher DocCHA: towards LLM-augmented interactive online diagnosis system. In Proceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Avignon, France, p. 609–619. External Links: Link Cited by: Long-Horizon Interactive Diagnosis.. Long et al. (2026) Z. Long, Z. Bao, and Z. Wei Strong reasoning isn’t enough: evaluating evidence elicitation in interactive diagnosis. External Links: 2601.19773, Link Cited by: Introduction, Long-Horizon Interactive Diagnosis.. Mo et al. (2026) K. Mo, S. Venkatayogi, C. Shaib, R. Kouzy, W. Xu, B. C. Wallace, and J. J. Li Faithfulness vs. safety: evaluating LLM behavior under counterfactual medical evidence. In Findings of the Association for Computational Linguistics: ACL 2026, San Diego, California, United States, p. 37053–37081. External Links: Document, Link Cited by: Counterfactual Reasoning in Medical Diagnosis.. Nagesh et al. (2023) S. Nagesh, N. Mishra, Y. Naamad, J. M. Rehg, M. A. Shah, and A. Wagner Explaining a machine learning decision to physicians via counterfactuals. In Proceedings of the Conference on Health, Inference, and Learning, Proceedings of Machine Learning Research, Vol. 209, p. 556–577. External Links: Link Cited by: Introduction, Counterfactual Reasoning in Medical Diagnosis.. National Academies of Sciences, Engineering, and Medicine (2015) National Academies of Sciences, Engineering, and Medicine Improving diagnosis in health care. The National Academies Press, Washington, DC. External Links: Document, Link Cited by: Introduction. OpenAI (2026a) OpenAI GPT-5.6: frontier intelligence that scales with your ambition. Note: https://openai.com/index/gpt-5-6/Accessed: 2026-07-28 Cited by: Metrics., Implementation Details.. OpenAI (2026b) OpenAI Introducing gpt-5.5. Note: https://openai.com/index/introducing-gpt-5-5/Accessed: 2026-07-28 Cited by: Implementation Details.. Ouyang et al. (2026) S. Ouyang, J. Yan, I. Hsu, Y. Chen, K. Jiang, Z. Wang, R. Han, L. T. Le, S. Daruki, X. Tang, V. Tirumalashetty, G. Lee, M. Rofouei, H. Lin, J. Han, C. Lee, and T. Pfister ReasoningBank: scaling agent self-evolving with reasoning memory. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Qiu et al. (2025) P. Qiu, C. Wu, J. Liu, Q. Zheng, Y. Liao, H. Wang, Y. Yue, Q. Fan, S. Zhen, J. Wang, J. Gu, Y. Wang, Y. Zhang, and W. Xie Evolving interactive diagnostic agents in a virtual clinical environment. External Links: 2510.24654, Link Cited by: Long-Horizon Interactive Diagnosis.. Qu and Färber (2026) Z. Qu and M. Färber MediEval: a unified medical benchmark for patient-contextual and knowledge-grounded reasoning in LLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, p. 16150–16164. External Links: Document, Link Cited by: Counterfactual Reasoning in Medical Diagnosis.. Ren et al. (2026) R. Ren, Y. Wang, Y. Liang, L. Luo, J. Liu, H. Wang, C. Feng, Y. Zhang, C. Miao, J. Wen, and W. X. Zhao Emulating clinician cognition via self-evolving deep clinical research. External Links: 2603.10677, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Ren et al. (2025) W. Ren, T. Zhao, L. Wang, T. Wang, and V. G. Honavar DiaLLMs: EHR-enhanced clinical conversational system for clinical test recommendation and diagnosis prediction. In Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, p. 25622–25635. External Links: Document, Link Cited by: Long-Horizon Interactive Diagnosis.. Richens et al. (2020) J. G. Richens, C. M. Lee, and S. Johri Improving the accuracy of medical diagnosis with causal machine learning. Nature Communications 11 (1), p. 3923. External Links: Document, Link Cited by: Introduction, Counterfactual Reasoning in Medical Diagnosis.. Rose et al. (2025) D. P. Rose, C. Hung, M. Lepri, I. Alqassem, K. Gashteovski, and C. Lawrence MEDDxAgent: a unified modular agent framework for explainable automatic differential diagnosis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, p. 13803–13826. External Links: Document, Link Cited by: Long-Horizon Interactive Diagnosis.. Sanghvi et al. (2026) A. Sanghvi, N. Akash, R. Imam, A. Sharma, and M. Jain MeDxAgent: multi-agent consultation for interactive medical diagnosis. External Links: 2606.03416, Link Cited by: Long-Horizon Interactive Diagnosis.. Schmidgall et al. (2026) S. Schmidgall, R. Ziaei, C. Harris, J. W. Kim, E. P. Reis, J. Jopling, and M. Moor AgentClinic: a multimodal benchmark for tool-using clinical AI agents. npj Digital Medicine 9 (1), p. 499. External Links: Document, Link Cited by: Introduction, Introduction, Long-Horizon Interactive Diagnosis., Problem Setup, Benchmarks., Metrics.. Shen et al. (2026) W. Shen, B. Jian, J. Li, C. Liu, J. Moll, X. Hu, D. Rueckert, H. B. Li, and J. Pan Evo-MedAgent: beyond one-shot diagnosis with agents that remember, reflect, and improve. External Links: 2604.14475, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Shinn et al. (2023) N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems, Vol. 36, p. 8634–8652. External Links: Document, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Wang et al. (2024) G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, and A. Anandkumar Voyager: an open-ended embodied agent with large language models. Transactions on Machine Learning Research. External Links: Link Cited by: Learning from Diagnostic Trajectories.. Wang et al. (2025) Z. Z. Wang, J. Mao, D. Fried, and G. Neubig Agent workflow memory. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, p. 63897–63911. External Links: Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Xiong et al. (2026) Z. Xiong, Y. Lin, W. Xie, P. He, Z. Liu, J. Tang, H. Lakkaraju, and Z. Xiang How memory management impacts LLM agents: an empirical study of experience-following behavior. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), San Diego, California, United States, p. 623–645. External Links: Document, Link Cited by: Introduction. Xu et al. (2022) R. Xu, Y. Yu, C. Zhang, M. K. Ali, J. C. Ho, and C. Yang Counterfactual and factual reasoning over hypergraphs for interpretable clinical predictions on ehr. In Proceedings of the 2nd Machine Learning for Health Symposium, Proceedings of Machine Learning Research, Vol. 193, p. 259–278. External Links: Link Cited by: Introduction, Counterfactual Reasoning in Medical Diagnosis.. You et al. (2026) Z. You, X. Chen, A. Vashishtha, S. Du, G. Erion-Barner, H. Mei, H. Peng, and Y. Guo Improving clinical diagnosis with counterfactual multi-agent reasoning. External Links: 2603.27820, Link Cited by: Counterfactual Reasoning in Medical Diagnosis.. Zhang et al. (2025) Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou Qwen3 Embedding: advancing text embedding and reranking through foundation models. External Links: 2506.05176, Link Cited by: Implementation Details.. Zhao et al. (2024) A. Zhao, D. Huang, Q. Xu, M. Lin, Y. Liu, and G. Huang ExpeL: LLM agents are experiential learners. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence, Vancouver, Canada, p. 19632–19642. External Links: Document, Link Cited by: Introduction, Learning from Diagnostic Trajectories.. Zheng et al. (2024) L. Zheng, R. Wang, X. Wang, and B. An Synapse: trajectory-as-exemplar prompting with memory for computer control. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: Learning from Diagnostic Trajectories..