Paper deep dive
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents
Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic approaches attempt to address these failures by reusing historical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may contain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG contrasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic impact through controlled counterfactual interventions, and organizes the resulting diagnosis--evidence--action relations into a structured graph. During inference, CDEG tracks the evolving patient evidence state to retrieve relevant diagnostic relations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out-of-distribution benchmarks with multiple doctor agent backbones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagnosis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions.
Tags
Links
- Source: https://arxiv.org/abs/2608.22899v1
- Canonical: https://arxiv.org/abs/2608.22899v1
Trouble viewing inline? Open PDF directly →
Full Text
48,997 characters extracted from source content.
Expand or collapse full text
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents Xiwei Dai ∗ , Zijie Meng ∗ , Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu † Zhejiang University Hangzhou, Zhejiang, China Abstract Unlike static medical question answering, long-horizon di- agnosis captures the sequential nature of clinical practice: evidence is progressively acquired, integrated, and evaluated over multiple rounds of interaction before reaching a final diagnosis. However, existing doctor agents often fail when critical evidence is either not acquired or not adequately incorporated into diagnostic reasoning. Recent agentic ap- proaches attempt to address these failures by reusing histor- ical trajectories or distilled memories. But their diagnostic gains remain constrained because such experience may con- tain noisy or incidental information and is typically reused without validating which evidence actually drives diagnostic decisions. To address this limitation, we introduce CDEG, a graph-based framework that learns reusable decision-critical evidence from historical diagnostic trajectories. CDEG con- trasts successful and failed trajectories from the same case to identify candidate evidence, validates their diagnostic im- pact through controlled counterfactual interventions, and or- ganizes the resulting diagnosis–evidence–action relations into a structured graph. During inference, CDEG tracks the evolv- ing patient evidence state to retrieve relevant diagnostic re- lations and selectively guide missing evidence acquisition or overlooked evidence reappraisal. Across in-domain and out- of-distribution benchmarks with multiple doctor agent back- bones, CDEG consistently improves diagnostic performance, achieving up to an 11.5% accuracy gain over vanilla agents. These results demonstrate that reliable long-horizon diagno- sis requires moving beyond trajectory-level experience reuse toward evidence-level learning of the factors that truly shape clinical decisions. Introduction Recent advances in medical language models have shifted diagnostic agents beyond static medical question answering toward long-horizon interactive diagnosis (Schmidgall et al. 2026; Almansoori, Kumar, and Cholakkal 2025; Li et al. 2024). Under partial observability, a doctor agent must ac- quire evidence from the patient environment through diag- nostic tools, integrate it across turns, and decide when it is sufficient for diagnosis (National Academies of Sciences, En- gineering, and Medicine 2015). Yet vanilla agents often fail to acquire relevant evidence or overlook observations already available (Long, Bao, and Wei 2026; Li et al. 2024; Hsu et al. 2026; Fang et al. 2026), as illustrated in Figure 1. Therefore, ∗ These authors contributed equally. † Corresponding author. Vanilla Long-Horizon Diagnostic AgentLong-Horizon Diagnostic Agent with CDEG Missing evidence: Lung auscultation findings Action: Lung examination Diffuse expiratory wheezing and prolonged expiration. Perform a complete lung examination. Most likely COPD exacerbation. ... CDEG Doctor Doctor Do you have any past medical history? I have had COPD for 8 years and use home oxygen. I have had worsening shortness of breath for three days. Patient Doctor Patient Exam Ignored evidence: Pancytopenia and hypocellular marrow. Most likely aplastic anemia. CDEG Doctor WBC 1.8↓; Hgb 6.4↓; Plt 28↓. No blasts; markedly hypocellular marrow. ... I have felt weak for two weeks and bruise easily. Patient Order CBC, peripheral smear, and bone marrow examination. Doctor Lab Do you have any past medical history? SpO₂88%; R 28. I have had COPD for 8 years and use home oxygen. Likely gastroesophageal reflux disease. Missing Evidence Decision-critical evidence remains missing ... I have had worsening shortness of breath for three days. Patient Doctor Check vital signs and oxygen saturation. Doctor Doctor Patient Exam WBC 1.8↓; Hgb 6.4↓; Plt 28↓. No blasts; markedly hypocellular marrow. Likely immune thrombocytopenia. Ignore Evidence Decision-critical evidence is ignored ... I have felt weak for two weeks and bruise easily. Patient Order CBC, peripheral smear, and bone marrow examination. Doctor Doctor Lab Figure 1: Long-horizon diagnosis with and without CDEG. The vanilla doctor agent may fail when decision-critical evi- dence is missing or ignored. With CDEG, the doctor agent is prompted to acquire missing evidence or reappraise observed evidence, enabling it to reach the correct diagnosis. a fundamental challenge in long-horizon diagnosis is how to identify and utilize the evidence that is truly critical for distinguishing diagnoses. Historical diagnostic trajectories record evidence, actions, and final decisions, offering reusable experience for future interactions. Existing methods retrieve prior trajectories or distill them into reflections, strategies, and structured mem- ories (Shinn et al. 2023; Zhao et al. 2024; Wang et al. 2025; Ren et al. 2026; Shen et al. 2026; Li, Du, and Guo 2026; Han et al. 2026). However, trajectories mix decision-relevant ev- idence with incidental observations, while outcomes do not reveal which evidence shaped the diagnosis. Recent meth- ods attempt to improve experience reuse through trajectory refinement, applicability-aware retrieval, and memory man- agement but still operate mainly at the trajectory or memory- unit level (Ouyang et al. 2026; Xiong et al. 2026). Thus, individual observations may be reused without establishing their diagnostic relevance. Counterfactual reasoning offers a potential way to re- arXiv:2608.22899v1 [cs.AI] 24 Aug 2026 solve this ambiguity (Richens, Lee, and Johri 2020; Nagesh et al. 2023; Xu et al. 2022). By changing one observation while keeping the remaining evidence unchanged, it tests whether diagnostic preference changes. We extend this prin- ciple to historical diagnostic trajectories and regard evidence as decision-critical only when a controlled intervention on its availability or use consistently changes diagnostic prefer- ence. This selects evidence for reuse by its effect on decisions, rather than its co-occurrence with a successful outcome. Building on this criterion, we introduce CDEG, a graph- based framework for learning reusable decision-critical ev- idence from historical diagnostic trajectories. CDEG con- trasts each failed trajectory with a similar successful trajec- tory from the same case, uses controlled counterfactual tests to determine whether candidate evidence supports or blocks the corresponding diagnostic revision, and organizes the resulting diagnosis–evidence–action relations into a struc- tured graph. During inference, CDEG matches the patient’s evolving evidence state to the graph to guide acquisition of missing evidence or reappraisal of overlooked observations. We evaluate CDEG on the in-domain (ID) MIMIC-IDx and out-of-distribution (OOD) AgentClinic benchmarks across four doctor-agent backbones (Johnson et al. 2023b,a, 2019; Schmidgall et al. 2026; Jin et al. 2021). CDEG improves av- erage accuracy by 7.63% and 8.88% on the two benchmarks, respectively, with gains of up to 11.5%, consistently outper- forming the vanilla agent and other experience-reuse meth- ods. These results show that evidence-level learning provides a more reliable basis for experience reuse in long-horizon di- agnosis. Our contributions are summarized as follows: • We formulate trajectory reuse as an evidence-level learn- ing problem and introduce CDEG to extract decision- critical evidence from historical diagnostic trajectories. • We develop a counterfactual validation method that con- trasts successful and failed trajectories from the same case and retains only evidence that demonstrably affects diagnostic preference under controlled interventions. • We construct a structured graph that uses the patient’s evolving evidence state to guide acquisition of missing evidence or reappraisal of overlooked evidence. • We evaluate CDEG on ID and OOD benchmarks with var- ious backbones, which consistently improves the vanilla agent and achieves the best diagnostic performance. Related Work Long-Horizon Interactive Diagnosis. Long-horizon in- teractive diagnosis requires a doctor agent to acquire and integrate clinical evidence under partial observability be- fore providing a final diagnosis. Interactive environments and benchmarks instantiate this setting through multi-turn patient interaction and diagnostic tools (Schmidgall et al. 2026; Al- mansoori, Kumar, and Cholakkal 2025; Jiang et al. 2025; Li et al. 2024). Recent methods improve evidence acquisition and diagnostic reasoning through targeted questioning, test selection, and iterative diagnosis (Qiu et al. 2025; Gao et al. 2026; Sanghvi et al. 2026; Hsu et al. 2026; Rose et al. 2025; Ren et al. 2025; Liu et al. 2025). Despite this progress, evi- dence acquisition remains challenging, and evidence already observed may still be ignored in the final diagnosis (Long, Bao, and Wei 2026; Fang et al. 2026). Learning from Diagnostic Trajectories. A complemen- tary line of works learns from prior trajectories to improve subsequent decisions. General agent methods retrieve tra- jectories as exemplars or distill them into verbal reflections, transferable lessons, reusable workflows, skill libraries, or reasoning memories (Zheng et al. 2024; Shinn et al. 2023; Zhao et al. 2024; Wang et al. 2025, 2024; Ouyang et al. 2026). Medical agents similarly retrieve prior cases or distill diagnostic trajectories into strategies and structured memo- ries (Ren et al. 2026; Shen et al. 2026; Li, Du, and Guo 2026; Han et al. 2026). However, not all such experience is equally relevant or reliable for subsequent diagnosis, and its quality is not always explicitly assessed before reuse. Counterfactual Reasoning in Medical Diagnosis. Coun- terfactual reasoning examines how diagnostic preference changes under controlled interventions on clinical evidence. Prior medical works use counterfactual inference and ev- idence perturbation for causal diagnosis, clinical predic- tion, and model explanation (Richens, Lee, and Johri 2020; Nagesh et al. 2023; Xu et al. 2022). Recent medical-LLM studies further modify symptoms, laboratory evidence, or diagnostic contexts to evaluate diagnostic anchoring and rea- soning over competing diagnoses (Bhasuran et al. 2026; Chen et al. 2026; Mo et al. 2026; You et al. 2026; Qu and Färber 2026). These works use counterfactual reasoning mainly for case-level diagnosis, explanation, or evaluation. Our work instead applies counterfactual validation during graph con- struction to assess candidate evidence mined from successful and failed trajectories before incorporating them into CDEG. Methods Problem Setup Long-horizon interactive diagnosis models clinical diagnosis as a multi-turn sequential decision-making process under partial observability, in which a doctor agent progressively acquires and integrates clinical evidence before providing a final diagnosis (Schmidgall et al. 2026; Li et al. 2024). Each clinical case is represented asx = (E x ,y ∗ x ), whereE x denotes the patient environment and y ∗ x is the reference diagnosis. The environment contains the complete case information but initially reveals only o 0 , with additional evidence becoming available through subsequent interaction. At interaction turn t, the doctor agent conditions on the interaction history h t and selects its next action a t : h t = o 0 , (a i ,o i ) t−1 i=1 , a t ∼ π θ (·| h t ). (1) where π θ denotes the doctor agent policy. The agent may interact with the patient, use diagnostic tools to review clin- ical history, obtain laboratory or imaging evidence, perform examinations, retrieve relevant knowledge, or provide a final diagnosis. Each action other than the final diagnostic action yields an observationo t , which is appended to the interaction history to form h t+1 for the next turn. The interaction terminates when the agent executes the final diagnostic action a T at turn T, yielding the trajectory: τ = o 0 , (a t ,o t ) T−1 t=1 ,a T .(2) Evidence-State InterventionCounterfactual Diagnostic Evidence Graph (CDEG) Construction � � + � � − Doctor Agent Patient Environment Candidate evidence mining Diagnosis Normalization & Evidence Mapping ... D t ... D t ... D t Reappraisal reconsider �∈� � in E s D s D t Blocking add �∈� � to E t D t D s Completion add �∈� � to E s D s D t Counterfactual Validation ... � � Ignored-evidence candidates ... � � Missing-evidence candidates ... � � Blocking-evidence candidates Graph Retrieval E Observed E Unobserved Proposed Diagnosis D cur No Intervention Acquire Missing Evidence Reappraise Ignored Evidence Matching Evidence State S t ... E 1 E 3 E t Selective Intervention Evidence-State Matching History Imaging Lab Retrieve Exam ... ... D s ... D s ... D s � 2 � � � � � 1 � � 2 � � � � � 1 � � 2 � � � � � 1 � Graph Consolidation Assessment Node Types Edge Types A Acquisition-Action Node E Evidence Node D Diagnosis Node A–E Edge E–D Edge D–D Edge CDEG E t E s E t E s E t E s D t D s E 2 E 3 E 1 E t ... ... D cur D t A 1 A 2 A 3 A t Figure 2: Overview of CDEG. CDEG first compares successful and failed trajectories from the same case to mine candidates that may explain their different diagnoses. It then tests whether adding missing evidence, reconsidering observed evidence, or introducing evidence against a revision changes diagnostic preference, and consolidates the validated evidence with diagnoses and acquisition actions into a graph. At inference time, CDEG retrieves relevant diagnostic edges based on the current evidence and guides the doctor agent to acquire missing evidence, reconsider ignored evidence, or avoid unnecessary intervention. Here, a T is the final diagnostic action, and D(τ) denotes the final diagnosis derived from the complete trajectory. Counterfactual Diagnostic Evidence Graph Construction Figure 2 provides an overview of CDEG construction and evidence-state intervention. CDEG construction proceeds in three stages. We first mine candidate evidence by comparing successful and failed trajectories from the same case, then validate each candidate by testing whether it changes diag- nostic preference, and finally organize the validated evidence with its associated diagnoses and acquisition actions into the graph. Candidate Evidence Mining. For each case x in the graph-construction split, we sample K independent rollouts of the doctor agent with the identical patient environmentE x , yielding a set of trajectoriesT x =τ (k) x K k=1 . We map the fi- nal diagnosis from each rollout to the same ICD-aligned label space as the reference diagnosis and partition the trajectories by diagnostic correctness: T + x =τ ∈T x | D(τ) = y ∗ x , T − x =τ ∈T x | D(τ)̸= y ∗ x . (3) With the patient case fixed across rollouts, we compare suc- cessful and failed trajectories to uncover evidence acquisition and utilization patterns that distinguish correct from incor- rect diagnoses. However, trajectories from the same case may reach iden- tical diagnoses through different interaction histories. We therefore pair each failed trajectory with the most similar successful trajectory based on the action–observation his- tory τ pre before the final diagnostic action. For each failed trajectory τ − , we select τ +∗ = arg max τ + ∈T + x cos φ(τ − pre ),φ(τ + pre ) ,(4) where φ(·) maps the action–observation history into a se- mantic embedding. The pair (τ − ,τ +∗ ) is retained only when this maximum similarity exceeds a predefined threshold. For each retained pair, we denote their final diagnoses as D s and D t , respectively. We further transform heterogeneous clinical observations in each retained pair to evidence atoms grounded in UMLS concepts (Bodenreider 2004). Each evidence atom represents a normalized clinical concept with its relevant attributes, in- cluding polarity, temporal status, numerical direction, and anatomical context. Observations mapped to the same evi- dence atom are merged while retaining their associated ac- quisition actions. Let E s and E t denote the evidence atoms extracted from the failed and successful trajectories, respectively. Based on their relationships, We derive three candidate evidence sets: C M = E t s , C I = E t ∩E s , C B = E s t . (5) where C M , C I , and C B denote missing-evidence, ignored- evidence, and blocking-evidence candidates, respectively. Evidence in C M is acquired in the successful trajectory but absent from the failed trajectory, representing potential decision-critical evidence supporting the target diagnosis. Evidence inC I appears in both trajectories, suggesting that the diagnostic discrepancy may stem from differences in ev- idence interpretation or utilization. Evidence inC B appears only in the failed trajectory and may support the source diag- nosis, indicating when revision toward the target diagnosis is unwarranted. Counterfactual Validation. Building on these candidate sets, we next evaluate which one can alter diagnostic pref- erence under a controlled counterfactual intervention. For each candidate, we apply only its corresponding intervention while keeping all other evidence unchanged. Given an evidence set, the verifier compares D s and D t and selects the diagnosis better supported by the available ev- idence. We measure whether the verifier’s preference changes after the intervention. To reduce order bias, all comparisons are repeated with the diagnosis order reversed, and a candi- date is retained only if the same preference shift occurs in both orders. Accordingly, we perform different counterfac- tual interventions for the three candidate categories. (1) Completion test. For each e ∈ C M , we add e to the failed-trajectory evidence E s . If the verifier’s preference changes from D s to D t , e is retained as validated missing evidence, indicating its potential role in supporting the target diagnosis. (2) Reappraisal test. For each e ∈ C I , we keep E s un- changed while explicitly exposing e during diagnosis com- parison. If the verifier’s preference shifts from D s to D t , e is retained as validated ignored evidence, suggesting that the evidence was available but insufficiently utilized. (3) Blocking test. For each e ∈ C B , we add e to the successful-trajectory evidence E t . If the verifier’s preference alters from D t to D s , e is retained as validated blocking evidence, indicating that revision toward the target diagnosis may be inappropriate when this evidence is present. Graph Consolidation. We organize the counterfactually validated evidence, together with its associated diagnoses and acquisition actions, into the Counterfactual Diagnostic Evidence Graph: G = (V D ,V E ,V A ,E AE ,E ED ,E D ),(6) where V D , V E , and V A denote diagnosis, evidence, and acquisition-action nodes, respectively.E AE maps acquisition actions to their resulting evidence,E ED links evidence to the diagnoses it supports, andE D contains directed diagnostic- revision relations. Each (D s ,D t ) ∈ E D represents a re- vision from the diagnosis D s in a failed trajectory to the diagnosis D t in its paired successful trajectory for the same case. Such an edge is added in the graph only when coun- terfactual validation identifies evidence that shifts diagnostic preference fromD s toD t . Validated missing and ignored ev- idence are linked to D t , while validated blocking evidence is connected toD s . Therefore, each diagnostic revision relation retains both the evidence that promotes correction towardD t and the evidence that favors retaining D s . Evidence-State Intervention At inference time, CDEG maintains an evolving patient evi- dence stateS t , which contains graph evidence nodes corre- sponding to the observations acquired up to interaction turn t. Newly obtained observations from the patient environment are mapped to graph evidence nodes through hybrid match- ing, which combines exact and semantic similarity matching. The resulting evidence nodes are added to S t and used for future graph retrieval. Graph Retrieval. Given the current evidence state S t , CDEG retrieves the most related diagnosis-to-diagnosis edge. Specifically, for each (D s ,D t ) ∈ E D , let E s→t de- note the evidence atoms associated with this edge. We mea- sure the evidence coverage of the edge at turn t as: cov t (D s ,D t ) = |S t ∩ E s→t | |E s→t | .(7) Before the doctor agent attempts to provide a final diagno- sis, CDEG evaluates all edges inE D and retrieves the one with the highest coverage if its score exceeds a predefined threshold. Once the doctor agent attempts to provide a final diagnosis, the proposed diagnosis is matched to a graph diag- nosis node D cur . CDEG then performs a localized retrieval over outgoing edges (D cur ,D t ) ∈ E D and selects the one with the highest coverage that exceeds the same threshold. Selective Intervention. For the retrieved edge (D s ,D t ), CDEG first determines whether the diagnostic revision is applicable to the current evidence stateS t . If evidence sup- porting D s has already been observed, CDEG does not in- tervene along this edge. Otherwise, CDEG checks whether any evidence associated withD t remains unobserved. If such evidence can be obtained through an available acquisition ac- tion, CDEG identifies a missing-evidence state and prompts the doctor agent to perform the corresponding acquisition ac- tion to obtain the missing evidence. If no further evidence can be acquired, CDEG considers diagnostic reappraisal when the doctor agent proposes D s . It checks whether evidence associated with D t is already present inS t . CDEG identifies an ignored-evidence state and prompts the doctor agent to reappraise the observed evidence before finalizing its diag- nosis. After each non-final action, newly acquired evidence is added toS t , and graph retrieval is repeated. When the doctor agent attempts to provide a final diagnosis, CDEG evaluates whether evidence acquisition or diagnostic reappraisal is re- quired. If an intervention is triggered, the final diagnostic action is deferred and the interaction continues; otherwise, the proposed diagnosis is returned as the final output. Experiments Experimental Setup Benchmarks. We evaluate CDEG on two long-horizon in- teractive diagnosis benchmarks. MIMIC-IDx serves as the ID benchmark and is constructed by integrating MIMIC- IV (Johnson et al. 2023b), MIMIC-IV-Note (Johnson et al. 2023a), and matched MIMIC-CXR records (Johnson et al. 2019). Each hospital admission is treated as a diagnostic case Benchmark SourceSetting# Cases MIMIC-IDx MIMIC-IV Graph Construction573 ID Evaluation200 AgentClinic MedQAOOD Evaluation107 Table 1: Benchmark statistics and experimental settings. in which the patient environment initially exposes limited clinical information and additional evidence becomes avail- able through subsequent interaction. Cases are partitioned at the patient level into disjoint graph-construction and eval- uation sets. For OOD evaluation, we use the MedQA sub- set of AgentClinic (Schmidgall et al. 2026), which converts MedQA cases (Jin et al. 2021) into text-based interactive diagnostic scenarios. The CDEG is constructed solely from the MIMIC-IDx graph-construction set, with the held-out MIMIC-IDx set and AgentClinic used exclusively for evalua- tion. Additional benchmark construction details are provided in the supplementary material, and benchmark statistics are summarized in Table 1. Baselines. We compare CDEG with the Vanilla Agent and three baselines. Self-Reflection prompts the doctor agent to reconsider its diagnostic reasoning from the current inter- action history before providing a final diagnosis. Trajectory RAG retrieves similar successful historical diagnostic tra- jectories as in-context examples. Flat Experience retrieves diagnostic experience distilled from historical trajectories as unstructured text. For each backbone and benchmark, all methods use the same patient environment, available tools, and interaction budget. Historical information is restricted to the graph-construction split, and retrieval-based baselines use the same retrieval encoder and retrieval budget. Metrics. We primarily evaluate diagnostic quality using two complementary metrics, Accuracy (Acc.) and Similarity (Sim.). Accuracy measures the proportion of predicted diag- noses that are clinically consistent with the reference diag- noses, which is determined by GPT-5.6-Sol (OpenAI 2026a) under a predefined clinical evaluation rubric (Schmidgall et al. 2026). Similarity captures the semantic proximity be- tween predicted and reference diagnoses and is computed as cosine similarity in the embedding space of MedEmbed- large-v0.1 (Balachandran 2024). We further evaluate how each enhanced agent affects the decision behavior of the vanilla agent using Correction (Corr.) (Hassoon et al. 2026) and Regression (Reg.) (Fang et al. 2026). For each backbone, different methods are com- pared with the vanilla agent on the same cases. Let N − V and N + V denote the numbers of cases diagnosed incorrectly and correctly by the vanilla agent, respectively. Let N fix denote the cases where the enhanced agent corrects errors made by the vanilla agent, and N reg denote the cases where the en- hanced agent changes correct predictions of the vanilla agent into incorrect ones. We define the two metrics as: Correction = N fix N − V ,Regression = N reg N + V . (8) Implementation Details. During graph construction, Gemini-3.1-Pro (Google 2026) and Gemini-3-Flash- Preview (Google 2025) were used to sample K = 4 inde- pendent trajectories for each case in the MIMIC-IDx graph- construction split. Qwen3-Embedding-0.6B (Zhang et al. 2025) was used as the trajectory encoder and as the retrieval encoder for all retrieval-based baselines. Each failed trajec- tory was paired with its most similar successful trajectory, and pairs were retained only when their cosine similarity ex- ceeded 0.7. GPT-5.5 (OpenAI 2026b) was used to map clin- ical observations to UMLS-grounded evidence atoms and to verify candidate evidence through counterfactual valida- tion. The resulting CDEG was frozen and shared across all doctor-agent backbones. During inference, we compared Gemini-3-Flash- Preview (Google 2025), Gemini-3.1-Pro (Google 2026), GPT-5.6-Luna (OpenAI 2026a), and Claude-Sonnet-5 (An- thropic 2026) as doctor agent backbones. Newly ac- quired observations are mapped to graph evidence nodes through hybrid matching, which combines exact matching with SapBERT-based semantic matching using a similarity threshold of 0.72 (Liu et al. 2021). Proposed diagnoses are mapped to graph diagnosis nodes using the same proce- dure with a threshold of 0.65. Diagnostic revision relations are retrieved when their evidence coverage exceeds 0.6. All methods use a maximum interaction budget of 30 turns. Ad- ditional prompts and implementation settings are provided in the supplementary material. Main Results As shown in Table 2, CDEG consistently improves diag- nostic performance across all doctor agent backbones and both benchmarks. It increases the average accuracy across various backbones by 7.63% on MIMIC-IDx and 8.88% on AgentClinic. In contrast, Trajectory RAG and Flat Experi- ence correct Vanilla Agent errors without producing con- sistent accuracy gains, suggesting that retrieving relevant historical experience alone is insufficient. The retained ex- perience must also be reliable and applied under an appro- priate evidence state. CDEG addresses both requirements by retaining evidence only when counterfactual intervention changes diagnostic preference and organizing the validated evidence into diagnostic revision relations for selective ac- quisition and reappraisal. Additionally, although CDEG is constructed only from Gemini-3-Flash-Preview and Gemini- 3.1-Pro-Preview trajectories, it still improves GPT-5.6-Luna and Claude-Sonnet-5 on both benchmarks. Its gains further transfer to the OOD evaluation set AgentClinic, suggesting that the validated evidence captures reusable diagnostic re- lations rather than reasoning patterns specific to a particular doctor-agent backbone or benchmark. Ablation Studies To further analyze the contributions of individual compo- nents, we conduct ablation studies on counterfactual valida- tion for structured experience, different candidate evidence sets, evidence-state intervention, and diagnostic experience scaling. All studies use the MIMIC-IDx evaluation set with Gemini-3-Flash-Preview as the doctor-agent backbone. BackboneMethod MIMIC-IDx (ID)AgentClinic (OOD)Avg. Acc. (%) Sim. Corr. (%) Acc. (%) Sim. Corr. (%) Acc. (%) Sim. Corr. (%) Gemini-3-Flash- Preview Vanilla Agent41.50 0.6584–60.75 0.8442–51.13 0.7513– + Self-Reflection47.00 0.682823.0861.680.851521.4354.34 0.767222.26 + Trajectory RAG 45.00 0.6726 22.2260.75 0.7904 30.9552.88 0.7315 26.59 + Flat Experience 48.500.6688 26.5061.680.8161 35.7155.090.7425 31.11 + CDEG53.00 0.6908 33.3368.22 0.8576 45.2460.61 0.7742 39.29 Gemini-3.1-Pro Vanilla Agent50.00 0.6733–66.36 0.8691–58.18 0.7712– + Self-Reflection47.00 0.6720 19.0066.36 0.8674 25.0056.68 0.7697 22.00 + Trajectory RAG 52.500.684628.0068.220.8690 30.5660.360.776829.28 + Flat Experience 50.00 0.6752 25.0067.29 0.872133.3358.65 0.7737 29.17 + CDEG59.00 0.7002 33.0071.96 0.8924 36.1165.48 0.7963 34.56 GPT-5.6-Luna Vanilla Agent50.00 0.6640–51.400.7474–50.70 0.7057– + Self-Reflection43.00 0.6613 16.0050.47 0.7397 17.3146.74 0.7005 16.66 + Trajectory RAG 48.50 0.6718 22.0047.66 0.7393 21.1548.08 0.7056 21.58 + Flat Experience 53.00 0.674124.0048.60 0.7340 25.0050.800.7041 24.50 + CDEG56.50 0.6942 33.0062.62 0.7952 36.5459.56 0.7447 34.77 Claude-Sonnet-5 Vanilla Agent57.50 0.6666–54.21 0.7779–55.86 0.7223– + Self-Reflection58.00 0.6672 24.7150.47 0.7531 24.4954.24 0.7102 24.60 + Trajectory RAG 63.00 0.6772 32.9455.140.7663 28.5759.070.7218 30.76 + Flat Experience 60.50 0.6669 29.4155.140.7629 22.4557.82 0.7149 25.93 + CDEG61.00 0.7035 30.5965.42 0.8331 46.9463.21 0.7683 38.77 Table 2: Diagnostic performance on MIMIC-IDx and AgentClinic. Avg. denotes the macro-average over the two benchmarks. The best and second-best results for each backbone are indicated by boldface and underlining. VariantAcc. (%) Sim. Corr. (%) CDEG53.00 0.6908 33.33 Counterfactual Validation for Structured Experience Flat Experience48.500.668826.50 w/o Counterfactual Validation 43.50 0.6481 24.79 Different Candidate Evidence Sets w/o Missing Evidence51.00 0.6865 25.64 w/o Ignored Evidence49.00 0.6875 25.64 w/o Blocking Evidence51.500.690430.77 Evidence Selective Intervention Acquisition Only43.50 0.6702 19.66 Reappraisal Only47.500.676929.91 Table 3: Ablation studies of CDEG on MIMIC-IDx using Gemini-3-Flash-Preview. CDEG is the full-model reference, and the best variant within each ablation group is underlined. Counterfactual Validation for Structured Experience As shown in Table 3, removing counterfactual validation re- duces accuracy from 53.00% to 43.50% and even performs worse than Flat Experience at 48.50%. This result shows that CDEG’s gains do not arise from graph structuring alone. The graph organizes trajectory-derived experience into structured diagnostic relations for efficient storage and retrieval, while the key advantage of CDEG lies in using counterfactual val- idation to determine which candidate evidence should be consolidated into the graph. By retaining only evidence that changes diagnostic preference, CDEG excludes noisy or inci- dental trajectory differences and concentrates diagnostically useful information in the graph. Different Candidate Evidence Sets As shown in Table 3, removing either missing or ignored evidence reduces Cor- rection from 33.33% to 25.64%, indicating that both are im- portant for broadening the coverage of decision-critical evi- dence represented in the graph. Without either, less validated evidence is available across different evidence states, weak- ening CDEG’s guidance for diagnostic correction. Remov- ing blocking evidence causes a smaller decline to 30.77%, suggesting that it helps CDEG recognize the boundary con- ditions of a diagnostic revision and reduce inappropriate in- terventions under incompatible evidence states. Evidence Selective Intervention As shown in Table 3, Reappraisal Only improves Correction by 10.25% over Ac- quisition Only, yet both underperform the complete CDEG. This pattern indicates that long-horizon diagnostic errors cannot be fully addressed by a single intervention. By dy- namically selecting between acquisition and reappraisal ac- cording to the current evidence state, CDEG can capture correction opportunities missed by either intervention alone. Diagnostic Experience Scaling As shown in Figure 3, in- creasing graph-construction data from 10% to 100% raises Correction from 19.66% to 33.33%, with Accuracy and Sim- ilarity improving in parallel. This trend suggests that addi- tional diagnostic trajectories broaden the coverage of coun- terfactually validated diagnostic revision relations, allowing 10%50%100% 40 44 48 52 56 Accuracy (%) 44.50 48.00 53.00 a Accuracy 10%50%100% 0.63 0.65 0.67 0.69 0.71 Similarity 0.6526 0.6642 0.6908 b Similarity 10%50%100% 0 5 10 15 20 25 30 35 40 Correction (%) 19.66 27.35 33.33 c Correction Figure 3: Effect of diagnostic experience scaling. Evidence Sufficiency # Cases Errors (%) Incomplete5950.43 Sufficient4437.60 Conflicting1411.97 Total117100.00 Table 4: Evidence states among incorrect trajectories pro- duced by vanilla agent on MIMIC-IDx. CDEG to provide relevant evidence guidance for a wider range of subsequent cases. Analysis Evidence States at Diagnostic Failure We audit 117 in- correct trajectories produced by the vanilla agent on MIMIC- IDx, and categorize each trajectory by its evidence suffi- ciency before final diagnostic commitment. The audit proto- col is provided in the supplementary material. As shown in Table 4, only 11.97% of errors occur under conflicting evi- dence. In contrast, 50.43% arise from incomplete decision- critical evidence acquisition, while another 37.60% occur de- spite such evidence already being available but insufficiently incorporated into the diagnosis. Together, these two failure modes account for 88.03% of all errors, showing that reliable long-horizon diagnosis hinges on effective management of decision-critical evidence across acquisition and use. CDEG explicitly addresses both failures by using the current evi- dence state to dynamically guide the doctor agent to acquire missing evidence or reappraise overlooked evidence. Evidence-State Intervention and Diagnostic Stability As shown in Figure 4, CDEG achieves the lowest average Regression among the evaluated methods, with 17.70% on MIMIC-IDx and 14.62% on AgentClinic, while maintain- ing higher Correction exhibited in Table 2. This combina- tion indicates that CDEG achieves a better balance between correcting diagnostic errors and preserving already correct diagnoses, rather than simply increasing the frequency of revision. This pattern is consistent with evidence-state in- tervention, which determines whether a retrieved diagnostic revision is applicable to the current evidence state before triggering acquisition or reappraisal. Diagnostic Gains Beyond Longer Interaction As shown in Figure 5, CDEG improves diagnostic performance with only 1.41 and 1.27 additional turns on MIMIC-IDx and AgentClinic, indicating more efficient use of the interaction Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 0 5 10 15 20 25 30 35 Regression rate (%) 19.28 15.00 20.00 16.52 a MIMIC-IDx Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 Regression rate (%) 16.92 9.86 12.73 18.97 b AgentClinic (OOD) Self-ReflectionTrajectory RAGFlat ExperienceCDEG Figure 4: Regression rates across doctor-agent backbones on MIMIC-IDx and AgentClinic. Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 0 2 4 6 8 10 12 14 Average turns 14.02 10.79 9.24 13.70 a MIMIC-IDx Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 Average turns 20.55 13.51 10.50 14.24 b AgentClinic (OOD) Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 0 2 4 6 8 10 12 14 Average tool calls 13.05 9.79 8.24 12.70 c Gemini 3 Flash Preview Gemini 3.1 Pro GPT-5.6 Luna Claude Sonnet 5 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 Average tool calls 19.57 12.51 9.50 13.24 d Vanilla AgentSelf-ReflectionTrajectory RAGFlat ExperienceCDEG Figure 5: Interaction turns and tool-call counts across doctor- agent backbones on MIMIC-IDx and AgentClinic. budget. Rather than indiscriminately extending the diagnos- tic process, CDEG uses the evidence state to selectively guide acquisition of missing evidence or reappraisal of observed evidence. This effect is particularly apparent for the GPT- 5.6-Luna backbone on AgentClinic, where CDEG improves Accuracy by 11.22% while reducing the average interaction length by 0.71 turns, suggesting that it reduces ineffective exploration and reaches an evidence state sufficient for diag- nosis more quickly. Tool-call counts exhibit the same pattern, further demonstrating that CDEG improves interaction effi- ciency through targeted evidence acquisition or reappraisal. Conclusion We presented CDEG, a graph-based framework that learns reusable decision-critical evidence by contrasting successful and failed diagnostic trajectories, validating candidate evi- dence through controlled counterfactual interventions, and organizing validated evidence, diagnoses, and actions into a structured graph that guides acquisition or reappraisal ac- cording to the evolving patient evidence state. Across ID and OOD benchmarks with four doctor-agent backbones, CDEG consistently improves vanilla agent diagnostic performance, and achieves the best performance. In the future, we will explore online graph evolution for continual experience val- idation and extend CDEG to broader clinical scenarios. References Almansoori, M.; Kumar, K.; and Cholakkal, H. 2025. MedA- gentSim: Self-Evolving Multi-Agent Simulations for Real- istic Clinical Interactions. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, vol- ume 15968 of Lecture Notes in Computer Science, 362–372. Springer Nature Switzerland. Anthropic. 2026. Claude Sonnet 5 System Card. https: //w.anthropic.com/claude-sonnet-5-system-card. Ac- cessed: 2026-07-28. Balachandran, A. 2024. MedEmbed: Medical-Focused Em- bedding Models. https://github.com/abhinand5/MedEmbed. MedEmbed-large-v0.1. Accessed: 2026-07-28. Bhasuran, B.; Prosperi, M.; Hanna, K.; Petrilli, J.; Wash- ington, C. J.; and He, Z. 2026. Evaluation of Causal Rea- soning for Large Language Models in Contextualized Clini- cal Scenarios of Laboratory Test Interpretation. npj Digital Medicine, 9(1): 487. Bodenreider, O. 2004. The Unified Medical Language Sys- tem (UMLS): Integrating Biomedical Terminology. Nucleic Acids Research, 32: D267–D270. Chen, W.; Huang, G.; Wang, W.; and Zhu, Z. 2026. MedE- inst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis. In Proceed- ings of the 64th Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), 39778– 39798. San Diego, California, United States: Association for Computational Linguistics. Fang, J.; Chen, R.; Yang, X.; Yu, J.; Xu, J.; Vinod, A.; Shi, W.; Chen, T.; Ji, H.; Zhai, C.; Ding, Y.; and Zhang, Y. 2026. Benchmarking Multi-turn Medical Diagnosis: Hold, Lure, and Self-Correction. arXiv:2604.04325. Gao, Y.; Zhou, X.; Li, Y.; Zhao, Y.; and Liu, R. 2026. MedEx- Agent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments. arXiv:2605.07058. Google. 2025. Gemini 3 Flash Preview. https://ai.google.dev/ gemini-api/docs/models/gemini-3-flash-preview. Accessed: 2026-07-28. Google. 2026. Gemini 3.1 Pro Preview. https://ai.google. dev/gemini-api/docs/models/gemini-3.1-pro-preview. Ac- cessed: 2026-07-28. Han, X.; Fan, Y.; Zhao, S.; Wang, H.; and Qin, B. 2026. GSEM: Graph-based Self-Evolving Memory for Experience Augmented Clinical Reasoning. arXiv:2603.22096. Hassoon, A.; Peng, X.; Irimia, R.; Lianjie, A.; Leo, H.; Ban- deira, A.; Woo, H. Y. J.; Dredze, M.; Abdulnour, R.-E.; Mc- Donald, K. M.; Peterson, S.; and Newman-Toker, D. 2026. Evaluating the AI Potential as a Safety Net for Diagnosis: A Novel Benchmark of Large Language Models in Correcting Diagnostic Errors. medRxiv preprint. Version 1. Hsu, H.-L.; Wang, Z.; Zhang, D.; Chen, N.-C.; Wang, J.; Ding, J.-E.; Hsu, C.-H.; Wang, G.; Liu, F.; Hung, F.-M.; Wu, C.; and Shen, L. 2026. MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs. arXiv:2605.07305. Jiang, Y.; Black, K. C.; Geng, G.; Park, D.; Zou, J.; Ng, A. Y.; and Chen, J. H. 2025. MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents. NEJM AI, 2(9): AIdbp2500144. Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2021. What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams. Applied Sciences, 11(14): 6421. Johnson, A.; Pollard, T.; Horng, S.; Celi, L. A.; and Mark, R. 2023a. MIMIC-IV-Note: Deidentified Free-Text Clinical Notes. PhysioNet. Version 2.2. Johnson, A. E. W.; Bulgarelli, L.; Shen, L.; Gayles, A.; Sham- mout, A.; Horng, S.; Pollard, T. J.; Hao, S.; Moody, B.; Gow, B.; Lehman, L.-w. H.; Celi, L. A.; and Mark, R. G. 2023b. MIMIC-IV, a Freely Accessible Electronic Health Record Dataset. Scientific Data, 10: 1. Johnson, A. E. W.; Pollard, T. J.; Berkowitz, S. J.; Green- baum, N. R.; Lungren, M. P.; Deng, C.-y.; Mark, R. G.; and Horng, S. 2019. MIMIC-CXR, a De-identified Publicly Available Database of Chest Radiographs with Free-Text Re- ports. Scientific Data, 6: 317. Li, B.; Du, S.; and Guo, Y. 2026. Joint Optimization of Reasoning and Dual-Memory for Self-Learning Diagnostic Agent. arXiv:2604.07269. Li, S. S.; Balachandran, V.; Feng, S.; Ilgen, J. S.; Pierson, E.; Koh, P. W.; and Tsvetkov, Y. 2024. MediQ: Question-Asking LLMs and a Benchmark for Reliable Interactive Clinical Reasoning. In Advances in Neural Information Processing Systems, volume 37, 28858–28888. Curran Associates, Inc. Liu, F.; Shareghi, E.; Meng, Z.; Basaldella, M.; and Col- lier, N. 2021. Self-Alignment Pretraining for Biomedical Entity Representations. In Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4228–4238. Online: Association for Computational Linguis- tics. Liu, X.; Sun, D.; Fung, Y.; Hakkani-Tur, D.; and Abdelzaher, T. F. 2025. DocCHA: Towards LLM-Augmented Interac- tive Online diagnosis System. In Proceedings of the 26th Annual Meeting of the Special Interest Group on Discourse and Dialogue, 609–619. Avignon, France: Association for Computational Linguistics. Long, Z.; Bao, Z.; and Wei, Z. 2026. Strong Reasoning Isn’t Enough: Evaluating Evidence Elicitation in Interactive Diagnosis. arXiv:2601.19773. Mo, K.; Venkatayogi, S.; Shaib, C.; Kouzy, R.; Xu, W.; Wal- lace, B. C.; and Li, J. J. 2026. Faithfulness vs. Safety: Evaluat- ing LLM Behavior Under Counterfactual Medical Evidence. In Findings of the Association for Computational Linguis- tics: ACL 2026, 37053–37081. San Diego, California, United States: Association for Computational Linguistics. Nagesh, S.; Mishra, N.; Naamad, Y.; Rehg, J. M.; Shah, M. A.; and Wagner, A. 2023. Explaining a Machine Learn- ing Decision to Physicians via Counterfactuals. In Proceed- ings of the Conference on Health, Inference, and Learning, volume 209 of Proceedings of Machine Learning Research, 556–577. PMLR. National Academies of Sciences, Engineering, and Medicine. 2015. Improving Diagnosis in Health Care. Wash- ington, DC: The National Academies Press. OpenAI. 2026a. GPT-5.6: Frontier Intelligence That Scales with Your Ambition. https://openai.com/index/gpt-5-6/. Ac- cessed: 2026-07-28. OpenAI. 2026b. Introducing GPT-5.5. https://openai.com/ index/introducing-gpt-5-5/. Accessed: 2026-07-28. Ouyang, S.; Yan, J.; Hsu, I.-H.; Chen, Y.; Jiang, K.; Wang, Z.; Han, R.; Le, L. T.; Daruki, S.; Tang, X.; Tirumalashetty, V.; Lee, G.; Rofouei, M.; Lin, H.; Han, J.; Lee, C.-Y.; and Pfis- ter, T. 2026. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. In The Fourteenth International Conference on Learning Representations. Qiu, P.; Wu, C.; Liu, J.; Zheng, Q.; Liao, Y.; Wang, H.; Yue, Y.; Fan, Q.; Zhen, S.; Wang, J.; Gu, J.; Wang, Y.; Zhang, Y.; and Xie, W. 2025. Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment. arXiv:2510.24654. Qu, Z.; and Färber, M. 2026. MediEval: A Unified Med- ical Benchmark for Patient-Contextual and Knowledge- Grounded Reasoning in LLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), 16150–16164. San Diego, California, United States: Association for Computational Linguistics. Ren, R.; Wang, Y.; Liang, Y.; Luo, L.; Liu, J.; Wang, H.; Feng, C.; Zhang, Y.; Miao, C.; Wen, J.-R.; and Zhao, W. X. 2026. Emulating Clinician Cognition via Self-Evolving Deep Clinical Research. arXiv:2603.10677. Ren, W.; Zhao, T.; Wang, L.; Wang, T.; and Honavar, V. G. 2025. DiaLLMs: EHR-Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction. In Findings of the Association for Computational Linguistics: ACL 2025, 25622–25635. Vienna, Austria: As- sociation for Computational Linguistics. Richens, J. G.; Lee, C. M.; and Johri, S. 2020. Improving the Accuracy of Medical Diagnosis with Causal Machine Learning. Nature Communications, 11(1): 3923. Rose, D. P.; Hung, C.-C.; Lepri, M.; Alqassem, I.; Gash- teovski, K.; and Lawrence, C. 2025. MEDDxAgent: A Uni- fied Modular Agent Framework for Explainable Automatic Differential Diagnosis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13803–13826. Vienna, Austria: Association for Computational Linguistics. Sanghvi, A.; Akash, N.; Imam, R.; Sharma, A.; and Jain, M. 2026. MeDxAgent: Multi-Agent Consultation for Interactive Medical Diagnosis. arXiv:2606.03416. Schmidgall, S.; Ziaei, R.; Harris, C.; Kim, J. W.; Reis, E. P.; Jopling, J.; and Moor, M. 2026. AgentClinic: A Multimodal Benchmark for Tool-Using Clinical AI Agents. npj Digital Medicine, 9(1): 499. Shen, W.; Jian, B.; Li, J.; Liu, C.; Moll, J.; Hu, X.; Rueckert, D.; Li, H. B.; and Pan, J. 2026. Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve. arXiv:2604.14475. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Re- inforcement Learning. In Advances in Neural Information Processing Systems, volume 36, 8634–8652. Curran Asso- ciates, Inc. Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2024. Voyager: An Open-Ended Embodied Agent with Large Language Models. Transactions on Machine Learning Research. Wang, Z. Z.; Mao, J.; Fried, D.; and Neubig, G. 2025. Agent Workflow Memory. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceed- ings of Machine Learning Research, 63897–63911. PMLR. Xiong, Z.; Lin, Y.; Xie, W.; He, P.; Liu, Z.; Tang, J.; Lakkaraju, H.; and Xiang, Z. 2026. How Memory Man- agement Impacts LLM Agents: An Empirical Study of Experience-Following Behavior. In Proceedings of the 64th Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), 623–645. San Diego, Cal- ifornia, United States: Association for Computational Lin- guistics. Xu, R.; Yu, Y.; Zhang, C.; Ali, M. K.; Ho, J. C.; and Yang, C. 2022. Counterfactual and Factual Reasoning over Hy- pergraphs for Interpretable Clinical Predictions on EHR. In Proceedings of the 2nd Machine Learning for Health Sym- posium, volume 193 of Proceedings of Machine Learning Research, 259–278. PMLR. You, Z.; Chen, X.; Vashishtha, A.; Du, S.; Erion-Barner, G.; Mei, H.; Peng, H.; and Guo, Y. 2026. Improving Clin- ical Diagnosis with Counterfactual Multi-Agent Reasoning. arXiv:2603.27820. Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; Huang, F.; and Zhou, J. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv:2506.05176. Zhao, A.; Huang, D.; Xu, Q.; Lin, M.; Liu, Y.-J.; and Huang, G. 2024. ExpeL: LLM Agents Are Experiential Learners. In Proceedings of the Thirty-Eighth AAAI Conference on Arti- ficial Intelligence, 19632–19642. Vancouver, Canada: AAAI Press. Zheng, L.; Wang, R.; Wang, X.; and An, B. 2024. Synapse: Trajectory-as-Exemplar Prompting with Memory for Com- puter Control. In The Twelfth International Conference on Learning Representations.