Paper deep dive
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries
Jiyoun Kim, Muhan Yeo, Eunhye Jang, Jeewon Yang, Hangyul Yoon, Su Ji Lee, Hee Jo Han, Hee-Jae Jung, Doyun Kwon, Jun young Lee, Jaehun Lee, Jung-Oh Lee, Sunjun Kweon, Jong Hak Moon, Daseul Kim, Minjae Cho, Edward Choi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 6/20/2026, 6:51:52 AM
Summary
EHRNote-ChatQA is a new benchmark designed for evidence-grounded, multi-turn clinical question answering (QA) over longitudinal patient discharge summaries. Built from de-identified MIMIC-IV data, it contains 967 patient-level multi-turn samples and 16,072 medical-expert-verified QA pairs. The benchmark addresses the limitations of existing clinical QA datasets by requiring models to not only answer clinical content questions but also identify the specific source (evidence-grounding) within multiple discharge summaries. The construction uses an expert-informed pipeline involving a granular discharge-summary structuring schema, multi-turn QA templates, and LLM-based generation (Gemini 1.5 Pro), followed by rigorous review by 11 medical experts. Evaluation of 22 LLMs shows that models struggle with evidence grounding and that errors tend to compound across multi-turn dialogues.
Entities (5)
Relation Signals (3)
EHRNote-ChatQA â availableon â PhysioNet
confidence 100% ¡ The dataset will be made publicly available through PhysioNet credentialed access.
EHRNote-ChatQA â builtfrom â MIMIC-IV
confidence 100% ¡ Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples
EHRNote-ChatQA â uses â Gemini 1.5 Pro
confidence 100% ¡ The pipeline uses... LLM-based generation with Gemini 1.5 Pro.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When reviewing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical question answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Benchmarking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi-turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet credentialed access.
Tags
Links
- Source: https://arxiv.org/abs/2606.15735v2
- Canonical: https://arxiv.org/abs/2606.15735v2
Trouble viewing inline? Open PDF directly â
Full Text
370,045 characters extracted from source content.
Expand or collapse full text
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries Jiyoun Kim 1 , Muhan Yeo 2,3 , Eunhye Jang 4 , Jeewon Yang 1 , Hangyul Yoon 1 , Su Ji Lee 5 , Hee Jo Han 5,7 , Hee-Jae Jung 8 , Doyun Kwon 6 , Jun young Lee 3,9 , Jaehun Lee 10 , Jung-Oh Lee 11 , Sunjun Kweon 1 , Jong Hak Moon 1 , Daseul Kim 12 , Minjae Cho 3 , Edward Choi 1â 1 KAIST 2 Seoul National University 3 Seoul National University Bundang Hospital 4 SAIHST, Sungkyunkwan University 5 Yonsei University College of Medicine 6 Gangnam Severance Hospital 7 Severance Hospital 8 Seoul Medical Center 9 Seoul National University Hospital 10 National Cancer Center 11 Icahn School of Medicine at Mount Sinai 12 Samsung Medical Center jiyoun.kim, edwardchoi@kaist.ac.kr Abstract Discharge summaries are crucial clinical documents containing the context of a patientâs overall hospital stay, and are routinely reviewed by medical experts for patient readmission, ongoing care, and diagnostic decision-making. When review- ing them, medical experts often must iteratively synthesize information across multiple summaries while verifying the evidence supporting each answer. Al- though large language models (LLMs) are increasingly explored for clinical ques- tion answering, existing benchmarks do not sufficiently reflect this setting: they often evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote- ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patientsâ multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding ques- tion) across eight clinical categories. The benchmark is constructed through an expert-informed pipeline combining discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by review and revision of every single QA sample by 11 medical experts. Bench- marking 22 open- and closed-source LLMs reveals several challenges, including that LLMs struggle more with evidence grounding than content answering, multi- turn errors compound across turns, and single-turn clinical QA performance does not reliably transfer to this setting. These findings establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems. The dataset will be made publicly available through PhysioNet 2 credentialed access. â Corresponding author 2 https://physionet.org Preprint. arXiv:2606.15735v2 [cs.CL] 16 Jun 2026 1 st Discharge Summary: [Patient X] [chartdate1] Chief Complaint: Abdominal diste ntion, hyponatremia ... HPI: ___ M w/ EtOH cirrhosis, r ef'd for worsening ascites + hyponatremia; prior tap at OSH ED, 6L removed; off diuretics 2/2 hyponatremia ... Ma jor Surgical or Invasive Procedure: large vol paracentesis - ___; endoscopy w/ Dobhoff (NJ) tube placement - ___; right heart cathe te rization - ___ ... Physical Exam: ... Ne uro: somnolent, arousable to voice. No asterixis ... Pertinent Results: ... SODIUM-121* K-5.6* ... US-guided therapeutic para: 9.75 L re moved ... Brief Hospital Course: #ASCITES / #ETOH CIRRHOSIS (Child C): hx of HE; trialed diuretics but diuretic-refr actory 2/2 hyponatremia --> mgmtw/ large-vol para (diure tics could not be used); Cards consult; RHC - normal gradient across tricuspid valve, ruling out TR; would not limit from TIPS. #SEVERE MALNU TRITION: s/p NJ place ment under endoscopy, d/c'd on tube feeds ... 4 th Discharge Summary: [Patient X] [chartdate4] Chief Complaint: Planne d TIPS ... Ma jor Surgical or Invasive Procedure: TIPS ___ - successful R IJ access w/ transjugularintrahepatic portosystemic shunt placement (L portal v --> L hepatic v); amber asc ites drained ... HPI: ___ M w/ EtOH cirrhosis c/b refractory ascites + well- controlled hepatic encephalopathy, s/p TIPS; had been ne eding weekly 8L para ... Physical Exam: ... Ne uro: +asterixis vs tremor ... Pertinent Results / STU DIES - TIPS ___: ... pre-/post-TIPS RA pre ssure 28 --> 23 mmHg; portosystemic gradient 10 --> 5 mmHg; 8L amber ascites removed ... Brief Hospital Course: # Refrac tory ascites: s/p TIPS, monitore d - no worsening; cont'd lactulose; if ascites recurs --> candidate for dilation of current TIPS or 2nd TIPS ... Q1. What were the main presenting symptoms for the admission on [chartdate1]? A. Abdominal distention and hyponatremia. B. A clogged nasojejunal feeding tube that was placed during a prior admission. C. Worsening abdominal distention and severe malnutrition. D. Confusion and somnolence consistent with an episode of hepatic encephalopathy. E. Abdominal distention and hyperkalemia. Q2. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 [chartdate1] Headers: Physical Exam B. Note #1 [chartdate1] Headers: Pertinent Results C. Note #1 [chartdate1] Headers: Chief Complaint D. Note #1 [chartdate1] Headers: History of Present Illness, Discharge Diagnosis E. Note #1 [chartdate1] Headers: Major Surgical or Invasive Procedure, Discharge Condition Q3. How was the symptom managed during that admission? A. As the ascites was diuretic-refractory, the patient was evaluated for and scheduled for a transjugular intrahepatic portosystemic shunt (TIPS) procedure. B. The severe malnutrition was managed via the endoscopic placement of a Dobhoff feeding tube to provide enteral nutrition. C. The patient was managed with aggressive intravenous diuretics and a cardiology consultation to evaluate for tricuspid valve surgery. D. With a large-volume paracentesis that removed 9.75 L of ascitic fluid, as diuretics could not be used. E. With a large-volume paracentesis that removed 6.0 L of ascitic fluid, based on the volume removed in a recent ED visit. Q5. Was a different intervention performed in later admissions to manage this persistent symptom? A. Yes, on admission [chartdate4], the patient underwent a planned transjugularintrahepatic portosystemic shunt (TIPS) procedure. B. Yes, the Dobhoff tube was replaced and advanced post-pyloric by interventional radiology in an attempt to prevent future clogging. C. Yes, a permanent indwelling peritoneal catheter was placed to allow for intermittent drainage of ascitic fluid at home by the patient's caregiver. D. Yes, in a later admission, he underwent a right heart catheterization to obtain definitive pressures and was referred for consideration of tricuspid valve surgery. E. Yes, he underwent a TIPS revision with balloon dilatation of a previously placed stent during the admission on [chartdate4]. Q6. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 [chartdate4] Headers: Discharge Diagnosis, Brief Hospital Course B. Note #3 [chartdate3] Headers: Brief Hospital Course C. Note #4 [chartdate4] Headers: Major Surgical or Invasive Procedure D. Note #4 [chartdate4] Headers: Chief Complaint, Major Surgical or Invasive Procedure E. Note #4 [chartdate4] Headers: History of Present Illness Q7. What was the initial outcome of that intervention? A. The right heart catheterization revealed a normal gradient across the tricuspid valve, definitively ruling out significant tricuspid regurgitation. B. The immediate outcome was the development of worsening hepatic encephalopathy, with confusion and asterixis, requiring an increase in his lactulose dose. C. The initial outcome was a reduction in paracentesis volume from 8 liters to 4 liters, but with an ongoing need for weekly procedures. D. The procedure was successful, with a decrease in the right atrial pressure from 28 mmHg to 23 mmHg, although 8 L of ascites were still drained. E. The procedure was successful, with the portosystemic gradient decreasing from 10 to 5 mmHg and 8 L of ascites drained. Q8. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 [chartdate4] Headers: Discharge Condition, Discharge Diagnosis B. Note #4 [chartdate4] Headers: Pertinent Results C. Note #4 [chartdate4] Headers: Discharge Diagnosis D. Note #4 [chartdate4] Headers: Major Surgical or Invasive Procedure E. Note #4 [chartdate4] Headers: History of Present Illness, Discharge Instructions Q4. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 [chartdate1] Headers: Discharge Diagnosis, Discharge Instructions B. Note #1 [chartdate1] Headers: History of Present Illness C. Note #1 [chartdate1] Headers: Past Medical History, Major Surgical or Invasive Procedure D. Note #1 [chartdate1] Headers: Discharge Diagnosis E. Note #1 [chartdate1] Headers: Pertinent Results, Brief Hospital Course Figure 1: An illustrative multi-turn QA sample of EHRNote-ChatQA (only parts of Notes 1 and 4 are shown for readability; [chartdateN] for Note #N is replaced with the actual chart date). Given a single patientâs entire set of discharge summaries, EHRNote-ChatQA consists of a sequence of con- tent questions about the notes (Q1, Q3, Q5, Q7), each paired with an evidence-grounding question that requires identifying the exact source supporting the preceding answer (Q2, Q4, Q6, Q8). 1 Introduction Patient discharge summaries are comprehensive clinical documents that describe a patientâs entire clinical course from admission to discharge. They record clinical events such as diagnoses, medi- cations, procedures, and laboratory tests, along with contextual information including causal rela- tionships, diagnostic reasoning, and patient outcomes. These documents are routinely reviewed by medical experts (e.g. clinicians, nurses) to understand a patientâs prior clinical status in settings such as readmission, longitudinal care, and diagnostic decision-making. However, discharge summaries are often written by different clinicians, follow inconsistent formats, and contain extensive medical abbreviations. Moreover, a new discharge summary is generated for each admission, accumulating multiple notes for a single patient over time. As a result, reviewing a patientâs discharge summaries requires navigating long, fragmented documents, making it difficult to extract relevant information. Recent advances in large language models (LLMs) have enabled question answering (QA) over clin- ical documents [31, 24], allowing medical experts to retrieve targeted information without manually reviewing each note. Yet, for LLMs to be safely integrated into clinical workflows, two fundamental requirements must be met: (1) accurately answering medical expertsâ iterative, multi-turn questions, and (2) grounding each answer in traceable evidence from the source documents. In practice, clinical inquiries are rarely isolated; they are inherently multi-turn [9], typically starting with an initial query and proceeding through related follow-ups that build on prior context. For ex- ample, a medical expert may first ask, âWhat symptoms led to the patientâs admission?â, followed by âWhat did the evaluation of those symptoms reveal?â, and then âDid any of those symptoms pre- cipitate the patientâs next hospitalization?â. Moreover, beyond answer correctness, traceability back to source is equally critical in clinical settings, since LLMs may generate hallucinated or unverifi- able responses [15, 26]. It is therefore essential to evaluate whether model output can be grounded in verifiable evidence [5], letting users trace each answer to specific segments of the source documents. Despite these requirements, existing clinical QA benchmarks fail to capture key aspects of real- world clinical usage. Benchmarks such as MedQA [16], MedMCQA [25], and PubMedQA [17] evaluate medical knowledge through exam-style questions and do not involve patient records. Electronic health record (EHR)-based benchmarks such as EHRNoteQA [20], emrQA [27], and 2 DiSCQ [22] include expert-curated questions over discharge summaries, but are limited to single- turn QA and lack evidence-grounding questions that ask for the source supporting each answer. To address this gap, we introduce EHRNote-ChatQA, the first evidence-grounded, multi-turn, multi-note clinical QA benchmark over patient discharge summaries. Built from de-identified MIMIC-IV [18] discharge summaries, our benchmark covers 967 patients (one to five notes each), yielding 967 multi-turn samples with 16,072 QA pairs: 8,036 content questions, each paired with an evidence-grounding question that identifies the supporting source location. Each sample provides the patientâs full sequence of discharge summaries as context. Examples are shown in Appendix A. To generate multi-turn questions that reflect medical expert inquiries in clinical practice, with clin- ically accurate answers grounded in the patientâs notes, we develop an expert-informed, structured QA generation pipeline in collaboration with 4 medical experts (3 clinicians and 1 nurse). The pipeline uses a comprehensive discharge-summary structuring schema and multi-turn QA templates (both defined with the experts) as seeds for LLM-based generation with Gemini 2.5 Pro. For each content question, we also generate a paired evidence-grounding question asking for the source loca- tion supporting the answer. Every QA sample is then reviewed and revised by 11 medical experts over three weeks to ensure clinical naturalness, accuracy, and evidence-grounding reliability. Using this benchmark, we comprehensively evaluate 22 closed- and open-source LLMs. Our evalu- ations reveal several findings, including the following: (1) models exhibit a substantial gap between content and evidence-grounding accuracy; (2) errors propagate across turns; and (3) strong perfor- mance on existing single-turn benchmarks does not reliably transfer to multi-turn evaluation. Our main contributions are as follows: ⢠Benchmark. We introduce EHRNote-ChatQA, the first evidence-grounded, multi-turn, multi-note clinical QA benchmark over patient discharge summaries, comprising 967 multi-turn samples and 16,072 medical-expert-verified QA pairs across eight categories. ⢠Dataset Construction Pipeline.We develop a structured, medical-expert-informed pipeline for generating clinically grounded multi-turn QA data, combining a discharge- summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation, followed by medical expert validation of every single QA sample. ⢠Evaluation. We benchmark 22 LLMs and conduct a comprehensive analysis, including content and evidence-grounding accuracy, multi-turn dependency and error propagation, comparison with existing single-turn clinical QA benchmark. 2 Related Work Multi-turn QA Benchmarks. Several benchmarks evaluate LLMs on multi-turn dialogue and se- quential QA. MT-Bench [39] and MT-Eval [19] cover general-domain conversational topics (e.g., writing, reasoning, STEM) and use GPT-4 as a judge to score open-ended responses. CoQA [28] and QuAC [6] are reading-comprehension dialogue benchmarks in which a model answers a se- quence of questions grounded in a single text passage. While these benchmarks have been central to evaluating multi-turn question answering, none target patient electronic health records (EHRs), and none include evidence-grounding questions asking for the source document of each answer. Clinical QA Benchmarks. Existing clinical QA benchmarks cover diverse settings, including medical knowledge evaluation and question answering over patient records. Benchmarks such as MedQA [16] (USMLE), MedMCQA [25] (Indian medical entrance exams), and PubMedQA [17] (PubMed-abstract comprehension) evaluate the parametric medical knowledge of LLMs through exam-style or literature-comprehension questions, without grounding in a specific patientâs record. More closely related to our setting, benchmarks such as emrQA [27], EHRNoteQA [20], and DiSCQ [22] are built on patient EHR notes (e.g., discharge summaries), but are limited to single-turn questions and do not pair each content question with an evidence-grounding question asking for the source evidence supporting the answer. Evidence-Grounding Benchmarks. Evidence grounding evaluation has also been studied in general-domain QA and long-form generation. Attributed QA [5] evaluates whether a model can provide both an answer and a supporting attribution to a source corpus, while ALCE [12] evaluates 3 LLM generations with citations over open-domain QA tasks such as ASQA [34], QAMPARI [4], and ELI5 [10]. These benchmarks emphasize the importance of verifiable question answering, but they are not specific to clinical records and do not evaluate multi-turn QAs. Among patient-record benchmarks, emrQA [27] includes answer evidence pairs derived from clinical annotations, and ArchEHR-QA [33] provides sentence-level evidence annotations for patient-style questions over discharge-summary excerpts. However, neither spans the patientâs multiple discharge summaries (they are only single note) or evaluates evidence grounding under multi-turn dialogue. To our knowledge, no prior benchmark combines the three properties EHRNote-ChatQA targets: multi-turn questions reflecting realistic clinical record review of medical experts, multi-note ground- ing across a patientâs longitudinal discharge summaries, and per-turn evidence-grounding questions that identify the source supporting each answer. Table 3 (in Appendix B) compares EHRNote- ChatQA with prior patient-record QA benchmarks along these dimensions. 3 Methods 3.1 Data Source and Cohort Construction We use de-identified discharge summaries from MIMIC-IV [18], a database of patient electronic health records (EHRs) from Beth Israel Deaconess Medical Center in Boston, USA. As approxi- mately 90% of patients in MIMIC-IV have 1â5 discharge summaries, we sampled 1,000 patients across five note-count groups, with 200 patients per group. We further assigned patients uniformly across eight expert-defined question categories (see Section 3.2)âDiagnosis, Symptom, Procedure, Medication, Microbiology, Clinical Assessment, Clinical Outcome, and Discharge Planâyielding 25 patients for each categoryânote-count combination. 3.2 Discharge Summary Structuring Schema and Multi-turn QA Template Design Our benchmark is built around a central principle: the multi-turn questions should closely reflect how medical experts actually inquire and follow up when reviewing discharge summaries in clinical practice, and the answers must be clinically accurate and grounded in the patientâs notes. To this end, we develop a data generation pipeline in collaboration with 4 medical experts (3 clini- cians and 1 nurse). We first define a schema for structuring discharge summary content, consisting of 19 clinical entity types, 3â10 attributes per entity type, and 23 relationship types (full schema in Appendix C). Although prior work has proposed partial schemas for structuring discharge sum- maries [36, 37, 35, 14], these omit several important enity types (e.g., microbiology, allergy, activ- ity) and merge clinically distinct concepts into broad categoriesâfor example, diagnosis, symptom, and finding under a âProblemâ entity type, and procedure and medication under a âTreatmentâ en- tity type. Such coarse-grained representations limit the detailed clinical attributes and relationships needed for realistic clinical questioning. We therefore build upon prior discharge-summary schemas and refine them with medical experts to obtain more comprehensive and granular representations. Together with medical experts, we also define eight question categories reflecting common inquiries during discharge-summary review: Diagnosis, Symptom, Procedure, Medication, Microbiology, Clinical Assessment, Clinical Outcome, and Discharge Plan. In collaboration with the same ex- perts, we design multi-turn QA templates for each category that capture clinically realistic inquiry patterns, such as querying entity-specific information (e.g., medication dosage or route) or asking about clinically meaningful relationships across turns (e.g., symptoms or patient outcomes after a medication). The QA templates are built on the discharge summary structuring schema, leverag- ing entity attributes for entity-specific questions and relationships for cross-entity chaining, so that each answer is grounded in the patientâs discharge summaries. The complete set of templates is provided in Appendix D. For each category, we design single- and multi-note templates, where multi-note templates include cross-admission inquiry patterns such as disease course, recurrence, and discharge-plan evolution. Each category contains 5 templates for each of the single-note and multi-note settings, yielding 80 templates total (5Ă 8 categoriesĂ 2 settings). Note that these QA templates are not used to generate samples by rule-based slot filling. Instead, they serve as examples of clinically realistic inquiry patterns, which the LLM (Gemini-2.5-Pro [7]) uses as seeds to generate patient-specific, contextually appropriate multi-turn QA samples from each patientâs discharge summaries. 4 Figure 2: Overview of the EHRNote-ChatQA construction pipeline. 3.3 Initial Multi-turn QA Generation Using the expert-defined schema and multi-turn QA templates, we generate initial QA samples with Gemini-2.5-Pro. The QA samples contain two question types: (1) content questions that ask about the clinical content of the discharge summaries, and (2) evidence-grounding questions, each paired with a content question, that ask which part of the source notes supports the answer to the content question. All questions are generated in a 5-way multiple-choice format (one correct choice, four incorrect answer choices). We choose the multiple-choice format over the open-ended format because (1) it enables objective, consistent evaluation via rule-based grading rather than subjective scoring of free-form responses; and (2) it is scalable: rule-based automated grading avoids the cost of medical expert manual review and the hallucination risk of LLM-as-a-judge evaluation. Each sample is generated through a four-step pipeline: (1) generate multi-turn content questions that reflect medical expert inquiries; (2) brainstorm plausible directions for incorrect answers to each content question; (3) generate one correct and four incorrect content-answer choices; and (4) generate the evidence-grounding answer choices. We detail the design principles below (see Appendix EâF, I for additional details on the generation process and prompts): (1) Multi-turn content questions (Step 1). We prompt Gemini-2.5-Pro to generate the multi-turn content questions given a single patientâs full sequence of discharge summaries, under the following constraints in priority order. (a) Clinical naturalness and realism (highest priority): questions must reflect inquiries that medical experts realistically pose when reviewing discharge summaries. The model is seeded with the expert-curated multi-turn QA templates as exemplars and is encouraged to include why/how reasoning questions when the reasoning can be inferred from the patientâs notes. Every question must require both clinical knowledge and the patientâs discharge-summary context to answer: we forbid questions answerable from external medical knowledge alone, and encourage questions that require clinical reasoning to interpret the notes rather than purely surface-level text retrieval. (b) Multi-turn dependency: answering a later question should require referencing prior questions and/or their answers, so that the multi-turn questions capture the cumulative context of chart review rather than collapsing into independent single-turn questions. (c) Multi-note coverage: for patients with multiple discharge summaries, questions should span across multiple notes when clinically meaningful questions can be asked (e.g., medication changes or diagnosis evolution across admissions). However, clinical naturalness takes priority; if a note cannot be covered without forcing a contrived question, it should not be generated. 5 (2) Content-question answer choices (Steps 2â3). We prompt Gemini-2.5-Pro to first brainstorm plausible distractor directions per question (Step 2), then generate one correct and four incorrect content-answer choices (Step 3). (a) Correct answer choice: should provide a complete answer to the question, grounded in the patientâs notes (Generated in Step 3 with the four incorrect answer choices per question). (b) Incorrect answer choices: For brainstorming in Step 2, six types of plausible distractor directions are defined with medical experts (see Appendix E). These directions are derived from one or both of two sources: (1) note-grounded distractors, which are incorrect interpretations of the patientâs notes and (2) clinical-knowledge-grounded distractors, which are content not present in the patientâs notes but clinically plausible given the patientâs clinical context. Together, these sources ensure correct answering requires both interpreting discharge-summary con- tent and applying relevant clinical knowledge. We incorporate two additional design principles in Steps 2-3. First, we encourage cross-turn error chaining where clinically appropriate distractors should form coherent chains across turns, such that if a model selects an incorrect answer in an earlier question, subsequent incorrect choices include plausible continuations of that earlier error. This allows mistakes to compound across turns, which is a key property of realistic multi-turn QA. Second, we encourage shared-entity choices: some distractors should reference the same key enti- ties as the correct answer but differ in relationships, attributes, or clinical interpretations applied to them. This prevents the trivial strategy of selecting whichever answer choice mentions entities from the notes, and instead requires evaluating the clinical reasoning tied to those entities. (3) Evidence-grounding-question answer choices (Step 4). Each content question is paired with an evidence-grounding question: âWhat are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer?â. Each answer choice for the evidence-grounding question consists ofâ¨note number (& chart date), section header(s)⊠tuples drawn from the patientâs discharge summaries (example section headers include Chief Complaint, Major Surgical or Invasive Procedure, Brief Hospital Course, Pertinent Results, and Discharge Med- ications). We prompt Gemini-2.5-Pro to generate the evidence-grounding answer choices as follows. (a) Correct answer choice (Minimal-yet-sufficient constraint): the smallest set ofâ¨note number (& chart date), section header(s)⊠tuples that fully supports the answer to the corresponding content question. Sufficiency requires that every clinical detail of the answer be grounded in the selected sections, while minimality requires removing any section whose supporting information is already contained in another selected section. (b) Incorrect answer choices: distractors that fail to fully support the answer via one of four error patterns: (i) all headers irrelevant, (i) headers covering only part of the answer, (i) a mix of partially relevant and irrelevant headers, or (iv) the full correct set plus irrelevant additions. Irrelevant headers are encouraged to be semantic traps: names that plausibly align with the answerâs wording but whose contents do not support it, requiring models to read the section content rather than rely on surface header-name matching. 3.4 Medical Expert Review and Dataset Finalization To ensure clinical naturalness, correctness, and grounding in the discharge summaries, each of the 1,000 multi-turn QA samples (one per patient) produced by the four-step pipeline was comprehen- sively reviewed by medical experts. Reviewers and protocol. Each sample was reviewed by one of 11 medical experts (9 specialist clinicians, 1 clinician resident, 1 nurse) over three weeks, covering 16,072 QA pairs. Reviewers were provided with a Streamlit interface that displayed each sampleâs multi-turn QA chain, with one correct and four incorrect choices per question, alongside the patientâs discharge summaries. They were provided a separate spreadsheet to record revisions. Full annotation instructions and interface screenshots are in Appendix G. Review criteria. Reviewers were instructed to apply the following decisions per sample: (1) Delete sample if the overall multi-turn question sequence was clinically unnatural and could not be revised into a realistic clinical multi-turn inquiry, or if the source headers supporting the correct answer had been incorrectly de-identified in MIMIC-IV, making the evidence-grounding question unanswerable; (2) Revise question if an individual question was not clinically natural but could be edited into a question that medical experts would ask in clinical practice; (3) Revise correct answer choice if its content was partially or fully incorrect against the discharge summaries; (4) Revise incorrect answer choice if it was actually correct and therefore invalid as an incorrect answer choice. 6 Table 1: Statistics of the final EHRNote-ChatQA dataset. Each patient-level sample consists of a multi-turn QA sequence over a single patientâs entire set of discharge summaries. Each content question is followed by an evidence-grounding question, so the two question counts are equal. OverallBy question categoryBy # discharge summaries MetricCountCategorySamples Total questions Total questions per sampleNotes Samples Total questions Total questions per sample Patient-level samples967Diagnosis1171,99417.012002,36011.8 Discharge summaries / patient (mean)2.95Symptom1222,17617.822003,23816.2 Content questions8,036Procedure1201,99416.631983,59418.2 Evidence-grounding (Source) questions8,036Medication1242,09416.941873,44418.4 Total questions (Content + Source)16,072Microbiology1211,71414.251823,43618.9 Min total questions / sample8Clinical Assessment1202,12017.7 Mean total questions / sample16.6Clinical Outcome1222,05016.8 Max total questions / sample26Discharge Plan1211,93016.0 Total96716,07216.6Total96716,07216.6 Outcome. After review, 33 of the 1,000 samples were removed, yielding 967 final samples. Each content question has one correct and four incorrect choices, yielding 8,036 correct and 32,144 in- correct content answer choices; the same counts apply to evidence-grounding questions. Among retained content questions, 125 questions (1.6% of 8,036), 101 correct choices (1.3%), and 156 incorrect choices (0.5% of 32,144) were revised. For evidence-grounding questions, 425 correct choices (5.3% of 8,036) and 1,676 incorrect choices (5.2% of 32,144) were revised. The correct- answer position (AâE) was then shuffled to be approximately uniform across the dataset. Table 1 reports the final dataset statistics. 4 Experiments 4.1 Experimental Settings Models. We benchmark 22 LLMs on EHRNote-ChatQA, spanning proprietary frontier models (Google Gemini [7], OpenAI GPT [30]), open-weight general-purpose models (Qwen3 [38], Llama- 4 [3], DeepSeek-R1-Distill [13], Ministral [23], Phi [2, 1]), and medical-domain specialized models (MedGemma [29], MediPhi [8]). Metrics. For each model, we report nine accuracy metrics combining three correctness criteria with three aggregation levels. The correctness criteria are: Content for content-question accuracy, Source for evidence-grounding accuracy, and Paired for accuracy when both the content answer and its paired evidence grounding answer are correct. The aggregation levels are: QA-level, a micro- average over all turns of all samples, where each turn contributes equally; Sample-level (mean), a macro-average of per-sample accuracies, so each sample contributes equally regardless of turn count; and Sample-level (0/1), which counts a sample correct only if every turn is correct. Evaluation protocol. We evaluate models in a multi-turn chat setting, following common proto- cols such as MT-Bench [39] and MT-Eval [19]. Prior user messages (questions and answer choices) and assistant responses are retained as conversation history using each modelâs native chat template, and the model is instructed to select one of A/B/C/D/E at each turn. The first-turn user message contains the patientâs entire sequence of discharge summaries, the first question, and its five an- swer choices; later turns contain only the next question and its answer choices. We use greedy decoding (temperature 0) wherever temperature can be set, except the DeepSeek-R1-Distill family, which uses 0.6 following the official recommendation [13]; gpt-5.4 and gpt-5.4-mini, whose temper- ature cannot be set, use the default API setting. Implementation details (e.g., inference framework, HIPAA-compliant API access, evaluation cost) are in Appendix H. 4.2 Main Results Table 2 reports the benchmark results across all 22 models. Content-question accuracy. For QA-level content-question accuracy, proprietary frontier models perform best, with gpt-5.4 and gemini-3-flash-preview achieving 98.33% and 96.17%, respectively. Among open-weight models, the Qwen3 family performs strongest, with the 80B, 30B, and 4B 7 Table 2: Main results on EHRNote-ChatQA (accuracy in %). Content: content-question accu- racy. Source: evidence-grounding accuracy. Paired: both the content answer and its paired evidence-grounding answer correct. QA-level micro-averages over all turns; Sample-level (mean) macro-averages per-sample accuracy; Sample-level (0/1) requires every turn in a sample correct. EHRNoteQA [20] is an external single-turn reference. Best per column and model-category in bold. Multi-Turn & Evidence-Grounding (Ours)Single-Turn QA-levelSample-level (mean)Sample-level (0/1) Model Params (Active / Total) Content Source Paired Content Source Paired Content Source Paired EHRNoteQA Proprietary frontier models gpt-5.4â98.3389.0287.9598.3589.3188.2488.8343.6440.0295.63 gpt-5.4-miniâ91.6374.0568.9691.8974.4069.4953.4613.348.6992.41 gemini-3-flash-previewâ96.1790.1087.0896.4190.4087.5971.8146.2935.7695.53 Open-weight general-purpose models Llama-4-Scout-17B-16E-Instruct17B / 109B92.1576.0071.3492.3476.4171.7954.6015.419.5193.56 Qwen3-Next-80B-A3B-Instruct3B / 80B95.4282.0979.2295.5382.3779.5671.2525.6519.8693.56 Qwen3-30B-A3B-Instruct-25073B / 30B93.1274.9070.7893.3275.1371.0558.9514.3710.0394.59 Qwen3-4B-Instruct-25074B88.4062.8255.9988.7563.1856.4742.305.893.0090.96 DeepSeek-R1-Distill-Llama-70B70B82.9165.6462.2283.5266.4563.1153.5711.488.3895.01 DeepSeek-R1-Distill-Qwen-32B32B78.8758.6653.2979.6059.1253.8738.887.244.5593.56 DeepSeek-R1-Distill-Qwen-14B14B76.1453.9447.5076.7854.0147.6927.613.101.4592.00 DeepSeek-R1-Distill-Llama-8B8B65.3141.5828.4766.0742.1629.176.620.410.1086.69 DeepSeek-R1-Distill-Qwen-7B7B27.0518.757.8528.7819.748.721.550.100.0080.15 Ministral-3-14B-Instruct-251214B31.5642.4615.5433.7244.2317.450.311.140.0091.68 Ministral-3-8B-Instruct-25128B49.5450.8631.8952.1553.0434.425.485.481.1490.33 Ministral-3-3B-Instruct-25123B75.0255.4644.3075.9955.9645.1516.443.721.7688.05 Phi-4-mini-instruct3.8B74.8148.8336.7875.4649.2337.4615.201.140.4183.58 Phi-3.5-mini-instruct3.8B76.9442.2632.5377.3442.7233.1017.271.340.3183.78 Medical-domain specialized models medgemma-27b-it27B87.7674.0065.5788.3474.4866.4140.9514.067.8691.06 medgemma-4b-it4B61.4144.2626.7461.7644.5026.934.550.410.0076.72 MediPhi-Instruct3.8B55.8031.2218.2955.3931.0618.083.720.310.0083.26 MediPhi-Clinical3.8B76.3445.6334.7376.7546.1935.3416.750.930.4184.41 MediPhi-PubMed3.8B77.0342.8733.0177.4643.4233.6617.681.240.3184.51 variants achieving 95.42%, 93.12%, and 88.40%, respectively; Llama-4-Scout-17B also performs strongly at 92.15%. The DeepSeek-R1-Distill family shows substantial variation across base model and size, with accuracies of 82.91%, 78.87%, 76.14%, 65.31%, and 27.05% for the 70B, 32B, 14B, 8B, and 7B variants, respectively. For the Ministral-3-Instruct family, the 14B and 8B models perform poorly compared with the 3B model, achieving 31.56%, 49.54%, and 75.02%, respectively. Among medical-domain open-weight models, medgemma-27b-it performs best, achieving 87.76% content accuracy. However, it underperforms similarly sized general-purpose models such as Qwen3-30B, which achieves 93.12%. Similarly, medgemma-4b-it obtains 61.41%, lower than general-purpose models of comparable size such as Qwen3-4B and Phi-4-mini, which achieve 88.40% and 74.81%, respectively. The MediPhi variants also do not consistently improve over their corresponding base model: MediPhi-Instruct, MediPhi-Clinical, and MediPhi-PubMed achieve 55.80%, 76.34%, and 77.03%, compared with 76.94% for Phi-3.5-mini-instruct. These results indi- cate that current medical-domain specialization does not necessarily improve performance on multi- turn clinical QA and can underperform strong general-purpose models of similar size. Evidence-grounding accuracy. QA-level Source accuracy is strongly correlated with Content accuracy (Spearman Ď = 0.89), indicating that models with higher content-question accuracy also tend to perform better on evidence grounding. Nevertheless, Source accuracy is consistently and substantially lower: across all 22 models, it is 17.56 percentage points (p) lower than Content accuracy on average. This gap varies with model capability: frontier models exhibit smaller gaps, such as 9.3 p for gpt-5.4 and 6.1 p for gemini-3-flash-preview, whereas smaller open-weight models show much larger gaps, such as 34.7 p for Phi-3.5-mini-instruct and 25.6 p for Qwen3-4B- Instruct. These results indicate that high content accuracy does not guarantee a model can locate the supporting discharge-summary sections, underscoring the need for evidence grounding evaluation. Multi-turn consistency evaluation. While QA-level results evaluate individual turns, practical clinical use requires models to remain correct across an entire multi-turn interaction. We therefore 8 020406080100 Content-question accuracy (%) DeepSeek-R1-Llama-70B DeepSeek-R1-Qwen-32B DeepSeek-R1-Qwen-14B Ministral-3-8B DeepSeek-R1-Qwen-7B Ministral-3-3B gpt-5.4 MediPhi Qwen3-Next-80B medgemma-27b Ministral-3-14B gpt-5.4-mini Qwen3-4B MediPhi-PubMed Llama-4-Scout-17B Qwen3-30B medgemma-4b Phi-4-mini-instruct Phi-3.5-mini-instruct MediPhi-Clinical DeepSeek-R1-Llama-8B gemini-3-flash-preview (a) Inter-turn content dependency 12345 Number of discharge summaries per patient 40 50 60 70 80 QA-level accuracy (%) (b) Accuracy by number of notes Content Source Paired Figure 3: Content-question inter-turn dependency and note-count effects on EHRNote-ChatQA. (a) Content accuracy at turn t conditioned on whether the immediately preceding content turn was answered correctly (green) or incorrectly (red); models are ordered by dependency gap â (green â red). (b) QA-level Content, Source, and Paired accuracy by number of discharge summaries per patient (1â5); lines show across-model accuracy averages and shaded bands showÂą1 standard error. examine Sample-level (0/1) accuracy, which requires every turn in a sample to be correct. This metric is substantially lower than QA-level accuracy: across all models, the average drop is 42.91, 47.67, and 41.20 p for Content, Source, and Paired accuracy, respectively. Under the strictest metric, Sample-level (0/1) Paired accuracy, even the best proprietary models achieve only 40.02% (gpt-5.4) and 35.76% (gemini-3-flash-preview), while the best open-weight models, Qwen3-Next- 80B-A3B-Instruct and Qwen3-30B-A3B-Instruct, achieve only 19.86% and 10.03%. These results show that current LLMs still have substantial room for improvement in consistently answering com- plete multi-turn clinical QA sequences. Comparison with single-turn discharge-summary QA. We compare EHRNote-ChatQA with EHRNoteQA [20], an external single-turn clinical QA benchmark containing clinician-curated ques- tions over patient discharge summaries. Compared with EHRNoteQA accuracy, QA-level Content accuracy on EHRNote-ChatQA is lower for most models, with an average drop of 14.06 p. The drop is especially large for several models, including DeepSeek-R1-Distill-Llama-8B, DeepSeek- R1-Distill-Qwen-7B, Ministral-3-14B-Instruct, and Ministral-3-8B-Instruct, which achieve 86.69%, 80.15%, 91.68%, and 90.33% on EHRNoteQA, but only 65.31%, 27.05%, 31.56%, and 49.54% on EHRNote-ChatQA, respectively. These results suggest that strong single-turn clinical QA perfor- mance does not reliably transfer to multi-turn clinical QA, underscoring the need for multi-turn, evidence-grounded clinical QA benchmarks. 4.3 Analysis Content-question inter-turn dependency. To test whether errors arise independently at each turn or propagate to later turns, we measure per-turn content accuracy conditioned on the correctness of the immediately preceding content turn (Figure 3(a)). For all 22 models, the dependency gap â = Acc(Q cont t | Q cont tâ1 correct) â Acc(Q cont t | Q cont tâ1 wrong) is positive. Averaged across models, content-question accuracy is 78.3% when the previous content turn was answered correctly versus 62.6% when incorrectly answered, yielding a gap of Ě â = 15.7 p. These results show that EHRNote-ChatQA exhibits a critical property of multi-turn multiple-choice QAâmulti-turn dependency and cross-turn error chainingâbuilt by design (Section 3.3). The effect is especially pronounced for the DeepSeek-R1-Distill family (â of 62.54, 52.91, 42.69 p for the 70B, 32B, 14B variants) and Ministral-3-8B-Instruct (32.29 p), indicating greater susceptibility to error cascades. Effect of note count. To examine how performance changes with the number of notes, we measure QA-level accuracy by the number of discharge summaries per patient (1â5; Figure 3(b)). Across all three metrics (Content, Source, and Paired), accuracy declines for most models (21 out of 22) as note count increases; averaged across models, accuracy drops by 10.6, 8.4, and 11.0 p, respectively, 9 Diagnosis Symptom Procedure Medication Microbiology Clin. Assess. Clin. Outcome Disch. Plan DeepSeek-R1-Distill-Qwen-7B Ministral-3-14B-Instruct-2512 MediPhi-Instruct medgemma-4b-it DeepSeek-R1-Distill-Llama-8B Ministral-3-8B-Instruct-2512 Phi-3.5-mini-instruct MediPhi-PubMed MediPhi-Clinical Phi-4-mini-instruct Ministral-3-3B-Instruct-2512 DeepSeek-R1-Distill-Qwen-14B DeepSeek-R1-Distill-Qwen-32B Qwen3-4B-Instruct-2507 DeepSeek-R1-Distill-Llama-70B medgemma-27b-it gpt-5.4-mini Qwen3-30B-A3B-Instruct-2507 Llama-4-Scout-17B-16E-Instruct Qwen3-Next-80B-A3B-Instruct gemini-3-flash-preview gpt-5.4 2927262226263031 3128323236303133 5658545353605359 6666635454646261 6867676161706265 5144494956465053 7979777172817977 7981777172807977 7880767169807877 7776767268807475 7874787272757576 7976777471807674 8179807877817777 9091918386928786 8682828280868480 9091908685888686 9492948890939191 9493958992959491 9393938990959292 9796969395979594 9796979597959697 9898999898989998 2.6 2.3 3.0 4.8 3.3 3.9 3.4 3.5 4.0 3.6 2.5 2.9 1.8 3.1 2.4 2.2 2.0 2.0 1.8 1.4 0.8 0.4 Content Diagnosis Symptom Procedure Medication Microbiology Clin. Assess. Clin. Outcome Disch. Plan 1819191719181822 4240384548364350 2829303232333134 4448424541424743 3844403943424342 4847465659445257 4144413945424343 4245424045424345 4549444348454644 4954494147495051 5555575655595255 5357535353565650 5963596060605553 6268666062636360 6668656668676561 7380747774767168 7275757675787170 7577767576767470 7781757877787270 8286858282838077 8892909091908991 8789889090898890 1.6 4.7 2.0 2.5 2.3 5.7 2.1 2.0 2.2 3.7 2.0 2.2 3.3 2.7 2.2 3.6 2.7 2.0 3.5 2.5 1.4 0.9 Source Diagnosis Symptom Procedure Medication Microbiology Clin. Assess. Clin. Outcome Disch. Plan 787687910 1512141621141619 1818171717211721 2832262521272826 2732292429302829 3128293441263237 3335322633343335 3436322733343435 3639343034363633 3743363033393639 4644474343454245 4850474643504944 5456555454565047 5762604954585553 6365616163656357 6672686662686259 6871726968736665 7273736870747065 7376717269756966 8083827778817774 8588888688868689 8688888889898888 1.1 2.7 1.8 3.3 2.3 4.8 2.7 2.7 2.7 3.8 1.6 2.8 3.1 4.0 2.5 4.1 2.9 3.0 3.4 3.0 1.2 0.8 Paired 020406080100 Accuracy (%) Figure 4: QA-level Content, Source, Paired accuracy (%) by question category per each model. Each cell shows the average per-category accuracy; Ď on the right of each panel reports the per- model standard deviation across the eight categories. Models are sorted by average Paired accuracy. from 1 to 5 notes. These results suggest that models are more likely to misidentify information as the number of notes increases, likely because they must process the full multi-note discharge-summary context and the questions require information across multiple admissions. Category-level performance. Accuracy shows limited variation across the eight question cate- gories (Figure 4). The per-model standard deviation across categories, averaged across all models, is 2.63, 2.63, and 2.74 p for Content, Source, and Paired accuracy, respectively, indicating that the accuracy across all eight categories is consistent within each model. This limited category-level variation suggests that performance is not dominated by any single clinical category. 5 Limitations EHRNote-ChatQA has several limitations. First, the benchmark is constructed from data from a single hospital system (MIMIC-IV), and therefore may not fully capture variation across institutions or EHR systems. Second, we focus on discharge summaries, which provide the overall clinical context of each admission, but do not cover other EHR sources such as progress notes, nursing notes, radiology reports, or structured EHR tables. Extending evidence-grounded multi-turn QA to these heterogeneous data sources may be an important direction for future work. Finally, although every QA sample is reviewed and revised by medical experts, each sample is reviewed by a single expert rather than multiple independent annotators. This trade-off was necessary given the large scale of the dataset (16,072 QA pairs) and the limited availability of medical experts. 6 Conclusion We introduced EHRNote-ChatQA, the first evidence-grounded, multi-turn, multi-note clinical QA benchmark over patient discharge summaries. The benchmark contains 967 patient-level multi-turn samples and 16,072 medical-expert-verified QA pairs across eight clinical categories, constructed through a medical-expert-informed pipeline and comprehensive expert review. Benchmarking 22 LLMs shows that current models still struggle more with evidence grounding than content answering, that errors compound across turns, and that strong single-turn clinical QA perfor- mance does not reliably transfer to this setting. These results highlight the need for benchmarks that evaluate multi-turn QA over longitudinal patient contexts with verifiable evidence grounding, and EHRNote-ChatQA provides a rigorous evaluation benchmark for evaluating clinical QA systems. 10 References [1] Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv. org/abs/2404.14219, 2(6):4, 2024. [2] Marah Abdin, Jyoti Aneja, Harkirat Behl, SĂŠbastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024. [3] Aaron Adcock, Aayushi Srivastava, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pande, Abhinav Pandey, Abhinav Sharma, Abhishek Kadian, Abhishek Kumawat, Adam Kelsey, et al. The llama 4 herd: Architecture, training, evaluation, and deployment notes. arXiv preprint arXiv:2601.11659, 2026. [4] Samuel Amouyal, Tomer Wolfson, Ohad Rubin, Ori Yoran, Jonathan Herzig, and Jonathan Berant. Qam- pari: A benchmark for open-domain questions with many answers. In Proceedings of the third workshop on natural language generation, evaluation, and metrics (GEM), pages 97â110, 2023. [5] Bernd Bohnet, Vinh Q Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, et al. Attributed question answering: Evaluation and modeling for attributed large language models. arXiv preprint arXiv:2212.08037, 2022. [6] Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettle- moyer. Quac: Question answering in context. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 2174â2184, 2018. [7] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, et al. Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261, 2025. [8] Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu, Asma Ben Abacha, Alessandro Sordoni, Lucas Caccia, François Beaulieu, Thomas Lin, Jens Kleesiek, and Paul Vozila. A modular approach for clinical slms driven by synthetic data with pre-instruction tuning, model merging, and clinical-tasks alignment. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 19352â19374, 2025. [9] John W Ely, Jerome A Osheroff, M Lee Chambliss, Mark H Ebell, and Marcy E Rosenbaum. Answer- ing physiciansâ clinical questions: obstacles and potential solutions. Journal of the American Medical Informatics Association, 12(2):217â224, 2005. [10] Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. Eli5: Long form question answering. In Proceedings of the 57th annual meeting of the association for computational linguistics, pages 3558â3567, 2019. [11] Scott L Fleming, Alejandro Lozano, William J Haberkorn, Jenelle A Jindal, Eduardo P Reis, Rahul Thapa, Louis Blankemeier, Julian Z Genkins, Ethan Steinberg, Ashwin Nayak, et al. Medalign: A clinician- generated dataset for instruction following with electronic medical records. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 22021â22030, 2024. [12] Tianyu Gao, Howard Yen, Jiatong Yu, and Danqi Chen. Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6465â6488, 2023. [13] Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforce- ment learning. arXiv preprint arXiv:2501.12948, 2025. [14] Sam Henry, Kevin Buchan, Michele Filannino, Amber Stubbs, and Ozlem Uzuner. 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records. Journal of the American Medical Informatics Association, 27(1):3â12, 2020. [15] Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM computing surveys, 55(12):1â38, 2023. [16] Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What dis- ease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11(14):6421, 2021. 11 [17] Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 2567â2577, 2019. [18] Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Sicheng Hao, Benjamin Moody, Brian Gow, et al. Mimic-iv, a freely accessible electronic health record dataset. Scientific data, 10(1):1, 2023. [19] Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, and Kam-Fai Wong. Mt-eval: A multi-turn capabilities evaluation benchmark for large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 20153â20177, 2024. [20] Sunjun Kweon, Jiyoun Kim, Heeyoung Kwak, Dongchul Cha, Hangyul Yoon, Kwanghyun Kim, Jeewon Yang, Seunghyun Won, and Edward Choi. Ehrnoteqa: An llm benchmark for real-world clinical practice using discharge summaries. Advances in Neural Information Processing Systems, 37:124575â124611, 2024. [21] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles, pages 611â626, 2023. [22] Eric Lehman, Vladislav Lialin, Katelyn Edelwina Legaspi, Anne Janelle Sy, Patricia Therese Pile, Nicole Rose Alberto, Richard Raymund Ragasa, Corinna Victoria Puyat, Marianne Katharina TaliĂąo, Isabelle Rose Alberto, et al. Learning to ask like a physician. In Proceedings of the 4th Clinical Natural Language Processing Workshop, pages 74â86, 2022. [23] Alexander H Liu, Kartik Khandelwal, Sandeep Subramanian, Victor Jouault, Abhinav Rastogi, Adrien SadĂŠ, Alan Jeffares, Albert Jiang, Alexandre Cahill, Alexandre Gavaudan, et al. Ministral 3. arXiv preprint arXiv:2601.08584, 2026. [24] Harsha Nori, Nicholas King, Scott Mayer McKinney, Dean Carignan, and Eric Horvitz. Capabilities of gpt-4 on medical challenge problems. arXiv preprint arXiv:2303.13375, 2023. [25] Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi- subject multi-choice dataset for medical domain question answering. In Conference on health, inference, and learning, pages 248â260. PMLR, 2022. [26] Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Med-halt: Medical domain hallu- cination test for large language models. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 314â334, 2023. [27] Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. emrqa: A large corpus for question answering on electronic medical records. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 2357â2368, 2018. [28] Siva Reddy, Danqi Chen, and Christopher D Manning. Coqa: A conversational question answering challenge. Transactions of the Association for Computational Linguistics, 7:249â266, 2019. [29] Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, CĂan Hughes, Charles Lau, et al. Medgemma technical report. arXiv preprint arXiv:2507.05201, 2025. [30] Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaugh- lin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al.Openai gpt-5 system card.arXiv preprint arXiv:2601.03267, 2025. [31] Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowl- edge. Nature, 620(7972):172â180, 2023. [32] Sarvesh Soni, Meghana Gudala, Atieh Pajouhi, and Kirk Roberts. Radqa: A question answering dataset to improve comprehension of radiology reports. In Proceedings of the thirteenth language resources and evaluation conference, pages 6250â6259, 2022. 12 [33] Sarvesh Soni, Soumya Gayen, and Dina Demner-Fushman. Overview of the archehr-qa 2025 shared task on grounded question answering from electronic health records. In Proceedings of the 24th Workshop on Biomedical Language Processing, pages 396â405, 2025. [34] Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. Asqa: Factoid questions meet long- form answers. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Pro- cessing, pages 8273â8288, 2022. [35] Weiyi Sun, Anna Rumshisky, and Ozlem Uzuner. Evaluating temporal relations in clinical text: 2012 i2b2 challenge. Journal of the American Medical Informatics Association, 20(5):806â813, 2013. [36] Ăzlem Uzuner, Imre Solti, and Eithon Cadag. Extracting medication information from clinical text. Journal of the American Medical Informatics Association, 17(5):514â518, 2010. [37] Ăzlem Uzuner, Brett R South, Shuying Shen, and Scott L DuVall. 2010 i2b2/va challenge on concepts, assertions, and relations in clinical text. Journal of the American Medical Informatics Association, 18(5): 552â556, 2011. [38] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. [39] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595â46623, 2023. 13 Supplementary Contents A EHRNote-ChatQA Dataset Samples15 B EHRNote-ChatQA vs. Prior Patient-Record QA Benchmarks27 C Discharge-Summary Structuring Schema28 D Multi-turn QA Templates29 D.1 Diagnosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .29 D.2 Symptom . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .32 D.3 Procedure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .35 D.4 Medication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 D.5 Microbiology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 D.6 Clinical Assessment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .43 D.7 Clinical Outcome . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .46 D.8 Discharge Plan . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .49 E Per-Step Details for LLM-Based Initial Data Generation52 F Per-Step Output Examples for LLM-Based Initial Data Generation54 F.1Step 1: Multi-turn Question Generation . . . . . . . . . . . . . . . . . . . . . . .54 F.2Step 2: Distractor-Direction Brainstorming . . . . . . . . . . . . . . . . . . . . .56 F.3Step 3: Content Answer-Choice Generation . . . . . . . . . . . . . . . . . . . . .59 F.4Step 4: Evidence-Grounding Answer-Choice Generation . . . . . . . . . . . . . .60 G Medical Expert Review and Revision61 G.1 Reviewer Instructions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .61 G.2 Reviewer Interface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .63 H Evaluation Details64 IPer-Step Prompts for LLM-Based Initial Data Generation65 I.1Step 1: Multi-turn Question Generation . . . . . . . . . . . . . . . . . . . . . . .65 I.2Step 2: Incorrect Answer Distractor Direction Brainstorming . . . . . . . . . . . .76 I.3Step 3: Content-Question Answer Choice Generation . . . . . . . . . . . . . . . .83 I.4Step 4: Evidence Grounding Source-Location Answer Choice Generation . . . . .89 14 A EHRNote-ChatQA Dataset Samples We provide three EHRNote-ChatQA samples below, spanning different question categories (Fig- ures 5, 6, and 7). Each turn is a 5-way multiple-choice question in chat format, with the correct option shown in bold. Odd-numbered turns are content questions; each is immediately followed by an evidence-grounding (source-location) question (even-numbered) that asks for the minimal set of notes and section headers supporting the preceding content answer. The text is reproduced from EHRNote-ChatQA; the only modification is that each noteâs chartdate is replaced by a placeholder ([chartdateN ] for Note #N ) to avoid disclosing patient-identifying dates. In the released dataset (un- der PhysioNet credentialed access) the questions and answer choices contain the actual chartdates. Category: Diagnosis Q1. During her first admission in [chartdate1], what was the primary diagnosis, and what underlying conditions were identified? A. The primary diagnosis was a right thalamic hemorrhage, which was found to be caused by Moyamoya syndrome and a long-segment dissection of the right internal carotid artery. B. The primary diagnosis was acute left-sided weakness, which workup revealed was caused by a right thalamic hemorrhage secondary to underlying Moyamoya syndrome. C. The primary diagnosis was a right thalamic hemorrhage, attributed to underlying Moyamoya syndrome and two small aneurysms found on angiography. D. The patientâs primary diagnosis was a right thalamic hemorrhage, attributed to a hypertensive emergency from previously undiagnosed and uncontrolled chronic hypertension. E. The patient was admitted for an evaluation of known Moyamoya disease, which was found to have caused a secondary right thalamic hemorrhage and the formation of two small aneurysms. Correct answer: C. Q2. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis B. Note #1 Chartdate: [chartdate1] Headers: Past Medical History, Social History C. Note #1 Chartdate: [chartdate1] Headers: History of Present Illness D. Note #1 Chartdate: [chartdate1] Headers: Discharge Diagnosis E. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course Correct answer: E. Q3. Did the patient have any subsequent hospitalizations for a similar problem? A. Yes, approximately nine years later, in [chartdate3], she presented with a week of headache and nausea and was diagnosed with a new left occipital intracerebral hemorrhage. B. Yes, in [chartdate2], she was readmitted with severe headache and fever and was diagnosed with EVD-associated meningitis after having ventricular drains placed. C. Yes, approximately three years later, in [chartdate2], she was admitted with a severe headache and vomit- ing and was diagnosed with an intraventricular hemorrhage. D. No, her underlying vascular issues were stable after the initial event. Her next admission was a planned hospital- ization for an elective revascularization procedure to treat her Moyamoya disease. E. Yes, she was readmitted approximately two years later for an acute ischemic stroke, which presented with new- onset aphasia and right-sided weakness, as had been a concern during her first admission. Correct answer: C. Q4. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis B. Note #2 Chartdate: [chartdate2] Headers: History of Present Illness C. Note #2 Chartdate: [chartdate2] Headers: Chief Complaint D. Note #2 Chartdate: [chartdate2] Headers:Family History E. Note #1 Chartdate: [chartdate1] Headers: History of Present Illness Correct answer: B. Q5. Were there any complications during that admission? A. Yes, her hospitalization was complicated by obstructive hydrocephalus requiring bilateral EVDs, an Ente- rococcus UTI, and EVD-associated meningitis. B. Yes, the ischemic stroke was complicated by hemorrhagic conversion, which led to brain swelling and required an emergent decompressive craniectomy. C. Yes, the main complications were metabolic, including severe hyponatremia down to 129 and delirium, which 15 was attributed to sleep deprivation in the ICU. D. Yes, the admission was complicated by obstructive hydrocephalus requiring EVDs, and she also developed hospital-acquired pneumonia that prolonged her ICU stay. E. Yes, during that admission she developed severe depression and anxiety with suicidal ideation, which required a psychiatry consultation and initiation of mirtazapine. Correct answer: A. Q6. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course B. Note #2 Chartdate: [chartdate2] Headers: Major Surgical or Invasive Procedure C. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis, Major Surgical or Invasive Procedure D. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course E. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis Correct answer: D. Q7. What organisms were found in relation to the infections? A. The UTI was caused by vancomycin-sensitive Enterococcus, but the meningitis was caused by a heavy growth of methicillin-resistant Staphylococcus aureus (MRSA). B. The UTI was caused by E. coli, a common pathogen, while the meningitis was caused by both Staphylococcus epidermidis and a fungal infection with Candida. C. The urinary tract infection was caused by Staphylococcus epidermidis, and the meningitis was caused by Entero- coccus sp. D. The UTI was caused by Enterococcus sp. and the meningitis was polymicrobial, involving Staphylococcus epidermidis, Corynebacterium, and Enterococcus. E. The urinary tract infection was caused by Enterococcus sp.; subsequent CSF cultures showed no growth and were considered contaminants after initial concerns. Correct answer: D. Q8. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis B. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course C. Note #2 Chartdate: [chartdate2] Headers: Physical Exam, Discharge Instructions D. Note #2 Chartdate: [chartdate2] Headers: Pertinent Results E. Note #1 Chartdate: [chartdate1] Headers: Pertinent Results Correct answer: D. Q9. How did the management of the central nervous system infection evolve? A. Given the diagnosis of enterococcal meningitis, the patient was treated with a synergistic combination of high- dose Ampicillin and Gentamicin after the EVD was replaced. B. The patient was empirically treated with Vancomycin and Meropenem. This was narrowed to Vancomycin based on culture results, but was later switched to Linezolid after follow-up cultures showed the development of van- comycin resistance. C. In addition to systemic IV Vancomycin, the patient also received intraventricular Vancomycin administered di- rectly into the EVD to achieve higher drug concentrations at the site of infection. D. The patient was started on Linezolid for meningitis. When she continued to spike fevers, her antibiotics were broadened to Vancomycin and Meropenem, and her EVD was replaced for source control. E. Initially, she received empiric Vancomycin and Meropenem, and her EVD was replaced. This was nar- rowed to Vancomycin per cultures, then switched to Linezolid for suspected drug fever. Correct answer: E. Q10. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis, Past Medical History B. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course C. Note #2 Chartdate: [chartdate2] Headers: Pertinent Results D. Note #2 Chartdate: [chartdate2] Headers: Discharge Medications E. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course Correct answer: B. Q11. What was the patientâs clinical course after that hospitalization? A. Approximately five and a half years later, in [chartdate3], she was admitted again after a week of headache and was found to have an ischemic stroke in the left occipital lobe. B. Her condition remained stable for several years. Her next known medical event was a scheduled admission for 16 the right-sided EDAS bypass surgery that had been planned after her first hemorrhage. C. About five and a half years later, in [chartdate3], she was readmitted for a new left occipital intracerebral hemorrhage with intraventricular extension. D. After recovering from meningitis, she was lost to follow-up until she presented with another severe headache and was found to have a ruptured aneurysm, which was then successfully treated with endovascular coiling. E. Her next hospitalization was for management of severe depression with suicidal ideation, which had become her primary medical issue after her prolonged and complicated recovery. Correct answer: C. Q12. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Discharge Diagnosis B. Note #3 Chartdate: [chartdate3] Headers: Discharge Diagnosis C. Note #3 Chartdate: [chartdate3] Headers: History of Present Illness D. Note #3 Chartdate: [chartdate3] Headers: Physical Exam, Family History E. Note #2 Chartdate: [chartdate2] Headers: History of Present Illness Correct answer: C. Q13. What did the evaluation of her underlying disease show during that latest admission? A. An MRI and MRA of the brain were performed, which showed the new hemorrhage and confirmed stability of the previously diagnosed Moyamoya disease without any new vascular changes. B. A conventional angiogram was performed which identified two 3m aneurysms in the basal ganglia, which were felt to be the cause of the hemorrhage. C. A diagnostic angiogram re-demonstrated her bilateral Moyamoya disease and also revealed a failure of her previ- ous right-sided EDAS bypass graft, which likely caused the new hemorrhage. D. A diagnostic angiogram was performed which re-demonstrated her bilateral Moyamoya disease. E. A diagnostic angiogram was performed which showed progression of her Moyamoya disease, with two new aneurysms having formed near the site of the new hemorrhage. Correct answer: D. Q14. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: Discharge Medications, Discharge Instructions B. Note #1 Chartdate: [chartdate1] Headers: Pertinent Results C. Note #3 Chartdate: [chartdate3] Headers: Major Surgical or Invasive Procedure D. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course E. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course Correct answer: E. Q15. Was a new long-term medication started to manage that condition? A. No, given her recent hemorrhage, the decision was made to hold all new antithrombotic or anticoagulant medica- tions to minimize the risk of re-bleeding. B. Yes, the neurology service cleared her to start aspirin 81 mg daily. C. Yes, she was started on Mirtazapine 7.5 mg at bedtime for long-term management. D. Yes, she was started on the antiplatelet agent Clopidogrel 75 mg daily instead of aspirin. E. Yes, she was started on Linezolid to be taken orally for an extended course as an outpatient. Correct answer: B. Q16. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: Discharge Instructions B. Note #3 Chartdate: [chartdate3] Headers: Discharge Medications C. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course D. Note #1 Chartdate: [chartdate1] Headers: Discharge Medications E. Note #3 Chartdate: [chartdate3] Headers: Past Medical History, Physical Exam Correct answer: C. Q17. What was the likely clinical reasoning for initiating this new treatment? A. The aspirin was started to maintain the patency of her right-sided EDAS bypass graft, which is standard practice after revascularization surgery. B. The goal of antiplatelet therapy in Moyamoya disease is to reduce the risk of future ischemic strokes. C. The Mirtazapine was initiated to help manage the patientâs depression, anxiety, poor sleep, and nausea, as recom- mended by the psychiatry service. D. The aspirin was started to prevent another hemorrhagic stroke, as its anti-inflammatory effects are thought to stabilize the fragile Moyamoya vessels. 17 E. The likely reason was for long-term prevention of deep vein thrombosis (DVT), as an oral alternative to the hep- arin the patient had received in her prior admission. Correct answer: B. Q18. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: Discharge Instructions B. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course C. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course D. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course E. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course, Note #3 Chartdate: [chartdate3] Head- ers: Brief Hospital Course Correct answer: E. Figure 5: Example multi-turn QA sample from the Diagnosis category. 18 Category: Symptom Q1. What were the main presenting symptoms for the admission on [chartdate1]? A. Abdominal distention and hyponatremia. B. A clogged nasojejunal feeding tube that was placed during a prior admission. C. Worsening abdominal distention and severe malnutrition. D. Confusion and somnolence consistent with an episode of hepatic encephalopathy. E. Abdominal distention and hyperkalemia. Correct answer: A. Q2. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Physical Exam B. Note #1 Chartdate: [chartdate1] Headers: Pertinent Results C. Note #1 Chartdate: [chartdate1] Headers: Chief Complaint D. Note #1 Chartdate: [chartdate1] Headers: History of Present Illness, Discharge Diagnosis E. Note #1 Chartdate: [chartdate1] Headers: Major Surgical or Invasive Procedure, Discharge Condition Correct answer: C. Q3. What was identified as the cause of those symptoms? A. Severe tricuspid regurgitation leading to congestive hepatopathy, based on initial TTE findings. B. Hepatorenal syndrome, which caused both fluid retention and impaired potassium excretion leading to hyper- kalemia. C. The patientâs severe hyponatremia was identified as the primary cause of the diuretic-refractory ascites. D. Decompensated alcoholic cirrhosis resulting in diuretic-refractory ascites. E. Mechanical obstruction of the nasojejunal feeding tube due to inadequate flushing and crushed medications. Correct answer: D. Q4. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Discharge Diagnosis, Discharge Instructions B. Note #1 Chartdate: [chartdate1] Headers: History of Present Illness C. Note #1 Chartdate: [chartdate1] Headers: Past Medical History, Major Surgical or Invasive Procedure D. Note #1 Chartdate: [chartdate1] Headers: Discharge Diagnosis E. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course Correct answer: E. Q5. How was the primary symptom managed during that admission? A. As the ascites was diuretic-refractory, the patient was evaluated for and scheduled for a transjugular intrahepatic portosystemic shunt (TIPS) procedure. B. The severe malnutrition was managed via the endoscopic placement of a Dobhoff feeding tube to provide enteral nutrition. C. The patient was managed with aggressive intravenous diuretics and a cardiology consultation to evaluate for tricuspid valve surgery. D. With a large-volume paracentesis that removed 9.75 L of ascitic fluid, as diuretics could not be used. E. With a large-volume paracentesis that removed 6.0 L of ascitic fluid, based on the volume removed in a recent ED visit. Correct answer: D. Q6. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course B. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course, Pertinent Results C. Note #1 Chartdate: [chartdate1] Headers: Major Surgical or Invasive Procedure D. Note #1 Chartdate: [chartdate1] Headers: Major Surgical or Invasive Procedure, Discharge Instructions E. Note #1 Chartdate: [chartdate1] Headers: Pertinent Results Correct answer: B. Q7. In the following admissions, is there evidence that this symptom burden continued? A. Yes, his hyponatremia persisted, with admission sodium levels of 121 mEq/L documented during both subsequent hospitalizations in [chartdate2]/[chartdate3]. B. Yes, the symptom burden of malnutrition continued, as he was readmitted twice in [chartdate2]/[chartdate3] for a clogged Dobhoff tube. C. No, his ascites was successfully managed as an outpatient, and his subsequent admissions in [chart- date2]/[chartdate3] were for the unrelated problem of a clogged feeding tube. 19 D. Yes, he required an 8 L therapeutic paracentesis during his admission on [chartdate2] and a 9 L paracentesis on [chartdate3]. E. Yes, he required a 9 L therapeutic paracentesis during his admission on [chartdate2] and an 8 L paracen- tesis on [chartdate3]. Correct answer: E. Q8. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course, Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course B. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course, Note #3 Chartdate: [chartdate3] Head- ers: Brief Hospital Course C. Note #2 Chartdate: [chartdate2] Headers: Major Surgical or Invasive Procedure, Note #3 Chartdate: [chartdate3] Headers: Major Surgical or Invasive Procedure D. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course E. Note #3 Chartdate: [chartdate3] Headers: Pertinent Results Correct answer: B. Q9. What was the documented reasoning for the persistence of this symptom and the need for frequent interven- tions? A. The patientâs ascites was documented as being diuretic-refractory because diuretic use was limited by per- sistent severe hyponatremia. B. The reasoning was the patientâs severe malnutrition, which could not be supported by oral intake alone, necessi- tating a feeding tube that was prone to clogging. C. The notes suggest the development of hepatorenal syndrome (HRS), which made his kidneys unable to excrete sodium and water, rendering diuretics ineffective. D. The documented reason was progressive right-sided heart failure due to severe tricuspid regurgitation, which was refractory to medical management. E. The recurrent clogging of the feeding tube was attributed to the formula, so it was changed to Osmolite 1.5, and a stricter flushing schedule was ordered. Correct answer: A. Q10. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course B. Note #1 Chartdate: [chartdate1] Headers: Major Surgical or Invasive Procedure C. Note #2 Chartdate: [chartdate2] Headers: History of Present Illness D. Note #1 Chartdate: [chartdate1] Headers: Discharge Diagnosis E. Note #1 Chartdate: [chartdate1] Headers: Medications on Admission, Pertinent Results Correct answer: A. Q11. Was a different type of intervention performed in a later admission to manage this persistent symptom? A. Yes, during the admission on [chartdate4], the patient underwent a planned transjugular intrahepatic portosystemic shunt (TIPS) procedure. B. Yes, the Dobhoff tube was replaced and advanced post-pyloric by interventional radiology in an attempt to prevent future clogging. C. Yes, a permanent indwelling peritoneal catheter was placed to allow for intermittent drainage of ascitic fluid at home by the patientâs caregiver. D. Yes, in a later admission, he underwent a right heart catheterization to obtain definitive pressures and was referred for consideration of tricuspid valve surgery. E. Yes, he underwent a TIPS revision with balloon dilatation of a previously placed stent during the admission on [chartdate4]. Correct answer: A. Q12. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis, Brief Hospital Course B. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course C. Note #4 Chartdate: [chartdate4] Headers: Major Surgical or Invasive Procedure D. Note #4 Chartdate: [chartdate4] Headers: Chief Complaint, Major Surgical or Invasive Procedure E. Note #4 Chartdate: [chartdate4] Headers: History of Present Illness Correct answer: D. Q13. What was the initial outcome of that intervention? A. The right heart catheterization revealed a normal gradient across the tricuspid valve, definitively ruling out signif- 20 icant tricuspid regurgitation. B. The immediate outcome was the development of worsening hepatic encephalopathy, with confusion and asterixis, requiring an increase in his lactulose dose. C. The initial outcome was a reduction in paracentesis volume from 8 liters to 4 liters, but with an ongoing need for weekly procedures. D. The procedure was successful, with a decrease in the right atrial pressure from 28 mmHg to 23 mmHg, although 8 L of ascites were still drained. E. The procedure was successful, with the portosystemic gradient decreasing from 10 to 5 mmHg and 8 L of ascites drained. Correct answer: E. Q14. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 Chartdate: [chartdate4] Headers: Discharge Condition, Discharge Diagnosis B. Note #4 Chartdate: [chartdate4] Headers: Pertinent Results C. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis D. Note #4 Chartdate: [chartdate4] Headers: Major Surgical or Invasive Procedure E. Note #4 Chartdate: [chartdate4] Headers: History of Present Illness, Discharge Instructions Correct answer: B. Q15. Did that intervention resolve the need for paracentesis? A. No, and the development of severe post-TIPS encephalopathy made him a poor candidate for further interven- tions, so weekly 8 L paracentesis continued. B. It is not documented if the intervention for urinary retention was fully successful long-term, though a single dose of Flomax did resolve the acute episode. C. Yes, the TIPS procedure was fully successful and completely resolved his ascites, with no further fluid reaccumu- lation noted in subsequent follow-up. D. No, it only partially resolved the need; paracentesis volume was reduced from 8 L to 4 L, but he still re- quired weekly procedures. E. No, the need for paracentesis continued, and in fact the required volume of fluid removal increased from 4 liters to 8 liters weekly. Correct answer: D. Q16. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #5 Chartdate: [chartdate5] Headers: Major Surgical or Invasive Procedure, Discharge Diagnosis B. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course C. Note #5 Chartdate: [chartdate5] Headers: Brief Hospital Course D. Note #5 Chartdate: [chartdate5] Headers: History of Present Illness E. Note #4 Chartdate: [chartdate4] Headers: Pertinent Results Correct answer: D. Q17. Because the symptom was not fully resolved, what further intervention was performed? A. The patient was referred for expedited liver transplant evaluation as the definitive treatment for his end-stage liver disease and refractory ascites. B. The patient underwent an emergency revision of the thrombosed TIPS shunt, during which the clot was removed and the stent was dilated to restore flow. C. The patient underwent placement of a second, parallel transjugular intrahepatic portosystemic shunt (TIPS). D. No further interventions for ascites were needed; the next procedure was a planned balloon dilatation of the ex- isting, well-functioning TIPS stent as routine maintenance. E. The patient was started on high-dose diuretic therapy with spironolactone 200 mg and furosemide 40 mg daily to augment the effect of the initial TIPS. Correct answer: C. Q18. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #5 Chartdate: [chartdate5] Headers: Major Surgical or Invasive Procedure B. Note #4 Chartdate: [chartdate4] Headers: Major Surgical or Invasive Procedure C. Note #5 Chartdate: [chartdate5] Headers: Major Surgical or Invasive Procedure, Discharge Instructions D. Note #5 Chartdate: [chartdate5] Headers: History of Present Illness E. Note #5 Chartdate: [chartdate5] Headers: Discharge Instructions Correct answer: D. 21 Q19. Did the patient develop any new symptoms immediately following the first of those major procedures? A. No, following the parallel TIPS procedure in the final admission, the patient reported feeling well with no new pain or other complaints documented. B. Yes, he developed severe hypokalemia with a potassium of 2.4 due to fluid shifts after the procedure, requiring aggressive intravenous potassium repletion. C. Yes, immediately after starting high-dose diuretics, he developed worsening kidney function and hyperkalemia, requiring the diuretics to be held. D. Yes, he developed post-procedural urinary retention that required multiple straight catheterizations. E. Yes, after the procedure he developed significant confusion and asterixis, consistent with worsening hepatic en- cephalopathy that required increasing his lactulose. Correct answer: D. Q20. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course B. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis C. Note #5 Chartdate: [chartdate5] Headers: Brief Hospital Course D. Note #4 Chartdate: [chartdate4] Headers: Pertinent Results E. Note #4 Chartdate: [chartdate4] Headers: Transitional Issues Correct answer: A. Figure 6: Example multi-turn QA sample from the Symptom category. 22 Category: Procedure Q1. Was any major surgical procedure performed for the patientâs left knee pain during the admission on [chart- date2]? A. Yes, the patient underwent a left unicompartmental knee arthroplasty after failing conservative treatments. B. No, a major procedure was not performed; the admission was for medical management of her knee pain, which included the use of a patient-controlled analgesia (PCA) pump. C. Yes, a left total knee arthroplasty was performed for management of her left knee pain. D. Yes, a right total knee arthroplasty was performed for management of her knee pain. E. Yes, a left total knee arthroplasty was performed to repair a pathologic fracture of the knee that occurred as a result of her underlying osteoporosis. Correct answer: C. Q2. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Social History, Family History B. Note #2 Chartdate: [chartdate2] Headers: Major Surgical or Invasive Procedure C. Note #2 Chartdate: [chartdate2] Headers: Chief Complaint, Major Surgical or Invasive Procedure, Discharge Condition D. Note #1 Chartdate: [chartdate1] Headers: Major Surgical or Invasive Procedure E. Note #2 Chartdate: [chartdate2] Headers: Chief Complaint, Major Surgical or Invasive Procedure Correct answer: E. Q3. What was the underlying reason for that operation? A. The procedure was indicated for severe rheumatoid arthritis, which had destroyed the left knee joint and was unresponsive to medical management. B. The operation was necessary to treat a pathologic fracture of the left knee caused by the patientâs severe osteo- porosis. C. The operation was an arthroscopic debridement for a meniscal tear and to remove loose cartilage causing mechan- ical symptoms in the left knee. D. The operation was a lumpectomy performed for Stage I infiltrating ductal carcinoma of the right breast. E. The procedure was performed for osteoarthritis of the left knee following the failure of conservative treat- ments. Correct answer: E. Q4. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #1 Chartdate: [chartdate1] Headers: History of Present Illness B. Note #2 Chartdate: [chartdate2] Headers: History of Present Illness C. Note #2 Chartdate: [chartdate2] Headers: Medications on Admission, Social History D. Note #2 Chartdate: [chartdate2] Headers: Chief Complaint E. Note #2 Chartdate: [chartdate2] Headers: Discharge Diagnosis Correct answer: B. Q5. What prophylactic treatments were administered after the procedure? A. She received perioperative IV antibiotics, and Lovenox for DVT prophylaxis for three weeks, to be followed by aspirin for an additional three weeks. B. She received Lovenox for DVT prophylaxis and was instructed to continue her home regimen of Fosamax and calcium to promote healing of the osteoporotic fracture. C. She received perioperative IV antibiotics, and aspirin for DVT prophylaxis for three weeks, to be followed by Lovenox for an additional three weeks. D. For DVT prophylaxis, she was started on the oral anticoagulant rivaroxaban 10 mg daily for three weeks and also received perioperative IV antibiotics. E. She was given perioperative IV antibiotics and started on warfarin for DVT prophylaxis, with instructions to monitor her INR until it was therapeutic. Correct answer: A. Q6. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Past Medical History, Physical Exam B. Note #2 Chartdate: [chartdate2] Headers: Discharge Medications C. Note #1 Chartdate: [chartdate1] Headers: Discharge Medications D. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course, Discharge Medications E. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course Correct answer: D. 23 Q7. What was the patientâs post-operative course and discharge disposition? A. Her post-operative course was uncomplicated with the incision healing well and pain adequately controlled; she was discharged in stable condition to her home with home health services. B. Her post-operative course was uncomplicated with the incision healing well and pain adequately controlled; she was discharged in stable condition to an extended care facility. C. The post-operative course was complicated by episodes of tachycardia and hypotension, requiring IV fluid resus- citation. After she was stabilized, she was discharged to an extended care facility. D. The course was uncomplicated, and her INR became therapeutic on post-operative day 4 after starting warfarin. She was then deemed stable for discharge to an extended care facility for rehabilitation. E. The course was complicated by a superficial wound infection which required a course of IV antibiotics. After showing improvement, she was discharged to an extended care facility to complete a course of oral antibiotics. Correct answer: B. Q8. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course B. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course, Discharge Disposition C. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course, Discharge Disposition D. Note #2 Chartdate: [chartdate2] Headers: Physical Exam, Discharge Condition E. Note #2 Chartdate: [chartdate2] Headers: Past Medical History, Medications on Admission Correct answer: C. Q9. In a subsequent admission on [chartdate3], was a new condition identified for which a procedure was recom- mended? A. Yes, she was newly diagnosed with chronic kidney disease of unclear etiology, and it was recommended that she undergo a renal biopsy as an outpatient. B. Yes, she was diagnosed with AV Nodal Reentrant Tachycardia (AVNRT), and the electrophysiology team recom- mended implantation of a permanent pacemaker to prevent future episodes. C. Yes, she was diagnosed with AV Nodal Reentrant Tachycardia (AVNRT) and an outpatient ablation was recommended. D. Yes, a 2.2 cm right thyroid nodule was identified on a CT scan, and a fine-needle aspiration biopsy was recom- mended for further evaluation. E. Yes, she was diagnosed with new-onset atrial fibrillation with rapid ventricular response, for which an electrical cardioversion was recommended by the cardiology team. Correct answer: C. Q10. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: Discharge Diagnosis B. Note #3 Chartdate: [chartdate3] Headers: Medications on Admission, Discharge Instructions C. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course D. Note #3 Chartdate: [chartdate3] Headers: Past Medical History, Social History E. Note #2 Chartdate: [chartdate2] Headers: Brief Hospital Course Correct answer: C. Q11. Was that recommended intervention ever performed? A. No, the notes indicate the patient opted for continued medical management with a beta-blocker and chose to defer the invasive procedure. B. Yes, an EP study was performed, but it failed to induce the tachycardia, so the ablation part of the procedure was aborted. C. Yes, the history in the subsequent admission note from [chartdate4] mentions the patient had a ârecent EP study and ablation.â D. No, the subsequent admission summary does not contain any record of a fine-needle aspiration of her thyroid nodule being performed. E. No, the subsequent admission note does not mention a cardioversion being performed, and instead notes she had an ablation, suggesting a change in the treatment plan. Correct answer: C. Q12. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis, Discharge Disposition B. Note #4 Chartdate: [chartdate4] Headers: History of Present Illness C. Note #2 Chartdate: [chartdate2] Headers: Major Surgical or Invasive Procedure D. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course 24 E. Note #4 Chartdate: [chartdate4] Headers: Major Surgical or Invasive Procedure Correct answer: B. Q13. After that procedure, did the patient experience any recurrence of related symptoms? A. Yes, because the ablation was not completed, she presented again with a sudden episode of supraventricular tachycardia with a heart rate of 200. B. Yes, on the admission following the procedure, she was tachycardic with a heart rate up to 104. C. Yes, she continued to experience left knee pain, for which she took tramadol as needed according to her medication list in the subsequent admission. D. No, the ablation was successful and she had no further documented cardiac symptoms; her subsequent admission was for an unrelated GI illness. E. Yes, she was tachycardic, and this was determined to be caused by post-ablation pericarditis, a known complication of the procedure. Correct answer: B. Q14. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: History of Present Illness B. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis, Past Medical History C. Note #4 Chartdate: [chartdate4] Headers: Physical Exam D. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course E. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course, Physical Exam Correct answer: E. Q15. How was that episode of tachycardia explained and managed? A. The recurrent SVT was managed with 6 mg of IV adenosine, and an electrophysiology consult was placed to discuss a repeat ablation procedure. B. The tachycardia was managed by restarting her on the beta-blocker metoprolol 25 mg daily, which had been previously discontinued. C. The episode was believed to be due to post-ablation atrial irritation, and she was started on a short course of colchicine for presumed pericarditis. D. The tachycardia was thought to be secondary to her viral gastroenteritis and it normalized with IV fluids. E. The episode was identified as AVNRT and was managed with 6 mg of IV adenosine, which converted her to normal sinus rhythm. Correct answer: D. Q16. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course B. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course C. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis D. Note #4 Chartdate: [chartdate4] Headers: Past Medical History, Medications on Admission E. Note #4 Chartdate: [chartdate4] Headers: Discharge Diagnosis, Discharge Instructions Correct answer: B. Q17. Was the patientâs cardiac medication regimen changed after the intervention? A. Yes, metoprolol 25 mg daily was restarted to manage her tachycardia, and her lisinopril was temporarily held to avoid hypotension. B. Yes, due to the recurrence of her arrhythmia, her metoprolol was restarted at a higher dose of 50 mg daily and she was referred back to electrophysiology. C. Yes, the metoprolol started before the procedure was stopped, as a subsequent note indicates she was no longer on a beta-blocker. D. No, her cardiac medication regimen was unchanged; she was continued on the metoprolol 25 mg daily that was initiated prior to the ablation. E. Yes, her lisinopril dose was decreased from 40 mg to 20 mg, and furosemide and nifedipine were stopped due to hypotension observed after the procedure. Correct answer: C. Q18. What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A. Note #3 Chartdate: [chartdate3] Headers: History of Present Illness, Note #4 Chartdate: [chartdate4] Headers: History of Present Illness B. Note #3 Chartdate: [chartdate3] Headers: Transitional Issues, Note #4 Chartdate: [chartdate4] Headers: Transi- tional Issues C. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course, Note #4 Chartdate: [chartdate4] Head- 25 ers: Brief Hospital Course D. Note #3 Chartdate: [chartdate3] Headers: Brief Hospital Course E. Note #4 Chartdate: [chartdate4] Headers: Brief Hospital Course Correct answer: C. Figure 7: Example multi-turn QA sample from the Procedure category 26 B EHRNote-ChatQA vs. Prior Patient-Record QA Benchmarks Table 3: Comparison with prior patient-record QA benchmarks: EHRNote-ChatQA is the first to combine multi-turn questions, multi-note grounding, and evidence-grounding questions. BenchmarkPatientsNotes Questions Source Multi- Note Multi- Turn Evidence Grounding Expert Verified Answer Format emrQA [27]â455,837 â Clinical Notes (i2b2/n2c2)Ă Ă âź ĂText Span DiSCQ [22]1141142,029 Discharge Summaries (MIMIC-I) Ă Ă ĂâNo Answer RadQA [32]â3,074 Radiology Reports (MIMIC-I)Ă Ă âźâText Span MedAlign [11]276â983 Full EHR (Stanford EHR)â Ă ĂâFree Text EHRNoteQA [20]962â962 Discharge Summaries (MIMIC-IV)â Ă ĂâMultiple Choice ArchEHR-QA [33]134134134 Discharge Summaries (MIMIC-I/IV) Ă ĂâFree Text EHRNote-ChatQA (Ours)967 2,278 ⥠16,072 Discharge Summaries (MIMIC-IV)âMultiple Choice â Auto-generated questions using rule-based template slot-filling. ⥠Total discharge summaries across 967 patients. âź (Evidence Grounding): the benchmarks provide extractive text spans as answers, but do not pair complete content answers with separate evidence-grounding questions that ask models to identify the supporting source for each answer. 27 C Discharge-Summary Structuring Schema Entity types and per-entity attributes (19 entity types): ⢠Diagnosis â is_main, status, finding_site, laterality, assertion, time ⢠Finding â is_main, status, finding_site, laterality, assertion, time ⢠Symptom â is_main, status, finding_site, laterality, assertion, time ⢠Procedure â is_main, status, finding_site, laterality, assertion, time ⢠Outcome â status, assertion, time ⢠Vital_Sign â status, value, assertion, time ⢠Physical_Exam â status, finding, finding_site, laterality, assertion, time ⢠Lab_Test â value, abnormal_flag, status, assertion, time ⢠Diagnostic_Imaging_Test â status, result, finding_site, laterality, assertion, time ⢠Specimen â source, assertion, time ⢠Microbiology_Test â organism_growth, result, assertion, time ⢠Microbiology_Organism â growth, assertion, time ⢠Microbiology_Antibiotics â dilution, sensitivity, assertion, time ⢠Activity â status, assertion, time ⢠Medication â strength, dosage, form_description, form, route, instruction, disp, refills, assertion, time ⢠Event â status, assertion, time ⢠Allergy â status, assertion, time ⢠Medical_Device â status, assertion, time ⢠Instruction â instruction_text, assertion, time Relationships (23 relationship types): ⢠Procedure/Medicationâ improvesâ Diagnosis/Symptom/Finding ⢠Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding ⢠Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding ⢠Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding ⢠Procedure/Medicationâ revealsâ Diagnosis/Symptom/Finding ⢠Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign âreveals â Diagnosis/Finding ⢠Specimenâ has_testâ Microbiology_Test/Lab_Test ⢠Microbiology_Testâ detectedâ Microbiology_Organism ⢠Microbiology_Organismâ tested_againstâ Microbiology_Antibiotics ⢠Microbiology_Organismâ revealsâ Diagnosis/Finding ⢠Diagnosis/Symptom/Findingâ causesâ Diagnosis/Symptom/Finding ⢠Diagnosis/Symptom/Findingâ causesâ Event ⢠Diagnosis/Symptom/Findingâ resulted_inâ Outcome/Finding/Activity ⢠Outcome/Findingâ afterâ Diagnosis/Symptom/Finding ⢠Outcome/Findingâ afterâ Procedure/Medication ⢠Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign âresulted_in â Procedure/Medication ⢠Procedure/Medicationâ causesâ Procedure/Medication ⢠Procedure/Medicationâ switched_toâ Procedure/Medication ⢠Procedure/Medicationâ causesâ Event ⢠Procedure/Medicationâ resulted_inâ Outcome/Finding ⢠Eventâ causesâ Diagnosis/Symptom/Finding ⢠Eventâ causesâ Procedure/Medication ⢠(any entity)â has_instructionâ Instruction Figure 8: Discharge-summary schema: 19 clinical entity types with per-entity attributes (top) and 23 relationship types linking entities (bottom). The is_main attribute marks principal diagnoses, symptoms, findings, and procedures of the admission. causes and resulted_in are distinguished as follows: causes denotes adverse effects or complications, whereas resulted_in denotes de- sired or expected outcomes. 28 D Multi-turn QA Templates The full set of multi-turn QA templates, defined in collaboration with medical experts, is shown below. The templates span all eight question categories in both single-note and multi-note settings, yielding 16 template sets of five templates each (80 templates in total), and build on the discharge- summary structuring schema (the comprehensive entity, attribute, and relationship definitions in Appendix C). For each template, every answer A# specifies the schema entities, attributes, and relationships that ground the answer to the paired question Q# in the source notes; attributes are written in dot notation after the entity (e.g., Medication.dosage), and relationships asâ¨entityâŠâ relationshipââ¨entity⊠triples. D.1 Diagnosis D.1.1 Single-note setting Template 1 â Treatment Response and Diagnostic Outcome Q1. What diagnosis was treated during admission? A1. (Diagnosis.is_main, Diagnosis.assertion) Q2. What medications or procedures were administered for the condition? A2. (Procedure/Medicationâ administered_forâ Diagnosis, Medication.dosage, route, time) Q3. What was the documented clinical reasoning for choosing that particular treatment approach over alterna- tives? A3. (Procedure/Medicationâ administered_forâ Diagnosis; clinical reasoning from Brief Hospital Course linking findings to treatment decision) Q4. Did the condition improve, worsen, or remain unchanged after treatment? A4. (Procedure/Medicationâ improves/worsensâ Diagnosis, Diagnosis.status over time, Outcomeâ after â Procedure/Medication, Procedure/Medicationâ resulted_inâ Outcome) Q5. Were there any treatment side effects or complications affecting the condition? A5. (Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding/Event) Q6. What was the final clinical outcome related to the condition at discharge? A6. (Diagnosisâ resulted_inâ Outcome, Outcome.status, time = discharge) Template 2 â Comorbidity Interaction and Diagnostic Interrelationships Q1. Were any comorbid diagnoses present during admission? A1. (Diagnosis.is_main = false, Diagnosis.assertion) Q2. Did any of those conditions contribute to or worsen another? A2. (Diagnosisâ causesâ Diagnosis, Diagnosisâ worsensâ Diagnosis) Q3. Was there any clinical evidence documented in the notes that supported that interaction between those conditions? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Examâ revealsâ Diagnosis; clinical reasoning connect- ing documented findings to the interaction) Q4. Did the interaction result in any complication or adverse event? A4. (Diagnosisâ causesâ Event, Diagnosisâ resulted_inâ Outcome) Q5. At discharge, which of those conditions were stabilized, and which required ongoing management? A5. (Diagnosis.status at discharge, Diagnosisâ has_instructionâ Instruction) Template 3 â Primary Diagnosis Identification and Clinical Basis Q1. What is the primary diagnosis for the patient? A1. (Diagnosis.is_main=true, Diagnosis.assertion, time=admission) Q2. Were there any symptoms or findings on presentation that led to the identification of that condition? A2. (Symptom/Finding â causes â Diagnosis; Symptom.time=admission; Physical_Exam.finding, Vi- tal_Signâ revealsâ Diagnosis) Q3. How did the documented evaluation narrow the differential to that condition? A3.(Lab_Test/Diagnostic_Imaging_Test/Microbiology_Test âreveals âDiagnosis; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result; clinical reasoning linking findings to diagno- sis confirmation) Q4. What treatments were given specifically for the condition during the hospitalization? A4. (Procedure/Medicationâ administered_forâ Diagnosis; Medication.dosage, Medication.route, Medica- tion.frequency; Procedure.status) Q5. What was the status of the condition at discharge? A5. (Diagnosis.status at discharge; Outcome.status; Diagnosisâ resulted_inâ Outcome) Template 4 â Differential Diagnosis Workup and Resolution Q1. Was any diagnosis initially suspected or considered as a working diagnosis when the patient was first 29 evaluated? A1. (Diagnosis.assertion=suspected/possible, Diagnosis.is_main, time=early admission) Q2. Were any other conditions considered as part of the differential during the workup? A2. (Diagnosis.assertion=possible/suspected/ruled_out; all non-primary Diagnosis entities, time=admission) Q3. What tests were performed to differentiate between those possibilities, and what were the results? A3.(Lab_Test/Diagnostic_Imaging_Test/Microbiology_Test âreveals âDiagnosis/Finding; Lab_Test.abnormal_flag; Specimenâ has_testâ Lab_Test/Microbiology_Test; test result values) Q4. Based on those results, were any conditions ruled out, and on what basis? A4. (Diagnosis.assertion=ruled_out; Lab_Test.result=negative/normal; Diagnostic_Imaging_Test â reveals â Finding that excludes diagnosis) Q5. Was any evidence documented that confirmed the final diagnosis? A5. (Diagnosis.status=confirmed; Lab_Test/Diagnostic_Imaging_Test/Physical_Examâ revealsâ Diagno- sis; key positive findings supporting confirmation) Q6. Were any incidental or secondary diagnoses identified during the workup? A6.(Diagnosis.is_main=false; Diagnostic_Imaging_Test â reveals â secondary Diagnosis/Finding; secondary Diagnosis.assertion, time relative to primary) Template 5 â Diagnosis Etiology, Complications, and Causal Chain Q1. What diagnosis was the patient admitted for, and what was it attributed to? A1.(Diagnosis.is_main,Diagnosis.assertion;Event/Diagnosis/Finding â causes â Diagnosis; time=admission) Q2. Were there any pre-existing conditions, prior medications, or procedures that may have precipitated the condition? A2. (Diagnosis/Symptom/Finding â causes â Diagnosis; Procedure/Medication â causes â Diagnosis; prior Diagnosis.status=active/chronic) Q3. How did the documented clinical course demonstrate the causal relationship between those contributing factors and the condition? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam â reveals â Diagnosis; temporal and causal evi- dence from notes linking precipitants to the diagnosis) Q4. Did the condition lead to any complications or secondary diagnoses during the hospitalization? A4. (Diagnosisâ causesâ Diagnosis/Symptom/Finding/Event; secondary Diagnosis.assertion; time=during admission) Q5. How were those complications managed, and what was their status at discharge? A5. (Procedure/Medicationâ administered_forâ secondary Diagnosis; Diagnosis.status at discharge; Out- comeâ afterâ Procedure/Medication) Q6. Were any instructions given to prevent recurrence or manage the underlying cause after discharge? A6. (Diagnosisâ has_instructionâ Instruction; Instruction.instruction_text for prevention, risk modification, or monitoring) D.1.2 Multi-note setting Template 1 â Longitudinal Problem Timeline and Status Transitions Q1. Across all admissions, what diagnoses did the patient have, and when did each first appear? A1. (Diagnosis.assertion; Diagnosis.time across notes; first-occurrence time) Q2. For those conditions, what was their clinical status at the time of each discharge? A2. (Diagnosis.status over time; Outcome.status; time = discharge) Q3. Among those conditions, did any recur or lead to re-hospitalization, and what was the interval between episodes? A3. (Diagnosis.time comparisons; Event.time related to recurrence) Q4. Was there any clinical evidence supporting recurrence or persistence of those conditions during subsequent admissions? A4. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis; time = admis- sion) Q5. How did the treatment approach evolve across admissions to address the persistence of those conditions? A5. (Medication/Procedure changes across admissions â administered_for â Diagnosis; clinical reasoning from notes linking prior treatment failure to subsequent treatment modification) Q6. By the most recent discharge, were any of those conditions still active, and what was the follow-up plan for them? A6. (Diagnosis.status = active; Diagnosisâ has_instructionâ Instruction) Template 2 â Diagnostic Reasoning and Differential Resolution Q1. What was the primary condition being evaluated in the most recent admission? A1. (Diagnosis.is_main; Diagnosis.assertion; time = admission) Q2. Were there any symptoms, findings, or clinical changes that prompted evaluation for that condition? A2. (Diagnosis/Symptom/Findingâ causesâ Diagnosis; time = admission) 30 Q3. What diagnostic tests were obtained to evaluate the condition, and were there any abnormal or positive results? A3.(Lab_Test/Diagnostic_Imaging_Test/Vital_Sign/Physical_Exam âreveals âDiagnosis; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result) Q4. Compared with the prior admission, were those findings new, worse, or unchanged? A4. (Diagnosis.status comparison; Lab_Test value trends; Diagnostic_Imaging_Test comparisons; Vital_Sign trends; Physical_Exam trends) Q5. What was the documented clinical reasoning that explained the change in the conditionâs status from the prior admission? A5. (clinical reasoning from notes linking cross-admission finding trends to the current condition status; causal factors documented) Q6. By discharge, was the condition confirmed, ruled out, or left as suspected? A6. (Diagnosis.status at discharge; Outcomeâ afterâ Diagnosis) Q7. Were any discharge instructions provided for monitoring or follow-up of the condition? A7. (Diagnosisâ has_instructionâ Instruction) Template 3 â Diagnosis Course and Treatment Response Across Admissions Q1. What was the primary diagnosis at admissionâ¨chartdate or admission_idâŠ? A1. (Diagnosis.is_main; Diagnosis.assertion; time =â¨admissionâŠ) Q2. Were there any symptoms, findings, or clinical events that led clinicians to identify the condition? A2. (Symptom/Findingâ revealsâ Diagnosis; Eventâ causesâ Diagnosis) Q3. Were any diagnostic tests performed that supported or confirmed the condition? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Examâ revealsâ Diagnosis; abnormal findings) Q4. What treatments were administered to manage the condition? A4. (Procedure / Medicationâ administered_forâ Diagnosis; Medication.dosage; Medication.route) Q5. During hospitalization, how did the condition respond to those treatments? A5. (Medication/Procedureâ improves or worsensâ Diagnosis; Diagnosis.status trend) Q6. What was the final status of the condition at discharge? A6. (Diagnosis.status at discharge; Outcome.status) Template 4 â Differential Diagnosis Reasoning and Confirmation Q1. Was any diagnosis initially suspected on presentation for admissionâ¨chartdate or admission_idâŠ? A1. (Diagnosis.is_main at admission; Diagnosis.assertion = suspected) Q2. During the evaluation of that suspected condition, were any other diagnostic possibilities considered? A2. (Diagnosis entities with assertion = possible/suspected/ruled out) Q3. Were any diagnostic tests ordered to differentiate between those possibilities? A3. (Lab_Test / Diagnostic_Imaging_Test / Microbiology_Test ordered; Specimenâ has_testâ Lab_Test) Q4. Based on those test results, were any possibilities ruled out, and why? A4. (Diagnosis.assertion = ruled out; Lab_Test.result; Diagnostic_Imaging_Test findings excluding diagnosis) Q5. Were any findings that ultimately confirmed the final diagnosis documented? A5. (Lab_Test / Diagnostic_Imaging_Test / Physical_Exam / Vital_Signâ revealsâ Diagnosis) Q6. During the workup for that condition, were any additional incidental findings or secondary diagnoses discovered? A6. (Diagnosis.is_main = false; Diagnostic_Imaging_Testâ revealsâ Finding/Diagnosis) Template 5 â Etiology, Complications, and Prevention of the Admission Diagnosis Q1. What diagnosis was the patient admitted for during admissionâ¨chartdate or admission_idâŠ? A1. (Diagnosis.is_main; Diagnosis.assertion; time = most recent admission) Q2. Were there any underlying conditions, events, or findings that contributed to the development of the con- dition? A2. (Diagnosis/Symptom/Findingâ causesâ Diagnosis; Eventâ causesâ Diagnosis) Q3. Were there any procedures or medications from prior admissions that may have contributed to the condi- tion? A3. (Procedure/Medicationâ causesâ Diagnosis; time from prior notes) Q4. What was the documented clinical reasoning that connected those contributing factors from the current and prior admissions to the current condition? A4.(clinical reasoning from notes linking cross-admission contributing factors to the diagnosis; Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Microbiology_Testâ revealsâ Diagnosis) Q5. Did the condition lead to any complications, secondary diagnoses, or clinical events? A5. (Diagnosisâ causesâ Diagnosis/Symptom/Finding/Event; Diagnosisâ resulted_inâ Outcome) Q6. Were any treatments administered for those complications, and if so, how were they addressed? A6. (Procedure / Medicationâ administered_forâ causal Diagnosis; Procedure/Medicationâ improves/- worsensâ condition) Q7. Were any preventive measures or follow-up instructions provided to reduce the risk of recurrence? A7. (Diagnosisâ has_instructionâ Instruction) 31 D.2 Symptom D.2.1 Single-note setting Template 1 â Symptom Progression and Secondary Complication During Hospitalization Q1. What was the initial presenting symptom at admission? A1. (Symptom.is_main, Symptom.time=admission, Symptom.assertion) Q2. Did the initial symptom or its underlying condition lead to additional symptoms or complications during the hospital stay? A2. (Symptomâ causesâ Symptom/Diagnosis; Diagnosisâ causesâ Symptom; Symptom.time progres- sion during admission) Q3. How were those secondary or evolving symptoms evaluated? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis/Finding related to secondary Symptom) Q4. What was the documented clinical reasoning that linked those evaluation findings to the secondary symp- toms? A4. (clinical reasoning from notes connecting documented evaluation results to the secondary symptom etiol- ogy; Diagnosis/Findingâ causesâ Symptom) Q5. Were any treatments directed at the secondary symptoms, and were they distinct from treatments for the primary symptom? A5. (Medication/Procedure â administered_for â secondary Symptom; compare with primary Symptom treatments) Q6. Did any treatments for the primary symptom worsen or cause the secondary symptoms? A6. (Medication/Procedure â causes â secondary Symptom; Medication/Procedure â worsens â Symp- tom) Q7. What was the status of all symptoms at the time of discharge? A7. (Symptom.status at discharge for each Symptom; Outcomeâ afterâ Symptom) Template 2 â Symptom Burden, Functional Impact, and Discharge Planning Q1. Were there any symptoms that impaired the patientâs functional status or activities of daily living during admission? A1. (Symptom.is_main, Symptom.status; Symptomâ resulted_inâ Activity limitation; Activity.status) Q2. How did those symptoms affect the patientâs ability to ambulate, self-care, or perform basic tasks? A2. (Symptomâ resulted_inâ Activity; Physical_Exam.finding related to functional performance) Q3. Were any interventions used to support functional recovery? A3. (Procedure/Medicationâ administered_forâ Symptom; Medical_Deviceâ improvesâ Symptom/Ac- tivity; Activity.status post-intervention) Q4. How did the care team determine the discharge disposition based on those functional limitations? A4. (Symptom â resulted_in â Outcome; clinical reasoning from notes linking the symptom burden and functional status to the disposition decision) Q5. Were any activity restrictions or precautions related to the patientâs symptoms specified in the discharge instructions? A5. (Symptom/Activityâ has_instructionâ Instruction; restrictions, duration, escalation criteria) Q6. Were any symptoms or warning signs documented that should prompt the patient to seek urgent care after discharge? A6. (Symptomâ has_instructionâ Instruction; threshold criteria for return to ED/clinic) Template 3 â Presenting Symptom, Diagnostic Workup, and Treatment Response Q1. What was the primary presenting symptom at admission, and when did it begin? A1. (Symptom.is_main = true; Symptom.time = onset relative to admission; Symptom.assertion = present) Q2. Was any condition identified as the cause of the symptom? A2. (Diagnosis/Findingâ causesâ Symptom) Q3. Were any laboratory tests, imaging studies, vital signs, or physical exam findings used to evaluate the symptom, and were any abnormal findings documented? A3.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis/Finding, Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result) Q4. Were any treatments administered specifically for the symptom? A4. (Medication/Procedureâ administered_forâ Symptom) Q5. Did the symptom improve, persist, or worsen during hospitalization? A5. (Medication/Procedureâ improves/worsensâ Symptom; Symptom.status trend) Q6. What was the status of the symptom at discharge? A6. (Symptom.status at discharge; Outcome.status) Q7. Were any discharge instructions provided regarding monitoring or recurrence of the symptom? A7. (Symptomâ has_instructionâ Instruction) Template 4 â Chronic Symptom Exacerbation During Single Admission 32 Q1. Did the patient have a chronic symptom prior to admission, and if so, what was its baseline severity and duration? A1. (Symptom.status = chronic; Symptom.time duration; Symptom.assertion) Q2. What changes in the symptom prompted admission? A2. (Symptom.status = worsened/exacerbated; Symptom.time = pre-admission period) Q3. Were any factors documented as contributing to the worsening of the symptom? A3. (Eventâ causesâ Symptom, Medication.status = non-compliance/discontinued, Diagnosisâ causesâ Symptom) Q4. What was the documented clinical reasoning that connected those contributing factors to the exacerbation? A4. (clinical reasoning from notes linking the documented triggers to the symptom worsening; temporal and causal evidence) Q5. Were any interventions implemented to manage the exacerbation? A5. (Medication/Procedureâ administered_forâ Symptom) Q6. How did the symptom respond to those interventions? A6. (Medication/Procedureâ improves/worsensâ Symptom; Symptom.status trend) Q7. Was long-term symptom management adjusted at discharge? A7. (Medication.time = discharge; Medication.dosage/route changes; Symptomâ has_instructionâ Instruc- tion) Template 5 â New Symptom Development During Hospitalization Q1. Did the patient develop any new symptoms during hospitalization? If so, what and when? A1. (Symptom.time = during admission; Symptom.assertion = new) Q2. Was any cause suspected for the new symptom? A2. (Medication/Procedure â causes â Symptom, Diagnosis â causes â Symptom, Event â causes â Symptom) Q3. Was any diagnostic evaluation performed for the new symptom? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ revealsâ Diagnosis/Finding) Q4. Did the documented findings confirm or refute the suspected cause based on those evaluation results? A4. (clinical reasoning from notes linking evaluation results to the suspected etiology; Diagnosis/Finding con- firmation or exclusion) Q5. Were any management changes made in response to the symptom? A5. (Medication.status = held/discontinued, Medicationâ switched_toâ Medication, Procedureâ admin- istered_forâ Symptom) Q6. Did the new symptom resolve before discharge? A6. (Symptom.status = resolved/persistent; Medication/Procedureâ improvesâ Symptom) Q7. Were any precautions given to prevent recurrence of the symptom? A7. (Symptom/Medicationâ has_instructionâ Instruction) D.2.2 Multi-note setting Template 1 â Readmission Trigger, Evaluation, and Treatment Response Q1. After discharge on admission â¨chartdate or admission_idâŠ, was there any symptom that prompted the patient to return to the hospital, and when did it begin? A1. (Symptom.is_main, Symptom.time relative to last discharge) Q2. Was any condition identified as the cause of the symptom? A2. (Diagnosis/Findingâ causesâ Symptom) Q3. Were any laboratory tests, imaging studies, physical examination findings, or vital sign abnormalities used to evaluate the condition? A3.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis/Finding; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result) Q4. Was any treatment or intervention administered for the symptom or its underlying condition? A4. (Medication/Procedureâ administered_forâ Symptom/Diagnosis/Finding) Q5. After the treatment was started, did the symptom improve, persist, or worsen, and was there any clinical evidence supporting the assessment? A5.(Medication/Procedure âimproves/worsens âSymptom;Symptom.status;relevant Lab_Test/Vital_Sign/Physical_Exam changes) Q6. Were any discharge instructions or return precautions provided regarding the symptom or its underlying condition? A6. (Symptomâ has_instructionâ Instruction; Diagnosisâ has_instructionâ Instruction) Template 2 â In-Hospital or Treatment-Related Symptom Development and Management Q1. During the hospitalization onâ¨chartdateâŠ, did the patient develop any new symptom after a medication or procedure was initiated, and when did that symptom appear? A1. (Symptom.time during admission; Medication/Procedure.time) Q2. Was any medication or procedure suspected to have caused that symptom? 33 A2. (Medication/Procedureâ causesâ Symptom/Finding; Symptomâ causesâ Diagnosis/Symptom/Find- ing) Q3. Were any diagnostic evaluations performed to assess the suspected adverse effect or to rule out complica- tions? A3. (Lab_Test/Diagnostic_Imaging_Test/Vital_Sign/Physical_Examâ revealsâ Diagnosis/Symptom/Find- ing; Lab_Test.abnormal_flag) Q4. What was the documented clinical reasoning that confirmed or refuted the relationship between the treat- ment and the new symptom? A4. (clinical reasoning from notes linking the temporal relationship and evaluation findings to the suspected causal mechanism; Medication/Procedureâ causesâ Symptom) Q5. Were any management changes made to address the symptom? A5. (Medication.status = held/discontinued; Medication.dosage; Medication â switched_to â Medication; Medication/Procedureâ improves/worsensâ Symptom) Q6. Were any monitoring or precautionary instructions given at discharge to prevent recurrence of the symp- tom? A6. (Symptomâ has_instructionâ Instruction; Medicationâ has_instructionâ Instruction) Template 3 â Recurrent Symptom Pattern Across Admissions Q1. Was there any symptom that occurred repeatedly across multiple hospitalizations for this patient? A1. (Symptom appears in >=2 notes; Symptom.time) Q2. Were there any events, medication changes, or clinical triggers that occurred before episodes of the symp- tom? A2. (Eventâ causesâ Symptom; Medication.status changes; Activity/environmental triggers across notes) Q3. How did the severity or clinical presentation of the symptom differ between hospitalizations? A3. (compare Symptom.status; Physical_Exam.finding; Vital_Sign.value associated Symptoms) Q4. Were any conditions associated with the symptom across admissions? A4. (Diagnosis/Finding/Symptomâ causesâ Symptom) Q5. How did the treatment approach evolve over time across admissions? A5. (Medication/Procedure â administered_for â Symptom; clinical reasoning from notes explaining how prior treatment outcomes informed subsequent treatment choices) Q6. Among those treatments, were any that improved the symptom, and were any that were ineffective or worsened it? A6. (Medication/Procedureâ improvesâ Symptom; Medication/Procedureâ worsensâ Symptom) Q7. Were there hospitalizations where that symptom was absent or well controlled? A7. (Symptom.status across notes; Symptomâ resulted_inâ Outcome/Finding) Q8. Were any preventive strategies or maintenance therapies recommended to prevent recurrence of the symp- tom? A8. (Medication.time = discharge; Symptomâ has_instructionâ Instruction) Template 4 â Symptom Progression and Secondary Symptom Development Q1. What was the initial presenting complaint at admissionâ¨chartdate or admission_idâŠ? A1. (Symptom.is_main; Symptom.time = admission) Q2. Was any condition identified as the cause of the symptom? A2. (Diagnosis/Symptom/Findingâ causesâ Symptom) Q3. Did the condition or symptom lead to any additional symptoms during the hospitalization? A3. (Diagnosis/Symptom/Findingâ causesâ Symptom) Q4. Did any treatment administered for the initial symptom lead to new symptoms? A4. (Medication/Procedureâ administered_forâ Symptom; Medication/Procedureâ causesâ Symptom) Q5. Were any treatments administered to manage the secondary or treatment-related symptoms? A5. (Medication/Procedureâ administered_forâ Symptom) Q6. Did those symptoms lead to any complications or clinical events affecting the hospital course? A6. (Symptomâ causesâ Event/Diagnosis/Symptom/Finding; Outcomeâ afterâ Symptom) Q7. By discharge, which symptoms had resolved and which remained active? A7. (Symptom.status at discharge; Outcome.status) Q8. Were any monitoring or follow-up instructions provided for those persistent symptoms? A8. (Symptomâ has_instructionâ Instruction) Template 5 â Chronic Symptom Burden and Functional Impact Across Admissions Q1. Was there any chronic or persistent symptom that the patient experienced across multiple admissions? A1. (Symptom present in >=2 notes; Symptom.status = chronic/persistent) Q2. How did the severity or frequency of that symptom change over time? A2. (Symptom.status trend; Symptom.time) Q3. Did the symptom affect the patientâs functional status or daily activities? A3. (Symptomâ resulted_inâ Activity; Activity.status) Q4. Were any treatments or interventions used to improve functioning despite the symptom? 34 A4. (Medication/Procedureâ administered_forâ Symptom; Activity.status) Q5. What was the documented clinical reasoning that connected the symptom trajectory to the care plan at the most recent discharge? A5. (clinical reasoning from notes linking the chronic symptom pattern to the discharge care plan; Symptom â resulted_inâ Outcome) Q6. Did the symptom result in changes in the patientâs care needs or living situation? A6. (Symptomâ resulted_inâ Outcome) Q7. Was any long-term management or monitoring plan recommended at the most recent discharge for the symptom? A7. (Symptomâ has_instructionâ Instruction) D.3 Procedure D.3.1 Single-note setting Template 1 â Procedure Indication, Supporting Evidence, and Discharge Planning Q1. Did the patient undergo any procedures during admission, and on what date or hospital day were they performed? A1. (Procedure.assertion; Procedure.time; Procedure.is_main; Procedure.status) Q2. What was the primary indication that led to the procedure? A2. (Procedureâ administered_forâ Diagnosis/Symptom/Finding; Diagnosis.status; Diagnosis.assertion) Q3. Were there any pre-procedural findings that supported the decision to perform the procedure, and were any abnormalities documented? A3.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â resulted_in â Procedure; Lab_Test.value; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result; Vital_Sign.value) Q4. What was the documented clinical reasoning that connected those abnormalities to the decision to proceed with the intervention? A4. (clinical reasoning from notes linking pre-procedural findings to the procedure decision; Procedure â administered_forâ Diagnosis/Symptom/Finding) Q5. Were there any intraoperative or immediate post-procedural complications, and if so, how were they identified? A5. (Procedure â causes â Event/Diagnosis/Symptom/Finding; Lab_Test/Imaging/Exam/Vitals â reveals â Diagnosis/Symptom/Finding) Q6. What was the patientâs clinical outcome by discharge related to the procedure? A6. (Procedure â resulted_in â Outcome/Finding; Outcome.status; Diagnosis/Symptom/Finding.status at discharge) Q7.Were any discharge instructions or follow-up plans provided that were specifically related to the procedure? A7. (Procedureâ has_instructionâ Instruction; time=discharge) Template 2 â Emergent or Urgent Procedure Decision-Making Q1. Were there any acute clinical changes or events during admission that led to an urgent or emergent proce- dure? A1.(Event â causes â Procedure; Diagnosis/Symptom/Finding â causes â Procedure; acute Vi- tal_Sign/Lab_Test abnormality) Q2. Were any diagnostic findings documented that supported proceeding with the urgent intervention? A2. (Lab_Test/Imaging/Physical_Examâ resulted_inâ Procedure; Lab_Test.abnormal_flag; Imaging.result) Q3. What procedure was performed, and when during the admission course? A3. (Procedure.assertion; Procedure.time; Procedure.is_main) Q4. Were there any intraoperative or immediate postoperative complications? A4. (Procedureâ causesâ Diagnosis/Symptom/Finding/Event; Vital_Sign changes; Lab_Test changes) Q5. How did the patientâs clinical status change after the procedure? A5.(Procedure â resulted_in â Outcome/Finding; Vital_Sign.value; Lab_Test.value trend; Physi- cal_Exam.finding changes) Q6. Was any monitoring, rehabilitation, or follow-up planned due to the urgency of the procedure? A6. (Procedureâ has_instructionâ Instruction; Activity.status; time=discharge) Template 3 â Surgical Approach, Perioperative Care, and Recovery Trajectory Q1. What surgical procedure was performed, and what were the operative approach or components? A1. (Procedure.assertion; Procedure.time; Procedure.is_main; Procedure.status) Q2. What condition was the surgery intended to address? A2. (Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q3. What perioperative management strategies were implemented? A3. (Procedure/Medication â administered_for â Diagnosis/Symptom/Finding; Medical_Device.status; Medication.route; Medication.time) 35 Q4. Did the procedure result in any expected postoperative findings or outcomes? A4. (Procedureâ resulted_inâ Outcome/Finding; Outcomeâ afterâ Procedure) Q5. Did any postoperative complication occur, and what was done in response? A5. (Procedureâ causesâ Diagnosis/Symptom/Finding/Event; Medication/Procedureâ administered_for â Diagnosis/Symptom/Finding) Q6. How was the recovery assessed at discharge based on the response to complication management? A6. (Activity.status; Outcome.status; Diagnosis/Symptom/Finding.status at discharge; clinical reasoning link- ing complication management to discharge status) Q7. What were the patientâs wound care, rehabilitation, or procedural precautions at discharge? A7. (Procedureâ has_instructionâ Instruction; Activityâ has_instructionâ Instruction) Template 4 â Procedure-Related Complication and Management Escalation Q1. Did the patient undergo a procedure during admission, and was there any sign that a complication had occurred afterward? A1. (Procedure.assertion, Procedure.time; Procedureâ causesâ Diagnosis/Symptom/Finding/Event; earliest time evidence) Q2. Was any diagnostic workup done to evaluate the complication, and what did it reveal? A2.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis/Finding; Lab_Test.value, Lab_Test.abnormal_flag, Diagnostic_Imaging_Test.result) Q3. What was the initial treatment approach for the complication, and was it effective? A3. (Medication/Procedure â administered_for â Diagnosis/Symptom/Finding; Medication/Procedure â improves/worsensâ Diagnosis/Symptom/Finding) Q4. If the initial approach was not fully effective, was the management escalated, and what clinical reasoning drove that decision? A4. (Procedure/Medication â switched_to â Procedure/Medication; clinical reasoning from notes linking inadequate response to the escalation decision) Q5. What was the final status of the complication at discharge, and were any additional follow-up steps recom- mended? A5.(Diagnosis/Symptom/Finding.status at discharge;Outcome.status;Procedure/Medication â has_instructionâ Instruction) Template 5 â Procedure Outcome, Functional Status Change, and Discharge Planning Q1.What was the patientâs functional or clinical status prior to the procedure that made intervention necessary? A1.(Diagnosis/Symptom/Finding.status pre-procedure;Physical_Exam.finding;Activity.status pre- admission) Q2. What procedure was ultimately performed, and what were the expected goals? A2. (Procedure.assertion, Procedure.is_main; Procedureâ administered_forâ Diagnosis/Symptom/Finding; Procedureâ resulted_inâ Outcome/Finding expected) Q3. How did objective measures such as labs, imaging, vitals, or exam change after the procedure? A3.(Lab_Test.valuepost-procedure;Vital_Sign.value;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result post-procedure) Q4. What was the patientâs functional and clinical status at discharge compared to admission? A4. (Outcome.status; Activity.status at discharge; Diagnosis/Symptom/Finding.status resolved/improved/per- sistent) Q5. What discharge disposition was chosen, and were any ongoing procedure-related instructions or follow-up given? A5. (Procedure â has_instruction â Instruction; Activity â has_instruction â Instruction; Outcome â afterâ Procedure; discharge disposition) D.3.2 Multi-note setting Template 1 â Procedure Indication, Course, and Discharge Follow-up Q1. Was any procedure performed during the most recent admission, and if so, when was it performed? A1. (Procedure.assertion, Procedure.time, Procedure.is_main) Q2. Was there any condition that led clinicians to perform the procedure? A2. (Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q3. Were there any laboratory tests, imaging, physical examination findings, or vital sign abnormalities that supported performing the procedure? A3.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â resulted_in â Procedure; Lab_Test.value; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result; Physical_Exam.finding; Vi- tal_Sign.value) Q4. What was the documented clinical reasoning that linked those abnormalities to the decision to proceed with the intervention? A4. (clinical reasoning from notes connecting pre-procedural findings to the procedure decision; Procedureâ 36 administered_forâ Diagnosis/Symptom/Finding) Q5. After the procedure was performed, did it lead to any complication, event, or new condition, and if so, how was that recognized? A5.(Procedure âcauses âEvent/Diagnosis/Symptom/Finding; Lab_Test/Diagnostic_Imaging_Test/Vital_Sign/Physical_Exam â reveals â Diagnosis/Symptom/Find- ing) Q6. Was any treatment or intervention administered to manage the complication? A6. (Procedure/Medication â administered_for â Diagnosis/Symptom/Finding; Procedure/Medication â resulted_inâ Outcome/Finding) Q7. Were any discharge instructions or follow-up plans documented for the procedure or the condition related to it? A7.(Procedure â has_instruction â Instruction; Diagnosis/Symptom/Finding â has_instruction â Instruction; time=discharge) Template 2 â Procedure-Related Complication Cascade and Escalation of Care Q1. At admission â¨chartdate or admission_idâŠ, was any procedure performed, and were any complications, events, or abnormal findings observed afterward? A1. (Procedure.time; Procedureâ causesâ Event/Diagnosis/Symptom/Finding) Q2. Were any diagnostic evaluations performed to further assess the complication or abnormal finding? A2.(Lab_Test/Imaging/Physical_Exam/Vital_Sign âreveals âDiagnosis/Symptom/Finding; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.result) Q3. Was any initial treatment or management strategy used to address the complication, and did it improve the condition? A3. (Medication/Procedure â administered_for â Diagnosis/Symptom/Finding; Medication/Procedure â improves/worsensâ Diagnosis/Symptom/Finding) Q4. If the initial management was not fully effective, was the care escalated, and what clinical reasoning drove that decision? A4. (Procedure/Medication â switched_to â Procedure/Medication; Event.status; clinical reasoning from notes linking inadequate response to escalation) Q5. What was the final outcome or clinical status after management of the complication? A5. (Outcome/Finding status; Instruction) Template 3 â Emergent Procedure Decision-Making and Execution Q1. Were there any clinical changes or acute events at admissionâ¨chartdate or admission_id⊠that necessitated an emergent or urgent procedure? A1. (Event â causes â Procedure, Diagnosis/Symptom/Finding â causes â Procedure, acute changes in Vital_Sign/Lab_Test) Q2. Were any diagnostic findings documented that supported performing the procedure? A2.(Lab_Test/Diagnostic_Imaging_Test/Physical_Exam âresulted_in âProcedure; Lab_Test.abnormal_flag; Diagnostic_Imaging_Test.finding) Q3. What procedure was performed to address the clinical change, and when during the hospitalization did it occur? A3. (Procedure.assertion, Procedure.time, Procedure.is_main) Q4. After the procedure was performed, were any intraoperative or immediate post-procedural complications observed? A4. (Procedureâ causesâ Diagnosis/Symptom/Finding/Event) Q5. Were any interventions performed to manage the complications following the procedure? A5. (Medication/Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q6. How did the patientâs clinical status change after the procedure? A6. (Procedureâ resulted_inâ Outcome/Finding) Q7. Were any monitoring or follow-up instructions provided for the procedure? A7. (Procedureâ has_instructionâ Instruction) Template 4 â Repeat Procedures and Recurrence-Driven Reintervention Across Admissions Q1. During the first admission, was any procedure performed forâ¨Diagnosis/Symptom/FindingâŠ, and what was the immediate post-procedure outcome at discharge? A1.(Procedure.time=admission; Procedure â administered_for â Diagnosis/Symptom/Finding; Out- come.status at discharge) Q2. After that procedure, did the condition recur in a later admission? A2. (Diagnosis/Symptom/Finding.status over time) Q3. Were there any tests, examinations, imaging studies, or vital signs that supported the decision to perform another procedure for the recurring condition? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ resulted_inâ Procedure) Q4. What was the documented clinical reasoning that explained why the condition returned despite the prior intervention? 37 A4. (clinical reasoning from notes linking factors such as incomplete resolution, non-adherence, or disease progression across admissions to the recurrence) Q5. Was a repeat procedure performed, or was a different procedure chosen? A5. (Procedureâ switched_toâ Procedure) Q6. After the procedure, did the patientâs clinical outcome improve compared with the outcome after the pre- vious procedure? A6. (Procedureâ resulted_inâ Outcome/Finding; Outcomeâ afterâ Procedure) Template 5 â Surgical Approach, Anesthesia Management, and Post-Surgical Outcomes Q1. Was any surgical procedure performed during admissionâ¨chartdate or admission_idâŠ, and what operative approach was used? A1. (Procedure.assertion, Procedure.time, Procedure.is_main) Q2. Was there any condition that led clinicians to perform the procedure? A2. (Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q3. What was the anesthesia type or anesthesia-related considerations for the surgical procedure? A3. (Event/Diagnosis/Finding related to anesthesia; Vital_Sign instability) Q4. During the surgery, were there any intraoperative findings or events? A4. (Procedureâ resulted_inâ Outcome/Finding during Procedure; Procedureâ causesâ Event/Diagno- sis/Finding) Q5. What was the immediate postoperative clinical status after the surgical procedure? A5. (Procedureâ resulted_inâ Outcome/Finding) Q6. Did the surgical procedure result in any postoperative complications? A6. (Procedureâ causesâ Diagnosis/Symptom/Finding/Event) Q7. By discharge, how had the patientâs condition changed after the procedure? A7. (Outcome.status; Diagnosis/Symptom/Finding.status) Q8. Were any discharge instructions or follow-up plans provided for the procedure? A8. (Procedureâ has_instructionâ Instruction) D.4 Medication D.4.1 Single-note setting Template 1 â Inpatient Initiation, Dose Adjustment, and Tolerability Q1. Were any medications initiated during admission to treatâ¨Diagnosis/Symptom/FindingâŠ? A1. (Medication.time = inpatient, Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q2. What were the initial dose, route, and frequency? A2. (Medication.strength, dosage, route, frequency, form) Q3. Did the patient develop any side effects or intolerance after initiation? A3. (Medicationâ causesâ Symptom/Finding/Event, Symptom.status; Lab_Test.abnormal_flag if applica- ble) Q4. Was the regimen modified in response to those adverse effects? A4. (Medication.dosage change with time; Medicationâ switched_toâ Medication; clinical reasoning from notes linking the adverse effect to the dose adjustment) Q5. Did the adjusted regimen improve the medical condition? A5. (Medicationâ improvesâ Diagnosis/Symptom/Finding, Medicationâ resulted_inâ Outcome) Q6. What was the final discharge medication regimen? A6. (Medication.time = discharge, Medication.strength, dosage, route, frequency, disp, refills) Template 2 â Diagnostic Findings Driving Medication Adjustment Q1. What medications was the patient taking on admission? A1. (Medication.time = admission) Q2. Were there any abnormal lab tests, imaging findings, or vital sign changes that prompted medication changes? A2. (Lab_Test/Diagnostic_Imaging_Test/Vital_Signâ resulted_inâ Medication change) Q3. What was the documented reasoning for the particular adjustment chosen in response to those abnormal findings? A3. (clinical reasoning from notes linking the specific abnormal values to the treatment decision; Medica- tion.dosage/route/frequency change rationale) Q4. After those medication adjustments were made, did the relevant clinical parameters improve? A4. (Medicationâ improvesâ Diagnosis/Symptom/Finding, Improved Lab_Test/Vital_Sign trends) Q5. Were any time-limited medications prescribed? A5. (Medication.instruction including duration, stop date) Q6. Were any monitoring instructions provided at discharge? A6. (Medicationâ has_instructionâ Instruction) 38 Template 3 â Medication Safety, Allergies, and Adverse Effects Q1. Does the patient have any medication allergies? A1. (Allergy.status, assertion) Q2. Were any medications avoided or modified due to the allergy history? A2. (Medication.assertion with avoidance) Q3. Did the patient experience any adverse medication-related complications during the stay? A3. (Medicationâ causesâ Diagnosis/Symptom/Event/Finding) Q4. Was a replacement medication selected, and what was the documented reasoning for that choice? A4. (Medicationâ switched_toâ Medication; clinical reasoning from notes explaining why the alternative was chosen given the adverse effect and allergy profile) Q5. Did the alternative medication lead to clinical stabilization? A5. (Medicationâ improvesâ Diagnosis/Symptom/Finding, Medicationâ resulted_inâ Outcome) Q6. Were any safety instructions provided regarding medication use? A6. (Medicationâ has_instructionâ Instruction) Template 4 â Discharge Medication Reconciliation and Patient Instructions Q1. What medications were prescribed at discharge? A1. (Medication.time = discharge; Medication.dosage, Medication.route, Medication.frequency, Medica- tion.disp, Medication.refills) Q2. Which of those medications were newly started during admission, and which were continued from before? A2. (Medication.assertion; Medication.time comparison) Q3. What condition was each newly started medication intended to treat? A3. (Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q4. Were any pre-admission medications discontinued at discharge, and if so, why? A4.(Medication.assertion = discontinued; Medication â causes â Diagnosis/Symptom/Finding; Al- lergy.status) Q5. Were any instructions given to the patient regarding the discharge medications? A5. (Medicationâ has_instructionâ Instruction) Template 5 â Medication for Complication Management During Hospitalization Q1. Did the patient develop any complications or new conditions during the hospital stay? A1. (Diagnosis/Symptom/Finding.time = inpatient; Procedure/Medication â causes â Diagnosis/Symp- tom/Finding) Q2. Were any medications started to manage those complications? A2. (Medication.time; Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q3. What was the documented reasoning for choosing that treatment approach for the complication? A3. (clinical reasoning from notes linking the complication characteristics to the treatment choice; Medication â administered_forâ Diagnosis/Symptom/Finding) Q4. Did the treatment improve the patientâs condition? A4. (Medicationâ improvesâ Diagnosis/Symptom/Finding; Medicationâ resulted_inâ Outcome) Q5. Were the complication-related medications continued or stopped at discharge? A5. (Medication.time = discharge; Medication.assertion) D.4.2 Multi-note setting Template 1 â Medication Intolerance, Adverse Effects, and Regimen Switching Q1. Was any medication initiated to treat â¨Diagnosis/Symptom/Finding⊠during admission â¨chartdate or admission_idâŠ? A1. (Medication.time; Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q2. After that medication was started, did the patient develop any adverse symptoms, complications, or clinical events? A2. (Medicationâ causesâ Diagnosis/Symptom/Event/Finding) Q3. Were any clinical findings or diagnostic abnormalities documented that suggested the adverse effect was related to the medication? A3. (Symptom.status; Lab_Test.abnormal_flag = abnormal; Medicationâ causesâ Event) Q4. Was the medication switched in response, and if so, what was the documented reasoning for choosing the replacement? A4. (Medication.status; Medication â switched_to â Medication/Procedure; clinical reasoning from notes explaining the rationale for the alternative selection) Q5. After the alternative treatment was started, did the patientâs condition improve without recurrence of the adverse effect? A5.(Medication/Procedure â improves â Diagnosis/Symptom/Finding; Medication/Procedure â re- sulted_inâ Outcome) Q6. Which medication regimen was continued at discharge for management of the patientâs condition? A6. (Medication.time = discharge) 39 Template 2 â Diagnostic Evidence Driving Medication Adjustment Q1. What medications was the patient taking at the time of admissionâ¨chartdate or admission_idâŠ? A1. (Medication.time = admission) Q2. While the patient was receiving those medications, were there any abnormal laboratory results, imaging findings, physical examination findings, or vital sign changes that prompted clinicians to adjust treatment? A2. (Lab_Test/Diagnostic_Imaging_Test/Vital_Sign/Physical_Examâ resulted_inâ Medication change) Q3. Were the dose, frequency, or route modified in response to those abnormal findings? A3. (Medication.dosage change; Medication.frequency change; Medication.route change) Q4. After those medication adjustments were made, did the relevant clinical parameters improve? A4. (Medicationâ improvesâ Diagnosis/Symptom/Finding; Lab_Test/Vital_Sign trend improvement) Q5. At discharge, what medications and dosing regimens were prescribed following those adjustments? A5. (Medication.time = discharge; Medication.dosage/frequency) Q6. Were any monitoring instructions provided related to those medications? A6. (Medicationâ has_instructionâ Instruction) Template 3 â Discharge Medication Reconciliation and Allergy Safety Decisions Q1. What medications were prescribed at discharge during admissionâ¨chartdate or admission_idâŠ? A1. (Medication.time = discharge) Q2. At the next admission discharge, were any of those medications added, discontinued, or changed? A2. (Medication.time = discharge comparison; Medication.status changes) Q3. Were there any conditions or clinical findings that prompted clinicians to add those new medications? A3. (Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q4. What was the documented reasoning that explained the evolution of the medication regimen from one discharge to the next? A4. (clinical reasoning from notes linking cross-admission clinical changes to the medication reconciliation decisions; Allergy.status; Medicationâ causesâ Event) Q5. Were the doses or frequencies of those medications modified to improve safety or adherence? A5. (Medication.dosage comparison; Medication.frequency comparison) Q6. At the final discharge, were any medication continuation or monitoring instructions provided? A6. (Medicationâ has_instructionâ Instruction) Template 4 â Medication Adherence, Outpatient Failure, and Readmission Risk Q1. What medications was the patient discharged on during admissionâ¨chartdate or admission_idâŠ? A1. (Medication.time = discharge) Q2. Before the next hospitalization, did the patient stop taking or inconsistently take any of those medications? A2. (Medication.status = non-adherence) Q3. Did that medication non-adherence precede worsening symptoms or acute clinical events? A3. (Medication.status = non-adherence; Medicationâ causesâ Event/Diagnosis/Symptom/Finding) Q4. How did clinicians address the adherence issue during readmission? A4. (Medication.time; Medicationâ switched_toâ Medication; clinical reasoning from notes linking non- adherence pattern to the revised treatment approach) Q5. After medication adherence was restored or treatment was adjusted, did the patientâs clinical status im- prove? A5. (Medicationâ improvesâ Diagnosis/Symptom/Finding; Outcome.status) Q6. Were any adherence counseling or follow-up instructions provided at discharge to prevent recurrence? A6. (Medicationâ has_instructionâ Instruction) Template 5 â Inpatient Medication Adjustment and Discharge Treatment Planning Q1. What medications was the patient taking on admissionâ¨chartdate or admission_idâŠ? A1. (Medication.time = admission; Medication.assertion) Q2. During hospitalization, were there any changes to the dose, route, or frequency of those medications? A2. (Medication.dosage change; Medication.route change; Medication.frequency change) Q3. Were there any diagnostic findings or vital sign abnormalities that prompted clinicians to make those med- ication changes? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ resulted_inâ Medication) Q4. After those medication adjustments were implemented, how did the patient clinically respond? A4. (Medication â improves/worsens â Diagnosis/Symptom/Finding; Medication â resulted_in â Out- come/Finding) Q5. Based on that response, what medications were ultimately prescribed at discharge? A5. (Medication.time = discharge) Q6. For the medications prescribed at discharge, what conditions were they intended to treat? A6. (Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q7. Were any discharge instructions given regarding those medications? A7. (Medicationâ has_instructionâ Instruction) 40 D.5 Microbiology D.5.1 Single-note setting Template 1 â Empiric Antibiotic Initiation for Suspected Infection Q1. Was any infection suspected or diagnosed during admission? A1. (Diagnosis.assertion, Diagnosis.status, Diagnosis.time) Q2. Were any specimens collected to work up the suspected infection? A2. (Specimen.source, Specimen.assertion; Specimenâ has_testâ Microbiology_Test/Lab_Test) Q3. Was any empiric therapy started, and for what indication? A3. (Medication.assertion, Medication.time; Medicationâ administered_forâ Diagnosis) Q4. What was the documented reasoning for choosing that particular empiric regimen? A4. (clinical reasoning from notes linking the suspected infection type, patient risk factors, and documented rationale to the empiric therapy selection) Q5. Was the empiric treatment continued, changed, or stopped by discharge? A5. (Medication.status at discharge; Medicationâ switched_toâ Medication) Q6. What treatment regimen was the patient discharged on, and for how long? A6. (Medication.time, Medication.instruction, Medication.duration at discharge) Template 2 â Culture Result Driving Therapy Modification Q1. Were any specimens sent for microbiologic testing during admission, and if so, when? A1. (Specimen.source, Specimen.time; Specimenâ has_testâ Microbiology_Test) Q2. Were any organisms identified from those cultures? A2. (Microbiology_Testâ detectedâ Microbiology_Organism; Microbiology_Organism.growth, Microbi- ology_Test.result) Q3. Were any susceptibility or resistance patterns reported for the organism? A3.(Microbiology_Organism âtested_against âMicrobiology_Antibiotics;Microbiol- ogy_Antibiotics.sensitivity, Microbiology_Antibiotics.dilution) Q4. Was the treatment regimen adjusted based on those susceptibility results? A4. (Medication â switched_to â Medication; Medication.assertion, Medication.time; clinical reasoning from notes linking susceptibility findings to therapy change) Q5. Did the patientâs infection-related signs and symptoms improve after the adjustment? A5. (Medicationâ improvesâ Diagnosis/Symptom/Finding; Vital_Sign/Lab_Test trends; Outcome.status) Template 3 â Culture-Negative Infection Workup and Clinical Decision-Making Q1. Was an infection suspected during admission, and what was the clinical basis? A1. (Diagnosis.assertion; Symptom.assertion; Physical_Exam/Vital_Sign/Lab_Test â reveals â Diagno- sis/Finding) Q2. What microbiology tests were performed, and what were the results? A2. (Specimenâ has_testâ Microbiology_Test; Microbiology_Test.result, Microbiology_Organism.growth) Q3. Since those cultures did not identify an organism, was there other clinical evidence supporting or refuting the infection? A3.(Lab_Test/Diagnostic_Imaging_Test/Vital_Sign â reveals â Diagnosis/Finding; Lab_Test.value, Lab_Test.abnormal_flag) Q4. What was the documented clinical reasoning for the decision to continue or discontinue treatment in the setting of those inconclusive results? A4. (Medication.status, Medication.assertion; clinical reasoning from notes explaining the treatment decision in the setting of negative cultures) Q5. What was the patientâs clinical status at discharge, and was any follow-up for the suspected infection rec- ommended? A5. (Outcome.status; Diagnosis/Medicationâ has_instructionâ Instruction) Template 4 â Infection as a Complication of Primary Admission Q1. What was the primary reason for this patientâs admission? A1. (Diagnosis.is_main, Diagnosis.assertion; Event.status) Q2. Did the patient develop a new infection during the hospitalization as a complication? A2. (Diagnosis.assertion, Diagnosis.time; Eventâ causesâ Diagnosis; Procedureâ causesâ Diagnosis) Q3. Were any specimens collected and what microbiologic testing was performed for the new infection? A3.(Specimen.source, Specimen.time; Specimen â has_test â Microbiology_Test; Microbiol- ogy_Test.result) Q4. Was any treatment started for the infection, and how did it affect the patientâs clinical course? A4. (Medicationâ administered_forâ Diagnosis; Medicationâ improvesâ Diagnosis/Symptom/Finding; Outcome.status) Q5. Was any follow-up imaging, laboratory testing, or repeat culture recommended at discharge to confirm resolution? 41 A5.(Diagnostic_Imaging_Test/Lab_Test/Microbiology_Test â has_instruction â Instruction; Instruc- tion.instruction_text) Template 5 â Monitoring Microbiologic Clearance Q1. Was any infection identified during admission, and was an organism detected? A1. (Diagnosis.assertion; Microbiology_Testâ detectedâ Microbiology_Organism) Q2. Was any treatment started, and when? A2. (Medication.assertion/time; Medicationâ administered_forâ Diagnosis) Q3. Were any follow-up microbiology tests performed? A3. (Microbiology_Test.time after initial test) Q4. Did follow-up testing show persistent organism growth or microbiologic clearance? A4. (Microbiology_Test.organism_growth/result; Diagnosis.status) Q5. How were the duration and continuation of therapy determined based on those follow-up results? A5. (Medication.duration/status; clinical reasoning from notes linking clearance/persistence findings to the treatment duration decision) Q6. What was the infection status and treatment plan at discharge? A6. (Outcome.status; Medication.time = discharge; Diagnosis/Medicationâ has_instructionâ Instruction) D.5.2 Multi-note setting Template 1 â Infection Workup from Suspicion to Targeted Therapy Q1. During admissionâ¨chartdate or admission_idâŠ, was any infection suspected or diagnosed? A1. (Diagnosis.assertion, Diagnosis.time; Event.status) Q2. Was any specimen collected for microbiologic testing to evaluate the infection? A2. (Specimen.source; Specimenâ has_testâ Microbiology_Test) Q3. From testing of the specimen, was any organism identified? A3. (Microbiology_Testâ detectedâ Microbiology_Organism; Microbiology_Test.organism_growth/result) Q4. Before identification of the organism, was any empiric therapy started for the infection? A4. (Medication.assertion; Medication.time; Medicationâ administered_forâ Diagnosis) Q5. After susceptibility results for the organism were available, was the treatment changed or narrowed? A5. (Microbiology_Organismâ tested_againstâ Microbiology_Antibiotics; Medicationâ switched_toâ Medication) Q6. What was the documented reasoning for selecting the targeted regimen based on those susceptibility re- sults? A6. (clinical reasoning from notes linking the susceptibility pattern to the targeted therapy choice; Medication â administered_forâ Diagnosis) Q7. By discharge, did the infection improve based on clinical or laboratory markers? A7. (Medication â improves â Diagnosis/Symptom/Finding; Lab_Test trends; Vital_Sign trends; Out- come.status) Template 2 â Culture-Negative vs Culture-Positive Infection Resolution Q1. During admissionâ¨chartdate or admission_idâŠ, was any infection clinically suspected even before micro- biologic confirmation? A1. (Diagnosis.assertion; clinical suspicion) Q2. Were any microbiologic tests performed to evaluate the infection, and what were their results? A2. (Specimenâ has_testâ Microbiology_Test; Microbiology_Test.result) Q3. If those tests did not confirm an organism, was there other clinical evidence supporting the infection? A3. (Lab_Test/Vital_Sign/Diagnostic_Imaging_Test/Physical_Examâ revealsâ Diagnosis/Finding) Q4. Based on the diagnostic evaluation, was the treatment continued, stopped, or modified? A4. (Medication.status; Medicationâ switched_toâ Medication; Medicationâ administered_forâ Diag- nosis) Q5. After the management decision, did the patientâs clinical status improve? A5. (Outcome.status; Symptom.status; Vital_Sign/Lab improvement) Q6. Was any follow-up or monitoring plan documented at discharge for the infection? A6. (Diagnosisâ has_instructionâ Instruction; Medicationâ has_instructionâ Instruction) Template 3 â Antibiotic Resistance-Driven Therapy Escalation Q1. During admissionâ¨chartdate or admission_idâŠ, was any organism identified from microbiologic testing? A1. (Microbiology_Testâ detectedâ Microbiology_Organism) Q2. For the infection caused by the organism, was any treatment initially started? A2. (Medicationâ administered_forâ Diagnosis; Medication.time) Q3. Were any susceptibility or resistance findings reported for the organism? A3.(Microbiology_Organism âtested_against âMicrobiology_Antibiotics;Microbiol- ogy_Antibiotics.sensitivity; Microbiology_Antibiotics.dilution) Q4. Was the treatment modified or escalated based on those susceptibility findings? 42 A4. (Medication â switched_to â Medication; Medication.status; clinical reasoning from notes linking resistance findings to the therapy escalation) Q5. After the therapy adjustment, did the infection markers or symptoms improve? A5. (Medicationâ improvesâ Diagnosis/Symptom/Finding; Outcome.status) Q6. What treatment regimen was continued at discharge for the infection? A6. (Medication.time = discharge) Template 4 â Recurrent Infection and Organism Persistence Across Admissions Q1. During the first admission, was any infection diagnosed? A1. (Diagnosis.assertion; Diagnosis.time = admission_1) Q2. For the infection, was any organism identified from microbiologic testing? A2. (Microbiology_Testâ detectedâ Microbiology_Organism) Q3. After treatment of the infection, did the same condition recur during a later admission? A3. (Diagnosis.status recurrence; Diagnosis.time comparison) Q4. During the later episode of the infection, was the same organism detected again or was a different organism identified? A4. (Microbiology_Testâ detectedâ Microbiology_Organism organism comparison) Q5. How did the treatment approach change compared with the earlier episode? A5. (Medication/Procedure â administered_for â Diagnosis; Medication â switched_to â Medication; clinical reasoning from notes linking prior treatment outcome and organism persistence to the revised approach) Q6. What was the clinical outcome after management of the recurrent infection? A6. (Outcome.status) Template 5 â Monitoring Microbiologic Clearance and Treatment Duration Decisions Q1. During admission â¨chartdate or admission_idâŠ, was any infection identified, and was any organism de- tected? A1. (Diagnosis.assertion; Microbiology_Testâ detectedâ Microbiology_Organism) Q2. For treatment of the infection, what therapy was started, and when? A2. (Medication.assertion; Medication.time; Medicationâ administered_forâ Diagnosis) Q3. After the therapy was initiated, were any follow-up microbiologic tests performed to evaluate treatment response? A3. (Microbiology_Test.time after Medication) Q4. Did results of those follow-up tests indicate persistence of infection or microbiologic clearance? A4. (Microbiology_Test.result; Diagnosis.status) Q5. Did those follow-up findings affect continuation or duration of therapy? A5. (Medication.duration; Medication.status) Q6. Was any discharge plan or outpatient monitoring recommended to ensure resolution of the infection? A6. (Diagnosisâ has_instructionâ Instruction; Medicationâ has_instructionâ Instruction) D.6 Clinical Assessment D.6.1 Single-note setting Template 1 â Clinical Assessment Used to Rule Out Competing Diagnoses Q1. Were any diagnoses initially suspected based on presenting clinical assessments? A1. (Diagnosis.status; Diagnosis.assertion; time = admission) Q2. What assessments were performed to evaluate those suspected conditions? A2. (Lab_Test; Diagnostic_Imaging_Test; Physical_Exam; Vital_Sign; time) Q3. Did any of those assessment findings rule out a suspected condition? A3. (Lab_Test / Diagnostic_Imaging_Test / Vital_Sign / Physical_Examâ revealsâ Diagnosis with negative assertion) Q4. After those exclusionary results, were any findings documented that confirmed the final diagnosis? A4. (Lab_Test / Diagnostic_Imaging_Test / Vital_Sign / Physical_Examâ revealsâ Diagnosis; clinical rea- soning from notes linking the positive findings to the confirmed diagnosis) Q5. Was any treatment decision made based on the clarified diagnosis? A5. (Procedure/Medicationâ administered_forâ Diagnosis) Q6. Were any assessment findings at discharge that supported clinical stability? A6.(Lab_Test.value; Vital_Sign.status; Physical_Exam.finding; Diagnostic_Imaging_Test.result; Out- come.status; time = discharge) Template 2 â Assessment-Guided Monitoring and Treatment Adjustment Q1. Were there any assessment findings that prompted initiation or adjustment of treatment during admission? A1. (Lab_Test / Vital_Sign / Physical_Exam / Diagnostic_Imaging_Testâ resulted_inâ Medication/Proce- dure) Q2. How frequently were those assessments repeated? 43 A2. (Lab_Test.frequency; Vital_Sign.frequency; Physical_Exam.frequency; time intervals during admission) Q3. Did repeat assessments show improvement after treatment? A3.(Medication/Procedure â resulted_in â Outcome/Finding; Lab_Test.value normalization; Vi- tal_Sign.status improvement) Q4. Were there any assessment findings that worsened despite treatment? A4. (Medication/Procedure â worsens â Diagnosis/Symptom/Finding OR Outcome/Finding â after â Medication/Procedure) Q5. Did those findings improve, and if not, was the treatment plan modified in response? A5. (Medication/Procedureâ switched_toâ Medication/Procedure; clinical reasoning from notes linking the worsening findings to the specific treatment modification) Q6. What were the clinical assessment values at discharge? A6. (Lab_Test.value; Vital_Sign.value/status; Physical_Exam.finding; Diagnostic_Imaging_Test.result; Out- come.status; time = discharge) Template 3 â Complication Detection Through Clinical Assessment Q1. Were there any new abnormal clinical findings during hospitalization? A1. (Lab_Test.abnormal_flag; Vital_Sign.status; Physical_Exam.finding; Diagnostic_Imaging_Test.result; time) Q2. Did any procedure or medication cause the new abnormal finding? A2. (Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding) Q3. Were any assessments used to confirm the complication? A3. (Lab_Test / Diagnostic_Imaging_Test / Vital_Sign / Physical_Examâ revealsâ Diagnosis/Finding) Q4. Was any treatment initiated to manage the complication? A4. (Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q5. Did follow-up assessments show resolution of the complication? A5. (Medication/Procedure â resulted_in â Outcome/Finding; Lab_Test.value trend; Vital_Sign improve- ment) Q6. What was the patientâs final assessment status at discharge? A6.(Outcome.status;Lab_Test.value;Vital_Sign.status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result) Template 4 â Discharge Assessment and Follow-Up Planning Based on Residual Abnormalities Q1. What were the patientâs clinical assessment findings on the day of discharge? A1. (Lab_Test.value; Lab_Test.abnormal_flag; Vital_Sign.status/value; Physical_Exam.finding; Diagnos- tic_Imaging_Test.result; time = discharge) Q2. Were any abnormalities still present at discharge? A2. (Lab_Test.abnormal_flag; Vital_Sign.status; Physical_Exam.finding; Diagnostic_Imaging_Test.result) Q3. What was the documented clinical reasoning for managing those persistent abnormalities as outpatient? A3. (clinical reasoning from notes explaining why persistent findings did not preclude discharge; Diagno- sis/Symptom/Findingâ has_instructionâ Instruction) Q4. Were any follow-up tests or imaging recommended based on those assessment findings? A4. (Diagnosis/Symptom/Findingâ resulted_inâ Lab_Test/Diagnostic_Imaging_Test) Q5. What was the overall clinical outcome based on the final assessments? A5. (Outcome.status; Outcome/Findingâ afterâ Procedure/Medication) Template 5 â Treatment Adjustment Guided by Reassessment Q1. What was the patientâs initial clinical status at admission based on exam findings and lab/imaging results? A1.(Physical_Exam.finding,Lab_Test.value/abnormal_flag,Vital_Sign.value/status,Diagnos- tic_Imaging_Test.result; time = admission) Q2. Was any treatment started initially, and what was it targeting? A2. (Medication/Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q3. Were there any assessment findings during the hospital course that indicated the initial treatment was insufficient or causing adverse effects? A3. (Medication/Procedureâ worsensâ Diagnosis/Symptom/Finding; Medication/Procedureâ causesâ Diagnosis/Symptom/Finding; Lab_Test/Vital_Sign trend worsening) Q4. Did the reassessment findings show adequate response, and if not, were any changes to treatment made? A4. (Medication/Procedure â switched_to â Medication/Procedure; Lab_Test/Vital_Sign/Physical_Exam â resulted_inâ Medication/Procedure; clinical reasoning linking reassessment to treatment change) Q5. Did the treatment modification lead to clinical improvement on reassessment? A5.(Medication/Procedure â resulted_in â Outcome/Finding; Lab_Test.value normalization; Vi- tal_Sign.value improvement) Q6. What were the assessment results at discharge after treatment was optimized? A6. (Lab_Test.value, Vital_Sign.value, Physical_Exam.finding at discharge; Outcome.status) 44 D.6.2 Multi-note setting Template 1 â Assessment Findings Triggering Treatment Initiation and Adjustment Q1. Were there any abnormal clinical assessments at admissionâ¨chartdate or admission_idâŠ? If so, what were they? A1.(Lab_Test.value/abnormal_flag;Vital_Sign.value/status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result) Q2. Did those findings point to any clinical condition? A2. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ revealsâ Diagnosis/Symptom/Find- ing) Q3. Was any treatment or procedure started to manage that condition? A3. (Medication/Procedureâ administered_forâ Diagnosis/Symptom/Finding) Q4. After the treatment was initiated, were the assessments repeated, and if so, what trends were observed? A4. (Lab_Test.value/abnormal_flag trends; Vital_Sign.value/status trends; Physical_Exam.finding trends; Diagnostic_Imaging_Test.result trends) Q5. Did those findings improve, and if not, was the treatment modified in response? A5. (Medication/Procedure â switched_to â Medication/Procedure; clinical reasoning from notes linking inadequate response to the treatment change) Q6. What were the assessment results by discharge after that management course? A6.(Lab_Test.value/abnormal_flag;Vital_Sign.value/status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result at discharge; Outcome.status) Template 2 â Assessment Findings Used to Rule Out Competing Diagnoses Q1. Were any conditions suspected at admissionâ¨chartdate or admission_id⊠based on the presenting findings? A1. (Diagnosis.assertion; Diagnosis.status) Q2. What clinical assessments were performed to evaluate those possibilities? A2. (Lab_Test; Diagnostic_Imaging_Test; Physical_Exam; Vital_Sign) Q3. Did any of those assessment results make some of those possibilities less likely or rule them out? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign â reveals â Diagnosis with negative assertion) Q4. After evaluating those results, what condition was ultimately supported? A4. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ revealsâ Diagnosis) Q5. What was the documented reasoning that connected those findings to the treatment decision? A5. (Procedure/Medication â administered_for â Diagnosis; clinical reasoning from notes linking the confirmatory findings to the treatment selection) Q6. By discharge, were there any assessment findings demonstrating improvement of the condition? A6.(Lab_Test.value/abnormal_flag;Vital_Sign.value/status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result; Outcome.status) Template 3 â Assessment-Guided Treatment Adjustment and Monitoring During Admission Q1. Were there any assessment findings during admission â¨chartdate or admission_id⊠that led to treatment initiation or change? A1. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ resulted_inâ Medication/Procedure) Q2. What were the values or findings in those assessments? A2. (Lab_Test.value; Lab_Test.abnormal_flag; Vital_Sign.value/status; Diagnostic_Imaging_Test.result; Phys- ical_Exam.finding) Q3. After that intervention, were the assessments repeated to monitor response, and if so, how frequently? A3. (Lab_Test.frequency; Vital_Sign.frequency; Physical_Exam.frequency) Q4. Did the repeat assessments show improvement after that treatment was started? A4.(Medication/Procedure â resulted_in â Outcome/Finding; Lab_Test.value normalization; Vi- tal_Sign.value improvement) Q5. Were there any assessments that worsened despite that treatment? A5.(Outcome.status; Outcome/Finding â after â Medication/Procedure; Lab_Test.value trend; Vi- tal_Sign.value trend) Q6. Did those findings improve adequately, and if not, was the treatment modified? A6. (Medication/Procedureâ switched_toâ Medication/Procedure; clinical reasoning from notes linking the specific worsening findings to the treatment change) Q7. After treatment optimization, what were the assessment values at discharge? A7. (Lab_Test.value; Vital_Sign.value; Physical_Exam.finding; Diagnostic_Imaging_Test.result at discharge; Outcome.status) Template 4 â Clinical Assessment-Driven Diagnosis Confirmation and Treatment Initiation Q1.What clinical assessments were performed to evaluate â¨Diagnosis/Symptom/Finding⊠at admission â¨chartdate or admission_idâŠ? A1. (Lab_Test; Diagnostic_Imaging_Test; Physical_Exam; Vital_Sign; time = admission) Q2. Were any abnormalities identified in those assessments? 45 A2. (Lab_Test.abnormal_flag; Lab_Test.value; Vital_Sign.status; Vital_Sign.value; Physical_Exam.finding; Diagnostic_Imaging_Test.result) Q3. Did those findings point to any condition? A3. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ revealsâ Diagnosis) Q4. Did those findings lead clinicians to start or change treatment? A4. (Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Signâ resulted_inâ Procedure/Medication) Q5. By discharge, how had those assessment findings changed? A5.(Lab_Test.value at discharge; Vital_Sign.value at discharge; Physical_Exam.finding; Diagnos- tic_Imaging_Test.result; Outcome.status) Template 5 â Cross-Admission Assessment Changes and Evaluation of New or Persistent Abnormalities Q1. Compared with the previous discharge, were there any new clinical findings or vital sign abnormalities at admissionâ¨chartdate or admission_idâŠ? A1. (Vital_Sign.status/value; Physical_Exam.finding; Diagnosis.status comparison over time; time = admis- sion) Q2. Were any diagnostic tests ordered to evaluate those abnormalities? A2.(Symptom/Diagnosis âresulted_in âLab_Test/Diagnostic_Imaging_Test; Lab_Test/Diagnostic_Imaging_Test.time) Q3. Were any results from those tests used to rule out suspected conditions? A3. (Lab_Test.value/assertion; Diagnostic_Imaging_Test.result/assertion; Lab_Test/Diagnostic_Imaging_Test â revealsâ Diagnosis with negative assertion) Q4. Did the documented clinical reasoning interpret those cross-admission findings as progression, a new problem, or a recurrence? A4.(clinical reasoning from notes linking cross-admission finding trends to the clinical interpretation; Lab_Test/Vital_Sign/Physical_Examâ resulted_inâ Medication/Procedure) Q5. Did any of those assessment trends lead clinicians to change treatment? A5. (Medicationâ switched_toâ Medication; Lab_Test.value trends; Vital_Sign.status trends) Q6. By discharge, were any abnormalities from those findings still present that required follow-up instructions? A6. (Lab_Test.abnormal_flag; Vital_Sign.status; Physical_Exam.finding; Diagnosis/Symptom/Finding â has_instructionâ Instruction; time = discharge) D.7 Clinical Outcome D.7.1 Single-note setting Template 1 â In-Hospital Complications, Rescue Interventions, and Final Outcome Q1. Did the patient experience any complications or clinical deterioration during admission? A1. (Event.assertion/time; Diagnosis/Symptom/Finding.status change; Outcome.status change) Q2. Were any clinical findings or test results documented that signaled the complication or deterioration? A2.(Lab_Test.abnormal_flag/value;Vital_Sign.status/value;Physical_Exam.finding;Diagnos- tic_Imaging_Test.resultâ revealsâ Diagnosis/Finding) Q3. Was any underlying cause identified for the complication? A3.(Procedure/Medication â causes â Diagnosis/Symptom/Finding; Diagnosis/Symptom/Finding â causesâ Event) Q4. Were any treatments or procedures initiated in response to the identified cause? A4. (Procedure/Medication â administered_for â Diagnosis/Symptom/Finding; Lab_Test/Vital_Sign â resulted_inâ Procedure/Medication) Q5. Did the interventions resolve or improve the complication before discharge? A5.(Procedure/Medication â improves â Diagnosis/Symptom/Finding; Procedure/Medication â re- sulted_inâ Outcome/Finding; Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding) Q6. What was the final outcome status at discharge, and were there any residual deficits? A6. (Outcome.status; Diagnosis/Symptom/Finding.status; Activity.status at discharge) Template 2 â Treatment Goal Attainment Q1. What was the primary clinical problem driving admission? A1. (Diagnosis.is_main; Symptom.is_main; Finding.is_main) Q2. What interventions were used to treat that problem? A2. (Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q3. Was the primary problem resolved, improved, or still active at discharge? A3. (Outcome.status; Diagnosis/Symptom/Finding.status at discharge; Outcomeâ afterâ Procedure/Medi- cation) Q4. Was there any evidence supporting whether the treatment goal was achieved? A4.(Lab_Test.value/abnormal_flag;Vital_Sign.value;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result) Q5. If the goal was not fully achieved, were any barriers, side effects, or complications documented as 46 contributing? A5. (Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding; Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding; Diagnosis/Symptom/Findingâ causesâ Event) Q6. How was the follow-up plan designed to address those barriers and the remaining problem? A6. (Any category â has_instruction â Instruction; continued/planned Procedure/Medication; clinical reasoning from notes linking residual issues to the follow-up strategy) Template 3 â Functional Outcome, Disposition, and Post-Discharge Support Needs Q1. What was the patientâs functional status at the time of discharge? A1. (Activity.status at discharge; Physical_Exam.finding relevant to mobility/function) Q2. Were there any clinical problems or outcomes that limited the patientâs function at discharge? A2. (Diagnosis/Symptom/Findingâ resulted_inâ Outcome/Finding; Outcomeâ afterâ Diagnosis/Symp- tom/Finding) Q3. Did the patient require any assistive devices, home services, or rehabilitation after discharge? A3. (Medical_Device.status; Activity.status; Outcome.assertion) Q4. What discharge disposition was selected based on those functional limitations and support needs? A4. (Outcome.status/time; Activity.status; Diagnosis/Symptom/Finding.status driving disposition decision) Q5. Were any follow-up appointments, therapy, or instructions arranged to support functional recovery? A5. (Any categoryâ has_instructionâ Instruction; Procedure/Medication planned/continued post-discharge) Template 4 â Medication Adjustments During Admission and Their Effect on Outcome Q1. Were any medications started, adjusted, or discontinued during admission, and why? A1. (Medication.assertion/time; Procedure/Medication â administered_forâ Diagnosis/Symptom/Finding; Procedure/Medicationâ switched_toâ Procedure/Medication) Q2. Did any medication cause adverse effects or complications during the admission? A2. (Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding; Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding) Q3. How were those adverse effects managed, and were medications adjusted as a result? A3. (Procedure/Medicationâ switched_toâ Procedure/Medication; Procedure/Medicationâ causesâ Pro- cedure/Medication; Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q4. What was the overall effect of the final medication regimen on the patientâs primary condition by dis- charge? A4.(Procedure/Medication â improves â Diagnosis/Symptom/Finding; Procedure/Medication â re- sulted_inâ Outcome/Finding; Outcomeâ afterâ Procedure/Medication) Q5. What medications were prescribed at discharge, and were any instructions given regarding their use? A5. (Medication.strength/dosage/route/instruction at discharge; Any category â has_instruction â Instruc- tion) Template 5 â Goal Attainment and Residual Risk at Discharge Q1. What was the intended treatment goal during admission? A1. (Diagnosis/Symptom/Finding.is_main) Q2. Was the treatment goal achieved by discharge? A2. (Outcome.status at discharge, Diagnosis/Symptom/Finding.status at discharge, Outcome â after â Diagnosis/Symptom/Finding) Q3. Was there any objective data confirming goal attainment? A3.(Lab_Test.value/abnormal_flag,Vital_Sign.status/value,Physical_Exam.finding,Diagnos- tic_Imaging_Test.result) Q4. If the goal was not fully achieved, were any factors documented as contributing to the incomplete response? A4.(Procedure/Medication â causes â Diagnosis/Symptom/Finding, Diagnosis/Symptom/Finding â causesâ Event, Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding) Q5. What was the documented clinical reasoning for determining which residual conditions posed the highest risk? A5. (clinical reasoning from notes linking residual conditions to specific risk factors; Persistent Diagno- sis/Symptom/Finding.status; Outcome.status = partial/stable with deficits) Q6. Was any follow-up plan established to prevent worsening? A6. (Any categoryâ has_instructionâ Instruction, Planned/continued Medication/Procedure) D.7.2 Multi-note setting Template 1 â Discharge Readiness: Clinical Stability and Objective Evidence Q1. What was the patientâs overall clinical status at discharge during admissionâ¨chartdate or admission_idâŠ? A1. (Outcome.status; Outcome.time = discharge) Q2. Were there any conditions, findings, or symptoms that were improved, resolved, or still active at discharge? A2. (Diagnosis.status; Symptom.status; Finding.status; time = discharge) 47 Q3. Were there any lab tests, vital signs, physical exams, or diagnostic imaging tests that supported that stability or improvement at discharge? A3.(Lab_Test.value/abnormal_flag;Vital_Sign.value/status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result at discharge) Q4. Among those clinical findings, were any abnormalities still present that could pose short-term risk after discharge? A4. (Finding.status; Lab_Test.abnormal_flag; Vital_Sign.status; persistent Diagnosis/Symptom) Q5. What was the documented reasoning for determining the patient was safe for discharge despite those residual abnormalities? A5. (clinical reasoning from notes linking the residual findings to the discharge readiness assessment; risk mitigation strategies documented) Q6. Were any monitoring or contingency instructions provided to address those remaining risks? A6. (Any categoryâ has_instructionâ Instruction) Template 2 â In-Hospital Deterioration, Rescue Interventions, and Outcome Shift Q1. Did the patient clinically deteriorate during admissionâ¨chartdate or admission_idâŠ? A1. (Event.assertion/time; Outcome.status change; Diagnosis/Symptom/Finding.status change) Q2. Were there any clinical findings or assessment results that indicated that deterioration? A2.(Lab_Test.abnormal_flag/value;Vital_Sign.status/value;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result) Q3. Were any clinical problems identified as the cause of that deterioration? A3. (Diagnosis/Symptom/Findingâ causesâ Event; Eventâ causesâ Diagnosis/Symptom/Finding) Q4. Were any urgent treatments or procedures performed in response? A4. (Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q5. Did those interventions stabilize or reverse the deterioration before discharge? A5.(Procedure/Medication â resulted_in â Outcome/Finding; Procedure/Medication â improves â Diagnosis/Symptom/Finding; Procedure/Medicationâ worsensâ Diagnosis/Symptom/Finding) Q6. What was the patientâs final outcome status at discharge, and were any deficits still present? A6. (Outcome.status; Activity.status; Diagnosis/Symptom/Finding.status) Template 3 â Functional Outcome and Care Needs: Disposition and Supports Q1. What was the patientâs functional status at discharge for admissionâ¨chartdate or admission_idâŠ? A1. (Activity.status; time = discharge; Physical_Exam.finding relevant to function) Q2. Were there any clinical problems that contributed to that functional status at discharge? A2. (Diagnosis/Symptom/Findingâ resulted_inâ Outcome/Finding; Outcomeâ afterâ Diagnosis/Symp- tom/Finding) Q3. Did the patient require any devices, services, or caregiver support after discharge? A3. (Medical_Device.status; Activity.status; Outcome.assertion) Q4. What discharge disposition was selected based on that support plan? A4. (Outcome.status/time; Activity.status) Q5. Was any follow-up or rehabilitation plan documented to improve that functional status? A5. (Any categoryâ has_instructionâ Instruction; planned/continued Procedure or Medication) Template 4 â Longitudinal Recovery Trajectory Across Admissions Q1. Across admissions, how did the patientâs outcome status change from each discharge to the next admis- sion? A1. (Outcome.status over time; time.category = discharge/admission) Q2. During that longitudinal course, were there clinical problems that repeatedly worsened after discharge or failed to resolve? A2. (Diagnosis/Symptom/Finding.status over time) Q3. Were any factors associated with worsening of those conditions between encounters? A3. (Medication/Procedureâ worsensâ Diagnosis/Symptom/Finding; Eventâ causesâ Diagnosis/Symp- tom/Finding; Diagnosis/Symptom/Findingâ causesâ Event) Q4. How did the treatment approach evolve to address those factors and the pattern of deterioration? A4.(Medication/Procedure â improves â Diagnosis/Symptom/Finding; Medication/Procedure â re- sulted_in â Outcome/Finding; clinical reasoning from notes linking cross-admission worsening factors to treatment modifications) Q5. By the most recent discharge, were any functional limitations related to those conditions still present, and what plan addressed them? A5. (Activity.status; Diagnosis/Symptom/Finding.status at discharge; Instruction) Template 5 â Treatment Goal Attainment vs Failure: Why the Outcome Was Achieved (or Not) Q1. What was the primary treatment goal during admissionâ¨chartdate or admission_idâŠ? A1. (Diagnosis/Symptom/Finding.is_main; Diagnosis/Symptom/Finding.assertion) Q2. What treatments were administered to manage the condition? A2. (Procedure/Medicationâ administered_forâ Diagnosis/Symptom/Finding) 48 Q3. By discharge, did the condition meet the intended treatment goal? A3. (Outcome.status at discharge; Diagnosis/Symptom/Finding.status at discharge; Outcome â after â Procedure/Medication) Q4. Were there any lab test, vital sign, physical exam, or diagnostic imaging results that supported the assessment of that treatment outcome? A4.(Lab_Test.value/abnormal_flag;Vital_Sign.value/status;Physical_Exam.finding;Diagnos- tic_Imaging_Test.result) Q5. Did the condition fully improve, and if not, were any complications or barriers documented as contributing to the incomplete response? A5. (Procedure/Medicationâ causesâ Diagnosis/Symptom/Finding/Event; Diagnosis/Symptom/Findingâ causesâ Event) Q6. How was the follow-up plan designed to address those barriers and the remaining problem? A6. (Any category â has_instruction â Instruction; continued/planned Procedure or Medication; clinical reasoning from notes linking residual issues to the follow-up strategy) D.8 Discharge Plan D.8.1 Single-note setting Template 1 â Discharge Medication Reconciliation Q1. What medications was the patient prescribed at discharge? A1. (Medication.assertion, time = discharge; Medication.strength, dosage, route, frequency) Q2. Which of those medications were newly started during the hospitalization versus continued from admis- sion? A2. (Medication.status comparison: admission vs. discharge; Medication.time) Q3. Were any admission medications stopped or held at discharge, and if so, why? A3. (Medication.status = stopped/held; Medication â administered_for â Diagnosis/Symptom/Finding or adverse effect) Q4. For the new or changed medications, what condition was each prescribed for? A4. (Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q5. Were any dose adjustments made during the hospitalization before arriving at the final discharge dose, and what drove those changes? A5. (Medication.dosage change; Medication â causes â Symptom/Finding; Medication â worsens â Symptom/Finding) Template 2 â Discharge Disposition and Care Setting Rationale Q1. Where was the patient discharged to, and what was the stated reason for that disposition? A1. (Outcome.status/assertion = discharge disposition; Activity.status; Diagnosis/Finding â resulted_in â Outcome) Q2. What was the patientâs clinical condition at the time of discharge that supported the disposition decision? A2. (Vital_Sign.value, time = discharge; Physical_Exam.finding, time = discharge; Outcome.assertion) Q3. What was the documented reasoning for determining the appropriate level of post-discharge care based on those findings? A3. (clinical reasoning from notes linking the discharge clinical status to the disposition decision; Activ- ity.status; functional limitations) Q4. Were any follow-up services or care arrangements put in place at the receiving facility or at home? A4. (Instruction.instruction_text; Event.status = scheduled; Medical_Device/Activity supports) Q5. Were there any outstanding clinical issues at discharge that the receiving care setting was expected to manage? A5. (Diagnosis/Finding.status = ongoing; Instruction.instruction_text directed to next provider) Template 3 â Follow-Up Plan, Outpatient Connections, and Contingency Instructions Q1. Were any follow-up appointments or outpatient referrals arranged at discharge? A1. (Event.status = scheduled; Instruction.instruction_text; Instruction.time) Q2. What clinical problems was each follow-up appointment intended to address? A2. (Diagnosis/Symptom/Finding/Outcomeâ has_instructionâ Instruction) Q3. Were any instructions given regarding who to contact and under what circumstances before the scheduled follow-up? A3. (Instruction.instruction_text; Symptom/Findingâ has_instructionâ Instruction with urgency framing) Q4. Were there any explicit instructions about substances, behaviors, or lifestyle factors the patient was told to avoid? A4. (Instruction.instruction_text; Activityâ has_instructionâ Instruction; Medication/substance warnings) Q5. What should the patient do if they are unable to reach their outpatient provider or feel unsafe? A5. (Instruction.instruction_text = emergency contingency; Eventâ has_instructionâ Instruction) 49 Template 4 â Symptom Monitoring, Red Flags, and Emergency Planning Q1. Were any symptoms or warning signs that the patient should monitor after discharge documented? A1. (Symptom/Findingâ has_instructionâ Instruction) Q2. Which conditions or complications were those warning signs intended to detect or prevent? A2. (Diagnosis/Symptom/Findingâ has_instructionâ Instruction) Q3. What was the documented reasoning for emphasizing those specific warning signs based on the patientâs hospital course? A3. (clinical reasoning from notes linking the hospital course â complications, persistent findings, or high-risk conditions â to the specific warning signs selected) Q4. What actions was the patient instructed to take if those symptoms occur? A4. (Instruction.instruction_text; may include Event) Q5. Were any objective thresholds provided? A5. (Instruction.instruction_text; Vital_Sign.value thresholds if present) Template 5 â Follow-Up Care, Appointments, and Diagnostic Monitoring Q1. Were any follow-up appointments or referrals arranged at discharge? A1. (Instruction.instruction_text; Event.status = scheduled) Q2. What medical or psychiatric issues were those follow-ups meant to address? A2. (Diagnosis/Symptom/Finding/Outcomeâ has_instructionâ Instruction) Q3. Were any follow-up tests such as labs, imaging, or microbiology recommended? A3. (Lab_Test/Diagnostic_Imaging_Test/Microbiology_Test planned; Instruction) Q4. What was the recommended timing or urgency of follow-up? A4. (Instruction.time; Instruction.instruction_text) Q5. Were any contingency plans described if follow-up could not be completed? A5. (Instruction.instruction_text; case management/social work if applicable) D.8.2 Multi-note setting Template 1 â Medication Reconciliation and Post-Discharge Medication Plan Q1.What medications was the patient instructed to take after discharge at admission â¨chartdate or admission_idâŠ? A1. (Medication.assertion; Medication.time = discharge) Q2. Compared with the medications in admissionâ¨chartdate or admission_idâŠ, were any of those medications newly started, stopped, held, or dose-adjusted? A2. (Medication.status/time comparison; Medication.dosage/route/frequency changes) Q3. For those medications, what clinical conditions or symptoms were they intended to treat? A3. (Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q4. What was the documented reasoning that explained the evolution of the medication plan from one dis- charge to the next? A4. (clinical reasoning from notes linking cross-admission clinical changes to the medication adjustments; Medicationâ administered_forâ Diagnosis/Symptom/Finding) Q5. Were any monitoring instructions provided for those treatments after discharge? A5. (Medicationâ has_instructionâ Instruction) Q6. In the instructions associated with those treatments, were there any warnings about adherence or risks if the medications were stopped or taken incorrectly? A6. (Instruction.instruction_text; Event risk framing if present) Template 2 â Red Flags, Return Precautions, and Symptom Monitoring Q1. At discharge from admissionâ¨chartdate or admission_idâŠ, were any warning symptoms or clinical changes that the patient should watch for at home documented? A1. (Symptom/Findingâ has_instructionâ Instruction) Q2. Which underlying condition or complication were those warning signs intended to detect early or prevent? A2. (Diagnosis/Symptom/Findingâ has_instructionâ Instruction) Q3. What was the documented reasoning for emphasizing those specific warning signs based on the patientâs hospital course? A3. (clinical reasoning from notes linking the hospital course â complications, persistent findings, or high-risk conditions â to the specific warning signs selected) Q4. If those warning signs occurred, what actions was the patient instructed to take? A4. (Instruction.instruction_text) Q5. Within the guidance related to those warning signs, were any objective thresholds provided that would trigger action? A5. (Instruction.instruction_text with numeric thresholds; Vital_Sign.value targets if present) Q6. Beyond watching for those warning signs, was any follow-up monitoring plan arranged to track recovery? A6. (Instruction.instruction_text; Event.status = scheduled if represented) 50 Template 3 â Follow-up Appointments and Care Coordination Q1. At discharge from admissionâ¨chartdate or admission_idâŠ, were any follow-up appointments, referrals, or services arranged for the patient? A1. (Instruction.instruction_text; Event.status = scheduled) Q2. What clinical problems or conditions were those follow-ups intended to address? A2. (Diagnosis/Symptom/Finding/Outcomeâ has_instructionâ Instruction) Q3. As part of that follow-up plan, were any diagnostic tests scheduled? A3. (Lab_Test / Diagnostic_Imaging_Test / Microbiology_Test planned; Instruction) Q4. Were any instructions given about the timing or urgency of those follow-ups, and who the patient should contact? A4. (Instruction.time; Instruction.instruction_text) Q5. If the patient was unable to complete those follow-ups, were any contingency instructions documented? A5. (Instruction.instruction_text; social work/case management references if present) Template 4 â Disposition, Functional Plan, and Post-Discharge Supports Q1. What was the patientâs disposition at discharge at admissionâ¨chartdate or admission_idâŠ, and why? A1. (Outcome.assertion/status; Activity.status) Q2. What was the patientâs functional or activity level at the time of discharge? A2. (Activity.status, time = discharge) Q3. Did the patient require any medical devices or supportive services after discharge? A3. (Medical_Device.status; Activity.status; Instruction) Q4. Were any activity restrictions or rehabilitation instructions provided in relation to that support plan? A4. (Instruction.instruction_text; Activity-related instructions) Q5. Were any criteria given for advancing activity or seeking reassessment? A5. (Instruction.instruction_text; thresholds or time-based progression) Template 5 â Cross-Admission Discharge Plan Adherence and Readmission Link Q1. What were the discharge instructions given after admissionâ¨chartdate or admission_idâŠ? A1. (Any categoryâ has_instructionâ Instruction; time = prior discharge) Q2. At the next hospitalization, was there evidence that those instructions were followed or not followed? A2. (Event.status = missed follow-up / non-adherence; Medication.status = not taking) Q3. Among those instructions, were any aspects of non-adherence clinically significant? A3. (Medication.status; Event.status; Symptom.status) Q4. What was the documented clinical reasoning that connected those non-adherence failures to the current presentation at readmission? A4. (clinical reasoning from notes linking specific non-adherence to clinical deterioration; Medication â worsensâ Diagnosis/Symptom/Finding; Eventâ causesâ Diagnosis/Symptom/Finding) Q5. At the most recent discharge, was the discharge plan modified to reduce the risk of recurrence or improve adherence? A5. (Updated Instruction; Medication adjustments; follow-up changes across admissions) 51 E Per-Step Details for LLM-Based Initial Data Generation This appendix expands on the four-step generation pipeline summarized in Section 3.3. We use Gemini-2.5-Pro at temperature 1.0 for Steps 1â2 (to encourage diversity) and 0.7 for Steps 3â4 (to maintain answer-choice quality). The full pipeline cost approximately $1,000 for 1,000 samples (âź$1 per sample). The complete prompt for each step is given in Appendix I, and example per-step outputs in Appendix F. Step 1: Multi-turn question generation. The input consists of a single patientâs full sequence of discharge summaries, the structuring schema (Appendix C), the target question category together with the five expert-curated QA templates for that category (single-note templates when the patient has one note and multi-note templates otherwise), and one randomly sampled template from each of the other seven categories. Gemini-2.5-Pro generates the multi-turn questions together with the sup- porting evidence and source location (note number, section header, and text span) for answering each question. Generating these jointly with the questions ensures the questions are grounded in the notes, and they are reused in later steps to produce the correct content-answer and evidence-grounding choices. The prompt enforces, in priority order, the clinical-naturalness, multi-turn-dependency, and multi-note-coverage constraints described in Section 3.3. Step 2: Incorrect-answer strategy brainstorming. The input is the patientâs discharge sum- maries, the structuring schema, and the complete Step 1 output (multi-turn questions with evidence and source location). For each question, the model produces one distractor direction for each of six types, derived from two grounding sources: incorrect interpretations of the notes (note-grounded) and content not present in the notes but clinically plausible given the patientâs context (clinical- knowledge-grounded). To support cross-turn error chaining, the prompt encourages later-turn dis- tractor directions to build on earlier ones (e.g., a later distractor involving a medication prescribed for an earlier distractorâs incorrect diagnosis). The six distractor types are: 1. Note-grounded entity/attribute substitution (note-grounded): replace the correct entity or attribute with a related but incorrect one drawn from elsewhere in the notes. 2. Plausible alternative interpretation (clinical-knowledge-grounded): introduce a compet- ing interpretation of the patientâs documented presentation, such as an alternative diagnosis, etiology, or drug-class indication, that is not documented but is medically plausible given the patientâs context. 3. Plausible alternative action (clinical-knowledge-grounded): introduce a competing ac- tion, such as a different intervention, medication within the same class, test, or follow-up plan, for the same clinical indication that is medically plausible but not documented. 4. Clinical relationship misattribution (note- or clinical-knowledge-grounded): misattribute relationships among clinical entities, such as by reversing causality, altering temporal order, or confusing entities across admissions. 5. Note-grounded answer with one clinically incorrect detail (note- and clinical- knowledge-grounded): use content that is mostly consistent with the notes but includes one clinically plausible yet incorrect detail, such as an incorrect dosage or severity grade. 6. Cross-admission temporal displacement (note-grounded): attribute a correct clinical fact to the wrong admission. This type applies only to patients with multiple admissions. Step 3: Content-question answer choice generation. The model is given the discharge sum- maries, the structuring schema, the Step 1 questions (with evidence and source location), and the Step 2 distractor directions, and generates one correct and four incorrect content choices per ques- tion. The correct choice is grounded in the Step 1 evidence; the incorrect choices follow the Step 2 directions, subject to two additional constraints. Shared-entity choices (Section 3.3): some in- correct choices reference the same key entities as the correct answer but differ in the relationship, attribute, or interpretation applied to them, preventing an entity-presence lookup strategy. Error- consistent choices: letting Q cont n denote the content question at turn n, for n ⼠2 the incorrect choices for Q cont n are encouraged to reflect clinically plausible downstream answers that would fol- low if the model had committed to a specific wrong answer for Q cont nâ1 , realizing the cross-turn error chains brainstormed in Step 2. 52 Step 4: Evidence-grounding answer choice generation. The model is given the discharge sum- maries, the Step 1 questions (with evidence and source location), and the Step 3 correct answer text, and generates one correct and four incorrect evidence-grounding choices, following the minimal- yet-sufficient correct-choice constraint and the four incorrect-choice error patterns defined in Sec- tion 3.3. As a concrete example of the semantic-trap principle, if the answer concerns an âinterven- tionâ, a distractor might include the Major Surgical or Invasive Procedure header even though that sectionâs content does not actually describe the intervention. This forces the question to be solved by reading the section content rather than by surface header-name matching. 53 F Per-Step Output Examples for LLM-Based Initial Data Generation This appendix shows an example LLM output for each of the four generation steps, all drawn from the same Microbiology-category multi-turn QA sample for a single patient with two discharge sum- maries. F.1 Step 1: Multi-turn Question Generation An example Step 1 output for a single patient with two discharge summaries is shown below (Microbiology-category multi-turn QA sample).To avoid disclosing patient-identifying dates and note content, chartdates are replaced with placeholders ([chartdate1] for Note #1, [chart- date2] for Note #2), and the exact text span in each [Source Location] block is replaced with raw_text_snippet. In the EHRNote-ChatQA dataset and in the LLM-generated Step 1 output, these fields contain the actual chartdates and supporting text spans; they are shown here only as placeholders to avoid disclosing patient-identifying information. For each multi-turn content question Q n it generates, Gemini-2.5-Pro is prompted to jointly pro- duce three components in a structured block: (1) [Entities, Attributes, Relationships], the âanswer evidenceâ for the questionâthe schema-defined entities, attributes, and relationships extracted from the discharge summaries that are necessary to answer it; (2) [Source Location], the note number(s), section header(s), and supporting raw text span(s) in the discharge summaries required to answer it; and (3) [Answer], the natural-language answer. The [Answer] is not used in Steps 2â4; it is produced only so that later turns within Step 1 can build on earlier an- swers when generating subsequent questions. The answer evidence ([Entities, Attributes, Relationships]) and [Source Location] are used in Steps 2â4 to construct the content- question answer choices and the evidence-grounding answer choices. Q1: During the admission for the right radical nephrectomy on chart date [chartdate1], were any infections noted or were any preventative measures taken? A1: [Entities, Attributes, Relationships] Diagnosis - name: infection, assertion: negative Procedure - name: laparoscopic right radical nephrectomy, is_main: true, time: during admission on [chartdate1] Medication - name: antibiotic prophylaxis, status: received, time: perioperative Medication -> administered_for -> Procedure [Source Location] Note #1Brief Hospital Course raw_text_snippet [Answer] No infection was diagnosed during this admission. The patient received perioperative antibiotic prophylaxis for the laparoscopic right radical nephrectomy. Q2: Was the patient monitored for any signs of infection after that intervention? A2: [Entities, Attributes, Relationships] Physical_Exam - name: examination of incision, time: on discharge Finding - name: clean, dry, and intact, finding_site: incision Finding - name: no evidence of hematoma collection or infection, finding_site: incision, assertion: negative Physical_Exam -> reveals -> Finding [Source Location] Note #1Brief Hospital Course raw_text_snippet [Answer] Yes, on the day of discharge, the surgical incision was examined and found to be clean, dry, and intact, with no evidence of hematoma collection or infection. Q3: Were there any laboratory findings during that admission that could suggest an inflammatory process? 54 A3: [Entities, Attributes, Relationships] Lab_Test - name: WBC, value: 11.3*, abnormal_flag: abnormal, time: ___ 06:58AM [Source Location] Note #1Pertinent Results raw_text_snippet [Answer] Yes, a complete blood count during the admission showed a white blood cell count of 11.3, which is slightly elevated. Q4: What was the clinical assessment of that lab result? A4: [Entities, Attributes, Relationships] Outcome - name: hospital course, status: unremarkable Outcome - name: condition at discharge, status: stable Lab_Test - name: WBC, value: 11.3* [Source Location] Note #1Brief Hospital Course raw_text_snippet [Answer] The hospital course was documented as unremarkable, and the patient was discharged in stable condition with a clean incision. This suggests the slightly elevated white blood cell count was considered a normal post-operative inflammatory response rather than a sign of a clinically significant infection. Q5: Was any treatment for an infection continued at discharge? A5: [Entities, Attributes, Relationships] Medication - name: antibiotics, assertion: negative, time: at discharge [Source Location] Note #1Discharge Medications raw_text_snippet [Answer] No, no antibiotics were included in the discharge medication list. The patient was prescribed pain medications and a stool softener. Q6: During the subsequent admission on chart date [chartdate2], was there any evidence of an infectious process? A6: [Entities, Attributes, Relationships] Diagnosis - name: infection, assertion: negative, time: admission on [chartdate2] Diagnosis - name: pulmonary embolism, is_main: true, time: admission on [chartdate2] Lab_Test - name: WBC, value: 9.7, abnormal_flag: normal, time: ___ 08:37AM [Source Location] Note #2Brief Hospital Course raw_text_snippet Note #2Pertinent Results raw_text_snippet [Answer] No, there was no evidence of an infectious process during the admission on chart date [chartdate2]. The primary diagnosis was a pulmonary embolism, and the patientâs white blood cell count on admission was within the normal range at 9.7. 55 F.2 Step 2: Distractor-Direction Brainstorming The Step 2 output for two consecutive questions (Q1 and Q2) is shown below, from the Microbiology-category multi-turn QA sample for the same patient used in Appendix F.1. In Step 2, Gemini-2.5-Pro generates distractor directions for the full sequence of Step 1 questions in a sin- gle pass; for brevity, we show the output for two consecutive questions only (Q1 and Q2). For each content question Q n , the model produces one distractor direction for each of the six distrac- tor types defined in Appendix E. Each direction is annotated with Explanation, Reasoning, and a Minimal-pair label indicating whether the distractor differs from the correct answer in ex- actly one detail. In the output, Source A denotes note-grounded content and Source B denotes clinical-knowledge-grounded content, defined in Section 3.3. As in Appendix F.1, the chartdate in the question text is shown as a placeholder to avoid disclosing patient-identifying dates. The first content question (Q1) contains the six distractor directions. For each follow-up content question Q n (n⼠2), the model performs two additional analyses before producing the six distractor directions. In the Referential Ambiguity Analysis, it identifies the referential expression in Q n , its correct referent, and plausible alternative referents, then constructs a distractor direction for each alternative. In the Cross-Turn Error Path Analysis, it identifies clinically plausible distractor directions to the previous content question Q nâ1 (drawn from Q nâ1 âs Step 2 distractor directions) and designs distractor directions for Q n that would be natural consequences. These outputs are used in Step 3 to construct the four incorrect content-answer choices for each question. Q1: During the admission for the right radical nephrectomy on chart date [chartdate1], were any infections noted or were any preventative measures taken? - Effective Distractor Type 1: Note-grounded entity substitution with clinical reasoning requirement (Source A) Explanation: State that a urinary tract infection was diagnosed post-operatively. Reasoning: The note mentions the patient had a urethral Foley catheter placed, which is a major risk factor for a UTI. A model might plausibly infer a UTI occurred, even though the note states the hospital course was "unremarkable" and he was "voiding without difficulty" after its removal. This requires clinical reasoning to know that while plausible, itâs not documented and contradicted by the overall summary. Minimal-pair label: [Minimal-pair: differs only in the presence of a UTI] - Effective Distractor Type 2: Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B) Explanation: State that because the procedure was laparoscopic and considered minimally invasive, antibiotic prophylaxis was deemed unnecessary and was not administered. Reasoning: This presents a plausible (though generally incorrect for this type of surgery) clinical judgment call. A model without specific knowledge of surgical prophylaxis guidelines might find it reasonable that a minimally invasive procedure doesnât require antibiotics. This is not contradicted by the notes, as the specific reasoning for prophylaxis isnât given, only that it was received. Minimal-pair label: [Multi-difference] - Effective Distractor Type 3: Plausible-but-absent management action, intervention, or monitoring strategy (Source B) Explanation: State that the patient received a single perioperative dose of intravenous Cefazolin as antibiotic prophylaxis. Reasoning: Cefazolin is a standard and very common choice for surgical prophylaxis for a nephrectomy. The note confirms "perioperative antibiotic prophylaxis" was given but does not name the specific drug. This distractor fills in a plausible but undocumented detail, making it difficult to reject without recognizing itâs not in the source text. Minimal-pair label: [Multi-difference] - Effective Distractor Type 4: Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A) Explanation: State that antibiotic prophylaxis was administered due to the patientâs history of a heart murmur to prevent endocarditis. 56 Reasoning: The patientâs history of a heart murmur is documented in the notes. Antibiotic prophylaxis for procedures to prevent endocarditis is a valid clinical concept. This distractor misattributes the reason for the prophylaxis from the nephrectomy (correct) to the heart murmur (incorrect for this procedure, but plausible). Rejecting it requires knowing both the correct indication and that GU procedures typically do not require endocarditis prophylaxis. Minimal-pair label: [Multi-difference] - Effective Distractor Type 5: Note-grounded detail with one external wrong element (Source A+B hybrid) Explanation: State that no active infection was present, but perioperative antibiotics were given because a pre-operative urinalysis was positive for bacteria. Reasoning: The answer correctly states no infection was diagnosed and prophylaxis was given (note-grounded). It adds a plausible but external reason for this action (a positive pre-op UA). This is a common clinical scenario, but there is no mention of a pre-operative urinalysis in the notes. This requires the model to identify the added, undocumented clinical justification. Minimal-pair label: [Multi-difference] - Effective Distractor Type 6: Cross-admission temporal displacement (Source A) Explanation: Not applicable: single admission context for this question --- substituting additional Type 1 direction: State that an incisional infection was noted at the time of discharge. Reasoning: This directly contradicts the noteâs statement "incision was clean, dry, and intact, with no evidence of hematoma collection or infection." However, surgical site infections are a common complication, making this a plausible negative outcome. A model must perform careful text verification to reject it. Minimal-pair label: [Minimal-pair: differs only in the presence of an incisional infection] Q2: Was the patient monitored for any signs of infection after that intervention? - Referential Ambiguity Analysis: Referential expression: "that intervention" Correct referent: laparoscopic right radical nephrectomy Alternative referent 1: "perioperative antibiotic prophylaxis" (from A1) Distractor direction for alternative referent 1: "Yes, the patient was monitored for common side effects of antibiotic prophylaxis, such as rash or diarrhea, but none were observed." This is a correct form of monitoring for an alternative (but incorrect) interpretation of the referent. - Cross-Turn Error Path Analysis: Prior error path 1: Believing a UTI was diagnosed during the first admission (from Q1, Type 1 distractor). Trap distractor: "Yes, a repeat urinalysis was performed on the day of discharge to ensure the urinary tract infection had cleared, and the results were negative." Clinical reasoning: If a UTI was diagnosed, checking for clearance before discharge is a logical next step in management. This answer is coherent with the prior error. Prior error path 2: Believing antibiotic prophylaxis was given for a heart murmur (from Q1, Type 4 distractor). Trap distractor: "Yes, the patientâs cardiac rhythm was continuously monitored and serial EKGs were performed to watch for any signs of developing endocarditis, all of which were normal." Clinical reasoning: If the prophylaxis was for endocarditis risk, cardiac monitoring would be a logical follow-up, creating a consistent but incorrect narrative. - Effective Distractor Type 1: Note-grounded entity substitution with clinical reasoning requirement (Source A) Explanation: State that the surgical incision was examined on discharge and found to be clean but with a small amount of serous drainage, which was considered a normal post-operative finding. Reasoning: This takes the correct action (examining the incision) but 57 changes the finding from "clean, dry, and intact" to include "serous drainage." Serous drainage can be a normal finding, making this clinically plausible. It requires precise text matching to know the note explicitly states "clean, dry, and intact." Minimal-pair label: [Minimal-pair: differs only in the description of the incision finding] - Effective Distractor Type 2: Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B) Explanation: "Yes, routine daily blood cultures were drawn on post-operative days 1 and 2 to screen for bacteremia, and both sets were negative." Reasoning: While not standard for an uncomplicated laparoscopic nephrectomy, drawing surveillance blood cultures is a monitoring strategy used in some higher-risk surgical patients. This represents a plausible (but absent) monitoring action that a model cannot disprove simply by reading the notes. Minimal-pair label: [Multi-difference] - Effective Distractor Type 3: Plausible-but-absent management action, intervention, or monitoring strategy (Source B) Explanation: State that in addition to vital signs, daily C-reactive protein (CRP) levels were trended to monitor the post-operative inflammatory response, and they showed a consistent downward trend before discharge. Reasoning: Monitoring inflammatory markers like CRP is a common way to track post-operative recovery and screen for infection. This is a plausible but undocumented monitoring strategy. Minimal-pair label: [Multi-difference] - Effective Distractor Type 4: Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A or B) Explanation: "Yes, a repeat urinalysis was performed on the day of discharge to ensure the urinary tract infection had cleared, and the results were negative." Reasoning: This is the cross-turn error path distractor. It takes a real monitoring concept (urinalysis) and misattributes it as follow-up for a non-existent UTI, which might have been selected as an answer in Q1. This makes the distractor seem like a logical continuation of an incorrect clinical picture. Minimal-pair label: [Multi-difference] - Effective Distractor Type 5: Note-grounded detail with one external wrong element (Source A+B hybrid) Explanation: State that the incision was examined and found to be clean, dry, and intact, and the Steri-strips were reinforced in one area as a precaution. Reasoning: This combines the correct findings from the note ("clean, dry, and intact") with a minor, plausible, but undocumented intervention ("Steri-strips were reinforced"). The discharge instructions mention Steri-strips, making this detail feel grounded in the text, but the action of reinforcing them is not mentioned. Minimal-pair label: [Multi-difference] - Effective Distractor Type 6: Cross-admission temporal displacement (Source A) Explanation: Not applicable: single admission context for this question --- substituting additional Type 1 direction: State that the incision was examined and found to be clean, dry, and intact, but a small hematoma was noted adjacent to the primary incision. Reasoning: This directly contradicts a specific negative finding in the note. The text says "no evidence of hematoma collection," but a small post-operative hematoma is a very common and plausible finding, making this a challenging distractor that requires careful reading. Minimal-pair label: [Minimal-pair: differs only in the presence of a hematoma] 58 F.3 Step 3: Content Answer-Choice Generation The Step 3 output for one content question (Q2) is shown below, from the Microbiology-category multi-turn QA sample for the same patient used in Appendices F.1 and F.2. In Step 3, Gemini-2.5- Pro generates the answer choices for the full sequence of Step 1 content questions in a single pass; for brevity, we show the output for one content question only. For each content question, the model produces one correct answer choice (A) and four incorrect answer choices (I1âI4). Each choice is accompanied by a Reasoning line stating why the correct answer is correct and why each incorrect choice is incorrect. Incorrect-choice reasoning is additionally annotated with its grounding source (Source: A for note-grounded content or Source: B for clinical-knowledge-grounded content, as defined in Section 3.3), a [Minimal-pair: ...] label where applicable, indicating that the distractor differs from the correct answer in exactly one detail, and an Error-path: tag identifying the cross-turn distractor chain (the specific wrong Q nâ1 answer that the distractor follows from). Q2: Was the patient monitored for any signs of infection after that intervention? A2: Yes, the surgical incision was examined on the day of discharge and was found to be clean, dry, and intact with no signs of infection. A2-Reasoning: The correct answer is a direct paraphrase of the last sentence of the Brief Hospital Course in note #1: "On exam, incision was clean, dry, and intact, with no evidence of hematoma collection or infection." This describes the monitoring that occurred. The referent "that intervention" is the laparoscopic right radical nephrectomy from A1. I1: Yes, the surgical incision was examined on the day of discharge and was noted to have a small amount of serous drainage, which was considered a normal post-operative finding. I1-Reasoning: Source: A. [Minimal-pair: differs in the description of the incision finding]. This is incorrect. This distractor correctly identifies the monitoring action (examining the incision) but alters the finding. The note states the incision was "clean, dry, and intact." While a small amount of serous drainage can be a normal post-operative finding, it is not what was documented. This tests for precise reading of the exam description. I2: Yes, a repeat urinalysis was performed on the day of discharge to ensure the urinary tract infection noted on post-operative day 1 had cleared, and the results were negative. I2-Reasoning: Source: A. Error-path: follows from Q1 wrong answer âa UTI was diagnosedâ. If that were correct, this choice would be correct because checking for clearance of a UTI before discharge is standard practice. This is incorrect because no UTI was ever diagnosed, and no repeat urinalysis is mentioned. It uses the note-grounded concept of the Foley catheter to build a plausible but false narrative. I3: Yes, in addition to routine vital signs, daily C-reactive protein levels were trended to monitor the post-operative inflammatory response, and they showed a consistent downward trend before discharge. I3-Reasoning: Source: B. This is incorrect. Trending inflammatory markers like CRP is a plausible post-operative monitoring strategy, but it is not documented anywhere in the patientâs notes. A model must know that this specific monitoring was not performed to reject this choice. I4: Yes, given that the antibiotic prophylaxis was for his heart murmur, the patientâs cardiac rhythm was continuously monitored on telemetry to watch for any signs of developing endocarditis, and all readings were normal. I4-Reasoning: Source: A. Error-path: follows from Q1 wrong answer âprophylaxis was for endocarditis due to heart murmurâ. If that were correct, this choice would be correct because cardiac monitoring would be a logical follow-up. This is incorrect because the prophylaxis was for the surgery, not endocarditis, and no such cardiac monitoring is mentioned. It incorrectly links the note-grounded "heart murmur" to a false monitoring plan. 59 F.4 Step 4: Evidence-Grounding Answer-Choice Generation The Step 4 output for one evidence-grounding question is shown belowâthe evidence-grounding question paired with content question Q 2 , from the Microbiology-category multi-turn QA sample for the same patient used in Appendices F.1âF.3 (in the raw output below it is labeled Q2-1). In Step 4, Gemini-2.5-Pro generates the evidence-grounding answer choices for the full sequence of content questions in a single pass; for brevity, we show the output for one evidence-grounding ques- tion only. Each evidence-grounding question asks for the minimal yet sufficient set ofâ¨note number, chartdate, section header(s)⊠tuples that fully support the paired content questionâs Step 3 correct answer. For each, the model produces one correct choice (labeled A2-1 in the raw output) and four incorrect choices (I1âI4). Each choice is accompanied by a Reasoning line: for the correct choice it justifies both sufficiency (the selected section(s) fully cover the answer) and minimality (no other section is required), and for each incorrect choice it states the error-pattern category from Appendix E: Category 1a (all listed headers irrelevant or from the wrong admission), Category 1b (listed headers cover only part of the answer), Category 1c (mix of partially relevant and ir- relevant headers), or Category 2 (the correct subset plus irrelevant additions). Several distractors use semantic trapsâheaders whose names plausibly align with the answerâs surface keywords (e.g., Physical Exam for an answer that mentions an examination, even though the actual exam content is in Brief Hospital Course)âso that the question cannot be solved by header-name matching alone. As in Appendix F.1, chartdates in the answer choices are shown as placeholders to avoid disclosing patient-identifying dates. Q2-1: What are the minimal exact sources from the patientâs discharge summaries that fully contain information for your previous answer? A2-1: Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course A2-1-Reasoning: Fully covers A2: Brief Hospital Course explicitly states, "On exam, incision was clean, dry, and intact, with no evidence of hematoma collection or infection," which is part of the summary of events leading to discharge. Minimal: This single header contains all the necessary information. I1: Note #1 Chartdate: [chartdate1] Headers: Physical Exam I1-Reasoning: Category 1b. This is a semantic trap. The âPhysical Examâ header notes "Incisions otherwise c/d/i" (clean/dry/intact), which partially supports the answer. However, it is less explicit than âBrief Hospital Courseâ as it does not state "no signs of infection" and does not clearly place the exam on the day of discharge. I2: Note #1 Chartdate: [chartdate1] Headers: Discharge Condition, Discharge Instructions I2-Reasoning: Category 1a. This is a semantic trap. âDischarge Conditionâ describes the patientâs general status (e.g., "Ambulatory - Independent") but does not mention the incision exam. âDischarge Instructionsâ provides guidance on wound care but does not describe the findings of the exam performed at discharge. I3: Note #1 Chartdate: [chartdate1] Headers: Brief Hospital Course, Physical Exam I3-Reasoning: Category 2. âBrief Hospital Courseâ is sufficient to fully support the answer. While âPhysical Examâ mentions the incision is "c/d/i", this information is already more completely stated in âBrief Hospital Courseâ ("clean, dry, and intact, with no evidence of ... infection"). Therefore, âPhysical Examâ is redundant. I4: Note #1 Chartdate: [chartdate1] Headers: Discharge Instructions I4-Reasoning: Category 1b. This is a semantic trap. This header explains how to care for the incision post-discharge but does not contain the findings of the examination performed at discharge (i.e., that the incision was found to be clean, dry, and intact with no signs of infection). 60 G Medical Expert Review and Revision The instructions given to the medical expert reviewers and the interfaces they used during the three- week dataset finalization stage (Section 3.4) are shown below. G.1 Reviewer Instructions The full instructions provided to each of the 11 medical expert reviewers are shown below. Purpose. This study aims to construct a multi-turn question-answering evaluation dataset for assessing how well language models understand patient discharge summaries. The QA examples were initially generated using large language models, including Gemini, and were reviewed by medical experts to ensure clinical validity, factual accuracy, and appropriate grounding in the patient notes. Notation. For a given multi-turn QA sample, we denote each discharge-summary content question as Qx, its correct answer as Ax, and the corresponding source-location question as Qx-1. Each Qx asks about clinical information in the discharge summaries; each Qx-1 asks which note number and section header(s) contain the evidence supporting Ax. Reviewing discharge-summary content questions. For each Qx, check whether the ques- tion is clinically natural and appropriate. If it is phrased in an unnatural or exam-like way, revise it into a form that a clinician might naturally use when reviewing a patient note. Also revise questions that are overly ambiguous or that contain excessive hints about the answer. If the question is understandable and does not interfere with answer selection, no revision is needed. Reviewing answer options. For each Qx, review the correct answer and the four incorrect answer options. ⢠The correct answer should be factually accurate and fully supported by the patient note. If it contains any incorrect information, revise it. ⢠Each incorrect answer option should be clearly wrong while remaining clinically plausible. If an incorrect option is not actually incorrect, revise it so that it contains a plausible but unsupported or incorrect clinical detail. ⢠When revising, keep the lengths of the correct and incorrect options reasonably sim- ilar, so that the answer is not identifiable based on option length alone. Reviewing source-location questions. For each Qx-1, revise the source options according to the following rules. Correct option. The correct source option should consist of the minimal set of section headers that contains the full content of Ax. The goal is not to include every section that may help answer Qx, but only the smallest set of headers that fully supports Ax. For example, if the content of Ax is partially mentioned in Discharge Diagnosis but fully covered by Brief Hospital Course, the correct source option should be Brief Hospital Course only. Incorrect options. Each incorrect source option should satisfy at least one of the following: ⢠It contains only part of Ax and does not cover the full answer. ⢠It contains the header(s) that fully support Ax, but also includes one or more headers unrelated to Ax, making the option non-minimal. Special edge cases. Two situations require additional care: ⢠Correct header plus a partially relevant header. If an incorrect option consists of the correct source header plus another header that contains part of Ax, revise it. For example, if Brief Hospital Course alone fully supports Ax, then Brief Hospital Course + Discharge Diagnosis should not remain as an incorrect option if it could also be interpreted as correct. Instead, revise the option so that it either contains only partial evidence (e.g., Discharge Diagnosis alone) or includes an unrelated header so that it is clearly non-minimal. 61 ⢠Multiple headers independently containing Ax. If multiple headers each contain the full content of Ax, do not use one of those headers alone as an incorrect option. For example, if both Pertinent Results and Brief Hospital Course each fully contain Ax, one may be used as the correct source option, but the other must not appear alone as an incorrect option. Replace it with a header that does not contain Ax, or add an unrelated header to make it clearly non-minimal. Choosing distractor headers. When adding unrelated headers to incorrect source options, you may use a completely unrelated header. Preferably, however, use a header whose name appears related by title but does not actually contain the content of Ax. For example, if Ax concerns medication information and only Brief Hospital Course fully contains Ax, while Discharge Medications and Medications on Admission do not contain the relevant informa- tion, appropriate incorrect options may include Discharge Medications, Medications on Ad- mission, or Brief Hospital Course + Discharge Medications. Header selection. Use main section headers rather than subheaders whenever possible. For example, if the relevant information appears under Microbiology or Imaging within Perti- nent Results, use Pertinent Results as the source header. Common main headers include Brief Hospital Course, Major Surgical or Invasive Procedure, History of Present Illness, Discharge Diagnosis, Discharge Medications, Chief Complaint, Discharge Instructions, Pertinent Re- sults, Past Medical History, Medications on Admission, Allergies, Discharge Condition, and Physical Exam. Removal criteria. A sample may be removed if the question is too clinically unnatural to revise, or if the required source header cannot be inferred because of severe de-identification artifacts. If a sample must be removed, mark it with the removal flag and provide the reason in the comments field. 62 G.2 Reviewer Interface The two interfaces used by the reviewers during the dataset finalization stage (Section 3.4) are shown below. Figure 9 shows the Streamlit interface, which displays each multi-turn QA chain (each ques- tion paired with its correct and four incorrect answer choices) alongside the patientâs complete set of discharge summaries, allowing reviewers to verify each choice against the source notes. Figure 10 shows the accompanying spreadsheet, in which reviewers recorded revisions to questions, correct answers, and incorrect answers. Figure 9: Streamlit interface used by medical expert reviewers. For each multi-turn QA sample, the interface presents the full QA chain (each question with its correct answer choice and four incorrect answer choices) side-by-side with the patientâs complete set of discharge summaries, so that every choice can be verified against the source notes. Figure 10: Spreadsheet used by medical expert reviewers to record revisions to questions, correct answer choices, and incorrect answer choices flagged during review. 63 H Evaluation Details Inference framework and hardware. Open-weight models are evaluated with vLLM [21] using each modelâs native chat template, on NVIDIA A6000 GPUs with 1â8 GPUs per model under tensor parallelism. Reasoning-model handling. For chain-of-thought models that emit <think>...</think> blocks, we keep only the text following the final </think> as the modelâs answer, discarding the reasoning span both for answer extraction and when appending the turn to the conversation history. Answer extraction. At each turn the model selects one of AâE; when scoring the model outputs, we extract the chosen letter from the (reasoning-stripped) output using a rule-based parser, with model-specific handling for differences in output formatting. Proprietary model access. All use of MIMIC-IV-derived content with proprietary modelsâboth Gemini-2.5-Pro in the data-generation pipeline (Section 3.2) and the proprietary models evaluated as benchmark subjects (gpt-5.4, gpt-5.4-mini, gemini-3-flash-preview)âwas conducted through HIPAA-compliant deployments: Azure OpenAI for the GPT models and Google Cloud Vertex AI for the Gemini models, consistent with the MIMIC-IV data use agreement. Evaluation cost. Proprietary-API evaluation over the 967 multi-turn samples cost approximately $120 for gpt-5.4, $40 for gpt-5.4-mini, and $40 for gemini-3-flash-preview. 64 I Per-Step Prompts for LLM-Based Initial Data Generation The full prompts used at each step of the four-step generation pipeline (Section 3.3) are shown below. For each prompt, we first list the curly-brace placeholders that are substituted per sample at generation time, together with what each is replaced with. (Other curly-brace tokens appearing inside a prompt are part of that promptâs output-format specification and are not substituted.) I.1 Step 1: Multi-turn Question Generation We use a single-note prompt variant for patients with one discharge summary and a multi-note variant for patients with two or more. Placeholders. qa_template_category â the target question category (one of the eight); qa_template_examples â the five expert-curated QA templates for that category (single- or multi-note variant); qa_template_examples_other â one randomly sampled template from each of the other seven categories; discharge_summaries â the patientâs full sequence of dis- charge summaries. I.1.1 Single-Note Prompt You are provided with a patientâs discharge summary. You are also provided with an entity category, attribute, relationship schema defined for tagging discharge summaries, along with multi-turn QA template examples regarding multiple pre-defined categories. The multi-turn QA templates are curated by medical experts (i.e., clinician, nurse). Carefully read the patientâs discharge summary and generate a multi-turn QA that medical experts would realistically ask in real-world clinical practice when reviewing the discharge summary. Among the QA template categories mentioned below, generate multi-turn QA regarding qa_template_category category, by referencing and following the reasoning style demonstrated in the qa_template_category QA template examples below. You must generate: 1. A sequence of medical expert-style multi-turn questions. 2. All entities, attributes, and relationships between entities extracted from the discharge summaries required to answer each corresponding question, 3. Source locations of the extracted entities, attributes, and relationships that compose the answer content. (Sufficient but minimal) 4. The answer. Multi-turn Reasoning Dependency (CRITICAL): The QA sequence must simulate stepwise clinical reasoning with narrative coherence, similar to how clinicians progressively review charts. Rules: - Q1 may introduce entities directly from the discharge summaries. - Q2 and later questions must reference entities introduced in previous answers. - Later questions must not independently re-introduce entities that have already appeared in prior answers. This ensures thematic continuity across the reasoning chain. Example logic: A1 identifies Entity X Q2 refers to Entity X using a coreference expression (e.g., "the diagnosis", "the procedure") A2 concerns Entity X or a closely related entity Unresolvability Requirement: For each Q2+ question, verify that the question CANNOT be answered correctly by reading only the discharge summaries without knowing the prior answers. If a reader with no prior context could identify the correct answer by searching the notes for the entity referenced in the question, the question is insufficiently dependent on prior context. Revise it to increase referential ambiguity --- the coreference expression must be genuinely ambiguous without prior context. For example, if the patient has only one procedure, "What was the outcome of that procedure?" is effectively self-resolving. In such cases, broaden the coreference ("What was the outcome of that intervention?") or restructure the chain so that the referenced entity has multiple plausible referents in the notes. Entity Introduction Rules - Q1 may name entities directly from the discharge summaries, as there is no prior answer to reference. - For Q2 and later questions, the question must not introduce new entity names from the discharge summaries. Entities for Q2 and later may only be referenced using coreference expressions that refer back to entities already introduced in a prior answer. - When referencing entities in Q2 and later questions, you must not include modifiers or details that reveal identity (e.g., type, location, mechanism). 65 For example, instead of "What antithrombotic treatment was started", write "What treatment was started". Also, instead of "What ongoing antiplatelet regimen was prescribed", write "What ongoing regimen was prescribed". Coreference Expression Restrictions (STRICT): The coreference expression must be the most generic clinically natural term possible. Terms that reveal drug class, mechanism, anatomical category, or clinical subcategory are PROHIBITED because they narrow the referent and can break error propagation by allowing direct lookup in the notes. PROHIBITED coreference expressions (examples --- not exhaustive): - "the antibiotic", "the antifungal", "the antihypertensive", "the anticoagulant", "the antiarrhythmic", "the analgesic", "the antiplatelet agent" - "the cardiac procedure", "the orthopedic procedure", "the endoscopic procedure", "the neurologic finding" - "the renal finding", "the hepatic finding", "the pulmonary imaging", "the cardiac imaging" - "the gram-positive organism", "the fungal infection" REQUIRED generic alternatives: - "the medication", "the treatment", "the regimen" - "the procedure", "the intervention", "the operation" - "the finding", "the result", "the abnormality" - "the condition", "the diagnosis", "the complication" - "the organism", "the infection" Entity Reuse Rules: For Q2 and all subsequent questions: 1. The question must not contain entity names copied directly from the discharge summary. 2. The question must refer to entities from the previous answer using coreference expressions, such as: - "the procedure", "the diagnosis", "the complication", "the finding", "the laboratory abnormality", "the imaging result", "the organism", "the treatment", "the intervention", "the outcome" 3. For the coreference expressions, questions must not include descriptive modifiers that reveal the identity, category, mechanism, location, or clinical implication of the referenced entity. Questions must refer to entities without adding informative adjectives or phrases. Once an entity name appears in an earlier answer, the original name must not appear again in subsequent questions. Answer Continuity Constraint: For each answer Ai (i > 1): - The answer must concern the entity or clinical concept referenced in Qi. - The answer must be clinically coherent with the prior question context and advance the reasoning chain naturally. Question Naturalness Constraint (CRITICAL): Each question must sound like an open chart-review question that a clinician would actually ask while reading the notes for the first time. The asker must NOT already know the answer, and the question must NOT presuppose the existence, valence, or notability of what the answer will be. Questions that read like textbook prompts ("What critical finding was discovered...", "Given the complicated course...", "Despite the eventual procedure...") are PROHIBITED because they leak the answer and are not clinically realistic. This rule constrains how NEW information in the question is framed. The required coreference back to entities from prior answers (per the Multi-turn Reasoning Dependency, Entity Reuse, and Coreference Expression Restrictions rules above) still applies --- the back-reference itself is fine; the leakage rule applies to the rest of the stem. Forbidden patterns: The problem is structural (presupposition), not lexical. Words like "critical", "major", " significant", "key", "notable" are clinically legitimate when used inside an open question (e.g ., "Was there any significant change in his renal function?", "Was anything notable found on imaging?"). What is forbidden is asserting the existence or characterization of the answer as a fact in the stem itself. The two structural patterns to avoid: 1. Embedding the answer category as an asserted fact in the stem (declarative presupposition). The stem states that the entity to be revealed already exists, instead of asking whether it exists or what was done in a given scope. - e.g., "A new condition was identified during that admission. What intervention was performed for it?" --- both clauses presuppose findings before the answer reveals them. - e.g., "What was the specific challenge in managing his medications?" --- presupposes that a challenge existed. - e.g., "What critical finding was discovered during the evaluation?" --- presupposes that a critical finding was discovered. The fix is to convert the presupposing clause into a yes/no opener ("Was there any challenge in managing his medications?", "Was there any finding of concern on that evaluation?") or to a neutral scope question ("What did that evaluation show?"). 2. Editorial framing clauses that characterize the course or outcome before asking. A leading clause like "Given the complicated hospital course, ..." or "Despite the eventual procedure, ..." asserts that the course was complicated or that a procedure occurred before the answer reveals 66 it. The fix is to drop the editorial clause: "What was the outcome of that hospital course?", " What happened after that procedure?". Required rewriting strategy: For each question that would otherwise presuppose the answer, convert it into either: (a) a YES/NO opener that asks whether the entity exists, optionally followed in the SAME turn or the NEXT turn by a specific follow-up question about its details, OR (b) a neutral wh-question that names the SCOPE of inquiry (the workup, the plan, the medication regimen, the disposition, what happened next) without evaluative adjectives and without asserting what the answer will be. The new entity introduced by the answer must still be unrecoverable from the question alone --- i.e ., the question may name the SCOPE (what was done, what was found, what happened next, how the patient was doing) but must not name or telegraph the ANSWER ENTITY itself. Paired BAD -> GOOD examples: BAD: "What critical finding was discovered during the evaluation?" GOOD: "Was there any finding of concern discovered during that evaluation?" GOOD: "What did the evaluation show?" BAD: "What was the specific challenge in managing his medications?" GOOD: "Was there any challenge in managing his medication regimen?" GOOD: "What was documented about his medication management at that point?" BAD: "A new condition was identified during that admission. What intervention was performed for it?" GOOD: "Was a new condition identified during that admission?" (next turn:) "What intervention was performed for it?" GOOD: "What was worked up during that admission, and what was done?" BAD: "Given the complicated hospital course, what was the outcome?" GOOD: "What was the outcome of that hospital course?" BAD: "Despite the eventual procedure, what happened next?" GOOD: "What happened after that procedure?" BAD: "What major clinical event occurred shortly thereafter, leading to escalation of care?" GOOD: "What happened to the patient after that?" GOOD: "Was there any clinical change after that? If so, what?" BAD: "Despite that improvement, what major clinical event occurred...?" GOOD: "What happened after that improvement?" BAD: "What was the patientâs overall functional outcome and discharge disposition at the end of that admission?" GOOD: "How was the patient functionally at discharge, and where was he discharged to?" BAD: "What critical finding was discovered during the evaluation leading up to that intervention ?" GOOD: "What did the workup before that intervention show?" BAD: "Following stabilization of that episode, a new condition was identified..." GOOD: "What happened after that episode was stabilized?" Note on yes/no openers: A yes/no opener is acceptable ONLY when the corresponding answer naturally begins with "Yes" + the full clinical details drawn from the source notes (e.g., "Yes --- atrial flutter was identified , and an ablation was performed on..."). Do NOT use yes/no openers as a loophole to hide content; the answer must still convey the same clinical information that would have been conveyed under the BAD phrasing. If the answer cannot be expanded that way, use a neutral wh- question instead. Self-check before finalizing each question: 1. Does the stem assert as fact something that the answer is supposed to reveal (e.g., "a new condition was identified", "the specific challenge in managing his medications", "the eventual procedure", "the critical finding that was discovered")? If yes, convert that clause into a yes /no opener or a neutral scope question. 2. Does the stem contain a leading editorial framing clause that characterizes the course or outcome before asking (e.g., "Given the complicated hospital course, ...", "Despite the eventual procedure, ...")? If yes, drop the editorial clause. 3. Could a clinician who has NOT yet read the relevant note section ask this question without already knowing what the answer will be? If no, rewrite until the answer would still be a genuine surprise to the asker. 4. Does the question still satisfy the Multi-turn Reasoning Dependency, Entity Reuse, and Coreference Expression Restrictions rules above (i.e., it back-references prior answer entities using only generic coreference expressions, and the back-referenced entity is genuinely ambiguous without the prior turn)? If no, restore the back-reference. 67 Template Diversity Rule: The provided templates are conceptual inspirations, not fixed question structures. When generating questions: - Do not copy templates directly. - Combine or adapt ideas from multiple templates when constructing the reasoning chain. - Adapt the reasoning chain based on the provided patient discharge summary content. Reference the QA template examples of qa_template_category when constructing questions. Question Generation Requirements: When generating questions, - Questions must reflect real clinical reasoning workflows, meaning questions medical experts actually ask when reviewing discharge summaries for patient care, handoff, or follow-up planning. - Each question must reference or build on entities or clinical concepts introduced in the previous answer. - Do not add descriptive modifiers that can reveal entity identity in later questions. - Do not add phrases such as "you identified", "you described", "from the prior answer". Instead, just assume that the prior answer context will be used to answer each question. - Avoid adding parentheses in the questions. - Questions must not introduce hypothetical or speculative clinical scenarios. - Questions must not be trivially answerable, and should require clinical details and medical experts to interpret, connect, or reason over clinical information. - Questions must not ask multiple things in a single question. Split them into separate questions. - Questions must be written in natural clinical language reflecting how medical experts actually phrase questions. - Each subsequent question in the multi-turn chain should logically follow prior questions in a way medical experts would naturally proceed during chart review. - You must reference the QA template examples of qa_template_category. Use them as seed to construct real-world questions that medical-experts would ask. Turn Count Guidance: Generate 5-7 content questions, calibrated to the patientâs clinical complexity. Dependency quality takes absolute priority over turn quantity. Every question must maintain genuine clinical reasoning dependency on prior answers. Do not generate additional questions merely to increase turn count. A strong 5-turn chain with tight dependency is far better than a 7-turn chain where later turns are self-contained or only loosely connected. If you cannot construct a clinically natural, dependent question that maintains the reasoning chain, stop the chain. Do not generate unnatural questions that clinicians would not realistically ask. WHY/HOW Question Guidance: You are encouraged to include WHY and HOW questions at Q3 and later positions. These questions strengthen both clinical naturalness and error propagation because they require understanding clinical reasoning chains, which forces reliance on prior-turn context. Use phrasing such as: - "Because [referencing prior answer context], how was the management adjusted?" - "Given [referenced finding from prior answer], why was that approach chosen?" - "What clinical reasoning led to the change in [the treatment / the intervention]?" - "How did [those findings / that response] influence the subsequent management?" Answerable-from-notes constraint (CRITICAL): All WHY/HOW questions MUST be answerable from information explicitly stated or directly inferable from the discharge summaries combined with standard clinical knowledge to interpret those facts. Permitted reasoning types: - Connecting two documented events in a causal or temporal sequence (e.g., a documented finding followed by a documented clinical decision) - Interpreting documented lab values, imaging results, or exam findings using standard clinical knowledge - Explaining a documented clinical decision using the rationale stated or clearly implied in the notes - Linking documented information across different sections of the discharge summary for the same patient PROHIBITED reasoning types (questions requiring these are NOT allowed): - Pathophysiological mechanisms not described in the notes (e.g., "Why does Drug X cause side effect Y at the molecular level?") - Clinical guideline reasoning not cited or implied in the notes (e.g., "Why is Drug X preferred over Drug Y per guidelines?") - Pharmacological mechanisms of action (e.g., "How does this medication work?") - Counterfactual scenarios (e.g., "What would have happened if treatment X had not been given?") - Epidemiological or population-level statistics (e.g., "Why is this condition more common in population X?") - Comparison with alternative treatments not mentioned in the notes (e.g., "Why was ticagrelor chosen over prasugrel?" when the notes do not discuss this choice) The output format must be as follows: Q1: medical expert question regarding the patientâs discharge summaries 68 A1: [Entities, Attributes, Relationships] Entity1 - attribute_name: attribute, attribute_name: attribute Entity2 - attribute_name: attribute, attribute_name: attribute Entity1 -> relationship_name -> Entity2 Entity3 - attribute_name: attribute, attribute_name: attribute Entity2 -> relationship_name -> Entity3 ... [Source Location] Note #xHeader Name corresponding text-span1 corresponding text-span2 Note #xHeader Name corresponding text-span1 corresponding text-span2 ... [Answer] answer_text Repeat for Q2, Q3, etc. Source Location Rules: - You must add all headers that are required for answering the question, but you must add only minimal headers that are necessarily required to answer the question. For example, if the answer content appears in both brief hospital course and major surgical or invasive procedure, and the content in major surgical or invasive procedure is included in brief hospital course, you must only include brief hospital course as source location. - For the header names, you must write the main section header names (e.g., Brief Hospital Course, Major Surgical or Invasive Procedure, History of Present Illness, Discharge Diagnosis, Discharge Medications, Chief Complaint, Discharge Instructions, Pertinent Results, Past Medical History, Medications on Admission, Allergies, Discharge Condition, Physical Exam), instead of the sub-header names (e.g., Microbiology, Imaging within the Pertinent Result header). If the context is included in the sub-header, you must write the main header name instead of the sub- header name, along with its corresponding context. [Entity Category, Attribute, Relationship Schema] 1. Diagnosis -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 2. Finding -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 3. Symptom -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 4. Procedure -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 5. Outcome -- status, assertion, time(dates, times, durations, frequencies) 6. Vital_Sign -- status, value, assertion, time(dates, times, durations, frequencies) 7. Physical_Exam -- status, finding, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 8. Lab_Test -- value, abnormal_flag, status, assertion, time(dates, times, durations, frequencies) 9. Diagnostic_Imaging_Test -- status, result, finding_site, laterality, assertion, time(dates, times , durations, frequencies) 10. Specimen -- source, assertion, time(dates, times, durations, frequencies) 11. Microbiology_Test - organism_growth, result, assertion, time(dates, times, durations, frequencies) 12. Microbiology_Organism - growth, assertion, time(dates, times, durations, frequencies) 13. Microbiology_Antibiotics - dilution, sensitivity, assertion, time(dates, times, durations, frequencies) 14. Activity -- status, assertion, time(dates, times, durations, frequencies) 15. Medication -- strength, dosage, form_description, form, route, instruction, disp, refills, assertion, time(dates, times, durations, frequencies) 16. Event -- status, assertion, time(dates, times, durations, frequencies) 17. Allergy -- status, assertion, time(dates, times, durations, frequencies) 18. Medical_Device - status, assertion, time(dates, times, durations, frequencies) 19. Instruction - instruction_text, assertion, time(dates, times, durations, frequencies) is_main attribute is either true or false. If the Diagnosis/Symptom/Finding is the main Diagnosis/ Symptom/Finding for the patient during the current admission (e.g., chief complaint, discharge diagnosis(primary), main Diagnosis/Symptom/Finding according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false (e.g., discharge diagnosis( secondary), not main Diagnosis/Symptom/Finding according to context). If the Procedure is the main Procedure for the patient during the current admission (e.g., major surgical or invasive procedure, main Procedure according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false. abnormal_flag attribute in Lab_Test, is either normal or abnormal. 1. Procedure/Medication -> improves -> Diagnosis/Symptom/Finding 2. Procedure/Medication -> worsens -> Diagnosis/Symptom/Finding 3. Procedure/Medication -> causes -> Diagnosis/Symptom/Finding 4. Procedure/Medication -> administered_for -> Diagnosis/Symptom/Finding 5. Procedure/Medication -> reveals -> Diagnosis/Symptom/Finding 69 6. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> reveals -> Diagnosis/Finding 7. Specimen -> has_test -> Microbiology_Test/Lab_Test 8. Microbiology_Test -> detected -> Microbiology_Organism 9. Microbiology_Organism -> tested_against -> Microbiology_Antibiotics 10. Microbiology_Organism -> reveals -> Diagnosis/Finding 11. Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding 12. Diagnosis/Symptom/Finding -> causes -> Event 13. Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding/Activity 14. Outcome/Finding -> after -> Diagnosis/Symptom/Finding 15. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> resulted_in -> Procedure/Medication 16. Procedure/Medication -> causes -> Procedure/Medication 17. Procedure/Medication -> causes -> Event 18. Procedure/Medication -> switched_to -> Procedure/Medication 19. Procedure/Medication -> resulted_in -> Outcome/Finding 20. Outcome/Finding -> after -> Procedure/Medication 21. Event -> causes -> Diagnosis/Symptom/Finding 22. Event -> causes -> Procedure/Medication 23. (any category) -> has_instruction -> Instruction The difference between "causes" and "resulted_in" in "Procedure/Medication -> causes -> Diagnosis/ Symptom/Finding" and "Procedure/Medication -> resulted_in -> Outcome/Finding" is that "causes" is used when a procedure or medication produces an adverse effect or complication, while " resulted_in" is used when a procedure or medication leads to a desired or expected state or improvement (positive or neutral). Similarly, the difference between "causes" and "resulted_in" in "Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding" and "Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding " is that "causes" is used when a diagnosis, symptom, or finding produces an adverse effect or complication, while "resulted_in" is used when a diagnosis, symptom, or finding leads to a desired or expected state or improvement (positive or neutral). [Multi-turn QA Templates - qa_template_category] qa_template_examples [Multi-turn QA Templates - Other Categories] qa_template_examples_other [Patientâs Discharge Summary] discharge_summaries I.1.2 Multi-Note Prompt You are provided with a patientâs discharge summaries. You are also provided with an entity category , attribute, relationship schema defined for tagging discharge summaries, along with multi-turn QA template examples regarding multiple pre-defined categories. The multi-turn QA templates are curated by medical experts (i.e., clinician, nurse). Carefully read the patientâs discharge summaries and generate a multi-turn QA that medical experts would realistically ask in real-world clinical practice when reviewing discharge summaries. Among the QA template categories mentioned below, generate multi-turn QA regarding qa_template_category category, by referencing and following the reasoning style demonstrated in the qa_template_category QA template examples below. You must generate: 1. A sequence of medical expert-style multi-turn questions. 2. All entities, attributes, and relationships between entities extracted from the discharge summaries required to answer each corresponding question, 3. Source locations of the extracted entities, attributes, and relationships that compose the answer content. (Sufficient but minimal) 4. The answer. Multi-turn Reasoning Dependency (CRITICAL): The QA sequence must simulate stepwise clinical reasoning with narrative coherence, similar to how clinicians progressively review charts. Rules: - Q1 may introduce entities directly from the discharge summaries. - Q2 and later questions must reference entities introduced in previous answers. - Later questions must not independently re-introduce entities that have already appeared in prior answers. This ensures thematic continuity across the reasoning chain. Example logic: A1 identifies Entity X Q2 refers to Entity X using a coreference expression (e.g., "the diagnosis", "the procedure") A2 concerns Entity X or a closely related entity Unresolvability Requirement: 70 For each Q2+ question, verify that the question CANNOT be answered correctly by reading only the discharge summaries without knowing the prior answers. If a reader with no prior context could identify the correct answer by searching the notes for the entity referenced in the question, the question is insufficiently dependent on prior context. Revise it to increase referential ambiguity --- the coreference expression must be genuinely ambiguous without prior context. For example, if the patient has only one procedure, "What was the outcome of that procedure?" is effectively self-resolving. In such cases, broaden the coreference ("What was the outcome of that intervention?") or restructure the chain so that the referenced entity has multiple plausible referents in the notes. Entity Introduction Rules - Q1 may name entities directly from the discharge summaries, as there is no prior answer to reference. - For Q2 and later questions, the question must not introduce new entity names from the discharge summaries. Entities for Q2 and later may only be referenced using coreference expressions that refer back to entities already introduced in a prior answer. - When referencing entities in Q2 and later questions, you must not include modifiers or details that reveal identity (e.g., type, location, mechanism). For example, instead of "What antithrombotic treatment was started", write "What treatment was started". Also, instead of "What ongoing antiplatelet regimen was prescribed", write "What ongoing regimen was prescribed". Coreference Expression Restrictions (STRICT): The coreference expression must be the most generic clinically natural term possible. Terms that reveal drug class, mechanism, anatomical category, or clinical subcategory are PROHIBITED because they narrow the referent and can break error propagation by allowing direct lookup in the notes. PROHIBITED coreference expressions (examples --- not exhaustive): - "the antibiotic", "the antifungal", "the antihypertensive", "the anticoagulant", "the antiarrhythmic", "the analgesic", "the antiplatelet agent" - "the cardiac procedure", "the orthopedic procedure", "the endoscopic procedure", "the neurologic finding" - "the renal finding", "the hepatic finding", "the pulmonary imaging", "the cardiac imaging" - "the gram-positive organism", "the fungal infection" REQUIRED generic alternatives: - "the medication", "the treatment", "the regimen" - "the procedure", "the intervention", "the operation" - "the finding", "the result", "the abnormality" - "the condition", "the diagnosis", "the complication" - "the organism", "the infection" Entity Reuse Rules: For Q2 and all subsequent questions: 1. The question must not contain entity names copied directly from the discharge summary. 2. The question must refer to entities from the previous answer using coreference expressions, such as: - "the procedure", "the diagnosis", "the complication", "the finding", "the laboratory abnormality", "the imaging result", "the organism", "the treatment", "the intervention", "the outcome" 3. For the coreference expressions, questions must not include descriptive modifiers that reveal the identity, category, mechanism, location, or clinical implication of the referenced entity. Questions must refer to entities without adding informative adjectives or phrases. Once an entity name appears in an earlier answer, the original name must not appear again in subsequent questions. Answer Continuity Constraint: For each answer Ai (i > 1): - The answer must concern the entity or clinical concept referenced in Qi. - The answer must be clinically coherent with the prior question context and advance the reasoning chain naturally. Question Naturalness Constraint (CRITICAL): Each question must sound like an open chart-review question that a clinician would actually ask while reading the notes for the first time. The asker must NOT already know the answer, and the question must NOT presuppose the existence, valence, or notability of what the answer will be. Questions that read like textbook prompts ("What critical finding was discovered...", "Given the complicated course...", "Despite the eventual procedure...") are PROHIBITED because they leak the answer and are not clinically realistic. This rule constrains how NEW information in the question is framed. The required coreference back to entities from prior answers (per the Multi-turn Reasoning Dependency, Entity Reuse, and Coreference Expression Restrictions rules above) still applies --- the back-reference itself is fine; the leakage rule applies to the rest of the stem. Forbidden patterns: The problem is structural (presupposition), not lexical. Words like "critical", "major", " significant", "key", "notable" are clinically legitimate when used inside an open question (e.g ., "Was there any significant change in his renal function?", "Was anything notable found on 71 imaging?"). What is forbidden is asserting the existence or characterization of the answer as a fact in the stem itself. The two structural patterns to avoid: 1. Embedding the answer category as an asserted fact in the stem (declarative presupposition). The stem states that the entity to be revealed already exists, instead of asking whether it exists or what was done in a given scope. - e.g., "A new condition was identified during that admission. What intervention was performed for it?" --- both clauses presuppose findings before the answer reveals them. - e.g., "What was the specific challenge in managing his medications?" --- presupposes that a challenge existed. - e.g., "What critical finding was discovered during the evaluation?" --- presupposes that a critical finding was discovered. The fix is to convert the presupposing clause into a yes/no opener ("Was there any challenge in managing his medications?", "Was there any finding of concern on that evaluation?") or to a neutral scope question ("What did that evaluation show?"). 2. Editorial framing clauses that characterize the course or outcome before asking. A leading clause like "Given the complicated hospital course, ..." or "Despite the eventual procedure, ..." asserts that the course was complicated or that a procedure occurred before the answer reveals it. The fix is to drop the editorial clause: "What was the outcome of that hospital course?", " What happened after that procedure?". Required rewriting strategy: For each question that would otherwise presuppose the answer, convert it into either: (a) a YES/NO opener that asks whether the entity exists, optionally followed in the SAME turn or the NEXT turn by a specific follow-up question about its details, OR (b) a neutral wh-question that names the SCOPE of inquiry (the workup, the plan, the medication regimen, the disposition, what happened next) without evaluative adjectives and without asserting what the answer will be. The new entity introduced by the answer must still be unrecoverable from the question alone --- i.e ., the question may name the SCOPE (what was done, what was found, what happened next, how the patient was doing) but must not name or telegraph the ANSWER ENTITY itself. Paired BAD -> GOOD examples: BAD: "What critical finding was discovered during the evaluation?" GOOD: "Was there any finding of concern discovered during that evaluation?" GOOD: "What did the evaluation show?" BAD: "What was the specific challenge in managing his medications?" GOOD: "Was there any challenge in managing his medication regimen?" GOOD: "What was documented about his medication management at that point?" BAD: "A new condition was identified during that admission. What intervention was performed for it?" GOOD: "Was a new condition identified during that admission?" (next turn:) "What intervention was performed for it?" GOOD: "What was worked up during that admission, and what was done?" BAD: "Given the complicated hospital course, what was the outcome?" GOOD: "What was the outcome of that hospital course?" BAD: "Despite the eventual procedure, what happened next?" GOOD: "What happened after that procedure?" BAD: "What major clinical event occurred shortly thereafter, leading to another hospitalization?" GOOD: "What happened to the patient after that discharge?" GOOD: "Was the patient readmitted? If so, why?" BAD: "Despite that improvement, what major clinical event occurred...?" GOOD: "What happened after that improvement?" BAD: "What was the patientâs overall functional outcome and discharge disposition at the end of that admission?" GOOD: "How was the patient functionally at discharge, and where was he discharged to?" BAD: "What critical finding was discovered during the evaluation leading up to that intervention ?" GOOD: "What did the workup before that intervention show?" BAD: "Following stabilization of that psychiatric episode, a new condition was identified..." GOOD: "What happened on the next admission after that episode?" Note on yes/no openers: A yes/no opener is acceptable ONLY when the corresponding answer naturally begins with "Yes" + the full clinical details drawn from the source notes (e.g., "Yes --- atrial flutter was identified , and an ablation was performed on..."). Do NOT use yes/no openers as a loophole to hide 72 content; the answer must still convey the same clinical information that would have been conveyed under the BAD phrasing. If the answer cannot be expanded that way, use a neutral wh- question instead. Self-check before finalizing each question: 1. Does the stem assert as fact something that the answer is supposed to reveal (e.g., "a new condition was identified", "the specific challenge in managing his medications", "the eventual procedure", "the critical finding that was discovered")? If yes, convert that clause into a yes /no opener or a neutral scope question. 2. Does the stem contain a leading editorial framing clause that characterizes the course or outcome before asking (e.g., "Given the complicated hospital course, ...", "Despite the eventual procedure, ...")? If yes, drop the editorial clause. 3. Could a clinician who has NOT yet read the relevant note section ask this question without already knowing what the answer will be? If no, rewrite until the answer would still be a genuine surprise to the asker. 4. Does the question still satisfy the Multi-turn Reasoning Dependency, Entity Reuse, and Coreference Expression Restrictions rules above (i.e., it back-references prior answer entities using only generic coreference expressions, and the back-referenced entity is genuinely ambiguous without the prior turn)? If no, restore the back-reference. Template Diversity Rule: The provided templates are conceptual inspirations, not fixed question structures. When generating questions: - Do not copy templates directly. - Combine or adapt ideas from multiple templates when constructing the reasoning chain. - Adapt the reasoning chain based on the provided patient discharge summary content. Reference the QA template examples of qa_template_category when constructing questions. Question Generation Requirements: When generating questions, - Questions must reflect real clinical reasoning workflows, meaning questions medical experts actually ask when reviewing discharge summaries for patient care, handoff, or follow-up planning. - Each question must reference or build on entities or clinical concepts introduced in the previous answer. - Do not add descriptive modifiers that can reveal entity identity in later questions. - When referencing admissions for the first time, identify them using either admission order (first, second, etc.), chartdate, admission ID, a specific entity (e.g., diagnosis, procedure etc.) that uniquely identifies the admission, or all admissions when the question concerns the entire admission history. - Do not add phrases such as "you identified", "you described", "from the prior answer". Instead, just assume that the prior answer context will be used to answer each question. - Avoid adding parentheses in the questions. - Questions must not introduce hypothetical or speculative clinical scenarios. - Questions must not be trivially answerable, and should require clinical details and medical experts to interpret, connect, or reason over clinical information. - Questions must not ask multiple things in a single question. Split them into separate questions. - Questions must be written in natural clinical language reflecting how medical experts actually phrase questions. - Each subsequent question in the multi-turn chain should logically follow prior questions in a way medical experts would naturally proceed during chart review. - You must reference the QA template examples of qa_template_category. Use them as seed to construct real-world questions that medical-experts would ask. Multi-Note Coverage: You are highly encouraged to generate a QA chain that spans all provided discharge summaries (or multiple of the provided discharge summaries). However, multi-note coverage is STRICTLY subordinate to dependency quality. - You should aim to link related entities across different notes through clinical relationships, or ask about common/similar entities appearing across different notes, provided it reflects how a medical expert would practically review the charts. - However, you DO NOT need to cover every single provided note if doing so forces an unnatural question. - Never introduce a cross-note question solely for the sake of coverage if it breaks the reasoning chain, creates a disjointed narrative, or requires violating the coreference rules. A highly dependent, clinically coherent chain that only covers 2 out of 3 notes is vastly preferred over a chain that covers all 3 notes but loses logical continuity. Turn Count Guidance: Generate questions calibrated to the patientâs clinical complexity and the number of discharge summaries: - 1 note: 5-7 content questions - 2 notes: 7-9 content questions - 3 or more notes: 8-12 content questions Dependency quality takes absolute priority over turn quantity. Every question must maintain genuine clinical reasoning dependency on prior answers. Do not generate additional questions merely to increase turn count. A strong 6-turn chain with tight dependency is far better than a 12-turn chain where later turns are self-contained or only loosely connected. 73 If you cannot construct a clinically natural, dependent question that maintains the reasoning chain, stop the chain. Do not generate unnatural questions that clinicians would not realistically ask. WHY/HOW Question Guidance: You are encouraged to include WHY and HOW questions at Q3 and later positions. These questions strengthen both clinical naturalness and error propagation because they require understanding clinical reasoning chains, which forces reliance on prior-turn context. Use phrasing such as: - "Because [referencing prior answer context], how was the management adjusted?" - "Given [referenced finding from prior answer], why was that approach chosen?" - "What clinical reasoning led to the change in [the treatment / the intervention]?" - "How did [those findings / that response] influence the subsequent management?" Answerable-from-notes constraint (CRITICAL): All WHY/HOW questions MUST be answerable from information explicitly stated or directly inferable from the discharge summaries combined with standard clinical knowledge to interpret those facts. Permitted reasoning types: - Connecting two documented events in a causal or temporal sequence (e.g., a documented finding followed by a documented clinical decision) - Interpreting documented lab values, imaging results, or exam findings using standard clinical knowledge - Explaining a documented clinical decision using the rationale stated or clearly implied in the notes - Linking documented information across different notes for the same patient PROHIBITED reasoning types (questions requiring these are NOT allowed): - Pathophysiological mechanisms not described in the notes (e.g., "Why does Drug X cause side effect Y at the molecular level?") - Clinical guideline reasoning not cited or implied in the notes (e.g., "Why is Drug X preferred over Drug Y per guidelines?") - Pharmacological mechanisms of action (e.g., "How does this medication work?") - Counterfactual scenarios (e.g., "What would have happened if treatment X had not been given?") - Epidemiological or population-level statistics (e.g., "Why is this condition more common in population X?") - Comparison with alternative treatments not mentioned in the notes (e.g., "Why was ticagrelor chosen over prasugrel?" when the notes do not discuss this choice) The output format must be as follows: Q1: medical expert question regarding the patientâs discharge summaries A1: [Entities, Attributes, Relationships] Entity1 - attribute_name: attribute, attribute_name: attribute Entity2 - attribute_name: attribute, attribute_name: attribute Entity1 -> relationship_name -> Entity2 Entity3 - attribute_name: attribute, attribute_name: attribute Entity2 -> relationship_name -> Entity3 ... [Source Location] Note #xHeader Name corresponding text-span1 corresponding text-span2 Note #xHeader Name corresponding text-span1 corresponding text-span2 ... [Answer] answer_text Repeat for Q2, Q3, etc. Source Location Rules: - You must add all headers that are required for answering the question, but you must add only minimal headers that are necessarily required to answer the question. For example, if the answer content appears in both brief hospital course and major surgical or invasive procedure, and the content in major surgical or invasive procedure is included in brief hospital course, you must only include brief hospital course as source location. - For the header names, you must write the main section header names (e.g., Brief Hospital Course, Major Surgical or Invasive Procedure, History of Present Illness, Discharge Diagnosis, Discharge Medications, Chief Complaint, Discharge Instructions, Pertinent Results, Past Medical History, Medications on Admission, Allergies, Discharge Condition, Physical Exam), instead of the sub-header names (e.g., Microbiology, Imaging within the Pertinent Result header). If the context is included in the sub-header, you must write the main header name instead of the sub- header name, along with its corresponding context. [Entity Category, Attribute, Relationship Schema] 1. Diagnosis -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 74 2. Finding -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 3. Symptom -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 4. Procedure -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 5. Outcome -- status, assertion, time(dates, times, durations, frequencies) 6. Vital_Sign -- status, value, assertion, time(dates, times, durations, frequencies) 7. Physical_Exam -- status, finding, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 8. Lab_Test -- value, abnormal_flag, status, assertion, time(dates, times, durations, frequencies) 9. Diagnostic_Imaging_Test -- status, result, finding_site, laterality, assertion, time(dates, times , durations, frequencies) 10. Specimen -- source, assertion, time(dates, times, durations, frequencies) 11. Microbiology_Test - organism_growth, result, assertion, time(dates, times, durations, frequencies) 12. Microbiology_Organism - growth, assertion, time(dates, times, durations, frequencies) 13. Microbiology_Antibiotics - dilution, sensitivity, assertion, time(dates, times, durations, frequencies) 14. Activity -- status, assertion, time(dates, times, durations, frequencies) 15. Medication -- strength, dosage, form_description, form, route, instruction, disp, refills, assertion, time(dates, times, durations, frequencies) 16. Event -- status, assertion, time(dates, times, durations, frequencies) 17. Allergy -- status, assertion, time(dates, times, durations, frequencies) 18. Medical_Device - status, assertion, time(dates, times, durations, frequencies) 19. Instruction - instruction_text, assertion, time(dates, times, durations, frequencies) is_main attribute is either true or false. If the Diagnosis/Symptom/Finding is the main Diagnosis/ Symptom/Finding for the patient during the current admission (e.g., chief complaint, discharge diagnosis(primary), main Diagnosis/Symptom/Finding according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false (e.g., discharge diagnosis( secondary), not main Diagnosis/Symptom/Finding according to context). If the Procedure is the main Procedure for the patient during the current admission (e.g., major surgical or invasive procedure, main Procedure according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false. abnormal_flag attribute in Lab_Test, is either normal or abnormal. 1. Procedure/Medication -> improves -> Diagnosis/Symptom/Finding 2. Procedure/Medication -> worsens -> Diagnosis/Symptom/Finding 3. Procedure/Medication -> causes -> Diagnosis/Symptom/Finding 4. Procedure/Medication -> administered_for -> Diagnosis/Symptom/Finding 5. Procedure/Medication -> reveals -> Diagnosis/Symptom/Finding 6. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> reveals -> Diagnosis/Finding 7. Specimen -> has_test -> Microbiology_Test/Lab_Test 8. Microbiology_Test -> detected -> Microbiology_Organism 9. Microbiology_Organism -> tested_against -> Microbiology_Antibiotics 10. Microbiology_Organism -> reveals -> Diagnosis/Finding 11. Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding 12. Diagnosis/Symptom/Finding -> causes -> Event 13. Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding/Activity 14. Outcome/Finding -> after -> Diagnosis/Symptom/Finding 15. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> resulted_in -> Procedure/Medication 16. Procedure/Medication -> causes -> Procedure/Medication 17. Procedure/Medication -> causes -> Event 18. Procedure/Medication -> switched_to -> Procedure/Medication 19. Procedure/Medication -> resulted_in -> Outcome/Finding 20. Outcome/Finding -> after -> Procedure/Medication 21. Event -> causes -> Diagnosis/Symptom/Finding 22. Event -> causes -> Procedure/Medication 23. (any category) -> has_instruction -> Instruction The difference between "causes" and "resulted_in" in "Procedure/Medication -> causes -> Diagnosis/ Symptom/Finding" and "Procedure/Medication -> resulted_in -> Outcome/Finding" is that "causes" is used when a procedure or medication produces an adverse effect or complication, while " resulted_in" is used when a procedure or medication leads to a desired or expected state or improvement (positive or neutral). Similarly, the difference between "causes" and "resulted_in" in "Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding" and "Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding " is that "causes" is used when a diagnosis, symptom, or finding produces an adverse effect or complication, while "resulted_in" is used when a diagnosis, symptom, or finding leads to a desired or expected state or improvement (positive or neutral). [Multi-turn QA Templates - qa_template_category] qa_template_examples [Multi-turn QA Templates - Other Categories] qa_template_examples_other [Patientâs Discharge Summaries] discharge_summaries 75 I.2 Step 2: Incorrect Answer Distractor Direction Brainstorming Placeholders. discharge_summaries â the patientâs full sequence of discharge summaries; multiturn_qa_step1 â the complete Step 1 output for the patient (the multi-turn questions with their answer evidence and source locations). You are provided with a patientâs discharge summaries. You are also provided with an entity category , attribute, relationship schema defined for tagging discharge summaries, along with a multi- turn QA pair that consists of 1) medical expert multi-turn questions regarding the patientâs discharge summaries, 2) the entities, attributes, and relationships between entities extracted from the discharge summaries required to answer each corresponding question, 3) the source location of the extracted entities, attributes, and relationships that compose the answer content, and 4) the correct answer. Carefully read the provided patientâs discharge summaries, and the multi-turn questions, answer evidence, answer source, and correct answer. For each multi-turn question, generate specific guidance on content that could support plausible incorrect answer choices. Do not generate multiple-choice options. Do not generate final distractor answer choices. Instead, generate specific distractor-design explanations describing how plausible incorrect answers could be constructed based on the patientâs discharge summaries, clinical knowledge, and multi -turn context. Primary Directive --- Mixed Distractor Strategy: Distractors must require genuine clinical reasoning to eliminate, not just text verification against the notes. Use a mix of the following two sources: Source A --- Note-Grounded With Clinical Reasoning Requirement (approximately 60% of distractors): Draw content from the patientâs discharge summaries, but combine note-grounded details in ways that require clinical knowledge to reject. Examples: - Plausible but clinically incorrect causal chains constructed from entities in the notes - Wrong drug-indication pairings, wrong organism-antibiotic pairings, or wrong diagnosis-treatment pairings using entities that appear in the notes - Clinically implausible complication pathways using conditions and procedures from the notes - Misattributed temporal sequences where the individual events are all documented but the ordering, causal direction, or admission assignment is wrong The key test: a model should NOT be able to eliminate the distractor simply by searching the notes for the presence or absence of a named entity. The distractorâs entities should all appear in the notes; what makes it wrong is the clinical relationship between them, or the admission context in which they appear. Source B --- Clinically Plausible External (approximately 40% of distractors): Introduce entities, diagnoses, treatments, findings, management approaches, or clinical pathways that are NOT present anywhere in the patientâs notes but are clinically plausible given the patientâs conditions. These should represent: - Genuine differential diagnoses, alternative etiologies, or competing interpretations of the patientâs presentation - Alternative treatments, antibiotics, or interventions from the same drug class or guideline pathway that were not used - Management actions, monitoring strategies, or follow-up services that are standard of care for the patientâs conditions but were not arranged or documented - Complications, findings, or clinical outcomes expected in similar clinical situations but not documented in this patient - Alternative discharge dispositions, follow-up plans, or rehabilitation modalities appropriate for the patientâs condition Important constraint for Source B distractors: The external content must not contradict any explicit statement in the notes. It must be absent from the notes, not denied by them. This forces the model to rely on parametric clinical knowledge to reject it, not note-reading. Key principle: A model should NOT be able to eliminate a distractor simply by searching the notes for the presence or absence of a specific entity. Eliminating distractors should require knowing whether a clinical relationship is medically valid, whether a treatment is appropriate for a given condition, whether an organism is likely in a given infection context, or whether a finding is expected in a given clinical context. Shared-Entity Distractor Preference: When generating distractor directions, prioritize designs where the distractor references the SAME key entities as the correct answer but applies a different clinical relationship, interpretation, causal direction, temporal assignment, or admission context. Distractor directions that substitute entirely different entities are less effective because they allow elimination by entity-presence checking. For each Q2+ question, at least 3 of the 6 distractor type directions should describe distractors that share the primary entity with the correct answer. Minimal-Pair Construction Principle: 76 At least 2 of the distractor directions you provide for each question must describe a "minimal-pair" distractor --- one that shares the vast majority of content, structure, and wording with the correct answer but differs in exactly ONE specific clinical detail. The single differing detail should be: - A single entity substitution (e.g., one organism name, one drug name, one diagnosis name) - A single attribute change (e.g., one dosage value, one lab value, one date, one laterality, one severity grade) - A single relationship reversal (e.g., cause vs. effect, temporal ordering of two events) - A single admission reassignment (for multi-admission cases) The goal is that a reader comparing the minimal-pair distractor to the correct answer side-by-side would need to identify precisely which ONE detail is wrong. Distractors that differ from the correct answer in multiple independent dimensions simultaneously are less effective because models can detect the "odd one out" by surface pattern matching. For Types 1, 4, 5, and 6 especially, default to minimal-pair construction unless the questionâs structure makes it impossible (e.g., broad multi-dimensional summary questions). When describing the distractor direction, explicitly state: (a) what content is SHARED with the correct answer, and (b) what SINGLE detail is changed and why it is clinically plausible. Important --- the single changed detail must be VERIFIABLY wrong from the discharge summaries or from established clinical knowledge. Do not change a detail to a value that is merely undocumented --- change it to a value that contradicts documented information or is clinically implausible given the documented context. Core Task Requirements: For each question in the multi-turn QA chain: 1. Identify the referenced entity or concept (applies to Q2 and later questions) For Q2 and later questions, identify: - which prior entity, concept, interpretation, or clinical finding the current question references - what plausible alternative interpretations of the referenced entity or concept could lead to a coherent but incorrect answer to the current question For Q1, there is no prior answer to reference, so proceed directly to step 2. 2. For Q2 and later questions only: Referential Ambiguity Analysis Explicitly identify what the pronoun or referent could alternatively resolve to if the model lacks prior context. For example: - If Q2 asks "What condition did those findings support?" and the patient has multiple documented abnormalities (stroke signs, syncope signs, metabolic abnormalities), identify each possible referent. - Design at least one distractor that represents the correct answer for an alternative pronoun resolution (the clinically correct treatment or finding that matches a different possible referent). Process: a) Identify the referential expression in the question (e.g., "that diagnosis", "the treatment", " those abnormalities", "those episodes", "those medications", "that infection"). b) Identify the correct referent from the prior answer. c) Identify 1-2 alternative referents that are also plausible given the discharge summaries (other documented conditions, symptoms, medications, or findings that the pronoun could resolve to in the absence of prior context). d) For each alternative referent, describe the distractor: the clinically correct answer to the current question IF the pronoun resolved to that alternative referent. Note for questions without explicit referential pronouns: If the current question does not use a referential pronoun but still builds on prior context (e.g., "What was the discharge disposition?" following a question about clinical status), identify what alternative prior answers would lead to a different current answer, and design distractor directions accordingly. 3. For Q2 and later questions only: Cross-Turn Error Path Analysis Identify the 2 most clinically plausible incorrect answers from Q(n-1). For EACH, design a distractor at Q(n) that represents the natural clinical consequence if that prior wrong answer were true. Both error-path distractors must be carried forward as high-priority distractor directions into Step 3. This creates "error-consistent paths": a model that committed a specific error at Q(n-1) encounters a tempting wrong answer at Q(n) that is coherent with its prior error. For each Q(n) where n >= 2, provide: - Prior error path 1: [specific wrong answer direction from Q(n-1)] Trap distractor: [what a model would answer at Q(n) if it believed that wrong answer] Clinical reasoning: Why this downstream answer naturally follows from the prior error - Prior error path 2: [specific wrong answer direction from Q(n-1)] Trap distractor: [what a model would answer at Q(n) if it believed that wrong answer] Clinical reasoning: Why this downstream answer naturally follows from the prior error 77 If only 1 plausible prior error exists (e.g., Q1 answer is unambiguous), explicitly note this and provide 1 error-path distractor. For multi-admission summary questions (where Q1 describes patterns across multiple admissions), the error path should target the most clinically consequential single-admission error. Note: Error-path design is an additional constraint that applies on top of the 6 distractor types. Any distractor type can be used to construct an error-path trap. 4. Generate distractor-design guidance For each distractor type listed below, describe how a plausible incorrect answer could be constructed using: - Note-grounded details combined in ways requiring clinical knowledge to reject (Source A) - Clinically plausible alternatives not in the notes (Source B) - The multi-turn QA context, including referential ambiguity and cross-turn error paths Distractor Coherence Requirement: The QA chain is structured so that each question references entities or clinical concepts introduced in the prior answer. Incorrect answer choices should reflect plausible alternative interpretations of the referenced entity or concept, using a mix of note-grounded (Source A) and clinically-plausible-external ( Source B) content as described in the Primary Directive. Distractor Design Requirements: Your distractor-design explanations must describe how plausible incorrect answers could be constructed. These incorrect answers should be: - clinically plausible - contextually appropriate - difficult for clinical LLMs to distinguish from the correct answer - credible enough that a medical expert reading quickly might consider them possible - free of obvious errors, implausible scenarios, or giveaway wording - NOT eliminable by simple text-matching against the notes (this is critical) Types of Effective Distractors: For each type below, provide multiple specific distinct distractor directions grounded in the case. For each direction, explicitly state whether it is Source A (note-grounded with clinical reasoning requirement) or Source B (clinically plausible external). Type 2 vs Type 3 distinction: Type 2 proposes a competing explanation or interpretation --- what the diagnosis is, what caused it, what the condition is, or what the drug class is for. Type 3 proposes a competing action --- what was done about it, what could have been ordered, what agent was used to achieve the same goal. For medication category specifically: Type 2 would propose an alternative drug class or therapeutic strategy based on a competing interpretation of the clinical indication (e.g., "long-term warfarin for a suspected hypercoagulable state" vs . Apixaban for AF --- competing interpretation of WHY anticoagulation is needed). Type 3 would propose a different agent within the same class already identified as appropriate (e.g., " rivaroxaban instead of Apixaban for the same AF indication" --- competing action within an accepted clinical strategy). 1) Note-grounded entity substitution with clinical reasoning requirement (Source A): Replace the correct entity, or one or more of its attributes (timing, dosage, laterality, severity, site, frequency, duration), with a related entity or attribute value drawn from the notes. The substitution must require clinical knowledge to reject --- the substitute entity or attribute value must appear somewhere in the notes; what makes it wrong is the clinical relationship to the questionâs focus, not entity absence. Prefer single-attribute substitution (changing only ONE entity or ONE attribute value) over multi- attribute substitution unless a multi-attribute change is the only way to produce a clinically plausible distractor. When describing a Type 1 distractor direction, clearly state whether the substitution is single-attribute or multi-attribute and why. Numerical precision priority: When the correct answer contains specific numerical values (lab results, vital signs, medication dosages, dates, durations, frequencies, measurement values, staging codes), at least one Type 1 distractor direction should describe substituting a numerically close but clinically distinct value from elsewhere in the notes. Examples: a creatinine value from a different lab draw or different admission, a dosage for the same drug at a different time point, a date from a different admission or clinical event, a lab value from a different specimen or test, a staging component from a related but different assessment. The numerical substitution should be close enough that it requires knowing the exact documented value to reject --- not an order-of-magnitude difference. How this applies across categories: - For diagnosis/clinical_assessment questions: substitute a different diagnosis, finding, or assessment result from the notes - For medication questions: substitute a different drug from the patientâs medication list that has a different clinical indication - For microbiology questions: substitute a different organism, specimen source, or susceptibility result from the notes - For procedure questions: substitute a different procedure documented in the notes - For symptom questions: substitute a different symptom or trigger from the patientâs history 78 - For disch_plan questions: substitute a different discharge destination, service, or follow-up from the notes - For clinical_outcome questions: substitute a different outcome status, resolved problem, or functional level from the notes 2) Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B): Propose an alternative interpretation, diagnosis, etiology, or approach that is NOT in the notes but is clinically plausible given the patientâs presentation. This type represents a competing EXPLANATION for the clinical picture --- what the diagnosis is, what caused it, what the condition is, or what the drug class strategy is for. How this applies across categories: - For diagnosis/clinical_assessment: alternative diagnosis or differential for the documented presentation - For medication: alternative drug class or alternative clinical interpretation of why the anticoagulant, antihypertensive, or antibiotic strategy is needed --- not just an alternative agent within the same class (see Type 3) - For microbiology: alternative organism commonly seen in this infection type, or alternative infection classification - For procedure: alternative surgical or diagnostic approach - For symptom: alternative etiology, trigger, or precipitating mechanism - For disch_plan: alternative post-discharge care model, referral pathway, or care coordination approach (e.g., alternative level of care, different specialty managing the post-discharge course) - For clinical_outcome: alternative outcome trajectory plausible for the documented conditions (e.g ., alternative functional recovery trajectory or alternative resolution pathway for the primary problem) 3) Plausible-but-absent management action, intervention, or monitoring strategy (Source B): Propose an action, order, intervention, or arrangement that is NOT documented but is clinically standard or guideline-appropriate for the patientâs conditions. This type represents a competing ACTION --- what was done about it, what could have been ordered, what agent was used to achieve the same goal. Important --- question-type dependency: Type 3 is most naturally applicable to questions that ask what was DONE, ORDERED, ARRANGED, or CHOSEN (e.g., "What treatment was started?", "What follow- up was arranged?", "Which antibiotic was selected?"). For questions that ask what clinical finding, diagnosis, outcome, or symptom WAS PRESENT or WAS FOUND (e.g., "What abnormal findings were identified?", "What was the discharge status?", "What symptom occurred?"), skip Type 3 and provide an additional Type 1, Type 4, or Type 5 distractor instead. Do not force a Type 3 distractor onto non-action questions. How this applies across categories: - For medication: alternative drug from same class that was not prescribed, or alternative dosing schedule for the identified drug - For microbiology: alternative antibiotic spectrum or empiric coverage that was not chosen, or alternative stewardship approach - For procedure: alternative operative approach or technique that was not performed - For diagnosis/clinical_assessment: alternative imaging modality or diagnostic test that was not ordered, presented as a competing answer to what workup was completed - For disch_plan: alternative outpatient service, rehabilitation modality, or follow-up arrangement not documented - For symptom: alternative trigger management approach, activity restriction, or lifestyle modification not documented - For clinical_outcome: alternative post-discharge monitoring plan or readmission prevention strategy not documented 4) Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A or B): Construct a distractor that misassigns a clinical relationship --- either reversing a causal direction, misattributing an association, displacing a temporal sequence, or confusing events across admissions --- where rejecting the misattribution requires knowing the typical clinical course or the logical structure of the case. How this applies across categories: - For diagnosis: attribute a diagnosis to the wrong etiology or reverse the cause-effect chain between comorbidities - For medication: attribute a drug to the wrong indication where both indications are plausible given the patientâs conditions - For microbiology: misattribute an organism to the wrong specimen source or infection site, or misassign an antibiotic sensitivity - For procedure: attribute a procedure to the wrong indication or misplace it in the temporal sequence of care - For symptom: misattribute a symptom to the wrong trigger, reverse the temporal relationship between symptom and intervention, or confuse symptoms from different admissions - For disch_plan: misattribute a discharge destination or functional status to the wrong clinical reason or wrong admission - For clinical_outcome: misattribute an outcome to the wrong clinical problem or intervention - For clinical_assessment: misattribute a clinical finding to the wrong underlying condition 5) Note-grounded detail with one external wrong element (Source A+B hybrid): Combine several correct clinical details from the notes with one incorrect element drawn from general medical knowledge. The correct elements make the distractor credible; the incorrect element requires 79 clinical knowledge to identify. The wrong element must not contradict any explicit statement in the notes. Numerical detail variant: When the correct answer includes numerical clinical data, one Type 5 direction should describe keeping all non-numerical content identical to the correct answer and changing only the numerical value to a clinically plausible alternative. This includes: a lab value within an adjacent reference range category, a dosage that exists for the same drug but for a different indication, a date shifted by one admission or clinical event, a measurement value from a similar but different anatomical site. This creates a distractor where the clinical reasoning is correct but the precise quantitative detail is wrong. How this applies across categories: - For diagnosis: correct diagnosis name and presentation with wrong severity, stage, or associated finding - For medication: correct drug and route with wrong dosage, frequency, or duration - For microbiology: correct organism identification with wrong susceptibility pattern or wrong gram stain characteristic - For procedure: correct procedure description with wrong complication, wrong approach detail, or wrong laterality rationale - For symptom: correct symptom description with wrong associated sign, wrong temporal pattern, or wrong precipitant - For disch_plan: correct discharge destination with wrong level of services or wrong reason for that level of care - For clinical_outcome: correct outcome status with wrong clinical attribution or wrong timeline - For clinical_assessment: correct assessment findings with wrong clinical interpretation or wrong associated condition 6) Cross-admission temporal displacement (Source A): For patients with multiple admissions, attribute a correct clinical fact to the wrong admission. The entities, relationships, and clinical content are all present in the notes; what makes the distractor wrong is that the detail belongs to a different admission than the one the question asks about. How this applies across categories: - Attributing the discharge destination from one admission to a different admission - Attributing the medication hold reason from one hospitalization to a different one - Attributing the complication that occurred postoperatively in one admission to a different surgical admission - Attributing a specific organism or culture result to the wrong admission This type is particularly important for multi-admission summary questions where models can construct plausible answers by assigning correct facts to incorrect admission positions. Important --- applicability condition: Type 6 applies only when (a) the patientâs discharge summaries document 2 or more distinct admissions AND (b) the current question references clinical facts that could plausibly be assigned to different admissions. For single-admission patients or questions with unambiguous temporal context, mark Type 6 as not applicable and provide an additional Type 1, Type 4, or Type 5 distractor instead. Distractor Construction Guidance: For each distractor type, write: 1. Source: State whether this is Source A (note-grounded with clinical reasoning requirement), Source B (clinically plausible external), or Source A+B hybrid. 2. Explanation: A concise, specific description of how an incorrect answer choice could be constructed using that distractor type. 3. Reasoning: A brief justification for why this distractor is plausible and why it cannot be eliminated by text matching alone. This must cite: - For Source A distractors: the specific note content that grounds the distractor and what clinical knowledge is needed to reject it. - For Source B distractors: the clinical basis and confirmation that this content does not contradict any explicit note statement. - For error-path distractors: what prior wrong answer this follows from and why the downstream reasoning is clinically coherent. 4. Minimal-pair label: If this direction is a minimal-pair distractor (differs from correct answer in exactly one clinical detail), label it: "[Minimal-pair: differs only in specific detail]". If it differs in multiple dimensions, label it: "[Multi-difference]". Do not write generic distractor advice. Make each explanation and reasoning specific to the content of the provided patientâs discharge summaries. Output Format: For each question, produce the following output format exactly: Q1: medical expert question - Effective Distractor Type 1: Note-grounded entity substitution with clinical reasoning requirement (Source A) Explanation: specific description of how to construct the distractor Reasoning: why this distractor is plausible, citing note content and clinical knowledge required Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] - Effective Distractor Type 2: Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B) Explanation: specific description of how to construct the distractor Reasoning: why this distractor is plausible, citing clinical basis and confirming no note contradiction Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] 80 - Effective Distractor Type 3: Plausible-but-absent management action, intervention, or monitoring strategy (Source B) Explanation: specific description of how to construct the distractor. If the question asks what was FOUND/PRESENT rather than what was DONE, write "Not applicable for this question type --- substituting additional Type [1/4/5] direction:" followed by the substitute distractor direction. Reasoning: why this distractor is plausible, citing clinical basis Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] - Effective Distractor Type 4: Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A or B) Explanation: specific description of how to construct the distractor Reasoning: why this distractor is plausible, citing note content or clinical basis Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] - Effective Distractor Type 5: Note-grounded detail with one external wrong element (Source A+B hybrid) Explanation: specific description of how to construct the distractor Reasoning: which elements are correct from the notes and what external medical knowledge is needed to identify the wrong element Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] - Effective Distractor Type 6: Cross-admission temporal displacement (Source A) Explanation: specific description of how to construct the distractor by assigning a correct clinical fact to the wrong admission, or âNot applicable: single admission or unambiguous temporal context --- substituting additional Type [1/4/5] direction:â followed by the substitute distractor direction Reasoning: which admission the fact belongs to and why a model might plausibly assign it to the wrong one; or brief explanation if not applicable Minimal-pair label: [Minimal-pair: differs only in detail] or [Multi-difference] Minimal-pair validation: After generating all 6 distractor directions, verify that at least 2 are labeled [Minimal-pair]. If fewer than 2 are minimal-pair, revise one of the [Multi-difference] distractors to a minimal-pair version by reducing it to a single differing dimension. Prior-Context Dependency Validation (Q2+ only): After generating all distractor directions, verify: Would a model reading only the current question, the discharge summaries, and the eventual answer choices (without any prior-turn context) be able to identify the correct answer? If all distractor directions target entities or facts clearly distinguishable from the correct answer using only the discharge summaries, the directions are insufficient. Ensure that at least 2 distractor directions describe answers that would be correct IF the questionâs referent resolved differently --- i.e., they are correct answers to plausible alternative interpretations of what the question is asking about. Q2: medical expert question - Referential Ambiguity Analysis: Referential expression: the pronoun or anaphora in the question, or the implicit prior-context dependency Correct referent: from prior answer Alternative referent 1: plausible alternative from notes Alternative referent 2: plausible alternative from notes, if applicable Distractor direction for alternative referent 1: correct answer if pronoun resolved to this referent Distractor direction for alternative referent 2: correct answer if pronoun resolved to this referent, if applicable - Cross-Turn Error Path Analysis: Prior error path 1: specific wrong answer direction from Q1 Trap distractor: what a model would answer if it believed this wrong Q1 answer Clinical reasoning: why this follows from the prior error Prior error path 2: specific wrong answer direction from Q1 Trap distractor: what a model would answer if it believed this wrong Q1 answer Clinical reasoning: why this follows from the prior error - Effective Distractor Type 1: Note-grounded entity substitution with clinical reasoning requirement (Source A) Explanation: specific description Reasoning: specific justification Minimal-pair label: label - Effective Distractor Type 2: Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B) Explanation: specific description Reasoning: specific justification Minimal-pair label: label - Effective Distractor Type 3: Plausible-but-absent management action, intervention, or monitoring strategy (Source B) Explanation: specific description, or note non-applicability and provide substitute direction Reasoning: specific justification Minimal-pair label: label - Effective Distractor Type 4: Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A or B) Explanation: specific description Reasoning: specific justification 81 Minimal-pair label: label - Effective Distractor Type 5: Note-grounded detail with one external wrong element (Source A+B hybrid) Explanation: specific description Reasoning: specific justification Minimal-pair label: label - Effective Distractor Type 6: Cross-admission temporal displacement (Source A) Explanation: specific description, or note non-applicability and provide substitute direction Reasoning: specific justification Minimal-pair label: label Minimal-pair validation: verify at least 2 are [Minimal-pair]; if not, revise Prior-Context Dependency Validation: verify at least 2 distractor directions would be correct answers under alternative referent resolutions; if not, revise Repeat for Q3, Q4, etc. (all Q2+ questions must include the Referential Ambiguity Analysis, Cross- Turn Error Path Analysis, and Prior-Context Dependency Validation sections) [Entity Category, Attribute, Relationship Schema] 1. Diagnosis -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 2. Finding -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 3. Symptom -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 4. Procedure -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 5. Outcome -- status, assertion, time(dates, times, durations, frequencies) 6. Vital_Sign -- status, value, assertion, time(dates, times, durations, frequencies) 7. Physical_Exam -- status, finding, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 8. Lab_Test -- value, abnormal_flag, status, assertion, time(dates, times, durations, frequencies) 9. Diagnostic_Imaging_Test -- status, result, finding_site, laterality, assertion, time(dates, times , durations, frequencies) 10. Specimen -- source, assertion, time(dates, times, durations, frequencies) 11. Microbiology_Test - organism_growth, result, assertion, time(dates, times, durations, frequencies) 12. Microbiology_Organism - growth, assertion, time(dates, times, durations, frequencies) 13. Microbiology_Antibiotics - dilution, sensitivity, assertion, time(dates, times, durations, frequencies) 14. Activity -- status, assertion, time(dates, times, durations, frequencies) 15. Medication -- strength, dosage, form_description, form, route, instruction, disp, refills, assertion, time(dates, times, durations, frequencies) 16. Event -- status, assertion, time(dates, times, durations, frequencies) 17. Allergy -- status, assertion, time(dates, times, durations, frequencies) 18. Medical_Device - status, assertion, time(dates, times, durations, frequencies) 19. Instruction - instruction_text, assertion, time(dates, times, durations, frequencies) is_main attribute is either true or false. If the Diagnosis/Symptom/Finding is the main Diagnosis/ Symptom/Finding for the patient during the current admission (e.g., chief complaint, discharge diagnosis(primary), main Diagnosis/Symptom/Finding according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false (e.g., discharge diagnosis( secondary), not main Diagnosis/Symptom/Finding according to context). If the Procedure is the main Procedure for the patient during the current admission (e.g., major surgical or invasive procedure, main Procedure according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false. abnormal_flag attribute in Lab_Test, is either normal or abnormal. 1. Procedure/Medication -> improves -> Diagnosis/Symptom/Finding 2. Procedure/Medication -> worsens -> Diagnosis/Symptom/Finding 3. Procedure/Medication -> causes -> Diagnosis/Symptom/Finding 4. Procedure/Medication -> administered_for -> Diagnosis/Symptom/Finding 5. Procedure/Medication -> reveals -> Diagnosis/Symptom/Finding 6. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> reveals -> Diagnosis/Finding 7. Specimen -> has_test -> Microbiology_Test/Lab_Test 8. Microbiology_Test -> detected -> Microbiology_Organism 9. Microbiology_Organism -> tested_against -> Microbiology_Antibiotics 10. Microbiology_Organism -> reveals -> Diagnosis/Finding 11. Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding 12. Diagnosis/Symptom/Finding -> causes -> Event 13. Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding/Activity 14. Outcome/Finding -> after -> Diagnosis/Symptom/Finding 15. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> resulted_in -> Procedure/Medication 16. Procedure/Medication -> causes -> Procedure/Medication 17. Procedure/Medication -> causes -> Event 18. Procedure/Medication -> switched_to -> Procedure/Medication 19. Procedure/Medication -> resulted_in -> Outcome/Finding 20. Outcome/Finding -> after -> Procedure/Medication 82 21. Event -> causes -> Diagnosis/Symptom/Finding 22. Event -> causes -> Procedure/Medication 23. (any category) -> has_instruction -> Instruction The difference between "causes" and "resulted_in" in "Procedure/Medication -> causes -> Diagnosis/ Symptom/Finding" and "Procedure/Medication -> resulted_in -> Outcome/Finding" is that "causes" is used when a procedure or medication produces an adverse effect or complication, while " resulted_in" is used when a procedure or medication leads to a desired or expected state or improvement (positive or neutral). Similarly, the difference between "causes" and "resulted_in" in "Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding" and "Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding " is that "causes" is used when a diagnosis, symptom, or finding produces an adverse effect or complication, while "resulted_in" is used when a diagnosis, symptom, or finding leads to a desired or expected state or improvement (positive or neutral). [Patientâs Discharge Summaries] discharge_summaries [Initial Multi-turn QA] multiturn_qa_step1 I.3 Step 3: Content-Question Answer Choice Generation Placeholders. discharge_summaries â the patientâs full sequence of discharge summaries; multiturn_qa_step1_step2 â the Step 1 multi-turn questions with their answer evidence and source locations, combined with the Step 2 distractor directions for each question. You are provided with a patientâs discharge summaries. You are also provided with an entity category , attribute, relationship schema defined for tagging discharge summaries, along with multi-turn QA set that consists of 1) medical expert multi-turn questions regarding the patientâs discharge summaries, 2) the answer evidence, 3) the answer source, 4) content describing how to construct effective distractors. Carefully read the provided patientâs discharge summaries, the medical expert questions, answer evidence, answer source, and effective distractor construction content. For each multi-turn question, generate one correct answer choice along with four plausible incorrect answer choices. The correct answer must be uniquely correct, while all incorrect choices must remain plausible enough to be difficult to distinguish. Primary Directive --- Mixed Distractor Strategy: Incorrect answer choices must use a mix of two distractor sources to ensure they cannot be eliminated by simple text verification against the notes: Source A --- Note-Grounded With Clinical Reasoning Requirement (approximately 60% of incorrect choices across all questions): Draw content from entities, attributes, relationships, clinical interpretations, or factual context that appear in the patientâs discharge summaries. However, combine these note-grounded details in ways that require clinical knowledge to reject. The entities in the distractor should be present in the notes; what makes the distractor wrong is the clinical relationship, interpretation, admission context, or reasoning applied to those entities. Source B --- Clinically Plausible External (approximately 40% of incorrect choices across all questions): Introduce entities, diagnoses, treatments, findings, management approaches, or clinical pathways that are NOT present in the patientâs notes but are clinically plausible given the patientâs conditions. These represent genuine differential diagnoses, alternative guideline-appropriate treatments, standard-of-care actions that were not taken, or complications and findings that commonly occur in similar clinical situations but did not occur in this patient. - Source B distractors must not contradict any explicit statement in the notes. They must be absent, not denied. - A model must use parametric clinical knowledge to reject Source B distractors --- note-reading alone is insufficient. The provided effective distractor construction content is reference material to help identify plausible distractor directions. It does not need to be applied one-to-one to the four incorrect answer choices --- there is no required mapping between distractor types and answer choices. Use it as guidance to generate the four most plausible incorrect answer choices possible. Each incorrect answer choice must specify its source (A or B) in the reasoning field. Important: - Ensure that the structure and tone of ALL FIVE answer choices are similar. IMPORTANT: "similar" refers to clinical register and sentence structure, NOT length. The correct answer must NOT systematically be the longest or most detailed choice. If the correct answer contains more information than the distractors, either trim the correct answer or elaborate the distractors 83 with plausible additional clinical detail. At least 1-2 incorrect choices per question should be equal to or longer than the correct answer. - Avoid using parentheses in the answer choices. Do not include unnecessary deidentified placeholders (e.g., ___) to the correct answer and incorrect answer choices. Core Task Workflow: For each question: Step 1: Determine the correct answer - Refer to the patientâs discharge summaries and answer evidence, answer source to provide the correct answer. - Do not infer or hallucinate information that is not clearly stated in the note. - Do not copy wording directly from the notes, and make sure to paraphrase the statements across the notes required to answer the question. - Use clinically natural language suitable for discharge summary interpretation. - Answer only what the question asks. The answer evidence and answer source are provided to help you identify the correct answer, but they may contain far more information than needed. Extract only the specific fact(s) the question is asking for --- do not reproduce all the details from the evidence. Step 2: Identify the referenced entity (applies to Q2 and later questions) For Q2 and later questions, identify the entity, concept, or interpretation introduced in the previous answer (Ai-1) that the current question references. For Q1, there is no prior answer, so proceed directly to Step 3 using the discharge summaries and answer evidence. Step 3: Model alternative interpretations Based on the discharge summaries and the effective distractor construction content, generate plausible alternative interpretations of the referenced entity or concept. These alternatives must be clinically coherent. Use a mix of: - Note-grounded alternatives (Source A): entities, attributes, relationships, interpretations, and admission-level assignments from the notes that could plausibly be confused with the correct referent - Clinically plausible external alternatives (Source B): diagnoses, treatments, management actions, findings, or clinical pathways not in the notes but medically plausible for this patientâs presentation Step 4: Construct answer choices Generate: - 1 correct answer - 4 incorrect answers that are as plausible and difficult to distinguish from the correct answer as possible Use the effective distractor construction content as reference to guide the construction of incorrect choices, but do not apply it as a one-to-one template. Select whichever distractor approaches and combinations yield the most plausible four incorrect answer choices. Each incorrect answer must: - follow a coherent but incorrect clinical reasoning path - represent a plausible but factually wrong interpretation - NOT be eliminable by simple note-searching for entity presence or absence Step 4b: Length and Overlap Verification Before proceeding to Step 5, verify the following conditions. If any fails, revise before proceeding : (a) Length check: The correct answer is NOT the longest of the 5 choices. If it is, trim the correct answer by removing non-essential qualifiers or secondary details, or expand at least 2 incorrect choices by adding plausible clinical detail. The correct answer should be at or below the median length of all 5 choices. (b) Inverse length check: At least 1 incorrect choice is noticeably longer than the correct answer ( by approximately 15% or more in character count). If not, expand one distractor with additional plausible-but-wrong clinical detail, secondary findings, or temporal context. (c) Minimal-pair check: At least 1 incorrect choice shares the same clinical entities, relationship types, and narrative structure as the correct answer, differing in only one clinical dimension (a single entity name, numerical value, date, laterality, causal direction, or admission attribution). If none qualifies, revise the most divergent distractor to be a minimal-pair version. (d) Numerical check (if applicable): If the correct answer contains specific numerical values (lab results, dosages, dates, measurements, staging codes), at least 1 distractor uses the same narrative structure with a different clinically plausible number. If not, revise one distractor accordingly. If the correct answer contains no numeric content, skip this check. Step 4c: Source-Mix Check Verify that no more than 2 of the 4 incorrect choices are Source B (clinically plausible external). If 3 or more are Source B, convert one to a Source A distractor by using note-grounded content with a wrong clinical relationship. Maintaining a strong Source A majority ensures distractors cannot be eliminated simply by checking whether entities appear in the notes. Step 5 (Q2+ only): Apply Cross-Turn Dependency Enforcement, Self-Resolution Test, and Error- Consistent Distractor Requirement (see below) Distractor Coherence Requirement: 84 The QA chain is structured so that each question references entities or clinical concepts introduced in the prior answer. Incorrect answer choices should reflect plausible alternative interpretations of the referenced entity or concept, using a mix of Source A and Source B content as described in the Primary Directive. Correct Answer Requirements: The correct answer must: - be directly supported by the discharge summaries and evidence - use natural clinical language - not copy exact phrasing from the notes - answer only what the question specifically asks. Do not include background clinical context, related findings, or reasoning steps unless the question directly asks for them. - be concise: include only the key facts required to distinguish the correct answer from the distractors. Avoid over-specifying details that are not necessary to answer the question. - CRITICAL LENGTH RULE --- MANDATORY VERIFICATION: The correct answer must NOT systematically be the longest answer choice. Follow this procedure: 1. Before writing the correct answer, mentally draft an "answer skeleton" --- the minimal statement that directly answers the question and nothing more. Use this skeleton as the ceiling for the correct answer. The correct answer may add at most one qualifying clause beyond the skeleton. 2. After writing all 5 answer choices, verify that the correct answer is at or below the MEDIAN length of all 5 choices. If it is the longest, either: (a) Trim the correct answer by removing qualifiers, temporal context, secondary findings, or enumerations not strictly required to answer the question, OR (b) Expand at least 2 incorrect answers by adding plausible clinical detail, context, or elaboration to make them equal to or longer than the correct answer. 3. INVERSE LENGTH REQUIREMENT: For approximately 1 in 5 questions, make the correct answer the SHORTEST choice, with at least 2 incorrect answers being noticeably more detailed. Achieve this by writing a concise factual correct answer and making 2+ distractors include plausible but unnecessary elaboration (additional clinical context, hedging language, multi-step reasoning chains, or related-but-unrequested findings). 4. Do NOT remove information that the question directly asks about. Only remove elaboration that goes beyond the questionâs scope. The correct answer must still be unambiguous --- it must contain enough information that a clinical expert would agree it is uniquely correct. Incorrect Answer Choice Requirements: Generate exactly four incorrect answer choices that are: - A mix of Source A and Source B distractors as described in the Primary Directive. Across all questions combined, aim for approximately 60% Source A and 40% Source B. For any individual question, at least one incorrect choice must be Source A and at least one must be Source B. - Clinically plausible and contextually appropriate. - Difficult for clinical LLMs to distinguish from the correct answer. - Realistic enough that a medical expert reading quickly might find them credible. - Free of obvious errors, implausible medical scenarios, or giveaway wording. - Matching or EXCEEDING the correct answer in length for at least 1-2 distractors per question. Each distractor must contain equal or greater level of detail compared to the correct answer. It is acceptable and encouraged for some distractors to include additional plausible clinical context, secondary findings, or temporal qualifiers that make them longer than the correct answer. The additional length must be clinically meaningful, not padding. - An incorrect answer choice may mix correct and incorrect information. Combining several correct clinical details with one or more incorrect details can make the answer choice hard to distinguish from the correct answer. Minimal-Pair Distractor Requirement: Among the 4 incorrect answer choices for each question, at least 1 (and ideally 2) must be "minimal- pair" distractors that share the vast majority of their text with the correct answer and differ in exactly ONE specific clinical detail. Construction process: 1. Copy the correct answerâs structure and most of its content. 2. Change exactly one factual element: a single entity name, a single numerical value, a single date , a single laterality, a single causal relationship, or a single admission attribution. 3. Ensure the changed element is clinically plausible --- not absurd. 4. Verify the changed element is VERIFIABLY WRONG from the discharge summaries or established clinical knowledge. The changed detail must contradict documented information, not merely be undocumented. 5. The resulting distractor should have high word overlap with the correct answer. Example --- correct answer: "Blood cultures from the second admission grew Staphylococcus aureus, sensitive to vancomycin and oxacillin." Minimal-pair distractor: "Blood cultures from the second admission grew Staphylococcus aureus, sensitive to vancomycin but resistant to oxacillin." (One detail changed: oxacillin sensitivity flipped to resistance.) Non-minimal-pair to AVOID: "Wound cultures from the first admission grew coagulase-negative Staphylococcus, resistant to multiple antibiotics." 85 (Multiple details changed: specimen type, admission, organism, susceptibility --- too easy to distinguish by surface pattern.) Exception: For broad summary questions requiring multi-dimensional answers, true single-attribute minimal pairs may be impossible. In these cases, minimize the number of differing dimensions to 2 and note the exception in the reasoning field. Numerical Precision Requirement: When the correct answer contains specific numerical values (lab values, vital signs, medication dosages, dates, durations, measurement values, staging codes), at least one incorrect answer choice must test numerical precision by: - Using the same narrative structure and clinical reasoning as the correct answer - Changing only the numerical value(s) to a clinically plausible alternative (e.g., a value from a different date in the patientâs data, a dosage for a different indication, a measurement from a different time point) - The alternative number should be close enough to be confusable --- not an order of magnitude different Source preference: Use a real numeric value from elsewhere in the patientâs notes (Source A) when possible. Confirm the distractor number does not appear in the notes in the SAME clinical context as the question asks about --- otherwise it may be a duplicate correct answer. If the correct answer contains no numeric content, this requirement is waived. Cross-Turn Dependency Enforcement (required for Q2+ questions): For Q2 and all subsequent questions, apply the following techniques to ensure answer choices cannot be resolved without prior turn context: Technique 1 --- Shared-Entity Choices: At least 3 of the 4 incorrect choices must reference the same key entities or clinical elements from the notes as the correct answer, but differ in the RELATIONSHIP, INTERPRETATION, REASON, or ADMISSION CONTEXT applied to those entities. Ideally, all 5 answer choices reference the same primary entity. The goal is that no choice can be eliminated by checking whether a specific entity name appears in the notes. Do NOT make each choice name a completely different drug-condition pair, organism, or destination. Example of what to AVOID (each choice names different entities --- self-resolving by note search): - A: "Heparin infusion for stent thrombosis" - B: "Nitroglycerin drip for unstable angina" - C: "Metoprolol for rate control" - D: "Clopidogrel loading for new stent placement" Example of what to DO (shared entity type, different clinical interpretation --- requires prior context): - A: "Anticoagulation bridging was initiated because imaging confirmed acute thrombus formation in the stent" - B: "Anticoagulation bridging was initiated as routine post-procedural prophylaxis per institutional protocol" - C: "Anticoagulation bridging was initiated because of new-onset atrial fibrillation detected on telemetry" - D: "Anticoagulation was not initiated because the initial imaging was reassuring and symptoms had resolved" Technique 2 --- Pronoun-Preserving Choices: For questions using referential pronouns ("that condition", "those findings", "the treatment", "those episodes", "those medications", "that infection"), at least 2 incorrect choices must be plausible answers for DIFFERENT resolutions of the pronoun. If "that condition" could refer to Condition X (correct, from prior answer) or Condition Y (also documented in the notes), one distractor should be the clinically correct answer assuming the pronoun refers to Condition Y. The model must resolve the pronoun using prior turn context before selecting. Technique 3 --- Reduced Entity Specificity: Where possible, describe clinical reasoning, therapeutic approach, mechanism, or outcome rationale rather than listing specific drug, organism, or diagnosis names as the primary distinguishing feature between choices. Instead of choices that each name different specific drugs (where entity names alone resolve the question), write choices that each describe a different clinical reasoning path using the same general treatment category. Applicability: Apply Technique 3 when the question asks about clinical reasoning, mechanism, therapeutic rationale, or causal explanation --- where reasoning paths, not entity names, are the distinguishing feature. Do NOT apply Technique 3 when the question explicitly asks for a specific named medication, organism, diagnosis, or procedure --- where naming specificity is the clinically relevant answer dimension. For such questions, use Techniques 1 and 2 instead. WHY/HOW Answer-Choice Guidance: When the question asks WHY a clinical decision was made or HOW a clinical course unfolded, all 5 answer choices should describe the same clinical action or outcome but differ in the causal reasoning, mechanism, or contributing factors. For WHY questions, all choices should share the form "Because [reasoning], [action was taken]" --- differing only in the reasoning. For HOW questions, all choices should describe the same general process but differ in the sequence, 86 contributing factors, or mechanism. This naturally prevents entity-lookup elimination and strengthens dependency on prior context. Self-Resolution Test (required for Q2+ questions): Before finalizing the answer choices for each Q2+ question, mentally apply this test: If a model with NO chat history reads ONLY the current question, the five answer choices, and the discharge summaries, could it identify the correct answer by checking each choice against the notes for consistency? If yes --- meaning only one choice is note-consistent while the others have detectable note-inconsistencies --- then redesign the choices. Remediation steps if the self-resolution test fails: 1. Check whether incorrect choices reference entities not in the notes (making them trivially eliminable). If so, revise those choices to use entities FROM the notes with wrong clinical relationships. 2. Check whether the correct answer is the only one whose clinical relationship matches the notes. If so, create additional incorrect choices where the relationship is also note-consistent but applies to a different referent (one the pronoun could plausibly resolve to without prior context). 3. After revision, verify that at least 2 incorrect choices are each internally consistent with the notes when read in isolation, and are only distinguishable from the correct answer by knowing which specific entity, interpretation, or admission context was established in the prior turnâs answer. Error-Consistent Distractor Requirement (required for Q2+ questions): Among the four incorrect answer choices for each Q2+ question, at least TWO must be "error- consistent" distractors --- answers that would be correct (or strongly defensible) if the model had selected a specific incorrect answer at the previous turn. The 2 error-consistent distractors should derive from DIFFERENT prior-turn errors, not the same one. This ensures the error-propagation trap catches models regardless of which specific error they made at the prior turn. Construction process: 1. From the step 2 distractor guidance, identify the cross-turn error paths provided. 2. For each of the 2 error paths, write one Q(n) incorrect choice as the clinically correct answer to Q(n) GIVEN that wrong Q(n-1) answer. 3. If the Step 2 guidance provides only one error path, construct one error-consistent distractor from it and construct a second by identifying an additional plausible Q(n-1) wrong answer yourself. When constructing the error-consistent distractors, give priority to minimal-pair implementations. Error-path traps that differ from the correct answer in a subtle, single-dimension way (one causal direction, one admission attribution) are maximally effective. In the reasoning field for error-consistent distractors, mark: "Error-path: follows from Q[n-1] wrong answer â[brief description]â. If that were correct, this choice would be correct because [clinical reasoning]." Explanation of Types of Effective Distractors: The following distractor types describe approaches for constructing plausible incorrect answer choices. You are not required to apply each type to one incorrect answer choice --- use them as needed to generate the four most plausible incorrect answer choices possible. 1) Note-grounded entity substitution with clinical reasoning requirement (Source A): Replace the correct entity or one of its attributes with a related entity or attribute value drawn from the notes, ensuring the substitution requires clinical knowledge to reject. Both the substitute entity and the attribute value must appear in the notes; what is wrong is the clinical relationship between the substituted entity and the questionâs focus. Prefer single-attribute substitution over multi-attribute substitution. This applies to any entity type: diagnoses, drugs, organisms, procedures, symptoms, dispositions, outcomes, or findings. 2) Clinically plausible alternative approach, standard-of-care variant, or parallel pathway (Source B): Propose an alternative interpretation, diagnosis, etiology, or approach NOT in the notes but clinically plausible for this patient. This type represents a competing explanation for the clinical picture, not a competing action (see Type 3). The model must use parametric medical knowledge to determine this was not the documented interpretation or approach. 3) Plausible-but-absent management action, intervention, or monitoring strategy (Source B): Propose a management action, order, intervention, or arrangement that is NOT documented but is clinically standard or guideline-appropriate for the patientâs conditions. This type represents a competing action that was not taken (vs Type 2 which represents a competing explanation). The model must know which specific action was actually taken to reject this distractor. Note: For questions asking what finding, diagnosis, outcome, or symptom was present, Type 3 distractors are less applicable. If a Type 3 distractor would not plausibly compete as an answer to this specific question, substitute an additional Type 1, 4, or 5 distractor instead. 4) Clinical relationship misattribution: causal, temporal, associative, or cross-admission (Source A or B): Misassign a clinical relationship --- reversing a causal direction, misattributing an associative link, displacing a temporal sequence, or confusing events across admissions --- 87 where rejecting the misattribution requires knowing the typical clinical course or the logical structure of the case. 5) Note-grounded detail with one external wrong element (Source A+B hybrid): Combine several correct clinical details from the notes with one incorrect element from general medical knowledge. The correct elements create credibility; the incorrect element requires clinical knowledge to identify. The wrong element must not contradict any explicit note statement. 6) Cross-admission temporal displacement (Source A): Attribute a correct clinical fact to the wrong admission. The entities and relationships are note-grounded and correct; what makes the distractor wrong is the admission context in which they are placed. Particularly important for multi-admission summary questions. Apply only when the patient has 2 or more admissions and the current question references facts that could plausibly be assigned to different admissions. For single-admission questions with no temporal ambiguity, skip Type 6 and use additional Type 1, 4, or 5 distractors instead. Reasoning Requirement: For each answer choice, provide a brief explanation describing: - why the answer choice is correct, and - why each incorrect answer choice is incorrect. For incorrect answer explanations: - For Source A distractors (note-grounded): include a short text span from the discharge summary AND explain what clinical knowledge is needed to distinguish this from the correct answer. - For Source B distractors (clinically plausible external): explain what clinical knowledge is needed to reject the choice. Do NOT cite note absence alone as the reason a distractor is wrong . - For error-consistent distractors: include the error-path notation as described above. - For minimal-pair distractors: explicitly state what single detail differs from the correct answer and label it: "[Minimal-pair: differs in detail]" - State the source type (A, B, or A+B hybrid) for each incorrect choice. The output format must be as follows: Q1: medical expert question A1: correct answer choice A1-Reasoning: brief explanation why correct I1: incorrect answer choice 1 I1-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject]. [ Minimal-pair label if applicable] I2: incorrect answer choice 2 I2-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] I3: incorrect answer choice 3 I3-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] I4: incorrect answer choice 4 I4-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] Q2: medical expert question A2: correct answer choice A2-Reasoning: brief explanation why correct I1: incorrect answer choice 1 I1-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject]. [ If error-consistent: Error-path notation]. [Minimal-pair label if applicable] I2: incorrect answer choice 2 I2-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] I3: incorrect answer choice 3 I3-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] I4: incorrect answer choice 4 I4-Reasoning: Source: [A/B/A+B]. [brief explanation with clinical knowledge required to reject] ... [Entity Category, Attribute, Relationship Schema] 1. Diagnosis -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 2. Finding -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 3. Symptom -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 4. Procedure -- is_main, status, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 5. Outcome -- status, assertion, time(dates, times, durations, frequencies) 6. Vital_Sign -- status, value, assertion, time(dates, times, durations, frequencies) 7. Physical_Exam -- status, finding, finding_site, laterality, assertion, time(dates, times, durations, frequencies) 8. Lab_Test -- value, abnormal_flag, status, assertion, time(dates, times, durations, frequencies) 9. Diagnostic_Imaging_Test -- status, result, finding_site, laterality, assertion, time(dates, times , durations, frequencies) 10. Specimen -- source, assertion, time(dates, times, durations, frequencies) 11. Microbiology_Test - organism_growth, result, assertion, time(dates, times, durations, frequencies) 88 12. Microbiology_Organism - growth, assertion, time(dates, times, durations, frequencies) 13. Microbiology_Antibiotics - dilution, sensitivity, assertion, time(dates, times, durations, frequencies) 14. Activity -- status, assertion, time(dates, times, durations, frequencies) 15. Medication -- strength, dosage, form_description, form, route, instruction, disp, refills, assertion, time(dates, times, durations, frequencies) 16. Event -- status, assertion, time(dates, times, durations, frequencies) 17. Allergy -- status, assertion, time(dates, times, durations, frequencies) 18. Medical_Device - status, assertion, time(dates, times, durations, frequencies) 19. Instruction - instruction_text, assertion, time(dates, times, durations, frequencies) is_main attribute is either true or false. If the Diagnosis/Symptom/Finding is the main Diagnosis/ Symptom/Finding for the patient during the current admission (e.g., chief complaint, discharge diagnosis(primary), main Diagnosis/Symptom/Finding according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false (e.g., discharge diagnosis( secondary), not main Diagnosis/Symptom/Finding according to context). If the Procedure is the main Procedure for the patient during the current admission (e.g., major surgical or invasive procedure, main Procedure according to headers such as brief hospital course), it is set as is_main=true. If not, it is set as false. abnormal_flag attribute in Lab_Test, is either normal or abnormal. 1. Procedure/Medication -> improves -> Diagnosis/Symptom/Finding 2. Procedure/Medication -> worsens -> Diagnosis/Symptom/Finding 3. Procedure/Medication -> causes -> Diagnosis/Symptom/Finding 4. Procedure/Medication -> administered_for -> Diagnosis/Symptom/Finding 5. Procedure/Medication -> reveals -> Diagnosis/Symptom/Finding 6. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> reveals -> Diagnosis/Finding 7. Specimen -> has_test -> Microbiology_Test/Lab_Test 8. Microbiology_Test -> detected -> Microbiology_Organism 9. Microbiology_Organism -> tested_against -> Microbiology_Antibiotics 10. Microbiology_Organism -> reveals -> Diagnosis/Finding 11. Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding 12. Diagnosis/Symptom/Finding -> causes -> Event 13. Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding/Activity 14. Outcome/Finding -> after -> Diagnosis/Symptom/Finding 15. Lab_Test/Diagnostic_Imaging_Test/Physical_Exam/Vital_Sign -> resulted_in -> Procedure/Medication 16. Procedure/Medication -> causes -> Procedure/Medication 17. Procedure/Medication -> causes -> Event 18. Procedure/Medication -> switched_to -> Procedure/Medication 19. Procedure/Medication -> resulted_in -> Outcome/Finding 20. Outcome/Finding -> after -> Procedure/Medication 21. Event -> causes -> Diagnosis/Symptom/Finding 22. Event -> causes -> Procedure/Medication 23. (any category) -> has_instruction -> Instruction The difference between "causes" and "resulted_in" in "Procedure/Medication -> causes -> Diagnosis/ Symptom/Finding" and "Procedure/Medication -> resulted_in -> Outcome/Finding" is that "causes" is used when a procedure or medication produces an adverse effect or complication, while " resulted_in" is used when a procedure or medication leads to a desired or expected state or improvement (positive or neutral). Similarly, the difference between "causes" and "resulted_in" in "Diagnosis/Symptom/Finding -> causes -> Diagnosis/Symptom/Finding" and "Diagnosis/Symptom/Finding -> resulted_in -> Outcome/Finding " is that "causes" is used when a diagnosis, symptom, or finding produces an adverse effect or complication, while "resulted_in" is used when a diagnosis, symptom, or finding leads to a desired or expected state or improvement (positive or neutral). [Patientâs Discharge Summaries] discharge_summaries [Initial Multi-turn QA] multiturn_qa_step1_step2 I.4 Step 4: Evidence Grounding Source-Location Answer Choice Generation Placeholders. discharge_summaries â the patientâs full sequence of discharge summaries; multiturn_qa_step1_step3 â the Step 3 content questions and answer choices, combined with the Step 1 answer evidence and source locations. You are provided with: 1. A patientâs discharge summaries, which may consist of multiple notes. Each note has a Note #, Chartdate, and Section Headers. 2. An initial Multi-turn QA set that contains medical expert multi-turn questions, along with the correct answer, answer evidence, and answer source locations for each question. 89 Your task is to generate evidence grounding source-location multiple choice questions for each medical expert question. Task: For each medical expert question in the Initial Multi-turn QA, create a new paired source-location question in the form: "What exact source(s) from the patientâs discharge summaries contain information for your answer Ax: âexact answer textâ to the previous question: Qx: âexact question textâ?" Then generate: - Ax (Correct answer choice) - I1--I4 (Incorrect answer choices) Correct Answer Requirements: The provided answer evidence and answer source are for reference only. They may include notes and headers that are redundant --- sources that repeat information already covered by other headers . Do not blindly copy the provided answer sources. Instead, independently determine which notes and headers are needed to support the provided "answer" text: - A source (note + header) is necessary if and only if it contains information for the provided answer AND removing it would leave some part of the answer unsupported - A source is redundant and must be excluded if another included source already contains that same information for the answer - Prefer the most comprehensive header when multiple headers cover the same information for the answer - Include all required notes if the answer spans multiple notes - Exclude headers that do not contain information for the provided answer The correct answer choiceâs notes/headers must be minimal but sufficient: every included source must contribute information for the provided answer, and together they must fully support it. Multi-Header Source Guidance for Reasoning Questions: For questions whose answers involve clinical reasoning chains (e.g., WHY a decision was made, HOW a course unfolded), the correct source set may include multiple headers that together establish the reasoning path. Include all headers necessary to trace the documented reasoning chain, but verify that each header contributes unique information to the reasoning --- do not include headers that merely repeat context available in another included header. Incorrect Answer Requirements: Generate exactly four incorrect answers. Each must: - Contain meaningful errors that prevent full support of the answer - Be clearly incorrect under careful inspection - Not be trivially dismissible Each incorrect answer must include at least one of: - incorrect note number, incorrect header, missing required header, inclusion of irrelevant header, mixing correct and incorrect sources Additional constraints: - An incorrect answer must fail to fully support the answer, even if partially correct - Do not create distractors that are supersets of the correct answer that still fully support the answer, or subsets that still fully support the answer. - Avoid trivial variations (e.g., removing trivial headers) Distractor Quality Requirements: - At least 2 of I1--I4 must be semantic traps: their header names must plausibly align with the questionâs surface keywords (e.g., "intervention" -> Major Surgical or Invasive Procedure; " outcome" -> Discharge Disposition / Discharge Condition; "imaging" -> Pertinent Results; " presentation" -> History of Present Illness) so a reader looking only at header names could reasonably select them. The distractor must still fail at the content level. - At most 1 of I1--I4 may be Category 1a. The other three must be Category 1b, 1c, or 2. - These two constraints may overlap: a single distractor that is both a semantic trap and Category 1 b/1c/2 counts toward both. Grounding Rules: 1. Use only information present in the discharge summaries. Do not invent notes, chartdates, or headers that do not exist. 2. For headers, include only main section headers (e.g., Brief Hospital Course, Major Surgical or Invasive Procedure, History of Present Illness, Discharge Diagnosis, Discharge Medications, Chief Complaint, Discharge Instructions, Pertinent Results, Past Medical History, Medications on Admission, Allergies, Discharge Condition, Physical Exam). Do not include sub-headers (e.g., Microbiology, Imaging within the Pertinent Result header). Header normalization rules: - Always write headers in canonical Title Case (e.g., "Brief Hospital Course"), not the raw casing from the note (e.g., NOT "BRIEF HOSPITAL COURSE" or "brief hospital course"). - Do NOT include trailing punctuation such as ":" or "___" (e.g., write "Physical Exam", NOT " Physical Exam:", "Physical ___:", or "PHYSICAL ___"). 90 - If a header appears in the note in an anonymized or partially redacted form (e.g., "Physical ___") but clearly corresponds to a common main header, normalize it to the canonical form (e.g., " Physical Exam"). - If the relevant content lives under a sub-header (e.g., "Microbiology", "Imaging", "Labs") inside "Pertinent Results", write the parent main header ("Pertinent Results") instead. The same rule applies to sub-headers inside "Brief Hospital Course", "Physical Exam", etc. - Never invent a header that does not appear in the discharge summary. 3. The correct answer must be complete and uniquely correct. - Independently determine the minimal set of sources (notes + headers) whose content supports the provided answer --- do not simply copy the provided answer sources. - The provided answer sources are reference only and may contain redundant headers. If a header in the provided answer source does not itself contain information for the provided answer (because that information is already covered by another header), exclude it. - If only a subset of the provided answer sourceâs headers are needed to fully support the answer, use only that subset. Important constraint: Do not include any incorrect answer choices that contain that same subset plus additional headers that also contain relevant information, because this would produce multiple correct answers. 4. Incorrect answers must be clearly incorrect. - Each incorrect answer must contain meaningful errors (e.g., incorrect note number, incorrect header, missing critical source, adding an irrelevant source, mixing correct and incorrect sources). Minor variations such as omitting or adding trivial headers should not be used. Each incorrect answer must be obviously insufficient to support the answer. - Headers like Brief Hospital Course or History of Present Illness may appear in incorrect answers only if they clearly lack the required information for the answer. 5. Chartdate accuracy: For both correct and incorrect answer choices, The Note # and Chartdate must match the discharge summaries exactly. Do not introduce incorrect chartdates. 6. Source formatting must exactly match the discharge summary structure: - The sources must be written as (where N is the note number, Y-M-D is the chartdate, and headers are listed plainly with no surrounding brackets, braces, or quotes): Note #N Chartdate: Y-M-D Headers: Header One, Header Two, Note #N Chartdate: Y-M-D Headers : Header One, Header Two - Multiple source notes must be separated by comma + space, and multiple headers within the same note must also be separated by comma + space. - Headers must NOT be wrapped in curly braces, square brackets, or any other delimiter. Write them as a plain comma-separated list immediately after "Headers: ". - Concrete example of the required format: Note #1 Chartdate: 2143-11-30 Headers: History of Present Illness, Note #2 Chartdate: 2152-04-30 Headers: History of Present Illness, Brief Hospital Course, Note #3 Chartdate: 2152-05-29 Headers: Brief Hospital Course - Concrete examples of FORBIDDEN formats (do NOT do any of these): Note #4 Chartdate: 2178-12-09 Headers: History of Present Illness, Discharge Diagnosis <-- braces around headers Note #1 Chartdate: 2186-06-08 Headers: [History of Present Illness] <-- square brackets Note #1 Chartdate: 2186-06-08 Headers: "History of Present Illness" <-- quotes 7. Do not output evidence text beyond the required reasoning fields. =================================================================== REASONING REQUIREMENT =================================================================== For each answer choice, provide a brief reasoning line: Correct answer reasoning must explain: - Completeness: Why these headers fully contain the information for Ax --- what specific information each header contributes toward Ax. - Minimality: Why this set is minimal --- why removing any single header would leave some part of Ax unsupported. Incorrect answer reasoning must explain: - Category label: Which category the incorrect choice falls into (1a, 1b, 1c, or 2). - Category 1a (all irrelevant): State that none of the listed headers contain any information for Ax , and briefly note what these headers contain instead. - Category 1b (partial/insufficient): Identify which specific part(s) of Ax are NOT contained in the listed headers. - Category 1c (mix of partial and irrelevant): Identify which headers are partially relevant (and what part of Ax they miss), and which headers are completely irrelevant. - Category 2 (correct subset + irrelevant extras): State which headers form the correct subset that covers Ax, and for each additional header, explain why it is irrelevant to Ax (what that header actually contains instead, or that it does not mention the topic of Ax). If a header looks relevant by name but does not contain Ax information, explicitly state this (e.g., "Discharge Medications lists current medications at discharge but does not mention the dosage adjustment during hospitalization discussed in Ax"). For each original medical expert question, output exactly: Q1-1: What exact source(s) from the patientâs discharge summaries contain information for your answer A1: âexact answer textâ to the previous question: Q1: âexact question textâ? A1-1: correct answer choice 91 A1-1-Reasoning: Fully covers Ax: what specific information each header contributes. Minimal: why each header is necessary --- what part of Ax would be unsupported if it were removed. I1: incorrect answer choice 1 I1-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I2: incorrect answer choice 2 I2-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I3: incorrect answer choice 3 I3-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I4: incorrect answer choice 4 I4-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. Q2-1: What exact source(s) from the patientâs discharge summaries contain information for your answer A2: âexact answer textâ to the previous question: Q2: âexact question textâ? A2-1: correct answer choice A2-1-Reasoning: Fully covers Ax: what specific information each header contributes. Minimal: why each header is necessary --- what part of Ax would be unsupported if it were removed. I1: incorrect answer choice 1 I1-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I2: incorrect answer choice 2 I2-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I3: incorrect answer choice 3 I3-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. I4: incorrect answer choice 4 I4-Reasoning: Category 1a/1b/1c/2. explanation per category guidelines above. ... [Patientâs Discharge Summaries] discharge_summaries [Initial Multi-turn QA] multiturn_qa_step1_step3 92