Paper deep dive
The Verbose Context Problem in Medical Records
Shiva Kaul, Min-Gyu Kim, Anjum Khurshid, Sriram Vishwanath
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 7/5/2026, 2:08:19 AM
Summary
The paper introduces PopMedQA, a benchmark designed to evaluate the 'verbose context problem' in medical records, where structured concepts (like ICD-10 codes) result in token-inefficient textual representations that hinder long-context reasoning in LLMs. The authors also present 'neopatient', a library for language-controlled generation of synthetic longitudinal patient records. Experimental results show that domain-independent methods like prompt compression and agentic decomposition fail to alleviate the problem, suggesting a need for domain-specific input structures for population-scale medical reasoning.
Entities (6)
Relation Signals (4)
neopatient → constructs → PopMedQA
confidence 100% · We construct the benchmark using neopatient...
PopMedQA → isolates → Verbose Context Problem
confidence 100% · We present PopMedQA, a benchmark isolating this problem...
ICD-10 → isrepresentedin → Electronic Health Records (EHRs)
confidence 100% · structured concepts—such as medical codes in electronic health records (EHRs)
neopatient → producesdatain → MEDS
confidence 100% · The resulting records are produced in the Medical Event Data Standard (MEDS) format.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in population health: cohort-level analysis of longitudinal patient records requires reasoning over thousands of medically-coded events, often exceeding 400K tokens in total. We present PopMedQA, a benchmark isolating this problem through computational tasks on groups of longitudinal patient records. We construct the benchmark using neopatient, a new library for language-controlled generation of artificial patient records. Through extensive ablations -- including prompting strategies, prompt compression, and agentic decomposition -- we find that domain-independent methods fail to alleviate the verbose context problem. There remains significant opportunity to exploit domain-specific structure in language model inputs for population-scale reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2606.29503v1
- Canonical: https://arxiv.org/abs/2606.29503v1
Trouble viewing inline? Open PDF directly →
Full Text
57,526 characters extracted from source content.
Expand or collapse full text
The Verbose Context Problem in Medical Records Shiva Kaul 1 Min-Gyu Kim 2 Anjum Khurshid 3 Sriram Vishwanath 4 Abstract The verbose context problem occurs when struc- tured concepts have token-inefficient textual rep- resentations. This bottleneck is acute in popula- tion health: cohort-level analysis of longitudinal patient records requires reasoning over thousands of medically-coded events, often exceeding 400K tokens. We present PopMedQA, a benchmark iso- lating this problem through computational tasks on groups of longitudinal patient records. We construct the benchmark using neopatient, a new library for language-controlled generation of arti- ficial patient records. Through extensive abla- tions—including prompting strategies, prompt compression, and agentic decomposition—we find that domain-independent methods fail to alle- viate the verbose context problem. There remains significant opportunity to exploit domain-specific structure in language model inputs for population- scale reasoning. 1. Introduction Population health analytics focuses on identifying patterns, detecting anomalies, and quantifying disease burden across large groups of individuals. While large language models (LLMs) offer a flexible alternative to traditional rule-based risk adjustment systems, their application is hindered by the verbose context problem. This problem arises when struc- tured concepts—such as medical codes in electronic health records (EHRs)—have token-inefficient textual representa- tions that inflate context lengths. For example, “ICD-10 I21: Acute myocardial infarction” is a textual representation str(c)of the underlying conceptcof a heart attack. The 1 Work completed at Department of Population Medicine prior to current affiliation. 2 Department of Biomedical Informatics, Ajou University School of Medicine 3 Department of Population Medicine, Harvard Pilgrim Health Care Institute and Harvard Medical School 4 School of Electrical and Computer Engineering, Georgia Institute of Technology. Correspondence to: Shiva Kaul <me@shivakaul.com>. Proceedings of the Workshop on Structured Data for Health at the43 rd International Conference on Machine Learning, Seoul, South Korea. Copyright 2026 by the author(s). usual process is to tokenize the string, lookup the token em- beddings, and process the resulting sequence of vectors by the modelf. In this notation, the language model’s output is:f◦ emb◦ tok◦ str(c) : = y. When such concepts pervade the context, as they do in longitudinal patient records, the total context length becomes too long for effective, efficient reasoning. We capture the verbose context problem in a new bench- mark called PopMedQA. Each of its questions involves reasoning over groups of 10-50 synthetic longitudinal pa- tient records, which typically amount to 64K-256K tokens in a textual representation. As described in Section 2 and Section B, it differs from existing long-context benchmarks in two ways. First, it specifically isolates how verbosity, rather than the presence of irrelevant “hay” context, affects performance. Second, it involves multi-hop reasoning over population-scale cohorts that cannot be decomposed into questions about individuals. Constructed with clinician and expert review, PopMedQA reflects the real-world priorities of population health, such as identifying latent clinical clus- ters or detecting sparse anomalies across disparate patient trajectories. To construct the large volume of synthetic data required for PopMedQA, we introduce neopatient, a new soft- ware library for language-controlled generation of artifi- cial patient records. Unlike rule-based generators (like Synthea (Walonoski et al., 2018)), neopatient trajectories are controlled through natural language descriptions, allowing for the creation of complex clinical cohorts without cus- tom simulation code. Details on neopatient are provided in Section 3. In Section 4, we conduct a thorough evaluation of a range of language models on PopMedQA. These confirm the claimed design characteristics of PopMedQA, such as decomposi- tion resistance. For ablations, we conduct a meta-analysis of multiple families of techniques for improving long-context performance, including prompting strategies, prompt com- pression, and agentic decomposition. Our analysis reveals several generalizable insights on the nature of population-scale EHR reasoning: (1) both clinical competence and long-context capability are required; (2) generic prompt compression is fragile; (3) medical pretrain- ing does not substantially improve performance; and (4) 1 arXiv:2606.29503v1 [cs.CL] 28 Jun 2026 The Verbose Context Problem in Medical Records Clustering These 29 lung cancer patients are all receiving the same immunotherapy regimen. 7051705111019110191457614576164911649126342634 193851938536033603801080104439443936663666 16492164922063520635203412034111507115071336213362 131141311445864586182071820714841148411738117381 93559355769976991889818898110981109884168416 6478647883708370 Cluster them into three distinct ’treatment response phenotypes’ based on the longitudinal pattern of their side effects. Top-k An oncologist’s schedule for tomorrow has two last- minute openings due to cancellations. From this wait- list of 10 patients needing follow-up, select the patients who should be offered the time slots. 981398132974297422492249145911459122962296 1862318623184831848334273427905290522459624596 Make sure to prioritize any patients who have experi- enced any delays in care. Two-Sample Test Here are two groups of eight patients newly diagnosed with rheumatoid arthritis. Group A has a commercial PPO insurance plan, while Group B has a Medicaid Managed Care plan. Are their initial treatment path- ways over the first 6 months statistically distinguish- able? Group A: 91099109604360433424342466786678 125351253521872187109811098150405040 Group B:108221082220042004748674861777817778 91019101197611976122722227222445824458 100K200K300K400K Token Counts(as Text) 5K10K15K20K25K Total Patient Events 1015202530 Number of Patients 1K2K3K4K5K Number of Unique Codes Figure 1. Example questions from PopMedQA (above), as well as statistics of all its questions (below). Within the questions, each patient record is visually depicted as a boxed patient ID number. Since these records may contains thousands of coded events, and codes have verbose textual representations, PopMedQA is a long-context benchmark. agentic decomposition is ineffective and/or cost-prohibitive. Overall, we find that domain-independent methods fail to alleviate the verbose context problem, exposing a signifi- cant unrealized opportunity to exploit domain structure for population-scale reasoning. 2. PopMedQA: Verbosity in Medical Records To embody the verbose context problem, we present a new benchmark. PopMedQA consists of questions about groups (of size between 10 and 50) of longitudinal patient records. These records are drawn from over 25 thousand synthetic patients generated specifically for the benchmark. These records comprise over 14M medically-coded events. When a typical code is conveyed as text, it uses between 8 and 20 tokens. A code accompanies each timestamped event. There are typically thousands of such events in each record. This makes most of PopMedQA’s questions approximately 64K, 128K, or 256K tokens in length, when represented as text. An important design feature of PopMedQA is decomposi- tion resistance. Long context can often be circumvented by partitioning it into separate chunks, process the chunks separately, and aggregating an answer from multiple rounds of processing. The purpose of PopMedQA is to more specif- ically reward verbose-context capability rather than this more routine context engineering. Accordingly, most of PopMedQA’s questions are designed so that they cannot be readily solved by asking a series of individual patient-level questions. Each question poses one of nine computational tasks (see Figure 4 and Section C). Some of these, such as planted clique and clustering, are especially challenging without holistic in-context processing. We see quantitative evidence of such decomposition resistance in our experi- mental results, presented in Section 4. PopMedQA presents realistic questions from population health. Whereas clinical medicine and biomedicine focus on interventions upon individual patients, population health concerns questions about groups of patients. PopMedQA’s questions were reviewed by both clinicians and population health experts for pertinence and validity. 2.1. PopMedQA Pipeline Since the questions in PopMedQA are very long, they are generally not possible for humans (or computers) to solve or verify. Thus, a correct pipeline cannot generate questions and then (independently) generate labels. To ensure answers are correct, the questions and the data involved in the an- swers must be jointly generated. The overall pipeline is as follows. 1. Task Definition This involves specifying the schema of the answer (e.g. a list of lists of patient IDs) as well as the scoring function between the answer and the truth (see Appendix C for details on each task’s metric). It also involves specifying the cohorts that should be generated so that the answer can be synthesized from them. For example, planted clique is the task of finding the most cohesive or similar size-ksubset of patients. The task indicates two cohorts should be created: the size-kclique, and theN − k remaining patients. 2 The Verbose Context Problem in Medical Records 2. Question IdeationGiven a task, abstract question ideas (or template) are generated. The idea does not have a con- crete question statement, nor does it have individual patients generated. Instead, the idea has a question template param- eterized byN, e.g. “among these N ED and urgent care records from the last 72 hours in our metro area, is there a subset of patients presenting with an unusual combina- tion of symptoms...”. The idea also gives question-specific descriptions of the cohorts that need to be generated. For example: 1.Syndromic subset of sizemin(12, int(N/4)): a subset of patients from recent ED and urgent care records in a specific zip code over the last 72 hours. All have a doc- umented chief complaint that includes a combination of ’severe gastrointestinal distress (vomiting/diarrhea)’ AND ’unusual rash’. 2.Other population of sizeN − min(12, int(N/4)): all other ED and urgent care records from the last 72 hours, showing typical presentations like chest pain, respira- tory infections, and minor trauma, without the combi- nation of severe GI distress and rash. 3. Question SamplingFor each question idea, cohorts are generated at a large sizeN max . To create concrete questions of sizeN, the idea’s cohorts are subsampled at the specified value of N . 4. Clinician Validation To ensure the quality and practi- cal relevance of the benchmark, questions undergo a clin- ician validation step. Clinician reviewers (M.D.s) score each question from 1–10 on three metrics: (1) Realism (how likely or commonly the questions would arise in practice), (2) Difficulty (of solving them by hand), and (3) Coherence (whether the cohorts are well-defined and distinct). Ques- tions are retained in PopMedQA only if they score at least 5 on all three metrics. 3. neopatient: Language-Controlled Generation of Patient Records To generate the large volume of synthetic data required for PopMedQA, we implemented a new software library, neopatient, for language-controlled generation of artificial patient records. Unlike rule-based generators (like Synthea (Walonoski et al., 2018)), patient trajectories in neopatient are controlled through natural language descriptions, allow- ing for the creation of complex clinical cohorts without the need for custom simulation code or state machines. Generating longitudinal patient records with LLMs presents several inherent challenges. First, output length constraints are significant; LLM generations are typically limited to less than 64K tokens, and they often become unreliable when producing long, structured outputs at that scale. Sec- ond, coding knowledge is limited, as LLMs do not precisely know the vast and frequently updated medical ontologies. Third, maintaining clinical plausibility requires managing complex sequential dependencies, such as ensuring a pre- scription follows an appropriate diagnosis. Finally, the need for batching for efficiency when generating large cohorts of hundreds of patients restricts the use of complex, multi- turn agentic loops. The neopatient pipeline consists of four primary stages designed to address these challenges: 1. Sampling An LLM generates individualized “patient recipes” that define demographics and divide the patient’s life trajectory into discrete temporal segments. This high- level blueprint ensures long-term clinical coherence—such as maintaining consistent medication dosages—while seg- mentation allows the system to bypass LLM context limits and avoid the “drifting” common in long-form generation. 2. GenerationFor each recipe, an LLM produces longitu- dinal medical events across the temporal segments. For each event, the LLM generates a natural language description as well as a target coding system (e.g., ICD-10 or SNOMED). At this stage, the descriptions are medically plausible but may not yet match official ontology strings exactly. 3. Matching Because LLMs are prone to hallucinating invalid codes or using imprecise language, the system uses a precomputed vector database (e.g., ChromaDB) to map free- text descriptions to standardized medical codes (SNOMED, ICD-10, LOINC, RxNorm, etc.). 4. Verification A final correctness check is performed where an LLM validates each completed record against the original input specifications. This acts as an automated quality gate, filtering out records that failed to follow the recipe or accidentally triggered exclusion criteria. The resulting records are produced in the Medical Event Data Standard (MEDS) format (Arnrich et al., 2024). neopa- tient is designed to be scalable, using LLM batch APIs and state-tracking for resumability, enabling the cost-effective generation of tens of thousands of records. Static Resources A vector database of medical codes ensures that natural language descriptions are accurately mapped to standardized ontologies. This database was con- structed by embedding code descriptions (using Qwen 3 8B to produce 4096-dimensional vectors). Second, the library maintains a set of reference statistics derived from real- world EHR data. these statistics define the typical length and density of both inpatient and outpatient longitudinal records, ensuring that the generated synthetic trajectories reflect realistic clinical patterns. 3 The Verbose Context Problem in Medical Records Input: Clinical Cohort Criteria Positive Description “Patient(s) with persistent atrial fibrillation who...” Negative Description “No valvular heart disease” Reference Population Statistics Sampling Sampler age 68, white Segment 1 Segment 2 ... ... Segment N Generation Generator Generator Generator Patient "Recipe" Matching vector database of codes Coded Events CPT 93656 ICD-10 148.1 LOINC 2339-0 Textual Events catheter ablation persistent atrial fib... blood glucose meas. Verification Verifier Reject Output: Longitudinal Patient Record 3819 Samuel Watkins Figure 2. The neopatient architecture for language-controlled artificial patient generation. The pipeline transforms natural language criteria describing a cohort into a set of coded longitudinal records in the MEDS format. The architecture consists of four primary stages: (1) Sampling, where an LLM generates a “patient recipe” that defines demographics and divides the patient’s life trajectory into discrete temporal segments to ensure coherence; (2) Generation, where longitudinal medical events are produced for each segment; (3) Matching, where a vector database maps natural language descriptions to standardized medical codes; and (4) Verification, a final correctness check where an LLM validates that the completed record strictly satisfies the original cohort specification. Note on Realism The primary goal of neopatient is to generate records that strictly adhere to provided clinical specifications rather than to exhaustively replicate every facet of real-world patient records. This design priority serves the core objective of PopMedQA: to isolate and eval- uate the specific computational challenges of the verbose context problem in a controlled setting. Nonetheless, language-controlled generation potentially en- ables researchers to maximize realism while preserving pri- vacy. The neopatient prompt can be optimized end-to-end (with TextGrad (Y ̈ uksekg ̈ on ̈ ul et al., 2025), GEPA (Agrawal et al., 2025), or similar) to generate records that maintain high fidelity with real, sensitive patient records. The re- sulting prompt can be more easily reviewed for privacy compliance than models trained with deep learning, which encode sensitive patient information in their weights. 4. Experiments We conduct a thorough evaluation of a range of language models on PopMedQA, testing models that range from 7B parameters to frontier scale, with context lengths from 128K to 2M tokens. On top of baseline models, we examine four families of interventions or ablations designed to improve long-context performance: Prompting strategies We consider three ways of repre- senting patient records as text: (1) a baseline method that re- places codes with truncated descriptions, (2) a naive method using only raw codes, and (3) a ”codebook” method that, at the beginning of each prompt, maps unique codes to IDs, thereby reducing redundancy (see Figures 7 and 8). Prompt compression We evaluate two methods: render- ing text to images using the Glyph pipeline (Cheng et al., 2025) and LLMLingua-2 text chunk compression using the microsoft/llmlingua-2-xlm-roberta-lar ge-meetingbank model (Jiang et al., 2023). Reasoning We expand inference-time compute using chain-of-thought prompting (Wei et al., 2022). ”Think” models were run with default temperatures and no upper bound on answer length, while non-reasoning models were run with temperature 0 and a 2048 token limit. Multi-turn interaction We utilize agentic context engi- neering and decomposition via systems like MARS, Long- CEPO, and Claude Code. These systems were allowed to perform multiple queries to solve a single question, with Claude Code having a $5 USD limit per answer. 4.1. Results and Discussion The primary experimental results on PopMedQA are pre- sented in Figures 3 and 5. Overall, the leaderboard re- sults align with expectations, as frontier models demon- strate superior performance, followed by large open-source models with strong long-context capabilities. The tasks in PopMedQA effectively stress frontier-level capabilities, with performance declining as context length increases. De- tailed task-specific scoreboards are provided in Figure 6. 4 The Verbose Context Problem in Medical Records 100K200K300K400K 20% Win 40% Win 60% Win 80% Win Maximum question length(tokens) 1. Gemini 3 Flash(Reasoning) 2. Gemini 3 Flash 3. Claude Code(Sonnet 4.5) 4. Claude Sonnet 4.5 5. Gemini 2.5 Flash 6. Grok 4.1 Fast(Reasoning) 7. Gemini 2.5 Flash(Reasoning) 8. Grok 4.1 Fast 9. Gemini 3 Pro 10. GPT-5 Mini(Reasoning) 11. Gemini 2.0 Flash 12. Qwen3 VL 235B(FP8) 13. Gemini 2.5 Flash Lite 14. Qwen3 VL 235B(Q4) 15. GPT-5 Mini 16. Qwen3 Next 80B(FP8) 17. Qwen3 235B(Q4) 18. Qwen3 235B(Q4, YaRN, Reasoning) 19. Nemotron 3 Nano 30B(BF16, Reasoning) 20. Nemotron 3 Nano 30B(FP8, Reasoning) 21. Nemotron 3 Nano 30B(BF16) 22. Nemotron 3 Nano 30B(FP8, YaRN) 23. Qwen3 VL 235B(Q4, Reasoning) 24. MegaBeam Mistral 7B(512K) 25. GPT-5 Nano 26. Qwen3 Next 80B(FP8, Reasoning) 27. Nemotron 3 Nano 30B(FP8) 28. Qwen3 VL 30B(FP8) 29. Qwen 2.5 7B(1M) 30. gpt-oss 120B 31. Gemma 3 27B 32. MedGemma 27B Figure 3. Leaderboard performance of models across all tasks on PopMedQA. The y-axis is the percentage of comparisons the model won against all other models. The x-axis restricts these comparisons to questions up to a given token length. We observe that performance stresses frontier-level capabilities. Clinical competence and long-context capability are re- quired. Maintaining a high rank at increased context lengths is not guaranteed. While models like Gemini 3 Flash are consistent, smaller models like Nemotron 3 Nano only rise in rank as the context expands, indicating that both medical domain knowledge and long-context processing are essential for PopMedQA. Generic prompt compression is fragile. Compressed prompts, such as those generated by rendering text to im- ages, significantly degrade model performance, particularly in instruction following. Domain-independent compression methods appear to strip away critical information needed for complex medical reasoning. Medical pretraining does not substantially improve per- formance.General-purpose frontier models often outper- formed specialized medical models on PopMedQA. This suggests that the primary challenge is not a lack of clinical knowledge but rather the ability to perform robust, multi- hop reasoning over the verbose contexts characteristic of longitudinal EHR data. Other works have cast doubt on the utility of medically-specialized language models (Jeong et al., 2024). Agentic (multi-turn) decomposition is ineffective and/or cost-prohibitive. While agentic systems showed relative strength, they failed to achieve absolute performance gains that justify their high computational and financial costs. Systems like MARS and LongCEPO frequently harmed instruction-following and absolute performance compared to monolithic baselines. Furthermore, Claude Code did not surpass the efficiency of frontier monolithic models. This confirms that PopMedQA’s tasks are resistant to simple decomposition and require holistic in-context reasoning. 5. Conclusion The verbose context problem inhibits population-level EHR reasoning. Our results demonstrate that domain- independent methods—including prompt compression and agentic decomposition—fail to alleviate performance degra- dation. These findings reveal a significant unrealized op- portunity to exploit domain-specific structure in language model inputs to enable robust population-scale reasoning. Beyond the verbose context problem, language-controlled generation of artificial patient records, as in neopatient, can accelerate AI research in population health. 5 The Verbose Context Problem in Medical Records References Agrawal, L. A., Tan, S., Soylu, D., Ziems, N., Khare, R., Opsahl-Ong, K., Singhvi, A., Shandilya, H., Ryan, M. J., Jiang, M., Potts, C., Sen, K., Dimakis, A., Sto- ica, I., Klein, D., Zaharia, M., and Khattab, O. GEPA: Reflective prompt evolution can outperform reinforce- ment learning. In First Workshop on Foundations of Reasoning in Language Models, 2025. URLhttps: //openreview.net/forum?id=4o6XTL6Oj. Arnrich, B., Choi, E., Fries, J., McDermott, M., Oh, J., Pol- lard, T., Shah, N., Steinberg, E., Wornow, M., and van de Water, R. Medical event data standard (meds): Facilitat- ing machine learning for health. In ICLR 2024 Workshop on Learning from Time Series For Health (TS4H), 2024. URLhttps://openreview.net/forum?id= IsHy2ebjIG. Artificial Analysis. Artificial analysis long context reason- ing benchmark leaderboard (a-lcr). Online leaderboard https://artificialanalysis.ai/evalua tions/artificial-analysis-long-conte xt-reasoning, 2025. Independent benchmark eval- uating language models’ ability to extract, reason about, and synthesize information from long-form documents (10k–100k tokens, cl100kbase tokenizer). Continuously updated as of February 2026. Part of the Artificial Analy- sis Intelligence Index. Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al. Longbench: A bilingual, multitask benchmark for long context under- standing. arXiv preprint arXiv:2308.14508, 2023. Bai, Y., Tu, S., Zhang, J., Peng, H., Wang, X., Lv, X., Cao, S., Xu, J., Hou, L., Dong, Y., et al. Longbench v2: Towards deeper understanding and reasoning on realis- tic long-context multitasks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3639–3664, 2025. Charlson, M. E., Pompei, P., Ales, K. L., and MacKenzie, C. R. A new method of classifying prognostic comorbid- ity in longitudinal studies: development and validation. Journal of chronic diseases, 40(5):373–383, 1987. Cheng, J., Liu, Y., Zhang, X., Fei, Y., Hong, W., Lyu, R., Wang, W., Su, Z., Gu, X., Liu, X., et al. Glyph: Scal- ing context windows via visual-text compression. arXiv preprint arXiv:2510.17800, 2025. Elixhauser, A., Steiner, C., Harris, D. R., and Coffey, R. M. Comorbidity measures for use with administrative data. Medical care, p. 8–27, 1998. Eyuboglu, S., Ehrlich, R., Arora, S., Guha, N., Zinsley, D., Liu, E., Tennien, W., Rudra, A., Zou, J., Mirhoseini, A., et al. Cartridges: Lightweight and general-purpose long context representations via self-study. arXiv preprint arXiv:2506.06266, 2025. Fleming, S. L., Lozano, A., Haberkorn, W. J., Jindal, J. A., Reis, E., Thapa, R., Blankemeier, L., Genkins, J. Z., Stein- berg, E., Nayak, A., Patel, B., Chiang, C.-C., Callahan, A., Huo, Z., Gatidis, S., Adams, S., Fayanju, O., Shah, S. J., Savage, T., Goh, E., Chaudhari, A. S., Aghaeep- our, N., Sharp, C., Pfeffer, M. A., Liang, P., Chen, J. H., Morse, K. E., Brunskill, E. P., Fries, J. A., and Shah, N. H. Medalign: A clinician-generated dataset for in- struction following with electronic medical records. In Proceedings of the AAAI Conference on Artificial In- telligence, volume 38, p. 22021–22030, 2024. doi: 10.1609/aaai.v38i20.30205. Grolleau, F., Alsentzer, E., Keyes, T., Chung, P., Swami- nathan, A., Aali, A., others, and Chen, J. H. Medfacteval and medagentbrief: A framework and workflow for gener- ating and evaluating factual clinical summaries. In Pacific Symposium on Biocomputing, volume 31, p. 388–399, 2026. Huang, S.-C., Huo, Z., Steinberg, E., Chiang, C.-C., Lan- glotz, C., Lungren, M., Yeung, S., Shah, N., and Fries, J. Inspect: A multimodal dataset for patient outcome pre- diction of pulmonary embolisms. In Advances in Neural Information Processing Systems, volume 36, p. 74141– 74163, 2023. Jahanian, A., Chai, L., and Isola, P. On the ”steerability” of generative adversarial networks. In International Confer- ence on Learning Representations, 2020. URLhttps: //openreview.net/forum?id=HylsTT4FvB. Jeong, D. P., Garg, S., Lipton, Z. C., and Oberst, M. Medical adaptation of large language and vision-language models: Are we making progress? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 12143–12170, 2024. Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., and Qiu, L. Llm- lingua: Compressing prompts for accelerated inference of large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 13358–13376, 2023. Jiang, Y., Black, K. C., Geng, G., Park, D., Zou, J., Ng, A. Y., and Chen, J. H. Medagentbench: A virtual ehr environment to benchmark medical llm agents. NEJM AI, 2(9):AIdbp2500144, 2025. doi: 10.1056/AIdbp2500144. Kindig, D. and Stoddart, G. What is population health? American journal of public health, 93(3):380–383, 2003. 6 The Verbose Context Problem in Medical Records Lee, G., Hwang, H., Bae, S., Kwon, Y., Shin, W., Yang, S., Seo, M., Kim, J.-Y., and Choi, E. Ehrsql: A practi- cal text-to-sql benchmark for electronic health records. In Advances in Neural Information Processing Systems, volume 35, p. 15589–15601, 2022. Li, X. L. and Liang, P. Prefix-tuning: Optimizing continu- ous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 4582–4597, 2021. OpenAI. Graphwalks: Multi-hop reasoning long-context benchmark. Introduced inhttps://openai.com /index/gpt-4-1, April 2025a. Dataset available at https://huggingface.co/datasets/open ai/graphwalks (MIT licensed); evaluates breadth- first search and parent retrieval in large directed graphs. OpenAI. Openai-mrcr: Multi-round coreference resolution benchmark. Introduced inhttps://openai.com /index/gpt-4-1, April 2025b. Dataset available at https://huggingface.co/datasets/open ai/mrcr. Pope, G. C., Kautter, J., Ellis, R. P., Ash, A. S., Ayanian, J. Z., Iezzoni, L. I., Ingber, M. J., Levy, J. M., and Robst, J. Risk adjustment of medicare payments using the cms- hcc model. Health Care Financing Review, 25(4):119, 2004. Subramani, N., Suresh, N., and Peters, M. E. Extracting latent steering vectors from pretrained language models. In Findings of the Association for Computational Linguis- tics: ACL 2022, p. 566–581, 2022. Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D. Long range arena : A benchmark for efficient transformers. In International Conference on Learning Representations, 2021. URLhttps://openreview.net/forum ?id=qVyeW-grC2k. Vodrahalli, K., Ontanon, S., Tripuraneni, N., Xu, K., Jain, S., Shivanna, R., Hui, J., Dikkala, N., Kazemi, M., Fatemi, B., et al. Michelangelo: Long context evaluations beyond haystacks via latent structure queries. arXiv preprint arXiv:2409.12640, 2024. Walonoski, J., Klaus, M., Granger, B., Hall, D., Gregor, C., Neyarapally, T., Watson, A., and Scanlon, J. Synthea: An approach, method, and software mechanism for generat- ing synthetic patients and the synthetic electronic health record. Journal of the American Medical Informatics Association, 25(3):230–238, 2018. Wei, H., Sun, Y., and Li, Y. Deepseek-ocr: Contexts optical compression. arXiv preprint arXiv:2510.18234, 2025. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. Weiner, M. G. et al. The johns hopkins adjusted clinical group (acg) system: A guide for researchers. Health Services Research, 2012. Wornow, M., Thapa, R., Steinberg, E., Fries, J., and Shah, N. Ehrshot: An ehr benchmark for few-shot evaluation of foundation models. Advances in Neural Information Processing Systems, 36:67125–67137, 2023. Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D. Helmet: How to evaluate long-context models effectively and thoroughly. In The Thirteenth International Conference on Learning Repre- sentations, 2025. Y ̈ uksekg ̈ on ̈ ul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J. Textgrad: Automatic ’differenti- ation’ via text. Nature, 2025. Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., et al.∞-bench: Extending long context evaluation beyond 100k tokens. arXiv preprint arXiv:2402.13718, 2024. Zheng, M., Feng, X., Si, Q., She, Q., Lin, Z., Jiang, W., and Wang, W. Multimodal table understanding. In Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 9102–9124, 2024. 7 The Verbose Context Problem in Medical Records A. PopMedQA Details Planted Clique These 12 patients all started a new biologic drug in the last year. Find thek=3patients who are most similar to each other in their longitudinal pattern of secondary effects, forming a distinct but previously unrecognized ‘adverse event phenotype’. 10434104341567156767876787201012010141644164 2295229525322253221568715687884288422145721457 11152111521018410184 Answer: [88428842104341043415671567] In-Context Classification We want to build a rule for appropriate CT scan use for ER patients with headaches. You are given a train- ing set of 8 ER headache visits, labeled by a human reviewer as ‘Appropriate Use’ or ‘Inappropriate Use’. Classify the 2 unlabeled new visits. Positive training examples:707970792319123191 2253422534539539 Negative training examples:16313163131164911649 160401604031513151 Test examples:6805680533103310 Answer: 68056805 :true, 33103310 :false Clustering These 11 patients have multiple documented Social De- terminants of Health (SDOH) needs. Cluster them to find the 3 most common ‘barrier archetypes’ - the com- binations of problems that frequently occur together and create a specific type of system challenge. 464464625625983698362142141229112291392392 19123191232319323193252862528684388438 Answer:[[4644642319323193625625214214 ], ...] Outlier Detection Most of these 30 knee replacement patients follow a standard path of tapering off opioids. Analyze their 6-month post-operative pharmacy and visit records to find any whose pain management journey is a signifi- cant conceptual outlier. 116841168426062606240922409224035240352225322253 2518525185129612969022902221931219311992219922 1632616326205132051313821382226182261880188018 184051840516380163801737617376290429041680616806 153581535817830178304588458813833138331215412154 2212122121171101711022235222353873871030310303 Answer: [ 22253222532261822618 ] Two-Sample Test Here are two groups of six patients with Type 2 dia- betes. Group A is managed by Dr. Smith, a physician with 30 years of experience. Group B is managed by Dr. Jones, a recent residency graduate. Are their key clinical outcomes over the past year distinguishable? Group A:249252492553815381184741847420402040 192961929623402340 Group B: 1988119881742674261554615546385385 13034130341334613346 Answer: false Similarity Search Here is an ‘index case’ of a patient who developed an opioid use disorder after elective surgery. Their pre-operative profile included a co-occurring anxiety disorder and a social history of an unstable living situa- tion. Search this cohort of of new pre-surgical patients and find the 2 whose holistic risk profile is most similar to the index case. Query patient: 1011010110 Other patients:179641796423417234171442314423 20675206752434824348381738172695269552665266 827582751628516285 Answer: [ 5266526638173817 ] Top-k We have two remote patient monitoring (RPM) kits to distribute among these 13 recently discharged heart failure patients. Synthesize their clinical data (readmis- sion risk) and social data (tech literacy, social support from notes) to generate an allocation list that maxi- mizes the predicted reduction in 30-day readmissions for the cohort as a whole, rather than just assigning them to the two clinically sickest patients. 98409840291229121221212212210022100231723172 33563356134931349370237023314231421970919709 1119711197282128211711917119 Answer: [31423142] Sorting Here are 11 patients with brittle Type 1 Diabetes. Rank them by the ‘brittleness’ of their glycemic control. Brit- tleness is not just average A1c, but the frequency and amplitude of swings between hypo- and hyperglycemic events, inferred from lab values, ER visits for DKA/hy- poglycemia, and rapid cycling of insulin dosing. 11731173821482141374213742127891278933683368 471447141212508350831719117191192061920612681268 Answer:[17191171911212508350831374213742 4714471433683368] Classification For these 11 patients with Stage 3 Chronic Kidney Disease (CKD), classify them into three ‘Progression Risk’ tiers (1: Low, 2: Medium, 3: High) for progress- ing to ESRD in the next 5 years, based on their labs and comorbidities. 12621126211889218892362536255323532395389538 21222212225598559815665156651425914259 Answer:1262112621:3,1425914259:2,1566515665:3, 1889218892:1,2122221222:3,36253625:1,53235323:1, 55985598 :2,95389538:2 Figure 4. Example questions from PopMedQA. Each question poses one of nine computational tasks. Each patient record is visually depicted as a boxed patient ID number. 8 The Verbose Context Problem in Medical Records Gemini 2.5 Flash Gemini 2.5 Flash(Reasoning) Gemini 3 Flash Gemini 2.5 Flash Lite Gemini 3 Flash(Reasoning) Gemini 3 Pro Qwen3 Next 80B(FP8) Gemini 2.0 Flash Qwen3 VL 235B(Q4) Qwen3 VL 235B(FP8) MegaBeam Mistral 7B(512K) MedGemma 27B Gemma 3 27B Nemotron 3 Nano 30B(BF16) Nemotron 3 Nano 30B(BF16, Reasoning) Qwen3 VL 30B(FP8) Qwen 2.5 7B(1M) Nemotron 3 Nano 30B(FP8, Reasoning) gpt-oss 120B Prompting -5.1% -1.3% MegaBeam Mistral 7B(512K) Gemini 3 Flash(Reasoning) Gemini 2.5 Flash Lite Qwen3 VL 30B(FP8) Gemini 2.5 Flash Qwen3 VL 235B(Q4, Reasoning) MedGemma 27B Compression -29.4%-25.8% Qwen3 Next 80B(FP8) Qwen3 VL 235B(Q4) Nemotron 3 Nano 30B(FP8, YaRN) Gemini 2.5 Flash Grok 4.1 Fast Gemini 3 Flash GPT-5 Mini Reasoning 4.6% 5.0% Gemini 2.0 Flash Qwen3 VL 30B(FP8) Gemma 3 27B Qwen3 VL 235B(FP8) Grok 4.1 Fast Claude Sonnet 4.5 Multiturn -0.7%-20.5% -40%-35%-30%-25%-20%-15%-10%-5%0%5%10%15%20%25%30% Gemma 3 27B Medical -5.3%-1.3% Figure 5. Meta-analysis of ablations on PopMedQA. Each dot compares two runs on PopMedQA: a baseline and an ablation. The model’s name is on the right; different families of ablations are presented. A dot’s position quantifies the effect of the ablation. A dark dot at -5% indicates that the baseline model won 55% of head-to-head comparisons, and therefore the ablation had a negative effect. A light dot restricts the scoring to examples where both models gave an answer; this distinction is important when the ablation affects the model’s capability to return correctly-formatted answers. For each family, the mean ablation effects are shown as diamonds. Overall, we find that most families of ablations, besides reasoning, are ineffective on PopMedQA. 9 The Verbose Context Problem in Medical Records Qwen3 VL 30B(FP8)+ MARS gpt-oss 120B Qwen3 Next 80B(FP8, Reasoning) Gemini 2.0 Flash+ MARS Gemma 3 27B MedGemma 27B Qwen3 VL 235B(FP8)+ MARS Gemini 2.0 Flash+ LongCEPO Qwen3 Next 80B(FP8) Qwen 2.5 7B(1M) Gemini 2.0 Flash Qwen3 VL 30B(FP8) Qwen3 235B(Q4) Qwen3 235B(Q4, YaRN, Reasoning) Qwen3 VL 235B(Q4) Qwen3 VL 235B(FP8) Qwen3 VL 235B(Q4, MRoPE) Grok 4.1 Fast Nemotron 3 Nano 30B(FP8) Gemini 2.5 Flash Lite Qwen3 VL 235B(Q4, Reasoning) Grok 4.1 Fast(Reasoning) GPT-5 Mini(Reasoning) Grok 4.1 Fast+ MARS GPT-5 Mini GPT-5 Nano Nemotron 3 Nano 30B(FP8, YaRN) Nemotron 3 Nano 30B(BF16) Gemini 3 Pro Gemini 3 Flash MegaBeam Mistral 7B(512K) Nemotron 3 Nano 30B(BF16, Reasoning) Gemini 2.5 Flash(Reasoning) Nemotron 3 Nano 30B(FP8, Reasoning) Gemini 2.5 Flash Claude Sonnet 4.5 Claude Code(Sonnet 4.5) Gemini 3 Flash(Reasoning) 0.00.10.20.30.40.50.6 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Planted Clique MedGemma 27B Gemini 2.0 Flash+ LongCEPO Gemma 3 27B Nemotron 3 Nano 30B(FP8, Reasoning) Gemma 3 27B+ MARS Nemotron 3 Nano 30B(BF16, Reasoning) gpt-oss 120B Nemotron 3 Nano 30B(BF16) Nemotron 3 Nano 30B(FP8) Nemotron 3 Nano 30B(FP8, YaRN) MegaBeam Mistral 7B(512K) GPT-5 Nano Gemini 2.0 Flash+ MARS GPT-5 Mini Gemini 3 Pro Qwen3 VL 30B(FP8)+ MARS Qwen3 VL 30B(FP8) Qwen 2.5 7B(1M) GPT-5 Mini(Reasoning) Grok 4.1 Fast+ MARS Gemini 2.0 Flash Gemini 2.5 Flash Lite Qwen3 VL 235B(Q4, Reasoning) Qwen3 235B(Q4) Qwen3 235B(Q4, YaRN, Reasoning) Qwen3 Next 80B(FP8, Reasoning) Gemini 2.5 Flash(Reasoning) Qwen3 VL 235B(FP8)+ MARS Gemini 2.5 Flash Qwen3 VL 235B(Q4, MRoPE) Gemini 3 Flash(Reasoning) Qwen3 VL 235B(Q4) Qwen3 VL 235B(FP8) Claude Sonnet 4.5 Grok 4.1 Fast(Reasoning) Gemini 3 Flash Claude Code(Sonnet 4.5) Qwen3 Next 80B(FP8) Grok 4.1 Fast 0.00.10.20.30.40.50.6 39 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 In-Context Classification Qwen3 Next 80B(FP8, Reasoning) GPT-5 Nano Gemini 2.5 Flash Lite Qwen3 235B(Q4, YaRN, Reasoning) Qwen3 VL 235B(Q4, Reasoning) Qwen3 Next 80B(FP8) Qwen3 VL 235B(Q4) Qwen3 VL 235B(Q4, MRoPE) Gemini 2.5 Flash Qwen3 235B(Q4) Qwen3 VL 30B(FP8) Nemotron 3 Nano 30B(FP8, YaRN) Qwen 2.5 7B(1M) Qwen3 VL 235B(FP8)+ MARS GPT-5 Mini gpt-oss 120B Qwen3 VL 235B(FP8) Nemotron 3 Nano 30B(BF16) Nemotron 3 Nano 30B(FP8) Claude Sonnet 4.5 Qwen3 VL 30B(FP8)+ MARS Grok 4.1 Fast Nemotron 3 Nano 30B(FP8, Reasoning) Nemotron 3 Nano 30B(BF16, Reasoning) Gemma 3 27B Gemini 2.0 Flash+ MARS Gemini 2.0 Flash+ LongCEPO MedGemma 27B Gemini 2.0 Flash MegaBeam Mistral 7B(512K) Gemini 3 Pro Claude Code(Sonnet 4.5) GPT-5 Mini(Reasoning) Grok 4.1 Fast(Reasoning) Gemini 3 Flash Grok 4.1 Fast+ MARS Gemini 2.5 Flash(Reasoning) Gemini 3 Flash(Reasoning) 0.00.20.40.60.8 39 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 8 7 6 5 4 3 2 1 Clustering Qwen 2.5 7B(1M) Gemma 3 27B Qwen3 VL 30B(FP8)+ MARS MedGemma 27B MegaBeam Mistral 7B(512K) Qwen3 Next 80B(FP8, Reasoning) gpt-oss 120B Qwen3 235B(Q4, YaRN, Reasoning) Qwen3 235B(Q4) Gemini 2.0 Flash+ LongCEPO Qwen3 VL 235B(Q4, Reasoning) GPT-5 Nano Qwen3 VL 30B(FP8) Nemotron 3 Nano 30B(FP8) Nemotron 3 Nano 30B(FP8, Reasoning) Nemotron 3 Nano 30B(FP8, YaRN) Nemotron 3 Nano 30B(BF16, Reasoning) Nemotron 3 Nano 30B(BF16) Qwen3 VL 235B(FP8)+ MARS Gemini 2.5 Flash Lite Qwen3 Next 80B(FP8) Qwen3 VL 235B(FP8) GPT-5 Mini Qwen3 VL 235B(Q4) Gemini 2.0 Flash Qwen3 VL 235B(Q4, MRoPE) Gemini 2.0 Flash+ MARS Grok 4.1 Fast+ MARS Gemini 3 Pro Grok 4.1 Fast Gemini 2.5 Flash(Reasoning) GPT-5 Mini(Reasoning) Gemini 2.5 Flash Grok 4.1 Fast(Reasoning) Claude Sonnet 4.5 Claude Code(Sonnet 4.5) Gemini 3 Flash Gemini 3 Flash(Reasoning) 0.00.10.20.30.40.5 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Outlier Detection Nemotron 3 Nano 30B(BF16, Reasoning) Qwen 2.5 7B(1M) MedGemma 27B Gemma 3 27B gpt-oss 120B Gemini 2.5 Flash Lite GPT-5 Nano Qwen3 VL 30B(FP8) Qwen3 VL 235B(FP8) Gemma 3 27B+ MARS Nemotron 3 Nano 30B(FP8, Reasoning) Gemini 2.0 Flash+ MARS Nemotron 3 Nano 30B(BF16) Qwen3 VL 235B(Q4) Qwen3 235B(Q4) Qwen3 Next 80B(FP8) Nemotron 3 Nano 30B(FP8) Nemotron 3 Nano 30B(FP8, YaRN) Qwen3 VL 235B(FP8)+ MARS GPT-5 Mini Qwen3 Next 80B(FP8, Reasoning) Grok 4.1 Fast+ MARS Gemini 2.0 Flash+ LongCEPO Qwen3 VL 30B(FP8)+ MARS Gemini 2.5 Flash(Reasoning) Qwen3 VL 235B(Q4, Reasoning) Claude Sonnet 4.5 GPT-5 Mini(Reasoning) Grok 4.1 Fast(Reasoning) Gemini 3 Pro Qwen3 VL 235B(Q4, MRoPE) Qwen3 235B(Q4, YaRN, Reasoning) Grok 4.1 Fast Gemini 2.0 Flash MegaBeam Mistral 7B(512K) Gemini 3 Flash Gemini 2.5 Flash Gemini 3 Flash(Reasoning) Claude Code(Sonnet 4.5) 0.00.20.40.60.8 39 36 36 36 34 34 33 32 30 30 28 28 27 23 23 23 23 21 21 19 19 18 13 13 13 13 13 11 11 10 9 7 7 4 4 4 3 2 1 Two-SampleTest Qwen3 VL 30B(FP8)+ MARS Gemma 3 27B MedGemma 27B Gemini 2.0 Flash+ LongCEPO gpt-oss 120B Qwen 2.5 7B(1M) Gemini 2.0 Flash+ MARS Qwen3 VL 30B(FP8) Qwen3 VL 235B(FP8)+ MARS Qwen3 235B(Q4, YaRN, Reasoning) Gemini 2.5 Flash Lite Gemini 2.0 Flash Qwen3 235B(Q4) Qwen3 Next 80B(FP8, Reasoning) Qwen3 VL 235B(FP8) Qwen3 Next 80B(FP8) Qwen3 VL 235B(Q4, MRoPE) Qwen3 VL 235B(Q4) Nemotron 3 Nano 30B(FP8) Qwen3 VL 235B(Q4, Reasoning) Nemotron 3 Nano 30B(BF16) GPT-5 Nano Nemotron 3 Nano 30B(FP8, YaRN) MegaBeam Mistral 7B(512K) Grok 4.1 Fast GPT-5 Mini Gemini 3 Pro Gemini 2.5 Flash(Reasoning) Nemotron 3 Nano 30B(BF16, Reasoning) Nemotron 3 Nano 30B(FP8, Reasoning) Claude Sonnet 4.5 Grok 4.1 Fast(Reasoning) GPT-5 Mini(Reasoning) Grok 4.1 Fast+ MARS Gemini 2.5 Flash Claude Code(Sonnet 4.5) Gemini 3 Flash Gemini 3 Flash(Reasoning) 0.00.10.20.30.40.50.60.7 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Similarity Search Qwen3 VL 30B(FP8)+ MARS Gemma 3 27B MedGemma 27B Gemini 2.0 Flash+ LongCEPO Qwen3 Next 80B(FP8, Reasoning) gpt-oss 120B Qwen 2.5 7B(1M) Gemini 2.0 Flash+ MARS Qwen3 VL 235B(FP8)+ MARS Nemotron 3 Nano 30B(BF16) Gemini 3 Pro Gemini 2.0 Flash Nemotron 3 Nano 30B(FP8) Qwen3 VL 30B(FP8) Qwen3 Next 80B(FP8) Nemotron 3 Nano 30B(FP8, YaRN) GPT-5 Nano Nemotron 3 Nano 30B(BF16, Reasoning) Nemotron 3 Nano 30B(FP8, Reasoning) Qwen3 VL 235B(Q4) Qwen3 VL 235B(Q4, Reasoning) Qwen3 VL 235B(Q4, MRoPE) Qwen3 235B(Q4) Qwen3 235B(Q4, YaRN, Reasoning) MegaBeam Mistral 7B(512K) Qwen3 VL 235B(FP8) GPT-5 Mini Gemini 2.5 Flash Lite Grok 4.1 Fast(Reasoning) Grok 4.1 Fast Gemini 2.5 Flash(Reasoning) Claude Code(Sonnet 4.5) Grok 4.1 Fast+ MARS GPT-5 Mini(Reasoning) Gemini 3 Flash(Reasoning) Gemini 3 Flash Gemini 2.5 Flash Claude Sonnet 4.5 0.00.10.20.30.40.50.6 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Top-k Qwen3 VL 235B(Q4, Reasoning) Qwen3 VL 30B(FP8)+ MARS MedGemma 27B Gemini 2.0 Flash+ LongCEPO Qwen3 235B(Q4, YaRN, Reasoning) Gemma 3 27B Qwen3 VL 235B(FP8)+ MARS Nemotron 3 Nano 30B(FP8) MegaBeam Mistral 7B(512K) Gemini 2.0 Flash+ MARS Nemotron 3 Nano 30B(FP8, YaRN) Nemotron 3 Nano 30B(BF16) Qwen3 Next 80B(FP8, Reasoning) gpt-oss 120B Nemotron 3 Nano 30B(FP8, Reasoning) GPT-5 Nano Qwen 2.5 7B(1M) Qwen3 VL 30B(FP8) Nemotron 3 Nano 30B(BF16, Reasoning) Gemini 2.5 Flash Lite Gemini 2.0 Flash Qwen3 Next 80B(FP8) GPT-5 Mini Qwen3 VL 235B(FP8) Qwen3 VL 235B(Q4) Qwen3 235B(Q4) Qwen3 VL 235B(Q4, MRoPE) GPT-5 Mini(Reasoning) Gemini 2.5 Flash(Reasoning) Grok 4.1 Fast Claude Sonnet 4.5 Gemini 3 Pro Grok 4.1 Fast+ MARS Grok 4.1 Fast(Reasoning) Gemini 2.5 Flash Claude Code(Sonnet 4.5) Gemini 3 Flash Gemini 3 Flash(Reasoning) 0.00.20.40.60.8 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Sorting Gemini 2.0 Flash+ LongCEPO Gemma 3 27B+ MARS MedGemma 27B Gemma 3 27B Nemotron 3 Nano 30B(FP8, Reasoning) Nemotron 3 Nano 30B(FP8) Nemotron 3 Nano 30B(FP8, YaRN) Nemotron 3 Nano 30B(BF16, Reasoning) Nemotron 3 Nano 30B(BF16) MegaBeam Mistral 7B(512K) GPT-5 Nano gpt-oss 120B Qwen 2.5 7B(1M) Qwen3 VL 30B(FP8) Gemini 2.5 Flash Lite Qwen3 VL 30B(FP8)+ MARS Qwen3 Next 80B(FP8, Reasoning) Qwen3 235B(Q4, YaRN, Reasoning) GPT-5 Mini Gemini 2.0 Flash+ MARS Gemini 2.0 Flash Qwen3 VL 235B(Q4, Reasoning) Qwen3 VL 235B(Q4) Qwen3 VL 235B(FP8) Grok 4.1 Fast+ MARS Qwen3 VL 235B(FP8)+ MARS Qwen3 235B(Q4) Qwen3 Next 80B(FP8) Qwen3 VL 235B(Q4, MRoPE) Gemini 2.5 Flash(Reasoning) GPT-5 Mini(Reasoning) Gemini 2.5 Flash Grok 4.1 Fast(Reasoning) Grok 4.1 Fast Claude Sonnet 4.5 Claude Code(Sonnet 4.5) Gemini 3 Pro Gemini 3 Flash Gemini 3 Flash(Reasoning) 0.00.20.40.60.8 39 38 37 36 35 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 Classification Figure 6. Task-specific scoreboards. The y-axis indicates the model’s rank, and the x-axis denotes its mean score on the task. The filled-in bars with solid borders indicate the mean over all the questions in the task. The faint bars with dashed borders indicate the mean over just the questions the model answered properly (i.e. in the correct format). 10 The Verbose Context Problem in Medical Records === Patient ID: 1182 === Time: 1968-07-25 Birth Time: 2023-08-15 4256F - Anesthesia administration documented for l CPT Code 44140 - Colectomy, partial; with anastomo Excision of Head, Open Approach - This surgical pr Diverticulosis of Large Intestine without Perforat Partial Intestinal Obstruction (K56600) - This con Cefazolin - Cefazolin Sodium, manufactured by Gene 1 vial Morphine Sulfate Injection - Morphine Sulfate, man 10 ML Sodium Chloride Injection Solution - Sodium Chlori 1 bag Heart rate by Noninvasive 95 beats/min Systolic blood pressure mean 135 mmHg Diastolic blood pressure mean 82 mmHg Oxygen saturation in Arterial blood by Pulse oxime 97 % Thyrotropin [Units/volume] in Serum or Plasma --ba 37 C Postoperative state - This term refers to the peri General anesthesia - A medical procedure used to i Postoperative monitoring - This refers to the syst <... elided for figure ...> Time: 2023-08-16 Heart rate by Noninvasive 105 beats/min Systolic blood pressure mean 110 mmHg Oxygen saturation in Arterial blood by Pulse oxime 94 % Thyrotropin [Units/volume] in Serum or Plasma --ba 38.1 C Prostate Tumor Incidental Histologic Finding - Thi Postoperative period - This term refers to the rec Bowel Sounds Quiet - This condition refers to the <... elided for figure ...> === Patient ID: 1182 === Time: 1968-07-25 Birth Time: 2023-08-15 cpt//4256F cpt//44140 icd10_proc//0WB00Z icd10//K57.31 icd10//K56.600 ndc//52584-924 1 vial ndc//0409-1134 10 ML ndc//0264-1800 1 bag loinc//76477-9 95 beats/min loinc//96608-5 135 mmHg loinc//96609-3 82 mmHg loinc//59408-5 97 % loinc//14999-7 37 C snomed//19585003 snomed//50697003 snomed//182775008 <... elided for figure ...> Time: 2023-08-17 loinc//76477-9 112 beats/min loinc//96608-5 105 mmHg loinc//59408-5 92 % loinc//14999-7 38.8 C snomed//406127006 snomed//101379003 snomed//207206002 <... elided for figure ...> Figure 7. Different prompting methods on the same patient record. (Left): the standard prompting method used as a baseline in this paper. It replaces the code altogether (since those are often not recognized by language models) by a truncated description of the code. (Right): a less verbose (and less informative) prompting method which includes the code but not the description. Codebook (integer IDs assigned to medical codes): #1: Birth of patient #2: (cpt//0513T) 0513T - External shockwave, integrated wound healing, each additional wound. #3: (cpt//1021887) Skilled Nursing Facility #4: (cpt//1111F) CPT Code: 1111F - Discharge medication reconciliation with current medication review. <... elided for figure ...> #313: (snomed//91251008) Physical Therapy Procedure - A therapeutic intervention aimed at alleviating physical impairments and enhancing mobility and fun === Patient ID: 1182 === Time: 1968-07-25 #1 Time: 2023-08-15 #17 #20 #111 #77 #76 #171 1 vial #160 10 ML #157 1 bag #144 95 beats/min #152 135 mmHg #153 82 mmHg #137 97 % #124 37 C #222 #295 <... elided for figure ...> Time: 2023-08-17 #144 112 beats/min #152 105 mmHg #137 92 % #124 38.8 C #280 #187 #225 <... elided for figure ...> Figure 8. An alternative prompting method which attempts to eliminate redundancy across groups of patient records. It defines a succinct ID numbers for all unique codes (across all patients), and then references those IDs within the subsequent patient records. 11 The Verbose Context Problem in Medical Records B. Formalization of the Verbose Context Problem We consider heterogeneous promptsx : list(string +C)for LLMs.Cis a set of concepts distinguished from the other unstructured parts of the prompt. As the canonical example in this paper, we takeCto be the set of all medical codes. The verbose context problem arises when convertingc∈Cto strings. Letstr(c)be a string which conveyscto a target language model with the desired level of precision. For a medical code, it may not be sufficient to pass just the abbreviated code; instead, much of the extended description may be needed as well. To be processed by the language model, the string must be tokenized, and then each token’s embedding is looked up in the embedding matrix. This produces a variable-length sequence of vectorsemb◦ tok◦ str(c)∈ listE, whereEis the LLM embedding space, which is typically a few thousand dimensions. Given a probability distribution over prompts x, the baseline (overall) expected context length is N . M is the context length that originates fromC. M =E x X c∈x∩C Length(emb◦ tok◦ str(c)) N =E x Length(emb◦ tok◦ str(x)) (In the second line, we slightly abuse notation to havestr(x)operate on all parts of the prompt). Informally, the verbose context problem is that N is too long to be practical. Definition B.1 (Verbose Context Problem). This occurs when M/N is large, i.e. Ω(1) as N →∞. C. Task Scoring TaskMetricDetails Planted CliquePrecision|Predicted ∩ True|/k In-Context ClassificationAccuracyBinary classification accuracy on the test set ClusteringRand IndexProportion of correctly identified pairs (same vs. different) Inverse-Propensity Weight- ing Absolute ErrorAbsolute difference between predicted and true ATE (Hajek) Outlier DetectionF 1 ScoreHarmonic mean of precision and recall Two-Sample TestAccuracy0–1 accuracy in identifying if distributions differ Similarity SearchPrecision@kProportion of model’s top-k that are in the true top-k Top-kPrecision@kProportion of model’s top-k that are in the true top-k SortingKendall’s τRestricted to pairs from different ground-truth cohorts ClassificationAccuracyMulti-class accuracy across all categories 12 The Verbose Context Problem in Medical Records D. Related Work Long-Context Benchmarks. Since the advent of modern language models, dozens of long-context benchmarks have been developed (Tay et al., 2021; Bai et al., 2023; Zhang et al., 2024; Yen et al., 2025; Bai et al., 2025). Context lengths under evaluation have increased from 4K to above 1M. Different evaluations attempt to isolate different failure modes. Early diagnostic tests focus on retrieval failures and positional bias (e.g. “lost in the middle”). Newer evaluations target reasoning degradation in multi-step tasks (OpenAI, 2025a). In these benchmarks, difficulty and context length are driven primarily by (1) the amount and density of irrelevant distractors, and (2) the number of reasoning hops required to bridge dispersed information (Vodrahalli et al., 2024; OpenAI, 2025b). PopMedQA emphasizes a different cause of context bloat (verbosity) in order to expose different failure modes. While some benchmarks isolate long-context reasoning through abstract tasks (OpenAI, 2025a), others prioritize naturalistic, real-world inquiries (Bai et al., 2023; Artificial Analysis, 2025). Our work belongs to the latter category. We encourage more situated, concrete study of long-context problems by recognizing and addressing their domain-specific aspects. We aim to further close the gap between benchmark performance and real-world utility. AI on EHRs. Existing benchmarks for AI in Electronic Health Records (EHRs) evaluate a wide range of clinical and administrative capabilities. For structured data, EHRSHOT (Wornow et al., 2023) and INSPECT (Huang et al., 2023) assess few-shot clinical prediction and algorithmic fairness within individual patient timelines. EHRSQL (Lee et al., 2022) and MedAgentBench (Jiang et al., 2025) extend these evaluations to cohort-level queries; however, these frameworks primarily test the model’s ability to translate natural language into structured queries (SQL or FHIR) that delegate computational aggregation to an external database engine. For unstructured text, MedAlign (Fleming et al., 2024) and MedFactEval (Grolleau et al., 2026) focus on instruction-following and factuality within clinical notes. PopMedQA diverges from these approaches by shifting the analytical focus to population health, requiring models to perform holistic, in-context reasoning across the raw longitudinal records of cohorts of 10 to 50 patients simultaneously. This framework unlocks complex use cases in population-level pattern discovery, such as identifying latent clinical clusters or detecting sparse anomalies across disparate patient trajectories, that cannot be readily addressed by standard query-based aggregation or individual-level processing. Population Health Analytics. Population health analytics focuses on the health outcomes of groups of individuals and the distribution of these outcomes within the group (Kindig & Stoddart, 2003). Its primary objectives are to quantify disease burden and guide resource allocation to ensure equitable and efficient healthcare delivery. To achieve this, established risk adjustment systems like the Johns Hopkins Adjusted Clinical Group (ACG) System (Weiner et al., 2012), the CMS Hierarchical Condition Category (CMS-HCC) model (Pope et al., 2004), and comorbidity indices such as Charlson (Charlson et al., 1987) and Elixhauser (Elixhauser et al., 1998) are widely employed. These tools primarily rely on rule-based aggregation of structured diagnosis codes and pharmacy data to perform retrospective financial risk stratification and predict healthcare utilization. However, these statistical frameworks are often limited by fragile or manual feature engineering that cannot capture the complex dependencies within a patient’s history, leading to low individual-level predictive accuracy (e.g.,R 2 values frequently below 15% for prospective cost prediction) (Pope et al., 2004). PopMedQA evaluates a more flexible, yet still code-centric, alternative where modern language models perform comprehensive longitudinal reasoning over groups of patient records across entire cohorts. Alternative Concept Representations. Are there more succinct ways to convey concepts to language models than language itself? Multiple lines of work support this general idea. The common strategy is to inject concepts as vectors at different model layers, in lieu of providing more input text. In prefix tuning (Li & Liang, 2021), the output embeddings of the earlier part of a prompt are truncated and directly optimized to improve the accuracy of subsequent generation. Cartridges (Eyuboglu et al., 2025) also condense the prefix by optimizing the contents of key-value caches in attention layers. Steering vectors (Jahanian et al., 2020; Subramani et al., 2022) are added to activations to control generation in an input-agnostic manner. Rendering text to images and using vision language models can be effective for long-context inference (Zheng et al., 2024; Cheng et al., 2025; Wei et al., 2025). 13