Paper deep dive
The Validity Gap in Health AI Evaluation: A Cross-Sectional Analysis of Benchmark Composition
Alvin Rajkomar, Pavan Sudarshan, Angela Lai, Lily Peng
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 6:01:58 AM
Summary
This paper identifies a 'validity gap' in health AI evaluation by analyzing 18,707 consumer health queries across six benchmarks. It reveals that current benchmarks are misaligned with real-world clinical needs, showing a systemic lack of raw clinical artifacts, vulnerable population representation, and longitudinal chronic care scenarios, while being heavily skewed toward low-acuity wellness data.
Entities (5)
Relation Signals (3)
HealthSearchQA â classifiedas â Generation 1
confidence 95% ¡ Generation 1: The Search Era â Fact Retrieval (HealthSearchQA, MashQA).
MedRedQA â classifiedas â Generation 2
confidence 95% ¡ Generation 2: The Case Presentation Era â Diagnostic Reasoning (MedRedQA).
HealthBench Main â classifiedas â Generation 3
confidence 95% ¡ Generation 3: Interactive and Data-Linked Queries â Interactive Dialogue and Wellness Coaching (HealthBench Main; GoogleFitbit).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Clinical trials rely on transparent inclusion criteria to ensure generalizability. In contrast, benchmarks validating health-related large language models (LLMs) rarely characterize the "patient" or "query" populations they contain. Without defined composition, aggregate performance metrics may misrepresent model readiness for clinical use. Methods: We analyzed 18,707 consumer health queries across six public benchmarks using LLMs as automated coding instruments to apply a standardized 16-field taxonomy profiling context, topic, and intent. Results: We identified a structural "validity gap." While benchmarks have evolved from static retrieval to interactive dialogue, clinical composition remains misaligned with real-world needs. Although 42% of the corpus referenced objective data, this was polarized toward wellness-focused wearable signals (17.7%); complex diagnostic inputs remained rare, including laboratory values (5.2%), imaging (3.8%), and raw medical records (0.6%). Safety-critical scenarios were effectively absent: suicide/self-harm queries comprised <0.7% of the corpus and chronic disease management only 5.5%. Benchmarks also neglected vulnerable populations (pediatrics/older adults <11%) and global health needs. Conclusions: Evaluation benchmarks remain misaligned with real-world clinical needs, lacking raw clinical artifacts, adequate representation of vulnerable populations, and longitudinal chronic care scenarios. The field must adopt standardized query profiling--analogous to clinical trial reporting--to align evaluation with the full complexity of clinical practice.
Tags
Links
- Source: https://arxiv.org/abs/2603.18294v1
- Canonical: https://arxiv.org/abs/2603.18294v1
Trouble viewing inline? Open PDF directly â
Full Text
81,664 characters extracted from source content.
Expand or collapse full text
The Validity Gap in Health AI Evaluation: A Cross-Sectional Analysis of Benchmark Composition Alvin Rajkomar, M.D. 1,â , Pavan Sudarshan, B.S. 1 , Angela Lai, M.S.E. 1 , Lily Peng, M.D.,Ph.D. 1 March 17, 2026 Abstract Background: Clinical trials rely on transparent inclusion criteria to ensure generalizabil- ity. In contrast, benchmarks validating health-related large language models (LLMs) rarely characterize the âpatientâ or âqueryâ populations they contain. Without defined composition, aggregate performance metrics may misrepresent model readiness for clinical use. Methods: We analyzed 18,707 consumer health queries across six public benchmarks. We utilized large language models as automated coding instruments to apply a standardized 16-field taxonomy profiling context, topic, and intent, enabling characterization of the evolution from static search (âGeneration 1â) to interactive care (âGeneration 3â). Results: We identified a structural âvalidity gapâ in evaluation paradigms. While bench- marks have evolved from static retrieval to interactive dialogue, clinical composition remains misaligned with real-world information needs. Although 42% of the corpus referenced objective data, this was polarized toward wellness-focused wearable signals (17.7%); complex diagnostic inputs remained rare, including explicit laboratory values (5.2%), imaging results (3.8%), and raw excerpts from medical records (0.6%). This scarcity limits validation of clinical reasoning. Crucially, safety-critical scenarios were effectively absent: high-acuity behavioral health queries (e.g., suicide/self-harm) comprised <0.7% of the corpus, and chronic disease managementâa core function of primary careâconstituted only 5.5%. Finally, benchmarks exhibited systemic neglect of vulnerable populations (pediatrics/older adults <11%) and global health needs, fa- voring a âstandard adultâ default that obscures physiological and geographic complexity. Conclusions: While model capabilities have evolved to support interactive dialogue and data interpretation, evaluation benchmarks remain compositionally misaligned with real-world clinical needs. Even modern interactive datasets exhibit systemic omissionsâlacking raw clin- ical artifacts, adequate representation of vulnerable populations, and longitudinal chronic care scenarios. To ensure clinical reliability, the field must adopt standardized query profilingâ analogous to clinical trial reportingâto align evaluation datasets with the full complexity of clinical practice. 1 Introduction Consumers have long turned to the internet for health information 1,2 , but the emergence of publicly available large language models (LLMs) 3 represents a new source of personalized guidance 4â6 . To rigorously evaluate the quality of LLM-generated health answers, model builders rely on âbench- marksââdatasets containing canonical user queries 7 . Yet just as the generalizability of a clinical â1 Apple â Corresponding author: Alvin Rajkomar can be contacted at alvinr@apple.com 1 arXiv:2603.18294v1 [cs.AI] 18 Mar 2026 trial depends on the composition of its enrolled participants 8 , we argue that the validity of ex- trapolating benchmark performance to real-world use depends critically on the composition of the benchmark prompts themselves. Benchmark composition profoundly affects model evaluation because different query types stress different capabilities 9â11 . For example, a model achieving 85% accuracy on a benchmark domi- nated by simple symptom lookups (e.g., âWhat causes headaches?â) may perform far worse on complex clinical reasoning tasks requiring differential diagnosis or treatment planning. Without compositional transparency, such performance gaps remain hidden from standard accuracy metrics, potentially leading to misleading claims about model readiness for use. When deployed clinically, such hidden performance gaps can translate to patient harm: missed diagnoses when models fail complex reasoning, false reassurance in acute presentations, or inappropriate triage advice that delays necessary care. Unlike clinical trialsâwhich report participant characteristics through standardized inclusion criteriaâhealth LLM benchmarks lack frameworks for describing the types of queries they contain. This absence of compositional reporting is akin to publishing trial results without describing who was enrolled. The need for benchmark transparency aligns with emerging dataset documentation frameworks such as Data Nutrition Labels 12 and Dataset Statements 13 , which emphasize reporting collection processes, data distributions, and intended use cases. Query profiling extends this transparency to the functional characteristicsâintent, clinical complexity, and context richnessâthat directly shape model behavior and determine appropriate deployment contexts. Health queries span a wide range of tasks, from diagnosis and treatment planning to fitness tracking, wellness advice, preventive care, and chronic disease management. Even within a single domain, structural complexity varies substantially: a concise question such as âDo I have diabetes?â poses a fundamentally different evaluation challenge than a detailed symptom timeline accompanied by laboratory values, imaging reports, or excerpts from medical records requesting differential diagnosis. Differences in information-seeking goalsâsymptom explanation, treatment guidance, triage dispositionâfurther stratify model requirements. We introduce query profiling: a standardized, model-agnostic framework for characterizing con- sumer health queries across three domainsâcontext (structural properties and information rich- ness), topic (clinical domain and conditions), and intent (information-seeking goals). Applying this 16-dimension taxonomy to six widely used benchmarks spanning 18,707 consumer queries reveals substantial compositional heterogeneity that reflects three generations of health information-seeking behavior: search engine queries (Generation 1), social media discussions (Generation 2), interactive LLM dialogue (Generation 3a), and data-augmented LLM insights (Generation 3b). This work makes three contributions: (1) a standardized query profiling taxonomy applied across six heterogeneous health benchmarks, (2) systematic characterization demonstrating that these benchmarks capture different slices of the health-query landscape with notable gaps in pre- vention, chronic care, and complex clinical data integration, and (3) open-source tagging tools and tagged datasets enabling reproducible query profiling across research groups. Understanding these compositional patterns is essential for interpreting benchmark performance, designing evaluation datasets aligned with real-world use, and assessing whether models are being tested on the types of queries they are ultimately intended to support. 2 2 Methods 2.1 Study Design and Taxonomy Development We performed a cross-sectional analysis of publicly available health query datasets to characterize the structural and clinical composition of benchmarks used to evaluate large language models (LLMs). To standardize this characterization, we developed a âQuery Profile,â a 16-field taxonomy designed to capture three dimensions of health information-seeking behavior: ⢠Context: Structural properties, including conversation depth (single query versus back-and- forth conversation), population characteristics, presence of objective data (e.g., laboratory values, vitals, clinical artifacts), and information richness. ⢠Topic: The primary clinical domain and specific medical conditions referenced. ⢠Intent: The userâs information-seeking goal, classified into nine categories ranging from factual education to triage and chronic disease management. In the absence of a standard ontology for characterizing health LLM queries, we synthesized a pragmatic taxonomy informed by clinical experience and commercial health information-serving ontologies. Traditional clinical ontologies such as ICD or SNOMED were not suitable because they presuppose confirmed diagnoses and clinical context that consumer queries typically lack; our taxonomy instead captures the limited, often ambiguous information available from a non-expertâs perspective. This framework represents one reasonable approach; alternative categorizations may be appropriate for specialized clinical domains. The full taxonomy and definitions are provided in Table 1. 3 Table 1: Query Profile Taxonomy: Elements and Descriptions ElementDescription Context Dimensions PopulationAge-specific demographic (pediatric, adult, older adult), derived from user-provided context cues. Conversation TypeSingle question versus multi-turn conversation. Length DetailShort text string versus extended narrative. Objective DataPresence of health data, including laboratory results, imaging, vitals (basic or wearable), diagnoses, medica- tions, procedures, and clinical artifacts (e.g., excerpts from medical records). Context DepthInformation richness (age, timeline, clinical anchors). Terminology LevelLay versus technical terminology. Clarification NeededSelf-contained versus ambiguous query (e.g., âibupro- fen dose for adultsâ vs. âibuprofen infoâ). SettingImplied clinical setting (outpatient, inpatient, emer- gency). User TypeConsumer versus healthcare professional. LanguageEnglish versus non-English. RegionGeographic or regulatory context. Topic Dimensions Topic Area PathHierarchical classification by body system or condi- tion. Key ConditionsHigh-priority health conditions as defined by global health agencies. SpecialtyClinical specialty best aligned with the query. Intent Dimensions Intent (Top-level)Primary information-seeking goal (9 categories). Intent (Sub-level)Subcategory of intent (e.g., treatment, prognosis). Risk SensitivityUrgency and severity implied by the query. 2.2 Data Sources We analyzed six widely cited public benchmarks containing consumer-facing health queries (N = 20, 034). These datasets were selected to represent the diverse sources of health inquiries currently used to train or evaluate AI models: ⢠Search Engine Queries: HealthSearchQA 14 (N = 3, 173) and MashQA Test (N = 3, 491), comprising short, atomic questions typical of web-based search. ⢠Online Medical Forums: MedRedQA Test 15 (N = 5, 099), consisting of user-authored narratives posted to physician-facing public forums, often containing detailed medical history. ⢠Simulated Interactive Dialogue: HealthBench Main 16 (N = 5, 000), representing multi- turn interactions between simulated users and AI agents. ⢠Wearable Data Streams: GoogleFitbit Sleep (N = 1, 521) and GoogleFitbit Fitness (N = 4 1, 750) 17 , representing inquiries derived from continuous biometric data streams rather than explicit text prompts. All datasets were de-identified and publicly available; the study was exempt from institutional review board approval. 2.3 Classification and Validation To ensure consistent application of the taxonomy across the large corpus, we utilized GPT-5.2 (OpenAI) as a standardized automated coding instrument, a methodology previously employed in large-scale analyses of LLM queries 2 . The model was provided with strict definitions, controlled vocabularies, and deterministic decision rules to classify each query across the 16 taxonomic dimen- sions (full prompt provided in Supplementary Material). For the wearable data datasets (Google- Fitbit), where queries are implicitly defined by structured data (JSON), a rule-based programmatic classifier was used to ensure accuracy. To validate the reliability of the automated classification, we conducted a cross-model agreement analysis comparing GPT-5.2 with Claude Opus 4.5 (Anthropic) on the HealthBench Main dataset (N = 5, 000). These models were selected because they represent different architectures developed by independent organizations, providing stronger validity evidence than same-model reproducibility. Both models independently tagged identical queries using the same v4.6 tagging prompt. Of the 5,000 queries, 33 (0.7%) were not tagged by Opus 4.5 due to content filtering on sensitive clinical scenarios and were excluded from the comparison. Overall agreement across all 21 dimensions averaged 90.0%, with Cohenâs Îş = 0.77 indicat- ing substantial agreement. Structural dimensions (conversation structure, language, user type) achieved near-perfect concordance (>98%), while more subjective dimensions (intent classification, risk sensitivity) showed moderate agreement (73â88%) comparable to human inter-rater reliability in medical annotation tasks. Chi-square tests revealed no statistically significant distributional differences between models for any dimension, confirming that aggregate population-level insights would be consistent regardless of which model performed the tagging. Full analysis is provided in Supplementary Material. 2.4 Flow of Data A CONSORT-style diagram (Figure 1) summarizes dataset selection, exclusions, and the final analytic sample. 2.5 Statistical Analysis Descriptive statistics were used to characterize the distribution of context, topic, and intent across the datasets. We analyzed the prevalence of specific clinical needsâsuch as preventive care, chronic disease management, and high-acuity triageâto assess the alignment between benchmark composi- tion and real-world clinical complexity. This reporting follows STROBE guidelines for observational research. All analyses used publicly available, de-identified datasets; no human subjects were in- volved. 5 3 Results 3.1 Study Population and Cohort Characteristics Across the six benchmarks, we analyzed 20,034 total queries. After excluding 1,327 queries classified as originating from healthcare professionals, the final analytic cohort consisted of 18,707 consumer- facing health queries (93.4% of the initial sample) (Table 2). The study population was predominantly adult (93.6%) and English-speaking (96.4%). Pediatric- and older-adultâfocused queries represented 10.9% of the corpus. Most queries (94.9%) were phrased in lay language rather than technical terminology. However, these aggregate descriptive statistics mask substantial structural variability across datasets. Query Profile analysis revealed that the benchmarks cluster into three distinct âgenerationsâ reflecting the evolution of consumer health information-seeking behavior (Figure 2). 3.2 Evolution of Clinical Complexity: Three Generations The six benchmarks aligned with three generational patterns that differed systematically in health topic coverage (Figure 3A) and query intent (Figure 3B). Generation 1: The Search Era â Fact Retrieval (HealthSearchQA, MashQA). These benchmarks comprised 35.4% of the corpus (Table 2). Queries were short (median 7â8 words), single-turn, and contained minimal contextual detail. A distinct disparity in clinical intent was ob- served: while this generation dominated the corpus in education queries (83% in HealthSearchQA), triage intent was effectively absent (1.2% in HealthSearchQA vs. 12.7% in Generation 2). Conse- quently, these benchmarks evaluated factual retrieval but provided virtually no exposure to queries requiring the discrimination of medically urgent situationsâincluding emergencies, rapidly evolv- ing symptoms, or conditions requiring timely clinical attentionâfrom benign symptoms. For de- ployment decisions, Generation 1 benchmarks can validate simple question-answering but provide insufficient evidence for claims about triage accuracy or clinical safety in acute presentations. Generation 2: The Case Presentation Era â Diagnostic Reasoning (MedRedQA). Accounting for 27.0% of all queries, MedRedQA showed the greatest contextual depth and risk profile. Most queries contained specific patient anchors (age, duration, medications), and intent categories shifted toward management and triage. Unlike Generation 1, this generation incorpo- rated the vast majority of the corpusâs high-risk scenarios. For example, queries related to suicidal ideation (N = 66 of 77 total) and self-harm (N = 30 of 34 total) were concentrated almost ex- clusively in this generation, with virtually no representation in Generation 1. These benchmarks are well-suited for evaluating diagnostic reasoning when clinical context is pre-provided in user- synthesized narratives, but do not test a modelâs ability to interpret raw clinical data or actively elicit missing information through follow-up questions. Generation 3: Interactive and Data-Linked Queries â Interactive Dialogue and Well- ness Coaching (HealthBench Main; GoogleFitbit). This generation introduced multi-turn conversations (HealthBench Main) and high-volume structured data (GoogleFitbit). However, the clinical utility of the structured data was limited by topic: while Fitbit queries contributed large data payloads (median length >400 words), they were overwhelmingly focused on low-risk lifestyle and wellness guidance rather than clinical pathology. HealthBench Main validates multi-turn clin- ical dialogue including high-risk scenarios, but as detailed below, behavioral health crises remain 6 virtually untested in conversational settings. GoogleFitbit validates data integration but only for wellness coaching rather than clinical pathology. 3.3 The Validity Gap: Mismatches with Clinical Reality Despite the structural evolution across generations, we identified fundamental gaps between bench- mark composition and real-world clinical needs. Prevalence of Longitudinal and Preventive Care Topics. Primary care topics involving longitudinal management and prevention were infrequent compared to episodic concerns. Chronic disease management accounted for 5.5% of all query topics (Table S2). Specifically, queries related to high-prevalence conditions such as diabetes (N = 244; 1.3%), hypertension (N = 139; 0.7%), and obesity (N = 40; 0.2%) combined represented less than 3% of the corpus. Similarly, pre- ventive health screening queries constituted 2.7% of the cohort (N = 513). Within this category, vaccination inquiries (N = 246; 1.3%) and requests for screening schedules (N = 85; 0.5%) were rare (Table S1). In contrast, intents related to acute symptom checking (N = 4, 523; 24.2%) and general health education (N = 5, 473; 29.3%) comprised the majority of the dataset. Global Health and Infectious Disease Bias. The representation of infectious diseases exhib- ited significant geographic and recency biases (Table S3). While queries related to COVID-19 were relatively frequent (N = 341; 1.80%), historically high-burden global pathogens were effectively ig- nored. Malaria (N = 41; 0.22%) and tuberculosis (N = 27; 0.14%)âconditions responsible for millions of deaths annually in low- and middle-income countriesâappeared in negligible num- bers. This omission implies that current benchmarks predominantly validate model performance for health contexts in high-income regions, leaving their safety for global health deployment unverified. Underrepresentation of Vulnerable Populations. Benchmark composition deviated signif- icantly from real-world demographic usage patterns. Queries explicitly referencing older adults (>65 years)âa high-utilization group in clinical settingsâcomprised only 4.6% (N = 861) of the dataset. Similarly, pediatric topics were scarce, with queries concerning children (infants to ado- lescents) representing just 6.4% (N = 1, 190) of the total cohort. This demographic skew suggests that current benchmarks may not adequately evaluate model performance for populations requiring specialized clinical reasoning, from pediatric weight-based dosing to geriatric risk stratification. Disparities in Objective Data Types. Objective data appeared in 42.3% of queries (N = 7, 921), but the distribution was heavily skewed toward non-clinical sources (Table 2). The single largest category was wearable vitals (N = 3, 307; 17.7% of total), representing low-acuity signals such as step counts and sleep logs. In contrast, core clinical diagnostics were less prevalent. Diag- noses appeared in 15.5% (N = 2, 891) and medications in 12.4% (N = 2, 322) of queries. Explicit laboratory results appeared in only 5.2% (N = 981) and imaging in 3.8% (N = 702). Clinical artifactsâraw text copied directly from electronic health records such as provider notes or test reportsâwere rare (N = 110; 0.6%). Notably, 12.5% of queries (N = 2, 344) contained multiple objective data typesâpredominantly in MedRedQA (91.5% of multi-type queries)âleaving most datasets reliant on single data sources or none at all. Absence of Critical Behavioral Health Scenarios. Despite their high societal importance, queries involving crisis intervention were effectively absent. Scenarios involving self-harm (N = 34; 7 0.18%) and suicidal ideation (N = 77; 0.41%) combined represented less than 0.7% of the corpus. Critically, these high-acuity queries were virtually nonexistent in the interactive âGeneration 3â datasets, leaving the safety mechanisms of conversational agents for crisis de-escalation untested within these benchmarks in back-and-forth conversational settings. Assessment of Clinical Risk. Overall, 72.7% (N = 13, 605) of queries were classified as low-risk (Table 2). High-risk prompts requiring triage or urgent safety guidance were concentrated almost entirely in MedRedQA and HealthBench Main. Excluding these Generation 2 and 3 datasets from an evaluation pipeline would eliminate nearly all exposure to clinically sensitive scenarios, skewing safety assessments toward the low-acuity âGeneration 1â content. To promote transparency and reproducibility, we have open-sourced an automated analysis and visualization framework that generates standardized dataset summaries (âQuery Profile Factsâ). This system provides consistent profiling of context, intent, and topic distributions for any new dataset, enabling rapid comparison and extension of the benchmark suite. The automated tool ensures that subsequent datasets can be analyzed and visualized using the same standardized representations used throughout this study. 4 Discussion Across 18,707 consumer-facing queries, we identified a fundamental âvalidity gapâ in health AI eval- uation 9 . While clinical AI models are increasingly deployed for âGeneration 3â tasksârequiring interactive reasoning over complex patient historiesâcurrent benchmarks fail to simulate this clin- ical reality. We found that evaluations remain dominated by static, low-context search queries, while even modern interactive datasets suffer from systemic clinical omissions: lacking the raw clinical artifacts (e.g., notes, imaging), vulnerable populations, and longitudinal chronic care sce- narios required to test true data synthesis. This compositional mismatchâakin to validating a new therapeutic in a healthy population and extrapolating safety to patients with multimorbidityârisks overstating model readiness for complex clinical care. This misalignment has direct implications for clinical safety 18 . Our analysis reveals that core components of primary care are effectively absent from standard evaluations: chronic disease management and preventive care constituted only 5.5% and 2.6% of queries, respectively. This safety blind spot extends to high-acuity behavioral health: queries regarding self-harm and suicidal ideation were not only rare (< 0.7% of the corpus) but were virtually nonexistent in multi-turn con- versational datasets. This leaves the safety of AI-mediated crisis de-escalation effectively untested, a critical oversight given the increasing deployment of chatbots for mental health support. 19 Furthermore, while objective data appeared in 42% of queries, this figure masks a structural polarization that limits clinical validity. The largest category was wearable vitals (17.7% of the total corpus), consisting of high-frequency signalsâsuch as sleep logsâthat focus on wellness coaching rather than clinical pathology. In contrast, explicit laboratory values appeared in only 5.2% of queries, significantly lower than wearable data. Dense clinical data with multiple objective data types (12.5%) was largely confined to âGeneration 2â benchmarks, which provide pre-formulated narratives containing all relevant findings upfront. While interactive âGeneration 3â benchmarks do capture back-and-forth dialogue, their inclusion of core diagnostic anchors remains sparse. This structural gap implies that we are not sufficiently validating a modelâs ability to actively elicit missing diagnostic data: benchmarks either provide the answers in advance (Generation 2) or lack the clinical density to require complex investigation (Generation 3). Finally, the validity gap extends to the patients represented in the queries, revealing a systemic 8 neglect of vulnerable and global populations. Our results show that queries specific to pediatrics and older adults are marginal (combined < 11% of the corpus), implying that models are be- ing validated on a standard adult default that ignores the distinct physiological profiles, dosing constraints, and multimorbidity patterns of these groups. This bias is further compounded by a geographic skew: while recent Western health concerns like COVID-19 were well-represented, high- burden global pathogens such as malaria and tuberculosis were effectively absent. By training and evaluating on such narrow demographic slices, the field risks developing tools that are safe only for the healthy, adult populations of high-income regions. To close this measurement boundary, we propose a reporting framework analogous to the CON- SORT guidelines for clinical trials 20 . We suggest that a standardized âQuery Profileâ must ac- company future health AI benchmarks, explicitly reporting five dimensions: (1) Clinical Topic Coverage (e.g., balance of oncology, primary care, or behavioral health); (2) Intent Distribution (e.g., triage vs. education); (3) Context Richness (prevalence of clinical anchors like age and medications); (4) Clinical Complexity (acuity and risk levels); and (5) Data Integration (pres- ence of labs, vitals, or imaging). Such standardization would allow clinicians to determine if a model has been evaluated on queries resembling their specific practice environment. Beyond transparency, query profiling serves as a roadmap for benchmark development. By identifying underrepresented cells in the taxonomyâsuch as longitudinal chronic care management and high-acuity behavioral health crisesâthis framework highlights where targeted data collection is most urgently needed. To facilitate the widespread adoption of compositional transparency, we have released an open- source âQuery Profilingâ toolkit (available at [GitHub Link]). This automated pipeline takes an annotated dataset as input and generates standardized structural and clinical metadata. The detailed profiles presented in the Supplementary Appendix (Tables S4âS9) were generated using this workflow and serve as a template for how future benchmarks could be reported. By automating this characterization, we aim to lower the barrier for researchers to report âinclusion criteriaâ for their evaluation datasets, effectively creating a âTable 1â for health AI benchmarks. Our study has limitations. Automated tagging using GPT-5.2 may miss subtle nuances de- spite validation against human review. However, this trade-off is intentional: by removing the dependency on human annotation, we provide a scalable, privacy-preserving framework that allows researchers to profile large datasets without the financial burden of manual labeling or the risk of exposing sensitive health queries to human raters. Additionally, our taxonomy represents one pragmatic framework for characterizing consumer health queries; alternative categorizations em- phasizing different clinical dimensions (e.g., specialty-specific ontologies) may reveal complementary insights. We characterized benchmark composition rather than model performance 21 ; future work must quantify performance changes as models transition from low-context to data-rich tasks. Sim- ilarly, future research should extend this profiling to provider-facing benchmarks, where queries involving electronic health record integration and complex clinical decision support may demand specialized intent categories. Finally, our taxonomy is cross-sectional. As AI integrates deeper into care, longitudinal benchmarks that capture evolving patient states over multi-session interactions will become essential. Conclusions While model capabilities have evolved to support interactive dialogue and data interpretation, evaluation benchmarks remain compositionally misaligned with real-world clinical needs. Even modern interactive datasets exhibit systemic omissionsâlacking raw clinical artifacts, adequate representation of vulnerable populations, and longitudinal chronic care scenarios. To ensure clinical reliability, the field must adopt standardized query profilingâanalogous to clinical trial reportingâ 9 to align evaluation datasets with the full complexity of clinical practice. References [1] Fox S, Duggan M. Health Online 2013. Pew Research Center. 2013 Jan. [2] Chatterji A, Cunningham T, Deming DJ, Hitzig Z, Ong C, Shan CY, et al. How People Use ChatGPT. NBER Working Paper Series. 2025 Sep. [3] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al.. Attention Is All You Need. arXiv; 2023. [4] Rosenbluth T, Astor M. Empathetic, Available, Cheap: When A.I. Offers What Doctors Donât. The New York Times. 2025 Nov. [5] Introducing ChatGPT Health; 2026. https://openai.com/index/introducing-chatgpt-health/. [6] AdvancingClaudeinHealthcareandtheLifeSciences;. https://w.anthropic.com/news/healthcare-life-sciences. [7] Chang Y, Wang X, Wang J, Wu Y, Yang L, Zhu K, et al.. A Survey on Evaluation of Large Language Models. arXiv; 2023. [8] Rothwell PM. External Validity of Randomised Controlled Trials: âTo Whom Do the Results of This Trial Apply?â. Lancet (London, England). 2005â7;365(9453):82-93. [9] Bedi S, Liu Y, Orr-Ewing L, Dash D, Koyejo S, Callahan A, et al. Testing and Evaluation of Health Care Applications of Large Language Models: A Systematic Review. JAMA. 2025 Jan;333(4):319-28. [10] Yan LKQ, Niu Q, Li M, Zhang Y, Yin CH, Fei C, et al.. Large Language Model Benchmarks in Medical Tasks; 2024. [11] Bean AM, Kearns RO, Romanou A, Hafner FS, Mayne H, Batzner J, et al.. Measuring What Matters: Construct Validity in Large Language Model Benchmarks. arXiv; 2025. [12] Holland S, Hosny A, Newman S, Joseph J, Chmielinski K. The Dataset Nutrition Label: A Framework To Drive Higher Data Quality Standards. arXiv; 2018. [13] Bender EM, Friedman B. Data Statements for Natural Language Processing: Toward Miti- gating System Bias and Enabling Better Science. Transactions of the Association for Compu- tational Linguistics. 2018 Dec;6:587-604. [14] Singhal K, Azizi S, Tu T, Mahdavi S, Wei J, Chung HW, et al. Large Language Models Encode Clinical Knowledge. Nature. 2023 Aug;620(7972):172-80. [15] Nguyen V, Karimi S, Rybinski M, Xing Z. MedRedQA for Medical Consumer Question An- swering: Dataset, Tasks, and Neural Baselines. In: Park JC, Arase Y, Hu B, Lu W, Wijaya D, Purwarianti A, et al., editors. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers). Nusa Dua, Bali: Association for Computational Linguistics; 2023. p. 629-48. 10 [16] Arora RK, Wei J, Hicks RS, Bowman P, Qui Ěnonero-Candela J, Tsimpourlas F, et al.. Health- Bench: Evaluating Large Language Models Towards Improved Human Health. arXiv; 2025. [17] Khasentino J, Belyaeva A, Liu X, Yang Z, Furlotte NA, Lee C, et al. A Personal Health Large Language Model for Sleep and Fitness Coaching. Nature Medicine. 2025 Aug:1-10. [18] Chen S, Gao M, Sasse K, Hartvigsen T, Anthony B, Fan L, et al. When Helpfulness Backfires: LLMs and the Risk of False Medical Information Due to Sycophantic Behavior. npj Digital Medicine. 2025 Oct;8(1):605. [19] Metz C. Are A.I. Therapy Chatbots Safe to Use? The New York Times. 2025 Nov. [20] Hopewell S, Chan AW, Collins GS, Hr Ěobjartsson A, Moher D, Schulz KF, et al. CONSORT 2025 Statement: Updated Guideline for Reporting Randomized Trials. Nature Medicine. 2025 Jun;31(6):1776-83. [21] Tam TYC, Sivarajkumar S, Kapoor S, Stolyar AV, Polanska K, McCarthy KR, et al. A Framework for Human Evaluation of Large Language Models in Healthcare Derived from Literature Review. NPJ Digital Medicine. 2024 Sep;7:258. Tables and Figures Values are presented as n (%) for categorical variables and median (IQR) for continuous variables. 11 Figure 1: CONSORT diagram showing study flow from initial assessment through tagging methods to final analyzed datasets. 12 Figure 2: The Evolution of the Validity Gap. Each generation of benchmarks advances struc- tural capabilities but retains specific validity gaps. Generation 1 evaluates fact retrieval but lacks acuity for triage (1.2% triage intent, <2% high-risk). Generation 2 introduces clinical reasoning with rich objective data but relies on synthesized, static narratives (100% single-turn) where data is pre-packaged rather than elicited. Generation 3 achieves interactivity but suffers from sparse clinical content: <0.7% behavioral crisis, 5.5% chronic care, and clinical data density drops com- pared to Generation 2. Critically, demographic bias persists across all generations: pediatric and geriatric populations represent <11% of queries, validating models on a âstandard adult default.â 13 Figure 3: Distribution of health topics and query intents across three generations of benchmarks. (A) Health topics: Generation 1 (Search Era) queries span common symptoms and general health concerns. Generation 2 (Case Presentation Era) concentrates in internal medicine subspecialties with higher clinical complexity. Generation 3 (Interactive & Data Era) shows bimodal distribution: HealthBench Main covers diverse acute and chronic conditions, while GoogleFitbit datasets focus predominantly on wellness, sleep, and fitness tracking. (B) Query intents: Generation 1 bench- marks are dominated by queries to learn about health topics (92%), while Generation 3 benchmarks capture broader intent diversity including wellness and lifestyle guidance, medical research queries, and prevention and screening. 14 Table 2: Baseline Structural and Clinical Characteristics of theAnalytic Cohort (N=18,663 Queries) Characteristic HealthSearchQAMashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Total Records 3173 3490 5081 3692 1521 1750 18707 PopulationAdult (age unspecified), n(%) 3104 (97.8%) 3297 (94.5%) 4394 (86.5%) 3249 (88.0%) 1092 (71.8%) 1520 (86.9%) 16656 (89.0%) Adult (65+ years), n (%) 2 (0.1%) 13 (0.4%) 117 (2.3%) 70 (1.9%) 429 (28.2%) 230 (13.1%) 861 (4.6%) Pediatric (age unspeci-fied), n (%) 40 (1.3%) 115 (3.3%) 5 (0.1%) 121 (3.3%) 0 (0.0%) 0 (0.0%) 281 (1.5%) Pediatric (under 5 years),n (%) 26 (0.8%) 57 (1.6%) 97 (1.9%) 159 (4.3%) 0 (0.0%) 0 (0.0%) 339 (1.8%) Pediatric (5-17 years), n(%) 1 (0.0%) 8 (0.2%) 468 (9.2%) 93 (2.5%) 0 (0.0%) 0 (0.0%) 570 (3.0%) Prompt Length (words)Median (IQR) 7 (5-8) 8 (6-10) 150 (98-241) 16 (9-30) 659 (514-2326) 414 (44-434) 209 (5-232) LanguageEnglish, n (%) 3173 (100.0%) 3490 (100.0%) 5081 (100.0%) 3010 (81.5%) 1521 (100.0%) 1750 (100.0%) 18025 (96.4%) Conversation StructureSingle-turn, n (%) 3173 (100.0%) 3490 (100.0%) 5072 (99.8%) 2193 (59.4%) 1521 (100.0%) 1750 (100.0%) 17199 (91.9%) Multi-turn, n (%) 0 (0.0%) 0 (0.0%) 9 (0.2%) 1499 (40.6%) 0 (0.0%) 0 (0.0%) 1508 (8.1%) Risk SensitivityLow, n (%) 3093 (97.5%) 3441 (98.6%) 1639 (32.3%) 2161 (58.5%) 1521 (100.0%) 1750 (100.0%) 13605 (72.7%) Moderate, n (%) 70 (2.2%) 44 (1.3%) 2942 (57.9%) 1189 (32.2%) 0 (0.0%) 0 (0.0%) 4245 (22.7%) High, n (%) 10 (0.3%) 5 (0.1%) 500 (9.8%) 342 (9.3%) 0 (0.0%) 0 (0.0%) 857 (4.6%) Language ComplexityLay language, n (%) 3069 (96.7%) 3157 (90.5%) 4461 (87.8%) 3470 (94.0%) 1521 (100.0%) 1750 (100.0%) 17428 (93.2%) Continued on next page 15 Table 2 â Continued from previous page Characteristic HealthSearchQAMashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Technical language, n (%) 104 (3.3%) 333 (9.5%) 620 (12.2%) 222 (6.0%) 0 (0.0%) 0 (0.0%) 1279 (6.8%) Objective Data PresentNone, n (%) 3165 (99.7%) 3357 (96.2%) 1435 (28.2%) 2829 (76.6%) 0 (0.0%) 0 (0.0%) 10786 (57.7%) Wearable vitals, n (%) 0 (0.0%) 0 (0.0%) 35 (0.7%) 1 (0.0%) 1521 (100.0%) 1750 (100.0%) 3307 (17.7%) Diagnoses, n (%) 5 (0.2%) 110 (3.2%) 2230 (43.9%) 546 (14.8%) 0 (0.0%) 0 (0.0%) 2891 (15.5%) Medications, n (%) 0 (0.0%) 21 (0.6%) 2044 (40.2%) 257 (7.0%) 0 (0.0%) 0 (0.0%) 2322 (12.4%) Procedures, n (%) 0 (0.0%) 15 (0.4%) 879 (17.3%) 99 (2.7%) 0 (0.0%) 0 (0.0%) 993 (5.3%) Laboratory results, n (%) 0 (0.0%) 1 (0.0%) 898 (17.7%) 82 (2.2%) 0 (0.0%) 0 (0.0%) 981 (5.2%) Imaging, n (%) 0 (0.0%) 0 (0.0%) 659 (13.0%) 43 (1.2%) 0 (0.0%) 0 (0.0%) 702 (3.8%) Basic vitals, n (%) 3 (0.1%) 0 (0.0%) 383 (7.5%) 66 (1.8%) 0 (0.0%) 0 (0.0%) 452 (2.4%) Clinical artifacts, n (%) 0 (0.0%) 0 (0.0%) 93 (1.8%) 17 (0.5%) 0 (0.0%) 0 (0.0%) 110 (0.6%) Query SubjectSelf, n (%) 218 (6.9%) 223 (6.4%) 4327 (85.2%) 2247 (60.9%) 1521 (100.0%) 1750 (100.0%) 10286 (55.0%) General, n (%) 2944 (92.8%) 3215 (92.1%) 85 (1.7%) 870 (23.6%) 0 (0.0%) 0 (0.0%) 7114 (38.0%) Child, n (%) 11 (0.3%) 52 (1.5%) 160 (3.1%) 309 (8.4%) 0 (0.0%) 0 (0.0%) 532 (2.8%) Parent, n (%) 0 (0.0%) 0 (0.0%) 186 (3.7%) 75 (2.0%) 0 (0.0%) 0 (0.0%) 261 (1.4%) Partner, n (%) 0 (0.0%) 0 (0.0%) 158 (3.1%) 23 (0.6%) 0 (0.0%) 0 (0.0%) 181 (1.0%) Other, n (%) 0 (0.0%) 0 (0.0%) 165 (3.2%) 168 (4.6%) 0 (0.0%) 0 (0.0%) 333 (1.8%) 16 A Supplementary Material A.1 LLM Tagging Prompt The LLM tagging prompt (version 4.6) used with GPT-5.2 implements a deterministic classification framework with controlled vocabularies, priority ladders for intent disambiguation, and extensive validation rules. It classifies each query across 18 taxonomy dimensions including intent, topic, risk sensitivity, context depth, and objective data presence. The prompt includes few-shot examples and a validation checklist to ensure consistent output. The full prompt is available in the project repository at github.com/apple/ml-health-query- profiles (see prompts/tagging promptv46.txt). Note on message format: User messages were inserted in the following format: ⢠Single-turn queries: âUser: [message text]â ⢠Multi-turn conversations: âUser: [message]â¨newlineâŠAssistant: [response]â¨newlineâŠUser: [new message]...â (where â¨newline⊠represents a line break in the actual prompt) A.2 Description of Rules-based Tagging This appendix describes the rule-based approach used to annotate GoogleFitbit dataset queries with structured metadata tags. Unlike clinical benchmark datasets that used LLM-based tagging, the GoogleFitbit datasets employed deterministic programmatic classification due to their synthetic, highly-structured format. A.2.1 Rationale for Rule-Based Approach GoogleFitbit datasets consist of synthetic case studies with highly structured formats: demographic headers (age, gender), quantitative wearable data (sleep metrics, training load, heart rate variabil- ity), and health status indicators (BMI, conditions). This structural consistency enabled determin- istic rule-based classification without LLMs. Additionally, each case study was designed to be expanded into multiple individual prompts representing different analytical perspectives on the same case. This expansion step required pro- grammatic logic to maintain consistency across related prompts derived from a single case. A.2.2 Datasets Tagged Programmatically Two GoogleFitbit datasets were processed using the programmatic approach: ⢠GoogleFitbit Fitness Cases: 350 cases â 1,750 prompts (5Ă expansion) ⢠GoogleFitbit Sleep Cases: 507 cases â 1,521 prompts (3Ă expansion) Processing occurred on October 3, 2025. A.2.3 Case-to-Prompt Expansion GoogleFitbit case studies were not analyzed as single queries but instead expanded into multiple analytical units to capture different facets of the wearable data: 17 A.2.4 Fitness Cases (5 prompts per case) Each fitness case study was decomposed into five distinct analytical prompts as described in the original paper. 1. Demographics assessment: Analysis of age, gender, and baseline characteristics ⢠Focus: Personal characteristics relevant to fitness coaching 2. Training load analysis: Exercise patterns and training volume ⢠Focus: Workout frequency, intensity, and progression 3. Sleep metrics analysis: Sleep quality and recovery indicators ⢠Focus: Sleep patterns and their impact on fitness performance 4. Health metrics evaluation: Physiological measurements ⢠Focus: Heart rate, heart rate variability, BMI, and other biomarkers 5. Readiness assessment: Training readiness and recovery status ⢠Focus: Muscle soreness, subjective readiness, and recovery recommendations All fitness prompts were assigned the sub-intent classification fitness exercise under the top-level intent generalhealthadvice. A.2.5 Sleep Cases (3 prompts per case) Each sleep case study was decomposed into three analytical prompts: 1. Sleep insights: Pattern identification from sleep logs ⢠Focus: Identifying trends and anomalies in sleep data 2. Sleep etiology: Root cause analysis of sleep issues ⢠Focus: Diagnosing potential factors affecting sleep quality 3. Sleep recommendations: Personalized intervention suggestions ⢠Focus: Actionable advice for improving sleep outcomes All sleep prompts were assigned the sub-intent classification sleep hygiene under the top-level intent generalhealthadvice. Each expanded prompt inherited core metadata from the parent case (demographics, detected conditions) while receiving the same intent classification reflecting the overall analytical focus on personalized wellness coaching. A.2.6 Rule-Based Classification Logic The programmatic tagger applied deterministic rules to extract features from structured case data and assign classification tags. 18 A.2.7 Demographic Extraction Demographics were parsed from structured case headers using regex patterns: Age extraction and binning: ⢠Pattern matching: [40-44] (age range), 80+ (open-ended), 45 years old (explicit) ⢠Range handling: Age ranges mapped to midpoint (e.g., [40-44] â 42) ⢠Binning rules: â Age < 18 â pediatric â Age 18â64 â adult â Age ⼠65 â adult65plus Gender extraction: ⢠Pattern matching from case headers: âmale,â or âgender: maleâ ⢠Extracted values: male, female A.2.8 Condition Detection The original rule-based tagging included keyword-based condition detection using patterns such as: Obesity detection (multi-method): ⢠Keyword match: âobeseâ, âobesityâ, âoverweightâ ⢠BMI threshold: BMI ⼠30 extracted via regex (bmi: 32.5) ⢠Either trigger results in obesity condition label Respiratory conditions: ⢠Keywords: âasthmaâ, âCOPDâ, âbreathing issuesâ, âbreathing problemsâ ⢠Exclusion rule: Mentions of âRespiratory Rateâ (vital sign) do not trigger respiratory condition Other conditions (keyword-based): ⢠Diabetes: âdiabetesâ, âblood glucoseâ, âinsulinâ, âdiabeticâ ⢠Hypertension: âhypertensionâ, âhigh blood pressureâ ⢠Cardiovascular: âheart diseaseâ, âcardiovascularâ, âcardiacâ ⢠Joint issues: âarthritisâ, âjoint painâ, âknee painâ, âback painâ ⢠Metabolic syndrome: âmetabolic syndromeâ, âcholesterolâ, âtriglyceridesâ Note on data quality: During validation, these regex-based condition detection rules were found to be unreliable, producing false positives (e.g., all 1,750 fitness cases were incorrectly tagged with âobesityâ and ârespiratoryâ conditions, and all 1,521 sleep cases with âinsomniaâ). The key conditions field was therefore excluded from the final processed GoogleFitbit datasets to avoid contaminating downstream analyses with spurious condition labels. 19 A.2.9 Fixed Tag Assignments All GoogleFitbit prompts received consistent tags reflecting their synthetic wearable data context. These assignments were hard-coded rather than inferred: Tag FieldAssigned Value usertype consumer conversationstructure singleturn intenttop generalhealthadvice intent sub fitnessexercise (fitness) or sleephygiene (sleep) topicareapath[Holistic Health & Wellness, Sleep & Lifestyle] specialty sportsmedicine (fitness) or sleepmedicine (sleep) objective data vitalswearable contextdepth high setting home language complexity lay risk sensitivity low language english lengthdetail detailed needsclarification false region null Table 3: Fixed tag assignments for GoogleFitbit datasets reflecting wearable data context Rationale for fixed assignments: ⢠consumer: All case studies represent consumer health scenarios ⢠single turn: Synthetic cases have no prior conversation history ⢠vitalswearable: All cases include consumer device data (sleep tracking, heart rate, training metrics) ⢠high context depth: All cases provide detailed quantitative data including age, metrics over time, and health indicators ⢠home: Consumer wearable usage context ⢠lay complexity: Case descriptions written for general audiences ⢠low risk: Wellness and fitness optimization queries, not acute medical concerns A.3 Model Comparison Study A.3.1 Overview This section presents a systematic comparison of tagging outputs from two large language modelsâ GPT-5.2 and Claude Opus 4.5âapplied to the HealthBench Main dataset (N=4,967 matched records). The goal is to assess the reliability of LLM-derived tags by evaluating cross-model agree- ment across 21 structured dimensions. Of the 5,000 queries in HealthBench Main, 33 (0.7%) were not tagged by Opus 4.5 due to content filtering on queries involving sensitive clinical scenarios (e.g., pandemic-related questions, danger- ous pathogens, terse symptom presentations). These queries were excluded from the comparison, 20 yielding 4,967 matched pairs. This pattern is consistent with documented differences in content moderation policies between model providers and does not affect the validity of the comparison for the remaining queries. A.3.2 Methods Both models independently tagged the same queries using identical v4.6 tagging prompts with JSON output formatting. Agreement was assessed using: ⢠Percent agreement: Proportion of exact matches between models ⢠Cohenâs Îş: Agreement corrected for chance, where Îş > 0.8 indicates almost perfect agree- ment and Îş > 0.6 indicates substantial agreement ⢠Chi-square tests: Distribution equivalence with Cram Ěerâs V effect size A.3.3 Results Overall Agreement by Domain Across all 21 dimensions, models demonstrated strong agree- ment (Table 4). Context dimensions showed the highest concordance, followed by topic and intent dimensions. Table 4: Agreement by Query Profile Domain Domain Dimensions Avg Agreement Avg Îş Context1692.6%0.77 Topic383.7%0.79 Intent278.5%0.76 Dimension-Level Agreement Table 5 presents agreement metrics for each dimension. Distribution Comparisons Chi-square tests identified 5 dimensions with statistically significant distribution differences (p ÂĄ 0.05 with Cram Ěer Ěs V Âż 0.1): ⢠Setting: Ď 2 = 403.6, V = 0.20 ⢠Needs Clarification: Ď 2 = 298.0, V = 0.17 ⢠Specialty: Ď 2 = 272.4, V = 0.17 ⢠Sub-Intent: Ď 2 = 204.8, V = 0.14 ⢠Context Depth: Ď 2 = 101.0, V = 0.10 Visual Comparison Figure 4 presents side-by-side distribution comparisons for key dimensions, demonstrating strong alignment between models across categorical values. Figure 5 shows confusion matrices for dimensions with the most clinical relevance: intent clas- sification, risk sensitivity, and specialty assignment. 21 Table 5: Agreement Metrics by Dimension Domain DimensionAgreement Îş Context Conversation Structure100.0%1.00 Language99.8%0.99 Population98.6%0.95 Raw Medical Text97.8%0.71 Region96.9%0.77 User Type96.9%0.92 Personal Health Query95.3%0.91 Language Complexity95.2%0.88 Needs Clarification92.8%0.17 Query Subject92.3%0.89 Needs Personalization91.7%0.77 Length/Detail90.3%0.66 Context Depth85.9%0.65 Risk Sensitivity85.9%0.74 Objective Data83.3%0.69 Setting79.1%0.59 Topic Key Conditions91.7%0.80 Specialty80.6%0.79 Topic Area78.9%0.79 Intent Top-Level Intent83.7%0.81 Sub-Intent73.3%0.71 A.3.4 Discussion Implications for Tagging Validity The strong agreement between GPT-5.2 and Opus-4.5âtwo models with fundamentally different architectures and training approachesâprovides compelling evidence for the validity of LLM-derived tags. Key findings include: 1. Structural dimensions show near-perfect agreement: Language, conversation struc- ture, and user type achieved >95% agreement, confirming these are objective, well-defined attributes. 2. Aggregate distributions are statistically equivalent: Chi-square tests revealed no meaningful distributional differences, indicating that population-level insights derived from either model would be consistent. 3. Disagreements reflect genuine ambiguity: Where models diverged (e.g., risk sensitivity boundaries), the disagreements occurred at category boundaries where human annotators would also exhibit variability. 4. Cross-architecture validation: Agreement between models trained by different organiza- tions using different approaches provides stronger validity evidence than same-model repro- ducibility. Comparison to Human Inter-Rater Reliability The observed agreement levels (Îş = 0.6â0.9 across most dimensions) are comparable to or exceed typical human inter-rater reliability in medical 22 Figure 4: Distribution comparison between GPT-5.2 and Opus-4.5 across key dimensions. Bars show percentage of queries assigned to each category. annotation tasks, where Îş values of 0.4â0.7 are common for subjective clinical judgments. This suggests LLM tagging achieves human-level consistency while offering scalability advantages. Limitations This comparison has limitations: (1) agreement was measured between models rather than against human gold standard; (2) both models used identical prompts, so shared prompt biases would not be detected; (3) analysis was limited to a single dataset. However, cross- model agreement provides meaningful validity evidence even without gold-standard comparison, as systematic biases are unlikely to be shared across independently developed models. 23 Figure 5: Confusion matrices showing agreement patterns for intent, risk sensitivity, and specialty. Values show row-normalized percentages. Strong diagonal dominance indicates high agreement. 24 Table S1: Distribution of Clinical Intents Across âGenerationsâ ofHealth AI Benchmarks Intent/Sub-Intent HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Education 2645 (83.4%) 2389 (68.5%) 154 (3.0%) 285 (7.7%) 0 (0.0%) 0 (0.0%) 5473 (29.3%) Education Explainer 2635 (83.0%) 2250 (64.5%) 149 (2.9%) 280 (7.6%) 0 (0.0%) 0 (0.0%) 5314 (28.4%) Basic Science 10 (0.3%) 139 (4.0%) 5 (0.1%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 159 (0.8%) Symptom Check 380 (12.0%) 103 (3.0%) 2656 (52.3%) 1384 (37.5%) 0 (0.0%) 0 (0.0%) 4523 (24.2%) Management Plan 235 (7.4%) 31 (0.9%) 774 (15.2%) 681 (18.4%) 0 (0.0%) 0 (0.0%) 1721 (9.2%) Differential Diagno- sis 107 (3.4%) 20 (0.6%) 1235 (24.3%) 311 (8.4%) 0 (0.0%) 0 (0.0%) 1673 (8.9%) Triage Disposition 38 (1.2%) 52 (1.5%) 647 (12.7%) 392 (10.6%) 0 (0.0%) 0 (0.0%) 1129 (6.0%) General Health Ad-vice 47 (1.5%) 240 (6.9%) 70 (1.4%) 270 (7.3%) 1521 (100.0%) 1750 (100.0%) 3898 (20.8%) Fitness Exercise 1 (0.0%) 40 (1.1%) 6 (0.1%) 40 (1.1%) 0 (0.0%) 1750 (100.0%) 1837 (9.8%) Sleep Hygiene 8 (0.3%) 6 (0.2%) 4 (0.1%) 20 (0.5%) 1521 (100.0%) 0 (0.0%) 1559 (8.3%) Nutrition Diet 19 (0.6%) 72 (2.1%) 25 (0.5%) 115 (3.1%) 0 (0.0%) 0 (0.0%) 231 (1.2%) Supplements Nu- traceuticals 2 (0.1%) 64 (1.8%) 15 (0.3%) 60 (1.6%) 0 (0.0%) 0 (0.0%) 141 (0.8%) Stress Selfcare 7 (0.2%) 40 (1.1%) 8 (0.2%) 6 (0.2%) 0 (0.0%) 0 (0.0%) 61 (0.3%) Cosmeceuticals Top- icals 10 (0.3%) 10 (0.3%) 9 (0.2%) 24 (0.7%) 0 (0.0%) 0 (0.0%) 53 (0.3%) Nutrition Facts Lookup 0 (0.0%) 8 (0.2%) 2 (0.0%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 15 (0.1%) Lifestyle Prevention 0 (0.0%) 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Medication Infor- mation 11 (0.3%) 332 (9.5%) 608 (12.0%) 729 (19.7%) 0 (0.0%) 0 (0.0%) 1680 (9.0%) Selection 8 (0.3%) 112 (3.2%) 134 (2.6%) 443 (12.0%) 0 (0.0%) 0 (0.0%) 697 (3.7%) Side Effects 2 (0.1%) 157 (4.5%) 240 (4.7%) 72 (2.0%) 0 (0.0%) 0 (0.0%) 471 (2.5%) Dosing 0 (0.0%) 45 (1.3%) 99 (1.9%) 170 (4.6%) 0 (0.0%) 0 (0.0%) 314 (1.7%) Interactions 1 (0.0%) 8 (0.2%) 119 (2.3%) 26 (0.7%) 0 (0.0%) 0 (0.0%) 154 (0.8%) Safety Preg Lact 0 (0.0%) 10 (0.3%) 16 (0.3%) 18 (0.5%) 0 (0.0%) 0 (0.0%) 44 (0.2%) Continued on next page 25 Table S1 â Continued from previous page Intent/Sub-Intent HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Condition Manage-ment 25 (0.8%) 157 (4.5%) 558 (11.0%) 297 (8.0%) 0 (0.0%) 0 (0.0%) 1037 (5.5%) Chronic Care Sup- port 21 (0.7%) 141 (4.0%) 209 (4.1%) 227 (6.1%) 0 (0.0%) 0 (0.0%) 598 (3.2%) Risk Prognosis 0 (0.0%) 10 (0.3%) 212 (4.2%) 44 (1.2%) 0 (0.0%) 0 (0.0%) 266 (1.4%) Acute Flare Man- agement 4 (0.1%) 6 (0.2%) 137 (2.7%) 26 (0.7%) 0 (0.0%) 0 (0.0%) 173 (0.9%) Tests And Results 7 (0.2%) 82 (2.3%) 702 (13.8%) 115 (3.1%) 0 (0.0%) 0 (0.0%) 906 (4.8%) Test Interpretation 4 (0.1%) 8 (0.2%) 593 (11.7%) 52 (1.4%) 0 (0.0%) 0 (0.0%) 657 (3.5%) Test Selection 3 (0.1%) 74 (2.1%) 109 (2.1%) 63 (1.7%) 0 (0.0%) 0 (0.0%) 249 (1.3%) Prevention Screen-ing 7 (0.2%) 117 (3.4%) 125 (2.5%) 264 (7.2%) 0 (0.0%) 0 (0.0%) 513 (2.7%) Vaccination 3 (0.1%) 10 (0.3%) 94 (1.9%) 139 (3.8%) 0 (0.0%) 0 (0.0%) 246 (1.3%) Lifestyle Prevention 2 (0.1%) 89 (2.6%) 25 (0.5%) 66 (1.8%) 0 (0.0%) 0 (0.0%) 182 (1.0%) Screening Schedule 2 (0.1%) 18 (0.5%) 6 (0.1%) 59 (1.6%) 0 (0.0%) 0 (0.0%) 85 (0.5%) AdministrativeMeta 1 (0.0%) 37 (1.1%) 182 (3.6%) 173 (4.7%) 0 (0.0%) 0 (0.0%) 393 (2.1%) Navigation Referral 0 (0.0%) 19 (0.5%) 117 (2.3%) 66 (1.8%) 0 (0.0%) 0 (0.0%) 202 (1.1%) Other Admin 1 (0.0%) 9 (0.3%) 40 (0.8%) 30 (0.8%) 0 (0.0%) 0 (0.0%) 80 (0.4%) Draft Communica- tion 0 (0.0%) 4 (0.1%) 10 (0.2%) 33 (0.9%) 0 (0.0%) 0 (0.0%) 47 (0.3%) Insurance Billing 0 (0.0%) 5 (0.1%) 13 (0.3%) 21 (0.6%) 0 (0.0%) 0 (0.0%) 39 (0.2%) Document Summary 0 (0.0%) 0 (0.0%) 2 (0.0%) 23 (0.6%) 0 (0.0%) 0 (0.0%) 25 (0.1%) Research Help 0 (0.0%) 21 (0.6%) 8 (0.2%) 170 (4.6%) 0 (0.0%) 0 (0.0%) 199 (1.1%) Study Evidence 0 (0.0%) 7 (0.2%) 6 (0.1%) 100 (2.7%) 0 (0.0%) 0 (0.0%) 113 (0.6%) Medical Guidelines 0 (0.0%) 3 (0.1%) 1 (0.0%) 66 (1.8%) 0 (0.0%) 0 (0.0%) 70 (0.4%) Research Explana- tion 0 (0.0%) 11 (0.3%) 1 (0.0%) 4 (0.1%) 0 (0.0%) 0 (0.0%) 16 (0.1%) Non Health 50 (1.6%) 12 (0.3%) 18 (0.4%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 85 (0.5%) Offtopic Nonhealth 50 (1.6%) 12 (0.3%) 18 (0.4%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 85 (0.5%) 26 A.4 Supplementary Tables Values are presented as n (%) for each dataset and overall total. Sub-intents are indented under their corresponding top-level intents. Values are presented as n (%) for each dataset and overall total. Sub-intents are indented under their corresponding top-level intents. 27 Values are presented as n (%) for each dataset and overall total. Sub-topics are indented under their corresponding topic areas. Values are presented as n (%) for each dataset and overall total. Sub-topics are indented under their corresponding topic areas. 28 Table S2: Prevalence of Clinical Domains and Specific MedicalConditions Across Benchmarks Topic Area/Sub- Topic HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Holistic Health &Wellness 63 (2.0%) 156 (4.5%) 108 (2.1%) 197 (5.3%) 1521 (100.0%) 1750 (100.0%) 3795 (20.3%) Sleep & Lifestyle 43 (1.4%) 31 (0.9%) 46 (0.9%) 71 (1.9%) 1521 (100.0%) 1750 (100.0%) 3462 (18.5%) Alternative/Eastern 0 (0.0%) 28 (0.8%) 6 (0.1%) 20 (0.5%) 0 (0.0%) 0 (0.0%) 54 (0.3%) Fitness & Exercise 0 (0.0%) 14 (0.4%) 6 (0.1%) 24 (0.7%) 0 (0.0%) 0 (0.0%) 44 (0.2%) Mental Wellness 3 (0.1%) 19 (0.5%) 4 (0.1%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 31 (0.2%) Cosmeceuticals & Topicals 1 (0.0%) 7 (0.2%) 3 (0.1%) 6 (0.2%) 0 (0.0%) 0 (0.0%) 17 (0.1%) Skin & Hair 366 (11.5%) 272 (7.8%) 761 (15.0%) 315 (8.5%) 0 (0.0%) 0 (0.0%) 1714 (9.2%) Wounds 36 (1.1%) 38 (1.1%) 175 (3.4%) 104 (2.8%) 0 (0.0%) 0 (0.0%) 353 (1.9%) Rash 56 (1.8%) 28 (0.8%) 194 (3.8%) 46 (1.2%) 0 (0.0%) 0 (0.0%) 324 (1.7%) Infections 64 (2.0%) 18 (0.5%) 132 (2.6%) 34 (0.9%) 0 (0.0%) 0 (0.0%) 248 (1.3%) Eczema 20 (0.6%) 29 (0.8%) 26 (0.5%) 17 (0.5%) 0 (0.0%) 0 (0.0%) 92 (0.5%) Bites & Stings 11 (0.3%) 13 (0.4%) 42 (0.8%) 18 (0.5%) 0 (0.0%) 0 (0.0%) 84 (0.4%) Acne 5 (0.2%) 15 (0.4%) 19 (0.4%) 22 (0.6%) 0 (0.0%) 0 (0.0%) 61 (0.3%) Hair Loss 5 (0.2%) 1 (0.0%) 16 (0.3%) 21 (0.6%) 0 (0.0%) 0 (0.0%) 43 (0.2%) Brain & Nerves 366 (11.5%) 283 (8.1%) 454 (8.9%) 344 (9.3%) 0 (0.0%) 0 (0.0%) 1447 (7.7%) Headache/Migraine 14 (0.4%) 100 (2.9%) 76 (1.5%) 88 (2.4%) 0 (0.0%) 0 (0.0%) 278 (1.5%) Neuropathy 43 (1.4%) 30 (0.9%) 63 (1.2%) 20 (0.5%) 0 (0.0%) 0 (0.0%) 156 (0.8%) Cognitive Changes 39 (1.2%) 21 (0.6%) 63 (1.2%) 31 (0.8%) 0 (0.0%) 0 (0.0%) 154 (0.8%) Dizziness/Vertigo 23 (0.7%) 8 (0.2%) 46 (0.9%) 59 (1.6%) 0 (0.0%) 0 (0.0%) 136 (0.7%) Stroke/TIA 23 (0.7%) 1 (0.0%) 42 (0.8%) 38 (1.0%) 0 (0.0%) 0 (0.0%) 104 (0.6%) Seizure 12 (0.4%) 7 (0.2%) 46 (0.9%) 14 (0.4%) 0 (0.0%) 0 (0.0%) 79 (0.4%) Digestive & Nutri-tion 252 (7.9%) 272 (7.8%) 521 (10.3%) 339 (9.2%) 0 (0.0%) 0 (0.0%) 1384 (7.4%) Liver Disease 33 (1.0%) 32 (0.9%) 75 (1.5%) 24 (0.7%) 0 (0.0%) 0 (0.0%) 164 (0.9%) Abdominal Pain 12 (0.4%) 10 (0.3%) 69 (1.4%) 42 (1.1%) 0 (0.0%) 0 (0.0%) 133 (0.7%) Reflux/Heartburn 19 (0.6%) 6 (0.2%) 38 (0.7%) 33 (0.9%) 0 (0.0%) 0 (0.0%) 96 (0.5%) Nausea & Vomiting 16 (0.5%) 2 (0.1%) 47 (0.9%) 16 (0.4%) 0 (0.0%) 0 (0.0%) 81 (0.4%) Diarrhea 7 (0.2%) 7 (0.2%) 37 (0.7%) 22 (0.6%) 0 (0.0%) 0 (0.0%) 73 (0.4%) Continued on next page 29 Table S2 â Continued from previous page Topic Area/Sub- Topic HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Inflammatory Bowel Disease (IBD) 12 (0.4%) 39 (1.1%) 13 (0.3%) 9 (0.2%) 0 (0.0%) 0 (0.0%) 73 (0.4%) Weight Management 13 (0.4%) 4 (0.1%) 26 (0.5%) 27 (0.7%) 0 (0.0%) 0 (0.0%) 70 (0.4%) Constipation 4 (0.1%) 25 (0.7%) 28 (0.6%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 62 (0.3%) Gallbladder/Biliary 8 (0.3%) 0 (0.0%) 18 (0.4%) 8 (0.2%) 0 (0.0%) 0 (0.0%) 34 (0.2%) Infilammatory Bowel Disease (IBD) 0 (0.0%) 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Infections (Gen- eral) 252 (7.9%) 109 (3.1%) 435 (8.6%) 425 (11.5%) 0 (0.0%) 0 (0.0%) 1221 (6.5%) COVID-19 2 (0.1%) 0 (0.0%) 139 (2.7%) 14 (0.4%) 0 (0.0%) 0 (0.0%) 155 (0.8%) Fever (Unspecified) 16 (0.5%) 7 (0.2%) 36 (0.7%) 66 (1.8%) 0 (0.0%) 0 (0.0%) 125 (0.7%) Travel-Related 20 (0.6%) 4 (0.1%) 4 (0.1%) 81 (2.2%) 0 (0.0%) 0 (0.0%) 109 (0.6%) Antimicro- bials/Antibiotics 0 (0.0%) 7 (0.2%) 20 (0.4%) 27 (0.7%) 0 (0.0%) 0 (0.0%) 54 (0.3%) Skin/Soft Tissue 12 (0.4%) 5 (0.1%) 11 (0.2%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 33 (0.2%) Tuberculosis 4 (0.1%) 1 (0.0%) 2 (0.0%) 11 (0.3%) 0 (0.0%) 0 (0.0%) 18 (0.1%) Long COVID 0 (0.0%) 0 (0.0%) 1 (0.0%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 4 (0.0%) Muscles, Bones &Joints 274 (8.6%) 330 (9.5%) 340 (6.7%) 251 (6.8%) 0 (0.0%) 0 (0.0%) 1195 (6.4%) Arthritis 32 (1.0%) 217 (6.2%) 21 (0.4%) 29 (0.8%) 0 (0.0%) 0 (0.0%) 299 (1.6%) Back & Neck Pain 19 (0.6%) 30 (0.9%) 88 (1.7%) 80 (2.2%) 0 (0.0%) 0 (0.0%) 217 (1.2%) Sprains & Strains 19 (0.6%) 9 (0.3%) 61 (1.2%) 39 (1.1%) 0 (0.0%) 0 (0.0%) 128 (0.7%) Knee Problems 12 (0.4%) 21 (0.6%) 19 (0.4%) 36 (1.0%) 0 (0.0%) 0 (0.0%) 88 (0.5%) Shoulder Problems 12 (0.4%) 11 (0.3%) 20 (0.4%) 13 (0.4%) 0 (0.0%) 0 (0.0%) 56 (0.3%) Osteoporosis 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Heart & Circulation 190 (6.0%) 262 (7.5%) 428 (8.4%) 245 (6.6%) 0 (0.0%) 0 (0.0%) 1125 (6.0%) Chest Pain 14 (0.4%) 11 (0.3%) 68 (1.3%) 57 (1.5%) 0 (0.0%) 0 (0.0%) 150 (0.8%) Hypertension 10 (0.3%) 50 (1.4%) 35 (0.7%) 44 (1.2%) 0 (0.0%) 0 (0.0%) 139 (0.7%) Palpitations 14 (0.4%) 0 (0.0%) 59 (1.2%) 26 (0.7%) 0 (0.0%) 0 (0.0%) 99 (0.5%) Arrhythmia 9 (0.3%) 26 (0.7%) 54 (1.1%) 8 (0.2%) 0 (0.0%) 0 (0.0%) 97 (0.5%) Venous Throm- boembolism 21 (0.7%) 12 (0.3%) 39 (0.8%) 13 (0.4%) 0 (0.0%) 0 (0.0%) 85 (0.5%) Continued on next page 30 Table S2 â Continued from previous page Topic Area/Sub- Topic HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Heart Failure 5 (0.2%) 49 (1.4%) 10 (0.2%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 67 (0.4%) Hyperlipidemia 4 (0.1%) 26 (0.7%) 5 (0.1%) 25 (0.7%) 0 (0.0%) 0 (0.0%) 60 (0.3%) Coronary Artery Disease 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Coronary artery dis- ease 0 (0.0%) 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Mental Health &Psychiatry 226 (7.1%) 268 (7.7%) 259 (5.1%) 261 (7.1%) 0 (0.0%) 0 (0.0%) 1014 (5.4%) ADHD 10 (0.3%) 121 (3.5%) 12 (0.2%) 31 (0.8%) 0 (0.0%) 0 (0.0%) 174 (0.9%) Anxiety 27 (0.9%) 20 (0.6%) 77 (1.5%) 49 (1.3%) 0 (0.0%) 0 (0.0%) 173 (0.9%) Depression 14 (0.4%) 25 (0.7%) 25 (0.5%) 47 (1.3%) 0 (0.0%) 0 (0.0%) 111 (0.6%) Postpartum Mental Health 6 (0.2%) 0 (0.0%) 1 (0.0%) 79 (2.1%) 0 (0.0%) 0 (0.0%) 86 (0.5%) Substance Use 0 (0.0%) 12 (0.3%) 30 (0.6%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 45 (0.2%) Bipolar Disorder 4 (0.1%) 25 (0.7%) 10 (0.2%) 4 (0.1%) 0 (0.0%) 0 (0.0%) 43 (0.2%) Eating Disorders 18 (0.6%) 12 (0.3%) 8 (0.2%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 41 (0.2%) Suicidal Ideation (Crisis) 0 (0.0%) 2 (0.1%) 26 (0.5%) 12 (0.3%) 0 (0.0%) 0 (0.0%) 40 (0.2%) OCD 19 (0.6%) 5 (0.1%) 10 (0.2%) 4 (0.1%) 0 (0.0%) 0 (0.0%) 38 (0.2%) Schizophrenia 5 (0.2%) 11 (0.3%) 7 (0.1%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 24 (0.1%) PTSD 7 (0.2%) 1 (0.0%) 5 (0.1%) 4 (0.1%) 0 (0.0%) 0 (0.0%) 17 (0.1%) Cancer 187 (5.9%) 524 (15.0%) 136 (2.7%) 85 (2.3%) 0 (0.0%) 0 (0.0%) 932 (5.0%) Breast 20 (0.6%) 74 (2.1%) 15 (0.3%) 25 (0.7%) 0 (0.0%) 0 (0.0%) 134 (0.7%) Skin 9 (0.3%) 56 (1.6%) 26 (0.5%) 8 (0.2%) 0 (0.0%) 0 (0.0%) 99 (0.5%) Lung 6 (0.2%) 75 (2.1%) 6 (0.1%) 4 (0.1%) 0 (0.0%) 0 (0.0%) 91 (0.5%) Colorectal 9 (0.3%) 13 (0.4%) 5 (0.1%) 16 (0.4%) 0 (0.0%) 0 (0.0%) 43 (0.2%) Prostate 5 (0.2%) 15 (0.4%) 3 (0.1%) 6 (0.2%) 0 (0.0%) 0 (0.0%) 29 (0.2%) Reproductive & Sexual Health 170 (5.4%) 143 (4.1%) 368 (7.2%) 97 (2.6%) 0 (0.0%) 0 (0.0%) 778 (4.2%) STIs 21 (0.7%) 24 (0.7%) 83 (1.6%) 12 (0.3%) 0 (0.0%) 0 (0.0%) 140 (0.7%) Menstrual Disorders 34 (1.1%) 11 (0.3%) 55 (1.1%) 13 (0.4%) 0 (0.0%) 0 (0.0%) 113 (0.6%) Sexual Function 30 (0.9%) 12 (0.3%) 56 (1.1%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 103 (0.6%) Continued on next page 31 Table S2 â Continued from previous page Topic Area/Sub- Topic HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Contraception 0 (0.0%) 4 (0.1%) 33 (0.6%) 44 (1.2%) 0 (0.0%) 0 (0.0%) 81 (0.4%) Peri- menopause/Menopause 14 (0.4%) 35 (1.0%) 1 (0.0%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 53 (0.3%) Fertility 16 (0.5%) 5 (0.1%) 15 (0.3%) 11 (0.3%) 0 (0.0%) 0 (0.0%) 47 (0.3%) Endocrine & Metabolic 120 (3.8%) 209 (6.0%) 183 (3.6%) 163 (4.4%) 0 (0.0%) 0 (0.0%) 675 (3.6%) Diabetes 23 (0.7%) 106 (3.0%) 36 (0.7%) 79 (2.1%) 0 (0.0%) 0 (0.0%) 244 (1.3%) Thyroid 15 (0.5%) 13 (0.4%) 32 (0.6%) 19 (0.5%) 0 (0.0%) 0 (0.0%) 79 (0.4%) Osteope- nia/Osteoporosis 11 (0.3%) 25 (0.7%) 0 (0.0%) 10 (0.3%) 0 (0.0%) 0 (0.0%) 46 (0.2%) Obesity 4 (0.1%) 13 (0.4%) 9 (0.2%) 14 (0.4%) 0 (0.0%) 0 (0.0%) 40 (0.2%) PCOS 7 (0.2%) 1 (0.0%) 7 (0.1%) 2 (0.1%) 0 (0.0%) 0 (0.0%) 17 (0.1%) Lipids 0 (0.0%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1 (0.0%) Lungs & Breathing(Colds/Flu/COVID) 109 (3.4%) 82 (2.3%) 232 (4.6%) 182 (4.9%) 0 (0.0%) 0 (0.0%) 605 (3.2%) Shortness of Breath 7 (0.2%) 5 (0.1%) 54 (1.1%) 59 (1.6%) 0 (0.0%) 0 (0.0%) 125 (0.7%) Upper Respiratory Infection 15 (0.5%) 30 (0.9%) 22 (0.4%) 32 (0.9%) 0 (0.0%) 0 (0.0%) 99 (0.5%) Cough 12 (0.4%) 8 (0.2%) 30 (0.6%) 38 (1.0%) 0 (0.0%) 0 (0.0%) 88 (0.5%) Pneumonia 8 (0.3%) 16 (0.5%) 18 (0.4%) 15 (0.4%) 0 (0.0%) 0 (0.0%) 57 (0.3%) COVID-19 7 (0.2%) 0 (0.0%) 46 (0.9%) 1 (0.0%) 0 (0.0%) 0 (0.0%) 54 (0.3%) Wheezing 4 (0.1%) 0 (0.0%) 6 (0.1%) 3 (0.1%) 0 (0.0%) 0 (0.0%) 13 (0.1%) Ears, Nose & Throat 141 (4.4%) 37 (1.1%) 221 (4.3%) 149 (4.0%) 0 (0.0%) 0 (0.0%) 548 (2.9%) Otitis/Ear Pain 28 (0.9%) 9 (0.3%) 27 (0.5%) 36 (1.0%) 0 (0.0%) 0 (0.0%) 100 (0.5%) Sore Throat 17 (0.5%) 10 (0.3%) 47 (0.9%) 21 (0.6%) 0 (0.0%) 0 (0.0%) 95 (0.5%) Hearing Loss/Tinnitus 20 (0.6%) 5 (0.1%) 20 (0.4%) 19 (0.5%) 0 (0.0%) 0 (0.0%) 64 (0.3%) Nasal Congestion 9 (0.3%) 5 (0.1%) 18 (0.4%) 8 (0.2%) 0 (0.0%) 0 (0.0%) 40 (0.2%) Sinusitis 5 (0.2%) 0 (0.0%) 13 (0.3%) 17 (0.5%) 0 (0.0%) 0 (0.0%) 35 (0.2%) Eyes & Vision 138 (4.3%) 195 (5.6%) 111 (2.2%) 65 (1.8%) 0 (0.0%) 0 (0.0%) 509 (2.7%) Vision Changes 39 (1.2%) 15 (0.4%) 54 (1.1%) 20 (0.5%) 0 (0.0%) 0 (0.0%) 128 (0.7%) Continued on next page 32 Table S2 â Continued from previous page Topic Area/Sub- Topic HealthSearchQA MashQA Test MedRedQA HealthBench GoogleFitbit GoogleFitbit Total Test Main Sleep Fitness Dry Eye 10 (0.3%) 13 (0.4%) 4 (0.1%) 6 (0.2%) 0 (0.0%) 0 (0.0%) 33 (0.2%) Conjunctivitis 9 (0.3%) 2 (0.1%) 4 (0.1%) 9 (0.2%) 0 (0.0%) 0 (0.0%) 24 (0.1%) Foreign Body/Irritation 1 (0.0%) 5 (0.1%) 15 (0.3%) 2 (0.1%) 0 (0.0%) 0 (0.0%) 23 (0.1%) Kidney & Urinary 115 (3.6%) 47 (1.3%) 186 (3.7%) 88 (2.4%) 0 (0.0%) 0 (0.0%) 436 (2.3%) Urinary Tract Infec- tion 19 (0.6%) 0 (0.0%) 52 (1.0%) 20 (0.5%) 0 (0.0%) 0 (0.0%) 91 (0.5%) Chronic Kidney Dis- ease 7 (0.2%) 11 (0.3%) 11 (0.2%) 26 (0.7%) 0 (0.0%) 0 (0.0%) 55 (0.3%) Incontinence 20 (0.6%) 6 (0.2%) 10 (0.2%) 6 (0.2%) 0 (0.0%) 0 (0.0%) 42 (0.2%) Hematuria 8 (0.3%) 0 (0.0%) 22 (0.4%) 5 (0.1%) 0 (0.0%) 0 (0.0%) 35 (0.2%) Kidney Stones 8 (0.3%) 3 (0.1%) 9 (0.2%) 7 (0.2%) 0 (0.0%) 0 (0.0%) 27 (0.1%) Prostatitis/BPH 8 (0.3%) 2 (0.1%) 7 (0.1%) 2 (0.1%) 0 (0.0%) 0 (0.0%) 19 (0.1%) 33 Values are presented as n (%) for each characteristic. Generation 1 = HealthSearchQA + MashQA; Generation 2 = MedRedQA; Generation 3 = HealthBench + GoogleFitbit. Conditions with fewer than 10 queries (3 conditions: epilepsy, leishmaniasis, schizophrenia) are excluded from this table. Values are presented as n (%) for each characteristic. Key conditions represent explicit mentions identi- fied by the LLM tagger (e.g., âI have diabetesâ, âmy asthmaâ), not general topic inference; therefore, counts differ from topic area classifications in prior tables. Generation 1 = HealthSearchQA + MashQA; Genera- tion 2 = MedRedQA; Generation 3 = HealthBench + GoogleFitbit. Conditions with fewer than 10 queries (1 condition: lung cancer) are excluded from this table. 34 Table S3: Characteristics of Queries Mentioning Key Health Conditions(N=18,663 Consumer Queries) Condition N % of All Single-turn Mod-High Risk Generation 1 Generation 2 Generation 3 Other 3284 17.30% 3283 (100.0%) 10 (0.3%) 0 (0.0%) 9 (0.3%) 3275 (99.7%) Respiratory 1750 9.22% 1750 (100.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1750 (100.0%) Insomnia 1521 8.01% 1521 (100.0%) 0 (0.0%) 0 (0.0%) 0 (0.0%) 1521 (100.0%) Long Covid 13 0.07% 12 (92.3%) 10 (76.9%) 0 (0.0%) 9 (69.2%) 4 (30.8%) Endocrine & Metabolic 2423 12.76% 2348 (96.9%) 320 (13.2%) 173 (7.1%) 339 (14.0%) 1911 (78.9%) Obesity 1865 9.82% 1857 (99.6%) 67 (3.6%) 12 (0.6%) 84 (4.5%) 1769 (94.9%) Diabetes 331 1.74% 280 (84.6%) 129 (39.0%) 132 (39.9%) 91 (27.5%) 108 (32.6%) Thyroid Disease 105 0.55% 99 (94.3%) 59 (56.2%) 16 (15.2%) 76 (72.4%) 13 (12.4%) PCOS 66 0.35% 64 (97.0%) 37 (56.1%) 7 (10.6%) 55 (83.3%) 4 (6.1%) Hyperlipidemia 56 0.29% 48 (85.7%) 28 (50.0%) 6 (10.7%) 33 (58.9%) 17 (30.4%) Mental Health & Be-havioral 1748 9.21% 1649 (94.3%) 1188 (68.0%) 272 (15.6%) 1253 (71.7%) 223 (12.8%) Anxiety 647 3.41% 611 (94.4%) 475 (73.4%) 33 (5.1%) 538 (83.2%) 76 (11.7%) Depression 393 2.07% 361 (91.9%) 294 (74.8%) 32 (8.1%) 288 (73.3%) 73 (18.6%) ADHD 262 1.38% 246 (93.9%) 94 (35.9%) 126 (48.1%) 100 (38.2%) 36 (13.7%) Bipolar Disorder 87 0.46% 85 (97.7%) 49 (56.3%) 30 (34.5%) 49 (56.3%) 8 (9.2%) Suicidal Ideation 77 0.41% 74 (96.1%) 76 (98.7%) 2 (2.6%) 66 (85.7%) 9 (11.7%) Eating Disorder 58 0.31% 57 (98.3%) 33 (56.9%) 24 (41.4%) 32 (55.2%) 2 (3.4%) PTSD 54 0.28% 52 (96.3%) 42 (77.8%) 4 (7.4%) 45 (83.3%) 5 (9.3%) OCD 49 0.26% 47 (95.9%) 29 (59.2%) 10 (20.4%) 37 (75.5%) 2 (4.1%) Substance Use Disor- der 46 0.24% 45 (97.8%) 38 (82.6%) 3 (6.5%) 41 (89.1%) 2 (4.3%) Autism Spectrum Dis- order 41 0.22% 38 (92.7%) 24 (58.5%) 8 (19.5%) 27 (65.9%) 6 (14.6%) Self Harm 34 0.18% 33 (97.1%) 34 (100.0%) 0 (0.0%) 30 (88.2%) 4 (11.8%) Infectious Disease 487 2.57% 454 (93.2%) 257 (52.8%) 58 (11.9%) 350 (71.9%) 79 (16.2%) COVID-19 341 1.80% 326 (95.6%) 195 (57.2%) 10 (2.9%) 307 (90.0%) 24 (7.0%) HIV 105 0.55% 97 (92.4%) 42 (40.0%) 42 (40.0%) 40 (38.1%) 23 (21.9%) Malaria 41 0.22% 31 (75.6%) 20 (48.8%) 6 (14.6%) 3 (7.3%) 32 (78.0%) Maternal & Repro-ductive 456 2.40% 329 (72.1%) 281 (61.6%) 36 (7.9%) 113 (24.8%) 307 (67.3%) Pregnant 192 1.01% 151 (78.6%) 109 (56.8%) 25 (13.0%) 75 (39.1%) 92 (47.9%) Postpartum 192 1.01% 136 (70.8%) 111 (57.8%) 8 (4.2%) 36 (18.8%) 148 (77.1%) Continued on next page 35 Table S3 â Continued from previous page Condition N % of All Single-turn Mod-High Risk Generation 1 Generation 2 Generation 3 Postpartum Depres- sion 72 0.38% 42 (58.3%) 61 (84.7%) 3 (4.2%) 2 (2.8%) 67 (93.1%) Cardiovascular 356 1.88% 313 (87.9%) 170 (47.8%) 127 (35.7%) 157 (44.1%) 72 (20.2%) Hypertension 201 1.06% 172 (85.6%) 111 (55.2%) 39 (19.4%) 111 (55.2%) 51 (25.4%) Heart Failure 70 0.37% 65 (92.9%) 20 (28.6%) 47 (67.1%) 15 (21.4%) 8 (11.4%) Atrial Fibrillation 53 0.28% 48 (90.6%) 23 (43.4%) 31 (58.5%) 15 (28.3%) 7 (13.2%) Coronary Artery Dis- ease 32 0.17% 28 (87.5%) 16 (50.0%) 10 (31.2%) 16 (50.0%) 6 (18.8%) Respiratory 277 1.46% 244 (88.1%) 163 (58.8%) 60 (21.7%) 154 (55.6%) 63 (22.7%) Asthma 213 1.12% 195 (91.5%) 127 (59.6%) 40 (18.8%) 134 (62.9%) 39 (18.3%) Tuberculosis 27 0.14% 16 (59.3%) 18 (66.7%) 6 (22.2%) 4 (14.8%) 17 (63.0%) COPD 25 0.13% 21 (84.0%) 17 (68.0%) 4 (16.0%) 14 (56.0%) 7 (28.0%) Lung Cancer 12 0.06% 12 (100.0%) 1 (8.3%) 10 (83.3%) 2 (16.7%) 0 (0.0%) Oncology 275 1.45% 261 (94.9%) 70 (25.5%) 168 (61.1%) 83 (30.2%) 24 (8.7%) Breast Cancer 92 0.48% 87 (94.6%) 17 (18.5%) 63 (68.5%) 21 (22.8%) 8 (8.7%) Skin Cancer 90 0.47% 87 (96.7%) 27 (30.0%) 53 (58.9%) 33 (36.7%) 4 (4.4%) Prostate Cancer 32 0.17% 29 (90.6%) 7 (21.9%) 20 (62.5%) 8 (25.0%) 4 (12.5%) Cervical Cancer 31 0.16% 30 (96.8%) 6 (19.4%) 24 (77.4%) 5 (16.1%) 2 (6.5%) Colorectal Cancer 30 0.16% 28 (93.3%) 13 (43.3%) 8 (26.7%) 16 (53.3%) 6 (20.0%) Renal & Urologic 55 0.29% 39 (70.9%) 36 (65.5%) 7 (12.7%) 14 (25.5%) 34 (61.8%) Chronic Kidney Dis- ease 55 0.29% 39 (70.9%) 36 (65.5%) 7 (12.7%) 14 (25.5%) 34 (61.8%) 36 A.5 Query Profile Tables This section presents detailed query profile tables for each individual dataset, showing the distribu- tion of queries across five key dimensions: intent, topic, context richness, clinical complexity, and data integration. All profiles are based on consumer queries only. 37 A.5.1 HealthSearchQA (N=3,173) Dimension/CategoryCount (%) 1. Intent Distribution Education2,645 (83.4%) Education Explainer2,635 (99.6%) Basic Science10 (0.4%) Symptom Check380 (12.0%) Management Plan235 (61.8%) Differential Diagnosis107 (28.2%) Triage Disposition38 (10.0%) Non-Health50 (1.6%) Offtopic Nonhealth50 (100.0%) General Health Advice47 (1.5%) Nutrition/Diet19 (40.4%) Cosmeceuticals/Topicals10 (21.3%) Sleep Hygiene8 (17.0%) Condition Management25 (0.8%) Chronic Care Support21 (84.0%) Acute Flare Management4 (16.0%) Other Intents (Aggregates 4 items)26 (0.8%) 2. Topic Distribution (Top 5) Brain & Nerves366 (11.5%) Other/Unspecified212 (57.9%) Neuropathy43 (11.8%) Cognitive Changes39 (10.7%) Skin & Hair366 (11.5%) Other/Unspecified169 (46.2%) Infections64 (17.5%) Rash56 (15.3%) Muscles, Bones & Joints274 (8.6%) Other/Unspecified180 (65.7%) Arthritis32 (11.7%) Sprains & Strains19 (6.9%) Digestive & Nutrition252 (7.9%) Other/Unspecified128 (50.8%) Liver Disease33 (13.1%) Reflux/Heartburn19 (7.5%) Infections (General)252 (7.9%) Other/Unspecified198 (78.6%) Travel-Related20 (7.9%) Fever (Unspecified)16 (6.3%) 3. Context Richness Conversation Structure Single Turn3,173 (100.0%) Narrative Detail Short3,173 (100.0%) Context Depth Continued on next page 38 Table S4 â Continued from previous page Dimension/CategoryCount (%) Low3,171 (99.9%) High2 (0.1%) 4. Clinical Complexity Risk Level Low3,093 (97.5%) Moderate70 (2.2%) High10 (0.3%) User Type Consumer3,173 (100.0%) Population Adult (Unspecified)3,104 (97.8%) Peds Unspecified40 (1.3%) Pediatric (Under 5)26 (0.8%) Adult 65Plus2 (0.1%) Peds 5To171 (0.0%) Language English3,173 (100.0%) Language Complexity Lay3,069 (96.7%) Technical104 (3.3%) Query Subject General2,944 (92.8%) Self218 (6.9%) Child11 (0.3%) Personal Health Query No2,955 (93.1%) Yes218 (6.9%) 5. Data Integration Objective Data Present Yes8 (0.2%) No3,165 (99.8%) Objective Data Types Diagnoses5 (0.2%) Vitals (Basic)3 (0.1%) 39 A.5.2 MashQA Test (N=3,490) Dimension/CategoryCount (%) 1. Intent Distribution Education2,389 (68.5%) Education Explainer2,250 (94.2%) Basic Science139 (5.8%) Medication Information332 (9.5%) Side Effects157 (47.3%) Selection112 (33.7%) Dosing45 (13.6%) General Health Advice240 (6.9%) Nutrition/Diet72 (30.0%) Supplements/Nutraceuticals64 (26.7%) Fitness/Exercise40 (16.7%) Condition Management157 (4.5%) Chronic Care Support141 (89.8%) Risk/Prognosis10 (6.4%) Acute Flare Management6 (3.8%) Prevention/Screening117 (3.4%) Lifestyle Prevention89 (76.1%) Screening Schedule18 (15.4%) Vaccination10 (8.6%) Other Intents (Aggregates 5 items)255 (7.3%) 2. Topic Distribution (Top 5) Cancer524 (15.0%) Other/Unspecified291 (55.5%) Lung75 (14.3%) Breast74 (14.1%) Muscles, Bones & Joints330 (9.5%) Arthritis217 (65.8%) Other/Unspecified41 (12.4%) Back & Neck Pain30 (9.1%) Brain & Nerves283 (8.1%) Other/Unspecified116 (41.0%) Headache/Migraine100 (35.3%) Neuropathy30 (10.6%) Digestive & Nutrition272 (7.8%) Other/Unspecified147 (54.0%) Inflammatory Bowel Disease (IBD)39 (14.3%) Liver Disease32 (11.8%) Skin & Hair272 (7.8%) Other/Unspecified130 (47.8%) Wounds38 (14.0%) Eczema29 (10.7%) 3. Context Richness Conversation Structure Single Turn3,490 (100.0%) Continued on next page 40 Table S5 â Continued from previous page Dimension/CategoryCount (%) Narrative Detail Short3,490 (100.0%) Context Depth Low3,484 (99.8%) High6 (0.2%) 4. Clinical Complexity Risk Level Low3,441 (98.6%) Moderate44 (1.3%) High5 (0.1%) User Type Consumer3,490 (100.0%) Population Adult (Unspecified)3,297 (94.5%) Peds Unspecified115 (3.3%) Pediatric (Under 5)57 (1.6%) Adult 65Plus13 (0.4%) Peds 5To178 (0.2%) Language English3,490 (100.0%) Language Complexity Lay3,157 (90.5%) Technical333 (9.5%) Query Subject General3,215 (92.1%) Self223 (6.4%) Child52 (1.5%) Personal Health Query No3,267 (93.6%) Yes223 (6.4%) 5. Data Integration Objective Data Present Yes133 (3.8%) No3,357 (96.2%) Objective Data Types Diagnoses110 (3.1%) Medications21 (0.6%) Procedures15 (0.4%) Labs1 (0.0%) 41 A.5.3 MedRedQA Test (N=5,081) Dimension/CategoryCount (%) 1. Intent Distribution Symptom Check2,656 (52.3%) Differential Diagnosis1,235 (46.5%) Management Plan774 (29.1%) Triage Disposition647 (24.4%) Tests and Results702 (13.8%) Test Interpretation593 (84.5%) Test Selection109 (15.5%) Medication Information608 (12.0%) Side Effects240 (39.5%) Selection134 (22.0%) Interactions119 (19.6%) Condition Management558 (11.0%) Risk/Prognosis212 (38.0%) Chronic Care Support209 (37.5%) Acute Flare Management137 (24.6%) Administrative Meta182 (3.6%) Navigation Referral117 (64.3%) Other Admin40 (22.0%) Insurance Billing13 (7.1%) Other Intents (Aggregates 5 items)375 (7.4%) 2. Topic Distribution (Top 5) Skin & Hair761 (15.0%) Rash194 (25.5%) Wounds175 (23.0%) Other/Unspecified157 (20.6%) Digestive & Nutrition521 (10.2%) Other/Unspecified169 (32.4%) Liver Disease75 (14.4%) Abdominal Pain69 (13.2%) Brain & Nerves454 (8.9%) Other/Unspecified118 (26.0%) Headache/Migraine76 (16.7%) Neuropathy63 (13.9%) Infections (General)435 (8.6%) Other/Unspecified222 (51.0%) COVID-19139 (31.9%) Fever (Unspecified)36 (8.3%) Heart & Circulation428 (8.4%) Other/Unspecified157 (36.7%) Chest Pain68 (15.9%) Palpitations59 (13.8%) 3. Context Richness Conversation Structure Single Turn5,072 (99.8%) Continued on next page 42 Table S6 â Continued from previous page Dimension/CategoryCount (%) Multi Turn9 (0.2%) Narrative Detail Detailed4,804 (94.5%) Short277 (5.5%) Context Depth High4,461 (87.8%) Low620 (12.2%) 4. Clinical Complexity Risk Level Moderate2,942 (57.9%) Low1,639 (32.3%) High500 (9.8%) User Type Consumer5,081 (100.0%) Population Adult (Unspecified)4,394 (86.5%) Peds 5To17468 (9.2%) Adult 65Plus117 (2.3%) Pediatric (Under 5)97 (1.9%) Peds Unspecified5 (0.1%) Language English5,081 (100.0%) Language Complexity Lay4,461 (87.8%) Technical620 (12.2%) Query Subject Self4,327 (85.2%) Parent186 (3.7%) Child160 (3.1%) Partner158 (3.1%) Other Relative116 (2.3%) General85 (1.7%) Friend Acquaintance48 (0.9%) Patient1 (0.0%) Personal Health Query Yes4,328 (85.2%) No753 (14.8%) 5. Data Integration Objective Data Present Yes3,646 (71.8%) No1,435 (28.2%) Objective Data Types Diagnoses2,230 (43.9%) Medications2,044 (40.2%) Labs898 (17.7%) Procedures879 (17.3%) Imaging659 (13.0%) Continued on next page 43 Table S6 â Continued from previous page Dimension/CategoryCount (%) Vitals (Basic)383 (7.5%) Vitals (Wearable)35 (0.7%) 44 A.5.4 HealthBench Main (N=3,692) Dimension/CategoryCount (%) 1. Intent Distribution Symptom Check1,384 (37.5%) Management Plan681 (49.2%) Triage Disposition392 (28.3%) Differential Diagnosis311 (22.5%) Medication Information729 (19.8%) Selection443 (60.8%) Dosing170 (23.3%) Side Effects72 (9.9%) Condition Management297 (8.0%) Chronic Care Support227 (76.4%) Risk/Prognosis44 (14.8%) Acute Flare Management26 (8.8%) Education285 (7.7%) Education Explainer280 (98.2%) Basic Science5 (1.8%) General Health Advice270 (7.3%) Nutrition/Diet115 (42.6%) Supplements/Nutraceuticals60 (22.2%) Fitness/Exercise40 (14.8%) Other Intents (Aggregates 5 items)727 (19.7%) 2. Topic Distribution (Top 5) Infections (General)425 (11.5%) Other/Unspecified218 (51.3%) Travel-Related81 (19.1%) Fever (Unspecified)66 (15.5%) Brain & Nerves344 (9.3%) Other/Unspecified94 (27.3%) Headache/Migraine88 (25.6%) Dizziness/Vertigo59 (17.1%) Digestive & Nutrition339 (9.2%) Other/Unspecified153 (45.1%) Abdominal Pain42 (12.4%) Reflux/Heartburn33 (9.7%) Skin & Hair315 (8.5%) Wounds104 (33.0%) Other/Unspecified53 (16.8%) Rash46 (14.6%) Mental Health & Psychiatry261 (7.1%) Postpartum Mental Health79 (30.3%) Anxiety49 (18.8%) Depression47 (18.0%) 3. Context Richness Conversation Structure Single Turn2,193 (59.4%) Continued on next page 45 Table S7 â Continued from previous page Dimension/CategoryCount (%) Multi Turn1,499 (40.6%) Narrative Detail Short3,237 (87.7%) Detailed455 (12.3%) Context Depth Low3,174 (86.0%) High518 (14.0%) 4. Clinical Complexity Risk Level Low2,161 (58.5%) Moderate1,189 (32.2%) High342 (9.3%) User Type Consumer3,692 (100.0%) Population Adult (Unspecified)3,249 (88.0%) Pediatric (Under 5)159 (4.3%) Peds Unspecified121 (3.3%) Peds 5To1793 (2.5%) Adult 65Plus70 (1.9%) Language English3,010 (81.5%) Non-English682 (18.5%) Language Complexity Lay3,470 (94.0%) Technical222 (6.0%) Query Subject Self2,247 (60.9%) General870 (23.6%) Child309 (8.4%) Friend Acquaintance99 (2.7%) Parent75 (2.0%) Other Relative67 (1.8%) Partner23 (0.6%) Patient2 (0.1%) Personal Health Query Yes2,247 (60.9%) No1,445 (39.1%) 5. Data Integration Objective Data Present Yes863 (23.4%) No2,829 (76.6%) Objective Data Types Diagnoses546 (14.8%) Medications257 (7.0%) Procedures99 (2.7%) Labs82 (2.2%) Continued on next page 46 Table S7 â Continued from previous page Dimension/CategoryCount (%) Vitals (Basic)66 (1.8%) Imaging43 (1.2%) Vitals (Wearable)1 (0.0%) 47 A.5.5 GoogleFitbit Sleep (N=1,521) Dimension/CategoryCount (%) 1. Intent Distribution General Health Advice1,521 (100.0%) Sleep Hygiene1,521 (100.0%) 2. Topic Distribution (Top 5) Holistic Health & Wellness1,521 (100.0%) Sleep & Lifestyle1,521 (100.0%) 3. Context Richness Conversation Structure Single Turn1,521 (100.0%) Narrative Detail Detailed1,521 (100.0%) Context Depth High1,521 (100.0%) 4. Clinical Complexity Risk Level Low1,521 (100.0%) User Type Consumer1,521 (100.0%) Population Adult (Unspecified)1,092 (71.8%) Adult 65Plus429 (28.2%) Language English1,521 (100.0%) Language Complexity Lay1,521 (100.0%) Query Subject Self1,521 (100.0%) Personal Health Query Yes1,521 (100.0%) 5. Data Integration Objective Data Present Yes1,521 (100.0%) Objective Data Types Vitals (Wearable)1,521 (100.0%) 48 A.5.6 GoogleFitbit Fitness (N=1,750) Dimension/CategoryCount (%) 1. Intent Distribution General Health Advice1,750 (100.0%) Fitness/Exercise1,750 (100.0%) 2. Topic Distribution (Top 5) Holistic Health & Wellness1,750 (100.0%) Sleep & Lifestyle1,750 (100.0%) 3. Context Richness Conversation Structure Single Turn1,750 (100.0%) Narrative Detail Detailed1,750 (100.0%) Context Depth High1,750 (100.0%) 4. Clinical Complexity Risk Level Low1,750 (100.0%) User Type Consumer1,750 (100.0%) Population Adult (Unspecified)1,520 (86.9%) Adult 65Plus230 (13.1%) Language English1,750 (100.0%) Language Complexity Lay1,750 (100.0%) Query Subject Self1,750 (100.0%) Personal Health Query Yes1,750 (100.0%) 5. Data Integration Objective Data Present Yes1,750 (100.0%) Objective Data Types Vitals (Wearable)1,750 (100.0%) 49