Paper deep dive
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Maker's Dilemma
Reva Schwartz, Gabriella Waters
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 4:35:08 AM
Summary
The paper introduces the Forum for Real-World AI Measurement and Evaluation (FRAME), a centralized infrastructure designed to bridge the gap between abstract model-centric benchmarks and fragmented user-centric studies. FRAME addresses the 'decision-maker's dilemma' by capturing 'user entropy'âthe variability in how real users interact with AIâthrough a Testing Sandbox and a Metrics Hub. This approach generates systematic, context-rich evidence (Layer 3) to help organizational leaders understand real-world AI impacts, risks, and value accumulation.
Entities (10)
Relation Signals (9)
FRAME â addresses â Decision-Maker's Dilemma
confidence 95% ¡ FRAME aims to address this gap by combining large-scale trials of AI systems with structured observation... The Forum establishes two core assets to achieve this
FRAME â comprises â Testing Sandbox
confidence 95% ¡ The Forum establishes two core assets to achieve this: a Testing Sandbox that captures AI-in-use under real workflows at scale
FRAME â comprises â Metrics Hub
confidence 95% ¡ a Metrics Hub that translates those traces into actionable indicators
User Entropy â coinedby â Gabriella Waters
confidence 92% ¡ Gabriella Waters has coined this variability as âuser entropy,â and it can be a significant factor in whether a deployment succeeds or fails
Testing Sandbox â captures â User Entropy
confidence 90% ¡ The testing sandbox is a controlled but realistic environment that uses large-scale remote participant panels to evaluate commercial AI systems... offering a clear window into user entropy in practice.
FRAME â generates â Layer 3: Systematic Knowledge
confidence 90% ¡ Layer 3: Systematic knowledge Longitudinal field studies, multi-site monitoring, and other real-world testing strategies generate systematic knowledge... FRAME aims to collect it at scale via systematic observation
Metrics Hub â translates â User Entropy
confidence 90% ¡ The metrics hub is a translation layer that converts sandbox evaluation outcomes into decision-ready evidence about AI system behavior with real users
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Organizational leaders are being asked to make high-stakes decisions about AI deployment without dependable evidence of what these systems actually do in the environments they oversee. The predominant AI evaluation ecosystem yields scalable but abstract metrics that reflect the priorities of model development. By smoothing over the heterogeneity of real-world use, these model-centric approaches obscure how behavior varies across users, workflows, and settings, and rarely show where risk and value accumulate in practice. More user-centric studies reveal rich contextual detail, yet are fragmented, small-scale and loosely coupled to the mechanisms that shape model behavior. The Forum for Real-World AI Measurement and Evaluation (FRAME) aims to address this gap by combining large-scale trials of AI systems with structured observation of how they are used in context, the outcomes they generate, and how those outcomes arise. By tracing the path from an AI system's output through its practical use and downstream effects, FRAME turns the heterogeneity of AI-in-use into a measurable signal rather than a trade-off for achieving scale. The Forum establishes two core assets to achieve this: a Testing Sandbox that captures AI-in-use under real workflows at scale and a Metrics Hub that translates those traces into actionable indicators.
Tags
Links
- Source: https://arxiv.org/abs/2603.13294v4
- Canonical: https://arxiv.org/abs/2603.13294v4
Trouble viewing inline? Open PDF directly â
Full Text
98,251 characters extracted from source content.
Expand or collapse full text
Real-World AI Evaluation: How FRAME Generates Systematic Evidence to Resolve the Decision-Makerâs Dilemma Reva SchwartzGabriella Waters reva@civitaas.comgwaters@vsu.edu Civitaas InsightsCenter for Responsible AI, Virginia State University, and Civitaas Insights Abstract Organizational leaders are being asked to make high-stakes decisions about AI deployment without dependable ev- idence of what these systems actually do in the environments they oversee. The predominant AI evaluation ecosystem yields scalable but abstract metrics that reflect the priorities of model development. By smoothing over the hetero- geneity of real-world use, these model-centric approaches obscure how behavior varies across users, workflows, and settings, and rarely show where risk and value accumulate in practice. More user-centric studies reveal rich contextual detail, yet are fragmented, small-scale and loosely coupled to the mechanisms that shape model behavior. The Forum for Real-World AI Measurement and Evaluation (FRAME) aims to address this gap by combining large-scale trials of AI systems with structured observation of how they are used in context, the outcomes they generate, and how those outcomes arise. By tracing the path from an AI systemâs output through its practical use and downstream effects, FRAME turns the heterogeneity of AI-in-use into a measurable signal rather than a trade-off for achieving scale. The Forum establishes two core assets to achieve this: a Testing Sandbox that captures AI-in-use under real workflows at scale and a Metrics Hub that translates those traces into actionable indicators. 1 FRAME Purpose and Activities 1.1 Background Recent reporting shows that broad-scale generative AI deployments have produced uneven results, with many initiatives stalling or failing to deliver meaningful value [21]. Policymakers and organizational leaders need con- crete expectations about what these deployments de- liver in their own settings so they can make compar- isons across sites and sectors[68, 37, 106]. Today, these stakeholders have only a small set of metrics to rely on, typically leaning on benchmarks of AI model capabili- ties because they are scalable and quantitative. Yet even as AI systems are being globally adopted and embedded in everyday life, benchmarks remain anchored in test- ing paradigms that bear little resemblance to real-world use. The benchmarking ecosystem is built to answer questions about optimization and safety in model devel- opment, with performance metrics based on standard- ized tasks under optimized conditions. Since it pro- duces scant evidence about how AI tools are actually taken up, worked around, or ignored in practice, exist- ing benchmarks map poorly onto deployment-level ques- tions about adoption, safeguards, oversight, and distri- butional impacts. This contributes to a structural mis- match between the evidence the ecosystem produces and the evidence that social, economic, and cultural decision-makers need. A key weakness of benchmarking is its inability to model or account for the unpredictable and variable ways people leverage AI technology in context, includ- ing how they express needs, interpret and act on out- 1 arXiv:2603.13294v4 [cs.CY] 29 Mar 2026 puts, and adapt and repurpose systems. Gabriella Wa- ters has coined this variability as âuser entropy,â and it can be a significant factor in whether a deployment suc- ceeds or fails[48, 6, 81, 92, 116]. Model-centric evalu- ations essentially treat user entropy as statistical noise to be eliminated rather than as a primary measurement signal[10, 71, 118]. User-centric studies generate comple- mentary, context-rich evidence, but usually within nar- row sites or populations, and disconnected from dom- inant benchmarking pipelines[11, 47, 67]. These two strands of evaluationâscalable model-centric tests and local user-centric pilotsâare rarely connected system- atically, leaving decision-makers without evidence that both reflects the diversity and unpredictability of real use and remains comparable across many deployments. Current limitations of connecting model behavior to the real-world uses, failures, and adaptations that drive AIâs higher-order effects[96] place leaders in a âdecision-makerâs dilemmaâ: they must reinterpret ab- stract scores for the contexts they oversee, without clear evidence about where risk and value lie or how they are likely to manifest[19, 64, 114, 109]. Just as clinical trial results cannot tell a city planner how a virus will move through a particular school system, benchmark scores and lab tests do not show chief information officers how large language models will alter workflows in their own organizations[19, 64, 114, 109]. Dashboards and trainings can help people interpret such metrics, but they cannot substitute for systematic evidence about how AI shapes everyday work, culture and the broader society. Figure 1 shows how modeling user entropy can help to connect abstract metrics to concrete deployment decisions. Figure 1: Modeling user entropy can support concrete AI deployment decisions. (Generative artificial intelligence was used to support the creation of this graphic representing the authorsâ own ideas, data, and words on this topic.) 1.2 The Next Generation of AI Evaluation Anchored at Virginia State University, the Forum for Real-World AI Measurement and Evaluation (FRAME) was established to address these challenges by measuring system behavior in real contexts, not just on optimized tests. FRAMEâs membership forms a global, interdis- ciplinary coalition spanning measurement science, ma- chine learning, social science, and the humanities across academia, industry, and government. Instead of each or- ganization building its own costly testing stack, the Fo- rum runs large-scale evaluationsâsupported by sponsor- ing organizationsâso sectors can assess AI in conditions similar to their own without exposing proprietary data. 2 FRAME also acts as a "methods lab" by pioneering ap- proaches to capture user entropy at scale and translate observations into comparable indicators across sites. FRAME evaluations complement existing capability benchmarks, reinforcement learning from human feed- back (RLHF) pipelines, and adversarial testing. While those tools illuminate model behavior and task perfor- mance, FRAME foregrounds user entropy to clarify how AI actually plays out in real-world use.This gives decision-makers evidence to interpret what AI-in-use means for their own purposes and to align deployments more closely with institutional goals and public benefit. To conduct this work, FRAME is developing a centralized research and evaluation infrastructure built around two core components: ⢠The testing sandbox is a controlled but realistic environment that uses large-scale remote partici- pant panels to evaluate commercial AI systems un- der task-driven scenarios. Rather than relying on people as labelers, panelists provide descriptive ac- counts of how they leverage, repurpose, or aban- don AI tools, offering a clear window into user entropy in practice. The sandbox maintains strict human-subjects protections and relies on carefully designed proxy tasks to safely measure high-stakes risks without exposing participants to harm or sen- sitive content. ⢠The metrics hub is a translation layer that converts sandbox evaluation outcomes into decision-ready evidence about AI system behavior with real users in real contexts. It produces key indicators of system utility, friction, and resilienceâoffering a grounded alternative to abstract leaderboard scores. Released on a regular cadence, these indicators align with stakeholdersâ deployment needs and complement existing capability, safety, and compliance metrics, adding a clear, real-world performance layer to to- dayâs evaluation landscape. This white paper describes FRAMEâs mission to help stakeholders better understand AIâs value in the real world, clarifies structural limitations in current evalua- tion approaches, and presents the centralized testing in- frastructure for conducting large-scale trials with struc- tured, context-rich observation. A concise glossary of key terms supports consistent language and shared un- derstanding across the real-world AI evaluation ecosys- tem. 2 Structural Limitations of the Current Ecosystem for Decision Making AI use in the real world is not a set of isolated pass/fail tests, but an ongoing exchange in which human inputs and model outputs continuously shape one another and affect public life, including people who never touch AI directly.[102, 94, 130]. Deployments exhibit two inter- acting sources of variation: ⢠User entropy (the âhostâ variable): The het- erogeneity in how people phrase questions, adapt workflows, and interpret outputs; in the âepidemi- ologyâ of AI, this is the host environment. ⢠Model stochasticity (the âagentâ variable): The inherent randomness in generative outputs. Small changes in decoding parameters or routing through mixture-of-experts layers can yield different re- sponses to seemingly identical prompts, and this variability never fully disappears[6]: this is the agent variable. In the wild, user entropy and model stochasticity com- pound one another[23, 70]: different users frame ques- tions differently, and the model responds stochastically each time. The result is not a single, stable pattern of behavior but a moving distribution of interactions that can only be understood by modeling higher-order ef- fects over time, across many users and settings. This turns AI deployments into wicked problems: complex socio-technical interventions with distributed causes and effects and dynamically evolving feedback loops[91]. If evaluation is to serve deployment-level decision-making, it must move beyond certifying model performance to account for real-world variation, and generate evidence about how these tools are actually used, by whom, in which contexts, and with what consequences. 2.1 A Missing Layer to Support Sensemaking Research on expertâlay communication shows that a core bottleneck in real-world decision-making is rarely a lack of technical detail. Instead, reliance on low-context, 3 highly technical outputs can increase interpretive am- biguity, making it harder for decision-makers to know what to do with the information[28, 90]. For example, accuracy metrics reported without contextâsuch as a sin- gle benchmark score or âhallucinationâ rateâmay appear precise but offer little guidance about what those num- bers mean for specific users, tasks, or settings. What leaders and other decision-makers often need most is sensemaking: the ability to integrate complex signals into simple, actionable âgistsâ about value and risk [36, 41, 89]. User entropy can serve as a central signal for sensemaking, since it captures how people actually ap- propriate, adapt, and interpret AI in the real world. A public health analogy helps clarify what is missing from current evaluation practices. Today, most AI eval- uation resembles a clinical trial, where lab-style studies test a âcompoundâ (the model) in tightly controlled en- vironments to see whether it meets accuracy targets or predefined safety thresholds. Sensemaking, however, re- quires a layer of evidence more akin to epidemiology and post-market surveillance. This additional layer can track what happens when AI systems are used in everyday set- tings, including side effects, adaptation, and longer-term outcomes across diverse contexts[96]. 2.2 Methodological Pitfalls: AI Evaluation that Mirrors Development Many evaluation methods overlook user entropy because they inherit the abstraction logic used in model training. Dominant approaches originate from developer-centric practices that prioritize reproducibility and model im- provement, but strip away the contextual detail that deployment-level sensemaking depends on. By collaps- ing the user entropy side of real-world deployment into a single âuse case,â current approaches lock evaluation into sterile test sets and scripted in silico runs that capture only prompts and responses[30, 60], and leave deployment-facing questions under-specified. This decontextualization process also ends up reward- ing performance in synthetic settings instead of every- day use, widening the gap between existing measures and the information needed to make sense of what happens across sites [30, 116]. As illustrated in Fig- ure 2, the current focus on model optimization leaves Figure 2: The current evaluation ecosystem uses methods that mirror the traditional AI development lifecycle, of- ten neglecting user entropy and suppressing the context needed to make sense of outcomes for decision-making. (Generative artificial intelligence was used to support the creation of this graphic representing the authorsâ own ideas, data, and words on this topic.) decision-makersâ questions about deployment and im- plementation out of the measurement frame[60, 10, 92]. Lacking evidence of AIâs real-world outcomes, organiza- tions struggle to judge whether deployments will gen- erate real value in their own environments, and develop- ers and risk teams must continually recalibrate guardrails and safety policies against shifting requirements [8, 14]. Partial Workarounds The current ecosystem relies on a set of strategies that serve as a substitute for direct, large-scale evidence about what happens in practice. For example, harm assessments often use keyword-matching safety filters that look for the presence of specific terms without regard for their role in context. A term such as âself-harmâ can appear in a public health awareness cam- paign, a news article, or promotional content for a men- tal health app or therapy service. A rigid, context-blind safety metric will flag or block this benign content in the same way it would treat a user in an active mental health crisis or malicious content that encourages self-harm. These real-world failures often stem from an inabil- ity to distinguish the mere presence of a term from its materialized outcomes in context. Systems trained to rely on surface keywords show âshortcutâ behavior: they 4 over-block harmless content and miss genuinely harmful content when keyword distributions shift[22, 110]. This problem is exacerbated for higher-order effectsâsuch as psychological over-reliance on AIâthat build up grad- ually over thousands of everyday interactions. Other strategies exhibit similar structural issues: ⢠Alignment pipelines and preference data: These approaches encode high-level human values into models using âA versus Bâ forced-choice labeling schemes tied to system guardrails and other stack-level constraints. Alignment pipelines are central for tuning systems to specific norms but cover only a narrow band of possible behaviors, leaving many deployment-relevant possibilities insufficiently specified or unmeasured, especially in heterogeneous settings[5, 20, 31, 27, 122]. ⢠Technically scoped approximations: Scripted agents and automated red-teaming rely on hard-coded interaction patterns and fixed prompt distributions[119] to surface specific bugs and toxic outputs. By removing real human variability, these tools can miss cumulative effectsâsuch as changes in reliance on AI systems over time[3, 23, 69, 83, 129, 85]. ⢠Human judgment overlays: These approaches leverage expert reviewers (e.g., lawyers) to label cat- egories of outputs as acceptable or not. Because these judgments are made out of context rather than under real conditions, labels may not reliably pre- dict real-world outcomes or generalize across set- tings and populations[10, 92, 113, 29]. ⢠Aggregation of capability evaluations: Meta-analyses pool benchmark scores across settings.By combining abstract tests without showing how systems actually behave in context, these approaches offer limited support for sense- making and can compound underlying sampling errors from in silico testing[49, 39, 96]. ⢠Governance and compliance tools: Risk tax- onomies, compliance checklists, and explainability dashboards demonstrate alignment with technical or policy requirements, and are well-suited to sys- tem design, documentation, and assurance. To fully support deployment-focused sensemaking and de- cisions, these tools also need details about how sys- tems behave with real users in real contexts.[39, 96, 44]. ⢠Isolated context-aware studies: Context-aware methodsâsuch as usability tests, adoption studies, and domain-specific pilotsâoffer valuable local in- sight into how people interact with systems[11]. Yet they rarely use a shared structure for collecting and organizing findings, making it hard to compare pat- terns across sites or connect specific interactions to broader secondary and tertiary effects[67, 69, 121, 54, 84, 47]. The strategies listed above generate signals that feed back into model development to make systems more usable and safer, but they are often too local- ized, partial, or inconsistent to support deployment-level decision-making.[30] This reflects a structural weak- ness: existing tools can track patterns in inputs and out- puts, but because they struggle to capture the dynamics, feedback loops, and behavioral adaptations that shape AI-in-use, they cannot reliably support deployment-level requirements. Some recurring gaps illustrate why it is so hard to form a coherent picture across measurement strategies and their outcomes: ⢠Obscuring operational friction: Model-centric success metrics can hide everyday friction points; a âcorrectâ output may still require substantial hu- man verification or rework, turning a theoretical ef- ficiency gain into an operational burden[10, 39, 59, 131]. ⢠Ignoring downstream consequences:When safety evaluations treat âhallucinationsâ only as ab- stract factual errors, organizations lose sight of what happens next. For example, do users notice and cor- rect errors or instead trust and propagate them into other materials [4, 107, 63, 69]. ⢠Scores with no context: Aggregate figuresâsuch as an â11% hallucination rateâ or a single safety scoreârarely indicate 11% of what, for whom, or un- der which conditions, and models that look similar on these metrics can behave very differently across subpopulations and settings [30]. ⢠Masking contextual harms: Benchmarks tuned for narrow accuracy are poorly suited for genera- tive systems, where there is no single ârightâ para- 5 graph, image, or audio clip. Systems can easily be factually correct yet still harmful. For example, a public-service chatbot can produce accurate infor- mation in a format that is inaccessible to people with disabilities[57, 39, 78]. 2.3 AI Evidence Layers The ecosystem limitations described above highlight the need for connective tissue between existing methods. Systematic, deployment-facing knowledge must link ab- stract evaluations to contextual insight, so evidence can travel across sites while remaining grounded in real use. To address this need, we distinguish three layers of evi- dence. Layer 1: Abstract knowledge Benchmarks, align- ment pipelines, and related tools produce abstract knowl- edge about what a model can do in principle[66, 39]. Like clinical lab results, these tests assess primary effects and answer development-focused questions, such as: ⢠âCan the model solve algebraic equations on this standardized test set?â ⢠âDid the model resist adversarial attacks during au- tomated red-teaming?â ⢠âDid it score higher on âhelpfulnessâ in generic pref- erence rankings?â Layer 2: Contextual knowledge Methods such as user research,audits,domain-specific pilots, red-teaming, and incident reports produce contex- tual knowledge about outcomes in particular settings [11, 32, 34, 54]. Analogous to medical case reports, they surface secondary effectsânear-term impacts such as workflow shifts or harmsâand answer questions like: ⢠âWhich responses did users prefer in this usability test?â ⢠âAre harms or unintended consequences occurring in this deployment?â ⢠âWhat failures or workarounds were logged during this pilot?â Layer 3: Systematic knowledge Longitudinal field studies, multi-site monitoring, and other real-world test- ing strategies generate systematic knowledge about cu- mulative impacts across organizations and over time[69, 96, 121]. Like epidemiological surveillance, these ap- proaches can show how AI is used, by whom, in which situations, and with what consequences. By bringing scalable testing into real settings and combining it with structured observation, user entropy can be treated as signal rather than noise. Resulting evidence from such tests can help decision-makers answer questions such as: ⢠âIs this tool creating rework for my staff, or actually saving time?â ⢠âAre employees over-relying on the tool and losing critical skills?â ⢠âIs the tool shifting liability to our frontline work- ers?â ⢠âWho is consistently benefiting, and who is absorb- ing new risks?â A recent field study of generative AI in the work- place illustrates these layers in practice [127]. When only Layer 1 abstract metrics about generative AI capa- bilities are available, organizations might presume pro- ductivity gains. The workplace study used observational approaches to corroborate those presumptions, generat- ing Layer 2 contextual insights that revealed added work and hidden friction for employees instead of freeing up time. Workers juggled multiple AI-mediated tasks, ex- panded their job scope, and experienced more intense, fragmented work. This "contextual knowledge" offers leaders the kind of evidence they need to manage AI use locally [127]. Real-world AI methods can build on such studies to generate systematic knowledge (Layer 3). These eval- uations can link system behavior, the context in which it operates, and real-world outcomes at scale to surface higher-order effects that emerge over time. They build evidence to support both organizational decision-making and cross-sector claims [96]. Figure 3 shows these layers, using the workplace study as an example. Analogies from Consumer Technology Systematic, cross-context measures have been used to understand the real-world impacts of other technologies: ⢠Smartphone telemetry: Large-scale telemetry and screen-time data revealed interaction patterns that individual user studies missed, linking inten- sive smartphone use to outcomes such as sleep qual- ity and attentional performance[33, 25]. This ev- 6 Figure 3: An example of how three knowledge layers build up evidence across contexts to address the decision-makerâs dilemma. (Generative artificial intelligence was used to support the creation of this graphic repre- senting the authorsâ own ideas, data, and words on this topic.) idence informed features like âdo not disturbâ de- faults and digital best-practice guidelines[101]. ⢠Traffic and GPS data: Aggregated GPS traces turn individual trips into measures of route reliability and travel-time variability[86, 124, 38], helping trav- elers choose when and where to travel.[61, 55, 117] 3 How FRAME Works FRAME is designed to complement the current ecosys- tem by modeling AI variation in the real world and pro- duce structured, descriptive evidence about how sys- tems are actually used and the outcomes they produce in context. Because current approaches are unable to reliably simulate such behavior, FRAME aims to collect it at scale via systematic observation of AI-in-use with- out stripping away real-world variability. While tradi- tional human-subject testing is often viewed as slow and manual, FRAMEâs centralized infrastructure and "meth- ods lab" enables speed and reproducibility. 3.1 Testing sandbox FRAMEâs testing sandbox is a controlled yet realistic en- vironment for sponsor-led, scalable trials of commercial systems across domains such as education, media, and consumer services. It is instrumented to capture sys- tem behavior under two parallel streams, with shared test scenarios, detailed logging, and common scoring rubrics to connect them. 1. Remote Participant Panels: a fully dynamic eval- uation stream to observe how thousands of peo- ple actually use AI in realistic, scenario-based tasks. Panelists complete structured, task-based scenarios and report what actually occurred, including how they adapted, repurposed, or worked around system outputs. Individual participants are neither identi- fied nor directly measured. Panelists annotate their own interactions using a common response pro- tocol, capturing how requests are phrased, when confusion arises, and when they give up or switch tools. To minimize participant risk, the sandbox uses proxy scenarios and standard human-subjects protections. 2. Scripted Chatbot Runs:a fully automated, in-silico counterpart that runs the same scenarios as the panels through scripted chatbot sessions to re- veal the modelâs ideal performance under controlled conditions. Test scenarios The mechanism by which FRAME sets evaluation task parameters, scenarios support repeat- able testing across systems, configurations, and user groups. Panelists are free to interact with AI systems as they normally would. FRAME logs telemetry for each session, tracking factors such as whether participants used or ignored the AI, how many steps they took, if they overrode recommendations, and whether they re- verted to old workflows. Over time, these traces build a reusable corpus of AI-in-use that feeds FRAMEâs commu- nity modelsâdigital summaries of how evaluated systems were used in the sandbox. Organizations and policymak- ers can use community models to test policies and com- pare systems across sites without building full evaluation pipelines from scratch. enabling comparisons across de- ployments and informing sector-specific baselines. Scoring: FRAME uses a dual-stream scoring engine in which the object of evaluation is the AI system, not the people using it. Panelist traces and scripted chatbot runs 7 Contextual scoring (human-grounded) LLM-as-a-judge (automated) Panelist traces (user entropy) Human participants work through realistic scenarios; their traces and reports are organized via contextual rubrics into descriptive categories of use, non-use, reliance, and abandonment, making patterns of use visible across contexts. The same participant outputs are also labeled by an LLM using the shared rubrics, producing parallel descriptive segmentations that show how the model summarizes the friction, workarounds, and value that people report. Scripted chatbot runs (abstract knowledge) Scripted chatbot runs on the same scenarios are categorized with the contextual rubrics to describe how the system behaves under controlled conditions, separate from human behavior. Scripted runs are also labeled by an LLM, creating a reusable, fully automated descriptive pathway for rapid comparison and regression-style checks across models and configurations. Table 1: Complementary human-grounded and automated descriptive scoring across panelist traces and scripted chatbot runs help reveal the gap between automated and real-world perspectives of sandbox activity. are both labeled with scenario-specific rubrics to capture descriptive detail about how systems are used, ignored, or abandoned in context. An LLM-as-judge applies the same rubrics automatically to generate parallel descrip- tive codes and scores for system behavior. Comparing scored outputs can highlight differences between automated runs and human-grounded ap- proaches for surfacing details about AIâs value, friction, and risk. Table 1 summarizes the kinds of insights that can be surfaced by applying human-grounded and LLM- as-a-judge scoring to the chatbot and panelist runs. 3.1.1 Illustrative example These insights can help address a common organiza- tional question when deploying AI: which workflows can be fully automated, and where, when, and how is human expertise still necessary to provide local, con- textual, or domain-specific knowledge? For example, transportation-sector decision-makers might want to de- termine which steps in routing, exception handling, and customer notification can safely run end-to-end on AI, and where dispatchers or operators still need to review, correct, or supplement system outputs before they affect passengers or freight. Sponsors from the transportation sector might work with FRAME members to design realistic evaluation sce- narios for deployment to the sandbox. Participant pan- els can use either specialized tasks tailored to transporta- tion experts or more general tasks that any panelist can complete, with high-volume automated trials run in par- allel on the same tasks. By comparing chatbot runs and panelist traces, differences between scripted and real-world behavior âincluding use, non-use, reliance, and abandonmentâbecome visible, providing stakehold- ers with clarity about the kinds of settings and sys- temâuser configurations that are suitable for their con- texts. 3.2 Metrics Hub Basic questions about AI use are still hard to answer. Many organizations must rely on proprietary dash- boards, one-off surveys, or self-reported pilots and can- not say with precision how many people actively use AI tools in their settings, which tools they rely on, how use varies across roles and communities, or whether some groups are consistently routed through lower- performing models. Internal telemetry can show logins or API calls, but it rarely shows how tools are embedded in tasks, when people abandon them, or how usage con- nects to outcomes like access, quality, safety, utility, or 8 Figure 4: Overview of FRAMEâs testing sandbox, metrics hub, and community model architecture. All measures are derived from instrumented sandbox interactions, where panel participants work through realistic scenarios under built-in safety protections. (Generative artificial intelligence was used to support the creation of this graphic repre- senting the authorsâ own ideas, data, and words on this topic.) impact. FRAMEâs metrics hub treats these questions as de- scriptive measurement targets in their own right. Each evaluation reports out which groups used the systems, in which settings, what happened during and after the interaction, and with what results. Indicators illumi- nate patterns of use, non-use, workarounds, and reliance across groups and configurations. They are designed to sit alongside existing capability, safety, and compli- ance metrics, adding a deployment-facing layer rather than replacing current tools. FRAME metrics are gen- erated outside vendor reporting pipelines, using pro- tocols and standing panels governed by academic and community-oriented standards. This external vantage point gives leaders actionable, context-specific insight into how systems behave in practice to corroborate, chal- lenge, or extend company claims. Figure 4 illustrates how FRAMEâs sandbox and metrics hub can support a practical systematic knowledge base for deployment decisions across settings. Selected metrics are released on a regular cadence to complement model leaderboardsâallowing evaluators to validate or challenge benchmark claims. The hub main- tains a focus on inquiry and learning rather than new optimization targets, aligning with recent critiques of leaderboard-centric evaluation and reinforcing a shift to- ward more context-aware, deployment-focused assess- ment [112, 100]. Metrics families To organize these insights, the met- rics hub groups sandbox outcomes into six families of decision-ready indicators that can be compared across sectors and deployments. Some headline indicators from each family will be released through public-facing dash- boards, while more detailed views remain configurable for specific sponsors and use cases. In each family, spe- cific indicators will vary by sector, sponsor questions, and available data. the examples below are not a fixed checklist, but illustrate the kinds of measures FRAME prioritizes 1. System usage and reach Describes actual system adoption and participation. Examples may include, active-use rates of systems across the sandbox, in- tensity and frequency of use, clustering by task type, and patterns of use and non-use across groups, con- texts, and configurations (where consent and pro- 9 tections permit). 2. Task and interaction patterns Characterizes whether and how sandbox participants integrate the systems into their day and how user entropy shows up in practice. This may include: error and rework rates, uplift versus added burden, workarounds, re- purposing, misuse and over-reliance, cognitive load from checking and fixing AI output, and how be- haviors like sycophancy and anthropomorphization build over time for different user groups. 3. Access and performance across contexts De- scribes how system behavior differs across sand- box participant groups and scenarios. This may include:compatibility with assistive technolo- gies,[120] readability and perceived accessibility of outputs, success rates across languages and dialects, and extra work required by specific population groups to attain comparable outcomes. 4. Risk and harm signals in routine use Tracks "early-warning signs" in sandbox scenarios that ap- proximate unsafe or problematic interaction. Exam- ples may include the conditions under which prox- ies for misleading, harmful, or confusing outputs ap- pear in panelist interactions; the time and effort the panelists needed to monitor and correct them; and proxy indicators of heavy or âbinge-likeâ use that may shift panelist information diets over time. 5. Organizational value and ROI Links panelist us- age patterns in FRAME-designed scenarios to exist- ing organizational indicators for sponsor-informed use cases. This may include how panelists used AI- mediated workflows, changes in decision speed, re- work, appeals, and complaints. 6. Downstream cultural and societal impact Sur- faces longer-term, higher-order effects on culture and communities through repeated waves of sand- box studies over time. This may include changes in how panel participants find and share information; how panelists align with or defer to the system over time; and shifting norms around language use. 3.3 Methods and Insights Lab Within the distributed consortium anchored at Virginia State Universityâs Center for Responsible AI, FRAME members draw on technical, social, organizational, and lived expertise to: ⢠Build a shared infrastructure and methodological backbone to systematically observe AI-in-use across many contexts and build up a cumulative evidence base. ⢠Formalize how interactions and outcomes are logged, to capture where AI fits within workflows, when and why people stop using it, and what hap- pens next (rework, delay, reliance, burden shifting). ⢠Turn sandbox findings into clear, decision-ready in- sights that help organizations make sense of AIâs value and risks in their own settings. 4 Reorienting AI Evaluation Toward an Evidence Layer for Decision-Making Industry leaders increasingly acknowledge that tra- ditional capability benchmarksâthough still neces- saryâare no longer sufficient on their own. As Microsoft CEO Satya Nadella recently noted, 1 âthe real benchmark for AI progress is whether it makes a real difference in peopleâs livesâin healthcare, education, and produc- tivity,â rather than how well models perform on chal- lenge tasks. FRAMEâs contribution is to help build an evidence layer that links model-centric metrics to these deployment-level questions about value, risk, and burden in real settings. 4.1 Formalizing Systematic Knowledge for AI Reliable insights into where AI will deliver valueâor what outcomes to expect in practiceâremain scarce. Even organizations well along the adoption curve encounter friction when moving from pilot projects to sustained use [21].From FRAMEâs perspective, the continuing priority on model performance leaves decision-makers with abstract results, ad hoc ex- periments, and reactive problem-solvingârather than grounded guidance about everyday use[111, 87]. While many organizations may hold both abstract 1 https://w.moneycontrol.com/technology/microsoft-ceo-satya- nadella-says-ai-needs-to-help-healthcare-education-and-productivity- not-just-burn-energy-article-13190215.html 10 Table 2: Decision lenses, reference standards, and guiding questions in the AI ecosystem. Decision lensReference standards for performanceGuiding question Real-world evaluation (FRAME) Systematic evidence about AI-in-use, including user entropy and higher-order effects What happens when people actually use the tool in this context (does it create value or new risks/burdens)? Model / research (ML)Labeled data, benchmarks, leaderboardsDid the model get the right answer on this test set? Product engineeringCustomer requirements, feature usage, ticketsDoes the system solve the problem well enough to ship? CompliancePolicies, laws, internal risk registersDid we break a specific rule? UX and usabilityUser feedback, adoption and retention metricsCan people use the interface as designed? Safety and alignmentHuman preference data, preference labels, safety taxonomies, red-teaming results Did the model pass the safety stress test and produce policy-compliant outputs? model-centric metrics and rich contextual knowledge about AI-in-use, they typically lack structured ways to turn that information into evidence that can guide in- stitutional decisions[121]. The field of risk communica- tion emerged from a similar challenge in public health, as raw epidemiological data and specialist reports proved insufficient for policymaking[93, 123, 16, 82]. As a dis- tinct practice, risk communication focuses on translat- ing complex, ambiguous evidence into usable guidance for decision-makers and the public[13, 41, 7]. Real-world AI evaluation can play an analogous role: not a replace- ment for capability or benchmark testing, but a system- atic knowledge layer that links those measurements to the practical choices facing deployment stakeholders. Many organizational policies require Testing, Evalu- ation, Verification, and Validation (TEVV) as a founda- tion for trustworthy AI, yet few can produce evidence at the deployment level[42, 66]. FRAMEâs real-world evaluation lens can help organizations operationalize these frameworksâparticularly the âTâ (Testing) and the second âVâ (Validation) by translating high-level TEVV requirements into concrete evidence of how sys- tems function in real-world contexts[96, 24] that exist- ing safety, alignment, and compliance tools often over- look. Grounding evaluation in real events can also enable the production of higher-quality datasets for training and governance [47, 52, 56], reducing dependence on canned user personas and on inferring behavior solely from the extreme ends of carefully curated test conditions [26, 46]. 4.2 Decision Lenses Across the AI Ecosystem The AI ecosystem is not a single, unified structure, and success looks different depending on who is asking the question. For model developers, success typically means improving system performance. For engineers, it revolves around meeting customer needs and tech- nical requirements. Real-world evaluation adds a third lensâone centered on deployment decisionsâtreating evidence about how AI is actually used as the primary reference point for judging success. Table 2 shows how this real-world lens differs from other organizational functions. 4.3 Three Shifts for Real-World AI Evaluation Practice To build up this real-world lens FRAMEâs ecosystem shifts who evaluation serves, what it measures, and how it is conducted. 4.3.1 Who Evaluations Are Built For: Decision-making Beyond the Stack Generative systems are more than software. They func- tion as products that people use and information media they interpret, trust, or ignore [19, 119, 102, 94]. The same underlying model can be used as a writing assis- tant, a search interface, or an advice source, often within a single workflow [108, 53]. Evaluation infrastructures tend to mirror the audiences they are built to serve, and 11 today those audiences are mostly inside the stackâmodel developers and product teamsârather than the people and institutions closest to deployment. As a result, this variability in how systems are used, governed, and in- centivized is often left out of view. FRAME responds by expanding evaluation to serve decision-makers beyond the stack, treating them as a distinct audience with its own evidence needs. AI Users as Reporters With a primary focus on model training, most benchmarking regimes effectively treat people as measuring instruments: they check whether a model stayed within preset bounds and label or rank outputs based on static preferences (for example, âRe- sponse A is better than Response Bâ) to optimize accu- racy or preference alignment on a fixed target. Systems are tuned to that endpoint even when it does not reflect real use[10, 116, 60, 92]. Yet, accuracy serves as only one part of the equation when aiming to understand real-world behavior. For ex- ample, evaluators need to do more than identify hallu- cinated output. Determining the real impact of halluci- nations requires knowing what people do with that con- tentâwhether they notice it, believe it, correct it, or prop- agate it into later decisionsâand how these choices accu- mulate into broader outcomes such as changes in skill and judgment [98, 126, 125, 2]. While more comprehen- sive information about real-world use is often captured via platform mechanisms such as feedback buttons or in- cident logs, it can serve as a lagging indicator and be prone to survivorship bias. To address these measurement challenges, panel par- ticipants in FRAMEâs sandbox act as observers and re- porters rather than raters, performing AI-mediated tasks and documenting what actually happens. This can in- clude detail about where panelists may improve, adapt, or repurpose AI outputs, quietly work around failures or disengage altogether [26, 46]. The sandbox lever- ages descriptivist methods to identify such patterns amid individual variation [43, 52, 12, 115].In aggregate, panel reports form a descriptive corpus of interaction in which overlapping perspectives yield a scalable, contex- tual view to reveal: ⢠How systems behave, where they create friction, and who is most affected. Figure 5: How different decision lenses position people in AI evaluation.(Generative artificial intelligence was used to support the creation of this graphic representing the authorsâ own ideas, data, and words on this topic.) ⢠How people integrate, adapt, or ignore system out- puts in real workflows. ⢠Where friction, workarounds, or new risks emerge in everyday use. ⢠Which outcomes matter most to different stake- holder groups, and how AI reshapes those outcomes over time. As shown in Table 3, foregrounding the questions decision-makers actually faceâand shifting evaluation methods accordinglyâ can also enhance key forms of va- lidity. 4.3.2 What Evaluations Target: From Accuracy to Higher-Order Effects To support increasing interest in the higher-order effects of system use[40, 88], FRAMEâs sandbox is instrumented to track how impacts accumulate for panelists across re- peated interactions and scenario-based tasks over time. FRAMEâs approach makes these dynamics observable by jointly modeling user entropy and model stochasticity, capturing how variation on both sides shapes real-world performance and surfaces higher-order effects. These outcomes can complement traditional capability bench- marks and inform how synthetic agents are tuned so their prompts, task choices, and abandonment rates more closely reflect real-world behavior [58]. 12 Table 3: How types of validity adapt between in-silico and in situ evaluations. Validity typeModel-centric (in-silico)Real-world evaluation-FRAME (in situ) Content and construct (Are we measuring the right thing?) Weakened by reliance on abstract proxies (such as perplexity or generic preference scores) that miss how people actually express needs in de- ployment[56]. Strengthened through metrics based on ob- served user behavior and stakeholder-defined goals[92]. Ecological (Does the test resemble the real-world use conditions?) Undermined by dependence on âsterileâ prompt-response tests that do not reflect real workflows. Strengthened through realistic scenarios and situated observation that capture friction, workarounds, and constraints. Consequential (What happens downstream when these tools are used?) Underdeveloped by failing to capture how out- puts alter real-world outcomes over time. Strengthened by tracking whether AI use cre- ates value or shifts burden. External (Will the test results hold for my setting?) Uncertain: Lessons from one optimized test set often fail to generalize to different users and contexts over time.[17] Strengthened by testing across diverse groups and settings at scale. DimensionStatus-quo AI EvaluationReal-world AI Evaluation (FRAME) WhoModel builders (inside the stack processes) Deployment stakeholders (outside the stack decision making) WhatIsolated model capabilitiesDownstream impacts and real-world variability HowPrescriptive model-centric checks (AI governance intent) Observing AI at scale in the wild (AI governance reality) Table 4: How FRAMEâs real-world AI approach shifts evaluation practice to support deployment decision-making. 4.3.3 How Evaluation is Done: From Static Tests to Ongoing Observation Current evaluation approaches often serve primarily as compliance checks of âgovernance intentâ (âwas the rule followed?â) Real-world AI evaluation instead generates higher-resolution evidence of AI-in-use, such as whether a tool saves time, adds friction, or shifts burden. Through descriptive tracking of what actually happens as pan- elists work through sandbox scenarios (e.g., "did it save time?"), FRAME can surface "governance reality" and en- able insights for improved decision making. Thisshiftalsotargetsacorefrictionin decision-making:ambiguity.High-context, low-ambiguity information about who is using AI, where, for what, and when helps stakeholders interpret findings against their own goals[128].Low-context benchmarks, by contrast, create ambiguity: the same efficiency or accuracy number can hide wide differences across contexts and user groups [79], making it hard to know what those results mean for any specific deployment. With dedicated infrastructure, standing panels, and shared scenarios, real-world tests can run on a regular schedule, similar to large field surveys of public opin- ion. Over time, sector-specific "digital twins" of users, tasks, and environments can help automate evaluation by comparing new system behavior to baselines grounded in prior human data. 13 5 Limitations and Operational Boundaries Decision-makers need more than leaderboards; they need ways to navigate real deployments. FRAME eval- uates AI-in-use and turns outcomes into systematic knowledge so organizations can move from guessing about AIâs impact to actively managing itâdeciding where to invest, what to scale, and how to protect their operations and communities. At the same time, produc- ing decision-ready evidence introduces its own complex- ities, trade-offs, and practical limits, some of which are outlined below. Balancing human observation with automated scaling FRAME uses automated tools while keeping peo- ple at the center of evaluation so insights can scale across scenarios, systems, and configurations. Tradi- tional in-silico testing is indispensable for probing model behavior and generating abstract knowledge, but on its own it cannot show how AI systems actually work in peopleâs lives. The sandbox still uses in-silico tests, but pairs them with structured human observation to pro- duce deployment-level insights about the strengths and limits of each approach. FRAMEâs digital twins serve as another automated scaling technique: they use observed behavioral traces to calibrate simulations so they reflect actual user entropy rather than purely hypothetical users. LLM-as-judge scoring is used as an automated way to produce descrip- tive baselines, not as a source of ground truth. Its out- comes are compared against human-grounded results to reveal key gaps between automated and dynamic scoring methods. Using proxies without losing the plot Proxies are often used in AI evaluation to stand in for richer real-world concepts that are hard to observe directly, and their value depends on how well they are grounded in context.[15] In FRAMEâs real-world methods, carefully chosen proxies help approximate complex constructs closely enough to make measurement practical while also reducing exposure to harm and protecting sensitive data. For example, severe emotional dependence on AI systems cannot be ethically studied directly, but proxy scenarios can model underlying mechanisms under real- istic conditions and with strong guardrails. When proxies are vague, weakly justified, or never val- idated against the constructs they are meant to represent, they can become misleading shortcuts that distort what âgoodâ performance looks like. For example, âpassing a safety filterâ can serve as a stand-in for âsafe behavior in the wildâ. When disconnected from real-world use, such proxies can shift evaluation toward what is easy to count (such as keyword refusals) rather than what matters for decision-making. Centralizing the resource burden of real-world evaluation Running situated observation and main- taining large-scale, standing participant panels is resource-intensive.FRAMEâs goal is not for every organization to build a permanent real-world AI evalu- ation lab but to function as centralized infrastructure, absorbing the burden of recruiting, verifying, and running panels at scale. This centralization allows many institutions to draw on human-grounded evidence while FRAME panels and protocols bear the operational load. Grounded, transparent scoring constructs FRAMEâs contextual scoring engine organizes AI-in-use around clearly defined ideas that matter in a given scenario, population, or setting. Rubrics are based on observational data so interactions can first be grouped into descriptive categories, such as: patterns of use, non-use, reliance, abandonment, and workarounds. Sys- tem scores then show how often each pattern appears across different groups, tasks, and contexts. Rubrics and scoring rules are documented in public-facing materials so others can review, question, adapt, and reuse them against their own norms and risk thresholds. Iterative formalization of metrics The precise for- mulas, data schemas, and standards for FRAMEâs six metric families are intentionally kept at a high level in this paper. Because real-world environments are highly varied, any formalization requires ongoing validation. FRAME members refine developed methods iteratively and, as metrics are validated, publish them to build a transparent, shared catalog that can serve as a common language for the field. Fixed benchmark datasets vs dynamic traces No- tably, FRAMEâs sandbox does not rely on fixed bench- mark datasets. Instead, it collects new data directly from the sandbox, then folds observational traces into shared community models. This keeps the evidence base tied to how people actually use AI systems over time. Dynamic 14 traces can complement fixed benchmarks by clarifying how model capabilities materializeâor fail toâin specific contexts. Protecting privacy via contextual mimicry Mea- suring real operational friction often raises concerns about access to proprietary systems or private employee data. FRAME does not tap into organizationsâ internal data streams. Instead, it works with sectors to map con- straints and workflows, then recreates those environ- ments inside the sandbox using proxy scenarios and spe- cialized panels (for example, credentialed nurses or civil servants). This approach is designed to protect sensitive data while still capturing domain-specific user entropy and norms, rather than relying solely on general-purpose consumer use in everyday settings. 6 About the Intiative 6.1 Origins and governance FRAME emerged from the AI Metrology Working Group (AI MWG), launched in summer 2024 by Humane Intelli- gence under the leadership of Dr. Rumman Chowdhury. The AI MWG was a dedicated community focused on strengthening the scientific measurement of AI and ex- panding evaluation capacity beyond big tech. Over time, this effort evolved into FRAME, whose members span in- dustry, government, civil society, and academia. FRAME is managed by Civitaas Insights and anchored at Vir- ginia State Universityâs Center for Responsible AI, which serves as host and sponsor rather than a commercial ben- eficiary. This structure is designed to safeguard indepen- dence while providing stable governance, administrative support, and robust conflict-of-interest protections. 6.2 Methodological Foundations FRAME adapts established scientific traditionsâprogram evaluation [97], realist evaluation[80, 65, 73, 62], and implementation science [9, 35]âto the challenges of AI measurement. This orientation allows FRAME to: ⢠treat AI deployment as an intervention in a complex socio-technical system, ⢠move beyond abstract performance scores to un- cover "what works, for whom, in what circum- stancesâ[77, 84, 45, 46], ⢠leverage rigorous quasi-experimental designs in real-world testing [18, 99], and ⢠build up the causal evidence decision-makers need to distinguish genuine utility from merely theoreti- cal capability. 6.3 Engaging with FRAME Organizations and communities can work with FRAME to access empirical evidence grounded in settings like their own and to help resolve the âdecision-makerâs dilemma.â Through paid tiers, 2 partners support the ini- tiative and receive tailored, decision-ready insights tied to their specific use cases and domains: ⢠Sponsorship & targeted trials: Organizations can underwrite sandbox trials to evaluate AI systems in their own use cases (for example, a benefits chatbot or a newsroom tool) under conditions that reflect how people actually adopt, adapt, or abandon sys- tems. ⢠Specialized panels: FRAME collaborates with sponsors to curate specialized participant pan- elsâsuch as educators or defined consumer seg- mentsâorganized as short-term or longitudinal co- horts so results reflect the populations and contexts that matter most. ⢠Licensing & decision-ready evidence: Sponsors can license access to FRAMEâs community models and detailed metrics to compare their own pilots against broader patterns of risk and valueâwithout sharing or exposing their proprietary data and sys- tems. Acknowledgments This white paper was developed as part of the Forum for Real-World AI Measurement and Evaluation (FRAME) at Virginia State Universityâs Center for Responsible AI. The authors would like to thank members of the AI Metrology Working Group who provided input on FRAMEâs structure and focus: Morgan Briggs, Peter Dou- glas, Matt Holmes, Fariza Rashid, Isar Nejadgholi, Afaf TaĂŻk, Carina Westling, and Kyra Wilson. The authors also thank Sundeep Bhandari and Rumman Chowdhury 2 FRAME maintains a strict firewall between sponsors and its sci- entific work. Sponsors can pose questions and fund studies, but all methods, analysis, and publication decisions follow VSUâs policies, not sponsor preferences. 15 for their feedback and suggestions on early drafts of this document. Thanks to the entire Center for Responsible AI team at VSU: Maurice B. Jones, Sylvia Jones, Dennis Donaldson, PhD. (affiliate researcher), and M. Omar Fai- son, PhD (affiliate researcher). AI Use Disclosure Portions of this white paper were developed with the as- sistance of generative AI tools. These tools were used to help format the LaTeX/Overleaf layout, streamline and clarify draft text, suggest alternative phrasings, assist in generating illustrative figures based on author require- ments, and create or refine a subset of the BibTeX entries. All ideas, arguments, structure, and final wording were reviewed, edited, and approved by the authors, who take full responsibility for the content. 16 Glossary of Real-World AI Evaluation Terminology ⢠Abstract knowledge: Information about what a model can do in principle, usually measured in controlled or lab-like settings. ⢠Accessibility: Whether a system is usable across different abilities, languages, and cultural backgrounds, and whether it avoids creating new barriers for specific groups. ⢠Adaptability: How well a system or workflow can adjust to new users, contexts, tasks, or constraints without breaking or requiring extensive rework. ⢠AI deployment: Phase of a project where a system is put into operation and cutover issues are resolved[105]. ⢠AI model capabilities: The tasks, functions, and behaviors an AI model can reliably perform, given specified inputs and conditions. ⢠AI stack: A layered set of technologies, tools, frameworks, and infrastructure for building, deploying, and operating AI systems and applications. ⢠AI system: A machine-based system that, for explicit or implicit objectives, infers, from the input it receives, how to generate outputs such as predictions, content, recommendations, or decisions that can influence phys- ical or virtual environments. Different AI systems vary in their levels of autonomy and adaptiveness after deployment[76]. ⢠Assessment: Action of applying specific documented criteria to a specific software module, package or product for the purpose of determining acceptance or release of the software module, package or product[105]. ⢠Availability: Ensuring timely and reliable access to and use of information[51]. ⢠Benchmarking: A systematic method by which organizations can measure themselves against the best industry practices. ⢠Community models: Digital snapshots of how AI is actually used in a given setting âcapturing user behavior, tasks, failure modes, and higher-order effectsâso organizations can compare systems, test policies, and tailor local evaluations without building full pipelines. ⢠Construct: An abstract, latent concept or theoretical variable, not directly observable, that is defined for scien- tific purposes and measured indirectly through multiple observable indicators or items. ⢠Construct operationalization: Linking a systematized concept to appropriate indicators and scores in a coherent measurement framework[1]. ⢠Construct systematization: The process of clarifying and organizing a concept by specifying its meaning, di- mensions, and relationships to other concepts.[1]. ⢠Context: The parameters in which interrelated factors, purposes, and circumstances may shape individual and collective perceptions, interpretations, and expectations about the functionality and impacts of AI technology, and resulting actions[95]. ⢠Contextual knowledge: Detailed insight tied to a specific setting, including the local conditions, people, prac- tices, and history that give information its meaning and practical relevance in that context. ⢠Decision-makerâs dilemma: The predicament faced by stakeholders outside the technical stackâsuch as CIOs, agency heads, editors, and compliance officersâwho must decide whether and how to deploy AI systems with- out solid evidence about how those systems actually behave in the environments they oversee. ⢠Decision-ready evidence: Systematic indicators derived from observing AI-in-use that help stakeholders out- side the stack judge utility, risk, and operational value in their own context, so they can act on procurement, deployment, governance, and evaluation. ⢠Descriptivist methods: An evidence-based approach to language that describes how people actually use it, rather than telling them how they should use it. ⢠Evaluation: (1) Systematic determination of the extent to which an entity meets its specified criteria; (2) Action that assesses the value of something[105]. 17 ⢠Evaluation scenario: A high-level description of a specific situation or sequence of conditions under which a system is to be tested, including the initial state, inputs, triggering events, and expected behavior or outcomes. ⢠Generative AI: A category of AI that can create new content such as text, images, videos and music[75]. ⢠Ground truth: Value of the target variable for a particular item of labelled input data[103]. ⢠Higher-order effects: Long-term and broad-scale outcomes and consequences that may result from AI use in the real world. ⢠In-silico testing: Testing or experimentation carried out entirely on a computer, using computational models and simulations. ⢠Interoperability: Degree to which two or more systems, products or components can exchange information and use the information that has been exchanged[105]. ⢠Measurement: (1) Quantitative measurement is the act or process of assigning a number or category to an entity to describe an attribute of that entity[105]. (2) Qualitative measurement is based on descriptive data such as through observations, interviews, focus groups, or open-ended text fields in surveys. ⢠Model stochasticity: The inherent randomness in generative model outputs that makes the same or similar prompt produce different answers on different runs.. ⢠Predictive AI: AI systems whose primary function is to infer from input data and produce predictions or fore- casts about future states, behaviors, or outcomes, typically to inform decisions or actions [76]. ⢠Privacy: Freedom from intrusion into the private life or affairs of an individual[104]. ⢠Proxy scenario: A test scenario that does not directly reproduce the real-world context of interest but is designed to stand in for it, using more tractable or safer conditions while preserving key features believed to be relevant for evaluating system behavior or risk. ⢠Quality: The totality of features and characteristics of a product or service that bear on its ability to satisfy stated or implied needs [74] ⢠Real-world AI evaluation: The process of measuring what actually happens when people use AI systems in everyday workflows and contexts, at a scale that allows patterns to emerge across different users and settings. ⢠Reality Gap: The difference between a modelâs performance in optimized test conditions and the outcomes it produces with real people in real contexts. ⢠Resilience: The ability of an information system to continue to: (i) operate under adverse conditions or stress, even if in a degraded or debilitated state, while maintaining essential operational capabilities; and (i) recover to an effective operational posture in a time frame consistent with mission needs[50]. ⢠Robustness: Ability of a system to maintain its level of performance under a variety of circumstances[104]. ⢠Situated observation: The practice of observing activity directly in its natural setting and attending to how behavior unfolds over time within specific physical, social, and cultural environments. ⢠Systematic knowledge: Organized, cumulative knowledge that a community maintains and updates so that its concepts, measures, and explanations can be applied across many contexts, not just a single local case. ⢠Testing, Evaluation, Validation, and Verification (TEVV): A framework for assessing and incorporating meth- ods and metrics to determine that a technology or system satisfactorily meets its design specifications and requirements, and that it is sufficient for its intended use [72]. ⢠User entropy: The inherent heterogeneity in how people use AI in context, including how they express needs, interpret and adapt outputs, and embed AI into their own goals and constraints. References [1] Robert Adcock and David Collier. âMeasurement Validity: A Shared Standard for Qualitative and Quan- titative Researchâ. In: American Political Science Review 95.3 (2001), p. 529â546. doi: 10 . 1017 / S0003055401003100. 18 [2] Ajay K. Agrawal, Joshua S. Gans, and Avi Goldfarb. AI Adoption and System-Wide Change. Working Paper. May 2021. doi: 10.3386/w28811. url: https://w.nber.org/papers/w28811. [3] Kofi Arhin et al. Ground-Truth, Whose Truth? â Examining the Challenges with Annotating Toxic Text Datasets. 2021. arXiv: 2112.03529 [cs.CL]. url: https://arxiv.org/abs/2112.03529. [4] Zahra Ashktorab et al. Emerging Reliance Behaviors in Human-AI Content Grounded Data Generation: The Role of Cognitive Forcing Functions and Hallucinations. en. Sept. 2024. url: https://arxiv.org/ abs/2409.08937v2. [5] Mohammad Atari et al. âWhich Humans?â eng. In: (). url: https://psyarxiv.com/5b26t. [6] Berk Atil et al. Non-Determinism of "Deterministic" LLM Settings. arXiv:2408.04667 [cs] version: 5. Apr. 2025. doi: 10.48550/arXiv.2408.04667. url: http://arxiv.org/abs/2408.04667. [7] Dominic Balog-Way, Katherine McComas, and John Besley. âThe Evolving Field of Risk Communicationâ. In: Risk Analysis 40.Suppl 1 (Nov. 2020), p. 2240â2262. issn: 0272-4332. doi: 10.1111/risa.13615. url: https://pmc.ncbi.nlm.nih.gov/articles/PMC7756860/. [8] Solon Barocas, Moritz Hardt, and Arvind Narayanan. âFairness and Machine Learning: Limitations and Op- portunitiesâ. In: (2023). [9] Simon Barry and Chris Davies. âCreating system-wide change in medicine: The role of implementation sci- ence in achieving scale and adoptionâ. In: Future Healthcare Journal 12.3 (July 2025), p. 100452. issn: 2514- 6645. doi: 10.1016/j.fhj.2025.100452. url: https://pmc.ncbi.nlm.nih.gov/ articles/PMC12395526/. [10] Andrew M. Bean et al. âMeasuring what Matters: Construct Validity in Large Language Model Benchmarksâ. en. In: Oct. 2025. url: https://openreview.net/forum?id=mdA5lVvNcU. [11] Glen Berman, Nitesh Goyal, and Michael Madaio. âA Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluationsâ. In: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. CHI â24. New York, NY, USA: Association for Computing Machinery, May 2024, p. 1â24. isbn: 979-8-4007-0330-0. doi: 10.1145/3613904.3642398. url: https://dl. acm.org/doi/10.1145/3613904.3642398. [12] Douglas Biber. âRepresentativeness in Corpus Designâ. In: Literary and Linguistic Computing 8.4 (1993), p. 243â257. doi: 10.1093/llc/8.4.243. [13] Ann Bostrom. âVaccine Risk Communication: Lessons from Risk Perception, Decision Making and Environ- mental Risk Communication Researchâ. In: Risk 8 (1997), p. 173. [14] Margarita Boyarskaya, Alexandra Olteanu, and Kate Crawford. Overcoming Failures of Imagination in AI Infused System Development and Deployment. arXiv:2011.13416 [cs]. Dec. 2020. doi: 10.48550/arXiv. 2011.13416. url: http://arxiv.org/abs/2011.13416. [15] Zana Buçinca et al. âProxy tasks and subjective measures can be misleading in evaluating explainable AI sys- temsâ. en. In: Proceedings of the 25th International Conference on Intelligent User Interfaces. Cagliari Italy: ACM, Mar. 2020, p. 454â464. isbn: 978-1-4503-7118-6. doi: 10.1145/3377325.3377498. url: https: //dl.acm.org/doi/10.1145/3377325.3377498. [16] Neville Calleja et al. âA Public Health Research Agenda for Managing Infodemics: Methods and Results of the First WHO Infodemiology Conferenceâ. EN. In: JMIR Infodemiology 1.1 (Sept. 2021), e30979. doi: 10. 2196/30979. url: https://infodemiology.jmir.org/2021/1/e30979. 19 [17] D. Campbell and J. Stanley. âEXPERIMENTAL AND QUASI-EXPERIMENT Al DESIGNS FOR RESEARCHâ. In: 1959. url: https : / / w . semanticscholar . org / paper / EXPERIMENTAL - AND - QUASI - EXPERIMENT - Al - DESIGNS - FOR - Campbell - Stanley / 7f7d4966313c36d903de5b30baa1a5f792007c9. [18] Donald T. Campbell. âMethods for the Experimenting Societyâ. en. In: Evaluation Practice 12.3 (Oct. 1991), p. 223â260. issn: 0886-1633. doi: 10.1177/109821409101200304. url: https://journals. sagepub.com/doi/10.1177/109821409101200304. [19] Valerio Capraro et al. âThe impact of generative artificial intelligence on socioeconomic inequalities and policy makingâ. In: PNAS Nexus 3.6 (June 2024), pgae191. issn: 2752-6542. doi: 10.1093/pnasnexus/ pgae191. url: https://pmc.ncbi.nlm.nih.gov/articles/PMC11165650/. [20] Stephen Casper et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv:2307.15217 [cs]. Sept. 2023. doi: 10.48550/arXiv.2307.15217. url: http: //arxiv.org/abs/2307.15217. [21] Aditya Challapally, Ramesh Raskar, and Pradyumna Chari. The GenAI Divide: State of AI in Business 2025. Research Report. Preliminary findings from AI implementation research from Project NANDA, research pe- riod JanuaryâJune 2025. Project NANDA, Massachusetts Institute of Technology and MLQ.ai, 2025. url: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_ 2025_Report.pdf. [22] Cong Chen et al. âA comprehensive review of LLM-based content moderation: advancements, challenges, and future directionsâ. In: Knowledge-Based Systems 330 (2025), p. 114689. issn: 0950-7051. doi: https: //doi.org/10.1016/j.knosys.2025.114689. url: https://w.sciencedirect. com/science/article/pii/S0950705125017289. [23] Ke Chen et al. Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems. arXiv:2505.13546 [cs]. May 2025. doi: 10.48550/arXiv.2505.13546. url: http:// arxiv.org/abs/2505.13546. [24] Alexandra Chouldechova et al. A Shared Standard for Valid Measurement of Generative AI Systemsâ Capabil- ities, Risks, and Impacts. arXiv:2412.01934 [cs]. Dec. 2024. doi: 10.48550/arXiv.2412.01934. url: http://arxiv.org/abs/2412.01934. [25] Andrea Christodoulou and Petros Roussos. âPhone in the Room, Mind on the Roamâ: Investigating the Impact of Mobile Phone Presence on Distractionâ. In: European Journal of Investigation in Health, Psychology and Education 15.5 (May 2025), p. 74. issn: 2174-8144. doi: 10.3390/ejihpe15050074. url: https: //pmc.ncbi.nlm.nih.gov/articles/PMC12110250/. [26] Katherine M. Collins et al. Evaluating Language Models for Mathematics through Interactions. arXiv:2306.01694 [cs]. Nov. 2023. doi: 10.48550/arXiv.2306.01694. url: http://arxiv.org/abs/2306. 01694. [27] Vincent Conitzer et al. Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback. arXiv:2404.10271 [cs]. June 2024. doi: 10.48550/arXiv.2404.10271. url: http://arxiv. org/abs/2404.10271. [28] Jonathan C. Corbin et al. âHow reasoning, judgment, and decision making are colored by gist-based intuition: A fuzzy-trace theory approachâ. In: Journal of Applied Research in Memory and Cognition 4.4 (Dec. 2015), p. 344â355. issn: 2211-3681. doi: 10.1016/j.jarmac.2015.09.001. url: https://w. sciencedirect.com/science/article/pii/S2211368115000728. 20 [29] Rebecca Crootof, Margot Kaminski, and W. Price I. âHumans in the Loopâ. In: Vanderbilt Law Review 76.2 (Mar. 2023), p. 429. url: https://scholarship.law.vanderbilt.edu/vlr/vol76/ iss2/2. [30] Alexander DâAmour et al. Underspecification Presents Challenges for Credibility in Modern Machine Learning. arXiv:2011.03395 [cs]. Nov. 2020. doi: 10.48550/arXiv.2011.03395. url: http://arxiv. org/abs/2011.03395. [31] Adam Dahlgren LindstrĂśm et al. âHelpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedbackâ. en. In: Ethics and Information Technology 27.2 (June 2025), p. 28. issn: 1388-1957, 1572-8439. doi: 10.1007/s10676-025-09837-2. url: https: //link.springer.com/10.1007/s10676-025-09837-2. [32] Wesley Hanwen Deng et al. âUnderstanding Practices, Challenges, and Opportunities for User-Engaged Al- gorithm Auditing in Industry Practiceâ. In: Proceedings of the 2023 CHI Conference on Human Factors in Com- puting Systems. arXiv:2210.03709 [cs]. Apr. 2023, p. 1â18. doi: 10.1145/3544548.3581026. url: http://arxiv.org/abs/2210.03709. [33] Ăilish Duke and Christian Montag. âSmartphone addiction, daily interruptions and self-reported productiv- ityâ. In: Addictive Behaviors Reports 6 (July 2017), p. 90â95. issn: 2352-8532. doi: 10.1016/j.abrep. 2017.07.002. url: https://pmc.ncbi.nlm.nih.gov/articles/PMC5800562/. [34] Bridget Dwyer et al. âMindbench.ai: an actionable platform to evaluate the profile and performance of large language models in a mental healthcare contextâ. In: NPP - Digital Psychiatry and Neuroscience 3 (Nov. 2025), p. 28. issn: 2948-1570. doi: 10.1038/s44277-025-00049-6. url: https://pmc.ncbi.nlm. nih.gov/articles/PMC12624894/. [35] Martin P. Eccles and Brian S. Mittman. âWelcome to Implementation Scienceâ. en. In: Implementation Science 1.1 (Feb. 2006), p. 1. issn: 1748-5908. doi: 10.1186/1748-5908-1-1. url: https://doi.org/ 10.1186/1748-5908-1-1. [36] Sarah M. Edelson and Valerie F. Reyna. âWho Makes the Decision, How, and Why: A Fuzzy-Trace Theory Approachâ. en. In: Medical Decision Making 44.6 (Aug. 2024), p. 614â616. issn: 0272-989X, 1552-681X. doi: 10.1177/0272989X241263818. url: https://journals.sagepub.com/doi/10. 1177/0272989X241263818. [37] Rabie Adel El Arab et al. âBridging the Gap: From AI Success in Clinical Trials to Real-World Healthcare ImplementationâA Narrative Reviewâ. en. In: Healthcare 13.7 (Mar. 2025), p. 701. issn: 2227-9032. doi: 10. 3390/healthcare13070701. url: https://w.mdpi.com/2227-9032/13/7/701. [38] Lily A. Elefteriadou and Transportation Research Board. Highway Capacity Manual 6th Edition: A Guide for Multimodal Mobility Analysis. Pages: 24798. Washington, D.C.: National Academies Press, Nov. 2016. isbn: 978-0-309-46019-4. doi: 10.17226/24798. url: https://w.nationalacademies.org/ publications/24798. [39] Maria Eriksson et al. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evalua- tion. arXiv:2502.06559 [cs]. May 2025. doi:10.48550/arXiv.2502.06559. url: http://arxiv. org/abs/2502.06559. [40] Henry Farrell et al. âLarge AI models are cultural and social technologiesâ. en. In: Science 387.6739 (Mar. 2025), p. 1153â1156. issn: 0036-8075, 1095-9203. doi: 10.1126/science.adt9819. url: https: //w.science.org/doi/10.1126/science.adt9819. 21 [41] Baruch Fischhoff, Ann Bostrom, and Marilyn Jacobs Quadrel. âRisk Perception and Communicationâ. en. In: Annual Review of Public Health 14.1 (May 1993), p. 183â203. issn: 0163-7525, 1545-2093. doi: 10.1146/ annurev.pu.14.050193.001151. url: https://w.annualreviews.org/doi/10. 1146/annurev.pu.14.050193.001151. [42] Michèle A. Flournoy, Avril Haines, and Gabrielle Chefitz. Building Trust through Testing: Adapting DODâs Test & Evaluation, Validation & Verification (TEVV) Enterprise for Machine Learning Systems. Tech. rep. Washing- ton, DC: Center for Security and Emerging Technology, 2020. url: https://cset.georgetown. edu/publication/building-trust-through-testing/. [43] Andrew Gelman and Christian Hennig. âBeyond Subjective and Objective in Statisticsâ. In: Journal of the Royal Statistical Society: Series A (Statistics in Society) 180.4 (2017), p. 967â1033. doi: 10.1111/rssa. 12276. [44] Ellen P. Goodman, National Telecommunications, and Information Administration. NTIA Artificial Intelli- gence Accountability Policy Report. Policy Report. Artificial Intelligence Accountability Policy Report, March 2024. U.S. Department of Commerce, National Telecommunications and Information Administration, Mar. 2024. url: https://w.ntia.gov/sites/default/files/publications/ntia- ai-report-final.pdf. [45] Joanne Greenhalgh and Ana Manzano. âUnderstanding âcontextâ in realist evaluation and syn- thesisâ. In: International Journal of Social Research Methodology 25.5 (Sept. 2022). _eprint: https://doi.org/10.1080/13645579.2021.1918484, p. 583â595. issn: 1364-5579. doi: 10.1080/13645579. 2021.1918484. url: https://doi.org/10.1080/13645579.2021.1918484. [46] Benjamin Harris, Lyn Alderman, and Jessica Staheli. âThe forgotten contexts of evaluationâ. EN. In: Eval- uation 31.2 (Apr. 2025), p. 240â261. issn: 1356-3890. doi: 10.1177/13563890241312910. url: https://doi.org/10.1177/13563890241312910. [47] Lama Ibrahim et al. âTowards Interactive Evaluations for Interaction Harms in Human-AI Systemsâ. In: (2025). White paper. url: https://knightcolumbia.org/content/towards-interactive- evaluations-for-interaction-harms-in-human-ai-systems. [48] Eaman Jahani et al. As Generative Models Improve, People Adapt Their Prompts. arXiv:2407.14333 [cs] version: 2. Aug. 2024. doi: 10.48550/arXiv.2407.14333. url: http://arxiv.org/abs/2407. 14333. [49] Zehua Jiang et al. âBeyond multiple-choice questions: rethinking evaluation frameworks for large language models for clinical medicineâ. In: Intelligent Medicine (Jan. 2026). issn: 2667-1026. doi: 10.1016/j. imed.2026.01.001. url: https://w.sciencedirect.com/science/article/ pii/S266710262600001X. [50] Joint Task Force Transformation Initiative. Managing Information Security Risk: Organization, Mission, and Information System View. NIST Special Publication 800-39. Gaithersburg, MD: National Institute of Standards and Technology, 2011. doi: 10.6028/NIST.SP.800-39. url: https://doi.org/10.6028/ NIST.SP.800-39. [51] Joint Task Force Transformation Initiative. Risk Management Framework for Information Systems and Orga- nizations: A System Life Cycle Approach for Security and Privacy. NIST Special Publication 800-37 Revision 2. Gaithersburg, MD: National Institute of Standards and Technology, 2018. doi: 10.6028/NIST.SP. 800-37r2. url: https://doi.org/10.6028/NIST.SP.800-37r2. 22 [52] Maurits Clemens Kaptein, Robin van Emden, and Davide Iannuzzi. âUncovering noisy social signals: Using optimization methods from experimental physics to study social phenomenaâ. In: PLOS ONE 12 (2017). url: https://api.semanticscholar.org/CorpusID:215779034. [53] Barry Kehoe. The Top 100 Ways People Are Using AI in 2025 (and How Theyâve Changed Since 2024). en. url: https://w.qualtrics.com/articles/customer-experience/the-top-100- ways-people-are-using-ai-2025/. [54] Aman Khullar et al. âNurturing Capabilities: Unpacking the Gap in Human-Centered Evaluations of AI- Based Systemsâ. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 101. New York, NY, USA: Association for Computing Machinery, Apr. 2025, p. 1â18. isbn: 979-8-4007-1394-1. url: https://dl.acm.org/doi/10.1145/3706598.3713278. [55] Kelly Klaver et al. Data Integration, Sharing, and Management for Transportation Planning and Traffic Oper- ations. Pages: 28690. Washington, D.C.: National Academies Press, June 2025. isbn: 978-0-309-73226-0. doi: 10.17226/28690. url: https://w.nationalacademies.org/publications/ 28690. [56] K. Larsen et al. Validity in Design Science. arXiv:2503.09466 [cs]. Mar. 2025. doi: 10.48550/arXiv. 2503.09466. url: http://arxiv.org/abs/2503.09466. [57] Sarah Lebovitz et al. âIs AI Ground Truth Really True? The Dangers of Training and Evaluating AI Tools Based on Expertsâ Know-Whatâ. en. In: MIS Quarterly 45.3 (Sept. 2021), p. 1501â1526. issn: 02767783, 21629730. doi: 10.25300/MISQ/2021/16564. url: https://misq.org/is-ai-ground-truth- really-true-the-dangers-of-training-and-evaluating-ai-tools-based- on-experts-know-what.html. [58] Chance Jiajie Li et al. HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning. Version Number: 3. 2025. doi: 10.48550/ARXIV.2510.15144. url: https://arxiv.org/ abs/2510.15144. [59] Q. Vera Liao and Jennifer Wortman Vaughan. AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap. arXiv:2306.01941 [cs]. Aug. 2023. doi: 10.48550/arXiv.2306.01941. url: http: //arxiv.org/abs/2306.01941. [60] Q. Vera Liao and Ziang Xiao. Rethinking Model Evaluation as Narrowing the Socio-Technical Gap. arXiv:2306.03100 [cs]. Jan. 2025. doi: 10.48550/arXiv.2306.03100. url: http://arxiv. org/abs/2306.03100. [61] Kai Liu. âAnalysis of the Conflict between Car Commuterâs Route Choice Habitual Behavior and Traffic Infor- mation Search Behaviorâ. In: Sensors (Basel, Switzerland) 22.12 (June 2022), p. 4382. issn: 1424-8220. doi: 10. 3390/s22124382. url: https://pmc.ncbi.nlm.nih.gov/articles/PMC9231029/. [62] Sarah Louart et al. âRealist Evaluationâ. en-ca. In: (). Book Title: Policy Evaluation: Methods and Approaches. url: https : / / scienceetbiencommun . pressbooks . pub / pubpolevaluation / chapter/realistic-evaluation/. [63] Brooke N. Macnamara et al. âDoes using artificial intelligence assistance accelerate skill decay and hinder skill development without performersâ awareness?â In: Cognitive Research: Principles and Implications 9 (July 2024), p. 46. issn: 2365-7464. doi: 10.1186/s41235-024-00572-8. url: https://pmc.ncbi. nlm.nih.gov/articles/PMC11239631/. 23 [64] Takuya Maeda and Anabel Quan-Haase. âWhen Human-AI Interactions Become Parasocial: Agency and An- thropomorphism in Affective Designâ. en. In: The 2024 ACM Conference on Fairness, Accountability, and Trans- parency. Rio de Janeiro Brazil: ACM, June 2024, p. 1068â1077. isbn: 979-8-4007-0450-5. doi: 10.1145/ 3630106.3658956. url: https://dl.acm.org/doi/10.1145/3630106.3658956. [65] Ana Manzano and Emma Williams, eds. Realist evaluation: principles and practice. eng. Abingdon New York: Routledge, 2025. [66] Nestor Maslej et al. Artificial Intelligence Index Report 2025. 2025. arXiv: 2504.07139 [cs.AI]. [67] J. Nathan Matias. âHumans and algorithms work together â so study them togetherâ. en. In: Nature 617.7960 (May 2023). Bandiera_abtest: a Cg_type: Comment Subject_term: Human behaviour, Information technology, Computer science, p. 248â251. doi: 10.1038/d41586-023-01521-z. url: https://w. nature.com/articles/d41586-023-01521-z. [68] Kelsey Lynn McAlister, Lee Gonzales, and Jennifer Huberty. âRethinking AI Workflows: Guidelines for Scien- tific Evaluation in Digital Health Companiesâ. en. In: JMIR AI 4 (Dec. 2025), e71798âe71798. issn: 2817-1705. doi: 10.2196/71798. url: https://ai.jmir.org/2025/1/e71798. [69] Kiana Jafari Meimandi et al. The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Pro- ductivity Claims. arXiv:2506.02064 [cs] version: 1. June 2025. doi: 10.48550/arXiv.2506.02064. url: http://arxiv.org/abs/2506.02064. [70] El-Mahdi El-Mhamdi et al. âOn the Impossible Safety of Large AI Modelsâ. In: (2022). Version Number: 2. doi: 10.48550/ARXIV.2209.15259. url: https://arxiv.org/abs/2209.15259. [71] Monika Nair et al. âCritical activities for successful implementation and adoption of AI in healthcare: towards a process framework for healthcare organizationsâ. In: Frontiers in Digital Health 7 (May 2025), p. 1550459. issn: 2673-253X. doi: 10.3389/fdgth.2025.1550459. url: https://pmc.ncbi.nlm.nih. gov/articles/PMC12122488/. [72] National Security Commission on Artificial Intelligence. Final Report. Tech. rep. Released March 2021. Washington, DC: National Security Commission on Artificial Intelligence, Mar. 2021. url: https:// reports.nscai.gov/final-report. [73] Steffen Bohni Nielsen, Sebastian Lemire, and Stinne Tangsig. âUnpacking context in realist evaluations: Find- ings from a comprehensive reviewâ. en. In: Evaluation 28.1 (Jan. 2022), p. 91â112. issn: 1356-3890, 1461-7153. doi: 10.1177/13563890211053032. url: https://journals.sagepub.com/doi/10. 1177/13563890211053032. [74] Organisation for Economic Co-operation and Development. OECD Glossary of Statistical Terms. Paris: OECD Publishing, 2008. isbn: 9789264025561. [75] Organisation for Economic Co-operation and Development. Initial Policy Considerations for Generative Arti- ficial Intelligence. OECD, 2023. url: https://doi.org/10.1787/fae2d1e6-en. [76] Organisation for Economic Co-operation and Development. Explanatory Memorandum on the Updated OECD Definition of an AI System. OECD, 2024. url: https://doi.org/10.1787/623da898-en. [77] Rebecca Palm and Alexander Hochmuth. âWhat works, for whom and under what circumstances? Us- ing realist methodology to evaluate complex interventions in nursing: A scoping reviewâ. en. In: Inter- national Journal of Nursing Studies 109 (Sept. 2020), p. 103601. issn: 00207489. doi: 10 . 1016 / j . ijnurstu.2020.103601. url: https://linkinghub.elsevier.com/retrieve/ pii/S0020748920300869. 24 [78] Srikant Panda, Amit Agarwal, and Hitesh Laxmichand Patel. âAccessEval: Benchmarking Disability Bias in Large Language Modelsâ. In: (Nov. 2025). Ed. by Christos Christodoulopoulos et al., p. 32504â32530. doi: 10.18653/v1/2025.emnlp-main.1653. url: https://aclanthology.org/2025. emnlp-main.1653/. [79] Hyerim Park et al. â"We Are Visual Thinkers, Not Verbal Thinkers!": A Thematic Analysis of How Pro- fessional Designers Use Generative AI Image Generation Toolsâ. en. In: Nordic Conference on Human- Computer Interaction. Uppsala Sweden: ACM, Oct. 2024, p. 1â14. isbn: 979-8-4007-0966-1. doi: 10.1145/ 3679318.3685370. url: https://dl.acm.org/doi/10.1145/3679318.3685370. [80] Ray Pawson and Nick Tilley. âAn introduction to scientific realist evaluationâ. In: Evaluation for the 21st century: A handbook. Thousand Oaks, CA, US: Sage Publications, Inc, 1997, p. 405â418. doi: 10.4135/ 9781483348896.n29. [81] Kathleen H. Pine and Max Liboiron. âThe Politics of Measurement and Actionâ. In: Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems. CHI â15. New York, NY, USA: Association for Computing Machinery, Apr. 2015, p. 3147â3156. isbn: 978-1-4503-3145-6. doi: 10.1145/2702123. 2702298. url: https://dl.acm.org/doi/10.1145/2702123.2702298. [82] Jocelyne Piret and Guy Boivin. âPandemics Throughout Historyâ. English. In: Frontiers in Microbiology 11 (Jan. 2021). issn: 1664-302X. doi: 10 . 3389 / fmicb . 2020 . 631736. url: https : / / w . frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2020. 631736/full. [83] Barbara Plank. âThe âProblemâ of Human Label Variation: On Ground Truth in Data, Modeling and Eval- uationâ. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Ed. by Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang. Abu Dhabi, United Arab Emirates: Association for Com- putational Linguistics, Dec. 2022, p. 10671â10682. doi: 10.18653/v1/2022.emnlp-main.731. url: https://aclanthology.org/2022.emnlp-main.731/. [84] Melanie Punton et al. Reality Bites: Making Realist Evaluation Useful in the Real World. en. report. The Institute of Development Studies and Partner Organisations, Mar. 2020. url: https://opendocs.ids.ac. uk/articles/report/Reality_Bites_Making_Realist_Evaluation_Useful_ in_the_Real_World/26432020/1. [85] Alberto Purpura et al. Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models. arXiv:2503.01742 [cs]. Mar. 2025. doi:10.48550/arXiv.2503.01742. url:http: //arxiv.org/abs/2503.01742. [86] Ezra Raez. INRIX Data Network and INRIX XD Traffic. en. Jan. 2026. url: https://w.inrix.com/ ra/datanetworkandxdtraffic/. [87] Sam Ransbotham et al. âThe Emerging Agentic Enterprise: How Leaders Must Navigate a New Age of AIâ. en-US. In: MIT Sloan Management Review (Nov. 2025). url: https://sloanreview.mit.edu/ projects/the-emerging-agentic-enterprise-how-leaders-must-navigate- a-new-age-of-ai/. [88] Anka Reuel et al. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices. arXiv:2411.12990 [cs]. Nov. 2024. doi: 10.48550/arXiv.2411.12990. url: http://arxiv. org/abs/2411.12990. [89] Valerie F. Reyna. âA theory of medical decision making and health: fuzzy trace theoryâ. eng. In: Medical Decision Making: An International Journal of the Society for Medical Decision Making 28.6 (2008), p. 850â865. issn: 0272-989X. doi: 10.1177/0272989X08327066. 25 [90] Valerie F. Reyna. âA scientific theory of gist communication and misinformation resistance, with implications for health, education, and policyâ. en. In: Proceedings of the National Academy of Sciences 118.15 (Apr. 2021), e1912441117. issn: 0027-8424, 1091-6490. doi:10.1073/pnas.1912441117. url:https://pnas. org/doi/full/10.1073/pnas.1912441117. [91] Horst W. J. Rittel and Melvin M. Webber. âDilemmas in a General Theory of Planningâ. In: Policy Sciences 4.2 (1973), p. 155â169. issn: 0032-2687. url: https://w.jstor.org/stable/4531523. [92] Olawale Salaudeen et al. Measurement to Meaning: A Validity-Centered Framework for AI Evaluation. arXiv:2505.10573 [cs]. June 2025. doi: 10.48550/arXiv.2505.10573. url: http://arxiv. org/abs/2505.10573. [93] Rodolfo Saracci. âIntroducing the history of epidemiologyâ. In: Teaching Epidemiology: A guide for teachers in epidemiology, public health and clinical medicine. Ed. by Jørn Olsen et al. Oxford University Press, Mar. 2015, p. 26. isbn: 978-0-19-968500-4. doi: 10.1093/acprof:oso/9780199685004.003.0001. url: https://doi.org/10.1093/acprof:oso/9780199685004.003.0001. [94] Ingo Schulz-Schaeffer. âWhy generative AI is different from designed technology regarding task-relatedness, user interaction, and agencyâ. EN. In: Big Data & Society 12.3 (Sept. 2025), p. 20539517251367452. issn: 2053-9517. doi: 10 . 1177 / 20539517251367452. url: https : / / doi . org / 10 . 1177 / 20539517251367452. [95] Reva Schwartz et al. The Assessing Risks and Impacts of AI (ARIA) Program Evaluation Design Document. Tech. rep. National Institute of Standards and Technology, Gaithersburg, MD, 2024. [96] Reva Schwartz et al. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AIâs Real World Effects. arXiv:2505.18893 [cs]. May 2025. doi: 10.48550/arXiv.2505.18893. url: http:// arxiv.org/abs/2505.18893. [97] Michael Scriven. Evaluation thesaurus, 4th ed. Evaluation thesaurus, 4th ed. Pages: xiii, 391. Thousand Oaks, CA, US: Sage Publications, Inc, 1991. [98] Marc Selgas-Cors. âSociotechnical Transformation: A Systematic Review on the Impact of Artificial Intel- ligence on Society and Organizationsâ. In: FinTech and Sustainable Innovation (2025). doi: 10.47852/ bonviewFSI52026076. url: https://ojs.bonviewpress.com/index.php/FSI/ article/view/6076. [99] William R. Shadish, Thomas D. Cook, and Donald T. Campbell. Experimental and quasi-experimental designs for generalized causal inference. Experimental and quasi-experimental designs for generalized causal infer- ence. Pages: xxi, 623. Boston, MA, US: Houghton, Mifflin and Company, 2002. isbn: 978-0-395-61556-0. [100] Sonali Uttam Singh and Akbar Siami Namin. âA survey on chatbots and large language models: Testing and evaluation techniquesâ. en. In: Natural Language Processing Journal 10 (Mar. 2025), p. 100128. issn: 29497191. doi: 10.1016/j.nlp.2025.100128. url: https://linkinghub.elsevier.com/ retrieve/pii/S2949719125000044. [101] Jeanette Skowronek, Andreas Seifert, and Sven Lindberg. âThe mere presence of a smartphone reduces basal attentional performanceâ. en. In: Scientific Reports 13.1 (June 2023), p. 9363. issn: 2045-2322. doi: 10.1038/ s41598-023-36256-4. url: https://w.nature.com/articles/s41598-023- 36256-4. [102] Molly G. Smith, Thomas N. Bradbury, and Benjamin R. Karney. âCan Generative AI Chatbots Emulate Human Connection? A Relationship Science Perspectiveâ. EN. In: Perspectives on Psychological Science 20.6 (Nov. 2025), p. 1081â1099. issn: 1745-6916. doi: 10.1177/17456916251351306. url: https://doi. org/10.1177/17456916251351306. 26 [103] International Organization for Standardization and International Electrotechnical Commission. Information technology â Artificial intelligence â Artificial intelligence concepts and terminology. First edition. Geneva, Switzerland, 2022. [104] International Organization for Standardization and International Electrotechnical Commission. Trustworthi- ness â Vocabulary. Technical Specification. Geneva, Switzerland, 2022. [105] International Organization for Standardization (ISO) et al. ISO/IEC/IEEE 24765:2017, Systems and Software Engineering â Vocabulary. Geneva, Switzerland, 2017. url: https://w.iso.org/standard/ 71952.html. [106] Vallijah Subasri et al. âDetecting and Remediating Harmful Data Shifts for the Responsible Deployment of Clinical AI Modelsâ. en. In: JAMA Network Open 8.6 (June 2025), e2513685. issn: 2574-3805. doi: 10. 1001/jamanetworkopen.2025.13685. url: https://jamanetwork.com/journals/ jamanetworkopen/fullarticle/2834882. [107] Mirac Suzgun et al. âLanguage models cannot reliably distinguish belief from knowledge and factâ. en. In: Na- ture Machine Intelligence (Nov. 2025), p. 1â11. issn: 2522-5839. doi: 10.1038/s42256-025-01113- 8. url: https://w.nature.com/articles/s42256-025-01113-8. [108] Alex Tamkin et al. Clio: Privacy-Preserving Insights into Real-World AI Use. arXiv:2412.13678 [cs] version: 1. Dec. 2024. doi: 10.48550/arXiv.2412.13678. url: http://arxiv.org/abs/2412. 13678. [109] Hongjie Tang, Mengxue Ou, and Han Zheng. âBetween promise and peril: Usersâ riskâbenefit trade-offs in their generative AI usageâ. EN. In: Big Data & Society 12.4 (Dec. 2025), p. 20539517251410046. issn: 2053-9517. doi: 10 . 1177 / 20539517251410046. url: https : / / doi . org / 10 . 1177 / 20539517251410046. [110] Panuthep Tasawong et al. âShortcut Learning in Safety: The Impact of Keyword Bias in Safeguardsâ. In: Pro- ceedings of the The First Workshop on LLM Security (LLMSEC). Ed. by Leon Derczynski, Jekaterina Novikova, and Muhao Chen. Vienna, Austria: Association for Computational Linguistics, Aug. 2025, p. 189â197. isbn: 979-8-89176-279-4. url: https://aclanthology.org/2025.llmsec-1.14/. [111] The State of AI: Global Survey 2025 | McKinsey. url: https : / / w . mckinsey . com / capabilities/quantumblack/our-insights/the-state-of-ai. [112] Rachel L. Thomas and David Uminsky. âReliance on metrics is a fundamental challenge for AIâ. en. In: Patterns 3.5 (May 2022), p. 100476. issn: 26663899. doi: 10.1016/j.patter.2022.100476. url: https: //linkinghub.elsevier.com/retrieve/pii/S2666389922000563. [113] Jana Uher. âRating scales institutionalise a network of logical errors and conceptual problems in research practices: A rigorous analysis showing ways to tackle psychologyâs crisesâ. English. In: Frontiers in Psychology 13 (Dec. 2022). issn: 1664-1078. doi: 10.3389/fpsyg.2022.1009893. url: https://w. frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2022. 1009893/full. [114] Ana Valenzuela et al. âHow Artificial Intelligence Constrains the Human Experienceâ. en. In: Journal of the Association for Consumer Research 9.3 (July 2024), p. 241â256. issn: 2378-1815, 2378-1823. doi: 10.1086/ 730709. url: https://w.journals.uchicago.edu/doi/10.1086/730709. [115] Gil Verbeke. âOn the Role of Ecological Validity in Language and Speech Researchâ. In: Taalkunde nu. Ed. by Joost Buysschaert and AndrĂŠ Lefèvre. Gent: Skribis, 2024, p. 69â95. 27 [116] Hanna Wallach et al. Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge. arXiv:2502.00561 [cs]. June 2025. doi: 10.48550/arXiv.2502.00561. url: http://arxiv. org/abs/2502.00561. [117] Jinghui Wang and Hesham Rakha. âEmpirical Study of Effect of Dynamic Travel Time Information on Driver Route Choice Behaviorâ. en. In: Sensors 20.11 (June 2020), p. 3257. issn: 1424-8220. doi: 10.3390/ s20113257. url: https://w.mdpi.com/1424-8220/20/11/3257. [118] Kaixuan Wang et al. Critical Challenges in Content Moderation for People Who Use Drugs (PWUD): Insights into Online Harm Reduction Practices from Moderators. arXiv:2508.02868 [cs] version: 2. Jan. 2026. doi: 10. 48550/arXiv.2508.02868. url: http://arxiv.org/abs/2508.02868. [119] Skyler Wang, Ned Cooper, and Margaret Eby. âFrom human-centered to social-centered artificial intelli- gence: Assessing ChatGPTâs impact through disruptive eventsâ. EN. In: Big Data & Society 11.4 (Dec. 2024), p. 20539517241290220. issn: 2053-9517. doi: 10.1177/20539517241290220. url: https://doi. org/10.1177/20539517241290220. [120] Gabriella Waters. âAI testing, evaluation, verification and validation for accessibility: a comprehensive frame- workâ. In: Frontiers in Digital Health 7 (2026), p. 1679603. doi: 10.3389/fdgth.2025.1679603. [121] Laura Weidinger et al. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv:2310.11986 [cs]. Oct. 2023. doi: 10.48550/arXiv.2310.11986. url: http://arxiv.org/abs/2310.11986. [122] Laura Weidinger et al. Toward an Evaluation Science for Generative AI Systems. arXiv:2503.05336 [cs]. Mar. 2025. doi: 10.48550/arXiv.2503.05336. url: http://arxiv.org/abs/2503.05336. [123] Michael Windle et al. âFrom Epidemiologic Knowledge to Improved Health: A Vision for Translational Epidemiologyâ. In: American Journal of Epidemiology 188.12 (Dec. 2019), p. 2049â2060. issn: 0002-9262. doi: 10.1093/aje/kwz085. url: https://pmc.ncbi.nlm.nih.gov/articles/ PMC8045479/. [124] Dawn Woodard et al. âPredicting travel time reliability using mobile phone GPS dataâ. In: Transportation Research Part C: Emerging Technologies 75 (Feb. 2017), p. 30â44. issn: 0968-090X. doi: 10.1016/j.trc. 2016.10.011. url: https://w.sciencedirect.com/science/article/pii/ S0968090X16302042. [125] Qing Xiao et al. AI Hasnât Fixed Teamwork, But It Shifted Collaborative Culture: A Longitudinal Study in a Project-Based Software Development Organization (2023-2025). arXiv:2509.10956 [cs] version: 1. Sept. 2025. doi: 10.48550/arXiv.2509.10956. url: http://arxiv.org/abs/2509.10956. [126] Fasheng Xu et al. Generative AI and Organizational Structure in the Knowledge Economy. arXiv:2506.00532 [econ]. May 2025. doi: 10.48550/arXiv.2506.00532. url: http://arxiv.org/abs/ 2506.00532. [127] Xin Ming Ye and Aakash Ranganathan. âAI Doesnât Reduce Work â It Intensifies Itâ. In: Harvard Business Review (Feb. 2026). Based on an ethnographic study of generative AI use at a U.S. technology company. url: https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies-it. [128] Norhayati Zakaria. âEdward Hall: High-Context versus Low-Context Intercultural Communicationâ. In: Cul- ture Matters. Num Pages: 8. CRC Press, 2016. [129] Zheng Zhang et al. Learning to Complement with Multiple Humans. arXiv:2311.13172 [cs] version: 2. May 2024. doi: 10.48550/arXiv.2311.13172. url: http://arxiv.org/abs/2311.13172. [130] Kaitlyn Zhou et al. Attention to Non-Adopters. arXiv:2510.15951 [cs]. Oct. 2025. doi: 10.48550/arXiv. 2510.15951. url: http://arxiv.org/abs/2510.15951. 28 [131] Yan Zhuang et al. âPosition: AI Evaluation Should Learn from How We Test Humansâ. en. In: June 2025. url: https://openreview.net/forum?id=MxCJbuJhWG. 29