Paper deep dive
Making AI Evaluation Deployment Relevant Through Context Specification
Matthew Holmes, Thiago Lacerda, Reva Schwartz
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 12:28:04 AM
Summary
The paper introduces 'context specification' as a foundational process for AI evaluation, shifting focus from model-centric benchmarks to deployment-relevant constructs. By systematizing stakeholder perspectives and operational realities into clear, measurable constructs, organizations can better assess AI systems' real-world impacts, risks, and value, ultimately informing deployment decision-making.
Entities (5)
Relation Signals (3)
Context Specification â informs â AI Evaluation
confidence 98% ¡ We argue that systematic context specification is the foundational step for evaluating AI in the real-world.
Context Specification â produces â Context Brief
confidence 95% ¡ The primary output of the process is the Context Brief
Context Specification â issubtypeof â Construct Systematization
confidence 90% ¡ Context specification can be understood as a subtype of construct systematization
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:With many organizations struggling to gain value from AI deployments, pressure to evaluate AI in an informed manner has intensified. Status quo AI evaluation approaches mask the operational realities that ultimately determine deployment success, making it difficult for decision makers outside the stack to know whether and how AI tools will deliver durable value. We introduce and describe context specification as a process to support and inform the deployment decision making process. Context specification turns diffuse stakeholder perspectives about what matters in a given setting into clear, named constructs: explicit definitions of the properties, behaviors, and outcomes that evaluations aim to capture, so they can be observed and measured in context. The process serves as a foundational roadmap for evaluating what AI systems are likely to do in the deployment contexts that organizations actually manage.
Tags
Links
- Source: https://arxiv.org/abs/2603.06811v1
- Canonical: https://arxiv.org/abs/2603.06811v1
Trouble viewing inline? Open PDF directly â
Full Text
43,998 characters extracted from source content.
Expand or collapse full text
Making AI Evaluation Deployment-Relevant Through Context Specification Matthew Holmes Intellect Frontier UK info@intellectfrontier.co.uk Thiago Lacerda Trustworks USA thiago@lacerda.com Reva Schwartz Civitaas Insights USA https://orcid.org/0000-0002-9012-6306 AbstractâWith many organizations struggling to gain value from AI deployments [1], pressure to evaluate AI in an informed manner has intensified. Status quo AI evaluation approaches mask the operational realities that ultimately determine deploy- ment success, making it difficult for decision-makers outside the stack to know whether and how AI tools will deliver durable value. We introduce and describe context specification as a process to support and inform the deployment decision making process. Context specification turns diffuse stakeholder perspectives about what matters in a given setting into clear, named constructsâexplicit definitions of the properties, behav- iors, and outcomes that evaluations aim to captureâso they can be observed and measured in context. The process serves as a foundational roadmap for evaluating what AI systems are likely to do in the deployment contexts that organizations actually manage. Index TermsâAI evaluation; context specification; real-world measurement; socio-technical systems; deployment decisions; I. INTRODUCTION AND MOTIVATION: FROM BENCHMARKS TO DECISION-GRADE EVALUATION As organizations adopt AI systems at increasing scale, decision makers are being asked to justify whether and how these systems should be deployed and used in their operational settings. The most consequential impacts of AI unfold in real contexts and over time, reshaping organizational processes and incentives and producing a growing list of downstream societal impacts [2]. Many of these impacts arise from how people adapt to, misuse, or over-rely on AI in specific settingsâfactors that status quo evaluation methods largely overlook as they remain oriented toward model capabilities and optimization [3]â[6]. Meanwhile, stakeholders outside the stack aim to answer decision-critical questions about AIâs real-world value and risk, such as: âWhat do we need to know about how this system behaves in our setting so that we can decide whether to adopt it, where it should be used, and what guardrails it requires?â When evaluation outcomes correlate poorly with what actually materializes in deployment environments, these stakeholders are left with little visibility into whether an AI deployment will generate genuine value or merely shift burdens to new parts of the workflow. In this paper, the term âdeploymentâ refers to the adoption and implementation of AI tools in usersâ settings. This work was conducted as part of the Forum for Real-World AI Mea- surement and Evaluation (FRAME) at Virginia State Universityâs Center for Responsible AI A. The Need for Well-Defined Constructs Addressing the gap between current evaluation outcomes and what stakeholders need requires a shift from generic performance metrics toward explicitly defining the outcomes and behaviors that matter in a given deployment environment, and translating them into concrete measurement targetsâalso referred to as âconstructsâ. This process of turning diffuse ideas into precise, measurable constructs is known as construct systematization and is widely used in other measurement domains [7]. Existing approaches to construct development for AI evalu- ation fall into several categories. Some benchmarks overlook constructs entirely, treating scores as self-interpreting without specifying what real-world concept they are meant to capture [3], [5]. Others import constructs from model tuning and optimization processes inside the AI stack, where labels like âreasoningâ or âhelpfulnessâ are treated as self-evident con- cepts rather than being fully defined for repeatable observation and measurement [7]. A complementary set of methods such as co-design, participatory techniques, and human-centered design are used to surface rich, locally grounded constructs around values, harms, and preferences, but primarily to support the design of systems rather than their evaluation [8]â[11]. A central challenge for these methods is whether the con- structs we use to design and develop AI bear a stable rela- tionship to what actually materializes in real-world settings, or whether internal metrics are drifting proxies that fail to account for downstream impacts on people and institutions. Constructs that connect well to real-world settings can be leveraged for evaluation, but even formalized constructs such as fairness remain under-examined with respect to whether model-level disparity measures reliably track who benefits or is burdened in the real-world [4], [12]â[15]. When measures do not faithfully track materialized outcomes, deployment and adoption decisions are effectively being made on numbers that look rigorous but offer weak guidance about when a system is safe to rely on, where it is likely to fail, or which communities will bear the risks. Well specified constructs give organizations a stable reference point for assessment, enabling a shift away from trend-driven experimentation toward a clearly articulated destination and set of landmarks for whether, when, and how to deploy AI technologies. arXiv:2603.06811v1 [cs.AI] 6 Mar 2026 B. Thesis and Problem Statement: Why Context Specification is Foundational Context specification can be understood as a subtype of con- struct systematization that remains explicitly focused on âwhat mattersâ to those in the deployment context. Rather than being driven by model developersâ internal objectives, context spec- ification is carried out for, by, and in coordination with stake- holders outside the AI stack [16]. In doing so, it directs con- struct work toward real-world evaluation targetsâhow systems actually perform in useârather than toward model-centric abstractions [7], [17]. To support valid deployment-relevant claimsâespecially about downstream outcomesâconstructs elicited through this process must be described in ways that are technically precise, neutral, collectively informed, and un- ambiguous so that different actors cannot reasonably interpret them in incompatible ways. Many organizations adopting AI lack dedicated evaluation capabilities and are therefore left to rely on whatever tools and metrics are readily available. Those tools are rarely de- signed for the questions and settings relevant to organizations deciding whether to deploy, adopt or use AI for their own use cases, so evaluation outcomes do not translate cleanly into answers about whether a system is safe, usable, or valuable in context [18]. At the same time, the evaluation community more broadly lacks a stable, rigorous way to define the real-world concepts of interest to AI adopters and users. Benchmarks provide performance snapshots under controlled conditions; participatory and design-oriented methods surface values and preferences; model-internal metrics track optimization objec- tives. Yet none of these, on their own, ensure that what is measured corresponds to what ultimately materializes in deployment. We argue that systematic context specification is the foun- dational step for evaluating AI in the real-world. It translates stakeholder prioritiesâwhat matters to them in practiceâand use contexts into evaluable constructsâexplicit, named de- scriptions of properties, behaviors, and outcomes that can be observed and measured in a given setting. By making these constructs explicit and linking them to candidate observables, context specification creates the conditions for systematic knowledge rather than isolated performance claims, and deliv- ers stakeholders the information they need to inform go/no-go decisions. This allows stakeholders outside the stack to gain direct knowledge about what matters in their own use cases so they can prioritize, manage, and govern AI deployments in ways that align with their goals and constraints, and to focus their resources on problems that actually materialize in practice rather than on abstract or hypothesized risks. Without systematic context specification, evaluation tends to collapse into three recurring problems. First, effects are misattributed: observed outcomes are interpreted as model per- formance when they may arise from workflow and contextual constraints, incentive structures, or humanâsystem interaction patterns. Second, evaluation relies on brittle proxies that appear rigorous but drift from the real-world phenomena they are meant to capture. Third, deployment decisions are made on metrics that lack a stable relationship to downstream impacts on people and institutions. Context specification addresses these problems not by prescribing particular metrics, but by clarifying what is being measured and why from the perspective of those in the deployment setting. It establishes the conceptual grounding required for later evaluation design, execution, and analysis. I. CONCEPTUAL FOUNDATION: WHAT CONTEXT SPECIFICATION YIELDS The preceding section explains why context specification is necessary. This section clarifies what it produces as seen in Figure 1. Unlike participatory methods that focus on what users want a technology to do or how it should be designed, systematic context specification uses stakeholder input to define and structure the deployment-relevant concepts that evaluations must measure. It clarifies what matters for an AI deployment in a particular setting so that assessments of utility, risk, and safety are tied to realistic operational matters rather than abstract desiderata. The process yields a structured set of outputs that serves as the bridge to evaluation design: ⢠Named stakeholder priorities, articulated in relation to a specific deployment setting and focused on the real- world impacts that matter to those who will live with the systemâs consequences. ⢠Evaluable constructs, defined as clear descriptions of properties, behaviors, or outcomes that can be meaning- fully observed in that setting. ⢠Context-of-use elements, including workflows, con- straints, incentive structures, institutional norms, and likely use variants. ⢠Linking mechanisms, which describe plausible pathways through which system behavior in context produces ob- servable outcomes. ⢠Candidate observables and evidence needs, distin- guishing what can be inferred from model outputs (in silico) and what requires observation in deployment (in situ). ⢠Explicit assumptions and uncertainties, identifying what is known, what is inferred, and what remains empirically open about which outcomes and behaviors matter in the deployment context. A linking mechanism is a description of how system be- havior unfolds when used by people to produce real-world outcomes. For example, in a hiring workflow, an AI ranking model may influence human reviewers not only through its output scores but also by shaping which applications are surfaced first. The mechanism may involve cognitive shortcuts (e.g., defaulting to top-ranked candidates), workload pressures, or organizational incentives. These pathways are not visible in model-level metrics but can be articulated during context specification to inform construct development and subsequent evaluation activities. The fact that AI systems shape human decision making is not inherently negative; the problem arises 2 Fig. 1. Context specification serves as the âContextualizeâ step in the CIRCLE real-world AI evaluation lifecycle from [16]. when these influences are poorly understood and unmodeled, leaving key aspects of AIâs real-world impacts unmeasured and unmanaged. By making constructs and mechanisms explicit, context specification delineates what evaluation must observe in order to support deployment-relevant claims. It therefore transforms diffuse foci into structured measurement targets, without yet selecting methods or metrics. I. METHOD: A DESCRIPTIVE PROCESS FOR SYSTEMATIC CONTEXT SPECIFICATION The method presented here is descriptive rather than pre- scriptive, and does not mandate particular controls or stan- dards. In contrast to participatory methods that aim to surface what stakeholders want an AI system to do, the goal of context specification is to determine the information stakeholders need about what AI means in their own setting: what changes it could provoke in workflows, roles, risks, and accountability, and how those changes influence whether, where, and how to deploy it at all. Context specification facilitates the collection of this in- formation and outlines a repeatable sequence that organiza- tions can use to clarify what matters in a given deployment setting for the purpose of evaluation. To effectively capture deployment-focused constructs, this sequence is best informed by those who: ⢠make or shape system deployment and adoption deci- sions, ⢠will use, be accountable for or be affected by evaluation findings, and ⢠stand to be empowered or disadvantaged by how evalua- tion criteria are set [19]â[21]. In line with theory of change approaches, the process is structured as Inputsâ Activitiesâ Outputsâ Outcomes, making the assumed links between what the system does in context and the outcomes that matter to stakeholders explicit [22], [23]. Teams conducting context specification activities should exhibit a blend of technical, measurement, participatory, and translational skills to effectively elicit local knowledge, anticipate evaluation requirements, and translate insights across the different disciplines in the AI evaluation ecosystem [9], [24], [25] A. Inputs Inputs ground the specification process in the realities of deployment rather than in abstract model capabilities defined from a development lens. They include detailed information about who is participating in the process (stakeholders and roles, including decision makers, system users and operators, potentially affected individuals, and oversight functions), AI system purpose and anticipated contexts of use (including foreseeable variants), associated operational constraints and institutional norms, relevant documentation and logs, and known regulatory or governance touch points. B. Activities Activities translate heterogeneous stakeholder inputs into a structured basis for measurement. They typically include (i) elicitation and synthesis to surface and refine stakeholder priorities and relevant contextual factors; (i) systematization in which these are grouped, filtered, and articulated as candidate systematized constructs; and (i) preliminary operationaliza- tion, in which plausible mechanisms and pathways are mapped to link system behaviors to observable indicators and down- stream outcomes [7], [17], [26]. 1) Elicitation modes: Elicitation is used to collect diffuse input about existing or future AI deployments from relevant stakeholder communities for evaluation design purposes. It can include in-person interviews and workshops, surveys, document review, and asynchronous exercises with relevant stakeholders that reflect organizational constraints and per- spectives on the role of AI in their setting. Effective elicitation collects both explicit knowledge (documented policies, proce- dures, formal requirements) and tacit knowledge (experiential insights, uncodified workflow realities) needed to understand âwhat mattersâ in a given setting and how that setting actually operates [27] LLM-driven approaches can be used to pre-structure the issue space before stakeholder engagement occurs, to stand in as a lightweight elicitation tool when direct engage- ment is constrained, or to iteratively prioritize and refine stakeholder feedback over time [28]. Automated extraction and synthesis can summarize large volumes of documented knowledge (for example, surfacing system requirements and constraints, or edge cases from policies, logs, or job post- ings) and propose candidate topics for discussion [29], [30]. However, tacit knowledgeâsuch as incentive structures, in- formal workarounds, and time pressureâtypically requires direct stakeholder engagement. The method therefore treats automated extraction and human elicitation as complementary rather than interchangeable. 3 Fig. 2.Context specification as the deployment-to-evaluation translation step: turning stakeholder priority items into evaluable constructs and evidence needs. C. Outputs The primary output of the process is the Context Brief, a structured artifact that clarifies and defines the priority items to be evaluated from the perspective of stakeholders in the deployment context. It includes stakeholder details, intended and likely use contexts, constraints and norms, decision points (e.g., pilot go/no-go), candidate mechanisms and effects, key uncertainties, and evidence needs (including whether they are best addressed in silico or in situ, framed as questions of observability). The brief also provides a ItemâConstructâ Indicator mapping that translates stakeholder-articulated prior- ity items about âwhat mattersâ into evaluable (or systematized) constructs. The context specification team maps the constructs to associated candidate indicators or prompts that can be tied to specific behaviors and outcomes, to inform the evaluation design process. These outputs do not constitute final metrics; rather, they establish the constructs and indicators that evaluations must address and explicitly demonstrate why those targets matter to create transparency around what evaluation aims to capture. D. Outcomes For organizations, context specification yields a more durable shift in how AI evaluation is understood and prac- ticed. First, it produces a clear, shared articulation of what matters to the relevant stakeholder community, including how success and harm are defined in a specific setting. Second, it generates a systematized set of constructs and associated indicators that provide a structured basis for selecting and combining evaluation methods, rather than relying on ad hoc or purely model-centric choices. Third, it creates an explicit record of key uncertainties and observability limits, clarifying which questions can be answered in silico, which require in situ observation, and where residual ambiguity will remain. Together, these outcomes equip stakeholders to design and sequence evaluations, interpret the resulting evidence in light of local priorities and constraints, and make downstream decisions such as AI pilot design, go/no-go determinations, scaling thresholds, or decommissioning. They also support institutional learning over time, as successive context briefs and evaluation cycles can be compared to track how items, constructs, and decision criteria evolve across deployments and sectors. E. Handoff to evaluation design choices The handoff to evaluation design occurs when the struc- tured outputs constrain method choice. Constructs and linking mechanisms identified during context specification determine which phenomena must be observed in situ, which require longitudinal study, and which may be adequately approximated under controlled conditions. In this way, context specification informsânot replacesâsubsequent evaluation planning. IV. EXAMPLE USE CASE We leverage an example use case to illustrate how context specification can produce structured inputs that inform orga- nizational governance and decision making about AI deploy- ments. Use Case:Publicly owned rail transport operator aiming to deploy an AI-driven HR screening system with chatbot functionality to support hiring decisions for Rail Operations Controller roles. The AI-driven system combines predictive components (such as ranking candidates based on skills or generating a pre- dicted job-success score for each applicant) with a generative chatbot interface that allows HR professionals to query and explore candidate information. This reflects a common real- world use case in which AI behavior and its consequences de- pend on how people and tools interact in a given organizational setting. In this case, task-specific predictions feed directly into ranked lists and scores, and indirectly shape how HR staff talk and think about candidates during hiring discussions [31]. Because AI-based HR screening tools have already shown patterns of harmful bias that can influence downstream decisions and trigger operational and regulatory concerns [32], this example highlights how model behavior, organizational practices, contextual constraints, and user perceptions jointly determine real-world impact. Context specification can drive AI evaluation design efforts that surface real-world risks and specify future measurement targetsâin contrast to status-quo evaluation methods that treat these contextual variables as noise to be eliminated. A. Methods We use Table I and Table IIto apply the methods outlined in Section I to the AI-driven HR use case, demonstrating how the rail provider could use context specification to identify what they need to know about how the AI tool will behave in their environment, so they can make more informed deploy- ment decisions. Table I displays example types of input, output, that may be generated during the context specification process. Inputs and activities are not fixed pairs; each is defined to be reusable. 4 Input Category Input Type Activities (Description) Example Output Organizational set-ting and constraints Org chart; basic HRmetrics; public com-pany information Process mapping of existingHR workflows and constraintsfor Rail Operations Controllerhiring, plus LLM-assisted search to gather HR policies,recruitment documents, and prior decisions Mid-sized rail operator with 1,000 employees, hiring for Rail Operations Controller roles spread across 20 regions;high-volume, time-pressured recruitment for safety-criticalposts System, use case,and scope Vendor documentationfor the screening model; procure- ment/business case; internal deployment note Wireframing to specify whatthe AI screening system andchatbot do, where they sit inthe hiring flow, and how sys-tem outputs are leveraged androuted across the organization Tool Info: Third-party HR screening tool that ranks applicantsfor Rail Operations Controller roles, marketed as improvingspeed and consistency in shortlisting; intended use limited topre-interview ranking and support for HR screeners Evaluation purposeand scope Pre-engagement mate-rials such as emailthreads, briefings, andinternal memos sum-marising why the re-view is being commis-sioned Interviews with HR staff, hir-ing managers, and senior lead-ers to clarify expectations, an-ticipated or already experi-enced system risks and bene-fits, and accountability struc-tures HR staff had made note of how the AI-enabled hiring processmight alter risks in shortlisting; senior leadership decided tocommission an evaluation to understand the risk and oppor-tunity surface of the AI-influenced Rail Operations Controllerhiring process Governanceand compliance expectations Existing internal policies, procedures, and governance documents (both AI-specific and general), plus any documented external regulatory or standardsobligations LLM-assisted search to iden-tify applicable policies andguidance to link them to con-crete constructs of interest suchas fairness, over-reliance, andproductivity The company had identified several compliance obligationsand governance standards for which understanding the de-ployment context and its potential impacts would be useful,motivating a closer look at how the AI-enabled hiring processoperates in practice Stakeholders and roles Org chart; documentsand emails related to the system procurement and how the systemâs usemay impact the hiringprocess; informal knowledge about whois involved or affected Mock-up of field testing forselected HR staff to see howscores and chat responses playout under realistic conditions A list of stakeholders and their roles is compiled, includingdecision-makers, day-to-day users, and affected groups (e.g.HR professionals, senior leadership, and Rail Operations Con-troller candidates) TABLE I C ONTEXT S PECIFICATION FOR AI-D RIVEN HR S CREENING : I NPUTS , A CTIVITIES , AND E XAMPLE O UTPUTS 5 and combinable, so that information can be elicited in ways that fit real-world context and constraints. Activities and inputs are selected using context-sensitive elicitation choices: methods are chosen to fit the specific orga- nizational setting, including existing workflows, constraints on time and resources, culture, and the particular properties of the AI system being evaluated, in line with community-centered norm-elicitation approaches [33].The activities for the use case are therefore selected in this spirit, reflecting the priorities surfaced by stakeholders rather than a rigid set of tasks. Table I displays how stakeholder-defined priorities from the use case can map to related constructs. Priority ItemsRelated Constructs Is this tool creating rework for my staff, or actually saving time? Productivity Are employees over-relying on the tool and losing critical skills? Over-reliance Is the tool shifting liability to our frontline workers? Accountability Who is consistently benefiting from these systems, and who is absorbing new risks? Fairness and equity TABLE I MAPPING STAKEHOLDER-ELICITED PRIORITIES TO RELATED CONSTRUCTS As the final artifact of the process, a brief sample of the Context Brief for the Rail Operations Controller hiring example follows: ⢠Stakeholders / roles: HR screeners; hiring manager; operations representative responsible for coordinator per- formance. ⢠Workflow / decision points: System ranks candidates; HR may accept, override, or escalate rankings; chatbot used for borderline or surprising candidates and during hiring meetings. ⢠Constraints / norms: High application volume; time pressure; limited transparency into system logic; emerg- ing norm of trusting top-ranked candidates by default. ⢠Linking mechanisms: Over-reliance on rankings under time pressure; chatbot summaries anchoring candidate framing; selection criteria potentially misaligned with disruption-management skills. ⢠Uncertainties / evidence needs: Frequency and direction of overrides; order of review (summary vs. application); whether overrides are recorded; gap between shortlisted skills and industry norms. By following the inputs and activities summarized in Ta- ble I, context specification for this use case yields three main outcomes: a clear articulation of what matters to HR staff, hiring managers, and affected candidates; a set of systematised constructs (such as productivity, over-reliance, shifting accountability, and fairness), and shared criteria for what counts as success or harm in AI-enabled Rail Operations Controller hiring for the stakeholder organization. These out- comes provide a structured basis for selecting and combining evaluation methods and an explicit record of key uncertainties and observability limits, enabling stakeholders to design and sequence evaluations, interpret evidence about the screening tool, and make downstream decisions about piloting, go/no-go, scaling, or decommissioning. V. EVALUATION DESIGN CHOICES AS TRADEOFFS Evaluation methods are not neutral instruments; they shape what can and cannot be observed. Once constructs and linking mechanisms are specified, method selection becomes a design decision involving tradeoffs between control and contextual richness. Highly controlled methods (e.g., simulation, benchmarking, scenario modeling) offer clarity and reproducibility but may abstract away the contextual conditions through which down- stream effects emerge. In contrast, high-context methods (e.g., field deployments, longitudinal observation, embedded stud- ies) capture richer interaction effects but introduce variability and may require additional resources to enforce experimental constraints. Context specification clarifies which constructs require which forms of evidence. For example, if a linking mechanism involves over-reliance driven by workflow pressures, purely in- silico methods will be insufficient and observational or field- based methods may be required. Conversely, constructs related to baseline system behavior independent of human interaction may be served by in silico testing methods alone. Rather than prescribing a fixed hierarchy of methods, this framework positions method selection as contingent on the constructs identified during context specification. By making these dependencies explicit, organizations can design evalua- tion strategies that are proportionate to the risks and uncer- tainties present in their deployment context. VI. DISCUSSION: LIMITATIONS AND FUTURE WORK 1) Example Use Case: While designed to be realistic, the hypothetical use case provided in this paper cannot stand in for empirical observations or fully capture the organizational dynamics, politics, or historical dependencies that shape real deployments. It also focuses on a single high-risk public-sector context where decision making may lean more on the side of governance and compliance as opposed to operational innovation. Future work should apply the approach in real deployments, to see how activity selection, elicitation, and context-brief production work under actual time, staffing, and budget con- straints. It should span multiple domains and risk profiles to test how far the approach transfers, which parts generalize, and where adaptation is needed for different institutional and sectoral settings. 2) Execution capability and culture: Real-World examples of AI impact assessment work suggest that organizations can lack the in-house expertise needed to design and run robust, 6 context-aware elicitation processes [34], weakening the quality of such assessments. The effectiveness of context specification requires clear elicitation protocols [35], the inclusion of stakeholders with appropriate local knowledge, and a broad set of socio-technical skills across the team conducting the exercise. In many or- ganizations, developing and applying this kind of capability remains a challenge. Further research should characterize the capability gap for conducting context specification across organizations and sectors, including typical needs for upskilling, role design, and external support. Empirical work should also examine how organizational politics, incentives, and culture limit what can be elicited, agreed, and recorded, and how this in turn affects the quality and defensibility of context briefs. 3) Data and documentation preconditions: The method presupposes that internal artifacts (for example, policies, pro- cess maps, and incident logs) exist and are of sufficient quality to support context specification. Reviews of AI documenta- tion and risk-governance practice find that documentation is frequently incomplete, inaccurate, or not up to date due to incentive, resourcing, and workflow constraints, limiting how far contextual analysis and risk management can go in practice [36]. Future work should examine how variability in internal data and documentation quality affects feasibility and outputs, including how organizations can scope, phase, or compensate when documentation is incomplete, inconsistent, or unreliable. 4) Construct maturity and measurement readiness: The ap- proach relies on systematized constructs (such as over-reliance and sycophancy) to translate contextual priorities into evalu- able targets. As these techniques are still nascent practice in the AI evaluation ecosystem, construct sets and associated measurement targets remain incomplete and may omit relevant risk or opportunity dimensions [7]. Finally, further work should refine and expand the construct set and associated measurement targets used in context spec- ification, testing them across domains and deployment con- ditions. This includes developing taxonomies, thresholds, and instrumentation that make socio-technical risk and opportunity surfaces more systematically identifiable and evaluable over time. VII. CONCLUSION The current AI evaluation ecosystem reflects its roots in system design and development. While useful for optimizing and comparing model capabilities, it provides limited insight for stakeholders deciding whether and how to deploy these technologies in their own settings. This paper introduces a context specification process for identifying âwhat mattersâ in deployment environments so that key constructs can be developed to guide system evaluation and inform deployment decisions. By connecting evaluation criteria to stakeholder- defined constructs, context specification creates a new pathway for decision-makers to obtain the information they need, facilitating responsible AI practice and enhancing the value of AI technologies in operational contexts. VIII. DEFINITIONS OF RELEVANT TERMS Construct: An abstract, latent concept or theoretical vari- able, not directly observable, that is defined for scientific purposes and measured indirectly through multiple observable indicators or items. Construct operationalization: Linking a systematized con- cept to appropriate indicators and scores in a coherent mea- surement framework [17]. Construct systematization: The process of clarifying and organizing a concept by specifying its meaning, dimensions, and relationships to other concepts [17]. Evaluation: (1) Systematic determination of the extent to which an entity meets its specified criteria; (2) Action that assesses the value of something [37]. In silico testing: Testing or experimentation carried out entirely on a computer, using computational models and simulations. IX. AI USE DISCLOSURE This manuscript was prepared with targeted assistance from AI technology. AI was used to (i) generate and format some BibTeX citation entries, (i) draft and refine LaTeX table code and cross-references, and (i) provide limited support for restructuring and line-editing sections of the text (e.g., smoothing phrasing, tightening paragraphs). All conceptual contributions, methodological design, analysis, arguments, and final editorial decisions were made by the authors, who also reviewed and edited all AI-assisted content for accuracy and appropriateness. REFERENCES [1] MLQ AI, âState of ai in business 2025,â 2025, version 0.1. [Online]. Available: https://mlq.ai/media/quarterly decks/v0.1State ofAIinBusiness2025Report.pdf [2] M. Selgas-Cors, âSociotechnical transformation: A systematic review on the impact of artificial intelligence on society and organizations,â FinTech and Sustainable Innovation, vol. 00, no. 00, p. 1â16, 2025. [Online]. Available: https://ojs.bonviewpress.com/index.php/FSI/article/ view/6076/1547 [3] A. M. Bean, R. O. Kearns, A. Romanou, F. S. Hafner, H. Mayne, J. Batzner, N. Foroutan, C. Schmitz, K. Korgul, H. Batra, O. Deb, E. Beharry, C. Emde, T. Foster, A. Gausen, M. Grandury, S. Han, V. Hofmann, L. Ibrahim, H. Kim, H. R. Kirk, F. Lin, G. K.-M. Liu, L. Luettgau, J. Magomere, J. Rystrøm, A. Sotnikova, Y. Yang, Y. Zhao, A. Bibi, A. Bosselut, R. Clark, A. Cohan, J. N. Foerster, Y. Gal, S. A. Hale, I. D. Raji, C. Summerfield, P. Torr, C. Ududec, L. Rocher, and A. Mahdi, âMeasuring what Matters: Construct Validity in Large Language Model Benchmarks,â Oct. 2025. [Online]. Available: https://openreview.net/forum?id=mdA5lVvNcU [4] H. Wallach, M. Desai, A. F. Cooper, A. Wang, C. Atalla, S. Barocas, S. L. Blodgett, A. Chouldechova, E. Corvi, P. A. Dow, J. Garcia- Gathright, A. Olteanu, N. Pangakis, S. Reed, E. Sheng, D. Vann, J. W. Vaughan, M. Vogel, H. Washington, and A. Z. Jacobs, âPosition: Evaluating Generative AI Systems Is a Social Science Measurement Challenge,â Jun. 2025, arXiv:2502.00561 [cs]. [Online]. Available: http://arxiv.org/abs/2502.00561 [5] O. Salaudeen, A. Reuel, A. Ahmed, S. Bedi, Z. Robertson, S. Sundar, B. Domingue, A. Wang, and S. Koyejo, âMeasurement to Meaning: A Validity-Centered Framework for AI Evaluation,â Jun. 2025, arXiv:2505.10573 [cs]. [Online]. Available: http://arxiv.org/abs/2505. 10573 7 [6] L. Weidinger, I. D. Raji, H. Wallach, M. Mitchell, A. Wang, O. Salaudeen, R. Bommasani, D. Ganguli, S. Koyejo, and W. Isaac, âToward an evaluation science for generative ai systems,â 2025. [Online]. Available: https://arxiv.org/abs/2503.05336 [7] A. Chouldechova, C. Atalla, S. Barocas, A. F. Cooper, E. Corvi, P. A. Dow, J. Garcia-Gathright, N. Pangakis, S. Reed, E. Sheng, D. Vann, M. Vogel, H. Washington, and H. Wallach, âA Shared Standard for Valid Measurement of Generative AI Systemsâ Capabilities, Risks, and Impacts,â Dec. 2024, arXiv:2412.01934 [cs]. [Online]. Available: http://arxiv.org/abs/2412.01934 [8] WorldEconomicForum,âAIValueAlignment:Guiding ArtificialIntelligenceTowardsSharedHumanGoals,â WorldEconomicForum,Geneva,Switzerland,Tech. Rep.[Online].Available:https://w.weforum.org/publications/ ai-value-alignment-guiding-artificial-intelligence-towards-shared-human-goals/ [9] F. Delgado, S. Yang, M. Madaio, and Q. Yang, âThe Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice,â in Equity and Access in Algorithms, Mechanisms, and Optimization. Boston MA USA: ACM, Oct. 2023, p. 1â23. [Online]. Available: https://dl.acm.org/doi/10.1145/3617694.3623261 [10] A. Mangold, J. Zietz, S. Weinhold, and S. Pannasch, âOn the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy,â Oct. 2025, arXiv:2510.12201 [cs]. [Online]. Available: http://arxiv.org/abs/2510.12201 [11] B. Friedman, P. Kahn Jr., and A. Borning, âValue Sensitive Design: The- ory and Methods,â in Proceedings of the 7th Conference on Designing Interactive Systems, 2002. [12] A. Jacobs and H. Wallach, âMeasurement and Fairness,â Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, p. 375â385, Mar. 2021. [Online]. Available: http: //arxiv.org/abs/1912.05511 [13] M. Sap, S. Gabriel, L. Qin, D. Jurafsky, N. A. Smith, and Y. Choi, âSocial Bias Frames: Reasoning about Social and Power Implications of Language,â in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.Online: Association for Computational Linguistics, 2020, p. 5477â5490. [Online]. Available: https://w.aclweb.org/anthology/2020.acl-main.486 [14] D. K. Mulligan, J. A. Kroll, N. Kohli, and R. Y. Wong, âThis Thing Called Fairness: Disciplinary Confusion Realizing a Value in Technology,â Proceedings of the ACM on Human-Computer Interaction, vol. 3, no. CSCW, p. 1â36, Nov. 2019. [Online]. Available: https://dl.acm.org/doi/10.1145/3359221 [15] S. Barocas, M. Hardt, and A. Narayanan, Fairness and Machine Learn- ing. MIT Press, 2023. [16] R. Schwartz, R. Chowdhury, A. Kundu, H. Frase, M. Fadaee, T. David, G. Waters, A. Taik, M. Briggs, P. Hall, S. Jain, K. Yee, S. Thomas, S. Bhandari, P. Duncan, A. Thompson, M. Carlyle, Q. Lu, M. Holmes, and T. Skeadas, âReality check: A new evaluation ecosystem is necessary to understand aiâs real world effects,â 2025. [Online]. Available: https://arxiv.org/abs/2505.18893 [17] R. Adcock and D. Collier, âMeasurement validity: A shared standard for qualitative and quantitative research,â American Political Science Review, vol. 95, no. 3, p. 529â546, 2001. [18] C. Yu, S. Engelmann, R. Cao, D. Ali, and O. Papakyriakopoulos, âHow should ai safety benchmarks benchmark safety?â 2026. [Online]. Available: https://arxiv.org/abs/2601.23112 [19] D. Leslie, C. Rinc Ě on, M. Briggs, A. Perini, S. Jayadeva, A. Borda, S. Bennett, C. Burr, M. Aitken, M. Katell, C. Fischer, J. Wong, and I. Kherroubi Garcia, âAi sustainability in practice part one: Foundations for sustainable ai projects,â 2023. [Online]. Available: https://zenodo.org/doi/10.5281/zenodo.10680113 [20] K. Bell and M. Reed, âThe tree of participation: a new model for inclusive decision-making,â Community Development Journal, vol. 57, no. 4, p. 595â614, 06 2021. [Online]. Available: https://doi.org/10.1093/cdj/bsab018 [21] A. Deshpande and H. Sharp, âResponsible ai systems: Who are the stakeholders?â in Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, ser. AIES â22.New York, NY, USA: Association for Computing Machinery, 2022, p. 227â236. [Online]. Available: https://doi.org/10.1145/3514094.3534187 [22] C. H. Weiss, âTheory-based evaluation: Past, present, and future,â New Directions for Evaluation, no. 76, p. 41â55, 1997. [23] A. Reuel and T. A. Undheim, âGenerative ai needs adaptive governance,â 2024. [Online]. Available: https://arxiv.org/abs/2406.04554 [24] S.Oduro,A.E.Marwick,C.Johnson,andE.Meyer, âTroublingtranslation:Sociotechnicalresearchinaipolicy andgovernance,âInternetPolicyReview,vol.14,no.4, 2025. [Online]. Available: https://policyreview.info/articles/analysis/ troubling-translation-ai-policy-and-governance [25] M. I. Maga Ě na and K. Shilton, âFrameworks, methods and shared tasks: Connecting participatory ai to trustworthy ai through a systematic review of global projects,â in Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, ser. FAccT â25. New York, NY, USA: Association for Computing Machinery, 2025, p. 2166â2179. [Online]. Available: https://doi.org/10.1145/3715275.3732148 [26] J. P. Connell and A. C. Kubisch, âApplying a theory of change approach to the evaluation of comprehensive community initiatives: Progress, prospects, and problems,â 1998. [Online]. Available: https: //api.semanticscholar.org/CorpusID:2320879 [27] O. A. Popoola, H. E. Adama, C. D. Okeke, and A. E. Akinoso, âAdvancements and innovations in requirements elicitation: Developing a comprehensive conceptual model,â World Journal of Advanced Research and Reviews, vol. 22, no. 1, p. 1209â1220, 2024. [Online]. Available: https://wjarr.com/sites/default/files/WJARR-2024-1202.pdf [28] L. Pasquale, A. Ragone, E. Piemontese, and A. A. Darban, âExploring the use of llms for requirements specification in an it consulting company,â 2025. [Online]. Available: https://arxiv.org/abs/2507.19113 [29] A. Leiva-Araos, B. Gana, H. Allende-Cid, J. Garc Ě Äąa, and M. J. Saikia, âLarge scale summarization using ensemble prompts and in context learning approaches,â Scientific Reports, vol. 15, no. 1, p. 10259, Mar. 2025. [Online]. Available: https://w.nature.com/articles/ s41598-025-94551-8 [30] W. M. Aly, T. H. A. Soliman, and A. M. AbdelAziz, âAn evaluation of large language models on text summarization tasks using prompt engineering techniques,â 2025. [Online]. Available: https://arxiv.org/abs/2507.05123 [31] J. Wang, A. Selbst, S. Barocas, and S. Venkatasubramanian, âDistinguishing task-specific and general-purpose ai in regulation,â 2026. [Online]. Available: https://arxiv.org/abs/2506.17347 [32] K. Wilson and A. Caliskan, âGender, race, and intersectional bias in resume screening via language model retrieval,â ArXiv, vol. abs/2407.20371, 2024. [33] B. Harris, L. Alderman, and J. Staheli, âThe forgotten contexts of evaluation,â Evaluation, vol. 31, no. 2, p. 240â261, 2025. [34] European Center for Not-for-Profit Law and Danish Institute for Human Rights, âA guide to fundamental rights impact assessments (fria),â The Hague, 2025, accessed 20 February 2026. [Online]. Available: https:// ecnl.org/publications/guide-fundamental-rights-impact-assessments-fria [35] A. Wahbeh, S. Sarnikar, and O. El-Gayar, âA socio-technical-based process for questionnaire development in requirements elicitation via interviews,â Requirements Engineering, vol. 25, 09 2020. [36] A. A. Winecoff and M. Bogen, âImproving governance outcomes through ai documentation: Bridging theory and practice,â 2024. [Online]. Available: https://arxiv.org/abs/2409.08960 [37] I. O. for Standardization (ISO), I. E. C. (IEC), I. of Electrical, and E. E. (IEEE), âISO/IEC/IEEE 24765:2017, systems and software engineering â vocabulary,â Geneva, Switzerland, 2017. [Online]. Available: https://w.iso.org/standard/71952.html 8