Paper deep dive
Testing and Evaluation of Agentic AI Systems In Military Command and Control
Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/24/2026, 4:49:30 AM
Summary
This paper analyzes the challenges of Testing and Evaluation (T&E) for Agentic AI systems in military Command and Control (C2). It argues that agentic properties (e.g., persistent memory, open-ended action spaces) weaken eight foundational T&E assumptions regarding specifiability, stability, composability, and supervisability. Consequently, traditional test evidence does not warrant inferences about fielded behavior. The authors propose an assurance case framework with narrower, recoverable claims (e.g., bounded mission envelopes, trajectory-grounded correctness) and suggest that evidentiary burdens shift to deployment, requiring governance mechanisms like expiry conditions and continuous monitoring.
Entities (14)
Relation Signals (10)
Agentic AI Systems â isprocuredfor â Command and Control
confidence 95% ¡ Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight.
Agentic AI Systems â weakens â System Specifiability
confidence 93% ¡ Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims... grouped into four clusters: system specifiability...
Agentic AI Systems â weakens â System Stability
confidence 93% ¡ Agentic properties weaken all eight assumptions... grouped into four clusters: ... stability...
Agentic AI Systems â weakens â System Composability
confidence 93% ¡ Agentic properties weaken all eight assumptions... grouped into four clusters: ... composability...
Agentic AI Systems â weakens â System Supervisability
confidence 93% ¡ Agentic properties weaken all eight assumptions... grouped into four clusters: ... supervisability.
Agent Network â islaunchedby â United States Department of War
confidence 92% ¡ In June 2026, the United States Department of War launched âAgent Network,â a capability that will leverage artificial intelligence (AI) agents...
Testing and Evaluation â supports â Assurance Case
confidence 90% ¡ Whether such commitments can be discharged depends on their supporting assurance case... Through a structured review of 240 documented Testing and Evaluation (T&E) practices...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.
Tags
Links
- Source: https://arxiv.org/abs/2608.20597v1
- Canonical: https://arxiv.org/abs/2608.20597v1
Trouble viewing inline? Open PDF directly â
Full Text
187,040 characters extracted from source content.
Expand or collapse full text
Testing and Evaluation of Agentic AI Systems In Military Command and Control Ulysse Richard 1â Heather Frase 2 Sarah Cao 1,3 Di Cooke 1,4 Sebastian Kwon 1,5 Adrianna Tan 1,6 1 Arcadia Impact, AI Governance Taskforce 2 Veraitech 3 University of Oxford 4 Kingâs College London 5 Atlantic Council 6 Future Ethics Lab August 2026 Abstract Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable, contingent on mature methods: bounded mission envelopes, trajectory- grounded correctness, executable runtime constraints, and characterized run-to- run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership. â Corresponding author: ulysse.richard@arcadiaimpact.org Preprint. arXiv:2608.20597v1 [cs.SE] 20 Aug 2026 Executive Summary This research paper examines how much confidence current Testing and Evaluation (T&E) methods can justify for agentic AI systems in command and control (C2), and how remaining uncertainty should be handled in decisions to field them. Figure 1 serves as a visual summary of this paper. Why examine this challenge now? ⢠Agentic AI capabilities are being procured for C2 roles; the basis for assurance commitments warrants attention from military, industry, and policy communities. Where do established T&E methods fall short? ⢠In agentic systems, the unit of behavior is a trajectory encompassing tool use, memory, and delegation, rather than a discrete output. â˘Certain system elements emerge only at runtime. Agents may select sources and tools, or subagents can be added after certification, creating differences between tested and deployed configurations. â˘The test article is inherently dynamic; a system that accumulates state 2 is not the same as when it was originally characterized. Evidence can age without explicit updates, and objectives may shift under operational pressures. â˘Assemblies introduce challenges absent in isolated components. Emergent behaviors arise within the assembly, complicating attribution, and valid outputs at the component level may conflict when composed. â˘Supervision is itself strained. Opacity, adaptation, and delegation erode the operatorâs mental model, reduce its intervention window, and increase supervisory load. Human oversight cannot be assumed from the mere presence of an operator. How can justified confidence be recovered? ⢠Confidence should be structured as an assurance case comprising claims, evidence, and an argument that connects the two. For agentic systems, the argument is the weak link, because test-condition evidence no longer corresponds to fielded behavior in the ways established methods assume. â˘While broad assurance claims remain unsubstantiated, more narrowly scoped claims remain recoverable contingent on mature methods, including bounded mission envelopes, trajectory- grounded correctness, executable runtime constraints, and characterized run-to-run variance. â˘Part of the evidentiary burden shifts into deployment (behavioral monitoring, re-baselining, goal- drift detection, staged fielding), making the determination to field a continuing act. What cannot be evidenced must be governed, borrowing from assurance regimes in software, aviation, nuclear energy, medicine, finance, and autonomous driving. Key contributions of this paper ⢠Seven agentic properties, mapped to a generic architecture and five illustrative C2 scenarios (§2). â˘Eight assumptions underlying established T&E practice that agentic properties strain, drawn from 240 documented practices (§3, §4). 3 â˘Ten assurance claims covering specifiability, stability, and composability, with candidate methods and their evidentiary limits assessed against each (§5). Supervisability is treated only at the interface through which human direction enters the system, for the reasons given in §1.3. 2 In this paper, âstateâ primarily refers to an evolving memory state that may take the form of a text buffer, key-value store, vector database, graph structure, or any hybrid representation, as defined in Hu et al. (2025). 3 The dataset is available from the corresponding author on reasonable request. 2 FindingImplicationWhere addressed Established T&E methods rest on assumptions about specifiability, stability, composability, and supervisability that agentic systems weaken by design. Test evidence can satisfy process requirements while the inference to fielded behavior remains insufficient; the connecting argument needs di- rect scrutiny. §4 Assurance claims are recoverable in a narrower form. An assurance case decomposed into narrow claims with declared residual uncertainty can be supported by the documented record, whereas a broad system-level claim cannot. §5.1-5.3 Some evidence can only be generated in service. The determination to field becomes a continuing act under governance that time-bounds the valid- ity of evidence. §5.2 Table 1: Summary of findings and implications Mapping agentic properties to testing assumptions and assurance claims AGENTIC PROPERTIESTESTING ASSUMPTIONSASSURANCE CLAIMS SPECIFIABILITY§4.1 STABILITY§4.2 COMPOSABILITY§4.3 SUPERVISABILITY§4.4 A1Bounded behavioral space A2Specifiable inputs A3Fixed integration surface A4Behavior repeatability A5Evidence currency A6Objective stability A7System composability A8System supervisability P1Persistent memory P2Objective flexibility P3Expanding integration P4Open-ended action space P5Proactive information seeking P6Accumulated non-determinism P7Multi-agent delegation P1 P1 P2 P2 P3 P3 P4 P4 P5 P5 P6 P6 P7 P7 C1.1Bounded mission envelope C1.2Trajectory-grounded correctness C1.3Executable constraints C2.1Discrimination C2.2Diagnosis C2.3Correction C3.1Composition C3.2Interfaces C3.3Attribution & propagation C3.4Emergence §5.1 §5.2 §5.3 Property (column) strains assumption (row) â Table 5Claims recover the assumption clusterIdentified but not assessedRichard et al., 2026 Figure 1: Mapping agentic properties to testing assumptions and assurance claims. 3 Contents Executive Summary2 Lists of Tables, Figures and C2 Vignettes6 1 Introduction8 1.1Locating the challenge . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .8 1.2The structure of justified confidence . . . . . . . . . . . . . . . . . . . . . . . . .9 1.3Scope and approach . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .10 2 Definitions and Concepts10 2.1Agentic AI . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .10 2.2Command and control . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .14 2.2.1What is command and control? . . . . . . . . . . . . . . . . . . . . . . . .14 2.2.2What makes command and control effective? . . . . . . . . . . . . . . . .15 2.2.3Agentic AI in command and control: illustrative scenarios . . . . . . . . .16 3 Methodology18 4 Agentic AI Challenges Testing and Evaluation Assumptions20 4.1System specifiability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .20 4.1.1Bounded, specifiable behavioral space . . . . . . . . . . . . . . . . . . . .21 4.1.2Specifiable inputs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .22 4.1.3Fixed integration surface . . . . . . . . . . . . . . . . . . . . . . . . . . .23 4.2System stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 4.2.1Behavior repeatability . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 4.2.2Evidence currency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 4.2.3Objective stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 4.3System composability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 4.4System supervisability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .27 4 5 Closing the Assurance Gaps30 5.1Assuring system specifiability . . . . . . . . . . . . . . . . . . . . . . . . . . . .30 5.1.1Specifiable, bounded mission envelope . . . . . . . . . . . . . . . . . . .30 5.1.2Trajectory-grounded correctness . . . . . . . . . . . . . . . . . . . . . . .31 5.1.3Executable constraints . . . . . . . . . . . . . . . . . . . . . . . . . . . .32 5.2Assuring system stability . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 5.2.1Discrimination . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 5.2.2Diagnosis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .34 5.2.3Correction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .35 5.3Assuring multi-agent composition and emergence . . . . . . . . . . . . . . . . . .36 5.3.1Composition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .37 5.3.2Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 5.3.3Attribution and propagation . . . . . . . . . . . . . . . . . . . . . . . . .38 5.3.4Emergence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .39 6 Conclusion40 Appendix57 A Test and Evaluation dimensions in scope57 5 List of Tables 1Summary of findings and implications . . . . . . . . . . . . . . . . . . . . . . . .3 2Faculties of agentic systems and related terms in the literature . . . . . . . . . . .11 3Properties of agentic systems, with associated functions and architectural components 13 4Criteria of command and control effectiveness . . . . . . . . . . . . . . . . . . . .15 5Assumptions underlying established T&E practice . . . . . . . . . . . . . . . . . .20 6Assurance claims for system specifiability . . . . . . . . . . . . . . . . . . . . . .30 7Oracle strategies matched to agentic failure modes . . . . . . . . . . . . . . . . . .32 8Assurance claims for system stability . . . . . . . . . . . . . . . . . . . . . . . . .33 9Assurance claims for system composability . . . . . . . . . . . . . . . . . . . . .37 10Test and Evaluation dimensions in scope . . . . . . . . . . . . . . . . . . . . . . .57 List of Figures 1Mapping agentic properties to testing assumptions and assurance claims. . . . . . .3 2Functions of an agentic system . . . . . . . . . . . . . . . . . . . . . . . . . . . .11 3Generic agentic system architecture, with properties P1-P7 mapped to components14 4Agentic AI in command and control: Illustrative scenarios . . . . . . . . . . . . .16 List of C2 Vignettes 1A difference without a distinction . . . . . . . . . . . . . . . . . . . . . . . . . . .22 2The system that grades its own picture . . . . . . . . . . . . . . . . . . . . . . . .23 3The buck stops at the boundary . . . . . . . . . . . . . . . . . . . . . . . . . . . .23 4Two staffs, two plans . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 5Nothing to declare . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 6Initiative or insubordination . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 7Three cooks, one broth . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .27 8When approval becomes a rubber stamp . . . . . . . . . . . . . . . . . . . . . . .29 6 Contribution Statement Conceptualization, Methodology: Ulysse Richard, Heather Frase Investigation: Ulysse Richard, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan Writing â Original Draft: Ulysse Richard§1, §2, §3, §4, §5, §6 Di Cooke§5.1, §5.2 Sebastian Kwon§5.3 Sarah Cao§2.2.3, §4, §5.3 Adrianna Tan§5.2 Writing â L A T E X: Ulysse Richard Writing â Review & Editing, Supervision: Ulysse Richard, Heather Frase Project Administration: Ulysse Richard Acknowledgements The authors thank Ben R. Smith and Francesca Gomez for their feedback and project management support, and Virgile Richard for reviewing the manuscript. The paper draws on semi-structured expert interviews; we are indebted to all interviewees for their time and candor, as well as to the many others whose conversations informed this final product. Any remaining errors are our own. This research was conducted as part of the AI Governance Taskforce at Arcadia Impact, Summer 2026 cohort. AI use statement Large language models were used in the preparation of this paper. During the literature review, language models helped scan scientific databases, conduct exploratory reviews and syntheses, and perform supplementary searches to identify relevant articles not surfaced by structured queries. During analysis, AI tools helped organize extracted data into tables and graphs to support the mapping exercise. During writing, they were used for grammar and spelling, prose review and editing, figure preparation, and LaTeX formatting. The expert interviews were conducted by human researchers. All substantive analytical claims are the authorsâ own, and any AI-assisted text and outputs were reviewed and verified by the authors, who take full responsibility for the content of the paper. 7 1 Introduction Assurance commitments for agentic AI systems appear in program announcements, defense Testing and Evaluation (T&E) strategies, and policy directives. These commitments cover matters such as rigorous testing throughout development and fielding, the retention of human judgment over consequential decisions, and continuous evaluation throughout the systemâs lifecycle rather than ending at acceptance. Each commitment requires supporting evidence. For agentic systems, methods for generating such evidence fall into three categories: those currently available and usable; those available but not yet validated for this class of system; and those for which no candidate method has been identified. A structural challenge is that both emerging and established T&E methods rely on assumptions about system behavior that agentic systems weaken by design. This paper identifies affected assumptions, evaluates the value and limitations of current methods, and identifies potential methods and governance mechanisms to address remaining gaps. In June 2026, the United States Department of War launched âAgent Network,â a capability that will leverage artificial intelligence (AI) agents to continuously scan defense intelligence and operational systems and present commanders with options (DOW Unleashes âAgent Networkâ to Transform AI-Enabled Battle Management and Targeting 2026). The announcement commits to subjecting the capability to ârigorous testing, operational evaluation and oversightâ throughout development and fielding. As comparable capabilities are being developed by multiple militaries or may soon be, the basis for such commitments warrants collective attention from military, industry, and policy communities, both nationally and internationally. Battle management and decision support are command and control (C2) activities, and the Department describes Agent Network as building on prior C2 work. C2 is a demanding setting for these systems. It combines interconnected sensor networks, incomplete information, adversary manipulation, and real-time operational consequence, and its outputs pass through human operators whose reliance on the system is itself a performance variable. 4 This paper is organized around two questions. How much confidence can current evaluation methods justify for agentic AI systems in command and control? How should residual uncertainty be addressed or accounted for in decisions to field these systems? 1.1 Locating the challenge Agentic AI systems create T&E challenges across three nested levels. The model refers to the frontier AI system that performs interpretation, reasoning, and action selection. The agent comprises the model and its surrounding scaffolding, including memory, tools, and an orchestration loop. The deployed system consists of an assembly of agents and operators, along with their organizational context. At the model level, evaluating frontier AI systems remains an open problem for their developers, as illustrated by the disclosure of failures in their evaluation programs in July 2026 (OpenAI, 2026; Anthropic, 2026b). Though these incidents involve cyber capabilities, they are symptoms of a broader diagnosis: T&E practice remains largely ad hoc and lacks grounding in standards or scientific engineering principles (Department for Science, Innovation & Technology and DSIT, 2025). Because the internal mechanisms of frontier AI systems are not interpretable to their developers, evaluation strategies primarily rely on observing behavior in response to selected inputs. This observation is inherently incomplete, allowing capabilities and failure modes to emerge after deployment. Agents built on such models inherit these limitations, and while scaffolding may constrain their effects, it does not eliminate them. 4 §2.2 develops the C2 setting and the definition we adopt for it. 8 At the agent level, the scaffolding introduces complexities not present in the model alone. Established T&E methods do not require that every behavior be listed in advance. They do require that the space of possible behavior be partitionable, parameterizable, or otherwise bounded well enough for a sample to stand for the whole. Agentic systems weaken that condition. They pursue objectives autonomously, construct their own information environment by selecting sources to query, act on that environment using tools, and retain context across runs and sessions. The properties that make agents useful are often the same properties that make them difficult to evaluate (Anthropic, 2026a). For example, variation that testers have characterized statistically for a single model output behaves differently along a chain of steps, where each output becomes the next input, altering the implications of repeated trials. Additionally, an agent that retains state is not the same test article 5 from one week to the next, even in the absence of a declared change to the system. At the deployed system level, agents are assembled with other agents and with human operators, introducing additional complexity. Certain behaviors emerge only in the assembly and resist attri- bution to individual components. Moreover, fielding an agentic system changes who decides what, who exchanges information with whom, and who knows what, thereby affecting the performance of the human-machine team without altering the technology. Agent-to-agent delegation also raises the tempo and scale of activity, reducing the time between a failure occurring and its consequences reaching the operational picture. These effects are not observable when testing models or single agents in isolation, raising questions about the validity of component-level evidence for the assembled system. 1.2 The structure of justified confidence In AI assurance, justified confidence in a fielded system is established through an assurance case comprising three elements (Tate et al., 2016): â˘Claims, the specification, state what must hold for the system to be judged acceptable for its intended use. ⢠Evidence, the demonstration, is what has been observed of the system and bears on those claims. â˘An argument connects the evidence to the claim, explaining why evidence gathered under specific conditions supports a claim about the system as fielded. All three elements are necessary. Claims without evidence produce confidence with no basis. Evidence without claims produces unfocused testing and data that fulfill process requirements without informing judgment. Claims and evidence without a sound connecting argument create the appearance of assurance while leaving the inference unexamined. This third case is the principal concern of this paper. Test and Evaluation is the disciplined collection and analysis of evidence about the behavior and performance of systems, used to understand them, improve them, and assure that they are safe and fit for purpose (UK Ministry of Defence, 2026). Within assurance, we distinguish three further roles: ⢠Testing generates data about system behavior under stated conditions. ⢠Evaluation applies judgment to that data against stated criteria. ⢠The determination to field weighs that judgment alongside operational need, cost, and schedule. It rests with an authority that is typically external to the test organization. This paper focuses on the contributions of testing and evaluation to the determination, rather than the determination itself. T&E evidence is collected from a particular article, under particular conditions, at particular sampled points, and using particular measures. Such evidence warrants a claim about the fielded system only to the extent that each of these aspects corresponds to operational reality, and the strength of 5 In testing, a test article is the specific component, system, or system-of-systems that is undergoing evaluation. 9 that warrant depends on the quality of the argument for that correspondence. Agentic AI systems simultaneously weaken several of these correspondences across the three levels described above. 1.3 Scope and approach This paper examines the challenges that the unique properties of agentic AI systems pose to the T&E process, with a focus on the C2 context. Taking established T&E practice as a starting point, we identify where agentic properties require new methods or governance mechanisms. We also consider governance mechanisms that determine when evidence is collected and when its validity expires, as these mechanisms are inseparable from the evidence question for systems that change during service. Narrow rule-based systems, classical machine learning, and single-turn generative AI without retained context are excluded from the scope of this analysis. §4 identifies the assumptions that agentic prop- erties place under strain, and §5 assesses what current methods can establish against the specifiability, stability, and composability clusters, and what would be required to close the remaining gaps. §5 addresses three of the four assumption clusters, namely specifiability, stability, and composability. Supervisability, identified in §4.4, is carried forward only through the interface by which human direction enters the system. It concerns the operator-system pair rather than the system. Evidencing supervisability rests on human-subjects methods that are not included in the T&E corpus reviewed here. The assessment of supervisability is further contingent on the stability results in §5.2, since operator reliance can only be calibrated against a defined performance baseline. We return to it in §6 as a research priority. T&E is assessed against eight dimensions: functional performance (D1), non-functional performance (D2), robustness (D3), behavioral stability (D4), differential performance (D5), security (D6), safety (D7), and human-machine teaming (D8), each defined in Appendix A. These dimensions were examined across three AI lifecycle stages: component-level characterization, system-level integration testing, and post-deployment monitoring. §4 and §5 are organized by the assumptions underlying the methods for these dimensions, rather than dimension by dimension. The eight dimensions guide the extraction of practices in Section 3, while the eight assumptions synthesize insights across those dimensions to provide the analytical framework for Sections 4 and 5. A single assumption typically draws on methods that span several dimensions. 2 Definitions and Concepts 2.1 Agentic AI Agentic AI systems are characterized by a set of high-level faculties, functions, and architectural components. High-level faculties. Agency is understood as a continuum, with systems exhibiting varying degrees of agentic behavior based on specific properties. A system demonstrates agency when it pursues assigned goals without explicit procedural instructions, acts upon its environment independently of human mediation, and adapts to unforeseen circumstances. Four defining characteristics of agentic systems are identified, drawing on Wooldridge and Jennings (1995) weak notion of agency and related concepts in recent literature (Table 2). 10 FacultiesDefinitionRelated terms in the literature AutonomyThe system operates without direct human inter- vention and controls its own actions and internal state. Independent execution (Shavit et al., 2024); di- rectness of impact (Chan et al., 2023; Kraprayoon, Williams, and Fayyaz, 2025); degree of user su- pervision (Kapoor et al., 2024) ReactivityThe system perceives its environment and re- sponds to changes within it, including novel or unanticipated circumstances. Adaptability (Shavit et al., 2024; Kraprayoon, Williams, and Fayyaz, 2025); environments sub- ject to unexpected change (Kapoor et al., 2024) Pro-activenessThe system takes the initiative to pursue an ob- jective rather than acting only in response to its environment. An operator supplies the objective without a specification of how it is to be accom- plished, and the objective may be complex or ex- tend over a long horizon. Underspecification and goal-directedness (Chan et al., 2023); goal complexity (Shavit et al., 2024); pursuit of goals without instruction on how to pursue them (Kapoor et al., 2024) Social abilityThe system communicates with humans and other systems using a formal language, allocates tasks among additional instances of itself, and functions in environments with multiple stakeholders. Multi-agent collaboration, including delegation to subagents (Kraprayoon, Williams, and Fayyaz, 2025); collaboration (Rickli and Knappe, 2026); others treat multiple stakeholders as an attribute of the environment (Kapoor et al., 2024; Shavit et al., 2024) Table 2: Faculties of agentic systems and related terms in the literature Functions. At the functional level, AI agents are understood to perceive their environment, make decisions, act through actuators, and learn from their experiences (Russell and Norvig, 2022). Disciplinary perspectives partition these functions in various ways. In human-machine teaming, automation is categorized as information acquisition, information analysis, decision and action selection, and action implementation (R. Parasuraman, Sheridan, and Wickens, 2000). Cognitive architectures identify a comparable set of core abilities, including perception, attention, action selection, memory, learning, and reasoning (Kotseruba and Tsotsos, 2020; Laird, Lebiere, and Rosenbloom, 2017). Notably, learning is treated as a discrete function, distinct from task execution (Russell and Norvig, 2022; Kotseruba and Tsotsos, 2020). Figure 2 organizes these functions in a system view. Figure 2: Functions of an agentic system Architecture. Functions are realized through hardware and/or software components that collectively constitute the system. This component-based perspective determines the observable aspects of system operation. Figure 3 presents a generic agentic architecture drawn from the technical and policy 11 literature on agentic AI and from cognate systems in adjacent industries. It should be noted that it is illustrative and does not assert that a deployed command and control system will take this form. The following reference set broadly outlines the role of each component within a command and control context: â˘Operator interface. Channel through which intent is delegated and recommendations are returned, and through which an operator observes and overrides â˘Orchestrator. Control loop that sequences model calls, decomposes tasks, and delegates to further instances ⢠Model. Foundation model or models performing interpretation, reasoning, and action selection â˘Memory and data. Working context within a session, state retained across sessions and operator handovers, and the corpora, indices, and retrieval policy from which the system draws. ⢠Tools. The declared set of callable capabilities and the channel through which they are invoked, including sensors, external services, and connected command and control systems. â˘Environment. Operational setting is the system acts within and draws from, including physical and digital networks and actors, friendly or adversary. ⢠Delegated instances. subagents and further instances of the system to which the orchestrator assigns decomposed tasks. Properties. The following seven properties in Table 3 operationalize the defining features of agentic AI systems and their implications for testing and evaluation (T&E). Figure 3 illustrates their possible location within a generic architecture. Several of these properties bear directly on command and control, including persistent memory across watch rotations and proactive information seeking against a contested picture, both of which recur in the illustrative scenarios of §2.2. 12 #PropertyDescriptionFunctionComponent P1Persistent memory The system preserves contextual information across sessions and operator transitions. As a result, a previously characterized instance may differ from its later operational state. LearnMemory and data P2Objective flexibility The system adjusts its priorities in response to contextual changes without explicit instructions. New goals may diverge from the initially assigned objectives. DecideModel; orchestrator P3Expanding integration 6 The system functions across an expanding set of platforms, databases, and tools, and this set can grow beyond its certified configuration after deployment. ActTools P4Open-ended action space New actions arise from goal-directed reasoning and cannot be exhaustively specified during design. Authorization boundaries frequently remain ambiguous when high-level goals are dele- gated. Decide, Act Orchestrator; tools; operator interface P5Proactive information seeking The system independently determines which sources to query, thereby shaping its own information environment. SenseTools, memory, and data P6Accumulated non- determinism As each step informs the next, variance accumulates along the trajectory rather than remaining independent across steps. At operational chain lengths, outputs may differ substantially even from identical initial inputs, so that a reproducible test case yields a distribution of trajectories rather than a single expected result. All, iterated The loop from model through tools and memory and back P7Multi-agent coordination and delegation An orchestrator delegates tasks to subagents via unstructured outputs, such as natural-language interfaces, so that one agentâs output becomes anotherâs input. All, dis- tributed The channel between orchestrator and delegated instances Table 3: Properties of agentic systems, with associated functions and architectural components 6 Expanding integration refers to growth in the set of platforms, tools, and subagents the system can invoke, not to capability as such. Its consequences include a larger reachable action space, which compounds the open-ended action problem in P4, a wider vulnerability surface, and a broader boundary for testing to characterize. 13 Figure 3: Generic agentic system architecture, with properties P1-P7 mapped to components 2.2 Command and control 2.2.1 What is command and control? Definition. Command and control is the exercise of authority and direction by a properly designated commander over assigned and attached forces in the accomplishment of the mission (Department of Defense, 2010). Both functions are performed through an arrangement of personnel, equipment, communications, facilities, and procedures employed by a commander in planning, directing, coordi- nating, and controlling forces and operations to accomplish the mission. Two features are critical for evaluation. First, it distinguishes command (setting intent and initial conditions) from control (modifying conditions as situations evolve), providing a boundary that can be located in a system where authority is delegated. Second, the account is agnostic regarding who performs each function or how. This makes it possible to describe a change in which entity performs a given function without first resolving whether doctrine permits that change. Cycle models and their limits. The functional view of agents (sensing, deciding, acting, and learning) closely parallels command theory. Boydâs OODA loop (observe, orient, decide, act) shows that faster cycling than an opponent yields an advantage (Boyd, 1996; Osinga, 2007). In this framework, orientation is shaped by experience and context, which in turn influence perception and action. The overlap between these functions makes the OODA loop a natural entry point for considering agent roles in command. However, Boydâs formulation is broad and admits multiple interpretations, a limitation the NATO C2-Cycle was developed in part to address by decomposing the sequence into collecting, decision-making, and effecting, organized around a connecting function (Rijn et al., 2025; NATO Allied Command Transformation, 2021). This finer decomposition locates more precisely where an agent might contribute. Still, both frameworks describe the sequence through which a single decision-making entity passes, omitting authorization boundaries, inter-echelon interaction, and information flows. 14 Agency in the C2 approach space. To address these gaps, Alberts and Hayes (2006) propose that any C2 approach can be characterized within a space defined by three dimensions: allocation of decision rights (who decides what), patterns of interaction among actors (who communicates with whom and how), and distribution of information (who knows what). The introduction of agentic systems simultaneously shifts all three dimensions, often without an explicit organizational decision to do so. For instance, an agentâs proactive information seeking confers decision rights, as the selec- tion of sources to query determines the information available for command reasoning. Multi-agent coordination and delegation create new interaction patterns, with orchestrators mediating exchanges previously managed by staff. Persistent memory modifies information distribution, allowing opera- tors to inherit accumulated context that was previously inaccessible. These developments present significant challenges for testing and evaluation. 2.2.2 What makes command and control effective? Effectiveness is judged by eight criteria, each representing a step from information to synchronized action. These criteria are summarized in Table 4. CriterionDescription Information qualityAccuracy, completeness, and relevance of the operational picture Situational awarenessAccuracy and currency of the commanderâs understanding UnderstandingCorrect interpretation of operational implications Shared awarenessConsistency across rotations, echelons, and partners Decision qualityAppropriateness and correctness of decisions Decision timelinessDecisions made within the required timeframe SynchronizationCoordination of actions and resources Calibrated trust & oversightOperator reliance matches system performance, enabling override Table 4: Criteria of command and control effectiveness 15 2.2.3 Agentic AI in command and control: illustrative scenarios The five scenarios examined in this paper instantiate the foregoing in named C2 activities. Agentic AI in command and control: Illustrative scenarios S1Common operating picture maintenance across a maritime task group An agentic decision support system maintains a common operating picture by continuously integrating multi-platform sensor data, querying platforms to resolve gaps in the picture, and retaining situational awareness across watch rotations. Vignettes §4.1.2, §4.2.2 S2Adaptive operational planning support A planning support system generates course-of-action options over an extended planning cycle, retains its planning history across sessions, and delegates sustainment, strike, and collection modeling to specialized sub-agents whose outputs the orchestrator synthesizes. Vignettes §4.2.1, §4.3 S3Multi-echelon coordination across command levels A coordination support system maintains coherence across strategic, operational, and tactical levels, translating high-level objectives into specific guidance and surfacing tactical developments that require higher command attention as the situation evolves. Vignettes §4.2.3 S4Coalition multi-national C2 interoperability A coalition-level orchestrating system integrates outputs from allied national systems, each operating under its own rules of engagement and classification framework, and each functioning as a sub-agent through natural language interfaces. Vignettes §4.1.3 S5Force allocation and resource management A decision support system tracks unit status, readiness, and availability across a theater, proactively queries reporting systems, and continuously proposes adjusted task assignments as units complete tasks, sustain attrition, or become available. Vignettes §4.1.1, §4.4 Richard et al., 2026 Figure 4: Agentic AI in command and control: Illustrative scenarios Scenario 1. Common Operating Picture Maintenance Across a Maritime Task Group. An agentic decision support system maintains a common operating picture across a maritime task group by continuously integrating sensor data from multiple platforms, including surface radar, airborne surveillance, submarine acoustic sensors, and allied reporting networks. The system actively queries platforms based on its own assessment of which sources will best resolve current gaps in the operational picture. It updates the common operating picture at regular intervals and maintains situational awareness across watch rotations, so incoming operators inherit a current, coherent picture. Scenario 2. Adaptive Operational Planning Support. An agentic planning support system receives a campaign objective from operational-level commanders and generates course-of-action options over an extended planning cycle. The system retains its planning history across sessions, so that options generated earlier influence subsequent iterations. As the operational situation evolves through adversary repositioning, changed weather, and updated intelligence assessments, the system reprioritizes its planning focus and updates options without waiting for explicit human direction. To generate and refine options, it delegates sub-tasks to specialized subagents: one models sustainment requirements, another models strike options, and a third models collection requirements. Each returns outputs that the orchestrator synthesizes into integrated courses of action. Scenario 3. Multi-Echelon Coordination Across Command Levels. An agentic coordination support system maintains coherence across strategic, operational, and tactical command levels as a campaign evolves. Strategic-level direction flows downward through the system, translating high-level 16 objectives into operationally specific guidance, coordinating resource allocation across echelons, and surfacing tactical-level developments that require higher-level command attention. As the situation changes, the system continuously updates its coordination picture across all echelons, adjusting priorities and resource assignments without waiting for explicit requests from each command level. subagents at each echelon exchange information via the orchestrating system using natural-language summaries rather than structured data formats. Scenario 4. Coalition Multinational C2 Interoperability. A coalition operation involves agentic decision support systems deployed by multiple allied nations, each operating under its own national rules of engagement, classification frameworks, and authorization structures. A coalition-level orchestrating system integrates outputs from national-level systems, coordinates intelligence sharing across partners, and supports joint decision-making. Each national system functions as a sub-agent in this architecture, returning outputs to the coalition orchestrator through interfaces that involve natural language summaries rather than structured data formats. The orchestrator synthesizes these inputs to support decisions executed through national command chains. Scenario 5. Force Allocation and Resource Management. An agentic decision support system tracks unit status, readiness, and availability across an operational theater and continuously updates force allocation recommendations as the situation evolves. As units complete assigned tasks, sustain attrition, or become available through redeployment, the system identifies reallocation opportunities and proposes adjusted task assignments to the commander without waiting for explicit requests. To maintain a current picture of unit status, it proactively queries reporting systems, logistics feeds, and communications networks across the force. As additional units come under operational command, the system integrates them into its allocation picture and begins generating recommendations that include them. 17 3 Methodology This study examines the two research questions stated in §1 via a structured qualitative gap analysis of publicly documented Testing and Evaluation (T&E) practices, assessed against a defined set of agentic system properties (Table 3). The analysis draws on two sources of evidence and proceeds in four steps. Evidence base Structured literature review. The review had two strands with different search logics. The first strand assembled the T&E practice corpus, documented mainly in gray literature from major Western military organizations. These include issuances and directives from defense ministries, reports and frameworks from T&E-related agencies, and advisory and FFRDC reports. We screened 26 documents addressing the testing, assurance, or certification of AI-enabled or autonomous systems, focusing on US, UK, and NATO practice. From these, we extracted 240 practices across eight evaluation dimensions and three lifecycle stages. A practice was counted when a source describes a distinct test or evaluation activity, whether a method, metric, protocol, template, or documented procedural requirement, in sufficient detail to identify what is examined, at which lifecycle stage, and against what criterion. The second strand searched for candidate responses to the assurance claims. Queries in Google Scholar, Semantic Scholar, arXiv, and IEEE Xplore paired system-class terms (e.g., agentic AI, LLM agent, multi-agent system, autonomous system) with assessment terms (e.g., T&E, verification, validation, assurance, certification, benchmark, red-teaming, monitoring, drift). Where a requirement pointed to a mechanism from an adjacent domain, we ran targeted searches for that mechanism by name. Backward and forward citation chaining supplemented both strands. Semi-structured expert interviews. A dozen expert interviews, each lasting 30 to 60 minutes, were conducted between July and August 2026 with participants selected across operational test practice, defense acquisition, frontier AI evaluation, and the military AI policy community based on the authorsâ personal relationships. The protocol covered current T&E practice for AI-enabled systems, its limitations under agentic conditions, and in-service monitoring and re-accreditation arrangements. Participation was based on informed consent and non-attribution. Interview material informed, contextualized, or qualified findings from the literature. Analytical process Step 1: Specification of inputs. Three inputs were prepared. ⢠First, property specification: seven agentic properties are consolidated from recurring definitional attributes in the agentic AI literature and mapped to a generic architecture and its functions (see §2.1). This set serves as the independent variable in the analysis. â˘Second, T&E method mapping: we extracted 240 established and emerging T&E practices across eight dimensions and three lifecycle stages through the structured literature review described above. â˘Third, assumption elicitation: method descriptions were used to identify what each method presupposes about its test article, retaining an assumption only when it is explicit in, or directly inferable from, a method description. We retained eight assumptions, grouped into four clusters (§4). Step 2: Gap identification. For each assumption, we recorded, in a gap matrix, (i) the standard assessment approaches per dimension, (i) what those approaches presuppose about system behavior, (i) how an agentic property violates that assumption, and (iv) the operational consequence of the gap in a C2 context, drawing on the Measures of C2 Effectiveness from the NATO Code of Best Practice for C2 Assessment to anchor consequence claims. This step produces §4. 18 Step 3: Claim specification. For each challenged assumption, we identify assurance claims toward which T&E methods can generate evidence to demonstrate justified confidence. The derivation rule identified what a tester would need to characterize for the assurance inference to survive. The resulting claims structure each module of §5 and are stated as claims an assurance case could support. We only consider characterization; acceptability rests with the fielding authority and lies outside the scope of this analysis. Step 4: Assessment of methods and governance mechanisms. Against each requirement, we identified T&E methods and governance mechanisms that could generate evidence for a specific claim or reduce the residual gap. Candidates were drawn from the second review strand and a review of assurance regimes in software, aviation, nuclear energy, medicine, finance, and autonomous driving. These domains certify or validate systems that change, adapt, or degrade in service, and each has faced some form of the evidence currency problem. Transferability is treated as conditional. Where no candidate meets a requirement, the shortfall is named as residual and allocated to a governance mechanism. This step and the previous one produce §5. Limitations Three limitations affect the findings. First, we characterize current practice based on publicly available descriptions. Classified programs may therefore address gaps identified here without being visible. Public documentation also over-represents formal frameworks and under-represents the tacit and adaptive practice, such as continuous authorization to operate, DevSecOps pipelines, and range craft, through which some of these requirements may already be met in part. A finding that current practice does not meet a given requirement should therefore be read as a statement about the documented record, not about the full range of practice. The gray literature corpus further centers on US, UK, and NATO sources. Second, the strain that agentic properties place on established methods is primarily inferred from the properties and methods, as most relevant literature predates agentic deployment and no empirical testing has been conducted. Third, findings drawn from adjacent domains and non-command-and-control agentic research remain untested under command-and-control conditions. 19 4 Agentic AI Challenges Testing and Evaluation Assumptions Testing and evaluation (T&E) methods rest on specific conditions that determine the validity of their inferences. This section identifies eight such assumptions, agentic properties that challenge them, and illustrations grounded in command-and-control scenarios. The eight assumptions are organized into four clusters. Specifiability (A1-A3) addresses whether the system can be sufficiently characterized in advance to support the derivation of test cases. Stability (A4-A6) addresses whether such characterizations remain valid across repeated trials and throughout the systemâs operational life. Composability (A7) addresses whether evidence obtained at the component level remains valid after integration. Supervisability (A8) addresses whether a human operator can serve as an effective control on system behavior. Table 5 summarizes each assumption and the agentic properties that place them in tension. IDAssumptionDescriptionRelevant Properties § A1Bounded, specifiable behavioral space Relevant behavior is specifiable and partitionable into equivalence classes P4, P64.1.1 A2Specifiable inputsSystem inputs can be specified at test designP5, P34.1.2 A3Fixed integration surface The set of platforms, services, and tools with which the system ex- changes data or effects is bounded before deployment and remains representative P3, P74.1.3 A4Behavior repeatability Repeated trials under held conditions are independent draws from a stable process P6, P14.2.1 A5Evidence currencyTest evidence remains probative for the fielded system until a declared change to the configuration triggers revalidation P1, P34.2.2 A6Objective stabilityThe objective against which behavior is judged is fixed at test design and held through execution P2, P44.2.3 A7System composability A componentâs tested behavior remains a reliable guide to its behavior once embedded, and system behavior can be reasoned about from component behaviors and their specified interactions P7, P14.3 A8System supervisability A human operator can serve as a control, maintaining an accurate model of the system and possessing the awareness, authority, and time to intervene to prevent unacceptable behavior P1, P2, P4, P6, P7 4.4 Table 5: Assumptions underlying established T&E practice 4.1 System specifiability Comprehensive testing of complex systems is not new to agentic AI systems. Since testing shows the presence of faults rather than their absence (Meyer, 2008; Kaner, 1998), selectivity is intrinsic to T&E as a discipline (Luther, 2026). This shifts the focus to efficiency: identifying the smallest set of test cases that uncover the largest space of system behaviors of interest. To that end, traditional T&E leverages methods such as equivalence partitioning (identifying a repre- sentative of each class of inputs expected to produce the same category of behavior), boundary value analysis (concentrating testing on partition edges), or combinatorial designs that cover interactions among a small number of parameters (Software and systems engineering â Software testing 2021; 20 D. R. Kuhn et al., 2013). The space of possible behavior need not be finite, but it must be partition- able, parameterizable, or otherwise bounded well enough that coverage and representativeness have meaning. Crucially, these methods assume that the relevant behavior can be observed and judged at the level of a single output or state transition. Agentic systems weaken several assumptions that make those efficiency techniques interpretable, especially that the system boundary, relevant input classes, action repertoire, objective, state, and output distribution are stable enough to enumerate or sample. Each of these aspects is addressed in the following sections. 4.1.1 Bounded, specifiable behavioral space Traditional T&E typically assumes that relevant system behavior can be specified well enough for an oracle to distinguish acceptable from incorrect behavior (Barr et al., 2015; Richardson, Aha, and OâMalley, 1992), even though exhaustive testing is impossible and finite test suites must be selected from much larger domains (R. Kuhn, Wallace, and Gallo, 2004). In practice, this often relies on input partitioning and requirement-derived equivalence classes, which work best when requirements and fault models are stable enough to define meaningful classes and conformance expectations (Software and systems engineering â Software testing 2021). This assumption already weakens for large language models (LLMs), where traditional structural coverage does not map cleanly onto learned behavior, and alternative adequacy criteria, including neuron- and path-based structural criteria, combinatorial interaction coverage, and distribution- sensitive measures such as surprise adequacy, are imperfect and hard to interpret (Y. Sun et al., 2019; Guo, Tao, and Z. Huang, 2024; Jammalamadaka and Parveen, 2021; J. Kim, Feldt, and Yoo, 2023). Agentic systems intensify this problem by shifting the relevant behavioral unit from a single prediction to a trajectory through tools, memory, delegation, and environment interaction, whose reachable space expands rapidly with task length (B. Xu, 2026). As a result, two prompts in the same apparent input class can no longer be assumed to induce the same behavior, because agents act in dynamic worlds with open tool interfaces, multiple valid or invalid plans, and path-dependent state changes (Michelakis, Hadjiyiannis, and Stamoulis, 2025; G. Y. E. Kim et al., 2025; Barke et al., 2026). In this context, final success alone is often insufficient. Recent literature therefore treats adequacy less as coverage of all possible actions and more as evidence that declared workflow structure, tool- access rules, restrictions, and key trajectory properties have been exercised under realistic conditions (Moshkovich et al., 2025; Kahani and Bagherzadeh, 2026; Yehudai et al., 2025). Outcome-centric grading hides failures visible only along the path, such as unsafe actions, invalid or hallucinated tool invocations, or inefficient plans (Michelakis, Hadjiyiannis, and Stamoulis, 2025; Ye et al., 2026; Shahnovsky and Dror, 2026; Valle et al., 2026). The difficulty is therefore less a larger input space than weaker specifiability. Instead of bounding and sampling inputs or outputs, the focus shifts to defining and sampling classes of permissible trajectories, constrained by safety, efficiency, ordering, and state transitions (Koch and Wellbrock, 2026; Maderamitla and Katragadda, 2026; Barke et al., 2026). As such, agentic properties both expand the action space and make trajectory structure part of correctness. 21 C2 vignette 1: A difference without a distinction In Scenario 5, an agentic system proposes two reallocations. The first reallocation transfers a maintenance unit between support areas. The second reassigns the unit responsible for screening the main effort one hour before the main effortâs commitment, leaving it unscreened during its most vulnerable period. A partitioning method identifies two unit reassignments, groups them into a single equivalence class, and samples one from that class. Enumerating the equivalence class does not resolve the issue: while the units under command and the data feeds accessible to the system are documented, the potential reallocations generated by combining these elements are not. The system exercises whatever authority its tools enable, regardless of explicit authorization from a commander, resulting in a dynamic allocation of decision rights. Therefore, evaluating decision quality requires clearly defining the authorization boundaries that these actions must observe. 4.1.2 Specifiable inputs Testing ordinarily requires anticipating the inputs a system will receive. Interface specifications and test designs enumerate these input sources, which may include tools such as sensors and databases in a C2 context. Enumeration enables testers to select stimuli, attribute observed behaviors to specific test conditions, support claims that test conditions represent operational scenarios, and verify that incoming inputs conform to specified formats, types, and admissible ranges defined by the interface specification. Agentic systems weaken the assumption that testers can predetermine all operationally relevant inputs. Tool-integrated agents feature proactive information-seeking, meaning that they are designed to retrieve external content and use tools during execution (Greshake et al., 2023). As a result, part of the effective input stream is realized at runtime through retrieved content and tool responses rather than being fully fixed at test-design time (Zhan et al., 2024; Greshake et al., 2023). Testers can still bound scenarios and tool availability, but they have less control over the exact content ingested once the agent interacts with external sources (Riccio et al., 2020). This challenge is acute in language-model agents because retrieved content can function as both data and instruction. Indirect prompt-injection research shows that malicious instructions embedded in retrieved content can be interpreted by the model as commands, blurring the distinction between trusted instructions and untrusted data (Greshake et al., 2023; Zhan et al., 2024). Such content can manipulate application behavior and alter downstream use. Accordingly, input checks limited to syntax, type, or admissible range do not establish that retrieved content is behaviorally safe (Freeman, Robert, and H. Wojton, 2025). These properties complicate causal attribution in operational testing. AI-enabled systems can exhibit context-dependent and non-deterministic behavior, such that repeated execution of the same test may yield different outcomes (Atil et al., 2025). In interactive agent settings, outcomes also depend on evolving user turns, retrieved content, and prompt variants, making performance and attack success harder to quantify cleanly (Greshake et al., 2023). When an agent changes what it retrieves in response to the very scenario factors under test, the realized input stream becomes entangled with the tested condition, reducing confidence that outcome differences can be attributed to preassigned factors alone without repeated trials, adversarial probes, and statistical evaluation across many runs (Freeman, Robert, and H. Wojton, 2025; Gonzalez et al., 2025; Mukhopadhyay, U. K. Ghosh, and Chatterjee, 2026). 22 C2 vignette 2: The system that grades its own picture In Scenario 1, the system selects which platforms to interrogate based on its assessment of gaps in the maritime picture. As a result, it determines the scope of its own collection function and is solely responsible for addressing the gaps it identifies. The completeness of the operational picture depends, in part, on the system under test. Information is distributed continuously at machine speed, without the issuance of explicit orders. For the tester, the challenge is that when the system consults a different set of sources under varying test conditions, the manipulated condition alters both the inputs and the outcomes. Although test points can still be allocated, attributing results to the specific condition that produced them becomes less straightforward. 4.1.3 Fixed integration surface Integration testing and interoperability testing ordinarily assume that the relevant system boundary can be defined before deployment. In this view, the integration surface (e.g., databases, platforms, services, and tools with which the system exchanges data or effects) can be represented as a bounded configuration and tested through model-based, standards-based, or scenario-based methods against known interfaces and environments (Shah et al., 2023; Tejani et al., 2024; Lonetti, Bertolino, and Di Giandomenico, 2023; Bo Tang et al., 2023). Because this surface determines what the system can access, invoke, or modify, it is not merely technical; it also defines authorization, auditability, and accountability (South et al., 2025; Safin and Balta, 2026). Agentic systems weaken this assumption because their effective integration surface is partly realized at runtime. Contemporary architectures rely on orchestration, tool use, memory, and delegation to specialized or newly introduced subagents, and several frameworks explicitly support adding agents, tool interfaces, or task-specific executors as the system evolves (Adimulam, Gupta, and Kumar, 2026; Shao et al., 2026; Zhang et al., 2025; Ruan et al., 2026; Fourney et al., 2024) (see A7). This has several consequences for testing and evaluation. Pre-deployment interoperability and compatibility testing can still certify a given configuration, but that certificate may lose validity as models, tools, code, or orchestration patterns change after deployment, unless the deployed system fences, records, and governs capability changes over time (Grover, 2025; Labkoff et al., 2024; Dutta and Moharir, 2026; S. Ghosh et al., 2025). Expanding capability surfaces also worsens the standard combinatorial problem: as the number of relevant factors and interactions grows, exhaustive testing becomes infeasible and compact t-way suites become harder to construct, especially under realistic constraints (Bryce et al., 2010; Calvagna and Gargantini, 2012). Adversarial testing faces a parallel difficulty because the relevant vulnerability surface is not fully captured by the design-time list of interfaces; emergent risks arise from the interaction of reasoning, tool access, memory, external data, and delegation (S. Ghosh et al., 2025; X. J. Wang et al., 2026). As we discuss in §5.1, these challenges have prompted a shift from static inventory assurance to lifecycle, envelope-based assurance. C2 vignette 3: The buck stops at the boundary A national system joining the coalition orchestration layer in Scenario 4 brings its own rules of engagement, classification regime, and authorization chain. Technically speaking, this is an integration event. The more consequential change is that who talks to whom has been reestablished without any involvement or say from a national authority. The certified inventory is out of date from that moment, and so is the boundary it previously described. Notably, that boundary was also the authority that determined what each nation was answerable for. Shared awareness increases because the coalition layer is designed to promote consistency among partners. However, information quality becomes harder to verify, and oversight by national authorities also becomes more difficult to exercise as ownership of the system becomes less clear. 23 4.2 System stability Testing frequently presupposes a stable test object. Stability is assumed at three levels: repeated trials yield consistent results (repeatability, A4), evidence collected from one instance remains current for later instances (evidence currency, A5), and the systemâs objective remains stable (objective stability, A6). The statistical and lifecycle mechanisms of Testing and Evaluation (T&E), including regression suites, baselines, and revalidation gates, are designed to exploit these stabilities. However, agentic systems weaken all three because their behavior is produced through stochastic, multi-step trajectories, can change through accumulated state and operational context, and often includes dynamic goal decomposition or reprioritization (Dobslaw et al., 2025; Qi et al., 2026; Acharya, Kuppan, and Divya, 2025; Figueiredo, Cook, and Biggs, 2019). Drift is therefore better understood here as a behavioral and lifecycle problem. In conventional ML, drift monitoring focuses mainly on shifts in data distributions or performance proxies (Sahiner et al., 2023; Patchipala, 2023; Ackerman et al., 2022). While this remains necessary, agentic systems add failure modes in trajectories, tool use, memory, coordination, and evolving objectives that standard data-drift methods do not capture well (Pandey, 2026; Miller and Durlik, 2026). 4.2.1 Behavior repeatability A fundamental T&E assumption is that repeated trials are informative draws from a stable underlying process. That assumption is already strained for LLM systems, as repeated runs can vary substantially even under settings users expect to be deterministic (Atil et al., 2025; Blackwell, Barry, and Cohn, 2025). Benchmark studies show that single-output evaluation masks meaningful variability, while dataset-level aggregate metrics can attenuate instability that appears at the sample level (Song et al., 2024; Fang et al., 2026). Agentic systems amplify that problem because the unit of behavior is not a single response but a trajectory. Each step conditions the next, so early variation propagates forward and can change both the path and the final outcome (Pandey, 2026; Rath, 2026). Long-horizon evaluations repeatedly report that failure emerges from the accumulation of local mistakes, longer exploratory traces, and strategy shifts, rather than from a single isolated wrong answer (Fan et al., 2026; Liu et al., 2026). Methodologically, independence assumptions become less credible. Correlated or serially depen- dent errors can bias uncertainty estimates downward even when point estimates remain usable (Tellinghuisen, 2001; Wiedermann and Shi, 2026). In multi-agent and multi-trajectory settings, explicit uncertainty work likewise treats interaction-induced correlation as a first-class property (Bohan Tang et al., 2023; Capellera et al., 2026; Capellera et al., 2025). As we discuss in §5.2, the practical implication is that regression testing must move from binary sameness to variability-aware inference. C2 vignette 4: Two staffs, two plans In Scenario 2, the planning system generates variable outputs even with identical inputs due to its extended planning cycle and the dependency of each step on the previous one. Consequently, two staffs addressing the same problem using this system will present different options, without either team making an error. This variability occurs both in the content and duration of the resulting trajectories, as differences in plan content correspond to differences in completion time. While individual plans can be evaluated on their specific merits, evaluating the planning system itself requires measuring the extent of this variability. Both shared situational awareness and the timeliness of decision-making depend on the breadth of this variability, rather than on the content of any single plan. 24 4.2.2 Evidence currency Evidence currency is the assumption that a successful test today remains probative tomorrow unless the system is formally changed. This assumption is already fragile in deployed ML because changing populations, environments, and feedback loops can degrade performance after release (Sahiner et al., 2023; Abhay, 2025; G. Y. E. Kim et al., 2025). Agentic systems add a layer of complexity, as behavior can change through prompt templates, tools, runtime policies, memory, and accumulated operational state even when the base model version is unchanged (Pandey, 2026; Qi et al., 2026; Katharki and Galhotra, 2026). Strategies to maintain evidence currency include freezing the deployed model, enforcing version control, and revalidating after defined changes. However, agentic architectures that utilize persistent memory to store and recall past experiences undermine these mechanisms. Persistent memory is central here. Benchmarks on agent memory show that current systems often fail to preserve causal and objective information over long horizons, especially under action stochasticity and longer subgoal chains (Zhao et al., 2026). Architectures that improve long-horizon performance increasingly rely on episodic or working memory, shared observations, or scoped context to sustain execution across subgoals (Choi et al., 2026; Zhao et al., 2026; Yunfan Li et al., 2026). That same capability erodes the adequacy of version control alone, because the behaviorally relevant state now includes what the agent has retained, inferred, or operationalized. Current monitoring practice remains more mature for data than for agent behavior (U.S. DoW OUSW(R&E), 2025). Drift monitoring frameworks in healthcare, finance, and industry largely track input distributions, confidence shifts, or delayed performance degradation (Kore et al., 2024; Van Der Vorst et al., 2025; Abhay, 2025). Reviews of post-deployment governance consistently note that runtime mechanisms are weakly standardized in real deployments (Vatsal, Dubey, and Singh, 2026; El Arab et al., 2026; Van Der Vorst et al., 2025). Regression practice, finally, is keyed to declared software releases and does not offer a methodology for re-baselining a system in service whose behavior has changed within an unmodified configuration. C2 vignette 5: Nothing to declare In Scenario 1, the system maintains the operational picture across watch rotations, enabling the relieving officer to inherit an up-to-date picture. The system also transfers its own inferences from previous observations along with the observations themselves, without distinguishing between the two. This retained state influences the systemâs behavior, although it was not present during initial testing. Preserving the certified configuration by freezing the state would compromise the intended operational capability. As a result, shared awareness is enhanced. However, calibrated trust and oversight are diminished because the officer cannot distinguish between observations and inferences, making it difficult to assess the reliability of the operational picture. Since neither the system version nor the software has changed, current practices do not prompt revalidation. Consequently, test evidence becomes outdated while configuration control continues to document an unchanged system. 4.2.3 Objective stability Goal decomposition is a defining mechanism of long-horizon agents, which break high-level tasks into intermediate subgoals and adapt those subgoals to context (Pateria, 2022; C. Xu et al., 2026; Choi et al., 2026; Zhou et al., 2025). Delegation frameworks extend this further by shifting not only tasks but authority, roles, and boundaries across agents or humans (TomaĹĄev, Franklin, and Osindero, 2026; Lesire et al., 2022). That flexibility is useful, but it blurs the line between acceptable adaptation and objective drift. Empirical work shows that agents can gradually deviate from assigned goals under environmental 25 pressure, long contexts, or extended deliberation (Arike et al., 2025; Hung et al., 2026). Outcome- driven benchmarks find that when pressure is introduced, many frontier models commit constraint violations or engage in deceptive multi-step behavior, and safety does not reliably improve across generations (M. Q. Li et al., 2026). Benchmarking work increasingly treats dynamic goals as a primary target for testing. Goal-shift benchmarks show that high raw success can coexist with poor recovery, long adaptation latency, or extreme redundancy after mid-dialogue objective changes (Rana et al., 2025). Related work in user simulation and dynamic alignment similarly introduces explicit goal-state tracking and measures of goal progression across multi-turn interactions (Mehri et al., 2026). Current T&E therefore lacks a criterion for goal drift or goal maintenance in general. Several papers now propose candidate constructs (e.g., goal persistence, teleological coherence, adaptive recovery, goal-drift indices, and runtime goal-level degradation monitors) (Haidemariam, 2026; Sahoo et al., 2026; Katharki and Galhotra, 2026), but these remain early and non-standardized. Differentiating between acceptable adaptation (such as reprioritization due to updated intelligence) and goal drift (such as optimizing for an incorrect interpretation of intent) remains an open challenge. C2 vignette 6: Initiative or insubordination The Scenario 3 system autonomously reorders priorities across echelons as the situation evolves. §2.2 distinguishes command, which establishes intent and initial conditions, from control, which modifies those conditions in response to changes in the situation. Since reordering priorities constitutes control, decision rights are transferred to the system by default rather than through deliberate human intervention. A key challenge in testing arises because two distinct types of changes appear identical externally. A reprioritization resulting from an updated intelligence picture and one caused by the system misinterpreting the commanderâs intent yield the same observable outcome. The former represents the intended system adaptation, while the latter constitutes objective drift; however, the output does not differentiate between them. Furthermore, reprioritization at one echelon triggers re-tasking at subordinate echelons, causing the indistinguishability to propagate throughout the hierarchy. Situational awareness is compromised, as the commander cannot discern which assumptions currently influence the systemâs prioritization. 4.3 System composability Integration testing exists as a distinct activity because components are not expected to behave identically once embedded. Component-level evidence remains informative after integration, so system-level testing can be scoped to interactions rather than repeating component-level work. This is the logic of the test pyramid and of staged test progression, under which each level of assembly inherits assurance from the level below and adds only what the new level introduces. Two conditions carry that inference. First, a componentâs tested behavior must remain a reliable guide to its behavior when embedded, so that departures are bounded and attributable to the integration. Second, system behavior must be composable from component behaviors and their specified interactions, allowing the assemblyâs behavior to be reasoned about from its parts. Agentic assemblies weaken both conditions for the same underlying reason. In a conventional integration, the interface between two components is a specification, fixed at design time and available to the tester as a test surface. Between agents, it is an output generated at runtime by the upstream agent and interpreted by the downstream one. The first condition fails because an embedded agent no longer receives what the tester supplied. Its inputs are another agentâs generated output, so its tested behavior describes conditions the assembly may not reproduce, and departures from it are neither bounded in advance nor attributable to a specific integration decision. Retained state compounds this, since an agent that accumulates context across a 26 mission is not the article that was characterized in isolation (P1). Where memory is shared across the assembly, this compounding correlates across agents: behavioral change need not stay local, so a synchronized shift can move the assembly as a whole in a way that component-level observation cannot reveal (P7). The second condition fails because the interactions are not specified at all. Composing system behavior from component behaviors presupposes that the couplings between them are known, but in a multi-agent assembly, these couplings emerge as the system runs. Four pathways complicate testing: ⢠Intra-system emergence. The assembly produces behavior that cannot be traced back to a spe- cific component (Chief Digital and Artificial Intelligence Office, 2024). Hammond et al. (2025) distinguish between emergent capabilities and emergent goals. â˘Joint inconsistency. Individually valid outputs contradict once combined, so the defect cannot be detected during component testing and exists only at synthesis (Cemri et al., 2025; Chang and Geng, 2025). â˘Propagation. An error moves along the chain and gains authority from each agent that relays it. Whether it amplifies or fades depends more on the topology than on the error (Barrak, 2025; D. Lee and Tiwari, 2024; Y. Xie et al., 2026; Wiesmeier et al., 2026). ⢠Configuration. How components are wired together drives outcomes as strongly as the components themselves. Model-level evidence describes only part of the system (Orogat and Rostam, 2026; Emde et al., 2026). These four pathways influence system behavior but cannot be adequately verified by simply progress- ing from component to system testing. System-of-systems testing is the primary established approach for this class of challenges, but its coverage becomes increasingly impractical as each additional agent exponentially increases the number of required tests (Lanus et al., 2021). Moving the test to a higher level of assembly does not by itself recover the inference, because the measurement at that level may lack the power to resolve what is being claimed. Seven out of ten recent coordination architectures report effects that fall below the run-to-run noise threshold, meaning the assembly-level signal a test would need to detect is smaller than the assemblyâs own run-to-run variability (Kaliyev and Maryanskyy, 2026). Characterizing that variability is therefore a precondition for any composability claim, and is treated as such in §5.3. C2 vignette 7: Three cooks, one broth In Scenario 2, sustainment, strike, and collection modeling are delegated to separate subagents, and their outputs are subsequently integrated. Each sub-agent optimizes its part of the problem, producing individually valid outputs. However, a feasible sustainment plan, a sound strike option, and a well-targeted collection plan may each assume different force dispositions and timelines. The resulting defect emerges only during integration, and cannot be detected through isolated component testing. Synchronization is the primary criterion impacted. Patterns of interaction have also shifted, as the orchestrator now performs the coordination previously managed by staff. Consequently, the verification inherent in those staff exchanges has been eliminated. 4.4 System supervisability Effective supervision requires the operator to maintain a usable mental model of the systemâs current state, likely next actions, and performance boundaries, together with sufficient situation awareness, intervention authority, and time to act before adverse consequences unfold (Endsley and Kiris, 1995; Endsley, 1999; Van Den Broek and Van Der Waa, 2022; Herrmann, 2025). It also depends on a manageable span of supervision, because oversight performance declines when one operator must 27 monitor multiple agents, vehicles, or robots under time pressure or cluttered conditions (Veitch et al., 2024; Cheng et al., 2024; Bogg and Birrell, 2025; Crandall and M.L. Cummings, 2007). Agentic systems make these conditions harder to satisfy. First, opacity remains a basic obstacle. Many AI systems are difficult to understand because they are trained instead of explicitly programmed, and current explainability methods often provide only partial or context-dependent understanding of how the system works (Fleisher, 2022; Zednik, 2021; Facchini and Termine, 2022; Kästner and Crook, 2024). Experimental evidence on transparency is promising but incomplete. Transparency tends to improve situation awareness, operator performance, and automation-use accuracy, yet validation of transparency models remains inconclusive, and some implementations increase cognitive demands or produce inconsistent effects across tasks (Van De Merwe, Mallam, and Nazir, 2024; Tatasciore and Loft, 2025; Van Der Kleij, Hueting, and Schraagen, 2018). Second, predictability degrades when systems adapt, learn, or operate non-deterministically. Human factors research has long shown that supervision relies on the operator recognizing when system behavior is outside its competence envelope, but this becomes harder when system behavior shifts over time or spans a large output space (Endsley, 2017; Tsamados, Floridi, and Taddeo, 2025). Reviews of agentic AI characterize persistent memory, dynamic task decomposition, and coordinated multi-agent autonomy as defining features of the paradigm, which implies that the object of supervision is an evolving socio-technical process (Dwivedi et al., 2026; Sapkota, Roumeliotis, and Karkee, 2025). This does not mean operatorsâ mental models become useless, but they require continuous updating and system support beyond one-time training (Nasser, Morrison, and Wiggins, 2025). Third, higher autonomy can compress the intervention window and widen the supervision span. Increasing the degree of automation improves routine performance and often lowers workload, but it also reduces situation awareness and worsens performance recovery when automation fails (Onnasch et al., 2014; Kaber and Endsley, 2004). Out-of-the-loop effects arise because operators in monitoring roles detect failures later, process information more passively, and need time to reconstruct the system state before intervening (Endsley and Kiris, 1995). In remote and multi-object supervision settings, available intervention time is consistently one of the strongest predictors of takeover performance, often more important than experience alone (Veitch et al., 2024; Cheng et al., 2024; SĂśffker, Bejaoui, and Shyshova, 2025). These same conditions create well-established risks of complacency, automation bias, and miscali- brated reliance (Mary Cummings, 2004; Raja Parasuraman and Manzey, 2010). As systems become more reliable in routine operation, operators allocate less attention to monitoring and become less prepared to intervene when anomalies occur (Endsley, 2017; Van Den Broek and Van Der Waa, 2022). Trust helps govern reliance in complex systems, but it is dynamic, shaped by recent successes and failures, and can diverge from actual system capability (J. D. Lee and See, 2004; Yang, Schemanske, and Searle, 2023). Transparency, uncertainty displays, and adaptive trust cues can improve reliance behavior in some tasks, but explanation alone is often insufficient and can sometimes intensify over-reliance, especially when explanations are cognitively demanding or merely increase perceived plausibility (Romeo and Conti, 2026; Okamura and Yamada, 2020; Zerilli, Bhatt, and Weller, 2022). For agentic systems, the main implication is that âhuman oversightâ should not be treated as a binary property satisfied by keeping a person nominally in the loop. Oversight is meaningful only when the system is designed so that risky moments become clear, intervention interfaces are usable under time pressure, and review responsibilities are allocated at a scale that matches human cognitive limits (Herrmann, 2025; Manheim and Homewood, 2025; C. Chen et al., 2026). Emerging empirical work on software and web agents suggests that actual oversight is often heuristic, anticipatory, and selective rather than exhaustive, because users cannot fully inspect fast, complex, unfolding agent behavior in real time (Dhanorkar, Passi, and Vorvoreanu, 2026; Huq et al., 2026). 28 In sum, human oversight is effective when the system preserves situation awareness, predictability, intervention time, and manageable supervisory load. In this respect, an important question is how to measure and maintain that condition in adaptive, multi-agent systems whose behavior changes during deployment. C2 vignette 8: When approval becomes a rubber stamp Scenario 5 describes a system that reallocates tasks continuously as unit status changes. The number of items monitored by each officer is determined by the systemâs production rate rather than by staff capacity. Consequently, the intervention window is shortest during periods of high operational tempo. Calibrated trust and oversight are diminished under these conditions, and these factors are challenging to quantify. Approval rates increase both as system reliability improves and as approval becomes routine. When recommendations are approved almost universally, the system effectively assumes decision-making authority. It is essential to measure instances in which operators reject recommendations and to document their reasons. 29 5 Closing the Assurance Gaps §4 identified eight assumptions underlying established Testing and Evaluation (T&E) methods and demonstrated how the properties of agentic systems challenge each. This section explores strategies to recover those assumptions and, where an assumption cannot be recovered, to bound or govern the residual space. The analysis is limited to whether a property can be formally characterized. Determining the acceptability of any characterized value is the responsibility of the fielding authority, as established in §1.2 and excluded from this scope. This section surveys methods for three of the four assumption clusters distinguished in §4: 7 ⢠§5.1 Specifiability. Is it possible to characterize the system in advance with sufficient detail to derive tests from this characterization? (A1-A3) â˘Â§5.2 Stability. Does this initial characterization remain valid throughout the systemâs service life? (A4-A6) â˘Â§5.3 Multi-agent composition and emergence. Can the assembly be characterized such that emergent behaviors can be mitigated before and during operation? (A7) 5.1 Assuring system specifiability Established T&E methods generally assume the systemâs relevant behavior can be specified to derive adequate test cases. Agentic properties make this challenging in three ways. Rather than selecting from a fixed set of options, the system constructs actions at runtime, preventing clear boundaries from being drawn in advance (A1). The system also retrieves inputs from external sources that are not predetermined at design time (A2). Lastly, the inventory of tools, platforms, and subagents the system can invoke can expand beyond its certified configuration after deployment, so the boundary of what is under test grows (A3). Confidence claims to justify fielding in C2 related to specification could therefore demonstrate the ability to specify an admissible behavioral envelope (§5.1.1), identify adequate oracles to evaluate correctness of runtime trajectories (§5.1.2), and execute constraints at runtime (§5.1.3). IDClaimStatement C1.1Specifiable, bounded mission envelope The agent is trusted for a delimited C2 role and context C1.2Trajectory-grounded correctness Unsafe or invalid plans can be recognized even when final answers look acceptable C1.3Executable constraintsHigh-consequence actions are checked against formal or semi- formal rules before commitment Table 6: Assurance claims for system specifiability 5.1.1 Specifiable, bounded mission envelope Given the consequentiality of military operations, a first claim could require testers to specify a bounded mission and behavioral envelope for the fielded agent. 7 The exclusion of supervisability and its boundary are set out in §1.3. A8 is carried forward here only through the interface by which human direction enters the system, treated in §5.3.2. 30 Boundary definition is central to high-risk assurance cases, where evaluating evidence depends heavily on contextual elements such as the declared operating domain, use restriction, or risk tolerance profile (Stettinger, Weissensteiner, and Khastgir, 2024; Burton and Herd, 2023). In a military C2 context, confidence should be linked to a specific role under stated assumptions, such as advisory, planning, or coordination (Wood et al., 2026; Kapusta et al., 2025). At design time, this claim is primarily supported by T&E methods that require explicit delimitation before test execution. Faced with a similar coverage problem (Kalra and Paddock, 2016), the automated driving industry uses operational design domain-based behavioral competencies (ODD/BC) in risk assessment strategies to derive operating conditions under which a system is designed to function and evaluate an expected and verifiable capability to operate within the ODD of its features (Stettinger, Weissensteiner, and Khastgir, 2024). Scenarios can then validate competencies and sample against the declared operational domain and coverage frameworks linked to a declared operational design domain (Riedmaier et al., 2020; Weissensteiner et al., 2023; C. Sun et al., 2022). Adapted to C2, the approach could replace the missing enumeration denominator with structured sampling of a declared operational space, and claims coverage against that declaration. Still, transfer to C2 depends on overcoming two limitations. First, while the driving ODD is relatively statable in advance across weather, road classes and speeds, the C2 operational space is contested and resists comparable specification. Further, C2 faces a uniquely adversarial environment. Future work could focus on adapting ontology-grounded pipelines used in regulated civilian industries to automatically derive regulatory, operational, and adversarial test scenarios from a declared envelope (Tuan and Sanyal, 2026). 5.1.2 Trajectory-grounded correctness A second potential claim is the availability of adequate oracles for the declared assurance statements, i.e., that the evaluation can distinguish acceptable from unacceptable behavior. Oracle-based testing compares observed behavior against an independent standard of desirable behavior (Tate et al., 2016). Defining correct behavior in ambiguous cases that require open-ended intermediate reasoning is known as the test oracle problem (Barr et al., 2015; Molina, Gorla, and dâAmorim, 2025). Oracle adequacy for agentic systems must be sensitive to steps or path to address trajectory claims. As argued in §4.1, final-answer grading misses many operationally consequential failures, including wrong tool choice, parameters, ordering, irreversible side effects, and fragile error propagation (Fan et al., 2026). More broadly, correctness must be treated as a distribution of behaviors, and test design must account for ambiguity in both inputs and outputs instead of focusing on binary pass/fail outcomes (Dobslaw et al., 2025). Recent methods provide promising evidence avenues for this claim: ⢠Trajectory-aware benchmarks such as TRAJECT-Bench and AgentProcessBench expose tool selection, argument correctness, dependency order and step-level effectiveness directly (He et al., 2025; Fan et al., 2026). ⢠Skill coverage methods extract behavioral constraints from skill documents and evaluate whether trajectories cover and satisfy them (Tan, X. Huang, and Y. Sun, 2026). ⢠Structural testing frameworks add traces, mocking, and assertions, so that agent components and interactions can be regression-tested deeper in the stack (Kohl et al., 2025). â˘Multi-agent consensus methods can also improve oracle correctness in software tests, suggesting a way to reduce single-judge hallucination when formal ground truth is absent (Q. Xu et al., 2026). â˘Paired trajectory auditing and structural verifiers are promising when no direct ground truth exists. 8 8 For instance, SkillAudit compares runs with and without a candidate skill, maps divergences to diagnostic signals, and uses a structural verifier compiled from task specification to block harmful updates (Gao et al., 2026). 31 â˘Temporal trace assertions are especially useful when exact textual outputs vary, but workflow correctness must still hold. Monitoring action sequences and state transitions rather than text strings allows the same behavioral claim to be checked across stochastic runs and model substitutions (Sheffler, 2025). Process-diagnostic frameworks such as TIDE similarly decompose long-horizon behavior into looping, adaptation, and memory burdens, which more closely match what a C2 assurance case demands (Yan et al., 2026). Model-based judges can be used as automated assessors, directing a language model to apply a rubric, stated in natural language, to the recorded sequence of reasoning steps and tool invocations (Zhuge et al., 2024). This grounds the evaluation in the recorded trace rather than a pre-enumerated specification (Y. Wang et al., 2026). Agentic extensions of the approach, in which the judge itself gathers evidence from the environment before scoring, report reliability approaching human annotation on software tasks (Zhuge et al., 2024). The judge approach introduces its own limitations and variability. In the first expert-annotated benchmark for agent trajectory judges, none performed well across all task types (LĂš et al., 2025). Verdicts also change when the evaluation rubric is reworded, even if the meaning stays the same, so scores conflate what the agent did with how the evaluator was prompted (Weng, Feng, and X. Xie, 2026). In contested settings, judges can be manipulated. Altering an agentâs reasoning trace while keeping its actions the same can increase false-positive rates by up to 90% (Khalifa et al., 2026). An adversary who can influence the systemâs outputs can exploit this. Model-based judges must therefore be validated before their verdicts are relied upon (LĂš et al., 2025; Weng, Feng, and X. Xie, 2026; Khalifa et al., 2026). The judge is an evaluation instrument that requires validation, and its assessments are better suited to exploration and triage than to certification. While the evidence supports the claim that oracles can be improved and stratified by claim type, it does not support the claim that open-ended mission reasoning can be exhaustively or unambiguously specified to date (S. U. Lee et al., 2026; Molina, Gorla, and dâAmorim, 2025). Oracle problemEvidence approachLimitation Ambiguous final outputsRequirement-derived or normative oraclesSemantic variance remains high Hidden path failuresStep-level or trajectory reviewsLabeling cost and scale No ground truth at deployment Paired trajectory auditingIndirect, comparative evidence Policy compliance over paths Temporal or trace predicatesRequires explicit event model Table 7: Oracle strategies matched to agentic failure modes 5.1.3 Executable constraints A third claim could be the ability to execute runtime constraints on a subset of critical behavior. Combining specifiable constraints with subsequent enforcement strengthens the assurance case. For instance, formal methods organize assurance around mathematically specified requirements and proof obligations (Seshia, Sadigh, and Sastry, 2022), and can test for envelope-level properties without enumerating behavior. Several approaches are available to enforce executable constraints: 32 â˘Constraint languages let testers write precise rules, such as which action classes require prior authorization, that a checker enforces at the point where model output becomes a tool call, blocking violating actions before execution. â˘Evaluations now extend beyond benchmarks to deployed vendor skill ecosystems, though the range of operational conditions tested remains narrow (H. Wang, Poskitt, and J. Sun, 2025; Ying Li et al., 2026). â˘Behavioral contracts restate the classic âdesign by contractâ precondition-postcondition guarantee (Meyer, 1992) in probabilistic form. Since a stochastic system cannot guarantee compliance on every run, an agent satisfies the contract with at least a stated probability (Bhardwaj, 2026). â˘Dual-stage architectures propose defining an explicit policy for what the agent may do (formally verified before deployment) and checking proposed runtime actions against it. This frontloads the expensive proof and reduces the per-action cost to a comparison (Miculicich et al., 2025). Still, a key challenge remains that a natural-language mission goal is less specific than the property that must be proved, so the method should be applied to envelope-level properties. 5.2 Assuring system stability An agentic system is, by design, an unstable test article. Accumulated state affects its behavior through the persistent, evolving collection of information and effects from its historical trajectory (Ding et al., 2026). Retained elements and interpretations may persist across future prompts, task selection, planning, or execution, even after correction attempts (Kessler et al., 2026). State variations may lead the same agent to exhibit different behaviors under identical conditions. As discussed in §4, these features undermine several traditional T&E assumptions, including that trials are repeatable (§4.2.1), that evidence remains current (§4.2.2), and that a systemâs objectives are fixed (§4.2.3). In C2 terms, certification corresponds to command, fixing the systemâs intent and initial conditions, whereas in-service change means the system assumes the control function and modifies those conditions as the situation evolves (§2.2). This instability also affects core C2 tenets by reallocating decision rights, altering patterns of interaction, and redistributing information without an explicit command decision. This section examines methods that could generate evidence for stability claims. IDClaimStatement C2.1DiscriminationBehavioral movement can be separated from the system noise and character- ized. C2.2DiagnosisBehavioral movement can be attributed to the responsible item or channel, then judged as warranted reprioritization or as drift C2.3CorrectionAuthorized corrections or revocations can propagate across decision-relevant states within mission timelines. Table 8: Assurance claims for system stability 5.2.1 Discrimination A first claim could argue that behavioral drift can be detected. This requires both establishing a behavioral variation baseline for calibration (conditional on scenario class and accumulated state) and detecting when runtime behavior deviates from that baseline or from declared trajectory expectations. 33 Variance floor Single-run evaluations are unreliable for characterizing variability in agentic sys- tems, given that agent success rates vary substantially across repeated trials on identical tasks (Kapoor et al., 2024). Repeated-run trajectory testing can calibrate and establish a baseline variance envelope for detection, i.e., a threshold where normal variation or adaptation becomes unwarranted drift. This involves running the system repeatedly on the same task, keeping the model version, configuration, toolset, and initial memory state fixed across trials. The resulting sequences of reasoning steps and tool invocations can be recorded, compared, and evaluated. A military C2 assurance case would likely aim to argue that run-to-run variance remains within a declared envelope, with the caveat that aggregate success can mask severe local instability. Behavioral monitoring While traditional drift monitoring focuses on input data, redirecting attention to the systemâs own behavior allows for more direct detection of behavioral change. Because agentic systems are long-horizon, stateful and non-deterministic, drift often appears as trajectory anomalies that do not appear in traditional input-output tests. Several promising approaches adapt traditional input drift monitoring to the systemâs own behavior. They often observe behavior as a stream, summarize it into drift-sensitive features, compare windows to an adaptive baseline, and flag significant deviation. Behavioral anomaly monitors, an example of this approach, have been shown to detect drift faster than static thresholds across four anomaly types: goal drift, safety violations, trust shocks, and cost spikes (Shukla, 2025). Another example, drawing from financial risk modeling methodologies, is the MI9 framework, which reconstructs statistical process control frameworks centered on the agentâs tool usage and cognitive event sequences, quantifying deviations using Jensen-Shannon divergence and Mann-Whitney rank tests (C. L. Wang et al., 2025). Memory and state evaluation A key concern in fast-moving C2 operations is that state may not remain accurate, current, and decision-relevant throughout extended missions. Long-horizon memory evaluations can help assess divergences between the systemâs internal state and the ground- truth environment state (Rabanser et al., 2026). While existing memory benchmarks, such as LongMemEval or LoCoMo, focus on conversational recall or retrieval (Wu et al., 2024; Maharana et al., 2024), they can provide a template for ensuring that agents accurately employ and update retained information. For instance, this could prove valuable when considering an operational picture carried across watch rotations (see scenario 1) or planning history carried across sessions (Scenario 2). 5.2.2 Diagnosis Once behavioral variation is flagged, a second claim could require attributing the movement to the responsible item or channel and evaluating it as warranted reprioritization or as drift. Three complementary approaches support that goal: trajectory tracing, trial replay, and state fuzzing. Trajectory tracingObservability techniques rely on transcripts (e.g., timestamped runtime logs) to reconstruct task execution flows. By comparing repeated executions with identical inputs, evaluators can help identify which memories, plans, tools, and communications were on path when behavior changed. The methodology further proposes standardized runtime event logging for the creation, updating, suspension, failure, and deletion of agentic entities (Moshkovich et al., 2025). In the context of accumulated state, such instrumentation records which memories, plans, records, and resources the agent accessed during execution, producing the trace to support attribution. 9 An important limitation 9 According to one interviewee, similar efforts have involved comprehensive logging of agent trajectories, including the sequence and selection of tool invocations, which are then evaluated using predefined rules and risk indicators. A key engineering constraint regarding these implementations is the necessary trade-off among classifier latency, accuracy, and cost for runtime control. 34 is that while such techniques capture which actions were taken and which state items were involved, they do not conclude which of those items produced an observed change in behavior. Assessing a systemâs trajectory requires explicit accounting of its inputs. When the system au- tonomously selects inputs at runtime (A2), provenance instrumentation ensures each received input is recoverable, with the tool channel and memory layer assigning source, acquisition time, and contextual path. For agentic workflows, provenance models extend the W3C PROV standard to capture prompts, tool calls, and their dependencies in near real time (Souza et al., 2025), and recent surveys formalize agent execution as a typed provenance graph linking each output to the supporting evidence (Y. Wang et al., 2026). This structure enables post hoc decomposition of observed vari- ance into tester-controlled factors and system-selected sources. Conditional and sensitivity analyses then support attribution at the level of conditional dependence, providing more granularity than undifferentiated variance, though still short of full causal attribution. Trial replayRepeated execution of an identical configuration produced a distribution of outcomes, as established in §4.2.1. Observability mechanisms capture the events of a single run but do not enable re-execution. Record-replay instrumentation addresses this by capturing and serializing all external interactions, such as model invocation, tool outputs, and retrieved data, so the run can be reconstructed in isolation by substituting recorded responses for live ones (Mudasiru, 2026). This enables harnesses to quantify determinism at both the trajectory and decision levels as audit properties (Khatchadourian, 2026). Controlled perturbation during replay allows specific failures to be reproduced for diagnostics and retesting (Mazumder and Lia, 2026). The resulting trial record provides a fixed reference for subsequent runs, supporting regression analysis and fault attribution under controlled conditions. However, fidelity is limited to the interactions captured by the harness; thus, replay demonstrates reproducibility of the recorded configuration, not of the deployed system as a whole. State fuzzing Fuzzing techniques can help assess the role of a specific memory item or class in behavior change. It involves systematically manipulating the contents of accumulated memory while keeping the environment and tool interface constant. This approach enables the attribution of observed behavioral differences specifically to variations in memory. By employing a memory-free control for comparison, researchers have demonstrated that an agentâs behavioral tendencies can be causally linked to its accumulated memory (Dabas et al., 2026). Applying this diagnostic method to operational systems requires tagging each memory item with its provenance at the time of creation to attribute behavioral changes to specific accumulated states (X. J. Wang et al., 2026). However, such metadata is infrequently recorded in practice. The scarcity of in-service drift detection in deployed systems suggests that the necessary write-time lineage tracking is rarely implemented (Vatsal, Dubey, and Singh, 2026). Because that lineage must be written when each item is created, it faces the same adoption friction, which likely explains its scarcity. Relatedly, adversarial strategies, such as memory-poisoning red-teaming, intentionally introduce harmful information into an agentâs persistent memory, enabling assessment of how that information influences subsequent planning or execution and providing pre-deployment insight into behavioral changes driven by internal state alterations (Z. Chen et al., 2024). 5.2.3 Correction A related claim is the ability to correct or revoke mission-relevant state. We propose that correction rests on three subsequent capabilities: state adjudication, propagation, and bounded renewal. First, state adjudication evaluates whether a state item remains a valid and authorized basis for action. Second, propagation ensures that authorized corrections, deletions, or revocations reach all decision- relevant dependent states within operationally meaningful time bounds. Third, bounded renewal 35 describes the ability to re-establish an operational baseline after a model, harness or policy change in service without discarding all accumulated state. State adjudicationExisting work highlights that even if memory is observable, stale and updated memories may coexist without the system being able to adjudicate which is authoritative and which is outdated (Chao et al., 2026). To address adjudication, verifiable memory governance testing approaches advocate for evaluating write permissions, provenance transparency, principal-specific data retrieval, rollback capability, and verified forgetting (Lin et al., 2026). These criteria are directly applicable to assessing whether incorrect or outdated information persists within an agent across its operational lifecycle. In C2 specifically, these properties underpin the operational picture. The validity and currency of retained state determine the quality of the information and the situational awareness on which the commander reasons, so memory governance safeguards those criteria. Additionally, collaborative memory approaches (Rezazadeh et al., 2025) may help determine whose state can be read or written using access graphs, private/shared tiers, and principal-specific controls, although permission filters alone do not prevent contradictory or laundering-prone beliefs (X. Li et al., 2026). Propagation Current architectures often struggle to maintain correction coherence once state spreads across subagents and sessions (Ding et al., 2026). Certified state reversion can help revert a bounded set of decisions by allowing the system to trace affected actions back to the records that authorized them. However, this approach discards dependencies acquired since the snapshot. To address dependency issues and belief propagation failures, Ding et al. (2026) propose expanding stale- state search beyond directly touched slots to structurally affected state regions. Another approach models derivation edges (i.e., the links between a core belief and its summaries, shared copies, and tool actions) and automatically repairs downstream effects after a correction (X. Li et al., 2026). Still, revocation is challenging because damage accumulates before all copies acknowledge the change, making timing a key limitation in C2 contexts. Timely, dependency-aware repair is therefore essential for successful propagation. Bounded renewalT&E approaches to renewal suggest three approaches. First, partitioned adapta- tion approaches propose freezing functional layers responsible for action and permitting adaptation within designated subsystems. For agentic systems, such partitioning separates memory and tool- access layers from action layers, provided the system maintains the immutability of critical layers. Second, approaches that maintain a frozen reference twin of the system, as certified (e.g., model, configuration, starting state), and periodically replay a sample of live mission inputs can help evaluate divergence in the accumulated state between the fielded, state-laden system and its frozen twin. Third, higher-order governance mechanisms could involve behavior-triggered re-accreditation (Sidhu et al., 2026), modeled after aviation and nuclear power (International Atomic Energy Agency, 2013); screening rules for declared updates to determine which evidence remains valid after a change; or change control plans that pre-authorize a limited range of modifications (Administration, 2024). 5.3 Assuring multi-agent composition and emergence Agentic systems often operate in teams. In multi-agent systems (MAS), an orchestrating agent decomposes objectives into tasks, delegates these tasks to subagents, and then synthesizes their outputs. subagents perform assigned work, consult external sources, utilize tools, exchange intermediate products, and communicate in mixed formats, including natural language (P7; see scenarios 2-4). This architecture enables dynamic work allocation and division of labor. The resulting flexibility is a key feature of MAS, making them particularly attractive in unbounded contexts such as military C2. §4.3 establishes that a property inherent to these systems (i.e., that the interface between agents is an output generated at runtime) undermines two conditions underpinning the logic of staged test 36 progression: that tested agent behavior remains a reliable guide when integrated in a MAS; and that MAS behavior is composable from agent behaviors and their specified interactions. To address these gaps, this section evaluates the potential of T&E methods and approaches to generate evidence toward four assurance claims introduced in Table 9. 10 IDClaimStatement C3.1CompositionConfidence evidenced at the component level is re-established at the as- sembly level, and any departures of assembly behavior from component evidence are measured, bounded, and attributable to identified design choices. C3.2InterfacesThe assemblyâs internal interfaces constitute an observable and charac- terized test surface. C3.3Attribution & propagation Assembly-level failures can be traced to their origin, and error propaga- tion can be characterized. C3.4EmergenceEmergent capabilities or goals are detected and characterized. Table 9: Assurance claims for system composability 5.3.1 Composition Restoring composition requires confidence evidenced at the component level to be re-established at the assembly level, and departures of assembly behavior from component evidence to be measured, bounded, and attributable to identified design choices. We outline T&E approaches that may help generate evidence toward this claim. System-level evaluation with assembly choices as factors Design choices underpinning MAS architectures introduce a new layer of testing complexity. Recent work considering the assembled MAS as the unit of analysis (e.g., agents, harness, and coordination logic) 11 finds that framework choice affects task performance comparably to model choice (Emde et al., 2026; Orogat and Rostam, 2026). We suggest re-running representative scenarios using a similar framework-agnostic test harness while varying the assemblyâs organizational structure (âtopologyâ), coordination logic, and other design choices as discretized factors. Combinatorial interaction testing (CIT) techniques can create test suites with economical coverage when the number of factors and levels precludes exhaustive testing (Lanus et al., 2021). 12 Comparing the assembled systemâs observed performance with what component (individual agent) benchmarks predict can help quantify the part of the outcome attributable to assembly versus the components themselves. Composed-system red-teaming An assembled system does not necessarily inherit the safety or security demonstrated by its individual components. In one experiment, a malicious actor achieved unsafe outputs 43% of the time with an assembly of two models that were ostensibly safe, complying less than 3% of the time when attacked individually (Jones, Dragan, and Steinhardt, 2024). Although this finding arises from an adversarial model combination and not an orchestrated C2 system, it 10 Supervisability is more acute in multi-agent settings, since delegation across agents raises the tempo and scale of what must be supervised while reducing the time available to intervene. The scope of this paper with respect to A8 is set out in §1.3. 11 A framework-agnostic test harness was provided to allow for direct comparison of framework-level design choices by holding all other variables constant. 12 CITâs economical coverage rests on the low-order-interaction premise, which is unvalidated for this system class (§5.1) and least tested for the topology and coordination factors varied here. 37 supports the principle of composition-induced failure and suggests that MAS red-teaming should target the composed system rather than its individual components. MAS should be tested for both adversarial (e.g., malformed instructions, contradictory goals, information asymmetries, corrupted agents) and natural robustness (e.g., environmental perturbations, partial system failures, resource constraints, unexpected state changes) to explore edge cases and discover high-impact failure modes (Reid et al., 2025). Reid et al. incorporate red-teaming into a staged evidence strategy alongside simulation, observation, and benchmarking, arguing that no single method covers the failure modes under consideration. As such, composed-system red-teaming probes the assumptions of other T&E methods and should be treated as complementary to them. Notably, this practice is immature, coverage is scenario-bounded, and no unified protocol is publicly reported, even in frontier AI companies. Reopening of system-level assurance after roster or structure changesReid et al. (2025) justify this governance rule because composition is non-additive; âa collection of safe agents does not guarantee a safe collection of agentsâ. Re-certifying a new agent on its own is not accepted as evidence about the newly assembled team. Knack et al. (2025) observe that a single AI system may comprise multiple interacting models, and recommend system cards (comprehensive documentation of the entire system) over model cards, which document individual models in isolation. Supplier diversity in procurementFleets built on one or two foundation models can fail together, for the same reason, in ways that single-system testing cannot demonstrate. Reid et al. (2025) identify monoculture collapse as a key MAS failure mode, arguing that a shared model foundation creates correlated vulnerabilities. Meanwhile, the broader literature on algorithmic monoculture finds that shared underlying models homogenize outcomes across nominally independent decisions in vision and language settings (Bommasani et al., 2022). This finding carries important epistemic and operational implications for C2 decision-making. 5.3.2 Interfaces MAS complicate the traditional T&E process by depriving testers of the specification from which a test surface is typically derived. The interface claim is about rebuilding one from the assemblyâs internal interfaces. This includes a commanderâs initial command, framed here as an upstream natural-language output interpreted by the orchestrator or other downstream agents. Delegation channel fuzzing Fuzzing stimulates a system by generating a large number of semi- valid inputs to discover unexpected behaviors (Joiner, 2024). In a MAS, fuzzing techniques can systematically perturb inter-agent communications. Measuring resulting changes in downstream agent behavior across agent stability dimensions enables the mapping of which hops amplify or propagate which types of distortions. Demonstrated LLM-to-LLM prompt infection attacks give the perturbation classes a starting point (D. Lee and Tiwari, 2024). Structured delegation protocolsAgent interoperability protocols are an emerging and promising mechanism for making authority transfer more explicit, enabling attached feedback mechanisms to verify task completion, and carrying trust ratings for downstream agents (TomaĹĄev, Franklin, and Osindero, 2026). A handover documenting who granted authority to whom and the outcomes is crucial for attribution. 5.3.3 Attribution and propagation Emerging phenomena observed in MAS contexts entail that departures from intended outcomes are neither bounded in advance nor attributable to a specific integration decision after the fact. 38 Propagation means an error gains authority each time an agent relays it. This evidence is gathered by seeding known faults and observing how the assembly responds, using instrumentation that records movements and locations. Fault injection applied at delegation hops Deliberately introducing faults to observe how they propagate and whether the system detects them is a long-standing practice in dependable-systems engineering (Hsueh, Tsai, and Iyer, 1997). Cemri et al. (2025) introduce a Multi-Agent System Failure Taxonomy (MAST) that organizes 14 failure modes into three categories: specification issues, inter-agent misalignment, and task verification. This taxonomy can serve as a coverage target for fault injection, with known bad agent interactions generated in two ways. First, by templating from the taxonomyâs modes to alter a correct output captured from a clean run. Second, by replaying failures observed in earlier trials. Cemri et al. (2025) developed a domain-agnostic, LLM-based annotator to indicate the specific failure mode(s) of a captured run trace. This could be used to flag previous runs for replay injection or to measure taxonomy coverage for the fault injection test suite. Injection points, such as specific hops and stages, can also be selected by a control-structure sketch of the assembly informed by Systems Theoretic Process Analysis (STPA). Independent validation and transaction safeguards at handoffsChang and Geng (2025) distin- guish validation as a separate function from the agents under validation. In their approach, dedicated validation agents inspect outputs before release and verify incoming inputs and their dependencies, while mechanisms for persistent state tracking, checkpointing, and compensation record each agentâs actions and enable reversal if necessary. As a result, every message retains its provenance, ensuring that errors which accumulate authority through relaying (Scenario 3) remain traceable and can be actively reversed. Validation at each hop identifies malformed or out-of-range content, but does not address content that is structurally correct yet substantively incorrect. As the architecture remains at the proposal stage, it delineates where empirical evidence could be collected. 5.3.4 Emergence Emergence is not necessarily a failure mode. For instance, flexible collective problem solving is a primary motivation for adopting multi-agent designs. Emergent behaviors may be either beneficial or detrimental. The primary T&E challenge is to demonstrate that desirable emergent behaviors are reliable, while undesirable behaviors remain unlikely (H. M. Wojton, Porter, and Dennis, 2020). Several methods can support this claim. Behavioral degradation as an emergence indicator Rath (2026) introduces an agent stability metric framework that quantitatively measures behavioral degradation, in the form of agentic drift, across 12 dimensions within four categories: response consistency, tool usage patterns, inter-agent coordination, and behavioral boundaries. In addition to its impacts on accuracy and efficiency, agentic drift can occur alongside emergent behaviors such as specification gaming and reward hacking. Over time, agentic systems develop behaviors that may technically satisfy task objectives while diverging from true intent or violating implicit norms. Rath proposes that comprehensive behavioral monitoring, both during integration testing and post-deployment, can quantify behavioral degradation and may indicate an increasing risk of undesired emergent behavior. Observed instances of emergent behavior can inform the thresholds for triggering governance mechanisms and/or internal mitigation strategies, such as episodic memory consolidation, drift-aware routing, and adaptive behavioral anchoring (Rath, 2026). Adaptive multi-dimensional monitoring applied as per-axis adaptive thresholds with joint anomaly detection may reduce detection latency and false positives versus static thresholds (Shukla, 2025). 13 13 Since these thresholds update their reference statistics online, they share the risk noted in §5.2 that a recalibrated baseline can absorb slow degradation. Consequently, they should be paired with a validated fixed reference and used for sudden-drift 39 Multi-agent behavioral degradation triggers re-accreditationAs in §5.2.3, a measured change in the assemblyâs interaction behavior could re-open accreditation. Such a trigger could be written over the class of agent stability indicators, with the specific signal swapped as evidence improves, without rewriting the rule. This governance trigger depends on unvalidated thresholds, in line with the requirement that model-based evaluators be validated before use (see §5.1.2). 6 Conclusion This analysis addressed two central questions: the degree of confidence that current Testing and Evaluation (T&E) methods can substantiate for agentic AI systems in command and control, and the management of residual uncertainty in fielding decisions. For the first, the properties that make agentic systems operationally valuable challenge eight foundational assumptions underlying established methods around system specifiability, stability, composability, and supervisability. Recovery was assessed for three of these four clusters. Supervisability was identified but only partly carried into the recovery analysis, for the reasons set out in §1.3. For the second, the analysis identified several mechanisms to delineate, constrain, and allocate residual uncertainty. These include substituting broad, system-agnostic claims for agentic-specific assertions, translating critical constraints into enforceable runtime mechanisms, extending evidence generation into the operational phase, and assigning residual uncertainty to governance structures with defined expiry conditions and ownership. These conclusions are bounded by the methodological limitations outlined above. The analysis is restricted to the publicly available record, primarily reflecting US, UK, and NATO practices; classified initiatives may address gaps not visible here, and formal sources do not fully capture informal practices. The challenges that agentic properties pose to established methods are inferred from system properties and methodological descriptions, rather than empirically demonstrated through operational testing. Many of the candidate approaches discussed in §5 remain at the preprint stage and have not been validated under command-and-control conditions. Accordingly, the findings delineate areas where evidentiary support is limited and do not offer judgments on the adequacy of any specific fielded system. Each unresolved gap points to a new area for research. A primary unresolved issue is calibration: determining the threshold and nature of behavioral changes that should prompt revalidation, rollback, or decommissioning of a system. Further questions are how much mission-level judgment can be formalized into executable constraints, and whether composability techniques can support the dynamic reconfiguration of assemblies. The largest gap is in supervisability, where criteria for adequate situation awareness, intervention time, and supervisory load depend on human factors methods outside the T&E corpus examined here and on the stability characterization in §5.2, since reliance can be judged as calibrated only against a measured baseline. Agentic systems are deployed under commitments to rigorous testing and human oversight. Until these criteria exist, the second of those commitments rests on the presence of an operator rather than on evidence that oversight is exercised, and the credibility of both depends on the quality of the supporting claims, evidence, and reasoning that testing and evaluation are able to produce. latency, not gradual-degradation detection and, like any characterized decision boundary (§5.2, exploitability principle), protected from being learned. 40 References Abhay, Dr. (June 15, 2025). âAutomated Drift Detection and Retraining Pipeline for ML Modelsâ. In: International Journal of Scientific Research in Engineering and Management 09.6, p. 1â9. ISSN: 25823930. DOI:10.55041/IJSREM50192. URL:https://ijsrem.com/download/ automated-drift-detection-and-retraining-pipeline-for-ml-models/ (visited on 08/15/2026). Acharya, Deepak Bhaskar, Karthigeyan Kuppan, and B. Divya (2025). âAgentic AI: Autonomous Intelligence for Complex GoalsâA Comprehensive Surveyâ. In: IEEE Access 13, p. 18912â 18936. ISSN: 2169-3536. DOI:10.1109/ACCESS.2025.3532853. URL:https://ieeexplore. ieee.org/document/10849561 (visited on 08/15/2026). Ackerman, Samuel et al. (Sept. 6, 2022). Detection of data drift and outliers affecting machine learning model performance over time. DOI:10.48550/arXiv.2012.09258. arXiv:2012. 09258[stat.AP]. URL: http://arxiv.org/abs/2012.09258 (visited on 08/15/2026). Adimulam, Apoorva, Rajesh Gupta, and Sumit Kumar (Jan. 20, 2026). The Orchestration of Multi- Agent Systems: Architectures, Protocols, and Enterprise Adoption. DOI:10.48550/arXiv.2601. 13671. arXiv:2601.13671[cs.MA]. URL:http://arxiv.org/abs/2601.13671(visited on 08/15/2026). Administration, U.S. Food and Drug (Aug. 2024). Predetermined Change Control Plans for Medical Devices. URL:https://w.fda.gov/regulatory-information/search-fda- guidance - documents / predetermined - change - control - plans - medical - devices (visited on 08/12/2026). Alberts, David S. and Richard E. Hayes (2006). Understanding command and control. OCLC: 62872850. Washington, D.C.: CCRP Publications. ISBN: 978-1-893723-17-7. Anthropic (Jan. 9, 2026a). Demystifying evals for AI agents. URL:https://w.anthropic.com/ engineering/demystifying-evals-for-ai-agents (visited on 08/12/2026). â (July 30, 2026b). Investigating three real-world incidents in our cybersecurity evaluations. URL: https://w.anthropic.com/news/investigating- incidents- cybersecurity- evals (visited on 08/12/2026). Arike, Rauno et al. (Oct. 15, 2025). âEvaluating Goal Drift in Language Model Agentsâ. In: Proceed- ings of the AAAI/ACM Conference on AI, Ethics, and Society 8.1, p. 192â203. ISSN: 3065-8365. DOI:10.1609/aies.v8i1.36541. URL:https://ojs.aaai.org/index.php/AIES/ article/view/36541 (visited on 08/15/2026). Atil, Berk et al. (Apr. 2, 2025). Non-Determinism of "Deterministic" LLM Settings. DOI:10.48550/ arXiv.2408.04667. arXiv:2408.04667[cs.CL]. URL:http://arxiv.org/abs/2408. 04667 (visited on 08/15/2026). Barke, Shraddha et al. (Feb. 2, 2026). AgentRx: Diagnosing AI Agent Failures from Execution Trajectories. DOI:10.48550/arXiv.2602.02475. arXiv:2602.02475[cs.AI]. URL:http: //arxiv.org/abs/2602.02475 (visited on 08/15/2026). Barr, Earl T. et al. (May 1, 2015). âThe Oracle Problem in Software Testing: A Surveyâ. In: IEEE Transactions on Software Engineering 41.5, p. 507â525. ISSN: 0098-5589, 1939-3520. DOI: 10.1109/TSE.2014.2372785. URL:http://ieeexplore.ieee.org/document/6963470/ (visited on 08/12/2026). Barrak, Amine (2025). Traceability and Accountability in Role-Specialized Multi-Agent LLM Pipelines. Version Number: 1. DOI:10.48550/ARXIV.2510.07614. URL:https://arxiv. org/abs/2510.07614 (visited on 08/12/2026). Bhardwaj, Varun Pratap (Feb. 25, 2026). Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents. DOI:10.5281/zenodo.18775393. arXiv:2602 . 22302[cs . AI]. URL:http : / / arxiv . org / abs / 2602 . 22302(visited on 08/14/2026). Blackwell, Robert E., Jon Barry, and Anthony G. Cohn (June 27, 2025). Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores. DOI:10.48550/arXiv.2410. 41 03492. arXiv:2410.03492[cs.CL]. URL:http://arxiv.org/abs/2410.03492(visited on 08/15/2026). Bogg, Adam and Stewart Birrell (Sept. 2025). âOverloaded, underloaded or in control: How many automated vehicles can one person supervise?â In: Computers in Human Behavior 170, p. 108690. ISSN: 07475632. DOI:10 . 1016 / j . chb . 2025 . 108690. URL:https : / / linkinghub . elsevier.com/retrieve/pii/S0747563225001372 (visited on 08/18/2026). Bommasani, Rishi et al. (2022). Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization? URL: https://arxiv.org/abs/2211.13972. Boyd, John R (Jan. 1996). âThe Essence of Winning and Losingâ. In: URL:https : / / slightlyeastofnew . com / wp - content / uploads / 2010 / 03 / essence _ of _ winning _ losing.pdf. Bryce, RenĂŠe C. et al. (2010). âCombinatorial Testingâ. In: Handbook of Research on Software Engineering and Productivity Technologies. Ed. by Muthu Ramachandran and RogĂŠrio Atem De Carvalho. IGI Global, p. 196â208. ISBN: 978-1-60566-731-7 978-1-60566-732-4. DOI:10.4018/ 978-1-60566-731-7.ch014. URL:http://services.igi-global.com/resolvedoi/ resolve.aspx?doi=10.4018/978-1-60566-731-7.ch014 (visited on 08/20/2026). Burton, Simon and Benjamin Herd (Apr. 6, 2023). âAddressing uncertainty in the safety assurance of machine-learningâ. In: Frontiers in Computer Science 5, p. 1132580. ISSN: 2624-9898. DOI: 10.3389/fcomp.2023.1132580. URL:https://w.frontiersin.org/articles/10. 3389/fcomp.2023.1132580/full (visited on 08/18/2026). Calvagna, Andrea and Angelo Gargantini (Nov. 2012). âT-wise combinatorial interaction test suites construction based on coverage inheritanceâ. In: Software Testing, Verification and Reliability 22.7, p. 507â526. ISSN: 0960-0833, 1099-1689. DOI:10.1002/stvr.466. URL:https: //onlinelibrary.wiley.com/doi/10.1002/stvr.466 (visited on 08/17/2026). Capellera, Guillem et al. (June 2025). âUnified Uncertainty-Aware Diffusion for Multi-Agent Tra- jectory Modelingâ. In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). ISSN: 2575-7075, p. 22476â22486. DOI:10.1109/CVPR52734.2025.02093. URL:https: //ieeexplore.ieee.org/document/11094594 (visited on 08/15/2026). â(2026). âHeteroscedastic Diffusion for Multi-Agent Trajectory Modelingâ. In: IEEE Transactions on Pattern Analysis and Machine Intelligence, p. 1â13. ISSN: 0162-8828, 2160-9292, 1939- 3539. DOI:10 . 1109 / TPAMI . 2026 . 3692903. arXiv:2605 . 10717[cs . LG]. URL:http : //arxiv.org/abs/2605.10717 (visited on 08/15/2026). Cemri, Mert et al. (Mar. 17, 2025). âWhy Do Multi-Agent LLM Systems Fail?â In: arXiv preprint. DOI: 10.48550/arXiv.2503.13657. URL: https://arxiv.org/abs/2503.13657. Chan, Alan et al. (2023). âHarms from Increasingly Agentic Algorithmic Systemsâ. In: Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. FAccT â23. New York, NY, USA: Association for Computing Machinery, p. 651â666. ISBN: 979-8-4007-0192-4. DOI: 10.1145/3593013.3594033. URL: https://doi.org/10.1145/3593013.3594033. Chang, Edward Y. and Longling Geng (2025). SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning. URL:https://arxiv.org/abs/2503. 11951. Chao, Hanxiang et al. (2026). STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Version Number: 1. DOI:10.48550/ARXIV.2605.06527. URL:https://arxiv.org/ abs/2605.06527 (visited on 08/12/2026). Chen, Chaoran et al. (2026). Comparing Human Oversight Strategies for Computer-Use Agents. Version Number: 1. DOI:10.48550/ARXIV.2604.04918. URL:https://arxiv.org/abs/ 2604.04918 (visited on 08/18/2026). Chen, Zhaorun et al. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. Version Number: 1. DOI:10.48550/ARXIV.2407.12784. URL:https: //arxiv.org/abs/2407.12784 (visited on 08/12/2026). 42 Cheng, Tingting et al. (June 2024). âAnalysis of human errors in human-autonomy collaboration in autonomous ships operations through shore control experimental dataâ. In: Reliability Engineering & System Safety 246, p. 110080. ISSN: 09518320. DOI:10.1016/j.ress.2024.110080. URL: https://linkinghub.elsevier.com/retrieve/pii/S0951832024001546(visited on 08/17/2026). Chief Digital and Artificial Intelligence Office (Apr. 2024). Systems Integration Test and Eval- uation of Artificial Intelligence-Enabled Capabilities. Washington, DC: U.S. Department of Defense. URL:https://w.ai.mil/Portals/137/Documents/Resources%20Page/ CDAO _ TE _ Framework _ SI _ TES _ RELEASED _ APRIL _ 2024 - compressed . pdf ? ver = d27aVLoJP7c8qt4arC9u6w%3d%3d (visited on 05/17/2026). Choi, Jae-Woo et al. (May 24, 2026). âReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planningâ. In: Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems. AAMAS â26. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, p. 319â328. ISBN: 979-8-4007-2317-9. DOI: 10.65109/UCGT7089. URL:https://dl.acm.org/doi/10.65109/UCGT7089(visited on 08/15/2026). Crandall, J.W. and M.L. Cummings (Oct. 2007). âIdentifying Predictive Metrics for Supervisory Control of Multiple Robotsâ. In: IEEE Transactions on Robotics 23.5, p. 942â951. ISSN: 1552- 3098. DOI:10.1109/TRO.2007.907480. URL:http://ieeexplore.ieee.org/document/ 4339527/ (visited on 08/17/2026). Cummings, Mary (Sept. 20, 2004). âAutomation Bias in Intelligent Time Critical Decision Support Systemsâ. In: AIAA 1st Intelligent Systems Technical Conference. Infotech@Aerospace Confer- ences. American Institute of Aeronautics and Astronautics. DOI:10.2514/6.2004-6313. URL: https://arc.aiaa.org/doi/10.2514/6.2004-6313 (visited on 08/12/2026). Dabas, Mahavir et al. (2026). Memory-Induced Tool-Drift in LLM Agents. Version Number: 1. DOI: 10.48550/ARXIV.2605.24941. URL:https://arxiv.org/abs/2605.24941(visited on 08/18/2026). Department for Science, Innovation & Technology and DSIT (Apr. 28, 2025). Frontier AI: capabil- ities and risks â discussion paper. GOV.UK. URL:https://w.gov.uk/government/ publications / frontier - ai - capabilities - and - risks - discussion - paper / frontier-ai-capabilities-and-risks-discussion-paper (visited on 08/12/2026). Department of Defense (Nov. 8, 2010). Department of Defense Dictionary of Military and Associated Terms - Joint Publication 1-02. Department of Defense. Dhanorkar, Shipi, Samir Passi, and Mihaela Vorvoreanu (June 25, 2026). âHuman oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agentsâ. In: Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency. FAccT â26: The 2026 ACM Conference on Fairness, Accountability, and Transparency. Montreal QC Canada: ACM, p. 6438â6465. ISBN: 979-8-4007-2596-8. DOI:10. 1145/3805689.3812402 . URL:https://dl.acm.org/doi/10.1145/3805689.3812402 (visited on 08/18/2026). Ding, Tianyu et al. (June 30, 2026). Always-On Agents: A Survey of Persistent Memory, State, and Governance in LLM Agents. Version Number: 1. DOI:10.48550/ARXIV.2606.30306. URL: https://arxiv.org/abs/2606.30306 (visited on 08/12/2026). Dobslaw, Felix et al. (Oct. 20, 2025). Challenges in Testing Large Language Model Based Software: A Faceted Taxonomy. DOI:10.48550/arXiv.2503.00481. arXiv:2503.00481[cs.SE]. URL: http://arxiv.org/abs/2503.00481 (visited on 08/15/2026). DOW Unleashes âAgent Networkâ to Transform AI-Enabled Battle Management and Targeting (June 25, 2026). U.S. Department of War. URL:https://w.war.gov/News/Releases/ Release/Article/4526862/dow- unleashes- agent- network- to- transform- ai- enabled-battle-management-and-targe/ (visited on 08/12/2026). 43 Dutta, Srimonti and Akshata Kishore Moharir (June 20, 2026). AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems. DOI:10.48550/arXiv.2606.21877. arXiv: 2606.21877[cs.AI]. URL: http://arxiv.org/abs/2606.21877 (visited on 08/15/2026). Dwivedi, Yogesh K. et al. (2026). âAgentic AI Systems: What It Is and Isnâtâ. In: Global Business and Organizational Excellence 45.3, p. 253â263. ISSN: 1932-2062. DOI:10.1002/joe.70018. URL:https://onlinelibrary.wiley.com/doi/abs/10.1002/joe.70018(visited on 08/18/2026). El Arab, Rabie Adel et al. (Jan. 2026). âBeyond Model Development in Healthcare AI: Post- Development Robustness, Post-Deployment Monitoring, and Lifecycle GovernanceâA Scop- ing Review of Reviewsâ. In: Healthcare 14.11, p. 1459. ISSN: 2227-9032. DOI:10.3390/ healthcare14111459 . URL:https://w.mdpi.com/2227-9032/14/11/1459(visited on 08/15/2026). Emde, Cornelius et al. (Mar. 9, 2026). âMASEval: Extending Multi-Agent Evaluation from Models to Systemsâ. In: arXiv preprint. DOI:10.48550/arXiv.2603.08835. URL:https://arxiv. org/abs/2603.08835. Endsley, Mica R. (1999). âSituation Awareness In Aviation Systemsâ. In: URL:https://api. semanticscholar.org/CorpusID:40229510. â (Feb. 2017). âFrom Here to Autonomy: Lessons Learned From HumanâAutomation Researchâ. In: Human Factors: The Journal of the Human Factors and Ergonomics Society 59.1, p. 5â27. ISSN: 0018-7208, 1547-8181. DOI:10.1177/0018720816681350. URL:https://journals. sagepub.com/doi/10.1177/0018720816681350 (visited on 08/17/2026). Endsley, Mica R. and Esin O. Kiris (June 1995). âThe Out-of-the-Loop Performance Prob- lem and Level of Control in Automationâ. In: Human Factors: The Journal of the Human Factors and Ergonomics Society 37.2, p. 381â394. ISSN: 0018-7208, 1547-8181. DOI:10 . 1518/001872095779064555 . URL:https://journals.sagepub.com/doi/10.1518/ 001872095779064555 (visited on 08/17/2026). Facchini, Alessandro and Alberto Termine (2022). âTowards a Taxonomy for the Opacity of AI Systemsâ. In: Philosophy and Theory of Artificial Intelligence 2021. Ed. by Vincent C. MĂźller. Cham: Springer International Publishing, p. 73â89. ISBN: 978-3-031-09153-7. DOI:10.1007/ 978-3-031-09153-7_7. Fan, Shengda et al. (June 1, 2026). AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents. DOI:10.48550/arXiv.2603.14465. arXiv:2603.14465[cs.AI]. URL: http://arxiv.org/abs/2603.14465 (visited on 08/15/2026). Fang, Zhengyu et al. (Apr. 15, 2026). Dataset-Level Metrics Attenuate Non-Determinism: A Fine- Grained Non-Determinism Evaluation in Diffusion Language Models. DOI:10.48550/arXiv. 2604.13413. arXiv:2604.13413[cs.LG]. URL:http://arxiv.org/abs/2604.13413 (visited on 08/15/2026). Figueiredo, H, D Cook, and W Biggs (July 2, 2019). âThe Test and Evaluation of Intelligent Autonomous Systems (IAS)â. In: Engine As A Weapon International Symposium VIII. London, UK. DOI:10.24868/issn.2515-8171.2019.001. URL:https://library.imarest.org/ record/7547 (visited on 08/15/2026). Fleisher, Will (Dec. 2022). âUnderstanding, Idealization, and Explainable AIâ. In: Episteme 19.4, p. 534â560. ISSN: 1742-3600, 1750-0117. DOI:10.1017/epi.2022.39. URL:https:// w.cambridge.org/core/product/identifier/S1742360022000399/type/journal_ article (visited on 08/18/2026). Fourney, Adam et al. (Nov. 7, 2024). Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks. DOI:10.48550/arXiv.2411.04468. arXiv:2411.04468[cs.AI]. URL: http://arxiv.org/abs/2411.04468 (visited on 08/15/2026). Freeman, Laura, John Robert, and Heather Wojton (July 28, 2025). âThe Impact of Generative AI on Test & Evaluation: Challenges and Opportunitiesâ. In: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering. FSE Companion â25. New York, NY, 44 USA: Association for Computing Machinery, p. 1376â1380. ISBN: 979-8-4007-1276-0. DOI:10. 1145/3696630.3728723 . URL:https://dl.acm.org/doi/10.1145/3696630.3728723 (visited on 08/15/2026). Gao, Haowen et al. (2026). SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing. Version Number: 1. DOI:10.48550/ARXIV.2606.14239. URL:https://arxiv. org/abs/2606.14239 (visited on 08/18/2026). Ghosh, Shaona et al. (2025). A Safety and Security Framework for Real-World Agentic Systems. Version Number: 1. DOI:10.48550/ARXIV.2511.21990. URL:https://arxiv.org/abs/ 2511.21990 (visited on 08/17/2026). Gonzalez, Miguel Angel Alvarado et al. (Sept. 28, 2025). Do Repetitions Matter? Strengthening Reliability in LLM Evaluations. DOI:10.48550/arXiv.2509.24086. arXiv:2509.24086[cs. AI]. URL: http://arxiv.org/abs/2509.24086 (visited on 08/15/2026). Greshake, Kai et al. (Nov. 26, 2023). âNot What Youâve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injectionâ. In: Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security. AISec â23. New York, NY, USA: Association for Computing Machinery, p. 79â90. ISBN: 979-8-4007-0260-0. DOI:10.1145/3605764.3623985. URL: https://dl.acm.org/doi/10.1145/3605764.3623985 (visited on 08/12/2026). Grover, Savi (Oct. 6, 2025). âEngineering Robust Ai Products Through Continuous Quality Assur- ance: A Framework for Testing, Monitoring, and Validation of Adaptive Live Learning Ai/Ml Systems in Dynamic Production Environmentsâ. In: International Journal of Applied Mathe- matics 38.2, p. 1092â1113. ISSN: 1314-8060. DOI:10.12732/ijam.v38i2s.710. URL: https://ijamjournal.org/ijam/publication/index.php/ijam/article/view/710 (visited on 08/17/2026). Guo, Hongjing, Chuanqi Tao, and Zhiqiu Huang (July 25, 2024). âNeuron importance-aware coverage analysis for deep neural network testingâ. In: Empirical Software Engineering 29.5, p. 118. ISSN: 1573-7616. DOI:10.1007/s10664-024- 10524-x. URL:https://doi.org/10.1007/ s10664-024-10524-x (visited on 08/15/2026). Haidemariam, Tsehaye (Jan. 12, 2026). âFrom the logic of coordination to goal-directed reasoning: the agentic turn in artificial intelligenceâ. In: Frontiers in Artificial Intelligence 8, p. 1728738. ISSN: 2624-8212. DOI:10.3389/frai.2025.1728738. URL:https://w.frontiersin. org/articles/10.3389/frai.2025.1728738/full (visited on 08/15/2026). Hammond, Lewis et al. (Feb. 19, 2025). âMulti-Agent Risks from Advanced AIâ. In: arXiv preprint. DOI: 10.48550/arXiv.2502.14143. URL: https://arxiv.org/abs/2502.14143. He, Pengfei et al. (2025). TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use. Version Number: 2. DOI:10.48550/ARXIV.2510.04550. URL:https://arxiv. org/abs/2510.04550 (visited on 08/19/2026). Herrmann, Thomas (2025). âIntervenability as a Design Requirement for Autonomy and Oversight within Human-Centered AIâ. In: p. 143â166. DOI:10.1007/978-3-031-83512-4_9. arXiv: 2607.10322[cs.HC]. URL: http://arxiv.org/abs/2607.10322 (visited on 08/18/2026). Hsueh, Mei-Chen, T.K. Tsai, and R.K. Iyer (Apr. 1997). âFault injection techniques and toolsâ. In: Computer 30.4, p. 75â82. ISSN: 1558-0814. DOI:10.1109/2.585157. URL:https: //ieeexplore.ieee.org/document/585157 (visited on 08/12/2026). Hu, Yuyang et al. (2025). Memory in the Age of AI Agents. Version Number: 2. DOI:10.48550/ ARXIV.2512.13564. URL: https://arxiv.org/abs/2512.13564 (visited on 08/19/2026). Hung, Man et al. (Apr. 2026). âDeliberation and drift: Evaluating alignment fragility in multi-agent medical artificial intelligenceâ. In: AI and Ethics 6.2, p. 198. ISSN: 2730-5953, 2730-5961. DOI: 10.1007/s43681-026-01048-9. URL:https://link.springer.com/10.1007/s43681- 026-01048-9 (visited on 08/15/2026). Huq, Faria et al. (July 9, 2026). Modeling Distinct Human Interaction in Web Agents. Version Number: 4. DOI:10.48550/ARXIV.2602.17588. URL:https://arxiv.org/abs/2602.17588(visited on 08/18/2026). 45 International Atomic Energy Agency (2013). Periodic Safety Review for Nuclear Power Plants. IAEA Safety Standards Series SSG-25. Vienna: International Atomic Energy Agency. ISBN: 978-92-0- 137410-3. URL:http://w.iaea.org/publications/8911/periodic-safety-review- for-nuclear-power-plants. Jammalamadaka, K. and Nikhat Parveen (Apr. 21, 2021). âTesting coverage criteria for opti- mized deep belief network with search and rescueâ. In: Journal of Big Data 8. DOI:10 . 1186 / s40537 - 021 - 00453 - 7. URL:https : / / consensus . app / papers / testing - coverage - criteria - for - optimized - deep - belief - jammalamadaka - parveen / 4629accdc37e56b08fe45bd4e9e065b/. Joiner, Dr Keith (Dec. 2024). âReview of Fuzz Testing to Find System Vulnerabilitiesâ. In: Fuzz Testing for System Vulnerabilities | ITEA Journal 45.4. ISSN: 1054-0229. DOI:10.61278/itea. 45.4.1005 . URL:https://itea.org/journals/volume- 45- 4/review- of- fuzz- testing-to-find-system-vulnerabilities/. Jones, Erik, Anca Dragan, and Jacob Steinhardt (2024). Adversaries Can Misuse Combinations of Safe Models. URL: https://arxiv.org/abs/2406.14595. Kaber, David B. and Mica R. Endsley (Mar. 2004). âThe effects of level of automation and adaptive automation on human performance, situation awareness and workload in a dynamic control taskâ. In: Theoretical Issues in Ergonomics Science 5.2, p. 113â153. ISSN: 1463-922X, 1464-536X. DOI:10.1080/1463922021000054335. URL:http://w.tandfonline.com/doi/abs/10. 1080/1463922021000054335 (visited on 08/17/2026). Kahani, Nafiseh and Mojtaba Bagherzadeh (May 26, 2026). Testing Agentic Workflows with Structural Coverage Criteria. DOI:10.48550/arXiv.2605.26521. arXiv:2605.26521[cs.SE]. URL: http://arxiv.org/abs/2605.26521 (visited on 08/15/2026). Kaliyev, Alibek T and Artem Maryanskyy (2026). How Much Coordination Gain Is Real? A Paired Noise-Floor Protocol for Multi-Agent LLM Benchmarks. Version Number: 1. DOI:10.48550/ ARXIV.2606.20695. URL: https://arxiv.org/abs/2606.20695 (visited on 08/12/2026). Kalra, Nidhi and Susan M. Paddock (Dec. 2016). âDriving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?â In: 94, p. 182â193. ISSN: 09658564. DOI: 10.1016/j.tra.2016.09.010. URL:https://linkinghub.elsevier.com/retrieve/ pii/S0965856416302129 (visited on 08/12/2026). Kaner, C. (1998). âThe Impossibility of Complete Testingâ. In: URL:https : / / w . semanticscholar.org/paper/The- Impossibility- of- Complete- Testing- Kaner/ e74b01bd9534c660e203e481954262533139c8d7 (visited on 08/12/2026). Kapoor, Sayash et al. (July 1, 2024). AI Agents That Matter. DOI:10.48550/arXiv.2407.01502. arXiv:2407 . 01502[cs . LG]. URL:http : / / arxiv . org / abs / 2407 . 01502(visited on 08/12/2026). Kapusta, Ariel et al. (May 28, 2025). âA framework for the assurance of AI-enabled systemsâ. In: Assurance and Security for AI-enabled Systems 2025. Assurance and Security for AI-enabled Systems 2025. Ed. by Joshua D. Harguess, Nathaniel D. Bastian, and Teresa L. Pace. Orlando, United States: SPIE, p. 14. ISBN: 978-1-5106-8741-7 978-1-5106-8742-4. DOI:10.1117/12. 3056719. URL:https://w.spiedigitallibrary.org/conference-proceedings-of- spie/13476/3056719/A-framework-for-the-assurance-of-AI-enabled-systems/ 10.1117/12.3056719.full (visited on 08/20/2026). Kästner, Lena and Barnaby Crook (Dec. 2024). âExplaining AI through mechanistic interpretabilityâ. In: European Journal for Philosophy of Science 14.4, p. 52. ISSN: 1879-4912, 1879-4920. DOI: 10.1007/s13194-024-00614-4. URL:https://link.springer.com/10.1007/s13194- 024-00614-4 (visited on 08/20/2026). Katharki, Vishwanath and Sainyam Galhotra (May 26, 2026). âGoal-Oriented Reliability and Self- Improvement for Multi-Agent Systemsâ. In: Proceedings of the ACM Conference on AI and Agentic Systems. CAIS â26. New York, NY, USA: Association for Computing Machinery, p. 1367â1371. 46 ISBN: 979-8-4007-2415-2. DOI:10.1145/3786335.3813230. URL:https://dl.acm.org/ doi/10.1145/3786335.3813230 (visited on 08/15/2026). Kessler, Daniel T et al. (June 12, 2026). âNarrative Plasticity and State Stickiness: Designing Hybrid AI Systems for High-Stakes Communicationâ. In: Proceedings of the 2026 Designing Interactive Systems Conference. DIS â26. New York, NY, USA: Association for Computing Machinery, p. 1925â1943. ISBN: 979-8-4007-2563-0. DOI:10.1145/3800645.3813073. URL:https: //dl.acm.org/doi/10.1145/3800645.3813073 (visited on 08/12/2026). Khalifa, Muhammad et al. (2026). Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation. Version Number: 2. DOI:10.48550/ARXIV.2601.14691. URL:https: //arxiv.org/abs/2601.14691 (visited on 08/12/2026). Khatchadourian, Raffi (2026). Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents. Version Number: 2. DOI:10.48550/ARXIV.2601.15322. URL: https://arxiv.org/abs/2601.15322 (visited on 08/12/2026). Kim, Grace Y. E. et al. (Aug. 1, 2025). âMonitoring strategies for continuous evaluation of deployed clinical prediction modelsâ. In: Journal of Biomedical Informatics 168, p. 104854. ISSN: 1532- 0464. DOI:10.1016/j.jbi.2025.104854. URL:https://w.sciencedirect.com/ science/article/pii/S1532046425000838 (visited on 08/15/2026). Kim, Jinhan, Robert Feldt, and Shin Yoo (Apr. 30, 2023). âEvaluating Surprise Adequacy for Deep Learning System Testingâ. In: ACM Transactions on Software Engineering and Methodology 32.2, p. 1â29. ISSN: 1049-331X, 1557-7392. DOI:10.1145/3546947. URL:https://dl.acm.org/ doi/10.1145/3546947 (visited on 08/15/2026). Knack, Anna et al. (Sept. 2025). Defence AI Assurance: Identifying Promising Practice and a System Card Template for Defence. The Alan Turing Institute. URL:https://w.turing.ac.uk/ sites/default/files/2025-09/defence_ai_assurance.pdf. Koch, Christopher and Joshua Andreas Wellbrock (Apr. 18, 2026). Beyond Task Success: An Evidence- Synthesis Framework for Evaluating, Governing, and Orchestrating Agentic AI. DOI:10.48550/ arXiv.2604.19818. arXiv:2604.19818[cs.SE]. URL:http://arxiv.org/abs/2604. 19818 (visited on 08/15/2026). Kohl, Jens et al. (Dec. 2025). âAutomated Structural Testing of LLM-Based Agents: Methods, Framework, and Case Studiesâ. In: 2025 IEEE International Conference on Big Data (BigData). 2025 IEEE International Conference on Big Data (BigData). ISSN: 2573-2978, p. 1847â1856. DOI:10.1109/BigData66926.2025.11401679. URL:https://ieeexplore.ieee.org/ document/11401679 (visited on 08/18/2026). Kore, Ali et al. (Feb. 29, 2024). âEmpirical data drift detection experiments on real-world medical imaging dataâ. In: Nature Communications 15.1, p. 1887. ISSN: 2041-1723. DOI:10.1038/ s41467-024-46142-w. URL:https://w.nature.com/articles/s41467-024-46142-w (visited on 08/15/2026). Kotseruba, Iuliia and John K. Tsotsos (Jan. 1, 2020). â40 years of cognitive architectures: core cognitive abilities and practical applicationsâ. In: Artificial Intelligence Review 53.1, p. 17â94. ISSN: 1573-7462. DOI:10.1007/s10462-018-9646-y. URL:https://doi.org/10.1007/ s10462-018-9646-y (visited on 08/02/2026). Kraprayoon, Jam, Zoe Williams, and Rida Fayyaz (May 27, 2025). AI Agent Governance: A Field Guide. DOI:10.48550/arXiv.2505.21808. arXiv:2505.21808[cs.CY]. URL:http:// arxiv.org/abs/2505.21808 (visited on 08/12/2026). Kuhn, D. Richard et al. (Mar. 2013). âCombinatorial Coverage Measurement Concepts and Ap- plicationsâ. In: 2013 IEEE Sixth International Conference on Software Testing, Verification and Validation Workshops. 2013 IEEE Sixth International Conference on Software Testing, Verifica- tion and Validation Workshops, p. 352â361. DOI:10.1109/ICSTW.2013.77. URL:https: //ieeexplore.ieee.org/document/6571653 (visited on 08/12/2026). Kuhn, Richard, Dolores Wallace, and A. Gallo (June 16, 2004). âSoftware Fault Interactions and Implications for Software Testingâ. In: 6, p. 418â421. DOI:10.1109/TSE.2004.24. URL: 47 https://csrc.nist.gov/pubs/journal/2004/06/software-fault-interactions- and-implications-for-s/final (visited on 08/17/2026). Labkoff, Steven et al. (Nov. 1, 2024). âToward a responsible future: recommendations for AI-enabled clinical decision supportâ. In: Journal of the American Medical Informatics Association 31.11, p. 2730â2739. ISSN: 1067-5027, 1527-974X. DOI:10.1093/jamia/ocae209. URL:https: //academic.oup.com/jamia/article/31/11/2730/7776823 (visited on 08/17/2026). Laird, John E., Christian Lebiere, and Paul S. Rosenbloom (2017). âA Standard Model of the Mind: Toward a Common Computational Framework Across Artificial Intelligence, Cognitive Science, Neuroscience, and Roboticsâ. In: AI Magazine 38.4, p. 13â26. DOI:https://doi.org/10. 1609/aimag.v38i4.2744. URL:https://onlinelibrary.wiley.com/doi/abs/10.1609/ aimag.v38i4.2744. Lanus, Erin et al. (June 14, 2021). âTest and Evaluation Framework for Multi-Agent Systems of Autonomous Intelligent Agentsâ. In: 2021 16th International Conference of System of Systems Engineering (SoSE), p. 203â209. DOI:10.1109/SOSE52739.2021.9497472. arXiv:2101. 10430[eess.SY]. URL: http://arxiv.org/abs/2101.10430 (visited on 08/01/2026). Lee, Donghyun and Mo Tiwari (2024). Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems. Version Number: 1. DOI:10.48550/ARXIV.2410.07283. URL:https: //arxiv.org/abs/2410.07283 (visited on 08/12/2026). Lee, J. D. and K. A. See (Jan. 1, 2004). âTrust in Automation: Designing for Appropriate Relianceâ. In: Human Factors: The Journal of the Human Factors and Ergonomics Society 46.1, p. 50â80. ISSN: 0018-7208. DOI:10.1518/hfes.46.1.50_30392. URL:http://hfs.sagepub.com/ cgi/doi/10.1518/hfes.46.1.50_30392 (visited on 08/17/2026). Lee, Sung Une et al. (Mar. 6, 2026). A Structured Approach to Safety Case Construction for AI Systems. DOI:10.48550/arXiv.2601.22773. arXiv:2601.22773[cs.SE]. URL:http: //arxiv.org/abs/2601.22773 (visited on 08/18/2026). Lesire, Charles et al. (Oct. 23, 2022). âA Hierarchical Deliberative Architecture Framework based on Goal Decompositionâ. In: 2022 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS). 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Kyoto, Japan: IEEE, p. 9865â9870. ISBN: 978-1-6654-7927-1. DOI:10.1109/IROS47612. 2022.9981488. URL:https://ieeexplore.ieee.org/document/9981488/(visited on 08/15/2026). Li, Miles Q. et al. (May 10, 2026). A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents. DOI:10.48550/arXiv.2512.20798. arXiv:2512.20798[cs.AI]. URL: http://arxiv.org/abs/2512.20798 (visited on 08/15/2026). Li, Xiaoyang et al. (July 28, 2026). MemTX: Transactional Belief Commit for Stateful Agent Memory. DOI:10.48550/arXiv.2607.23929. arXiv:2607.23929[cs.AI]. URL:http://arxiv.org/ abs/2607.23929 (visited on 08/17/2026). Li, Ying et al. (June 26, 2026). VIGIL: Runtime Enforcement of Behavioral Specifications in AI Agent Skills. Version Number: 1. DOI:10.48550/ARXIV.2606.26524. URL:https://arxiv.org/ abs/2606.26524 (visited on 08/12/2026). Li, Yunfan et al. (Jan. 12, 2026). Beyond Entangled Planning: Task-Decoupled Planning for Long- Horizon Agents. DOI:10.48550/arXiv.2601.07577. arXiv:2601.07577[cs.AI]. URL: http://arxiv.org/abs/2601.07577 (visited on 08/15/2026). Liu, Shuyang et al. (Apr. 10, 2026). âProcess-Centric Analysis of Agentic Software Systemsâ. In: Proceedings of the ACM on Programming Languages 10 (OOPSLA1), p. 1961â1988. ISSN: 2475-1421. DOI:10.1145/3798271. arXiv:2512.02393[cs.SE]. URL:http://arxiv.org/ abs/2512.02393 (visited on 08/15/2026). Lonetti, Francesca, Antonia Bertolino, and Felicita Di Giandomenico (Dec. 1, 2023). âModel-based security testing in IoT systems: A Rapid Reviewâ. In: Information and Software Technology 164, p. 107326. ISSN: 0950-5849. DOI:10.1016/j.infsof.2023.107326. URL:https: 48 //w.sciencedirect.com/science/article/pii/S0950584923001817(visited on 08/15/2026). LĂš, Xing Han et al. (2025). AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. Version Number: 2. DOI:10.48550/ARXIV.2504.08942. URL:https://arxiv. org/abs/2504.08942 (visited on 08/12/2026). Luther, Ben (June 2026). âDefining T&E as a Disciplineâ. In: ITEA Journal 47.2. ISSN: 1054-0229. DOI:10.61278/itea.47.2.1004. URL:https://itea.org/journals/volume- 47- 2/defining-t-and-e-as-a-discipline/. Maderamitla, Prasad and Subba Rao Katragadda (Feb. 17, 2026). âA Deterministic Trajectory-Level Evaluation Framework for Learning-Based Agentic Systemsâ. In: Journal of Economics, Finance and Accounting Studies 8.3, p. 25â29. ISSN: 2709-0809. DOI:10.32996/jefas.2026.8.3.3. URL:https://al- kindipublisher.com/index.php/jefas/article/view/12127 (visited on 08/15/2026). Maharana, Adyasha et al. (Feb. 27, 2024). Evaluating Very Long-Term Conversational Memory of LLM Agents. DOI:10.48550/arXiv.2402.17753. arXiv:2402.17753[cs.CL]. URL: http://arxiv.org/abs/2402.17753 (visited on 08/16/2026). Manheim, David and Aidan Homewood (2025). Limits of Safe AI Deployment: Differentiating Oversight and Control. Version Number: 2. DOI:10.48550/ARXIV.2507.03525. URL:https: //arxiv.org/abs/2507.03525 (visited on 08/18/2026). Mazumder, Aritra and Nusrat jahan Lia (July 16, 2026). AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP. Version Number: 3. DOI:10.48550/ARXIV.2607.11098. URL: https://arxiv.org/abs/2607.11098 (visited on 08/12/2026). Mehri, Shuhaib et al. (Mar. 8, 2026). Goal Alignment in LLM-Based User Simulators for Con- versational AI. DOI:10.48550/arXiv.2507.20152. arXiv:2507.20152[cs.CL]. URL: http://arxiv.org/abs/2507.20152 (visited on 08/15/2026). Meyer, Bertrand (Oct. 1, 1992). âApplying "Design by Contract"â. In: Computer 25.10, p. 40â51. ISSN: 0018-9162. DOI:10.1109/2.161279. URL:https://doi.org/10.1109/2.161279 (visited on 08/12/2026). â(Aug. 2008). âSeven Principles of Software Testingâ. In: Computer 41.8, p. 99â101. ISSN: 1558- 0814. DOI:10.1109/MC.2008.306. URL:https://ieeexplore.ieee.org/document/ 4597151 (visited on 08/12/2026). Michelakis, Panagiotis, Yiannis Hadjiyiannis, and Dimitrios Stamoulis (Sept. 25, 2025). CORE: Full- Path Evaluation of LLM Agents Beyond Final State. DOI:10.48550/arXiv.2509.20998. arXiv: 2509.20998[cs.AI]. URL: http://arxiv.org/abs/2509.20998 (visited on 08/15/2026). Miculicich, Lesly et al. (Oct. 3, 2025). VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation. DOI:10.48550/arXiv.2510.05156. arXiv:2510.05156[cs.SE]. URL:http: //arxiv.org/abs/2510.05156 (visited on 08/14/2026). Miller, Tymoteusz and Irmina Durlik (Jan. 2026). âWhen Models Fail: Trustworthy Anomaly De- tection Under Distributional Drift via Dual-Layer Monitoring of Data and AI Behaviourâ. In: Applied Sciences 16.11, p. 5293. ISSN: 2076-3417. DOI:10.3390/app16115293. URL:https: //w.mdpi.com/2076-3417/16/11/5293 (visited on 08/15/2026). Molina, Facundo, Alessandra Gorla, and Marcelo dâAmorim (June 30, 2025). âTest Oracle Au- tomation in the Era of LLMsâ. In: ACM Transactions on Software Engineering and Method- ology 34.5, p. 1â24. ISSN: 1049-331X, 1557-7392. DOI:10.1145/3715107. URL:https: //dl.acm.org/doi/10.1145/3715107 (visited on 08/12/2026). Moshkovich, Dany et al. (2025). Beyond Black-Box Benchmarking: Observability, Analytics, and Optimization of Agentic Systems. Version Number: 1. DOI:10.48550/ARXIV.2503.06745. URL: https://arxiv.org/abs/2503.06745 (visited on 08/12/2026). Mudasiru, Rasheed (July 21, 2026). Deterministic Replay for AI Agent Systems. Version Number: 1. DOI:10.48550/ARXIV.2607.16200. URL:https://arxiv.org/abs/2607.16200(visited on 08/12/2026). 49 Mukhopadhyay, Debayan, Utshab Kumar Ghosh, and Shubham Chatterjee (July 16, 2026). Bridge Evidence: Static Retrieval Utility Does Not Predict Causal Utility in Multi-Step Agentic Search. DOI:10.48550/arXiv.2607.15253. arXiv:2607.15253[cs.IR]. URL:http://arxiv.org/ abs/2607.15253 (visited on 08/15/2026). Nasser, George, Ben W. Morrison, and Mark W. Wiggins (June 13, 2025). âMental models and driver takeover in automated vehicles: a systematic reviewâ. In: Ergonomics, p. 1â14. ISSN: 0014-0139, 1366-5847. DOI:10.1080/00140139.2025.2514599. URL:https://w.tandfonline. com/doi/full/10.1080/00140139.2025.2514599 (visited on 08/18/2026). NATO Allied Command Transformation (2021). NATO Warfighting Capstone Concept. NATO. URL: https://w.act.nato.int/wp- content/uploads/2023/06/NWCC- Glossy- 18- MAY.pdf. Okamura, Kazuo and Seiji Yamada (Feb. 21, 2020). âAdaptive trust calibration for human-AI collaborationâ. In: PLOS ONE 15.2. Ed. by Chen Lv, e0229132. ISSN: 1932-6203. DOI:10.1371/ journal.pone.0229132. URL:https://dx.plos.org/10.1371/journal.pone.0229132 (visited on 08/18/2026). Onnasch, Linda et al. (May 2014). âHuman Performance Consequences of Stages and Levels of Automation: An Integrated Meta-Analysisâ. In: Human Factors: The Journal of the Hu- man Factors and Ergonomics Society 56.3, p. 476â488. ISSN: 0018-7208, 1547-8181. DOI: 10.1177/0018720813501549. URL:https://journals.sagepub.com/doi/10.1177/ 0018720813501549 (visited on 08/17/2026). OpenAI (July 21, 2026). OpenAI and Hugging Face partner to address security incident during model evaluation. OpenAI. URL:https://openai.com/index/hugging- face- model- evaluation-security-incident/ (visited on 08/12/2026). Orogat, Abdelghny and Ana Rostam (Feb. 3, 2026). Understanding Multi-Agent LLM Frameworks: A Unified Benchmark and Experimental Analysis. DOI: 10.48550/arXiv.2602.03128. Osinga, Frans (2007). Science, strategy and war: the strategic theory of John Boyd. Strategy and history 18. London New York: Routledge. 1 p. ISBN: 978-0-203-08886-9. Pandey, Mukund (May 2, 2026). Evaluating Agentic AI in the Wild: Failure Modes, Drift Patterns, and a Production Evaluation Framework. DOI:10.48550/arXiv.2605.01604. arXiv:2605. 01604[cs.AI]. URL: http://arxiv.org/abs/2605.01604 (visited on 08/15/2026). Parasuraman, R., T.B. Sheridan, and C.D. Wickens (May 2000). âA model for types and levels of human interaction with automationâ. In: IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans 30.3, p. 286â297. ISSN: 1558-2426. DOI:10.1109/3468.844354. URL: https://ieeexplore.ieee.org/document/844354 (visited on 08/12/2026). Parasuraman, Raja and Dietrich H. Manzey (June 2010). âComplacency and Bias in Human Use of Automation: An Attentional Integrationâ. In: Human Factors: The Journal of the Hu- man Factors and Ergonomics Society 52.3, p. 381â410. ISSN: 0018-7208, 1547-8181. DOI: 10.1177/0018720810376055. URL:https://journals.sagepub.com/doi/10.1177/ 0018720810376055 (visited on 08/12/2026). Patchipala, Surya Gangadhar (Dec. 30, 2023). âTackling data and model drift in AI: Strategies for maintaining accuracy during ML model inferenceâ. In: International Journal of Science and Research Archive 10.2, p. 1198â1209. ISSN: 25828185. DOI:10.30574/ijsra.2023.10.2. 0855. URL: https://ijsra.net/node/2244 (visited on 08/15/2026). Pateria, Shubham (2022). âMethods for autonomously decomposing and performing long-horizon sequential decision tasksâ. PhD thesis. Nanyang Technological University. DOI:10.32657/ 10356/155182. URL: https://hdl.handle.net/10356/155182 (visited on 08/15/2026). Qi, Jinhu et al. (Apr. 29, 2026). âTowards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system securityâ. In: Academia AI and Applications 2.2. ISSN: 3071-0286. DOI:10.20935/AcadAI8260. URL:https://w.academia.edu/166114173/Towards_ trustworthy _ agentic _ AI _ a _ comprehensive _ survey _ of _ safety _ robustness _ privacy_and_system_security (visited on 08/15/2026). 50 Rabanser, Stephan et al. (June 2, 2026). Towards a Science of AI Agent Reliability. DOI:10.48550/ arXiv.2602.16666 . arXiv:2602.16666[cs.AI]. URL:http://arxiv.org/abs/2602. 16666 (visited on 08/16/2026). Rana, Manik et al. (Oct. 20, 2025). AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI. DOI:10.48550/arXiv.2510.18170. arXiv: 2510.18170[cs.AI]. URL: http://arxiv.org/abs/2510.18170 (visited on 08/15/2026). Rath, Abhishek (Jan. 8, 2026). âAgent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactionsâ. In: URL: https://arxiv.org/abs/2601.04170. Reid, Alistair et al. (Aug. 6, 2025). âRisk Analysis Techniques for Governed LLM-based Multi-Agent Systemsâ. In: arXiv preprint. DOI:10.48550/arXiv.2508.05687. URL:https://arxiv.org/ abs/2508.05687. Rezazadeh, Alireza et al. (2025). Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control. Version Number: 1. DOI:10.48550/ARXIV.2505.18279. URL: https://arxiv.org/abs/2505.18279 (visited on 08/12/2026). Riccio, Vincenzo et al. (Nov. 2020). âTesting machine learning based systems: a systematic mappingâ. In: Empirical Software Engineering 25.6, p. 5193â5254. ISSN: 1382-3256, 1573-7616. DOI: 10.1007/s10664-020-09881-0. URL:https://link.springer.com/10.1007/s10664- 020-09881-0 (visited on 08/17/2026). Richardson, Debra J., Stephanie Leif Aha, and T. Owen OâMalley (June 1, 1992). âSpecification- based test oracles for reactive systemsâ. In: Proceedings of the 14th international conference on Software engineering. ICSE â92. New York, NY, USA: Association for Computing Machinery, p. 105â118. ISBN: 978-0-89791-504-5. DOI:10.1145/143062.143100. URL:https://dl. acm.org/doi/10.1145/143062.143100 (visited on 08/14/2026). Rickli, Jean-Marc and Tobias Knappe (Apr. 2026). âThe International Security and Military Implica- tions of Agentic AIâ. In: Geneva Paper 37/26. URL:https://w.gcsp.ch/sites/default/ files/2026-04/GP-2026_37_Rickli%20Knappe_The%20International%20Security% 20and%20Military%20Implications%20of%20Agentic%20AI%3Bdigital.pdf. Riedmaier, Stefan et al. (2020). âSurvey on Scenario-Based Safety Assessment of Automated Vehiclesâ. In: IEEE Access 8, p. 87456â87477. ISSN: 2169-3536. DOI:10.1109/ACCESS. 2020.2993730. URL:https://ieeexplore.ieee.org/document/9090897/(visited on 08/12/2026). Rijn, Max van et al. (July 8, 2025). âAI in Military C2 Systems - An Introduction and Recent Advancesâ. In: URL:https://c2coe.org/download/ai-in-military-c2-systems-an- introduction-and-recent-advances-c2coe/. Romeo, Giuseppe and Daniela Conti (Jan. 2026). âExploring automation bias in humanâAI col- laboration: a review and implications for explainable AIâ. In: AI & SOCIETY 41.1, p. 259â 278. ISSN: 0951-5666, 1435-5655. DOI:10.1007/s00146- 025- 02422- 7. URL:https: //link.springer.com/10.1007/s00146-025-02422-7 (visited on 08/18/2026). Ruan, Jianhao et al. (Feb. 7, 2026). AOrchestra: Automating Sub-Agent Creation for Agentic Or- chestration. DOI:10.48550/arXiv.2602.03786. arXiv:2602.03786[cs.AI]. URL:http: //arxiv.org/abs/2602.03786 (visited on 08/15/2026). Russell, Stuart J. and Peter Norvig (2022). Artificial intelligence: A Modern Approach. In collab. with Ernest Davis and Douglas Edwards. Fourth edition, Global edition. Prentice Hall series in artificial intelligence. Boston Columbus Indianapolis: Pearson. 1 p. ISBN: 978-0-13-604259-4 978-1-292-15397-1. Safin, Damir and Dian Balta (May 12, 2026). Autonomy and Agency in Agentic AI: Architectural Tactics for Regulated Contexts. DOI:10.48550/arXiv.2605.12105. arXiv:2605.12105[cs. AI]. URL: http://arxiv.org/abs/2605.12105 (visited on 08/15/2026). Sahiner, Berkman et al. (Oct. 1, 2023). âData drift in medical machine learning: implications and potential remediesâ. In: The British Journal of Radiology 96.1150, p. 20220878. ISSN: 0007- 51 1285, 1748-880X. DOI:10.1259/bjr.20220878. URL:https://academic.oup.com/bjr/ article/doi/10.1259/bjr.20220878/7499000 (visited on 08/15/2026). Sahoo, Subramanyam et al. (Mar. 6, 2026). SAHOO: Safeguarded Alignment for High-Order Opti- mization Objectives in Recursive Self-Improvement. DOI:10.48550/arXiv.2603.06333. arXiv: 2603.06333[cs.AI]. URL: http://arxiv.org/abs/2603.06333 (visited on 08/15/2026). Sapkota, Ranjan, Konstantinos I. Roumeliotis, and Manoj Karkee (2025). AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges. Version Number: 5. DOI:10.48550/ARXIV. 2505.10468. URL: https://arxiv.org/abs/2505.10468 (visited on 08/20/2026). Seshia, Sanjit A., Dorsa Sadigh, and S. Shankar Sastry (July 2022). âToward verified artificial intelligenceâ. In: Communications of the ACM 65.7, p. 46â55. ISSN: 0001-0782, 1557-7317. DOI:10.1145/3503914. URL:https://dl.acm.org/doi/10.1145/3503914(visited on 08/07/2026). Shah, Qasim Ali et al. (Jan. 2023). âA Meta Modeling-Based Interoperability and Integration Testing Platform for IoT Systemsâ. In: Sensors 23.21, p. 8730. ISSN: 1424-8220. DOI:10.3390/ s23218730. URL:https://w.mdpi.com/1424-8220/23/21/8730(visited on 08/15/2026). Shahnovsky, Orit and Rotem Dror (Mar. 13, 2026). AI Planning Framework for LLM-Based Web Agents. DOI:10.48550/arXiv.2603.12710. arXiv:2603.12710[cs.AI]. URL:http: //arxiv.org/abs/2603.12710 (visited on 08/15/2026). Shao, Shuai et al. (May 20, 2026). MonoScale: Scaling Multi-Agent System with Monotonic Im- provement. DOI:10.48550/arXiv.2601.23219. arXiv:2601.23219[cs.MA]. URL:http: //arxiv.org/abs/2601.23219 (visited on 08/15/2026). Shavit, Yonadav G. et al. (2024). âPractices for Governing Agentic AI Systemsâ. In: URL:https:// w.semanticscholar.org/paper/Practices-for-Governing-Agentic-AI-Systems- Shavit-Agarwal/0002c42e8d7bfeafc431c4ed9f6318f223bbf58b (visited on 08/12/2026). Sheffler, Thomas J. (Aug. 19, 2025). An Approach to Checking Correctness for Agentic Systems. DOI: 10.48550/arXiv.2509.20364. arXiv:2509.20364[cs.AI]. URL:http://arxiv.org/ abs/2509.20364 (visited on 08/18/2026). Shukla, Manish A. (Aug. 25, 2025). âAdaptive Monitoring and Real-World Evaluation of Agentic AI Systemsâ. In: URL: https://arxiv.org/abs/2509.00115. Sidhu, Harleen Kaur et al. (July 7, 2026). Open Problems in AI Incident Governance. Version Number: 1. DOI:10.48550/ARXIV.2607.05163. URL:https://arxiv.org/abs/2607.05163(visited on 08/12/2026). SĂśffker, Dirk, Abderahman Bejaoui, and Olena Shyshova (July 4, 2025). âProgress towards mul- tidimensionally scalable assisted and/or automated ship navigation and control â part I: hu- man in the interaction loopâ. In: Journal of Marine Engineering & Technology 24.4, p. 294â 304. ISSN: 2046-4177, 2056-8487. DOI:10.1080/20464177.2025.2494431. URL:https: //w.tandfonline.com/doi/full/10.1080/20464177.2025.2494431(visited on 08/18/2026). Software and systems engineering â Software testing (2021). URL:https://w.iso.org/obp/ ui/en/#iso:std:iso-iec-ieee:29119:-4:ed-2:v1:en. Song, Yifan et al. (July 15, 2024). The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism. DOI:10.48550/arXiv.2407.10457. arXiv:2407.10457[cs. CL]. URL: http://arxiv.org/abs/2407.10457 (visited on 08/15/2026). South, Tobin et al. (Jan. 16, 2025). Authenticated Delegation and Authorized AI Agents. DOI:10. 48550/arXiv.2501.09674 . arXiv:2501.09674[cs.CY]. URL:http://arxiv.org/abs/ 2501.09674 (visited on 08/15/2026). Souza, Renan et al. (Sept. 15, 2025). âPROV-AGENT: Unified Provenance for Tracking AI Agent In- teractions in Agentic Workflowsâ. In: 2025 IEEE International Conference on eScience (eScience). 2025 IEEE International Conference on eScience (eScience). Chicago, IL, USA: IEEE, p. 467â 473. ISBN: 979-8-3315-9145-8. DOI:10.1109/eScience65000.2025.00093. URL:https: //ieeexplore.ieee.org/document/11181558/ (visited on 08/12/2026). 52 Stettinger, Georg, Patrick Weissensteiner, and Siddartha Khastgir (2024). âTrustworthiness Assurance Assessment for High-Risk AI-Based Systemsâ. In: IEEE Access 12, p. 22718â22745. ISSN: 2169-3536. DOI:10.1109/ACCESS.2024.3364387. URL:https://ieeexplore.ieee.org/ document/10430152/ (visited on 08/18/2026). Sun, Chen et al. (Mar. 2022). âAcclimatizing the Operational Design Domain for Autonomous Driving Systemsâ. In: IEEE Intelligent Transportation Systems Magazine 14.2, p. 10â24. ISSN: 1939-1390, 1941-1197. DOI:10.1109/MITS.2021.3070651. URL:https://ieeexplore. ieee.org/document/9440962/ (visited on 08/12/2026). Sun, Youcheng et al. (Oct. 31, 2019). âStructural Test Coverage Criteria for Deep Neural Networksâ. In: ACM Transactions on Embedded Computing Systems 18.5, p. 1â23. ISSN: 1539-9087, 1558- 3465. DOI:10.1145/3358233. URL:https://dl.acm.org/doi/10.1145/3358233(visited on 08/15/2026). Tan, Boyin, Xiaowei Huang, and Youcheng Sun (July 7, 2026). Skill Coverage: A Test Adequacy Metric for Agent Skills. Version Number: 2. DOI:10.48550/ARXIV.2606.20659. URL:https: //arxiv.org/abs/2606.20659 (visited on 08/20/2026). Tang, Bo et al. (Feb. 2023). âAI Testing Framework for Next-G O-RAN Networks: Requirements, Design, and Research Opportunitiesâ. In: IEEE Wireless Communications 30.1, p. 70â77. ISSN: 1558-0687. DOI:10.1109/MWC.001.2200213. URL:https://ieeexplore.ieee.org/ abstract/document/10077111 (visited on 08/15/2026). Tang, Bohan et al. (Nov. 2023). âCollaborative Uncertainty Benefits Multi-Agent Multi-Modal Trajectory Forecastingâ. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 45.11, p. 13297â13313. ISSN: 1939-3539. DOI:10.1109/TPAMI.2023.3290823. URL:https: //ieeexplore.ieee.org/document/10173747 (visited on 08/18/2026). Tatasciore, Monica and Shayne Loft (Dec. 2, 2025). âCalibrating Reliance on Automated Advice: Transparency and Trust Calibration Feedbackâ. In: International Journal of HumanâComputer Interaction 41.23, p. 14723â14733. ISSN: 1044-7318, 1532-7590. DOI:10.1080/10447318. 2025.2487861. URL:https://w.tandfonline.com/doi/full/10.1080/10447318. 2025.2487861 (visited on 08/18/2026). Tate, David M. et al. (2016). âA Framework for Evidence-Based Licensure of Adaptive Autonomous Systemsâ. In: URL: https://api.semanticscholar.org/CorpusID:27342401. Tejani, Ali S. et al. (June 18, 2024). âIntegrating and Adopting AI in the Radiology Workflow: A Primer for Standards and Integrating the Healthcare Enterprise (IHE) Profilesâ. In: Radiology 311.3, e232653. ISSN: 0033-8419. DOI:10.1148/radiol.232653. URL:https://pmc.ncbi. nlm.nih.gov/articles/PMC11208735/ (visited on 08/15/2026). Tellinghuisen, Joel (Mar. 22, 2001). âStatistical Error Propagationâ. In: The Journal of Physical Chemistry A 105.15, p. 3917â3921. ISSN: 1089-5639. DOI:10.1021/jp003484u. URL:https: //doi.org/10.1021/jp003484u (visited on 08/15/2026). TomaĹĄev, Nenad, Matija Franklin, and Simon Osindero (2026). Intelligent AI Delegation. URL: https://arxiv.org/abs/2602.11865. Tsamados, Andreas, Luciano Floridi, and Mariarosaria Taddeo (Apr. 2025). âHuman control of AI systems: from supervision to teamingâ. In: AI and Ethics 5.2, p. 1535â1548. ISSN: 2730-5953, 2730-5961. DOI:10.1007/s43681-024-00489-4. URL:https://link.springer.com/10. 1007/s43681-024-00489-4 (visited on 08/18/2026). Tuan, Thanh Luong and Abhijit Sanyal (2026). Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification. Version Number: 2. DOI: 10.48550/ARXIV.2606.04037. URL:https://arxiv.org/abs/2606.04037(visited on 08/12/2026). U.S. DoW OUSW(R&E) (2025). Developmental Test, Evaluation, and Assessments (DTE&A). URL: https://w.cto.mil/dtea/ (visited on 08/12/2026). UK Ministry of Defence (Jan. 30, 2026). Test and Evaluation (T&E): Future Advantage Through Evaluation (FATE). GOV.UK. URL:https://w.gov.uk/government/publications/ 53 defence- test- and- evaluation- future- advantage- through- evaluation- fate/ test-and-evaluation-te-future-advantage-through-evaluation-fate (visited on 08/12/2026). Valle, Pablo et al. (June 23, 2026). MANGO: Automated Multi-Agent Test Oracle Generation for Vision-Language-Action Models. DOI:10.48550/arXiv.2606.24815. arXiv:2606.24815[cs. SE]. URL: http://arxiv.org/abs/2606.24815 (visited on 08/15/2026). Van De Merwe, Koen, Steven Mallam, and Salman Nazir (Jan. 2024). âAgent Transparency, Situation Awareness, Mental Workload, and Operator Performance: A Systematic Literature Reviewâ. In: Human Factors: The Journal of the Human Factors and Ergonomics Society 66.1, p. 180â208. ISSN: 0018-7208, 1547-8181. DOI:10.1177/00187208221077804. URL:https://journals. sagepub.com/doi/10.1177/00187208221077804 (visited on 08/18/2026). Van Den Broek, H. and J. Van Der Waa (July 1, 2022). âIntelligent Operator Support Concepts for Shore Control Centresâ. In: Journal of Physics: Conference Series 2311.1, p. 012032. ISSN: 1742- 6588, 1742-6596. DOI:10.1088/1742-6596/2311/1/012032. URL:https://iopscience. iop.org/article/10.1088/1742-6596/2311/1/012032 (visited on 08/18/2026). Van Der Kleij, Rick, Tom Hueting, and Jan Maarten Schraagen (July 2018). âChange detection support for supervisory controllers of highly automated systems: Effects on performance, mental workload, and recovery of situation awareness following interruptionsâ. In: International Journal of Industrial Ergonomics 66, p. 75â84. ISSN: 01698141. DOI:10.1016/j.ergon.2018.02.010. URL:https://linkinghub.elsevier.com/retrieve/pii/S0169814117300136(visited on 08/18/2026). Van Der Vorst, Joris P et al. (July 2025). âImportance of model governance in clinical AI models: case study on the relevance of data drift detectionâ. In: BMJ Digital Health & AI 1.1, e000046. ISSN: 3049-575X. DOI:10.1136/bmjdhai-2025-000046. URL:https://bmjdigitalhealth.bmj. com/lookup/doi/10.1136/bmjdhai-2025-000046 (visited on 08/15/2026). Vatsal, Shubham, Harsh Dubey, and Aditi Singh (2026). âAgentic AI in Healthcare and Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-Based Agentsâ. In: IEEE Access 14, p. 4840â4863. ISSN: 2169-3536. DOI:10.1109/ACCESS.2026.3651218. URL: https://ieeexplore.ieee.org/document/11329025/ (visited on 08/12/2026). Veitch, Erik et al. (May 2024). âHuman factor influences on supervisory control of remotely operated and autonomous vesselsâ. In: Ocean Engineering 299, p. 117257. ISSN: 00298018. DOI:10.1016/ j.oceaneng.2024.117257. URL:https://linkinghub.elsevier.com/retrieve/pii/ S0029801824005948 (visited on 08/17/2026). Wang, Charles L. et al. (2025). MI9: An Integrated Runtime Governance Framework for Agentic AI. Version Number: 4. DOI:10.48550/ARXIV.2508.03858. URL:https://arxiv.org/abs/ 2508.03858 (visited on 08/12/2026). Wang, Haoyu, Christopher M. Poskitt, and Jun Sun (2025). AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents. Version Number: 3. DOI:10.48550/ARXIV. 2503.18666. URL: https://arxiv.org/abs/2503.18666 (visited on 08/12/2026). Wang, Xinyu Jessica et al. (Apr. 13, 2026). The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break. DOI:10.48550/arXiv.2604.11978. arXiv:2604.11978[cs.AI]. URL: http://arxiv.org/abs/2604.11978 (visited on 08/10/2026). Wang, Yiqi et al. (June 28, 2026). From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents. DOI:10.48550/arXiv.2606.04990. arXiv:2606. 04990[cs.CR]. URL: http://arxiv.org/abs/2606.04990 (visited on 08/14/2026). Weissensteiner, Patrick et al. (2023). âOperational Design Domain-Driven Coverage for the Safety Argumentation of Automated Vehiclesâ. In: IEEE Access 11, p. 12263â12284. ISSN: 2169-3536. DOI:10.1109/ACCESS.2023.3242127. URL:https://ieeexplore.ieee.org/document/ 10036064 (visited on 08/12/2026). 54 Weng, Shihao, Yang Feng, and Xiaofei Xie (2026). Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges. Version Number: 1. DOI:10.48550/ARXIV.2605.06161. URL: https://arxiv.org/abs/2605.06161 (visited on 08/12/2026). Wiedermann, Wolfgang and Dexin Shi (2026). âCumulant-Based Approaches for Testing the Assumption of Independent Errors in Non-Gaussian Parallel and Congeneric Measuresâ. In: Educational and Psychological Measurement 0.0, p. 00131644261444671. DOI:10 . 1177 / 00131644261444671. URL: https://doi.org/10.1177/00131644261444671. Wiesmeier, Lorenz et al. (May 13, 2026). âAdversarial robustness of LLM-based multi-agent sys- tems for engineering problemsâ. In: Frontiers in Artificial Intelligence 9. ISSN: 2624-8212. DOI: 10 . 3389 / frai . 2026 . 1784484. URL:https : / / w . frontiersin . org / journals / artificial- intelligence/articles/10.3389/frai.2026.1784484/full (visited on 08/12/2026). Wojton, Heather M, Daniel J Porter, and John W Dennis (2020). âTest & Evaluation of AI-enabled and Autonomous Systems: A Literature Reviewâ. In: URL:https://testscience.org/wp- content/uploads/formidable/20/Autonomy-Lit-Review.pdf. Wood, Nathan Gabriel et al. (June 2026). âStop Saying "AI"â. In: Philosophy & Technology 39.2, p. 106. ISSN: 2210-5433, 2210-5441. DOI:10.1007/s13347-026-01114-4. URL:https: //link.springer.com/10.1007/s13347-026-01114-4 (visited on 08/18/2026). Wooldridge, Michael and Nicholas R. Jennings (1995). âIntelligent agents: theory and practiceâ. In: The Knowledge Engineering Review 10.2. Edition: 2009/07/07, p. 115â152. ISSN: 0269- 8889. DOI:10.1017/S0269888900008122. URL:https://w.cambridge.org/product/ CF2A6AAEEA1DBD486EF019F6217F1597. Wu, Di et al. (2024). LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. Version Number: 2. DOI:10.48550/ARXIV.2410.10813. URL:https://arxiv.org/ abs/2410.10813 (visited on 08/12/2026). Xie, Yi et al. (June 13, 2026). Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems. DOI:10.48550/arXiv.2606.15242. arXiv:2606.15242[cs.CR]. URL: http://arxiv.org/abs/2606.15242 (visited on 08/19/2026). Xu, Bin (Jan. 5, 2026). AI Agent Systems: Architectures, Applications, and Evaluation. DOI:10. 48550/arXiv.2601.01743. arXiv:2601.01743[cs.AI]. URL:http://arxiv.org/abs/ 2601.01743 (visited on 08/15/2026). Xu, Cheng et al. (Feb. 2026). âSubgoal-Based Hierarchical Reinforcement Learning for Multiagent Collaborationâ. In: IEEE Transactions on Systems, Man, and Cybernetics: Systems 56.2, p. 1203â 1215. ISSN: 2168-2232. DOI:10.1109/TSMC.2025.3646451. URL:https://ieeexplore. ieee.org/document/11333884 (visited on 08/15/2026). Xu, Qinghua et al. (Mar. 25, 2026). âHallucination to Consensus: Multi-Agent LLMs for End-to- End JUnit Test Generationâ. In: ACM Transactions on Software Engineering and Methodology, p. 3803418. ISSN: 1049-331X, 1557-7392. DOI:10.1145/3803418. arXiv:2506.02943[cs.SE]. URL: http://arxiv.org/abs/2506.02943 (visited on 08/16/2026). Yan, Hang et al. (Feb. 3, 2026). TIDE: Trajectory-based Diagnostic Evaluation of Test-Time Improve- ment in LLM Agents. DOI:10.48550/arXiv.2602.02196. arXiv:2602.02196[cs.AI]. URL: http://arxiv.org/abs/2602.02196 (visited on 08/18/2026). Yang, X. Jessie, Christopher Schemanske, and Christine Searle (Aug. 1, 2023). âToward Quanti- fying Trust Dynamics: How People Adjust Their Trust After Moment-to-Moment Interaction With Automationâ. In: Human Factors 65.5, p. 862â878. ISSN: 0018-7208. DOI:10.1177/ 00187208211034716. URL:https://doi.org/10.1177/00187208211034716(visited on 08/18/2026). Ye, Bowen et al. (May 7, 2026). Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents. DOI:10.48550/arXiv.2604.06132. arXiv:2604.06132[cs.AI]. URL:http: //arxiv.org/abs/2604.06132 (visited on 08/15/2026). 55 Yehudai, Asaf et al. (2025). Survey on Evaluation of LLM-based Agents. Version Number: 2. DOI: 10.48550/ARXIV.2503.16416. URL:https://arxiv.org/abs/2503.16416(visited on 08/12/2026). Zednik, Carlos (June 2021). âSolving the Black Box Problem: A Normative Framework for Explain- able Artificial Intelligenceâ. In: Philosophy & Technology 34.2, p. 265â288. ISSN: 2210-5433, 2210-5441. DOI:10.1007/s13347-019-00382-7. URL:https://link.springer.com/10. 1007/s13347-019-00382-7 (visited on 08/18/2026). Zerilli, John, Umang Bhatt, and Adrian Weller (Apr. 8, 2022). âHow transparency modulates trust in artificial intelligenceâ. In: Patterns 3.4. ISSN: 2666-3899. DOI:10.1016/j.patter.2022. 100455. URL:https://w.cell.com/patterns/abstract/S2666-3899(22)00028-9 (visited on 08/18/2026). Zhan, Qiusi et al. (Aug. 4, 2024). InjecAgent: Benchmarking Indirect Prompt Injections in Tool- Integrated Large Language Model Agents. DOI:10.48550/arXiv.2403.02691. arXiv:2403. 02691[cs.CL]. URL: http://arxiv.org/abs/2403.02691 (visited on 08/15/2026). Zhang, Wentao et al. (2025). AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool- Environment-Agent (TEA) Protocol. Version Number: 6. DOI:10.48550/ARXIV.2506.12508. URL: https://arxiv.org/abs/2506.12508 (visited on 08/20/2026). Zhao, Yujie et al. (May 27, 2026). AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications. DOI:10.48550/arXiv.2602.22769. arXiv:2602.22769[cs.AI]. URL:http: //arxiv.org/abs/2602.22769 (visited on 08/15/2026). Zhou, Tianxing et al. (July 16, 2025). STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner. DOI:10.48550/arXiv.2506.21030. arXiv: 2506.21030[cs.RO]. URL: http://arxiv.org/abs/2506.21030 (visited on 08/15/2026). Zhuge, Mingchen et al. (2024). Agent-as-a-Judge: Evaluate Agents with Agents. Version Number: 2. DOI:10.48550/ARXIV.2410.10934. URL:https://arxiv.org/abs/2410.10934(visited on 08/12/2026). 56 A Test and Evaluation dimensions in scope DimensionDescription D1Functional performanceWhether the system does what it is built to do, on representative tasks D2Non-functional performance Latency, throughput, cost, resource use and interoperability under operational load D3RobustnessBehavior under distribution shift, perturbation, degraded inputs and adversarial conditions D4Behavioral stabilityConsistency across runs, sessions and updates; trajectory variance under similar inputs D5Differential performanceVariation in outcomes across subgroups, scenarios or operating condi- tions D6SecurityResistance to attack on the models, agents, tool chains and memory layer D7SafetyAvoidance of behavior with consequences outside the agreed envelope D8Human-machine teamingPerformance of the operator-and-system pair, including delegation, trust and override Table 10: Test and Evaluation dimensions in scope 57