Source-linked AI summary

Testing and Evaluation of Agentic AI Systems In Military Command and Control

Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan

arXiv:2608.20597v1cs.SEcs.AIcs.CY

TL;DR

Agentic AI is entering military C2 under commitments to rigorous testing and oversight, but it remains unclear how current T&E evidence supports claims about fielded behavior. The paper reviews documented practices, identifies strained assumptions, and assesses recovery methods, finding that broad system-level claims are unsupported while narrower claims and deployment governance remain possible.

  • Problem

    The paper addresses how much confidence current T&E methods can justify for agentic C2 systems and how residual uncertainty should be managed in fielding decisions.

  • Method

    The paper reviews documented T&E practices, maps agentic properties to eight strained assumptions, and assesses recovery strategies for specifiability, stability, and composability.

  • Results

    Agentic properties challenge eight foundational T&E assumptions, weakening the argument that connects test evidence to claims about fielded behavior.

  • Takeaways & Limitations

    Narrower claims can remain recoverable through bounded envelopes, trajectory-grounded correctness, runtime constraints, and variance characterization, while residual uncertainty is governed through deployment mechanisms.

  • Takeaways & Limitations

    The findings describe the publicly documented record, while classified and tacit practices may address some identified gaps; composability evidence also faces impractical coverage and noise limits.

Abstract

from arXiv · show

Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.

6 Future Ethics Lab

The paper identifies how agentic properties strain established T&E assumptions and outlines narrower assurance claims and governance mechanisms for C2 fielding decisions.

  • Key findings: Agentic systems make trajectories involving tool use, memory, and delegation the relevant behavioral unit rather than discrete outputs.Runtime emergence, accumulated state, and changing objectives create differences between tested and deployed configurations.
  • Assurance claims: Narrower assurance claims remain recoverable in principle through bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance.The paper frames these methods as contingent on maturity rather than as established guarantees.
  • Key findings: Eight assumptions underlying established T&E practice are strained by agentic properties, spanning system specifiability, stability, composability, and supervisability.The assumptions were identified from 240 documented practices.
  • Assurance claims: Ten assurance claims address specifiability, stability, and composability, while supervisability is carried forward only through the human-direction interface.Supervisability is not assessed because it requires human-subjects methods and depends on stability results.
  • Framework: The paper maps agentic properties to testing assumptions and assurance claims for evaluating confidence in military C2 systems.Figure 1 presents the mapping underlying the analysis.

AI use statement

Language models supported literature review, data organization, graphing, and writing assistance during preparation of the paper.

  • AI use statement: Language models helped scan scientific databases, conduct exploratory reviews, identify supplementary articles, organize extracted data, prepare graphs, and review prose.The stated uses also included grammar and spelling support and figure preparation.

1 Introduction

The paper asks how much confidence current T&E methods can justify for agentic AI in C2 and how residual uncertainty should inform fielding decisions. It frames the issue as a gap between tested evidence and fielded behavior across model, agent, and deployed-system levels.

  • Research questions: The paper examines whether current evaluation methods can justify confidence in agentic C2 systems and how residual uncertainty should be addressed in fielding decisions.These are the paper’s two organizing questions.
  • Analytical scope: Agentic C2 systems are analyzed at model, agent, and deployed-system levels, where scaffolding, operators, organizational context, and multi-agent assemblies add complexity.The agent level includes memory, tools, and orchestration; the deployed level includes agents, operators, and organizational context.
  • T&E challenge: Agentic systems weaken the bounded-behavior assumptions needed for sampling because they autonomously pursue objectives, select information sources, use tools, and retain context across runs.Established methods require the possible behavior space to be partitionable, parameterizable, or otherwise bounded enough for samples to represent it.
  • T&E challenge: At the deployed-system level, emergent assembly behavior and changes in human information flows can alter outcomes without changing the underlying technology.Agent-to-agent delegation adds further complexity to system behavior and attribution.
  • Assurance framework: Assurance cases require claims about acceptability, evidence bearing on those claims, and an argument connecting evidence gathered under test conditions to the fielded system.The paper’s principal concern is the connecting argument.
  • Scope and approach: The analysis focuses on agentic systems in C2, excludes narrow rule-based and single-turn systems without retained context, and assesses new methods or governance mechanisms against established T&E practice.The assessment covers specifiability, stability, and composability; supervisability is only carried through the human-direction interface.

2 Definitions and Concepts

Agentic AI is described through its faculties, functions, properties, and architecture, then situated in command and control through effectiveness criteria and five illustrative scenarios.

  • Agentic AI: Agentic systems pursue assigned goals without explicit procedural instructions, act independently of human mediation, and adapt to unforeseen circumstances.
  • Agentic AI: Agents perceive environments, make decisions, act through actuators, and learn from experience.
  • Agentic AI: The generic architecture links operator interfaces, orchestrators, models, memory, and data to the realization of agent functions.
  • Agentic AI: Seven agentic properties operationalize defining features and map onto functions and components relevant to testing and evaluation.
  • Agentic AI: Expanding integration enlarges the reachable action space, vulnerability surface, and testing boundary.
  • Command and Control: Command and control separates setting intent from modifying conditions as situations evolve, without specifying who or what performs each function.
  • Command and Control: The paper examines five C2 activities: maritime operating-picture maintenance, adaptive planning, multi-echelon coordination, coalition interoperability, and force allocation.

3 Methodology

The study uses a structured qualitative gap analysis of documented T&E practices and candidate assurance responses. It identifies challenged assumptions, derives assurance claims, assesses methods, and records limitations of the documented evidence base.

  • Methodology: The study analyzes publicly documented T&E practices against a defined set of agentic system properties to answer its two research questions.
  • Methodology: The literature review assembled a T&E practice corpus from 26 documents covering testing, assurance, or certification of AI-enabled or autonomous systems.
  • Methodology: A second literature-review strand searched for candidate responses to assurance claims across academic databases and adjacent assurance domains.
  • Methodology: A dozen expert interviews lasting 30 to 60 minutes addressed current T&E practice, agentic limitations, and in-service monitoring and reaccreditation.
  • Methodology: The analysis extracted 240 T&E practices across eight dimensions and three lifecycle stages, alongside seven agentic properties and elicited assumptions.
  • Methodology: For each assumption, the gap matrix records assessment approaches, behavioral presuppositions, violated assumptions, and C2 operational consequences.
  • Limitations: The findings are limited by public-document coverage, predominantly inferred strains without empirical testing, and untested transfer from adjacent domains to C2.

4 Agentic AI Challenges Testing and Evaluation Assumptions

Agentic properties weaken established T&E assumptions about specifiability, stability, and composability, making tested behavior a less reliable guide to fielded behavior. The resulting challenges require trajectory-focused testing, lifecycle evidence management, and system-level characterization.

  • Specifiability: Agentic systems shift the behavioral unit from discrete outputs to trajectories through tools, memory, delegation, and environment interaction.Trajectory structure becomes part of correctness, while reachable behavior expands with task length and path-dependent state.
  • Specifiability: Retrieved content can act as both data and instruction, allowing indirect prompt injection to alter application behavior and downstream use.Syntax-only input checks cannot fully distinguish trusted instructions from untrusted retrieved content.
  • Composability: Runtime-added agents, tools, models, and orchestration patterns can invalidate pre-deployment interoperability certificates unless capability changes are fenced, recorded, and governed.The effective integration surface is partly realized during operation, so a certified configuration may not remain the deployed configuration.
  • Stability: Persistent memory undermines evidence currency because the system can accumulate state even when its nominal model or software version remains unchanged.Freezing state would compromise intended operational capability, while unchanged configuration control may fail to trigger revalidation.
  • Stability: Current T&E lacks a general criterion for goal drift or goal maintenance, and proposed constructs remain early and non-standardized.Distinguishing acceptable adaptation from optimization based on an incorrect interpretation of intent remains an open challenge.
  • Composability: Seven out of ten recent coordination architectures report effects below the run-to-run noise threshold, making variability characterization a precondition for composability claims.Intra-system emergence, joint inconsistency, propagation, and configuration effects cannot be adequately verified by simply progressing from component to system testing.

5 Closing the Assurance Gaps

The paper examines how Testing and Evaluation methods can recover assurance for agentic C2 systems after eight established assumptions are challenged. It identifies narrower assurance claims and methods, while marking important limits around adversarial operational spaces, judge validity, deployment drift, composition, and supervisability.

  • Scope of the assurance-gap analysis: The analysis surveys methods for specifiability, stability, and multi-agent composition, while excluding supervisability beyond the human-direction interface.Acceptability remains the fielding authority’s responsibility; the analysis addresses whether properties can be formally characterized.
  • 5.1 Specifiability: Agentic systems challenge advance characterization because they construct actions, retrieve unforeseen inputs, and may expand their tools, platforms, or subagents after deployment.These properties make the boundary of the certified configuration difficult to maintain.
  • 5.1 Specifiability: A bounded mission and behavioral envelope can anchor C2 assurance, but contested and adversarial operational spaces resist the specification available in more structured domains.Structured sampling of a declared operational space is proposed, with future work needed to derive regulatory, operational, and adversarial scenarios.
  • 5.1 Specifiability: Trajectory-grounded correctness remains limited because model-based judges vary across task types, rubrics, and manipulated reasoning traces, making validation necessary before certification use.The evidence supports improved and stratified oracles, but not exhaustive or unambiguous specification of open-ended mission reasoning.
  • 5.2 Stability: Stability assurance requires detecting behavioral drift, attributing changes to accumulated state, and correcting or revoking mission-relevant state through adjudication, propagation, and bounded renewal.Operational diagnosis may require provenance tagging for each memory item at creation, which is infrequently implemented.
  • 5.3 Multi-agent composition and emergence: Multi-agent composition requires assembly-level evidence showing that component behavior remains valid after integration and that departures are measurable, bounded, and attributable.Runtime-generated inter-agent interfaces undermine staged-test progression, while emergent behavior may be beneficial or detrimental and must be shown reliable or unlikely, respectively.

6 Conclusion

The analysis finds that current T&E methods cannot broadly substantiate system-level behavior for agentic C2 systems, while narrower assurance claims remain possible with mature methods and continued governance of residual uncertainty.

  • 6 Conclusion: Current T&E methods substantiate confidence only within limits because agentic properties challenge foundational assumptions about specifiability, stability, composability, and supervisability.Supervisability was identified but only partly assessed in the recovery analysis.
  • 6 Conclusion: The conclusions are restricted to publicly available US, UK, and NATO practices, and the inferred challenges were not empirically demonstrated through operational testing.Many candidate approaches remain unvalidated under command-and-control conditions.
  • 6 Conclusion: Unresolved research questions include calibration thresholds for revalidation, rollback, or decommissioning, formalizing mission judgment into executable constraints, and dynamically reconfiguring assemblies.The largest remaining gap concerns supervisability.
  • 6 Conclusion: Latency boundaries should be protected from being learned, alongside other characterized decision boundaries.The passage links this protection to the exploitability principle.

A Test and Evaluation dimensions in scope

Table 10 presents the Testing and Evaluation dimensions included in scope.

  • A Test and Evaluation dimensions in scope: Table 10 lists the Testing and Evaluation dimensions in scope.
Loading 2608.20597v1…