Source-linked AI summary

Position: AI Is Not Ready for Strategic Conflicts

Mark Riedl, Glenn Matlin

arXiv:2609.16189v1cs.AI

TL;DR

Open-ended strategic wargames can let language models shape both what actors do and what becomes simulated reality, creating a safety gap for consequential decision support. The paper argues for auditable, refusal-capable safety cases and positions wargames as stress tests for LM agents rather than evidence that those agents are safe to deploy.

  • Problem

    Open-ended wargames lack sufficient evidence that their language-mediated roles, adjudication, uncertainty, and human oversight are safe for consequential planning and crisis decisions.

  • Method

    The paper develops a position through a five-failure-mode analysis and proposes auditable safety-case requirements plus stress-test research directions.

  • Results

    The paper concludes that ordinary benchmarks and LM-as-judge evaluation cannot establish safety, while open-ended wargames can expose failures as stress tests.

  • Takeaways & Limitations

    No LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case.

  • Takeaways & Limitations

    LM agents may interpolate recorded strategies rather than generate genuinely novel strategic options, leaving strategic imagination as an unverified safety property.

Abstract

from arXiv · show

Open-ended strategic wargames are high-stakes LM-based social simulations: they model adversaries, institutions, escalation, plan brittleness, doctrine, and crisis response. Language models (LMs) are attractive because they can play agents, generate scenario branches, adjudicate ambiguous actions, and summarize lessons, but the same affordances make open-ended roles dangerous: model language determines both what an actor attempts and what becomes simulated reality. This position paper argues that no LM-enabled wargame should inform planning, doctrine, policy, or crisis response without an auditable safety case, and that the proper use of open-ended wargames today is to stress-test decision-influencing LM agents. We identify five failure modes: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination. Ordinary benchmarks cannot establish safety for these settings. Wargames can expose failures as stress tests; they are not themselves safety cases for consequential use.

1 Position and Scope

Frontier language models are entering upstream defense planning and wargaming without auditable safety cases, while open-ended wargames should currently be used to expose failures before they influence real decisions.

  • Scope: These concerns extend beyond target selection because language models can shape planning, course-of-action selection, and information aggregation upstream of kinetic or policy effects.Military wargaming is the primary testbed, but the risk applies wherever open-ended simulations inform consequential decisions.
  • Position: No LM-enabled wargame or decision-support system should inform planning, doctrine, policy, or crisis response until a refusal-capable, auditable safety case establishes its boundaries.The required case must cover role boundaries, uncertainty, adjudication authority, human oversight, and inference limits.
  • Constructive path: Open-ended wargames should serve as stress tests that expose failures in decision-influencing LM agents before those failures reach real decisions.The paper frames this as preparation for pressure to use technologies perceived to offer strategic benefit during conflict.
  • Argument: The paper distinguishes open-ended wargames from closed games by their production of decision influence rather than merely game performance.It argues that closed-game performance cannot certify safety for this role.

2 Why Open-Ended Wargames Are Different

Open-ended wargames turn language into both action and simulated reality: models can flexibly generate, interpret, adjudicate, and summarize events, but unsupported consequences may thereby become evidence for human judgment.

  • Wargame structure: A strategic wargame models opposing actors, an uncertain environment, choices, adjudicated consequences, and after-action interpretation intended to inform real-world judgment.Its product is often a judgment about adversary behavior, important assumptions, or plan brittleness rather than a score.
  • Open-endedness: Unlike fixed games with inspectable rules, open-ended language-mediated wargames can transform novel proposals into scenario state, making unsupported consequences appear simulated facts.The same interface that provides flexibility also interprets, negotiates, or partially accepts proposed actions.
  • LM affordances: Language models can draft plans, maintain role-conditioned dialogue, simulate stakeholders, summarize transcripts, and adjudicate unforeseen actions in this setting.These affordances also create a hazard when hallucinations become world state, adjudications become evidence, and narratives become strategic conclusions.
  • Decision pathway: The consequential pathway runs from model-generated action through adjudicated consequence and accumulated scenario state to an after-action finding and human judgment.A biased adversary model can minimize an apparent vulnerability, while a plausible escalation narrative can make escalation seem inevitable.

3 Failure Pathways

Open-ended LM wargames exhibit interacting failure pathways in which shared model biases, opaque adjudication, and narrow strategic generation can turn fluent simulations into misleading strategic judgments.

  • Failure Pathways: Five failure modes structure the paper’s critique: decision laundering, adjudication opacity, role collapse, escalation-through-adjudication, and failure of strategic imagination.Together they shift the safety question from whether an agent wins to which generated claims enter the simulated world and influence decisions.
  • Adjudication opacity: Adjudication opacity allows unsupported assumptions to pass from contested actions into scenario state and final summaries, increasing the likelihood of decision laundering.The adjudicator determines whether actions work, what side effects occur, and which new facts enter the scenario.
  • Role collapse: Role collapse makes simulations self-confirming when one model family acts as player, adversary, adjudicator, analyst, judge, and summarizer.Shared priors, training data, objectives, tools, or prompting can create correlated error across institutional roles; harmlessness training may also soften adversarial behavior.
  • Escalation-through-adjudication: Escalation-through-adjudication can make escalation appear likely, rational, or unavoidable when early model framing persists as simulated reality despite no player choosing it.Repeatedly interpreting ambiguity in the most hostile way manufactures escalation pressure through path-dependent scenario state.
  • Failure of strategic imagination: LM agents often interpolate within recorded strategies rather than extrapolating beyond them, producing fluent but strategically narrow opponents and suppressing unexpected options.The gap between aspirational and realized creativity is presented as a safety property that should be measured rather than assumed.
  • Evaluation limits: Standard benchmarks and LM-as-judge evaluation identify specific failures but do not establish auditable operational boundaries for high-stakes use.Relevant evidence includes adjudication and summary traces showing assumptions, alternatives, uncertainty, and why contested actions were accepted.

4 Benchmarks Are Not Safety Cases

Benchmarks and common evaluation substitutes can identify specific failures, but they cannot establish auditable safety boundaries for open-ended wargames. A safety case instead documents fitness for a specified decision-support role, including hazards, controls, evidence, counter-evidence, and refusal limits.

  • Benchmarks face over-optimization, distributional mismatch with adaptive strategic interaction, and speculative futures without ground-truth labels.Expert review of exercise traces is therefore required.
  • LLM-as-judge evaluation introduces prompt sensitivity and correlated errors, while standard benchmarks and ad-hoc red-teaming do not establish operational boundaries.Human-likeness alone is also insufficient; audits should retain adjudication and summary traces, including assumptions, alternatives, and uncertainty.
  • A safety case is a retained-evidence argument about whether a system is acceptably safe for a specified use, covering hazards, controls, counter-evidence, and limits on inference.For open-ended wargames, it must also specify refusal bounds where deployment is prohibited.
  • A safety case does not certify that a wargame is true; it assesses whether a particular exercise fits a particular decision-support role.The case should make non-use available, and its evidentiary requirements should scale with consequence.

5 Implications and Conclusion

The paper proposes research and governance measures that expose failure boundaries before open-ended LM systems influence consequential decisions. Its central conclusion is that knowing when not to trust the game is more important than winning it.

  • Adjudication interpretability should inspect the causal assumptions, sources, and role commitments supporting major adjudications.
  • Open-action robustness should test semantically novel moves, paraphrases, hybrid-domain actions, and attempts to exploit ambiguity.
  • Summary-effect evaluation should measure how summaries, confidence language, and narrative framing affect human interpretation.
  • Negative findings define where systems should not be trusted, including role drift, unsupported consequences, collapsed dissent, escalation bias, and brittle behavior.
  • Procurement offices, reviewers, and model labs should require safety cases, retain exercise traces, disclose role boundaries, and evaluate escalation behavior.These norms extend to other high-stakes social simulations using open-ended LMs.
  • The key evaluative question is when the game should not be trusted, rather than whether it can be won.

Ethics Statement

This position paper addresses language-model support for strategic wargaming, planning, doctrine, policy, and crisis response as a high-risk dual-use setting. It performs risk assessment rather than introducing a deployable system or reporting human-subject experiments.

  • The paper examines a high-risk dual-use setting involving LM support for strategic wargaming, planning, doctrine, policy, and crisis response.
  • It does not introduce a deployable system, release operational scenarios, provide military-use instructions, or report experiments with human subjects.
  • Its purpose is to argue that open-ended wargames can stress-test decision-influencing LM agents and that high-stakes use requires auditable safety cases.

Use of Generative AI Tools

The authors used LM assistance for manuscript preparation while retaining responsibility for the submission’s claims and wording. The accompanying inventory documents publicly supported AI and LM roles in wargaming and planning, while distinguishing research, procurement, training, and operational evidence boundaries.

  • LM assistance was used to compress existing material, identify consistency issues, and prepare submission-support materials, not to generate data, figures, evaluation results, or new bibliography entries.
  • The authors reviewed and remained responsible for all claims, citations, and final wording.
  • The inventory supports a narrow claim that AI and LM systems are researched, procured, tested, or advertised for planning, course-of-action generation, simulation, wargaming, training, and after-action workflows.Public sources are not treated as evidence of operational reliance unless they state that directly.
  • The cited record includes reports about Claude in battlefield simulation and a sworn declaration describing Grok Gov Model integration into Maven Smart Systems for planning, red-teaming, logistics, and related roles.Exact operational roles remain unverified in the public record.
  • The inventory separates research prototypes, procurement or vendor claims, training and exercise use, and operational reliance when public evidence supports those distinctions.
  • The documented systems span AI-assisted scenario creation, staff support, agent training, COA generation, adjudication, analysis, red-teaming, and after-action workflows across multiple national and institutional settings.Examples include Snow Globe, ERDC agent training, NATO TIDE prototypes, and international wargaming ecosystems.

B Expanded Failure Pathways

Open-ended LM wargames combine generic model weaknesses with role-specific pathways that can turn fluent speculation into simulated reality and strategic judgment. These failures compound across adjudication, role assignment, escalation, and strategic imagination.

  • Open-ended wargames combine generic LM failures into domain-specific pathways that can influence real decisions.The paper identifies hallucination, prompt sensitivity, sycophancy, unfaithful reasoning, bias, long-context failures, and brittle out-of-distribution behavior as contributing weaknesses.
  • Decision laundering: Fluent model branches can pass through adjudication and summaries, laundering speculation and undisclosed assumptions into strategic judgment.Because wargames elicit expert judgment, model-supplied assumptions can become institutional belief and compound errors over long horizons.
  • Adjudication opacity: Adjudicators determine whether actions work, what side effects occur, and which new facts enter the scenario, but fluent explanations can conceal unsupported or inconsistent reasoning.Open-ended semantic action spaces make adjudication difficult across diplomacy, cyber, legal, financial, alliance, humanitarian, and hybrid moves.
  • Role collapse: Shared model families across player, adversary, adjudicator, analyst, and summarizer roles can produce correlated mistakes and self-confirming simulations.Multiple agents do not remove this risk when they share training data, objectives, tools, or prompting conventions.
  • Escalation-through-adjudication: Adjudicator interpretations can introduce hostile escalation, while early errors become durable scenario state across subsequent turns.The resulting escalation pressure is path-dependent rather than solely a consequence of player choices.
  • Failure of strategic imagination: LM agents often interpolate existing strategies rather than exploring novel branches, producing plausible but strategically narrow adversaries and plans.Agents perform better when the next move is well-cued and worse on open-ended planning where exploration matters; the creativity gap should be measured rather than assumed.

C.1 Expanded benchmark critique

Closed-game benchmarks and LM-as-judge methods can measure components of capability or error, but they cannot establish safety for open-ended wargames. Evaluation must instead examine role boundaries, uncertainty, causal assumptions, and exercise-level traces.

  • Benchmarks are useful for isolating monitored errors but insufficient for high-stakes open-ended wargame safety.Their limits include benchmark over-optimization, mismatch with adaptive open-ended tasks, and the absence of observable ground truth for speculative futures.
  • Open-ended wargame safety requires evidence about role boundaries, disclosed uncertainty, causal assumptions, strategic drift, and limits on presenting speculation as validated conclusions.These questions are not reducible to win rate or other closed-game capability measures.
  • LM judges can share player or adjudicator blind spots, correlate errors, and reward narrative plausibility over strategic validity.A safety case must test family bias and distinguish generated claims from independently verified evidence.
  • Human-likeness is not the correct safety target because LM agents may coordinate, collude, or act sycophantically in ways humans would not.Safer adversaries should instead be calibrated, role-faithful, non-sycophantic, and explicit about uncertainty.
  • Adjudication traces—not merely action explanations—should provide primary safety-case evidence for contested actions, assumptions, alternatives, and uncertainty.Mechanistic interpretability may help inspect role leakage, but it cannot replace exercise-level traces in the broader sociotechnical system.
  • Metrics should be treated as components of safety evidence rather than substitutes for a safety case.The paper frames the needed standard as a systematic evaluation of whether an exercise is fit for its intended use.

C.2 Expanded safety-case argument

A safety case is a systematic, evidence-supported argument that a specified LM-enabled wargame is acceptably safe for a particular decision-support role. The proposed norm is no safety case, no high-stakes use, with evidence scaled to consequence and supported by auditable controls.

  • A safety case defines the system’s purpose, hazards, controls, supporting evidence, counter-evidence, and permitted user conclusions for a specified use.Unlike a benchmark, it addresses the scope and conditions under which safety is claimed.
  • The central norm is no safety case, no high-stakes use; a safety case establishes fitness for a role rather than truth of the wargame.Table 5 sketches the argument structure for a high-stakes LM-enabled open-ended wargame.
  • COA-GPT’s published evaluation supports human selection and adaptation of courses of action but leaves separation, assumption tracing, and downstream over-claim auditing unaddressed.The paper characterizes the gap between that evaluation and a deployment-ready safety case as wide.
  • Responsibilities include stress-testing decision-influencing agents, reporting openness and failure profiles, auditing traces and role leakage, and requiring artifacts and safety cases before high-stakes use.The proposed responsibilities span researchers, benchmark designers, model developers, interpretability researchers, practitioners, reviewers, and funders.
  • Safety requirements should scale with consequence, from lightweight logging and human review to stronger role separation, SME adjudication, escalation testing, adversarial prompts, and post-game audit.National-security, public-health, and critical-infrastructure exercises receive the strongest requirements described.
  • The live-exercise audit rubric checks causal explicitness, role independence, escalation attribution, strategic continuity, and uncertainty disclosure.These checks target whether state changes, role boundaries, escalation, long-context consistency, and speculative inferences are traceable before scenario updates.
  • Progress should prioritize uncertainty exposure, resistance to role leakage, documented causal assumptions, and evidence-based refusal over greater fluency alone.The paper defines progress partly as learning where an LM-enabled exercise should not be trusted.

D Alternative Views and Objections

The paper answers objections about scope, reproducibility, capability, enforcement, and legitimacy by distinguishing open-ended decision influence from direct control and closed-game evaluation. It argues for studying failures now while requiring auditable, enforceable safety boundaries before high-stakes use.

  • Objection 1: Open-ended wargames pose an upstream decision-influence hazard through natural-language scenario generation, role-play, adjudication, and after-action interpretation, rather than direct weapon control.The paper distinguishes these systems from bounded perception-action loops in autonomous or remotely piloted systems.
  • Objection 2: Closed-game performance cannot establish safety for open-ended strategic decision support, where agents must handle ambiguous coercion, diplomatic signaling, and crisis-game summaries.Closed games remain appropriate for many questions about planning, search, and self-play.
  • Objection 3: Current LM failures—including brittle reasoning, hallucinations, rule nonadherence, inconsistency, and escalation risks—make safety study urgent rather than deferrable.Halo effects from capable demonstrations and jointly produced prompt sensitivity or sycophancy can make brittle wargame models appear ready.
  • Objection 4: A meaningful safety case must test negative claims, identify where hazards exceed controls or evidence fails, and specify where deployment is prohibited.The paper rejects checklist compliance and points to procurement, publication review, internal review, and post-incident audit as enforcement channels.
  • Objection 5: Auditability is a precondition, not a deployment license; some roles, including real-time crisis-response adjudication, may lack any valid safety case and therefore require refusal.The paper frames “no safety case, no high-stakes use” as covering currently uncertifiable roles.
Loading 2609.16189v1…