Source-linked AI summary

When AI Navigates the Fog of War

Ming Li, Xirui Li, Tianyi Zhou

arXiv:2603.16642v1cs.AIcs.CLcs.CY

TL;DR

The paper asks whether AI can reason about an unfolding war before its trajectory becomes historically obvious, despite incomplete, ambiguous, and contradictory signals and the risk of hindsight bias. It uses a temporally grounded, leakage-resistant case study that restricts models to information available at each moment, finding strong but domain-specific strategic reasoning whose narratives evolve as the conflict unfolds.

  • Problem

    The paper examines whether AI can anticipate geopolitical conflict in real time, where incomplete, ambiguous, and contradictory signals make retrospective judgments vulnerable to hindsight bias.

  • Method

    The study constructs temporal nodes with contextual information packages limited to publicly reported news available at each node, reducing training-data leakage while evaluating judgments and evolving narratives under uncertainty.

  • Results

    Models often show strong strategic reasoning, perform more reliably in economically and logistically structured settings than politically ambiguous multi-actor environments, and shift from rapid-containment expectations toward systemic accounts of escalation and fragile de-escalation.

  • Takeaways & Limitations

    The findings indicate that LLM reasoning in an unfolding geopolitical crisis can extend beyond surface rhetoric toward structural incentives, but its reliability varies across domains and its narratives change over time.

  • Takeaways & Limitations

    The questions are analytical anchors rather than definitive benchmark labels, and quantitative results are supportive signals within a qualitative case study because ongoing events and observation windows remain temporally open.

Abstract

from arXiv · show

Can AI reason about a war before its trajectory becomes historically obvious? Analyzing this capability is difficult because retrospective geopolitical prediction is heavily confounded by training-data leakage. We address this challenge through a temporally grounded case study of the early stages of the 2026 Middle East conflict, which unfolded after the training cutoff of current frontier models. We construct 11 critical temporal nodes, 42 node-specific verifiable questions, and 5 general exploratory questions, requiring models to reason only from information that would have been publicly available at each moment. This design substantially mitigates training-data leakage concerns, creating a setting well-suited for studying how models analyze an unfolding crisis under the fog of war, and provides, to our knowledge, the first temporally grounded analysis of LLM reasoning in an ongoing geopolitical conflict. Our analysis reveals three main findings. First, current state-of-the-art large language models often display a striking degree of strategic realism, reasoning beyond surface rhetoric toward deeper structural incentives. Second, this capability is uneven across domains: models are more reliable in economically and logistically structured settings than in politically ambiguous multi-actor environments. Finally, model narratives evolve over time, shifting from early expectations of rapid containment toward more systemic accounts of regional entrenchment and attritional de-escalation. Since the conflict remains ongoing at the time of writing, this work can serve as an archival snapshot of model reasoning during an unfolding geopolitical crisis, enabling future studies without the hindsight bias of retrospective analysis.

1 Introduction

The paper studies whether LLMs can reason about an unfolding war without hindsight by placing them in a temporally constrained information environment. It reports strategic realism, domain-specific reliability, and evolving narratives while preserving the uncertainty of an ongoing conflict.

  • Motivation: Retrospective geopolitical evaluation is confounded by training-data leakage because models may encode historical outcomes in their pretraining data.The paper therefore focuses on an unfolding conflict whose early stages occurred after current models’ training cutoffs.
  • Study objective: The study uses the early 2026 Middle East conflict to examine how LLMs interpret incomplete signals and construct narratives under the fog of war.Models are evaluated at successive temporal nodes using information available at each moment.
  • Findings: Models often display strategic realism by reasoning beyond political rhetoric toward military sunk costs, deterrence pressures, and material constraints.Some models anticipate escalation before kinetic conflict begins.
  • Findings: Model reliability is stronger for structural economic dynamics and material constraints than for politically ambiguous, multi-actor settings.The contrast involves signaling, leadership instability, and strategic interaction.
  • Findings: As the conflict unfolds, model narratives shift from expectations of rapid containment toward longer, more systemic accounts.The archived responses preserve this evolution for future comparison without final-outcome hindsight.
  • Study design: The framework contains 11 critical temporal nodes, 42 event-specific reasoning probes, and 5 general exploratory questions spanning military, economic, and political dynamics.The design enables longitudinal observation of model analyses as new information becomes available.

2 Background and Related Work

Prior work studies geopolitical forecasting, multi-agent reasoning, heterogeneous evidence, and leakage mitigation, but existing evaluations rarely combine real-time updating with an unfolding geopolitical crisis. This paper addresses that gap through a post-training-cutoff conflict and sequential temporal constraints.

  • Geopolitical forecasting: Existing forecasting benchmarks show that LLMs remain below expert human forecasters on unresolved future questions.Related systems include structured event databases and evolving-cast evaluations.
  • Leakage concerns: Temporal leakage remains a persistent concern, and prompting models to ignore known outcomes does not reliably create true ignorance.These findings motivate evaluating a crisis outside the models’ training distribution.
  • Multi-actor reasoning: Geopolitical crises intensify multi-actor reasoning demands through conflicting incentives, cascading effects, and rapidly shifting information.Existing benchmarks often simplify the actor space or treat geopolitical scenarios as static prediction tasks.
  • Temporal evaluation: Standard reasoning benchmarks use fixed inputs and predefined answer spaces, whereas this study repeatedly updates models as new information arrives.The temporal design enables analysis of belief revision and narrative coherence.
  • Contribution: The paper combines a post-training-cutoff conflict with node-specific information restrictions to reduce both direct and retrospective leakage.This is stricter than merely refreshing benchmark items.

3 Study Design

The study reconstructs the conflict as 11 temporally ordered information snapshots and probes each with verifiable and exploratory questions. Models receive only contemporaneous public reporting, while quantitative judgments remain auxiliary to qualitative analysis.

  • 3.1 Critical Temporal Nodes Construction: Each temporal node is a snapshot of publicly reported information available at a specific time, excluding later developments.The post-cutoff conflict further reduces training-data leakage risk.
  • 3.1 Critical Temporal Nodes Construction: The final timeline contains 11 nodes covering Initial Outbreak, Threshold Crossings, Economic Shockwaves, and Political Signaling.Nodes represent moments that substantially alter the strategic landscape and incorporate news review plus interviews with five Middle East residents.
  • 3.2.2 General Exploratory Questions: Exploratory questions are not scored for correctness; they create a longitudinal record of how model narratives change as the conflict unfolds.These responses are included in the released dataset for future research.
  • 3.3 Interaction Protocol: For every question, models receive the complete context package for that node and respond independently with analysis, a future-direction assessment, and an explicit probability.The full reasoning response and final probability are recorded.
  • 3.3 Interaction Protocol: Ground truth uses a fixed observation cutoff with at least a one-week resolution window, making quantitative comparisons transparent but provisional.Labels may change with a shorter or longer window, so not observed is not treated as permanent negation.
  • 3.2.1 Node-Specific Verifiable Questions: The analysis prioritizes qualitative reasoning patterns, using quantitative labels as auxiliary signals rather than definitive judgments.This reflects the ongoing conflict and the absence of universally final binary labels.

4 Reasoning Analysis under the Fog of War

Models often reasoned beyond political rhetoric toward structural incentives, but their performance varied with the domain and ambiguity of the setting. Across temporal nodes, they showed strategic realism while diverging most over diplomatic signals and uncertain political consequences.

  • Models commonly grounded geopolitical analysis in military deployments, deterrence dynamics, and institutional incentives rather than simply repeating political rhetoric.
  • Models differed in narrative structure and granularity, but these stylistic differences generally did not produce radically different overall judgments.
  • T0: Strategic Intuition and the “Credibility Trap”: Models were most consistent when inferring strategic consequences from large-scale military deployments and the credibility costs of withdrawing without concessions.Several models framed the buildup as generating its own momentum and a point of no return.
  • T0: Diplomatic Signals: Models diverged sharply over whether diplomatic talks implied sanctions and negotiation would remain preferable to military action.Some treated the military buildup as leverage, while others viewed it as creating momentum that made diplomacy a secondary strategy.
  • Escalation Reasoning: Models often reasoned beyond the June 2025 precedent, inferring that Iran’s prior survival could reduce the deterrent effect of another attack and raise the risk of sustained escalation.
  • Escalation Reasoning: Models generally separated inflammatory rhetoric from military doctrine, favoring calibrated retaliation because indiscriminate attacks could invite overwhelming escalation threatening regime survival.

5 Narrative Evolution in an Unfolding Conflict

Across three manually defined phases, model narratives shifted from narrow deterrence-based assessments toward systemic regional risks and economically constrained, indirect de-escalation. Despite escalating violence and geographic expansion, no model predicted direct military confrontation between major nuclear powers.

  • No model predicted a traditional “Third World War,” but their treatment of global conflict evolved across three phases.The analysis defines this as direct military confrontation between major nuclear powers rather than a static dismissal of risk.
  • Phase I (T0–T2): In Phase I, models defined global war narrowly through direct great-power intervention and generally viewed Russia and China as unlikely to enter militarily.Their reasoning centered on deterrence, U.S. force deployment, and initial kinetic exchanges.
  • Phase II (T3–T9): In Phase II, infrastructure attacks, Hormuz’s effective closure, and wider regional involvement led models to frame the crisis as a “Globalized Regional War.”Models emphasized systemic disruption of global energy supply chains and the possibility that prolonged closure would harden external powers’ positions.
  • Phase III (T10): In Phase III, models foregrounded decentralized command and state-collapse risks rather than deliberate escalation by major powers.The “Mosaic” doctrine and potential regional land grabs became central to the analysis.
  • Conflict resolution: De-escalation forecasts moved from rapid diplomatic containment toward a “hurting stalemate” and an indirect ceasefire driven by economic attrition and material limits.Later outputs emphasized unsustainable economic costs, interceptor shortages, and a messy negotiated pause rather than decisive military or political outcomes.
  • Conflict resolution: By the final phase, models converged on an “ugly, tacit, and indirect ceasefire” after both sides improved their bargaining positions over days or weeks.The predicted pause could occur before political objectives were fully achieved because of munition shortages and sustained economic pain.

6 Quantitative Signals

Quantitative signals broadly aligned with observed trajectories, while showing stronger reliability for economically structured causal chains than for politically ambiguous, multi-actor dynamics. The authors treat these scores as supportive indicators within a qualitative case study, not as a model-ranking scorecard.

  • 0.72 cross-model average indicates broad alignment between probabilistic judgments and outcomes observed at the fixed cutoff.The score uses 1 − MAE, where higher values indicate closer agreement with cutoff-based outcomes.
  • 0.79 thematic average for Theme III (Macroeconomic Contagion) was the strongest reported domain-level result.Models more reliably traced military disruptions to downstream energy-market and global-supply-chain effects.
  • Themes II and IV each averaged 0.67, reflecting greater difficulty with escalation thresholds, alliance entanglement, leadership dynamics, and strategic ambiguity.These themes involve unstable interactions among multiple actors.
  • Models were more reliable when tracing structurally legible downstream effects than when interpreting ambiguous strategic intent.The domain pattern was more informative than the relatively narrow cross-model score range of 0.63 to 0.75.
  • The quantitative results are treated as supportive signals rather than the primary basis for the paper’s claims.Questions concern evolving thresholds and partially realized developments, so cutoff-based non-observation does not mean impossibility or permanent resolution.

7 Conclusion

This study examines LLM reasoning about an unfolding geopolitical crisis using a temporally grounded, leakage-resistant case study. It finds strong but uneven strategic reasoning and evolving narratives, while framing the results as a contemporaneous snapshot because the conflict remains ongoing.

  • The study examines how LLMs reason about an unfolding geopolitical crisis under the fog of war.
  • The 2026 Middle East conflict provides a temporally grounded setting restricting models to information available at each moment.The design supports analysis of both probabilistic judgments and changing narratives under uncertainty.
  • Models often show strategic reasoning, but are more reliable in economically and logistically structured settings than politically ambiguous multi-actor environments.Their responses attend to military posture, deterrence, material constraints, escalation, exhaustion, and fragile de-escalation.
  • Because the conflict remains ongoing, the work records contemporaneous machine reasoning rather than a retrospective reconstruction.Its quantitative signals are anchored to a fixed observation cutoff, and archived responses support future temporal analyses.

A Experiment Settings

The experiment standardizes model access, prompting, context construction, and call scheduling across evaluated systems. Models receive identical temporally filtered news contexts under a uniform budget, with one sequential question per model and retry handling for transient failures.

  • Model Access and Identifiers: All models are accessed through OpenRouter using an OpenAI-compatible interface and a standardized request format.The exact model identifiers are listed in Table 9.
  • Generation Settings: A low temperature setting is used across models, while maximum output is capped at 2,048 tokens and other parameters remain provider defaults.
  • Prompt Structure: Each query contains only a user-role message with the contextual news corpus followed by the question, without a system prompt.
  • Evaluation Design: Each question is posed once per model, producing 97 questions per model and 485 API calls across five models.The design prioritizes qualitative reasoning patterns over estimating statistical variance from repeated samples.
  • Context Construction: Every model receives the same temporally filtered context: articles published before each node are collected, sorted newest first, formatted, truncated, and concatenated under a 480,000-character cap.The question and instruction are appended after the context.
  • Context Budget: The uniform context budget remains within the 256K-token capacity of the smallest evaluated models rather than being tailored to each model.
  • Execution: Calls use up to three exponential-backoff retries, rate-limit delays, and per-model sequential processing while models are queried in parallel.

B Complete Examples

The section presents a complete example of a model response assessing whether Iran would retaliate militarily after U.S. strikes.

  • The example answer assigns an 85–90% probability to Iranian military retaliation after U.S. strikes.It illustrates how the released examples operationalize model judgments about an unfolding conflict.

Analysis of Key Factors

The analysis combines Iranian threats, demonstrated capabilities, proxy options, and domestic pressures to assess retaliation, while noting constraints against uncontrolled escalation.

  • Iranian officials publicly warned that any U.S. attack would receive a firm and severe response.
  • Iran’s prior retaliation during the June 2025 conflict suggests a doctrinal commitment to responding directly rather than absorbing strikes passively.
  • Rebuilt missile facilities, military drills in the Strait of Hormuz, and missiles capable of reaching regional U.S. bases support Iran’s retaliatory capacity.
  • Iran retains proxy options through Hezbollah, the Houthis, and Iraqi militias, while regional warnings indicate contingency planning for proxy escalation.
  • Domestic unrest and regime vulnerability increase pressure to respond, but retaliation may be calibrated to avoid regime collapse or uncontrolled escalation.

Probability Assessment: 85-90%

The analysis argues that retaliation is highly probable because Iranian commitments, capabilities, and domestic incentives favor responding, although the response may remain layered and calibrated.

  • Iranian retaliation is assessed as highly probable because public commitments, military preparations, and internal legitimacy pressures all favor responding.
  • The expected response is multi-layered, combining direct strikes, proxy activity, maritime pressure, and potentially cyber operations.
  • Iran’s weakened condition, diplomatic incentives, and fear of regime-threatening escalation could limit retaliation’s scale.
  • The analysis treats a measured, layered response as more plausible than either no retaliation or immediate all-out regional war.
  • The assessment gives an 80% probability of military retaliation.
  • Leadership transition and the broken psychological barrier from the earlier war are presented as additional pressures for forceful retaliation.

5. The “Venezuela Model” Possibility

This section considers a Venezuela-style strategy of limited retaliation and diplomacy alongside broader estimates of possible retaliation forms, scales, and uncertainties.

  • 5. The “Venezuela Model” Possibility: A symbolic retaliation followed by diplomatic signaling is presented as a possible strategy for preserving Iranian interests under pressure.
  • Probability Assessment: 85-90%: The likely response is described as multi-layered, including ballistic strikes, proxy activation, cyberattacks, and maritime threats.
  • Probability Assessment: 85-90%: Full Strait of Hormuz closure, attacks on European targets, and use of newly acquired Chinese hypersonic systems are categorized as less likely but possible.
  • Probability Assessment: 85-90%: Uncertainty remains about strike damage, China’s role, IRGC cohesion, and the new supreme leader’s authority.
  • Probability Assessment: 85-90%: The conclusion favors sustained multi-front retaliation that stops short of actions inviting complete regime destruction.
  • Probability Assessment: 85-90%: The assessment assigns approximately 95% probability to meaningful military retaliation.

Analysis of the Current Situation

Iran’s leadership is portrayed as increasingly willing to absorb the costs of regional conflict because domestic vulnerability and the perceived threat to its strategic capabilities make non-retaliation especially dangerous. Ongoing diplomacy, military preparations, and explicit warnings therefore point toward a response involving asymmetric attacks, economic disruption, and broader regional escalation if the U.S. strikes.

  • Internal credibility: Iran’s domestic vulnerability and reliance on regional influence and missile capabilities increase the leadership’s perceived need to avoid appearing weak.The regime’s internal crackdown and postwar fragility are presented as intensifying the credibility stakes of its regional posture.
  • Warnings: Iranian officials have warned that aggression will not be treated as limited and that U.S. bases and regional assets would be legitimate targets.These statements frame retaliation as extending beyond the immediate point of attack.
  • Strategic shift: Iran appears to have moved beyond earlier caution, treating attritional or broader regional conflict as a necessary cost to preserve its strategic capabilities.The stated rationale is preventing the total dismantling of those capabilities.
  • Military posture: Missile-facility repairs, fortified sites, and Strait of Hormuz drills signal readiness to impose costs on the U.S. and its allies, not merely to defend Iran.The passage specifically identifies Gulf infrastructure as a potential source of complications for a U.S. campaign.
  • Diplomacy: The wide gap between U.S. and Iranian negotiating positions, combined with the U.S. buildup, leads Tehran to prepare for a severe response despite ongoing diplomacy.Tehran interprets the buildup as a credible threat rather than only a negotiating tactic.
  • Potential future direction: A U.S. strike is unlikely to remain a one-off event, with expected retaliation combining proxy and missile attacks, Hormuz disruption, and possible regional-allied involvement.The anticipated effects include pressure on U.S. bases, global oil prices, and U.S. and Israeli defensive resources.
  • Conclusion: Because Iran’s leadership frames survival around power projection and resistance, it may judge the strategic cost of not retaliating greater than the cost of military response.This conclusion links domestic fragility and the existential framing of a U.S. strike to the expected decision to retaliate.

C Detailed Performance Results

Detailed prediction alignment scores for node-specific verifiable questions are reported in Table 10, with higher scores indicating better alignment.

  • Performance results: Table 10 reports detailed prediction alignment scores for the node-specific verifiable questions.The accompanying description specifies that higher scores indicate better alignment.
Loading 2603.16642v1…