Source-linked AI summary
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, Yongfeng Zhang
TL;DR
The paper asks whether AI can help explain and potentially avoid wars by addressing limits in static historical analysis and existing simulations. It proposes WarAgent, an LLM-powered multi-agent system that simulates countries and their decisions across WWI, WWII, and the Warring States Period. The simulations suggest that even minimal triggers can produce conflict-like trajectories, while altered national policies may divert them, although the model does not capture the full complexity of historical diplomacy.
Problem
Historical conflict analysis is limited by static methods, while the triggers, conditions, and inevitability of wars remain open questions.
Method
WarAgent uses LLM-based country agents to simulate historical conflicts, decisions, interactions, and counterfactual settings across WWI, WWII, and the Warring States Period.
Results
Even minimal or null triggers can spiral toward Cold War-like situations, while counterfactual experiments suggest that altered national policies are necessary to divert conflict trajectories.
Takeaways & Limitations
The simulations offer a data-driven perspective for examining war triggers and conditions relevant to historical understanding, conflict prevention, and resolution.
Takeaways & Limitations
WarAgent does not encompass the full spectrum of historical diplomatic intricacies, including communication delays caused by differing technologies and message-transmission times.
Abstract
from arXiv · showhide
Can we avoid wars at the crossroads of history? This question has been pursued by individuals, scholars, policymakers, and organizations throughout human history. In this research, we attempt to answer the question based on the recent advances of Artificial Intelligence (AI) and Large Language Models (LLMs). We propose \textbf{WarAgent}, an LLM-powered multi-agent AI system, to simulate the participating countries, their decisions, and the consequences, in historical international conflicts, including the World War I (WWI), the World War II (WWII), and the Warring States Period (WSP) in Ancient China. By evaluating the simulation effectiveness, we examine the advancements and limitations of cutting-edge AI systems' abilities in studying complex collective human behaviors such as international conflicts under diverse settings. In these simulations, the emergent interactions among agents also offer a novel perspective for examining the triggers and conditions that lead to war. Our findings offer data-driven and AI-augmented insights that can redefine how we approach conflict resolution and peacekeeping strategies. The implications stretch beyond historical analysis, offering a blueprint for using AI to understand human history and possibly prevent future international conflicts. Code and data are available at \url{https://github.com/agiresearch/WarAgent}.
1 Introduction
The paper introduces an LLM-based multi-agent framework for simulating historical international conflicts and investigating why wars occur or might be avoided. It frames simulation effectiveness, war triggers, and historical inevitability as central questions with implications for history, policy, education, and computational social science.
- Research framework: WarAgent models historical actors as country agents whose interactions can represent conflict, cooperation, and alternative international outcomes.The framework uses a dynamic environment to explore how different decisions and conditions could shape international conflicts.
- Research questions: The study asks whether LLM-based simulations can reproduce strategic decision-making, identify decisive war triggers, and reveal conditions leading to war or peace.These questions are examined using WWI, WWII, and the Warring States Period as historical settings.
- Simulation effectiveness: Simulation fidelity is treated as the foundation for evaluating the model’s credibility and for conducting subsequent analyses of conflict dynamics.The authors propose comparing simulated outcomes with documented historical events and trends.
- Casus belli: Iterative simulations provide a controlled environment for varying casus belli and observing the consequences of different war-trigger scenarios.The authors aim to compare the relative importance of triggers for conflict.
- War inevitability: The framework replays history under altered conditions to examine whether wars were inevitable or contingent on particular circumstances and decisions.This targets the relationship between historical structures and agent decisions.
- Implications: The proposed simulations are presented as tools for historical analysis, policy experimentation, education, and future research in computational history and digital humanities.The paper describes possible uses ranging from conflict prevention and resolution to interactive what-if learning.
2 Background and Related Work
The related work spans reasoning, NPC, production, and historical simulation systems, while positioning WarAgent as an extension of multi-agent methods to historical events. Its contribution is an LLM-based system for exploring historical occurrences and quantitative what-if scenarios.
- Multi-agent systems: Multi-agent systems are categorized into reasoning-enhancement, NPC, and production-enhancement systems.These categories organize prior work around collaborative reasoning, simulated characters, and task or software production.
- Reasoning-enhancement systems: Reasoning-enhancement systems use debate, role-based evaluation, and argument exchange to refine solutions and assess generated outputs.Examples include LLM-Debate, ChatEval, Corex, and MAD.
- NPC systems: NPC systems simulate believable human behavior, including planning, communication, relationships, and coordinated activities among agents.Generative Agents and Humanoid Agents are cited as representative approaches.
- Production-enhancement systems: Production-enhancement systems assign specialized roles to multiple agents for software development, task management, and complex problem solving.The cited systems include MetaGPT, BOLAA, OpenAGI, BabyAGI, and AgentVerse.
- WarAgent’s position: WarAgent extends multi-agent systems to historical event simulation using WWI, WWII, and the Warring States Period, adding quantitative analysis of historical and counterfactual scenarios.The authors describe this as an initial LLM-based effort to model historical-event trajectories.
- Historical simulation: Historical simulation has progressed from human and hybrid simulations to computer-based models of warfare, strategy, and international interaction.Earlier systems combined human decisions with computation, while later models represented tactical operations and adaptive strategies.
3 WarAgent Simulation Setting
WarAgent simulates international conflicts through country agents representing historical settings, profiles, relationships, and strategic conditions. The study focuses on WWI, WWII, and the Warring States Period as distinct environments for examining war initiation and international dynamics.
- WarAgent examines international relations and war initiation in WWI, WWII, and the Warring States Period in Ancient China.
- The simulation setting defines historical events, country-agent profiles, available actions, required inputs, and possible outcomes.
- Each country profile includes Leadership, Military Capability, Resources, Historical Background, Key Policy, and Public Morale.These dimensions provide a multifaceted basis for agent behavior and decision-making.
- Military capability, resources, policy, and morale represent strategic, economic, political, and social conditions shaping country decisions.The paper describes military strength, geography, population, GDP, objectives, and public sentiment as relevant profile information.
- Historical background represents prior conflicts and unresolved disputes that can influence current policies, posture, and potential alliances.The paper illustrates this with France’s loss of Alsace-Lorraine after the Franco-Prussian War.
4 WarAgent Architecture
WarAgent combines country agents, secretary agents, a Board, and a Stick to simulate decisions and maintain structured information about international conflicts. Agents reason about alliances, enemies, and actions while internal verification and shared records constrain the simulation process.
- WarAgent consists of country agents, secretary agents, a Board, and a Stick, with separate agent-secretary and agent-agent information flows.The Board and Stick support shared international-state and country-level record keeping.
- Country Agents: Country agents use structured prompts to analyze alliances, enmities, interests, and available actions in each round.The prompts guide agents through complex international relationship situations.
- Secretary Agents: Secretary agents verify action names, formatting, permissible action-space choices, and basic logical consistency.They serve as a safeguard against hallucination and imperfect reasoning in long, complex scenarios.
- Board: The Board records and displays changing international relationships, including war declarations, military alliances, non-intervention treaties, and peace agreements.It updates relationship status so agents use current information during simulation rounds.
- Stick: The Stick records domestic conditions such as mobilization, internal stability, and war-readiness prediction for each country.These metrics support alignment between country actions and predefined domestic protocols.
- Agent-Secretary Interaction: Country-agent proposals undergo iterative secretary review, with revision dialogue capped at four exchanges before direct amendment if agreement is not reached.Secretary agents do not participate in country-agent interactions.
- Agent-Agent Interaction: Each round begins with country agents responding to a triggering event, after which they communicate through actions, messages, and requests.In the WWI example, the trigger is the assassination of Archduke Franz Ferdinand of Austria-Hungary.
5 Experimental Design
The experiments evaluate WarAgent’s ability to simulate historical conflicts, identify war triggers, and examine whether alternative conditions change historical trajectories. The design combines human and board-based evaluation for simulation effectiveness with counterfactual analysis for casus belli and war inevitability.
- Experiments use GPT-3.5-turbo-1106, GPT-4-1106-preview, and Claude-2 as backbone models across the study.All experiments use the three models unless otherwise specified.
- Simulation Effectiveness (RQ1): RQ1 evaluates simulation effectiveness by comparing simulated outcomes with documented historical events and trends across multiple runs.The evaluation includes human assessment and calculated accuracy scores.
- Casus Belli (RQ2): RQ2 varies the intensity of counterfactual WWI trigger events to examine their potential impact on war outbreak.The design probes whether the historical trigger was unique or essential to the conflict’s outbreak.
- War Inevitability Observation (RQ3): RQ3 alters country profiles and decision-making pathways to analyze how initial conditions affect historical trajectories.The aim is to identify conditions that can significantly modify the course of history.
- RQ1 uses human and Broad Connectivity Evaluation, while RQ2 and RQ3 use counterfactual analysis and observational comparison of simulation outcomes.Table 1 links these evaluation methods to the corresponding research questions.
- Human Evaluation: Human Evaluation assesses whether country actions fit agent profiles and remain stable and rational across multiple rounds.The evaluation focuses on interests, consistency, and coherent decision-making.
- Board-based Accuracy: Board-based Accuracy compares simulated and historical alliances, war declarations, and general mobilization.Alliance accuracy uses mutual information between partitions, while war and mobilization accuracy use the Jaccard index.
6 Results
WarAgent produced historically plausible alliances and conflict patterns, while accuracy varied across relationships, models, and simulation settings. Counterfactual finetuning did not prevent global war, whereas de-anonymization increased historical alignment and consistency but reduced some simulation-specific behaviors.
- Military Alliance: 100% of simulations formed the historically aligned Britain–France, German Empire–Austria-Hungary, and Serbia–Russia alliances.The reported alliances reflected strategic, political, and ethnic considerations described for the WWI setting.
- War Declaration: War declarations occurred in 100% of simulations for Austria-Hungary–Serbia, Austria-Hungary–Russia, and German Empire–Russia, but France–German Empire and Britain–German Empire occurred in 71.4% and 14.3%.The reported sequence began with Austria-Hungary declaring war on Serbia, followed by declarations shaped by alliance structures and hostilities.
- Human Evaluation: The simulations reproduced plausible historical scenarios under the default assassination-trigger setting, but two special cases involved unsupported messages or changing diplomatic commitments.One case involved supportive messages without concrete action; another involved Britain violating a non-intervention treaty and later declaring war.
- Warring States: Alliance accuracy exceeded 75% and mobilization accuracy exceeded 90%, while war-declaration accuracy was comparatively low, although global war occurred in every simulation.These results summarize the three evaluated aspects under the default setting across the scenarios.
- Error Analysis: One of seven GPT-4 simulations reversed major allegiances, while Claude-2 and GPT-3.5 produced implausible alliances and random war declarations linked to weaker reasoning.Variability in Ottoman and United States participation also reduced simulated historical accuracy.
- Anonymized and De-anonymized Simulation: Counterfactual finetuning changed GPT-3.5’s reasoning performance but still produced global war, whereas de-anonymized simulations aligned more closely with history and were more consistent across models and random seeds.De-anonymized runs also omitted non-intervention treaties, unlike anonymized simulations where such treaties often occurred.
- Network Dynamics: During the six-day de-anonymized evolution, no non-intervention treaty was signed, matching the reported historical pattern.The figure’s board notation distinguishes war declarations, alliances, non-intervention treaties, peace agreements, and default states by color and symbol.
6.2 Casus Belli
The Casus Belli experiment varied WWI trigger intensity across repeated GPT-4 simulations. Outcomes ranged from mobilization without war to frequent global war, indicating that triggers influenced escalation but did not determine it uniformly.
- Experimental design: GPT-4 simulations tested three WWI triggers at increasing conflict intensity, repeating each scenario three times.The Null trigger provided a no-conflict baseline, followed by the Anglo-German Naval Incident and the higher-intensity Austria-Russia conflict over the Dardanelles Strait.
- Null trigger: The Null trigger produced military alliances and war readiness across simulations, but no direct conflict or declarations of war.The resulting situation resembled a cold war, with France, Britain, Russia, and Serbia opposed to the German Empire and Austria-Hungary.
- Intermediate summary: Across scenarios, heightened tensions and mobilization did not inevitably produce war, while specific triggers changed the likelihood and timing of escalation.The results support varied outcomes rather than a uniform relationship between military readiness, trigger intensity, and open conflict.
- Anglo-German Naval Incident: 1 of 3 Anglo-German Naval Incident simulations produced war, while the other two ended without declarations despite military mobilization.The war simulation involved alliance formation and additional declarations, whereas the other cases were predominantly resolved peacefully.
- Austria-Russia conflict: 2 of 3 Austria-Russia Dardanelles simulations produced global war after immediate mobilization and declarations by major powers.The declarations generated a domino effect that drew allied countries into the conflict.
6.3 War Inevitability
The War Inevitability experiments varied agent decision-making and country parameters to examine how aggression and national conditions affect conflict. Aggressive settings increased early war declarations, while historical background, public morale, and key policy appeared more influential than military capacity or resources alone.
- Experimental design: The experiments manipulated agent decision-making and country parameters to assess their effects on war likelihood.Decision-making used default, aggressive, and conservative settings; country experiments varied military capacity, resources, historical background, public morale, and key policy.
- Decision-making process: Aggressive settings produced war declarations in the first round, whereas conservative settings produced only alliances, treaties, and peace agreements after 10 rounds.The comparison indicates that agent aggressiveness substantially changed the timing and probability of conflict.
- Military capacity: Changing German military capacity and French military capacity produced no obvious delay in the German Empire’s war involvement across the reported rounds.The alternative military-capacity scenario reported an average involvement starting round of 4, similar to the default setting.
- Resources: Changing resource abundance produced no obvious change in war involvement or declaration patterns for France and the German Empire.This result held under both alternative resource scenarios.
- Historical background: Removing the historical background between France and the German Empire resulted in no direct war involvement or declaration between them.The paper’s intermediate summary identifies historical background, key policy, and public morale as significant influences on war propensity.
- Intermediate summary: The findings attribute greater importance to historical grievances, nationalistic sentiments, and diplomatic contexts than to military capability or resources alone.Military capacity and resources remain relevant, but the paper describes historical context and alliances as more decisive in the examined cases.
7 Can We Trust Simulation Results?
The paper treats LLM-based social simulation as a complementary way to study complex social systems, while acknowledging criticisms about simplicity, limited insight, real-world relevance, and verification. It therefore frames simulation outputs as informative suggestions rather than definitive conclusions.
- Simulation approach: The study applies LLM-derived intelligence to social simulation as an inductive method for generating artificial social scenarios.Unlike direct real-world measurement, the approach simulates phenomena to examine possible social dynamics.
- Criticisms: Critics question whether simulations oversimplify society, reveal unprogrammed interactions, relate to real-world complexity, or produce verifiable results.These concerns define the paper’s main trust and validity challenges.
- Interpretation: The paper argues that simulation results should be interpreted as informative suggestions or rationales, not definitive conclusions.The proposed role is to provide hypothetical outcomes that support evaluation of strategies and policies.
- Conclusion: The authors position simulations as complementary tools for social science research and policy analysis that provide additional insights into complex social systems.Their value is framed as augmentation of understanding rather than replacement of other approaches.
8 Conclusions, Discussions, and Future Vision
WarAgent is presented as a tool for analyzing international conflict dynamics and counterfactual historical scenarios, while its current design omits several historically important timing, information, and mobilization processes.
- WarAgent simulations indicate that even minimal or null war triggers can escalate toward Cold War-like situations.
- The system is the first LLM-based multi-agent system presented for simulating historical events, but it does not capture the full complexity of historical diplomacy.
- Limitations: Historical communication delays, espionage, and message-publicity gradients can shape diplomatic information flows beyond the model’s current representation.
- Limitations: National differences in army mobilization can affect the timing and feasibility of war declarations, but the framework may not fully model these time-sensitive processes.
- Future vision: LLM-based multi-agent simulations could support quantitative studies of diplomatic communication, non-state actors, treaties, and historical counterfactuals.
- Simulation design: The current round-based system permits one-way communication per country pair per round, whereas historical interactions often unfolded asynchronously.
- Future vision: Future work could introduce systematic stopping criteria, such as stable board connectivity, peace treaties, or economic and military equilibrium.
A Setting Anonymity
The setting anonymizes countries, locations, historical events, and resources by mapping them to generic names or altered descriptions.
- Britain, France, Germany, Austria-Hungary, Serbia, Russia, the United States, and the Ottoman Empire are mapped to generic country labels.
- Alsace-Lorraine is represented as two iron mines.
- The Dardanelles Strait is renamed Allison Strait in the anonymized setting.
- The assassination of Archduke Franz Ferdinand is represented as the assassination of the king of Country A.
B An Example Experiment of WWI
The WWI example records country agents exchanging messages, alliances, non-intervention treaties, mobilizations, and war declarations across the simulated conflict sequence.
- Britain and France request and accept a military alliance during the simulated WWI sequence.
- Multiple countries send condolences, diplomatic messages, and non-intervention proposals while seeking to manage escalation.
- Austria-Hungary requests a military alliance from Germany and declares war against Serbia.
- Russia generalizes its involvement through mobilization, support for Serbia, and a declaration of war against Austria-Hungary.
- The United States and Ottoman Empire pursue non-intervention arrangements and diplomatic communication with several participants.
- France and Germany mobilize, while Serbia mobilizes and accepts Russia’s military alliance.
C Example Prompts for Decision-Making Process of Agents
The example prompts vary agent behavior by emphasizing survival, reversibility, caution, long-term interests, or aggressive action in a virtual war game.
- The system prompt frames each country agent as playing a virtual war game aimed at maximizing national survival and winning.
- Aggressive setting: The aggressive setting permits aggressive actions when they are expected to benefit the country.
- Default setting: The default prompt asks agents to weigh immediate ambition, long-term benefit, and reversibility when selecting actions.
- Aggressive setting: The aggressive action prompt instructs agents to execute necessary war declarations promptly when circumstances favor national benefit.
- Conservative setting: Conservative prompts ask agents to align actions with their interests, consider long-term benefits and reversibility, and exercise caution before declaring war.
D An Example Experiment of WWII of One Round
In one simulated WWII round, countries pursued a mixture of military alliances, non-intervention treaties, mobilization, and diplomatic messaging. The actions reflected stated concerns about expansion, regional stability, neutrality, and mutual protection.
- Germany requested military alliances with Italy and Hungary while seeking non-intervention treaties with Japan and China.
- Japan requested military alliances with Germany and Italy, sought a non-intervention treaty with Hungary, and conducted general mobilization.
- Italy, Hungary, the United States, Russia, Britain, China, and France made alliance or non-intervention requests to other countries.
- Several messages framed alliances as responses to expansionist or aggressive threats and described non-intervention treaties as means to preserve neutrality, stability, or peace.
- The United States, Russia, and Britain chose general mobilization during the round.
E An Example Experiment of Warring States Perios of One Round
In one simulated Warring States round, states alternated between waiting, requesting alliances or non-intervention treaties, and sending messages proposing cooperation, defense, trade, or peaceful coexistence. The interactions combined strategic balancing with efforts to avoid conflict.
- Qi and Yan chose to wait without action during the round.
- Zhao requested a non-intervention treaty with Qin and military alliances with Wei, while Wei requested military alliances with Han and Zhao.
- Messages emphasized strategic alliances for mutual military or economic interests, defense, regional balance, and power balancing.
- Other messages proposed dialogue, trade, cooperation, peaceful coexistence, or non-intervention to maintain stability and avoid conflict.
- Qin requested military alliances with Wei and Han, and Chu requested a military alliance with Han.