Source-linked AI summary

Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents

Batu El, Jinhee Paeng, Fatih Dinc, Shiye Su, Mete Erdogan, Aneesh Pappu, Haotian Ye, Wanjia Zhao, Surya Ganguli, James Zou

arXiv:2608.16578v1cs.AIcs.MAcs.SI

TL;DR

Predicting how interacting AI agents collectively evolve—and when interaction improves decisions or produces polarization and bias—is important for effective, aligned multi-agent systems. Studying over 10,000 agent communities, the paper develops a statistical-mechanics model that predicts trajectories and collective outcomes, finding conviction buildup, improved objective accuracy, and frequent rightward drift on subjective political questions.

  • Problem

    The paper asks how interacting AI-agent communities evolve and when communication improves collective decisions versus producing polarization, disagreement, or shared-bias amplification.

  • Method

    The authors study language-model communities revising opinions over fixed social networks and model updates as stochastic responses to social pressure.

  • Results

    Across models, tasks, and networks, interaction increased conviction, improved objective-question accuracy, shifted many subjective political opinions rightward, and the model predicted trajectories and group outcomes.

  • Takeaways & Limitations

    Collective behavior of interacting AI agents follows compact, predictive dynamical laws that connect recurring regimes and population outcomes to interpretable interaction parameters.

  • Takeaways & Limitations

    The method simplifies interactions by discarding the content of the natural-language messages that mediate influence.

Abstract

from arXiv · show

AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.

1 Introduction

The paper studies how interacting language-model agents develop structured collective behavior through repeated opinion exchange. It combines large-scale simulations with a statistical-mechanics model that predicts and interprets these dynamics.

  • Motivation: Interacting agents can combine complementary information but may also produce herding, polarization, disagreement, oscillation, and amplified shared biases.These collective outcomes become increasingly important as agents grow more numerous and autonomous.
  • Study design: The study simulates around 10,000 language-model agent communities revising opinions over repeated interactions across objective mathematics and subjective political questions.Experiments vary language models, questions, communication networks, and episodes, with agents assigned distinct personas or expertise.
  • Empirical findings: Individual opinion trajectories form recurring archetypes, while communities organize into three regimes: indifference, polarization, and consensus.The individual patterns include frozen opinions, one-time switches, reversals followed by returns, and repeated oscillations.
  • Model: The theoretical model represents binary opinions updated to minimize social pressure, combining intrinsic predispositions with interactions on a social graph.This framework is designed to predict and interpret the observed collective behavior.
  • Model evaluation and contribution: The model predicts collective behavior on held-out questions, outperforms standard baselines, generalizes to unseen network families, and approximately reproduces population-level outcome distributions.The results support compact dynamical laws for forecasting population dynamics and relating system-level outcomes to interpretable interaction parameters.

2 Setup

The study models heterogeneous language-model agents that repeatedly exchange messages and update opinions over a signed communication network. It compares objective mathematics questions, where personas encode expertise, with subjective political statements, where personas encode personal preferences.

  • Personas and Questions: The experiments use objective binary-choice MATH problems and subjective political statements answered as agreement or disagreement.Objective-question personas encode domain expertise from worked ground-truth solutions, whereas subjective-question personas draw on profiles of real individuals.
  • Agents: Each agent has a fixed persona, processes messages from connected agents, and samples an opinion and supporting message at every timestep.Personas introduce heterogeneity through variation in preferences for subjective questions and competence for objective questions.
  • Communication Network: The communication network is encoded by J ∈ {−1, 0, +1}^N×N, distinguishing friendly, absent, and unfriendly ties that regulate information exchange and trust.Friendly and unfriendly messages are routed into separate receiver inboxes according to each connection’s sign.
  • Procedure: Agents advance in synchronous rounds, sample each vote K = 5 times to reduce noise, and use the resulting opinion to compose a short message with a supporting reason.The averaged vote takes one of six evenly spaced values in [−1, 1], while the message state follows the sign of that average.
  • Opinion update: At each round, agents regenerate opinions from their personas, the question, and refreshed inboxes containing only the latest messages, propagating influence through J.The initial opinion depends only on the persona and question, and the system evolves to a terminal state after T timesteps.

3 Empirical Characterization of Agents’ Collective Dynamics

Across 9,600 simulated communities and a broader dataset exceeding 10,000 groups, agents display diverse individual and collective trajectories that organize into recurring archetypes and three regimes. Interaction builds conviction, shifts communities toward consensus or polarization, improves objective-question answers, and tends to move subjective opinions rightward.

  • Dataset and setup: The empirical analysis covers 9,600 simulated communities, while the full dataset exceeds 10,000 groups across models, questions, network structures, and episodes.The section analyzes continuous opinions over 8 interaction rounds.
  • Individual dynamics: Individual trajectories fall into four exhaustive archetypes: Frozen agents, Switchers, Intermittent agents, and Oscillators, defined by their number of opinion switches.Frozen agents make zero switches; Switchers make one; Intermittent agents make two; and Oscillators make more than two.
  • Group dynamics: Group trajectories exhibit five archetypes, with Divergence and Majority Switch reaching 11–12% in GPT-4o-mini and Qwen3.5-9B.The group classification uses initial and final net opinions together with a split band around zero.
  • Collective regimes: Communities begin predominantly indifferent and progressively transition toward consensus or polarization as conviction rises over time.Consensus combines high absolute net opinion with high conviction; polarization combines near-zero net opinion with high conviction; indifference has low conviction and near-zero net opinion.
  • Question-dependent outcomes: On objective questions, community majorities tend to move toward the correct answer, with incorrect-to-correct switches more common than correct-to-incorrect switches across all models.The initial majority vote is measured before interaction, and subsequent message exchange shifts the community toward correctness.
  • Question-dependent outcomes: On subjective political questions, majority opinions are more likely to shift from left to right than from right to left over time.Because subjective questions lack a well-defined correct answer, the analysis tracks political-directional drift instead.

4 Statistical Mechanics Model of Agents’ Behavior

The paper models opinion formation as agents stochastically minimizing social pressure while balancing signed social influence against intrinsic predispositions. A three-coupling extension captures connection, concordant, and discordant effects, with parameters fitted from observed one-step opinion transitions.

  • Model assumptions: Agents’ opinions are modeled as a system in which individuals adjust their states to minimize perceived social pressure.The interaction coefficient J_ij represents agent j’s influence on agent i, while the model assumes individuals seek lower social pressure.
  • Model assumptions: The energy function combines social-relationship consistency with intrinsic fields that bias agents toward their individual predispositions.J_ij ∈ {+1, 0, −1} encodes social influence, and g_i ∈ R captures each agent’s issue-specific predisposition.
  • Opinion dynamics: An agent’s next opinion is a logistic function of signed peer pressure scaled by β and an intrinsic persona field.The peer-pressure term is ∑j J_ij s_j(t), while g_i represents the agent’s leaning absent peer influence.
  • Three Couplings: The three-coupling model separates the effects of connection existence, concordant influence, and discordant influence.β0 measures connection presence, β+ concordant influence, and β− discordant influence; friendly-neighbor weight is β+ + β0, whereas unfriendly-neighbor weight is β0 − β−.
  • Fitting: The model fits social couplings and embedding-based intrinsic fields by minimizing cross-entropy on observed one-step opinion transitions.The three-coupling fit first estimates β0, then jointly estimates β+ and β− with β0 fixed.

5 Applications of the Statistical Physics Model

The fitted three-coupling statistical-physics model predicts agent trajectories on held-out questions and unseen graph families while reproducing aggregate group-archetype distributions. Its fitted dynamics explain conviction buildup, consensus preference, and truth-seeking through below-critical-temperature operation and asymmetric social influence.

  • Prediction: The discrete update with three couplings is best across Table 1’s prediction columns, reaching 75–86 balanced accuracy for one-step predictions and 61–77 for rollouts.Splitting the interaction parameter β matters: a single β can approach random chance in some model–question-type combinations.
  • Generalization: The three-coupling rule generalizes across graph families, with in- and out-of-distribution performance differing by at most 2.4 points and ranking best in 15 of 16 Table 2 columns.The evaluation includes held-out random graphs, low-rank graphs, square lattices, and triangular lattices.
  • Group archetypes: Model rollouts reproduce group-archetype shares closely, with mean absolute deviations of ∼3 points for objective questions and ∼5 points for subjective questions.The largest single gap is ∼10 points, while the two dominant classes are correct in every cell.
  • Mechanisms: The fitted operating point lies below the critical temperature in every dataset–model cell, making indifferent states unstable and causing conviction to grow over rounds.Aligned neighbors raise an agent’s local field, making agreement with neighbors resistant to change when stochasticity is low.
  • Mechanisms: Friendly influence dominates unfriendly influence because β+ consistently exceeds β−, favoring consensus over polarization.A stable split requires discordant connections strong enough to hold two groups apart.
  • Mechanisms: Correct neighbors exert stronger concordant and discordant influence, creating a truth-seeking asymmetry that moves communities toward the correct answer.The repulsive discordant channel also contributes to truth-seeking by pushing agents away from wrong answers.

6 Related Works

The paper connects its model to statistical mechanics, including Ising energies, Glauber updates, and Edwards–Anderson ordering. It also situates the work among applications to collective systems, LLM social simulations, and multi-agent reasoning.

  • Statistical Mechanics: The model uses the Ising energy function, Glauber opinion updates, and the Edwards–Anderson parameter to distinguish polarization from consensus.These foundations come from statistical mechanics of interacting spins and spin glasses.
  • Statistical Mechanics: Energy-based systems performing collective computation trace back to Hopfield networks.The cited origins include Hopfield networks and subsequent work by Amit et al.
  • Applications of Statistical Mechanics: The Ising model has been applied to spin glasses, flocks, neural systems, and protein sequences.These applications address distinct low-energy states, collective order, and descriptive biological models.
  • Social Simulations: LLMs are increasingly used to simulate experimental subjects, virtual communities, and collective phenomena such as opinion dynamics.This literature demonstrates that LLM communities can represent complex behaviors.
  • Multi-agent Systems: Multi-agent LLM systems improve reasoning through interaction, including debate and role-conditioned discussion, with increasingly automated prompt, workflow, and graph design.Earlier systems manually specified interactions, whereas newer systems search over interaction structures.

7 Discussion and Conclusion … C Additional Observations

The discussion identifies structured collective dynamics in interacting language-model communities, while outlining simplifying assumptions, extensions, and implications for anticipating and designing multi-agent systems. The appendices document setup details, additional observations, and methodological extensions.

  • 7 Discussion and Conclusion: 10,000 simulated groups exhibited structured collective dynamics across models, tasks, and communication networks, with repeated interaction increasing conviction and ordering initially indifferent groups.Collective accuracy improved on objective questions, while three of four models showed rightward political drift on subjective questions.
  • 7 Discussion and Conclusion: The setup assumes one shared binary question, fixed symmetric communication, and updates based only on current inbox messages rather than complete interaction histories.The method also discards much of the natural-language message content.
  • 7 Discussion and Conclusion: Future setup extensions include vector-valued opinions, simultaneous discussion of multiple statements, and communication patterns that change across timesteps.These extensions broaden the representation of opinions and the dynamics of interaction.
  • 7 Discussion and Conclusion: Future methodological extensions include Potts-type vector opinions and incorporating message content into the local field through message embeddings.The proposed message treatment draws on message passing in graph neural networks.
  • 7 Discussion and Conclusion: A theory of collective dynamics could anticipate multi-agent failure modes and inform system design if collective outcomes are predictable.The discussion frames this as increasingly important as AI agents become more widespread and interact routinely.
  • Appendix: The appendix documents setup details covering tasks, communication networks, models, prompts, and message sampling.These materials are listed under Appendix and its setup subsections.
  • C Additional Observations: Additional observations address political lean, label bias, characteristic-regime thresholds, rare group trajectories, and communication effects.The listed methodological appendix sections also cover local updates, continuous extensions, fitting, and baselines.

D Method Details

The method-details material includes analyses of raw accuracy, asynchronous updates, temperature sweeps, group-archetype predictability, and opinion flow.

  • Method Details: The listed method details cover raw accuracy and continuous-model utility, asynchronous updates, temperature sweeps, group-archetype predictability, and opinion flow.These topics appear as subsections E.1–E.5 in the supplied passage.

E Additional Results · A Terms and Notation · B Setup Details

The supplied appendix excerpts document the paper’s terminology and notation and show prompt templates spanning both task types and formats. Full prompts are referenced as appearing in Appendix B.4.

  • A Terms and Notation: Table S1 presents the paper’s terms and notation.
  • B Setup Details: Figure S1 covers prompt templates for objective and subjective tasks.
  • B Setup Details: The prompt templates include both opinion and message formats.
  • B Setup Details: The setup distinguishes two task types and two prompt formats.The task types are objective and subjective; the formats are opinion and message.
  • B Setup Details: Figure S1 directs readers to Appendix B.4 for the full prompts.

B.1 Tasks … C.1 Political Lean of Interacting Agents

The study evaluates interacting language-model agents on binary objective mathematics and subjective political tasks across varied communication networks, models, personas, prompts, and message-sampling procedures. Political interactions generally shift communities rightward, with model-specific differences in the magnitude and symmetry of opinion changes.

  • B.1 Tasks: Subjective tasks convert political statements into Agree (+1) or Disagree (−1) choices and use personas spanning demographic, attitudinal, political, and behavioral attributes.The dataset is filtered for high-entropy questions after binarization.
  • B.1 Tasks: Objective tasks convert MATH competition problems into binary choices with randomly assigned correct-answer positions and distractors to reduce positional bias.Objective expert personas receive ground-truth mathematical solutions rather than human-style demographic profiles.
  • B.2 Communication Networks: Communication networks comprise random signed graphs, square and triangular lattices, and rank- and frustration-controlled graphs designed to probe how graph structure shapes dynamics.Random graphs use a fixed edge budget, while the lattices are all-positive and reused across train and test.
  • B.3 Models: Experiments independently use four language models: gpt-4o-mini, gemma-3n-e4b-it, qwen3.5-9b, and llama-3.1-8b-instruct.Persona and question embeddings are computed with text-embedding-3-small.
  • B.4 Prompts: Prompts separately generate direct opinions and messages while conditioning agents on detailed personas describing demographic and attitudinal characteristics.Opinion generation explicitly requests a direct judgment without chain-of-thought.
  • B.5 Message Sampling: The standard procedure samples one message per agent per round and delivers it to all neighbors, whereas earlier GPT-4o-mini and Gemma-3n-E4B runs sampled messages separately for each neighbor.Both procedures draw from the same conditional message distribution.
  • C Additional Observations: Figure S6 analyzes political lean across 1,920 groups using label-balanced subjective statements from signed random graphs and square and triangular lattices.The analysis covers training- and test-set questions and tracks the fraction of communities with net right-leaning opinions.
  • C.1 Political Lean of Interacting Agents: Over eight rounds, three models drift right: Gemma-3n-E4B rises from 75% to 96%, Qwen3.5-9B from 52% to 67%, and GPT-4o-mini from 30% to 37%.Left→right switches exceed right→left switches for these models, while Llama-3.1-8B-Instruct remains near its initial 53% and is close to symmetric.

C.2 Label Bias … C.5 Effect of Communication Networks on Trajectories

The supplementary analyses identify label bias, define thresholds for indifference, polarization, and consensus, characterize rare trajectory patterns, and show that communication networks can alter opinion trajectories. Together, these results extend the paper’s account of collective dynamics beyond typical group outcomes.

  • C.2 Label Bias: Label bias can move communities toward an answer option regardless of its meaning, even when objective questions are balanced between +1 and −1.The analysis spans objective and subjective questions across signed random graphs and square and triangular lattices.
  • C.2 Label Bias: 39% to 16%: Llama-3.1-8B-Instruct groups answering +1 declined despite objective questions having equally frequent correct labels.This change cannot be attributed to truth seeking because the objective questions are label balanced.
  • C.3 Thresholds for Characteristic Regimes: The conviction threshold c⋆ is the pooled 1/3 quantile, classifying communities below it as indifferent.Conviction thresholds use the native range c ∈[0, 1].
  • C.3 Thresholds for Characteristic Regimes: The net-opinion threshold n⋆ is the median |n(t)| among committed communities, separating polarization from consensus symmetrically.Polarization has |n(t)| ≤ n⋆, while consensus has |n(t)| > n⋆, with |n| ∈[0, 1].
  • C.4 Rare Group Trajectories: Rare trajectories include backtrack, recovery, hesitation, and flipping, distinguished by reversals, weakening, or repeated changes in majority direction.Strong-majority episodes use |n(t)| ≥0.5, while the split band is ±0.2.
  • C.4 Rare Group Trajectories: Backtrack and recovery involve temporary or reversed strong majorities, whereas hesitation weakens and re-consolidates on the same side and flipping repeatedly changes sides.These patterns are not fully distinguished by first-versus-last-round classification.
  • C.5 Effect of Communication Networks on Trajectories: Different communication networks can produce different trajectories for the same question, as illustrated by random-matrix and square-lattice graphs.The comparison tracks couplings, net opinion n(t), and group trajectories across four episodes.

D Method Details … E.1 Raw Accuracy and the Utility of the Continuous Model

The appendices derive synchronous and continuous-time opinion dynamics from a Boltzmann energy model, then specify fitting, baselines, evaluation, and rollout procedures. Raw accuracy favors the continuous rule over the discrete rule because rare flips make its fitted update rate trade sensitivity for precision on common non-flips.

  • D.1 Local Update Rule: The local update probability is a logistic function of peer pressure and an agent’s predisposition under the Boltzmann distribution.Only rescaled parameter products are identifiable from opinion data, so the model fits them directly and reports temperature T = 1.
  • D.2 Continuous Extension: The continuous extension lets agents revise independently at random times, with each agent updating during dt with probability dt/τ.The synchronous rule is recovered in the limit ε →1 after discretizing rounds, where ε = 1 −e−dt/τ.
  • D.2.1 Flip Rate: The continuous model’s flip rate combines an independent update rate with a state-dependent logistic flip probability.Its master equation expresses probability conservation as inflow from and outflow to states differing by one agent flip.
  • D.2.3 Mean-Field ODE: The exact master equation is reduced to agent-level mean opinions by a mean-field approximation that neglects correlations between agents.The resulting closed ODEs describe each opinion relaxing on timescale τ toward tanh(β(Jm)i + gi).
  • D.3 Fitting: Fitted methods share a 16-dimensional persona–question feature block and add population, signed-network, or three-coupling interaction terms.The continuous one-coupling design has 18 parameters because ε is additional, while continuous three-coupling models also include ε.
  • D.4 Baselines: Baselines isolate class imbalance, persistence, intrinsic preferences, and social influence, including an interaction-free model and a population-average mean-field model.The mean-field baseline has 17 learned parameters and tests whether signed-network structure adds predictive value.
  • D.5 Balanced Accuracy: Balanced accuracy averages performance across four flip/stay and opinion-label groups, preventing rare flips or label imbalance from dominating evaluation.Rollout evaluation applies the same decomposition to predictions from rolled-out trajectories.
  • D.6 Rollouts; E.1 Raw Accuracy and the Utility of the Continuous Model: Deterministic rollouts use modal predictions for trajectory tracking, whereas stochastic rollouts sample agent updates to compare individual and group archetype distributions.On synchronously updating communities, the continuous rule outperforms the discrete rule under raw accuracy, but not flip-balanced accuracy, because rare flips favor precision on non-flips.

E.2 Asynchronous-update experiment

The appendix tests asynchronous agent updates using the main experiment’s questions, agents, and interaction graphs, with independent update schedules over eight timesteps. Under this setup, the continuous model outperforms the discrete model by a larger margin, while rollout accuracies are evaluated against four standard baselines.

  • Generating the update schedule: Asynchronous experiments use the same questions, 32 agents, and random interaction graphs as the main experiment, running each trajectory for T = 8 timesteps.Each agent updates independently with probability 0.5/1000 per cell across 1000 cells per timestep, averaging 0.5 updates per step.
  • Running the dynamics: When scheduled, an agent reads its current inbox, sends its current answer to neighbors, and samples a new answer using majority vote over K = 5 samples.Agents begin at time 0 with empty inboxes, as in the synchronous experiment.
  • Prediction under asynchronous dynamics: Rollout accuracies are evaluated for the asynchronous setup against Majority Class (M), Persistence (P), Interaction-Free (IF), and Mean-Field (MF) baselines.GPT-4o-mini is used as the language model, and the model is fitted and evaluated on the asynchronous setup.
  • Results: The continuous model outperforms the discrete model with a larger margin under asynchronous updates.This result is reported for the asynchronous setup described in the appendix.

E.3 Temperature Sweep Experiment · E.4 Predictability of Group Archetypes · E.5 Flow of Opinion

The appendices detail a fitted-model temperature sweep, assess whether group archetypes are predictable across independent episodes, and visualize opinion flow against continuous-model rollouts. Together, these analyses specify the experimental procedures and compare observed dynamics with model predictions.

  • E.3 Temperature Sweep Experiment: The temperature experiment modifies the fitted model’s temperature rather than the language model’s sampling temperature.At T = 1, the sweep recovers the fitted model itself.
  • E.3 Temperature Sweep Experiment: 80 objective and 40 subjective groups per model are generated from four random graphs and the test question bank, with four independent episodes each.The setup covers every test-set question and four in-distribution random graphs.
  • E.3 Temperature Sweep Experiment: The sweep runs 41 log-spaced temperatures from 0.05 to 20.16 for 500 synchronous steps, retaining the final 300 after discarding 200.At each step, opinions are sampled according to σ(h_i/T).
  • E.3 Temperature Sweep Experiment: Each group’s critical temperature T_c is estimated from the peak of susceptibility χ, using a parabolic fit around the maximizing grid point.The fitted vertex resolves T_c below the spacing of the 41-point temperature grid.
  • E.4 Predictability of Group Archetypes: Group-archetype predictability is tested by using one of four independently sampled trajectories for prediction and the remaining three to form a confusion matrix.The analysis uses test questions and four random graphs for each question–graph pair.
  • E.5 Flow of Opinion: Opinion-flow analysis divides 32 agents into equal subgroups of 16 based on their round-0 opinions, with subgroup A defined as the initially higher side.Ties are resolved using subsequent rounds.
  • E.5 Flow of Opinion: Observed subgroup trajectories are compared with rollouts from the fitted continuous three-coupling rule initialized from the same state on held-out questions.The analysis records subgroup net opinions on a 220 × 220 grid and visualizes real paths in white versus model paths in pink.
Loading 2608.16578v1…