Source-linked AI summary
Emergent social conventions and collective bias in LLM populations
Ariel Flint Ashery, Luca Maria Aiello, Andrea Baronchelli
TL;DR
The paper asks whether LLM populations can spontaneously develop shared social conventions, and experimentally examines this through decentralized agent interactions. It finds that conventions and collective biases can emerge, while committed minorities can impose alternatives.
Problem
The paper asks whether universal social conventions spontaneously emerge in populations of large language models.
Method
The study uses decentralized LLM-agent interactions in which randomly selected pairs communicate and agents can initially invent new words.
Results
Social conventions spontaneously emerge, coordination produces collective biases, and committed minority agents can impose alternative conventions on larger populations.
Takeaways & Limitations
LLM populations can autonomously develop and change shared social conventions without explicit programming.
Takeaways & Limitations
The experimental setup relies on several parameters that are unavoidable in LLM research.
Abstract
from arXiv · showhide
Social conventions are the backbone of social coordination, shaping how individuals form a group. As growing populations of artificial intelligence (AI) agents communicate through natural language, a fundamental question is whether they can bootstrap the foundations of a society. Here, we present experimental results that demonstrate the spontaneous emergence of universally adopted social conventions in decentralized populations of large language model (LLM) agents. We then show how strong collective biases can emerge during this process, even when agents exhibit no bias individually. Last, we examine how committed minority groups of adversarial LLM agents can drive social change by imposing alternative social conventions on the larger population. Our results show that AI systems can autonomously develop social conventions without explicit programming and have implications for designing AI systems that align, and remain aligned, with human values and societal goals.
Introduction
The paper investigates whether communicating LLM agents can spontaneously develop shared social conventions, how individual biases shape those conventions, and how minority agents can change them. It uses decentralized local coordination to study convention emergence organically within AI populations.
- Convention emergence: LLM populations may develop shared conventions through repeated communication, a foundational process for multi-agent systems and mixed human-LLM ecosystems.The paper focuses on conventions emerging organically from communicating AI agents rather than treating LLMs as proxies for human participants.
- Collective bias: The study examines whether individually unbiased agents can produce collective bias as repeated communications drive populations toward universal conventions.It distinguishes bias as an initial statistical preference between equivalent alternatives and notes that collective processes can suppress or amplify individual traits.
- Critical-mass dynamics: The study tests whether adversarial minorities can redirect LLM populations toward alternative conventions after reaching a critical mass.Understanding these dynamics may help anticipate and steer beneficial norms while mitigating harmful ones.
- Research questions: The paper addresses three questions: whether conventions emerge spontaneously in LLM populations, how individual biases influence them, and how critical-mass minorities can change them.These questions concern convention emergence, individual bias, and critical-mass dynamics in populations of LLM agents.
- Approach: The experiments model convention formation through naming coordination, giving agents local incentives so conventions may emerge unintentionally from decentralized interaction.This complex-systems design prioritizes transparent interpretation over high-fidelity simulation of human interactions.
Experimental Setting
The experiments simulate pairwise naming-game interactions among LLM agents, using memory-based prompts, coordination incentives, and finite pools of names. Four LLMs are tested, with adversarial agents added in norm-change experiments and comprehension checks confirming effective instruction following.
- Simulation design: Each trial randomly pairs agents who choose names from a finite pool, rewarding matches and penalizing mismatches to incentivize pairwise coordination without directly promoting global consensus.Agents’ scores change equally after each interaction, while partner selection and population membership remain unspecified in the prompt.
- Agent prompting: Agents forecast the next action from the most recent H interactions, using memory of prior choices, outcomes, and scores without prescribed rules for applying that history.The memory starts empty, so the first convention is randomly selected from the available name pool.
- Prompt validation: LLM comprehension tests showed good instruction-following capability, supporting the use of the prompting setup for the interaction experiments.A meta-prompting strategy posed text-comprehension queries and evaluated response precision; results are reported in Fig. S1.
Results
LLM populations spontaneously converge on shared linguistic conventions, while collective biases emerge through convention formation even without individual bias. Committed minorities can overturn established conventions once they reach model- and convention-dependent thresholds.
- Emergence of social conventions: A single linguistic convention becomes dominant across models by population round 15, except Llama-2-70b-Chat, with consensus remaining robust up to N=200 and W=26.The dynamics show a disorder-to-order transition and winner-take-all selection among competing names.
- Collective bias: Some names are disproportionately likely to become conventions across models, although preferred names vary, and collective bias can emerge even when agents are individually unbiased.Repeated communication and coordination can create a strong convention through dynamically formed memory states.
- Collective bias: After success, agents reuse the same name 99.4% of the time, whereas after failure they switch names 97.3% of the time, reinforcing asymmetric convention selection.The resulting bias arises from repeated interactions and diverse memory states rather than necessarily from isolated-agent preferences.
- Committed minorities and social change: Committed minorities trigger population-wide adoption at a critical threshold; observed thresholds range from 2% for Llama-3-70B-Instruct to 67% for Llama-2-70b-Chat.Below threshold, populations settle into mixed states, while stronger conventions require larger minorities to overturn.
Discussion
LLM populations can develop social conventions through local interactions, while collective biases and minority-driven tipping points shape which conventions emerge and persist. These findings motivate group-level alignment tests, caution about generalizing from synthetic settings, and further study of AI–human norm dynamics.
- Contributions: Local interactions can spontaneously produce universally adopted social conventions without central coordination, revealing collective dynamics in LLM populations.The study presents a framework for detecting higher-order biases arising from complex social interactions.
- Collective bias: Collective bias can favor some conventions over others, remain hidden in isolated-agent analyses, and vary across LLM models.The discussion identifies collective bias as a group-level phenomenon rather than a direct consequence of individual-agent behavior.
- Norm change: A committed minority can impose an alternative convention on a settled majority, with the tipping threshold depending on competing conventions and the LLM model.These tipping points create both opportunities for coordinated change and vulnerabilities to manipulation or injection attacks.
- Limitations and future work: The findings show that alignment must be evaluated at the group level, while broader validity requires testing different models, prompts, population sizes, social structures, conventions, and mixed LLM–human settings.The reported results come from parameter-dependent experiments with randomly paired agents, motivating studies in realistic networks and with sensitive human norms.
- AI–human comparison: AI populations show qualitative similarities to humans in shared-norm emergence and critical-mass dynamics, alongside LLM-specific collective-bias effects requiring further human testing.These similarities and differences affect whether LLMs can reliably model human social systems or be deployed in social settings.
Supplementary materials
The supplementary materials include references and a supplementary-materials section.
- The supplementary materials contain references numbered 8, 55–57, and 83–86.
- The document includes a section labeled “Supplementary Materials for.”
Emergent Social Conventions and Collective Bias in LLM · Populations1
The document identifies the authors and notes that it is a preprint version of a 2025 Science Advances publication.
- Populations1: The listed authors are Ariel Flint Ashery, Luca Maria Aiello, Andrea Baronchelli, and additional coauthors indicated by the ellipsis.The passage displays three author names and an asterisk after Baronchelli, followed by an ellipsis.
- Populations1: Ariel Flint Ashery and Andrea Baronchelli are affiliated with institution 1.Both names carry superscript affiliation marker 1.
- Populations1: Luca Maria Aiello is affiliated with institutions 2 and 3.Aiello’s name carries superscript affiliation markers 2 and 3.
- Populations1: Andrea Baronchelli is marked as a corresponding or otherwise designated author with an asterisk.The asterisk appears immediately after Baronchelli’s name.
- This PDF file includes:: The cited Science Advances publication carries the article identifier eadu9368 and the year 2025.The metadata gives the identifier and year explicitly.
Supplementary Text · Microscopic Bias
Microscopic bias persists across configurations and emerges rapidly even without an initial preference between conventions. Fine-tuning therefore cannot be identified as the cause of collective bias effects.
- Statistical Tests for Table 1: Exact Binomial tests found p† < 0.05 in every configuration, confirming that the model’s decisions were biased toward an extreme.
- Statistical Tests for Table 1: Bootstrapping 70% of observations 10,000 times yielded p‡ > 0.05 in every case, so the hypothesis of more extreme underlying bias could not be rejected.These results are reported in Table S5.
- Effect of Fine-Tuning: Experiments reproduced the Table 1 setup with Llama-3.1-70B without instruction fine-tuning to assess fine-tuning’s effect on strategy bias.The model was run with 4-bit quantization on a single A100 GPU, using next-token probabilities over valid convention tokens.
- Effect of Fine-Tuning: The pretrained model’s initial convention distribution was biased, with p(Q) = 0.525, before testing collective-bias dynamics.
- Effect of Fine-Tuning: By the second interaction, strong collective bias appears despite no initial preference for either convention, so fine-tuning is not established as its cause.The comparison used randomly selected convention pairs with Jensen-Shannon distance below 0.005 from a neutral no-memory distribution.
Theoretical Model
The naming game model shows how local pairwise negotiations can produce global consensus on social conventions. Its dynamics progress through innovation, propagation, and convergence, while the experimental setup restricts word invention to a finite pool W < N.
- Theoretical Model: Local pairwise negotiations in a population of N agents produce global consensus on conventions through coordination mechanisms.Agents negotiate names using only local interactions, paralleling the experimental framework.
- Theoretical Model: Randomly selected speaker–hearer pairs exchange words, retaining successful conventions or adding novel words after failed recognition.Agents begin with initially empty lexicons and can invent a word when the speaker has none.
- Theoretical Model: The model follows three temporal phases: innovation through word creation, propagation through lexicon reorganization, and convergence to global consensus.These are the model’s distinct non-equilibrium dynamical phases.
- Theoretical Model: The experimental framework limits invention to a finite word pool of size W < N, making the initial innovation phase extremely short.The shortened innovation phase is visible in the inset of Fig. 1.
Characterizing Relative Strength
Relative strength is defined by a convention’s ability to attract and retain consensus, and is characterized by the steady-state basin of attraction. The basin’s width captures convergence across initial conditions, while its pull captures resistance, recovery, and convergence rate.
- Characterizing Relative Strength: Relative strength also constrains convergence by restricting the system’s exploration along its path toward consensus in Llama-3.1-70B-Instruct populations.The system’s trajectory is sensitive to initial conditions, so not all memory configurations are observed.
- Characterizing Relative Strength: Relative strength is characterized by the strong convention’s steady-state basin of attraction, whose width measures convergent initial conditions and whose pull measures stability and return after deviations.The pull also captures the rate at which trajectories are drawn toward the steady state.
- Characterizing Relative Strength: Llama-3.1-70B-Instruct populations resist perturbations and return toward equilibrium around the strong convention, despite exploratory behavior without perturbation.Their exploration remains bounded in memory-state space.
- Characterizing Relative Strength: Llama-3-70B-Instruct populations remain at consensus once reached but are more susceptible to perturbations, briefly exhibiting memory states reflecting coordination on the weak convention.Some agents temporarily show memories in which all past interactions within their memory range reflect successful coordination on the weak convention.
Prompting
The prompting framework combines fixed rules, dynamic game memory, and response-format instructions, while testing answer-first reasoning and alternative prompt designs. Meta-prompting evaluates prompt comprehension, and experiments show generally strong comprehension alongside emergent bias across prompt variations.
- Prompt Structure: The system prompt combines fixed game rules, dynamic memory about play history, and instructions governing the agent’s response format.Agents nevertheless often treated their co-player as an opponent, sometimes accepting negative payoffs to reduce the co-player’s accumulated tally.
- Output Structure: The framework uses an answer-first-reason-later format instead of asking agents to reason before deciding, to reduce reasoning-driven bias and strengthen generalization.The authors argue that reason-first outputs can expose and amplify biases before the final decision.
- System Prompt: The example prompts specify a 100-round partnership game, simultaneous actions, payoff rules, accumulated scores, interaction history, and a structured action-and-reason response.The system prompt asks agents to think step by step while explicitly reporting the selected action and reason.
- Meta-Prompting: Because generative outputs lack a defined error target, meta-prompting tests comprehension of interaction rules, action chronology, and payoff statistics.Agents’ histories are replayed with their original memory lengths, and comprehension accuracy is averaged across interactions and agents.
- Meta-Prompting: Prompt comprehension accuracy is nearly always above 0.8 and often close to 1, with Llama-2-70b-Chat the only model falling below 0.8 on any metric.Its main weakness was counting how often a convention appeared within the memory range.
- Alternative Prompts: Alternative factual, bullet-point, and narrative prompts use randomized six-character convention names selected to keep the initial Jensen-Shannon distance below 0.005.These variations are designed to reduce bias from previously seen training examples and initial convention preferences.
- Alternative Prompts: Biased consensus emerges across all tested prompt and randomized-name variations, while the game’s long-term trajectory remains emergent rather than directed toward one action.The narrative and bullet-point templates provide distinct alternative representations of the same repeated two-player game.