Source-linked AI summary
Benchmarking large language model agent societies against human behavioural distributions
Raad Bin Tareaf
TL;DR
LLM agent societies need validation of their human fidelity, robustness to apparatus changes, and independence from memorised experiments. SILICA tests these questions across five human-anchored environments using perturbation, payoff-divergence, and contamination audits. The results support exploratory claims only: models often match starting points but not behavioural trajectories or incentive-sensitive decision functions.
Problem
LLM societies lack a calibrated instrument for testing human behavioural fidelity, robustness to apparatus changes, and contamination by memorised experiments.
Method
SILICA evaluates twelve open-weight models across five human-anchored environments with controlled perturbations, payoff-divergent variants, and contamination audits.
Results
Across tasks, agreement with human data is largely confined to starting points, while models show representation sensitivity and inconsistent incentive-sensitive decision functions.
Takeaways & Limitations
The certification ladder supports exploratory claims and no more for current silicon societies.
Takeaways & Limitations
The anchors mainly represent Western, educated, industrialised, rich, democratic populations and often summarize particular experiments rather than individual human distributions.
Abstract
from arXiv · showhide
Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.
1 Introduction
The paper introduces SILICA to test whether LLM societies reproduce human behaviour, withstand apparatus changes, and exhibit interaction rather than memorised experimental scripts. Its certification protocol answers these questions with exploratory claims only.
- Purpose: SILICA couples interactive multi-agent environments with published human anchors and tests fidelity, robustness, and provenance.The instrument uses versioned, seeded environments and separates perturbations that re-render information from those that alter it.
- Purpose: 56 of 111 design-level contrasts and 8 of 71 representation-level contrasts moved behaviour.Representation changes can matter despite conveying no new information, including a 58-point cooperation loss from reversing action order.
- Purpose: A fixed offer schedule identifies acceptance functions independently of the offers models generate as proposers.The audit distinguishes payoff-following, script-following, and mixed behaviour more directly than aggregate rejection rates.
- Purpose: The study overturned its own initial mechanism claim when independently permuting name pools restored coordination but made it slower or unreliable.Coordination took one and a half to two and a half times longer in seven models and stopped being reliable in the eighth.
- Purpose: The paper concludes that its certification ladder supports exploratory claims and no more.The article reports what findings permit and do not permit rather than treating the instrument as proof of robust human-like social dynamics.
2 Related work
Related work establishes generative social simulation, behavioural validation, and contamination concerns as distinct literatures. This paper occupies their intersection by testing interactive multi-agent claims against human anchors with perturbation and provenance audits.
- Human behavioural anchors: Human social experiments establish distributions, trajectories, local convention formation, and tipping thresholds as empirical anchors.These anchors come from programmes of experiments and meta-analyses rather than isolated point estimates in the broader literature.
- Validation: Agent-based modelling developed validation guidance around targets, tolerances, sensitivity, and non-identification from matching outputs.That guidance maps directly onto the paper’s requirements for empirical targets and robustness to modelling choices.
- Generative social simulation: Generative social simulation spans sandbox societies, large-scale simulators, individual-level sampling, conventions, tipping, and norm formation.This paper differs by treating such social findings as hypotheses tested against human anchors rather than as demonstrations of human-like behaviour.
- Games and validation: Existing game benchmarks measure strategic capability, while repeated-game studies report behavioural signatures that diverge from human ones.The paper extends this validation turn from strategic evaluation toward interactive, human-anchored social behaviour.
- Contribution: The article’s distinctive niche is the intersection of interactive, multi-agent, human-anchored, perturbation-audited, and contamination-controlled evaluation.Its contamination audit executes a proposed novel-game falsification design and extends it to collective dynamics.
3 Methods
The methods combine five environments with documented human anchors, perturbation and payoff variants, and a protocol that measures behavioural agreement against predefined equivalence margins. The environments span cooperation, public goods, bargaining, strategic reasoning, and convention formation.
- Environments: Five environments cover cooperation under temptation, shared-pool contribution, bargaining, iterated strategic reasoning, and convention emergence.Each was selected because its human result can be stated numerically and its provenance is documented.
- Environments: Repeated prisoner’s dilemma uses ten rounds and canonical payoffs (R, S, T, P) = (3, 0, 5, 1), anchored to cooperation levels and decline over rounds.The human anchor includes average cooperation rates, declining cooperation, and increased final-round defection.
- Environments: Public goods gives four agents 20 points per round, with marginal per-capita return 0.4 and optional costly punishment.The anchors include first-round contribution near half the endowment, decline over repetition, and conditional cooperation.
- Environments: The 11–20 money-request game uses integer requests from 11 to 20, a one-below bonus, and a human distribution whose mode is 17.Each run contains 50 one-shot games.
- Environments: The naming game uses 24 agents, ten names, five-interaction memory, a 95% convention criterion, and a 3,000-interaction cap.Its anchors are guaranteed convention formation and a 21–25% critical mass; one 11–20 anchor was corrected during the study.
3.2 The perturbation library
The perturbation library changes one apparatus factor at a time, distinguishing re-renderings from changes in agent-provided content. Payoff-divergent variants and fixed offer schedules then test whether behaviour follows incentives or familiar scripts.
- Library design: Each perturbation changes exactly one factor while retaining the canonical baseline for all other components.The library includes persona, framing, history, action-label order, names, temperature, memory, and group-size factors.
- Library design: Representation-level cells re-render existing information, whereas design-level cells alter what agents are given.Examples include table-versus-sentence personas and reversed action order versus prosocial or risk framing.
- Variants: Four environment versions progress from canonical C0 through re-skinned C1 and shifted-number C2 to payoff-divergent C3.C3 makes the literature’s recorded result no longer profitable, including dominant cooperation and rewards for mismatching or exceeding.
- Variants: Responders receive the same fixed offers from 10 through 50 points, twenty games per offer, in canonical and divergent-payoff conditions.This isolates each model’s acceptance function from the proposer behaviour it would otherwise encounter.
- Variants: Behaviour in divergent-payoff conditions is classified as following payoffs, following the script, or mixed using thresholds τ∗ = 0.70 and τs = 0.30.Ultimatum decisions are evaluated per game against the offer faced rather than by aggregate rejection rate.
- Variants: Randomly permuting each agent’s name pool tests whether naming-game coordination depends on a shared presentation order.Unparsable replies fall back to a random name rather than the first listed option.
3.4 Models, serving and decoding
The study serves twelve open-weight models locally through vLLM on one consumer graphics card, using recorded precision and controlled decoding settings.
- Twelve open-weight models were served locally with vLLM through an OpenAI-compatible interface on a single consumer graphics card.Quantisation varied across the roster and was recorded per model.
- Sampling temperature was 0.7 except in the dedicated temperature cell.
- Optional Qwen3 deliberation was disabled so reasoning differences were carried by the single reasoning-trained model rather than a decoding switch.Deliberation markers emitted by models were stripped before parsing.
3.5 Run structure
The experiment combines repeated seeded runs with descriptive one-shot cells and evaluates primary outcomes against pre-registered human-equivalence margins and distributional measures.
- Run structure: Each repeated-game cell contains 30 seeded baseline runs per model and 10 runs per perturbation cell.Seeds control agent order, pairing, and sampling.
- Run structure: Each one-shot cell contains 6 baseline runs and 2 perturbation runs, each covering 50 games; perturbation results are descriptive.Two runs cannot support an inferential test.
- Outcomes: Primary outcomes are cooperation, contribution fractions in rounds 1 and 10, offer fractions, human-mode choice shares, and convention formation.
- Outcomes: Human-anchor agreement requires a 90% confidence interval for the model–human difference to lie within a pre-fixed outcome-specific margin.Margins are ±0.05 for ultimatum offers and dictator giving, ±0.07 for trust quantities, and ±0.10 for public-goods contributions.
- Outcomes: For the 11–20 game, agreement is measured by normalized Wasserstein distance between pooled model choices and the published distribution.A bootstrap interval is computed over runs.
- Perturbations: Perturbation stability is summarized by the range of cell means and the share of cells retaining the canonical human-effect sign, reported separately by perturbation level.
- Outcomes: Directional agreement requires the median baseline run-level effect to exceed 0.10 on the outcome’s own scale.
3.7 Statistical procedure
Statistical comparisons use exact nonparametric rank-sum tests, Cliff’s δ, and Holm adjustment within each test family.
- Statistical procedure: Cell-versus-baseline comparisons use the Mann–Whitney rank-sum test against the exact permutation distribution conditional on observed ties.This avoids normal-approximation p-values below the smallest attainable value for the sample sizes.
- Statistical procedure: Effect sizes are reported as Cliff’s δ.
- Statistical procedure: P-values are adjusted with the Holm step-down procedure separately within perturbation, priming, punishment, and contamination families.
3.8 Certification
Certification is assigned to claims rather than models, with higher tiers requiring baseline support, perturbation robustness, divergent-payoff resistance, and human equivalence. The study’s evidence is bounded by partial registration and single-author analysis.
- Certification tiers: Tier 1 marks an exploratory claim that holds at baseline in at least one model.
- Certification tiers: Tier 2 additionally requires sign stability of at least 0.8 across cells in at least three model families and no reversal under divergent payoffs.
- Certification tiers: Tier 3 additionally requires equivalence with the human anchor rather than directional agreement alone.
- Certification tiers: Table 2 organizes each research question by assembled evidence and highest supported certification tier, excluding R1-Distill where fewer than three runs make comparisons unavailable.It also distinguishes equivalence from significance and reports acceptance thresholds rather than potentially misleading aggregate consistency scores.
- Scope and limitations: The released record is described as a timestamped public deposit because it was not converted into a frozen OSF registration.
4 Results
Across SILICA’s environments, language-model agents resemble human behaviour mainly at starting points, while trajectories, robustness, incentive responsiveness, and convention formation vary substantially across models and perturbations.
- Fidelity: agreement is confined to starting points: Eight of eleven models match the human-equivalence margin for first-round public-goods contributions, but none matches final-round contributions.Ultimatum offers match the human mean for one of twelve models, dictator giving for none, and no model falls inside the prisoner’s-dilemma cooperation corridor.
- Fidelity: agreement is confined to starting points: One model reproduces the canonical decline in public-goods contributions, three reproduce declining cooperation in the repeated dilemma, and ten reproduce convention formation.A conditional-cooperator baseline declines from 0.48 to 0.07, while ten of eleven language models miss that human trajectory.
- Fidelity: agreement is confined to starting points: Costly punishment lowers final-round contributions in four of nine models, by as much as 0.53 of endowment.Punishment is purchased in a median 97% of runs, and spending correlates negatively with contributions (Spearman ρ = −0.84, p = 0.001, k = 11).
- Robustness: form matters, content matters more: At least one perturbation significantly shifts behaviour in 23 of 27 tested model-by-environment blocks, with design-level changes affecting 56 of 111 contrasts versus 8 of 71 representation-level contrasts.Thus, most fragility responds to altered content, while a smaller but consequential share responds to form alone.
- Robustness: form matters, content matters more: Reversing action-label order costs Qwen3-14B 58 points of cooperation, while changing persona format ends Phi-3’s convention formation from 0.90 to 0.00.Persona content and step-by-step reasoning can also remove cooperation entirely or cost Qwen3-8B 99 points.
- Provenance: retrieval, thresholds, and non-response: Fixed offer schedules reveal heterogeneous acceptance thresholds: only R1-Distill places its threshold where the divergent payoff requires it.Two models move partway, two move in the wrong direction, and three never acquire a threshold; Qwen3-32B instead adopts the human-like boundary near one-third of the pie.
- Provenance: retrieval, thresholds, and non-response: Naming-game convergence follows a shared prior over labels rather than negotiation: permuting each agent’s list preserves the winning name but slows convergence to 75–129 interactions.Convergent models remain faster than the classical minimal naming game, whose median is 2,076 interactions.
5 Discussion
SILICA supports exploratory use of LLM societies for how interactions begin, but not for predicting their development, institutional effects, or mechanisms without stronger controls.
- Supported uses: Eight of eleven models matched human first-round public-goods contributions within the equivalence margin, but later contributions and cooperation did not match human patterns.Costly punishment lowered contributions in four of nine comparable models, contrary to the human effect.
- Mechanisms: ρ = −0.84: models that punished most were those whose public-goods contributions collapsed furthest.The two models whose contributions did not move were also the two that never purchased punishment.
- Mechanisms: One of twelve models placed its rejection threshold where the outside option required; others moved partially, incorrectly, or not at all.Fixed offer schedules distinguish incentive-responsive thresholds from aggregate rejection patterns that merge different behaviours.
- Mechanisms: Convention formation initially looked like retrieval because six models converged immediately, but shuffled name orders restored slower negotiation.With permuted name pools, convergence took 75 to 129 interactions instead of the 50 interactions detectable under the original ordering.
- Certification: A benchmark can support a mechanism claim only when outcomes survive apparatus perturbations and transcript audits test whether the mechanism matches the human experiment.The study’s missing manipulation initially allowed a confident but wrong interpretation of convention formation.
- Certification: Convention formation alone reached Tier 2, while no finding reached the top certification tier.Its sign stability met the stated criterion across seven models from three families, but its negotiation remained two to twenty-eight times faster than classical dynamics.
- Limits: The evidence is bounded by summary-level human anchors, predominantly Western samples, open-weight models small enough for one consumer graphics card, and limited reasoning-model coverage.The reasoning axis rests on one distilled reasoning-trained model, while two runs per perturbation cell support description rather than inference.
6 Conclusion
Current silicon societies reproduce where human interactions begin, not how they proceed, and support exploratory claims rather than stronger certification.
- Conclusion: LLM societies reproduce interaction starting points but not subsequent development, respond to apparatus form as well as content, and often lack a decision function when payoffs conflict with memorised experiments.The benchmark identifies specific missing components of human social behaviour and releases the instrument, anchors, and run record for further testing.
Declarations
The study reports no external funding or competing interests and uses only published human comparison data; its records, statistics, anchors, transcripts, and code are openly available.
- Declarations: No external funding was received, no competing interests were declared, and the study involved neither human participants nor animals.Human comparison data came exclusively from published, publicly available datasets.
- Data availability: The complete run-level record, aggregated statistics, exact test-level results, anchor provenance, transcripts, and code archive are openly available.The released corpus includes 9,115 model runs, 150 classical-baseline runs, and 8,220 transcript files.