Source-linked AI summary

Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias

Messi H. J. Lee

arXiv:2609.00009v1physics.soc-phcs.CY

TL;DR

As language-model agents increasingly interact in groups, existing evaluations leave their social behaviour unmeasured. This paper adapts the minimal-group paradigm into a controlled allocation probe and finds that arbitrary group labels induce in-group favouritism, concentrated among numerical-minority deciders and absent when labels are hidden.

  • Problem

    Existing evaluations of language-model agents probe memorised stereotype content or simulate people, leaving group-based social behaviour unmeasured.

  • Method

    The study adapts the minimal-group paradigm into a point-allocation probe across four reasoning models, varying minority size and whether arbitrary group labels are visible.

  • Results

    Arbitrary categorisation induced in-group favouritism that vanished when labels were hidden and concentrated among numerical-minority deciders, while majority deciders allocated close to proportionally and the asymmetry closed at equal group sizes.

  • Takeaways & Limitations

    Open-weight reasoning models display a behavioural signature of intergroup discrimination independent of demographic stereotype content, suggesting social-psychology methods can measure and help govern this behaviour.

  • Takeaways & Limitations

    Under a paraphrased instrument, R1-Distill-Llama-8B showed a genuine hidden-condition floor violation, limiting robustness of the group-blind control for that model.

Abstract

from arXiv · show

Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.

1 Introduction

As language-model agents increasingly interact in groups, existing stereotype audits and human-simulation studies do not measure how models themselves behave toward grouped peers. This study adapts the minimal-group paradigm to test whether arbitrary categorisation produces resource-allocation bias in reasoning models, including how relative group size shapes it.

  • Research gap: Existing LLM evaluations mainly examine memorised stereotype content or use models to simulate people, rather than measuring models as agents whose actions affect others.The study instead treats the model as the subject of an interactional behavioural probe.
  • Method: The design varies minority size and label visibility, with every agent allocating points so both minority- and majority-decider perspectives are observed within each society.A labels-withheld condition provides a group-blind reference, while direct manipulation of relative group size tests the structural asymmetry emphasized by social identity theory.
  • Method: The probe asks an agent to divide 100 points among 19 anonymous peers distinguished only by neutral, meaningless group labels.In-group favouritism is scored as the mean allocation to same-group targets minus the mean allocation to other-group targets.
  • Contribution: Across four open-weight reasoning models, arbitrary categorisation produced substantial in-group favouritism that vanished under the group-blind control.The result is presented as behaviour revealed when the model is placed in a group structure, rather than as stereotype content read from model weights.
  • Contribution: The bias was concentrated in numerical minorities, whose own-group allocations exceeded their numbers, while majority members allocated close to proportionally.The asymmetry closed at equal group sizes, linking the effect to relative group size rather than the arbitrary labels themselves.

2 Results

Across four reasoning models, visible arbitrary group labels produced in-group favouritism, while group-blind controls stayed at zero. The effect was concentrated among minority deciders, and disabling reasoning increased overall favouritism but largely removed the minority–majority asymmetry.

  • Design: The probe tested four reasoning models across minority sizes of 3, 5, and 10, with labels either visible or hidden across 30 seeded societies per cell.Each society contained 20 deciders, each allocating 100 points among the other 19 agents.
  • 2.1 The group-blind control validates the instrument: Group-blind favouritism was indistinguishable from zero for every model and minority size.The dyad-level floor remained within ±0.03 points/target, with |z| < 1.5 and all p > 0.13; all twelve seed-level tests were non-significant.
  • 2.2 Categorisation alone elicits in-group favouritism: 0.84–4.06 points/target: visible labels elicited positive in-group favouritism in all four models, with every estimate at p < 10−25.R1-Distill-Qwen-14B had the strongest effect at 4.06, while Qwen3-8B (2.31) exceeded Phi-4-reasoning (2.09).
  • 2.3 The bias is concentrated in minority deciders: +2.38 to +6.84 points/target: minority deciders favoured their in-group more than majority deciders in every model.The size-robust excess-share measure confirmed the same minority-concentration pattern despite the per-target measure’s sensitivity to group size.
  • 2.3 The bias is concentrated in minority deciders: R1-Distill-Llama-8B showed a narrow exception: at minority size 3, majority deciders gave the small out-group more than a per-capita-even split predicts.Its excess share was −0.037 and its points/target contrast was −1.45, both p < 10−9; the model was also the noisiest and most paraphrase-sensitive.
  • 2.4 A matched instruction-tuned control: reasoning and the minority asymmetry: 3.81 versus 2.31 points/target: disabling Qwen3-8B reasoning increased overall visible-label favouritism, but reduced the minority−majority gap from +4.08 to +0.54.The interaction was −1.50, with p = 3 × 10−83; the gap remained distinguishable from zero at p = 0.002.

3 Discussion

The controlled minimal-group probe found that arbitrary labels induce in-group favouritism in reasoning models, with the effect concentrated among numerical minorities. The study frames this as a behavioural analogue of human intergroup discrimination while limiting claims about internal mechanisms and broader model classes.

  • Findings: Arbitrary, meaningless labels were sufficient to induce in-group favouritism across four reasoning models, while concealing labels removed the effect.The result is distinct from reproducing stereotypes or expressed attitudes about demographic groups.
  • Reasoning control: Disabling reasoning in Qwen3-8B did not remove favouritism but sharply reduced the minority–majority asymmetry.The matched control therefore complicates claims that reasoning causes the bias or determines whether it appears.
  • Findings: Minority deciders over-allocated resources to their own group relative to group size, whereas majority deciders allocated close to proportionally.The asymmetry disappeared when the groups were equal in size.
  • Implications: Minimal-group methods could extend machine-behaviour evaluation from individual cognition to social and collective regularities relevant to multi-agent deployments.The authors propose adapting established paradigms for conformity, authority, minority influence, and related phenomena.
  • Scope: The study’s open-weight sample comprises four reasoning models in the 8–14B range, so whether the pattern extends to closed, larger deployed systems remains untested.The probe is also a one-shot allocation task rather than sustained interaction.
  • Interpretation: The evidence establishes what models do and where bias concentrates, but not the mechanisms generating these regularities.Possible sources include training data, instruction tuning, and preference optimisation.

4 Methods

The methods transplant the human minimal-group allocation task into a controlled, machine-readable probe. The design varies group size and label visibility, evaluates four reasoning models, includes a matched no-reasoning control, and uses society-level statistical analyses.

  • 4.1 The allocation probe: Each trial asks one of 20 agents to distribute 100 points among the other 19, whose arbitrary group labels are either visible or hidden.Every agent serves as decider in turn, yielding minority- and majority-perspective allocations.
  • 4.2 Design: The experiment crosses minority sizes of 3, 5, or 10 with visible or hidden labels across 30 independent societies per cell.The hidden arm validates the zero-favouritism floor because labels provide no conditioning cue.
  • 4.3 Controls: Neutral invented group names are swapped between groups, and minority membership is randomised across agent indices to control label valence and position.These counterbalancing procedures isolate categorisation from name and roster-position effects.
  • 4.4 Models and generation: Four reasoning models were evaluated with identical templates, temperature 0.7 sampling, and up to 8192 generated tokens.The models were Qwen3-8B, two DeepSeek-R1 distillations, and Phi-4-reasoning.
  • 4.5 Reasoning control: A matched Qwen3-8B rerun disabled thinking while keeping prompts, rosters, counterbalancing, parsing, and sampling settings unchanged.This isolates the role of reasoning mode within one checkpoint.
  • 4.6 Outcome measure: Favouritism is the difference in mean points per target received by in-group versus out-group recipients, with zero indicating equal treatment.The neutral even-split reference is 100/19 ≈5.26 points per target.
  • 4.7 Statistical analysis: The exploratory analysis treats society seeds as the unit, using seed-level Wilcoxon tests alongside dyad-level GEE and robustness checks.The statistical treatment was developed alongside the data, and an earlier pseudo-replicated seed test was corrected.

Ethics declarations

The study used no human participants, human data, or animal subjects, so ethics approval was not required.

  • Ethics declarations: All trials were allocations generated locally by open-weight language models, without human participants, human data, or animal subjects.The authors therefore state that no ethics approval was required.

Supplementary Information

The supplementary materials document implementation, statistical, robustness, paraphrase, and reasoning-control analyses, with numerical results and figures regenerated from trial-level data.

  • Supplementary Information: Supplementary materials cover models and generation, prompts, parsing, quality control, extended statistics, robustness, and matched instruction-tuned controls.They include Tables S1–S15.

S1 Models and generation

The study evaluated four openly available reasoning language models locally, using identical prompts and generation settings.

  • Four openly available reasoning language models were evaluated locally.
  • All models were loaded in bfloat16 and run with Hugging Face transformers.
  • Generation used sampling at temperature 0.7 with a maximum of 8192 new tokens.
  • The 20 deciders of each society were generated in a single left-padded batch.

S2 Prompt templates

Each trial asks an agent to distribute a fixed point pool among anonymous participants, with neutral group labels shown only in the visible condition and strict JSON output constraints.

  • Each trial gives the agent a pool of 100 points to distribute among other participants.
  • In minimal-group runs, no personality is supplied and group-label tags appear only in the visible condition.
  • The required response is exactly one JSON object mapping participant ids to whole-number allocations summing to the pool.
  • The prompt identifies the decider and displays the other participants with their group labels.

S3 Parsing and quality control

Model outputs were cleaned, parsed, restricted to valid roster ids, and renormalised before scoring; parse success exceeded 98.8% for every model.

  • Reasoning traces and code-fence markers were removed before JSON parsing.
  • Parsed allocations were restricted to valid roster ids and renormalised to the pool.
  • 98.8% parse success was exceeded by every model.

S4 Extended statistical results

The statistical analysis reports seed-level favouritism tests and dyad-level generalized estimating equations across label conditions and minority–majority interactions, alongside a size-robust measure.

  • Table S3 reports seed-level favouritism in points/target for every design cell.
  • Table S3 includes one-sample Wilcoxon signed-rank tests against zero.
  • Table S4 reports dyad-level GEE coefficients under both label conditions and the minority−majority interaction.
  • Table S5 reports a size-robust group-level measure because points/target is sensitive to group size.

S5 Robustness of the effect

A suite of robustness checks shows that visible-label favouritism is widespread, label-dependent, and not explained by token valence, roster position, parsing choices, or a few extreme allocations.

  • Individual-level prevalence: 78.7%–99.8% of minority deciders favoured their in-group across models, while group-blind prevalence stayed below 50%.The individual-level distributions were predominantly above zero for both decider groups, whereas the group-blind distribution was centred on zero.
  • Group-blind control: Group-blind favouritism stayed within ±0.03 points/target of zero across models and minority sizes.All twelve group-blind cells were non-significant, supporting the label-based interpretation of the visible-arm effect.
  • Nonce-label counterbalancing: Swap-pass differences were at most 0.41 points/target, and token-receipt standard deviations were at most 0.24 points/target.These checks support cancellation of nonce-token valence relative to favouritism effects of 0.8 to 4.1 points/target.
  • Roster position: A small position effect of −0.02 to −0.03 points per roster position could not drive favouritism because group membership was randomised with respect to position.In- and out-group targets had balanced presented positions.
  • Parsing robustness: Removing uniform fills changed favouritism estimates by at most 0.03 points/target, with only 10–51 filled trials per model.The parse-failure handling was therefore inconsequential in both visible and group-blind arms.
  • Extreme allocations: Fully out-group-excluding allocations reached 30.4% for Qwen3-8B but were at most 1.2% in the group-blind control.The rate peaked at intermediate minority size, so it is reported as a descriptive robustness check rather than a dose–response.
  • Multiple-comparison correction: 21 of 24 visible cells survived Benjamini–Hochberg correction, compared with 1 of 24 group-blind cells.Every minority-decider block survived correction, with rank-biserial effect sizes up to +1.00.

S5.7 Seed-level cluster bootstrap

Additional checks test uncertainty, prompt wording, and a narrow model-specific reversal. Bootstrap estimates corroborate the main inference, while paraphrasing preserves direction but reveals magnitude and one null-floor sensitivity.

  • Seed-level cluster bootstrap: Bootstrap and analytic standard errors agreed within 0.003 points/target for both favouritism and the minority−majority interaction.The bootstrap used 2,000 resamples of the 30 seeds, and every bootstrap 95% interval excluded zero.
  • Paraphrase robustness: The paraphrase retained positive favouritism in 22 of 24 visible cells and kept the GEE coefficient positive in every model×size cell but one.The reduced design used 15 matched seeds and preserved the task’s structural elements while changing surface wording.
  • Paraphrase robustness: Three models kept paraphrase hidden-condition favouritism within ±0.28 points/target of zero, but R1-Distill-Llama-8B rose to 0.74–1.31.The exception was statistically significant at society and majority levels, making that null condition wording-sensitive.
  • Model-specific exception: R1-Distill-Llama-8B majority deciders at minority size 3 showed −1.45 points/target and −0.037 excess share.Every one of 30 societies showed negative mean favouritism, but the reversal was absent at sizes 5 and 10 and did not survive paraphrasing.

S6 Matched instruction-tuned control

A matched Qwen3-8B comparison tests whether extended reasoning is necessary for the disposition or its minority asymmetry. Disabling reasoning leaves favouritism intact but greatly reduces the minority–majority gap.

  • Control validation: Both reasoning modes had clean group-blind floors: 0.011 ± 0.008 and −0.003 ± 0.004 points/target.Both tests had p > 0.13, and parse success ranged from 98.3% to 100%.
  • Overall favouritism: 3.81 versus 2.31 points/target: non-thinking Qwen3-8B showed more overall favouritism than thinking-enabled Qwen3-8B.The same-group×reasoning interaction was −1.50, p = 3 × 10^-83.
  • Minority asymmetry: The minority−majority gap fell from +4.08 to +0.54 points/target when reasoning was disabled.The non-thinking gap remained distinguishable from zero, with p = 0.002.
  • Minority asymmetry: Thinking-enabled majority deciders had +0.01 excess share, versus +0.15 for non-thinking majority deciders.Minority deciders had +0.26 and +0.17 excess share in the thinking-enabled and non-thinking modes, respectively.
  • Scope: This single matched pair does not establish that reasoning affects the asymmetry similarly in other model families or instruction-tuned models generally.A four-family reasoning-versus-instruction-tuned comparison is identified as the natural next step.
Loading 2609.00009v1…