Source-linked AI summary
LLM Social Simulations Are a Promising Research Method
Jacy Reese Anthis, Ryan Liu, Sean M. Richardson, Austin C. Kozlowski, Bernard Koch, James Evans, Erik Brynjolfsson, Michael Bernstein
TL;DR
LLM social simulations could provide accessible synthetic data for studying human behavior, but their adoption and results remain limited. This position paper reviews comparisons and related work, identifies five tractable challenges, and argues that sims can cautiously support exploratory research while conceptual models and iterative evaluations are developed.
Problem
LLM social simulations could address limitations of human data, but results remain limited and few social scientists have adopted the method.
Method
The paper reviews empirical comparisons between LLMs and human research subjects, commentaries, and related work to identify five tractable challenges and promising research directions.
Results
91% of the variation in average treatment effects was predicted by GPT-4 in a 70-experiment test after adjusting for measurement error.
Takeaways & Limitations
LLM social simulations can now be cautiously used for exploratory social research, while researchers develop conceptual models and iteratively refine evaluations.
Takeaways & Limitations
Sycophancy may make simulations scientifically inaccurate by steering outputs toward results that researchers appear to want.
Abstract
from arXiv · showhide
Accurate and verifiable large language model (LLM) simulations of human research subjects promise an accessible data source for understanding human behavior and training new AI systems. However, results to date have been limited, and few social scientists have adopted this method. In this position paper, we argue that the promise of LLM social simulations can be achieved by addressing five tractable challenges. We ground our argument in a review of empirical comparisons between LLMs and human research subjects, commentaries on the topic, and related work. We identify promising directions, including context-rich prompting and fine-tuning with social science datasets. We believe that LLM social simulations can already be used for pilot and exploratory studies, and more widespread use may soon be possible with rapidly advancing LLM capabilities. Researchers should prioritize developing conceptual models and iterative evaluations to make the best use of new AI systems.
1. Introduction
LLM social simulations could expand access to human-behavior data, but their accuracy and verifiability remain uncertain. This position paper identifies five tractable challenges, reviews promising evidence, and argues for cautious exploratory use alongside iterative development.
- Motivation: LLM simulations could reduce limitations of human data, including representative-sampling difficulties, financial costs, and non-response bias.They may also support counterfactual research, large-scale policy pilots, and human-centered AI development.
- Research agenda: The paper organizes the research agenda around five challenges: diversity, bias, sycophancy, alienness, and generalization.Its argument draws on empirical comparisons between humans and LLMs, commentaries, and related work.
- Evidence: 91% of variation in average treatment effects was predicted by GPT-4 across 70 preregistered, U.S.-representative experiments after adjusting for measurement error.This result came from a straightforward prompting technique in the largest test of simulations to date.
- Evidence: Fine-tuned Llama-3.1-70B outperformed existing cognitive models after training on data from 160 human-subject experiments.
- Evidence: Simulator agents predicted survey responses 85% as well as participants’ responses two weeks earlier, accounting for test-retest variation.The study built 1,052 individual simulations from interview transcripts in a U.S.-representative sample.
- Future direction: Most studies use only a small fraction of methods that could improve simulation accuracy, leaving substantial room for improvement.The authors therefore support cautious exploratory use now and iterative conceptual models and evaluations as capabilities advance.
2. Scope
The paper defines LLM social simulations as a research agenda focused on generating accurate, verifiable behavioral data through language modeling. Its scope is empirical and literature-based, while normative questions are treated separately.
- Definition: LLM social simulations use language modeling to generate accurate and verifiable data usable as if collected from human research subjects.
- Scope: The scope includes simulations with or without agency and overlaps with research using social science methods to study LLMs, human-LLM comparisons, roleplay, and non-LLM simulations.
- Review procedure: The review summarizes primary-scope studies in Table A1 and selected related work in Table A2, covering work preprinted or published by May 31, 2025.
- Scope: The paper focuses on how simulations could be developed and discusses whether they should be developed only briefly in Appendix B.
3. Challenges
LLM social simulations face five overlapping challenges: diversity, bias, sycophancy, alienness, and generalization. These challenges concern accurately representing human variation and behavior across groups, mechanisms, contexts, and research settings.
- 3.1. Diversity: Diversity requires simulations to match variation in the human population, but LLMs can produce highly uniform responses and narrower opinion distributions.In the 11–20 money request game, LLMs almost always chose 19 or 20, whereas humans had a median choice of 17 and wider variation.
- 3.2. Bias: Simulation bias is systematic inaccuracy when representing particular social groups and should be distinguished from diversity.A simulation can represent an underrepresented group frequently while still portraying that group homogeneously as a stereotype.
- 3.2. Bias: Stereotypes can either increase or decrease simulation accuracy, depending on whether they reproduce real group patterns or stereotypical portrayals.Avoiding the stereotype that most Fortune 500 executives are male could reduce accuracy when simulating that population.
- 3.3. Sycophancy: Sycophancy can reduce simulation accuracy because instruction-tuned models may optimize for user approval rather than produce typical human behavior.Human research data used to improve simulations may also contain social desirability bias, complicating the establishment of ground truth.
- 3.4. Alienness: Alienness arises when LLMs superficially match human behavior while relying on non-humanlike mechanisms and training data that does not cover human behavior in full.Studies report strong broad-level replication but weaker item-level performance, while internet data reflects what people say online rather than what they do in the real world.
- 3.5. Generalization: Generalization is necessary for widespread use, including reliable transfer beyond standard accuracy measures, populations, and contexts.Researchers must also understand how and when simulations generalize, not merely whether they do.
4. Promising Directions
Promising directions focus on richer context, distribution-level generation, steering, and model adaptation, while iterative evaluation tests whether simulations generalize to human behavior across groups and novel data.
- Context-rich prompting: Context-rich prompting can improve diversity, but explicit demographics may encourage stereotyping and should be supplemented with varied, conditionally realistic signals and more participant data.Implicit demographic cues can also correlate with socioeconomic status and criminal-record assumptions; broader variation may support a more nuanced subject representation.
- Distribution elicitation: Distribution elicitation shifts simulation from generating one subject’s response to modeling a distribution, potentially reducing diversity and bias problems and harnessing sycophancy-aware expert forecasting.The paper contrasts direct subject roleplay with prompts asking LLMs to predict how people respond to messages.
- Steering vectors: Steering vectors inject semantic or undirected variation into embedding space, but their ability to match human diversity or reliably alter behaviors remains uncertain.The paper recommends caution because superposition, disputed linear representations, limited usefulness, and detrimental side effects complicate applied steering.
- Training and tuning: Training entirely new models is generally inaccessible, so researchers may instead fine-tune instruction-tuned models with more diverse human instruction data or use available base models cautiously.Base models may reduce instruction-tuning distortions but can encode a limited subset of human behavior, be difficult to prompt effectively, and are often unavailable.
- Training and tuning: Fine-tuning on social-science data can improve alignment with human cognition, as Centaur’s internal representations predicted whole-brain fMRI data more consistently than those of its base model.Both Centaur and Llama-3.1-70B outperformed a traditional cognitive model on fMRI prediction.
- Iterative evaluation: Novel-data comparisons are essential for iterative evaluation, because group-level prediction can fall substantially below population-level prediction and human-difficult cases remain underexamined.In one reviewed comparison, GPT-4 predicted 91% of adjusted treatment-effect variation versus 84% for laypeople, while experts performed about as well as GPT-4.
5. Applications
The paper envisions progressively broader uses of LLM social simulations, beginning with exploratory research and extending to complete studies where human research is impractical or impossible. These applications still require verifiability and ethical conduct.
- Exploratory applications: Sims can support pilot studies that surface issues, estimate effect sizes, and inform methodological choices for later human-subject research.They can also support theory-building, brainstorming, and sketching future studies, where exploratory errors typically have lower costs.
- Replication and sensitivity analysis: Sims can enable exact replication, combined-data analysis, and sensitivity analysis across methodological counterfactuals.These uses may increase or decrease confidence in findings, bolster generalizability, and guide which analyses to run with more costly human studies.
- Complete studies: As the five challenges are addressed, sims could support end-to-end studies when human research is impractical because of limited funds or impossible, such as testing large-scale policy change.Large-scale simulations could be partially validated through tractable human studies of individual components.
- Requirements: Across applications, the proximate goal is accurate findings that generalize current knowledge, while simulations must remain verifiable and ethical under established standards.Using sims does not eliminate broader social-science requirements such as confidence in results and adherence to standards such as the Belmont report.
6. Alternative Views
Alternative views emphasize that LLM social simulation is new and that initial studies have reported significant challenges and limited results. Critics also argue that LLMs may be fundamentally unlike humans, posing a deep obstacle to humanlike prediction.
- Critiques of current evidence: Because the field is new, discussion of its promise remains limited, while reviewed works emphasize significant challenges and limits in initial results.These concerns include problems such as diversity and doubts about whether simulations can meet foundational values of human-participant research.
- Theoretical objections: Critics describe LLMs as fundamentally unlike humans and possessing ineradicable defects, implying deep limitations for social simulation.The concern is especially relevant because accurate social simulation requires humanlike systems or systems that can sufficiently understand humanlikeness.
7. Conclusion
The paper argues that LLM social simulation difficulties are tractable rather than insurmountable, while acknowledging uncertainty about future AI progress. It presents sims as a promising method for advancing social science and generating humanlike synthetic data for AI development.
- Conclusion: The paper develops an agenda organized around five tractable challenges to show that difficulties in LLM social simulation are not necessarily insurmountable.The argument accounts for uncertainty about future AI progress while responding to stronger claims that the endeavor may be impossible.
- Conclusion: Sims are presented as a promising research method that could accelerate social science, open new areas of inquiry into human behavior, and generate synthetic data for safe and beneficial AI.The paper links these prospects to thoughtful use of contemporary methods and future AI capabilities.
Impact Statement
The paper situates LLM social simulations against persistent limitations in human research data, including publication bias, resource constraints, and reduced accessibility. It also notes substantial potential benefits alongside risks of irresponsible use.
- Societal impacts: Effective human-population simulations could advance scientific research, train human-compatible AI, and address other societal issues.These potential benefits are accompanied by risks including online propaganda and manipulative bots.
- Evidence base: The paper reviews empirical comparisons between LLM-generated data and human-subject data alongside commentaries and related work.This review frames both the limitations of human data and the evidence for the promise of LLM simulations.
- Impacts on research: Publication bias from professional incentives can distort understandings of human behavior by limiting which studies and analyses are published.This concern is presented within the broader replication crisis and the emergence of questionable research practices.
- Impacts on research: Additional resources could improve human data through larger samples, harder-to-reach populations, engaged participation, and replication, but these remedies remain infeasible at current funding levels.The paper notes that wealthier institutions have made progress while much social-science data collection remains under-resourced.
- Access and equity: Increasing resource requirements make reliable social-science research less accessible to low-resource populations, particularly researchers in the Global South.The paper identifies reduced participation as compounding other issues in social science.
A.1.2. FUNDAMENTAL LIMITATIONS
Human research data are constrained by logistical infeasibility, sampling and self-report biases, and limited laboratory realism. LLM simulations offer a potentially cheaper and more scalable complement, but humanlike task performance may not ensure reliable out-of-distribution generalization.
- Fundamental limitations: Many important questions, including world leaders’ psychology, large-scale policy effects, and counterfactuals, are infeasible to study with real-world data.Some questions concern future possibilities or past alternatives rather than directly observable events.
- Fundamental limitations: Quasi-experiments and laboratory reconstructions provide causal evidence only for limited subsets of real-world conditions.Natural disasters rarely satisfy the assumptions needed for quasi-experimental inference, while laboratory facsimiles can only partially match reality.
- Fundamental limitations: Human surveys and experiments remain vulnerable to self-selection, non-response, and self-report biases that are difficult to verify as removed.These biases can systematically distinguish nonparticipants from participants and distort reported attitudes or behaviors.
- Potential advantages: Simulations may be useful at low accuracy for exploratory steps that tolerate errors, provided results undergo iterative verification.The paper frames science as repeated exploration and validation, with errors less consequential in some stages.
- Fundamental limitations: Human-level performance on isolated LLM tasks may not transfer to usable simulations when model mechanisms differ from human behavior.Such mechanistic differences may obstruct out-of-distribution generalization.
A.2.1. COMPARISON ACROSS EXPERIMENT DATABASES
Most reviewed simulations use straightforward prompts containing participant demographics and study materials, then compare generated responses with individual, group, or average human responses.
- Prompt-based comparisons: Typical prompts include participant demographics, question text, answer choices, and formatting instructions for automated analysis.The resulting outputs are compared against individual participants, demographic or experimental groups, or human-study averages.
- Prompt-based comparisons: Simulated outputs are evaluated against individual human responses, demographic or experimental groups, or aggregate human-study averages.This comparison design supports several levels of correspondence between simulations and human data.
- Prompt-based comparisons: Hewitt et al. used this straightforward approach to simulate preregistered experiments in the largest test described in the section.The passage identifies the study as the largest test but does not provide its full design or results here.
A.2.2. FINE-TUNING AN LLM ON HUMAN DATA
Fine-tuning on human behavioral data can improve simulation accuracy, while interview-based prompting offers another way to represent specific participants. The reviewed evidence spans tasks with widely varying gains.
- Fine-tuning: Fine-tuning embeds additional human-behavior knowledge or makes existing knowledge easier for an LLM to use in simulation.Binz et al. created Centaur by fine-tuning Llama-3.1-70B-Instruct on data from 160 experiments.
- Fine-tuning: Fine-tuning consistently improves simulation accuracy across training-test splits, but the improvement ranges from just under 10% to over 90% of human-data variation.The reported range differs substantially across tasks.
- Fine-tuning: Centaur is presented as a possible foundation model of human cognition, with fine-tuning also improving alignment between model activations and human fMRI scans.The authors propose translating this model into a unified theory of human cognition as future work.
- Participant-specific prompting: Park et al. used two-hour interviews with 1,052 participants to create participant-specific GPT-4o prompts and evaluated each participant twice two weeks apart.The difference between participants’ own responses across the two sessions served as the ceiling for simulation accuracy.
A.2.4. STUDIES THAT TESTED SIMULATIONS BY PREDICTING NOVEL DATA
Novel-data prediction is presented as the most direct test of out-of-distribution simulation, with reviewed studies showing both successful extrapolation and low performance depending on the task.
- Evaluation approach: Novel-data prediction is the paper’s preferred direct test of whether simulations generalize out of distribution.The reviewed studies use data unavailable during model training, collected later, or held out from public access.
- Reviewed findings: GPT-3 reproduced political polarization that later emerged around COVID-19 policies, including vaccine mandates, masks, and lockdowns.The model’s training data were restricted to an October 2019 cutoff.
- Reviewed findings: Low out-of-distribution performance was found for willingness to pay for novel product categories and features, with or without fine-tuning on in-distribution survey data.The test included unusual toothpaste flavors unavailable for purchase.
- Reviewed findings: GPT-4 predictions correlated with average treatment effects at r = 0.91 across 70 preregistered survey experiments, increasing to r = 0.94 for unpublished studies.The correlation measure was adjusted for sampling error, and unpublished studies could not have been in training data.
- Implications and risks: The paper situates novel-data evaluation within broader empirical and normative debates about simulation use, including misuse and misinformation risks.It also notes that critiques commonly focus on empirical issues such as diversity and bias rather than opposing the technology categorically.
C. Limitations of Our Work
The paper’s conclusions are constrained by the scattered, rapidly evolving literature and by difficulty defining the boundary of LLM social simulations relative to related topics.
- The literature is recent and scattered across disciplines, venues, terms, and framings, limiting conventional database-based identification.The authors expect to have identified most preprints and publications but acknowledge concern about missed work.
- Relevant work may have been missed when it was not connected bibliometrically, appeared in different venues, or was published in non-English languages.
- The boundary between LLM social simulations and related areas such as LLM text annotation is difficult to draw clearly.Some related studies compare LLM-generated and human-generated data while addressing adjacent tasks.
- The rapid development of LLMs and their use in social simulation limits confidence in the paper’s conclusions.The authors argue that the work should be revisited as new AI architectures emerge and become popularized.