Source-linked AI summary

Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?

John J. Horton, Apostolos Filippas, Benjamin S. Manning

arXiv:2301.07543v2econ.GN

TL;DR

The paper asks whether LLMs can provide useful computational models of human economic behavior despite conceptual and causal-inference concerns. It uses dozens of LLMs as simulated subjects in replications and extensions of established experiments, finding core qualitative similarities to human results and deviations that may inform future research. The authors therefore characterize Homo silicus as theory in flexibly executable form, while retaining the need for empirical confirmation.

  • Problem

    The paper examines whether AI simulations can be practically valuable for generating insights about actual humans, given concerns about their conceptual foundations and causal interpretation.

  • Method

    The paper uses dozens of LLMs as simulated experimental subjects in scenarios derived from five established economic experiments.

  • Results

    The simulations recover core qualitative patterns from analogous human experiments and surface deviations that may be informative for future research.

  • Takeaways & Limitations

    Homo silicus is best understood as theory in flexibly executable form, offering a framework for exploring economic behavior through simulation.

  • Takeaways & Limitations

    AI simulations do not by themselves establish conclusions about human behavior because causal analysis is challenging and results still require empirical confirmation.

Abstract

from arXiv · show

We argue that newly-developed large language models (LLMs), because of how they are trained and designed, are implicit computational models of humans -- a Homo silicus. LLMs can be used like economists use Homo economicus: they can be given endowments, information, preferences, and so on, and then their behavior can be explored in scenarios via simulation. Experiments using this approach, derived from Charness and Rabin (2002), Kahneman et al. (1986), Samuelson and Zeckhauser (1988), Oprea (2024b), and Horton (2025), show qualitatively similar results to the original, and when they differ, it is often generative for future research. We discuss potential applications, conceptual issues, and why this approach can inform the study of humans.

1 Introduction

The paper frames LLMs as implicit computational models of humans, or “Homo silicus,” that can simulate economic scenarios like economists’ Homo economicus. Across five experiments, these simulations recover qualitative human patterns while offering flexible, inexpensive, and potentially generative research tools.

  • Conceptual foundation: Homo silicus can receive endowments, information, and preferences, enter natural-language scenarios, and have its behavior explored through computational simulation.Unlike fixed-parameter mathematical models, prompts can be adjusted to incorporate conventional economic theory and other behavioral assumptions.
  • Limitations: Homo silicus remains a flawed model whose fidelity depends on prompts and whose training data may produce brittle or memorized responses.The paper discusses concerns involving imperfect training sources, opacity, generalizability, and the possibility that simulations appear robust while regurgitating problematic results.
  • Conceptual foundation: LLMs are computational models of humans with social information learned from economic activity, decision-making, social preferences, and published research.Their training includes textual discussions of economic matters and codified facts, theory, and empirical results.
  • Experiments: Across five simulated experiments, the paper compares AI behavior with human-subject findings from established economic studies.The experiments are selected for simple implementation and clear qualitative results suitable for comparison.
  • Experiments: In the pricing, allocation, framing, and labor scenarios, AI agents reproduce several qualitative patterns, including political effects on fairness, preference-sensitive allocations, status quo bias, and minimum-wage effects.Higher gouging is viewed more negatively; endowed political views matter; agents can select equitable, efficient, or self-interested outcomes; status quo effects vary across models; and minimum wages shift hiring toward more experienced applicants and higher wages.
  • Implications: The simulations are inexpensive to run and can be replicated, shared, and updated, while deviations from human results may generate future research questions.The paper emphasizes practical usefulness but maintains that AI-experiment results require empirical confirmation about actual humans.

2 Experiments

Across five experiments, LLM-based agents recapitulate several human behavioral patterns while allowing researchers to vary prompts, personas, models, and scenarios. The experiments also show that calibrated personas can improve prediction of human responses and that agents reproduce status quo bias, prospect-theory-like patterns, and labor-market responses.

  • Fairness and price gouging: Larger price increases are judged less fair, while right-leaning agents are more accepting of price gouging across many experimental permutations.Left-leaning agents consistently judge the original price increase as “Unfair” or “Very Unfair.”
  • Social preferences: Persona instructions produce consistent choices: efficiency-oriented agents maximize combined payoffs, inequity-averse agents reduce payoff differences, and self-interested agents maximize their own payoffs.Baseline behavior differs across models, but assigned personas generally guide responses in the predicted direction.
  • Status quo bias: AI agents exhibit strong status quo bias, with allocations farther from the status quo less likely to be chosen, broadly matching human subjects.The pattern remains relevant despite substantial post-training fine-tuning.
  • Prospect-theory-like behavior: Across five models and thousands of simulations, lottery and mirror conditions generally move together, while lower-assigned mathematical ability elicits a more pronounced fourfold pattern.More capable models usually select expected value unless prompted to be “very bad at math.”
  • Prospect-theory-like behavior: AI agents’ responses vary systematically with mathematical-ability personas in a way that aligns with prospect theory, although the authors note possible exposure to lottery literature.The authors interpret the relationship as evidence that models may encode human trade-offs learned from training examples.
  • Labor-labor substitution: A minimum wage increased the hired worker’s wage by about $1.3/hour when no reference wage was provided.The experiment also examines effects on the hired worker’s experience.

3 Why and when we can learn from Homo silicus

LLMs can simulate human behavior because their training data and training processes capture explicit social-science knowledge and implicit patterns of human behavior. Homo silicus extends conventional economic modeling through flexible, theory-informed simulations that can generate predictions and guide research.

  • Why LLMs can simulate human behavior: LLMs draw social information from social-science literature and non-scholarly human-generated text depicting communication, decisions, interactions, and perceptions.The two paths provide explicit theories and findings alongside behavioral examples across diverse contexts.
  • Why LLMs can simulate human behavior: Their corpora, architectures, and post-training processes enable generalizable behavioral patterns that support predictions across settings.The paper presents this as the basis for LLMs’ broad predictive capacity.
  • Applications: Homo silicus enables many rich, low-cost simulations that can produce generative insights and help target expensive data collection where information has highest marginal value.The authors frame this as a faster form of learning by doing that can support testable predictions.
  • Flexibility: Unlike conventional economic theory, Homo silicus can generate predictions in virtually any setting described in natural language and flexibly execute theory-based instructions.The paper characterizes economic theory as supplying instructions and LLMs as supplying flexible execution capacity.
  • Evidence from simulations: Theory-grounded prompts improve fit to human behavior, whereas atheoretical or scientifically meaningless traits generally do not improve predictive accuracy.This pattern appears in Charness and Rabin’s games and supports a connection between interpretable theory and predictive performance.
  • Using Homo silicus to navigate uncharted territory: AI simulations can corroborate sound social science or reject poor propositions, but they can also produce false positives and false negatives when navigating uncharted settings.The framework uses a confusion-matrix view and recommends calibration against available human ground truth before generalizing to nearby unlabeled settings.

4 Conceptual issues

The paper addresses concerns about whether LLMs are suitable for social science, including imperfect training data, memorization, performativity, incoherent world models, and causal identification. It argues that these concerns constrain interpretation but do not eliminate the value of theory-informed, flexible simulations as research tools.

  • Training-data concerns: LLM training data are difficult to curate and may disproportionately represent people who create public writing rather than the broader human population.The paper also notes that social information sources are imperfect and shaped by distinct, often opaque selection mechanisms.
  • Training-data concerns: Because LLMs are trained on statements rather than actions, critics question whether they capture revealed preferences, although internet text also records reasoning, intentions, decisions, purchases, and ratings.The authors argue that the stated-versus-revealed-preference critique is only superficially persuasive because training data contain implicit and explicit social information.
  • Memorization and performativity: LLMs need not merely reproduce memorized examples: their flexibility and usefulness across novel games suggest predictions more like a student applying concepts than one recalling isolated answers.The paper contrasts brittle memorization with applying a concept such as Nash equilibrium to games with different payoffs or action spaces.
  • Memorization and performativity: Performativity can be beneficial when an AI agent follows a correct theory, because internalizing and applying predictive social-science relationships can improve behavioral modeling.The authors therefore treat memorization and performativity as potentially useful rather than intrinsically problematic.
  • World-model coherence: Incoherent world models can undermine credibility, as illustrated by poorer predictions when street detours are introduced, while broader and more heterogeneous training is associated with stronger performance.The paper presents this as an unresolved concern whose importance for the most capable models remains unclear.
  • Flexibility and policy change: Homo silicus offers flexible natural-language programming, maps preferences and beliefs to actions, and can expose reasoning, expanding the scope of simulations beyond conventional agent-based models.This flexibility may also reduce susceptibility to the Lucas critique because agents can reason about environmental changes rather than follow fixed rules.
  • Causal inference: Causal analysis with LLM experiments is difficult because an experiment uses a single model and because model responses can depend on contextual assumptions such as reference wages.The authors recommend additional simulations to probe sparsity, identify threats to generalizability, and run checks at negligible marginal cost.

5 Conclusion

The paper presents LLMs as Homo silicus: flexible computational models that can simulate economic behavior and recover qualitative human-experiment patterns. It emphasizes practical research uses, open reproducibility, and the need to map when simulations faithfully track humans.

  • Dozens of LLMs used as simulated experimental subjects recover core qualitative patterns from analogous human-subject experiments.
  • When AI simulations deviate from human results, those deviations can be informative for future research.
  • Homo silicus is theory in flexibly executable form, linking conceptual economic models to manipulable simulations.
  • AI simulations can rapidly explore ideas, stress-test experimental designs, and calibrate power calculations before costly human data collection.
  • Open-source software standardizes AI-experiment design, permits model substitution, and supports updating results as models evolve.
  • The central limitation is an incomplete map of settings and assumptions in which AI simulations are high-fidelity representations of Homo sapiens.
  • Automated scientific loops could propose hypotheses, instantiate agents, run and score simulations, iterate, and accelerate research productivity.

A Figures from the original draft

The appendix presents figures and a table from the original draft, covering simulated choices, fairness opinions, safety budgets, and minimum-wage effects. These materials organize AI-agent outcomes by framing, scenario, and worker outcome.

  • Figure A1 displays Charness and Rabin (2002) choices by model type and endowed personality.
  • Figure A2 presents the snow-shovel price-gouging question with endowed political views and reports opinions by scenario.
  • Figure A3 shows the distribution of preferred car-safety budgets by status quo framing.
  • Table A1 reports minimum-wage effects on observed wages and hired-worker attributes, including wage and experience outcomes.

B A brief primer on large language models

The primer explains how LLMs generate responses from learned text distributions and how post-training incorporates human feedback and instruction following. It also describes temperature as a control over response determinism and stochasticity.

  • LLMs generate responses by repeatedly predicting the next word from the prompt and previously predicted words.
  • Pre-training teaches an LLM to predict text sequences and produces a conditional probability distribution shaped by architecture and hyperparameters.
  • Post-training can include supervised fine-tuning, instruction-tuning, and reinforcement learning with human feedback.
  • In RLHF, human evaluations train a reward model, and reinforcement learning updates the LLM toward responses receiving favorable predicted feedback.
  • Further fine-tuning can make responses more consistent with a particular persona or writing style.
  • At temperature zero, the model deterministically outputs its highest-probability response; higher temperatures make outputs more stochastic and uniform.
  • LLMs operate on tokens, which may be sentences, words, word parts, or characters.

C Additional Figures and Tables

The appendix documents the simulation materials through linked Jupyter notebooks and additional regression and comparison tables. These resources support replication and examination of fairness effects across political beliefs and price changes.

  • Figure A4 shows a linked Jupyter-notebook screenshot for running AI simulations.
  • The notebooks share instructions, comments, and links for sharing or downloading experiments to support replication and exploration.
  • Table A2 estimates how political beliefs and price-hike levels affect AI agents’ fairness assessments across four experiment permutations.
  • The regression specification uses the snow-shovel price change and political beliefs as independent variables, with fairness assessed on a 1–4 scale.
Loading 2301.07543v2…