Source-linked AI summary

This human study did not involve human subjects: Validating LLM simulations as behavioral evidence

Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw

arXiv:2602.15785v1cs.AI

TL;DR

The paper addresses when LLM simulations can support valid inference about human behavior, distinguishing heuristic uses from statistically calibrated methods. It argues that heuristic validation lacks formal guarantees, whereas calibration can preserve validity under explicit assumptions, while both approaches remain limited by how well LLMs approximate relevant populations.

  • Problem

    Researchers lack clear guidance on when LLM simulations can support valid inference about human behavior, especially when heuristic agreement may conceal systematic bias.

  • Method

    The paper contrasts heuristic validation with statistical calibration, examining their assumptions and suitability for exploratory versus confirmatory behavioral research.

  • Results

    Heuristic approaches cannot guarantee the absence of systematic bias, whereas statistical calibration can preserve validity under explicit assumptions and improve precision using auxiliary human data.

  • Takeaways & Limitations

    LLM simulations should be evaluated according to the scientific claim being made, with calibrated methods supporting confirmatory inference and heuristic uses serving narrower exploratory purposes.

  • Takeaways & Limitations

    Calibration methods currently produce only modest precision gains, and their effectiveness depends on assumptions and the quality of LLM approximation to the relevant population.

Abstract

from arXiv · show

A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations support valid inference about human behavior. We contrast two strategies for obtaining valid estimates of causal effects and clarify the assumptions under which each is suitable for exploratory versus confirmatory research. Heuristic approaches seek to establish that simulated and observed human behavior are interchangeable through prompt engineering, model fine-tuning, and other repair strategies designed to reduce LLM-induced inaccuracies. While useful for many exploratory tasks, heuristic approaches lack the formal statistical guarantees typically required for confirmatory research. In contrast, statistical calibration combines auxiliary human data with statistical adjustments to account for discrepancies between observed and simulated responses. Under explicit assumptions, statistical calibration preserves validity and provides more precise estimates of causal effects at lower cost than experiments that rely solely on human participants. Yet the potential of both approaches depends on how well LLMs approximate the relevant populations. We consider what opportunities are overlooked when researchers focus myopically on substituting LLMs for human participants in a study.

1. Introduction

LLM simulations offer fast, inexpensive behavioral data, but researchers must distinguish exploratory uses from claims requiring valid human inference. The paper contrasts heuristic validation, exploratory simulation, and statistical calibration, emphasizing that each supports different scientific claims.

  • LLM simulations: LLMs can simulate responses using demographic characteristics, experimental conditions, or archives of individuals’ prior information.These approaches aim to generate behavioral responses without directly recruiting human participants for every study.
  • Heuristic validation: Heuristic validation treats human and LLM subjects as interchangeable when responses appear to proxy human behavior in related settings.Examples include matching trust-game effects or consumer demand curves, and proposed applications include hard-to-reach groups and longer questionnaires.
  • Exploratory simulation: Simulate-then-validate uses LLMs for exploratory research to prioritize hypotheses, diagnose design problems, and select promising human studies.Simulated results are not treated as definitive evidence.
  • Statistical calibration: Statistical calibration combines human observations with larger LLM-predicted samples and adjusts for discrepancies under explicit assumptions.The approach can improve precision and reduce costs relative to studies relying only on human participants.
  • Validity and scope: The paper argues that heuristic validation cannot guarantee the absence of systematic LLM bias, whereas calibrated methods provide a framework for valid inference under stated conditions.It also examines how focusing only on replacing human participants may overlook broader opportunities for LLMs in behavioral science.

2. Behavioral Research Context

The paper studies behavioral research as inference about a human data-generating process from controlled scenarios, treatments, and subject-level characteristics. Researchers use statistical procedures to estimate parameters and test hypotheses under explicit assumptions.

  • Research setting: A controlled behavioral study probes a human data-generating process with scenarios representing experimental prompts or survey instruments.The resulting dataset contains observations of inputs and outcomes.
  • Observed data: Each observation includes prompt text, treatment indicators, and subject-level observables such as demographic information.The observations are assumed to be independent and identically distributed.
  • Inference: The paper focuses on hypothesis-driven research assessing support for population regularities through estimates, uncertainty, and decision rules.These procedures connect sample results to target parameters such as population means.
  • Research goals: Confirmatory research tests theory-derived hypotheses, whereas discovery-oriented research builds theory through more open-ended investigation.This distinction matters for deciding how simulated evidence should be used.
  • LLM measurements: The framework allows datasets to replace some human measurements with LLM responses, including shared human-LLM observations and larger LLM-only collections.These data structures support both validation and hypothesis generation or testing.

3. Heuristic validation: From partial evidence of fidelity to generalization

Heuristic validation uses partial agreement between LLM and human results to justify generalization from validated settings to related scenarios without statistical guarantees. Evidence of distorted distributions, subgroup disparities, memorization risk, and brittle reasoning limits that generalization.

  • Validate-then-simulate: Validate-then-simulate uses shared human-LLM observations to support LLM-only estimates as proxies for human estimates without statistical adjustment.The larger LLM-only dataset supplies responses for scenarios where human ground truth is unavailable.
  • Heuristic validation: Heuristic validation compares jointly labeled human and LLM data, then applies or presumes the model is suitable for related scenarios lacking human data.The approach relies on implicit partial validation rather than formal guarantees.
  • Generalization claims: Heuristic claims range from near generalization across altered study details to broader assertions about predicting human behavior in novel experiments.The breadth of the generalization target varies substantially across applications.
  • Systematic distortion: LLM and human effect sizes can correlate moderately or strongly while LLM effects are larger and response distributions have lower variance, affecting uncertainty and standardized-effect inference.These distortions matter for measures such as Cohen’s d, probability of superiority, and significance tests.
  • Group differences: Replication fidelity may be lower for historically disadvantaged, older, or internet-underrepresented groups, including a drop from 77% to 42% for socially sensitive main effects.Such asymmetries constrain claims that LLM behavior broadly represents human populations.
  • Memorization: Memorization can make apparent similarity depend on prompt overlap with training data, yet only one of 53 reviewed studies compared predictions on novel scenarios with human data.This creates a direct threat to claims of generalization beyond studied prompts.
  • Misleading generalization: LLMs can perform well on one test of a concept but poorly on another that humans would find easy, revealing brittle or “potemkin” understanding.This pattern appears in compositional, spatial, temporal, and prompt-rephrasing tasks.
  • Generated rationales: LLM explanations of decisions are fragile, can omit systematically influential inputs, and vary in faithfulness across tasks and models.Researchers should therefore be cautious when using generated rationales to corroborate behavioral mechanisms.

4. Necessary conditions for valid inference from LLM surrogates

Valid inference from LLM surrogates requires more than demonstrating small prediction errors: researchers must establish generalization beyond validated scenarios and preserve the assumptions identifying the target parameter. Without these conditions, even accurate-looking predictions can bias downstream estimates.

  • Conditions for valid inference: Heuristic validation therefore rests on implicit closeness arguments, whereas valid substitution requires explicit conditions for scenario generalization and parameter identification.These requirements determine when LLM responses can be treated as interchangeable with human responses without LLM-contributed bias.
  • Generalization across scenarios: The generalization target is the population distribution over study scenarios, not merely new subjects facing the exact validated stimuli.Claims about human behavior may extend to new stimuli of the same type, making assumptions about scenario sampling central to what can be inferred.
  • Generalization across scenarios: No Training Leakage requires sampled validation scenarios to support unbiased extrapolation to the researcher’s larger population of relevant scenarios.If all relevant prompts are observed, there is no generalization problem; alternatively, researchers can define a finite scenario population and randomly sample jointly labeled scenarios.
  • Parameter identification: Preservation of Necessary Assumptions for Parameter Identification is required for downstream inferences to remain valid when human responses are replaced by LLM responses.This condition concerns parameters such as means, treatment effects, and regression coefficients, whose identification depends on population moment conditions.
  • Parameter identification: 90% prediction accuracy can still yield roughly 30% relative regression bias and confidence intervals covering the true value only about 40% of the time.Systematic prediction errors correlated with a covariate can distort the estimated coefficient even when mean measurement error is zero.
  • Parameter identification: No Training Leakage alone is insufficient because small, bounded errors can remain correlated with covariates and non-trivially bias parameter estimation.Valid identification additionally requires the LLM-based moment function to have the same expectation as the human-response moment function for any parameter value.

5. Statistical calibration: Explicit accounting for bias in inference

Statistical calibration combines human observations with larger LLM-predicted samples and adjusts for prediction errors, enabling valid parameter estimates under explicit assumptions without requiring accurate or unbiased LLM predictions. Empirical applications find only modest precision gains, while validity and efficiency remain bounded by data, predictability, and calibration-method limitations.

  • Calibration framework: Statistical calibration combines typically smaller human samples with larger LLM-predicted samples to estimate population means, regression coefficients, and other parameters.The LLM is treated as a cost-effective but imperfect information source about human behavior.
  • Calibration framework: Shared human and LLM outcomes estimate prediction errors and a rectifier term that calibrates either a human-only or LLM-only base estimator.The resulting estimators target the human parameter rather than treating LLM predictions as automatically interchangeable with human data.
  • Validity guarantees: Under the stated construction and assumptions, calibrated estimators are consistent and asymptotically unbiased without assuming that the LLM predictor is accurate or unbiased.Different methods vary in their base estimator, human-label sampling assumptions, and use of shared data for adjustment.
  • Calibration methods: PPI improves precision by combining gold-standard human observations with surrogate predictions, provided predicted-input and human-input features share a distribution and the predictor is independent of both samples.Appropriate tuning makes the combined estimate at least as precise as human measurements alone.
  • Validity guarantees: Under standard regularity conditions, PPI and cross-fitting yield consistent, asymptotically normal estimators with asymptotically valid confidence intervals.Cross-fitting learns bias on held-out shared-data folds before debiasing predictions across the larger LLM dataset.
  • Empirical performance: Empirical gains remain modest: augmenting 10,000 human decisions with 100,000 LLM predictions increased effective sample size by about 13%, while another study found gains up to 14%.These results illustrate that large LLM-augmented datasets do not automatically produce proportionally large precision improvements.

6. Language models and the role of simulation in behavioral science

LLM simulations can make exploratory behavioral research more efficient by screening designs, generating hypotheses, and expanding the search over outcomes and features. Their usefulness is constrained by model-based bias, especially when effects are small or LLMs overestimate significance.

  • Simulate-then-validate: Simulate-then-validate uses AI surrogates to screen many study designs, prioritize promising hypotheses, and reserve human testing for selected candidates.This approach is intended for exploratory discovery rather than treating simulated results as definitive evidence.
  • Simulate-then-validate: LLM-based screening may improve exploratory search when only a small fraction of possible designs produce effects consistent with a broad theory.The paper illustrates this with mindfulness interventions and alternative operationalizations of study designs.
  • Simulate-then-validate: Low-powered human pilots can produce noisy estimates and false positives, whereas LLMs can cheaply generate many measurements on the same design, though with model-based bias.The efficiency gain comes from scale, not guaranteed accuracy.
  • Simulate-then-validate: LLMs can overestimate effect sizes and produce significant results for human designs that were not statistically significant, limiting their reliability for ranking candidate designs.Across 12 causal paths in four experiments, estimates correctly predicted the effect sign 10 times; failures involved two smaller effects.
  • Causal discovery: Beyond measurement, LLMs may support cheaper search over outcomes, inputs, and interpretable features in text-based behavioral studies.Proposed methods include theme discovery, prediction-based completeness measures, and feature-level causal analysis.
  • Mechanism discovery: LLM simulations can also help generate theories of human mechanisms by probing model representations and suggesting follow-up studies for human participants.These interpretability-based applications remain nascent in human behavioral science.

7. Conclusions

The paper distinguishes heuristic uses of LLM surrogates from statistically calibrated inference. Heuristics can support exploratory work, but confirmatory claims require explicit assumptions and calibration guarantees, while broader adoption may narrow the questions researchers pursue.

  • Heuristic versus calibrated inference: Heuristic validation does not provide statistical guarantees that LLM behavior generalizes from validated tasks to the tasks of interest.It relies on implicit or partial validation instead of formal guarantees.
  • Heuristic versus calibrated inference: Valid plug-in use requires explicit conditions, including no leakage between test and training data and preservation of critical moment conditions for parameter identification.The paper contrasts these requirements with reliance on face plausibility alone.
  • Exploratory use: Using AI surrogates to pre-test exploratory designs carries less risk than using them for confirmatory inference, but efficiency gains and ranking reliability are not guaranteed.The paper calls for domain-specific corroboration of how reliably surrogates rank experimental scenarios.
  • Scope and research selection: Heavy adoption of AI surrogates could create selection effects by favoring research domains that are easier for LLMs to simulate.Text-based reactions may be easier to model than paralinguistic signals or uncontrolled field contingencies.
  • Scope and research selection: This domain-selection risk could produce an illusion of exploratory breadth while weakening attention to socially or policy-relevant questions that are harder to simulate.The paper compares this concern with earlier narrowing caused by readily available survey-experiment infrastructure.

8. Appendix

The appendix formalizes a research context as a joint distribution over which prompt strings enter the study and which appear in model training. It then states assumptions about this sampling process.

  • Research context: The formal setup represents prompts as strings S from a space of possible scenarios, with structured covariates Z containing treatment and subject-level observables.The same string space is used to describe study prompts and possible training-data strings.
  • Research context: String-level indicators record whether each prompt appears in the study or in the model’s training data.C marks study inclusion, while T marks training-data inclusion.
  • Research context: A research context Q is defined as a joint distribution over study-prompt inclusion and training-data inclusion indicators.Q is used for string-level statements, while inference statements take expectations over the distribution of analysis inputs.
  • Assumptions: The appendix introduces assumptions on how strings are sampled within a research context before deriving generalization results.Ludwig et al. are identified as the source of two sampling assumptions.

2. The marginal number of study prompts is unaffected by conditioning on T: EQ

Valid LLM substitution requires more than close behavioral resemblance: it depends on explicit conditions governing training leakage, preserved analysis moments, and treatment-related bias. The section also contrasts substitution with calibration methods that can correct bias and improve precision under stated assumptions.

  • Model setup: The LLM is modeled as a deterministic black-box generator that maps a text string to its most likely response.This assumption defines the simulation mechanism used in the formal framework.
  • 8.1.1 Requirements for LLM substitution: Valid substitution requires no training leakage between the model and study prompts, so validation performance generalizes to the target scenario population.The relevant leakage condition is defined at the prompt-string level.
  • 8.1.1 Requirements for LLM substitution: Substitution also requires preserving the population moments used by the downstream estimator, not merely achieving small pointwise prediction error.Systematic errors related to analysis covariates can bias parameter estimates even when predictions are individually accurate.
  • 8.2 Alternative causal inference treatment (Perkowski, 2025): The formal framework treats human outcomes as generated from potential-outcome models and the LLM as producing noisy approximations to those outcomes.The model includes individual characteristics, idiosyncratic noise, individual-specific simulation bias, treatment-specific prompt bias, and residual response noise.
  • 8.2 Alternative causal inference treatment (Perkowski, 2025): Unbiased ATE estimation is possible when individual-specific simulation bias is treatment-invariant, because that bias cancels when treated and control simulations use the same individuals.This requires that the LLM not systematically favor one treatment condition over the other.
  • 8.2 Alternative causal inference treatment (Perkowski, 2025): With nonadditive interactions between individual-level and treatment-specific biases, treatment invariance alone is insufficient and an additional population-average interaction restriction is required.The interaction must vanish for identification under the extended model.
  • 8.3 Example of correlated bias in downstream estimation: Even mean-zero overall measurement error can shift a regression slope when errors align with treatment or another analysis covariate.The example represents the LLM outcome as a proxy measurement and shows that treatment-aligned errors distort downstream estimation.
  • 8.4 Estimating PPI parameter λ for a population mean: Prediction-powered inference centers the corrected estimate at the human-targeted parameter and can achieve precision no worse, and often better, than a human-only estimator.For the population mean, the result holds under adequate sample sizes and applies to any prediction algorithm f and λ.
Loading 2602.15785v1…