Source-linked AI summary
Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates
Daehwan Ahn, Chengfeng Mao, Dokyun Lee
TL;DR
The paper evaluates whether LLMs can act as respondent-specific human surrogates rather than item-mean predictors. Across residual-prediction, persona-swap, distributional, and adaptation analyses, respondent-specific prediction remains weak and model outputs distort human response distributions.
Problem
The paper addresses whether LLM predictions capture respondent-specific deviations from item-level human averages, rather than only stable item means.
Method
The paper evaluates per-respondent prediction, de-meaned residuals, persona reassignments, Wasserstein-2 distributional decompositions, and fine-tuned model variants.
Results
3.05% de-meaned R2 under human-mean centering remains far below the 53.6% reliability-based reference, while fine-tuning improves item aggregation but can reduce individual-level signal.
Takeaways & Limitations
LLM surrogates can approximate item-level averages, but weak residual prediction and distributional shape distortion limit their use for replacing individual humans.
Takeaways & Limitations
The 53.6% benchmark applies specifically to de-meaned analyses; raw response prediction has a higher ceiling because stable item means add reliable variance.
Abstract
from arXiv · showhide
LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.
Supporting Information for: Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human
The supporting information identifies Daehwan Ahn, Chengfeng Mao, and Dokyun “DK” Lee as authors.
- Daehwan Ahn, Chengfeng Mao, and Dokyun “DK” Lee are listed as authors.
SI Text
The paper separates human response structure from LLM prediction-error structure and defines notation for respondents, items, reliability, and variance components.
- The notation distinguishes respondent, item, and occasion indices from reliability and prediction-error components.p indexes respondents, i indexes items, and t indexes test or retest occasions.
- The notation includes absolute prediction error and the respondent-specific de-meaned R2 used for individual prediction.The de-meaned metric is computed from matched respondent–item responses and predictions after item means are removed.
- Response-model terms describe human response structure, whereas prediction-error terms describe structure in LLM prediction error.Primed symbols mark reliability-based partitions of prediction-error residuals.
POMP Scaling
The primary analyses rescale eligible ordered, bounded responses to a common 0–100 POMP scale, while excluding unsuitable response types.
- POMP rescales a response to 100 × (x − a)/(b − a) using its valid lower and upper endpoints.The transformation reduces mechanical influence from scale width in pooled analyses.
- POMP standardizes scale width without making different constructs interchangeable.
- Primary POMP analyses include direct responses with meaningful ordered, bounded endpoints supplied by source metadata.The analysis excludes nominal categories, background variables, computed scores, objective-correctness measures, open estimates without defined endpoints, and unresolved invalid codes.
De-Meaning Procedure for Residual Analysis
The residual analysis removes each item’s human mean from both human responses and LLM predictions, then evaluates whether remaining deviations track respondent-specific differences.
- Human-mean centering subtracts each item’s human mean from both the human response and LLM prediction.The procedure targets individual-level deviations after removing item-level structure.
- 3.05% de-meaned R2 was obtained in the primary Megastudy analysis under human-mean centering.This indicates weak respondent-specific prediction after human item means are removed.
- 3.87% de-meaned R2 resulted from own-mean centering for the same primary prompt and model specification.Own-mean centering removes calibration differences, including systematic offsets between LLM and human item means.
- 4.44% was the strongest non-fine-tuned Megastudy result under human-mean centering, below the 53.6% reliability-based reference.Weak respondent-specific prediction remained under either centering choice.
Persona Swap Procedure
The procedure estimates a de-meaned individual-prediction ceiling from human test–retest reliability and evaluates predictions against respondent-specific deviations from item means.
- Persona swap control: 10,000 within-study persona permutations reassigned descriptions across individuals while keeping survey items, prompts, and recorded responses fixed.Each permutation’s de-meaned R2 was compared with the true-assignment result.
- Reliability benchmark: 2,057 Twin-2K-500 participants contributed matched Wave 1–3 and Wave 4 responses across 126 survey items.The resulting within-item test–retest reliability was rtt = 0.536, with bootstrap 95% CI [0.506, 0.565].
- De-meaning: After subtracting each item’s human mean, the analysis retains respondent-specific variation while removing the item main effect.The item-centered response is defined as dpit = zpit − z̄·i.
- Reliability benchmark: R2 ceiling = rtt = 0.536 serves as the reliability-based upper-bound reference for de-meaned individual predictions.The benchmark represents the stable share of item-centered response variance under the reliability model.
- Scope of ceiling: The 53.6% ceiling applies specifically to de-meaned predictions, whereas raw-response prediction would have a higher ceiling because item means add reliable variance.The de-meaned benchmark is appropriate for claims about respondent-specific deviations from item-level averages.
Variance Decomposition Procedure
The variance procedure decomposes prediction error with REML mixed-effects models, then uses external reliability to separate stable person-by-item structure from transient noise.
- Model decomposition: REML mixed-effects models decompose prediction error in the primary POMP-scaled Megastudy analysis containing 133 items.Person is the grouping factor and item is a random variance component.
- Model decomposition: The main G-theory decomposition operates on single-occasion POMP-scaled prediction error rather than repeated observed responses.This distinction determines which variance components are being estimated.
- Residual partition: External Survey reliability partitions the residual because δpi combines stable person-by-item structure with one-time transient response noise.The Survey and Megastudy use the same participant pool but different item sets.
- Variance estimates: 86.41% of total prediction-error variance is residual variance in the 133-item POMP-scaled REML decomposition.This is reported as σ²δ in Table S2.
- Variance estimates: 42.4% is the estimated person-by-item variance share at rtt = 0.536, exceeding the 4.93% person main effect by at least 8.4×.Across the full 95% CI for rtt, the person-by-item estimate ranges from 41.3% to 46.7%.
- Assumption: Applying rtt = 0.536 from survey items to Megastudy behavioral items assumes comparable de-meaned-response reliability across item types.The assumption may be affected by differences in context dependence, novelty, or related properties.
Cross-Family and Adaptation Protocols
The comparison design spans multiple model families and adaptation protocols, with identical prompts for five cross-family model rows but distinct protocols for other models.
- Cross-family comparisons: Five cross-family rows used identical prompts for GPT-5, DeepSeek R1, Gemini 2.5 Flash, Gemini 3 Pro, and Llama 3.1 70B.The comparison also included GPT-4.1, Centaur, and other adaptation rows.
- Protocol differences: GPT-4.1 used an enrichment-gradient prompt template, while Centaur used its original protocol rather than the shared cross-family prompt.Megastudy rows required at least five included items.
- Adaptation protocols: Socrates supervised-fine-tuning and direct-preference-optimization variants used the SocSci210 corpus and study-level seen/held-out metadata.Both were full-parameter adaptations of Qwen2.5-14B.
- Adaptation protocols: Fine-tuned GPT-4.1 used a stratified person-level split by sex and age group, with pooled de-meaned R2 reported separately.The associated result is reported in the main table and Table S3.
Supplementary Results
Supplementary analyses document the analytic rows, item-mean tracking, and distributional collapse across four datasets, showing that LLM outputs match means while using narrower and distorted response distributions.
- Supplementary analyses: The supplementary structure covers item-mean tracking, distributional collapse, shape distortion, variance decomposition, and intervention checks.It also documents model sources and reported metrics.
- Analytic rows: Supplementary tables map source items or outcomes to primary and comparison analytic rows using pooled de-meaned R2 unless otherwise stated.The main text reports primary pooled de-meaned R2 values.
- Item-mean tracking: The primary Megastudy full-persona prompt achieved Pearson r = 0.777 for item-mean agreement on 133 POMP-scaled items.Spearman ρ was 0.721 after POMP scaling on the same primary row.
- Distributional results: Across all four datasets, LLM outputs showed variance compression, repertoire collapse, and elevated Wasserstein-1 distances relative to human response distributions.These metrics consistently indicate distributional collapse despite approximate item-level means.
- Distributional results: Median SD ratios were 0.50 in SocSci210, 0.65 in the Megastudy, 0.57 in the Survey, and 0.99 in ANES.The SD ratio compares LLM and human per-item response spread.
- Distributional results: Effective-category ratios quantify repertoire collapse by estimating how many response options each distribution effectively uses.LLM responses cluster in a restricted response-scale region while human responses span the full range.
Shape Distortion Beyond Variance Compression
The paper tests whether LLM distributional collapse is merely uniform variance compression. Across complementary diagnostics, LLM outputs also show item-level shape distortion beyond location and spread differences.
- Compression-Null Simulation: Compression-only nulls preserved item-level skewness more strongly than actual LLM outputs in the Megastudy, SocSci210, and Survey.Null-versus-LLM skewness correlations were 0.88 versus 0.49 in the Megastudy, 0.58 versus 0.31 in SocSci210, and 0.69 versus 0.48 in the Survey.
- Conclusion: Together, the diagnostics show that LLM distributional differences cannot be explained by variance compression alone.
- Skewness and Kurtosis Invariance: Raw moment correlations remained below compression-null benchmarks, indicating altered distributional shape beyond location-scale shifts.
- Wasserstein-2 Decomposition: Wasserstein-2 decomposition partitions squared distance into additive location, size, and shape components after comparing each item’s human and model distributions.The shape term depends on quantile-function similarity after matching mean and spread; linear mean-variance calibration cannot remove it.
Error Profile vs Item-Mean-Only Predictor
The paper compares LLM individual predictions with an item-mean-only benchmark and examines whether apparent signal reflects item-level averages rather than respondent-specific information. Persona reassignment and fine-tuning analyses further test this interpretation.
- Error Profile: The LLM’s itemwise individual-prediction RMSE was strongly associated with the error profile of an item-mean-only predictor, r = 0.726 across 133 POMP-scaled items.The item-mean-only predictor gives every respondent the human sample mean for that item, so its error equals human dispersion around the item mean.
- Baseline Comparison: In the primary POMP-scaled Megastudy analysis, the LLM underperformed a leave-one-out human item-mean baseline on per-person prediction correlations.The baseline uses other respondents’ mean response for each item and contains no respondent-specific information.
- Persona Swap Control: Correct persona–respondent matching yielded de-meaned R2 = 3.05% with human-mean centering and 3.87% with own-mean centering.The result comes from the Megastudy permutation-control analysis.
- Persona Swap Control: The largest human-mean-centered permutation value was 0.028%, separating the matched result from the null maximum by 107.7-fold with empirical p = 0.0001.The null used 10,000 within-study persona reassignments while preserving prompt length, structure, and demographic-category proportions.
- Fine-Tuning: Socrates fine-tuning produced R2 = 7.57% on seen studies versus 0.73% on held-out studies, with gains concentrated in seen-study metadata.The split was at the study level, and the labels do not establish model pretraining exposure.
Model-Family and Adaptation Coverage
The paper combines literature classification with model and adaptation analyses to assess whether stronger models, fine-tuning, or broader validation practices close the individual-prediction gap. The evidence indicates that these approaches do not eliminate it.
- Adaptation Results: Fine-tuned variants did not eliminate the individual-prediction gap.Fine-tuned GPT-4.1 had centered r = 0.100.
- Adaptation Results: In the Survey, fine-tuned GPT-4.1-mini improved item-level aggregation from r = 0.859 to 0.992 while centered r fell from 0.038 to −0.042.
- Adaptation Results: On the same 126-item Survey set, POMP pooled de-meaned R2 fell from 6.16% to 0.31% after fine-tuning.The fine-tuned cohort included 1,557 persons and 149,498 pairs.
- Literature Classification: The literature review classified surrogate papers by paper type, stance, and strongest reported fidelity level rather than terminology alone.Individual-level coding required a per-respondent validation metric computed across multiple responses from that respondent.
- Scope and Limitations: The literature pool is restricted to papers proposing, deploying, or empirically testing LLMs as human surrogates and is not a field-wide distribution of views.
SI Figures
The supplementary figures document three related patterns: narrower LLM response spread, fewer effective response categories, and error profiles resembling an item-mean-only predictor. Representative distributions also illustrate shape distortion beyond variance compression.
- Figure S1: Figure S1 plots human response SD against LLM prediction SD; points below the diagonal indicate compression and red crosses indicate zero-variance outputs.Median SD ratios were 0.50 in SocSci210, 0.65 in the Megastudy, 0.57 in the Survey, and 0.99 in ANES.
- Figure S2: Figure S2 compares human and LLM response frequencies and reports fewer effective LLM categories for 100.0% of Megastudy, 98.9% of SocSci210, 93.9% of Survey, and 75.0% of ANES items or outcomes.The aggregate frequency panels show LLMs avoiding scale extremes.
- Figure S3: Figure S3 compares LLM itemwise individual-prediction RMSE with an item-mean-only predictor’s error across 133 POMP-scaled items.The reported association is r = 0.726.
- Figure S4: Figure S4 shows representative human and LLM item-response distributions, while Table S6b provides quantitative checks for shape distortion.The checks are compression-null skewness, raw moment correlations, and Wasserstein-2 decomposition.
SI Tables
The SI tables document dataset scope, pooled de-meaned R2 definitions, variance-decomposition sensitivity, model and prompt coverage, and distributional-compression diagnostics. Together, they clarify how the paper evaluates respondent-specific prediction and robustness across analytic choices.
- Literature classification: 63 surrogate papers are classified as 44 empirical advocates, 9 non-empirical advocates, and 10 critiques.The classification also codes whether studies assess means, distributions, individual prediction, or paradigm-level advocacy.
- Dataset and analytic coverage: The SI inventories primary and comparison rows for Megastudy, SocSci210, Survey, and ANES, including outcome-level inclusion rules and special ANES-code handling.SocSci210 rows require at least 20 matched human–LLM pairs per outcome after model-output filters; ANES code 99 is excluded before POMP scaling.
- Metric and comparison design: Pooled de-meaned R2 is the squared Pearson correlation between human and model deviations after subtracting each item’s human mean.Comparison rows alter item sets, scale handling, inclusion rules, or ANES ideology handling.
- Variance decomposition: 4.93% of prediction-error variance is attributed to stable person effects, while the residual accounts for 86.41% in the primary Megastudy decomposition.The residual is further partitioned into stable person-by-item interaction and transient error using test–retest reliability.
- Variance sensitivity: 8.4× is the minimum person-by-item-to-person-main-effect ratio at the lower Survey reliability bound, versus 4.6× across the broader sensitivity range.The primary reliability reference is rtt = 0.536, with a 95% confidence interval of [0.506, 0.565].
- Distributional and model diagnostics: Distributional diagnostics report response spread, effective-category usage, Wasserstein distance, and shape distortion, while model checks summarize centered respondent correlations and item-mean tracking.SD ratio is defined as SDLLM/SDHuman, and values below 1 indicate compressed response distributions.