Source-linked AI summary

When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability

Steven Wang, Kyle Hunt, Shaojie Tang, Kenneth Joseph

arXiv:2609.07987v1cs.AIstat.AP

TL;DR

LLM-based digital twins may reduce repeated human measurement, but existing evaluations provide limited evidence that they preserve valid inference. The paper introduces statistical substitutability and evaluates it through four dimensions using prediction-powered and mixed-subject inference. Across its evaluations, twins can reproduce average human effects while providing weak or unstable respondent-level information for reducing human-data requirements.

  • Problem

    Existing digital-twin evaluations emphasize held-out prediction and behavioral replication without establishing whether human measurement can be reduced while preserving valid inference for a target estimand.

  • Method

    The paper evaluates statistical substitutability through aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations using prediction-powered inference.

  • Results

    Across the evaluated settings, digital twins can reproduce average human effects while providing little or unstable information about individual deviations needed for reliable human-data savings.

  • Takeaways & Limitations

    Confirmatory use should depend on whether twin predictions reduce uncertainty about human quantities, not merely whether they reproduce human means, distributions, or effects.

  • Takeaways & Limitations

    The findings are limited by one retrospective Twin-2K baseline, matched two-condition estimands, one political domain in the Moore-Berg extension, and unmeasured profile-collection costs.

Abstract

from arXiv · show

LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.

1 Introduction

The paper asks whether LLM-based digital twins can reduce human measurement while preserving valid inference, distinguishing inferential usefulness from behavioral resemblance. It introduces statistical substitutability and evaluates aggregate fidelity, respondent-level signal, finite-sample recovery, and transport stability.

  • Motivation and contribution: Existing synthetic respondents can recover human means and treatment-effect directions but may compress variation, misrepresent subgroups, and overstate effects.
  • Motivation and contribution: Statistical substitutability measures whether digital-twin predictions can reduce human measurement for a prespecified estimand while preserving valid inference.It differs from response accuracy, distributional similarity, and aggregate effect replication.
  • Empirical design: Twin-2K evaluates 12 reconstructed behavioral studies by comparing aggregate replication with respondent-level signal and prediction-assisted performance.The design asks whether strong held-out prediction and behavioral replication provide enough information to reduce future human measurement.
  • Empirical design: The Moore-Berg extension provides a more favorable test because party identity is directly related to the estimands and permits comparisons across models and respondent representations.
  • Motivation and contribution: The framework combines aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations.It uses prediction-powered inference and mixed-subject methods rather than treating twin responses as human observations.
  • Implication: The study shifts evaluation from behavioral resemblance toward inferential usefulness when judging AI-generated participants as scientific evidence.The paper frames this shift as relevant to AI-enabled behavioral research and human-AI workflows.

2 Related Work

Prior work shows that synthetic respondents can match aggregate patterns while failing to preserve person-level variation, and that calibration can convert paired signal into valid efficiency gains. What remains unresolved is whether current digital twins reliably provide enough estimand-specific, stable signal to reduce human measurement.

  • Behavioral fidelity and paired signal: Synthetic participants can reproduce some experimental findings and treatment-effect directions, but interactions are harder to reproduce and effect magnitudes may be distorted.
  • Behavioral fidelity and paired signal: Aggregate or distributional alignment does not guarantee paired respondent-level prediction, because reassigned responses can preserve population effects while losing individual correspondence.
  • Digital twins and the open gap: Richer respondent histories and behavior-specific training can improve prediction in some settings, making digital twins a favorable test of person-specific signal.
  • Digital twins and the open gap: Digital-twin benchmarks mainly measure held-out response prediction or behavioral-pattern reproduction rather than how much future human measurement can be avoided for a downstream estimand.
  • Statistical calibration: Prediction-powered inference uses labeled validation outcomes to correct prediction error and improve precision when predictions carry useful signal.Related methods establish valid uses of imperfect predictions but do not show that current digital twins satisfy the conditions reliably enough to reduce human measurement.
  • Statistical calibration: Transporting calibration from Moore-Berg respondents to Twin-2K respondents requires assumptions about population overlap and stability of the human-model error relationship.
  • Open research question: The paper evaluates aggregate fidelity, paired signal, finite-sample recovery, and cross-population stability because behavioral realism alone is insufficient for statistical substitutability.

3 Research Design and Methods

The research design evaluates whether digital-twin predictions can reduce human measurement for a specified estimand while preserving valid inference and reference precision. It separates aggregate fidelity from respondent-level signal, finite-sample feasibility, and population validity, using prediction-powered inference to assess achievable savings.

  • Evaluation criterion: Statistical substitutability requires a reduced-human design to preserve valid inference for the target estimand while attaining at least the human-only reference precision.The workflow specifies the estimand, target population, and precision benchmark before evaluating twin predictions.
  • Prediction-powered inference: Twin responses enter as auxiliary predictions, while observed human outcomes remain the inferential baseline in prediction-powered estimation.PPI combines a smaller human validation sample with prediction-only observations and corrects model error using paired outcomes.
  • Evaluation dimensions: The framework separately evaluates aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations.Aggregate agreement alone does not establish that predictions track individual departures from relevant group means.
  • Twin signal: The value of a digital twin depends primarily on paired association with human outcomes, not aggregate fidelity or predictive accuracy alone.For a population mean, the idealized variance-reduction ceiling approaches ρ^2 as the prediction-only sample becomes large relative to the human sample.
  • Twin-2K benchmark: For two-condition effects, respondent-level signal is measured after removing condition means, because reproducing an aggregate treatment effect may not identify individual departures.A twin can match the treatment contrast while having near-zero residualized human-twin correlation.
  • Empirical evaluation: The design tests human-label budgets of 25, 50, 100, 115, 250, and 500 with 500 repetitions to compare aggregate agreement and paired signal.The Moore-Berg extension also examines calibration transport, but its target sample lacks human outcomes needed to establish target substitutability or accuracy.

4 Results

Across both evaluations, digital twins often reproduce aggregate human patterns without providing the respondent-level signal needed for valid human-data savings. Newer models and calibration improve some dimensions, but gains in statistical substitutability remain modest and inconsistent.

  • Twin-2K results: Framing reproduces a strong aggregate effect but has almost no respondent-level signal.The human standardized effect is 0.761, the twin effect is 0.941, and residualized correlation is 0.003.
  • Twin-2K results: 0.260 is the largest Twin-2K residualized correlation, yet Proportion Dominance reverses the human effect direction.Anchoring also shows that matching direction or signal alone does not ensure matching effect magnitude.
  • Twin-2K results: The median residualized correlation is 0.028, with 10 of 12 contrasts below |𝜌| = 0.10.That threshold corresponds to less than 1% idealized variance reduction; the strongest task reaches only 6.7%.
  • Twin-2K results: Aggregate replication can occur without the paired respondent-level information needed to reduce uncertainty about the human estimand.This is the central negative finding for most Twin-2K tasks.
  • Finite-sample feasibility: Finite-sample PPI requires stronger correlations than the theoretical ceilings provide: approximately 0.316–0.318 versus a maximum observed 0.260.At the 90% label fraction, no realized-study repetition produced a finite prediction-only requirement, and more aggressive reductions also failed.
  • Finite-sample feasibility: In the reduced-size sensitivity design, Proportion Dominance succeeds in 17.8% of repetitions, while Myside Bias succeeds in only 3.0%.Among successful Proportion Dominance repetitions, conditional variance reduction is 2.9%.
  • Moore-Berg results: Later models improve aggregate source replication, but respondent-level signal and efficiency gains do not improve uniformly.Mean absolute source deviation falls across all three model families, while Qwen’s mean residualized correlation declines from 0.195 to 0.187; the highest oracle fixed variance reduction is 6.9%.
  • Moore-Berg results: Model rankings differ across aggregate deviation and respondent-level correlation, so no single metric identifies the most useful model.GPT-5.4 has the highest mean residualized correlation at 0.222, whereas Qwen3.8 has the lowest mean absolute deviation at 15.8.

5 Discussion

The discussion distinguishes behavioral fidelity from statistical substitutability: reproducing human averages does not ensure useful respondent-level information or human-data savings. The paper therefore argues for evaluating digital twins by their inferential usefulness, while recognizing scope boundaries and exploratory applications.

  • Core distinction: Aggregate fidelity and statistical substitutability are distinct: a twin can reproduce an average effect while providing little information about individual departures.Framing closely reproduces the standardized human effect, yet its residualized human-twin correlation is near zero.
  • Core distinction: Finite-sample feasibility depends on the full design, because limited validation data and prediction-only pools can eliminate gains suggested by theoretical efficiency ceilings.A nonzero residualized correlation alone does not establish an achievable reduction in human measurement.
  • What improves—and what does not: Stronger models, richer respondent representations, and calibration can improve aggregate alignment without reliably improving respondent-level signal or inferential efficiency.Human calibration corrects error but cannot create paired signal.
  • Implications: Weak paired signal does not make digital twins useless; they may still support hypothesis generation, treatment screening, and instrument development.These uses remain exploratory rather than confirmatory substitutes for human measurement.
  • Implications: Confirmatory use requires estimand-specific evaluation with independent human validation and assessment of respondent-level departures and attainable precision.The paired human-twin data support evaluation of prediction error, respondent-level tracking, and precision.
  • Scope and limitations: Transporting calibration across populations requires evidence that the human-model error relationship remains stable in the target population.Without human outcomes in the Twin-2K target sample, target accuracy and statistical substitutability cannot be directly validated.

A PPI Variance and Sample Requirement Details

This appendix derives the variance-reduction ceiling and finite prediction-only sample requirement for prediction-assisted estimation. It distinguishes optimized, fixed, and adaptive coefficients and states when population-level guarantees do or do not apply.

  • Variance reduction: The variance-reduction ceiling is obtained by optimizing the coefficient multiplying twin predictions relative to human-only variance.The unconstrained variance-minimizing coefficient is then restricted to [0, 1] when power tuning is imposed.
  • Variance reduction: Setting the coefficient to zero returns the human-only estimator, so the constrained population optimum cannot be less precise than human-only estimation.This guarantee applies to the optimized population procedure, not necessarily to fixed-coefficient implementations.
  • Interpretation: The effective-sample-size multiplier expresses prediction-assisted precision as the equivalent number of human-only observations.Under the inverse-sample-size approximation, it reports the human-only sample needed to match the prediction-assisted design.
  • Finite-sample planning: A finite prediction-only requirement exists only when the estimated-correlation threshold is met and enough eligible prediction-only observations remain after labeling.The design must also satisfy the available-pool condition in the empirical benchmark.

B Moore-Berg Transported PPI Coefficients and Efficiency

The Moore-Berg appendix defines transported PPI estimands and compares unit, adaptive, and oracle fixed coefficients. The available signal does not automatically yield feasible variance reduction, and transport requires assumptions about cross-population stability.

  • Transport setup: The transported analysis estimates a human source contrast, a model-predicted source contrast, and a model-predicted target contrast across Moore-Berg and Twin-2K samples.The source contains human outcomes and predictions, whereas the target contains only model predictions.
  • Coefficient choices: The unit estimator fixes the prediction weight at one, while the adaptive estimator re-estimates it from each labeled source subset.The adaptive coefficient can be noisy when the labeled subset is small or prediction contrast has little stable variance.
  • Coefficient choices: The oracle fixed estimator uses the complete source sample to diagnose potential efficiency but is not feasible under a low-label budget.It asks how much efficiency the signal could provide if the coefficient were known precisely.
  • Results: At n_L = 115, available oracle signal is modest across several models, and unit or adaptive procedures do not automatically convert it into feasible variance reduction.Table B.1 reports the cross-model snapshot for transported efficiency.

B.1 Why adaptive weighting can fail under transport

Under source-to-target transport, adaptive weighting can turn useful prediction signal into unstable or harmful variance adjustments because coefficient error is amplified by changing prediction contrasts. In the common-fields analysis, feasible adaptive estimation does not reliably recover the information available under oracle coefficients.

  • A nonzero residualized correlation does not guarantee variance reduction for unit or adaptive estimators.
  • GPT-5.4 achieved positive unit-coefficient variance reduction at every reported transported budget, ranging from 2.5% to 5.5%.
  • Coefficient error is multiplied by the difference between target and source prediction contrasts, which can be substantial under transport.
  • The adaptive coefficient was negative throughout the budget grid for five models, while GPT-4.1-mini stayed near zero.
  • Oracle fixed coefficients remained positive for all six models, but feasible coefficient estimation did not reliably recover these modest gains.

C Reproducibility and Documentation

The study documents its data, model-generation, validation, analysis, and human-verification procedures. Repositories preserve the code, prompts, metadata, outputs, and related research files supporting reproducibility.

  • Twin-2K uses 500 repetitions per reference design and fraction, with labeled and prediction-only respondents disjoint by identifier.
  • The documentation identifies released predictions, six Moore-Berg models, prompt regimes, generation and validation workflows, verification procedures, and observed failure modes.
  • LLMs served as research infrastructure and implementation tools, while human judgment remained central to study selection, estimand reconstruction, design, verification, and interpretation.
  • Run-specific model strings, parameters, timestamps, retries, raw outputs, and metadata are preserved, but model-generation comparisons are descriptive rather than controlled scaling experiments.
  • The pipeline validates structured outputs, bounds numeric responses to the intended 0–100 scale, retains errors and raw outputs, and deduplicates successful records.
  • Prospective studies must protect the labeled validation sample because synthetic or assisted panel respondents could compromise calibration and evaluation.

D.5 Failure Modes, Risks, and Mitigations

The study identifies technical, model, privacy, and data-integrity risks and pairs them with validation, reporting, mitigation, and accountability procedures. It emphasizes that newer models may improve one target while worsening another.

  • Invalid, incomplete, missing, or out-of-range responses are addressed through required-key validation, parse flags, numeric bounding, retries, and exclusion of unsuccessful records.
  • Duplicate retry rows are controlled by deduplicating respondent identifiers and retaining the latest valid record.
  • Newer models may improve one validation target while worsening another, so aggregate, individual, and inferential metrics should be reported separately.
  • Rich person-specific profiles may reveal sensitive attributes or misrepresent individuals and groups, requiring anonymized identifiers, approved data resources, and human accountability.
  • GenAI assisted research execution but did not replace human responsibility for evidence selection, statistical design, verification, interpretation, or authorship.

E Twin-2K Study Scope

The Twin-2K scope centers on primary two-condition behavioral contrasts with predefined estimands and respondent-information regimes. Several tasks are excluded because they do not share the common estimand structure, not because of model performance.

  • Twin-2K’s primary contrasts use a common two-condition estimand, with four behavioral quantities defined by perceived out-party bias.
  • False Consensus, Nonseparability of Risk and Benefit, and other listed tasks are excluded because they lack the common two-condition estimand.
  • The first two estimands measure Democrats’ overestimation of Republican prejudice and dehumanization, while the other two reverse the parties.
  • Human reference values for the four estimands are 23.95, 33.20, 25.43, and 37.62.
  • Target prompt regimes vary respondent information from party identity alone to structured common fields and broader persona summaries.

G Finite-Sample Sensitivity

This section examines finite-sample sensitivity under a Cohen d=0.5 condition at a 90% human-label fraction, using success criteria that combine prediction-only sizing, disjoint capacity, and PPI estimation.

  • Table G.1 reports sensitivity results for Cohen d=0.5 at the 90% human-label fraction.
  • Success requires a finite prediction-only sample size, sufficient remaining disjoint capacity, and successful PPI estimation.
  • Conditional variance reduction is computed only over successful repetitions and is unstable when success is rare.

H Full Moore-Berg Component Diagnostics

Figure H.1 presents humanity-rating components for six models in the common-fields regime, comparing survey-weighted human source means with raw and calibrated twin predictions. Because target human outcomes are unobserved, movement toward the source reference does not establish target accuracy.

  • Figure H.1 compares humanity-rating components across six models in the common-fields regime.
  • Human bars represent survey-weighted Moore-Berg source means, while raw and calibrated bars represent Twin-2K target predictions.
  • Calibrated bars apply source corrections estimated from 500 draws with 115 labeled Moore-Berg respondents per draw.
  • Each Δ reports the within-panel party gap, with panels including Republicans' and Democrats' actual ratings and Republicans' meta-ratings.
  • Target human outcomes are unobserved, so movement toward the source reference does not establish target accuracy.
Loading 2609.07987v1…