Source-linked AI summary

Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments

Ho Ting Hung, Nachiket Midha, Victor Y. Wu, Yiwen Zhang

arXiv:2609.04243v1cs.MAcs.CYstat.ME

TL;DR

The paper asks whether synthetic agents can reproduce the multi-dimensional human preference patterns studied in conjoint experiments, a question with limited evidence despite growing interest in these agents. It replicates published conjoint studies and compares synthetic and human results across representational, inferential, and procedural dimensions. Results are uneven and do not support general replacement, indicating that synthetic-agent validity is claim-dependent and hierarchical.

  • Problem

    Whether synthetic agents can be applied specifically in conjoint experiments remains underexplored, despite growing interest in using them in survey experiments.

  • Method

    The paper replicates published conjoint studies and evaluates synthetic-agent results against human data across representational correspondence, inferential correspondence, and procedural stability.

  • Results

    Overall, the results do not strongly support a general claim that synthetic agents can replace human participants in conjoint experiments.

  • Takeaways & Limitations

    Synthetic agents are better treated through a claim-dependent validation hierarchy rather than as general substitutes for human samples.

  • Takeaways & Limitations

    Because replicated studies are published and widely cited, aggregate agreement may partly reflect recall of documented results, making the findings an upper bound of synthetic performance.

Abstract

from arXiv · show

Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.

1. Introduction

This paper examines whether synthetic agents can replace human participants in conjoint experiments, where multi-dimensional preferences and trade-offs are central. Replicating published studies, it finds uneven correspondence across aggregate distributions, estimates, individual choices, subgroups, and model stability.

  • Research approach: The paper evaluates synthetic participants by replicating conjoint experiments published in major political science journals and comparing their results with original human data.The comparison examines how stable and similar synthetic results are relative to human participants.
  • Interpretive framework: Synthetic-agent validity is claim-dependent and hierarchical rather than a binary question of whether agents can replace human respondents.Different uses require different levels of evidence, from aggregate reproduction to more demanding individual and inferential correspondence.
  • Core findings: Synthetic agents often approximate marginal attribute-level distributions and sometimes recover estimate directions, but fail on several more demanding correspondence tests.Failures include full joint profiles, individual-level choices, precise effect magnitudes, subgroup heterogeneity, and stability across models.
  • Implications: The uneven findings do not support reliable substitution of human samples in conjoint experiments.The paper argues that apparent aggregate agreement may partly reflect recall of documented findings, yet synthetic agents still fail to align consistently with human samples.
  • Implications: The paper proposes a validation hierarchy whose evidentiary requirements vary with intended use before high-stakes deployment.Its contribution is to establish evidence about substitutability in conjoint experiments and map the innovation’s boundaries.

2. Literature Review

The literature frames conjoint experiments as tools for studying multi-dimensional preferences through attribute combinations and causal estimands beyond standard treatment effects. It also presents synthetic agents as efficient but potentially biased proxies whose aggregate agreement may not extend to distributions, effect sizes, or generalizable conclusions.

  • Conjoint Experiments: Conjoint experiments manipulate multiple attribute levels, combine them into profiles, and present those profiles in choice sets.Their many possible attribute combinations make estimating treatment effects for every complete combination infeasible.
  • Conjoint Experiments: The average marginal component effect estimates an attribute level’s average effect relative to another level while averaging over the joint distribution of other attributes.Conditional AMCEs restrict this comparison to profiles sharing specified conditioning-attribute values.
  • Synthetic Agents: Synthetic agents may reduce survey fatigue, attention constraints, recruitment costs, and validation burdens while enabling diverse response simulations.Prior work also highlights their potential for representing aggregate human preferences and exploring intersectionality.
  • Synthetic Agents: Critics argue that language models can collapse diverse perspectives into modal opinions, which is problematic when political science studies preference distributions rather than only averages.Other concerns include data and algorithmic bias, hallucinations, static value lock-in, and limited training data.
  • Synthetic Agents: Synthetic agents can reflect human preferences to some extent, but may overestimate or underestimate effect sizes and produce different results under configuration choices.Model, hyperparameter, scale, and prompting decisions can drive significantly different conclusions.
  • Research gap: These concerns motivate evaluating synthetic agents in conjoint experiments beyond a binary replacement question.The paper develops a multi-level, claim-dependent framework for assessing this intersection.

3. Hypotheses

The paper preregisters three dimensions for evaluating whether synthetic agents mirror human conjoint responses: representational correspondence, inferential correspondence, and procedural stability. The hypotheses progressively require similarity in distributions, estimates, and robustness across analytic choices.

  • Representational correspondence: The first dimension asks whether synthetic agents reproduce human choice distributions at joint, marginal, and individual levels.Marginal agreement may coexist with different joint profiles, while the strongest benchmark requires reproducing individual respondents’ choices.
  • Representational correspondence: Hypothesis 1 predicts similar overall choice distributions and relative attribute weights between synthetic replications and original human studies.The design distinguishes broad feature popularity from how attributes are combined in complete profiles.
  • Inferential correspondence: The second dimension asks whether synthetic responses support the same substantive and statistical inferences as human responses.The benchmark progresses from estimate direction to relative ordering, magnitude, uncertainty, significance, and heterogeneous effects.
  • Inferential correspondence: Hypothesis 2 predicts similarity between synthetic and human estimates in direction, magnitude, and significance.Matching signs alone is weaker than matching effect sizes, uncertainty, and the broader ordering of effects.
  • Procedural stability: The third dimension tests whether synthetic estimates remain stable across reasonable analytic choices in the generation procedure.This guards against results that depend on a favorable model or parameter setting when no specification is uniquely justified.
  • Procedural stability: Hypothesis 3 predicts robust and stable synthetic estimates across various analytic choices.The absence of an external basis for selecting one uniquely appropriate specification makes robustness testing essential.

4. Research Design

The study replicates six published conjoint studies across 13 experimental setups, comparing synthetic agents with matched human samples on distributional similarity, causal-estimate alignment, and stability across analytic choices.

  • Scope: Six published conjoint studies provide 13 experimental setups for benchmarking synthetic agents against human data.The replication design follows Clayton et al. (2026).
  • Scope: The benchmark uses GPT-4o, GPT-4o mini, Llama 3.2 (3B), Llama 3.3 (70B), and Gemini 2.5 Flash.The study avoids overly advanced models to target realistic and affordable research practices.
  • Synthetic-agent design: A one-to-one persona-mirroring strategy matches synthetic agents to human demographic information and profile encounters to minimize design variance.The baseline temperature is T = 0.5, balancing diversity and predictability.
  • Distributional correspondence: Distributional correspondence is assessed with joint Wasserstein distance, marginal Hellinger distance, and Pearson correlations of attribute-level choice shares.The Wasserstein ground distance uses Manhattan distance between one-hot-encoded attribute profiles, preserving informativeness in sparse conjoint spaces.
  • Inference and testing: The design uses respondent-level bootstrapping and permutation tests, with 2,000 bootstrap samples and Benjamini-Hochberg-adjusted p-values.The permutation framework calibrates the Wasserstein comparison against sampling noise and a bias-corrected null distribution.
  • Inferential correspondence: Inferential correspondence compares AMCE, MM, cAMCE, and other estimands using estimate correlations, RMSE, sign agreement, and overlap with human 95% confidence intervals.Weighted least squares reproduces original survey-weighted specifications where applicable.
  • Procedural stability: Procedural stability decomposes estimate variability across models, temperatures, and model-by-temperature combinations, distinguishing analytic-choice variability from sampling error.The stability ratio uses R_m = 1 as the threshold where between-setting variability exceeds typical within-setting estimation uncertainty.
  • Inference and testing: Interpretation thresholds are pre-registered for Hypothesis 1, while the RMSE threshold treats 0.05 as substantively large relative to typical human conjoint effects.A methodological limitation is that some heterogeneous-effect analyses require survey responses omitted from the simulations, and original studies may use different estimands.

5. Results

Across replicated conjoint studies, synthetic agents often approximate aggregate distributions and effect directions but show uneven magnitude agreement and substantial instability across models and analytic settings.

  • Distributional correspondence: Across all 13 setups and five models, marginal Hellinger-distance and Pearson-r confidence intervals satisfy their thresholds, but no individual-level combination passes.Llama 3.2 performs worse than the other models, especially at T = 0.5.
  • Interpretive caveat: Non-rejection in the permutation tests does not establish distributional equivalence because sparse profile support can inflate the null distribution and reduce test power.Rejections provide stronger evidence of divergence than non-rejections provide evidence of similarity.
  • Inferential correspondence: Synthetic estimates frequently match human AMCE directions while missing their magnitudes: human-CI overlap rarely exceeds 0.7 and RMSE often exceeds 0.05.The discrepancy is especially pronounced for Bechtel and Scheve (2013) and Hankinson (2018).
  • Inferential correspondence: AMCE alignment is relatively stable for Hainmueller and Hopkins (2015), with lower RMSE across models and temperatures, although no setting passes the threshold.Similar patterns hold for MM estimates.
  • Inferential correspondence: Heterogeneous-effect alignment is weaker than aggregate AMCE alignment: sign agreement can be strong while Pearson-r, confidence-interval overlap, and RMSE remain uneven or poor.Teele et al. (2018) exemplifies strong sign agreement but much weaker confidence-interval overlap.
  • Procedural stability: Model choices generate instability exceeding ordinary estimation uncertainty for many estimates, with R_m > 1 across all studies and sometimes R_m much greater than 1.Temperature alone generally contributes less instability than model choice or model-by-temperature interactions.

6. Discussion

The results do not support a general replacement claim, but they reveal a hierarchy: synthetic agents perform best on broad aggregate benchmarks and weaken as evaluations target joint distributions, effect magnitudes, subgroup heterogeneity, and stability.

  • Claim-dependent hierarchy: Synthetic agents perform best in broad aggregate benchmarks, but their performance weakens as evaluation targets become more demanding.The hierarchy extends from low-stakes aggregate exploration to inferential and respondent-level substitution, which require increasingly accurate magnitudes, uncertainty, and individual variation.
  • Representational correspondence: Marginal attribute-level alignment can coexist with failures to reproduce joint distributions and individual-level choice structure.LLMs may match average preferences while collapsing heterogeneous perspectives toward modal responses, obscuring covariance and attribute trade-offs.
  • Implications: A similar preference distribution or favorable statistic does not establish a similar statistical estimate or justify respondent replacement.The authors argue that synthetic agents are better understood as imperfect aggregate simulators and that the discipline needs systematic boundary mapping before treating them as robust human substitutes.
  • Inferential correspondence: Synthetic agents sometimes recover the direction of human effects, but they are much less reliable for precise effect magnitudes and subgroup-specific preference structures.Strong sign agreement can therefore coexist with estimates whose magnitudes remain far from human estimates, while conditional AMCE heterogeneity is weaker than aggregate AMCE alignment.
  • Procedural stability: Stability varies across studies, estimands, analytic choices, models, and temperature settings, so the same conjoint design can yield different conclusions.No model is consistently superior across studies, estimands, and metrics; fixed-model temperature contributes little to instability without making estimates robust.

7. Conclusion

The paper finds that synthetic agents can approximate broad marginal patterns and sometimes recover preference directions, but remain unreliable for more demanding conjoint targets and as human substitutes. Their usefulness is therefore claim- and use-case-dependent, with potential value for exploratory and robustness-oriented work.

  • Synthetic agents often approximate broad marginal patterns and sometimes recover the general direction of human preferences.Their strongest performance appears at the most aggregated level.
  • They perform poorly on full joint distributions, individual choices, precise effect magnitudes, subgroup heterogeneity, and stability across settings.These weaknesses extend beyond basic directional agreement and vary across studies, estimands, and analytic choices.
  • Directional similarity is insufficient when estimates may be too large, too small, or unstable enough to produce misleading substantive conclusions.
  • Evidence at a lower validation level cannot justify replacement claims at a higher level of correspondence.Reproducing aggregate outputs or obtaining sign agreement does not establish reliable individual-level or inferential equivalence.
  • Synthetic agents may still support robustness checks, exploratory analyses, replication diagnostics, and simulated pilots before costly human deployment.Their capacity to reflect aggregate preferences can help identify published findings that merit further investigation and provide an initial aggregate picture.
  • The conclusions are bounded by typical conjoint designs, exclusion of advanced chain-of-thought models, and possible recall of documented results from training data.Observed aggregate alignment may represent an upper bound rather than a typical expectation, while some heterogeneity analyses require additional variables or simulations.

C. Appendix: Choice Extraction

Choice extraction is generally valid, but invalid responses arise from reluctance or parsing edge cases, with Llama 3.3 showing particular difficulty.

  • Most study-model combinations have perfect valid-response rates, so invalid responses are a minority overall.
  • Invalid responses mainly reflect reluctance to answer or difficult-to-parse edge cases that the analysis ignores.
  • Llama 3.3 struggles more with valid extraction and sometimes outputs only “Choice: PROFILE” without identifying the selected profile.

D. Appendix: Wasserstein Distance Calibration

Permutation calibration evaluates whether model–human distributional divergence exceeds finite-sample and support-sparsity baselines. Divergence remains common despite sparse empirical support, including in a high-density benchmark.

  • The calibration has limited power under sparse support, so true population divergence can be nearly indistinguishable from baseline sampling noise.
  • 12 of 13 replicated study setups exhibit effective empirical support sparsity, with mean observation-to-support ratios of 1.00-2.93.
  • Support sparsity is especially severe in several setups, including approximately 48% singletons in Arias and Blair (2022) and 25-32% in Teele et al. and Hankinson.
  • 73.3% of cases show significant calibrated model–human distributional divergence, while only 3.03% are flagged by raw W1 but cleared by permutation.
  • Divergence is not simply a sparsity artifact: Bechtel and Scheve (2013) has dense support, yet every model fails the Wasserstein test with W_excess > 0 and p < 0.001.
  • The analysis uses dense marginal attribute-choice distributions and cluster-level resampling to address metric and dependence problems in the preregistered analysis.

G. Appendix: Additional Heterogeneity Analysis

Additional heterogeneity analyses show relatively strong aggregate alignment but weak individual alignment across conditional, profile, and subset estimands. Hankinson (2018) is especially unstable across subgroup conditions.

  • Conditional MM, profile MM, subset AMCE, and subset cAMCE reproduce the main pattern of strong aggregate but weak individual alignment.
  • Hankinson (2018) has extremely wide intervals and sharply varying performance across subgroup conditions.
  • Subset cAMCE can produce median Pearson r below 0, including instances reported for Ono and Burden (2019).

H. Appendix: Additional Stability Ratio Results

Additional stability-ratio results show that temperature sensitivity often appears low in point estimates, but uncertainty remains substantial around some ratios.

  • Temperature-based stability-ratio confidence intervals often cross the threshold rather than falling decisively below it.
  • Point estimates suggest low temperature sensitivity, while uncertainty around some temperature-based ratios remains non-negligible.
  • Figures G.1-G.4 report heterogeneity across conditional MM, profile MM, subset AMCE, and subset cAMCE analyses.

I. Appendix: Additional Variance Analysis

Additional variance analyses show that synthetic conjoint results vary across settings and analytic choices. Between-setting variance and estimate ranges are especially pronounced for model-related settings.

  • Between-setting variance estimates are consistently non-zero for many attribute-level estimates, especially in model and model-by-temperature settings.
  • Together, the random-effects variance estimates and stability ratios show that synthetic conjoint results are sensitive to analytic choices.
  • Model and model-by-temperature settings generally produce wider estimate ranges across analytic choices than temperature-only settings.
Loading 2609.04243v1…