Source-linked AI summary
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
Mantas Lukauskas, Viktorija Šarkauskaitė
TL;DR
Existing LLM survey evaluations emphasize plausible individual answers rather than whether synthetic respondents preserve human psychometric structure. This paper introduces a Lithuanian benchmark spanning 37 models and finds that LLMs reproduce relationship direction but exhibit substantial psychometric fidelity failures, including outperforming Gaussian-copula baselines.
Problem
Existing benchmarks largely use English US or UK instruments, leaving unclear whether LLMs reproduce the joint psychometric distribution of real human samples.
Method
The paper evaluates 37 LLMs on an open Lithuanian organisational-psychology benchmark using a multidimensional PSS framework and human and statistical comparison anchors.
Results
LLMs reproduce the qualitative direction of human psychometric relationships but generate over-coherent, low-variance, demographically stereotyped, and inter-model-homogeneous respondents.
Takeaways & Limitations
LLM synthetic respondents should not be treated as psychometrically faithful replacements for human survey data.
Takeaways & Limitations
The human reference is a single Lithuanian organisational sample, limiting external validity to other countries, sectors, and constructs.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM "crowd" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.
1 Introduction
This paper argues that evaluating LLMs as synthetic survey respondents requires psychometric rather than item-level plausibility, testing whether they preserve human covariance, latent structure, reliability, mediation, and demographic effects. It introduces an open Lithuanian benchmark and a preregistered evaluation framework to identify where synthetic respondents diverge from human data and the practical costs of substitution.
- Evaluation design: The study evaluates 37 LLMs across persona disclosure, presentation, reasoning-effort, cross-language, and counterfactual demographic conditions against five non-LLM baselines and a held-out human ceiling.The framework also uses respondent-bootstrap confidence intervals, Holm–Bonferroni-adjusted p-values, and an item-permutation null for Tucker’s phi.
- Motivation: Psychometric validity requires items to co-vary correctly, reflect latent factors, yield reliable scores, and relate to other constructs as theory predicts.Matching item means alone can leave covariance structures wrong, inflate internal-consistency statistics, and produce different latent-variable solutions.
- Benchmark: The benchmark uses 263 Lithuanian employees, 68 items, 12 subscales, three validated instruments, and documented Attitudes→Engagement→Performance mediation.The dataset includes ATC, UWES-17, IWPQ, and a 15-item change-engagement scale.
- Downstream consequences: The study assesses practical substitution costs through response-style bias, synthetic-to-human predictive validity, mediation negative controls, cohort-stratified fairness, inter-LLM clustering, and ensemble analysis.These diagnostics extend evaluation beyond raw fidelity to acquiescence, extreme- and midpoint-responding, fabricated indirect effects, and model agreement.
- Central finding: LLMs can reproduce the theoretical direction of human psychometric relationships but tend to generate over-coherent, low-variance, and demographically stereotyped respondents.The paper therefore examines which fidelity dimensions fail and how failures depend on persona conditioning, rather than classifying survey simulation as simply good or bad.
2 Related work
Prior studies show that LLMs can generate plausibly human responses and reproduce some qualitative human patterns, but also exhibit systematic demographic and identity-group biases. This work extends the literature by auditing psychometric properties across distributions, reliability, mediation, demographic effects, factor structure, and prompting conditions.
- LLMs as synthetic respondents: Prior studies established that LLMs can act as synthetic respondents, replicate behavioral and psychological experiments, and reproduce qualitative human attitude patterns [2] [1] [17].These studies primarily assessed single items, marginal effects, or aggregate outcomes rather than full psychometric structure.
- Limits of LLM survey simulation: Existing audits find systematic divergence from target populations, including liberal, college-educated US views and generic or unreliable representations of identity groups [9] [24].The literature particularly highlights failures along intersectional dimensions and collapse of single-attribute personas into generic profiles.
- Limits of LLM survey simulation: This work isolates failures across joint distributions, correlations, reliability, mediation, demographic effects, and factor structure while testing persona conditioning, presentation, reasoning effort, and stochastic stability.It therefore evaluates psychometric properties rather than only whether individual responses appear plausible.
- Persona prompting and synthetic-data pipelines: Persona designs range from narrative scaffolding and structured profile prompts to category labels and reasoning traces, with profile granularity and other design choices interacting [17] [2] [9] [24].The related work frames prompting as a design space rather than a single “answer as a person” instruction.
- Counterfactual fairness and stereotype amplification: Adjacent methods perturb demographic attributes to study stereotype sensitivity, test cross-language measurement invariance, and quantify reliability, mediation, covariance similarity, factor congruence, and invariance.These literatures motivate counterfactual audits and the paper’s use of established psychometric criteria, including an item-permutation null for Tucker’s φ.
3 Dataset: an open Lithuanian organisational psychology survey
This open dataset contains 263 Lithuanian employees’ de-identified organisational-psychology survey responses, collected under informed consent and spanning diverse demographic and employment contexts. It combines validated work-performance, engagement, and change-attitude instruments with profile cards and a reproduced partial-mediation benchmark.
- Human sample and provenance: The dataset comprises 263 Lithuanian business-organisation employees recruited through an anonymous online questionnaire in March–April 2020 under informed consent.The record-level data are de-identified and released with permission under CC-BY-NC-4.0.
- Human sample and provenance: The sample is female-skewed, spans ages 19–62, includes 33% managers, and covers multiple sectors and organisation sizes.Women comprise 76% of respondents and men 24%; 67% are non-managerial.
- Measures: The instrument battery measures individual work performance, work engagement, and attitudes toward organisational change using Lithuanian-translated, multi-subscale Likert instruments.IWPQ includes 18 items and three subscales, while UWES-17 includes 17 items and three subscales; the translations follow established Lithuanian usage or double-translation procedures.
- Mediation reproduction: The pipeline reproduces the published Attitudes → Engagement → Performance partial mediation, with an indirect effect of ab = 0.16 and 95% bootstrap CI [0.10, 0.24].The direct path is positive but smaller than the indirect path, and the indirect effect is used to score direction, magnitude, and significance matches in the LLM-side mediation component.
- Profile cards: LLM conditioning uses profile cards containing 11 demographic and organisational fields rather than respondents’ item responses.The cards include gender, age, education, role, tenure, sector, organisation size, and two change-context ratings, and are intended for public re-use.
4 Method
The study evaluates 37 LLMs as synthetic survey respondents using controlled persona, presentation, language, sampling, and reasoning conditions. It scores each model on six psychometric dimensions against statistical baselines and a held-out human ceiling.
- Persona conditions: C3 is the headline persona condition, providing gender, age, role, and education without exposing free-text or item responses.The five-level disclosure ladder ranges from C0 with no profile through progressively richer demographic profiles; C3 is justified as the most informative fair test and is checked for actual conditioning.
- Experimental conditions: M1 presents all 68 items together as the primary mode, while M2 and M3 split responses by instrument or subscale to test whether cross-instrument context supports coherent measurement.Headline runs use Lithuanian prompts; English is a matched cross-language ablation using identical respondent profiles.
- Sampling and generation: Each model-condition cell samples n = 100 stratified respondents matched to the human joint gender × age-band × role distribution.Stratification improves coverage of lower-frequency demographic cells, while the headline grid uses one repeat per cell and minimal reasoning effort.
- Psychometric evaluation: The benchmark scores each model-condition cell on six dimensions spanning distributions, correlations, reliability, mediation, demographic effects, and construct-level fidelity.Construct fidelity uses factor alignment, Tucker’s φ with an item-permutation null, bifactor statistics, and inter-subscale correlation gaps.
- Baselines and ceiling: PSS is benchmarked against statistical generators and an 80/20 held-out human-vs-human ceiling of around 0.83.The baselines include marginal sampling and MVN covariance-preserving synthesis; the human ceiling estimates the maximum attainable score at the study’s sample size.
5 Results
At C3, LLMs approach but do not match human psychometric fidelity: statistical baselines outperform them on sample-driven structure, while LLM-specific gains are limited to conditioning-sensitive components. Additional analyses reveal presentation sensitivity, chance-inflated factor congruence, and systematic over-coherence in synthetic responses.
- Headline results: PSS 0.714 was the best LLM score, below the held-out human ceiling of 0.825; Gaussian-copula and MVN baselines matched c = 0.95, r = 0.99, versus c = 0.52, r = 0.94 for the best LLM.The copula reached combined PSS 0.688, close to the best LLM at 0.714, while LLMs gained ground mainly on mediation and demographic-effect components.
- Validation and scaling: AUC 0.999 was the median separation of LLMs from humans, whereas Gaussian-copula and MVN baselines were indistinguishable from humans at AUC 0.39 and 0.52.As sample size increased, the copula improved from PSS 0.69 at n=10 to 0.82 at n=200, while the best LLM plateaued near 0.62 by n≈50.
- Conditioning and presentation: Persona disclosure improved mean per-model PSS by 0.18 from C0 to C3, but narrative C9 presentation shifted PSS by a median absolute 0.06 and correlated only 0.31 with structured C3 rankings.The C3 leaderboard is therefore conditional on structured presentation, and small rank gaps are sensitive to how profiles are rendered.
- Factor congruence: Permutation-corrected Tucker’s φ showed genuine factor preservation for IWPQ (mean φ=0.821; 32/37 rejections), but only 29/37 UWES models rejected the null and ChangeEng was weakest (φ=0.458; 22/37).UWES had mean null φ=0.771, placing eight models with raw φ up to 0.856 inside the chance distribution; some averages also exclude nonconverged solutions.
- Discriminant and latent validity: HTMT violations affected 7.7/12 within-instrument subscale pairs in LLM samples versus 4/12 in humans, indicating systematic over-coherence.At C3, LLMs also increased general-factor explained-common-variance by ∆ECV = +0.08 to +0.13 and collapsed Change-engagement specific-factor reliability by ∆ωh= −0.93.
6 Discussion
The discussion concludes that LLMs recover broad psychometric directions but fail as human-survey replacements because statistical baselines better preserve sample-driven structure and persona conditioning introduces systematic distortions. It recommends baseline-anchored, ceiling-calibrated, permutation-corrected evaluation and cautious use limited to pilot direction-of-effect analyses.
- Interpretation of psychometric relationships: 36 of 37 models recover the human mediation direction under C3 persona conditioning, although standardised path coefficients are typically attenuated relative to humans.The exception, grok-4-20, produces a negative indirect path.
- Failure modes: σLLM/σH is 0.53 on IWPQ items and 0.44 on UWES items for the median model, indicating range restriction that inflates reliability and distorts correlations.LLMs use the full Likert range less freely than humans, even with an anti-social-desirability instruction.
- Failure modes: UWES falls inside the item-permutation null for 8 of 37 models despite φ ≥0.85, showing that raw Tucker’s φ can overstate factor-structure preservation.The discussion therefore treats permutation correction as necessary for similarity metrics whose alignment is searched greedily.
- Memorisation probe: 4.7% is the worst-case high-recall rate, 22/37 models have zero recall, and recall-PSS correlation is 0.00, ruling out verbatim memorisation as the leaderboard driver.Observed probe behaviours include refusal, confabulation, and leaked English chain-of-thought.
- Implications for measurement: ∼0.03 PSS units separate the best LLM from the Gaussian-copula baseline, while PSS varies by ∼0.35 across 37 models, making LLM-only leaderboards misleading.The Gaussian copula preserves distribution, correlation, and reliability without understanding the constructs, whereas LLM contributions appear mainly in mediation and demographic effects.
7 Limitations and future work
The benchmark’s limitations concern external validity, self-report and instrument exposure, sample size, reasoning settings, model drift, and the scope of causal and construct generalization. Future work should replicate the study across languages and instruments, enlarge repeated samples, and test longitudinal, open-text, and counterfactual stability.
- Limitations: External validity is limited by one Lithuanian organisational sample (n = 263), although LT↔EN drift was bounded at |∆| = 0.26 Likert points with per-item Pearson 0.889.A same-respondent human cross-language test was unavailable; follow-up replication should include non-Indo-European samples such as Korean or Mandarin.
- Limitations: Self-report ground truth prevents separating LLM over-reporting of work performance from human over-reporting, while the IWPQ reference inherits documented social-desirability bias.The study cannot quantify how much of the LLM–human gap on task performance reflects either source of bias.
- Limitations: Memorization is unlikely to drive the leaderboard, but paraphrased exposure to the publicly available instruments remains possible.The worst-case high-recall rate was 4.7%, 22/37 models had exactly zero recall, and recall-rate–PSS rank correlation was 0.00; the instruments are publicly available through original publications and thesis.
- Limitations: Headline estimates use n = 100 and R = 1, with a 0.046 PSS-unit bootstrap CI mean width; larger n and repeated administrations would better resolve within-top-5 differences.The proposed rerun at n = 263 and R = 3 would cost roughly 8× more, while reasoning effort was fixed at minimal and provider model versions may drift.
- Future work: The causal interpretation is limited to algorithmic asymmetry under controlled persona prompting, and documented failure modes may not generalize beyond organisational-psychology constructs.Future work should use novel instruments and non-Indo-European samples, extend to open-text responses, and test longitudinal drift and within-respondent demographic-counterfactual sensitivity.
8 Broader impacts
Using LLM-generated respondents as human substitutes risks inflated effects, stereotype amplification, and the entrenchment of synthetic biases in future training data. The framework directly surfaces these failure modes and supports comparable evaluation against human ground truth.
- Risks: Inflated Cronbach’s α and inter-item correlations can overestimate effects, producing under-powered human follow-up studies and over-confident published conclusions.These risks arise from over-coherence and range restriction relative to the human reference.
- Risks: Education-driven effects averaged |d| = 0.56, while gender effects were often near the resampling-noise floor despite being null in humans.Using synthetic respondents for fairness audits without counterfactual testing can certify biased systems as bias-free.
- Risks: Republishing synthetic respondents as human-equivalent data risks baking over-coherence, range restriction, and stereotype amplification into future LLM training corpora.The resulting contamination could make these failures unrecoverable from the LLM-versus-human gap.
- Mitigations enabled by our framework: The framework flags over-confident sample-driven components, stereotype amplification, training-data leakage, and homogeneity-by-averaging through complementary diagnostics.Statistical baselines, counterfactual swaps, memorisation probes, the inter-LLM PSS matrix, and ensemble results each target a distinct failure mode.
- Mitigations enabled by our framework: The released harness, dataset, configurations, and generation logs will enable benchmarking new LLMs against the same human ground truth without rerunning prior models.Deployment-specific protocols can then be evaluated using the same framework.
9 Conclusion … D Tucker’s φ permutation null distribution
The paper introduces an open Lithuanian organisational-psychology benchmark that evaluates LLM synthetic respondents through psychometric structure rather than single-item plausibility or aggregate effect sizes. It also provides reproducible tooling, validated English translations, and a permutation-based significance test for Tucker’s φ.
- 9 Conclusion: The benchmark evaluates 37 LLMs on psychometric structure and finds that they reproduce the qualitative direction of human psychometric relationships.The lineup spans major proprietary and open-weight families released between 2024-Q4 and mid-2026.
- 9 Conclusion: The de-identified dataset, configurations, evaluation harness, statistical baselines, and per-call generation logs will be released upon peer review publication.The paper identifies the referenced file paths as part of that forthcoming release.
- A Parser and validation pipeline: The pipeline parses, repairs, and stores structured LLM responses, resolving >95% of parsing failures in pilot runs after stricter re-issues.It stores both raw text and parsed integer dictionaries in generation-log parquet files.
- B Reproducibility, resumability, and incremental extension: The full workflow runs through four commands, supports offline mock-provider smoke tests, and records configurations in YAML files.The commands prepare data, generate synthetic responses, run analyses, and create figures.
- B Reproducibility, resumability, and incremental extension: Adding a model requires a one-line configuration edit, while rerunning analysis computes the PSS leaderboard without repeating prior models’ analysis.New calls are merged into the existing parquet, and previously completed model analyses are reused.
- C Instrument English translations: The English instrument translations yield an LT↔EN per-item Pearson correlation of 0.889 with mean |∆|=0.263 Likert points.The translations cover all 68 items, response anchors, and instructions and were validated against published English versions where available.
- D Tucker’s φ permutation null distribution: The Tucker’s φ procedure tests whether observed factor congruence exceeds alignment-search noise by permuting item labels in the LLM loading matrix.The permutation p-value uses (k + 1)/(K + 1), where k counts null draws meeting or exceeding the observed value.
- D Tucker’s φ permutation null distribution: Table 3 summarizes Tucker’s φ permutation tests by instrument, including mean null φ and the number of models exceeding the null’s 95th percentile.The full per-model, per-instrument results are provided in construct_tuckers_phi_permutation.csv.
E Power analysis
The power analysis indicates that the 37-model PSS leaderboard was well-powered to detect the observed gaps, while counterfactual power differed sharply across demographic axes. Education-swap effects were detectable, but the smaller gender-swap effect should be treated as directional rather than per-cell significant.
- Power analysis: d≥0.48 was detectable at n=100 for the 37-model PSS leaderboard’s 666 pairs under Holm–Bonferroni correction, below the observed PSS gaps.The calculation used α=0.05 and power 0.80.
- Power analysis: |d|=0.56 for education swaps was well-powered across the counterfactual contrast family of 3 axes × 12 subscales.The family used Holm–Bonferroni-corrected α for the corresponding comparisons.
- Power analysis: |d|=0.12 for gender swaps was underpowered and should be interpreted directionally rather than as a per-cell-significant finding.The counterfactual contrast family comprised 3 axes × 12 subscales.
F Counterfactual, ablation, cohort, inter-LLM tables and figures
Across counterfactual, structural, and downstream tests, LLM-generated survey data remains distinguishable from human data and fails to reproduce key psychometric and predictive properties. Statistical generators outperform LLMs as synthetic respondents, and increasing sample size or prompting does not close the gap.
- Counterfactual amplification: Education swaps dominate counterfactual amplification, role swaps are intermediate, and gender swaps are small and usually at or below the paired-d noise floor of approximately 0.09.The comparison uses mean absolute paired Cohen’s d across 12 composite subscales.
- Mediation and steerability: LLMs fabricate 3 of 10 placebo mediation paths, while explicit debiasing improves surface response style but leaves overall PSS flat because the correlation component does not move.The mediation result flags placebo paths whose LLM confidence intervals exclude zero although the human confidence intervals include zero; the steerability result comes from a six-model subset at n=100.
- Human-vs-synthetic distinguishability: Every LLM is separated near-perfectly from humans by discriminators, with median AUC 0.999, whereas Gaussian-copula and MVN baselines are statistically indistinguishable.The discriminator uses held-out human-versus-synthetic respondents, with chance at AUC 0.5.
- Sample-size scaling: The copula baseline’s PSS3 rises from 0.69 at n=10 to 0.82 at n=200, while the best LLM plateaus near 0.62 by n≈50, widening the fidelity gap from 0.11 to 0.19.Collecting more synthetic respondents therefore does not close the gap to a purely statistical generator.
- Latent structure: LLMs inflate general-factor ECV across instruments while collapsing specific-factor reliability, including Change-engagement ∆ωh = −0.94, indicating over-coherence.The bifactor comparison is between the LLM mean at C3 and humans.
G Robustness ablations: numerical detail · H Psychometrics primer for machine-learning readers
The robustness checks show that persona recall, response consistency, and PSS rankings are generally not artifacts of ordering or formatting, although some models fail through genuine response degeneracy or incomplete conditions. The psychometrics primer explains why preserving item dependencies, reliability, factor structure, mediation, discriminant validity, invariance, and general-factor structure matters beyond matching item means.
- G Robustness ablations: numerical detail: Inter-model PSS clustering forms within-vendor groups for Anthropic, Google, and OpenAI, with llama-4-maverick and grok-4-20 as outliers.Cohort-stratified PSS is also summarized using a model-by-stratum heatmap and each model’s worst-axis disparity.
- G Robustness ablations: numerical detail: All 37 models show negative within-respondent forward-versus-reverse correlations on reverse-keyed items, as expected from faithful responding.This result covers the reverse-keyed Dunham and change-engagement items.
- G Robustness ablations: numerical detail: grok-4-20 and llama-4-maverick rank low because of genuine response degeneracy, while missing C9 conditions reduce effective n and widen bootstrap CIs.grok-4-20 collapses toward a near-constant response vector with r≈0 on reliability; gpt-oss-120b, the grok-4-1-fast pair, and llama-4-scout lack C9.
- H Psychometrics primer for machine-learning readers: Psychometric usefulness depends on joint item dependencies, because composites, α and ω reliability, mediation paths, and demographic Cohen’s d values use correlations, covariances, or composite variation.Matching item means alone therefore does not ensure valid downstream analysis.
- H Psychometrics primer for machine-learning readers: Tucker’s φ compares aligned factor-loading vectors, but conventional φ thresholds can be inflated by chance in high-dimensional instruments; α and ω provide complementary reliability estimates.The primer notes fair factor match at φ ∈[0.85, 0.94] and identical structure at φ ≥0.95, while the LLM-versus-human reliability gap is essentially identical across α and ω.
- H Psychometrics primer for machine-learning readers: The primer defines mediation through direct and indirect paths, HTMT for discriminant validity, measurement invariance across groups, and bifactor ECV and ωh for general-factor structure.HTMT values < 0.85 are conventionally adequate; invariance proceeds through configural, metric, and scalar tests, while ECV > 0.7 is sometimes treated as evidence of essential unidimensionality.
I Compute and cost breakdown
The headline 37-model experiment used 18,500 calls and cost approximately $136, while the complete experimental program cost approximately $390 and took about 25 wall-clock hours. Analysis ran on a single CPU node, with Tucker permutation testing the heaviest operation.
- Headline grid: 18,500 calls across 37 models, five conditions, and 100 respondents consumed approximately 105 million tokens and cost approximately $136 in about six hours.The headline grid used one repeat per respondent and provider-specific concurrency limits.
- Test–retest stability: The test–retest stability run added approximately 20,400 calls, cost approximately $150, and required 6–8 hours, licensing the R=1 headline design.It comprised approximately 17,500 completed cells across the lineup and included retries.
- Total cost: Approximately $390 and 25 wall-clock hours covered the full experiments, dominated by the headline grid and test–retest stability run.The disk cache makes reruns of completed cells cost nothing, enabling inexpensive incremental lineup extensions.
- Analysis-side compute: The Tucker-permutation null was the heaviest analysis operation, performing 74,000 factor analyses in approximately 40 minutes on a single CPU node.Bootstrap confidence intervals were the second-heaviest analysis, taking approximately 15 minutes.
J Worked qualitative example: a single ATC item
A single ATC item illustrates that top-ranked models reproduce human mean and variance shifts across personas, whereas amplification and low-PSS models respectively exaggerate or ignore those differences.
- gpt-5.4-mini: gpt-5.4-mini captures both shifts: Persona A responses center on 4, while Persona B is more variable and lower on average.Persona A has 8 of 10 responses at 4 and two 5s; Persona B spans 2–5 with four 4s, three 3s, two 5s, and one 2.
- claude-sonnet-4-6: claude-sonnet-4-6 amplifies the persona gap, producing a mean difference of ∼2 Likert points versus ∼0.5 among humans.It gives Persona A mostly 5s and Persona B mostly 3s, creating a stereotype-amplification pattern tied to education, role, and sector.
- grok-4-1-fast-non-reasoning: grok-4-1-fast-non-reasoning gives nearly identical distributions to both personas, indicating weak conditioning on the persona block and explaining its low PSS.Both personas have a modal response of 3, despite substantially different profile information.
- Interpretation: The qualitative example matches the PSS ranking: strong models track both between-persona mean shifts and within-persona variance shifts, amplifiers exaggerate means, and low-PSS models track neither.This links the item-level behavior directly to the quantitative leaderboard patterns.
K Negative results
Four negative experiments show that common decoding, persona, reasoning-trace, and prompt-conditioning interventions do not reliably improve psychometric similarity. Richer personas reduce PSS, while debiasing and few-shot examples leave the overall gap largely unchanged.
- Negative result 1: temperature: PSS showed no temperature effect beyond sampling noise across six decoding temperatures tested on five models.The sweep used n=30 respondents per cell, with default-temperature controls for claude-opus-4-7 and the GPT-5 family.
- Negative result 2: free-text persona: Free-text personas reduced PSS by 0.04–0.07 units across a five-model subset rather than improving conditioning.The intervention used one-paragraph LLM-generated biographies derived from C3 profile fields.
- Negative result 3: chain-of-thought reasoning trace as data: Reasoning traces contained no measurable psychometric content and were dominated by surface-level persona rationalisations.The traces did not provide a measurable side-channel signal of psychometric reasoning.
- Negative result 4: prompt-based debiasing and few-shot conditioning: Debiasing and few-shot conditioning left overall PSS essentially flat, indicating that the gap is structural rather than a prompting artifact.Debiasing changed response style but not correlation, while three human answer vectors were almost entirely ignored in few-shot conditioning.
L PSS weight sensitivity analysis … Maintenance
PSS rankings remain stable across alternative weightings, while the Gaussian-copula baseline retains its advantage on sample-driven components. The paper documents ethical safeguards, dataset provenance and processing, intended uses, licensing, and maintenance procedures for the Lithuanian survey release.
- L PSS weight sensitivity analysis: The Gaussian-copula baseline beats every LLM on sample-driven components under every weighting, while the best LLM remains essentially tied with the copula on five-component PSS.Equal-weight rankings correlate 0.99 with the default ranking, and the copula’s PSS changes from 0.688 to 0.692.
- M Extended ethics statement: The human dataset was collected with informed consent covering academic re-use of de-identified aggregated data, and no new human data was collected for this benchmark.The original data collector oversees the de-identified record-level release.
- M Extended ethics statement: The release removes free-text and direct identifiers, uses coarse demographic variables with at least three respondents per non-empty joint cell, and prohibits re-identification under CC-BY-NC-4.0.The dataset contains no names, emails, phones, IP addresses, or exact employers; its eight demographic fields bound re-identification risk.
- M Extended ethics statement: LLM responses are synthetic and expose only de-identified demographic fields, change-context ordinal ratings, and no human-respondent free text.The benchmark’s estimated footprint is approximately $240, 18 hours, 50 kWh, and 24 kg CO2-equivalent.
- M Extended ethics statement: The benchmark may be misused to certify unfit synthetic-respondent pipelines, so released materials specify scope conditions for positive evidence and limitations.This addresses the bias-entrenchment harm identified in Section 8.
- Motivation; Composition: The Lithuanian dataset contains 263 fully completed employee responses to a 68-item questionnaire covering organisational change, work engagement, and self-rated performance.It is a March–April 2020 convenience sample with coarse demographics and 1.4% item-level missingness.
- Collection process; Pre-processing / cleaning / labeling: The survey was collected online with institutional supervision, informed consent, no compensation, and documented Likert recoding, field harmonisation, composite scoring, and reproducible processed files.The raw CSV is included verbatim alongside item-level, demographic, and composite-score data.
- Uses; Distribution; Maintenance; N Datasheet for the Lithuanian Organisational Psychology Survey (2020): The dataset supports instrument validation, response-style and cross-cultural studies, and survey-tool benchmarking; it is publicly redistributed under CC BY-NC 4.0 with versioned maintenance and repository-based errata.The raw and processed data follow the Gebru et al. datasheet schema and preserve the original thesis provenance.