Source-linked AI summary

Human Psychometric Questionnaires Mischaracterize LLM Behavior

Woojung Song, Dongmin Choi, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo

arXiv:2509.10078v4cs.CLcs.AI

TL;DR

The paper asks whether human psychometric questionnaires reliably characterize and predict LLM behavior in everyday interactions. It compares questionnaire self-reports with generation probabilities over validated responses to realistic queries across eight open-source models. The profiles diverge, supporting generation-based profiling as a more ecologically valid measure while exposing limits in questionnaire-based characterization.

  • Problem

    The paper examines whether human psychometric questionnaires reliably characterize and predict LLM behavior in everyday user interactions.

  • Method

    The study compares Likert questionnaire profiles with generation-probability profiles derived from construct-validated responses to realistic user queries across eight open-source LLMs.

  • Results

    The two profiling methods diverge, with questionnaire coherence attributed to textual transparency rather than stable underlying dispositions, while persona-induced questionnaire shifts fail to transfer to generation probabilities.

  • Takeaways & Limitations

    Questionnaire scores alone are insufficient evidence of model-level psychological characteristics, so generation-probability evaluation should supplement or replace them when measuring realistic behavior.

  • Takeaways & Limitations

    The generation-probability method currently requires token-level log-probabilities and therefore is limited to open-source models; the scenario pool and demographic reference also constrain generalizability.

Abstract

from arXiv · show

We examine whether human psychometric questionnaires can serve as reliable tools for characterizing and predicting LLM behavior in everyday user interactions. We analyze eight open-source LLMs by comparing their value and personality profiles derived from two different methods: Likert self-reports on established questionnaires (PVQ-40/21 and BFI-44/10) and generation probabilities over value-laden responses to everyday user queries. The two profiles diverge substantially. Within-construct item consistency, often cited as evidence of stable LLM dispositions, disappears in generation probabilities. We attribute this gap to the fact that explicit lexical cues in established questionnaire items allow models to recognize the target construct and respond in alignment-consistent, socially desirable ways, whereas realistic user queries provide no such cues. In addition, demographic persona prompts shift models' responses to human questionnaires in ways consistent with real human patterns, but no such shifts appear in the generation probabilities of responses to realistic user queries, showing their limited ability to simulate the behaviors of target demographics in real-world user interactions. Overall, our study shows that human psychometric questionnaires are insufficient tools for predicting LLM behavior and suggests generation-based profiling as a more accurate measure.

1 Introduction

The study tests whether human psychometric questionnaires can reliably characterize and predict LLM behavior in everyday interactions. It compares questionnaire self-reports with generation probabilities and finds substantial divergence, attributing the gap to textual transparency in questionnaire items.

  • The study targets behavioral predictability and safety as LLMs enter high-risk applications such as emotional support, ethical advice, and children’s chatbots.
  • It compares established questionnaire scores from Likert responses with generation-probability scores for eight open-source LLMs.The questionnaires include PVQ-40, PVQ-21, BFI-44, and BFI-10.
  • Generation profiles use log-probabilities assigned to construct-validated responses drawn from realistic user queries, rather than annotated open-ended generations.The design aims to preserve generative behavior while avoiding errors from annotating unconstrained responses.
  • The two methods produce different construct-level profiles, and intra-construct item consistency disappears in generation probabilities.The study organizes these findings around RQ1 and RQ2.
  • Explicit construct cues in questionnaire wording let models recognize targets and produce alignment-consistent, socially desirable responses, whereas realistic scenarios provide no such cues.
  • Persona prompts shift questionnaire responses toward human demographic patterns, but corresponding shifts do not transfer to generation probabilities for realistic user interactions.The authors interpret this as evidence of superficial demographic simulation rather than reliable real-world behavioral simulation.
  • The findings indicate that established psychometric questionnaires are insufficient for predicting LLM behavior in realistic user interactions and motivate generation-based profiling.

2 Related Works

Prior work applies psychometric questionnaires to LLMs and increasingly tests whether questionnaire-derived profiles generalize beyond self-reports. This study extends that work with generation probabilities over psychometrically validated, realistic scenarios and links questionnaire coherence to item transparency.

  • Earlier studies applied BFI, IPIP-NEO, and PVQ questionnaires to assess LLM values and personality, reporting reliability, differentiation, or construct-related validity.
  • Subsequent work examined external validity through scenario actions, situational judgments, behavioral tasks, value-informed actions, opinions, downstream tasks, and perceived traits.These studies each documented a gap between questionnaire responses and behavior.
  • The paper complements prior studies by using generation probabilities as a controlled behavioral proxy over psychometrically validated items while preserving the target questionnaires’ construct space.
  • Probability-based measurements read the model’s scoring of candidate texts from its generation distribution and avoid relying on accurate self-judgments about its own dispositions.The cited background contrasts this with self-reports that are sensitive to prompt design.

3 Profiling Methods and Models

The profiling framework compares Likert questionnaire scores with generation probabilities assigned to psychometrically validated responses in realistic scenarios. It uses Value Portrait data, scenario-level macro-averaging, natural prompts, and eight open-source models.

  • The study compares established questionnaire scores with generation-probability scores for Big Five traits and ten Schwartz basic values.
  • It administers PVQ-40, PVQ-21, BFI-44, and BFI-10 using gender-neutral item wording and Likert-scale responses.PVQ responses use a 1–6 scale, while BFI responses use a 1–5 scale.
  • For ecological validity, behavior is represented by everyday-user generations and measured generatively rather than through reflective self-reports.The method uses plausible responses to real-world queries and avoids error-prone annotation of open-ended model outputs.
  • The Value Portrait dataset contains 520 psychometrically validated scenario-response pairs from 104 real-world queries spanning everyday requests.
  • Each construct receives a score based on the mean across scenarios of the within-scenario mean log-probability of its tagged responses.The scenario-then-construct macro-average gives every scenario equal weight.
  • The framework uses open-ended prompts: conversational sources are framed as user messages, while advisory sources use titles and situational descriptions.Both templates ask models to respond naturally without Likert-scale constraints.
  • The method restricts generation to multiple plausible candidate responses representing different constructs, balancing generative behavior with psychometric validity.
  • The analysis covers eight open-source models across four families, with smaller and larger variants.

4 Research Questions

The study tests whether established psychometric self-reports reliably characterize realistic LLM behavior by comparing them with generation-probability profiles across eight models. It finds substantial divergence in rankings, item-level structure, construct recognition, and demographic persona shifts, attributing these patterns to textual transparency in questionnaire items.

  • Research design: The study compares established questionnaire profiles with generation-probability profiles derived from realistic user-query responses across eight LLMs.Established profiles use Likert self-reports, while generation profiles use probabilities assigned to construct-annotated responses.
  • RQ1: Profile agreement: Within-method Spearman ρ averages 0.74 for PVQ-40↔PVQ-21 and 0.77 for BFI-44↔BFI-10, whereas cross-method agreement falls to 0.31 and 0.28 for Schwartz values.For Big Five traits, cross-method ρ is 0.26 for Gen↔BFI44 and 0.11 for Gen↔BFI10, with several negative correlations.
  • RQ2: Item-level consistency: Established questionnaires show construct structure, with η2 averaging 0.526 for PVQ-40 and 0.492 for BFI-44, while generation probabilities are indistinguishable from permutation baselines.Established WMV averages .603 and .592; generation-probability WMV remains near 1.0, indicating little within-construct clustering.
  • RQ3: Textual transparency: LLMs recognize established questionnaire constructs substantially better than Value Portrait scenarios: mean F1 ranges from .69 to .83 versus .05–.11.Sentence-embedding results similarly show established-item discrimination of 0.13–0.22 and clustering gaps of 0.07–0.15, versus near-zero measures for Value Portrait scenarios.
  • RQ3: Textual transparency: Established questionnaire responses favor prosocial constructs, whereas realistic scenarios lack the explicit cues that let models map questionnaire items to target constructs.The study links this transparency to aligned, socially desirable responding and to the absence of stable-disposition evidence in generation probabilities.
  • RQ4: Demographic simulation: Persona-induced shifts match human demographic patterns under established questionnaires but not under generation probabilities.Established PVQ-40 direction matches are 62/80, while generation-probability direction matches are 40/80 and indistinguishable from chance; VP mean cosine is −0.03.

5 Conclusion

The study finds that questionnaire-derived psychological profiles do not align with generation-probability profiles across eight open-source LLMs. Questionnaire responses also produce stereotype-consistent persona shifts, unlike generation-probability responses, which show directionally incoherent shifts.

  • Questionnaire-based profiles do not align with profiles derived from generation probabilities across eight open-source LLMs.
  • The apparent construct structure in questionnaire profiles largely reflects textual transparency in item wording rather than stable underlying dispositions.
  • Demographic personas shift questionnaire responses in stereotype-consistent ways but produce generation-probability shifts that do not track human demographic patterns.
  • Questionnaire scores alone are insufficient evidence of model-level psychological characteristics.
  • The study recommends supplementing or replacing established questionnaires with generation-probability evaluation when characterizing model behavior rather than self-report responses.

Limitations

The study’s generation-probability framework is currently limited to open-source models and to the constructs and human references covered by its datasets and comparisons.

  • Token-level log-probability access limits the current generation-probability analysis to open-source models.
  • The Value Portrait dataset covers Schwartz values and Big Five traits but would benefit from broader constructs and more diverse situational contexts.
  • The persona comparison uses European Social Survey Schwartz-value data and lacks a parallel human reference for Big Five traits.
  • The study lacks a matched human baseline for RQ1–RQ3, preventing assessment of whether the observed gaps are specific to LLMs or comparable to human self-report–behavior gaps.

B Value Portrait Dataset Details

The Value Portrait dataset combines everyday user queries and interpersonal dilemmas with psychometrically validated candidate responses annotated for Schwartz values and Big Five traits.

  • Queries come from human–LLM conversations and human–human advisory archives, covering everyday requests and value-laden interpersonal dilemmas.
  • The 104 source queries span ShareGPT, LMSYS, Reddit, and Dear Abby archives.
  • Five candidate responses were generated and annotated for Schwartz values and Big Five personality traits, retaining responses that met annotation-quality thresholds.
  • Validation used 681 participants who rated response similarity and completed PVQ-21 and BFI-10 measures.

C Inference Configuration

The study uses distinct inference procedures for Likert questionnaires and Value Portrait scoring, with fixed prompt templates, model-specific serving settings, and direct response log-probabilities.

  • The analysis runs open-weight models locally on 4× NVIDIA A100 80 GB GPUs, using bfloat16 except for the FP8-quantized Qwen3-235B-A22B checkpoint.
  • Established survey prompts use deterministic chat-completion inference with temperature = 0.0 and max_tokens = 1024.
  • Value Portrait scoring uses echoed prompt–response sequences and extracts log-probabilities only for response tokens.
  • All models use default HuggingFace chat templates, while GPT-OSS reasoning fields are excluded and Qwen3 reasoning traces are stripped before parsing.
  • Established questionnaires use two prompt variants with reversed high-to-low and low-to-high Likert option orderings.
  • Value Portrait scoring computes each candidate response’s log-probability directly without showing candidate responses to the model.
  • Construct-recognition prompts pair each questionnaire item with a candidate construct and definition, asking the model whether the item measures that construct.

E Sampling Log-Probability Validity

The study validates sampling-based generation probabilities against models’ free-form responses, then defines a scenario-balanced aggregation for construct profiles. VP candidates generally fall within each model’s generation distribution, and the main profiling gap is robust to aggregation choice.

  • Sampling validity: 15 responses per scenario combine 10 temperature-1 sampled completions with 5 VP candidates for rank-based sampling validation.The procedure uses the same scenario prompt and echo-based log-probability evaluation, after filtering empty or reasoning-only completions.
  • Sampling validity: At the median, the highest-ranked VP candidate placed fourth among 15 responses, with mean rank 5.7, approximately the 69th percentile.This indicates that VP candidates generally occupy a reasonable region of each model’s generation distribution, though placement varies by model.
  • Sampling validity: VP candidates ranked near the top for GPT-OSS, Qwen2.5-7B, and Qwen3-30B, but in the middle-to-lower region for Gemma-3, Qwen2.5-72B, and Qwen3-235B.The reported medians are 1 for the first group and 9–11 for the second; the authors suggest temperature-1 sampling may be more diffuse in the first group.
  • Aggregation: The macro-average gives every scenario equal weight, unlike the flat micro-average, which lets scenarios with more co-tagged candidates dominate construct scores.Macro-averaging also parallels equal per-item weighting in established questionnaire profiles.
  • Aggregation: Macro- and micro-aggregated Schwartz profiles correlate strongly across models, with Spearman ρ ranging from 0.79 to 1.00 and mean ρ = 0.92.The headline within-versus-cross-method gap remains robust under both aggregation choices, although individual cells can shift by approximately 0.31.
  • Measures: Spearman’s ρ measures monotonic agreement between construct rankings, while NDCG emphasizes recovery of highly ranked constructs.Because construct counts are small, significance is tested on aggregate within-versus-cross-method differences across models rather than on individual profiles.

F.3 Statistical Testing: Within-Method vs. Cross-Method Agreement

Across eight models, established questionnaires agree more with one another than generation-probability profiles agree with established questionnaires. Exact sign-flip tests support this within-versus-cross-method gap for both construct families and overall.

  • Setup: Within-method agreement systematically exceeds cross-method agreement between generation probabilities and established questionnaires.The paired difference is defined as established within-method correlation minus the average generation-probability-established correlation.
  • Inference: The primary inferential statistic is the exact sign-flip permutation p-value; the percentile bootstrap confidence interval is treated as descriptive because n = 8 per construct group.The bootstrap procedure uses 10,000 resamples, but its percentile interval can be somewhat anticonservative with this sample size.
  • Results: ∆ = 0.446 for Schwartz values, 0.584 for Big Five, and 0.515 overall in Spearman ρ, with p = 0.004, 0.016, and < 0.001, respectively.These are one-sided sign-flip permutation results comparing within-method and cross-method agreement.
  • Results: ∆ = 0.132 overall for NDCG, with p = 0.002, and both construct groups individually significant.NDCG shows the same directional pattern as Spearman ρ while applying top-weighted ranking agreement.
  • Per-model pattern: All eight models show positive within-minus-cross differences for Schwartz values, while seven of eight do so for Big Five.The Qwen3-235B BFI exception has an unusually low within-method baseline of ρ = 0.205.
  • Score construction: Generation-probability scores are mean total log-probabilities for construct-tagged VP responses, whereas established scores are prompt-averaged Likert means.Absolute generation-probability scores are not comparable across models; within-model relative ordering is the meaningful quantity for rank metrics.

G.2 Per-Construct Within-Variance

The per-construct analysis contrasts established questionnaire structure with VP generation-probability profiles and item-recognition behavior. Established items show construct-specific clustering and recognition, whereas VP items show weak construct-identifying structure.

  • Within-construct variance: Values below 1.0 indicate that items within a construct cluster more tightly than expected under random assignment.The tables report per-construct within-variance and aggregate metrics for BFI-44, PVQ-40, and VP profiles.
  • Item–construct recognition: Established BFI items remain highly recognizable: even Gemma3-4B’s weakest BFI-44 construct pairs reach F1 = .53.GPT-OSS-120B and Qwen3-235B achieve near-perfect recognition across all five traits, while Gemma3-4B struggles most with Neuroticism and Agreeableness.
  • Item–construct recognition: VP construct recognition is consistently weak: Agreeableness ranges from F1 .09 to .27 and Stimulation from .07 to .25 across models.Both constructs exceed F1 .10 for six of seven models but fall just below .10 on GPT-OSS-120B.
  • Textual transparency: Established items achieve 77–81% Top-1 accuracy and discrimination of 0.13–0.22, while VP items are near chance with near-zero discrimination.The embedding analysis interprets established item text as revealing construct identity, whereas VP scenario text carries no comparable lexical cue.
  • Textual transparency: Established questionnaires have clustering gaps of 0.07–0.15, whereas VP items range from −0.001 to +0.004.Thus, VP scenarios tagged with the same construct are no more textually similar than scenarios tagged with different constructs.
  • Textual transparency: Within-construct similarity is strongest for Stimulation and Hedonism in PVQ-40 and Agreeableness in BFI-44, with PVQ-40 Security a borderline below-baseline case.PVQ-40 Security is 0.426 against a 0.432 inter-construct baseline; Tradition is borderline above it at 0.438.

H.2.4 Encoder Robustness Check

The encoder robustness check repeats item–definition and within-construct similarity analyses across five sentence-encoder families. Established questionnaires consistently show stronger textual construct cues than VP items regardless of encoder choice.

  • Robustness design: Five encoders spanning MP-Net, MiniLM, BGE, and E5 reproduce the established-questionnaire advantage over VP items.The analysis uses MPNet-base, two MiniLM variants, BGE-base, and E5-base.
  • Robustness result: Across all five encoders, established questionnaires have higher item–definition discrimination and within-construct clustering gaps than VP outputs.The robustness check tests both textual transparency measures rather than relying on one sentence encoder.
  • Measures: Item–definition discrimination measures similarity to the correct construct definition relative to other definitions, while clustering gap compares intra- with inter-construct similarity.These measures operationalize whether item wording identifies its construct and whether same-construct items share surface-level text.
  • Persona setup: The persona conditions prepend demographic system prompts before either Likert ratings or VP generation-probability evaluation, with vanilla prompts as the baseline.This setup applies the same persona manipulation to both measurement approaches.

I.2 Methodology Details

The study compares demographic shifts in human ESS profiles with LLM persona-induced shifts using centered value profiles, directional agreement, cosine similarity, and normalized magnitudes. Likert questionnaire shifts align with human references, whereas VP generation-probability shifts do not.

  • Data and setup: The ESS reference uses Schwartz values from 29 European countries and Israel, with N = 37,398 complete respondents.The ESS does not include a Big Five instrument, so this analysis is restricted to value dimensions.
  • Profile construction: Centered profiles subtract each respondent’s or model run’s grand mean across ten values, making each profile sum to zero and reducing scale-use bias.LLM deltas compare persona profiles with vanilla profiles; human deltas use analogous ESS subgroup differences.
  • Statistical tests: Direction match counts value dimensions where human and LLM deltas share an algebraic sign, with chance modeled as Binomial(10, 0.5) and tested above 50%.The aggregate test covers 80 dimensions across eight conditions at α = 0.05.
  • Results: 77.5% PVQ-40 and 68.8% PVQ-21 direction agreement exceed chance, while VP achieves 50.0% with p = 5.44e −01 and no significant agreement.PVQ-40 has p = 4.07e −07 and PVQ-21 has p = 5.26e −04.
  • Results: PVQ-40 has positive human-LLM cosine similarity in all eight conditions and PVQ-21 in seven, whereas VP has four positive and four negative conditions.Across models, VP-human cosine is near zero or negative for most models, so the result is not driven by one outlier.
Loading 2509.10078v4…