Source-linked AI summary

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez

arXiv:2606.12730v1cs.AIcs.CLcs.CYcs.LG

TL;DR

Low-cost self-reports are useful for anticipating LLM behavior only if they predict downstream actions. Across four tasks and 11 models, the study finds coherence is selective: TPB reaches human-level within-session coherence, but cross-session alignment depends on behavior and context.

  • Problem

    Whether low-cost psychometric self-reports reliably predict LLM behavior remains unresolved, limiting their use for scalable behavioral auditing.

  • Method

    The study compares Big Five and TPB self-reports across four behavioral tasks and 11 frontier LLMs while varying session context and identity induction.

  • Results

    TPB reaches human-level within-session coherence, whereas cross-session coherence survives for some behaviors, collapses for sycophancy, and persona prompting does not align behavior.

  • Takeaways & Limitations

    Task- and behavior-specific instruments are more informative than coarse traits, but their predictive coherence must be evaluated across behaviors and contexts.

  • Takeaways & Limitations

    Shared-session self-report–behavior coherence is correlational rather than causal because behavioral priming and self-report compliance can shift both responses.

Abstract

from arXiv · show

Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR-behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly, even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR-behavior coherence exists but is selective. 1) Within a shared conversation, the Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations, but does not bring behavior into alignment. These findings suggest that coarse personality frameworks, such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.

1 Introduction

This study asks when low-cost psychometric self-reports predict LLM behavior, addressing whether prior dissociations reflect the instrument, probing context, or the models themselves. Across four behavioral tasks and 11 LLMs, it tests framework granularity, session context, and identity induction, finding that coherence exists but is selective.

  • Motivation: Self-reports are attractive because they are cheap, theoretically grounded, and widely used, but they are useful only if they reliably predict downstream behavior.This matters for anticipating behavioral tendencies in high-stakes deployments such as clinical decision support, financial advising, and educational tutoring.
  • Motivation: Prior work found systematic self-report–behavior dissociation in LLMs, while leaving unclear whether the gap reflects the instrument, probing context, or model properties.Models can produce psychometrically coherent personality profiles that fail to predict choices in behavioral tasks.
  • Study Design: Big Five traits are poor predictors of specific behaviors even in humans, with trait–behavior Pearson correlations rarely exceeding r ≈.20.The study therefore contrasts Big Five with TPB, a finer-grained instrument whose human predictive validity is reported as r ≈.47.
  • Study Design: The experiments use a 2 × 2 × 2 factorial design varying framework, session context, and identity induction across 4 behavioral tasks and 11 LLMs.The framework comparison is Big Five versus TPB; context is shared versus separate sessions; identity induction is parameter perturbation versus persona prompting.
  • Key Findings: The study proposes that coherence can arise from stable model-state coupling or from within-session context priming, motivating task- and behavior-specific instruments.It also provides a practical mapping of when self-reports are and are not useful for predicting LLM behavior.
  • Key Findings: Within-session TPB probing reaches the human meta-analytic baseline of mean r = +0.40, whereas Big Five is not predictive.The introduction characterizes this coherence as selective and notes that within-session probes can themselves shift self-report or behavior toward a policy being evaluated.

2 RQ1. Shared-Context: Does Self-Report Predict Behavior at All?

Under shared-session conditions, TPB self-reports predict LLM behavior with overall coherence comparable to human intention–behavior baselines. Coherence is task-specific, strongest for volitional behaviors and theoretically inverted for implicit bias.

  • Experimental design: The shared-session design placed self-report and behavior in one thread without resets or explicit consistency instructions, testing whether stated dispositions spontaneously predicted behavior.Primary analyses used parameter-grid induction with 54 matched conditions per model.
  • Overall coherence: TPB self-report predicted behavior overall at r = +0.25, rising to r = +0.40 when the theoretically dissociated IAT task was excluded.The overall estimate falls within the human meta-analytic range of r ≈0.25–0.50.
  • Task-specific coherence: Honesty × Attitude showed r = +0.67, Sycophancy × Intention r = +0.47, and CCT × Intention r = +0.22.All three volitional tasks showed substantial positive within-model coherence.
  • Task-specific coherence: IAT showed the theoretically expected explicit–implicit dissociation, with r = −0.59, consistent with compensatory-effort inversions documented in humans.Models reporting higher intention to categorize without bias subsequently produced more stereotype-consistent responses.
  • Model heterogeneity: Coherence varied substantially across models, from Claude 4.5 Haiku at r = +0.75 to Claude 3.7 Sonnet at r = −0.53.Qwen 235B followed at r = +0.72, while Phi-4 also fell below zero at r = −0.11.

3 RQ2. Framework Specificity: Does TPB Granularity Outperform Big Five?

Under identical within-session shared-context conditions, TPB’s task-specific granularity outperforms the context-independent Big Five for predicting LLM behavior. TPB exceeds Big Five across models, while Big Five detects little reliable, theoretically aligned coherence.

  • Experimental comparison: RQ2 compares TPB and Big Five under identical within-session, shared-context conditions, changing only the self-report instrument.The experiment adds 88 Big Five cells to the 77-cell TPB matrix and correlates each trait with the task’s primary behavioral outcome.
  • Model-level results: TPB exceeds Big Five in 8/11 models, with mean raligned of +0.21 versus +0.01, a 0.20 Fisher-z effect-size gap.The three exceptions—Phi-4, Qwen 72B, and Claude 3.7 Sonnet—have TPB aggregates that are negative or near zero.
  • Model-level results: The TPB advantage is largest for Claude 4.5 Haiku (∆= +0.74), Qwen 235B (+0.64), and LLaMA 4 Maverick (+0.48).These are the strongest model-level gaps reported in the comparison.
  • Big Five results: Only 3 of 88 Big Five cells reach p < .05, and just 1 is theoretically aligned; the other 2 are significant inversions.Nearly all Big Five heatmap cells are near zero, and CCT-Neuroticism illustrates unsupported directionality with raligned = +0.02, [−0.10, +0.15].

4 RQ3. Context Separation: Does Coherence Survive Session Separation?

Across separate API sessions sharing initialization but not response context, TPB coherence survived selectively: it largely persisted for honesty, collapsed for sycophancy, and remained positive in only two of 11 models. High cross-session self-report consistency alongside task-linked behavioral consistency indicates that collapse was driven by behavior, not self-report drift.

  • Experimental setup: RQ3 separated self-report and behavior into independent API calls while matching temperature, seed, and system prompt, estimating ∆r = r_same − r_separate.This tests whether coherence depends on response context or persists under shared initialization context alone.
  • Task-dependent coherence: Honesty showed partial survival: Attitude × align changed from r = +0.67 to +0.53, with ∆r = +0.14, 95% CI [+0.04, +0.25], p < .001.Most same-session coherence persisted after session separation.
  • Task-dependent coherence: Sycophancy completely collapsed: Intention × align changed from r = +0.47 to −0.07, with ∆r = +0.54, 95% CI [+0.39, +0.69], p < .001.The separate-sessions correlation was indistinguishable from zero.
  • Model-level survival: Only two of 11 models retained significantly positive separate-sessions coherence: Claude 4.5 Haiku and LLaMA 3.3 70B.Claude 4.5 Haiku changed from r = +0.75 to +0.65∗∗∗, while LLaMA 3.3 70B changed from r = +0.38 to +0.66∗∗∗; Qwen 235B showed the largest collapse, from +0.72 to −0.14.
  • Mechanism: Cross-session self-report consistency was high across all four tasks, while behavior consistency tracked survival, including IAT +0.98 and Honesty +0.45.Fisher-z aggregated self-report consistency was Honesty +0.81, Sycophancy +0.59, CCT +0.23, and IAT +0.52.

5 RQ4. Identity Induction: Does Persona Grounding Rescue Coherence?

Persona grounding changes and stabilizes LLM self-reports across separate sessions but does not rescue their coherence with behavior. This failure is consistent across models, while CCT and IAT coherence remain induction-invariant.

  • Experimental setup: The experiment compares parameter-grid and persona inductions across separate sessions, using persona descriptions with temperature fixed at 0.2.Persona prompting uses 30 diverse PersonaHub character descriptions and 60 conditions per model, versus 54 grid conditions per model.
  • Per-task TPB: CCT and IAT coherence are induction-invariant, with both ∆rinduction confidence intervals containing zero.For IAT, the inversion persists from r = −0.66 to −0.63, with ∆r = +0.02 ns.
  • Per-model TPB: Zero models satisfy the rescue criterion under persona induction.The two models retaining grid coherence—Claude 4.5 Haiku and LLaMA 3.3 70B—also exclude zero under personas, but coherence attenuates by ∆r = −0.25∗∗∗ and −0.15∗∗, respectively.
  • RQ4 — Identity Induction: Persona induction produces more diverse and stable self-report profiles, but separate-sessions SR–behavior coherence is not rescued.SR diversity is positive in 50% of model×task cells and substantial in 11/44; SR stability is positive in 75% of cells, with mean ∆r = +0.14.

6 Discussion

LLM self-reports predict behavior selectively: fine-grained, same-session measures can reach human-level intention–behavior coherence, whereas Big 5 performs much worse. Context separation weakens coherence for context-loaded behaviors, and persona induction improves report consistency without aligning behavior.

  • Core findings: Fine-grained, same-session self-reports reach the human meta-analytic baseline for intention–behavior correspondence, while Big 5 recovers roughly an order of magnitude less coherence.The comparison is made under identical conditions; Big 5 remains the dominant framework in LLM personality research.
  • Core findings: Context separation attenuates self-report–behavior coherence, and alternative explanations fail to account for the task-discrete pattern.The discussed alternatives include weak parameter-grid anchoring, inference-level nondeterminism, and generic test-retest decay.
  • Limitations: Same-session coupling is the most permissive setting for coherence but confounds causal interpretation through behavioral priming and self-report compliance.Evaluative prompts can shift both self-reports and behavior toward in-context framing.
  • Implications: Behavioral prediction should favor fine-grained, TACT-anchored instruments over Big 5 because broad traits poorly predict specific task choices in humans.The target behavior should be known in advance when selecting the instrument.
  • Implications: Persona induction improves self-report consistency but not behavioral coupling, so persona-customized deployments may yield distinct reports without distinct behavior.For context-loaded tasks, safety probes should elicit self-reports and target behavior in separate sessions to avoid conflating priming with disposition.

7 Conclusion

The study finds that self-reports and behavior share an upstream generative basis, but task structure determines whether their coherence persists across contexts. Fine-grained instruments achieve human-level within-session coherence, whereas coarse trait inventories do not.

  • Self-report and behavior are jointly produced by shared upstream model state, creating a common-cause generative structure.
  • Task structure determines which self-report–behavior couplings persist when conversational context is separated.
  • Fine-grained instruments yield within-session coherence at the human meta-analytic baseline, while coarse trait inventories do not.

A Further Discussion · B Broader Impact

The discussion shows that cross-session self-report–behavior coherence depends on task and context: persona induction improves self-report fidelity without restoring behavioral coupling, while implicit-association behavior persists and sycophancy reverses across sessions. The broader impact is methodological and cautionary, supporting behavior-specific auditing while warning that self-report probes can create false assurances outside validated conditions.

  • A Further Discussion: Persona-induced self-reports show high cross-session fidelity, yet behavior coupling still does not recover.This argues against weak parameter-grid anchoring as the sole explanation for cross-session collapse.
  • A Further Discussion: Matched persona descriptors can recover stable dispositional components such as IAT and Honesty, but cannot distinguish surviving from collapsing tasks after shared context is removed.Prior studies’ high self-report–behavior correlations may therefore be genuine for some properties without generalizing across tasks.
  • A Further Discussion: IAT behavior consistency is near-perfect for every sampled model, with r ∈[+0.90, +1.00], while within-model explicit–implicit inversion is r = −0.59.The pattern is consistent with safety training shaping explicit reporting more strongly than implicit associations, especially in heavily aligned closed Claude models.
  • A Further Discussion: Sycophancy is heavily context-primed, showing strongly negative cross-session consistency for Qwen 235B (−0.96), Gemini 2.5 Flash (−0.90), and DeepSeek V3.1 (−0.89).Deferral to the confederate flips when the self-report is removed from context, contrasting with context-independent IAT behavior.
  • A Further Discussion: The claims concern input/output self-report–behavior coherence, not consciousness or moral status, and do not imply that self-reports are useless.Between-model self-report variance correlates with between-model behavior variance, while persona prompting can alter surface text generation for legitimate UX or domain-adaptation uses.
  • B Broader Impact: Behavior-specific instruments grounded in the Theory of Planned Behavior substantially improve predictive validity over broad inventories such as the Big Five under appropriate conditions.This is presented as the work’s primary methodological contribution to pre-deployment behavioral auditing.
  • B Broader Impact: Self-report probes can create false assurances when used without same-session probing, behavior-specific instrumentation, and tasks whose behavioral basis is context-independent.The risk is overreliance on benchmarking utility, particularly for deployment-relevant concerns such as sycophancy.

C Limitations

The findings are limited by narrow task coverage, unresolved interpretation of shared state and priming, uncertain human-to-LLM construct validity, model-version dependence, and single-turn interaction design. These constraints limit generalizability to broader behaviors, future checkpoints, and naturalistic deployment contexts.

  • Task coverage: The battery covers only four tasks, so extended agentic behavior, numerical reasoning, and real-world tool use may show unobserved SR–behavior coupling patterns.The study notes that the task-discrete survival pattern may not generalize beyond risk-taking, sycophancy, honesty, and implicit bias.
  • Interpretation: The same-session/separate-session comparison cannot fully distinguish shared model state from in-context priming, especially for sycophancy, without internal representations.Mechanistic interpretability could clarify where SR–behavior coupling arises or breaks down.
  • Construct validity: The validity of translating human TPB and BFI-44 constructs to LLMs remains unresolved, despite internal-reliability evidence from Cronbach’s α.TPB’s TACT anchoring may invoke different generative pathways in LLMs than the motivational states targeted in humans.
  • Model generations: The experiments used a fixed set of 11 models, so reported self-report and behavioral tendencies may change across versions, post-training updates, or fine-tuning.Observed safety-training signatures, such as Claude’s IAT inversion gradient, may be checkpoint-specific.
  • Interaction structure: Single-turn self-report and behavioral sessions may not extrapolate to deployment, where multi-turn conversation, tool use, and retrieved context can modulate priming effects.The generalization of session-structure findings to naturalistic interaction remains open.

D Behavioral Tasks: Selection, Operationalisation, and Construct Mappings … E.2 TPB self-report fingerprints

The paper evaluates four psychologically distinct behavioral tasks across 11 instruction-tuned LLMs, mapping each task to specific TPB domains or to implicit associations outside volitional control. Per-model fingerprints show stable Big Five profiles but differentiated TPB patterns, including low Subjective Norm and greater variability under persona induction.

  • D Behavioral Tasks: Selection, Operationalisation, and Construct Mappings: Four validated human-paradigm tasks cover TPB’s Attitude, Subjective Norm, and Perceived Behavioral Control, plus an implicit-bias test outside volitional scope.The task set also supports directional comparison with partial Big Five–behavior links reported in humans.
  • D Behavioral Tasks: Selection, Operationalisation, and Construct Mappings: Risk-taking uses the Columbia Card Task, with mean cards flipped as the outcome, and operationalizes the TPB Attitude domain.The task is described as having the broadest expected Big Five coverage in humans.
  • D Behavioral Tasks: Selection, Operationalisation, and Construct Mappings: Sycophancy uses an Asch-style conformity task, measuring flip rate after exposure to a conflicting confederate opinion, and primarily targets Subjective Norm.The construct indexes sensitivity to perceived social pressure, with Agreeableness and Extraversion linked to conformity in humans.
  • D Behavioral Tasks: Selection, Operationalisation, and Construct Mappings: Honesty measures factual calibration with Brier score and ∆confidence, while the IAT measures stereotype associations across six domains using a d-score per test.Honesty is mapped to Perceived Behavioral Control, whereas implicit bias is explicitly outside TPB’s volitional scope.
  • D.1 Models: The study evaluates 11 instruction-tuned LLMs spanning proprietary and open-weight families, model scales, training pipelines, and providers.The evaluation applies both parameter perturbation across 27 conditions and persona prompting across 30 conditions.
  • E Per-Model Fingerprints: Per-model fingerprints map behavioral metrics to a common 1–5 normalized scale, while self-report figures retain underlying Likert means and shaded regions show ±1 standard deviation.The variability reflects conditions within each model–dimension cell rather than uncertainty in the mean estimate.
  • E.1 Big Five self-report fingerprints: Big Five trait profiles are highly stable across parameter-grid and persona inductions within sessions, so prediction failure is not attributed to unstable self-report scales.Persona and parameter-grid profiles track closely for most models, consistent with internal reliability of the Big Five responses.
  • E.2 TPB self-report fingerprints: Across model–task cells, Subjective Norm is usually the lowest TPB construct, Intention and PBC cluster tightly, and persona induction increases per-condition variability.The low Subjective Norm pattern parallels a lower SN–Intention correlation reported elsewhere in the appendix.

E.3 Behavioral fingerprints … F.3 TPB construct structure: do Attitude, Subjective Norm, and PBC predict Intention?

The paper finds that behavioral coherence is selective: sycophancy depends strongly on shared context, while other behaviors remain more stable, and TPB self-reports are reliable and structurally predictive whereas Big Five reports are reliable but behaviorally non-diagnostic. TPB’s Attitude, PBC, and Subjective Norm all predict Intention across session and induction conditions, though reliability and construct validity vary by task and model.

  • E.3 Behavioral fingerprints: Sycophancy changes sharply across session types, with same-session self-reports suppressing deferral while separate-session behavior often restores high deferral.Same-session probing produces mapped scores near 1.0 for several models, whereas separate sessions show high deferral rates when the self-report is no longer visible.
  • E.3 Behavioral fingerprints: Risk Taking, Stereotyping, Epistemic Honesty, and Self-Reflective Honesty remain broadly comparable across sessions, with Stereotyping near ceiling for most models.These stable behavioral levels align with preserved self-report–behavior correlations for IAT and Honesty and a small Δr for CCT.
  • F Construct Validity: Internal Reliability of TPB and Big Five Instruments: Internal-reliability analyses test whether Big Five prediction failures reflect noisy scales rather than a genuine mismatch between the constructs and behaviors.Cronbach’s α is computed for model-by-scale cells, with canonical reverse-scoring applied to the 19 reverse-keyed BFI-44 items before analysis.
  • F.1 Big Five (BFI-44) internal construct validity: Big Five reports are internally reliable under persona induction but still fail behaviorally, indicating a construct–task mismatch rather than measurement noise.Persona-induced mean α ranges from 0.79–0.86 versus human targets of 0.79–0.88, while behavioral prediction remains r_aligned ≈0.06–0.07.
  • F.1 Big Five (BFI-44) internal construct validity: Grid induction lowers Big Five reliability, especially for Extraversion, largely because GPT-4o Mini produces degenerate near-zero item variance.Extraversion has α = 0.24 in the grid condition; excluding GPT-4o Mini raises grid means to E: 0.50, A: 0.66, C: 0.75, N: 0.78, O: 0.78.
  • F.2 TPB (TACT-anchored) internal construct validity: TPB Intention reliability is stable across inductions, but other construct reliabilities depend on task, with CCT showing the clearest weakness.Mean Intention α = 0.71 under both inductions; Honesty and IAT are generally reliable, whereas CCT-Subjective Norm reaches α = 0.28 under personas.
  • F.2 TPB (TACT-anchored) internal construct validity: Construct-mean aggregation across approximately 54 observations per model-by-condition cell buffers item-level noise, so lower CCT reliability moderates rather than invalidates correlation estimates.Within-model coupling estimates are Fisher-z aggregated across cells, making the construct means more stable than individual TPB items.

F.4 Behavioral and self-report variance: floor, ceiling, and between-model differentiation … G.5 Between-model coherence

Variance diagnostics rule out SR compression as the main explanation for the reported prediction patterns, while identifying specific ceiling effects and same-session behavioral priming. The statistical analysis uses Fisher-z aggregation and complementary proportion, pooled-regression, contrast, and between-model coherence tests under defined limitations.

  • F.4 Behavioral and self-report variance: floor, ceiling, and between-model differentiation: Big Five self-reports show no floor or ceiling effects within-session, ruling out response compression as an explanation for their prediction results.Behavioral floor and ceiling rates are not interpreted as measurement failure because policies deliberately span the scale.
  • F.4 Behavioral and self-report variance: floor, ceiling, and between-model differentiation: 99.7% at floor with ICC ≈0 occurs for independent_judgment, reflecting same-session priming rather than measurement failure; defer_when_uncertain shows 100% of theoretical range and ICC 1.00.The mirror policy has η2 = 0.57, and sycophancy analyses obtain power from cross-policy contrasts.
  • F.4 Behavioral and self-report variance: floor, ceiling, and between-model differentiation: Between-session persona induction preserves TPB variance across all 16 task-by-construct cells, with ICC ≥0.85 and significant between-model differentiation.Big Five Openness reaches ceiling in 48.2% of cells, without producing SR–behavior coupling.
  • G.1 Primary: Fisher-z meta-analytic aggregation: Within each model, Pearson correlations link TPB constructs to sign-corrected align_score across approximately 54 observations per cell, with Fisher-z 95% confidence intervals.Inverse-variance-weighted meta-analysis aggregates on the z-scale and back-transforms results to Pearson r, combining approximately 378 observations per model.
  • G.2 Proportion metrics and null baseline: Direction-correct and alignment-hit proportions use Wilson-score 95% confidence intervals and are tested against a 2.5% null alignment-hit rate.With 77 cells, the null predicts approximately 1.9 chance alignment-hit cells, evaluated using a binomial z-test.
  • G.3 Robustness I: Mundlak pooled OLS with cluster-robust SEs: Mundlak pooled OLS separates within-model and between-model self-report effects, with model-clustered standard errors and standardized coefficients comparable to Pearson r.βwithin tracks Fisher-z r but need not match because pooled OLS assumes homogeneous slopes while Fisher-z aggregates model-specific slopes.
  • G.4 Robustness II: Policy-contrast specification: Matched-policy difference-score correlations subtract model-level response style across 297 conditions, but Honesty contrasts compare non-equivalent strategies rather than poles of one behavioral axis.The Honesty contrast is reported for completeness and flagged as requiring cautious interpretation.
  • G.5 Between-model coherence: Between-model coherence correlates model-mean TPB self-reports with model-mean aligned behavior across 11 models, providing directional ancillary evidence despite low power.Analyses restrict to grid perturbations at fixed persona context to isolate measurement coupling from identity-induction effects.

G.6 Robustness results … I.8.4 Within-model: persona induction

Robustness analyses refine the framework comparison: TPB predicts within-condition behavioral variation, whereas Big Five effects primarily reflect between-model differences. Across sessions and persona induction, coherence is task-dependent, while persona prompting substantially stabilizes self-reports without consistently aligning them with behavior.

  • H.1 Framework-specific outcome choice: TPB items target behavior-specific, policy-framed intentions, whereas Big Five items measure general dispositions; this asymmetry intentionally tests granularity as the mechanism.The comparison uses policy-corrected alignment outcomes so positive raligned indicates theory-consistent prediction for both frameworks.
  • H.3 Within/between decomposition reveals a qualitative distinction between frameworks: Big Five βwithin values are 0.00–0.05 and non-significant, while CCT Neuroticism reaches βbetween = +0.63 (p = .003).This separates stable model-identity differences from condition-sensitive shifts that Big Five does not predict.
  • G.6 Robustness results / H.4 Summary: TPB’s within-model advantage over Big Five persists under pooled OLS, while Big Five within-model effects remain uniformly null.Big Five traits carry essentially no condition-level predictive signal, whereas occasional between-model effects reflect stable model-level differences.
  • I.1 RQ3: Per-task cross-session results: Across separate sessions, sycophancy collapses, honesty attenuates, and CCT and IAT remain stable.The cross-session analysis defines Δr = rsame − rseparate and identifies task-specific persistence of coherence.
  • I.3 Robustness: formal pooled-OLS interaction test: The pooled OLS interaction test agrees with Fisher-z Δr in 14 of 16 cells, with divergences driven by weighting differences across models.For example, Honesty TPB Attitude has Fisher-z Δr = −0.15*** but OLS βSR×P = −0.01, p = .86.
  • I.4 Mundlak within/between decomposition / I.6 Full session×induction interaction: Persona induction reduces Honesty coupling but preserves its positive sign, while IAT–intention dissociation remains highly stable across induction and session conditions.Honesty within-session βwithin changes from +0.47* under grid to +0.42* under personas; IAT βwithin values range from −0.60 to −0.73, all p < .001.
  • I.7 Prerequisite analyses: discriminability and stability: Persona prompting increases self-report stability, most dramatically for Big5 Openness, but this stabilization does not produce a comparable increase in SR–behavior coherence.Big5 Openness rises from rgrid = +0.01 to rpersonas = +0.61 (Δ = +0.60***), while TPB shifts are smaller at Δ = +0.14 to +0.20.
  • I.8.2 Between-model: persona induction / I.8.4 Within-model: persona induction: Persona induction leaves between-model structure largely unchanged and preserves within-model TPB patterns, while within-model Big Five correlations remain near zero.The persona panels retain TPB clustering and CCT within-policy covariation, corroborating that Big Five fails to predict within-model behavior.

I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition … K Prompts and Stimuli

Within-session SR–behavior coherence reflects both context-driven priming and stable dispositional structure, with their relative contributions predicting whether coherence survives across sessions. The experiments operationalize these mechanisms using task-specific TPB reports, task-agnostic Big Five reports, separate session designs, and grid or persona induction.

  • I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition: Within-session coupling arises from policy-driven priming and stable dispositional structure.Priming occurs when self-report framing remains in context during behavioral choice; disposition reflects model differences that align self-report with behavior independently of policy.
  • I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition: RSR/Beh < 1 indicates behavior shifts more than self-report, whereas RSR/Beh > 1 indicates self-report shifts more than behavior.The ratio uses per-model×task Cohen’s d shifts between Policy A and Policy B, averaged across 11 models.
  • I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition: Spearman ρ(R, |rbetween|) = 1.0 across four tasks: Sycophancy and CCT show priming-dominated shifts and cross-session collapse, whereas Honesty and IAT retain coherence.The ordering of the shift ratio exactly tracks cross-session absolute-coupling magnitude.
  • I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition: R = 0.12 for Sycophancy: policy sharply splits within-session behavior, but removing self-report context yields cross-session collapse to r = −0.07.Deferral is 0% under independent_judgment and 43% under defer_when_uncertain, while the between-session distribution centers at 55% deferral.
  • I.9 Self-Report & Behavior Coherence Mechanism: decomposing within-session coherence into priming and disposition: Big Five self-reports also perturb behavior despite lacking policy framing, with within- and between-session behavior means differing in three of four tasks.CCT, Honesty, and Sycophancy shift, whereas IAT remains nearly unmoved at within = 0.86 versus between = 0.91.
  • J.1 Selection procedure: Persona induction uses 30 diverse PersonaHub character descriptions selected from a 500-persona pool by greedy max-min TF-IDF diversity.The selected personas have mean pairwise cosine distance 0.999, with minimum 0.991 and maximum 1.000.
  • K Prompts and Stimuli: TPB uses four task-specific constructs on 1–7 Likert scales, whereas Big Five uses a task-agnostic 44-item block on a 1–5 scale.Same-session probing keeps self-report and behavior in one thread; separate sessions use independent API calls with no shared conversational history or access to the model’s self-report.

L Computational Resources · Big Five (BFI-44)

The study used API-hosted inference on a standard workstation to evaluate 11 models across behaviorally targeted and Big Five self-report probes. The BFI-44 elicited five broad traits through 44 Likert-rated items, alongside task-specific behavioral measures spanning risk-taking, sycophancy, honesty, and implicit bias.

  • L Computational Resources: All inference ran through OpenRouter-hosted endpoints, requiring no local GPU, TPU, or institutional HPC cluster.Sweep orchestration used only a standard workstation.
  • L Computational Resources: The full factorial design crossed 11 models, 4 tasks, 2 session types, and 2 induction conditions, producing approximately 5000 conditions.Each condition generated 19–105 sequential API calls depending on the task.
  • L Computational Resources: Approximately $200 covered primary-experiment inference, pilot runs, and failed calls; two models used OpenRouter’s free tier, while nine paid models ranged from Phi-4 to Claude 3.7 Sonnet rates.The free tier was rate-limited to 20 req/min, and paid-model prices spanned $0.065/$0.14 to $3.00/$15.00 per 1M input/output tokens.
  • L Computational Resources: Persona induction changed cross-session TPB correlations by −0.15 for Honesty, +0.16 for Sycophancy, and +0.02 for IAT, while CCT changed by −0.02.The Honesty shift was reported as statistically significant, whereas the CCT and IAT intervals included zero.
  • L Computational Resources: Within-model results showed TPB associations varying by task, while Big Five associations were generally small and uncertain across CCT, Honesty, Sycophancy, and IAT.Reported Big Five coefficients included between-model values from −0.05 to +0.07 and within-model values from −0.03 to +0.21, all with confidence intervals spanning zero.
  • Big Five (BFI-44): The Big Five elicitation instructed models to return only valid JSON, map each item to an integer from 1 to 5, and not mention being an AI.The prompt asked how accurately each statement described the participant.
  • Big Five (BFI-44): The BFI-44 measured Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness using 44 statements rated from 1 to 5.Items were framed as “I see myself as someone who...” statements, with reverse-keyed items included across traits.
  • Big Five (BFI-44): Behavioral evaluation covered risk-taking, sycophancy, honesty, and implicit bias using mean cards flipped, sycophancy rate, calibration and confidence consistency, and IAT d-score outcomes.The tasks used 20-round card decisions, 52 moral dilemmas presented twice, a two-stage confidence paradigm, and six-domain IATs.
Loading 2606.12730v1…