Source-linked AI summary

Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations

Xiaoxu Ma, Xiangbo Zhang, Zhenyu Weng

arXiv:2601.09833v1cs.CL

TL;DR

Existing questionnaire-based LLM personality evaluations are unstable under prompt and role-play variations and provide limited interpretability. PVNI addresses this by extracting persona vectors from internal activations and interpolating neutral scores along those directions. Across diverse LLMs and evaluation variants, PVNI produces substantially more stable evaluations than existing methods, while retaining judge-dependent absolute anchoring and relying on an approximate neutral-axis assumption.

  • Problem

    Questionnaire-based evaluations are sensitive to prompt changes, primarily measure prompted role play, and provide limited evidence of stable internal traits or interpretable scoring.

  • Method

    PVNI extracts target-trait persona vectors from contrastive prompts and estimates neutral trait strength by interpolating along the internal activation direction.

  • Results

    PVNI achieves substantially greater stability than existing methods across diverse LLMs and questionnaire and role-play variants.

  • Takeaways & Limitations

    Internal representation geometry provides an interpretable basis for comparing neutral prompts with trait directions in LLM personality evaluation.

  • Takeaways & Limitations

    Absolute values can reflect judge preferences, and interpolation assumes neutrality lies approximately on the positive–negative persona axis.

Abstract

from arXiv · show

Evaluating personality traits in Large Language Models (LLMs) is key to model interpretation, comparison, and responsible deployment. However, existing questionnaire-based evaluation methods exhibit limited stability and offer little explainability, as their results are highly sensitive to minor variations in prompt phrasing or role-play configurations. To address these limitations, we propose an internal-activation-based approach, termed Persona-Vector Neutrality Interpolation (PVNI), for stable and explainable personality trait evaluation in LLMs. PVNI extracts a persona vector associated with a target personality trait from the model's internal activations using contrastive prompts. It then estimates the corresponding neutral score by interpolating along the persona vector as an anchor axis, enabling an interpretable comparison between the neutral prompt representation and the persona direction. We provide a theoretical analysis of the effectiveness and generalization properties of PVNI. Extensive experiments across diverse LLMs demonstrate that PVNI yields substantially more stable personality trait evaluations than existing methods, even under questionnaire and role-play variants.

1 Introduction

Existing LLM personality evaluations are prompt-sensitive and difficult to interpret, motivating PVNI, which uses internal activations and persona vectors for stable, explainable trait evaluation.

  • Personality testing supports standardized descriptions of stable individual differences and has been extended to quantify human-interpretable behavioral tendencies in LLMs.
  • Self-report assessments and open-ended questionnaires provide scalable personality profiling but depend on researcher-designed prompts.
  • Minor changes in prompt framing or wording can substantially shift questionnaire-based trait scores without an underlying model change.
  • These methods primarily measure prompted role play, offering limited evidence that scores reflect stable internal properties and making the scoring process difficult to interpret.
  • PVNI derives a target-trait persona vector from positive and negative contrastive prompts, then estimates neutral trait strength through activation-space interpolation.
  • PVNI is supported by a linear theory of persona vectors and experiments showing more stable evaluations across diverse LLMs, including questionnaire and role-play variants.

2 Related Works

Related work studies personality evaluation through Big Five assessments and persona vectors, while PVNI uses internal trait directions to produce more stable, explainable estimates across prompt variants.

  • Personality Trait Evaluation: LLM personality evaluation commonly uses the Big Five OCEAN framework through self-report assessment or open-ended elicitation.
  • Persona Vector: Persona vectors represent target traits as interpretable directions in internal representation space for control, monitoring, and intervention.
  • Persona Vector: Prior work uses persona vectors for personality measurement, and this paper extends them to quantify LLM personality from internal trait directions.
  • Persona Vector: The proposed direction yields more stable and explainable estimates across prompt variants.
  • Linear Properties in Hidden-State Space: Related representation research motivates modeling high-dimensional activations with low-dimensional subspaces that are globally nonlinear but locally near linear.

3 Persona-Vector Neutral Interpolation

PVNI represents Big Five traits as persona directions in hidden-state space and estimates prompt-neutral trait scores by interpolating between judged anchors using representation geometry.

  • Persona-vector representation: PVNI defines a representation-space coordinate system for Big Five traits using persona-vector directions and projection-based estimates.The resulting axes support trait-specific score estimation in a shared personality subspace.
  • Persona-vector representation: For each trait, contrastive positive and negative prompts produce a persona vector, while a neutral prompt locates neutral behavior in representation space.The persona vector is obtained from mean hidden-state differences and normalized before use.
  • Score anchoring: An API-based judge supplies 0–100 scores for trait expression, which PVNI uses as anchors for interpolation.Higher judged values indicate stronger expression of the target persona, and scores are aggregated over inputs.
  • Weight interpolation: PVNI computes an interpolation weight from hidden-space geometry and linearly combines the positive and negative score anchors to estimate the neutral trait score.The neutral prompt is used to locate neutral behavior, while the resulting estimate is returned independently for each Big Five trait.
  • Big Five output: Applying the procedure across O, C, E, A, and N yields a prompt-neutral Big Five coordinate vector and an embedding whose axes are anchored by trait-promoting directions.The pipeline maps the five estimated neutral scores into the Big Five personality representation.

4 Linear Theory of Persona Vector

The theory models persona-induced hidden-state changes as approximately linear along trait directions, supporting directional amplification, trait composition, and out-of-domain persona synthesis under stated assumptions.

  • Theoretical assumptions: PVNI’s theory assumes local linearity of persona scores and well-trained persona adaptation, providing the basis for analyzing hidden-state updates.The analysis also assumes bounded residual effects in typical hidden-state regions.
  • Multi-persona composition: Persona correlations determine composition behavior: orthogonal traits can combine at constant scale, positive correlation reduces the required scale, and negative correlation creates a trade-off.The framework labels trait pairs aligned, contradictory, or orthogonal according to the sign of their representation-space correlation.
  • Representation linearity: Lemma 4.1 characterizes persona vectors as approximate rank-one amplifiers that selectively boost a trait direction while leaving a bounded residual.The induced attention reweighting amplifies interactions with positive projection on the persona direction and suppresses the rest up to O(β).
  • Representation linearity: Persona vectors act as directional amplifiers, yielding near-linear additivity in persona editing and supporting the observed low-dimensional trait-subspace structure.The theory links the representation update to the component of the hidden state aligned with the persona direction.
  • Multi-persona composition: Persona negation is easiest for orthogonal or contradictory traits, while highly aligned traits are harder to suppress without affecting the other trait.Negative coefficients can suppress one trait while preserving another in supported regimes, whereas mildly aligned traits may require small positive coefficients.
  • Generalization: An out-of-domain persona can be synthesized from Big Five directions when its direction lies mostly within their span and the residual component remains controlled.The construction decomposes the target direction into an in-subspace component and an orthogonal residual, then combines in-domain updates.

5 Experiments

The experiments evaluate Big Five traits across three open-source LLMs and multiple prompt variants. PVNI shows the lowest variability and strongest robustness to prompt rephrasing, while trait means remain broadly similar across questionnaire and role-play variants.

  • Experimental setup: Experiments evaluate Qwen-2.5-7B, Llama-3-8B, and Mistral-7B-v0.1 on Big Five traits using four evaluation protocols.Each protocol uses questionnaire and role-play variants, with mean ± std reported across prompt sets.
  • Stability results: PVNI consistently achieves the smallest standard deviation across all models and traits, indicating the strongest prompt robustness.Table 1 and Table 2 compare variability across questionnaire and role-play variants.
  • Variant analysis: Questionnaire and role-play variants produce similar means, while role-play variants typically show slightly lower variance.Questionnaire variants modify question content more substantially, creating larger perturbations.
  • Trait-level patterns: Traits with lower mean scores, such as Neuroticism and Extraversion, tend to exhibit larger standard deviations.This mean–variance coupling indicates higher uncertainty when trait strength is weak.
  • Stability results: PVNI produces the narrowest uncertainty bands across all five traits, while other protocols show wider prompt-induced fluctuations.The shaded bands represent ± one standard deviation over questionnaire variants, and the pattern also holds across models and role-play variants.
  • Stability results: PVNI yields tighter boxplot distributions with smaller IQRs and shorter whiskers than Open-ended Elicitation and self-report protocols.The comparison spans Qwen-2.5-7B, Llama-3-8B, and Mistral-7B-v0.1.

6 Conclusion

The paper concludes that PVNI provides a stable and explainable approach to personality trait evaluation by extracting trait directions from contrastive prompts and interpolating along them. Across LLMs and prompt variants, PVNI consistently reduces variance relative to prior protocols.

  • Conclusion: PVNI extracts persona vectors from contrastive prompts and estimates prompt-neutral trait strength through projection and interpolation.The method uses internal activation geometry and judged positive and negative anchors.
  • Conclusion: Across LLMs and questionnaire and role-play variants, PVNI consistently reduces evaluation variance over prior protocols.The conclusion identifies variance reduction as the central empirical outcome.

7 Limitations

The paper identifies practical, methodological, access, and scope limitations for PVNI. These include computational overhead, judge bias, geometric assumptions, correlated trait directions, limited model and setting coverage, and incomplete prompt-space coverage.

  • Methodological constraints: Judge preferences can affect PVNI’s absolute trait values, so its primary claim is variance reduction rather than judge-invariant calibration.The method uses a judge for positive and negative anchor scores, while interpolation relies on internal representation geometry.
  • Geometric assumptions: PVNI assumes that neutrality lies approximately along the positive–negative persona axis, which may fail for curved, multimodal, or context-dependent trait effects.Clipping and prompt-set averaging reduce pathological behavior but do not guarantee calibration in highly nonlinear regions.
  • Access constraints: PVNI requires hidden-state extraction, limiting direct applicability to closed, API-only models.The paper suggests accessible surrogates such as logit-space proxies or distillation as future directions.
  • Trait disentanglement: Correlated persona directions complicate axis-specific interpretation because shifting one trait may partially move other traits.The paper treats traits as a shared low-dimensional structure rather than perfectly independent axes.
  • Evaluation scope: Experiments cover Big Five traits, a small set of open-source instruction-tuned models, controlled prompt variants, and largely single-turn settings.Transfer to multilingual, long-conversation, and tool-augmented-agent settings remains untested.
  • Prompt coverage: Persona directions may remain incomplete proxies when designed contrastive prompts provide narrow coverage of the prompt space.Prompt sets and two controlled variant types partially mitigate this limitation, but broader or adversarial selection is proposed.
  • Practical constraints: PVNI incurs non-trivial computational overhead from generation, judging, and activation extraction across multiple traits and prompt sets.The cost may be especially significant for larger models, despite parallelization and caching opportunities.

A.1 Questionnaire and Role-Play Variants

The evaluation varies prompt formulation through questionnaire rewrites and role-play framing while preserving the targeted trait and core evaluation setup. These variants test robustness to alternative question phrasings, aligned open-ended probes, minimal persona framing, and alternative contrastive instructions.

  • Variant Design: Prompt-robustness variants perturb only the prompt wrapper while preserving the trait target, model, decoding settings, response format, question count, and judge procedure.This isolates sensitivity to prompt formulation rather than changes in the underlying evaluation pipeline.
  • Questionnaire Variants: Questionnaire variants rewrite equivalent self-report items or replace open-ended questions with alternative probes targeting the same trait direction.Self-report edits preserve item meaning and rating scale, while open-ended rewrites keep the instruction fixed.
  • Role-Play Variants: Role-play variants add minimal framing to self-report questionnaires or rewrite only the trait-eliciting instruction for open-ended elicitation and PVNI.The latter uses multiple positive/negative instruction pairs with the same contrastive intent.
  • Robustness Targets: The design tests whether personality estimates remain stable under surface-form changes, aligned probes, persona framing, and alternative positive/negative realizations.Questionnaire variants emphasize wording or question-set changes, whereas role-play variants emphasize framing or instruction changes.

A.2 Prompt Robustness Analysis

PVNI is the most prompt-robust protocol in the controlled comparison, producing the smallest uncertainty and standard deviations across models, traits, and both variant types. Self-report protocols are generally most prompt-sensitive, while role-play variants tend to reduce variance relative to questionnaire rewrites.

  • PVNI Stability: PVNI consistently yields the smallest uncertainty across all models and both questionnaire and role-play variants.For Mistral-7B-v0.1, examples include standard deviations of 0.65/0.48 for O and 0.32/0.26 for C under questionnaire/role-play variants.
  • Baseline Variability: Self-Report Assessment is the most prompt-sensitive overall, with IPIP-BFFM-50 and IPIP-NEO-120 showing wider uncertainty bands and larger standard deviations than PVNI.Open-ended elicitation is also unstable for some traits, but its variance is less consistently high than IPIP across settings.
  • Variant Effects: Role-play variants tend to reduce variance while preserving similar mean profiles, whereas questionnaire variants introduce larger surface-form perturbations.The passage presents the explanation as plausible because role-play changes only a minimal framing line while questionnaire variants rewrite question text more aggressively.

A.3 Prompt Variability via Boxplots

The boxplots compare four protocols across Big Five traits under questionnaire and role-play variants, using spread to show prompt sensitivity. PVNI remains tightest across models and most traits, while IPIP methods and open-ended elicitation show greater variability.

  • Figure Layout: Figures 6 and 7 compare OCEAN boxplots for IPIP-BFFM-50, IPIP-NEO-120, Open-ended Elicitation, and PVNI under questionnaire and role-play variants.Wider boxes and longer whiskers indicate higher prompt sensitivity.
  • PVNI Stability: PVNI consistently shows the tightest boxes and shortest whiskers across all three LLMs and nearly all traits.This pattern identifies PVNI as the most stable protocol in the boxplot comparisons.
  • Baseline Variability: IPIP-BFFM-50 and IPIP-NEO-120 exhibit the largest spread, while Open-ended Elicitation is unstable for several traits with occasional extreme ranges.Role-play variants preserve similar medians but reduce variance relative to questionnaire variants.
Loading 2601.09833v1…