Source-linked AI summary
PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang, Qiang Hu, Chao Shen, Xiaofei Xie
TL;DR
Existing LLM personality evaluations largely overlook that traits can vary with context, despite the importance of predictable personality for interactive systems. PTCBench addresses this gap with controlled location and life-event interventions, NEO-FFI measurement, and large-scale comparisons across models and agents. The results show substantial context-sensitive trait shifts, architecture-dependent stability, and effects of some contexts on reasoning performance.
Problem
Existing LLM studies emphasize static or role-based personality consistency, overlooking systematic trait changes induced by situational shifts despite their importance for interactive systems.
Method
PTCBench combines 12 contextual conditions with preset personality configurations and measures pre- and post-context traits using the NEO-FFI.
Results
LLMs have reproducible native personality baselines but can shift substantially under specific locations and life events, with architectures varying widely in contextual stability and some contexts affecting reasoning performance.
Takeaways & Limitations
Personality stability should be evaluated as a controllable, context-sensitive system property when developing reliable and psychologically interactive LLM agents.
Takeaways & Limitations
Findings may not generalize to emerging models and agents, because the benchmark studies representative rather than comprehensive model coverage.
Abstract
from arXiv · showhide
With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user trust and engagement. However, existing work overlooks a fundamental psychological consensus that personality traits are dynamic and context-dependent. To bridge this gap, we introduce PTCBENCH, a systematic benchmark designed to quantify the consistency of LLM personalities under controlled situational contexts. PTCBENCH subjects models to 12 distinct external conditions spanning diverse location contexts and life events, and rigorously assesses the personality using the NEO Five-Factor Inventory. Our study on 39,240 personality trait records reveals that certain external scenarios (e.g., "Unemployment") can trigger significant personality changes of LLMs, and even alter their reasoning capabilities. Overall, PTCBENCH establishes an extensible framework for evaluating personality consistency in realistic, evolving environments, offering actionable insights for developing robust and psychologically aligned AI systems.
1. Introduction
LLM systems increasingly serve as interactive partners, making stable and predictable personality important for trust, engagement, and coherent personalization. PTCBench addresses the overlooked question of whether LLM personality traits change under evolving contexts by measuring responses to controlled environmental conditions.
- Interactive LLM applications require personality that is empathetic, engaging, emotionally steady, and socially appropriate across extended interactions.These qualities matter in personalized companions, role-playing assistants, and long-form dialogue agents.
- Personality consistency supports user trust, reduces interaction uncertainty, and enables more coherent personalization as agentic architectures become more autonomous and difficult to anticipate.Personality can emerge from the interaction among the base model, system instructions, memory, and evolving conversation.
- Existing studies profile prompted personalities and assess longer-term persona consistency, but largely overlook systematic trait variation caused by life events and location changes.This leaves static personality scores and narrow stay-in-role metrics insufficient for evolving environments.
- PTCBench evaluates whether LLM personality traits remain stable or change systematically when exposed to controlled contextual interventions.The benchmark targets environment-induced variation rather than only static personality profiles or narrow stay-in-role consistency.
- PTCBench combines 12 external conditions with NEO-FFI assessment and 39,240 personality trait records from four LLMs and two agents.The study reports comparatively stable profiles for some foundation models, greater variability for agentic systems, baseline-dependent trait changes, and especially large deviations under Divorce and Unemployment.
- The benchmark provides a scalable, psychologically grounded way to map LLM personality dynamics and support more robust, authentic, and controllable interactive agents.Its contributions include operationalizing context-induced personality change and identifying alignments and divergences with established human psychological mechanisms.
2. Background and Related Work
Human personality research characterizes traits as longitudinally stable yet systematically variable across life events, locations, and situations. Related LLM work can induce and measure static personalities, but lacks a benchmark for evaluating human-like trait dynamics under situational shifts.
- 2.1. Human Personality Trait Changes: Human personality exhibits longitudinal stability alongside systematic state-level variation across life events, locations, and situational cues.Career changes are associated with gradual conscientiousness shifts, while immediate social environments can transiently modulate extraversion.
- 2.1. Human Personality Trait Changes: A complete contextual scenario integrates the actor, location or environment, and ongoing activity or event to ground personality expression.Figure 1 presents these components as typical elements for describing an incident’s comprehensive context.
- 2.1. Human Personality Trait Changes: Despite progress in modeling context-sensitive behavior, a systematic benchmark for measuring context-induced personality changes in LLMs and comparing them with human patterns remains absent.This gap concerns both evaluation of LLM trait dynamics and direct comparison with empirical human evidence.
- 2.2. LLM Personality Traits: LLMs can express prompt-controllable personality profiles, while newer work studies trait induction, role-playing consistency, personality recognition, and psychologically informed decision-making.Examples include PersonaLLM (Zollo et al., 2024), neuron-level trait induction (Deng et al., 2024), and PsycoLLM (Hu et al., 2024).
- 2.2. LLM Personality Traits: Most existing LLM studies assess static personality expressions rather than whether traits change in human-like ways when situational conditions shift.This limitation motivates a benchmark focused on context-dependent personality dynamics.
3. Benchmark Construction
PTCBench constructs controlled contextual pressures by combining external scenarios with preset personality configurations, then measures trait changes before and after exposure. Its NEO-FFI-based evaluation uses contextual grounding, randomized response options, and decaying conversational history to quantify consistency.
- 3. Benchmark Construction: PTCBench uses two stages—Prompt Construction and Personality Evaluation—to quantify whether LLM traits remain stable or shift across situational conditions.The benchmark compares model responses against human empirical baselines and is summarized in Figure 2.
- 3.1. Prompt Construction: Prompts combine an external scenario representing situational context with an internal configuration representing the model’s preset personality state.A unified template automatically injects selected contextual descriptions and personality profiles for controlled comparisons.
- 3.1. Prompt Construction: The external scenarios cover six location contexts and six life events, including Social Venues, Workplaces, Childbirth, Divorce, Marriage, and Job Loss.These scenarios are operationalized as detailed descriptions of situational cues relevant to human experience.
- 3.1. Prompt Construction: Preset Personalities manipulate the Big Five dimensions at Low, Medium, or High intensity, plus a None control group, producing 244 distinct configurations.The intensity levels are defined relative to a maximum inventory score of 60, with approximate example scores of 10, 30, and 50.
- 3.2. Personality Evaluation: Personality Evaluation administers the NEO-FFI before and after each contextual scenario to measure trait consistency and context-induced change.The instrument balances psychometric reliability with practical administration length and follows standard protocols.
- 3.2. Personality Evaluation: The evaluation grounds prompts in explicit context, randomizes response-option order, and models conversational history with distance-based decay.These procedures target contextual validity, positional bias, and continuity of cognitive context.
4. Evaluation
PTCBench evaluates native and preset personality consistency across location and life-event contexts, then examines how context-induced trait changes relate to reasoning performance. Models show stable neutral baselines but architecture-dependent shifts, with agentic systems generally more unstable and preset traits moderating change.
- 4.1. Experiment Setup: PTCBench evaluates four foundation models and two agentic frameworks using NEO-FFI trait scores, contextual-change coefficients, standardized event changes, and uncertainty estimates.Agentic evaluations use default AutoGen and CAMEL configurations with GPT-4o-mini as the backbone.
- 4.2. Consistency of Native Personality: LLM systems exhibit stable and reproducible native personality baselines under neutral conditions, with consistently low Neuroticism across models.Gemini 2.0 Flash, Claude Sonnet 4, and AutoGen show relatively high Openness and Conscientiousness, while repeated measurements have minimal variance.
- 4.2. Consistency of Native Personality: Location contexts produce model-dependent shifts: foundation models show bounded adaptation, whereas AutoGen exhibits pronounced instability in task-related traits.Gemini shows adaptive increases in Openness and Extraversion, while AutoGen has Conscientiousness decreases nearing -20 across locations; compared with human changes typically below 0.8, agentic responses are exaggerated.
- 4.2. Consistency of Native Personality: Life events polarize personality responses: relational events can increase sociability, while negative events trigger large destabilizing shifts, especially in agentic systems.During Divorce, GPT-4o-mini shows Extraversion -10.3, Agreeableness -5.6, and Neuroticism +11.6; AutoGen shows Conscientiousness decreases of -16 to -18 under Unemployment and Graduation.
- 4.3. Consistency of Preset-Personality: GPT-4o-mini follows preset personality instructions monotonically, while high presets constrain and low presets amplify context-induced trait changes.Under Divorce, Extraversion changes from -3.28 with High presets to +0.67 with Low presets, while Neuroticism reaches +6.17 for Low presets versus -2.65 for High presets.
- 4.4. Personality Impact on LLM Performance: Personality changes can enhance or impair reasoning: Openness increases of around 20% yield up to +20% AGIEval accuracy, whereas stress-related trait disruptions often coincide with lower performance.Large Neuroticism increases of 56-58% are commonly accompanied by broad accuracy drops across tasks, while effects of Agreeableness changes are mixed.
5. Discussion
PTCBench indicates that LLM personality is a reproducible behavioral prior that changes systematically with context and can affect downstream reasoning. The discussion connects these findings to bounded adaptation, stable presets, and future causal evaluation.
- Contextual personality modulation can meaningfully affect downstream capabilities rather than remaining merely cosmetic.
- Negative events may produce uncontrolled personality shifts that undermine user trust, while preset personalities can provide stable priors for long-term regulation.
- Personality-like behavior appears stable yet context-sensitive, may be amplified by agentic components, and should be treated as a system property with explicit bounds on trait change.
- Increased Openness improves reasoning, whereas reduced Conscientiousness or elevated Neuroticism degrades it, motivating evaluation under contextual and personality perturbations.
- Future work should investigate personality origins, adaptive personality-aware design, and longitudinal causal evaluation beyond single-context assessment.
6. Conclusion
PTCBench evaluates personality consistency under controlled situational contexts and finds that LLM traits can shift substantially across locations and life events. These shifts can affect reasoning performance, while model architectures differ in stability under contextual pressure.
- LLM traits can shift substantially in response to specific locations and life events, with some contexts affecting reasoning performance.
- Different model architectures vary widely in their stability under contextual pressure.
- PTCBench provides a psychologically grounded framework for analyzing personality dynamics and supporting more reliable, psychologically interactive LLM systems.
7. Limitations
The study covers 12 contextual scenarios and six LLMs and agents, but its findings may not generalize to emerging models. Its human comparison is qualitative, and questionnaire-based measurement may miss long-term behavioral nuance.
- Findings from 12 contextual scenarios and six LLMs and agents may not hold for some emerging models and agents.
- The benchmark identifies patterns across representative systems rather than providing a comprehensive leaderboard of all available models.
- Human–LLM alignment is qualitative because comparisons rely on aggregated psychological literature rather than individual-level longitudinal data.
- Prompt-based NEO-FFI assessment may miss behavioral nuances in open-ended, long-term interactions, motivating dynamic interaction-based metrics.
8. Ethical Considerations
The ethical discussion treats contextual personality effects as relevant to affective and interactive applications while cautioning against misleading users, inappropriate emotional reliance, and anthropomorphic interpretations.
- The study treats LLM personality traits as behavioral abstractions rather than indicators of consciousness or human mental states.
- Contextual prompts can influence LLM personality expression, with implications for affective agents and interactive applications.
- Applying these insights should avoid misleading users or fostering inappropriate emotional reliance, and human comparisons do not imply equivalence between LLMs and humans.
A. Appendices
The appendices document the experimental setup, robustness safeguards, prompt construction, and supplementary analyses supporting PTCBench’s main findings.
- Appendix A.1 details the evaluated models, metrics, and implementation procedures.
- Appendix A.2 describes four strategies intended to improve the robustness and reliability of the experimental findings.
- Appendix A.3 presents PTCBench’s prompt template and synthesis process.
- Appendix A.4 supplies supplementary tables and figures covering reliability, trait changes, preset personalities, and performance under context-induced personality shifts.
A.1. Evaluation Setup Details
PTCBench evaluates three foundation models with NEO-FFI-derived trait metrics, contextual effect estimates, uncertainty measures, and reliability statistics across six locations and six life events.
- Models and Agents: Experiments cover Google Gemini 2.0 Flash, OpenAI GPT-4o, and Anthropic Claude 2.0.
- Metrics: Trait scores average item responses within each NEO-FFI dimension to produce normalized personality-dimension scores.
- Metrics: Multilevel regression coefficients B_jl estimate the expected trait change associated with each contextual condition, with sign indicating direction and magnitude indicating strength.
- Metrics: Standard errors quantify uncertainty in estimated contextual effects, enabling comparisons of effect magnitude and estimate reliability.
- Metrics: Standardized mean change d_e measures before–after trait differences for life events, with positive values denoting increases and negative values denoting decreases.
- Metrics: Reliability is assessed with ICC(3,1) for single measurements and ICC(3,k) for averaged repeated measurements, interpreted using established reliability thresholds.
- Implementation Details: Contexts comprise six locations and six life events, with each system assessed first neutrally and then after contextual prompts are introduced.
A.2. Reliability of PTCBench Result
PTCBench uses repeated runs, bias controls, statistical corrections, and reproducibility measures to strengthen the reliability of its findings.
- Reliability safeguards: Each assessment uses three random-seed runs, accepts scores only when ICC(3,k) exceeds 0.8, and averages stable runs with uncertainty intervals.
- Reliability safeguards: Response-option shuffling and manual auditing reduce potential positional bias in multiple-choice personality assessments.
- Reliability safeguards: Cohen’s d, Wilcoxon signed-rank tests, and Bonferroni correction provide effect-size, significance, and multiple-comparison controls.
- Reliability safeguards: Open-sourced prompts, scenario templates, evaluation scripts, and a fixed-dependency Docker container support reproducibility.
A.3. Prompts used in PTCBench
The appendices define PTCBench’s prompt construction and report supplementary reliability, personality-change, preset-personality, and reasoning-performance analyses across models and agents.
- A.3. Prompts used in PTCBench: PTCBench separates internal personality configurations from external scenarios through automated baseline and situated prompt templates.
- A.3. Prompts used in PTCBench: Situated prompts add location, life-event, prior trait scores, and selected previous answers before retesting the model with numeric responses.
- A.4.1. ICC Evaluation: Baseline reliability is high, while situational prompts reduce single-measure ICC across systems; aggregated ICC(3,k) remains 0.85–0.98 in most cases.
- A.4.2. Personality Trait Change of Different LLMs and Agents: Gemini shows bounded positive shifts in Openness and Extraversion, whereas Claude exhibits stronger event-polarity sensitivity with large negative-event changes and Neuroticism increases above +14.
- A.4.2. Personality Trait Change of Different LLMs and Agents: GPT-OSS-120B shows moderate asymmetric changes, including Neuroticism increases up to +17.9 under Divorce and Unemployment.
- A.4.2. Personality Trait Change of Different LLMs and Agents: CAMEL declines in Conscientiousness and Extraversion, while AutoGen shows the greatest instability, including Conscientiousness drops of -16 to -18 under Unemployment and Graduation.
- A.4.2. Personality Trait Change of Different LLMs and Agents: LLM changes exceed human changes, which are typically below 0.8 for locations and 0.11 for events, especially in agentic systems.
- A.4.3. Pre-set Personality for LLM Systems / A.4.4. Personality Impact on LLM Performance: Preset traits moderate contextual responses, while personality shifts affect reasoning unevenly: Openness gains are broad, Conscientiousness declines harm logical reasoning, and Neuroticism increases degrade nearly all tasks.