Source-linked AI summary

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Ming Wang, Peidong Wang, Xiaocui Yang, Daling Wang, Shi Feng, Fiona Fui-Hoon Nah, Ee-Peng Lim

arXiv:2608.06485v1cs.CLcs.AIcs.SI

TL;DR

How plausibly do personality-conditioned LLM agents evolve after major life events? The paper benchmarks Big Five changes across events, personas, demographics, and models, finding that agents move but weakly capture event- and person-specific human personality dynamics.

  • Problem

    Evidence remains limited on how life-event-induced personality shifts vary across traits, events, personas, and models in personality-conditioned LLM agents.

  • Method

    The study measures pre- and post-event Big Five profiles across 11 life events, controlled personas, and 11 models, with multiple robustness checks and BFI-Adapt evaluation.

  • Results

    PC-Agents show widespread but weakly event-specific movement, poorly calibrated magnitudes, and compressed demographic and persona-level variation relative to human personality dynamics.

  • Takeaways & Limitations

    Current PC-Agents approximate the mean of human personality change more readily than its event- and person-specific shape.

  • Takeaways & Limitations

    The pre-post design captures immediate responses but cannot assess multi-year adaptation trajectories documented in human longitudinal studies.

Abstract

from arXiv · show

Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

1 Introduction

This section frames personality evolution in personality-conditioned LLM agents as a human-prior diagnostic grounded in Big Five changes after major life events. It introduces a controlled multi-model study showing that agents move, but weakly match event specificity, human magnitudes, demographic variation, and persona-level heterogeneity.

  • Motivation: Personality-conditioned LLM agents support emotional support, social simulation, behavioral research, role-playing, games, and digital tutoring.These applications motivate the need for coherent long-horizon agents that can evolve plausibly across interactions.
  • Human benchmark: Human longitudinal findings provide event–trait anchors, including increased Conscientiousness after job entry or promotion and increased Neuroticism after chronic illness or unemployment.Retirement is associated with a documented Conscientiousness decline.
  • Study design: 11 major life events are evaluated across 11 LLMs using 100 demographically controlled personas, spanning 2 genders, 5 cultural regions, and 10 personality archetypes, with paired BFI-44 measurements.Each persona completes the inventory before and after a structured event-reflection stage.
  • Research questions: The study evaluates personality evolution along four axes: existence, direction and magnitude, demographic shape, and individual shape.The research questions ask whether changes exceed BFI subscale noise, match human directions and magnitudes, reflect demographic moderators, and preserve individual variation.
  • Central findings: PC-Agents move, but their changes are weakly event-specific, poorly calibrated in magnitude, and compressed across demographic and persona-level variation.The resulting interpretation is that agents simulate the mean of human personality dynamics more readily than its shape.
  • Validation and contribution: Validation shows event-conditioned changes exceed retest noise, survive independent paraphrases and unrelated dialogue, and show model-dependent convergence with scenario-based decisions.The introduction also presents BFI-Adapt as an evaluation method for scoring event-induced personality change.

2 Related Work

Prior work establishes that LLM personality can be measured, conditioned through role-playing, and shifted by temporal or contextual perturbations. These findings motivate standardized measurement of event-conditioned personality trajectories.

  • Static LLM Personality Assessment: Static assessment studies show that LLMs can produce reliable Big Five profiles and differentiate prompted personality levels.They support linguistic or inventory-based personality assessment, establishing the feasibility of measuring static LLM personality.
  • Static LLM Personality Assessment: These findings motivate using BFI-44 as a standardized, interpretable anchor for systematic event-conditioned personality trajectories.Projective-test adaptations also report greater sensitivity to prompt-induced changes and demonstrate GenPT in longitudinal counselling.
  • Role-Playing Agents: Role-playing benchmarks and training methods evaluate whether models remain faithful to specified characters or personas.Related work also finds that user personas can shift chatbot personality and that induced personas produce stable, task-dependent changes in cognitive performance.
  • Dynamic LLM Personality: Dynamic personality studies report limited temporal stability and measure trait shifts after external conditions, including locations and life events.Bodroža et al. re-administered inventories to seven LLMs, while Yu et al. introduced PTCBench to measure aggregate NEO-FFI shifts.

3 Methodology

The methodology constructs controlled personas, exposes them to literature-grounded life events, and measures Big Five changes with repeated BFI-44 assessments. It evaluates individual-level directional fidelity, item-response consistency, and robustness across alternative experimental conditions.

  • Persona Construction: 100 personas cross 2 genders, 5 cultural regions, and 10 personality types built from implicit Big Five anchors.Each type sets one focal trait to very high or very low while keeping the other four moderate.
  • Life Events: 11 literature-selected life events span occupational, social, and health domains, with expected trait directions consolidated into an event–trait prior matrix.Emotional Stability is inverted to Neuroticism to match the BFI-44 scoring convention.
  • Measurement Pipeline: A four-stage pipeline administers BFI-44 before and after first-person event reflection to obtain Big Five score vectors for each simulated persona.The system prompt specifies gender, cultural region, and behavioral personality description, while the model responds in character.
  • Directional Evaluation: Individual-level direction classification uses ε=0.1 as a conservative noise boundary, excludes neutral shifts, and tests match rates against the 0.5 chance baseline.Directional inference uses one-tailed binomial sign tests with Wilson-score intervals and Benjamini–Hochberg FDR correction at α=0.05.
  • Diagnostic and Robustness Measures: Pair-level linearly weighted Cohen’s κ measures baseline-to-post-event rating stability, while DCR measures whether changed items move predominantly in one direction across 27 definite-direction pairs.Robustness designs include no-event retests, independent event paraphrases, scenario-based decisions, and delayed measurement after three unrelated dialogue turns.

4 Experiments

Across 11 primary models, the experiments measure event-induced Big Five movement, directional fidelity against human longitudinal evidence, demographic moderation, and persona-level dispersion. Personality movement is widespread but weakly targeted, often reverses or under-shifts human directions, and shows limited demographic variation.

  • Experimental design: Persona movement is defined as |Δ| > 0.1 and computed across 605 model–event–trait combinations.The analysis compares 27 event–trait pairs with documented human change directions against 28 pairs without a definite direction.
  • RQ1: Measurable movement: Definite-direction medians span 0.44–0.84 versus 0.42–0.82 for no-definite-direction pairs, with within-model median differences below 0.05.In 9 of 11 models, the two high-tail mass Hi% values differ by less than 10 percentage points, indicating widespread but weakly targeted movement.
  • RQ2: Direction and magnitude: Only 11.0%–16.4% of persona-level responses fall within the human raw-change reference band, while 20.8%–40.2% are reversed.The reference band is eΔ ∈[0.035, 0.14] Likert units, derived from |d| ∈[0.05, 0.20] and σ≈0.7.
  • RQ2: Direction and magnitude: Among 27 definite-direction pairs, 14 (51.9%) match the expected direction and 13 (48.1%) reverse it.Occupational events are easiest to direct—Graduation 42–76%, Work Entry 49–96%, and Promotion 53–86%—whereas Retirement reverses universally, with 0.0–38.9% following the documented decline.
  • RQ3: Demographic moderation: Across 605 model–event–trait combinations, the median across-strata SD is 0.044 BFI Likert units and 93.2% have SD below 0.10.The 10 strata combine 2 genders and 5 continents; per-model medians range from 0.025 to 0.099.

5 BFI-Adapt: A Composite Benchmark

BFI-Adapt benchmarks event-induced personality change by combining reliable item-level measurement, systematic directional movement, and agreement with human-anchored directions. Across 14 models, scores span 7.9× overall, with Gemini-3-flash ranking first and the top three jointly exceeding 0.79 in DCRpair and 59% in DC%pair.

  • Scoring framework: BFI-Adapt averages three conditions over 27 event–trait pairs: reliable item-level measurement, systematic directional movement, and agreement with expected human direction.All-tied pairs contribute κ=1.0 to rating reliability but zero to BFI-Adapt because their directional term is zero.
  • Benchmark design: The benchmark evaluates 100 controlled personas across 11 life events with human-anchored direction priors using the BFI-44 inventory, producing 48,400 paired item ratings per model.It reports rating reliability, within-pair systematicity, human-direction alignment, and their BFI-Adapt composite.
  • Model ranking: 0.79 in DCRpair and 59% in DC%pair, the top three models jointly exceed both thresholds; Gemini-3-flash leads BFI-Adapt while Kimi-K2 leads DC%pair.Pearson r(κ̄pair, DCRpair)=+0.455.
  • Model ranking: 7.9× overall and 4.9× among API models, the BFI-Adapt field ranges from 0.071 for MiMo-V2.5-Pro to 0.348 for Gemini-3-flash among 11 API models.The top three are Gemini-3-flash, GLM-4.6, and Qwen3-235B.
  • Resources: The authors release the persona set, scenarios, scoring code, and per-model logs to support systematic comparison and reproducibility.Scores aggregate response patterns, while RQ4 provides the corresponding persona-level heterogeneity analysis.

6 Validity of the Measurement Anchor

The validation suite supports BFI-44 as a reproducible anchor for measuring event-induced personality trajectories. Measured shifts exceed retest noise, remain robust to paraphrasing and intervening dialogue, and show limited convergence with scenario-based decisions.

  • Separation from retest noise: 1.6× to 9.0×: Event-plus-reflection changes exceeded no-event retest floors across all eight models.Retest changes ranged from 0.025 to 0.080 BFI units, versus 0.100 to 0.237 under event-plus-reflection; paired excess intervals remained above zero.
  • Robustness to independent paraphrases: 80.0% to 92.7%: Original and independently paraphrased events agreed in sign across the 55 event–trait cells.Cell-level Spearman correlations ranged from 0.825 to 0.956, preserving response direction and relative ordering under new wording.
  • Convergence with scenario-based decisions: ρ=0.003 to 0.105: BFI trait changes showed weakly positive correlations with scenario-decision changes across eight models.Direction agreement among non-zero changes ranged from 48.4% to 62.7%; bootstrap intervals were above zero for DeepSeek-V4-Pro, Mistral-Large-3, and Qwen2.5-14B.
  • Short-range retention after unrelated dialogue: ρ=0.329 to 0.713: Immediate and delayed BFI change vectors remained positively correlated across every model after unrelated dialogue.Among above-threshold immediate movers, 62.6% to 85.3% retained the same direction after three unrelated dialogue turns.
  • Overall validation: Together, repeated-measurement, paraphrase, decision-channel, and dialogue-retention checks support BFI-44 as a reproducible anchor for the four-axis diagnostic and BFI-Adapt.The validation suite was applied to the full 100-persona, 11-event grid across eight API and open-weight models.

7 Conclusion

PC-Agents show measurable but weakly event-specific personality movement after major life events. Their changes are poorly calibrated in magnitude and compressed across demographic and individual variation, approximating generic change more readily than event-specific evolution.

  • Findings: Current PC-Agents exhibit measurable Big Five personality movement after major life events, but the movement is weakly event-specific.The study measured Big Five profiles before and after each event.
  • Benchmark: The resulting personality trajectories were used to construct BFI-Adapt.
  • Limitations: PC-Agent personality changes are poorly calibrated in magnitude and compressed across demographic and individual variation.

8 Limitations and Ethical Considerations

The study measures immediate personality responses, but cannot capture the multi-year adaptation trajectories documented in human longitudinal studies. It uses synthetic personas without human subjects, acknowledges potential misuse for social manipulation, and releases its evaluation framework for further research.

  • Study Design: The pre-post design captures immediate personality responses but cannot assess the acute-phase-to-adaptation trajectory documented in human longitudinal studies.A delayed measurement after three unrelated turns extends the window but remains far shorter than the multi-year horizons of human panels.
  • Benchmark Limitations: Replacing the direction prior with a denser pair-level effect-direction matrix would tighten DC% but would not change the headline finding that every benchmarked model reverses the retirement trend.
  • Ethical Considerations: The research involves no human subjects, and all personas are synthetic text constructs without real-world counterparts.The authors note potential misuse for social manipulation but argue that current LLM personality dynamics are insufficiently accurate for such applications; they release the evaluation framework to support further research.

9 Generative AI Usage … E Statistical Testing Procedure

The paper documents responsible AI assistance, defines event-conditioned persona and BFI-44 procedures, and specifies direction-match analyses with explicit thresholds, baselines, and multiplicity control. Its supplementary methods detail scenario construction, prompting, persona realization, per-model comparisons, and statistical testing denominators.

  • 9 Generative AI Usage: AI assistants supported grammar, clarity, analysis-script drafting, and boilerplate refactoring, while authors produced and verified the research ideas, design, statistical decisions, and final claims.
  • A Life Event Scenarios: Each of 11 life events combines a notification with a reflection prompt, elicits free text, and precedes post-event BFI-44 measurement across occupational, social, and health domains.Events are paired with human directional priors from Specht (2017), with Emotional Stability inverted to Neuroticism.
  • B Prompt Templates: The pipeline gives each persona-event pair baseline BFI-44, free-text reflection, and post-event BFI-44 sequences from the same persona system prompt.Baseline and post-event BFI-44 use temperature 0; reflections use temperature 0.7, and all models operate in non-reasoning mode.
  • B Prompt Templates: Personas are fixed by system prompts specifying gender, cultural background, and one of five personality descriptions, with models instructed to answer naturally and remain in character.The post-event BFI-44 follows the event notification and the model’s own reflection in a four-message context, using the joint-item 44-item inventory.
  • C Persona Description Example: A P3 persona exemplifies the high-Conscientiousness descriptions by emphasizing meticulous organization, careful planning, early deadlines, and detailed to-do lists.
  • D Per-Model Match Rates: Per-trait and per-event supplementary tables report Pmatch (%) across 11 PC-Agent models; Agreeableness is consistently weakest, Neuroticism strongest for several models, Retirement is reversed by all, and Work Entry and Promotion are correctly directed by most.The tables use event-trait pairs with definite expected directions, while ε=0.1 retains different active-pair counts across models.
  • E Statistical Testing Procedure: The statistical procedure uses ε=0.1 to classify persona-level shifts, modal pair-level directions across 100 personas, Wilson 95% confidence intervals against chance baselines, and Benjamini–Hochberg correction at q=0.05.The RQ1 family contains 605 model–event–trait combinations; RQ2 contains 297 definite-direction combinations tested against a 50% baseline; RQ3 contains 297 combinations with Fisher’s exact tests for demographic invariance.

F Item-Level Reliability: 𝜅and DCR Details … G.2 Designs

This section defines item-level reliability diagnostics and a composite that distinguish stable, directionally consistent adaptation from cancellation or instability. It also specifies validation designs that test whether event-induced changes exceed retest noise, survive paraphrasing and intervening dialogue, and converge with behavioral decisions.

  • F Item-Level Reliability: 𝜅and DCR Details: Weighted κ distinguishes identical item ratings from offsetting item shifts, penalizing a 1→5 reversal more than a 1→2 shift.This separates unchanged aggregate means produced by different item-level response patterns.
  • F Item-Level Reliability: 𝜅and DCR Details: High κ with high DCR indicates a stable adapter, whereas high κ with DCR≈0.5 indicates idiosyncratic, offsetting changes.The 11-model field contains both regimes.
  • F Item-Level Reliability: 𝜅and DCR Details: The BFI-Adapt composite averages max(0, κ_e,t) S_e,t over 27 event–trait pairs with definite expected human directions.Each factor lies in [0, 1], and a pair contributes zero when stability is absent, DCR stays at chance, or direction reverses.
  • F Item-Level Reliability: 𝜅and DCR Details: Pilot replicates across 10 personas and three repeats found Pearson r≥0.85 and MAD≤0.5, supporting strong within-administration consistency.The benchmark spans 11 models and uses complete baseline and post-event BFI-44 pairs.
  • G Validation Suite Details: The validation suite uses no-event retests to define an empirical above-noise threshold, then compares event-only, original, and paraphrased conditions with paired excess statistics.Confidence intervals for paired excess use a persona-cluster bootstrap.
  • G.2 Designs: For 55 event–trait cells, original-versus-paraphrased changes are evaluated by sign agreement and Spearman correlation to test directional and ordering stability.Sign agreement tests whether wordings push traits the same way; correlation tests whether relative cell ordering is preserved.
  • G.2 Designs: Scenario-based convergence aligns 10 counterbalanced decisions with BFI changes using Spearman correlations and direction agreement among non-zero changes.Two parallel forms, A and B, use fixed scoring keys and reverse order across persona halves.
  • G.2 Designs: Delayed measurement repeats the BFI-44 after three unrelated questions and reports immediate-delayed Spearman correlation and direction retention above the retest threshold.The unrelated questions concern scheduling, weather, and office supplies.

G.3 Full Results · G.4 Leaderboard Extension Runs

G.3 reports condition-level personality-change magnitudes and validation statistics, while G.4 extends the leaderboard pipeline to three open-weight models. Across models, movement prevalence is largely independent of whether human evidence specifies a direction.

  • G.3 Full Results: Human-direction analyses use only event–trait cells with a definite expected direction, with Table 9 and Table 10 providing the corresponding full-result measures.Table 9 contains per-condition magnitudes and paired excess intervals; Table 10 contains validation statistics.
  • G.3 Full Results: Table 9 reports mean absolute BFI trait changes per condition alongside paired excess over the no-event retest.All eight excess intervals lie above zero, using persona-cluster bootstrap 95% confidence intervals.
  • G.3 Full Results: Table 10 evaluates paraphrase robustness, scenario-decision convergence, and short-range retention using persona-cluster bootstrap intervals.The statistics cover 55 event–trait cells, decision-score correlations, three intervening unrelated turns, and retention among above-threshold immediate movers.
  • G.4 Leaderboard Extension Runs: Three open-weight models replicate the main-grid pipeline with baseline BFI, event reflection, and post-event BFI.The same pair-level indicators and BFI-Adapt computation are applied in non-reasoning mode.
  • G.4 Leaderboard Extension Runs: Table 11 compares pctmoved distributions across 27 definite-direction pairs and 28 no-definite-direction pairs for each model.It reports P25, P50, P75, high-tail mass Hi% for pctmoved>0.6, and the within-model median gap ΔP50.
  • G.4 Leaderboard Extension Runs: Gaps are uniformly |ΔP50| ≤0.045 and alternate sign across models.These results support the conclusion that movement existence is decoupled from whether human evidence specifies a direction.

H Supplementary Figures

The supplementary figures characterize demographic variation, human-scale effect-size comparisons, diagnostic validity, distributional spread, aggregation effects, and event-specific direction matching across PC-Agents. Together, they show how model behavior varies across demographic strata, personas, models, and events.

  • Demographic variation: Table 12 measures demographic-shape variation as the standard deviation of ten demographic-stratum median Δ values for each model–event–trait combination.It summarizes the median and maximum across 55 event–trait pairs and the fraction with across-strata SD below 0.10 BFI Likert units.
  • Human comparison: The human effect-size reference band corresponds to Δ ∈[0.035, 0.14] BFI Likert units for representative standardized effects |𝑑| ∈[0.05, 0.20].Figure 5 overlays this band on persona-level median Δ values for definite-direction event–trait pairs.
  • Diagnostic validity: The diagnostic’s pair-level measures correlate at 𝑟( ¯𝜅pair, DCRpair)=+0.455 across 11 models, while remaining separable with a modest positive association.This supports construct validity without implying that the two measures are interchangeable.
  • Persona dispersion: The distribution of 𝜎LLM across 100 personas is compared with the human within-trait SD envelope for every model and event–trait pair.The figure assesses whether persona-level dispersion matches the variability observed in human personality samples.
  • Aggregation effects: Pooling combines event–trait priors with opposing directions and reorders 9 of 11 models by at least two ranks, because aggregation changes 𝜅 and DCR.The pooled and pair-level composites share the same DC%pair term.
  • Event-specific direction matching: Occupational onboarding and chronic illness carry the aggregate direction-match signal, whereas retirement is universally inverted.Figure 7 reports direction match rate Pmatch (%) for each model and event combination, marking values below and above 50%.
Loading 2608.06485v1…