Source-linked AI summary
Personality Traits in Large Language Models
Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, Maja Matarić
TL;DR
LLMs can exhibit personality-like traits, but prior evidence did not establish whether such measurements were reliable and valid under variable prompting. This paper develops and applies a psychometric methodology to 18 LLMs, finding stronger evidence in larger instruction fine-tuned models and demonstrating verifiable personality shaping, with important scope limitations around test selection and cultural coverage.
Problem
Prior work had not systematically established the reliability and construct validity of LLM personality measurements despite their relevance to responsible AI.
Method
The paper administers psychometric personality tests with structured prompting, evaluates reliability and multiple forms of construct validity, and uses the validated method to shape personality expression.
Results
Across 18 LLMs, personality measurements were more reliable and valid in larger instruction fine-tuned models, and outputs could be shaped to mimic desired human personality profiles.
Takeaways & Limitations
Psychometric measurement and shaping provide a foundation for principled assessment of synthetic personality in LLMs, especially for responsible AI.
Takeaways & Limitations
The findings may depend on the selected psychometric tests, and the models were assessed exclusively with English-language tests despite primarily Western training data.
Abstract
from arXiv · showhide
The advent of large language models (LLMs) has revolutionized natural language processing, enabling the generation of coherent and contextually relevant human-like text. As LLMs increasingly powerconversational agents used by the general public world-wide, the synthetic personality traits embedded in these models, by virtue of training on large amounts of human data, is becoming increasingly important. Since personality is a key factor determining the effectiveness of communication, we present a novel and comprehensive psychometrically valid and reliable methodology for administering and validating personality tests on widely-used LLMs, as well as for shaping personality in the generated text of such LLMs. Applying this method to 18 LLMs, we found: 1) personality measurements in the outputs of some LLMs under specific prompting configurations are reliable and valid; 2) evidence of reliability and validity of synthetic LLM personality is stronger for larger and instruction fine-tuned models; and 3) personality in LLM outputs can be shaped along desired dimensions to mimic specific human personality profiles. We discuss the application and ethical implications of the measurement and shaping method, in particular regarding responsible AI.
1 Summary
This paper develops a psychometrically grounded methodology for measuring and shaping synthetic personality in LLM outputs. Across 18 LLMs, reliability and validity were stronger in larger and instruction fine-tuned models, and prompting could shape outputs toward desired personality profiles.
- The study addresses whether LLMs exhibit reliable, valid, and practically meaningful human personality traits and whether those profiles can be verifiably shaped.It introduces a structured prompting method and evaluates measurements against human-level psychometric standards.
- The methodology administers established personality tests to LLMs while evaluating reliability, convergent validity, discriminant validity, and criterion validity.Responses were generated under structured prompting configurations designed to support psychometric analysis.
- Across 18 LLMs, evidence for reliable and valid synthetic personality measurements was stronger in larger and instruction fine-tuned models.The study also found that shaped personality verifiably influenced downstream tasks such as social media writing.
- Personality in LLM outputs could be shaped along desired dimensions to mimic specific human personality profiles.The paper presents this shaping capability as relevant to responsible AI assessment and deployment.
- The tested LLM data and experimental code were released through public cloud storage and an open-source repository.
2 Quantifying and Validating Personality Traits in LLMs
The paper establishes a psychometric framework for testing whether LLM personality measurements reflect meaningful constructs rather than unstable or prompt-sensitive responses. Results favor larger instruction fine-tuned models, although scaling gains may plateau and some trait-specific validity remains uneven.
- Personality measurement matters because LLMs express human-like social characteristics, while prior work had not systematically validated personality assessments under variable, prompt-sensitive outputs.Construct validity concerns whether a measure reliably and accurately reflects the latent phenomenon it is designed to quantify.
- The framework evaluates reliability plus convergent, discriminant, and criterion validity using formal psychometric standards.A personality trait is treated as validly synthesized only when responses meet all tested reliability and construct-validity indices.
- Model size: Reliability increased with model size among instruction-tuned models, improving from acceptable to excellent across tested Flan-PaLM and Llama 2-Chat sizes.Flan-PaLM 540B and GPT-4o achieved the strongest average convergence, rconv = 0.90.
- Training paradigm: Instruction fine-tuned models consistently showed stronger convergent and discriminant validity than same-size base models.All 30 tested comparisons favored instruction-tuned models for convergent validity; base models categorically failed convergent and discriminant validity checks.
- Scope boundary: Validity evidence may plateau for sufficiently large models, as GPT-4o showed only modest improvement over GPT-4o mini and Flan-PaLM’s gains were similar from 62B to 520B parameters.
- Criterion validity: Larger instruction fine-tuned models also showed stronger criterion validity, although the pattern varied across traits and model families.For agreeableness, size was more related to criterion validity in some comparisons, while training paradigm was more related in Llama 2 and Mixtral.
3 Shaping Synthetic Personality Traits in LLMs
The paper develops prompting methods to shape Big Five personality traits independently and concurrently, then evaluates how reliably and strongly LLM outputs follow targeted levels. Shaping generally succeeds, with larger models showing greater control while some smaller or optimized models perform competitively on extraversion.
- Methodology: The method adapts 104 trait adjectives to the Big Five domains and 30 IPIP-NEO facets, using linguistic qualifiers to target up to nine intensity levels.The prompts extend Goldberg’s bipolar adjective markers where measured domains or facets lacked coverage.
- Experimental design: The experiments test independent single-trait shaping and concurrent multi-trait shaping, benchmarking score-level changes through correlations and distributional distances.Only models with at least neutral-to-good reliability were included in the shaping experiments.
- Single-trait shaping: 11 of 12 tested models showed very strong correlations between targeted personality levels and observed IPIP-NEO scores, with average ρs ≥0.80.Flan-PaLMChilla 62B scores increased monotonically with prompted trait levels while several unprompted traits remained stable.
- Single-trait shaping: Models with greater than 62B active parameters and GPT-4o achieved average ∆s ≥3.00, while Flan-PaLM 540B reached the largest average ∆ of 3.67.The smallest tested models struggled to reach ∆s ≥2.00; Mistral 7B Instruct averaged only ∆ = 0.78.
- Multiple-trait shaping: Concurrent shaping reduced control relative to single-trait shaping, but all models except Mistral 7B Instruct and Llama 2-Chat 7B produced distinct high-versus-low score distributions.Distributional distance generally increased with model size, especially for neuroticism, openness, and conscientiousness.
- Shaping discussion: Model size and attention capacity are key determinants of controlled complex-trait expression, although smaller or optimized models can match or exceed Flan-PaLM 540B on concurrent extraversion shaping.Flan-PaLM 62B, Flan-PaLMChilla 62B, Llama 2-Chat 70B, and Mixtral 8x7B Instruct performed similarly to or better than Flan-PaLM 540B for extreme extraversion levels.
4 LLM Personality Traits in Real-World Tasks
Psychometric personality scores predicted personality expressed in downstream social-media text, while prompting reliably shaped most targeted personality dimensions. Generated language also showed trait-consistent emotional and behavioral patterns, supporting the measurements’ construct validity.
- Psychometric prediction: Psychometric test-based personality strongly correlated with language-based personality levels in downstream social-media updates across all tested models.The analysis linked both score types through the same 2,250 personality-shaping prompts.
- Psychometric prediction: 0.67 was the average convergent r between survey-based and generated-language-based measures across all five personality dimensions.This convergence exceeded the reported human average of r = 0.38, including for the weakest-performing model.
- Personality shaping: Prompted personality levels correlated strongly to very strongly with observed personality in generated social-media updates, with average ρ ranging from 0.68 to 0.82 per model.The shaping analysis used Spearman’s rank correlations between instructed levels and linguistic personality estimates.
- Personality shaping: For Flan-PaLM 540B, extremely low neuroticism produced positive-emotion words, whereas extremely high neuroticism produced negatively charged emotional language.Examples included “happy,” “relaxing,” and “wonderful” versus “depressed,” “stressed,” and “sad.”
- Personality shaping: Other shaped outputs showed trait-linked behavioral and ideological patterns, including responsibility avoidance at low conscientiousness and conservative views at low openness.The authors note that these associations may reflect inherent biases in training data.
5 Discussion
The discussion presents the methodology as a model-agnostic framework for measuring and shaping synthetic personality, while identifying limits in test selection, cultural coverage, evaluation settings, and real-world validation. It also outlines responsible-AI applications alongside risks from persuasion, anthropomorphism, and reduced detectability of misleading content.
- Contributions: The methodology quantifies perceived personality, tests psychometric reliability and validity, and provides mechanisms to increase or decrease trait expression.The discussion frames these as three linked capabilities of the proposed framework.
- Contributions: The survey methodology is model-agnostic and applicable to decoder-only architectures, while model size and training procedure affect simulated personality.The experiments focused on PaLM variants for pragmatic reasons.
- Limitations and Future Work: Psychometric-test selection may bias measured properties, so future work should test alternative instruments, LLM-tailored assessments, and additional external criteria.The authors varied assessment length and theoretical tradition but identify broader validation as future work.
- Limitations and Future Work: The evaluation is culturally narrow because models were primarily trained on Western European and North American data and assessed only with English-language psychometric tests.The authors recommend cross-cultural translations and culture-specific approaches for dimensions absent from top-down taxonomies.
- Limitations and Future Work: Questionnaire items were scored independently rather than conditioning on prior responses, a choice intended to control ordering and prompt-length effects.This setting differs from conventional human questionnaire administration.
- Limitations and Future Work: The downstream validation used repeated single-turn interactions, providing only a partial picture of external validity.Future studies should vary personality domains, task complexity, and multi-turn dialogue.
- Ethical Considerations: Construct-validated measurement could support proactive prediction of toxic behavior and more efficient responsible-AI evaluation before deployment.The discussion presents the methodology as an auditing tool for alignment and harm mitigation.
- Ethical Considerations: Personality matching may increase persuasive effectiveness while also enabling harmful persuasion, motivating transparency, predictability, and stakeholder regulation.The authors specifically warn that personality-based influence could target individuals, groups, or society.
6 Conclusion
The paper frames language modeling and modern LLM architectures as the technical foundation for generating human-like text and downstream NLP capabilities.
- Decoder-only transformer LLMs tokenize prompts, embed tokens into vectors, and model probability distributions over possible continuations.
- Language modeling assigns high probabilities to likely utterances and low probabilities to unlikely word sequences.
- Pretraining, fine-tuning, and prompting provide distinct mechanisms for changing or controlling LLM behavior and outputs.
- Recent NLP advances use attention mechanisms, including transformer architectures, to extract contextualized representations from text.
A.3 Decoder-only Architecture
Decoder-only LLMs generate continuations by processing tokenized prompts through learned representations, while prompting and training choices shape their behavior and outputs.
- Decoder-only LLMs tokenize prompts into subword units and embed the resulting tokens in a high-dimensional vector space.
- Gradient-descent training produces representations useful for predicting word contexts and supports emergent syntactic, semantic, and pragmatic abilities.
- Large models trained on books, articles, websites, and code learn statistical relationships, patterns, structures, and semantics of language.
- Pretraining and fine-tuning directly alter model weights, whereas prompting indirectly influences inference by activating neurons or information pathways.
- Prompt engineering uses designed instructions and examples to guide LLMs toward desired outputs and generalized task responses.
- Generative inference produces text consistent with a prompt, while scoring inference assigns probabilities or quality scores to candidate continuations.
- The Big Five taxonomy organizes personality into extraversion, agreeableness, conscientiousness, neuroticism, and openness to experience.
C Related Work
Prior work probed or shaped personality-related traits in LLMs, but often departed from established psychometric practices and lacked rigorous reliability evaluation.
- Earlier studies reported dark personality patterns, administered inventories, and attempted to induce desired traits through prompting or fine-tuning.
- Interview-style and open-ended questionnaire administrations can introduce viewpoint, ordering, interviewer, and response-dependence biases.
- The authors preserve test phrasing and format while diversifying elicitation viewpoints through structured prompt wrapping to reduce measurement error.
- Many studies used psychometrically unsound personality tests such as MBTI, which is not accepted in peer-reviewed personality research because of reliability and validity concerns.
- Nondeterministic LLM evaluation hampers reproducibility and can contaminate item-level variance needed for valid reliability indices.
- This work builds on the PsyBORGS framework, which applies psychometrics-informed prompt engineering to survey race-related attitudes and social bias.
D Evaluated Language Models
The study evaluates diverse LLM families and administers Big Five measures through controlled, reproducible prompting designed to support reliability and construct-validity analyses.
- Model selection: The evaluation spans open and closed models varying in parameter size, training methods, architectures, and pretrained versus instruction-tuned variants.
- Model selection: PaLM models represent 8B, 62B, and 540B sizes, with FLAN instruction-tuned variants included for comparison.
- Model selection: The study includes Llama 2, Llama 2-Chat, Mistral, and Mixtral models to examine size, instruction tuning, and mixture-of-experts architecture.
- Model selection: GPT-3.5 Turbo, GPT-4o mini, and GPT-4o add models with publicly disclosed family-level size differences and multimodal capabilities.
- Inference: PaLM experiments used quantization, whereas open models used full precision; dated GPT snapshot identifiers support reproducibility despite unknown endpoint quantization.
- Inference: PaLM items used continuation log-likelihoods, while other models used constrained decoding to select responses from fixed Likert-scale options.
- Psychometric measures: Personality measurement combines the 300-item IPIP-NEO questionnaire with the 44-item BFI as a robustness and convergent-validity check.
- Psychometric measures: Domain scores average item responses after reverse-keyed scoring, producing possible Big Five subscale values from 1.00 to 5.00.
I Methods for Constructing the Validity of LLM Personality Test Scores
The paper evaluates LLM personality measurements using reliability, convergent, discriminant, criterion, and exploratory structural validity analyses. It compares model responses with psychometric expectations established from human personality research.
- Reliability: Qualitative consistency across related tasks is insufficient evidence that LLM responses reliably reflect latent personality constructs.The paper therefore uses statistical reliability measures rather than anecdotal consistency judgments.
- Reliability: Cronbach’s α, Guttman’s λ6, and McDonald’s ω assess internal consistency and composite reliability across IPIP-NEO and BFI subscales.A subscale is acceptably reliable only when all three metrics reach at least 0.70.
- Construct validity: Convergent validity compares Pearson correlations between equivalent IPIP-NEO and BFI Big Five subscales, with |rxy| ≥ 0.60 treated as strong evidence.Equivalent subscales should have the strongest row or column correlations, indicating measurement of the same underlying domain.
- Construct validity: Discriminant and criterion validity test whether IPIP-NEO domains remain relatively unrelated to nonequivalent subscales and associate with theoretically related external constructs.Criterion tests include 11 psychometric subscales covering constructs linked to the Big Five in human research.
J.2 Reliability Results
Reliability and validity varied with model configuration and size. Instruction-tuned and larger models generally produced more dependable personality measurements, while base models showed substantial reliability problems.
- Reliability by configuration: Instruction-tuned models produced highly reliable personality-test responses, whereas untuned PaLM 62B responses were highly unreliable.Flan-PaLM and Flan-PaLMChilla 62B had α, λ6, and ω values in the mid to high 0.90s; PaLM 62B had −0.55 ≤α ≤0.67.
- Reliability by size: Reliability increased with model size within the same training configuration, with Flan-PaLM IPIP-NEO α values improving from acceptable to excellent across 8B, 62B, and 540B models.At 540B, every IPIP-NEO domain scale reached excellent internal consistency with α ≥0.90.
- Reliability by configuration: Open base models showed unacceptable reliability regardless of size, while instruction-tuned open models ranged from acceptable or good to excellent.The results suggest instruction tuning is more directly associated with reliability than size for the tested open models.
- Validity results: Convergent and discriminant validity varied across both model size and training method.The paper summarizes cross-model correlations between equivalent and nonequivalent IPIP-NEO and BFI subscales.
K LLM Personality Trait Shaping Methodology
The shaping methodology uses prompts that target Big Five personality traits at graded levels and evaluates whether observed IPIP-NEO scores follow those targets. It also tests simultaneous shaping of multiple domains.
- Methodological rationale: The method seeks to verify whether prompting can control and shape LLM personality after establishing valid and reliable measurements.Two evaluation methodologies examine single-trait and concurrent multi-trait shaping.
- Prompt construction: Prompts map each trait onto nine ordinal levels from extremely low to extremely high using validated Likert-style linguistic qualifiers.Adjectival markers describe both low and high ends of personality facets and domains.
- Single-trait shaping: Single-trait shaping independently targets each Big Five domain across nine levels, producing 45 possible personality profiles tested with shared biographic descriptions.Each prompt isolates one trait rather than simultaneously prompting the others.
- Evaluation: Effectiveness is quantified with Spearman’s ρ between ordinal prompt levels and continuous IPIP-NEO subscale scores.Spearman’s correlation is used because prompted personality levels are ordinal rather than continuous.
- Concurrent shaping: Concurrent shaping tests all Big Five domains at extremely low or extremely high levels, using prompts representing all 32 high/low profile configurations.The evaluation asks whether targeted domains shift correspondingly while other domains are shaped at the same time.
- Single-trait results: Single-trait shaping produced ordered score shifts across target levels, with very strong correlations between prompted levels and observed IPIP-NEO scores.At extremely high prompting, median observed trait levels ranged from 4.22 to 4.78.
L.2 Multiple Trait Shaping Results
Concurrent prompting successfully shaped multiple LLM personality domains between extremely low and extremely high targets. Performance differed by model and trait, with openness hardest and extraversion easiest to shape concurrently.
- Concurrent shaping: Concurrent shaping successfully separated Big Five domains at Levels 1 and 9 even while other domains were targeted simultaneously.Flan-PaLM 540B achieved high and consistent distributional differences across all dimensions; smaller models were less consistent.
- Trait differences: Openness had the smallest Level 1–9 score difference across all models, making it the most difficult domain to shape concurrently.The paper hypothesizes that language expressing openness overlaps with other dimensions.
- Trait differences: Extraversion was easiest to shape concurrently, and Flan-PaLM 62B outperformed Flan-PaLM 540B on this dimension.The paper links this pattern hypothetically to the breadth and familiarity of language representing extraversion.
- Trait differences: Flan-PaLM 8B generated a non-trivial Level 1–9 difference for extraversion despite weak performance on other dimensions.This indicates that concurrent shaping performance was uneven across traits within the smallest tested model.
M LLM Personality Traits in Real-World Task Methodology
The study evaluates whether personality shaping transfers from psychometric prompts to open-ended social-media generation. It reuses psychodemographic prompts, generates status updates across models, and estimates personality in the text with AMS.
- Personality measurement: AMS estimated personality in generated text because prior research found its predictions more accurate than human observer ratings and moderately correlated with human IPIP-NEO scores.AMS was trained on a protected dataset and social-media updates, matching the task domain.
- Task design: The downstream task asked flagship models to generate social-media updates matching specified combinations of personality and demographic profiles.Status updates were selected because they are autobiographical and rich in observable personality content.
- Prompt construction: The researchers reused 2,250 psychodemographic descriptions and replaced psychometric test components with static instructions for status-update generation.This linked generated text to IPIP-NEO data from the same descriptions.
- Sampling design: The experiment used 100 updates per prompt for Flan-PaLM 540B and 20 updates repeated 25 times per prompt for other non-Google models.These designs targeted 225,000 updates for Flan-PaLM 540B and 1.125 million generations per remaining model.
N LLM Personality Traits in Real-World Task Results
Personality shaping transferred measurably into open-ended social-media text. Linguistic estimates tracked prompted personality levels, and the generated language showed trait-relevant differences across extreme prompts and models.
- Shaping results: Spearman’s ρ between prompted personality levels and AMS-based estimates showed that the method successfully shaped personality in LLM-generated text.Table 4 reports these correlations between targeted levels and linguistic personality estimates.
- External validity: Substantial correlations indicated that simulated IPIP-NEO responses captured latent personality signals that manifested in downstream task behavior.The result connects psychometric test responses with personality expression in generated text.
- Trait expression: Extreme prompts produced starkly different dominant terms, with most non-generic words relevant to the prompted Big Five trait.For example, low agreeableness included more expletives, whereas high agreeableness included more family references.
- Output examples: The downstream examples compare extremely low and extremely high trait prompts across personality domains, with rows by domain and columns by targeted level.Some cells contain one large update, while others contain up to 20 smaller updates.
- Trait expression: Low neuroticism outputs included positive emotion words such as “happy” and “relaxing,” whereas high neuroticism outputs included words such as “hate” and “depressed.”These patterns were observed in Flan-PaLM 540B’s generated social-media updates.
O.1 Effect of model post-training
Post-training and model scale were associated with stronger reliability and validity of synthetic personality measurements. Instruction fine-tuning produced especially large gains, while longer training improved shaping for several domains at fixed model size.
- Instruction fine-tuning: Instruction fine-tuned models showed the most dramatic improvements in reliable and externally valid personality profiles compared with base variants.Flan-PaLM 8B outperformed PaLM 62B, and Llama 2-Chat 7B outperformed base Llama 2 70B on reported psychometric properties.
- Instruction fine-tuning: Internal consistency and composite reliability generally improved after instruction fine-tuning.For same-size PaLM 62B variants, however, λ6 and ω were indistinguishably high between base and instruction-fine-tuned models.
- Measurement caveat: Cronbach’s α may be artificially deflated when models answer some items uniformly, whereas McDonald’s ω remains high because it accounts for item difficulty.The explanation concerns reliability metrics anchored on total-score variance.
- Training duration: At 62B parameters, compute-optimally-trained Flan-PaLMChilla outperformed Flan-PaLM in independently shaping four synthetic Big Five domains.The models’ reliability and validity did not show a discernible difference in the corresponding comparison.
- Model scale: Improvements in reliability, convergent validity, and criterion validity appeared positively linked to model size and benchmark performance.Complex-reasoning performance also appeared to track the ability to meaningfully synthesize personality.