Source-linked AI summary
PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits
Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, Jad Kabbara
TL;DR
The paper addresses limited evidence about whether personalized LLM agents consistently express assigned personality traits. It evaluates GPT-3.5 and GPT-4 personas through Big Five Inventory responses, story writing, psycholinguistic analysis, and human and automatic evaluation, finding aligned self-reports and linguistic patterns, while AI-authorship awareness reduces personality-prediction accuracy.
Problem
Limited research evaluates whether personalized LLM agents accurately and consistently reflect specific personality traits in their behavior and language.
Method
The study creates GPT-3.5 and GPT-4 personas with Big Five profiles, administers the BFI and story-writing task, and evaluates outputs with psycholinguistic, human, and automatic analyses.
Results
LLM personas consistently tailor BFI answers and writing to assigned traits, while human judges predict traits with varying accuracy that decreases when AI authorship is disclosed.
Takeaways & Limitations
LLM personas can express assigned personality profiles through self-reports and representative linguistic behavior, and human perception of those traits is affected by AI-authorship awareness.
Takeaways & Limitations
The study primarily focuses on closed GPT models, while LLaMA 2 outputs were unsuitable for human evaluation because of repetition, instruction-following failures, and explicit personality references.
Abstract
from arXiv · showhide
Despite the many use cases for large language models (LLMs) in creating personalized chatbots, there has been limited research on evaluating the extent to which the behaviors of personalized LLMs accurately and consistently reflect specific personality traits. We consider studying the behavior of LLM-based agents which we refer to as LLM personas and present a case study with GPT-3.5 and GPT-4 to investigate whether LLMs can generate content that aligns with their assigned personality profiles. To this end, we simulate distinct LLM personas based on the Big Five personality model, have them complete the 44-item Big Five Inventory (BFI) personality test and a story writing task, and then assess their essays with automatic and human evaluations. Results show that LLM personas' self-reported BFI scores are consistent with their designated personality types, with large effect sizes observed across five traits. Additionally, LLM personas' writings have emerging representative linguistic patterns for personality traits when compared with a human writing corpus. Furthermore, human evaluation shows that humans can perceive some personality traits with an accuracy of up to 80%. Interestingly, the accuracy drops significantly when the annotators were informed of AI authorship.
1 Introduction
The paper examines whether LLM-based agents can express assigned Big Five personality traits consistently and whether those traits are detectable in their generated narratives. It addresses this gap through personality testing, story analysis, and human and LLM evaluation.
- Motivation: Personalized LLM agents are increasingly used and studied as believable human-like characters, but their consistent human-like behavior remains an unsubstantiated assumption.Prior work has explored generative agents, personality expression, benchmarks, prompting, and editing, without fully establishing consistent trait expression.
- Research focus: The study investigates whether LLM personas can express Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience.It uses the Big Five model and defines personas through prompted trait configurations.
- Research questions: The paper first tests whether LLM personas reflect assigned personalities on the 44-item Big Five Inventory, then examines whether their narratives indicate those traits.The narrative task follows the initial personality assessment because the authors report promising results in that first exploration.
- Research questions: It asks what linguistic patterns appear in persona-generated stories, how humans and LLMs evaluate those stories, and whether they can infer the writers’ personality traits.These questions cover linguistic behavior, story quality, and personality perception.
2 Experiment Design
The experiment creates Big Five-conditioned personas, has them complete a personality inventory and write stories, and evaluates the resulting text through linguistic, human, and LLM-based analyses.
- Workflow: The study creates personas with distinct Big Five traits, administers the BFI, prompts story writing, analyzes language with LIWC, and evaluates stories with human and LLM raters.The workflow also asks evaluators to infer the personas’ assigned traits from their stories.
- Persona construction: 320 personas are simulated for each model family by using 10 LLM personas for every combination of binary Big Five personality types.The main reported models are GPT-3.5 and GPT-4, with temperature set to 0.7 and other parameters at default settings.
- Personality assessment: Each persona completes the 44-item BFI, and responses are aggregated into five personality scores.Only responses following the required “(x) y” format are accepted, with agreement recorded on a 1–5 scale.
- Story generation: Personas are prompted to write personal stories in 800 words without explicitly mentioning their personality traits.The restriction is intended to prevent direct revelation of hidden attributes during subsequent assessment.
- Linguistic analysis: LIWC-225 features are correlated with assigned traits and compared with human Essays-dataset writing to examine shared linguistic markers.The human corpus contains short essays with self-reported binary Big Five traits, although its stream-of-consciousness prompt differs from the study’s personal-story prompt.
- Evaluation: Human and LLM raters assess 32 GPT-4 stories across six dimensions and predict the writer’s Big Five traits on 1–5 scales.Human raters evaluate readability, personalness, redundancy, cohesiveness, likeability, and believability; stories are sampled to exclude explicit trait mentions.
3 Results
The results indicate that assigned Big Five traits shape personas’ self-reported assessments and story language, while human and LLM evaluations reveal both perceptible personality signals and source-dependent judgments.
- 3.1 RQ1: Behavior in BFI Assessment: The BFI results show statistically significant differences across all five personality traits, with large effect sizes for both GPT-3.5 and GPT-4 personas.For GPT-3.5, reported effects include EXT d = 7.81, AGR d = 5.93, CON d = 1.56, and NEU d = 1.83; the GPT-4 values are partially reported in the supplied passage.
- 3.2 RQ2: Linguistic Patterns in Writing: Assigned personality types considerably influence personas’ linguistic styles, including curiosity-related language for Openness and anxiety, negative-tone, and mental-health lexicons for Neuroticism.The analysis uses LIWC features and point-biserial correlations between linguistic features and assigned binary personality types.
- 3.2 RQ2: Linguistic Patterns in Writing: GPT-4 personas show greater overlap with human linguistic correlations than GPT-3.5 personas, especially for Conscientiousness and Openness.Overlap is 11/31 versus 1/31 for Conscientiousness and 17/36 versus 2/36 for Openness, comparing GPT-4 with GPT-3.5 respectively.
- 3.3 RQ3: Story Evaluation: GPT-4 stories receive high ratings for readability, cohesiveness, and believability, while human evaluators rate personalness highly but likeability lower.The evaluations suggest fluent, cohesive, believable, and personal stories that are not necessarily engaging or enjoyable to read.
- 3.3 RQ3: Story Evaluation: Informing human evaluators that stories are AI-generated significantly lowers perceived personalness, while other human ratings remain comparatively consistent.GPT-3.5 raters also rate readability higher and personalness lower when informed; GPT-4 ratings vary minimally across conditions.
- 3.4.1 Personality Prediction: Human personality-prediction accuracy reaches 0.84 for Extraversion and 0.69 for Agreeableness with majority voting, whereas GPT-4 reaches 0.97 for Extraversion individually.Human accuracy decreases when evaluators know the stories are AI-generated, and informed human–BFI correlations weaken, with Openness becoming non-significant.
4 Related Work
Prior work studies human-like LLM behavior, personality prediction, and personalized dialogue, but has left the linguistic behavior and human perception of personality-conditioned LLM content comparatively underexplored.
- 4.2 LLMs as Simulated Agents: Research has examined LLMs as emerging agents capable of human-like reasoning, role-playing, and social-science experimentation through prompting.The related work frames these systems as increasingly agentic while noting a remaining gap concerning their personality traits and effects.
- 4.3 Personality in NLP: Personality research in NLP includes text-based personality prediction, digital-footprint prediction, personalized dialogue generation, and stylistic transfer.These lines of work provide background for studying personality-conditioned language generation and perception.
- 4.3 Personality in NLP: Previous work had not examined the linguistic behavior of LLM personas or human perception of personality-conditioned LLM content together.The paper addresses this gap using story evaluation and personality prediction with both human and LLM evaluators.
5 Conclusion
The study finds that GPT-3.5 and GPT-4 personas can express assigned personality traits across BFI responses and writing, with recognizable linguistic patterns. Human judgments of personality vary by trait and decline when AI authorship is disclosed, while general story-quality judgments remain stable.
- GPT-3.5 and GPT-4 personas consistently tailor BFI answers and writing features to their assigned personality traits.
- Each personality trait is associated with distinct representative linguistic behavior in persona-generated stories.
- Human evaluators’ judgments of readability, redundancy, cohesiveness, likeability, and believability remain stable regardless of AI-authorship awareness.
- Human judges can predict expressed personality traits with varying accuracy across traits, but accuracy decreases when AI authorship is disclosed.
Limitations
The study’s evidence is constrained by its model coverage, dataset size, task and language settings, and subjective personality-perception evaluation. These boundaries limit how broadly its findings can be generalized beyond the tested conditions.
- The study mostly evaluates closed GPT models; LLaMA 2 was excluded from human evaluation because its outputs repeated content and followed instructions poorly.
- The dataset is relatively small because larger human evaluations would be costly, although analyses use 160 instances per trait rather than 10 per personality type.
- The evaluation covers personality assessment and writing in English, not naturalistic interaction, collaboration, or other languages.
- Personality perception is subjective, and the effects of annotators’ personality and background on prediction accuracy require deeper investigation.
Ethical Considerations
The study follows ethical procedures for human annotation while emphasizing transparency and the potential misuse of personalized LLMs. Its findings also motivate disclosure because awareness of AI authorship changes evaluators’ responses.
- Ethical procedures: The study received IRB Exempt status, followed the ACL Code of Ethics, and compensated Prolific annotators at $15 per hour.The researchers also provided study instructions and prompts in the appendix or GitHub repository.
- Risks and transparency: Personalized LLMs may be misused to target individuals, communities, or societies, despite their ability to generate enticing interactions.The paper avoids taking a general stance on AI applications while calling for transparency about AI usage.
- Risks and transparency: Human evaluators reported lower personalness when informed that stories were AI-generated, while readability, redundancy, cohesiveness, likeability, and believability remained consistent.The authors connect this result to the importance of ethical disclosure to users.
- Study scope: The paper frames story writing as a vehicle for scientific inquiry into LLM expressivity and human perception rather than a specific application.It urges stakeholders to remain vigilant and mitigate potential AI misuse.
C Personality Ratings
The study compares personality ratings from humans and LLM evaluators for GPT-4-generated stories. Ratings are centered around a neutral midpoint for balanced positive and negative personality labels, but evaluator behavior differs by model and trait.
- Experimental setup: 32 personas represent 32 personality types, with 16 positive-label and 16 negative-label personas for each personality.This design would ideally produce average ratings close to 3.
- Evaluation results: GPT-3.5 and GPT-4 evaluator averages are closer to 3 than human averages for Extraversion.For traits other than Extraversion, GPT-4 averages appear consistently farther from 3 than human and GPT-3.5 averages.
- Evaluation results: GPT-3.5 and GPT-4 evaluators rate GPT-4-generated stories across five personality traits using mean Likert scores and standard deviations.The table distinguishes informed and uninformed conditions for human or LLM evaluators and uses temperature 0.
- Agreement analysis: Inter-annotator agreement is reported separately for six story-quality metrics and five personality traits using Krippendorff’s α.The subscript percentage denotes the average share agreeing with the most-voted rating.
D Story Evaluation Details
The story evaluation filters out explicit personality terms, recruits English-speaking U.S. Prolific workers, and collects quality, personality, and optional comment judgments. Agreement among annotators is low, and the study documents the participant sample and evaluation materials.
- Story selection: A lexicon-based filter removes stories containing explicit personality terms before human evaluation.The lexicons cover variants of Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience.
- Human evaluation: The study recruits U.S.-based Prolific workers whose first language is English and whose approval rates range from 99% to 100%.Thirty-two stories are divided into four batches of eight.
- Human evaluation: Each annotator rates readability, personalness, redundancy, cohesiveness, likeability, and believability, then answers five personality questions.The study also provides consent materials and an optional comment section.
- Agreement and participants: Low inter-annotator agreement occurs across both six story-quality metrics and five personality traits.The paper reports this challenge as consistent with earlier social-computing labeling findings.
- Agreement and participants: The evaluation sample includes 39 unique participants, all living in the United States, with age, sex, and ethnicity distributions reported separately.Thirty-seven participants were born in the USA and two in Nigeria.
E.1 Temperature
The study tests GPT-3.5 and GPT-4 evaluators at different temperatures for GPT-4-generated stories. Higher temperature is associated with lower ratings and greater variance, so the experiment uses temperature 0 for reproducibility.
- Temperature effects: Larger temperature produces greater variance among three LLM evaluators.This indicates that sampling temperature affects both rating level and evaluator consistency.
- Experimental choice: The experiment sets temperature to 0 to make evaluation results more deterministic and reproducible.Table 8 reports mean Likert ratings and standard deviations for GPT-4-generated stories at different temperatures.
- Experimental setup: The temperature comparison evaluates GPT-4-generated personal stories using mean Likert scales and standard deviations for each attribute.The table covers LLM evaluation results under different temperature settings.
F LLaMA 2 Results in BFI Scores
LLaMA 2 personas showed statistically significant differences across all Big Five dimensions, but with smaller effect sizes than GPT personas. Their writings also exhibited trait-linked LIWC patterns that varied by personality profile.
- BFI scores: Statistically significant differences appeared across all personality dimensions in LLaMA 2’s BFI assessment, but effect sizes were much smaller than GPT results.The high and low labels represent the binary personality traits for each dimension.
- Linguistic patterns: Spearman’s ρ analysis linked LLaMA 2 personas’ five-point BFI results with LIWC features.The analysis reports correlations between LIWC features and the personas’ BFI scores.
- Linguistic patterns: Extroverted and agreeable personas used more positive, social, prosocial, or affiliative language, while showing fewer conflict-related features.Reported associations include social and prosocial language for extroversion and reduced conflict language for agreeableness.
G.2 GPT-4 Personas
GPT-4 persona writings displayed trait-associated linguistic patterns across social orientation, affect, analytical language, and lexical complexity. The reported correlations connected personality profiles with distinct narrative features.
- Analytical language and openness: GPT-4 persona writings contained more analytical thinking and insight-related language, with open-minded personas using longer words and more words per sentence.Analytical and insight features correlated at ρ = 0.27 and ρ = 0.25; BigWord and WPS correlated at ρ = 0.29 and ρ = 0.26.
- Extroversion: Extroverted GPT-4 personas used more positive tone, social references, affiliations, exclamations, and drive-related words.Reported correlations included tone: ρ = 0.47, affiliation: ρ = 0.48, exclam: ρ = 0.23, and Drive: ρ = 0.34.
- Agreeableness: Agreeable personas used more positive emotion, friendship, affiliation, and “we” language, alongside fewer work and conflict references.The reported associations included friend: ρ = 0.47, work: ρ = −0.48, conflict: ρ = −0.31, and we: ρ = 0.20.