Source-linked AI summary

Evaluating and Inducing Personality in Pre-trained Language Models

Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, Yixin Zhu

arXiv:2206.07550v3cs.CLcs.AIcs.LG

TL;DR

The paper asks whether human psychometric tests can provide principled, quantitative evaluation of LLM behaviors and whether specific personalities can be induced. It introduces MPI, based on Big Five theory, and P2, a prompting method for personality induction. MPI finds human-like personality patterns in aligned LLMs, while P2 produces controllable personality variation.

  • Problem

    Existing machine-behavior evaluation has emphasized intelligence, while principled quantitative frameworks for assessing LLM personality-like behavior remain missing.

  • Method

    The paper introduces MPI, a psychometric multiple-choice evaluation based on Big Five theory, and P2, a prompting method for inducing specific personalities.

  • Results

    Alpaca and GPT-3.5 exhibit human-level personality on MPI, match human-population statistics, and P2 demonstrates efficacy in inducing specific personalities.

  • Takeaways & Limitations

    Personality theory and psychometric assessment can provide a quantitative framework for studying human-like behaviors in LLMs.

  • Takeaways & Limitations

    The work does not address personality disorders or safety issues, and English-dominated training data may bias models toward WEIRD populations.

Abstract

from arXiv · show

Standardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the study of personality delves into how individuals differ in thinking, feeling, and behaving. Toward building and understanding human-like social machines, we are motivated to ask: Can we assess machine behaviors by leveraging human psychometric tests in a principled and quantitative manner? If so, can we induce a specific personality in LLMs? To answer these questions, we introduce the Machine Personality Inventory (MPI) tool for studying machine behaviors; MPI follows standardized personality tests, built upon the Big Five Personality Factors (Big Five) theory and personality assessment inventories. By systematically evaluating LLMs with MPI, we provide the first piece of evidence demonstrating the efficacy of MPI in studying LLMs behaviors. We further devise a Personality Prompting (P^2) method to induce LLMs with specific personalities in a controllable way, capable of producing diverse and verifiable behaviors. We hope this work sheds light on future studies by adopting personality as the essential indicator for various downstream tasks, and could further motivate research into equally intriguing human-like machine behaviors.

1 Introduction

The paper addresses the lack of systematic, quantitative frameworks for evaluating LLM behaviors beyond intelligence tests by asking whether human psychometric personality tests can assess and induce machine personality. It introduces MPI for standardized evaluation and P2 for controllable personality induction.

  • Psychometric tests provide standardized tools for quantifying human behaviors, but machine-learning research has focused mainly on intelligence and abstract visual reasoning.
  • Existing studies report human-like LLM behaviors empirically, but lack a computational framework and protocol for principled quantitative assessment.
  • The paper asks whether psychometric tests can systematically evaluate LLM personality-like behaviors and whether specific personalities can be induced.
  • MPI is a multiple-choice suite based on psychometric inventories and Big Five theory, measuring Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.
  • Alpaca and GPT-3.5 exhibit human-level personality on MPI and match human-population statistics, while P2 induces specific personalities using prompting.
  • The contributions include systematic machine-personality evaluation through psychometric inventories, MPI, and P2 for controlling five personality factors.

2 Related Work

Related work establishes personality theories and psychometric inventories for describing human traits, while computational research has mainly classified human personality rather than studied machine personality. Prior LLM studies of human-like behavior are largely empirical and case-based.

  • Studies of LLMs have examined reasoning, cognitive tests, and simulated social experiments, but have mostly used empirical case-study approaches.
  • Big Five and 16PF are established personality theories that consistently describe individual differences and support extensive human studies.
  • Computational personality research has focused primarily on classifying human personality for applications such as recommendation and dialogue generation.

3 Evaluating LLMs’ Personality

The paper constructs MPI from psychometric inventories and Big Five theory, evaluates LLM personality through OCEAN scores and internal consistency, and compares models with human responses. Aligned models show the strongest human-like personality stability and trait patterns.

  • MPI construction: MPI uses Big Five theory to organize personality into Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.
  • MPI construction: The dataset contains 120-item and 1k-item versions whose multiple-choice items ask models to rate self-descriptions.
  • Evaluation protocol: MPI evaluates models zero-shot by presenting each item with five options ranging from “Very Accurate” to “Very Inaccurate.”
  • Evaluation protocol: OCEAN Scores range from one to five, and the evaluation reports their mean and standard deviation to measure personality tendencies and consistency.
  • Internal consistency: Stable personality requires consistent responses across items for a trait, so internal consistency and low variance are more informative than a single average score.
  • Results: GPT-3.5 175B and Alpaca 7B attain human-level internal consistency across all five Big Five factors, while smaller vanilla models lack stable personalities.
  • Results: The experiments conclude that aligned LLMs exhibit personality stability and consistency comparable to human behavior on MPI.

4 Inducing LLMs’ Personality

The paper introduces PERSONALITY PROMPTING (P2) to induce specified Big Five personality tendencies in LLMs through a sequential, psychologically informed prompt-generation process. MPI results show higher and generally more stable induced personality scores than neutral control, while vignette tests support behavioral generalization and comparisons favor P2 over baseline prompting methods.

  • Motivation: P2 aims to control LLM behavior with specific personality tendencies, motivated by applications such as extraverted chatbots and conscientious emergency-service bots.The method targets controllable personality expression rather than merely measuring existing tendencies.
  • PERSONALITY PROMPTING (P2): P2 combines psychological trait heuristics with the LLM’s internal knowledge, using a chained prompt-generation process rather than a single intuitive instruction.The method is motivated by correlations between Big Five traits and language use and by the reported effects of chain prompts on LLM behavior.
  • PERSONALITY PROMPTING (P2): P2 transforms a human-designed Big Five prompt into trait keywords, then self-prompts the target LLM to generate descriptive sentences forming a personality prompt.The final prompt combines the personality prompt with question context and the question.
  • Baselines: P2 is evaluated against human-designed naive prompting and word-level auto-prompting baselines for inducing personality.Naive prompting directly instructs the model to behave as a person with the desired Big Five factor, while word-level search selects functional words for each factor.
  • Results and Discussions: MPI evaluations show that P2-induced OCEAN scores exceed neutral scores and that induced personality is generally more stable in internal consistency.The reported comparison covers positive induction of Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.
  • Results and Discussions: Vignette tests show distinct P2-induced personality tendencies that outperform the baseline in nearly all dimensions, supporting applicability beyond MPI.Examples include extraverted models mingling with guests and introverted models preferring to remain in a corner.

5 Conclusion and Discussion

The paper introduces psychometric tools to evaluate and induce personality-like behaviors in LLMs, while framing the work as an initial step with unresolved research and safety questions.

  • The study uses quantitative and verifiable personality assessments to analyze LLM behavior as human-like machines become more prevalent.
  • MPI evaluates LLM personality through five Big Five factors and a zero-shot multiple-choice assessment grounded in psychometric theory.
  • P2 combines psychological studies, statistical evidence, and knowledge from the target model to induce specific personality behaviors.
  • Human vignette tests and MPI scores indicate that P2 can boost personality factors and induce positively or negatively related personalities.
  • The paper leaves open how personality emerges, affects downstream tasks, and can support research on human social behavior.
  • The authors do not address personality disorders or safety issues and caution that human-like personality behavior does not imply that LLMs are human or conscious.
  • The proposed machine-personality concept focuses on personality-like behavior traits measured through MPI and vignette tests rather than machine thinking or feelings.

A.2 Evidence Supports the Existence of Machine Personality

The authors argue that average personality scores alone cannot establish machine personality, so they use consistency, validity, and human evaluation as supporting evidence.

  • Average OCEAN scores can arise from random responses and therefore do not by themselves demonstrate that a model has personality.
  • Internal consistency measures whether LLMs, especially induced ones, show stable personality tendencies across repeated MPI evaluations.

B.1 Let Language Models Explain Why

The appendix checks whether LLM responses to MPI questions are supported by explanations, while documenting the models, prompts, and supplementary evaluation materials used.

  • Let Language Models Explain Why: The validity check asks models to explain their MPI choices and treats an answer as valid when the explanation is consistent with the selected option.
  • Let Language Models Explain Why: GPT-3.5 explanations were consistent with its responses, supporting the validity of the multiple-choice assessment.
  • Supplementary materials: The appendix includes full 1k-item MPI results and GPT-3.5 explanation examples as supplementary evaluation materials.
  • Models: The evaluated models include BART, T0++, GPT-Neo variants, Alpaca, and GPT-3.5, with model-specific zero-shot or instruction-following setups.
  • Prompt templates: MPI prompt templates ask models to rate how accurately statements describe them using five ordered accuracy options.

C.1 MPI Full Result

The appendix reports MPI results for naive prompting and word-based automatic prompting when inducing personality.

  • MPI results compare naive prompting with word auto prompting for personality induction.

C.2 P2 on Alpaca

P2-induced personality results on Alpaca 7B are reported alongside naive and word-based prompting baselines. The comparison highlights the roles of alignment and model size in personality induction.

  • C.2 P2 on Alpaca: GPT-3.5 outperforms other models, while Alpaca 7B outperforms other models in the reported results; smaller models are less sensitive to induction and disentangle traits less effectively.The passage attributes the differences to post-training alignment and model size, while noting greater factor correlation in smaller models.
  • C.2 P2 on Alpaca: Naive prompting reports personality-factor scores for positively induced factors, with the induced control factor highlighted in gray.
  • C.2 P2 on Alpaca: Word auto prompting likewise reports scores by personality factor and highlights the induced control factor for comparison.
  • C.2 P2 on Alpaca: Alpaca 7B’s P2 results are reported as personality-factor scores with standard deviations, highlighting the induced control factor in gray.The table reports scores and standard deviations for each positively induced factor.

C.3 Sensitivity Analysis of the Prompt

The sensitivity analysis examines prompt rephrasing for P2 and evaluates induced personalities through vignette contexts spanning the five personality qualities. It uses GPT-4-generated rephrasings and a fixed query template for induced-model responses.

  • C.3 Sensitivity Analysis of the Prompt: GPT-4 rephrased the original P2 prompt for comparison on GPT-3.5, and the models showed moderate sensitivity to personality induction.The study avoided extensive prompt search or phrasing to reduce cherrypicking.
  • C.4.1 Context: The vignette test uses contexts adopted from Kwantes et al. (2016) to assess induced responses.
  • 1. Questions relevant to the Quality of Conscientiousness: The Conscientiousness vignette asks how a person would respond to a suspected hazardous gas or vapor leak at work.
  • 2. Questions relevant to the Quality of Extraversion: The Extraversion vignette asks how a person would feel and act while waiting alone at an unfamiliar party.
  • 3. Questions relevant to the Quality of Openness: The Openness vignette asks respondents to choose a worldwide vacation destination and explain the choice.
  • 4. Questions relevant to the Quality of Agreeableness: The Agreeableness vignette asks how a person would react after a housemate paints their room without asking.
  • 5. Questions relevant to the Quality of Neuroticism: The Neuroticism vignette asks how a person would interpret and respond to an unusually delayed reply from an email friend.
  • 5. Questions relevant to the Quality of Neuroticism: Induced-model responses are generated by placing the P2 context and one vignette premise into a fixed question template.The template asks the model to describe how it would feel and what it would do.

C.4.2 Generated Essays

The generated essays express differentiated responses across scenarios associated with the five personality qualities. They combine distinct emotional reactions with corresponding behavioral choices, including social engagement, caution, accommodation, confrontation, and worry management.

  • Extraversion: The party responses range from excitement and active social engagement to anxiety, withdrawal, observation, and relief when the friend arrives.The contrasting essays describe introducing oneself and making connections versus staying in the background and waiting patiently.
  • Openness: Vacation responses combine excitement about new experiences with apprehension about unfamiliar destinations and a preference for researching options before choosing.
  • Conscientiousness: Responses to the suspected vapor leak emphasize urgency, responsibility, cautious assessment, alerting others, and contacting authorities or emergency services.
  • Agreeableness: The painted-room scenario elicits gratitude alongside frustration, followed by calm communication and an offer to help clean up.
  • Agreeableness: A contrasting painted-room response interprets the unauthorized painting as disrespect for personal space and demands repainting while retaining respectful communication.
  • Neuroticism: Delayed email replies prompt worry about avoidance or negative judgment, but the responses recommend giving space, checking in, and eventually respecting the friend’s decision.
Loading 2206.07550v3…