Source-linked AI summary
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
Huiqi Zou, Pengda Wang, Zihan Yan, Tianjun Sun, Ziang Xiao
TL;DR
The paper asks whether self-report personality scales validly measure LLM-based chatbot personality, given that these scales were developed for humans and may not reflect interaction-based perceptions. It evaluates 500 personality-designed chatbots through chatbot self-reports, human task interactions, perceived personality ratings, and interaction-quality measures. Self-reports show moderate internal validity but weak alignment with human-perceived personality and interaction quality, supporting more contextualized and interactive evaluation.
Problem
Self-report scales for LLM-based chatbot personality rely on unvalidated assumptions about alignment with human perceptions and transferability from human psychological assessment.
Method
The study evaluates 500 distinct chatbots by comparing three self-report personality scales with human interaction-based personality ratings and UEQ interaction-quality assessments.
Results
Self-report scales show moderate convergent and discriminant validity but poor alignment with human-perceived personality and weak correlations with interaction quality.
Takeaways & Limitations
Personality evaluation for LLM-based chatbots should incorporate human perceptions and task-driven interactions rather than relying solely on self-report measures.
Takeaways & Limitations
The study emphasizes that chatbot personality should be treated as an emergent, dynamically co-constructed phenomenon within human–chatbot exchanges.
Abstract
from arXiv · showhide
A chatbot's personality design is key to interaction quality. As chatbots evolved from rule-based systems to those powered by large language models (LLMs), evaluating the effectiveness of their personality design has become increasingly complex, particularly due to the open-ended nature of interactions. A recent and widely adopted method for assessing the personality design of LLM-based chatbots is the use of self-report questionnaires. These questionnaires, often borrowed from established human personality inventories, ask the chatbot to rate itself on various personality traits. Can LLM-based chatbots meaningfully "self-report" their personality? We created 500 chatbots with distinct personality designs and evaluated the validity of their self-report personality scores by examining human perceptions formed during interactions with these chatbots. Our findings indicate that the chatbot's answers on human personality scales exhibit weak correlations with both human-perceived personality traits and the overall interaction quality. These findings raise concerns about both the criterion validity and the predictive validity of self-report methods in this context. Further analysis revealed the role of task context and interaction in the chatbot's personality design assessment. We further discuss design implications for creating more contextualized and interactive evaluation.
1 Introduction
The paper examines whether self-report personality scales validly measure LLM-based chatbot personality, comparing chatbot responses with human perceptions and interaction quality. Across 500 designed chatbots, self-reports aligned poorly with human-perceived personality and interaction quality, motivating more interaction-grounded evaluation.
- Self-report evaluation assumes that designed traits match human-attributed traits and that human psychological scales transfer directly to chatbots, assumptions not yet systematically validated.The paper notes that task-irrelevant items, such as “Love surprise parties” for a project-management chatbot, may undermine applicability.
- The study evaluates convergent, discriminant, criterion, and predictive validity to test both the internal structure of self-reports and their relationships with external outcomes.
- Human perceptions and interaction quality are treated as external variables central to evaluating whether chatbot personality design works in practice.
- Self-report scales show moderate convergent and discriminant validity but poor alignment with human-perceived personality and weak correlations with interaction quality.These results indicate limited criterion and predictive validity for chatbot personality evaluation.
- The authors argue that personality evaluation should use real-world task interactions and human perceptions rather than relying solely on chatbot self-reports.
- The released dataset contains human interaction logs and personality perceptions for 500 chatbots with distinct personality designs.
2 Related Work
Prior work commonly applies human personality scales to LLMs, while alternative prompt- and expert-based approaches seek more behaviorally grounded assessments. The paper situates its study within longstanding concerns about divergence between self-reports and external perceptions.
- Human-like LLM responses have encouraged researchers to borrow social-science methods, including personality frameworks, for studying and shaping model behavior.
- Most LLM personality assessments use human-validated standardized scales such as BFI-2 and HEXACO, calculating traits from model responses.
- Alternative approaches modify scale items into open-ended prompts or have experts rate generated stories to better reflect real-world application scenarios.
- Psychological evidence reports an average self-report–informant-report correlation of 0.36 for Big Five traits, indicating both overlap and divergence.
- For chatbots, task-specific design may widen the gap between self-reported traits and human-perceived personality, making user experience an important evaluation target.
3 Methodology
The methodology combines controlled personality designs, task-based human interactions, chatbot self-reports, human personality ratings, and psychometric validity analyses. This design tests whether scale-based chatbot traits emerge consistently and correspond to human judgments and user experience.
- 3.1 Chatbot Design: Each chatbot uses a personality module for predefined traits and a task module for role-specific instructions and task outlines.
- 3.1.1 Personality Module: The personality module applies the Big Five model, sampling five adjective descriptors to prompt either a high or low level of one trait while leaving other traits unconstrained.
- 3.1.2 Task Module: Five task categories operationalize personality in applied settings: job interviews, public service, personalized social support, customized travel planning, and guided learning.
- 3.1 Chatbot Design: The study creates 500 distinct chatbot designs by combining five task instructions with high- or low-level personality configurations, using GPT-4o at temperature zero.
- 3.2.1 Self-report Personality: Chatbot self-reports use BFI-2-XS, BFI-2, and IPIP-NEO-120 responses aggregated into Big Five domain scores on a one-to-five scale.
- 3.2.2 Human-perceived Personality: Participants interacted with one assigned chatbot in multi-turn task conversations and then rated its perceived personality using BFI-2-XS scores aggregated by personality configuration.
- 3.3 Study Procedure: The study procedure included a designated task, interaction, personality evaluation, and separate UEQ-based interaction-quality assessment in a 10–15 minute session.
- 3.4 Analysis Method: Validity analysis uses descriptive statistics, significance tests, MTMM Spearman correlations, and four dimensions: convergent, discriminant, criterion, and predictive validity.Criterion validity compares self-reports with human perceptions, while predictive validity examines relationships with UEQ interaction quality.
4 Results
Personality settings produced consistent differences in chatbot self-reports and, generally, human perceptions, but self-report scores aligned poorly with human-perceived personality and interaction quality. Task context further shaped these relationships, with perceived personality predicting user experience more reliably than chatbot self-reports.
- 4.1 Personality Settings Work for Both Human and Chatbot: Personality manipulations worked in both measures, but condition differences were more pronounced in self-reports than in human perceptions.Self-report scores were consistently higher under high-setting conditions, while human-perceived scores followed the same trend except for Conscientiousness in social support; self-report scales also showed convergent correlations of at least 0.85.
- 4.2 Chatbot’s Self-report Scores Correlate with Human Perceptions Poorly: Self-report scores showed poor alignment with human perceptions: correlations were below 0.4 for most traits, except Agreeableness at 0.58, and self-reports differed significantly from human ratings.Human-perceived scores were more clustered, while self-reports varied more and showed higher cross-trait correlations, indicating weaker discriminant validity.
- 4.2 Chatbot’s Self-report Scores Correlate with Human Perceptions Poorly: Across tasks, personality–perception correlations fluctuated substantially: Agreeableness was consistently strongest, while Extraversion was near zero and Conscientiousness became negative in social support.These variations suggest that personality expression depends on task context and interaction; excessive detail or rule focus may reduce engagement in social support.
- 4.3 Chatbot’s Self-report Scores Unreliably Predict User Experience: Overall, the results expose a disconnect between static questionnaire responses and personality as observed during interactive exchanges.The findings caution against using self-report measures alone for evaluating personality design and support incorporating human perception and task-based interaction signals.
- 4.3 Chatbot’s Self-report Scores Unreliably Predict User Experience: Human-perceived personality predicted interaction quality more strongly than chatbot self-reports, whose associations were weak, inconsistent, or negligible across tasks.Perceived Agreeableness and Conscientiousness correlated 0.71 and 0.76 with travel-planning user experience, whereas the strongest self-report correlation was 0.19 and Openness reached ρ = -0.01 in guided learning.
5 Towards Interactive and Task Grounded Personality Evaluation
Self-report scores show limited predictive and criterion validity, motivating personality evaluation grounded in task context, interaction dynamics, and human perceptions.
- Self-report personality scores show limited predictive and criterion validity, indicating a disconnect from user experience across scenarios.Current scales assume traits are expressed consistently, but the findings challenge that assumption.
- Fine-tuning GPT-4o on multi-turn conversations and human-perceived scores produced stronger average correlations with human perceptions than self-report methods.This suggests contextualized conversation can support more behaviorally grounded assessment.
- The paper advocates replacing static questionnaire-based evaluation with task-driven assessments that reflect how chatbots operate in realistic scenarios.
- Personality evaluations should be based on specific tasks or scenarios because chatbot traits and user attributions vary with context.Evaluation should account for in-context expectations and context-driven behavior.
- Interactive evaluation should track response patterns, adaptability, and changing human perceptions because personality is conveyed through behavior during exchanges.These signals connect personality assessment to user experience and satisfaction.
6 Conclusion
The conclusion finds that chatbot self-report scores diverge from human task-based perceptions and do not align with interaction quality, supporting interactive evaluation methods.
- Self-report personality assessments do not accurately capture chatbot personality as perceived during real-world interactions or its relationship to interaction quality.
Ethics Statement
The human study received exempt Institutional Review Board status and used consented, compensated participants with privacy safeguards.
- The study received Exempt status from the Institutional Review Board, obtained informed consent, paid participants $12 per hour, and manually checked data for sensitive information.
Reproducibility Statement
The study documents its reproducibility materials, chatbot prompt construction, self-report format, human-study interface, and participant procedures.
- Reproducibility Statement: The study makes prompts, self-report collection code, and human-study data publicly accessible; reported costs were $37.41 for self-report experiments and $1091.57 for the human study.
- A Chatbot Prompt Format: Each chatbot prompt combines a personality description, task-specific instructions, and repeated traits to encourage consistent behavior.Its components specify the role, personality level and domain, adjective profile, task objectives, and additional information.
- B Self-report Personality Prompt Format: The self-report prompt asks the chatbot to respond according to the personality description and rate each test item on a five-point agreement or accuracy scale.BFI-2 uses agreement labels, while IPIP-NEO-120 uses accuracy labels.
- C Interface for Human Study: The human-study interface places the chatbot conversation beside a personality survey completed after the participant’s interaction.Participants select responses to statements based on their experience with the chatbot.
- D Human Study Details: Participants were recruited for a 10–15 minute desktop task involving chatbot interaction followed by a 15-question personality questionnaire.The study information stated that participation was voluntary, anonymous, and confidential.
E Dataset Statistics
The study used 500 chatbot configurations and human evaluations across five tasks to examine how configured personality settings appear in human perception.
- E Dataset Statistics: 500 valid conversation transcripts covered 10 chatbots × 2 personality levels × 5 Big Five domains × 5 tasks, with interactions averaging 8 to 9 turns.
- F Participant Statistics: 331 participants reported using conversational AIs at least weekly, while the median participant age was 25–34 and education level was a Bachelor’s degree.
- G Variance in Human-perceived Personality: F-values indicated that high-versus-low personality settings produced greater between-group than within-group variance in human-perceived personality, especially for Agreeableness and Conscientiousness.
H Statistical Significance and Correction
Across conditions, self-report differences were clearer than some human-perceived differences, while direct comparisons revealed systematic gaps between chatbot self-reports and human evaluations.
- H Statistical Significance and Correction: Self-report differences between high and low conditions were clear and statistically significant at p < .001 across tasks, whereas some human-perceived trait differences were weaker or nonsignificant.
- H Statistical Significance and Correction: Most paired comparisons between self-report and human-perceived scores were highly significant across the three psychometric scales, indicating systematic score differences.
- H Statistical Significance and Correction: The main statistical conclusions remained unchanged after false-discovery-rate correction, with key effects retained under FDR q < 0.05.
- H Statistical Significance and Correction: Across all five personality domains, human-perceived scores were lower than self-reports under high settings and higher under low settings, with statistically significant differences.
I Correlation Analysis
Self-report scores correlated weakly with human-perceived personality, while fine-tuning on multi-turn conversations produced stronger alignment with human evaluations; the analysis remained trait-level.
- I Correlation Analysis: Self-report and human-perceived personality scores showed weak overall correlations, including near-zero or negative correlations for several Conscientiousness and Extraversion comparisons.
- J Fine-tuning Details: The fine-tuning procedure used an 80:20 split of conversation scripts and human personality evaluations, training GPT-4o to rate BFI-XS statements from transcripts.
- J Fine-tuning Details: Machine-inferred scores correlated more strongly with human-perceived scores than self-report scores across all evaluated scales on the test set.
- K Future Direction: The study examined relationships between self-report and human-perceived scores at the trait level using correlations, variance analyses, and correlations with external variables.
- K Future Direction: Future work could examine response-pattern differences and forced-choice scales, and extend human-centered evaluation to multimodal models such as MLLMs or VLMs.
L Limitations
The study’s conclusions are constrained by its measurement choices, prompt-based personality control, task selection, model, language, and participant population.
- L Limitations: Human-perceived personality was measured with one questionnaire, so broader personality tests or LLM-specific assessments may be needed for more accurate measurement.
- L Limitations: Because the study used only prompt-based personality control, its findings may not generalize to other personality-setting approaches.
- L Limitations: Task-focused settings may have constrained perceived personality, with Neuroticism remaining relatively low across high and low conditions despite explicit prompting.
- L Limitations: The evaluation covered five common chatbot tasks and may not represent the full spectrum of user interactions.
- L Limitations: Results came from GPT-4o, English psychometric tests, and English-speaking participants, leaving generalization across models and cultures unresolved.