Source-linked AI summary

Emotional Intelligence of Large Language Models

Xuena Wang, Xueting Li, Zi Yin, Yue Wu, Liu Jia

arXiv:2307.09042v2cs.AI

TL;DR

LLM alignment with human emotions and values has not been systematically evaluated despite its relevance to communication and real-world use. The paper develops a standardized emotion-understanding test for humans and LLMs and finds that most tested models achieve above-average EQ scores, though their performance patterns can differ from humans.

  • Problem

    LLMs’ alignment with human emotions and values has received limited systematic evaluation despite its relevance to broad real-world use.

  • Method

    The study develops the standardized SECEU test, comprising realistic scenarios across school, family, and social contexts, to measure emotion understanding in humans and LLMs.

  • Results

    Most tested LLMs achieved above-average EQ scores, although significant performance differences were observed across models.

  • Takeaways & Limitations

    The SECEU provides a psychometric approach for evaluating LLM emotional intelligence and human-like characteristics.

  • Takeaways & Limitations

    Existing theory-of-mind task designs may not meet reliability and validity standards for general assessment and may instead discriminate autism spectrum disorder.

Abstract

from arXiv · show

Large Language Models (LLMs) have demonstrated remarkable abilities across numerous disciplines, primarily assessed through tasks in language generation, knowledge utilization, and complex reasoning. However, their alignment with human emotions and values, which is critical for real-world applications, has not been systematically evaluated. Here, we assessed LLMs' Emotional Intelligence (EI), encompassing emotion recognition, interpretation, and understanding, which is necessary for effective communication and social interactions. Specifically, we first developed a novel psychometric assessment focusing on Emotion Understanding (EU), a core component of EI, suitable for both humans and LLMs. This test requires evaluating complex emotions (e.g., surprised, joyful, puzzled, proud) in realistic scenarios (e.g., despite feeling underperformed, John surprisingly achieved a top score). With a reference frame constructed from over 500 adults, we tested a variety of mainstream LLMs. Most achieved above-average EQ scores, with GPT-4 exceeding 89% of human participants with an EQ of 117. Interestingly, a multivariate pattern analysis revealed that some LLMs apparently did not reply on the human-like mechanism to achieve human-level performance, as their representational patterns were qualitatively distinct from humans. In addition, we discussed the impact of factors such as model size, training method, and architecture on LLMs' EQ. In summary, our study presents one of the first psychometric evaluations of the human-like characteristics of LLMs, which may shed light on the future development of LLMs aiming for both high intellectual and emotional intelligence. Project website: https://emotional-intelligence.github.io/

Introduction

LLMs show strong intellectual abilities, but their emotional intelligence has been less systematically assessed, partly because Theory of Mind tasks are heterogeneous and insufficiently discriminative for standardized EI measurement. This study addresses that gap by introducing the SECEU, a standardized emotion-understanding test for humans and LLMs, normed on more than 500 young adults and applied across mainstream models.

  • Motivation: LLMs’ empathy has been studied less systematically than their performance in logic-based and other intellectual tasks.Previous Theory of Mind results improved from 70% for text-davinci-002 to 100% for GPT-4 with in-context learning, but empathy investigations remained comparatively scarce.
  • Limitations of prior assessment: Theory of Mind is unsuitable as a standardized EI test because its heterogeneous content may not meet psychometric reliability and validity standards.The task spans abilities from false-belief understanding to pragmatic reasoning.
  • Limitations of prior assessment: Theory of Mind is also generally too simple for typical humans, making it more useful for diagnosing EI-related disorders than discriminating performance in the general population.Standardized EI tests such as MSCEIT therefore do not include the Theory of Mind task.
  • Study rationale: Emotion understanding—recognizing, interpreting, and understanding emotions in social contexts—is a fundamental EI component supporting communication, empathy, and social interaction.Because LLMs lack internal emotional states or experiences, EU assessment focuses on their understanding and interpretation of social context.
  • Contribution: The study developed the Situational Evaluation of Complex Emotional Understanding (SECEU), a standardized EU test suitable for both humans and LLMs.The test was designed to evaluate LLM emotion understanding without assuming internal emotional experiences.
  • Study design: More than 500 young adults established the SECEU norm, which was used to standardize scores from diverse mainstream LLMs for direct human comparison.The study also compared multivariate response patterns between LLMs and human participants to assess representation similarities.

Results

The SECEU showed strong discrimination, reliability, and validity in human samples, enabling standardized comparison of mainstream LLMs with humans. GPT-4 achieved an EQ of 117, exceeding 89% of human participants, while some models could not complete the test even with prompts.

  • Human test validation: The SECEU discriminated individual EU ability, with scores ranging from 1.40 to 6.29 (M = 2.79, SD = 0.822).Smaller Euclidean distances from consensus scores indicated better EU ability.
  • Human test validation: Cronbach’s α = 0.94 demonstrated high internal consistency across the 40-item SECEU.The test showed no evidence of ceiling or floor effects, with mean item distances from 2.19 to 3.32.
  • Human test validation: All three high-EI experts exceeded at least 73% of the population, while their average score exceeded 99%.These results supported the validity of consensus scoring for standardizing the SECEU.
  • LLM performance: LLaMA, Fastchat, and RWKV-v4 were unable to complete the SECEU even with prompt assistance.Other models completed the test under prompting, and analyses retained each model’s best performance under ideal conditions.
  • LLM performance: GPT-4’s EQ score is 117, exceeding 89% of human participants.LLM scores were standardized against human norms using Euclidean distances from human standard scores.

LLaMA

LLaMA-based models generally scored below the OpenAI GPT series, with performance ranging from above-average EQ to failure to complete the test. Their representational similarity also varied substantially, from human-like patterns in Koala to qualitatively different patterns in Alpaca and Vicuna.

  • EQ performance: Alpaca and Vicuna achieved the highest LLaMA-based EQ scores, at 104 and 105, respectively.
  • EQ performance: 83 was Koala’s EQ score, surpassing 13% of humans, while base LLaMA was unable to complete the test.
  • Representational similarity: Koala showed the highest human representational similarity among evaluated models, with r = 0.43, p < 0.01, exceeding 93% of human participants.
  • Representational similarity: Alpaca and Vicuna differed qualitatively from human representational patterns, with r = 0.03 and r = -0.02, respectively.Despite their above-average EQ scores, these models likely employed mechanisms qualitatively different from human processes.

Discussion

The study introduces a psychometric test of LLM emotional understanding and finds that most tested models achieve above-average EQ, despite substantial differences in performance and human-like representations. Model size, training, and architecture appear related to EQ, but limited samples and descriptive analyses constrain conclusions and motivate broader, causal, multimodal evaluations.

  • Contribution: The SECEU provided a valid and reliable psychometric test of emotional understanding for evaluating LLM emotional intelligence.The study addresses the scarcity of tests examining LLM alignment with human values and needs.
  • Findings: Most tested LLMs achieved above-average EQ scores, but individual differences were substantial and some reached human-level performance through non-human-like representations.Some models’ representational patterns diverged significantly from human patterns, suggesting qualitatively different underlying mechanisms.
  • Influencing factors: Architecture was associated with performance: Transformer models generally performed well, whereas RNN-based RWKV-v4 failed to complete the test and decoder-only models outperformed encoder-decoder models.Oasst and Alpaca also yielded similar scores despite architectural differences, demonstrating the potential importance of fine-tuning.
  • Limitations and future work: The findings are biased and inconclusive because only a limited number of LLMs were tested, and the examined factors do not establish causal relationships.Future work should broaden EI assessments beyond EU, include multimodal emotional cues and cognitive factors, and manipulate model size or training to test causality.

Methods

The study assessed human participants and mainstream LLMs with the 40-item SECEU test, quantifying human alignment through standardized scores and comparing response patterns with human representations.

  • Representation analysis: Item-wise Pearson correlations compared LLM discriminability patterns with a human template derived from averaged human multi-item patterns.Human-to-human similarity formed the norm, while LLM-to-human similarity assessed representational correspondence.

2.0. Intelligence, 33(3), 285–305. https://doi.org/10.1016/j.intell.2004.11.003

The study materials, including the SECEU test, standardized scores, norms, and prompts, are publicly available, while human raw data require a reasonable request. The listed authors divided responsibilities across test development, translation, participant testing, analysis, LLM evaluation, prompting, and manuscript preparation.

  • Materials availability: The SECEU test, standardized scores, norm, prompts, and human-participant test code are available at the project website.The English and Chinese versions of the SECEU test are included.
  • Materials availability: Human-participant raw data are available from the corresponding author upon reasonable request.
  • Author contributions: Authors developed, translated, and implemented the SECEU test, constructed the EQ norm, and analyzed human-participant data.X.L. developed the test; X.L. and X.W. translated it; Z.Y. built the online test; and X.W. and J.L. contributed to testing, norm construction, and analysis.
  • Author contributions: The team evaluated LLMs, wrote prompts, and prepared the manuscript through collaborative writing, suggestions, and revisions.Y.W. performed the LLM testing, Z.Y. wrote the prompts, and X.W. and J.L. wrote the manuscript.

Declaration of conflicting interests

The author declared no potential conflicts of interest regarding the research, authorship, or publication of the article.

  • The author declared no potential conflicts of interest concerning the research, authorship, and/or publication.

Supplementary

Prompting substantially affected LLM performance on emotional-intelligence tests, with most models requiring Two-shot Chain of Thought prompts while GPT-3.5-turbo achieved an EQ of 94 with Zero-shot prompts. Step-by-Step Thinking increased GPT-3.5-turbo’s human-model correlation, whereas combined prompting could increase pattern similarity without improving EQ scores.

  • Prompting effects: Most LLMs were unable to complete the task without Two-shot Chain of Thought prompts.The passages attribute this requirement to limitations in long-term memory and context understanding.
  • Prompting effects: Step-by-Step Thinking prompts did not improve DaVinci, Curie, Babbage, or text-davinci-002 performance.The passages suggest that limited instruction-following ability may explain the lack of improvement, although the explanation for text-davinci-002 is speculative.
  • Prompting effects: Step-by-Step Thinking prompts increased the correlation between humans and GPT-3.5-turbo.This indicated progress in the model’s ability to mimic human emotional understanding and thought processes.
  • Prompting effects: Combined Two-shot Chain of Thought Reasoning and Step-by-Step Thinking prompts increased pattern similarity but did not raise EQ scores for GPT-3.5-turbo, text-davinci-001, or text-davinci-003.The varied responses across models underscore the relevance of architecture, training data, fine-tuning, and optimization objectives.

LLaMA

The LLaMA section reports failed evaluations for Fastchat, RWKV-v4, and several zero-shot prompting conditions, alongside prompted results summarized by SECEU, EQ, pattern similarity, and human-performance percentages. The table footnote defines failures, percentile percentages, and Pearson-correlation-based pattern similarity.

  • LLaMA: Fastchat and RWKV-v4 both failed the evaluation.The source marks each model as “FAILED.”
  • LLaMA: Prompt-assisted results are reported using SECEU score, EQ score, pattern similarity, and the percentage of humans scoring below the LLM.Table S2 presents these measures, while the footnote defines the percentage as the proportion of humans whose performance was below the LLM’s.
  • LLaMA: Zero-shot and zero-shot + step-by-step conditions are each marked FAILED.Four listed zero-shot conditions receive failure labels.
  • LLaMA: Pattern similarity is indexed by the Pearson correlation coefficient r, with significance markers * for p < 0.05 and ** for p < 0.01.These definitions apply to the prompted results table.
Loading 2307.09042v2…