Source-linked AI summary

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models

Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao

arXiv:2501.16566v2cs.HC

TL;DR

MLLM-based emotion understanding needs richer descriptive datasets and multimodal-centric modeling because closed-set labels and existing annotation strategies are limited. The paper introduces MER-Caption, AffectGPT, and MER-UniBench, and reports robust cross-task performance while acknowledging possible annotation inaccuracies from automatic labeling.

  • Problem

    Emotion understanding is limited by closed-set emotion categories, scarce large-scale descriptive datasets, and insufficiently multimodal evaluation and modeling frameworks.

  • Method

    The paper constructs MER-Caption with model-led human-assisted annotation, develops AffectGPT with pre-fusion multimodal integration, and introduces MER-UniBench with tailored metrics.

  • Results

    AffectGPT significantly outperforms existing MLLMs across MER-UniBench evaluations, while MER-Caption complements existing datasets with large-scale descriptive annotations.

  • Takeaways & Limitations

    The dataset, model, and benchmark provide a foundation for evaluating and advancing MLLM-based emotion understanding.

  • Takeaways & Limitations

    MER-Caption+ may contain inaccurate descriptions because its automatic annotation strategy lacks manual checks, motivating further post-filtering.

Abstract

from arXiv · show

The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level, from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT's robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.

1. Introduction

Traditional closed-set emotion classification cannot capture diverse, nuanced, and coexisting affective states, while MLLMs enable descriptive emotion understanding. This paper addresses dataset, multimodal-fusion, and evaluation gaps through MER-Caption, AffectGPT, and MER-UniBench.

  • Motivation: Closed-set taxonomies fail to capture the diverse and nuanced emotional expressions found in real-world scenarios.Cultural idioms, context-dependent metaphors, and personalized behaviors contribute to this diversity.
  • Motivation: MLLMs can describe complex, coexisting emotional states in natural language and generate emotion categories beyond basic emotions.Their broad vocabulary supports more descriptive emotion modeling than traditional discriminative approaches.
  • Research gaps: Large-scale datasets with intensive descriptive annotations remain scarce because manual annotation is costly and model-only annotation can lack label quality.Existing human-model collaboration strategies also face scalability challenges.
  • Contributions: MER-Caption adopts a model-led, human-assisted strategy, while MER-UniBench provides tailored metrics and comprehensive evaluation across MLLM-based emotion-understanding tasks.The benchmark is designed for free-form natural-language outputs and typical MER tasks.
  • Contributions: AffectGPT uses additional pre-fusion operations to enhance multimodal integration for emotion understanding.The model is designed to address limitations of leaving multimodal fusion entirely to the language model.
  • Results: Over 9% performance improvement over existing MLLMs is reported for AffectGPT in extensive experiments.The paper presents this result as evidence of AffectGPT’s effectiveness.

2. MER-Caption: Dataset Construction

MER-Caption combines model-generated descriptions with human priors and filtering to balance dataset scale with annotation quality. Its construction includes guided description generation, automatic annotation, and multilevel filtering of mismatched or low-quality samples.

  • Motivation: Descriptive datasets offer more diverse emotion labels, but existing construction strategies trade off scalability, completeness, and annotation quality.Manual annotation is costly and can produce incomplete descriptions, whereas model-only annotation may lack human proofreading.
  • Annotation strategy: MER-Caption uses a model-led, human-assisted strategy in which human priors guide description generation and sample filtering.The process targets a balance between label quality and dataset size.
  • Dataset: MER-Caption contains 115K coarse-labeled samples and 31K fine-labeled samples created from previously unlabeled data.The dataset is constructed through automatic annotation after the guided selection and generation process.
  • Description generation: The pipeline selects base models using human-prior-guided preliminary experiments and combines multimodal generation with automatic annotation.The supplied passages describe model selection through fine-grained preliminary labeling and subsequent automated dataset creation.
  • Sample filtering: A two-level filtering process removes mismatched audio-video samples and other problematic descriptions before dataset construction is finalized.The process addresses unverified generated descriptions and data that do not match the target emotion-analysis setting.
  • Sample filtering: High-level filtering uses multiple multimodal classifiers, majority voting, and disagreement checks between extracted and model-based labels.Descriptions are removed when their extracted labels disagree with labels obtained through model-based crowdsourcing.

3. AffectGPT: Model Design

AffectGPT extends mainstream multimodal language-model architectures with pre-fusion operations that integrate audio and video before the LLM, targeting emotion understanding. The model offers Q-Former-based and attention-based fusion, balancing temporal preservation, compression, and computational efficiency.

  • Mainstream Architecture: Existing AV-LLMs extract modality-specific features and generally leave cross-modal interaction to the LLM, which is insufficient for multimodal emotion recognition.The mainstream pipeline encodes modalities, projects them into LLM-compatible tokens, and autoregressively generates responses.
  • Pre-fusion Operation: AffectGPT moves cross-modal interaction outside the LLM through pre-fusion operations applied by default to audio and video latent features.Experiments using projected features instead decreased performance.
  • Q-Former: Q-Former-based pre-fusion concatenates audio and video features temporally, then uses learnable query tokens and cross-attention to distill multimodal content.The fused representation preserves temporal information and is compressed into K query-token features.
  • Attention: Attention-based pre-fusion averages each modality's features, computes attention weights over the compressed representations, and produces a fused feature.Unlike Q-Former, this module directly compresses temporal information before multimodal fusion.
  • Efficiency: Q-Former and attention mechanisms provide cross-modal interaction with substantially lower computational cost than LLMs.Q-Former distills multimodal content into query tokens, whereas attention dynamically weights modalities from multimodal inputs.

4. MER-UniBench: Evaluation Benchmark

MER-UniBench evaluates MLLM-based emotion understanding across fine-grained emotion recognition, basic emotion recognition, and sentiment analysis using metrics tailored to free-form outputs. Its evaluation groups synonymous emotion labels, accommodates variable-length predictions, and averages results across emotion wheels.

  • Benchmark scope: MER-UniBench covers fine-grained emotion recognition, basic emotion recognition, and sentiment analysis with specialized metrics for MLLM outputs.The benchmark includes OV-MERD+ for fine-grained recognition, four datasets for basic recognition, and four datasets for sentiment analysis.
  • Fine-grained Emotion Recognition: Fine-grained evaluation reduces synonym effects by mapping inflected forms, synonyms, and emotion-wheel labels into unified groups.The three levels map forms such as happier and happiness to happy, synonyms such as joyful to happy, and outer-wheel labels to inner labels.
  • Fine-grained Emotion Recognition: Fine-grained metrics use set-based evaluation for samples with variable numbers of true and predicted emotion labels, then average results across emotion wheels.The framework computes set-level metrics and averages the results from different emotion wheels for ranking.
  • Basic Emotion Recognition: Basic emotion recognition uses hit rate because MLLMs may output multiple labels while datasets provide one majority-voted basic-emotion label.A prediction receives credit when the true basic label is included after mapping free-form outputs to the basic-emotion space.
  • Evaluation limitation: The benchmark does not yet evaluate whether additional predicted labels are incorrect because basic-emotion datasets lack fine-grained reference labels.Such labels may represent valid fine-grained emotions outside the basic categories, leaving their evaluation as future work.
  • Sentiment Analysis: Sentiment analysis maps floating-point annotations to negative or positive polarity according to whether scores are below or above zero.The task uses CMU-MOSI, CMU-MOSEI, CH-SIMS, and CH-SIMS v2, with sentiment labels extracted from MLLM outputs.

5. Results and Discussion

AffectGPT is evaluated on MER-UniBench against existing MLLMs, with analyses of dataset quality, multimodal inputs, encoder and LLM choices, LoRA, and pre-fusion. Results generally attribute performance gains to MER-Caption and the framework rather than particular backbone choices.

  • Effectiveness of MER-Caption: MER-Caption achieves excellent performance when training-data choice is varied under otherwise consistent settings, complementing existing datasets.Existing datasets are limited by insufficient emotion focus, scale, or annotation quality.
  • Ablation Study on MER-Caption: Two-level filtering improves performance over no filtering or low-level filtering alone, indicating that dataset quality matters alongside quantity.Fewer training samples do not necessarily produce worse performance.
  • Ablation Study on Model: Pre-fusion operations generally improve performance, supporting separate treatment of cross-modal interactions outside the LLM.AffectGPT implements Q-Former-based and attention-based pre-fusion operations.
  • Analysis of Input Impact: Multimodal inputs outperform unimodal inputs, while face inputs slightly outperform frame inputs and combining them adds no further improvement.The default inputs are audio, face, and text.
  • Backbone Choices: LLM choice has limited impact, and encoder choice has minimal impact, indicating that AffectGPT’s performance is driven mainly by its dataset and framework.CLIP ViT marginally outperforms EVA CLIP and DINOv2 for video; ImageBind is slightly inferior for audio.
  • Role of LoRA in LLMs: LoRA fine-tuning improves performance over no LoRA, but increasing its rank adds computational cost without significant performance gains.

6. Conclusion

The paper advances MLLM-based emotion understanding through a descriptive dataset, a pre-fusion model, and a benchmark tailored to free-form outputs. Experiments validate these components.

  • The paper introduces MER-Caption, AffectGPT, and a benchmark with metrics tailored to free-form natural-language emotion outputs.

Impact Statements

The paper situates multimodal emotion recognition as relevant to human-computer interaction and reviews how emotion datasets and models have evolved from categorical to descriptive approaches. It also states the work’s social-impact and ethics positions.

  • Social Impact: Emotion recognition can support human-computer interaction and applications including education and psychological counseling.
  • Ethics Statement: The paper states that it uses previously unlabeled MER2024 data with permission and releases the dataset under CC BY-NC 4.0.
  • Related Work: Categorical emotion datasets may not fully capture diverse or coexisting emotions, motivating descriptive natural-language datasets.
  • Related Work: Descriptive emotion datasets require generative models, including LLM- and MLLM-based frameworks.

B. Implementation Details

The implementation uses frozen unimodal encoders and LLM weights with trainable projection, pre-fusion, and LoRA components. Evaluation includes MLLM outputs and prompts for emotion-label extraction and sentiment classification.

  • Model configuration: CLIP ViT-L and HUBERT-L serve as the visual and acoustic encoders, respectively.
  • Model configuration: The training setup fine-tunes only the LLM's LoRA module, projector, and pre-fusion branch while freezing the LLM and unimodal encoders.This reduces GPU memory usage and speeds up training.
  • Model configuration: Qwen-2.5 is used as the LLM, with LoRA rank set to 16 by default.
  • Evaluation procedure: The evaluation extracts open-ended emotion labels from MLLM outputs before applying a second prompt to classify sentiment as positive, negative, or neutral.Traditional accuracy and F1 are unsuitable for unrestricted emotion outputs.

F. Choice of Description Generation Strategy

The description-generation strategy combines audio, visual, and textual cues using GPT-3.5 after selecting SALMONN and Chat-UniVi as unimodal cue generators. The authors report that multimodal integration can improve performance, while acknowledging that combined results were not used for model selection.

  • Generation strategy: GPT-3.5 integrates audio and video cues from candidate models with text content to generate emotional descriptions.The integration prompt asks the model to analyze and explain how textual, acoustic, and visual clues indicate emotion.
  • Preliminary experiments: The combined models generally outperform either individual model in preliminary experiments measured by Fs.Fs is chosen as the primary metric because it considers both accuracy and completeness.
  • Model selection caveat: Combined results are not used for model selection; individual-model performance determines the choices of Chat-UniVi and SALMONN.The combination experiments are intended to demonstrate the benefits of integrating multimodal cues.
  • Multimodal conflict: GPT-3.5 is reported to provide reasonable responses when audio, video, and text convey conflicting emotions.

H. Dataset Comparison

MER-Caption provides detailed descriptions and multiple emotion labels per sample, while most benchmark datasets use more constrained emotion or sentiment annotations. MER-UniBench covers three MER tasks and focuses on single-person videos.

  • MER-Caption comparison: MER-Caption samples contain detailed descriptions and rich emotion labels, with Figure 7 comparing description lengths and labels per sample.
  • Video duration: Most MER-Caption videos last between 2 and 5 seconds.
  • Benchmark scope: Most MER-UniBench datasets focus on single-person videos, and MER-UniBench evaluates fine-grained emotion recognition, basic emotion recognition, and sentiment analysis.
  • Dataset annotations: OV-MERD+ supports a variable number of fine-grained emotions per sample without restricting labels to predefined taxonomies.It extends the original OV-MERD dataset, which contained 332 samples.
  • Dataset annotations: IEMOCAP, MELD, and MER2023/MER2024 use predefined categorical emotion labels, whereas CMU-MOSI, CMU-MOSEI, CH-SIMS, and CH-SIMS v2 annotate sentiment intensity.The sentiment ranges are -3 to +3 for CMU-MOSI/CMU-MOSEI and -1 to 1 for CH-SIMS variants.
  • Emotion representation: The paper uses five emotion wheels derived from previous research because no universal definition of the emotion wheel exists.

L. Main Results

Table 12 reports complete results across datasets and multiple metrics, with primary metrics highlighted and averaged in the final column. The authors state that these results verify AffectGPT's effectiveness in multimodal emotion understanding.

  • AffectGPT's complete results across datasets verify its effectiveness in multimodal emotion understanding.Table 12 highlights primary metrics and reports their average in the final column.

M. Ablation Study on MER-Caption

The ablation study evaluates MER-Caption through controlled dataset comparisons and examines how sampled-frame counts affect model performance across two input configurations.

  • Dataset comparison: Table 13 compares datasets while keeping the model architecture and experimental setup fixed, changing only the training dataset.The comparison is intended to assess MER-Caption’s effectiveness for emotion understanding.
  • Dataset comparison: Existing datasets are described as providing insufficient attention to emotion tasks or lacking high-quality emotion descriptions, whereas MER-Caption addresses these issues.
  • Evaluation setup: Table 12 identifies audio, video, and text inputs and reports the mean of each dataset’s primary metric across datasets.
  • Sampling-frame ablation: The frame-sampling experiment compares face-only with face-text inputs across sampling counts from 2 to 64 frames.The paper’s default is 8 frames per video.
  • Sampling-frame ablation: Using too few frames produces a noticeable performance decline in the reported sampling experiment.The passage specifically identifies fewer than 2 frames as an example of too few frames.
Loading 2501.16566v2…