Source-linked AI summary

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

Yi Zheng, Yifan Xu, Yan Zhou, Hejia Chen, Chunyu Qiang, Xiaoqiang Liu, Xiaohan Li, Shenze Huang, Yue Zhang, Guoying Zhao, Pengfei Wan

arXiv:2608.20905v1cs.CV

TL;DR

Existing multimodal dialogue datasets often lack sufficient emotion annotation, diversity, scale, and natural face-to-face capture. EmotionDialogCN addresses these gaps through a large-scale audiovisual-emotional dataset and low-interference collection framework. Its emotion distribution deviation is 0.64 versus MultiDialog’s 5.65, while evaluations report stronger facial quality, body stability, and interaction naturalness.

  • Problem

    Existing multimodal dialogue datasets often have inadequate emotion annotations, limited emotional diversity, small scale, and webcam-related visual distortions.

  • Method

    EmotionDialogCN combines scenario-driven improvised dialogues with a low-interference audiovisual recording and performance-authenticity framework.

  • Results

    Emotion distribution deviation is 0.64 for EmotionDialogCN versus 5.65 for MultiDialog, and comprehensive evaluations report stronger facial quality, body stability, and interaction naturalness.

  • Takeaways & Limitations

    EmotionDialogCN provides a large-scale foundation for emotionally expressive and socially grounded auditory-visual dialogue systems.

  • Takeaways & Limitations

    The dataset is limited to Mandarin-speaking regions because recordings from non-native English speakers were excluded, restricting cross-cultural generalizability.

Abstract

from arXiv · show

Face-to-face audiovisual interaction is central to human communication, conveying rich emotional and social cues. However, existing multimodal dialogue datasets remain limited by inadequate emotion annotations, poor emotional diversity, and small scale. We introduce EmotionDialogCN, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication. It contains 21,880 dialogue sessions performed by 119 professional actors across 20 everyday scenarios, covering 18 emotion categories with over 400 hours of recordings, the largest and most comprehensive dataset of its kind. A novel data collection framework minimizes equipment interference, enabling natural and nuanced emotional expressions. EmotionDialogCN achieves an emotion distribution deviation of 0.64 from real human emotion statistics (versus 5.65 for prior datasets) and consistent subject framing (52-59% frame occupancy). Together, these properties translate into stable unimodal and multimodal performance across acoustic, lexical, and visual modalities, with fusion results further underscoring strong multimodal alignment and cross-modal complementarity.

1 Introduction

EmotionDialogCN addresses limitations in existing multimodal dialogue datasets by collecting large-scale, emotionally diverse, natural face-to-face audiovisual interactions. Its recording and performance framework is designed to preserve authentic visual, acoustic, and emotional expression.

  • Motivation: Existing datasets commonly suffer from inaccurate emotion annotations, limited emotion categories, small scale, and webcam-induced visual distortion.These shortcomings motivate a larger, better-annotated, more emotionally diverse resource for face-to-face interaction.
  • Scientific contribution: The dataset provides a large-scale foundation for multimodal emotional dialogue generation in realistic face-to-face interaction scenarios.The paper characterizes it as a practical framework and resource for emotionally expressive auditory-visual dialogue systems.
  • Collection framework: A low-interference recording environment uses professional cameras and human-eye-approximating focal lengths to reduce visual distortion and preserve natural interaction.The setup is intended to minimize equipment effects on actors’ performance and capture faithful acoustic and visual signals.
  • Collection framework: Improvised dialogues, real-time equipment supervision, and audio monitoring support spontaneous, expressive, and contextually appropriate emotional performances.Actors receive scripts containing contextual information, prompts, and intended emotional tone.
  • Dataset contribution: EmotionDialogCN contains 21,880 dialogue sessions by 119 professional actors across 20 scenarios and 18 emotion categories, totaling 400 hours.The dataset targets spontaneous, emotionally rich everyday interactions.

2 Related Work

Prior multimodal dialogue datasets provide valuable audiovisual and emotional resources but remain constrained by participant, linguistic, scene, annotation, or recording limitations. EmotionDialogCN is positioned as stronger in expression diversity, actor scale, data scale, and recording quality.

  • Single-speaker and controlled datasets: Early controlled audiovisual datasets supported talking-face models but were limited in linguistic diversity and participant variety.These datasets generally emphasized high-quality capture in laboratory environments.
  • In-the-wild interaction datasets: Movie and television interaction datasets provide diverse multi-subject scenes, but frontal facial capture of both participants is not guaranteed.Complex scenes make consistent face-to-face visual coverage difficult.
  • Dyadic datasets: EmotionTalk, NoXi, MultiDialog, and MARS focus on dyadic interaction in controlled environments, addressing some data-quality needs.These datasets capture communication between two participants but differ in scale, annotation, or recording design.
  • Comparison: EmotionDialogCN excels in expression diversity, actor scale, data scale, and recording quality relative to the compared datasets.The comparison distinguishes single and dyadic datasets using S and D.
  • Motivation: Prior research links mediated interaction with altered facial-expression dynamics, disrupted eye contact, and weaker emotional connection compared with in-person communication.These findings reinforce the relevance of carefully designed face-to-face audiovisual collection.

3 Method

The method combines scenario-driven prompt generation, improvised actor performance, and a low-interference audiovisual recording environment. Quality control and broad performer coverage are used to produce natural, emotionally diverse dialogue data.

  • Prompted dialogue generation: Dialogue scenarios span 20 everyday themes and diverse emotional states, with actors improvising from contextual information, intended tone, and content prompts.Actors build seamless interactions rather than reciting prompts verbatim.
  • Prompted dialogue generation: The prompt pipeline defines scenarios and emotional tones, generates conversation-oriented prompts with GPT-4, and iteratively adjusts emotional coverage.Human review by designers, supervisors, and actors filters unsuitable wording and supports authenticity and ethical compliance.
  • Dataset scale: Applying the prompt strategies produced 21,880 dialogue sessions with substantial emotional diversity.The resulting corpus is presented as a systematic and scalable resource for multimodal dialogue collection and modeling.
  • Recording environment: Actors perform face-to-face while camera and teleprompter placement is adjusted to preserve natural sightlines, framing, and comfortable eye alignment.Subjective actor evaluations identified chest-height camera placement while seated as an effective balance.
  • Recording environment: Vertical 4K cameras with human-eye-approximating focal lengths improve facial and upper-body completeness while reducing perspective distortion.Dedicated microphones, separate audio tracks, and acoustic treatment support clean audiovisual capture.
  • Quality control: On-site supervision, sound monitoring, manual review, and removal of corrupted or noisy segments support broadcast-grade data quality.Actors’ creative freedom is retained to enhance naturalness and emotional authenticity.

4 Result

EmotionDialogCN combines broad emotional, topical, and visual diversity with stable unimodal and multimodal recognition performance. Its emotion frequencies also more closely match natural human statistics than MultiDialog, while recordings provide consistent framing and perceptually reliable signals.

  • Data Diversity Analysis: EmotionDialogCN covers 18 core emotions across 20 topical categories with balanced emotional distribution and high semantic variability.
  • Data Diversity Analysis: 119 professional actors provide diversity in voices, intonation, speaking styles, hairstyles, and clothing across recordings.
  • Body Completeness and Stability: 52%–59% frame occupancy provides consistent subject framing, while conventional webcams often introduce distortion and omit upper-body information.
  • Unimodal Emotion Recognition: Unimodal experiments show consistently high and stable accuracy across acoustic, textual, and visual modalities.
  • Multimodal Emotion Recognition: All evaluated multimodal fusion models achieve consistently high and stable performance, indicating strong multimodal alignment and emotional diversity.
  • Emotion Distribution Analysis: 0.64 average absolute distribution deviation for EmotionDialogCN is substantially lower than MultiDialog’s 5.65, indicating closer alignment with natural human emotion frequencies.
  • Subjective Consistency Analysis: A two-sample t-test found a significant difference between low- and high-variance annotation groups (t = −47.49, p < 0.001), supporting annotation reliability and internal consistency.

5 Conclusion

The paper introduces EmotionDialogCN and its data-collection framework for emotionally expressive, socially grounded audiovisual dialogue systems. The dataset supports future emotion-aware and multimodal dialogue applications, but its Mandarin-speaking regional scope limits cross-cultural generalizability and it is not advised for direct clinical or neuroscience use.

  • Conclusion: EmotionDialogCN is a large-scale, high-quality multimodal dataset designed for emotionally expressive and socially grounded audiovisual dialogue systems.
  • Conclusion: Comprehensive evaluations report improved facial quality, body stability, and overall interaction naturalness relative to existing datasets.
  • Future Work: Future work includes emotion-aware response generation, emotion recognition, multimodal dialogue generation, empathetic virtual agents, and adaptive learning systems.
  • Limitations: Excluding recordings from non-native English speakers leaves the dataset limited to Mandarin-speaking regions and restricts cross-cultural generalizability.
  • Limitations: The authors advise against direct use of the dataset in clinical or neuroscience applications pending official usage guidance.
Loading 2608.20905v1…