Source-linked AI summary

SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild

Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen, Maximilian Schmitt, Fabien Ringeval, Jing Han, Vedhas Pandit, Antoine Toisoul, Bjorn Schuller, Kam Star, Elnar Hajiyev, Maja Pantic

arXiv:1901.02839v2cs.HCcs.AIcs.CV

TL;DR

Existing emotion datasets often rely on controlled settings and limited demographic or task diversity, motivating richer real-world evidence. The paper introduces SEWA DB, a multilingual in-the-wild audio-visual dataset with broad annotations and baseline evaluations. It provides a benchmark resource for automatic affect and behaviour analysis, subject to academic-use access restrictions.

  • Problem

    Existing datasets often use controlled settings, induced behaviour, and limited demographic diversity, restricting their suitability for in-the-wild affect analysis.

  • Method

    The paper constructs SEWA DB from unconstrained audio-visual recordings and combines multimodal annotations with early-fusion baseline modelling.

  • Results

    SEWA DB contains 1990 clips comprising 1600 minutes of advert reactions and 1057 minutes of video-chat recordings, alongside baseline FAU, valence, arousal, and liking/disliking experiments.

  • Takeaways & Limitations

    SEWA DB provides a benchmark for automatic analysis of audio-visual behaviour in the wild and cross-task affect estimation.

  • Takeaways & Limitations

    Access is restricted to academic use and requires researchers to sign an end-user licence agreement.

Abstract

from arXiv · show

Natural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life. Accurately annotated real-world data are the crux in devising such systems. However, existing databases usually consider controlled settings, low demographic variability, and a single task. In this paper, we introduce the SEWA database of more than 2000 minutes of audio-visual data of 398 people coming from six cultures, 50% female, and uniformly spanning the age range of 18 to 65 years old. Subjects were recorded in two different contexts: while watching adverts and while discussing adverts in a video chat. The database includes rich annotations of the recordings in terms of facial landmarks, facial action units (FAU), various vocalisations, mirroring, and continuously valued valence, arousal, liking, agreement, and prototypic examples of (dis)liking. This database aims to be an extremely valuable resource for researchers in affective computing and automatic human sensing and is expected to push forward the research in human behaviour analysis, including cultural studies. Along with the database, we provide extensive baseline experiments for automatic FAU detection and automatic valence, arousal and (dis)liking intensity estimation.

1 INTRODUCTION

Existing affect datasets often lack realistic recording conditions, spontaneous interaction, demographic diversity, and comprehensive annotations. SEWA DB addresses these gaps with multilingual in-the-wild recordings, broad participant coverage, rich multimodal labels, and baseline experiments.

  • Research gaps: Existing datasets are predominantly collected under controlled conditions, limiting generalisation to in-the-wild behavioural recordings.They commonly use controlled noise, reverberation, illumination, cameras, and limited verbal content.
  • Research gaps: Many datasets use induced rather than spontaneous behaviour, whose dynamics differ from natural facial expressions and are often not modelled.The paper highlights differences in timing, velocity, frequency, and temporal inter-dependencies between gestures.
  • Research gaps: Existing approaches typically analyse one interactant, despite behavioural influence between interlocutors and the importance of mimicry, rapport, and sentiment.The paper identifies simultaneous analysis of both interacting parties as an unmet need in webcam-mediated interaction.
  • Research gaps: Available databases are commonly culture-specific and lack large-scale support for studying cultural effects on emotional expression and communication.Examples include corpora focused on German, UK, French-speaking, or Greek participants.
  • SEWA DB contribution: SEWA DB provides spontaneous audio-visual recordings captured in unconstrained real-world environments using standard webcams and microphones.The database is multilingual and includes participants spanning six cultural groups, five age groups, and both genders.
  • SEWA DB contribution: SEWA DB combines facial, vocal, verbal, affective, and social annotations, including FAUs, landmarks, valence, arousal, liking, agreement, and mimicry.This breadth supports simultaneous study of multiple aspects of affect and relationships among demographics, language, affect, and behaviour.
  • SEWA DB contribution: The paper provides exhaustive baseline experiments for FAU detection and valence, arousal, and liking/disliking estimation.
  • SEWA DB contribution: SEWA DB is publicly available as a resource for affective computing, automatic human sensing, and cross-cultural human behaviour analysis.

2 STATE-OF-THE-ART IN AUDIO-VISUAL EMOTION

Audio-visual emotion research depends on how affect is elicited, represented, and annotated, while existing conversational corpora often remain constrained by setting, language, demographics, or annotation breadth. The paper positions spontaneous interaction and multidimensional continuous annotations as important directions for more naturalistic analysis.

  • Elicitation methods: Emotion datasets use posed, induced, or spontaneous elicitation, with elicitation type significantly shaping affect models.
  • Elicitation methods: Posed data provide control and known target emotions but do not represent fully natural expressions for training in-the-wild systems.
  • Elicitation methods: Induced methods elicit reactions through standardised stimuli or controlled human-computer interactions while recording participants’ responses.
  • Elicitation methods: Spontaneous human interactions offer fully naturalistic emotional displays, and SEWA is presented as the first in-the-wild collection of annotated interactive behaviour.
  • Emotion representation: Continuous multidimensional models describe affect more finely than a small set of emotion categories, including arousal and valence.
  • Existing corpora: Existing dyadic corpora often use controlled tasks, limited demographic variability, and primarily English recordings.
  • Existing corpora: The reviewed corpora vary in scope, including culturally diverse interaction data, freely led English conversations, conflict intensity, and film-derived valence-arousal clips.

3 SEWA DATABASE

SEWA DB is designed to provide sufficient labelled data for robust automatic understanding of human behaviour. The paper describes its collection process, diverse annotations, recording statistics, and searchable web resource.

  • Database aim: SEWA DB aims to facilitate robust automatic machine understanding of human behaviour through sufficient labelled examples.

3.1 Data collection

SEWA collected paired participants’ audio-visual behaviour through a dedicated website in two stages: advertisement viewing followed by video-chat discussion. The design targeted naturalistic, cross-cultural interaction under participants’ own recording conditions.

  • Participant setup: Participants were paired by cultural background, age, and gender and completed demographic, personality, and familiarity questionnaires.
  • Recording setup: A WebRTC/OpenTok website played adverts, enabled video chat, and synchronously recorded webcam and microphone data on participants’ own computers.
  • Advertisement viewing: Each participant watched four approximately 60-second adverts selected to elicit amusement, empathy, liking, and boredom.The adverts used visuals and music without dialogue to support consistent understanding across cultures.
  • Advertisement viewing: After viewing the adverts, participants self-reported their emotional state and sentiment toward each advert.
  • Video-chat discussion: After the fourth advert, paired participants discussed it through video chat for an average of three minutes.The discussion elicited further reactions and opinions about the advert and product.

3.2 The Data Statistics and subject demographics

SEWA comprises 398 participants from six cultures, with near-balanced gender representation and coverage across five adult age groups. The collection includes 1990 audio-visual clips spanning advert reactions and video-chat interactions.

  • 398 subjects participated across six cultural backgrounds: British, German, Hungarian, Serbian, Greek, and Chinese.
  • 201 males and 197 females produced a near-balanced gender ratio of 1.020.
  • Participants were divided into five age groups from 18–29 through 60+, with the youngest group most numerous.
  • 1990 clips were collected, including four advert-watching clips and one video-chat clip per subject.
  • The recordings comprise 1600 minutes of advert reactions and 1057 minutes of video-chat data.

3.3 Data annotation

SEWA uses broad, multimodal annotation covering facial, vocal, verbal, gestural, emotional, and social behaviours. Because manual annotation was costly at this scale, several components were produced through semi-automatic procedures with targeted manual correction.

  • The database annotates facial landmarks, acoustic descriptors, hand and head gestures, facial action units, verbal and vocal cues, affect dimensions, templates, agreement, and mimicry.
  • The fully annotated basic SEWA dataset contains 538 short video-chat segments selected across emotional states and cultures.
  • Facial landmarks: 49-point facial landmarks were annotated for every basic-dataset segment.
  • Facial landmarks: 95.1% of Chehra tracking results were accurate without correction, while the remaining 18,099 frames received manual annotation.
  • Gestures: Hand gestures used five labels, while head-gesture annotation identified 282 nod sequences and 122 head-shake sequences.
  • Facial action units: Five sentiment-relevant AUs—AU1, AU2, AU4, AU12, and AU17—were annotated semi-automatically using trained detectors and manual false-positive removal.

3.4 Database availability

SEWA is distributed through an online portal with filtering by recording characteristics and annotation availability. Access is restricted to academic researchers who accept the database’s usage agreement.

  • The online portal provides search filters for demographics and available annotation types.
  • The database is intended for academic use only, with commercial and other non-academic uses prohibited by the participant consent terms.
  • Researchers must sign an EULA, and authorized data transfers are protected with SSL encryption.

4 BASELINE EXPERIMENTS

The paper reports baseline experiments for action-unit detection and continuous estimation of valence, arousal, and liking/disliking. The baselines span conventional regressors, recurrent neural networks, and a deep convolutional model.

  • The baseline experiments address action-unit detection and valence, arousal, and liking/disliking estimation.
  • The regression baselines use a linear-kernel Support Vector Regressor and a Random Forest Regressor.
  • An LSTM recurrent neural network serves as a sequential baseline capable of learning long-range contextual information.
  • A ResNet-18 deep convolutional network was trained for valence and arousal using either RMSE or CCC loss.

4.1 Feature extraction

The experiments combine geometric and appearance-based video features with acoustic descriptors, using early fusion to create multimodal inputs. Audio features include both broad COMPARE and compact GEMAPS sets, while baseline models use summarized COMPARE descriptors.

  • Video features: Video features combine dense SIFT appearance descriptors around normalized facial landmarks with landmark-based geometric shape features.Landmarks are normalized for translation, scaling, and rotation before feature extraction.
  • Audio features: Audio features use frame-wise low-level descriptors extracted with openSMILE at a 10 ms step size.The feature sets include COMPARE and GEMAPS acoustic descriptors.
  • Audio features: COMPARE contains 65 low-level descriptors spanning spectral, cepstral, prosodic, and voice-quality information, whereas GEMAPS contains 18 descriptors selected for affective voice analysis.COMPARE applies deltas and functionals to frame-level descriptors over the audio signal.
  • Feature fusion: Early fusion concatenates feature vectors from different modalities into one input vector for simple, reproducible multimodal experiments.The study does not use late or mid-level fusion.
  • Baseline representation: Baseline experiments summarize COMPARE descriptors over six-second blocks using means and standard deviations, producing 130-dimensional vectors.The study uses COMPARE because GEMAPS features were mostly redundant and did not yield superior average results.

4.2 Experimental setting

The study evaluates affect recognition in mixed-culture and culture-specific settings using person-independent partitions. It uses separate metrics for action-unit classification and continuous affect regression, with both conventional and deep-learning splits.

  • Experimental settings: Experiments include a mixed-culture, person-independent setting and six culture-specific, person-independent settings.The culture-specific settings separately divide each culture into training, validation, and testing data.
  • Data partitioning: Conventional experiments use subject-independent training, development, and test partitions with a 3:1:1 ratio balanced for age, gender, and selection criteria.Model parameters are selected by grid search on the development set.
  • Deep experiments: Deep experiments use subject-independent training, development, and test partitions with an 8:1:1 ratio.Valence and arousal models optimize either RMSE or CCC, with CCC optimization yielding better results.
  • Performance measures: Action-unit detection is evaluated as classification with F1, while valence, arousal, and liking/disliking estimation are evaluated as regression.F1 is used because action-unit data commonly contain imbalanced positive and negative samples.
  • Performance measures: Continuous affect estimation reports Pearson correlation and concordance correlation coefficient metrics.CORR and CCC compare predictions with corresponding ground-truth label series.

4.3 Experimental results

Baseline results show that performance depends on modality, culture, model, and target dimension. Geometric and appearance features support action-unit and valence estimation, audio supports arousal, and deep models particularly improve arousal estimation.

  • Action unit detection: AU 12 reaches an F1-score of 0.618 with feature fusion and an SVM classifier, while the detectors remain insufficient for fully automatic AU detection.Landmark features generally outperform texture features, but results are lower than overlapping FERA2015 baselines in controlled recordings.
  • Modality effects: Valence is better estimated from video than audio, whereas arousal is better predicted from audio than video.This modality pattern appears in both annotation and feature comparisons reported by the experiments.
  • Regression models: SVM generally outperforms Random Forest, which outperforms LSTM, but LSTM achieves a liking-prediction CCC of 0.254 versus 0.194 for SVR and 0.087 for RF.The exception occurs for liking/disliking prediction using audio features and audio-plus-video annotations.
  • Target dimensions: Liking/disliking performance is generally lower than valence and arousal performance, and may benefit from additional multimodal data.The paper links this difficulty to the content-related nature of liking/disliking and limited linguistic information in acoustic cues.
  • Cultural effects: For culture-specific results, Hungarian valence reaches CCC 0.495 with video features, while German arousal reaches CCC 0.694 with audio features.The strongest fused-feature arousal result is CCC 0.501 for culture 2.
  • Deep models: ResNet-18 outperforms other deep baseline models, especially when optimizing CCC directly.Deep representations particularly improve arousal, while valence remains accurately predicted from geometric facial features at higher computational cost.

5 CONCLUSION

SEWA DB is a publicly available benchmark for automatic audio-visual behaviour analysis in the wild, combining recordings from multiple cultures with extensive behavioural and affective annotations. The paper also provides baseline experiments for action-unit detection and valence, arousal, and liking/disliking prediction.

  • SEWA DB was made publicly available as a benchmark for automatic analysis of audio-visual behaviour in the wild.
  • SEWA DB contains 204 experiment sessions with 408 subjects from six cultural backgrounds and 1525 minutes of advert-reaction plus 568 minutes of video-chat recordings.The recorded cultures are British, German, Hungarian, Greek, Serbian, and Chinese.
  • The authors provide exhaustive baseline experiments for action-unit detection and valence, arousal, and liking/disliking prediction.These experiments are intended to provide comparison benchmarks for affect estimation.
  • The corpus is positioned as a resource for psychological hypothesis testing and for advancing automatic sentiment analysis in the wild.
Loading 1901.02839v2…