Source-linked AI summary
K-EmoCon, a multimodal sensor dataset for continuous emotion recognition in naturalistic conversations
Cheul Young Park, Narae Cha, Soowon Kang, Auk Kim, Ahsan Habib Khandoker, Leontios Hadjileontiadis, Alice Oh, Yong Jeong, Uichin Lee
TL;DR
Existing emotion datasets provide limited support for studying idiosyncratic emotions in naturalistic social interactions. K-EmoCon constructs a multimodal dataset from paired debates with continuous annotations and self, partner, and observer perspectives. It provides a publicly available basis for multiperspective emotion assessment, while its participant demographics and debate setting constrain generalizability.
Problem
Existing emotion datasets often come from constrained environments and therefore provide limited support for studying emotions arising in the wild.
Method
K-EmoCon combines physiological and audiovisual measurements with continuous emotion annotations collected during paired social-issue debates from self, partner, and external-observer perspectives.
Results
K-EmoCon is presented as the first dataset with emotion annotations from all three available perspectives: the subject, debate partner, and external observers.
Takeaways & Limitations
The dataset supports examining cases where perceptions of emotions do not match across participants and observers.
Takeaways & Limitations
The young, highly educated, predominantly Asian participant and rater sample may not generalize well to different ethnic or age groups.
Abstract
from arXiv · showhide
Recognizing emotions during social interactions has many potential applications with the popularization of low-cost mobile sensors, but a challenge remains with the lack of naturalistic affective interaction data. Most existing emotion datasets do not support studying idiosyncratic emotions arising in the wild as they were collected in constrained environments. Therefore, studying emotions in the context of social interactions requires a novel dataset, and K-EmoCon is such a multimodal dataset with comprehensive annotations of continuous emotions during naturalistic conversations. The dataset contains multimodal measurements, including audiovisual recordings, EEG, and peripheral physiological signals, acquired with off-the-shelf devices from 16 sessions of approximately 10-minute long paired debates on a social issue. Distinct from previous datasets, it includes emotion annotations from all three available perspectives: self, debate partner, and external observers. Raters annotated emotional displays at intervals of every 5 seconds while viewing the debate footage, in terms of arousal-valence and 18 additional categorical emotions. The resulting K-EmoCon is the first publicly available emotion dataset accommodating the multiperspective assessment of emotions during social interactions.
1 Background & Summary
Emotion recognition needs data that captures elusive, context-dependent emotions in realistic interactions, because laboratory, media, and crowdsourced datasets may generalize poorly to the wild. K-EmoCon addresses this gap with multimodal recordings and annotations from multiple perspectives during paired social-issue debates.
- Emotions are difficult to measure because facial expressions can be compound, relative, and misleading rather than discrete signals.
- Laboratory emotion datasets offer experimental control but may generalize poorly because they often contain intense prototypical expressions from limited populations.
- Media-content and crowdsourced datasets increase sample size and subject diversity, but generalizability remains an issue.
- K-EmoCon contains physiological signals, audiovisual footage, and continuous emotion annotations from 32 subjects in 16 paired social-issue debates.
- K-EmoCon includes emotion annotations from the subject, debate partner, and external observers, enabling multiperspective assessment.
2 Methods
K-EmoCon is designed for naturalistic social-interaction emotion research by combining multimodal sensing with annotations from the subject, partner, and external observers. Its debate setting and perspective distinctions support studying differing emotion perceptions and recognition.
- Dataset design: The dataset is intended to extend research on whether multiple perspectives improve automatic emotion recognition.
- Dataset design: K-EmoCon combines self, partner, and external-observer annotations to examine how emotional expressions are perceived from multiple perspectives.The subject reports felt emotions, the partner has contextual knowledge of the interaction, and external observers lack its exact context.
- Dataset design: K-EmoCon further supports investigation of perception mismatches by distinguishing observers according to whether they have contextual knowledge of the emotion-generating situation.
- Data collection setting: Data were collected in semi-structured, turn-taking debates between randomly assigned partners, a setting chosen to approximate workplace-like social interaction and elicit naturally arising emotions.Its formality and spontaneity were expected to expose less pronounced emotions that partners might misperceive.
- Data collection apparatus: The apparatus combines low-cost wearable physiological sensors with audiovisual recordings to support reproducible, expandable study of emotion-perception mismatches in the wild.The dataset was also intended to support research on wearable biosignals for affective communication.
- Data collection procedure: Participants annotated felt emotions from third-person recordings rather than head-mounted first-person footage, which may have made the measurement less naturalistic.
3 Data Records
K-EmoCon packages multimodal recordings, physiological measurements, synchronized metadata, and continuous emotion annotations from 16 paired debates. Its records span raw sensor files, audiovisual data, annotation perspectives, timestamps, and signal-quality tables.
- Annotations: External-rater metadata lists five raters, while aggregated external annotations combine their labels through majority voting.External annotation files use participant and rater identifiers, and aggregated files store the resulting vote.
- Dataset summary: 16 paired debates produced 172.92 minutes of dyadic interaction with physiological signals, audiovisual recordings, and three-perspective emotion annotations.The dataset includes annotations from subjects, debate partners, and external observers.
- Data quality: Quality tables report file durations, zeros, outliers, and completeness ratios for the physiological recordings.Completeness is computed from total values after accounting for outliers and zeros, with separate formulas for E4 and NeuroSky/Polar data.
- Audiovisual records: Audiovisual records include 16 WAV debate audios and MP4 participant recordings, with filenames encoding participant identities or recording duration.Audio boundaries correspond to debate timestamps in subjects.csv.
- Physiological records: NeuroSky and Polar files provide attention, brainwave, meditation, and heart-rate measurements, while E4 files provide accelerometer, BVP, EDA, heart-rate, IBI, and temperature data.The records document sampling rates, units, EEG frequency bands, and conversions such as HR from IBI.
- Annotations: Emotion annotations are organized into self, partner, external, and aggregated-external directories, with annotations acquired every 5 seconds.External annotations are provided per rater, alongside majority-vote aggregates.
4 Technical Validation
Technical validation characterizes annotation distributions, inter-rater reliability, signal quality, and missing data. Emotion annotations are strongly imbalanced, while reliability is generally low and depends on preprocessing choices and annotation perspective.
- Emotion distributions: Emotion annotations are predominantly neutral or concentrated in two categorical affective categories.Likert-scale emotions are biased toward neutral, while categorical annotations are concentrated mainly in concentration and none.
- Inter-rater reliability: Krippendorff’s alpha measured inter-rater reliability across seven ordinal emotions and four combinations of annotation perspectives.The analysis compared self–partner, self–external, partner–external, and all-perspective agreement.
- Annotation caveats: Scaled emotions lacked a zero-neutral reference for five categories, producing widely varying interpretations among participants and raters.Arousal and valence were centered at zero, whereas cheerful, happy, angry, nervous, and sad used scales from 1 to 4.
- Inter-rater reliability: Mode subtraction was used to measure agreement in relative emotional changes rather than absolute ratings and mitigate spuriously low alpha values.Annotation values were interpreted as ordinal, and column modes were treated as neutral baselines before reliability computation.
- Inter-rater reliability: 20 of 32 participants showed higher self–external than self–partner arousal reliability, with mean difference 0.143 and standard deviation 0.322.Alpha coefficients remained low overall, and the perspective difference may indicate differing emotion perceptions, although further validation is required.
- Signal quality and missing data: Physiological measurements were examined for quality, but E4 data from four participants were excluded because of device malfunction.Most files were more than 95% complete, although additional missingness affected IBI, EDA, NeuroSky, and Polar HR measurements.
5 Usage Notes
Usage of K-EmoCon should account for measurement artifacts, debate-specific emotional regulation, demographic narrowness, and unmeasured interpersonal and participant variables. These factors constrain interpretation and generalization of analyses using the dataset.
- Potential usage: The dataset can support analyses of physiological markers related to emotion regulation and models of individual emotional profiles from sensor-based recordings.Suggested applications include examining cognitive appraisal and applying machine-learning methods to physiological and behavioral data.
- Data collection apparatus: Contact-based EEG sensors may contain noise from frowning or eye movements, and other devices may have similar systematic errors.These artifacts can affect physiological measurements collected during the debates.
- Data collection context: The turn-taking debate may have suppressed emotional expressions and deflated agreement between self-reports and partner or external perceptions.This limitation may not apply to the same extent in more natural interactions in the wild.
- Demographics: The participant sample was young, highly educated, and predominantly Asian, limiting generalization across ethnicities and age groups.Participants and raters were between 19 and 36 years old.
- Unaccounted variables: Rapport, spoken-English competence, and familiarity with the debate topic were unaccounted variables that may affect cross-perspective emotion mismatches.These variables may contribute to variance in the observed mismatch between emotional perceptions.
6 Code Availability
The project provides code for key preprocessing, annotation, reliability, and visualization utilities, while withholding privacy-sensitive raw log-level data and its preprocessing code.
- Available code: Python code implements outlier detection, majority voting, mode subtraction, utility functions, and heatmap generation.The repository also uses the Krippendorff Python package for alpha computation and Python 3.6.9.
- Availability constraints: Raw log-level data and the SQL-to-CSV preprocessing code are not publicly available because they contain privacy-sensitive information outside the agreed sharing boundary.Users may contact the corresponding authors for further assistance or information.