Source-linked AI summary

A dataset of continuous affect annotations and physiological signals for emotion analysis

Karan Sharma, Claudio Castellini, Egon L. van den Broek, Alin Albu-Schaeffer, Friedhelm Schwenker

arXiv:1812.02782v1cs.HCcs.LG

TL;DR

Emotion assessment is difficult because internal emotional states cannot be directly inspected in realistic settings. CASE addresses this gap with continuous simultaneous valence-arousal annotation and synchronized physiological recordings from 30 participants watching validated videos, while exploratory analyses show matching annotation and physiological patterns.

  • Problem

    Internal emotions are inaccessible to external systems, making real-time emotion assessment difficult and motivating measurable signals paired with subjective annotations.

  • Method

    CASE collects continuous simultaneous valence-arousal annotations with JERI alongside physiological recordings while participants watch video stimuli.

  • Results

    Exploratory analyses found that physiological feature patterns matched joystick annotations across video types, while prior analyses validated annotation differences and usability.

  • Takeaways & Limitations

    CASE provides a dataset for studying relations between physiological responses and continuous emotional annotations in affective computing and psychology.

  • Takeaways & Limitations

    The raw log files and preprocessing code were not released because the raw files lacked video identifiers.

Abstract

from arXiv · show

From a computational viewpoint, emotions continue to be intriguingly hard to understand. In research, direct, real-time inspection in realistic settings is not possible. Discrete, indirect, post-hoc recordings are therefore the norm. As a result, proper emotion assessment remains a problematic issue. The Continuously Annotated Signals of Emotion (CASE) dataset provides a solution as it focusses on real-time continuous annotation of emotions, as experienced by the participants, while watching various videos. For this purpose, a novel, intuitive joystick-based annotation interface was developed, that allowed for simultaneous reporting of valence and arousal, that are instead often annotated independently. In parallel, eight high quality, synchronized physiological recordings (1000 Hz, 16-bit ADC) were made of ECG, BVP, EMG (3x), GSR (or EDA), respiration and skin temperature. The dataset consists of the physiological and annotation data from 30 participants, 15 male and 15 female, who watched several validated video-stimuli. The validity of the emotion induction, as exemplified by the annotation and physiological data, is also presented.

Background & Summary

Affective computing addresses the difficulty of inferring inaccessible internal emotions by combining measurable signals, subjective annotations, and predictive modelling. CASE contributes a dataset with physiological data and continuous, simultaneous valence-arousal annotations from participants watching videos.

  • Affective computing develops machines that recognise, interpret, and adapt to human emotions.
  • The standard pipeline acquires measurable indicators and subjective annotations, then models their relation to predict emotional states.
  • Continuous-annotation datasets address limitations of discrete emotion labels, but commonly annotate valence and arousal separately with mouse-based tools.
  • CASE records physiological signals and continuous emotional annotations from 30 subjects watching video stimuli while reporting their experience with JERI.
  • CASE is presented as the first dataset featuring continuous and simultaneous annotation of valence and arousal.

Methods

The study used a within-subjects experiment in which 30 volunteers watched and annotated multiple video stimuli. Video order was varied between participants to reduce carry-over effects, with interleaved blue-screen rest periods.

  • 30 volunteers, comprising 15 males and 15 females aged 22–37, participated in the data-collection experiment.
  • The within-subjects design required every participant to watch and annotate all video stimuli.
  • Video sequences differed for each participant through pseudo-random ordering to avoid carry-over effects.
  • The setup showed participants watching videos while annotating with the joystick-based JERI interface.
  • Two-minute blue-screen intervals separated videos, isolating emotional responses and allowing participants to rest.

Experiment Protocol

Participants received experiment instructions, physiological sensors, and training on the annotation procedure before viewing stimuli. The interface enabled continuous two-dimensional emotion traces during videos selected to elicit four emotional states.

  • Participants received oral and written instructions, signed informed consent, and were introduced to the 2D circumplex model before annotation.
  • The interface supplemented valence and arousal axes with Self-Assessment-Manikin guides and produced a continuous 2D trace during each video.
  • Figure 2 displays one participant’s annotations across videos and the ‘scary-2’ annotations from participants p1–p5.
  • The experiment targeted amusing, boring, relaxing, and scary states using video stimuli previously used in other studies.
  • Table 1 documents each video’s source, label, identifier, intended valence-arousal attributes, and duration.

Sensors & Instruments

The experiment combined physiological sensors with a participant-controlled joystick and a digitisation system. The instruments measured cardiac, electrodermal, respiratory, temperature, and facial or upper-back muscle activity.

  • The sensor selection followed prevalence in affective-computing datasets and applications, manufacturer recommendations, or compatibility with the acquisition setup.
  • The ECG measured electrical activity from three chest electrodes arranged in a triangular configuration.
  • The BVP/PPG sensor used reflected light variation at the non-dominant middle finger to measure cardiac activity.
  • The GSR/EDA sensor measured sweat-related changes in skin electrical conductance using electrodes on the index and ring fingers.
  • Respiration, skin temperature, and muscle activity were measured with chest, finger, and three EMG sensors targeting facial and trapezius muscles.
  • A participant-controlled joystick provided proprioceptive feedback intended to reduce cognitive load during simultaneous viewing and annotation.

Data Acquisition

The experiment integrated video playback, joystick annotation, and physiological acquisition through LabVIEW, with arrows and line styles distinguishing data flow and acquisition tasks.

  • LabVIEW directly managed video playback, the annotation interface, and physiological data acquisition components.
  • Figure 3 uses arrows to show data-flow direction and solid versus dotted lines to distinguish primary and secondary acquisition tasks.
  • Annotation data was acquired at 20 Hz, while physiological data was acquired at 1000 Hz.

Data Preprocessing

The preprocessing pipeline converted raw logs into standardized, video-labeled CSV records, correcting sensor scales and timing through interpolation while preserving non-interpolated data.

  • The preprocessing workflow iteratively processed each participant’s log files after determining video durations once.
  • Raw sensor voltages were transformed using Table 2 equations, and annotation values were rescaled from [−26225 . . 26225] to [0.5 . . 9.5].
  • Linear interpolation generated 1 ms physiological and 50 ms annotation sampling intervals to address logging latencies.
  • Video durations and sequence lookups were used to assign video-IDs to both interpolated and non-interpolated log files.
  • Processed physiological and annotation data were saved as separate CSV files for broad programming and scientific-computing accessibility.

Code availability

The experiment and preprocessing code were not released because they depend on specific equipment or unavailable raw logs, although researchers may request the materials.

  • The LabVIEW experiment and acquisition code was withheld because it is highly specific to the sensors and equipment used.
  • Preprocessing code was not released because the raw log files it processes are excluded from the dataset due to missing video-IDs.
  • The raw data and preprocessing code are available to interested researchers upon request.

Data Records

CASE is distributed as an archive containing interpolated and non-interpolated records, metadata, README files, participant and video information, and separate annotation and physiological CSV data.

  • The dataset is hosted as a single figshare archive with interpolated, non-interpolated, and metadata directories.
  • Metadata: Metadata includes participant demographics and video-sequence ordering, while video records provide durations, source links, URLs, and retrieval time-windows.
  • Directory organization: Interpolated and non-interpolated directories contain separate annotations and physiological subdirectories with similar structures and naming conventions.
  • Annotation records: Each annotation directory contains 30 participant CSV files recording timestamps, valence, arousal, and video identifiers.
  • Physiological records: Each physiological directory contains 30 participant CSV files with timestamps, eight transformed sensor outputs, and video identifiers.
  • Time alignment: Annotation and physiological timestamps share a common logging-computer clock but use 50 ms and 1 ms sampling intervals, respectively.

Technical Validation

Technical validation examined selected physiological features and continuous valence–arousal annotations across video stimuli. The analyses found stimulus-specific physiological patterns and concurrent clustering between physiological and annotation data.

  • Annotation Data: The annotation data’s reliability and usability were previously evaluated through exploratory analyses, MANOVA, and System Usability Scale ratings.Scary videos showed low valence and high arousal, while amusing videos showed relatively high valence and medium arousal; the setup received excellent usability ratings.
  • Feature Extraction: Feature extraction segmented each participant’s recordings by video and computed physiological features from each video-specific chunk.The selected-feature analysis used one predominantly used feature per sensor and mean values across each video chunk.
  • Feature Extraction: Figure 4 summarizes distributions of selected features and mean valence–arousal annotations across video types using violin plots, IQR box plots, and mean markers.The figure’s yellow diamonds mark distribution means.
  • Feature Extraction: Scary videos produced high SCR and elevated HR, whereas amusing videos produced accelerated respiration and greater zygomaticus activity.Boring and relaxing videos elicited similar values across features; reported arousal was higher for scary videos, while zygomaticus activity followed valence patterns.
  • Feature Extraction: PCA projected the selected Z-scored physiological features into two dimensions for comparison with mean valence and arousal across video labels.The scatter plots include one-standard-deviation ellipses for each video type.
  • Feature Extraction: Physiological and annotation data formed concurrent clusters, with scary and amusing videos occupying distinct regions and boring and relaxing videos grouped at low arousal.The authors describe this as an initial investigation requiring more rigorous analysis.

Usage Notes

The dataset cannot include the source videos because of copyright restrictions. Supporting code is available to interested researchers upon request.

  • Videos: The source videos are not directly distributed because of copyright issues.Links to currently hosted copies are provided for examining emotional content and potentially replicating the experiment.
  • Videos: The provided video links may become unusable in the future, potentially complicating access and replication.The authors invite users to contact them for assistance acquiring or editing the videos if links fail.
  • Implementation: Technical-validation code was developed in MATLAB 2014b and R 3.3.3 and is available to interested researchers upon request.Feature extraction used MATLAB tools including TEAP and Pan–Tompkins QRS detection, while PCA used R’s prcomp function.

Data Citations

The paper cites the CASE dataset as a 2018 data resource containing the study’s physiological and annotation recordings.

  • Data Citations: The CASE dataset is cited as Sharma et al. (2018) and is available as a downloadable dataset.The citation provides the dataset name, authors, year, and download address.
Loading 1812.02782v1…