Source-linked AI summary

EmoBank: Studying the Impact of Annotation Perspective and Representation Format on Dimensional Emotion Analysis

Sven Buechel, Udo Hahn

arXiv:2205.01996v1cs.CLcs.AIcs.LG

TL;DR

Emotion analysis lacked large, reliable resources using psychologically grounded models and distinguishing writer from reader perspectives. EmoBank addresses this with a genre-balanced VAD corpus, bi-perspectival annotations, and a bi-representational subset. Reader ratings show better correlation-based agreement and greater emotionality, while mappings between formats achieve near-human performance.

  • Problem

    Existing VA(D)-annotated corpora were rare, small, or limited in reliability, and fine-grained resources lacked adequate writer-reader perspective distinctions.

  • Method

    EmoBank constructs a genre-balanced VAD corpus with writer and reader annotations, plus a subset pairing VAD with categorical Basic Emotion labels.

  • Results

    Reader ratings yield better correlation-based IAA and higher emotionality, while automatic mappings between dimensional and categorical formats achieve near-human performance.

  • Takeaways & Limitations

    EmoBank supports separate analysis of writer and reader emotion and enables comparison between dimensional and categorical emotion formats.

  • Takeaways & Limitations

    Statistical significance is rare in this setup because the number of cases is based on the number of raters.

Abstract

from arXiv · show

We describe EmoBank, a corpus of 10k English sentences balancing multiple genres, which we annotated with dimensional emotion metadata in the Valence-Arousal-Dominance (VAD) representation format. EmoBank excels with a bi-perspectival and bi-representational design. On the one hand, we distinguish between writer's and reader's emotions, on the other hand, a subset of the corpus complements dimensional VAD annotations with categorical ones based on Basic Emotions. We find evidence for the supremacy of the reader's perspective in terms of IAA and rating intensity, and achieve close-to-human performance when mapping between dimensional and categorical formats.

1 Introduction

Fine-grained emotion analysis still lacks psychologically adequate models and resources that distinguish writer and reader perspectives. EmoBank addresses both gaps with a genre-balanced, VAD-based corpus combining bi-perspectival and partially bi-representational annotations.

  • Fine-grained affective-language research has moved beyond binary polarity toward multiple classes and real-valued sentiment scores.
  • Psychologically adequate emotion models and separate writer-versus-reader perspectives remain insufficiently resourced.
  • EmoBank is a large-scale corpus based on Valence-Arousal-Dominance and balanced across genres.
  • EmoBank distinguishes writer and reader emotions through bi-perspectival annotation.
  • A subset combines VAD ratings with Ekman Basic Emotion annotations, enabling mappings between dimensional and categorical formats.

2 Related Work

Emotion analysis uses categorical and dimensional representations, with VAD offering independent dimensions for more comparable fine-grained analysis. Existing dimensional corpora are scarce or small, while writer-reader distinctions address increasingly detailed affective interpretation.

  • Dimensional models represent emotion using Valence, Arousal, and Dominance in a three-dimensional real-valued space.Valence corresponds to polarity, Arousal to calmness or excitement, and Dominance to perceived control.
  • Existing VA(D)-annotated corpora are rare, small, or limited in reliability, including resources with 120, 2,895, and 2,009 sentences.
  • Categorical emotion studies often use incompatible category subsets, limiting comparability across studies.
  • VAD dimensions are designed as independent, so results remain comparable dimension-wise even when a dimension such as Dominance is absent.
  • Fine-grained analysis increasingly separates emotion expressed by writers from emotion evoked in readers.A sentence can be neutral from a writer’s viewpoint while evoking adverse emotions in readers.
  • Prior work modeled writer and reader emotions or used reader reactions as proxies, but effects on annotation quality had received limited investigation.

3 Corpus Design and Creation

EmoBank was assembled from genre-diverse English sources and annotated through crowdsourcing for both writer and reader VAD perspectives. A subset also supports categorical comparison, while quality control removed a heavily overrepresented rating.

  • Corpus selection targeted several genres and domains of general English to complement social-media and review-focused resources.
  • The corpus combines categories from MASC with SemEval-2007 Affective Text, which contributes Basic Emotion annotations.
  • The corpus was annotated using crowdsourcing through CrowdFlower, selected for its quality-control mechanisms and accessibility.
  • Each sentence received independent writer- and reader-perspective tasks using modified 5-point SAM scales, with linguistic clues for writers and average-reader judgments for readers.
  • Five annotators produced ratings for each perspective and each VAD dimension, yielding 30 ratings per sentence.
  • About 10% of ratings for each task were removed after the (1, 1, 1) rating was judged heavily overrepresented and psychologically improbable.
  • EmoBank became, to the authors’ knowledge, the largest corpus for dimensional emotion models and one of the largest gold standards for emotion formats.

4 Analysis and Results

Across VAD dimensions, READER annotations achieved higher correlation-based IAA and greater emotionality than WRITER annotations, while also showing higher error. Error differences tracked emotionality differences closely, without disproportionate error at equal emotionality.

  • Inter-annotator agreement: IAA exceeded r > .6 for both perspectives, while READER produced significantly higher correlation and higher error than WRITER.The significance pattern was p < .05 for Valence in r and for all dimensions in MAE.
  • Rating emotionality: READER ratings were significantly more emotional than WRITER ratings, using distance from the neutral rating 3 averaged across VAD dimensions.Emotionality was treated as a complementary quality criterion to IAA.
  • Error and emotionality: The analysis computed sentence-wise error separately for WRITER and READER, then compared perspective differences in error and emotionality sentence by sentence.V, A, and D were represented as m × n matrices of ratings, with m sentences and n annotators.
  • Error and emotionality: r = .718 linked differences in sentence emotionality with differences in annotation error between the two perspectives.The regression line passed through the origin, with intercept p = .992, indicating no average error difference when emotionality differences were absent.

5 Mapping between Emotion Formats

The study tests whether VAD scores can predict Basic Emotion annotations, comparing writer, reader, and combined perspectives. Mapping approaches approach human agreement, with reader-based features outperforming writer-based features and combined features performing best.

  • Experimental setup: 10-fold cross-validation trains one k-nearest-neighbor model per Basic Emotion using writer, reader, or combined VAD features.The models use the bi-representational SE07 subset and select hyperparameters through cross-validation.
  • Results: Reader-based VAD features outperform writer-based features for mapping dimensional scores to categorical Basic Emotions.The comparison uses correlations between model predictions and categorical annotations.
  • Results: Combined writer-and-reader VAD features perform even better, matching or exceeding average human agreement.The comparison is made against human IAA reported for the categorical annotations.
  • Results: For Joy, Anger, Sadness, and Fear, the best models average above human agreement, while human IAA is below .5 for Disgust and Surprise.Human agreement varies substantially across Basic Emotion categories.

6 Conclusion

EmoBank is a genre-balanced, large-scale VAD corpus with writer-reader and dimensional-categorical annotations. Reader ratings show higher agreement and intensity, while automatic format mapping reaches near-human performance.

  • Corpus design: EmoBank is a genre-balanced large-scale corpus using the dimensional VAD model and providing writer and reader emotion annotations.A subset also contains categorical Basic Emotion and VAD ratings.
  • Annotation perspectives: The reader perspective yields better inter-annotator agreement and more emotional ratings than the writer perspective.
  • Format mapping: Automatic mapping between categorical and dimensional formats achieves near-human performance using standard machine-learning techniques.

A Appendix: Instructions

The appendix documents two annotation tasks based on Bradley and Lang’s VAD instructions, with adaptations for crowdworkers and distinct writer-reader perspectives.

  • Task design: The two tasks annotate author emotion and reader emotion using instructions adapted from Bradley and Lang’s VAD lexicon.
  • Task design: The tasks replace 9-point VAD scales with 5-point scales to reduce crowdworkers’ cognitive load.
  • Perspective-specific instructions: Writer-perspective instructions provide linguistic hints, whereas reader-perspective instructions ask about an average reader’s emotion rather than personal feelings.

Expressing Emotion

The Expressing Emotion instructions ask annotators to rate the author’s expressed feelings on independent Pleasure, Arousal, and Control scales. They define each scale’s endpoints and provide heuristics for consistent judgments.

  • Scale framework: Annotators rate the author’s expressed feeling on three independent SAM scales: Pleasure, Arousal, and Control.The task concerns how the author feels while writing, not how the reader feels.
  • Pleasure: Pleasure ranges from unhappiness, annoyance, or boredom to happiness, satisfaction, contentment, or hope.
  • Arousal: Arousal ranges from relaxation and calmness to stimulation, excitement, frenzy, nervousness, or arousal.
  • Control: Control ranges from submissiveness and being controlled to feeling influential, autonomous, dominant, or controlling.
  • Rating heuristics: Capitalization, exclamation marks, and swearing may indicate high Arousal, while commands may indicate high Control.
  • Rating heuristics: Annotators should focus on the author’s feeling, choose the most likely interpretation, use the full scale range, and avoid overinterpreting sentences.

Evoking Emotion

The study measures how people would feel when reading sentences using three independent SAM scales: Pleasure, Arousal, and Control. Participants are instructed to rate average reader reactions rather than personal feelings or presumed writer emotions.

  • SAM measures Pleasure, Arousal, and Control as three independent multiple-choice scales for rating how people feel when reading each sentence.Pleasure ranges from unhappy to happy, Arousal from calm to excited, and Control from controlled to in control.
  • Participants imagine how people would react on average while ignoring their personal opinions, beliefs, and experiences.
  • Pleasure captures feelings from unhappy, annoyed, or bored to happy, pleased, contented, or hopeful.The scale uses intermediate choices for intermediate levels and extreme options for completely unhappy or completely happy reactions.
  • Arousal captures feelings from relaxed, calm, or sleepy to stimulated, excited, frenzied, or wide-awake.The middle option represents feeling neither excited nor calm, with intermediate options indicating degrees between the endpoints.
  • Control captures whether people feel controlled, influenced, or submissive versus in control, influential, dominant, or autonomous.The scale contrasts being controlled by a situation with feeling controlling or in control.
  • Raters should focus on how people in general would feel after reading a sentence, choose the most likely interpretation, and use the full scale range.They are also asked to work rapidly rather than overinterpret individual sentences.
Loading 2205.01996v1…