Source-linked AI summary

Crowdsourcing a Word-Emotion Association Lexicon

Saif M. Mohammad, Peter D. Turney

arXiv:1308.6297v1cs.CL

TL;DR

Emotion analysis has lacked large, high-coverage emotion lexicons because traditional annotation is costly. This paper uses Mechanical Turk crowdsourcing to construct EmoLex and develops quality controls for sense-level annotation. The resulting resource covers more than 10,000 word–sense pairs, with eight-emotion and polarity annotations.

  • Problem

    High-quality, high-coverage emotion lexicons are scarce, while traditional expert annotation requires substantial cost and manual effort.

  • Method

    The paper crowdsources term–emotion annotation through Mechanical Turk, using word-choice questions to detect erroneous or unqualified annotations and convey the requested word sense.

  • Results

    EmoLex contains more than 10,000 word–sense pairs annotated for eight basic emotions and positive, negative, or neutral semantic orientation.

  • Takeaways & Limitations

    Crowdsourcing yields a large term–emotion resource whose annotations are reported as high quality and include emotion associations, emotion terms, co-occurring emotions, and polarity.

  • Takeaways & Limitations

    The annotation framework is limited to Plutchik’s eight basic emotions, selected partly because broader emotion schemes would be expensive and difficult to annotate.

Abstract

from arXiv · show

Even though considerable attention has been given to the polarity of words (positive and negative) and the creation of large polarity lexicons, research in emotion analysis has had to rely on limited and small emotion lexicons. In this paper we show how the combined strength and wisdom of the crowds can be used to generate a large, high-quality, word-emotion and word-polarity association lexicon quickly and inexpensively. We enumerate the challenges in emotion annotation in a crowdsourcing scenario and propose solutions to address them. Most notably, in addition to questions about emotions associated with terms, we show how the inclusion of a word choice question can discourage malicious data entry, help identify instances where the annotator may not be familiar with the target term (allowing us to reject such annotations), and help obtain annotations at sense level (rather than at word level). We conducted experiments on how to formulate the emotion-annotation questions, and show that asking if a term is associated with an emotion leads to markedly higher inter-annotator agreement than that obtained by asking if a term evokes an emotion.

1. INTRODUCTION

Emotion analysis seeks to identify opinions, feelings, and other private states in large volumes of text, but emotion lexicons remain limited because traditional annotation is costly. The paper uses crowdsourcing to build a substantially larger lexicon while addressing annotation-quality challenges.

  • Emotion analysis examines opinions and private states expressed toward target entities such as companies, products, policies, people, and countries.
  • Limited emotion resources reflect the high cost and manual effort of using hand-picked expert annotators.
  • Mechanical Turk enables Internet-based annotation, but tasks require checks to discourage, reject, and re-annotate random or erroneous responses.
  • The paper compiles EmoLex, an English term–emotion lexicon created through manual annotation on Amazon’s Mechanical Turk.
  • The study investigates question wording, annotation difficulty, part-of-speech patterns, annotator agreement, and relationships between emotion and polarity.
  • EmoLex contains close to 10,000 terms and targets eight basic emotions, including frequent nouns, verbs, adjectives, adverbs, and bigrams.

2. APPLICATIONS

Emotion analysis supports applications ranging from customer-relations management to search, dialogue, tutoring, and social or literary analysis. The paper illustrates these uses through emotion-aware customer-relations systems.

  • Emotion recognition can support customer-relations decisions based on states such as dissatisfaction, satisfaction, sadness, trust, anticipation, and anger.
  • Emotion analysis can track attitudes toward politicians, movies, products, countries, and other target entities.
  • Emotion-aware search can distinguish emotions associated with products, while dialogue systems can respond to users’ different emotional states.
  • Intelligent tutoring systems can manage learners’ emotional states, with some support for better and faster learning in positive emotional states.
  • Other applications include analyzing communication, assisting emotional writing, depicting emotions in novels, and interpreting headline-evoked emotions.
  • Customer-relations systems use emotional information alongside interactions with customers, prospects, and business partners to support satisfaction and corrective action.

3. EMOTIONS

The paper adopts Plutchik’s model of eight basic emotions as its annotation framework. This choice makes annotation more manageable while acknowledging that alternative emotion taxonomies exist.

  • Emotion theories differ on which emotions are basic and on how instinctual and cognitive emotions relate.
  • Ekman proposes six basic emotions, whereas Plutchik proposes eight by adding trust and anticipation.
  • Plutchik organizes emotions in a wheel in which radius represents intensity and spatial relationships represent similarity or contrast.
  • The paper selects Plutchik’s eight emotions because annotating hundreds of emotions would be expensive and difficult for annotators.
  • The authors do not claim Plutchik’s categories are more fundamental than other categorizations, but cite their research basis, broader coverage, and future empirical verification.

4. RELATED WORK

Prior work has emphasized polarity and small emotion resources, while emotion analysis has used varied lexicon-construction and computational approaches. The paper situates its crowdsourced lexicon within this broader landscape.

  • Sentiment-analysis research has concentrated on positive and negative polarity, whereas comparatively less work has addressed emotion lexicons and emotional text analysis.
  • The WordNet Affect Lexicon contains a few hundred emotion-annotated words derived from seed words and their WordNet synonyms.
  • Plutchik’s wheel represents adjacent emotions as similar, opposite emotions as contrasting, radius as intensity, and intermediate spaces as primary dyads.
  • Automatic emotion-analysis systems identify emotion words, exploit co-occurrence with seed words, use hand-coded rules, or apply machine learning.
  • Recent work often studies Ekman’s six emotions, while other projects annotate complex emotions such as politeness, embarrassment, persuasion, and deception.
  • Emotion analysis has been applied across blogs, fairy tales, novels, chat, and other communication domains, with emotional expressiveness varying by domain and medium.

5. TARGET TERMS

The target terms were selected from the Macquarie Thesaurus and filtered for frequency, part of speech, and limited ambiguity to support word- and sense-level emotion annotation.

  • The Macquarie Thesaurus supplied unigrams and bigrams, including over 57,000 word types and more than 40,000 commonly used phrases.
  • The study selected the 200 most frequent terms for each of four parts of speech among unigrams and bigrams, except for 187 qualifying adverb bigrams.
  • Terms occurring in multiple Macquarie categories were excluded from these sets, treating categories as coarse senses.
  • The selection also included 640 Ekman-subset WordNet Affect Lexicon word–sense pairs with at most two thesaurus-category senses.

6. MECHANICAL TURK

The annotation workflow used Amazon Mechanical Turk to divide emotion-labeling work into independently solvable HITs completed by multiple Turkers.

  • Mechanical Turk requesters break annotation tasks into small, independently solvable units called HITs and upload them to the platform.
  • Requesters specify task keywords, compensation per HIT, and the number of annotators assigned to each HIT.
  • Responses submitted by Turkers are called assignments, and Turkers typically find tasks using keywords and minimum compensation preferences.
  • Each target term received annotations from five different Turkers, with no Turker allowed to attempt multiple assignments for the same term.

7. ISSUES WITH CROWDSOURCING AND EMOTION ANNOTATION

Crowdsourcing offers speed and low cost but creates quality-control, recruitment, language, sense-disambiguation, and task-design challenges for emotion annotation.

  • Quality control is the foremost crowdsourcing challenge because tasks may attract cheaters or malicious annotators who submit incorrect information.
  • Crowdsourcing participation depends on how interesting workers find the task and how attractive they find its compensation.
  • The annotation requires native or fluent English speakers but no special skills, because such speakers can identify emotions associated with words.
  • Different word senses can evoke different emotions, making sense-inventory choice and granularity central annotation challenges.
  • Long definitions increase reading time and reduce annotations obtainable within a budget, while annotators should label only words whose meanings they know.

8. OUR APPROACH

The approach combines a preliminary word-choice question with emotion, polarity, and related survey questions, using automated sense-based item generation and quality checks. A pilot comparison found higher agreement for asking whether terms are associated with emotions than whether they evoke them.

  • Word-choice screening: A four-option word-choice question precedes emotion questions to convey the intended sense without long definitions and to filter unfamiliar or random responses.The correct option is a synonym for one target sense; with three distractors, an unfamiliar or random annotator has a 75% chance of answering incorrectly.
  • Word-choice screening: The word-choice alternatives are generated from the Macquarie Thesaurus using the target category’s head word, three irrelevant distractors, and coarse senses.
  • Question formulation: A pilot on 2,100 terms compared questions asking whether a word is associated with an emotion against questions asking whether it evokes an emotion.
  • Survey implementation: The HITs used multiple-choice questions, included instructions for fluent English speakers and check-question rejection, and paid $0.04 per HIT.
  • Annotation questions: The survey asks about positive and negative polarity, eight emotions, and whether the target itself is an emotion.The eight emotions are joy, sadness, fear, anger, trust, disgust, surprise, and anticipation.

9. ANNOTATION STATISTICS AND POST-PROCESSING

The annotation project scaled from a pilot to a larger batch, then used validation and annotator-quality filtering to produce a master set of reliable assignments at modest cost.

  • Annotation scale and timing: The two annotation batches covered about 2,100 terms in one week and about 8,000 HITs in two weeks.The paper notes that annotation time was not linearly proportional to the number of HITs.
  • Post-processing and master set: About 2,666 of 50,850 assignments were discarded and rejected because they included at least one unanswered question.The assignments came from 10,170 terms with five assignments each.
  • Post-processing and master set: More than 95% of remaining assignments answered the word-choice question correctly, and assignments with incorrect answers were discarded.Annotators scoring below 66.67% on word-choice questions were treated as having attempted to violate the instructions.
  • Post-processing and master set: Annotations from 111 Turkers more than two standard deviations from the mean agreement probability were discarded as outliers.Agreement was measured by each annotator’s maximum-likelihood probability of agreeing with the majority on emotion questions.
  • Post-processing and master set: 8,883 of 10,170 terms remained after post-processing, each with at least three valid assignments.The master set contained 38,726 assignments from about 2,216 Turkers, averaging about 4.4 assignments per term.
  • Annotation scale and timing: The total annotation cost was about US$2,100, including Amazon fees and dual annotation of the pilot set.

10. ANALYSIS OF EMOTION ANNOTATIONS

The analysis characterizes emotion associations in EmoLex, including intensity distributions, polarity links, cross-lexicon comparisons, and annotator agreement. It also examines how annotation design and category boundaries shape reliability.

  • Emotion prevalence: 5% of the 8,883 target terms strongly evoke joy, while 22.5% strongly evoke at least one of the eight basic emotions.Strong emotion was determined from the strongest emotion expressed by each target.
  • Emotion prevalence: 9.3% of the 8,883 target terms refer directly to emotions rather than merely being associated with them.This distinction came from the Q12 analysis.
  • Parts of speech: Adjectives (68%) and adverbs (67%) are most often emotive; trust and joy are the most common associated emotions at 16% each.Nouns are most associated with trust (16%), whereas adjectives are most associated with joy (29%).
  • Polarity and lexicon comparisons: Negative General Inquirer words are mostly associated with anger, fear, disgust, and sadness, whereas positive words are associated with anticipation, joy, and trust.EmoLex-WAL comparisons likewise show substantial alignment between source emotion categories and Turker associations.
  • Agreement: For almost 60% of terms, at least four annotators agree at four intensity levels; at two levels, more than 60% have unanimous agreement and almost 85% have at least four agreeing annotators.The two-level conversion groups no/weak assignments as non-emotive and moderate/strong assignments as emotive.
  • Agreement: Average Fleiss’s κ is 0.29, indicating fair agreement; six emotions have fair agreement, while anticipation and trust have slight agreement.Anger and sadness have the highest κ values among the eight emotions.
  • Agreement: Agreement is constrained by out-of-context words, graded associations near bin boundaries, and possible personal interpretations of evoked emotion.The authors speculate that “associated” prompts elicit more widely accepted judgments than “evoked” prompts, which may draw on personal experience.

11. ANALYSIS OF POLARITY ANNOTATIONS

The polarity annotations show substantial agreement, with negative polarity more consistently identified than positive polarity. The lexicon also reveals systematic polarity patterns across term categories and existing resources.

  • Annotation consolidation: Four-level polarity annotations were consolidated into majority classes of no, weak, moderate, and strong polarity, and also converted into evaluative versus non-evaluative bins.
  • 30.1% of target terms were strongly positive or strongly negative.
  • Comparison with existing resources: Turkers generally reproduced General Inquirer polarity labels, while 12% of GI-neutral terms were marked negative and 30% positive.
  • Emotion–polarity patterns: Anger, disgust, fear, and sadness terms tended to be negative, joy terms were positive, and surprise terms were more than twice as likely to be positive than negative.
  • Agreement: More than 55% of terms had unanimous agreement at two polarity levels, and more than 80% had at least four annotators agreeing.
  • Agreement: Negative polarity annotations had markedly higher agreement than positive polarity annotations.The authors associate this difference with a fuzzier boundary between positive and neutral than between negative and neutral.

12. CONCLUSIONS

The paper presents EmoLex as a large, crowdsourced word–emotion resource and describes quality-control mechanisms for producing sense-level annotations. It reports broad coverage, high annotation quality against gold data, and polarity labels for all terms.

  • Contributions: EmoLex contains more than 10,000 word–sense pairs, each annotated for association with eight basic emotions.The lexicon was created using Amazon’s Mechanical Turk.
  • Crowdsourcing quality control: Word choice questions detect erroneous or malicious annotations, reject unqualified Turkers, and convey the intended word sense.
  • Validation: The authors compared a lexicon subset with existing gold-standard data and report that the annotations were high quality.
  • Additional annotations: All 10,170 lexicon terms were annotated with positive, negative, or neutral semantic orientation.
  • Additional analyses: The lexicon also identifies 826 terms that directly refer to emotions and examines simultaneous emotion associations and associations among high-frequency words.

13. FUTURE DIRECTIONS

Future work targets broader coverage, improved annotation efficiency, multilingual and cross-cultural resources, and better use of word-sense and contextual information. The authors also identify unresolved challenges in interpreting emotions in text.

  • Coverage and multilingual expansion: Future work includes expanding lexicon coverage, creating resources in other languages, and studying cross-cultural and cross-language differences.
  • Future analyses: The authors plan to study emotion variation among near-synonyms and across different senses of polysemous words.
  • Annotation improvement: MaxDiff would ask annotators to identify the most- and least-associated items among groups, producing more efficient comparative judgments and potentially higher agreement.
  • Sense and context: Because annotations are at word-sense level, accurate word-sense disambiguation is needed to use them fully.
  • Interpretation challenges: Emotion analysis must determine who experiences an emotion, what evokes it, and how multiple entities in a passage may have different emotions.
  • Interpretation challenges: Negation remains challenging because not sad does not usually mean happy, while not happy can often mean sad.
Loading 1308.6297v1…