Source-linked AI summary

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

Shamsuddeen Hassan Muhammad, Nedjma Ousidhoum, Idris Abdulmumin, Jan Philip Wahle, Terry Ruas, Meriem Beloucif, Christine de Kock, Nirmal Surange, Daniela Teodorescu, Ibrahim Said Ahmad, David Ifeoluwa Adelani, Alham Fikri Aji, Felermino D. M. A. Ali, Ilseyar Alimova, Vladimir Araujo, Nikolay Babakov, Naomi Baes, Ana-Maria Bucur, Andiswa Bukula, Guanqun Cao, Rodrigo Tufino Cardenas, Rendi Chevi, Chiamaka Ijeoma Chukwuneke, Alexandra Ciobotaru, Daryna Dementieva, Murja Sani Gadanya, Robert Geislinger, Bela Gipp, Oumaima Hourrane, Oana Ignat, Falalu Ibrahim Lawan, Rooweither Mabuya, Rahmad Mahendra, Vukosi Marivate, Alexander Panchenko, Andrew Piper, Charles Henrique Porto Ferreira, Vitaly Protasov, Samuel Rutunda, Manish Shrivastava, Aura Cristina Udrea, Lilian Diana Awuor Wanzare, Sophie Wu, Florian Valentin Wunderlich, Hanif Muhammad Zhafran, Tianhui Zhang, Yi Zhou, Saif M. Mohammad

arXiv:2502.11926v4cs.CL

TL;DR

Emotion-recognition research lacks high-quality multilingual resources, particularly for under-resourced languages. BRIGHTER addresses this gap with fluent-speaker, multi-label emotion datasets for 28 languages and reports broad evaluations across languages, domains, tasks, and prompting conditions. The results show persistent difficulty for LLMs, especially on under-resourced languages, alongside strong dependence on prompt wording, language, and few-shot examples.

  • Problem

    Emotion-recognition research has concentrated on high-resource languages because under-served languages often lack annotated datasets, leaving a major low-resource-language research gap.

  • Method

    BRIGHTER constructs multi-label emotion datasets for 28 languages from diverse domains, using fluent-speaker annotation and reporting monolingual and crosslingual emotion and intensity evaluations.

  • Results

    LLMs still struggle to predict perceived emotions and intensity, especially for under-resourced languages, and performance depends strongly on prompt wording, language, and few-shot shots.

  • Takeaways & Limitations

    BRIGHTER publicly releases datasets, annotation guidelines, and individual labels as a step toward expanding research on multilingual emotion recognition.

  • Takeaways & Limitations

    BRIGHTER does not claim to represent speakers’ true emotions, fully represent language use, or cover all possible emotions, and some low-resource languages have limited data sources.

Abstract

from arXiv · show

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition--an umbrella term for several NLP tasks--impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities in research efforts and proposed solutions, particularly for under-resourced languages, which often lack high-quality annotated datasets. In this paper, we present BRIGHTER--a collection of multi-labeled, emotion-annotated datasets in 28 different languages and across several domains. BRIGHTER primarily covers low-resource languages from Africa, Asia, Eastern Europe, and Latin America, with instances labeled by fluent speakers. We highlight the challenges related to the data collection and annotation processes, and then report experimental results for monolingual and crosslingual multi-label emotion identification, as well as emotion intensity recognition. We analyse the variability in performance across languages and text domains, both with and without the use of LLMs, and show that the BRIGHTER datasets represent a meaningful step towards addressing the gap in text-based emotion recognition.

1 Introduction

Emotion recognition addresses subtle, subjective language use and supports applications across NLP and related fields. BRIGHTER targets the resulting resource gap by providing manually annotated, multi-label emotion datasets for 28 languages, especially under-resourced ones.

  • Motivation: Emotion recognition concerns perceived emotions in text and supports applications including healthcare, dialogue systems, computational social science, and digital humanities.The paper uses the term for what most people think the speaker may have felt from a sentence or short text snippet.
  • Research gap: Most emotion-recognition research has focused on high-resource languages, partly because datasets for under-served languages are unavailable.This shortage is particularly pronounced for low-resource languages.
  • Contribution: BRIGHTER introduces manually annotated emotion datasets for 28 languages, containing nearly 100,000 instances from speeches, social media, news, literature, and reviews.The languages span seven language families and are predominantly low-resource, while also including mid- to high-resource languages such as English.
  • Contribution: Each BRIGHTER instance is multi-labeled across six emotion classes plus neutral and includes four intensity levels from 0 to 3.Instances are curated and annotated by fluent speakers.
  • Findings: The paper reports that LLMs perform significantly better with English prompts for low-resource languages and publicly releases the datasets.The authors present the release and local-community involvement as steps toward more inclusive digital tools.

2 The BRIGHTER Dataset Collection

BRIGHTER combines multilingual, domain-diverse data collection with fluent-speaker annotation, multi-label emotion judgments, intensity ratings, and reliability checks. The resulting datasets show substantial cross-language variation while generally achieving high annotation reliability.

  • Data collection: BRIGHTER uses data sources, collection strategies, and annotation procedures tailored to textual-data availability and access to fluent annotators across 28 languages.The collection combines multiple sources when resources are scarce and includes social media, narratives, speeches, literature, news, and generated text.
  • Data collection: The collection includes translated and manually corrected material, newly elicited sentences, and quality-approved ChatGPT-generated instances for Hindi and Marathi.A small Hindi section was automatically translated to Marathi and then corrected by native speakers.
  • Preprocessing: Before annotation, the authors remove duplicates, invisible characters, garbled encoding, incorrectly rendered emoticons, excessive expletives, and dehumanising language.All texts are anonymised.
  • Annotation: Annotators select all applicable perceived emotions from six categories, treat unselected instances as neutral, and rate selected emotions on a four-point intensity scale.The intensity scale runs from 0 for no emotion through 3 for high intensity.
  • Reliability: SHCMP estimates annotation reliability by repeatedly splitting annotations into n bins and averaging the proportion of instances receiving matching class labels.Overall SHCMP scores exceed 60% for n = 2, indicating reliable annotations.
  • Aggregation: Final emotion labels require at least two annotators to select an emotion and an average score above the threshold T = 0.5.Final intensity is the rounded-up average of selected scores and is assigned only where most instances have at least five annotators.
  • Data statistics: Emotion distributions vary substantially across languages because the datasets draw on different data sources, with insufficient representation excluding disgust in English and surprise in Afrikaans.Most languages include all six predefined emotion categories.

3 Experiments

BRIGHTER evaluates multi-label emotion classification and intensity prediction across 28 languages using multilingual models, large language models, and crosslingual transfer. Results show substantial variation by language, with weaker performance in low-resource settings and stronger transfer when languages are represented in pretraining.

  • Test sets range from approximately 1,000 to nearly 3,000 instances, while datasets without training data are reserved for crosslingual testing.
  • The experiments measure multilabel emotion classification with macro F1 and intensity prediction with Pearson correlation using MLMs and LLMs.
  • LLM emotion classification is challenging even for high-resource languages and performs worse for low-resource languages; Dolly-v2-12B is worst, while Qwen2.5-72B is best on average.
  • 27.44 is the maximum reported performance for yor, while hin, mar, and tat perform best overall, partly because their test data contain many single-labeled instances.
  • Crosslingual performance depends on transfer languages and model pretraining, with same-family training sometimes surpassing few-shot results and Niger-Congo languages benefiting least.
  • Multilingual models transfer more effectively to languages seen during pretraining, but often produce random or unreliable outputs for languages absent from training data.
  • DeepSeek-R1-70B outperforms other models in most intensity-prediction languages, including improvements exceeding 36 points for low-resource vernaculars such as arq.

4 Analysis

The analysis examines prompt wording, few-shot examples, pass@k generation, and prompt language as sources of LLM performance variation. Performance generally improves with more examples, larger k, and English prompts, although exceptions remain.

  • Prompt wording substantially affects LLM performance when equivalent English test texts are presented through different paraphrases.
  • Performance improves with more few-shot examples, then generally plateaus at 4 shots, suggesting that 4 to 8 shots may provide stable results.
  • Increasing k consistently improves pass@k performance, with DeepSeek-R1-70B exceeding F-score 90 at k = 8 while model rankings remain unchanged.
  • LLMs generally perform better with English prompts than target-language prompts, especially in low-resource languages, except arq with Qwen2.5-72B prompted in MSA.

5 Related Work

Related work has expanded from sentiment analysis toward specific emotion recognition, but multilingual resources remain limited in coverage and annotation. BRIGHTER addresses this gap with emotion-labeled data spanning 28 languages and combining simultaneous emotions with intensity.

  • Appraisal and constructed-emotion theories describe emotions as evaluations or conceptual constructs shaped by personal experience and the brain.
  • NLP research shifted from positive, negative, or neutral sentiment toward detecting specific emotions such as anger, fear, joy, and sadness.
  • Existing datasets cover several non-English languages, but multilingual resources remain limited and may rely heavily on translated data.
  • No multilingual resource previously captured simultaneous emotions and intensity across languages, motivating BRIGHTER’s 28-language collection.

6 Conclusion

BRIGHTER provides multi-labeled emotion-recognition datasets in 28 languages, with fluent-speaker annotation and intensity labels for 10 datasets. LLM evaluations show persistent difficulty with perceived emotions and intensity, especially in under-resourced languages, alongside sensitivity to prompting conditions.

  • BRIGHTER contains multi-labeled emotion-recognition datasets in 28 languages, collected and annotated by fluent speakers.
  • 10 datasets additionally include emotion-intensity annotations.
  • LLMs still struggle to predict perceived emotions and their intensity levels, especially for under-resourced languages.
  • LLM performance depends strongly on prompt wording, prompt language, and the number of few-shot examples.

Limitations

The authors limit BRIGHTER’s claims about representativeness and coverage, while noting that some low-resource languages have limited data sources. These constraints make the datasets unsuitable for data-intensive tasks but leave them useful as a starting point.

  • BRIGHTER does not claim to capture speakers’ true emotions, fully represent language use, or cover every possible emotion.
  • Limited data sources in some low-resource languages mean BRIGHTER cannot support tasks requiring large amounts of language-specific data.
  • Despite these constraints, the datasets remain a good starting point for research.

Ethical Considerations

The paper treats perceived emotion as distinct from actual emotion and acknowledges biases and incomplete coverage in the data. It also restricts high-risk uses and warns that systems built from the datasets may be unreliable under individual-instance and domain-shift conditions.

  • BRIGHTER focuses on perceived emotions because short text cannot establish with absolute certainty how someone feels.
  • The data may reflect biases from text-based communication and annotators’ internalised biases.
  • The datasets may not fully capture language usage, and some inappropriate instances may have been overlooked.
  • Commercial or state-actor use in high-risk applications is prohibited unless dataset creators explicitly approve it.
  • Systems developed using BRIGHTER may be unreliable for individual instances and sensitive to domain shifts.
  • Critical decisions about individuals, including health-related decisions, require appropriate expert oversight.
  • Annotators were compensated at rates exceeding the local minimum wage.

B Data sources

The annotation resources combine diverse language-specific sources with guidelines for multi-label emotion classification and four-level intensity ratings. A pilot across languages refined instructions for selecting applicable labels and handling complex emotions.

  • Data sources: The datasets draw on speeches, social media, news, literature, reviews, and newly created or translated text across multiple languages.Sources include Reddit, YouTube, Twitter, Weibo, BBC news, a translated Algerian novel, and manually drafted or ChatGPT-generated instances.
  • Annotation guidelines: Annotators classify texts into emotion categories for training models and studying how emotions are conveyed through language.The annotation guide emphasizes that emotions may be inferred even when not explicitly stated.
  • Annotation guidelines: Intensity is rated on four levels: 0 no emotion, 1 slight, 2 moderate, and 3 high emotion.The guide provides examples illustrating different intensity levels.
  • Annotation guidelines: A pilot annotation led to clarifications that annotators should select all applicable labels and may assign multiple labels to complex emotions.Bitterness or jealousy in Algerian Arabic could involve both anger and sadness.

D SCHMP Calculation

SHCMP measures annotation consistency by repeatedly splitting annotations, comparing item-level classes or score bins, and averaging the resulting match proportions.

  • D SCHMP Calculation: The annotated items are randomly divided into two equal subsets, with probabilistic tie-breaking for odd-sized datasets.The subsets are denoted A1 and A2.
  • D SCHMP Calculation: For each item, the method assigns scores from both subsets and derives classes C1(x_i) and C2(x_i).These class assignments provide the basis for comparing the two annotation subsets.
  • D SCHMP Calculation: Continuous scores in [−3, 3] are divided into equal-sized bins, with bin size b = 6/#Bins.Scores from A1 and A2 are assigned to bins c1 and c2.
  • D SCHMP Calculation: An item counts as consistent when its two bin indices differ by less than 1, meaning they are in the same or adjacent bins.The match indicator is M(x_i) = 1 if |c1 − c2| < 1 and 0 otherwise.
  • D SCHMP Calculation: SHCMP is the percentage of matched items, and the process is repeated across k random splits before averaging.The total matches are computed across all items before converting the proportion to a percentage.

E Experimental Settings

The experiments compare LLM and MLM settings for emotion classification, using deterministic LLM decoding and standardized training and evaluation procedures. The reported tables and prompt examples cover monolingual performance, prompt variants, and anger recognition tracks.

  • E Experimental Settings: LLMs use default HuggingFace parameters except temperature 0 and top-k 1 for deterministic output.Top-k ablations in Figure 5c use temperature 0.7.
  • E Experimental Settings: LLMs and MLMs are trained for 2 epochs with learning rate 1e-5 and evaluated on the test set.The same training duration and learning rate are stated for both model types.
  • E Experimental Settings: Table 5 reports Average F1-Macro for monolingual multi-label emotion classification within each language.The table highlights the best results.
  • E Experimental Settings: The monolingual ablation study compares prompt variants listed in Table 6.The prompt examples are organized around anger assessment.
  • E Experimental Settings: Track A asks whether anger is present, whereas Track B asks for an anger level from 0 to 3.Figures 7 and 8 show few-shot prompt templates for the two tracks.
Loading 2502.11926v4…