Source-linked AI summary
Yor-Sarc: A gold-standard dataset for sarcasm detection in a low-resource African language
Toheeb Aduramomi Jimoh, Tabea De Wille, Nikola S. Nikolov
TL;DR
Sarcasm detection lacks annotated resources in low-resource languages, particularly African languages, despite requiring pragmatic and cultural interpretation. The paper introduces Yor-Sarc, a 436-instance Yorùbá dataset annotated by three native speakers with culturally grounded guidelines and agreement analysis. It reports substantial to almost perfect agreement, preserves majority-agreement cases as soft labels, and identifies limited discourse-domain coverage as its main scope boundary.
Problem
Sarcasm detection requires pragmatic and cultural interpretation, but low-resource African languages lack high-quality manually annotated datasets.
Method
The paper constructs a 436-instance Yorùbá corpus with culturally grounded three-speaker annotation and complementary inter-annotator agreement analyses.
Results
Fleiss’ κ = 0.7660, with 83.3% unanimous consensus and agreement ranging from substantial to almost perfect across reported measures.
Takeaways & Limitations
The agreement analysis supports Yor-Sarc as a reliable resource for computational sarcasm research and preserves difficult cases as a training signal through soft labels.
Takeaways & Limitations
The dataset covers social media and news but does not exhaustively represent Yorùbá discourse contexts such as face-to-face conversation or dialogue.
Abstract
from arXiv · showhide
Sarcasm detection poses a fundamental challenge in computational semantics, requiring models to resolve disparities between literal and intended meaning. The challenge is amplified in low-resource languages where annotated datasets are scarce or nonexistent. We present \textbf{Yor-Sarc}, the first gold-standard dataset for sarcasm detection in Yorùbá, a tonal Niger-Congo language spoken by over $50$ million people. The dataset comprises 436 instances annotated by three native speakers from diverse dialectal backgrounds using an annotation protocol specifically designed for Yorùbá sarcasm by taking culture into account. This protocol incorporates context-sensitive interpretation and community-informed guidelines and is accompanied by a comprehensive analysis of inter-annotator agreement to support replication in other African languages. Substantial to almost perfect agreement was achieved (Fleiss' $κ= 0.7660$; pairwise Cohen's $κ= 0.6732$--$0.8743$), with $83.3\%$ unanimous consensus. One annotator pair achieved almost perfect agreement ($κ= 0.8743$; $93.8\%$ raw agreement), exceeding a number of reported benchmarks for English sarcasm research works. The remaining $16.7\%$ majority-agreement cases are preserved as soft labels for uncertainty-aware modelling. Yor-Sarc\footnote{https://github.com/toheebadura/yor-sarc} is expected to facilitate research on semantic interpretation and culturally informed NLP for low-resource African languages.
1 Introduction
Sarcasm detection is difficult because intended meaning depends on pragmatic, literary, and cultural cues, while low-resource African languages lack high-quality annotated datasets. Yor-Sarc addresses this gap with a culturally grounded, multi-annotator gold-standard corpus for Yorùbá.
- Sarcasm detection requires resolving discrepancies between literal and intended meaning using subtle pragmatic and cultural cues.It is also frequently confused with related figurative phenomena such as irony and metaphor.
- Low-resource language sarcasm research remains limited, especially in Africa, where high-quality manually annotated datasets are scarce.Existing African NLP resources have made progress on sentiment and other classification tasks, but the sarcasm resource gap remains substantial.
- Yor-Sarc introduces the first, to the best of the authors’ knowledge, manually annotated sarcasm dataset for Yorùbá.The corpus contains 436 Standard Yorùbá texts from social media, news, video captions, and crowdsourced examples, annotated independently by three native speakers.
- The work contributes a publicly available gold-standard corpus, a culturally grounded three-annotator protocol, agreement analysis, and corpus-level source and class analysis.The protocol is intended to inform similar annotation efforts in other African languages.
- The paper proceeds from related work to dataset and annotation procedures, agreement analysis, and conclusions covering limitations and future directions.
2 Related works
Sarcasm research has progressed from English rule-based classification to neural, transformer, dialogue, and multimodal approaches, but African NLP resources remain concentrated on foundational tasks and sentiment.
- Early sarcasm detection treated English sentences as classification problems using lexical, syntactic, and sentiment-based incongruity features.The field later moved toward deep learning and transformer-based methods for semantic and pragmatic cues.
- SARC and MUStARD enabled sarcasm research in Reddit and multimodal dialogue settings, while newer work incorporates visual and acoustic signals.
- African NLP has primarily developed resources for language identification, part-of-speech tagging, named entity recognition, and sentiment analysis.
- AfriSenti provides over 110,000 annotated tweets across 14 African languages, including Yorùbá, but targets coarse-grained sentiment rather than sarcasm.
3 Dataset, Multi-Annotator Framework and Inter-Annotator Agreement
Yor-Sarc comprises 436 Yorùbá instances annotated independently by three native speakers under structured guidelines, with agreement, uncertainty, bias, and disagreement analyzed through complementary measures.
- Dataset: The dataset contains 436 instances from six sources, led by BBC News Yorùbá at 65.4% (n = 285) and social media at 28.5%.Instagram, X/Twitter, Facebook, YouTube, and crowdsourced contributions provide additional informal or gap-filling data.
- Multi-Annotator Framework: Three native Yorùbá speakers independently labeled all 436 instances using guidelines refined through a 20-example pilot and subsequent discussion.The protocol covers diverse dialectal backgrounds and binary sarcastic versus non-sarcastic judgments without consultation.
- Pairwise Agreement: Cohen’s kappa measures chance-corrected agreement for each annotator pair, with observed agreement Po and chance-expected agreement Pe.The analysis reports all three pairwise values, their average, and 95% bootstrap confidence intervals.
- Multi-Rater Agreement: Fleiss’ kappa measures overall chance-corrected agreement across all three annotators using per-instance counts, mean observed agreement, and marginal expected agreement.The framework uses r = 3 raters and C = 2 categories.
- Agreement and Uncertainty: Instances are categorized by unanimity, majority agreement, or disagreement to quantify agreement patterns and derive soft labels.Soft-label derivation preserves uncertainty information rather than retaining only majority-vote outcomes.
- Annotator Bias: Annotator bias is measured as deviation from consensus, with positive values indicating more sarcastic labeling and negative values indicating more conservative labeling.Cross-annotator consistency is summarized by the standard deviation σ_bias.
- Uncertainty: Binary entropy identifies uncertainty, with H(k) = 0 for unanimous cases and H(k) ≈ 0.92 for 2-1 splits.Instances with H(k) > 0.5 are treated as hard cases for evaluation.
- Disagreement Analysis: Row-normalized confusion matrices compare conditional label assignments between annotator pairs, distinguishing symmetric disagreement from systematic bias.Balanced off-diagonals indicate symmetry, whereas imbalanced off-diagonals indicate bias.
4 Inter-Annotator Agreement Analysis
Yor-Sarc annotations show substantial overall reliability, near-unanimous instance-level consensus, and preserved disagreement that captures uncertainty in sarcasm judgments. Agreement variation is linked to annotators’ differing interpretive thresholds, while the reported metrics exceed several published benchmarks.
- 4.1 Overall Agreement Metrics: Fleiss’ κ = 0.7660 indicates substantial overall agreement, while pairwise Cohen’s κ ranges from 0.6732 to 0.8743.The highest pairwise value reaches the almost-perfect agreement range.
- 4.1 Overall Agreement Metrics: κ = 0.8743 and 93.81% raw agreement across 436 instances make A1–A2 the strongest annotator pair.This pairwise result exceeds reported English and code-mixed sarcasm benchmarks.
- 4.2 Agreement Distribution and Consensus Patterns: 83.26% of instances received unanimous agreement, while 16.74% showed two-against-one majority patterns.The instance-level analysis distinguishes clear consensus from interpretive variation.
- 4.3 Soft Labels and Uncertainty Preservation: Soft labels preserve disagreement as uncertainty: 83.9% of instances lie at extreme values, while 16.1% receive intermediate values of 0.333 or 0.667.The approach supports uncertainty-aware model training rather than eliminating disagreement through adjudication.
- 4.4 Annotator Behavior and Pairwise Patterns: Annotators differed in sarcastic-label rates, from A3’s 30.96% to A2’s 45.87%, reflecting distinct interpretive thresholds rather than inconsistent guideline application.A1–A2’s threshold difference was 4.81 percentage points, compared with 14.91 between A2 and A3.
- 4.4 Annotator Behavior and Pairwise Patterns: A2–A3 errors were asymmetric, with A2 marking 42 instances sarcastic that A3 did not, versus 27 in the reverse direction.The asymmetry is consistent with A2’s more liberal threshold and larger sarcastic-label rate.
- 4.5 Benchmark Comparison and Quality Assessment: Yor-Sarc’s average pairwise κ of 0.7671 exceeds prior English sarcasm annotation results, including κ = 0.56–0.62 and κ = 0.67.The A1–A2 pair also exceeds the comparable κ = 0.81 value reported for a high-confidence subset.
5 Conclusion and Future Directions
Yor-Sarc establishes a gold-standard Yorùbá sarcasm resource and shows that native speakers can reliably annotate this culturally and pragmatically complex phenomenon. Its agreement analysis supports preserving disagreement as soft labels for future uncertainty-aware modelling.
- Yor-Sarc is a gold-standard manually annotated sarcasm dataset intended to advance sarcasm detection in low-resource African languages.
- Fleiss’ κ = 0.7660 shows substantial agreement, while native speakers with shared cultural backgrounds reliably distinguish subtle intended meanings in a semantically complex task.
- 83.3% unanimity indicates that native speakers reliably recognize markers of Yorùbá sarcasm, including lexical hyperbole.
- 16.7% of majority-agreement cases reflect differing evidence thresholds for identifying sarcasm when markers are present.
- Soft labels preserve the scalar nature of sarcasm instead of imposing an artificial binary categorization.
- The study treats disagreement as a training signal, noting that future work will expand Yor-Sarc and conduct advanced sarcasm detection experiments.
Limitation
The dataset covers social media and news media but does not exhaustively represent all Yorùbá discourse contexts. Expanding domain coverage while maintaining agreement quality is identified as a worthwhile direction.
- The dataset’s only observed limitation is incomplete domain coverage across Yorùbá discourse contexts.
- Its instances span social media and news media but exclude contexts such as face-to-face conversation and dialogue.
- Future expansion should broaden domain coverage while maintaining annotation agreement quality.
Ethics statement
The study reports compliance with ethical guidelines for human-subjects research. Public data were obtained under applicable permissions, and crowdsourced contributors gave informed consent.
- Public BBC News Yorùbá and social-media data were collected from sources permitting public distribution and were handled in compliance with platform terms of service.
- Crowdsourced participants gave informed consent through an ethically approved online survey and explicitly permitted research use of their examples.
A.1 Label Distribution Across Annotators
Annotators’ sarcastic-label rates range from 30.96% to 45.87%, indicating threshold differences while retaining balanced class distributions. Figure 5 presents these label distributions for each annotator.
- 30.96% (A3) to 45.87% (A2) is the range of sarcastic-label rates across annotators.
- The differing rates reflect annotator threshold differences while maintaining balanced class distributions.
- Figure 5 shows the label distributions for each annotator.
A.2 Annotator Labelling Bias
Annotators showed different labelling thresholds, with A2 the most liberal and A3 the most conservative.
- 45.87% and 30.96%: A2 was the most liberal annotator, while A3 was the most conservative.
A.3 Agreement Level Distribution
Agreement was high overall, with 83.3% unanimity and no complete disagreements; sarcastic instances had slightly higher consensus than non-sarcastic ones.
- 83.3% unanimity was observed, with zero complete disagreements.
- 85.5% versus 81.5%: sarcastic instances had slightly higher consensus than non-sarcastic instances.