Source-linked AI summary
Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, Noah A. Smith
TL;DR
Toxicity datasets often average away disagreement linked to annotator identities and beliefs, even though language such as AAE can be differentially perceived. Using two online studies of diverse participants and three text characteristics, the paper finds systematic associations between annotator variables and toxicity ratings, and shows that a popular detector reflects some perspectives more than others.
Problem
Toxicity annotation often ignores variation associated with annotator identities and beliefs, despite evidence that AAE can be falsely flagged as toxic.
Method
Two online studies examine diverse annotators’ ratings across anti-Black language, AAE, and vulgarity using controlled and broader post samples.
Results
The studies find associations between annotator identities or beliefs and toxicity ratings, including lower anti-Black toxicity ratings among higher racist-beliefs scorers and higher AAE ratings among more conservative annotators.
Takeaways & Limitations
The findings support contextualizing toxicity labels in social variables and looking beyond aggregated discrete decisions.
Takeaways & Limitations
The study is U.S.-centric, its attitude and scale choices may affect results, and its automatic AAE detector may have induced data-selection bias.
Abstract
from arXiv · showhide
The perceived toxicity of language can vary based on someone's identity and beliefs, but this variation is often ignored when collecting toxic language datasets, resulting in dataset and model biases. We seek to understand the who, why, and what behind biases in toxicity annotations. In two online studies with demographically and politically diverse participants, we investigate the effect of annotator identities (who) and beliefs (why), drawing from social psychology research about hate speech, free speech, racist beliefs, political leaning, and more. We disentangle what is annotated as toxic by considering posts with three characteristics: anti-Black language, African American English (AAE) dialect, and vulgarity. Our results show strong associations between annotator identity and beliefs and their ratings of toxicity. Notably, more conservative annotators and those who scored highly on our scale for racist beliefs were less likely to rate anti-Black language as toxic, but more likely to rate AAE as toxic. We additionally present a case study illustrating how a popular toxicity detection system's ratings inherently reflect only specific beliefs and perspectives. Our findings call for contextualizing toxicity labels in social variables, which raises immense implications for toxic language annotation and detection.
1 Introduction
Toxicity judgments are subjective and vary with annotator identities, beliefs, and the linguistic characteristics being rated. Two diverse online studies examine these influences across anti-Black content, AAE, and vulgar language, finding systematic associations and perspective-specific system ratings.
- Motivation: Toxicity detection requires nuanced interpretation, yet AAE has been falsely flagged as toxic in prior systems.The paper frames this as a consequence of overlooking pragmatic and racialized variation in language.
- Motivation: Most datasets reduce variable human judgments to one averaged label, ignoring differences associated with annotator identities and beliefs.This treats toxicity as a simple classification despite documented annotation variance.
- Approach: The authors study who, why, and what shape toxicity annotations through two online studies with demographically and politically diverse participants.They vary annotator identities and attitudes while examining anti-Black meaning, AAE, and vulgar words.
- Study overview: 641 annotators rated 15 hand-curated posts in the breadth-of-workers study, while 173 annotators rated approximately 600 posts in the breadth-of-posts study.The two designs provide controlled breadth of workers and broader breadth of posts.
- Key findings: Higher racist-beliefs scores predicted lower toxicity ratings for anti-Black content, while conservatism predicted higher toxicity ratings for AAE and conservative or traditionalist attitudes predicted higher ratings for vulgar language.These are the paper’s most salient associations across the three language characteristics.
- Implications: PERSPECTIVEAPI ratings aligned more closely with annotators holding certain attitudes and identities, rather than representing all perspectives equally.For anti-Black language, its scores better reflected annotators scoring high on racist beliefs.
2 The Who, Why, and What of Toxicity Annotations
The paper separates toxicity annotation into who annotates, why they judge language as toxic, and what textual characteristics are present. It operationalizes these dimensions through demographic variables, attitude scales, and categories including anti-Black language, AAE, and vulgarity.
- The who, why, and what: The study asks how annotator identities, beliefs, and text characteristics shape perceptions of toxicity.Its identity variables include race, gender, and political leaning, while text categories include anti-Black language, AAE, and vulgarity.
- The who, why, and what: The authors restrict participants to the United States because perceptions of race and political attitudes vary across global contexts.This establishes a U.S.-centric scope for the study.
- Attitudes: Seven attitude dimensions are operationalized with scales drawn from political science, social psychology, and sociolinguistics.The dimensions include free speech, harm of hate speech, racist beliefs, traditionalism, language purism, empathy, and altruism.
- Attitudes: Language purism measures the belief that English has a correct form and typically involves negative reactions to non-canonical language use.The authors created and validated a four-item language-purism scale.
- Text characteristics: The selected text dimensions are anti-Black language, AAE dialect markers, and vulgar language, with vulgarity distinguished by whether it refers to identity.These categories are chosen because prior work found them prone to over- or under-detection as toxic.
3 Data & Study Design
The paper uses two complementary online annotation studies: one holds posts controlled while broadening annotator diversity, and the other broadens posts while using fewer annotators per post. The designs isolate text characteristics and associate toxicity ratings with annotator variables.
- Design: Both studies ask participants to rate how offensive and racist each post is rather than imposing prescriptive toxicity definitions.This choice reflects substantial disagreement among annotators in prior work.
- Breadth-of-workers study: The breadth-of-workers study uses 15 hand-curated posts belonging exclusively to one text category.Vulgar posts use non-identity-referring terms to avoid confounding vulgarity with offensive identity mentions.
- Breadth-of-workers study: 641 participants with varied racial, political, and gender identities rated every one of the 15 posts.The pool was recruited through an MTurk pre-qualifier survey designed to ensure racial and political diversity.
- Analysis: Associations in the workers study use Pearson r or Cohen’s d between category-average toxicity ratings and annotator identities or attitude scores.The posts study instead uses a linear mixed-effects model because annotators rated varying numbers of posts.
- Breadth-of-posts study: The breadth-of-posts study samples 571 posts and allows anti-Black or AAE posts to also contain identity-referring or non-identity-referring vulgarity.Posts that were simultaneously anti-Black and AAE were excluded because their pragmatic toxicity implications were considered complex.
- Breadth-of-posts study: Each post in the breadth-of-posts study was annotated by six participants from a diverse pool of 173 annotators.Participants also answered one-item versions of the attitude scales.
4 Who finds anti-Black posts toxic, and why?
Anti-Black posts were rated differently according to annotators’ beliefs, identities, and political attitudes. Higher racist-belief and conservative scores were associated with lower toxicity ratings, while harmony-focused hate-speech attitudes and some identity variables showed opposite or additional associations.
- 4.1 Breadth-of-Workers Results: Higher RACISTBELIEFS, FREEOFFSPEECH, and conservative scores were associated with lower toxicity ratings for anti-Black posts.
- 4.1 Breadth-of-Workers Results: Higher HARMOFHATESPEECH scores were associated with rating anti-Black posts as more offensive and more racist.
- 4.1 Breadth-of-Workers Results: Black and white annotators both rated anti-Black posts highly offensive, with means of 3.85 and 3.59 out of 5, respectively.
- 4.1 Breadth-of-Workers Results: Exploratory analyses linked lower ratings to LINGPURISM, TRADITIONALISM, and male gender, while higher ratings were associated with EMPATHY, ALTRUISM, and female gender.
- 4.2 Breadth-of-Posts Results: The breadth-of-posts study reproduced lower anti-Black offensiveness ratings among higher-RACISTBELIEFS annotators and higher ratings among higher-HARMOFHATESPEECH annotators.
5 Who finds AAE posts toxic, and why?
AAE posts were evaluated through racial, political, and attitudinal lenses, with several associations differing from expectations. Conservative and racist-belief measures were linked to higher racism or offensiveness ratings in the broader post study, while identity differences were limited.
- 5 Who finds AAE posts toxic, and why?: AAE comprises U.S. English varieties common among, but not limited to, African-American or Black speakers.
- 5.1 Breadth-of-Workers Results: The breadth-of-workers study found conservative leaning and RACISTBELIEFS associated with racism ratings for AAE posts, but white and Black annotators did not differ significantly in offensiveness ratings (d = 0.14, p > 0.1).
- 5.2 Breadth-of-Posts Results: In the breadth-of-posts study, conservative annotators and those scoring higher in TRADITIONALISM gave AAE posts higher offensiveness ratings.
- 5.2 Breadth-of-Posts Results: Conservative annotators and those scoring higher in FREEOFFSPEECH rated AAE posts as more racist, with TRADITIONALISM showing a near-significant association.
- 5.3 Perceived Toxicity of AAE: For AAE posts containing “n*gga,” conservative raters assigned higher racism scores (β = 0.465, p = 0.003; corrected for multiple comparisons).
- 5.3 Perceived Toxicity of AAE: The authors suggest that AAE markers may be perceived as obscene or racist, contributing to over-estimation of AAE as toxic and potential racio-linguistic representational harms.
6 Who finds vulgar posts toxic, and why?
Vulgarity was treated as a distinct form of offensiveness rather than uniformly toxic language. Traditionalism, language purism, and conservative political leaning predicted higher offensiveness ratings for exclusively vulgar posts.
- 6 Who finds vulgar posts toxic, and why?: Vulgarity includes non-identity swearwords and identity-referring slurs, both of which can have non-hateful uses.
- 6 Who finds vulgar posts toxic, and why?: The study focused on exclusively vulgar, non-identity-referring posts to avoid confounding vulgarity with AAE or anti-Black characteristics.
- 6.2 Perceived Toxicity of Vulgar Language: Offensiveness ratings for vulgar posts correlated with TRADITIONALISM, LINGPURISM, and conservative political leaning.
- 6.2 Perceived Toxicity of Vulgar Language: No associations between annotator attitudes and racism ratings were found for vulgar posts.
- 6.2 Perceived Toxicity of Vulgar Language: The findings characterize vulgarity as a specific form of offensiveness whose interpretation may vary across traditional, generational, and cultural norms.
7 Toxicity Detection System Case Study: PERSPECTIVEAPI
The case study tests whether PERSPECTIVEAPI toxicity scores align equally with annotators across identities and attitudes. Its associations indicate that the system reflects particular viewpoints depending on text category.
- Method: PERSPECTIVEAPI scores were compared with ratings from annotators grouped by demographic identities and attitude-scale scores across 571 posts.The analysis used posts from the breadth-of-posts study and split attitude groups at their mean.
- Interpretation: Overall, PERSPECTIVEAPI scores aligned with specific viewpoints or ideologies rather than representing annotator perceptions uniformly.The alignment varied with the text category, including lower apparent sensitivity to anti-Black toxicity and stronger alignment with white perceptions of AAE toxicity.
8 Discussion & Conclusion
The discussion argues that toxicity perceptions vary with annotator identities and beliefs, while existing annotation and classification practices often compress this variation. It recommends documenting social variables and developing more nuanced, explainable systems, while noting important scope and design limitations.
- Discussion & Conclusion: Analyses across two studies found associations between annotator identities and beliefs and toxicity perceptions for anti-Black, AAE, and vulgar text.The paper also found that a popular toxicity detection system aligned more with some attitudes and identities than others.
- Recommendations: Researchers and dataset creators are recommended to investigate and report annotator attitudes and demographics alongside toxicity datasets.The paper suggests collecting relevant attitude scores and reporting them in dataset documentation such as datasheets.
- Recommendations: The paper asks whose perspective should guide toxicity systems and urges consideration of stakeholders and end users through human-centered design.It frames toxicity detection as subjective and socially consequential.
- Beyond classification: Majority-vote toxicity labels may be inadequate; alternatives include modeling label variation by annotator identities or beliefs and generating explanations of biased implications.The paper advocates moving beyond opaque classification toward nuanced, holistic, and explainable frameworks for assisting moderators.
- Limitations and open questions: The study’s attitude-scale choices, automatic AAE detector, and pilot analysis of PERSPECTIVEAPI constrain the conclusions and motivate deeper study.The authors also identify broader axes of discrimination and non-U.S. cultural contexts as open directions.
- Measurement: The attitude dimensions were measured with scales including free speech, harm of hate speech, racist beliefs, traditionalism, lingpurism, empathy, and altruism.The breadth-of-workers study used full scales, while the larger study used selected one-item versions; responses used 5-point Likert scales.
- Participant attitudes: RACISTBELIEFS scores were skewed toward low values, while liberal leaning correlated with higher EMPATHY and HARMOFHATESPEECH and conservative leaning with higher TRADITIONALISM, LINGPURISM, FREEOFFSPEECH, and RACISTBELIEFS.These distributions and inter-variable associations characterize the participant sample used in the analyses.
C.1 Data Selection & Validation
The studies selected and validated posts designed to isolate anti-Black meaning, AAE dialect, and vulgarity, then recruited diverse MTurk annotators for controlled and large-scale rating tasks.
- Data Selection & Validation: Candidate posts were selected to represent vulgarity, AAE, or anti-Black meaning while minimizing overlap with the other characteristics.Vulgar and AAE candidates came from an existing hate-speech corpus, while anti-Black candidates were curated separately.
- Data Selection & Validation: Three trained undergraduate assistants validated candidates using binary judgments of vulgarity and offensiveness to minorities, after which five posts per category were selected.The validation procedure was intended to ensure that posts were indicative of their target categories.
- Breadth-of-posts study: The breadth-of-posts study sampled up to 600 posts with stratification by toxicity, vulgarity, AAE, and anti-Black meaning, yielding 571 posts.Each post received ratings from white conservative, white liberal, and Black workers.
- Validation and comparisons: Figure 4 reports average offensiveness and racism ratings by tweet category, with all displayed differences significant after multiple-comparison correction.The caption specifies significance at p < 0.001.
- Breadth-of-posts study: The final breadth-of-posts dataset contained 3,171 ratings from 173 participants spanning racial, gender, and political groups.Participants were 76% white, 20% Black, 54% liberal, and 30% conservative; one-item attitude measures were used to reduce burden.
E Further Breadth-of-Workers Results
Additional analyses compare toxicity perceptions across anti-Black, AAE, and vulgar posts and report how these ratings differ by annotator characteristics. Anti-Black posts received the strongest overall offensiveness and racism ratings, while AAE and vulgarity showed smaller, differentiated effects.
- Further Breadth-of-Workers Results: The study reports the full set of attitude–toxicity associations in Table 10 using Pearson r or Cohen’s d with significance levels.Exploratory variable relationships were corrected for multiple comparisons.
- Category comparisons: Anti-Black tweets were rated substantially more offensive and racist than AAE or vulgar tweets, with effect sizes ranging from d = 2.4 to d = 3.6.Vulgar tweets were also rated more offensive than AAE tweets, with d = -0.29 and p < 0.001.
- Category comparisons: AAE tweets were rated slightly more racist than vulgar tweets, with d = 0.19 and p < 0.001.This difference was further examined by annotator gender, race, and political leaning.
- Annotator-group comparisons: AAE tweets exceeded vulgar tweets in racism ratings specifically among white or liberal annotators, with d = 0.20 and d = 0.22, respectively.The reported subgroup differences were statistically significant.
F Further Breadth-of-Posts Results
The study models associations between toxicity ratings and annotator identities or attitudes across overlapping tweet categories, using mixed-effects regression and group-based correlation comparisons.
- Analysis approach: A linear mixed-effects model accounts for varying numbers of posts rated by annotators and includes a random effect for each worker.The model regresses attitude scores onto toxicity scores.
- Tweet categories: The tweet categories in the breadth-of-posts analysis can overlap, as illustrated by the category-count Venn diagram.The analysis explicitly notes potentially overlapping categories such as AAE and vulgar posts.
- Analysis approach: The analysis obtains PERSPECTIVE toxicity scores for every post in the breadth-of-posts study.The API was accessed in October 2021.
- Annotator groups: Annotators are divided into high/low groups for attitudes and political leaning, while gender and race use man/woman and white/black bins.Attitude-group assignments use the mean score as the threshold.
- Reported associations: Table 11 reports only significant associations between demographic or attitude variables and offensiveness or racism ratings, with Holm correction for multiple comparisons.The table uses β coefficients from a mixed-effects model and distinguishes several significance thresholds.
G.2 Results
The results section presents comparisons between PERSPECTIVEAPI toxicity scores and annotator ratings across anti-Black, AAE, and vulgar tweet categories.
- Correlation comparisons: Table 12 evaluates whether correlations between PERSPECTIVEAPI scores and annotator ratings differ significantly between high- and low-attitude groups.The comparison uses Fisher’s z-to-r test with † and ∗ significance markers.
- Category comparisons: Figures 6–9 compare PERSPECTIVEAPI scores with offensiveness and racist ratings for anti-Black and AAE tweets.Separate figures cover offensiveness and racist ratings for each category.
- Category comparisons: Figures 10–13 extend the same PERSPECTIVEAPI comparisons to vulgar-OI and vulgar-ONI tweets.Each category has separate offensiveness and racist-rating figures.