Source-linked AI summary

Language (Technology) is Power: A Critical Survey of "Bias" in NLP

Su Lin Blodgett, Solon Barocas, Hal Daumé, Hanna Wallach

arXiv:2005.14050v2cs.CLcs.CY

TL;DR

This paper surveys 146 papers on “bias” in NLP systems to examine their motivations and quantitative techniques. It finds vague, inconsistent motivations and techniques poorly matched to them, then proposes recommendations grounded in language’s relationship to social hierarchies and affected communities.

  • Problem

    Analyzing “bias” is inherently normative, but many NLP papers do not explain which system behaviors are harmful, how, to whom, or why.

  • Method

    The authors survey 146 written-text NLP papers, categorizing their motivations and quantitative techniques using an adapted taxonomy of allocational and representational harms.

  • Results

    The surveyed papers’ motivations are often vague, inconsistent, and lacking normative reasoning, while their techniques are poorly matched to those motivations and weakly grounded in relevant literature outside NLP.

  • Takeaways & Limitations

    Future bias research should examine language and social hierarchies, articulate conceptualizations of bias, and engage more deeply with communities affected by NLP systems.

  • Takeaways & Limitations

    The analysis recognizes that different social groups may experience NLP systems differently because of their social positions and relationships with those systems.

Abstract

from arXiv · show

We survey 146 papers analyzing "bias" in NLP systems, finding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process. We further find that these papers' proposed quantitative techniques for measuring or mitigating "bias" are poorly matched to their motivations and do not engage with the relevant literature outside of NLP. Based on these findings, we describe the beginnings of a path forward by proposing three recommendations that should guide work analyzing "bias" in NLP systems. These recommendations rest on a greater recognition of the relationships between language and social hierarchies, encouraging researchers and practitioners to articulate their conceptualizations of "bias"---i.e., what kinds of system behaviors are harmful, in what ways, to whom, and why, as well as the normative reasoning underlying these statements---and to center work around the lived experiences of members of communities affected by NLP systems, while interrogating and reimagining the power relations between technologists and such communities.

1 Introduction

NLP bias research spans many systems and behaviors, but the survey finds that its motivations often leave harm, affected groups, and normative reasoning unspecified. The authors propose three recommendations centered on social hierarchies, explicit conceptualizations of bias, and engagement with affected communities.

  • Research analyzes “bias” across embedding spaces and tasks including language modeling, coreference resolution, machine translation, sentiment analysis, and toxicity detection.
  • Analyzing “bias” is inherently normative because it deems some system behaviors good and others harmful.
  • The term “bias” describes behaviors that may harm different groups in different ways or for different reasons, even within the same task.
  • The survey finds motivations are often vague and inconsistent, while many techniques are poorly matched to them and not comparable.
  • The authors recommend examining language and social hierarchies, articulating what harms whom and why, and engaging affected communities.

2 Method

The authors identify 146 written-text NLP papers through keyword searches, citation traversal, and manual inspection, then categorize their motivations and techniques using an adapted harms taxonomy. The survey spans multiple tasks and categories, with counts that may overlap because papers can cover multiple items.

  • The survey identifies 146 papers by searching the ACL Anthology for “bias” or “fairness” before May 2020 and excluding unrelated uses and speech papers.
  • Citation-graph traversal and manual inspection of major machine learning, HCI, web, and arXiv venues were used to find additional relevant papers.
  • The authors record the NLP tasks covered, noting that task counts do not sum to 146 because some papers cover multiple tasks.
  • They read each paper to categorize motivations and quantitative techniques using an adapted taxonomy distinguishing allocational and representational harms.
  • The survey presents category counts and illustrative examples, while noting that multiple motivations or techniques can make counts exceed unique-paper totals.

3 Findings

The findings show that NLP bias research often lacks clear normative reasoning and uses techniques that do not align with its stated harms. The survey also identifies narrow attention to predictions and weak engagement with relevant external literature.

  • Motivations: 33% of papers state multiple motivations, while 16% state only vague motivations or no motivations at all.
  • Motivations: 32% of papers lack apparent normative concerns, focusing instead on performance or learned correlations.
  • Motivations: Even clear motivations often do not explain why a behavior is harmful, how it causes harm, or who is affected.
  • Motivations: Papers studying the same task often conceptualize bias differently, leading to different techniques or inconsistent motivations.
  • Motivations: Papers sometimes conflate immediate representational harms with imagined downstream allocational harms; 16% name both in their motivations.
  • Techniques: The proposed techniques generally do not engage relevant literature outside NLP, although stereotyping work sometimes draws on social psychology or established stereotypes.
  • Techniques: 21% of papers include allocational harms in their motivations, but only four propose techniques measuring or mitigating allocational harms.
  • Techniques: Nearly all papers focus on system predictions as potential bias sources, while few examine task definitions, annotation guidelines, or evaluation metrics.

4 A path forward

The paper proposes a path forward for analyzing “bias” in NLP by grounding research in language’s relationship to social hierarchies, making normative reasoning explicit, and centering affected communities.

  • The authors propose three recommendations: engage relevant literature outside NLP, articulate why behaviors are harmful, and center affected communities’ lived experiences.They also call for interrogating and reimagining power relations between technologists and affected communities.
  • 4.1 Language and social hierarchies: Relevant scholarship shows that language labels, ideologies, and contested meanings can reinforce or challenge social hierarchies.The paper draws on sociolinguistics and related fields to connect language practices with stereotypes, oppression, and struggles over material and symbolic resources.
  • 4.1 Language and social hierarchies: Research should examine how social hierarchies, language ideologies, and NLP systems are coproduced across development and deployment.This grounding is intended to explain representational harms and prevent researchers from measuring only what is convenient rather than normatively concerning.
  • 4.2 Conceptualizations of “bias”: Researchers should state what system behaviors are harmful, in what ways, to whom, and why, including the normative reasoning behind those judgments.The recommendation responds to inconsistent conceptualizations of “bias,” including across papers studying the same task.
  • 4.3 Language use in practice: Work should investigate language use in practice through affected communities’ lived experiences and reconsider technologists’ power over development and deployment decisions.The paper notes that incremental fairness techniques can preserve these power relations by assuming systems should continue to exist and leaving decisions with technologists.

5 Case study

The AAE case study shows that NLP systems can perform differently on text associated with African-American English, while existing analyses often fail to situate those differences within racial and linguistic hierarchies.

  • Part-of-speech taggers, language identification systems, and dependency parsers work less well on text containing features associated with AAE.The case study also discusses toxicity detection systems scoring tweets containing AAE-associated features as more offensive.
  • Studies of AAE-related “racial bias” differ in conceptualization, with some focusing only on performance differences and others adding reasoning about stigmatization.The authors contrast papers that treat AAE as a challenging variety with papers that connect system scores to broader social harms.
  • The surveyed papers do not engage with literature on AAE, U.S. racial hierarchies, and raciolinguistic ideologies, leaving these systems socially unsituated.The authors argue that AAE must be understood through questions about its speakers and how they are viewed.
  • Toxicity detection disparities matter not only for system performance but also because they can reproduce stigmatization and disenfranchisement of AAE speakers.The paper connects re-stigmatizing AAE with language ideologies that portray it as ungrammatical, uneducated, and offensive.
  • The authors argue that measuring or mitigating NLP “bias” remains incomplete unless it interrogates structural conditions and power relations affecting racialized communities.They emphasize research on how system development and deployment produce harms and how AAE speakers interact with systems not designed for them.

6 Conclusion

The survey finds recurring weaknesses in NLP “bias” research and proposes three recommendations grounded in the relationship between language and social hierarchies.

  • The survey of 146 papers finds vague, inconsistent motivations and quantitative techniques poorly matched to them and disconnected from relevant literature outside NLP.The authors present three recommendations as a path forward.

A Appendix

The appendix provides examples of papers’ motivations and techniques across several NLP tasks.

  • Table 3 gives examples of motivations and techniques used across several NLP tasks.

A.1 Categorization details

The survey uses explicit inclusion rules to categorize papers’ tasks and motivations, while documenting several exclusions and classification decisions.

  • Task and motivation criteria: A paper counted as covering an NLP task when it analyzed “bias” with respect to that task, not merely its overall performance.A paper examining mitigation in embeddings and its impact on sentiment analysis counted as covering both tasks; performance degradation alone counted only as embeddings.
  • Task and motivation criteria: Motivations included descriptions of the problem motivating a paper or quantitative technique, including normative reasoning.
  • Exclusions and classification decisions: The survey excluded GAP Shared Task papers from motivation and technique counts because shared-task papers generally do not engage with “bias” critically.The authors considered this understandable given the nature of shared tasks.
  • Exclusions and classification decisions: Dabas et al. (2020) was excluded because the surveyors could not determine what its fairness user study measured.
  • Exclusions and classification decisions: Liu et al. (2019) was categorized as “Questionable correlations” because another sentence supplied additional detail; without it, the motivation would have been “Vague/unstated.”

A.2 Full categorization: Motivations

The motivation categorization spans six categories, including allocational harms, stereotyping, other representational harms, questionable correlations, and vague or unstated motivations, with examples documented in Table 3.

  • Motivation categories: Table 3 organizes examples of papers’ motivations into categories such as allocational harms, stereotyping, other representational harms, and questionable correlations.The table caption states that bolded text in quoted passages determines the categorizations.
  • Motivation categories: The listed examples demonstrate that the survey tracks distinct motivation categories across the papers rather than treating all uses of “bias” as equivalent.
  • Motivation categories: The categorization also includes vague or unstated motivations and surveys, frameworks, and meta-analyses.

B Full categorization: Techniques

The technique categorization groups proposed approaches for measuring or mitigating “bias” into categories paralleling the harms and correlations identified in the survey.

  • Technique categories: The technique categories include allocational harms, stereotyping, other representational harms, and questionable correlations.
  • Technique categories: The categorization also lists surveys, frameworks, and meta-analyses among the forms of work included in the technique overview.
Loading 2005.14050v2…