Source-linked AI summary

A Survey of Race, Racism, and Anti-Racism in NLP

Anjalie Field, Su Lin Blodgett, Zeerak Waseem, Yulia Tsvetkov

arXiv:2106.11410v2cs.CL

TL;DR

Race and racial bias remain minimally explored in NLP, despite the field’s ties to language and racialized social hierarchies. The paper surveys 79 ACL Anthology papers and finds racialized differences across NLP systems and research practices, while identifying narrow task coverage, simplified race categories, and limited inclusion of marginalized people.

  • Problem

    Race and racial bias have been minimally explored in NLP, leaving limited examination of how NLP research and systems engage with racialized differences.

  • Method

    The paper surveys 79 ACL Anthology papers and synthesizes how race and racial bias are addressed across NLP pipelines, research practices, and related fields.

  • Results

    NLP systems and research practices produce differences along racialized lines, with representational harms and performance gaps appearing throughout NLP pipelines.

  • Takeaways & Limitations

    NLP research should explicitly incorporate race, examine the social hierarchies its systems uphold, and engage directly with marginalized communities of color.

  • Takeaways & Limitations

    The survey is primarily grounded in U.S. research, so future studies should ground racialization analyses in appropriate geocultural contexts.

Abstract

from arXiv · show

Despite inextricable ties between race and language, little work has considered race in NLP research and development. In this work, we survey 79 papers from the ACL anthology that mention race. These papers reveal various types of race-related bias in all stages of NLP model development, highlighting the need for proactive consideration of how NLP systems can uphold racial hierarchies. However, persistent gaps in research on race and NLP remain: race has been siloed as a niche topic and remains ignored in many NLP tasks; most work operationalizes race as a fixed single-dimensional variable with a ground-truth label, which risks reinforcing differences produced by historical racism; and the voices of historically marginalized people are nearly absent in NLP literature. By identifying where and how NLP literature has and has not considered race, especially in comparison to related fields, our work calls for inclusion and racial justice in NLP research practices.

1 Introduction

Race and language are mutually shaped through racialized hierarchies and the communication of stereotypes, yet race and racial bias remain minimally explored in NLP. This survey examines how racial bias appears across NLP pipelines and compares NLP’s engagement with race to related fields.

  • Race and language are mutually constructed through racial hierarchies, while language communicates and perpetuates stereotypes and prejudice.
  • Questions of race and racial bias have been minimally explored in NLP, despite substantial attention to racial bias in related areas of computing.
  • The survey examines 79 ACL Anthology papers mentioning race, racial, or racism and identifies racial biases across all stages of NLP model pipelines.
  • The authors compare NLP with HCI, fairness in machine learning, linguistics, and related NLP scholarship while maintaining an explicit focus on race and racism.
  • The survey finds that NLP systems and research practices produce differences along racialized lines and calls for greater inclusion and racial justice.

2 What is race?

Race is shaped by historical, social, political, and colonial forces rather than biological differences, and its meaning varies across social contexts. Race is also multidimensional, while this survey’s analysis is primarily grounded in the United States.

  • Race is understood as a social construct shaped by historical events, social forces, political power, and colonial conquest rather than natural biological differences.
  • Race can refer to racial identity, observed race, or reflected race, and these categorizations may differ across schemas and contexts.
  • Racial categories have no universal ground-truth labels because their meanings depend on historical and social context.
  • The survey primarily analyzes U.S. research, while noting that race and racism are global constructs requiring appropriate geocultural grounding elsewhere.

3 Survey of NLP literature on race

The survey analyzes 79 ACL Anthology papers and related literature to trace racial bias across data, labels, models, outputs, and social analyses. Across these stages, the literature identifies racialized disparities but also narrow coverage and interpretive limitations.

  • 3.1 ACL Anthology papers about race: The final analysis set contains 79 papers after excluding papers that only mention race incidentally or use racism solely as an abusive-language label.
  • 3.2 NLP systems encode racial bias: The survey organizes evidence around five pipeline stages: data, data labels, models, model outputs, and social analyses of outputs.
  • 3.2 NLP systems encode racial bias: Parsing systems trained on White Mainstream American English perform poorly on African American English, while data authorship and representation shape whom models serve.
  • 3.2 NLP systems encode racial bias: Annotation schemas and instructions can introduce racial bias, including through parsing standards and annotators’ differing judgments of AAE-associated tweets.
  • 3.2 NLP systems encode racial bias: Different model architectures and instances can produce different racial-bias patterns, but the factors driving disparate performance require further investigation.
  • 3.2 NLP systems encode racial bias: Abusive-language classifiers disproportionately flag AAE-associated text and identity terms, while GPT produces more negative sentiment for AAE-like inputs.
  • 3.2 NLP systems encode racial bias: Embedding associations do not always correlate with human opinions, and racial analyses remain vulnerable to interpretive bias and uncommon researcher self-disclosure.
  • 3.2 NLP systems encode racial bias: When researchers have looked for racial biases in NLP systems, they have usually found them, motivating proactive attention to data collection, annotation, use, and interpretation.

4 Limitations in where and how NLP operationalizes race

The survey finds that NLP research engages race through narrow datasets, fixed classification schemes, and mostly single-dimensional analyses, leaving many tasks and potential harms unexamined.

  • 4.1 Common data sets are narrow in scope: Race research in NLP relies on a limited range of datasets, often failing to represent race’s multiple dimensions or the people affected by models.Common sources include inferred race tweets, name lists, and identity terms; these do not capture self-identified race or all relevant harms.
  • 4.1 Common data sets are narrow in scope: Names and hand-selected identity terms can miss bias and generalize poorly across languages, domains, geographies, and time.Evidence from related work shows that performance gaps may remain after overt identity indicators are removed.
  • 4.2 Classification schemes operationalize race as a fixed, single-dimensional U.S.-census label: U.S.-census categories are repeatedly reused without sufficient contextual justification, risking the presentation of historically produced racial divisions as natural.The survey identifies U.S.-centrism and English-language focus as recurring features of this operationalization.
  • 4.2 Classification schemes operationalize race as a fixed, single-dimensional U.S.-census label: Most surveyed work analyzes race or gender separately, overlooking people who experience marginalization across intersecting axes.The survey contrasts analyses centered on Black men or white women with the underexamined experiences of Black women.
  • 4.3 Race is ignored in many NLP tasks: Race remains unexamined in common NLP applications such as machine translation, summarization, and question answering.The survey therefore calls for broader task coverage, multidimensional datasets, and context-sensitive racial categories.

5 NLP propagates marginalization of racialized people

NLP marginalizes racialized people through data practices, researcher underrepresentation, and deployed systems whose harms can fall on the communities they purport to serve.

  • 5.1 People generate data: Diverse datasets can create privacy and profiling risks when researchers overlook consent, usage limits, privacy, or communities’ wishes about data use.The Diversity in Faces dataset is presented as an example: publicly available Flickr photos were used without explicit consent, prompting privacy-related allegations.
  • 5.2 People build models: NLP research underrepresents people from Africa and Black or African American communities, while demographic statistics do not capture lived experiences such as microaggressions and invisible diversity labor.The paper reports 5 of 2,695 author affiliations from Africa in five major 2018 conferences and fewer than 4% of relevant U.S. computer-science doctorates awarded to Black or African American students.
  • 5.2 People build models: Interest convergence can favor incremental, surface-level bias studies over research that would require fundamental changes to the field.The paper connects this risk to institutional incentives and the interests of people in positions of power.
  • 5.3 People use models: Researchers often leave the racial harms of applications unaddressed, including demographic prediction, prison-term prediction, and low-resource NLP that may increase technological inequality.The paper links these concerns to technology’s history of policing racial minorities and to structural racism and differences in lived experience.
  • 5.3 People use models: Abusive-language classifiers can falsely label identity terms and African American English as toxic, potentially censoring the people these systems aim to help.The paper presents this as a case where failing to consider affected people produces harmful deployment outcomes.
  • 5.4 Engaging affected communities: Participatory design can amplify affected communities’ voices, but it must be context-specific, long-term, and genuine rather than participation-washing.The paper presents participatory design as a useful but non-sufficient approach.

6 Discussion

The discussion argues that NLP must explicitly incorporate race because systems can uphold racial hierarchies through harms, narrow race definitions, and exclusion of marginalized people.

  • 6 Discussion: The authors conclude that NLP research should explicitly incorporate race because technical systems can uphold or dismantle social relations.The paper frames non-intervention as a risk of perpetuating or exacerbating prevalent social systems such as racism.
  • 6 Discussion: NLP research identifies representational harms and performance gaps, restricts race to narrow tasks and definitions, and excludes underrepresented people as technology producers and consumers.The paper notes that related concerns may also apply to socioeconomic class, disability, and sexual orientation.
  • 6 Discussion: The paper calls for direct engagement with marginalized communities of color through approaches such as participatory and value-sensitive design.It points to organizations including Black in AI, Masakhane, Data for Black Lives, and the Algorithmic Justice League as partnership starting points.
  • 6 Discussion: No single dataset, model, or guideline can solve racism in NLP, and the paper does not address every racialized effect of NLP research.The discussion specifically notes environmental costs of large language models and their disproportionate effects on marginalized communities.

7 Ethical Considerations

The authors situate their analysis within U.S. and U.K./European academic contexts and acknowledge that their perspectives may not represent viewpoints outside their institutions and experiences.

  • 7 Ethical Considerations: The authors’ U.S. and U.K./European institutional and personal contexts shape the work, while viewpoints outside those experiences may not be fully represented.They note that some authors identify as people of color and that all authors are situated within exclusionary academic research practices.

A ACL Anthology Venues

The survey distinguishes ACL and non-ACL publication venues using separate event lists.

  • ACL venues include major conferences, Findings, workshops, special-interest groups, and related ACL events.
  • Non-ACL venues include conferences and workshops such as COLING, LREC, PACLIC, RANLP, and IJCNLP.

B Additional Survey Metrics

The survey reports publication, venue, and race-categorization patterns across the dataset. Race-related work increased in recent years, appeared substantially in workshops, and commonly used binary Black/white categories identified through names or keywords.

  • The survey reports publication counts by year and venue alongside how papers operationalized race.
  • Race-related publication increased in 2019 and 2020 alongside broader NLP growth and increasing attention to social issues.
  • 46.8% of surveyed papers were published in workshops, while the authors found no evidence that race research was siloed into particular venues.
  • Most papers focused on binary Black/white racial categories, with few considering less commonly selected census categories such as Native American or Pacific Islander.
  • Names and explicit racial keywords were each used by 10 papers as the most common methods for identifying people’s race.

C Full List of Surveyed Papers

The surveyed-paper list spans NLP venues and years, including work on abusive language, text representations, social media, text generation, and bias.

  • The surveyed papers cover abusive language, text representations, social science and media, text generation, and other NLP applications.
  • The papers use varied study types, including corpus collection and analysis, bias detection, debiasing, model development, and surveys or position papers.
Loading 2106.11410v2…