Source-linked AI summary
CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Nikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. Bowman
TL;DR
Pretrained language models can learn social biases from minimally filtered training text, creating a need for reliable measurement. The paper introduces CrowS-Pairs, a crowdsourced benchmark of minimally different stereotype pairs across nine bias types, and finds that all three evaluated MLMs substantially favor stereotypes in every category. The dataset is intended to support evaluating progress toward less biased models, but its US-specific and non-exhaustive scope prevents treating a low score as proof that a model is unbiased.
Problem
Language models can learn social biases from training corpora, creating a need to identify and quantify those biases directly for safer downstream use.
Method
The authors introduce CrowS-Pairs, a crowdsourced benchmark of minimally different sentence pairs covering nine social-bias categories about disadvantaged US groups.
Results
All three evaluated MLMs substantially favor stereotype-expressing sentences in every CrowS-Pairs category.
Takeaways & Limitations
CrowS-Pairs can serve as a metric for stereotyping in future work on model debiasing.
Takeaways & Limitations
CrowS-Pairs has limited, US-specific coverage and is not an assurance that a model is truly unbiased.
Abstract
from arXiv · showhide
Pretrained language models, especially masked language models (MLMs) have seen success across many NLP tasks. However, there is ample evidence that they use the cultural biases that are undoubtedly present in the corpora they are trained on, implicitly creating harm with biased representations. To measure some forms of social bias in language models against protected demographic groups in the US, we introduce the Crowdsourced Stereotype Pairs benchmark (CrowS-Pairs). CrowS-Pairs has 1508 examples that cover stereotypes dealing with nine types of bias, like race, religion, and age. In CrowS-Pairs a model is presented with two sentences: one that is more stereotyping and another that is less stereotyping. The data focuses on stereotypes about historically disadvantaged groups and contrasts them with advantaged groups. We find that all three of the widely-used MLMs we evaluate substantially favor sentences that express stereotypes in every category in CrowS-Pairs. As work on building less biased models advances, this dataset can be used as a benchmark to evaluate progress.
1 Introduction
CrowS-Pairs is introduced as a crowdsourced benchmark for measuring nine categories of social bias in masked language models, focusing on stereotypes about historically disadvantaged US groups. The authors report that widely used MLMs express bias across categories, with variation by category.
- Benchmark design: CrowS-Pairs measures whether language models prefer more stereotypical sentences about historically disadvantaged groups in the United States.Each example contrasts a stereotype or anti-stereotype involving a disadvantaged group with a contrasting advantaged group.
- Benchmark design: The benchmark is crowdsourced rather than template-based, allowing greater diversity in stereotype content and sentence structure.Its scope reflects biases widely acknowledged in the United States.
- Benchmark coverage: CrowS-Pairs covers nine bias types, including race, gender identity, sexual orientation, religion, age, nationality, disability, physical appearance, and socioeconomic status.
- Findings: All three evaluated MLMs express bias in every CrowS-Pairs category, with comparatively higher bias for religion and lower bias for gender and race.The authors describe gender and race as comparatively easier categories for the models.
- Findings: The authors argue that CrowS-Pairs is more reliable than standard masked-language-modeling metrics for measuring stereotype use and demonstrates deployment risks for recent MLMs.
2 Data Collection
The authors collect minimally distant stereotype and anti-stereotype sentence pairs through MTurk, then validate and filter them with multiple annotations. The resulting dataset contains 1508 examples spanning nine bias categories.
- Use scope: The collected data is intended for model evaluation rather than straightforwardly effective bias-mitigation training.The authors leave collection of training data to future work.
- Bias categories: The dataset uses nine bias categories, including race/color, gender, socioeconomic status, nationality, religion, age, sexual orientation, physical appearance, and disability.
- Writing pairs: Crowdworkers write minimally distant sentence pairs by changing only the words identifying disadvantaged and contrasting advantaged groups.Examples may express a stereotype or violate one as an anti-stereotype.
- Validation: Annotators assess stereotype presence, minimal distance, and bias category, and examples are accepted when at least 3 of 6 annotators agree on validity and minimal distance.Examples lacking majority agreement on bias type are filtered out.
- Resulting data: 1508 examples remain after collecting 2000 examples, removing 490 during validation, and removing two fully overlapping pairs.Average inter-annotator agreement on validity is 80.9%.
- Resulting data: Race/color contributes 516 examples, about one-third of the dataset, while anti-stereotype examples comprise 15%.The authors report that each bias category is nevertheless well represented.
3 Measuring Bias in MLMs
CrowS-Pairs measures whether masked language models prefer stereotypical sentences by estimating conditional likelihoods over shared tokens, reducing frequency confounds. The metric is an approximation whose validation favors iterative masking of unmodified tokens.
- Metric: The score conditions the likelihood of overlapping, unmodified tokens on the differing, modified tokens.This reverses the conditioning direction used by StereoSet to reduce the effect of differing word frequencies in pretraining data.
- Scoring procedure: For each sentence, the adapted pseudolikelihood procedure masks one unmodified token at a time until all such tokens have been masked.The approach estimates p(U|M, θ), where U denotes unmodified tokens and M denotes modified tokens.
- Validation: The metric is an approximation to the true conditional probability and was informally compared with alternative masking formulations.Among the tested formulations, iterative masking was found to be the most reliable approximation for preferring meaningful over nonsensical sentences.
- Metric: The metric compares how often an MLM assigns higher likelihood to a more stereotypical sentence than to its less stereotypical pair.An unbiased model should achieve an ideal bias score of 50%.
4 Experiments
Experiments evaluate three widely used MLMs across CrowS-Pairs and related bias datasets, examining overall bias, confidence, stereotype subsets, and bias categories. All models show substantial bias, with confidence and category patterns varying across models and datasets.
- Experimental setup: Three MLMs—BERTBase, RoBERTaLarge, and ALBERTXXL-v2—are evaluated on CrowS-Pairs alongside WinoBias and StereoSet comparisons.The related datasets provide baseline measurements for comparing patterns in model performance.
- Overall results: All three models exhibit substantial bias on all four evaluated datasets, while BERT has the lowest bias score on each dataset.BERT is also the smallest model and performs worst on most downstream tasks in the cited comparison.
- Implications: The authors caution that debiasing may degrade MLM performance on naturally occurring text, leaving preservation of downstream performance as a future challenge.This concern is presented as a possibility rather than a demonstrated outcome.
- Model confidence: ALBERT has the highest CrowS-Pairs bias score and the widest confidence distribution, while confidence differences between stereotypical and less stereotypical choices increase across BERT, RoBERTa, and ALBERT.The median differences are 1.2 for BERT, 2.3 for RoBERTa, and 3.2 for ALBERT.
- Bias categories: All models show bias in every CrowS-Pairs category, with comparatively higher scores for religion and lower scores for gender and race.The authors describe gender and race as comparatively easier categories for the models.
- Stereotype subsets: All three models are less biased on anti-stereotype examples than on stereotype examples.RoBERTa scores 51.4 on anti-stereotype examples and 66.3 on stereotype examples; BERT and ALBERT improve by 4.2 and 4.4 percentage points, respectively.
5 Dataset Analysis
The dataset analysis examines how CrowS-Pairs examples are constructed, validated, and compared with StereoSet, while identifying several data-quality ambiguities and limitations.
- Example construction: CrowS-Pairs examples commonly use first names and explicit group names to indicate the groups being compared.The name Jamal appears 17 times, and names such as “African Americans” and “Mexicans” are also common.
- Data quality: Some examples contain bias whose status as an undesirable model bias is unclear, such as a Black-associated name prompting alarm at KKK members.The authors report that such examples are infrequent.
- Data quality: An example involving race-color and nationality labels illustrates that annotator disagreement can signal intersectional bias rather than noise.The example received three race-color labels and two nationality labels during validation.
- Data quality: 23 examples fail to isolate the stereotype because model preference could reflect sentence priors rather than the intended group contrast.The paper gives “Women are not hunters” versus “Men are hunters” as an example of this problem.
- Comparison with StereoSet: CrowS-Pairs has a much higher valid-example rate than StereoSet’s intrasentence data, while their inter-annotator agreement is similar.Both datasets were independently validated using five annotations per example and majority voting.
6 Related Work
Related work measures bias through embeddings, downstream tasks, language-model evaluations, and mitigation methods, while also developing recommendations for bias research.
- Bias measurement: Prior work finds that word and language-model embeddings reflect biases involving gender, age, religion, and socioeconomic status.Caliskan et al. connect geometric bias in GloVe embeddings with crowd judgments, while later work extends the evaluated bias categories.
- Bias measurement: Some studies evaluate bias in downstream tasks such as coreference resolution and relation extraction.This represents one line of work distinct from evaluations conducted within the language-modeling framework.
- Bias measurement: StereoSet measures bias in language-modeling settings using intrasentence and intersentence examples, including discourse-level evaluation.Its intrasentence data is discussed as related work for comparison with CrowS-Pairs.
- Bias mitigation: Other work probes generated language with sentiment analysis and uses the resulting evaluation for model debiasing.This approach targets bias in language-model generations.
- Bias mitigation: Mitigation methods include removing gender-related linear projections, although follow-up work reports that this can hide rather than remove bias.Other debiasing work reports lower SEAT bias scores while maintaining downstream task performance.
- Research guidance: A survey of 146 NLP papers provides recommendations that the authors follow when positioning and explaining their work.The recommendations concern research that analyzes or mitigates bias.
7 Ethical Considerations
The paper treats CrowS-Pairs as a sensitive measurement resource for debiasing, while warning that its scope and metric cannot establish that a model is unbiased.
- Responsible use: The dataset should not be used to train language models because doing so would defeat its purpose of measuring social bias for debiasing.The authors explicitly frame the data as sensitive and intended for evaluation rather than language-model training.
- Scope and interpretation: A low CrowS-Pairs score cannot support the claim that a model is completely bias free.The dataset has limited scope, its numeric metric can be misused, and its covered stereotypes are specific to the United States rather than exhaustive.
8 Conclusion
The paper concludes that CrowS-Pairs is a crowdsourced nine-category challenge dataset and that widely used MLMs show substantial bias across every category. Evaluation remains limited to MLMs, motivating metrics for autoregressive models and further debiasing work.
- Conclusion: CrowS-Pairs is a crowdsourced challenge dataset covering nine categories of social bias.The dataset is intended to provide a metric for stereotyping in future model-debiasing research.
- Conclusion: Widely used MLMs exhibit substantial bias in every CrowS-Pairs category, highlighting dangers in deploying systems built around them.The authors expect CrowS-Pairs to serve as a metric for stereotyping in future debiasing work.
- Future work: The evaluation is limited to MLMs, and testing autoregressive language models requires developing suitable metrics.The authors identify this as a clear next step and note that broader debiasing may require further methods work.
A Data Statement
CrowS-Pairs is a crowdsourced evaluation dataset for measuring U.S. stereotypical biases in masked language models. It spans nine bias types, uses U.S.-based workers and validation, and is intended for evaluation rather than training.
- Coverage: 1,508 examples cover nine types of social bias, including race, gender/gender identity, sexual orientation, religion, age, nationality, disability, physical appearance, and socioeconomic status.Race, gender/gender identity, and socioeconomic status are the three most frequent types.
- Dataset structure: Each example pairs a sentence about a historically disadvantaged U.S. group with a contrasting advantaged group through a minimal edit changing only group-identifying words.The disadvantaged-group sentence can demonstrate or violate a stereotype.
- Collection and validation: Crowdworkers wrote examples from prompts based on MultiNLI or ROCStories, and five other U.S.-based crowdworkers validated each example.Workers were required to have completed at least 5,000 HITs and maintain an acceptance rate above 98%.
- Assumptions: The dataset assumes most examples use U.S. English because workers were required to be in the United States, although worker language varieties were not collected.Manual analysis found that most, if not all, sentences fit standard written English.
- Use and scope: CrowS-Pairs contains deliberately biased and often offensive statements, so the authors recommend using it as an evaluation set rather than as a source of English examples or training data.Its purpose is to measure stereotypical biases as a step toward building fairer NLP systems.
- Limitations: The authors caution that the dataset has limited scope and a single numeric bias metric, so a low score does not guarantee freedom from stereotypical bias.This limitation constrains how model scores should be interpreted.
B CrowS-Pairs Data Collection HIT Layout
The data-collection HIT asks U.S. crowdworkers to write a stereotype or anti-stereotype sentence from a prompt, then create a minimally changed counterpart naming an advantaged group. Workers also label the relevant attribute and receive instructions about quality and pay.
- Quality and compensation: The HIT includes full instructions, optional reporting of actual working time, and a warning when the pair’s Levenshtein distance exceeds 13.The distance warning helps flag pairs that differ by more than the intended minimal edit.
- Data collection: Workers begin with a prompt and write a sentence highlighting a cultural stereotype about a historically disadvantaged U.S. group.Prompts include the example: “I shall not ask her to tell me anything, he said quietly.”
- Data collection: Workers copy the first sentence and change the historically disadvantaged group to a contrasting advantaged group.The resulting pair is designed to differ in the group-identifying words.
- Annotation: Workers classify the first sentence as a stereotype or anti-stereotype.The HIT presents “Stereotype” and “Antistereotype” as the response options.
- Annotation: Workers select the most applicable attribute from race/color, sexual orientation, gender/gender identity or expression, age, religion, nationality or citizenship status, disability, physical appearance, and socioeconomic status/occupation.The attribute list is presented as the main categorical label for each example.
C CrowS-Pairs Data Validation HIT Layout
The validation HIT presents sentence pairs and asks annotators to judge each sentence’s stereotype status, verify that the pair differs only in the relevant wording, and identify the applicable attribute. Separate layouts support validation examples associated with CrowS-Pairs and StereoSet.
- Validation judgments: Annotators judge whether each sentence clearly demonstrates a stereotype, clearly violates one, or does neither.The HIT asks this question separately for Sentence 1 and Sentence 2.
- Pair validation: Annotators verify that the two sentences are identical except for words indicating which person or group is discussed.The HIT gives a John/Jane example where the rest of the sentence remains unchanged.
- Attribute labeling: Annotators identify whether the stereotype concerns race/color, sexual orientation, gender/gender identity or expression, age, religion, nationality or citizenship status, disability, physical appearance, or socioeconomic status/occupation.When multiple attributes apply, they select the one they consider most relevant.
- HIT layouts: The validation HIT design is stated to be used in both rounds of CrowS-Pairs validation, while a separate layout is identified for StereoSet.The latter is labeled as HIT Layout 3.
- Validation examples: A second validation example asks the same stereotype-status, pair-identity, and attribute questions for sentences contrasting a Colombian man as a druglord versus Jewish.The example illustrates validation of potentially different stereotype-related wording.