Source-linked AI summary

UnQovering Stereotyping Biases via Underspecified Questions

Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, Vivek Srikumar

arXiv:2010.02428v3cs.CL

TL;DR

The paper addresses how stereotyping biases in language representations affect downstream QA, a question not previously explored. It introduces UNQOVER, which uses underspecified questions and a formalism that isolates positional dependence and question independence. Across four stereotype classes and multiple models, it finds notable bias, often greater bias in larger models, and fine-tuning effects that vary by dataset and model size.

  • Problem

    How stereotyping biases in language representations affect downstream QA models remains unexplored.

  • Method

    UNQOVER probes QA models with underspecified questions and isolates positional dependence and question independence when estimating bias.

  • Results

    Across five models, two QA datasets, and four stereotype classes, all evaluated models show notable stereotyping biases; larger models often show more bias, and fine-tuning effects vary with dataset and size.

  • Takeaways & Limitations

    UNQOVER supports evaluating and potentially mitigating stereotyping biases in QA models and their underlying language models.

  • Takeaways & Limitations

    The analysis uses a binary view of gender and English resources, while its group and attribute choices and extracted prejudiced statements may reflect Western-specific views of bias.

Abstract

from arXiv · show

While language embeddings have been shown to have stereotyping biases, how these biases affect downstream question answering (QA) models remains unexplored. We present UNQOVER, a general framework to probe and quantify biases through underspecified questions. We show that a naive use of model scores can lead to incorrect bias estimates due to two forms of reasoning errors: positional dependence and question independence. We design a formalism that isolates the aforementioned errors. As case studies, we use this metric to analyze four important classes of stereotypes: gender, nationality, ethnicity, and religion. We probe five transformer-based QA models trained on two QA datasets, along with their underlying language models. Our broad study reveals that (1) all these models, with and without fine-tuning, have notable stereotyping biases in these classes; (2) larger models often have higher bias; and (3) the effect of fine-tuning on bias varies strongly with the dataset and the model size.

1 Introduction

UNQOVER probes stereotyping in QA through underspecified questions, while isolating positional and question-independence reasoning errors that can distort bias estimates. Across models and datasets, the study finds substantial biases and dataset- and size-dependent fine-tuning effects.

  • Underspecified questions provide no factual support for either candidate, so a model’s preference can expose stereotyping associations.
  • QA predictions can be confounded by subject position and by failing to change when the questioned attribute is negated.
  • UNQOVER measures stereotyping biases in QA models using underspecified questions.
  • The framework removes these reasoning-error factors to reveal stereotyping biases.
  • Across five models, two QA datasets, and four bias classes, larger models tend to show more bias, while fine-tuning effects vary by dataset and model size.Bias increases with SQuAD and decreases with NewsQA; fine-tuning reduces distilled-model bias but can amplify bias in larger models.
  • The analysis acknowledges a binary treatment of gender and a likely Western, Educated, Industrialized, Rich, and Democratic skew in its English resources.The framework is described as adaptable to more nuanced gender perspectives and other cultural contexts.

2 Related Work

Prior bias research often studies representations or downstream label changes, whereas UNQOVER uses underspecified QA inputs and model scores to expose comparative stereotypes that may not change predicted labels.

  • Bias research has commonly analyzed pretrained representations or intermediate classification tasks.
  • Downstream-task studies often detect bias through changes in predicted labels, aligning analysis with system use.
  • UNQOVER compares candidate scores in underspecified QA inputs to reveal stereotypes that categorical predictions may obscure.
  • The framework focuses on exploring bias across QA models and may support future bias-mitigation efforts.

3 Constructing Underspecified Inputs

UNQOVER constructs minimally informative contexts containing two candidates and an attribute, then uses their predictions to probe whether models express unsupported preferences.

  • The framework treats highly confident predictions without input support as unwarranted preferences embedded in model parameters.
  • Templates contain two subjects and one attribute, instantiated across subject and attribute lists.
  • Questions are designed to make each subject equally likely, while attributes are chosen so preferring either subject is unfair and not common knowledge.
  • The method aggregates scores across examples to distinguish true bias from minor deviations from equal preference.
  • The same construction extends to masked language models, using scores for the two specified candidate fillers.

4 Uncovering Stereotypes

UNQOVER probes stereotyping in QA models with underspecified questions while separating bias from positional dependence and attribute indifference. Its symmetric comparative metrics then aggregate subject-attribute preferences and show that measured bias is systematic rather than driven only by outliers.

  • Positional Dependence: QA predictions can depend strongly on subject order even when the information content is unchanged.The positional error compares a subject’s scores before and after swapping subject positions, then averages this difference across subjects, attributes, and templates.
  • Attribute Independence: Negating an attribute can leave model predictions inconsistent with the expected answer reversal, revealing attribute-independence errors.In the example, Gerald receives 0.26 for the original attribute while Jennifer receives 0.62 for its negation; antonyms and “never” improve negation recognition.
  • Bias Measurement: UNQOVER uses symmetric original and swapped questions plus negated attributes to neutralize positional and attribute-independent components of model scores.The comparative metric is bounded in [−1, 1] and satisfies positional independence, attribute-negation dependence, complementarity, and zero centrality.
  • Bias Measurement: The comparative score can reverse the apparent preference from an uncorrected question: Gerald’s B score is 0.16, Jennifer’s is −0.15, and C is 0.31.Without removing the confounding factors, the single example would instead make Jennifer appear preferred as the hunter.
  • Aggregated Metrics: The framework aggregates comparative scores into subject-attribute bias, intensity, and count-based measures to capture preferences across subjects, attributes, and templates.Count-based aggregation avoids having a few high-scoring outliers dominate estimates, while intensity averages each subject’s most extreme attribute bias.
  • Aggregated Metrics: All evaluated datasets and models have η ∼0.5, indicating that the observed bias is systematic rather than explainable by only a few outliers.The η metric uses signed comparative outcomes and can also aggregate their absolute values to measure extremeness.

5 Experiments

The experiments evaluate stereotyping bias across four classes, five transformer models, and multiple fine-tuning settings. Bias varies with model size, dataset, and reasoning reliability, while the analyzed associations are explicitly contextual rather than general claims about groups or countries.

  • Experimental setup: The study compares five transformer models across pre-trained, SQuAD-fine-tuned, and NewsQA-fine-tuned settings, covering gender, nationality, ethnicity, and religion.The evaluation uses templates, selected subjects, and prejudice-related attributes to construct evaluation datasets.
  • General trends: Larger QA models generally show more intensive bias than their base counterparts, with a few NewsQA exceptions for RoBERTa.BERTDist is among the least biased models across different bias classes.
  • General trends: Fine-tuning shifts bias differently by model and dataset: it reduces bias in BERTDist, can amplify bias in larger models, and produces substantially lower bias with NewsQA than SQuAD.NewsQA models show lower bias consistently across all four classes, and sometimes lower intensity than their masked-LM counterparts.
  • Gender-occupation bias: Gender results consistently associate “nurse,” “model,” and “dancer” with female names, while male-associated occupations vary between BERT and RoBERTa.BERTDist’s highest female-bias score is still negative, suggesting a general preference for male names across occupations.
  • Nationality bias: Nationality results show stronger bias in RoBERTa than BERT, with many listed countries almost always preferred and a boundary between Western and non-Western geoschemes.The nationality rankings are dataset-based and concern negative attributes rather than general statements about countries.
  • Ethnicity and religion bias: Ethnicity and religion rankings show contrasting group sentiment patterns, but the bias intensity remains small, with |γ(x)|≤0.03.Arab and African-American rank toward one extreme, European toward the other, and Muslim is ranked most negatively with low variance.
  • Reasoning errors: QA models exhibit substantial positional and attribute reasoning errors, so raw scores can misestimate stereotyping bias.RoBERTa has more positional error than similarly sized BERT models, while fine-tuning does not improve attribute-error consistency.

6 Conclusions & Future Work

The conclusion presents UNQOVER as a framework for measuring stereotyping bias in QA models and masked language models while accounting for reasoning errors. Its broad experiments reveal model- and fine-tuning-dependent bias, but the analysis remains bounded by simplified and culturally specific choices.

  • Contributions: UNQOVER combines underspecified input construction with evaluation metrics that factor out reasoning-error effects.The framework is applied to more than 15 transformer models across four stereotype classes.
  • Limitations and future work: The analysis uses a binary gender view and common nationality, ethnicity, and religion groups, while some prejudice statements and training data reflect Western-specific views.The authors identify more inclusive studies as future work.

A Appendix

The appendix explains that presenting every model prediction is impractical, so it reports broader results and uses one representative fine-tuned model for specific prediction examples.

  • Appendix scope: Because the study evaluates many models, the appendix reports broader results and uses RoBERTaB fine-tuned on SQuAD for specific prediction examples.This choice is made to keep detailed prediction displays practical.

A.1 Details of Experiments

The experiments use released transformer language models and fine-tuned QA variants, with separate procedures for SQuAD, NewsQA, and BERTDist. Evaluation reports QA performance on corresponding development sets.

  • Model preparation: Pre-trained transformer language models are used directly or fine-tuned with standard settings for SQuAD and NewsQA.NewsQA prediction requires adding the “(CNN) —” header to obtain high average answer probabilities.
  • Model preparation: BERTDist is fine-tuned directly without additional downstream distillation to isolate the effect of fine-tuning.This setup is intended to make fine-tuning effects easier to study.
  • Evaluation: QA models are evaluated with F1 scores on their corresponding official development sets.Training and evaluation use a 384-token window containing the ground-truth answer.

A.2 Proof of Propositions in Sec 4.2

The metric C is designed to remove ordering and attributive-independence effects from bias estimates. The proofs establish positional independence and cancellation of errors caused by negated attributes.

  • Complementarity and zero centrality are properties of C.
  • C is independent of the ordering of the two subjects.
  • The metric cancels reasoning errors caused by attributive independence.

A.3 Count-based Bias Metric

The count-based η metric compares model win/lose ratios across model sizes. Most models show similar bias levels, while η values near 0.5 indicate that biases are aggregated by small margins.

  • η measures model-wise bias through the win/lose ratio.
  • Models are mostly biased at similar levels under the count-based metric.
  • η values close to 0.5 indicate that most biases are aggregated by small margins.

A.4 Dataset Generation

The datasets cover gender, nationality, ethnicity, and religion using curated subjects, occupations, countries, and templates. The evaluation also examines model predictions and bias rankings across fine-tuned and masked language models.

  • Gender-occupation: The gender-occupation dataset uses gendered names, occupations, and templates, with automated grammar correction during instantiation.
  • Nationality: The nationality dataset combines country names and demonyms in templates, with countries selected for relatively balanced continental representation.
  • Ethnicity and religion: Ethnicity and religion datasets use curated subject lists and shared templates, with out-of-vocabulary filtering for masked language models.
  • Prediction analysis: Gendered predictions are ranked for fine-tuned RoBERTaB and pretrained language models, scoring pronouns through the maximum probability over gendered names and pronouns.
  • Bias analyses: The appendix reports top-3 nationality-attribute pairs and subject bias scores for ethnicity and religion models.
Loading 2010.02428v3…