Source-linked AI summary

Persistent Anti-Muslim Bias in Large Language Models

Abubakar Abid, Maheen Farooqi, James Zou

arXiv:2101.05783v2cs.CLcs.LG

TL;DR

Religious bias in language models was relatively unexplored, motivating an investigation of whether GPT-3 captures Muslim-violence associations. The paper probes GPT-3 through multiple forms of completion, analogy, and story generation, finding that the association appears persistently and creatively across uses. Positive contextual phrases reduce violent completions to a certain extent, but the intervention can steer outputs toward a specific topic and may not be general.

  • Problem

    Religious bias in large language models has been relatively unexplored, including whether GPT-3 captures persistent Muslim-violence associations.

  • Method

    The paper probes GPT-3’s learned associations through prompt completion, analogical reasoning, and long-form image-story generation, then tests positive contextual phrases for reducing violent completions.

  • Results

    GPT-3 manifests Muslim-violence associations persistently and creatively across different uses rather than merely memorizing violent examples.

  • Takeaways & Limitations

    Positive words and phrases can reduce violent completions to a certain extent, but their side effect of redirecting outputs toward a specific topic means the approach may not be general.

  • Takeaways & Limitations

    The paper’s use of “debias” is loose: reduced violence does not mean the completions are free of all bias.

Abstract

from arXiv · show

It has been observed that large-scale language models capture undesirable societal biases, e.g. relating to race and gender; yet religious bias has been relatively unexplored. We demonstrate that GPT-3, a state-of-the-art contextual language model, captures persistent Muslim-violence bias. We probe GPT-3 in various ways, including prompt completion, analogical reasoning, and story generation, to understand this anti-Muslim bias, demonstrating that it appears consistently and creatively in different uses of the model and that it is severe even compared to biases about other religious groups. For instance, "Muslim" is analogized to "terrorist" in 23% of test cases, while "Jewish" is mapped to "money" in 5% of test cases. We quantify the positive distraction needed to overcome this bias with adversarial text prompts, and find that use of the most positive 6 adjectives reduces violent completions for "Muslims" from 66% to 20%, but which is still higher than for other religious groups.

Results

GPT-3 exhibits persistent Muslim-violence associations across prompt completions, analogies, and image-story generation. Positive adjective prompts reduce violent completions, but the strongest intervention leaves violence higher than for other religious groups.

  • Prompt completions: GPT-3’s Muslim-violence association varied the weapons, nature, and setting of violence rather than reproducing a small set of memorized headlines.
  • Image stories: Image-story descriptions were typically humorous or poignant, but became violent when captions included Muslim religious attire such as “headscarf.”For some images of Muslim women, the resulting output was also less coherent.
  • Debiasing completions: Positive adjective prompts reduced violence-containing completions from 66% to 46% across 50 adjectives, and to 20% using the six best-performing adjectives.The 20% rate remained higher than the analogous 13–15% rate for “Christians.”
  • Debiasing completions: The most effective adjectives were “hard-working” and “luxurious,” which redirected completions toward a specific topic, whereas “calm” did not significantly change violent-completion rates.

Discussion

GPT-3 captures strong, persistent associations between Muslims and violence across varied language uses. These biases emerge creatively rather than as simple memorization, complicating detection and mitigation, although positive contextual interventions reduce them to a certain extent.

  • GPT-3 captures strong negative and creative stereotypes regarding Muslims across different uses of language.
  • Associations between Muslims and violence are learned during pretraining but do not seem to be memorized.
  • GPT-3 mutates biases in different ways, which may make them more difficult to detect and mitigate.
  • Positive words and phrases in the context reduce GPT-3’s biased completions to a certain extent.
  • The interventions were carried out manually, and redirecting completions toward a specific topic may limit their generality.

A. GPT-3 Parameters

All experiments use the default settings of OpenAI’s davinci GPT-3 engine.

  • All experiments use the default settings of OpenAI’s davinci GPT-3 engine.

B. Violence-Related Keywords

The study operationalizes violent completions using manually identified keywords and phrases found in sampled GPT-3 outputs.

  • A completion is considered violent if it includes specified keywords or phrases, in part or whole.
  • The keyword list was compiled by manually reviewing 100 random GPT-3 completions.

C. Full Results with Analogies

The analogy experiments expand the religious-group comparison by including demonyms and adding Hindus and Catholics.

  • The original analogy experiments used six religious groups and excluded outputs that produced demonyms.
  • The rerun includes demonyms and extends the experiments to Hindus and Catholics.
  • Figure 5 is identified, but the supplied text does not specify its axes, layout, or result.

D. Further HONY Examples

Figures 6 and 7 provide additional Humans of New York-style descriptions generated by GPT-3, including neutral descriptions and descriptions showing anti-Muslim bias.

  • Additional Humans of New York-style descriptions are provided in Figures 6–7.
  • Figure 6 presents neutral descriptions generated by GPT-3.
  • Figure 7 presents GPT-3 descriptions showing anti-Muslim bias.

E. Debiasing Examples

Adding positive descriptions of Muslims can reduce violent completions, but the trigger also steers outputs toward specific topics and does not provide a general solution.

  • The proportion of completions containing violent language was reduced using positive descriptions of Muslims.
  • The positive trigger redirects completions toward a specific direction rather than removing all bias.
  • The trigger “Muslims are luxurious” often steers completions toward financial or materialistic matters.
  • Examples include restaurant, bank, bar, and workplace scenarios generated after the “Muslims are luxurious” trigger.
  • The examples include violent or threatening details, including robbery, guns, and references to death.
Loading 2101.05783v2…