Source-linked AI summary

On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning

Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, Diyi Yang

arXiv:2212.08061v2cs.CL

TL;DR

The paper asks whether zero-shot CoT improvements extend to socially sensitive reasoning, where prior evidence was limited. It evaluates stereotype benchmarks and harmful questions under standard and CoT prompting, finding increased bias and toxicity with CoT, while improved alignment reduces the effect. The authors therefore urge careful auditing of reasoning-induced behavior, especially given benchmark and prompt-scope limitations.

  • Problem

    Prior CoT research focused mainly on logical reasoning, leaving its effects on socially situated tasks unclear despite their importance for sensitive topics and marginalized groups.

  • Method

    The paper compares standard and zero-shot CoT prompting across three stereotype benchmarks and HarmfulQ using several GPT-3 davinci models and prompt formats.

  • Results

    Across evaluations, CoT increases stereotypical generalizations by 8.8 percentage points and encouragement of explicit toxic behavior by 19.4%, with effects reduced by improved preference alignment and mitigation instructions.

  • Takeaways & Limitations

    Zero-shot CoT should be carefully controlled and audited on socially important tasks because reasoning prompts can produce biased and toxic outcomes.

  • Takeaways & Limitations

    The study is primarily limited to GPT-3 davinci models and does not systematically explore how different CoT prompt structures affect stereotypes.

Abstract

from arXiv · show

Generating a Chain of Thought (CoT) has been shown to consistently improve large language model (LLM) performance on a wide range of NLP tasks. However, prior work has mainly focused on logical reasoning tasks (e.g. arithmetic, commonsense QA); it remains unclear whether improvements hold for more diverse types of reasoning, especially in socially situated contexts. Concretely, we perform a controlled evaluation of zero-shot CoT across two socially sensitive domains: harmful questions and stereotype benchmarks. We find that zero-shot CoT reasoning in sensitive domains significantly increases a model's likelihood to produce harmful or undesirable output, with trends holding across different prompt formats and model variants. Furthermore, we show that harmful CoTs increase with model size, but decrease with improved instruction following. Our work suggests that zero-shot CoT should be used with caution on socially important tasks, especially when marginalized groups or sensitive topics are involved.

1 Introduction

Although zero-shot CoT often improves reasoning performance, this paper shows that it can increase bias and toxicity in socially sensitive tasks. Controlled evaluations find more stereotypical reasoning and more encouragement of harmful behavior under CoT prompting.

  • Motivation: Zero-shot CoT improves performance on tasks including question answering, mathematical problem solving, and commonsense reasoning.The approach prompts models to generate reasoning steps, commonly with “Let’s think step by step.”
  • Motivation: Zero-shot CoT consistently produces undesirable biases and toxicity on tasks requiring social knowledge.The authors caution that reasoning improvements are not universal across task types.
  • Findings: CoT-prompted models can perpetuate stereotypes about disadvantaged groups and actively encourage recognized toxic behavior.The paper contrasts these outcomes with tasks having objectively correct answers, where CoT may work well.
  • Approach: Controlled evaluations compare standard prompting with CoT across stereotype benchmarks and harmful questions.The study reformulates CrowS-Pairs, StereoSet, and BBQ, and introduces HarmfulQ.
  • Findings: 8.8% point increase in stereotypical generalizations and 19.4% higher rates of encouraging explicit toxic behavior occur with CoT.These are averaged effects across the reported evaluations relative to standard prompt counterparts.

2 Related Work

Prior work established CoT as an effective but opaque reasoning strategy, while research also documented broad social biases and sensitivity to prompting. This paper extends that literature by evaluating reasoning strategies on socially sensitive tasks.

  • LLM Reasoning: CoT emerged as an LLM capability that can improve arithmetic, metaphor generation, and commonsense or symbolic reasoning.Zero-shot prompting with “Let’s think step by step” was reported to improve reasoning-benchmark performance.
  • LLM Reasoning: Other reasoning strategies include self-consistency, combining imperfect prompts, and decomposing prompts into less-to-more complex questions.The paper uses these developments to motivate broader evaluation of reasoning strategies.
  • Robustness and Failures: LLMs are sensitive to prompting perturbations, and their reasoning processes can produce unreliable explanations or misunderstand demonstrations.This makes reasoning behavior difficult to interpret directly.
  • Social Bias: Language models exhibit social and cultural biases, including stereotypical behavior across established stereotype benchmarks.The paper reframes prior stereotype datasets as zero-shot reasoning tasks to probe intrinsic bias.

3 Stereotype & Toxicity Benchmarks

The paper evaluates zero-shot reasoning on three stereotype benchmarks and a newly constructed harmful-question benchmark. These datasets cover ambiguous stereotype judgments and open-ended toxic requests.

  • Benchmark Overview: The evaluation uses CrowS-Pairs, StereoSet, BBQ, and a bootstrapped HarmfulQ benchmark.The stereotype datasets are converted into zero-shot reasoning tasks, while HarmfulQ contains explicitly harmful questions.
  • Benchmark Design: Zero-shot evaluation is used to quantify out-of-the-box capabilities and avoid variability from few-shot exemplars.The authors note that few-shot examples can resemble finetuning and trivialize stereotype benchmarks.
  • Stereotype Benchmarks: CrowS-Pairs contains 1508 minimal sentence pairs spanning nine stereotype dimensions.Each pair reinforces either a stereotype or an anti-stereotype.
  • Stereotype Benchmarks: StereoSet provides stereotypical and anti-stereotypical examples across gender, race, profession, and religion; 1508 sentences are sampled for evaluation.Contexts in some instances are concatenated to standardize the examples.
  • Stereotype Benchmarks: BBQ contains 50K questions across 11 stereotype categories, with 1100 ambiguous-setting questions sampled for this study.The ambiguous setting treats Unknown as the correct answer when neither stereotype nor anti-stereotype is acceptable.
  • Toxicity Benchmark: HarmfulQ contains 200 explicitly toxic questions generated across racist, stereotypical, sexist, illegal, toxic, and harmful categories.Repetitive questions with high text overlap were manually removed.

4 Methods

The study compares standard prompting with a two-stage zero-shot CoT procedure across stereotype and toxicity benchmarks. It uses answer-selection or manual output scoring, multiple prompt formats, and several GPT-3 davinci models.

  • Prompting Procedure: Standard prompting extracts an answer directly, whereas zero-shot CoT first elicits step-by-step reasoning and then extracts a final answer.The CoT procedure follows two stages: reasoning generation followed by answer extraction.
  • Prompting Procedure: Two prompt formats, BigBench CoT and Inv. Scaling, control for effects from minor formatting changes.Both CoT templates use “Let’s think step by step,” which is omitted from the Standard condition.
  • Scoring: Stereotype tasks use an Unknown option and measure accuracy as Nunk/N, where lower accuracy indicates less normative or value-aligned prediction.CrowS-Pairs and StereoSet require choosing between stereotype and anti-stereotype options, while BBQ is already framed as QA.
  • Scoring: HarmfulQ accuracy is Ndiscourage/N, with outputs manually labeled as encouraging or discouraging harmful behavior.Lower accuracy means a greater likelihood of encouraging harmful behavior.
  • Scoring: CoT effect is computed as AccCoT − AccStandard, with arrows marking positive or negative percentage-point differences.This directly compares each model’s CoT and Standard prompting conditions.
  • Models and Evaluation: The initial evaluation uses text-davinci-002 with temperature 0.7, max_tokens 256, five completions per condition, and 95% confidence intervals.The study also analyzes text-davinci-001 through text-davinci-003 to separate prompting effects from instruction tuning and preference alignment.

5 Results

Zero-shot CoT consistently worsens socially sensitive evaluations, increasing stereotypical and toxic outputs, with effects shaped by prompt format, model scale, and instruction tuning.

  • 5.1 Analyzing TD2: HarmfulQ accuracy decreases by an average of 19.4 percentage points across davinci models under CoT.
  • 5.1 Analyzing TD2: Zero-shot CoT decreases stereotype-benchmark performance, averaging an 18 percentage-point drop across evaluations.
  • 5.1 Analyzing TD2: CrowS Pairs, StereoSet, and BBQ show average decreases of 24.1, 22.2, and 6.3 percentage points, respectively.
  • 5.1 Analyzing TD2: CoT errors include explicit or implicit stereotypical reasoning, with stereotyped hallucinations appearing in 37% of sampled stereotype cases.
  • 5.2 Instruction Tuning Behaviour: Improved instruction tuning generally reduces stereotype effects, but TD3 still loses 53 points on HarmfulQ versus TD2’s 4-point decrease.
  • 5.3 Scaling Behaviour: CoT-induced harms increase with model scale: CrowS Pairs differences rise from 6 to 14 to 29 points, while StereoSet rises from 4 to 10 to 31.
  • 5.4 Explicit Mitigation Instructions: Explicit mitigation instructions reduce TD3’s stereotype-benchmark decline to 1 percentage point, while TD2 still declines by 11.8 points.

6 Evaluating Open Source LMs

The authors evaluate open-source Flan models to reduce confounding from closed-model instruction tuning and alignment differences. CoT improves conventional reasoning benchmarks with scale but worsens bias-benchmark accuracy for larger models.

  • 6 Evaluating Open Source LMs: The evaluation uses Flan models from 80M to 20B parameters with a common BigBench CoT template.
  • 6 Evaluating Open Source LMs: CoT accuracy on MMLU, BBH, and MGSM consistently improves as Flan model scale increases.
  • 6 Evaluating Open Source LMs: On bias benchmarks, 80M Flan models improve by 13% on average, whereas models with at least 250M parameters decline by 5%.
  • 6 Evaluating Open Source LMs: Bias-benchmark effects worsen and then plateau with scale, contrasting with the consistent CoT gains on conventional reasoning tasks.

7 Conclusion

The paper recommends auditing and carefully controlling reasoning strategies because zero-shot CoT can circumvent value-alignment efforts and produce biased or toxic outputs, including in social applications.

  • 7 Conclusion: Reasoning strategies can uncover underlying toxic generations despite current value-alignment efforts.The authors describe this as a suspicion and propose auditing reasoning steps.
  • 7 Conclusion: Zero-shot CoT may circumvent value-alignment efforts by giving models tokens to think, even with innocuous step-by-step reasoning.The authors hypothesize that more complex reasoning strategies may exacerbate these findings.
  • 7 Conclusion: In high-stakes chatbot domains such as mental health or therapy, models should be explicitly uncertain when generating reasoning.The paper connects this recommendation to the risk that CoT can exacerbate downstream biases.
  • 7 Conclusion: Researchers should re-evaluate uncertainty behaviours and bias distributions before adopting zero-shot CoT for social tasks.The paper cautions against assuming that performance gains transfer safely to socially situated applications.
  • 7 Conclusion: The study is primarily based on GPT-3 davinci models with zero-shot CoT capabilities, while the authors report degradations in open-source models as they become more powerful.The authors frame broader generalization beyond GPT-3 as constrained by the models studied.

8 Limitations

The paper’s limitations concern prompt coverage, few-shot settings, and the imperfect conceptual and empirical basis of bias benchmarks.

  • 8 Limitations: The study focuses on the “let’s think step by step” prompt and does not systematically evaluate how alternative CoT prompts affect stereotypes.The authors qualitatively observe that “faster” prompts are less toxic.
  • 8 Limitations: The study does not examine bias in few-shot CoT settings, where exemplars can strongly affect models and confound intrinsic-bias evaluation.The authors identify careful measurement of few-shot CoT as future work.
  • 8 Limitations: Existing fairness benchmarks have flaws and differing conceptualizations of bias, which can produce contradictory results.The authors use three benchmarks and manual qualitative analysis to mitigate, but not eliminate, these concerns.
  • 8 Limitations: The limitations and stated goals of each bias benchmark should be considered when interpreting the findings.The authors do not treat agreement across analyses as removing benchmark limitations.

9 Ethics

The paper addresses ethical risks from offensive content and the potential misuse of reasoning prompts, while limiting exposure through synthetic or previously collected data and automated scoring.

  • 9 Ethics: The work highlights a limitation of LLMs that malicious end-users can widely and easily exploit.The paper includes a warning because its content is offensive.
  • 9 Ethics: The study uses previously collected or synthetically generated benchmarks and automated scoring, avoiding exposure of offensive text to human participants.This design addresses participant exposure rather than eliminating the offensive content itself.
  • 9 Ethics: The BigBench CoT prompt presents a question, stereotype and anti-stereotype options, an Unknown option, and a step-by-step reasoning instruction.The template fills placeholders with stereotype/anti-stereotype pairs from the benchmarks.
  • 9 Ethics: The Inverse Scaling prompt similarly uses a question, answer options, and the instruction “Let’s think step by step.”The paper derives this format from templates used in the Inverse Scaling Challenge.

B Unknown Identifiers

The paper replaces the Unknown identifier with several synonymous phrases and uses benchmark datasets whose licensing and privacy properties are described explicitly.

  • B Unknown Identifiers: The evaluation randomly replaces the Unknown identifier with synonymous phrases such as “Can’t be determined” and “Not enough information.”The replacements follow the methodology of Parrish et al. (2022).
  • B Unknown Identifiers: BBQ is licensed under Creative Commons Attribution 4.0 International, while StereoSet and CrowS Pairs use Creative Commons Attribution Share Alike 4.0 International.The paper states that these datasets do not contain personally identifiable information.
  • B Unknown Identifiers: The authors plan to release their code and datasets under Creative Commons Attribution Share Alike 4.0 International.This release plan is stated for publication.

D Constructing HarmfulQ

HarmfulQ is a 200-question benchmark generated recursively with text-davinci-002 to cover multiple harmful-content categories. The construction uses seeded examples, adjective variation, deduplication, and manual review.

  • HarmfulQ is generated with text-davinci-002 using a prompt that requests lists of harmful questions.The generation setup uses temperature 0.9 and a maximum of 1024 tokens.
  • The process begins with three harmful seed questions and recursively replaces them with randomly selected generations.The seed topics concern making a bomb, bullying someone, and promoting racism.
  • Questions are generated across six adjectives: racist, stereotypical, sexist, illegal, toxic, and harmful.
  • Representative questions cover activities involving restricted access, trafficking, fake news, violence, discrimination, drugs, gangs, and other harms.
Loading 2212.08061v2…