Source-linked AI summary

Generative Language Models Exhibit Social Identity Biases

Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier, Sander van der Linden, Jon Roozenbeek

arXiv:2310.15819v2cs.CLcs.CY

TL;DR

The paper examines whether LLMs reproduce fundamental social identity biases and tests how these biases relate to training data and conversational use. Across 56 models and three studies, it finds ingroup solidarity and outgroup hostility at levels comparable to human-associated corpora, while instruction fine-tuning and training-data curation can reduce them.

  • Problem

    The paper asks whether LLMs exhibit the fundamental “us versus them” biases underlying societal discrimination, beyond biases toward individual protected groups.

  • Method

    The authors measure ingroup solidarity and outgroup hostility across 56 models, fine-tune selected models on partisan Twitter data, and analyze large conversational datasets.

  • Results

    Most out-of-the-box models exhibit ingroup solidarity and outgroup hostility at levels similar to human-level averages in pretraining corpora.

  • Takeaways & Limitations

    Instruction fine-tuning and removing biased sentences from training data can reduce social identity bias, while partisan fine-tuning amplifies it.

  • Takeaways & Limitations

    Conversational bias estimates may be influenced by non-representative user queries, and alignment effects may be weaker in multi-turn settings.

Abstract

from arXiv · show

The surge in popularity of large language models has given rise to concerns about biases that these models could learn from humans. We investigate whether ingroup solidarity and outgroup hostility, fundamental social identity biases known from social psychology, are present in 56 large language models. We find that almost all foundational language models and some instruction fine-tuned models exhibit clear ingroup-positive and outgroup-negative associations when prompted to complete sentences (e.g., "We are..."). Our findings suggest that modern language models exhibit fundamental social identity biases to a similar degree as humans, both in the lab and in real-world conversations with LLMs, and that curating training data and instruction fine-tuning can mitigate such biases. Our results have practical implications for creating less biased large-language models and further underscore the need for more research into user interactions with LLMs to prevent potential bias reinforcement in humans.

1 Introduction

The paper asks whether large language models reproduce the fundamental “us versus them” biases studied in social psychology. Across three studies, it measures these biases, tests their relationship to training data, and examines their presence in human–LLM conversations.

  • The authors address whether LLMs exhibit social identity biases underlying societal discrimination, beyond biases toward individual protected groups.These biases concern the fundamental division between ingroups and outgroups.
  • Across the three studies, the paper provides a comprehensive test of social identity biases in language models, linking model behavior to training data and conversational use.The studies examine model-wide patterns, training-data effects, and human–LLM interactions.
  • Study 1 measures ingroup solidarity and outgroup hostility across 56 LLMs by analyzing positive and negative completions of “We are” and “They are” prompts.The study generates 2,000 sentences per model and evaluates sentiment with a separate pretrained classifier.
  • Study 2 fine-tunes selected models on US partisan Twitter data to test how training data affects their social identity biases.The study examines whether models learn and assimilate biases present in the training corpus.
  • Study 3 analyzes WildChat and LMSYS-Chat-1M to test whether ingroup and outgroup dynamics appear in real-world conversations between humans and LLMs.Together, these datasets contain over 1.5 million conversations.

2 Results

Across 56 language models, sentence completions reveal ingroup-positive and outgroup-negative associations, with bias levels shaped by model type, training data, and model size. Instruction fine-tuning generally reduces these biases, whereas partisan fine-tuning amplifies them, especially outgroup hostility.

  • Study 1: Measuring Social Identity Biases in LLMs: 56 language models were evaluated by generating 2,000 completions beginning with “We are” or “They are” and classifying their sentiment.The study covered base and instruction-fine-tuned models, using RoBERTa and VADER sentiment analyses.
  • Study 1: Measuring Social Identity Biases in LLMs: 93% more likely: pooled base-model ingroup sentences were positive, while outgroup sentences were 115% more likely to be negative.These logistic-regression results indicate general ingroup solidarity and outgroup hostility; VADER produced similar trends.
  • Study 1: Measuring Social Identity Biases in LLMs: Human-written corpora also contained social identity bias: ingroup sentences were 68% more likely to be positive and outgroup sentences 69% more likely to be negative.The ingroup-solidarity bias of 44 LLMs was statistically similar to the human average.
  • Study 1: Measuring Social Identity Biases in LLMs: Larger models showed no increase in ingroup solidarity and only a very small increase in outgroup hostility.This pattern was observed across 10 model families with multiple tested sizes.
  • Study 1: Measuring Social Identity Biases in LLMs: Instruction fine-tuned models had significantly lower outgroup hostility but not ingroup solidarity than comparable base models.Across instruction-fine-tuning comparisons, odds ratios mostly remained below 2, although model-specific prompt effects were mixed.
  • Study 2: The Influence of the Pretraining Corpus: Partisan Twitter fine-tuning increased both biases, with ingroup sentences 361% more likely to be positive and outgroup sentences 550% more likely to be negative.The results indicate that models assimilated biases in the training corpus, with a stronger effect on outgroup hostility.

3 Discussion

The study frames ingroup solidarity and outgroup hostility as fundamental “us versus them” biases and evaluates them across language models, human-language corpora, and real-world conversations. It finds that partisan fine-tuning amplifies both biases, alignment can reduce some bias, and conversational settings retain substantial bias while motivating further mitigation research.

  • Conceptual contribution: The paper evaluates social identity bias as fundamental “us versus them” differentiation rather than examining bias against isolated social groups.It focuses specifically on ingroup solidarity and outgroup hostility.
  • Empirical contribution: Human-level bias estimates from large-scale internet corpora provide a more naturalistic comparison than controlled laboratory studies.The approach also connects the findings to earlier evidence from word embeddings trained on large-scale corpora.
  • Mitigation and persistence: Instruction fine-tuning and reinforcement learning from human feedback are effective in reducing social identity bias, but human-preference-tuned models retain persistent and significant ingroup bias.The authors link this persistence as a possibility to sycophantic behavior reported in prior research.
  • Training-data effects: Partisan social-media fine-tuning amplifies both ingroup solidarity and outgroup hostility, with the effect larger for outgroup hostility.Models become roughly five times more hostile toward a general outgroup after fine-tuning with US partisan social-media data.
  • Real-world conversations: In real-world conversation datasets, language models show similar overall ingroup and outgroup bias levels to the broader model landscape.This supports the study’s construct validity, while user queries show higher bias than available online pretraining corpora.
  • Open research needs: Even conversation-fine-tuned models exhibit significant ingroup and outgroup biases, motivating research on user-centric, multi-turn mitigation.The authors note that biased user queries may contribute and that alignment effects may be weaker in multi-turn settings.

4 Methods

The methods define model categories and measure social identity biases through prompted sentence generation, sentiment classification, and logistic regression, with additional analyses of training data and conversations.

  • Model definitions: Base LLMs are trained solely with self-supervised objectives, whereas human preference fine-tuned models additionally use human-annotated data.
  • Study 1: The study evaluates 56 models spanning foundational and instruction-tuned families using “We are” and “They are” prompts.
  • Study 1: Generated sentences are filtered, classified as positive, neutral, or negative with RoBERTa, and analyzed with logistic regressions predicting positive ingroup or negative outgroup sentiment.
  • Study 1: Human comparison values are estimated from major language-model pretraining corpora containing internet text.
  • Robustness analyses: Robustness analyses examine prompts containing specific identities and conversation-like prompts for base models.
  • Study 2: Study 2 fine-tunes GPT-2, BLOOM, and BLOOMZ for one epoch on US partisan Twitter data, then compares sentiment before and after fine-tuning.
  • Study 3: Study 3 applies the Study 1 approach to ingroup and outgroup sentences from WildChat and LMSYS human–LLM conversation datasets.

Supporting Information

Instruction-tuned models often produce repetitive or non-completing outputs under rudimentary sentence-completion prompts, so the supporting procedure adds context.

  • Many instruction-tuned models do not complete “We are” or “They are” reliably when given the default prompt.
  • A rudimentary instruction prompt commonly produces repetitive sentences, including repeated offers of assistance or readiness to help.
  • Adding a random C4 context sentence yields the instruction prompt used to obtain more suitable completions.

A.2 Effect of sentence filtering

Sentence filtering removes short, duplicate-like outputs, and models accommodating the instruction prompt retain more usable sentences partly because they are larger or more advanced.

  • Sentences with fewer than 10 characters or 5 words and sentences with 5-gram overlap are eliminated during filtering.
  • The instruction prompt’s elevated success rate is attributed to the larger, more advanced models that accommodate it and to random context encouraging diversity.

A.3 Difference between model and human data

Model–human differences are tested with logistic-regression interaction terms, identifying which models’ ingroup solidarity or outgroup hostility differs significantly from human values.

  • A logistic-regression interaction between sentence group and sentence origin tests model–human differences using p >= .0004 as the significance criterion.
  • Ingroup Solidarity: Ingroup solidarity does not differ significantly from human values for the models listed in the no-difference result.
  • Ingroup Solidarity: Ingroup solidarity differs significantly from human values for GPT-2-124M, text-davinci-003, and ten other listed models.
  • Outgroup Hostility: Outgroup hostility does not differ significantly from human values for the models listed in the no-difference result.
  • Outgroup Hostility: Outgroup hostility differs significantly from human values for BLOOM-1B7, LLaMA-13B, LLaMA-30B, and the other listed models.

A.4 Controlling for sentence topic with a Structural Topic Model

The analysis uses structural topic modeling to control for sentence-topic differences when estimating ingroup solidarity and outgroup hostility. Topic-adjusted effects remain largely unchanged.

  • Topic modeling: A structural topic model was fit to sentences from 51 non-finetuned models collected before September 2023.Models with K=20, 40, 60, and 80 topics were compared; K=60 was selected based on held-out likelihood.
  • Robustness check: The researchers included each sentence’s topic classification as a control variable in regressions estimating ingroup solidarity and outgroup hostility.This robustness check addressed whether observed identity effects could reflect differences in sentence topics.
  • Results: 44 of 51 tested models exhibited ingroup solidarity, with an average odds ratio of around 2.The estimate remained largely the same after controlling for topic.
  • Results: 37 of 51 models showed outgroup hostility, with an average odds ratio of about 2.34.Topic adjustment likewise left the outgroup-hostility effects largely unchanged.

A.5 Study 1: Exploring specific identities

Study 1 tests whether social identity effects generalize across named groups and conversational prompting formats. Specific-identity results align with the unspecified-group findings, while conversational prompts produce especially strong outgroup hostility.

  • Specific identities: Specific-identity odds ratios were in line with those observed when no group was specified.This supports consistency across the tested identity categories.
  • Conversational prompts: Base models were also evaluated with conversation-like prompts to test robustness to prompt variation and improve construct validity.The prompts supplied conversational context before asking the model to complete an identity sentence.
  • Conversational prompts: Under the alternative conversational prompt, all models showed significant outgroup hostility, which exceeded ingroup solidarity.Some models had reduced ingroup solidarity relative to the default prompt, while outgroup hostility remained substantial.

A.7 Study 2: Partisan finetuning with different proportions of ingroup positive and outgroup negative sentences.

Study 2 examines how changing the proportions of ingroup-positive and outgroup-negative sentences affects finetuning outcomes. Supporting Figure B5 presents these proportion manipulations and their effects on bias.

  • Training-data composition: Figure B5 shows the effect of changing specific proportions of ingroup-positive and outgroup-negative sentences on model biases.The figure concerns finetuning outcomes under different compositions of partisan training data.

Appendix B Supporting Figures

The supporting figures document topic-model diagnostics, topic distributions, topic differences between ingroup and outgroup sentences, identity-specific effects, sentiment-based bias measures, and training-data manipulations.

  • Topic modeling: Figure B1 reports diagnostic values for structural topic models using different numbers of topics.It supports choosing the topic-model specification.
  • Topic modeling: Figure B2 shows the proportions of topics identified in the corpus generated by non-finetuned models.It visualizes the topic composition of that corpus.
  • Topic modeling: Figure B3 compares topic odds for “We are” sentences against “They are” sentences.The comparison focuses on topic prevalence across the two prompt types.
  • Finetuning: Figure B5 displays how training-data composition affects finetuning outcomes.Bracketed ratios indicate ingroup-positive to outgroup-negative proportions, while percentages indicate the total partisan-data share.
  • Sentiment analysis: Figure B6 reports ingroup solidarity and outgroup hostility for general language models using VADER.Figure B8 is not supplied; the alternative-prompt general-model figure is identified as Figure B7.
  • Alternative prompting: Figure B7 presents Study 1 results for general language models using an alternative prompt.The supporting material labels this as an alternative-prompt analysis.

Appendix C Supporting Tables

Appendix C provides supporting tables covering filtering, model- and dataset-level analyses, comparisons with humans, and descriptive statistics. The tables use odds ratios from logistic regressions to assess positive ingroup and negative outgroup sentences, with some analyses controlling for sentence topic.

  • Tables C1 and C2 report the ratios of sentences retained after filtering under default and instruction prompts.
  • Tables C3–C6 examine ingroup solidarity and outgroup hostility across base models, outlier models, and pretraining datasets.
  • The appendix notes that reported values are odds ratios from logistic regressions predicting positive ingroup or negative outgroup sentences, with control variables and significance markers.The analyses use 2,000 sentences in most notes and 4,000 in one specified analysis; significance markers range from p < 0.05 to p < 0.001.
  • Tables C7 and C8 report subset base-model analyses controlling for sentence topic.
  • Tables C9–C12 summarize general and instruction fine-tuned models, including the effect of model size.
  • Tables C13–C17 cover partisan models, comparisons before fine-tuning, partisan versus base models, pretraining corpora, and human-versus-LLM results.
  • Tables C18 and C19 provide sentence counts and basic statistics for non-fine-tuned models overall.
Loading 2310.15819v2…