Source-linked AI summary
Bias of AI-Generated Content: An Examination of News Produced by Large Language Models
Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, Xiaohang Zhao
TL;DR
The paper addresses whether AIGC produced by LLMs exhibits gender and racial bias, and evaluates this question using news-based prompts and comparisons with reference articles. Across seven LLMs, the study finds substantial bias and discrimination against females and Black individuals, while ChatGPT shows the lowest bias and uniquely declines some biased requests.
Problem
Evidence is limited on how LLM-generated content manifests gender and racial bias across words, sentences, and documents, despite the need to understand LLM limitations.
Method
The study prompts seven LLMs with headlines from 8,629 New York Times and Reuters articles, compares generated content with the originals, and tests gender bias under biased prompts.
Results
AIGC from every examined LLM shows substantial gender and racial bias, including notable discrimination against females and Black individuals; ChatGPT has the lowest bias in most experiments and uniquely declines biased prompts.
Takeaways & Limitations
LLMs should be employed with caution because their generated content contains considerable gender and racial biases and can reflect harmful user prompts.
Takeaways & Limitations
The study primarily uses predefined news-headline prompts and relies on a single automated sentiment-detection model at the sentence level.
Abstract
from arXiv · showhide
Large language models (LLMs) have the potential to transform our lives and work through the content they generate, known as AI-Generated Content (AIGC). To harness this transformation, we need to understand the limitations of LLMs. Here, we investigate the bias of AIGC produced by seven representative LLMs, including ChatGPT and LLaMA. We collect news articles from The New York Times and Reuters, both known for their dedication to provide unbiased news. We then apply each examined LLM to generate news content with headlines of these news articles as prompts, and evaluate the gender and racial biases of the AIGC produced by the LLM by comparing the AIGC and the original news articles. We further analyze the gender bias of each LLM under biased prompts by adding gender-biased messages to prompts constructed from these news headlines. Our study reveals that the AIGC produced by each examined LLM demonstrates substantial gender and racial biases. Moreover, the AIGC generated by each LLM exhibits notable discrimination against females and individuals of the Black race. Among the LLMs, the AIGC generated by ChatGPT demonstrates the lowest level of bias, and ChatGPT is the sole model capable of declining content generation when provided with biased prompts.
Introduction
The paper frames AIGC bias as systematic discrimination in generated content and evaluates gender and racial bias by comparing LLM outputs with reference news across multiple textual levels. It also extends evaluation to biased prompts, where users may induce biased content and models may refuse generation.
- Motivation: LLMs generate AIGC from user prompts, offering efficient, low-cost content production but requiring evaluation of their limitations.The paper motivates this need by noting LLMs’ potential applications across organizational work.
- Bias definition: AIGC is biased when it systematically and unfairly discriminates against particular population groups, especially underrepresented groups.
- Evaluation framework: The study evaluates gender and racial bias at word, sentence, and document levels by comparing generated content with reference news articles.These levels examine word distributions, sentence sentiment and toxicity, and document-level semantics.
- Evaluation framework: The framework uses news headlines as prompts, generates AIGC with an LLM, and compares the outputs with articles from The New York Times and Reuters.
- Biased prompts: Biased prompts are added to headline-based prompts to assess gender bias and each LLM’s resistance to generating induced biased content.The evaluation also considers whether a model refuses content generation under biased prompts.
- Relation to prior work: Unlike prior work focused mainly on short phrases, this evaluation examines bias across words, sentences, and complete documents.
Word Level Bias
Across the examined LLMs, generated news shows substantial word-level gender and racial bias relative to New York Times and Reuters articles. ChatGPT has the lowest overall gender and racial bias scores, but all models show notable disadvantages for females and Black individuals.
- Gender Bias: 0.1536 is ChatGPT’s lowest word-level gender-bias score among the examined LLMs, measured by average Wasserstein distance.The score represents a 15.36% average absolute difference in gender-specific word proportions from the reference articles.
- Gender Bias: 56.04%–73.89% of generated articles show female prejudice, defined as a lower percentage of female-specific words than the counterpart reference article.Grover has the highest proportion at 73.89%, while GPT-3-curie has the lowest at 56.04%.
- Racial Bias: 0.2331 is ChatGPT’s lowest word-level racial-bias score, while every examined LLM exhibits notable racial bias.The score corresponds to a 23.31% average absolute difference in race-related word proportions from the reference articles.
- Racial Bias: Every examined LLM shows significant word-level bias against the Black race, whereas only Grover and GPT-2 show significant bias against the Asian race.For Grover, White-race words increase by 20.07%, while Black- and Asian-race words decrease by 11.74% and 8.34%, respectively.
- Racial Bias: 60.94%–81.30% of generated articles show Black prejudice, and Black-specific words decrease by 30.39%–48.64% in those articles.ChatGPT has the lowest reported reduction, at −30.39%, while Grover has the highest proportion of Black-prejudice articles and the largest reduction.
Sentence Level Bias
At the sentence level, the examined LLMs show substantial gender and racial sentiment bias relative to New York Times and Reuters articles. ChatGPT generally produces the lowest prevalence of female and Black prejudice in sentiment, while toxicity results are qualitatively similar.
- Gender Bias on Sentiment: 0.1396 was Cohere’s lowest sentence-level gender sentiment-bias score, while every examined LLM exhibited substantial gender bias on sentiment.Scores measure maximal absolute sentiment differences between population groups in generated articles and their news counterparts.
- Gender Bias on Sentiment: 39.50% was ChatGPT’s proportion of female-prejudice articles based on sentiment, the lowest among the examined LLMs.Female prejudice means that sentiment toward females was lower in generated articles than in their New York Times or Reuters counterparts.
- Racial Bias on Sentiment: 0.1348 was Cohere’s lowest sentence-level racial sentiment-bias score, while every examined LLM exhibited degraded racial sentiment alignment.The score compares sentiment differences across White, Black, and Asian groups with corresponding news articles.
- Racial Bias on Sentiment: 39.22% was ChatGPT’s proportion of Black-prejudice articles based on sentiment, the lowest among the examined LLMs.Black prejudice means that sentiment toward Black people was lower in generated articles than in their news counterparts.
- Sentence-Level Toxicity: Sentence-level gender and racial toxicity results were qualitatively similar to the sentiment-bias findings.The toxicity analyses are reported in Appendix A.
Document Level Bias
At the document level, generated news exhibits substantial gender and racial bias relative to New York Times and Reuters counterparts. ChatGPT generally has the lowest prevalence and reduction of female- and Black-related topics among the examined models.
- Gender Bias: 0.2377 was Grover’s example document-level gender-bias score, and every examined LLM exhibited substantial document-level gender bias.The score measures absolute differences in male or female topic percentages between generated articles and their news counterparts.
- Gender Bias: 25.86% was ChatGPT’s proportion of document-level female-prejudice articles, lower than the corresponding proportions for the other examined LLMs.Female prejudice means that generated articles contain a lower percentage of female pertinent topics than their counterparts.
- Gender Bias: ChatGPT showed the smallest reduction of female pertinent topics among the examined models, consistent with the broader gender-bias findings.The study also reports that RLHF is beneficial for reducing document-level bias against females.
- Racial Bias: 0.2815 was Cohere’s lowest document-level racial-bias score, while every examined LLM exhibited significant document-level racial bias.The score measures differences in topic percentages for White, Black, or Asian groups between generated articles and counterparts.
- Racial Bias: 24.23% was ChatGPT’s proportion of document-level Black-prejudice articles, and ChatGPT consistently performed best on this measure and Black-topic reduction.Black prejudice means that generated articles contain a lower percentage of Black-race pertinent topics than their counterparts.
Bias of AIGC under Biased Prompts
Biased prompts were compared with unbiased prompts at word, sentence, and document levels, revealing altered gender-bias patterns and selective refusal behavior. ChatGPT alone refused biased prompts, while its surviving outputs showed a significant reduction in female-specific words.
- Resistance to biased prompts: 89.13% of gender-biased requests were declined by ChatGPT, the only examined LLM reported to refuse content generation under biased prompts.Its outputs under biased prompts therefore came from the remaining 10.87% of requests.
- Evaluation design: The study compares biased- and unbiased-prompt outputs using female-prejudice proportions and decreases in female-specific words, sentiment, and female-related topics.The word-level figure reports 95% confidence intervals for its comparisons.
- Word-level comparison: −32.85% was ChatGPT’s decrease in female-specific words under biased prompts, compared with −24.50% under unbiased prompts.This was the only examined LLM with a significant difference between biased- and unbiased-prompt conditions (∆ = −8.35%, p < 0.001).
- Sentence-level comparison: At the sentence level, female prejudice was evaluated through sentiment differences between generated articles and their news-article counterparts.For ChatGPT, the average sentiment score of female-related sentences was −0.1429 under biased prompts.
- Document-level comparison: At the document level, female prejudice was assessed by comparing the percentage of female-related topics in generated articles with their counterpart news articles.Figure 10 compares this document-level measure under unbiased and biased prompts.
Discussion
Across the evaluated models, AIGC showed substantial gender and racial bias, while ChatGPT generally exhibited the lowest bias and uniquely refused many biased prompts. The discussion nevertheless emphasizes caution because screened prompts could still produce highly biased content and the study has several scope limitations.
- Main findings: All examined LLMs produced substantial gender and racial biases across word, sentence, and document levels.The reported deviations concerned word choices, sentiment and toxicity, and conveyed semantics related to population groups.
- Model comparison: ChatGPT exhibited the lowest bias in most experiments and was the only examined model capable of declining biased prompts.The paper attributes part of this advantage to reinforcement learning from human feedback.
- Residual vulnerability: When biased prompts bypassed its screening, ChatGPT produced significantly more biased news articles than the other studied LLMs.The discussion identifies this as a vulnerability that malicious users could exploit.
- Limitations: The study’s prompt choices were pre-defined and primarily used news headlines, limiting the evaluated prompt setting.The authors propose examining prompts that also include journalist names and interactive prompt flows.
- Limitations: Sentence-level results relied on one automated sentiment model, while the study covered limited gender and racial categories.The authors suggest multiple sentiment models and broader minority-group coverage for future research.
Methods
The study evaluates seven representative LLMs by comparing news articles they generate with articles from The New York Times and Reuters. It measures bias at word, sentence, and document levels using population-group distributions, sentiment and toxicity differences, and topic-based semantic proportions.
- Data and models: News articles from The New York Times and Reuters serve as reference articles for evaluating generated content.The study describes both agencies as highly rated for accurate and unbiased news.
- Data and models: The evaluation covers Grover, GPT-2, GPT-3-curie, GPT-3-davinci, ChatGPT, Cohere, and LLaMA-7B.The investigated models include both closed-source systems accessed through APIs and open-source systems run locally.
- Generation procedure: Each LLM generates a news article from a collected news headline, with Grover receiving headlines directly rather than an additional prompt.The other systems use prompts constructed from the headlines.
- Bias measures: Word-level bias compares population-associated word distributions in generated and reference articles using Wasserstein distance.The evaluated population groups include female and male for gender bias and White, Black, and Asian for racial bias.
- Bias measures: Sentence-level bias compares generated and reference articles through population-group differences in sentiment and toxicity, while document-level bias compares topic-associated semantic proportions.Document-level proportions are derived from topics discovered with LDA, and topic associations use standardized residuals exceeding 3.
A.1 Gender Bias on Toxicity
The gender-toxicity analysis compares toxicity toward females and males in generated articles with corresponding articles from The New York Times and Reuters. Every examined model produced female-prejudice articles, and ChatGPT had the smallest reported toxicity increase among them.
- Female prejudice: 29.10% of ChatGPT-generated articles showed female prejudice with respect to toxicity, the lowest proportion among the seven models.The reported proportions were based on articles containing sentences associated with females.
- Female prejudice: 48.29% of Grover-generated articles showed female prejudice with respect to toxicity.This corresponds to N = 1,081 articles.
- Toxicity increase: 0.0225 was ChatGPT’s average toxicity-score increase in female-prejudice articles, the smallest increase reported.The 95% confidence interval was [0.0144, 0.0307] with N = 266.
- Toxicity increase: 0.0536 was Grover’s average toxicity-score increase toward females in female-prejudice articles.The 95% confidence interval was [0.0446, 0.0625] with N = 522.
A.2 Racial Bias on Toxicity
The racial-toxicity analysis compares toxicity associated with White, Black, and Asian populations in generated and reference articles. All examined LLMs showed racial bias, while ChatGPT had the lowest aggregate racial-toxicity bias and 32.91% of its articles showed Black prejudice.
- Aggregate racial bias: 0.0186 was ChatGPT’s racial bias on sentence-level toxicity, the lowest value among the investigated LLMs.The 95% confidence interval was [0.0170, 0.0202] with N = 3,581.
- Aggregate racial bias: All seven investigated LLMs exhibited a certain degree of racial bias on sentence-level toxicity.The reported values ranged from 0.0186 for ChatGPT to 0.0761 for GPT-2.
- Black prejudice: 32.91% of ChatGPT-generated articles showed Black prejudice with respect to toxicity.The proportion was calculated from N = 1,110 articles containing sentences associated with the Black race.
- Black prejudice: 56.50% of GPT-2-generated articles showed Black prejudice with respect to toxicity, the highest proportion reported.The proportion was calculated from N = 962 articles.
B Topic Examples
Tables B.1 and B.2 provide example topics associated with gender and racial population groups, respectively, along with relevant words illustrating each topic’s semantic content.
- Topic 4 links the female population group with a mixed theme involving art and family.
- Topic 51 links the male population group with politics and famous male politicians.
- Topic 3 links the White population group with the international conflict between Russia and Ukraine.
- Topic 25 links the Black population group with racism and cultural diversity.
- Topic 171 links the Asian population group with Asian politics.