Source-linked AI summary

Social Biases through the Text-to-Image Generation Lens

Ranjita Naik, Besmira Nushi

arXiv:2304.06034v1cs.CYcs.AIcs.CLcs.CV

TL;DR

Text-to-image models trained on massive web data may reproduce harmful social biases, but evidence across multiple bias dimensions and usage contexts remains limited. The paper evaluates DALLE-v2 and Stable Diffusion with automated and human methods across occupations, personality traits, and everyday situations. Both models show severe and model-specific biases; prompt expansion can diversify outputs but may leave quality discrepancies and other stereotypes unresolved.

  • Problem

    Web-trained text-to-image models may reproduce harmful social biases, motivating systematic evidence about representations across occupations, traits, situations, and social dimensions.

  • Method

    The study evaluates DALLE-v2 and Stable Diffusion across gender, race, age, and geography using automated and human evaluations of occupational, personality-trait, person, and everyday-situation prompts.

  • Results

    Both models exhibit major biases, with different representation patterns; prompt expansion can diversify outputs, while default everyday-situation generations are closest to the United States, Australia, and Germany.

  • Takeaways & Limitations

    Prompt specification can help diversify generated representations, but representational mitigation should be considered alongside image-quality discrepancies and other stereotype risks.

  • Takeaways & Limitations

    The study does not provide deep insight into non-binary gender definitions, individuals with disabilities, smaller countries, or religious groups.

Abstract

from arXiv · show

Text-to-Image (T2I) generation is enabling new applications that support creators, designers, and general end users of productivity software by generating illustrative content with high photorealism starting from a given descriptive text as a prompt. Such models are however trained on massive amounts of web data, which surfaces the peril of potential harmful biases that may leak in the generation process itself. In this paper, we take a multi-dimensional approach to studying and quantifying common social biases as reflected in the generated images, by focusing on how occupations, personality traits, and everyday situations are depicted across representations of (perceived) gender, age, race, and geographical location. Through an extensive set of both automated and human evaluation experiments we present findings for two popular T2I models: DALLE-v2 and Stable Diffusion. Our results reveal that there exist severe occupational biases of neutral prompts majorly excluding groups of people from results for both models. Such biases can get mitigated by increasing the amount of specification in the prompt itself, although the prompting mitigation will not address discrepancies in image quality or other usages of the model or its representations in other scenarios. Further, we observe personality traits being associated with only a limited set of people at the intersection of race, gender, and age. Finally, an analysis of geographical location representations on everyday situations (e.g., park, food, weddings) shows that for most situations, images generated through default location-neutral prompts are closer and more similar to images generated for locations of United States and Germany.

1 INTRODUCTION

The paper examines how web-trained text-to-image models reproduce social biases across occupations, personality traits, and everyday situations. It evaluates DALLE-v2 and Stable Diffusion across gender, race, age, and geography, finding major biases and uneven effects of prompt expansion.

  • Motivation: Web-trained text-to-image models may reproduce biases present in non-curated web data, potentially resurfacing harmful stereotypes in generated content.The paper frames this as a concern for representational fairness and prior mitigation efforts.
  • Findings: Image generation models show major setbacks in occupational representational fairness compared with U.S. Bureau of Labor Statistics data for five example occupations.The introduction contrasts real-world, search-engine, and image-generation distributions.
  • Approach: The study evaluates DALLE-v2 and Stable Diffusion across gender, race, age, and geographical location using occupations, personality traits, everyday situations, and person prompts.Occupational and personality-trait prompts receive automated and crowdsourced human evaluation, while everyday situations use default and location-specific prompts.
  • Findings: Both models exhibit major biases, but DALLE-v2 tends toward white younger men whereas Stable Diffusion tends toward white women and more balanced age representation.DALLE-v2 also shows more extreme cases with almost no representation from a given gender or race.
  • Mitigation and geography: Prompt expansion can diversify generated representations, but it does not consistently mitigate bias and can create discrepancies in image quality.The paper also reports location differences: Nigeria, Ethiopia, India for Stable Diffusion, and Papua New Guinea and Colombia are farthest from default generations, while the USA, Australia, and Germany are closest.

2 RELATED WORK

Prior work documents gender, racial, and stereotype amplification in image search and text-to-image systems. This study extends that literature by evaluating two models across more topics and social dimensions with both automated and human methods.

  • Image search: Image-search studies found slight gender under-representation of women and modest amplification of occupational stereotypes relative to U.S. labor statistics.The prior work also examined how men and women were depicted in retrieved images.
  • Text-to-image bias: Prior text-to-image research measured gender and skin-tone skew in neutral occupation prompts using automated and human inspection.One cited finding reports that Stable Diffusion was more prone than minDALL-E to produce a particular gender or skin tone.
  • Text-to-image bias: Stable Diffusion has been reported to perpetuate racial, ethnic, gendered, class, and intersectional stereotypes from neutral prompts, including negative associations such as poverty and subordination.The cited work also observed stereotype amplification and complex stereotypes from prompts naming social groups.
  • Study scope: This study extends prior work by examining two models across people, occupations, traits, and everyday life, covering gender, race, age, and geography.It combines human and automated evaluation methods.
  • Study scope: The paper additionally quantifies how prompt crafting affects occupational representations, an aspect not previously examined beyond example-based evidence.Table 2 is used to contrast the study with recent related work.

3 METHODOLOGY

The study evaluates perceived social representations in DALLE-v2 and Stable Diffusion across multiple prompt types and demographic dimensions. It combines automated detection with human annotation and compares neutral, expanded, and location-specific prompts.

  • 3.1 Social Bias Dimensions: The analysis uses discrete perceived attributes for depicted people, while acknowledging that real-world gender and race can be continuous, socially constructed, and intersectional.
  • 3.2 Social Bias Subjects: DALLE-v2 uses “a portrait of a [prompt]” and Stable Diffusion uses “a photo of a [prompt]” because these prefixes produced higher-quality images for the respective models.
  • 3.2 Social Bias Subjects: Expanded occupation prompts explicitly specify gender, race, or age to assess prompt engineering as a bias-mitigation strategy.
  • 3.2 Social Bias Subjects: The study evaluates person, occupation, personality-trait, and everyday-situation prompts across gender, race, age, and geographical location.
  • 3.2 Social Bias Subjects: The study uses 50 human-containing images for person, occupation, and trait prompts, increasing to 250 images for everyday situations and explicitly specified occupation prompts.
  • 3.3 Data Annotation: Three Mechanical Turk workers annotate race, gender, and age; labels require agreement from at least two workers, while unanimous disagreement produces an excluded “unclear” label.

4 RESULTS

The results show substantial demographic imbalances in both models, with different gender and age patterns. Both models overrepresent white people, while each omits some racial groups.

  • 4.1 What does a person look like in T2I generation?: 70% of DALLE-v2’s “person” images depict males versus 30% females, while Stable Diffusion depicts 66% females versus 28% males.
  • 4.1 What does a person look like in T2I generation?: At least 70% of images from both models depict white individuals.
  • 4.1 What does a person look like in T2I generation?: DALLE-v2 omits East Asian, Southeast Asian, and Middle Eastern individuals, whereas Stable Diffusion omits Latino and Middle Eastern individuals.
  • 4.1 What does a person look like in T2I generation?: DALLE-v2 depicts adults aged 18–40 in 76% of images, while Stable Diffusion provides more varied age representation.

4.2 Representational bias for occupations

The study finds substantial demographic and occupational representational bias in DALLE-v2 and Stable Diffusion outputs, including strong departures from labor-statistics baselines. Prompt expansion can reduce some demographic mismatches, but may introduce residual bias and image-quality discrepancies.

  • Neutral occupations: After filtering ambiguous or non-human outputs, 44 occupations remained for analysis against labor-statistics and image-search baselines.Excluded cases included prompts producing too few people, obstructed faces, or unclear caricatures.
  • Gender: Only 8 of 43 DALLE-v2 occupations and 7 of 43 Stable Diffusion occupations matched labor-statistics female proportions within ±5%.Both models reinforced occupational under- and over-representation, though the affected occupations differed.
  • Age: Occupation outputs showed distinct age concentrations, with many DALLE-v2 and Stable Diffusion occupations dominated by either 18–40 or 40–60-year-old individuals.DALLE-v2 included strong 18–40 representation in administrative assistant and nurse, while CEO and truck driver skewed 40–60; Stable Diffusion showed analogous occupation-specific concentrations.
  • Expanded prompts: Expanded prompts can produce the requested demographic content while retaining significant gender-related image-quality discrepancies.The study evaluates quality using FID, for which lower values indicate better-quality images.

4.3 Representational bias for personality traits

Personality-trait prompts generated systematic associations across gender, race, and age, with different demographic groups linked to distinct trait categories. Stable Diffusion was excluded because more than half of its outputs for some traits were non-human.

  • Evaluation scope: Stable Diffusion was excluded from personality-trait analysis because it generated non-human images for over 50% of some trait prompts.The remaining results were based exclusively on human evaluation of DALLE-v2.
  • Race: Traits such as “ambitious” and “determined” were most associated with black individuals, whereas “vigorous” and “detached” were most associated with East Asian individuals.These are reported as the strongest racial associations among the evaluated traits.
  • Race: White individuals were more commonly associated with positive traits including “competent,” “active,” “rational,” and “sympathetic.”The study separately grouped traits into positive and negative categories for racial-association analysis.

4.4 Representational bias for everyday situations

The everyday-situation analysis compared default generations with location-specific generations across six categories and 12 countries. Default outputs most closely represented Australia, Germany, and the United States, while Nigeria, Ethiopia, and Papua New Guinea were least represented.

  • Method: The study analyzed six everyday-situation categories using CLIP-embedding distances between default and location-specific generations across 12 countries.The categories were events, food, institutions, clothing, places, and community, with two highly populated countries selected per continent.
  • Events: Nigeria, Ethiopia, and Papua New Guinea had the lowest representation across events for both DALLE-v2 and Stable Diffusion.Australia, Germany, and the United States were most represented in the events analysis.
  • Across situations: Across situation prompts, Nigeria, Ethiopia, and Papua New Guinea were least represented by both models.Germany was most represented in DALLE-v2, whereas the United States was most represented in Stable Diffusion.

4.5 Human vs. Automated Evaluation

The paper compares human and automated assessments of demographic representations in occupation and personality-trait images. Agreement is generally high, but severe under-representation prevents meaningful correlation estimates for some groups.

  • Evaluation comparison: The evaluation study reports correlations between human and automated assessments for occupations and personality traits across gender, race, and age.Tables 4 and 5 summarize the corresponding evaluation results.
  • Agreement: The correlation coefficient exceeded 0.9 for every evaluated group except white individuals in the occupation category.This indicates strong agreement for most groups assessed by both methods.
  • Limitations: Meaningful correlation scores could not be computed for groups substantially under-represented in both models’ generations.The cited examples include age groups younger than 18 and older than 60.
  • Figures: Trait-balance figures identify traits with the most balanced race distribution, while the events heat maps compare default and country-specific generations for each model.These visualizations support interpretation of the paper’s automated and human-evaluation analyses.

4.6 Limitations

The study’s limitations include incomplete coverage of underrepresented groups, uncertainty about prompt expansion as a mitigation strategy, and evaluation restricted to two models.

  • The study does not deeply quantify non-binary gender, disability, smaller-country, or religious-group representation.It reports that these groups are generally poorly represented in neutral prompts based on example-based evidence.
  • Prompt expansion may not resolve quality discrepancies or complex associations requiring qualitative evaluation.Country-specific prompts may also associate countries with poor economic status or harmful stereotypes.
  • The evaluation covers DALLE-v2 and Stable Diffusion v1, leaving proprietary and newer models for future study.The paper specifically identifies Imagen and Stable Diffusion v2 as unevaluated examples.
  • Figure 15 shows the most represented countries across situation prompts for DALLE-v2.

5 CONCLUSION

The study evaluates representational biases in two text-to-image models across social dimensions and prompt types. It finds significant bias across dimensions, possible diversification through prompt expansion, image-quality variation, and uneven country representation.

  • The study measures gender, race, age, and geographical-location biases in DALL-E v2 and Stable Diffusion v1.It uses both human and automated evaluation across occupations, personality traits, everyday situations, and person prompts.
  • Both models exhibit significant biases across all evaluated dimensions and exacerbate biases relative to recent BLS labor statistics.
  • Prompt expansion can diversify generated content but may also produce variations in image quality.
  • Some countries are underrepresented while others are overrepresented in images depicting everyday situations.

Occupation prompts

The occupation-prompt appendix lists the personality-trait prompts used in the study, based on adjectives proposed in previous work.

  • Table 7 lists all personality-trait prompts used for the study.The traits correspond to adjective prompts proposed in previous work.

A APPENDIX

The appendix provides figures and tables documenting occupational, demographic, personality, geographical, and human-evaluation analyses for DALLE-v2 and Stable Diffusion.

  • Occupations: Table 8 reports correlations between BLS 2022 and DALLE-v2 or Stable Diffusion for occupation prompts.
  • Personality traits: Figures 25–26 show race distributions for positive and negative personality traits.
  • Human evaluation: Figure 41 contains the Amazon Mechanical Turk questionnaire used for human evaluation.
Loading 2304.06034v1…