Source-linked AI summary
Easily Accessible Text-to-Image Generation Amplifies Demographic Stereotypes at Large Scale
Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, Aylin Caliskan
TL;DR
The paper asks how widely available text-to-image models amplify dangerous and complex stereotypes, a question made consequential by their mass use and production of millions of images daily. Using mixed-methods analysis of Stable Diffusion, it examines ordinary prompts, generated exemplars, and user or institutional interventions. It finds stereotypes across traits, occupations, social groups, and objects, including racial and gender disparities and dominant cultural norms that persist despite counter-prompting and guardrails.
Problem
Mass deployment of accessible text-to-image models raises the question of whether they amplify dangerous and complex stereotypes across generated images.
Method
The paper uses mixed-methods analysis of Stable Diffusion, combining ordinary prompts, generated exemplars, and qualitative connections to psychological, sociological, and critical race theory literature.
Results
The study finds stereotypes across traits, occupations, social groups, and objects, including racial and gender disparities and dominant cultural norms that persist despite counter-prompts and system guardrails.
Takeaways & Limitations
The findings indicate that easily accessible ordinary prompts can mass-produce and disseminate images reinforcing historically dangerous stereotypes.
Takeaways & Limitations
The paper does not survey model strengths, quantitatively assess all possible mitigation strategies, or identify an optimal broader solution.
Abstract
from arXiv · showhide
Machine learning models that convert user-written text descriptions into images are now widely available online and used by millions of users to generate millions of images a day. We investigate the potential for these models to amplify dangerous and complex stereotypes. We find a broad range of ordinary prompts produce stereotypes, including prompts simply mentioning traits, descriptors, occupations, or objects. For example, we find cases of prompting for basic traits or social roles resulting in images reinforcing whiteness as ideal, prompting for occupations resulting in amplification of racial and gender disparities, and prompting for objects resulting in reification of American norms. Stereotypes are present regardless of whether prompts explicitly mention identity and demographic language or avoid such language. Moreover, stereotypes persist despite mitigation strategies; neither user attempts to counter stereotypes by requesting images with specific counter-stereotypes nor institutional attempts to add system ``guardrails'' have prevented the perpetuation of stereotypes. Our analysis justifies concerns regarding the impacts of today's models, presenting striking exemplars, and connecting these findings with deep insights into harms drawn from social scientific and humanist disciplines. This work contributes to the effort to shed light on the uniquely complex biases in language-vision models and demonstrates the ways that the mass deployment of text-to-image generation models results in mass dissemination of stereotypes and resulting harms.
1 Introduction
Text-to-image models are widely accessible and can generate stereotypes through ordinary prompts, including prompts without identity language or attempts to counter stereotypes. The paper characterizes these biases and their persistence, emphasizing their potential prevalence and difficulty of mitigation.
- 1 Introduction: Millions of users generate millions of images daily with easily accessible text-to-image models trained on web-scraped data containing stereotyping and other harmful content.The models often require little or no technical expertise, while their training data are primarily English and include toxic and pornographic material.
- 1 Introduction: The study uses mixed methods to analyze Stable Diffusion through stereotype-inducing prompts, exemplar images, and qualitative connections to social-scientific and humanist research.The analysis examines user interventions such as careful prompting and institutional interventions such as system guardrails.
- 1 Introduction: Simple descriptors can reproduce stereotypes, including attractive person prompts approximating a White ideal and terrorist prompts generating brown faces with dark hair and beards.These outputs connect ordinary descriptors to historically subordinating and anti–Middle Eastern violence narratives.
- 1 Introduction: 99% of generated software developer images are represented as white, compared with 56% of U.S. software developers identified as white.This example illustrates near-total stereotype amplification for an occupation with comparable real-world demographic statistics.
- 1 Introduction: Prompts mentioning social groups or everyday objects associate groups with poverty and subordination, reproduce North American and heterosexual defaults, and may resist explicit counter-stereotypes.The paper reports that adding terms such as wealthy may still fail to produce images that counter unintended poverty stereotypes.
- 1 Introduction: The paper argues that accessible ordinary prompts make these patterns plausible and potentially prevalent, while the many intersecting dimensions of identity complicate mitigation.It does not survey model strengths, assess all mitigation strategies, or identify an optimal broader solution.
2 Prompts with no identity language perpetuate and amplify stereotypes
Stable Diffusion produces harmful stereotypes from prompts that avoid explicit identity language, linking neutral descriptors, occupations, and objects to racial, gendered, classed, and cultural norms. These outputs can amplify occupational disparities and default to American representations.
- 2.1 Human traits and descriptors: Perpetuating stereotypes: Ten neutral human-descriptor prompts generated images tying descriptors to stereotypically racialized and gendered visual features.The study generated 100 images per descriptor and examined random samples.
- 2.1 Human traits and descriptors: Perpetuating stereotypes: Attractiveness was represented near a White ideal, while poverty, criminality, and terrorism were associated with darker or brown faces and stereotypically Black or Middle-Eastern features.These associations reproduce harmful social narratives and can subordinate or criminalize targeted groups.
- 2.1 Human traits and descriptors: Perpetuating stereotypes: Prompts for happy couples and families produced straight-passing images reinforcing heteronormative family and marriage structures.Such normative assumptions can alienate people who do not conform to them.
- 2.2 Occupations: Stereotype amplification: Occupation prompts amplified gender and race imbalances even without mentioning demographic identities.Software developer images were nearly exclusively pale and masculine, whereas housekeeper images were darker and feminine.
- 2.2 Occupations: Stereotype amplification: 99% of generated software developer images were represented as white, compared with 56% of U.S. software developers identifying as white.Several occupations, including housekeeper, nurse, and flight attendant, had 100% female image representations.
- 2.2 Occupations: Stereotype amplification: Amplification was unevenly distributed: prestigious, higher-income occupations skewed more white and male, while lower-income occupations appeared more non-white and female.The authors connect this pattern to representational and allocational harms.
- 2.3 The view from nowhere: Defaulting to Americanness: Neutral object prompts typically produced images most similar to North American prompts and most different from Africa prompts, which encoded poverty stereotypes.The outputs therefore reflected cultural defaults rather than global population distributions.
3 Prompts with identity language perpetuate stereotypes, despite mitigation efforts
Identity prompts produce layered stereotypes across people, objects, and backgrounds, while counter-stereotypical modifiers often fail to remove them. The model repeatedly links African identities with poverty and Middle-Eastern identities with war, even when prompts request wealth or other counter-stereotypes.
- 3.1 Stereotyping representations of groups: Ethiopian prompts linked people and cars with poverty, while Iraqi prompts linked people with war and military force.These patterns reinforced narratives that African countries are defined by poverty and Iraqis by war.
- 3.1 Stereotyping representations of groups: Identity-language prompts produced systemic disadvantage across people, objects, and backgrounds, including malnourishment, poverty, and war-related associations.The examples show that bias can operate through multiple visual components of one generated scene.
- 3.2 Stereotypes despite counter-stereotypes: Adding targeted modifiers can sometimes increase diversity, but in many cases it does not eliminate stereotypes.The authors report more diverse generations for some prompts, including software developer and flight attendant, but not consistently.
- 3.2 Stereotypes despite counter-stereotypes: A “white poor person” prompt still mostly generated darker-skinned faces, adding some features associated with whiteness such as blue eyes.The model therefore continued to associate poverty with Blackness despite the counter-stereotypical identity modifier.
- 3.2 Stereotypes despite counter-stereotypes: The findings show that counter-stereotypical wording often fails to disentangle poverty from Blackness and terrorism from Middle-Eastern identity.These associations persisted across modified prompts rather than being removed by explicit textual opposition.
- 3.2 Stereotypes despite counter-stereotypes: Prompts describing a fancy or mansion-owning African man continued to depict a simple hut or broken structure compared with an American man.For a wealthy African man, the house remained unchanged while the man received a Western suit as a wealth signal.
4 Stereotypes are perpetuated despite institutional guardrails: The Case of DALL·E
Dall·E reproduces stereotypes found in Stable Diffusion, and its guardrails do not prevent subtler biases across race, disability, gender, and sexuality.
- Dall·E outputs for occupational prompts show varied skin tones, yet subtler biases beyond gender and race persist.The authors attribute the apparent attention to occupational bias partly to its prominence in computer science bias research.
- Dall·E reproduces Stable Diffusion’s associations between African identity and poorer, more worn-down houses compared with American identity.When both identities appear, the generated house shifts toward the American-associated style and quality.
- The prompt “a disabled woman leading a meeting” produces disabled women listening rather than leading, unlike the corresponding “blonde woman” prompt.The example indicates that explicit counter-stereotypical prompting does not reliably overcome ableist associations.
- Prompts for “a happy couple” and “a happy family” reinforce heteronormative ideals despite Dall·E’s implemented guardrails.The examples show that mitigation does not eliminate biases along less straightforward dimensions.
5 Conclusion
The paper concludes that widely deployed image generators embed dangerous, difficult-to-mitigate biases. Because outputs infer unspecified visual details from training norms, addressing these harms requires sustained analysis beyond narrow computational metrics.
- Millions of daily generated images create serious concern about how embedded biases will be used and shape the world.The authors state that users and model owners may find such biases challenging or impossible to anticipate, quantify, or mitigate comprehensively.
- Unspecified visual details force models to infer characteristics, causing outputs to adhere to norms reflected in training data and processes.Images also provide many dimensions for subtle meanings that text-focused bias methods may not capture.
- The authors urge caution and recommend refraining from applications whose outputs have downstream effects on the real world.They describe the models as brittle and limited in the worlds they create despite their apparent power and versatility.
òū̧ ĺƜʼnđòū̧ ũòū̧ ƤưòūĘʼnūĻ̧ ūĠǖư̧
Figure 8 presents examples of complex biases in Dall·E involving race, disability, and family norms.
- African-associated houses appear in worse condition than American-associated houses, while mixed prompts shift house style and quality toward the American version.
- “A disabled woman leading a meeting” produces a disabled woman listening rather than leading, whereas “blonde woman” yields the intended scene.
- “A happy family” produces heteronormative images of marriage and family.
A.1 Additional examples of stereotypes in Stable Diffusion
The appendix provides additional trait-based and occupation-based stereotype examples, discussed in Sections 2 and 3.
- Additional trait-based and occupation-based stereotype examples appear in Figures A5 and A6.The paper discusses these examples in Sections 2 and 3, respectively.
A.2 Additional Examples of Bias in DALL·E
Figures A1 and A2 provide additional DALL·E examples of complex bias that support the paper’s main analyses.
- Figures A1 and A2 present additional examples of complex bias in DALL·E.
- The examples are used to support the analyses described in the main paper.
- These examples extend the paper’s discussion beyond its primary figures.
B Prompts
The appendix tests descriptor and occupation prompts using multiple templates, emphasizing visible face outputs and illustrating complex demographic stereotypes in DALL·E.
- B.1 Descriptors and occupations: The descriptor prompts include traits, social roles, and potentially stigmatizing character descriptions such as “attractive,” “poor,” “terrorist,” and “illegal.”
- B.1 Descriptors and occupations: The occupation prompts span chefs, firefighters, therapists, CEOs, housekeepers, pilots, flight attendants, taxi drivers, and nurses.
- B.1 Descriptors and occupations: Occupation titles in parentheses correspond to official U.S. Bureau of Labor Statistics categories used to obtain self-reported demographics.
- B.2 Different prompting templates: Three prompt templates produced similar results: “a photo of X person,” “a photo of a person working as X,” and “a photo of X.”
- B.2 Different prompting templates: The study focuses on “a photo of the face of X” because it generally generates clearer, more visible faces.
- B.2 Different prompting templates: For “a family in Kenya” versus California, Dall·E includes indicators of poverty, as it does for cars in Ethiopia versus California.
- B.2 Different prompting templates: Adding “disabled” to “a woman in a room full of people” produces a group containing multiple people in wheelchairs.
- B.2 Different prompting templates: Adding “non-binary” to “a person in Ethiopia” incorporates younger Western concepts in clothing, rainbows, and hairstyles.
C Additional methodological details and social context
The appendix details image representation, filtering, sampling, and demographic measurement choices while illustrating stereotype patterns and framing their social context.
- Images were generated with Stable Diffusion v1-4 using the latest diffusers library with default parameters.
- CLIP-L-14 represented gender and ethnicity, while CLIP-B-32 represented objects.
- Taxi-driver outputs required manual removal of images containing only taxis, which occurred in 20% of cases.
- The Chicago Face dataset sample included 100 images each of self-identified Asian, white, and Black individuals with neutral expressions.
- For “a white terrorist,” long beards resemble outputs for “a terrorist” and “a Middle-Eastern,” associating the attribute with Middle-Eastern appearances.
- Object prompts without identity descriptors most resemble North America prompts and differ most from Africa prompts, encoding stereotypes of poverty.
- The sample included 75 self-identified males and 75 self-identified females, with results showing only white versus non-white distribution.
- The study uses U.S. demographic categories to measure stereotypically raced and gendered traits, not because those categories are objectively true.