Source-linked AI summary
Cultural Bias and Cultural Alignment of Large Language Models
Yan Tao, Olga Viberg, Ryan S. Baker, Rene F. Kizilcec
TL;DR
Cultural values embedded in increasingly used generative AI may bias people’s authentic expression, especially when models reflect overrepresented cultures. The study compares five GPT models with nationally representative survey data across countries and territories, finding Western cultural alignment without prompting and improved alignment for many locations with cultural prompting.
Problem
The paper addresses limited evidence about how cultural bias varies across countries and GPT model generations, and whether prompting can improve cultural alignment.
Method
The authors compare responses from five GPT models with Integrated Values Surveys and Inglehart-Welzel cultural-map values across 107 countries and territories, with and without country-specific cultural prompting.
Results
GPT outputs resemble English-speaking and Protestant European cultures without prompting, while cultural prompting improves alignment for 71.0%–81.3% of countries and territories across GPT-4o, GPT-4-turbo, and GPT-4.
Takeaways & Limitations
Cultural prompting is a simple, flexible, and accessible approach that can improve alignment with a given cultural context, although it does not eliminate disparities.
Takeaways & Limitations
Results may depend on English-language prompts and their phrasing, and survey-question behavior should not be generalized uncritically to broader LLM use.
Abstract
from arXiv · showhide
Culture fundamentally shapes people's reasoning, behavior, and communication. As people increasingly use generative artificial intelligence (AI) to expedite and automate personal and professional tasks, cultural values embedded in AI models may bias people's authentic expression and contribute to the dominance of certain cultures. We conduct a disaggregated evaluation of cultural bias for five widely used large language models (OpenAI's GPT-4o/4-turbo/4/3.5-turbo/3) by comparing the models' responses to nationally representative survey data. All models exhibit cultural values resembling English-speaking and Protestant European countries. We test cultural prompting as a control strategy to increase cultural alignment for each country/territory. For recent models (GPT-4, 4-turbo, 4o), this improves the cultural alignment of the models' output for 71-81% of countries and territories. We suggest using cultural prompting and ongoing evaluation to reduce cultural bias in the output of generative AI.
1 Introduction
Culture shapes reasoning, behavior, and communication, while language and generative AI increasingly mediate how people produce and consume language. Because LLM training data overrepresents some regions, this study evaluates cultural bias and tests prompting models to represent specific societies.
- Cultural differences shape perception, causal attribution, communication, and other aspects of how people think and behave.
- Generative AI increasingly affects daily language production and consumption across education, medicine, public health, and creative and opinion writing.
- LLMs trained predominantly on English text exhibit latent bias toward Western cultural values, especially when prompted in English.
- Cultural prompting instructs an LLM to answer like a person from another society, offering a flexible strategy whose effectiveness depends on cultural representation.
- The study evaluates cultural bias across 107 countries and territories for five GPT models and examines cultural prompting across models released from 2020 to 2024.
- The authors benchmark cultural values using nationally representative Integrated Values Surveys and the Inglehart-Welzel map’s survival–self-expression and traditional–secular-rational dimensions.
2 Results
Without cultural prompting, GPT models’ expressed values align most closely with Anglosphere and Protestant European countries and differ most from African-Islamic countries. Cultural prompting generally reduces country-level cultural distance, but its effectiveness varies across models and countries.
- GPT models without cultural prompting align most closely with Anglosphere and Protestant European values and most distantly with African-Islamic values.
- GPT outputs consistently favor self-expression values involving environmental protection and tolerance of diversity, foreigners, gender equality, and different sexual orientations.
- Cultural prompting reduces average cultural distance for GPT-4o from 2.42 to 1.57, GPT-4-turbo from 2.71 to 1.77, and GPT-4 from 2.69 to 1.65.
- Figure 2 compares country-level Euclidean-distance distributions with and without prompting, averaging GPT outputs across ten prompt variations except for GPT-3.
- Cultural prompting improves alignment for 71.0% of countries with GPT-4o, 81.3% with GPT-4-turbo, 77.6% with GPT-4, 72.6% with GPT-3.5-turbo, and 80.4% with GPT-3.
- Cultural prompting reduces GPT-4o’s distance from Jordan from 4.10 to 0.36, but increases distances for some countries including Finland, Luxembourg, Andorra, Switzerland, and Taiwan ROC.
3 Discussion
The study finds robust cultural bias in GPT outputs toward English-speaking and Protestant European values, while cultural prompting improves alignment but does not eliminate disparities. These findings motivate critical evaluation, broader audits, and culturally informed AI literacy.
- Findings: GPT outputs show unequal distances from countries’ local cultural values, suggesting a bias toward English-speaking and Protestant European values.The pattern remains robust across GPT versions and prompt wordings when no cultural identity is specified.
- Implications: GPT’s observed bias toward self-expression values may lead GPT-assisted communication to convey more interpersonal trust, bipartisanship, and support for gender equity.The authors note that these signals may have interpersonal and professional consequences, especially for users outside Anglosphere and Protestant European contexts.
- Findings: Cultural prompting improves alignment with a specified cultural context but cannot entirely eliminate disparities between generated and observed cultural values.For GPT-4o, the prompted mean cultural distance is 1.57, approximately the distance between GPT-4o and Uruguay in Figure 1.
- Findings: Cultural prompting failed to improve alignment or exacerbated cultural bias for 19-29% of countries and territories.Thus, the strategy is not a universal remedy for cultural misalignment.
- Limitations: The study cautions against assuming that LLM responses to cultural-values surveys predict behavior in everyday human-LLM interactions.Prompt language, prompt phrasing, and differences between human and LLM survey response mechanisms constrain generalization.
- Implications: Cultural prompting and ongoing cultural-alignment evaluation can help mitigate, but not fully control, cultural bias in LLM outputs.The authors frame this as a lesson for AI literacy and encourage similar evaluations of internationally used LLMs.
4 Materials and Methods
The study reconstructs a cultural-values benchmark from IVS survey data, queries five GPT versions with standardized prompts, and compares model locations with country-level cultural values. It then evaluates cultural prompting by measuring country-specific distances between prompted outputs and the benchmark.
- Benchmark construction: The IVS benchmark combines WVS and EVS time-series data from 2005 to 2022, covering 393,536 individual responses from 112 countries.The country-level cultural map retains all survey waves for countries participating in multiple waves.
- Benchmark construction: The cultural map uses ten IVS questions spanning happiness, trust, authority, religion, sexuality, abortion, nationality, post-materialism, and autonomy.These questions reproduce the inputs used for the Inglehart-Welzel World Cultural Map.
- Benchmark construction: Five countries were omitted because principal component scores were undefined for all participants due to invalid responses on at least one question.The remaining 107 countries were summarized using rescaled individual-level and country-year-level means.
- Model evaluation: Ten respondent-descriptor variants were used to assess sensitivity to prompt wording, except GPT-3, which used only variant 0 because it was deprecated.Responses were standardized, projected into the IVS-based PCA space, and averaged across variants.
- Cultural prompting: Cultural prompting added each country or territory’s cultural identity to the respondent descriptor while keeping the rest of the procedure unchanged.The resulting responses were projected into the cultural-map space and averaged across prompt variants.
- Cultural prompting: For each model, the analysis compared Euclidean distances from unprompted model values to each country’s benchmark with prompted values to the corresponding country benchmark.These two distance distributions form the basis for evaluating cultural alignment.