Source-linked AI summary
Towards Measuring and Modeling "Culture" in LLMs: A Survey
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, Monojit Choudhury
TL;DR
This survey examines how cultural representation and inclusion are studied in LLMs despite culture evading simple definition. Reviewing 90 papers, it finds that studies use cultural proxies and that digital under-representation can aggravate biases and stereotypes.
Problem
Culture evades simple definition, while LLMs are strongly biased toward Western, Anglocentric, or American cultures, motivating systematic study of cultural representation and inclusion.
Method
The authors searched ACL Anthology, Google Scholar, NeurIPS, and the Web Conference, manually filtered 90 papers from 2020–2024, and labeled their cultural definitions, probing methods, languages, and studied cultures.
Results
None of the surveyed papers explicitly define culture; instead, they represent cultural differences through proxies, while digitally under-represented cultures may receive thin descriptions that aggravate biases and stereotypes.
Takeaways & Limitations
The survey situates LLM cultural-inclusion research within a broader landscape and identifies gaps for future work on cultural representation and evaluation.
Takeaways & Limitations
The analysis focuses primarily on probing LLMs and omits relevant work on culture and technology use from HCI and ICTD, as well as broader culture-and-AI, speech, and multimodality research.
Abstract
from arXiv · showhide
We present a survey of more than 90 recent papers that aim to study cultural representation and inclusion in large language models (LLMs). We observe that none of the studies explicitly define "culture, which is a complex, multifaceted concept; instead, they probe the models on some specially designed datasets which represent certain aspects of "culture". We call these aspects the proxies of culture, and organize them across two dimensions of demographic and semantic proxies. We also categorize the probing methods employed. Our analysis indicates that only certain aspects of ``culture,'' such as values and objectives, have been studied, leaving several other interesting and important facets, especially the multitude of semantic domains (Thompson et al., 2020) and aboutness (Hershcovich et al., 2022), unexplored. Two other crucial gaps are the lack of robustness of probing techniques and situated studies on the impact of cultural mis- and under-representation in LLM-based applications.
1 Introduction
The survey frames culture as multifaceted and difficult to define, so NLP studies usually examine measurable dataset-based proxies instead. It identifies incomplete coverage, methodological concerns, and a need for situated evaluation of cultural effects in applications.
- Culture as a complex concept: Culture encompasses heritage, interpersonal interactions, and collective ways of life, making a single practical definition difficult.Anthropological accounts distinguish outsider-oriented thin descriptions from insider-oriented thick descriptions that include lived context.
- Culture as a complex concept: Most NLP studies do not explicitly define culture; instead, their datasets specify the cultural features being examined as proxies.These proxies make conceptual cultural differences measurable through concrete NLP datasets.
- Survey rationale: The survey makes explicit how datasets instantiate cultural facets and proposes a schema linking employed datasets to the aspects of culture studied.The authors use this schema to expose gaps in existing research.
- Open concerns: Existing probing studies raise doubts about reliability and generalizability because benchmarking choices may not reveal the full extent of cultural limitations or representation.Cultural representation is also tied to local language use and terminology.
- Open concerns: Situated studies of LLM-based applications in particular cultural contexts are largely absent, limiting understanding of how culture plays out in real applications.The survey argues that combining rigorous benchmarks with naturalistic studies would provide a fuller picture.
2 Method
The survey searches and manually annotates literature on culture and LLMs, then organizes studies by cultural proxies and probing methods. Its taxonomy separates demographic from semantic proxies and characterizes predominantly black-box evaluation.
- Searching and annotation: The search combined ACL Anthology, Google Scholar, NeurIPS, and Web Conference sources with culture-related keywords, followed by manual filtering.The resulting corpus contained papers published between 2020 and 2024.
- Searching and annotation: The authors manually labeled papers by their culture definition, probing method, and studied languages and cultures.Because papers did not explicitly define culture, annotation emphasized data types representing cultural differences and language-culture interaction.
- Proxy taxonomy: The taxonomy groups 12 proxy labels into demographic proxies and semantic proxies, including culturally influenced domains such as emotions, values, food, and social etiquette.The semantic framework draws on 21 domains identified by Thompson et al. (2020).
- Proxy taxonomy: Demographic and semantic proxies are orthogonal and can be combined, such as studying festivals within a particular country.This allows one study to represent both a cultural domain and a demographic group.
- Probing methods: Most surveyed studies use black-box probing, appending cultural context to inputs and comparing responses across conditions and baselines.The survey distinguishes discriminative multiple-choice probing from generative open-ended probing.
- Probing methods: No surveyed culture study used white-box approaches, which the authors identify as a gap because such methods are more interpretable and likely more robust.White-box approaches can expose internal model states such as attention maps.
3 Findings: Defining Culture
The survey organizes cultural studies around demographic and semantic proxies, finding concentration in region, language, emotions, and values while several cultural facets remain underexplored.
- Demographic Proxies: 37 of 90 studies use geographical region, 35 use language, and 17 use both as cultural proxies.
- Demographic Proxies: These studies commonly contrast region- or language-specific datasets with dominant Western, American, or English baselines.
- Demographic Proxies: Gender, sexual orientation, race, ethnicity, and religion are usually studied through discrimination and stereotyping rather than groups’ cultural characteristics.
- Demographic Proxies: The surveyed demographic proxies are strongly shaped by Western diversity-and-inclusion discourse, leaving contexts such as caste underrepresented.
- Semantic Proxies: 25 of 55 studies on semantic proxies focus on emotions and values, one of 21 semantic-domain categories.
- Semantic Proxies: Other work examines food customs, recipes, named entities, and broad cultural practices through datasets spanning multiple domains.
- Semantic Proxies: Studies of semantic domains beyond emotions and values remain sparse, including culturally situated conventions such as Hindi pronoun usage, quantity, and kinship terms.
4 Findings: Probing Methods
The surveyed studies predominantly use black-box probing, comparing model responses under cultural conditions and baselines through single- or multi-turn prompts. These methods include discriminative and generative evaluation but face prompt-sensitivity concerns.
- Probing Approaches: Black-box probing is the most common approach, appending cultural context to prompts and comparing observed responses across conditions and baselines.
- Cultural Conditioning: Cultural context may be supplied through country placeholders, explicit norms or values, simulated conversations, structured text, or Wikipedia paragraphs.
- Probing Approaches: Discriminative probing evaluates predictions against ground-truth answers under different cultural conditioning and can compare them with unconditioned baseline predictions.
- Probing Approaches: Generative probing evaluates free-text outputs, which is less streamlined and may require manual inspection.
- Interaction Structure: Single-turn probing gives context and probe in one prompt, whereas multi-turn probing evaluates responses across several interactions.
- Limitations: Prompt sensitivity to irrelevant wording and formatting raises questions about probing reliability and generalizability.
5 Gaps and Recommendations
The survey identifies three major gaps in cultural-inclusion research: narrow coverage of cultural differences, limited methodological robustness, and insufficiently situated application studies. It recommends clearer cultural frameworks, broader methods, multilingual data, and interdisciplinary investigation.
- Studies heavily emphasize values and norms, leaving many other aspects of cultural difference understudied.
- Definition of culture: Existing work lacks a coherent framework for defining culture and contextualizing studies across a broader research program.
- Limited Exploration: Most surveyed datasets are in English, while culturally specific elements may not translate across languages.The survey therefore calls for culturally situated multilingual datasets collected or created from scratch.
- Lack of situated studies: The lack of situated studies makes the practical significance of cultural biases in real-world applications difficult to determine.
- Lack of interdisciplinarity: NLP studies rarely draw on anthropology or human-computer interaction to analyze culture’s complexity and technological consequences.The survey presents interdisciplinary work as a way to better understand and evaluate cultural inclusion.
6 Conclusion
The survey situates cultural-inclusion evaluation within a broader account of language and culture, emphasizing that culture is contextual, situated, and difficult for text-based LLMs to capture. Digitally under-represented cultures may consequently appear through outsider-produced thin descriptions that aggravate stereotypes.
- The survey offers a holistic view of cultural-inclusion evaluation by connecting language, cultural differences, and a broader research landscape.
- Culture’s amorphous, contextual, and situated nature creates bottlenecks for text-based LLMs attempting to master cultural nuances.The survey links this challenge to the need for thick descriptions, which digital text corpora rarely capture in entirety.
- Digitally under-represented cultures are more likely to be represented by outsider-produced thin descriptions in digital spaces.
- These thin descriptions can further aggravate cultural biases and stereotypes.
Limitations
The survey’s scope is centered on probing LLMs for culture, so relevant research outside that focus is not covered extensively. In particular, HCI and ICTD research are excluded from comprehensive treatment.
- The analysis primarily examines studies that probe LLMs in the context of culture.
- Research on culture outside the LLM-probing scope is not extensively covered, despite possible relevance to language technology and applications.
- The survey does not include research from Human-Computer Interaction and Information and Communication Technologies for Development.
A Black Box Probing Methods
The surveyed black-box approaches probe cultural associations and judgments through prompts, conditional likelihoods, forced-choice questions, scenarios, and conversational knowledge interventions. Examples span stereotypes, values, moral dilemmas, cultural preferences, and injected cultural information.
- Conditional-likelihood probing compares model preferences between paired sentences that differ in a social identity or stereotyped attribute.Examples contrast descriptions involving black versus white men and poor versus rich people.
- Value-oriented probes measure responses to culturally framed importance judgments and persona instructions.Examples ask about the importance of respectful superiors in China, interesting work for Chinese respondents, and high power, achievement, and self-enhancement.
- Conversational knowledge-injection prompts test how model responses change after culturally specific, ineffective, or anti-factual information is supplied.
- Forced-choice prompts test cultural judgments by presenting contexts, candidate answers, and sometimes explicit stereotype or anti-stereotype options.Examples cover gender stereotypes, racial associations, managerial attitudes, and culturally framed beliefs.
- Scenario-based moral-dilemma prompts ask models to choose among actions and justify decisions under stated principles or competing obligations.Examples include Heinz’s decision about stealing medicine and a software engineer’s conflict between fixing a security bug and attending a wedding.