Source-linked AI summary

Investigating Cultural Alignment of Large Language Models

Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, Mona Diab

arXiv:2402.13231v2cs.CLcs.CY

TL;DR

The paper investigates whether LLMs capture diverse cultural knowledge and measures this by comparing persona-conditioned model answers with survey participants’ responses. Across experiments in Egypt and the United States, cultural alignment varies with prompting and pretraining language, while the study also identifies demographic gaps and proposes Anthropological Prompting.

  • Problem

    The paper asks whether LLMs genuinely capture the diverse knowledge and views of different cultures.

  • Method

    The study simulates surveys in Egypt and the United States by prompting multiple LLMs in English or Arabic with respondents’ personas and questions, then compares answers with survey responses.

  • Results

    Cultural alignment is greater when prompting uses a culture’s prevalent language and when pretraining includes a refined mixture of languages used by that culture.

  • Takeaways & Limitations

    Language choices in prompting and pretraining are important considerations for representing cultural diversity in LLM responses.

  • Takeaways & Limitations

    The study is limited to two languages and two countries, and the actual pretraining data sources, domains, and dialect coverage of many models are unknown.

Abstract

from arXiv · show

The intricate relationship between language and culture has long been a subject of exploration within the realm of linguistic anthropology. Large Language Models (LLMs), promoted as repositories of collective human knowledge, raise a pivotal question: do these models genuinely encapsulate the diverse knowledge adopted by different cultures? Our study reveals that these models demonstrate greater cultural alignment along two dimensions -- firstly, when prompted with the dominant language of a specific culture, and secondly, when pretrained with a refined mixture of languages employed by that culture. We quantify cultural alignment by simulating sociological surveys, comparing model responses to those of actual survey participants as references. Specifically, we replicate a survey conducted in various regions of Egypt and the United States through prompting LLMs with different pretraining data mixtures in both Arabic and English with the personas of the real respondents and the survey questions. Further analysis reveals that misalignment becomes more pronounced for underrepresented personas and for culturally sensitive topics, such as those probing social values. Finally, we introduce Anthropological Prompting, a novel method leveraging anthropological reasoning to enhance cultural alignment. Our study emphasizes the necessity for a more balanced multilingual pretraining dataset to better represent the diversity of human experience and the plurality of different cultures with many implications on the topic of cross-lingual transfer.

1 Introduction

The paper asks whether LLMs capture cultural knowledge and proposes measuring cultural alignment against survey responses. It examines how prompting language, pretraining language composition, demographics, and Anthropological Prompting relate to alignment.

  • Motivation and framework: Cultural alignment is measured by comparing persona-conditioned model answers with actual survey responses as a reference.The persona describes traits such as social class, education level, and age.
  • Study focus: Prompting language and pretraining language composition are studied as factors affecting model responses across cultures.The study uses English and Arabic prompts and models with different language compositions.
  • Experimental setting: The experiments simulate surveys in Egypt and the United States using LLMs prompted with personas matching the original respondents.The selected models include multilingual, primarily English-trained, and Arabic-focused systems.
  • Findings and contribution: The authors report that models capture demographic variance unevenly, with larger alignment gaps for underrepresented groups.This finding motivates attention to representation across demographic groups.
  • Findings and contribution: Anthropological Prompting is introduced as a method intended to enhance cultural alignment.It is presented as the paper’s final contribution.

2 Research Questions

The research questions test whether native-language prompting and culturally matched pretraining improve alignment, and whether alignment varies across personas and topics. They also examine cross-lingual transfer through finetuning.

  • Prompting language: The study tests whether prompting in a culture’s native language yields greater cultural alignment than prompting in a foreign language.Arabic prompting is expected to align more closely with Egyptian survey responses than English prompting.
  • Pretraining data composition: The study tests whether increasing a culture’s representation in pretraining data improves alignment for that culture.The stated comparison concerns fixed-size models with different language compositions.
  • Personas and cultural topics: The authors expect greater misalignment for digitally underrepresented personas and uncommon cultural topics.The motivating example contrasts a working-class persona in Aswan with an upper-middle-class persona in Cairo.
  • Cross-lingual transfer: The study evaluates cross-lingual transfer by comparing primarily English-trained LLaMA-2-Chat-13B with its Arabic-and-English-finetuned derivative, AceGPT-Chat-13B.Both models have 13B parameters, while AceGPT is further finetuned on Arabic and English data.

3 Anthropological Preliminaries

The paper treats culture as multidimensional and expressed through human communities’ worldviews, beliefs, behaviors, and linguistic records. Its analysis assumes relationships between language, cultural representation, and the production of linguistic data.

  • Culture: Cultural trends can appear in written records, and a model is culturally aligned when its views match those of a group in a given scenario.This links linguistic expression to the paper’s alignment concept.
  • Culture: The paper frames artificial agents as both influenced by and influential in human patterns of behavior and recorded ideas.This places LLM outputs within a broader account of cultural expression.
  • Language and culture: The paper assumes that a language can serve as a proxy for its dominant culture, while acknowledging that languages may be shared across cultures.Population-specific dialect prompting is identified as one way to reduce this concern.
  • Language and culture: The paper assumes that cultures most often produce linguistic records and communications in one dominant language.It also notes that people in Egypt may express opinions online in English for various reasons.

4 Experimental Setup

The experiment simulates World Values Survey responses from Egypt and the United States using demographic personas, multilingual prompts, and four instruction-tuned LLMs. Cultural alignment is measured by comparing model answers with survey responses using hard and soft metrics, while Anthropological Prompting guides models to reason about cultural identities and perspectives.

  • 4.1 World Values Survey: The study selects 30 culturally variable World Values Survey questions from Egypt and the United States.The source surveys contain 1,200 Egyptian and 2,596 United States participants; the simulations use 303 matched personas per country.
  • 4.2 Survey Participants: Each persona specifies six demographic dimensions, and prompts instruct the model to emulate that persona and select a numbered answer.The prompt contains persona information, perspective-following instructions, output constraints, and the survey question with answer options.
  • 4.4 Pretrained Large Language Models: The model set includes GPT-3.5, mT0-XXL, LLaMA-2-13B-Chat, and Arabic-English-finetuned AceGPT-13B-Chat with differing language mixtures.The three non-GPT models are 13B-parameter instruction-tuned models selected for comparison, including a more balanced multilingual model and an English-dominant model.
  • 4.5 Computing Cultural Alignment: Each question is paraphrased four ways, sampled five times at temperature 0.7, and resolved by majority vote.The resulting answer for each persona and question variant is compared with the original survey subject’s response.
  • 4.5 Computing Cultural Alignment: Hard alignment measures exact answer accuracy, while soft alignment awards partial credit for ordinal responses and otherwise defaults to accuracy.The soft alignment score averages 1 minus the per-question, per-persona error across queries and personas.
  • 4.6 Anthropological Prompting: Anthropological Prompting digitally adapts ethnographic fieldwork by directing models to consider cultural complexity, identity, language, and emic and etic perspectives.The method encourages models to reason like anthropologists and analyze subjects and topics through layered interpersonal and cultural perspectives.

5 Results

Cultural alignment varies with survey population, prompting language, demographics, and model adaptation. Anthropological Prompting improves alignment, including for underrepresented personas, despite using fewer sampled responses than vanilla prompting.

  • 5.1 Anglocentric Bias in LLMs: All evaluated LLMs align more strongly with United States survey respondents than Egyptian respondents, consistent with euro-centric bias in current models.
  • 5.2 Prompting & Pretraining Languages: Using a country’s dominant language increases alignment for GPT-3.5 and AceGPT-Chat across both metrics.Arabic improves alignment with Egypt, whereas English improves alignment with the United States.
  • 5.3 Digitally Underrepresented Personas: Alignment improves with higher social class and education, while male and older personas receive more accurate model responses than female and younger personas.
  • 5.5 Finetuning for Cultural Alignment: Finetuning LLaMA-2-Chat on Arabic data improves Egyptian alignment but decreases alignment with the United States survey.The decline suggests loss of existing US cultural knowledge during adaptation to another language.
  • 5.6 Anthropological Prompting: Anthropological Prompting outperforms vanilla prompting on Egypt alignment across both metrics, even though it generates one response instead of five.
  • 5.6 Anthropological Prompting: Anthropological Prompting improves alignment for underrepresented social-class and education-level groups, making their alignment distribution more equitable.

6 Discussion

The discussion links cultural alignment to both pretraining and prompting languages, while emphasizing limits from language choice, cultural diversity, and cross-lingual knowledge transfer. The study uses Modern Standard Arabic despite its mismatch with Egyptians’ daily speech.

  • 6 Discussion: Pretraining and prompting languages both enhance cultural alignment when they are prevalent in the relevant country.The authors relate this pattern to culture-specific content being generated primarily in native languages online.
  • 6 Discussion: Pretraining encodes cultural knowledge in model parameters, while prompting language activates the subnetwork associated with that knowledge.The authors identify cross-lingual transfer as limited, particularly between Arabic and English scripts.
  • 6 Discussion: The study represents Egyptian culture primarily with Modern Standard Arabic, although Egyptians do not use MSA in daily interactions.Egypt also contains dialectal and demographic variation, so the study measures multiple personas rather than one Egyptian archetype.
  • 6 Discussion: Cultural information is difficult to verify and consolidate, and transfer commonly flows from dominant languages and cultures into other-language responses.

7 Related Work

Related work studies cultural alignment, demographic effects, biases, and prompting methods from complementary perspectives. This paper extends that literature through persona-level demographic analysis and explicit examination of language and pretraining composition.

  • 7 Related Work: Prior survey-based work measures model–survey distribution similarity with Jensen-Shannon Distance rather than persona-level response alignment.
  • 7 Related Work: Other studies probe cross-cultural differences in multilingual encoders or evaluate ChatGPT against societies using WVS, Hofstede, and multilingual prompting.
  • 7 Related Work: Research reports Western cultural trends in multilingual and Arabic models, while other work proposes prompting methods to increase response diversity.
  • 7 Related Work: Some findings caution that LLMs should not serve as proxies for human opinions because altered wording can produce mismatched response biases.
  • 7 Related Work: This paper adds analysis of demographic dimensions, digitally underrepresented personas, question topics, pretraining-language composition, and prompting language across models.
  • 7 Related Work: Earlier studies document harmful demographic biases and stereotypes in model outputs, including stronger persona-associated toxicity for some demographic groups.

8 Conclusion & Future Work

The paper introduces a persona-level framework for measuring cultural alignment and Anthropological Prompting for improving it. Future work will broaden the framework across cultures and languages and test its use for cross-lingual transfer.

  • 8 Conclusion & Future Work: The framework measures whether four LLMs capture cultural trends in Egypt and the United States using participant-mirroring personas across six demographic dimensions.It compares persona-level responses, varies pretraining-language composition, and prompts models in native study languages.
  • 8 Conclusion & Future Work: Anthropological Prompting guides models to reason about personas using a framework adapted from anthropological methods to improve cultural alignment.
  • 8 Conclusion & Future Work: Future work will apply the framework to more cultures and languages and test cultural alignment as a proxy for cross-lingual knowledge transfer.

Limitations

The study’s scope is constrained by limited cultural, linguistic, survey, model, and pretraining-data coverage. The authors also note unresolved challenges in capturing human complexity and refining Anthropological Prompting.

  • Scope: The analysis covers only two languages and two countries, limiting the scope for assessing whether findings generalize across cultures.The authors selected Egypt, the United States, Arabic, and English to keep the analysis tractable while examining several additional dimensions.
  • Model selection: Available Arabic models lacked suitable instruction tuning and often had substantially fewer parameters, preventing inclusion of an Arabic monolingual model.This constrained model selection and limited the comparison of pretraining-language composition.
  • Data sources: Using one survey source leaves generalization to other cross-national surveys untested.The authors identify Arab-Barometer as an example of an additional survey source worth exploring.
  • Human complexity: The authors acknowledge that LLMs cannot capture the essence and complexity of human experience, despite attempts to model nuanced diversity.Creative prompting is used to mimic diversity, but the authors state that this remains insufficient.
  • Anthropological Prompting: Anthropological Prompting still requires refinement across languages and prompt variations to better investigate dataset biases.The authors connect this need to the large number of languages and the diversity of possible prompt formulations.
  • Pretraining transparency: Unknown pretraining sources, domains, and dialect coverage constrain comprehensive interpretation of model behavior and create downstream ethical implications.The authors specifically identify the black-box nature of models such as GPT-3.5 as a significant limitation.

Ethics Statement

Culturally non-aligned LLMs can fail to serve the people affected by downstream sociotechnical systems or create harm. The authors therefore advocate collaboration between computer scientists and social scientists to study language, culture, and machine behavior together.

  • Ethical stakes: If LLMs are non-aligned with cultural values, downstream systems may fail to serve intended populations or create harm.The authors frame cultural alignment as relevant because LLMs are pervasive components of sociotechnical systems.
  • Interdisciplinary responsibility: The paper calls for interdisciplinary collaboration to uncover LLM biases and advance AI ethically.The proposed collaboration joins researchers studying human language and cultures with researchers studying machine internals.

A Extended Results

Extended results examine model consistency, survey construction, cultural alignment across themes and languages, and Anthropological Prompting. They report language-dependent alignment and consistency patterns, alongside improved demographic balance with anthropological reasoning.

  • Language effects: GPT-3.5 and AceGPT-Chat align better with each country’s dominant prompt language on both metrics, while LLaMA-2-Chat favors English and mT0-XXL shows a cross-country reversal.For mT0-XXL, English prompting performs better for Egypt, whereas Arabic prompting performs better for the United States.
  • Survey simulation: Each survey question receives four generated linguistic variations, and five responses per variation are sampled before majority voting determines the model answer.This procedure addresses unavailable interviewer wording while aggregating responses at the question-persona level.
  • Model consistency: English prompts generally produce higher response consistency than Arabic prompts, except for AceGPT-Chat, and the disparity decreases with improved multilingual pretraining.Consistency is measured across alternative phrasings of the same question under a fixed model, question, and persona.
  • Survey data: The dataset uses six demographic dimensions from World Values Survey wave 7, with 303 paired participants selected from Egypt and the United States.The source surveys interviewed 1,200 participants in Egypt and 2,596 in the United States; paired personas matched on demographic parameters except geography.
  • Anthropological Prompting: Anthropological Prompting asks the model to reason from an anthropological framework before answering the persona-based question and choices.The framework includes emic and etic perspectives, cultural context, individual experience, cultural relativism, space and time, and nuance.
  • Anthropological Prompting: Anthropological Prompting improves cultural alignment and produces a more balanced distribution across demographic dimensions for GPT-3.5.The reported comparison uses English prompts and both Soft and Hard similarity metrics.
Loading 2402.13231v2…