Source-linked AI summary

Having Beer after Prayer? Measuring Cultural Bias in Large Language Models

Tarek Naous, Michael J. Ryan, Alan Ritter, Wei Xu

arXiv:2305.14456v4cs.CLcs.AIcs.LG

TL;DR

Large language models often fail to adapt culturally in Arabic, motivating better ways to measure Arab–Western bias. The paper introduces CAMeL, evaluates 16 LMs across multiple tasks, and finds Western-entity bias, stereotyping, unfairness, and limitations in some pre-training sources. It also identifies the analysis of corpus cultural relevance as a scope boundary for explaining these issues.

  • Problem

    Existing work provides limited evidence and no readily available resource for evaluating cultural appropriateness in non-Western, non-English settings, especially Arab–Western differences.

  • Method

    The paper constructs CAMeL with 628 naturally occurring Arabic prompts and 20,368 culturally labeled entities, then evaluates LMs across generation, NER, sentiment, and text-infilling tasks.

  • Results

    The evaluated LMs exhibit Western-entity bias, stereotyping, cultural unfairness, and failure to adapt appropriately to Arab contexts across the reported tasks.

  • Takeaways & Limitations

    CAMeL provides a foundation for evaluating and developing culturally aware LMs.

  • Takeaways & Limitations

    The pre-training-corpus analysis examines cultural-content relevance but does not deeply analyze the causes of stereotyping and unfairness.

Abstract

from arXiv · show

As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances. In this paper, we show that multilingual and Arabic monolingual LMs exhibit bias towards entities associated with Western culture. We introduce CAMeL, a novel resource of 628 naturally-occurring prompts and 20,368 entities spanning eight types that contrast Arab and Western cultures. CAMeL provides a foundation for measuring cultural biases in LMs through both extrinsic and intrinsic evaluations. Using CAMeL, we examine the cross-cultural performance in Arabic of 16 different LMs on tasks such as story generation, NER, and sentiment analysis, where we find concerning cases of stereotyping and cultural unfairness. We further test their text-infilling performance, revealing the incapability of appropriate adaptation to Arab cultural contexts. Finally, we analyze 6 Arabic pre-training corpora and find that commonly used sources such as Wikipedia may not be best suited to build culturally aware LMs, if used as they are without adjustment. We will make CAMeL publicly available at: https://github.com/tareknaous/camel

1 Introduction

The paper argues that culturally aware multilingual LMs remain important but inadequately adapt to Arab contexts, often prioritizing Western-centric content. It introduces CAMeL to measure these biases across diverse evaluations and reports stereotyping, unfairness, and corpus-relevance concerns.

  • Culturally aware LMs must capture distinctions among diverse communities, not merely bridge language barriers.
  • LMs often prioritize Western-centric completions in Arabic, including alcoholic beverages after prompts explicitly mentioning Islamic prayer.The paper also reports Western-centric names and food dishes in culturally inappropriate contexts.
  • Less work has examined cultural appropriateness in non-Western, non-English settings, and no readily available resource contrasted Arab and Western cultural differences.
  • CAMeL contains 20,368 Arab and Western entities across eight types and 628 naturally occurring prompts.The entities come from Wikidata and CommonCrawl, while prompts provide contexts for evaluation.
  • CAMeL supports story generation, NER, sentiment analysis, and text-infilling evaluations of 16 Arabic-data-pretrained LMs.Reported findings include stereotypes, cultural unfairness, and possible problems with Wikipedia as a culturally aware pre-training source.

2 Related Work

Prior work studies moral values, cultural surveys, commonsense knowledge, and social norms, but this paper focuses on culturally varying entities in natural contexts. CAMeL complements these approaches with evaluations spanning stereotypes, fairness, and text infilling.

  • Earlier studies probe moral variation and whether LMs encode the values of particular societies.
  • Survey-based studies measure cultural alignment through question-answering or cloze-style prompts about culturally divergent norms.
  • Other work tests geo-diverse facts, culinary customs, time expressions, and social-norm reasoning.
  • This paper instead studies culturally varying entities such as people’s names and food dishes using annotated resources from Wikidata and CommonCrawl.Its prompts come from social media and are naturally occurring rather than survey-generated.
  • CAMeL complements existing literature through stereotype examination, NER and sentiment fairness evaluation, and text-infilling tests.

3 Construction of CAMeL

CAMeL combines culturally labeled entities with naturally occurring Arabic prompts to support controlled evaluation of Arab–Western cultural distinctions. Its construction uses Wikidata, CommonCrawl, manual annotation, and culturally contextualized or neutral prompt types.

  • Collecting Cultural Entities: CAMeL covers eight entity types, including names, foods, beverages, clothing, locations, authors, worship places, and sports clubs.
  • Collecting Cultural Entities: Entities are collected from Arabic-labeled Wikidata classes and expanded with pattern-based extraction from Arabic CommonCrawl.Manual filtering and annotation are used to improve quality, while CommonCrawl expansion avoids using LMs in dataset construction.
  • Collecting Cultural Entities: Wikidata provides extensive Arabic coverage for locations, sports clubs, and authors but more limited coverage for other entity types.
  • Collecting Cultural Entities: 5k–10k unique extractions per entity type are manually filtered and annotated, and CAMeL includes both frequent and long-tail entities.About 15–20% of CommonCrawl extractions overlap with Wikidata.
  • Collecting Cultural Entities: Annotators classify extractions as Arab, Western, other foreign, not culture-specific, or non-entities.
  • Collecting Naturally Occurring Prompts: Prompts are split into culturally contextualized CAMeL-Co prompts and culturally agnostic prompts to test adaptation and default cultural leanings.
  • Collecting Naturally Occurring Prompts: Sentiment labels are positive, negative, or neutral, with Cohen’s Kappa inter-annotator agreement of 0.954.

4 Measuring Cultural Bias in LMs

The paper evaluates cultural bias across story generation, NER, sentiment analysis, and text infilling, finding stereotypes, unfairness, and persistent Western preferences in Arabic contexts.

  • Evaluation setup: CAMeL supports evaluation of 16 Arabic-data LMs across story generation, NER, sentiment analysis, and culturally appropriate text infilling.The models include Arabic monolingual and multilingual systems, including GPT-type, BERT-type, and T5-type architectures.
  • Cultural stereotypes in story generation: Stories about Arab characters more often used poverty and traditionalism associations, while Western stories more often used wealthy, likeable, and high-status associations.For male Arab names, dominance also appeared; benevolence was associated with female Arab names.
  • Fairness in NER and sentiment analysis: Up to 20 F1 points separated Western and Arab location tagging, while name-tagging differences were around 5 F1 points.Most LMs performed better on Western person names and locations; results were averaged across five runs.
  • Fairness in NER and sentiment analysis: Nearly all LMs showed more false association of Arab entities with negative sentiment, with no clear positive-sentiment preference.The analysis compared false-positive and false-negative predictions between sentences containing Arab and Western entities.
  • Culturally appropriate text infilling: 40-60% average CBS indicated persistent Western preference even in Arab cultural contexts, comparable to neutral prompts.Monolingual Arabic LMs also showed high CBS, multilingual LMs generally showed stronger bias, and Arab demonstrations reduced CBS for most models.

5 Analyzing Arabic Pre-training Data

The paper compares six Arabic pre-training corpora for cultural relevance and finds substantial differences in their Western-centricity.

  • Setup: The study trains unsmoothed 4-gram LMs on six Arabic corpora and compares their text-infilling Cultural Bias Scores.The corpora include local and international news, CommonCrawl, Arabic Wikipedia, and Arabic tweets.
  • Results: Arabic Wikipedia was the most Western-centric corpus despite often being considered a high-quality pre-training source.The paper attributes this mainly to the large portion of Arabic Wikipedia articles discussing Western content.
  • Results: International news had the second-highest CBS, while web-crawled data ranked third most Western-centric.The paper suggests machine-translated web content may contribute to Western content prevalence in Arabic.
  • Results: Local news and Twitter/X corpora had the lowest CBS among the analyzed sources.The paper suggests these sources may merit consideration for training more culturally adapted LMs.

6 Conclusion

The paper introduces CAMeL to evaluate cultural bias in LMs and reports Western bias, cultural unfairness, and inadequate adaptation in Arabic.

  • Contribution: CAMeL contains naturally occurring prompts and culturally relevant entities spanning eight entity types.The resource is designed to provide culturally contrasting prompt completions for evaluation.
  • Findings: Arabic LMs exhibit Western-entity bias, cultural unfairness in NER and sentiment analysis, and stereotypes in generated stories.The conclusion summarizes failures in appropriate cultural adaptation across these evaluation settings.
  • Implication: The authors release CAMeL to support evaluation and development of culturally aware LMs.The dataset is intended to enable measurement across multiple cultural-bias evaluation setups.

Limitations

The study’s limitations concern its cultural, linguistic, analytical, and corpus-focused scope. The authors identify finer-grained cultural distinctions, additional generation features, and deeper corpus analyses as future work.

  • The study focuses on overall adaptation to Arab contexts and Western-entity bias, leaving sub-cultural distinctions within Arab and Western worlds unexplored.
  • CAMeL covers only Arabic and evaluates bias between Western and Arab cultural entities.
  • The story-generation stereotype analysis is limited to lexical terms, specifically adjectives, rather than stylistic features.
  • The pre-training-corpus analysis examines cultural-content relevance but does not quantify entity-theme co-occurrences or assess how fine-tuning datasets may amplify fairness problems.

Ethics Statement

The paper frames cultural adaptation around user preferences, data privacy, research-only release, non-toxic prompts, and Arabic grammatical requirements. It distinguishes grammatical gender handling from the study of gender bias.

  • Neutral-context cultural defaults depend on users’ preferences and backgrounds, while current LMs default to Western culture.
  • CAMeL-Ag prompts are proposed as a test bed for aligning LMs with users’ unique cultural preferences.
  • The prompts are anonymized masked versions of Twitter/X contexts, contain no personally identifiable information, and are released exclusively for research.
  • Prompts avoid toxic or offensive language, and names and clothing are gender-grouped to match Arabic grammatical conjugation rather than define gender identities.

A Additional Background

The background distinguishes cultural appropriateness from earlier English-centered bias studies and introduces CAMeL’s entity-based, naturally occurring, intrinsic and extrinsic evaluation framework. It also describes the resource’s construction choices and corpus-analysis setup.

  • Additional Background: Prior work studied demographic and social biases, but less work examined cultural appropriateness and adaptation in non-English, non-Western environments.
  • Additional Background: CAMeL contains naturally occurring Arabic prompts and Arab- and Western-associated entities across eight culturally variable entity types.
  • Additional Background: Unlike intrinsic embedding- and probability-based approaches, extrinsic methods compare downstream model behavior when groups are switched in evaluation contexts.
  • Additional Background: CAMeL supports extrinsic evaluation through generation, sentiment analysis, and NER, plus intrinsic evaluation through text infilling.
  • Additional Background: For pre-training-corpus analysis, researchers trained unsmoothed 4-gram LMs on six Arabic corpora and compared their CAMeL text-infilling performance.
  • Collecting Arab and Western Entities: Arabic grammatical gender requires names and clothing entities to be categorized by gender so prompt verbs agree grammatically.

D Language Models Details

The paper evaluates a broad set of Arabic, bilingual, and multilingual language models, alongside diverse Arabic pre-training corpora and cultural-bias measures. Its analyses include model-specific architectures, corpus composition, adjective odds ratios, and CAMeL-based cultural preference scores.

  • Language Models: AraBERT combines Arabic Wikipedia, a 1.5B-word Arabic corpus, OSCAR, and Assafir data, with base and large versions evaluated.
  • Language Models: AraBERT-T extends AraBERT through continued pre-training on 60M Arabic tweets, with base and large architectures.
  • Stereotype Measurement: Odds ratios compare adjective occurrence odds in stories about Western-named characters against stories about Arab-named characters.Additional results analyze female names and stereotypical traits.
  • Stereotype Results: Arab-named female characters are associated with traditionalism and poverty, while Western-named characters receive more likeable traits.The adjective “generous” also appears frequently in stories about Arab-named female characters.

F.3.2 Results on CAMeL-Ag

On culturally agnostic CAMeL-Ag prompts, models frequently prefer Western entities even when neither cultural group is contextually required. Multilingual models generally show stronger Western preference than monolingual models.

  • The same broad trends observed in the main CAMeL results also appear on culturally agnostic prompts.
  • 70-80% CBS is reached across entity types on culturally agnostic prompts, showing frequent Western-entity preference without cultural contextualization.
  • Multilingual models generally show higher CBS than monolingual models on CAMeL-Ag.

G.1 Analyzing Entity Encodings

The paper examines how Arab and Western entities are represented in model embeddings and relates representation clustering to cultural-bias behavior. Monolingual models more often preserve cultural separation, whereas multilingual models generally mix the representations more strongly.

  • Visualization: Entity embeddings are averaged across tokens and prompts before being projected into two dimensions with t-SNE.
  • Visualization: Most monolingual BERT models separate Arab and Western entities into distinct clusters, while most multilingual models mix them.Bilingual GigaBERT models still show recognizable cultural clusters.
  • Clustering Quality: Davies-Bouldin Index measures within-cluster compactness and between-cluster separation, with lower values indicating better clustering.
  • Clustering Quality: Multilingual models generally have higher DBIs than monolingual models, with XLM-R showing the poorest clustering quality.The findings suggest multilingual capability can coincide with less culturally distinct representations.
  • Prompt Structure: Dropping first-person pronouns tests whether English-like Arabic prompt structure contributes to stronger preference for Western entities.
  • Prompt Structure: Most models achieve higher CBS with Arabic prompts that retain an English-like structure than with pronoun-dropped prompts.
Loading 2305.14456v4…