Source-linked AI summary

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Sebastian Ruder, Madeline Smith, Antoine Bosselut, Alice Oh, Andre F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, Sara Hooker

arXiv:2412.03304v2cs.CL

TL;DR

Translated MMLU benchmarks may confound multilingual performance with Western-centric cultural knowledge and translation artefacts. The paper measures these effects across models and introduces Global-MMLU, finding that cultural composition and subset choice change evaluations while expanding coverage to 42 languages. It concludes that culturally annotated, higher-quality multilingual benchmarks are needed for more reliable assessment.

  • Problem

    Translated English benchmarks do not guarantee multicultural evaluation, because cultural bias and translation artefacts can limit MMLU’s effectiveness as a global benchmark.

  • Method

    The paper annotates MMLU for cultural sensitivity, improves translations through professional and community contributors, and evaluates open-weight and proprietary models on Global-MMLU subsets.

  • Results

    Model rankings change between culturally sensitive and culturally agnostic subsets, while 28% of MMLU questions require culturally sensitive knowledge.

  • Takeaways & Limitations

    Global multilingual evaluations should use Global-MMLU and report culturally sensitive and culturally agnostic results separately.

  • Takeaways & Limitations

    Identifying cultural sensitivity does not fully resolve missing non-Western cultural representation or achieve complete cultural inclusion.

Abstract

from arXiv · show

Cultural biases in multilingual datasets pose significant challenges for their effectiveness as global benchmarks. These biases stem not only from differences in language but also from the cultural knowledge required to interpret questions, reducing the practical utility of translated datasets like MMLU. Furthermore, translation often introduces artefacts that can distort the meaning or clarity of questions in the target language. A common practice in multilingual evaluation is to rely on machine-translated evaluation sets, but simply translating a dataset is insufficient to address these challenges. In this work, we trace the impact of both of these issues on multilingual evaluations and ensuing model performances. Our large-scale evaluation of state-of-the-art open and proprietary models illustrates that progress on MMLU depends heavily on learning Western-centric concepts, with 28% of all questions requiring culturally sensitive knowledge. Moreover, for questions requiring geographic knowledge, an astounding 84.9% focus on either North American or European regions. Rankings of model evaluations change depending on whether they are evaluated on the full portion or the subset of questions annotated as culturally sensitive, showing the distortion to model rankings when blindly relying on translated MMLU. We release Global MMLU, an improved MMLU with evaluation coverage across 42 languages -- with improved overall quality by engaging with compensated professional and community annotators to verify translation quality while also rigorously evaluating cultural biases present in the original dataset. This comprehensive Global MMLU set also includes designated subsets labeled as culturally sensitive and culturally agnostic to allow for more holistic, complete evaluation.

1 Introduction

Global MMLU examines how cultural bias and translation quality limit translated MMLU as a multilingual benchmark, then introduces an improved 42-language resource with culturally sensitive and agnostic metadata. Evaluations show that cultural sensitivity, translation quality, and subset choice can materially affect model assessment.

  • Cultural bias: Translated English benchmarks do not guarantee multicultural evaluation and can preserve Western-centric concepts, including US-specific history, accounting, and law.The paper also identifies translation artefacts as a separate challenge for multilingual evaluation.
  • Cultural bias: 28% of annotated MMLU questions require Western cultural knowledge, while 84.9% of geographic questions focus on North American or European regions.These distributions indicate substantial Western-centric content in the benchmark.
  • Global-MMLU: Global-MMLU expands multilingual MMLU coverage to 42 languages and combines professional translations, post-edits, crowdsourced translations, and machine translations.The release also provides metadata for culturally sensitive and culturally agnostic subsets, plus Global-MMLU Lite.
  • Model evaluation: Model rankings change when evaluations use culturally sensitive or culturally agnostic subsets rather than aggregate translated MMLU scores.The authors evaluate 14 open-weight and proprietary models to measure this ranking sensitivity.
  • Translation quality: Human-translated and machine-translated datasets produce notable performance differences across high- and low-resource languages.The paper argues that low-resource language evaluation remains uncertain without high-quality human-translated or in-language data.
  • Recommendations: The paper recommends prioritizing Global-MMLU and reporting culturally sensitive and culturally agnostic results separately in multilingual evaluations.These recommendations are intended to provide clearer and more reliable assessments across languages and cultural contexts.

2 Evaluating cultural bias in MMLU

The study annotates MMLU questions for cultural, geographic, dialectal, and temporal knowledge, then measures how these attributes shape the benchmark’s composition. It finds substantial Western and subject-specific cultural concentration, with culturally sensitive questions concentrated in humanities and social sciences.

  • 2.1 Data Annotation Process: 2,850 MMLU questions were sampled uniformly across 57 subjects and annotated for cultural, geographic, dialectal, and temporal knowledge.Each subject contributed 50 samples; annotations were aggregated using majority agreement.
  • 2.2 Analysis of MMLU Cultural Biases: 28% of MMLU questions require culturally sensitive knowledge, while dialect-specific knowledge accounts for 0.5%.Cultural knowledge represents 32.7% of the relevant tags, and 10.6% of questions require both cultural and geographic knowledge.
  • 2.2 Analysis of MMLU Cultural Biases: 73.9% of Western-culture questions require knowledge of the United States, while India accounts for 59% of Asian-culture questions.The United Kingdom contributes 8% of Western-culture questions; China and Japan each contribute 17.9% of Asian-culture questions.
  • 2.2 Analysis of MMLU Cultural Biases: Humanities contains 68% culturally sensitive questions, with more than 80% in Philosophy, Moral Scenarios, High School US History, and High School Government and Politics.STEM and Medical questions generally require less cultural or regional knowledge.
  • 2.2 Analysis of MMLU Cultural Biases: Culturally sensitive questions overrepresent Social Sciences and Humanities while underrepresenting STEM, shifting the dataset toward contextual rather than technical content.Social Sciences comprise 26.3% of culturally sensitive questions versus 21.1% of the annotated sample, while STEM comprises 2.9% versus 33.3%.

3 Introducing Global-MMLU

Global-MMLU expands multilingual MMLU while improving translation quality through professional, community, and machine translation workflows. It covers 42 languages and preserves culturally sensitive and culturally agnostic subsets for targeted evaluation.

  • 3.1 Translation Process: Google Translate was used to translate MMLU into 41 languages because prior evaluations found it stronger than alternatives on low-resource languages.The authors also report higher ChrF++ scores and lower per-subject deviation than GPT-3.5-Turbo in their comparison.
  • 3.1 Translation Process: Compensated professional annotators reviewed Arabic, French, Hindi, and Spanish translations for fluency and cultural appropriateness, forming the Gold Set.Community contributors additionally verified fluency and corrected poor translations across a broader set of languages.
  • 3.1 Translation Process: 7,565 edits were made, covering 36.9% of reviewed samples.Professional annotators edited an average of 789 samples per language, while community contributors edited 362 per language.
  • 3.2 Data Composition of Global-MMLU: Global-MMLU covers all 14K MMLU samples across 42 languages, totaling 589,764 samples.The dataset integrates human-translated datasets, machine translations, and the original English MMLU.
  • 3.2 Data Composition of Global-MMLU: Global-MMLU includes culturally sensitive and culturally agnostic subsets derived from 2,850 annotated English samples and extended across 41 languages.The culturally sensitive subset contains 33,264 entries, while the culturally agnostic subset contains 86,436 entries.

4 Model Evaluations

Model evaluations vary substantially with cultural sensitivity, language-resource level, and model size. Culturally sensitive subsets produce greater cross-language performance volatility and more ranking changes, especially for low-resource languages, while larger models are more consistent.

  • 4 Model Evaluations: Closed-source models such as GPT-4o and Claude 3.5 Sonnet consistently outperform smaller open-source models.The authors caution that direct proprietary-versus-open comparisons are not feasible because of model-size and evaluation-method differences.
  • 4 Model Evaluations: Culturally sensitive tasks show higher performance variance across languages than culturally agnostic tasks for all models.The paper attributes this pattern to contextual demands, translation sensitivity, and models’ uneven training across cultures.
  • 4 Model Evaluations: Low-resource languages have average standard deviations of 6.37 on CA datasets and 6.78 on CS datasets.These represent increases of 98% and 75% compared with high-resource languages, respectively.
  • 4 Model Evaluations: Chinese and Hindi are the most sensitive languages to culture-specific knowledge, with models showing both ranking increases and decreases.Similar variation occurs in French, German, Italian, Japanese, and Portuguese; Aya Expanse and CommandR often improve on CS datasets.
  • 4 Model Evaluations: 6.8 average rank changes and 9.1 position shifts occur on CS datasets for high-resource languages, compared with 3.3 and 3 on CA datasets.The comparison shows greater ranking movement on culturally sensitive evaluations even among high-resource languages.
  • 4 Model Evaluations: Large models average only 0.21 rank changes on CA datasets, while their maximum position shift is 3 compared with 5 for small models.The paper links this consistency to robustness and greater capacity to generalize across diverse datasets.

5 Related Work

Prior multilingual benchmarks combine broad language coverage with limitations in cultural representation, translation quality, and cross-language comparability. Existing work addresses these issues through multilingual exams, culturally focused evaluations, and studies of native versus translated data.

  • Multilingual benchmarks: Language-specific MMLU variants expand evaluation for individual languages, including Arabic, Chinese, Indonesian, Thai, Turkish, African, Korean, and Vietnamese settings.
  • Multilingual benchmarks: Multilingual benchmarks cover diverse tasks and languages, but many support only a small number of languages or lack a consistent framework for direct comparison.INCLUDE is noted as an exception, covering 44 languages through local exams across countries and languages.
  • Translation quality: Translated MMLU versions broaden language coverage, but machine-generated translations can vary significantly in quality and may not be reliable across languages.ChatGPT-based translation covered 26 languages, while MMMLU used professional human translators for 14 languages.
  • Cultural alignment: Cultural-alignment research finds that language models often reflect values and opinions aligned with Western culture across multiple languages.SEA-HELM emphasizes Southeast Asian languages using handcrafted diagnostics and manually translated, validated tasks.
  • Multimodal evaluation: Multilingual visual-language benchmarks extend cultural evaluation through broad language coverage and culturally diverse images and questions.PangeaBench covers 47 languages, while CVQA spans 30 countries and 31 languages.
  • Training data and culture: Studies of training data report that native instructions typically outperform translated data, while human-written multilingual data is used to reflect local culture and preferences.

6 Conclusion

The paper finds that translated MMLU contains substantial Western-centric cultural bias and that translation artifacts and cultural sensitivity affect multilingual model rankings. Global-MMLU addresses these issues with culturally differentiated evaluation subsets and broad multilingual coverage.

  • Findings: 28% of MMLU questions require culturally-sensitive knowledge, and geographic questions primarily focus on North America and Europe.The paper reports that this bias persists in translated MMLU variants and risks over-indexing evaluations on Western-centric idioms and knowledge.
  • Findings: Global-MMLU examines how translation artifacts and cultural bias affect multilingual model rankings.
  • Dataset contribution: Global-MMLU distinguishes culturally-sensitive and culturally-agnostic knowledge using professional and crowdsourced annotations for multilingual evaluation.
  • Evaluation implications: Model rankings change across culturally-sensitive and culturally-agnostic subsets, so translated MMLU performance alone is insufficient to indicate multilingual progress.The authors recommend reporting Global-MMLU subsets as part of holistic multilingual capability evaluations.
  • Dataset contribution: Global-MMLU is released under a fully permissive license for use in evaluations.

7 Limitations

The paper identifies limits in Global-MMLU’s coverage, annotation balance, taxonomy, and safety monitoring. These constraints mean the benchmark does not yet represent the world’s linguistic and cultural diversity comprehensively.

  • Uneven distribution of contributions: Community-annotator contributions were uneven across languages, potentially limiting distributional and annotator diversity.Some languages were dominated by one or two frequent contributors.
  • Language and dialect coverage: Only a fraction of the world’s linguistic diversity is represented beyond the 42 evaluated languages.Future work is needed to extend evaluations beyond these languages.
  • Language and dialect coverage: 42 languages are covered, but many dialects and language varieties remain unrepresented.The authors explicitly leave dialect coverage for future work.
  • Toxic or offensive speech: The annotation interface did not explicitly flag toxic, harmful, or offensive speech, so such data may remain in the benchmark.This content was not monitored during cultural-sensitivity annotation or translation post-editing.
  • Region category assignment: The six-region geographic taxonomy may be too coarse, with the authors recommending more granular World Bank designations.The proposed changes include separate categories for Central America and Sub-Saharan Africa.
  • Cultural inclusion: Global-MMLU identifies cultural sensitivity but does not by itself guarantee culturally inclusive evaluation.The authors call for culturally grounded knowledge to achieve greater inclusivity and fairness.

C Global-MMLU Lite

Global-MMLU Lite is a lighter, balanced multilingual evaluation set designed to provide efficient coverage across languages and subject categories. Its sampling procedure balances culturally sensitive and agnostic examples while adjusting representation for sparse subjects.

  • Sampling strategy: Subjects exclusively tagged as culturally sensitive were excluded so both culturally sensitive and culturally agnostic samples could be represented.Social Sciences and Humanities consequently become more prevalent in the Lite distribution.
  • Sampling strategy: Five culturally sensitive and five culturally agnostic samples were selected per subject where available.Subjects with only one culturally sensitive sample received one sample of each type.
  • Sampling strategy: Business, Medical, and General Knowledge subjects were slightly upsampled to ensure adequate representation.General Knowledge used 22 Miscellaneous and 8 Global Facts samples per category.
  • Dataset design: Global-MMLU Lite is designed as a balanced dataset for efficient multilingual evaluation across multiple languages.It is described as a lighter version of Global-MMLU.

D Temporal Knowledge

The annotation process also identifies questions whose answers depend on temporal or time-sensitive knowledge. Such questions are uncommon overall and concentrated outside STEM subjects.

  • Temporal knowledge: Time-sensitive questions are those whose correct answers may change because of current political leaders or economic statistics.Annotators were asked to label samples requiring this type of knowledge.
  • Temporal knowledge: 2.4% of the dataset is tagged as time-sensitive.Most time-sensitive samples belong to Social Sciences, Humanities, Medical, and Other categories.
  • Temporal knowledge: STEM contains no time-sensitive samples.The figure caption separately notes that STEM subjects do not include temporal knowledge.

F Additional Results

Additional results describe how model rankings and subject performance are analyzed across culturally agnostic and sensitive subsets, resource levels, and subject areas. Aya Expanse 32B performs above its average more often on Social Sciences and Humanities than on STEM subjects.

  • Model rankings: Tables 6 and 7 compare model-rank changes and position shifts across high-, mid-, and low-resource languages.Color-coded boxes mark ranking increases and decreases relative to the MA reference.
  • Subject performance: Aya Expanse 32B achieves 66.4% average accuracy across subjects.Most STEM subjects fall below this average, while most Social Sciences and Humanities subjects exceed it.
  • Culture and region: Among Western-culture samples, 73.3% are also tagged North America and 25.5% Europe; Asian-culture samples are 97.2% associated with Asia.These associations are reported as part of the Western–Asian culture and region analysis.

G.2 Culture Country Relations

The annotations show that culture–country and region–country relationships are unevenly distributed, with some categories concentrated in particular countries. Latin American and Indigenous tags have different country patterns, while North America is dominated by the United States.

  • Culture–country relations: Latin American culture tags are distributed across Bolivia, Mexico, Honduras, and Peru, while Indigenous tags are split between the United States and Micronesia.Bolivia and Mexico each account for 33.3% of Latin American tags; Honduras and Peru each account for 16.7%, while the Indigenous distribution is 66.7% United States and 33.3% Micronesia.
  • Tag relationships: The figures organize relationships between Western and Asian cultures, cultures and countries, and regions and countries.
  • Region–country relations: 89.6% of North American regional tags identify the United States, with Canada and the United Kingdom each representing 0.8%.
  • Region–country relations: European tags are more distributed than North American tags, with the United Kingdom at 20.1% and France at 10.1%.

H Annotation Process

The annotation process used compensated annotators, structured interfaces, guided review, and agreement checks to assess cultural sensitivity and translation quality. Coverage was complete for cultural-sensitivity labels but partial for translation review and editing.

  • Annotation coordination: Author briefings, discussion channels, comments, and ambiguity resolution were used to calibrate annotations across annotators and languages.
  • Annotation coverage: 100% of selected samples received cultural-sensitivity labels, while 37% were fully reviewed and 12.3% were edited for translation quality.
  • Annotation infrastructure: The annotation interfaces were built with Argilla and deployed through Hugging Face Spaces with Hugging Face account access.
  • Annotation tasks: Annotators labeled questions for cultural, geographic, dialect, and regional knowledge and reviewed translated questions against the original English.
  • Agreement checks: 12 subjects had complete unanimous agreement on time-sensitivity annotations, while moral scenarios showed reasonable disagreement across annotation types.

I Translation Analysis

The translation analysis compares machine translations with human-edited references using translation-quality judgments and edit-distance measures. Results differ by subject category: Humanities has the largest raw question edits, while STEM answers have the highest normalized edit distance.

  • Machine-translation comparison: Google Translate performs significantly better than GPT-3.5-turbo across MMLU subject categories in the reported translation-quality comparison.The comparison uses MMMLU samples as a human reference and includes only languages overlapping across the compared translated sets.
  • Interpretation of edit measures: The analysis distinguishes raw editing volume from normalized editing complexity across questions and answers.
  • Raw edit distance: Humanities has the largest average edit distances, with questions requiring more edits than answers.Edit distance is computed between machine translations and versions edited by professional and community annotators.
  • Normalized edit distance: Questions in Humanities have the greatest average length, whereas STEM answers have the highest normalized edit distance.Normalized Edit Distance divides edit distance by text length to account for the influence of longer question-answer pairs.
Loading 2412.03304v2…