Source-linked AI summary
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li
TL;DR
Existing LLM benchmarks provide limited coverage of multilingual, cross-linguistic reasoning and can suffer from inconsistent translation quality. MMLU-ProX addresses this gap with a 29-language benchmark and expert-verified translation pipeline, then evaluates 36 LLMs. The results show strong high-resource-language performance but marked declines for low-resource languages, with reported gaps of up to 24.3%.
Problem
Existing benchmarks have limited multilingual coverage or translation quality, restricting comprehensive evaluation of cross-linguistic reasoning.
Method
MMLU-ProX extends MMLU-Pro to 29 languages and uses an LLM-driven translation framework with expert verification before evaluating 36 LLMs.
Results
LLMs perform well in high-resource languages but decline in low-resource languages, with gaps of up to 24.3%.
Takeaways & Limitations
MMLU-ProX supports broader assessment of multilingual reasoning and highlights limitations in current LLM global accessibility and fairness.
Takeaways & Limitations
Language coverage is constrained by budget, leaving expansion to additional, especially extremely low-resource, languages for future work.
Abstract
from arXiv · showhide
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it challenging to comprehensively assess LLMs' performance in the multilingual setting. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-linguistic comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, with gaps of up to 24.3%. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.
1 Introduction
MMLU-ProX addresses limited multilingual benchmark coverage and inconsistent translation quality by extending reasoning-focused evaluation across 29 languages with expert-verified translations. It supports cross-lingual reasoning assessment and systematic analysis of multilingual LLM disparities.
- Existing multilingual evaluation is limited by language coverage and translation quality, while monolingual benchmarks provide narrow linguistic insight.
- Each language version contains 11,829 questions, with a 658-question lite version for efficient evaluation.
- A semi-automated translation framework combines SOTA LLM translation with expert verification to improve expression, terminology consistency, and cultural appropriateness.
- The study evaluates 36 recent LLMs using 0-shot and 5-shot chain-of-thought prompting across open-weight and proprietary models.
- The analysis finds significant multilingual capability disparities, highlighting limitations in global contexts and the need for stronger fairness evaluations.
2 Related Work
Prior multilingual benchmarks trade off language coverage, translation consistency, and reasoning difficulty. MMLU-ProX is positioned against these limitations by combining broad multilingual coverage with challenging reasoning evaluation.
- Table 1 compares multilingual benchmarks by whether they include chain-of-thought and parallel data.
- Single-language benchmarks assess expert reasoning rigorously but provide limited evidence about multilingual performance.
- MGSM and XCOPA provide multilingual coverage but restrict evaluation to narrow reasoning formats such as mathematics or causal inference.
- Global-MMLU covers 42 languages through human-machine hybrid translation but has inconsistent translation quality and limited reasoning difficulty.
- MMLU-Pro increases reasoning complexity and distractor options over original MMLU, but remains English-centric.
3 Benchmark
MMLU-ProX constructs a 29-language benchmark through curation, translation, verification, and expert review. Its development emphasizes deduplication, preservation of technical content, and human assessment of translation quality.
- Benchmark scope: MMLU-ProX covers 29 specified languages, extending MMLU-Pro to typologically diverse linguistic settings.
- Pipeline: The data pipeline comprises data curation, translation, external model verification, and expert review.
- Data curation: Curation removes or merges duplicate questions and manually corrects grammatical, hyphenation, symbol, and related inconsistencies.
- Translation: Claude Sonnet 3.7 translates the benchmark using prompts that preserve terminology, cultural appropriateness, LaTeX, formulas, code, symbols, numerical relationships, and formatting.
- Expert verification: Expert verification samples 20 items from each of 14 disciplines across 15 languages, with two native-speaking translators rating accuracy, fluency, and completeness from 1–5.
- Expert verification: Categories scoring below 3 from both translators are retranslated; only Yoruba law required this modification, while other categories averaged at least 4.
- Quality results: Expert evaluations report consistently high translation quality across resource groups, including Wolof, Yoruba, and Nepali.
- Resources: Development cost approaches $80,000 at market rates, including translation, testing, expert verification, and computation.
4 Experiments
MMLU-ProX evaluates 36 open-weight and proprietary LLMs across 29 languages, revealing strong overall performance from leading models but substantial disparities across languages and language families. Reasoning-enhanced models improve multilingual performance, especially in lower-resource settings.
- Experimental setup: 36 LLMs are evaluated across all 29 MMLU-ProX languages, with Table 3 reporting CoT performance and selected model-family averages.The evaluated models include open-weight and proprietary systems spanning varied architectures, parameter scales, and training paradigms.
- Overall performance: 75.2% is DeepSeek-R1’s average across languages, followed by GPT-4.1 at 72.7% and DeepSeek-V3 at 70.5%.The results also report that larger models generally outperform smaller counterparts.
- Reasoning-enhanced models: 4.7% is DeepSeek-R1’s average improvement over DeepSeek-V3, with gains of 11.3% on Wolof and 9.3% on Yoruba.The comparison concerns reasoning-focused versus standard models and shows larger gains in these low-resource languages.
- Reasoning-enhanced models: 80.7%, 80.7%, and 80.9% are Qwen3-235B with thinking mode’s scores on English, Spanish, and Italian, respectively.The thinking-mode variant reaches state-of-the-art performance on these Western European languages.
- Language disparities: 0.6% to 58.6% is the performance range reported for Wolof, while Western European languages exceed 75% for some models.African languages show the lowest performance overall, whereas Western European languages consistently perform strongly.
5 Analysis
The analysis examines model scaling, prompting, and evaluation efficiency. Larger models and demonstrations benefit multilingual performance unevenly, while the lite benchmark closely preserves full-benchmark outcomes and model rankings.
- 5.1 Model Size: 59.9% is the Qwen3 dense 32B model’s average accuracy, a 17.9% absolute gain over the 4B model.The largest increase occurs from 8B to 14B at +8.0%, while later gains are more modest.
- 5.1 Model Size: 20.5% is Wolof’s improvement with Qwen3 scaling, compared with 12.6% for English and 14.4% for Russian.Low-resource languages generally benefit more from increased model size, and Zulu shows meaningful gains only for the largest models.
- 5.2 Prompting Strategies: 5-shot prompting generally improves performance, but gains vary across languages and model families.GPT-4.1 gains +3.7% on English, while low-resource African languages benefit more substantially; Qwen3-30B in thinking mode changes less between prompting styles.
- 5.3 Full and Lite Versions: 658 items per language comprise the lite version, compared with 11,829 questions per language in the full version.Both versions reserve 70 validation questions for few-shot prompt construction, leaving 588 and 11,759 assessment questions respectively.
- 5.3 Full and Lite Versions: 1.5% and 1.1% are the lite-versus-full gaps for DeepSeek-R1 and GPT-4.1, while DeepSeek-V3 differs by 0.4%.The lite version preserves model rankings almost perfectly and shows under-1% differences for Wolof.
6 Conclusion
MMLU-ProX extends reasoning-focused multilingual evaluation to 29 diverse languages and combines LLM translation with expert verification. Evaluating 36 state-of-the-art LLMs reveals substantial multilingual performance disparities and supports the goal of more equitable global access.
- Benchmark contribution: LLMs and expert verification form the semi-automatic translation approach used to ensure quality across languages.The paper pairs this benchmark construction with a comprehensive evaluation of 36 state-of-the-art LLMs.
- Conclusion: 36 state-of-the-art LLMs reveal significant performance disparities in the multilingual setting.The stated aim is to promote equitable accessibility of LLMs in the global context.
Limitations
The benchmark has scope, verification, and modality limitations that constrain how broadly its results can be generalized. It covers 29 languages, uses partial expert verification, and evaluates textual inputs only.
- 29 languages are included, while adding more—especially extremely low-resource languages—remains future work because of budget constraints.
- Expert verification was not feasible across all languages and subject areas because of resource constraints.
- Automated translation may still introduce subtle quality errors, particularly in complex or domain-specific content.
- The benchmark focuses solely on textual inputs and does not account for multimodal contexts.
C Performance Patterns across Language Groups
Performance varies substantially across language groups. Western and Eastern European languages generally perform strongly, while African and some South Asian languages show larger disparities and lower scores.
- Western European languages typically score 70-80% for top-performing models, with Qwen3-235B with thinking exceeding 77% across the family.It reaches 80.9% for Italian and 80.7% for both English and Spanish.
- South Asian performance varies by language, with Hindi ranging from 58.4% to 78.7% and Telugu generally performing lower across models.
- Indonesian reaches 81.3% with DeepSeek-R1, while Chinese varies from 53.4% to 75.5% across models.
- African languages show the largest disparities: Wolof ranges from 0.6% to 58.6%, Yoruba from 3.9% to 57.0%, and Zulu from 11.5% to 67.3%.
- Eastern European languages exhibit comparable performance, with Qwen3-235B-Think achieving over 77%.
D Translation Pipeline Analysis
The translation framework is evaluated against reasoning-based translation and human translators using professional scoring. The comparison reports that the reasoning-based method is only marginally inferior to the framework.
- The evaluation compares reasoning-based translation with human translators for English-to-Japanese translation using Accuracy, Fluency, and Completeness scores out of 5.
- The reasoning-based method achieves translation quality only marginally inferior to the proposed translation framework.
E Expert Verification Guidance
Expert verification guidance evaluates translated assessment content for accuracy, fluency, and completeness. The pipeline also specifies format-preserving translation instructions and structured JSON output.
- Expert annotators assess translation quality using Accuracy, Fluency, and Completeness criteria on a 1-5 scale.
- A score of 5 requires accurate technical terminology, natural target-language expression, and full retention of the source meaning and details.
- Expert verification was conducted in 15 selected languages, with results reported in Table 5.
- The framework preserves LaTeX, mathematical formulas, programming code, option counts, and English JSON keys while translating question and option content.
G Detailed Evaluation Results
This section presents additional MMLU-ProX evaluation results across 5-shot and zero-shot settings, including both full and lite versions.
- Additional MMLU-ProX results are reported in Tables 6–9.The tables cover full and lite benchmark variants under 5-shot and zero-shot evaluation.
- Table 6 reports MMLU-ProX 5-shot results.
- Table 7 reports MMLU-ProX zero-shot results.
- Tables 8 and 9 report MMLU-ProX Lite results under 5-shot and zero-shot settings, respectively.