Source-linked AI summary

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, Jingren Zhou

arXiv:2504.18428v4cs.CL

TL;DR

Multilingual mathematical reasoning remains underexplored because existing datasets are often too simple and challenging benchmarks are largely English-only. PolyMath addresses this gap with a four-level, 18-language benchmark and evaluates advanced LLMs across performance and linguistic behavior. The results show limited absolute performance, substantial cross-language variation, language inconsistency and thinking-length differences, and potential benefits from controlling the reasoning language.

  • Problem

    Existing multilingual mathematical datasets are too simple for advanced LLMs, while most challenging benchmarks are English-only, limiting study of multilingual mathematical reasoning.

  • Method

    PolyMath organizes mathematical problems into four difficulty levels using Thought Depth and Knowledge Breadth, provides 18 language versions, and uses expert-calibrated translations.

  • Results

    54.6 and 52.2 are the benchmark scores of Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, with about 40% accuracy at the highest level and performance varying by up to 10 points across languages.

  • Takeaways & Limitations

    PolyMath reveals multilingual reasoning challenges involving cross-language performance, input-output language consistency, thinking length, and potential benefits from explicit language control.

Abstract

from arXiv · show

In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs. We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning: (1) Reasoning performance varies widely across languages for current LLMs; (2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance; (3) The thinking length differs significantly by language for current LLMs. Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs.

1 Introduction

PolyMath addresses the lack of challenging multilingual benchmarks for studying mathematical reasoning by spanning four difficulty levels, 18 languages, and high-quality expert-calibrated translations. Evaluations reveal substantial cross-language performance variation, inconsistent output languages, and language-dependent thinking lengths, while explicit language control can affect performance.

  • Motivation: Existing multilingual mathematical datasets are often too simple, while challenging benchmarks are mostly English-only, leaving multilingual reasoning underexplored.This motivates creating a challenging multilingual benchmark for current reasoning LLMs.
  • Benchmark Design: PolyMath spans four difficulty levels from K-12 to Olympiad and advanced frontier mathematics, with 125 problems per language at each level.The levels use Thought Depth and Knowledge Breadth to cover broad mathematical difficulty and domains.
  • Benchmark Design: Each problem is available in 18 parallel language versions covering major language families and over 75% of the world’s native speakers.The benchmark includes both high-resource and low-resource languages.
  • Benchmark Design: Language experts calibrate PolyMath translations rather than relying directly on LLM-generated outputs, preserving mathematical terminology and logical clarity.This addresses the specialized accuracy demands of mathematical translation.
  • Findings: 54.6 and 52.2 are the benchmark scores of Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, while highest-level accuracy is about 40%.Performance also varies by up to 10 points across languages, even at low accuracy settings.
  • Findings: Reasoning models show lower input-output language consistency, language-dependent thinking lengths, and potential performance effects from explicit language control.Forcing reasoning in English improves performance, whereas forcing adherence to the input language typically performs worse.

2 Construction of PolyMath Benchmark

PolyMath is constructed as a multilingual benchmark spanning four difficulty levels, 18 languages, and diverse mathematical domains. Its human-verified translation and metadata processes aim to preserve problem meaning while enabling cross-lingual evaluation.

  • Difficulty design: PolyMath partitions mathematical difficulty using Thought Depth and Knowledge Breadth across four levels.Existing learning-stage and problem-source categories are used within these broader levels.
  • Data collection: 500 problems are distributed evenly across four levels, with 125 problems per level.Sources range from MGSM and university or competition exams to Olympiads and frontier mathematics.
  • Quality control: Human experts assign each problem’s level, with two additional experts confirming every tag.This replaces LLM-based difficulty tagging and is intended to improve assessment professionalism and accuracy.
  • Language coverage: 18 parallel language versions cover major language families and over 75% of the world’s native speakers.English problems are translated into the other language versions after collection.
  • Translation annotation: The translation pipeline combines GPT-4o pre-translation, expert term extraction, and language-expert calibration.Calibration checks terminology, formulas, and meaning across all 18 parallel versions.
  • Benchmark statistics: PolyMath maintains diverse domains at every level, while semantic embeddings separate low-level problems most clearly from the other three levels.The remaining levels are distinguishable but partially overlap in the T-SNE visualization.

3 Experiments

The experiments evaluate non-reasoning and reasoning LLMs across four levels and 18 languages using accuracy and difficulty-weighted benchmark scores. Results show stronger high-level performance and greater difficulty stability for reasoning models, alongside substantial language gaps.

  • Setup: 16 models are evaluated across non-reasoning and reasoning categories using four levels and 18 languages.The setup uses language-specific answer-extraction prompts and reports ACC for each model and language.
  • Absolute performances: Qwen-3-235B-A22B-Thinking consistently leads across levels, while reasoning models approach 40% ACC at the top level.Some non-reasoning models nearly fail at the top level, whereas reasoning models can underperform them at low difficulty.
  • Performances across levels: All models decline as difficulty rises, but reasoning models degrade more gradually than non-reasoning models.Examples include Gemini-2.0-flash-thinking falling 87.3 →59.4 →42.9 →22.3 and ChatGPT-4o-latest falling 91.6 →40.8 →20.4 →13.7.
  • Language gaps: Higher difficulty levels retain language ranges around 10%, with Qwen-QwQ-32B reaching nearly 20%.These gaps exceed the 0.5–1.5% ACC fluctuations observed across repeated runs.
  • Language gaps: Stronger reasoning models tend to exhibit larger language gaps at the same difficulty level.Qwen-QwQ-32B consistently has the highest standard deviation and range across the three higher levels.

4 Further Analysis

The analysis examines language consistency, language control, and thinking length across multilingual reasoning tasks. Reasoning models show lower and more variable language consistency, while explicit English control can improve performance and thinking-length patterns differ across model types and languages.

  • 4.1 Input-Output Language Consistency: Non-reasoning LLMs exhibit near-perfect input-output language consistency, whereas reasoning LLMs consistently show lower consistency, especially during thinking.Qwen-QwQ-32B and Deepseek-R1-671B remain around 40% for both thinking and answers, while Claude-3.7-sonnet-thinking reaches about 90% in answers.
  • 4.1 Input-Output Language Consistency: For Qwen-QwQ-32B and Deepseek-R1-671B, lower language consistency is strongly negatively correlated with accuracy, with inconsistent responses predominantly using English or Chinese.The results suggest that the language used for thinking and answering can influence reasoning performance to some extent.
  • 4.2 Language Control: Forcing English responses yields the best overall performance and minimizes language disparities, improving Qwen-QwQ-32B scores in Bengali from 36.8 to 39.0 and Portuguese from 39.4 to 47.5.Forcing the response language to match the query instead produces the poorest performance, including Arabic declining from 40.0 to 25.9 and Bengali from 36.8 to 32.8.
  • 4.2 Language Control: When forced to respond in English, reasoning models follow the instruction in over 90% of thinking parts and nearly 100% of answer parts, although English thinking remains low for Chinese and Russian inputs.Qwen-QwQ-32B produces English thinking in 33.2% of Chinese contexts and 36.2% of Russian contexts.
  • 4.3 Thinking Length Across Languages: Reasoning models increase thinking length with difficulty, ranging at the top level from 6k–8k tokens for OpenAI-o1-mini and o3-mini to 20k–30k for Claude-3.7-sonnet-thinking.Non-reasoning models typically stabilize between 1k and 2k tokens across the last three levels.
  • 4.3 Thinking Length Across Languages: Reasoning LLMs show relatively stable thinking lengths across languages, while non-reasoning models exhibit much larger cross-language differences.At the top level, reasoning-model maximum-to-minimum ratios range from 1.21 to 1.45, compared with 6.30 for Deepseek-v3, 7.92 for Llama3.3-70B-Instruct, and 2.32 for Qwen-2.5-Max.
  • 4.3 Thinking Length Across Languages: As difficulty increases, thinking-length and performance correlations become less evident across languages, and longer thinking does not consistently improve reasoning.Only some models show relatively strong positive correlations, while most show little, no, or negative correlation.

5 Related Work

The related-work discussion positions PolyMath against existing multilingual mathematical benchmarks and multilingual reasoning research. It argues that PolyMath offers greater difficulty, quality, scale, and multidimensional coverage for evaluating current reasoning LLMs.

  • Multilingual Mathematical Reasoning Benchmarks: Existing multilingual benchmarks are limited because MGSM and MSVAMP are too easy, while MT-AIME relies on full LLM translation and limited per-language samples.These limitations reduce their suitability for evaluating advanced reasoning LLMs reliably.
  • Multilingual Mathematical Reasoning Benchmarks: PolyMath matches the reasoning level of advanced LLMs while maintaining high data quality and scale, making it a more challenging and reliable multilingual benchmark.Table 10 compares PolyMath with prior benchmarks across multiple dimensions.
  • Multilingual Research Challenges in Current LLMs: Multilingual reasoning remains underexplored, with language mixing, language alignment, and low-resource language support identified as continuing challenges.PolyMath’s experiments empirically support these challenges and position the benchmark as a tool for driving progress.
  • Multilingual Research Challenges in Current LLMs: PolyMath reports substantial performance variation across languages, input-output language inconsistencies, and differing thinking-length patterns among reasoning LLMs.The paper also finds potential performance benefits from explicit output-language control.

A.1 Annotator Background

The annotator team combines linguistically trained second-language experts, native mathematical speakers, professional translators, and language-specialist graduate students. The staffing strategy varies by language according to collaborator availability and expertise.

  • Annotator Background: The project primarily collaborates with second-language experts holding linguistics degrees and specialized translation experience, while using native speakers when low-resource languages are harder to staff.This approach prioritizes professional translation expertise alongside native-language knowledge.
  • Annotator Background: The first author performs mathematical-term extraction and Chinese calibration.The author is described as a Chinese computer-science Ph.D. candidate with strong mathematical and competition experience.
  • Annotator Background: Professional translation teams and language experts support Arabic, French, Italian, Japanese, Korean, Spanish, Thai, and Vietnamese.The experts hold degrees in their languages and experience with large-scale scientific or literary translations.
  • Annotator Background: Graduate students specializing in German, Indonesian, Portuguese, and Russian provide language-specific calibration.The paper identifies contributors from universities including Tongji, Beijing Language and Culture University, Beijing Foreign Studies University, and Peking University.
  • Annotator Background: Native speakers with mathematical backgrounds directly annotate Bengali, Malay, Swahili, and Telugu.These languages receive dedicated native-speaker calibration.

A.2 Annotation Guidance

The annotation guidance focuses on accurate terminology, fluent logical expression, exact formula migration, and complete alignment with the English source. Annotators also flag translation and fluency issues for later analysis.

  • Annotation Guidance: Nested conditions should be reorganized into natural, fluent sentences so the problem’s logic is easy to understand.Annotators are instructed to prioritize readable expression rather than literal but awkward ordering.
  • Annotation Guidance: Mathematical formulas inside supported environments must remain unchanged during translation, with corrections based on the English version when migration errors occur.The guidance explicitly covers environments such as $$, \[ \], \( \), and [asy] [/asy].
  • Annotation Guidance: Translated text must contain all information from the English source without redundant additions.The guidance treats complete source-target alignment as a required quality criterion.
  • Annotation Guidance: Annotators may flag term errors, fluency issues, and other remarks for statistical analysis.These fields are optional and accompany the modified translations.
  • Annotation Guidance: Annotators must ensure mathematical terms are accurate, logic is smooth and concise, and formula blocks are migrated exactly.The platform provides the English source, terms requiring attention, and calibrated Chinese translations for some annotators.

B.1 Specific Data Source

PolyMath documents its data sources and multilingual metadata in dedicated tables.

  • Table 11 lists the data sources for each PolyMath difficulty level.
  • Table 12 records PolyMath metadata across all languages.

B.3 Data Domain

PolyMath uses level-specific domain classification: low-level questions require no classification, medium-level questions use specific knowledge-point subdomains, and high- and top-level questions remain broadly categorized.

  • Domain classification by difficulty: Low-level questions come from K-12 mathematics, so no domain classification is applied.
  • Domain classification by difficulty: Medium-level questions, mainly from exams and exercises, are subdivided into specific knowledge-point subdomains.
  • Domain classification by difficulty: High- and top-level questions cover broad, diverse knowledge points and are not further subdivided into subdomains.
  • Domain statistics: Tables 13 and 14 provide detailed domain statistics for the medium, high, and top levels.

C Experimental Settings

The experimental settings document the evaluated models and the instruction prompts used in the main experiments, including prompts for controlling the output language.

  • Experimental resources: Table 15 lists paper citations and source URLs for all models used in the experiments.
  • Instruction prompts: Figure 8 presents instruction prompts appended after the input query in the main experiments.
  • Language control: Figure 9 presents prompts used for language control.

C.4 Sampling Details

To reduce decoding instability, the experiments use sampling-based evaluation with 16 repeated trials per model, language, and difficulty level, while reporting accuracy variation across runs.

  • Decoding procedure: Sampling-based decoding is used because greedy decoding can cause instability and repetition in reasoning LLMs.
  • Accuracy computation: Average accuracy preserves 0.8-point granularity by averaging correct-answer counts, rounding to the nearest integer, and dividing by the total number of samples.
  • Variation analysis: Most accuracy standard deviations fall between 0.5 and 1.5, corresponding to variation below two questions, or 0.8% per problem.
  • Variation analysis: Standard deviations are reported separately for low, medium, high, and top levels in Tables 16–19.

D.1 Consistent Input-Output Language

The supplied passages contain multilingual mathematical solutions and explicit language-pair labels, but they do not directly report benchmark-level consistency measurements. They instead show language-specific reasoning traces, intermediate derivations, and final-answer claims across several problems.

  • Language annotations: Language-pair annotations explicitly mark Swahili, Chinese, Korean, and French problem-answer combinations.These labels appear alongside separate mathematical solution traces.
  • Swahili reasoning: The Swahili traces derive optimization conditions for f(x)=(x+a)(x+b), including monotonicity and equality-based maximization.The passages state that positive a,b make f strictly increasing for x⩾0 and motivate equal xi values.
  • Chinese reasoning: The Chinese traces apply a Borcea–Voisin construction to compute the maximum h1,1 of a smooth threefold from a K3 surface and a genus-2 curve.The supplied passages introduce the involutions, quotient construction, and Hodge-number calculation setup, but do not provide the final maximum.
  • Korean reasoning: The Korean traces test whether quartic polynomials over finite fields satisfy a surjectivity condition for p=2, 3, and 5.The supplied examples conclude that the condition fails for each of those three primes.
  • Japanese reasoning: The Japanese traces analyze a 2024-row, 2023-column board and conclude that one unused column can reach the final row.They describe 2022 monsters occupying distinct rows and at most one per column, leaving one column without a monster.

D.2 Inconsistent Input-Output Language

PolyMath examples show that reasoning models may use a different language for thinking than for the input or answer. The supplied cases include Korean input and answer with English thinking, Japanese input and answer with Chinese thinking, and Spanish input with English thinking and answer.

  • Japanese input and answer can be paired with Chinese thinking.
  • Spanish input can be followed by English thinking and an English answer.
  • Korean input and answer can be paired with English thinking.
Loading 2504.18428v4…