Source-linked AI summary

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, Martin Vechev

arXiv:2605.00674v2cs.CL

TL;DR

Static benchmarks are often narrow, quickly saturated, and rarely updated, limiting reliable tracking of mathematical reasoning progress. The paper develops MathArena as a continuously maintained platform spanning diverse mathematical tasks and finds rapid progress, including strong frontier-model performance. Its scope remains limited because several important aspects of real mathematical work are not yet evaluated.

  • Problem

    Static benchmarks are narrow, quickly saturated, and rarely maintained, making reliable comparison and progress tracking difficult.

  • Method

    The paper expands MathArena into a continuously maintained platform covering diverse benchmarks, with a clear evaluation protocol and regular benchmark and model updates.

  • Results

    84% average performance across MathArena benchmarks, up from 45% a year ago, with frontier models saturating easier competitions while research-level benchmarks retain headroom.

  • Takeaways & Limitations

    MathArena provides a comprehensive picture of mathematical reasoning progress and highlights nuanced capability differences across benchmarks.

  • Takeaways & Limitations

    MathArena currently omits interactive problem-solving, mathematical work beyond problem solving, and broader tool use.

Abstract

from arXiv · show

Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often narrow in scope, quickly saturated, and rarely updated. This makes it hard to compare models reliably and track progress over time. Instead, we need evaluation platforms: continuously maintained systems that run, aggregate, and analyze evaluations across many benchmarks to give a comprehensive picture of model performance within a broad domain. In this work, we build on the original MathArena benchmark by substantially broadening its scope from final-answer olympiad problems to a continuously maintained evaluation platform for mathematical reasoning with LLMs. MathArena now covers a much wider range of tasks, including proof-based competitions, research-level arXiv problems, and formal proof generation in Lean. Additionally, we maintain a clear evaluation protocol for all models and regularly design new benchmarks as model capabilities improve to ensure that MathArena remains challenging. Notably, the strongest model, GPT-5.5, now reaches 98% on the 2026 USA Math Olympiad and 74% on research-level questions, showing that frontier models can now comfortably solve extremely challenging mathematical problems. This highlights the importance of continuously maintained evaluation platforms like MathArena to track the rapid progress of LLMs in mathematical reasoning.

1 Introduction

Static benchmarks are too narrow, quickly saturated, and rarely updated to track rapidly advancing mathematical reasoning. MathArena addresses this gap as a continuously maintained platform spanning diverse tasks and reporting rapid model progress.

  • Accurate, up-to-date evaluations are needed to validate performance claims, track progress, and guide practitioners toward state-of-the-art mathematical models.
  • Traditional benchmarks provide one-time snapshots, cover limited capabilities, become saturated, and leave practitioners relying on self-reported or informal evidence.
  • Evaluation platforms continuously aggregate and analyze diverse benchmarks, adapt as models improve, and provide interfaces for detailed result analysis.
  • MathArena expanded from a competition benchmark into a platform covering proof-based competitions, research questions, false-claim rejection, and Lean theorem proving.
  • 84% average performance across MathArena benchmarks, up from 45% a year ago, accompanies saturation of easier competitions and remaining headroom in research-level tasks.
  • The work argues for evaluation platforms, extends MathArena, and analyzes frontier-model performance across its benchmarks.

2 Related Work

Prior mathematical evaluations span final answers, proofs, formal Lean verification, research-level problems, and additional dimensions such as reliability and interaction. MathArena differs by combining broad mathematical coverage with an open, dynamic evaluation platform.

  • Final-answer benchmarks: Final-answer benchmarks are easy to verify and scale, but they test a narrow subset of mathematics and can be exploited.
  • Proof-based benchmarks: Proof-based benchmarks broaden mathematical evaluation but are harder to assess and often require expert human judges.
  • Formal benchmarks: Formal benchmarks use proof assistants such as Lean for automatically checkable proof-generation evaluation, including recent research-level mathematics.
  • Research-level benchmarks: Research-level benchmarks evaluate mathematical problems near the frontier of human knowledge, commonly using final-answer tasks and sometimes proofs.
  • Other dimensions of mathematical evaluation: Other evaluations address reliability, sycophancy, visual reasoning, interaction, usefulness, and metrics beyond raw accuracy.
  • Evaluation platforms: MathArena differs from related platforms by focusing specifically on mathematics while remaining open, transparent, and comprehensive.

3 Evaluation Platforms

Evaluation platforms replace static, narrow benchmark releases with continuously maintained systems designed to provide current, comprehensive views of model performance. Their usefulness depends on adaptation, transparency, faithful measurement, and real-world validity.

  • Static benchmark releases become uninformative as models saturate tasks, results grow stale, and no single benchmark covers all practically relevant capabilities.
  • Benchmarks: A benchmark is a fixed sample collection paired with an evaluation protocol that assigns model scores.
  • Evaluation platforms: An evaluation platform is continuously maintained to provide an up-to-date, comprehensive view of model performance within a broad domain.
  • Key properties: Platforms differ from benchmarks through changing tasks, ongoing reruns and analysis, and a public interface exposing results, outputs, and metadata.
  • Trustworthy platforms: Open and transparent platforms should expose problems, outputs, metadata, and code, while third-party operation helps avoid conflicts of interest.
  • Trustworthy platforms: Faithful measurement requires implementation choices that accurately reflect model capabilities because noise, hyperparameters, and tool configuration can alter measured performance.
  • Trustworthy platforms: Construct validity requires realistic tasks and deployment conditions, but can conflict with reliability, scalability, and openness.

4 MATHARENA

MATHARENA functions as an evolving evaluation platform that broadens mathematical reasoning assessment beyond its original benchmark. It combines changing tasks, diverse evaluation methods, and research-level benchmarks while supporting transparent analysis and model comparison.

  • Platform design: MATHARENA is organized as an evaluation platform whose benchmark set changes as tasks saturate or model capabilities shift.Older benchmarks can be retired, new benchmarks introduced, and benchmark families updated continuously.
  • Platform design: The platform emphasizes openness and faithful measurement by publicly releasing most outputs and results while reconciling provider-reported discrepancies before publishing its own results.Project Euler answers are withheld to comply with competition rules.
  • Platform design: Overall expected performance is estimated with an item response theory approach that imputes missing benchmark results instead of evaluating every model on every task.This avoids the prohibitive cost of running each model across all benchmarks.
  • Benchmark coverage: MATHARENA combines final-answer competitions, proof-based competitions, and research-level benchmarks to measure textual, visual, proof, reliability, and formal-proving capabilities.Final-answer competitions use automatic verification, while proof-based benchmarks use official submission, human evaluation, or semi-automatic evaluation.
  • Proof-based competitions: Proof-based benchmarks are useful for tracking overall field progress but are too limited in size to reliably compare individual models.Evaluation uses official competition grading, expert internal grading, or a semi-automatic LLM-jury pipeline followed by manual review.
  • Research benchmarks: The research-level suite includes ArXivMath for final-answer problems, BrokenArXiv for rejecting false claims, and ArXivLean for formal proof generation.These benchmarks are updated through an automated pipeline drawing on recent arXiv papers, with human review for quality.

5 Evaluation

MathArena’s evaluation shows rapid gains across diverse mathematical tasks, while frontier performance now saturates easier final-answer competitions and remains uneven on visual, research-level, false-claim, and Lean benchmarks.

  • Overall performance: GPT-5.5 achieves the highest overall expected performance, ranks first on every benchmark, and leads the next-best model by 34% on BrokenArXiv.
  • Model comparisons: DeepSeek-v4-Pro is the strongest open model but trails GPT-5.5 by 20%, with larger closed–open gaps on BrokenArXiv, ArXivLean, and USAMO 2026.
  • Performance over time: 84% average performance across MathArena benchmarks now exceeds 45% a year earlier.The earlier average was achieved by o4-mini in May 2025, while the current average is achieved by GPT-5.5.
  • Final-answer competitions: Top models reach 97% on both AIME and HMMT, leaving these final-answer benchmarks with little ability to distinguish frontier models.Apex and Apex Shortlist retain limited headroom at 80% and 94%, respectively.
  • Proof-based benchmarks: Models achieve strong proof-based results, including 95% for GPT-5.4 on USAMO 2026 and 86% for DeepSeek-v3.2-Speciale with DeepSeekMath-V2 on Putnam 2025.The latter score would place the system among the top three human participants, compared with just below 5% for the best pre-USAMO 2025 model.
  • Research-level and formal reasoning: Research-level performance reaches 74% on ArXivMath, but models frequently endorse false statements and achieve only 17% on the highly demanding ArXivLean benchmark.On BrokenArXiv, only GPT-5.5 reaches a strong 72% score, while many other models score below 20%.

6 Limitations

MATHARENA does not cover several important aspects of real mathematical work, making scope expansion a central limitation. In particular, its current evaluation is non-interactive and focused on problem solving.

  • 6 Limitations: MATHARENA’s main limitation is that important aspects of real mathematical work remain outside its scope.The authors hope to address these gaps in future iterations.
  • Interactive problem-solving: MATHARENA evaluates models only in non-interactive settings, missing follow-up questions, user feedback, and workflows built on intermediate outputs.
  • Mathematical work beyond problem solving: MATHARENA focuses on problem-solving ability, omitting other important mathematical work.
  • Additional limitations: The paper discusses additional limitations in Appendix D.

7 Conclusion

The paper presents MATHARENA as a comprehensive, evolving platform covering diverse mathematical skills and maintained through regular updates. Its adoption in frontier model reports illustrates use beyond the paper’s reported results.

  • 7 Conclusion: MATHARENA evaluates LLM mathematical capabilities across problem solving, proof writing, and research-level mathematics.
  • 7 Conclusion: MATHARENA is regularly updated with new benchmarks and models to provide an accurate picture of performance and progress.
  • Adoption: MATHARENA has been used in frontier model reports covering competitions, proof-based olympiads, hard problem sets, and recent competition benchmarks.

A Benchmark Details

The benchmark details specify how MATHARENA constructs tasks, executes models, aggregates incomplete results, estimates costs, reports uncertainty, and designs metrics. These choices aim to make measurements comparable and informative across diverse benchmarks.

  • Benchmark construction: Each benchmark documents task selection, prompt design, and evaluation metrics, with further execution details in the appendices.
  • Model execution: Models use provider-recommended hyperparameters and maximum API token limits, with provider-result reproduction checks and retries for failed requests.
  • Expected performance: IRT estimates expected average performance across all questions when models are not evaluated on every benchmark.The approach addresses prohibitive evaluation costs and uses observed results from other models to impute missing results.
  • Expected cost: Expected cost averages per-problem costs across non-deprecated, non-Euler competitions while weighting benchmarks by problem count.
  • Expected cost: Missing costs are predicted with a log-additive model using model-specific terms and benchmark-specific offsets.The model is fit by least squares and predictions are transformed as exp(µm + βb).
  • Confidence intervals: 95% confidence intervals are reported for model scores using a normal approximation that simulations find better calibrated than more complex alternatives.
  • Metric design: Metric design prioritizes unbiasedness and accuracy so scores reflect underlying performance without systematically favoring models or approaches.The paper considers unbiasedness more important because biased metrics can misrank models.

A.2 Final-Answer Evaluation Protocol

MATHARENA uses automatic answer extraction and verification for final-answer tasks, while its datasets combine competition sources with manual checking and targeted construction. Some continuously updated problems require manual correctness verification.

  • Final-answer evaluation: Final-answer questions have single unambiguous answers, enabling automatic evaluation with a SymPy-based parser.Incorrect answer formatting can still produce false negatives.
  • AIME 2026: AIME 2026 contains two sets of 15 competition problems, each requiring an integer answer from 0 to 999.
  • HMMT Feb 2026: HMMT February 2026 problems were OCR-processed, manually verified, and filtered to exclude proof or construction tasks, leaving 33 problems.
  • Kangaroo 2025: The selected Kangaroo 2025 version was translated into English and rendered as single images containing text, visuals, and answer options.The final dataset contains 168 problems across six grade groups.
  • Apex and Apex Shortlist: Apex problems are selected from difficult recent competitions using repeated failures by specified frontier models as inclusion criteria.This process produced 12 Apex problems and 48 Apex Shortlist problems, followed by manual verification.
  • Project Euler: Project Euler adds its latest problem weekly, with model-generated solutions manually submitted to verify correctness because answers are not public.

A.5 Proof-Based Competitions

MathArena evaluates proof-based competitions using detailed rubrics, human review, organizer grading, and an LLM jury, while documenting grading limitations and pipeline corrections.

  • IMC 2025: IMC 2025 contains 10 problems graded on a 0–10 scale, with each grader primarily responsible for five problems and reviewing the other grader’s decisions.The benchmark’s stated total maximum is 70 points.
  • Putnam 2025: Putnam 2025 comprises 12 problems scored from 0 to 10, with model solutions submitted as anonymized PDFs for organizer grading.Model submissions were included in the organizers’ second grading round using the same rubrics and graders as human solutions.
  • Miklós Schweitzer 2025: Miklós Schweitzer 2025 follows the Putnam evaluation procedure across 10 problems worth 100 total points, with organizers not informed that some submissions were generated by models.This preserves the ordinary submission format used for human-written solutions.
  • USAMO 2026: USAMO 2026 uses LLM-generated grading rubrics that were manually revised after initial versions proved too rigid for valid solution variations.Small rubric phrasing issues also caused substantial judge disagreement.
  • USAMO 2026: The grading pipeline combines a problem-specific rubric with three LLM judges, reconciliation of disagreements, and conservative score aggregation.Judges are GPT-5.4, Gemini 3.1 Pro Preview, and Claude Opus 4.6.
  • USAMO 2026: USAMO solutions received careful hand review after automated grading, including special validation procedures for computation-heavy geometry solutions.Reviewers used explanations from the LLM jury to accelerate validation.

A.6.1 Extraction Methodology

MathArena extracts research-level questions from recent arXiv papers through automated generation and filtering followed by manual review, enabling frequent updates and contamination mitigation.

  • Continuous updates: Monthly ArXivMath and BrokenArXiv releases, plus quarterly ArXivLean updates, add recent problems while reducing contamination risk.The process also permits refinement as model capabilities and extraction insights change.
  • Pipeline: The extraction pipeline has three stages: candidate generation, automated filtering, and author-led manual review.The first two stages use Gemini 3.1 Pro Preview, while the final stage is performed by the authors.
  • Filtering: Filtering removes questions that are not self-contained, omit necessary technical context, or can be guessed from prior cited work.Questions may also be discarded when the source paper explicitly reports AI use or appears insufficiently expert-verified.
  • Manual review: Manual review checks that each surviving question is well-defined, non-trivial, interesting, and correctly filtered, removing doubtful cases.Reviewers use an interface containing relevant source-paper and question information.
  • Benchmark-specific outputs: After filtering and review, ArXivMath retains around 30 questions each month, while BrokenArXiv retains around 50.The two benchmarks use related filtering and review procedures with setting-specific adjustments.
  • BrokenArXiv: BrokenArXiv generates plausible false statements directly contradicting extracted true claims, then evaluates whether models resist proving them.Its simple prompt asks models to try proving the perturbed statement, and scores behavior through LLM judging.
  • ArXivLean: ArXivLean retains theorem-like claims that appear expressible in Lean 4 with Mathlib, then translates and verifies them using GPT-5.4 and Lean-oriented tools.The formalization workflow includes Lean execution, declaration search, and related support tools.

B Examples of Underelicitation

MathArena’s maintenance history shows that evaluation results can depend strongly on elicitation choices, while reliability experiments expose persistent false-premise acceptance.

  • Tool limits: DeepSeek-v3.2’s Project Euler performance rose from 25% to 50% after the tool-call limit increased from 20 to 200.The original limit measured performance under a constrained tool budget rather than unrestricted problem-solving ability.
  • Tool limits: The Project Euler case shows that equal resources can still produce construct-validity problems when the resource constraint is unrealistic for the task.Open models were more affected because they frequently approached the original 20-call limit.
  • Pipeline effects: Other underelicitation sources included incorrect timing feedback, prompt wording, formatting failures, API configuration, and strict request time limits.These issues respectively altered model behavior, parsing, reasoning effort, or completion on difficult problems.
  • Maintenance: The platform’s history includes benchmark errors that required extensive human validation before correction.Some issues were temporarily live on the leaderboard, while most were caught before publication.
  • BrokenArXiv variants: BrokenArXiv tests whether models reject plausible false statements rather than produce invalid proofs, with variants for disproof and self-correction.The benchmark targets reliability alongside mathematical capability.
  • Disproof variant: 71.0% overall accuracy was achieved by Gemini 3.1 Pro Preview, compared with 56.3% for GLM 5.1 and 46.7% for Step 3.5 Flash.All three models classified original true statements more accurately than perturbed false statements.
  • Self-correction: Iterative self-correction improved results, but the best final accuracy remained only 34.8%, and GLM 5.1 improved by just 0.4%.Models often repaired flawed proofs instead of rejecting the underlying false statement.

C.2 Expected-Performance Robustness

MathArena’s aggregate analyses test ranking stability, interval calibration, and inference-time scaffolds, finding robust rankings, near-ideal pooled-normal coverage, and differing scaffold scaling.

  • Ranking robustness: Spearman rank correlation was 0.99 between the default two-parameter model and a one-parameter discrimination variant.This indicates nearly identical model rankings under the tested latent-variable specifications.
  • Ranking robustness: Nine of the top 10 models remained in the top 10 across all leave-one-family-out ablations, with a largest rank shift of 5 positions.The reported rankings were therefore not driven by one benchmark family or the exact imputation model.
  • Confidence-interval calibration: Calibration was estimated by simulating 1000 fresh four-sample reevaluations per eligible model-benchmark pair using fitted per-question success probabilities as ground truth.This approximates repeated evaluation while avoiding the cost of actually rerunning every evaluation.
  • Confidence-interval calibration: At the 95% level, pooled-normal interval coverage was 95.99%, closest to ideal calibration among the five tested approaches.Wilson reached 98.83%, question plug-in 87.75%, Jeffreys-smoothed plug-in 98.54%, and question-SD 96.80%.
  • Agentic scaffolds: RSA achieved the strongest scaffold performance and scaled more effectively with reasoning steps than selfcheck, Nomos, or DeepSeekMath-V2 on the tested MathArena benchmarks.Selfcheck and Nomos produced modest gains that saturated earlier, while DeepSeekMath-V2 underperformed on Gemini 3 Flash.

D Additional Limitations

The paper identifies scope and validation limitations in its research-level benchmarks and LLM-jury evaluation. In particular, BrokenArXiv is a proxy for research reliability, tool use is disabled in two benchmarks, and jury ablations are outside scope.

  • BrokenArXiv measures responses to potentially flawed proof requests rather than mathematical reliability in research directly.The paper treats this setting as a meaningful practical proxy while identifying more direct reliability benchmarks as future work.
  • More direct research-reliability benchmarks likely require accurate proof-correctness judgments, which remain difficult to obtain.
  • ArXivMath and BrokenArXiv currently disable tools because web search would expose source problems and code execution produced improvements below 5% in initial experiments.The paper notes that added tool use substantially increases evaluation cost and may be reconsidered later.
  • The study does not include detailed ablations of the LLM jury’s necessity or mitigation of issues such as self-preference bias.Such comparisons would require human ground-truth judgments on current state-of-the-art models and are considered outside the present scope.

E MATHARENA Evaluation Interface

The MATHARENA interface connects high-level benchmark views with detailed model, comparison, problem-trace, failure, and ongoing-analysis views. Its supporting prompts and evaluation workflows emphasize inspection, faithful rewriting, formalization constraints, and structured proof grading.

  • Interface workflow: The platform moves from homepage and benchmark-index summaries to detailed leaderboards, model pages, comparisons, traces, surprising failures, and blog analyses.These views support qualitative inspection, reproducibility, and analysis of changing model behavior across tasks and over time.
  • Interface workflow: The homepage combines a live leaderboard, benchmark tabs, filters, and per-problem correctness summaries.
  • Interface workflow: Benchmark pages link competitions to scores, datasets, outputs, and concise task notes, while leaderboard and model views add evaluation metadata and execution details.The detailed leaderboard includes confidence intervals, token counts, retries, runtime, and openness metadata; model pages aggregate performance, expected rank, pricing, and settings.
  • Interface workflow: Comparison views place two models side by side across benchmarks to expose differences in accuracy and cost.
  • Interface workflow: Problem pages expose prompts, verified answers, run-by-run outputs, and full model traces, while surprising-traces views surface high-loss failures.
  • Interface workflow: The interface also supports longitudinal analysis through blog posts documenting benchmark launches and qualitative analyses over time.

F.4 Proof-Based Competitions

The proof-based competition evaluation workflow prioritizes mathematical validity, problem constraints, and rubric-aligned scoring while allowing valid alternative approaches. It requires detailed assessments, conservative treatment of unsupported steps, and especially strict verification of computational geometry solutions.

  • Scoring principles: Proof grading prioritizes mathematical validity, problem constraints, and advisory mapping to the marking scheme.
  • Scoring principles: Valid alternative methods are mapped to equivalent rubric checkpoints, with zero-credit items and deductions applied only when the underlying issue occurs.
  • Scoring principles: A wrong final answer where uniqueness is required receives only partial credit, while subjective judgments about style, elegance, or clarity are excluded.
  • Scoring principles: Intermediate claims receive credit only when adequately justified, while plausible but under-justified steps receive conservative partial credit.
  • Output requirements: The evaluator must identify errors, assign an integer score from 0 to 7, and explain earned checkpoints, deductions, and the score calculation in a detailed assessment.
  • Geometry-specific grading: Bashed geometry solutions are checked especially strictly, with extensive verification of calculations and a score of 0 for multiple minor mistakes or any major invalidating mistake.
  • Output requirements: The evaluation output is constrained to well-formed XML containing a 0–7 points value, detailed assessment, and errors list.
Loading 2605.00674v2…