Source-linked AI summary

ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions

Guangxiang Zhao, Qilong Shi, Xusen Xiao, Wenpu Liu, Yaoming Li, Linfeng Hao, Shuyang Hou, Zijian Guo, Xinrui Zhang, Yuntian Zhao, Zhengyang Wang, Wenrui Liu, Yuhan Wu, Tong Yang, Lin Sun, Xiangzheng Zhang

arXiv:2608.30517v1cs.AIcs.CL

TL;DR

Saturated and contamination-prone benchmarks make genuine scientific reasoning difficult to measure, especially for long, visual, partially correct problems. ScienceArena builds an expert-audited benchmark from thirteen science competitions, calibrates LLM judges against medalist grading, and evaluates fourteen recent models with interleaved solving. Strong models reach medal-equivalent scores on several public international exams, while chemistry and long-horizon consistency remain bottlenecks.

  • Problem

    Scientific reasoning remains difficult to measure because saturated benchmarks and short-answer formats inadequately test long, visual, multi-step, partially correct solutions.

  • Method

    ScienceArena converts official materials from thirteen science competitions into medalist-audited structured items and calibrates an LLM-as-judge system against medalist grading.

  • Results

    The strongest of fourteen evaluated LLMs obtain medal-equivalent rubric scores on several public international exams, with chemistry and long-horizon consistency remaining bottlenecks.

  • Takeaways & Limitations

    Medalist grading comments identify visual grounding, structure fidelity, and global problem control as persistent diagnostic boundaries beyond missing terminology.

  • Takeaways & Limitations

    Results cover theory components, may not generalize across scientific curricula, and rely on costly expert calibration currently focused on IPhO and IChO.

Abstract

from arXiv · show

Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympiad medalists. To scale evaluation beyond costly human grading, we calibrate LLM-as-judge against medalist ground truth on archived answers from five models across IPhO and IChO; two strong judges stay within one point of expert total scores. Medalist notes show that failures often stem from visual grounding, structure fidelity, and global problem control rather than missing terminology. Evaluating fourteen recent LLMs with interleaved solving, we find that top models obtain medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain key bottlenecks. We provide an interactive \href{https://science-arena.onrender.com/}{demo}.

1 Introduction

ScienceArena addresses saturated, contamination-prone scientific reasoning evaluation with an expert-audited olympiad benchmark and calibrated rubric-based scoring. It evaluates fourteen recent LLMs using interleaved solving and identifies persistent domain-specific bottlenecks.

  • Motivation: ScienceArena targets long, visual, partially correct scientific problems that short-answer and multiple-choice benchmarks cannot fully assess.Olympiad tasks provide expert-authored problems, official solutions, and scoring schemes for evaluating derivations, interpretation, consistency, and partial credit.
  • Benchmark: The benchmark covers thirteen physics, chemistry, and biology competition datasets from 2023–2026, preserving original exam structure and rubric metadata.Official materials are converted into structured JSON with OCR and multimodal parsing, then audited by olympiad medalists.
  • Evaluation: An expert-calibrated LLM-as-judge system is developed to scale rubric-faithful grading beyond repeated human evaluation.The system is calibrated against medalist scores on archived IPhO and IChO answers from five models.
  • Evaluation: Interleaved prompting generally improves performance over solving an entire multi-part problem at once, so it is used for the main evaluation.The protocol preserves prior context while presenting subparts sequentially.
  • Findings: Across fourteen recent LLMs, the strongest systems reach medal-equivalent rubric scores on several public international exams, while chemistry and long-horizon consistency remain bottlenecks.Expert comments additionally diagnose visual grounding, structure fidelity, and unsupported intermediate steps.

2 The ScienceArena Framework

ScienceArena combines expert-audited olympiad materials with calibrated rubric-based judging and interleaved solving for scalable evaluation of open-ended scientific reasoning. Its framework preserves multi-part structure, visual evidence, official solutions, and partial-credit rubrics while validating automated scores against medalist grading.

  • Benchmark construction: ScienceArena contains thirteen public olympiad-style competitions spanning physics, chemistry, and biology, with linked turns preserving each problem’s original hierarchy.Each subproblem includes text, figure context, official answer, rubric, and maximum score.
  • Benchmark construction: A five-stage digitization pipeline converts official problems, figures, solutions, and marking schemes into structured fields audited by subject medalists.Audits check completeness, figure alignment, answer fidelity, and point totals before unified export.
  • Expert-calibrated LLM-as-judge system: The LLM-as-judge pipeline receives frozen problem content, official solutions, rubrics, and candidate responses, then returns validated sub-scores, evidence, deductions, confidence, and totals.Post-processing validates JSON, score bounds, maximum scores, and turn-level aggregation; objective-key items are scored deterministically.
  • Expert-calibrated LLM-as-judge system: Two judge models stayed within one point of medalist total scores for every archived model–exam pair across IPhO 2025 and IChO 2025.Item-level replay also showed strong correlations with expert scores after rubric packaging, structured sub-scores, and edge-case post-processing were added.
  • Solving protocols: Interleaved prompting answers one subpart at a time while retaining prior context, and it won all 16 displayed model–exam comparisons against one-shot solving.The protocol was used for main open-ended evaluations, while IBO was excluded because it is objective-key scored.

3 Experimental Results and Diagnostics

ScienceArena shows that frontier-model performance reaches medal-equivalent levels on some public olympiad exams but varies sharply by domain, task structure, and evaluation conditions. Diagnostics attribute remaining weaknesses to visual grounding, exact chemical structures, global physical-model control, and judge limitations.

  • Overall performance: 29.45/30 placed Gemini 3.1 Pro slightly above the best human theory score on IPhO 2025.
  • Overall performance: 48.95/60 put Gemini 3.1 Pro above the IChO gold threshold but below the human winner score of 57.1.
  • Overall performance: Gemini 3.1 Pro led equal-weighted averages in physics at 89.5%, chemistry at 86.3%, and biology at 95.1%.
  • Modality diagnostics: 22.90 percentage points separated native images from no visual information, with chemistry showing a 5.01-point native-image advantage.
  • Subfield diagnostics: Chemistry is bottlenecked by structure fidelity: stereochemistry scored 34.7%, versus 78.4% for physical chemistry.
  • Subfield diagnostics: Physics errors often arise from an incorrect early approximation or variable definition propagating through later subparts, despite locally coherent derivations.
  • Overall performance: 431.00/453 was Gemini 3.1 Pro’s best IBO score, while seven models exceeded the IBO human-winner reference.
  • Evaluation diagnostics: The calibrated judge agreed closely with medalist scores, but structure-heavy chemistry can require expert adjudication when prose does not uniquely specify a molecule.

4 Conclusion

ScienceArena provides a provenance-preserving, rubric-faithful benchmark for evaluating frontier LLMs on open-ended olympiad science problems. Its results show medal-equivalent performance alongside domain-specific diagnostic failures and caution that many measurements are not strict uncontaminated holdouts.

  • ScienceArena spans thirteen recent competitions and preserves multi-step interactions, original images, source provenance, and rubric-faithful scoring.
  • Recent models reach medal-equivalent rubric scores on public international exams, with Gemini 3.1 Pro matching the IPhO 2025 winner reference under text-only interleaving.
  • Physics, chemistry, and biology retain distinct bottlenecks in global regime control, exact structures and stereochemistry, and multi-statement data consistency.
  • Because many results postdate the exams, they measure current capability rather than strict uncontaminated holdout performance.

Limitations

ScienceArena’s scope and evaluation design impose several important boundaries. Results may not generalize across curricula, experimental components, contests, or diagnostic taxonomies, and some uncertainty sources remain coupled.

  • Results may not generalize to all scientific curricula despite spanning thirteen competitions.
  • The evaluation covers theory components only because experimental sections require physical apparatus and proctoring.
  • Expert grading currently calibrates the judge on IPhO and IChO, while contest-specific expert spot checks remain strongest for high-stakes comparisons elsewhere.
  • Repeated calls jointly measure solver-generation and fixed-route judge variation rather than separating them.
  • The model-assisted subfield taxonomy may reflect classifier priors, so alternative decompositions could produce different diagnostic slices.

A Additional Failure Cases

Additional cases across physics, chemistry, and biology show that strong models can fail when exact visual, structural, or experimental-design commitments are required. These failures often preserve plausible reasoning while missing the rubric-scored object.

  • Physics: Physics: GPT-5’s theory was correct, but binding to the wrong point on the correct curve reduced credit to 0.3/0.5.
  • Chemistry: Chemistry: plausible Trigonoliimine C rearrangement prose omitted the official intermediates and mismatched the sequence, receiving 0/6.
  • Biology: Biology: several strong models confused two similar guppy mate-choice controls, returning 2 4 3 1 instead of the official 2 4 1 3.

B Full Subfield Diagnostics

The diagnostic presentation supports within-domain comparisons, but its underlying table does not provide question-level classification counts. Consequently, unique-question totals are unavailable and should not be inferred from model–turn counts.

  • Question-level classification maps are unavailable, so unique-question counts are not inferred from N.

C Human Expert Notes and Ability Boundary

Human grading notes reveal that scientific performance depends on more than fluent local derivations or terminology. Models lose credit through visual grounding, incorrect global physical models, structure-inexact chemistry, and insufficiently explicit commitments.

  • Ability Boundary: Human graders distinguish local symbolic derivation from broader physical control, including graph binding, curve selection, and uncertainty propagation.In IPhO S1 B2, GPT-5’s correct theory but wrong curve point earned 0.3/0.5.
  • Ability Boundary: In IPhO S2 C1–C2, GPT-5 received 0/1 and 0/1 after missing a conserved total-force constraint and necessary case analysis.
  • Ability Boundary: Interleaving gives each subpart focused context, but cannot correct a mistaken regime choice once the conversation has anchored on it.
  • Chemistry: Chemistry grading requires exact structures: plausible reaction logic still fails when the carbon skeleton or requested molecular form is wrong.
  • Grading: Partial credit rewards explicit, checkable commitments; vague structural prose may receive zero when a unique molecule cannot be reconstructed.
  • Judge Calibration: Expert comments inform judge gates for answer format, structure specificity, and subpart-by-subpart score allocation.

C.1 Ability Axis Glossary

The ability-axis glossary translates medalist grading comments into concrete, recurring forms of scientific credit and deduction, especially across physics reasoning tasks.

  • C.1 Ability Axis Glossary: The taxonomy defines ability axes from recurring credit and deduction patterns in medalist grading.These axes are intended to represent concrete, reusable evaluation dimensions.
  • C.1 Ability Axis Glossary: Symbolic derivation measures multi-step algebraic manipulation and equation-based reasoning.
  • C.1 Ability Axis Glossary: Numeric+Units measures numerical accuracy alongside unit and scale discipline, including dimensional checks and sanity bounds.
  • C.1 Ability Axis Glossary: Global Coherence measures whether assumptions and variables remain consistent across coupled subparts and changing regimes.
  • C.1 Ability Axis Glossary: Physics abilities include symbolic derivation, numerical computation with units, approximation and limiting cases, graph reading, plotting, global coherence, data inference, and visual grounding.The axes cover both mathematical reasoning and interpretation of visual or empirical information.

C.2 Domain-specific ability profiles

Domain-specific profiles show that frontier models are strongest on local symbolic manipulation and routine numerics, while weaknesses cluster in visual grounding, structure fidelity, evidence binding, and long-horizon consistency.

  • Physics: Physics profiles emphasize symbolic derivation, numerical computation with units, approximation, graph reading, plotting, global coherence, data inference, and visual grounding.
  • Physics: In IPhO 2025, both profiled models are strongest on local symbolic manipulation and routine numerics, while errors concentrate in global coherence and visual grounding.Performance varies on tasks requiring regime switching.
  • Chemistry: In IChO 2025, Gemini 2.5 Pro is more uniformly strong on quantitative and analytical axes, whereas GPT-5 is less consistent when structure, stereochemistry, and quantitative constraints must be satisfied together.
  • Chemistry: Chemistry profiles cover quantitative physical chemistry, spectroscopy, organic structures, stereochemistry, inorganic reasoning, coordination, biochemistry, and biosynthesis tracing.
  • Cross-domain: Across the profiles, remaining weaknesses concentrate in evidence binding, representational precision, and long-horizon consistency rather than simple recall.
  • Biology: Biology profiles include molecular biology, genetics, physiology, experimental design and statistics, data or figure interpretation, and an other category.

E 2026 Evaluation and Calibration Audit

The 2026 audit verifies route completeness, modality separation, scoring consistency, and judge-study scope across the controlled evaluation and calibration pipeline.

  • Evaluation completeness: The four 2026 runs contain 153 linked turns, 2,142 sealed model–turn terminals, and complete aggregates for all 14 routes.The runs include 23 IPhO, 68 IChO, 28 NBPhO, and 34 USAPhO linked turns.
  • Modality audit: Payload audits verify direct official PNG blocks for vision routes and zero image blocks with frozen question-only staticization for text routes.The route assignments and frozen rubric criteria are shared across all four contests.
  • Uncertainty analysis: The 10,000 paired resamples sample whole problem sequences, normalize scores by explicit maxima, and equally average contest percentages.Intervals capture sequence and item sampling uncertainty, while neighboring-system rankings may remain statistically unresolved.
  • Modality audit: The transport audit confirms that the designed visual-evidence difference is not caused by unintended route leakage.
  • Judge studies: The frozen-answer judge study separates four primary vision comparisons from two modality-sensitivity text conditions rather than pooling them.
  • Calibration audit: Archived calibration replay contains 1,050 judgment rows over 525 expert rows, with scope-aligned agreement excluding a mismatched multipart Q2.3 comparison.
  • Calibration audit: The audit preserves provenance-addressed expert-score rows and supports item-level analysis of judge–expert discrepancies.

F Additional Temporal and Judge Sensitivity Checks

Additional checks find no direct memorization signal in a small temporal probe and close judge agreement on a demanding chemistry subset, while clarifying the benchmark’s scope and reproducibility boundaries.

  • Temporal probe: The temporal probe elicited 69/72 perturbed targets, with 0/72 stale original-answer substitutions and 0/72 recalls in a strict no-question control.The probe finds no direct memorization signal but cannot establish absence of training exposure.
  • Judge sensitivity: Four image-aware judges reached MAEs of 0.66 to 1.24 points on 50 IChO structure and stereochemistry rows.At least 0.5-point over-credit occurred in 5, 8, 9, and 21 rows across the judges; multi-judge agreement or expert adjudication remains the strongest safeguard for low-confidence structural cases.
  • Scope and reproducibility: ScienceArena spans thirteen physics, chemistry, and biology competitions and preserves source provenance while using theory-only evaluation where experiments require apparatus or proctoring.
  • Scope and reproducibility: Release dates are recorded descriptively, and unknown timestamps are not treated as strict time-release holdouts because timing alone cannot prove training absence.
  • Scope and reproducibility: IBO 2023 is objectively scorable, enabling automatic evaluation without judge-model mediation.
Loading 2608.30517v1…