Source-linked AI summary

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

Miles Wang, Robi Lin, Kat Hu, Joy Jiao, Neil Chowdhury, Ethan Chang, Tejal Patwardhan

arXiv:2601.21165v1cs.AIcs.CYcs.LG

TL;DR

Existing science benchmarks have become increasingly saturated as models improve, motivating harder evaluations of expert-level scientific reasoning. FrontierScience addresses this gap with expert-authored Olympiad and PhD-level Research tasks, including rubric-based assessment of open-ended reasoning. Initial results show strong performance on Olympiad problems but substantial headroom on research-style work.

  • Problem

    As models rapidly improve at reasoning, existing science benchmarks need a new generation capable of keeping pace with progress.

  • Method

    FrontierScience combines expert-authored Olympiad and PhD-level Research problems with constrained answer grading and 10-point rubrics that assess intermediate reasoning.

  • Results

    Frontier AI systems have progressed substantially on expert-level questions, particularly self-contained Olympiad problems, but remain far from saturation on research-style tasks.

  • Takeaways & Limitations

    FrontierScience provides a broader diagnostic of expert-level scientific reasoning strengths and weaknesses across constrained and open-ended tasks.

  • Takeaways & Limitations

    FrontierScience uses constrained problem statements, so it evaluates completing research tasks less than proposing novel research directions, hypotheses, or ideas.

Abstract

from arXiv · show

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. FrontierScience addresses this gap through two complementary tracks: (1) Olympiad, consisting of international olympiad problems at the level of IPhO, IChO, and IBO, and (2) Research, consisting of PhD-level, open-ended problems representative of sub-tasks in scientific research. FrontierScience contains several hundred questions (including 160 in the open-sourced gold set) covering subfields across physics, chemistry, and biology, from quantum electrodynamics to synthetic organic chemistry. All Olympiad problems are originally produced by international Olympiad medalists and national team coaches to ensure standards of difficulty, originality, and factuality. All Research problems are research sub-tasks written and verified by PhD scientists (doctoral candidates, postdoctoral researchers, or professors). For Research, we introduce a granular rubric-based evaluation framework to assess model capabilities throughout the process of solving a research task, rather than judging only a standalone final answer.

1 INTRODUCTION

FrontierScience addresses the need for unsaturated science benchmarks by combining expert-authored Olympiad problems with PhD-level Research subproblems. Its initial evaluations show strong performance on Olympiad tasks but substantial remaining difficulty on research-style work.

  • Recent improvements have made a new generation of science benchmarks necessary to keep pace with model reasoning progress.
  • FrontierScience contains difficult, verifiable, original questions written and verified by subject-matter experts across physics, chemistry, and biology.
  • FrontierScience-Olympiad: The Olympiad track uses short-answer science problems designed by international olympiad medalists to assess precise scientific reasoning.
  • FrontierScience-Research: The Research track uses PhD-scientist-designed subproblems that resemble tasks encountered during original research.
  • FrontierScience combines constrained Olympiad grading with open-ended Research evaluation using expert-designed rubrics to diagnose scientific reasoning strengths and weaknesses.
  • GPT-5.2 scored 77% on Olympiad and 25% on Research, while frontier systems remain far from saturation on research-style work.

2 BENCHMARK CONSTRUCTION

FrontierScience was constructed through expert-authored problem development, multi-stage review, novelty and difficulty controls, and track-specific grading procedures. Its design balances verifiability for Olympiad tasks with rubric-based assessment of intermediate reasoning in open-ended Research tasks.

  • Expert problem creation: FrontierScience-Olympiad problems were created with 42 former international medalists or national team coaches, representing 108 olympiad medals.
  • Expert problem creation: FrontierScience-Research problems were created with 45 qualified scientists, including postdoctoral researchers, professors, and doctoral candidates, across broad scientific disciplines.
  • Data collection pipeline: Tasks passed through Creation, Review, Resolution, and Revision, with independent experts checking questions, solutions, rubrics, and guideline alignment.
  • Quality criteria: Problems were required to be novel and authentically grounded, while preliminary model evaluation was used to discard or modify tasks that models solved too easily.
  • Grading: Olympiad answers are designed for numeric, algebraic, or fuzzy-string verification, whereas Research tasks use experimental rubrics that score intermediate reasoning as well as final answers.
  • Dataset composition: The open-sourced gold set was filtered from over 500 Olympiad and over 200 Research questions to 100 and 60 questions, respectively.
  • Grading: Each Research question includes an expert-crafted solution path and is graded by judge models using a 10-point rubric, with seven points considered a suitable solution.
  • Dataset composition: The gold Research set is equally split among physics, chemistry, and biology, while the Olympiad set is more weighted toward physics and chemistry because verifiable answers are easier to develop there.

3 EXPERIMENTS

FrontierScience evaluates frontier models on diverse Olympiad and open-ended Research tasks using model-based judging and repeated trials. GPT-5.2 leads overall, while models perform better on chemistry than other subjects and still show substantial weaknesses in Research-style tasks.

  • Evaluation setup: The evaluation covers FrontierScience-Olympiad and FrontierScience-Research across several frontier models, using high reasoning effort and no browsing.GPT-5.2 was evaluated at xhigh reasoning effort; the other reasoning models used high effort.
  • Main results: 77% on Olympiad and 25% on Research: GPT-5.2 is the top-performing model overall, while GPT-5 ties it on Research at 25%.Gemini 3 Pro is comparable to GPT-5.2 on Olympiad at 76%.
  • Evaluation setup: Research responses receive rubric-based scores from a GPT-5 judge, while Olympiad answers are judged against the actual answer for equivalence.Research scores reflect rubric points, whereas Olympiad judging compares expressions, numbers, or phrases.
  • Main results: Across subjects, models perform best on chemistry, followed by physics and biology for Olympiad and biology and physics for Research.Transcript analysis identifies reasoning or logic errors, niche-concept failures, calculation errors, and factual inaccuracies.
  • Main results: The Olympiad split spans topics from biochemistry to quantum mechanics, illustrating broad scientific coverage.The figure caption presents the split as diverse across scientific topics.
  • Main results: Research evaluations average 30 independent trials and count responses earning at least seven rubric points as correct.Olympiad evaluations average 20 independent trials.

4 DISCUSSION

FrontierScience advances evaluation of scientific reasoning but remains constrained by its Q&A design, rubric-based judging, text-only format, and lack of human baselines. Model performance also varies by subject and increases with reasoning effort for GPT-5.2.

  • Limitations: FrontierScience does not evaluate proposing novel research directions, hypotheses, or ideas because its constrained problems focus on completing defined research tasks.The Research set measures more open-ended reasoning than the Olympiad set, but remains an autogradable Q&A evaluation.
  • Limitations: Rubric-based Research evaluation is less objective than single-expression or numerical equivalence checking and depends on the model judge’s capabilities.The authors attempted to improve reliability through strict guidelines, verification, and consistency with human grading.
  • Limitations: Text-only problems omit image, video, and interaction with reality, including wet-lab work that often forms part of scientific research.The authors identify modalities beyond text as more representative of scientific research.
  • Discussion: 67.5% to 77.1% on Olympiad and 18% to 25% on Research as GPT-5.2 receives more test-time tokens.OpenAI o3 performs marginally worse at high than medium reasoning effort on the Research set.
  • Limitations: Human baselines were not performed, leaving expert consensus baselines for these specialized questions to future work.The authors note that the specialization of the questions may require domain experts to establish a meaningful baseline.
  • Discussion: Scientific reasoning evaluations should continue developing robust, practical benchmarks relevant to scientific progress.The paper frames research and practical evaluations as important for building long-standing, directly relevant evaluations.

5 RELATED WORK

FrontierScience builds on science benchmarks that increasingly use open-ended evaluation, while emphasizing novel expert-authored reasoning tasks and rubric-based assessment of research subtasks.

  • Existing science benchmarks: Earlier benchmarks such as MMLU, GPQA, and ScienceQA primarily use multiple-choice or single-answer formats to measure knowledge and basic reasoning.These benchmarks have contributed to understanding scientific knowledge and basic reasoning, but largely target retrieval or recognition of established concepts.
  • Open-ended evaluation: OlympiadBench introduced high-school-level open-ended Science Olympiad questions, whereas FrontierScience extends this format with international Olympiad medalists and broader scientific coverage.The paper identifies contamination concerns because OlympiadBench collects pre-existing mathematics and physics questions.
  • Research-task benchmarks: FrontierScience complements CritPt’s verifiable-checkpoint physics benchmark by evaluating more open-ended research subtasks across physics, chemistry, and biology.It also differs from LAB-Bench, which emphasizes multiple-choice biology questions relevant to practical workflows.
  • Rubric-based evaluation: FrontierScience extends rubric-based evaluation toward scientific research, addressing the limited insight provided by final-answer correctness alone.Prior work introduced rubric-based assessment for qualities such as naturalness and conciseness, and HealthBench applied the format in a real-world domain.

A FULL SAMPLE RESEARCH PROBLEMS

The appendix presents sample FrontierScience-Research problems from biology and physics, illustrating the benchmark’s scientific problem coverage.

  • Physics: Figure 9 presents a sample FrontierScience-Research physics problem.
  • Biology: Figure 10 presents a sample FrontierScience-Research biology problem.

B EVALUATION PROMPTS

The evaluation prompts use different judging procedures for Research and Olympiad tasks: rubric-point scoring for open-ended answers and equivalence checking for concise reference answers.

  • Evaluation setup: The paper uses a GPT-5 thinking judge at high reasoning effort to evaluate model responses for both benchmark sets.Research judging produces rubric-point totals, while Olympiad judging compares attempted answers with actual answers.
  • FrontierScience-Olympiad Judge Model Prompt: The Olympiad judge compares an attempted answer with a reference answer and evaluates whether they fully match or are otherwise equivalent.Reference answers may be numbers, algebraic expressions, chemical formulas, compound names, or specific phrases.
  • FrontierScience-Olympiad Judge Model Prompt: Olympiad answers are marked correct when equivalent expressions, numbers within 1 decimal-place rounding, names, formulas, or units match the reference answer.Answers that are not equivalent to the reference answer are marked incorrect.
  • FrontierScience-Olympiad Judge Model Prompt: The Olympiad prompt requires the judge to output either VERDICT: CORRECT or VERDICT: INCORRECT as the final line.
  • FrontierScience-Research Judge Model Prompt: The Research judge receives the problem, attempted answer, and a rubric totaling up to 10 points, then returns the absolute points earned.The prompt instructs the judge to evaluate strictly against the rubric and report a numerical verdict.
  • FrontierScience-Research Judge Model Prompt: The Research prompt requires step-by-step consideration of each rubric item before reporting the total score in a final VERDICT line.

C PROBLEM REQUIREMENTS

The benchmark’s problem requirements emphasize unambiguous, verifiable tasks with explicitly defined information and objectively assessable answers. Research problems target complex scientific reasoning, while Olympiad problems use tightly constrained answer formats and international-olympiad difficulty.

  • Research Problem Guidelines: Research problems must explicitly define all necessary background, variables, notation, and assumptions, while providing experts and models with the same information.
  • Research Problem Guidelines: Research rubrics must use independent, objective, affirmative criteria with specific pass/fail conditions and defined variables and acronyms.
  • Research Problem Guidelines: Research questions should require complex reasoning and typically take 3–5 hours to draft, while testing problem solving rather than prose, search, or recency.
  • Olympiad Problem Guidelines: Olympiad tasks must provide all variables, units, and required information, yielding a single numeric or algebraic expression or a fuzzy string-matchable biology answer.
Loading 2601.21165v1…