Source-linked AI summary
FrontierChallenge: Evaluating Scientific Workflow Completion
Liangcai Su, Zhaopeng Feng, Zhuo Chen, Zhen Zhang, Xiang Lin, Ruilin Li, Handuo Zhang, Ning Wang, Kailong Wen, Yueqi Guo, Feng Xing, Yiling Guo, Chenxiong Qian, Simon Shaolei Du, Lidong Bing, Xinyu Wang
TL;DR
Existing benchmarks often do not test whether agents can complete heterogeneous scientific workflows with multiple required deliverables. FrontierChallenge evaluates this capability across 300 workflows, releasing 97 tasks and measuring both contract-level completion and partial progress. The best configurations completed only 20 of 97 tasks, while high partial scores remained poorly aligned with complete delivery.
Problem
Existing benchmarks often evaluate final answers, interaction traces, single programs, or single-discipline workflows rather than heterogeneous scientific work with multiple required deliverables.
Method
The benchmark defines fixed-input, multi-stage workflows with heterogeneous deliverable contracts and task-specific executable evaluation, then evaluates twelve frontier models with multiple agent scaffolds.
Results
20 of 97 tasks were completed by the best-performing configurations, yielding a 20.6% Pass Rate despite a highest Avg. Score of 87.9.
Takeaways & Limitations
Complete, contract-level delivery is a distinct unresolved capability requiring explicit contract tracking, cross-artifact validation, and evidence-based completion checks.
Takeaways & Limitations
The six domains are descriptive slices of the released evaluation set, not probability samples of their scientific fields, so domain performance differences should not be interpreted as intrinsic disciplinary difficulty rankings.
Abstract
from arXiv · showhide
Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows. In this paper, we release and evaluate 97 of these tasks, spanning quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment. Each task provides fixed inputs and specifies a bundle of required scientific deliverables. We evaluate twelve frontier models with three agent scaffolds. Pass Rate measures the fraction of tasks satisfying the full-completion criterion, while Avg. Score captures partial progress. Each of the best-performing configurations completed only 20 of the 97 released tasks, yielding a Pass Rate of 20.6%. Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment: Avg. Scores reached 87.6 and 94.9, but the highest Pass Rates were only 4% and 0%. Among non-passing Claude Code trajectories, 75.5% still ended with language claiming completion. These findings show that neither high partial scores nor confident claims of completion reliably indicate that a scientific task has been fully delivered, highlighting the need to evaluate end-to-end workflow execution and the completeness of scientific deliverables together.
1 Introduction
FrontierChallenge addresses whether agents can complete heterogeneous, multi-stage scientific workflows and deliver all required artifacts, rather than merely produce plausible answers. It introduces a cross-domain benchmark and evaluates complete delivery separately from partial progress.
- Existing benchmarks often evaluate final answers, interaction traces, single programs, or workflows from one discipline, leaving heterogeneous multi-deliverable completion insufficiently characterized.
- FrontierChallenge asks whether an agent can independently execute a specified workflow from fixed data through final deliverables while satisfying the complete task contract.
- Pass Rate measures complete contract satisfaction, whereas Avg. Score measures partial progress and does not count incomplete high-scoring submissions as passes.
- 20 of 97 tasks were completed by the best configurations, producing a 20.6% Pass Rate despite a highest Avg. Score of 87.9.
- 300 workflows span six scientific domains, with 97 publicly released and 203 retained as an internal held-out set.
- Each task uses fixed inputs, heterogeneous deliverables, and a task-specific Grader, evaluating complete submitted artifact bundles rather than single answers.
2 Related Work
Existing benchmarks span expert knowledge, tool use, scientific data analysis, and end-to-end research workflows. However, many still evaluate a single answer, decision set, or self-contained program rather than heterogeneous deliverables.
- General benchmarks evaluate difficult expert questions, tool use, computer interaction, software engineering, and long-horizon professional work.
- Scientific benchmarks target biological reasoning, biological data problems, expert analytical decisions, or executable analyses derived from literature.
- Table 1 summarizes core design characteristics of related benchmarks using solid dots for primary characteristics and dashes otherwise.
- Many scientific benchmarks assess an answer, decision set, or single self-contained program instead of a heterogeneous deliverable set.
- End-to-end benchmarks evaluate computational reproduction, research replication, scientific software use, interactive research environments, multi-step tool use, or bioinformatics and biomedical artifacts.
3 FrontierChallenge
FrontierChallenge collects realistic, diverse, verifiable scientific workflows and packages them with fixed inputs, execution environments, deliverable contracts, and executable evaluation. The released set contains 97 tasks from a curated 300-workflow collection, while domain composition limits cross-domain interpretations.
- Task collection: Tasks come from analysis, computation, simulation, and research-delivery processes requiring domain knowledge, specialized software, experimental data, or engineering environments.
- Screening and quality control: The collection requires representative professional practice, dependency-aware end-to-end complexity, workflow diversity, and verifiable outputs.
- Released task composition: Figure 1 depicts domain and workflow-family composition for the 97 released tasks and shows the retained Hard/Medium split in its center bars.
- Screening and quality control: Every included task has fixed inputs, a defined execution environment, a complete deliverable contract, and an executable evaluation procedure.
- Standardization and packaging: Each task package aligns the scientific objective, fixed inputs, available software and tools, required deliverables, and successful-completion evaluation.
- Release and evaluation set: All 300 workflows passed quality-control checks; 97 non-GPU-required tasks were randomly released for experiments, while 203 remain held out.
- Released task composition: The released set contains 97 tasks across six domains and 21 workflow families, with reports, structured data, figures, executable code, and simulation products among its deliverables.
- Scope limitation: Domain slices are not probability samples, so performance differences should not be interpreted as intrinsic rankings of disciplinary difficulty.
4 Experiments
The experiments are organized to address three research questions. The passage establishes the study’s question-driven experimental structure without specifying the questions themselves.
- The authors conducted experiments for the paper’s evaluation.
- The experimental program addresses research questions rather than an unspecified objective.
- The study identifies three research questions guiding its experiments.
2. RQ2: How does scientific workflow performance vary across domains?
Across domains, scientific workflow performance showed a persistent separation between partial progress and strict complete delivery. Quantum chemistry and molecular dynamics had the strongest pass rates, while analytical chemistry and electrochemistry/environment combined high average scores with very low or zero completion.
- Overall performance: 20.6% was the highest overall Pass Rate, while the highest Avg. Score reached 87.9 across evaluated configurations.Codex with GPT-5.6 Sol achieved the highest Avg. Score, and Claude Code with Grok 4.6 shared the highest Pass Rate.
- Overall performance: Eight configurations scored above 80 Avg. Score, but none completed more than 20.6% of tasks under the strict criterion.High-scoring submissions often still missed at least one required artifact or other contract requirement.
- Domain-level performance: 60% was the highest domain Pass Rate in quantum chemistry, while molecular dynamics reached 38%.Claude Code with Grok 4.6 led quantum chemistry; three configurations shared the molecular-dynamics maximum, while Terra had the highest molecular-dynamics Avg. Score at 93.7.
- Domain-level performance: 94.9 was the maximum electrochemistry/environment Avg. Score, yet every configuration had a 0% Pass Rate; analytical chemistry reached 87.6 and only 4% completion.Materials characterization also showed divergence, with Avg. Scores up to 88.1 but no Pass Rate above 9%.
- Domain-level performance: Representative configurations had distinct domain profiles, so aggregate rankings did not capture domain-specific strengths and weaknesses.GPT-5.6 Sol led Avg. Score in four domains, while Grok 4.6 led quantum-chemistry Pass Rate and shared the molecular-dynamics lead.
- Failure Mode Analysis: 75.5% of non-passing Claude Code trajectories used completion language, whereas raw tool errors occurred in both 80.7% of non-passing and 94.2% of passing trajectories.Failure-signature frequencies also varied by domain, partly reflecting differences in task contracts and evaluator composition; the analyses were descriptive and did not assign unique causes.
5 Conclusion
FrontierChallenge evaluates whether agents can complete specified, multi-stage scientific workflows and deliver mutually consistent artifacts rather than merely plausible answers. The findings show that reliable completion remains distinct and unresolved, motivating explicit contract tracking, cross-artifact validation, and evidence-based checks.
- FrontierChallenge evaluates multi-stage workflow completion and mutually consistent artifact delivery rather than merely plausible answers.
- 20.6% was the highest Pass Rate among evaluated configurations, despite a highest Avg. Score of 87.9.
- Partial progress translated especially poorly into complete delivery in analytical chemistry and electrochemistry/environment.
- 75.5% of non-passing Claude Code trajectories ended with completion language, while tool errors were frequent in both passing and non-passing runs.
- More reliable scientific agents will require explicit contract tracking, cross-artifact validation, and evidence-based completion checks.
A Contributors
The paper lists its contributors and identifies equal contributions and project leadership.
- Liangcai Su, Zhaopeng Feng, Zhuo Chen, and Zhen Zhang are among the listed contributors.
- The paper marks some contributors as having contributed equally and identifies a project lead.
B Illustrative Task Case
The illustrative cases expose the benchmark unit as a scientific objective, frozen inputs, a required workflow, and an artifact contract. They describe what must be completed and handed off without revealing solutions, scores, evaluator internals, or reference artifacts.
- Each illustrative case specifies a scientific objective, frozen input bundle, required workflow, and artifact contract.
- The examples characterize required work and handoff contents rather than task answers.
B.1 Cell-migration wound-healing assay
The wound-healing case is an image-analysis and statistical workflow using 18 bright-field microscopy images across two groups, three time points, and three biological replicates. Experimenter-drawn white contours identify the unmigrated region, while the required handoff includes segmentation, migration analysis, statistical tests, reproducible code, tables, plots, and a report.
- Task and visible evidence: The case uses 18 bright-field microscopy images spanning control and experimental groups, three time points, and three biological replicates.
- Task and visible evidence: Figure 6 displays one complete time course from each group, with experimenter-drawn white outlines marking the unmigrated region in the source images.
- Expected handoff: The system must segment the marked region in all images, inspect quality, pair time points by group and replicate, and compute 12-hour and 24-hour migration rates relative to matching 0-hour images.
- Expected handoff: Separate between-group tests are required at both follow-up times.
- Expected handoff: The checkable handoff includes a reproducible script, image-level and group-statistics tables, two labeled summary plots, and a methodological report.
B.2 TLC monitoring of a Suzuki coupling
The Suzuki-coupling TLC task starts from a structured, eight-lane UV254 plate with tilted geometry and requires a complete, reproducible analysis-to-endpoint handoff. The visible plate is unmodified, so all annotations, integrations, compositions, progress estimates, and endpoint decisions must be produced by the system.
- Task and visible evidence: The frozen input is an eight-lane UV254 TLC plate containing standards, a co-spot, and reaction samples from 0 to 120 minutes.The supplied data also include plate geometry, compound-reference Rf windows, response factors, integration settings, and the endpoint rule.
- Task and visible evidence: Tilted spotting and solvent-front lines require lane-specific Rf computation from local geometry.Global geometric assumptions would not satisfy the stated measurement setup.
- Required handoff: The system must detect and integrate spots after local background correction, assign compounds, apply response correction, reconstruct reaction progress, and recommend an endpoint.Assignments use the standards and co-spot, while the endpoint follows the supplied rule.
- Required handoff: The handoff requires a reproducible script, traceable QC and analysis tables, an annotated plate, lane profiles, and a report.Required tables cover plate QC, spot-level results, lane composition, reaction progress, and endpoint determination.
- Required handoff: Figure 7 shows only the agent-visible input plate, without detected boundaries, Rf assignments, corrected intensities, compositions, or an endpoint decision.The figure therefore documents the starting evidence rather than an analytical result.