Source-linked AI summary
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang, Jinxin Tan, Yuelin Hu, Tu Guo, Jiejun Zhang, Yuhan Wang, Wenbo Li, Yinuo Gao, Shuo Chen, Weiqi Jiang, Yayun Hu, Zixing Lei, Xianghe Pang, Zexi Liu, Yuzhi Zhang, Linfeng Zhang, Kun Chen, Wei Wang, Weinan E, Siheng Chen
TL;DR
Existing benchmarks inadequately evaluate the exploratory planning, adaptation, and procedural complexity required for autonomous scientific research. PRL-BENCH addresses this gap by converting 100 Physical Review Letters papers into expert-validated, long-horizon physics research tasks with objective verification, and finds that frontier models remain substantially below the demands of autonomous research, with the best overall score below 50.
Problem
Existing benchmarks largely use predefined objectives and solution pathways, limiting evaluation of autonomous planning, adaptation, and exploration in realistic scientific research.
Method
PRL-BENCH constructs expert-validated theoretical and computational physics tasks from 100 Physical Review Letters papers, emphasizing exploration-oriented formulation, long-horizon workflows, and objective verifiability.
Results
The best overall frontier-model score remains below 50, with conceptual and formulaic errors dominant and exploration and derivations unstable.
Takeaways & Limitations
Long-horizon reasoning, adaptive methodology selection, and coordination of multi-step workflows remain fundamental challenges for current systems.
Takeaways & Limitations
Tasks provide richer background information and well-defined objectives than authentic research, reducing open-ended exploration and omitting explicit falsification of incorrect hypotheses.
Abstract
from arXiv · showhide
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain confined to domain knowledge comprehension and complex reasoning, failing to evaluate the exploratory nature and procedural complexity of real-world research. In this work, we present research-oriented evaluations in theoretical and computational physics, a natural testbed with comprehensive domain knowledge, complex reasoning, and verifiable end-to-end workflows without reliance on experiments. Here we introduce PRL-Bench (Physics Research by LLMs), a benchmark designed to systematically map the capability boundaries of LLMs in executing end-to-end physics research. Constructed from 100 curated papers from the latest issues of Physical Review Letters since August 2025 and validated by domain experts, PRL-Bench covers five major theory- and computation-intensive subfields of modern physics: astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics. Each task in the benchmark is designed to replicate the core properties of authentic scientific research, including exploration-oriented formulation, long-horizon workflows, and objective verifiability, thereby reconstructing the essential reasoning processes and research workflows of real physics research. Evaluation across frontier models shows that performance remains limited, with the best overall score below 50, revealing a pronounced gap between current LLM capabilities and the demands of real scientific research. PRL-Bench serves a reliable testbed for accessing next generation AI scientists advancing AI systems toward autonomous scientific discovery.
1. Introduction
Agentic science requires systems to autonomously plan, adapt, and explore across end-to-end research workflows, capabilities that existing benchmarks only weakly assess. PRL-BENCH addresses this gap with expert-validated, research-oriented physics tasks, while frontier-model performance remains below the demands of autonomous research.
- Motivation: Existing benchmarks test domain knowledge and reasoning in well-defined problems but provide limited insight into autonomous planning, adaptation, and exploration.Their objectives, solution pathways, and reasoning processes are predetermined.
- Motivation: Theoretical and computational physics offers rigorous reasoning, heterogeneous tools, and verifiable end-to-end workflows without experimental dependence.Existing physics benchmarks still rely largely on short, clear-path tasks rather than authentic long-horizon research.
- Benchmark contribution: PRL-BENCH converts 100 Physical Review Letters papers into open-pathway, long-horizon research tasks spanning five physics subfields and validated by more than ten domain experts.Expert cross-validation checks consistency with the source papers’ underlying physics.
- Evaluation findings: Even the strongest frontier models score well below 50 overall, with failures dominated by conceptual and formulaic errors and unstable exploration and derivations.The reported gap reflects limitations in domain knowledge, derivation stability, numerical reliability, and long-horizon adaptation.
2. Related Work
Scientific benchmarks have progressed from closed-ended knowledge questions to harder reasoning tasks and research-oriented evaluations. However, existing resources remain limited in exploratory realism, physics coverage, scale, or faithful end-to-end reproduction.
- General science benchmarks: Early science benchmarks emphasized closed-ended question answering, while later benchmarks added complex reasoning and advanced domain knowledge.Humanity’s Last Exam is described as difficult and comprehensive but still lacking the exploratory nature of scientific research.
- Research-oriented evaluation: Frontier Science introduced research-oriented evaluation, but its physics component contains only 20 questions and does not effectively cover condensed matter or high-energy physics.Its limited scale restricts coverage of frontier physics subfields.
- Physics benchmarks: Physics-specific benchmarks generally use short, clear-path tasks, while PRBench focuses on reproducing detailed implementations and results from original physics studies.These designs address different aspects of physics research than exploratory, long-horizon task execution.
- AI scientists: General-purpose AI scientist systems have emerged as AI’s role expands from isolated scientific subtasks toward potentially autonomous research.Examples include AI co-scientist, Robin, and Kosmos.
3. Benchmark
PRL-BENCH builds a five-subfield physics benchmark from 100 theory- and computation-focused papers, using exploratory, long-horizon tasks with objective checks. Its design combines heterogeneous subtasks, explicit rubrics, and verifiable answers to assess methodological adaptation.
- Corpus: The benchmark source corpus contains 100 Physical Review Letters papers selected for theoretical derivation and numerical computation.Experimental studies and tasks requiring large-scale datasets, substantial resources, or specialized software are excluded.
- Subfields: PRL-BENCH spans astrophysics, condensed matter physics, high-energy physics, quantum information, and statistical physics.These areas cover physical phenomena from cosmological structures to microscopic quantum regimes and use distinct methodological approaches.
- Subfields: The benchmark evaluates whether models can adapt appropriate methodological strategies across major physical subfields in end-to-end research.The target is broader than robust reasoning alone.
- Task design: Tasks preserve exploration-oriented formulation, long-horizon workflows, and objective verifiability rather than closed-form, single-path problem solving.They provide scientific motivation and concrete objectives while leaving solution pathways implicit.
- Task design: Each task combines relatively independent analytical and computational subtasks under a shared scientific objective, reducing error propagation.This structure supports assessment of capability boundaries across heterogeneous research activities.
- Evaluation: Answers provide verifiable numerical values, formulas, or discrete judgments, while rubrics decompose subtasks into reasoning steps and checkpoints.Together they support reproducible, fine-grained evaluation of long-horizon exploration.
4. Evaluation
PRL-BENCH evaluates frontier LLMs under unified tool-use and scoring procedures across physics subfields, revealing low overall performance and distinct error patterns. Conceptual and formulaic mistakes dominate, while derivation, calculation, and incomplete-response failures expose additional limitations in research-oriented reasoning.
- Evaluation setup: Each problem is executed five times per model, averaged, and scored from 0–100 by GPT-5 based on final-answer correctness and rubric-matching intermediate results.Models receive a code interpreter, while search tools are disabled to prevent information leakage and support impartial evaluation.
- Overall performance: 44.27 is the best overall score, leaving even frontier models well below 50 on PRL-BENCH.The result indicates that multi-step derivation, numerical validation, and autonomous planning remain major bottlenecks.
- Model comparison: Gemini-3.1-Pro achieves the highest overall score and leads in multiple subfields, while Qwen-3.5-Plus ranks second and Kimi-K2.5 trails behind.GPT-5.4, Claude-Opus-4.6, and Doubao-Seed-2.0-Pro form a broadly comparable middle tier.
- Subfield comparison: Most models perform worse in Astrophysics and Statistical Physics than in Condensed Matter, High-Energy Physics, and Quantum Information.The authors infer that greater heterogeneity and weaker standardization reduce canonical training coverage and reusable reasoning templates in Astro and Stat.
- Error analysis: Formulaic or conceptual errors account for roughly 45–55% of global errors, including 0.4697 for GPT-5.4, 0.5079 for Gemini-3.1-Pro, and 0.5562 for Doubao-Seed-2.0-Pro.The pattern reflects inappropriate physical models or formulas and is especially pronounced in Condensed Matter.
- Error analysis: Derivation errors are typically ≈0.08–0.13 globally, calculation errors ≈0.20–0.30, and Claude-Opus-4.6 has 0.6393 incomplete or unsupported responses globally.Derivation errors become more prominent in HEP, while calculation errors remain non-trivial but are not dominant; Claude’s incompleteness is linked to unstable long-horizon trajectories.
5. Conclusion
PRL-BENCH is a research-oriented benchmark built to evaluate LLMs on realistic physics research workflows rather than closed-form problems. Its evaluation finds a substantial capability gap, arising from combined weaknesses in knowledge, reasoning stability, numerical reliability, and long-horizon adaptation.
- Benchmark contribution: PRL-BENCH emphasizes exploration-oriented tasks, long-horizon reasoning, and heterogeneous tool integration to reflect authentic scientific inquiry.It is designed to evaluate capability boundaries in realistic physics research settings.
- Benchmark construction: 100 curated Physical Review Letters papers support a benchmark spanning five physics subfields and combining analytical with computational task components.The benchmark was validated by domain experts.
- Findings: The performance gap reflects combined deficiencies in domain knowledge, derivation stability, numerical reliability, and long-horizon task adaptation rather than a single failure mode.Conceptual and formulaic errors, unstable trajectories, and incomplete solutions indicate limited robustness in research planning and exploration.
- Implications: Long-horizon reasoning, adaptive methodology selection, and coordination of multi-step workflows remain fundamental challenges for current systems.PRL-BENCH provides a rigorous and scalable testbed for future AI-scientist research.
6. Limitations and Future Work
PRL-BENCH acknowledges limitations in task openness, annotation reliability, and disciplinary categorization. Future work targets more open formulation, hypothesis testing, broader coverage, and diverse research paradigms.
- Limitations: Richer background information makes objectives and answers verifiable but partially reduces the difficulty of open-ended exploration.The benchmark also omits explicit falsification of incorrect hypotheses, despite its importance in real scientific reasoning.
- Limitations: Expert validation does not eliminate possible annotation imperfections.The authors plan iterative refinement through expert review and community feedback.
- Limitations: The five-subfield division is approximate because problems such as quantum many-body systems span multiple domains.Strict categorization may therefore not fully capture interdisciplinary research.
- Future Work: Future work will increase task openness, add hypothesis generation and falsification, and extend the benchmark across broader domains and research paradigms.
Appendix A: Full Sample Task in PRL-BENCH
The sample task studies tensor-network simulation of lattice gauge theories in (2+1) dimensions through multiple theory-specific subtasks and thermal observables. Its reference answers combine numerical results with symmetry, conservation, and dynamical interpretations.
- Introduction: Tensor-network methods are presented as a classical strategy for real-time lattice-gauge simulation that avoids the sign problem.The sample focuses on the more challenging (2+1)d setting with PEPS, higher-dimensional entanglement, and exact gauge constraints.
- Model Setup: A gauge-invariant PEPS ansatz uses vertex and link tensors while enforcing gauge invariance through Gauss-law constraints.
- Subtasks: The task spans pure Z3, odd Z2, hard-core-boson, and finite-temperature gauge-theory subtasks on specified lattices and parameter settings.It requests derivatives, staggered observables, critical fields, plaquettes, total energy, and dynamical interpretation.
- Reference Answers: 0.2421075221119777 is obtained for both squared staggered observables, with C4 symmetry enforcing their equality.
- Reference Answers: The rubric identifies total-energy conservation during real-time evolution and infers vison propagation from spatiotemporal plaquette dynamics.
Appendix B: Evaluation Prompt
The evaluation prompt scores candidate answers against reference answers and rubrics at the subtask level. It requires normalized scoring, one error category per subtask, concise reasons, and strict output formatting.
- Inputs: The judge receives the problem, reference answer and rubrics, and candidate answer as its evaluation inputs.
- Scoring: Scores are normalized to 0–100 by summing predefined points for correctly satisfied final-answer and intermediate rubric items.Credit is limited to responses that sufficiently match the rubric requirements.
- Error Classification: Each subtask receives exactly one primary error type from the defined categories: formulaic or conceptual, derivation, calculation, incomplete, or correct.
- Output Requirements: Evaluation occurs per subtask and records correctness, one error type, and a brief one-to-three-sentence reason.Correct subtasks must use the Correct error type, and judgments must remain grounded in the rubric.
- Output Format: The output must match the required subtask count and schema, with no additional text or fields.