Source-linked AI summary
Expert-validated STEM QA
Kihwan Han, Saurabh Patil, Chinmayee Shukla, Abhinav Sharma, Marko Pavlovic, Anshuman Lall, Mahesh Joshi
TL;DR
Existing STEM benchmarks have limited headroom and shortcomings in format, taxonomy, and answer validation, motivating more reliable expert-created evaluations. The paper introduces Expert-validated STEM QA, a consensus-reviewed dataset across four STEM domains, and reports low benchmark performance alongside gains after post-training on a private version.
Problem
Existing STEM benchmarks can be saturated, taxonomically skewed, misaligned with scientific use, and affected by inaccurate or ambiguous answers, limiting evaluation reliability.
Method
The study constructs a 398-instance Physics, Chemistry, Biology, and Mathematics dataset using balanced taxonomy design, quality-driven contributor selection, and consensus-based multi-layer expert review.
Results
Frontier and open-source models scored below 25% overall, while post-training on a private 2,000-instance version increased open-source model performance by 15% of baseline on HLE-verified STEM (p=0.045).
Takeaways & Limitations
The dataset provides a human-created, expert-validated benchmark and shows potential utility for training models for scientific reasoning.
Takeaways & Limitations
The study could not systematically evaluate model reasoning processes, used only 398 public instances, and fine-tuned a non-reasoning model because reasoning traces were unavailable.
Abstract
from arXiv · showhide
Recent advancements in AI are helping scientists achieve breakthroughs in fields such as mathematics, medicine, and materials sciences. New evaluation datasets for AI models contribute to such advancement in AI. In the STEM domain, frontier models have consumed most of the available online data, creating the need for human-created datasets that codify the knowledge of leading experts in the domain. There are several STEM datasets available for the research community in this field. However, there are some gaps in these datasets, leaving room for improvement. Examples of gaps include (1) saturation in model performance on these datasets, leaving no head-room for meaningful evaluations, (2) skewed taxonomy distributions, (3) multiple choice question format that is misaligned with how scientists use AI in the real world, and (4) inaccurate answers and rationales partially led by a contest-based data collection and a time-bound review process. In this study, we present 'Expert-validated STEM QA', a high-quality, expert-validated STEM dataset (N=398) in Physics, Chemistry, Biology, and Mathematics, created by 241 domain experts. We (1) carefully designed a taxonomy with balanced distribution, (2) vetted question contributors with quality-driven incentive, (3) conducted multiple rounds of reviews with revisions validated by domain experts based on consensus, and (4) created the dataset in verifiable question and answer format. Our study demonstrated low performance ($<25\%$) of frontier AI models on the dataset as a benchmark. Post-training on a separate, private version of the dataset (N=2,000) increased performance of the open source model by $15\%$ relative to the baseline model (p=0.045) on the STEM subset of HLE-verified dataset, indicating potential utility of the dataset for model training. We have open-sourced a portion of our dataset for the AI research community.
1 INTRODUCTION
Existing STEM benchmarks have saturation, format, taxonomy, leakage, and accuracy limitations. Expert-validated STEM QA addresses these gaps with an expert-reviewed dataset spanning four STEM domains.
- Existing benchmarks can approach saturation, use formats misaligned with scientific practice, and exhibit skewed taxonomies or restricted domain coverage.
- Inaccurate answers and ambiguous questions can make benchmark scores misrepresent model capability and inflate apparent confidence in scientific responses.
- Expert-validated STEM QA contains 398 questions across Physics, Chemistry, Biology, and Mathematics.
- The dataset uses a carefully designed taxonomy, quality-driven incentives, and consensus-based multi-layer expert review.
- The study describes dataset generation, evaluates model benchmarking utility, and examines post-training for improving scientific reasoning.
2 MATERIALS AND METHODS
The study designs, annotates, reviews, and evaluates a four-domain STEM question-answer dataset through structured guidelines and multi-stage quality control. It also tests model performance, consistency, and post-training effects using external evaluation data.
- Dataset design: Each instance is specified as a question, explanation, and answer in a verifiable question-and-answer format across Physics, Chemistry, Biology, and Mathematics.
- Annotation: Annotator guidelines define project goals, instance structure, taxonomy, metadata, and difficulty calibrated against mid-tier model responses.
- Dataset design: The two-level taxonomy combines prior benchmarks, journal research areas, key papers, and subject-matter-expert discussions to support topical diversity.
- Annotation: Metadata records web-searchability evidence and model-response evaluations distinguishing reasoning limitations from question ambiguity.
- Review: Review rubrics decompose quality criteria into atomic binary judgments to support consistent expert assessment of answer accuracy and other requirements.
- Review: Annotators revise submissions through multiple rounds involving agentic reviewers, reviewers, and team leads.
- Benchmark evaluation: Ten model variants from three proprietary and one open-source model family were evaluated with varying reasoning effort levels.
3 RESULTS
The dataset spans four STEM domains with expert-informed coverage, while model evaluations reveal low scores, overconfidence, and potential gains from post-training.
- 3.1 Annotator, reviewer, and team lead statistics: 80% of the 241 annotators, reviewers, and team leads earned PhDs in their expertise.
- 3.2 Dataset generation results: 398 examples were balanced across Physics, Chemistry, Mathematics, and Biology, although Mathematics and Biology had concentrated L1 taxonomy coverage.The concentration was attributed to the expertise represented in the annotator pool and, for Biology, to molecular biology’s broad conceptual scope.
- 3.2 Dataset generation results: Chemistry primarily drove the upper long tails in question and response token counts, while Biology had relatively low combined-response token counts.
- 3.3 Model performance on the dataset: All 10 models scored below 25%, with higher reasoning effort associated with higher scores and open-source models relatively weaker in Chemistry.Proprietary models outperformed open-source models overall, while larger open-source models scored higher.
- 3.3 Model performance on the dataset: Median confidence exceeded 80% for nearly all model configurations despite low scores, indicating frequent high-confidence incorrect responses.GPT 5.2 at medium reasoning effort was the stated exception.
- 3.3 Model performance on the dataset: Neither open-source model’s pass@k curve showed a steep elbow, implying that many attempts would be needed to approach its performance ceiling.The 27B model’s pass@k curve remained consistently higher than the 9B model’s across eight trials.
- 3.4 Post-training results: 15% of baseline performance improvement followed fine-tuning Qwen 3.5 9B on a separate proprietary dataset, with statistical significance at p=0.045.The gain was driven primarily by the verifiable question-and-answer subset, which improved by 22% of baseline performance.
4 DISCUSSIONS
The dataset combines expert validation, leakage mitigation, and holistic model assessment to benchmark STEM reasoning. Results show low model performance, potential training utility, and important limits on interpretation.
- <25%: Frontier and open source models scored below one-quarter of total scores on the dataset.
- 241 domain experts, 80% with PhDs, contributed to the dataset’s consensus-driven review process.
- Verified search-proof evidence and a private dataset portion were used to mitigate leakage and support objective generalizability measurement.
- Expert verification and substantial headroom favor questions requiring genuine multi-step reasoning over questions designed to exploit model weaknesses.
- Model evaluation should extend beyond answer accuracy to include hallucination and reasoning efficiency.
- The study could not systematically evaluate reasoning processes across models, and its post-training interpretation is limited by data volume and training constraints.
A.1.1 Physics
The Physics taxonomy spans multiple research areas, including astrophysics, condensed matter physics, and electromagnetism and photonics.
- Physics coverage includes Astrophysics, Condensed Matter Physics, and Electromagnetism & Photonics categories.
A.1.2 Chemistry
The Chemistry taxonomy covers analytical, biochemical, general, electrochemical, inorganic, and polymer chemistry topics.
- Chemistry coverage includes Analytical Chemistry, Biochemistry, General Chemistry, Electrochemistry, Inorganic Chemistry, and Polymer Chemistry.
A.1.3 Mathematics
The Mathematics taxonomy includes Algebra and Analysis, with topics spanning algebraic structures, functions, approximation, and differential equations.
- Mathematics coverage includes Algebra and Analysis categories with topics such as representation theory, analytic functions, polynomial approximation, and differential equations.
A.1.4 Biology
The Biology taxonomy spans biochemistry, cell biology, genetics, immunology, microbiology, neurobiology, and proteomics, with multiple specialized subtopics listed across these areas.
- Biology: Biology includes Protein Biochemistry under Biochemistry and several specialized areas within Cell Biology.Cell Biology subtopics include cancer cell growth, cellular signaling, cellular transport, and lineage fate specification.
- Biology: Genetics covers ten subtopics, including epigenetics, gene expression, gene function, and molecular genetics.The listed areas also include linkage and gene mapping, mutations, non-Mendelian genetics, population genetics, protein function, proteomics, and stress response.
- Biology: Immunology includes immunotherapy and in vitro kinetics of cytokine secretion, while Microbiology includes host–pathogen interactions in plants, structural biology, and virology.
- Biology: Neurobiology covers brain tumor biology, molecular mechanisms of synaptic transmission, and NMR-based brain diagnostics.
- Biology: Proteomics includes protein biochemistry in plant–pathogen interactions and signal transduction profiling.
A.2 Prompts for model response and grading, and label for post-training
The appendix documents prompts, review materials, credentials, token-count distributions, example responses, and additional post-training contingency tables.
- A.2 Prompts for model response and grading, and label for post-training: Figure S1 presents the prompt used to generate model responses during scoring, with the variable to substitute shown in cyan.
- A.2 Prompts for model response and grading, and label for post-training: Figure S2 presents the prompt used to grade model responses, with variable names to substitute shown in cyan.
- A.2 Prompts for model response and grading, and label for post-training: Figure S3 presents the prompt used for model response generation during post-training, with the variable to substitute shown in cyan.
- A.2 Prompts for model response and grading, and label for post-training: Figure S4 labels post-training and identifies the variable name to substitute in cyan.
- A.3 Review rubrics and A.4 Annotator, reviewer, and team lead credentials: Table S1 contains the review rubrics, while Table S2 breaks down credentials for annotators, reviewers, and team leads.
- A.5 Token count distribution: Figure S5 shows completion-token-count distributions across models after excluding examples without model responses.
- A.6 Example model responses: Figure S6 shows example reasoning snapshots from two open-source models and highlights repeated phrases and expressions of confusion in sky blue.
- A.7 Additional post-training results: Tables S3–S5 provide baseline-versus-post-training correct/incorrect contingency tables for HLE-verified STEM overall, MCQ, and VQA.