Source-linked AI summary
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
Sean Wu, Pan Lu, Yupeng Chen, Jonathan Bragg, Yutaro Yamada, Peter Clark, David Clifton, Philip Torr, James Zou, Junchi Yu
TL;DR
It remains unclear whether strong scientific reasoning enables reliable forecasts of future scientific advances. Using CUSP to evaluate frontier AI models, the paper finds plausible mechanistic reasoning but limited reliability in forecasting feasibility, realization pathways, and timing.
Problem
Existing scientific reasoning evaluations provide limited evidence about whether AI models can reliably anticipate scientific advances beyond their knowledge cutoff.
Method
The paper introduces CUSP, a temporally grounded suite of 4,760 verifiable scientific events across eight disciplines and four event-level forecasting dimensions.
Results
Across frontier models, mechanistic forecasting is strongest, while feasibility, solution-design alignment, temporal prediction, and realized pathways are substantially less reliable.
Takeaways & Limitations
Scientific reasoning and scientific forecasting should be evaluated as distinct capabilities when using AI for research prioritization and scientific decision-making.
Takeaways & Limitations
CUSP evaluates event-level forecasting, so whether these limitations extend to entire fields, long-term trajectories, or scientific paradigms remains open.
Abstract
from arXiv · showhide
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances. Here we show that strong scientific reasoning does not reliably translate into accurate forecasting of future scientific advances. To study this question, we introduce CUSP, a temporally grounded evaluation suite for event-level scientific forecasting across eight scientific disciplines. Across six frontier AI models, we observe a striking asymmetry in forecasting performance together with systematic error patterns. Models often identify plausible mechanisms underlying future scientific advances, yet perform near chance on feasibility assessment, generate solution strategies that only weakly align with realized advances, and systematically predict scientific advances later than they become publicly observable. Providing additional pre-cutoff scientific knowledge improves performance but does not eliminate these forecasting limitations. These findings suggest that current AI systems possess substantial retrospective scientific competence but limited forward-looking predictive capability. Scientific forecasting should therefore be evaluated as a complementary dimension of AI scientific capability when deploying AI systems for research prioritization and scientific decision-making.
Introduction
This study introduces CUSP to evaluate whether frontier AI models can forecast scientific advances beyond their knowledge cutoffs. Across six models, strong mechanistic reasoning coexists with near-chance feasibility assessment, weak alignment between proposed and realized solutions, and systematic temporal misprediction.
- Evaluation framework: CUSP comprises 4,760 verifiable scientific events across 8 disciplines for temporally grounded, event-level scientific forecasting.The suite evaluates four complementary forecasting dimensions, including feasibility assessment and mechanistic forecasting.
- Forecasting asymmetry: Across six frontier AI models, performance is strongest on mechanistic forecasting but weaker on solution design, feasibility assessment, and temporal prediction.Models often produce technically detailed and scientifically plausible reasoning while remaining less reliable about whether, when, and how advances will occur.
- Forecasting asymmetry: Feasibility assessment remains near chance, while generated solution strategies only weakly align with the approaches underlying realized scientific advances.These limitations are reported consistently across the evaluated frontier AI models.
- Knowledge and forecasting: Additional pre-cutoff scientific knowledge consistently improves forecasting performance but does not eliminate substantial limitations.The findings indicate that forecasting requires more than access to and reasoning over existing scientific knowledge.
- Systematic errors: Models exhibit persistent response biases, substantial overconfidence, and a consistent tendency to predict advances later than they become publicly observable.These systematic failures suggest limitations in how current AI systems form calibrated expectations about future scientific advances.
Results
Across frontier AI models, scientific forecasting is strongest for identifying plausible mechanisms but substantially weaker for predicting whether advances will occur, how they will be realized, and when they will become observable. Additional historical knowledge improves performance, yet large forecasting gaps and structured response biases persist.
- Overall forecasting asymmetry: GPT-5.4 reached 79.2% on mechanistic forecasting versus the 25% random-guess baseline, but only 5.04/10 on open-ended solution design.Mechanistic forecasting was the strongest dimension across evaluated models, while solution design was considerably weaker.
- Overall forecasting asymmetry: 49.1% GPT-5.4 accuracy on Binary prediction remained near chance and below the 57.03% always-No baseline; frontier models ranged from 43.5% to 52.6%.Temporal prediction showed similar limitations, with LLaMA 3.3 achieving the highest Date score at 0.500.
- Solution design: Alignment was the lowest-scoring FRQ dimension, reaching only 3.3/10 for GPT-5.4 despite feasibility scores ranging from 4.8 to 6.0 across models.The gap indicates that failures more strongly concern matching realized advances than constructing scientifically feasible strategies.
- Solution design: From GPT-4o to GPT-5.4, specificity rose from 2.9 to 6.2 and novelty from 2.4 to 4.9, while alignment increased only from 2.2 to 3.3 and feasibility remained 4.8–6.0.Recent model improvements therefore favored richer and more creative solutions over alignment with realized scientific advances.
- Knowledge access: For post-cutoff events, GPT-5.4’s temporal forecasting gap was more than six times its knowledge gap: ∆fore = 0.44 versus ∆know = 0.07.Additional pre-cutoff information consistently improved performance, but substantial forecasting limitations persisted when relevant historical knowledge was available.
- Human comparison and error patterns: Humans outperformed GPT-4o under identical information conditions, achieving 73.0% versus 52.5% on binary questions, 67.0% versus 40.0% on MCQs, and 0.65 versus 0.54 on Date score.Models also showed structured feasibility-response priors: LLaMA 3.3 and GPT-5.4 favored affirmative responses, while most other models favored negative predictions.
Discussion
CUSP shows that frontier AI systems’ scientific reasoning does not reliably translate into event-level forecasting of future scientific advances. Their systematic forecasting errors have practical implications, while the findings remain bounded to the event-level setting studied.
- Core findings: CUSP evaluates 4,760 scientific events across eight disciplines and finds substantial limitations in current frontier AI systems’ event-level scientific forecasting.The results indicate that these limitations cannot be explained solely by access to scientific knowledge and include consistent systematic error patterns.
- Core findings: Scientific reasoning and scientific forecasting should be treated as distinct evaluations because forecasting requires expectations about scientific outcomes that have not yet occurred.Existing benchmarks mainly assess reasoning about scientific knowledge through question answering, retrieval, and problem solving.
- Practical implications: Systematic forecasting failures may cause overconfidence, bias go/no-go decisions through persistent response priors, and delay investment through temporal conservatism.These patterns could lead systems to overestimate uncertain opportunities, mis-rank research directions, and predict rapidly developing areas later than they emerge.
- Limitations and future directions: CUSP’s conclusions are limited to event-level scientific forecasting, leaving their applicability to entire fields, long-term trajectories, and paradigms unresolved.Future research should develop better-calibrated expectations, examine broader scientific foresight, and investigate mechanisms underlying systematic forecasting failures.
Methods · The CUSP suite
CUSP is a temporally grounded, multidisciplinary suite that evaluates scientific forecasting under controlled knowledge cutoffs using verifiable event-level tasks. It combines retrospective forecasting benchmarks with a prospective Time Capsule for analyzing how models predict future scientific progress.
- The CUSP suite: CUSP uses a temporally stratified corpus of scientific events from January 2024 to March 2026, with domain-specific criteria requiring incorporated events to represent verifiable, definitively resolved advances.Natural-science events come from high-impact peer-reviewed publications, while multiple scholarly databases establish the earliest observed manuscript date as the knowledge boundary.
- Question Types and Synthesis: CUSP decomposes scientific forecasting into four dimensions and constructs four complementary tasks for each accepted milestone.Binary prediction tests feasibility and calibration; multiple choice probes mechanistic forecasting; and free response evaluates generative solution design.
- Question Types and Synthesis: Task construction removes post-cutoff identifiers and narratives, then uses independent LLM judging and human expert review to verify fidelity, objective verifiability, and reliable targets.Each abstract is decomposed into a problem statement, technical approach, and results summary before validation.
- CUSP Time Capsule Construction: CUSP Time Capsule introduces prospective questions whose real-world outcomes are not yet known at evaluation time but are designed to become authoritatively and unambiguously verifiable.Because ground truth is unavailable, it is used to analyze prediction consistency and confidence rather than accuracy.
- CUSP Time Capsule Construction: Time Capsule questions span scientific benchmarks, institutional recognitions, future technological events, and AI capability forecasting conditioned on current state-of-the-art results.Human experts curate the questions for relevance and verifiability, extending CUSP from retrospective evaluation to prospective analysis.
- Key Statistics of CUSP: 4,760 scientific events generate 17,429 structured forecasting tasks across eight top-level scientific domains and 4,245 distinct subcategories.Biology and artificial intelligence are the largest represented domains, with 1,234 and 1,141 papers respectively.
- Comparison to Related Evaluation: CUSP differs from general-world forecasting and retrospective scientific-reasoning benchmarks by combining scientific grounding with temporal knowledge constraints and verifiable discovery-linked targets.The related approaches do not evaluate scientific forecasting under temporal knowledge constraints.
1 Model Evaluation
CUSP uses a two-track evaluation framework that combines deterministic forecasting scores with rubric-based assessment of scientific reasoning. The framework evaluates task-specific outcomes and temporally valid free-response proposals, including explicit detection of post-cutoff information.
- Evaluation framework: CUSP combines deterministic outcome scoring with rubric-based scientific reasoning evaluation to assess both forecast correctness and generated scientific reasoning.The design addresses the possibility that correct final answers may rely on flawed, unfaithful, or post-hoc reasoning.
- Task coverage: Each scientific advance can include binary classification, multiple-choice, free-response, and date-prediction tasks, with scoring applied only to tasks present for that event.Binary classification uses original and negation-perturbed statements and reports a merged score averaged across both variants.
- Track I. Deterministic outcome evaluation: Track I scores forecasting accuracy deterministically using exact binary agreement, semantic matching for MCQs when needed, and exponential-decay distance for date predictions.The date-prediction metric is exp(−0.1d), where d represents the distance from the ground-truth date.
- Track II. Free-response scientific reasoning evaluation: Track II evaluates open-ended solution strategies with a rubric-based LLM judge, because multiple valid solutions make exact-match scoring inappropriate.The protocol uses GPT-5.4-mini augmented with agentic web search under a strict temporal cutoff.
- Track II. Free-response scientific reasoning evaluation: FRQ scoring first detects post-cutoff information and withholds valid forecasting credit from contaminated responses before evaluating non-contaminated proposal quality.Proposal quality is assessed across alignment, specificity, novelty, and feasibility.
Author contributions statement
The authors describe equal contributions, study conception and evaluation-suite design, and distributed responsibilities across data curation, experimentation, analysis, discussion, and manuscript revision.
- Sean Wu and Pan Lu contributed equally, while Junchi Yu conceived the study with input from Philip Torr and James Zou.
- Sean Wu, Pan Lu, and Junchi Yu designed the evaluation suite; Wu led data curation, task synthesis, validation, and experiments with contributions from Pan Lu and Yupeng Chen.
- Peter Clark, Yutaro Yamada, and Jonathan Bragg guided analysis and evaluation design, while David Clifton contributed to scientific discussion and manuscript revision.
Supplementary Information for Scientific reasoning does not reliably translate into scientific forecasting in frontier AI … A.3 Task synthesis procedure
CUSP constructs standardized scientific-forecasting tasks from rigorously filtered milestones across natural science, AI, and leaderboard data. Its synthesis pipeline separates each milestone into problem, approach, and measurable results before generating multiple evaluation formats.
- A.1 Data acquisition and source construction: Natural-science milestones come from publication logs of Nature, Science, and Cell, restricted to high-impact peer-reviewed journals.The corpus targets foundational breakthroughs in biology, chemistry, and physics as rigorously validated empirical discoveries.
- A.1 Data acquisition and source construction: CUSP includes high-visibility AI papers from weekly community lists and Hugging Face Top Papers, ranked using an age-adjusted hybrid impact score.The score is upvotes + 5 × [citations/(months old + 1)].
- A.1 Data acquisition and source construction: Leaderboard targets include GPQA Diamond, MMLU-Pro, Humanity’s Last Exam, and domain-specific benchmarks that provide standardized, time-resolved capability measurements.These targets test whether models can extrapolate rapid machine-learning progress.
- A.2 Corpus Extraction and Automated Filtering: Milestones are extracted from top-tier papers using only titles, publication metadata, and abstracts to ensure standardized and reproducible representations.Abstracts summarize primary claims, methods, and quantitative outcomes while remaining uniformly accessible across venues.
- A.2 Corpus Extraction and Automated Filtering: An LLM-based pipeline retains only candidate milestones containing at least one verifiable, measurable outcome, producing well-defined prediction targets.Extraction identifies concrete results or capabilities, while filtering applies a strict binary acceptance decision.
- A.2 Corpus Extraction and Automated Filtering: Domain-aware criteria require validated benchmark advances or explicit performance metrics in computational research and measurable quantities in experimental sciences.The pipeline also rejects purely descriptive, speculative, or review-oriented abstracts to improve predictive-evaluation suitability.
- A.3 Task synthesis procedure: Each accepted milestone is decomposed into a problem statement, technical approach, and results-and-metrics field before task generation.Novel acronyms, method names, and system names are avoided so historical-cutoff models cannot recognize discoveries through post-cutoff terminology.
- A.3 Task synthesis procedure: The pipeline generates five formats: binary feasibility, perturbed binary, multiple-choice approach selection, free-response strategy, and date prediction questions.These formats assess feasibility, robustness to plausible alternatives, technical-approach identification, implementation proposals, and timing judgments.
A.4 Key Statistics and Distributional Properties
CUSP is a temporally grounded, diverse benchmark spanning 27 months, 17,429 validated tasks, eight scientific domains, and 4,245 subcategories. Its task formats and linguistic complexity are intentionally non-uniform, reflecting strict validation and increasing reasoning demands.
- Temporal Distribution: CUSP spans January 2024–March 2026 across 27 active months and includes 4,760 timestamped milestones, averaging 176.3 papers monthly (± 51.5).All 27 months are represented, enabling evaluation across short- and long-term forecasting horizons.
- Task Composition: 17,429 validated task instances cover four task types, with non-uniform distributions caused by filtering for verifiability, faithfulness, and logical consistency.CUSP prioritizes task reliability and scientific validity over uniformity.
- Label Distribution: Binary questions have an overall yes/no label distribution of approximately 0.75, while MCQ answer options are shuffled to uniformly distribute correct answers.Binary ground truth is yes for standard questions and no for perturbed questions; MCQs initially assign the correct answer to option A.
- Question Length and Complexity: Average task lengths are 29.2 words for binary questions, 36.6 for MCQs, 41.8 for FRQs, and 70.6 for problem statements.Problem statements provide rich scientific context, FRQs require open-ended synthesis, and MCQs demand discriminative reasoning over expert-level distractors.
- Domain and Subcategory Diversity: CUSP spans eight scientific domains and 4,245 subcategories, led by biology with 1,234 papers and artificial intelligence with 1,141 papers.The remaining listed domain counts are medicine (746), neuroscience (403), materials science (375), physics (359), environmental science (235), chemistry (203), and other domains (64).
B Benchmark validation
CUSP uses an automated validation framework with an independent LLM judge to assess generated questions against their source abstracts. Validation checks task-specific faithfulness, evaluability, technical validity, and perturbation quality, followed by fine-grained filtering of invalid components.
- Validation framework: CUSP questions are validated against source abstracts by Grok-3, a model distinct from the GPT-4o system used for question generation.The framework is designed to maintain question faithfulness and quality as CUSP evolves with newly published discoveries.
- Binary and Perturbed Binary Questions: Binary questions are assessed for faithfulness to concrete claims and verifiability as objectively evaluable yes/no statements.Faithfulness checks entities, conditions, outcomes, and quantitative details, while verifiability rejects vague or underspecified claims.
- Binary and Perturbed Binary Questions: Perturbed binary questions must introduce meaningful, non-trivial modifications that are no longer directly supported by the source abstract.Examples include changing thresholds or adding unmet constraints, preventing trivial or paraphrased perturbations.
- Multiple-Choice Questions (MCQ): MCQ validation checks problem faithfulness, support for the correct technical approach, and plausibility without direct support for distractors.The process evaluates whether the question captures the scientific challenge and whether answer choices align with mechanisms supported or implied by the abstract.
- Free-Response Questions (FRQ): For FRQs, validation focuses on whether the problem context and background faithfully reflect the source abstract because multiple valid solutions make direct answer validation ambiguous.All criteria use structured LLM judgments, enabling field-level removal of invalid components while preserving the rest of each sample.
C Forecasting in a time capsule … E Extended Results
CUSP Time Capsule evaluates whether frontier models can forecast future scientific and AI progress beyond April 2026, while CUSP distinguishes forward-looking discovery prediction from retrospective scientific reasoning. Extended results show shared expectations of continued capability growth and emissions increases, but substantial variation in predicted magnitude and saturation timing across models.
- C Forecasting in a time capsule: CUSP Time Capsule asks frontier models to extrapolate scientific breakthroughs and AI progress beyond April 2026 using open-ended forecasting and capability prediction tasks.The evaluation covers both scientific forecasting and benchmark-based capability prediction.
- C Forecasting in a time capsule: All models forecast 2027 global CO2 emissions above 2025 levels, with GPT-5.4 predicting a slightly larger increase and LLaMA 3.3 and GPT-OSS the steepest growth.Claude S4.5, DeepSeek R1, and GPT-4o produce more conservative estimates close to the historical trend.
- C Forecasting in a time capsule: Across AI capability benchmarks, models expect continued gains through 2026-2027, but projected magnitudes vary substantially, with GPT-5.4 the most optimistic.The forecasts include Humanity’s Last Exam, GPQA Diamond, and MMMLU46, with especially optimistic GPT-5.4 predictions for 2027-10.
- C Forecasting in a time capsule: GPQA Diamond and MMMLU forecasts cluster near the upper performance bound, whereas DeepSeek R1 predicts smaller gains and earlier plateaus.The clustering suggests models expect GPQA Diamond and MMMLU to saturate within the next few generations.
- C Forecasting in a time capsule: CUSP Time Capsule is intended to reveal implicit scientific priors and future-oriented world models alongside factual recall and forecasting ability.The framework is presented as supporting analysis of multiple dimensions of frontier AI systems’ future-oriented knowledge.
- D.1 Comparison to Related Benchmarks: Prior forecasting benchmarks assess future-event prediction, while scientific reasoning benchmarks test expert knowledge, hypothesis generation, and structured reasoning over content whose answers are already known.ForecastBench36, FutureX37, FOReCAst38, and PROPHET39 represent forecasting efforts; Humanity’s Last Exam26, AstaBench25, PreScience28, ResearchBench27, and ScienceQA23 represent scientific benchmarks.
- D.1 Comparison to Related Benchmarks: CUSP differs by deriving tasks from real peer-reviewed breakthroughs, imposing pre-milestone temporal cutoffs, and requiring forward-looking prediction across multiple tasks.This design distinguishes CUSP from both general forecasting benchmarks and retrospective scientific reasoning evaluations.
- D.2 AI for Science: AI-for-Science systems increasingly support discovery workflows, but prior approaches typically depend on human researchers to define the problems and directions explored.This dependence leaves open the question of whether AI systems can independently identify promising scientific directions.
E.1 Full Web-search Results … F.1 Human evaluation of Dataset Validation
The supplementary analyses show that models can match broad scientific themes while missing exact constraints, overextending weak evidence, and misestimating the timing of advances. Results also reveal domain-specific FRQ limitations, calibration and response-bias issues, and stronger-than-human dataset-validation rigor from LLM judges.
- E.1 Full Web-search Results: GPT-5.4 produced false positives by overlooking an inserted < 2 nm constraint in a metal–oxide interaction forecast.The model matched broad scientific themes but incorrectly confirmed that the threshold was met.
- E.1 Full Web-search Results: GPT-5.4 also extrapolated from evidence supporting 85% to predict that a system would cross the rigid 90% threshold.This was classified as speculative hallucination or overconfident extrapolation.
- E.2 FRQ Sub-dimension Score Analysis: FRQ evaluation profiles Alignment, Specificity, and Novelty on 0–10 scales, with a positive Spec. −Align. gap indicating technically detailed responses that miss the paper’s method.This gap is identified as a signature of plausible-sounding hallucination.
- E.3 Results by Research Area: Chemistry and Physics consistently showed lower FRQ alignment, reflecting greater domain specificity and fewer overlapping concepts with general pretraining data.The finding is reported in the research-area analysis.
- GPT-4o: DeepSeek R1 correctly anticipated GLM-4.5’s RL-plus-MoE breakthrough and selected the eventual hybrid MoE-style mechanism, but predicted 2026-07 instead of 2025-08.Its binary prediction was Yes with confidence 0.70, while its mechanism prediction was B with confidence 0.85.
- E.4 Model Bias and Confidence: Most models were overconfident on MCQ, while all models were overconfident on date prediction and binary tasks; GPT-5.4 and Claude S4.5 remained relatively well calibrated on MCQ.ECE was computed over 10 equal-width bins, with lower ECE indicating better calibration.
- F.1 Human evaluation of Dataset Validation: Ten graduate-level researchers evaluated benchmark questions under the same keep-or-remove criteria, and LLM judges were on average more rigorous in removing invalid examples while retaining clean, verifiable ones.Evaluators had expertise spanning artificial intelligence, materials science, and chemistry.
F.2 Human Evaluation of LLM Judge … G.2 Verification Examples and Prompts
Human evaluation found moderate agreement between the AI judge and human FRQ scoring, while the verification pipeline used strict validators to assess faithfulness, objective verifiability, perturbation validity, and question quality. The supplied verification example illustrates rejection of a biological abstract because it lacked a concrete experimental result or measurable quantity.
- F.2 Human Evaluation of LLM Judge: 60 examples were evaluated by three human evaluators using the same rubric as the LLM judge across GPT-4o and GPT-OSS.The evaluators included two Computer Science PhDs and one postdoctoral scholar.
- F.2 Human Evaluation of LLM Judge: r = 0.34 (p < 0.01) Pearson correlation and ρ = 0.33 Spearman correlation accompanied the AI judge’s human FRQ scores.The mean absolute error was 0.75 points on a 0–10 scale.
- F.2 Human Evaluation of LLM Judge: +0.26 points was the AI judge’s mean Bland–Altman bias, with 95% limits of agreement spanning [−1.65, +2.17] points.The positive bias indicates that the AI judge was marginally more generous than human evaluators.
- G.2 Verification Examples and Prompts: The example elephant-whisker abstract was rejected because it did not provide a concrete experimental result or measurable biological quantity.The example was labeled Biology, dated February 2026, and marked Rejected under CUSP filtering.
- G.2 Verification Examples and Prompts: The verifiability validator requires claims to be concrete enough for objective yes/no judgment and rejects vague language without measurable criteria.Examples of acceptable claims specify metrics, benchmarks, thresholds, or evaluation protocols.
- G.2 Verification Examples and Prompts: The perturbation validator passes only meaningful changes to salient details that break direct support from the source, rejecting paraphrases and cosmetic rewordings.Changed elements include numbers, thresholds, entities, outcomes, time, scope, and conditions.
- G Benchmark Verification: MCQ validation separately assesses stem faithfulness and verifiability, correct-answer support, and distractor plausibility without judging unrelated components.Forecasting dates are ignored when evaluating the scientific core, while unsupported mechanisms, vague claims, or trivial distractors cause failure.
G.3 Human Vs AI in Benchmark Validation · H CUSP Evaluation Details · H.1 Evaluation System Prompts
The paper filters benchmark questions for source fidelity, then evaluates CUSP forecasts through deterministic outcome scoring and leakage gating. Its judges score scientific responses on alignment, specificity, novelty, and feasibility while auditing exact post-cutoff entity names for contamination.
- G.3 Human Vs AI in Benchmark Validation: Human validation removed questions with unsupported numerical thresholds or vague performance claims, while retaining concrete, objectively verifiable benchmark criteria.Examples include unsupported 30% token-cost and less-than-2% performance claims, versus a specific MMIU threshold of 55.7% accuracy.
- H CUSP Evaluation Details: CUSP’s deterministic track scores task outcomes by exact match, applies date decay using e−0.1|∆mo|, and adds rubric scores for forecasting-question alignment, specificity, novelty, and feasibility.The rubric is applied when a forecasting-response question is present.
- H CUSP Evaluation Details: Its leakage-gating track applies only to forecasting-response questions and uses a web-search judge to detect verbatim post-cutoff entities, withholding forecasting credit from contaminated responses.Responses judged contaminated receive no forecasting credit.
- H.1 Evaluation System Prompts: The contamination auditor checks only names copied verbatim from the LLM response, verifies whether each was first released after the cutoff, and requires that it could not have been independently invented.It immediately passes responses with no proper nouns, model names, paper titles, system names, or dataset names.
- H.1 Evaluation System Prompts: The auditor explicitly excludes methodology matches, correct predictions, and coincidentally matching numerical predictions from leakage, and returns “unclear” when evidence is uncertain.Web-search results are never treated as part of the LLM response.
I Example Benchmark Items … L Knowledge and Forecasting Gap Results
CUSP combines temporally grounded scientific forecasting items with benchmark-creation procedures and example responses spanning scientific milestones, technical breakthroughs, and proposed solutions. Its supplementary analyses distinguish gains from added pre-cutoff knowledge from gains enabled by retrospective post-cutoff information.
- I Example Benchmark Items: CUSP includes binary, perturbed-binary, technical multiple-choice, free-response, and date-prediction items tied to concrete scientific advances and target dates.Examples cover electromagnetic-interference shielding, unified biomolecular modeling, nanobody design, and Humanity’s Last Exam performance.
- J Benchmark Creation Criteria: Benchmark creation filters retain abstracts describing concrete methods, verifiable outcomes, measurable improvements, or defined capability milestones validated against recognized standards.The criteria span technical, material, biological, experimental, and theoretical breakthroughs, including independently reproducible or benchmarked results.
- J Benchmark Creation Criteria: The generation procedures separately extract outcomes, technical approaches, and problem statements before producing binary questions, counterfactual claims, expert MCQs, and free-response prompts.The instructions require measurable results in forecasting questions, preserve benchmark names in counterfactuals, and keep solution terminology out of question stems where specified.
- K Example FRQ Responses: Example free-response evaluations compare model proposals with realized technical approaches for repository generation, breath-condensate monitoring, echocardiography, and continuous depth estimation.The examples feature repository graphs, portable microfluidic-electrochemical sensing, multi-view cardiac modeling, and implicit neural depth fields.
- K Example FRQ Responses: The repository-generation example reports 81.5% functional coverage and a 69.7 metric, with performance described relative to Claude Code and other baselines.The passage states 3.9times the strongest baseline (Claude Code) and about 64times other baselines, while also describing dependency modeling and near-linear planning scaling.
- L Knowledge and Forecasting Gap Results: Supplementary analyses define the knowledge gap as improvement from pre-cutoff information and the forecasting gap as additional gain from post-cutoff information.The table states that the two gaps sum to the total web-search improvement over baseline and reports stratifications by citation-count quartile across binary, perturbed-binary, MCQ, date, and FRQ metrics.