Source-linked AI summary

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See

arXiv:2609.07611v1cs.AIcs.CL

TL;DR

Scientific ideation benchmarks typically ask models to synthesize from curated references, leaving literature exploration and its role in ideation undermeasured. AgentIdeaBench compares matched static observation with active, tool-mediated exploration using literature-verified critics. Active evaluation reveals more capability headroom, scales about twice as fast, and improves grounding dimensions without changing measured originality, while Scientific World Modeling remains capability-gated and unconfirmed.

  • Problem

    Existing ideation evaluations largely use static, curated reference sets, whereas autonomous scientific work interleaves literature retrieval with reasoning and revision.

  • Method

    AgentIdeaBench evaluates matched Static and Active settings across multidisciplinary subfields, scoring hypotheses with critics that verify originality against retrieved prior art.

  • Results

    Active evaluation reveals capability headroom, scales about twice as fast as Static, and improves feasibility, clarity, and specificity while measured originality remains unchanged.

  • Takeaways & Limitations

    AgentIdeaBench provides an agent-era measurement basis showing that active exploration better differentiates capability and grounds hypotheses, especially for stronger models.

  • Takeaways & Limitations

    Originality validation is limited, and the Static–Active contrast also bundles multi-turn interaction and tool-use competence with retrieval control.

Abstract

from arXiv · show

Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.

1 Introduction

Scientific ideation remains a critical, underdeveloped part of autonomous AI science, while prevailing evaluations largely replace literature exploration with one-pass synthesis from curated references. AgentIdeaBench addresses this gap by matching static observation against active, tool-mediated exploration and finds that active evaluation reveals capability headroom that static scoring conceals.

  • Motivation: Scientific ideation sets the direction and bounds of downstream scientific execution but remains underdeveloped and difficult to evaluate rigorously.The paper frames ideation as a foundational step whose quality constrains what later research activities can achieve.
  • Motivation: Existing ideation benchmarks usually provide a fixed reference set for one-pass synthesis, unlike research workflows that interleave retrieval, reading decisions, and hypothesis revision.This static-observation paradigm measures recombination of supplied references rather than whether a model can gather its own evidence.
  • Benchmark contribution: AgentIdeaBench makes literature-access mode the experimental variable by comparing matched Static and Active settings with shared prompts, formats, and scoring.Static provides curated references, whereas Active gives only the subfield name and a search tool under a fixed interaction budget.
  • Main findings: Active evaluation improves about twice as fast as Static across knowledge cutoffs, at +1.16 versus +0.54 points per year, and its gain is capability-gated.Static ability predicts Active gain at r=+0.69; stronger models benefit more, while weaker models can be hurt.
  • Scientific World Modeling: Scientific World Modeling helps smaller backbones, but its benefit shrinks with backbone strength and is unconfirmed after multiple-comparison and compute-matched controls.The paper presents SWM as an exploratory, capability-gated direction rather than an established improvement.
  • Main findings: Agent-controlled retrieval raises feasibility, clarity, and specificity while leaving measured originality unchanged under the literature-verified critics.The reported originality change is −0.14 and is not significant.

2 Related Work

Related work evaluates scientific ideation through human studies, curated-reference benchmarks, keyword prompts, data-induced hypotheses, and structured generation methods. AgentIdeaBench targets recurring gaps by letting models choose supporting literature and grounding novelty judgments in retrieved prior art.

  • Autonomous AI scientists: End-to-end AI scientist systems increasingly automate execution and writing, but ideation is still often supplied by human insight or predefined task specifications.AgentIdeaBench targets this under-measured step by asking whether models can gather evidence and originate hypotheses without that scaffolding.
  • Ideation benchmarks: Prior ideation work spans human-preference studies, curated-reference benchmarks, keyword-based divergent-thinking tests, data-induced hypotheses, and retrieval or critique methods.These approaches differ in inputs and evaluation targets but generally do not provide the full matched contrast introduced here.
  • Limitations of prior work: Two recurring limitations are that models either read references assembled for them or use no literature, and they do not decide what to read during generation.The related-work discussion identifies retrieval control as a missing part of existing ideation evaluation.

3 AGENTIDEABENCH

AgentIdeaBench is a matched, multidisciplinary benchmark whose central contrast is whether models receive curated references or gather evidence through their own tool calls. It combines active exploration with literature-grounded originality scoring and broad task coverage.

  • Benchmark design: AgentIdeaBench compares curated Static observation with agent-controlled Active retrieval while holding subfields, output format, and scoring matched.Its primary contrast is who controls retrieval of supporting literature.
  • Tasks and subfields: The benchmark covers 100 narrow subfields across computer science, physics, biology, chemistry, and medicine, with 40 densely scored subfields.Reference sets are assembled from Semantic Scholar, and target literature is selected to be unseen at training time for nearly all evaluated models.
  • Static and Active settings: Static supplies curated titles and abstracts, whereas Active supplies only the subfield name and a search/fetch tool under a fixed call budget before identical hypothesis generation.Both settings require one 80–150-word hypothesis paragraph.
  • Scope of the contrast: The Active–Static contrast bundles retrieval control with multi-turn interaction and tool-use competence rather than isolating retrieval control alone.Recall-only and replay tracks are used to bound this attribution issue.
  • Literature-grounded scoring: Each hypothesis receives 1–10 scores on originality, feasibility, clarity, impact, and specificity from critics using literature retrieved before judging.The critics must justify originality against retrieved evidence, while coherence and boilerplate checks limit keyword-stuffed proposals.
  • Validation: Human-landmark checks place rewritten landmark papers 1.00–2.24 points above discipline-level Active averages, while CORE-7 averages 7.71 above every model’s Active mean.The scorer was calibrated through a seven-round study.

4 Experimental Setup

The experimental setup fixes the roster, generation and scoring procedures, controls, and statistical tests before comparing matched Static and Active outputs. Analyses use paired cells, bootstrap uncertainty, family-aware checks, and control tracks to bound interpretation.

  • Roster: The study evaluates 35 models in total, including 33 with matched Static–Active outputs; headline analyses use 28 matched primary-roster models.The primary roster is predominantly open-weight, with five closed-source Gemini models held out for replication.
  • Scoring: Critics score hypotheses using an originality-dominant weighted aggregation after literature retrieval and originality justification.The scoring pipeline uses three open-weight critics, trimmed aggregation, and weights O:2, I:1.5, F:1, C:0.5, S:0.5.
  • Controls: Recall-only and replay controls separate the effects of merely having references from those of agent-controlled retrieval.The controls compare Static, recall-only, and Active on shared cells and refeed Active-surfaced references in passive Static format.
  • Statistical tests: Headline differences are paired over shared model-and-subfield cells, with Wilcoxon tests and bootstrap 95% confidence intervals from 5000 resamples.Additional checks address family clustering, leave-one-family-out stability, within-model permutation, disjoint subfield halves, and discrimination criteria.

5 Results and Analysis

Active evaluation reveals capability headroom that Static conceals, separating models more effectively and producing gains that increase with model capability. The gains primarily improve grounding-related dimensions rather than measured originality, while robustness checks identify scope and attribution limits.

  • 5.1 Active Evaluation Improves Model Separation: Active evaluation carries 4.4× the between-model score variance of Static across 28 primary-roster paired models.It distinguishes 62% of top-half pairs against Static’s 12%.
  • 5.1 Active Evaluation Improves Model Separation: 19 of 28 models improve under Active, with GLM-5.1 gaining +1.21 while Gemma-2 27B loses −0.82.The strongest models gain over a point, whereas several weak models are hurt.
  • 5.2 Active Gains Increase with Model Capability: The Active−Static gain correlates with static ability at Pearson r=+0.69, with weakest models losing −0.18 and strongest models gaining +0.76.The correlation replicates on five held-out Gemini models at within-family r=0.88.
  • 5.2 Active Gains Increase with Model Capability: Active scores scale +1.16 points per knowledge-cutoff year versus Static’s +0.54, with a track interaction of +0.62/year.Active is steeper in all five disciplines, and the interaction has 95% CI [0.41, 0.94].
  • 5.3 The Gain Lands in Grounding: Active retrieval raises feasibility by +1.21, clarity by +0.63, specificity by +0.58, and impact by +0.23, while originality remains flat at −0.14.The originality change is not significant, with 15/28 models improving at sign test p=0.85.
  • 5.4 Alternative Explanations and Robustness: Active exceeds Recall by +0.23, whereas Static does not significantly exceed Recall and Replay preserves only a small part of the Active−Static gap.Replay−Static is +0.08 at p=0.21, while Active−Replay is +0.26 at p<10−3.
  • 5.4 Alternative Explanations and Robustness: The benchmark’s cross-discipline claims are limited because score discrimination is strongest in computer science and physics and weakest in dense-literature fields.Additional checks also indicate that exploration breadth and target-paper memorization do not explain the score gain or cutoff trend.

6 Scientific World Modeling

Scientific World Modeling adds a generation-time thought experiment that critiques draft novelty and feasibility while preserving a novel core. Its apparent benefit is capability-gated and remains unconfirmed after correction and compute-matched controls.

  • Method: Active inference motivates Scientific World Modeling as a generation-time procedure for pressure-testing hypotheses beyond retrieval.The procedure uses structured feedback to critique novelty and feasibility separately while protecting a novel core.
  • Results: +0.61 is the qwen-9b gain from the strongest SWM design, compared with +0.36 for qwen-27b.The estimated benefit shrinks as backbone capability increases.
  • Results: +0.14 and −0.06 are the changes for the two frontier deepseek-v4 backbones, with p>0.6.The frontier results are consistent with no measurable SWM benefit.
  • Results: +0.47 originality and +0.67 feasibility are the reported qwen improvements where SWM helps.The gains occur together, while the value of explicit thought experiments falls for stronger backbones.
  • Controls and limitations: The pooled SWM gain does not survive Holm correction, and a compute-matched best-of-3 baseline meets or exceeds it.The paper therefore reports structured thought experiments as a capability-gated direction that remains unconfirmed.

7 Conclusion

The conclusion presents AgentIdeaBench as an agentic evaluation of scientific ideation that separates handed-over reading from self-directed evidence gathering. It reports that active retrieval exposes capability differences and improves grounding-related dimensions without changing measured originality.

  • Conclusion: AgentIdeaBench evaluates scientific ideation by comparing ideas formed from handed-over reading lists with ideas supported by self-gathered evidence.Both settings are scored by literature-verified critics.
  • Conclusion: Active retrieval raises feasibility, clarity, and specificity while leaving measured originality unchanged under the critic.The value of retrieval is capability-gated: it helps strong models and hurts weak ones.
  • Conclusion: Active scaling continues beyond the local plateau where Static evaluation levels off.The conclusion also reports that retrieval equalizes idea diversity without equalizing quality.

Limitations

The paper’s conclusions are constrained by LLM-based scoring, confounding in the Static–Active comparison, and an unconfirmed generation-time result. These limitations narrow how the reported effects should be interpreted.

  • Scoring: Every dimension score is an LLM judgment, and mitigation does not constitute validation.Originality has the weakest human-agreement support, so originality claims specifically concern measured originality under this critic.
  • Scoring: The most important missing validation is an expert study in which domain researchers rate idea novelty directly.The paper identifies this as validation it has not run.
  • Attribution: The Active condition bundles agent-controlled retrieval with multi-turn interaction and tool-use competence.Without a condition that supplies agent-selected references without agent-driven turns, retrieval control cannot be isolated.
  • Generation-time modeling: Scientific World Modeling is unconfirmed because its benefit is confined to part of a small backbone set and clears neither corrections nor tested baselines.The paper also leaves open whether a different loop would help stronger models.

Ethics Statement

The paper describes artifact handling, voluntary human annotation, risks of acting on untested hypotheses, and limits in the benchmark’s literature coverage. It also cautions that centroid cosine is not a valid measure of retrieval similarity or diversity within a subfield.

  • Data and artifacts: Released artifacts contain generated hypotheses, critic scores, and retrieval identifiers, but no redistributed publisher full texts.The literature side uses bibliographic metadata, titles, and abstracts returned by the Semantic Scholar API.
  • Human participation: The agreement study used a small number of voluntary unpaid NLP researchers and collected ratings, optional comments, and timestamps without demographic data.Annotators saw anonymized hypotheses with model identities and critic scores withheld.
  • Risks and intended use: The benchmark’s hypotheses are untested and may be incorrect, unoriginal, or unsafe to act on, especially in medicine, chemistry, and biology.The authors say scores should not gate funding, publication, or priority claims, or substitute for expert review.
  • Risks and intended use: The models and retrieval index over-represent English-language, well-indexed, highly cited work.This limits how broadly the benchmark’s ideas and prior-art coverage should be interpreted.
  • Metric critique: 0.032 mean Jaccard overlap and 0.868 centroid cosine show that cosine similarity can remain high despite nearly disjoint reference sets.The paper attributes the apparent similarity to embedding anisotropy and mean-pooling, not retrieval similarity or diversity.
  • Metric critique: Static and Active sets have effective-distinct-counts of 4.08 and 8.31, respectively, indicating roughly twice the measured diversity.Set-level metrics, rather than centroid cosine, separate the regimes more cleanly.

B Judge validity and robustness

The benchmark’s judge and protocol checks support robust comparisons across models, while diversity and world-modeling analyses characterize what active exploration changes. Validation is strongest for critic agreement and weaker for originality and some disciplinary rankings.

  • Robustness: r=0.65–0.79 for the capability gate across alternative weightings, with p<10^-3 and Spearman ≥0.988 ranking agreement.The gate remains under equal, originality-only, and reduced-dimension weighting schemes.
  • Robustness: 0.95–0.99 inter-judge ranking agreement persists across single judges, and excluding self-graded model families leaves results essentially unchanged.The ensemble-level capability gate is therefore not dependent on one judge or apparent self-preference.
  • Robustness: κ=0.97 for binary pairwise critic agreement and κ=0.81 when ties are allowed, with 5.7% position-bias flips and 93% majority–single-critic agreement.The pairwise protocol used 100 hypothesis pairs in both presentation orders.
  • Diversity: Active exploration increases within-cell diversity across Vendi, cosine distance, self-ROUGE-L distinctness, and distinct-2, all at Wilcoxon p<10^-6.Across n=2793 paired cells, Vendi rises 1.943→2.092 and self-ROUGE-L falls 0.298→0.199.
  • Scientific World Modeling: The SWM variant adds SIMULATE before FINAL, while its structured schema diagnoses mechanism consistency, thought experiments, weaknesses, repairs, knowledge gaps, and next probes.A deterministic referee searches, refines, or accepts according to inconsistency, missing facts, and severity thresholds.
  • Scientific World Modeling: S5 suppresses feasibility repair below severity 2, but gating does not recover saturated backbones and its pilot’s positive sign was a small-sample artifact.The documented negative result preserves the originality cost observed for always-on repair.

D Critic design and calibration

The critic combines literature-grounded originality assessment with calibration and safeguards against scoring artifacts. Validation separates models from controls and anchors most clearly in concept-driven fields, while dense-literature disciplines remain softer.

  • Calibration: The literature-verified critic is supported by a seven-round calibration study because benchmark validity rests on it.Its purpose is to ground scoring in literature-based validation rather than uncalibrated judgment.
  • Calibration: 0% of rewritten landmark hypotheses exceeded 6.5 under the naive critic, while their mean was 4.78 versus 5.20 for all models, exposing a scorer-related ceiling.Fifteen ICLR-2026 Oral anchors averaged 5.54, only +0.34 over models, with 0% above 6.5.
  • Critic design: The critic’s two main failure modes are hindsight suppression of known landmarks and over-crediting fluent recombinations without findable prior art.Detail density can also reward named entities and precise restatements of known mechanisms, a bias strong models exploit.
  • Critic design: Date-filtered prior-art retrieval prevents retrieved ideas from being called novel, while revised clarity and specificity target precision of the new contribution.The rubric also calibrates feasibility for theory and algorithm ideas and ties impact to evidence tiers.
  • Validation: Spearman 0.805 between selection and held-out critics rules out same-family self-grading in the reported check.This provides an additional validation signal beyond the primary critic scores.
  • Critic design: Coherence and boilerplate caps change mean Originality by under 0.05 on three backbones but by −0.39 on qwen3.5-9b.The caps target incoherent, keyword-stuffed, or template-like proposals and rarely bind on coherent proposals.
  • Domain validity: Top-10 anchor–model gaps are +1.50 in computer science and +0.88 in physics, versus +0.43 in medicine, +0.38 in chemistry, and +0.28 in biology.The benchmark treats computer science and physics as more reliable regimes than experiment-heavy disciplines.
  • Domain validity: The cutoff-axis findings are not explained by memorizing target papers: retained sources postdate nearly all model cutoffs, and cutoff-gap regressions are flat.Static is +0.007/year at p=0.75 and Active is −0.012/year at p=0.85 in the model-demeaned regression.

F Tool-budget saturation

Across five open-weight backbones, additional Active tool calls produce bounded and backbone-specific gains rather than a universal monotonic improvement. The operating budget of 10 lies within the jointly flat interval for all backbones, but it understates one backbone and utilization does not explain the pattern.

  • Saturation: From 5 to 10 calls, every backbone is individually flat, with |d| ≤0.15 and p ≥0.30.From 1 to 5 calls, qwen3.5-27b gains +0.72 [+0.25, +1.20] and mimo-v2.5 gains +0.81 [+0.39, +1.18].
  • Saturation: From 10 to 20 calls, qwen3.5-9b gains +0.45 [+0.20, +0.74], while the other four remain within [−0.31, +0.11].Across six budgets, total spread ranges from 0.22 for gemma-4-31b to 1.07 for mimo-v2.5.
  • Saturation: The jointly flat interval is 5 to 10 calls, but budget 10 understates qwen3.5-9b rather than being flat for every backbone.gemma-4-31b is the only backbone whose best cell is the operating budget of 10.
  • Utilization: At budget 20, utilization spans 21% to 38% and does not align with benefit: qwen3.5-9b gains while the heavier-spending mimo-v2.5 declines by −0.31.No backbone comes close to the cap, so score flattening is not explained by models uniformly exhausting the available budget.
  • Protocol failures: Protocol failures increase for qwen3.5-397b from 0% at budgets 1, 2, and 5 to 24% at budget 20, motivating its exclusion from the SWM probe.Failures scatter across papers, leaving all 30 cells with all 10 papers.

K Recall-only control (closed-book)

The recall-only control shows that passive reference presence does not explain the Active advantage: Static is comparable to Recall, while Active performs significantly better. A replay decomposition further attributes the gain to agentic retrieval process rather than passively delivered reference content.

  • Active scored 5.05 versus 4.82 for Recall and 4.72 for Static across 280 shared cells.
  • Static−Recall was −0.09, non-significant at p=0.11, whereas Active−Recall was +0.23 at p<10^-3.
  • Active−Static was +0.33 at p<10^-3, confirming the Active advantage over both comparison regimes.
  • Replay−Static was +0.08 and non-significant, while Active−Replay was +0.26 at p<10^-3.
  • At the model level, Replay−Static was +0.08 at p=0.19, compared with Active−Replay of +0.25 at p=0.02.
  • The decomposition leaves multi-turn interaction bundled with tool-execution competence inside the process term.

N Compute-matched baseline for SWM

At matched compute, naive best-of-3 resampling outperforms the structured SWM loop, so the observed SWM gain cannot be attributed to its aggregation structure. An illustrative probe suggests frontier models may already perform some SWM-like checks internally, while mid-capability models benefit more from the external structure, though the probe has important limitations.

  • Compute-matched baseline: S4b uses roughly 21 LLM calls per rollout versus about 7 for single Active, motivating comparison with a three-sample best-of-3 baseline.S4b averages 8.4 outer turns plus approximately 13 internal expert-panel calls, while single Active averages 5.8 tool turns plus final synthesis.
  • Results: 0.79 points: best-of-3 exceeds single Active with win rate 0.64, p<10−3, and 95% CI [+0.66, +0.93].The resampling advantage is positive across all four backbones, ranging from +0.65 to +0.95.
  • Results: −0.55 points: S4b falls below best-of-3 with win rate 0.35, p<10−3, and 95% CI [−0.73, −0.37].The difference is significant on three of four backbones; qwen-9b is statistically tied at −0.05 and p=0.92.
  • Results: +0.245 points: S4b exceeds single Active with win rate 0.57, p=0.042, and 95% CI [+0.03, +0.46].This nominal gain is much smaller than the +0.79 produced by compute-matched naive resampling.
  • Interpretation: Because best-of-3 matches or exceeds S4b on every backbone, the SWM gain cannot be attributed to its aggregation structure.The best-of-3 comparison is an oracle upper bound because it selects using the same critic score later reported and is unavailable at inference.
  • Illustrative mechanism probe: SWM helps mid-capability qwen backbones but leaves frontier deepseek-v4 backbones unchanged, plausibly because frontier models already perform novelty and mechanism checks internally.A small thinking-enabled probe found deepseek-v4-pro spontaneously auditing prior work and testing mechanisms, whereas qwen3.5-9b generated more candidate ideas but rarely stress-tested its selection.
Loading 2609.07611v1…