Source-linked AI summary

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

Ziyue Wang, Aomufei Yuan, Yiran Yao, Linli Yao, Hongyao Zuo, Ziwen Gong, Yuanxin Liu, Shicheng Li, Yishuo Cai, Tong Yang, Xu Sun, Xiaohui Li, Haoli Bai

arXiv:2608.22948v1cs.CLcs.AI

TL;DR

Existing research-ideation evaluations lack a shared decision rule: free-form judgments can depend on style and position, while future-paper scoring rewards one realized trajectory. Lit2Test addresses this by evaluating six-field proposals that precommit a falsifying outcome in 200 prospective paper neighborhoods through blind, order-audited pairwise judgments. It recovers a strict model ordering in all 10,000 bootstrap replicates, with separation attributed to test and metric design rather than surface fluency.

  • Problem

    Existing ideation benchmarks do not measure whether a model can state what observation would prove its research idea wrong, while current paradigms lack a shared decision rule.

  • Method

    Lit2Test uses a six-field contract for grounded, executable proposals and evaluates four models on 200 prospective real-paper neighborhoods without future-paper answer keys.

  • Results

    A strict model ordering is recovered in all 10,000 case-level bootstrap replicates, with separation driven by proposed test and metric design above a falsifiability floor.

  • Takeaways & Limitations

    Lit2Test provides a versioned, audit-oriented benchmark for evaluating prospective research-proposal testability and releases its construction and audit artifacts.

  • Takeaways & Limitations

    Human calibration covers only 20 of 200 neighborhoods, the canonical judge remains a single LLM per version, and the protocol does not establish downstream experimental success.

Abstract

from arXiv · show

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.

1 Introduction

Lit2Test addresses the lack of a shared decision rule for evaluating research ideas by requiring proposals to specify what observation would falsify them. It combines this six-field contract with prospective, order-audited comparisons and reliability controls.

  • Motivation: A falsifiable proposal precommits a controlled comparison, decisive metric, and rejecting outcome, making quality experimentally decidable rather than merely stylistic.The benchmark measures whether models can turn a literature tension into a test.
  • Motivation: Free-form judging and future-paper scoring fail differently: the former tracks style and presentation position, while the latter rewards one realized research trajectory.Figure 1 shows the paradigms producing conflicting preferences on the same long-context QA neighborhood.
  • Benchmark design: The benchmark requires a six-field contract covering a literature gap, hypothesis, minimal test, decisive metric, and bidirectional supporting and falsifying outcomes.Its rubric favors small, executable designs and excludes realized execution outcomes.
  • Scale and reliability: The protocol audits ordering robustness, capability drivers, and human corroboration through folded aggregation, controls, diagnostics, and bounded calibration.These components are presented as empirical findings and reliability checks for the instrument.

2 Benchmark Construction

Lit2Test constructs a prospective benchmark from fixed real-paper neighborhoods and asks each model to produce a common six-field proposal. Pairwise comparisons therefore evaluate grounded, executable test design without future-paper answer keys.

  • Context construction: The benchmark uses 200 fixed literature contexts, each a real four-paper neighborhood containing a cross-paper tension that a small experiment could adjudicate.The contexts are drawn from 800 unique publications organized into five batches, with provenance and deduplication audits.
  • Task definition: Each model receives the same context and returns P = (literature_gap, hypothesis, minimal_test, decisive_metric, supporting_result, falsifying_result).The task covers literature synthesis, gap identification, hypothesis formulation, and executable minimal-test design.
  • Task definition: The six fields create a common decision structure: grounding ties proposals to the neighborhood, hypotheses specify claims and conditions, and minimal tests normalize proposal granularity.The contract is intended to make free-form ideas judgeable rather than reward generic brainstorming or the largest agenda.
  • Proposal generation: Four participant models generate one native six-field proposal per neighborhood under shared prompting and settings, yielding 800 proposals.Schema validation and bounded retries address malformed outputs before comparison.
  • Pair formation: The four proposals create 200 × 6 = 1,200 canonical pairs, each judged in original and reverse presentation orders.An orientation gate verifies that reverse tasks genuinely swap the two proposal contents.

3 Evaluation Protocol

The evaluation protocol uses blind pairwise judgments, folds reversed presentations into case-level outcomes, and aggregates only order-stable comparisons. Diagnostic controls and explicit scope limits calibrate what the resulting ranking can support.

  • Blind pairwise judgment: A fixed non-participant judge compares two anonymized six-field proposals for the same neighborhood and returns A, B, or TIE under an explicit rubric.Holistic verdicts are the canonical observable because they match the ordinal comparison used for aggregation.
  • Order folding: Each canonical pair is judged in both presentation orders; agreement yields an order-stable case, while a reversal or tie in either orientation makes it order-sensitive.The canonical pair, not the ordered row, remains the statistical unit.
  • Aggregation: Only order-stable outcomes feed Bradley–Terry estimation and Condorcet summaries, with uncertainty estimated by resampling canonical cases in 10,000 bootstrap replicates.This avoids treating forward and reverse judgments as independent observations.
  • Diagnostic controls: The protocol audits construct validity with dimension-decomposed scoring, hidden naive baselines, same-source rendering controls, and single-field corruption checks.The controls test rubric alignment, real-versus-naive discrimination, formatting effects, and sensitivity to obvious defects.
  • Scope and limitations: Human labels are calibration evidence on a stratified subset, while benchmark-wide human validation and downstream experimental success remain outside the protocol’s scope.The canonical judge remains single per benchmark version, and execution success is not established by measured testability.

4 Experiments and Empirical Findings

Lit2Test recovers a strict and highly robust ordering of four models, while controlled audits indicate that judgments track proposal substance rather than presentation style. Human calibration corroborates the aggregate ranking and tier structure, but modest inter-annotator agreement limits claims of exhaustive validation.

  • RQ1: Ordering and robustness: 950 of 1,200 canonical cases are order-stable, while 250 order-sensitive cases are isolated rather than included in the aggregate ordering.Treating sensitive cases as ties leaves the ordering unchanged; adversarial assignment preserves the two-tier structure.
  • RQ1: Ordering and robustness: GPT-5.2 > Claude Sonnet 4.6 > GLM-5 > DeepSeek-V3.2, with the full ordering recovered in all 10,000 case-level bootstrap replicates.The ordering remains identical under neighborhood-cluster resampling and across all five construction batches.
  • RQ2: Drivers of the ordering: Minimality/feasibility provides the strongest natural dimension-level diagnostic signal, while falsifiability acts primarily as an admissibility floor.Falsifiability yields non-tied dimension verdicts in only 33/180 judgments and is decisive in 1/180 comparisons, although manipulation checks show the judge enforces it when degraded.
  • RQ2: Drivers of the ordering: All 40/40 ordered comparisons per dimension favor clean proposals after clear corruption of grounding, decisive metric, or falsifiability.Mean target-score drops are 1.73, 2.00, and 1.95 respectively, with near-zero non-target drift.
  • RQ2: Drivers of the ordering: 0 clean wins, 40 ties, and 0 sham wins per dimension show that style-matched rewriting alone does not move the judge.After sham adjustment, clean proposals retain positive preferences for grounding, decisive metric, and falsifiability, with confidence intervals excluding zero.
  • RQ3: Human corroboration and its boundary: 87.2% of decisive stable human majorities agree with the judge, and 88.3% of human bootstrap rankings differ from the judge's ordering by at most one inversion.Humans detect naive controls in 11/12 judgments, while inter-annotator agreement remains modest: Krippendorff’s α = 0.238 for winner selection and 0.127 for neighborhood screening.
  • Qualitative case studies: Qualitative cases trace comparisons to contract fields: resource matching, falsifying outcomes, mechanism-isolating metrics, and order sensitivity determine the adjudication.A presentation reversal causes folding to mark one minimality-versus-grounding comparison as order-sensitive rather than forcing a winner.

5 Related Work

Lit2Test differs from neighboring ideation benchmarks by combining a prospective six-field proposal contract with order-audited pairwise evaluation. Its unit requires a minimal test, adjudicating metric, and explicit supporting and falsifying outcomes.

  • Closest benchmarks: Neighboring benchmarks evaluate open-ended idea quality, realized-trajectory recovery, or data-explanatory hypothesis generation.These targets use different scoring objects and reference signals.
  • Falsifiability: Lit2Test adapts Registered Report-style precommitment into a model-facing unit requiring bidirectional outcomes at proposal time.The contract places falsifiability inside the judged answer rather than only in generation or execution.
  • Measurement foundations: Pairwise comparison, order reversal, human calibration, and ordinal aggregation are established primitives rather than isolated contributions of Lit2Test.The paper claims novelty in composing these primitives around its evaluation unit.
  • Lit2Test positioning: The benchmark combines a complete six-field contract with prospective, order-audited pairwise comparison.This combination addresses the gap between systems with falsification machinery and evaluations with careful reliability protocols.

6 Discussion and Limitations

The benchmark supports conclusions about measurement reliability and testability within its studied setting, while leaving several validation and generalization questions open. The authors identify bounded human coverage, a single canonical judge, limited audit scope, absent execution outcomes, and an ML-focused domain.

  • Limitations: Human calibration covers 20 of 200 neighborhoods with three annotators, supporting aggregate conclusions but not benchmark-wide validation.The authors identify expanded human coverage as the direct remedy.
  • Limitations: The canonical judge remains a single LLM per version despite reliability audits and reproduction by a second independent judge.The limitation concerns judge multiplicity, not the absence of auditing.
  • Limitations: Measured testability is a prerequisite for downstream success, not a predictor of it, because the benchmark includes no execution of proposed tests.Connecting prospective testability to execution-based validation is listed as future work.
  • Scope: The 200 neighborhoods come from ML-adjacent literature, so broader domains and temporal splits remain extensions rather than redesigns.The pipeline is described as replicable for those extensions.
  • Disclosure: AI assistants supported writing polish, figure drafting, and literature cross-checking, with the authors stating that they verified all content.The statement describes assistance and author responsibility.

7 Conclusion

Lit2Test addresses the absence of a shared decision rule for evaluating research ideas by requiring six-field, falsifiable next-step tests. The benchmark is prospective, auditable, and released with its construction and audit artifacts.

  • 7 Conclusion: Existing ideation benchmarks leave unmeasured whether models can state what would prove their own ideas wrong.Lit2Test targets this capability rather than generic quality, realized-trajectory recovery, or forecast performance.
  • 7 Conclusion: Each proposal contains a literature gap, hypothesis, minimal test, decisive metric, supporting result, and falsifying result.The benchmark requires these six fields in a single JSON object.
  • 7 Conclusion: The supporting result confirms the hypothesis, whereas the falsifying result specifies the observation that would reject it.These two fields make the proposal’s decision rule bidirectional.
  • 7 Conclusion: The benchmark uses one minimal falsifiable next-step test from a fixed literature context and requires valid JSON with exactly six fields.The generation prompt supplies research context, an open problem, resource constraints, and paper materials.
  • 7 Conclusion: The running example is a four-paper ICLR neighborhood on long-context question answering, including LooGLE, ALR2, ChatQA 2, and LongPack.Its open problem asks whether retrieve-then-reason helps beyond additional retrieval under public-resource constraints.
  • 7 Conclusion: The dataset comprises five construction batches of 40 contexts drawn from recent OpenReview/ICLR-adjacent machine-learning literature.Deduplication prevents paper repeats across contexts, while source matching verifies topical coherence.
  • 7 Conclusion: The design, construction, and code snapshots predate the IdeaSpark posting, but the main experiment finished two days afterward.The authors therefore do not claim that the experimental results predate the posting.

C.1 Six-Pair Folded Detail

Table 3 reports complete folded head-to-head outcomes from the primary judge, separating order-stable wins and losses from order-sensitive cases. The resulting Condorcet relations are strict and consistent across construction batches.

  • C.1 Six-Pair Folded Detail: W counts order-stable wins, L counts order-stable losses, and T counts order-sensitive cases for the row model.These counts define the table’s folded head-to-head layout.
  • C.1 Six-Pair Folded Detail: Table 3 reports six folded pairwise outcomes from the primary judge.Each pair contains 200 folded cases.
  • C.1 Six-Pair Folded Detail: Each higher-ranked model strictly wins its head-to-head against every lower-ranked model, with no Condorcet cycle.The ordering is consistent across all five construction batches.

C.2 Tie-Sensitivity Analysis

The ranking remains robust when order-sensitive cases are treated conservatively or adversarially, while cluster resampling widens uncertainty without changing the modal ranking. The sham-adjusted audit separates content degradation from style effects, though construction checks show moderate validator agreement and limited naturalness leakage.

  • Tie-Sensitivity Analysis: 1.0 bootstrap recovery preserves the reference ordering when all 250 order-sensitive cases are treated as ties.All six head-to-head majorities remain strict, and the Condorcet and Bradley–Terry structures are unchanged.
  • Tie-Sensitivity Analysis: The two-tier structure survives the adversarial assignment of every sensitive case to the lower-ranked model, although both within-tier adjacencies flip.The non-adjacent relations retain margins of at least 86/200.
  • Context-Cluster Bootstrap: 11–27% wider confidence intervals under context-cluster bootstrap reflect within-context correlation, but the ranking is recovered in all 10,000 replicates.The cluster bootstrap resamples 200 neighborhoods while retaining all six pairs per neighborhood.
  • Sham-Adjusted Preference: The sham-adjusted preference estimates content sensitivity by subtracting clean wins against style-matched sham edits from clean wins against subtle corruptions.Positive values indicate sensitivity to degradation net of style effects.
  • Sham-Adjusted Preference: 0 clean wins, 40 ties, and 0 sham wins per dimension pass the sham-equivalence gate before interpreting adjusted preferences.This control checks that style-matched rewriting alone does not move the judge.
  • Construction Imperfections: Validator agreement on corruption targeting is 65.6%, with 62 of 180 final judgments requiring third-call tie-breaking.The authors attribute this moderate agreement to ambiguity in naturalistic defects.
  • Construction Imperfections: 3.9% of non-unknown screening responses identify the edited side, indicating limited stylistic leakage that the sham contrast is designed to bound.The leakage rate is 14/362.

D.4 Falsifier On/Off Ablation

The falsifier field functions as an auditability constraint rather than an independent quality shortcut. Dimension-level aggregation largely matches holistic judgments, while human agreement remains modest despite stable leave-one-annotator-out conclusions.

  • Falsifier On/Off Ablation: Filling the falsifying-result field alone does not make the judge prefer an otherwise identical proposal.The field enables contract checking rather than directly improving generation quality.
  • Dimension-Decomposed Audit: 152/180 ordered judgments (84.4%) agree between structured dimension verdicts and the canonical holistic verdict.Equal-weight score-sum aggregation agrees on 145/180, while a top-3 variant agrees on 170/180.
  • Human Calibration: Three annotators evaluated 20 stratified neighborhoods containing 90 real model pairs and 4 hidden real-versus-naive controls.Their rubric mirrored the judge rubric across five proposal dimensions.
  • Human Calibration: Krippendorff’s α is 0.238 for winner selection and 0.127 for neighborhood screening, indicating modest human agreement.Majority aggregation and case-level bootstrap partially absorb individual disagreement.
  • Human Calibration: 86.7–87.2% stable-case agreement and an unchanged Bradley–Terry ordering persist when each annotator is removed in turn.The GPT-5.2/Claude Sonnet 4.6 versus GLM-5/DeepSeek-V3.2 tier split also remains unchanged.

E.5 Second-Judge Replication

An independent judge reproduces Lit2Test’s model ordering, supporting the ranking’s reliability beyond the canonical judge. The appendix situates this replication alongside diagnostic controls, human calibration, and the benchmark’s broader comparison framework.

  • Diagnostic Controls and Human Calibration: Table 5 consolidates diagnostic controls and human calibration that audit the benchmark’s measurement protocol.These include hidden controls, corruption checks, structured-versus-holistic comparisons, and human corroboration.
  • Second-Judge Replication: The independent judge reproduces the ranking GPT-5.2 > Claude Sonnet 4.6 > GLM-5 > DeepSeek-V3.2.The replication uses all 2,400 ordered comparisons with the same prompt and pair content.
  • Second-Judge Replication: The independent judge yields 980 order-stable and 220 order-sensitive cases, compared with 950 and 250 for the primary judge.Ordered-judgment agreement on the primary judge’s stable subset is 86.1%.
  • Benchmark Positioning: The benchmark compares 27 systems across 16 dimensions, including decision rules, answer keys, evaluation methods, reliability primitives, and falsifiability treatment.The full comparison matrix is provided as a CSV survey artifact.
  • Key Finding: Lit2Test combines a bidirectional decision rule with order-audited pairwise judgment without using a future-paper answer key.The authors identify this combination as absent from surveyed systems.

G.2 Randomness and Reproducibility

The reproducibility strategy acknowledges nondeterministic closed-source APIs while making downstream analyses deterministic from archived outputs. The paper also limits human-study and benchmark-use claims to their stated scope.

  • Randomness and Reproducibility: Closed-source model APIs do not guarantee bit-exact reproducibility across calls.The authors therefore fix and document construction and bootstrap seeds, archive raw API responses, and rerun analyses from frozen outputs.
  • Randomness and Reproducibility: Archived API responses and deterministic analysis scripts make downstream results fully reproducible from frozen outputs.Documented seeds include 20260620, 20260721, and 20260729.
  • Study Scope: The human study uses three senior computer-science undergraduates and collects no personally identifying information beyond annotation files.Annotators were compensated at standard research-assistant rates.
  • Intended Use: Lit2Test is intended as a diagnostic benchmark, not a sole criterion for evaluating researchers, funding proposals, or academic merit.Its source papers are publicly available on OpenReview, with no proprietary or restricted-access documents included.
Loading 2608.22948v1…