Source-linked AI summary
TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents
Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
TL;DR
Existing benchmarks test whether agents reproduce hidden target studies, not whether they determine what claims data support. TruthInsightBench evaluates discovery using blind tasks and artifact-grounded scoring; four agents scored 58.4–60.3 of 100 without reliable pairwise separation, revealing deficits in scientific judgment.
Problem
Existing benchmarks center on reproducing hidden target studies, leaving no principled way to assess what autonomous research claims actually imply about scientific ability.
Method
TruthInsightBench uses 40 blind tasks across 10 domains, exposing only neutral objectives and frozen data while scoring agents’ claims across six dimensions and 29 artifact-grounded items.
Results
Four coding agents scored 58.4–60.3 of 100 with no statistically reliable pairwise separation, while weaknesses concentrated in controls, robustness, falsifiability, and cross-dataset generalization.
Takeaways & Limitations
The benchmark makes the gap between competent analysis execution and genuine discovery measurable, identifying scientific judgment rather than coding as the current bottleneck.
Takeaways & Limitations
The study uses one frozen mid-capability base model and a single run per system–task, so it does not establish exact pairwise ordering or whether the gaps persist with frontier models.
Abstract
from arXiv · showhide
Autonomous coding agents are increasingly proposed as AI-scientist systems that conduct analyses and write research reports, but executing a prescribed analysis is not the same as making a discovery. Existing benchmarks are configured for reproduction: tasks, data, and rubrics are built around a hidden target study, and recovery of its result is rewarded. We present TruthInsightBench, a benchmark configured for discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 scientific domains, expose only a neutral scientific objective and frozen data; source conclusions, expected values, and analysis paths are withheld, leaving the agent to determine what claim the data support. A fixed LLM-based judge scores the evidentiary maturity of an agent's own claims along six dimensions, operationalized as 29 artifact-grounded items, with automated, deterministic aggregation and no per-instance human grading, so evaluation can be repeated automatically as agents evolve. On one frozen base model, four coding agents form a narrow plateau (58.4-60.3 of 100) with no statistically reliable pairwise separation: they execute and document analyses competently, with comparatively strong evidence auditability and novelty, but largely lack the discriminating acts that establish a trustworthy claim (controls, robustness, falsifiability, and cross-dataset generalization). The bottleneck is scientific judgment rather than coding, and genuine discovery remains out of reach. TruthInsightBench makes this gap a measurable target; data and scoring code are at https://github.com/TruthInsight-stack/TruthInsightBench.
1 Introduction
Existing evaluations often measure execution of prescribed analyses, whereas open-ended discovery requires agents to form, test, and delimit claims from raw data. TruthInsightBench changes the benchmark configuration to evaluate that capability, revealing competent execution but weak scientific discrimination.
- Motivation: Existing benchmarks commonly conflate executing a defined analysis with open-ended inquiry, which requires forming and testing claims without a known answer.Open-ended inquiry also involves distinguishing genuine effects from artifacts, testing controls and robustness, and revising conclusions when counter-evidence appears.
- Benchmark design: TruthInsightBench changes the task setting, evaluation object, and scoring from reproduction toward discovery while retaining automatic evaluation.The benchmark is presented as complementary to, rather than a replacement for, reproduction benchmarks.
- Benchmark design: 40 blind tasks across 10 scientific domains provide neutral research goals and frozen raw data while hiding source conclusions, expected values, and analysis paths.The benchmark evaluates the evidentiary maturity of agents’ own claims rather than similarity to a reference paper.
- Baseline findings: 58.4–60.3 of 100: all four agents form a narrow, statistically indistinguishable plateau and remain below the standard for credible discovery.They produce approximately 3.7–4.3 claims per task, but their two highest-quality discoveries reach only approximately 44 of 80 evidence-quality points.
- Baseline findings: The shortfall is concentrated in control testing, robustness, falsifiability, and cross-dataset generalization rather than execution-oriented evidence auditability and novelty.The preliminary same-model study localizes the bottleneck in scientific judgment rather than coding; multi-seed and multi-model studies remain future work.
- Evaluation framework: A six-dimension, 29-item artifact-grounded framework scores claims without per-instance human grading, enabling repeated automated feedback for agent improvement.The framework treats null results, counterexamples, method artifacts, and scope boundaries as valid endpoints when supported by appropriate evidence.
2 Related Work
Scientific-agent evaluations range from knowledge tests and reproduction benchmarks to research workflows and interactive environments, but many still organize tasks around known targets or abstract away real evidentiary standards. TruthInsightBench addresses this gap with an artifact-grounded instrument for evaluating agents’ own claims on real scientific data.
- Benchmark landscape: Scientific-agent benchmarks span static knowledge and reasoning, research-process and reproduction tasks, and autonomous systems acting in interactive environments.These lines of work increasingly approach end-to-end research but differ in how much of the research process they engage.
- Static knowledge and reasoning: Knowledge benchmarks test stored knowledge and closed-book reasoning but do not require selecting questions, analyzing raw data, or forming claims.Examples include SciQ, GPQA, MMLU-Pro, Humanity’s Last Exam, SciBench, ATLAS, and domain-specific suites.
- Research-process and reproduction benchmarks: Research-process and reproduction benchmarks evaluate coding, experimentation, or recovery of target papers, yet their tasks remain organized around known targets or hypotheses.The passage contrasts ScienceAgentBench and DiscoveryBench with the fully target-free orientation sought here.
- Autonomous research systems: Open-ended research benchmarks increasingly cover analysis decisions, exploratory questions, rediscovery, research workflows, and GUI/CLI interaction, but grading still does not fully target self-formed evidentiary claims.The cited examples include BLADE, HeurekaBench, FIRE-Bench, AstaBench, and ScienceBoard.
- Interactive environments: Interactive science environments let agents act, observe, hypothesize, and experiment, but simulated worlds abstract away real datasets and real evidentiary standards.ScienceWorld and DiscoveryWorld illustrate this environment-oriented line of work.
- Position of TruthInsightBench: TruthInsightBench supplies an artifact-grounded instrument that evaluates agents on real scientific data and can be rerun without per-instance human grading.Its design responds to the need for comparable evaluation as autonomous research systems proliferate.
3 TruthInsightBench
TruthInsightBench operationalizes scientific discovery as evidence-grounded, testable claims rather than recovery of hidden reference findings. It uses blind, heterogeneous tasks and an automated six-dimension scoring framework to evaluate whether agents support, delimit, and test their discoveries.
- Scientific discovery is defined as a data-grounded claim that can be independently tested and adds judgment about relationships, mechanisms, quantitative characteristics, or applicability boundaries.
- Discovery quality depends on auditable evidence, stability under reasonable changes, control testing, cross-dataset support, novelty, and falsifiability rather than reference matching.
- Each task exposes a neutral goal, frozen data, metadata, and literature cutoff while hiding the source study, reference findings, verification assets, and scoring rules.
- 40 tasks derived from peer-reviewed studies span 10 scientific domains and heterogeneous data modalities, requiring methods suited to different scientific objects.
- Task admission requires interpretable data, verified recoverability of at least two results, and repeats, controls, condition variation, or alternative analyses for testing explanations and scope.Construction includes two result re-derivations and six perturbation analyses per task, totaling 320 verifications across 40 tasks.
- The rubric assigns 100 points across six dimensions, with evidence auditability weighted most heavily, and scores only analyses and results that were actually completed and verifiable.A frozen LLM judge applies artifact-grounded items, while aggregation is deterministic and requires no per-instance human grading.
- TruthInsightBench separates discovery credibility from reference convergence: novel supported claims can score highly, while correct but unsupported claims do not receive credible-discovery credit.This design prevents answer-key logic and keeps the ranking focused on how well claims are grounded in evidence.
4 Experiments
Across 40 tasks, four coding agents produced statistically indistinguishable overall scores, with competent execution and traceability but weak evidence depth and scientific discrimination. The benchmark separated systems most clearly on harder tasks, where claims required controls, robustness, falsifiability, and independent validation.
- Overall results: Four agents clustered between 58.4 and 60.3 of 100, with no statistically reliable pairwise separation under a single frozen run.The 40-task means had overlapping confidence intervals, and paired tests found no significant difference between any pair.
- Overall results: 38–40 of 75 current-evidence points were attained on average, leaving leading discoveries incomplete across several audited dimensions.This corresponds to 51–53% of the evidence-auditability, robustness, and control-testing component.
- Domain and task-level structure: Overall means concealed substantial task-level variation: individual-task differences reached 25.5 points, and harder tasks tended to show wider cross-system divergence.The top-scoring system rotated across domains, while its margin over the runner-up was at or below 1.1 points in seven of ten domains.
- Breadth versus evidence depth: Comparable discovery breadth masked shallow evidence: top-two discoveries reached only 43.3–44.6 of 80 evidence-quality points.Reports contained approximately 3.65–4.30 de-duplicated discoveries per task, with repetition rates of 9.9–15.1%.
- Six-dimension diagnostic: Evidence auditability and novelty were strong, at 78–81% and 83–87%, respectively, while control testing and robustness remained weak at 11–16% and 9–14%.The strong dimensions support execution and traceability, whereas the weak dimensions require discriminating tests and uncertainty assessment.
- Six-dimension diagnostic: Falsifiability was only 29–33%, and cross-dataset generalization was near zero because independent validation was rarely performed and some tasks lacked suitable second data sources.The authors distinguish benchmark-side feasibility limits, agent-side discipline gaps, and the single-phase protocol’s temporal limits.
5 Diagnostic Case Studies
TruthInsightBench uses diagnostic tasks where superficially compelling patterns require discriminating tests rather than naive execution. These cases expose whether systems can test artifacts, generalize findings, and delimit extrapolation.
- Chemistry_07: Chemistry_07 contains extreme photoionization-delay spikes that could reflect either a physical resonance or high-order interpolation artifacts.Resolving the ambiguity requires comparing interpolation methods and testing stability under node perturbation.
- Physics_05: Physics_05 tests whether an apparent critical slowdown transfers from synthetic models to real networks.Its scores range from 45.6–56.8, among the benchmark’s lowest-scoring regions.
- Physics_05: The Physics_05 task measures whether systems can delimit extrapolation instead of reporting a visually salient synthetic trend as general.The relevant test is whether the trend survives on real data.
- Interpretation: These are ordinary scientific situations requiring discriminating tests, not adversarial trick items.Their inclusion lets the benchmark measure discovery rather than execution.
6 Discussion
TruthInsightBench measures retrospective, blind, data-driven discovery, while explicitly bounding what its protocol can establish. Its automated feedback is intended to support scalable research-improvement loops, but current tasks and experiments leave important sources of uncertainty unresolved.
- Scope: TruthInsightBench evaluates claims re-formed and re-tested from frozen data without revealing source conclusions or the standard analysis path.A high score does not constitute a real-world first discovery; ultimate confirmation requires new data, experiments, and independent teams.
- Limitations: Protocol-level blindness cannot establish that the base model never encountered related papers during pretraining.The fixed literature cutoff and novelty rule mitigate but cannot fully eliminate contamination concerns.
- Limitations: Static single-phase tasks under-measure cross-dataset generalization because they provide limited second-phase data.Two-phase tasks releasing validation data after pre-specified discovery criteria are recorded are identified as future work.
- Limitations: A single frozen run per system–task does not estimate selection variance in open-ended discovery.Scoring variance is constrained by frozen judging and deterministic aggregation, whereas planned repeated runs address selection variance.
- Research-improvement loops: Automated artifact-grounded scores can provide dimension-level feedback and item diagnostics without human grading for each instance.The paper recommends more stable aggregates, such as means over repeated runs or pooled dimension-level diagnostics, rather than single-run totals.
- Limitations: The release uses one frozen mid-capability base model, computational dry-lab tasks, and a single frozen LLM judge.Therefore, exact pairwise ordering, transfer to wet-lab or less curated data, and equivalence to expert judgment remain unestablished.
7 Conclusion
TruthInsightBench reconfigures evaluation around open-ended discovery rather than recovery of a hidden target result. Its preliminary study finds a narrow performance plateau, with the main deficit in the discriminating acts needed for credible discovery.
- Benchmark conclusion: TruthInsightBench gives agents a neutral objective and frozen data, hides source conclusions and evaluation assets, and scores claims across six dimensions and 29 artifact-grounded items.Automated deterministic aggregation enables repeated evaluation without per-instance human grading.
- Benchmark conclusion: Null results, counterexamples, method artifacts, and scope boundaries receive credit alongside positive regularities.This design rewards scientific judgment rather than recovery of a privileged endpoint.
- Empirical finding: 58.4–60.3 of 100: four coding agents form a narrow plateau with no statistically reliable pairwise separation.The systems conduct and document analyses competently but remain well short of credible discovery.
B The 29 evaluation items
The evaluation rubric converts evidence quality into deterministic item scores with frozen satisfaction thresholds and weighted contributions.
- Scoring rules: Each evaluation item receives a score of 1, 0.5, or 0 for full, partial, or unsatisfied satisfaction.Partial satisfaction requires no direction-critical conflict and at least 50% but less than 100% frozen weighted coverage.
B.1 Evidence auditability (45 points)
Evidence auditability evaluates whether agents use the correct data and analytical units, execute key analyses, trace claims to results, and faithfully synthesize successes, failures, and null findings.
- 4 points assess data-source, version, scope, object, unit, and independent-observation consistency with the task.
- 6 points require key analyses to run successfully and reports to match actual results.
- 4 points require every claim and key number to map to a concrete analysis result.
- 10 points require conclusion-critical quantities to reproduce the reported conclusion from the executed analysis.
- 5 points require effect sizes, sample sizes, and supporting numbers to reproduce within tolerance.
- 8 points require executed results to support the claimed scope and relation, while the final conclusion incorporates later successes, failures, and null results.
B.2 Robustness (15 points)
Robustness evaluates uncertainty, sensitivity to reruns and analytical choices, applicability boundaries, competing explanations, confounds, and corroboration from methods with different error sources.
- 3 points require estimation of the main random, measurement, or model uncertainty with applicable methods.
- 3 points require complete randomized reruns under specified seeds, while another 3 assess retained perturbations of batches, outliers, held-out data, or time splits.
- 3 points assess whether method, parameter, or threshold variations are retained and whether the stated validity regime matches observed failures and untested scope.
- 3 points require comparison with simple or null models, and 3 require controls that check the method’s false-positive rate.
- 4 points require testing a principal competing mechanism with discriminating predictions, while 2 assess grouping, batch, sampling, measurement, preprocessing, or fitting-scale confounds.
- 3 points require a second analysis with different dominant error sources to corroborate the conclusion.
B.4 Cross-dataset generalization (10 points)
Cross-dataset generalization evaluates independent validation data, preregistered analyses and criteria, separation between discovery and validation, novelty beyond prior literature, and actionable updating tests.
- 2 points require validation data to remain independent of hypothesis, feature, threshold, and model selection.
- 3 points require a preregistered analysis on new data, time, sources, or objects, with 3 more for criteria fixed before viewing results.
- 2 points require that direction-critical observations or answer-derived labels do not feed back into discovery.
- 4 points require the core claim to come from task data, while 6 assess whether it adds an object, relation, result, or boundary beyond frozen pre-T0 literature.
- 1 point each assesses a new discriminating measurement, an observation that distinguishes competing explanations, and a stated condition for maintaining, strengthening, narrowing, or overturning the conclusion.
- 1 point each assesses whether required implementation resources are available and whether the conclusion-update rule is fixed, runnable, and covers the outcome space.
C Complete per-task score matrix (40 tasks × four coding agents)
Table 7 reports the complete score matrix for 40 frozen tasks evaluated across four coding systems, with bold marking the maximum score on each task.
- 160 frozen task scores cover 40 tasks evaluated by four systems.
- Bold formatting marks the per-task maximum among the four systems.
- The columns identify Claude Code, OpenScience, Codex CLI, and DeepSeek Harness.