Source-linked AI summary

ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment

Qiuyu Tian, Haojie Yin, Yingce Xia, Youyong Kong, Zequn Liu

arXiv:2606.00644v2cs.AI

TL;DR

LLM agents need to make research decisions before future evidence exists, but existing benchmarks do not test this historical, open-ended judgement. ForeSci builds a temporally controlled benchmark with hidden post-cutoff validation and evaluates multiple agent designs. Agentic workflows often improve traceability and some evidence-grounded metrics, but performance varies by decision family and agents can cite relevant evidence while choosing the wrong research object.

  • Problem

    Existing benchmarks do not test whether agents can make open-ended research decisions from evidence available at a specific historical moment.

  • Method

    ForeSci uses 500 cutoff-controlled tasks across four AI domains and four decision families, with offline historical knowledge bases, hidden future targets, and pre-cutoff backbones.

  • Results

    Agentic workflows often improve traceability and evidence-grounded metrics, but no method is uniformly best across backbones, task families, and evaluation signals.

  • Takeaways & Limitations

    Separating factual support, future-target alignment, traceability, and persuasiveness makes evidence-decision decoupling and distinct failure modes measurable.

  • Takeaways & Limitations

    ForeSci’s conclusions are bounded to four fast-moving AI areas and four decision families and do not establish universal rankings across scientific domains, languages, or time horizons.

Abstract

from arXiv · show

AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce ForeSci, a temporally controlled benchmark for evaluating whether LLM agents can make such forward-looking research judgements from historical evidence. ForeSci contains 500 tasks across four fast-moving AI domains and four decision families. Each task is paired with a cutoff-aligned offline knowledge base; post-cutoff papers are hidden during generation and used only for validation. To avoid random future-event prediction, tasks are derived from pre-cutoff taxonomy branches and evidence signals, and answer-generation backbones are selected to precede the task cutoffs. We evaluate native LLMs, Hybrid RAG, and three research-agent adaptations across four backbones. Results show that explicit evidence organization improves traceability and factual support, but gains depend strongly on the decision family. Diagnostics reveal a recurring evidence-decision decoupling: agents may cite relevant evidence while forecasting the wrong research object. ForeSci turns forward-looking AI research judgement into a controlled benchmark for evaluating research agents as decision-making systems.

1 Introduction

ForeSci addresses whether LLM agents can make defensible, evidence-grounded research decisions about an unwritten future using only historically available evidence. It introduces a temporally controlled benchmark that evaluates these judgements across diverse decision families and signals.

  • Motivation: Existing benchmarks do not ask agents to make open-ended research decisions from evidence available at a specific historical moment.Prior work emphasizes paper questions, synthesis, tool use, workflows, or future-paper components rather than choosing bottlenecks, agendas, or venues under a historical cutoff.
  • Benchmark: ForeSci contains 500 tasks across four fast-moving AI domains and four decision families, paired with cutoff-aligned offline knowledge bases.The decision families include direction forecasting, bottleneck–opportunity discovery, strategic research planning, and venue-aware research positioning.
  • Temporal control: Tasks are grounded in pre-cutoff taxonomy branches, evidence records, and method-evolution signals, while post-cutoff evidence remains hidden until evaluation.Answer-generation backbones are also selected to precede the task cutoffs, limiting hindsight and future leakage.
  • Evaluation: The benchmark evaluates factual support, future-target alignment, evidence traceability, and reviewer persuasiveness across native, retrieval-based, and agentic systems.These signals separate whether an answer is factually supported, aligned with future targets, grounded in visible evidence, and persuasive as a research judgement.
  • Findings: Agent-style methods improve traceability and factuality, but the strongest method depends on the decision family.Diagnostics identify evidence-decision decoupling, where agents cite relevant evidence while forecasting the wrong research object, causal role, or intervention.

2 Related Work

ForeSci extends research-agent and temporal-forecasting benchmarks toward open-ended scientific decisions made from historically available evidence. Its distinctive setting combines strict temporal control with scholarly evidence and hidden post-cutoff supervision.

  • Research-agent benchmarks: Existing autonomous-research benchmarks measure retrieval, tool use, synthesis, scientific reasoning, or workflow execution rather than forward-looking research decisions.ForeSci targets the decision layer beyond paper search or summary.
  • Temporal evaluation: Temporal integrity prevents systems from benefiting from hindsight, leakage, or later-stabilized terminology during foresight evaluation.Prior time-sliced benchmarks motivate strict cutoffs for future-oriented reasoning.
  • ForeSci’s distinction: ForeSci focuses on future-oriented scientific decision-making in fast-moving AI subfields rather than general-domain event prediction.The answer must transform cutoff-visible scholarly evidence into an open-ended research judgement.
  • ForeSci’s distinction: The benchmark pairs a cutoff-aligned offline knowledge base with hidden post-cutoff supervision for research-agent outputs.This extends temporal control from event prediction to decisions such as trajectories, bottlenecks, plans, and venue positioning.

3 The ForeSci Framework

ForeSci formulates forward-looking research judgement as a cutoff-conditioned decision problem: systems use only historical evidence, while future literature supplies hidden validation targets. The framework constructs these tasks from filtered corpora, temporal taxonomies, evidence assets, and four decision families.

  • Problem formulation: A ForeSci instance is defined by a question, cutoff date, cutoff-aligned knowledge base, and required task family.Systems return an answer using only the provided historical knowledge base, while post-cutoff targets are reserved for evaluation.
  • Problem formulation: ForeSci prevents leakage by using pre-cutoff backbones, disabling web search, and restricting external support to the cutoff-aligned knowledge base.The withheld post-cutoff targets are accessible only during evaluation.
  • Task families: The framework covers Direction Forecasting, Bottleneck–Opportunity Discovery, Strategic Research Planning, and Venue-Conditioned Positioning.These families respectively predict trajectories, identify bottleneck-enabled opportunities, rank plans, and position projects for venue communities.
  • Data collection and filtering: The corpus spans LLM agents, fine-tuning and post-training, RAG and retrieval structuring, and visual generative modeling, with domain and benchmark-core filtering before chronological truncation.Core papers provide representative contributions and future-facing signals, while support papers preserve relevant context.
  • Taxonomy and evidence assets: Temporal taxonomy induction represents evolving research landscapes as nodes and method-evolution edges grounded in cutoff-visible literature.Human experts check taxonomy support strength and temporal validity to preserve temporal causality.
  • Taxonomy and evidence assets: Node evidence records connect research subdirections to representative and supporting papers describing pre-cutoff problems, methods, evaluations, limitations, and contributions.Derived signals include candidate directions, method development, bottlenecks, feasibility and risk, and venue-community metadata.
  • Task construction: Human–LLM collaboration turns taxonomy-derived evidence into benchmark instances, with experts checking cutoff validity, leakage risk, clarity, and grounding before approval.The process is intended to produce expert-validated foresight challenges rather than merely reproduce taxonomy structure.

4 Evaluation

ForeSci evaluates research judgements through four complementary metrics that compare answers with hidden future targets and visible pre-cutoff evidence. It tests native, retrieval-based, and offline-adapted agentic systems across multiple backbones.

  • Metrics: Four metrics assess factuality, future-target alignment, evidence traceability, and reviewer persuasiveness.Together they evaluate future facts, task-specific decisions, visible-evidence grounding, and the quality of the presented judgement.
  • Metrics: Prediction Factuality reports claim-level F1 between answer atomic claims and task-relevant hidden validation claims.The metric follows an atomic-fact approach and derives its claim bank from hidden future targets.
  • Metrics: Future-Target Alignment compares prediction claims with hidden claim banks or computes deterministic ranking alignment for ordered decisions.Direction Forecasting and Bottleneck–Opportunity Discovery use similarity, while planning and venue positioning use preferred rankings.
  • Metrics: Evidence Traceability scores relevant use of pre-cutoff evidence, support for the stated decision, and avoidance of unsupported reasoning jumps on a normalized [0, 1] rubric.Reviewer Persuasiveness separately scores decision quality, mechanistic and comparative reasoning, clarity, and risk awareness.
  • Systems: The evaluation compares a native LLM, Hybrid RAG, and three offline-adapted agentic systems using constrained retrieval, tools, memory, and task-specific answer schemas.All systems operate within the offline knowledge base across four LLM backbones.

5 Results

ForeSci results show that agentic methods improve evidence-grounded metrics inconsistently, with performance varying by backbone and decision family. Error analysis identifies evidence-decision drift as a key failure mode: answers may be traceable yet target the wrong research decision.

  • Overall results: Agent-style methods generally improve Fact, FTA, and Trace, but gains do not consistently improve Reviewer Persuasiveness.The strongest agent is competitive with or better than Native LLM and Hybrid RAG on Fact and FTA, while all three agents improve Trace over Hybrid RAG.
  • Overall results: No agent is uniformly strongest across metrics, backbones, and task families, and some settings show no clear advantage over the native backbone.Method rankings vary by task family, so additional retrieval and tool use do not automatically improve foresight performance.
  • Family-dependent failures: Strategic Planning has the highest low-score rates on Fact and FTA, showing that failures are strongly family-dependent.Matching both the ranked decision and its supporting facts is particularly difficult in this family.
  • Evidence-to-decision drift: Causal-role drift lowers Fact by 1.13 standard deviations, while scope/granularity and intervention-mode drift lower FTA by 1.22 and 1.12 standard deviations, respectively.Persuasiveness also declines under severe drift, whereas Trace is more weakly coupled and varies by drift type.
  • High traceability but high drift: High-Trace answers with low FTA show higher drift severity across all four bias types, demonstrating that local evidence support can coexist with the wrong decision object.One example has Trace 0.920 but Prediction Factuality 0.200 and FTA 0.355 because its venue framing is misaligned with the reference target.
  • Prospective forecasting: The prospective pipeline generates transparent forecast artifacts from cutoff-visible literature, although the showcased future outcomes are not scored at writing time.A proof-of-concept package contains 12 prediction-only questions for a 2026-05-16 to 2026-08-15 forecast window.

6 Conclusion

ForeSci evaluates forward-looking AI research judgement by pairing cutoff-controlled evidence with hidden future targets. Its results show that traceability gains do not guarantee correct decisions, making evidence-decision decoupling measurable across research-agent systems.

  • Conclusion: ForeSci evaluates 500 cutoff-controlled tasks across four decision families using offline knowledge bases and hidden post-cutoff validation targets.The benchmark separates factual support, future-target alignment, traceability, and reviewer-style persuasiveness.
  • Conclusion: Agentic workflows often improve traceability and some evidence-grounded metrics, but no method is uniformly best across backbones, task families, and evaluation signals.The benchmark therefore evaluates research agents as decision-making systems rather than literature interfaces alone.
  • Conclusion: Evidence-decision decoupling occurs when agents cite relevant evidence yet choose the wrong research object, causal role, intervention mode, or time horizon.ForeSci makes these distinct failures measurable through complementary evaluation signals.

Limitations

ForeSci’s conclusions are bounded by its controlled benchmark setting and by evaluation choices that approximate, rather than directly measure, scientific value.

  • The findings apply to four fast-moving AI areas, four decision families, and the benchmark’s specified time horizons rather than all scientific settings.The paper cautions against treating results as a universal ranking across domains, languages, or horizons.
  • The benchmark emphasizes paper-visible signals and cannot fully capture tacit community knowledge, unpublished work, private reviewer expectations, or downstream adoption.
  • Evaluation depends on hidden post-cutoff targets and LLM-as-judge metrics, whose rubric-based persuasiveness scores remain approximations rather than direct measurements of scientific value.Family-conditioned rubrics, repeated judging, cross-backbone comparisons, and diagnostic audits are used to reduce over-interpretation.

Ethical Considerations

ForeSci is intended for diagnostic evaluation of research-assistant systems, not autonomous decisions about venues, peer review, hiring, funding, or research priorities.

  • The released benchmark should not automate peer review, venue selection, hiring, funding, or research-prioritization decisions without human oversight.The paper warns that poorly calibrated forecasts could encourage premature convergence or overstate evidential support.

Code and Data Availability

ForeSci describes a cutoff-aligned, auditable construction pipeline that combines automated drafting with expert verification across public scholarly artifacts.

  • The corpus is built by broad domain harvesting, metadata normalization, deduplication, relevance screening, and a stricter benchmark-core screen.Less-central relevant papers are retained as support, while noisy or borderline cases are excluded or retained only for audit.
  • Temporal slicing preserves recent research changes, while adaptive taxonomy expansion exposes emerging subdirections across tasks, methods, datasets, evaluation, and applications.Induced nodes must be grounded in evidence records before supporting benchmark construction.
  • Candidate directions require pre-cutoff support, a clear and evaluable decision, and separability from hidden future validation evidence.Ambiguous, duplicated, overly broad, overly narrow, or weakly grounded candidates are revised or removed.
  • Method-development and bottleneck signals encode trajectory relations, recurring limitations, evaluation gaps, reliability or safety concerns, and technical risks.These signals support comparisons of directions and mechanism-level change using full-text evidence where appropriate.
  • Venue profiles summarize contribution styles, maturity expectations, reviewer risks, evidence requirements, and compatible venue families for venue-cycle judgments.
  • LLMs draft construction records, after which human experts verify support, temporal validity, specificity, leakage risk, and duplication before release.The same audit process is applied before task release.

B.7 Task Curation and Artifact Separation

Task curation combines historical decision premises with post-cutoff validation separation, while supporting artifacts document corpus evolution and publication-calendar context.

  • Task Curation: Tasks are revised when their decision is underspecified, validation evidence mismatches the requested judgment, or multiple items express the same decision.
  • Artifact Separation: Table A1 reports task counts by domain and horizon type, with three-month and six-month values drawn from task metadata and venue tasks using venue-cycle timing.
  • Artifact Separation: Figure A1 counts unique normalized papers per domain-month from January 2023 through the March 2026 cutoff across all four domains.
  • Artifact Separation: Each released task contains a public question, cutoff, forecast window, family instructions, answer requirements, and pre-cutoff support packet, while future targets remain evaluation-only.
  • Artifact Separation: Publication-volume curves are cutoff-dependent background variables rather than direct measures of foresight difficulty, with recurring patterns broadly matching major venue cycles.
  • Artifact Separation: Corpus evolution shows differing domain scale and breadth, increasingly specialized taxonomy descendants, and recurring bottlenecks becoming concrete mechanisms over time.

C.1 Metric Calibration Details

ForeSci calibrates a family-conditioned metric stack for evaluating factual support, future-target alignment, traceability, and persuasiveness in open-ended research decisions.

  • Metric suite: ForeSci reports Prediction Factuality, Future-Target Alignment, Evidence Traceability Score, and Reviewer Persuasiveness as complementary evaluation signals.The protocol uses claim-level support, hidden future targets, evidence linkage, and rubric-based research judgment.
  • Prediction Factuality: Prediction Factuality is claim-level F1 over answer-claim support and coverage of hidden task-relevant claim units.Answer claims receive support labels, while hidden claims receive coverage labels; precision and recall are retained as intermediate quantities.
  • Future-Target Alignment: Future-Target Alignment uses reference-guided similarity for Direction and Bottleneck tasks, but deterministic ranking alignment for Planning and Venue tasks.Ranking-aware FTA combines top-item, position, and pairwise-order checks into precision and recall terms, then reports their F1.
  • Evidence Traceability: Evidence Traceability combines external evidence linkage, support specificity, and answer-internal trace using weights 0.50, 0.25, and 0.25.Traceability aggregates exclude methods without an attached external support artifact.
  • Reviewer Persuasiveness: Reviewer Persuasiveness scores research judgments with family-specific criteria plus generic clarity, mechanistic reasoning, comparative reasoning, and uncertainty or risk awareness.The rubric evaluates the task, answer, visible evidence, and hidden future target through a virtual reviewer.
  • Validation: Human validation supports the reliability of both Reviewer Persuasiveness and Prediction Factuality as decision-critical evaluation metrics.Validation covers 400 long-form answers for persuasiveness and 240 answers for claim extraction quality.

E.2 Scalar Metric Stability

The scalar metric stack is evaluated for run-to-run stability, with Future-Target Alignment the most stable metric and Evidence Traceability the most variable. Close within-family rankings for moderately variable metrics should be interpreted cautiously.

  • Stability protocol: Five repeated DeepSeek-V4 evaluations measure metric variability on fixed candidate answers across 500 expanded task–method rows.The study combines one formal evaluation with four independent replicate runs and reports standard deviations, ranges, and row-level variation.
  • Metric stability: Future-Target Alignment is the most stable metric, while Prediction Factuality and Reviewer Persuasiveness show moderate variance.Planning and Venue FTA are deterministic ranking-aware scores; Bottleneck and Direction FTA use reference-guided embedding similarity.
  • Metric stability: Evidence Traceability has the largest variance, although its variance remains moderate and acceptable for rubric-based evidence assessment.Native LLM rows are excluded from traceability analyses because they lack external support artifacts.
  • Interpretation: Close within-family method rankings should be treated as small-margin differences for Prediction Factuality and Reviewer Persuasiveness.The stability analysis supports family-level comparisons more strongly than fine-grained rankings within a family.

F.1 Detailed Error Analysis

The error analysis finds that failures concentrate in family-specific decision mismatches rather than a single global weakness. It also shows that factuality, future-target alignment, traceability, and persuasiveness capture related but nonredundant aspects of answer quality.

  • Failure concentration: Strategic Planning is hardest, with low Prediction Factuality of 0.315 and low Future-Target Alignment of 0.512.A wrong top priority can invalidate an otherwise plausible plan.
  • Metric conflicts: The refreshed Fact–FTA audit finds 131 large disagreements, including 112 rows with an absolute gap of at least 0.5 and 19 high-FTA/low-Fact cases.The high-FTA/low-Fact cases are mostly in Direction Forecasting, where answers approach the future target space but miss the exact mechanism, trajectory label, or scope.
  • Planning diagnostics: Top-priority errors range from 0.468 for ResearchAgent-style answers to 0.586 for Native LLM answers.These answers may justify individual candidates while failing to support the full global ordering.
  • Venue diagnostics: Venue low-trace cases are often topical rather than venue-specific evidence chains, with Hybrid RAG showing a 0.312 low-trace rate for this pattern.This indicates that relevant evidence can still fail to support venue-conditioned positioning.
  • Metric coupling: Fact–FTA correlation is 0.816 for Strategic Planning and 0.743 for Venue Positioning, versus 0.292 for Bottleneck–Opportunity and 0.368 for Direction Forecasting.FTA versus Persuasiveness is only 0.274 for Venue Positioning, while method and backbone effects are smaller than family effects.
  • Scope boundary: ForeSci's prospective forecasting example has no post-hoc ground truth or judge/evaluation scores, so it illustrates the construction pipeline rather than providing a scored benchmark result.The benchmark's hidden future supervision remains the basis for formal evaluation.
Loading 2606.00644v2…