Source-linked AI summary

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

Jiayu Wang, Yifei Ming, Riya Dulepet, Qinglin Chen, Austin Xu, Zixuan Ke, Frederic Sala, Aws Albarghouthi, Caiming Xiong, Shafiq Joty

arXiv:2510.14240v5cs.AI

TL;DR

Deep research requires benchmarks that reflect dynamic, user-centered, multi-faceted information needs and evaluations that reliably assess citation-grounded long-form reports. The paper introduces LiveResearchBench and DeepEval, then evaluates 17 systems, finding that agents can gather and organize information but still struggle with citation reliability and analytical depth.

  • Problem

    Existing deep-research benchmarks often lack realistic, dynamic, unambiguous, multi-faceted tasks, while long-form reports remain difficult to evaluate reliably across multiple quality dimensions.

  • Method

    The paper builds LiveResearchBench with 100 expert-curated tasks and checklists, and introduces DeepEval for content- and report-level evaluation of citation-grounded reports.

  • Results

    Most evaluated systems still struggle with citation reliability and analytical depth despite being able to gather and organize information.

  • Takeaways & Limitations

    LiveResearchBench and DeepEval provide a rigorous foundation for evaluating deep research and identifying improvements needed for more insightful reports.

  • Takeaways & Limitations

    The benchmark comparison includes prior tasks with limited support for long-form answers and some search-intensive tasks with low reasoning load.

Abstract

from arXiv · show

Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for agentic systems. To rigorously evaluate this ability, four principles are essential: tasks should be (1) user-centric, reflecting realistic information needs, (2) dynamic, requiring up-to-date information beyond parametric knowledge, (3) unambiguous, ensuring consistent interpretation across users, and (4) multi-faceted and search-intensive, requiring search over numerous web sources and in-depth analysis. Existing benchmarks fall short of these principles, often focusing on narrow domains or posing ambiguous questions that hinder fair comparison. Guided by these principles, we introduce LiveResearchBench, a benchmark of 100 expert-curated tasks spanning daily life, enterprise, and academia, each requiring extensive, dynamic, real-time web search and synthesis. Built with over 1,500 hours of human labor, LiveResearchBench provides a rigorous basis for systematic evaluation. To evaluate citation-grounded long-form reports, we introduce DeepEval, a comprehensive suite covering both content- and report-level quality, including coverage, presentation, citation accuracy and association, consistency and depth of analysis. DeepEval integrates four complementary evaluation protocols, each designed to ensure stable assessment and high agreement with human judgments. Using LiveResearchBench and DeepEval, we conduct a comprehensive evaluation of 17 frontier deep research systems, including single-agent web search, single-agent deep research, and multi-agent systems. Our analysis reveals current strengths, recurring failure modes, and key system components needed to advance reliable, insightful deep research. Our code is available at: https://github.com/SalesforceAIResearch/LiveResearchBench.

1 INTRODUCTION

LiveResearchBench addresses shortcomings in deep-research evaluation by defining realistic, dynamic, unambiguous, multi-faceted tasks and pairing them with DeepEval for fine-grained report assessment.

  • Benchmark motivation and contributions: Existing benchmarks often use static, domain-specific, coarse-grained, or ambiguous tasks that omit intended audience, output format, or scope.These limitations make consistent interpretation and fair evaluation of deep-research systems difficult.
  • Benchmark motivation and contributions: LiveResearchBench introduces 100 expert-curated tasks across diverse domains, designed around user needs, dynamic information, unambiguous specifications, and intensive search.The benchmark was constructed with over 1,500 hours of human labor and includes detailed checklists.
  • Benchmark motivation and contributions: Evaluating open-ended reports is difficult because real-time queries lack fixed ground truth and require simultaneous assessment of coverage, reasoning, evidence use, and presentation.Human evaluation is costly and hard to scale, while naive LLM judges can produce inconsistent results.
  • Benchmark motivation and contributions: DeepEval evaluates long-form reports across coverage, analysis depth, citation association, factual accuracy, presentation, organization, and logical and factual consistency.Its dimensions use tailored checklist-based, pointwise, pairwise, or rubric-tree protocols for scalable and stable assessment.
  • Benchmark motivation and contributions: The study evaluates 17 open-source and proprietary single- and multi-agent systems and reports systematic vulnerabilities, with most models unable to write insightful reports.This evaluation is positioned as a comprehensive assessment of current deep-research capabilities.

2 RELATED WORK

Prior deep-research benchmarks address pieces of information seeking or long-form evaluation but commonly remain narrow, static, short-form, closed-ended, or single-dimensional.

  • Agentic system context: Single-agent systems place tool-use decisions in one model, whereas deep-research systems are evaluated across broader architectural categories.The paper contrasts single-agent approaches with multi-agent systems in its system taxonomy.
  • Deep research benchmarks: Existing benchmarks commonly target domain-specific, short-form or closed-ended, and static research tasks rather than broad, evolving deep-research outputs.The related benchmarks include specialized related-work generation, short open-ended questions, and closed-ended information-seeking tasks.
  • Long-form answer evaluation: Prior long-form evaluation studies often isolate one dimension, such as sub-question coverage, claim verification, structure, or overall quality.DeepEval is presented as addressing this fragmentation with broader coverage of report-quality dimensions.

3 LIVERESEARCHBENCH

LiveResearchBench is a user-centered benchmark built from realistic, dynamic, unambiguous, and search-intensive research needs. Its expert-curated queries and verified checklists support consistent evaluation across diverse domains and task categories.

  • 3.1 TASK DESIGN PRINCIPLES: Its tasks are designed to reflect realistic user needs while specifying scope, audience, and format to reduce ambiguity and improve interpretability.The benchmark also emphasizes broad, real-time information gathering rather than static, narrowly specified questions.
  • 3.2 BENCHMARK OVERVIEW: LiveResearchBench contains 100 expert-curated questions spanning seven domains and ten task categories, each paired with detailed checklists.The benchmark covers areas including science, business, healthcare, market analysis, technical support, and decision support.
  • 3.3 BENCHMARK CONSTRUCTION: The benchmark’s six-stage construction pipeline combines user research, expert drafting, model-generated clarification questions, human refinement, and GPT-5 checklist generation.User interviews and surveys determine realistic task and domain distributions before experts draft and refine the queries.
  • 3.3 BENCHMARK CONSTRUCTION: GPT-5-generated checklists decompose each query into unit questions that test whether reports address essential requirements, enabling consistent coverage evaluation.For example, checklists can separately test market years, geography, and requested metrics within one task.
  • 3.4 DATA VERIFICATION: Independent expert assessment, two quality-control rounds, and final cross-checking verify the questions and checklists before dataset release.Annotators label query and checklist-item quality under detailed guidelines, while later experts resolve conflicts and finalize the dataset.

4 DEEPEVAL: A COMPREHENSIVE EVALUATION SUITE FOR DEEP RESEARCH

DeepEval evaluates long-form deep-research reports across complementary content- and report-level dimensions. It combines metric-specific protocols and multi-judge assessment because open-ended, real-time reports cannot be reliably evaluated by string matching or a single rating.

  • 4 DEEPEVAL OVERVIEW: DeepEval assesses presentation, consistency, coverage, analysis depth, citation traceability, and citation accuracy for generated research reports.These dimensions address both report-level quality and the correctness, completeness, and evidential grounding of content.
  • EVALUATION PROTOCOL: Single-rating evaluation produced below 60% human agreement and run-to-run score differences exceeding 50 points, motivating more specialized protocols.The reported instability occurred even with Gemini-2.5 Pro and GPT-5 judges.
  • AGENT-ENSEMBLE-AS-A-JUDGE: A multi-judge ensemble is adopted because pilot results found Gemini 2.5 Pro most aligned with human preferences, followed by GPT-5, while Claude 4 Sonnet was inconsistent.The ensemble is intended to mitigate inductive bias from relying on one model.
  • METRICS: Coverage uses verified checklist items as binary unit tests, while analysis depth compares reports across reasoning granularity, layered insights, critique, evidence use, and insight density.Coverage averages checklist success across items and queries; depth uses pairwise comparisons with 1–5 scores.
  • CITATION METRICS: Citation traceability checks whether factual claims link to sources, whereas citation accuracy verifies whether linked URLs are accessible and genuinely support those claims.Citation accuracy therefore complements association by checking evidential entailment rather than merely the presence of a link.

5 MAIN RESULTS AND ANALYSIS

The evaluation shows systematic trade-offs across report length, coverage, consistency, citation association, presentation, analysis depth, and citation correctness. Multi-agent systems lead average performance and presentation, but no system consistently combines breadth, depth, coherence, and grounded citation quality.

  • Multi-agent systems achieve the highest family-average score at 69.5, ahead of single-agent web systems at 62.8.Open Deep Research has the highest individual average at 73.6, followed by GPT-5 at 72.7.
  • Single-agent web models lead factual and logical consistency at 69.7, while multi-agent systems lead citation association at 61.9.Gemini 2.5 Pro achieves the highest overall consistency score at 76.5.
  • Multi-agent systems lead presentation with an average score of 77.7, but polished organization remains weakly correlated with citation association and consistency.Open Deep Research and Deerflow+ achieve presentation scores of 81.0 and 78.8, respectively.
  • Grok-4 Heavy Deep Research achieves the highest Coverage & Comprehensiveness score at 89.3, but scaling retrieval and system complexity strains memory capacity.The passage identifies system-wide memory architectures and compression strategies as the resulting open challenge.
  • Only Deerflow+ and Gemini Deep Research exceed Open Deep Research in analysis depth, while long reports often fail to synthesize insights across sources.The analysis-depth comparison uses win rate over Open Deep Research with GPT-5 as the backbone.
  • All evaluated top systems produce non-trivial citation errors, with unsupported claims driving many errors in wide information search.The evaluated systems are GPT-5, Grok-4 Deep Research, and Open Deep Research, measured using E1, E2, and E3 error counts.

6 CONCLUSION

The paper concludes that LiveResearchBench and DeepEval provide a rigorous basis for evaluating dynamic deep-research tasks and citation-grounded reports. Its evaluation framework combines complementary protocols for coverage, presentation, consistency, and related report-quality dimensions.

  • 6 CONCLUSION: LiveResearchBench and DeepEval establish a rigorous foundation for benchmarking citation-grounded deep research.The benchmark contains 100 expert-curated dynamic tasks, while DeepEval spans six report-quality dimensions.
  • Presentation & Organization: Presentation and organization are evaluated with binary checklist scoring using runtime-filled report and query fields.The presentation checklist is reused across research questions.
  • Consistency and Citation Traceability: Pointwise additive evaluation identifies concrete factual or logical inconsistencies and citation-traceability errors, with adjustable penalties for detected issues.The consistency rubric separately excludes ordinary factual accuracy and permits repeated use of the same source across claims.
  • Coverage & Comprehensiveness: Checklist-based evaluation tests whether reports fully deliver every requested component of each research task.Any missing, incorrect, or incomplete requested component yields a zero for that checklist item.

A.4 ANALYSIS DEPTH

DeepEval measures analysis depth through pairwise comparison of two reports. It scores five dimensions covering reasoning, implications, critique, evidence use, and substantive density.

  • A.4 ANALYSIS DEPTH: Analysis depth is evaluated by comparing Report A and Report B side by side.The evaluator decides which report demonstrates greater depth of analysis.
  • A.4 ANALYSIS DEPTH: Five dimensions assess mechanisms, implications, critique, evidence-to-argument links, and substantive analysis per token.The dimensions distinguish surface description from causal chains, deeper implications, limitation probing, connected evidence, and concentrated substance.
  • A.4 ANALYSIS DEPTH: Each report receives a total depth score from 0–25 by summing its five dimension scores.Every dimension is scored from 0 to 5.

B.1 BENCHMARK COMPARISON

LiveResearchBench is compared with prior benchmarks across task, domain, dynamism, verification, and evaluation dimensions. The paper argues that existing benchmarks do not satisfy all criteria simultaneously, whereas LiveResearchBench combines them.

  • B.1 BENCHMARK COMPARISON: No existing benchmark satisfies all comparison criteria simultaneously.The comparison covers human-verified rubrics or answers, expert curation, open-ended design, domain coverage, dynamism, and evaluation methodology.
  • B.1 BENCHMARK COMPARISON: LiveResearchBench combines human-verified rubrics, expert-curated open-ended multidomain queries, time-varying scenarios, and ensemble judging.These design choices are presented as addressing consistency and reliability in benchmark evaluation.
  • B.1 BENCHMARK COMPARISON: Prior long-form evaluation efforts commonly isolate a single quality dimension, whereas DeepEval evaluates multiple content- and report-level dimensions.The comparison contrasts single-call holistic scoring and specialized benchmarks with DeepEval’s broader evaluation design.

C ALIGNMENT OF LLM JUDGES WITH HUMAN EXPERTS

The study examines agreement among LLM judges and alignment with human experts across report-quality dimensions. It also reports that targeted context and citation controls improved DEERFLOW’s evaluation stability and report quality.

  • C ALIGNMENT OF LLM JUDGES WITH HUMAN EXPERTS: 95% of human judgments agreed with LLM judges on the analysis-depth agreement set, while humans preferred Gemini 2.5 Pro or GPT-5 in 92.5% of disagreement cases.The analysis-depth study used 40 agreement-set and 40 disagreement-set samples.
  • C ALIGNMENT OF LLM JUDGES WITH HUMAN EXPERTS: Pointwise additive evaluation achieved 82% agreement with human judgments across 128 detected inconsistencies.The paper reports this scheme as substantially better aligned than direct scoring for consistency and citation traceability.
  • C ALIGNMENT OF LLM JUDGES WITH HUMAN EXPERTS: 87.1% of human citation-accuracy judgments agreed with either Gemini or GPT-5 on claim–URL support decisions.The study sampled 200 claim, URL, and judge-decision pairs.
  • C ALIGNMENT OF LLM JUDGES WITH HUMAN EXPERTS: DEERFLOW+ completed the full evaluation suite without token-limit failures after adding context management and inline-citation validation.The enhanced system also showed better evidence retention, formatting, factual consistency, and citation-related presentation checks.

(b) Deerflow+ (ours)

This section presents a structured, source-based account of visual-art evolution across regions, periods, materials, institutions, and historiographic frameworks. Its evidence base is uneven, with no populated periods for Indigenous American art and incomplete Ancient and Contemporary coverage.

  • Key Points: Cross-regional trade and imperial politics mediated artisans, motifs, and technologies across Eurasia, shaping ceramics, textiles, glass, and architectural vocabularies.The account highlights Silk Road, Indian Ocean, and Venetian-Islamic connections between the 9th and 17th centuries.
  • Key Points: Technical innovations—including porcelain, stonepaste, luster painting, muqarnas, silk weaving, ukiyo-e printing, and European print processes—generated new stylistic systems and markets.The report links these developments to China, the Islamic world, Ottoman workshops, Japan, and nineteenth-century Europe.
  • Key Points: Market and patronage ecologies moved court idioms into broader consumption and global circuits, influencing European modernisms through Japonisme.Examples include Ottoman silks, ukiyo-e, imperial academies, palace workshops, urban publishers, and commercial galleries.
  • Key Points: The report treats debates over figuration, aniconism, and iconoclasm as locally negotiated across sacred and secular spheres rather than as uniform bans.This interpretation emphasizes the interdependence of law, piety, and patronage in Byzantine-Islamic transitions.
  • Analytical Frameworks and Normalized Data: The report organizes visual-art history across Europe, East Asia, the Islamic world, and Indigenous America using normalized timelines, comparative matrices, and cross-cultural transmission pathways.It draws on sources including the Metropolitan Museum of Art’s Heilbrunn Timeline, specialized essays, and collection entries.
  • Analytical Frameworks and Normalized Data: The evidence base robustly documents 9th–19th-century developments but lacks substantive Indigenous American entries and remains incomplete for Ancient and Contemporary endpoints.Claims are constrained to the provided sources, with missing areas explicitly marked as information not provided.

E EVALUATING CITATION ACCURACY WITH RUBRIC TREE

The rubric-tree framework evaluates citation accuracy by grouping claims by source link and checking accessibility, relevance, and support. It reduces repeated retrieval while distinguishing invalid, irrelevant, and unsupported citations.

  • E EVALUATING CITATION ACCURACY WITH RUBRIC TREE: The evaluation identifies each claim and linked URL, checks URL accessibility, and assesses whether accessible content supports the claim.This process categorizes citation failures as invalid links, irrelevant links, or unsupported claims.
  • E EVALUATING CITATION ACCURACY WITH RUBRIC TREE: Grouping claims associated with the same link requires only one web retrieval, reducing the cost of verifying extensive citations.A coarse relevance check first examines webpage metadata or its initial content before support assessment.
  • E EVALUATING CITATION ACCURACY WITH RUBRIC TREE: All evaluated systems produce non-trivial citation errors, with unsupported claims more common than invalid or irrelevant URLs.The result indicates that web access does not eliminate citation hallucinations.
  • E EVALUATING CITATION ACCURACY WITH RUBRIC TREE: Open Deep Research averages 91.9 unsupported-claim errors per market-analysis report, illustrating the severity of citation problems on search-intensive tasks.Market analysis is identified as the more difficult task category in this comparison.
  • E EVALUATING CITATION ACCURACY WITH RUBRIC TREE: Table 7 reports average E1, E2, and E3 errors for leading systems on market-analysis and wide-information-search tasks, with fewer errors preferred.E1 denotes invalid URLs, E2 irrelevant URLs, and E3 unsupported claims.
  • F LIVERESEARCHBENCH DEMONSTRATION: The appendix provides concrete LiveResearchBench queries in Tables 8–10 to demonstrate the benchmark’s task categories.These examples complement the citation-accuracy evaluation framework.

G ERROR PATTERN EXAMPLES

The appendix illustrates structural and citation-related failure patterns that can reduce the clarity and reliability of generated research reports. These observations motivate tailored DeepEval metrics and include both citation-counting conventions and annotation procedures.

  • G ERROR PATTERN EXAMPLES: The pilot study identifies mismatched references, missing links, inconsistent formats, uncited bibliography entries, broken tables, out-of-order references, embedded citations, and hallucinated information.These examples span both citation structure and broader report presentation.
  • G ERROR PATTERN EXAMPLES: E1 and E2 are counted once per problematic URL, whereas E3 is counted once per cited source that fails to support a claim.Thus, a claim with several unsupported cited sources can contribute multiple E3 errors.
  • G ERROR PATTERN EXAMPLES: The examples include reports on health, stress and burnout, ESG reporting, banking, and other applied topics, showing varied contexts for presentation and citation failures.The appendix also includes representative labels such as “Embedded Citations Breaking Text Flow” and “Inconsistent Citation Format.”
  • G ERROR PATTERN EXAMPLES: Professional annotators with advanced degrees and 2–6 years of annotation and analysis experience supported data annotation and verification.Their expertise covered multimodal evaluation and quality control for advanced annotation pipelines.
  • G ERROR PATTERN EXAMPLES: Benchmark design used an online survey spanning enterprise professionals, researchers, and general users to collect questions and preferences for deep-research tasks.The participant groups represented 27%, 24%, and 49% of respondents, respectively.
  • G ERROR PATTERN EXAMPLES: The annotation workflow evaluates both deep-research query quality and checklist quality, with one CSV row representing each grading criterion.The stated goal is to assess report comprehensiveness and coverage through query-linked checklists.

J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY

This task defines criteria for judging deep-research queries and checklist items, emphasizing clarity, scope, audience fit, verifiability, importance, and non-redundancy. Examples show how overly broad or exhaustive requirements should be revised.

  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: Deep-research queries are assessed for answerability, clarity, completeness, appropriate format, and target-audience fit.Useful clarifiers include temporal bounds, geographic scope, specific entities, report organization, and expected reader expertise.
  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: Checklist items are judged for relevance, verifiability, importance, and non-redundancy against the research question.These criteria ensure that checklist items test meaningful and objectively assessable requirements.
  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: The electric-vehicle market query is rated appropriate because it specifies a temporal scope, U.S. geography, relevant entities, and answerable subquestions.Its market-research framing does not require a single narrowly defined audience.
  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: The query about LLM evolution and current trends needs modification because it lacks a defined subject scope, temporal bounds, audience, and precise expectations.Its broad wording leaves unclear whether the intended readers are researchers, practitioners, or the general public.
  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: A checklist item asking for all electric-vehicle companies is inappropriate because exhaustive market completeness lacks authoritative ground truth and is not reasonably verifiable.The recommended alternative focuses on the specific entities named in the query.
  • J.1.1 TASK 1: DEEP RESEARCH QUERY QUALITY: Annotation guidance recommends avoiding unverifiable completeness, redundancy, scope creep, and criteria that are not grounded in the query.Each query is rated once, while checklist items are evaluated individually.

K RESULTS BREAKDOWN

Presentation quality is strongest in high-level organization and simple citation or figure behaviors, while citation correctness, formatting, and reference alignment vary substantially across systems.

  • High-level coherence, citation placement, and figure/table validity are generally reliable, whereas citation correctness and formatting remain difficult.P1, P7, and P8 are strong across systems; P2–P6, P9, and P10 show wider variation.
  • 31%, 32%, and 24% are the P6 scores for GPT-5, GPT-4.1, and GPT-5 Mini, respectively.The shared provider-level pattern illustrates correlated weaknesses among sibling models.
  • Multi-agent wrappers often improve P4 by explicitly mapping in-text citations to references, while P10 numbering accuracy varies by pipeline.Some single-agent web systems exhibit steadier numbering accuracy.
Loading 2510.14240v5…