Source-linked AI summary
DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report
Ruizhe Li, Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao
TL;DR
Deep-research benchmarks do not consistently test evidence analysis and coherent reporting, and their evaluation criteria can be coarse or misaligned with human experts. DeepResearch Bench II addresses this gap with expert-derived, fine-grained rubrics across grounded research tasks, finding that even the strongest agents satisfy fewer than 50% of rubrics. The benchmark therefore provides a more verifiable evaluation of information recall, analysis, and presentation while exposing a substantial gap from human experts.
Problem
Existing benchmarks inadequately test open-ended evidence analysis and reporting, while coarse or LLM-defined criteria can be difficult to verify and misaligned with human expertise.
Method
DeepResearch Bench II builds 132 tasks across 22 domains and 9,430 fine-grained binary rubrics from expert reports through LLM extraction, filtering, manual cleaning, and domain-expert refinement.
Results
Even the strongest evaluated agents satisfy fewer than 50% of the rubrics, with especially large deficits in Information Recall and Analysis.
Takeaways & Limitations
The benchmark provides a grounded, verifiable framework for assessing information recall, analysis, and presentation and for quantifying current agents’ gap from human experts.
Takeaways & Limitations
Prompt-based restrictions did not fully prevent source-article leakage, which may make reported scores differ from agents’ actual performance.
Abstract
from arXiv · showhide
Deep Research Systems (DRS) aim to help users search the web, synthesize information, and deliver comprehensive investigative reports. However, how to rigorously evaluate these systems remains under-explored. Existing deep-research benchmarks often fall into two failure modes. Some do not adequately test a system's ability to analyze evidence and write coherent reports. Others rely on evaluation criteria that are either overly coarse or directly defined by LLMs (or both), leading to scores that can be biased relative to human experts and are hard to verify or interpret. To address these issues, we introduce Deep Research Bench II, a new benchmark for evaluating DRS-generated reports. It contains 132 grounded research tasks across 22 domains; for each task, a system must produce a long-form research report that is evaluated by a set of 9430 fine-grained binary rubrics in total, covering three dimensions: information recall, analysis, and presentation. All rubrics are derived from carefully selected expert-written investigative articles and are constructed through a four-stage LLM+human pipeline that combines automatic extraction with over 400 human-hours of expert review, ensuring that the criteria are atomic, verifiable, and aligned with human expert judgment. We evaluate several state-of-the-art deep-research systems on Deep Research Bench II and find that even the strongest models satisfy fewer than 50% of the rubrics, revealing a substantial gap between current DRSs and human experts.
1 Introduction
DeepResearch Bench II addresses shortcomings in deep-research evaluation by grounding fine-grained rubrics in expert reports and assessing recall, analysis, and presentation. Its evaluation of leading agents reveals a substantial gap from human experts.
- 132 tasks across 22 domains are built from high-quality, expert-written investigative reports and filtered through quality checks.
- 9,430 fine-grained binary rubrics assess Information Recall, Analysis, and Presentation through an LLM-based end-to-end judging protocol.The rubrics encode concrete factual or inferential requirements and provide dimension-wise scores for each task.
- Even the strongest evaluated agents fail to pass more than 50% of the rubrics, with especially large deficits in Information Recall and Analysis.The study also includes human–LLM agreement analyses and examines systematic weaknesses in existing models.
- The benchmark combines real expert reports, verifiable rubrics, and a three-dimensional framework to evaluate deep research agents comprehensively.
2 Related Work
Prior benchmarks either focus on fixed-answer retrieval or evaluate open-ended reports with criteria that can be coarse, unverifiable, or misaligned with human expertise. DeepResearch Bench II instead compares evaluation schemes and emphasizes expert-grounded, fine-grained criteria.
- Fixed-answer benchmarks test retrieval of specific entities or numbers but only partially reflect open-ended research needs.
- LLM-defined criteria can misalign with human experts, while coarse human-written rubrics may allow seemingly correct hallucinations to pass.
- Open-ended benchmarks assess report qualities such as comprehensiveness, insight, and citation accuracy but do not verify recalled information like fixed-answer benchmarks.
- Table 1 compares methods by real-world topics, expert-sourced rubrics, internal-knowledge verifiability, and average rubrics per task, with OURS meeting all criteria.
- Rubric granularity determines whether judges can verify content accurately, especially for deep-research tasks that cannot be checked against internal model knowledge.
3 Methodology
DeepResearch Bench II builds research tasks from expert-written reports and evaluates generated reports with fine-grained, verifiable rubrics. Its four-stage construction process combines LLM extraction with manual and domain-expert refinement.
- Source Article Selection: 132 expert-written articles were retained as source materials for constructing research tasks and extracting ground-truth rubrics.
- Task and Rubric Design Protocol: Each task requires information collection and non-trivial analysis, while time-sensitive investigations specify the source article’s temporal scope.
- Task and Rubric Design Protocol: Rubrics are binary, atomic criteria that encode essential factual or inferential requirements and are aggregated into three evaluation dimensions.
- Task and Rubric Construction Pipeline: The rubric-construction pipeline consists of LLM extraction, self-evaluation iteration, manual revision, and expert review and refinement.
- Task and Rubric Construction Pipeline: Over 400 human-hours of domain-expert review help ensure that the resulting tasks and rubrics are grounded, verifiable, fine-grained, and aligned with human evaluation standards.
4 Experiment
The experiments benchmark diverse deep-research agents across rubric dimensions and examine evaluation-pipeline choices. OpenAI-GPT-o3 leads overall, but every agent shows substantial weaknesses across retrieval, analysis, or presentation.
- Main Results: The evaluated systems exhibit substantial disparities in long-context synthesis, structured reasoning, and task-aligned report generation.
- Main Results: OpenAI-GPT-o3 Deep Research achieves the highest Information Recall and overall aggregated score, while even the best agent satisfies fewer than half the rubrics.
- Main Results: Gemini-3-Pro and Gemini-2.5-Pro show strong Information Recall and Analysis, whereas Grok Deep Search leads Presentation but trails in retrieval and reasoning.
- Main Results: Tongyi Deep Research ranks near the bottom across dimensions, with weaker information-seeking behavior and less effective organization and communication.
- Rubric Batch Size: Batch size 50 offers the best balance between evaluation cost and quality, delivering strong accuracy while keeping overhead manageable.
- Evaluator Model Choice: Gemini-2.5-Pro shows the highest alignment with human annotations across both reported metrics and is selected as the evaluator.
5 Analysis and Discussion
Robustness analyses find generally stable performance across languages and topics, while source-article access remains an evaluation risk and user-adaptive presentation is not yet fully realized.
- Language Effects: Most models show no significant language-based performance differences, while Grok shows a marginal difference (p = 0.034).
- Topic Effects: All evaluated agents show no significant topic-based performance variation across Health, Finance & Business, Software Development, and other research topics (p > 0.05).
- Source Article Leakage: Prompt-based restrictions cannot guarantee that closed-source models avoid accessing source articles and may affect evaluation performance.
- Future Directions: User-adaptive presentation remains an open research direction because prompts alone do not reliably convey users’ cognitive levels and background knowledge.
6 Conclusion
Deep Research Bench II is presented as a comprehensive, more robust benchmark for evaluating deep research agents across information recall, analysis, and presentation. The authors argue that it can provide a more accurate measure of model capabilities and facilitate more effective research tools.
- Deep Research Bench II evaluates deep research agents across Information Recall, Analysis, and Presentation.
- The benchmark is designed around real-world user needs and deconstructed research tasks.
- The authors argue that the benchmark provides a more accurate measure of model capabilities and can facilitate more effective research tools.
7 Limitations
The benchmark has limitations involving source-article leakage, subjectivity in human annotations, and incomplete assessment of personalized presentation. These constraints may affect evaluation results and leave personalization for future work.
- Prompt-based restrictions did not fully prevent source articles from leaking into closed-source model search results.
- Subjective judgments by human annotators could introduce bias into the benchmark’s final results.
- The presentation dimension assesses report formatting and layout but not personalization to users’ knowledge backgrounds and preferences.
- Improving personalized presentation requires integration with advances in agent memory.
8 Potential Risks
The benchmark uses licensed human-expert articles under specified Creative Commons terms, but commercial use as training data remains a potential intellectual-property risk.
- The benchmark incorporates human-expert articles licensed under CC-by-4.0 or CC-BY-4.0-NC and follows those licensing terms.
- Commercial entities could potentially use the benchmark as training data, creating an intellectual-property concern.
A.1 Task Topic and Language Statistics
The benchmark reports task-topic and language distributions, alongside rubric-count distributions across its three evaluation dimensions.
- Table 7 organizes tasks by topic and reports English, Chinese, and total counts.
- Tasks average 52.902 InfoRecall rubrics, 12.773 Analysis rubrics, and 5.652 Presentation rubrics.
- Figure 5 shows the frequency distribution of rubric counts per task for InfoRecall, Analysis, and Presentation.
B Detailed Result
Table 8 presents model performance across thematic categories.
- Table 8 reports model performance across thematic categories.
C Source Article List
The source-article list contains research articles spanning diverse domains, including finance, health, technology, policy, and the sciences.
- It also includes articles on health, aging, insurance, and medical monitoring.
- The benchmark includes source articles on finance and investment topics, such as portfolio diversification, green bonds, and sovereign wealth funds.
- Technology-focused sources cover deep learning, quantum technologies, batteries, edge computing, and software development.
- Additional sources address materials science, biology, animal navigation, corrosion, energy recovery, and government or social policy.
D Source Article Leakage Rate
The evaluation attempts to prevent source-article exposure through blocked lists and secondary inspection, while the benchmark prompt defines three scoring dimensions and structured task requirements.
- The benchmark uses a blocked list and manually checks generated reports for references to source articles.
- The leakage rate is non-zero for every model, but all models except Qwen remain within 5%.
- Deepresearch is scored on information recall, analysis, and presentation.
- Tasks require models to retrieve information from the open internet and produce reports resembling the human reference article or a specified section.
- The prompt requires structured coverage of biography, theoretical contributions, international roles, disciplinary impact, and major works.