Source-linked AI summary

DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity

Joey Zhong, Hao Zhang, Clare Southern, Jeremy Yang, Thomas Wang, Kate Jung, Shu Zhang, Denis Yarats, Johnny Ho, Jerry Ma

arXiv:2602.11685v1cs.LGcs.AI

TL;DR

Deep research systems need evaluation that reflects complex, open-ended work across diverse domains and information sources. DRACO addresses this gap with a 100-task benchmark built from anonymized production requests and graded by task-specific rubrics. Across its evaluation, Perplexity Deep Research achieves the strongest overall performance, while the benchmark identifies remaining headroom in factual accuracy.

  • Problem

    Evaluating deep research systems requires benchmarks that reflect realistic tasks across domains, regions, heterogeneous sources, and multiple capabilities.

  • Method

    DRACO samples, filters, and augments anonymized production requests into 100 tasks spanning 10 domains and 40 countries, then evaluates outputs with task-specific rubrics.

  • Results

    Perplexity Deep Research consistently achieves the strongest overall score and pass rate across domains and rubric categories.

  • Takeaways & Limitations

    DRACO provides a cross-domain evaluation of deep research quality across factual accuracy, analytical breadth and depth, presentation quality, and citation quality.

  • Takeaways & Limitations

    The benchmark still has gaps relative to current and future real-world use, and rubric creation relies heavily on human experts with substantial variation across judges.

Abstract

from arXiv · show

We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sources from 40 countries, originate from anonymized real-world usage patterns within a large-scale deep research system. Tasks are sampled from a de-identified dataset of Perplexity Deep Research requests, then filtered and augmented to ensure that the tasks are anonymized, open-ended and complex, objectively evaluable, and representative of the broad scope of real-world deep research use cases. Outputs are graded against task-specific rubrics along four dimensions: factual accuracy (accuracy), breadth and depth of analysis (including completeness), presentation quality (including objectivity), and citation quality. DRACO is publicly available at https://hf.co/datasets/perplexity-ai/draco.

1 Introduction

DRACO addresses the difficulty of evaluating deep research systems across realistic, diverse, open-ended tasks. It introduces a 100-task benchmark with expert-grounded rubrics covering accuracy, analytical quality, presentation, and citations.

  • Motivation: Deep research systems decompose complex queries, retrieve diverse information iteratively, and synthesize structured, cited reports.These systems support multi-step planning, evidence verification, conflict resolution, and gap identification.
  • Motivation: Evaluating deep research requires realistic tasks spanning domains, regions, information sources, and underlying capabilities.This combination creates a curse of dimensionality for benchmark construction.
  • DRACO: 100 tasks span 10 domains and require information from 40 countries, with tasks drawn from actual user requests and paired with expert-grounded rubrics.The tasks are sampled from tens of millions of requests, then filtered and augmented for privacy, rigor, and representativeness.
  • Evaluation: Outputs are graded for factual accuracy, breadth and depth of analysis, presentation quality, and citation quality.These dimensions jointly assess correctness, completeness, objectivity, and sourcing.
  • Results: Perplexity Deep Research consistently achieves the strongest overall score and pass rate across domains and rubric categories.The evaluation compares publicly available versions of OpenAI Deep Research, Gemini Deep Research, Claude Opus, and Perplexity Deep Research.

2 Related Work

DRACO is positioned as an open-ended benchmark designed to better reflect real deep research needs than closed-ended evaluations. Its construction uses production tasks that are reformulated, augmented, and filtered for rigorous assessment.

  • Benchmark landscape: Many existing benchmarks use closed-ended tasks whose solutions can be checked deterministically, whereas real deep research often requires judgment and open-ended analysis.Such benchmarks still test retrieval, synthesis, planning, and reasoning, but do not fully represent open-ended research work.
  • Benchmark landscape: DRACO compares open-ended benchmarks by production origin, authorship, domain coverage, and expert-designed rubrics.The comparison identifies dimensions relevant to how closely benchmarks reflect authentic research use.
  • DRACO: DRACO constructs tasks from actual Perplexity Deep Research requests, reformulating and augmenting them to protect privacy and create challenging, realistic evaluations.The pipeline is designed to automate fresh task generation while retaining human review as a safety and quality gate.

3 Task Construction

DRACO builds its benchmark by transforming difficult production queries into anonymous, bounded, objective, and challenging research tasks. The pipeline samples broadly, augments query context and scope, filters for evaluability, and manually curates the final set.

  • Pipeline: DRACO sources production queries, then reformulates, augments, and filters them before expert verification.The process is summarized in Figure 1 and is designed to produce anonymous, well-specified tasks demanding open-ended analysis.
  • Sampling: 1,000 high-difficulty English queries from September–October 2025 were sampled across 10 general and specialized domains.Difficulty was proxied by negative user sentiment or an explicit thumbs-down rating on a prior response.
  • Pre-processing: An automated preprocessing stage removes personally identifiable information and reduces ambiguity without exposing raw queries to human analysts.The preprocessing operates end-to-end through an LLM-based pipeline.
  • Augmentation: Query augmentation adds user context and broadens scope through longer time horizons, comparisons, and geographic variation.These changes turn ambiguous queries into defined tasks with consistent evaluation criteria.
  • Filtering and curation: Augmented queries are filtered for objectivity, tractability, and difficulty before 100 are sampled and manually reviewed by domain experts.Objectivity requires measurable criteria, tractability requires bounded scope, and difficulty requires nontrivial information gathering and multi-step reasoning.

4 Rubric and Grading

DRACO uses expert-built, iteratively reviewed rubrics to assess deep-research outputs across four weighted quality axes. Its grading aggregates criterion-level verdicts into normalized scores and pass rates.

  • 4.1 Rubric: Rubrics are designed by domain experts through initial drafting, iterative review, saturation testing, and final quality assurance.Tasks exceeding 90% during saturation testing are revised; about 45% return for revision then, and about 10% do so during final review.
  • 4.1 Rubric: Each task averages 39.3 criteria, with 20.5 targeting factual accuracy and the remainder covering analysis, presentation, and citation quality.The four axes allocate approximately 22% to analysis, 14% to presentation, and 12% to citation quality.
  • 4.1 Rubric: Rubrics include positive criteria and negative criteria that penalize undesirable properties such as unsupported claims.Of 3,934 total criteria, 415 are negative; negative criteria are especially prevalent in Presentation Quality.
  • 4.1 Rubric: Evaluation granularity varies by domain, from 30.2 criteria for Needle in a Haystack to 47.6 for Finance.Law and Medicine have the highest proportions of negative criteria, reflecting heightened scrutiny of errors in these domains.
  • 4.1 Rubric: Factual Accuracy dominates every domain, while Finance and Academic tasks require the most criteria and Citation Quality varies from 5.8 in Academic to 3.0 in Medicine.Finance and Academic average 47.6 and 41.6 criteria per task, respectively; Needle in a Haystack averages 30.2.
  • 4.2 Grading: For each criterion, an open-source LLM judge returns a binary MET or UNMET verdict with a short justification.MET contributes the criterion weight, whereas UNMET contributes zero; negative weights can penalize undesirable properties.

5 Experiments and Results

The experiments evaluate deep research systems on a 100-task benchmark using rubric-based LLM judging, reporting overall performance, efficiency, domain coverage, and rubric-axis results. Perplexity Deep Research leads overall and across domains and rubric categories, with substantial variation in latency, token usage, and performance gaps.

  • Evaluation setup: 100 tasks were evaluated with per-criterion LLM judging, aggregated into overall scores, normalized scores, and pass rates.The benchmark used five independent grading runs for normalized scores, with standard deviations measuring variability across runs.
  • Overall results: 70.5% is Perplexity Deep Research’s normalized score with Opus 4.6, ahead of Gemini Deep Research at 59.0%, OpenAI o3 at 52.1%, and OpenAI o4-mini at 41.9%.The reported ordering is consistent with the overall pass-rate pattern.
  • Efficiency: 245.3 seconds is Perplexity Deep Research’s average latency, while OpenAI Deep Research o3 reaches 1808.1 seconds and Perplexity uses 778,711 average input tokens.OpenAI o3 and Gemini produce more output tokens, but their lower normalized scores show that longer outputs do not necessarily yield higher benchmark performance.
  • Results by domain: Perplexity Deep Research ranks first across all ten domains, with strongest absolute scores in Law at 90.2% and Academic at 82.8%.The gap over the second-best model is largest in Finance at 21.6 percentage points and smallest in Law at 1.6 percentage points.
  • Results by rubric axis: Perplexity leads all four rubric axes: Factual Accuracy 67.9%, Breadth and Depth of Analysis 73.1%, Presentation Quality 90.3%, and Citation Quality 64.6%.The largest gaps over the second-best model occur in Breadth and Depth of Analysis at 13.2 percentage points and Factual Accuracy at 10.1 percentage points.

6 Discussion

DRACO is designed to measure deep research systems against authentic, complex research needs, while acknowledging limits in generalization, evaluation scalability, and component-level diagnosis.

  • DRACO bridges AI evaluation and authentic research needs through a benchmark derived from real-world production tasks.
  • 6.1 Generalization: The benchmark remains limited because current tasks do not fully capture present or future system use.
  • 6.1 Generalization: Single-turn evaluation excludes capabilities such as asking relevant clarifying questions.
  • 6.1 Generalization: The benchmark is static, so it may not fully generalize to future deep research applications.
  • 6.1 Generalization: Text-to-text evaluation omits multimodal verification as agents begin processing images and videos.
  • Human expert involvement makes rubric creation costly, while different judges produce substantial variation in score magnitudes.
  • 6.3 Attribution and Decomposition: End-to-end scoring does not isolate retrieval, source selection, planning, or synthesis failures, motivating component-level evaluation and targeted ablations.
  • DRACO is offered as a foundation for measuring and improving deep research systems in production settings.

A Alternative LLM Judges

The alternative-judge analysis tests whether DRACO's system rankings depend on the evaluator model. Absolute scores differ across judges, but relative ordering remains stable.

  • Three LLM judges—Gemini-3-Pro, GPT-5.2, and Sonnet-4.5—were used to assess robustness across five systems and Claude Opus versions.
  • GPT-5.2 consistently assigned lower absolute scores than the other judges.
  • Relative system ordering remained stable across all three judges.
  • Table 15 reports normalized scores, with bold marking the best result and underlining the second-best non-Perplexity result.
  • Judge configurations varied in reasoning effort and temperature to reduce variability or follow model-specific guidance.

B Query Augmentation Examples

The augmentation examples show how raw research requests are transformed into bounded, comparative, and context-rich tasks across multiple domains.

  • Finance augmentation expands a broad industrial-automation query into a Saudi Arabia study covering market metrics, policy, projects, and sales implications.
  • A photography request specifies three medium-format systems and compares synchronization, tethering, skin tones, workflow speed, and lens costs.
  • The historiography query broadens analysis across four regions and combines archaeological, linguistic, manuscript, and geopolitical perspectives.
  • The deepfake query adds video and audio methods, cross-dataset generalization, multimodal analysis, foundation models, privacy, deployment, ethics, and regulation.
  • The agriculture query compares mega-farm expansion and resistance across Ukraine, Brazil, Saudi-linked investments in the United States, and China.
  • The code-completion query frames interface timing and quality as comparisons across tools and developer-experience groups.

C Prompts

The prompts operationalize privacy-safe preprocessing, query augmentation, and objective filtering to turn raw requests into challenging, evaluable research tasks.

  • C.1 Pre-Processing Prompt: Preprocessing removes personally identifiable information using generic placeholders for direct, financial, digital, physical, professional, and medical identifiers.
  • C.1 Pre-Processing Prompt: Ambiguity resolution adds explicit referents, assumed context, precise interpretations, temporal and geographic bounds, quantities, and research scope.
  • C.1 Pre-Processing Prompt: The preprocessing output records the processed query, removed PII, and added clarifications while preserving the original intent.
  • C.1 Pre-Processing Prompt: Examples show clarification of temporal scope, technical comparisons, evidence standards, intervention categories, and named alignment methods.
  • C.2 Augmentation Prompts: Augmentation prompts add authoritative sources, explicit date ranges, professional personas, comparable entities, geographic regions, and deliverable formats when useful.
  • C.2 Augmentation Prompts: Augmentation examples specify date ranges, peer entities, geographic scope, persona, and output format through chained prompt steps.
  • C.3 Filtering Prompt: Filtering requires evaluation questions to be objective, bounded, and difficult enough to demand nontrivial information gathering and synthesis.
Loading 2602.11685v1…