Source-linked AI summary

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li

arXiv:2608.21374v1cs.AI

TL;DR

Automatically generated literature reviews are difficult to evaluate because research utility depends on expert judgment beyond reference-overlap metrics. The paper introduces LitReview Arena, where matched experts compare anonymized drafts across five dimensions, and reports persistent gaps to human drafts alongside improved expert alignment from a calibrated evaluator.

  • Problem

    Existing review evaluations rely heavily on static rubrics and proxy signals, while many criteria determining research utility are non-verifiable and intuition-based.

  • Method

    LitReview Arena collects blind pairwise preferences from AI paper-writing experts matched to topics, producing five dimension-wise outcomes for LitReviewBench and calibrating LitJudge.

  • Results

    Current systems win only 23.0% of decisive matches against human drafts on overall utility, while LitJudge improves expert-ranking alignment from ρ ≈0.467 to ρ ≈0.792 on D5.

  • Takeaways & Limitations

    The benchmark identifies coherent landscape organization and non-obvious research-direction discovery as key areas where literature review agents still need progress.

  • Takeaways & Limitations

    Broader validation across fields with different evidential norms remains future work, and uncritical use of expert-calibrated preferences may reinforce mainstream styles or fashionable topics.

Abstract

from arXiv · show

Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspects of research utility depend on expert judgment rather than reference-overlap metrics. We introduce LitReview Arena, a battle-style evaluation platform with a structured protocol tailored to literature review quality: domain experts with AI paper-writing experience compare anonymized drafts, are matched to topics within their expertise, and provide dimension-wise outcomes over five literature-review-specific criteria. From this protocol, we collect approximately 3k expert judgments, each containing five dimension-wise outcomes, and show that even the strongest current systems win only 23.0% of decisive matches against human drafts on overall utility, while agentic LLMs such as Sonar Deep Research substantially outperform base language models by over 60%. We further find that existing LLM-as-a-judge methods are substantially misaligned with human experts (Spearman's rho=0.467), especially on synthesis-heavy criteria such as paper structure and research suggestions. Using the collected preference data, we provide an expert-calibrated evaluator, LitJudge, which improves alignment to Spearman's rho=0.78, comparable to inter-expert consistency; code and data are publicly available at https://github.com/VanellopeAsher/LitReview-Arena.

1. Introduction

LitReview Arena addresses the difficulty of evaluating literature reviews by using expert, blind, dimension-wise comparisons tailored to scholarly synthesis. Its benchmark shows substantial gaps between current systems and human drafts, while calibrated judging improves alignment with expert preferences.

  • Evaluation gap: Existing review evaluations emphasize retrieval, organization, reference correctness, coverage, or overlap, but do not fully capture non-verifiable, intuition-based research utility.The paper frames structure and gap identification as central challenges beyond surface completeness.
  • Protocol: The protocol recruits AI paper-writing experts, derives topics from real surveys, matches annotators to relevant expertise, and evaluates five literature-review-specific dimensions.These design choices target reliable and reproducible expert preference data.
  • Findings: 3k expert judgments show that non-human systems win only 23.0% of decisive matches against human drafts on overall utility.Paper Structure and Research Suggestions trail human drafts by roughly 200 points.
  • Findings: Agentic models substantially outperform pure language models across dimensions, but consume roughly 15× the token budget for the same topic.The result identifies a performance–test-time-compute trade-off.
  • Evaluator: LitReviewBench calibration improves expert-ranking alignment for LitJudge from ρ ≈0.467 to ρ ≈0.792 on D5.The evaluator is intended for low-cost offline assessment and preference-based training signals.

2. Related Work

Prior literature-review systems and benchmarks improve retrieval, coverage, fluency, and citation handling, but often under-measure expert-level structure and gap quality. LitReviewBench combines arena-based expert preferences with a frozen, versioned benchmark and calibrated evaluator to address that gap.

  • Automated review generation: End-to-end review systems collect papers, organize outlines, and generate cited narratives, while evaluations largely emphasize operationalizable criteria such as references and coherence.These systems remain closer to search and integration tools than expert-level synthesizers.
  • Automated review generation: Coverage and fluency improvements do not necessarily produce a defensible global structure or non-generic, researcher-useful gap narrative.The paper identifies organizing axes and actionable gaps as under-specified qualities.
  • Static benchmarks: Existing benchmarks skew toward scalable signals such as topical relevance, summary similarity, citation matching, and claim-level factuality.These signals are easier to measure reliably than expert-valued scholarly synthesis.
  • Static benchmarks: LLM-as-a-judge methods can drift with prompts, overweight surface features, and only partially match expert preferences on scientific literature-grounded tasks.The paper therefore treats off-the-shelf judging setups as insufficient primary development signals.
  • Arena-based evaluation: LitReviewBench centers landscape structure and gap quality while retaining citation-aware dimensions, using topic-matched researchers and an evaluator calibrated to expert trade-offs.It complements rather than replaces existing benchmarks.
  • Arena-based evaluation: Arena advantages are converted into a frozen, versioned dataset with standardized per-dimension outcomes for reproducible offline measurement.The benchmark preserves topic diversity and pairwise expert preference while reducing dependence on live arena evaluation.

3. The LitReviewBench Construction

LitReviewBench is constructed by defining survey-grounded topics, collecting matched expert comparisons of anonymized drafts, and freezing the resulting five-dimensional arena outcomes into a reproducible benchmark. Its records support diagnostic analysis and standardized offline leaderboards.

  • Construction stages: The benchmark defines survey-draft generation topics from high-quality survey literature, maps them into a consistent taxonomy, and freezes arena logs into a versioned offline dataset.The construction has three stages: task and taxonomy definition, expert preference collection, and benchmark freezing.
  • Task definition: The task requires drafts to list relevant papers, organize a field into an interpretable landscape, and surface gaps that can inform next steps.This operationalizes usefulness to researchers beyond paper retrieval.
  • Topic construction: Topics are extracted from recent, highly cited OpenAlex survey papers and normalized into a consistent survey-prompt form.The source pool contains more than 3,000 survey-style papers published from 2022 to 2025 with over 50 citations.
  • Topic construction: AAAI-based field and subfield tags support expertise matching during annotation and stratified generalization analysis.The taxonomy is refined for modern AI.
  • Arena protocol: Each topic presents two anonymized, randomly ordered drafts side by side for expert comparison, stored as a battle record with the query, drafts, and outcomes.The battle format supports Bradley–Terry and Elo-style aggregation.
  • Expert annotation: Experts are matched to topics within their fields, receive shared instructions, and judge drafts using blind comparisons with written reasoning and safeguards against surface-fluency bias.Matching is especially important for review structure and gap-quality judgments.
  • Evaluation dimensions: Every battle records five four-way votes—A, B, Tie, or Both Bad—covering Literature Coverage, Claim Support, Paper Structure, Research Suggestions Quality, and Overall Utility.D1–D4 provide diagnostic signals alongside the overall D5 preference.
  • Offline benchmark: Frozen benchmark instances preserve the query, paired drafts, D1–D5 outcomes, and aggregation metadata, with explicit Tie and Both Bad labels for standardized leaderboards.Versioning is intended to support reproducibility over time.

4. Results on SOTA Models and Agents

Expert-preference evaluation shows that human drafts remain substantially ahead of current models, while agentic systems trade higher performance for much greater computation. The largest weaknesses involve synthesis, organization, and actionable research direction rather than basic coverage alone.

  • Overall Performance: 208 of 904 decisive matches, or 23.0%, were won by state-of-the-art models and agents against human drafts on Overall Utility (D5).Human drafts rank first across all five dimensions.
  • Overall Performance: GPT-5.2 leads non-human systems on Literature Coverage, Claim Support, and Paper Structure, while Sonar Deep Research leads on Research Suggestions but still trails humans.The standings indicate persistent difficulty with field organization and actionable direction-setting.
  • Efficiency: 122.3K tokens per query is the average cost of agentic systems, compared with 8.1K for standalone models, a 15× difference.Search and synthesis produce comprehensive gains over raw models but do not close the utility gap to expert expectations.
  • Overall Performance: Open Deep Research and SurveyForge outperform average agentic and language-model baselines but remain below human drafts across all five dimensions.These systems are designed specifically for literature review generation.
  • Capability Gaps: Paper Structure (D3) and Research Suggestions (D4) show the largest deficits and correlate most strongly with Overall Utility, at 0.99 and 0.96 respectively.Literature Coverage correlates at 0.90, while observed failure modes also include incomplete coverage and weak claim support.
  • Capability Gaps: Increasing draft length or interaction turns does not automatically improve D3 or D4 scores; experts reward drafts that make latent relationships legible.Reference overlap and checklist-style coverage can therefore miss the reasoning gap exposed by the arena format.

5. Meta-Evaluation of LLM-Based Judges

The meta-evaluation finds that uncalibrated LLM judges align with experts on retrieval-oriented coverage more than on synthesis-heavy review qualities. Their stable cross-model preferences can nevertheless diverge systematically from expert judgments, especially when comparing human and machine drafts.

  • Alignment by Dimension: Qwen correlates moderately with experts on Literature Coverage (D1), with Spearman’s ρ = 0.552.This dimension primarily tests retrieval recall and recognition of canonical papers.
  • Alignment by Dimension: On Claim Support (D2), Paper Structure (D3), and Research Suggestions (D4), the uncalibrated judge struggles to distinguish high-quality analysis from plausible but incorrect statements.Performance degrades as evaluation shifts from retrieval toward synthesis.
  • Systematic Bias: On Overall Utility (D5), the automated judge assigns humans an Elo score of 310 while boosting GPT-5.2 to 2490, reversing the expert ranking.The result suggests model-like fluency is being conflated with quality in uncalibrated human–machine comparisons.
  • Consistency and Reliability: Expert agreement reaches 0.861 on Overall Utility (D5), while judge-to-judge agreement remains above 0.7 across dimensions despite poor expert alignment.The pattern indicates systematic shared AI-side preferences rather than random judge noise.
  • Consistency and Reliability: Judge agreement can reflect shared preferences for surface fluency and formatting while under-weighting grounded claims and insightful research directions.These shared biases help explain why high LLM–LLM consistency does not establish expert-level reliability.
  • Implications: Naive LLM judging is an insufficient substitute for expert preference, and the same naive-versus-calibrated gap appears across GPT, Claude, and Llama as well as Qwen.Uncalibrated leaderboards can be deceptively stable yet inaccurate, motivating task-specific calibration.

6. Closing the Loop: Expert-Aligned Evaluator

LitJudge uses LitReviewBench preferences to calibrate automated evaluation with examples matched to review structure, topic, and expert-derived research gaps. Calibration most improves synthesis-oriented alignment and supports repeated offline evaluation at low marginal cost.

  • Evaluator Design: LitJudge conditions on case context matched to the test instance in structure and topic, with diversity-aware gap anchors from expert-written drafts for D4.The evaluator is designed to address dimensions where standard LLM judges diverge most from experts.
  • Context Construction: Structure-similar demonstrations use paragraph-relationship networks, content-similar cases support D1/D2, and expert-only gap anchors support D4.Up to three demonstrations per group are assembled for the in-context packet.
  • Calibration Results: Calibration improves Literature Coverage (D1) alignment from 0.552 to 0.576, Claim Support (D2) from 0.442 to 0.673, and Paper Structure (D3) from 0.467 to 0.649.Coverage changes modestly while synthesis dimensions improve substantially.
  • Calibration Results: Research Suggestions (D4) alignment rises from 0.430 to 0.842, consistent with the use of expert-derived gap anchors.This is the largest reported dimension-wise improvement.
  • Calibration Results: Overall Utility (D5) alignment increases from 0.467 to 0.792, indicating that calibration helps recover the expert holistic preference signal.The improvement is reported alongside stronger gains on synthesis dimensions.
  • Generalization and Utility: LitJudge outperforms naive judging across GPT, Claude, Llama, and Qwen backbones and beats a same-example-count random few-shot control.The evaluator enables repeated offline evaluation that tracks expert tradeoffs with low marginal cost.

7. Discussion

The discussion identifies synthesis—not surface coverage—as the central bottleneck for literature-review agents and evaluation, while outlining scalable expert-calibrated benchmarking and important scope limits.

  • Evaluation bottleneck: Uncalibrated LLM judges track experts moderately well on Literature Coverage but become substantially misaligned on synthesis-heavy Paper Structure and Research Suggestions.The contrast indicates that measuring scholarly organization and actionable gaps is harder than verifying retrieved literature.
  • Model capabilities: Agentic and specialized literature-review systems improve over base language models, but their largest remaining gaps to human drafts are on Paper Structure and Research Suggestions.Search, decomposition, and iterative synthesis help, yet more computation or retrieved papers alone is insufficient for meaningful conceptual structure and non-obvious directions.
  • Benchmark design: Relative preference votes, explicit Tie and BothBad outcomes, and frozen arena logs make expert review preferences more consistent, repeatable, and usable for offline evaluation.The live arena can also refresh the benchmark as topics, systems, and expert expectations change.
  • Calibrated evaluation: LitJudge uses structure-matched cases and diversity-aware expert-derived gap anchors to support low-cost offline iteration while reducing concentration on recurring research directions.The anchor-selection goal is to preserve agreement with expert judgment while lowering academic echo-chamber risk.
  • Limitations: The main benchmark is anchored in AI topics, and broader validation across fields with different evidential norms remains future work despite initial biology-pilot evidence.Biomedicine and law may additionally require stricter factual, safety, and citation audits.
  • Limitations: Expert-calibrated preferences may entrench familiar rhetorical forms or fashionable directions, and the benchmark does not cover long-horizon living-review maintenance.Suggested complements include factual audits, citation-grounding checks, and longitudinal update tasks.

8. Conclusion

The paper introduces LitReview Arena, LitReviewBench, and LitJudge to align literature-review-agent evaluation with expert judgment and researcher needs. Results show persistent human–system gaps, especially in coherent organization and non-obvious research insights.

  • LitReview Arena is a battle-style peer-review platform, while LitReviewBench is an offline benchmark distilled from arena logs.
  • 3k expert judgments show that current state-of-the-art systems lag behind human drafts, particularly in organizing coherent landscapes and generating non-obvious research insights.
  • LitJudge uses structure-matched cases and diversity-aware expert-derived anchors to improve agreement with experts over uncalibrated LLM judges.
  • The benchmark indicates that future agents should move beyond broad paper collection toward defensible synthesis, grounded claims, and field-level reasoning.

Impact Statement

The work aims to make evaluation of scientific literature-review generators more transparent, reproducible, and aligned with expert judgment. It also warns that automated evaluators should supplement rather than replace human and factual review.

  • The benchmark and LitJudge may help researchers identify weaknesses in coverage, claim support, structure, and research suggestions before deployment in research workflows.
  • LitJudge should serve as a development and triage signal rather than a replacement for human expert review.
  • Preference-based evaluation should be combined with factual checks of claims and citations because automated evaluators may be over-relied upon or reinforce dominant research styles.

A. Appendix

The appendix defines the battle-record schema, expert comparison dimensions, scoring conventions, and evaluator controls, and reports specialized-system and biology-pilot results.

  • Evaluation outcomes: Each dimension uses one of four outcomes: A, B, Tie, or BothBad.
  • A.1. Expert Evaluation Protocol: Experts compare anonymized drafts side by side with randomized order, judge dimensions independently, provide written reasoning, and avoid treating fluency, citation count, or generic future work as sufficient quality.
  • A.2. Data Schema (Battle Record): Each battle record stores a normalized query, two draft responses, five dimension labels, taxonomy tags, a pseudonymized annotator identifier, and optional audit metadata.
  • A.3. Aggregation and Scoring Conventions: Bradley–Terry and Elo aggregate dimension-wise preferences, treating ties as half wins or half scores for both drafts.
  • Evaluation dimensions: The five dimensions assess Overall Utility, Literature Coverage, Claim Support, Paper Structure, and Research Suggestions Quality.
  • A.5. Calibrated Expert-Aligned Evaluator: LitJudge uses structure-similar, content-similar, and gap-anchor examples, while controls show its gains exceed random few-shot prompting across multiple evaluator backbones.
  • Results: Specialized literature-review systems outperform average agentic and language-model baselines but remain below human drafts on every dimension.
  • Biology pilot: The biology pilot preserves the high-level ordering of human drafts over agentic models over language models, with the largest gaps on Paper Structure and Research Suggestions.
Loading 2608.21374v1…