Source-linked AI summary

NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment

Wenqing Wu, Yi Zhao, Yuzhuo Wang, Siyou Li, Juexi Shao, Yunfei Long, Chengzhi Zhang

arXiv:2604.11543v1cs.CLcs.AIcs.DLcs.IR

TL;DR

Assessing academic novelty is increasingly difficult for reviewers, while existing methods lack dedicated, semantically grounded evaluation of free-form novelty assessments. NovBench introduces a four-dimensional benchmark for comparing LLM-generated novelty evaluations, reporting analyses across general and specialized models under varied prompting conditions.

  • Problem

    Existing methods lack dedicated, reliable assessment of free-form, aspect-specific LLM-generated novelty evaluations, despite mounting pressure on peer review.

  • Method

    NovBench combines paper-introduction novelty descriptions with reviewer evaluations and scores generated assessments across four interpretable dimensions.

  • Results

    The study analyzes general and specialized LLM performance across all four dimensions under varying prompting conditions.

  • Takeaways & Limitations

    NovBench provides a useful reference for automated academic novelty assessment and LLM-based evaluation.

  • Takeaways & Limitations

    The benchmark draws on a limited set of NLP venues, which may restrict generalizability to broader research domains with different review formats and rubrics.

Abstract

from arXiv · show

Novelty is a core requirement in academic publishing and a central focus of peer review, yet the growing volume of submissions has placed increasing pressure on human reviewers. While large language models (LLMs), including those fine-tuned on peer review data, have shown promise in generating review comments, the absence of a dedicated benchmark has limited systematic evaluation of their ability to assess research novelty. To address this gap, we introduce NovBench, the first large-scale benchmark designed to evaluate LLMs' capability to generate novelty evaluations in support of human peer review. NovBench comprises 1,684 paper-review pairs from a leading NLP conference, including novelty descriptions extracted from paper introductions and corresponding expert-written novelty evaluations. We focus on both sources because the introduction provides a standardized and explicit articulation of novelty claims, while expert-written novelty evaluations constitute one of the current gold standards of human judgment. Furthermore, we propose a four-dimensional evaluation framework (including Relevance, Correctness, Coverage, and Clarity) to assess the quality of LLM-generated novelty evaluations. Extensive experiments on both general and specialized LLMs under different prompting strategies reveal that current models exhibit limited understanding of scientific novelty, and that fine--tuned models often suffer from instruction-following deficiencies. These findings underscore the need for targeted fine-tuning strategies that jointly improve novelty comprehension and instruction adherence.

1 Introduction

The introduction identifies increasing pressure on peer review and shortcomings in existing novelty-evaluation metrics, motivating a structured benchmark and interpretable, semantics-aware framework. It also positions systematic evaluation of general and specialized LLMs as a basis for understanding gaps from human judgment and improving AI-assisted peer review.

  • Motivation: The explosive growth of submissions and limited reviewer availability make robust novelty assessment increasingly difficult for peer review.Novelty assessment is described as a core peer-review function, but reviewer-resource constraints intensify the challenge.
  • Motivation: ROUGE, BLEU, BERTScore, and LLM-as-judge approaches inadequately assess the semantic adequacy and aspect-specific correctness of free-form novelty evaluations.The introduction characterizes these approaches as surface-level or non-transparent, limiting reliable evaluation.
  • Contributions: The paper introduces a structured novelty-evaluation benchmark pairing reviewers’ novelty assessments with novelty descriptions from paper introductions.This resource is presented as a foundation for future research on novelty evaluation.
  • Contributions: It proposes a four-dimensional, interpretable, semantics-aware framework for evaluating generated novelty assessments.The framework is presented as a contribution alongside the benchmark, with the supplied passage emphasizing interpretability and semantic awareness.
  • Contributions: The study systematically benchmarks general and specialized LLMs, analyzes divergence from human judgments, and identifies factors and behavioral patterns relevant to reliable AI-assisted peer review.The analysis is intended to guide future development of higher-quality and more interpretable novelty-review text.

2 Related Work

Prior work has advanced automated scholarly review and LLM-based novelty research, but existing evaluations remain focused on overall review performance, scores, or human judgment rather than dedicated assessment of generated novelty text.

  • Automated Scholarly Paper Review: Automated scholarly paper review initially emphasized paper-rating recommendation, while later studies fine-tuned pretrained models to generate paper reviews (Kang et al., 2018; Li et al., 2020; Wang et al., 2020; Yuan et al., 2022; Yuan and Liu, 2022).
  • LLM-Based Review Generation: LLM peer-review studies report meaningful feedback generation, but models often lack critical analysis despite their powerful text-generation capabilities (Liang et al., 2024; Du et al., 2024).
  • Research Gap: Despite broad efforts to improve review generation, LLM effectiveness on fine-grained aspects of papers, especially novelty, remains insufficiently addressed (Yu et al., 2024b; Gao et al., 2024; Idahl and Ahmadi, 2025; Weng et al., 2025; Zhu et al., 2025b; Chang et al., 2025).
  • Scientific Novelty: Recent work studies LLM-based paper-novelty assessment and novel-idea generation, recognizing novelty as a key measure of academic quality and contribution (Huang et al., 2025; Liu et al., 2025a; Lin et al., 2025; Liu et al., 2025b; Wu et al., 2025a,b; Tan et al., 2026; Shahid et al., 2025; Su et al.).
  • Evaluation Gap: Existing novelty evaluations largely use quantitative scores or human evaluation, leaving dedicated assessment of generated textual novelty evaluations underdeveloped and often relying on “LLM as a judge.”

3 NovBench

NovBench is a benchmark that combines author-stated novelty descriptions with human reviewer evaluations to assess LLM-generated novelty judgments. It defines a structured sentiment-based task and evaluates outputs using relevance, correctness, coverage, and clarity.

  • Dataset construction: The dataset pipeline extracts introduction novelty descriptions, identifies novelty-related review evaluations, and structures them by sentiment polarity.GPT-4o is prompted to remove redundant evaluations when multiple reviewers express overlapping judgments.
  • Dataset construction: NovBench automatically annotates each paper with introduction-based novelty descriptions and human reviewer novelty evaluations, capturing both author claims and independent judgments.The full benchmark differs from the manually annotated COLING 2020 subset, which primarily supports controlled model selection for novelty-description extraction.
  • Evaluation framework: NovBench evaluates generated assessments on Relevance, Correctness, Coverage, and Clarity using source alignment, reviewer sentiment agreement, expert-point coverage, and grounded fluency.Relevance uses average sentence-level Information Matching Scores; Correctness compares sentiment distributions; Coverage measures expert-identified novelty points; Clarity combines keyword grounding with elaboration and fluency.

4 Experiments

Experiments evaluate general-purpose and specialized LLMs on NovBench using deterministic inference and zero-shot, few-shot, and RAG prompting. Results indicate advantages for closed-source and specialized models, with performance shaped by model size and fine-tuning strategy.

  • Experimental Setup: The study evaluates 11 general-purpose and eight peer-review-specialized LLMs using greedy decoding with a 4096-token limit and zero-shot, few-shot, or RAG prompting.Closed-source models were accessed through official APIs, whereas open-source models were run locally from HuggingFace.
  • Results: Table 2 shows stronger performance for closed-source general LLMs, while specialized models generally outperform similarly sized general models across prompting settings.The specialized-model advantage depends on fine-tuning approaches, as illustrated by CycleReviewer-8B and SEA-S.
  • Results: Performance generally improves with model size, although exceptions among general-purpose LLMs suggest that larger models can over-interpret under strict evaluation constraints.The proposed explanation is that stronger reasoning and generation can cause distributional drift.
  • Human Evaluation: Human validation used 100 randomly selected samples in controlled pairwise judgments based on the paper’s introduction novelty description and human reviewer evaluation.Evaluators selected which of two model outputs was the higher-quality novelty evaluation.

5 Result Analysis

LLMs show only surface-level novelty understanding: they perform relatively well on relevance and clarity but struggle with fine-grained correctness, coverage, and instruction following. Performance varies by prompting, model specialization, contribution type, and reviewer disagreement, while remaining stable across generations, publication years, and controlled perturbations.

  • Prompt Tuning: Few-shot prompting improves Coverage and Correctness but reduces Relevance, whereas zero-shot prompting yields the highest Relevance score, only 3.6983.The results indicate a trade-off between leveraging human-evaluated examples and preserving alignment with explicitly stated novelty.
  • Model Comparison: Specialized models offer only marginal overall advantages, with CycleReviewer-70B and SEA models maintaining comparable Relevance while outperforming general-purpose models across specific training prompts.Better-performing specialized models also achieve higher Correctness, suggesting that fine-tuning helps them learn human expressive and structural patterns, though mixed sentiment limits performance.
  • Case Analysis: GPT-4o and SEA-S identify core innovation claims well, but they also exaggerate contributions, force negative judgments, introduce unsupported details, and produce templated analyses.Their positive evaluations are largely grounded in the methodologies and contributions explicitly stated in the source papers.
  • 5 Result Analysis: LLMs achieve strong Clarity but struggle with fine-grained Relevance and Coverage, often diverging from human reviewers’ assessments of novelty breadth.Clarity largely reflects information extraction rather than deeper novelty understanding, while retrieval augmentation and prompting alone do not resolve relevance limitations.
  • Instruction Following: Some peer-review-fine-tuned models exhibit severe instruction-following failures, especially Reviewer2, while larger models and those trained on noisy, inconsistent data appear better equipped to follow instructions.The specialized models’ deficiencies can substantially decrease performance and are not confined to a single prompt.
  • Additional Analyses: Performance remains stable across model generations, publication years, and controlled input perturbations, while models perform better on resource papers and align more with high-confidence reviews.These patterns argue against memorization or temporal leakage and show that evaluation difficulty varies by contribution type and reviewer disagreement.

6 Conclusion

The paper introduces NovBench, a controlled benchmark that systematically evaluates LLMs’ academic novelty assessments across four dimensions. It analyzes general and specialized LLMs under varied prompting conditions and proposes extending the benchmark across venues and domains.

  • 6 Conclusion: NovBench systematically evaluates LLMs’ ability to assess academic paper novelty using four quality dimensions in a controlled, homogeneous setting.The setting is designed to ensure reliability and isolate the novelty-assessment task.
  • 6 Conclusion: The study analyzes novelty evaluations from general and specialized LLMs across varying prompting conditions and all four evaluation dimensions.These analyses provide insights intended to guide future development of LLM-based novelty assessment.
  • 6 Conclusion: Future work will extend NovBench to additional venues through the same data-construction pipeline, enabling study of cross-venue and domain generalization.

Limitations

The study is limited by its reliance on introductions, potentially biased and narrow venue coverage, simple prompting, and incomplete modeling of novelty types and evaluation credibility. Despite these constraints, it provides a useful reference for automated academic novelty assessment and LLM-based evaluation.

  • Using only paper introductions may omit detailed content needed to fully support novelty evaluations.The introduction contains primary novelty claims, but the study does not analyze full paper text.
  • Accepted-paper-heavy COLING and EMNLP data may introduce selection bias, while limited NLP venue coverage restricts generalizability across disciplines and review formats.ICLR and NeurIPS differ in review formats, scoring rubrics, and interdisciplinary scope.
  • The study uses simple prompting, omits advanced prompting and multi-agent architectures, and does not incorporate numerical confidence scores despite concerns about reviewer-comment credibility.These omissions limit analysis of more sophisticated generation strategies and calibrated evaluation.
  • The analysis does not distinguish novelty types, and more robust evaluation methods are still needed despite the effectiveness of the proposed metrics.EMNLP emphasizes methodological novelty, but the study does not model different novelty categories.
  • Future work should develop finer-grained taxonomies, analyze hallucination patterns, and investigate model aggregation through ensembling or multi-agent methods.These directions are proposed within the study’s evaluation framework.
  • Despite these limitations, the study provides a useful reference for automated academic novelty assessment and LLM-based evaluation.

Ethics Statement · A Supplement of Automatic Extraction of Novelty Descriptions

The study uses publicly available, non-identifiable peer-review reports and frames LLMs as controlled assistants for novelty assessment rather than replacements for human reviewers. A supplement manually annotates novelty descriptions and finds context-prompted GPT-5 performs best among evaluated extraction settings.

  • Ethics Statement: All data are openly available peer-review reports without additional personally identifiable information, and the study collected no new personal data.The authors state that the analysis poses no additional privacy-leakage or harm risk to authors or reviewers.
  • Ethics Statement: The work evaluates LLM assistance for the specific task of novelty analysis under controlled, transparent conditions rather than promoting automated peer review as a replacement for experts.The intended benefits are potentially reducing reviewer workload and providing complementary perspectives.
  • Ethics Statement: The authors acknowledge risks of over-reliance, bias amplification, and misuse, positioning the study as empirical evidence for ethical discussion rather than advocacy for autonomous reviewers.This limitation qualifies the intended role of LLMs in peer review.
  • A Supplement of Automatic Extraction of Novelty Descriptions: Manual annotation covered 87 COLING 2020 papers, identifying 533 novelty-description sentences among 2,300 total sentences.Two experienced reviewers performed the annotation, achieving Cohen’s κ inter-rater agreement of 0.831; extraction was framed as binary classification.
  • A Supplement of Automatic Extraction of Novelty Descriptions: Context prompting yielded the best extraction performance across all models, with GPT-5 reaching Accuracy 0.89 and Macro F1 score 0.84.The authors consequently selected context-prompted GPT-5 for subsequent novelty-description extraction.
  • A Supplement of Automatic Extraction of Novelty Descriptions: The extraction benchmark compared zero-shot, few-shot, step-by-step, and in-context learning prompts across various LLMs.The zero-shot setup is illustrated in Figure 5, while the reported results appear in Figure 9.

B Supplement of Automatic Extraction of Novelty Evaluations · C Supplement of Sentiment-Based Normalization of Novelty Evaluations

The supplements describe automatic extraction of novelty evaluations and sentiment-based normalization, enabling binary classification testing and consistent comparison of human-written and LLM-generated feedback. The extraction benchmark uses annotated novelty and non-novelty review comments, while normalization removes redundancy and organizes feedback by sentiment polarity.

  • B Supplement of Automatic Extraction of Novelty Evaluations: The extraction resource contains 493 novelty-related comments from Lu et al. (Lu et al., 2025) and 500 randomly selected non-novelty evaluations.
  • B Supplement of Automatic Extraction of Novelty Evaluations: The task is framed as binary classification, requiring models to judge whether each review sentence expresses a novelty evaluation.
  • B Supplement of Automatic Extraction of Novelty Evaluations: A few-shot prompt supports extraction of novelty descriptions from peer-review text.
  • C Supplement of Sentiment-Based Normalization of Novelty Evaluations: GPT-4o normalization deduplicates semantically similar comments, consolidates them into concise statements, and categorizes them by sentiment polarity.
  • C Supplement of Sentiment-Based Normalization of Novelty Evaluations: This normalization produces consistent, non-redundant evaluative statements for more reliable automatic evaluation of LLM-generated novelty descriptions.

D Supplement of Agreement Evaluation · E Experiment Implementation Details · F Supplemental Analysis of Instruction-Following Deficiencies in Specialized Review Generation Models

The appendices document human agreement evaluation, NovBench implementation details, and instruction-following failures in specialized review models. They use expert pairwise judgments and majority-vote agreement metrics, evaluate multiple prompting strategies, and identify repetitive or empty outputs after fine-tuning.

  • D Supplement of Agreement Evaluation: Four NLP experts independently compare which of two models produces the higher-quality novelty evaluation for each sample.The evaluators include two Ph.D. students, an Associate Professor, and a Lecturer; the appendix also supplies detailed instructions, examples, and guidelines.
  • D Supplement of Agreement Evaluation: Human preference is aggregated by majority vote, excluding samples without a strict majority from agreement computation.The agreement score measures whether an automatic metric selects the same preferred model as the aggregated human judgment.
  • E Experiment Implementation Details: NovBench testing evaluates general and specialized LLMs with zero-shot, few-shot, and retrieval-augmented-generation prompting strategies.The RAG condition supplies extracted novelty descriptions together with retrieved context from relevant academic papers.
  • E Experiment Implementation Details: RAG retrieves the five most relevant titles and abstracts for each NovBench paper from ACL, EMNLP, and NAACL proceedings published during 2019–2022.Retrieval uses each NovBench paper’s abstract as the query against a locally stored ACL Anthology database.
  • E Experiment Implementation Details: The appendix details eight fine-tuned LLMs trained on peer-review data spanning ICLR, NLPeer, NeurIPS, and multiple research fields.The models use different backbones, including Mistral-Nemo-12B, Qwen2.5-Instruct-72B, Phi-4, and Mistral-7B-Instruct-v0.2.
  • E Experiment Implementation Details: Inference runs on A100 80GB or H100 80GB GPUs, with larger models distributed across two GPUs.Models from 8B through 32B generally use one A100, 70B-scale models use two A100s, and gpt-oss-120B uses two H100s.
  • F Supplemental Analysis of Instruction-Following Deficiencies in Specialized Review Generation Models: CycleReviewer-8b generates repetitive evaluations, while DeepReviewer produces null or empty evaluations, revealing severe operational failures after peer-review fine-tuning.These failures are documented in Figure 19 and extend beyond the deficiencies reported in Section 5.2.

G Case Studies Comparing Human and LLM-Generated Novelty Evaluations · H Additional Analyses

The case studies examine how GPT-4o and SEA-S generate novelty evaluations relative to human reviewers and extracted novelty descriptions. The supplied passages do not provide substantive content for the additional analyses section.

  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: Five case studies were selected for analysis.
  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: Each case study includes a novelty description extracted from the paper introduction.
  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: Figures 9 and 10 report the performance of various LLMs under different prompts for novelty description and novelty evaluation extraction, respectively.
  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: Each case study includes the corresponding novelty evaluation provided by a human reviewer.
  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: Each case study also includes novelty evaluations generated by GPT-4o and SEA-S.
  • G Case Studies Comparing Human and LLM-Generated Novelty Evaluations: The case-study examples are illustrated in Figures 20, 21, 22, 23, and 24.

H.1 Memorization and Temporal Analysis · H.2 Analysis by Paper Type · H.3 Alignment under Reviewer Disagreement

Across temporal, paper-type, and reviewer-disagreement analyses, model behavior is not primarily explained by memorization or temporal leakage, while evaluation difficulty and reviewer alignment vary systematically with paper characteristics and confidence. Models perform better on resource papers than methodological papers and align more closely with high-confidence reviews.

  • H.1 Memorization and Temporal Analysis: GPT-3.5 performs competitively among general LLMs, suggesting novelty-evaluation performance is not primarily driven by access to more recent training data.GPT-3.5 was released before EMNLP 2023 and was evaluated under the same prompting settings as the other models.
  • H.1 Memorization and Temporal Analysis: Performance shows no substantial differences between COLING 2020 and EMNLP 2023 datasets, providing no evidence of a publication-year effect.The cross-year comparison is reported in Table 4.
  • H.1 Memorization and Temporal Analysis: Models express uncertainty when asked to continue review sentences, suggesting that they do not reproduce exact memorized continuations.The test prompts models to continue review sentences as a direct probe for verbatim memorization.
  • H.1 Memorization and Temporal Analysis: Performance remains largely stable after novelty descriptions are paraphrased or partially deleted, while the combined evidence attributes behavior to intrinsic evaluation difficulty rather than memorization or temporal leakage.The perturbation experiments modify inputs through “change” and “del” conditions, and the overall conclusion combines these results with the temporal and continuation tests.
  • H.2 Analysis by Paper Type: Models consistently perform better on resource papers than methodological papers, likely because resource contributions are more explicit whereas methodological novelty requires more nuanced reasoning.Tables 6 and 7 report results for methodological and resource papers, respectively.
  • H.2 Analysis by Paper Type: Model rankings remain broadly consistent across paper categories despite paper characteristics affecting evaluation difficulty.The paper-type analysis groups benchmark results by methodological and resource papers using representative models from the main results.
  • H.3 Alignment under Reviewer Disagreement: Models consistently achieve higher similarity to high-confidence than low-confidence reviews when reviewer confidence differs substantially, indicating non-arbitrary alignment under disagreement.The analysis uses samples with a confidence gap ≥3 and compares semantic similarity to high- and low-confidence review groups.
Loading 2604.11543v1…