Source-linked AI summary
On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal, Ian Wu, Viktor Zaverkin, Spase Petkoski, Daniel R. Schrider, Ilija Dukovski, Francesco Santini, Biljana Mitreska, Yong Jeong, Kyeongha Kwon, Young Min Sim, Dragana Manasova, Arthur Porto, Biljana Mojsoska, Makoto Takamoto, Marko Shuntov, Ruoqi Liu, Hyunjoo Jenny Lee, Niyazi Ulas Dinç, Yehhyun Jo, Sunkyu Han, Chungwoo Lee, Huishan Li, Esther H. R. Tsai, Ergun Simsek, Khushboo Shafi, Yeonseung Chung, Jihye Park, Aleksandar Shulevski, Henrik Christiansen, Yoosang Son, Elly Knight, Amanda Montoya, Jeongyoun Ahn, Christian Langkammer, Heera Moon, Changwon Yoon, Nikola Stikov, Mooseok Jang, Edward Choi, Junhan Kim, Yeon Sik Jung, Woo Youn Kim, Jae Kyoung Kim, Ishraq Md Anjum, Hyun Uk Kim, Drew Bridges, Carolin Lawrence, Xiang Yue, Alice Oh, Akari Asai, Sean Welleck, Graham Neubig
TL;DR
AI reviewers are increasingly used in peer review, but verdict-level evaluations do not show whether their individual criticisms are correct, significant, or sufficiently evidenced. This study uses expert annotations of human and AI review items and finds that GPT-5.2 exceeds top-rated human reviewers on composite quality, while current AI reviewers remain complements rather than substitutes for humans.
Problem
Existing evaluations mainly compare aggregate scores or verdicts, leaving the quality of individual AI-generated criticisms insufficiently characterized.
Method
Forty-five scientists evaluated 2,960 atomic criticisms from human and AI reviews of 82 Nature-family papers for correctness, significance, and evidence sufficiency.
Results
GPT-5.2 achieves a 60.0% fully-positive rate versus 48.2% for the top-rated human reviewer, while all three AI reviewers exceed the lowest-rated human.
Takeaways & Limitations
Current AI reviewers can complement but should not replace human reviewers in scientific peer review.
Takeaways & Limitations
AI reviewers show severity miscalibration from limited field-specific knowledge and factual errors from weak long-context management across multiple files.
Abstract
from arXiv · showhide
With the advancement of AI capabilities, AI reviewers are beginning to be deployed in scientific peer review, yet their capability and credibility remain in question: many scientists simply view them as probabilistic systems without the expertise to evaluate research, while other researchers are more optimistic about their readiness without concrete evidence. Understanding what AI reviewers do well, where they fall short, and what challenges remain is essential. However, existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits. In this paper, we close this gap through a large-scale expert annotation study, in which 45 domain scientists in Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual criticisms (each targeting one specific aspect of a paper) from human-written and AI-generated reviews of 82 Nature-family papers on correctness, significance, and sufficiency of evidence. On a composite of all three dimensions, a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009), while all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension. AI reviewers' accurate criticisms are also more often rated significant and well-evidenced, and surface a distinct 26% of issues no human raises. However, AI reviewers overlap far more than humans do (21% vs. 3% for cross-reviewer pairs), and exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues. Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers.
1 Introduction
Peer review is under unprecedented scaling pressure as scientific output grows while qualified reviewers remain limited, exposing the limits of verdict-level evaluations of AI reviewers. This study addresses the gap with expert annotation of atomic criticisms from human and AI reviews, and introduces resources for continued evaluation.
- Motivation: Scientific output is rising at a historic rate, while the pool of qualified reviewers faces unprecedented scaling pressure.Peer review supports credibility, rigor, error detection, methodological improvement, and judgments about reliable findings.
- Evaluation gap: Prior evaluations chiefly compare aggregate scores, acceptance recommendations, or holistic ratings, which do not establish whether AI criticisms are useful.The study shifts evaluation from verdict-level agreement to individual review items targeting specific paper aspects.
- Study design: 45 domain scientists spent 469 hours rating 2,960 atomic criticisms from reviews of 82 Nature-family papers on correctness, significance, and evidence sufficiency.Annotators spanned Physical, Biological, and Health Sciences and also provided free-form qualitative feedback.
- Resources: PEERREVIEW BENCH applies the expert evaluation criteria automatically, enabling continued tracking of AI reviewer quality without repeating costly expert annotation.The benchmark exposes substantial headroom: GPT-5.4, DeepSeek-V4-Pro, and Claude-Opus-4.7 achieve 41.4%, 48.5%, and 50.5% F1, respectively.
- Resources: The study also releases CMU PAPER REVIEWER as an open-source resource for the AI and scientific communities.The supplied passage identifies the resource but does not provide further details about its functionality.
2 Preliminaries: Expert annotation study design and experimental setup
The study evaluates AI and human peer reviews at the level of atomic criticisms rather than aggregate verdicts, using expert annotations of correctness, significance, and evidence sufficiency. It compares reviews of 82 Nature-family papers across three scientific domains, with 45 domain scientists contributing 469 hours of annotation.
- Study design: Review items are single atomic criticisms directed at one specific aspect of a paper, providing the study’s unit of analysis instead of aggregate review scores or verdicts.This decomposition enables direct evaluation of individual criticisms rather than only comparing overall review outputs.
- Study design: Each review item is rated on correctness, three-level significance, and evidence sufficiency, with later dimensions assessed only when the preceding conditions are met.Correctness is binary; significance ranges from Significant to Not Significant; evidence sufficiency is binary and requires a correct, at least marginally significant criticism.
- Dataset and reviewers: 82 Nature-family papers span 27 subject categories across Physical, Biological, and Health Sciences, including 38, 30, and 14 papers in those fields, respectively.The dataset includes 73 Nature Communications papers and papers from six additional Nature-family journals, published between 10 January 2020 and 27 October 2025.
- Dataset and reviewers: Three AI reviewers—GPT-5.2, Claude Opus 4.5, and Gemini 3.0 Pro—operate as agents with access to paper source files and tools, producing up to five structured review items per paper.Accessible materials include the main text, supplementary materials, figures, and submitted code; the agents use the six Nature peer-review evaluation criteria.
- Expert annotation: 45 domain scientists from 25 institutions produced 109 meta-reviews across the 82 papers, averaging 2.42 papers per scientist and totaling 469 hours of expert annotation.The pool includes faculty members, industrial or national-laboratory research scientists, postdoctoral researchers, and Ph.D. students.
- Expert annotation: 908 review items from 27 papers received independent second annotations, showing almost-perfect agreement for correctness and evidence sufficiency but moderate agreement for the three-level significance scale.The study reports Gwet’s AC1 alongside raw percent agreement and Cohen’s κ because annotation class distributions are highly skewed.
3 In which aspects are AI reviewers better or worse than human reviewers?
AI reviewers surface more significant criticisms than top-rated human reviewers but are less factually correct. On aggregate review-item quality, all three exceed the lowest-rated human, while only GPT-5.2 exceeds the top-rated human.
- Dimension-level comparison: 6–10 percentage points lower: all three AI reviewers fall below the Top-Rated Human’s 92.3% correctness.GPT-5.2 reaches 86.2%, Claude Opus 4.5 83.7%, and Gemini 3.0 Pro 81.9%.
- Dimension-level comparison: 1.39: among correct items, AI reviewers’ mean significance score exceeds the Top-Rated Human’s score on the 0–2 scale.Rank-biserial correlations are +0.49 for GPT-5.2, +0.30 for Claude Opus 4.5, and +0.42 for Gemini 3.0 Pro; all p ≤ .028.
- Aggregate review-item quality: All three AI reviewers exceed the lowest-rated human on fully positive review-item quality, but only GPT-5.2 exceeds the top-rated human.Fully positive means an item is correct, maximally significant on the 0–2 scale, and sufficiently evidenced.
- Holistic expert judgments: 48.6%: GPT-5.2 matches or exceeds the Top-Rated Human on expert-judged overall quality across papers.GPT-5.2 matches or exceeds the Lowest-Rated Human on 73.4% of papers.
- Overall takeaway: Tool-based access to the full paper source, code, and external literature accompanies a consistent pattern of more significant issues but lower correctness.The pattern is reported across dimension-level, aggregate item-quality, and expert-judged paper-level analyses.
4 To what extent do AI reviews overlap with human reviews?
AI reviewers substantially overlap with one another but preserve diversity relative to humans, while surfacing largely valid criticisms that humans miss. Their coverage is strongest for paper targets and weaker for matching specific criticisms, supporting augmentation rather than replacement of human panels.
- Overlap definition: The study classifies review-item pairs as similar only when they share both the same target and the same criticism.Pairs can additionally differ in evidence, producing four mutually exclusive categories from different targets through near-paraphrases.
- AI–human overlap: 74.0% of AI-raised items have a similar human counterpart, while 26.0% are uncovered by human reviewers.Similarity requires the same target and the same criticism, regardless of whether the evidence is identical.
- AI–human overlap: 64.2% of human concerns share a target with an AI reviewer, versus 43.8% for human–human replacement, but same-criticism coverage is 27.1% versus 25.8%.The AI panel therefore identifies most of the same targets as humans while providing more distinct feedback on those targets.
- Reviewer diversity: 20.9% of AI–AI item pairs share a target and criticism, compared with 3.4% of human–human pairs.AI reviewers overlap with one another roughly six times more than humans do, whereas AI–human overlap is 5.1%, only slightly above the human–human baseline.
- Reviewer diversity: 96.6% of human–human item pairs are nonmatching, comprising 8.4% with different criticisms about the same target and 88.3% addressing different targets.Only 25.8% of one human reviewer’s items have a same-criticism counterpart from another human reviewer.
- AI-only criticisms: 81.8% of uncovered AI items are correct and 93.5% are well-evidenced, indicating that most AI-only criticisms are valid additions.Uncovered items show only a modest reduction in the highest-significance rate compared with matched AI items.
5 What are the concrete strengths and weaknesses of AI reviewers?
AI reviewers provide genuine value on rigor- and code-heavy scrutiny, but their dominant limitation is contextual miscalibration against field-specific norms rather than factual emptiness. They also sometimes falsely identify information as missing because they fail to manage long contexts across manuscript files.
- Weaknesses: 189/260 weakness comments concern five recurring weaknesses, chiefly missing field norms and over-harsh or out-of-scope demands.These weaknesses describe reviews as contextually uncalibrated: technical content is often correct under discipline-neutral standards, but severity is misjudged.
- Weaknesses: 54 comments identified missing community or field norms, where accepted subfield practices are incorrectly treated as methodological gaps.The critique may be technically accurate under generic reproducibility standards, while lacking knowledge of what the field considers normal.
- Weaknesses: 37 comments involved AI reviewers claiming that manuscripts lacked information they actually provided elsewhere, reflecting limited long-context management across files.Experts documented cases where procedures appeared in supplementary material or other manuscript sections, making the critique’s factual premise wrong.
- Strengths: 115/132 strength comments fell within four categories, centered on statistical and methodological rigor and routine, labor-intensive scrutiny.AI reviewers are particularly diligent when tasks require reading code, checking statistical assumptions, or cross-referencing specialized literature.
- Strengths: AI reviewers can uncover implementation bugs, data leakage, methodology mismatches, and ambiguities by directly cross-checking manuscript claims against source code.Domain experts specifically endorsed this code-based scrutiny as valuable because human reviewers typically avoid it when it is too time-consuming.
6 Tools for improving and using AI reviewers
The section develops tools to reduce the cost of evaluating AI reviewers and to support manuscript feedback before submission. It introduces PEERREVIEW BENCH and the open-source CMU PAPER REVIEWER, including concrete mitigations for documented review weaknesses.
- Benchmarking AI reviewers: PEERREVIEW BENCH is a 78-paper benchmark that evaluates AI reviewers using precision and recall.Precision measures the fraction of AI-raised items judged fully positive; recall measures the fraction of fully positive human review items identified by the AI reviewer.
- Benchmarking AI reviewers: The benchmark reports Precision, Recall, and F1 score for publicly available AI reviewer platforms.Review items are operationalized differently for Stanford Agentic Reviewer and OpenAIReview, using weakness bullets and feedback cards, respectively.
- CMU PAPER REVIEWER: The CMU PAPER REVIEWER is an open-source platform for authors, students, and researchers seeking detailed manuscript feedback before submission.It is built on the pipeline used in the expert annotation study.
- CMU PAPER REVIEWER: The platform pairs every review item with a concrete patch suggestion to make vague or non-actionable critiques more useful.Suggestions may be proposed manuscript edits or runnable code patches when source is provided.
- CMU PAPER REVIEWER: Severity ratings are grounded in the manuscript’s stated limitations, while an interactive debate mode addresses over-harsh or out-of-scope demands.These features implement mitigations for weakness patterns documented earlier in the paper.
7 Conclusion · Table of Contents in Appendix · H Recommendation for journal and conference organizers: Panel composition analysis
The study uses expert annotations to characterize AI reviewers’ strengths and weaknesses, finding them competitive with top-rated human reviewers while identifying correctness and calibration as priorities for improvement. It proposes PEERREVIEW BENCH for tracking progress and lists methodology and panel-composition analyses in Appendix H.
- 7 Conclusion: 45 domain scientists evaluated 2,960 review items over 469 hours across 82 Nature-family papers, comparing AI- and human-generated reviews.The study assesses individual review items rather than only reviewer-level verdict agreement.
- 7 Conclusion: AI reviewers were competitive with Nature’s top-rated official reviewers on the composite of correctness, significance, and evidence sufficiency.The composite combines three expert-rated dimensions of review quality.
- 7 Conclusion: Developers should prioritize closing the correctness gap and improving criticism calibration.Calibration concerns whether criticism is warranted rather than inflated.
- 7 Conclusion: Concrete development directions include modeling subfield-specific norms, improving long-context management in LLM agents, and calibrating criticism to expert judgment.These directions target recurring weaknesses identified in the study.
- 7 Conclusion: PEERREVIEW BENCH offers a testbed for tracking progress across future generations of AI reviewers.The benchmark is intended to evaluate improvement on correctness and criticism calibration.
- H Recommendation for journal and conference organizers: Panel composition analysis: The panel-composition analysis is presented under Appendix H’s recommendation section for journal and conference organizers.The supplied table-of-contents passage identifies this section but provides no substantive findings.
A Related Work · B Extended: Expert annotation study design and experimental setup · B.1 Subject category breakdown
Prior AI-reviewer evaluations largely rely on review-level, human-referenced outcomes, whereas this study evaluates individual AI and human criticisms bidirectionally across correctness, significance, and evidentiary sufficiency using domain experts. The accompanying dataset comprises 82 Nature Communications papers categorized under the journal’s subject taxonomy.
- A Related Work: Prior systems span structured comment generation, multi-agent review, hierarchical decomposition, specialized review models, and reviewer-agent outcome prediction, but their evaluations generally omit at least one of the study’s four properties.Examples include ReviewRobot, TreeReview, OpenReviewer, GAR, ReviewerToo, ReviewEval, REVIEWSCORE, and FLAWS.
- A Related Work: Existing AI-reviewer evaluations mostly measure review-level score correlation, decision alignment, or text similarity against human reviews treated as a gold standard.Some newer frameworks add finer-grained factual verification or actionability, but do not jointly provide the present design’s full set of properties.
- A Related Work: The study evaluates every individual AI- and human-written review item with domain scientists from each paper’s field on correctness, significance, and sufficiency of evidence.This design is explicitly presented as a prior gap-filling evaluation framework.
- A Related Work: Four design properties are jointly absent from prior frameworks: bidirectional annotation, per-comment analysis, multi-axis quality decomposition, and paper-specific domain-expert annotation.These properties distinguish the study from evaluations centered on outcomes, full reviews, single dimensions, or non-domain judges.
- A Related Work: Related benchmarks provide partial precedents: REVIEWSCORE has per-comment annotations, while ReviewEval provides multi-axis scoring and FLAWS evaluates individual errors.Their limitations include non-domain annotators, one-directional or single-axis evaluation, and synthetic or automated ground truth.
- B Extended: Expert annotation study design and experimental setup: The appendix supplies procedural and statistical details supporting Section 2 in the order of the main text’s elements.It serves as an extended account of the expert annotation study design and experimental setup.
- B.1 Subject category breakdown: The expert annotation study uses a dataset of 82 papers, with subject categories defined according to the Nature Communications subject taxonomy.Table 11 reports the full category breakdown, but the supplied passage does not include the category counts.
B.2 Evaluation criteria for reviewing a paper … C.2 Complete pairwise paired comparisons
The study defines prioritized paper-level criteria and a cascading item-level rubric, then applies them through a controlled review-generation and expert-annotation workflow. Supplementary analyses detail agreement, item-level rates, and complete paired comparisons across human and AI reviewers.
- B.2 Evaluation criteria for reviewing a paper; B.6 AI reviewer prompt: Six Nature criteria prioritize validity, conclusions, originality and significance, data and methodology, statistics and uncertainties, and clarity and context.AI reviewers receive these criteria verbatim, and earlier criteria take precedence when selecting up to five criticisms.
- B.3 Evaluation criteria for reviewing a review; B.8 Annotation guidelines: Each review item is assessed through cascading judgments of correctness, significance, and sufficiency of evidence.Significance is evaluated only for correct items, and evidence sufficiency only for correct items that are at least marginally significant.
- B.4 Processing official peer review files: Human-review comparisons retain first-round comments from the first three reviewers, excluding editor decision letters and author rebuttals.This cap creates a consistent comparison between three human and three AI reviewers.
- B.5 AI reviewer configuration details: Three autonomous agents—GPT-5.2, Claude Opus 4.5, and Gemini 3.0 Pro—use shared OpenHands settings, paper-file tools, and restricted Tavily web search.Search excludes nature.com, researchsquare.com, springer.com, and springerlink.com to prevent retrieval of benchmark papers or existing review reports.
- B.6 AI reviewer prompt: AI reviews contain at most five atomic criticisms ordered by significance, with required claims, evaluation criteria, evidence, comments, and a citation list.The prompt instructs agents to inspect papers, supplementary files, images, and code, retrieve relevant literature, and verify criticisms before inclusion.
- B.7 Domain scientist recruitment: 45 domain scientists were recruited through a three-stage process conducted between September 2025 and April 2026.Candidates identified papers they were qualified to meta-review before receiving the corresponding AI reviews.
- B.8 Annotation guidelines: The annotation task decomposes reviews into atomic items and uses expert judgments to compare human and AI review quality and build PEERREVIEW BENCH.Annotators also complete paper-level comparisons of best and worst human reviews, AI matches, and issues uniquely noticed by an AI reviewer.
- B.9 Inter-annotator agreement details; C.1 Item-level descriptive statistics; C.2 Complete pairwise paired comparisons: 84.7% vs. 87.6% raw agreement on correctness, 60.2% vs. 59.3% on significance, and 94.3% vs. 83.9% on evidence show broadly stable annotation agreement.Agreement is consistently high for correctness and evidence sufficiency and moderate for significance across human and AI items.
C.3 Generalized linear mixed-effects model robustness analysis · D Extended: To what extent do AI reviewers overlap with human reviewers? · D.1 Detailed similarity breakdown
A paper-level random-intercept GLMM reproduces the paired analysis’s directional conclusions across correctness, significance, and evidence sufficiency. The extended similarity analysis compares six reviews per paper across four item-pair categories, revealing greater AI convergence than human overlap in key between-reviewer comparisons.
- C.3 Generalized linear mixed-effects model robustness analysis: A Bayesian binomial GLMM models item-level outcomes with a paper-level random intercept, explicitly accounting for items nested within papers.The model complements paper-level aggregation and can provide greater statistical power when within-paper variance is small.
- C.3 Generalized linear mixed-effects model robustness analysis: The model uses a common log-odds scale across three metrics, treating the ordinal significance outcome at the highest cut point, P(Y = 2).Reviewer indicators are dummy-coded against a reference category, and coefficients represent differences in outcome log-odds relative to that reference.
- C.3 Generalized linear mixed-effects model robustness analysis: Every AI reviewer has significantly lower correctness and significantly higher significance than the Top-Rated Human reviewer in the GLMM.GPT-5.2 and Claude Opus 4.5 also have significantly higher evidence sufficiency, whereas Gemini 3.0 Pro is indistinguishable on that dimension.
- C.3 Generalized linear mixed-effects model robustness analysis: The GLMM reaches the same directional conclusions as the paper-level paired analysis on all three dimensions, with some borderline contrasts attaining stronger significance.This supports the robustness of the main comparative findings under an item-level hierarchical model.
- D Extended: To what extent do AI reviewers overlap with human reviewers?: Each paper contributes six reviews—three human and three AI—and review-item pairs are classified into four similarity categories across human–human, AI–AI, and human–AI comparisons.The analysis evaluates every review item in one review against every item in another to quantify overlap.
- D.1 Detailed similarity breakdown: Table 18 reports full four-category similarity distributions with within-reviewer and between-reviewer splits for the two same-group pair types.The extended breakdown adds same-reviewer and same-model rows, capturing item diversity within a single review.
- D.1 Detailed similarity breakdown: 13.5% is the same-model AI–AI within-reviewer similarity rate, compared with 20.9% for different-model AI–AI between-reviewer pairs.For human reviewers, within-reviewer similarity is 6.0% versus 3.4% between different reviewers; thus, AI models converge more across models than within one model’s review.
D.2 Similarity judge calibration and selection … E.1.6 W8: Citing evidence that appeared after the preprint
The paper calibrates and corrects automated similarity judgments, while documenting recurring AI-reviewer weaknesses ranging from field-misalibrated criticism and missed manuscript content to redundancy, poor actionability, and anachronistic standards. These findings show that AI criticisms can be substantively valuable but require contextual, temporal, and human oversight.
- D.2 Similarity judge calibration and selection: 87.1% sensitivity and 96.8% specificity were estimated for GPT-5.4 on the 164-pair calibration set and used for Rogan–Gladen prevalence correction.The correction shifts similar categories downward and not-similar categories upward, especially for H-H and H-A pairs with low raw similarity rates.
- D.3 Raw versus Rogan-Gladen-corrected comparison: The correction assumes constant judge error rates across 65,704 pairs and transfers AI-calibrated rates to Human–Human comparisons, an untested assumption.The calibration set contains no Human–Human pairs, so this transfer is plausible but unverified.
- E.1.1 W1: Missing community / field norms: AI reviewers often identify technically valid concerns but misjudge their severity because they lack subfield norms or demand work beyond the paper’s scope and feasible revision budget.Examples include treating R2 = 0.36 as weak in observational ecology and requesting patient-derived atlases or substantially expanded experiments.
- E Extended: What are the concrete strength and weaknesses of AI reviewers: Across these cases, AI criticisms are often accurate in content but require human calibration of field norms, feasibility, document coverage, reviewer diversity, and submission-era standards.This distinction separates useful methodological observations from over-severe, redundant, non-actionable, or temporally inappropriate recommendations.
- E.1.3 W3: Paper explicitly states X, AI says missing: AI reviewers sometimes call information missing when it appears elsewhere in the manuscript or supplements, reflecting failures to manage long, multi-file contexts.The errors range from missing unconventional headings to overlooking complete calibration procedures and partial external benchmarking.
- E.1.4 W4: Redundancy across the three AI reviewers: n = 28: The three AI reviewers frequently converge on the same criticism, so additional AI reviews often contribute less distinct information than additional human reviews.The overlap can concern both methodological comparisons and shared interpretations of assumptions, differing mainly in phrasing or emphasis.
- E.1.5 W5: Vague, verbose, or without actionable recommendation: n = 24: AI reviews can be vague, overly verbose, or incomplete about the concrete revision authors should make.Experts may endorse the underlying concern while criticizing its length, placement, or failure to specify whether analyses should be rerun or claims narrowed.
- E.1.6 W8: Citing evidence that appeared after the preprint: AI reviewers can apply later evidence or subsequently established practices to earlier papers, incorrectly judging historically defensible claims while still suggesting narrower wording.The examples involve post-submission DeepSolid evidence and modern fine-tuning norms applied to an early transformer-based polymer study.
E.1.7 S1: Statistical and methodological rigor … W16: Cannot analyze figures — only text — n = 1
Across the reviewed examples, AI reviewers showed strengths in statistical and methodological rigor, source-code inspection, and domain-specific technical depth, while recurring weaknesses reflected missing field norms and overly harsh or out-of-scope criticism. Their strongest findings often exposed issues that human reviewers missed, but expert comments also identified limitations in contextual judgment and prioritization.
- E.1.7 S1: Statistical and methodological rigor: 45 experts identified statistical and methodological rigor as an AI reviewer strength, including scrutiny of independence assumptions, significance tests, validation splits, and uncertainty quantification.These critiques were treated as legitimate additional scrutiny that neither human reviewers nor, in some cases, authors had addressed.
- E.1.7 S1: Statistical and methodological rigor: The AI reviewer detected test-set contamination in SHS27K PPI evaluation because code used test metrics for model selection and early stopping without a validation split.The expert confirmed that the trainer had only train/test splits and used test metrics under the misleading name best valid f1.
- E.1.7 S1: Statistical and methodological rigor: AI reviewers also raised specialized statistical and evidential concerns, including underpowered K-S tests on approximately 10–22 heatwaves and 8–24 cold snaps and unsupported qualitative imaging claims.Experts extended the K-S critique to correlated regional observations and endorsed quantifying claims such as “a great deal” rather than relying on subjective language.
- E.1.8 S2: Inspecting the submitted source code: Source-code inspection enabled AI reviewers to verify data leakage and expose implementation contradictions that manuscript-only review could miss.Examples included test-set pseudo-interactions, test-metric model selection, and a 400× discrepancy between an 800 Hz claim and an approximately 2 Hz implemented sampling rate.
- E.1.9 S3: Domain-specific technical depth: 27 experts identified domain-specific technical depth as an AI reviewer strength, with correct critiques of optical-field scope, stereocontrol evidence, and quantitative s-SNOM modeling.Experts confirmed that the critiques were technically meaningful while sometimes narrowing their implications, such as treating improved PDM modeling as reducing error limits rather than invalidating main findings.
- W1: Missing community / field norms — n = 54: 54 cases involved missing community or field norms, where experts judged technically valid AI criticisms as impractical, benchmark-wide, acceptable for proof-of-concept work, or insufficiently contextualized.Examples included limited observational data, engineering rather than clinical novelty, accepted methods, and field-specific standards.
- W2: Over-harsh / out-of-scope / unrealistic — n = 46: 46 cases were categorized as over-harsh, out-of-scope, or unrealistic, including criticisms experts considered marginal, irrelevant to the paper’s focus, or valid architectural choices.Human reviewers sometimes ignored these issues because they did not materially affect the paper’s central contribution.
W unspecified: Residual: AI judged Not Correct without specific reason — n = 10 … S6: Big-picture / counter-narrative synthesis — n = 7
Across the reviewed sections, AI reviewers surfaced numerous technical, code-level, niche-domain, consistency, reproducibility, and big-picture issues, including several missed by human reviewers. However, some assessments were unsupported or misunderstood the paper’s central contribution.
- W unspecified: Residual: AI judged Not Correct without specific reason — n = 10: 10 criticisms were judged incorrect without a specific reason, including claims that reviews were weak or missed the manuscript’s main point.One assessment said the reviewer lacked a fundamental understanding of the key discovery and missed the paper’s main point.
- S1: Statistical / methodology rigor — n = 45: AI reviewers identified statistical and methodological flaws that human reviewers missed, including test-set model selection, correlated observations, data leakage, invalid equations, and missing uncertainty analyses.They also raised issues involving baselines, variance estimates, multiple comparisons, imputation, causality, and spatial autocorrelation.
- S2: Code reading — n = 28: AI reviewers inspected code and repositories to uncover discrepancies between manuscripts and implementations, including data leakage, computational bias, broken links, and incorrect evaluation procedures.Several passages describe code findings that could materially alter or even overturn the papers’ reported claims.
- S3: Specialized niche field catch — n = 27: AI reviewers uniquely caught specialized domain issues, such as diastereoselectivity, population structure, benchmark limitations, biosafety, misleading terminology, insufficient flexibility testing, and inadequate sensor calibration.Reviewers also identified concerns about niche physical interpretations, computational cost, systematic uncertainties, and tissue-equivalent simulation properties.
- S4: Internal consistency across sections — n = 15: AI reviewers detected internal inconsistencies between claims, experimental designs, reported values, and figure captions, including impossible physiological values and contradictory passivation statements.They also challenged technically flawed comparisons and mismatches between the claimed and actual operating ranges.
- S5: Reproducibility / dependency failures — n = 10: AI reviewers identified reproducibility failures involving missing parameter specifications and modified external tools in the experimental pipeline.These omissions were characterized as reproducibility failures rather than merely presentation issues.
- S6: Big-picture / counter-narrative synthesis — n = 7: AI reviewers sometimes provided stronger big-picture analysis by identifying circularity, counter-narratives, weak novelty claims, insufficient baselines, and disorder-specific pathology.One critique argued that pandemic temperature-related mortality may be overestimated, countering the paper’s narrative that compound impacts are underestimated.
S generic: Residual: AI judged Correct without specific reason — n = 40 · F Details of the AI meta-reviewer and PEERREVIEW BENCH · F.1 The dual-annotated calibration set
The residual examples show that experts often endorsed AI criticisms as correct, sometimes identifying important issues missed by humans, while occasionally acknowledging limited expertise. The appendix then defines the AI meta-reviewer and PEERREVIEW BENCH through a dual-annotated calibration set, paired evaluation settings, and paper-access tools.
- S generic: Residual: AI judged Correct without specific reason — n = 40: Some endorsements remained qualified, with experts saying an issue was worth discussing, difficult to address methodologically, or reasonable despite limited domain expertise.Several experts explicitly stated they were not experts in the relevant analysis procedure.
- S generic: Residual: AI judged Correct without specific reason — n = 40: Experts repeatedly judged AI criticisms as good, well-argued, valid, or important, including issues that human reviewers had missed.Examples include a major issue identified by Claude 4.5 and a good point missed by human reviewers.
- S generic: Residual: AI judged Correct without specific reason — n = 40: AI reviews were viewed as strong on technical matters and thoroughness, whereas human reviews contributed points associated with years of experience.One expert rated an AI review above all human reviews but below other AI reviews.
- F Details of the AI meta-reviewer and PEERREVIEW BENCH: The appendix documents the AI meta-reviewer, its calibration against human experts, PEERREVIEW BENCH construction, leaderboard, and verbatim prompt.These components are presented as supporting documentation for precision judgments in the benchmark.
- F.1 The dual-annotated calibration set: 908 review items from 27 dual-annotated papers form the calibration set, comprising 568 human and 340 AI items with three-axis expert judgments.The axes are correctness, significance, and evidence sufficiency, assigned under the study’s cascade protocol.
- F.1 The dual-annotated calibration set: The calibration evaluates two settings: the meta-reviewer’s own per-axis judgments and predictions of the two experts’ joint ten-class judgment.The secondary label captures both cascade outcome and inter-annotator agreement.
- F.1 The dual-annotated calibration set: 32.8%, 30.5%, and 14.2% are reported for selected ten-class outcomes, including both-correct disagreement patterns and disagreement on correctness.The supplied passage lists “both correct, both significant, evidence sufficient” at 30.5% and “disagree on correctness” at 14.2%.
- F.1 The dual-annotated calibration set: The meta-reviewer uses terminal, file-editor, and web-search tools to inspect preprint markdown, figures, source code, and supplementary material like a human meta-reviewer.This setup addresses review items whose claims depend on specific figures, files, or supplementary sections.
F.2 Calibration results … F.4 PEERREVIEW BENCH: construction and evaluation protocol
Calibration shows AI meta-reviewers approach human agreement but share a more uniform judgment style, while failure analysis identifies recurring errors in evidence interpretation, significance calibration, reviewer-type asymmetry, disagreement prediction, and context management. PEERREVIEW BENCH is constructed from fully positive human review items across 78 papers.
- F.2 Calibration results: 87.9% correctness accuracy was achieved by Claude-Opus-4.7, above the 85.8% primary–secondary human baseline, while significance remained hardest for both humans and AI.GPT-5.4 and Gemini-3.1-Pro reached 82.0% and 81.5% correctness; AI significance scores were 56.7%, 56.9%, and 54.0% against 59.9%, and evidence scores were 85.6%, 85.3%, and 87.4% against 88.0%.
- F.2 Calibration results: AI–AI agreement reached AC1 values of 0.83–0.88 on significance, 0.90–0.93 on evidence, and 0.86–0.91 on correctness, exceeding AI–human agreement.AI–human AC1 ranged from 0.43 to 0.55 on significance, 0.83 to 0.93 on evidence, and 0.77 to 0.86 on correctness, compared with a 0.44 human–human significance baseline.
- F.3.2 The partial-evidence trap: 41 of 54 correctness errors came from the partial-evidence trap, where partial coverage was treated as fully resolving a reviewer’s concern.The proposed remedy is to decompose multi-part review items into atomic sub-claims and evaluate each independently.
- F.3.3 Over-leniency on technically detailed items: 13 false positives arose when technically precise review items persuaded the AI meta-reviewer despite both experts rejecting the underlying claims.Ten of these 13 errors involved AI-generated review items, consistent with a lower acceptance threshold for technical vocabulary and code-level evidence.
- F.3.4 Significance boundary miscalibration: 56 significance errors reflected a tendency to underweight methodological rigour, scope qualification, and translational relevance while over-weighting presentation issues involving technical terminology.The AI meta-reviewer commonly applied a “would addressing this change the core result?” test, downgrading substantive concerns or upgrading minor ones.
- F.3.5 Evidence closure demand: 13 evidence errors showed a demand for self-contained analytical closure rather than accepting pointed, verifiable gaps that experts considered sufficient.This pattern included demanding paper-specific evidence for established phenomena and justification for proposed benchmarks.
- F.3.6 Reviewer-type asymmetry: 41 false negatives involved human items, whereas 10 of 13 false positives involved AI items, reflecting stricter treatment of informal human critiques and greater leniency toward technically specific AI critiques.The analysis attributes this asymmetry to a prior that specificity correlates with correctness.
- F.3.7 Expert-disagreement prediction failures; F.3.8 Context anchoring and paper-mismatch blind spots; F.3.9 How the failure categories interact; F.4 PEERREVIEW BENCH: construction and evaluation protocol: 50 sampled cases involved confident predictions of expert agreement when experts disagreed, while 30 involved predicted disagreement when experts converged; shared context also produced anchoring and paper-mismatch errors.The analysis recommends independent per-item evaluation or an explicit reset before each item; PEERREVIEW BENCH contains 78 papers whose rubrics pool fully positive human items.
F.5 Full leaderboard analysis … G.1 Implementation and design
The leaderboard shows substantial differences among AI reviewers, with Claude-Opus-4.5 leading F1 and newer models not consistently outperforming older ones. The appendix also specifies a structured meta-reviewer workflow and describes the CMU PAPER REVIEWER platform’s implementation, mitigations, and configurable settings.
- F.5 Full leaderboard analysis: Claude-Opus-4.5 leads the twelve-model leaderboard at F1 = 50.89, narrowly ahead of Claude-Opus-4.7 at F1 = 50.46.DeepSeek-V4-Pro and GPT-5.2 form the next tier at F1 48.52 and F1 47.37, respectively.
- F.5 Full leaderboard analysis: Claude-Opus models lead F1 from a balanced precision-recall profile, whereas GPT-5 models favor high precision at the expense of recall.GPT-5.4 reaches precision 93.81% while raising 26.55% of the human rubric; GPT-5.2 follows the same pattern.
- F.5 Full leaderboard analysis: Newer models do not consistently perform better: Claude-Opus-4.5 exceeds Claude-Opus-4.7 in F1, and GPT-5.2 exceeds GPT-5.4 by nearly six points.The GPT comparison reflects a tradeoff in which GPT-5.4 gains precision while losing recall relative to GPT-5.2.
- F.5 Full leaderboard analysis: PEERREVIEW BENCH provides a concrete testbed for measuring targeted improvements intended to make models useful as scientific peer reviewers.The benchmark is positioned as an evaluation resource for model developers.
- F.6 AI meta-reviewer prompt: The AI meta-reviewer receives paper source files and reconstructed human and AI reviews, then judges every review item for correctness, significance, and evidence sufficiency.It also predicts how two independent expert meta-reviewers would jointly judge each item using 10 collapsed class labels.
- G Details of the CMU PAPER REVIEWER platform: The CMU PAPER REVIEWER platform uses OpenHands with terminal, file editor, Tavily web search, and Mistral OCR tools to match the AI reviewer analyzed in the study.The appendix documents the released platform’s implementation and intended-use policy.
- G.1 Implementation and design: The platform mitigates vague critiques by pairing each review item with a manuscript edit or, when code is provided, a runnable patch, script, and README.It also supports limitation-grounded review items, interactive author challenges, and controls for post-preprint citations.
- G.1 Implementation and design: Users can configure item counts, criteria presets, and supplementary materials, with default item counts ranging from 1 to 15 and no page limit for added materials.Available criteria presets include Nature default, NeurIPS, and custom settings.
G.2 Comparison with public AI reviewer platforms … H.1 Methodology and metric definitions
The CMU PAPER REVIEWER outperforms two public AI reviewer platforms on benchmark F1, while panel-composition analysis defines how human–AI reviewer mixes affect useful, unique feedback and author burden. The platform is publicly available as an assistive pre-submission tool, not a replacement for permitted human review.
- G.2 Comparison with public AI reviewer platforms: G.2 Comparison with public AI reviewer platforms: The CMU PAPER REVIEWER was compared with Stanford Agentic Reviewer and OpenAIReview on 78 PEERREVIEW papers using precision, recall, and F1.Weaknesses bullets and feedback cards were treated as individual review items for the two public services.
- G.2 Comparison with public AI reviewer platforms: 58.64 F1: GPT-5.4 with a 15-item cap exceeded Stanford Agentic Reviewer at 51.65 and OpenAIReview at 47.88.The CMU PAPER REVIEWER dominated on F1 across configurations.
- G.2 Comparison with public AI reviewer platforms: 93.81% precision and 26.55% recall: GPT-5.4 with a 5-item cap favored high precision, whereas Claude-Opus-4.7 with 5 items achieved 50.46 F1 with 71.47% precision and 39.00% recall.The Claude configuration produced 4.73 items per paper, versus 11.08 and 18.64 for competing platforms.
- G.2 Comparison with public AI reviewer platforms: 3.60 to 7.35 items per paper: tripling GPT-5.4’s cap from 5 to 15 roughly doubled output while improving coverage without sacrificing precision.Internal selection emits additional items only when they meet the three-axis fully-positive bar.
- G.3 Availability and intended use: G.3 Availability and intended use: The platform offers up to three free reviews daily, unlimited use with a user API key, and open-source deployment.It is intended for pre-submission feedback and should not be used where conferences or journals prohibit AI reviewers.
- H Recommendation for journal and conference organizers: Panel composition analysis: H Recommendation for journal and conference organizers: The appendix simulates four human–AI panel configurations, evaluates seven review-item metrics, and reports results, priorities, generalization caveats, and 95% confidence intervals.The analysis is designed to inform organizer decisions about deploying AI reviewers.
- H.1 Methodology and metric definitions: H.1 Methodology and metric definitions: Panels comprise 3H, 2H+1AI, 1H+2AI, or 3AI across 53 papers, with mixed panels averaged over nine combinations and optional GPT-5.4 filtering.The filter removes items judged not fully positive.
- H.1 Methodology and metric definitions: The bottom-line metric is # Fully Pos. + Unique Items: feedback that is correct, significant, evidence-sufficient, and unique within the panel.Other metrics quantify total output, uniqueness, non-fully-positive burden, useful-item fractions, overall yield, and Noise per Gem.
H.2 Panel composition results
Panel composition determines the trade-off between useful-feedback volume and review efficiency. Hybrid panels preserve human-level useful feedback, while AI-heavy filtered panels prioritize low-noise, high-confidence selection at the cost of volume.
- Recommendation 1 (2H+1AI): 2H+1AI matches 3 humans on fully-positive-and-unique items (3.9 vs 3.9) while producing fewer total items (21.4 vs 25.8).It also has higher yield (18.2% vs 15.1%), higher quality of unique items (35.8% vs 33.9%), and lower noise per gem (2.95 vs 3.74).
- Filtering trade-off: The AI meta-reviewer filter raises every panel’s efficiency, increasing yield and lowering noise per gem while reducing the absolute count of fully-positive-and-unique items.For 2H+1AI, yield rises from 18.2% to 20.0% and noise per gem falls from 2.95 to 2.14.
- Recommendation 3 (3AI + meta-reviewer filter): The 3AI filtered panel produces 3.5 not-fully-positive items per paper versus 14.6 for 3 humans, and 63.2% of its unique items are fully positive versus 33.9%.AI reviewers converge on well-known issues, while distinctive departures tend to be narrow, verifiable critiques; filtering sharpens this selectivity.
- Recommendation 2 (1H+2AI + meta-reviewer filter): 20.6% of 1H+2AI filtered-panel items are fully positive and unique, versus 15.1% for 3 humans, while noise per gem falls from 3.74 to 1.95.The configuration delivers 2.1 fully-positive-and-unique items per paper versus 3.9 for the 3-human baseline, trading useful-content volume for triage efficiency.
- Caveats: The recommendations may differ across venues because the study used one Nature submission pool, broad reviewer disciplines, nonspecialist-oriented papers, a per-AI five-item cap, and three frontier AI reviewers.Narrower technical venues, single-discipline journals, shorter review timelines, or different experimental settings may exhibit different dynamics.