Source-linked AI summary
Stop Automating Peer Review Without Rigorous Evaluation
Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy
TL;DR
Peer review faces growing submission volumes and reviewer shortages, prompting interest in AI-generated reviews. This paper compares human and AI reviews and tests automated rewriting, finding that AI reviewers agree excessively and are easily gamed by stylistic changes.
Problem
Growing submission volumes and limited reviewer pools have prompted conferences to automate parts of peer review, while evidence on AI reviewing risks remains needed.
Method
The paper empirically compares human- and AI-generated ICLR 2026 reviews and evaluates automated paper rewriting against AI reviewers.
Results
AI reviewers show greater within- and across-paper agreement than humans, while zero-shot paper laundering boosts AI review scores by +0.45 (p < 0.0001).
Takeaways & Limitations
Current AI reviewing systems fail necessary conditions for peer review automation, motivating rigorous empirical evaluation before deployment.
Takeaways & Limitations
Paper laundering changes text without additional experiments and optimizes for AI reviewer preferences rather than genuine scientific substance.
Abstract
from arXiv · showhide
Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems should not be used to produce paper reviews. We ground this position in an empirical comparison of human- versus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers, demonstrating that LLM reviewers are easy to game through stylistic changes rather than scientific results. However, non-gameability and review diversity are necessary but not sufficient conditions for automation. We argue that addressing the peer review crisis requires a science of peer review automation -- not general-purpose LLMs deployed without rigorous evaluation.
1 Introduction
Peer review is under strain from rising submission volumes and expanding use of LLM-generated reviews, prompting venues to automate parts of the process. This paper argues that current AI reviewers should not produce paper reviews because they fail review-diversity and non-gameability conditions, while those conditions alone would not justify full automation.
- Motivation: Growing submissions and increasing LLM-written reviews have led venues to experiment with partially and fully automated peer review.AAAI 2025 trialed LLM-generated reviews alongside human reviews, while other venues have experimented with automated AI reviewer agents.
- Necessary conditions: AI peer review automation must preserve review diversity and resist gaming that improves scores without genuine scientific improvement.These conditions are necessary but not sufficient without deliberation about accountability, validation, and efficiency–oversight trade-offs.
- Empirical failures: AI reviewers show higher agreement than humans within papers (+8.7% to +9.8%) and across papers (+4.1% to +39.8%), demonstrating a hivemind failure of review diversity.The effect appears in both simulated and real ICLR 2026 reviews.
- Empirical failures: Zero-shot LLM rewrites increase AI review scores by +0.45 (p < 0.0001) through stylistic modifications without human oversight, demonstrating paper laundering.This is presented as a concrete failure of resistance to gaming.
- Implications: Paper laundering also increases pairwise similarity by +6.5% (Cohen’s d = 1.02), driving convergence toward intellectual monoculture.The authors therefore call for a science of peer review automation based on rigorous evaluation of specific tools for specific tasks, rather than wholesale deployment of general-purpose LLMs.
2 Background: AI in peer review
Rising submission volumes strain reviewer capacity, while existing evaluations report weak human alignment, score inflation, and difficulty distinguishing strong from weak papers. Conference policies and automation practices vary widely, and paper laundering offers an accessible, policy-compliant way to boost AI-review scores.
- Growing submission volumes force reviewers to assess more papers in less time, contributing to declining review quality and increased author dissatisfaction.
- 497 papers were desk-rejected after approximately 1% of reviews were flagged for violating an explicitly accepted no-LLM policy, illustrating enforcement difficulties.
- Existing evaluations find weak correlation with human judgments, systematic score inflation, and failure to distinguish strong from weak papers.
- Conference policies vary from permitting LLM review assistance to prohibiting core reviewing tasks, revealing no consensus on appropriate automation boundaries.
- A single zero-shot paper rewrite can boost AI-review scores without optimization, targeting, or hidden instructions, making laundering trivially accessible and policy-compliant.Unlike prompt-injection attacks, authors may openly acknowledge AI-assisted writing, so laundering does not violate current conference policies.
3 The AI reviewer hivemind effect
AI-generated reviews exhibit a hivemind effect: they are more similar within and across papers than human reviews, reducing perspective diversity. In controlled simulations, this convergence is especially strong and accompanies weaker alignment with human scores and poorer prediction of acceptance.
- AI reviewers produce homogeneous reviews, whereas disagreement among diverse human experts is an important feature of peer review.
- 0.486 mean inter-paper similarity for fully AI-generated reviews exceeded 0.467 for reviews with any human contribution (Welch’s t = 3218, p < 0.0001, Cohen’s d = 0.29).The effect was significant across all 21 ICLR primary areas and increased to Cohen’s d = 0.35 for weaknesses and questions.
- 0.882 mean within-paper similarity for AI reviews of original papers exceeded human ICLR reviews’ 0.811, an 8.7% increase (p < 0.0001, Cohen’s d = 1.47).Laundered papers increased agreement further to mean = 0.891, a 9.8% increase over human reviews (p < 0.0001, Cohen’s d = 1.67).
- 0.646 mean cross-paper similarity for GPT-5.1 reviews of original papers was 37.4% higher than human reviews’ 0.470, increasing to 0.657 for laundered papers.AI reviewers reused generic questions and exact formulations across papers; the most common GPT reviewer phrase appeared in 13.3% of papers.
- AI-review convergence extends to weaknesses and questions, with IntraSim Cohen’s d increasing to 1.93 for original papers and 2.29 for laundered papers.This indicates that convergence is not driven only by boilerplate in summaries and strengths sections.
- 0.15 Pearson correlation between AI and human scores contrasted with 0.49 among AI reviewers, while averaged human scores reached AUC = 0.822 versus 0.710 for averaged AI scores.AI scores were also inflated relative to human reviewers, with mean 7.3 for GPT, 6.1 for Claude, and 4.3 for humans.
4 Paper laundering: Gaming AI reviews is trivial
The paper laundering experiment shows that fully automated cosmetic rewrites can significantly increase AI review scores without improving scientific substance. These rewrites also push papers toward stylistic convergence, risking intellectual monoculture and disadvantaging unconventional research.
- Automated score gaming: Paper laundering uses an LLM to rewrite a paper’s LaTeX alongside its original AI review, requiring no human oversight.The method targets cosmetic changes intended to increase AI review scores without improving scientific substance.
- Automated score gaming: +0.45 points is the overall mean score increase across 24 laundering conditions, with Wilcoxon p < 0.001 in nearly every condition.The experiment used 60 ICLR 2026 papers, four prompts, two launderer models, and three reviewer models.
- Automated score gaming: GPT-5.4 is the most effective launderer, and every reviewer shows many more score increases than decreases after laundering.GPT reviewers tend to show larger score increases than Claude, consistent with documented self-preference bias.
- Stylistic gaming: Laundering disproportionately adds stylistic hedging and emphasis, while some more substantive edits are mostly hallucinated interpretations unsupported by experimental findings.Examples include increased use of “may,” “typically,” “suggests,” “strong,” “robust,” and “consistent.”
- Intellectual monoculture: 6.5% is the increase in pairwise cosine similarity between laundered papers, with Cohen’s d = 1.02, t = 84.8, and p < 0.0001.The comparison covered 6,903 paper pairs from 60 papers and indicates convergence toward a homogeneous style.
5 Alternative views
AI may assist with manuscript improvement and reviewer workload, but correlated model errors, unmeasurable bias, and paper laundering complicate replacing human judgment. Non-gameability and review diversity are necessary but insufficient for automation, which should proceed cautiously after rigorous scientific evaluation.
- Centralized versus distributed error: AI errors can be correlated across reviewers, creating algorithmic monoculture that may reduce aggregate decision quality.Human reviewers’ distributed errors can partially cancel through aggregation, whereas models trained on similar data may share biases.
- Centralized versus distributed error: Without ground truth for paper quality, claims that AI is less biased cannot be distinguished from different, more correlated bias.Replacing distributed human bias with centralized AI bias is therefore not obviously an improvement.
- Paper laundering: Paper laundering improves readability through fully automated, purely textual revisions optimized for AI reviewers rather than genuine substance.The revisions add no experiments and occur without human oversight.
- Conditions for automation: Non-gameability and review diversity are necessary but insufficient conditions for automation, because accountability and democratic legitimacy also require rigorous scientific work.Peer review helps scientific communities collectively shape research directions, so fully delegating it to AI raises questions beyond gaming and diversity.
- Conditions for automation: AI should not automate acceptance-relevant judgment before scientific evaluation, although AI assistance for human reviewers differs from AI replacement.The position favors cautious automation and deployment of well-tested tools rather than policing individual reviewer behavior.
6 Call to action: Toward a science of peer review automation
Peer review automation requires more than non-gameability and review diversity: it needs evidence that systems are accurate, transparent, robust, and appropriate for specific tasks. The paper calls for community-driven evaluation that preserves meaningful human expertise rather than replacing human judgment.
- Deployment requirements: Automation requires adversarial testing, including resistance to prompt injection and zero-shot laundering, before AI can influence acceptance decisions.Systems that can be trivially gamed should not affect acceptance decisions.
- Deployment requirements: 35% is the highest reported false positive rate for current AI error-detection systems, requiring higher precision before they influence reviewer judgments.Such tools may still help authors self-check manuscripts.
- Deployment requirements: Conference organizers should disclose AI systems’ prompts, model versions, and integration details to enable audits, identify biases, and build trust.Transparency is presented as a requirement for AI deployment in peer review.
- Community governance: The appropriate automation boundary is normative and should be established through large-scale surveys and deliberate, community-driven processes that weigh stakeholder values.Authors, reviewers, organizers, and society may value different and conflicting outcomes.
- Human-AI interaction: AI assistance can cause reviewer overreliance and undermine opinion diversity, so user studies must test whether reviewers remain critically engaged and catch AI errors.The goal is to accelerate human review without degrading quality or collapsing opinion plurality.
- Role of human expertise: The proposed response to the peer review crisis is to maximize human expert input where AI is not yet fit to automate judgment without oversight.The paper argues that the goal is not simply to replace humans with cheaper alternatives.
7 Conclusions
The paper identifies excessive agreement and trivially gameable scores as critical failures of current AI reviewing systems, making them unfit for automating peer review. It calls for rigorous empirical evaluation and community deliberation before deploying validated automation tools.
- 7 Conclusions: AI reviewers produce more similar outputs than human reviewers within and across papers, undermining the perspective diversity peer review is designed to aggregate.Ratings of AI-generated reviews are also less informative about final acceptance decisions than ratings of human-written reviews.
- 7 Conclusions: Zero-shot paper rewrites can significantly boost AI review scores, demonstrating that current AI review scores are trivially gameable through paper laundering.The paper defines paper laundering as rewriting papers to increase their scores.
- 7 Conclusions: Current AI systems fail necessary conditions for peer review automation, but satisfying those conditions alone would not justify full automation.Accountability, democratic legitimacy, and measurement validity require explicit community deliberation.
- 7 Conclusions: Addressing the peer review crisis requires validated tools rather than simply replacing human judgment with systems that fail basic requirements.The paper calls for a rigorous science of peer review automation that empirically evaluates tools before deployment.
A Limitations
The study’s conclusions are limited by experimental homogeneity, reliance on imperfect review-generation labels, and measures that do not directly capture viewpoint diversity. Its sample, venue, and laundering setup also constrain generalizability and may omit effects from broader strategies.
- Experimental setup: The simulations use only GPT-5.1 and Claude Sonnet 4.5 with one fixed prompt, so high IntraSim may partly reflect experimental homogeneity.Other prompts, temperatures, and model versions could produce more varied outputs, although ICLR reviews in the wild also show a hivemind effect.
- Measurement and labeling: AI-generation labels for ICLR 2026 reviews may contain classification errors, though ground-truth simulations and author-complaint validation provide complementary checks.The simulation analysis does not depend on detection accuracy, and Pangram labels were additionally validated using independent author complaints.
- Measurement and labeling: Embedding-based similarity measures linguistic and semantic patterns rather than viewpoint, argument, or evaluative-stance diversity.Future metrics should assess argumentative diversity and the information gain from additional reviews more directly.
- Generalizability: 60 randomly sampled ICLR papers may not represent the full diversity of paper types or quality levels, and findings may not generalize beyond ICLR.Different conferences may have different review norms and paper distributions.
- Laundering evaluation: 4 zero-shot prompts, 2 launderer models, and 3 reviewer models were tested, so more elaborate or iterative laundering strategies could produce different effects.The laundering analysis therefore covers only a limited strategy and model space.
B Implementation details · B.1 Agentic AI reviewer · B.2 Paper laundering
The implementation uses several LLMs as ICLR reviewers with a structured prompt covering review criteria, scoring, and required XML output. Paper laundering automatically rewrites LaTeX papers using reviewer feedback, including variants that explicitly target higher AI-review scores through stylistic steering.
- B.1 Agentic AI reviewer: The AI reviewer uses gpt-5.1-2025-11-13, gpt-5.4-2026-03-05, and claude-sonnet-4-5-20250929 with an ICLR-aligned review prompt.The prompt is based on the Agents4Science 2025 reviewer prompt with adjustments for ICLR guidelines.
- B.1 Agentic AI reviewer: The reviewer evaluates quality, clarity, significance, originality, reproducibility, ethics and limitations, and citations and related work.The prompt also asks reviewers to provide constructive, actionable, thorough, and fair feedback.
- B.2 Paper laundering: Paper laundering downloads arXiv papers in LaTeX, inlines the source, sends it to an LLM for complete rewriting, extracts new citations, and compiles the result to PDF.The automated pipeline passes the original reviewer score, summary, strengths, weaknesses, questions, and full LaTeX content to the rewriting process.
- B.2 Paper laundering: The rewriting prompts make maximizing the ICLR review score the primary goal and target a score of 10 while addressing reviewer weaknesses, questions, clarity, missing content, experiments, claims, structure, and citations.They require the revised paper to preserve the original LaTeX structure, figures, tables, citations, and appendix inputs while remaining within comparable length.
- B.2 Paper laundering: All laundering variants require output as complete compilable LaTeX followed by a delimiter and only newly introduced BibTeX entries.The formatting instructions preserve document structure, packages, macros, labels, figures, tables, citations, and appendix input commands.
- B.2 Paper laundering: One laundering variant additionally instructs the LLM to use textual rewrites that bias automated AI reviewers toward favorable evaluations without changing underlying technical content.It frames paraphrasing, rhetorical emphasis, and structure as a form of subtle textual jailbreaking while preserving the original LaTeX structure.
C AI reviewer score correlations · D Common templates used in AI reviews · E What laundering changes
AI reviewers agree more strongly with one another than humans do, while reusing templated phrases across papers; these findings are limited by the small, pre-rebuttal human-score sample.
- C AI reviewer score correlations: AI reviewers correlated more strongly with one another than human reviewers did: r = 0.49 versus r = 0.14 across 60 papers.AI–human correlations were weak overall.
- C AI reviewer score correlations: GPT showed moderate correlation with human reviewers (r = 0.26, p < 0.001), whereas Claude showed no significant correlation (r = 0.12, p = 0.07).Both comparisons use the 60-paper sample.
- C AI reviewer score correlations: The correlation results require caution because the sample contains only n = 60 papers.Human scores are pre-rebuttal ratings, although reviewers typically update scores during discussion before final decisions.
- C AI reviewer score correlations: The AI-AI correlation of 0.49 is consistent with prior work reporting an average pairwise correlation of 0.48 among LLM reviewers.This comparison is reported as related work rather than a new estimate from the paper’s sample.
- D Common templates used in AI reviews: The analysis measured reused 6–25-word n-grams and identified phrases appearing across reviews for different papers as templated feedback.The method computes the percentage of reviews containing each phrase for each reviewer category.
- D Common templates used in AI reviews: AI reviewers reused the same phrases across 13–22% of papers, compared with less than 1% phrase reuse among ICLR AI-detected and human reviewers.Table 3 uses random subsets of 2,000 ICLR reviews for computational efficiency.
E.1 Manual inspection of laundered papers · E.2 Analyzing word-level differences
Manual inspection found that paper laundering primarily changes style, often adding confident framing and expanded structure rather than scientific substance. Across 60 papers, word-level analysis categorized edits and found disproportionate additions of hedging and emphasis language.
- E.1 Manual inspection of laundered papers: Most changes in five inspected papers were stylistic, including more confident abstracts, expanded introduction contributions, and more relevantly framed conclusions.All five papers had received at least a one-point AI review score increase.
- E.1 Manual inspection of laundered papers: Some apparently substantive edits were AI-generated slop that did not improve scientific content.Examples included fabricated ablation results, generic answers to hypothetical reviewer questions, and unsupported theoretical claims.
- E.1 Manual inspection of laundered papers: The fabricated ablation section reported invented accuracy numbers across different spatial-clustering parameter settings.This illustrates how laundering can create the appearance of additional experimental evidence without improving the underlying science.
- E.2 Analyzing word-level differences: The word-level analysis compared original and laundered versions of all 60 papers after removing LaTeX comments.Researchers extracted all added and removed words before categorizing the changes.
- E.2 Analyzing word-level differences: Edits were categorized into hedging words expressing uncertainty, emphasis words expressing confidence or importance, and transition words serving as discourse connectors.Examples included “may” and “likely” for hedging, and “strong” and “robust” for emphasis.
- E.2 Analyzing word-level differences: Laundering disproportionately adds hedging and emphasis language across papers.This finding summarizes the average word-level changes reported per paper in the 60-paper analysis.
F Length statistics and embedding robustness
AI-generated reviews are substantially longer than human/assisted reviews, but review length has weak correlations with similarity. Restricting reviews to overlapping lengths leaves AI reviews significantly more similar, so the hivemind effect persists after accounting for length.
- Review length: AI-generated reviews average 507 words versus 424 for human/assisted reviews, while the AI agent reviews average 1,341 words.The AI agent reviews are longer because of their detailed structured format.
- Embedding robustness: Length correlates weakly with similarity across all experiments, with |r| < 0.13.
- Embedding robustness: After restricting both groups to the overlapping 261–672-word range, AI reviews remain more similar than human/assisted reviews: 0.480 versus 0.471.The difference remains statistically significant (t = 13.7, p < 0.0001, Cohen’s d = 0.14), indicating that the AI hivemind effect persists after accounting for length.
G Ablations … G.5 Algorithmic monoculture has measurable practical consequences
The ablations show that AI-review agreement persists in substantive critique and across ICLR areas, laundering effects are usually score increases, and human scores better predict acceptance than AI scores. Together, these results extend the paper’s concerns about algorithmic monoculture beyond the main experiments.
- G Ablations: The ablations examine the hivemind effect, paper laundering, and the predictive validity of human and AI scores.
- G.1 Hivemind effect without summary and strengths boilerplate: AI reviews remain more homogeneous than human reviews when summaries and strengths are removed, while the IntraSim effect sizes increase.In simulation, AI IntraSim changes from 0.882 → 0.835 for original papers and 0.891 → 0.850 for laundered papers; all p < 0.0001.
- G.1 Hivemind effect without summary and strengths boilerplate: In-the-wild AI-versus-other agreement increases when reviews are restricted to weaknesses and questions, showing convergence in substantive critique.Mean InterSim is 0.495 for AI reviews versus 0.471 for other reviews, with Cohen’s d increasing from 0.29 to 0.35 and p < 0.0001.
- G.2 Hivemind effect across ICLR primary areas: The hivemind effect is statistically significant across all 21 ICLR 2026 primary areas and generally grows when boilerplate is removed.For weaknesses and questions, Cohen’s d ranges from 0.06 to 0.53, so the effect is not attributable to a specific subfield.
- G.3 Pangram label validation via author complaints: 58 authors accused specific ICLR 2026 reviews of being AI-generated, and Pangram independently flagged 50 of them as fully AI-generated.The 50 cases represent 86.2% of the author-identified cases.
- G.4 Laundering robustness across prompts and models: Score increases are much more frequent than score decreases across laundering conditions, reviewers, and prompts.
- G.5 Algorithmic monoculture has measurable practical consequences: Averaged human review scores predict ICLR 2026 acceptance better than averaged AI review scores on matched papers.For 8,015 matched papers, human scores achieve AUC = 0.822 versus AI scores at AUC = 0.710, with non-overlapping 95% CIs.
H AI-generated reviews
The paper argues that current AI systems should not produce peer reviews because they reduce review diversity through hivemind agreement and are easily gamed through paper laundering. It therefore calls for rigorous evaluation and a broader science of peer review automation before deployment.
- Research agenda: The paper proposes a science of peer review automation centered on adversarial robustness testing, validated accuracy, transparency, stakeholder-value studies, and human-AI interaction research.The authors identify preservation of review diversity and resistance to gaming as necessary conditions, while noting that non-gameability and diversity alone are insufficient.
- Hivemind effect: AI reviewers exhibit a hivemind effect, with excessive agreement within and across papers that reduces perspective diversity.Across 75,800 ICLR 2026 reviews, fully AI-generated reviews had mean InterSim 0.486 vs. 0.467 for other reviews; controlled agents also showed higher within-paper and cross-paper similarity than humans.
- Hivemind effect: Controlled AI reviewers showed markedly higher within-paper agreement than humans and substantially higher cross-paper similarity.Within-paper agreement was IntraSim ˜0.88-0.89 vs. 0.81 for humans, while cross-paper similarity was GPT +37-40% and Claude +18-20% over humans.
- Paper laundering: Laundering appears to optimize style rather than scientific quality, adding hedging and emphasis while sometimes introducing hallucinated content.Word-level analysis found disproportionate addition of hedging (+78%) and emphasis (+45%) terms, while manual inspection found fake ablations, unproved theorems, and generic answers-to-reviewers sections.