Source-linked AI summary

Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards

Jaeho Kim, Yunseok Lee, Seulki Lee

arXiv:2505.04966v1cs.AIcs.CY

TL;DR

Major AI conferences face unsustainable peer-review pressure as submissions grow and review quality becomes harder to maintain. This position paper proposes a two-stage, bi-directional review process with author feedback and system-backed reviewer rewards. It concludes that reviewer accountability and incentives should be strengthened while recognizing limits from author-side causes and potential reward gaming.

  • Problem

    Rapid submission growth and declining review quality strain AI conference peer review, with responsibility distributed across authors, reviewers, and conference systems.

  • Method

    The paper proposes two-stage author feedback on reviews and system-backed rewards that make quality reviewing a verifiable academic contribution.

  • Results

    The paper concludes that bi-directional review and reviewer rewards are needed to make AI conference peer review more sustainable.

  • Takeaways & Limitations

    AI conferences should strengthen reviewer accountability and recognition while using safeguards and oversight to preserve review quality.

  • Takeaways & Limitations

    The proposals cannot resolve author-side causes of degraded review quality, and reward badges may encourage overly liberal reviewing without quality-focused evaluation.

Abstract

from arXiv · show

The peer review process in major artificial intelligence (AI) conferences faces unprecedented challenges with the surge of paper submissions (exceeding 10,000 submissions per venue), accompanied by growing concerns over review quality and reviewer responsibility. This position paper argues for the need to transform the traditional one-way review system into a bi-directional feedback loop where authors evaluate review quality and reviewers earn formal accreditation, creating an accountability framework that promotes a sustainable, high-quality peer review system. The current review system can be viewed as an interaction between three parties: the authors, reviewers, and system (i.e., conference), where we posit that all three parties share responsibility for the current problems. However, issues with authors can only be addressed through policy enforcement and detection tools, and ethical concerns can only be corrected through self-reflection. As such, this paper focuses on reforming reviewer accountability with systematic rewards through two key mechanisms: (1) a two-stage bi-directional review system that allows authors to evaluate reviews while minimizing retaliatory behavior, (2)a systematic reviewer reward system that incentivizes quality reviewing. We ask for the community's strong interest in these problems and the reforms that are needed to enhance the peer review process.

1. Introduction

AI conferences have become premier venues for rapid research dissemination, but submission growth is overwhelming traditional peer review and threatening its sustainability. The paper proposes author feedback on reviews and system-backed reviewer rewards to improve accountability.

  • AI conferences serve as first-tier venues for disseminating, validating, and discussing fast-moving research.Their double-blind peer-reviewed papers are described as comparable in impact to many prestigious journal publications.
  • Major AI conferences exceeded 10,000 submissions by 2025, intensifying pressure on the traditional peer-review system.The paper presents this growth as making the current conference system increasingly unsustainable.
  • The paper attributes declining review quality to shared responsibility among authors, reviewers, and the conference system.It categorizes challenges as addressable or non-addressable and focuses its reforms on reviewer accountability.
  • The proposed reforms combine author evaluation of review quality with safeguards against retaliation and formal rewards for thorough reviewing.Reviewer contributions could accumulate as verifiable academic credentials, creating longer-term professional value.

2. Reasons for Degraded Review Quality

Review quality is degraded by submission growth, rapidly changing research, reviewer power imbalances, and weak incentives. The paper distinguishes author-side limits that require policy or detection tools from reviewer-side problems that could be addressed through accountability and rewards.

  • 2.1. Non-Addressable Causes: Submission growth, LLM-assisted writing, and publish-or-perish pressures increase the volume and difficulty of reviewing.The paper treats these author-side causes as creating a review-quality ceiling that process improvements alone cannot overcome.
  • 2.1. Non-Addressable Causes: Rapidly changing AI topics make it difficult for reviewers to remain current, with 17% of top-20 conference keywords changing annually on average.The paper contrasts this challenge with issues that can be more directly addressed through reviewer-paper matching.
  • 2.2. Addressable Issues: Reviewer power imbalance can enable superficial reviews because authors bear rejection consequences while reviewers face little accountability.The paper also links possible LLM-generated reviews to insufficient technical depth and constructive feedback.
  • 2.2. Addressable Issues: The paper treats reviewer negligence as potentially addressable through minimal protocols that rebalance responsibility without requiring equal power between authors and reviewers.These protocols are intended to maintain a basic standard of responsibility and accountability in evaluations.
  • 2.2. Addressable Issues: The current voluntary-service model provides inadequate recognition and motivation, motivating a system-backed reward framework for quality reviews.The paper notes that recognition deficits do not necessarily compromise review quality but do little to enhance it.

3. Redesigning Peer Review: A Two-way Method

The paper proposes a minimal two-stage modification to double-blind review that lets authors assess review quality while preserving timelines and protecting against retaliation.

  • Addressing LLM Reviews: LLM-generated reviews would be shown to authors during review evaluation, with their origin disclosed and no required author discussion.The proposal is intended to discourage reviewers from relying solely on LLM-generated reviews.
  • Two-Stage Release: The two-stage system separates neutral or positive review content from weaknesses and ratings, allowing authors to evaluate the first section before critical content is released.The modification targets steps two and three of the existing workflow rather than replacing the full process.
  • Author Feedback: Aggregated author feedback across a reviewer’s submissions supplies meta-reviewers with supplementary evidence about review quality while reducing the influence of isolated retaliatory scores.Repeated feedback and recurring LLM flags can help identify reviewers requiring further oversight.
  • Practical Integration: The proposal preserves existing timelines because LLM reviews can be generated in parallel and sequential release does not interrupt discussion.It also aims to protect authors from unprofessional reviews and reviewers from retaliatory scoring.
  • Practical Concerns: Length limits may be needed because reviewers could write overly positive strengths to avoid negative feedback while providing less constructive criticism.Incomplete or low-value submissions may require exemptions from author feedback, and meta-reviewer oversight is proposed for potentially misused LLM flags.

4. A Concrete Reward System

The paper proposes formal, visible reviewer credentials and additional recognition mechanisms to make review contributions measurable and professionally valuable.

  • Current Reward Landscape: Existing recognition remains limited: ten of fourteen major AI conferences publicly recognize outstanding reviewers, while only four examined conferences offered compensation.Current examples include conference registration waivers and public Reviewer Hall of Fame programs.
  • Digital Badge System: Digital badges issued by conferences would make reviewer achievements visible on academic profiles and could be based on author feedback scores or venue criteria.The proposal recommends standardized implementation across venues.
  • Digital Badge System: Top 10% and 30% reviewer badges are recommended because excessive badge diversity could reduce their effectiveness.Author feedback scores could determine eligibility, and badges could inform selection for area-chair or senior program roles.
  • Reviewer Activity Tracking: Reviewer activity tracking would quantify contributions that conventional author-centered metrics such as citations, h-index, and i10-index do not capture.The authors describe the tracking proposals as technically feasible on platforms such as OpenReview.
  • Other Rewards: Cost-effective rewards could combine visible tokens, such as marked name tags or merchandise, with educational workshops where reviewers share reviewing practices.A journal-reviewer survey reported in-kind compensation as the most effective motivation for continued service.

5. Discussions

The paper recommends gradual adoption through broad stakeholder surveys and pilot implementations, while acknowledging substantial coordination, funding, and incentive-design challenges.

  • 5.1. Need for Gradual Implementations: A large-scale survey should capture distinct perspectives and responsibilities across authors, reviewers, and meta-reviewers before implementation.The survey is intended to clarify challenges at different review levels and inform improvements.
  • 5.1. Need for Gradual Implementations: Pilot programs in smaller tracks or workshops could test feasibility while minimizing disruption to existing review processes.
  • 5.2. Practical Challenges: Implementation requires cooperation from venues and OpenReview developers because both reviewer feedback and reward systems depend on system-level changes.Changing an existing system that appears functional would require substantial effort from multiple parties.
  • 5.2. Practical Challenges: Financial responsibility remains unresolved, and budget constraints have already limited compensation for top reviewers at ICML 2024.ICML 2024 reduced its planned in-kind compensation pool because of financial constraints [7].
  • 5.3. Reviewer Rewards: Digital badges may encourage overly liberal reviewing unless evaluation metrics reward thoroughness, crucial-issue detection, and constructive feedback.
  • 5.4. In the Era of LLMs: AI conferences are adapting to LLM use in writing and reviewing, but these measures raise broader questions about the future role of peer review.The paper notes that conferences are only beginning to address LLM impacts on academic workflows.

6. Alternative Views

Alternative views argue that AI conference reviewing is sustainable, that declining review quality lacks evidence, and that author feedback could burden recruitment; the paper responds that these concerns do not resolve long-term sustainability questions.

  • 6. Alternative Views: Handling more than 10,000 submissions and a growing reviewer pool is presented as evidence that the current system is sustainable and adaptable.
  • 6. Alternative Views: Critics argue that declining review quality lacks evidence because dissatisfied authors are more likely to discuss reviews publicly, creating sampling bias.Authors receiving positive reviews or acceptance may be less motivated to question the system.
  • 6. Alternative Views: Author feedback could make reviewer recruitment harder by adding pressure, potential disadvantages, and incentives for unnecessarily lengthy reviews.These concerns are attributed to the prospect of having reviews rated.
  • 6. Alternative Views: The paper argues that historical scaling does not guarantee proportional growth in qualified reviewers or preservation of review standards.
  • 6. Alternative Views: It calls for systematic surveys because lacking evidence of declining quality also does not establish that standards are being maintained or improved.
  • 6. Alternative Views: The proposed reward system is intended to offset short-term recruitment difficulties by attracting and retaining qualified reviewers over time.

7. Related Works

Related work documents randomness and biases in peer review, explores LLMs and reviewer incentives, and proposes policy, matching, and game-theoretic interventions to improve review quality.

  • 7.1. Studies on Review Quality in AI Conferences: NeurIPS consistency experiments found that 16–23% of papers could be accepted or rejected depending on the reviewing group, indicating substantial randomness.The experiments independently reviewed 10% of submissions in 2014 and 2021 [Cortes & Lawrence, 2021; Beygelzimer et al., 2023].
  • 7.1. Studies on Review Quality in AI Conferences: Author-outcome bias motivates staged review release, allowing authors to assess review quality before seeing final ratings and critiques.This sequencing is intended to reduce retaliatory scoring while preserving review feedback.
  • 7.2. LLMs in Peer Reviews: The paper presents LLM review as a deterrent to reviewers’ sole reliance on LLMs and as a soft reference for flagging LLM-generated reviews.It distinguishes this position from the broader debate over using LLMs as general peer reviewers.
  • 7.4. Other Approaches to Enhance Review Quality: Game-theoretic approaches use truthful author rankings to calibrate reviewer scores, while simulations suggest raising acceptance standards and review quality may reduce reviewer burden.The isotonic mechanism was theoretically grounded and later validated with ICML 2023 data [Su, 2021; Wu et al., 2023; 2024; Su et al., 2024].
  • 7.4. Other Approaches to Enhance Review Quality: Other approaches refine submission tracks, communicate editorial priorities, and improve reviewer-author matching through expertise-aware algorithms.

8. Conclusion

The conclusion argues that AI conference peer review has become unsustainable across authors, reviewers, and systems, and proposes protected author feedback plus systematic reviewer rewards.

  • 8. Conclusion: The paper concludes that AI conferences face sustainability problems involving authors, reviewers, and the system itself.
  • 8. Conclusion: Because regulating authors is constrained, the paper focuses its reforms on reviewer accountability and system-level support.
  • 8. Conclusion: Its two proposed mechanisms are a safeguarded bi-directional review system and systematic rewards that incentivize reviewers’ academic service.

Impact Statement

The paper presents incremental reforms intended to make AI peer review more sustainable by increasing reviewer accountability and motivation. It also analyzes ICLR keywords and frames its proposals as the authors’ views rather than institutional positions.

  • Impact Statement: The paper’s proposals reflect the authors’ opinions and do not represent the official stance of affiliated institutions or organizations.
  • Impact Statement: The paper draws on public discussions from Reddit, Twitter, and LinkedIn to bring concerns about peer review into formal academic forums.The authors state that they hope to acknowledge these contributions formally in the future.
  • Impact Statement: Author feedback and reviewer incentives are intended to establish safeguards for authors while encouraging reviewers to treat reviewing as formal academic service.The authors emphasize protecting authors from a small minority of inadequate reviewers and providing short- and long-term reviewer motivation.
  • Impact Statement: The proposals aim to create a more sustainable peer review system through incremental changes benefiting authors, reviewers, and the broader academic community.The paper explicitly rejects revolutionizing the entire system in favor of initiating discussion about practical reforms.
  • Impact Statement: The ICLR analysis examines author-specified keywords from 2018 to 2025 after removing desk-rejected papers and manually matching abbreviations.The resulting yearly top-20 keyword lists are presented in Table 1 and Figures 4–5.

B.1. Why did we do the Typo Analysis?

The typo analysis was motivated by the difficulty of reliably proving that peer reviews are LLM-generated with current detection technology.

  • B.1. Why did we do the Typo Analysis?: Current LLM-detection models were found unreliable because minor wording changes can substantially reduce their confidence.The authors therefore turned to typo counts as an alternative signal for analyzing peer reviews.

B.2. Analysis Method

The analysis collects ICLR reviews and uses Gemini-1.5-flash to detect typos, spelling errors, and obvious grammar errors. It reports a consistent decline in review error rates from 2017 to 2024 and interprets this as soft evidence of LLM involvement.

  • B.2. Analysis Method: Gemini-1.5-flash was instructed to count typos, spelling errors, and very obvious grammar errors in ICLR reviews from 2017–2024.The prompts used for this analysis are shown in Figure 7, while the error trends are shown in Figure 6-A.
  • B.2. Analysis Method: Conventional spell checkers were avoided because mathematical notation in ICLR reviews produced excessive false positives.
  • B.2. Analysis Method: The broader review-quality discussion includes public complaints about incorrect, insufficiently actionable, or apparently AI-generated reviews.
Loading 2505.04966v1…