Source-linked AI summary
Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
Shaochen Zhong
TL;DR
ML peer review faces rising submission volume and widely shared frustrations, while guidance alone appears insufficient. This position paper evaluates existing mechanisms and proposes enforceable safeguards paired with a spendable cross-conference credit system to support a more sustainable review ecosystem.
Problem
ML peer review faces rising submission volume, reciprocal reviewing obligations, and widely shared frustrations despite modern review platforms and substantial scholarly participation.
Method
The paper assesses existing conference mechanisms and proposes enforceable procedural guardrails with OpenReview Points earned through reviewing and protected by anti-abuse rules.
Results
The paper argues that sustainable ML peer-review improvement requires enforceable, fine-grained procedural safeguards combined with a spendable cross-conference credit system.
Takeaways & Limitations
The proposed system is intended to make good review practices earn credits that can be spent on conference perks or additional review resources.
Takeaways & Limitations
The paper lacks numerical experiment results and relies instead on hypothetical discussion and three case studies as a compromise for real-world anchoring.
Abstract
from arXiv · showhide
With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offsetting pressure of publication fees, the machine learning (ML) community has one of the largest scholarly presences among all scientific fields. And yet, \textbf{almost \textit{everyone} has \textit{many} unpleasant things to share about their review experience.} Worse, there is little public space to seriously discuss, let alone debate, what makes a review system effective or how it might be improved.\quad In this position paper, we expand our discussion from two core problems: \textit{How can we reasonably limit submission volume?} and \textit{How can we incentivize good and discourage bad reviewing?} We first assess the strengths and shortcomings of existing attempts to address such problems. Specifically, we present four takes on some popular conference mechanisms and propose two alternative designs for improvement.\quad Our general position is that meaningful improvement in ML peer review won't come from polite best-practice suggestions tucked into Calls for Papers or Reviewer Guidelines: it requires \textbf{enforceable yet fine-grained procedural safeguards} paired with \textbf{a currency-like credit system (e.g., our proposed \textit{OpenReview Points})}. ML practitioners can ``earn'' such points by contributing good review practices, and ``spend'' them across one or multiple major conferences to redeem different kinds of ``perks,'' such as complimentary registration or the right to request additional review resources.
1. Introduction
The paper argues that ML peer review cannot be sustainably improved through polite guidance alone. It instead requires enforceable, fine-grained safeguards paired with a spendable credit system across conferences.
- Contribution: Polite requests in Calls for Papers or Reviewer Guidelines are unlikely to improve peer review sustainably.The paper rejects optimistic guidance as sufficient for building a sustainable review ecosystem.
- Contribution: The paper identifies two core challenges and advocates enforceable procedural guardrails combined with an across-conference credit system.The credits are spendable, making them distinct from purely advisory review guidance.
- Motivation: ML peer review has scaled to tens of thousands of submissions at a single conference, alongside OpenReview and reciprocal reviewing obligations.These developments increase both the visibility of ML scholarship and the need to match review supply with demand.
1. How can we reasonably limit submission volume? · 2. How can we incentivize good and discourage bad reviewing?
The paper examines submission volume and review quality as root causes of unpleasantness in ML peer review, evaluating existing conference mechanisms and proposing enforceable safeguards alongside OpenReview Points. It advocates an adaptable direction to make review failures rarer, less painful, more accountable, and more sustainable rather than claiming to perfect peer review.
- 2. How can we incentivize good and discourage bad reviewing?: The paper identifies submission volume and reviewing quality as two root causes of widespread unpleasantness in ML peer review.It first develops background on these issues before assessing mitigation attempts used by ML conferences.
- 1. How can we reasonably limit submission volume?: It assesses existing attempts to mitigate submission-volume and review-quality problems across several ML conferences.The paper presents its own takes on these measures before introducing alternative mechanisms.
- 2. How can we incentivize good and discourage bad reviewing?: The paper proposes fine-grained procedural safeguards enforceable at scale to address shortcomings in ML peer review.These safeguards are presented as one of two new mechanisms for improving review processes.
- 2. How can we incentivize good and discourage bad reviewing?: OpenReview Points would let researchers earn and spend reviewing efforts in tangible ways across major conferences and review cycles.The proposed credit system is intended to provide a flexible mechanism for recognizing and using review contributions.
- 2. How can we incentivize good and discourage bad reviewing?: The proposed mechanisms are intended to address peer-review shortcomings effectively while remaining flexible enough for use across conference contexts.The paper presents this flexibility as an important reason the mechanisms could have a fair chance of helping.
- 2. How can we incentivize good and discourage bad reviewing?: The paper does not claim to perfect ML peer review, instead aiming to make its failures rarer, less painful, more accountable, and more sustainable.It explicitly rejects presenting its approach as a complete solution.
- 2. How can we incentivize good and discourage bad reviewing?: Rather than prescribing one rulebook for every conference, the paper advocates a general direction that organizers can explore and adapt to their needs.The proposed direction combines procedural and credit-based mechanisms without requiring uniform adoption.
2. Root Causes
ML peer review is strained by submission growth, limited reviewing capacity, and conference-level physical constraints that make workload allocation unsustainable. Weak oversight, feedback loops, metrics, and incentives also allow poor reviewing to persist without meaningful correction or improvement.
- Submission pressure: Overwhelming submission volume, driven by rapid field growth and increasingly accessible AI-assisted research, creates cascading demand for reviewers, ACs, and SACs.The resulting workload increases across successive levels of conference administration.
- Capacity constraints: In-person presentation guarantees for accepted papers impose physical capacity limits, potentially excluding borderline papers and creating unreasonable SAC assignments.The passage describes SAC-level content review as unsustainable given the number of submissions relative to SACs.
- Weak oversight: Reviewers, ACs, and SACs can submit dismissive, inconsistent, or low-effort assessments with little oversight unless their conduct triggers formal intervention.The system provides few mechanisms to correct or surface such behavior.
- Weak incentives: Without routinized feedback, transparent metrics, or positive incentives, peer review neither rewards exemplary stewardship nor deters poor practices.The lack of feedback also leaves participants with few avenues to learn how to improve.
3. Why Existing Fixes Fall Short
Existing fixes provide limited relief because submission caps do not address distributed submission incentives, mandatory reciprocal reviewing can degrade review quality, and desk rejection is too coarse for widespread poor reviewing. The paper instead argues for proportional enforcement and incentives that motivate capable, willing reviewers.
- Submission caps: Submission caps offer at best marginal relief because submission pressure is distributed across the community rather than concentrated among a few hyper-prolific lead authors.The authors suspect per-author quotas may trim auxiliary byline authors without reducing whether papers are submitted.
- Retaliatory sanctions: Desk rejection meaningfully enforces review safeguards but is too coarse to address the long tail of irresponsible behaviors short of abandonment.The paper calls for graduated, proportional penalties spanning the spectrum from infractions to severe violations.
- Reciprocal reviewing: 100% reciprocal reviewer recruitment assumes every eligible author can and will provide thoughtful reviews, but auxiliary contributors and specialized consultants may be ill-suited to general-purpose ML reviewing.Mandatory review also removes the ability to decline when contributors know they cannot meaningfully help.
- Reciprocal reviewing: Forced reviewing can produce rushed, templatized, or disengaged reviews as reviewers optimize for formally completing assigned duties rather than investing meaningful effort.The authors link this dynamic to reviewers’ limited motivation and bandwidth when participation is compelled.
- Alternative design: A workable alternative combines sensible opt-outs with accountability for reviewing obligations and additional willing reviewers to cover resulting gaps.The authors argue that better review depends on motivating people with bandwidth and genuine enthusiasm, including capable scholars who currently abstain.
4. Our Proposal: Fine-Grained Procedural Guardrails with a Currency-Like Incentive System
The proposal combines enforceable, fine-grained procedural safeguards with a community-wide, cross-conference OpenReview Points economy that lets participants earn, spend, and track credits. Organizers gain flexible incentives and penalties to reward good-faith contributions, deter abuse, and shape review behavior beyond blunt policies.
- Core design: The system pairs enforceable safeguards at different granularities with a community-wide, cross-conference credit economy, giving organizers both a “stick” and a “carrot.”The proposed economy is intended to provide incentives while preserving flexibility across conferences.
- Core design: OpenReview Points make review participation accountable by giving contributors something to earn, spend, and track.The current ecosystem relies on reviewers performing duties diligently under an honor system.
- Earning and spending: Review contributions earn points: a standard review earns 1 point, an emergency review 2, and “outstanding reviewer” recognition adds 3 points.These values are hypothetical and would require sophisticated balancing for a real economy.
- Earning and spending: Points can be spent on flexibility and privileges, including opting out of a review for 5 points, requesting an additional expert reviewer for 50, or redeeming free registration for 100.Authors can also spend 10 points to exempt a co-author from reciprocal reviewing obligations.
- Behavioral incentives: Because credits can be awarded or deducted for many behaviors, the system creates a broad design space for influencing participation beyond universal reciprocal reviewing or desk rejection.Point incentives could encourage extra effort, internal reviewer discussion, AC investigation, and detailed author feedback.
- Enforcement: Enforcement should deter fraud, point farming, and bulk low-effort reviewing through fine-grained, composable primitives rather than blunt desk rejection.Examples include role-based duty delegation, upper limits, voting-based penalties, and dynamic pricing; points must be earned through labor, not bought.
5. Alternative Views
The paper addresses concerns that a credit system could gamify peer review, create inequities and bureaucracy, or be misused. It argues that contribution-based credits, safeguards, controlled circulation, and broader participation can preserve accountability while improving review incentives.
- Gamification and bureaucracy: Credit systems may seem gamifying or bureaucratic, but peer review already operates through incentives, making the concern a question of design rather than novelty.The authors frame credits as an explicit version of incentives already embedded in peer review.
- Fairness and participation: Points are earned through labor rather than status, while exemptions can redirect incentives toward reviewers with the bandwidth and motivation to contribute quality reviews.The proposal also seeks greater participation from non-author researchers, who may have more bandwidth and fewer conflicts from their own submissions.
- Safeguards and penalties: Voting-based penalties require peer and area-chair agreement, appeals, and graduated deductions; positive credit awards can instead create feedback loops that help reviewers improve.The authors argue that limited false positives are acceptable because point deductions are less severe than desk rejection or submission bans.
- Credit economy design: Transfer restrictions preserve the link between points and personal service, while team-based redemptions remain possible; expiration and discounted redemption windows can limit inflation and hoarding.The authors distinguish points from private wealth, treating them more like records of community service or professional credentials.
- Adoption and perk allocation: Although full effectiveness requires adoption by multiple major conferences, existing reviewer perks could be allocated more objectively by ranking credits earned across papers or venues.The mechanism is presented as compatible with existing complimentary-registration resources rather than requiring conferences to provide more of them.
6. Recommended Practices / Call to Action
The paper calls for a flexible, conference-adaptable credit framework rather than a prescriptive rulebook. It recommends phased perk deployment, evidence-based point calibration, and shared measurement to support responsible iteration and interoperability.
- Flexible framework: Conferences should adapt the credit system to their own circumstances rather than follow a single prescriptive rulebook.The framework treats conference panels and authors as decision-makers who determine how credits and perks are offered and spent.
- Perk rollout: Early adopters should initially redistribute existing perks, enabling comparison with historical metrics and reducing risks from introducing new benefits.Existing perks include complimentary registration and emergency reviewer invitations; reuse also provides baseline data for testing new perks later.
- Perk rollout: New perks should be introduced gradually, with one-off tests offering faster evidence about whether incentives improve borderline-case resolution and the reviewer–AC pipeline.A one-off right for top point-earners to request an additional reviewer can expire at the end of the conference cycle.
- Point calibration: Point values should be calibrated from available perks and contributor percentiles, using one completed regular review as the ecosystem’s base unit.The paper avoids exact contribution values because insufficient empirical data currently supports setting them responsibly.
- Evaluation and interoperability: Conferences should coordinate point values, monitor key metrics, and publish post-conference statistics to improve implementations and encourage adoption of a shared credit currency.Reported measures can include reviewer–AC interaction and whether additional exchanges increase AC confidence in decisions.
7. Limitations
The proposal lacks numerical experiments and instead argues that review-mechanism discussions are best evaluated hypothetically. Its scope is largely corrective, leaving the root cause and fundamental scaling challenge of ML submission volume unresolved.
- 7. Limitations: The proposal’s lack of numerical experiment results is a limitation, though hypothetical analysis is presented as more meaningful than large-scale LLM-powered simulation.The authors argue that conference mechanisms cannot be rewound for A/B testing and that simulations across multiple conferences add little value because many factors compound.
- 7. Limitations: The proposal rationalizes an already-large submission pool rather than addressing the root cause of exploding ML submission volume.The authors characterize the work as largely corrective and reactive in scope.
- 7. Limitations: A credit system does not resolve the fundamental scaling challenge created by shifting review from expert committees toward committees of authors.The authors note that authors are often recruited by force as submission volume grows.
A. Related Works · B. Three Case Studies Where We Served as Reviewers
The paper situates its credit-based review proposal among mechanism, incentive, platform, and empirical studies, emphasizing enforceable currency-like systems over voluntary fixes. It then introduces three reviewer case studies as practical illustrations of how small reviewer initiatives can help.
- A. Related Works: Kim et al. (2025) similarly advocates reviewer feedback loops and rewards, while the Isotonic Mechanism compares author rankings with reviewers’ mean scores to trigger intervention when gaps are excessive.The Isotonic Mechanism may request an additional reviewer when author and reviewer assessments diverge too sharply.
- A. Related Works: Shah (2022) surveys peer-review problems across expertise mismatch, dishonesty, and miscalibration, proposing separate computational remedies tailored to each dimension.Examples include randomized assignment to mitigate dishonest bidding and machine learning to address commensuration bias.
- A. Related Works: Rogers & Augenstein (2020) and Zhang et al. (2022) examine incentive conflicts and persistent resubmission, considering policies such as better matching, more tracks, and track-specific formats.These works focus on diagnosing or modeling review and resubmission dynamics rather than presenting the paper’s credit-system mechanism.
- A. Related Works: Credit-based studies share the paper’s view that voluntary fixes fall short, but Francia et al. (2026) targets journal review delays whereas this work targets ill-suited reviews and quality.The settings differ because journals face reviewer scarcity and rolling queues, while ML conferences have comparatively ample reviewers and fixed, batched cycles.
- A. Related Works: Francia et al. (2026) mainly reduces rewards for late or weak reviews, whereas the paper seeks finer-grained defenses against subtle misconduct and malicious gaming.Under Francia et al., balances typically become negative only for an undelivered review; the paper argues that desk rejection is too coarse for long-tail misconduct.
- A. Related Works: Gasparyan et al. (2015) argues that no single financial or nonfinancial incentive is proven effective alone and supports combined rewards and reviewer credits.The paper differentiates its proposed perks by offering conference registration or privileges such as requesting additional review resources.
- A. Related Works: ReviewerCredits and Publons recognize review labor through durable, portable credit, but generally operate as voluntary third-party services rather than venue-linked decision mechanisms.The paper’s proposal ties credit to concrete stakes within conferences and shapes behavior after admission, unlike vouch’s access-control focus.
- B. Three Case Studies Where We Served as Reviewers: Because the position paper lacks real conference data, it presents three case studies from similar top ML conferences to illustrate how small reviewer initiatives can help.The authors state that these cases are not intended to highlight their own conduct.
B.1. Case 1: Reviewers criticizing matters outside the paper’s scope.
The case illustrates how reviewers can criticize benchmark papers for limitations outside their intended scope, and why rule-based policies alone cannot reliably identify severely low-quality reviews.
- Case 1: Reviewers criticizing matters outside the paper’s scope.: The paper was ultimately accepted after the authors raised concerns about the reviewer’s unreasonable, out-of-scope evaluation to the AC.The authors recommended disregarding the review or encouraging the reviewer to revise it.
- Case 1: Reviewers criticizing matters outside the paper’s scope.: A reviewer criticized the benchmark for low-performing methods, even though benchmarking is intended to reveal when methods fail.The authors argued that such performance is not a weakness of a dataset-proposing or benchmark paper.
- Case 1: Reviewers criticizing matters outside the paper’s scope.: The reviewer also criticized the benchmark for not capturing all possible real-world application scenarios.The authors viewed this as a boilerplate concern applicable to any dataset and analogous to criticizing a method for not evaluating every dataset.
- Case 1: Reviewers criticizing matters outside the paper’s scope.: Simple policies such as “inactive →desk rejection” fail to capture severely low-quality reviews, motivating finer-grained internal reviewer analysis.The authors argue that stronger safeguards are needed so positive impact scales beyond a few desk-rejected papers.
B.2. Case 2: Reviewers asking for particular experiments after the rebuttal deadline.
This case shows that reviewers raised potentially meritorious experiment requests after the rebuttal process had effectively failed, leaving authors without a meaningful response channel. The authors argue that enforceable deadlines and credit-based rewards or penalties could improve procedural fairness and elicit productive reviewer discussion.
- Case 2: Reviewers asking for particular experiments after the rebuttal deadline.: Reviewers identified two potentially meritorious but procedurally mistimed requests: a comparison against combined prior methods and an investigatory study of the proposed method.The requests were raised after the rebuttal deadline rather than through the designated exchange process.
- Case 2: Reviewers asking for particular experiments after the rebuttal deadline.: Only 2/4 reviewers provided the required exchange, with all engaged reviewers leaning positive, while one reviewer engaged substantively only late.The authors had submitted their initial rebuttal earlier, but the late substantive engagement limited the procedural value of the requested comparison.
- Case 2: Reviewers asking for particular experiments after the rebuttal deadline.: More than five comments exchanged with the reviewers contributed to the eventual decision, showing that internal discussion can help resolve borderline papers when reviewers receive a push to initiate it.The authors use this anecdote to motivate their credit system as a way to elicit productive discussions.
- Case 2: Reviewers asking for particular experiments after the rebuttal deadline.: Enforceable reply deadlines are necessary to preserve a meaningful rebuttal channel, while rewards and penalties are needed to support procedural fairness.The authors argue that reviewers could leave authors unanswered or reply late because such misconduct carries virtually no penalty.
B.3. Case 3: Authors presenting unsupportive results while claiming otherwise.
Internal reviewer discussion exposed results that contradicted the authors’ verbal claims, helping reviewers and the AC recognize weaknesses that contributed to rejection. The case also suggests that reviewer discussions are rare but can meaningfully shape paper evaluation when initiated.
- B.3. Case 3: Internal discussion revealed that the component’s dataset2 accuracy dropped significantly when trained only on dataset1, contrary to the authors’ verbal claims.The reviewer described this as a disadvantaged result.
- B.3. Case 3: After training on both datasets, dataset1 performance improved only slightly over the baseline while generating many more tokens.The cited review characterized this as a small gain accompanied by substantially higher output cost.
- B.3. Case 3: The system achieved only a small accuracy gain on its specifically trained task at higher cost, suggesting task-specific behavior that generalized poorly.The reviewer concluded that the added experiment made the work more negative despite appreciating the authors’ transparency.
- B.3. Case 3: The paper was ultimately rejected after reviewer discussion, with the AC citing concerns raised during that exchange.The reviewer exchanged views with the authors of the case and reached an agreement before the rejection.
- B.3. Case 3: Internal reviewer discussion appears rare despite its value, while initiating such exchanges can meaningfully facilitate evaluation and shape outcomes.The authors acknowledge lacking direct statistical evidence and describe the conclusion as based on three anecdotal case studies.