Source-linked AI summary
Looking for Bidding Teammates: A Game-Theoretic Model of Stranger Collusion in Peer Review
Jinming Xing, Charlotte Brian
TL;DR
Open recruitment lets strangers form peer-review collusion despite prior models assuming trusted colleagues. The paper develops a four-stage game and calibrated conference simulation, finding that one-shot reciprocation unravels, persistence requires enforcement, and current detectors struggle against adaptive attacks.
Problem
Prior peer-review collusion models assume established, trusted groups and do not explain how strangers recruited online form and sustain reciprocal bidding arrangements.
Method
The paper combines a four-stage extensive-form game of recruitment, information exchange, hidden bidding, and reviewing with calibrated end-to-end conference simulations.
Results
F1 = 0.322 is the highest detector performance tested against camouflage, ring distribution, and affinity manipulation, while a two-person arrangement adds 3.2 percentage points to acceptance probability.
Takeaways & Limitations
Reciprocal inflation is unsustainable in the one-shot game but can persist through observed reviews, repeated deadlines, and reputation, making enforcement rather than payoff magnitude decisive.
Takeaways & Limitations
The simulation leaves several parameters unconstrained, and its best-response assumption models an upper-end adversary rather than a typical colluder.
Abstract
from arXiv · showhide
Paper bidding is the entry point to reviewer assignment at large CS conferences: reviewers declare interest, combined by an optimizer with automated affinity scores. Reviewers who have never met recruit each other online, exchange identifiers, bid on each other's papers, and reciprocate with inflated scores. Existing collusion models assume already-trusted colleagues; open recruitment removes that assumption and the mechanism that made such arrangements work. We give the first game-theoretic model of collusion \emph{formation} in peer review, the \emph{Mutual Bidding Dilemma}: a four-stage game covering recruitment, exchange of identifiers under risk of being reported, unverifiable bidding, and reciprocal reviewing. The model predicts the arrangement cannot form: once assigned a partner's paper, writing the inflated review is pure cost, since the benefit depends on the partner's own decision. Reciprocation is never individually rational, for any payoffs, and the arrangement unwinds. What closes the gap is enforcement, not incentives: authors see their own reviews, deadlines recur every few months, and the group remembers who reciprocated. We derive the condition under which inflation is sustainable, the detection rate above which no partnership survives, and show effort enters both. On a calibrated end-to-end conference simulation, no detector we test exceeds $F_1 = 0.322$ against an attacker who camouflages bids, spreads them around a ring, and manipulates affinity; a two-person arrangement is worth $3.2$ points of acceptance probability; and the harm is distributional, not aggregate: $70$ honest papers are displaced while mean quality moves by only $0.002$, so no summary statistic reveals it. Randomized assignment is the one defense reaching enforcement itself, making a partner who never bid indistinguishable from one who bid and lost.
I. INTRODUCTION
Open recruitment turns peer-review collusion into a formation problem: strangers publicly seek partners, exchange identifiers, manipulate bids, and rely on reciprocal inflated reviews. The paper models this previously unaddressed process and its downstream assignment effects.
- I. INTRODUCTION: Bids influence optimizer-based assignment alongside automated affinity scores, making cheap, unverified declarations consequential when positive bids are sparse.The optimizer combines bids with affinity, and an eager bid can outweigh a sizable affinity deficit.
- I. INTRODUCTION: Public recruitment threads show strangers soliciting reciprocal bidding partners for named conferences and repeating the practice each submission cycle.Responders sort by subject area, disclose sting risk, and sometimes state the reviewing-stage dilemma directly.
- I. INTRODUCTION: Prior collusion models take trusted groups as given, whereas this paper models how strangers form a collusive relationship.Open recruitment removes shared history, institutional bonds, and prior expectations of future contact.
- I. INTRODUCTION: Four formation problems distinguish stranger collusion: first-mover exposure, sting risk, unverifiable bidding, and the difficulty of sustaining reciprocal reviewing.The first identifier reveals authorship and intent, while private bids make non-assignment compatible with either cooperation or betrayal.
- I. INTRODUCTION: The paper contributes a four-stage formation model and a calibrated simulation measuring acceptance effects, displaced honest papers, and defenses.The game covers recruitment, information exchange, hidden bidding, and reciprocal reviewing.
B. Manipulation, mitigation, and detection
The paper frames peer-review manipulation as a sequential game whose hidden bidding and observable reviewing create different enforcement problems. It contrasts existing defenses and derives an unraveling result for the one-shot interaction.
- B. Manipulation, mitigation, and detection: Existing mitigations either cap pairwise assignment probability, forbid short cycles, or analyze bidding and text-matching signals, while collusive groups are generally taken as given.Cycle-free constraints can be evaded by larger rings, and randomized assignment trades assignment quality for reduced matching probability.
- B. The four-stage game: The Mutual Bidding Dilemma is a four-stage extensive-form game covering recruitment signaling, identifier exchange under sting risk, hidden bidding, and reciprocal reviewing.Bids are private and assignment occurs only probabilistically, creating a confounded signal when a partner is not assigned.
- B. The four-stage game: At Stage 4, honest reviewing strictly dominates inflation for every parameter setting because the inflated-review benefit depends on the partner’s separate action.The Stage 4 game is a Prisoner’s Dilemma, with mutual honesty as its unique Nash equilibrium.
C. Payoffs
The payoff analysis separates the value of entering collusion from the incentive to reciprocate after assignment. It concludes that enforcement, rather than payoff magnitude alone, determines whether stranger collusion can persist.
- C. Payoffs: Under mutual cooperation, collusion is worth entering only when the per-round net gain uC = g − w exceeds recruitment costs cs + µπs.The expected payoff includes the gain from reciprocal inflation and the costs of signaling and possible sanction.
- C. Payoffs: Increasing the value of publication raises the incentive to enter collusion but does not make reciprocation individually rational once assigned.The reviewing benefit cancels from the comparison because it depends on the partner’s action, while effort and sanction remain costs.
- C. Payoffs: In the one-shot benchmark, the unique equilibrium is no recruitment, no identifier exchange, no collusive bidding, and honest reviewing for all b, ρ, q, and π.The result assumes cs > 0, cb > 0, and κ + ce > 0.
- C. Payoffs: Observed collusion therefore requires enforcement outside the one-shot stage game, such as monitoring and repeated interaction.The paper identifies reviews returned to authors, recurring deadlines, and reputation as the relevant ecosystem features.
B. Where the enforcement comes from
Repeated interaction and observable betrayal can enforce reciprocal inflation, but hidden non-bidding creates a noisy signal that makes tolerance rules fragile. Randomized assignment lowers the chance of matching colluders and thereby weakens enforcement against free-riding.
- B. Where the enforcement comes from: Reviews returned to authors make betrayal observable, while recurring conference deadlines create a repeated relationship summarized by δr.Under grim trigger, cooperation continues only when no observed betrayal has occurred.
- B. Where the enforcement comes from: Mutual cooperation is sustainable under grim trigger exactly when the discounted continuation benefit covers the effort cost of inflation.The one-shot deviation saves review effort while retaining the partner’s already-committed benefit for that round.
- B. Where the enforcement comes from: δ⋆ increases with d, π, ce, and cb, but decreases with b and ρ.Lower effort cost expands the stable region because ce → 0 lowers the sustainability threshold.
- B. Where the enforcement comes from: Hidden non-bidding is never observed directly, so enforcement must punish non-assignment despite its occurring with probability 1−ρ under honest cooperation.This turns Stage 3 into a noisy-monitoring problem.
- B. Where the enforcement comes from: A k-strike rule is feasible only when kmin ≤ kmax; as ρ → 0, the required tolerance diverges while the admissible tolerance shrinks.Randomization therefore weakens deterrence by making honest partnerships harder to distinguish from free-riding.
E. Who reveals first, and when it is worth posting
Recruitment is governed by the continuation value of a genuine partnership: strangers reciprocate when that value exceeds the reporting payoff, and public posting becomes worthwhile once enough readers make a successful match likely. Enforcement ultimately depends on detection and assignment probabilities that determine whether cooperation can persist.
- E. Who reveals first, and when it is worth posting: At Stage 2, a responder reciprocates rather than reports if and only if the partnership continuation value W is at least R.The first mover reveals an identifier only when the expected continuation value offsets the risk of being reported.
- E. Who reveals first, and when it is worth posting: The effective lever for stranger recruitment is W, because increasing the continuation value lowers the required responder cooperation probability.The sanction L saturates because it appears in both terms of the responder’s comparison.
- E. Who reveals first, and when it is worth posting: n⋆ increases with monitoring intensity μ and the sanction πs attached to an attributed post.An initiator posts when the audience is large enough that expected recruitment gains exceed posting costs.
- E. Who reveals first, and when it is worth posting: Public recruitment groups exceed the posting threshold at plausible response rates, while raising μ is the model’s only pre-bidding intervention.The result links platform-scale exposure directly to the feasibility of recruiting strangers.
- E. Who reveals first, and when it is worth posting: Mutual cooperation is sustainable only below a detection threshold, and no discount factor sustains collusion once d > (b −ce −cb/ρ)/π.The boundary trades publication benefit against detection probability at a rate determined by sanctions and relationship durability.
V. ATTACK STRATEGY TAXONOMY
The attack taxonomy expands from naive bidding to camouflage, distributed rings, affinity manipulation, and increasingly subtle reviewing styles. The evaluation combines these strategies against detector families in a calibrated simulator, where each detector is defeated by a corresponding cheap countermeasure, while direct measurement is limited by unavailable conference data.
- V. ATTACK STRATEGY TAXONOMY: The adversary strategy space includes naive, camouflaged, distributed-ring, and affinity-aware bidding choices.These strategies manipulate bid concentration, graph structure, topical affinity, or the defensibility of assignments.
- V. ATTACK STRATEGY TAXONOMY: Camouflage makes collusive bids statistically ordinary by embedding them among genuine bids on topically adjacent papers.The strategy defeats frequency and density tests at the cost of selecting cover papers.
- V. ATTACK STRATEGY TAXONOMY: A distributed ring replaces a two-cycle with a directed cycle, keeping each member’s collusive bid count below frequency-test thresholds.The trade-off is greater coordination and a longer monitoring chain.
- V. ATTACK STRATEGY TAXONOMY: Affinity-aware attacks raise assignment probability and make collusive matches appear defensible by aligning bids with manipulated text-based similarity.The optimizer combines affinity and bids, so a strong match can mask the collusive signal.
- V. ATTACK STRATEGY TAXONOMY: Detector families target bid frequency, dense subgraphs, cycles, and similarity, but combined countermeasures leave colluders indistinguishable from honest reviewers in the simulation.The simulator generates submissions and reviewers, computes noisy affinity, injects strategic collusion, solves assignment, and models inflated scores.
- V. ATTACK STRATEGY TAXONOMY: The evaluation relies on an end-to-end simulator because conferences release neither bidding data nor assignment logs or affinity scores.Calibration uses publicly available counts, loads, score distributions, acceptance rates, and simulated malicious-bidding patterns.
B. Experiment 1: equilibrium validation
The calibrated population simulation validates the model’s comparative statics and quantifies attack, downstream, and defense outcomes. Collusion becomes easier as effort falls, evasion remains effective, and randomization most strongly reduces the combined strategy while preserving high honest affinity.
- Equilibrium validation: The measured detection transition occurs at d = 0.079, below the closed-form threshold d⋆ = 0.113 because congestion raises effective detection.The measured q transition is q = 0.627 versus q⋆ = 0.571 because congestion lowers W.
- Equilibrium validation: At machine-assisted effort ce = 0.08 b, 91.9% of the population sustains collusion, compared with 68.2% at hand-written effort and 95.2% at ce = 0.The optimal ring size increases from m⋆ = 4.6 to m⋆ = 8.6 across the same effort range.
- Detection evasion: No detector family exceeds F1 = 0.322 against any strategy; distributed rings reduce cycle-detection F1 from 0.156 to 0.032.Bid-frequency detection scores 0.000 against every dyadic strategy after topical screening.
- Downstream impact: A bare dyad raises a participant’s acceptance probability from 25.0% to 28.2%, a gain of 3.2 percentage points, while ring size 4 peaks at 4.8 points.The peak reflects competition for review slots and added detection exposure beyond the interior optimum.
- Downstream impact: At a 5% colluder fraction, the combined strategy displaces 70 honest papers while mean accepted quality falls by only 0.002.The two quality distributions remain visually indistinguishable because swapped papers cluster near the acceptance threshold.
- Defense effectiveness: Randomization reduces the combined strategy’s success from 0.673 to 0.311 while retaining 96.8% honest affinity, and makes the enforcement-feasible set empty.Cycle-free assignment barely changes the combined strategy, from 0.673 to 0.665, while retaining 100.0% optimal honest affinity.
VII. DISCUSSION
The discussion identifies detection, assignment probability, monitoring, sanctions, and effort as distinct levers, but prioritizes transparent randomized assignment and credible monitoring. The paper’s simulation-based estimates support threshold directions more strongly than absolute magnitudes.
- A. Levers available to a program chair: Detection matters only above d⋆ = 0.113, while current detector evaluations remain at or below F1 = 0.322.The discussion recommends evaluating detection against the threshold rather than relative improvement below it.
- A. Levers available to a program chair: Assignment probability ρ is prioritized because it directly affects sustainability and separately prevents colluders from policing one another.The discussion argues that venues should publish a per-pair cap so non-assignment can undermine enforcement expectations.
- A. Levers available to a program chair: Monitoring public recruitment channels can deter participation if the policy and sanctions are credible, even without catching every post.Monitoring acts before a collusive bid exists and raises the threshold for everyone reading a recruitment post.
- A. Levers available to a program chair: Effort cost is difficult to control because author-drafted reviews can reach ce ≈ 0 without machine assistance.Sanction magnitude and reported-first-mover loss have saturating effects in the model.
- Practical implications: The paper recommends randomized assignment with an explicit published per-pair cap, plus monitoring and multi-signal detection.The proposed detection system fuses bidding, scoring, assignment, and authorship signals because each individual signal is evadable.
- Scope and limitations: The quantitative conclusions rely on simulation because venues do not release the bidding, assignment, or affinity data needed for direct measurement.The simulation constrains observable conference properties but leaves several behavioral and payoff parameters unconstrained, emphasizing thresholds and comparative statics.
APPENDIX A PROOFS
The proof solves the four-stage game by backward induction and shows that honest reviewing strictly dominates inflation after assignment, causing collusive bidding, information exchange, and recruitment to unravel.
- Stage 4: H strictly dominates I after assignment because inflation costs κ + ce in either case, independently of the partner’s action and publication benefit b.The unique Nash equilibrium is (H, H), and its dominance argument does not depend on b.
- Stage 3: With honest reviewing anticipated, bidding on a partner’s paper yields expected benefit ρ·0 but costs cb, so not bidding is strictly optimal.This conclusion holds at every Stage 3 information set and does not depend on beliefs about whether the partner bid.
- Stage 2: Because a formed partnership has zero continuation value, revealing an identifier is strictly worse than withdrawing whenever q < 1 and L > 0.At q = 1, revealing is only weakly optimal and still produces no collusive bid.
- Stage 1: Posting is strictly worse than not posting because its expected payoff is bounded above by a negative signaling and sanction cost.The initiator’s payoff is at most [1 −(1 −λ)n] · 0 −cs −µπs < 0.
- Conclusion: Each backward-induction step is strict, so the profile is the unique subgame-perfect equilibrium for all values of b, ρ, q, and π.The result establishes that enforcement, rather than payoff magnitude, constrains formation.
E. Proof of Proposition 6
The enforcement analysis derives when repeated collusion survives noisy assignment signals and how partnership formation and ring size respond to continuation value, loss risk, effort, and detection.
- Enforcement: k ≤ ln δ⋆ / ln δr bounds the number of tolerated strikes, with k = 1 recovering the one-strike theorem.The bound follows from requiring δr ≥ δ⋆ under the k-strike rule.
- Enforcement: False termination requires k ≥ ln ε / ln(1 −ρ), because cooperation can produce k consecutive non-assignments by chance.The rule fires on a run of k non-assignments, whose probability is (1 −ρ)^k.
- Enforcement: As ρ →0+, both the minimum strike threshold kmin and the sustainability threshold δ⋆ diverge, so cooperation fails outright.Lower assignment probability simultaneously worsens false termination and raises the cost term cb/(ρb).
- Formation: A responder reciprocates iff W ≥ R, while a first mover reveals an identifier iff q ≥ q⋆ = L/(W + L).The threshold q⋆ falls as partnership continuation value W rises and increases as reporting loss L rises.
- Ring size: The optimal ring size is m⋆ = α/(2ρπβ), decreasing with effort cost ce and detector sensitivity β and increasing with publication benefit b.In practice the operative choice is min{Li, ⌊m⋆⌉}.
I. Proof of Theorem 10
The proof calibrates the simulation’s population and attack construction, then formalizes how heterogeneous reviewers, topical affinity, bid strategies, and camouflaging shape collusive rings.
- Simulation setup: The simulation sweeps game parameters while calibrating venue-scale parameters to published conference statistics, leaving several assumptions unconstrained by public evidence.Table V distinguishes calibrated venue parameters from swept game parameters and assumed quantities.
- Population generation: Papers and reviewers are generated over 50 subject areas with specialized reviewer mixtures and author-linked paper topics, creating over-bid and under-bid regions.Reviewer and submission popularity vectors are drawn separately, so reviewer supply does not track submission volume area by area.
- Affinity and quality: Affinity combines cosine similarity with Gaussian noise, while latent quality θp is independent of topics so displacement can be measured against ground truth.The simulator detects authorship conflicts automatically but does not observe latent quality in real datasets.
- Bidding: Honest bidding is probabilistic and capped at 8 eager bids, with participation calibrated to sparse real-venue bid concentration.The targets produce a mean near six eager bids per paper and roughly a quarter of papers receiving two or fewer.
- Ring formation: Rings require high mutual affinity and are grown from admissible candidates, preserving topical adjacency while modeling strangers rather than random pairings.Uniform random pairing would understate collusive strategies because distant bids lose on affinity and appear conspicuous.
- Attack strategies: Camouflaging adds r = 4 cover bids ranked by reviewer affinity plus similarity to targets, while affinity-aware attacks add 0.15 to targeted affinity.The additive affinity boost is a modeling limitation because it fixes the gain rather than deriving it from permissible text edits.
D. Assignment
The assignment pipeline solves a sparsified integral matching problem, implements randomized and repaired assignment variants, generates reproducible reviews, and evaluates detector families under fixed investigation budgets and parameter sensitivity.
- Assignment: The assignment linear program retains the 30 highest-scoring reviewers per paper plus every bid pair, excluding authorship conflicts.Keeping every bid pair prevents sparsification from silently removing the collusive attack being measured.
- Assignment: Coverage requires k = 3 reviews per paper and load permits Li = 5, with total unimodularity guaranteeing an integral optimum.HiGHS returns the optimum, which is read as an assignment when x > 0.5.
- Assignment variants: Randomized assignment samples an integral assignment from fractional marginals with a per-pair cap of 0.35, while repair can perturb feasibility and marginals.The cap is enforced approximately rather than exactly.
- Review simulation: Honest review scores are reproducible because noise is a deterministic hash of the seed, reviewer, and paper rather than sequential random draws.This keeps honest scores identical between collusive and baseline runs that share the generated conference.
- Detection: Detector families score bid isolation, dense reviewer-author subgraphs, directed cycles, and affinity anomalies, then threshold at a common 1 − FPR quantile.The fixed investigation budget makes detector families comparable without assuming an intrinsic score scale.
- Sensitivity analysis: Across spread assumptions, d transitions remain within [0.076, 0.081] and q transitions within [0.573, 0.650], while ce transitions range over [0.280, 0.472].The effort transition is most sensitive to dispersion in b because effort enters as ce/b.
B. Detector operating points
Detector performance depends sharply on the operating point for naive attackers but remains weak against combined camouflage and distributed strategies. The simulation also shows that acceptance-rate changes scale displacement without removing colluders’ advantage, while randomized assignment targets the enforcement constraint.
- Detector robustness: 0.228 is the best F1 across all four detector families and five budgets against the combined strategy, below the 0.322 result in Table IV.The distributed strategy remains lower because longer rings produce no counted cycle and camouflage leaves no frequency signal.
- Detector operating points: Naive bidding rises from F1 = 0.156 at 5% to 0.567 at 10%, because only 6.8% of honest reviewers carry any two-cycle.The 5% convention falls just below the threshold discontinuity; examining 10% helps naive and camouflaged dyads but not the described sophisticated strategies.
- Acceptance rate: 2.4 percentage points at 15% and 4.7 at 40% are the acceptance boosts, while displaced honest papers rise from 39 to 68.A larger accepted set creates more decisions to disturb, but colluding submissions retain an advantage at both rates.
- Audience size: At 1% response, the minimum audience is n⋆ = 4.9 readers, falling to 0.2 at 20%, so audience size does not restrain public recruitment.The binding condition is Γ > cs + µπs, making venue monitoring intensity the model’s pre-bid lever.
- Defense implications: Randomized assignment lowers ρ and can make a rule unable to deter free-riding while preserving honest partnerships, directly degrading colluders’ monitoring.The feasibility frontier is plotted over (ρ, δr), and the infeasible region widens as δr falls.
- Detector operating points: The bid-frequency family reaches precision 1.000 against distributed and combined strategies, but recall is only 0.041 and 0.068.Its F1 values are 0.079 and 0.127, illustrating why precision alone can conceal sparse detection.
- Camouflage: As cover bids per collusive bid increase from 1 to 8, cycle-detection F1 rises from 0.205 to 0.416 while collusive bid share falls from 0.258 to 0.078.Camouflage adds incidental graph edges and enlarges the bid neighborhood, but assignment success falls from 0.296 to 0.232.
C. Cycle detection beyond two
Cycle detection weakens as the allowed cycle length grows because longer cycles are common in honest bidding and distribute attackers’ signal. Temporal dispersion remains an unevaluated boundary because existing detectors ignore bid order and venues do not publish timestamps.
- Cycle bounds: A three-member ring improves from F1 = 0.032 at bound 2 to 0.248 at bound 3, then falls to 0.144 at bound 4.Matching the bound does not guarantee detection because longer bounds admit more honest cycles.
- Cycle bounds: Honest reviewers carrying a cycle rise from 0.068 at bound 2 to 0.180 at bound 4, reducing separability under a fixed false-positive budget.Longer cycles become ordinary in honest bidding, diluting the collusive signal.
- Cycle bounds: At bound 4, a six-member ring still reaches only F1 = 0.095 while enumeration time rises from 0.01 to 0.61 seconds per instance.Longer rings cost attackers one additional recruited partner but impose growing computational and false-positive costs on the venue.
- Temporal dispersion: The four evaluated detector families cannot measure temporal dispersion because they read bid sets, graphs, or affinity outcomes, all invariant to bid arrival order.A timing-based detector would require a calibrated honest-bid arrival process, but venues publish no bid timestamps.
- Scope and ethics: The evaluation uses public recruitment threads as evidence that the phenomenon exists and as motivation, not as input to the analysis.The collection involved no private messages, covert group entry, or interaction with posters.
- Disclosure: The attack taxonomy includes stronger strategies than prior detector evaluations because camouflage and ring topologies are openly discussed and affinity manipulation is published.The paper frames this disclosure as necessary for evaluating defenses against a specified adversary.
- Model limitations: The model fixes a pair across rounds and a single venue, while treating detection probability d as exogenous despite adaptive adversaries.Cross-venue collusion and the detector–strategy fixed point are discussed but not modeled.
- Future work: Future work includes multi-signal detection, adaptive adversaries, cross-venue collusion, expertise–familiarity tradeoffs, and stylometric analysis of author-drafted reviews.Bid timing and bid volume are also identified as untested signals requiring additional data or calibrated arrival processes.