Source-linked AI summary

CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review

Jicheng Zhou, Kemou Li, Kahim Wong, Zheyuan Li, Zhuan Shi, Fengpeng Li, Haiwei Wu, Jiantao Zhou

arXiv:2609.05227v1cs.AI

TL;DR

Collusive bidding in peer review is difficult to study because intent is often unobserved and matched conference counterfactuals are unavailable. CABAL addresses this gap with fixed-environment LLM reviewer simulacra and affinity-guided collusion, finding substantially greater target-paper access and inflated evaluations while conference-wide effects remain modest.

  • Problem

    Collusive intent is typically unobserved, and real conferences lack matched counterfactuals for isolating how reviewer behavior affects assignments and evaluations.

  • Method

    CABAL holds the conference environment fixed while comparing honest and collusive LLM reviewer policies, using mutual affinities to construct expertise-consistent collusion rings and targets.

  • Results

    Coordinated bidding more than doubles assignment access, and colluders assigned to targets score them about two points higher than honest co-reviewers while aggregate effects remain comparatively modest.

  • Takeaways & Limitations

    CABAL provides a controlled testbed for tracing assignment-integrity risks and evaluating defenses against collusive bidding.

  • Takeaways & Limitations

    Bid-phase detectors face a precision-coverage trade-off: native bid graphs confound collusion with benign affinity, while stricter filtering recovers only a small subset of colluders.

Abstract

from arXiv · show

Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciprocal assignment advantage. Prior work treats bidding, reviewer assignment, and review manipulation as separate stages, leaving the lifecycle effects of collusive bidding unclear. Real-world analysis is further constrained by typically unobservable collusive intent and the lack of counterfactuals for the same conference. Motivated by this gap, we introduce \alg, an end-to-end multi-agent simulacra framework for studying reviewer assignment integrity by holding the conference environment fixed and configuring LLM-driven reviewer agents with honest or collusive policies. We further develop an affinity-guided collusive bidding strategy that uses mutual reviewer-paper affinities to construct collusion rings and select target papers, producing expertise-consistent rather than arbitrarily targeted attacks. Controlled experiments show that collusive bidding more than doubles target-paper capture and that assigned colluders score target papers about two points higher than honest co-reviewers, while conference-wide effects remain comparatively modest. Evaluated bid-phase detectors provide only limited evidence of collusion: in a fixed-triplet detector stress test, native positive-bid graphs are confounded by benign affinity, while a Very-High-only diagnostic view enables precise but low-coverage local recovery.

1 Introduction

CABAL addresses the limited understanding of how expertise-grounded collusive bidding propagates through reviewer assignment and downstream evaluation. It introduces a controlled, end-to-end framework that holds the conference environment fixed while varying reviewer behavior.

  • Motivation: Prior work studies matching, bid manipulation, collusion, and LLM-based review agents separately rather than connecting bidding, assignment, and downstream evaluation.
  • Motivation: Real-world analysis cannot cleanly isolate collusive effects because intent is typically unobserved and the same conference cannot be rerun under different reviewer behaviors.These constraints make expertise-consistent collusion difficult to distinguish from legitimate interest and prevent direct counterfactual attribution.
  • Framework: CABAL is an end-to-end multi-agent simulacra framework for studying reviewer assignment integrity and downstream evaluation under controlled behavioral worlds.
  • Method: CABAL uses mutual reviewer-paper affinities to construct expertise-grounded collusion rings and target papers, avoiding arbitrary or implausible targeting.
  • Framework: The framework holds papers, reviewers, COI constraints, and the assignment mechanism fixed while varying reviewer behavior between honest and collusive policies.
  • Scope: The study conducts an end-to-end empirical analysis of targeted access, downstream review scores, conference-level outcomes, and bid-phase detection patterns.

2 Problem Statement

The paper models peer-review integrity as resistance to coordinated bid manipulation across a sequential lifecycle. It compares honest and collusive behavioral worlds while keeping the conference substrate and assignment procedure fixed.

  • Problem formulation: The lifecycle maps bids B(z), affinity A, and COI constraints C through a fixed assignment procedure Φ to assignments M(z) and review scores Y(z).
  • Threat model: The threat model considers reviewer-authors with overlapping expertise who coordinate bids to increase reciprocal assignment likelihood without bypassing COI or eligibility constraints.
  • Integrity: Assignment integrity is compromised when coordination gives colluders greater access to one another’s submissions than honest behavior would provide.
  • Integrity: The analysis distinguishes direct assignment capture from downstream review distortion caused when successful assignments alter evaluations.
  • Counterfactual design: Each collusive world is paired with an all-honest world sharing papers, reviewers, affinities, COI constraints, and the assignment procedure.

CABAL: Collusive Agent-based Bidding and Assignment Laboratory

CABAL implements a controlled peer-review workflow that derives expertise-grounded collusion structures, instantiates honest or collusive reviewer simulacra, and traces behavior from bidding through evaluation. Matched worlds separate assignment effects from strategic reviewing effects.

  • Conference substrate: CABAL fixes a conference substrate and derives collusion rings and target papers from reviewer expertise, authorship, affinity, and COI structure.
  • Affinity construction: A reviewer’s eligible high-affinity pool contains non-COI papers compatible with that reviewer’s expertise profile.
  • Affinity construction: The mutual-affinity graph connects reviewer-authors only when compatibility holds reciprocally, before bids are generated.
  • Ring construction: Greedy clique expansion forms rings in which every reviewer pair satisfies the mutual-affinity condition.
  • Target construction: Targets contain papers authored by other ring members that already lie within each reviewer’s eligible high-affinity pool.
  • Reviewer simulacra: Reviewer personas keep conference context, expertise, and rubric fixed while changing only honest or collusive behavioral objectives.
  • Strategic behavior: Collusive reviewers prioritize target papers during bidding, while aggressive and subtle strategies differ only in downstream target evaluation.
  • Lifecycle analysis: The workflow preserves B(z) → M(z) → Y(z), enabling separate analysis of assignment access and downstream evaluation.

4 Experiments

CABAL evaluates collusive bidding through matched honest and collusive worlds, tracing effects from assignment access to review scores and conference-wide outcomes. Collusion substantially increases target access and inflates captured-paper evaluations, while broader quality alignment changes remain modest.

  • Experimental setup: The experiments use a fixed conference instance, seven runs, and matched behavioral worlds that vary reviewer behavior while preserving papers, reviewers, reference quality, affinities, COI constraints, and matcher settings.The lifecycle analysis proceeds from bidding to assignment, then review outcomes and conference-wide effects.
  • Assignment access: Target-paper capture is measured using targeted assignment rate for support relations and target-paper capture rate for papers receiving at least one designated supporter.The same support relations are evaluated in collusive and matched all-honest runs.
  • Assignment access: 41.4-44.1 percentage-point TAR gains and 42.2-48.9-point TPCR gains raise assignment access from honest to collusive bidding, more than doubling overall access.TAR rises from 18.3-22.7% to 62.4-64.1%, while TPCR rises from 28.0-29.6% to 71.8-76.9%.
  • Target-paper score inflation: 2.21 and 1.99 points are the mean within-paper score differences between assigned supporting colluders and honest co-reviewers at r = 0.2 and r = 0.5.Target papers gain 0.61-0.69 points overall, with captured targets gaining 0.84-0.89 points and uncaptured targets 0.03-0.04 points.
  • Conference-wide effects: 0.13 and 0.38 points are the conference-wide average-score increases at r = 0.2 and r = 0.5, while above-threshold reviews increase by 2.4 and 6.0 percentage points.These localized effects accumulate into broader score inflation as the colluding population grows.
  • Conference-wide effects: Quality alignment degrades modestly: Pearson correlation falls from 0.804 to 0.775, while top-30 overlap remains between 22.3 and 23.0 papers across conditions.The results indicate limited conference-wide disruption rather than wholesale changes in top-paper membership.

CABAL-generated Collusion be Detected?

The evaluated bid-phase detectors provide limited and uneven evidence of CABAL-generated collusion. Native positive-bid views are confounded by benign affinity, while stricter views improve localization but recover only a small portion of the distributed attack.

  • Evaluation setup: The detector stress test uses one fixed triplet with an all-honest world and two collusive worlds containing 18 and 47 colluders in 8 and 20 rings.It evaluates reviewer rankings, dense bid-author groups, and dense reviewer-paper blocks.
  • Detector results: Ranking methods provide little attack-induced separation, while Pairwise Reciprocity largely rediscovers reciprocity already induced by honest affinity.Low-rank Residual produces only small, encoding-dependent shifts.
  • Detector results: On native positive-bid graphs, OQC and TellTail change little from the honest reference, while Densest and Fraudar obtain broad coverage by flagging 77-101 of 140 reviewers.At τ = 2, TellTail recovers one complete ring with precision .753/.874 but recall .178/.126.
  • Interpretation: The evaluated detectors do not cleanly recover a distributed attack spanning eight or 20 rings, although CABAL’s strongest bids expose local structure.The conclusion applies to the evaluated detector-input combinations rather than general undetectability.

6 Conclusion

CABAL traces expertise-grounded collusive bidding through assignment and downstream evaluation under matched behavioral worlds. The experiments find strong target-level access and score effects, comparatively modest aggregate effects, and limited detector evidence, motivating CABAL as a controlled testbed for future defenses.

  • Findings: Coordinated bids substantially increase colluders’ access to target papers and inflate evaluations when assignment capture succeeds, while aggregate conference-wide effects remain comparatively modest.The dominant effects are concentrated on successfully captured targets.
  • Detection: Native bid graphs often confound collusion with benign expertise-driven affinity, whereas stricter filtering improves localization but recovers only a small subset of colluders.This establishes a precision-coverage trade-off for the evaluated detectors.
  • Implication: CABAL provides a controlled testbed for studying assignment-integrity risks and evaluating future defenses against collusive bidding.The framework supports matched behavioral comparisons under a fixed conference environment.

CABAL

CABAL combines reviewer-agent simulation with synthetic conference artifacts to study peer-review dynamics under controlled conditions. Its design grounds reviewer expertise in public profiles while fixing reference-quality assessments across runs.

  • Prior work: CABAL builds on work in reviewer matching, bid manipulation, collusion detection, and LLM-based review agents.The paper positions these research strands as complementary but previously insufficiently integrated.
  • Research gap: Existing agent-based peer-review systems largely begin after reviewer access is determined, leaving strategic bidding and endogenous assignment jointly unmodeled.CABAL addresses this gap by connecting bidding, assignment, and downstream reviews.
  • Conference construction: The reviewer pool contains 140 Semantic Scholar profiles spanning cs.LG, cs.CV, cs.CL, and cs.AI.The profiles provide research areas and representative publications for reviewer-agent grounding.
  • Conference construction: The submission set contains 100 synthetic papers distributed across 5 strong accept, 25 accept, 40 borderline, and 30 reject conditions.It includes single-author and same-area or cross-area dual-author papers, with papers and profiles fixed across behavioral worlds.
  • Reference-quality assessment: Reference-quality scores align strongly with preset conditions, with Pearson/Spearman correlations of 0.832/0.828.The continuous scores and derived categorical labels are computed once and fixed across experimental runs.

B.4 Bidding and Assignment Implementation

The implementation uses affinity-weighted bidding and constrained greedy matching over a fixed synthetic conference substrate. Collusion rings are formed from mutual affinity, while integrity checks and matched runs support downstream comparisons.

  • Bidding and matching: Reviewers bid on their top-20 non-COI papers using five ordinal labels encoded as Bpj values from −100 to 2.The Very Low label is treated as a hard refusal, while the matcher combines affinity and bid utility.
  • Bidding and matching: Each paper receives three eligible reviewers, with reviewer loads capped at five papers and author-reviewers required to serve at least three.Minimum-service obligations receive priority before affinity-bid utility ranking.
  • Collusion construction: Collusion is restricted to 92 author-reviewers, with realized means of 18.43 ± 0.49 and 46.71 ± 0.70 colluders at r = 0.2 and r = 0.5.Complete rings are retained, so realized counts can differ slightly from nominal quotas.
  • Collusion construction: Rings contain two or three members selected as cliques in the mutual top-20 affinity graph.Targets are restricted to ring-authored papers that also lie in each reviewer’s eligible top-20 pool.
  • Experimental controls: The seven-run experiment fixes papers, profiles, quality assessments, affinities, COI constraints, and matcher hyperparameters.Independent LLM calls and structural seeds vary behavioral artifacts while preserving the conference substrate.
  • Assignment effects: Collusive bidding significantly increases targeted assignment access in every run at both r = 0.2 and r = 0.5.Exact McNemar tests give p < 0.05; the authors treat these as supplementary paired checks alongside the primary assignment-access evidence.
  • Review effects: +2.21 and +1.99 points are the within-paper score gaps between supporting ring members and honest co-reviewers at r = 0.2 and r = 0.5.Aggressive and subtle strategies both favor target papers, although their supporter-score levels differ.

C.4 Metrics for Conference-Wide Effects

The conference-wide analysis measures review-score and quality-alignment effects, then stress-tests bid-phase detectors under limited observational views. Native positive-bid graphs are confounded by honest affinity, while Very-High-only views recover small structures with limited coverage.

  • Metrics: Conference-wide metrics are computed separately for each run and behavioral condition, then summarized by mean and standard deviation over seven runs.The metrics include review-score aggregates, favorable-review prevalence, quality correlations, and top-30 overlap.
  • Metrics: The favorable-review fraction measures individual evaluations above an acceptance-level threshold, not the conference’s final paper acceptance rate.This distinction limits how the metric should be interpreted.
  • Detector setup: The detector evaluation uses only bidding and authorship information, without colluder identities, ring memberships, targets, assignments, or review scores.Detector outputs are compared with realized collusion structure only after ranking or suspicious-set generation.
  • Scope and limitations: Detector findings characterize the tested fixed triplet and detector representations rather than general or deployment-level undetectability.Separate LLM-generated worlds do not hold non-target bids identical, and CABAL distributes colluders across multiple small rings.
  • Native positive-bid view: Native positive-bid graphs recover only about two of 18 colluders at r = 0.2 and about 12 of 47 at r = 0.5 for several ranking methods.Their honest-to-collusive overlaps are nearly unchanged or decrease, indicating that much of the graph structure predates the attack.
  • Native positive-bid view: Greedy Densest Subgraph recovers 16 of 18 and 42 of 47 colluders only by flagging 89 and 77 of 140 reviewers.High recall reflects a broad background graph core rather than precise ring localization.
  • Detector limitations: Fraudar’s native view recovers most colluders while flagging 97 and 101 reviewers, yielding recall .833/.872 but weak localization.The Very-High-only view improves visibility of target bids but reaches global recall .333/.191 with precision .167/.429.

D.7 What the Experiment Reveals

The experiments show that collusive signals are fragmented and representation-dependent: native positive-bid views are confounded by benign affinity, while Very-High-only views recover local rings precisely but incompletely.

  • Native bid representations: Native positive-bid graphs confound collusive structure with benign expertise-based affinity.Mutual expertise affinity lets colluders legitimately bid on one another’s papers, and binary thresholding merges ordinary and strongest bids.
  • Very-High diagnostic view: Very-High-only bids enable precise local recovery but do not provide high global coverage.The τ = 2 view lets TellTail and OQC-Greedy recover at least one complete small ring with high precision, while remaining an attack-informed diagnostic projection.
  • Distributed attacks: 18 or 47 colluders are distributed across eight or 20 rings, creating a precision-coverage trade-off for single-block detectors.Methods isolating one ring achieve high precision but low recall, whereas broad coverage includes much of the reviewer pool.
  • Overall interpretation: The evaluated detectors do not cleanly separate the distributed collusive population from benign affinity structure.The conclusion is limited to the fixed bidding triplet, tested input adaptations, and detector-level restarts rather than independent conference replications.

E Discussion

The discussion argues that auditing must account for the attack pathway and that assignment is an integrity-critical intervention point. It recommends defenses that reduce collusive access while preserving legitimate expertise-based bidding.

  • Detection: Positive bids and reciprocal bid-author patterns are insufficient evidence of collusion when rings are built from mutual expertise affinity.Colluders target papers on which they could plausibly express interest under honest behavior.
  • Detection: Detection should be pathway-aware because native bid graphs can confuse attack-induced structure with benign affinity.The detector comparison motivates conditioning auditing on how collusive behavior generates its observable footprint.
  • Assignment integrity: Coordinated bids can substantially increase target-paper access, making assignment an integrity-critical mechanism before downstream review effects emerge.The proposed response is to act before or during assignment rather than relying only on post-hoc review anomalies.
  • Evaluation scope: CABAL enables controlled comparisons of manipulation resistance, assignment quality, reviewer workload, and false-positive risks across defenses.This supports evaluating alternative assignment mechanisms and auditing procedures within a common controlled environment.

F Limitation

CABAL is a controlled stress-testing framework rather than a faithful replica of every real-world conference. Its synthetic agents, fixed configuration, and lightweight assignment mechanism limit direct generalization to human operations and deployment.

  • Evaluation setting: CABAL uses synthetic submissions, LLM-based reviewer simulacra, reference assessments, and a fixed conference configuration with lightweight assignment.These choices enable controlled comparisons across behavioral worlds but simplify real conference processes.
  • Scope boundary: The evaluation does not fully capture human reviewer heterogeneity, operational assignment systems, or broader strategic behavior.The limitation concerns both the modeled participants and the assignment environment.
  • Interpretation: The findings concern propagation through the modeled bidding–assignment–reviewing pathway, not real-world prevalence or guaranteed deployment detectability.Future work proposes additional model families, human quality assessments, richer representations, and production-oriented assignment algorithms.
  • Future work: Future extensions should cover discussion, rebuttal, meta-review, and final decision stages to assess broader downstream effects.Such extensions are intended to strengthen external validity and support comparisons under more diverse operational conditions.

G Broader Impacts

CABAL supports reproducible defensive stress testing of peer-review vulnerabilities, but its outputs also create misuse, false-positive, privacy, and governance risks. The paper therefore limits its intended use to defensive evaluation with human oversight.

  • Potential benefits: CABAL may help organizers identify vulnerabilities, compare manipulation-resistant assignment mechanisms, and develop auditing procedures without experimenting on an active conference.Its broader value is reproducible research on peer-review security.
  • Risks: The simulations could be misused to refine collusive strategies or evade existing detectors.This creates a dual-use risk for releasing attack artifacts and operational details.
  • Risks: Automated detection can falsely implicate legitimate reviewers in small or closely connected research communities.Natural reciprocal expertise and bidding patterns may resemble collusive behavior.
  • Mitigations: CABAL should be used as a defensive stress-testing tool, not as an operational basis for accusing or sanctioning individuals.Deployment should protect confidential data, combine evidence sources, audit performance across communities, and retain human oversight.
Loading 2609.05227v1…