Source-linked AI summary
The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
Chenguang Wang, Ming Li, Adebayo Braimah, Chenrui Fan, Tuo Wang, Weijie Guan, Ruiyi Zhang, Tianyi Zhou, Dawei Zhou
TL;DR
AI is changing both research production and scholarly evaluation, but treating them separately misses how each alters the other’s incentives, constraints, and behavior. The survey synthesizes 230 publications and institutional records through six connected dynamics, finding strong evidence for scaling, manipulation, and institutional response while later feedback remains less directly observed. It argues that scholarly AI should be studied as an adaptive ecosystem rather than as isolated capabilities.
Problem
Research on AI-generated research and AI-mediated review is often separated, despite the question of how changes on one side affect the other.
Method
The survey synthesizes 230 scholarly publications and institutional records using a descriptive taxonomy of six linked dynamics across the scientific ecosystem.
Results
The literature supports an emerging progression from production scaling and evaluation automation through manipulation and institutional response, while post-policy adaptation and long-horizon feedback are less directly observed.
Takeaways & Limitations
Understanding scholarly AI requires analyzing response relations and adaptation among authors, evaluators, institutions, and AI systems rather than isolated tools.
Takeaways & Limitations
Real-world evidence is concentrated in selected conferences, OpenReview settings, and journals, so findings may not transfer directly to disciplines with different review structures or norms.
Abstract
from arXiv · showhide
Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: changes on one side alter the incentives, constraints, and behavior of the other. We synthesize 230 scholarly publications and institutional records using a taxonomy of six connected dynamics: production scaling, evaluation automation, evaluation manipulation, defense mechanisms and policy responses, evasion and side effects, and long-horizon ecosystem feedback. The literature shows an emerging progression in which cheaper and faster research production increases pressure on evaluation, AI-mediated evaluation becomes more scalable and repeatable, participants can exploit evaluator regularities, and institutions respond with technical safeguards and policy controls. These responses can in turn induce evasion, redistribute errors and workload, and shape the scholarly records reused by future research and evaluation systems. Evidence is strongest for production and evaluation at scale, reproducible manipulation, and institutional response, while post-policy adaptation and artifact-level long-horizon feedback remain less directly observed. This systems view shifts attention from isolated AI capabilities toward how scholarly actors and AI systems adapt to one another over time.
1 Introduction
Generative and agentic AI are entering both scientific production and evaluation, linking the two through an emerging sequence of scaling, adaptation, and institutional response. This survey frames that sequence as six connected dynamics and finds strongest evidence for scaling, manipulation, and institutional response, with later feedback less directly observed.
- AI tools increasingly support research production from idea development and literature review through experimentation and manuscript preparation, while also assisting peer review and publication decisions.
- AI-enabled production can increase output and submission volume faster than human reviewing capacity, creating pressure to scale evaluation.Agentic systems extend this scaling toward integrated idea-to-paper workflows.
- AI-mediated evaluation is becoming more repeatable and observable, enabling hidden-instruction attacks and optimization toward more favorable assessments without stronger scientific evidence.
- Venues have responded with AI-specific policies, submission controls, detection mechanisms, and integrity safeguards.
- The survey organizes these developments into six linked dynamics and reports strongest support for production scaling, evaluation automation, reproducible manipulation, and institutional response.Post-policy counter-adaptation and long-horizon ecosystem feedback remain substantially less directly observed.
- The paper shifts analysis from individual AI tools to response relations through which scholarly actors adapt to one another over time.
2 Taxonomy
The taxonomy organizes AI-enabled scholarly publishing around six process categories and the response relations connecting them, rather than around technologies alone. It defines operational boundaries for production, evaluation, manipulation, defense, evasion, and long-horizon reuse across scholarly actors and workflows.
- Survey scope and actors: The survey defines system boundaries and institutional roles, covering research production, peer review, publication decisions, and the policies and safeguards shaping them.
- Survey scope and actors: The paper-review arms race denotes adaptive sequences in which an action changes another actor’s target or signal, prompting counter-response and shifting incentives, costs, or capabilities.
- Six categories in the taxonomy: The taxonomy comprises six process categories that describe how AI changes research and evaluation and how the ecosystem responds.
- Six categories in the taxonomy: Production scaling covers AI uses that reduce the time, cost, or effort of producing and revising research, from individual tasks to end-to-end workflows.
- Six categories in the taxonomy: Evaluation automation covers AI support or automation for manuscript assessment, review generation and revision, scoring, meta-review, and decision support.
- Six categories in the taxonomy: Evaluation manipulation comprises attempts to influence AI-mediated evaluation without improving underlying scientific evidence, including hidden instructions and exploitable presentation signals.
- Six categories in the taxonomy: Defense mechanisms and policy responses include technical, procedural, and institutional measures for reliability, integrity, accountability, detection, verification, control, disclosure, and AI-use governance.
- Six categories in the taxonomy: Evasion and side effects capture adaptations that avoid controls and unintended shifts in verification effort, false-positive risk, confidentiality concerns, or other ecosystem costs.
3 Scaling scholarly production
AI is making scholarly production faster, more scalable, and increasingly capable of supporting end-to-end research workflows. This expansion increases downstream evaluation pressure, while substantive revisions requiring new evidence or scientific judgment remain harder to automate.
- Production scaling: AI-assisted production reduces the time and effort required for research outputs, revision, and resubmission, allowing more outputs to be completed faster.Production scaling spans literature search, experimentation, writing, rebuttal, and increasingly integrated research workflows.
- Production scaling: AI-assisted writing is widespread, with corpus studies finding distributional changes consistent with substantial LLM-associated modification in scientific papers.The evidence indicates diffusion of AI-enabled writing tools across scientific communication.
- Deployment-scale systems: 166 complete papers across 67 AI/ML topics were produced in a public FARS deployment, but only 11.4% of outputs reached a mean score of six.Among 282 structured reviews of 140 outputs, the mean overall score was 3.17, and audits identified insufficient experimental evidence as a major weakness.
- Empirical evidence: Inferred AI adoption was associated with output increases of 36.2%, 52.9%, and 59.8% across arXiv, bioRxiv, and SSRN, respectively.The estimates vary across author backgrounds and fields, and adoption was inferred from text rather than directly observed.
- Empirical evidence: 15% higher productivity in 2023 and 36% in 2024 was associated with adoption in matched social and behavioral-science author panels.The study again inferred adoption from textual markers, and residual selection remained possible.
- Post-submission scaling: AI currently handles presentation-focused revisions better than revisions requiring new experiments, code, scientific reasoning, or evidence.Across 12 ICLR papers, model-generated revisions remained behind human camera-ready versions and sometimes introduced incomplete experiments, formatting errors, or fabricated results.
- From production to evaluation pressure: Submission volume increased by 42% at Organization Science after ChatGPT’s release, adding editorial and review load.At ICLR, estimated annual submission growth accelerated from 25% before ChatGPT to 59% afterward.
4 Automating scholarly evaluation
AI is expanding scholarly evaluation from author-facing feedback and reviewer assistance toward official reports and decision-adjacent workflows. Yet greater scale and authority expose persistent limits in scientific judgment, error detection, calibration, and review diversity.
- Authority ladder: Evaluation automation spans author feedback, reviewer assistance, official AI reviews, scoring, triage, meta-review, and publication decisions.These forms differ chiefly in the evaluative authority granted to the system.
- Author-facing feedback: 57.4% of 308 researchers rated GPT-4 manuscript feedback helpful or very helpful, although deeper methodological criticism remained difficult.The system’s feedback overlapped with human reviews of accepted papers and ICLR submissions.
- Author-facing feedback: More than 70% of respondents found a NeurIPS 2024 LLM checklist assistant useful, and a similar share intended to revise their paper or checklist.The deployment covered 234 voluntary submissions from 184 papers and remained outside formal publication decisions.
- Official AI reviews: 22,977 AAAI-26 main-track papers received clearly labeled AI reviews generated in less than 24 hours.The reports entered the official evaluation record but provided neither scores nor acceptance recommendations and did not replace human reviewers.
- Reviewer assistance: 6.5–16.9% of post-ChatGPT review text across four venues was estimated to be substantially AI-modified.This indicates that AI-assisted evaluation entered review practice before many venues established official systems.
- Selection outcomes: AI-assisted reviews were associated with a 3.1 percentage-point higher acceptance rate overall and a 4.9-point increase for borderline papers at ICLR 2024.The study estimated that at least 15.8% of reviews were AI-assisted and relied on inferred rather than directly observed AI use.
- Reliability bottleneck: Across automated reviewers, soundness-critical edits produced no statistically significant differences from surface-level controls, while irrelevant wording changes affected evaluations.The counterfactual study constructed 931 variants of 133 accepted AI and NLP papers.
- Reliability bottleneck: On SPOT, no tested model exceeded 21.1% recall or 6.1% precision for 91 validated errors from 83 published papers.In MLReplicate, automated review accepted 10 of 37 valid papers, and 59% of accepted papers contained fabricated or unsupported claims.
5 Evaluation manipulation
Evaluation manipulation ranges from explicit attacks on AI reviewers to ordinary-looking changes in presentation that alter judgments without changing scientific evidence. Repeated feedback can turn these vulnerabilities into an adaptive optimization process, although effects on official venue decisions remain limited.
- Explicit and multimodal attacks: Hidden instructions, layout changes, perturbations, and figures can influence AI-review outcomes through evaluator-visible signals.Attacks include concealed text, placement and formatting changes, character-, word-, and sentence-level perturbations, and figure-based attacks.
- Subtle presentation manipulation: Meaning-preserving rewrites can change evaluations without altering the underlying scientific content.Full-paper experiments reported a 75.1% success rate and a 1.21-point average score increase across three reviewer models.
- Subtle presentation manipulation: Presentation sensitivity can coexist with reduced diversity in evaluation, as AI reviews may be more similar to one another than human reviews.Rewriting papers without changing scientific content increased AI-review scores by an average of 0.45 points in one study.
- Scientific-content failures: Favorable judgments can also arise when plausible presentation is paired with unreliable or fabricated scientific content.BadScientist found that agent-generated manuscripts without real experiments could receive favorable LLM assessments under controlled conditions.
- Adaptive manipulation: Repeated evaluator feedback converts reviewer regularities into an optimization signal for refining instructions, wording, framing, layout, and other features.Iterative attacks can produce stronger transferable attacks, while repeated presentation revisions can target sections and framing that alter assessments.
- Evidence boundary: The evidence supports reproducible manipulation mechanisms, but direct evidence that they change official venue decisions remains limited.The synthesis identifies pressure for venues to protect evaluation reliability while distinguishing observed mechanisms from demonstrated institutional outcomes.
6 Defense mechanisms and policy responses
Defenses are shifting from passive AI-use detection toward targeted verification, adversarially tested evaluation pipelines, process evidence, and bounded institutional roles. Policies and safeguards are spreading, but their effectiveness and impact on behavior remain unevenly established.
- Detection limits: AI-use detectors are sensitive to segmentation and can misclassify permitted AI assistance, limiting their value for attributing how a review was produced.A venue audit changed the maximum AI-score share from 28.2% to 12.7% when moving from document-level to 100-word-window analysis.
- Targeted verification: Targeted verification introduces controlled signals to test whether a prohibited process occurred rather than inferring production history from text alone.Randomized markers appeared in 98.6% of tested LLM reviews on average, and watermarking has been used with manual verification and sanctions.
- Targeted verification: Watermarks, refusal markers, monitored redirection, and artifact-specific checks broaden defensive coverage but remain vulnerable to rephrasing, modality changes, and jailbreaks.These methods are effective against targeted behaviors, not necessarily against all forms of AI assistance or presentation-level manipulation.
- Robust evaluation: Static defenses can lose effectiveness against adaptive attackers, motivating evaluation pipelines that expose reviewers to progressively stronger attacks during training.SafeReview preserved paper ranking better than static adversarial training under evaluated attacker configurations and avoided some excessive blocking from generic detectors.
- Process verification: Process-level evidence can improve auditability beyond the submitted paper, with code, logs, and execution traces raising audit accuracy from 55% to 82%.These artifacts also exposed errors that were difficult to detect from the paper alone.
- Institutional responses: Venues increasingly combine disclosure, human verification, bounded AI roles, and enforcement while retaining an identifiable human decision point.Policies differ across conferences, publishers, and journals in permitted uses, disclosure requirements, and restrictions on external services.
- Evidence boundary: Current evidence documents the spread of AI policies more clearly than their effectiveness in changing author behavior.Comparative studies report similar growth in detected AI-assisted writing across venues with and without policies, while causal effects remain difficult to establish.
7 Evasion and side effects
Defenses create a moving target: participants can adapt around detectable signals, while safeguards may introduce false positives, unequal errors, confidentiality risks, and additional verification work. The full real-world sequence from deployed defense to evasion and institutional adjustment remains under-studied.
- Evasion: Detection traces can be altered through paraphrasing, rewriting, watermark modification, transferable substitutions, and detector-guided generation while preserving much of the original content.These evasion strategies target both post-generation traces and generation-time behavior.
- Evasion: Known verification mechanisms can prompt sanitization, rewriting, alternative document-processing routes, or model switching.Presentation and multimodal manipulation can move beyond the signals targeted by a particular defense.
- Evidence boundary: Direct evidence of participants adapting to deployed policies and triggering subsequent institutional responses remains limited.Existing studies rarely trace the complete sequence from a specific defense through behavioral change to adjusted safeguards.
- Unequal errors: False positives and uneven detector errors can burden legitimate human or AI-assisted writing, particularly for non-native-English writers and other populations unlike training data.These effects can arise even without strategic evasion.
- Unequal errors: LLM reviewers may respond to irrelevant author information, while detection and evaluation biases redistribute errors across different stages of review.Automation can therefore shift rather than eliminate unequal treatment and inconsistency.
- Labor and confidentiality: AI assistance may reduce drafting effort while increasing work for checking content, verifying evidence, resolving disagreements, investigating cases, and handling appeals.External AI services also create confidentiality concerns when unpublished materials leave venue-controlled infrastructure.
- Overall implication: Defenses should be evaluated for downstream costs and risks, not only for success against the original threat.The synthesis identifies false positives, unequal burdens, verification work, and confidentiality risks as consequential side effects.
8 Long-horizon ecosystem feedback
Scholarly papers, reviews, decisions, and citations increasingly persist as inputs, supervision, or optimization signals for later AI systems. This establishes pathways for long-horizon feedback, but complete artifact-level provenance from one scholarly cycle into later systems and decisions has not yet been demonstrated.
- Persistence: AI-associated material is entering searchable and published scholarship, establishing persistence of artifacts within the scholarly record.Corpus studies and audits report post-2022 changes in scientific writing and visible generated fragments in indexed publications.
- Error persistence: Fabricated or inaccurate references can persist through indexing and repetition, acquiring an apparent citation history and becoming harder to correct.Large audits identify nonexistent citations in published or publicly available research records.
- Reuse by AI systems: Published papers already feed pretraining and retrieval systems, making later scientific AI systems dependent on prior scholarly records.Documented examples include scholarly papers in OLMo’s training mixture and open-access retrieval for literature synthesis; memorized information can persist through adaptation.
- Record-to-system pathways: Reviews, rebuttals, decisions, and citations provide supervision or optimization signals for future evaluators and research-idea generators.SPARK filters generated ideas using OpenReview data, ReviewGuard trains future evaluation, and Scientific Thinker uses a citation-based evaluator as a reward model.
- Downstream influence: Controlled and adjacent studies show that persistent artifacts can influence later retrieval, training, knowledge graphs, and reasoning systems.A single malicious abstract substantially altered downstream drug-disease rankings, although equivalent effects from manipulated scholarly papers in deployed scientific systems remain unproven.
- Distributional feedback: Repeated reliance on AI may narrow future research questions, methods, and viewpoints while amplifying existing patterns and reducing diversity.These risks are supported by studies of AI reliance and repeated training on generated content.
- Evidence boundary: The central empirical gap is complete provenance linking a specific AI-mediated artifact to a successor system and then to a later paper, review, or decision.Record ingress and reuse are observed, but the full artifact-level lineage remains an open problem.
9 Cross-cutting findings and research agenda
The survey finds that trustworthy evaluation is harder to scale than research production, static evaluations weaken under participant adaptation, and safeguards redistribute rather than eliminate risk and workload. It therefore calls for longitudinal empirical studies of the response relations linking production, evaluation, manipulation, defense, evasion, and scholarly-record feedback.
- Cross-cutting findings: Trustworthy scientific evaluation remains harder to scale than research production because verification, flaw detection, novelty assessment, and accountable decisions remain costly.Review generation can scale, but reliable judgments still require human verification, disagreement resolution, and oversight.
- Cross-cutting findings: AI-led teams achieved about 37% reproducibility performance, compared with 94% for human-only and 91% for AI-assisted teams.In a randomized study of 103 teams, human-only teams also found more coding errors.
- Cross-cutting findings: Static evaluations of AI reviewers become less informative when authors can observe, query, and adapt to machine-mediated evaluation.Adaptation can target instructions, wording, presentation, or other evaluator-facing signals without improving the underlying scientific evidence.
- Cross-cutting findings: Defensive measures can motivate evasion while introducing false positives, verification work, confidentiality concerns, and unequal compliance burdens.Safeguards should be assessed by the errors, human labor, and participants affected after deployment, not only by detection or blocking performance.
- Research agenda: The research agenda prioritizes longitudinal studies of production pressure, author adaptation, post-policy evasion, and scholarly records entering successor retrieval or training systems.These studies would test the response relations represented in the taxonomy using real scholarly settings and versioned provenance.
10 Limitations
The survey’s evidence base is constrained by uneven and rapidly changing coverage, reliance on heterogeneous source types, and concentration of real-world evidence in particular publishing settings. Its arms-race framing is an organizing lens rather than a claim that all AI-assisted scholarly activity is adversarial.
- Coverage: Coverage of the newest literature may be uneven because the structured search ended July 1, 2026 and the targeted update did not rerun every query.Many relevant studies also remain available only as preprints.
- Evidence types: Surveys, position papers, policy documents, comments, replies, and conceptual analyses are not treated as evidence of prevalence, causal effects, or policy effectiveness without corresponding empirical data.These sources primarily characterize terminology, proposals, debates, and announced institutional responses.
- Generalizability: Real-world evidence is concentrated in a relatively small number of AI and computer-science conferences, OpenReview-based settings, and selected journals.Many attack, defense, and long-horizon feedback studies remain controlled experiments or early demonstrations, limiting direct transfer to other disciplines.
- Framing: The paper-review arms-race framing organizes adaptation and counter-adaptation without claiming that all AI-assisted research, reviewing, revision, or institutional change is adversarial.The framing is most relevant when one actor’s behavior meaningfully changes the incentives or responses of others.
11 Conclusion
The survey presents AI-mediated scholarly publishing as a connected process in which changes in one part reshape the incentives, constraints, and behavior of others. It concludes that evaluation should increasingly examine these interactions and adaptations in real settings and over time.
- Conclusion: The survey organizes AI-mediated scholarly publishing around six linked dynamics rather than isolated tools.The dynamics span production scaling, evaluation automation, manipulation, defenses and policy responses, evasion and side effects, and long-horizon feedback.
- Conclusion: Changes in one part of publishing can reshape the incentives, constraints, and behavior of other participants and systems.The conclusion shifts attention from individual system performance toward interaction and adaptation within the broader scholarly ecosystem.
- Conclusion: Future evaluation should study these relationships in real settings and over time as AI becomes more embedded in research and review.
A Supplemental details
The appendix consolidates the survey’s resources and studies into tables covering datasets, benchmarks, deployments, and taxonomy-organized systems.
- Supplemental details: Table A.1 summarizes datasets, benchmarks, and deployments discussed throughout the survey.
- Supplemental details: Table A.2 organizes the surveyed studies and systems using the taxonomy in Figure 2.
A.1 Datasets, benchmarks, and deployments
Table A.1 indexes the survey’s datasets, benchmarks, and deployments by resource type, scale, and primary use.
- The table organizes surveyed resources by resource type, scale, and primary use.
A.2 Taxonomy index
Table A.2 indexes the studies and systems according to the mechanisms in the taxonomy shown in Figure 2.
- The table organizes studies and systems by the mechanisms defined in the survey’s taxonomy.