Source-linked AI summary
Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025
Samar Ansari
TL;DR
AI-generated fabricated citations can evade expert peer review, but the mechanisms behind this failure are insufficiently characterized. This study analyzes 100 verified hallucinations from NeurIPS 2025, develops a failure-mode taxonomy, and finds that compound deception is pervasive, motivating multi-attribute citation verification.
Problem
Fabricated citations to nonexistent sources can evade review by 3–5 experts, while current peer review does not systematically verify citation existence.
Method
The study analyzes 100 GPTZero-identified and human-verified hallucinated citations from 4,841 NeurIPS 2025 papers using a failure-mode taxonomy.
Results
100% of hallucinations exhibited compound failure modes, with Semantic Hallucination in 63% and Identifier Hijacking in 29% of citations.
Takeaways & Limitations
Effective prevention requires automated verification that cross-checks authors, titles, venues, dates, and identifiers rather than relying on link-checking alone.
Takeaways & Limitations
The analysis covers only citations flagged by GPTZero and cannot determine author intent; its category distribution may not generalize beyond NeurIPS 2025 or AI/ML.
Abstract
from arXiv · showhide
Large language models (LLMs) are increasingly used in academic writing workflows, yet they frequently hallucinate by generating citations to sources that do not exist. This study analyzes 100 AI-generated hallucinated citations that appeared in papers accepted by the 2025 Conference on Neural Information Processing Systems (NeurIPS), one of the world's most prestigious AI conferences. Despite review by 3-5 expert researchers per paper, these fabricated citations evaded detection, appearing in 53 published papers (approx. 1% of all accepted papers). We develop a five-category taxonomy that classifies hallucinations by their failure mode: Total Fabrication (66%), Partial Attribute Corruption (27%), Identifier Hijacking (4%), Placeholder Hallucination (2%), and Semantic Hallucination (1%). Our analysis reveals a critical finding: every hallucination (100%) exhibited compound failure modes. The distribution of secondary characteristics was dominated by Semantic Hallucination (63%) and Identifier Hijacking (29%), which often appeared alongside Total Fabrication to create a veneer of plausibility and false verifiability. These compound structures exploit multiple verification heuristics simultaneously, explaining why peer review fails to detect them. The distribution exhibits a bimodal pattern: 92% of contaminated papers contain 1-2 hallucinations (minimal AI use) while 8% contain 4-13 hallucinations (heavy reliance). These findings demonstrate that current peer review processes do not include effective citation verification and that the problem extends beyond NeurIPS to other major conferences, government reports, and professional consulting. We propose mandatory automated citation verification at submission as an implementable solution to prevent fabricated citations from becoming normalized in scientific literature.
1 Introduction
The paper examines fabricated citations that entered NeurIPS 2025 despite expert peer review and situates them within a broader problem of LLM-generated misinformation. It develops a failure-mode taxonomy to explain how these citations evade detection and motivates systematic verification.
- The problem: 53 NeurIPS 2025 papers, approximately 1% of accepted papers, contained fabricated citations despite review by 3–5 expert researchers.The citations referred to non-existent sources with invented authors, titles, or publication details.
- Why it matters: Fabricated citations undermine scholarly evidence, reproducibility, and the citation graph by creating unverifiable claims and false linkages.Readers may search for non-existent papers, while future research inherits contaminated references.
- Broader scope: The problem extends beyond NeurIPS to ICLR submissions, government reports, and consulting outputs involving LLM writing workflows and inadequate verification.Reported cases include over 50 ICLR hallucinations, corrected government reports, and AUD 98,000 in consulting refunds.
- Research gap: The study addresses limited systematic analysis of how AI-generated fabricated citations exploit weaknesses in peer review.It asks whether hallucinations differ in detectability and which failure modes allow them to pass review.
- Approach: The authors analyze 100 hallucinated citations by classifying how each deviates from legitimate scholarly practice and succeeds at evading detection.The taxonomy focuses on mechanisms rather than merely labeling citations as fake.
- Contribution: Secondary characteristics were dominated by Semantic Hallucination at 63% and Identifier Hijacking at 29%, often accompanying Total Fabrication.These combinations create plausible wording and false verifiability through links to unrelated papers.
2 Methods
The study analyzes GPTZero-identified citations from NeurIPS 2025 and manually codes each verified hallucination using a five-category taxonomy. Primary and secondary codes capture both the dominant failure characteristic and additional deception mechanisms.
- Data source: GPTZero scanned 4,841 of 5,290 accepted NeurIPS 2025 papers using web, database, and DOI/URL checks, followed by human verification.The dataset included citations judged probable hallucinations rather than archival or indexing anomalies.
- Data source: The resulting dataset contained 100 verified hallucinated citations across 53 papers, each previously reviewed by 3–5 expert reviewers.The analysis used the complete citation text, paper title, and GPTZero diagnostic notes.
- Taxonomy: The taxonomy distinguishes Total Fabrication, Partial Attribute Corruption, Identifier Hijacking, Semantic Hallucination, and Placeholder Hallucination.Categories separate wholesale invention from strategic corruption or hijacking of real scholarly metadata.
- Coding: Each citation received a primary code for its most prominent failure characteristic and a secondary code for additional deception mechanisms.For mixed cases, the primary code reflected the feature most evident to a human reviewer attempting verification.
- Coding: The author manually classified all citations over three days in a structured spreadsheet recording category codes, reasoning, and GPTZero evidence.The procedure used observable features rather than subjective interpretation.
3 Results
The primary-code results are dominated by Total Fabrication, while the remaining categories are much less frequent. Representative examples span detection difficulty from obvious placeholders to sophisticated identifier hijacking.
- Primary failure modes: 66% of the 100 hallucinated citations were Total Fabrications, indicating wholesale invention rather than corruption of real metadata.Partial Attribute Corruption accounted for 27% of cases.
- Primary failure modes: Identifier Hijacking accounted for 4%, Placeholder Hallucination for 2%, and Semantic Hallucination for 1% of primary codes.These categories were less common but potentially more sophisticated failure modes.
- Detection difficulty: The examples range in detection difficulty from trivial placeholder text to identifier hijacking that creates false verifiability.They illustrate how different hallucination types exploit aspects of citation verification.
3.2 Total Fabrication (TF): Complete Invention
Total Fabrication is the dominant primary failure mode, comprising 66% of hallucinated citations and involving citations invented wholesale. Examples show that professionally formatted, technically plausible references may have no corresponding publication.
- Definition: Total Fabrication citations have no correspondence to real scholarly work, including fabricated authors, titles, venues, and identifiers.They represent complete invention rather than alteration of an existing source.
- Example: A citation with generic authors, a plausible journal, and fabricated DOI and URL was entirely invented despite professional formatting.Verification found no matching article in the cited journal volume.
- Example: A technically coherent event-camera title sounded appropriate for its field but corresponded to no ICASSP 2023 publication.The same pattern recurred with domain-appropriate titles that had no real scholarly source.
- Prevalence: 66% of hallucinated citations were Total Fabrications, making wholesale invention the predominant failure pattern.This exceeds the share of Partial Attribute Corruption, which accounted for most remaining cases at 27%.
- Propagation: One fabricated citation later appeared verbatim in another paper’s references before being corrected in later versions.This suggests the hallucination may not have originated with the NeurIPS author.
3.3 Partial Attribute Corruption (PAC): Strategic Blending
Partial Attribute Corruption combines real and fabricated citation elements, exploiting familiar authors or partial recognition to appear legitimate while failing detailed verification. Examples show that incorrect bibliographic details and corrupted author or venue information can evade superficial checks.
- Partial Attribute Corruption combines real and fabricated elements, creating citations that appear legitimate initially but fail detailed verification.
- Familiar authors can mask incorrect titles, volumes, issues, and page numbers when reviewers verify only author lists or publication years.The cited authors collaborated on a real 2020 paper, but the listed bibliographic details were incorrect.
- Corrupted author lists and publication venues can misrepresent real papers while recognizable names reduce reviewer scrutiny.The cited paper omitted two authors, added one, and was published at ICLR 2024 rather than EMNLP 2023.
3.4 Identifier Hijacking (IH): False Verifiability
Identifier Hijacking uses valid scholarly identifiers that resolve to real papers whose metadata does not match the citation. This creates false verifiability because a working link can pass a basic check while failing substantive content verification.
- Identifier Hijacking provides valid arXiv IDs or DOIs that link to real papers despite mismatched citation metadata.
- A valid arXiv identifier can conceal a completely different title and author list from a reviewer who checks only whether the link works.
- All 4 primary Identifier Hijacking cases involved valid identifiers pointing to unrelated papers.
3.5 Placeholder Hallucination (PH): Generation Failures
Placeholder Hallucinations expose incomplete AI-generated citations through unfilled template text and missing identifiers. Despite their obviousness, two such citations appeared in published NeurIPS papers without basic verification.
- Placeholder Hallucinations occur when a model fails to complete citation generation.
- “Firstname Lastname” and “URL or arXiv ID to be updated” reveal unfilled author and identifier fields.
- Two placeholder citations appeared in published NeurIPS papers despite their obvious incompleteness.
3.6 Semantic Hallucination (SH): Plausible Fabrication
Semantic Hallucinations invent conceptually appropriate titles that fit a research context but correspond to no actual publication. Their plausibility can derive from domain knowledge, recognizable researchers, and professionally phrased topics.
- Semantic Hallucinations invent conceptually appropriate titles that do not correspond to real papers.
- A citation about mechanistic interpretability can sound professionally appropriate while pairing a known researcher with a nonexistent specific title.
- The failure mode fits the research context closely while corresponding to no actual publication.
3.7 Distribution Across Papers
Hallucinations were concentrated in papers with minimal AI use, but a small group showed extensive reliance, including one paper with 13 fabricated citations.
- 49 of 53 contaminated papers contained 1–2 hallucinations, while 4 outlier papers contained 4–13 hallucinations.The distribution had a mean of 1.89, a median of 2, and a range of 1–13 hallucinations per paper.
- 13 fabricated citations appeared in the paper with the highest hallucination count, demonstrating that citation fabrication can be extensive rather than isolated.The paper’s pattern suggests reliance on AI tools throughout the writing process rather than occasional consultation.
3.8 Compound Failure Modes: A Critical Finding
Every hallucination exhibited multiple failure modes, with semantic plausibility and false verifiability layered onto fabricated or corrupted citation metadata. This compound structure defeats isolated checks and requires cross-verification of the complete citation.
- 100% of hallucinations exhibited compound failure modes involving multiple deception mechanisms simultaneously.The classification assigned primary and secondary characteristics to reveal these layered patterns.
- 66% of citations were Total Fabrications, while Semantic Hallucination appeared as a secondary characteristic in 63% of all citations.Among Total Fabrications, 50 cases (76%) used semantic plausibility to make invented titles appear domain-appropriate.
- 29% of cases included Identifier Hijacking as a secondary characteristic, typically layered onto Total Fabrication or Partial Attribute Corruption.This pattern indicates that working links were generally used alongside other deception mechanisms rather than alone.
- A fabricated citation can combine invented metadata with a valid but mismatched identifier, creating both semantic plausibility and false verifiability.The example uses a real arXiv identifier attached to different authors and a different title.
- Reliable detection requires confirming that authors, title, venue, date, and identifier all match the same actual source.Checking only one attribute can miss citations whose individual components appear partially legitimate.
4 Discussion
The analysis finds that fabricated citations evade elite peer review through compound deception, combining plausible semantics, familiar metadata, and misleading identifiers. It also identifies systematic scope, proposes automated verification, and acknowledges limits in detection, classification, and generalizability.
- Compound failure structure: 100% of analyzed hallucinations combined multiple failure modes, most commonly Total Fabrication with Semantic Hallucination.The dominant compound pattern pairs wholesale invention with plausible terminology or titles.
- Compound failure structure: 63% of citations used Semantic Hallucination as a secondary characteristic, exploiting plausible titles that pass superficial “sounds right” checks.This pattern provides semantic plausibility despite zero verifiable content.
- Compound failure structure: 29% used Identifier Hijacking as a secondary characteristic, linking to real but unrelated papers and creating false verifiability.A working identifier can pass a link check while the citation’s metadata remains mismatched.
- Implications for peer review: Comprehensive cross-verification must confirm that authors, title, venue, date, and identifier refer to the same source, but this is not standard peer-review practice.The evidence indicates that multiple expert reviewers missed even obvious placeholders and errors.
- Generative mechanism: Two-thirds of hallucinations involved wholesale invention rather than corruption of real sources, contradicting a retrieval-and-modification hypothesis.The analysis attributes this pattern to next-token prediction rather than retrieval of bibliographic data.
- Scope and limitations: The problem extends beyond NeurIPS to ICLR submissions, government reports, and consulting outputs, while its observed distribution may not generalize across domains or tools.The authors also report taxonomy ambiguity and acknowledge that the dataset comes from one conference and one discipline.
- Proposed intervention: The paper proposes mandatory automated citation verification at submission using existing tools and APIs.The proposed checks are intended to verify citation details before fabricated references enter the literature.
5 Conclusion
The study finds that fabricated citations evade peer review through compound failure modes and argues that automated, multi-attribute verification is an implementable response. It identifies institutional verification infrastructure, rather than individual reviewers, as the central gap requiring action.
- 66% of hallucinations were Total Fabrications, while 63% incorporated semantic plausibility and 29% incorporated identifier hijacking.The secondary characteristics often accompanied wholesale fabrication to create false verifiability.
- Current peer review lacks systematic citation verification, allowing even trivially detectable placeholder errors to pass unnoticed.Reviewers prioritize methodological rigor, novelty, and experimental results rather than confirming that cited sources exist.
- 100% of documented hallucinations exhibited compound failure modes that exploit multiple verification heuristics simultaneously.These fabrications combine features such as semantic plausibility, familiar names, and working links.
- Effective detection requires simultaneously checking citation existence, metadata consistency, identifiers, and semantic plausibility.Simple link-checking is insufficient because compound hallucinations can include working links and plausible metadata.
- The proposed response is mandatory automated citation verification at submission, matched to the documented compound failure structure.The recommended process combines existence checks, metadata matching, identifier validation, and human review of semantically suspicious citations.
- The authors frame the primary obstacle as institutional inertia and warn that relying on existing peer review risks normalizing fabricated citations.They characterize the problem as a failure of verification infrastructure to keep pace with AI capabilities, not an indictment of individual reviewers or authors.