Fake Citations and AI‑Generated Papers Flood Peer Review: Evidence, Impact, and Mitigation
The crisis in a nutshell
Two reviewers flagged fabricated citations and author lists in 22 submissions to NeurIPS, WACV, and the TerraBytes workshop; 68 % of the papers showed clear signs of LLM‑generated slop, and two papers with invented authors were still accepted for oral talks. The prevalence of hallucinated references is now documented across tens of thousands of publications, and AI‑generated reviews are becoming common.
Scope of the problem
| Venue | Reviewer | Flagged papers | Total reviewed |
|---|---|---|---|
| NeurIPS (Datasets & Benchmarks) | Caleb | 2 of 5 | 5 |
| NeurIPS (Position Paper) | Isaac | 2 of 2 | 2 |
| TerraBytes (ECCV workshop) | Caleb | 1 of 2 | 2 |
| Isaac | 4 of 5 | 5 | |
| WACV | Caleb | 3 of 4 | 4 |
| Isaac | 3 of 4 | 4 | |
| Total | Caleb | 6 of 11 (55 %) | 11 |
| Isaac | 9 of 11 (82 %) | 11 |
These numbers are a micro‑sample of a larger trend:
- Nature (April 2026) estimated tens of thousands of 2025 papers likely contain invalid AI‑generated references.
- Zhao et al. (arXiv 2026‑05‑07) counted ~146,900 hallucinated citations in 2025 across arXiv, bioRxiv, SSRN, and PubMed Central, with 85.3 % persisting into the published version.
- The Lancet (2026) found the share of papers with at least one fabricated reference rose from 1 in 2,828 (2023) to 1 in 277 (early 2026).
- Ansari (arXiv 2026‑02‑05) audited 100 fabricated citations from NeurIPS 2025; every citation passed 3‑5 expert reviewers, and 53 papers (≈1 % of acceptances) remain in the proceedings.
AI‑generated reviews are also widespread
- Pangram (2025) detected that 21 % of ICLR 2026 reviews were fully AI‑generated, with over half showing some AI involvement.
- Gartenberg et al. (2026) reported a 42 % post‑ChatGPT surge in submissions to Organization Science and that >30 % of its peer reviews now use AI.
- ICML 2026 identified 795 reviews (~1 % of all) that violated a “no‑LLM” policy.
- Li et al. (2026) showed adversarially rewritten abstracts can boost AI‑reviewer scores by +1.31 (Gemini 3 Flash) or +0.88 (GPT 5.4 Mini) in 38 % of attacks.
What reviewers observed
Worst offenders
Isaac: Two submissions swapped real authors for invented names; both were accepted for oral presentations on the condition that the references be fixed.
Caleb: A 53‑page NeurIPS paper filled with hallucinated jargon and nonsensical citations; another 40‑page paper contained embedded LLM notes next to each citation.
Relative severity of fake citations vs. fake authors
Both reviewers agreed that fabricated author lists are as damaging as entirely bogus references because they signal a lack of basic diligence and raise doubts about the entire manuscript.
Reliable “LLM‑written” tell‑tale signs
| Rank | Caleb’s tells | Isaac’s tells |
|---|---|---|
| 1 | Repetitive phrasing (“It’s not X, it’s Y”), bold formatting, em‑dashes | Metric mismatches between text and tables |
| 2 | Overly dense, hard‑to‑parse sentences | Verbose discussion sections that merely restate numbers |
| 3 | Over‑hyping results | Unusual use of colons/semicolons, marketing‑style language |
How reviewing workflows have changed
- Bibliography first: Reviewers now scan the reference list immediately after the abstract, often using the
bib‑auditskill. - Increased desk‑reject focus: Reviewers spend more time hunting for fabricated references to decide whether a paper is worthy of full review.
- Time allocation: Most reviewers report that a substantial fraction of their effort is now spent on detecting AI artifacts rather than evaluating scientific merit, raising sustainability concerns.
Mitigation tool: bib‑audit
bib‑audit is a Claude‑based plugin that:
- Parses
.bib,.bbl, or PDF bibliographies into structured fields. - Resolves each entry against Crossref, arXiv, DataCite, and Semantic Scholar.
- Flags, worst‑first, non‑existent works, invented authors, mismatched metadata, and formatting errors.
- Provides advisory warnings for entries lacking DOIs or arXiv IDs.
Installation (CLI)
claude plugin marketplace add isaaccley/skills
claude plugin install bib-audit@isaaccley-skills
Run the skill on a bibliography before submission; it can also be integrated into CI pipelines as a read‑only pre‑submission gate.
Policy landscape
- NeurIPS 2025‑2026: Allows AI‑assisted reviewing only through a sanctioned experiment; reviewers must not share submission excerpts with external LLMs.
- WACV 2026‑2027: Labels AI‑generated reviews as “highly irresponsible” and sanctions reviewers who break the rule.
- ECCV 2026: Prohibits any LLM use for writing reviews or sharing substantial submission material.
Community perspectives (selected HN comments)
"At this point papers are written by AI, reviewed by AI, and read by AI – we are automating humans out of the academic loop." – nneonneo
"Fake citations should be treated like plagiarism; otherwise the incentive structure remains broken." – DarkUranium
"If conferences required authors to upload reference PDFs, many hallucinations could be caught automatically." – leikarnes
"We need a desk‑reject mechanism or flagging tool; disclosure policies alone won’t stop bad actors." – Isaac
Who bears responsibility?
Both reviewers argue that authors are the primary culprits because they choose to submit low‑effort, AI‑generated manuscripts. However, they also note that advisors, program chairs, and conference organizers must enforce stricter desk‑reject criteria and provide tooling (e.g., bib‑audit) to protect reviewers from slop.
Takeaways for researchers and reviewers
- Never submit a paper with unverified references. Use automated bibliography audits and manual checks before upload.
- If you use LLMs, edit aggressively. Verify every generated citation, metric, and statement against original sources.
- Reviewers should treat fabricated citations as a ground‑for‑desk‑reject. Allocate a small, early‑stage bibliography scan to avoid wasted effort.
- Conferences need enforceable policies and tooling. Disclosure alone is insufficient; automated flagging and clear penalties are required.
References (human‑verified)
- Naddaf & Quill, Nature 2026 – Hallucinated citations polluting literature.
- Zhao et al., arXiv 2605.07723 – Large‑scale evidence of non‑existent citations.
- Topaz et al., The Lancet 2026 – Fabricated citations across 2.5 M biomedical papers.
- Ansari, arXiv 2602.05930 – Taxonomy of 100 fabricated citations at NeurIPS 2025.
- Emi, Pangram blog 2025 – 21 % of ICLR 2026 reviews AI‑generated.
- Li et al., arXiv 2606.10159 – Gaming AI‑assisted peer reviews.
- Gartenberg et al., Organization Science 2026 – AI surge in submissions and reviews.
- Kamath, ICML blog 2026 – Violations of LLM review policies.
- Lipton & Steinhardt, Queue 2019 – Early critique of “mathiness” in ML papers.
Tools and policies referenced
bib-auditplugin: https://github.com/isaaccley/skills/tree/main/plugins/bib-audit- John Owens’s bibliography error guide: https://www.ece.ucdavis.edu/~jowens/biberrors.html
- NeurIPS LLM policy (2025) and AI‑review experiment (2026)
- WACV reviewer guidelines (2026, 2027)
- ECCV 2026 reviewing policies