Caleb and Isaac found fabricated citations in most submissions to top ML venues, and audits now show those references are crossing into the published record.
When a reviewer for a top machine learning venue opened a recent submission and scrolled to the bibliography, several of the cited papers did not exist. The titles were there. The author names were there. The references were not. Over the rest of the season, that reviewer and one colleague worked through 22 submissions to NeurIPS 2025, WACV, and the TerraBytes workshop at ECCV. They flagged 15, about 68 percent, for fabricated citations, fabricated author lists, or text unmistakably written by a large language model. Their first-person notebook is now the most concrete on-the-record description of what review looks like when the review queue itself is being automated.
The two reviewers, who go by Caleb and Isaac, recorded per-track counts that held in every venue they touched. Per-reviewer totals: Caleb 6 of 11, Isaac 9 of 11. Per track: NeurIPS Position Paper 2 of 2, NeurIPS Datasets and Benchmarks 2 of 5, WACV combined 6 of 8, TerraBytes combined 5 of 7. They also released a bib-audit skill for automated citation verification, framed as a remediation tool for the next submission cycle.
That is the local evidence. The systemic evidence arrived at the same time. A Zhao et al. preprint scanning 2025 output across arXiv, bioRxiv, SSRN, and PubMed Central estimates roughly 146,900 hallucinated citations in a single year. A Lancet audit of 2.5 million biomedical papers found the share carrying at least one fabricated reference rose sixfold in two years, from 1 in 2,828 in 2023 to 1 in 458 in 2025, and 1 in 277 in early 2026. A Nature analysis in April described tens of thousands of 2025 publications as "probably" containing invalid AI-generated references. Pangram, an AI-detection vendor with a commercial interest in the headline, estimated 21 percent of ICLR reviews as AI-generated, a number to treat as one estimate among several rather than ground truth.
The mechanism is what makes this different from a cheating story. Ansari's 2026 forensic sample traced 100 hallucinated citations into 53 NeurIPS 2025 papers, about 1 percent of acceptances. Every one of those 53 passed three to five expert reviewers, and all 53 remain in the proceedings. The same dynamic shows up in Zhao et al.: roughly 85.3 percent of bioRxiv's hallucinated citations persist into the published versions of those papers. The slop is not staying in the rejected queue. It is crossing into the published record, where the next round of papers will cite it as if it were real.
The pattern has a name: citation-graph integrity debt. A fabricated reference does not stay a single bad citation. It gets inherited by a downstream paper, which gets cited in turn, and the original fabrication becomes a load-bearing node in someone else's argument. Volunteer reviewers, spending their own nights on a third or fourth review of the month, are not going to catch every one. The question is what the next layer of review looks like when the default assumption has to shift.
The venues are starting to answer. NeurIPS 2026 is running a formal AI-assisted review experiment with public documentation of how LLM assistance is permitted and disclosed. The ICML 2026 blog called out ongoing violations of its LLM review policy and named venues' obligations to enforce what they publish. The community discussion on Hacker News framed the same dynamic more bluntly: papers written by AI, reviewed by AI, read and summarized by AI, with the human academic loop increasingly ornamental.
The open question, for the field rather than the headline, is whether the response can match the persistence rate. Roughly 85 percent of bioRxiv's hallucinated citations survive into publication, which means the next round of reviews is inheriting fabricated citations as standard inputs. The bib-audit skill, the formal NeurIPS experiment, and the ICML policy blog are the constructive handles. None of them is the answer yet. The first honest step is to admit that the review queue is no longer the bottleneck, and that the published record is the next surface to defend.