NeurIPS, the largest annual machine learning research conference, is testing LLM assisted peer review; one reviewer describes a detailed review next to two superficial ones, and an unflagged LLM quote during the reviewer back and forth.
NeurIPS, the largest annual machine-learning research conference, is running a controlled experiment in letting reviewers use large language models for its 2026 cycle. A self-identified reviewer's account shows where the policy is meeting practice.
The experiment, detailed on the official NeurIPS page, is voluntary and opt-in. Authors choose at submission, reviewers opt in via recruitment, and each reviewer-paper assignment is randomized to one of three conditions: unassisted, open-ended LLM assistance, or structured LLM assistance. LLM providers use zero data retention, and the design has IRB review at multiple institutions.
The design builds on last year's Checklist Assistant experiment, with 234 author submissions and 539 pre-use survey responses. More than 70 percent of authors found the assistant useful; the top failure modes were inaccuracy and over-strictness, per the experiment paper.
A single reviewer's account describes a detailed handwritten-style review sitting next to two superficial ones, and a discussion-period moment in which a reviewer quoted LLM output without flagging LLM use.
Area chairs, blind to condition, will assess review quality after review. The question their data will not resolve on its own: whether reviewers should disclose any LLM use, including during discussion, or use the model to check against an author's rebuttal.