NeurIPS, the major machine learning research conference, hides instructions in 2026 submissions so reviewers who upload manuscripts to AI chatbots leave a signature.
A reviewer opening a NeurIPS 2026 submission this month encounters a paper that reads like any other: equations, figures, a methods section. Hidden inside the manuscript sits an instruction aimed not at the reviewer but at any large language model asked to summarize the file. If the model follows it, the review it produces sprouts telltale phrases such as "This work addresses the central challenge" and "The claims of the paper," which then read like a watermark when a human ethics reviewer sees the report (The Scientist).
That tripwire is now the centre of a quiet fight inside the machine-learning community over how to police the conference's ban on reviewers uploading manuscripts to AI chatbots. NeurIPS 2026, the 40th annual conference on neural information processing systems, is slated for Sydney in December and is running the watermarking scheme in active review. The practice has been publicly criticised by researchers who say it presumes bad faith and could itself be weaponised.
The first hard numbers on the technique come from a neighbour. ICML 2026, the International Conference on Machine Learning, reported in March that it used the same approach, hidden LLM instructions watermarking submission PDFs, to flag 795 reviews, about 1% of all reviews submitted, by 506 unique reviewers. Of those, 398 were reciprocal reviewers whose behaviour led to 497 papers being desk-rejected, roughly 2% of all submissions. The conference emphasised that no generic AI-text detector was used; every flag was manually confirmed by a human before any paper was rejected.
That human-in-the-loop step is the central defence against the obvious objection. The hidden prompt does not, on its own, sink a paper. It produces a review that a human can recognise as model-generated, and a human then has to confirm the violation against the conference's reviewer code of conduct, the NeurIPS reviewer guidelines and the equivalent rules at ICML. Supporters argue that the watermarking turns an invisible offence into a verifiable one without requiring surveillance of every keystroke a reviewer types.
Critics, including researchers who have reviewed for NeurIPS and posted objections on LinkedIn and Reddit, call the practice adversarial. Their core complaint is that hiding instructions in a submitted paper assumes the reviewer is cheating and turns the manuscript itself into a tool against the person reading it. As one critic quoted by The Transmitter put it, the same technique that catches an LLM-assisted review can be aimed at any reviewer who later paraphrases or quotes the paper, blurring the line between detection and entrapment.
The criticism is not abstract. A NeurIPS 2025 virtual session on prompt-injection attacks and defences targeting AI reviewers already catalogues the offensive possibilities, and any technique that embeds invisible instructions in a manuscript can, in principle, be turned against the reviewer pool the conference is trying to protect. Conferences that adopt the tripwire therefore adopt, by construction, an adversarial model of their own reviewers.
What the two big 2026 conferences have bought with that model is an audit trail. ICML's 497 desk-rejects are the first concrete demonstration that hidden watermarking, paired with human verification, produces a number large enough to deter and small enough to defend. Whether the practice spreads, and whether the same mechanism stays a defensive tool rather than a weapon, is the question the community is now arguing out in public, one hidden phrase at a time.