A flagship natural language and AI research conference is testing what happens when AI systems participate in reviewing submitted papers.
EMNLP 2026, a flagship venue for natural-language and AI research, is wiring AI reviewers into a slice of its own peer-review pipeline, on an opt-in basis, and treating the process itself as the experiment.
Authors submitting through ARR can choose whether their paper enters the AI-reviewing track and can change that choice later. AI reviews are not the verdict: a human reviewer remains the decision-maker, and authors get a feedback channel to flag problems with the AI reviews they receive.
Only one of the two AI systems is fully identified: ReviewerToo, a multi-stage agent built by researchers at Mila, ServiceNow Research, HEC Montréal, Polytechnique Montréal, Université de Montréal, and Unified Sciences, running GPT-OSS-120B in inference mode on a dedicated server. On 1,963 held-out ICLR 2025 submissions, the system reached 81.8% accept/reject agreement versus 83.9% for the average human reviewer, and an LLM judge rated its reviews higher quality than the human average. The study runs under UT Austin IRB protocol STUDY00007931.
The headline number is not an EMNLP result; it is ICLR-2025 validation. The conference is measuring whether authors trust the process enough to opt in, what they flag, and what guardrails the field normalizes before AI reviews are read as authoritative.