A Princeton led blind test—sending AI agents' papers back to the original human authors for review—gave today's top AI research agents six days, $3,000 in AI service credits, and a GPU budget.
A new paper from a Princeton-led team separates one of AI's load-bearing promises into two parts. The result, posted as arXiv 2607.27191 by Peter Kirgis, Sayash Kapoor, and collaborators, is a specific negative on the second half.
The team ran what they call a "shadow evaluation." They pulled real research questions from unpublished NeurIPS 2026 papers, where an agent could not have memorized the answers, and asked current frontier agents to write full papers in response. Each agent was given six days, $3,000 in Anthropic API credits, its own virtual computer, a GPU budget, and open web access. Two of the resulting papers were then sent back to the original human authors for blind review.
The result, as reported by MIT Technology Review: both were rejected. The authors judged the agents capable on engineering, including literature review, running experiments, and compiling results, and "unambiguously bad" on the research itself. Specifically, the agents explored poorly, committed to unpromising approaches, could not backtrack when those approaches failed, did not use feedback to revise their direction, and misused the compute, time, and token budgets they were given.
The finding is a snapshot of today's agents on today's questions. The falsifier is straightforward: a future agent that passes author review in the same protocol.
Why this lands now. The recursive self-improvement loop, in which an AI system helps train or design a better AI system, is the load-bearing assumption under several recent high-profile claims. Anthropic's June essay "When AI Builds Itself" argued the loop is closer than skeptics think. Jack Clark's Import AI newsletter has tracked the same direction. OpenAI said in July that GPT-5.6 Sol helped post-train a smaller model. And the widely circulated AI-2027 forecast scenario assumes the loop tightens on a near-term timeline.
The Princeton result does not contradict any of those claims directly. It does name a specific failure mode the loop will have to clear: not raw capability, but the exploration-and-backtracking behavior that defines open-ended research. A useful reader lens: when a vendor says its model is helping build the next model, ask whether the work is in the engineering loop (likely improving) or the research loop (per this snapshot, not yet).
The Download, in brief: