AI is starting to help build the next, more capable model. A new research agenda is building the public machinery to measure, slow, and verify that race from outside the labs.
Anthropic disclosed two numbers in a recent Wired report. One of its models now performs 26 percent of its own AI research, up from zero at the start of 2026. Anthropic also says it spent 6 percent of its compute budget this year on figuring out how to make its AI safer. Both numbers are self-reported. No independent auditor has checked them. There is no public dataset that says whether the underlying capability is moving at the pace Anthropic claims.
The capability being measured is recursive self-improvement, the process in which AI systems help design and train the next, more capable AI. It is the capability that leading labs now describe as defining the frontier, the most capable and largest-scale models in the field. The measurement of recursive self-improvement sits almost entirely inside the labs themselves.
A new research agenda, Pacing the Frontier and led by University of Toronto AI researcher Raymond Douglas, treats that situation as a research problem. "We need to start treating this as a research problem," Douglas told Wired. "We don't really understand what our options even are or what they will do." The agenda does not endorse a pause. It does not endorse the industry counter that any slowdown is impossible. Its premise is that slowing frontier AI is an engineering problem with engineering solutions, and that the solutions need to be designed and tested in the open.
Measurement is one part of the agenda. Vals AI has launched a Recursive Self-Improvement Index that aims to put a public number on the same capability Anthropic just disclosed internally. Hardware-level interventions are another. A September 2025 arXiv preprint sketches embedded off-switches for AI compute, hardware-level interrupts that can throttle a model in operation. A RAND working paper maps the same governance idea onto the US export-control system, specifically the 3A090 and 4A090 classification numbers that govern how many high-end GPUs a given buyer can receive and where those chips physically end up. The mechanism is the same in each case: a frontier model cannot be governed from the outside unless the outside can see and touch the hardware it runs on.
Public-interest institutions are a third. The UK's AI Security Institute is the closest existing body with a mandate to evaluate frontier models before deployment. Its existence is the proof of concept for outside oversight. Other governments are funding similar work. None of it is at the scale of the labs whose models it would evaluate.
The argument for external oversight rests on the labs' self-interest, not in spite of it. Anthropic's 6 percent figure is exactly the kind of number that an outside evaluator would stress-test. Right now no one is doing the stress-testing.
The watch item is which gets solved first: the measurement, the hardware-level interrupt, or the public-interest institution. Each is solvable in principle. None is solved today.