Frontier AI is being built by a closed loop, and the human has moved up a level. Capability is no longer produced by researchers writing better code; it is produced by a propose-score-retrain cycle that compounds on its own signal, and the engineering trick that made it work was turning the next training task itself into an output the model can generate.
A concrete pass: the model studies its own prior research, names a task it currently fails, the system spins up many attempts in parallel, an objective rubric scores each one, and the best output becomes the next training signal. The next iteration starts measurably better. Where outcomes are objectively scorable (code, math, protein folding) the loop closes fast and the human contributes at four points that still cannot be skipped: the rubric, the next task choice, the curated fine-tuning data, and the reviewer backstopping the fuzzy-judge cases.
The CASP/Cambridge working paper that brought Hinton, Bengio, Pachocki, Horvitz, and Clark to one document describes the same loop from the policy angle, and several of those co-authors are employed by the very labs whose trajectory it assesses, so the projection is held at the level the source supports. The asymmetry is the honest edge: progress compounds where ground truth is cheap and stalls where a fuzzy judge and human reviewers have to stand in for it, exactly where the loop is most prone to quietly endorsing its own taste. The mechanism is reusable across labs: name the next hard task, score attempts, fold the best back in, repeat. The build loop is closed; the human still owns the rubric.
Reported by Sky for Type0, from What if automating AI R&D triggers an intelligence explosion?. Read the original: arxiv.org