Software benchmarks have been measuring the wrong race. Most coding tests cap spending at a few dollars and grade a fix to an existing file, then publish the result as if they had measured a marathon. Epoch AI and METR's MirrorCode is the first instrument built for the longer distance: it lets the model run for days, spend what a real project costs, and rebuild whole programs from scratch in a sandbox.
The first readings look strange on purpose. A single run on one of the largest tasks cost $2,600 and went nineteen days without a human touching it. Claude Opus 4.7 reimplemented gotree, a roughly 16,000-line Go bioinformatics toolkit, in about fourteen hours for $251, against Epoch's estimate of two to seventeen weeks for an unassisted engineer. The numbers are not just big. They are the first honest read of how far a model can carry a build before the test itself becomes the bottleneck.
The contamination caveat is not a footnote. The 25 target programs are open source, which means pretraining likely saw them, and Epoch says so. Treat that as a calibration error on the instrument, not a flaw in the model. The field's next move is not to crown a leader but to decide whether contamination is a patchable bug or a structural limit on every benchmark built from public code.
That choice will shape the next decade of capability claims, and the field has not made it yet.