AI can finish a day's coding task but still can't run a coffee shop. The gate is the kind of work, not its length — the test for any AGI (human level artificial general intelligence) headline.
Rob Wiblin had filmed an episode arguing for long AGI timelines. By the time it shipped, the public mood had already moved past him. That is the framing of his 2026 retrospective episode, and the gap between his footage and the moment it landed is the year's clearest example of why the AI forecasting community kept changing its mind. Wiblin walks back through seven exhibits for and against acceleration, then pokes four honest holes in the acceleration case. The mechanism the recaps keep skipping is the kind of work, not the length of the task.
The acceleration case starts with METR's task-completion horizon: the human-expert task duration at which an agent hits 50% or 80% reliability. That horizon was doubling on a short cadence through 2024 and into 2025. Capabilities curves looked vertical, and a budget line that started at two-minute chores had reached multi-hour work in roughly two years. On the revenue side, Anthropic closed its Series H round at a $47B run-rate in late May 2026, with a confidential S-1 draft filing following days later; Simon Willison flagged the run-rate the same week. The 8,400% annualised growth Wiblin uses in Exhibit 1 is an extrapolation of those company-attributed numbers.
The task-completion horizon stopped doubling. New model releases added capability in places the benchmarks did not measure. The surprise Wiblin flags in Exhibit 3 is the plateau, not the jump. METR's methodology page is candid about which models get periodic re-evaluation and which get skipped, which is part of why a "doubling every X months" line travels further than the underlying data warrants. The accelerationists had been forecasting on the slope; the slope had changed.
The case for self-improvement is Exhibit 4. According to Anthropic's Institute essay, Claude now writes 80% of the company's code, and engineers on average ship roughly 8x as much code per quarter as they did from 2021 to 2025. The same essay is careful: recursive self-improvement is "not inevitable," but "could come sooner than most institutions are prepared for." It is a company-attributed case for compounding capability, made by the company that benefits if the case is true. Read it that way.
Exhibit 6 is the open math result OpenAI announced in 2026. The Wiblin transcript walks through the announcement but does not point to a specific paper or arXiv link, so the strongest version of the claim is "OpenAI made a breakthrough on a famous mathematics puzzle," not "OpenAI solved X." That distinction is the difference between reporting and a hype echo.
The counter-case is the clean-versus-messy gap. AI can now finish a day's software task, draft a multi-step report, or run a small process under a clear rubric. It still cannot run a coffee shop. Wiblin points to vending-bench, the AI Village, and Project Vend cafe as the live probes where AI managers fail at exactly the kind of work that mixes long-horizon planning, physical-world recovery from novel failures, and shifting customer demand. The AI cafe experiment is the load-bearing anecdote of the year: capability on a benchmark is not the same as capability in a building with a sink.
Exhibit 7 is the inference-scaling turn. Toby Ord argues that the shift from pre-training-compute scaling to inference-compute scaling reshapes AI governance: it lowers the strategic importance of open-weight models, changes frontier-lab business models, and complicates the training-compute-threshold governance lane that has dominated policy debate. The size of the effect is contested, but the direction is not.
The audit's residue is four unresolved forecasting questions, and the reader can use them. The first is how much of the recent revenue growth is durable demand versus one-time enterprise pull-through. The second is whether the plateau in task length is a measurement artefact or a real ceiling. The third is whether recursive self-improvement compounds or stalls. The fourth is whether long-horizon training closes the clean-versus-messy gap inside two years, ten, or never.
The simplest reader-side test: any future AGI headline that talks about task length without naming the kind of work being measured is using the wrong yardstick. The gate is the work, not the duration. Wiblin's audit hands the reader that test, and the next round of "AGI in 18 months" or "AGI in 50 years" headlines is where it gets used.