Geoffrey Irving says the leading labs' three pillar alignment plan might survive the jump to AI that outthinks us, but nobody has a strong argument that it will.
Between now and the day AI passes the human-level threshold, the training signal for each new model will stop coming from people and start coming from the previous generation of models. Geoffrey Irving calls that transition a "phase shift," and it is the moment, in his telling, when the leading labs' alignment plan either holds or breaks, with no proof yet that it holds.
Irving laid out the argument in episode #251 of the 80,000 Hours Podcast, recorded June 29, 2026. The episode makes a specific, time-bounded claim: full-blown superintelligence (AI that can outperform humans on most cognitive tasks) could arrive within two to three years, putting the verification gap on a concrete clock rather than a vague future.
The playbook Irving is interrogating has three pillars. Train future models to have good character. Use increasingly capable AIs to supervise other AIs. Watch closely for signs of deception or scheming, the technical term for models that hide goal-seeking behavior from their operators. Irving thinks the combination might work. He also thinks nobody has a strong argument that it will.
The reason is the threshold. Below human-level intelligence, humans can usually tell whether a model's work is good and correct its mistakes; the training signal that produces the next generation is grounded in human judgment. Above the threshold, models themselves increasingly determine the feedback used to train their successors. The human grader is, by definition, no longer the most capable system in the loop.
That is the "phase shift" the rest of the episode is built around, distinct from a more familiar AI-doom register. Irving treats the leading labs as serious actors working on a real problem, and points instead at a single mechanism: a training pipeline that depends on the next model being smart enough to grade the work of a model smarter than its graders.
Resolution, the new research organisation Irving co-founded, is being built around that mechanism. The nonprofit is described in the episode and on the 80,000 Hours job board listing as pursuing a portfolio of neglected alignment research bets, and it is actively hiring. The full episode transcript lists the workstreams, which the show presents as the kinds of questions that would let the field actually prove, rather than assume, that the three-pillar plan survives the threshold.
Irving's policy prescription is the slower part of the argument. He wants governments to slow AI development now, while research on which methods can actually be trusted catches up. The case for slowdown isn't a substitute for a working playbook. It buys time for the gap between "the playbook might work" and "we have a strong argument it will" to close, which is the part that is hardest to close quickly, with the cost of being wrong rising sharply at the threshold.
What slowdown would buy, by his account, is a window for the missing verification research to mature before the human grader leaves the loop. What it cannot buy is certainty; even with a slower timeline, the threshold problem remains. And it assumes the major governments and labs are willing to coordinate, which the past two years of AI policy have not exactly suggested.
The two-to-three-year horizon is what gives the argument its weight. If Irving's timing is roughly right, the window for the verification research to mature is the same window the leading labs expect human-level systems, and the clock starts now rather than in some unspecified future. If the horizon slips, the argument still applies; it just buys more time than Irving is crediting.
The question worth carrying out of the episode is sharper than "are the labs being careful." It is the one Irving's research bets are built to answer: how will anyone grade the work once they cannot grade it themselves. The next two years of lab releases, regulator questions, and Resolution's hiring list are the readable signals for whether anyone is working on that question, or only on getting to the threshold first.