Google's engineering post claims a 1.69× speedup on its TPU v6e AI chip at 1440p. The underlying idea is older, and the speedup is smaller than the academic paper behind it.
A few seconds of AI-generated video at HD or 2K resolution still costs more than most independent teams can afford to run. The reason is not that the math is hard. Most of the math the model runs is wasted.
Inside a video model, a layer called self-attention compares every token against every other token. Comparisons grow quadratically, and a 1440p video carries roughly four times as many tokens as a 720p video. As resolution climbs, that layer's share of per-layer latency rises from 55.5% at 720p to 88.2% at 1440p, per a Google engineering post published Sept. 30.
The fix is structural. Spatial heads decide what is in this frame and only need to look locally. Temporal heads decide how this frame relates to the next and only need to look across a small spatial region. A sparse attention schedule skips the rest, running on a kernel called Splash Attention on TPU v6e.
The result on 8× TPU v6e: 1440p end-to-end denoising drops from 2,471 seconds to 1,461, a 1.69× speedup that saves more than 16 minutes per video at 24.05 dB PSNR. The idea is not new. The Sparse VideoGen paper reported up to 2.28× on CogVideoX-v1.5 and 2.33× on HunyuanVideo on different hardware. Google's number is smaller, and it is the company's own measurement on its own silicon, with no independent benchmark yet.
If the optimization holds outside Google's stack, the stakes are concrete: cheaper inference means smaller product teams and open-source projects can ship video generation without hyperscaler budgets. That is the real test, and it has not been run in public.