The UK AI Security Institute found downloadable models match closed frontier systems on narrow cyber tasks within 4 7 months, shorter than most enterprise planning cycles.
The UK AI Security Institute has measured the gap between the best commercial AI systems and the best publicly downloadable ones on cybersecurity tasks: four to seven months on narrow capability tests, and roughly seven months on chained, multi-step hacking operations. The AISI blog post published this month is the first regulator-side measurement of where that lag stands in 2026, and the agency frames it bluntly: defenders have a "short window" before today's frontier cyber capabilities become accessible without the same safeguards that proprietary labs apply.
For a CISO or a procurement officer, the operational meaning is sharper than any benchmark name. Tasks that today require a paid top-tier commercial AI model could be done with a free downloadable model inside a year. The chained, multi-step operations that today's red teams still pay to run on closed systems are only a few months further behind. Enterprise patch and procurement cycles run longer than that, which is the actual story the measurement is telling.
AISI ran 70 narrow cyber capability evaluations and one long-horizon cyber range it calls "The Last Ones." On the narrow evals, GLM-5.2 performed most like Claude Opus 4.6, which Anthropic released about 4.3 months earlier. DeepSeek V4-Pro sat between Claude Opus 4.5 and GPT-5, models released roughly six and ten months before. AISI's note earlier in 2025 had the gap closer to six to ten months across most of those tests, so the agency is documenting a real compression in 2026, not a single-quarter fluctuation.
The long-horizon result is the one that should keep security leaders from declaring parity. On "The Last Ones," a chained cyberrange that mimics realistic intrusion paths rather than isolated exploits, GLM-5.2 reached Claude Opus 4.5 (released less than seven months earlier), and DeepSeek V4-Pro fell below Sonnet 4.5 (released seven months earlier). AISI describes this as evidence of "big model smell," its phrase for open models that match on narrow benchmarks while still lagging on generalization. The narrower the test, the smaller the lag. The more the test looks like an actual intrusion, the larger the lag.
AISI also flagged Kimi K3, the model Moonshot AI has been positioning as "Open Frontier Intelligence" on its own tech blog, and said it will run the same cyber evaluations once Moonshot releases the weights. Until those weights land, any specific capability claim about K3 is speculative. The K3 blog is positioning, not a measured result. The watch item is the day AISI posts a Kimi K3 cyber eval, because that post will reset the lag number again.
Google DeepMind CEO Demis Hassabis published a framework piece this month arguing for a coordinated international response to what he frames as the dawning of a new age of AI capability. A separate arXiv preprint on distributed attacks against persistent-state AI control adds a research-side control-and-adversarial anchor: as open models approach frontier capability on the same tasks, the surface for distributed, coordinated attacks on the systems that manage them grows in proportion. Hassabis's piece is CEO opinion, not peer-reviewed, and the AISI post is a government benchmark, not a policy directive. Read together, the two are converging on the same operational point: the diffusion clock is shorter than the policy clock.
What changes for a defender in the next six to twelve months? Three things, all anchored to the measured lag. Treat any assumption that frontier cyber tooling is a paid-only capability as a planning horizon under seven months, not a year. Expect red-team assumptions calibrated to closed systems to be obsolete at the next public model release, and build a refresh cadence around that. Make open-weight evaluations part of vendor due diligence, not a separate research lane: if a public model is within months of your paid vendor on the relevant eval, the procurement question is no longer about raw capability.
The next data point on this clock is the Kimi K3 weights. Moonshot has not released them as of the issue date. The day they drop, AISI's evaluation queue is the one to watch.