The joint UK US pre release test found Kimi K3 reached step 17 of 32 on a simulated attack path, vs. 28.5 for U.S. leaders. The Beijing based Moonshot plans a public download by July 27.
The UK AI Security Institute and the U.S. Center for AI Standards and Innovation jointly tested Moonshot AI's new model Kimi K3 for cyber capability before its planned open-weight release. They found Kimi K3 trails the most recent frontier U.S. cyber-capable models on the published benchmarks, and that its safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during the evaluation.
On "The Last Ones," a 32-step simulated corporate-network attack scenario, Kimi K3 reached on average step 17, while the most cyber-capable U.S. models averaged 28.5 steps, according to the joint assessment. Kimi K3 also performed above GLM-5.2 on the same evaluations.
Its overall cyber capability has a wider confidence interval because Kimi K3 was estimated from a single benchmark (ExploitBench, 41 tasks). Moonshot's hosting setup constrained UK AISI and CAISI to a selective set of cyber evaluations. The U.S. closed-weight models were tested with system-level safeguards disabled to reduce refusals and enable measurement of maximal capability; publicly available versions of those models retain those safeguards. That asymmetry means the capability gap narrows materially in deployed comparisons.
Kimi K3 was released July 16, 2026, with the open-weight release slated by July 27, 2026, per Moonshot AI's blog.