Kimi K3, Moonshot's new AI model, is pitched for long running coding and AI agents that use tools. Parameters are a rough measure of a model's internal wiring, and the capability claims are still vendor asserted until independent benchmarks land.
A Chinese AI lab says it is preparing to release the first openly downloadable model at roughly three trillion parameters, with weights dropping on HuggingFace on July 27. Moonshot AI's Kimi-K3 release page calls the model the "next generation of open frontier models" and the "world's first open 3T-class model."
Moonshot pitches the model for long-running coding and AI agents that click through tools, with an extended context window aimed at "repository-scale code understanding." "Open" means downloadable weights anyone can run or audit, and "parameters" measure how much a model can hold in its wiring, not how well it reasons. The vendor names a custom architecture, Kimi Delta Attention with "Attention Residuals," plus native tool calling, browsing, and multi-step planning.
The capability claims are not yet measured. A practitioner thread on Hacker News puts the model at "mxfp4 native" with roughly 1.5 TB of VRAM to host, sitting at the practical limit of 8x B200 GPUs and realistically needing 16x B200s for full context. The same thread cites AISI cybersec benchmarks that rank K3 above GLM-5.2 but well behind state-of-the-art closed models.
Four gates turn Moonshot's pitch into measured facts: an independent coding benchmark such as HumanEval or SWE-bench, a real agentic tool-use evaluation rather than a self-reported demo, a technical paper or model card documenting Kimi Delta Attention beyond a marketing line, and license terms that confirm the weights are truly downloadable.
Kimi-K3 is a descendant of an active coding-focused line; K2.5, K2.6, and the K2.7-Code specialization already sit on HuggingFace. The July 27 drop tests whether the open-weights frontier can credibly scale past two trillion parameters.