Chinese lab Moonshot AI's unreleased Kimi K3 may have appeared on a public leaderboard as 'Kivine', where one response took 35 minutes. The slow, capable result is a new shape for the frontier model race.
A Minecraft-like game, rendered live on a public test arena, took roughly 35 minutes for one of the year's most capable AI models to generate. The frontier model race is now being decided on how long a user will wait, and what they get in return.
On July 15, 2026, an anonymous model called Kivine appeared on LMArena, the crowdsourced leaderboard where users pit models against each other head-to-head. Within hours, X user @Lentils80 linked it to Kimi K3, a flagship that Chinese lab Moonshot AI had not yet announced. The post triggered a community-wide attribution chain that ran through @testingcatalog and other accounts. Twenty-four hours later, Moonshot officially launched Kimi K3. The same-day timing and the model's behavior left the community confident the anonymous checkpoint was a pre-release test of K3, though Moonshot has not publicly confirmed the link.
K3's architecture is a Mixture-of-Experts design, meaning only a slice of the model activates on any given token, with 2.8 trillion total parameters, 104 billion active per token, 896 routed experts, and 16 of those experts firing per token alongside 2 shared ones, according to launch coverage. Its context window is exactly 1,048,576 tokens, a million tokens of working memory that fits roughly a long codebase, a research report, or a stack of PDFs in one prompt. On independent leaderboards K3 ranked second on the Vals AI index, third on Artificial Analysis's Intelligence Index (behind Claude Fable and GPT-5.6 Sol Max while being cheaper to run), and first on Frontend Code Arena, as catalogued by Interconnects. That posture is the highest an open-weight model has held on these boards since DeepSeek R1.
That is the same arc on which the 35-minute figure sits. To get Fable-tier frontend quality, the model has to think longer per token. The trade is no longer hidden in benchmark bars. It is the headline.
The Minecraft figure is a tester-reported number, not a formal measurement, and the underlying task varies. What it does show is the consumer-product cost of the quality gain. A 35-minute response is unusable for chat, marginal for code review, and tolerable only for offline creative work where the user can walk away. That third category is what Moonshot is targeting.
The economics are the part the wire coverage tends to skip. On Artificial Analysis's Intelligence Index, K3 lands third while running cheaper per token than either Fable or GPT-5.6 Sol Max, per Interconnects. For a buyer, that combination of open weights and lower per-token cost at near-frontier quality is the actual disruption. The 35-minute latency is the consumer cost. The dollar cost per token is the procurement story, and procurement is where open-weight pressure on closed labs usually starts.
Open weights followed on July 27, 2026, when 96 Safetensors shards landed on Hugging Face under a custom Kimi K3 License. The release makes K3 the first open-source model in the 3-trillion-parameter class and the first time a Chinese frontier lab has shipped full weights for a flagship at this scale. Western wire coverage in the same window, including the BBC, Bloomberg, and the New York Times, treated the launch as a geopolitical milestone, with the usual backdrop of US export controls and Alibaba and Tencent's Moonshot backing.
Nathan Lambert of Interconnects calls K3 the closest an open-weight model has been to the frontier since DeepSeek R1, narrowing the US-to-China frontier gap from a debated six to nine months down to an estimated three to five. That is a claim, not a measurement, and it depends on which leaderboard you trust. The benchmark community is still working out how to score models that take half an hour to answer.
If a flagship can be stealth-tested on a public leaderboard and reach credible parity with closed frontier models within a day, the launch event itself is becoming a footnote. The signal worth tracking is what a model does on the arena before anyone has heard of it. The harder question is what kind of product a 35-minute flagship is actually built for, and who is patient enough to use it.