Raffi Krikorian's switch from Anthropic's Claude to Moonshot's Kimi K3 sits at the front of a quiet wave of U.S. developers and companies moving to cheaper Chinese AI releases.
Mozilla's chief technology officer Raffi Krikorian used to run most of his team's routine work on Anthropic's Claude, the U.S. AI lab's flagship assistant. This month he moved them to Kimi K3, a 2.8-trillion-parameter open-weight model from Beijing-based Moonshot AI. He told the GAS Indian Times it felt snappier. The bill was smaller. That is the door through which the rest of the story walks.
K3 is the latest in a half-year run of Chinese frontier releases that have closed the price gap with U.S. labs. Z.ai shipped GLM-5.2 in June; Alibaba put up a Qwen3.8 Max preview in July; DeepSeek's V4 previews have been circulating since April. K3 stands out because it is built around MoE sparsity, a structure with 896 expert sub-networks inside one model and only 16 firing on any given prompt, roughly 1.8% of the total. The same trick that lets Chinese teams train on limited compute is what makes inference cheap for the buyer: most of the model is dormant most of the time, so the per-task bill drops.
Tom's Hardware reports the full open-weight release landed on July 27. K3 carries a one-million-token context window and native vision input. Bleap's review puts K3 at 57 on the Artificial Analysis Intelligence Index, within three points of Claude Fable 5's 60. K3 also posts a +732 Elo jump over its predecessor K2.6 on long-horizon knowledge work, Simon Willison notes.
The pricing is the part that matters for procurement. Per million tokens, K3 lists at about $3 input and $15 output, Sonnet-tier rather than Opus. Bleap's review puts cost-per-task at $0.94 against Claude Opus 4.8's $1.80, nearly half. For a CTO running an internal coding assistant or a search pipeline at scale, the math is the math.
Krikorian is not alone. Coinbase has publicly swapped to Chinese AI models as a cost move, and a wave of independent developers and small U.S. teams has done the same, trading a small benchmark gap for a large bill cut. The pattern crosses industries because the variable that flipped, per-task cost at near-frontier capability, is the same variable procurement teams already measure.
That is also why the U.S. government is no longer treating this as a product story. The Trump administration has formally accused Moonshot of building K3 through "covert" distillation, a technique in which one model is trained to imitate another by generating cheap synthetic data from the target, effectively copying its behavior without access to its weights. Anthropic has made similar claims. Beijing has called the allegations "groundless." None of the accusations are proven. They sit in the same news cycle as Treasury Secretary Scott Bessent's warning of further sanctions on Chinese AI, and Moonshot's own disclosure that its training run used export-grade Nvidia silicon plus an unnamed alternative GPU vendor.
The chip piece is the real policy lever. China's frontier labs are working around U.S. compute limits through MoE sparsity, the same architecture choice that makes K3 cheap to run is also what makes it trainable on a constrained supply of high-end GPUs. If Moonshot can credibly say it built a 2.8-trillion-parameter model on a hardware pool that Washington thought it had throttled, the policy assumption that compute caps delay parity collapses.
The engineering decision and the policy response have collided in one week. Krikorian made his choice at the per-task dollar level. Coinbase did the same. The administration's accusation tried to push that choice into the national-security lane. The procurement math is already running in the opposite direction.
The next falsifier is concrete: if the dollar-per-task advantage does not generalize beyond K3, the pattern is anecdote. If a second U.S. financial, healthcare, or infrastructure company names a Chinese model as its default by the end of the third quarter, the pattern is procurement, and the policy fight catches up to it.