The unit of frontier intelligence is no longer how big a model is. It is how much of that model actually runs on any given question. That is a quiet rewrite of the entire AI cost stack, and it just priced itself into the market twice in one week.
Thinking Machines' Inkling-Small is a 276-billion-parameter mixture-of-experts model that spends compute on only 12 billion parameters per answer. The wire will frame this as a model release. The mechanism is bigger: training cost still tracks total parameters, but the price any buyer pays now appears to track active parameters — the two curves may have just split.
OpenAI's GPT-5.6 cuts the same week make the point in price. The Luna tier dropped 80 percent and Terra 20 percent, while the Sol tier, the largest, kept its throughput gains. Pricing appears to be following the same curve as the architecture — small active paths, cheaper per query, same frontier reach — though direct pricing-policy evidence from Thinking Machines was not sourced. The number that matters on Thinking Machines' release is the 23-to-1 ratio between total and active parameters, and the pattern underneath it points to a market in which active parameters set the price.
Anyone reading a model announcement from here on should ask one question. What runs per query, not what was trained. The frontier just became a per-call economy, and the model size on the press release is the wrong number to compare.