The cost of running AI is on the same curve as electricity and broadband. The real question is who captures the value when it does.
A Microsoft engineer reportedly got an email this summer telling him to ration his AI usage. The cost of running AI, the line item the industry calls inference, has become a procurement problem at the world's largest software buyer. John Furrier, the co-founder of trade publication SiliconANGLE, argues that's the wrong direction. In a piece published this week, he says inference should become as cheap and abundant as broadband, electricity, or cloud storage, and that the market will only fully open up when it does. The cost curve backs part of the case. The redistribution that follows is where the thesis gets contested.
Introl's December 2025 update on inference unit economics puts GPT-4-class performance at roughly $0.40 per million tokens today, down from about $20 in late 2022, a decline of close to 50x in three years, faster than PC compute fell in the 1990s or bandwidth fell during the dotcom buildout. Cloud rents for NVIDIA H100 chips, the previous-generation accelerator that still anchors most production inference fleets, stabilized at $2.85 to $3.50 per hour after dropping 64% to 75% from their 2023 and 2024 peaks. DeepSeek's published API pricing sits at the lower bound of that curve, and the company's roughly 90% list-price cut against Western incumbents is what pulled the rest of the market down with it.
Zylos's April 2026 research on AI agent compute markets calls the shift the "Inference Flip": cumulative global spending on running models overtook training spend early this year. Inference is now about 85% of enterprise AI budgets and roughly two-thirds of total AI compute spend. The discussion is no longer a forecast about whether inference will dominate the cost stack. It already does.
Cheaper inference doesn't democratize the value of AI. It migrates it. Zylos estimates that an agentic workload, a multi-step task that chains dozens or hundreds of model calls, costs $0.10 to $1.00 to complete, against about $0.001 for a single chatbot exchange. A 100x to 1,000x multiplier per task means the buyer is no longer purchasing tokens. The buyer is purchasing a workflow, and the bill is now bounded by how the workflow is shaped, not by the per-token rate. The same Introl data shows operational levers compounding the curve: quantization, a technique that runs models at lower numerical precision, cuts 60% to 70% off the cost of serving a given model, and speculative decoding, in which a small model drafts an answer and a larger one verifies it, can cut latency 2x to 3x. Those gains accrue to whoever controls the serving stack, not the customer.
The supply side is doing its part to push prices down. NVIDIA's blog on open-source inference on Blackwell, the company's current-generation accelerator, makes the vendor case that the lowest cost-per-token today is achieved with open-weight models on its newest silicon. That is also the configuration that compresses reseller margins fastest, since open weights remove the licensing moat that protected premium API providers through 2024 and 2025.
This is where Furrier's analogy earns its keep, used carefully. A $27 truffle is a luxury item with a small market. A 99-cent chocolate bar is a mass-market product with a much larger one. The analogy is that inference should follow the chocolate path: cheap, abundant, everywhere. The market for the underlying AI does expand on that path. The market for selling inference at a markup does not. Per-token resellers, premium-priced API providers, and the inference layer of any cloud whose margin depended on a scarcity premium are the first casualties. The model-plus-data owners, hyperscalers with a power moat, and the platforms that own agentic workflow shapes are the beneficiaries.
The falsifier is physical. The Inference Flip runs on real electrons, real advanced packaging, and real high-bandwidth memory, the stacked DRAM that feeds modern accelerators. Each of those supply chains is concentrated, and each is now a binding constraint on how fast the chocolate bar can actually get cheaper. If power, advanced packaging, or HBM becomes the bottleneck before the cost curve flattens, the commoditization track stalls and the market stays shaped like the truffle side of Furrier's analogy for longer than the curve implies.
The Microsoft rationing email, if it landed the way Furrier describes, is the visible symptom. The argument isn't whether inference gets cheaper. It is. The argument is who is positioned when it does.