Dwarkesh Patel argues that AI compute (the chips and time spent running every model) could rise 10 15x in cost as smarter systems shift work from training (building the model) to inference (running it for users), breaking the 'AI gets cheaper every
The mechanism turns on where compute is actually consumed. Epoch AI's frontier-labs analysis shows that the labs training the largest models don't consume the majority of AI compute. The bulk of demand sits elsewhere: large-scale inference workloads, reinforcement-learning post-training, and agentic tasks where a model reasons, calls a tool, reads the result, and reasons again. Each step adds tokens to the bill. Each smarter model adds more steps, not fewer, because the work a user can offload to the model grows with capability.
The supply side matches. SemiAnalysis frames the value-capture shift toward model labs: as the model becomes the bottleneck product, the model captures the margin, and the chip vendor doesn't. Compute providers no longer compete the price of inference down, because the customer paying the bill is the lab selling the finished product. The price of a token rises toward what the application can bear, not toward what the silicon costs.
That breaks the curve most readers have internalized. The "AI gets cheaper every year" intuition was built on training-cost declines: bigger chips, more efficient architectures, falling FLOPs per parameter. Training is a one-shot bill. Inference is a recurring meter, and a recurring meter doesn't respond to hardware efficiency the way a one-shot training run does. A 40% gain in chip efficiency, the kind hardware roadmaps still promise, doesn't offset a 5x rise in tokens per task, and tokens per task is what has been climbing.
The bill lands unevenly. Frontier labs can absorb it because they sell finished products rather than raw compute, and they pass it through as subscription and API prices. Independent builders, mid-market SaaS companies, and any team that buys model access by the token feel the meter running faster than their runway. AIWeekly's alert puts the 10-15x range on the table as the magnitude Dwarkesh is modeling. The exact phrasing in the canonical written post is the version to quote once the full text is in hand.
The constructive read is that scaling-only is not a strategy. The cost curve for raw compute has bent. The cost curve for the next unit of model capability may be steeper than the prior decade implied, and the builders who plan around it (caching, routing cheap models to cheap tasks, treating inference as a budgeted line item rather than a free good) will outlast the ones who don't.
The next empirical test is the next frontier release. Whether its inference cost per useful task rises, falls, or holds, and whether the chips built for it arrive on the schedule their vendors promised, will set the read on whether Dwarkesh's number is a forecast, a floor, or an overstatement. Both data points land before the end of 2026.