Per token cost is falling, but token consumption is forecast to grow 24x by 2030, and the gap between the two curves is why AI bills keep surprising buyers.
A finance lead at a mid-size software company rolls out an internal AI agent project. The first invoice arrives, and nobody, not the vendor and not the buyer, can reliably forecast next month's spend. The reason is a math problem, not a sales one.
The unit being billed is the token, the chunk of text a large language model reads and writes. The same prompt can produce different responses, different models produce different responses, and agentic systems stack multiple agents, multiplying the variance per task. Token consumption is not a meter reading. It behaves closer to a probability distribution.
That would be a tolerable quirk if usage were flat. It is not. Goldman Sachs forecasts token consumption will grow roughly 24x between 2026 and 2030, to about 120 quadrillion tokens a month, as companies shift from chat-style assistants to AI agents that do work in the background. The forecast is a bank projection rather than measured usage, and Goldman ties the gain to industry cash flow rather than to any single customer's bill. It is still the cleanest available anchor for what is about to hit enterprise buyers.
Per-token cost has plummeted in recent years, the product of inference efficiency gains, model distillation, and a competitive race among OpenAI, Google, Anthropic, and a long list of open-weights challengers. The total bill has not fallen with it. The two curves have decoupled: the unit gets cheaper while consumption scales faster than the price drop can offset. Microsoft, Google, and Anthropic have invested hundreds of billions of dollars in developing the LLMs behind ChatGPT, Claude, and Gemini, and the vendors have every incentive to keep that engine running at full throttle. Falling unit prices, in this market, are an adoption subsidy, not a margin cut.
Agentic systems make the variance worse, not better. A single chat query bills one model for one turn. An agent doing a task can call a planner model, a coder model, and a critic model in sequence, and re-loop when the critic rejects the output. Each loop burns tokens on work the user never sees. Multiply that across a workforce of agents and the bill stops tracking prompts and starts tracking the model's own self-correction rate.
That is the gap buyers are starting to push back on. Simon Gooch of identity management vendor Saviynt argues long-term fixed-price AI commitments do not make sense because the underlying economics are moving too fast to model. The same frustration surfaces inside the vendors' largest customers: Uber's chief operating officer has publicly flagged runaway token spending on Anthropic's Claude Code, and Microsoft's AI head has said Anthropic's services are too expensive to use at scale. The pushback is not coming from the budget hawks in the corner office. It is coming from the engineers and operators whose product roadmap runs through the models.
The portable mental model is this: when the unit of cost behaves like a probability distribution rather than a price tag, no sales motion can rescue the bill. Goldman Sachs has put a date on the scale-up, 2030 at around 120 quadrillion tokens a month. The industry's ability to price its own product is the part still missing.