The 85% cache share — re reads of text the model has already processed in the same session to keep its working memory coherent — points to high bandwidth memory, not GPUs, as the next build constraint.
OpenRouter, one of the largest AI model-routing platforms, reported this month that the customers consuming inference compute at the highest rate are no longer humans. They are other AI agents. In a 7-day average for the week of August 10, 2026, agent-driven traffic hit roughly 7.3 trillion tokens while human-driven traffic stayed near 1.4 trillion tokens, a 5x ratio that did not happen by accident. The company's head of insights, Peter Walker, first surfaced the inversion in February 2026, when agent volume first crossed over.
OpenRouter is a gateway that sits in front of dozens of large language models, so when a developer routes a request through it, OpenRouter can see whether the caller is a person typing in a chat box, an autonomous agent in a tool-using loop, or some mix of both. To make that distinction, OpenRouter applies a 7-signal weighted composite (tool call rate, turn count, gap timing) and assigns each API key a label. The "agent" bucket is the one that just blew up.
Since the February crossover, agent token volume on OpenRouter has grown roughly 14x while human volume grew about 2.8x, per Walker's underlying data. A fourth bucket, "mixed" traffic that is partly human and partly agent, grew 4.7x over the same window. Daniel Newman, CEO of the Futurum Group, amplified the chart on X and forecast the ratio will hit 10x "and then higher and higher." Newman's projection is a CEO call, not a measurement, and the "10x" should be read as commentary, not data.
The 5x figure is also platform-specific. OpenRouter is one gateway, and the trajectory is not monotonic, with dips in April and July. But the shape is being corroborated elsewhere. A call-center consultancy testing DeepSeek on rented Nvidia H200 GPUs told Tom's Hardware it saw a similar cache-dominant pattern in its own agent logs. McKinsey's 2026 State of AI survey found 40% of large organizations now report scaling AI agents, up from 27% a year earlier. Agentic traffic is climbing on more than one front.
The 5x headline, though, is the smaller story. The bigger one is what those 5x tokens actually are. According to an Andreessen Horowitz chart built on OpenRouter's data, more than 85% of agent tokens come from cached prompts, text the model has already processed in the same session and is rereading to keep its working memory coherent. Agents do not start cold. They open a long task, call a tool, wait for a result, and reread the prior context to decide what to do next. Each reread counts as a token. Multiply that by a million simultaneous agents and the cached-prompt share explodes.
A token served from cache is much cheaper than a fresh one, since OpenRouter and similar gateways price cached tokens lower, so the revenue picture does not scale 1:1 with the volume picture. What does scale is the demand for high-bandwidth memory. Every cached prompt has to live somewhere the GPU can reach in a few hundred nanoseconds, which means it sits in HBM stacks on the same package as the accelerator. As agent sessions stretch longer and the working context per agent grows, the KV cache, the index the model uses to look up prior tokens, is outgrowing GPU HBM capacity. a16z is a reported OpenRouter investor, so the chart is interested, not independent; the underlying trend, though, is what the chip physics imply.
This is the next constraint to watch. The first wave of inference build-out was about how many GPUs to order. The next one is about how much HBM to bind to each GPU, how aggressively to invest in advanced packaging that puts memory closer to compute, and whether near-memory compute architectures pay off. The HBM market was already tight going into 2026. A 5x agent ratio, with 85% of it cached, makes the constraint tighter and the demand more structural. The industry's job now is to give agents somewhere fast to put what they have already said.
The OpenRouter data is not the global inference market, and the McKinsey 40% is a self-reported survey. Both can overstate. But the mechanism is consistent with what enterprise builders report and with what the chip physics demand. Agents are not just talking to themselves. They are reading themselves, and the next build dollar lands there.