AI is splitting into millions of small, task specific models. The $100 billion question is who runs the switchboard that picks the right one for each job.
Fireworks AI, a four-year-old company that runs other companies' AI models rather than training one of its own, raised $1.5 billion at a $17 billion valuation in the week before a recent episode of 20VC aired. The firm says it has crossed $1 billion in annual revenue with roughly 200 employees. The bet underneath those numbers is structural: that the next dollar of AI infrastructure spending goes to running specialized models in production, not to training bigger generalist ones.
Lin Qiao, Fireworks' co-founder and CEO, made the case on 20VC that token costs will fall 10x and usage will rise 100x over the next two years. Tokens are the units AI companies charge for, the way phone companies once charged per minute. The shape of the claim is the point: it says demand is being throttled by price, not by appetite, and that the throttle is about to come off.
The bet is that demand does not settle on one generalist brain. It settles on millions of small, specialized models. Each one does a single job: a customer-support bot, a code reviewer, an image-tagger, a contract-clause extractor. Each is often fine-tuned on a few thousand examples and runs at a fraction of the cost of a frontier model on that specific task. Lin Qiao has made the same case on Sequoia's Training Data podcast and to Turing Post. The Observer profiled Fireworks in July as the visible proxy for the inference side of the market.
A market built on millions of models has a problem none of them can solve alone. Pick the wrong model for a customer-support ticket and the answer is wrong. Pick a frontier model for a tagging job and the bill is an order of magnitude higher than it needs to be. Something has to choose, in milliseconds, which model to call for which request. That something is the routing layer, and Lin Qiao put a number on it: a $100 billion market, in the 20VC discussion, if the specialized-models thesis plays out. The number is a speaker estimate, not an analyst consensus.
The routing layer does not exist as a clean category today. It is partly inside inference platforms like Fireworks, partly inside model aggregators, partly inside the API gateways enterprises build themselves. Whoever owns it sets the price the end user pays and decides which model gets the traffic. That position is more valuable than any single model in the stack, because it is the only one that gets to see every request.
OpenAI and Anthropic have the brand and the frontier capability. The 20VC frame is that they are overvalued at current multiples because the market they dominate is going to be sliced up by routing. Lin Qiao is a self-interested party: she runs one of the companies that would benefit if the thesis is right. Read the claim as the Fireworks thesis, not as a settled view of the labs' business model. The structural shift underneath it is real, even if her number on it is one data point.
The falsifier is visible too. If generalist models keep winning on capability, or if enterprise buyers prefer one-vendor simplicity over cost optimization, the routing layer stays fragmented and the specialized-model thesis becomes a cost story rather than a market-structure one. The next 18 months will be the test. If token costs fall 10x and specialized models become the production default, the routing question stops being theoretical. Whoever owns the switchboard collects the toll on every call.