Fireworks AI says it runs 40 trillion AI tokens a day — the small chunks of text that AI models read and write — more than OpenAI.
Lin Qiao says her company processes roughly 40 trillion tokens a day, the small chunks of text that AI models read and write. She framed the number, on a recent Weights & Biases podcast episode, as exceeding OpenAI's daily API volume. Treat the figure as a CEO-attributed claim, not an audited benchmark. The mechanism behind it is the actual story.
Fireworks runs in a layer of the AI stack that does not get much public attention: the inference vendor, the company that hosts already-trained models for other companies' products, handling traffic, latency, and customization at production scale. It is the part of the stack that turns a research checkpoint into a feature inside someone else's app. Model labs like OpenAI, Anthropic, and Google train the underlying systems; inference vendors are what run them once they are deployed.
According to BigGo Finance's writeup of the conversation, Lin said 95% of the traffic on Fireworks runs through customer-customized or fine-tuned models, not stock checkpoints. Customers take an open-weights model, adjust it on their own data, and deploy the result through Fireworks' infrastructure. The 40T figure, on that read, is largely a measure of how much customized inference is now flowing through third-party platforms, and how much of the AI value chain is migrating from training to serving.
Investors spent the past year asking when the hundreds of billions in AI capex would produce profits. Per BigGo Finance's reporting, Fireworks closed a $1.5B Series D in 2025, betting the answer is yes, on the serving side, not the training side. The pitch is straightforward. Companies do not want a raw model. They want a model tuned to their data running reliably at low latency. That is a different business from building the underlying model, and it scales with how many products ship AI features, not with how many new models are trained.
Lin built Fireworks with six other PyTorch alumni, including co-founders from the framework's core team, and led infrastructure at Meta AI before that. Fireworks has framed itself less as an inference provider and more as a specialized intelligence platform, a positioning that signals the company sees the inference layer as a market of its own rather than a side service. The host of the podcast, Lukas Biewald, presses Lin on these points. The conversation is two people with deep AI infrastructure experience discussing the space.
Daily token totals are not a standardized metric. Different vendors count prompts only, completions only, or both. A direct comparison to OpenAI and Google depends on whose definition of a token you trust. The 95% customization share is Fireworks' own platform attribution, and may overstate fine-tuning in the strict sense versus off-the-shelf open-weights deployments routed through the same platform. BigGo Finance is summarizing a podcast, not conducting an independent audit.
If a substantial share of that 40T is genuinely customized inference, the comparison to OpenAI is the wrong one. The right one is to the model labs' training compute budgets, which the inference layer is now rivaling in customer dollars. The 40T figure is a hook. The 95% figure is the business.