Decagon says 90% of its production AI runs on open source models, while open source's share of enterprise LLM spend has dropped from 19% to 11%, a gap the company pins on per workflow model customization and response speed, not cost.
Decagon, the AI customer-service company, runs 90% of its production AI on open-source models. That's not a cost story. It's a fine-tuning story.
Decagon's CEO Jesse Zhang made the claim on an a16z podcast, in the same conversation where he noted that open-source had fallen to 11% of total enterprise LLM spend, from 19% a year earlier. The two numbers are not a contradiction. They are the same story.
Workload share and spend share are splitting because the parts of enterprise AI that scale in production are the parts frontier APIs cannot serve: the workflows that need model weights they can edit, latency they can tune, and fine-tuning that goes deeper than any public endpoint allows.
In a blog post that drew the a16z invitation, Zhang frames the gap as a deployment problem, not a price problem. Decagon sells AI customer-service agents to large enterprises. The agents have to ship answers in single-digit seconds for the conversation to feel like service rather than a queue. An eight-second response is the kind of delay that pushes a customer to a second channel. Most open-source models, run on the right hardware, clear that bar. Frontier APIs from Anthropic and OpenAI cannot be tuned below a multi-second floor at Decagon's traffic, and they cannot be fine-tuned to the per-task granularity the company's customers need.
The 8-second customer-service turn is the unit of failure, and the bar does not show up in any benchmark. It is also not a constraint that prompt engineering or retrieval-augmented generation can solve from outside the model weights.
Zhang's broader claim is that frontier and open-source are two phases of one lifecycle. Frontier models are the discovery and prototyping layer, where new capabilities show up first and new agentic patterns get tested. Open-source is the production layer, where the cost-per-call math gets done, where fine-tuning gets scoped to a specific workflow, and where latency gets engineered into the contract.
The dashboard data backs the split, with caveats. On Vercel's AI gateway, a developer platform that routes AI traffic for application builders, according to TechCrunch, DeepSeek processes just over a third of total token volume, but Anthropic still captures more than half of the platform's total AI spend. Z.ai's GLM-5.2 has jumped to fourth on the same dashboard; Nvidia's Nemotron is positioned to surge. On OpenRouter, an API router that aggregates multiple model providers, DeepSeek V4 Flash processes roughly 5.3 trillion tokens a week against Opus 4.8's just-over-2 trillion, and Opus 4.8 averages about $1.37 per million tokens against V4 Flash's roughly $0.06, a gap of about 23x.
Anthropic still collects more spend than the token volume implies, so the spend share is not collapsing. The workload question is separate, and the enterprise AI builder's answer there is different from the spend-share answer.
Decagon's third-party credibility on the open-source workload claim is thin. The 90% figure is a first-party self-report. The 11% spend share is the company's read of a market TechCrunch did not directly measure, and the a16z X thread that publicized the conversation is a podcast host's summary, not a transcript. A third of the Fortune 500 has, per Decagon, verified accounts on Hugging Face. The claim has the texture of a category bet, not a measurement.
The enterprise AI market is answering the open-source-versus-frontier question at every layer instead of once. At the top of the spend stack, frontier APIs still collect rent. At the bottom of the latency stack, open-source is the only answer that ships.
The watch item is whether frontier labs can move the two-second latency floor and offer fine-tuning at the per-workflow depth their customers actually want. If they can, the 90% workload figure starts to look like a temporary deployment decision rather than a structural shift. If they cannot, the 11% spend share starts to look like a static tax on the part of the stack frontier APIs can still serve.