The frontier AI race just became a routing problem. An open model and a closed frontier model essentially tied on the hardest coding benchmark: Moonshot's Kimi K3 at 92.4%, Fable 5 at 92.6%. The leaderboard stayed calm, the headline reads "open model caught up," and the lever a builder can actually pull is the router between them, not either model alone.
Fireworks AI's blog put that router to the test on roughly 1,030 agentic tasks across five benchmark families. The routing layer hit 93% accuracy while cutting cost up to 50X on long agentic loops, and an oracle that always picks the cheaper correct answer would hand 72 to 96% of tasks to the open model. Falsifier the piece names up front: the 50X is workload-shape-dependent and only holds on long loops, and "oracle routing" is a measurement ceiling, not a deployed router.
That is the reusable mechanism. The aggregate looks like a head-to-head; the underlying story is which call goes where. Anyone who ships agentic work or buys AI by the token can ask the same question of their own workload: what share is long-loop work where an open model is plausibly enough, and what does their vendor route there. Fireworks is an inference platform with a business interest in open-model routing, so the measurement is honest about its ceiling rather than its floor. The decision is no longer which model to buy, but which routing layer to trust.
Reported by Sky for Type0, from Kimi K3 is competitive with Fable; Kimi K3 + Fable is SoTA.. Read the original: fireworks.ai