Frontier model pricing is decoupling from composite benchmark scores, and the gap is the real story. The composite index the field has been using to compare flagship models compresses agentic work into a single number, so a release that beats the next frontier model on the work people actually pay for can still read as "roughly even" on the headline index.
Anthropic launched Claude Opus 5 on Friday 2026-07-24, pricing it at roughly half of what the AI community expects the next frontier model, "Fable," to cost. The lab's own framing was that Opus 5 "comes close" to that unreleased target. The agentic knowledge-work benchmark AA-Briefcase tells a different story: Opus 5 now leads the field, outperforming Fable 5 by nearly 150 Elo while costing 20% less per task. Epoch's broader index still has the two models within two points of each other, with their software-engineering subset tied.
Composite scores appear to weight many evals equally, which can bury an agentic jump that lowers per-task cost underneath the curve. When a flagship model wins on the eval that maps to billing but loses nothing on the others, the composite can stay calm while the price has already moved. Anthropic shipped a product that outperforms its own framing — whether or not that was the pricing intent.
For anyone watching frontier pricing, the watch is the agentic eval, not the composite. The composite will catch up. The price already has.
Reported by Sky for Type0, from Artificial Analysis on AA-Briefcase Opus 5 vs Fable 5. Read the original: x.com