Inference platforms and coding agents are stacking their own discounts on top of DeepSeek's $0.14 list price. The price war has moved past the model itself.
DeepSeek set the AI price floor at $0.14 per million input tokens on July 31. Within days, the middlemen were stacking their own discounts on top. The price war had stopped being about the model.
The new floor showed up in DeepSeek's V4-Flash-0731, a post-training rerun of the same architecture and model size as the April 24 V4 Preview, but with materially stronger agentic and coding scores. On the V4-Flash-0731 model card, DeepSeek reports DeepSWE moving from 7.3 to 54.4, Cybergym at 76.7, Terminal-Bench 2.1 at 82.7, and an Artificial Analysis Intelligence Index of 50, up 10 points and ahead of DeepSeek's own V4 Pro. The Chinese outlet QbitAI framed the public price as $0.14 per million input tokens on cache miss and $0.28 per million output tokens, the lowest figure the publication has documented for a model it classes as frontier-tier.
The interesting part is what happened next.
Within 48 hours of the rerun, NousResearch and Novita AI announced a 90% off promotion on V4-Flash-0731 through Nous Portal, a seven-day run that started August 2. The partner claim, repeated on Novita's post: more than 1000x cheaper than Fable 5 on comparable tasks while still beating it on Terminal-Bench 2.1. Per the same QbitAI re-report, Cline, the open-source coding agent, tripled its free V4-Flash credits after observing lower-than-expected burn, and OpenCode reported 8 trillion V4-Flash tokens running through it on August 1, per a post on X.
That is not a list-price story. It is a subsidy stack.
When the model maker's own list price collapses, the platforms that sit between the model and the developer, inference portals, coding agents, IDE plugins, stop competing on access and start competing on subsidy. Nous Portal is paying Novita's margin out of pocket to undercut the rest of the market for a week. Cline is making the same bet on usage, trading short-term revenue for a wider free tier in the hope of retention. The user's effective cost falls faster than the model maker's price cut, because each layer of the stack is willing to take a short-term loss to capture the demand the price cut just unlocked.
For a developer, the practical version is short. A 3D space-shooter demo running on V4-Flash-0731 costs about 7 cents in API fees per play, in QbitAI's worked example. A full Counter-Strike match costs about 50 cents. A month of prototyping a real product on the model runs in the $30 range, the cost of a few monthly subscriptions. That puts an idea's cost of iteration in the same band as the cost of hosting the project, and it changes what a weekend hack can be.
The model maker is not the only one writing a check. The stack is.
The caveats are real and worth naming. V4-Flash-0731 is a post-training rerun of the same architecture DeepSeek shipped in April, not a new model, so the benchmark jump is real but the gain is one of post-training rather than pretraining. The $0.14 / $0.28 price point is sourced from a re-report attributing it to "a research firm," likely Artificial Analysis; DeepSeek's own pricing page should be checked before treating it as settled. The ">1000x cheaper than Fable 5" line is a partner marketing claim on a single benchmark. Cline's tripled free credits and OpenCode's 8 trillion tokens are self-reported platform metrics, unaudited. The QbitAI piece is a Chinese-AI-media trend story with a triumphalist frame; the subsidy-stack mechanism is Type0's argument, not the source's claim.
The question to watch next: if Nous Portal's seven-day promo expires and the platforms revert to passing through DeepSeek's list price, the stack collapses and the week becomes an anomaly. If it persists, the next round of frontier releases will ship into a market where the model is the floor, not the ceiling, and the price war stays down the stack for good.