DeepSeek's open source model, running on two Nvidia boxes in a Beijing bar, is forcing OpenAI and Anthropic into a price war over the AI that powers ChatGPT style assistants.
At the AGI Bar in Beijing's Zhongguancun district, owner Song De keeps two Nvidia workstations humming behind the counter. The boxes run DeepSeek's V4 Flash, a Chinese-built large language model (the same kind of technology that powers chatbots like ChatGPT), and serve free AI tokens to anyone who walks in. "Mostly foam," customers joke about the beer. The tokens are the real draw.
This is the human face of a quiet shift in the AI economy. OpenAI, the US lab behind ChatGPT, and Anthropic, the maker of Claude, are now competing on cost rather than capability. The pressure comes from Chinese open-source models, meaning AI systems whose internal settings are published so anyone can run them on their own hardware, that have crossed a "good enough" threshold for the work real products actually ship.
OpenAI made the new math visible last week when it cut prices on its GPT-5.6 lineup by as much as 80%, according to the company's own announcement. The smaller "Luna" tier dropped sharply; the heavier "Terra" tier dropped 20%. OpenAI framed it as passing efficiency gains to customers. The framing is true. It is also a defensive margin move in a market where the comparison point now sits in Hangzhou.
Anthropic answered the same week with Claude Sonnet 5, positioned as close to its previous flagship Opus 4.8 on the agent-style work that AI carries out on its own (booking tools, multi-step planning, browser navigation), and at lower prices. The published pricing page shows the gradient: the current top tier, "Fable 5," lists at $10 per million input tokens and $50 per million output tokens, while the older Opus 4.x line sits at $5 and $25. Sonnet 5 lands below the older flagship, not above it. That is a different pricing posture than Anthropic has used before.
The pressure shows up in usage data, not just pricing pages. DeepSeek's V4 Flash topped OpenRouter's weekly ranking with 7.22 trillion tokens routed through the developer-facing inference marketplace in the week ending August 5, according to OpenRouter's own writeup. Developers vote with their API calls. The votes are not going to the most expensive model on the menu.
Independent analyst Jack Gold of J.Gold Associates told AFP that "older models" and "small language models" can handle much of the work frontier systems were assumed to monopolize. Song De, the AGI Bar owner, said the same thing in different words: he swapped a frontier "bells-and-whistles" model for V4 Flash because the cheaper one was good enough for what his customers actually asked it to do.
Two caveats shape the read. DeepSeek's plan to raise programmer-tier prices "significantly" is a single-source claim carried by the AFP wire, not yet echoed in a DeepSeek primary statement; if the increase lands, it would test whether open-weight economics are sustainable at the current price floor. The second is that an 80% headline cut is not the same as an 80% margin reset. OpenAI and Anthropic are still funding the next training run. Cutting inference prices is what you do when you can no longer charge a scarcity premium for them.
The next milestone to watch is Anthropic's Opus-line pricing later this quarter. If the company holds the line above Sonnet 5, the price war has tiers. If it does not, the AI race has become a margin contest with the cheapest useful model on top, and US frontier labs are no longer competing on who has the best one.