Amazon's $1.8M Claude overrun is the visible symptom. The structural failure is that pay per AI call pricing killed the built in budget alarms enterprise IT used to rely on.
Amazon ran an internal project to match author names to product listings on its ecommerce site. It used Anthropic's Claude Sonnet to do the matching. The project cost about $1.8 million, ran 860% over its allocated budget, and went undetected for roughly five months. The Financial Times first reported the overrun; the figures have been relayed by Futurism and corroborated within 72 hours by Tom's Hardware, The Next Web, Cybernews, and ExtremeTech.
The author-matching project was not an isolated case. FT reported two other Amazon AI projects that ran materially over budget: a financial-auditing tool about $541,000 over, and a delivery-speed optimization effort about $134,000 over. Different teams, different problems, same shape: a token-billed AI tool with no real-time cost guardrails, deployed against a task cheap to spec out and expensive to actually finish.
That is the structural lesson. Token billing is the variable-cost pricing model AI providers now default to. Every prompt costs a fraction of a cent, every response costs more, and the invoice arrives at the end of the month. The legacy alternative, flat per-seat software subscriptions, produced a built-in cost-feedback loop. A team that added licenses showed up in procurement. A team that spun up a new cloud instance tripped a budget alarm. A team that added API calls under a flat enterprise contract did not. Under token pricing, cost is decoupled from the act of using the tool. The user does not pay. The team does not see the meter. Finance sees the total only after the run finishes.
A senior Amazon employee told FT that "it's difficult to figure out how much anything [AI related] costs." Senior engineers told the paper that mistakes that were once "trivially cheap" are now "catastrophically expensive." Both observations describe the same shift: the unit of waste got smaller (one API call), the aggregate got larger (millions), and the time-to-detection stayed the same as it was in the flat-license era, which is to say far too slow for usage-based pricing.
Amazon's internal culture amplified the gap. Staff ran a "tokenmaxxing" usage leaderboard, a gamified ranking of who was burning the most model tokens. The internal meter measured engagement, not cost. FT separately reported the leaderboard was shut down in May. The cost overrun the leaderboard was incenting had already landed.
Leadership posture added a second layer. In January 2026, Amazon announced 16,000 layoffs while publicly attributing the move to AI "efficiency gains." An Amazon spokesperson pushed back on FT's framing: "Cherry-picking small, isolated examples where teams are learning from one another and portraying them as business as usual doesn't reflect how teams across Amazon are using AI." The pushback works on the narrow point: a few over-budget projects do not mean Amazon's whole AI program is broken. It works less well on the broader one. Detection lag, internal usage pressure, and AI-as-layoff-rationale form a coherent incentive problem, not a string of one-team mistakes.
The same failure mode is showing up elsewhere. Earlier reporting described an unnamed company running roughly $500 million in Claude usage in a single month. The figure is unverified and the company unnamed, but the order of magnitude is the point. When the largest enterprises on the planet can run nine-figure single-model bills in a single billing cycle, the question is not whether other Amazon-scale overruns are happening. It is which finance teams will be next to discover them.
Amazon's $1.8M invoice landed in May. The companies that beat Amazon to the punch will be the ones that built per-team cost dashboards before the meter started, not after the invoice arrived.