AI data center spend is on track for $1.1T by 2027, and Wharton's Jessica Wachter says breaking even by 2030 needs an 'eye opening' productivity jump. OpenAI's nonprofit, meanwhile, is funding bids for failed biotech data.
Hyperscalers are projected to spend roughly $1.1 trillion on AI data centers by 2027, and a Wharton finance professor's model says breaking even by 2030 will need an "eye-opening" jump in productivity growth.
Jessica Wachter, a finance professor at the University of Pennsylvania's Wharton School, published the modeling last week in an NBER working paper, What Investment Data Implies about the AI Transition. The trillion-dollar figure is a projection of total AI-infrastructure spend across Microsoft, Google, Amazon, and Meta through 2027, not a single commitment from any one firm. MIT Technology Review's newsletter covered the work alongside two unrelated items, and the break-even question is the load-bearing one.
The framing is "eye-opening" because the implied productivity growth is large enough to look implausible. Wachter's point is that the capex bill is large enough that the standard "AI saves time" narrative does not pay for it on its own. The companies need AI to change how other industries operate, not just how their own engineers write code.
The OpenAI Foundation, the nonprofit parent of the for-profit lab, announced the same day that it is backing a parallel bet on the same constraint. The nonprofit-vs-for-profit distinction matters because the foundation can fund work the company cannot. Policy analyst Ruxandra Teslo's proposal: bid at the bankruptcy proceedings of failed biotechs, harvest the regulatory filings, manufacturing protocols, and safety data those companies filed before they went under, and curate that material into what she calls "biotech's lost archive." The Public Data for Health program funds the curation work, not the bankrupt companies themselves.
Read together, the shared shape is a scarcity of returns. The hyperscalers have committed more capital than their current customers can repay. The frontier models, the most capable systems currently being trained and deployed, need to convert that capital into returns and are running out of high-quality training data. The OpenAI Foundation's bet is that the bottleneck is the data layer, and that the data already exists inside companies that failed for reasons unrelated to the data itself.
The fair objection is that the two storylines may simply be unrelated. Capex break-even is a demand-side return question; a bankruptcy-data acquisition is a supply-side question. "Scarcity" may be coincidence, not mechanism. If the synthesis does not hold, the most that can be said is that both items sit at the same AI inflection point, not that they share a single constraint.
If the synthesis does hold, the framing cuts against a common read of the AI buildout as a simple arms race. The question becomes which lab turns the inputs it already has into capability faster, and how quickly Wachter's productivity curve actually shows up in the rest of the economy.
The same MIT Technology Review digest also flagged an AI-extinction roundtable with the magazine's AI desk: Niall Firth, Will Douglas Heaven, and Grace Huckins, available on demand to subscribers. The recording belongs in the same conversation about AI's constraints, but a short article cannot do the argument justice, and the extinction framing is not what the data- and capital-side questions turn on.
The watch items are dated. Wachter's break-even clock runs to 2030. The OpenAI Foundation's data-curation bets take longer than a model release to pay off: bankruptcy proceedings are slow, regulatory filings are large, and the dataset has to be cleaned and licensed before it trains anything. If the hyperscaler capex curve keeps climbing past the 2027 projection and the bankruptcy-data pipeline does not produce a usable corpus within roughly two model cycles, the trillion-dollar bet will be tested against a productivity curve that has not yet appeared.