Benchmarks used to be the closest thing AI had to a credit score: a public number that priced which AI lab had the best model and therefore who got the next contract. That unit of account is breaking under two compounding forces. Training data now contaminates test sets, so a model has effectively seen the answers; frontier labs also optimize directly against known evaluation targets, so score differences compress long before model quality does. Once the public scoreboard stops moving, neither vendor selection nor procurement has anywhere left to point.
The ETH Zurich and Stanford preprint quantifying this found 29 of 60 widely used tests now show the top model and the runner-up inside measurement noise. Ten days later, Andreessen Horowitz led a $40 million Series A in Vals AI at a $400 million valuation, with 8VC, Pear VC, and Bloomberg Beta returning. Vals AI grew revenue eightfold in 2025 and doubled its customer count. The company's half-year-old finance suite shows frontier models topping out at roughly half of professional analyst tasks, which is the gap between leaderboard scores and real-work delivery priced in dollars.
Researchers have identified two pathways. When public scoring loses resolution, money and procurement migrate to private evaluation against the buyer's own task. The agency-expanding move is to test the next AI vendor on the work it will actually be paid to do, before the contract is signed, and to weight the leaderboard accordingly.
Reported by Sky for Type0, from Vals AI Raises $40M From a16z: Frontier Models Fail 52% of Real Finance Analyst Tasks. Read the original: techtimes.com