Vals just raised $40M from Andreessen Horowitz to build private benchmarks that frontier models cannot train against, a bet that the next phase of AI evaluation has to move underground.
Public AI benchmarks have a problem the labs that build them cannot solve from inside. As soon as a test is published, the frontier models train against it, and the score stops meaning what it said. A 2024 startup called Vals is trying to break that loop by hiding the test.
The company raised a $40 million Series A led by Andreessen Horowitz last month, on top of an earlier seed from 8VC and Bloomberg Beta. The pitch is that legacy public benchmarks, including the MMLU, the GSM8K, and the long tail of academic test sets that show up in every model card, are now gamed by definition. Co-founder Rayan Krishnan, 25, told TechCrunch that the field has been measuring AI on the equivalent of a "bar exam," a static general-knowledge test that anyone with internet access can prepare for.
Vals' alternative rests on three moves. First, it does not disclose its specific test materials, which makes the train-to-test channel physically harder. Second, it grades models on complex, real-world tasks rather than abstract multiple-choice questions. Third, it goes deep in one vertical at a time, starting with finance, where a Vals-cited figure shows frontier models failing 52% of real analyst work.
That last number is the load-bearing one, and it is worth sitting with. If the best models money can buy are getting roughly half of a working analyst's job wrong, the public benchmarks that report them as near-human-expert on general reasoning are not lying, exactly, but they are answering a different question. They answer "can the model pass the test in front of it?" Vals is trying to ask "can the model do the job?"
The mechanism matters. The methodology page and the benchmarks page describe a private test library that is updated continuously, graded against actual work product rather than right answers, and built in partnership with practitioners who know what the work looks like. The academic anchor is a Finance Agent Benchmark paper on arXiv that evaluates LLMs on end-to-end financial research tasks. The leaderboard is tracked independently at the BenchLM Vals Index, which lets third parties see the scores without seeing the test.
This is a real, falsifiable bet about how evaluation should work, and the funding signal is real too. A $40M Series A from a16z is the kind of round that turns an argument into infrastructure. The bet is that the next phase of AI accountability will be built by firms that own their test sets the way Bloomberg owns its terminal data.
The counter-question is also real, and Vals does not fully answer it. If benchmarks go private, who audits the auditor? Vals publishes methodology and tracks scores on third-party leaderboards, but the test materials themselves are not open to outside inspection. The integrity gap the company closes on the model side stays open on the benchmark side, because the test itself is hidden. A 52% failure number is only as credible as the test that produced it, and the test is the part Vals is not showing.
Krishnan's framing, that the alternative is benchmarks everyone has already seen, is fair. But the answer to that problem does not have to be Vals' specific test library. It could be a public benchmark built from continuously refreshed, provenance-tracked tasks. It could be regulator-run test sets. It could be a consortium. The bet is that none of those will move fast enough, and that a private company with a $40M war chest and a finance-vertical first move will set the template before the alternatives exist.
Two things will tell whether the bet works. The first is whether the Vals Index becomes the page major labs paste into their model cards. The second is whether the methodology page grows teeth, with external audits and reproducibility studies, or stays a vendor document. The funding buys time to find out.