Vals Raises $40M to Grade AI Models in Secret
The benchmarking startup Vals closed a $40 million Series A led by Andreessen Horowitz, on a simple premise: publish your test and labs will train against it. Vals keeps its questions confidential and scores models on real work in law, finance, coding and biosecurity. The labs being graded are also the ones paying.
Vals, a two-year-old startup that tests what AI models can actually do, has raised a $40 million Series A led by Andreessen Horowitz. The company was founded in 2024 and had previously taken a seed round from 8VC and Bloomberg Beta. Its pitch is narrow and increasingly hard to argue with: the public benchmark, as an instrument, is broken, and it is broken for a boring reason.
The reason is that a published test is a published answer key. Once a benchmark is public, its questions leak into training data — sometimes deliberately, often just because the whole web gets scraped — and a rising score stops meaning the model got better at the underlying skill. Vals' response is to keep its test materials confidential and never release them. Models are scored on their ability to finish domain-specific work in law, finance, coding, cybersecurity, biosecurity and mental health, among other areas. "What we're doing is actually looking at what are the real impacts of the models," co-founder Rayan Krishnan said. "Can they do work that produces a product of the same quality as a human within every domain?"
Krishnan is 25, and interned at Palantir before working at Microsoft and Stanford's artificial intelligence lab. The business has grown on the back of demand that barely existed when he started: headcount went from eight to 25 in nine months, revenue is eight times what it was a year ago, and the company plans another 10 to 15 hires and a larger office. It has also launched a programme offering evaluations to federal agencies, which is a different kind of customer from a model lab and a useful hedge.
The obvious objection is the one Krishnan answers head-on, and his answer is worth examining rather than accepting. Model developers pay Vals to be tested, which he likens to "a student might pay the College Board to take the SAT." The analogy holds in one respect — the College Board is paid by the tested and is still broadly trusted — and strains in another, because the College Board publishes retired exams and its scoring methodology, while Vals' entire value proposition rests on not publishing either. A confidential test graded by a vendor the test-taker pays is asking for a lot of trust, and the mechanism that would normally supply it, independent replication, is the one thing the model rules out.
That tension is not a reason to dismiss the company so much as the thing to watch. Krishnan's bet is that the demand for third-party measurement is about to stop being a research-community nicety: as AI companies go public, he argues, rigorous outside benchmarking becomes material to investment decisions and eventually to public filings. That is a real market, and it is one where an auditor's independence is the product. It also means Vals is trying to occupy the role of auditor at the same moment that the labs are beginning to embed evaluators of their own — a crowded position to claim, and one where the first serious dispute over a score will settle more than any funding round.
For readers trying to work out which model to use, none of this yet replaces reading the numbers with suspicion. Public benchmark tables remain the only figures most people can actually see, which is why we keep running head-to-head comparisons rather than quoting a single index. Vals is betting that within a couple of years, the score that matters will be one nobody outside the building can see.
Want AI news before everyone else?
The morning's most important AI stories, straight to your inbox. No fluff.