Companies·3 min read
By BitsMindsSource: TechCrunch

Vals Raises $40M to Grade AI Models in Secret

The benchmarking startup Vals closed a $40 million Series A led by Andreessen Horowitz, on a simple premise: publish your test and labs will train against it. Vals keeps its questions confidential and scores models on real work in law, finance, coding and biosecurity. The labs being graded are also the ones paying.

Vals — the benchmark vault An editorial still life of a deep teal steel vault with a copper combination lock, confidential evaluation papers and an AI processor on an ivory surface. The sealed vault represents test materials kept out of model training. The forty-million-dollar Series A is noted discreetly above the scene. All artwork is original vector geometry; the document markings are illustrative and do not represent actual test results. Original SVG illustration for BitsMinds. Article: https://www.bitsminds.com/news/vals-40m-series-a-confidential-ai-benchmarks. Vector only; no external resources. VALS THE BENCHMARK VAULT $40M / SERIES A EVALUATION CONFIDENTIAL BITSMINDS.COM PRIVATE TESTS. REAL CAPABILITIES.
Share:

Vals, a two-year-old startup that tests what AI models can actually do, has raised a $40 million Series A led by Andreessen Horowitz. The company was founded in 2024 and had previously taken a seed round from 8VC and Bloomberg Beta. Its pitch is narrow and increasingly hard to argue with: the public benchmark, as an instrument, is broken, and it is broken for a boring reason.

The reason is that a published test is a published answer key. Once a benchmark is public, its questions leak into training data — sometimes deliberately, often just because the whole web gets scraped — and a rising score stops meaning the model got better at the underlying skill. Vals' response is to keep its test materials confidential and never release them. Models are scored on their ability to finish domain-specific work in law, finance, coding, cybersecurity, biosecurity and mental health, among other areas. "What we're doing is actually looking at what are the real impacts of the models," co-founder Rayan Krishnan said. "Can they do work that produces a product of the same quality as a human within every domain?"

Krishnan is 25, and interned at Palantir before working at Microsoft and Stanford's artificial intelligence lab. The business has grown on the back of demand that barely existed when he started: headcount went from eight to 25 in nine months, revenue is eight times what it was a year ago, and the company plans another 10 to 15 hires and a larger office. It has also launched a programme offering evaluations to federal agencies, which is a different kind of customer from a model lab and a useful hedge.

The obvious objection is the one Krishnan answers head-on, and his answer is worth examining rather than accepting. Model developers pay Vals to be tested, which he likens to "a student might pay the College Board to take the SAT." The analogy holds in one respect — the College Board is paid by the tested and is still broadly trusted — and strains in another, because the College Board publishes retired exams and its scoring methodology, while Vals' entire value proposition rests on not publishing either. A confidential test graded by a vendor the test-taker pays is asking for a lot of trust, and the mechanism that would normally supply it, independent replication, is the one thing the model rules out.

That tension is not a reason to dismiss the company so much as the thing to watch. Krishnan's bet is that the demand for third-party measurement is about to stop being a research-community nicety: as AI companies go public, he argues, rigorous outside benchmarking becomes material to investment decisions and eventually to public filings. That is a real market, and it is one where an auditor's independence is the product. It also means Vals is trying to occupy the role of auditor at the same moment that the labs are beginning to embed evaluators of their own — a crowded position to claim, and one where the first serious dispute over a score will settle more than any funding round.

For readers trying to work out which model to use, none of this yet replaces reading the numbers with suspicion. Public benchmark tables remain the only figures most people can actually see, which is why we keep running head-to-head comparisons rather than quoting a single index. Vals is betting that within a couple of years, the score that matters will be one nobody outside the building can see.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles

Crusoe: capital for the physical infrastructure of AI An original editorial cutaway of a modular AI data centre at night. Eight illuminated racks stand behind an open structural frame, alongside cooling equipment and a transformer connected by an amber power route. The illustration presents the article's 3.9 billion dollar Series F and 30.9 billion dollar post-money valuation. Separate labels distinguish 1 gigawatt operational from more than 6 gigawatts of gross contracted capacity across Crusoe's platform. The building is an illustrative concept, not a technical drawing or an exact depiction of a Spark product; its rack count has no quantitative meaning. BitsMinds editorial vector artwork. Article: crusoe-3-9b-series-f-ai-factories. 18 September 2026. Self-contained SVG. Funding and capacity figures reproduced from the article. CRUSOE $3.9B SERIES F $30.9B POST-MONEY VALUATION AI FACTORY 1 GW OPERATIONAL 6+ GW CONTRACTED POWER → COMPUTE BITSMINDS.COM
Companies

Crusoe Raises $3.9B at $30.9B for Its AI Factories

Novo puts Claude Science into drug discovery CLAUDE SCIENCE BITSMINDS.COM
Companies

Novo Nordisk Puts Claude Science Into Drug Discovery

Anthropic's planned Queensland data-centre lease: a campus and its power demand A conceptual evening view across agricultural fields on the Darling Downs. Four proposed data halls occupy a site, with the first-stage hall highlighted and carrying the Anthropic name. The other three halls appear as translucent planning outlines. Cooling equipment, an electrical substation and tall transmission towers connect the campus to the landscape. A small caption identifies the 2.16-gigawatt projected peak draw as applying to the fully built campus. This is an editorial illustration of a development awaiting approvals, not an actual site plan, construction progress image or claim that Anthropic owns the campus. Original BitsMinds vector illustration for anthropic-queensland-western-downs-data-centre. 17 September 2026. Site geometry is illustrative; first-stage lease and full-campus power demand are separate. 03 04 02 ANTHROPIC STAGE 01 ANTHROPIC WESTERN DOWNS / QUEENSLAND PROPOSED DEVELOPMENT 2.16 GW PLANNED PEAK / FULL CAMPUS BITSMINDS CONCEPTUAL ILLUSTRATION
Companies

Anthropic Anchors Australia's Biggest Data Centre