Skip to main content
Live
Main content

Vals raises $40M from Andreessen Horowitz to remake AI benchmarking

The two-year-old startup keeps its tests private, sells evaluations to model makers, and grew revenue 8x year over year.

Jaeden Schafer
Editor in Chief · · 5 min read
Vals raises $40M from Andreessen Horowitz to remake AI benchmarking

Vals, a two-year-old AI evaluation startup, raised $40 million in a Series A led by Andreessen Horowitz last month, betting that the industry's current benchmarking regime is broken and that private, industry-specific tests will replace it. Revenue at the San Francisco company is running 8x what it was a year ago, and its staff has tripled from 8 people at the start of the year to 25. The round followed a 2025 seed from 8VC and Bloomberg Beta.

Co-founder Rayan Krishnan, 25, previously interned at Palantir and worked for Microsoft and Stanford's AI lab as an undergraduate. He argues that widely cited academic benchmarks have failed to track how quickly frontier models are advancing, leaving buyers with little reliable signal about what a given model can actually do in production.

We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance.
Rayan Krishnan, Vals co-founder

The pitch to model developers is that public benchmarks are gameable. When a test's questions are freely available online, they inevitably leak into training data, and scores stop measuring capability and start measuring memorization. Vals does not publicly disclose its specific test materials, which is meant to make contamination structurally impossible.

Key facts

  • 01Vals raised $40M in a Series A led by Andreessen Horowitz last month, following a 2025 seed round from 8VC and Bloomberg Beta.
  • 02Revenue is currently 8x what it was last year, and headcount tripled from 8 to 25 people since January.
  • 03The company was founded in 2024 by 25-year-old Rayan Krishnan, a Stanford graduate who previously interned at Palantir and Microsoft.
  • 04Vals keeps its test materials private to prevent training-set contamination, unlike most academic benchmarks.
  • 05New benchmarks cover recursive self-improvement, cybersecurity, biosecurity, mental health, and Geneva Convention compliance.

The second pitch is task realism. Rather than measuring whether a model can pass a bar exam or answer trivia, Vals scores models on their ability to complete domain-specific work in law, finance, and coding, comparing model output to what a human professional would produce.

Krishnan describes the evaluation target as the actual product of the work, not an abstract proxy for intelligence. The company also scores for negative outcomes — what happens when a model behaves badly in the wild — rather than only positive completions.

The business model is unusual: Anthropic, OpenAI, and other model providers pay Vals to test their own systems. Krishnan compares it to a student paying the College Board to sit the SAT. A poor score costs the vendor short-term marketing material but gives them a private map of where to improve, and a strong score becomes a defensible external validation buyers can trust.

The benchmark surface area is widening quickly. Beyond law and finance, Vals has built evaluations for recursive self-improvement, mental health, cybersecurity, biosecurity, and even a test of whether models correctly apply the Geneva Convention to scenarios in the law of armed conflict.

The company also recently launched a program supplying model evaluations to federal agencies, a market where procurement officers need defensible, third-party numbers before signing contracts with frontier labs. Krishnan said Vals plans to add another 10 to 15 people and move into a substantially larger office. The current headquarters is a two-floor space on San Francisco's Folsom Street inside an old brick building that was a brewery a century ago.

Related · from this week
AI-linked Asian stocks slump after lab CEOs call to slow AI development
Jaeden Schafer · 4 min read →

The skeptical read on paid benchmarking is that a vendor who pays for its own report card has leverage over the grader, and Vals will need to publish enough methodology and comparative results to convince buyers the scores are independent. There is also a competitive question: Scale AI, Epoch, and several academic groups are moving into private, capability-oriented evaluations, and no single provider has locked in the role of trusted third party.

Krishnan is betting that evaluations become financial infrastructure as AI companies enter public markets. SpaceX went public, Anthropic is slated to list later this year, and he expects OpenAI to follow. Benchmarks that hold up to auditor and investor scrutiny will underpin S-1 filings, revenue-guidance calls, and capital-allocation decisions across the sector — and whichever firm sets that standard first captures a durable position at the center of the AI economy.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Business

AI-linked Asian stocks slump after lab CEOs call to slow AI development
Business

AI-linked Asian stocks slump after lab CEOs call to slow AI development

Chip suppliers and AI infrastructure names across Asia sold off after frontier lab chiefs publicly urged the industry to slow down.

Jaeden Schafer4 min read
AI spend per employee fell 10% at top firms in August, Ramp data shows
Business

AI spend per employee fell 10% at top firms in August, Ramp data shows

The top 1% of Ramp customers cut AI spend per employee to $7,205 as OpenAI and Anthropic price wars bite into token revenue.

Jaeden Schafer5 min read
Microsoft logo
Business

Microsoft coaches sales team to pitch against OpenAI and Anthropic

Executives at an internal FY27 strategy meeting told salespeople to frame Copilot as faster and more secure than Claude and rival models.

Jaeden Schafer4 min read