Vals, founded in 2024, raised $40 million in Series A funding led by Andreessen Horowitz to address a critical gap in AI model evaluation: benchmarks that actually measure real-world performance rather than gaming vectors.
The problem is straightforward. Most existing AI benchmarks—many designed years ago—are easily exploited. Companies train models directly against publicly known test sets, inflating benchmark scores without delivering equivalent capability on tasks that matter. Academic benchmarks have largely failed to keep pace with rapid model advances, leaving enterprise buyers with no reliable way to validate vendor claims.
Vals' approach: proprietary, task-specific evaluations across verticals like law, finance, and coding. The company does not disclose its test materials, preventing the training-against-benchmarks problem. Assessments measure whether a model can produce work comparable to a human expert in each domain, accounting for both quality and failure modes.
"Benchmarks should exist to confirm models perform as advertised," said Rayan Krishnan, Vals' 25-year-old co-founder, who previously worked at Palantir, Microsoft, and Stanford's AI lab. Krishnan spotted the benchmarking gap as capable new models entered the market faster than evaluation standards could adapt.
The competitive moat is structural. Once Vals establishes credibility as a neutral validator, enterprises will demand its certifications as a purchasing signal. Vendors, in turn, will pay for assessment and re-assessment cycles as they release new model versions. That creates recurring revenue with high switching costs—customers and vendors alike become locked into the standard.
The funding follows a seed round led by 8VC and Bloomberg Beta. Vals operates from San Francisco's Folsom Street.