Vals aims to make AI benchmarking more trustworthy with $40M Series A

Vals, a two-year-old startup, is working to create more reliable AI benchmarks by keeping its test materials private and focusing on complex tasks. The company recently raised $40 million in a Series A round led by Andreessen Horowitz. Its co-founder argues that traditional academic benchmarks fail to measure the capabilities of modern AI models.
Vals' approach centers on keeping test materials confidential, preventing AI developers from training models against known benchmarks. The startup evaluates models on industry-specific tasks in law, finance, and coding, rather than general knowledge tests. Krishnan, who gained experience at Palantir, Microsoft, and Stanford's AI lab, founded the company after observing that academic benchmarks lagged behind rapid model development.
The company's evaluations extend into emerging areas including recursive self-improvement, mental health, cybersecurity, biosecurity, and even application of the Geneva Convention. Vals generates revenue by charging companies to test their models, a model Krishnan compares to paying for standardized testing. The startup reports revenue eight times higher than last year and has grown from eight employees at the start of the year.
Vals' private benchmarking approach could reshape how AI capabilities are verified across industries. If adopted widely, it may give enterprises and regulators more reliable signals about model performance, potentially influencing purchasing decisions and deployment choices. However, the reliance on private benchmarks could also concentrate evaluation power in a single company, raising questions about transparency and accountability in how AI systems are assessed and compared.