AI benchmarking has become more than just metrics—it’s a public relations cornerstone. Models proving their worth through benchmark scores get headlines and investment. Yet many existing benchmarks struggle to keep up with today’s AI advances, leaving gaps that companies are all too willing to exploit.
Enter Vals, a benchmarking startup founded in 2024, now backed by venture heavyweights including a new $40 million Series A led by Andreessen Horowitz. Vals is positioning itself as the next-generation bridge between hype and accountability in AI model evaluation. It offers companies an alternative to outdated metrics that are no longer sufficient for assessing the real-world complexity AI systems face.
Why traditional benchmarks are no longer enough
The problem Vals identifies isn’t new: legacy benchmarks often fail to keep pace with fast-moving model improvements, and because many are public, they’re vulnerable to exploitation—companies can train models on known test data. What’s more, these evaluations tend to focus on abstract knowledge rather than specific, domain-intensive tasks.
Vals targets this gap. Its core evaluation approach combines industry-specific tasks—such as those in law, finance, and coding—with unseen test data that isn’t disclosed publicly. This prevents companies from cheating by training directly for the test. Alongside performance, Vals weighs potential negative outcomes to measure how AI systems handle misuse or unintended consequences.
Stretching into ethical and high-stakes terrain
Beyond traditional sectors, Vals is branching into more sensitive areas of AI capability. Its evaluations now encompass domains like cybersecurity, biosecurity, mental health, and global norms—including benchmarking model behavior against Geneva Convention law. Vals is also exploring recursive self-improvement as a metric for model risk. This suggests an ambition to cover not just what models do well, but what they might do poorly if deployed at scale.
The company has grown rapidly: from a team of eight at the start of the year to 25, with plans to expand further. Revenue has jumped to eight times what it was a year ago. To support its evaluations, Vals began offering gently restricted programs for federal agencies. As its reach grows, so too does its influence over investment decisions and public disclosures in the AI space.
Building trust, guiding development
Vals is asking companies to voluntarily submit their AI models for assessment—with the understanding that not every result might shine. Paying for validation is like investing in third-party certification. The comparisons with standardized tests like the SAT are intentional: the goal is clear, rigorous data about where models stand and what they can and cannot do.
Founder Rayan Krishnan—a former Stanford undergrad with experience at Microsoft and Palantir—views these evaluations as central to how AI models will be discussed by investors, regulators, and the public. As more AI companies go public, report to regulators, or enter highly regulated spaces, they’ll need trustworthy, meaningful benchmarks—not just scoring systems built for PR.
This isn’t just about giving CEOs bragging rights. Vals aims to reshape expectations for what benchmarking means in AI. By focusing on unseen tests, domain specificity, and risk alongside performance, it may mark a turning point in what it means to trust model claims. Watch as its methods ripple into regulation and industry norms.