Benchmarks sell models. That is the problem Vals is trying to solve.
The Bay Area startup, founded in 2024, argues that the tests the industry uses to prove capability were built for an earlier generation of models and have become easy to outwit. “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance,” co-founder Rayan Krishnan told TechCrunch.
Krishnan, 25, previously interned at Palantir and worked for Microsoft and Stanford’s AI lab as an undergraduate. He raised a seed round led by 8VC and Bloomberg Beta, then a $40M Series A led by Andreessen Horowitz last month after a stretch of rapid growth.
His pitch is that evaluation should verify what companies advertise rather than rank abstract intelligence. Vals argues the gap matters more now that models sit inside hiring, medicine and finance, where a strong score can shape procurement and public trust.
The company now occupies two floors of a former brewery building on Folsom Street in San Francisco, sharing the space with other startups. Its bet is that as capability claims grow louder, someone has to check them, and that buyers will pay for the check.