{"id":263945,"date":"2026-09-19T13:00:00","date_gmt":"2026-09-19T13:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/19\/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking-techcrunch\/"},"modified":"2026-09-19T13:00:00","modified_gmt":"2026-09-19T13:00:00","slug":"vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/09\/19\/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking-techcrunch\/","title":{"rendered":"Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">Benchmarking has become the industry norm for how AI companies validate their models\u2019 capabilities and, when the metrics swing in their favor, stand out from competitors and advertise their superiority. In other words, good benchmarks pretty much always mean good PR.<\/p>\n<p class=\"wp-block-paragraph\">Unfortunately, companies have also figured out how to outwit legacy benchmarking systems \u2014 many of which <a rel=\"nofollow noopener\" href=\"https:\/\/www.technologyreview.com\/2026\/03\/31\/1134833\/ai-benchmarks-are-broken-heres-what-we-need-instead\/\" target=\"_blank\">are older<\/a>, and not built to measure the capabilities of modern models. <\/p>\n<p class=\"wp-block-paragraph\">Vals, a startup <a rel=\"nofollow noopener\" href=\"https:\/\/www.bloomberg.com\/news\/newsletters\/2024-04-11\/this-startup-is-trying-to-test-how-well-ai-models-actually-work\" target=\"_blank\">formed in 2024<\/a>, says that it is on a mission to fix this very imperfect system. In the span of less than two years, the company has established itself as a notable presence in the tech industry and, last year it managed to secure a seed round led by 8VC and Bloomberg Beta. Then, last month, after a period of rapid growth, it <a rel=\"nofollow noopener\" href=\"https:\/\/www.citybiz.co\/article\/890247\/vals-ai-raises-40m-to-expand-independent-ai-benchmarking\/\" target=\"_blank\">raised $40 million<\/a> in a series A led by Andreessen Horowitz.<\/p>\n<p class=\"wp-block-paragraph\">Rayan Krishnan, the company\u2019s 25-year-old co-founder, previously interned at Palantir, and, as an undergraduate at Stanford, worked for Microsoft and the school\u2019s much lauded artificial intelligence lab. Krishnan says Vals was born from his own observations about how benchmarking was falling behind the advances of the industry it was designed to measure. <\/p>\n<p class=\"wp-block-paragraph\">\u201cWe were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance,\u201d Krishnan shares. With AI being integrated into every part of society, benchmarks should really exist to verify that models can do what companies advertise they can do, Krishnan said.<\/p>\n<p class=\"wp-block-paragraph\">Last week, the young founder showed me around his company\u2019s two-floor office on San Francisco\u2019s Folsom Street \u2014 an old brick building that, a century ago, <a rel=\"nofollow noopener\" href=\"https:\/\/noehill.com\/sf\/landmarks\/sf199.asp\" target=\"_blank\">served as the site of a large brewery<\/a>. Instead of an industrial output of beer, the historical structure is now home to a number of different startups looking to ship the future of the tech industry.<\/p>\n<p class=\"wp-block-paragraph\">\u201cHistorically, I think evaluation has been done to evaluate intelligence in a very abstract way,\u201d Krishnan tells me. \u201cLike, do models know enough information to be able to take a bar exam type test?\u201d<\/p>\n<p class=\"wp-block-paragraph\">Here, Vals seeks to differentiate itself. While many benchmarking systems offer tests that are publicly available (this can allow a company to train its model against those tests, thus arguably <a rel=\"nofollow noopener\" href=\"https:\/\/www.nytimes.com\/2024\/04\/15\/technology\/ai-models-measurement.html\" target=\"_blank\">cheating on their exam<\/a>), Vals doesn\u2019t publicly disclose its specific test materials. Instead of measuring an AI model\u2019s general knowledge, Vals also evaluates models on their ability to complete complex tasks associated with specific industries like law, finance, and coding. <\/p>\n<p class=\"wp-block-paragraph\">\u201cWhat we\u2019re doing is actually looking at what are the real impacts of the models,\u201d said Krishnan. \u201cCan they do work that produces a product of the same quality as a human within every domain?\u201d<\/p>\n<p class=\"wp-block-paragraph\">The idea is to check not just for positive outcomes but also for negative ones, he says. The hope is to analyze how, \u201cif these models ran wild in the world, what the negative implications would be.\u201d<\/p>\n<p class=\"wp-block-paragraph\">The capabilities that Vals is measuring are growing. In additional to more traditional industries, the startup continues to push into more unique terrain. \u201cWe have a benchmark on recursive self improvement. We\u2019re doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention,\u201d Krishnan shares. <\/p>\n<p class=\"wp-block-paragraph\">Companies pay Vals to test their models, which can be an odd concept to wrap your head around. Why would a company pay to learn its model isn\u2019t performing well? But having an effective measurement helps companies troubleshoot and improve over time. Krishnan compares their revenue model to how a student might pay the College Board to take the SAT. <\/p>\n<p class=\"wp-block-paragraph\">In turn, these evaluations are becoming key decision-making factors for companies looking to acquire new AI models.<\/p>\n<p class=\"wp-block-paragraph\">The startup <a rel=\"nofollow\" href=\"https:\/\/x.com\/ValsAI\/status\/2087917239966290168\" target=\"_blank\">recently revealed<\/a> that its revenue is currently eight times what it was last year. Its staff is also growing. Vals, which started the year with only eight people, has already tripled to a team of 25. Krishnan said that as the startup grows, the plan is to relocate to a significantly bigger office, as well as to bring on an additional 10 to 15 people. The company also recently <a rel=\"nofollow noopener\" href=\"https:\/\/www.youtube.com\/watch?v=iUZ6c1YkFPc\" target=\"_blank\">launched a program<\/a> centered around providing model evaluations to federal agencies.<\/p>\n<p class=\"wp-block-paragraph\">Krishnan sees his company\u2019s system of benchmarking as the future of how AI companies think about growing their businesses and establishing public trust.<\/p>\n<p class=\"wp-block-paragraph\">\u201cAI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings or talk about the prospective investments they\u2019re going to make in AI,\u201d he said. <\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/09\/19\/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Benchmarking has become the industry norm for how AI companies validate their models\u2019 capabilities and, when the metrics swing in their favor, stand out from competitors and advertise their superiority. In other words, good benchmarks pretty much always mean good PR. Unfortunately, companies have also figured out how to outwit legacy benchmarking systems \u2014 many [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":263946,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-263945","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263945","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=263945"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/263945\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/263946"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=263945"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=263945"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=263945"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}