a16z Show
a16z Show

Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan

September 9, 2026

AI Summary

5 min read

In early 2024, when Meta released Llama 4, something strange happened. On all the major public benchmarks—where the questions and grading rubrics were open source—the model showed incredible capabilities. But on Val’s held-out private benchmarks, the model was actually underperforming. That disconnect, between what labs self-report and what independent testing reveals, is the core problem that Val founder and CEO Rayan Krishnan came to solve. In this episode of the a16z Show, Krishnan joins Ben Horowitz and Jennifer Lee to explain why public benchmarks can give a distorted picture of AI progress, how independent evaluation works in practice, and why this is becoming an existential question for enterprises spending heavily on AI.

Why Public Benchmarks Break

Continue reading the full summary in the app — free to try.

Read Full Summary →

Free • No credit card required

What you'll learn

  • 1 (00:00) **Introduction: The Need for Independent AI Evaluation** - Hosts set up the problem: public benchmarks are insufficient and a new methodology is needed to accurately measure model progress.
  • 2 (02:18) **The Founding Insight: Why Public Benchmarks Fail** - Krishnan recounts his background in building benchmarks and the early-2024 realization that third-party evaluation was necessary.
  • 3 (04:26) **Historical Analogies for AI Testing** - Krishnan draws parallels to other trillion-dollar industries that required independent testing groups.
  • 4 (04:56) **The Six-Hour Pre-Release Testing Window** - Krishnan describes the intense process of running evaluations just before a model launch.
  • 5 (06:13) **The "AI-Complete" Problem of Evaluation** - The host asks how to deal with the fact that we haven't agreed on how to evaluate human intelligence, and models are good at hacking benchmarks.
  • 6 (07:10) **Ben Horowitz on the Fuzzy Nature of AI Standards** - Horowitz compares defining AI capability to the MPAA film rating system, which is inherently fuzzy and changes over time.
  • 7 (08:46) **Avoiding the "Enron" Problem: The Decision to Not Sell Training Data** - Krishnan explains a key early decision: never sell training data to labs.

+ Full timestamped outline available in the app

Guests on this episode

Show Notes

a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better?

As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks.

They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.


Resources:

Follow Rayan Krishnan on X: https://x.com/RayanKrishnan

Follow Ben Horowitz on X: https://x.com/bhorowitz

Follow Jennifer Li on X: https://x.com/JenniferHli
 

Stay Updated:

Find a16z on YouTube: YouTube

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Show on Spotify

Listen to the a16z Show on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

a16z Show

More from this podcast

a16z Show →