
Sep 9, 2026 · 39m, Episode 1193
Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
From The a16z Show
— Player
?t=90 · ?t=1:30 · ?t=1h2m3s
— Episode notes
a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a model is getting better? As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even...
— Timestamp deep links
Share any moment by clicking "Copy @ timestamp" in the player above. Supported formats: ?t=90 (seconds), ?t=1:30 (mm:ss), ?t=1h2m3s (hms).