AI Summary
5 min readWhy Medical AI Needs a Referee
Hundreds of millions of people ask ChatGPT questions about their health every day, and doctors are increasingly using AI tools like OpenEvidence in clinical practice. But as Protege co-founder and chief scientific officer Engy Ziedan points out, when it comes to ensuring those answers are safe and correct, the answer to "who is checking?" is "everyone and no one." Foundation model builders do internal safety work, and every vertical AI vendor produces a brochure claiming to be best. But there is no independent party looking beyond catastrophic failures to catch the subtler, harder-to-detect forms of misalignment that could quietly harm patients.
The Evaluation Gap: Why Benchmarks Aren't Enough
Ziedan, a healthcare economist whose work has been cited by the CDC, sees a fundamental measurement problem in medical AI. A model can ace thousands of test questions—scoring 92% on licensing exams—and still fail at real clinical tasks, performing at only 45% in actual practice. The gap exists because benchmarks test general knowledge, not the specific, high-stakes scenarios patients actually care about.
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of a16z Show
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (00:00) **The Medical AI Measurement Problem** - Introduces the core tension: models can ace benchmarks but fail in real clinical settings, and no independent referee exists to verify safety and alignment.
- 2 (02:17) **Origin Story: From Pandemic Data Need to AI Data Mission** - Engy describes meeting the Protege team and the founding insight that models would be limited by training data.
- 3 (05:07) **Evolving Customer Needs: From Consumer to Enterprise** - Describes how Protege's customers shifted from startups to foundation models, and now demand enterprise-grade evaluations.
- 4 (06:45) **Why Evals Matter: The Economic Case for a Referee** - Explains why evaluations are critical for pricing, trust, and market equilibrium in AI.
- 5 (09:55) **Catastrophic vs. Subtle Failure: The Harder Problem** - Distinguishes between easy-to-detect catastrophic errors and the harder-to-catch subtle bias and misalignment.
- 6 (12:55) **Who Is Watching the Answers? Everyone and No One** - Answers the question of who ensures AI safety today, highlighting the gap between vendor claims and independent oversight.
- 7 (15:19) **The Benchmark Problem: Sensitivity, Contamination, and Misleading Rankings** - Explains why current benchmarks are unreliable and can produce contradictory results.
+ Full timestamped outline available in the app
Show Notes
Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark doesn't necessarily mean a model is ready for the hospital.
Engy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it's prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up.
The conversation also gets into Protege's role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.
Resources:
Read our insights piece: https://www.a16z.news/p/the-oracle-problem-an-invisible-bottleneck
Follow Engy Ziedan on X: https://x.com/engyziedan
Follow Daisy Wolf on X: https://x.com/daisydwolf
Follow Eva Steinman on X: https://x.com/evajsteinman
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
More from this podcast
a16z Show →