AI Summary
5 min readIn 2026, Apple raised prices on iPads and MacBooks by up to $500, and the Apple TV now costs $200. The reason, as David Pierce explains on The Vergecast, is an "era of AI driven shortages of memory and storage." This is the backdrop for a conversation with Alex Reisner of The Atlantic, who has spent years investigating the raw material that makes AI possible: training data. The episode is built around a simple but often overlooked premise: understanding how AI models are trained—and on what—is essential to understanding the technology itself, and the industry is deeply conflicted about letting anyone see how it works.
Why the data matters more than the model
Reisner argues that the training data is arguably the most important aspect of any AI model. "If you train it on 1950s jazz," he says, "that model will be very good at generating music that sounds a lot like 1950s jazz." The model's name—ChatGPT, Claude—is almost irrelevant. "You could make an argument that the right name for a model is actually the description of the data it was trained on, because that is the description of its capabilities."
Continue reading the full summary in the app — free to try.
Read Full Summary →Free • No credit card required
Never miss an episode of The Vergecast
Get every new episode summarized in your inbox — free, ~5 minutes to read.
No spam. Unsubscribe anytime.
What you'll learn
- 1 (03:50) **Interview begins** - David Pierce introduces Alex Reisner and frames the discussion around training data
- 2 (04:24) **Why training data matters** - Explains that data determines a model's capabilities and outputs
- 3 (05:27) **Why companies guard training data** - Competitive advantage claims versus fear of creator backlash
- 4 (07:09) **Alex's investigative process** - Reverse-engineering datasets by monitoring developer forums and research papers
- 5 (08:36) **Open source transparency** - Groups like EleutherAI and Hugging Face publish data sources more openly
- 6 (10:01) **Common Crawl's role** - Nonprofit that scrapes and releases massive web archives monthly
- 7 (10:34) **Origins of large datasets** - Most are built specifically for AI rather than repurposed from other uses
+ Full timestamped outline available in the app
Guests on this episode
Show Notes
Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments. All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade.
Further reading:
Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed.
We love hearing from you! Email your questions and thoughts to [email protected] or call us at 866-VERGE11.
Learn more about your ad choices. Visit podcastchoices.com/adchoices
More from this podcast
The Vergecast →