The Vergecast - How to train your data

Episode Date: June 25, 2026

Training data is the raw material of the AI industry. Claude, ChatGPT, Gemini, and the rest are built on top of oceans of stuff. What is that stuff? Books. Blog posts. YouTube videos. Reddit comments.... All of it and more, in virtually incomprehensible quantities. Alex Reisner, a staff writer at The Atlantic who has been investigating training data, explains how AI companies get all this data, why they'd really prefer you not know what's in it, and whether training data could ever be a fair trade. Further reading: Apple raises prices on Macs, iPads, and more by hundreds of dollars | The Verge⁠ ⁠Disney agrees to pay $50 million to YouTube TV and DirecTV subscribers | The Verge⁠ Two handlebars are better than one, right? | The Verge⁠ At Least 15 Million YouTube Videos Have Been Snatched by AI Companies⁠⁠ ⁠⁠The Hypocrisy at the Heart of the AI Industry ⁠⁠ ⁠⁠The Millions of Songs Mashed Into AI-Generated Music⁠⁠ ⁠⁠Common Crawl Is Doing the AI Industry’s Dirty Work⁠⁠ Subscribe to The Verge for unlimited access to theverge.com, subscriber-exclusive newsletters, and our ad-free podcast feed. We love hearing from you! Email your questions and thoughts to vergecast@theverge.com or call us at 866-VERGE11. Learn more about your ad choices. Visit podcastchoices.com/adchoices

Transcript
Discussion (0)
Starting point is 00:00:02 Hello and welcome to the Vergecast, the flagship podcast of music that sounds eerily but not exactly like other music. I'm your friend David Pierce, and today on the show we're talking about training data. It's the raw materials of everything that is AI. And we probably don't understand it or talk about it enough. I'm talking to Alex Reisner, who's a staff writer at the Atlantic, and over the last couple of years, he has spent a lot of time investigating training data and how all of these books and all of these articles and all of these YouTube videos and all of these songs are compiled into these gigantic data sets that AI companies then use to form the basis of their models. I think the way that these models are created and the sources of these data has a lot to do with the way that we feel about AI, particularly generative AI as a creative expression.
Starting point is 00:00:50 And understanding how this data works, where it comes from, and how it gets used, I think is really important. So we're going to have Alex on. We're going to dig into it. I'm very excited about it. But first, here's everything else happening on The Verge today. This is 90 seconds on the verge for Thursday, June 25th, 2026. Apple just raised the prices of a huge number of its most important products. iPads and MacBooks in particular are now up anywhere from $100 to $500,
Starting point is 00:01:15 and even like the Apple TV is way up. It's now $200, roughly the price of 900 Amazon Fire Six. We knew this was coming, but this is still a big moment. Apple is as well-managed a supply change. company as anybody and exists at really high margins. If it can't keep prices down in this era of AI-driven shortages of memory and storage, nobody can. And I suspect this is going to get even worse.
Starting point is 00:01:39 My advice from yesterday still very much holds, which is Prime Day is not over. Go get your deals while you still can. And here's a way to help you pay for all of those more expensive gadgets. Go get some money from Disney. If you have YouTube TV or DirecTV stream, you might be eligible for a cash payout from a new $50 million settlement. The case started in 2022, and the argument was essentially that because Disney owns ESPN and Hulu, which are both powerful streaming services in their own right, Disney was able to drive up the prices of all of its rivals to unnecessary heights.
Starting point is 00:02:10 Live TV is lucrative and competitive and is going to keep being messy in this way. This will not end things. But hey, go get some of that $50 million. And finally, a quick PSA to advertisers everywhere. Yes, I get it. The whole idea of using AI creative and AI targeting to put your A. AI ads in front of AI people is exciting, I guess. But maybe check your ads to make sure that you're not, I don't know,
Starting point is 00:02:33 advertising a bike with two sets of handlebars like REI did this week. That may not go over so well with people. Just a little free advice from me to you. You can read more about all of this at the verge.com. That's 90 seconds on the verge for Thursday, June 25th. Support for the show comes from Service Now. AI is moving fast across the enterprise. But without visibility, it's just chaos.
Starting point is 00:02:54 different tools, different models, different teams using AI in completely different ways. ServiceNow turns that chaos into control. With the AI control tower, you see all your AI across the business in one place. What it's doing, what it's done, and what it's about to do. So you stay in control. To put AI to work for people, visit servicenow.com. Support for this show comes from fetch pet insurance. Do you have a pet?
Starting point is 00:03:25 Every six seconds, a pet owner in the U.S. gets hit with a vet bill of over $1,000. And it's almost always an unwelcome surprise. That's where Fetch Pet Insurance comes in. Fetch is the most complete pet insurance. Get paid back up to 90% of vet bills. You can use any vet in the U.S. and Canada. All vets are in network. Go to fetchpet.com slash save right now for your free quote.
Starting point is 00:03:52 That's fetchpet.com slash save. safe. All right, let's talk training data. Alex Reisner from The Atlantic is here. Hi, Alex. Hey, David. Thank you for having me. Very excited to talk to you. I think we have talked a lot on this show about why people feel the way that they feel about AI and the sort of instinctive reactions to the whole idea of AI. And a theory that I have is that a lot of it is about training data. And I am, I just want to talk about that. You've done a lot of work on this stuff and you've been investigating how AI models get trained for a long time now. But I want to just lay a little bit of groundwork here. And just to ask the very obvious question up front, why does it matter what data
Starting point is 00:04:35 is used to train these products? What difference does it make what's inside of these models? I think it is potentially the most important aspect of a model is what it's trained on. I mean, if you take a model and you train it on, let's say it generates music and you train it on 1950s jazz, that model will be very good at generating music that sounds a lot like 1950s jazz. If you train it on a recent hip-hop, it's going to generate music that sounds like recent hip-hop. You know, these models have names like GPT and Claude, but I think you could make an argument that the right name for a model is actually the description of the day. it was trained on because that is a description of its capabilities. That's what it can output. And so I think that the training data is really, it's really fundamental to the model, maybe more than the architecture to some degree.
Starting point is 00:05:33 Interesting. And I think, I mean, that kind of answers to some extent my next question, which is why is training data such a closely guarded secret from these companies? And it seems to me that there is both a straightforward business answer and maybe be a slightly more nefarious answer as to why these companies would be so secretive. But what is your read on why this is such a closely guarded secret, what data is being used to train these models? Yeah. I mean, the companies have argued that they need to keep this secret because the data that they
Starting point is 00:06:07 have selected to train on is their competitive advantage, right? Like Anthropic has done a better job at selecting data than Google and OpenAI. And if they were to let that come out in a court case, or be public in some way they would lose their competitive advantage. You know, there's another pretty obvious reason, which is that they have gone about acquiring a lot of this data in ways that the people who've created the data, the authors of the books and the creators of the videos and the music would not be happy about. And in a lot of cases, they just don't know that their work is being used. And when they find out they're not happy about it. And I think it's a conversation that the AI companies have tried to just avoid having.
Starting point is 00:06:51 So a big part of what you've been up to recently has been sort of reverse engineering this process to like peel the model apart to figure out what data is inside of it. And it seems to me that you've had to do sort of the reverse of what all of these companies have done, which is go figure out where these mass sources of books or web articles or videos or songs are. So tell me a little bit about your process. How do you go about finding these things that are kind of otherwise closely held secrets? Yeah, my process is actually maybe not that different from the company's process. I am a programmer by Trav. I worked in tech for 20 years. I built websites and apps and did some statistics work.
Starting point is 00:07:38 And so I've been aware of these models for a long time. And a certain point I started hanging out in the forums where AI development. developers were hanging out and talking about the work and, you know, they're talking about what data they're using to train these things. And so that was a useful source. And I've also been reading a lot of their research papers. You know, selecting the data is really challenging. There's a lot of effort that they put into it. They have to come up with some notion of what is high quality data, like what do and don't we want in the model.
Starting point is 00:08:10 And they write papers about that. You know, I think they want to be involved in this conversation. that's happening about training data within the industry. One thing that's been really helpful to me is that there's a open source AI development community, and they believe that the work should be done more out in the open. And I think a lot of them are doing really good and important work, and they have really interesting things to say about AI just philosophically and socially. And so they, you know, they're pretty transparent, places like Allen, AI and Illuther
Starting point is 00:08:41 about what they're using to train. And so that's been helpful. But then even at the companies, you know, AI, the AI world is a little bit like academia in that it is good for your career to publish papers. And the companies don't like that, but they also acknowledge that they have to let the employees publish something. And so the lawyers will go over it and tell them what they can and can't say. And over time, they've clamped down more. And so the companies are revealing less through the research papers.
Starting point is 00:09:10 But yeah, a lot of my research is just from reading a Google paper, you know, for example, for this last article, where they said we trained on, you know, tens of millions of songs. Yeah, I remember not that long ago Apple going through a big
Starting point is 00:09:24 cultural issue with this, with its AI team, because Apple is so secretive and so reluctant to let any of its work be public that all of its researchers are like, well, if you're not going to let us publish,
Starting point is 00:09:33 we don't want to work here. And that push and pull seems like it has, it has kind of morphed in a bunch of different directions over time, but interesting to hear that it is definitely, it is,
Starting point is 00:09:43 everybody is retrenching a little bit as this space gets hotter and hotter. Yeah, it's not 20-21 anymore. The research papers read very differently now than they did a few years ago. That's really interesting. But tell me a little bit about these databases. One thing I think I had not given enough thought to until I started really reading your work is that there's a real business in making and maintaining these databases.
Starting point is 00:10:08 Like one company you wrote about Common Crawl, I think, is like a thoroughly fascinating player in this. And they essentially, as far as I can tell, just crawl the internet and make it available to whoever wants it, which is weird and complicated. And we should probably come back and talk about crime and crawl sometime. But it seems like if you look around a little, these gigantic databases of books and articles and videos exist. Like, are people making these and selling them? Why do these databases exist? Mainly, well, I mean, for AI training, why else would they exist? It's a huge, you know, it's an extremely labor-intensive process to find, you know,
Starting point is 00:10:52 in one sense, you just go and download all of Library Genesis or all of Anas Archive. That's one sort of naive thing you can do, but the companies realize pretty early on that they need to filter the stuff pretty carefully. So the organization you just mentioned, Common Crawl, yeah, they've been crawling the web since the late 2000s, maybe 2009 or something like that. And they just make the whole thing available every month. They've scraped a few more 100 million web pages. And it's available to anyone who wants to do any kind of research with it.
Starting point is 00:11:30 In fact, it's mostly AI researchers who are using it. But all the early large language models were trained on Common Crawl. If you go back and read those early open AI papers, everyone was training on Common Crawl and at first, and the models were terrible because if you train a model on the whole internet, it's just, you know, there's, you get all the, it says all the junk that people say on the internet along with, along with the intelligent things. Yeah. Mostly junk. Yeah. And so it's statistically mostly junk. Statistically, yeah, mostly junk. And I think the early large record models were proof of that. But yeah, I think there, you know, Common Crawl is a nonprofit. So, you know, You know, they would argue it's not a big business. They do get a lot of money from AI companies and AI investors. But yeah, I think this, the topic of training data selection, the challenge of selecting the right data for a model is still really hard.
Starting point is 00:12:30 The AI companies, I would say, are still have a very primitive understanding of what data will make their model better. It's an area of research that I think even they at this stage are not very good at. They do it mainly by trial and error as far as I can tell. Interesting. Again, so the reason you're asking, why do these data sets exist? I think it's people trying to share what they've learned from, you know, curating datasets in different ways and training models with them. Yeah, I mean, part of the reason I ask is, I think, one of the things I have come to believe about the AI
Starting point is 00:13:09 industry is that this shift that went from AI research being fundamentally a research thing. Like if you go, if you go way back, Open AI was basically a research organization, right? And you talk about Common Crawl, and I think it's, it's early users were largely researchers. And it was, these things were academic things for academic purposes. And from what I understand, the kind of rules of the road are different, right? that like if you want to make copies of a bunch of things for academic purposes, these things are generally considered less problematic, right? But if you then do it and become a trillion-dollar company on the back of it,
Starting point is 00:13:48 people are going to rightly feel differently about the way that you went about getting that information. And just the speed with which AI commercialized, all of these companies just moved so fast from, we are essentially an academic thing to, oh, my God, we're making so much money. everyone's filthy rich, that it feels like they just hoped everybody would ignore the ways in which they got this information. And so that's why I'm particularly fascinated by where these data sets even come from, because it does seem like in many cases like you're saying, it's not that there is some gigantic business in being the one to sell the songs to somebody.
Starting point is 00:14:28 It's that stuff is being compiled for other purposes. It's just that now the main purpose, because of the sheer volume of work is AI research. All this stuff has been sort of co-opted from every other purpose to AI. And it just all feels so concentrated now. Does that feel right to you? Does that make sense? Yeah, I think that I agree with most of that. I do think that I'm not sure how much these data sets were really collected for other purposes.
Starting point is 00:14:58 Common Crawl likes to talk about. I think they're one case. They probably have the strongest argument that their data could be used for other purposes. But when you go back, they've been cited by over 10,000 papers. I didn't read all 10,000, but I read a lot of them. And they are mostly AI, right? And it's early, a lot of it is stuff that people wouldn't mind as much as with generative AI, right? Like, Common Crawl, I think without Common Crawl, you know, AI translation tools might not be as good as they are.
Starting point is 00:15:32 I think it was really a huge help because they scraped web pages the same page in multiple languages, and people were able to train translation models based on that. So that was helpful. But the thing that, you know, there is still a, what I would call it, data laundering network where the AI companies are still relying on, they'll do a collaboration with the university. and they'll have the university download millions of images to train a model or download millions of articles to train a model. And the AI company can say, like, well, we didn't do it.
Starting point is 00:16:11 This was like an academic thing. You know, the same goes, Common Crawl is not the only nonprofit that's like doing a lot of this scraping for the AI industry. One of the data sets I reported on in the music, the article about music training data is this organization based in Europe called Lyon. have a data set of 12 million songs from YouTube. So anyway, this is like, is it academic? Like, not really.
Starting point is 00:16:38 Like, this is, you know, technically, yeah, there's universities and nonprofits, but they're all receiving money from the AI industry. I'm Seth Matlins. My new show, Creator Destroy Reimagining Marketing Explores how every decision a company makes, not just the marketing ones, but the HR, IR, pricing, org design, and planning ones. The ones most don't consider marketing at all contribute to either creating value or destroying it. Each week I sit down with CMOs, CEOs, founders, cultural thinkers, the people building, breaking, and reimagining how businesses grow or don't for conversations
Starting point is 00:17:12 about what creates value and what destroys it. It's a business show, it's a marketing show. Creator destroys the show that argues. They've always been the same thing from the Vox Media Podcast Network and the Wisdomist Company. New episodes drop weekly on YouTube and your favorite podcast app. This is kind of a diversion, but I'm so struck in reading all of your work by how often YouTube appears as just it is everybody's favorite source for everything. Like, it's what, you know, it's what Open AI used allegedly to create Whisper. It's what a lot of the music stuff is using. It's what there's a lot of stuff based on videos. Like, what, what is your sense of YouTube's role as an AI training force? because it seems to be everywhere. Yeah, that's accurate. YouTube is an extremely common source.
Starting point is 00:18:09 I think one reason is there are tools for downloading from YouTube that are really, that work really well. They're really easy to use. And it's pretty common for AI developers to just use those tools. And it's just kind of a, it's become a custom. But also, you know, and that's, that includes, stuff on YouTube is just less protected, I think, is one way of saying it. There's music, you know, if you're a musician, your song might be on Spotify, but Spotify's
Starting point is 00:18:35 website has digital rights management protections is really hard to download from Spotify. It's much easier to get the same song from YouTube and so many songs are also on YouTube. So I think it's just ease of downloading. YouTube obviously would say out loud that this is not allowed, right? That it violates systems or service. And yet it seems to have either it can't or it just has it. done anything to stop this, really. Yeah, that's a question that's been in the back of my head for a long time,
Starting point is 00:19:08 and which I've asked YouTube in which they don't really answer. They have said that they consider it a violation of their terms of service to be downloading their videos. But yeah, they haven't, you know, years have passed. And it's just as easy to download from YouTube now as it was a few years ago. Yeah, I downloaded a YouTube video this morning. It is just a thing you can do. Yeah.
Starting point is 00:19:28 It's really true. You mentioned the sort of data laundering stuff, but it also seems to me that more and more the people doing the AI training are just completely unapologetic about it. You quoted Rich Screnta, the CEO of Common Crawl, who literally said to you, like if you don't let AI robots crawl your data,
Starting point is 00:19:54 you essentially don't exist. You're going to miss out on the future of the internet, And I think about Mark Andresen, even a couple of years ago, being like, none of this would work if we couldn't just take the training data that we needed. There is this almost like manifest destiny sense of the AI industry that we can have this data because the thing that we're doing is so important that we must have it no matter what. What is your sense of the trend there? Because it occurs to me that that's happening even as the backlash from people who hate the experience and feeling of AI, in part because of the way that this stuff is trained, just keeps getting worse.
Starting point is 00:20:32 These things are just running away from each other at, like, record speeds. Yeah, I, that's a huge question. I think it, you know, I think it has something to do with the fact that we just have not done a great job in this country with establishing the value of data. And who should be able to have data, right? This is something that privacy advocates have talked about. Jaron Lanier, I think, was one of the earliest people to be talking about this.
Starting point is 00:21:05 I think he wrote in 2007 that you should get paid, like you were being surveilled, basically, and companies are taking, that they're monitoring everything you do, they're generating data from your online activity. That data is extremely valuable to them. It might seem like nothing to you, but you should be getting paid for that. People thought that was crazy back then, and I think a lot of people still think that's crazy now. But, you know, taking people's music to build models that generate songs that compete with them is just the next version of that.
Starting point is 00:21:38 And it's just going to keep going. If we don't acknowledge that this data is incredibly valuable and figure out a way to write laws around that or just have, you know, better business practices or something, or something, this is just going to get worse. This is going to be more exploitation and more, more mining. The next step, I think a lot of people have perceived for a while to be synthetic data, right? That eventually we're going to get AI models that are so good at making new things that then we can use those things to train new models. And eventually they don't need existing recorded music, that synthetic data is the future and that's how we get to everything. And especially now, I mean, you look at Spotify and there are tons of AI generated songs on Spotify that some people are listening to. there are a billion AI generated podcast out there.
Starting point is 00:22:30 Like the content is being made. Is that the next phase of training data? Is synthetic data coming into its own in such a way that we're going to start to see the next generations of these models built on the things made by the last generations of these models? Absolutely not. Really? I don't think that there's any evidence that that actually works. I think when AI companies talk about training on synthetic data, They choose their words very carefully, and they always exaggerate the extent to which is happening.
Starting point is 00:23:03 There's a lot of research out there on a phenomenon called model collapse, which is what happens when you train a model on its own outputs. It very quickly, it doesn't get better. It very quickly degrades. And it's not hard to see why that could be the case, right? Like AI is kind of an averaging machine. It's finding statistical average between different types. types of content and putting that into some new, new kind of more average type of content. And there's just not enough weirdness or interestingness or something like that.
Starting point is 00:23:38 There's some quality in the work that humans do that's not in the work that AI does. And I think that's actually proved by the model collapse phenomenon. So I'm amazed that AI companies are going around talking about synthetic data still. There's so much evidence that it doesn't work. So is there a next untapped place filled with data? I mean, these models keep getting bigger, they keep meeting more data.
Starting point is 00:24:03 They got to go somewhere, right? Is there a next sort of unturned place to go for these companies? I think they just pay people to make it. I think there's already a gigantic industry of writers who are writing for AI, musicians who are making music for AI. You know, after I published this story, I got an email from a company that is doing this. They claim to have paid creators over $10 million just to make things for AI training.
Starting point is 00:24:37 It's very strange, but this is the next, like, AI as your audience is the next frontier for creators. It is a deeply strange thing to think about. But also, like, you know, we're going to put this on YouTube and it's going in there anyway. So who knows? Maybe this is all of it. our destinies no matter what. I certainly hope not, and I don't think so. I think we can, I think we can have a conversation and arrive a more reasonable future for ourselves and the culture. I'm with you on that. All right, well, Alex, thank you so much for being here. Really
Starting point is 00:25:12 appreciate it. Thank you, David. That's great. All right, that's it for the show. Thank you to Alex for being here. And thank you, as always, for watching and listening. If you have thoughts, questions, feedback of any kind. If you have a favorite AI song you want to send me, just not the Puerto Rico song, I already know that one. Send me all your favorite AI songs. No judgment. I just want to hear all of them. You can always hit up the hotline. 866, verge 1-1. You can send us an email, Vergecast at theverge.com. We love hearing from you. And as a reminder, the best thing you can do to support all of this is to subscribe to the verge.org.org.com slash subscribe. It gets you all of our podcasts, add free, including this one. It gets you all of our coverage. Terrence O'Brien on our team has been.
Starting point is 00:25:50 been covering Suno a lot and doing a really terrific job. Lots to come on that. Subscribe to the Verge. I think it's a pretty good website. The Vergecast is Verge production and part of the Vox Media Podcast Network. This show is produced by Josh Kahas, Eric Gomez, Brandon Kiefer, Travis Larchuk, and Aaron Laccio. We'll see you tomorrow. Rock and roll.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.