Taylor Lorenz’s Power User - AI Is Destroying Millions of Books (But Not For The Reason You Think)

Episode Date: August 5, 2026

Is AI destroying human literature? SUPPORT MY WORK: Buy a paid subscription to my newsletter at https://www.usermag.co      Support my work on Patreon: http://patreon.com/taylorlorenz     �...�   You might have seen the viral videos of a "book guillotine" slicing the spines off thousands of books in a massive warehouse. Last week, the internet exploded after videos surfaced showing Anthropic cutting the spines off thousands of books to train AI models. Social media quickly filled with claims that AI companies were destroying rare books, erasing human knowledge, and building a future where artificial intelligence replaces libraries. But is that actually what happened?In this episode of Power User, I sit down with tech policy expert Derek Slater, who worked on the original Google Books project, to explain why Anthropic is actually scanning books and destroying physical copies afterward, how decades of copyright law (not just AI) created this bizarre situation, and how we can actually save the books!! We discuss:Why Anthropic buys books by the palletWhether rare books are actually being destroyedHow Google Books handled digitizationWhy copyright law incentivizes destroying booksThe Internet Archive lawsuitsFair use and AI trainingThe future of libraries in the AI eraWhy the viral outrage missed the bigger story

Transcript
Discussion (0)
Starting point is 00:00:00 This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales, using automation, analytics, and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most. Learn more at Accenture.com slash Spotify.
Starting point is 00:00:28 hear that it's your money calling it wants a promotion elevate your savings with the scotia high interest savings account always earn high regular interest rates that grow the more you save and invest conditions apply visit scotia bank.com slash h-i-sa to learn more scotia bank you're richer than you think in copyright statutory damages can be up to $150,000 per work so if you're digitizing millions of works it's not just like a slap on the wrist, right? It's guess wrong, go out of business. Over the past week, you might have seen videos or images like this online of these massive
Starting point is 00:01:13 warehouses full of books and what looks like almost like a book guillotine slicing off the spines before the pages are scanned into this giant AI reader. These books are being scanned, then destroyed as part of Anthropics quest to harvest massive amounts of training data for its AI models. But a lot of people seem really confused about why they have to destroy the books. There have been tons of viral posts claiming things like AI companies want to destroy all human art because they genuinely want everyone to be joyless capitalist slaves in a future where AI occupies the actual fun hobbies. Another post claimed that Anthropic wanted to effectively destructively scan all the books in the world. There's been tons of misinformation claiming that the books that they're scanning and destroying are rare. One big account tweeted,
Starting point is 00:02:02 AI companies are buying rare books training on them than destroying the originals forever. This is straight out of 1984. But as usual, the truth about what's happening is actually a lot more nuanced. Number one, as rare book collectors have pointed out themselves, genuinely rare books are not sold by the pallets. And this digitization process where books are destroyed was also not pioneered by AI companies. The truth of the situation is a lot more complicated. And I would argue the real villain in this scenario is actually our copyright system.
Starting point is 00:02:36 Derek Slater is the co-founder of Proteus Strategies. He's also an iconic copyright and tech lawyer who worked in the original Google Books project. He's a legal expert who knows the ins and outs of what's actually happening. And I'm so excited that he's joining me to break down the anthropic book destroying scandal and talk about what those of us who are book lovers can do to actually save the books. Derek, welcome to Power User. Thanks for having me. So Derek, just to set the scene, I feel like people have seen a lot of these really inflammatory videos online.
Starting point is 00:03:06 Everyone is freaking out. And it's sometimes hard to tell kind of what's happening. We see these warehouses full of books. We see this like guillotine type thing, like slicing them up. What is Anthropic doing with all of these books? So they're acquiring these books so that they can scan them, digitize the text, and then use that text to train their model. I don't know exactly which books they are, but, you know, they're probably not particularly old books. because those old books might be less relevant to their training.
Starting point is 00:03:32 They also might already be in the public domain and they might have been digitized, so they might have other sources. So these are books that are probably in the discard pile, you know, are sort of things that might have been thrown out or discarded otherwise, and they can get relatively cheaply in order to get a large number of text, which then they can use to turn the models. Yeah, I saw that they were buying them by the palette.
Starting point is 00:03:51 So they don't even know like what's in them. Basically, it's like these books were going to go to the junkyard, probably or be destroyed in other ways. and so they're sort of able to buy them in bulk. I saw people saying that actually in Hollywood, you can buy books similarly for like movie sets and stuff, where you could just sort of get lots of books for pennies on the dollar.
Starting point is 00:04:10 That's right. This is a challenge, I think, for people in the book preservation, you know, the space, libraries and so on of what do you do when you run out of space and you can't necessarily afford to buy more archival space? You gotta find somebody to put them in some, that means they're sort of, you have to sort of prioritize and put them into a discard bin and they end up in these sort of places. What Anthropic has said is,
Starting point is 00:04:28 that these are not like rare valuable books. So I think people, and we can kind of get into like, what makes a rare book, you know, like what books are out of print. I think there are a lot of older books that are maybe out of print that end up in these giant swaths of book. Like they're sort of buying these books in mass bulk and they're not necessarily adjudicating which books are in there. They're just sort of scanning it.
Starting point is 00:04:49 But I would say that the suppliers are making those adjudications. And you know, at some point, somebody went through and made a judgment of like, this book is no longer worth keeping. And so I think. you'll see a lot of commentary online of people being like, while they're burning the library of Alexandria, like we're losing all of human knowledge. And I'm just wondering like, I don't know how much truth is there to that
Starting point is 00:05:08 because yeah, sure, maybe nobody needs an old like Windows 3.0, you know, book. But also that those types of books can be very valuable to archivists and sort of historians. This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales, Using automation, analytics, and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most.
Starting point is 00:05:42 Learn more at Accenture.com slash Spotify. Summer heats up in FX's The Shards. Set in 80s, Los Angeles, the Shards follows a group of beautiful, privileged prep school students who are preyed on by The Trollers. a serial killer targeting teenagers across the city. This fever dream of youth, beauty, sex, and mystery is a must see. FX's The Shards. Now streaming on Hulu on Disney Plus. Sign up now at Disneyplus.com.
Starting point is 00:06:18 Hear that? It's your money calling. It wants a promotion. Elevate your savings with the Scotia high-interest savings account. Always earn high regular interest rates that grow the more you save and invest. conditions apply, visit Scotiabank.com slash H-I-S-A to learn more. Scotia Bank, you're richer than you think. The shame of all this is that we could have built the library band Alexandria,
Starting point is 00:06:46 but it has been kept locked up by old laws, bad laws, you know, pre-existing business structures that we haven't been over to overcome yet. I don't have any reason to believe that they're focusing on rare books, let alone focusing on destroying them, right? In many cases, there will be lots of other copies of these books. Or there might be ones in libraries or digitized elsewhere. I can't say that for everything, but I have not seen any evidence that suggests that's what they're doing level on it. As you said, intent to do it.
Starting point is 00:07:10 So Anthropic is not the first tech company to scan massive amounts of books and destroy them. Decades ago, there was this effort called the Google Books Project. And I feel like this was one of the first efforts to really digitize books. And the precedents that were set around it ended up having a really big impact on how companies today can treat different books and interact with, I guess, the publishing industry. What did it aim to do and how did it ultimately play out? So the Google Books project was started, you know, over 20 years ago now. The idea was digitize books and make them searchable like the web was searchable. So Google started digitizing books and you could search through them. It would show, you know, short snippets from book,
Starting point is 00:07:52 not a full page, not paragraphs, et cetera, et cetera, and place it to go, you know, buy the book or get it from your library if you could. They got those books in two ways. One was they worked with publishers and rights holders to acquire them. And in those cases, the publishers or the rights holders would say, actually, you can show full pages from the book in some cases. And so you could sort of like have the experience of flipping through the book in the bookstore. The other way, which addressed the vast majority of books in the Google Books database, was partnerships with libraries. So they went to libraries, University of Michigan, Harvard, you know, Stanford, et cetera, And so at these partnerships where they received books from the libraries, from those university libraries.
Starting point is 00:08:30 They digitized the book for Google's use. And then they provided the physical book and a digital copy back to the library. So the books were not destroyed in that case. So it was different in that sense. Google Books is not a product today that most people have heard about or thought about. And it is core to what Google is doing in AI. They have digitized over 40 million books. No one has anywhere close to that many books.
Starting point is 00:08:53 and the extent books are really important in the capabilities of these models. That's a huge advantage. And that's something that has been really big for Google, even if that initial product didn't work out. So I think what Anthropic was trying to do is figure out, well, how do I compete? Well, I have to go buy books of my own. And they have, you know, there, a lot of the same people who started the Google Books project. And thus, I expect to be quite diligent in terms of how they think about what they're buying and, you know, what they're preserving and whatnot. But they're trying to catch up and sort of get to a point where they can compete.
Starting point is 00:09:20 If you're watching this video and you like my work, please support me on Patreon via the link below or buy a paid subscription to my tech and online culture newsletter at usermag.co. That's usermag.com. I don't have any long-term brand partnerships and a lot of my content is effectively demonetized. I've lost major brand deals for speaking out on certain issues and for challenging power. As you can imagine, advertisers are not exactly eager to work with somebody who covers a lot of the topics that I cover and talks about the things that I talk about. These videos I make are entirely funded by you and I can't continue to make them without your support. So if you get any value out of the videos that I create and you want me to be able to
Starting point is 00:09:58 create more, please support me on Patreon or Substack via the links below. On Patreon, I do bonus episodes, monthly Q&A live streams, and post frequent updates about my work. My Substack newsletter gives you a bi-weekly roundup of everything that I'm seeing and reading and paying attention to online. You can also get my newsletter on Patreon. Once again, the links to everything are below in the description. Every dollar of your support makes such a difference. I think it's important to note that like Google is not destroying them in a way.
Starting point is 00:10:24 And I think that's what's like making people so upset. It is this horrible visual to see a book, you know, and have its spine sliced off. Like there's something sad and tragic about it. And I think that that's what people are so angry about is like, it'd be one thing if they were scanning all these books. I'm sure some people would take issue as they took issue with Google. But it's the like method in which they're destroying them that seems sort of unnecessary. So why is it that Anthropic can't say simply do what Google did with Google books.
Starting point is 00:10:51 So there are two reasons here. One is sort of efficiency costs, the other is legal. There are lots of commercial activities, commercial efforts to scan books for various reasons, including for like disability uses and other accessibility uses. In those cases, taking the scan where you slice it off and you just feed the pages in is way cheaper and more efficient. And if you have a student who is print disabled, you know, is blind,
Starting point is 00:11:13 and needs a chemistry book in a week, you just want to digitize it quickly. And that's not quite the anthropic situation. There are certainly like libraries that will do, might slice off the binding, feed them through, and then they may rebind it at some point, but in many cases I think ultimately, those slips of paper just actually sit in rubber bands in a basement. They're not always rebound.
Starting point is 00:11:31 I'm not saying that to sort of excuse one way there, just to sort of set the facts. The real key then is the law. So Anthropic, like many AI developers, has been sued for copyright infringement, including for their use of books. Again, there were two sources of books that Anthropic had.
Starting point is 00:11:45 One was they downloaded some books, off the internet. There's these different databases, books three, Libgen that contain books that were digitized by somebody else. They don't necessarily know the provenance of them, and they use them. And the judge in the Anthropic case had serious reservations about the way they went about doing that. And I'd say, you know, it depends on how you read it, but at the very least that they had these and they didn't secure these files was bad. So the other way, and this is the way where the judge said, this is okay. They went to bookstores. They bought books and they made a one-for-one replacement. So they had an analog copy, they digitized it.
Starting point is 00:12:19 And at the end, they had one digital copy, not one digital one analog. And that one for one replacement is what the judge said was okay. And given the extraordinary damages in copyright, remember, Anthropics settled the case for $1.5 billion. And there are other plaintiffs are still out there. It is totally logical to say, we're not going to take that risk. Basically, the law, copyright law, has put them in a bind. For them to use books, they are incentivized strongly to do this. Well, because if they don't destroy the book,
Starting point is 00:12:45 the legal liability is so insanely large, especially for so many books that they effectively have to destroy the book. Yes, in copyright, statutory damages can be up to $150,000 per work. So if you're digitizing millions of works, right, it's guess wrong, go out of business. It's not just like a slap on the wrist. So I'm kind of wondering, like, what is the judge's thinking in that? Is it just like, hey, we're making a one for one replacement. If you have a duplicate, then you have this duplicate out there in the market that's going to affect. the market somehow? Like, what is the kind of the framework that he used to make this determination?
Starting point is 00:13:21 Obviously, we can't get inside his mind, but what do you think? I think it's sort of about, well, then you would have two instora one. There'd be a substitution if somebody's still reading that physical copy and so on and so forth. Remember, all of this is happening within the bounds of copyright law. Copyright says, you made a copy. That could be infringement. Okay, fair use. We got to go through this four-factor test. Impact on the market for that book is one of them. Was it a transformative use? Well, if there's still a copy, was it? How transformative was it, et cetera? So given that that's what the judge said they could do,
Starting point is 00:13:49 I totally understand why Anthropic would not want to take the risk of coming up with some other ways of doing it. Right, because any other kind of novel way that they tried to do this and tried to save these books could open them up to such liability. And it's sort of untested waters here, not to be so sympathetic with, you know, this big company. But to me, that makes a lot of sense. And I feel like this is why, I mean, I saw people tweeting, like, If you're mad about what's happening at Anthropic,
Starting point is 00:14:15 you should be just as mad at like Disney and some of these other companies that really fought for the copyright system that we have today. So a couple things to know about copyright. I talked about one already. Damages can be huge. It's not just actual damages, these huge statutory damages. If you're working with millions and millions of works,
Starting point is 00:14:29 that's huge. The other things that copyright applies when you commit a thing into a material object, right? You scribble something on a napkin, on a piece of paper. That is copyrighted as soon as you did it. So everything under the sun is roughly copyrighted. And so then if you think about it, But, you know, books, again, most books that I've ever been created are not in print,
Starting point is 00:14:47 or not in commerce. Maybe you could buy them out as an out of print or an out of commerce bookstore, right? You generally can't find and identify and contact and license from the rights holder because they're gone. I mean, it's true for, you know, lots of books, but I'd say particularly in the 20th century, if we want to be able to preserve and use books in the 20th century, get a license is not an answer. When it comes to copyright, there is no, like, registry or database or phone book of finding
Starting point is 00:15:11 rights holders. A lot of time, authors and publishers may disagree on who own the rights, but even if you know this is the rights holder, finding them, contacting them, is simply not feasible. And this has been, this is not just sort of a back of the envelope. This is something of a copyright office is reviewed time and time again. It came up in the Google Book Search litigation. The authors and publishers themselves firmly agreed with this, that licensing is just not a solution for the vast majority of all books that have ever been published. Yeah, I mean, I feel like there's so many books that are out of print or the publisher doesn't even exist anymore. or it was self-published and that person has passed away.
Starting point is 00:15:44 I mean, it's quite complicated. And also, of course, publishers do these side deals with other imprints. There's the paperback version, the print, you know, there's so much. Well, I mean, think of it this way, particularly like 20th century books that were not born digital. There's a physical cost for each one you print and try and distribute. And if you're not going to recover that cost, you don't print the book. It's kind of that simple. And most books don't make that much money.
Starting point is 00:16:08 There's a lot of people who just sell one or two copies of something or a small number. copies and then that's it. So those ones just go out of commerce and they go away and yes, there's no way to contact that rights holder and figure out how to get a license it if you want to do that. Are there any kind of like novel things that you think Anthropic could do to be compliant and not pay billions of dollars in fines while also preserving these books? Again, nobody wants to see these books destroyed. Of course, Anthropic wants to save money, but is there any way that they could potentially sue or go through some sort of like legal system to ensure that they could protect these books. Here's a reason why that would be really hard, and then here's a direction
Starting point is 00:16:44 one could look at. It's really hard in a sense that there's nobody to just go and ask, right? We can't go knock on the corporate office's door and just like get a letter that says, actually this is okay. You can't go to Congress and easily pass a law. There's no easy way, aside from like litigating a multi, multi, multi, multi million dollar case that you would get that answer. On the other side, and we did talk about Google books, right? And they have their digital copy, they give the physical copy back to the library. Were there different facts here? Yes, or rather, I think there could be different facts that lead to different results. I think another thing to understand here about copyright is these sorts of cases turn around the doctrine called fair use,
Starting point is 00:17:20 which is a case-by-case analysis. So even if I had a case that was kind of on all fours with what we're talking about, I couldn't guarantee you because some little fact might be different. With that in mind, though, I think one of the things that did come up in the Google Books litigation was the security controls that Google had over the copies and that the libraries had over the copies. The courts looked at both, you know, what Google did in terms of who had access to the digital copy. What were the security controls? What was it used for? Same thing the libraries got sued by the Authors Guild. And the court said, huh, the libraries also have good security controls. They don't list just anybody into their data center or their data repositories and do whatever. They're careful about this.
Starting point is 00:17:57 And the courts, you know, to be clear, said, if there were different security controls, I might have a different answer. But on the basis of these facts, this. this way where they gave the fiscal copyback was okay. Thus, I could imagine a court ruling similarly, and it's a risky proposition that I could not guarantee you out of the gate you would win on. And thus, I understand why a company like Anthropic would make that decision, even though I have a lot of sympathy for people saying, you know, these are not just sort of dusty old relics that they're in the waist bin. I love that people are ranting and raving about preservation of knowledge. That's great. The thing we should do then is reform our copyright laws and to support libraries and memory institutions.
Starting point is 00:18:32 That's the way we do that. I think people are right. Companies are never going to be the ones that we should trust to do that historical knowledge preservation. That's a job for libraries, for the Internet Archive, for others. They need your support. So what I would say, my plea to people would be, do your rants, do your rants, then go donate to the Internet Archive.
Starting point is 00:18:49 Go donate to your local library. Get that done. Well, speaking of the Internet Archive, the Internet Archive is also being sued because it has been this repository that AI companies have been able to scrape data from. And so these media companies are. now trying to cut off access to the Internet Archive. I feel like the Internet Archive has been under huge attack just for doing the work of archiving.
Starting point is 00:19:10 There are actually two things that have gone on with them. One is really close to what we were just talking about. The Internet Archive for a long time has been digitizing and preserving books, including both books that are out of copyright and some in copyright. They've made them accessible for people who are reading disabled, print disabled, the blind. Again, under special laws that do allow that,
Starting point is 00:19:28 but not for anybody else. And over the years, in the past, they started this program where they lend a book out for a period of time. You get full access to the book. That's just one to one, and that copy of the book store off the shelves. You can't go look at the physical one. It's preserved, but it's in the basement somewhere, right? So it's still one to one. Preserve the physical copy, but only make the digital copy accessible. And they got sued for that and they lost. So that's like even more evidence, I feel like that a similar situation could happen to Anthropic. Yes. Now, they were making the full text of the book available to a reader,
Starting point is 00:19:59 which is different than training a model. And this, again, it gets very case- by case. I think the archive should have won that case too, but they didn't, and you're right. There's a lot of uncertainty in this space. That is thing one. Thing two is what you're talking about of, yes, news publishers and others have accused the archive of letting people train on their archives. I've not seen any actual evidence that that's the case. That seems to me to be a red herring of what's going on. And I think it's sad that whether it's organizations like the New York Times, where journalists rely on the archive, or just places like Reddit that should know better about open access to knowledge or shutting them down rather than, you know, working with them to come up with a way where that archive can persist and they can still manage their works with respect to AI trading.
Starting point is 00:20:39 Yeah, absolutely. I know I feel like the internet archive is just such an important tool for journalists, for archivists, et cetera. There are these claims online that I kind of want to demystify a little bit because I think people are very angry about all this stuff. And they're saying, well, Anthropic is intentionally destroying the books. The goal is to destroy the books that we rely on. on AI for all knowledge. Have you seen any evidence of the company sort of intentionally destroying books in order to, I guess, like, yeah, force people to not rely on their AI models for this knowledge instead? No, I have seen no evidence of that. It seems also pretty far-fetched, given what Anthropic is actually trying to do. There could actually be lots of good reasons to keep some copies around because you might want to re-scan or do other things. I think this is about efficiency
Starting point is 00:21:26 and the law. And you're right. They're a company, so they're focused on efficiency in ways public institutions aren't, but that's why we need public institutions too. It's a yes and. Copyright is a, if not the big barrier here. And it's important to stress this does not have to be a sort of binary, U.N. I, whatever situation, because let's be clear, out of print books, they're not making any money for the rights holder. They're not getting read. They're not getting used for new transformative purpose. They're just laying fallow. That's just dead weight loss. That is deadweight loss of our shared corpus of human knowledge. That's bad.
Starting point is 00:21:58 I'm glad that these things are now getting used in this way that helps. And sure, we should also be finding ways to preserve knowledge in other ways. This episode is brought to you by Accenture. When your advertising operations fall out of sync, everything else follows. Spotify and Accenture are working together to reinvent the rhythm of ad sales, using automation, analytics, and smarter workflows to simplify campaign delivery and access better data across the business. The result? Less time spent on operations, more time connecting brands with the moments and fandoms that matter most.
Starting point is 00:22:30 Learn more at Accenture.com slash Spotify. The first ever All Electric 2026 Subaru Trail Seeker is the EV for the Trail Obsessed, with up to 444 kilometers on a full charge and DC fast charging from 10 to 80% in about 30 minutes in ideal conditions. Plus, ample ground clearance and symmetrical all. all-wheel drive, make the 2026 trail seeker the most capable EV Subaru has ever built. Test drive it at your local Subaru dealer or visit Subaru.ca. Yeah, I think that's been sort of driving a lot of these misconceptions, this idea of rare books.
Starting point is 00:23:09 When I think when you think of rare books, you think of like somebody was saying, you know, imagine this 500-year-old book that survived the Civil War and all this stuff and now it gets sliced off and destroyed by an AI company. It's like, that's a very like compelling image. but again, as somebody that worked in the warehouse of a used bookstore, like, that's not most books. And I don't know that people are reading these anyway. So like, yeah, I wish that we could archive more.
Starting point is 00:23:30 I think these archivists and historians deserve lots more funding and all of that, but still, we have a glut of books. Yeah. And to be clear, if you're talking about the 500-year-old book, that's in the public domain. That can be digitized and preserved. There's no problem there. We're talking about this slice of knowledge that, yes,
Starting point is 00:23:45 because of the Disney's and others who argued for copyright expansion, have locked up the 20th century. Copyright attaches when you fail, fix it in a tangible medium, right? You scribble on a piece of paper, and it lasts for life plus, you know, decades and the author, right? At this point, that means things from before 1930 or so, they're in the public domain and they, we don't have the same problems. Now, it still might be efficient to do that sort of, quote-unquote, destructive scanning, but like that book also could be digitized and made available in other ways because it's in the public domain. Right, right. I think that's
Starting point is 00:24:12 really important to distinguish. I think a lot of people don't realize that books that are effectively before the 20th century are in the public domain and there are probably scans, they can be, you know, retained, I guess. Well, Derek, thank you so much for joining me to break all of this down. I really appreciate your time. My pleasure, Taylor. Thank you.
Starting point is 00:24:29 All right, that's it for this week's episode of Power User. If you like my work, please, please, buy a paid subscription to my substack at Usermag.co. That's Usermag.com or support me on Patreon. I really don't have any sort of consistent advertising on this channel, so my work is entirely supported by you. I cannot continue to report on these things and try to bring nuance to these tech policy conversations
Starting point is 00:24:50 without your support. I'll be back next week with a brand new episode of Power Use. Ready to take your investing knowledge to pro-level. This is Fidelity Connects, your daily edge in the markets. Get deep insights on real-time market topics that may impact your investment portfolio. Listen to Fidelity Connects on Spotify today and power your next move tomorrow. Sir, see you then.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.