Plain English with Derek Thompson - AI Slop Is Breaking the Internet. Can We Save It?

Episode Date: September 1, 2026

So many people seem to hate AI, and it’s not just because it might displace millions of jobs, destroy the world, or deepen distrust in big corporations and government. The reason might actually be s...impler than we think: People don't like AI because it is so often used to make things that are fake. From sham biographies and bogus social media posts to AI-generated journalism, the internet has become inundated with “AI slop.” So how do we begin to clean up the mess? That’s where AI detection company Pangram comes in. Derek talks with Pangram founder and CEO Max Spero about the seduction of AI writing, how his program detects it, and the moral implications of publicly calling out those who use it. Subscribe to our YouTube channel here:https://www.youtube.com/@PlainEnglishwithDerekThompson If you have questions, observations, or ideas for future episodes, email us at PlainEnglish@Spotify.com. Host: Derek Thompson Guest: Max Spero Producer: Devon Baroldi Additional Production Support: Ben Glicksman Learn more about your ad choices. Visit podcastchoices.com/adchoices

Transcript
Discussion (0)
Starting point is 00:00:05 I've been doing shows recently about why so many Americans say they hate AI. I think it's not just that AI chief executives have promised that the technology will displace tens of millions of jobs. It's not just that AI optimists variably claim that it will destroy the world and or cure cancer, like some kind of cosmic coin being flipped to determine the future of existence. It's not just that its physical manifestation, the data center is a big ugly box. It's not just that a lot of Americans don't trust big corporations or government, and therefore certainly don't trust big corporations working hand and glove with government to build this thing, and it's not just that they think AI videos are corny or annoying or tawdry or artless or all for at once.
Starting point is 00:00:46 And the reason I think that a lot of people just don't like AI is so obvious that it is literally the first word, the A for artificial. People don't like AI because it is so often used to make things that are fake, fake inspirational monologues, on LinkedIn, fake posts on Twitter, fake photos on Instagram, fake videos on TikTok. The Wall Street Journal recently reported that AI has, quote, plunge the book industry into utter chaos. Continuing, the spectacular implosions of big deals over suspected AI use are forcing a reckoning over creativity, trust, and the future of the industry, end quote. And here is the New York Times just today. Quote, the worst of artificial intelligence slop has cluttered Google searches with recipes for glue-topped pizza,
Starting point is 00:01:38 Amazon listings with sham biographies, and Facebook feeds with images of shrimp Jesus. Silicon Valley wants to take out the trash. LinkedIn said in July that AI slop is a top priority for all of us and developed a button for users to report it. Researchers from YouTube's parent company, Google, described how a video service wiped 130 channels of low-quality content off its platform over six months.
Starting point is 00:02:06 End quote. For many people on the internet, the best medicine for AI slop is Pangram. Pangram is an AI detection app. Here's how it works. You read some piece of writing that seems to you like AI Slop. Maybe it's an email you got, or a grant that you're reviewing,
Starting point is 00:02:24 or a memo that you read, a post that you saw. You copy the text. You go to PANgram, you paste it inside a text box, you press a button and voila, it tells you what percent of that writing is AI, from 100 percent human to 100 percent artificial. In my corner of the world, which is to say journalism and academia
Starting point is 00:02:44 and writing in general, Pangram has become nothing less than a cultural force, a meme even. I've seen prominent politicians, chief executives, and public intellectuals, exposed as AI slop slingers. As the author Tim Rekwarth recently reported, the publisher Hachette
Starting point is 00:03:02 recently pulled a horror novel after detectors flagged it as substantially AI generated. The New York Times fired a book critic who seemed to be using AI and the Atlantic reported that a modern love column in the Times seemed to be half AI generated.
Starting point is 00:03:19 Today's guest is Pan Graham founder and chief executive Mark Spiro. We talk about the seduction of AI writing and the effort to clean up the internet from its mountains and mountains of AI slop. I'm Derek Thompson. This is Plain English.
Starting point is 00:04:02 Max, Spiro, thank you for joining us. Hey, thanks for having me. So tell me a little bit about your background. What were you studying in college? What did you do right after college? And how did you come up with the idea for PanGram? Yeah, so I studied computer science in college. I loved programming.
Starting point is 00:04:20 I did a bunch of things around robotics. I really just loved making things happen in the real world. And then I kind of got into machine learning because I was interested in robotics. And that's like kind of like the natural next step is like how do you, instead of like hard coding a trajectory, how do you make an algorithm that can learn and be more natural? And so ultimately what ended up happening is I went to Google, working, training big like machine learning systems of basically for ads, essentially. and I kind of learned how to do ML there. And then I went to Nero, self-driving car company. And I was really just like, chat GPT came out during my tenure at Nero.
Starting point is 00:05:04 And I realized, like, this is like really going to be important technology. And it's going to change the world in a lot of ways. Not all of them necessarily good. And then tell me, like, what you wanted to do with this company. Like, what did you initially see? see as the problem that PANGram was going to solve back in 2022 or 2023? I think early on I was thinking a lot about bad actors and what they could do with AI as essentially this infinite volume of plausible text. I was thinking somebody could go, you know, write thousands,
Starting point is 00:05:45 hundreds of thousands of fake reviews. Somebody could go, like, create, like, totally fabricated news websites. That was sort of like political misinformation, disinformation, could just be supercharged with AI. It's mostly what I was thinking about. So I think it's funny because most people, I think, locate the scale of AI risk at the level of apocalypse. Like, AI will destroy all the jobs.
Starting point is 00:06:09 AI will kill the people. The idea that AI will publish a lot of DREC on Twitter and LinkedIn and doesn't, like, quite rise the level of an existential risk. Why was it so important to you that we be able to detect AI writing? Like, what did you see as the real risk here? Yeah, I mean, I think, like, obviously humanity is not going to go extinct if we can't tell AI writing from human writing. But I think what we are doing is we're trying to save the Internet.
Starting point is 00:06:42 We're trying to save the written word and communication. I think if you look at the scale of AI and bot traffic on the internet, it's gone from a small fraction of human traffic to recently pass 50-50. So like 50% of the internet traffic is bots. And I think it's not that long before it's going to be 99% of internet traffic is bots and not just readers either, but writers, communicators, agents that are trying to get things done on behalf of their owners. So I think we're very much at risk of this idea of dead internet theory
Starting point is 00:07:18 where the internet is just like this echo chamber of bots talking to bots and it's going to be impossible to parse out the signal from the noise when there's so much AI chatter. I like that answer. I would add that I'm worried about the relationship between humans and humans. Like something happens to the way that you can speak to another person on the internet authentically, when you plausibly believe that their output might just be chat GPT or Claude.
Starting point is 00:07:50 I mean, we've historically had this association whereby if it's written, then a human wrote it. If it's written by and signed by Derek Thompson, therefore Derek has in his head an understanding of what he wrote. But this idea that like everyone suddenly has an AI ghostwriter means that like we don't know who we're dealing with when we read. certain posts. And that to me, like, scrambles and weakens the relationship, like, between people on the internet, which further deadens the dead internet theory that you're pointing to. Yeah, no, totally. I think previously, well-written text was plausible. And, like, I see something
Starting point is 00:08:32 that's well-written. Looks like a lot of work effort was put into it. Now that's like, I see something that's well-written. And then I have to question, well, is this AI generated? Did this person actually put any time into writing this in the first place, which I think is kind of just like that sucks in general that I even have to ask that question. And just in the last few days, the Wall Street Journal reported that, quote, AI has plunged the book publishing industry into utter chaos, end quote, just this morning, the New York Times reported that Spotify, LinkedIn, and others are, quote, trying to dig out of a digital sewage heap full of low quality content made by artificial intelligence.
Starting point is 00:09:08 within that article, they cited PANGram, which reported that AI was used to create nearly half of the posts with more than 50 words on X or Twitter. I mean, that is crazy, this idea that if you read a post with more than 50 words on Twitter, you can flip a coin and heads means it was written by AI or odds worse. I'm not sure if that got misquoted somehow, but on Twitter, So our study found that 29% of long-form content, so 250 words or more, 29% on Twitter was AI generated. And then of short-form content, 50 to 250 words, it was about 9%. I say, okay, well, good to have fact-checked in your Times here in a way that is live. How, are there other ways that you can describe, like, how bad the problem is,
Starting point is 00:10:01 the problem of more and more of what we're reading on the internet, essentially being written by bots? Yeah, I mean, I think it's just, it's not just the social platforms. I think they obviously have this high degree, this big problem. LinkedIn had 41% of their long form content was AI generated. But if you look at the internet at large, we did this study with the Stanford and the internet archive, which found that in May of last year, it was already 40% of the internet was AI generated. And I think this number is just going to keep expanding. There's people trying to pollute the internet for different reasons, whether it's they want their own narratives in the LLM training data or they're trying to win at SEO. And now it's easier than ever to produce AI content for search results. So your original thesis was that you wanted to stop, quote, bad actors. Yes. From filling the internet with AI track. Now that you're in the trenches, right, now that we're living in the reality of 2026,
Starting point is 00:11:06 almost four years after the release of chat GPT. Is it your conclusion that most AI content is in fact written by bad actors, so to speak? Or is the problem more complex that, like, we invented an output machine that just spits out humanoid sentences. And so a lot of people with, like, busy days are deciding to just, like, write that email with AI,
Starting point is 00:11:32 write that LinkedIn post with AI, write that memo to the boss with AI, so that in a way that is both less pernicious and more pernicious, you don't really need to be a bad actor to make AI output the way that you write. This is just how busy people are filling paragraphs these days. Yeah, I don't think I fully understood how the economic incentives would work. So it's not just bad actors, but it's sort of just like everyone. It's so cheap to produce the written word.
Starting point is 00:12:07 And then, like, why is LinkedIn more full of AI content than, say, like, Twitter? I think it's because there's this incentive that people feel. It's like, oh, posting on LinkedIn regularly is going to improve my career aspects, my career prospects in some way. It's going to make me more likely to get a good job because I'm going to be visible to people who are relevant in my career. And so obviously, if posting more has an advantage to your career, then people are going to find the easiest way to post more, which ends up being using AI. I do think that's pretty much the explanation. Like, I think that sometimes it is used explicitly to lie. But I also think a lot of people are just like, I don't know, like, when I post something, number goes up.
Starting point is 00:12:55 So I'm just going to post more and that posting is more likely or easier, cheaper to come from chat GPT than it is from my own fingers. And today's algorithms, they also just really incentivize volume and quantity over quality. And I think this is something that we're hopefully going to see change in the coming years as quantity is free, but quality is really the limiting factor. So, PANGram, as I explain in the open, you take text, you put into PANGram, the program comes back with a rating between 100% AI and 100% human. without spilling state secrets here, what can you tell us about how Pangram works? So Pangram is a classifier model. That means what it's doing is it's looking at text
Starting point is 00:13:40 and saying, I believe this is AI generated, I believe it's AI-assisted, or I believe it's human-written. So it's not doing any sort of generation. It's not an L-LM or it's not like chat GPT. But what it's doing is we have, have a training set of human writing. So writing that we know was written by a human. And for example, like a five-star review, Yelp review about Denny's, or on the other side, like a 500-word essay on
Starting point is 00:14:12 Moby Dick. And then in each case, we're picking an AI model and we're asking it to write something similar. So we're going to ask ChatGBT, write a five-star review about Denny's. I'm going to ask Claude, write a 500 word essay on Moby Dick. And then what Pangram is doing is then comparing the human and the AI items, which are similar in topic, but different in style. And we're learning the difference
Starting point is 00:14:40 in how AI produces text versus how humans do. If you look at the broad field of AI systems, there's models that will discriminate. For example, like Waymo is, it has a bunch of cameras. it has this pedestrian detection system where it takes in an image input and then it's going to say
Starting point is 00:15:01 either like nope there's no pedestrians or yes there's a pedestrian here they are and so that's probably closer to what pangram is than like chat gbt which is more like producing text like pangram isn't producing anything it's really just like making a judgment based on an input it's the same way that
Starting point is 00:15:21 when we were first learning about the text the chat GPT and Claude were slurping up. It was clear that they were getting a lot of old books. They were getting a lot of Reddit. And the output was essentially trained toward old books and Reddit. So that if, for example, you wanted Chat Chapti to write a sonnet like Shakespeare, there were so much Shakespearean sonnet content that they had slurped up. That was very easy for them to output that.
Starting point is 00:15:51 Whereas some other questions were a little bit harder for you. Chachb-T because there wasn't enough of that or as much of it in their training set. Can I ask a similar question about Pancram's training set? Is there a certain kind of writing that you guys have trained Pangram on in abundance versus other kinds of writing that is harder to get or less frequent in your training set? I would say there's a little bit, this concept called topic drift. So our entire human data set is from pre-2020. It's from before chat GPT was released.
Starting point is 00:16:32 And so from 2022 to 26, I think language has changed a little bit, but it's mostly the same. But one of the big things that has actually changed is how people talk about AI. Like there's very little writing about AI and LLMs, and there's a lot of terms today that didn't exist four years ago. So I think a lot of what we're actually, that's like a big question for us is like, how do we confirm that pangram still does well on writing about AI? Because we don't have this in our test set to prove that we have low false positive right here. So something we have done is just take writing about like different algorithms or something like that. And then replacing the world the word algorithm with AI or LLM to try and like simulate.
Starting point is 00:17:24 modern text without actually having access to confirmed human written modern text. And why should people trust Pangram? What is the best third-party empirical evidence that it actually works without throwing out a bunch of false positives or false negatives? So there's a really good study by the University of Chicago. And what they found was that they tested a whole bunch of AI detectors, some open source ones, some commercial ones, and they found that pangrum was the best by a large margin. They tested almost 2,000 human texts and found zero false positives. And they found that pangram was just
Starting point is 00:18:06 like highly accurate across all of these different domains. And so I think that was a really good one. And honestly, part of the reason that we started building some credibility. Before that, we were just trying to ask people to trust us. And it wasn't necessarily working that well, because there's so many AI detectors out there that say, trust us, and then ends up actually sucking it like says the Declaration of Independence is AI or something like that.
Starting point is 00:18:31 But I think Pangram is different because we've put so many resources behind training this really great model and being able to tell you at a high granularity, what degree of AI is there in this text? I'm going to talk a little bit about AI writing and what makes AI writing AI-ish.
Starting point is 00:18:51 And I feel like you have the perfect, like, God's eye view to help me think through this question that obsesses me as a writer. You guys, you know, I talked a little bit about this last week. What would you say are the hallmarks of AI writing? And how have those hallmarks changed, if at all, in the last four years? Yeah, so I think, like, the early chat GPT and AI models, they were just, like, fairly biased. they were trained on this like instruction tuned data set and then they would just like
Starting point is 00:19:24 way overuse some words and phrases so example like delve tapestry like intricate they like really loved these individual words and so like if you see the word delve kind of in an out of place setting that was an immediate red flag to some people
Starting point is 00:19:41 like oh this is probably just AI content and then maybe like a year and a half later they were able to hammer out these like word level um inconsistencies but there were still patterns that were showing up for example like the it's not just x but y pattern well like it's it's called a negative parallelism and and lMs love it because it's like kind of like impactful it has a lot of like impact on the reader or like these like sentence constructions that use m dashes
Starting point is 00:20:16 because these were also associated with good writing. And now I think they've even hammered some of these out. And so kind of like how I tell, it's, I'm just relying on like longer context signals. Like it's at the like paragraph or sentence level. I could like look at the shape of the text and tell you like, oh, that looks like it's chat GPT or Claude. I want you to say more about that
Starting point is 00:20:47 because I love this theory that the scale of AI's tell has changed over time that AI's tell used to exist at the scale of the word, delve, or the punctuation, the M-Dash. And then AI got a little bit better at writing and the scale was raised to the level of the sentence
Starting point is 00:21:08 and it fell in love with some sentences that were then identified as being AI tells. Like it's not X but Y, of parallelism, which, embarrassingly, I used to use a lot, but now I can't use anymore because it's such an A-I-Tel. It's good writing, but yeah.
Starting point is 00:21:22 It's good, it's good in moderation. It's good explanatory writing at a sentence level. And this gets at something that I'm gesturing towards right now. And now you're saying that tell is more at the level of the sentence. Something I've noticed is that long pieces of AI writing often try to summarize and re-sumorize and re-sumorize and re-sumorize.
Starting point is 00:21:54 And so the tell is at the scale of the sentence. It's like, or at the scale of the paragraph, I should say. Like, they'll be like, this is genuinely important. This is the main course. The bottom line is, it's like every sentence. The honest answer. It's trying to outdo the previous one in terms of summarizing what it's trying to say better. What is that?
Starting point is 00:22:13 How would you describe what's going on there? Yeah, yeah. I think the way I put it the other day was like AI wants to try and make every sentence. It says the most impressive sentence. Like, it's just really trying to impress the reader. And I think partly this is just due to the way these things are tuned. So, like, we typically AI models are trained in two stages. So the first stage is the pre-training, and this is when it's predicting the next token.
Starting point is 00:22:47 It's looking at internet data and saying, I think this word comes next. And then the second stage is reinforcement learning, where now it's already good at predicting the next token. And instead, we're changing the, giving it a new reward function, which is like, give it a reward if it's judged as producing good writing or the correct answer. And so this reinforcement learning can help basically like it turns up a positive signal to 11. So like obviously like I think people like when they read an answer and they feel that it was like a little bit impressive. They're like, oh, this this is meaningful and then it summarized it.
Starting point is 00:23:32 And then you take that reward signal and then you ratchet it up and then eventually now what it's doing is it's summarizing itself. over and over because that's how it knows it can make its answer the most legible to the human. And it's like making these sentences really like impressive and trying to make them meaningful because it knows it'll get rated as a better writer if it does this. Yeah, I don't know. Does that make sense? No, I think I think you're describing it perfectly. That's exactly what it does. It's like, it's like, it's like a tall poppy syndrome thing going on with AI writing. We're like every sentence is trying to stand out taller than the one that came before it. And it's, it's,
Starting point is 00:24:10 drives me crazy as a lover of language because that's not how good writing works at any appropriate length. Like, I remember, I just thought of this as you were giving that answer. This is not in my notes. When I was writing my first long features for the Atlantic, my first cover stories for the magazine, I was working with an editor named Dom Peck. And my first drafts, he said, he said, you're signposting two months. early in the essay, trying to bang people over the head with,
Starting point is 00:24:47 this is why this essay is important, this is why this anecdote's important, these are the six things I'm going to explain to you. He said, that might make sense at the scale of like, whatever, a thousand words. But in a 7,000-word essay, readers want to go on a little bit of a journey. They want to get a little bit lost. They want you to tell them a story and leave, like, the candy. at the end of that story and then there's a grand reveal about
Starting point is 00:25:15 oh no, this is what the story means it's not what you thought it meant. And this idea of writing going from being signpost, signpost, signpost, this is what I'm saying, this is what I'm saying, to this is a story. You're walking down a path and you don't know where that bend in the road
Starting point is 00:25:30 is headed toward. That's long-form writing. And that, to me, is what AI is terrible at. Because AI, in part because of the reinforcement learning with human feedback, that you were describing, because it's so often being used to summarize or to write short posts, it might be overtrained to or overtuned to whatever the word is. Excellent synopsies, great summaries where every sentence is just like, I can sum up the best. No, I can sum up the best.
Starting point is 00:26:00 But there's no sense of that talented novelist or even talented creative nonfiction writer's sense of, I'm going to take you on a story here. And you, you're going to take you on a story here. and you'll only understand its imports by the time we get to the end. That seems totally lost on the models for now, is what I hear you generally saying. Yeah, yeah, 100%. Yeah, I think it's just like, it's simply not trained for this sort of task. And maybe even if it is, it doesn't have enough training data, right? Like the number of 7,000-word articles is very small.
Starting point is 00:26:34 The number of novels is like also relatively vanishingly small. compared to like internet posts. And so I think the, yeah, and then even on the scale of novels, it's like, what makes this novel good? Well, it's so long that, like, I'm not entirely sure that LLMs are actually able to, like, pick up the wise. What makes it novel grade is change. The characters change. The stories change. The emotions change.
Starting point is 00:27:05 And so the most brilliant, the novelist that I like the best, are masters at being able to create scenes and characters in the first 10 pages that have undergone, in most cases, something quite wrenching by the end of the story that moves you. But that requires, I think, a confident ability to manage meaning across tens of thousands of words that a summary-making machine is not going to be as good at.
Starting point is 00:27:36 Now, having said all of that, I'm going to complicate this whole thesis by referring to lots of studies, which I'm sure you've seen, that suggests that there are readers and even MFA writers who cannot tell the difference between human writing and AI-assisted writing, especially when the AI is told quite explicitly, you know, write like William Faulkner, right like David Foster Wallace, it becomes considerably better at training toward that individual style. What do you make of that? What do you make of the fact that even quite sophisticated readers today, according to some seemingly quite good studies, truly cannot tell the difference between AI and human writing? Look, I think the studies are good, but I think they're still looking at these like localized small short passages of text.
Starting point is 00:28:28 So we're comparing at most a couple hundred words of AI text versus a couple hundred words of famous author text. And so I think, like, at this scale, AI is actually probably maybe a little bit superhuman at being able to portray something clearly and concisely in a way that is, like, impressive to the reader. But I think the problem is the, like, long context writing. Like, it can't do this over the course of a novel. It still, like, has, how do I describe it? I totally, I'll try. of you out there. I mean, like, what you're saying is it can write a Cormac McCarthy sentence, but it can't write blood Meridian. Exactly. And the ability to write a Cormac McCarthy sentence
Starting point is 00:29:14 that is spooky and has a lot of ands and some esoteric words and has some haunting biblical quality to it, it can write plausibly Cormick-ish sentences. But the ability to hold those sentences in a mind and create through them the, ha-ha, tapestry of Blood Meridian, that's really, really different than the ability to mimic McCarthy at the sentence level. Like that seems to me to be the distinction that you're drawing out here. Exactly. And again, this is something that I don't want to say
Starting point is 00:29:45 AI is never going to be able to do this because historically, if you say that, you're going to be proven wrong eventually. But I think today's AIs are really just like not even close to there. And why is, just to tie bone in this particular section, can you explain how PANGram is able even at this short passage level to detect A.I. Ness
Starting point is 00:30:13 that these expert readers and writers themselves cannot. Yeah, I mean, I think there's a lot of information hidden in natural language, and I think Pangram is kind of able to pick up on this. So, like, if you think about a way that you might phrase a sentence, you might have kind of this wide distribution of, like, I could phrase it in five different ways. and AI might have a slightly different distribution
Starting point is 00:30:41 where it's going to prefer like ways two and three out of the five. And so I think over time, over the course of a document, we can build up increasingly high confidence that something was written by AI because of the ways that the decisions are made consistent with how an AI makes these decisions.
Starting point is 00:31:00 I know that's a little bit vague. No, what it makes you think of is like AI, it makes me think that like Pangram is able to see like a mathematical structure and language that I can't in a weird way. I think it's true, yeah. I mean, it's trained on so many millions of AI documents that, like, you and I would not be able to read all of the Pangram training data in our lifetimes. And so I think that's sort of part of it is like you can pick up much greater patterns
Starting point is 00:31:29 the more you study and the more you see examples. And with machine learning, we're able to do this at a, you. a great scale that is able to make a panggram, which is essentially superhuman at AI detection. It's better than most people. And that's just because it's seen more data. Did you know Uber has a range of safety features for riders? Like the Share My Trip feature that lets you send your live location to the people who matter most, your spouse, your kids, your best friend, so they can track your ride and make sure you get where you're going. But the safety doesn't stop there. Uber requires every driver to pass a third.
Starting point is 00:32:07 background check before they can start driving. This consists of a multi-step screening process that checks for impaired driving or criminal offenses followed by annual background checks each and every year moving forward. Share My Trip and annual driver screenings are just a few of Uber's many safety features that put safety at every turn. Learn more at uber.com slash safety. Annual driving history reruns do not apply in New York City. Let's talk about pangram as a cultural force, as a meme, as it's effectively become.
Starting point is 00:32:40 Like, you know, I'm, I spend more time on Twitter certainly than Instagram or LinkedIn, but there are certain stories where prominent pieces of writing will be discovered or identified as AI, and people will take that pangram screenshot of 100% AI, and it'll just suddenly be everywhere. How do you feel about the pangram rating becoming a kind of, internet meme? I mean, I think it's emblematic of this greater backlash against AI content. I think for a long time, either like people thought other people couldn't tell, you know,
Starting point is 00:33:17 people are automating their LinkedIn or whatever. And basically it kind of just felt to a lot of people like they're just having AI shoved in their faces like everywhere they turn. And so I think this is part of why the backlash has been so great. people feel so strongly. And I think Pangram is really powerful for a lot of these people as a way to say, it's not just my intuition. It's like, there's some third party accurate thing that is also helping me give judgment and say that this is AI generated. Yeah, everyone wants a referee. I mean, you know, you want some outside referee to tell you what's right. So your model
Starting point is 00:33:57 is very effective. It's not perfect. And some folks really do get publicly shamed by the use of pangram, possibly for reasons that are good, possibly for reasons that are bad. But even if pangram is, let's say, 99% accurate, and it might be an order of magnitude more accurate, but let's say 99% accurate, if it's used 100,000 times, then one should expect that there's going to be 1,000 errors. And a thousand errors is a lot. Does it bother you that the technology might be used to publicly shame people who don't deserve it? Because in fact, they did write what is being labeled AI writing. So two things.
Starting point is 00:34:41 So first, our error rate, our false positive rate, which is how often we say that something human written is actually AI, is about one in 10,000. So that means of every 10,000 pieces of human writing we scan, one of them is going to be AI. So I said 99% in my question, but it's actually 99.99%. Correct. Yeah. With that said, we scan enough text that we're going to have dozens of false positives every day. And that's sort of just a fact of life. I think there's a couple reasons that I consider this fine. Obviously, we want to still drive this down. But I think the big thing is that Pangram has already has a lower false positive rate than your average person. what we've kind of seen in the art world is people will kind of go witch hunt and be like, hey, this digital art looks like
Starting point is 00:35:40 AI and they might go really deep into like the hair is wrong or like, you know, the fingers messed up. And it's actually just like the artist made a mistake. It's not actually AI generated. And so I feel like that does really suck. And it does kind of happen for text too. But I think the larger thing that we're seeing like on a broader scale is that like
Starting point is 00:36:02 people, people's reputation exists in, like, everybody has a corpus of work and text output that they produce. It's not just a single story. And so everyone sort of has this reputation of like, oh, well, I've put out like a hundred pieces of human content before. Maybe this one is a false positive if it says that a few sentences are AI. Whereas if somebody's last three substack posts for AI, and then the next one's also AI, then I think it's like increasingly unlikely that this is a false positive. I want to describe something that went down on Twitter a few months ago. So someone identified a column by a Guardian sports journalist that they said sounded like artificial intelligence.
Starting point is 00:36:47 And the Guardian released a statement that said, no, this is how the writer has always sounded. And then you took about 900 of that journalist's articles and you ran them through PanGram and you published a time series showing that their writing had become quite, quote, increasingly reliant on AI, end quote. When you look back on that, on that post that you made and that time series that you published, what part of you thinks like
Starting point is 00:37:13 that's good noble behavior that we should do, that we should encourage, and what part of you worries that if everyone behaves like this forever, it's going to feel a little bit witch-hunting on the internet going forward? Look, I mean, I think if the Guardian is going out there and telling people, hey, this detector is completely wrong, it's making a bunch of errors.
Starting point is 00:37:41 I think I can, like, reasonably go defend myself and say, like, it's not, like, this is not how the writer has always sounded, because there's actually a point in time where they started using AI, and you can see it. It's, like, very visible. This writer's still around. they're still I don't want to name names but they're still writing with AI you can still look at their recent pieces
Starting point is 00:38:11 and you know like if if the Guardian is happy with this writer's output then like I'm not going to say like that's fine with me but I really do think that this one is AI generated like this
Starting point is 00:38:26 it's not But it's not an isolated incident, and I think it's very clear after looking at the data that it's not a false positive either. I'm of two minds. On the one hand, I'm like dispositionally against public shaming, but I also think your behavior was a bit akin to policing. And policing is good, I think, when the rules being violated are worth policing. And I am pretty concerned about AI writing for probably more reasons than I can even remember to name in the next 30 seconds. I mentioned the fact that I think it reduces the authenticity with which people deal with
Starting point is 00:39:06 each other on the internet because suddenly their output can plausibly be that of a bot that is communicating information that they don't even know in their own heads. I worry about the brains of young people who never learned to write because they simply learn how to prompt Jack Chbett to write all of their essays for them. And they have this logoreic genie that they can just, you know, throw wishes at. So I think there's like values worth protecting. And values worth protecting requires some kind of police presence, right? What values do you think you're protecting?
Starting point is 00:39:44 Like what are the values that are important to you that are worth publishing a time series like this? I think we're really, I value human authenticity and human-to-human connection. And I think in a sense, AI threatens to replace a lot of this, especially when combined with capitalism and monetary incentives. So if you consider a world where it is taboo to ever call someone out for AI writing, what do you think major news and media organizations are going to do? They're going to start to let journalists go and push everyone else to output 10 times as many articles with AI because it's kind of taboo to call people out on this. And I think that monetary
Starting point is 00:40:28 incentive, which is here and is here to stay, means that we need an incentive on the other side, which is people who are vocally anti-AI taxed and pro-humanity. I want to talk about three scenarios that could threaten your company in several different ways. The first scenario is that the gap between AI and human writing closes. And the obvious way that this would happen is that AI gets better and better and better at writing, quote-unquote, like a person. But the less obvious, and maybe just as plausible way this could happen, is that humans get worse at writing like humans. We write more and more like AI. And you sort of have this pincere movement where like the convergence of AI and human writing ultimately closes this gap between what is identifiably AI writing.
Starting point is 00:41:20 which is identifiably human writing. Do you put any stock into that scenario that human writing itself could like AI-Fi itself? I mean, I think the human language does shift. I don't know if you've ever felt yourself saying something that feels like an AI-ism. I just said AI-Fi, which is not even a word.
Starting point is 00:41:42 So maybe even that is an AI-ism. I don't even know. Yeah, I don't know. But, I mean, obviously the human language is changing, but I think it changes at such a like small, long scale. The scale that like language is changing for humans happens on a much longer time horizon than AI text changing. So I think we really should only probably only be looking at the AI text because that's what's changing really, really rapidly. I think there is a good argument that AI writing will be as good as human writing, possibly even surpass human rights.
Starting point is 00:42:19 writing within the next decade. And so the question will then become, well, do we care if something is human authored or not and why? And my argument is we still are going to care a ton. I think we're still going to care whether, like, the internet's just going to be full of bots. The internet's going to be, internet traffic is going to be 99% bots and 1% humans. I'm going to care that I can get information that I couldn't get from clot or chat GPT. To do this, I need to be hearing from a real human, not a bot. I think in general, we are in very ordinary ways, impressed by activities that we know robots can do better than humans, but we're only impressed when the human can do it. I was just thinking about this in a really random scenario, which is pitching in baseball.
Starting point is 00:43:14 Do we think that a robot can throw a ball faster than 105 miles per hour? Well, the robot can probably launch a rocket faster than like 100,000 miles an hour. So it could certainly throw a ball faster than 105 miles an hour. But it's not interesting. No one wants to watch a robot pitch for the Milwaukee Brewers or the Pittsburgh Pirates. They want to watch Paul Skeens or Mizierowski. They want to watch the human throw the fastest ever fastball thrown. by a human. That's the only thing that has athletic or artistic value. And so it's strange to think
Starting point is 00:43:52 about that invading the purely artistic space of painting and writing, we're already maybe seeing it in the world of mathematics where I think it's just in short order, I think people are going to recognize that AI is a better mathematician than a human mathematician living. But I think there's still going to be a part of us that will be specifically and uniquely impressed when achievements are human achievements rather than just technological achievements. We recognize that difference in so many other contexts. We connect with the story behind it too. I think that really I look at the art world a lot where the camera came out and like photography was invented and this was a new technology. And previously, the best thing you could have to faithful reproduction was
Starting point is 00:44:43 like a painting. And then suddenly a photograph can do better along basically every axis. And so what happened to art? Like painting was not completely supplanted by photography. Instead, artists found different ways to find meaning and express themselves, which was not just simply going farther down the path of realism. But instead, they're expressing their humanity in a different way. Yeah. Thank God for that. All my favorite painters are between roughly 18, 1870 and 1925, so all after the invention of the photograph in the camera. The second threat that I can imagine coming to your company isn't just that the gap between human and AI writing closes, but also that people just stop caring.
Starting point is 00:45:29 Are you worried about the end of people caring? I think at least in the medium term, no. I think there's just, again, like, you know, I've been talking about. this like backlash against AI but I think like beyond that there's just a lot of like societal reasons why we're going to care about human text like like I I still find it meaningful that like like you know if if somebody writes me a personal note that it was written by them and it's not just like an AI slop like AI personalized note and I think like the more AI proliferates and becomes more common.
Starting point is 00:46:13 Like, I don't think that sort of thing is going away. Like, like, there's, there are just reasons to prefer human writing that are not, again, like, not quality. And we're going to prefer human text, like, more and more, the more AI text there is. The last touch of the company is more specifically technical. There's this concept in artificial intelligence called The BitterPill, which says that these general models with a lot of compute can often do just about any task better than a narrow model. And that might speak to a world in which Chachiti or Anthropic build the best possible
Starting point is 00:46:56 pangram, or even that the biggest platform companies like Apple or Google, just create something on platform that automates the pangram function such that, you know, if you get a text, an iMessage, from someone on Apple. There's something inside the Apple software that just pings an AI little flag around it. What are you guys doing to prepare yourself for the possibility or the reality that now all the big guys are recognizing
Starting point is 00:47:26 that a lot of consumers are sick to death of the internet being drowned in AI slop and they're coming up with their own medicine. How are you guys preparing for that future of competition? Yeah, so I think the bitter lesson, like the biggest takeaway that an AI researchers should take from this is that more compute and more data will solve a problem more efficiently than any sort of like human expertise here. So like it doesn't matter how much of an expert I am in AI slop.
Starting point is 00:47:59 If I just have more data than anyone else, I can make a better AI detection model. And so I think we're trying to be on the right side of history with this lesson by being, having the biggest model, having the most data. And so far, I think, like, we kind of are that, except for the big labs. Like, Anthropic and Open AI obviously have infinitely more data than us in terms of AI-generated content. But I think there's still a lot of value as well in Pangram being this, like, external third party.
Starting point is 00:48:35 Like, like, I think people will trust us because I'm, I'm not anthropic telling you this wasn't AI generated or this was. I am, you know, separate and independent from OpenAI, Anthropic, Google. And because of that, I can make the best judgment for myself rather than like over or under attributing the AI models. Interesting. Right. It's like the referees in the NFL, the NBA, not belonging to the players union. It's like they're a third thing.
Starting point is 00:49:04 They're not the owners and not the players. They're a third entity and therefore can maintain some kind of authoritative objective. independent. Exactly. So Anthropic recently announced that they're coming out with watermarking, so they're going to be able to tell you if AI touched a piece of text, because the watermark, you know, when it produces its output, the watermark will essentially apply its statistical invisible pattern so that somebody can then read that text, put it through the watermark checker, and say, oh, this text was generated or came out of Claude. However, I think this has the risk of over-attributing authorship to Claude,
Starting point is 00:49:46 especially when somebody, when there's like truly a mixed authorship document where somebody spent a lot of time working on it and then worked with Claude, whereas I think Pangram will be able to do a better job correctly attributing that Claude wrote these parts and a human wrote these parts. This is exactly what's giving my last question to you because I obviously read the news about Claude's watermarking,
Starting point is 00:50:10 And, you know, I've said this in other context. I'm an independent writer. I don't have a copy editing team the same way that I had at The Atlantic to check for my voluminous little errors of, you know, repeated words and lost commas, dropped words, et cetera. And so I have, since I started my substack, taken my essays and put them into Chachbhbt or Claude, and said, said, edit, in line, in bold, all grammar, typos, and errors of clarity or something like that. And then it'll give back essentially my column, but in bold. I've never copy-pasted it. I always, I have it in bold so I can put two windows.
Starting point is 00:50:55 I'm just going to get in more detail than a listener's asked for. But I can create two screens on my computer, one that is substack and the other that is, let's say, Claude, and then just scroll through them and make the edits that are emboldened. But this suggests that, like, right, if I just like, if I copy past, that clawed piece into, you know, any CMS, it would have an AI watermark that would falsely identify what is 99.99% my writing as AI writing, which is something that I wouldn't want. So that if I am, in fact, describing this technology accurately, that's not something that I'd be interested.
Starting point is 00:51:35 Or that I can imagine a lot of people being scared off of that technology, I should say. Yeah, definitely. I think it kind of depends on how visible the watermark will be. There's a lot of questions we have. Here, it could be that if I ask Claude to really only correct my grammatical issues, then there actually won't be enough room for Claude to play with the text in a way that activates its watermark. So it kind of remains to be seen here. But I think it's actually great that you went into detail on this because I think there's a lot of people who, just like paste their text in chat chPD and say make it better and then they don't realize that chat chPD actually just rewrote 80% of their text and they're like hey most ideas are still me it just made it better and then they and then they're like surprised when pangram says this is like 80 plus percent AI they're like no but the ideas were myself but whereas i think if you're using it more for like in bold you change cloud tells you change this wording here add this period here. I think you're a lot more like actively engaging with it and like understanding
Starting point is 00:52:42 the edits that it's making. So feel free to make news here if you want, but you guys are already integrated with substack, which publishes all of my work on the internet. Are there other similar deals that you are similar integrations that you're looking to get started? So we also have a Chrome extension. So even if we don't have a deal with any individual platform, Pangram will still go and label posts on Reddit, Twitter, LinkedIn, Medium Substack as AI Mixed or Human, which I think is really valuable
Starting point is 00:53:15 for somebody who's just navigating an internet that's increasingly full of AI slop. Next Bureau. Thank you very much. Thanks so much for having me.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.