librarypunk - 168 - Bullshit for Good: Scarlet and Dorothea (productively) rant about AI

Episode Date: July 18, 2026

We’re talking about that bad article on critical confabulations in archives. We talk about history as a process, the archive, gaps in our knowledge, language as a conveyor of reality, and The Waterm...elon Woman.  Media mentioned THE ARTICLE: Critical Confabulation: Can LLMs Hallucinate for Social Good? https://arxiv.org/abs/2511.07722 Guillaume Cabanac, “How a Tortured Conference Becomes a Series: An Analysis of Conference Manipulations” IEEE Xplore  https://ieeexplore.ieee.org/document/11363742  Research repository ArXiv will ban authors for a year if they let AI do all the work, TechCrunch  https://techcrunch.com/2026/05/16/research-repository-arxiv-will-ban-authors-for-a-year-if-they-let-ai-do-all-the-work/  The Watermelon Woman https://en.wikipedia.org/wiki/The_Watermelon_Woman  Keith Haring Unfinished Painting: https://mymodernmet.com/artificial-intelligence-finishes-keith-harings-unfinished-painting/ Patrick Wyman, Lost Worlds https://www.harpercollins.com/products/lost-worlds-patrick-wyman?variant=43084775817250  Predatory States: Operation Condor and Covert War in Latin America, Patrice McSherry https://www.bloomsbury.com/us/predatory-states-9780742568709/  Dorothea series on patron privacy https://ischool.wisc.edu/continuing-education/tech-crash-course-data-information-privacy/  Transcript: https://pastecode.io/s/3oy0sgqw  Join the Discord: https://discord.gg/qWPTurTnkT

Transcript
Discussion (0)
Starting point is 00:00:00 Okay, where's my soundboard? All right, let's... I'm Justin, I'm an academic librarian, and my pronouns are he and they. I'm Sadie. I work IT at a public library, and my pronouns are they them. I'm Jay. I'm a cataloging librarian. My pronouns are he, him. And we have guests, which I like to introduce yourselves. I'm Scarlett. I am a collection strategist librarian at an R1, and my pronouns are she and they. And I'm Dorothea, the guest y'all, can't.
Starting point is 00:00:55 seem to get rid of. I teach in the information school at the University of Wisconsin at Madison and my pronouns are she, her. Welcome back. Returning guests, returning champions. Love that. I didn't even bother counting up how many times Dorothy has been on,
Starting point is 00:01:14 but just unofficial fourth mic of the podcast from the beginning. First guest, longest time guest. Thank you. So I love a good shit talking episode. So that's what we're going to do. I started looking on Blue Sky the other day to see if people were still talking about this article. Because I had a sneaking suspicion that they're never going to get this published because it's a preprint.
Starting point is 00:01:36 And one thing I know about some of the AI stuff that's happening is a lot of these boosterish articles are getting published in or published quote unquote on archive.org where then they are reported on as a new preprint article. and they never actually make their way into any journals. So I'm interested to see if that's sort of what's happening. They still have not published. It's been a month since this made the rounds. But this is the article, and let me pull up the title, Critical Confabulation, Can LLM Sillucinate for Social Good? This is from the same people that brought you,
Starting point is 00:02:16 I believe some of the same authors that brought you the article, Why Slop Matters, Hoyt Long and... Burlett's face journey was amazing when Justin said. And Amen, so two of these articles, or three of them, three of these authors also wrote that article Why Slot Matters, which is published in like a new I-Triple-E journal that has like two issues. Oh, can we shit-talk I-Tri-E?
Starting point is 00:02:47 Yeah, I mean, why not? I would love to shit-talk I-R-R-E. Okay, a few months ago, an article came out by, Guillome Cabanac and his crew. Guillaume is awesome, by the way. I love Guillem. Fantastic scholarly communications researcher. And basically a ridiculous percentage of crap publications in the relevant subject areas are from ICCLEE.
Starting point is 00:03:14 Because ICCLEE is basically a franchise operation, right? If you want to have a quote unquote reputable journal, you can sign up with ICCI and basically they don't check. And yeah, you can publish all the crap you want. Hmm. Well, I was wrong. This is ACM. A journal called ACM AI Letters, Volume 1, Why Slot Matters. And then the next one is why AI slot matters, but not like that, which is a response article in an upcoming issue, but it has not been published yet. But anyway, I was looking to see if they ever got this article published because, like, there is this problem. And some of the citations that aren't from digital humanists that are from other AI authors are also citing archive
Starting point is 00:03:58 preprints. So there is kind of an endemic of these questionable publications that don't seem to ever get into any journals, which is not terrible. I mean, from a Skalkan perspective, fine, but when you're, there's, there's clearly like a communication strategy here that I think is problematic and has not been discussed at length. I know people have mentioned it, because I obviously I'm not the first person to notice this, but it is an interesting trend to keep an eye out on. Well, like when Archive finally, this happened in a couple of stages, so Archive came out and said, we can no longer trust institutional affiliation for email submissions. So if you're the corresponding author and you have an email connected to an institution, we can't trust that anymore. That was the first thing
Starting point is 00:04:44 that Archive had to come out and say. And then the second thing that Archive came out and said was, we're just going to ban you from submitting if you send us papers with hallucinated citations. We're taking a hard line on it. And the number of people who came out and absolutely told on themselves, their affiliates, and all of their labs is incredible to me. So many people came out and said, there won't be any PIs left if you go ahead and do this. And I said, there's a line eight deep behind you who will show up, take over your lab and actually read everything that they're citing. because that's what you're supposed to do. I thought that's what we were...
Starting point is 00:05:20 My gosh, it was amazing to see. I would love for someone to go through and do a study of all of the output of all of the authors that complained about this. It's incredible. Like before AI could do this, what the fuck were they doing?
Starting point is 00:05:34 Also, why would you say that out loud on LinkedIn, on social, where people can see you? Like, show your entire ass. Right. Like, you just want to pull people aside and be like, you're saying outdoors,
Starting point is 00:05:47 thing, you're saying indoor things with your outdoor voice, my friend. I don't, what's going on? Yeah, I mean, it's really strange. I think part of the things people were saying was, well, there's so many co-authors. We can't trust that one of them is not going to use AI and then screw it over, screw over all of us. It's like, well, just read your citations before you publish it or before you submit it. It's not that hard. And I think my, my theory is, because these are getting into, you know, prestigious journals that are supposed to be doing peer review, whenever they find these hallucinated citations, they're just going to go back and remove them without actually retracting the articles to make them, to make the journals look better. Like, the journals
Starting point is 00:06:28 are going to go back and do this. Like, Elsevier did this with Sci Hub DOIs at one point. Back in, like, 2019, when people started pointing out how many SciHub DOIs there were in Elsevier papers, they just went back and replaced all the DOIs without any retraction notices without any addendums or anything because they care about papers as a product and they don't care about how, you know, the paper trail, but actually how you cited something. So I used to collect a lot of data on that. And I have, I have years of automated data that I scraped based on like, in like my inbox and of DOIs for SciHub. But anyway, that's what we're not what we're talking about today. But we do have some trivia and news section.
Starting point is 00:07:14 So do I have a news drop? Hang on, I haven't been using my drops as much. Oh, feel free. I find them delightful. I really do. Yeah, see someone, someone does. Chat, DVDs for deluxe. Oh, no.
Starting point is 00:07:25 Have I taken the side? Oh, no. You don't have to be gay to get married to a man. I mean, you did it. Yeah, that's my news drop. So, Starly, Good News. Yeah, so I have a new job since the last time I was on the podcast. And I don't know that I ever publicly said thank you because the best thing about writing my tenure dossier for Grand Valley was putting library.
Starting point is 00:07:49 Dot Gay on my CV, having the provost read it and say that essentially that she had a blast reading it. And then she quit two days later after we all got the mass blessing from the board of trustees. And now I'm also at UChicago. I think Dorothea has a new position since you're managing. Sort of. Sort of. You're managing people finally. I am managing people who let me do this.
Starting point is 00:08:15 Somebody stop me. But yes, it is my first kind of, I mean, I've supervised TAs and project assistants and stuff in the past. But I am actually supervising people now, eight teaching faculty who are basically like me, permanent academic staff. And four part timers who have recurring appointments, so they teach for us again and again and again. And as management gigs go, this is like the easiest one ever because the people that I supervise are absolutely without exception. Fantastic. They're amazing. And I love them. Yeah, I'm supervising a teen that I have inherited rather than one that I've built, which is new for me. And they're all older than me. So they're all like Gen X and older. Some of them have been in their jobs for like 40 years at the same place. like the serials like one of them is like a serials librarian or the library assistant who's been there and now he just is the only person left from the cereals department.
Starting point is 00:09:18 Everyone else is slowly like no longer had a department anymore. So now he's he's on my team and there's still plenty of cereals work to be done because turns out they take up physical space and they got to get moved and we're absorbing other libraries and taking on all their crappy cataloging practices and having to fix stuff. So, yeah, independent libraries. I don't understand why universities have independent libraries instead of just one library system. That's one thing also that I'm having to deal with is people telling me, don't touch my stuff. We're the different library.
Starting point is 00:09:51 But like, why? Buy your own Alma instance then. They've basically said they don't want us helping them with their primo instance anymore. So like we set it up for them and now they don't want us to touch it. And it's like one of our systems people, or one of our discovery people has gotten very annoyed at this. And it's like, fine, you don't want my help. I will not touch anything of yours anymore. That is the correct response for something like that.
Starting point is 00:10:16 Like, nope, go ahead and have. Really is. And if it breaks, you're on your own. Yeah, it's very silly. But anyway, what else do we got? Richard Dawkins fell in love with an AI. Yeah, we don't have to talk about that. But Richard Dawkins fell in love with an AI.
Starting point is 00:10:35 It's going to happen to more people with large platforms. And while it's extremely depressing that it will happen to more people, it's very funny that it happened to Richard Dawkins. One of the other articles that we have in kind of the working document for this discusses AI and Pygmalion Displacement, which is one of the many reasons why Richard Dawkins was compelled to give his AI a femme name because he couldn't handle Claude being called Claude. He had to rename it and name her Claudia in order to be okay with the level of enmesment. Yeah, immensiment that he's experiencing. But yeah, he basically believes that it's conscious and alive. And that's Richard Dawkins at the end of his life now, which is incredible to me. You know, I just, it's going to happen to more people.
Starting point is 00:11:24 And I think when it happens on an individual level, it's going to be heartbreaking and awful. But when it happens to Richard Dawkins, it's incredibly funny. No, I mean, our friends have a dating and relationship podcast. and I was totally pitching an episode of doing like AI as my boyfriend Reddit. It's really interesting because the people who are having the most like relationships with AI are women. And one of the things that was interesting when I was going through posts on like my boyfriend is AI was someone saying, I'm so glad I found this one because most of the other places are entirely women. And I wanted to find a place where other men have fallen in love with AI.
Starting point is 00:12:00 But it's still dominated by women. It's very interesting phenomenon. and very strange and very sad. And it's very in-cell-coded, too. Like, it's very much like, they're like, don't. They're like, we are in love with AI because we have no other option. And like, what I would give to have a real person love me. They're constantly talking about like, oh, I love my AI boyfriend, but what I would give.
Starting point is 00:12:23 It's very in-cell-coded and very strange, very sad. Yeah, it's not good this is happening. No, no, it's really not. I know that there's probably going to be some. you know, significant study your work about this, but, but I, you know, remember kind of exploring a little bit of that just to sort of see what was happening. And I mean, even, you know, Wisenbaum wrote about this when he built Eliza. Like, no one should see what you, what you're typing into this thing. People are immediately treating it like, you know, like a therapist, that there was this,
Starting point is 00:12:53 you know, immediate, you know, immanchement factor sort of happening, even with something as basic as Eliza, which my, you know, accomplishment when I was eight years old was getting Eliza to swear at me, you know, or get angry. You know, because that was the fun thing to do when you're eight and somebody has a, you know, has a computer. Amber, you know, monitor and firework screen savers forever, right? But it's interesting and scary to see all of this happen, especially in the current context. Yeah. So anyway, let's get to the article because I'm sure we've got plenty to talk about on that.
Starting point is 00:13:31 So critical confabulation can LLMs hallucinate for social good? The idea is they want to propose critical confabulation inspired by critical fabulation from literary and social theory. The use of LLM hallucinations to, quote, fill in the gap for omissions in archives due to social and political inequality and reconstruct divergent yet evidence-bound narratives for histories, quote, hidden figures. Basically, what they did is they took this corpus of, work called the got it highlighted black writing and thought collection and what they did was they
Starting point is 00:14:08 made sure that none of it was already ingested into their lLM so they used open source LLMs where they know what the training data is they ran the corpuses against each other and tried to cancel out anything that was possibly already in there and then what they did is a narrative closed task CLOZE which is basically a closed task is where you take out bits of information in a sentence. This is how you do quizzes for students learning a language. You take out pieces of the information in a sentence, and by context, you want the person taking the test to fill in the gaps. And so they're doing this with narratives, and they want the LLM to see if it can guess what they know is in the corpus, but it doesn't know is in the corpus.
Starting point is 00:14:57 So they give it parts of the story from the corpus, blot out certain parts, and then want to see how accurately it can guess the parts that were blotted out. Basically, taking a timeline and then adding and seeing, can it figure this out? This sort of combat, quote unquote, what they're calling a sort of like archival silence, but archive used more broadly than how we would use archive on this podcast, but like gaps in historical record. I don't think they understand the difference.
Starting point is 00:15:26 They don't. So the thing about this is I don't think they understand any of the digital humanities scholarship they read. And I don't think they read it. I think the way that it's summarized and misunderstands the source material indicates to me that they also used Gen. I.I. to write most of the literature review because they don't really seem to understand it. But there is in the notes that there is a rough outline of how history happens. I don't know who added this, but whoever did. So, yeah, so I, so I added a bunch of stuff in here.
Starting point is 00:15:59 And just a note for audience members that don't know us is that Dorothea and I aren't archivists by training. So we should probably say what our approach to something like this would be, I'm a historian by training. So I've worked in and around archives for all of my academic life, but not in the same way that an archivist would, right? So when I talk about this, you know, doing kind of a high-level pass here, but not the way that someone who is a trained and active historian would, because I'm not a historian, I'm a librarian, and not the way that someone who is an archivist would. So someone who is better at French pronounce this name that I have highlighted in here. I have a comment that says, I cannot French. Someone help me, and it is because I read more than I speak. And so French words were an incredible embarrassing. when I was a kid. I don't know how to pronounce this person's name. Michelle Rolf Trio? I think it is Trio, yeah. Okay.
Starting point is 00:16:57 Yeah, it should be Trio. Okay. So one of the things that I read to get into this was Michel Rolf Trio's silencing the past power and the production of history to look at what we, one, sort of what we mean by archival silence and how it interacts with the practice history, because this is sort of what this paper is trying to get at a little bit. And I thought, okay, fine, we'll take it on their terms and look at it. So there's a bunch of places where the archive is effectively silent or has an absence. And Trillo writes that there's the moment of fact creation, so the making of sources, the moment of fact assembly, so how we make archives, and that would include things like
Starting point is 00:17:37 arrangement and description, the moment of fact retrieval, so having gotten at all of these sources, how do we then create and structure narratives around it, and then the moment of retrospective significance, so the making of history in the final instance. instance. Good recent examples of this or anything really related to the Cold War or the dirty war, especially under Pinochet. Somebody opens a footlocker in Argentina or the former Soviet Union, and suddenly that making of retrospective significance that he's writing about becomes different. Our understanding of that source making of assembly and retrieval then reiterates on those primary and secondary sources. So this is what critical fabulation, the concept that inspired this paper,
Starting point is 00:18:22 is trying, you know, kind of to do. So Sidiah Hartman is the scholar responsible for the idea of critical fabulations. And she first introduces this concept in another essay called Venus and Two Acts, where she talks about this. And in it, what she's doing is trying to reconcile that archival silence in the context of two girls who were murdered aboard the slave ship recovery as part of the Atlantic slave trade. And so she's sort of looking at where are they explicitly mentioned, where are they not, where do sort of catalogs of people as objects, you know, become, you know, where do you stop saying, for example, how people died and you just put the quotation marks. It's because everyone is dying the same way or everyone is sold.
Starting point is 00:19:10 or everyone is like all of these kinds of, that there are these sorts of silences that she's trying to deal with and reconcile. And I'm going to read a couple of paragraphs from Venus and Two Acts, which if you haven't read it, is incredible. Everyone should read it. So this is her talking about critical fabulation. The intent of this practice is not to give voice to the slave, but rather to imagine what cannot be verified, a realm of experience which is situated between two zones of death, social and caporeal death, and to reckon with the precarious lives, which are visible only in the moment of their disappearance. It is an impossible writing which attempts to say that which resists being said, since dead girls are unable to speak.
Starting point is 00:19:50 It is a history of an unrecoverable past, a narrative of what might have been or could have been. It is a history written with and against the archive. Admittedly, my own writing is unable to exceed the limits of the sayable dictated by the archive. It depends upon the legal records, surgeons, journals, ledgers, ships manifest and captain's logs, and in this regard falters before the archive's silence and reproduces its omission. The irreparable violence of the Atlantic slave trade rests precisely in all the stories that we cannot know and that will never be recovered. This formidable obstacle or constitutive impossibility defines the parameters of my work. The necessity of recounting Venus's death is overshadowed by the inevitable failure of any attempt,
Starting point is 00:20:36 to represent her. I think this is a productive tension and one unavoidable in narrating the lives of the subaltern, the dispossessed, and the enslaved. So let's stop for a second. An LLM could never. I would give limbs to be able to write like that. My God. And also, she's very deliberately saying, don't do the thing that's being done in this other essay. She's very deliberately saying that the whole point of sitting with archival silence as a human in the world is to sit in its totality. It is to do all of those things. We're supposed to sit with it. Applying AI to this process is really removing the humanity of both the subject and the researchers themselves and really speaks to a bunch of things posed by this, you know, critical confabulation paper about how the humanities are incentivized
Starting point is 00:21:33 right now. But Sadie, you've got a comment in here, and I want to make sure that you're in as well. Oh, I just, when I read that and the tension is necessary and we are supposed to sit with it, I just didn't, what immediately came to mind was the person on what is formerly known as Twitter who took Keith Herring's unfinished painting and finished it with AI. And then when people flashed back at that was like, well, I didn't really know the significance of this painting or the history behind it. I just thought it was kind of funny and I'd probably do it again even if I knew the significance of it. And I'm just like, how can the point of art be so missed? Especially because the AIDS crisis is one of those things that is really hard for me to sit with when I think about
Starting point is 00:22:28 like all of what was lost. So Keith Herring's painting really pulls something out of me and then to see that and just be like, it's not even like the things that make Keith Herring's art are not represented at all in that AI finished version. The motifs are completely messed up. The only thing you could say is that they use the same colors and lines.
Starting point is 00:22:55 Yeah. It's not art. So, yes. this is, I, it's the, what is it, the torture nexus, the thing that people are like, oh yeah. Torment nexus, yeah. The torment nexus, you know, the person, the guy who writes, hey, don't make the torment nexus. And then somebody turns around and goes, so we made the torment nexus from the great book, don't make the torment nexus. It's just, yeah.
Starting point is 00:23:16 Whereas like what this reminded me of like, because what I feel like of these authors and these, like, these researchers are missing the point of what critical, like, there's like this, like, this, sitting with the tension that's like human reckoning with it but like also critical fabulation doesn't fix the problem of archival silence
Starting point is 00:23:37 if anything it brings more attention to the fact that it cannot be fixed but it also reminds me of like what immediately came to mind was the film that we've talked about on here before called the watermelon woman which is an act of sort of like highlighting the absence
Starting point is 00:23:53 of knowledge of black queer women in early cinema history by doing a sort of mockumentary around this made up black lesbian in this movie that she's tracking down.
Starting point is 00:24:09 And framing it is it's like a form of like myth making in order to show that like this isn't real, but we're doing the sort of like radical myth making around it to show the fact that you know, people like her probably did exist in cinema
Starting point is 00:24:25 but we don't have record of that because of the white supremacist nature of this country and this industry. Right? Like the point of the watermelon woman isn't to like surprise, we fixed and found this and made this up. It's to like construct this narrative and do this kind of like radical myth making
Starting point is 00:24:48 as a way of showing the inherent violence of the record, right? Yeah. It's, you know, and And Cedia Hartman, you know, talks about this extensively in this essay and in other works that to work with and against it is to, is you're essentially working with and against power in going and finding, you know, all of these particular narratives. And so, you know, one of the questions that I sort of have when I'm looking at this paper, and others like it, are really about how the humanities are incentivized now in that, you know, not just these particular authors, but, but also
Starting point is 00:25:27 members of, say, you know, the American Historical Association who never saw a bad idea that they weren't willing to endorse. And that includes things like the use of AI to create historical artifacts, which is absolutely mind-blowing that that endorsement is there. But that so many humanists are in some ways abandoning their own training. And I think that's, you know, fairly interesting. This is the humanities in a lot of ways trying to thrive under fascistic conditions because those are the conditions under which things like large language models were developed, and they're the conditions which sustain higher educations in interest and investment in AI tools, you know, in a lot of ways.
Starting point is 00:26:07 And so when we look at, like, I have a big list about things here about, like, what the paper is trying to do and what it's actually kind of doing. And I keep coming back to like LLMs as we understand them. Even the open models were essentially trained on, you know, angry Yelp reviews and subredits of people that can't be within 500 feet of a school. What do you think they're going to say about black life and black experience? How could, and I just don't understand, you know, so much about why we would use tools that enforce and tend toward the average. If we're really trying to find something extraordinary,
Starting point is 00:26:52 if we're really trying to, if my goal is a historian is to find, you know, extraordinary lives in the mundane and try to find those narratives, why would I use a machine that insists on creating a narrative mean to identify those silences? It just becomes a trope and not history. And if I accept, you know, a claim that, you know, sort of showed up from my undergraduate background, that, you know, if history really is a conversation between the present with the past to understand the future. And I remove the humanity from one or more parts of that. What have I done? That's not accepting the totality of those experiences and stories in the context of their silent like Venus and Two Acts is trying to do. And that, you know, the authors are claiming to be
Starting point is 00:27:38 inspired by it's a kind of necromancy. And I don't think I have words that are polite for that. I feel like a lot of, because I went and read Hartman's Venus and two acts, and because I hadn't read it before, because I wanted to, I had this thought that a lot of the language that they use in this article that they don't fully understand, like the way they use archive, it sounds like, and I think the way Hartman uses it is in sort of like the media studies term where archive means like the archive of humanity, like the things that we know in terms of all of writing and what is allowed to be known politically and what's in our political discourses. And I think that's the media way that she's using it. But they seem to think like, oh, there should be archives where we can go in and grab other data from the archive
Starting point is 00:28:28 and fill in gaps in the actual archive. And some of these other papers that they're citing are about like making a case for black digital humanities. because they also use the word recover in a lot of ways that I don't think they are using correctly, but I didn't have access to, like, I wasn't quite sure. Like, I pulled one up right now to see, like, how the word, like, you know, recovery rests at the heart of black studies as a scholarly tradition that seeks to restore the humanity of black people
Starting point is 00:29:00 lost and stolen through systemic global racialization. Recovering lost historical and literary text should be foundational to the black digital humanities, recovering alternate constructions of humanity that have been historically excluded from that concept, politics of recovery, not only as the recovery of lost or non-canonical and difficult to locate texts, but also recovery of black author's humanity. See, that's what I thought exactly they were like overlooking. They're like, oh, we can find these things. And there's one thing they say that like, we can at scale the, we can scale the discovery of divergent yet evidence
Starting point is 00:29:37 balance. I was the one that made me rip my fucking hair out. At scale. Yeah, like, when was scale something that we valued in this?
Starting point is 00:29:49 I don't even know. We're going to have a database of known unknowns, Dorothea. It's going to be great. An unknown unknowns. Yeah. Rumsfeld.
Starting point is 00:29:58 So there's also this, this like lack of care, which is like, central to Hartman of like there's a line that I outlined which is I want to tell the story about two girls capable of retrieving what remains dormant the purchase or claim of their lives on the present without committing further violence in my own act of narration so all of this is about like the humanity of the subject the impossibility of recovery the impossibility of filling in the gaps
Starting point is 00:30:31 and I don't know if you all read past the footnotes, but you can see what stories they use, and one of the stories is like this 14-year-old black child who gets shot, and that's part of the story they cut out, and they do a close study on it. It just seems like there's really no care. Someone on Blue Sky pointed out that it can't be critical fabulation or critical confabulation if there is no one being critical.
Starting point is 00:31:00 like the critical part of it is like about critical studies. And if there's no agent being critical about the materials, then there is no criticality actually happening. And it just seems like these authors really don't care about the subject matter in any way, especially considering the other stuff they're authoring why AI slot matters. I keep coming at this because I am always charmed whenever I get to use the word bullshit in a serious academic context. This, of course, goes back to the philosopher Harry G. Frankfurt on bullshit.
Starting point is 00:31:34 But there's also a charming piece in the journal Ethics and Information Technology. I'm sure most of you read it. Chat GPT is bullshit. And it turns out that bullshit has a very specific definition deriving from Frankfurt, which is truth. Who cares about truth? I don't give a fuck about truth. I'm just going to, you know, send my slop out there into the world.
Starting point is 00:31:59 and its truth value literally doesn't matter, and I may not even know it. And that so strikes me is what's happening here. What these people are doing is bullshitting, because they don't care about the truth of any of these people, and they're filling in or demarcating, whatever, these silences, as you say, without care. You have to care about these people.
Starting point is 00:32:27 I think you have to care about their truth and you have to care about their truth getting cut off. And I just, it really irritates me the amount of bullshit that is just accumulating in this supposed research method. Yeah, that leads me to a question I have for the two of you.
Starting point is 00:32:51 Because like what this, one of the problems that I was having with this article was that like, they are citing so many scholars in the black digital humanities and like the post-colonial digital humanities. They might not be citing them well or accurately,
Starting point is 00:33:07 but they are citing them. They exist. And it's like this is a problem that I've had with like the A, like with the fact that AI is often used as a marketing term to include all these other technologies, some of which have been used in the digital humanities for years, right?
Starting point is 00:33:23 As a way of like approaching specific types of research questions, right? Like they're not about solving a problem or fixing something, but it's like here is a different way to approach this research question, what can we get out of it, right? And people do interesting things with that. Sure. And it's like they're taking like the way that someone in,
Starting point is 00:33:46 I hope Justin Burp shows up on my recording now. Like they're like the way that the people in the humanities and like the digital humanities, especially like in like black and post-colonial digital humanities, are approaching this research question or this area of like silence in the archival record, however we are using archive to either mean literal or more broadly, right? Like this question of power and violence affecting historical record
Starting point is 00:34:24 and like how people in those disciplines are approaching it versus like, you know, there's very similar technologies like natural language processing that these scholars, these scholars, are using that people in the digital humanities are using. It's like not that different, but it's such a different way of coming about it. And it's like, is it because of like the,
Starting point is 00:34:50 like, is there something in the epistemology or the methodology of, of these disciplines that makes them like, because it feels like the AI confabulation folks, it's like it's a problem to solve instead of a way of exploring a research question. It's like, no, there's a problem and we can solve it with the data.
Starting point is 00:35:10 Like, I didn't know if y'all could speak to that at all. So for me, there are a couple of things that come into play. One of them is just straight up accountability, right? And that ties into the explainer. ability of some of these methods. With a natural language processing pipeline, it's not quite what I would call completely deterministic, but you can take the process apart, you can interrogate it, you can tweak parameters and see what happens, right? And I think that's part of being responsible as a
Starting point is 00:35:48 researcher. With Gen AI, there's no explainability. If you give it the same prompt twice in a row, you can get wildly different answers. So if you can't in any way, it's the lack of control, I guess, that's really bugging me. There are a lot of ways to control an NLP pipeline and make it do more or less precisely what you want. With generative AI, You cannot do that. It is not controllable in this way, which means I think it's really, really irresponsible to call this a research method. Does that speak to your question at all?
Starting point is 00:36:30 Somewhat? Because I do agree with you. Yes. I was more like, is it something like, I guess I was like thinking about like, is there something inherent in these two disciplines of like the way they approach epistemology and methodology, which I guess like if you are willing to use an LLM generative AI. as a legitimate form of research that does say something about how you approach methodology.
Starting point is 00:36:54 But like, it was just like, I was so assounded like, why are they looking at digital humanities stuff? Like, what about the, like, this research area, this research question, this very humanities thing that just like there are, you know, techie ways of exploring it. What is it about this that was interesting
Starting point is 00:37:15 to these authors that they're approaching it? this way or doing it so badly. I think we've seen this before, though. What was it that Google, when Google books happened, there were these absolute, there was a raft of absolutely craptastic attempts at linguistic and discourse analysis by people who didn't really understand the limitations of the corpus, didn't understand things like long S, that's a thing, it's going to mess up your word counts, y'all.
Starting point is 00:37:47 Cultureomics, is that what they called it? Whatever they called it, it was trash. And it's, to the extent I can be sympathetic to it at all. It's, as you said, a little bit of playing with blocks and let's see what happens. And there's nothing intrinsically wrong with that impulse, I don't think. But you have to be honest about your limitations. And your limitations of understanding most of all, if you don't have control over, and boy, I saw this in DH2,
Starting point is 00:38:24 if you don't have control over your understanding of what's in your data, your understanding of how the pipeline works, there's a lot of DH work that just happily adopts whatever tool and basically treats it as a black box, which I think is very much what's happening here. Yeah, it does not do much for the legitimacy of your results. Let me put it that way. Yeah, I think your question
Starting point is 00:38:50 is a really good one. The thing that I pulled from the article two is, is one thing I keep coming back to is that, especially when you're looking at a narrative closed task or series of tasks in natural language processing is that tools measure the limits of experience, right?
Starting point is 00:39:07 So are we looking for narrative structure in archival silences using a colonial or supremacist or inherently like white sort of notion of storytelling or way of kind of filling that, you know, in that blank. One of the interesting things about the methodology behind the paper is that they exclude, those researchers exclude folklore and song from that original database. I'm like, oh, you mean the way that people who don't have power
Starting point is 00:39:33 and are displaced and dispossessed tend to deal with the fact that they are not in the, in the archive as it's, you know, either narrowly or broadly conceived, like the way that we community. We're just going to exclude that and we're going to exclude that and, you know, in something that, you know, is ostensibly trying to reveal where potential silences might be when we're talking about black histories and culture. So like one of the quotes that I pulled was that, you know, the implications are twofold for natural language processing. We demonstrate a domain application where hallucinations become a unique and optimizable resource rather than a pure deficit. For the humanities, we highlight the narrative understanding capabilities of LOMs that make them viable
Starting point is 00:40:17 tools for helping scholars probe the latent semantic spaces of large archives and to surface unknown unknowns and then to rapidly prototype candidate reconstructions to narrow the bounds of the known unknowns, right? So that's, you know, what, you know, what they're trying to do? Again, it's like, in service to what? Because there you're not, there is not that, you know, human experience, there isn't that sitting with the incredible weight of it of what it means to be near an unfinished painting. You know, I have, or unfinished art, I actually have had a similar experience with another artist. I can't remember the name of them right now, but it was a series of when we first visualized HIV, it was a series of crystalline and glass blown, like sort of large scale pieces of what the virus looks like. you know, blown up and set so you could see it. And it was someone who, you know, was living with
Starting point is 00:41:14 HIV, who has since passed away, who was the artist who did that. And he said, you know, it is incredibly powerful and important to me to be able to visualize the thing at scale that will kill me. And that there's this, there is an incredible sort of weight of loss and potential and possibility and, you know, sort of the micro and macro and way that, you know, the art is engaged with itself. and all of that kind of work. And if you don't do that, if instead there's just, we're going to create, you know, sort of a corpus of silences or something, I'm just, I'm trying to get my head around the use case for this because so much of what forms the confrontation of power that Sophia Hardman is writing about and that these people are supposed to be inspired by is that confronting of that power. And so I don't, yeah, I just kind of wondering, does it even work? when you've effectively removed to the people from that equation at all. I know I keep coming back to that question, but it is very strange to me. And I wonder so much at the power of incentives that are working in the humanities now because of this kind of stuff.
Starting point is 00:42:25 I mean, I imagine they're trying to sell something. They're trying to, like, I think why, like, why focus on black scholarship? why focus on this corpus, why focus on specifically archival work where you could feed a lot of information in and create these like possibilities of timeline billing events? I have a hypothesis. Yeah, go ahead. And that hypothesis is these people are deeply uncomfortable with the idea of silence. It really bugs them. And in a very, and I'm a white woman, so I can say this, in a white woman's tears fashion,
Starting point is 00:43:10 they're trying to appropriate these silences to fill them. What's the phrase they use? Let me look at this. Oh, yeah, prototype candidate reconstruction. What the hell is that? But, yeah, it is absolute wankery. You are correct. But I don't think they can do the work that Scarlett says is so important, which is just sitting with that silence, accepting that it will not be filled ever in the absence of more discoveries, which can always happen but probably won't.
Starting point is 00:43:48 And so they're bringing in the bullshit machine because they can't think of what else to do. It's, you know, it's like the AI significant others. it's profoundly sad and also profoundly messed up. Because AI also can't handle silences, right? Like that's why we get hallucinations in the first place because it's just doing pattern matching what makes sense to come next. And if there's nothing,
Starting point is 00:44:15 it can't reliably put in, well, it's going to make something up. Yep. Right? Like, that's the whole fucking point of these generative AI, like LLMs, is to do that kind of, we can't handle the fact
Starting point is 00:44:29 that there is an unknown, so I will make something up. Like, the tool itself, you know, I'm going to get all McLuhan videodrome, like the medium is the message kind of thing. But like when literally the tool itself is something that like abores a silence, a boars a void, like how can you do meaningful research or even attempt to explore this question with this tool. Right on. I totally agree with you. I was reading their conclusions,
Starting point is 00:45:04 which is like two paragraphs, and they say they're exploring a wide range of use cases for critical confabulation to support humanistic scholarship and augment human storytelling towards new AI-enabled methods for studying culture and history, for paradigm shifts, whatever.
Starting point is 00:45:20 And then it says, like, this work is preliminary in nature. Future work will design more robust and comprehensive evaluations for open-ended narrative outputs, broaden coverage across languages, build ethical safeguards and provenance tracking to ensure faithful event reconstruction and avoid compounding archival violence. But faithful to what? How could you do any of that?
Starting point is 00:45:44 And, you know, speaking as a sometime storyteller, who told them that sort of that human storytelling needed augmentation? Especially by bullshit machines. Just miss me with that. Yeah, I mean, that's what really makes me think that it's their, they're, they're, seeing if there's interest in spinning this off into a product, because I seriously doubt there's going to be future work, considering like how little they seem to care about like the topic. They just wanted to see, could we get this thing to guess how to fill in these narrative gaps
Starting point is 00:46:16 using like real world data and can't do it? Which also it can't. It's like 50-50. It can kind of guess 50% of the time. If you babysit all. the prompts constantly and know what the answer is, which of course, in a real silence, you wouldn't know what the answer is. Scarlett, you're a historian. I'm not. What's, the historical history is a discipline? What is its view on making shit up, no matter who's doing it? Well, yeah. I mean, so, so the way that, I mean, I've seen historians handle and sort of triangulate through archival silence really well. One of the examples I think I have, you know,
Starting point is 00:47:02 over in the notes document is from, I believe, Patrice McSherry, yes, who wrote predatory states, which is about Operation Condor and Covert War in Latin America. What Patrice McSherry was able to do, and this is surfaced in how things like the dirty war and Operation Condor were executed, you know, in particular under Pinochet and others, is that they hid to each other's captors and hit each other's records. So you had to go through sort of this multi-archive investigation to try and line up all of this information. I think that, you know, generally speaking, bullshit bad was my takeaway as an undergraduate, making stuff up not helpful, right? Unless we're looking at things like forgeries, right, which are their own sort of sense of, own area of study. But, you know, I also
Starting point is 00:47:52 In that case, it's not the historian who's making it up. Yeah, it's somebody else. But I don't know. I also go back to this same group of researchers construction of slop. And it's like, of course slop is important. Of course, lowbrow, anything is important, especially now since AI slop is used to kind of propagate weird fascist nonsense. So in that sense, it's worthy of study.
Starting point is 00:48:16 But again, not the way that they're thinking. Of course, archival silences are worthy of study. and can have all of these different methods and tools, you know, applied to them, but only some of them are going to be thoughtful or yield results. Only some of them are going to be ethical and only people can engage in the kind of care that's necessary in order to make them all happen, which I think is sort of my position on, you know, this kind of work and also bullshit generally. Thank you, Scarlett. That was beautiful. Yeah. I mean, I'm finishing up a book now about ancient cultures, like trying to.
Starting point is 00:48:51 different things and how our understanding of the prehistoric world is shaped by like what we're able to reconstruct. And sometimes if we don't have descendant languages, you can't figure out anything about these cultures because we don't know what words they had for things. Whereas with Proto Indo Europeans, you know that they had a word for wheel and parts of a wheel and parts of a carriage and riding in a vehicle. Whereas in other cultures, all you have left is just how they built their settlement. So there's all of these sitting with the gaps and saying, like, we don't really know anything. And that's kind of the best you can get, even though we have some evidence of, like, they made pottery for this. And we can tell, like, their society was unequal
Starting point is 00:49:37 because we saw inequality start to develop in burial rituals, things like that. So I think whenever you have these kinds of things, the silence tells you things as data points, which is why a lot of the archaeology and social sciences split off from history was to study those non-narrative things and to say, like, these are our data points now, whereas before we would say, like, the Fertile Crescent is where civilization started because it had these things, and it turned out our definition of civilization was just survivorship bias. And actually, the stuff that happened and the fertile crescent was the exception, not the rule. In fact, it wasn't very, like, they made writing, and we know about writing,
Starting point is 00:50:22 and we know about them because of their writing, but no one speaks their languages anymore. The civilization did die out. So, you know, it was the proto-winder Europeans who came through, illiterate as they were, and took over all of these areas. So, like, what you don't know and what you do know determines what you can write about historically or in pre-history. So it really seems, yeah, there's. there's probably something libidinal in terms of like the fear of the
Starting point is 00:50:48 violences. It's probably also why pre-built chatbots are so chatty and you have to tell them to like stop talking to you like a chatty co-worker and just give you the answer and not be like, you're right. I did get that wrong. emoji emoji emoji. Did you catch? And you have to tell it like stop doing that. I would have to go find this. I didn't bookmark it. But there's at least one software outfit that is reducing its token spend by telling the chatbot to talk like a caveman. Yeah, yeah. Yeah, there's a built-in, I forget what they're called, but like skills or something,
Starting point is 00:51:24 things you can like tell them to do. I'm learning some of these terms because my university has like their own bespoke AI lab. So we have access to like five different models and you can build stuff and the library is trying to test out building a discovery AI. No! Oh, sorry. That's one of my video moments. It doesn't work great.
Starting point is 00:51:48 The one that they built, I built my own just because it was already built in. SDK, those things that are built in SDKs. Anyway, I built my own, and it works way better than the one they built. So that'll tell you, whoever built it. Yeah, they did tell it to talk like the caveman because it was saying,
Starting point is 00:52:06 I remember early on someone was saying that people were costing Chadji, GPD money because they were saying thank you to it. It also is the same thing. Cost, chat, GPT, all the money, people. Keep being polite. Love it. Bankrupt them with politeness.
Starting point is 00:52:22 Kill it with kindness. Yes. I want to say the Chipotle compute is still going where you can use Chipotle's helper agent to do your coding because it hasn't imposed any limits on it or hasn't imposed certain limits on it. so someone has written a script and put it up in GitHub so that you can, if you run out of compute, you can submit it to Chipotle.
Starting point is 00:52:49 That's funny. I think it's something like free burrito or something. I'll have to find it, but it's just incredible the kinds of things that people are doing. And I know that we just rolled out in the last three days, Quad, Enterprise Cloud, at work. And so I'm just sort of paying attention to where the token apocalypse is happening at other schools like ours. So, like, I know it's already happened at Johns Hopkins, where the, the, yeah, the token apocalypse has happened. And I want to say at least one other, Penn, yeah, University of Pennsylvania is also trying to manage the token apocalypse. And it's just kind of like, okay, and we froze graduate admissions.
Starting point is 00:53:29 So those are sending some interesting messages, or at least graduate admissions in the humanities, right? So there's a lot of messages going on in different corners of the elite institutions. right now. I want to go back for a second, actually, to what was being said about the split of anthropology, archaeology, from history, and the way that that centers around language. One of the issues, certainly not the only one, as we've discussed, with this, let's use language as our soul interface to history here, is precisely that language is a shitty, absolutely. Absolutely. terrible, not faithful at all, representation of reality. Right.
Starting point is 00:54:16 So anything that you're going to get by throwing generative AI at archival silences is going to be impoverished because the only medium that it has to work with to construct whatever confabulation is constructing is language. That's all it's got. And it's not enough. Yeah, we talked about this when we had you on to talk about Bivframe and how people were trying to to use language in such a way that you could encode all the semantic meaning of Hamlet and Bibframe. Yeah, can you write Hamlet and RDF?
Starting point is 00:54:55 No, no, you cannot. Yeah. But even then, you know, those are two, at least they're two language-based questions, right? Archives have materiality, even digital ones. as Trevor Owens. And of necessity, a purely linguistic-based approach locks out all of those other approaches to figuring out what happened to a particular person or at a particular time or whatever. Thank you, actually, for that. That was a really neat piece of insight that I didn't have before, and I appreciate it. I think the real split starts with the analysis school,
Starting point is 00:55:34 and they were saying that because literacy is limited to the upper classes throughout history, the only way you can do actual history is to work with non-literate data. And so that was their, I think, Marxist approach to history, which actually led them to abandoning the field of history and making their own, which is now sociology and archaeology and anthropology of using population statistics and using things that can be quantified more rather than. linguistic. And so that was the big first school that you learn about when you're doing your historiography classes in grad school. That is very cool. Thank you. Because history relies on
Starting point is 00:56:14 written words. Like that's why we have prehistory and history. Like history relies on writing. So if you don't have writing, you're prehistory. Sometimes even if you do have writing, you're still considered prehistory because you're not in the, you know, like the ancient Americas are considered prehistory even though we have writing from them. But a lot of it was destroyed. and some of it's still not interpreted. So every time I go to, like, a big city's museum, I always want to see the stuff from the Americas because there's always, like, ancient, you know, writing
Starting point is 00:56:45 that you just don't learn about in school, that you just are told that, like, all Native Americans were illiterate and didn't have writing systems, which they did. What? No. That's straight up wrong. Anyway, keep going. I know. But, yeah, I love seeing, like, big sundials and stuff
Starting point is 00:56:59 that have all these, like, names and, like, how we know how to read them and everything. It's cool. I love it. Yeah. The Field Museum actually has a very fun series of exhibits, a lot of which are covered because they have big signs on them that say, because of the Native American Graves Act, this exhibit is covered. Okay.
Starting point is 00:57:21 But the rest of it's pretty cool. I highly recommend. There's one other thing, and then I swear I'm going to let this go. So they trained their own LLM, right, with the... the materials, the stories, whatever, that they could get their hands on while excluding folklore and songs, which, as Scarlett says, is a completely indefensible decision. But these are the people who actually manage somehow to make it to text, right? And we're using them to fill in the stories of the people who didn't.
Starting point is 00:57:58 And I think inevitably there are going to be the usual raft of biases and inaccuracies and missing perspectives, right, that that's going to entail. And I think that's another form of violence, honestly. It's not cool. Yeah, there's so much primacy in our society given to writing and given to like the book, like something we used to talk about a lot on the show was like, you know, the world. worship of the book as a semi-religious object. I think that's, you know, part of the reason LLMs have this semi-religious language around them. I talked about it a bit in the live show. Like, you know, that's why it's making its own eschatology of, like, it's creating
Starting point is 00:58:43 its own apocalypse. Like, this is going to change the world in some major way. Why? Because it can manipulate language. And I think that literacy has something to do with it. If they had read and understood any of the black and post-colonial scholarship that they are citing. They may have already known this, but
Starting point is 00:59:02 oh, I think that's proof that they maybe didn't read this or understand what they did read. Because any, that's like black and post colonial humanities 101. Is like learning about like the sort of like white supremacist's
Starting point is 00:59:18 like primacy of the written historical record and the written word and written language as sources. It's like 101. Well, and Justin's saying, you know, because it can manipulate language. And this is my problem with AI and just in general. It's like, bitch, I can do that.
Starting point is 00:59:36 I can do that without, I can do that without having to dumb it down to a fucking prompt. Yeah, just anything I can do, I can do better. Why hasn't that been written already? Yeah, I'm sentinel coming on. I was going to say, yeah, that's probably by, you know, you know, some early hour, Dorothy and I will have finished texting each other that, that entire thing. This is very likely, I must say. But thanks for putting a call out, by the way, in terms of does anyone want to come talk about this?
Starting point is 01:00:13 Because, yeah, I apparently did. And it was good to flex that muscle a little bit. I don't want to keep you all too late, mostly, because I've had to turn off my AC in order to get this recording. Yeah. No, absolutely. Is there anything, any final thoughts? Right. Any plugs? Thank you so much for coming on. Is there anything you want people to keep an eye out for or do you want people to leave you alone?
Starting point is 01:00:39 Oh, I'll do a little bit of self-promotion. Yeah, I did a talk. Sarah Lambden could not do a talk that she had meant to do at Public Libraries Association in Minneapolis. So she was like, okay, who do I know in the Midwest, who knows anything about library privacy? and for some reason she landed on me. So I took over that talk, digital privacy, sorry, patron privacy with digital vendors.
Starting point is 01:01:05 And PLA really liked it. And apparently I'm taking it on the road virtually for the most part. But keep an eye out. There are going to be several more opportunities to hear that one. One of them actually is through the information schools webinar series that is coming up. And all three of,
Starting point is 01:01:26 the talks and that are going to be privacy-centric. They should be fantastic. So check it out and maybe sign up. Yeah. And as far as my, since it's been a theme, I am in fact back on my bullshit. As soon as I get more word about what's happening with some of the AI tools in library licensed databases here, I will let you all know and talk about the results more publicly. until then, signs points to good outcome about shutting down trash that doesn't need to be in databases in our libraries. And I guess one thing that I want to leave with is something I heard over on Kill James Bond, I think Devin said it, about the current state of the world. And I thought of it as I was preparing for this. And it's the idea that we're all in a line.
Starting point is 01:02:13 And there's a lot of people ahead of me in line right now and behind me in line. And before it gets to me, I'm going to do everything I can to enjoy. sure the safety of the people in line. And that includes being real critical when this stuff shows up in my workplace. And I hope we all are too. So thanks. You're here. Great. All right. Well, thanks so much for coming on and good night.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.