Latent Space: The AI Engineer Podcast - Academia is for Ambition — Alex Zhang, MIT
Episode Date: October 2, 2026Last call for regular tickets for AI Engineer NYC! As an exclusive for Latent Space subscribers, the first 30 of you can take a 30% off code if it helps - for new tickets only, no refunds! See you in ...2 weeks!While we tend to cover industry on the pod, every so often we celebrate a clearly emerging superstar PhD. In 2024 we featured Shunyu Yao, who went on to build Operator at OpenAI and is now Chief AI Scientist of Tencent. In 2025 we featured Jack Morris, who went on to cofound Engram at $600m and is now a leading voice on continual learning. This year we are proud to feature the work of Alex Zhang of MIT.From GPU kernels and KernelBench to Recursive Language Models, Mismanaged Geniuses, and massive multi-agent swarms, Alex Zhang is exploring how much capability we’re leaving on the table by wrapping increasingly powerful models in primitive systems. RLMs took over the timeline early this year:and an RLM based harness was the first to ~solve ARC-AGI-3 before OpenAI’s Astra:and is even today, influencing new research that has more extreme implications than RLMs:We go deep on GPU Mode and AI-written kernels, research taste and why academics should take bets industry labs won’t, GEV and alternatives to the standard autoregressive language model, and the idea of harnesses as compositional generalizers. Alex explains RLMs, context offloading, programmatic subagent calling, Prime Agent, persistent subagents, and why the “language model” of the future may actually be an invisible swarm of agents underneath a simple interface. We also discuss OpenAI’s massive agent experiments, Kimi swarms, open-ended research at Sakana AI, speculative programmatic tool calling, capability overhang, Neuralese, and where Alex thinks the next big research opportunities may lie.We discuss:* Why AI-generated GPU kernels still leave substantial room for human expertise* How one expert insight can potentially replace enormous amounts of brute-force token search* Why PhD students should take research bets that initially look trivial, weird, or pointless* What SWE-bench, RLMs, ReAct, and Quiet-STaR reveal about research taste* GEV and why a language model does not have to mean an autoregressive text-to-text decoder* Why Claude Code, Codex, and Pi are structurally more similar than they look* How harness design can improve compositional generalization across tasks and domains* RLMs: context offloading, code execution, recursive subagents, and shared memory* Prime Agent, continual harnesses, and persistent agent-to-agent communication* Why the model you query in the future may secretly be an entire swarm or scaffold* OpenAI’s 10,000-agent experiment, 130B output tokens, and ~$40M-equivalent problem solving* Why much of an agent swarm may be wasted search — and why convergence is still hard* Kimi versus OpenAI and different approaches to multi-agent systems* Open-endedness, Sakana AI, and finding hidden gems in enormous amounts of generated work* Why current frontier models may already have a large capability overhang* Speculative programmatic tool calling and overlapping tool execution with generation* Whether English, code, or an entirely new “Neuralese” constrains how models reason* AI for science, fast-moving benchmarks, and how Alex chooses what research problems to bet onAlex Zhang* Website: alexzhang13.github.io* X: @a1zhangTimestamps00:00:00 Introduction00:00:49 GPU Mode, KernelBench, and AI-Written Kernels00:07:38 Human Expertise vs. Brute-Force AI Search00:13:20 Research Taste and Taking Big Bets00:19:28 GEV and Rethinking the Language Model00:29:03 Video Game Agents and the Harness Problem00:31:01 Why Claude Code, Codex, and Pi Are So Similar00:36:42 Harnesses as Compositional Generalizers00:44:24 RLMs Explained00:52:01 Prime Agent and Persistent Subagents00:57:41 RLMs in the Wild01:00:30 OpenAI Swarms and the Future of Language Models01:07:26 Open-Endedness and Sakana AI01:15:52 Kimi vs. OpenAI Agent Swarms01:20:06 Capability Overhang and Speculative Tool Calling01:28:19 Neuralese, Future Research, and AI for ScienceTranscriptIntroduction: Alex Zhang, RLMs, and GPU ModeSwyx [00:00:00]: All right, we’re here in the studio with Alex Zhang, I guess most famously of, RLMs, but you have a few other affiliations. Welcome to the show.Alex Zhang [00:00:12]: Yeah. Thank you for having me.Swyx [00:00:13]: Yeah. I guess GPU Mode as well?Alex Zhang [00:00:15]: Yes, GPU Mode as well.Swyx [00:00:16]: You were shepherded in by Mark Saroufim. Not everyone gets that kind of welcome.Alex Zhang [00:00:19]: Yep. Yeah. Yeah. I’m very close to all the people in GPU Mode, so yeahSwyx [00:00:23]: YeahAlex Zhang [00:00:23]: We often end up working together in various capacities, like even beyond just GPU Mode itself, so.Swyx [00:00:29]: Yeah. Can we explain, so people who are not that close Don’t know about this. It’s just a-- it’s a Discord. It used to be focused on, I guess, CUDA Mode, and then generalized a little bit. it was started by Mark.Alex Zhang [00:00:41]: Yep.Swyx [00:00:42]: It was basically like. To me, it’s like the hiring pipeline of the PyTorch team.Swyx [00:00:45]: And then you left PyTorch.Alex Zhang [00:00:47]: Yep.Alex Zhang [00:00:49]: Yeah. So it used to be, I think, it started actually around when I was in college, in like 2023. I think it was started by Mark, Andreas, and Jeremy Howard. The original premise was just like, it was a GPU-- or it was a Discord dedicated to learning how to write GPU kernels, and they had, like, lectures. That was basically the extent of it. and I got interested in it because I was writing GPU kernels. It was actually out of, like, pure chance. I was interning at Snapchat at the time, andFrom CUDA Mode to GPU ModeSwyx [00:01:22]: Rexis.Alex Zhang [00:01:23]: Yeah, I was very bored with Rexis. So, they had a project where, like, they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper.Swyx [00:01:35]: Yes, we’ve covered it on Paper Club.Alex Zhang [00:01:36]: Yes, yeah. So I was interested in whether or not you could write specialized kernels for it at Snapchat. it didn’t. Nothing really came of it, but I joined GPU Mode. At the time, it was called CUDA Mode, I think for, like, legal reasons or something, they changed the name. But I met Mark, I met Matei, I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn, which was now what you see as the leaderboard today. But the general idea was like. I think all of us had this, like, intuition that GPU programming is, like, very similar to if you guys have done, like, competitive programming. It’s a, it’s a not. I don’t mean to say, like, they’re transferable skills.Popcorn, KernelBench, and Automating GPU KernelsSwyx [00:02:17]: You have constraints. You code golf a little bit.Alex Zhang [00:02:19]: Yep.Swyx [00:02:19]: Yeah.Alex Zhang [00:02:19]: Yeah. And there’s like. There’s actually a surprisingly small space of optimizations that people do. and there’s actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that, like, if you had enough data, like, in the same way that Codeforces, there’s millions of problems. If you could do this with GPU code, like, you could scale and automate kind of GPU kernel development, which for researchers is a huge deal. ‘Cause I think one of the bigger bottlenecks. Like, if you look at, like, Mamba, for example, like, they release the paper with kernels because otherwise, like, you can’t really use it in any meaningful way? And not everyone has, like, a Tri Dao on their team. So we’re very interested in this. KernelBench kind of spawned from that too, of like, can we get LLMs to automate, GPU kernel code? And I think that was like a. It was a very fun time. It was like between college and my PhD, and yeah, I had a really pleasant time doing stuff with GPU Mode. Now I kind of just help with the lectures sometimes. I’m not as involved, and I think in general, like, we don’t have as many competitions as we used to. But, yeah, I still keep in touch a lot with everyone there.Swyx [00:03:27]: Is there a friendly rivalry? Because I think the previous community that used to do this was like MLSys, MLPerfAlex Zhang [00:03:32]: YeahSwyx [00:03:32]: Kind of thing. Is there a friendly rivalry? Is this like just new generation MLPerf, or what’s going on?Alex Zhang [00:03:38]: The nice thing about GPU Mode is that it is also a community in the sense that, like, a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that. And, like, the competitions are, like, somewhat, not secondary, but, like, you can participate in them to learn. I think with a lot of, like, MLSys, MLPerf kind of benchmarks, like, for the most part, like, only serious labs and companies participate. Like, seriously in them, at least. That was my understanding of it. I could be wrong. But I think also beyond GPU Mode now, one thing that has been really exciting is there’s a lot more websites and, like, people that work on hosting competitions. Like, I think there’s this.Alex Zhang [00:04:21]: I think there’s this website called, like, LeetGPU or something, and it’s like leet code for GPU problems.Swyx [00:04:26]: Wow.Alex Zhang [00:04:26]: There’s, like, other ones that I’ve. Like, we’ve, we’ve seen. Like, there’s many that have kind of spawned and, like, talked on GPU Mode, and like, it’s very exciting in the sense that I think GPU programming used to be super niche, like when I was interested in it. And the only reason I got interested in it was Tri Dao gave a talk at Princeton because he was, applying for faculty. he is faculty there now, but I listened to his talk on FlashAttention in like 2023, and I was like, “Wow, this is like the coolest thing ever.”Alex Zhang [00:04:56]: And I was like, “This is like. This is what everyone should be working on.” I guess, like, vLLM and stuff had come out too, and it was like, “Oh, we should be writing kernels.” But now it’s like, everyone writes kernels. Like, everyone. It’s, it’s. I think it’s actually almost saturated in some sense, as a field.AI-Written Kernels and the Verification GapVibhu [00:05:11]: Any interesting takes for people that wanna get into it? So I think one of the biggest news is GPT-5.6 Wrote more efficient kernels, so Terra and, Luna could be 80% cheaper.Alex Zhang [00:05:24]: Yeah.Vibhu [00:05:25]: And then we’ve seen other competitions where people are, like, setting records, and they’re like, “We’re doing some auto research loop,” and these are people that don’t have a backgroundAlex Zhang [00:05:34]: YepVibhu [00:05:34]: In any kernel writing, right?Alex Zhang [00:05:36]: Yeah. So even on the GPU Mode leaderboard, if you look at, like, a lot of the recent problems, almost all the solutions are AI generated. However, you’ll notice on the leaderboard. So there’s this guy named Gauners who’s, like, a very, like, regular member of GPU Mode. We’ve always known for a long time that he’s, like, a super cracked, like, GPU kernel writer. One thing we discovered on this leaderboard is, like, almost ev-- Like, he also used AI to help him with these solutions, but for the most part, like, he helped prompt and move it in certain directions. we found that, like, his kernel was, like, basically the only one in, like, the top 10 that was actually stable in, like, actual, like. end-to-end systems. And it does bring into question, like, it’s not. Like, GPU kernels have a verification problem. Like, we’ve kind of known this. It’s been a problem since KernelBench was released. Like, there’s a lot of reward hacking that goes on. But you al-- Yeah, you also notice, like, the lines of code is a lot smaller, butVibhu [00:06:33]: Yeah, I was gonna ask, is that noticeable, or is it justAlex Zhang [00:06:36]: Yeah, no, it’s, it’sVibhu [00:06:36]: OkayAlex Zhang [00:06:36]: It’s definitely, like, very important, and I think, like, it’s, it’s really interesting that still there’s a lot of alpha in being good at writing GPU kernels.Vibhu [00:06:44]: Okay, so there is a gap from verifyingAlex Zhang [00:06:45]: There definitely is, yeah. I think, like. And this applies to a lot of AI systems as well. Like, I think, even with the most recent, like, math proofs and stuff, like, it doesn’t necessarily mean mathematicians are obsolete. these companies still hire mathematicians, like, to do, whether it be, like, data labeling work or even just, like, steering the models to solve problems. Like, there is still a lot of alpha in being, like, knowledgeable in these things, so.Swyx [00:07:12]: Is it just knowledge, or is it also there is just more planning, and is there, an emergent style of planning that works better?Alex Zhang [00:07:22]: I think it’s, it’s a mix of. Maybe this is what you mean, like intuition forSwyx [00:07:27]: Something like thatAlex Zhang [00:07:28]: How to solve the problems.Swyx [00:07:29]: Like, for example, I always diagram my code.Alex Zhang [00:07:31]: Yeah.Swyx [00:07:31]: Right?Alex Zhang [00:07:31]: Yeah.Swyx [00:07:31]: And then, like, if there’s a part of the diagram I don’t understand, I work until I understand it. Otherwise, I, it’s not allowed.Alex Zhang [00:07:37]: Yeah.Swyx [00:07:38]: Yeah.Alex Zhang [00:07:38]: So I think it’s, like, it’s a mix of those things of, like, the people who work. Like, the people who know how to look at these problems and how to solve them, like, also know how to use AI to do them. Because, like, you’re acting as a very strong verifier. Like, if you are knowled- or if what to do and you are also. Like, I think the thing that we’ve kind of discovered with all these agent swarms and things like this is, like, when you throw enough compute at a problem, you, like, can sufficiently explore solutions to that problem. But oftentimes, like, maybe you can burn, like, 100 billion or a trillion tokens on something, but if you bring in someone who knows something about the problem, they can uncover something for the model that would, like, erase that one trillion token spend. it’s not, it’s not super clear, like, what exactly the trends are here. But I think, like, there are so many problems in the wild still right now that we want to solve, and, like, we can’t afford to just always, throw as much compute as possible at it. Like, there is still an efficiency aspect of all of these things that is super important.Speed-of-Light Limits, Memory, and MegakernelsSwyx [00:08:41]: Is there, like, a theoretical right answer that you just calculate based on physics, and then you just get close to the physics limit?Alex Zhang [00:08:49]: Yes. So for GPU kernels, you can compute. It’s actually not that easy to compute sometimes, like, depending on how complex the problem is. Like, for matrix multiplication, it’s very easy to compute, this, like, speed-of-light kind of, estimate of what the fastest kernel can be. And, like, this is also assuming, like, maybe all of your, all your data starts on the CPU, or maybe it starts in DRAM, on the GPU, et cetera. Like, this changes these numbers slightly, butSwyx [00:09:18]: The transfers and all these things, yeah.Alex Zhang [00:09:20]: Yeah. But I will say, like, it’s not clear, though, like, in a lot of cases if it’s even possible to hit this theoretical number, if that makes sense. Like, this is assuming, like, perfect overlapping and transfer of data, and, like, there’s maybe some bottleneck that you can’t get around. But often, the kernels are not even close. Like, that we write are not nearly close enough to this number to be, like, meaningful at all.Swyx [00:09:43]: Yeah. And is it speed that matters? Do you also care about, obviously memory, whichAlex Zhang [00:09:48]: MmSwyx [00:09:48]: Feeds into speed? Do you care about power consumption? So one of my, one of our top pods of the year was Geoff Dean, who was like, “Actually, I just tracked the microjoules or, like, the nanojoules, picojoules.”Alex Zhang [00:09:59]: Yeah, it’s often picojoules today.Swyx [00:10:01]: Picojoules.Alex Zhang [00:10:01]: Yeah.Swyx [00:10:01]: Do you care about that?Alex Zhang [00:10:03]: So I don’t.Alex Zhang [00:10:04]: Yeah. I guess maybe I’m not, I’m not asSwyx [00:10:06]: But everything here is speed, right? LikeAlex Zhang [00:10:07]: Everything here is speedSwyx [00:10:08]: Nobody’s counting picojoules.Alex Zhang [00:10:09]: But I-- There’s a caveat here, which is, I think, like, there is speed in the context of a single kernel, and there is speed in the context of a larger problem, like maybe the N10 model. Because, like, one thing to consider, and this is why it’s important to talk about what speed-of-light is referring to, because in these cases for the kernels, like, we always start with everything in, like HBM, for example, right? But you can imagine that, like, an end-to-end like an end-to-end model, what you might wanna do between two layers is, like, you might sacrifice the speed of the first operation to keep things in the cache for the second operation. And, like, these are things that, like, you can’t really get out of, in isolation with, like, these kinds of kernels. And, people call this, like, the fusion or, like, theSwyx [00:11:01]: MegakernelAlex Zhang [00:11:02]: Kernel fusion problem. Yeah, or, like, megakernel stuff. And it generally only applies, like, when you are, like, memory-bound in most cases. But this is something that, like, also there is this question of, like, as these models get better, like, should we just be generating like, megakernels? Is that, like, what we want?Vibhu [00:11:20]: What’s, what’s your take?Alex Zhang [00:11:21]: I think that this is really difficult because you need the data to do this. And I think, like, I have yet to see an example in the wild of, like, we bootstrap the ability to solve a very difficult class of problems without any examples. and I think, like, the other reason why I think maybe this isn’t that interesting is that at the level of an individual kernel, A, like, they’re not that, they’re not as complex, but B, you’re somewhat confident that there’s not as much structure in a single kernel. But, like, in a megakernel, like, I would be more inclined to believe that, like, a compiler would be better here. Like, some compiler over, like, higher level- Ops makes sense, because in general, like actually, I think mega kernels are very, like the pieces are very composable of like the individual kernels. There’s some areas where you might wanna do like weird fusions and everything, but in general, I think these are cases that like a compiler can probably handle. And there is a company that’s working on this from what I understand that has given some talks on GPU mode as well.Swyx [00:12:28]: Yeah, I wanna basically cluster all the GPU mode discussions here because obviously there’s other partsAlex Zhang [00:12:32]: Right.Swyx [00:12:32]: That we need to move on to.Vibhu [00:12:33]: I think there is something to plug. You guys do host a lot of really good lectures. They’re all on YouTube. People can follow. And youAlex Zhang [00:12:39]: YesVibhu [00:12:39]: Lead quite a bit of it. You’re still quite involved.Alex Zhang [00:12:41]: I used to. sometimes I still do. I think they’re mostly Mark. Mark is the one who usually does them. Matei does sometimes as well, but, yeah, I highly recommend them. They are extremely good resources. Like, I think it’s kind of crazy how much people share on there, so yeah.Swyx [00:12:59]: ‘Cause like if you’re there, like you’re very, like you’re exactly the right audience?Alex Zhang [00:13:03]: Yes, exactly.Swyx [00:13:03]: Like this isn’t gonna reach the mainstream.Alex Zhang [00:13:04]: And there’s a lot of like introductory material as well, that we’ve put on, that I think is useful for people.Benchmarks, Princeton, and Research TasteSwyx [00:13:10]: KernelBench was kind of influential. I just wanna see like, that was last year.Alex Zhang [00:13:13]: Yep.Swyx [00:13:14]: What other ongoing work do you wanna shout out that people should pay attention to? ‘Cause obviously you’re involved in thisAlex Zhang [00:13:20]: YepSwyx [00:13:20]: Field.Alex Zhang [00:13:20]: Yeah. I will give maybe the background story of like I am actually involved in a lot of benchmarks, or I used to be, maybe prior to my PhD. It started because I was at Princeton. I worked with the SWE-bench team there.Swyx [00:13:33]: John, CarlosAlex Zhang [00:13:33]: John, CarlosSwyx [00:13:34]: OfirAlex Zhang [00:13:34]: And Ofir. They’re all great. LikeSwyx [00:13:36]: KarthikAlex Zhang [00:13:36]: I love them. Yeah.Swyx [00:13:37]: There’s basically this, like I think people don’t understand how much benchmarks come from the same groupAlex Zhang [00:13:42]: YeahSwyx [00:13:42]: At Princeton.Alex Zhang [00:13:44]: It is crazy.Swyx [00:13:45]: Do Xun Yu?Alex Zhang [00:13:45]: Yes. Yeah.Swyx [00:13:46]: We had him on a pod before. Now he’s like running Tencent.Alex Zhang [00:13:48]: Yeah, now he’s like, he’s like a superstar. when I met him, so he was advising my friend Michael Tang, who is now at Anthropic, but they worked together a lot. We were like the two undergrads in Karthik’s lab. I-- And then some others joined later as well. But yeah, Xun Yu is great. I did not know he was like such a superstar until like later on, like after I left, butSwyx [00:14:12]: Yeah. like, okay, so there are very few PhD students. Like yours is like the next one. Like once a year, we feature someone like who is like basically entire PhD, has been like on target.Swyx [00:14:25]: There’s not that many of them. Xun Yu was like clearly one of them. And, Jack Morris is another one. And like, we talked, before the show, we talked about research taste.Swyx [00:14:33]: Right? Like somehow some grad students just have a very blessed career where like, yeah, mostly like, yep, this is like going to stick around, relevant, everyone should know this.Alex Zhang [00:14:42]: Yeah.Swyx [00:14:42]: And then others, just nothing.Alex Zhang [00:14:44]: I think this is also true of like even people within like industry labs as well. I think it’s just like grad students are a lot more visible. So you just see, like you see, like, there are some people whoSwyx [00:14:56]: Yeah, you can publishAlex Zhang [00:14:57]: Who really like get lucky and like, or it’s, it’s a mix of being lucky and also being very smart and things like that. I think like with research taste as well, like I think it gets developed through opportunities, at least in my case. Like I got-- I was very fortunate to have like taken the path that I took, like working at Princeton and then like later, like finding my like Omar at MIT. Like he’s a fantastic advisor. I will say, though, I find that the most successful research from grad students or like in academia comes when people care about problems that maybe like most people in industry are not looking at. I think this is the issue that like a lot of grad students work on things that benefit, like that look good to an industry lab. Like for example, they’ll work on some, like some benchmark that’s really popular now. I think benchmarks in 2023 were a very different story than benchmarks now. There are a lot of people that work on like harnesses and like meta-harnesses and like specific harnesses for XYZ task. And when you really think about it, the reason someone would work on this is like maybe there’s like a clear goal shaped around the models that we have today of like, this is what I want to see. But like, I’ll give, I’ll give the like RLM, like the recursive language model paper as an example, because I think like it’s a super simple idea. I think when it came out as well, like there were a lot of people that were like, when they see something like that, they’re like, “What is even the purpose of this?”Swyx [00:16:27]: Or like too cool.Alex Zhang [00:16:28]: Yeah. Like why, like what? This is just subagents or something, right?Alex Zhang [00:16:32]: And I think it’s like when you get a reaction like that, it’s almost like a good sign in the sense that like it’s clear that people aren’t thinking about what the purpose of this is. And I will give another example of like SWE-bench. When SWE-bench came out, Ofir loves to tell this story. When it came out, like nobody cared. Like everybody was like, “This is an impossible task. Like why would we ever even consider this as a benchmark?” And it wasn’t until Devin came out that everyone was like, “Whoa, like this is something we wanna hill climb.” And I think this is, this rings true for. You tend to see that a lot of ideas. I think like the. My favorite, I guess, example of this is Eric Seligman’s work, with like STaR and like Quiet-STaR. Like I think when you read the paper, at least when I first read the paper, I was like, “Is this not like an obvious idea?” Or maybe not. I don’t know. I was like, “Oh, this seems really simple.” Or like chain of thought, and the same thing. Or like Xun Yu’s react. It’s like, okay, like, yeah, sure. But then like when you really think about it’s like why. What is the value of the paper? And I think it comes from like, it tells a bit of a story as to like what you want the field to look like. And that is something that it’s very hard to do this in academia because if you look at all these papers, Quiet-STaR, ReAct, RLMs, SWE-bench, none of these papers are. It’s not like a GPT-6 Astro release? It’s not like everyone’s like, “Oh my gosh, like I’m gonna use this now and this is the best thing in the world.” Like academia just can’t afford to do this, at least right now. I. There’s a whole slew of reasons why I think that should change, but I think it’s like. If you don’t have. As a PhD student, I think you’re in such a unique position where you can work on literally whatever you want for the most part. If you’re not taking advantage of that and working on things that, like, nobody cares about or, like, people see as some trivial thing, like, “Oh, I thought about this, but, like, I don’t use it,” I just think, like, in the end, the research is just never gonna be that interesting because you kind of need to take big bets if you’re gonna be in academia. Because otherwise, I think, like, just go to an industry lab. Like, they have tons of resources, tons of talent. Why constrain yourself in an area where you don’t have a lot of resources and, like, there’s not even that many people around? And I think it’s just. it literally just comes down to, like, big bets. like, you just have to take big bets, and, like, a lot of them will fail? Like, that’s just. it’s, it’s natural. But I thinkAlex Zhang [00:18:55]: That is, as a PhD student, like, that’s the biggest advantage you have over any single person at another lab because you don’t have to deal with bureaucracy and all these other things.Swyx [00:19:07]: Fair enough.Alex Zhang [00:19:07]: Yeah.Swyx [00:19:08]: I ask a lot of people this question, and usually they hand-wave away. So I think. I appreciate that you’re actually giving a thoughtful response on Like, no, like, this is your unfair advantage because everything else is biased against you, basically.Jev and Breaking the Autoregressive Decoder ParadigmAlex Zhang [00:19:20]: Yeah, exactly. And so, like, it’s honestly. I will bring up Jev as an example because it’s, it’s not an academicSwyx [00:19:26]: Wow, okay.Alex Zhang [00:19:27]: It’s not an academic project.Swyx [00:19:28]: Yes.Alex Zhang [00:19:28]: I want to bring this up because this also happened with RLMs and, it happens with many other works. Like, things get overhyped, right? To an extent, like, something gets overhyped and then people are like, “Why is this overhyped?” Like, “This is trivial. This is stupid.” And I saw the same thing with Jev because I think the release was like. there is this whole thing about, like, academics, or they’re not an academic group, but, like, people have to do branding and they have to, like, kind of market their research. And so, like, I understand, but I think there was a lot of discourse about Jev just being, like, something we’ve known for years. And I think it’s kind of missing the point of, like, why is such a system so interesting? It’s why is it not just some stupid NLP classifier that, like, we’ve, we’ve been doing, back in our intro ML classes or something? I think what’s really interesting about Jev is that it kind of opens up this question of, are language models correct? Like, in the form that they’re in, can we consider a different design space other than text-to-text? Because what they’re doing is they’re basically saying like, “I will take advantage of this language model backbone. Like, I know it captures a lot of information about language, but I’m going to change the output space of the model to give you a trade-off, which is I will do very fast inference over.” Like, if you have some prior about this problem, like, let’s say I only need to make a binary classification. Am I gonna ask my language model to do this and pay, like, a 400X cost? Like, no. That’s-- it’s, like, silly, right? and I think for the longest time, because the labs are the only places that control, you’re never gonna use something other than, like, GPT-4 or GPT-6 or Fable because they’re the best models. But because of that, like, people have gotten kind of accustomed to this idea that a language model is just a autoregressive decoder. Like, we have accepted this. And I think when RLMs came out, it was the same thing. Like, one of the comments, like a very frequent criticism I got was like, “This is not a language model.” Or like, “When I look at this, like, I thought it was a new architecture, but it’s actually not.” And my response to that is like, “Well, a language model is just modeling language. It doesn’t have to be this transformer decoder,”? and Jev is really interesting in that, like, we now have a newAlex Zhang [00:21:49]: Thing to tune, which is like, what is the output space and how does this affect inference latency? and I think we can actually start asking this about various parts of the language model itself. we are seeing this too with, like, loop transformers. It’s a similar idea of a lot of the attention around it was like, “This is a silly idea.” Like, “Why? Who cares about this?” But it’s like, it is a simple idea, but it’s actually. it opens up a whole new set of questions that I think, like, especially if you’re a PhD student, these are the things that you wanna answer. Because I think it’s like we don’t know. For Jev, for example, we don’t know how far we can take this. for loop transformers, we also don’t know how far we can take this. What if you loop only a subset of the model? what if you route to, like, only. like you have some router to different parts of the model? Like, can you mimic what you do in a harness inside of the model? What can you bridge between a harness choice and a model architecture choice? These are all questions I think that open up with works like this, and that’s, like, really exciting. Jev in particular, when I saw it, I was like, “This is actually really useful for RLMs.” Like, I think it’s, it’s. it makes sense ‘cause the biggest bottleneck in RLMs or swarms or systems like these is they’re slow. When you do multiple language model calls all the time, you’re not distributing your compute correctly because, like, maybe there’s something trivial that you just want a simple model to do, but you can’t do it because your language model is just this bulky thing? So I’m very excited. I think we will start to see new types of models emerge beyond just the bog-standard frontier model, and that is like. there’s so many things that you can do with these, like, new trade-offs.Swyx [00:23:34]: I’ll also shout out Thinky with their interaction models.Alex Zhang [00:23:36]: Yes. Yeah. Another great example.Swyx [00:23:38]: Yeah. So, like, basically try to break the paradigm from sequence to sequence, decoder only, and just literally do anything else.Alex Zhang [00:23:46]: Yeah.Vibhu [00:23:46]: I think there’s, there’s a level of if you’re trying to compete, you’re not gonna compete with a Frontier lab doing an autoregressiveAlex Zhang [00:23:54]: No.Vibhu [00:23:54]: Decoder on. Like, the amount of compute scaling resources they Even Thinky will not. okay, there may be one of the handful that can, but you’re not really gonna do much in that at least.Alex Zhang [00:24:06]: I don’t know too much about Thinky, or I don’t wanna say anything either, but it’s like if their strategy is just to replicate OpenAI or Anthropic, like that’s a horrible strategy.Alex Zhang [00:24:15]: Because, well, because, like, they just don’t have. Like, you kinda just have to think of it in terms of, like, what advantage do you have? And if you’re going to use the same setup. I’m sure they’re not, but it’s like if you’re going to do the same setup, like you’re basically competing on the things you can control, which is data and compute, and obviously they cannot compete with the Frontier Labs on that. So yeah, it makes sense that, like, if you are a neo lab, like. Actually, I don’t know if you would consider them to be a neo lab, but I guess, likeVibhu [00:24:40]: Yeah. That’s why they’re, they’re in there.Alex Zhang [00:24:42]: I guess they’re kind of a weird one, yeah.Vibhu [00:24:43]: They’re in their list. They shipped Inkling. Like, they can’t.Alex Zhang [00:24:46]: Anything other than OpenAI or Anthropic, maybe like Meta and GDM, like you just, you gotta do something else? Like, it just. It’s the sad reality, but I think. I actually think it’s a good thing. I’m very happy that, like, scale works and these companies will just keep doing it because it opens up, like, potentially new players, like if they uncover something really interesting. Because I sort of have my doubts that this is, like, seriously going on at Frontier Labs, ‘cause it’s like why would you do that? like, why would you take the risk of allocating a large amount of compute to new bets when the old bet already works? SoVibhu [00:25:24]: And I think that’s what spins off a lot of neo labs, right?Alex Zhang [00:25:27]: Yeah.Vibhu [00:25:27]: You have a side bet and you don’t get compute, and you’re likeAlex Zhang [00:25:30]: YeahVibhu [00:25:30]: “Okay, I’ll go, I’ll go do that.”Alex Zhang [00:25:31]: Exactly.Vibhu [00:25:31]: And, your example of the potential upside is something like Jev, which is X hundred times cheaper, comes out, and maybe it is language model.Calibration, Fast Classification, and New Model Trade-offsAlex Zhang [00:25:41]: Yeah.Vibhu [00:25:41]: In this case, it’s just different.Alex Zhang [00:25:42]: Yeah.Swyx [00:25:43]: Yeah.Swyx [00:25:44]: So no speculation on what Jev actually is?Alex Zhang [00:25:46]: I guess I have some guesses for what it might be. I have seen some people say like, “Oh, it’s like a diffusion thing.” I guess thatSwyx [00:25:56]: Which is the parallel decode, right?Alex Zhang [00:25:58]: Yeah, parallel decode. Honestly, I think regardless of what it actually is, ‘cause I think you can. I’ve seen some, like, open-source replications of it. What is really exciting to me about what they did is I’m not entirely sure what their optimization objective was and how they trained it. And I think, like, this is a thing for RLMs that we’ve also been thinking about, which is like, okay, like RLMs are a very simple idea. If I come out with this paper, like anyone can use it now. But what distinguishes The actual value of an RLM is whether or not you can train it properly, and whether or not maybe you can mold some architecture around the system to make it really good. And that’s something that, like, I’m actively working on, I guess. But I think for them, like, they figured out a way to train the system, which is completely non-trivial. Like, I actually don’t really know how they did it. And I’ve seen some comparisons online of, some people are claiming they used Qwen, or they post-trained on top of Qwen, but every open-source Qwen that you use is gonna be worse, ‘cause whatever they did to train it clearly works very well. And so that’s, that’s very exciting.Swyx [00:27:04]: There’s one element of calibration Which, is a rare topic that I don’t think people even knew about or understood. We covered it with, our conversation with Clementine Foley of Hugging Face, and she used to run the evals, at Hugging Face, which is basically the idea that, models are attuned to give you the most likely next token. But, they’re gonna lie to you when you ask them, “How confident are you?” Because they’re just gonna give you the most likely next answer instead of, like, actually, like, no, let’s calibrate. Like, I am actually fifty percent sure, or I am twenty percent sure, and, like, let’s try to calibrate that. I would say, like, if anything, I think that actually that’s pretty easy to generate synthetic data around Because you can sort of see the truth and then synthetically generate a bunch of answers, have Jev classify it, and then compare with ground truth.Alex Zhang [00:27:52]: Oh, I see.Swyx [00:27:52]: That would be my reverse engineering of this.Alex Zhang [00:27:54]: Yeah.Swyx [00:27:54]: I’ve actually. I think calibration is probably the under. Like, people are just using it as a very fast classifier But they’re actually not even using the probability or calibration estimates.Alex Zhang [00:28:04]: Yeah.Vibhu [00:28:04]: I think it’s also still just misunderstood to reiterate. When you ask a model, “How confident are you?” it will spew out what, forty-three percent. the big delta is this is a grounded classification, right?Alex Zhang [00:28:16]: Yeah. Yeah, I’m, I’m very excited to see what people do with this model. is it gonna solve everything? Like, no, of course not. But I think it solves a class of problems that we traditionally struggled with, which is low-latency things. So I love the examples with games. That’s actually like. ISwyx [00:28:34]: Yeah, the Doom exampleAlex Zhang [00:28:35]: YeahSwyx [00:28:35]: Was very good.Alex Zhang [00:28:35]: I have a benchmark on language models playing video games. I’ve always been fascinated by whether you can have an intelligent system play new games and things of this nature. And so I think it’s really cool that they have sort of a unique way to do this, to capture language and understanding in, like, a fast, a very fast model.Language Models Playing Video GamesVibhu [00:28:59]: Oh, while you’re on the topic, anything you wanna point out for video games?Alex Zhang [00:29:02]: Oh, yeah.Vibhu [00:29:03]: This is. you did do a benchmark on any project, right?Alex Zhang [00:29:05]: So, yeah. I guess these numbers are very outdatedVibhu [00:29:08]: Yeah.Alex Zhang [00:29:08]: Because a lot of the models are very different. And I’ve seen actually people run. There are some folks out there that are running newer models on these games, which is really cool. I guess the general premise of this benchmark was we just want to see if vision language models are good enough at just, like, plugging into games with the latency constraint included. ‘Cause this actually. I came out with this right after Claude Plays Pokémon came out.Alex Zhang [00:29:33]: So this was, like, two years ago, which I guess is, like, ancient now. But I think what’s really cool about this suite of tasks, it’s very diverse in terms of what games they are. And also, I think most of the games are games that people know or, like, have seen before. I saw, yeah, Jeff playing Doom. I will say I don’t think. I think they were just playing, like, really simple levels and stuff. But honestly, like, most models still can’t really do. Or I don’t actually think any models can solve these games very meaningfully. Like, there are some games that they can. I think I’ve seen Astra be able to solve the Kirby game. And we also, for this benchmark, we intentionally designed a really minimal harness. And I will get to this point about harnesses because I think there’s a whole conversation to be had about, like, what is the value of a harness? Like, what is even the purpose of a harness for a model? But in general, like, I think it’s. yeah, I hope to see very quickly or very soon, like, all of these games beaten by newer models.Vibhu [00:30:35]: Yeah, it’s interesting. Like, the old Cloud Place Pokemon, they, like, read state from RAM and saw what tiles are walkable and whatnot. We did a podcast with them A long time ago.Alex Zhang [00:30:45]: Gotcha.Vibhu [00:30:46]: Yeah. Just fun.Swyx [00:30:46]: Yeah, and it’s similar. Like, Jeff doesn’t have visionAlex Zhang [00:30:48]: YesSwyx [00:30:48]: So you have to kind of feed in,Harnesses as Compositional GeneralizersAlex Zhang [00:30:50]: YeahSwyx [00:30:50]: These, like, game state and all these things. let’s go right into the harness stuffAlex Zhang [00:30:54]: AwesomeSwyx [00:30:54]: Because you brought it up. Language model harnesses are compositional generalizers.Alex Zhang [00:30:59]: Yes.Vibhu [00:31:00]: You struggled to read that one.Alex Zhang [00:31:01]: Explain. Yes. Okay. So I have been a little unsatisfied maybe with how people think about harnesses, because people compare like, “Oh, like, I love Claude Code, I love Codex, I love Pi.” Like, “No, I love Oh My Pi, I love Prime Agent.” To be honest, I think all of them are the same. Most of the design decisions or, like, the design choices around these harnesses are the same. Maybe Prime Agent is a little bit different because it’s, like, inherently an RLM. But in general, like, I think we can be a lot more creative with harnesses. And what by that is if we think about this from the perspective of what exactly is the harness doing for the model? Well, basically, when you’re trying to solve a problem and you want to use a language model to solve it, like a very difficult task, one thing that we have discovered is that next token prediction is a really awkward form to do a lot of these tasks. So for example, take SWE-bench. When you’re navigating a code base, like, are you going to be able to figure out how to do all of this with a single language model call? Like, you just say, “Solve code,” or like, “Solve my query over this code base.” No. So we rely on a harness to help you do these kinds of things. And I think what is interesting is, like, a harness is a very opinionated program over how you want a language model to be form fit over a problem. And I bring this up because when we think about the, what a harness is doing, we should really think about, like, what choices in the harness let the language model solve this task? And can I actually just have a language model that just does this? Because a harness, if you think about it, now that loop transformers are a thing, I think what’s really interesting about it is you can model a looped transformer in some ways, like, with a harness as well, right? You’re just looping over the model. Now you can say like, “Oh, I’m not decoding,” so it’s, like, a little bit different. But in general, we, for whatever reason, have stuck with the same model architecture choice forever. And I. And there’s many arguments for why, but clearly, like, we are now training models around harness rollouts. And so there is this very awkward way of doing training over harnesses, which is that we train a language model to act within a harness, but, like, now it’s like a really long, maybe, like, multiple agent rollout that we’re doing. And there’s, like, really awkward, hacky ways of doing this. So what this blog talks about is like, well, one way you can think about what is going on here is if the harness is basically helping the model solve a particular task, can different harness design choices actually do something a little bit more meaningful beyond just, “Here are some tool calls that will help you. Here is a way to grep through your code base.” And so this actually. The idea for this blog came with the RLM idea as well. We just didn’t package it that way. And I think this is actually true. There are many other ideas around RLMs that, like, we will be coming out with, but were all there from the beginning. these are all design decisions around. I think with what is. What I like about the RLM is that there were many iterations and versions of different abstractions that I was interested in doing, and ultimately the RLM made the most sense. But there’s a lot of reasons that aren’t public as to why that’s the case. you will see. So in this example, one of the things that we see with an RLM is that if you sufficiently offload context and write. ask the model to write code over that context, you get this really weird but useful property, which is that when the model recognizes during training how to solve a task, it turns out that the solution to many tasks is very similar across, like, tasks where you don’t even. Like, it’s not even that clear to you that the solutions are similar. So in this example, we have, like, a retrieval task and we have, like, an aggregation task, and they’re very different query. Like, the domain is just completely different. And when you train a regular language model over these two tasks, the trajectories look very different. And so what you’re relying when you. Like, let’s say you use Pi or Claude Code or something, which is not in this blog, but we do have these results. You’ll find that, like, these harnesses distinguish too much between these problems, even though the solutions are the same. And so one thing that we find when training RLMs is that, like. When it sees these problems, it’s the same. And the reason it sees these problems as the same is the sub-agent sees different problems, but the sub-agent is solving an easier sub-task, and so you’re confident it’s smart enough to do it. But for the base overall strategy, they end up looking the same. And so when you train on the left task, for example, it can immediately solve the right task. And so if you go down to, like, the plots that we have, one thing you’ll kind of observe is that the. as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks, because the strategy is basically the same. You’re just modifying, like, a length variable. And this actually also holds for tasks that are different, and they’re, it’s not even-- they’re not different across length. They’re completely different tasks, math tasks versus writing tasks. But the solution, the, like, meta high-level solution is the same. And so when you train the RLM on one of them, it generalizes this behavior to the second one. And there’s no magic here. I guess maybe that’s the thing that I wanna kind of stress. LikeVibhu [00:36:38]: How would you kind of verbalize what they are learning? So I think in here you say you train it on short tasks, they generalize to stuff 8–30x longer.Alex Zhang [00:36:47]: Yeah.Vibhu [00:36:48]: They are learning how to solve these type of problems, or what’s the, what’s the core thing they’re actually learning?Alex Zhang [00:36:52]: Yeah. They’re learning how to solve these types of problems at a certain length. And it turns out that when you take the strategy that they learned, it is directly transferable to the longer length. Like, they’reVibhu [00:37:05]: Yeah.Alex Zhang [00:37:05]: Effectively the same program. And this is something that is really exciting because what this kind of implies is that when you have a corpus of data or environments that you train your model on, the hope is that, like, A, you can train on less environments and generalize to more than what existing models can do through or harnesses through existing kind of, like, naive training. But B, also, you still wanna use all the data you have. So when you train on these tasks, like, hopefully it generalizes to a wider class of problems. And why this is also even more exciting, at least in the context of RLMs or any recursively calling system, is this argument holds inductively. So, like, I’ll give you an example, because I said, I claimed that competitive programming and GPU optimization use very similar skill sets. The model can. the harness potentially, you might have to nudge it a certain way, but it can learn that, like, “Okay, how I’m gonna go about solving this GPU programming task is very similar to what I learned for competitive programming. So I’m gonna list out a set of solutions. I’ll, I’ll, like, spawn subagents to list out promising solutions, and then I’ll, like, write this loop to go through and check these solutions, maybe evolve them, and, like, evolve them against a verifier.” And between these two tasks, this looks the same. But what the subagents are doing are maybe, like, unique and something, like, different. But even what the subagents are solving might actually also be of the same form, right? Because it’s like a, it’s a recursive argument. And so what I’m trying to get at with this whole blog post is just that, like, we should rethink what the role of the harness is, because harnesses can actually greatly increase the generalization capability of your model and the amount of data that it’s given. And this is not exclusive to RLMs. I think there is a wide class of harnesses that are yet to be discovered that actually can yield similar properties. And to extend this argument a little bit further, I think what you can also extrapolate from this is, like, if I look at an RLM, what are the components of an RLM that are actually necessary, and can I actually just directly train a model to do this? Can I train a model to act as an RLM implicitly in its forward pass? It’s a really weird thing to think about because, like, you might say like, “Oh, code is non-differentiable, blah.” But there are many approximations of this behavior that we will start to uncover. And I think, like, we will see beyond just, like, I’m gonna design a new coding harness that uses a special form of compaction or something. I think we can be a lot more creative here. Like, there’s, there’s so much we can do with these language models that I think we are just not doing. And I’m, like, very excited about this because I think, like, I think we can get serious gains from very opinionated and good harness design that lends itself better to scale. And what by this is, like, the RLM, for example, is a very primitive inductive bias. Like, there’s nothing super special about the design other than the fact that it’s very different than what we currently do. But this may potentially scale much better with, like, the data and the environments that we have available to us.Vibhu [00:40:20]: I guess the, opposite thing that people would probably ask is current harnesses Are very generalized towards coding, which people see works for a lot of domains. Cloud code is being used for design, presentationsAlex Zhang [00:40:33]: Yep.Vibhu [00:40:33]: Everything. MuseSpark, Grokbot.Vibhu [00:40:36]: These are very simple, non-opinionated harnesses that are good at code, and that is also scaling out. what’s the example of how we improve those, I guess?Alex Zhang [00:40:48]: Yeah. let me bring up another paper, which came out very recently. It’s like the harness tax paper. I think it’s by Arena. I really like this paper because it puts forward a prior that I had, which is basically that, like, mostVibhu [00:41:07]: So it confirms the prior.Alex Zhang [00:41:08]: Yeah. Like, most harness choices don’t matter becauseVibhu [00:41:12]: Yeah.Alex Zhang [00:41:13]: All of these harnesses are the same. But I will say, like, Grokbot, for example, is actually quite different, I think, from my understanding, than how some of these other harnesses have been designed, and I like that a lot. and I think it’s clear from here at least that, like, I’m pretty sure. Anthropic or OpenAI are exclusively training on their harnesses. They’re probably not training on their competitor’s harness. I’d assume not, because I don’t know why they would do that. ButVibhu [00:41:38]: But, this is a thing you see in open models, right? Like, Qwen is really good at using open code.Alex Zhang [00:41:43]: Yes.Vibhu [00:41:43]: They need to train in harnesses. Old Gemmas were notoriously bad at this.Alex Zhang [00:41:47]: Yeah.Vibhu [00:41:48]: Models are good, but you need to train in a harness.Alex Zhang [00:41:50]: I think, though, as models get smarter, or, like, as they get better, this distinction becomes, not that important in the sense that, like, if you take Astra and you put it inside of open code, like, it’s not gonna go crazy, because I think it’s, like, just sufficiently good. And so I say this because the only benefits between these different harnesses is just cost, for the most part. And I think, like, what I’m getting at, Sugru, is if you plug these models into RLMs, though, they’re not that good still. They’re okay. And I think it’s mainly because the types, like the class of harness that we are training around is this class of harness, this, like, pi loop, this, like. I like to call it trajectory as a prompt, which just means, like, you keep the whole trajectory of the rollout as the context that your main model is using. Even if you use subagents, it’s still, like, kind of this form. And I think we’re going to. If we want to explore new harnesses, like, there needs to be teams that are dedicated to actually running meaningful experiments over, like, scaling out new harnesses, like maybe post-train scaling out on different harness designs. Like, I think we can actually get very meaningful knowledge or gains from doing this kind of thing, whether it’s an RLM or whether it’s something different. And that’s exciting, ‘cause I think, for example, if you train a lot on. Fable for a long time was the best model for RLMs because they had dynamic workflows, and it was pretty obvious that, like, this was a capability that was somewhat trained in. Even if the model was still, like, a little dumb, like, in the RLM harness, it still worked a lot better than other models did. Astra is now also, like, good enough at doing these things. ButSwyx [00:43:33]: Wait, is this where we see that Fable is the best for RLMs, or is there some otherAlex Zhang [00:43:38]: Oh, no, these are all internal results, I guess.Alex Zhang [00:43:40]: Yeah, I don’t, I don’t have themSwyx [00:43:41]: OkayAlex Zhang [00:43:42]: Public right now. But in general, like, I think you can, You can very easily tell that we have not optimized for RLM, like, workflows yet. And I think if we get models that do this correctly, they will be a lot more efficient as well. you can kind of just think through, like, why this is the case, right?Vibhu [00:44:03]: I think this is the point where you have to give the ten-second what are RLMs.What Is an RLM?Alex Zhang [00:44:07]: Oh, yes.Vibhu [00:44:07]: Because there’s a lot of listeners here thatSwyx [00:44:09]: Yes. We’re, we’re assuming a lot of knowledge.Alex Zhang [00:44:10]: Yes.Swyx [00:44:11]: Also, I think you. But you have set some contextVibhu [00:44:13]: YesSwyx [00:44:13]: So you can. Like, with everything we just saidVibhu [00:44:15]: YesSwyx [00:44:15]: Can we have a clean, crisp definition of RLMs?Alex Zhang [00:44:18]: Yes. Okay. I want to go back to the blog, theVibhu [00:44:21]: YesAlex Zhang [00:44:21]: The compositional generalizers blog. This one. Okay.Swyx [00:44:24]: Okay.Alex Zhang [00:44:24]: This is, like, the best. I wish I had this in the original paper. An RLM is basically just a harness design where the only tool in the harness is code, which is this programmatic subagent calling thing, where it has the option to call itself as a tool, and it has other tools. But all of these things are functions in code, and the context that it’s dealing with is always stored in some memory inside of this code environment. So this could be a file system. Like, this could be, like. I’ll give you an example, Prime Agent. The trajectory of Prime Agent, like the context, even the. when you compact and do all these things, is stored on disk. So the model can always reference its original context, even if it’s compacted, and all of its tools are run inside of, let’s say, like, a Python REPL or a Bash REPL. And so this. it’s like this very primitive abstraction. And I would say, like, where most harnesses differ is, A, context offloading is not done that, like that. if you look at Prime Agent, by the way, Prime Agent does context offloading, but not all the way. So, like, it still maintains the standard cod code, Codex loop of, like, trajectory as a prompt where you compact, but it has the additional kind of, like, the context is offloaded, and it only has. The unique point of Prime Agent is that the only tool is IPython. So this is, like, the very kind of generic abstraction around, like, RLMs. Yeah.Vibhu [00:45:59]: Concretely, what’s the core thing RLMs are trying to solve?Long Context, Composition, and Locally In-Distribution TasksAlex Zhang [00:46:02]: Yeah.Vibhu [00:46:02]: At one point, I think when it first came out, it was context.Vibhu [00:46:06]: I’ll pass the question.Alex Zhang [00:46:06]: Yeah. So the original problem was harnesses have a really bad time dealing with long context. They typically were used to only really do them for, like, specific things like code. Like, they could deal with your code base because it was trained on it. But now it’s more around what this blog is talking about, which is Compositionality and the fact that, like, I think harnesses. We want to have language model systems that have much more control over the actions they make at every step. And what by this is tool calls are very limited because you have to invoke them every turn. Like, you have to invoke tool A, then tool B, then tool C, and there’s no central context that you can kind of draw back from. And RLMs are specifically, like, designed around composition and having, like, a central context that you can always draw from. and this context is, like, designed around the existing language models. Another, like, very similar example actually in design is, like, agent swarms, for example, the Hugging Face incident. Like, these agent swarms have, like, a message board that they learn to communicate over. And this message board, in some sense, is the shared context that they, like, act over. And RLMs basically say that, like, the best way to communicate through this is in code. Like, you write the code to do this, and it’s, it’s because these models are so good at writing code. like, we wanna take advantage of that fact.Swyx [00:47:37]: And then for this compositional thing, up to and including generating your own harness, specific for the task.Alex Zhang [00:47:44]: Exactly.Swyx [00:47:45]: Right?Alex Zhang [00:47:46]: I think we will start to see that if you go up to, this figure. Okay. So we talk about this idea of locally in-distribution tasks for a harness, and it is like a, an idea on top of, like, in-distribution tasks. When we think about language models, an in-distribution task is just a task where, like, the prompt is something that the model has either seen before or, like, has seen some version of it. Most harnesses work like the one on the right, which is they keep appending the trajectory as a prompt. And so eventually, unless you’re Anthropic or OpenAI and you train on, like, kind of these, like, user trajectories, most of these things end up being out of distribution for the most part. But locally in distribution is basically the compositional argument of if an RLM breaks down its computation into, like, kind of a meta-harness of sorts or, like, a program that involves subagents that, like, look at a local problem, every individual language model call over the course of this task is in distribution, even if the entire task itself is out of distribution. And this is a very desirable property, I think, for obvious reasons. Like, if every task is in distribution for each individual language model call, you will probably get to the right answer. so.Swyx [00:49:00]: The logical limit of RLMs is RLLMs where, like, you not just, you don’t just write the harness, you also train a custom model forTraining RLMs and Smarter HarnessesAlex Zhang [00:49:11]: Exactly.Swyx [00:49:11]: You collect data, everything.Alex Zhang [00:49:14]: Yeah.Swyx [00:49:14]: Like, it’s a fully automated AI researcher inside of your harness.Alex Zhang [00:49:16]: Yeah. We will see where the training of RLMs goes. I will say, as an academic, I am not working on this at MIT, or at least in the scaled sense, because I can’t afford to. but there are companies out there that are working on this. I think Prime and Select is very clearly working on this, and it’s very cool. Like, I’m, I’m very excited to see. Maybe we’ll observe, I don’t know, but maybe we’ll observe better, like, post-training scaling laws with when you train around a smart harness. Maybe we’ll even see smarter harnesses that come out and, like, they work better around these kinds of principles.Swyx [00:49:50]: What is a smarter harness? Like, that doesn’t mean anything. You just said they’re all the same.Alex Zhang [00:49:54]: No. What, more of what is, like, Claude Code, Codex, Pi, et cetera, are all the same in that, like, when you break down the logic of the harness, it’s, like, virtually the same thing.Swyx [00:50:06]: Yeah, two calls in a loop orAlex Zhang [00:50:07]: YeahSwyx [00:50:08]: Whatever.Alex Zhang [00:50:08]: But With RLMs and with other harness abstractions, it looks very different. And this is where I think you really distinguish. it’s, it’s in the same way that, like, I think with language model architecture choices, a lot of architecture choices end up kind of looking the same when you, like, scale it out or, like, it. The differences end up being, like, somewhat minor in terms of. for a lab it’s not minor, but, maybe one model converges better than the other one, like, slightly. But in general, like, if you were. So for example, pre-training scaling laws only hold because the architecture choices we have are somewhat stable right now. But if you were to completely change the architecture, pre-training scaling laws probably don’t hold. Or, like, these kinds of. This, like, power law is gonna look very different. And it’s like the same thing with harnesses. Like, I think all the harnesses we have right now, for the most part, roughly look the same, but there are some exceptions to this, I think, that are coming out.Swyx [00:51:06]: I was gonna say, I actually, one of the things that I’ve been more interested by, like, talking about PhD students who take big risks, is that people have been. People also pursuing the other side, which is pre-training scaling laws don’t hold if you change data. right now it’s just raw, unstructured text, corpus of internet. What if you had a better data representation to train on? Then yes, your scaling law would change as well.Alex Zhang [00:51:28]: Yeah.Swyx [00:51:28]: So there’s architecture, there’s data, and, whatever else, you can think about. Well, so I just wanna get back to this. it all makes sense. It’s, it’s very interesting how you sort of recurse up and down the stack from, like, very conceptual to, like, not like, well, this is where we are today.Prime Agent and Opinionated Harness DesignSwyx [00:51:43]: But, like, yeah, obviously, it can scale up and down. I guess, I’m curious, how did you start working with Prime? Is Prime taking on more work with this? Is this their answer to Hermes agents?Alex Zhang [00:51:56]: Mm.Swyx [00:51:56]: You mentioned Grokbot is a little bit different. I just wanted to, like, namecheck all these guys and get your thoughts on each.Alex Zhang [00:52:01]: Yeah. I got involved with Prime, after they released a blog post, by the way, not affiliated with me at all, about, like, how they believed RLMs were kind of the future. And I had a friend that was working there, GPU Mode, Matei. Like, we got in touch, and I think I agreed with a lot of the researchers there and, like, what they believed about harness design. Like, I was very impressed, I think, that, like, they understood the purpose of the RLM paper, which is not necessarily just to say that, like, we’re solving long context tasks, but actually, like, we want more opinionated harness designs.Swyx [00:52:39]: Yeah. There’s always, like, the result of the paper that you choose to highlightAlex Zhang [00:52:42]: YesSwyx [00:52:43]: Versus the actual point.Alex Zhang [00:52:44]: Yes. As I would love to talk about, like, the incentives of academia and, like, the things around, like, why it’s kind of flawed and all the issues, and we’ll get back to that. Yeah. So anyways, I love the guys at Prime. So we kind of had been. After we decided to work together, we decided to look into training in RLM and also build this kind of RLM harness and kinda see where we can take it. That is how, like, Prime Agent came about, and I think the reception for Prime Agent has been pretty good. Like, the one thing I was worried about with Prime Agent is that none of them, at least at the time when we were building it, none of the models were that good at doing RLM stuff. So this was, like, pre-Fable, pre-Astra.Swyx [00:53:27]: I guess, I think to take a step back, can you explain what Prime Agent is, how it’s different thanAlex Zhang [00:53:32]: YeahSwyx [00:53:33]: A traditional, Claude Code, what people would expect harness?Alex Zhang [00:53:36]: Yes. So Prime Agent, I think I mentioned this a little bit earlierSwyx [00:53:40]: Yeah, there was the diagram. YeahAlex Zhang [00:53:41]: Is basically. it is a. A harness on top of Pi, like Pi Mono, which is-- Pi Mono, for context, is like the, like aSwyx [00:53:50]: Core agentAlex Zhang [00:53:51]: A minimalistSwyx [00:53:51]: Yeah.Alex Zhang [00:53:52]: Yeah, like harness. I use Pi as the reference for everything because I think all other harnesses are basically just Pi.Alex Zhang [00:53:57]: But it is Pi, except we explicitly restrict IPython to be the only tool that’s available to it. Every other tool gets loaded in as, like, a Python module, or like a Bash kind of script that it can run. So it uses the core RLM abstraction on top of Pi, and then it also has this continual harness thing, which is, Seth, he’s another PhD student. This is a thing that he used to get language model harnesses to play games. Like, so he worked a lot with Joel, who is the, like, Gemini plays Pokemon guy. And continual harness is also, by the way, very simple. I quite like it. It basically is this design, principle around, like, what parts of the harness can you let the harness itself modify? There are certain pieces that, like, you’ll let it modify its own skills, the subagents available to it, what the system prompt to the model is. And continual harness is available basically as a tool inside of the IPython kernel. And so that’s what Prime Agent is, like, how, what it’s designed around. Everything else in Prime Agent is likeSwyx [00:55:05]: Standard.Alex Zhang [00:55:06]: Standard.Swyx [00:55:06]: Standard.Alex Zhang [00:55:06]: Right? Yeah. I think what is, what I really liked about it, and we got kind of lucky, is that, like, a lot of the new frontier models actually work really well inside. And actually, even a lot of the open-source models work really well, at least some of the newer ones. And there is another thing in Prime Agent I should highlight, which is that, like, we have a very particular agent-to-agent communication system or, like, framework, which is because RLMs tend to spawn many subagents, we want a way for subagents to communicate with maybe the root or with each other. And so there are some design decisions around, like, what each subagent is allowed to talk to, how it does it. Again, everything is in code, so it writes the code to do this kind of communication, which I think is really cool. And then there’s, I guess, persistent subagents is another thing that was kind of added, which is the subagents, they can last beyond, like, the standard runtime of the actual, like, original agent. And you can go into that subagent, you can prompt it more, like, you have more visibility and flexibility into what is kind of going on.Swyx [00:56:10]: This is my number one pain with Codex right now. They, their subagents are just very ephemeral And they actively discourage you from using it for long-running things.Alex Zhang [00:56:17]: Yeah. Yeah. Which I think it makes sense. Yeah. ISwyx [00:56:21]: So the trick is just externalize to a file system, right?Alex Zhang [00:56:24]: Yes. Yeah. That’sSwyx [00:56:25]: Like, that’s the trick.Alex Zhang [00:56:25]: That is the big trick.Alex Zhang [00:56:27]: Yeah.Swyx [00:56:27]: And well, and also, like, force everything to run through code. trust the modelAlex Zhang [00:56:30]: YepSwyx [00:56:30]: That can write code, and it’s gonna write its own harnesses itself. So is Prime gonna take on, like, training, post-training custom models for this? Is this a one-off collaboration between you guys, that’s it? Like, what’sAlex Zhang [00:56:41]: Yeah. They are training a model, intern-- I think they were pretty public about this actuallySwyx [00:56:46]: YeahAlex Zhang [00:56:46]: Back in March.Swyx [00:56:47]: Clearly it is their business.Alex Zhang [00:56:49]: Yeah.Swyx [00:56:49]: Yeah, so.Alex Zhang [00:56:50]: Yeah. they’re, they’re showing that they can train it on their kind of hosted training stack. But no, so for model training, I’m, I’m not involved with them on that. The main reason is just I have other things in the PhD I wanna work on. I think, like, there are many other big bets to take,Swyx [00:57:03]: OohAlex Zhang [00:57:03]: Outside of just RLMs. someSwyx [00:57:06]: OohAlex Zhang [00:57:06]: Some I don’t know how much I can share yet. but in general, like, I think, I actually think one of the luxuries of being a PhD student, genuinely, is that there’s so many big bets to take. most of them will probably yield nothing, but it’s a really exciting time to be in research, especially because I think most progress in the field has been a little bit boring. Like, I’m not saying the outcome is boring, but the process of doing these things tends to be quite boring. and so there is kind of this question of, like, what do we wanna do next? butSwyx [00:57:39]: YeahAlex Zhang [00:57:39]: We can talk about that later.Third-Party RLM Work: Harvey, Headlong, DSPy, and ARC-AGI-3Swyx [00:57:41]: Yeah.Alex Zhang [00:57:41]: Yeah.Swyx [00:57:41]: Okay, I wanna close out a little bit more of your research, and then we can,Alex Zhang [00:57:44]: Cool. YepSwyx [00:57:44]: Start putting it out. since you released RLM, a lot of excitement about it. Any secondary third-party work that you wanna shout out as, like, that you guys should take a look at this?Alex Zhang [00:57:53]: Oh, yeah. So, Harvey, the legal AI company, released a blog post, not affiliated at all, but they post-trained an RLM on their, like, legal work, which often involves a lot of, like, sifting through documents and kind of looking through, like, a variety of, specific information that maybe is not so easy to retrieve with, like, a pure retrieval system. And they show, like, really good results. It’s very exciting. I was shocked that they worked on this. They did not tell me, so when this came out, I was like, “Oh, that’s awesome.” So there’s this one I think is super cool and what they’re doing there. I think this is a collaboration with Base 10, by the way, as well.Swyx [00:58:33]: Yes, this was Base 10.Alex Zhang [00:58:34]: Headlong, which is law, the Law Institute’s kind of. it is their, like, persistently running harness. it’s very cool thatSwyx [00:58:42]: Oh, they renamed it? They used to call it something else.Alex Zhang [00:58:45]: It was like AutoSwyx [00:58:46]: Terminus.Alex Zhang [00:58:46]: Yeah, I know. They’ve gone through. Yeah.Swyx [00:58:49]: All right.Alex Zhang [00:58:49]: So this is Andy Konwinski’s big project. it’s super cool. I love Andy. I don’t want to downplay what they’re doing because they’re using the RLM abstraction, but they’re doing something much cooler than the RLM, which is like they have a system that kind of what they call, like, thinks persistently. So even when you don’t query it has a way to, think through problems that it has in its context.Swyx [00:59:16]: Oh, so it’s just like a always-on type thing.Alex Zhang [00:59:18]: It’s like an always-on thing, but it’s, like, not that expensive. Like, they control the token costs, to make sure it’s not, like, burning through all your credits. This is super cool. I’m trying to think. There are many. Actually, if you go to the RLM, GitHub page, there’s a bunch of things I’ve linked, below. There’s a ton of really cool kind of things that people have been doing. Axe is another really cool one that I think it’s just by this one guy. It’s like a harness around DSPy and RLMs. DSPy also has an RLM. Oh, the last thing I’ll shout out is on ARC-AGI-3, I believe, there were a lot harnesses on their, like, Kaggle competition, like the official one, not the, like, public primates, like one that, or like what people have evaluated on. They all, like, claim to use or they reference, like TUFA, for example, some form or some inspired form of the RLN abstraction in their harness, which is really cool. I think it’s, This is where-- this is exactly the setting where you would see a lot of benefits from composition and using code and combining, like, neuro symbolic systems with AI. And soSwyx [01:00:27]: Yeah.Alex Zhang [01:00:27]: Yeah. Very cool.Swyx [01:00:28]: We love a good neuro symbolic reference.Agent Swarms, Unsolved Math, and What the User Should SeeAlex Zhang [01:00:30]: Yeah.Swyx [01:00:30]: You also, mentioning ARC-AGI-3, OpenAI comes out and says, “We’re at 99.9% on this.”Alex Zhang [01:00:36]: Yep.Swyx [01:00:37]: They also say, “We solved Navier–Stokes. We just threw a model at it.”Swyx [01:00:40]: There’s some debate around whether or not it’s just model.Swyx [01:00:44]: Are they using an RLM? Do?Alex Zhang [01:00:46]: I would guess probably not, unless you say, like. I’ll, I’ll be, I’ll be careful here because, people debate what is an RLM, what is not an RLM. It’s somewhat clear that what they used is some kind of swarm of agents with a shared, some shared context, like some shared file system. And, like, this is very much in the spirit of RLM stuff, but I think there’s a, there’s a lot of, like, more clever things that they did that’s not maybe related to the RLM itself. I agree a little bit with the idea that, like, a harness is not that necessary for what they did. The way that I would put this is that I think a model, like a GPT-6 Astra type thing, is technically smart enough, conditioned on the right information, to come up with a proof for these very difficult problems. Now, how you get to that information is a giant question mark. And in their case, it probably came down to, like, a very long search over, like, many of these sub-age-- or many of these, like, agents in the swarm and maybe also, like, researchers cond-- I’m, I’m actually not sure about this part, but putting in, like, their kind of intuition as to, like, what you should explore and things like this. And ultimately, like, this produced some information that some agent was able to take to finish the proof. And so in that sense, like, I think, was the harness that important? No. And I think what this is pointing at is, like, the specific details of a harness do not really matter, and I think that’s also what that, what the harness task paper is pointing at, which is that, like, beyond the user’s feeling of the harness, realistically all that matters is just, like, how are you composing these agents in a meaningful way to get to the final answer? And maybe that’s what, like, swarms and all these things are really about. And so from my POV at least, if we start to think about, like, for user use cases, what do we want out of harnesses and things like this? Like,Alex Zhang [01:02:51]: We want to take the good parts out of these, like, the Claude codes, the Codexes, like the stream that people like to see. But, like, under the hood, whatever is running can be some really weird, complicated swarm of agents that, like, ultimately come up with an answer. The user doesn’t wanna see that, though, obviously, right? Like, it’s, it’s not legible information. And so I-- this was another kind of thing in the spirit of RLMs, like recursive language model. It sounds like it’s a language model and, but it’s not a language model architecture. But the reason for this is, like, I think we will start to see in the future probably one day, and I wrote a blog about this, what we think of as a language model, like the thing that we query, might actually be like a swarm or like a scaffold or like some Weird harness design that scales very well, but the user just doesn’t see it. Ultimately, all the user sees is some front-end version of this harness. And yeah, I think it’s a relatively safe bet at least to make that this is what we will see.Swyx [01:03:48]: Yeah.Alex Zhang [01:03:49]: And this maybe goes back to the limitations of the base transformer. Like, obviously if you just took a base transformer and you said, like, “Solve Navier–Stokes,” or something, it’s not gonna do it. Like, yeah, we all know this is not what’s gonna happen. But yeah, I think this is maybe the more interesting part. and maybe the claims around, like, did the harness matter is more around this, of, like, just arbitrarily pointing models, like, or agents at a growing kind of context of information maybe is just enough to solve very difficult problems. that I can buy.Swyx [01:04:20]: While there’s a lot in there, I do wanna say, the amount that OpenAI spent is semi-public. It’s, 10,000 agents in 88 hours.Alex Zhang [01:04:29]: Oh, yeah.Swyx [01:04:29]: 130 billion output tokens, which is estimated to be about 40 million dollars in public pricing.Alex Zhang [01:04:34]: Surprisingly, actually, like, less than I thought.Swyx [01:04:37]: Yeah, not that much.Alex Zhang [01:04:38]: Yeah. Yeah.Vibhu [01:04:38]: 130 billion output for the final, but as you said, there was a lot of context being passed around.Vibhu [01:04:44]: It’s more than double that in just the total agent messages being sent.Alex Zhang [01:04:47]: Yeah.Swyx [01:04:48]: Yeah.Alex Zhang [01:04:48]: Yeah. Yeah.Swyx [01:04:49]: I think, one thing I was honored to bring up also was Cursor as far as, like, swarm stuff is concerned.Swarm Architectures, Coordination, and Token EfficiencySwyx [01:04:53]: This is slightly older, meaning February, which is ancient.Alex Zhang [01:04:57]: Whoa.Swyx [01:04:57]: But if you scroll down all the way to the final sort of multi-agent architecture that they arrived at, it was basically an org chart of a normal software team. One thing I’m thinking about, because I basically. There’s, like, this, gather all function That you have to do with subagents or. It’s very similar to GPU programming, actually.Alex Zhang [01:05:16]: Yeah. Yeah. Yeah.Swyx [01:05:17]: And so, that’s a bottleneck. This is a bottleneck.Swyx [01:05:21]: If there’s one main guy, that’s coordinating all the sub guys, then they have to, like, gather again and then re-coordinate.Swyx [01:05:29]: That’s slow. That’s, that’s crappy. what a true swarm should be is everyone is just their own person.Alex Zhang [01:05:35]: Yeah. Yeah.Swyx [01:05:36]: Right?Alex Zhang [01:05:36]: Well, I agree with this, and I think that there is a question to be had, though. Let me give an analogy, which is like, when would you use compaction and when would you use an RLM? And there are many settings where an RLM can solve maybe a more difficult task than compaction can, but you would prefer to use compaction in most cases because it’s cheaper and it’s quicker. And I think in the context of agent swarms, there is a similar thing going on of, like, I’m. Fairly certain that, like, 95% of the swarm is entirely useless, or, like, what it’s exploring is entirely u-- You’re just burning tokens. Versus in this setup, maybe not so much. I’m not sure. maybe it’s also the case here.Swyx [01:06:17]: Everyone has a job. This is your board.Alex Zhang [01:06:19]: Yeah.Vibhu [01:06:19]: I think at some level that’s how problems are framed, right?Alex Zhang [01:06:23]: Yes.Vibhu [01:06:23]: So, like, if you have a search problem and you’re spanning out a bunch of subagents to do search, there’s gonna be a lot of useless information, right? There is one retrieval answer That you’re getting and you’re spanning off, but that is consciously understood, right?Alex Zhang [01:06:36]: Yes. But there is kind of this question of, like, what is appropriate to solve for which problems? Like, what design-- in theory, OpenAI can use-- can package up this API and they’ll call it swarms, and then they’ll give it to you and they’ll be like, “Point this at any problem and we’ll give you a solution.” But maybeSwyx [01:06:53]: Yeah, it’s called, it’s called pro, right?Alex Zhang [01:06:54]: Yeah. maybe you’ll have to pay like 40 million dollars to get a result.Alex Zhang [01:06:57]: And it’s like, well-- But it’s exciting. I will say, like, it is-- it’s very exciting that we even have the option to point 40 million dollars at a problem and solve it.Vibhu [01:07:07]: Yes.Alex Zhang [01:07:08]: But there is still kind of, a lot of research to be done in this area around, like, what is necessary. Like, what do we want to do? What design do we want? we probably don’t want everything to be a swarm, but, like, where do we draw the line? Like, can the agent design that or decide that for itself? Et cetera. So.Open-Endedness and Research Without a Fixed ObjectiveSwyx [01:07:26]: Yeah. and then, just a quick check in case you have opinions on this. Have you looked much into open-endedness as a general category of problems?Alex Zhang [01:07:35]: OhSwyx [01:07:35]: Meaning no prompts, just go.Alex Zhang [01:07:38]: A little bit. so I was at Sakana for a summer, right after, or I guess right before my PhD, and that’s something that they work on a lot there. And I think there’s a lot of people even at, like, Recursive Super Intel-- There’s many of them now.Swyx [01:07:53]: Yes, we just had Richard Socher on.Alex Zhang [01:07:55]: Oh, yes. Yeah. So, like, Richard’s company and then also. Actually, wait, that might be the s-- it might be the same company. I don’t remember. Is Tim Lautenschlager alsoSwyx [01:08:03]: Yeah.Alex Zhang [01:08:03]: Okay. Yes, that company.Swyx [01:08:04]: He’s the main co-founder. He used to be head of open-endedness for Google.Alex Zhang [01:08:07]: Yeah.Swyx [01:08:07]: Yeah.Alex Zhang [01:08:08]: So I think with open-endedness problems, like, I view them as somewhat similar to even, like, unsolved math problems. Maybe that’s a weird way of framing it, but, like, I think a lot of the techniques in terms of, like, how people like, approach them are kind of the same. Like, evolutionary search is, like, very similar to just launching swarms of agents and hoping that, like, they come up with, like, an interesting s-- And this is what, like, AlphaEvolve and some of these other works did, like a year or two ago. But I think what maybe is not, And I’m not sure if this is what you were alluding to, but I think what’s not super clear in open-endedness style search problems is do we frame all of them as basically like an unsolved, very difficult problem, or, like, where the objective is clear? If that’s not the case, I still don’t know yet entirely what the value of it is. maybe you have other opinions. Like, I don’t have too many opinions on this, but at least from my time at Sakana, like, I got the sense that, like, we ultimately still kind of wanna approach things the way that, like, say, OpenAI approached Navier–Stokes. We want. We-- There’s still a lot of nudging in certain directions that we want to have to, like, get to the point where we have something interesting.Swyx [01:09:29]: My version of it is, like, maybe it’s a split between basic science and applied science. Basic science, you’re researching for researching’s sake.Swyx [01:09:36]: You just wanna understand things better. I have no idea if, like, there will be any application at all whatsoever, but, that, And then applied, you have a goal.Swyx [01:09:45]: You’re, you’re trying to minimize loss in some way? and so, what I really, think, in terms of, like, the big bets that people have, what if there was no prompt? Like.Alex Zhang [01:09:57]: I see.Vibhu [01:09:58]: You just pick domain and let itSwyx [01:09:59]: Like, you just, like, you just spawned in this, like, swarm of things and you’re like, “Hey, what’s up, guys? Like, what you guys working on?”Alex Zhang [01:10:03]: Yeah.Swyx [01:10:03]: And, like, you just decideAlex Zhang [01:10:05]: I seeSwyx [01:10:05]: Like, this is an interesting problem.Alex Zhang [01:10:06]: The biggest issue that, like. And maybe there’s a clear path to this, but when I was there at least, the biggest issue that people had with open-endedness was like, how do you ultimately choose at the end? How do you pick out the interesting stuff? Because when there is no. Like, maybe the agent comes up with a goal, but in a lot of cases, like, what they. And they have something called Fugu, I think, which is like a It’s like a model router type thing that was, like, inspired at least by this idea of, like, let’s pick a problem where maybe we can pick out the best solutions to something. In this case, it’s like pick the best model for this problem.Swyx [01:10:41]: You’re the first person to connect model routing to open-endedness.Alex Zhang [01:10:43]: No, yeah. But, so I bring this up because I think, like, with open-endedness, like, just generally the issue is, like, when we have this giant corpus of, like, slop, like, how do we sift throughSwyx [01:10:57]: YesAlex Zhang [01:10:57]: And find, like, the hidden gems? And, like, the solution to open-endedness really just letting models run forever and, like, finding. Like, just doing data gen-- just doing super high throughput data generation and then, like, asking agents to go through and, like, find meaningful things. Like, I’m not sure. Maybe that’s sufficient? Like, that would be, that would be cool.Swyx [01:11:21]: To me, it’s, like, very interesting as a counter to basically all of machine learning Where you have a goal, to have no goal.Alex Zhang [01:11:28]: Yeah. Yeah.Swyx [01:11:29]: But, or, like, an ill-defined goal that you’re like, “Well, what about this goal?” And you’re like, “Well, okay, maybe.” And then you, like, sort of research more and you findAlex Zhang [01:11:36]: YeahSwyx [01:11:36]: That is an interesting goal. ‘Cause, like, I think, like, finding the objective function, like you said, like, Jeff found an objective function That was interesting that no one was exploring.Alex Zhang [01:11:43]: Yep. Yeah.Swyx [01:11:44]: I think that is, like, similar to your message about grad students as well. Like, you stay in school because you are. you want to pursue open-endedness. If you want to, profit max and, like, join the, escape the permanent underclass, then you join a lab.Swyx [01:11:59]: Right?Alex Zhang [01:11:59]: Yeah. Yeah. It’s funny. I feel like I don’t, I don’t hear this discourse a lot. I’m, I’m in the East Coast, so it’s, like, a very different type of. But then when I. whenever I come here, it’s like, that’s always, like, the topic of discussion.Swyx [01:12:12]: You cannot pay rent without doing this.Alex Zhang [01:12:13]: Yeah.Swyx [01:12:15]: Yeah, you’re getting priced out, guys.Vibhu [01:12:16]: Yeah.Swyx [01:12:17]: Okay. So, yeah, there’s, there’s all that. I don’t know if you wanna-- if it’s relevant, enough to talk about the mismanaged geniuses, which you were pulling up.Vibhu [01:12:25]: No, it’s just on your blog. But I will poke on, Sakana.Swyx [01:12:28]: Oh, Sakana? Oh, okay.Vibhu [01:12:28]: Yeah. So they did in their, blog post, I talked to them about this as well. So one of the cool results of Ultra is they basically just let it loose on automated data science researchSakana AI and Weird Research BetsVibhu [01:12:40]: With little to no human intervention. It’s just kind of makingSwyx [01:12:44]: Yeah, so this is auto research, which is a little bit more open-ended, and there’s degrees of open-endedness, and I agree with that.Vibhu [01:12:51]: Yeah, separate than auto research with objective, this is justSwyx [01:12:54]: YeahVibhu [01:12:54]: Do stuff. But, it’s cool. They’re, they’re working on it for those that are interested.Swyx [01:12:58]: While you’re bringing it up, actually, what is your take on Sakana? Like, what are they doing apart from being, the Japan one?Alex Zhang [01:13:04]: Yeah.Alex Zhang [01:13:05]: I actually love the people there. Like, I think they have a really smart team. and it makes sense. it branched off from, like, an earlier team at GDM, which was also kind of, I guess, doing this kind of, like, open-ended evolutionary research style stuff. What I liked about my experience there, at least, was that they did have that, like, mishap back, I forget, at this point when, but I think, likeSwyx [01:13:31]: You’re talking about AI scientists?Alex Zhang [01:13:32]: No, the GPU kernel.Swyx [01:13:34]: Oh, okay, yes.Alex Zhang [01:13:35]: That one. Yeah.Swyx [01:13:36]: People cannot forgive them for that. Yes.Alex Zhang [01:13:37]: Yeah. And I guess, like, the AI scientists, like, there’s, there’s some criticisms of it that I don’t, I don’t work on that, so I have no kind of take on it. But I think in general, like, what I like about them at least is that they’re a little bit more of a researchy type lab. So, like, they don’t operate in the same space as, like, OpenAI or Anthropic. Like, for sure, like, definitely no. They do not. it’s pretty obvious probably that, at least when I was there, they do not have a big competitor model or something that, like, that everyone is using. But I think they kind of operate in some ways as, like, a PhD lab, which is cool. Like, and I think, like, David Ha is, like, he’s, he’s really smart. Like, I think he has a good sense of, like. Also, I think the market in Japan is also a little bit different for AI, and, like, who they’re targeting is slightly different than maybe what we’re used to here. But yeah, I like that they take kind of. A lot of their research is kind of weird, I think, when people view it? And I like that. Like, I think it’sSwyx [01:14:34]: We should have more weirdness, yes.Alex Zhang [01:14:35]: Exactly.Swyx [01:14:36]: And you said different market. Just, is it, like, enterprise?Alex Zhang [01:14:39]: Like, the way it works there is a bit different, like how deals happen and stuff like that.Vibhu [01:14:44]: They do have a. I guess this page is originally in Japanese, but they do have a modelAlex Zhang [01:14:49]: OhVibhu [01:14:49]: Specialized for the Japanese market.Swyx [01:14:51]: So you didn’t know that.Alex Zhang [01:14:51]: I didn’t know that.Vibhu [01:14:52]: I didn’t know it too.Alex Zhang [01:14:53]: Did not know this.Vibhu [01:14:53]: I also have personal friends that know the team.Alex Zhang [01:14:56]: Yeah.Vibhu [01:14:56]: So there is a. Even from the sense of a way that you speak culturally Responses are tuned towards that. This is not like it’s frontier on benchmarks. It is a cultural appropriate model for them, and then they have, like, chat and all that.Alex Zhang [01:15:11]: Yeah.Vibhu [01:15:12]: But, to mirror your point, there’s also, like, how should education look like? And someone wants to work on it, and they’re a very PhD lab of, “Do your thing. Why not? We have money. Go research.”Swyx [01:15:22]: Oh, they say it’s a Kimi fine-tune. That’s nice.Vibhu [01:15:24]: Oh, there you go. Kimi.Swyx [01:15:25]: Good for them.Swyx [01:15:26]: Yeah, and, speaking of Kimi, right, like another, just a grad student that spit out and like, yeah, I just have this, like, Kimi delta attention that wants-- that I wanna work on.Vibhu [01:15:36]: Yeah. Yeah.Swyx [01:15:36]: And, like, somehow managed to make Moonshot. Don’t understand it still.Alex Zhang [01:15:41]: Yeah. Well, he’s, he’s super cracked, at least my understanding. I think in general, like, a lot of, a lot of the Chinese labs have done really cool work.Kimi Swarms, Dynamic Workflows, and ConvergenceSwyx [01:15:51]: Yeah.Alex Zhang [01:15:51]: Like, yeah.Vibhu [01:15:52]: Any thoughts on Kimi agent swarms?Alex Zhang [01:15:55]: Yes. one thing I will say is whatever OpenAI is doing with their agent swarm is, like, clearly the right thing to do. You have to kind of think about it this way. Like, nothing, especially without, like, a very smart harness design, which I don’t, I don’t think anyone really has so far, Nothing comes so easily for free. For example, like, the agent swarm design is not something you can just take for granted. Like, it’s not like GPT-6 Astra is just super smart and then it just got agent swarms running well. They clearly trained. the Hugging Face incident was them training a system to be like a swarm. and I think, like, clearly they’ve done something really well to the point where you can throw 40 million dollars and solve an unsolved problem. And I think with the Kimi agent swarm thing, like, at least from when I read it just came off as like, this is interesting, but I don’t actually know whether or not this can solve anything novel.Swyx [01:16:59]: Yeah, they just. They were like, “It makes spreadsheets for you.”Alex Zhang [01:17:01]: They kind of were just like, yeah, like, here is a, here is a swarm that, like, kind of does stuff, and it’s cool. but. And I will say the same thing with dynamic workflows. I actually kind of think the dynamic workflows release was sort of a flop. I don’t know how you guys feel about it, but I think it’s like. my understanding is, like, it’s not used that often or, likeSwyx [01:17:21]: It’s just very expensive.Alex Zhang [01:17:22]: It’s too expensive and likeSwyx [01:17:22]: It’s ultra code. It’s basically like, take over my bed.Alex Zhang [01:17:26]: And it doesn’t I’ve tried it, and, like, it doesn’t act in the way that, like. Again, I’m not, I’m not the biggest OpenAI, like, stan or something, but I think whatever they did was very impressive. Like, they somehow managed to get a way for this swarm to actually actSwyx [01:17:41]: I seeAlex Zhang [01:17:42]: Towards a goal. And yeah.Swyx [01:17:44]: I see.Alex Zhang [01:17:44]: It’s very difficult.Swyx [01:17:45]: I see. So, like, efficiency of the multi-agent swarm is the objective function here.Alex Zhang [01:17:51]: Yeah.Swyx [01:17:51]: Right?Alex Zhang [01:17:51]: Yeah.Swyx [01:17:52]: Like, how much of this is slop? Like, this is a lot of slop.Alex Zhang [01:17:55]: Yeah, I think so.Swyx [01:17:55]: OpenAI is less slop.Alex Zhang [01:17:56]: We take for granted what it means for a swarm to converge to an answer.Swyx [01:18:00]: Yeah.Alex Zhang [01:18:00]: It’s just like. It’s not something we take for granted.Swyx [01:18:02]: Yeah. Yeah. We’ve, we’ve done one pod with Noam Brown and, like, hisAlex Zhang [01:18:06]: Oh, yes, I remember. YeahSwyx [01:18:06]: His thing. His whole thing was like, okay, like, we’ve worked on a lot of, like, competitive agents. we’re working on collaborative agents.Alex Zhang [01:18:13]: Yeah.Swyx [01:18:13]: And, like, that’s now called a swarm.Gemini, GDM, and Harness Engineering at ScaleAlex Zhang [01:18:15]: Yeah.Vibhu [01:18:15]: I would do a quick poke. Do you have any thoughts on Gemini? They were also IMO gold. Like, there was a time where they were getting agents to reason for a long time and.Alex Zhang [01:18:26]: Yeah,Vibhu [01:18:28]: Is it too old to think about?Alex Zhang [01:18:29]: No. I think it’s a little bit blown out of proportion. Like, Gemini. For, again, I, so I should preface by saying I haven’t worked at any of these places, so take this with a grain of salt, right?Vibhu [01:18:39]: You have strong opinions on agent harnessesAlex Zhang [01:18:42]: YeahVibhu [01:18:42]: And, this wasAlex Zhang [01:18:43]: But, let me just say this first. This work was really impressive. I think what they showed here was, like, they took a time when the models weren’t that goodVibhu [01:18:53]: YesAlex Zhang [01:18:53]: And they managed to be very smart about, like, what the harness does. I remember for this at least, like, yeah, like AlphaGeometry, I guess that was a year before this, but it was very cool. They took it to the max, and they, like, designed, I don’t know. I’d say I’m not too big on competitive math, but I think, like, GDM, it’s sort of a shame. Like, everyone I’ve talked to about GDM kind of has the same opinion, which is that it’s way too, like, bureaucratic. Whatever is they have the talent and the resources to do almost anything, but, like, I don’t know, until they figure that part out, like. Nothing against anti-gravity, for example, but, like, I don’t know anybody that uses anti-gravity. And so I’ve tried it once, and it’s. I don’t see a reason to switch to it. and I think for whatever reason, like, they’ve been struggling with this, so, yeah.Swyx [01:19:40]: Yeah. Well, a lot of people dogged on Meta for a long time until they startedAlex Zhang [01:19:44]: Yeah, and they recoveredSwyx [01:19:44]: Coming out. And, like, I think, Google’s going through that phase right now.Alex Zhang [01:19:47]: Yeah.Swyx [01:19:48]: And, it’s, it’s just you gotta stay alive andAlex Zhang [01:19:51]: Yeah.Swyx [01:19:52]: I wanna focus back onAlex Zhang [01:19:53]: YeahSwyx [01:19:53]: Just, like, your thoughts, just general. we can talk about speculative PTCSpeculative PTC and Parallel Tool ExecutionSwyx [01:19:58]: Mismanaged geniuses, or just, like, throw away all this and just talk about whatever else.Alex Zhang [01:20:03]: Okay, let’s, let’s talk about mismanaged genius for a little bit.Swyx [01:20:06]: Yeah.Alex Zhang [01:20:06]: I, the only comment I’ll say on speculative PTC is that it’s a really simple idea. It’s almost, like, obvious that this should be done, and, like, there’s not much more to talk about it. Like, I think it’s just, like, you should just use it for, like, coding. Like, anything with programmatic agent calling, like RLMs or Kodak, like, yeah, it’s like, it’s like a no-brainer.Vibhu [01:20:25]: What’s the, for people that haven’t read itAlex Zhang [01:20:27]: YeahVibhu [01:20:27]: What’s the one-liner for people?Alex Zhang [01:20:29]: The simple thing is when the model is, writing its code or, like, even as, like, after it finishes writing the code, a lot of tools tend to be, like, sequential or, like, you have to wait on them, so you should just launch them in advance. Like, if you’re able to jit compile this code, you can probably figure out, like, even though it’s, like, kind of variables and stuff, like, you can figure out, likeSwyx [01:20:52]: Yeah, statically analyze.Alex Zhang [01:20:53]: Yeah. SoVibhu [01:20:54]: Speculation.Alex Zhang [01:20:55]: Yeah. There is this, Someone pointed to me some actually, like, academics have, especially PL, like programming languages people, have some, like, very kind of cool ways of doing this. And so, like, at some point maybe I’ll, I’ll, I’ll, like, work on this.Swyx [01:21:09]: I guess mostly you have to change language, because if you are in JavaScript, Python, you can’t do this.Alex Zhang [01:21:13]: Yes. Yeah.Swyx [01:21:14]: So, like, Haskell, yes. what’s, what’s the, what’s the normal one that’s, that’s not Haskell?Alex Zhang [01:21:21]: Lisp.Swyx [01:21:22]: Lisp, OCaml.Alex Zhang [01:21:24]: OCaml, oh, yeah.Swyx [01:21:25]: Yeah, any functional languageAlex Zhang [01:21:25]: YeahSwyx [01:21:25]: You can actually, like, pipeline this.Alex Zhang [01:21:27]: Yeah.Swyx [01:21:27]: So Effect-TS if you wanna do TypeScript.Alex Zhang [01:21:29]: Yeah.Swyx [01:21:30]: Okay, we can switch over to,Vibhu [01:21:32]: I like this diagram.Alex Zhang [01:21:33]: Yeah.Swyx [01:21:34]: Which is like your, you guys’ whole thesis, right?Capability Overhang: Reliable Long-Running WorkSwyx [01:21:36]: Like, that, to me this is, like, kind of like a restatement, but maybe I’re missing something ofAlex Zhang [01:21:41]: YeahSwyx [01:21:41]: Like, well, work on better harnesses or, like, your models actually are capable a lot more if you try harder, so this is a skill issue.Alex Zhang [01:21:48]: Yeah. Yeah, basically. I think there’s one thing I want to see. I appreciate that there’s a big focus on, like, jagged intelligence, because it paints a big picture of, like, we can do this if we really set our minds on it. But I kind of wish. And maybe someone in academia should do this. Like, really just sit down and think about, like, if I took Astra, even the current frontier models are not good enough at, like, doing a particular job over, let’s say, the span of a month consistently and well. And I think this is, like, a stupid problem. Like, I genuinely think we can solve this. You don’t need to be a frontier lab and, like, do all this, like, fancy stuff for your IPO. Like, I think, like, these models are so smart that even if it’s, like, a silly way, I think that it genuinely is a skill issue of you can get a model to be as good as, let’s say, like, just some 18-year-old high school kidAlex Zhang [01:22:50]: Doing some job. I think it’s, like, ridiculous that we can’t do that. And it’s. Part of the reason is, like, the format of a language model is not really amenable to that, but I think you can shape a harness around it and do it. And I think, like, this in itself is, like, ignoring the RLM stuff, ignoring all the, like, what abstractions should we use? Like, I just think someone can design a harness that can do this. Like, I, that’s. Maybe it’sSwyx [01:23:13]: When you say do this, do what?Alex Zhang [01:23:15]: Do long-running but simple tasks, and do them reliably.Swyx [01:23:20]: Okay.Vibhu [01:23:20]: What’s an example? So is this different than, like, pick your favorite company, Harvey, for example Using LLMs to do legal work, or what’s the.Alex Zhang [01:23:30]: I guess it’s kind of like that, except if the bottleneck was not, like, certain legal knowledge or something. Like, I don’t know. Let’s say,Vibhu [01:23:38]: I guess, like, the examples, people can take models and build pipelines or whatever and Have an agent repeatedly do whatever it has they want, right?Alex Zhang [01:23:47]: Yeah. So for example, like, if I wanted a general system that I could kind of talk. I can talk to it like I would talk to an intern and basically just ask it to do. to explore some small thing. So maybe an example of this is, like, very silly auto research is maybe an example of this, of, like, not necessarily finding super novel solutions, but at least optimizing all of the easy parts of any problem. They often end up being over-indexed for, like, ML training and things like that. But yeah, I don’t know. maybe that’s, that’s, that’s not, like, super clear, but there is a lot of People’s general workflows where you probably could just vibe code up. Some specific harness to help you do, like automate this thing. Some examples are like automating, finding, like research papers and stuff like that. But usually people will design like a specialized agent to help them do this kind of thing. Or like they’ll vibe put a harness, like, and just run it or like their Slack bot or something. But I almost think there’s just like a standard form, like just a harness that you just plug in. Like you don’t-- It doesn’t need to be designed for finding papers or fi-- Like, you just kind of tell it, find this for me, and you like plug it into that setting. What I’m getting at is that I think there’s a lot of easy things that can be automated. AndVibhu [01:25:13]: Is this like a hypothesis or a point around like capability overhang? Like even if we paused, there’s still a lot of impact to be had with current state of models?Alex Zhang [01:25:22]: In some sense, yes. Like, I guess what I’m presenting is the easiest form of this. But what this is kind of saying is that, like, we have jagged intelligence on a lot of things. Like, for example, models are like disproportionately good at coding and math. This is saying like we can translate those abilities to many other things. So like for example, if you took-- Usually if you take someone who was an IMO gold or something and you kind of apply them to a lot of different problem-solving domains, they can figure it out. I don’t actually know if this applies to models. For example, like in GPU code optimization, one very interesting question is whe-- if you were to take out all of the GPU programming data, like from a model, but it was re-- it was like as good as Astra is now, just without, like with that taken out, would it be able to still optimize GPU kernels? like would it be able to learn in context roughly what it needs to learn and then like have some pipeline or like come up with some solution to solving like optimization tasks? And I think like there’s like a mismatch between, like if you took a human that was as smart or like knew as much as Astra, there is a mismatch between what that human can do and what Astra can do, maybe around a harness. and I think we can actually approximate the human a lot more.Continual Learning and General Problem SolvingSwyx [01:26:42]: To me, it sounds very approximate to the continual learning problem. I think you’reAlex Zhang [01:26:46]: That’s the best example. Yes.Alex Zhang [01:26:48]: Yes.Swyx [01:26:48]: Why didn’t you just say that then?Alex Zhang [01:26:49]: Yeah. I guess IVibhu [01:26:50]: I was like, I was like thinking, could I just blurt out some words like learning?Alex Zhang [01:26:54]: I’m careful with that. But yes, IVibhu [01:26:57]: You have a very eye.Alex Zhang [01:26:58]: Maybe, Maybe, like a bit of a, I don’t know, like aSwyx [01:27:03]: No. So I think you’re a very, yeah, I don’t know, I don’t know your undergrad actually. Are you like a math person generally, orAlex Zhang [01:27:09]: A little, yeah. Yeah.Vibhu [01:27:09]: Did you study math?Alex Zhang [01:27:11]: Yeah, that’s what I wanted to do at least.Swyx [01:27:12]: Like a category theory type of, abstraction where you think in categories and then you have to like then translate down to the specific. And, but then you like actually really care more about the category.Swyx [01:27:24]: And like that’s the communication error because like everyone’s listening for the specific, but actually trying to also, convey the general.Alex Zhang [01:27:31]: Yeah.Swyx [01:27:32]: Which is hard. I don’t really know. you can, maybe use like a shorthand of like, “Okay, I’m at level two and then I’m gonna go up to level three Then come back to level two.” That we should have say some like epistemic, like shorthand for like this kind of thing.Alex Zhang [01:27:45]: Yeah.Swyx [01:27:45]: Because it’s hard. Like you’re, you’re compressing a lot into word after, like sequential word decoding.Swyx [01:27:51]: Should we convert to Neuralese? is there like a, a better form of output than English or, Python or JavaScript? I don’t know.Neuralese, Programming Languages, and Diffusion ThinkingAlex Zhang [01:28:06]: Yeah.Swyx [01:28:06]: This is very kind of like a s**t post, but like people have speculated about like what is the native language that people want to out-- that models want to output?Swyx [01:28:14]: Some people say binary. That’s, that’s Marc Andreessen’s thing. I don’t know. Yeah.Swyx [01:28:19]: PTX?Alex Zhang [01:28:20]: Let’s say a mix of English and Python. And I only say this because The capability of a model is somewhat a reflection of what we train them on. So we still want like. Yeah, I don’t really buy the binary argument. I guess I, like, I understand, but it’s likeSwyx [01:28:40]: Yeah. You wanna model the world in some way. I think the,Alex Zhang [01:28:43]: Yeah.Swyx [01:28:43]: One thing I’ll, I’ll bring up is always, which I always do in this kind of conversation is Sapir–Whorf Which is you, if you choose English, you will have locked into however long English has been around, which is, let’s say five hundred years, which is not that long. Like actually Like what you, the language that you speak constrains how you think.Swyx [01:28:59]: And if you learn a different language, for example, someone, in Chinese, we don’t have tenses. I don’t know if you. I actually didn’t know that. And I speak Chinese.Alex Zhang [01:29:09]: Oh. I did know that, but my Chinese is not great.Swyx [01:29:12]: Okay.Alex Zhang [01:29:13]: Yeah.Swyx [01:29:13]: Yeah. Or like, in, let’s say in Japanese or, I know, I forget what language it is. Like in Korean, everyone you speak to, yeah, you have to like acknowledge social status.Alex Zhang [01:29:23]: Yeah.Swyx [01:29:24]: But it’s a different dimension than gender, right? Like, it just like, it just influences everything you do. when I take Ling 101, apparently there’s a, there’s a, there’s a language in Africa where like there’s a vegetable gender. yeah, right? Like just like you have. Or like Eskimos know no word for snow or whatever. Like, anyway, so like the language that you adopt affects your thinking. And if you Choose to output your chain of thought in English, you are biasing towards whatever English solves. I don’t know what the sort of prior of English is.Alex Zhang [01:29:51]: That’s interesting. I did not think of it that way.Vibhu [01:29:55]: At some level it’s interesting, right? So you’re right on language. a lot of model chain of thought also fluctuates language.Vibhu [01:30:03]: The obvious example is Chinese models Speaking in English might still reason in Chinese. but at the same level, most models are very capable multilingually.Vibhu [01:30:13]: And that adaptation we can see, you can add in languages. You don’t get that much from adding a whole language, butSwyx [01:30:18]: Yeah.Vibhu [01:30:18]: They will reason interchangedSwyx [01:30:21]: Yeah. So we’re all autoregressive. but also likeVibhu [01:30:23]: Right.Swyx [01:30:23]: Let’s say German, like, subject-object, agreement, you have to put the verb at the end, which is very super annoying, like very famously. yeah, right. You don’t know what you’re doing until the end where you’re like, “Oh, that Mess of nouns and then the verb.” Well, the most classic one that most people be-- have heard of is Arrival, where, they have the heptapods where they think, that time is like flat to them. So they think in, they output entire sentences at one shot. so it’s, this is closest to, like, the difference between autoregression and diffusion.Swyx [01:30:55]: We talk in autoregression. What if you could talk in diffusion Where things just resolve over time?Alex Zhang [01:31:01]: I see. I see.Swyx [01:31:02]: But like the whole idea shows up at once.Alex Zhang [01:31:05]: I see.Swyx [01:31:07]: So that is a drastically differenting, language, but it is a language.Alex Zhang [01:31:10]: I see. Oh, that’s really interesting. That’s really interesting.Swyx [01:31:12]: Which, like, machines could speak, that we, probably will never speak, but like, yeah, machines don’t care.Alex Zhang [01:31:19]: Maybe this is a huge tangentSwyx [01:31:20]: YeahAlex Zhang [01:31:20]: But are there not, like, things inherently that are reasoning chains that are inherently autoregressive?Swyx [01:31:30]: Yeah, time.Alex Zhang [01:31:30]: Sometimes, like. Yeah. Or like, yeah.Swyx [01:31:32]: Yeah. Something happens first, then something else happens.Alex Zhang [01:31:34]: Even like, yeah, anything in code, for example, like has to be causal in some-- usually at least has to be causal.Swyx [01:31:40]: Well, no. so it’s a. Then you have to. Then you’re not exploring enough,Alex Zhang [01:31:44]: That’s true. YeahSwyx [01:31:45]: Programming language theory, where, everything is like pure functional and like completely relational And, you sort of abstract away the solver that translates the relationships that is, are always true into code. So I, yeah, I feel like this is maybe a little bit too out of my depth.What Comes After RLMs?Alex Zhang [01:32:01]: No, it’s interesting though. Yeah.Swyx [01:32:01]: But I love languages Whether it’s coding or human, and I do think a lot about how that affects reasoning and the boundaries of what we can do. I don’t need to go too much beyond that. I don’t know if you have any other thoughts. my closing question was gonna be, you have all these research, directions that you wanna do. You’re, you had a GPU mode phase. You had a, RLMs phase.Swyx [01:32:25]: Presumably you have other stuff planned, which is why you’re not, doubling down on that. By the way, I notice that it is interesting how you guys do start with the GPU side, and then you migrate towards the zero gradient side, it’s, which is what Shenyou called it. Doesn’t that feel less legit than messing with GPUs?Alex Zhang [01:32:44]: Yeah, I guess in the sense that, like. So you did bring up that, like, I like to think about things in, like, a math-oriented way.Swyx [01:32:51]: Category, yeah.Alex Zhang [01:32:52]: And it’s, like, very uncomfortable sometimes to be working on, like, harnesses and agents because it’s so.Swyx [01:32:57]: Because you think all harnesses are the same.Alex Zhang [01:32:58]: Yeah. So it’s like super fuzzy.Swyx [01:33:00]: So like, you just, like, two new ideas in harnesses. Got it.Alex Zhang [01:33:01]: Yeah. It’s, it’s It’s, it’s also just, like, empirically it’s hard to, like, verify a lot of findings, at least with the compute that we have available to us. But the reason why I think I’ve moved on to a lot of these problems is I think actually this is where most of, like, the innovation is yet to happen. To me, like, the GPU level is a means to exploring other ideas. Like, you want to, for example, like, get good at writing kernels or, like, even automate writing kernels for the sake of a broader goal of, like, I want to explore ideas where I’m not bottlenecked by systems challenges. In that sense, like, I guess a lot of what’s written there is all harness stuff, but I am also interested in things at the model level as well. but I’ll just leave it at that.Swyx [01:33:48]: Okay.Alex Zhang [01:33:48]: Yeah.Swyx [01:33:48]: That’s a good hint. anything, any. If people wanna reach out to you, what are you looking for help on? What do you want collaborators on? any sort of calls to action?Collaborating on Research and Choosing Big BetsAlex Zhang [01:33:58]: Yeah. So, I guess there’s nothing I have in particular where I feel like I need to work with someone on, unless it’s, like. unless it’s with a company before, like, compute or, like, with. to talk with other people about it. But I will say I’m not. I’m never opposed to working on ideas with other people. I get reached out to a lot by often undergrads or even, like, other students.Swyx [01:34:22]: Podcasters.Alex Zhang [01:34:23]: Podcasters. and usually I feel like I get an email that’s something along the lines of, like, “I really like RLMs.” Like, “I wanna work together.” and I feel like ISwyx [01:34:33]: Yeah, that’s a bad reach out, right?Alex Zhang [01:34:35]: Yes, yeah.Swyx [01:34:35]: The worst is like, “Can I pick your brain?”Alex Zhang [01:34:36]: Yeah.Swyx [01:34:37]: And like, “On what?” Like, “Read my paper, dude.” Like.Alex Zhang [01:34:39]: They’ll like, they’ll be like, “I read your paper,” in quotes, like, “Recursive language models,” or like, “Prime Agent,” like a self-improving RLM harness or something. and it’s kinda like I like. I really like people that are opinionated, even if we disagree. I think if you have strong opinions and are able to, like, think through why you think those opinions are right or wrong, ‘cause usually it’s, it’s hard to actually tell. But, like, you have strong convictions about certain problems. Like, I’m, I’m always happy to, like, chat and even, like, potentially work on something together. I have, like, no limit to who or, like, what I would like to work on. So, yeah.Swyx [01:35:12]: No limit?Alex Zhang [01:35:13]: Yeah. I. in the era of agents, I think there’s a lot more work you can do, like, bandwidth-wise. So I. Yeah. I think in general, like, I am not hard to impress, but I think it just takes a little bit of effort toSwyx [01:35:30]: YeahAlex Zhang [01:35:30]: Kind of. Yeah, knowSwyx [01:35:32]: YeahAlex Zhang [01:35:32]: Know what you want.Swyx [01:35:32]: It’s very clear. and like, when you see a new thing come out, well executed, good, simple idea, then, like, get thatAlex Zhang [01:35:40]: YeahSwyx [01:35:40]: Immediately gets your attention, right?Alex Zhang [01:35:41]: I get excited. Yeah.Swyx [01:35:42]: It’s actually, like, not that hard to get the same attention that all the Frontier Lab guysAlex Zhang [01:35:45]: YeahSwyx [01:35:45]: Because they are looking for you. you just have to put yourself out there, right?Alex Zhang [01:35:48]: Exactly, yeah.Swyx [01:35:49]: But yeah, it’s true. I will say, I think human attention very scarce right now, and I do struggle with, like, the number of projects I have going on.Swyx [01:35:57]: And, I don’t know how to manage it. I don’t think agents are helping at all.Swyx [01:36:00]: Like, I will just prompt it and. I’ll prompt a thing and then never look at it.Swyx [01:36:03]: Right? Like, which is very common.Alex Zhang [01:36:05]: Yeah.Swyx [01:36:05]: Yeah, and that sucks.Alex Zhang [01:36:06]: I guess, maybe the. one of the smaller differences in, like. actually, maybe you were doing research. I’m not sure. But for me at least, like, I’ll have maybe, like, 10 or 15 different ideas that I wanna do, but the thing is, like, most of them are bad.Swyx [01:36:20]: Yeah.Alex Zhang [01:36:20]: And this also maybe is true even for someone that reaches out to me. Like, maybe the idea is actually bad, but it looks interesting to me. And so, like, we can spend, like, some time looking into it, and if, like, we feel like there’s actually something there, like, then we’ll. we should take the next few weeks and just really pursue it. And, like, this is my style with. This is why I love the PhD, by the way, because there are times when I’m just thinking about problems, like, maybe on a run or just, like, playing tennis or something. Like, I’m not, I’m not working, I guess. But it’s like those are the most fun times, and then when I, like, really am convicted about something, I’ll just, like, drop everything and just do it.Alex Zhang [01:36:52]: Like, just spend, like, all my time thinking and working on that problem. And then, once you get to the point where, like, you can just run experiments, then it’s, it’s, it’s kind of easy coasting again, so.Science as the Next FrontierSwyx [01:37:03]: Yeah, sorry. This isAlex Zhang [01:37:04]: Yeah. No problemSwyx [01:37:04]: Like the, for the fourth last question, which is like, I think a lot of people are also thinking about science as the next frontier, like physical sciences Bio, math even. how do you separate, like, I guess, let’s say your choice of projects that is applicable for industry And then maybe it’s part-- choice project is just, like, science?Alex Zhang [01:37:27]: I actually worked on, like, AI for bio stuff before I started my PhD. The field has changed a lot since thenSwyx [01:37:33]: YeahAlex Zhang [01:37:34]: I should say.Swyx [01:37:34]: ‘Cause it’s, it’s like, it used to be a theoretical, like, of course, what do you mean? Like, I have one path and thenAlex Zhang [01:37:39]: YeahSwyx [01:37:39]: I chose that. But now a lot of people are crossing over.Alex Zhang [01:37:41]: Yeah.Swyx [01:37:41]: And like, so we have started a science pod to just Cover those thingsAlex Zhang [01:37:44]: Oh, wowSwyx [01:37:45]: Because a lot of engineers are like, “Well, actually there’s, Tractable problems there.”Alex Zhang [01:37:50]: Yeah. I will preface by saying my understanding of a lot of these topics is probably pretty limited. But I think, like, if I find out either because someone reaches out or, like, I look at a problem and I’m like, “Hey, like, some design principles that we use or that we’re thinking about right now actually make a lot of sense for this problem,” I get excited about those as well. But I think it’s harder. I don’t know. I think with. I think science, especially like empirical or, like, applied science has very long, like, what is it called? Like, feedback loops or whatever.Swyx [01:38:23]: Yeah, it converges to a robotics question.Alex Zhang [01:38:25]: Yeah. to me, like, also this aspect of, like, what is worth spending and betting my time on now? Because, like, maybe I spend a lot of time on this problem and then, like, in six months, like, a different solution, kind of like maybe a new model comes out and it’s like, “Oh, it’s way better for this.” And so I do have to be careful. I, like, you have to be conscious about, like, where you think things might be going.Swyx [01:38:47]: Yeah, exactly.Alex Zhang [01:38:47]: So, yeah.Swyx [01:38:48]: Publish cycle.Alex Zhang [01:38:49]: Yeah. SoVibhu [01:38:51]: Which then I can just tell ARC-AGI-3, it’s saturated. We did it.Alex Zhang [01:38:54]: Yeah, ARC-AGI-3 got saturated in less than a year, so it’s kind of ? Like, it’s. I don’t know. Like, if you were a lab picking that problem, like, you’re probably kind of sad now ‘cause actuallySwyx [01:39:04]: Yeah. Well, so exactly. That’s why knowledge work, gaming, all these things are saturated.Alex Zhang [01:39:08]: Yeah.Swyx [01:39:08]: Now they’re actually the frontier is science. SoAlex Zhang [01:39:09]: Yeah.Vibhu [01:39:10]: Knowledge work is saturated.Swyx [01:39:12]: Yeah. GDP val is like 80 something, 90 something.Swyx [01:39:16]: Like, there’s, there’s 90 to 100% that was obviously gonna getAlex Zhang [01:39:20]: It’s gonna be really hard to tellSwyx [01:39:21]: 10 years. But, like, well, the next low-hanging fruit is gonna beAlex Zhang [01:39:25]: It makes senseSwyx [01:39:25]: The other stuff.Alex Zhang [01:39:26]: Yeah. Maybe I’ll think about that more. I actually, I haven’t given too much thought toSwyx [01:39:30]: I’m just trying to guess your next direction, actually.Alex Zhang [01:39:32]: No. I will say because I think especially at MIT, like, it’s, it’s, it-- Yeah, there’s a lot of really talented scientists there, like, in the natural sciences. And I think it’s, it’s a little bit, like, sacrilegious almost to be like, “I’m gonna figure out, like, your problem.” LikeSwyx [01:39:45]: No, that’s how, that’s how it’s done.Vibhu [01:39:46]: It’s great.Alex Zhang [01:39:46]: Oh, no, I know. Yeah.Swyx [01:39:47]: So I interviewed Yitai who did the IMO thing. He’s just likeAlex Zhang [01:39:50]: Oh, yes. Oh, yeah.Swyx [01:39:50]: “I’ve never, I’ve never been to IMO. I don’t even know what it is.” It’s a skill model, dude.Vibhu [01:39:55]: Yeah.Swyx [01:39:56]: Which is, like, very disrespectful, but like, whatever. But that’s why, yeah.Vibhu [01:40:00]: Yeah. At some point, like, you have to respect, like, okay, the progress is being made.Vibhu [01:40:05]: Like, number is getting output, right?Alex Zhang [01:40:07]: True. Yeah, true.Swyx [01:40:08]: Yeah. that is a very big lesson. It’s very interesting, like, ‘cause the mathematicians are responding this way to NARI systems right now.Alex Zhang [01:40:13]: Right. Yeah.Swyx [01:40:14]: Like, Terry Turnstow is likeVibhu [01:40:15]: TurnstowSwyx [01:40:16]: “No, like, let’s not, let’s not use AI.”Alex Zhang [01:40:17]: Yeah.Swyx [01:40:17]: I’m like, “Mm, I don’t know.”Alex Zhang [01:40:19]: Well, yeah. I think that whole thing is kind of weird ‘cause I feel like, I feel like they would have had a stronger case if a lot of them didn’t work with OpenAI before, like, all this happened.Swyx [01:40:30]: No, that’s ad hominem, and they’re really trying to stay away from that. Then so what, right?Alex Zhang [01:40:34]: So what? Yeah.Swyx [01:40:35]: Like, I don’t know. Like, so what? They got. they’ve, they’ve collaborated. I collaborate with people that I don’tAlex Zhang [01:40:39]: I guess it’s trueSwyx [01:40:39]: I don’t agree with orClosing: Research, Academia, and What Comes NextAlex Zhang [01:40:40]: That’s trueSwyx [01:40:40]: Whatever.Alex Zhang [01:40:41]: Yeah. That’s true.Swyx [01:40:41]: Or, like, I did a thing and then now I regret that. I changed my mind. whatever.Alex Zhang [01:40:45]: Yeah.Swyx [01:40:45]: So I’ll defend their right to say that. But like, yeah, a lot of people are reasonably disagreeing with them.Alex Zhang [01:40:50]: Yeah.Swyx [01:40:51]: Okay, cool. thanks for your joining us. Congrats on, your success so far. I can’t believe you’re still not done with your PhD.Alex Zhang [01:40:58]: Well, it’s year two?Vibhu [01:41:00]: Yeah. Can’t believe we did this podcast without going through the RL paper.Swyx [01:41:05]: He had a, he had a definition.Alex Zhang [01:41:07]: I think the paper is more about, like, empirical results. Like, the actual idea is quite simple.Vibhu [01:41:11]: Yeah. And you’ve talked about it many times.Alex Zhang [01:41:13]: Yeah, at this point. I think there’s, there’s more interesting things to look over now, so.Vibhu [01:41:17]: Cool.Alex Zhang [01:41:17]: Yeah.Vibhu [01:41:18]: Well, we’re excited to see what you do next.Alex Zhang [01:41:19]: Thank you so much. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Transcript
Discussion (0)
All right, we're here in the studio with Alex Jang.
I guess most famously of RLMs,
but you have a few other affiliations.
Welcome to the show.
Yeah, thank you for having it.
Yeah, I guess GPU mode as well.
Yes, GPU mode as well.
You were shepherded in by Mark Saraphn.
Not everyone gets that kind of well-term.
Yep, yeah.
I'm very, very close to all the people in GP mode.
So, yeah, we often end up working together in various capacities,
like even beyond just GPU mode itself.
Yeah.
Can we explain that?
So people who are not that close don't know about this,
It's a Discord.
It used to be focused on, I guess, Kuda mode, and then then generalize a little bit.
It was started by Mark.
Yep.
It was basically like, to me it's like the hiring pipeline of the Pyttorch team.
And then he left Pyttorch.
Yep.
Yep.
Yep.
Yeah.
So it used to be, I think, it started actually around when I was in college in like 2023.
I think it was started by Mark, Andreas, and Jeremy Howard.
The original premise was just like, it was a GPU or is a Discord dedicated to learning how to write GPU kernels and they had like lectures.
That was basically the extent of it.
And I got interested in it because I was writing GPU kernels.
It was actually out of like pure chance.
I was interning at Snapchat at the time.
Rexis.
Yeah, I was very bored with Rexis.
So they had a project where like they were interested in writing.
It was this paper called Infinite Attention.
It was like a Google paper.
Yes, we've covered it on Paper Club.
Yes, yeah.
So I was interested in whether or not you could write specialized kernels for it at Snapchat.
It didn't, nothing really came of it.
But I joined GPU mode at the time.
It was called Kuda Mode, I think for like legal reasons or something.
They changed the name.
But I met Mark.
I met Mattay.
I met a bunch of other people that were very involved in the community.
And then Mark had pitched this idea called Popcorn.
which was now what you see as the leaderboard today.
But the general idea was like,
I think all of us had this like intuition that GPU programming is like very similar to,
if you guys have done like competitive programming.
It's a, it's not, I don't mean to say like they're transferable skills.
You've called golf a little bit.
Yeah, yeah, yeah.
And there's like, there's actually a surprisingly small space of optimizations that people do.
And there's actually not that many kernels per se that people are interested in optimizing.
And so we kind of had this thought.
that like if you had enough data like in the same way that code forces there's millions of problems
if you could do this with GPU code like you could scale and automate kind of GPU kernel
development which for researchers is a huge deal because I think one of the bigger bottlenecks like
if you look like mamba for example like they release the paper with kernels because otherwise
like you can't really use it in any meaningful way you know and not everyone has like a treat
out on their team so we're very interested in this uh kernel bench kind of
spawn from that too of like, can we get LLMs to automate GPU kernel code? And I think that was like,
it was a very, very fun time. It was like between college and my PhD. And yeah, I had a really
pleasant time doing stuff with GPU mode. Now I kind of just help with the lectures sometimes.
I'm not as involved. And I think in general, like, we don't have as many competitions as we used to.
But yeah, I still keep in touch a lot with everyone there.
Is there a friendly rivalry? Because I think the previous community that used to do
this like MLSys, MLPIRF, kind of thing.
It's a friendly rivalry?
Is this like just new generation MLPIRF or what's going on?
The nice thing about GPU mode is that it is also a community in the sense that like a lot
of the lectures are very easy enough for a beginner to follow and ask questions and things
like that.
And like the competitions are like somewhat, not secondary, but like you can participate in
them to learn.
I think with a lot of like MLSys, MLPERF kind of benchmarks, like for the most, you know,
part, like only serious labs and companies participate, like seriously in them at least. That was my
understanding of it. I could be wrong. But I think also beyond GPU mode now, one thing that has
been really exciting is there's a lot more websites and like people that work on hosting competitions.
Like I think there's this, I think there's this website called like Leet GPU or something and it's like
leak code for GPU problems. There's like other ones too that I've seen like there's
many that have kind of spawned and, like, talked on GPU mode.
And, like, it's very exciting in the sense that I think GPU programming used to be super niche.
Like, when I was interested in it.
And the only reason I got interested in it was Triedau gave a talk at Princeton because he was applying for faculty.
I mean, he is faculty there now.
But I listened to his talk on flash attention in, like, 23.
And I was like, wow, this is, like, the coolest thing ever.
And I was like, this is what everyone should be working on.
I guess like VLM and stuff had come out too, and it was like, oh, you know, we should be writing kernels.
But now it's like, you know, everyone writes kernels.
Like everyone, it's, I think it's actually almost saturated in some sense as a field.
Any interesting takes for people that want to get into it?
So I think one of the biggest news is GPT5.6 wrote more efficient colonels.
So Terra and Luna could be 80% cheaper.
Yeah.
And then we've seen other competitions where people are like setting.
records and they're like, we're doing some auto research loop and these are people that don't
have a background in any kernel writing, right? Yeah. So even on the GPU mode leaderboard,
if you look at like a lot of the recent problems, almost all the solutions are AI generated.
However, you'll notice on the leaderboard, so there's this guy named Gow Nurse, who's like a very,
very, like, regular member of GPU mode. We've always known for a long time that he's like a super,
super cracked, like, GPU kernel writer.
One thing we discovered on this leaderboard is, like, almost every, he also used AI to
help him with these solutions, but for the most part, like, he helped prompt and move it in
certain directions.
We found that, like, his kernel was, like, basically the only one in, like, the top
10 that was actually stable in, like, actual, like, end-to-end systems.
And it does bring into question, like, it's not, like, I mean, GPU kernels have a
verification problem.
Like, we've kind of known this.
It's been a problem since Colonel Bench was released.
Like, there's a lot of reward hacking that goes on.
But you all, yeah, you also notice, like, the lines of code is a lot smaller.
Yeah, I was going to ask, is that noticeable or is it just?
Yeah, no, it's definitely, like, very important.
And I think, like, it's really interesting that still there's a lot of alpha in being good at brain GPU kernels.
There definitely is, yeah.
I think, like, and this applies to a lot of AI systems as well.
Like, I think, you know, even with the most recent, like, math proofs.
and stuff. Like, it doesn't necessarily mean mathematicians are obsolete. I mean, these companies
still hire mathematicians, like, to do, whether it be, like, data labeling work or even just, like,
steering the models to solve problems. Like, there is still a lot of alpha in being, like, knowledgeable
in these things. So. Is it just knowledge or is it also there's just more planning? And is there
an emergent style of planning that works better? I think it's, it's a mix of, of, you know,
Maybe this is what you mean, like intuition for how to solve the problems.
Like, for example, I always diagram my code.
Yeah.
Right.
And then, like, if there's a part of the diagram, I don't understand, I work until I understand it.
Otherwise, it's not allowed.
Yeah.
Yeah.
So I think it's like, it's a mix of those things of, like, the people who work, like, the people
who know how to look at these problems and how to solve them, like, also know how to use AI to do them.
Because, like, you're acting as a very strong verifier.
Like, if you are knowledge or if you know what to do and you.
And you're also, like, I think the thing that we've kind of discovered with all these agent swarms and things like this is like when you throw enough compute at a problem, you like can sufficiently explore solutions to that problem.
But oftentimes, like, you know, maybe you can burn like a hundred billion or a trillion tokens on something.
But if you bring in someone who knows something about the problem, they can uncover something for the model that would like erase that one trillion token spent.
I mean, it's not super clear, like, what exactly the trends are here.
But I think, like, there are so many problems in the wild still right now that we want to solve.
And, like, we can't afford to just always, you know, throw as much compute as possible at it.
Like, there is still an efficiency aspect of all of these things that is super, super important.
Is there, like, a theoretical right answer that you can just calculate based on physics?
And then you just get close to the physics limit?
Yes.
So for GPU kernels, you can compute.
It's actually not that easy to compute sometimes, like, depending on how complex the problem is.
Like, for matrix multiplication, it's very easy to compute.
This, like, speed of light kind of estimate of what the fastest kernel can be.
And, like, this is also assuming, like, you know, maybe all of your, all your data starts on the CPU,
or maybe it starts in DRAM, on the GPU, et cetera.
Like, this changes these numbers slightly.
The transfers and all these things.
Yeah.
But I will say, like, it's not clear, though, like, in a lot of cases, if it's even possible to hit this theoretical number, if that makes sense.
Like, this is assuming, like, perfect overlapping and transfer of data and, like, there's maybe some bottleneck that you can't get around.
But often the kernels are not even close, like, that we write, are not nearly close enough to this number to be, like, meaningful at all.
Yeah.
And is it speed that matters?
Do you also care about, obviously, memory, which feeds into speed?
Do you care about power consumption?
So one of our top pods of the year was Jeff Dean,
who was like, actually, I just tracked the microjoules or like the nanojoules,
picojoules.
Yeah, it's often picojoules.
Do you care about that?
So I don't.
I guess maybe I'm not as...
But everything here is speed, right?
Everybody's here is speed, right? Nobody's counting pico's jewels.
But there's a caveat here, which is, I think, like,
there is speed in the context of a single kernel,
and there's speed in the context of a large...
problem, like maybe the N10 model, because like one thing to consider, and this is why it's
important to talk about what speed of light is referring to, because in these cases for the
kernels, like, we always start with everything in, like, HBM, for example, right?
But you can imagine that, like, an end-to-end model, what you might want to do between two
layers is like you might sacrifice the speed of the first operation to keep things in the cash
for the second operation. And like these are things that like you can't really get out of
isolation with like these kinds of kernels. And you know, people call this like the fusion
problem or like the mega-cernel problem. Yeah, or like mega-cernal stuff. And it generally only
applies like when you are like memory bound in most cases. But this is something that like also there
there is this question of like, as these models get better, like, should we just be generating
like megachronnels? Is that like what we want? What's your take? I think that this is really
difficult because you need the data to do this. And I think like, I have yet to see an example
in the wild of like we bootstrap the ability to solve a very difficult class of problems
without any examples.
And I think, like, the other reason why I think maybe this isn't that interesting is that at the level of an individual kernel, A, like, they're not that, they're not as complex, but B, you're somewhat confident that there's not as much structure in a single kernel.
But, like, in a mega kernel, like, I would be more inclined to believe that, like, a compiler would be better here.
Like, some compiler over, like, higher level ops makes sense.
because in general, actually, I think mega kernels are very, like,
the pieces are very composable of, like, the individual kernels.
There's some areas where you might want to do, like, weird fusions and everything,
but in general, I think these are cases that, like, a compiler can probably handle.
And there is a company that's working on this from what I understand
that has given some talks on GPU mode as well.
Yeah, I want to basically cluster all the GPU mode discussions here
because obviously there's other parts that we need to move on to.
I think there is something to plug.
You guys do host a lot of really good lectures.
they're all on YouTube, people can follow.
And you lead quite a bit of it?
You're still quite involved?
I used to.
Sometimes I still do.
I think they're mostly Mark.
Mark is the one who usually does them.
Matei does sometimes as well.
But yeah, I highly recommend them.
They are extremely good resources.
Like, I think it's kind of crazy how much people share on there.
So, yeah.
Because, like, if you're there, like, you're exactly the right audience, you know?
Yes, exactly.
This isn't going to reach the mainstores.
And there's a lot of, like, introductory material as well.
well that we've put on that I think is useful for people.
Colonel Bench was kind of influential.
I just want to see, like, you know, that was last year.
Yep. Yep.
What other ongoing work do you want to shout out that people should pay attention to
because obviously you're involved in this field?
Yeah, I will give maybe the background story of like, I am actually involved in a lot of
benchmarks where I used to be, maybe prior to my PhD.
It started because I was at Princeton.
I worked with the Sweet Bench team there.
John and Carlos and Ophir.
They're all great.
Cartake.
I love them.
There's basically this, I think people don't understand how much benchmarks come from the same group.
Yeah.
At Princeton.
It is crazy.
Do you know Shun You?
Yes.
Yeah.
We had it on a pod before.
Now he's like running tensing.
Yeah.
Now he's like a superstar.
When I met him, so he was advising my friend Michael, Michael Tang, who is now an anthropic.
But they work together a lot.
We were like the two undergrads in Carthic's lab.
And then some others joined later as well.
But, yeah, Shenu is great.
I did not know he was, like, such a superstar until, like, later on, like, after I left.
But, yeah, I mean, like, okay, so there are very few PhD students.
Like, yours is, like, the next one.
Like, once a year, we feature someone, like, who is, like, basically entire PhD has been, like, on target.
There's not that many of them.
Shannu was, like, clearly one of them.
And, you know, Jack Morris is another one.
And, like, you know, we talked, before the show, we talked about research tastes, right?
Like, somehow, some grad students just have a very blessed.
career where like, yeah, mostly like, yep, yep, yep, yep, yep, this is like going to stick
around relevant, everyone should know this. Yeah. And then others, just nothing. I think this is also
true of like even people within like industry labs as well. I think it's just like grad students are
a lot more visible. So you just see like, you know, there are some people who you can publish who really
like get lucky in like or I mean, it's a mix of being lucky and also being very smart and things like that.
I think like with with research tastes as well like I think it it gets it gets developed through
opportunities at least in my case like I got I was very fortunate to have like taken the path that I
took like working at Princeton and then like later like finding my advice like Omar at MIT like he's
a fantastic advisor I will say though I find that the most successful research from grad students or like
in academia comes when people care about problems that maybe like
most people in industry are not looking at. I think this is the issue that, like, a lot of grad
students work on things that benefit, like, that look good to an industry lab. Like, for example,
they'll work on some, like, some benchmark. It's really popular now. I mean, I think benchmarks in
23 were a very different story than benchmarks now. There are a lot of people that work on, like,
harnesses and, like, meta-harnesses and, like, specific harnesses for XYZ task. And when you really think
about it. The reason someone would work on this is like maybe there's like a clear goal shaped around
the models that we have today of like, this is what I want to see. But like I'll give, I'll give the,
like the RLM, like the recursive language model paper as an example because I think like it's a
super, super simple idea. I think when it came out as well, like there were a lot of people that were like,
when they see something like that, they're like, what is even the purpose of this? Or like, yeah,
like why, like what this is just subagents or something, right?
And I think it's like when you get a reaction like that, it's almost like a good sign in the sense that like it's clear that people aren't thinking about what the purpose of this is.
And I will give another example of like Sweet Bench.
When Swimch came out, I mean, Ophir loves to tell this story.
When it came out, like nobody cared.
Like everybody was like, this is an impossible task.
Like, why would we ever even consider this as a benchmark?
And it wasn't until Devin came out that everyone was like, whoa, like this is something we want to hill climb.
And I think this is, this, this rings true for, you tend to see that a lot of ideas.
I think like the, my favorite, I guess, example of this is Eric Selkman's work with like Star and like Quiet Star.
Like I think when you read the paper, at least when I first read the paper, I was like, is this not like an obvious idea?
Or maybe not.
I don't know.
I mean, I was like, oh, this seems really simple or like chain of thought.
And the same thing.
Or like, should use react.
It's like, okay, like, yeah, I mean, sure.
But then, like, when you really think about it, it's like, why, what is the value of the paper?
I think it comes from, like, it tells a bit of a story as to, like, what you want the field to look like.
And that is something that it's very hard to do this in academia because if you look at all these papers,
Quiet Star, React, RLMs, Sweet Bench, none of these papers are, it's not like a GPT6 Astro release.
You know, it's not like, oh my gosh, like, I'm going to use this now.
This is the best thing in the world.
Like, academia just can't afford to do this, at least right now.
I mean, there's a whole slew of reasons why I think that should change.
But I think it's like, if you don't have, as a PhD student, I think you're in such a unique position where you can work on literally whatever you want for the most part.
If you're not taking advantage of that and working on things that, like, nobody cares about or, like, people see as some trivial thing, like, oh, I thought about this, but, like, I don't use it.
I just think, like, in the end, the research is just never going to be that interesting because,
you kind of need to take big bets if you're going to be in academia.
Because otherwise, I think, like, just go to an industry lab.
Like, you know, they have tons of resources, tons of talent.
Why constrain yourself in an area where you don't have a lot of resources
and, like, there's not even that many people around?
And I think it's just, it literally just comes down to, like, big bets.
But you just have to take big bets and, like, a lot of them will fail, you know?
Like, that's just, it's natural.
But I think that is, as a PhD student,
Like, that's the biggest advantage you have over any single person at another lab because you don't have to deal with bureaucracy and all these other things.
Fair enough.
Yeah?
I ask a lot of people this question and usually they hand wave away.
So I think I appreciate that you're actually giving a thoughtful response on like, no, like, this is your unfair advantage because everything else is biased against you, basically.
Yeah, yeah, exactly.
And so, like, it's honestly, I will bring up Jeff as an example because it's not an academic project.
I want to bring this up because this also happened with RLMs and it happens with many other works, like things get overhyped, right?
To an extent, like something gets overhyped and then people are like, why is this overhyped?
Like, this is trivial.
This is stupid.
And I saw the same thing with Jev because I think the release was like, I mean, there is this whole thing about like academics or they're not an academic group, but like people have to do branding and they have to like kind of market their research and so like I understand.
But I think there is a lot of discourse about Jev just being like something we've known for years.
And I think it's kind of missing the point of like, why is such a system so interesting?
It's why is it not just some stupid MLP classifier that like we've we've been doing, you know, back in our intro ML classes or something?
I think what's really interesting about Jev is that it kind of opens up this question of, are language models correct?
Like, in the form that they're in, can we consider a different design space other than text to text?
What they're doing is they're basically saying, like, I will take advantage of this language model backbone.
Like, I know it captures a lot of information about language, but I'm going to change the output space of the model to give you a tradeoff, which is I will do very, very fast inference over, like, if you have some prior about this problem, like, let's say I only need to make a binary class.
classification. Am I going to ask my language model to do this and pay like a 400x cost? No, it's like
silly, right? And I think for the longest time, because the labs are the only places that control,
you're never going to use something other than like GPD6 or Fable because they're the best models.
But because of that, like people have gotten kind of accustomed to this idea that a language model
is just a auto-regressive decoder. Like we have accepted this. And I think when
RLMs came out, it was the same thing. Like, one of the comments, like, a very frequent criticism I got was like, this is not a language model. Or like, when I look at this, like, I thought it was a new architecture, but it's actually not. And my response to that is like, well, a language model is just modeling language. It doesn't have to be this transformer decoder, you know? And Jeb is really, really interesting in that, like, we now have a new thing to tune, which is like, what is the output space and how to
does this affect inference latency? And I think we can actually start asking this about various parts
of the language model itself. We are seeing this too with like loop transformers. It's a similar
idea of a lot of the attention around it was like, this is a silly idea. Like why, who cares about this?
But it's like, it is a simple idea, but it's actually, it opens up a whole new set of questions
that I think like, especially if you're a PhD student, these are the things that you want to answer.
Because I think it's like, we don't know, for Jev, for example, we don't know how far we can take this.
For loop transformers, we also don't know how far we can take this.
What if you loop only a subset of the model?
What if you route to like only, like you have some router to different parts of the model?
Like, can you mimic what you do in a harness inside of the model?
What can you bridge between a harness choice and a model architecture choice?
These are all questions, I think, that open up with works like this, and that's, like, really exciting.
Jeb in particular, when I saw it, I was like, this is actually really useful for RLMs.
Like, I think it's, it makes sense because the biggest bottleneck in RLMs or swarms or systems like these is, they're slow.
When you do multiple language model calls all the time, you're not distributing your compute correctly because, like, maybe there's something trivial that you just want a simple model to do, but you can't do it because your language model is just this bulky thing, you know?
So I'm very excited.
I think we will start to see new types of models emerge beyond just the Boggs standard frontier model.
And that is like, there's so many things that you can do with these like new tradeoffs.
I also shout out, Thanky, with their interaction models.
Yes, yeah, yeah, another great example.
Yeah, so like basically try to break the paradigm from sequence to sequence, decoder only, and just literally do anything else.
Yeah.
I mean, I think there's a level of, if you're trying to compete,
you're not going to compete with a frontier lab doing an auto-aggressive decoderone.
Like, the amount of compute-scaling resources they have, even Thinky will not, you know,
okay, there may be one of the handful that can, but you're not really going to do much in that at least.
I don't know too much about Thinky or I don't want to say anything either,
but it's like if their strategy is just to replicate opening I or Anthropic, like that's a horrible strategy.
because, like, I mean, they just don't have, like, you kind of just have to think of it in terms of, like, what advantage do you have?
And if you're going to use the same setup, I'm sure they're not, but it's like if you're going to do the same setup, like, you're basically competing on the things you can control, which is data and compute.
And obviously, they cannot compete with the frontier labs on that.
So, yeah, it makes sense that, like, if you are a neelab, like, actually, I don't know if you would consider them to be a neilab, but I guess like.
That's what they're in there.
I guess they're kind of a weird one.
Yeah.
In that list, they ship inkling, like, they can't.
Anything other than open-air, anthropic, maybe like, meta and GDM, like, you just, you got to do something else, you know?
Like, it's just, it's the sad reality, but I think, I actually think it's a good thing.
I'm very happy that, like, scale works and these companies will just keep doing it because it opens up, like, potentially new players, like, if they uncover something, really, really interesting.
because I sort of have my doubts that this is like seriously going on at Frontier Labs
because it's like, why would you do that?
You know, like, why would you take the risk of allocating a large amount of compute
to new bets when the old bet already works?
And I think that's what spins off a lot of Neo Labs, right?
You have a side bet and you don't get compute and you're like, okay, I'll go.
I'll go to it.
And, you know, your example of the potential upside is something like Jev, which is
X hundred times cheaper, you know, comes out and maybe it is language model.
Well, in this case, it's just different.
Yeah, yeah, yeah.
Yeah.
So no speculation on what Jeff actually is.
I guess I have some guesses for what it might be.
I have seen some people say like, oh, it's like a diffusion thing.
I guess that...
Which is the parallel decode.
Yeah, parallel decode.
Honestly, I think regardless of what it actually is, because I think you can, I've seen some
like open source replications of it, what is really exciting to me about what they did is I'm
not entirely sure.
what their optimization objective was and how they trained it. And I think, like, this is a thing for
RLMs that we've also been thinking about, which is like, okay, like RLMs are a very simple idea.
If I come out with this paper, like anyone can use it now. But what distinguishes the actual
value of an RLM is whether or not you can train it properly and whether or not maybe you can mold
some architecture around this system to make it really, really good. And that's something that, like,
I mean, I'm actively working on, I guess.
But I think for them, like, they figured out a way to train the system, which is completely
non-trivial.
Like, I actually don't really know how they did it.
And I've seen some comparisons online of, you know, some people are claiming they used
Quinn or they post-trained on top of Quinn.
But every open source quen that you use is going to be worse.
Because whatever they did to train it clearly works very well.
And so that's very exciting.
There's one element of calibration, which is a rare topic that I think.
don't think people even knew about or understood.
We covered it with our conversation with Clementine Forre of Hugging Face and she used to run
the evels at Hugging Face, which is basically the idea that, you know, models are attuned
to give you the most likely next token, but they're going to lie to you when you ask them,
how confident are you?
Because they're just going to give you the most likely next answer.
Instead of like, actually, like, no, let's calibrate.
Like, I am actually 50% sure.
I am 20% sure.
and like let's try to calibrate that.
I would say like if anything,
I think that actually that's pretty easy
to generate synthetic data around
because you can sort of see the truth
and then synthetically generate a bunch of answers,
have Jeff classify it and then compare it with ground truth.
That would be my reverse engineering.
I've actually, I think calibration is probably the under,
like people are just using it as a very fast classifier,
but they're actually not even using the probability
or calibration estimates.
Yeah, I think it's also still just misunderstood
to reiterate.
When you ask a model, how confident are you?
It will spew out 43%.
The big delta is this is a grounded classification, right?
Yeah, I'm very excited to see what people do with this model.
I mean, is it going to solve everything?
Like, no, of course not.
But I think it solves a class of problems that we traditionally struggled with,
which is low-latency things.
So I love the examples with games.
That's actually like, I mean, I...
Yeah, the Doom example.
Very good.
I have a benchmark on language models playing video games.
I've always been fascinated by whether you can have an intelligent system, play new games and things of this nature.
And so I think it's really, really cool that they have sort of a unique way to do this to capture language and understanding in like a fast, a very, very fast model.
Well, while you're on the topic, anything you want to point out video games.
Oh, yeah, yeah.
I mean, you know, you did do a benchmark on any.
Any friend of you, yeah, yeah.
I guess these numbers are very outdated,
because a lot of the models are very different.
And I've seen actually people run,
there are some folks out there that are running newer models on these games,
which is really cool.
I guess the general premise of this benchmark was we just want to see
if vision language models are good enough at just, like,
plugging into games with the latency constraint included.
Because this actually, I came out with this right after Claude plays Pokemon came out.
So this was like two years ago, which I guess is like ancient now.
But I think what's really, really cool about this suite of tasks, it's very diverse in terms of what games they are.
And also, I think most of the games are games that people know or like have seen before.
I saw, yeah, Jev playing Doom.
I will say I don't think, I think they were just playing like really simple levels and stuff.
But honestly, like most models still can't really do or I don't actually think any models can solve these games.
very meaningfully. Like, there are some games that they can. I think I've seen Astra be able to solve
the Kirby game. And we also, for this benchmark, we intentionally designed a really minimal
harness. And I will get to this point about harnesses because I think there's a whole conversation
to be had about, like, what is the value of a harness? Like, what is even the purpose of a harness
for a model? But in general, like, I think it's, yeah, I hope to see very quickly or very soon,
like, all of these games be in by newer models.
Yeah, it's interesting.
Like the old, old Cloud Place Pokemon, they like red state from RAM and saw what tiles are walkable and whatnot.
We did a podcast with them a long, long time ago.
Gotcha.
Yeah, I mean, and it's similar, right?
Like, Jeff doesn't have vision, so you have to kind of feed in these, like, game state and all these things.
I mean, let's go right into the harness stuff because you brought it up.
Language model harnesses a compositional generalizes.
Yes.
To read that one.
Yes.
Okay.
Okay.
So I have been a little unsatisfied maybe with how people think about harnesses.
Because people compare like, oh, like I love Cloud Code.
I love Codex.
I love Pi.
Like, no, I love Oh, my Pie.
I love Prime Agent.
To be honest, I think all of them are the same.
Most of the design decisions or like the design choices around these harnesses are the same.
Maybe Prime Agent is a little bit different because it's like inherently an RLM.
But in general, like, I think we can be a lot more creative.
with harnesses. And what I mean by that is, if we think about this from the perspective of,
what exactly is the harness doing for the model? Well, basically, when you're trying to solve
a problem and you want to use a language model to solve it, like a very difficult task,
one thing that we have discovered is that next token prediction is a really, really awkward form
to do a lot of these tasks. So, for example, take SweetBench. When you're navigating a code base,
Like, are you going to be able to figure out how to do all of this with a single language model call?
Like, you just say, solve code, or, like, solve my query over this code base.
No.
So we rely on a harness to help you do these kinds of things.
And I think what is interesting is, like, a harness is a very, very opinionated program over how you want a language model to be form fit over a problem.
And I bring this up because when we think about what a harness is doing, we should really think about, like, what choices in the harness let the language model solve this task?
And can I actually just have a language model that just does this?
Because a harness, if you think about it, now that looped transformers are a thing, I think what's really interesting about it is you can model a looped transformer in some ways, like with a harness as well, right?
You're just looping over the model.
Now you can say like, oh, I'm not decoding,
so it's like a little bit different.
But in general, we for whatever reason,
have stuck with the same model architecture choice forever.
And there's many arguments for why.
But clearly, like, we are now training models
around harness rollouts.
And so there is this very awkward way of doing training over harnesses,
which is that we train a language model
to act within a harness.
But like, now it's like a really long,
maybe like multiple agent rollout that we're doing.
And there's like really awkward hacky ways of doing this.
So what this blog talks about is like, well, one way you can think about what is going on here
is if the harness is basically helping the model solve a particular task, can different
harness design choices actually do something a little bit more meaningful beyond just here are
some tool calls that will help you?
Here is a way to grep through your codebase.
And so this actually, the idea for this blog came with the RLM idea as well.
We just didn't package it that way.
And I think this is actually true.
There are many other ideas around RLMs that like we will be coming out with, but
we're all there from the beginning.
I mean, these are all designed decisions around.
I think with what is, what I like about the RLM is that there were many, many, many
iterations and versions of different abstractions that I was interested in doing.
And ultimately, the RLM made the most sense.
But there's a lot of reasons that aren't public as to why that's the case.
I mean, you will see.
So in this example, one of the things that we see with an RLM is that if you sufficiently
offload context and ask the model to write code over that context, you get this really, really
weird but useful property, which is that when the model recognizes during training how to
solve a task, it turns out that the solution to many tasks is very similar across.
like tasks where you don't even, like, it's not even that clear to you that the solutions are similar.
So in this example, we have like a retrieval task and we have like an aggregation task.
And they're very, very different query.
Like the domain is just completely different.
And when you train a regular language model over these two tasks, the trajectories look
very different.
And so what you're relying when you, like, let's say you use pie or cloud code or something,
which is not in this blog, but we do have these results, you'll find that like these harnesses,
distinguish too much between these problems, even though the solutions are the same.
And so one thing that we find when training RLMs is that, like, when it sees these problems
as the same, and the reason it sees these problems as the same is the sub-agent sees different
problems, but the sub-agent is solving an easier subtask, and so you're confident it's
smart enough to do it.
But for the base overall strategy, they end up looking the same.
And so when you train on the left task, for example, it can immediately solve the right task.
And so if you go down to like the plots that we have, one thing you'll kind of observe is that the, as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks because the strategy is basically the same.
You're just modifying like a length variable.
And this actually also holds for tasks that are different.
And it's not even, they're not different across length.
They're completely different tasks, math tasks versus writing tasks.
But the solution, the like meta, high level solution is the same.
And so when you train the RLM on one of them, it generalizes this behavior to the second one.
And there's no magic here.
I guess maybe that's the thing that I want to kind of stress.
How would you kind of verbalize what they're learning?
So I think in here you say you train on short tasks, they generalize the stuff 8 to 30 times longer.
They're learning how to solve these type of problems.
Or what's the core thing they're actually learning?
Yeah.
They're learning how to solve these types of problems at a certain length.
And it turns out that when you take this strategy that they learned, it is directly transferable to the longer length.
Like, they're effectively the same program.
And this is something that is really exciting because what this kind of implies is that when you have a corpus of data or environments that you train your model on, the hope is that, like, A, you can train on less environments and,
generalize some more than what existing models can do through
harnesses through existing kind of like naive training.
But B, also, you still want to use all the data that you have.
So when you train on these tasks, like hopefully it generalizes to a wider class of
problems.
And why this is also even more exciting, at least in the context of RLMs or any recursively
calling system, is this argument holds inductively.
So like, I'll give an example, because I said, I claimed that competitive programming
and GPU optimization use very similar skill sets.
The model can, the harness potentially,
I mean, you might have to nudge it a certain way,
but it can learn that like,
okay, how I'm going to go about solving
this GPU programming task
is very similar to what I learned
for competitive programming.
So I'm going to list out a set of solutions.
I'll like spawn subagents to list out promising solutions
and then I'll like write this loop
to go through and check these solutions,
maybe evolve them,
and like evolve them against a verifier.
And between these two tasks, this looks the same.
But what the sub-agents are doing are maybe unique
and something like different.
But even what the sub-agents are solving
might actually also be of the same form, right?
Because it's like a recursive argument.
And so what I'm trying to get at with this whole blog post
is just that like we should rethink
what the role of the harness is.
Because harnesses can actually greatly increase
the generalization capability of your model
and the amount of data that it's given.
And this is not exclusive to RLons.
I think there is a wide class of harnesses
that are yet to be discovered
that actually can yield similar properties.
And to extend this argument a little bit further,
I think what you can also extrapolate from this
is like, if I look at an RLM,
what are the components of an RLM
that are actually necessary?
And can I actually just directly train a model
to do this?
Can I train a model to act as an
RLM implicitly in its forward pass.
It's a really weird thing to think about because, like, you know, you might say like,
oh, code is non-differential, blah, blah, blah, blah, blah.
But there are many approximations of this behavior that we will start to uncover.
And I think, like, we will see beyond just like I'm going to design a new coding harness
that uses a special form of compaction or something.
I think we can be a lot more creative here.
Like, there's so much we can do with these language models that I think we are
are just not doing. And I'm like very excited about this because I think like I think we can get
serious gains from very, very opinionated and good harness design that lends itself better to
scale. And what I mean by this is like, you know, the RLM for example is a very primitive,
inductive bias. Like there's nothing super special about the design other than the fact that it's
very different than what we currently do. But this may potentially scale much better with like
the data and environments that we have available to us.
I guess the opposite thing that people would probably ask is
current harnesses are very generalized towards coding,
which people see works for a lot of domains.
Cloud code is being used for design presentations, everything.
Muse Spark, Grockbot, these are very simple,
non-opinionated harnesses that are good at code,
and that is also scaling out.
What's the example of how we improve
dollars, I guess.
Yeah.
Let me bring up another paper, which came out very recently.
It's like the harness tax paper.
I think it's by Arena.
I really like this paper because it puts forward a prior that I had, which is basically
that like most.
It confirms a prior.
Yeah.
Most harness choices don't matter because all of these harnesses are the same.
But I also like Grockbot, for example, is actually.
quite different, I think, from my understanding than how some of these other harnesses have been
designed. And I like that a lot. And I think it's clear from here, at least, that, like,
I'm pretty sure Anthropic or Open AI are exclusively training on their harnesses. They're
probably not training on their competitors' harness. I mean, I'd assume not because I don't know
why they would do that. But, I mean, you know, this is a thing you see in Open models, right? Like,
Quinn is really good at using OpenCode. They need to train in harnesses. Old Gemmas were notoriously bad at this.
models are good, but you need to train in A harness.
I think, though, as models get smarter or, like, as they get better, this distinction
becomes not that important in the sense that, like, if you take Astra and you put it
inside of open code, like, it's not going to go crazy because I think it's, like, just sufficiently
good.
And so I say this because the only benefits between these different harnesses is just cost,
for the most part.
And I think, like, what I'm getting at Higger is if you plug.
these models into RLMs, though, they're not that good still. They're okay. And I think it's mainly
because the types, like the class of harness that we are training around is this class of harness,
this like pie loop, this like, I like to call it trajectory as a prompt, which just means like
you keep the whole trajectory of the rollout as the context that your main model is using.
Even if you use sub-agents, it's still like kind of this form. And I think we're going to,
If we want to explore new harnesses, there needs to be teams that are dedicated to actually
running meaningful experiments over like scaling out new harnesses, like maybe post-trained
scaling out on different harness designs.
Like I think we can actually get very meaningful knowledge or gains from doing this kind
of thing, whether it's an RLM or whether it's something different.
And that's exciting because I think, for example, if you train a lot on Fable for a long time
was the best model for RLMs because they had dynamic workflows.
And it was pretty obvious that, like,
this was a capability that was somewhat trained in.
Even if the model was still, like, a little dumb, like in the RLM harness,
it still worked a lot better than other models did.
Astra is now also, like, good enough at doing these things.
But is this where we see that Fable is the best for RLMs?
Or is there some other?
Oh, no, these are all internal results, I guess.
Yeah, I don't have them public right now.
But in general, I think you can very easily tell that we have not optimized for RLM like workflows yet.
And I think if we get models that do this correctly, they will be a lot more efficient as well.
I mean, you can kind of just think through like why this is the case, right?
I think this is the point where you have to give the 10 second what are RLLLF.
Oh, yes.
Because there's a lot of listeners here that are assuming a lot of knowledge.
Also, I think, but you have set some context.
So you can, like, with everything we just said, can we have a clean, crisp definition of RLM?
Yes.
Okay.
I want to go back to the blog, the compositional generalizer's blog.
This one.
Okay.
This is like the best.
I wish I had this in the original paper.
An RLM is basically just a harness design where the only tool in the harness is code,
which is this programmatic sub-agent calling thing, where it has the object.
option to call itself as a tool, and it has other tools. But all of these things are functions
in code. And the context that it's dealing with is always stored in some memory inside of this
code environment. So this could be a file system. Like this could be like, I'll give an example,
Prime Agent. The trajectory of Prime Agent, like the context, even the when you compact and do all
these things, is stored on disk. So the model can always reference its original context, even if it's
compacted, and all of its tools are run inside of, let's say, like a Python repel or a bash
repel. And so it's like this very primitive abstraction. And I would say like where most
harnesses differ is a context offloading is not done that like that. I mean, if you look at
Prime Agent, by the way, Prime Agent does context offloading, but not all the way. So like it still
maintains the standard codcode code codex loop of like trajectory as a prompt where you compact,
but it has the additional kind of like the context is offloaded and it only has the unique
point of Prime Agent is that the only tool is iPython. So this is like the very kind of
generic abstraction around like RLMs. Concretely, what's the core thing our LMs are trying to
solve? At one point I think when it first came out it was context. I'll pass the question.
So the original problem was harnesses have a really bad time dealing with long context.
They typically were used to only really do them for like specific things like code.
Like, you know, they could deal with your code base because it was trained on it.
But now it's more around what this blog is talking about, which is compositionality.
And the fact that like I think harnesses, we want to have language model systems that have much more
control over the actions they make at every step. And what I mean by this is tool calls are very
limited because you have to invoke them every turn. Like you have to invoke tool A, then tool B,
then tool C. And there's no central context that you can kind of draw back from. And RLMs are
specifically like designed around composition and having like a central context that you can
always draw from. And this context is like designed around the existing language.
models. Another, like, very similar example, actually in design is, like, agent swarms,
for example, the hugging face incident. Like, these agent swarms have, like, a message board
that they learn to communicate over. And this message board, in some sense, is the shared
context that they, like, act over. And RLMs basically say that, like, the best way to communicate
through this is in code. Like, you write the code to do this. And it's because these models are so
good at writing code.
Like we want to take advantage of that fact.
And then for this compositional thing, up to when including generating your own harness,
you know, specific for the task.
Exactly, exactly.
I think we will start to see that if you go up to this figure.
Okay.
So we talk about this idea of locally in distribution tasks for a harness.
And it is like an idea on top of like in distribution tasks when we think about,
language models, an indistribution task is just a task where, like, the prompt is something that
the model has either seen before or, like, has seen some version of it. Most harnesses work,
like the one on the right, which is they keep appending the trajectory as a prompt. And so eventually,
unless you're anthropic or open AI and you train on, like, kind of these, like, user trajectories,
most of these things end up being out of distribution for the most part. But locally in distribution
is basically the compositional argument of if an RLM breaks down its computation, it's computation,
into like a kind of a meta harness of sorts
or like a program that involves subagents
that like look at a local problem,
every individual language model call
over the course of this task is in distribution,
even if the entire task itself is out of distribution.
And this is a very desirable property.
I think for obvious reasons,
like if every task is in distribution
for each individual language model call,
you will probably get to the right answer.
So, I mean, the logical limit
of RLMs is RLMs, where like, you're not just, you don't just write the harness,
you also train a custom model.
Exactly.
You collect data, everything.
Like it's a fully automated AI researcher inside of your harness.
Yeah.
We will see where the training of RLMs goes.
I will say as an academic, I am not working on this at MIT, or at least in the scaled sense,
because I can't afford to.
But there are companies out there that are working on this.
I think Prime and Select is very clearly working on.
this and it's very cool like I'm very excited to see maybe we'll observe I don't know but maybe we'll
observe better like post-training scaling walls with when you train around a smart harness maybe
we'll even see smarter harnesses that come out and like they they work better around these kinds
of principles what is a smarter harness like that doesn't mean anything you just did it all the same
no more of what I mean is like Claude Codex pie etc are all the same in that like when you break down
the logic of the harness,
it's like virtually the same thing.
Yeah, two goals in the loop.
Yeah, whatever.
But with RLMs and with other harness abstractions,
it looks very different.
And this is where I think you really distinguish.
I mean, it's in the same way that like,
I think with language model architecture choices,
a lot of architecture choices end up kind of looking the same
when you like scale it out or like it.
The differences end up being like somewhat minor
in terms of, I mean, for a lab, it's not minor, but, you know, maybe one model converges better than the other one, like, slightly.
But in general, like, if you were, so for example, pre-training scaling laws only hold because the architecture choices we have are somewhat stable right now.
But if you were to completely change the architecture, pre-training scaling walls probably don't hold.
Or, like, these kinds of, this, like power law is going to look very different.
And it's like a same thing with harnesses.
Like, I think all the harnesses we have right now, for the most part, roughly look the same.
same. But there are some exceptions to this, I think, that are coming out. I was going to say,
actually, one of the things that I've been more interested by, you know, talking about PhD students
who take big risk, is that people have been, people also pursuing the other side, which is pre-training
scaling laws don't hold if you change data. Right now, it's just raw and structured text,
a corpus of internet. What if you had a better data representation to train on? Then, yes, your scaling
law would change as well. Yeah. So there's architecture, there's data, and, you know, whatever else.
can think about. Well, so I just want to get back to this. It's all makes sense. It's very interesting
how you sort of recurs up and down the stack, from like very conceptual to like, no, like,
well, this is where we are today. But like, yeah, obviously you can scale up and down. I guess,
I'm curious, how did you start working with Prime? Is Prime taking on more work with this?
Is this their answer to Hermes agents? You mentioned Grogbaugh. It's a little bit different.
I just wanted to like name check all these guys and get your thoughts on each.
Yeah, yeah. I got involved with Prime after they released a blog post, by the way, not affiliated with me at all, about like how they believed RLMs were kind of the future. And I had a friend that was working there, GP mode, Mate, like, we got in touch. And I think I agreed with a lot of the researchers there and like what they believed about harness design. Like I was, I was very impressed, I think that like they understood the purpose.
of the Arlen paper, which is not necessarily just to say that, like, you know, we're solving
long context tasks, but actually, like, we want more opinionated harness designs. Yeah. There's always,
like, the result of the paper that you choose to highlight. Yes. Versus the actual point.
Yes. I would love to talk about, like, the incentives of academia and, like, the things around,
like, why it's kind of flawed and all the issues. And we'll get back to that. Yeah. So, so anyways,
You know, I love the guys at Prime.
So we kind of had been, after we decided to work together,
we decided to look into training in RLM
and also build this kind of RLM harness
and kind of see where we can take it.
That is how Prime Agent came about.
And I think the reception for Prime Agent
has been pretty good.
Like the one thing I was worried about with Prime Agent
is that none of them, at least at the time,
when we were building it,
none of the models were that good at doing RLM stuff.
So this was like pre-fable pre-astra.
I think to take a step back, can you explain what prime agent is, how it's different than a traditional, you know, cloud code, what people would expect harness?
Yes.
So Prime Agent, I think I mentioned this a little bit earlier.
Yeah, there's the diagram.
Yeah.
It is basically a harness on top of pie, like pie mono, which is pie mono for context is like the like a minimalist.
Yeah, like harness.
I use pie as the reference for everything because.
I think all other harnesses are basically just pi.
But it is pi, except we explicitly restrict IPython to be the only tool that's available to it.
Every other tool gets loaded in as like a Python module or like a like a bash kind of script that it can run.
So it uses the core RLM abstraction on top of Pi.
And then it also has this continual harness thing, which is Seth.
He's another PhD student.
This is a thing that he used to get language model harnesses to play games.
So he worked a lot with Joel, who is the Gemini Place Pokemon guy.
And continual harness is also, by the way, very simple.
I quite like it.
It basically is this design principle around like what parts of the harness can you let the harness itself modify.
There are certain pieces that you'll let it modify its own skills, the subagents available to it,
what the system prompt to the model is,
and continual harness is available
basically as a tool inside of the IPython kernel.
And so that's what Prime Agent is,
like how, what is designed around.
Everything else in Prime Agent is like standard, right?
Yeah, I think what is what I really liked about,
and we got kind of lucky,
is that like a lot of the new frontier models
actually work really well inside.
And actually, even a lot of the open source models
work really well, at least some of the newer ones.
And there is another thing in prime agent I should highlight, which is that like, we have a very particular agent-to-agent communication system or like framework, which is because RLMs tend to spawn many sub-agents, we want a way for sub-agents to communicate with maybe the root or with each other.
And so there are some design decisions around like what each sub-agent is allowed to talk to, how it does it.
Again, everything is in code.
So it writes the code to do this kind of communication, which I think is really, really cool.
And then there's, I guess, persistent subagents is another thing that was kind of added,
which is subagents, they can last beyond, like, the standard runtime of the actual, like, original agent.
And you can go into that subagent.
You can prompt it more.
Like, you have more visibility and flexibility into what is kind of going on.
This is my number one pain with codex right now.
Their subagents are just very ephemeral.
And they actively discourage you from using it for a long.
learning things. Yeah, yeah, yeah, which I think it makes sense. Yeah. So, I mean, the trick is just
externalized to a file system, right? Yes, yeah. That's the trick. That is the big trick.
Yeah. And also like force everything and run through code, you know, trust the model that can
write code and it's going to write its own harnesses itself. So it's prime going to take on like
training, post-training custom models for it is? Is this a one-off collaboration between you guys?
That's it? Like, what's? Yeah, they are training a model in turn. I think they were, they were pretty
public about this actually back in March.
I mean, clearly it is their business.
Yeah, yeah, yeah.
Yeah, I mean, they're showing that they can train it on their kind of hosted training stack.
But no, so for model training, I'm not involved with them on that.
The main reason is just I have other things in the PhD.
I want to work on.
I think, like, there are many other big bets to take outside of just RLMs.
Some, I don't know how much I can share yet.
But in general, like, I think, I actually think one of the luxuries of being a PhD student genuinely is that there's so
any big bets to take. Most of them will probably yield nothing. But it's a really, really exciting
time to be in research, especially because I think most progress in the field has been a little bit
boring. Like, I'm not saying the outcome is boring, but the process of doing these things tends to
be quite boring. And so there is kind of this question of like, you know, what do we want to do next?
But we can talk about that later. Yeah. Okay, I want to close out a little bit more of your research
and then we can put it in and out.
Since you released RLM,
a lot of excitement about it,
any secondary third-party work
that you want to shout out as like,
that you guys should take a look at this.
Oh, yeah, yeah, yeah.
So Harvey, the legal AI company,
released a blog post,
not affiliated at all,
but they post-trained an RLM
on their like legal work,
which often involves a lot of like sifting through documents
and kind of looking through
like a variety of specific information
that maybe is not so easy to retrieve with like a pure retrieval system and they show like really,
really good results. It's very exciting. I mean, I was, I was shocked that that they worked on this.
They did not tell me. So when this came out, I was like, oh, that's awesome. So there's this one,
I think it's super, super cool. And what they're doing there, I think this is a collaboration with base 10,
by the way, as well. Headlong, which is law, the Laud Institute's kind of, it is their like
persistently running harness. It's very cool. Oh, they're renamed.
it. They used to call it something else.
It was like auto.
Terminous.
Yeah, I know.
They've gone through, yeah.
So this is Andy Knewinski's big project.
It's super, super cool.
I love Andy.
I don't want to downplay what they're doing because they're using the RLM abstraction,
but they're doing something much cooler than the RLM, which is like they have a system that
kind of what they call like thinks persistently.
So even when you don't query it, it has a way to,
think through problems that it has in its context.
Oh, so it's just like an always on type thing.
It's like an always on thing, but it's like not that expensive.
Like they control the token costs to make sure it's not like burning through all your credits.
This is super cool.
I'm trying to think there are many.
Actually, if you go to the RLM GitHub page, there's a bunch of things I've linked below.
There's a ton of really, really cool kind of things that people have been doing.
Ax is another really cool one that I think it's just by this one guy.
It's like a harness around DSPye and RLMs.
DSPi also has an RLM.
Oh, the last thing I'll shout out is on Arc AGI 3, I believe, there were a lot of harnesses on their, like, Kaggle competition, like, the official one.
Not the, like, public prime-inslect one that or, like, what people have evaluated on.
They all, like, claim to use, or they reference, like, Tufa, for example, some form or some
inspired form of the RLN abstraction in their harness, which is really, really cool.
I think it's, this is where, I mean, this is exactly the setting where you would see a lot of
benefits from composition and using code and combining, like, neuro-symbolic systems with,
with AI.
Yeah.
Yeah.
Very, very cool.
We love a good newer symbolic reference.
Yeah.
You also mentioning RKGI3, opening icons on it says we're 99.9% on this.
They also say we solve Neviya Stokes.
We just throw a model at it.
there's some debate around whether or not
it's just model
or are they using an RLM? Do you know?
I mean, I would guess probably not
unless you say, like, I'll be careful here
because, you know, people debate
what is an RLM, what is not an RLM,
it's somewhat clear that what they used
is some kind of swarm of agents
with a shared, some shared context,
like some shared file system.
And like, this is very much in the spirit
of RLM stuff, but I think there's
a lot of like, you know, more
clever things that they did that's not maybe related to the Arlem itself. I agree a little bit with
the idea that a harness is not that necessary for what they did. The way that I would put this
is that I think a model, like a GPC6 astra type thing, is technically smart enough, conditioned
on the right information to come up with a proof for these very, very difficult.
problems. Now, how you get to that information is a giant question mark. And in their case,
it probably came down to, like, a very long search over, like, many of these sub-age,
or many of these, like, agents in the swarm and maybe also, like, researchers can, I'm actually
not sure about this part, but putting in, like, their kind of intuition as to, like, what you
should explore and things like this. And ultimately, like, this produced some information that some agent was
able to take to finish the proof. And so in that sense, like, I think was the harness that important?
No. And I think what this is pointing at is like these specific details of a harness do not really
matter. And I think that's also what that what the harness paper is pointing at, which is that like
beyond the user's feeling of the harness, realistically, all that matters is just like, how are you
composing these agents in a meaningful way to get to the final answer? And maybe that's what like
swarms and all these things are are really about and so from my POV at least if we if we start to think
about like for user use cases what do we want out of harnesses and and and things like this like
we want to take the good parts out of these like the cloud codes the codexes like the stream
that people like to see but like under the hood whatever is running can be some really weird
complicated swarm of agents that like ultimately come up with an answer the user doesn't want to
see that, though, obviously, right? Like, it's not legible information. And so I, you know,
this was another kind of thing in the spirit of RLM's, like, recursive language model. It sounds
like it's a language model, but it's not a language model architecture. But the reason for this is,
like, I think we will start to see in the future probably one day, and I wrote a blog about
this, what we think of as a language model, like the thing that we query might actually be like a swarm
or like a scaffold or like some weird harness design
that scales very well,
but the user just doesn't see it.
Ultimately,
all the user sees is some front-end version of this harness.
And yeah,
I think it's a relatively safe bet,
at least,
to make that this is what we will see.
And this maybe goes back to the limitations
of the base transformer.
Like, you know,
obviously if you just took a base transformer
and you said, like,
solve Navier Stokes or something,
it's not going to do it.
Like, yeah, we all know this is not what's going to happen.
But yeah, I think this is maybe the more interesting part.
And maybe the claims around like, you know, did the harness matter is more around this of like just arbitrarily pointing models like or agents at a growing kind of context of information.
Maybe it's just enough to solve very difficult problems that I can buy.
Well, there's a lot in there.
I do want to say the amount that OpenA has spent is semi-public.
It's 10,000 agents in 88 hours, 130 billion open tokens, which is estimated to be about.
$40 million in public pricing.
Surprisingly actually, like, less than I know.
Yeah, not that much.
Yeah, $130 billion output for the final.
But as you said, there was a lot of context being passed around.
It's more than double that in just the total agent messages being said.
Yeah.
Yeah, yeah.
I think one thing I was honor to bring out also was cursor as far as like swarm stuff is
concerned.
This is slightly older, meaning February, which is ancient.
But if you scroll down all the way to the final sort of multi-agent architecture that
they arrived at. It was basically an org chart
of a normal software team.
One thing I'm thinking about
because I basically, there's like this
gather all
function that you have to do with
subagents or, you know, it's very similar
to GPU programming actually. Yeah, yeah, yeah.
And so
that's a bottleneck.
This is a bottleneck. If there's one main guy
it's coordinating all the subguides
then they have to like, you know, gather
again and then re-coordinate.
That's slow. That's crappy.
what a true swarm should be is everyone is just their own person.
Yeah, yeah.
Right.
Well, I agree with this, and I think that there is a question to be had, though.
Let me give an analogy, which is like, when would you use compaction and when would you use an RLM?
And there are many settings where an RLM can solve maybe a more difficult task than compaction can, but you would prefer to use compaction in most cases because it's cheaper and it's quicker.
And I think in the context of agent swarms, there is a similar thing going on of like, I'm fairly certain that like 95% of the swarm is entirely useless or like what it's exploring is entirely, you're just burning tokens.
Versus in this setup, maybe not so much.
I'm not sure.
I mean, maybe it's also the case here.
Everyone has a job.
There's a jury board.
I think at some level that's how problems are framed, right?
So like if you have a search problem and you're spanning out a bunch of sub-agents to do search, there's a good.
there's going to be a lot of useless information, right?
There's one retrieval answer that you're getting and you're spanning off.
But that is consciously understood, right?
Yes.
But there is kind of this question of, like, what is appropriate to solve for which problems?
Like, what design, you know, in theory, open AI can use, can package up this API and they'll call it swarms.
And then they'll give it to you.
And they'll be like, point this at any problem and we'll give you a solution.
Yeah, it's called pro.
Yeah, maybe you'll have to pay like $40 million to get a result.
And it's like, well, but it's exciting.
I will say, like, it is, it's very exciting that we even have the option to point $40 million at a problem and solve it.
Yes.
But there is, there is still kind of a lot of research to be done in this area around, like, you know, what is necessary.
Like, what do we want to do?
What design do we want?
I mean, we probably don't want everything to be a swarm, but like, where do we draw the line?
Like, can the agent design that or decide that for itself, et cetera, et cetera.
So.
Yeah.
And then just a quick check in case you have opinions on this.
Have you looked much into open-endedness as a general category of problems?
Meaning no problems, just go.
A little bit.
So I was at Sikana for a summer right after, or I guess right before my PhD,
and that's something that they work on a lot there.
And I think there's a lot of people, even at like recursive superintosh, there's many of them now.
Yes, we just had Richard Sletcher on.
Oh, yes.
Yeah.
like Richard's company and then and then and then also actually wait that might be this it might be
the same company I don't remember is Tim Rock to Shell also yeah yeah yes that he's the main co-founder
is being head of open-endedness for Google yeah yeah so I think with open-endedness problems like
I view them as somewhat similar to even like unsolved math problems maybe that's a weird way of
framing it but like I think a lot of the techniques in terms of like how people like approach them
are kind of the same.
Like, evolutionary research is, like, very similar to just launching swarms of agents
and hoping that, like, they come up with, like, an interesting, and this is what, like,
Alf evolved and some of these other works did, like, a year or two ago.
But I think what maybe is not, and I'm not sure if this is what you were alluding to,
but I think what's not super clear in open-endedness style search problems is, do we frame all
them as basically like an unsolved very difficult problem.
Or like where the objective is clear.
If that's not the case, I still don't know yet entirely what the value of it is.
Maybe you have other opinions.
Like I don't have too many opinions on this.
But at least from my time at Saganah, like I got the sense that like we ultimately still
kind of want to approach things the way that like say open AI approach snobier stokes.
We want, there's still a lot of nudging in certain directions that we want to.
to have to like get to the point where we have something interesting.
My version of it is like maybe it's a split between basic signs and applied science.
Basic sign, you're researching for researching sick.
You just want to understand things better.
I have no idea if like there will be any application at all whatsoever.
But I mean that and then applied you have a goal.
You're trying to minimize loss in some way, you know?
And so what I really think, you know, in terms of like the big bets that people have,
what if there was no prompt?
Like,
I see it.
You know what I mean?
You just picked domain.
Like you just like,
you just spawned in this like swarm of things and you're like, hey, what's up guys?
Like what do you're working on?
Yeah.
And like you just decide.
I see.
This is an interesting problem.
The biggest issue that like, maybe there's a clear path to this.
But when I was there at least, the biggest issue that people had with open end and this was like,
how do you ultimately choose at the end?
How do you pick out the interesting stuff?
Because when there is no, like maybe the agent comes up with a goal.
But in a.
lot of cases, like what they, and they have something called Fulu, I think, which is like a, it's like a model router type thing that was like inspired at least by this idea of like, let's pick a problem where maybe we can pick out the best solutions to something. In this case, it's like pick the best model for this problem. You're the first person to connect model routing to open into this. No, no, yeah. So I bring this up because I think like with open endedness, like just generally the issue is like when we have this giant corpus of like slop, like, I'm. I bring this up because I think like with open endedness like just generally the issue is like when we have this giant corpus of like slop, like,
How do we sift through and find, like, the hidden gems?
And, like, the solution to open-endendness really just letting models run forever and, like, finding, like, just doing data gen.
Just doing super, super, super high throughput data generation.
And then, like, asking agents to go through and, like, find meaningful things.
Like, I'm not sure.
Maybe.
Maybe that's sufficient, you know?
Like, that would be cool.
To me, it's very interesting as a counter to basically all the machine learning where you have a goal to have no goal.
Yeah.
Or like an ill-defined goal that you're like, well, what about this goal?
And you're like, well, okay, maybe.
And then you sort of research more and you find that that is an interesting goal.
Because I think like finding your objective function.
Like you said, like Jeff found an objective function that was interesting and no one was exploring.
Yeah.
I think that is like, you know, similar to your message about grad students as well.
Like, you know, you stay in school because you want to pursue what?
open-endedness. If you want to, you know, profit max and like, you know, join the,
escape the permanent underclass, then you join a lab. Right? Right. Yeah. Yeah. It's funny. I,
I feel like I don't, I don't hear this discourse a lot. I mean, I'm in the East Coast. So it's like a very
different type of, but then whenever I come here, it's like, that's always like the topic of
discussion. You cannot pay a rent without doing this. Yeah. We're getting priced out, guys.
Okay, so, so yeah, there's, there's all that. I don't know if you want to, if it's
relevant enough to talk about the mismanaged geniuses, which you were pulling on.
No, no, it's just on your blogs.
But I will poke on a...
Oh, so they did in their blog post.
I talked to them about this as well.
So one of the cool results of Ultra is they basically just let it loose on automated data science research with little to no human intervention.
It's just kind of making...
Yeah, so this is auto research, which is a little bit more open-ended.
And there's degrees of open-endedness.
And I agree with that.
Yeah, separate than auto research with objective.
This is just do stuff.
But, you know, it's cool.
They're working on it for those that are interested.
While you're bringing it up, actually, what is your take on Sakana?
Like, what are they doing, apart from being, you know, the Japan one?
Yeah, yeah, yeah.
I mean, I actually love the people there.
Like, I think they have a really, really smart team.
And it makes sense.
I mean, it branched off from like an earlier team at GDM, which was also kind of, I guess, doing this kind of like
open-end evolutionary research style stuff.
What I liked about my experience there, at least,
was that they did have that, like, mishap back,
I forget, at this point one.
But I think like...
You're talking about AI scientists.
No, the GPU kernel.
Yeah, okay.
That was up.
Yeah, yeah.
People cannot forgive them for that.
Yeah, yeah.
And I guess, like, the AI scientists, like,
I mean, there's some criticism of it that,
I mean, I don't work on that,
so I have no kind of take on it.
But I think in general, like,
what I like about them at least is that they're a little bit more
of a researchy type lab.
So, like, they don't operate in the same space as, like, open AI or anthropic.
Like, for sure.
Like, definitely, no.
They do not, I mean, it's pretty obvious probably that at least when I was there,
they do not have a big competitor model or something that, like, you know, that everyone
is, is using.
But I think they kind of operate in some ways as, like, a PhD lab, which is cool.
Like, and I, I think, like, David Haugh is, like, he's really, really smart.
Like, I think he has a good sense of, like, also, I think the, you.
The market in Japan is also a little bit different for AI.
And, like, who they're targeting is slightly different than maybe what we're used to here.
But, yeah, I like that they take kind of, a lot of their research is kind of weird, I think, when people view it, you know.
And I like that.
Like, I think it's...
You should have more weirdness, yes.
Exactly, exactly.
And if you say different market, just is it like enterprise?
Like, the way it works there is a bit different, like how deals happen and stuff like that.
They do have a, I mean, I guess this page is originally in Japanese, but they do have a model.
specialized for the Japanese
You didn't know that
I didn't know that
I didn't know it too
I also have personal
friends and know the team
so there is
even from the sense
of a way that
you speak culturally
you know responses are tuned
towards that
this is not like
it's frontier on benchmarks
it is a cultural
appropriate model
for them
and then they have like chat
and all that
but to mirror your point
you know
there's also like
how should education
look like
and someone wants to work on it
and they're very PhD
D-Lab of do your thing.
Why not? We have money. Go research.
They say it's a Kimi fine tune. That's nice.
Oh, there you go.
Good for them.
Yeah, I mean, and you know, speaking of Kimi, right?
Like another, you know, just a grad student that split out and like, yeah, I just have
this like Kimi Delta attention that ones that I want to work on.
Yeah.
And like somehow managed to make Moonshot.
No one understand it still.
Yeah, yeah.
Well, he's super cracked.
At least my understanding.
I mean, I think in general, like a lot of, a lot of.
of the Chinese labs have done really, really cool work.
Yeah.
Any thoughts on Kimi agent swarms?
Yes.
One thing I will say is whatever Open AI is doing with their agent swarm is like clearly the right thing to do.
You have to kind of think about it this way.
Like nothing, especially without like a very, very smart harness design, which I don't think anyone really has so far.
Nothing comes so easily for free.
For example, like, the agent swarm design is not something you can just take for granted.
Like, it's not like GPT6 Astra.
It's just super smart.
And then it just got agent swarms running well.
They clearly trained.
I mean, the hugging face incident was them training a system to be like a swarm.
And I think, like, clearly they've done something really well to the point where you can throw $40 million and solve an unsolved problem.
And I think with the Kimmy agent swarm thing, like,
at least from when I read it,
it just came off as like,
this is interesting,
but I don't actually know
whether or not this can solve anything novel.
Yeah, they just,
they were like,
it makes spreadsheets,
they kind of were just like,
yeah,
like,
you know,
here is a swarm
that,
like,
kind of does stuff.
And it's cool.
But,
and I will say,
same thing with dynamic workflows.
I actually kind of think
the dynamic workflows
release was sort of a flop.
I don't know how,
how you guys feel about it,
but I think it's like,
my understanding is,
Like, it's not used that often.
It's just very expensive.
It's too expensive.
It's ultra-code.
It's basically, like, you know, take over my bag.
And it doesn't act.
I've tried it.
And, like, it doesn't act in the way that, like, again, I'm not the biggest open AI, like, you know, Stan or something.
But I think whatever they did was very impressive.
Like, they somehow managed to get away for this swarm to actually act towards a goal.
And, yeah, it's very difficult.
I see.
So, like, efficiency of the multi-agent swarm.
is the objective function here.
Yeah.
Like, how much of this is slopped?
Like, this is a lot of slop.
I think so.
I think we take for granted what it means for a swarm to converge to an answer.
It's just like not something we can take for granted.
Yeah.
We've done one part with Noam Brown and like his whole thing was like, okay, like we've worked
on a lot of like competitive agents.
We're working on collaborative agents.
Yeah.
And like that's now called a swarm.
Yeah.
I would do a quick poke.
Would you have any thoughts on Gemini?
They were also IMO gold.
Like, you know, there was a time where they were getting agents to reason for a long time.
Yeah, I mean.
Is it too old to think about?
No, no, I think it's a little bit blown out of proportion.
Like, Gemini, again, so I should preface by saying, I haven't worked at any of these places.
So take this with a grain of salt.
I mean, you have strong opinions on agent harnesses.
Yeah.
And, you know, this was.
But let me just say this first.
This work was really impressive.
I think what they showed here was like they took a time when the models weren't that good.
And they managed to be very, very smart about what the harness does.
I remember for this at least like, yeah, like alpha geometry.
I guess that was a year before this, but it was very cool.
They took it to the max and they like designed.
I don't know.
I'm not too big on competitive math.
But I think like GDM, it's sort of a shame.
Like, everyone I've talked to about GDM kind of has the same opinion, which is that it's way too, like, bureaucratic.
Whatever is good.
They have the talent and the resources to do almost anything, but like, I don't know, until they figure that part out, like, nothing against anti-gravity, for example.
But, like, I don't know anybody that uses anti-gravity.
So I've tried it once, and it's, I don't see a reason to switch to it.
And I think for whatever reason, like, they've been struggling with this.
Yeah, well, you know, a lot of people dog them matter for a long time until they started.
And they recovered.
Coming out.
And like, I think Google's going through that phase right now.
Yeah.
And, you know, it's just you got to stay alive.
Yeah, yeah, yeah.
I want to focus back on just like your, your thoughts, just general.
We can talk about speculative PCC, mismanaged geniuses,
or just like throw away all this and just talk about whatever else.
Okay, let's talk about mismanaged news for a little bit.
Yeah.
The only comment I'll say on speculative PCC is that it's a really simple idea.
It's almost like obvious that this should be done.
and like there's not much more to talk about it.
Like I think it's just like, you should just use it for like coding,
like anything with programmatic agent calling like RLMs or Kodak,
like yeah, it's like a no-brainer.
What's the, you know, for people that haven't read it,
what's the one-liner?
The simple thing is when the model is writing its code
or like even as like after it finishes writing the code,
a lot of tools tend to be like sequential or like you have to wait on them.
so you should just launch them in advance.
Like if you're able to Jit compile this code,
you can probably figure out like,
even though it's like kind of variables and stuff,
like you can figure out like how exactly analyze.
Yeah, yeah.
So speculation.
Yeah, there is this,
someone pointed to me,
some actually like academics have,
especially PL like programming languages,
people have some like very kind of cool ways of doing this.
And so like, you know, at some point maybe I'll like work on this.
Mostly you have to change language because if you are in JavaScript Python,
you can't do this.
Yes.
Yeah, yeah, yeah.
So like Haskell, yes.
What's the normal one that's not Haskell?
Lisp.
Ocamel.
Any functional language you can actually pipeline this.
So effect TES if you want to do text.
Yeah.
Okay, we can switch over to this match.
Which is like your, you guys' whole thesis, right?
Like that, I mean, to me this is kind of like a restatement,
but maybe I'm missing something of like, well, we're going to better
harnesses or like, you know, your models actually are capable a lot more if you try harder.
So this is a skilled issue.
Yeah. Yeah, basically.
I mean, I think there's one thing I want to see.
I appreciate that there's a big focus on like jagged intelligence because it paints a big
picture of like, we can do this if we really set our minds on it.
But I kind of wish and maybe someone in academia should do this.
Like really just sit down and think about like if I took.
hook Astra, even the current frontier models are not good enough at like doing a particular job
over, let's say, the span of a month consistently and well. And I think this is like a stupid
problem. Like I genuinely think we can solve this. You don't need to be a frontier lab and like
do all this like, you know, fancy stuff for your IPO. Like I think like these models are so smart
that even if it's like a silly way, I think that it genuinely is a skill issue of you can get a model
to be as good as, let's say, like, just some 18-year-old high school kid doing some job.
I think it's, like, ridiculous that we can't do that.
And it's part of the reason is, like, the format of a language model is not really amenable to that.
But I think you can shape a harness around it and do it.
And I think, like, this in itself is, like, ignoring the RLM stuff, ignoring all the, like, you know,
what abstraction should we use?
Like, I just think someone can design a harness that can do this.
Like, I mean, that's maybe you say do this, do what?
Do long running but simple tasks and do them reliably.
What's an example?
So is this different than like, you know, pick your favorite company, Harvey, for example, using LLMs to do legal work?
Or what's the?
I mean, I guess it's kind of like that except if the bottleneck was not like certain legal knowledge or something.
Like, I don't know.
Let's say.
I mean, I guess like the examples, you know, people can take models and build pipelines.
plans or whatever and have an agent repeatedly do whatever test they want.
Yeah, yeah. So, for example, like, if I wanted a general system that I could kind of talk,
I can talk to it like I would talk to an intern and basically just ask it to do, to explore
some small thing. So maybe an example of this is like very silly auto research is maybe an
example of this of like not necessarily finding super, super novel solutions, but
at least optimizing all of the easy parts of any problem.
They often end up being over-indexed for like ML training and things like that.
But yeah, I don't know.
Maybe that's not like super, super clear.
But there is a lot of people's general workflows where you probably could just
vide code up some specific harness to help you do, like, automate this thing.
Some examples are like automating, finding, like, research papers and stuff like that.
But usually people will design like a specialized agent to help them do this kind of thing.
Or like the vibe code of harness, like, and just run it or like their SlackBot or something.
But I almost think there's just like a standard form, like just a harness that you just plug in.
Like you don't, it doesn't need to be designed for finding papers or like you just kind of tell it, find this for me and you like plug it into that setting.
What I'm getting at is that I think there's a lot of easy things that can be automated.
Is this like a hypothesis or a point around the capability overhang?
Like even if we paused, there's still a lot of impact to be had with current state of models?
In some sense, yes.
I guess what I'm presenting is the easiest form of this.
But what this is kind of saying is that like we have jagged intelligence on a lot of things.
Like, for example, models are like disproportionately good at coding and math.
This is saying, like, we can translate those abilities to many other things.
So, like, for example, if you took usually, if you take someone who was an IMO gold or something
and you kind of apply them to a lot of different problem solving domains, they can figure it out.
I don't actually know this applies to models.
For example, like in GPU code optimization, one very interesting question is if you were to take out
all of the GPU programming data, like, from a model, but it was, it was, like, as good as Astra is now,
just without, like, without taken out, would it be able to still optimize GPU kernels?
Like, would it be able to learn in context roughly what it needs to learn and then, like,
have some pipeline or, like, or, like, come up with some solution to solving, like,
optimization tasks? And I think, like, there's, like, a mismatch between, like, if you took a human
that was as smart or, like, knew as much as Astra, there is a mismatch between what that human can do
and what Astra can do, maybe a round of harness.
And I think we can actually approximate the human a lot more.
To me, it sounds very approximate to the continual learning problem.
That's the best example.
Yes.
Yes.
Why didn't you just say that?
Yeah, yeah, I guess.
I was like thinking, could I just blurt out some words?
I'm careful with that, but yes.
I have a very, maybe a bit of a, I don't know, like that.
So I think you're a very, I don't know, I don't know your undergrad actually.
Are you like a math person?
A little.
Yeah.
Yeah.
Yeah.
Yeah.
Yeah, that's what I wanted to do, at least.
Like a category theory type of abstraction where you're thinking categories and then you have to like then translate down to the specific.
And but then you're like actually really care more about the category.
And like that's the communication error because like everyone's listening for the specific but actually trying to also, you know, convey the general.
Yeah.
which is hard
I don't really know
you can maybe use like a shorthand of like
okay I'm at level two
and then I'm gonna go up to level three
and go back to level two
that we should have some of like epistemic
like shorthand for like this kind of thing
because it's hard
like you're compressing a lot into
word after like sequential word decoding
should convert to neuralese
you know is there like a
better form of
output than English
or
Python or JavaScript.
I don't know.
This is kind of like a shitposts,
but people have speculated about like
what is the native language that people want to
that models want to output.
Some people say a binary.
That's Mark and Jensen's thing.
I don't know. Yeah.
PtX.
Let's say a mix of English and Python.
And I only say this because
the capability of a model
is somewhat a reflection of
what we train them on.
So we still want like, like, yeah, I don't really buy the binary argument.
I guess I like, I understand, but it's like, uh, like.
Yeah, you want to model the world in some way.
I think the one thing I'll bring up is always, which I always do in this kind of conversation,
it's Sapir Wharf, which is if you choose English, you will have locked into
however long English has been around, which is, let's say, 500 years, which is not that
long.
Like, actually, you know, like, what you, the language that you speak constraints how
you think.
And if you learn a different language, for example, someone, in Chinese, we don't have tenses.
I don't know if you, I actually didn't know that.
And I speak Chinese.
Oh, oh, I did know that, but my Chinese is not great.
Okay.
Yeah.
Or like in, let's say in Japanese or, I don't know, I forget what language it is.
Like in Korean, everyone you speak to, you have to like acknowledge social status.
Yeah, yeah, yeah, yeah.
It's a different dimension than gender, right?
It just influences everything you do.
When I take Ling 101,
apparently there's a language in Africa
where there's a vegetable gender.
Right?
Or like SMOL is no word for snow or whatever.
Anyway, so the language that you adopt affects your thinking.
And if you choose to output your chain of thought in English,
you are biasing towards whatever English solves.
I don't know what the sort of prior of English is.
That's interesting.
I did not think of it that way.
I mean, at some level, it's interesting, right?
So you're right on language,
a lot of model chain of thought also fluctuates language.
The obvious example is Chinese models.
Speaking in English might still reason in Chinese.
But at the same level, most models are very capable multilingual.
And that adaptation we can see.
You can add in languages.
You don't get that much from adding a whole language.
But they'll reason interchange.
Yeah, so we're all all regressive.
But also, let's say German, like, you know, subject, object.
agreements, you have to put the verb at the end, which is very super annoying, like, very
famously.
Yeah, right?
You don't know what you're doing until the end where you're like, oh, that mess of nouns and
then the verb.
Well, the most classic one that most people have heard of is arrival, where they have
the heptopods, where they think the time is, like, flat to them.
So they think, they output entire sentences at one shot.
So it's as closest to the difference between auto-regression and diffusion.
We talk in auto-regression.
What if you could talk in diffusion?
Where things just resolve over time.
I see.
I see.
But the whole idea shows up at once.
Ah.
I see.
So that is a drastically different language, but it is a language.
I see.
That's really interesting.
That's really interesting.
Which machines could speak that we probably will never speak, but, like, machines don't care.
Maybe this is a huge tangent.
But are there not, like, things inherently?
that are reasoning chains
that are inherently
auto-regressive.
Yeah, time's like, yeah, or like, yeah, yeah.
Yeah, something happens first,
then something else happens.
Even like, yeah, anything in code, for example,
like has to be causal in some,
usually at least, has to be causal.
Well, no.
So then you have to,
then you're not exploring enough
program and language theory
where everything is like pure functional
and like completely relational
and you sort of abstract away the solver
that translates the relationship,
ships that is always true into code.
So, yeah, I feel like this is maybe a little bit too out of my death.
But I love languages, whether it's coding or human,
and I do think a lot about how that affects reasoning and the boundaries of what we can do.
I don't need to go too much beyond that.
I don't know if you have any other thoughts.
My closing question is going to be, you have all these research directions that you want to do.
You know, you had a GPU mode phase.
you had a RLM's phase,
presumably you have other stuff planned,
which is why you're not doubling down on that.
By the way, I notice that it is interesting
how you guys do start with the GPU side
and then you migrate towards the zero gradient side,
which is what you called it.
Doesn't that feel less legit than messing with GPUs?
Yeah, I guess in the sense that like,
so you did bring up that, like,
I like to think about things in like a math-oriented way.
And it's like very uncomfortable sometimes to be working on like harnesses and agents because it's so.
Because you think all harnesses are the same.
It's like super fuzzy.
You're like two new ideas in harnesses.
Yeah.
It's also just like like empirically it's hard to like verify a lot of findings at least with with the compute that we have available to us.
But the reason why I think I've moved on to a lot of these problems is I think actually this is where most of like the innovation is yet to happen.
To me, like, the GPU level is a means to exploring other ideas.
Like, you want to, for example, like, get good at writing kernels or, like, even automate
writing kernels for the sake of a broader goal of, like, I want to explore ideas where I'm not
bottlenecked by systems challenges.
In that sense, like, you know, I guess a lot of what's written there is all harness stuff,
but I am also interested in things at the model level as well.
But I'll just leave it at that.
Okay.
That's a good hint.
anything any if people want to reach out to you what are you looking for help on what do you want
collaborators on any sort of calls to action yeah so um i guess there's nothing i have in particular
where i feel like i need to work with someone on unless it's like unless it's with a company
before like compute or like with to talk with other people about it but i will say i i'm not
i'm never opposed to working on ideas with other people i get reached out to a lot
by often undergrads or even like other students.
Podcasters.
Podcasters.
And usually I feel like I get an email that's something along the lines of like I really like
RLMs, like I want to work together.
And I feel like I.
Yeah, that's a bad at reach out, right?
Yes.
The worst is like, can I pick your brain?
Yeah.
And like, I don't know what?
Like read my paper, dude.
Like they'll be like, I read your paper in quotes like recursive language models or like
Prime Agent like a self-improving RLM harness or something.
And it's kind of like I like, I like, I really like people.
people that are opinionated, even if we disagree, I think if you have strong opinions and are
able to, like, think through why you think those opinions are right or wrong, because usually
it's hard to actually tell. But, like, you have strong convictions about certain problems. Like,
I'm always happy to, like, chat and even, like, potentially work on something together. I have, like,
no limit to who or, like, what I would like to work on. So, yeah, yeah, yeah. I, you know,
in the era of agents, I think there's a lot more work.
you can do, you know, like, like bandwidth-wise.
So I, yeah, I think in general, like, I am not hard to impress,
but I think it just takes a little bit of effort to kind of, yeah, know what you want.
It's very clear.
I mean, and like when you see a new thing come out, well-executed, good, simple idea,
then, like, immediately gets your attention, right?
It's actually, like, not that hard to get the same attention at all the Frontier Lab guys.
Because they are looking for you.
You just have to put yourself out there, right?
Exactly.
Yeah.
But yeah, I will say I think human attention is very scarves right now.
And I do struggle with like the number of projects I have going on.
And I don't know how to manage it.
I don't think agents are helping at all.
Like I'll just prompt it and I'll prompt the thing and then never look at it.
Right?
Like, which is very common.
Yeah.
Yeah.
And that sucks.
I guess maybe that one of the smaller differences in like, I mean, actually maybe you're doing research.
I'm not sure.
But for me at least like I all have maybe like 10 or 15 different ideas that I want to do.
but the thing is like most of them are bad.
And this also maybe is true even for someone that reaches out to me.
Like maybe the idea is actually bad, but it looks interesting to me.
And so like we can spend like some time looking into it.
And if like we feel like there's actually something there,
then we should take the next few weeks and just really pursue it.
And like this is my style with, this is why I love the PhD, by the way,
because there are times when I'm just thinking about problems,
like maybe on a run or just like playing tennis or something.
Like I'm not working, I guess.
But it's like those are the most fun times.
And then when I like really am convicted about something,
I'll just like drop everything and just do it.
Like just spend like all my time thinking and working on that problem.
And then I, you know, once you get to the point where like you can just run experiments,
then it's it's kind of easy coasting again.
So.
Yeah.
Sorry.
This is like the fourth last question.
Which is like, I think a lot of people are also thinking about science as an experimenter like physical sciences,
bio, math, even.
How do you separate, like, I guess, let's say,
your choice of projects that is applicable for industry,
and then maybe a choice project that's just, like, science?
I actually worked on, like, AI for bio stuff
before I started my PhD.
The field has changed a lot since then, I should say.
Because it used to be a theoretical, like,
of course, what do you mean?
Like, you know, I have one path and then I chose that.
But now a lot of people crossing over.
Yeah.
And so we have started.
science pod to just cover those things. Oh, well. Because a lot of engineers are like, well, actually
there's tractable problems there. Yeah. You know, I will preface by saying my understanding of a lot of
these topics is probably pretty limited. But I think like if I find out either because someone
reaches out or like I look at a problem and I'm like, hey, like some design principles that we use
or that we're thinking about right now actually make a lot of sense for this problem. I get excited
about those as well. But I think I think it's harder. I don't know. I think with
I think science, especially like empirical or like applied science, has very, very long, like, what is it called?
Like feedback loops or whatever.
Yeah, it converges to a robotics question.
Yeah, yeah.
I mean, to me, like, also this aspect of like what is worth spending and betting my time on now?
Because, like, maybe I spend a lot of time on this problem and then, like, in six months, like a different solution kind of like maybe a new model comes out and it's like, oh, it's way better for this.
And so I do have to be careful.
Like, you know, you have to be conscious about, like, where you think things might be going.
So.
Publish cycle.
Yeah, yeah, yeah.
So I can just, oh, RKGI3, it's saturated.
We did it.
Yeah, I mean, RKGI3 got saturated in less than a year.
So it's kind of, you know, like, it's, I don't know.
Like, if you were allowed picking that problem, like, you're probably kind of sad now because.
Well, exactly.
That's like knowledge work, gaming, all these things are saturated.
Now, actually, the frontier is science.
Yeah.
Knowledge work is saturated.
Yeah, GDP vows like 80-something, 90-something.
I mean, like, you know, there's 90 to 100%
that's obviously going to take the next 10 years,
but like, well, the next low-hanging fruit is going to be...
It makes sense.
Yeah, yeah, yeah, yeah.
Maybe I'll think about that more.
I actually, I haven't given too much thought to.
I'm just trying to guess your next direction, actually.
No, no. I will say because I think, I think especially at MIT,
like, it's, there's a lot of really talented scientists there, like, in the natural
sciences and I think it's it's a little bit like sacrilegious almost to be like I'm going to figure out like
your problem like no that's that's great oh no I know yeah so I interviewed I Te who did the IMO thing he's like I've never
I've never been to IMO I don't even know what it is skill model dude yeah yeah uh which is like very
disrespectful about like whatever that's like at some point like you have to respect like okay
you know the the progress is being made like number is getting output right sure yeah I mean
That is very bitter a lesson.
It's very interesting because the mathematicians are responding this way to Naryostov right now.
Like, Terry turns towel is like, no.
Like, let's not use AI.
Yeah.
I'm like, I don't know.
Well, yeah.
I mean, I think that, that whole thing is kind of weird because I feel like, I feel like they would have a stronger case if a lot of them didn't work with open AI before.
Like all this happened.
No, that's ad hominem.
And they're really trying to stay away from that.
So what, right?
So what?
Like, I don't know.
Like, so what?
They got, you know, they've collaborated.
I'll collaborate with people that I don't agree with or whatever.
Like, I did a thing and then now I regret that.
I changed my mind.
Whatever.
So I'll defend their right to say that.
But like, yeah, a lot of people are reasonably disagreeing with them.
Yeah.
Okay, cool.
Thanks for your joining us.
Congrats on your success so far.
I can't believe you're still not done with a PhD.
Oh, it's your two, you know.
I can't believe we did this podcast without going through the RL and paper.
He had a definition.
I think the paper is more.
about like empirical results like the actual idea is quite simple yeah and you've talked about it
many times yeah yeah at this point i think there's there's more interesting things to to look over now
so cool yeah well we're excited to see what you do next thank you so much
