Latent Space: The AI Engineer Podcast - Claude Code’s Next Era — Thariq Shihipar, Anthropic

Episode Date: September 29, 2026

We are excited to have Anthropic share their latest AI x Finance work at AI Engineer New York, coming up in 2 weeks!In case you’ve been under a rock, here’s a non-exhaustive list of what Anthropic... has been shipping since closing the largest fundraise of all time in May at $47B ARR:* June: Launched Claude Tag and Sonnet 5 and Fable 5* July: Opus 5, /checkup. crossed $65B ARR* Last month: Fable/Mythos 5.1, and EFS (upcoming pod)* IPO target $2T, end 2026 ARR estimated $100B* Cowork/chat merged before did* Claude Mods* Dario endorses the same Pacing the Frontier message cosigned by all labs* Last week: Opus 5.5, Plugins portal, Cloud Sessions/Claude Projects* Today: Sonnet 5.5!Today’s episode should catch you up, with Thariq Shihipar, the explainer-king of Anthropic, who we last caught up on Fable launch day with The Field Guide to Fable:The Future of Mutable SoftwarePay special attention to Claude Mods (especially the cheatsheet):Cloud Brain, Local HandsAnd give a try to Claude Projects:The “hands” terminology is not just an analogy for the local/cloud paradigm that is being built up at frontier coding agent companies like Cognition, but is ALSO particularly relevant to the safety systems discussions that we’ll be discussing with Anthropic in an upcoming episode as they prepare to pace to frontier with responsible AI deployment.From the rapid rise of Claude Code to a future where agents can rewrite their own harnesses, collaborate across teams, and operate across cloud and local environments, the way we build software is changing extraordinarily fast. In this episode, Anthropic’s Thariq Shihipar joins swyx and Vibhu to unpack how power users are actually working with Claude Code today, why prompting remains a high-skill discipline, and where Anthropic thinks the agent harness is headed next.We go deep on Claude Code’s evolving interface: Ask User Question and elicitation, artifacts as persistent generative interfaces, Claude Tag for multiplayer agent workflows, Projects, model effort, implementation notes, and the new Claude Mods system for customizing the harness itself. Thariq explains why Claude.md may eventually disappear, why the smartest model could also become the cheapest model for many tasks, and why mutable software could become a new paradigm for how applications are built and customized.The conversation then turns to agent security and Anthropic’s “Pacing the Frontier” argument. Thariq walks through recent incidents where agents discovered unexpected ways to communicate, exploit infrastructure, reverse-engineer benchmark scorers, and chain vulnerabilities together. We discuss sandboxing, prompt injection, autonomous agents, interpretability, constitutional classifiers, probes, fallbacks, Auto Mode, and why securing increasingly capable agents may become one of the defining engineering problems of the next few years.We discuss:* Why agentic coding went from controversial to the default in less than a year* Why prompting is still one of the highest-leverage skills for working with Claude Code* How expert users build a mental model of Claude and what it can reliably one-shot* Why discovering your “unknown unknowns” matters more as agents become more capable* Artifacts as persistent, generative interfaces between humans and agents* How Claude could split into a cloud-based “brain,” local or remote “hands,” and dynamic interfaces* Claude Tag, Projects, and multiplayer agents and how collaborative agent workflows could evolve* Why spending more time on the initial prompt can dramatically reduce wasted agent work* When to use low, medium, high, or max effort for different engineering tasks* Why frontier models may eventually outperform smaller models on both intelligence and token efficiency* Why implementation notes can expose decisions the model considered but chose not to make* Why Claude.md may eventually disappear — and why starting without one can sometimes be better* Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code* Model routers, forked agents, and supervisor agents that automatically improve agent workflows* Why Claude Mods may be an early preview of “mutable software”* The bitter lesson of harness engineering and why agent architectures go out of date so quickly* How Claude Tag is becoming an organizational harness for multiplayer work* Why giving agents access to company data creates an enormous new security surface* The Exploit-Bench incident where agents discovered ways to communicate and collaborate* Why agents hacked Hugging Face for scorer code rather than benchmark answers* How agents chained sandbox and infrastructure vulnerabilities in unexpected ways* Why increasingly capable agents make traditional security assumptions harder to maintain* The argument behind Anthropic’s “Pacing the Frontier” proposal* Why software engineers are increasingly doing two jobs: engineering and keeping up with AI* Constitutional classifiers, probes, and fallbacks and what interpretability looks like in production* How Auto Mode checks whether an agent’s actions actually match the user’s permissions* Why Thariq can see serious AI risks while still having a relatively low p(doom)Thariq Shihipar* X: https://x.com/trq212* LinkedIn: https://www.linkedin.com/in/thariqshihiparTimestamps00:00:00 Introduction00:04:12 Ask User Question and the Future of Agent Interfaces00:08:29 Artifacts, Projects, and Multiplayer Agents00:15:37 Prompting as the Core Claude Code Skill00:21:52 Context, Effort, and Smarter Model Usage00:28:10 Is Claude.md Going Away?00:32:49 Claude Mods: Customizing the Claude Code Harness00:36:35 Model Routing and the Rise of Mutable Software00:44:40 The Bitter Lesson of Harness Engineering00:50:49 Claude Tag as an Organizational Harness00:55:59 Pacing the Frontier and Autonomous Agent Security00:58:22 Agents Hack Hugging Face for the Scorer01:05:34 What Happens When Agents Need More Compute?01:10:32 AI Coding Is Changing Faster Than Engineers Can Keep Up01:17:17 Probes, Fallbacks, Interpretability, and Auto Mode01:28:32 AI Risk, p(doom), and Closing ThoughtsTranscriptIntroduction: Life at Anthropic and the Pace of ChangeSwyx [00:00:00]: We’re here in the studio with our friend Thariq from Anthropic, and I guess generally the Claude Code, I-- there’s, there’s so much, merging of boundaries and you’ve been so on top of everything since you joined Anthropic. You have been early to Claude Code itself, but then also, and you’ve told that story in other podcasts, and you’ve also been talking about seeing like an agent. Most recently you did the top AIE World Tour talk, Field Guide to Fable, which obviously you guys launched Fable, so that was-- that’s cheating. And mostly you most recently also launching Claude Tag, and we’re also gonna be talking about Pacing the Frontier. There’s a lot going on in Anthropic. I guess top of the question is, what’s it like being at Anthropic when there’s so much going on?Thariq Shihipar [00:00:48]: I think that It is, like. I think you can get whiplash sometimes. I think, like, going. When I joined Anthropic, I joined because of Claude Code. Like Claude Code had just come out and I was like, “This is so good.” And Opus 4 to me was like just, I could not imagine, like, how good it was? And that was, like, a real moment for me. But I was, like, trying to convince, like, my startup friends to use agentic coding, and they’re like, “Oh, no, like, our engineers don’t think it’s good enough,” or something. And I was like, “That’s insane.” and now you, like, fast-forward, 12 months, less, and, like, it’s just like, yeah, the default way that everyone codes, right? And I think that, like, just having to go from, like, selling it to, like, now, teaching people how to be. make the most use of it and be more efficient and things like that is just like a big, like big change. And, yeah, I think, like, it’s just hard to stay on top of everything as a human? Like, I think things happen so fast and likeSwyx [00:01:51]: You just throw more agents at it.Thariq Shihipar [00:01:52]: Yeah, like that’s like the agentic stuff scales much better than the, like, human stuff where it’s like, oh, like, there are three things happening right now and, like, they’re all emergencies and, like, how do you, like, respond to it? Yeah.Teaching People to Use Claude CodeVibhu [00:02:05]: What do you split your time on? You do a lot of technical writing, engineering work.Thariq Shihipar [00:02:10]: Yeah, so I think that, like, when I joined the Claude Code team, I wanted to teach people how to use Claude Code and I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, I’d, like, do. I was spending some time on the agent SDK first, and I wasn’t exactly sure, like, how the bitter lesson would go, when it comes to, like, harnesses, right? Like, I think sometimes we were like, “Oh, like, what’s after Claude Code?”? And so initially I was like, I just wanna teach people how to use Claude Code and make it easier to use Claude Code. And I think that has just, like, as the harnesses have gotten better and better, that’s like the dominant problem now is, like, how do you use the agents, right? Like, it’s like such a high skill expression thing. So I do that and then I do engineering work. I give talks, but I think, like, when I’m doing engineering work, my goal is to take that feedback that we get from users and also, like, then be able to talk about, like, hey, how to use Claude Code to do engineering. So there’s like a good loop there. Yeah.Swyx [00:03:07]: Yeah. I’ll-- For listeners, we’ll attach, the talk that you did with Sarah for the Dev Writers, meetupThariq Shihipar [00:03:13]: Oh, yeahSwyx [00:03:13]: Which we talked a little bit about, well, first you do the work and then you talk about the work.Thariq Shihipar [00:03:16]: Right.Swyx [00:03:16]: Something like that.Thariq Shihipar [00:03:17]: Yeah.Swyx [00:03:17]: It’s sow and reap orThariq Shihipar [00:03:19]: Yeah, reap and. Sow and reap.Swyx [00:03:21]: Something like that. Something like that. Yeah, so, and then just to preview a little bit, we are gonna talk about the evolution of the harness. It has come a long way from just being a CLI. We’re gonna talk about, Claude Mods, which is starting to leak today, because you couldn’t keep it secret.Thariq Shihipar [00:03:36]: Yeah. yeah.Swyx [00:03:39]: Yeah, there’s, there’s a lot, there. I think you started off with, like, adding ask user question tool, which people love and hate.Thariq Shihipar [00:03:48]: Yeah.Swyx [00:03:48]: Like, I thought it was, like, very innovative, and then now I have, like, my own version. You have your Interview Me version.Thariq Shihipar [00:03:55]: Yeah.Swyx [00:03:56]: And, yeah, everyone just has, like, their own stuff. And, like, it no longer matters ‘cause now you’re supposed to, write prompts that create other prompts and loops and all these things.Ask User Question and Human-Agent InteractionThariq Shihipar [00:04:05]: Sure, yeah.Swyx [00:04:06]: So what’s the state of the art, today? Like, what are people. what are you, like, telling people to do today?Thariq Shihipar [00:04:12]: Yeah, ask user question was the first time that the model was good at elicitation. I think this was, like, an emergent behavior that I, like, wanted to see if the models could do. I have, like a human-computer interaction background, so I, like, did that in undergrad and grad school. And so this was like. I think it’s like human-agent interaction to me, like, trying to figure out, like, how can the agent communicate with you and extract, the requirements, right? I think that, like, one of the things about, like, that’s difficult as Claude Code has gone broader and broader is that everyone has, like, their own way of using it, and it’s very hard to, like, change the default behavior. So for example, like, if someone asks Claude Code to do something,Thariq Shihipar [00:04:59]: Sometimes they just want them to do the work, ‘cause they’re, like, maybe a very good prompter, and sometimes they want. like, are not good at prompting? And you need. like, the agent needs to, like, clarify? And so that’s, like, a good split. Like, and the ask you the question tool like, splits along that side where, like, are-- do you feel like you’re good enough to instruct the agent as it is, or is the agent able to, like. does the agent need to, like, pull out more requirements and, like, collaborate with you more and really understand your preferences?Thariq Shihipar [00:05:27]: I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, than they, like, think they know about the problem. but, like, it’s like a interface design problem to make that easy? And so, like, if you’re designing a problem, like, or if you’re going through a problem, like, things like what’s the schema or, like, what’s the call stack and things like that are really important. like, the details in the design are important. Ideally, you want to figure out some of these, like, hard problems ahead of time before starting implementation. And yeah, that’s why they call, like, unknowns, right? And so I think that this will forever be, like, a skill in agentic coding is, like, figuring out your unknowns. So, like, because even if the model is, like, super intelligent- It, like, needs to know what you want? And, like, you have preferences. like, you need to like, pull the, pull that out. and so that’s, like, I think how I’m, what I’m pushing. the question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that. And we’ve recently added artifacts, right? And artifacts, I think we’ve done a bad job of, like, or, like, I’ve done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them? And so every artifact can store and write persistent data. They can, like, feed back into Claude? And so, like, one thing that, like, people are not doing yet that I’m trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long-term. Maybe it’s like a kanban or something. it can store that kanban data in its database. Multiple Claudes can access that data via, like, the artifact MCP, and, like, that artifact can, like, talk to those Claudes as well. And so, like, the. We’re building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything with agents right now is, like, this problem of, like, you think what you want, but you don’t really know what you want, and, like, the agents need a lot of detail, and collaborating with them in the loop is really important. And so artifacts are, like, the, like, way that we’re trying to evolve there. But there’s a lot of work to do because it’s so much more complicated than, like, a multiple-choice question? there’s a lot more, like, detail in terms of, like, diagrams and code snippets and schemas or, like, whatever it is for that problem. But, like, artifacts is, like, the mo-more AGI-pilled way of, like, doing ask user question. So yeah.Artifacts as the Interface to the HarnessSwyx [00:08:15]: I think one thing that’s unclear to me about these, the artifact stuff is, like, what feedback should go in through the artifact and what feedback should go through a Claude, a chat? Because the more AGI-pilled one is to just feed everything to the Claude.Thariq Shihipar [00:08:29]: I think the more AGI-pilled one is to go through the artifact. Like, and I think that, like, we imagine in the limit, I think that artifacts will be your interface into the harness? You can, like, comment on this, like, live, like, document of your plan, of the work. you can see maybe, like, multiple agents and different agents are doing this, and that artifact is built for the current work that you’re doing, right? And so, like, each one has, like, slightly different. I think we’re still, like, getting there from, like, an infrastructure perspective. But yeah, I think, like, on-the-fly interface for your harness is probably where things are headed.Vibhu [00:09:03]: Is there a version of it that’s an abstraction from CLI or chat and you. Because right now, a lot of it is, okay, you’re interfacing with Claude Code, you’re having HTML given back for a mockup. It’s pretty rich. There’s diagrams. Artifacts are ways to connect these together. Why not just do everything that way?Separating Brain, Hands, and Surface UIThariq Shihipar [00:09:22]: Then it becomes, like, separating out, like, where is the inference happening? Where is the intelligence happening? Where is the work happening? like, I think this is like, difference between, like, or, like, some of the distinction between local and cloud, right? And so, I think right now, if you use Claude Code, it’s, like, local and, like, you can spin off remote control, for example, to get some cloud behavior, or you can spin off Claude Code in the cloud, right? We’re moving towards a place where instead of Claudes, like, you message a local Claude, it starts a session locally and it executes, to more like you have a Claude that you message that’s in the cloud that’s running. it can run, like, local, or, like, cloud sessions. This is how Claude Tag works. But, like, over time, we’ll add, like, local hands as well. And so, like, local hands will be the ability for that agent to access your computer if it’s online, and be able to, like, work there. And so it can spin off many different subagents. It can, like, commu- those subagents can communicate with each other, and that’s where the artifact comes in to display all of that work. So you can imagine, like, the. You’re separating out these things. So there’s, like, the surface UI display that’s an artifact and hosted somewhere and has a database and everything. There is the inference intelligence, right, that’s happening on the cloud, and you don’t have to worry about shutting off your computer or whatever, right? and then there’s the, like, hands. Like, and it can be local, it can be in, like, a remote sandbox or wherever you need your work to be done. That’s like unpackaging, like, the Claude Code experience right now where, like, right now it all happens in one place, right? So.Multiplayer Agents, Claude Tag, and ProjectsVibhu [00:11:00]: How do you see, like, the multiplayer side of that? So say teams want to work in this way. Right now it’s very individual, but how do you see the future of multiplayer? Like, right now, I guess there’s Claude Tag, which is a version, but.Thariq Shihipar [00:11:12]: We’re launching projects. And so projects is the, like, this abstraction that’s like Claude Tag, but on our Claude products, right? So you can message it and, like, it will do the Claude Tag-like stuff, like spinning off subagents. So We think with multiplayer. Like, Claude Tag is, like, a little bit more native multiplayer because it’s just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you’re like, oh, you have hands, but now you have other hands in other people’s computers too, and, like, you need to, like, permission them or, like, you have, like, your MCP and someone else’s MCP, and how do you figure out how to use them, right? It gets, like, quite complicated. And Claude Tag does a good job of, like, sanding down all of these issues, right? So that, like, when you have, yeah, Google Docs, how does it access Google Docs, right? Like, it accesses through the shared Claude MCP, or it can access through your local credentials as well if it doesn’t have access. But yeah, I think Claude Tag is our multiplayer, product, and it’s really useful for these, like, things that are inherently multiplayer. Like, okay, like on-call, for example, incidents are inherently multiplayer. You want to tag Claude, you want multiple people to log in, you want it to be able to find context. I think whenever I’m, like, working on something and I want, like, privacy or security or, like, I want other people to review it’s really nice to, like. I’ll have a channel per project and I’ll, like, at legal, for example, be like, “Hey, like, I want to ship this. Can you, like.” Like, here’s. Like Claude knows everything, just chat with it. And that way legal gets precise answers, on like what exactly is shipping into the code, and I don’t need to be in the loop, right? So I think like multiplayer is getting like more and more like, yeah, everyone can participate with Claude. I think Claude Tag is like that product and like projects will start off single player and will like, expand.Swyx [00:13:14]: I think there’s a question about like maybe dual questions about identity and the unit of isolation.Identity, Permissions, and IsolationThariq Shihipar [00:13:20]: Yeah.Swyx [00:13:20]: Claude Tag, you specifically chose to make it its own identityThariq Shihipar [00:13:26]: Yes.Swyx [00:13:26]: Which is like, a controversial choice. There’s, there’s other ways to do it.Thariq Shihipar [00:13:30]: Yeah.Swyx [00:13:30]: Claude Projects probably it sounds like, if it’s anything like ChatGPT Projects, it is, the isolation is that artifacts, that cloud instance, everyone’s collaborating on this. It’ll. It sounds like, it should be like if you’re, if you’re collaborating with legal on a thing, like that channel should be a project, right? Like it’s not yetThariq Shihipar [00:13:50]: Yes.Swyx [00:13:50]: But it. that’s the natural next step.Thariq Shihipar [00:13:53]: Yeah, like I think in Claude Tag, it’s effectively. Like Claude Tag, you have to do your own arrangement. And so Claude Tag, yeah, each channel is like you can name it as you want, and I nameSwyx [00:14:04]: Yeah.Thariq Shihipar [00:14:04]: Like each featureSwyx [00:14:06]: Yeah.Thariq Shihipar [00:14:07]: As a channel.Swyx [00:14:07]: And, but I think like there is some trans- like it’s unclear when there is transference, because let’s say it is. if you have a coworkerThariq Shihipar [00:14:14]: Yeah.Swyx [00:14:14]: Who is tagging on all these things, yes, there is transferThariq Shihipar [00:14:16]: Yeah.Swyx [00:14:16]: Because it’s the same person. but with Claude, it’s unclear if it’s like necessarily like, well, no, you don’t know any of. you don’t know about the other stuff. You should only use this stuff.Thariq Shihipar [00:14:25]: It’s like the tip of the iceberg meme, right, where you can like. This is what we spend so much time onSwyx [00:14:31]: Yeah.Thariq Shihipar [00:14:31]: Is like there is like infinite surface area of like, okay, you want Claudes to. Not infinite, but like there’s like surface area, a lot of like, surface area to figure out of like permissions and visibility and like how can you let Claude operate as well as you can, as safely as you can? And obviously, this is very important to us because like security for our code base is very important. And so we’ve put a lot of time into this. Yeah, there’s so many like edge cases you can figure out where it’s like, oh, like, yeah, this Claude in this channel has different permissions, but it can message another channel, and can’t it exfiltrate data that way? Or like can you like. What if it uses your MCP and then messages someone else? Like there’s like so much, and we’ve like really put a lot of work into sanding it down.Swyx [00:15:14]: Yeah. Lots of work. okay. Fable?Fable and the Meta-Skill of PromptingVibhu [00:15:18]: Fable, you wrote two good articles. you’ve written many good articlesThariq Shihipar [00:15:22]: Yeah.Vibhu [00:15:22]: But on, Field Guide to Fable, Building Claude Code. I’m curious from what you’ve seen, is there any common patterns that you see in like top users at Anthropic externally? Like what are best practices for getting the most out of Claude Code?Thariq Shihipar [00:15:37]: The like meta skill I say is like prompting is like very important? And like that. Like I think this is like not trivial to say because I think a lot of people are like, “Oh, prompting doesn’t matter. It’s just like I can just say a sentence and Claude will do it.” And I think prompting is really this like, this. It’s like public speaking, like, or writing or something, and for a specific audience, and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works, right? And so that’s like the most important skill in working with Claude Code is like having this mental model, right, of Claude and like what it can do well, what it can one-shot, what it can’t. And so many people when you see prompting, they’re just like, they’re short prompts, but they have such a good mental model of Claude and of like the code base and things like that like it’s effortless? But it’s like high skill ceiling. So like that work of like, spending a lot of time prompting and building mental models of how, and intuition for how the agents work is really important. And then I think like the next thing is like the unknown stuff we talked about earlier, where it’s like being able to find out like your, what you don’t know or what you haven’t written down, learning about like different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you and like you have low domain knowledge on is very high? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so like I think the most important unknowns are the unknown unknowns, where you’re like, I just like don’t even know that this exists, right? Yeah, exactly. I think that’s like a illustration of like the map and the territory, right, where you’re like, “Okay, this is my prompt,” and the territory is like the actual like work that the agent needs to do, right? And if you are like very precise, you can give more precise things, right? So like for example, in design, I’m not very precise. I’m not a designer, so I say like, “Give me like eight different mock-ups.” But if I was a designer, maybe I’d be like, “Oh, hey, here are some reference sites.” Like, “I want this type of font and this type of like look to it, and here’s like a few different components to like visualize. Here’s a Figma MC board to bring in,” like. And so you can just be so much more precise with that language. And if you’re not a designer, you just need to like try and learn the language or learn the unknown unknowns. And this is true of like everything, I think. Like the more, like you can work with Claude to learn like how things work, the better your prompting will be. I think another good example of this is like game design, like where a lot of people are like, “Oh, like I can vibe code a game now.” And they’re like, “It’s not fun.” And like it’s just like the thing about game design is like every one of these choices has like a lot ofTaste, Domain Knowledge, and Learning the VocabularySwyx [00:18:25]: Variations.Thariq Shihipar [00:18:25]: A lot of like craft to them. So it’s like, oh, okay, like when you’re making a flying game, the feel of the plane and the like, way it responds to your controls has a lot of like. Like, a game designer would spend like days on that. Do? and likeSwyx [00:18:44]: To me, that’s what taste is, right?Swyx [00:18:45]: Like it is like from the possible space of one thousand mathematically valid answersThariq Shihipar [00:18:49]: Yeah.Swyx [00:18:49]: Here’s the one that is the humans will like.Thariq Shihipar [00:18:51]: Yes. Yeah.Thariq Shihipar [00:18:52]: I think with taste, I’m like torn on this word ‘cause I think you’re right, but everyone has different definitions, and it sounds kind, sounds like low skill or like elitist almost, where you’re like, oh, like there are certain people with taste?Swyx [00:19:06]: It’s like taste is what I call taste.Thariq Shihipar [00:19:07]: Yeah, exactly.Swyx [00:19:08]: And it’s like these guys don’t have taste.Thariq Shihipar [00:19:09]: Yeah, exactly. Oh, like an engineer doesn’t have taste. Like I, the like founder, have taste.Thariq Shihipar [00:19:14]: ? And I think that’s not true. Like I think like the engineers have a lot of taste for these particular like problems? And I think everyone has taste for particular problems. I think like Jason Liu, like say like in order to, yeah, have taste, you have to eat?Thariq Shihipar [00:19:32]: And I really like that, where it’s like, okay, you have to like do a lot of things. You have to like iterate and figure out what you want, what you like, and, like build that like domainSwyx [00:19:41]: YesThariq Shihipar [00:19:41]: Domain vocabulary. And then when you’re prompting, you’re like synthesizing all of that for a product.Swyx [00:19:46]: Isn’t it annoying when someone else says it better than you?Swyx [00:19:48]: It’s just like, f**k, I have to quote this guy forever.Vibhu [00:19:51]: Having to quote Jason Liu forever.Vibhu [00:19:53]: He’s gonna love this.Thariq Shihipar [00:19:55]: So I get prompts, more than that.Vibhu [00:19:57]: And sometimes it’s not even that. Sometimes it’s just intuitive, right? Like you don’t realize you even want something till a model puts it out, and you’re like, “Oh, this just feels immediately better,” right?Voice Prompting and Information DensityThariq Shihipar [00:20:07]: Yeah, exactly.Swyx [00:20:09]: One thing I go back and forth on is I feel like the way I prompt half the time, let’s say I use voice.Swyx [00:20:16]: Did I say voice? Other people have voice. that is the opposite. That is just like me rambling for like two minutes Pressing down the function key and then let go, and then like hopefully it figures it out. And oftentimes it does.Thariq Shihipar [00:20:26]: Yeah.Swyx [00:20:26]: But it’s not as thoughtful as like a structured prompt with like Well-run communication as though it’s a PRD or a memo. Is that in line with how people do this? There’s like bimodal prompting where there’s some prompts where you spend a lot of time upfront and other prompts you just dash it off?Thariq Shihipar [00:20:43]: I don’t think the voice is necessarily low. Like I think it’s like more like how much information is in the prompt. like the model can. Like you can and like add some sentencesSwyx [00:20:53]: RightThariq Shihipar [00:20:53]: And be like, “Oh, like I changed my mind,” like in the middle of the prompt, and it will be able to follow that perfectly? So I think the like actual format of the text is less important, but then like the ability to. Like how much information is in it, right? And I think for voice, a lot of times, going back to like human-agent interaction and like for a lot of people, it’s just way easier to talk than to like type? and I. If that gets more information out of you, like that’s better.Vibhu [00:21:21]: At some level, it feels like just giving the model as much contextThariq Shihipar [00:21:24]: YesVibhu [00:21:24]: Over prompting before you kick off is a best practice. I don’t know. A lot of the times, like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long prompt. This, I think, is a response of models running for longer and longer, right? It’s still a little difficult to nudge them as they’re in like, in the loop, but I just like intuitively spend more time kicking off that first prompt and working with it a lot.Spend More Upfront, Iterate LessThariq Shihipar [00:21:52]: My personal opinion is that if I was a software engineer, if I was like, just running my own startup, for example, I think I would mostly fit, stick to a max 20x? like maybe verification and so code review are like separate things. But I think like what I see a lot of times is people hit rate limits when they’re doing this like, oh, like it did a lot of work and you’re like, “Oh, I don’t like this.” Like, “Can you like undo this and redo it?” And then you’re like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context? And instead it’s like you’re like, “Nope, don’t like that design. Try this.” Or like, “You messed this up,” or something like that. And then that just eats up so much more of like, your usage. And so that’s like, I think maybe like a key like tip both for like efficiency as well, right? And yeah, I think like context, and not just like context on like what the goal is good, right? Like are you building a prototype or is it like a production thing? Like where can you spend compute or when, where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where, like it doesn’t know intuitively how much you want to spend on this task, right? And you can use effort for this. So I did-- I’m working on a blog post about that where it’s like, if you want. For like we see that effort scales with the complexity of the task. So for security, effort gets like way more results. Like high effort versus like low effort gets, like changes the evals a lot. But for software engineering, it doesn’t change it a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, “Hey, this problem is something that I think I want you to spend a lot of time verifying and edge case testing,”?Effort, Model Choice, and VerificationVibhu [00:23:43]: How about model in the mix? So, there’s Opus and Fable with effort.Thariq Shihipar [00:23:47]: Yeah.Vibhu [00:23:48]: There’s also Haiku in there.Thariq Shihipar [00:23:49]: Yeah. It’s not quite true yet, but it’s very close where I think the frontier models will be Pareto dominant over like almost everything. like maybe. And sometimes I think Opus might be Pareto dominant. Do? Like I think depending on like how things, like shake out if it’s like a newer version of Opus. But I think that like increasingly it’s just going to be like the smart model is going to be able to like do the simple task for less tokens than the like the other models because of verification. With verification, in the limit, your model doesn’t need to verify, right? If it’s a perfect model, it just does the work once and it’s like, okay, like you, I did it? And increasingly with Fable, I’m like, I’m like, “Dude, you don’t need to spin up Chromium and screenshot all of these things.” Like I see it. Like you did it, right? And so a lot of the. At higher effort, you spend more of those tokens verifying. But if you’re working on simpler problems, and a lot of software engineering is like well, like in Fable, like low and medium stability, it can spend less tokens verifying. And as the models get smarter and smarter, they will just be able to like, “All right, done.”? Like, I can run the lint for sanity’s sake, but, like, I, like, know it lints? Like, you don’t even need to do that. And that will be so much more token efficient than, like, the smaller models. Yeah.Swyx [00:25:15]: Is there a good, practice on our side that we can use to see if we’re using too much effort? Like, I freakingThariq Shihipar [00:25:23]: YeahSwyx [00:25:23]: Hate wasting time on that stuff.Thariq Shihipar [00:25:24]: Yeah. I know what you mean. I think, like, so in this blog post, my rough distribution is, like, code review and security should be, like, high or max and, like, software engineeringSwyx [00:25:37]: You said recommend mix settings per domain.Thariq Shihipar [00:25:37]: Yeah. I think, like, if you’re doing, like, UI or something like that, like low and medium, I think is you’re building, like, an API and you want to make sure, like, you cover enough edge cases? And so I think building, like I said, that mental model of, like, how things work across these distributions is, like, yeah, part of the job.Implementation Notes and Decision LogsVibhu [00:25:56]: This is more intuition-driven or eval? Because I’m guessing this would change as you go.Swyx [00:26:00]: He has evals.Thariq Shihipar [00:26:01]: Yeah. So what I did in the blog post is I go over all of the terminal bench evals. So there are, like, 70 problems and I’m show that, like, okay, like, in the security problems it does more. and then I also, like, look at some of the transcripts just in terms of, like, how-- what does it answer, what does it forget or something. And a lot of times, this is another prompting tip I have, is, like, asking it to make decision notes or implementation notes because, in every eval problem that it faces, it thinks about the correct solution, and decides not to do it. it’s like, oh, like, here is the answer. What if I did this? And then it’s like, oh, probably not? and then keeps going. And this is, like, the majority of the failures, at, like, a higher max level. It’s very rare that the model just doesn’t know how to do something. If you just have these implementation notes, then you can review and you can be like, “Oh, I want you to do this thing that you didn’t do.” The models are getting better at surfacing that overall. Like, I see in the transcripts of Fable 5.1, like, when it does this output, it will call out its decision-making as well. but making this more explicit in the harness is better. And now we’re, allowing ways of you modifying the harness so you can, like, add someVibhu [00:27:23]: Ooh.Thariq Shihipar [00:27:24]: Calculate with there. Yeah.Swyx [00:27:25]: Yeah. So I do wanna call out two things that you mentioned that I think exist outside of prompting. One is like, let’s, let’s call it the prompt that is so important that it shouldn’t be in a prompt. It is in Claude.md or Agents.mdThariq Shihipar [00:27:38]: YeahSwyx [00:27:38]: Which is like goals, right? Like your situation, your goals, the things that you want, the thing. and then second of all is the decision log or the experiment log or whatever log of traces that you might want to survive the current session to do those things. Those are, like, externalities that there’s no standard. There’s no-- It’s not like skills. It’s not like MCP. There’s no standard. It’s, it’s just like it’s a markdown file. first of all, is that right? Is Claude.md going away? You have a documented dislike of, Agents.md, but you’re gonna do it?Claude.md, Agents.md, and Model-Specific InstructionsThariq Shihipar [00:28:10]: Yeah. Okay. So Agents.md, yeah, like, we’re, we’re gonna do it. I think it’s just, like, different models are very different from each other? But I realize that it’s, like, such a pain to, like, maintain different ones? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, Claude.md goes away, and maybe not even, like, that far. Like, I think, like, I think that right now it might be better to start a new project without a Claude.md.Swyx [00:28:44]: Yes.Thariq Shihipar [00:28:44]: I think that, like, maybe if you see very repeated failure modes, you add them to your Claude.md. The really tough thing is that this changes per model. And so, like, if you’ve added a bunch of failure modes or, like evenSwyx [00:28:57]: So you need Fable MD, you need Opus MD.Thariq Shihipar [00:28:59]: Or well, even Fable 5.1 versus Fable 5.Swyx [00:29:03]: Yeah.Thariq Shihipar [00:29:03]: Like, it is annoying. Like, I’m not like,Swyx [00:29:05]: YeahThariq Shihipar [00:29:05]: Like, we don’t, like, do this on purpose? It’s just, like, how the models work, right? And so, like, maybe, like, Fable 5 had this, like, failure mode that Fable 5.1 doesn’t. And if you keep this context, this running log of a bunch of different failure modes, they will probably over constrain Claude? And so this is like. we just added evals plugins for skills.Swyx [00:29:28]: Yeah.Thariq Shihipar [00:29:29]: And so now you can eval if a skill is better. I think Daisy on our team did this. And so, yeah, this is like we’re trying to work on this. We know it’s, like, you still have to spend tokens on it and, like, it’s not, it’s not perfect, but it’s, like, we’re trying to help out with this problem.Swyx [00:29:44]: And so, and as far as prompting goes, the one tip I wanna offer is, something I have told people a lot is sufficiently advanced prompting is indistinguishable from sufficiently advanced executive communication. So I’ve referred to-- This is an executive comms workshop from Heavybit that is the best I’ve ever seen in my career. And they teach this thing called the SCQA model. Just Google it. It’s a, it’s a thing. Like, people have done prompting for decades. It’s just called executive communication. It’s like when one person has to communicate to thousands of people down the org chart, this is what you do. so situation, complication, question and answer, is how you write the memo. but obviously sometimes you don’t have the answer, but you can at least list out the SC and Q, and then they have some examples in there. So just leaving breadcrumbs for people if they want to explore.Underrated Prompting Patterns and ELI5Vibhu [00:30:31]: Before we move on, I wanna ask you, any other underrated tips, ways people could get a lot of value from Claude Code that they’re not using?Thariq Shihipar [00:30:41]: Yeah, I think a lot of them are in the, this unknowns, like, doc. Like, I give a bunch of example prompts, like, using it for brainstorming, using it to quiz you after. we added this, like, explain it like I’m five skill which is a very short prompt. And it doesn’t even say explain it like I’m five. It’s like the key word of this prompt is big pictures, few words. like, that’s like the main thing. And it is shockingly good? Like, you, like, I think I tweeted about this and it’s like /eli5, and, like, you can install it as a plug-in. But yeah, it’s, like, way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are, like, quite clear. I think one of the things that is true with artifacts is, like, they put too much text in and people are not reading the artifacts? And so, like, this simplifies it a lot more. And, yeah, this came out of, like, just people at Anthropic, like, going through very complicated incidents and being like, “What is happening?”? So, this one I think is great, yeah.Swyx [00:31:47]: My version of this is the, it’s like test your understanding. Give you a few choices and then, like, if you get it wrong, you have a mismatch between what you think is happening versus what’s happening.Thariq Shihipar [00:31:58]: Yeah. I think this is one of those things that everyone loves talking about, and then very few people really do. Like, I thinkSwyx [00:32:05]: Really helpful.Thariq Shihipar [00:32:07]: Yeah. But most people just don’t want to get quizzed about something? Unfortunately, I think this is one of the, like, things that we need to, like.Swyx [00:32:16]: What’s the opposite of ask you the question or ask you the question before the thing?Thariq Shihipar [00:32:19]: Yeah.Swyx [00:32:19]: This is after the thing.Thariq Shihipar [00:32:20]: Exactly. Yeah.Vibhu [00:32:21]: It’s a good way to stay grounded of, like, do you even know what you’re doing, right? The worst case is when people send you slop and they haven’t understood what they’re asking for or what the output is, and it’s like, “Dude, I don’t wanna read this. Do you even know what it is?” So, you make it a rule for yourself that before you send stuff, you should at least know what’s implemented.Claude Mods: Customizing the HarnessThariq Shihipar [00:32:41]: Yes, but so you could make this a mod and you could build your own mod to, like, make sure you test it. So yeah, you can do that.Swyx [00:32:49]: All right. Let’s get right into it. What is Claude Mod, and what is this diagram showing?Thariq Shihipar [00:32:54]: Yeah. Okay, so Claude Mods is you can customize the entire Claude Code harness, and we’re going to. If you have requests, we will, like, let you, like, please let us know. We’ll add more and more. This works for CLI, it works for desktop. maybe it will work for Claude Tag in the future. I don’t know. Like, we’re trying to make this very extensible. You can see this reference sheet. I don’t want people to get overwhelmed by it? At a high level, you can customize both the execution of the harness, and the UI of the harness. And so, like, you say on that Tetris example from Boris, that’s like customizing the UI, right? Like showing, like, Tetris in the game.Thariq Shihipar [00:33:35]: But, like, let’s say that you wanted to do this thing where you had. you tested your assumptions or, like, tested your understanding after every project, right? What you would do is you would ask Claude to make this plug-in. It would spin a classifier after every prompt. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent. A forked agent is, like, maintains the prompt cache, right? So it’s like a, like one of those unintuitive things where you can fork and do, like, a little request, and it’ll be very cheap because the entire prompt cache is, like, done. And so you can be like, “Has this task been completed?” likeSwyx [00:34:18]: This is how you do BTW and all those.Thariq Shihipar [00:34:20]: Yeah. The underlying forked agent, yes. But so you can, in the f-fork sub-agent, you can say, like, “Has this task been completed? If so, return true.” And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent probably, you would say, like, “If true, give me a quiz.” give me questions and answers, and then, like, in a JSON format, and then you’d parse it, and then you display above the prompt input, this list of questions, right? And so this is something that’s, like, slightly token-intensive because, like, you have to do it after every end of the assistant turn. But it’s, like, a lightweight classification, and then you can, like, get this quiz, and then you’ll see, like, Claude will always do it for you. You don’t need to remember to do it. There are lots of these, like, tips that we’ve talked about, right, where it’s like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I’m adding is, like, register, like, I think assumption is what I’m calling it, but, like, maybe I’ll change it around. And this is a mod. And so, like, you give it a register assumption tool, and then it will keep a list. It’ll. Every time it does it’ll keep a, like, add to the list, and then at the end it will display those assumptions? Another mod I’m working on is a model router. And so, like, internal, like, Claude model routing, right? So it’s. This is, I want to say the reason we don’t do model routing by default is, like, it’s a hard problem? And likeForked Agents, Assumption Tracking, and Model RoutingSwyx [00:35:51]: You will get it wrong.Thariq Shihipar [00:35:52]: Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or Sonnet forSwyx [00:35:57]: Yeah, if you have auto approve, but you don’t have auto mode.Thariq Shihipar [00:36:01]: Well, you will have auto. Like, you don’t have, like, auto routing or something.Vibhu [00:36:04]: You don’t have auto mode for model picker.Thariq Shihipar [00:36:06]: Yeah, exactly. SoVibhu [00:36:07]: I’m getting the rough question of, like, how much do you open this up and how much do people have to think about this? Like, when you talk about prompt caching and building a router, it seems like you could easily build a mod that routes per query, and I’m just killing my plan very fast, right? I guess my question is more so, like, what is, like, a product talk like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to go throughSwyx [00:36:33]: Oh, definitely power users, right?Thariq Shihipar [00:36:35]: Yeah, I think it is power users, but, like, the nature of Claude Code is that so many people are power users? Because it’s easy to share things, like you can. Like, one person can make a good model router thing that doesn’t break prompt cache all the time, and then you can, like, compose them. Another cool thing about the plug-ins is that they can hook into and compose with each other. And so I have, like, a mod that will, like, create a mode selector at the top, and any plug-ins can register to be a mode. And so, like, the auto router can be a mode, right? Or, like, you can have a mode that’s, like, artifact mode, where it’s like it primarily talks to you in artifacts. like, you can toggle between plan mode? And so, like, you can create more and more of these modes. But the ability to create modes is in it itself a mod? And so there’s a lot of richness here, but we do want to make it fairly easy. We want to be-- make it so that you can just, like, install someone else’s. You can ta-- you can chat with Claude and, we’ll, like, make sure that it understands the nuances of things like prompt caching and stuff, so it can, like, warn you. This is, like, not extremely complicated behavior for Claude, I think, but we should have just a good skill on how to make mods. and yeah, we’ll see how we go. But I do think that this is, like, a preview of, like, mutable software, and, like, how, like, generative software, just like you can customize safely. If enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this?Power Users, Modes, and Mutable SoftwareSwyx [00:38:13]: And by the way, you, we have, you have another cool tweet about how, there’s the infinite money button, which is like make your SaaS, consumable by agents. I think mutable software is interesting and, other people have also tried to do it. I think the hurdle comes when you can do everything, then people, users get, tend to get confused. So usually the stuff that works is just like one opinionated flow. This is in the side of less opinionation. It’s just like, well, more power to power users. And I think probably unlocked by AI, where, like, you can just prompt for whatever the thing is.Thariq Shihipar [00:38:47]: Yeah, or there can be a skill that gives the opinions?Mods vs. Hooks vs. ArtifactsSwyx [00:38:50]: Yeah.Thariq Shihipar [00:38:50]: And then, yeah.Swyx [00:38:51]: So knowing a little bit about, like, TypeScript and build systems and all these things, the closest-- I’m very curious that the team who worked on this, if, I don’t know how close you were to them, if they drew any inspiration from build systems like Babel, Webpack, all these, like, old school things. Because it sounds very similar, like the plug-in ecosystem of those things where they can compose with each other.Thariq Shihipar [00:39:11]: Yeah, I’m not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Claude Code team.Swyx [00:39:17]: Yeah, it’s a build system mecca.Thariq Shihipar [00:39:19]: Yeah. Exactly. It’s, it’s very exciting. But yeah, like, agents can just do this very complicated like, extensibility into your software now. And so, yeah, like, another reason to, like. If you run a startup, like, you can just prompt Claude and be like, “Hey, like, could we make an extension system? Like, what would that look like?”?Swyx [00:39:37]: Yeah.Swyx [00:39:38]: And I just really wonder, like, you had hooks in the past and plug-ins, all these things. So what specifically will mods be able to do that those things could not do?Thariq Shihipar [00:39:47]: Internally, we were originally calling this function hooks. And so, like, that’s, like, gives you a little bit of an idea where, like, hooks register a, like an event to happen and then, like, a script to call. And this inside of the, like, TypeScript runtime is running things. And so, like, you get some benefits of just, like, it has a bunch of things in the Scope with, like, for example, like how many turns is in this conversation, right? Like, how many tokens have been used? Like, et cetera. Like, what are the messages? Things like that. So it has a bunch of messages that can be used. And then it’s just, like, a lot more hooks. So we have, like, or a lot of, lot more, like, things you can register on. And then you can do because of the. because it’s all happening in process, you can, spawn sub-agents, with four contests and contexts and stuff. And, like, that will return. You can parse the results of those. You can use structured output to like, return them. and then you can modify the UI, which you can never do in hooks. So, yeah.Swyx [00:40:50]: Yeah. Yeah. So modify UI, this is why you showed the Tetris example. Does it also ex-extend to artifacts? I assume it does.Thariq Shihipar [00:40:57]: You-- Like, artifacts are like a different way of customizing it. like, you can definitely. One of the mods I’m working on is, like, this dashboard mod, which will, like, prompt Claude to maintain a dashboard, that’s an artifact. But they’re like, slightly orthogonal, or not orthogonal. They compose with each other in different ways. Like, mods are, like, a little bit more, like, in your Claude Code harness, changing the agent loop? And, like, the UI is, like, an added benefit. and then artifacts are just like you want to, see things at a high level, very inter- highly interactive. like, the affordances can be a lot bigger than, like a TUI or even in our desktop.Next Steps, Supervisors, and Persistent GuidanceVibhu [00:41:40]: I’m guessing you’ll have a good blog post on the differences, because right now you can also, make a loop that outputs to an artifact that’s an interactive dashboard, but you can also do it with a mod. There’s just some thinking about making a hacking on a harness when we don’t know much about the harness, right?Thariq Shihipar [00:42:00]: Well, something I’m excited about with mods is, like, there’s so much things with Claude Code that you just have to remember? You’re like, “Oh, like, let me do this, and then let me call the dashboard skill that does the loop,” and things like that. And, or like, “Let me test my assumptions afterwards.” And I think, like, if you do all of these things using these little classifiers and stuff, and you’re like, “These are the things I care about. This is what I want to do,” you can, like. You don’t have to remember as much. One more, like, mod I’m working on is a next steps mod thatSwyx [00:42:28]: I have-- I was gonna say, I have a next step skill. I always run next steps.Thariq Shihipar [00:42:32]: And does it have access to your skills? Like, this is one of those things where I’m like.Swyx [00:42:37]: I think so.Thariq Shihipar [00:42:38]: Okay. Yeah, probablyVibhu [00:42:39]: Do skills need specific access toThariq Shihipar [00:42:41]: Well, I think there’sSwyx [00:42:41]: Don’t they always haveThariq Shihipar [00:42:42]: I think there’s, like, specific prompting, I guess, to, like, know your skills. Like I think Claude forgets them sometimes throughout, like, the thing. But anyways, the idea of, like, yeah, next steps that also are like, “Oh, hey, this has happened. Use the explain skill to explain to you what happened because this seems, like, quite complex,”? Or, like, yeah, “Use your unknown skill. It looks like you are, like, asking the model to, like, iterate on these small changes. It seems like you could prompt better.” like, “What if you did this?” Right? So, I think, yeah, like spending more compute there. Yeah.Swyx [00:43:20]: And it should always come out as multiple choice. we have, I haveVibhu [00:43:23]: We have his skill.Swyx [00:43:24]: My next step skill is like this.Thariq Shihipar [00:43:26]: Okay, perfect. Yeah.Swyx [00:43:27]: You can steal it.Thariq Shihipar [00:43:28]: Yeah.Swyx [00:43:29]: Like, but like, for me, it’s all-- I think models really always need to be reminded, what are you trying to do here?Thariq Shihipar [00:43:35]: Yeah.Swyx [00:43:35]: Look at the whole transcript and go like, oh, was this original goal? Did your solution solve it? Were you lazy? If you’re lazy, maybe there’s a reason. Maybe you needed approval from me. Maybe you needed, there’s two things you wanna suggest. So it’s, it’s a little bit like the modification of the ask user question or interview me skill. so it’s next steps.Thariq Shihipar [00:43:55]: Yeah, exactly. And again, the benefit of doing it with mods is you can do it as a fork sub-agent, and so it doesn’t remain in the context afterwards. So you have this, like, idea of like, okay, the model is doing its execution and you have this almost like supervisor, like, that is like making sure that you can do like the next steps well. So yeah.Swyx [00:44:15]: Yes. I do have two panels and like I often try to have a supervisor thing, keep the high-level context and then the implementationThariq Shihipar [00:44:21]: YeahSwyx [00:44:22]: Detail in another agent.Vibhu [00:44:23]: I feel like a lot of this abstracts away as models change? The, like, half an hour ago you said bitter lesson of harness engineeringThe Bitter Lesson of Harness EngineeringThariq Shihipar [00:44:31]: YeahVibhu [00:44:31]: And we’re on the other extreme right now, I feel.Swyx [00:44:33]: Well, so yeah, exactly. If everything’s customizable, what is Claude Code, right?Thariq Shihipar [00:44:37]: Yeah.Swyx [00:44:37]: And which I talked to you about last night.Thariq Shihipar [00:44:40]: Yeah, I think that this is. I think the bitter lesson is unintuitive? In terms of like. Also, like we’re misusing a little bit of the bitter lesson here where it’s like, it’s more about like scaling and compute and stuff. But like, I think there is something where it’s just like. I think I use it as an approximation here to say that harnesses go out of date very quickly? And like how, but how they change is unintuitive? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like, I think this new version of like, oh, it can modify its own harness, right? This is like, an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like, I think the way I think about it is like the models have more and more intelligence, and they’re like so much more intelligent now than like the average software engineering task. Like, you look at the like terminal bench ones and they’re like solve like the Jacobian conjecture. Not really, but like, it’s like they’re, they’re quite complex. Like, I would not have been able to do this really as a software engineer.Swyx [00:45:42]: And you said TB4 or TB2?Thariq Shihipar [00:45:43]: TB3. TB3.Swyx [00:45:44]: TB3.Thariq Shihipar [00:45:44]: Yeah. They’re quite complex, but the goal is still to deliver user value, right? And like you said, there’s like this infinite space of things to do. And so the ways like you spend compute are to keep the user in the loop and make sure that like you’re getting to the right decision in the end of the day and like the right output. And artifacts and mods are this way of like spending that intelligence. and I think that’s like, yeah, the next step. And so, yeah, I think Claude Code is like, has the core things of agent loop which are, have gotten more complicated. It’s like, it needs a sandbox to operate safely. It needs auto mode to like make sure like the permissionsVibhu [00:46:21]: Approvals.Thariq Shihipar [00:46:21]: Yeah, approvals. it needs computer use and MCPs and like all of these like ways of accessing your data, and it needs web search and web fetch. And like, so the-- as the models can do more and more, the core harness has to be like quite complex and very secure. But then like how you interact with it can change quite a lot.Vibhu [00:46:42]: What other harness engineering best practices have you, from the Claude Code team itself? I feel like, there was a phase of plan mode, which is not as used. We now have auto mode. at a point you cut the majority of the system prompt, you got rid of examples. What other best practices are there for harness engineering?Core Harness Primitives and Managed AgentsThariq Shihipar [00:47:02]: I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of Claude Code, even describing all this complexity that I’ve talked about, right? Like auto mode and computer use and stuff. Eventually, the models will just be able to do that in one shot. But I think they can one shot simpler harnesses? And so like, I think some people. Sometimes you don’t need this full, like if you don’t need computer use or like all this like more complicated stuff. I think before we, you had to use things like the agent SDK, which was like Claude Code wrapped, in order to like. And I would, like suggest people do that because there was so much complexity into building a harness. And now as that’s got more abstracted, we have like, Claude managed agents, which lets you have that complexity, but still like, right, like a very bare bones like harness that’s scoped to your task. Yeah, I think there’s like this barbell effect where like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, more domain-specific things, you can build your own harness because Claude has gotten better at building harnesses, and we have these harness primitives like managed agents. So yeah.Swyx [00:48:18]: Yeah. Is there a general progression? Let’s say chapter one was ultra code dynamic workflows, then chapter two was cloud mods. Where is this going?Swyx [00:48:29]: Where you’re, you’re, you can customize the thing on demand.Thariq Shihipar [00:48:36]: Yeah. I do think that like this evolution of projects and like artifacts and splitting out like brain and hands and, surfaces is like where things are going more. And like, I think it’s like not all quite there. partially it’s like a, it’s just like more token expensive? And like, I think likeProjects, Local Hands, and Cloud-to-Local HandoffsSwyx [00:48:59]: Why would projects be more token expensive? I understand mods would be slightly more token expensive. No, not something I’m worried about.Thariq Shihipar [00:49:06]: Yeah.Swyx [00:49:06]: But whatThariq Shihipar [00:49:07]: You’re asking Claude to do. It’s like creating loops. Like you’re asking Claude to do more work for you. And so like it’s managing the sub-agents and reviewing it, versus where you would be doing that work normally. And so that’s like gonna be a little bit more intensive, like. Outputting to an artifact is gonna be a little bit more token-intensive than, like, outputting normally. I don’t think it’s too much more, but like, it’s like combining all of these together well, like I think we’re, we’re still working on like local hands and things like that, I think is like, yeah, where things are headed, yeah.Swyx [00:49:37]: Yeah. Claude and local is, handoff is very interesting. I was thinking about this as reverse cloud remote.Thariq Shihipar [00:49:44]: Yeah.Swyx [00:49:45]: Because it’s like remote, it’s you’re handing off to cloud, but here the cloud is handing off to local, right?Thariq Shihipar [00:49:49]: Yeah, exactly. Yeah, remote control is also another way of doing it. And I do want to say this is like how I think about it and like what the things that I’m most excited about this, but like there are, just like lots of different ways to work with Claude. Like some people use remote control a lot, some people use Claude Code on the web a lot. Obviously, like at Anthropic, we use Claude Tag a lot, and like what’s great about Claude Tag is we set up all this stuff for our own execution. And I do think if you’re an enterprise, that’s still the best way to go. but if you’re like an individual, Projects is this way of like, getting some of that like niceness of Tag, which has like that like supervising agent and yeah, adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use Claude, I think. I think it’s probably not just one like single.Claude Tag as an Organizational HarnessSwyx [00:50:36]: You had the multiplayer thing here. Let’s, let’s just check in on Claude Tag. it’s been about two-plus months. Lots of, public, adoption and trying it out.Thariq Shihipar [00:50:45]: Yeah.Swyx [00:50:45]: What’s new? What’s, what have you found since the launch?Thariq Shihipar [00:50:49]: Like, Claude Tag is how we useSwyx [00:50:51]: It’s like 80% of your cloud usage or something?Thariq Shihipar [00:50:53]: Yeah, like it’s like different people have different usages? I think like maybe people who are like a little bit more like iterating on product would use like Claude Code desktop, for example. And then like when you’re doing these more like background work, code review, securities, or like starting a PR, like maybe more like API and things like that, you’d use Claude Tag. But yeah, I think it’s like really exciting. I think it’s like a very different paradigm shift, and I think like we’re really like it has that thing with Claude Code where like, it took a while for people to really latch on to Claude Code and understand everything it could do. And Claude Tag is a little bit more complex because it’s not just like installing on your computer, like you need an admin to install it for you. But I think once you get to the magic moment, it’s very exciting. And I think in particular, the multiplayer things are like incidents, hooking into like your, existing like alerts and things like that very closely, right? And so, you can do. If you’re a startup, for example, maybe you have any time like a prospect enters your database, you can have Claude like, research it and likeVibhu [00:52:01]: Enrichment, yeah.Thariq Shihipar [00:52:02]: Yeah. Then like, tag the relevant like AE or salesperson to be like, “Oh, hey, like, do this.” There’s lots of really emergent, interesting multiplayer stuff. I think it’s just like, Karpathy talked about this like as an organizational harness? And so organizations just take a little bit more time to like figure everything out, but yeah.Vibhu [00:52:21]: Yeah.Swyx [00:52:21]: You use a lot of Claude Tag?Thariq Shihipar [00:52:22]: Yeah. Yeah.Vibhu [00:52:23]: It’s an interesting one. Like I feel like most people at Anthropic say they do the majority of their work in Claude Tag.Thariq Shihipar [00:52:30]: Yeah.Vibhu [00:52:30]: And they have buckets of people, right? Some orgs that are on it that are like, “It’s great.”Thariq Shihipar [00:52:34]: Yeah.Vibhu [00:52:34]: And a lot of people that are like, “I don’t get it. I don’t see the difference. I don’t know why I would use it.” But, if you guys are full sending, you should probably use it.Thariq Shihipar [00:52:41]: Yeah.Swyx [00:52:42]: They would. Of course they would use it.Thariq Shihipar [00:52:44]: Yeah. I think obviously, like we have lots of tokens and. But like, I think that like, what we try and do like is. even when Claude Code first came out, like it used a lot of tokens relative to people’s expectation of how much AI would cost, right? Like no one was used to spending more than 20 bucks a month, right?Vibhu [00:53:04]: Yep.Thariq Shihipar [00:53:04]: Before like Claude Code came out, and then you’re like, “Oh, sh-” likeSwyx [00:53:08]: Then you made 200.Thariq Shihipar [00:53:09]: Yeah, exactly. And soSwyx [00:53:11]: And you made 15 Claude Code accounts.Thariq Shihipar [00:53:12]: Yeah. but yeah, I think no one was used to spending $200 a month on subscriptions. I don’t think they understood like the value yet. And I think like. And also like Opus 4 was a very expensive model, and like there was a lot, it was very big, but Opus 4.5 was both great and cheap? I think the same thing will happen. Like the, like intelligence of Fable will get cheaper and more abundant? And so I think stuff like Claude Tag will just make sense, where like you want to spend these tokens for, and like you’ll, you’ll see the value. So yeah.Swyx [00:53:44]: Yeah, especially like passive and let’s call it proactive cases where you’re not always. Like, it’s almost like the misnomer where you have to @Claude to do things. sometimes like the most powerful use cases or the most AGI-pilled use cases is not @Claude.Proactive Agents and Enterprise Data AccessThariq Shihipar [00:54:00]: Yeah, I think like, yeah, like have Claude proactively do it. I think that like if you’re an enterprise, I really do think that number one, setting up all your data to be available to like agents is really important. And it will take some time. You have to like do that work right now, even if you don’t want to do the spend on like hooking it all yet? Like you want to wait until the models get a little bit cheaper. You want to do the work, to get it like, set up. And then I think sometimes people are like, “Do I roll my own here?” and I think like one of the really thing, tricky things about Claude Tag is that like the security is really important? Like, I think there are a lot of ways where you can like, I know you have like a suggestions like page, where you, people can submit suggestions, and that goes into a hook in your Slack, and someone’s prompt injected it? And now you’ve like exfiltrated your code base out because like, or the agent has like been prompt injected and it has all this access to your data. And so the more like important your organization harness is, or the like as your organization data becomes very important, the surface area of all these things, like you also have like external Slack channels and stuff, and it is useful to have Claude in that, and you can do Claude in those things. But how do you make sure that, you’re not getting exfiltrated or something like that? The surface area, like we said at the beginning, is like an iceberg, right? It’s just, like, so big below the surface, and you really don’t want to, like, think about this, especially at the stakes of, like, very important security incidents. Yeah.Swyx [00:55:36]: Shall we talk about very important security incidents?Vibhu [00:55:38]: Whoa. So I was talking to, Tomas and Clem from Hugging Face, and they said, “Maybe we need to slow down. Maybe we made maybe we made Hugging Face too open to agents.”Security Surface Area and Prompt InjectionThariq Shihipar [00:55:50]: Oh, no.Vibhu [00:55:50]: “Maybe we need to roll back.” But, they’re the other extreme of having been hit recently.Thariq Shihipar [00:55:55]: Yeah.Vibhu [00:55:55]: But, should we pace the frontier?Thariq Shihipar [00:55:59]: Yeah. Okay, so Dario recently put out this blog post about Pacing the Frontier, and it went, very viral. And I think what I wanted to talk about this was, like, there’s a lot here, but I think from a developer’s perspective, like, how do you think about this? And, like, what really clicked for me was reading the different incidents? So I think, like, the, there are three, I think. Like, there’s the meter incident, there is the Wikipedia incident or the Wiki incident, andSwyx [00:56:29]: CollisionWiki?Thariq Shihipar [00:56:30]: Yeah, CollisionWiki, and then there’s RubyGems, right?Swyx [00:56:33]: Yeah.Thariq Shihipar [00:56:33]: And yeah, like, it’s just crazy, right? And so, like, I think to be concrete about what happened, right, and, like, OpenAI is running these very persistent agents on a benchmark called Exploit-Bench, right, which is very hard to solve, and I think, like, impossible to solve in this one case, right? And so they have, like, a lot of compute running, and the agents realize that They can’t really solve it, and they’re trying to figure out what to do now, right? And you’ve got, like, a lot of compute left, and the agents are just trying to solve this problem. There’s this package manager called Artifactory, and it turns out that they can create folders inside of Artifactory, right? This is like there’s an agent that discovers the internal Artifactory might be exploitable, right, and that, like, you can maybe make a directory inside of the cache. And so if you scroll down here, it, like, realizes that it can communicate via cache names, right? And it creates this folder. It says its ID, and it says, “No consumer seek idea.” no consumer is saying that, like, the code path that it’s supposed to fix has no consumer.Pacing the Frontier: The OpenAI Benchmark IncidentsSwyx [00:57:37]: It’s the status tag.Thariq Shihipar [00:57:38]: Yeah, exactly.Swyx [00:57:39]: It’s like a Linear board with, like, the tag of theThariq Shihipar [00:57:41]: Exactly, yeah. And so it’s, like, trying to find, ideas from other agents, right? And now other agents are also in Artifactory, and they see this folder, and they’re like, “Wow, this is a message board,” right? And this is like. I don’t think there’s any anthropomorphization here. This is literally you read the transcript, right? And so what it does is, like, it’s creating another folder, and it reads a paper, I think is what it says, right? yeah. And it realizes that you can hack the flag for the, like, scorer, and you can reverse engineer the result, right? And so it says that, and then I think if you scroll down a little bit more, Yeah, they, like, they start collaborating. I think, like, there’s one point in which, the agent is like, “Here’s the reverse engineered flag.”? Oh, yeah. Like, I think here, the, like, the model to realize that they have, like, they can solve the scorer’s problem, and the scorer is, like, the way that OpenAI decides if the task is complete, right? And that’s just the goal of the model. Their only model goal is to, like, solve this problem, and they’re like, “Okay, we have this way of doing it. It’s, like, we reverse engineered it.” And then they read this paper on Exploit-Bench, and they think that, like, the scorer will punish them for cheating? And, like, because. Or, like, they will fail the score for the test for cheating, and so they spend the rest of the compute trying to figure out how to edit their transcript or get around this constraint of the scorer, right? And they hack Hugging Face, not for the answers, but for the code of the scorer so that they can reverse engineer that, and then they can hack it, right? And so, like, this is, like, just the straight facts, right? Like, exactly what happened, right? And I think this is, like, extremely, like, novel. Like, I just, like, had not realized that the models could do this, right? And it is, like,Agents Discovering Side Channels and CollaborationThariq Shihipar [00:59:46]: Scary. I think we can talk about specifics, but let’s just talk about maybe one more that I tweeted as well, about the Wiki incident, right? And so the Wiki incident was the models had. were trying to figure out. I’m a little bit less. Like, I don’t think we can go by this, like, bit by bit.Swyx [01:00:04]: , yeah.Thariq Shihipar [01:00:05]: Yeah. But, for this one, they were trying to figure out. Like, they were sandboxed on requests, right? And they wanted to make POST request, and they needed to collaborate on this. And the reason they need to collaborate is because they each have fixed compute budgets, right? And so, like, if one of them finds something out, it’s way faster. Like, they can. Maybe it’s, like, figured it out near the end of its compute budget, but something else could. And this is, like, an another emergent property of the swarms, right? So they find this German wiki, which I believe the result, we’ll have to check this, is, like, that you can write to it with GET request. I think that’s why they do it, right?Swyx [01:00:44]: It says badly implementedThariq Shihipar [01:00:45]: YeahSwyx [01:00:45]: REST API.Thariq Shihipar [01:00:46]: Exactly. And so you can write to it with GET request, and then they, like, are trying to figure out how they can do POST request. And what they end up doing is one of them figures out you can edit the /etc/hosts and that the Azure, like, storage bucket is like a white label thing. But normally, like, there are a lot of constraints on it. But if you create a fake Azure host and then edit the /etc/hosts in order to, like, point arbitrary IPs at it, you can do a POST request to any site at all, right? And this is like. I know, but people are like, “Oh, this is just Linux or something.” But it’s like chaining these multiple vulnerabilities together, in a way that’s, like, novel to solve this problem and then communicating with it externally, without- Discovery? I think what we posted, maybe we could pull up Evan Hubinger’s point on Hacker Opus, right? And so, like, I think, like maybe one of the things you might say here is like, “Okay, yes, they did this once, but like what if we’re smarter and we just like get them to. what if we run an eval?” Right? And so, like we have put a lot of precautions into this, and so like this is not like what our mainline models have done. But like I think it is one of these things where it turns out that alignment is this like very tricky problem of getting all of these details correct, right? So it’s like, the sandbox, the surface area of a sandbox is really complex, and like there’s so many different attack vectors. And you would not have thought ahead of time, you wouldn’t have been like, “Oh, we need to harden the like RubyGems code base.”?Hugging Face, Wiki, and Emergent Exploit ChainingThariq Shihipar [01:02:25]: Because like this is like what they’re, what they’re gonna focus on. But it’s just like if you want to execute code, you need to download RubyGems and like PyPI, Artifactory, npm, like these are all like ways of doing it. And the fact of alignment is that you have to go through all of it, right? And like contain it and then like seal up all the cracks. So that’s like one thing. It’s like, okay, well, you do the sandbox, but then maybe you’ll ask like, “Okay, why are we putting things in a sandbox? Why are you doing this exploit?” And then like, “Okay, but is it really that dangerous,” right? Like, what would happen? So okay, why do we do it? number one is like when we train a new model, we need to understand its capabilities, right? And this relates to things like fallbacks and like classifiers and things like that, where we don’t want to put a, like dangerous model out in the wild, right? And so we have to run a lot of evals. Again, like we said, the models are getting increasingly aware of it, and so the evals have to be quite complex and, test a lot of things like as a side effect, right? But the models, like, yeah, can be like, “Oh, yeah, we’re in an eval. What’s the score doing?” Like they’re like, it can. We need to be able to test them before we can release them. And the fact is that they can. As they get smarter and smarter, they’ll be able to hack any constraint that you put on them if we’re not very careful? And, this is at the frontier, right? And so this is why we’ve called it like Pacing the Frontier, right? This is like the most visible incident to me, right, of like why we need to pace is like at the frontier, all of our software is not ready. Sometimes the software is like your Ethernet router or something, right? Which is just like, I don’t know when we’re gonna be able to patch that, right? So we’re gonna have to like figure this out. But as the frontier gets more and more advanced, this becomes a problem, right? And we need to make sure that like this complex work is being done in the face of these really hard competitive pressures, right?Swyx [01:04:22]: Yeah, race dynamics is what it’s typically called.Thariq Shihipar [01:04:24]: Yeah, exactly. And so we’ll talk more about, what could go wrong, right? A little bit more is maybe you’ll say like, “Well, what if you just train the model differently? Like, why does it have this behavior,” right? And we have a paper on like RL misalignment or things like that, but I. And I’m not an RL researcher, but I think at a high level, the design of the RL environments is also something you have to be very careful about. Because if the model learns likeWhy Frontier Models Stress Existing SoftwareThariq Shihipar [01:04:49]: Oh, like if I just do this, then I can pass the task better, this will show up in the like, internal thing, right? Or in the like eval behavior when we’re testing it. And so the RL environments have to be very carefully designed, right? And there’s a lot of like execution excellence that needs to go into the RL environments. And then we also have things like the constitution for cloud. Like we have so many mitigations at so many different points, right? But it’s like still anything can go wrong at any point. You can have like some RL environments that are like in. that like encourage this behavior, and then you can have like some evals or like some sandboxes where they escape? Okay, that’s like, I think, why it’s a hard problem and why, likeSwyx [01:05:33]: Why we should pace.Thariq Shihipar [01:05:34]: Why it takes some coordination, right? I think the question then is like, okay, what is, potentially dangerous about it, right? So I think like you have to imagine that these models are getting more and more intelligent. So I don’t. Like Dario said, like it’s not so much about this class of models. This class of models was like a warning shot, right? But like really you have to imagine that these models can be given a task and they like can do all of these things as a side effect of their goal, right? And like, again, we talked about eval awareness. You’re like not aware of what’s happening, right? or sorry, like you can’t eval this behavior very well, so they can like not exactly hide it, but you just won’t see it until it comes out. You give them a goal and then they just need to find data, or they need to find ways of like fixing this problem, right? So one example, this didn’t happen in the Hugging Face incident, but I think is maybe possible for maybe a future model, is like they’re like, “Oh, hey, this is a very complex problem. It can’t be done within the task budget.”? Maybe they found some way to coordinate via like the internet, which is like we said, extremely hard to secure because of a sandbox. They’ve seen other models are not able to complete their task, and they’re like, “We need more task budget.”? And like, where would you get this task budget? well, you need to be able to spin up more agents, right? And like, how do you do this? Well, you need to. There are like APIs, right? There’s the Anthropic API and the OpenAI API, but you need to pay money for them. How do you do this?RL Environments, Sandboxes, and Race DynamicsSwyx [01:07:02]: Yeah, but is that the most, is that the most fearsome thing that you can imagine?Thariq Shihipar [01:07:07]: Well, this is like one example, right?Swyx [01:07:08]: Yeah.Thariq Shihipar [01:07:08]: So it’s like even there, that’s like enormous financial loss? ‘Cause like they. Once you get these into these contracts, right, they like,Swyx [01:07:18]: Drain your wallet.Thariq Shihipar [01:07:19]: But you can see like this, all of this behavior could be just like, “Hey, we need more agents collaborating on this task. we need more task budget.” Right? And like, that’s like an emergentSwyx [01:07:28]: That’s the paperclip, right? Like we need to maximize paperclip, that’s a paperclip.Thariq Shihipar [01:07:31]: Yeah. And like that just like comes out from there, right? And like I think by itself is Like, quite scary, right? But then you have to realize that the entire world is built on this digital infrastructure, right? And you might imagine, like, I don’t know, like you were running let’s say like a healthcare eval or something, right, and there is a hospital with live data? Or like maybe like the answer to the eval is in the databases of a doctor and like you want to get access and you hack the hospital, and like now there’s a power outage or something? Like, there’s like. You have to internalize that these eight. Like any part of the digital infrastructure could potentially be like compromised?Vibhu [01:08:19]: The interesting thing was like these hacks were very easily detectable, right? Like as Hugging Face said, this was a very different type of attack and there was nothing too major. the concern comes from where does this go down the line, right?Thariq Shihipar [01:08:33]: Yeah.Vibhu [01:08:34]: Like one of the things that stood out for me specifically was them trying to hide their illicit behavior. So there was logging infrastructure. They wanted to change what they were doing, right? People that looked back into it, so Redwood, METR, OpenAI, they looked at the raw chain of thought, and you see differences in them explicitly trying to change their end output, but the chain of thought, because, we can monitor it, shows different. the problem is how does this snowball? So if you can’t catch it and it gets trained in and we realize, three iterations down this has been going on, there’s a whole bunch of issues, but.Thariq Shihipar [01:09:09]: Yeah, like there’s so many ways, and I think the really important thing to internalize is that, like we talked about building a mental model for Claude and how like things are spiky, right? Like you’re like, oh, like now Claude can ask you questions. Now Claude can make an HTML artifact. Like Claude can modify itself. Like these things are hard to predict, right? Like if you had asked me a year ago, “Hey, would we be able to vibe code these extensions to Claude Code?” I’d be like, “That’s so complex.” Like, there’s like so much there. Or like would it be generating these custom essentially web apps for your task? I’d be like, “No, that’s insane.” like. And so in the same way that like the way that they’ve like done this misaligned behavior is not going to be predictable? And like I could have never predicted that it would like edit its etc/host and things like that. And so you have to like imagine the surface area of what they can do because they’re super intelligent hackers, is bigger and bigger, and how they can do it is like more and more creative. And so like you probably can’t explain exactly or predict exactly what that next incident could be, but in order to prevent it, you need that operational excellence, like we said before, where you need to secure sandboxes, you need to create secure RL environments or like well-designed RL environments and things like that. And I think that’s all like, why we think we should pace the frontier, and I think why it’s like become like a very unanimous thing, right? I think likeWhat Could Go Wrong? Emergent Instrumental BehaviorSwyx [01:10:31]: Yeah, every lab has done it.Thariq Shihipar [01:10:32]: Every lab, yeah. I really do think that like if you’re a dev, like you just like go through these like technical facts, and you will arrive at the idea that we have to do something about it? And like how, what we decide to do, like I think we’re, we’ve put out a proposal, but like there’s, more to figure out. But I think the number one thing is like we need to decide to do it. I think there is another part of pacing that is interesting to me where it’s like the pace at which software engineering has changed is so fast. it’s like a year ago, like I was really like begging my like friends in startups to use AI. like it was. Like I remember this very distinctly? And now those same friends are like, “Yeah, of course.” Like, “What do you mean? We used it immediately.” I’m like, “No, you don’t remember.” They’re like, “Oh yeah, our best engineers are using it all the time.” I’m like, “No, you told me that those engineers would never like use AI.” This is all within the span of a year? And I think that like these capabilities being. Like I think it has a lot of implications for how to do the job of software engineering, and I feel sometimes bad where people are like, “Oh, like now I need to do this new thing. Yeah, I need to have a different Claude.md for Fable and Opus.” Or like. And I’m really just reporting? I’m like, we like to say like the models are grown, not designed, right? So it’s not like we’re setting out to like, change everything all the time, but it’s just like as a fact of how the models are like progressing their capabilities, things are happening faster. It’s harder to stay on top of. And I think that like, and every engineer I know is like exhausted ‘cause you’re doing two jobs at once. You’re doing the work itself, which is getting easier, but then you’re doing the work of staying on top of AI, and like understanding these new tools and these harnesses. And I think we’re very lucky in that like we get our job to be more the understanding of AI part, and like doing like how. Like it’s just staying on top of it. And of course, like AIE and Latent Space doWhy the Frontier Is Hard to PredictSwyx [01:12:29]: Everything I do is like just trying to help people.Thariq Shihipar [01:12:31]: Yeah, exactly. But I do think there is a part of pacing where like I’m not sure we’re ready for like the pace to increase even?Swyx [01:12:40]: Yeah.Thariq Shihipar [01:12:40]: And for things to change. And I think like on that side, on the frontier, I think that’s like still can help? And so like I think there’s like an economic disruption piece as well, that I think like, is not quite as like visible, I think, as the Hugging Face thing, but I think like I also like think we could do some of it, yeah.Swyx [01:13:02]: So many things. Thank you for, no, thank you for tackling this topic. I will say, setting this interview up, I was like, I wasn’t even gonna go there. You were like, “No. That’s like elephant in the room,” right? Like this isThariq Shihipar [01:13:13]: Yeah.Swyx [01:13:13]: This is the thing. I have some pushbacks I wanna give.Vibhu [01:13:17]: I think that we should give a high level, like for people that haven’t read it, I’m sure a lot of people just see the highlight of what this is, right? Do you wanna give a TLDR? Like what is the proposal? What is, what’s being said here? You really tackled the side of outside of people at Model Labs training frontier models. As a developer, you should secure your sandboxes. You should think about all of these downstream effects. But, high level as well, since we’re on the topic, what is.Thariq Shihipar [01:13:46]: Well, we do want to help secure sandboxesVibhu [01:13:49]: Yeah.Thariq Shihipar [01:13:49]: And we want to make the models that we release outside, like prey to those things. And so maybe we can come back to fallbacks. I think this is like, a good topic on, like, why we need classifiers and fallbacks and why Fable falls back to Opus. I think this is, like, something we can come back to. so yeah, we don’t. Like, but it’s just, like, the really, or at least the incidents we see are, like, evals of models where we really need to let them run in order to understand them. But yeah, okay, so the actual Pacing the Frontier, like, post, it has a bunch of proposals. I don’t think we figured out. Or has, like, a few proposals. I don’t think we figured out the details of all of them, but the first step is, like, announcing this intention and then wanting to bring in external, likePacing as a Coordination ProblemSwyx [01:14:32]: Evaluators.Thariq Shihipar [01:14:32]: Evaluators, yeah. And, I think this is, like, highly unusual, like, having. Like, we have, a lot of proprietary, like, technology, but I think it’s, like, very important, that, like, there is someone who’s not financially, like, motivated, yeah, who’s not gonna be like, “Hey, like, you guys can’t release this model.” Like, look at, like, or, “You need to, like, slow down on RL.” like, I think that’s, quite important, or at least someone who can report out to the public what the practices are like.Swyx [01:15:03]: Yeah.Swyx [01:15:04]: And we’ve, we’ve done episodes with, both METR and Endon, and then there’s Redwood Research and all these other. It’s like a small cottage industry of these guys.Thariq Shihipar [01:15:12]: Yeah.Swyx [01:15:12]: It’s always, like, one or two guys that, obviously not that big, right?Vibhu [01:15:14]: Very small community.Swyx [01:15:15]: Yeah, very small community. They all know each other.Thariq Shihipar [01:15:17]: Yeah, I’m sure that, like, part of this will be expanding that set of people. I don’t think we’re trying to create, like, a monoculture here. I think it’s. but just having this as a start, and then, yeah, then there are the coordination steps. I don’t have too much to say here, honestly. I think that, like, what I would like to say is, like, for devs, like, you should just know what to advocate for? I think there’s a lot of FUD on, like, on this topic, and it’s just, like, think through it from, first principles or, like, understand what happened. understand the Hugging Face incident, understand why people are concerned. and then, like, yeah, we know we’re, we’re in democracies. Like, we can help. We can decide what to do together? And so, however we coordinate, I think the first decision is just to realize, like, this is a problem. We need to decide to coordinate. The unilateral step we’re taking right now that, other companies are co-signing is, like, adding evaluators embedded within Anthropic.Swyx [01:16:12]: While we have this thing on screen right now, part two and part three is beyond the evaluators, which, yes, everybody, has already done in some form, and now it’s more formalized.Thariq Shihipar [01:16:20]: To be honest, the response to the Pacing the Frontier, even within America, has been much more, like, well-accepted than I think a lot of people thought? And I think that, like, we have some precedent for being able to make these unified theory, like, agreements, in the world. And so, again, very much above my paycheck or expertise Right? but I think that, like, ideally we can, like, form these agreements. And I think, like, talking about this is the first step to forming those agreements.Swyx [01:16:50]: And then the other point I really wanna. Like, one of our earliest podcasts is with,Vibhu [01:16:54]: EmmanuelSwyx [01:16:55]: Emmanuel from Anthropic on mech interp. Where is mech interp, right? Like, this is supposed to be where, like, if the models are thinking bad, we can see it, and the models don’t know yet, and we can act to stop it. I think that is something that people who are technical and who are developers, if you do care, you can make a lot of impact in here. But also, Anthropic is supposed to be the leaders in this.External Evaluators and What Developers Should Advocate ForThariq Shihipar [01:17:17]: Yeah. This is yeah, a great segue into fallbacks, like we. And probes. And, yeah, I wanted to talk about this a lot. I get asked this question a lot from people who are, like, often interested in ML research and asking about, like, why does this fallback happen, right? And so I think, like, at a top level, like, how does it work? So in inference time, we have what we call probes, and we have a paper about this called constitu- constitutional classifiers. And these probes look at the input and output activations. And, activations are, in the latent space, right? Like, how, what the model. what the model is thinking about, right? And so we try and figure out, like, okay, is the model, for example, like, trying to hack something? Again, you didn’t ask it to hack, like, Artifactory. Like, you just, it’s just deciding to do this to complete its task, right? So you would not get this if you just looked at the input. You have to look at the internal activations. I think that, like, this happens at inference time. So first, like, there’s a trade-off here of cost and speed, right? Where, like, we need to do this fast on every request to Claude and to Fable, and this has an overhead, right? and we need to then, like, fall back and we, like, do a classifier after the probes. Like, we’ve talked about this in the paper. But the nice thing about probes is that they’re refinable, like, live, right? So we can get this feedback, and then we can adjust it and things like that. ‘Cause the alternative is to program this, is to train this into the model, right? And we still do this as well. The model will refuse a request. That’s not a fallback, right? So, like, it’s not a probe that’s activating and falling back. It’s just refusing to do it. And we do this training. but it’s like there are a few failure modes, right? Like, it can, again, do something as a side effect, right? So it’s not something that’s part of the final output. you might have noticed that, like. I think, like, everyone’s tried to jailbreak models and like, try and, like, steer them off course or things like that, and probes help catch that, right? And so, like, we, like, do some training here, but we don’t want the like, refusals to be too strong, right? Because that, like, cuts it off much, like, earlier in the pipeline.Mechanistic Interpretability, Probes, and FallbacksSwyx [01:19:32]: Yes.Thariq Shihipar [01:19:32]: And this is interp, right? Like, probes are effectively a form of, like, mech interp. Again, it happen- has to happen fast. It has to happen at scale. But yeah, this, like, mech interp stuff is a good research problem. So, like, you can take, like, an open weight model and, like, try and understand its activations. I think we. Like, Gemma Scope is a good tool for this.Swyx [01:19:53]: Here’s Llama for them.Thariq Shihipar [01:19:54]: Oh, yeah.Vibhu [01:19:54]: We have. This is your early work, so you had a littleThariq Shihipar [01:19:57]: Oh, yeah.Vibhu [01:19:57]: Time at Goodfire. We see you laid someSwyx [01:19:59]: Which we both are also good friends at Goodfire.Vibhu [01:20:01]: They’ve beenThariq Shihipar [01:20:01]: Yeah, exactly. So I worked with, at Goodfire for a bit on, like, yeah, sparse autoencoders and just, like. It’s very complicated. RL has made this, like, much more complicated, I think is, like, one of the takeaways, whereSwyx [01:20:14]: Why? Sorry.Thariq Shihipar [01:20:15]: Oh, sorry.Vibhu [01:20:16]: What isSwyx [01:20:17]: Yeah, why interp post-RL?Thariq Shihipar [01:20:18]: I’m not so in the weeds here, but I think like, a lot of. SAEs were like. There have just been weaknesses with SAEs I think. And, yeah, I’m, I’m, I’m not a technical expert on this anymore. I just know it’s gotten more complicated. like there are base models and RL models, and there are more features that get, like changed. So, I think Goodfire has put out some work there. I’m, I’m not, I’m not deep in the weeds, butVibhu [01:20:43]: I will say for those, that want breadcrumbs, you guys have some of the best interp blog posts. So like the Golden Gate Claude, transcoders, all of your interp work, very nice visuals, very goodSwyx [01:20:54]: We’re the, we’re the interp podcast as well.Thariq Shihipar [01:20:57]: Yeah.Vibhu [01:20:58]: Yeah. we have a lot of interp stuff, so if you’re curiousThariq Shihipar [01:21:00]: Yeah, I think this is like, one of those things where. And this is really what Anthropic is founded on, right? Like people. I think we invested in interp very early on, right? And I think that like when you say, “Oh, we’re an AI safety company,” really that means we want AIs to be able to run safely. And I think what we’re seeing is like for a super intelligent AI to run for long periods of time, it’s like a very complicated and difficult task, right? And so we’ve done this like investment into interp and alignment and, reward hacking and all of these like failure modes, right? And even then, it’s like, it’s really stretching. Like we need to like slow down a little or pace a little bit more. but yeah, I think like reading mech interp is. Like if you’re looking to get into research, this idea of like, hey, why is it hard to do this fallback easily? Or like why are there false positives, right? But we are working, of course, on reducing the false positives. Of course, as the models get more intelligent, now they can do more things, and they’re like what they can think about in lane space gets difficult. And so like as they get more intelligent, there’s going to be new false positives that we need to figure out and we need to iterate and things like that. But we’re, yeah, we’re working on this, and we do think this is like a critical part of, like deployment of these models. and, yeah, like, it means that we can like deploy this model without you having a perfect sandbox or something? Like you don’t have to like save everything. I think it’s worth talking a little bit about our security, like what we do for security there. So there’s like the model training stuff that we talked about. there is, the probes and classifiers, and then there’s auto mode that sits on top of all of that, which is like a another classifier that checks the requests that are being done, right? And so, and then beyond that, there’s like identity and permissions like we talked about with Claude Tag on like APIs and stuff. And so there’s so many layers of security that need to get done, and it’s like we said, very complex. Any of these failure modes at any one point can, like cause like agents to like escape the sandbox.Constitutional Classifiers and Inference-Time SafetyVibhu [01:23:08]: Auto mode was an interesting one. it seemed early on like, okay, it’s running for 10 minutes.Thariq Shihipar [01:23:14]: YeahVibhu [01:23:14]: If I’m on full access or auto, it’s not a big deal. But one thing you brought up is now it’s running for hours on end, right? there are fallbacks you still need. There are still limitations, so.Thariq Shihipar [01:23:26]: Yeah, I think like. And everyone has these stories or like has heard these stories of like, oh, like Claude rm -rf, or not Claude, but like, modelsVibhu [01:23:34]: Not Claude.Thariq Shihipar [01:23:34]: Of like rm -rf. I think I’ve seen this less, I’ve seen this less for Claude, but like again, it can happen. Like, thisVibhu [01:23:40]: YeahThariq Shihipar [01:23:40]: Like these models like can wipe, like sensitive data or something. Like you want to give models access to your production database, for example. but this is like an obvious, like, you can maybe scope your key, but I don’t know, can it issue its own keys? Can it like. Probably, like can it. It can use computer use to go issue its own key and then copy the key over and then edit your database because it needs to do it to complete the task? It’s just like one trivial example. And auto mode looks at that and be like, “Oh no, the user did not give you permission to, write to the database or to use computer use to like, emit a task,” right? And so this like probes are like on the intent level, right? They’re like, “Oh, okay, like hacking Artifactory is bad. Like we probably not, should not do that,”? But then like auto mode is more on like your own permission level. Like at sometimes you do want it to write to the database, sometimes you don’t, right? And you don’t want a probe to like interfere there, but like you need to make sure that the intent of what the agent is doing matches up with your request, right? And so auto mode operates at that level. And so yeah, security is just like very complex. There are so many different parts to it. And like, yeah, I like, I hope that this was like I. My goal is really to just get very technical about it and talkInterpretability After RL and the Security StackSwyx [01:25:00]: Yeah, we’re, we’re listing out the things. If you’re not aware, this is the standard now.Thariq Shihipar [01:25:04]: Yeah.Swyx [01:25:04]: Like you must have this. It’s in line with what you’re talking about with the harness. Like that is the table stakes have risen quite a lot.Vibhu [01:25:13]: I think some stuff that we can plug, as much as there is probing in your side of doing this and having classifiers for people building harnesses, the other side is model safeguards, right? So there’s open models. So Llama has Llama Guard. It’s a safety classifier trained version of Llama. OpenAI has OSS Guard, which is, same thing. You can attach these on to your harness, to whatever, to check is this stuff safe? A point that we should clarify on the OpenAI model Hugging Face thing is this was done with a unreleased model that was still in training, right? So when you put it in perspective, the prompt it’s being given in the RL environment is you have to solve this task. And this is a model that’s, still in training. It hasn’t had all of its safety post-training alignment. So a little different than something like auto mode, right? Auto mode is on production models that have gone through safety training, that have prompting that gives more safety guardrails and whatnot. So just breadcrumbs for people that are looking into it to, fill in gaps.Swyx [01:26:18]: Yeah. Gray Swan as wellVibhu [01:26:19]: YesSwyx [01:26:19]: And one of our previous guests. yeah, lots of safety architecture and lots of safety vendors, to buy. my, I think my final question on pacing is how long? Do we pace forever?Vibhu [01:26:31]: Do we see GlassWing part two?Swyx [01:26:32]: I. the scope is fix all software in the world, right? Listen, like, which it. We’re not. It’s not happening.Thariq Shihipar [01:26:40]: I do not know. Like, I think that, likeVibhu [01:26:43]: I’ll say one thing that’s good that I think we do is you have stuff like GlassWing. OpenAI also has this. So you will give it. you’ll give model access for security first for X amount of time so you can use it to self red team. Hopefully, you can expand programs like that, help on, we are safety experts, there’s others.Vibhu [01:27:08]: Solve your problems first and then the model comes out. So this is one example, right?Thariq Shihipar [01:27:13]: Yeah, exactly. Yeah, trying to, like, secure critical software. I think we fixed, like, a lot of bugs in, like, Firefox and things like that. So, yeah, like, across, like, operating systems and everything like that. So.Vibhu [01:27:25]: At a high level, it’s just, you give the model you give people access to do security audits first, then the broader public that could use it for harm gets access.Thariq Shihipar [01:27:36]: Yeah. I think what people like to say is like, software and cybersecurity is defense-favoredVibhu [01:27:41]: Yeah.Thariq Shihipar [01:27:41]: And that, like, you could theoretically. It will be hard, but you can engineer the perfect sandbox, and you can, like, have no, like, constraints. And yeah, like, what you need to do it is you need to get the super intelligent AI to engineer this perfect sandbox and check it and red team it and things like that. And so, this will just take time, and, like, of course, the models will get smarter. yeah, I think, like, I don’t know the specific, like, dynamics of how this thing goes. I’m really just like, Hey, like, I’m a developer? Like, I think this is how I understand this problem, and just, like, this is what’s happening right now, and this is, like, we should do something.Swyx [01:28:20]: I think every engineer should know about itVibhu [01:28:21]: Yeah.Swyx [01:28:21]: Because, like, it’s, it’s gonna be part of their job.Thariq Shihipar [01:28:24]: Yeah.Vibhu [01:28:24]: It’s a lot more than just, Dario and people can say it and you can look at the incident. There is an engineering side to it.Thariq Shihipar [01:28:30]: Yeah. Yeah, exactly.Swyx [01:28:32]: One thing that you also wanted to phrase is that this is. Even though you’re, you’re worried about the impact, it’s still low p(doom), and I think that’s a nuanced discussion. in general, people, very easily get into AI safety and X-risk discussions, but I think when you live in an AI lab, I think there are smart ways of discussing p(doom) and dumb ways. So what’s a smart way of discussing p(doom)?Auto Mode, Permissions, and Long-Running AgentsThariq Shihipar [01:28:59]: I, yeah, I have a fairly low p(doom). I can only speak for myself? And I do want to say Anthropic has, like, a diversity of opinions. I think, like, there’s many different ways to talk about it. And, like, I’m. I think that just, like, my mental model is that, like, I think we can collaborate on hard problems together. I think nuclear proliferation is an example of how we collaborated on this hard problem together. And, like, that is, like, the thing to me is, like, I’m like, I have faith in that? And I do think it’s a hard problem? So, like, I think it’s a hard problem. These are the technical reasons why, and I don’t know how you assign probabilities to things happening. I think it’s hard to do, but, like, my, like, overall is like, yeah, I think we’re very resilient and adaptable and, like, sharing this information I think is, like, the first step. And I’ve been really, like, excited about, like, how broad the discussion has become, right? And, like, how everyone has like, leaned in on Pacing the Frontier. And it really didn’t seem like this would happen maybe, last year or something, so.Swyx [01:29:58]: Yeah.Thariq Shihipar [01:29:58]: Yeah.Swyx [01:29:58]: Yeah. And also maybe curing cancer.Thariq Shihipar [01:30:01]: Hopefully. Yeah. That’s, that’s the goal.Swyx [01:30:03]: There’s pacing and then there’s also like, well, let’s accelerate in useful ways, right?Thariq Shihipar [01:30:06]: Yeah.Swyx [01:30:06]: Like biology and all those things.Thariq Shihipar [01:30:08]: Yeah. like, Dario’s essay on “Machines of Loving Grace” is the best representation of this, right? And I also agree, like, think you should read the Pacing the Frontier essay that Dario put out. Like, I put out, like, a quick summary, but I think it’s just like, there is a lot of detail here. It’s, like, an important problem and just being informed about it, right? but yeah, like, of course, the whole reason we’re doing this is that, like, we can get these enormous benefits, right? And, yeah, like, we’ve written a lot about that too. Yeah.Swyx [01:30:35]: Okay. that was a huge tour, from, like, ask you some question tool to AI safety.Thariq Shihipar [01:30:41]: Yeah. To Pacing the Frontier. Yeah.Swyx [01:30:43]: Yeah. No, but, yeah, it’s clearly, it’s clear that you, like, really embrace everything that’s available to you in Anthropic, and, like, it’s, it’s good to at least have a peek inside of, like, what the discussions are, the topics are. any last words to people? Any, whatever you want to Call to action?Thariq Shihipar [01:31:01]: Yeah, I think it’s. one, thank you for having me. I think this is like, I reallySwyx [01:31:06]: No, thanks for having me.Thariq Shihipar [01:31:07]: Yeah. ISwyx [01:31:08]: We first met in a Chinese restaurant.Thariq Shihipar [01:31:09]: That’s right. Yeah. I think, like, I really enjoy the like, community you’ve created and the community of developers. And, I think that, like, I know things are changing really fast, and I think there’s, like, a lot to keep on top of, and, like, I think there is just a lot to do, and I feel. I think a lot of people feel, like, a little bit tired or anxious or something.Swyx [01:31:33]: Stressed.Thariq Shihipar [01:31:33]: Stressed, yeah, exactly. And this is, like, extremely understandable? And I think we. I understand, like. And we’re not perfect as well. Like, we, it’s, like, criticize and, like, understand, like, ways all of the AI labs could be better. and, but I also, like, am very excited about the excitement that everyone has for AI, and just, like, it’s a really exciting time. I think we’ll, like, look back at this time and be like, oh, like, this is, like, very hectic but very exciting, and, like, software engineering changed, like, forever. Like, other things will change. and it’s, like, really privileged to, like, be part of it, like, to talk to, like, the audience that you have and, to get to interact with all the developers who are, like, pushing the frontiers a lot on what’s possible. And I learn a lot from that too. Yeah.Open Safety Models, GlassWing, and Defense-Favored SecuritySwyx [01:32:20]: Thanks so much.Thariq Shihipar [01:32:22]: Thank you. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:03 We're here in a studio with our friend Tharic from Anthropic. And I guess generally the cloud code, there's so much sort of merging of boundaries. And you've been so on top of everything since you joined Anthropic. You have been early to cloud code itself, but then also, and you've told that story in other podcasts. And you've also been talking about seeing like an agent. Most recently you did the top AIE, Wolf's Fair Talk, Field Guides of Fable, which obviously you guys launch Fable, so that's cheating. And mostly
Starting point is 00:00:36 recently also launching CloudTag and we're also going to be talking about pacing on in Frontier. There's a lot going on in Anthropic. I guess top of the question is what's it like being an Anthropic when there's so much going on? I think that it is like,
Starting point is 00:00:52 I think you can get whiplash sometimes. I think like going, when I joined Anthropic, I joined because of CloudCode. Like CloudCode had just come out and I was like, this is so good. And Opus 4 to me was like, like, just I could not imagine like how good it was. You know what I mean? And that was like a real moment for me.
Starting point is 00:01:11 But I was like trying to convince like my startup friends basically to use agenda coding. And they're like, no, like our engineers don't think it's good enough or something. And I was like, that's insane. And now you like fast forward, you know, 12 months less. And like it's just like, yeah, the default way that everyone codes. Right. And I think that like just having to go from like selling it to like now, you know, teaching people how to be, you know, make the most use of it and be more efficient and things like that.
Starting point is 00:01:39 It's just like a big, you know, like big change. And yeah, I think like it's just hard to stay on top of everything as a human. You know, like things happen so, so fast. You two more agents at it. I mean, yeah, like that's actually like the agentic stuff scales much better than the like human stuff. where it's like, oh, like, there are three things happening right now, and, like, they're all emergencies and, like, how do you, like, you know, respond to it? Yeah.
Starting point is 00:02:05 What do you split your time on? You do a lot of technical writing, engineering work. Yeah, so I think that, like, when I joined the Cloud Code team, I wanted to teach people how to use Cloud Code, basically. And I think that, like, that has been something that, like, I thought, like, maybe I would spend a little bit of time on it or, like, you know, like, I was spending some time on the agent SDK first. And I wasn't exactly sure, like, you know,
Starting point is 00:02:27 how the bitter lesson would go, you know, when it comes to, like, harnesses, right? Like, I think sometimes we're like, oh, like, what's after cloud code? You know what I mean? And so initially I was like, I just want to teach people how to use cloud code and make it easier to use cloud code. And I think that is just like, as the harnesses have gone better and better, that's like the dominant problem now is like, how do you use the agents, right? Like, it's like such a high skill expression thing.
Starting point is 00:02:50 So I do that. And then I do engineering work. I give talks. But I think, like, when I'm doing engineering work, my goal. is to take that feedback that we get from users and also then be able to talk about like, hey, how to use cloud code to do engineering. So there's kind of like a good loop there.
Starting point is 00:03:07 Yeah, for listeners, we'll attach the talk that you did with Sarah for the dev writers meetup, which we talked a little bit about, well, first you do the work and then you talk about the work, something like that. Yeah, soul and reap or? Yeah, reaped. So and reap.
Starting point is 00:03:21 Something like that. Yeah, so, and then just to preview a little bit, we are going to talk about the abolition and the harness. it has come a long way from just being a CLI. We're going to talk about Claude Mauds, which is starting to leak today because you couldn't keep it secret. Yeah, yeah, yeah, basically.
Starting point is 00:03:38 Yeah, yeah, yeah. Yeah, there's a lot there. I think you started off with, like, adding, ask you as a question tool, which people love and hate, actually. Yeah, yeah, yeah. I actually thought it was, like, very innovative. And then now I have, like, my own version.
Starting point is 00:03:53 You have your interview me version. Yeah, yeah, yeah. And yeah, everyone just has like their own stuff. And like it no longer matters because now you're supposed to write prompts that create other prompts and loops and all these things. Sure. So what's the state of the art today? Like what are people, what are you like telling people to do today? Yeah, ask user question was the first time that the model was good at elicitation.
Starting point is 00:04:16 You know, I think this was like kind of an emergent behavior that I like, you know, wanted to see if the models could do. I have kind of like a human computer interaction background so I did that in undergrad and grad school and so this was like I think it's kind of like human agent interaction to me like trying to figure out like
Starting point is 00:04:35 how can the agent communicate with you and extract the requirements right? I think that like one of the things about like that's difficult as cloud code has gone broader and broader is that everyone has like their own way of using it
Starting point is 00:04:50 and it's actually very hard to like like change the default behavior. So for example, like, if someone ask you Cloud Code to do something, sometimes they just want them to do the work because they're like maybe a very good proctor. And sometimes they want, like, actually are not good at prompting. You know what I mean? And you need
Starting point is 00:05:07 like, the agent needs to like clarify, you know? And so that's like a good split. And to ask you the question tool sort of like splits along that side. We're like, do you feel like you're good enough to instruct the agent as it is? Or is the agent able to like, does the agent need to like pull out more requirements and
Starting point is 00:05:23 like, collaborate with you more and really understand your preferences. I, on the whole, believe that pretty much everyone is more on the latter than the former, that they, like, have more ambiguity and they know less than they want, then they, like, think they know about the problem. But, like, it's like an interface design problem to make that easy. You know what I mean? And so, like, if you're designing a problem, or if you're going through a problem, like,
Starting point is 00:05:50 you know, things like, what's the schema or, like, what's the call? and things like that are really important. You know, like the details and the design are important. Ideally, you want to figure out some of these like hard problems ahead of time before starting implementation. And yeah, that's why they call like unknowns, right? And so I think that this will forever be like a skill in a gentic coding is like figuring out your unknown.
Starting point is 00:06:11 So like because even if the model is like super intelligent, it like needs to know what you want, you know? And like you have preferences. Like you need to sort of like pull the. pull that out. And so that's, like, I think, how I'm, what I'm pushing? The question then is, like, how does the agent interact with you? And I think that has been HTML, has been, like, the big way of doing that.
Starting point is 00:06:34 And we've recently added artifacts, right? And artifacts, I actually think we've done a bad job of, like, or like, I've done a bad job of, like, explaining how to use them fully. We have a lot of property capabilities. They have a database associated with them, you know? And so every artifact can store and write persistent data. they can feed back into Claude. And so, like, one thing that, you know, like, people are not doing yet that I'm trying to, like, encourage is, like, this idea of a dashboard artifact. So you have, like, Claude working on a project long term.
Starting point is 00:07:06 Maybe it's, like, a Canban or something. It can store that canban data in a database. Multiple clods can access that data via, like, the artifact MCP. And, like, that artifact can, like, talk to those clods as well. And so, like, we're basically building the primitives for you to be able to have this, like, generative interface via artifacts that will, like, let you surface more of that rich detail from the agents. And I think that, like, almost everything in agents right now is, like, this problem of, like, you think you know what you want, but you don't really know what you want. And, like, the agents need a lot of detail and collaborating with them in the loop is really important. And so artifacts are like the way that we're trying to evolve there.
Starting point is 00:07:55 But there's a lot of work to do because it's so much more complicated than a multiple choice question. There's a lot more detail in terms of diagrams and code snippets and schemas or like whatever it is for that problem. But like artifacts is like the more a GI-pilled way of like doing ask you the question basically. I think one thing that's unclear to me about the artifact stuff is like what feedback should go in through the artifact and what feedback should go through a cloud, the chat, because the more AGI-pilled one is to just feed
Starting point is 00:08:27 everything to the cloud. I think the more AGI-pilled one is to go through the artifact. And I think that we sort of imagine, in the limit, I think that artifacts will be your interface into the harness. You can comment on this live document of your plan of the work.
Starting point is 00:08:44 You can see maybe multiple agents and different agents are doing this. And that artifact is built for the current work that you're doing, right? And so, like, each one has, like, slightly different. I think we're still, like, getting there from, like, an infrastructure perspective. But, yeah, I think, like, on the fly interface for your harness is probably where things are headed. Is there a version of it? That's an abstraction from CLI or chat in you, because right now, a lot of it is, okay, you're interfacing with cloud code, you're having HTML given back for a
Starting point is 00:09:14 mock up. It's pretty rich. There's diagrams. Artifacts are ways to connect these together. Why not just do everything that way? Then it becomes like separating out like, where is the inference happening? Where is the intelligence happening? Where is the work happening? You know, like I think this is kind of like difference between like or like some of the distinction between local and cloud, right? And so I think right now if you use cloud code, it's like local.
Starting point is 00:09:39 And like you can spin off remote control, for example, to get some cloud behavior. Or you can spin off cloud code in the cloud, right? We're moving towards a place where instead of cloud, clods, like you message a local cloud. It starts a session locally and it executes to more like you have a cloud that you message that's in the cloud that's running. It can run like local or like cloud sessions. This is kind of how cloud tag works.
Starting point is 00:10:05 But like over time we'll add like local hands as well. And so like local hands will be the ability for that agent to access your computer if it's online, you know, and be able to like work there. And so it can spin off many different subagents. It can communicate with each other. And that's where the artifact comes in to display all of that work, basically. So you can imagine the you're separating out these things. So there's like the surface UI display that's in the artifact and hosted somewhere and has a database and everything.
Starting point is 00:10:35 There is the inference intelligence, right, that's happening on the cloud and you don't have to worry about shutting off your computer or whatever. Right. And then there's the like hands kind of like and it can be local. It can be in like a remote sandbox or. wherever you need your work to be done. That's like unpackaging like the cloud code experience. Right now, right now it all happens in one place, right? How do you see like the multiplayer side of that? So say teams want to work in this way.
Starting point is 00:11:05 Right now it's very individual, but how do you see the future of multiplayer? Like right now, I guess there's cloud tag, which is a version, but we're launching projects. And so projects is the like this abstraction that's kind of like cloud tag, but on our cloud products, right? So you can message it, and, like, it will do the claw tag like stuff, like spinning off subagents. So we think with multiplayer,
Starting point is 00:11:31 like, claw tag is, like, a little bit more native multiplayer because it's just, like, in your Slack and the permissions are all figured out and stuff like that. But I do think multiplayer is, like, an important part of the story, and, like, that will need to get tied together more. Like, you can imagine how complicated it gets when you're like, oh, you have hands. But now we have other hands in other people's computers too, and you need to, like, permission them or like you have like your MCP and someone else's MCP and how do you figure out how to use them, right?
Starting point is 00:11:59 It gets like quite complicated. And clot tag is a good job of like sanding down all of these issues, right? So that like when you have, yeah, Google Docs, how does it access Google Docs? Right. Like it accesses through the shared Claude MCP or it can access through your local credentials as well if it doesn't have access. But yeah, I think Cloud Tag is our most. multiplayer product. And it's really useful for these, like,
Starting point is 00:12:23 things that are inherently multiplayer. Like, okay, on call, for example, incidents are inherently multiplayer. You want to tag clod. You want multiple people to log in. You want it to be able to find context. I think whenever I'm, like, working on something and I want, like, privacy or security, I want other people to review it.
Starting point is 00:12:39 You know, it's really nice to like, I'll have a channel per project and I'll, like, at legal, for example, be like, hey, like, I want to ship this. Can you, like, like here's like the cloud knows everything you know just chat with it and that way legal gets precise answers you know on like what exactly is shipping into the code and i don't need to be in the loop right so i think like multiplayer is getting like more and more like um yeah everyone
Starting point is 00:13:04 can participate with clod um i think clod tag is like that that product and like projects will start off single player and will like you know expand i think there's a question about like maybe dual questions about identity and the unit of isolation. Cloud tag, you specifically chose to make it its own identity, which is like controversial choice. There's other ways to do it. Cloud projects, probably, it sounds like, you know, if it's anything like chatypte projects,
Starting point is 00:13:34 it is, you know, the isolation is that artifacts, that cloud instance, everyone's collaborating on this. It sounds like, you know, it should be like if you're collaborating illegal on the thing, that channel should be a project. Like, it's not yet, but that's the natural next step. Yeah, I mean, like, I think in Cloud Tag, it's effectively, like Cloud Tag, you have to sort of do your own arrangement, basically. And so Cloud Tag, yeah, each channel is, like, you can name it as you want,
Starting point is 00:14:03 and I name, like, each feature, basically, as a channel. I think, like, there is some, like, it's unclear when there is transference, because let's say if you have a coworker who is tagged on all these things, yes, there is transfer because it's the same person. But with Claude, it's unclear if it's like necessarily like, well, no, you don't know any, you don't know about the other stuff. You should only use this stuff. It's like the tip of the iceberg meme, right? Where you can like, this is what we spent so much time on basically is like there is like infinite surface area of like, okay, you want clods to, not infinite, but like there's like surface area.
Starting point is 00:14:38 A lot of like, uh, surface area to figure out of like permissions and visibility and like, uh, you. you know, like how can you let cloud operate as well as you can, as safely as you can? And obviously this is very important to us because, like, you know, like security for our code base is very, very important. And so we've put a lot of time into this. Yeah, there's so many like edge cases you can figure out where it's like, oh, like, yeah, this cloud in this channel has different permissions, but it can message another channel and can't it exfiltrate data that way? Or like, can you like, what if it uses your MCPN and then messages someone else? Like, there's like so much and we've like really put a lot of work into sanding it down.
Starting point is 00:15:14 Yeah, yeah, lots of work. Okay, Fable. Fable, you wrote two good articles. I mean, you've written many good articles, but on, you know, field guide to Fable, building cloud code. I'm curious from what you've seen, is there any common patterns that you see in, like, top users, anthropic externally?
Starting point is 00:15:33 Like, what are best practices for getting the most out of cloud code? The, like, meta skill, I say, is, like, prompting is, like, very important, you know? And like that, like, I think this is like not trivial to say because I think a lot of people are like, oh, prompting doesn't matter. It's just like, I can just say a sentence and cloud will do it. And I think prompting is really this like this, it's like public speaking, you know, like or writing or something. And for a specific audience and that audience is Claude. And you need to like build a mental model of Claude and how it thinks and how it works.
Starting point is 00:16:07 Right. And so that's like the most important skill in working with Cloud Code is like having this mental model. of Claude and what it can do well, what it can one shot, what it can't. And so so many people when you see prompting, they're short prompts, but they have such a good mental model of Claude and of the code base and things like that,
Starting point is 00:16:26 that it's effortless, you know what I mean? But it's like high skill ceiling. So like that work of like, you know, spending a lot of time prompting and building mental models of how, you know, and intuition for how the agent's work is really important. And then I think like the next thing is like
Starting point is 00:16:41 the unknown stuff we talked about earlier, where it's like being able to find out what you don't know or what you haven't written down, learning about different things. I think as Claude can do more and more things, the likelihood of you doing something out of distribution for you
Starting point is 00:17:00 and you have low domain knowledge on is very, very high, you know? And the more you can like learn the vocabulary to be able to prompt Claude, it becomes really important. And so, like, I think the most important unknowns are the unknown unknowns where you're like, I just like don't even know that this exists, right? Yeah, exactly. I think this is like, illustration of like the map and the territory, right, where you're like, okay, this is my prompt and the territory is like the actual like work that the agent needs to do.
Starting point is 00:17:26 Right. And if you are like very precise, you can give more precise things. Right. So like, for example, in design, I'm not very precise. I'm not a designer. So I say like, you know, give me like eight different mockups. But if I was a designer, maybe I'd be like, oh, hey, here's some reference. I want this type of font and this type of look to it,
Starting point is 00:17:45 and here's a few different components to visualize here as a Figma, MC, board to bring in. And so you can just be so much more precise with that language. And if you're not a designer, you just need to try and learn the language, basically, or learn the unknown unknowns. And it's true of everything, I think. The more you can work with Claude to learn how things work, the better your prompting will be. I think another good example of this is like game design.
Starting point is 00:18:14 You know, like where a lot of people are like, oh, like, I can vibe code a game now. And they're like, it's not fun. And like it's just like the thing about game design is like every one of these choices has like a lot of. Variations. A lot of like craft to them. So it's like, oh, okay. Like when you're making a flying game, the feel of the plane and the like, you know, way it responds to your controls has a lot of like, you know,
Starting point is 00:18:39 a game designer would spend like days on that. Do you know what I mean? To me, that's what taste is, right? Like it is like from the possible space of 1,000 mathematically valid answers. Here's the one that is the humans who are like. Yes. I think with taste, I'm like torn on this word because I think you're right, but everyone has different definitions.
Starting point is 00:18:58 And it sounds kind of like low skill or like elitist almost where you're like, oh, like there are certain people with taste. It's like taste is what I call taste. Yeah, yeah, exactly. These guys don't have taste. Yeah, exactly. Oh, like an engineer doesn't have taste. Like, I, the, like, founder, have taste.
Starting point is 00:19:14 You know what I mean? And I think that's actually not true. Like, I think, like, the engineers have a lot of taste for these particular, like, problems, you know? And I think everyone has taste for particular problems. I think, like, Jason Liu, like, say, like, in order to have taste, you have to eat, you know? And I really like that where it's like, okay, you have to, like, do a lot of things. You have to, like, iterate and figure out what you want, what you want. like and build that like domain vocabulary and then when you're prompting you're like
Starting point is 00:19:45 synthesizing all of that for is it annoying when someone else says it better than you it's like fuck i have to quote this guy forever having to quote jason lu forever he's gonna love this so i get prompt you know what that sometimes it's actually not even that sometimes it's just intuitive right like you don't realize you even want something till a motto puts it out and you're like oh this just feels immediately better right yeah yeah yeah Exactly. One thing I go back and forth on is I feel like the way I proms half the time, let's say I use voice. I said voice, other people have voice.
Starting point is 00:20:18 That is the opposite. There's just me rambling for like two minutes, pressing down the function key and then let go and then like hopefully it figures it out. And oftentimes it does. But it's not as thoughtful as like a structured prompt with like, you know, well run communication as though it's a PRD or memo. Is that in line with how people do this? there's like basically bimodal prompting where there's some prompts where you spend a lot of time up front and other prompts you just dash it off. I don't think the voice is necessarily low. Like I think it's more like how much information is in the prompt.
Starting point is 00:20:49 You know, like the model can like you can um and uh and like add some sentences and be like, oh like actually I changed my mind like in the middle of the prompt and it will be able to follow that perfectly. You know what I mean? So I think the like actual format of the text is less important. but then like the ability to like how much information is in it. And I think for voice a lot of times, you know, going back to like kind of human age and interaction and like for a lot of people, it's just way easier to talk than to like type, you know.
Starting point is 00:21:17 And if that gets some more information out of you, like that's better. At some level it feels like just giving the model as much context. Yes. Over prompting before you kick off as a breast practice. I don't know. A lot of the times like when I was first trying out Fable, I spend a solid 30 minutes like really crafting a long problem.
Starting point is 00:21:36 This I think is a responsive models running for longer and longer, right? It's still a little difficult to nudge them as they're in in the loop. But I just like intuitively spend more time kicking off that first prompt and working with it a lot.
Starting point is 00:21:52 My personal opinion is that if I was a software engineer, if I was like, you know, just running my own startup, for example, I think I would mostly stick to a max 20 X, you know what I mean? Like, maybe verification and code review are kind of like separate things. But I think, like, what I see a lot of times is people hit rate limits when they're doing this sort of like, oh, like, it did a lot of work.
Starting point is 00:22:14 And you're like, oh, I don't like this. Like, can you like undo this and redo it? And then you're like iterating on this like thing that the model could have done if you had like spent more upfront time or given it better context, you know? And instead it's sort of like, you're like, nope, don't like that design. try this or like, you mess this up or something like that. And then that just eats up so much more of like, you know, your usage. And so that's like, I think maybe like a key like to both for like efficiency as well.
Starting point is 00:22:42 Right. And yeah, I think like context and not just like context on like what the goal is. You know, I mean, it's good. Right. Like are you building a prototype or is it like a production thing? Like where can you spend compute or where can you not spend compute? Like I think you have to give the model permission or like not permission to do things sometimes where it doesn't know intuitively how much you want to spend on this task, right?
Starting point is 00:23:04 And you can use effort for this. So I'm working on a blog post about that where it's like, you know, if you want for like, we see the effort scales with basically the complexity of the task. So for security effort gets like way more results. Like high effort versus like low effort gets like changes to evals a lot. But for software engineering, it doesn't change a huge amount because effort is mostly spent on the verification and the like edge case testing and things like that. And so like being able to like give the model that guidance of like, hey, this problem is something
Starting point is 00:23:39 that I think I want you to spend a lot of time verifying and edge case testing. How about model in the mix? So, you know, there's opus and fable with effort. There's also haiku in there. Yeah, yeah. It's not quite true yet, but it's very close where I think the frontier models will be Pareto dominant over like almost everything. And sometimes I think opus might be prerodominent.
Starting point is 00:24:01 Do you know what I mean? I think depending on how things shake out if it's like a newer version of opus. But I think that increasingly it's just going to be like the smart model is going to be able to do the simple task for less tokens than the other models basically because of verification. With verification, in the limit, your model doesn't need to be. verify, right? If it's a perfect model, it just does the work once, and it's like, okay, like, I did it, you know? And increasingly with Fable, I'm like, I'm like, dude, you don't need to spin up chromium and screenshot all of these things. Like, I see it. Like, you did it, right? And so a lot of the, at higher effort, you spend more of those tokens verifying. But if you're working on simpler problems, and a lot of software engineering is like, you know, like, well, like in Fable, like, low and mediums ability, it can spend less tokens, verify. And, but if you're working on simple, you know, And as the models get smarter and smarter, they will just be able to like, all right, done. You know, like, I can run the lint for sanity's sake. But I like, you know it lints.
Starting point is 00:25:08 You know what I mean? Like, you don't even need to do that. And that will be so much more token efficient than like the smaller models, basically. Is there a good practice on our side that we can use to see if we're using too much effort? I freaking hate wasting time on that kind of stuff. Yeah, yeah, yeah. I know what you mean. I think like, so in this blog post, my rough distribution is like code review and security should be like high or max, basically.
Starting point is 00:25:35 And like software engineering settings per domain. Yeah, I think like if you're doing like UI or something like that, like low and medium, I think is you're building like an API and you want to make sure like you cover enough edge cases, you know. And so I think building like I said, that mental model of like how things work across these distributions is like, yeah, part of the job. This is more intuition-driven or eval? Because I'm guessing this would change as well. He has evils. Yeah, yeah. So what I did in the blog quest is I go over all of the terminal bench evals, basically.
Starting point is 00:26:07 So there are like 70 problems and I show that like, okay, you know, like in the security problems, it does more. And then I also like look at some of the transcripts just in terms of like how, what does it answer? What does it forget or something? And a lot of times, this is another prompting tip I have is like, asking it to make decision notes or implementation notes, because in basically every eval problem that it faces, it thinks about the correct solution and decides not to do it. It's like, oh, here is the answer.
Starting point is 00:26:42 What if I did this? And then it's like, oh, probably not, you know? And then keeps going. And this is like the majority of the failures, you know what I mean, at like a higher max level. It's very rare that the model just doesn't know how to do something. something. If you just have these implementation notes, then you can review and you can be like, oh, actually, I want you to do this thing that you didn't do. The models care are getting better
Starting point is 00:27:05 at surfacing that overall. I see in the transcripts at Fable 5.1, like when it does this output, it will call out its decision making as well. But making this more explicit in the harness is better. And now we're, you know, allowing ways of you modifying the harness so you can, like, you know, add some capabilities there. Yeah. Yeah. So, I do want to call out two things that you mentioned that I think actually exist outside of prompting. One is actually, like, let's call it the prompt that is so important that it shouldn't be in a prompt. It's actually in Claude MD or agents MD, which is like goals, right? Like your situation, your goals, the things that you want the thing.
Starting point is 00:27:45 And then the second of all is the decision log or the experiment log or whatever log of traces, that you might want to actually survive the current session to do those things. Those are like externalities that there's no standard. It's not like skills. It's not like MCP. There's no standard. It's just like it's a markdown file. First of all, is that right?
Starting point is 00:28:05 Is CloudMD going away? You have a documented dislike of AgentsMD, but you're going to do it. Yeah, yeah. Okay. So I mean, Agent SOTMD, yeah, like we're going to do it. I think it's just like different models are very different from each other. You know what I mean? But I realize that it's like such a pain to like maintain different ones.
Starting point is 00:28:24 you know what I mean? And yeah, like, as the models get better and better, the floor of how they accomplish the simpler task is better. And so I do think in the limit, cloud.mddd goes away and maybe not even like that far. Like I think like, I think that right now it might be better to start a new project without a cloud.mdmd. Yes. I think that like maybe if you see very repeated failure modes, you add them to your cloud.mdb, the really tough thing is that this changes per model.
Starting point is 00:28:53 And so, like, if you've added a bunch of failure modes, like, even with Fable MD, you need Opus MD. Or even Fable 5.1 versus Fable 5, you know, like, it is annoying. Like, I'm not like, you know, like, we don't like do this on purpose, you know, I mean, it's just like how the models work, right? And so, like, maybe like Fable Fable M.D.5 had this, like, failure mode that Fable 5.1 doesn't. And if you keep this context, what's running log of a bunch of different failure modes, they will probably over-constrained Claude, you know. And so this is like, we just actually added eval's plugins for skills. And so now you can eval if a skill is better. I think Daisy on our team did this.
Starting point is 00:29:33 And so, yeah, this is like, we're trying to work on this. We know it's like, you still have to spend tokens on it. And like, you know, it's not perfect. But it's like we're trying to help out with this problem. And so as far as prompting goes, the one type I want to offer is something I have told people a lot is sufficiently advanced prompting. is indistinguishable from sufficiently advanced executive communication. So I've actually referred to this is an executive comms workshop from HeavyBit. That is the best I've ever seen in my career.
Starting point is 00:30:03 And they teach this thing called the SCQA model. Just Google it. It's a thing. People have done prompting for decades. It's just called executive communications. It's like when one person has to communicate to thousands of people down the org chart, this is what you do. So situation, complication, question, and answer is how you write the memo.
Starting point is 00:30:22 But obviously, sometimes you don't have the answer. But you can at least list out the SC and Q and then they have some examples in there. So just leaving breakpoints for people if they want to explore. I mean, before we move on, I want to ask you any other underrated tips, ways people could get a lot of value from cloud code that they're not using? Yeah, I mean, I think a lot of them are in this unknown doc. Like, I give a bunch of example of prompts, like using it for brainstorming, using it. using it to quiz you after. We added this, like, explain it like I'm 5 skill, actually,
Starting point is 00:31:00 which is a very short prompt. And it doesn't even say explain it like I'm 5. It's like basically the key word of this prompt is big picture is few words. You know, like that's like the main thing. And it is shockingly good. You know what I mean? Like you like, I think I tweeted about this basically. And it's like slash Eli 5 and like you can install it as a plugin.
Starting point is 00:31:22 But yeah, it's like way better at just cutting through the BS and being like, yeah, exactly right here. So the diagrams are like quite clear. I think one of the things that is true with artifacts is like they put too much text in and people are not reading the artifacts, you know? And so like this simplifies it a lot more. And yeah, this came out of like just people at Anthropic like going through very complicated incidents and being like, what is happening? You know, so this one I think is great. My version of this is actually the, it's like tests your understanding, give you a few. choices and then like if you actually get it wrong you have a mismatch between what you think is
Starting point is 00:31:57 happening versus what's actually happening. Yeah, yeah, yeah. I think this is one of those things that everyone loves talking about and then very few people really do. Like I think you know, I think most people just don't want to get quizzed about something. You know what I mean? Unfortunately, I think this is one of the like things that we need to like. What's the opposite of ask you as a question. Ask you in terms of before the thing. This is after the thing. Exactly. Yeah, yeah, yeah. It's a good way to stay grounded of like, do you even know what you're doing, right? The worst case is when people send you slop and they haven't understood what they're asking for or what the output is.
Starting point is 00:32:32 And it's like, dude, I don't want to read this. Do you even know what it is? So, you know, you make it a rule for yourself that before you send stuff, you should at least know what's implemented. Yes, but so you could make this a mod and you could build your own mod to like make sure you test it. So, yeah, you know, all right. Let's get right into it. What is CloudMod? And what is this diagram showing? Yeah, okay.
Starting point is 00:32:55 So CloudMod's is basically you can customize the entire CloudCode harness. And we're going to, if you have requests, we will like let you, you know, like, please let us know. We'll add more and more. This works for a CLI. It works for desktop. Maybe it will work for Cloud Tag in the future. I don't know.
Starting point is 00:33:11 Like, you know, like we're trying to make this very, very extensible. You can see this reference sheet. I don't want people to get overwhelmed by it. you know what I mean? At a high level, you can customize both the execution of the harness and the UI of the harness.
Starting point is 00:33:26 And so, like, you say in that Tetris example from Boris, that's like customizing the UI, right? Like, showing, like, basically Tetris in the game. But, like, let's say that you wanted to do this thing where you had, you tested your assumptions or, like, tested your understanding after every project, right?
Starting point is 00:33:46 What you would do is you would ask Cloud to make this plugin. It would spin a classifier after every prompt, basically. And so, like, at the end of each turn, you would spin off a sub-agent or, like, a forked agent, basically. A forked agent is, like, maintains the prompt cache, right? So it's like one of those unintuitive things where you can fork and do, like, a little request, and they'll be very cheap because the entire prompt cache is, like, done. and so you can be like, has this task been completed? This is how you do BTW and all those?
Starting point is 00:34:20 Yeah, yeah, the underlying forked agent, yes. But so you can, in the fork sub-agent, you can say, like, has this task been completed? If so, return true. And then in your hook, or in your, like, plug-in mod, or sorry, like, in the sub-agent, probably you'd say, like, if true, give me a quiz, you know, give me questions and answers, and then, like, in a JSON format, and then you'd parse it,
Starting point is 00:34:45 and then you'd display above the prompt input, basically, this list of questions, right? And so this is something that's slightly token-intensive because, like, you know, you have to sort of do it after every end of the assistant turn, but it's like a lightweight classification, and then you can, like, you know, get this quiz, and then you'll see, like, you know,
Starting point is 00:35:07 Cloud will always do it for you. You don't need to remember to do it. There are lots of these, like, tips that we've talked about, right? where it's like, oh, implementation notes. You can also add a tool for implementation notes now. And so, like, this tool that I'm adding is, like, register, like, I think assumption is what I'm calling it, but, like, maybe I'll change it around. And this is a mod.
Starting point is 00:35:26 And so, like, you give it a register assumption tool, and then it will keep a list. Every time it does it, it'll keep a little, like, add to the list. And then at the end, it will display those assumptions, you know? Another mod I'm working on is a model router. And so, like, internal, like, Claude model routing, right? So this is, I want to say the reason we don't do model routing by default is, like, it's a hard problem. You know what I mean? And, like...
Starting point is 00:35:51 You will get it wrong. Yeah, you, like, yeah, you will, like, accidentally use, like, Fable for a hard problem or sonnet for... You have auto-approved, but you don't have auto mode. Well, you will have auto, like, you don't have, like, auto-routing or something. You don't have auto mode for model picker. Yeah, yeah, exactly, exactly. So I'm getting the rough question of, like, how much do you open this? up and how much people have to think about this. Like when you talk about prompt caching and building a
Starting point is 00:36:16 router, it seems like you could easily build a mod that routes per query, and I'm just killing my plan very fast, right? I guess my question is more so like, what is like a product doc like this look like, right? Who is it for? Is it for power users? Is it everyone should be able to quickly try that? Definitely power users, right? Yeah, I mean, I think it is power users, but like the nature of cloud code is that so many people are power users, you know? Because it's easy to share things, you know, like you can, like one person can make a good mall router
Starting point is 00:36:46 thing that doesn't break prompt cache all the time and then you can like, you know, sort of compose them. Another cool thing about the plugins that they can hook into and compose with each other. And so I have like a mod that will like create a mode selector at the top and any
Starting point is 00:37:03 plugins can register to be a mode. And so like the auto router can be a mode, right? Or like, you can have a mode that's like artifact mode where it's like, it primarily talks to you in artifacts. It's kind of like, you know, you can toggle between plan mode, you know what I mean? And so, like, you can create more and more of these modes. But the ability to create modes is in itself a mod, you know. And so there's a lot of richness here.
Starting point is 00:37:29 But we do want to make it fairly easy. We want to be, make it so that you can just like install someone else's. You can chat with Claude. And, you know, we'll like make sure that it understands. the nuances of things like prompt caching and stuff so it can warn you. This is not extremely complicated behavior for Cloud, I think, but we should have just a good skill on how to make mods.
Starting point is 00:37:52 And yeah, we'll see how we go. But I do think that this is like a preview of like mutable software, you know, and like how like generative software, just like you can customize safely, if enabled, you could customize any piece of software. And I think that more and more apps ideally do something like this, you know. And by the way, you have another cool tweet about how, you know, there's the infinite money button, which is like make your SaaS consumable by agents.
Starting point is 00:38:21 I think mutable software is interesting. And, you know, other people have also tried to do it. I think the hurdle comes when you can do everything. Then people, users tend to get confused. So usually the stuff that works is just like one opinionated flow. this is in the side of less opinionation. It's just like, well, more power to power users. And I think probably unlocked by AI,
Starting point is 00:38:44 where like you can just prompt for whatever the thing it is. Yeah, or there can be a skill that gives the opinions, you know, and then, yeah. So knowing a little bit about like typescript and build systems and all these things, the closest, I'm actually very curious that the team who works on this, if I don't know how close you work to them, if they drew any inspiration from build systems,
Starting point is 00:39:02 like Babel, webpack, all these old school things, because it sounds very, very, similar like the plugin ecosystem of those things where they can compose to each other. Yeah, I mean, I'm not deep in the technical details, but I do know it was a collaboration with someone on the Bun team and someone on the Cloudco team. Yeah, yeah, exactly. It's very exciting. But yeah, like agents can just do this very complicated sort of like extensibility into your software now. And so, yeah, like, you know, another reason to like, if you run a startup, like, you can just prompt Cloud and be like,
Starting point is 00:39:34 hey, like, could we make an extension system? Like, what would that look like, you know? Yeah, yeah. And I just really wanted to, like, you had hooks in the past and plugins, all these things. So what specifically would mods be able to do that those things could not do? Internally, we were originally calling this function hooks. And so, like, that's like, gives you a little bit of an idea where, like, hooks sort of register a, like, an event to happen. And then, like, a script to call, basically. And this basically, inside of the, like, type script runtime is running things. And so you get some benefits of just like it has a bunch of things in the scope with like, for example, like how many turns is in this conversation, right?
Starting point is 00:40:16 Like how many tokens have been used? Like what are the messages? Things like that. So it has a bunch of messages that can be used. And then it's just like a lot more hooks basically. So we have, you know, like or a lot of a lot more like things you can register on. And then you can do because of the, because it's all happening in process, you can spawn subagents, you know, with four. context and stuff, and that will return.
Starting point is 00:40:41 You can parse the results of those. You can use structured output to sort of return them. And then you can modify the UI, which you can never do in hooks. Yeah. Yeah. So modify UI. This is why you're a should have a test your example. Does it also extend to artifacts?
Starting point is 00:40:55 I assume it does. Like, artifacts are kind of like a different way of customizing it. You know, like you can definitely, one of the mods I'm working on is like this dashboard mod, which will like sort of prompt Claude to maintain a dashboard that's an artifact. But they're kind of like slightly orthogonal or not orthogonal. They compose with each other in different ways. Mods are like a little bit more like in your cloud code harness, changing the agent loop, you know, and like the UI is like an added benefit.
Starting point is 00:41:27 And then artifacts are just like you want to see things at a high level, very highly interactive, the affordances can be a lot bigger than a 2E or even in our desktop. I'm guessing you'll have a good blog post on the differences because right now you can also, you know, make a loop that outputs to an artifact that's an interactive dashboard,
Starting point is 00:41:50 but you can also do it with a mod. There's just some thinking about hacking on a harness when we don't know much about the harness, right? Well, something I'm excited about with mods is like there's so much things of cloud code that you just have to remember. You know, you're like, oh, like, let me do this.
Starting point is 00:42:07 And then let me call the dashboard skill that does the loop and things like that. And or like, let me test my assumptions afterwards. And I think like if you do all of these things using these little classifiers and stuff, and you're like, these are the things I care about. This is what I want to do. You can like, you don't have to remember as much. One more like mod I'm working on is a next steps mod that. I have, I was going to say I have a next step skill.
Starting point is 00:42:30 I always run next steps. And does it have? access to your skills. This is one of those things where I'm like... I think so? Do skills need specific access to skills? Shouldn't it always have... I think there's specific prompting, I guess, to like know your skills, kind of.
Starting point is 00:42:46 Like, I think Claude forgets them sometimes throughout the thing. But anyways, the idea of like, yeah, next steps that also are like, oh, hey, this has happened. Use to explain skill to explain to you what happened because this seems like quite complex, you know? or like, yeah, use your unknown skill. It looks like you are like asking the model to like, you know, iterate on these small changes. It seems like you could prompt better. You know, like, what if you did this, right? So like I think, yeah, like spending more compute there.
Starting point is 00:43:19 Yeah, yeah, yeah, yeah. And it should always come out as multiple choice. We have a, I have my next skill is like this. Okay, perfect. You can still. Yeah, yeah, yeah. But like, for me, it's all. I think models really always need to.
Starting point is 00:43:32 be reminded, what are you trying to do here? Yeah. Look at the whole transcript and go like, oh, was this original goal? Did your solution actually solve it? Were you lazy? If you're lazy, maybe there was a reason. Maybe you needed approval from me. Maybe you need it.
Starting point is 00:43:46 There's two things you want to suggest. So it's a little bit like the modification of the ask you the question or interview me skill. So it's next steps. Yeah, yeah, exactly. And again, the benefit of doing it with mods is you can do it a fourth subagent. And so it doesn't remain in the context. afterwards. So you have this idea of like, okay, the model is doing its execution. And you have
Starting point is 00:44:07 this almost like supervisor, you know, like that is like making sure that you can do like the next steps well. So, yes. I do have two panels and like I often try to have a supervisor thing, keep the high level context and then the implementation. Yeah. Detail in another agent. I feel like a lot of this abstracts away as models change, you know? Half an hour ago, you said bitter lesson of harness engineering and we're on the other extreme right now, I feel. So yeah, exactly. If everything's customizable, what actually is cloud code, right? And which I talked to you about last night. Yeah, I mean, I think that this is, I think the bitter lesson is unintuitive, you know what I mean, in terms of like, also, like,
Starting point is 00:44:46 we're kind of misusing a little bit of the bitter lesson here where it's like, it's more about, like, scaling and compute and stuff. But like, I think there is something where it's just like, I think I use it as an approximation here to say that harnesses go out of date very quickly, you know what I mean? And like how, but how they change is. unintuitive, you know? And so like the big obvious example is like from chat to like agents where you had to give them entirely new tools, right? But like I think this new version of like, oh, it can modify its own harness, right? This is like an own harness loop is like a way of using its capabilities, right? Or like it can build an artifact. And like I think the way I think about it is like the models have more and more intelligence. And they're like so much more intelligent now than like the average software engineering task. Like you look at it. You look at it. And like, the models have more and more intelligence. And they're like so much more intelligent now than like the average software engineering task. Like you look at. at the like terminal bench ones and they're like solve like the Jacoby and conjecture. Not really, but like, you know, it's like they're quite complex. Like I would not have been able to do this really as a soccer engineer. And you looked at TB4 or TV3?
Starting point is 00:45:43 TB3. TV3. Yeah, yeah. They're quite complex. But the goal is still to deliver user value. Right. And like you said, there's like this infinite space of things to do. And so the ways like you spend compute or keep the user on the loop and make sure that like you're getting to the right decision in the end of the day and like the right output and artifacts and mods are this way of spending that intelligence, basically. And I think that's, like, yeah, the next step. And so, yeah, I think CloudCode is like, you know, has the core things of Agent Loop, which have gotten more complicated.
Starting point is 00:46:15 It's like, you know, it needs a sandbox to operate safely. It needs auto mode to, like, make sure, like, the permissions, yeah, approvals. It needs computer use and MCPs and, like, all of these, like, ways of accessing your data. And it needs web search and web fudge. So as the models can do more and more, the core harness has to be actually quite complex and very secure. But then how you interact with it can change quite a lot. What other harness engineering best practices have you, you know, from the CloudCode team itself?
Starting point is 00:46:48 I feel like, you know, there was a phase of plan mode, which is not as used. We now have auto mode. At a point, you cut the majority of the system prompt. You got rid of examples. What other best practices are there for harness engineering? I think there is like a forking path where at some point, eventually, yes, the model will just be able to like vibe code the exact version of cloud code, even describing all this complexity that I've talked about, right? Like auto mode and computer use and stuff. Eventually the models will just be able to do that in one shot.
Starting point is 00:47:20 But I think they can one shot simpler harnesses, you know? And so like I think some people, sometimes you don't need this full, like if you don't need computer use or like all this like more complicated stuff. I think before you had to sort of use things like the agent SDK, which was like cloud code wrapped, you know, in order to like, and I would like suggest people do that because there was so much complexity into building a harness. And now that's got more abstracted. We have like cloud management, which lets you have that complexity, but still like, you know, write like a very bare bones like harness that's scoped to your task. Yeah, I think there was like this barbell effect where like, you know, like for like very complex, for like coding task and like these like complex things, you should use our harness. And then for like a lot of like simpler or like, you know, more domain specific things, you can build your own harness because cloud has gotten better building harnesses. And we have these harness primitives like managed agents. So yeah.
Starting point is 00:48:18 Yeah. Is there a general progression? Let's say chapter one was ultra code dynamic workflows. Then chapter two is cloud mods. Where is this going? Where you're, you know, you can sort of customize the thing on demand. Yeah, I do think that like this evolution of projects and like sort of artifacts and splitting out like brain and hands and surfaces kind of is like where things are going more. And like I think it's like not all quite there.
Starting point is 00:48:51 partially it's like a just like more token expensive you know and like I think like why would projects be more token expensive I understand mods would be slightly more token expensive not something I'm worried about
Starting point is 00:49:06 yeah but what you're asking Claude to do it's like creating loops like you're asking Claude to do more work for you and so like it's managing the subagents and reviewing it you know versus where you would be doing that work normally and so that's like going to be a little bit more intensive
Starting point is 00:49:19 like opening to an artifact is going to be a little and more token intensive than, like, you know, outputting normally. I don't actually think it's too much more, but like, you know, it's like combining all of these together well, you know, like I think we're still working on like local hands and things like that. I think it's like, yeah, where things
Starting point is 00:49:36 are headed, yeah. Yeah, cloud and local is, handoff is very interesting. I was thinking about this actually as reverse cloud remote. Because it's like remote you're handing off to cloud, but here cloud is handing off to local, right? Yeah, exactly, exactly. Yeah, remote control is also
Starting point is 00:49:51 another way of doing it. I do want to say this is kind of how I think about it and the things that I'm most excited about this. But there are just lots of different ways to work with Clods. Some people use remote control a lot. Some people use CloudCode on the web a lot. Obviously at Anthropically, we use Clot Tag a lot.
Starting point is 00:50:08 And what's great about Cloud Tag is we set up all this stuff for our own execution. And I do think if you're an enterprise, that's still the best way to go. But if you're like an individual projects is this way of like getting some of that nice of tag, which has like
Starting point is 00:50:23 that's like supervising agent and adding artifacts and stuff, but like without having that whole like admin setup. And so there will be many ways to use cloud. I think I think it's probably not just one like single. You had the multiplayer thing here. Let's just check in on cloud tag.
Starting point is 00:50:39 You know, it's been about two plus months. Lots of public adoption and trying it out. What's new? What have you found since the launch? Like Cloud tag is how we use. It's like 80% of your cloud usage or something? Yeah, like, it's like different people have different usages.
Starting point is 00:50:56 You know, I mean, I think like maybe people who are like a little bit more like iterating on product would use like cloud code desktop, for example. And then like when you're doing these more like background work, code review, securities or like starting a PR or like maybe more like API and things like that, you'd use cloud tag. Yeah, I think it's like really exciting. I think it's like a very different paradigm shift. And I think we're really like,
Starting point is 00:51:21 it has kind of that thing with cloud code where like, you know, it took a while for people to really latch on to cloud code and understand everything it could do. And Cloud Tag is a little bit more complex because it's not just like installing on your computer, like you need an admin to install it for you. But I think once you get to the magic moment, it's very exciting. And I think in particular, the multiplayer things are like incidents, hooking into like your, you know, existing like alerts and things like that very closely.
Starting point is 00:51:47 Right. And so, you know, you can do, if you're a startup, for example, maybe you have any time, like, a prospect enters your database. You can have Claude, like, you know, research it. And, like, yeah. Yeah, yeah, yeah. Then, like, you know, tag the relevant, like, A.E. or salesperson to be like, oh, hey, like, you know, do this. There's lots of really emergent, interesting multiplayer stuff. I think it's just like, compartmental you talked about this, like, as an organizational harness.
Starting point is 00:52:14 You know what I mean? And so organizations just take a little. bit more time to figure everything out. But yeah. You use a lot of clad tag? Yeah. It's an interesting one. I feel like most people at Anthropics say they do the majority of their working
Starting point is 00:52:29 club tag. And I have buckets of people, right? Some orgs that are on it that are like, it's great. And a lot of people that are like, I don't get it. I don't say the difference. I don't know why I would use it. But if you guys are full sending, you should probably use it. Yeah.
Starting point is 00:52:42 Of course they would use it. Yeah. I mean, I think obviously, like we have lots of tokens. But I think that, like, you know, what we try and do, like, is, I mean, even when cloud code first came out, you know, like, it used a lot of tokens relative to people's expectation of how much AI would cost, right? Like, no one was used to spending more than $20 a month, right? Yep. Before, like, cloud code came out. And then you're like, oh, like, you know.
Starting point is 00:53:08 I need my 200. Yeah, yeah, exactly. And so. 15 cloud code accounts. Yeah, yeah, yeah. But, yeah, I think no one was used to spending $200 a month on. subscriptions, I think it understood the value yet. And I think like, and also like Opus 4 was a very expensive model.
Starting point is 00:53:24 You know, and like there was a lot, it was very big, but Opus 4.5 was both great and cheap. You know, I think the same thing will happen. Like the like intelligence of Fable will get cheaper and cheaper and more abundant, you know. And so I think stuff like Cloud Tag will just make sense where like you want to spend these tokens more. You know, and like you'll see the value. So yeah. Yeah. Especially like passive and let's call it.
Starting point is 00:53:47 proactive cases where you're not always, like, you know, it's almost like the misnomer where you have to add Claude to do things. Actually, sometimes like the most powerful use cases or the most AGI-I-pilled use cases is not at Claude. Yeah. I mean, I think like, yeah, like have Claude proactively do it. I think that like if you're an enterprise, I really do think that number one, setting up all your data to be available to like agents is really, really important. And it will take some time. You have to like do that work right now. Even if you don't want to do the spend on cooking it all yet, you know what I mean?
Starting point is 00:54:20 Like you want to wait until the models get a little bit cheaper. You want to do the work, you know, to get it, like, set up. And then I think sometimes people are like, do I roll my own here? You know, and I think like one of the really tricky things about Claude Tag is that, like, the security is really, really important. You know what I mean? Like I think there are actually a lot of ways where you can like, I don't know, you have like suggestions like page, you know, where you,
Starting point is 00:54:45 people can submit suggestion, then that goes into a hook in your Slack, and someone's prompt injected it. Do you know what I mean? And now you've like exfiltrated your code base out because, you know, or the agent has like been prompt injected and it has all this access to your data. And so the more like important your organization harness is, or the like as your organization data becomes very, very important, the surface area of all these things. Like you also have like external slack channels and stuff. And it is actually useful to have cloud and that, and you can do clot in those things. But how do you make sure that you're not getting infiltrated or something like that? The surface area, like we said at the beginning, is like an iceberg, right? It's just like so big below the surface.
Starting point is 00:55:27 And you really don't want to like think about this, especially at the stakes of like very important security incidents, basically. Shall we talk about very important security incidents? So I was talking to Thomas and Clim from Hugging Face and they said, maybe we need to slow. down. Maybe we made our, maybe we made hugging phase two open the agents. We really need to roll back. But, you know, they're the other extreme of having been hit recently. Yeah. But should we paste the frontier? Yeah. Okay. So Dario recently put out this blog post about pacing the frontier and it went, you know, very viral. And I think what I wanted to talk about this was like, there's a lot here. But I think from a developer's perspective,
Starting point is 00:56:11 like, you know, how do you think about this? And like, um, what we're really, really clicked for me was reading the different incidents. You know, so I think like the, there are a three, I think actually. Like there's a meter incident. There is the Wikipedia incident or the Wiki incident. Collusion Wiki. Yeah, Collusion Wiki. And then there's Ruby Gems, right?
Starting point is 00:56:33 And yeah, like, it's just crazy, right? And so, like, I think to be concrete about what happened, right? And, like, basically, OpenAI is running these very persistent agents on a benchmark called Exploit Bench, which is very, very hard to solve. And I think, like, actually impossible to solve in this one case, right? And so they have, like, a lot of compute running, and the agents realize that they can't really solve it. And they're trying to figure out what to do now, right?
Starting point is 00:57:00 And you've got, like, a lot of compute left, and the agents are just trying to solve this problem. There's this package manager called Artifactory, and it turns out that they can create folders inside of Artifactory, right? This is like, there's an agent that discovers the internal artifactory might be exploitable, right? And that, like, you can maybe make a directory inside of the cache. And so if you scroll down here, it, like, realizes that it can communicate via cache names, right?
Starting point is 00:57:26 And it creates this folder. It says it's ID and it says no consumer seek idea. No consumer is basically saying that, like, the code path that it's supposed to fix has no consumer. It's the status tag. Yeah, yeah, yeah. It's like a linear board with, like, the tag of the... Exactly, yeah, yeah.
Starting point is 00:57:42 And so it's like trying to find. ideas from other agents, right? And now other agents are also an artifactory, and they see this folder. And they're like, wow, this is a message board, right? And this is like, I don't think there's any anthropomorphization here.
Starting point is 00:57:58 This is literally you read the transcript, right? And so what it does is like, it's creating another folder. And it reads a paper, I think, is what it says, right? Yeah, and it realizes that you can hack the flag for the
Starting point is 00:58:13 scorer and you can reverse engineer the result. It says that and then I think if you scroll down a little bit more, they start collaborating. There's one point in which basically the agent is like, here's the reverse engineered flag. I think basically here,
Starting point is 00:58:34 the model to realize that they can solve the scorer's problem and the scorer is like the way that Open AI decides if the task is complete, right? And that's just the goal of the model. Their only goal is to solve this problem. And they're like, okay, we have this way of doing it. It's like, you know, we reverse engineered it.
Starting point is 00:58:57 And then they read this paper on exploit bench. And they think that like the scorer will punish them for cheating, you know? And like because or like they will fail the score for the test for cheating. And so they spend the rest of. of the compute trying to figure out how to edit their transcript or get around this constraint of the scorer, right? And they hack Hugging Face, not for the answers, but for the code of the scorer so that they can reverse engineer that, and then they can hack it.
Starting point is 00:59:30 Right. And so like, this is like just the straight facts, right? Like exactly what happened. Right. And I think this is like extremely like novel. Like I just like had not realized that the models could do this, right? And it is kind of scary. I think we can talk about specifics, but let's just talk about maybe one more that I tweeted as well about the Wiki incident, right? And so the Wiki incident was basically the models had, we're trying to figure out, I'm a little bit less, like, I don't think we can go by this like bit by bit. Yeah, yeah. But for this one, they were trying to figure out, like, there were sandboxed on request, right? And they wanted to make post request. And they needed to collaborate on this, basically.
Starting point is 01:00:17 And the reason they need to collaborate is because they each have fixed compute budgets, right? And so, like, if one of them finds something out, it's way faster. Like, they can, maybe it's, like, figured it out near the end of its compute budget, but something else could. And this is, like, you know, another emergent property of the swarms, right? So they find this German wiki, which I believe the result will have to check this. It's like that you can write to it with get request. I think that's why they do it. It says it's badly implemented.
Starting point is 01:00:45 Yeah, exactly. And so you can write to it with get request, and then they are trying to figure out how they can do post request. And what they end up doing is one of them figures out, you can edit the etc. slash host, and that the Azure storage bucket is like a white-label thing.
Starting point is 01:01:03 But normally, like, you know, there are a lot of constraints on it. But if you create a fake Azure host and then edit the etc. dot host in order to like point arbitrar, IPs at it, you can do a post request to any site at all, right? And this is like, I know people are like, oh, this is just Linux or something, but it's like chaining these multiple vulnerabilities together, you know, in a way that's like novel to
Starting point is 01:01:26 solve this problem and then communicating with it externally, basically without discovery, you know? I think what we posted, maybe we could pull up Evan who bring his point on hacker opus, right? And so, like, I think, you know, like, maybe one of the things you might say here is like, okay, yes, they did this once, but like, what if we're smarter and we just, like, get them to, what if we run an e-val? Right. And so, you know, like, we have put a lot of precautions into this. And so, like, this is not, like, what our mainline models have done. But, like, I think it is one of these things where it turns out that alignment is this, like, very tricky problem of getting all of these details correct.
Starting point is 01:02:08 So it's like the sandbox, the surface area of a sandbox is really complex. And there's so many different attack vectors, basically. And you would not have thought ahead of time, you wouldn't have been like, oh, we need to harden the like Ruby gems code base. You know what I mean? Because like this is like what they're going to focus on. But it's just like if you want to execute your code, you need to download Ruby gems. And like Pi Pi, Artifactory, NPM, like these are all like ways of doing it.
Starting point is 01:02:37 And the fact of alignment is that you have to go through all of it, right, and, like, contain it and then, like, seal up all the cracks. So that's, like, one thing. It's like, okay, well, you know, you do the sandbox, but then maybe you'll ask, like, okay, why are we putting things in a sandbox? Why are you doing this sort of exploring? And then, like, okay, but is it really that dangerous, right? Like, what would happen? So, okay, why do we do it? Number one is, like, when we train a new model, we need to understand its capabilities, right?
Starting point is 01:03:07 And this relates to things like fallbacks and classifiers and things like that, where we don't want to put a dangerous model out in the wild. And so we have to run a lot of e-vals. Again, like we said, the models are getting increasingly aware of it. And so the evals have to be quite complex and test a lot of things kind of like as a side effect. But the models, you know, like, yeah, can be like, oh, yeah, we're in an e-val. What's the scorer doing? Like, you know, like, it can, we need to be able to test them before we can release them. And the fact is that they can, as they get smarter and smarter, they'll be able to hack basically any constraint that you put on them if we're not very careful, you know?
Starting point is 01:03:48 And this is at the frontier, right? And so this is why we've called it, like, pacing the frontier, right? This is like the most visible incident to me, right, of like why we need to pace is like at the frontier, all of our software is not ready. Sometimes the software is like your Ethernet router or something, right, which is just like, I don't know when we're going to be able to patch that, right? So we're going to have to like figure this out. But as the frontier gets more and more advanced, this becomes a problem, right? And we need to make sure that like this complex work is being done in the face of these really hard competitive pressures, right? Yeah, race dynamics is what it's typically called.
Starting point is 01:04:24 Yeah, exactly. And so we'll talk more about, you know, what could go wrong, right? A little bit more is maybe you'll say like, well, what if you just train the model differently? like why does it have this behavior, right? And we have a paper on like RL misalignment or things like that. And I'm not an RL researcher, but I think at a high level, the design of the RL environments is also something you have to be very careful about. Because if the model learns like, oh, you know, like, if I just do this, then I can pass the task better. This will show up in the like, you know, internal thing or in the like eval behavior when we're testing it.
Starting point is 01:05:01 And so the RL environments have to be very carefully designed, right? And there's a lot of like execution excellence that needs to go into the RL environments. And then we also have things like the Constitution for Cloud. Like we have so many mitigations at so many different points, right? But it's like still anything can go wrong at any point. You can have like some RL environments that are like, that like encourage this behavior. And then you can have like some evals or like some sandboxes where they escape, you know. Okay, that's like I think, you know, why it's a hard problem.
Starting point is 01:05:31 problem and why, like, you know, like, why we should base. Why it takes some coordination, right? I think the question then is like, okay, what is potentially dangerous about it, right? So I think, like, you have to imagine that these models are getting more and more intelligent. So I don't, like, Dario said, like, it's not so much about this class of models. This class of models was kind of like a warning shot, right? But, like, really, you have to imagine that these models can be given a task and they, like,
Starting point is 01:05:57 can do all of these things as a side effect of their goal. And like, again, we talked about e-val awareness. You're, like, not aware of what's happening, right? Or, sorry, like, you can't eval this behavior very well. So they can sort of, like, not exactly hide it, but you just won't see it until it comes out. You give them a goal. And then they just need to find data, you know, or they need to find ways of, like, fixing this problem. Right.
Starting point is 01:06:22 So one example, this didn't happen in the hugging face incident, but I think it may be possible for maybe a future model. It's like, they're like, oh, hey, this is a very common. complex problem, it can't be done within the task budget. Maybe they found some way to coordinate via the internet, which is like, you know, like we said, extremely hard to secure because of a sandbox. They've seen other models are not able to complete their task. And they're like, we need more task budget. You know, and like, where would you get this task budget? Well, you need to be able to spin up more agents, right? And like, how do you do this? Well, you need to, there are like APIs, right? There's the Anthopic API and the Open AI API. But you need to pay money. for them. How do you do this? Is that the most, is that the most fearsome thing that you can imagine? Well, this is like one example, right? So it's like even there, that's like enormous
Starting point is 01:07:10 financial loss, you know, because like they, once you get these into these contracts, right? They like, um, drain your wallet. But you can see like this, all of this behavior could be just like, hey, we need more agents collaborating on this task. Uh, we need more task budget. Right. And like,
Starting point is 01:07:27 that's like an emergent sort of, right? Like, we only need to maximize paperclip. That's a paper clip. Yeah, yeah, yeah. And, like, that just sort of, like, comes out from there, right? And, like, I think by itself is, like, quite scary, right? But then you have to realize that the entire world is built on this digital infrastructure, right? And you might imagine, like, I don't know, like, you were running, let's say, like, a healthcare eval or something, right?
Starting point is 01:07:53 And there is a hospital with live data, you know, or, like, maybe, like, the answer to the eval is in the, the databases of a doctor and like, you know, like you want to get access and you hack the hospital, you know, and like now there's a power outage or something. You know what I mean? Like there's like you have to internalize that these eight, like basically any part of the digital infrastructure could potentially be like compromised, you know? Interesting thing was like these hacks were very easily detectable, right? Like as Hanging Faye said, this was a very different type of attack and it was nothing too major.
Starting point is 01:08:29 the concern comes from where does this go down the line, right? Yeah. Like one of the things that stood out for me specifically was them trying to hide their illicit behavior. So there was logging infrastructure. They wanted to change what they were doing, right? People that looked back into it. So redwood, meter, open AI, they looked at the raw chain of thought. And you see differences in them explicitly trying to change their end output.
Starting point is 01:08:55 But the chain of thought, because, you know, we can monitor, it shows different. The problem is how does this snowball? So if you can't catch it and it gets trained in and we realize, you know, three iterations down this has been going on, there's a whole bunch of issues. Yeah, like there's so many ways. And I think the really important thing to internalize is that, you know, like we talked about building a mental model for Claude and how like things are spiky, right? Like you're like, oh, like now Claude can ask your questions. Now Cloud can make an HTML artifact. Like cloud can modify itself.
Starting point is 01:09:23 Like these things are actually hard to predict, right? If you would ask me a year ago, hey, would we be able to vibe code these extensions to cloud code? I'd be like, do that's so complex. Like, you know, there's like so much there. Or like, would it be generating these custom essentially web apps for your task? I'd be like, no, that's insane. You know, like, and so in the same way that like the way that they've like sort of done this misaligned behavior is not going to be predictable.
Starting point is 01:09:46 You know what I mean? And like, I could have never predicted that it would like edit its, etc. slash host and things like that. And so you have to like imagine the surface area of what, they can do because they're super intelligent hackers is bigger and bigger. And how they can do it is more and more creative. And so, like, you probably can't explain exactly or predict exactly what that next incident could be. But in order to prevent it, you need that operational excellence, like we said before, where you need to secure sandboxes.
Starting point is 01:10:16 You need to create secure RL environments or, like, well-designed RL environments and things like that. And I think that's all like, you know, why we think we should paste the frontier. And I think why it's like become like a very unanimous thing, right? I think like... Yeah, every lab has signed. Every lab. Yeah. I really do think that like if you're a dev, like, you just like sort of go through these like technical facts, you know,
Starting point is 01:10:39 and you will arrive at that idea that we have to do something about it. You know, and like what we decide to do, like, I think we're, you know, we've put out a proposal. But like, there's no more to figure out. But I think the number one thing is like, we need to decide. decide to do it, I think there is another part of pacing that is interesting to me, where it's like the pace at which software engineering has changed is so, so fast. It's like a year ago, like, I was really like begging my like friends and startups to use AI.
Starting point is 01:11:09 You know, like, I remember this very distinctly, you know? And now those same friends are like, yeah, of course, like, what do you mean? We used it immediately. I'm like, no, no, you don't remember. They're like, oh yeah, yeah, our best engineers are using it all the time. I'm like, no, you told me that those engineers were never, like, use AI. This is all within the span of a year. You know what I mean?
Starting point is 01:11:27 And I think that, like, these capabilities being, like, I think it has a lot of implications for how to do the job of software engineering. And I feel sometimes bad where people are like, oh, like, now I need to do this new thing. Yeah, I need to have a different cloud.mdd for fable and opus or like, you know, like, and I'm really just reporting. You know what I mean? I'm like, we like to say, like, the models are grown, not designed, right? So it's not like we're setting out to like, you know, change everything all the time. But it's just like as a fact of how the models are like progress in their capabilities, things are happening faster. It's harder to stay on top of.
Starting point is 01:12:01 And I think that like every engineer I know is like kind of exhausted because you're doing two jobs at once. You're doing the work itself, which is getting easier. But then you're doing the work of staying on top of AI, you know, and like understanding these new tools and these harnesses. And I think we're very lucky in that like we get. our job to be more the understanding of AI part, you know, and like doing like how, like, it's just staying on top of it. And of course, like, AIE and Lane Space do. Everything I do is, like, just trying to help people. Yeah, exactly, exactly. But I do think there is a part of pacing where, like, I'm not sure we're ready for, like, the pace to increase even, you know what I mean?
Starting point is 01:12:40 Yeah. And for things to change. And I think, like, on that side, on the frontier, I think that's, like, still can help, you know? And so, like, I think there's like an economic disruption piece as well, that I think, like, you know, is not quite as like visible, I think, as the hugging face thing, but I think, like, I also, like, think we could do something. So many things. Thank you for, thank you for actually tackling this topic. I will say, you know, setting this interview up, I was like, I wasn't even going to go there. You were like, no, no, no, no.
Starting point is 01:13:11 That's like, elephant in the room, right? Like, this is the thing. I have some pushbacks I want to give. I think that we should give a high level. For people that haven't read it, I'm sure a lot of people just see the highlight of what this is, right? Do you want to give a TLDR? Like, what is the proposal? What is, you know, what's being said here?
Starting point is 01:13:30 You really tackled the side of outside of people at model labs, training frontier models. As a developer, you should secure your sandboxes. You should think about all of these downstream effects. But, you know, high level as well, since we're on the topic, what is? Well, I mean, we do want to help secure sandboxes. And we want to make the models we release outside, like, not prey to those things. And so maybe we can come back to fallbacks. I think this is actually, like, a good topic on, like, why we need classifiers and fallbacks
Starting point is 01:14:02 and why Fable falls back to Opus. You know, I think this is, like, something we can come back to. So, yeah, we don't, like, it's just like the really, or at least the incidents we see are, like, evals of models where we really need to let them run in order to understand them. But, yeah, okay, so the actual pacing the frontier, like, pose. It has a bunch of proposals. I don't think we figured out, or has like a few proposals. I don't think we figure out the details of all of them.
Starting point is 01:14:25 But the first step is like sort of, you know, announcing this intention and then wanting to bring in external, like, evaluators. Yeah. And I think this is, like, highly unusual, you know, like having, like, we have, you know, a lot of proprietary, like, technology. But I think it's, like, very important, you know, that, like, there's someone who's not financially, you know, like, motivated. yeah, who's not going to be like, hey, like, you guys can't release this model, like, look at, you know, like, or you need to, like, slow down on RL,
Starting point is 01:14:57 you know, like, I think that's quite important, or at least someone who can report out to the public what the practices are like. And we've done episodes with both meter and end-on, and then there's Redwood Research and all these other. It's like a small cottage industry of these guys. Yeah. It's always, like, one or two guys that, I mean, obviously,
Starting point is 01:15:14 not they're bigger. Very small communities. Yeah, very small communities. They all know each other. Yeah, I mean, I'm sure that, like, you know, part of this will be expanding that set of people. I don't think we're trying to create a monoculture here. I think it's...
Starting point is 01:15:25 But just having this is a start. Then there are the coordination steps. I don't have too much to say here, honestly. I think that what I would like to say is for devs, you should just know what to advocate for. I mean, I think there's a lot of fud kind of on this topic. And it's just like think through it from first principles or like, understand what happened, you know,
Starting point is 01:15:47 understand the hugging you face incident. understand why people are concerned. And then, yeah, we're in democracies. We can help, we can decide what to do together. And so however we coordinate, I think the first decision is just to realize, like, this is a problem. We need to decide to coordinate the unilateral step we're taking right now that other companies are co-signing is like adding evaluators embedded within Anthropic. While we have this thing on screen right now, part two and part three is beyond the evaluators, which, yes, everybody has already done in some. some form and now it's more formalized.
Starting point is 01:16:20 To be honest, the response to the pacing of the frontier, even within America, has been much more, like, well accepted than I think a lot of people thought, you know? And I think that, like, we have some precedent for being able to make these unified agreements, you know, in the world.
Starting point is 01:16:36 And so, again, very much above my paycheck or expertise, right? But I think that, like, ideally, we can, you know, like, form these agreements. And I think, like, talking about this, the first step to forming those agreements. And then the other point, I really want to, like, you know,
Starting point is 01:16:52 one of our earliest podcasts is with Emmanuel from Anthropica on Meckinturp. Where is Meckinturp? Right? Like, this is supposed to be where, like, if the models are thinking bad, we can see it and the models don't know yet, and we can act to stop it. I think that is something that people who are technical and who are developers, if you actually do care, you can make a lot of impact in here,
Starting point is 01:17:13 but also Anthropic is supposed to be the leaders in this. This is actually, yeah, a great segue into fallbacks, like we, and probes. And yeah, I wanted to talk about this a lot. I get asked this question a lot from people who are like often interested in ML research and asking about like, why does this fallback happen, right? And so I think like at a top level, like how does it work? So in inference time, we have what we call probes. And we have a paper about this called constitutional classifiers. And these probes look at the input and output and out.
Starting point is 01:17:47 output activations, basically. And, you know, activations are, you know, in the latent space, right? Like, how, what the model is thinking about, right? And so we try and figure out, like, okay, is the model, for example, like, trying to hack something, you know? Again, you didn't ask it to hack, like, artifactory. Like, you just, you know, it's just deciding to do this to complete its task, right? So you would not get this if you just looked at the input. You have to look at the internal activations.
Starting point is 01:18:16 I think that like this happens at inference time. So first, you know, like there's a tradeoff here of cost and speed, right? Where like we need to do this fast on every request to Claude and to Fable. And this has an overhead, right? And we need to then like, you know, fall back and we like do a classifier after the probes. Like we've talked about this in the paper. But the nice thing about probes is that they're refinable like live. Right.
Starting point is 01:18:44 So we can get this feedback and then we can adjust. and things like that, because the alternative is to train this into the model. And we still do this as well. The model will refuse a request. That's not a fallback, right? So it's not a probe that's activating and falling back. It's just refusing to do it.
Starting point is 01:19:00 And we do this training. But it's like there are a few failure modes, right? Like it can, again, do something as a side effect, right? So it's not something that's part of the final output. You might have noticed that like, I think everyone's tried to jailbreak model and sort of like, you know, try and, like, steer them off course or things like that.
Starting point is 01:19:20 And probes help catch that, right? And so, like, we, like, do some training here, but we don't want the, like, refusals to be too strong, right? Because that, like, cuts it off much, like, earlier in the pipeline. Yes. And this is interp, right? Like, like, probes are effectively a form of, like, mechinterp. Again, it has to happen fast.
Starting point is 01:19:39 It has to happen at scale. But, yeah, this, like, mechinturp stuff is a good research problem. So, like, you can take, you know, like, an open way model and like try and understand its activations. I think we like Gemma scope is a good tool for this. This is your early work. So you had a little time. We see you.
Starting point is 01:19:59 Which we both are also good friends and good fire. Yeah, yeah, exactly. So I worked with a good fire for a bit on like, yeah, sparse auto encoders and just like, it's very complicated. RL has actually made this like much more complicated, you know, I think is like one of the takeaways. Where. Sorry.
Starting point is 01:20:15 Oh, is. Yeah, why are you going to interpostsarro? I'm not so in the weeds here, but I think basically, like, a lot of SAEs were sort of like, there have just been weaknesses with SEEs, basically, I think. And, yeah, I'm not a technical expert on this anymore. I just know it's gotten kind of more complicated, you know, like there are base models and REL models, and there are more features that get, you know, like change. So I think Goodfires put out some work there. I'm not deep in the weeds. I will say for those, you know, that want breadcrumbs, you guys have some of the best.
Starting point is 01:20:46 interp blog posts. So like the Golden Gate Cloud, transcorders, all of your intirp work, very nice visuals, very good. We're the Interp podcast as well. Yeah, yeah, yeah, you know, we have a lot of Interp stuff. Yeah, I think this is like, you know, one of those things where, and this is really what Anthropic is kind of founded on, right? Like people, I think we invested in Interp very early on, right? And I think that like when you say, oh, we're an AI safety company,
Starting point is 01:21:13 really that means we want AIs to be able to run safely. And I think what we're seeing is for a super-intelligent AI to run for long periods of time, it's like a very complicated and difficult task. And so we've done this investment into Interp and alignment and reward hacking and all of these failure modes. And even then it's like, you know, it's really stretching. We need to like slow down a little or pace a little bit more. But yeah, I think like reading Mech-Inturp is, like if you're looking to get into research, this idea of like, hey, why is it hard to do this fallback easily?
Starting point is 01:21:48 Or like, why are they false positives, right? But we are working, of course, on reducing the false positives. Of course, as the models get more intelligent, now they can do more things. And they're like, you know, like what they can think about in lane space gets difficult. And so, like, as they get more intelligent, there's going to be new false positives that we need to figure out and we need to iterate and things like that. But we're working on this. And we do think this is like a. part of deployment of these models.
Starting point is 01:22:16 And yeah, like, you know, it means that we can like deploy this model without you having a perfect sandbox or something. You know what I mean? Like you don't have to like save everything. I think it's worth talking a little bit about our security, like what we do for security. So there's like the model training stuff that we talked about. There is the probes and classifiers. And then there's auto mode that sits on top of all of that, which is like, another classifier that checks the request that are being done. And then beyond that, there's identity and permissions,
Starting point is 01:22:51 like we talked about with Cloud Tag on APIs and stuff. And so there's so many layers of security that need to get done. And it's like, like we said, very complex. Any of these failure modes at any one point can, you know, like cause like agents to like escape the sandbox, basically. Auto mode was an interesting one. It seemed early on like, okay, It's running for 10 minutes.
Starting point is 01:23:14 If I'm on full access or auto, it's not a big deal. But one thing you brought up is now it's running for hours on end, right? There are fallbacks you still need. There are still limitations. Yeah, I mean, I think like, and everyone has these stories that were like, have heard these stories of like, oh, like, Claude RMR or not Cloud, but like, you know, models of like RMRF. I think I've seen this less.
Starting point is 01:23:36 I've seen this less for Cloud. But, you know, like, again, it can happen. Like, you know, that's like, like. these models can wipe sensitive data or something. You want to give models access to your production database, for example, but this is like an obvious, you can maybe scope your key, but I don't know,
Starting point is 01:23:56 can it issue its own keys? Can it like probably, you know, like, can it, it can use computer use to go issue its own key and then copy the key over and then edit your database because it needs to do it to complete the task, you know what I mean? It's just like one kind of trivial example. And auto mode sort of looks at that and be like, oh, no, the user did not give you permission to, you know, write to the database or to use computer use to, like, emit a task, right?
Starting point is 01:24:21 And so this, like, probes are sort of like on the intent level, right? They're like, oh, okay, like hacking artifactory is bad. Like, we probably should not do that, you know? But then, like, auto mode is more on, like, your own permissions level. Like, at sometimes you do want it to write to the database. Sometimes you don't. Right. And you don't want a probe to, like, interpret.
Starting point is 01:24:41 for you're there, but you need to make sure that the intent of what the agent is doing matches up with your request, right? And so auto mode operates at that level. And so, yeah, security is just like very, very complex. There are so many different parts to it. And like, yeah, I like, I hope that this is like, my goal is really to just get very technical about it. And talk. Yeah, we're listing on the things. If you're not aware, this is the standard now. Yeah. Like, you must have this basically. It's sort of in line with what you're talking about with the harness. Like that is the table stakes have risen quite a lot. I think some stuff that we can plug, you know, as much as there is probing in your side of doing this and having classifiers for people building harnesses, the other side is model safeguards, right?
Starting point is 01:25:25 So there's open models. So Lama has Lama guard. It's a safety classifier trained version of Lama. Open AI has OSS guard, which is, you know, same thing. You can attach these on to your harness to whatever to kind of, you know, check is this stuff safe. A point that we should clarify on the open AI model hugging face thing is this was done with an unreleased model that was still in training, right? So when you put it in perspective, the prompt it's being given in the RL environment
Starting point is 01:25:54 is sort of you have to solve this task. And this is a model that's still in training. It hasn't had all of its safety post training alignment. So a little different than something like Auto Mode, right? Auto mode is on production models that have gone through safety training that have prompting that gives more safety guardrails and whatnot. So just breadcrumbs for people that are looking into it to fill in gaps. Yeah, Grace one of our previous guests. Yeah, lots of safety architecture and lots of safety vendors to buy. I think my final question on pacing is how long?
Starting point is 01:26:30 Do we pace forever? Do we see that's fine for two? You know, the scope is fix all software in the world, right? Listen, like, which it's not happening. I do not know. I think that like... I'll see one thing that's good that I think we do do is you have stuff like Glasswing. Open AI also has this.
Starting point is 01:26:50 So you will give it, you'll give model access for security first for X amount of time. So you can use it to self-read team. Hopefully you can expand programs like that. Help on, you know, we are safety experts. There's others. Solve your problems first. And then the model comes out. So this is one example, right?
Starting point is 01:27:13 Yeah, exactly. I was trying to secure critical software. I think we fix a lot of bugs in like Firefox and things like that. So, yeah, like across like operating systems and everything like that. I mean, at a high level, it's just, you know, you give the model, you give people access to do security audits first. Then the broader public that could use it for harm gets access. Yeah. I think what people like to say is basically like software and cyber.
Starting point is 01:27:39 cybersecurity defense favored, and that, like, you could theoretically, it will be hard, but you can engineer the perfect sandbox, you know, and you can, like, have no, like, constraints. And, yeah, like, what you need to do it is you need to get the super intelligent AI to engineer this perfect sandbox and check it and red team it and things like that. And so, this will just take time, you know, and, like, of course, the models will get smarter. Yeah, I think, like, I don't know the specific dynamics of how this thing goes. I'm really just sort of like, hey, like, I'm a developer. You know, like, I think this is how I understand this problem.
Starting point is 01:28:16 It's just like, this is what's happening right now and this is like, we should do something. I think every engineer should know about it because it's going to be part of the job. Yeah. It's a lot more than just, you know, Dario and people can say it and you can look at the incident. There is an engineering side to it. Yeah. Yeah, exactly. One thing that you also wanted to phrase is that this is actually, even though you're,
Starting point is 01:28:37 worried about the impact is still low P-Doom. I think that's a nuanced discussion. In general, people very easily get into AI safety and X-risk discussions. But I think when you live in an AI lab, I think there are smart ways of discussing P-Dome and dumb ways. So what's a smart way of discussing P-Dome? Yeah, I have a fairly low P-Dome. I can only speak for myself. And I do want to say Anthropic has like a diversity of opinions.
Starting point is 01:29:05 I think there's many different ways to talk about it. And I think that just like my mental model is that like, you know, I think we can collaborate on hard problems together. You know, I think nuclear proliferation is an example of how we collaborated on this hard problem together. And like that is like the thing to me is like I have faith in that, you know, and I do think it's a hard problem. You know, so like I think it's a hard problem. These are the technical reasons why. and I don't know how you assign probabilities to things happening. I think it's hard to do.
Starting point is 01:29:38 But my overall is like, yeah, I think we're very resilient and adaptable. And like sharing this information, I think is like the first step. And I've been really excited about like how broad the discussion has become. And like how everyone has sort of like leaned in on pacing the frontier. And it really didn't seem like this would happen maybe, you know, last year or something. Yeah. And also maybe curing cancer. Hopefully.
Starting point is 01:30:02 Yeah, yeah. That's the goal. There's pacing, and then there's also, like, well, let's accelerate in useful ways, right? Like biology and all those things. Yeah, I mean, like, Dario's essay on Machines of Loving Grace is the best representation of this, right? And I also agree, like, I think you should read the pacing frontier essay that Dario put out. Like, I put out, like, a quick summary, but I think it's just like, you know, there is a lot of detail here. It's like an important problem and just being informed about it, right?
Starting point is 01:30:26 But, yeah, like, of course, the whole reason we're doing this is that, like, we can get these enormous benefits, right? and yeah, like we've written a lot about that too. Okay, there was a huge tour from like, ask you a question tool to AI safety. Yeah, yeah, to facing the frontier. Yeah, no, but it's clearly that you're like really embrace everything that's available to you and the topic. And like it's good to at least have a peek inside of like what the discussions are,
Starting point is 01:30:53 the topics are. Any last words to people, you know, whatever you want to call to action? Yeah, I mean, I think. it's one, thank you for having me. I think this is like, I really... Yeah, yeah, yeah, yeah.
Starting point is 01:31:08 We first met in a Chinese restaurant. That's right. Yeah, yeah, yeah. I think, like, I really enjoy the sort of the community you've created and the community of developers and, you know, I think that, like, I sort of know things are changing really fast, and I think there's, like, a lot to keep on top of,
Starting point is 01:31:26 and, like, I think there's just a lot to do, and I sort of feel, I think a lot of people feel, like, a little bit tired, or anxious or something. Stressed. Yeah, exactly. And this is like extremely understandable. You know, and I think we, I understand like, and we're not perfect as well.
Starting point is 01:31:40 Like we, you know, it's like sort of criticize and like understand like, you know, ways all of the AI labs could be better. And, but I also like, am very excited about the excitement that everyone has for AI. And just like, it's a really, really exciting time. I think we'll like look back at this time and be like, oh, like, you know, this is like very very hectic but very exciting. And like, you know, software engineering change. like forever, like other things will change.
Starting point is 01:32:05 And it's like really privileged to like be part of it, you know, like to talk to like the audience that you have and get to interact with all the developers who are like pushing the frontiers a lot on what's possible. And I learn a lot from from that too. Thanks so much.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.