Y Combinator Startup Podcast - Robot-Use Agents: Why General-Purpose Models May Win in Robotics

Episode Date: September 26, 2026

One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the ...era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training.In this episode of Decoded, we're joined by the founders of Waddle Labs and RoboCurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents. Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.

Transcript
Discussion (0)
Starting point is 00:00:00 One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now frontier researchers are showing that this includes controlling robots. This led MIT professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents, where general purpose models could make different robots more capable. So today, Francois and I invited the founders of Waddle Labs and Roboccurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs. Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your company is focused on. I'm Hamei. I'm from Wadol Labs together with Vincent.
Starting point is 00:00:43 We work on building LLM that control robots. And we do this by doing two things. Building a harness that allows the LLMs to do this very effectively, collecting data and using that data to train better LLMs. Hi, I'm Jay, a co-founder of Roboccurve. We are an Eval's company for physical AI. We measure everything, any robot, any model. including LLMs and also classical approaches such as VALAs, vision language action models and world action models.
Starting point is 00:01:11 And we evaluate all kinds of embodiments like hands, grippers, arms, humanoid, quadrupeds, all kinds of stuff. So in the last few weeks, videos from both of you guys went quite viral on Twitter and I think inspired. I think both of them are actually quoted in that Philip Isola essay talking about robot use agents. The videos showed things like LOM's being able to unscrew caps, being able to communicate between multiple robots, being able to do tasks like uncapping a pen, for example. And so I thought maybe what would be fun is for Francois and I to dig in with you guys on some of the research that led to this moment to begin. And then maybe we can also show some demonstrations of what this can actually do. To start, why don't we talk about some of the early research around transfer between language models and robot control? So I want to talk about some of the early work on pre-training language models for decision-making,
Starting point is 00:02:01 and then also the RT2 paper, which established the VLA. So maybe, Jay, do you want to tell us a little bit about some of this early work involving being able to use language models to do any kind of robot control? One of the earliest successful approaches of using AI on robots is the RT2 paper, where they use a pre-trained language model on web, text, and images, and use that to control robots. And the interesting thing about that is that it's a fine-tuned version of a language model, and instead of outputting English, for example, they just output what we call an end-effector pose,
Starting point is 00:02:41 which is the coordinates that you can then translate into joint commands that can control the robots. In some cases, in some sense, it's very similar to what we see now with large language models, just that instead of fine-tuning, the models are good enough to just do it out of the box. And this is actually quite similar to, like, the mapping that I have in my head is kind of like the COT moment. And so, like, when we did GSM-8K and you said, Susie had $8, she spent $5, how much does she have now,
Starting point is 00:03:11 it had to output like four hash marks and then the answer and then EOS. There was no ability for it to chain of thought. And so then we allowed it to actually have, like, okay, let me go, eight minus five is three, and so like, hash, hash, hash, three. And so you allowed it to do this kind of thinking before it actually gave an action or an answer. And similarly now, like VLA is where basically like an RT2 is basically like it has to output an action. There's no, like I can't allocate more compute for a more complex task. And then now I have this code chain of thought thing that I can do.
Starting point is 00:03:46 And I can say, okay, even if you were using the example where you're not actually outputing the code, just giving it to Astra and let it think, think, think, think, and then output in action. Similarly, now, like, and then with code, it's a more Kolmograv complexity or lower Kalmograv complexity, code length to just use code and really compact code to just like, here's the code. I totally agree with you on that. I feel like at the end of the day, it's a lot about a bit of less, right? If you give the agents or you give the AI model, more autonomy, and you if you unshackle it a bit more and give it more resources, it can actually do a lot of the things that we fine tune it to do. I'm just going to hop in here as well. I think the really
Starting point is 00:04:28 interesting thing about the beta lesson here is that like VLA's by design architecturally, they're built on top of language models as well, right? They can potentially reason they can write code. So perhaps the bit of lesson here isn't necessarily what architecture you build necessarily, but it's what data is most useful, right? Like people have been nagging at this VLA data bottleneck for years now, and we're seeing very, very slow progress. And so the lesson here is maybe we have a modality of data that we know works very well, and this idea of transferring across different modalities, which Hamming and I are super excited about, like how do you take something that's traditionally out of distribution,
Starting point is 00:05:05 labor robotics, and make it something that's in distribution, like maybe the way of harnessing the better lesson is saying, let's pick a data for which we know this modality is pretty blessed. there's a lot of data, these language, what these models understand it well, and then use this as a way of unlocking a lot of other domains as well. Yeah. And to add to that, the reason why RT2 is such a good model is because it is using a language model that taps into the modality of all the data that the language model is trained on. So all those web images and web text actually improves the VLA compared to just training a specific robotics foundation. model without any pre-training.
Starting point is 00:05:45 So you're benefiting in the case of RT2, these early VLA approaches, from pre-training. I guess what's the limitation in the, I guess, action-taking, fine-tuning approach that you can now get around if you can directly write code. Like, you're referring to like a difference in data there. What exactly is that difference in data you're referring to? There's a few things that models have got much better at since RT2. Like one, I mean, they're much better tool use, for example. So, you know, one kind of data is now these models can, well, I mean, they write code much better.
Starting point is 00:06:18 So now they can write complex policies as code. They're also much better using tools to explore the kinds of environments they're in, like what arms they have access to. And a lot of this comes from in context learning. Like the key difference between like these VLAs and like GPD6 or LLM isn't necessarily the architecture, but kind of the approach you take towards training. We want to be bitter lesson pill, right? We want to benefit from all kinds of data. We want to pour in computer-use data into our robot models.
Starting point is 00:06:45 You want to put encoding data into our robot models. But when you do that, you just end up getting what we think of as like these general-purpose LLM agents. And so to me, it's like, okay, why not just build a really good foundational LLM? And I use that to control robots rather than training like a model that's specifically for robotics and that's more dependent on just robot data. On that note, maybe one of the things that I think if we could rewind the clock a couple years may have been a good sign that we were trending in a good direction to this, is research
Starting point is 00:07:13 on code as policies, right? Do you guys want to talk a little bit about when the research community refers to code as policies, what exactly that means? And I think it's worth putting this in the context of when this paper came out, which was coding agents just starting to work. I think this was still in the era in which people were like putting comments and auto-filling Python code blocks, which, you know, now we think of as ancient history, but this was like two years ago.
Starting point is 00:07:35 So, yeah, how does code as policies work? And how did that inspire some of which you guys are now seeing? was possible. I'd probably rewind back to Voyager, where Voyager was like the one of the first, what does a coding agent mean? Coding agent requires good tool use and on the fly tool creation. That's called code. Okay, you have Python built ins. Let's just say I have only Python built ins. Those are tools. And I have sort and I have like if and I have four and I have while and I have all these things. And I need to like use those tools to assemble a new tool called a new francois dot pie. right and like now that's a new tool
Starting point is 00:08:10 and I get to use that that was Voyager and like yeah I'm not going to say Voyager was the first one to do this Voyager was the most popular one to do this for Minecraft and then they created tools that they could invoke on the fly to help them play the game better and compressed thinking and experience into a new tool that I can later invoke
Starting point is 00:08:27 and then I think because everyone in 24 had this insight was like okay if we just pour if you know Darry and Sam both all realized if I pour all of my heart and soul and resources and compute and intelligence into getting better at coding, we will have lift off. Then I can automate the ML engineer, then I can have lift off and I'll have all the things. I'll solve all the things. And so then that's what everyone did. And I don't
Starting point is 00:08:53 think that they had the insight. I don't think they were so insightful to know that, oh, if I did this, then we will have LMs that would be good enough to write policies for robotics. And I can displace all of robotics with actual coding agents. And I could displace artists with, like, you know, coding agents that are writing JavaScript to make beautiful, like, art. Like, I'm not sure that they had that insight, if you honest. Could we talk a little bit about what some of these early CODIS policy methods were even able to do?
Starting point is 00:09:21 So even, you know, this is well ahead of us having coding agents that are widely available and people intuitively understanding how they work. It's well ahead of people using REL methods to actually train those coding agents to be really good at tool calling. What did some of those early ones, maybe even before Voyage or, like, code as policies, which I think the paper came out in 2022 at the end, right around when Chat Chb-T came out. What was the capability of those compared to what you can see now? Yeah, those Codas policy papers, especially ones for the Google DeepMind team, were incredible
Starting point is 00:09:48 because what they did is they created these kind of functions, like pick up an object, lift up or like move to certain posts. They provided this list of functions to a coding agent in a form of like literally Python functions. And then decoding agent will be able to write code involving these functions to then control the robot to do very complex hacks. And what was very surprising, or what was excelled was that coding agents can do this very one shot. They did not need additional robot data in order to work with this code
Starting point is 00:10:20 because they're already trained on so much coding data. They already have a sense of like how to, what to do first, what to do second in order to like move a block into a bowl, for example. And I think it's this one shot ability, this in context, exploration ability that really motivated a lot of later work to continue exploring, including us, to continue exploring how we can apply LLMs to robotics. I know, Francois, you have this framework we've talked about of, you know, the various ways that learning can happen.
Starting point is 00:10:46 There's going to be learning that's embedded into the weights versus being context learning and so on. Do you want to quickly talk about that? And maybe we can think about where all the research over the last few years has fit into that framework, especially the direction it seems to be going. I mean, I've been working on this experiment and actually I haven't really crystallized it into a paper yet. But basically it's like,
Starting point is 00:11:04 what is the most efficient way from an intelligence sample perspective to input a learning, let's say, in this context, a state action results or state action reward tuple
Starting point is 00:11:17 back into the policy. There's ICL. There is... So it's in context learning. In context learning where like I just like append it. This is how most people are using LMs. They're just like, oh no, don't do it like this.
Starting point is 00:11:31 Do it like this. And it's like, okay. And then it's still in the context. And then I can remember, I can iterate. And that really only works if you train the LLM or at least post-trained the LLM to be able to learn and improve. And that self-refines a reflection, like literature from way back that, like, actually allowed multi-turn into the training set. If you don't have that, it doesn't learn, it doesn't not any kind of learn. But even then, there's a limit to how good ICL can go.
Starting point is 00:11:56 So I've done this experiment where, like, you take an LLM that we've trained on. and I've held out a task, let's just say GSMAK for simplicity. And then I ICL it and I want to measure on the Val Set how much it improves on a per sample basis. And it improves greatly, very cheaply. It doesn't cost much. I don't have the, there's no SGD, so I don't have any flops. And so very quickly I can adapt and improve on my Val set. Example, example, example, one is non-monitonic improvement, which is wild.
Starting point is 00:12:26 So it gets worse, it gets better, it gets worse, it gets better, like pretty aggressively. Number two is that it caps out very quickly. And so after like 20, 30, maybe 40 examples, it is basically saturated and more examples back into the context, it's not improved. You're like you're constrained by the model's ability to intelligently use all of its context. Exactly. From the post-training. How many multi-turned did it actually get to be able to improve and do self-reflection? And then the most important one for surely can't do is after it hits the context window, context length of the model that it was trained on.
Starting point is 00:12:58 If it was trained in 100,000, really that means you have a context window. of like 50,000. Once you exceed 50,000, you don't improve anymore. You actually just get worse because the model can't attend over everything. There's a really good example you said, which I didn't really think about, is like, well, if I can rag or I can load in my act into active memory, and this is very prime agent continual harness kind of like thinking where I can pull stuff in on the fly, similar examples, then it's like me like, hmm, I have an exam, I have an exam with a textbook. I'm going to do better if I have the textbook to look up. as reference.
Starting point is 00:13:32 And so that is definitely going to work. And it's still, and it also scales, I would imagine much better. And then the last two paradigms would be Laura with rank one or rank two or whatever. Laura with rank 10 or rank 100. And then all the way to full SFTRL. And if you're Tesla and you have all the data,
Starting point is 00:13:55 infinite data, and you're learning, doing self-driving car by ICL, what are you doing? Like, are you kidding? It's like, clearly we wouldn't do that, right? And if your figure or something like that.
Starting point is 00:14:05 But it's amazing how good in low data regimes you can do with ICL and then even cooler where you can compress the ICL into tool use, into tools, or like compress it into learnings. And that's the hierarchy that we haven't really figured out that you guys are probably on the forefront of. Yeah, do you guys want to elaborate on that a little bit? Because I imagine that's a central part of how you guys think about building a lot. There's a few things that are quite interesting here, I think, to unpack.
Starting point is 00:14:29 Yeah, the first thing is we've been sort of thinking about this idea of a harness, almost as a form of, like, domain specificity. Like when you, for example, deploy a robot in a new environment, maybe it's in a wet lab and it just needs to do a lot of tests you picking up, for example. You can learn a lot of this in context, but then one way of consolidating this context for future agents is sort of packaging each of these skills that you learn into specific programs, for example. Like writing out these skills and write out these memories is a form of consolidating. It's like a form of dissolution from past experience for your future agents to use. So this is quite interesting. The other thing that's quite interesting is it almost seems like there's some relationship to the broader meta-learning literature throughout in machine learning history.
Starting point is 00:15:18 Like you have this sort of bigger model that then programs a smaller model to do things. And one of the very interesting things is that, yes, these smaller models are going to learn in context. they're sort of wrapped inside the harness and they're doing the task. You can make them maybe smaller. You can make them run faster as long as the bigger model can sort of do this domain specialization of your harness quite well. So this is something that we're pretty interested in right now. And last point as well, I think the really interesting theoretical question is how far
Starting point is 00:15:45 in context learning can get you. And I think it's really anyone's guess as to that. Like there's papers that show that these models sort of can approximate gradient descent during in context learning as well. So it's this, this line. between weight space and symbolic space, I think is a little bit blurry. YC's next batch is now taking applications. Got a startup in you, apply at Ycombinator.com slash apply.
Starting point is 00:16:08 It's never too early, and filling out the app will level up your idea. Okay, back to the video. So what exactly are we watching here? We're watching Asra controlling the arms to pick a block off the table and put it into the bone. So here, Astra is using just these camera inputs, and presumably it's aware of the robot that it's controlling, and it's going to write code that controls a robot at its individual joints. Yes, so it does look at all the cameras.
Starting point is 00:16:34 It has all the camera feeds. Instead of code, it's more like a tool call. So it sends commands to the robot to control and put the block into the bowl. So now in this situation, we're using Aster directly, but if we were using, say, Waddle's harness, what's the kind of difference between directly using Aster to control this versus doing this through Waddle's API. Sometimes directly, like having a
Starting point is 00:16:58 quick command the post a robot should go to, is not the optimal tool to use. If it's a repetitive task, you don't want Astra in the loop. Maybe you want to write code. It can run repeatedly very, very fast. Or if it's a task, you've done a similar task before, you should be able to call a skill that has compiled and use that skill to do the task faster
Starting point is 00:17:16 and handle edge cases better. So here we can see it seems to be approaching in on this guy. Oh, very dramatic. Yeah. You can notice that. notice that the latency is a bit slow because we are bottlenecked by the latency of Asha. But if you look at the trends, the latency of these models are improving very rapidly. Right.
Starting point is 00:17:34 So one thing that we saw is that for Fabo class LLMs, their latency is improving by around 2x per month, which is very, very fast. Very fast. Yeah. If the trends continue, we could get real-time control by end of the year. So now that we saw this task go from end to end, why don't we walk through what that actually entailed. So there's a coding agent in the loop here running this. Yes. What are the steps that the coding agent would have taken? And let's kind of contrast
Starting point is 00:18:00 the direct coding agent version versus the wattle harness in the loop version. Yes. So what just happened was that in a series of turns, Astra received the images from the cameras and it outputed the end-effecter pose for the robot to go to. And it does it over multiple returns to complete the task. This is less of code as policies, but more of tool calls. The unscreen part is pretty repetitive. So you could actually automate a lot of that using code, as well as actually approaching the bottle, picking it up. Like a lot of these things are pretty deterministic once you've done a task quite a few times. The interesting part is where you build in the variation into the code as policy graph. And so we have a few
Starting point is 00:18:47 points of variation, this can come from, for example, when you detect an object, maybe you use a VLM as far as your tool call. Or when something fails, for example, how do you check that it's failed? How do you then potentially do something else, depending on the failure? Like, there's kind of more flexible responses. We tend to put a VLM inside a loop to ensure that, you know, while the coding graph is sort of deterministic, there's points of variation that allow it to generalize. The biggest insight I think I've had a change in worldview on how to perform machine learning and how to get the rest of the distance on AGI was like a lot of conversations I've had with Francis Chalet.
Starting point is 00:19:23 And in 2020, maybe even 2018 when he did on the measure of intelligence, he talked a lot about transduction versus is just wrong versus program induction. And transduction just means I'm learning a function theta that maps from X's to Y's. And why is that wrong? It's just slow. And it's information inefficient. And so to go from X's to Y's, I need a lot of pairs of X's of Y's. And if I have a very small amount, then you need a lot of inductive bias.
Starting point is 00:19:50 And what you need is a good generator to map from X's to Y. So now my theta takes in three N pairs of X's and Y's and emits the function that maps from X to Y's. And that's what code is. And so that's like if you give me a coding interview, even a whiteboard, you say, okay, here's a coding problem. Here's some examples. Okay, cool. write the function. And then we are emitting an F that will map from X to Y. And this is quite interesting, right? Because I think, I mean, this was a lot of the traditional program synthesis
Starting point is 00:20:22 literature. And I feel like a lot of reason why those methods didn't work as well was because that inductive bias, as you said, right? You're shifting the difficulty of the problem from finding that mapping to find the right inductive biases to map it to this like smaller code to then do the mapping. But finding that set of inductive biases is super hard. But maybe that's what Astra is buying us. Like, even in Coxide, there's a giant shift in the literature from, you know, specifically neurosymbolic methods to just using Astra to write code to map. And so maybe that's a cool way to think about it. It is neurosomboic. That is neurosimolic. We have neurons and then they're emitting symbols. Yeah. I think there's actually a pretty good
Starting point is 00:20:57 segue to a kind of final topic, which is around, inspired by this paper also from Philippa Sola that he titles the Platonic Representation Hypothesis, right? Where platonic, because as a reference to Plato's Cave and the idea that, you know, they're seeing shadows of various realities. And I think here the point he's making is that there's extensive evidence to show that language models and various representations of different data actually learn these distance mappings between similar things under different training policies that are overlapping. And so like the idea there would be that as you train these systems on increased large amounts of data, they converge to a consistent mapping of the world. And maybe this
Starting point is 00:21:45 implies that we would expect language models to get better and better at things over time that would make them useful for new tasks like robot control. But maybe, I imagine this representation hypothesis is actually pretty central to all of your guys' worldview or your view or your view of your companies. And so maybe why don't you guys tell me a little about how you think about this? And also, maybe we can use that to make some guesses as to why we think models like Astra seem to be so much better at robot control than previous models and we're, you know, make some predictions from where we might be going here. Yeah, I think one way to make this concrete for the robotics models versus language models
Starting point is 00:22:19 debate is that the very, very strong language models will have very similar representations of the world with very strong robotics models. And if that's the case, then if you have a really strong language model, you also have a really strong robotics model. using if we believe in the platonic representation hypothesis. And if that's true, then bitter lesson is the best manifestation of the bitter lesson, right? Because you just need one really strong model, regardless of architecture, and they would be upperforming any specific models that is slightly weaker. The interesting thing that's maybe not entirely intuitive,
Starting point is 00:22:59 and maybe I'd be curious to hear you guys all think about, is why specifically it seems like these newest models, specifically Astra seems to be so much better at spatial intelligence. You know, we've all seen the demos online of, you know, controlling blender and making great 3D images. Now it seems like, you know, you showed Jay in a benchmark that is like a meaningful step function improvement on certain tasks. Where do you guys think the pre-training and post-training, I guess, of those models? What has likely changed about opening eyes approach there that made them so much better compared to models even six months earlier? Like, you know, that's still we're benefiting from the better lesson and we're still extremely good at code.
Starting point is 00:23:32 and we're still solving math problems, but haven't quite cracked the spatial intelligence needed here. I think what Astra does incredibly well is its vision capabilities. It was probably pre-trained on way more computer-use data than ever before. It's already pre-trained on so much like CAD data. And it's like all of these kinds of data probably teach the model like similar understanding of like physical world as like a lot of robot data might. So the computer use data is an interesting point because it's not totally intuitive, I think,
Starting point is 00:24:01 why computer use data is useful for understanding the physical world and that it tends to be like a computer where you're clicking around. Say more about why you think that is a big unlock because I totally buy they train more computer use data than ever before. Right. I think immediately if you were just training a model for computer use, that would not be able to apply for robotics. But if you feed a computer use data into a big model like Astra,
Starting point is 00:24:22 and by computer use data, I mean like you drag a cursor around on a screen to orbit some cat object in order to design and blender, this tells you how to reason about spaces. It at least tells you about like top down, left, right, all these concepts that you need to control a robot. Right. And I think that's why this data helps so much for making these LLM so much better at robot use. I think like from the other direction, like a lot of people in a robotist community were trying more and more different kinds of data as well, like egocentric videos. Kind of went from just teleoperation to like, you know, a broader range of this data because it's not just robot data that can teach a model how to use a robot.
Starting point is 00:24:59 So if we take this to the extreme, it's like, why not feed every kind of data coding, computer use, egocentric into the same model? I think that's how we get to the most capable, like, robot use agent. And it's so funny that like, you know, if you go to 1980s, X-Park or whatever, like, we literally made the graphical user interface to be more like the physical world so that we could interface with it. And like, what we ended up doing is building an environment that was actually helpful for robotics to learn how to use the physical world. Right? We have file systems. We have like files. We have folders. We have, you know, the GUIs for like solid works and like auto desks to like spin around stuff to make it similar to the physical world. And then like we couldn't get robots to work in the physical world. So we just trained on that. And then now it works. Right. Exactly. There were incredible papers, I think from Princeton and also CloudPace Robot touches on this. It's like if you design a right harness where you make the tools that can face it with the robot look like computer used tools. Like you have a, you have a, you have a, you know, and drag a cursor around to control what a robot goes, that improves how well the LLM is able to perform on these physical tasks. Where do you guys see, you know,
Starting point is 00:26:07 given reasonable guesses as to where the base models are going to continue improving and your guys' own investments in either evaluating these models or building harnesses around them, what do you think is going to be capabilities that we maybe now see as challenging to do, but that are going to be increasingly possible or even trivial very few months from now? Like, you know, six months ago, even the demonstrations that I've seen you guys post on your Twitter, I think would have been kind of mind-blowing to imagine coming from an L.M. There's some consensus within the Frontier Labs and also in the Robotics Foundation models companies
Starting point is 00:26:37 that we will have general-purpose robots within the next two years or even earlier. And this is something that society is probably unaware of or even unprepared for. And when we say general-purpose robots, we mean something like if you give any natural language instruction, it can do what a competent teenager could do with their hands. It's kind of like the chat GPT moment, but for robotics in terms of capabilities where it can generate through unseen tasks and unseen environments. For you guys, for the Wild Labs guys, what does that mean for what you guys are building? I think we're to want to take steps to get there in about two years' time. I think there's so many challenges that are very visible. For example, like
Starting point is 00:27:17 let people online talk about latency. This is a big problem. If you just have Astra being in the loop thinking at every step. This is really slow. It's not going to be economically, like, useful. So how do you, you know, consolidate kind of a first pass by Astra, like, this context or learning into like a faster skill or policy that you can then run repeatedly at incredibly high throughput? I think it's the mappings to humans will get more and more like this. where like the optimal thing to do may be to put a new S-A-R state action reward back into context because it's very quick. But then there needs to be some like of go to sleep for a while, like almost everything
Starting point is 00:27:58 that is intelligent sleeps. Like tell me an intelligent system that doesn't sleep, right, in some way. And then during sleep, compression happens. The weird thing that happens from your hippocampus and it shortwave ripples to both lobes and like there's weird, you know, pass from memories that were compressed throughout the day. to train the weights. And similarly, maybe the right thing to do is similar to dagger a data segregation framework in our classic RL where you're going, you're collecting a bunch of data, and then
Starting point is 00:28:25 you're somewhat reflecting on it, and then you're using it to update your weight file, and then you have distillation from those experiences back into an updated weight file that maybe isn't Astra, but maybe is your own models. It sounds very dream coder-esque, but a lot of those specific tools, I think that people used to build, to take DreamCoder, for example, right? It has this library of skills, and then during the sleep phase, it basically refacted everything into more compact representation and so on. I mean, we're seeing sort of similar things just happen not as rigidly as before in the space
Starting point is 00:28:55 of programs, but now it's maybe like maybe you want to refactor traces. Maybe you want to refactor skills. A lot of these robotic code as policy papers, right, they have this growing library of skills and as anyone's guess really how you print them, how you organize them and things like that. So, yeah, I think as Waddle goes forward, like how you manage. our growing context of skills, of deployment data, all that is going to be quite interesting. I think with that, I think this was an excellent discussion. Thank you so much, Vincent, Hanming, and Jay, for being here.
Starting point is 00:29:24 I'm very excited for all the incredible robotics advancements. I think we're going to see over the next few years, and I think you're totally right, Jay, that I don't think broader society is totally aware of how much is coming. It's going to be a very incredible few years to come. Thank you so much.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.