The a16z Show - Fei-Fei Li on Spatial Intelligence and Robotics

Episode Date: July 28, 2026

Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI's biggest unsolved problems: how to give machines a true understanding of the physical world.... Martin Casado sits down with Fei-Fei Li, co-founder and CEO of World Labs, creator of ImageNet, and pioneer of spatial intelligence, alongside Yunzhu Li, co-founder of SceniX and assistant professor at Columbia University. They discuss why World Labs acquired SceniX, how simulation can unlock the next generation of robotics, and why training robots may require a fundamentally different approach than training language models. The conversation explores real-to-sim-to-real pipelines, world models, robotics foundation models, evaluation, synthetic data, and why the future of AI depends not just on understanding language—but on understanding and interacting with the physical world.   Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Yunzhu Li on X: https://x.com/YunzhuLiYZ Follow Martin Casado on X: https://x.com/martin_casado Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 We are building the next frontier of AI, which is what we call spatial intelligence. As cynics, we are developing what we call a real to seam to real pipeline. We can replace all the data or the evaluation we need in the real environment by using the data that can generate at a scalable way in our digital world. Think about human intelligence. We do a lot of simulation in our head. You know, why there's a very important role simulation
Starting point is 00:00:30 place that real-world data doesn't play, which is counterfactual reasoning. What we are building is a consistent world. Consistence both over space, over time, over different viewpoints, and over different type of interactions. My North Star is I won the robot work. The world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces. Do you believe we'll ever be able to build robots that have the power,
Starting point is 00:01:00 efficiency of a human being. How far away are we from this? Is this like five years or this is like never? The TLDR is... Language models transformed how AI understands words. The next frontier is teaching AI to understand and act within the physical world. Following World Labs acquisition of Cinex, Martin Casado sits down with Fei-Fei Li and Yun Ju-li
Starting point is 00:01:23 to unpack the vision behind the deal. They discuss spatial intelligence, world models, simulation, and simulation, and why solving robotics will require a new generation of AI, built for three-dimensional reasoning, not just language. All right. Well, it's great to have you both here. So, Faye, for the listeners that may not have the background, maybe you can give an overview of what World Labs does.
Starting point is 00:01:49 Yeah, well, World Lab is a two-year-old startup. I think we should just recognize it's a frontier model lab. We are building the next frontier of AI, which is what we call spatial in time. And spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces, whether it's physical or virtual. And of course, a means to an end towards spatial intelligence is building large world models. And that's what World Labs is mostly focused on. Yeah, so you've been saying this since the very beginning, which is the machine's ability to perceive and reason about spaces and act on spaces. But I always had the assumption that the acting on spaces was some long-distance shooter thing, but now you're acquiring a robotics company.
Starting point is 00:02:44 And so maybe talk a little bit about the timeliness of this and the intentions. Yeah. So first of all, it doesn't just take robotics to act within spaces or to interact, right? I mean, look at the creative field, whether it's VFX or gaming or design. Many use cases, you can create and act within virtual spaces. A world-app thesis has always been that the world we live in can be multiverse, that we create technology to allow people, builders, developers to act within different spaces. Having said that, the ability to act within the physical space, is one of the most exciting and most profoundly important capability of the future AI world.
Starting point is 00:03:33 So robotics is very much that. So World Lab has always believed that robotics is an important application as well as use case of spatial intelligence and world modeling. So by joining force with inviting CynX and Cynx team to World Labs, is part of our long-term vision and mission. We've always committed to that. Amazing. So, Yunchu, you're the co-founder of Scenics. So maybe provide everyone with a quick overview of your background and what Cynics does?
Starting point is 00:04:09 Yeah, so I'm Yunchu. So I'm currently co-founder of Cynics and also assistant professor at Columbia University. Wow. So my research started from my PhD at MIT and then postdoc with Pfei. Really? Yes.
Starting point is 00:04:24 That's great. The world is small. It's very small. Throughout my career, my goal has been very simple, trying to help the robots better perceive and interact with the physical world. So I'm a very practical person. I want my robot to work in a real physical environment. So for cynics, the unique opportunity we see is that there has been a lot of bottlenecks. Right now, we see faced by the developments of general purpose robots,
Starting point is 00:04:48 especially around training and also around evaluations. So as cynics, we are developing what we call a real to-sumvirate. seem to real pipeline. We're going to map the real environments into the digital world that has the best alignments with the real environments. By alignment, we mean that whatever happens in the digital world is also going to happen in the real environments, such that we can replace all the data or the evaluation we need in the real
Starting point is 00:05:14 environment by using the data that can generate at a scalable way in our digital world. So that is how everything started in Synix. We put together a very, very strong and best teams around robotics, robot learning, and also simulation and rendering, trying to build this real-to-syn to real stack to solve some of the key bottlenecks. It's amazing that you two work together.
Starting point is 00:05:35 Yeah, and there is a funny story here, because you would think because we work together, he was my amazing poster, we've been talking about the cynics and World Lab integration for a long time. It's actually not true. They came into World Labs as a customer. Really? When we released the first version of our generated model
Starting point is 00:05:55 called Marble last winter around November, December. Cynics just sign up. No kidding. Is it a customer? Yes. And I didn't even know what it was. And then I realized this is Windrew's company. I called Windrew.
Starting point is 00:06:12 I'm like, wow, this is your company. And then we realized there's so much synergy. Maybe Safi just quickly describe what Marble is. Yeah, Marble is the code name for the base model that World Lab is being used. training and iterating on the fundamental capability right now of marble that is publicly released is to take a prompt. It can be an image, it can be a few images, or a text, and turn that into a geometrically consistent world that can be represented in 3D geometry, whether it's Gaussian
Starting point is 00:06:48 splat or mesh. Really what Cynix team is doing is trying to solve this extremely difficult problem in robotics, which is the lack of data. The lack of data in training, the lack of data in evaluation, this is very, very different from language models, where data is abundant on the internet. And we know that in order for robotics to work, we have to somehow unlock the power of scaling law. But where does that come from? This is something that, it's a profound problem that everybody is battling with in robotics. It would actually be great to talk about this energy. You have put together a very, very talented team.
Starting point is 00:07:29 You have put together a very talented team. And so to what extent is there overlap? To what extent is this an extension? Maybe talk a little bit about that. Yeah, that's how complementary it is. It's actually the TLDR is very complementary and with the shared mission. So, Windr is one of the three technical co-founders. The other two are Changsie-Zen, another Columbia professor who has been a world-class
Starting point is 00:07:52 technologist in simulation. And Changsia has his background in also VFX. He worked at Weta. He worked at Tencent. He's being an entrepreneur. Then there's Sonny Hu, who is a phenomenal engineering leader, who was also in a startup that was acquired by Amazon many years ago. So he worked in many different tech stacks in the computer vision field in Amazon.
Starting point is 00:08:18 So when we started talking more seriously, I recognize that a couple of things, that Cynix has from a talent point of view is extremely complementary to world labs. One is obviously Andrew's incredible thought leadership and just technical prowess in robotics, right? So from really, from hardware, full-stack robotics, and even when he was by postdoc at Stanford, at that time, you already had your faculty offer.
Starting point is 00:08:49 So you were there only for one year. I wanted you for more than one year, but he had to go become a happy, the real job. So he was a full-stack researcher in robotics, from modeling to hardware. And of course, Yunru and his student as cynics was that pool of talent world life hasn't had yet. Then on the Changsci side is just an incredible simulation capability, right? He's such a senior researcher and technologists in simulation. What World Labs is doing is very much interfacing the world of simulations.
Starting point is 00:09:30 So I think what they don't have, obviously, is on the generative model side, as well as the computer vision 3D reconstruction side, we're also very strong at World Labs. So that's a technology that Cynics needs. So together, these two sides come together and make it much more complete. Faye's motivation in this is like this is an extent. and a compliment to get into robotics. Having been in your situation, which is deciding when to sell a company,
Starting point is 00:10:00 it would be great to hear from you and how you think about joining World Labs and kind of the fit there and why you made the decision to do it. Yeah. So at the very beginning, we were deciding, okay, do you want to just keep going? But after chatting with Fifei,
Starting point is 00:10:14 after seeing all the synergies that are happening in the middle, it just makes perfect sense for the forces to join each other. So in any sense, as cynics, what we have been doing is real to seem to real. is to do dense reconstruction of the environment. So we capture the appearance of the environment, geometry of the environment, and also the dynamics of the environment,
Starting point is 00:10:33 meaning how the environment is going to change when you apply actions. So this dense reconstruction right now is still a little bit on the heavier side. And what World Labs right now has been doing involves a lot of profound capabilities around sparse reconstruction and generations. So we see a lot of opportunities
Starting point is 00:10:51 of leveraging like a marble and other capabilities as World Labs in order to do very efficient reconstructions and modeling of the environment. So can we expect a foundation model for robotics from World Labs? World Labs is building a foundation model. As you know, Martin, we're building a base model. And as the technology has been evolving, some of the most exciting base models are Omni models, right?
Starting point is 00:11:19 They take multimodal input, they have multimodal outputs. And what is a foundation model for robotics? It's very likely going to involve actions. It's very likely going to involve the output of actions in addition to the state of the world. We're definitely not ruling this out. Yeah, great. So for example, for the foundation models,
Starting point is 00:11:44 it essentially needs to be a multimodal model. So it has to take into account from text, image, depths, and different kind of modalities. And action is a very, very important part. parts of that modalities. So if you think about frame actions as the input, that is essentially a forward similar. That is going to predict how the environment is going to change
Starting point is 00:12:03 when you apply a specific action. When the action is output, this is essentially a policy model. That is trying to predict giving a specific goal, what should be the action you take in the real environment to get you closer to that goal. So this kind of only models actually can benefit a lot and actually provide huge amount of values
Starting point is 00:12:21 for the robotics communities in trying to understand how to make. model the environments and at the same time how to act in the environments. And this can also act as a backbone for you to find you into specific robotic applications to making sure it's really live up to the
Starting point is 00:12:35 reliability and efficiency that expected by the clients. You know, if you don't mind kind of a lay investor question, I see a lot of robotics companies and a very popular approach right now for the robotics companies that come in is like we'll use a video model,
Starting point is 00:12:51 you know, and like, you know, that's the predominant method where this is, you know, 3D and simulation. It's a very different approach. And so maybe you could contrast this popular approach of just using video only versus kind of what the ambition here is. Yeah. So in order to create words with the robot can learn,
Starting point is 00:13:10 the words, as I mentioned, is to capture the essential structure of the problem. And one of the very important and necessary requirements for those words will be consistency. So that is where I actually see there's very, very strong synergies with marble, because what we are building is a consistent world. Consistence both over space, over time,
Starting point is 00:13:30 over different viewpoints, and over different type of interactions. And marble, the generated words from marble, is also provide an infrastructure, a component of that entire words that we believe is necessary for the robot turner. Imagine if a robot's pushing an object forwards, the object just magically disappear,
Starting point is 00:13:46 which has been a problem of many of the existing video prediction models. It's one provides good enough signal for the robot to know what is the right thing to do. But obviously right now there has been a lot of investigation on building better and better and stronger and stronger, like video models. So we actually see a way where some of the infrastructure we build can provide as initial momentums. And to go in through this data flywheel of going from this like a more simulation-driven models into like a robot policy models, which is going to do the execution in the real environment, collecting new data, the data will, come back in, where the model doesn't necessarily have to be physics only or learning only,
Starting point is 00:14:25 but somewhere in the middle, which be able to capture the essential structure of the problem, but at the same time, be able to scale and become better and better as you accumulate more data. You know, I've worked now Sefei very closely for a while, and you've always had this North Star, which has driven this, and, you know, you've articulated variously as kind of 3D and in a number of other ways. And I'm just wondering for you, is there also a sense? a more philosophical North Star or you're more the pragmatic like I am. I build the system, like do the thing.
Starting point is 00:14:58 My North Star is to make robots work. Amazing, yeah. In the real environment. I'm a very practical person. I want the robot to work. One interesting thing that's actually coming from my collaborations with Phoebe during a postdoc, we are building this kind of benchmark. We actually send out surveys asking the general public,
Starting point is 00:15:14 what do you want the robots to do for them. Among the southern tasks we collected, one third of the tasks, are about cleaning. People just don't like to do those, like, dial and dirty tasks. And those are the scenarios. We really want to make sure we have robotic solutions to deal with.
Starting point is 00:15:30 One thing I really like about cynics, Martin, especially continuing your question, there's a lot of robotics companies building models and all that. One thing I truly like about cynics is Rindrew and his co-founders have such an incredibly pragmatic approach to robotics.
Starting point is 00:15:50 They, especially they come from academia, right? Sunny doesn't, but in Zhu and Chanxi come from academia, but their first instinct is work with design partners and customers in real industry, whether it's labs at industry labs or warehouses or electronics assembly. Electronics assembly. That is such a refreshing, actually, a refreshing way of,
Starting point is 00:16:20 for approaching robotics. And that really true made me very excited to work with them. This is for you, and you? But I'll just be this is personal curiosity, which is it seems to me that for robotics, you have to be pretty exact. I mean, not perfect, but pretty close. But for the creative use cases, which Boralep's
Starting point is 00:16:37 done a lot of, you kind of don't need to because, you know, I mean, even sometimes like being wrong is stylistic or intentional or whatever. And so from a technical perspective, what is the challenge here for reconciles? these two things, or do they never get reconciled? Like, there will always be two points in the design space. So they will be reconciled in the long terms, of course.
Starting point is 00:17:00 And modeling of the environments doesn't have to be perfect. The model doesn't have to be perfect in robotics. And by the way, again, this is pure curiosity, but is there like a bit more formal way to say that? Like, what does that mean not to be perfect? It has to be pretty close. So let me put in this way. For example, models over the developments of all different kind of robotic applications
Starting point is 00:17:22 has been a very important cornerstone. If you look at all the existing robotic applications, like Plain, Jones, Rumba, or even for quadrapad robots, bipedal robots, model has been the way for them to actually work and be able to transfer from simulation to the real environments. But if you look at those locomotion robots, like quadruped robots, bipedal robots, they can walking on snows, they can walking on bushes, but you don't need to have a simulator. They can simulate all the bushes and snow. like very precise.
Starting point is 00:17:49 You need to have a simulation that captures the essential structure of the problem and do a whole different kind of randomizations inside the digital environments. So that is what we're aiming for. So basically with synics and together with word labs, we're trying to investigate what is the level of fidelity we need to model the messy and mess of worlds besides the robots.
Starting point is 00:18:09 As I said, we'll be able to transfer the robotic systems training the simulated environment and digital worlds back into the real scenarios. As an investor, I've heard, other researchers, say like Sergey Levine, say simulation will always eventually deviate from the physical world and real world data collection is absolutely critical. And so maybe talk a little bit about like the viability of this approach where simulation is a cornerstone as opposed to some other approach.
Starting point is 00:18:38 So they don't contradict with each other. So if you think about the simulation, simulation essentially trying to predict how the environment is going to change when you are. the actions. And this is essentially a model of the world that doesn't necessarily have to be pure physics. It can be a combination between both physics and also learning. We are collecting real-world data. We will be using those real-world data.
Starting point is 00:19:02 It's just at different stages of this, like, a data fly wheel. Maybe at the very beginning, we have stronger emphasized on we have more physics to making sure we have the right consistency and right structure for us to learn the world, for us to train the robot policies. But as we accumulate more and more data, both through data collection and also through the collaboration with our clients,
Starting point is 00:19:22 we'll have the data that will be moving towards more towards more learning-based, like modeling of the environments. So this kind of transition and also this kind of data file is really enabling factors of both getting the best of both physics and the geometry and consistency, as well as all the power and magics from the data and computers. I want to add to this and be slightly philosophical here, is there isn't a binary choice between simulation or no simulation. All this come in together to make robotics work.
Starting point is 00:20:00 Think about human intelligence. We do a lot of simulation in our head. You know, why? There's a very important role simulation plays that real-world data doesn't play, which is counterfactual reasoning, is that you play out events, that hasn't happened or cannot happen, or you don't have enough data to make it happen in real world.
Starting point is 00:20:23 And while you play it out, you learn how to act in it. Humans do this all the time. We probably don't, you know, we just, I know you were at World Cups. I was at the World Cup. Congratulations to Spain winning. I'm sure in the planning of every game, there is simulation, whether it's digital or on the whiteboard or whatever,
Starting point is 00:20:46 that simulation, the role simulation plays is counterfactual reasoning. And that's really important in robotics because we just do not have, cannot possibly have, enough real-world data for that. Here's a real-life example, the industry of self-driving cars. Waymo has officially said they use billions of hours of simulation.
Starting point is 00:21:10 And actually, Waymo is more simulation-heavy than just real-world data heavy. So these are real examples. And as you know, Martin and Andrew, too, cars are the simplest kind of robots. Yeah, 2D, yeah. Yeah. So clearly, simulation plays a huge role in robotic learning. I also want to add to that.
Starting point is 00:21:32 So, like, there are, if you put things more specific, simulation can provide two levels of benefits. The first one is reliability, and the second one is efficiency. So for reliability, if you're thinking about a robotic system working with a reliability, in the real environment. You need data to provide systematic coverage of all the state space and variations that robot mining control. That's how you can learn of how that is robust. So with simulation, you can do systematic randomizations and control
Starting point is 00:22:01 and variations of lighting, frictions, geometries, object types, and also all different kind of physical parameters to making sure you have sufficient coverage of the state space. So this is what can give the robotic systems reliability. And second is about efficient. So right now, many people are doing teleoperation. And if you look at many of the teleoperation device, imagining all the actual skeletons you are using,
Starting point is 00:22:25 you're actually collecting the data at a speed that is actually slower than human actually doing the task. But for many of our clients, human speed to them is not good enough. They want faster than human speeds. So for the robot to move faster, it's not as simple as just drives the robot faster, because the gravity doesn't change. But in simulation, you can do systematic speed up of the brain.
Starting point is 00:22:46 of the robots' behaviors to train the robots such that it considers all the dynamics, changes of the environments. So this is what can give like our clients for them efficiency. So both for the reliability and efficiency, though there are some kind of like a very unique like values where simulation can provide. You've talked about the technology and the platform, what it does. Let me talk about the specific use cases people use it for. There are essential like two specific use cases, especially around both training and also around
Starting point is 00:23:16 evaluations. Starting from the evaluations. So evaluation is something like people often overlooked in the robotics. But if you are tuning like a robotic models, you have to know how well it works and that is the only source of information for you to iterate.
Starting point is 00:23:32 By the way, every AI person really understands what evals are and uses it all the time. Non-AI people, it often means something a little different. So maybe it's even worth just describing specifically what you mean by evaluation. Okay. So So what I mean by evaluation is we'll be able to understand for this specific checkpoints, how will does it perform?
Starting point is 00:23:52 Does it perform, for example, 95% of the time or 99.9% of the time? And the key criteria people use in industry is, how long does it take? How long in work-clock time does it take for you to distinguish between a checkpoints that is 90% from a checkpoint that is 92 points? And if you only do that in the real environment, that's just take so long. for you to do the distinguishments. And if you really think about also the robotic evaluations, right now people are doing in the real environments,
Starting point is 00:24:23 the iteration speeds is multiple orders of magnitude slower than iterations of those language models. So not only is like the robotic tasks very varied, very diverse. Oh, you actually have to do the thing. Yeah, yeah, yeah, right. Like atoms have to move through space. Yes, exactly. But only there are the laws of physics have to be obeyed.
Starting point is 00:24:45 watch those robotics videos, every video has like 10x, 8x, because it moves so slowly. Exactly. So not only is slow, it's dangerous, it's costly, but at the same time, the speed is also like multiple hours of magnitude is like slower. So some of our clients actually needs this digital environment that can be used to evaluate their robotic like systems. And because our digital environment has proven alignments with the real words, So meaning whatever happens in the theme
Starting point is 00:25:15 is also likely to happen in the real environment. If a checkpoint is working better in the simulation, it's also highly likely to also work better in the real environments, as we have also been discussed in the blog post. So that actually gives our clients very strong confidence in actually using the data, using the signal from the digital environment
Starting point is 00:25:33 to do scalable, safe, and much faster evaluations of their robotic systems. Great. So that is on the evaluation. Then on the training. So on the training side, So basically, like I also mentioned, it's about controlability. So you want to control all the different possible variations of states, parameters,
Starting point is 00:25:51 lighting, frictions, physical parameters, like even object geometry, object types. So you want to make sure you have sufficient coverage of all different kinds of scenarios, such as we will be able to generate, like, an informative data for your robots to be robust. And this is just going to be so hard to do just in the real environment. Like we discussed, if you do title operation, the speed at which you are collecting data, is slow, you're also limited by how many robots you have, how many tele-operation device you have.
Starting point is 00:26:18 There's like a whole different kind of like challenges around all the data operations around it. But in simulation, everything can be controllable, everything can be systematic, and everything can be understood at a level where you know exactly and making claims about exactly what distribution
Starting point is 00:26:35 you have covered. To develop confidence about within the distribution, we know the robot will work. So those kind of confidence and efficiency and scalability is something that our clients also value to use our digital words for the training of robotic systems.
Starting point is 00:26:50 Here's a crazy thing. Even before Cynics and we are talking, our inbound customers for Marble were already seeing this kind of demands. We just cannot serve these customers, but we are already getting a lot of phone calls from robotics, early stage robotics companies
Starting point is 00:27:08 who are developing their models all the way to downstream, very pragmatic use cases and we're seeing these needs. When people hear you're going into robotics, what they're going to envision is you're pulling out a 3D printer and you're going to be making hardware and then you're going to be programming the brain of a robot and sticking it in the robot and then you've got a robot.
Starting point is 00:27:33 And I don't think that's what you guys are talking about here. So maybe talk about where this fits in the life cycle of creating a robot and where you will end and where the rest of the ecosystem will begin. So what we have been building, you can imagine is a infrastructure like with the software around these infrastructures for people to, for them, build words,
Starting point is 00:27:57 such that robot can learn and evaluate. And these infrastructures is naturally model agnostic and embodiments agnostic. So I just want to be very clear, just because this is actually a very subtle, I mean, for you it's obvious, but it's a very subtle point, which is, from what you said,
Starting point is 00:28:13 that's not building a robot. It's building an environment which another company can place their robot brain to navigate and to learn. Yeah. So for our customers right now, they have all different kinds of robots. Some are using single robot arms.
Starting point is 00:28:28 Some are using bimail. Some are using a fixed arm. Some are using mobile manipulators. Some are using grippers. Some are using some more elaborate versions of the only factors. So our platform right now is just naturally embodiment agnostic.
Starting point is 00:28:40 We can very easily integrate different kind of robotic embaliments, be able to put them into the words we generated, we digitalized, such as we will be able to give those individual robots capabilities of doing the right tasks and at the right levels of reliability and efficiency in the real environments. And we are also, for example, model agnostic. So we can just using the data generated by our words to train different models, either from scratch or doing post-training of existing foundation models, like vision language action models or word action models.
Starting point is 00:29:12 So to us, it doesn't matter. We just want to making sure we have the infrastructure, we have all the words such as the robot can work reliably in the real environment. You know, you have told me that you think a lot of the predictions around humanoids were a little bit aggressive and were likely to see more constrained rollouts like warehouses or whatever. Can you talk a little bit about that and how that impacts what you're going to be tackling here at like World Labs. So that's a very good question.
Starting point is 00:29:42 So if you look at, for example, all the progressions of robotic applications in the real environments, it has always followed the trend going from fully structured environments into semi-structured environments and then into unstructured environments. For fully structured environments,
Starting point is 00:29:58 what do we mean is that you have knowledge and control over all the configurations within the environments. Like factories. Like factories or, for example, car main factory in lines. Those has been automated for decades. Yeah, yeah, yeah.
Starting point is 00:30:09 And then you have, for example, semi-structural environments, which you have certain controls over the environments, for example, like the Amazon warehouses, or, for example, like restaurants, hotels, where you have certain control over the environment to just make the task easier for your robots. But there are obviously many other, like, objects.
Starting point is 00:30:28 Or, for example, clothes, those are the objects. You don't have control. And then for the unstructured environments, it's like your home and my house. Those is, I would say, the grand chat. Especially my house, trust me.
Starting point is 00:30:40 Three dollars, five-year-old. Yes. Dogs. Exactly. If you're thinking about where does the robustness coming from? Robustness coming from a sufficient coverage of the scenarios and robots might encounter. So it's so much easier and more approachable at least like right now to focus more on the semi-structed environments before we move on to fully unstructured environments. So we will move into that direction.
Starting point is 00:31:04 It's just we want to take a more sustainable and more realistic approach. towards there. I think your point here is that humanoid mimics human body. And evolution has optimized human body for unstructured environment. And so our fingers, our legs are not the best apparatus to do one thing. For example, if our only goal as a species is to climb trees, we will not have this body necessarily, right? So we'll have different kind of fingers. But what humans end up having are evolved, evolved into is this body shape that can be very general, but not necessarily best
Starting point is 00:31:50 at everything. And that is for the survival of unstructured environment. But from a business point of view, from a pragmatic technology point of view, that this unstructured environment and a generalized body is actually the hardest problem to solve. It's not necessarily even the right way to solve the problem. It's we specialize, so we take more specialized body to solve a narrower problem. But the challenge for cynics is that to be more body agnostic so that their infrastructure can serve different bodies and different semi-structured environments.
Starting point is 00:32:34 A common lens to look at exactly this question is an economic lens, right? Which is, you compare it to like generative LLMs, they can create pros or a code 10,000 times faster than a human being, a bunch cheaper than a human being. So the economic case makes sense because our brains aren't very efficient at that. However, our brains and our bodies are very efficient at 3D navigation, right? You know, like movies of the world are picking things up. And so this is just a prediction question,
Starting point is 00:33:04 but do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power efficiency of a human being when it comes to menial tasks? So let's say just basically, you know, minimum wage or something like that. Like how far away are we from this? Is this like five years or this is like never? I think it's going to take a very long time.
Starting point is 00:33:27 So if you really think about like robots in the real environments, in the end, it will. it will always be a system. So every working robot in the real environment is a system work. It's new need to be very mindful and thoughtful about how the system are coming together. The hardware, the software, the brain, even like to the details of, for example, what's the friction coefficients of your fingers? So there's a lot of things you have to consider to make these things a reality.
Starting point is 00:33:52 And it will take iterations. But what I am excited about is that I have always been at the state of the arts of robot learning and also trying to push the state of the art forward. But the state of the arts always moving faster than I expected. So what I'm focusing on and trying to investigate right now is very different from, for example, when I started my PhD. So this is a speak to how fast the whole ecosystem has been evolving and all the moving pieces start coming together or building these robotic systems. But we also have to be calibrated about our predictions. So we will see a lot of progress.
Starting point is 00:34:28 But to achieve, for example, human level efficiency and components, it will take longer. Martin, the hardest thing in today's AI is to have the right measured optimism. Right. It's totally true. Yeah. I mean, even LLMs does not have human brain efficiency. Human brain operates on 30 watts.
Starting point is 00:34:50 Yeah, that's true. So we are far from that. But I mean, performance to power it may be close, right? In narrow tasks, like software engineering. Like generating and image or software engineering, then it is, right? Yeah, I think so. I don't think we're anywhere close on the cost of robotics. This has changed how you think about your, like, strategically the level of ambition that your team can go after.
Starting point is 00:35:15 I mean, does it change that, or is it still very much in line what you expected to do when you started? It definitely changed, though, to jackfries in a very profound manners. So we see a lot of unlock in be able to do this whole process, do the modeling of the environment, you are much more efficient and much more scalable manners, especially in partner together with Word Labs. And I also want to add to Fayevei, if you think about, for example, the current states of the language models. So those are models that's with incredible capabilities.
Starting point is 00:35:45 But still, you don't just blank trust it to book your flight tickets or make your hotel reservations. You still, hopefully there's still a person who's reading the output from those language models. But that is very different from how people and will be using, for them robotic models. Based on robotic models, out of the box, the robots has to work reliably
Starting point is 00:36:04 in the real environment. So, and we don't even have the data. We don't even have all the necessary infrastructures around those for the robots to just out of the box work reliable in the real environments. So for that reasons, be able to create this digital words, this scalable digital words where the robot can learn
Starting point is 00:36:21 and evaluate within. Yes, it's going to unlock so much more potentials for being able to replace all the, like a costly, and the unsafe data in the real environments with the data generated from the words for the robots to be able to do scalable learning and evaluations. You know, I've seen many of these kind of integrations.
Starting point is 00:36:41 They actually work very well at this stage when they have this much alignment, which is great. But there's always like this question of, do you integrate now into what's happening now or do you keep things quite separate and provide kind of like a long-term trajectory that will, you know, be realized, you know, in the year timeframe.
Starting point is 00:37:00 How are you thinking about this, Fei, Fei. Is this something that integrates right away, or is this kind of a separate longer term? This is a great question. I think at this point, you know, Yun-Ju, Changxi, Sunny, Justin Ben,
Starting point is 00:37:12 and I have been talking about this. At this point, we are going to take it thoughtfully. We're not rushing to integrate everything from code base to teams because I think Cynix does have a very, well-thought, and I wouldn't call it, a standalone completely, but fairly
Starting point is 00:37:35 contained tech stack as well as their customers, as well as the kind of products they're building. We're going to take time. We definitely will, we're already on the simulation side, as well as the potential base model, action condition model side. We already are starting to talk. And also, they are using Marble as an internal customer. So we will be integrating, but we're not rushing to blend the team as like a full salad bowl.
Starting point is 00:38:11 How are you thinking about geography so this with Cynx move? Is it going to stay in the same place? We're going to, Vindra is going to move. Oh, well, welcome. Yeah. Florence to the Renaissance. Perfect. I think we, our labs is officially becoming a bi-coastal company,
Starting point is 00:38:28 where the headquarters is in San Francisco. I live in Palo Alto. I feel like I'm in a different state. But I'm actually excited that we're going to have an office in New York that can help us to attract talent on East Coast. And also we have been talking about making sure that in both offices we set up the robots so that we get to basically test us.
Starting point is 00:38:58 and mature our engineering stack so that we can work with robots remotely because we have to do that for our customers anyway. So maybe just to be very concrete, Faefe, maybe let's just pencil out. What is the perfect success case in two years? Like what product do you have, who's engaging with it, how do they use it?
Starting point is 00:39:17 Just the crisp, like what this becomes. I would be very happy that Cynix team, a world app's team, will have valid. validated customers in a small number of important vertical use cases, where our system, our infrastructure has proven to be truly beneficial to their automation needs. And these customers became our lighthouse examples to scale our business. How early, let's say someone listening to this is running a robotics company,
Starting point is 00:40:05 at what stage should they engage with World Labs? Is it really early on? Is it somewhere in the middle? So right now for our customers, because we are building this kind of real-to-themed real pipelines, where the simulation is essentially the words, we're going to provide the training and evaluation grounds. Some customers, they need only the real-to-thin part. They want to digitalize the task they care about and be able to do the evaluations of their robotic systems.
Starting point is 00:40:34 Some customers need this real-to-same to this entire pipeline, such as they will be able to have policies like running on their hardwors. So our platform is also designed in a way that is flexible, depending on what our clients needs. And at the same time, the clients who are working with are actually pretty close to the deployments, like a stage. So basically, they are working on very, very practical tasks. Those tasks, when we have robotic solutions that are there, can just create value immediately. And they have at least like tens or hundreds of like this kind of situations.
Starting point is 00:41:09 They are thinking about to do the automations for. So as like together with Word Labs, we'll be able to develop reliable solutions for those scenarios. As we have already shows, we have a number of scenarios already instantiated in our blog post. And we'll be able to further our investigation to see how they can actually solve the key requirements and also constraints faced by the real-world deployment. Great. I want to be very specific about this. Is it ever too late or too early to call World Labs if you're a robotics company? No. We want everybody to call us. We want to learn about your use case. Wonderful. If you're listening to this and you're anywhere close to a robotics
Starting point is 00:41:47 project or robotics company, please track World Labs. Yes. Thank you. Definitely open for business. Yeah, we are open for business. Not too early. All right. If you're doing robotics, called World Labs. Thank you both very much for coming. Thank you.
Starting point is 00:42:04 Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcast, and Spotify. Follow us on X at A16Z
Starting point is 00:42:21 and subscribe to our substack at A16Z. Substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments
Starting point is 00:42:46 in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures. closures.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.