The a16z Show - Fei-Fei Li on Spatial Intelligence and Robotics
Episode Date: July 28, 2026Last week, World Labs announced its acquisition of SceniX, bringing together two teams working on one of AI's biggest unsolved problems: how to give machines a true understanding of the physical world.... Martin Casado sits down with Fei-Fei Li, co-founder and CEO of World Labs, creator of ImageNet, and pioneer of spatial intelligence, alongside Yunzhu Li, co-founder of SceniX and assistant professor at Columbia University. They discuss why World Labs acquired SceniX, how simulation can unlock the next generation of robotics, and why training robots may require a fundamentally different approach than training language models. The conversation explores real-to-sim-to-real pipelines, world models, robotics foundation models, evaluation, synthetic data, and why the future of AI depends not just on understanding language—but on understanding and interacting with the physical world. Resources: Follow Fei-Fei Li on X: https://x.com/drfeifei Follow Yunzhu Li on X: https://x.com/YunzhuLiYZ Follow Martin Casado on X: https://x.com/martin_casado Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
We are building the next frontier of AI,
which is what we call spatial intelligence.
As cynics, we are developing what we call a real to seam to real pipeline.
We can replace all the data or the evaluation we need in the real environment
by using the data that can generate at a scalable way in our digital world.
Think about human intelligence.
We do a lot of simulation in our head.
You know, why there's a very important role simulation
place that real-world data doesn't play, which is counterfactual reasoning.
What we are building is a consistent world.
Consistence both over space, over time, over different viewpoints, and over different type
of interactions.
My North Star is I won the robot work.
The world we live in can be multiverse, that we create technology to allow people, builders,
developers to act within different spaces.
Do you believe we'll ever be able to build robots that have the power,
efficiency of a human being.
How far away are we from this?
Is this like five years or this is like never?
The TLDR is...
Language models transformed how AI understands words.
The next frontier is teaching AI to understand and act within the physical world.
Following World Labs acquisition of Cinex,
Martin Casado sits down with Fei-Fei Li and Yun Ju-li
to unpack the vision behind the deal.
They discuss spatial intelligence, world models, simulation, and simulation,
and why solving robotics will require a new generation of AI, built for three-dimensional reasoning,
not just language.
All right.
Well, it's great to have you both here.
So, Faye, for the listeners that may not have the background, maybe you can give an overview
of what World Labs does.
Yeah, well, World Lab is a two-year-old startup.
I think we should just recognize it's a frontier model lab.
We are building the next frontier of AI, which is what we call spatial in time.
And spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces, whether it's physical or virtual.
And of course, a means to an end towards spatial intelligence is building large world models.
And that's what World Labs is mostly focused on.
Yeah, so you've been saying this since the very beginning, which is the machine's ability to perceive and reason about spaces and act on spaces.
But I always had the assumption that the acting on spaces was some long-distance shooter thing, but now you're acquiring a robotics company.
And so maybe talk a little bit about the timeliness of this and the intentions.
Yeah. So first of all, it doesn't just take robotics to act within spaces or to interact, right?
I mean, look at the creative field, whether it's VFX or gaming or design.
Many use cases, you can create and act within virtual spaces.
A world-app thesis has always been that the world we live in can be multiverse,
that we create technology to allow people, builders, developers to act within different spaces.
Having said that, the ability to act within the physical space,
is one of the most exciting and most profoundly important capability of the future AI world.
So robotics is very much that. So World Lab has always believed that robotics is an important application as well as use case of spatial intelligence and world modeling.
So by joining force with inviting CynX and Cynx team to World Labs,
is part of our long-term vision and mission.
We've always committed to that.
Amazing.
So, Yunchu, you're the co-founder of Scenics.
So maybe provide everyone with a quick overview of your background
and what Cynics does?
Yeah, so I'm Yunchu.
So I'm currently co-founder of Cynics
and also assistant professor at Columbia University.
Wow.
So my research started from my PhD at MIT
and then postdoc with Pfei.
Really?
Yes.
That's great.
The world is small.
It's very small. Throughout my career, my goal has been very simple,
trying to help the robots better perceive and interact with the physical world.
So I'm a very practical person.
I want my robot to work in a real physical environment.
So for cynics, the unique opportunity we see is that there has been a lot of bottlenecks.
Right now, we see faced by the developments of general purpose robots,
especially around training and also around evaluations.
So as cynics, we are developing what we call a real to-sumvirate.
seem to real pipeline.
We're going to map the real environments into the digital world
that has the best alignments with the real environments.
By alignment, we mean that whatever happens in the digital world
is also going to happen in the real environments,
such that we can replace all the data or the evaluation we need in the real
environment by using the data that can generate at a scalable way
in our digital world.
So that is how everything started in Synix.
We put together a very, very strong and best teams
around robotics, robot learning, and also simulation and rendering,
trying to build this real-to-syn to real stack
to solve some of the key bottlenecks.
It's amazing that you two work together.
Yeah, and there is a funny story here,
because you would think because we work together,
he was my amazing poster,
we've been talking about the cynics and World Lab integration for a long time.
It's actually not true.
They came into World Labs as a customer.
Really?
When we released the first version of our generated model
called Marble last winter around November, December.
Cynics just sign up.
No kidding.
Is it a customer?
Yes.
And I didn't even know what it was.
And then I realized this is Windrew's company.
I called Windrew.
I'm like, wow, this is your company.
And then we realized there's so much synergy.
Maybe Safi just quickly describe what Marble is.
Yeah, Marble is the code name for the base model
that World Lab is being used.
training and iterating on the fundamental capability right now of marble that is publicly released
is to take a prompt. It can be an image, it can be a few images, or a text, and turn that into a
geometrically consistent world that can be represented in 3D geometry, whether it's Gaussian
splat or mesh. Really what Cynix team is doing is trying to solve this extremely
difficult problem in robotics, which is the lack of data. The lack of data in training,
the lack of data in evaluation, this is very, very different from language models, where data
is abundant on the internet. And we know that in order for robotics to work, we have to somehow
unlock the power of scaling law. But where does that come from? This is something that,
it's a profound problem that everybody is battling with in robotics.
It would actually be great to talk about this energy.
You have put together a very, very talented team.
You have put together a very talented team.
And so to what extent is there overlap?
To what extent is this an extension?
Maybe talk a little bit about that.
Yeah, that's how complementary it is.
It's actually the TLDR is very complementary and with the shared mission.
So, Windr is one of the three technical co-founders.
The other two are Changsie-Zen, another Columbia professor who has been a world-class
technologist in simulation.
And Changsia has his background in also VFX.
He worked at Weta.
He worked at Tencent.
He's being an entrepreneur.
Then there's Sonny Hu, who is a phenomenal engineering leader,
who was also in a startup that was acquired by Amazon many years ago.
So he worked in many different tech stacks in the computer vision field in Amazon.
So when we started talking more seriously, I recognize that a couple of things,
that Cynix has from a talent point of view
is extremely complementary to world labs.
One is obviously Andrew's incredible thought leadership
and just technical prowess in robotics, right?
So from really, from hardware, full-stack robotics,
and even when he was by postdoc at Stanford,
at that time, you already had your faculty offer.
So you were there only for one year.
I wanted you for more than one year,
but he had to go become a happy,
the real job. So he was a full-stack researcher in robotics, from modeling to hardware.
And of course, Yunru and his student as cynics was that pool of talent world life hasn't had yet.
Then on the Changsci side is just an incredible simulation capability, right?
He's such a senior researcher and technologists in simulation.
What World Labs is doing is very much interfacing the world of simulations.
So I think what they don't have, obviously, is on the generative model side,
as well as the computer vision 3D reconstruction side, we're also very strong at World Labs.
So that's a technology that Cynics needs.
So together, these two sides come together and make it much more complete.
Faye's motivation in this is like this is an extent.
and a compliment to get into robotics.
Having been in your situation,
which is deciding when to sell a company,
it would be great to hear from you
and how you think about joining World Labs
and kind of the fit there
and why you made the decision to do it.
Yeah.
So at the very beginning,
we were deciding, okay, do you want to just keep going?
But after chatting with Fifei,
after seeing all the synergies that are happening in the middle,
it just makes perfect sense for the forces to join each other.
So in any sense, as cynics,
what we have been doing is real to seem to real.
is to do dense reconstruction of the environment.
So we capture the appearance of the environment,
geometry of the environment,
and also the dynamics of the environment,
meaning how the environment is going to change
when you apply actions.
So this dense reconstruction right now
is still a little bit on the heavier side.
And what World Labs right now has been doing
involves a lot of profound capabilities
around sparse reconstruction and generations.
So we see a lot of opportunities
of leveraging like a marble and other capabilities
as World Labs in order to do very efficient reconstructions
and modeling of the environment.
So can we expect a foundation model for robotics from World Labs?
World Labs is building a foundation model.
As you know, Martin, we're building a base model.
And as the technology has been evolving,
some of the most exciting base models are Omni models, right?
They take multimodal input, they have multimodal outputs.
And what is a foundation model for robotics?
It's very likely going to involve actions.
It's very likely going to involve the output of actions
in addition to the state of the world.
We're definitely not ruling this out.
Yeah, great.
So for example, for the foundation models,
it essentially needs to be a multimodal model.
So it has to take into account from text, image, depths,
and different kind of modalities.
And action is a very, very important part.
parts of that modalities.
So if you think about frame actions as the input,
that is essentially a forward similar.
That is going to predict how the environment is going to change
when you apply a specific action.
When the action is output,
this is essentially a policy model.
That is trying to predict giving a specific goal,
what should be the action you take in the real environment
to get you closer to that goal.
So this kind of only models actually can benefit a lot
and actually provide huge amount of values
for the robotics communities
in trying to understand how to make.
model the environments and at the same time
how to act in the environments.
And this can also act as a backbone
for you to find you into
specific robotic applications to making sure
it's really live up to the
reliability and efficiency that
expected by the clients.
You know, if you don't mind
kind of a lay investor question,
I see a lot of robotics companies
and a very popular approach right now
for the robotics companies that come in
is like we'll use a video model,
you know, and like,
you know, that's the predominant method
where this is, you know, 3D and simulation.
It's a very different approach.
And so maybe you could contrast
this popular approach of just using video only
versus kind of what the ambition here is.
Yeah. So in order to create words with the robot can learn,
the words, as I mentioned,
is to capture the essential structure of the problem.
And one of the very important and necessary requirements
for those words will be consistency.
So that is where I actually see
there's very, very strong synergies with marble,
because what we are building is a consistent world.
Consistence both over space, over time,
over different viewpoints,
and over different type of interactions.
And marble, the generated words from marble,
is also provide an infrastructure,
a component of that entire words
that we believe is necessary for the robot turner.
Imagine if a robot's pushing an object forwards,
the object just magically disappear,
which has been a problem of many of the existing
video prediction models.
It's one provides good enough signal for the robot to know what is the right thing to do.
But obviously right now there has been a lot of investigation on building better and better and stronger and stronger, like video models.
So we actually see a way where some of the infrastructure we build can provide as initial momentums.
And to go in through this data flywheel of going from this like a more simulation-driven models into like a robot policy models,
which is going to do the execution in the real environment, collecting new data, the data will,
come back in, where the model doesn't necessarily have to be physics only or learning only,
but somewhere in the middle, which be able to capture the essential structure of the problem,
but at the same time, be able to scale and become better and better as you accumulate more data.
You know, I've worked now Sefei very closely for a while, and you've always had this North Star,
which has driven this, and, you know, you've articulated variously as kind of 3D and in a number
of other ways.
And I'm just wondering for you, is there also a sense?
a more philosophical North Star or you're more the pragmatic like I am.
I build the system, like do the thing.
My North Star is to make robots work.
Amazing, yeah.
In the real environment.
I'm a very practical person.
I want the robot to work.
One interesting thing that's actually coming from my collaborations with Phoebe during
a postdoc, we are building this kind of benchmark.
We actually send out surveys asking the general public,
what do you want the robots to do for them.
Among the southern tasks we collected, one third of the tasks,
are about cleaning.
People just don't like to do those, like,
dial and dirty tasks.
And those are the scenarios.
We really want to make sure we have robotic solutions
to deal with.
One thing I really like about cynics,
Martin, especially continuing your question,
there's a lot of robotics companies
building models and all that.
One thing I truly like about cynics
is Rindrew and his co-founders
have such an incredibly pragmatic approach
to robotics.
They, especially they come from academia, right?
Sunny doesn't, but in Zhu and Chanxi come from academia,
but their first instinct is work with design partners and customers
in real industry, whether it's labs at industry labs
or warehouses or electronics assembly.
Electronics assembly.
That is such a refreshing, actually,
a refreshing way of,
for approaching robotics.
And that really true made me very excited to work with them.
This is for you, and you?
But I'll just be this is personal curiosity,
which is it seems to me that for robotics,
you have to be pretty exact.
I mean, not perfect, but pretty close.
But for the creative use cases, which Boralep's
done a lot of, you kind of don't need to because, you know,
I mean, even sometimes like being wrong is stylistic
or intentional or whatever.
And so from a technical perspective,
what is the challenge here for reconciles?
these two things, or do they never get reconciled?
Like, there will always be two points in the design space.
So they will be reconciled in the long terms, of course.
And modeling of the environments doesn't have to be perfect.
The model doesn't have to be perfect in robotics.
And by the way, again, this is pure curiosity,
but is there like a bit more formal way to say that?
Like, what does that mean not to be perfect?
It has to be pretty close.
So let me put in this way.
For example, models over the developments of all different kind of robotic applications
has been a very important cornerstone.
If you look at all the existing robotic applications, like Plain, Jones, Rumba, or even
for quadrapad robots, bipedal robots, model has been the way for them to actually work
and be able to transfer from simulation to the real environments.
But if you look at those locomotion robots, like quadruped robots, bipedal robots, they
can walking on snows, they can walking on bushes, but you don't need to have a simulator.
They can simulate all the bushes and snow.
like very precise.
You need to have a simulation
that captures the essential structure of the problem
and do a whole different kind of randomizations
inside the digital environments.
So that is what we're aiming for.
So basically with synics and together with word labs,
we're trying to investigate what is the level of fidelity
we need to model the messy and mess of worlds besides the robots.
As I said, we'll be able to transfer the robotic systems
training the simulated environment and digital worlds
back into the real scenarios.
As an investor, I've heard,
other researchers, say like Sergey Levine, say simulation will always eventually deviate from the physical world
and real world data collection is absolutely critical.
And so maybe talk a little bit about like the viability of this approach where simulation is a cornerstone
as opposed to some other approach.
So they don't contradict with each other.
So if you think about the simulation, simulation essentially trying to predict how the environment is going to change when you are.
the actions.
And this is essentially a model of the world
that doesn't necessarily have to be pure physics.
It can be a combination between both physics and also learning.
We are collecting real-world data.
We will be using those real-world data.
It's just at different stages of this, like, a data fly wheel.
Maybe at the very beginning,
we have stronger emphasized on we have more physics
to making sure we have the right consistency
and right structure for us to learn the world,
for us to train the robot policies.
But as we accumulate more and more data,
both through data collection and also through the collaboration with our clients,
we'll have the data that will be moving towards more towards more learning-based,
like modeling of the environments.
So this kind of transition and also this kind of data file is really enabling factors
of both getting the best of both physics and the geometry and consistency,
as well as all the power and magics from the data and computers.
I want to add to this and be slightly philosophical here,
is there isn't a binary choice between simulation or no simulation.
All this come in together to make robotics work.
Think about human intelligence.
We do a lot of simulation in our head.
You know, why?
There's a very important role simulation plays that real-world data doesn't play,
which is counterfactual reasoning,
is that you play out events,
that hasn't happened or cannot happen,
or you don't have enough data to make it happen in real world.
And while you play it out, you learn how to act in it.
Humans do this all the time.
We probably don't, you know, we just,
I know you were at World Cups.
I was at the World Cup.
Congratulations to Spain winning.
I'm sure in the planning of every game,
there is simulation, whether it's digital or on the whiteboard or whatever,
that simulation,
the role simulation plays is counterfactual reasoning.
And that's really important in robotics
because we just do not have, cannot possibly have,
enough real-world data for that.
Here's a real-life example,
the industry of self-driving cars.
Waymo has officially said they use billions of hours of simulation.
And actually, Waymo is more simulation-heavy
than just real-world data heavy.
So these are real examples.
And as you know, Martin and Andrew, too, cars are the simplest kind of robots.
Yeah, 2D, yeah.
Yeah.
So clearly, simulation plays a huge role in robotic learning.
I also want to add to that.
So, like, there are, if you put things more specific, simulation can provide two levels of benefits.
The first one is reliability, and the second one is efficiency.
So for reliability, if you're thinking about a robotic system working with a reliability,
in the real environment.
You need data to provide systematic coverage
of all the state space and variations that robot mining control.
That's how you can learn of how that is robust.
So with simulation, you can do systematic randomizations and control
and variations of lighting, frictions, geometries, object types,
and also all different kind of physical parameters
to making sure you have sufficient coverage of the state space.
So this is what can give the robotic systems reliability.
And second is about efficient.
So right now, many people are doing teleoperation.
And if you look at many of the teleoperation device,
imagining all the actual skeletons you are using,
you're actually collecting the data at a speed that is actually slower
than human actually doing the task.
But for many of our clients, human speed to them is not good enough.
They want faster than human speeds.
So for the robot to move faster,
it's not as simple as just drives the robot faster,
because the gravity doesn't change.
But in simulation, you can do systematic speed up of the brain.
of the robots' behaviors to train the robots such that it considers all the dynamics,
changes of the environments.
So this is what can give like our clients for them efficiency.
So both for the reliability and efficiency, though there are some kind of like a very unique
like values where simulation can provide.
You've talked about the technology and the platform, what it does.
Let me talk about the specific use cases people use it for.
There are essential like two specific use cases, especially around both training and also around
evaluations.
Starting from the evaluations.
So evaluation is something like people
often overlooked in
the robotics. But if you
are tuning like a robotic models, you have
to know how well it works and that is the
only source of information for you to iterate.
By the way, every
AI person really understands what evals are
and uses it all the time. Non-AI people,
it often means something a little different. So maybe it's even
worth just describing specifically what
you mean by evaluation. Okay. So
So what I mean by evaluation is we'll be able to understand for this specific checkpoints,
how will does it perform?
Does it perform, for example, 95% of the time or 99.9% of the time?
And the key criteria people use in industry is, how long does it take?
How long in work-clock time does it take for you to distinguish between a checkpoints that is 90%
from a checkpoint that is 92 points?
And if you only do that in the real environment, that's just take so long.
for you to do the distinguishments.
And if you really think about also the robotic evaluations,
right now people are doing in the real environments,
the iteration speeds is multiple orders of magnitude slower
than iterations of those language models.
So not only is like the robotic tasks very varied, very diverse.
Oh, you actually have to do the thing.
Yeah, yeah, yeah, right.
Like atoms have to move through space.
Yes, exactly.
But only there are the laws of physics have to be obeyed.
watch those robotics videos, every video has like 10x, 8x,
because it moves so slowly.
Exactly. So not only is slow, it's dangerous, it's costly,
but at the same time, the speed is also like multiple hours of magnitude is like slower.
So some of our clients actually needs this digital environment
that can be used to evaluate their robotic like systems.
And because our digital environment has proven alignments with the real words,
So meaning whatever happens in the theme
is also likely to happen in the real environment.
If a checkpoint is working better in the simulation,
it's also highly likely to also work better
in the real environments,
as we have also been discussed in the blog post.
So that actually gives our clients very strong confidence
in actually using the data,
using the signal from the digital environment
to do scalable, safe,
and much faster evaluations of their robotic systems.
Great.
So that is on the evaluation.
Then on the training.
So on the training side,
So basically, like I also mentioned, it's about controlability.
So you want to control all the different possible variations of states, parameters,
lighting, frictions, physical parameters, like even object geometry, object types.
So you want to make sure you have sufficient coverage of all different kinds of scenarios,
such as we will be able to generate, like, an informative data for your robots to be robust.
And this is just going to be so hard to do just in the real environment.
Like we discussed, if you do title operation, the speed at which you are collecting data,
is slow, you're also limited by
how many robots you have,
how many tele-operation device you have.
There's like a whole different kind of
like challenges around all the
data operations around it. But in simulation,
everything can be controllable,
everything can be systematic,
and everything can be understood at a level
where you know exactly
and making claims about exactly what distribution
you have covered.
To develop confidence about
within the distribution, we know the robot will work.
So those kind of confidence
and efficiency and scalability
is something that our clients also value
to use our digital words
for the training of robotic systems.
Here's a crazy thing.
Even before Cynics and we are talking,
our inbound customers for Marble
were already seeing this kind of demands.
We just cannot serve these customers,
but we are already getting a lot of phone calls
from robotics,
early stage robotics companies
who are developing their models
all the way to downstream,
very pragmatic use cases and we're seeing these needs.
When people hear you're going into robotics,
what they're going to envision is you're pulling out a 3D printer
and you're going to be making hardware
and then you're going to be programming the brain of a robot
and sticking it in the robot and then you've got a robot.
And I don't think that's what you guys are talking about here.
So maybe talk about where this fits in the life cycle of creating a robot
and where you will end
and where the rest of the ecosystem will begin.
So what we have been building,
you can imagine is a infrastructure
like with the software around these infrastructures
for people to, for them, build words,
such that robot can learn and evaluate.
And these infrastructures is naturally model agnostic
and embodiments agnostic.
So I just want to be very clear,
just because this is actually a very subtle,
I mean, for you it's obvious,
but it's a very subtle point,
which is, from what you said,
that's not building a robot.
It's building an environment
which another company can place their robot brain
to navigate and to learn.
Yeah.
So for our customers right now,
they have all different kinds of robots.
Some are using single robot arms.
Some are using bimail.
Some are using a fixed arm.
Some are using mobile manipulators.
Some are using grippers.
Some are using some more elaborate versions
of the only factors.
So our platform right now
is just naturally embodiment agnostic.
We can very easily integrate different kind of robotic embaliments,
be able to put them into the words we generated, we digitalized,
such as we will be able to give those individual robots capabilities
of doing the right tasks and at the right levels of reliability and efficiency in the real environments.
And we are also, for example, model agnostic.
So we can just using the data generated by our words to train different models,
either from scratch or doing post-training of existing foundation models,
like vision language action models or word action models.
So to us, it doesn't matter.
We just want to making sure we have the infrastructure,
we have all the words such as the robot can work reliably in the real environment.
You know, you have told me that you think a lot of the predictions around humanoids
were a little bit aggressive and were likely to see more constrained rollouts like warehouses or whatever.
Can you talk a little bit about that and how that impacts what you're going to be tackling here at
like World Labs.
So that's a very good question.
So if you look at, for example,
all the progressions
of robotic applications in the real environments,
it has always followed the trend
going from fully structured environments
into semi-structured environments
and then into unstructured environments.
For fully structured environments,
what do we mean is that you have knowledge
and control over all the
configurations within the environments.
Like factories.
Like factories or, for example, car main factory in lines.
Those has been automated
for decades.
Yeah, yeah, yeah.
And then you have, for example,
semi-structural environments,
which you have certain controls over the environments,
for example, like the Amazon warehouses,
or, for example, like restaurants, hotels,
where you have certain control over the environment
to just make the task easier for your robots.
But there are obviously many other, like, objects.
Or, for example, clothes,
those are the objects.
You don't have control.
And then for the unstructured environments,
it's like your home and my house.
Those is, I would say,
the grand chat.
Especially my house, trust me.
Three dollars, five-year-old.
Yes.
Dogs.
Exactly.
If you're thinking about where does the robustness coming from?
Robustness coming from a sufficient coverage of the scenarios and robots might encounter.
So it's so much easier and more approachable at least like right now to focus more on the semi-structed environments before we move on to fully unstructured environments.
So we will move into that direction.
It's just we want to take a more sustainable and more realistic approach.
towards there.
I think your point here is that humanoid mimics human body.
And evolution has optimized human body for unstructured environment.
And so our fingers, our legs are not the best apparatus to do one thing.
For example, if our only goal as a species is to climb trees, we will not have this body
necessarily, right? So we'll have different kind of fingers. But what humans end up having
are evolved, evolved into is this body shape that can be very general, but not necessarily best
at everything. And that is for the survival of unstructured environment. But from a business
point of view, from a pragmatic technology point of view, that this unstructured environment
and a generalized body is actually the hardest problem to solve.
It's not necessarily even the right way to solve the problem.
It's we specialize, so we take more specialized body to solve a narrower problem.
But the challenge for cynics is that to be more body agnostic
so that their infrastructure can serve different bodies
and different semi-structured environments.
A common lens to look at exactly this question is an economic lens, right?
Which is, you compare it to like generative LLMs,
they can create pros or a code 10,000 times faster than a human being,
a bunch cheaper than a human being.
So the economic case makes sense because our brains aren't very efficient at that.
However, our brains and our bodies are very efficient at 3D navigation, right?
You know, like movies of the world are picking things up.
And so this is just a prediction question,
but do you believe we'll ever be able to build robots,
at least in the foreseeable future,
that have the power efficiency of a human being
when it comes to menial tasks?
So let's say just basically, you know, minimum wage or something like that.
Like how far away are we from this?
Is this like five years or this is like never?
I think it's going to take a very long time.
So if you really think about like robots in the real environments,
in the end, it will.
it will always be a system.
So every working robot in the real environment is a system work.
It's new need to be very mindful and thoughtful about how the system are coming together.
The hardware, the software, the brain, even like to the details of, for example,
what's the friction coefficients of your fingers?
So there's a lot of things you have to consider to make these things a reality.
And it will take iterations.
But what I am excited about is that I have always been at the state of the arts of robot
learning and also trying to push the state of the art forward.
But the state of the arts always moving faster than I expected.
So what I'm focusing on and trying to investigate right now is very different from, for example, when I started my PhD.
So this is a speak to how fast the whole ecosystem has been evolving and all the moving pieces start coming together or building these robotic systems.
But we also have to be calibrated about our predictions.
So we will see a lot of progress.
But to achieve, for example, human level efficiency and components,
it will take longer.
Martin, the hardest thing in today's AI is to have the right measured optimism.
Right.
It's totally true.
Yeah.
I mean, even LLMs does not have human brain efficiency.
Human brain operates on 30 watts.
Yeah, that's true.
So we are far from that.
But I mean, performance to power it may be close, right?
In narrow tasks, like software engineering.
Like generating and image or software engineering, then it is, right?
Yeah, I think so.
I don't think we're anywhere close on the cost of robotics.
This has changed how you think about your, like, strategically the level of ambition that your team can go after.
I mean, does it change that, or is it still very much in line what you expected to do when you started?
It definitely changed, though, to jackfries in a very profound manners.
So we see a lot of unlock in be able to do this whole process, do the modeling of the environment,
you are much more efficient and much more scalable manners,
especially in partner together with Word Labs.
And I also want to add to Fayevei,
if you think about, for example, the current states of the language models.
So those are models that's with incredible capabilities.
But still, you don't just blank trust it to book your flight tickets
or make your hotel reservations.
You still, hopefully there's still a person who's reading the output from those language models.
But that is very different from how people and will be using,
for them robotic models.
Based on robotic models,
out of the box,
the robots has to work reliably
in the real environment.
So, and we don't even have the data.
We don't even have all the necessary infrastructures
around those for the robots to just out of the box
work reliable in the real environments.
So for that reasons,
be able to create this digital words,
this scalable digital words where the robot can learn
and evaluate within.
Yes, it's going to unlock so much more potentials
for being able to replace all the,
like a costly,
and the unsafe data in the real environments
with the data generated from the words
for the robots to be able to do scalable learning and evaluations.
You know, I've seen many of these kind of integrations.
They actually work very well at this stage
when they have this much alignment, which is great.
But there's always like this question of,
do you integrate now into what's happening now
or do you keep things quite separate
and provide kind of like a long-term trajectory
that will, you know, be realized, you know,
in the year timeframe.
How are you thinking about this,
Fei, Fei.
Is this something that integrates right away,
or is this kind of a separate longer term?
This is a great question.
I think at this point,
you know,
Yun-Ju, Changxi, Sunny, Justin Ben,
and I have been talking about this.
At this point,
we are going to take it thoughtfully.
We're not rushing to integrate
everything from code base to teams
because I think Cynix does have a very,
well-thought, and I wouldn't call it,
a standalone completely, but fairly
contained tech stack as well as
their customers, as well as the kind of
products they're building. We're going to take time.
We definitely will, we're already on the simulation side, as well as
the potential base model, action condition model side.
We already are starting to talk. And also, they
are using Marble as an internal customer.
So we will be integrating, but we're not rushing to blend the team as like a full salad bowl.
How are you thinking about geography so this with Cynx move?
Is it going to stay in the same place?
We're going to, Vindra is going to move.
Oh, well, welcome.
Yeah.
Florence to the Renaissance.
Perfect.
I think we, our labs is officially becoming a bi-coastal company,
where the headquarters is in San Francisco.
I live in Palo Alto.
I feel like I'm in a different state.
But I'm actually excited that we're going to have an office in New York
that can help us to attract talent on East Coast.
And also we have been talking about making sure
that in both offices we set up the robots
so that we get to basically test us.
and mature our engineering stack
so that we can work with robots remotely
because we have to do that for our customers anyway.
So maybe just to be very concrete, Faefe,
maybe let's just pencil out.
What is the perfect success case in two years?
Like what product do you have,
who's engaging with it, how do they use it?
Just the crisp, like what this becomes.
I would be very happy that Cynix team,
a world app's team,
will have valid.
validated customers in a small number of important vertical use cases,
where our system, our infrastructure has proven to be truly beneficial to their
automation needs. And these customers became our lighthouse examples to scale our business.
How early, let's say someone listening to this is running a robotics company,
at what stage should they engage with World Labs?
Is it really early on?
Is it somewhere in the middle?
So right now for our customers, because we are building this kind of real-to-themed real pipelines,
where the simulation is essentially the words,
we're going to provide the training and evaluation grounds.
Some customers, they need only the real-to-thin part.
They want to digitalize the task they care about and be able to do the evaluations of their robotic systems.
Some customers need this real-to-same to this entire pipeline,
such as they will be able to have policies like running on their hardwors.
So our platform is also designed in a way that is flexible, depending on what our clients needs.
And at the same time, the clients who are working with are actually pretty close to the deployments, like a stage.
So basically, they are working on very, very practical tasks.
Those tasks, when we have robotic solutions that are there,
can just create value immediately.
And they have at least like tens or hundreds of like this kind of situations.
They are thinking about to do the automations for.
So as like together with Word Labs, we'll be able to develop reliable solutions for those scenarios.
As we have already shows, we have a number of scenarios already instantiated in our blog post.
And we'll be able to further our investigation to see how they can actually
solve the key requirements and also constraints faced by the real-world deployment.
Great. I want to be very specific about this. Is it ever too late or too early to call
World Labs if you're a robotics company? No. We want everybody to call us. We want to learn about
your use case. Wonderful. If you're listening to this and you're anywhere close to a robotics
project or robotics company, please track World Labs. Yes. Thank you. Definitely open for business.
Yeah, we are open for business.
Not too early.
All right.
If you're doing robotics,
called World Labs.
Thank you both very much for coming.
Thank you.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode,
be sure to like, comment, subscribe,
leave us a rating or review
and share it with your friends and family.
For more episodes,
go to YouTube, Apple Podcast, and Spotify.
Follow us on X at A16Z
and subscribe to our substack at A16Z.
Substack.com.
Thanks again for listening and I'll see you in the next episode.
As a reminder, the content here is for informational purposes only.
Should not be taken as legal business, tax, or investment advice,
or be used to evaluate any investment or security
and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments
in the companies discussed in this podcast.
For more details, including a link to our investments,
please see A16Z.com forward slash disclosures.
closures.
