Y Combinator Startup Podcast - Waymo Co-CEO Dmitri Dolgov: "Move Fast And Ship Safely"

Episode Date: August 4, 2026

Waymo’s first autonomous demo took eighteen months. The product took fifteen years. Today, the Waymo Driver runs 500,000 trips a week — four million fully autonomous miles across fifteen cities, w...ith 17 times fewer serious-injury crashes than human drivers.At Startup School 2026, Waymo co-CEO Dmitri Dolgov shares the seven lessons behind that journey, from bridging the gap between a demo and a real product to building systems that can safely operate in the physical world.Transcript: https://www.ycrootaccess.com/p/dmitri-dolgov-seven-lessons-from

Transcript
Discussion (0)
Starting point is 00:00:06 Good afternoon, everyone. It's great to be here. Now, we talk a lot about AI that lives on your screen, lives in the digital world. And today I'd like to talk to you about a different kind of AI that we've been building at Waymo, AI that lives in the real physical world. How many of you, by the way, have been in a Waymo? Just raise your arms. Wow. Okay. That is a... impressive, especially I understand many of you are out of town. The folks who are visiting and have not had a chance to check out Waymo, I hope while you're
Starting point is 00:00:47 here in the Bay Area, give it a try. So this being a startup school, I structured this presentation as a sequence of lessons, seven lessons that we've learned over the years at Waymo around what it takes to build and safely ship today's most important. mature application of AI in the physical world, the Waymo driver. Let me start with a short video. This is a clip from a ride that I recently took in a Waymo with my kids. So as you see here, we're moving forward or proceeding through an intersection,
Starting point is 00:01:31 and a couple of human drivers just decide to cut in right in front of us. And the Waymo driver reacted safely, reacted smoothly, in fact so much so that the kids, my kids were preoccupied in the back seat, they didn't even notice that anything happened. And to me, this was a pretty powerful moment. I've been working on this technology and this product for close to two decades, and, you know, it just did something fairly important. It acted safely. It kept my kids safe. It kept everybody safe. And nobody noticed. And that I think will be a bit of a theme in general when it comes to physical AI, that the best AI moments will look like nothing happened. It's just the task got done safely and smoothly.
Starting point is 00:02:21 And these sort of moments where the Waymo driver kept everyone safe are happening daily across our fleet. Today, the Waymo driver is serving around 500 trips per week and driving over 4 million fully autonomous miles every week in 15 cities across the United States. For a comparison, that's over 300 years every week of an average American driver per year. And the Waymo driver is accomplishing that with a superhuman safety record. So what does it take to build and deploy an AI agent in the physical world at scale? Now in Silicon Valley, there's a common mantra to move fast and break things.
Starting point is 00:03:15 However, when you're dealing with atoms instead of bits, breaking things is not really okay. So the thing you have to do is to move fast and ship safely. And that's a much more difficult thing to do. You have to build systems that are robust from day one. You have to build AI models and you have to build training recipes where safety is the foundation and not an afterthought, not an add-on. And by the way, the problem itself, a physical AI is different from digital AI. There are four main gaps that you have to contend with if you're building AI for the physical
Starting point is 00:03:56 world versus the digital world. First, there is the cost of error gaps. You have a language model or a chatbot or a co-pilot and it makes a mistake. Usually it costs you a retry. In the physical world, the cost of a mistake can be measured in human lives, not tokens. There's simply not an undo and a retry button. Secondly, you have the latency gap. And typically when you're running a VLM or a digital
Starting point is 00:04:30 assistant, it can take many seconds, sometimes minutes to come back with an answer to you. A car, traveling at freeway speeds, moves about 100 feet in one second. So there milliseconds really matter, and you have to run all of your inference, make all of your decisions on board a compute that fits in a trunk of your car. Next, there's the data gap. Digital AI had the internet. of this wonderful immense cache of pre-labeled human knowledge and human thought that we've ever assembled. There's no digitized version of the internet for the physical world. And lastly, there's the validation gap. In digital AI, often you can ship something that's good enough.
Starting point is 00:05:26 And then you'll let your user product, they find the edge cases, and that allows you to deploy on day one practically at unlimited scale. And then you can just iterate on hill climate quality from there. In physical AI, the situation is different. Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your first autonomous mile. Now, at the same time, when you're dealing with physical AI,
Starting point is 00:06:04 the actual experience of having your agent in the real world is invaluable and it's irreplaceable. These systems are not just something that you can build in the lab, get it perfect, and then deploy at full scale overnight. So given those two factors, you really need to super clearly and super crisp define the operating conditions and the deployment parameters of your agent, and then build a rigorous framework to guide your deployments so that you can scale in a responsible manner. And this is absolutely critical. This is how you earn trust from your customers, from the communities, from the regulators, and yourself. So at Wayne, we see these gaps, of course, in the context of
Starting point is 00:06:55 of autonomous vehicles, but these gaps will show up in practically any sort of non-trivial physical agent that we will deploy in some shape or form. And driving is simply the first domain where AI has crossed these four gaps at scale with the public interacting with our product. So let's dive into those lessons that we've learned over the years at Waymo from working on this problem
Starting point is 00:07:20 and talk about how we address those gaps. I have seven lessons in this talk they're all technical. There's a lot more that goes into building a company and building a product, but today I'll just focus on the technical aspects of building AI for the physical world. And each one of those lessons, I think by itself,
Starting point is 00:07:40 will not be exactly earth-shattering. A lot of it will overlap with likely things you've heard elsewhere. But I hope that the grounding of these lessons in our experience and some of the nuance that I can add about how they showed up in our experience of deploying a physical agent in the, and scaling it safely will be interesting and useful for many of you who are in the space as you build your product,
Starting point is 00:08:05 as you build their startup. So let's dive in. The first lesson has to do with this massive, frustrating, sometimes soul-crushing difference between a demo and a real product. And a working demo is 1% at best of the work that you have to do. The many nines of performance,
Starting point is 00:08:27 the many nines of reliability that followed, that's where the real work happens. And if you're a founder in the room, chances are you are focused on getting that first prototype, that first demo off the ground. And when you hit that first version of a system that works, that first 90%, when the demo actually works, it feels incredible.
Starting point is 00:08:47 You feel like you solved it, the sky is the limit, you're extrapolating forward. And in our world, we hit that first milestone, at first 90% back around 2010. So when this project started in 2009, before we started building the system, we set a couple of pretty ambitious goals for ourselves. One was to drive 100,000 miles in autonomous mode.
Starting point is 00:09:17 The second goal was to drive 10 routes. Each one was 100 miles long, chosen to cover a variety of conditions across the Bay Area, and we had to do each one from beginning to end without a human intervention. We had at the time a team of about a dozen engineers, and we accomplished both of these goals in about a year and a half. And keep in mind, this was well before any of the AI breakthroughs, before Convness, before Transformers, before BLMs, before any of the stuff that we talk about today. And yet, you know, we got it done. And kind of by demo standards, we driving, autonomous driving was solved
Starting point is 00:09:54 in 2010. We handled everything. We handled, we could drive during the day, during the night. We handled traffic, pedestrians, cyclists, traffic lights, construction zones on freeways, on surface streets. So we were quote unquote capability complete. And at the time, we felt like we're on top of the world. But then we quickly ran as we started building towards a product. We quickly ran into a brutal reality that there's a massive difference between doing something once or driving 10 routes once and building a scalable service with nobody behind the wheel. It took us about 10 more years to begin providing a service, and then five more years to scale to half a million trips per week.
Starting point is 00:10:40 So the demo took 18 months, the product took about 15 years. But now we're scaling exponentially. To date, we've served well over 20 million fully autonomous trips, and we've driven well over 200 million fully autonomous miles. And we have rider-only vehicles operating in 15 cities across the United States. And we're scaling exponentially. It took us 15 years to get to that first 100 million miles and about seven months to drive the next 100 million.
Starting point is 00:11:13 It took us about eight years to go from the time when we started our initial rider-only operation to the time when we were serving riders in four cities. Earlier this year, we launched four cities in just one day. So why does bridging that gap from demo to product takes along? Because there's this harsh engineering reality that you can't really cheat, that reliability and performance lives on this exponential letter of nines. So getting to that first 90%, or 99%, that's the easy part.
Starting point is 00:11:52 But then every next nine that you want to add, that takes about 10 times more effort. So you need to know up front exactly how many nines your product actually needs. So demo might need 1-9, an assist product or co-pilot might need a few, but a fully autonomous AI agent that we're going to be putting out in the physical world that engages with the public and with kids running around, that needs a whole stack of them. And at scale, the long tail is the problem space, is your entire problem statement. When you drive millions of miles per week,
Starting point is 00:12:31 a rare event that might happen once in a million miles, that just becomes your daily reality. And getting those next nines means doing something different every time. So you don't get to say six nines of performance or reliability by doing the same thing that you did for, you know, to achieve the first two, but longer. You have to do fundamentally different things.
Starting point is 00:12:54 It requires a fundamentally different approach. For example, we can take reliability. You can get to the first couple of nines by just doing proper engineering and doing some bug fixes. But to get to the next few, you need to invest in fundamentally different approaches. You need to build fully redundant systems, have tiered fallback architectures and so forth and so on.
Starting point is 00:13:18 And the same thing holds for the performance of AI models. So what that actually means is that in this space, it's incredibly easy to get started, but it can be excruciatingly difficult to get to the real product. And that effect is only amplified with every wave of technological breakthroughs. And that naturally leads to hype cycles. So every AI breakthrough from deep learning to ConvNets to Transformers, BLAMs, you name it, it makes it that much easier to get started. Your demos, your prototypes, they get 100 times easier. But the tail, that's where the hard problems are,
Starting point is 00:13:59 that moves much less. It moves, but the effect is muted. And that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products. And the recurring mistake of every cycle is spending on the demo what you should be saving for the nines. Now, this being a startup school, the last thing I want to do is throw too much cold water on the
Starting point is 00:14:21 magic and the excitement of those early days. This time is absolutely magical. It's amazing. Cherish it, leverage it. But the key is to remain honest about the product that you're building, the number of nines and performance and reliability that that product demands, and not cutting corners to get there. Otherwise, you might be in for a pretty rude awakening later.
Starting point is 00:14:43 So count your nines before you count your demo views. And this brings us to the second lesson. Once you know how many nines your product actually needs, it fundamentally dictates the architecture and the core technical approach that you need to pursue. Now, every technology has a performance versus effort curve, right? They all tend to start fairly steep and go up, and then they flatten out. And as I just mentioned, every other nine gets an order of magnitude more difficult.
Starting point is 00:15:16 So common failure mode is picking the tech that gives you the fastest early ramp, riding that steep curve, feeling like you're winning, projecting that, you know, steep slope into the future and feeling like the sky is the limit, and then hitting the plateau and discovering that the technology path that you picked actually flattens out way before the performance that is required by your product. Now you might still choose to be, at least for a while, on that steep curve for a variety of practical reasons. Maybe you want to prototype, something or demo something or build something in service of learning, but be honest with yourself where you're building for the purpose of a demo, for the purpose of learning, or towards
Starting point is 00:16:01 an actual product. So let's take an example from our domain, autonomous vehicle and sensing. There's been a long-standing debate about what kind of sensors do you actually need for autonomous driving? Naturally, more sensors means higher performance, but also means high complexity. So humans Of course, you can drive with just eyes, so there's that proof of existence. Now, and if the goal were to just approximately match human performance or to build an assist product, that's a very reasonable way to go. However, if you are targeting full autonomy
Starting point is 00:16:39 and you're targeting superhuman, strongly superhuman performance, you find that weak sensing just leads to a safety curve that flattens out way too early. So at Waymo, we've taken an approach where we use multiple sensing modalities. sensing modalities. We use cameras, lighters, and radars, and they all complement each other. Cameras give you high resolution and color, but they're passive, and they degrade in darkness and glare. LiDar gives you a direct measurement of the 3D structure of the world around you, and radar is very good at punching through environmental conditions and weather, like fog or rain or snow, and it directly can measure velocity using Doppler.
Starting point is 00:17:20 Lighter and radar are active sensors, so that means they see just as well in pitch darkness or, for example, when driving into a blinding sunset. And these different sensing modalities, of course, they're not backups to each other. In our stack, each modality has an encoder and the information from all of those sensors get fused into a single view of the world around us that is much more precise and generally vastly superior to what you get with any one sensor. So let me show you a few examples. Here's the scene, our Waymo is driving in a dust storm in Phoenix.
Starting point is 00:18:01 So what you see here is what the scene looks like to our fairly advanced high-resolution and high-dynamic range camera. It's very close to what a human would see in the same conditions, which is not much. And here on the right is what the lighter sees for the exact same frame. And you can much more clearly see that there's a pedestrian standing on the side of the road. So if they were to step onto the road, that early detection can make a really big difference in how the situation plays out and the safety of everyone involved. Here's another example at night, driving along, and there are a couple of pedestrians who are about to jump onto the road over a concrete construction barrier. Again, at the bottom, you see the camera, really can't see much, and the lighter view at the top.
Starting point is 00:18:49 Again, lighter versus camera. Here's another example. There are a couple of dogs chasing a bowl and a couple of kids chasing the dogs. And big difference. Here's what it looks like to the camera. Here's the lighter and the early detection of the kids is off to the side and there are no headlights. There are no lamps there. It's complete darkness.
Starting point is 00:19:15 So it makes a big difference. Or think about what happens when something physically obstructs the view of your sensors. If you don't have redundancy in sensing, you can have a single leaf land on your sensors and bring your robot to a full stop. So you need redundancy. Redundancy, of course, does not necessarily mean multiple sensing modalities, but if you need redundancy anyway, you might as well benefit from the complementary physics of the different sensing modalities in the nominal case.
Starting point is 00:19:52 So here's a video of one of our cars that picked up a leaf or actually, I think, a full branch of a tree, that our wipers were unable to shake, and the car detected that. And because we have sensing redundancy, it safely was able to get back to the depot for proper cleaning. So specifically, when it comes to hardware, do not anchor to today. components prices. We're on the sixth generation of the Waymo driver, the Waymo Hardware Suite today, and with every generation the hardware not only delivered amazing capability but we're able to drastically simplify and radically reduce the cost of the hardware as well. So betting your company, betting your approach
Starting point is 00:20:43 on today's hardware prices is just betting your company on a number that has a fairly short shelf life and is going to expire. So hardware would change. Many components will get commoditized and drop in price. So design for that future and be ready to upgrade. And then brings us to the next lesson, lesson number three. Technology moves incredibly fast, especially nowadays. So you need to be ready to ride those tech waves and do that repeatedly.
Starting point is 00:21:17 And when you do, have to not only think about the wins and performance and the wins and performance, and the wins and capability, you have to be very mindful about unification and simplification. Over the years, we've seen a number of major breakthroughs and technology, a lot of them around AI, and with every wave of innovation, we pretty much rebuild the Waymo driver
Starting point is 00:21:41 around that major wave of AI breakthroughs. And we often push the state of the art in those areas forward ourselves. We leveraged ConvNs around 2013 for computer vision and perception. Then when Transformers came about around 2017, we bet big on them for perception and for the task of behavior prediction and decision making and planning. Turns out the task of driving is not that dissimilar from the task of modeling language.
Starting point is 00:22:10 Because of the social aspects of driving, you're kind of having a conversation with other dynamic actors in the world, but you're doing that in the space, kind of body language of your agent, your car, as opposed to just the language of words. And you operate in sequences and local continuity matters, but so does global context. And today we're leveraging the latest in VLMs and Frontier World Models.
Starting point is 00:22:34 Now, using the latest tech for capability and performance wins, I don't want to say it's easy, but it can be reasonably straightforward doing applied research in isolation or starting a Tiger team to prototype some new technology is not the most difficult part. There's many companies, many teams that are excellent in this. The much harder muscle to build is to carry that bleeding edge research into production and deploy it in a safety critical environment without regressions. And do it without breaking stride on the scaling of your product. And adding capability, again, is not the hardest part, but adding capability while at the same time reducing fragmentation and reducing complexity, that is really important. And finally, the hard muscle to build as a company is to be able to do that repeatedly
Starting point is 00:23:28 through multiple waves of technical innovation or technical breakthroughs. So on this front, I have two bits of advice. The first one, when the technology, a new technology shows up, you know, it can be very exciting, very tempting to kick off a new effort, a tiger team to pursue it. And that's great. You should absolutely do that. However, when you do, it's very important that you consider what you would do after. Under a success scenario, let's say that effort succeeds, you should be very clear on what the path of that new innovation is for your company, for your entire product, for your entire system.
Starting point is 00:24:04 Oftentimes I've seen a failure mode where a project, a very difficult technical project succeeds, and then there's a dead end. That can be very wasteful, that can be completely deflating. The second bit of advice I have here is when pursuing new tech, again, don't just ask what does this new tech give me in terms of capability and performance. Also ask, has it simplified my stack? And has it led to fragmentation or unification? So set your launch bar to demand both breakthrough performance and at the same time radical simplification and unification. And this exact philosophy and this muscle that we've built at Waymo over the years is what produced our latest. core technology. And the heart of it is the Waymo Foundation model. Now the Waymo
Starting point is 00:24:53 Foundation model is a multimodal world action language model. It's kind of a mouthful, so let me unpack the ingredients. It's a multimodal model because it is able to process these multi-modal sensor inputs, cameras, Liders, and radar. It's a world model because it inherently understands how the world works. The physics, the dynamics, as well as the social and semantic aspect of it. It's an action model because we are not just passively observing how the world evolves. We're an active participant. So the model needs to understand the effects of our actions of our agent on the world and be able to tell the good ones from bad ones. And finally, it's aligned with language and that allows us to unlock general world knowledge from visual language models.
Starting point is 00:25:43 and that's incredibly useful in the long tail of rare semantic situations. So more specifically, this is what the architecture looks like. It's kind of your typical encoder-decoder architecture. The encoder part takes in the multimodal sensing and compresses it or encodes it into an efficient representation that retains all of the relevant data, all of the relevant information for the generative part or the decoder. It's an end-to-end model,
Starting point is 00:26:17 which has a couple of nice properties. It allows us to effectively back-propagate the gradient from the task that we actually care about all the way to the early layers of the model. And it allows the encoder to learn the right, rich representations for what the generative part needs to solve the task. It uses a system one, system two,
Starting point is 00:26:39 think fast, think slow architecture. and it leverages the general world knowledge of VLMs for efficient learning of semantic tasks. So let's dive deeper. First, the Think Fast Path. That part fuses the raw data from our cameras, our lighters, our raiders, and that allows for split second safety critical decisions.
Starting point is 00:27:00 So you can think of it as kind of your driving instincts. This is what allows the car to break instantly if, let's say, a pedestrian runs into the road or a cyclist that's nearby, svers into your path. This is like the, if you will, the lizard brain of your agent that deals with a lot of geometric tasks and can react in milliseconds. Second is the slow path. That's the part that's responsible for the more complex semantic and scene-level understanding type tasks. And these sort of tasks, these things don't typically change in milliseconds. So there you can afford a bit more latency and you can trade the
Starting point is 00:27:41 off for higher capability and higher levels of reasoning. So for example, if the Waymo driver encounters a situation when there's a vehicle, let's say it's on fire on the side of the road, the fast path might just see it as a generic obstacle and reason that the path ahead of us is clear. And this is where the slow path comes in, and that path can use deep semantic reasoning reasoning to understand the semantics of that object, the car being in fire, and the broader scene context, and that allows our driver to decide to take a very different action or a different route entirely, even if geometrically the path ahead of us is very clear.
Starting point is 00:28:27 And finally, there's the generate component. That's the decoder. That's the component that understands and can produce behavior. It understands how other actors behave, and it allows us to make predictions, and and plan our own driving decisions. And our Waymo Foundation model powers the Waymo driver that runs on different generations of hardware and runs on different vehicle platforms.
Starting point is 00:28:52 You have our fifth generation, the sixth generation, the JLR IPAs, the Ohai, and the Hyundai Ionic, and in the future, we'll power different products and different commercial applications like trucking and personally owned vehicles. So by leveraging this, Leveraging the strategy of focusing on the high-capacity foundation of upward model, we're able to move a lot of complexity upstream to that large, shared foundation.
Starting point is 00:29:21 And that allows us to make that specialization layer that's running on the car a pretty lightweight. And that in turn allows us to speed up the development process. So the most important muscle in this lesson is for your company to not just leverage, the tech of the day, but have the ability and build that muscle to repeatedly ride those tech waves and pulling the results of that innovation into production without regression, without breaking stride and deployment and scaling, and without drowning complexity. So let's move the next lesson. There is a well-known lesson in the AI community that general methods that leverage massive compute and massive data
Starting point is 00:30:11 will always beat methods that rely on handcrafted, engineered human knowledge. That's the so-called bitter lesson that Richard Sutton published and formulated in 2019. And we have lived this and we have seen this in every wave of technical breakthroughs. Each time the bitter lesson holds methods that scale best with compute, with data, they always went out. And by the way, this is one of the reasons why we bet on the approach of building the foundation model. There is a well-known property that if you bet on high-capacity model and you use your data and your compute on that, you just get better scaling laws and then you distill into smaller, more efficient models that are running on your agent in real time. You just get better scaling laws as opposed to just focusing on the smaller models directly.
Starting point is 00:31:07 So one nuanced area where this lesson shows up is the use of structure in your models. And depending on how you use your structure, you can end up on either side of the bitter lesson. Essentially, structure that fight scale will always lose. And structure the channel scale always wins. In particular, this comes up around the discussion of end-to-end models. As I mentioned, an N2N model has some very nice properties. You back property gradient from the final tasks all the way through the model, and it allows the API between the encoder and decoder to use rich learn representations.
Starting point is 00:31:53 And those are the easiest models to build and train. You know, you can start the architectures are known. You can start with doing some imitation learning and kind of a black box N2N model will give you very rapid progress, and you will ride that very initial steep part of the curve. And for some products, that's enough. But if you need to reach superhuman levels of performance in a fully autonomous agent,
Starting point is 00:32:17 in a safety critical environment, just doing kind of that basic vanilla end-to-end is not enough. And this is where a structure comes in. And the key question here is, does the structure boost scale or does it fight it? Does it limit and constrain your solution space? Or does it help you scale without loss of generality? So let me give you an example.
Starting point is 00:32:41 Let me illustrate this point with kind of a simple thought exercise and a toy problem. Imagine you're building a robot that will play the game of goal. And you want it to play the game in the physical world. So you have a camera that's observing the board and you have an actuator that will actually move the pieces around. Now one way you can build such a robot, is to have an end-to-end system that goes directly from pixels to actuation. And maybe you train it by giving it some videos of how humans play the game. And that could be a very interesting research exercise.
Starting point is 00:33:16 However, if your goal was to build the world's best playing Go-Robo, but that's probably not the most efficient way to go. And the reason for that is that there is a very simple, intermediate representation that captures completely the state of the game, the state of the task that you're trying to solve, as 19 by 19 board. And that gives you a fully observable and complete state of the world that you care about, at least for the game playing part.
Starting point is 00:33:46 So leveraging that structure, it doesn't limit your model. It doesn't constrain your solution space, but it gives you a very helpful way to scale. Now, that of course was a toy example. Anything that's not trivial that you're trying to deploy in the physical world will not have that property. And the fact that such a simple, clean engineered representation doesn't exist in the physical world is the whole reason why we need end-to-end systems and learned representations and learned embeddings. But in the physical world, that does exist structure. You have laws of physics, you have rules of the road, you have objects that behave in reasonably predictable ways,
Starting point is 00:34:22 and you can use that structure in addition to the learned representations to boost your performance. performance, simplify validation, and at the end of the day, just get better scaling laws. And this is the approach that we are pursuing at Waymo, which we call structure augmented end-to-end. So we go beyond the basic vanilla end-to-end by augmenting the learn embeddings with materialized structure representations. And then gives us a few very important advantages. So first is validation at inference time. Now, because the model isn't just the black box where sensors go in and, you know, actuation commands go out, we can create a very powerful correctness and safety validation layer that you can run in real time when the agent is deployed on our vehicles.
Starting point is 00:35:12 And this is really important for any agent that's operating in the physical world. Secondly, we get great wins in efficiency when it comes to large-scale training and evaluation of the generative part of the model, the decoder. If all you have, is a black box end-to-end system, you are forced to do all of your evaluation and all of your training in the end-to-end setup all the way from sensors to decisions to actuation. Having that intermediate structured representation
Starting point is 00:35:42 allows you to kind of mix and match. You can do some training at larger scale and some evaluation in the space of those compact structure representations and some in the full space of end-to-end from sensors to decisions. And finally, we get strong, very very, verifiable feedback signals for both evaluation and for training, training recipes to support things like reinforcement learning.
Starting point is 00:36:06 That additional materialized structure just gives you much more powerful tools for evaluation, for metrics, as well as crafting your loss function or reinforcement learning recipes. So the lesson here is to bet on a system that's maximally learned and minimally constrained and leverage structure intentionally to boost performance and scaling laws both in training and an evaluation. Now that raises the question of how do you actually train and evaluate your physical AI agent? And that brings us to the next lesson.
Starting point is 00:36:41 To build and safely deploy an agent in the physical world, it is absolutely critical to have a good large-scale, realistic, high-fidelity simulator. Now, there's two ways you can do training and evaluation. You can do open loop and you can do closed loop. In open loop, kind of passively observing input-output pairs, and you can use that for evaluation, for training, imitation, learning works like that.
Starting point is 00:37:09 Evaluation takes usually the shape of, if you are finding yourself in this situation, what would you do, and then you score that. And that's in contrast with closed-loop, where in closed-loop, you take an action, you see the effect that that action has in the world, then you update through your sensors, the view of the world, you take another action and so forth and so on,
Starting point is 00:37:31 and you evaluate and you train on those sequence of actions and sequence of world evolutions. Now the ability to take an action and evaluate that counterfactual is absolutely vital for building and deploying safety critical agents in the physical world. So a real simulator is how you do that. And a real simulator isn't just some lightweight tooling that sits next to your AI, it is a big AI model in of itself. And the problem of building a good realistic simulator is just as hard as building the agent itself.
Starting point is 00:38:13 So the AI behind the simulator really needs to understand how the world works, the physics, the semantics, the traffic, the weather, and so on and so forth. And the quality of that simulator has to be high enough so that it doesn't only look good, but it's sufficient to train, and evaluate with high confidence an agent that you're gonna be putting in the world
Starting point is 00:38:33 in a safety critical environment. So in other words, you have to build a highly accurate generative world model. And at Waymo, for years, we've been building what we called behavioral world models, and we're doing that way before the term world models even became popular. And now in the era of N-to-end models,
Starting point is 00:38:57 you also need on top of behavioral realism, You need sensing realism as well. And in fact, building an N2N model has been fairly easy for quite a while now. But evaluating it in closed loop, that was the hard part of the problem. So we've moved on to building sensing world models. And because we're using the structure augmented representation in our models, we can also leverage that structure
Starting point is 00:39:24 in our simulation. Our behavior world model operates in the space of structure intermediate representations and the tightly coupled sensor world model then produces realistic sensor simulations. Our world model leverages the great work of Google Deep Mines on Genie 3 and that gives us the ability to produce controllable and highly realistic scenarios both in the behavioral as well as sensing aspects. And that in turn allows us to not just evaluate our agent and train new versions of our agent in situations that we've previously
Starting point is 00:40:07 encountered, but it allows us to train and evaluate in purely synthetic rare scenarios that we've never seen in the real world. So what you're seeing here is not just the generated video, it's a generative, a full generous simulation of the Waymo driver in operating closed loop. So here we're simulating what would happen if it came across a car that was stopped in a lane on the freeway. And you can go further to that. Here's a plane that's landing on a freeway in front of us.
Starting point is 00:40:39 Where you can simulate an elephant on the loose walking through the intersection, snow on the Golden Gate Bridge, or a dinosaur walking around. So the lesson here is that closed-loop simulation is absolutely required for evaluation and is extremely valuable for training of your physical AI agents. So you need highly realistic large-scale simulation to train and evaluate. And this brings us to lesson number six. When you're dealing with a problem of that complexity,
Starting point is 00:41:17 you can't just build a model on call it a day. You have to build an entire ecosystem. And then you also need a flywheel that power is it? Because to make this work at scale, you can't just build the agent. And one AI, you need to build three. You're building the agent. For us, that's the driver that drives the car. You also have the simulator, which is that virtual playground for the agent to learn in. And then you have the critic. And the critic is what rigorously evaluates and judges the performance of the agent and tells it how to improve. And the good news is that the fundamental reasoning and the generative capabilities of all three of those are shared,
Starting point is 00:42:03 and that's why, in our case, they're based on the same foundation world model. Now, once you have these three pillars, you can create an incredibly powerful flywheel to accelerate your progress. So a deployment of your agent in the real world generates data. that data then grounds the simulator and makes it more realistic. The simulator generates harder age cases for the critic to score and for the agent to learn from. So the agent gets smarter, gets deployed in the physical world, generate more data, and that powers the flywheel and accelerates progress. But a flywheel, of course, will spin in any direction or in place.
Starting point is 00:42:51 So in order to make it go in the direction you want, you need to guide it. by metrics. And then brings us to the final lesson that your model is really table stakes, but eval and metrics, that's your most important. That's your strategic mode. So build your eval before you build your technology. Build your eval and your metrics before you build your product.
Starting point is 00:43:16 If you can't quantitatively define what good enough means, you're not really building a product you're just iterating on your demo. So nowadays, The best model architectures are fairly well known, and new ideas tend to proliferate fairly quickly. Data is incredibly important, but without good metrics, you're just flying blind.
Starting point is 00:43:42 You aren't leveraging the best data, and you can't really evaluate the ROI on making changes to it. So really, eval and metrics, that's your foundation, and that's what steers your whole tech stack. But for physical AI agents, model level evaluation is not enough. When you're putting an AI agent into the physical world, your e-val and your validation needs to go much deeper and much broader. You need to evaluate and validate every component of your system from the physical layer to the
Starting point is 00:44:19 behavioral layer that's running on board in the physical world as well as the off-board components. all of the operational processes around it. So for us, we call that the safety and readiness framework, and we spent years building and refining it, and that's what guides our development and our deployment and our scaling, and I consider that to be one of our most important assets. Again, the reason it's important is because in the physical world, trust is everything.
Starting point is 00:44:54 And Eval and metrics is how you go about earning that trust. You don't just win trust by talking about the clever technical solution or the clever state-of-the-art architecture of your models. Or doing a flash a demo. You earn it gradually day by day in the field by relentlessly proving that your system is safe and that your system works. And of course you can just prove that to yourself behind closed doors. and this is exactly why we openly publish our safety data and our safety ongoing safety research. So then that earned trust becomes your ultimate business advantage.
Starting point is 00:45:37 Your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate. So when you zoom out and look at this playbook as a whole, you realize that none of these lessons works alone. So the nine set your bar and ensure that you pick the right technology
Starting point is 00:46:05 and the right technical approach so that you don't get stuck on the local minimum. Then intentional use of structure to boost scaling and the ability to arrive technical waves of innovation. It helps you get to the right level of nines. And your AI ecosystem with the agent, the simulator, and the critic guided by your Eval and metrics, That's what allows you to build that powerful flywheel,
Starting point is 00:46:28 and that's how all of these effects compound. And it's this playbook that we've been refining over the years is what allows us to achieve the strongly superhuman safety performance of the Waymo driver. This is a snapshot of the latest safety data we've released is based on over 220 million fully autonomous miles, and we're seeing there that in the areas where we operate, the Waymo driver is about 17 times
Starting point is 00:46:55 better than human drivers when it comes to crashes that cause serious injury. And that really matters because today, somewhere in the world, every 26 seconds, someone loses their lives on a road to a crash event. And on the current scale, what that means is that Waymo is preventing a serious injury every eight dates. And this isn't just a metric on a dashboard. that means that someone's loved one got to walk through the front of the door at the end of the day, safe and unharmed. So these are just the early safety benefits of AI in the physical world, and they will only grow from there. If you look at the broader landscape, the opportunity here is absolutely massive.
Starting point is 00:47:44 Physical AI right now is where digital AI was a few years ago, and we have all of the right ingredients to go after it. We have degenerative world models. We have the architectures. We have affordable compute and sensing. We have proven scaling laws, and we have a real product operating at scale. And the last decade of AI happened in the digital world. I think the next decade will also happen in the physical world. And for those of you who decide to build in this space, good luck.
Starting point is 00:48:13 Have fun. And remember who you're building for. Your mission and your customers, that's what matters. otherwise tech is just a science project. And the end of the day, as exciting as exhilarating the tech is, nothing really beats the joy of making a difference in people's lives. What are we doing? We're in our first ever Waymo.
Starting point is 00:48:40 And what does it mean when we're in a Waymo? It means that there is nobody driving this thing. And this is a fully autonomous Waymo ride. I cannot believe this. The car did a better job than if somebody was driving. The truck was over the yellow line, so the Waymo break and moved to the side. It knew how to pronounce my name. Oh, my God.
Starting point is 00:49:10 It's nice. I love it. I'll never forget this. Never.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.