Science Friday - Creating 'world models' for robots + An AI math shakeup
Episode Date: August 19, 2026When you ask an LLM like ChatGPT or Claude a question, the model goes through its massive amount of training data and guesses the answer by mathematically predicting the word most likely to appear nex...t in a sentence. This model, experts say, will not work well for technology designed to navigate the physical world. Something like a robot that works in a warehouse will instead require a “world model” that can understand spatial surroundings, like the stuff we walk by or bang into. But what is a world model, exactly? And how do you train AI to recognize what the real world looks like? Host Ira Flatow checks in with tech journalist Joanna Stern, who’s seen the early days of these models up close, even in her own home. Then, we check in on the math world, where frontier AI models have made meaningful progress on decades-old problems. Mathematician Emily Riehl gives us the big picture on how significant these results actually are. Guests: Joanna Stern is a tech journalist who writes newsletters and creates videos for New Things Media. Dr. Emily Riehl is a professor of mathematics at Johns Hopkins University. Transcript will be available after the show airs on sciencefriday.com. Subscribe to this podcast. Follow our show on Instagram, TikTok, Facebook, and Bluesky @scifri and sign up for our newsletters. Got a science question that’s keeping you up at night? Call us: 877-472-4374 Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
Hey, I'm Ira, and you're listening to Science Friday.
When you ask chat GPT or claude a question,
the model goes through its massive amount of training data,
makes guesses on the answer,
mathematically predicting the next most likely word to appear in a sentence, for example.
This model experts say will not work well in a world that's not based on textual information.
Instead, they argue, you need a completely different approach,
a world model that can understand the physical world around us, stuff we walk by or we bang into,
and enable humanoid robots to work in warehouses and improve self-driving cars.
It's an area of AI development that's raised billions over the past few years,
with big players like Amazon and Google's Deep Mind showing more and more interest.
So what does a world model mean, actually?
And how do you train the AI to recognize with the real world?
world looks like. My next guest has seen the early days of these models up close, even in her home,
and is here to tell us how ready for prime time they are. Joanna Stern is a tech journalist who writes
newsletters and creates videos for New Things Media. Welcome to Science Friday, Joanna. Hello, thank you
for having me, Ira. You're welcome. Okay, tell me more about why tech companies want this different
approach to AI? Well, if you think about the physical world and the type of AI, we are thinking we're
going to get in the physical world, best examples would be self-driving cars and humanoid robots or
robots of any kind. You can see why just studying the text of the world and even just studying
basic images of the world is not going to be enough for a car to get you from point A to point B,
or it's not going to be enough for a humanoid robot to step into your house and start doing the laundry.
So a company like Amazon that has to move a lot of stuff around could really benefit from this kind of robot.
Sure. And Amazon has over a million robots already functioning in some of their warehouses and in their partner's warehouses.
And so the big difference there, though, in terms of what Amazon's talking about in robotics, is that those robots are largely industrial work-type robots.
right? They have a very, very regimented, simple environment that never changes. Think about it. If you are
taking things from one box and putting them into another box, the box size never really changes.
They're usually located in the same exact spot. And so this is really simple to program robots.
That's very different than, say, a self-driving car where, sure, you're still going from point A to point B, but a bicycle just went in,
between that road that wasn't there before, or they just set up a road closure sign that wasn't
there before. And so these things are constantly changing in our physical world.
The home is even worse, right? I mean, I don't know about your home, but my home is changing
every day. Someone moves something in the refrigerator. Somebody moves a plate. My kids have put
their toys all over the ground. And so is video very helpful? Is that what we're talking about here?
Well, so there's some debate about how it's going to be best to train these robots.
One way is through something called teleoperation, where the robot actually is in the home.
There's cameras and sensors on the robot.
And somebody is controlling that robot with a virtual reality headset someplace else.
Right?
They're wearing arm controllers.
And they are controlling, basically, as a puppet, the humanoid robot that's in the house.
And I've seen this demoed a number of times now.
So the robot isn't doing it autonomously.
And so this idea of teleoperation is that we are training and gathering data of the movements of the robot.
We're also gathering video and sensor data from the robot.
And that over time can start to train.
There's another idea that's becoming pretty popular.
And that's we use video data.
We take lots of video off YouTube.
We take lots of video from humans doing this in the real world and we train it.
And in fact, last month, I did a story on a company called MicroAGI out of Germany.
And they have a new app called Shift.
And you can say, I want to earn money by filming myself, doing things in my house, doing the dishes, vacuuming, cleaning the counters.
They will pay me to do these things.
That's catch is you have to wear an iPhone on your head, some sort of smartphone on your head, or now they've started to make hats with embedded cameras.
So I tried this.
I tried it for about three weeks.
I will say I was not very successful at it.
Turns out you have to really, the footage has to show both hands in the frame.
And you've got to kind of really show it very carefully.
But look, I did make some money folding my laundry, putting my dishes in, which is more to say that, you know, look, it was, I would say everyone in the house was happy that I was finally contributing.
Well, actually, watch your video.
I saw your video doing that, and it looked very interesting.
And of course, as you say, to train the robot that has to see physically what you're doing with your hands
so that it could tell the robot what to do with its hands.
Exactly.
The big thing, and I kind of challenged the CEO on this, I said, well, look, I filmed five hours.
You're only paying me for three hours.
He said, well, yeah, because only three hours is useful to me.
right? Only three hours showed my hands both in frame and my hands doing something useful that they think they can then go train the robot on.
In fact, they sent me the footage where they apply sort of the 3D modeling to my hands.
And you can see, okay, the robot has to see these weird kind of like almost skeleton-like lines in your hands doing things for it to start to make sense of what is happening.
You know, if I were to go backwards engineer things, if I were a Luddite, I might say, you know, instead of a robot trying to do this, my 15-year-old will do this for next to nothing, don't have to teach anything. Why go with a robot on this?
It's true. It's true. But look, I will say there's many times in the week where I have a big pile of laundry that I need to fold. And I would rather do something different. Last summer, I actually had a company come to my head.
house and set up their laundry folding contraption. And that wasn't a humanoid robot. That was really
these two robotic arms. You can kind of picture, you know, when you go to an arcade and they have the
claw machine that you can put the money in and it will pick up a stuffed animal or some other toy.
The robotic arms looked a little bit like that. And so these two robotic arms sat over the table,
had some cameras. It was a big sort of contraption set up. And these two robotic arms just on their
own were folding the laundry. And this, it didn't work great, but it was really great to see because I
could see, first of all, it improve over time. And it gave me a sense of where we are at right now.
I mean, it was only folding t-shirts. The company now says it folds more things. Sometimes it would
take five minutes to fold one t-shirt, right? This gives us a good idea of where we are at in this
moment. And look, this was last year and things have improved now. But this is how, this is how
the progress we need to see and why this data is so important to these companies and why there are
companies specifically popping up that just collect the data to sell to the robotics companies.
Right. And that brings me to, I'm glad you brought that point up because I'm sure there are
privacy experts who are not thrilled about people sending thousands of hours of footage inside
their own homes. Totally. And I thought about that, obviously, before embarking on the experiment I did in
that video with that company microagii, though they gave me a pretty long list of the ways
they're protecting privacy. One of them being that they blur out any identifiable information
they might see in your footage. So, for instance, they sent me back a video where some writing
on a t-shirt was blurred out. And I said, why is this blurred out? They said, well, we blur any text
that we see. Okay, it was my kids like, you know, Pokemon T-shirt, but that's fine. But these
companies are taking steps towards that, but look, you're agreeing to do this. And if you think the
trade-off of putting some cameras on your head and showing them your house is worth the $20 they might
pay you an hour, well, that's your decision to make. How do you see this panning out? Do you think
this model is going to catch on like the large language models did? Or do you think this is just
the beginning? I mean, folding a shirt for five minutes has a long way to go. Well, that's one of the
reasons I love covering this space right now. I feel like we saw this big breakthrough in large
language models. We've seen how that really disrupted society and technology and all of the
things we see playing out with AI right now. And there's a lot of hype about how this is coming
right now. People should not believe the hype about a humanoid robot coming to their house right
now. It is not coming right now. I don't think it's coming in the next five years. What we're watching
play out is this need for more data and this need for better, we haven't talked about it here,
but what we need is better and safer robots, too.
I don't want to allow a hundred-pound robot in my house that can't stand up, right?
That has a risk of falling down.
So there are a lot of things that have to happen in the next number of years.
I think it will happen.
I can't give you a time frame, but I can tell you it's not happening this year or next year.
All right.
That's a good place to end it, Joanna.
Good luck in your research.
Yeah, I'm going to keep on.
I'm keeping on this topic.
I'm going to be, you know, with the robots for many, many years.
and they'll remember when they're finally working,
they'll say, this was the woman who told the world.
So when the singularity comes, they'll keep you in mind.
That's right.
Joanna Stern, tech journalist who writes newsletters
and creates videos for New Things Media.
After the break, we're adding up the drama happening
in the AI math world.
Stay with us.
Now we're turning to news in the math world.
In the last month, AI models have disproved
a handful of conjectures.
Open AI announced that it has
had solved 10 long-standing problems with an advanced unreleased model.
Some mathematicians said the results were impressive,
while others said they were overblown and accused the announcement of sloppy attribution.
Play nice in the sandbox, folks.
So what do mathematicians make of all of this?
We thought it was a good time to check in with our mathematical referee,
Dr. Emily Reel, Professor of Mathematics at Johns Hopkins University,
who's joined us before to walk us through the AI Math World,
and she's back with us from Baltimore, Maryland.
Welcome back, Emily.
Thanks for having me.
Nice to have you.
All right, let's get into this.
When we last had you on in March, you talked about AI also.
Are things progressing slower or faster than you thought?
I think most mathematicians would say this summer has been very surprising
at the pace of improvement in the models.
I think the first one that really caught my eye was a disproof of something called the unit distance conjecture, which was a 1946 conjecture of Paul Erdisch, celebrated Hungarian commentatorialist.
I would say that there were relatively few areas where AI had made important contributions, and that is changing.
AI is still not contributing to all research areas in mathematics,
but in certain areas it is certainly making progress.
It seems to do best in areas where the problems are easy to state,
if not easy to solve,
and maybe as less skilled in the areas
where the problems are sort of harder to understand the statements of.
Right.
You know, like I said, OpenAI announced their new unreleased model
also solved 10 longstanding problems.
What's been the reaction from mathematicians, and what do you think of them?
So the announcement came in the form of a 250-page PDF, entirely written by AI, so not with any human mathematical input that we can tell.
Together with some further verification of each of the 10 solutions in the form of something called a lean-proof.
Lean here refers to computer-proof assistant that can automate some of the refereeing process.
Maybe the first thing to say is that when a new mathematical breakthrough happens, it usually takes the community some time to respond.
And if the new breakthrough involves a delicate mathematical argument that might take dozens or even hundreds of pages,
then it can take sometimes even a year or even more for the community to really verify the
correctness of the solution and also understand the impact down the field.
I'm trying to think of what a mathematical prompt to AI would look like.
I mean, it must be huge, right?
You might think that, but, you know, one of the surprises is that often these prompts are really
short and relatively naive.
A recent example is the Dennets-Garman's conjecture, which is a question in graph theory.
And here the prompting was sort of four extremely naive prompts, telling the model which problem to work on and to keep working until it found a counter-example.
So, you know, what do we make of this?
I mean, you know, one phenomenon is that there are a lot of folks out there who love math but are not getting paid to do math full-time, who are kind of rediscovering their love of math by pushing the frontier of research knowledge in this way.
way. You could imagine anybody who was aware of that problem could write that prompt and, you know,
then generate a solution. So in one sense, this is expanding the number of people who are
contributing to the mathematics research enterprise. And we've seen this with the Erdisch
problems too. The Erdisch problems were sort of famous for attracting non-specialists who could
contribute in a productive way. And those non-specialists are sort of superpowered now.
with the latest AI models.
Did AI have a particular strength that it allowed it to do this?
There's been some discussion of that.
So there are some mathematical problems that essentially involve finding a needle in a haystack.
And it does seem that AI, for whatever reason, is good at finding unusual mathematical objects
with unusual properties.
It's not as good as explaining to us how it found them.
The sort of interpretation work for some of these counter examples has been done by humans.
Yeah, you know, my math teacher always said, show your work, right?
Speaking of the latest AI models, are you getting a clearer picture on what the relationship between mathematics and AI models will look like in the future?
I think there is a lot that is currently in flux. You know, one open question is how much,
AI can actually help us understand mathematics as opposed to just tell us what is true and false.
You know, right now it still seems like there's a lot of work for humans to do to sort of map out the mathematical universe.
So you can think of the mathematical universe as some sort of vast building of unbounded expanse that mathematicians have spent millennia, both exploring and renovating.
This building is so large. There's no architect who has seen the floor plan for everything.
There's no contractor who's directing everybody to do the work.
Instead, there are sort of specialists in different areas who are interested in different
problems who are going off down some dark corridor and trying to open the doors that are stuck
to get inside.
So when there's a door that's stuck, you know, a problem that we don't know immediately how
to solve.
It's not clear immediately how to get through.
Maybe it just needs a shove in the right corner.
Maybe it needs somebody to find a key.
Maybe it needs somebody to create a new key.
And what the AI models are doing right now is helping us get in some of the way.
of these doors that mathematicians were having trouble breaking through before.
But what mathematicians still have to do once they've opened up a new door is see if they
can turn on any lights in there, maybe renovate a bit to make sure the ceiling doesn't
collapse, see whether this door leads to just a closet or maybe a whole entire new wing
of the building where there will be new mathematical discoveries around the road.
And right now, at least, AI is not contributing to those sort of larger enterprises of
expanding the mathematical universe.
Well, Emily, it's always fun having you coming on the show.
Thank you for taking time to be with us today.
Thanks very much.
Dr. Emily Real, professor of mathematics at Johns Hopkins University in Baltimore, Maryland.
This episode was produced by Dee Peter Schmidt.
I'm Ira Flato.
Thanks for listening.
