The Pomp Podcast - Inside Google's Billion Dollar Bet To Win The AI Race | Logan Kilpatrick
Episode Date: September 15, 2026Logan Kilpatrick is a member of the technical staff at Google DeepMind. In this conversation, we break down whether Google is actually behind in the AI race, the strategy behind Gemini 4's massive pre...-training run, and how DeepMind decides between chasing general intelligence versus building specialized products. We also discuss China's open-source AI labs and why measuring real progress toward AGI might be harder than building the models themselves.=====================Arch Public is an agentic trading platform that automates investment strategies across Stocks, Commodities, ETFs and Crypto. Whether you’re rotating into AI & Gold, allocating to the S&P 500, or accumulating Bitcoin, Arch Public executes your plan 24/7 without ever taking custody of your assets or funds. Sign up today at https://www.archpublic.com, and start your FREE automated trading strategy! =====================TOKEN2049 returns to Singapore on October 7–8 at Marina Bay Sands. The world's largest crypto event. 25,000 attendees, 300 speakers, 1,000 side events and the whole industry in one place for two days, into the F1 weekend. Get 10% off your ticket with code POMP10 at https://token2049.com/singapore=====================Simple Mining makes Bitcoin mining simple and accessible for everyone. We offer a premium white glove hosting service, helping you maximize the profitability of Bitcoin mining. For more information on Simple Mining or to get started mining Bitcoin, visit https://www.simplemining.io/pomp=====================0:00 - Intro0:51 - Is Google behind in the AI race?2:47 - Frontier commitment & the Gemini 4 pre-training run8:31 - Build vs. buy: Google's AI acquisition strategy13:13 - Chinese AI labs, competitors & filtering the hype19:03 - The ambition problem & where Google chooses to compete21:39 - General intelligence vs. specialized AI products35:01 - Inside DeepMind: Genome research & the innovation flywheel41:52 - Kaggle & the race to actually measure AI progress48:25 - The AI data economy: why data is the new bottleneck
Transcript
Discussion (0)
Okay, when I sell my business, I want the best tax and investment advice.
I want to help my kids, and I want to give back to the community.
Ooh, then it's the vacation of a lifetime.
I wonder if my out of office has a forever setting.
An IG Private Wealth advisor creates the clarity you need with plans that harmonize your business,
your family, and your dreams.
Get financial advice that puts you at the center.
Find your advisor at IG Private Wealth.com.
Visit BetMGM Casino and check out the newest exclusive.
The Price is Right Fortune Pick.
BetMDM and GameSense remind you to play responsibly.
19 plus to wager.
Ontario only.
Please play responsibly.
If you have questions or concerns about your gambling or someone close to you,
please contact Connects Ontario at 1-866-531-2,600 to speak to an advisor.
Free of charge.
BetMGM operates pursuant to an operating agreement with Eye Gaming, Ontario.
I think we are laser focused right now at the frontier.
We're seeing all these early signs of recursive self-improvement.
I think the other labs are seeing this as well.
And so I think it's like underscoring the value of like.
Today's conversation is with Logan Kilpatrick.
He's a member of the technical staff at Google DeepMind.
In this conversation, we talk about the AI industry, the model lab wars.
What's going on at Google?
Are they actually committed to the frontier?
How are they investing capital internally?
What are the specific products and the decision-making that's going on inside of Google?
Google DeepMind? How are the AI efforts at some of the other labs actually affecting their decision-making?
And then how should you as an individual think about benchmarking, coding agent, different types of bots,
and many other aspects of the AI industry that everyone's talking about?
Logan is somebody who has worked at a number of different companies in the industry.
He's very well-versed not only in what Google's doing, but how the industry is developing.
I think you'll find this conversation very valuable.
Here's my conversation with Logan Kilpatrick.
All right, Logan, everyone thinks that Google is behind an AI race.
What do you think? What is your response to the critics who believe that Google maybe is not near
where they should be? Yeah, it's a good question. And I think it is, I think my reflection of the last
two and a half years is like, I think it's a fair criticism because people expect a lot of Google.
This is what I try to remind myself. It's like, you know, it's not people trying to be rude
saying that Google is behind. It's like Google's an incredible company, have such a story legacy.
People expect high things of us. I think the tension point.
for us is you look at this portfolio of stuff that we're doing. We're actually talking off camera about this,
everything from like, you know, genomics work to weather and science to, you know, new frontier
models with Gemini 4 to we just released a bunch of new audio models, et cetera, et cetera. Like,
I have this firm conviction that like Google and deep mind, we have the world's best portfolio of stuff.
The tension is like the portfolio spikes in different ways. And sort of this is a very natural thing.
I don't think that's like an excuse to not be at the frontier.
And so I think I think we are like laser focused right now at the frontier.
We're seeing all these early signs of recursive self improvement.
I think the other labs are seeing this as well.
And so I think it's like underscoring the value of like being at the frontier.
And so hopefully we'll see that with with Gemini 4.
But an immense amount of progress.
I think you've seen this with the Gemini 3.5, 3.6, 3.7, 3.8 lineup of like in literally like three to four week increments,
sometimes less, sometimes a little bit more,
we're seeing very reasonable progress.
And again, this is like the early signs of this recursive self-improvement loop.
So hopefully we'll see that sort of like translate over in the same way to Gemini 4.
And it'll be our largest, most ambitious pre-training run so far.
So I think you'll, it'll sort of get us back in contention with some of the frontier labs.
Talk a little bit more about this, like commitment to the frontier, right?
I think that is one of the things that people have always wondered is,
You can go after general intelligence.
You can go after specialized workflows.
You guys have a business to run.
You get a lot of cash.
But also, you're making a lot of bets on this kind of being the future of the company.
But I do think that when I speak to people at Google, maybe there's more of a
commitment to the frontier than people realize kind of outside of the company.
Yeah.
It's part of why I had this conversation with Cori and he was sort of like making this verbal
commitment to the frontier.
Cori is our SVP of DeepMind leads the organization.
now and is incredible. I love working with him. And he was sort of making this verbal commitment.
It's obviously been everyone's perspective internally for a long time. So he was sort of doing it to
answer a lot of these questions of like, do we even care about it? And I think this stems from like
folks asking the question about, you know, why haven't we landed the pro model? Why are we pushing
on flash? Do we only care about flash models now because we're shipping flash? And the reality is like
we were just seeing a lot of progress. There was a lot of juice to squeeze. The recipe that we had was like
working extremely well with the flash models.
And so there was this like really tight iteration loop.
People were getting excited about the progress and they're like,
hey, let's focus and make that happen.
But I do think like the one of the things that makes me most excited about
Google and the reason that I'm here and the reason that I'm doing this work is
because the front, having a frontier model and the commitment to the frontier so
deeply permeates across Google's business.
You look at like every single thing that we're doing as Google, the way that, you know,
our 13 plus billion user products are sort of like touching the world in different ways from
workspace to everything else we're doing across search. Having a frontier model is core and
fundamental and a direct accelerant of every single component of that business, not to mention
the things that are like more exploratory. Like what is having, I mean, we're seeing signs of
this from Open AI and others. Like, what does having a frontier model mean for frontier drug
discovery and being able to cure cancer? And like, obviously there's a correlation between
those things. It's like a not immediate correlation as, you know, having a bunch of products
with billions of users, but like there is going to be some correlation. I think that correlation is
going to increase over time. So it is like the business is set up to like fundamentally depend on
having a frontier model. And so you go ask everyone, I think this is to your comment. Like,
it is the most important thing. There's nothing else. We're not thinking about like, hey, let's go make
small cheap models that are great for everyone else. Like we do that because there's use
cases in Google that support it, but the aim is to be at the frontier and being at the frontier
will enable us to do all the other things.
As you go through that decision making, it is kind of an interesting thing.
Most businesses, they would look at like, what's the problem we face?
Okay, let's go build technology to solve that problem.
When you're committed to the frontier, though, there's some version of just like, let's go
build the smartest model possible and then we'll figure out how to apply it.
And then once you get into the application of that, you know, kind of general intelligence,
it becomes, okay, do we do it in a general way where people all throughout the company can
just ping it or customers can do that.
Or do we go build specialized workflows and kind of go through that path?
Maybe just walk me through the decision tree, like,
bring us in the room as you guys are thinking through some of this stuff.
How do you guys decide where to put resources and like maybe what the sequence of events is,
at least, you know, aspirationally for you.
Yeah, I'll make a general comment, which is, you know,
Sam Altman has that famous interview where he's like being interviewed by people a long time ago
ago and they're like, so what do you, what's the plan to make money off this thing?
And he's like, I don't know, we're going to make the smart
this thing possible and then we're going to ask it how to make money or something like that.
And, you know, a little bit of a tongue and cheek answer. But I think that's actually like
less of what we've been trying to do at Google from the sense of like, we actually know where
the models will create tons of value for our customers in the world. We have all these products.
We have all these services. We have the fastest growing cloud business in the world, et cetera,
et cetera. So like, it's very obvious what the commercial application is. You don't have to think
too deeply about that. I think to answer the question specifically about like, what are the
set of tradeoffs and how is the resource allocation being made.
I think that's where this really sign of sort of like recursive self-improvement is coming
from.
And I think there's a lot of resources being focused on like, how do we actually make the models
better at coding and research and science and sort of the core work needed to make like
further breakthroughs and accelerations of model progress.
And I think folks are very interested in that because they're just like such a large
economic opportunity and then having a great model will then enable us to do all the other
things that we want to do.
So there's a huge, I think everybody is very code-pilled, very science-pilled, very focused on that
right now.
And I think we were like a little, it's not that we relate to the game because folks I think
knew it was important, but I think you, it's like, yeah, in hindsight, everything is much more
clear.
Like I think in hindsight now is obvious, like we should have, you know, from an order of magnitude
of resource allocation, probably
put more into coding sooner. And like, that makes sense. And we had a bunch of other stuff that we were
doing, which those things actually turned out quite well. Like, you know, a good example of this is,
you know, nanobanana, like a great, incredible image model that sort of took the world by storm,
had this massive impact for our consumer products and a bunch of other parts of the business. And like,
you know, to make that model, it took research and compute and time and sort of and had this huge
impact. And like, was it in hindsight right to do that versus doing something on coding? Like,
I don't know that there's actually a clear right answer,
but those are the types of tradeoffs that are actually a lot easier to analyze in hindsight.
And it's harder to know in the moment whether you're making that right decision.
But we do that reflection and sort of introspection, which is important.
If you look at some of the companies, I think Open AI has been very acquisitive in trying
to go after some of these verticals.
Obviously, SpaceX AI recently went and bought Cursor and then they launched Grockbot,
they've got a lot of the coding stuff.
It does feel like there's different strategies.
There's a little bit of chess getting there.
Some people say, hey, look, we're going to focus on certain things.
We're going to build it from scratch.
And that's kind of our ethos and DNA.
Google and other areas outside of AI has done both.
They've built some things, but they've also acquired things.
How are you guys thinking about if you feel like you're behind maybe in coding or other areas?
Will you guys go buy stuff or is the focus internally on like, let's go build?
We know how to do it.
We've got the resources.
It's just a focus and that kind of strategy thing.
Yeah, we've definitely done both in this context.
And so we had, I think the Gemini-CLA launched
early, maybe almost like a year and a half or two years ago,
and sort of had some early traction.
I think a few million users using that product.
And then we also went and did the sort of acu-hire of the windsurf team,
which is now part of cognition, whom we are talking about off-camera.
But yeah, and so we have a bunch of those folks.
They've sort of been building this anti-gravity product internally,
both for our internal engineers, but also for external users and have seen,
actually I think one of the biggest impacts has been the internal acceleration.
And so they're really, really deeply focused on like,
how do we build a great product for engineers
and actually even non-engineers now inside of Google,
make it work really well.
And then that sort of will translate to a great product
externally eventually.
And so it's been cool to see us do both things.
I think it's one of the interesting strategic advantages of Google
is we get to take many shots on goal.
And so yeah, I think having a bunch of different coding products.
And it's also like, I think what's most interesting
on this thread is the ecosystem,
has evolved. I think about this all the time, like, the coding product that you would go to market with today,
or the sort of like developer or whatever, like, knowledge, even more generally knowledge work products
that you're going to market with today looks so different than what it was even 12 months ago. And actually,
like, GrockBot's a perfect example of this. Like, it's not obvious to me that, like, the Grockbox style
product would have worked 12 months ago. I think it works now because the models are good enough.
But, like, 12 months ago, you actually, like, needed all of this, like,
developer UI scaffolding of like, you know, give me all these extra features and buttons and
things because like the model is not really that smart and I need to wield it and turn it and
sort of critique it in these very, very specific ways. And so I think one of the, my observation of this and
my point of this is like, I think there's going to be many of these opportunities to take shots on
goal as the model progress continues. It like unlocks a new paradigm. And it feels like, you know,
Grakbod and Instinct and Muse and a bunch of these new products that are coming to market right now are like the evidence of like the models have crossed another chasm where the product experience you can now build is fundamentally different than the one you were building 12 months ago.
And this is true for developers.
It's true at a bunch of other verticals as well.
With Gemini 4, you guys have talked about this large pre-training run that you're doing.
What should we take away from that?
Are there specific things that you can share in terms of what that's going to look like?
Yeah, I think the thing to take away from this is like the commitment to the frontier.
Doing large pre-training runs is extremely expensive. It is like a large order of magnitude of investment.
It's like I don't the numbers are quite large.
You want to tell us?
I actually don't even know what the numbers are at the top of my head just, but I can do some of the math in my head and like it's a lot.
And I think it's important. I think the other point of this actually is like pre-training has been a significant
strength for deep mind in the past.
And so I think this is like one of the areas where I think we have like some of the best
talent in the world.
We have like actually, this is like the infrastructure scale of Google is an advantage in
this context.
Like the TPU fleet is an advantage in this context.
A bunch of the data infrastructure stuff we have is an advantage.
And so I think the point of telling people about this pre-training running is to tell people
that like it's not like we're rolling over and playing dead.
Like we are pushing the frontier.
everybody's working as hard as humanly possible,
we'll hopefully see a bunch of incredible results
from this new pre-training run.
And actually, interestingly, you do the model comparison today
and all of our, I think this pre-training run
will very specifically get us to the category
that we need to be at to be competitive
with where our competitors are at.
And so I think the proof will be in the pudding
when we hopefully launch this model
and customers get their hands on it.
But I think that's the expectation
and that's the hope right now.
When you say competitors, I think most people will think about Open AI, Anthropic,
you know, GROC or SpaceX AI, et cetera.
Do you guys worry at all about like the Chinese open source, open weight type model players?
Do you worry about maybe other competitors that aren't one of those three companies?
Yeah, I think what's so interesting right now is like the space feels incredibly dynamic.
And so I think, I mean, I think all I personally have never understood all these memes of like,
I don't think about the competitor.
I'm like, I think about our competitors because they're all incredible companies.
And like, I think we would be wrong to not be thinking about what they're doing and, you know, examining, are they making the right decisions?
Are the things that we could be doing differently?
Still sort of like knowing what our core focus is.
And so spend a lot of time looking at like what are the things that folks are doing.
And obviously the Chinese model labs have done an incredible job so far.
It's like there's clearly a bunch of like question marks as far as IP stuff, model training stuff.
But with the set of constraints they have all things considered, they've seemingly done a pretty solid job.
And like there's clearly research innovation they're doing as well.
It's not like they're just copying what everyone else is doing.
There's like actual frontier research happening.
So don't want to discount them as a competitor.
It's also clear that like startups and companies want to use models that they can host themselves.
Like I think there's like a there's like a philosophical question of like, oh, how do you, you know, what's the, what's the feeling about these labs in China open sourcing these models?
and doing the thing they're doing.
And then there's like a practical business question of like,
customers want these type of models.
They want to, like I talk to startups all the time.
Startups want to be able to take the way to the models
and customize them for the use cases that they care about.
And so there's a huge market there.
There's a huge opportunity.
And so I think the question is like, will we see like US open source labs
and like in videos, you know, spinning up these types of efforts?
We have some of this on the smaller on device model side with Gemma.
I think we'll see like Reflection AI, a bunch of other folks, like, take shots at like,
can you actually produce frontier open weight models?
But I think the cool, actually the cool thing for all of us is that just how competitive it is.
Like the fact that those labs are able to like stand in a similar regard at all to these large
companies in the U.S. is actually, I think, a good thing for all of us right now.
And so there's a there's a huge amount of competition that's pushing everyone to be better.
Today's episode is brought to you by Token 2049.
The largest conference in crypto is back.
Token 2049 will host 25,000 people, 300 speakers, and 1,000 plus side events in Singapore on October 7th and 8th at Marina Bay Sands.
The speaker list is absolutely stacked.
Shane Copeland from Pollymarket, Jeff Yon from Hyperliquid.
Adina Friedman from NASDAQ, Arthur Hayes, Balaji, and Eric Trump.
Crypto and traditional finance in the same building, which tells you a lot about where this is going.
then the conference runs right into F1 weekend,
so the whole thing turns into one giant week.
If you're headed to Singapore, use code Pomp 10 for 10% off your ticket.
Token 2049, October 7th and 8th in Singapore.
Go check them out in the link in the description.
When you think about playing chess,
you definitely got to understand what your opponent is doing.
I agree with you that only focusing on your pieces does not help you win the game.
With that said, though, it does feel like there is a lot of question marks
about how some of this stuff is getting done.
If you think of some of the math problems that have recently been solved,
it was, hey, did the models train on other people's questions?
Was there some peeking at data that people thought was private?
There's some questions now about was Kimmy actually passing some of their queries
just to Claude to answer versus Kimmy doing it themselves.
And it's very difficult, you know, at least for me,
but I think many of people to understand what is real and what is just like Twitter fodder
or ex fodder and where people just say, you know, they like they like the drama.
It's almost like the TMZ of.
of the AI industry, right?
Like what's the new thing that we could all,
you know, grab hold of for the day?
How do you personally think through, you know,
where to spend your time in terms of like your attention?
Because it's happening so fast,
there's so many different things,
you know, even if we just think over the last week or so,
you've got everyone from Paul Tudor Jones putting out op-edge,
you've got, you know, Jensen talking about AGI,
you've got AGI, you've got like all these components,
unless you figured out how to get one in 24 hours in a day,
you know, you don't have as much time.
So what is your process to do that?
Yeah, I think this is actually an interesting point that you're making, which is,
and I think this has, I think the trend has changed over time in like the level of signal to noise.
I think there's actually just a lot more noise these days.
And so I do think it is like a, it is a muscle that you have to build to sort of filter as much of this stuff as possible.
And like, for me, it means like I am definitely passively consuming a bunch of this stuff and trying to engage.
and things, but like, I'm trying to stay focus.
Like, we need to be at the frontier.
The most important thing that I can do is, like,
help us go build better models.
There's a bunch of stuff to stay on top of and make sure that, like,
we're reacting to the right things that are happening and being proactive where it's needed.
But, like, I think it's really easy to get caught up in all the crap
that's happening in the world right now.
And, like, I think my advice to people is, like, filter out as much as possible.
Be more intentional about how you spend your time because, like,
there's a lot of noise, and it's not always clear to me that, like, the noise is actually
translating to any amount of signal.
And so, yeah, it's like there's, yeah, there's very specific cases where this is true.
Like, you know, the hugging face open AI situation is like a good example of, like,
lots of noise, there's definitely signal there.
There's something to be learned.
There's something to understand.
There's a lot of cases, though, we're like, this is not the case.
And so, yeah, trying to, try and be intentional about the places where there's actual
signal.
If you almost take that same issue or challenge.
and flip it, the other side of that is like there's a lot of opportunity cost, given that the
cost or barrier to build things has come down so much, you now have access to superhuman
intelligence. You can vibe code things. You know, you can have the bots go and build companies or,
you know, kind of run parts of your business. Like, it does feel like not only are there more distractions,
but if you get distracted, the opportunity cost is higher than ever. And so how do you think about,
your role internally, you've worked inside of Open AI, you've worked at Google,
like maybe like, what are some of the things you've picked up and how you're navigating
the productivity side of this as well?
Yeah, it's so true.
And I think actually the thing that I struggle most with now is like, um,
it's this like level of ambition problem, which is like, I used to be able to be like,
oh, I'll just go like do this thing. It's going to be small and concise and well-scoped.
And now it's like, shit, I actually like, if I go do this, like, this could be
be a billion dollar opportunity for us.
So I have to take it quite seriously.
And that weighs on me.
And I'm having to spend more time to be thoughtful about,
is this the opportunity that we really want to go after as a team?
Because there's so much opportunity everywhere.
I think for me, this goes back to like, I try to be extremely principled about what are the things
that Google is well positioned to compete in?
And there's a lot of things that we're not well positioned to compete in.
There's definitely some that we are well positioned to compete in.
And we like, we have structural.
manages with Google Workspace and Google Cloud and distribution and things like that.
And so that's sort of my filtering mechanism on the product side when we think about, like,
what are the opportunities to go after?
Like, I don't want to go after everything.
I want to go after things in which there's like a natural lift because we have other assets
inside of Google that will actually contribute to the success of these things.
And so this is what we've done in AI Studio.
We have all these deep integrations with Google Cloud and all this stuff that like,
no other product team in the world can actually do because they're not inside of Google
building this product.
And so it means that we're competing in some of these categories, but we're running a
playbook that only we can actually run.
And so we'll see in the fullness of time of was that the right playbook?
Does it actually make sense?
Maybe we should have just been doing the things everybody else we're doing.
But I think it's in, I'm trying to keep that filtering mechanism very top of mind as we're
making the decisions.
And there's just so many cool things that Google has that make this, like, fun.
And so I feel like I'm not limited by this at the moment.
One of the aspects of the AI industry that is just intellectually stimulating,
I think, for you, me, many other people, is there's a level of strategy that is being played out.
So it's not just like, can you get the hardware and the software to do certain things,
you know, kind of create magic or turn sand into intelligence?
Like that is obviously very difficult and plenty of challenges there.
But the strategy side, we were talking previously
that most of the frontier models are pursuing
general intelligence some form or fashion.
But then there's a bunch of startups that are saying,
well, what if I take a specialized workflow approach?
And if you think of what we've been building with Sylvia,
this idea of, well, if we go and we build a bunch of proprietary
technology that from model routers to harnesses to data pipelines
and our own models, et cetera, it does feel like there's
almost a point on each application of AI.
where you kind of have to decide,
do we go after general intelligence
or do we go after the specialized workflows?
And Harvey, Sylvia, there's many players, I think,
that are seeing a lot of traction in specialized workflows.
But what I find fascinating about Google
is you guys have multiple applications
where you have to make that decision over and over again.
Like, do we go and build the specialized workflows
or can we just use the general purpose model?
How are you navigating that?
Like, for each one of these use cases,
it's almost like you guys may be making the decision
more than anyone else in the world.
Yeah, actually, I've got two points for you on this.
One of the thought exercises that I am continually proposing to our team internally is in five years, do we expect, and maybe five years is the wrong time horizon.
But like in five to ten years, do we expect Google to have 10,000 products or two or three products?
And I think this gets to this like vertical workflows versus sort of like general intelligence.
And so I think there's like clear signal in the market that customers want vertical applications.
They don't like, you know, you think about like why apps are so successful and like that's sort of there's this this user behavior pattern, which is like, hey, I think as a user in terms of like using a particular application or a tool in real life.
I want to go swat a fly.
I go get a fly swatter.
I want to go drink water.
I get a cup.
I don't like go to this like Oracle all encompassing.
tool that can morph to do.
It's not something that like we intrinsically have grown up and sort of evolved as
humans to understand.
And so I do think there's this like really deep rooted muscle memory.
And I think the tension point will be given that extremely deep rooted muscle memory.
Does it, is that enough of a sticking point that will like keep these vertical applications
alive in a world where alive and thriving in a world?
where the general purpose thing can actually do the same stuff.
And so I think that will be the most interesting.
And this is where I think this like, what's the intersection of like AI and new hardware,
consumer hardware devices, I think is going to be really interesting.
As people change the way that they work with software and with technology, like, you imagine
you will want like new form factors because the form factors we have right now are sort of
a little bit more of these like verticalized experiences.
I do think there's a separate edge of this, which is,
And this is true in all these vertical domains.
The vertical domains are successful also because somebody is focused.
And I deeply believe this.
Like, you know, startups are always worried about like, oh, is the big company going to come after me?
And it's like, you can always do a better job than the big company with like very few exceptions because you're focused and more deep on some vertical that your problem that your customers have that like no one else is going after.
And so I do think it's like an edge to have that vertical mess.
which is really interesting.
I think the other point that I wanted to make,
and I'm curious actually what you think about this,
I think there's all this conversation of, like,
general intelligence.
And something that's been, like,
that's been very top of mind is I think the labs have historically
described general intelligence as if, like,
they would build the general intelligence themselves.
And that, like, seemingly the general intelligence would then be powered by,
like, end to end, almost the models that one of these labs,
creates. I think there's something really interesting about this future where, like, if we really
had general intelligence, you wouldn't expect that the general intelligence that Google creates
is only using Google product services and models. You'd expect, like, hey, if that's generally
intelligent enough to know that, like, this other thing that some other company created can do something
that our thing can't do or can do it better, humans are generally intelligent enough to figure that
out. And so I think it actually, it's going to add a lot of, I think, on this power.
to general intelligence as like the model labs go and continue down this direction,
I think it's going to add a lot of like confusion to even like understand what that really ends up
becoming because I think it's going to look a lot, it's going to look a lot more chaotic, I think,
and how this like these general intelligence systems play out. And I think this like beautiful vision
of something that can just like do anything and everything for you.
It's interesting you talk about this. So before we talk about the general internal,
let's talk about like a microcosm of this. And you know, the problem that I've been thinking the most
about for the last year and a half of Sylvia.
But for those that don't know, the product,
you basically come in, you attach your financial accounts,
you put up your private investments,
and you start talking to Sylvia,
but we have chosen to go to specialized workflows,
and we have done a whole bunch of very innovative things,
I think in terms of the A harness,
the memory and file system, the model routers,
that fine-tuning, et cetera.
But if you take like the model router,
one of the perceived advantages of not being a Frontier Lab,
is that you should be able to route,
queries to any of the model lab models, right?
And so what are the odds that Google is going to route to
open air anthropic?
It's not zero, but it's not 90% either, right?
And so same thing, I think with each one of the labs is like,
what is the incentive for them to keep the queries within
their family of models versus the ability to act more as like
a third party and actually route across?
I don't think we've really seen how everyone's going to play that.
And so as a third party, you're like,
well, I don't really care, right?
I just want the best level of intelligence
at the lowest cost that answers the question for our user.
And so I do think there's some of those things also
where people are trying to figure out,
not just like, if you then extrapolate this out
to like general intelligence, the user doesn't care
if it's a Google product or not.
Maybe there's some like ethical or moral things
that maybe they align more with,
but for the most part, they just use a product
because it's the best one.
And is that going to actually be built by one company
or is it going to be kind of a,
a bundling.
I mean, you know, what's the saying is like the world is just bundling and unbundling over and over
and over again.
And so, you know, it is a very difficult thing to predict because I think we're so early
in this journey that like every week, so it's got a new model that seemed to leapfrog everybody
else.
And then you're trying to predict how consumers are going to interface with this stuff.
And maybe like the last example I'll give is in my own life, I have been using for the last
couple of weeks, Grock bought professionally and instinct personally. They actually do a lot of the same
stuff. But to your point about like you get the cup for water and you know, you get the fly swatter to
swat the fly, like I just kind of have in my head, you know, okay, instinct when I got a personal
question and Grock when I'm doing something professionally, it's probably pretty dumb, you know,
like if they do the same thing, like why don't you just use the same product? But it then goes to like,
okay, well, now you're starting to see four or five others come to market.
And as somebody who likes to be an early adopter, you're like, well, I should try those,
but then what about the context? What about the memory? How do I port that over? And there's like this,
you know, kind of like user journey. We're all learning of, is it worth the time to go try the new thing
if I don't have some kind of shared memory? Or, you know, especially you get network effects where
like you're a wife and you have shared memory somewhere, then how do you interface with something?
And so like, it's almost like more questions than answers right now.
And I think that's probably why, you know, you and I and so many other people are so excited about this, right?
Yeah, yeah.
I think you're right.
And I love this like the bundling and unbundling analogy because I think there's another version of this is like what's old as new again or what's new as old.
Or whatever the expression is.
And like actually you see this with what's so interesting about Grockpot and instinct is like this like form factors.
Like what's old from a form factor is not like message.
Chat was one of the original ones.
You had all these like mechack thoughts like in the 80s and 90s or whatever it was.
And then like that came back and then boom, all of a sudden, what's old is new again.
And then the same thing is now true for messaging.
It's like we all use all these messaging apps and then it's like now all of a sudden
the hottest form factor for AI is like the messaging.
And so I think it's an interesting exercise of like it actually in all of these domains.
I think the reason people feel this way is because like they're accustomed to this experience,
getting some customer to like adapt to some few.
futuristic new thing is actually extremely difficult to do and takes a really long time.
You want people to go to some form factor they're familiar with.
I think about this all the time for like, how do you actually get consumers or users to
go and adopt new technology?
It's like you want to make it feel familiar.
And I have this hypothesis that like messaging like in actual like chat apps, like has,
it seems so unlikely that that's not going to be the dominant form factor.
a few years from now.
It's surprising to me, it hasn't been more dominant.
And I think it's because actually the operating system,
like messaging app owners are like just at the cusp of this.
But you'd expect, you know, like Apple and Android and WhatsApp,
et cetera to like really lean into that form factor.
And they already have where all the communication is happening and
you throw some agents in there.
And like, you know, it feels like it's a natural place.
It does feel like there could be some platformers.
Now, I think that the people developing these products are obviously partnering with and trying to prevent that.
But if you wake up and you've got a chat bot that's very popular and all of a sudden you're blocked on Apple's system, that would be a big problem.
Right. And so, you know, I don't know if that's really an Apple's best interest to do that stuff, but I do think that there's some folks who are kind of thinking through that.
The other aspect, though, around, you know, the kind of chat bots and the messaging, I do agree that it's an interface that we all are very comfortable with.
I use it on a daily basis with these bots.
But I have not yet become a very big voice user.
And I have a lot of friends that are voice-pilled.
They're walking around.
They're like whispering in their microphones or whatever
or sitting at their desk.
Do you use voice or like, what is it maybe the adoption
if you had to predict it inside of like the AI team at Google
in terms of people who are fat fingers on a keyboard
versus using voice?
That's a good question, actually.
And I've been flowed between this.
There's like, actually, what I found is like the best use case for me for voice is when I'm,
like, doing some sort of demo in front of other people.
And that way, I don't have to, like, fumble typing things and I can't spell and all that
stuff.
And so just voice, like, straight in is like way faster.
It makes the point.
It's like much, much more succinct.
But there is something about, I think it's a, and I'm sure there's like good, you know,
neuroscience research out there that explains this.
Like, I think it is like people manifest thoughts in different ways and like the
physical manifestation of thoughts either coming audibly or like through tactile typing or even
writing.
Like to me, I feel like I have like different thoughts depending on the sort of expression form.
And so like the way that I spend a lot of time talking and doing stuff at work and I spend
less time writing sometimes.
And so it's like actually quite helpful for me to like pulls me into a different mode of
thinking when I start to write just because of the form factor.
And so I think we'll see actually more of that as well.
And that's, I think the, you know, obviously people.
like audibly speaking, and it's a helpful way to think through things as well.
But I think we'll see the sort of buckets of these different types of thinking
manifest from how you interact with AI as well.
It is interesting.
I think the science shows the single best way to remember something is to physically write it down,
like with your hand.
Next would be typing, right?
And third is just kind of hear it and don't do anything.
But maybe there is something about not just the memory, but also the ideation.
Right.
you know, we definitely know from science that walking outside,
kind of the act of moving, you know, forward,
does a lot of ideation, showering, right?
There's many kind of examples throughout history,
people who just wanted to shower and thought of things.
I shower twice a day for this reason.
It's not for hygiene.
It's just for, I mean, not tongue in cheek,
but sometimes, honestly, because like you do just have headspace.
It's great.
Yeah.
And look, part of it is like, are we just so all terminally online
that just like the shower is the only place
that the phone doesn't go?
Or is it like there is something about, you know,
even 50 years ago before people had, you know,
super computers in their pocket, the shower did lead to new ideas, right?
I think the shower was just cold 50 years ago.
And so people were just being shocked and, you know,
having new ideas probably.
I love it.
Let's talk about DeepMind more specifically.
You know, the work there obviously has been very broad
for a very long time.
We mentioned a little bit about the genome kind of project
and the work that's being done there, I'm pretty surprised that just how large it is,
but it still doesn't get maybe the respect that it deserves.
Can you talk a little bit about some of what's going on there?
Yeah, no, 100%.
I think I'm not an expert on all the science stuff that we're doing,
but it's incredible to see the progress.
I think across the way that I would frame this is deep mine is split up in a couple of different
ways.
They're sort of like foundational Gemini, and there's like,
we want to make the best frontier model and a bunch of different sizes of that
model and sort of all of the different modalities that work in mainline Gemini.
And then there's a whole science unit.
And inside the science unit, there's everything from the genome project, a bunch of the
alpha-fold stuff.
There's a bunch of like science things related to like biology.
There's a bunch of other science stuff related to like weather and mathematics.
And that whole portfolio is like also at the frontier of doing all these really interesting
problems that no one else is doing.
And then has all these very unique collaborations actually with folks like isomorphic
labs, which is the, our sort of like drug discovery company inside of Google and Deep Mind
that Demis is the CEO of.
And, you know, they do all these deep collaborations.
And so it's a really interesting way for them to like not only solve the problem and make
progress in the problem from a foundational research perspective, but then actually have like
the applied side of it as well.
And so I think it's this like unique flywheel that exists inside of deep mind itself where like we're creating a bunch of the frontier innovation or doing all this interesting science work.
And then it actually has an application.
It's not like we're just doing it for the sake of doing it.
And I think this was actually the lesson from from Alpha Fold, which was like, hey, we did all this really, really interesting work.
It was super interesting.
But we're actually doing it like to solve a scientific grand challenge less because like we had somewhere where it was an immediate commercial application.
of, but it's like, it became very clear, like, hey, there's all these commercial applications.
We're opening this up.
The scientists are all using it.
And so I think the general philosophy is, like, do this frontier science work across genome,
across weather, et cetera, and then actually have a place to apply it to inside of Google.
And then actually, most interestingly, take a bunch of the lessons in learning and data
and other things and upstream those back into the mainline Gemini model.
Because ultimately, like, the mainline Gemini model is going to become better at a bunch of those
things than those individual domain-specific models.
And so you need to make sure the flywheel also goes back to there.
And so that's why having it under like a single roof actually makes sense.
And we see the cross-pollination between these things.
We've seen historically like all of these interesting like alpha proof with mathematics
trickle back to directly increasing the reasoning capabilities of the model for
mathematics in the mainline Gemini model.
We've seen this for cyber now.
It's not an alpha project, but it's a similar domain.
or like cyber capabilities as we push the frontier on cyber directly correlate to like models having better coding capabilities.
And so there's all these other examples where like this flywheel spins.
And I'll make one comment, which is I think people talk about the flywheels stuff like this as if it's like a magical thing that just like works.
There is an immense amount of effort and energy that is required to actually this the flywheel does not like you don't spin it.
And then it's like a hamster wheel.
It's like you are manually pulling it and like forcing it to work because like you know that the outcome is going to be great.
But like I have to remind myself this and our teams this internally because you think of this, this magical thing that's always spinning and you just throw things into it.
That's not how it works.
It's a lot of effort and energy to make the thing actually move.
Today's episode is brought to you by Arch Public.
Arch Public has just expanded its agentic trading platform beyond crypto.
So pay attention.
This is a big one.
Now they are automating strategies across stocks, commodities, and ETFs, and I think that this is going to be huge.
You can now automatically take profits when one market hits new all-time highs and rotate that capital into other markets showing more opportunity.
Whether you're rotating capital into AI stocks, gold, if you're investing in the S&P 500, or you're accumulating Bitcoin,
ArchPublic brings real discipline and automation to your investment strategy.
Additionally, they've launched a powerful new tax loss harvesting tool.
With crypto being so volatile and its exemption from the wash sale rule,
Archpublic can offset gains with losses without compromising your long-term positions.
It's exactly what every serious investor does.
Institutional-grade automation that works across every major asset class.
There's no more emotional trading, no more missing tax opportunities,
just smarter, hands-free execution of your preferred strategies.
Go to archpublic.com right now.
Connect with their team, set up a time, bring your account if you'd like,
and then you can learn what automated trading can do for you.
Archpublic.com.
Today's episode is brought to you by simple mining.
Bitcoin mining has a reputation for being complicated, risky, and hard to evaluate
as a real investment.
If you're considering mining in 2026, what actually matters is in headline profitability.
It's uptime, repairs, and whether the operation is run like a real business.
That's why I've been using Simple Mining.
They're based in Cedar Falls, Iowa, and they run a white glove hosting operation where you
own your miners.
You choose your own pool, and you have Bitcoin sent directly to your wallet.
They were featured on the Inc. 5,000 list as the fastest growing company in Iowa with over 40,000 machines under management.
What stands out to me is execution.
They have the number one rated ASIC repair center, and for the first 12 months, repairs are included.
If mining margins get tight, you can pause with no penalties.
And if you want to resize or upgrade your fleet, there's a marketplace to resell equipment instead of being stuck.
To help people think it through whether mining actually makes sense right now,
they put together a short resource called the 2026 Bitcoin Mining Blueprint.
It walks through the five mistakes investors make when allocating the mining,
and they also explain how to avoid them before deploying capital.
If it sounds interesting to you, you can get it for free at simplemining.io slash pomp.
That's simplemining.io slash pomp.
Go check it out today and see if you should get into the mining game.
It's funny.
I thought AI was just going to solve all our problems.
We would have no jobs.
You know, we'd just be like hanging out at the beach, but I have said it over and over again.
Every single person I know is working harder today than they've ever worked in their career.
And a lot of that I think is just they feel like it's a big moment.
You've got to kind of accelerate to be able to capture kind of your piece of it.
But at the same time, I think that people are inspired, right?
That there is this element of imagine if you can be part of a team that accomplishes, you know,
XYZ thing.
And, you know, obviously deep mind is a big part of that.
Before we let you go, let's talk about is it Kegel or Kagle?
How do you actually pronounce this?
Cagle.
Cagle.
Cagle.
All right.
Explain a little bit as to what it is.
You're now running Cagle.
And maybe kind of like what your vision for the product.
Yeah, I think the sort of the perspective, actually historical context, Cagle is sort of a startup
Google acquired in, I think 2016 or 2017.
He's done a bunch of interesting stuff inside of Google.
We sort of brought the team over into, to be part of my team earlier this year.
And the sort of the basic hypothesis for this is like,
model progress itself, and like this is like such a important point to underscore,
model progress is gated by our ability to measure progress.
You cannot make progress on something that you aren't able to measure.
And so actually, as you see, one of the most interesting things in the last like three or
four weeks is like you look at Fable 5.1, you look at Astra, you look at hopefully Gemini
4 as it lands in the market.
Like these models are saturating all of the available benchmarks.
And so now you sort of sit there and you're like, okay, well, where do we go?
Like, we don't, there isn't a bunch of problems that are difficult that sort of, that we can
actually measure and like scientifically continue to hill climb.
And so the mission for the, for the Kaggle team and for building this platform is like,
we want to build the most open benchmark and evaluation platform in the world so that people can
come together and collaborate on, on like, all of these extremely difficult frontier benchmarks
and challenges and competitions so that we can actually see.
difficult problems that models can't yet, have not yet saturated,
and we know what we can actually measure.
And I think there's like a bunch of nuance bits of this,
like the everyday, and I say the everyday person in quotes,
because like I'm sure the everyday person is not going to,
but like people who care about this technology,
being able to show up on a platform and have a voice and like,
and a say in, you know, how are we measuring progress towards AGI?
What are the things that we should care about?
What are the types of tasks and bench
that actually prove these things.
What are the for my company?
What are the things that you as a company building Sylvia actually care about?
What are the capabilities you wish you had in a model that would unlock entirely new sectors,
entirely new geographies, entirely new use cases for your customers?
And having a place where you can articulate that in a way that this is my, I had this epiphany
a year and a half ago where I sat in years of
like customer conversations where sort of you would take a customer and they'd say,
hey, I wish the models could do this and here's an example of that.
And then you'd see a researcher sort of, you know, with a blank stare because like it's
the work to translate sort of one anecdotal example into something that like an AI researcher
can actually take action on is like it's on two ends of the spires.
It's impossible.
It's not capable.
And so you have all this great feedback coming in from customers and you can't take action
on it from a model perspective.
And so the exercise is like, how do you get people
to speak the same language?
The language that the researchers speak,
the language of model improvement is in the form of benchmarks.
You need to be able to measure something
to make scientifically rigorous progress on it.
And I think the world is like slowly starting
to wake up to this fact.
And I want to help accelerate this because progress
is bounded by our ability to measure progress.
And so yeah, excited.
We're like definitely in the early
these stages of this, we'll have lots more stuff to share soon, but trying to get the world
building more benchmarks so that we can make progress for the stuff that like real people actually
care about, not like a bunch of academic stuff that people don't care about, but like use cases
that like you personally and your company have and every other startup and company has is really
important. I think this is like one of the problems of the decade. We were talking earlier about,
you know, one of the things that we've been talking internally quite a bit about that I still
I don't have an answer for it.
I don't know if anyone does.
But when you look at these benchmarks, you know,
you may see on a scale of one to 100,
somebody comes in an 82 and somebody comes in at 78.
And you're like, all right, well, I know 82's a higher number
than 78, so like they're quote-a-quote better.
But to the naked eye, does that actually a difference
that the human user can even tell?
What does that mean?
The four percentage points is kind of like a benchmark we invented, right?
And like, is it real?
Is it not?
Is it noticeable?
Does it improve accuracy?
or like, whatever the thing is.
And to me, you know, the work you guys are doing there,
but just more broadly as an industry, 10 years from now,
we'll probably have an excellent answer.
We'll be like how you...
But, Pam, I think the nuance to this,
and this is why the building, the platform,
and the transparency matters so much,
because all of the detail is in exactly what are the four tasks
that are different between those things.
And so here's a great example of this.
Like, if you haven't spent any time looking at benchmarks before,
Like, the more time you spend, the more you realize the things that we're measuring quality on and the things that people are talking about are fucking crazy.
Like, none of it makes sense.
Like, for example, and here's like one specific example.
There's some of these new coding benchmarks.
I won't name names because I said that's crazy.
Not a lot of disparage these folks because I think they're doing a reasonable job.
But like some of the new coding benchmarks have like 6% of tasks on like programming, like,
called ZIG, ZIG.
Nobody's ever heard of ZIG before.
This is not a programming language that any engineer at any company is actually using.
I'm sure some people are using it.
But the 6% difference could be like the quality on ZIG and like maybe some model
happened to get access to some data and whatever this language is.
But like that doesn't matter for 99.9% of startups.
Nobody cares about this thing.
And I'm sure they have a good reason for including that data.
But like being able to do this like,
introspection of like not just there's five percentage points difference between these two models,
but like, why is there a five percentage point? Does that actually matter for me as a business,
for me as a developer, for me as a user of this model? I think it's the whole game. And like the
products and surfaces and like even the benchmarks themselves don't do this right now. They sort
of show as in this empirical thing that, you know, your 75% should be the same way that I perceive
75%, which is completely not true. And so I think it's like a fundamental problem with the
way the things are set up right now. It does feel like on one hand, personalized benchmarks are going to
become a thing. Yeah, I don't know how somebody smarter than me will figure that out, but like that
obviously is going to be, you know, important, especially for businesses that are kind of like,
hey, here's my specific ramifications. The second thing, though, is we came out at Sylvia and we showed
that a lot of the harness work and things that we had done made the Sylvia product more accurate
than the frontier labs at answering tax-related questions.
And to me, I'm like, you know, more business-minded,
not as technical as the engineering team.
I'm like, great, you know, we are higher, we are better,
we are more accurate, you know, et cetera.
And immediately the engineering team was like,
we better publish the evals, you know, the rubric,
we better publish the user parameters.
Like there was this entire effort as to like,
how much can we publish an open source without actually giving away
things that would be considered, you know,
very important IP related type things.
And we went through a strategic kind of debate internally as to like,
there was definitely some stuff that we published that we could have not published
and it would have maybe given us a little bit more of an advantage.
But it was like if you're not a frontier model and you come out and you say that you're,
you know, more accurate on something, you almost have to like open source more to let people
validate it themselves.
And so I do think that if you're a frontier model, like you kind of don't care if people
believe you or not because you're just like,
you know, here's the evals, whatever, but the users care, right?
Like the kind of what you're talking about, I think, is a different situation.
And so it's less the academic application for, you know, who's got the best model.
And it's more about like, I'm a business or I'm a user.
I'm trying to evaluate which one of these things I should use for my specific use case.
I mean, the benchmarking industry is going to be, you know, significantly bigger than it is today.
And obviously, you guys have kind of a lead there in what you're doing.
Yeah.
Same thing with the data industry.
That's what's most interesting is all this, this data moment.
is, and we don't need to talk deeply about it, but like, it's having this, like, crazy,
I'm sure you're seeing this on the startup side. It's just, like, absolutely ridiculous.
And I think the framing of this is like, 2023, the question was like, does the recipe work?
Do we have the recipe to get to general intelligence to get to this sort of like AGI thing in the future?
And I think we de-risk the recipe and we've made a few tweaks over the last few years,
but like generally de-risk the recipe.
Then it was obvious like, oh, shit, there's not enough compute in the world.
let's blast hundreds of billions of dollars into getting compute online,
that is going to be the blocker.
And sort of like to keep scaling up, we need more compute, et cetera, et cetera.
All the labs have not done that.
We now have enough compute that's coming online.
There'll be further investment.
But generally people know that that's something that needs to be solved.
It's now all data bound.
Like the data to make progress on the model does not exist in the world.
Like it is data that has to actually be created net new that doesn't exist
or is coming from like even startups.
I think there's like a huge wave of like startups that are going into these like exclusive data licensing agreements with like data providers or model labs.
And like the, you know, if you want to make progress, it's all data bound.
And so it's like this like data business is very tied to this benchmark ecosystem is like very tied to ultimately model progress at the end of the day.
And so it's very interesting to see like how quickly these things are like spiking up into the right.
I am very biased, but I'm an investor in Micro 1.
I think they've done a fantastic job on the data side.
But I'm also an investor in a company called Sunset.
In Sunset, they started out as a company to help other companies shut down.
So if you have a startup, it doesn't work.
It's a pain in the ass, right?
You've got to get the lawyers involved.
You've got to figure out how do I save as much money as possible to get back to investors,
but I also have like, you know, kind of a responsible way to wind down.
And that's where they started.
I don't even know if they could spell AI at the time.
I loved them, but that was not their focus.
Well, all of a sudden, they realize, like, there is a unmonetized asset that these companies have,
which is, like, all the Slack messages in Google Drive, you know, just like the corporate data.
And could they basically, at the point of shutdown, buy that data from the company,
which creates a new asset that then can help them get more money back for their investors?
They have to clean it and structure it.
And, you know, kind of do all these things to make it a usable form.
But then they can turn around and they can then sell it to the model.
And so you almost have this like beautiful thing where like a normal company would never want to sell that data because they were worried about all the competitive components and all the stuff.
But if you're shutting your business down, you're like looking under the couch cushions for a couple of, you know, pennies.
You're like, hey, wherever we can find.
Oh, you want to buy our Slack messages?
We were just going to delete them.
So like here, knock yourself out.
And so to your point, like I do think that this has happened.
I've seen a ton of startups in all blue collar work, you know, medical, et cetera.
they're just trying to figure out how do we go and find data sets and no one else has yet,
and then turn around and let's go and use it for robotics, model training, whatever.
I don't know how big that thing can be, but it feels like we haven't even scratched the surface
of what that whole industry is going to look like.
It's going to be massive.
I think actually the most difficult part of this, and this is the part that's still like a dark
art, is like having data is not necessarily the problem at the moment.
The problem is, like, getting the data into a format that the model labs can actually use or that the data vendors can actually, some of the data vendors are now doing a bunch of this stuff.
But it's the most difficult part because the raw Slack messages, like, there's like an immense amount of work and labor that's involved in like taking that and finding some way to take that data and like make it actually usable from a model improvement perspective and like rigorously can go and like increase quality in some dimension.
So there's like, I think there's even just like that business of like helping companies understand how they can actually make what's the value of their data and all that stuff.
I think is a is a really, really difficult problem that it feels like we're still early in trying to solve.
And it's like fundamentally correlated with like we can do that.
We'll see more model progress.
100%.
I think my takeaway from this conversation, Google, the frontier commitment is real.
You guys got a lot of stuff going on.
I think you guys are doing a great job.
If people want to connect with you or find you online,
we're sure you send them.
X.
I'll see you on X.
Ping me.
I'm also on LinkedIn if you want to ask more boring questions.
All right, my friend.
Thank you very much for doing this.
We'll do it again in the future.
I love it.
Thank you for having me.
It was a fun conversation.
