This Week in Startups - What's Next for AI Infrastructure with Amin Vahdat | AI Basics with Google Cloud
Episode Date: May 1, 2025In this episode of AI Basics, Jason sits down with Amin Vahdat, VP of ML at Google Cloud, to unpack the mind-blowing infrastructure behind modern AI. They dive into how Google’s TPUs power massive q...ueries, why 2025 is the “Year of Inference,” and how startups can now build what once felt impossible. From real-time agents to exponential speed gains, this is a look inside the AI engine that’s rewriting the future.*Timestamps:(0:00) Jason introduces today’s guest Amin Vahdat(3:18) Data movement implications for founders and historical bandwidth perspective(5:29) The shift to inference and AI infrastructure trends in startups and enterprises(8:40) Evolution of productivity and potential of low-code/no-code development(11:20) AI infrastructure pricing, cost efficiency, and historical innovation(17:53) Google's TPU technology and infrastructure scale(23:21) Building AI agents for startup evaluation and supervised associate agents(26:08) Documenting decisions for AI learning and early AI agent development*Uncover more valuable insights from AI leaders in Google Cloud's 'Future of AI: Perspectives for Startups' report. Discover what 23 AI industry leaders think about the future of AI—and how it impacts your business. Read their perspectives here: https://goo.gle/futureofai*Check out all of the Startup Basics episodes here: https://thisweekinstartups.com/basicsCheck out Google Cloud: https://cloud.google.com/*Follow Amin:LinkedIn: https://www.linkedin.com/in/vahdat/?trk=public_post_feed-actor-name*Follow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanis*Follow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com
Transcript
Discussion (0)
Welcome back to another episode of AI basics brought to you in partnership with our friends at Google Cloud.
Google Cloud just published a fantastic report.
It's called The Future of AI Perspectives for Startups.
It features insights from 23 leading AI experts.
And today on the program, we're delighted to have Amin Vod.
He's VP and GM of Machine Learning and Cloud AI over at Google Cloud.
Amin Fadat has worked on scaling the internet for decades.
And we're going to talk today about TPP.
GPUs, just how the infrastructure is scaling. I'm in, but I wanted to do it from the application
layer backwards. Last night, I had my high fidelity headphones on. I was listening to some
groovy music and I was like, tell me more artists like Bob Dylan and Mark Knopfler, give me an
analysis by decade. And I realized I had sent Google Gemini's deep research to do essentially
what would have been like a PhD, you know, thesis.
Yes.
And it returned it.
And I would have never known about this music.
My God, the joy that that gave me finding new artists.
And this is a silly example.
But when I clicked on, like, show your work and I watched it work, I was astounded
at how many threads were fired at once.
So I want to help people conceptualize when you do a deep research and you know, switch
your model to something like that. What's going on in the background on the hardware layer going through,
I think what is the largest cluster of computers ever assembled, which is Google's cloud?
It's a great question. And I mean, I'll start to take one step back that, you know, today,
when you do a, you know, quote unquote, normal web search, actually somewhere between a thousand
and 10,000 standard computers, putting aside any AI mode or AI overview or any of the new features
they're essentially collaborating together to comb over an index of the entire internet to give you
good answers for your point, Corey.
Now, the example that you're going to, though, it probably is the same 1,000 to 10,000 standard
servers so that we've grown to love over the past couple of decades.
But in addition to that, there is now a incredible pool of custom accelerators at Google for
deep research.
It's going to be our TPUs, our tensor processing units.
and essentially think of these as being able to pack in the general purpose computing power of 100 servers
into one chip.
Incredible.
Now compose those out for your queries.
I'm going to give you an estimate.
There were probably 256 of these chips, each packing in 100 servers worth of compute,
coordinating together.
So that would be, again, do the simple math, right?
We're talking about tens of thousands of server equivalents, doing the work to,
answer just a part of your question. And so for your specific example, the amazing thing is,
and as you said, show your work, it's not just one query. It's not just one request and response.
It is iterating in real time. So there might have been, depending on the complexity of your request,
10, 20, or 100 sub queries that have to get run and then compose together to give you a really good
answer. Wow. So let's talk about how that data moves, because I think this will influence. This will
inform what founders will think about building. So I'm trying to sort of open up their minds a little bit
of just what's capable today. We'll get into what's going to be what we'll be capable of next year
and maybe in two or three years so we can skate to where the puck is going. But this is
kind of mind-blowing how much compute you can put to work and at such an affordable level that it's
almost like, you know, when you and I started our career, we were talking pre-show about floppy disk
and stuff like that. If somebody told you, like, you have infinite bandwidth and infinite storage
in the 80s or 90s, our minds probably wouldn't be able to comprehend that. So let's think about
how this data moves around, how quickly it moves around. We know now that there's, you know,
we're talking about thousands and thousands of what would be web servers at your disposable,
answering queries. How should a developer or even a founder who's a, you know, quote-unquote idea
person, comprehend how to take
this amount of compute, the amount of data storage, and the bandwidth we have available now,
and turn that into a product.
Yeah.
Is there like an exercise we can do to open our minds to these possibilities?
The exponentials here are astounding, and I did enjoy a pre-show conversation.
I mean, an anecdote I'll share with you that you're well aware of, but, you know, in the 80s
and early 90s when the internet was in its early stages, the overhead of you signing your name
to your email in terms of bandwidth was not.
noticeable. Like in other words, hey, don't type too many characters because that might hurt the
internet connection. Like, don't sign the Jason. Don't sign the mean. Now where are we?
Right. Real-time video, you know, seven streams to your house in real-time. So really what I would say
to founders is imagine the transformation. And the way that things are advancing, whether it's bandwidth,
whether it's storage, whether it's raw compute capability, probably your imagination can be achieved.
Now, in a particular day or month or year, you're going to have to think about making it efficient, but it is just exploding.
Really, levels I've never seen before.
So a lot of the power now is making its way into inference.
Or folks who don't know, you're using a lot of the infrastructure to build a model.
Okay, people are building models all over the place.
You can go on Hugging Face.
You can see all the incredible activity there.
But inference, being able to process queries faster,
faster in real time. How is that changed just in the last year or two? And then how would that
change in the next year or two? Because, you know, we were sitting here, I don't know, it was 18 months
ago. And back to bandwidth, it was writing a response. I'm moving my fingers to fully across the screen.
Left to right. Like, you know, we haven't had that experience since dial up modems. Now it's just,
boom, it just shows up the entire thousand words just appears on your screen instantly, which is kind of mind-blowing.
But when we do these deep researches, you're watching it work and it's working in parallel.
It's putting everything together.
It might take a couple of minutes.
And now there's just like really neat feature where it's like, hey, we're working on it.
We're going to send you an alert, which I think is just awesome, like a great little user feature there.
So talk to me about inference, how that's changed in just the last 18 months.
And then maybe we put our binoculars on here and where it will be in 18 months.
Yeah, it's a good question.
I mean, I call the moment that we're in right now, the age of inference.
2025 is the year of inference.
And exactly, as you said, we've moved from the primary focus being about training and building the models to shifting that focus to serving the models and really making them super useful for our customers across the world.
And as you said, we went from you being able to almost read faster than the happen to it appearing instantaneously.
So what that means now, though, is we're making the jobs harder.
We're doing research.
We're doing lots of queries.
Sometimes it might take minutes.
It's a time right now where people are really focused on efficiency.
I mean, I can tell you it would not be uncommon here at Google for us to make things twice as fast
than three months.
Wow.
And then do it the next three months and then do it three months after that.
And as you know, with exponential compounding, all of a sudden you have 10x and all of a sudden
you have 20x and et cetera.
So it's just and then the hardware gets faster than, you know, 12 months after that.
So it's not just the software work.
It's the hardware works, the management.
So really think in terms of the capabilities of this model.
is totally taking off.
And a lot of the work shifting from training time to thinking time, inference time,
as you said, in response to your particular question.
What are you seeing from startups right now and even enterprises in terms of the jobs
they're throwing at this massive amount of infrastructure that's being built at a blistering pace?
Yes.
On, in some cases, you know, what we used to do commodity hardware.
Now we're doing sort of premium hardware, but it will become commoditized.
Again, I assume.
what kind of jobs are people throwing at the hardware now?
And are there instances where they're overwhelming it or, you know, otherwise hitting a breaking point?
Because I haven't heard from any startups.
The bottleneck is now, you know, infrastructure and availability.
That was like the discussion 12 months ago.
And now I don't hear it anymore.
You know, I think that we've, as I said, things are getting so much more efficient so quickly.
I think infrastructure remains at a premium from the perspective of there's just so many companies
out there with so many great ideas. But what we're saying right this moment is around productivity
and productivity starting with engineering productivity or employee productivity. The hottest
areas right now are around code, software engineering. And in other words, the kinds of things that
people are now able to do. I mean, the anecdote I like to share is somewhere around 12 months ago,
these models had trouble counting the number of R's in strawberry.
I asked the model how many R's in strawberry,
and some of the models would get the answer wrong.
Now they're generating working code,
and they're actually finding bugs in real complex systems.
So in other words, the productivity boost for engineering disciplines,
science disciplines, math, etc., really unbelievable.
So the capabilities of these things are just taking off.
So let's take a little detour here into development.
startups have been constrained by a couple of things over the last couple of decades,
and I've been investing in them for just over a decade,
but, you know,
it was building them and reporting on them before that.
And in the early days, you know, getting the money together, you know,
three, four, five million dollars to launch a product and time.
These were blockers.
It took three or four million dollars and, call it 18, 24 months to get your product tomorrow.
Cloud computing came out.
People started working remote. There were more developers, but still developers became the blocker
very quickly after people, we had cloud computing. Are we going to see what happened with
racking and stacking servers as the blocker? Then the blocker became the talent, the developers.
Is that blocker going away? Do you believe in this vibe coding moment where the 96% of humans
who don't write code, some percentage of them are going to be able to write code or speak English
and have code written? Is that actually going to happen?
And to what extent?
I think more people are going to be able to write code for sure.
But I think that actually, it really think of it as not needing fewer developers.
It's multiplying the productivity and capability of the developers.
I mean, your analogy is spot on.
In the end, we were limited by racking and stacking servers.
Cloud computing came along.
And all of a sudden, near infinite capacity became available to startups.
Today, though, probably where we're constrained, actually, across the board for talent,
whether it's large company or startup is the really transfer.
formative engineering talents.
Got it.
There's a finite set of them.
Now, if we could make them more productive, which is the goal, I mean, it's going to be
incredible because now you're not going to be limited by finding those people.
You're going to be limited by your imagination.
Which is pretty crazy.
I want to talk, you know, brass hacks on pricing and costs.
We're obviously seeing a lot of people standing up a lot of hardware and a lot of compute.
How is the pricing dropping for, you know, the simple tasks?
you know, that, man, this hardware can kind of do queries, can do deep research. It seems like
it's, I don't know what percentage the cost has dropped each of the last, say, two years,
and then where you expect it to drop over the next two. But for things like storage,
it seemed like storage would drop 5% a year, sometimes 10. It wasn't like some incredible pace
where you went from a 1 terabyte hard drive to a 50 terabyte hard drive. In fact, that doesn't
exist as a concept yet. I think maybe we're up to 14 terabytes.
I don't know.
People stop counting, which is a good indicator of how fast even that is moving.
But how fast is queries, tokens, how fast is that plummeting in costs?
Yeah, so this is something that we're proud of at Google.
In other words, there have been some external studies on this recently, where we are really
driving the frontier in what we refer to as a unit cost of intelligence.
And so what I mean by that is you pick your quality target.
We probably have a model that hits that quality target.
how much you pay normalized to the quality at Google for our models is at that frontier where you really can't push beyond it.
So whether it's the lower cost fastest models to the higher cost, most sort of quality, highest quality models, we're leading there.
And it's because we are, as I said, 2X efficiency improvements in three months is not uncommon.
We're, of course, passing all those savings on to our customers.
So it's not 5% a year.
it might be 300, you know, factor three reduction.
Oh, my God.
In a year.
It's so interesting because we did have startups who are like, I don't think I can do that.
And do that means some function, store this many videos, let people do this many queries, et cetera.
Again, 12 months ago, 24 months ago.
And that question has stopped when they come talk to their investors.
They're not like, can we get another million dollars because we're going to spend
another 100,000 a month on this, they're there now not able to utilize the infrastructure
as fast as is being built out or made more efficient. That's a very interesting. Have you ever seen
that in our careers as technologists? It go from the infrastructure being the blocker to the
idea people and the dreamers not being able to utilize the infrastructure and it just flip in 24
months? Not at this level. I mean, I think that's sort of back in the heyday of the early days
the Internet where actually you were talking about hard drives, even servers, there was a time
where, as you remember, 12, 18 months past, things went down by a factor two. I mean, amazing,
right? This was the hey, Dave Moore's Law in 2008, et cetera. Today, though 12 months pass, it's a factor
10. Right. So humans aren't very good at thinking in terms of those massive exponentials.
But, you know, physics is not exponential, you know, in the human experience. Like, if, if you became a faster runner because you trained every day and you perfected your diet and you had the greatest trainer in the world, you know, you'd be shaving your 12 minute mile to 10 to 8 over, you know, whatever number of months or years to do that. We just, you're right. We don't actually think of the, we can't conceptualize these things. But one way to think about it is to look backwards at what happened during the period you're talking about.
During the period you're talking about, I remember it very well because Chad Hurley was doing YouTube.
Yes.
And the price that you would pay for having a video go viral on the internet was whoever was hosting your video would turn off your server because you had used up your bandwidth allocation.
So posting a video meant you got a $3,000 hosting bill.
You maxed it out.
They turned it off.
Yes.
So it was this whole concept that video.
Video hosting had to be paid.
Then there was this little company Google that was like, hey, maybe we should start a second product.
And Lori Park was like, hey, I got an invite for you, J-Cal.
Check out this Gmail.
You get three invites yourself.
Those three invites became worth like $1,000 at the peak because there was, I don't know,
what was the first Gmail storage?
Five gigs of storage.
It was something crazy at that time.
Maybe it was 10 gigs.
At some point they said infinite storage, right?
And then at some point, again, whoop, infinite.
Flickr.
This incredible photo sharing app.
If you just pause for a second as an entrepreneur and realize, Gmail Flickr YouTube were not considered viable businesses because of the cost posting.
Exactly.
And they had to charge users.
And so they were gated by this thing.
And then a bunch of entrepreneurs are like, you know what?
We're going to go unlimited and let's see what happens.
And everybody thought YouTube would just wouldn't be fundable as a startup.
And I think we know how that sort of went.
So we can actually think about that ourselves.
Yeah, I'm looking it up right now.
One gigabyte.
April 1st, 2004.
I remember Lori giving me my invite back that one gigabyte.
That's right.
It was crazy.
You thought it was an April Fool's joke.
It wasn't.
It was launch.
Yes.
I think it was April 1st.
Yeah, exactly.
It was.
And the interesting thing is I was, I mean, I'm name dropping here like crazy,
but I remember talking to Larry and I said, how can I delete emails in Gmail?
He's like, you just archive.
And I was like, wait, what do you mean?
Like, what is archive?
I mean, it's like, well, you just put it away, but it's still there.
And I was like, that doesn't make sense.
You either delete it or you keep it.
He's like this archive, new thing.
And I said, no, I like to delete them.
If I don't, if I read it, I processed it, it gets deleted.
Then I don't have to worry about the storage.
And he's like, you're not going to have to worry about storage anymore.
Just archive it.
And then he was like, why are you putting things in folders?
Because I was showing them how he was just a gym.
I said, don't put it in folders.
Just search.
You're wasting your own time.
And you know, you think about how prescient that is.
That's how you have to think as an.
entrepreneur now.
Yes.
I'm trying to get out
like what our business
was that are constrained.
As I would say as an entrepreneur,
when you are saying
this is impossible,
actually write down the equation
that leads you to think that.
And then divided by 10.
Divided by 10.
Divided by 100.
Yes.
And then is it still impossible?
And maybe this,
even when you divide it by 100,
because probably in two years
you get to divided by 100.
Yeah.
So what is Google doing
in the TPU space?
You guys created this
concept, correct? Like these were papers that came out of, I believe, Deep Mind or another unit.
I don't know the history exactly.
Google Brain, that's right, Transformers.
Right. So maybe you talk a little bit about the history of Transformers here,
and then what Google is doing with TPUs, and why that's going to change compute in some major way over the coming years.
Yeah, I mean, it's a story I love, and I think that it goes back to free Gen.
Gen AI for sure.
2013, we did this thought exercise that said,
as Larry and others, our founders, our want to do,
there was this question that Jeff Dean,
currently our chief scientist, formerly senior fellow,
one of our leading technical folks,
he asked this question,
what if every user of Google in 2013
wanted to interact with Google via voice for 30 seconds a day?
Voice recognition would then need to run in our data centers
for 30 seconds.
it turned out that we would need to build two more Googles
just to support that one use case, 30 seconds of voice in it.
And as you've noted, Google's already a pretty big infrastructure.
So it's not like we were starting from a small base.
We had one of the largest infrastructures in the world.
We would have to triple it to support this one use case.
Now, it turned out, though, that the operations needed for voice recognition
were very predictable and regular, very large matrix multiplications.
And so we realized that if we could, that CPUs only had so much efficiency in them.
For general purpose, central processing units could only go so far.
So we essentially then invented the tensor processing unit that could do matrix multiplications
a hundred times more efficiently.
The work behind voice recognition than a general purpose computer.
We built this very quickly, deployed it, and we enabled a use case that didn't, wasn't
possible before. In other words, what we were really excited about is we made something that was
impossible, possible. And this wound up going at scale for lots of different use cases. That
was a serving inference case that we talked about. It went to training. It then at least partially
contributed to a breakthrough like transformers because we had so much computing power available
to us. And that was a real enabler. It's taken off. We've had seven generations of TPUs since then.
And each one is not just 5% better than the last. This current one,
Are current TPU pods 10x more capable than the previous generation TPU?
Again, not just 5%, 10%, 15%, or even two times, 10 times.
Right.
So, again, it's just exploding.
It's so crazy to think about just when you're, and this is one of the great things about,
you know, companies reaching a billion users.
And listen, a number of companies have done it.
And I think Google has six products that have reached a billion users each.
Gmail search, Chrome, Android.
YouTube.
Has Docs?
Has Google Docs?
And that Sweet hit a billion?
I don't know.
I don't have it out of my fingertips,
but I think many of those,
by the way, are two billion.
In other words,
talk about doublings.
It's really continuing to take off.
And so this is where, like,
you know, you do hit a roadblock
because you're not just releasing something
to 10 million or 1 million or 100 million people.
You have to actually solve it
for the 2 billion people using those
really interesting problems.
Williams Gibson said something really interesting in one of his books.
The street finds its own use for technology, right?
Exactly.
We build things as technologists.
It goes out there and then they decide what's going to happen.
What are you seeing out there on the streets?
You guys get a unique view since you're providing to, you know,
I guess it's low millions of startups or, you know, enterprises now with these tools.
and on consumer basis, you know, billions, what do you see people do that you didn't expect
or that is just kind of weird, odd?
You know, that's what I look for as an entrepreneur and as an investor.
I look for those projects that people go, non-consensus, never going to work.
Why would people need a meditation app?
We invested in calm.
Nobody needs their own personal driver.
You're doing a car service company, Uber, a stock trading app that doesn't
charge people Robin Hood. Like these I, Airbnb we weren't in, but that was a weird one,
you know, millions of people going to sleep on people's couches. What are you seeing out there that's
weird, odd, or otherworldly, you don't have to name the company, but just things people are doing
with the tools. Yeah, I don't know if I see a lot weird. I see a lot exciting. And I would say that the
big thing that I see very exciting is how things are progressing with agents. In other words,
now to your point of, hey, I'm going to come back and give you an answer in a couple of minutes.
These agents are now able to actually invoke code,
invoke perhaps interact with other agents.
So the scope of what AI can do is going, again,
far beyond left to right to not just,
I'm going to generate lots of answers,
but I'm actually going to take action based on some of those answers
on your behalf.
That's going to be fun.
Yeah, exactly.
And so I think that this is going to be more in the early,
but not so early stages of that,
think that we're going to see these agents really explode, and the creativity behind them is also
really, really heartening. I am studying everybody in our venture firm on our podcasting teams.
I study what they do that's repetitive. And I say to them, ADD, automate, deprecate, delegate.
Like, what? Yeah, like, do we have, is there a reason we're doing this? Let's just have that
fundamental question. If we're not doing, just deprecate it. Okay, now we're left with, like, can we
automate it, or can we delegate it? We'll use an Athena assistant, you know, somebody in the
lowest cost place in the world, you know, where they have great college educated people who can do it,
or an outside firm. No, delegate our accounting legal to outsource other firm. So then that leaves
automate. And really interesting how some of the young people have working for me have got their
heads around this already. Yes. And the simple thing, I'll give you but one example. We have 20,000 people
apply for funding.
Wow.
Crazy.
It's like second only
to Y Combinators program.
There's a lot of people coming in
and they send us a bunch of information
and then we have literally seven researchers
sit there, full-time people,
and categorize the startup,
make sure that what they tell us they're doing,
oh, it's a SaaS company,
it's marketplace company,
one founder, it's three founders.
And I have a system like,
if they have these 13 qualities,
that's what I'm up to,
these are like reasons to get excited
or maybe get more curious about the startup.
So we're building an agent now
who we're kind of architect.
protecting, how do we get as clean information in, clean up the information we have, check it
with other sources, look at other competitive startups, and build the dossier of this startup.
You know, basically the deal memo, it's actually getting close, which means we can process
more companies and not miss companies, which is really sins of omissions in venture.
If you miss Google, you miss YouTube, you miss Instagram, whatever you missed, defines your career.
And so how close am I to having this associate agent?
A supervised associate agent like this, you're very close.
And in fact, it's not just you give it the 13 categories that you're interested in.
It'll come back to you and say, you know what?
I think you might have missed three.
Yeah, blind spots, right?
Yeah.
And I look back at all the successful deals you've had and the less successful deals you've had.
Here's what I found that you might not have thought about.
What I love about that is, you know, the AI is,
is going to be brutally can.
Yep, exactly.
There's not going to be like,
I wonder if this is going to hurt his feelings that like,
we actually had this incredible company, you know, apply and we missed it.
You know, somebody might not tell me that because they don't want to make me feel bad.
Yeah, I was going to be like, hey, dummy.
I missed it.
And here's why.
And here's why.
Yeah, because this is, there's a really good book, super forecasting.
I don't know if you ever read it, but basically tells you how to become a forecaster.
And one of the key things is writing down why you made the decision right as you're making the decision and after you've made it.
So if AI can just bring us the stuff, tell it the decision it would make, and then we make our decision, the reinforcement learning that could occur.
Exactly.
But we're only in what, the second inning of agents you think if it was a nine inning game?
First or second, for sure.
Yep.
First or second is where I would put it.
Yeah.
Listen, this has been amazing.
I got to have you on the pot again.
We love it.
Anytime you got like three or four.
for really interesting use cases that you guys discover.
Come back on and let's break them down.
This has been amazing.
Thanks again to Amin Vodat for joining us here on the AI Basic series.
Go to this week in startups.com slash basics.
You'll see all the basic series in one place.
For more insights, take a minute.
Learn some more.
Go to Google Cloud's Future of AI Perspectus for Startups Report.
Let me give you the URL.
It's also in the show notes and everywhere.
G-O-O-D-G-L-E slash Future of AI.
That's that Google short URL, G-O.glo.g.Le. slash future of AI, you're going to get predictions, real-world examples, startup advice, and you're going to discover what the top AI leaders have to say about the future of AI and its impact on your business. Thanks again for listening, and we'll see you next time on AI basics.
