This Week in Startups - What's Next for AI Infrastructure with Amin Vahdat | AI Basics with Google Cloud

Episode Date: May 1, 2025

In this episode of AI Basics, Jason sits down with Amin Vahdat, VP of ML at Google Cloud, to unpack the mind-blowing infrastructure behind modern AI. They dive into how Google’s TPUs power massive q...ueries, why 2025 is the “Year of Inference,” and how startups can now build what once felt impossible. From real-time agents to exponential speed gains, this is a look inside the AI engine that’s rewriting the future.*Timestamps:(0:00) Jason introduces today’s guest Amin Vahdat(3:18) Data movement implications for founders and historical bandwidth perspective(5:29) The shift to inference and AI infrastructure trends in startups and enterprises(8:40) Evolution of productivity and potential of low-code/no-code development(11:20) AI infrastructure pricing, cost efficiency, and historical innovation(17:53) Google's TPU technology and infrastructure scale(23:21) Building AI agents for startup evaluation and supervised associate agents(26:08) Documenting decisions for AI learning and early AI agent development*Uncover more valuable insights from AI leaders in Google Cloud's 'Future of AI: Perspectives for Startups' report. Discover what 23 AI industry leaders think about the future of AI—and how it impacts your business. Read their perspectives here: https://goo.gle/futureofai*Check out all of the Startup Basics episodes here: https://thisweekinstartups.com/basicsCheck out Google Cloud: https://cloud.google.com/*Follow Amin:LinkedIn: https://www.linkedin.com/in/vahdat/?trk=public_post_feed-actor-name*Follow Jason:X: https://twitter.com/JasonLinkedIn: https://www.linkedin.com/in/jasoncalacanis*Follow TWiST:Twitter: https://twitter.com/TWiStartupsYouTube: https://www.youtube.com/thisweekinInstagram: https://www.instagram.com/thisweekinstartupsTikTok: https://www.tiktok.com/@thisweekinstartupsSubstack: https://twistartups.substack.com

Transcript
Discussion (0)
Starting point is 00:00:03 Welcome back to another episode of AI basics brought to you in partnership with our friends at Google Cloud. Google Cloud just published a fantastic report. It's called The Future of AI Perspectives for Startups. It features insights from 23 leading AI experts. And today on the program, we're delighted to have Amin Vod. He's VP and GM of Machine Learning and Cloud AI over at Google Cloud. Amin Fadat has worked on scaling the internet for decades. And we're going to talk today about TPP.
Starting point is 00:00:33 GPUs, just how the infrastructure is scaling. I'm in, but I wanted to do it from the application layer backwards. Last night, I had my high fidelity headphones on. I was listening to some groovy music and I was like, tell me more artists like Bob Dylan and Mark Knopfler, give me an analysis by decade. And I realized I had sent Google Gemini's deep research to do essentially what would have been like a PhD, you know, thesis. Yes. And it returned it. And I would have never known about this music.
Starting point is 00:01:09 My God, the joy that that gave me finding new artists. And this is a silly example. But when I clicked on, like, show your work and I watched it work, I was astounded at how many threads were fired at once. So I want to help people conceptualize when you do a deep research and you know, switch your model to something like that. What's going on in the background on the hardware layer going through, I think what is the largest cluster of computers ever assembled, which is Google's cloud? It's a great question. And I mean, I'll start to take one step back that, you know, today,
Starting point is 00:01:44 when you do a, you know, quote unquote, normal web search, actually somewhere between a thousand and 10,000 standard computers, putting aside any AI mode or AI overview or any of the new features they're essentially collaborating together to comb over an index of the entire internet to give you good answers for your point, Corey. Now, the example that you're going to, though, it probably is the same 1,000 to 10,000 standard servers so that we've grown to love over the past couple of decades. But in addition to that, there is now a incredible pool of custom accelerators at Google for deep research.
Starting point is 00:02:22 It's going to be our TPUs, our tensor processing units. and essentially think of these as being able to pack in the general purpose computing power of 100 servers into one chip. Incredible. Now compose those out for your queries. I'm going to give you an estimate. There were probably 256 of these chips, each packing in 100 servers worth of compute, coordinating together.
Starting point is 00:02:46 So that would be, again, do the simple math, right? We're talking about tens of thousands of server equivalents, doing the work to, answer just a part of your question. And so for your specific example, the amazing thing is, and as you said, show your work, it's not just one query. It's not just one request and response. It is iterating in real time. So there might have been, depending on the complexity of your request, 10, 20, or 100 sub queries that have to get run and then compose together to give you a really good answer. Wow. So let's talk about how that data moves, because I think this will influence. This will inform what founders will think about building. So I'm trying to sort of open up their minds a little bit
Starting point is 00:03:30 of just what's capable today. We'll get into what's going to be what we'll be capable of next year and maybe in two or three years so we can skate to where the puck is going. But this is kind of mind-blowing how much compute you can put to work and at such an affordable level that it's almost like, you know, when you and I started our career, we were talking pre-show about floppy disk and stuff like that. If somebody told you, like, you have infinite bandwidth and infinite storage in the 80s or 90s, our minds probably wouldn't be able to comprehend that. So let's think about how this data moves around, how quickly it moves around. We know now that there's, you know, we're talking about thousands and thousands of what would be web servers at your disposable,
Starting point is 00:04:13 answering queries. How should a developer or even a founder who's a, you know, quote-unquote idea person, comprehend how to take this amount of compute, the amount of data storage, and the bandwidth we have available now, and turn that into a product. Yeah. Is there like an exercise we can do to open our minds to these possibilities? The exponentials here are astounding, and I did enjoy a pre-show conversation. I mean, an anecdote I'll share with you that you're well aware of, but, you know, in the 80s
Starting point is 00:04:43 and early 90s when the internet was in its early stages, the overhead of you signing your name to your email in terms of bandwidth was not. noticeable. Like in other words, hey, don't type too many characters because that might hurt the internet connection. Like, don't sign the Jason. Don't sign the mean. Now where are we? Right. Real-time video, you know, seven streams to your house in real-time. So really what I would say to founders is imagine the transformation. And the way that things are advancing, whether it's bandwidth, whether it's storage, whether it's raw compute capability, probably your imagination can be achieved. Now, in a particular day or month or year, you're going to have to think about making it efficient, but it is just exploding.
Starting point is 00:05:27 Really, levels I've never seen before. So a lot of the power now is making its way into inference. Or folks who don't know, you're using a lot of the infrastructure to build a model. Okay, people are building models all over the place. You can go on Hugging Face. You can see all the incredible activity there. But inference, being able to process queries faster, faster in real time. How is that changed just in the last year or two? And then how would that
Starting point is 00:05:55 change in the next year or two? Because, you know, we were sitting here, I don't know, it was 18 months ago. And back to bandwidth, it was writing a response. I'm moving my fingers to fully across the screen. Left to right. Like, you know, we haven't had that experience since dial up modems. Now it's just, boom, it just shows up the entire thousand words just appears on your screen instantly, which is kind of mind-blowing. But when we do these deep researches, you're watching it work and it's working in parallel. It's putting everything together. It might take a couple of minutes. And now there's just like really neat feature where it's like, hey, we're working on it.
Starting point is 00:06:30 We're going to send you an alert, which I think is just awesome, like a great little user feature there. So talk to me about inference, how that's changed in just the last 18 months. And then maybe we put our binoculars on here and where it will be in 18 months. Yeah, it's a good question. I mean, I call the moment that we're in right now, the age of inference. 2025 is the year of inference. And exactly, as you said, we've moved from the primary focus being about training and building the models to shifting that focus to serving the models and really making them super useful for our customers across the world. And as you said, we went from you being able to almost read faster than the happen to it appearing instantaneously.
Starting point is 00:07:11 So what that means now, though, is we're making the jobs harder. We're doing research. We're doing lots of queries. Sometimes it might take minutes. It's a time right now where people are really focused on efficiency. I mean, I can tell you it would not be uncommon here at Google for us to make things twice as fast than three months. Wow.
Starting point is 00:07:27 And then do it the next three months and then do it three months after that. And as you know, with exponential compounding, all of a sudden you have 10x and all of a sudden you have 20x and et cetera. So it's just and then the hardware gets faster than, you know, 12 months after that. So it's not just the software work. It's the hardware works, the management. So really think in terms of the capabilities of this model. is totally taking off.
Starting point is 00:07:49 And a lot of the work shifting from training time to thinking time, inference time, as you said, in response to your particular question. What are you seeing from startups right now and even enterprises in terms of the jobs they're throwing at this massive amount of infrastructure that's being built at a blistering pace? Yes. On, in some cases, you know, what we used to do commodity hardware. Now we're doing sort of premium hardware, but it will become commoditized. Again, I assume.
Starting point is 00:08:16 what kind of jobs are people throwing at the hardware now? And are there instances where they're overwhelming it or, you know, otherwise hitting a breaking point? Because I haven't heard from any startups. The bottleneck is now, you know, infrastructure and availability. That was like the discussion 12 months ago. And now I don't hear it anymore. You know, I think that we've, as I said, things are getting so much more efficient so quickly. I think infrastructure remains at a premium from the perspective of there's just so many companies
Starting point is 00:08:48 out there with so many great ideas. But what we're saying right this moment is around productivity and productivity starting with engineering productivity or employee productivity. The hottest areas right now are around code, software engineering. And in other words, the kinds of things that people are now able to do. I mean, the anecdote I like to share is somewhere around 12 months ago, these models had trouble counting the number of R's in strawberry. I asked the model how many R's in strawberry, and some of the models would get the answer wrong. Now they're generating working code,
Starting point is 00:09:19 and they're actually finding bugs in real complex systems. So in other words, the productivity boost for engineering disciplines, science disciplines, math, etc., really unbelievable. So the capabilities of these things are just taking off. So let's take a little detour here into development. startups have been constrained by a couple of things over the last couple of decades, and I've been investing in them for just over a decade, but, you know,
Starting point is 00:09:49 it was building them and reporting on them before that. And in the early days, you know, getting the money together, you know, three, four, five million dollars to launch a product and time. These were blockers. It took three or four million dollars and, call it 18, 24 months to get your product tomorrow. Cloud computing came out. People started working remote. There were more developers, but still developers became the blocker very quickly after people, we had cloud computing. Are we going to see what happened with
Starting point is 00:10:17 racking and stacking servers as the blocker? Then the blocker became the talent, the developers. Is that blocker going away? Do you believe in this vibe coding moment where the 96% of humans who don't write code, some percentage of them are going to be able to write code or speak English and have code written? Is that actually going to happen? And to what extent? I think more people are going to be able to write code for sure. But I think that actually, it really think of it as not needing fewer developers. It's multiplying the productivity and capability of the developers.
Starting point is 00:10:50 I mean, your analogy is spot on. In the end, we were limited by racking and stacking servers. Cloud computing came along. And all of a sudden, near infinite capacity became available to startups. Today, though, probably where we're constrained, actually, across the board for talent, whether it's large company or startup is the really transfer. formative engineering talents. Got it.
Starting point is 00:11:10 There's a finite set of them. Now, if we could make them more productive, which is the goal, I mean, it's going to be incredible because now you're not going to be limited by finding those people. You're going to be limited by your imagination. Which is pretty crazy. I want to talk, you know, brass hacks on pricing and costs. We're obviously seeing a lot of people standing up a lot of hardware and a lot of compute. How is the pricing dropping for, you know, the simple tasks?
Starting point is 00:11:37 you know, that, man, this hardware can kind of do queries, can do deep research. It seems like it's, I don't know what percentage the cost has dropped each of the last, say, two years, and then where you expect it to drop over the next two. But for things like storage, it seemed like storage would drop 5% a year, sometimes 10. It wasn't like some incredible pace where you went from a 1 terabyte hard drive to a 50 terabyte hard drive. In fact, that doesn't exist as a concept yet. I think maybe we're up to 14 terabytes. I don't know. People stop counting, which is a good indicator of how fast even that is moving.
Starting point is 00:12:12 But how fast is queries, tokens, how fast is that plummeting in costs? Yeah, so this is something that we're proud of at Google. In other words, there have been some external studies on this recently, where we are really driving the frontier in what we refer to as a unit cost of intelligence. And so what I mean by that is you pick your quality target. We probably have a model that hits that quality target. how much you pay normalized to the quality at Google for our models is at that frontier where you really can't push beyond it. So whether it's the lower cost fastest models to the higher cost, most sort of quality, highest quality models, we're leading there.
Starting point is 00:12:54 And it's because we are, as I said, 2X efficiency improvements in three months is not uncommon. We're, of course, passing all those savings on to our customers. So it's not 5% a year. it might be 300, you know, factor three reduction. Oh, my God. In a year. It's so interesting because we did have startups who are like, I don't think I can do that. And do that means some function, store this many videos, let people do this many queries, et cetera.
Starting point is 00:13:24 Again, 12 months ago, 24 months ago. And that question has stopped when they come talk to their investors. They're not like, can we get another million dollars because we're going to spend another 100,000 a month on this, they're there now not able to utilize the infrastructure as fast as is being built out or made more efficient. That's a very interesting. Have you ever seen that in our careers as technologists? It go from the infrastructure being the blocker to the idea people and the dreamers not being able to utilize the infrastructure and it just flip in 24 months? Not at this level. I mean, I think that's sort of back in the heyday of the early days
Starting point is 00:14:06 the Internet where actually you were talking about hard drives, even servers, there was a time where, as you remember, 12, 18 months past, things went down by a factor two. I mean, amazing, right? This was the hey, Dave Moore's Law in 2008, et cetera. Today, though 12 months pass, it's a factor 10. Right. So humans aren't very good at thinking in terms of those massive exponentials. But, you know, physics is not exponential, you know, in the human experience. Like, if, if you became a faster runner because you trained every day and you perfected your diet and you had the greatest trainer in the world, you know, you'd be shaving your 12 minute mile to 10 to 8 over, you know, whatever number of months or years to do that. We just, you're right. We don't actually think of the, we can't conceptualize these things. But one way to think about it is to look backwards at what happened during the period you're talking about. During the period you're talking about, I remember it very well because Chad Hurley was doing YouTube. Yes. And the price that you would pay for having a video go viral on the internet was whoever was hosting your video would turn off your server because you had used up your bandwidth allocation.
Starting point is 00:15:21 So posting a video meant you got a $3,000 hosting bill. You maxed it out. They turned it off. Yes. So it was this whole concept that video. Video hosting had to be paid. Then there was this little company Google that was like, hey, maybe we should start a second product. And Lori Park was like, hey, I got an invite for you, J-Cal.
Starting point is 00:15:39 Check out this Gmail. You get three invites yourself. Those three invites became worth like $1,000 at the peak because there was, I don't know, what was the first Gmail storage? Five gigs of storage. It was something crazy at that time. Maybe it was 10 gigs. At some point they said infinite storage, right?
Starting point is 00:15:56 And then at some point, again, whoop, infinite. Flickr. This incredible photo sharing app. If you just pause for a second as an entrepreneur and realize, Gmail Flickr YouTube were not considered viable businesses because of the cost posting. Exactly. And they had to charge users. And so they were gated by this thing. And then a bunch of entrepreneurs are like, you know what?
Starting point is 00:16:22 We're going to go unlimited and let's see what happens. And everybody thought YouTube would just wouldn't be fundable as a startup. And I think we know how that sort of went. So we can actually think about that ourselves. Yeah, I'm looking it up right now. One gigabyte. April 1st, 2004. I remember Lori giving me my invite back that one gigabyte.
Starting point is 00:16:41 That's right. It was crazy. You thought it was an April Fool's joke. It wasn't. It was launch. Yes. I think it was April 1st. Yeah, exactly.
Starting point is 00:16:48 It was. And the interesting thing is I was, I mean, I'm name dropping here like crazy, but I remember talking to Larry and I said, how can I delete emails in Gmail? He's like, you just archive. And I was like, wait, what do you mean? Like, what is archive? I mean, it's like, well, you just put it away, but it's still there. And I was like, that doesn't make sense.
Starting point is 00:17:03 You either delete it or you keep it. He's like this archive, new thing. And I said, no, I like to delete them. If I don't, if I read it, I processed it, it gets deleted. Then I don't have to worry about the storage. And he's like, you're not going to have to worry about storage anymore. Just archive it. And then he was like, why are you putting things in folders?
Starting point is 00:17:19 Because I was showing them how he was just a gym. I said, don't put it in folders. Just search. You're wasting your own time. And you know, you think about how prescient that is. That's how you have to think as an. entrepreneur now. Yes.
Starting point is 00:17:31 I'm trying to get out like what our business was that are constrained. As I would say as an entrepreneur, when you are saying this is impossible, actually write down the equation that leads you to think that.
Starting point is 00:17:42 And then divided by 10. Divided by 10. Divided by 100. Yes. And then is it still impossible? And maybe this, even when you divide it by 100, because probably in two years
Starting point is 00:17:51 you get to divided by 100. Yeah. So what is Google doing in the TPU space? You guys created this concept, correct? Like these were papers that came out of, I believe, Deep Mind or another unit. I don't know the history exactly. Google Brain, that's right, Transformers.
Starting point is 00:18:09 Right. So maybe you talk a little bit about the history of Transformers here, and then what Google is doing with TPUs, and why that's going to change compute in some major way over the coming years. Yeah, I mean, it's a story I love, and I think that it goes back to free Gen. Gen AI for sure. 2013, we did this thought exercise that said, as Larry and others, our founders, our want to do, there was this question that Jeff Dean, currently our chief scientist, formerly senior fellow,
Starting point is 00:18:41 one of our leading technical folks, he asked this question, what if every user of Google in 2013 wanted to interact with Google via voice for 30 seconds a day? Voice recognition would then need to run in our data centers for 30 seconds. it turned out that we would need to build two more Googles just to support that one use case, 30 seconds of voice in it.
Starting point is 00:19:03 And as you've noted, Google's already a pretty big infrastructure. So it's not like we were starting from a small base. We had one of the largest infrastructures in the world. We would have to triple it to support this one use case. Now, it turned out, though, that the operations needed for voice recognition were very predictable and regular, very large matrix multiplications. And so we realized that if we could, that CPUs only had so much efficiency in them. For general purpose, central processing units could only go so far.
Starting point is 00:19:31 So we essentially then invented the tensor processing unit that could do matrix multiplications a hundred times more efficiently. The work behind voice recognition than a general purpose computer. We built this very quickly, deployed it, and we enabled a use case that didn't, wasn't possible before. In other words, what we were really excited about is we made something that was impossible, possible. And this wound up going at scale for lots of different use cases. That was a serving inference case that we talked about. It went to training. It then at least partially contributed to a breakthrough like transformers because we had so much computing power available
Starting point is 00:20:11 to us. And that was a real enabler. It's taken off. We've had seven generations of TPUs since then. And each one is not just 5% better than the last. This current one, Are current TPU pods 10x more capable than the previous generation TPU? Again, not just 5%, 10%, 15%, or even two times, 10 times. Right. So, again, it's just exploding. It's so crazy to think about just when you're, and this is one of the great things about, you know, companies reaching a billion users.
Starting point is 00:20:42 And listen, a number of companies have done it. And I think Google has six products that have reached a billion users each. Gmail search, Chrome, Android. YouTube. Has Docs? Has Google Docs? And that Sweet hit a billion? I don't know.
Starting point is 00:20:58 I don't have it out of my fingertips, but I think many of those, by the way, are two billion. In other words, talk about doublings. It's really continuing to take off. And so this is where, like, you know, you do hit a roadblock
Starting point is 00:21:10 because you're not just releasing something to 10 million or 1 million or 100 million people. You have to actually solve it for the 2 billion people using those really interesting problems. Williams Gibson said something really interesting in one of his books. The street finds its own use for technology, right? Exactly.
Starting point is 00:21:29 We build things as technologists. It goes out there and then they decide what's going to happen. What are you seeing out there on the streets? You guys get a unique view since you're providing to, you know, I guess it's low millions of startups or, you know, enterprises now with these tools. and on consumer basis, you know, billions, what do you see people do that you didn't expect or that is just kind of weird, odd? You know, that's what I look for as an entrepreneur and as an investor.
Starting point is 00:22:03 I look for those projects that people go, non-consensus, never going to work. Why would people need a meditation app? We invested in calm. Nobody needs their own personal driver. You're doing a car service company, Uber, a stock trading app that doesn't charge people Robin Hood. Like these I, Airbnb we weren't in, but that was a weird one, you know, millions of people going to sleep on people's couches. What are you seeing out there that's weird, odd, or otherworldly, you don't have to name the company, but just things people are doing
Starting point is 00:22:29 with the tools. Yeah, I don't know if I see a lot weird. I see a lot exciting. And I would say that the big thing that I see very exciting is how things are progressing with agents. In other words, now to your point of, hey, I'm going to come back and give you an answer in a couple of minutes. These agents are now able to actually invoke code, invoke perhaps interact with other agents. So the scope of what AI can do is going, again, far beyond left to right to not just, I'm going to generate lots of answers,
Starting point is 00:23:01 but I'm actually going to take action based on some of those answers on your behalf. That's going to be fun. Yeah, exactly. And so I think that this is going to be more in the early, but not so early stages of that, think that we're going to see these agents really explode, and the creativity behind them is also really, really heartening. I am studying everybody in our venture firm on our podcasting teams.
Starting point is 00:23:27 I study what they do that's repetitive. And I say to them, ADD, automate, deprecate, delegate. Like, what? Yeah, like, do we have, is there a reason we're doing this? Let's just have that fundamental question. If we're not doing, just deprecate it. Okay, now we're left with, like, can we automate it, or can we delegate it? We'll use an Athena assistant, you know, somebody in the lowest cost place in the world, you know, where they have great college educated people who can do it, or an outside firm. No, delegate our accounting legal to outsource other firm. So then that leaves automate. And really interesting how some of the young people have working for me have got their heads around this already. Yes. And the simple thing, I'll give you but one example. We have 20,000 people
Starting point is 00:24:15 apply for funding. Wow. Crazy. It's like second only to Y Combinators program. There's a lot of people coming in and they send us a bunch of information and then we have literally seven researchers
Starting point is 00:24:25 sit there, full-time people, and categorize the startup, make sure that what they tell us they're doing, oh, it's a SaaS company, it's marketplace company, one founder, it's three founders. And I have a system like, if they have these 13 qualities,
Starting point is 00:24:37 that's what I'm up to, these are like reasons to get excited or maybe get more curious about the startup. So we're building an agent now who we're kind of architect. protecting, how do we get as clean information in, clean up the information we have, check it with other sources, look at other competitive startups, and build the dossier of this startup. You know, basically the deal memo, it's actually getting close, which means we can process
Starting point is 00:25:03 more companies and not miss companies, which is really sins of omissions in venture. If you miss Google, you miss YouTube, you miss Instagram, whatever you missed, defines your career. And so how close am I to having this associate agent? A supervised associate agent like this, you're very close. And in fact, it's not just you give it the 13 categories that you're interested in. It'll come back to you and say, you know what? I think you might have missed three. Yeah, blind spots, right?
Starting point is 00:25:31 Yeah. And I look back at all the successful deals you've had and the less successful deals you've had. Here's what I found that you might not have thought about. What I love about that is, you know, the AI is, is going to be brutally can. Yep, exactly. There's not going to be like, I wonder if this is going to hurt his feelings that like,
Starting point is 00:25:50 we actually had this incredible company, you know, apply and we missed it. You know, somebody might not tell me that because they don't want to make me feel bad. Yeah, I was going to be like, hey, dummy. I missed it. And here's why. And here's why. Yeah, because this is, there's a really good book, super forecasting. I don't know if you ever read it, but basically tells you how to become a forecaster.
Starting point is 00:26:11 And one of the key things is writing down why you made the decision right as you're making the decision and after you've made it. So if AI can just bring us the stuff, tell it the decision it would make, and then we make our decision, the reinforcement learning that could occur. Exactly. But we're only in what, the second inning of agents you think if it was a nine inning game? First or second, for sure. Yep. First or second is where I would put it. Yeah.
Starting point is 00:26:36 Listen, this has been amazing. I got to have you on the pot again. We love it. Anytime you got like three or four. for really interesting use cases that you guys discover. Come back on and let's break them down. This has been amazing. Thanks again to Amin Vodat for joining us here on the AI Basic series.
Starting point is 00:26:53 Go to this week in startups.com slash basics. You'll see all the basic series in one place. For more insights, take a minute. Learn some more. Go to Google Cloud's Future of AI Perspectus for Startups Report. Let me give you the URL. It's also in the show notes and everywhere. G-O-O-D-G-L-E slash Future of AI.
Starting point is 00:27:11 That's that Google short URL, G-O.glo.g.Le. slash future of AI, you're going to get predictions, real-world examples, startup advice, and you're going to discover what the top AI leaders have to say about the future of AI and its impact on your business. Thanks again for listening, and we'll see you next time on AI basics.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.