Odd Lots - Inside the Battle for Chips That Will Power Artificial Intelligence
Episode Date: May 8, 2023Nobody knows for sure who is going to make all the money when it comes to artificial intelligence. Will it be the incumbent tech giants? Will it be startups? What will the business models look like? I...t's all up in the air. One thing is clear though — AI requires a lot of computing power and that means demand for semiconductors. Right now, Nvidia has been a huge winner in the space, with their chips powering both the training of AI models (like ChatGPT) and the inference (the results of a query.) But others want in on the action as well. So how big will this market be? Can other companies gain a foothold and "chip away" at Nvidia's dominance? On this episode we speak with Bernstein semiconductor analyst Stacy Rasgon about this rapidly growing space and who has a shot to win it.See omnystudio.com/listener for privacy information.
Transcript
Discussion (0)
Thanks for listening to OddLots. Follow the show on Amazon Music for more future episodes or just ask Alexa play the podcast, OddLots on Amazon Music.
Today's show is brought to you by Vanguard. To all the financial advisors listening, let's talk bonds for a minute.
Capturing value and fixed income is not easy. Bond markets are massive, murky, and let's be real. Lots of firms throw a couple flashy funds your way and call it a day.
But not Vanguard. At Vanguard, institutional quality isn't a tagline. It's a very big. It's a firm. It's a firm. It's a couple flashy funds. It's a lot. It's
commitment to your clients. We're talking top grade products across the board of over 80 bond funds,
actively managed by a 200-person global squad of sector specialists, analysts, and traders.
These folks live and breathe fixed income. So if you're looking to give your clients consistent
results year in and year out, go see the record for yourself at vanguard.com slash audio.
That's vanguard.com slash audio. All investing is subject to risk vanguard marketing corporation
distributor.
Hello and welcome to another episode of the Oblop's podcast. I'm Joe Wisenthal.
And I'm Tracy Allo.
Tracy, I'm not sure if you've heard anyone talking about it or anything, but have you heard
about like this sort of AI thing people have been discussed?
Oh, you know what? I discovered this really cool new thing called Chat GPT.
Oh yeah, I saw that website too.
Yeah. Have you tried it?
I tried to, yeah, I had to like write a poem for me.
She's a pretty cool technology.
We should probably learn more about it.
Yeah, I think we should. No, okay. All right. Obviously, we're being facetious and joking, but everyone has been talking about AI and these new sort of natural language interfaces that allow you to ask questions or generate all different types of texts and things like that. It feels like everyone is very excited about that space.
Every, like, almost every conversation. To put it mildly.
Like, I went out with some friends that I hadn't seen in a long time. Like, I was in a bar.
last night. And like the conversation like turned to AI within like two minutes. And I've got to
talk about the experiments they did. But yes, there is a lot. It's basically like this like wall of
noise. And everyone's been talking about actually but us because I don't think we have done as far as I
can recall like an AI episode. We don't want to just add to the noise and get another sort of
tune stroke around. But obviously there's a lot there for us to discuss. Totally. And I'm sure this
will be the first of many episodes, but one of the ways that it fits into sort of classic
odd lots lore is via semiconductors, right? If you think about what chat GPT, for instance, is doing,
it's taking words and transforming them into numbers and then spitting those words back out at you.
And the thing that enables it to do that, semiconductors, chips.
Right. So here's like the four things I think I know about this.
And so this is that A, training the AI models so that they can do that is a computationally intensive process.
B, each query is much more computationally intensive than, say, a Google search.
Three, the company that's absolutely crushing the space and printing money because of this is NVIDIA.
Yeah.
And four, there is a general scarcity of computing powers so that even if you and I, like, were brilliant mathematicians and AI theorists,
etc. If we wanted to start a Chad GPT competitor, just getting access to the computing power
in order to do that would not be trivial, even if we had tons of money. Outside of that...
I'm going to buy an out-of-business crypto mine and take all the chips out of that. Tracy, they've
already been bought. Someone got that. But that's it. That's basically the extent of my understanding
of the nexus between this AI and chips. And I suspect there's more to know.
Well, I also think having a conversation about semiconductors and AI is a really good way to understand the underlying technology of both those things.
So that's what I'm hoping for out of this conversation.
All right.
Well, you mentioned we've been doing, we've done lots of chips episodes in the past.
So we're going to go back to the future or something like that.
We're going to go back to our first episode, our first guest, where we started exploring chips episodes.
I think it was the first one that we did sometime, maybe in early 2021.
We're going to be speaking with Stacey Razgan, managing director and senior analyst of U.S.
semiconductors and semiconductor capital equipment at Bernstein Research, someone who's great at breaking
all this stuff down, has been doing a lot of research on this question now. So, Stacey,
thank you so much for coming back on odd lots. I am so happy to be back. Thank you so much for
having me. All right. So I'm going to start with just sort of like not even a business question,
but a sort of semiconductor design question, which is,
this company in Vida, like for years, I just sort of knew them, is like they were the company that made graphics cards for video games. And then for a while they got there like, oh, and they're also good for crypto mining. And they were very popular for a while in Ethereum mining when it used proof of work. And now my understanding is everyone wants their chips for AI purposes. And we'll get into all that. But just to start, what is it about the design of their chips that makes them naturally suited for these other things? A company,
that started in graphics cards that makes them naturally suited for these things like AI in a way,
apparently that other chip makers, like say, in Intel, their chips do not seem to be as used for this
space. Yeah. So let me step back. Yeah, sure. If the question, if the question is totally flawed in its
premise, then feel free to say your question is totally flawed. Let me step back. So I'd say the idea of like
using compute in artificial intelligence has obviously been around for a long,
long time.
And actually the AI industry has been through a number of what they call AI winters over the
years where people would get really excited about this and then they would do work.
And then it would just turn out it wasn't working.
And pretty much it was just because the compute capacity and capabilities of the
hardware at the time wasn't really up to the task.
And so interest would wane and you'd go through this winter period.
And a while back, I don't know, 10, 15 years ago, whenever it was.
it was sort of discovered that the types of calculations that are used for neural networks
and machine learning, it turns out they are very similar to the kinds of mathematics that
are used for graphics processing and graphics rendering. As it turns out, it's primarily matrix
multiplication. And we'll probably get into this call on this call a little bit in terms of how
these machine learning models and everything actually work. But at the end of the day, really,
it comes down to like really, really large amounts of matrix multiplication in parallel operations.
And as it turned out, the GPU, the graphics of processing unit was quite suitable.
Okay, before you go on, and maybe we'll get into this in hour three of this conversation.
No, we're not going to go that long. But what is matrix multiplication?
Yeah, so I don't know how many of your listeners here have had linear algebra or anything.
But a matrix is just like an array of numbers.
Think about like a square array of numbers.
Okay. And matrix multiplications, I've got two of these arrays and I'm multiplying them together. And it's not as simple as the kind of math or multiplication that maybe you're typically used to. But it can be done. And it turns out there are some of these characteristics of these kinds of matrices. Number of these matrices can be really big and there's like lots and lots of operations that need to happen. And this stuff needs to happen like quite rapidly. And again, I'm grossly simplifying here for the listeners. But when you're working through these kinds of, um,
machine learning models, that's really what you're doing.
It's a bunch of different matrices, a bunch of different arrays of numbers that contain
all of the different parameters and things.
But we should probably step up a bit and talk about what we actually mean when we talk
about machine learning and models and all kinds of things.
But at the end of the day, you have these really large arrays of numbers that have to
get multiplied together in many cases over and over again many, many times.
And it turns into a very, very large compute problem.
And it's something that the GPU architecture can actually do really, really efficiently,
much more efficiently than you could say on a traditional CPU.
And so as it turns out, the GPU has become a good architecture for this.
Now, what Nvidia is done on top of this, not only with having the hardware,
is they've also built a really massive software ecosystem around all of this.
They have their software is called Kuta.
Think about it's kind of like the software, the programming environment,
like the parallel programming environment for these GPS.
And they've layered on all kinds of other libraries and STKs
and everything on top of that that actually makes this relatively easy
to use and to deploy and to deliver.
And so they built up not just the hardware,
but also the software around this.
And it's given them a really, really sort of like massive gap
versus like a lot of the other competitors
that are now trying to get into this market as well.
And so it's funny if you look at Nvidia as a stock.
I mean, today and this morning,
it's about, oh, I don't know, $260 or $270 a share.
This was a $10 to $20 stock forever.
And in fact, they did a $4.4.1 stock split recently.
So that'd be more like, you know, like a $2.50 to $5 stock on today's basis for years and years and years.
And just the magnitude of the growth that we've had with these guys over the last like five or 10 years,
particularly around their data center business and artificial intelligence and everything has just been like quite remarkable.
And so the earnings have gone through the roof and clearly the multiple that you're placing on those earnings has gone through the roof because the view is that the opportunity here is massive and that we're early and there's a lot of runway ahead of us.
And the stocks, I mean, it's had its ups and downs, but in general, it's been a home run.
I definitely want to ask you about where we are in the sort of semiconductor stock price cycle.
But before we get into that, you know, I will also bite on the really basic question that you already alluded to.
But how does machine learning slash AI actually work?
You mentioned this idea of, I guess, processing a bunch of data in parallel versus, I guess, old-style computing where it would be sequential.
But, like, talk to us about what is actually happening here and how does it fit into the semiconductor space?
You bet.
So let me first abstract this up and give you a really contrived example, just sort of simplistically about what's going on.
and then we can go a little bit more into the actual details of what's happening.
But let's imagine you want to have some kind of a neural network.
By the machine learning is typically done with something called a neural network.
And I'll talk about what that is in a moment.
But let's just imagine, for example, you want to build an artificial intelligence, a neural network to recognize pictures of cats, it's just saying.
Okay.
So let's imagine, I've got this black box sitting in front of me.
And it's got a slot on one side where I'm taking pictures and I'm feeding them in.
it's got a display on the other side which tells me yes it's a cat or no it's not and on the side of the box there are a billion knobs that you can turn okay and and they'll change various parameters of this model that right now are inside the black box don't worry about what those parameters are but there's there's knobs that can change them and so effectively what you're doing when you're training the same by the when you have the artificial dogs what you have is you have this big black box you need to train it to do a specific thing and so effectively what you're doing
task. That's what we're going to talk about in a moment. That's called training. And then once it's
trained, you need to use it for, for whatever task you've traded for. That task is called inference.
You got to do the other thing, training and inference. So the training, here's what you go. I got my
box with a slot and the display and a billion knobs. So what I do for the training process,
effectively is I take a picture and a known picture. So I know if it's a cat or not. I feed it
into the box and I look at the display and it tells me, yes, it's a cat or yes, it's not and it
probably gets it wrong. So then what I do is I turn some of the knobs and I feed another
picture in and then I turn some of the knobs and I'm basically tuning all of the parameters
and sort of measuring how accurate is this network at doing its task at recognizing is this a
picture of a cat or is it not. And I keep feeding pictures in known pictures, known data set,
and I keep playing with all the knobs until the accuracy of the thing is wherever I want it to be.
So yes, it's decided that now it's very good at recognizing is this a picture of a cat or is it not.
At that point, my model, my box is trained.
I now walk all of those knobs in place.
I don't move them anymore.
And now I use it.
Now I can just feed in pictures and it'll tell me, yes, it's a cat or yes, it's not.
And so the process of training this model is, that's really what it's about.
It's about varying all of the parameters.
And by the way, these models can have billions or hundreds of billions or even more of parameters that can be changed.
And that's the process of training.
You're basically trying to optimize this sort of situation.
I'm changing the parameters a little bit at a time such that I can optimize the response of this thing, such that I can get the performance of it, the accuracy of the network to be high.
So that's the training process.
And it is very, very compute intensive because you can imagine if I've got a billion different knob.
that I'm turning on, trying to optimize the output that takes a lot of compute.
The inference process, once that's at all that is much less compute-intensive.
Because I'm not changing anything.
I'm just applying the network as it is to whatever data that I'm feeding in at that point.
I'm not changing anything.
But I may be doing a lot more.
The difference is with the inference, I may be using it all the time, whereas once I've
trained the model, I've trained it.
So it's more like a one-and-done versus like a continual use sort of thing.
Since we're getting into sort of the economics of training versus inference,
A, is there sort of any way to get a sense of like, I'd say Tracy and me start odd lodge GPT?
It's a competitor to chat, a competitor to OpenAI.
Like, what are we thinking of in terms of just that scale?
How much we're spending to compute on the training part, then how much are recurring costs in terms of inference are?
And then I'm also just curious, like, also like, I know you said the inference is much cheaper,
but how much cheaper is it versus, say, asking Google a question?
How much more expensive is it?
How much more expensive is a Chad GPT query or an odd-lodge GPT query versus just a normal Google search?
Yeah, no, you get.
And by the way, when I say cheaper, it's like for any given single use, right?
Again, if I've got like a hundred billion different inference activities, maybe it's not.
It's still expensive.
Right.
Yeah.
But I first want to talk about just really quickly about like, so that this is my big abstract, contrived example about what's going on.
If I go just a little bit deeper about what this thing is, like, let's talk just briefly about a neural network and then I will get true question.
But it kind of influences it.
So what is a neural network?
If I was to draw like a representation of a neural network for you, what I would do is I would have a bunch of circles.
Each of those circles would be a neuron.
And I wish I was there, I could draw a picture for you.
But imagine like a picture.
Send a picture. After you're done, send a picture and we'll run it with the episode.
We'll run it with the episode.
Okay, okay.
I can do that.
Your hand drawn, your hand drawn explanation of how it's.
On a napkin.
These are very easy to find.
But anyways, but imagine, like, I've got, like, a group of circles.
I've got, like, a column, you know, in column one with, like, three circles.
And then column two, I've got, I don't, three or four circles.
In column three, I've got some circles.
These are my neurons.
And imagine I've got arrows that are connecting each circle to the circles in one row
to all of the circles in the next row.
Those are my connections between my neurons.
So you can see it looks like kind of a net or a network, okay?
And so within each circle, I've got some of what's called an activation function.
So what each circle does is it takes an input, the arrow that's coming into it.
And it has to decide, based on those inputs, do I send an output out out the other side or not?
Right.
So there's some certain threshold.
If the inputs reach some amount of threshold, the neuron will fire, just like the neuron in your brain.
Okay.
Each neuron can have more than one input coming in from more than one neuron in the preview.
These are called layers, by the way, these rows of circles, can have more than one input from the different neurons in the previous layer.
And that the neuron can weight those different inputs differently.
It can say, you know, from this one neuron, I'm going to give that a 50% weight.
And from the other neuron, I'll only wait at 20%.
I'm not going to take the full signal.
So those are called the weights of the network.
And so each neuron has inputs coming in and outputs going out.
And each of those inputs and outputs will have a weight associated with it.
So those are, when I remember I talked about those knobs, those parameters.
Yeah.
Those weights are one set of parameters.
And then within each neuron, there's basically there's a certain threshold with all those,
all those signals coming in when you add them up, if they reach a certain threshold,
then the neuron fires.
Okay.
So that threshold is called the bias and you can tune that.
Like I can have a really sensitive neuron where if the bias doesn't, I don't need a lot of signal
coming in to make it fire.
I can have a neuron that's less sensitive.
I need a lot of signal coming in before it'll fire.
that's called a bias that that that's also a parameter so those are the parameters that you're
setting the structure of the network itself the number of neurons and the number of layers and
everything that's that's sort of set and then you're trying to determine these weights and biases
and again just just a level set you check gpT which which is i'm getting excited about
has 175 billion separate parameters that they get set during their during the training process
okay so that's that's kind of what's what's going on
Today's show is brought to you by Vanguard.
To all the financial advisors listening, let's talk bonds for a minute.
Capturing value and fixed income is not easy.
Bond markets are massive, murky, and let's be real.
Lots of firms throw a couple flashy funds your way and call it a day.
But not Vanguard.
At Vanguard, institutional quality isn't a tagline.
It's a commitment to your clients.
We're talking top-grade products across the board of over 80 bond funds,
actively managed by a 200-person global squad of sector specialists,
analysts and traders.
These folks live and breathe
fixed income.
So if you're looking to give your clients
consistent results year in and year out,
go see the record for yourself
at vanguard.com slash audio.
That's vanguard.com
slash audio.
All investing in subject to risk
Vanguard Marketing Corporation distributor.
They say abs are made in the kitchen.
Cool, but who has time for
three hours of meal prep and a fridge full
of Tupperware?
That's why I started using Factor.
Factor delivers fresh, never frozen,
ready-to-eat-eatian designed for balanced science-backed nutrition.
No prep, no cleanup.
Just heat, eat, and move on.
I've got gym days, work days, super long days,
and Factor keeps me on track without slowing me down.
It's real food, great flavor,
and the kind of meals that actually support the work I'm putting in.
We're talking chicken pesto, steak with veggies, roasted salmon,
meals that feel like they were cooked for you, not by you.
So, yeah, abs might be made in the kitchen.
But thanks to Factor, I don't have to be in the kitchen.
Right now, get 11 meals, free shipping and free sides for life.
Hurry, this offer won't last long.
Go to FactorMeals.ca and use code pro.
That's 11 meals, free shipping and free sides for life, but only with the code pro at factormeals.
Factor, Canada's number one ready to eat meal delivery service.
Before you talk about economics, can I just ask?
So one of the things about the technology is it's supposed to be iterative, right?
Like, it's learning as it goes along.
Can you talk just briefly maybe about how it's incorporating, like, new inputs as it develops?
Yeah.
So when you training, let's talk about training now.
So when you train the network, it happens on a static dataset.
Okay?
So you have to start with a data set.
Right.
And in terms of chat GPT, that is, you know, it has a large corpus of data that it was trained on.
It was a lot of data from the, there's a lot of data from the, you know, it was a lot of data from
internet and from other sources, right?
Basically, we've trained the smartest.
Like, whole of the internet, wasn't it?
But also, I think like, I'm not exactly, a lot of Reddit.
So it's like we've, right?
Like, it's like, we've trained as like the greatest brain of all time is like Reddit
pills.
Now it talks like a 17 year old boy.
So there's a lot of data.
And so you ask like sort of how does that data get, you know, incorporated into, so
I don't want to get too, I'm already getting too complicated.
I don't want to get too complicated.
Let me talk about how to standard training works.
And then we can talk about chat chief.
because that uses a different kind of model.
It's called a transformer model.
But anyway,
but when I'm training this,
so what happens is I feed this stuff.
There's a process called,
it's called back propagation.
Basically what you do is you sort of feed this stuff through,
through the network itself,
and then you work it backwards.
Basically, what you're doing is you're measuring the output
against a known response.
I want to sort of, you know,
that's my cat texture.
Is it a cat or is it not a cat?
I'm trying to minimize the difference between them.
I want it to be accurate, right?
So what you sort of do is you roll a certain step through the network, right?
You measure the output against the known, what it should be.
And then there's a process that's called back propagation.
What you're doing, you're actually, you're calculated what's called the gradients of all of these things.
You're basically looking at sort of like the rate of change of these different parameters,
and you sort of work the network backwards.
And that gradient that you're calculating kind of tells you how much to adjust each parameter.
So you work it back.
And then you work it forward again and then you work it backward.
And then you work it forward and you work it backward.
And then you do that until you've converged,
that the network itself is accurate to wherever you want it to be to be accurate at.
So that's, again, I'm grossly simplifying here.
I'm trying to keep this as high level as possible.
But that's kind of way you're doing.
And just in terms of the amount of could be sort of trained check GPT.
And check deep, they've actually released all the,
the details of the network, like how many layers and what's the dimension and like parameters,
all this stuff so we can do this math. It turns out to take about three times 10 to the 23rd
operations to train it. And so just, that's 300 sextillion operations it took to train chat GPT.
Now, in terms of how much it costs, so chat TV was, they kind of said this. It was trained
on 10,000 Nvidia, what they called V100. That's the voltage chip. That's a chip that's several years old
for NVIDIA, but it was trained on supposedly about 10,000 of these.
And we did some of this math ourselves.
I was coming out more like three or four thousand, but there's a ton of other
assumptions you have to make.
And your 10,000 seems to be the right order of magnitude for that part.
That part of the time cost about, you know, I don't know, $8,000.
And so the number that was kind of tossed out with something like $80 million to
train chat GPT one time.
Wow.
I think that's, I don't know, 80 million dollars doesn't seem like that much to me.
Well, no, so this is.
I get it, but like there are a lot of companies that could spend that have $80,000.
I actually agree with it.
We're jumping ahead.
But my take is that for large language models, and we can talk about these different things,
but for large language, I was like Jack CPD, I actually think inference is a bigger opportunity.
And you're kind of getting to the heart of it.
It's because inference scales directly, the more queries I run.
Because you only have to train once and that's done.
And that's 80 million.
Or even if you're training it more than once.
And again, do your question, Tracy, like you can add to the data set and retrain it.
But if I've already got the info, let's say I'm training it every two weeks.
Okay.
Yeah.
That'd be training it like 24 or 25 times a year, but I've got the infrastructure that is in place
already.
Right.
Right.
To do that.
And so the training tan will be more around how many different entities actually
develop these models and how many models each do they develop and how often do they train
those models?
And importantly, how big do the models get?
Because this is one of the things.
Chat, GPDD is big, but GPT4, which they,
or at least now is even bigger.
They haven't talked about specs,
but I wouldn't be surprised.
ChatGPD4 is rumored to have over a trillion parameters.
I could very well like.
And we're very early into this.
Like these models are going to keep getting bigger and bigger and bigger.
And so that's how I think the training market,
the training tem will be growing.
It's a function of the number of trainings of all these models
we're doing every year in the size of these models.
And the models will get big.
But in your view, the big money is going to be made on the inference.
So let's talk about.
I think so talk about what happens then and your sort of sense of the side.
I don't know.
Yeah, just talk to us about the inference part and the economics.
You bet.
Chat Chb-T in these large language models, it's a new type of model.
It's called a transformer model.
There's a bunch of compute steps that have to happen.
There's also a step in there that helps it map the relation, capture the relationship
between, you know, by the way, if you've ever used chat chabit, you know, you type in like
a query and do.
a box and it returns a response. So that query is broken into what are called tokens. It's basically
thinking, do you think about a token is kind of like a word or a group of words sort of, but the transformer
model has something it's called a self-attention mechanism. And what that does is it captures the
relationship between those different tokens and the input sequence based on the training data
that it has. And that's how it knows what it's really doing, it's predictive text. It knows,
based on this query, I'm going to start the response with this word and based on this word
and this query and my data said, I know these other words typically follow. And it kind of constructs
the response from that. And so our math suggests that for like a typical query response
called like, you know, 500 tokens or maybe 2,000 words, it was something like 400 quadrillion
operations needed to accomplish something like that. And so you can size this up because I know
for like an Nvidia GPU and you can do it for different GPUs.
I know how many operations per second each GPU can run.
And I know how much these GPS ballpark kind of cost.
And so then you've got to assume like, well, okay, how many queries per day are you going to do?
And you can come up with a number.
And I mean, frankly, the number can be as big as you want.
It depends on how many queries.
But I think a Tam, you know, at least in the multiple tens of billions of dollars is not
unreasonable, if not more.
And just to level set, I mean, it gets to your Google question.
And Google does about 10 billion searches a day, give or take.
I think a lot of people have been looking at that level as part of like, you know,
like the end-all be all for where this could go.
I'll be honest.
Like I understand why people are, especially internet investors,
are concerned that large language models and things like chat GPD can start to disrupt search.
I'm not exactly sure that search is the right proxy personally.
It feels kind of limiting to me.
I mean, you can imagine, I've watched a little too much Star Trek, I guess.
But I mean, you can imagine, you know, you have like a virtual.
assist in the ceiling, I'm calling out to it. And, you know, it doesn't have to be just search on
my screen. I could have it in my car. Right. I could have, you know, I call up American Airlines
that change my airline tickets, and it's a chat bot that's talking to me. So this could be very
big. And by the way, I think to guess, by the way, the one problem with this, probably a calculation,
it's kind of static. Like the cost is sort of an output rather than an input. I think to drive adoption,
cost will come down. And we've already seen that. Like, Nvidia has,
has a new product that's called Hopper,
which is like two generations past those V-100s
that I was talking about, past the Volta generation.
The cost per query to do this or the cost for training on Hopper
is much lower than Volta because it's a much more efficient part.
That's a good thing, though.
It's Tamacretive.
It will drive adoption.
InVVVVVIA actually has specific products specifically designed to do this kind of thing.
And Hopper has specific blocks on it that actually helped
with the training and inference on these kind of large language models.
And so I actually think over time is the efficiency gets better and better, you're going to drive adoption more and more.
I think this is a big thing.
And remember, we're still really early.
Chat GPD only showed up in November.
Yeah, it's crazy, isn't it?
It's really early still.
Well, just on that note, can you draw directly the connection between the software and the hardware here?
Because I think at this point, probably everyone listening has tried ChatGPT, and you're used to seeing it as a sort of, you know, it's an interface on the
internet and you type stuff into it and it spits something out. But like, where do the semiconductors
actually come in when we're talking about crunching these enormous data sets? And what makes a,
you kind of touched on this a little bit with Invidia, but what makes a semiconductor better at
doing AI versus more traditional computational processes? Yeah. Yeah. You bet. So to answer that second
question, I think AI is really much more around parallel processing. And in particular thing,
It's this kind of matrix map.
It's a single class of calculations that these things do very, very efficiently and do very, very well.
And they do them much more efficiently than a CPU that performs a little more serially versus parallel.
You just couldn't run this stuff on CPUs.
But don't get me wrong, you do some of, we've been talking about inference on large language models.
There's all kinds of inference.
Inference workloads range from very simplistic to very, very, very complex.
Again, my, you know, cat recognition example was very simplistic.
Something like this, or frankly, something like autonomous driving, that is an inference
activity, but is a hugely computationally intense inference activity.
And so there's still a lot of inference today that actually happens.
In fact, most inference today actually happens on CPUs.
But I'd say the types of things that you're trying to do are getting more and more complex
and CPUs are getting less and less viable for that kind of, for that kind of math.
And so that's kind of the difference between GPUs and other types of parallel offerings versus like a CPU.
I should say, by the way, GPU is not the only way to do this.
Google, for example, has their own AI chips.
They call them a TPU, tensor processing unit.
One thing I really like about talking to Stacey, two things is, A, I think he comes up with better versions of our questions than we do.
One thing about the question you just asked him didn't actually ask.
He's always like, all right, that's a good question.
but let me actually reframe the question to get a better response.
So I appreciate that.
And he also anticipates because I literally, like on my computer right now,
I had Google cloud tensor processing units because that was my next question.
And also in part because I think yesterday the information reported that Microsoft is also.
So why don't you talk to us about that, these other and what are they competing directly with the technology?
Yeah.
Yeah, you bet.
So Google is, by the good, this is not new.
Google has been doing their own chips.
for seven or eight years. It is not new.
But they have what they call a TPU, and they use it extensively for their own internal workloads.
Absolutely. Amazon has their own chips. They have a training chip that's called, you know,
kind of hysterically, it's called Traneum. They have an inference chip. It's called Interferentia.
Microsoft apparently is working on their own. My feeling is every hypersaler is working on
their own chip, particularly for their own internal workloads. And that is an area,
We talked about Nvidia software mode.
Like Google doesn't need Nvidia's software mode.
They're not running Kuda.
They're just running TensorFlow.
Okay.
And doing their thing.
They don't need Kuda.
Anything, however, that is facing an end customer,
like an enterprise like end customer, like on a public cloud,
like a customer going to AWS and renting, you know,
compute power, that tends to be GPS because customers don't have Google's
sophistication.
They really do need the software ecosystem that's built around.
So for example, I can go to Google Cloud.
I can actually rent a TPU instance.
It can be done.
Nobody really does it.
And actually, if you look how they're priced, typically, it's actually more expensive
than have the way that Google's pricing GPUs on Google Cloud.
It's similar for Amazon and others.
And so I do think that all the hyper-skillers are working on their own,
and there is certainly a place for that, especially for their own internal workloads.
Anything that's facing a customer that, that Nvidia-GPU,
ecosystem is really kind of grown up. So actually, just to clarify, because that point is really
interesting, that for like, if again, Tracy and I want to launch OddLod's GPT, part of the issue
would be not necessarily the hardware, the silicon, but actually that Nvidia's software suite built
around it would make it much easier for us to sort of start and use on Nvidia for training our
model. Yeah, yes, it would. And they've built a lot. And it's funny, you can go listen to, like,
in Vydea's announcement in their analyst days and things. And there's as much about software as they
are about hardware. So not only if they continue to extend like the basic, like the Kuda ecosystem,
they've layered all kinds of other application specific things on top of it. So they've got what
they call Rapids, which is for enterprise machine learning. They've got a library package called
Isaacs, which is for automation and robotics. They've got a package called Clara, which is specifically
for medical imaging and diagnostics.
They've got something called Ku Quantum,
which is actually for quantum computer simulations.
They've got something for drug discovery.
So they're layering all these things on top, right,
depending on your application.
They've got internal teams that are working on this.
It's not just throwing the software out there.
They've got people there that can actually help you work or work
and come along with it.
They're doing other things easier.
So they actually just launched a cloud service.
And this is with Google and Oracle and Google and Microsoft,
where you can almost, they'll do like a fully provisioned
Nvidia AI supercomputer in the cloud.
So, because like, you know, they sell these AI servers and they can cost hundreds of
thousands of dollars apiece.
If you want now, you can just go to Oracle Cloud or Google Cloud or whatever.
And you can sort of rent a fully provisioned Nvidia supercomputer sitting in the cloud that they'll,
all you got to do is access it right through a web browser.
This was going to be.
And they'll just make it super easy.
This was going to be my next question, actually.
So I take the point about software.
but like what do the AI supercomputers actually look like nowadays?
Like is there a physical thing in a giant data center somewhere?
Oh yeah.
Are they mostly like cloud-based or what does this look like?
Like walk us through the ecosystem.
So Nvidia sells something they call it a DGX.
It's a box.
I mean, it's, I don't know what it's, when is it two feet?
I don't know what the dimensions are two feet by two feet or something like that.
It's got eight GPUs and two CPUs and a bunch of memory and a bunch of network.
They've got their own, like, you know, they bought a company called Melanox a while a while back that did networking hardware.
So it's got a bunch of proprietary network.
Because that's, by that's something else we haven't talked about.
It's not just enough to have the computer, the compute.
These models are so big.
They don't fit on a single CPU.
So you have to be able to network all this stuff together.
Yeah.
Right.
And so they've got networking in there.
And they have this, this box.
And then you can, you can stack a whole bunch of boxes together.
Like, Nvidia has their own internal supercomputer.
It's on the, it's fairly high on the top 500 list.
They call it Celine.
it's a bunch of these DGX
like servers that they make all just like
stacked together effectively
and they sell for the older generation
their prior generation was called Ampere
and that box sold for $199,000
I don't believe they've released pricing
on the Hopper version but I know for the Hopper GPU
it costs two to three X what Amper cost
the prior generation so
So this actually raises a separate question
to me which is
okay there's the price and it exists
and you could go to, you could theoretically go and use Google's TensorFlow-based cloud,
or is it available?
Like, or is, because I sort of get the impression that, like,
for some of the technology that people want to use,
it's not available at any price and that there is a actual, is that real or not?
It seems to be.
So we're, like, so they're new generation, which is called Hopper,
which, like I said, has characteristics of it that make it very attractive,
especially for these kind of like chat, GPT, large language models,
is in tight supply.
We're at the very beginning
of that product cycle.
They just launched it
like when it lasts
like a couple of quarters.
And so that ramp up takes time.
And it does seem like
they are seeing accelerated demand
because of this kinds of stuff.
And so yeah,
I think supply is tight.
We've heard stories about GPU shortages
at Microsoft
and the cloud vendors.
I think there was a Bloomberg store
the other day that said
these things were selling
for like $40,000 on eBay or something.
I took a look at some of those listings.
They looked a little shady to me.
But yeah, it's tight.
You have to remember, these parts are very complicated, so the lead times to actually have more made.
It takes a while.
Wait, so just on this, no, I joked about this in the intro, but, you know, could I buy, like, a Bitcoin mining facility and take all that computer processing power and, like, convert it into something that could be used for AI?
Is that a possibility?
You could.
The Bitcoin stuff, at least a lot of the Bitcoin stuff was done that was with GPS.
Those were still mostly gaming GPUs.
people were buying gaming GPUs and repurposing them for Bitcoin and Ethereum,
mostly Ethereum mining.
Yeah, they're not nearly as compute efficient as the data center parts, right?
But I mean, in theory, yeah, you could get, you know, gaming GPUs if you could and stringently
to get, but it would be prohibitive, right?
And even now, most of that stuff's cleared out, I think, as Joe said.
But the math is somewhat similar.
I'd say for for these kinds of models, though, again, like a hopper in video's new
data center product has, they have something they call it a transformer.
engine. What it really does is it allowed you to do the training at a slightly lower precision
than, it lets you do it at 8-bit floating point versus 16-bit. So it lets you get higher performance.
And then there's another process. There's like a conversion process. Sometimes it has to go
when you go from training to inference, it's something like quantization. And with these transformer
engines, you don't have to do that. So it increases the efficiency, which you wouldn't get by
picking some random GPU. Where is Intel in this story? Well, so let's talk about the other
competitive options that are out there.
So we talked about some of the captive
silicon and hyperscalers.
That is there and it is real and they're all
building their own and they've been doing it forever and it hasn't
slowed anything down the slightest because we're
still early and the opportunity is big.
By the way, I will say, I don't worry
to lead with it. I don't worry so
much about competition at this point because
think about it. Invidia's
running into their data center business right now. It's something
like $15 billion a year. That's where it is.
It's growing but that's where it is.
So Jensen, Invidia's CEO, like,
likes to throw out big numbers. And he threw out, I think he said for silicon and hardware
Tam in the data center, he thought that their Tam overtime was $300 billion. And it seemed kind
of crazy, although I would say like it's seeming a little less and less crazy every day.
But if you thought the Tam was $300 billion or $100 billion or like whatever, and the run rating
at $15 billion, there's tons of headroom. Competition doesn't really matter. And that's what we've
seen. We've seen competition. But there's so much opportunity, like who cares, right?
Right, versus like if he thought it was a $20 billion,
Tam, like they would have a problem like already today.
So that's why I don't worry too much because I think the opportunity is still very,
very large relative to where they're running into the business today.
In terms of other competitors, though, say yes,
you mentioned, let's talk about AMD first because A&D actually makes GPUs.
They make data center GPUs.
They don't sell very many of them.
Their current product is something called the MI250, and they've sold de minimis, basically.
And in fact, you know, when the China,
China sanctions were put on.
And we didn't talk about that, but the U.S.
stopped allowing like high-end AI chips from being shipped to China.
Right.
The MI 250A&E part was on the list, but it didn't affect them at all because they weren't selling any.
So their sales were zero.
They've got another product coming out at the following that's called the MI300.
And people have been getting kind of excited about A&B.
They've been sort of looking to play it as kind of like the poor man's Nvidia.
I'll be honest.
I don't think it's the poor man's invidia.
And Vigia is doing, you know, close to $4 billion a quarter in data center revenues.
I don't know that I see anything like that with the MI300.
In AMD, as far as I tell, has not even released any sort of specifications for what it looks like at this point.
But that is an option.
And some people would say there's maybe some truth that this is, you know, if you want an alternative,
A&D will present an alternative.
And if the opportunity is really that big, they'll get some.
They'll probably get some if you have that.
You have Intel.
So Intel's got a few things.
On their CPUs, their current version is called Sapphire Rapids.
It has AI-specific accelerators for inference,
not so much maybe for this kind of stuff,
but for general inference activities,
they're trying to play up the capabilities of their CPU on that.
Fine.
And why are they doing that?
It's because their accelerator roadmap isn't so good.
So they have a GPU roadmap.
The code name for it was Pontaveccio.
And they've kind of gutted that roadmap.
So the follow-on product was something called Rialto Bridge
that they've since canceled.
and one of the Pontaventio products recently they just canceled.
And Pontaventio originally was designed for the Aurora supercomputer,
and it was massively late.
I mean, so they took, how much was it?
It was something like a $300 billion charge.
I think it was at the end of 2021.
It was either end of 20 or end of 2021,
where they basically gave it away.
It was so late.
So that's how late they were.
They also have another product.
They bought an Israeli AI company.
called Hibana.
And Hibona has a product called Gaudi.
It's not a GPU exactly, but it's like a specific accelerator technology.
And Amazon bought some of them, and they sell a little bit.
But again, versus Intel's total revenues, it's de minimis.
So they're not really there.
There's also a bunch of startups.
And the problem with most of the startups is their story tends to be something like, you know,
we have a product that's 10 times as good as Nvidia.
And the issue is, with every generation, Nvidia has something that's 10 times as good
is Nvidia and they have the software ecosystem that goes with it.
By the way, neither AMD nor Intel nor most of the startups have anything remotely resembling
Nvidia software.
So that's another huge issue that all of them are facing.
There's a few startups that have some niche success.
One of the one that's probably gotten the most, you know, attention is called cerebrus or cerebrus.
And their whole thing, they make a chip.
It's imagine taking a 300 millimeter silicon wafer and it's inscribing a square on it.
That's their chip.
It's like one chip for wafer.
And so you can put very large models onto these chips,
and they've been deploying them for those kinds of things.
But again, the software becomes an issue,
but they've had a little bit of success.
There's some other names that you've got Grock and some others,
I think they're that are still out there.
And then there's a company go Tens Forge, which is interesting,
not because of so far what they're doing because it's early,
but it's run now by Jim Keller.
Do you guys know who Jim Keller is?
I do not.
Jim Keller was, he's sort of like a star chip designer.
He designed it.
Apple's first custom processor.
He designed AMD's Zen and Epic Roadnefts that they've been that they've been taking a lot of share with.
He was even at Tesla for a while and at Intel.
And so he's now running 10 store it.
And they do it's our risk five.
Risk five is another type of architecture.
And they do they do an AI chip.
So Jim is running that.
So can I just ask based on that?
I mean, how like CAPEX intensive is developing chips that are well suited for AI versus other types of chips?
And then secondly, like, where do the improvements come from or what are the like improvements focused on?
Is it speed or like scale given the data sets involved in the parallel processes that you described?
Yeah.
So it's a few things.
So in terms of CAPEX intensive, these are mostly design companies.
So they don't have a lot of CAPEX.
It's certainly R&D intensive.
So maybe that's what you're getting.
And Viniest spends make many billions of dollars a year on R&D.
and Nvidia has a little bit of advantage too
because it's effectively the same architecture
between data center and gaming
so they've got other volume
effectively to sort of amortize some of those
investments over. Although now, I mean,
this year I mean, data center is probably 60%
of Nvidia's revenues now. So I mean,
Nvidia is sort of the center of data center
is a center of gravity for Nvidia now.
But it's very R&D intensive
and probably getting more so. And you've got
folks all up and down the value chain that are investing.
You're both the Silicon guys and
the cloud guys and the customers and
and everything else. But I mean, that's kind of where we are. In terms of what you're looking for,
so there's a few things. You're looking for performance, and I'm training quite often that comes
down to like time to train. So I've got a model, like some of these models, I mean, you could
imagine could take weeks or months historically to train, right? And that's a problem.
Like you want it to be faster. So if you can get that down, you know, to weeks or, you know,
to days or hours, that would be better. So that's one thing clearly that they work on.
I don't want to
You know something notice
Yeah go ahead
Finish your thought
Then I have it slightly
Oh yeah
The other thing that was talking
There's something around like
Scale out
So basically remember I said
You're you're connecting
Lots and lots of these chips together
So for example
If I if I increase the number of chips
By 10x
Does my trading time go back down by like a factor of 10
Or is it like by factor of 2
So like ideally
Ideally you would want like linear scaling
Right
I want like as I add resources
It scales linearly
So this is kind of getting
was going to get into my next question, actually. And, you know, we can talk in another, with someone
else about certain, like, AI fantasy, doomed scenario. But I'm not an AI. No. I'm not an AI architect
expert. I'm a Dundex engineer. So I could just say, you may want to get an AI. No, I know.
But I am curious, though, because I do think it relates to this question, which is that, okay, like,
with each one, like, GPT5, and they're going to, like, keep adding more knobs on the box, et cetera.
Like, and is your perception that this sort of quality of the output is growing exponentially,
or is it the kind of thing where it's like GPT4, you know, there's a lot more knobs and they got a big jump from GPT3,
GPT5 will be way more knobs, but like, is it going to be marginally better?
Like, what is the sort of like, where are we in the sort of like, what does the shape of the output curve look like?
and this sort of like cost of, you know, these chip developments in terms of getting there.
I don't know.
So there's a couple of things.
So first of all, when you're talking about large language,
where I was accuracy is sort of a nebulous term because it's not just accuracy.
It's like, like, it's also capability.
Like what can it do?
We can show what chat GPT and GPD4 can do.
And also like, I think as you're going forward, you talk about the trajectories here,
it's not just text, right?
We're talking text to text right, but there's also text to images.
And if anybody played with like Dali,
You know, it's generating images from a text prompt.
And now we've got like video.
What is it?
Was it mid-summer?
Is that what it's called mid-Journey?
Yeah.
I can't remember.
Mid-Journey, yeah.
So it's creating like video prompts.
I mean, so like the, like text is just scrapped is just the tip of the iceberg, I think,
in terms of what we're going to need in terms of people.
They're never going to get to where they could have three people having a conversation with voices.
Why?
That's not like Tracy, Joe and Stacy.
Why?
No, I'm just kidding.
No, I'm just kidding.
It feels like we have one more year of this job.
Now, one of the dangers, clearly, and maybe this gets to capability.
So one thing with chat GPT is it's very, very good.
This is where I should worry about my job because it's very good about it's sounding like it knows what it's talking about where maybe it doesn't.
Maybe I should be worried about my job.
And accuracy, I think, is a big issue.
But you have to remember.
So, but like on this accuracy question, like, I assume, you know, like, self,
driving cars. Like when people were really hyped about them 10 years ago, they're like, oh, it's 95%
solid. We just have a little bit more and then it's solid. And then 10 years later,
10 years later, it feels like they haven't made any progress on that final 5%. Yeah, I mean,
these things are always a power law. So this is my question when we talk about accuracy or
these things like, are we at the point where like, is it going to be the kind of thing where it's like,
yeah, GPT5 will definitely be better than GPT4, but it will be like 90s.
36% of the way there.
Well, again, let me separate out
let me separate an accuracy
from capability again.
So if there's an accuracy, you have to remember
like it, the model has
no idea what accurate even
means. It doesn't, remember,
these things are not actually intelligent. I know there's
a lot of worry about like what they go like, like,
like AI, like articles with general intelligence,
right? I don't think this is it, this is
predictive text. Yeah. That's all. The model
doesn't know if it's if it's spewing
bull crap or truth. It has no idea. It's just
predicting the next word in the thing.
And it's because of what it's trained on.
So you need to add on maybe other kinds of things to ensure accuracy,
maybe to put guardrails or things, things like that.
You may need to very carefully, like more harsh like your input like datasets and things
like that.
I think that's a problem now.
I think it'll get solved.
There's enough to, but like, and this has already been an issue.
And you can take it like the other, like the, I don't know if it's the converse of it or
not, but things like deep fakes, people are deliberately trying to use AI to deceive.
I mean, this is just human nature.
this is why we have problems.
But I think they can work through that.
In terms of capabilities.
Oh, no, go ahead.
I think it's really interesting to look at like sort of similar,
like a response like to a similar prompt between like chat GPT and GPD4.
And like what people are getting out of GPD4,
it's miles ahead of like some of the stuff that chat GPD,
which was trained on GPT3, the model, that what it was what is delivering in terms of nuance.
Right.
And color and everything else.
I mean, and I think that's going to continue.
I wouldn't be, and already, you're on the point where these things can already pass the
Turing test.
Oh, yeah.
Right?
It can be very difficult to know if it's, you know, if I'm putting the question of accuracy
aside from it, it's very difficult to know for some of these things if you didn't
know any better whether it was coming from a real person or not.
And I think it's going to get like harder and harder to tell, like whether, you know,
even if it's not, you know, quote, unquote, really thinking, it's going to be hard for us
to tell what's really going on.
That is sort of like other interesting, you know, implications for what this might be.
of the next five years or 10 years.
They say abs are made in the kitchen.
Cool, but who has time for three hours of meal prep
and a fridge full of Tupperware?
That's why I started using Factor.
Factor delivers fresh, never frozen,
ready to eat meals that are dietitian designed
for balanced science-backed nutrition.
No prep, no cleanup.
Just heat, eat, and move on.
I've got gym days, work days,
super long days, and Factor keeps me on track without slowing me down.
It's real food, great flavor.
and the kind of meals that actually support the work I'm putting in.
We're talking chicken pesto, steak with veggies,
roasted salmon, meals that feel like they were cooked for you, not by you.
So, yeah, abs might be made in the kitchen.
But thanks to Factor, I don't have to be in the kitchen.
Right now, get 11 meals, free shipping, and free sides for life.
Hurry, this offer won't last long.
Go to FactorMeals.com.C.A. and use code pro.
That's 11 meals, free shipping, and free sides for life,
but only with the code pro at FactorMeals.C.A.
Factor. Canada's number one ready-to-eat meal delivery service.
If Bell Fib TV is now streaming, is it still TV?
Is it still TV if there's no TV box?
If I can stream all my favorite channels and pause and record shows, that's TV, right?
A new era of Fibb TV.
It's streaming, but it's still TV.
Well, glad that's settled.
Bell, Connection is everything.
Just going back to the stock prices, I mean, we mentioned the NVIDIA chart, which is
quite a lot, although not, it hasn't reached its peak back in 2021.
The Sox Index is recovering, but, you know, still below.
And Intel, I mean, I won't even mention.
But like, where are we in the semiconductor cycle?
Because it feels like on the one hand, there's talk about excess capacity.
and orders starting to fall.
But on the other hand,
there is this real excitement
about the future
in the form of AI.
Yes, yes.
So semis in general
were pretty lousy last year.
They've had a very strong
year-to-date performance
and sectors up,
which is sectors up,
you know,
20-2% year-to-date,
quite a bit above the overall market.
And the reason is,
to your point,
that we've been in a cycle,
numbers have been coming down.
And we may have talked about this last time,
I don't remember,
but semiconductor investors,
turns out the best time to buy stocks in general is after numbers come down but before they hit
bottoms. Like if you could buy them right before the last cut, if you could have perfect foresight.
You never know when that is, but numbers have cut, but numbers are cut down a lot.
So estimates, forward estimates for the industry peaked last June. And they are down over 30%,
like 35% since that point. It's actually the largest negative earnings revision we've had
probably since the financial crisis.
Wow.
and people are looking for, you know, playing the bottoming theme and that hopefully things get better into the second half.
You know, we get hopefully China reopening.
And you've got markets like, and this relates to Intel, like PCs and things where, you know, we've now corrected kind of, we're back like more on a pre-COVID run rate for PCs versus where we were.
And the CPUs, which were massively overshipping at the peak, they're now undershipping.
And so we're in that inventory flush part of the cycle.
And so people have been sort of playing the space.
for that like second half recovery.
Now, all that being said, if you look at the overall industry,
if you look at numbers in the second half,
they're actually above seasonal.
So people are starting to bake in that cyclical recovery
to the numbers.
And if you look at inventories, just overall in the space,
they are ludicrously high.
I've actually never seen them this high before.
So we've had some inventory correction,
but we may have not, we may just be getting started there.
And if you look at valuations,
I think the sector's trading
that's something like a 30% premium
with the S&P 500.
which is the largest premium we've had, again,
probably since things normalized after the tech bubble,
or after the financial crisis, at least.
So people have been playing this backup recovery,
but yeah, we better get it.
As it relates to some of the individual stocks,
like you mentioned, Intel.
It's funny, I think you guys may not know this.
I just upgraded Intel.
Oh.
How come?
The title of the note was,
we hate this call.
And I meant,
I desperately,
would like the standard.
It was, and it was not a, we like an Intel call.
It was just, I think that they, that they're now undershipping in PCs by a wide margin.
And I think for the first time in a while, the second half street numbers might actually
be too low.
So that's, it's not like a super compelling call, but I felt uncomfortable pushes.
Although they report earnings next week, I may be kicking myself.
Like, we'll see.
InVIDIA, however, so it's clearly, you know, you're right, it hasn't reached its prior
peak from a stock price base.
And the reasons the numbers have come down a lot.
I mean, let's be honest.
The gaming business was inflated significantly by crypto.
Right.
And so that's all come out, right?
And then, you know, with data center, he had some impacts from China.
China in general was weak.
And then we had some of the export controls that they had to work their way around.
So you had some issues there.
Now, all of that being said, graphics cards in gaming, we talked about some of these inventory corrections.
Graphics cards actually corrected the most and the most rapidly.
those have already hit bottom and they're growing again.
And Nvidia's got a product cycle there that they just kicked off.
The new cards are called Lovelace.
And they look really good, especially behind,
and they're starting to fill out like the rest of the stack.
So gaming's okay.
And then in data center, again, this, you know,
this generative AI has really caught everybody's fancy.
Yeah.
And Nvidia had a data center.
And they're saying they were at the beginning of a product cycle in data center.
And, you know, they had an event a couple weeks ago,
their GTC event where they actually basically,
I mean, directly said we're seeing upside from generative AI.
even now, right? So people were buying Nvidia on those, on that thesis. And like the last time the
stock hit these peaks, at least in terms of valuation, the issue is we were at the peak of their
product cycles and numbers came down. This time, valuations kind of went back to where they were
at those peaks, but at the beginning of the product cycles and numbers are probably going up,
knock down. So that's why. Stacey, I joked at the beginning that we could talk about, about this for
three hours and I'm sure we could.
I'm sure we're. There's such a deep area.
But that was a great overview of just like the state of competition, the state of play,
and the economics of this, in a very good way for us to sort of enter talking about AI stuff
more broadly.
Thank you so much for coming back on online.
My pleasure.
Anytime you guys want me here, just let me know.
All right.
We'll have you back next week for Intel.
All right.
Take care of Stacy.
Thanks, Stacy.
Thank you.
Bye-bye.
I really like talking this.
Stacey. He's really good at explaining complicated things. Yeah, I know he made a point of saying that he's not an AI expert, but I thought he did a pretty good job of explaining it. I do think the trajectory of how all this, I mean, this is such an obvious thing to say, but it's going to be really interesting to watch and how businesses adapt to this. And what's kind of fascinating to me is that we're already seeing that differentiation play out in the market with Nvidia shares up quite a bit. And Intel, which.
which is seen as not as competitive in the space down quite a bit.
I was really interested in some of his points about software in particular.
And so you think, okay, I haven't realized that.
Yeah, like, I mean, I would, you know, like sometimes I see, like someone will post on Twitter.
It's like, look at this cool thing in video just rolled out where they can make your face look like something else or whatever.
But thinking about like how important that is in terms of like, okay, you and I want to start an AI company and new ideas.
for a large language model or something specific, we have a model to train, there's going to be
a big advantage going with the company that has this huge wealth of libraries and codebases
and specific tools around specific industries as opposed to it seems like where some of the
other competitors are, or it's just much more technically challenging to even like use the chips
if they exist like Google's TPUs.
Totally. The other thing that caught my attention, and I know these are very very,
different spaces in many ways, but there's so much of the terminology and, like, that's very
reminiscent of crypto. So just the idea of like an AI winter and a crypto winter. And you can see,
I mean, you can see the pivot happening right now from like crypto people moving into AI. So
that's going to be interesting to watch play out. Like how much of it is hype,
classic sort of garment hype cycle versus the real thing. But, you know, two things I would
absolutely, you know, so two things I think would be interesting. It'd be interesting to go back
like past AI summers.
Like what were some past periods
which people thought
that we made this breakthrough
and then what happened?
So that might be an interesting.
And then the other thing is like,
look like, you know,
in 2023,
I have never actually like found a reason
I've ever felt compelled
to like need to use a blockchain for something.
And I get use out of Chad GPT on something
like almost every day.
And so for example,
we recently did on episode,
you know,
yeah,
Look, we'll do an episode now of a question at the end of like, oh, what is the difference?
Like yesterday, you know, we recently did an episode on like lending and so it's like,
oh, what's the difference sort of structurally between the leverage loan market and the private debt market?
It's like, this might be an interesting question for a chat GPT.
And like I got this like very useful, clear answer from it that like I couldn't have gotten
perhaps as easily from a Google search.
So I do think like some of these hype cycles like are really useful.
But like I am already in my daily life and very.
rudimentary reasons getting use out of this technology in a way that I cannot say for anything
related to like Web 3.
No, that is very true.
And, you know, the fact that this only came out a few months ago and everyone has been talking
about it and experimenting with it kind of speaks for itself.
Shall we leave it there?
Let's leave it there.
This has been another episode of the Odd Thoughts podcast.
I'm Tracy Allaway.
You can follow me on Twitter at Tracy Allaway.
And I'm Joe Wisenthal.
You can follow me on Twitter at the stalwart.
Follow our guest.
Stacey Razkin.
at S. Razgen.
Follow our producers,
Carmen Rodriguez at Carmen Armin
and Dash O'Bennett at DashBot.
And check out all of our podcasts at Bloomberg
under the handle at podcasts.
And for more Oddlots content,
go to Bloomberg.com slash Oddlots.
We blog, we post transcripts.
We have a newsletter.
And check out the OddLod's Discord,
people, listeners chatting 24-7
about all the things we talk about here.
We even have an AI-specific room
that's really fun.
And semi-conduct.
semiconductor room and so people chatting about these things. I even solicited some questions for today from
that group. So it's really fun. I like hang out there. You should go to discord.g.g. slash
office. Thanks for listening. The news doesn't stop on the weekends. Context changes constantly.
And now Bloomberg is the place to stay on top of it all. Hi, I'm David Gurra. Join us every Saturday and
Sunday for the new Bloomberg this weekend. I'm Christina Ruffini. We'll bring you the latest headlines
in-depth analysis and big interviews.
All the stories that hit home on your days off.
And I'm Lisa Mateo.
Watch and listen to Bloomberg this weekend
for thoughtful, enlightening conversations
about business, lifestyle, people, and culture.
On Saturday mornings, we put the past week's events
into context, examining what happened
in the markets and the world.
That on Sundays, we speak with journalists,
columnists, and key political figures
to prepare you for the week ahead.
Join us as soon as you wake up
and bring us with you wherever your weekend plans take you.
us on Bloomberg Television. Listen on Bloomberg Radio, stream the show live on the Bloomberg business app,
or listen to the podcast. That's Bloomberg this weekend. Saturdays and Sundays starting at 7 a.m.
Eastern. Make us part of your weekend routine on Bloomberg Television, radio, and wherever you get your
podcasts. What separates good leaders from transformational ones? I'm Jessica Chen, and in season two of
Leading By Example, we'll sit down with executives like Grace Chen of Bertie Gray,
to find out.
It's important to understand where you spike,
but also really acknowledge where you don't
and find people who can fill those gaps.
Listen to leading by example,
executives making an impact
on the IHeart radio app, Apple Podcast,
or wherever you get your podcasts.
