Grey Beards on Systems - 176: GreyBeards talk recent news at Future of Memory & Storage Conference with Jim Handy, Objective Analysis

Episode Date: August 22, 2026

At the FMS 2026 Conference, discussions centered on the AI supercycle and its impact on the IT industry, particularly through HBM and DRAM shortages. Innovations like KV cache offload and MRAM advance...ments were also highlighted. Despite rising prices, demand for memory remains high, especially for AI applications.

Transcript
Discussion (0)
Starting point is 00:00:00 Hey, everybody, Ray LaKazey here. Jason Collier here. Welcome to another sponsored episode of the Greybirds on Storage Podcasts show where we get Greybirds bloggers together with storage. This is the vendors to discuss upcoming products, technologies, and trends affecting the data center today. We have with us today, Jim Handy, analyst with objective analysis, focused on SSD and Memory Tech. Jim's been on our show many times before, has been our go-to guy in SSD and Memory Tech forever. We were both at the FMS conference last week in Santa Clara.
Starting point is 00:00:40 So, Jim, the keynotes had a lot on AI inference, how AI inference is driving generational changes to memory and storage heartkey. And it was an interesting session on the AI's super cycle impact of supply chains. But I must have missed a bunch. What did you see of interest at the show? Well, Ray, thanks for having me on, first of all. But the things that were most interested to me was just watching the keynotes. as far as the presentations that happened in multiple tracks at the same time, there were just an awful lot of them.
Starting point is 00:01:15 And it was hard to get straight which ones I wanted to go to. Right, right, right. Yeah, well, the keynotes were pretty informative as they were. I mean, God, I saw a lot about how AI is changing a world. Yeah, I didn't see an awful lot of new product introduction stuff, but mostly just people talking about AI and what they're going to be doing with it. probably the most significant thing was Sandisk and S.K. Hynix both talking about their high bandwidth flash. Right, right, right. Well, yeah, and last year they talked a little bit about how high bandwidth flash, but this year you could start seeing some of it being fleshed out and stuff like that.
Starting point is 00:01:58 Yeah, yeah, there was some real stuff. They finally, you know, agreed upon a spec for the thing. and something Who is right, mine would need 100 million iaps at 512 bytes in iap? Does that make any sense on a chip effectively? M.2 chip or whatever. God, it's insane.
Starting point is 00:02:18 I can tell you. Yeah? Yeah, I would imagine so. What are they going to do with this stuff, Jason? Yeah, I mean, well, the reality is like AI is driving all of this stuff. And, you know, kind of one of the things that we're, you know, see and, you know, is, you know, talking about that AI super cycles. It's like the explosive demand for HBM.
Starting point is 00:02:45 Okay. And it's all about how do you process, you know, it's all about getting processing the first token out there. And the reality is when you look at a lot of these memory producers, the reason that they're, they're shifting a lot of production over to HBM is because it's a much higher margin product for them. Yeah. Yeah, yeah. The money's there and the demand is there. And it seems like it's price and elastic. Yeah. Tell me about it.
Starting point is 00:03:13 It's like, I can go off on a tangent on price elasticity, so don't get me started on that. Well, it was actually the session on the super cycle supply chain impacts. I really talked about, you know, how the hypers and their ilk are buying up fab capacity. literally buying a fab line capacity. And, you know, there's no, there's, there's, there's, they don't care about price. They just don't care. There's, there's something kind of funny going on with that whole hyperscalary thing is that they're in a race to outspend each other.
Starting point is 00:03:50 And you listen to them talk and they talk about tens of billions of dollars and they talk about megawatts, hundreds of megawatts, gigawatts of, you know, hour consumption. But you never hear them say. we're going to, you know, process a billion tokens. They just don't, don't measure themselves by that. They measure themselves by their spending. Yeah, yeah. Yeah, I think it's going to be a really interesting topic over the next, you know, over the next year.
Starting point is 00:04:20 And I think basically tokenization is going to effectively be the next commodity market to monitor. Yeah, it would make more sense than just how many tens of billions of dollars you spend. I mean, you know, it's got to be tempting for all of the, hardware suppliers to, you know, when somebody says, I'm going to increase my spending from $10 billion to $20 billion a quarter. It's got to be tempting for them to say, okay, well, I'll take 80% of that. Somebody flashed up on a screen that's that the AI spending and may have been 2030 was projected to be $7 trillion. And it's like the fourth biggest GDP. Yeah. There was somebody. I think it was a third actually. I think it was like.
Starting point is 00:05:02 It could have been third. Yeah, it was like the U.S. and then China and then. Yeah, $7 trillion of data centers spent. I think we ought to be in a data center business here. I just want somebody to accidentally cut me a check. Well, that's true. That helps a lot. So there was some talk about new HBM technologies coming online.
Starting point is 00:05:25 I think was HBM 4E and maybe five stuff like that. So that's all what? increasing the bandwidth, increasing the transactions per second, that sort of thing that these things
Starting point is 00:05:38 can hold, maybe even capacity. I didn't see a lot on the capacity end, I guess. And, you know, that's something that's been a little bit
Starting point is 00:05:46 of a problem with DRAM. DRAM chips haven't been shrinking very much for the past 10 years. Really? Yeah.
Starting point is 00:05:52 And so because of that, the density of the chips hasn't gotten a whole lot bigger. And, you know, that's an issue. And so what
Starting point is 00:06:00 they're doing with HB. right now is that they're looking at different ways of packaging it to shorten the length of the wires that go from the processor to, or I'm sorry, from the DRAM to the processor chip. Because it goes both ways. Yeah. You know, and they're talking about what's called hybrid bonding, which is a way to put the chips inside the HBM closer to each other. Jesus. Yeah.
Starting point is 00:06:29 We're talking microns here. And isn't Samsung talking about their like Z-HBM stuff that they were talking about where they were talking about vertically stacking HBM on top of AI accelerators? Yeah, yeah. And that's something where Samsung has done an awful lot of experimentation. They've got their saint process, which is putting an SRM chip on top of a processor chip. And I know that AMD also has, I think it's called their V-Cache or something like that. Yeah, you can...
Starting point is 00:07:00 3D cache. Is that what the stack DRAM was all about? I mean, I thought it was a standard DRAM. It's actually stacking HBM on top of the compute. Yeah, it's all the like the chiplet technology of basically doing basically 3D stacking and shiplets on top of each other. Right. Yeah.
Starting point is 00:07:17 But, you know, it's in the case of HBM, it's stacking a stack on top of the processor. Yeah. Yeah. Interesting. Interesting. Yeah. It gives all kinds of new opportunities for mechanical failures. Like thermal mechanical failures?
Starting point is 00:07:33 Yeah, because something that is an issue is that you'll have these chips that, you know, the one in the center warms up and cools down while the ones on either side are, you know, just cool. And so they end up acting like a bimetallic strip. And, you know, one of them will push to bend the whole stack one way or the other. You know, some clever programmer is going to do the same thing they did with core memory back in the 60s. There were people who would program little loops in their software, and it would make the cores rattle on their wires. And they could play music by, you know, choosing how many times they flipped a bit.
Starting point is 00:08:12 I don't know if you want to do that sort of thing on these GPUs. Yeah, no, I think you could with the flexing that would go on inside the HPM. I think got hot enough. Well, yeah, well, that's the thing, too, right? It's like basically trying to maintain a core temperature across the last. layers. And so then that makes it more efficient for the cooling process. And then,
Starting point is 00:08:35 you know, and then I mean, a big thing, you know, is, is, you know, I've,
Starting point is 00:08:40 I've looked at this on kind of some of the large scale systems that we do is like, now how to get direct liquid cooling into literally all components of the system. Right. And, and this is all basically between, you know, memory, SSD,
Starting point is 00:08:56 all that stuff is like, how do you get, you know, kind of a more efficient cooling practice into, you know, all of these components, especially when you started doing that 3D stacking piece. Yeah. There was a, there was a, you know, I think last year, Solidime talked about their liquid cooled SSD, and I think this year was another vendor out there talking about liquid cooling SSDs. The memory tier SSD, what is this?
Starting point is 00:09:28 I mean, is this something is going to be sitting on the, on the accelerator? I couldn't answer that, but I will say that most SSD companies that are selling enterprise SSDs are talking about liquid cooling. And, okay, SSD reads don't consume an awful lot of energy, but SSD rights are huge energy hogs. And so that mostly goes into applications where there's a whole lot of rights. activity going on. And Jason, correct me if I'm wrong, but my understanding is that inference is not that right intensive, whereas training is like 50-50 read right. Yeah. Yeah. Yeah. It's only when you get into reinforcement learning that basically the
Starting point is 00:10:15 inferencing component. And then that kind of is model dependent and kind of application model dependent. But yeah, you're totally right. I mean, that's that's the, that's the classic, you know, kind of the raid five you know hey read you're reading 75% of the time right right they said that like the weights would be almost like worm memory so once it was written into the flash that it would only be read after that point forever yeah and then the the i guess it's the kv cash was going to be pretty much written and appended to possibly but read lots and lots of times and stuff like that Yeah, I mean, mutability on that one, it's basically meant to be, you know, kind of immutable right once as KB Cash is that that's the whole idea behind it. So, Jim, what's the difference between an append right and a random right from a NAN perspective?
Starting point is 00:11:12 I thought it was always, you know, you had to flash a page and you had to be, you have these free pages around and stuff like that. Yeah, unfortunately, I don't know what an append right is. I think you're talking software language and I suppose. It could be. It could be. I mean, you know, so random rights says I can write anywhere on the drive itself or the memory itself. But an append right says I will only append more data to the current page. Maybe that's how it works because the page is being written to. It's only, you know, it's only going to be appended to. So it doesn't have to be, it doesn't have as much table functionality, I guess. I don't know. Yeah, that actually is the case with more with blocks than with pages. With pages, you know, the pages end up being whatever, 4K bytes, 16K bytes. And so typically, the SSD will gather enough information for 16k bytes in a DRAM and then blurp out all of that 16K bytes into a page of NAM flash all at once. And, you know, I guess that what the DRAM is doing could be called appended. I've heard it called coalescing. Right, right, right.
Starting point is 00:12:26 I heard a new term called P-SLC, which I had never heard before. Yeah, it's a new term for an old concept. It's basically, okay, there's almost no difference between SLC, MLC, TLC, and QLC, Nant Flash. You've basically got the same kind of bits, but you're writing. different numbers of voltage levels onto them. You know, the higher the number of bits per cell than the larger the number of voltage levels you write on. You do that by converting your digital bits to an analog voltage level,
Starting point is 00:13:02 programmed by analog voltage level into the same bit. And then when you read the bit out, then you read an analog voltage level, you convert it back to digital. So, you know, PSLC is when you take an MLC or TLC or QLC flash, and you say, okay, well, I'm just going to bypass that A-to-D and DDA conversion part of the whole equation and use the bits as if they were SLC. And they really are SLC. So it doesn't, you know, it's not some big invention or anything like that.
Starting point is 00:13:35 So it's a lot faster. It bypasses a lot of conversion of logic, I guess. It is. And, you know, in the SSDs that do that, what happens is that it gobbles up the over-provisioning at twice as fast or three times as fast or four times as fast as it would normally get consumed if you were, you know, writing in TLC or QLC. Right. But, you know, it's, it's eventually that's going to catch up with you.
Starting point is 00:14:03 And what you've got to do is somewhere is, as one of your background tasks, you've got to convert the pseudo-SLC data into TLC to be able to regain. the over-provisioning that you used to have. Right, right, right. What seemed to me, I'm probably wrong about this, that the PSLC was sort of a, configured as sort of a cache to the rest of the storage and that sort of thing.
Starting point is 00:14:30 But, yeah. It is. It is. It just allows you to write into the flash media a little bit faster. And, you know, what that does is it reduces the amount of DRAM that you need. So you've heard of a host memory buffer, HMB. HMB, no. So there's a DRAM buffer in the SSDs, I guess, right? No, it's not. It's actually in the host, which is why it's called host memory buffers.
Starting point is 00:14:54 Basically, you, you, instead of using the DRAM in the SSD, you economize on the SSD by leaving out the DRAM. And then you mooch a part of the memory in the host. And you say, okay, you're going to pretend you're the DRAM inside the SSD. So it does the accumulation or whatever. whatever, of the block or something. And also the, I'm trying to remember what the acronym is for it right now. But anyway, the remapping of the NAND flash because of the fact that you're constantly moving things around in the NAN flash so you can do wear leveling. Huh, huh, huh.
Starting point is 00:15:29 I would take a lot of coordination between the host software and the SSD, I imagine. Yeah, it does. But, you know, it's done with a driver and it's not too complicated about a concept. and the major hyperscalers all pretty much do that so that they can build really cheap SSDs and just use their servers to manage it. I see, I see. You can do that in a closed system.
Starting point is 00:15:56 It's really good to do without the shelf software. Right, right, right, right, right, right. You know, the supply chain with the AI super cycle, you talked a lot about obviously HBO, being the high margin solution and that everybody's converting their DRAM and NAN lines to HBM to take advantage of that. But he made mention of other things like PCBs being starting to become supply constraint and the sophistication of PCBs had to go up in order to support AI.
Starting point is 00:16:32 Did you see any of that, Jim? I didn't, but I'm not surprised. You know, when DRAM went into a shortage back in. 2025, then, you know, it was first, first they started rerouting stuff over to HBM and then it kind of stole wafers away from DDR manufacturer. And, you know, that tightened up. But then NAND started going into a shortage. And that was because the guys were buying the DRAM.
Starting point is 00:17:03 We're also buying NAN flash to go with it. And they were also buying DDR to go along with it. And so I'd like to tell people that even if HBM, didn't exist, we would have a DRAM and NAN flash shortage right now. Right. And then the NAN flash shortage got worse when HBMs suffered from the same thing. Well, the suppliers don't think it's suffering. But basically, the eight hard drives manufacturers, Western Digital C8, Toshiba,
Starting point is 00:17:31 they all found out that they didn't have the capacity to support this growth in AI. Right, right. Then that ended up causing some people to use SSDs in places where they would have used hard drives. And so it worsened the SSD shortage. You know, you got this great snowballing. Well, you know, also the wafers that are being used for invidia processors are taking away from TSM's capacity to produce wafer. And so all of a sudden, everything is basically falling into short supply. So you talk about the PC boards.
Starting point is 00:18:09 I didn't know about that, but, you know, it just makes jokes. Everything's going into a shortage because of this great enormous spending binge. The guy said that, obviously, the AI servers are consuming PC boards. And B, the PCBs for accelerators have, you know, higher numbers of layers and higher frequency interconnects and all that's causing PCB supply chain problems. And then he started talking about components, things like power management ICs and ceramic capacitors and resistors and inductors. All those are starting to become short on supply because of the data center buildout. It's just insane. Yeah.
Starting point is 00:18:56 You know, and the capacity constraints on this, HBM is actually sold out through the end of 2027. up. And, you know, it's a, you know, according to many forecasts, you know, so, and, and conventional DRAM and nan, they're tight because the FABs and the advanced packaging are prioritized for the AI products. Right. So, you know, the reality as a result of that is basically there's these sharp, you know, price increases because of the, of the supply shortage, right? The DRAM contract prices are up as, much as like 80 to 95 percent and so yeah it was like you know in the long term you know like LTAs like becoming common like the memory revenue projected is is is dramatic growth and so some are
Starting point is 00:19:50 showing the DRAM revenue rising hundreds of of percentage points year over year in in 2026 and and but yeah this whole this whole memory super cycle is It's like, like I said, it's already, like, HBM has already sold out through 2027, and, and there is a high likelihood that it's, you know, in the next couple of months it's going to be sold out through 2028. You know, a guy was doing a supply, it was from IBM, and he said that he thought maybe the, the constraints would start to release around 2028, which seems to be pretty obscene. Yeah. And like I said, a lot of that shift in, a lot of shift in manufacturing is because the, like basically that HBM is a much higher margin product, right? Yeah.
Starting point is 00:20:37 Yeah. It's something that I said earlier was that even if HBM were not an issue, the way that the hyperscalers are spending is they would have consumed so much DDR that it would have still caused a shortage. Yeah. Yeah. Yeah. That's one of those interesting areas too where I still think, you know, like CXL's got,
Starting point is 00:20:56 got a playpoint there, right? It was a lot about CXL at the show. I was surprised. Yeah, no. Meta published a paper recently where they told of something that I've heard is happening at pretty much all the hyperscalers, that they are no longer buying a whole bunch of memory and throwing away servers with the older memory in them. So they've got all of these servers that they're retiring that are DDR4 based. They're scavenging the DDR4 out of those servers, flopping it in a piece, in a CXL. card and then adding that to their DDR5 based servers. Yeah.
Starting point is 00:21:38 Yeah. And effectively doing a memory tearing component, right? Yeah. They're doing the memory tearing from a technology standpoint, but from a business standpoint, they're filling the gap, the DDR5. The DDR5 can't do, right? Because you don't have enough to supply. It was actually, you know, there was kind of an interesting play to some extent at the show
Starting point is 00:22:00 about, you know, how are you going to do the KV cash? Is it going to be flash? Is it going to be high bandwidth flash? Is it going to be CXL memory? And all those are various solutions to this, you know, context memory and KV Clash. KV. Cash. I like KV Clash.
Starting point is 00:22:17 That's awesome, dude. No, no. Don't go there, man. Trademark that one. Oh, God. You guys are tough on me. So, yeah, so the CXL memory pooling is as a KV cash solution. you know is is kind of interesting so yeah yeah the other thing that was that to show obviously is
Starting point is 00:22:38 UA Link was there talking about you know supporting up to a thousand GPUs in one pod I don't know what you call a thing anymore I said a thousand GPUs I mean what are you got to seek that on an iceberg with a nuclear SMR next door to power the damn thing I mean Jesus you guys need to come out to a data center I'll give you guys a tour. Yeah, I'd love to. I'd love to, AMD. Yeah, yeah. People start talking about M-RAM. I mean, M-RAM's been at the show ever since I've been there. But it's never been a big thing. It's always this niche little thing, you know, they're talking megabits per chip and maybe they got to a gigabit chip or something like that. But people are actually thinking about M-RAM as possibly a solution
Starting point is 00:23:27 to some of those problems here. I wouldn't count on that as, you know, as something that would help a shortage. But M-RAM is gaining in popularity for a different reason, and that's because when you go to these shrinking processes, you know, 14 nanometers and smaller, they're all fin-fet-based, and you can't make Norflash on a fin-fed. And most microcontrollers have some amount of Norflash in them. And so any microcontroller that you want to use one of these advanced processes on, you know, typically a very elaborate microcontroller is going to need some other kind of memory or an external memory for its because it can't use nor flash for code storage.
Starting point is 00:24:14 And so MRAM has come along to fill that void in those particular applications, but it's not something you see a lot of in the data center. So the microcontrollers of the IoT devices, that sort of stuff. That's where it would play out. Yeah. You do see it used in outer space. Again, because of rad hardening? Yeah, exactly. And, you know, the major memory technologies all are, have a penchant for flipping bits when there's radiation.
Starting point is 00:24:46 Yeah. Yeah. Huh. I just had, you know, it's been talked about forever, but it's never been, you know, something of interest at a high level. You know, it exists and its capacity is improving and stuff like that. You're talking about something that's been around forever. My favorite is ferroelectric memory, f-ramm. Yes.
Starting point is 00:25:07 I didn't hear about f-ramm at the show. No, no. But, you know, it is a technology that's working out there. Infinion, Fujitsu, Texas instruments, you know, they've all got various chips that use it. And it's been around since, I think, 1951. Oh, God. Yeah. Interesting.
Starting point is 00:25:32 I tell you one of the other areas where I think this is actually going to be a forcing function for good is it's going to force people because there's been a lot of sloppy code written. There's been a lot of sloppy code written because basically memory was effectively not a bottleneck. It wasn't, you know, it was, it was cheap enough that you could, you could write lousy code and do some, you know, interesting, you know, just let the system take care of it, right? And I think this is going to be a forcing function for people to write more efficient code. You know, and honestly, one of the recent acquisitions that, that AMD made was, you know, a company called Maxed, where they are doing, you think. think of it as like more an intelligent version of swap, right, where you're actually utilizing AI, finding hot blocks in memory, and then taking those components and, you know, that are less hot in memory. Because what you find is, like, a lot of the software is just like, yeah,
Starting point is 00:26:47 25% of that memory is actually being used, but it's like claiming, you know, like 75% of 80%. So you can take that, tear it off, you know, tear it off to those lower, you know, either, you know, be it CXL or or even, you know, SSD tiers, basically being able to tear that stuff off so that you can actually get more efficient utilization out of, out of that memory component. And, you know, like we saw this like memory thing coming. and, you know, it was, we thought it was imperative to make some, you know, investment decisions based on, you know, how do you more efficiently utilize this memory without driving, you know, cost through the roof for a system that's, you know, already, you know, not a cheap system, especially if you're looking at helios, right?
Starting point is 00:27:45 Yeah. How do you make it more efficient? You wouldn't think AI inference could be a player in a virtual memory select page selection loop, but I guess it is. At some point. Yeah, Ray, I have a very different viewpoint than you on that, and that is that I think that what AI is doing, you know, and all of the learning that is going on right now, we'll end up transferring down to very low levels. where there will be places where people right now are coding, you know, very deterministic things that they don't need to be deterministic. And you'll end up finding things like, you know, health monitoring watches that have AI inside
Starting point is 00:28:32 that nobody knows about, but, you know, it does something to optimize something or other. Yeah, yeah. And for years, I've been, you know, kind of watching the way SSD controllers work because they have to do an awful lot of packaging data back and forth. The session that Fado did on their keynote was so inspiring. It says, like, Nan is the worst technology in the world, but it happens to be cheap. Yeah, yeah. But yeah, there will be AI inside of SSDs.
Starting point is 00:29:04 There already has been, except they haven't been successful efforts. Yeah. And, you know, there will be AI in an awful lot of other things. Something that I like to remind people is that you can run AI algorithms, in today's MCUs, it'll run really, really, really slowly. And, you know, and you want to limit yourself, you know, to maybe five input parameters or something like that. You know, it's supposed to billions or trillions of parameters.
Starting point is 00:29:30 Yeah, layers and stuff like that. Yeah, yeah. But, you know, there might be certain cases where that's a good, good approach, and it's a better approach than just hard coding something. You could see ML Perf. They've started to do some benchmarks for, IoT devices as well as mobile devices and things of that nature. So you could see it starting to move in that direction. Yeah, but probably today it's used in, you know, for doing big,
Starting point is 00:29:58 complicated things. And I picture it being used for doing small, simple things later on. Right, right, right, right, right, right, right. So what do you think about the show? I mean, it seemed like it was about the same size as Expo floors as last year. The new guys seem to have a different attack on a couple of things, I thought. You know, some things were good. Some things were not so good, I guess. Yeah. And I think that's going to happen anytime you have a change of leadership.
Starting point is 00:30:32 Yeah. You know, what charmed me was that it didn't fall apart at the seams. Typically during a transition, that happens. Right, right. So, I mean, obviously, I think there's a need for the show. And that's been present for quite a while. But yeah, yeah. They seemed like the transition went okay, I guess.
Starting point is 00:30:54 Yeah, I think it went pretty well. Yeah, yeah, yeah. Yeah, today all the presentations went online. So, you know, anybody who was at the show can download to their heart's content. Right, right, right, right, right. Well, yeah, it can only take so many sessions with one body, you know. So a lot of the other sessions were probably pretty interesting as well. Yeah.
Starting point is 00:31:18 And if you're like me, you know, you've got somebody saying, hey, Ray, you're here. You made it. Let's talk. And you end up missing one of the sessions that you had intended to go to. Yeah. Yeah. Well, I also try to talk to the vendors of interest to try to generate some business and stuff like that. So there's always that.
Starting point is 00:31:38 So one of their technologies did you see it to show, Jim, that were kind of of interest? Hmm. Naturally, everything I talked about was interesting. Of course. Of course. So, yeah, that's CXL. That's hyperscalers. That's whatever new memory technologies. But, yeah, you know, there, let's see. It was surprising to me.
Starting point is 00:32:04 There wasn't a lot of talk about capacity increases. In the last couple of shows, like 100 terabytes, 200 terabytes of drive. And they were talking about, you know, doubling the last. layers of NAND and all that stuff. All that's sort of fell by the wayside this time. Yeah, people aren't talking too much about layer count and NAN Flash anymore. And it's it for a very strange reason. It's because the staircases are getting too big. Explain that, Jim. Okay. So, so the way 3D NAND works is that they've got a certain thing that's a staircase, which is you've got these layers. And now the question is, how do you attach a
Starting point is 00:32:42 wire to each one of the layers to communicate with it. And the way that they do it is by creating a staircase. And so if you have, you know, let's just say for arguments sake, you've got a 32 layer and end and it's going to have 32 steps on the staircase. And so it's going to take up a certain amount of room. You double that to a 64 layer in and all of a sudden you've got 64 steps on the staircase and it becomes twice as large. And so that's become an issue. I'd so these guys that I literally have 200 layers, right? I mean... Yeah, and they've got 200 steps in their staircase.
Starting point is 00:33:17 But that's not the only problem. The other problem that they have is that each one of those layers adds capacitance inside the chip. And so that capacitance needs to have... The larger the capacitance gets, the more current you need to drive it. And they got to the point where the current requirements were strong enough that they needed to use these huge transistors to turn them on and off. And that ended up determining the size of the chip.
Starting point is 00:33:47 It wasn't the number of bits and the thing. It was the size of those transistors. And so that's why Sandusk and Toshiba with their Bix 8 and Bix 10 generations of flash have decided to start using something that Chinese manufacturer, oh, come on, I was going to say the YMTC. I almost said the DRAM manufacturer's name. But YMTC started it. They called it stacking spelled with an X instead of an S. And, you know, Western Digital and Kioxi are doing their own version of that. But basically, eventually everybody's going to be doing that just because the layer counts are getting high enough that they need to have.
Starting point is 00:34:31 Those big drivers that you get by using a logic process for the transistors. You know, what they do is they make two wafers. One wafers got all the bits on it. and the other wafer's got the drive transistors. Oh, God. And a chiplet interconnects, right? No, no. Actually, they do what's called hybrid bonding.
Starting point is 00:34:49 They take these two wafers and bond them to each other face to face before they split them apart into individual chips. Oh, kidding. Yeah. So, and then I'm sure that's going to effectively drive up what the overall, you know, wattage that is the drive is going to be consuming, which is, you know, when you're talking to AI data, centers. Yeah, sure.
Starting point is 00:35:12 I'd throw another bunch of watts on there. Yeah, you know, Ray was talking about, I think they were called SMDs, the nuclear, little nuclear power. Yeah, yeah, nuclear reactors. SMRs. Yeah, yeah.
Starting point is 00:35:26 Maybe, maybe what we need to do is to have micro reactors that are attached to each SSD or each GD. Please, please. That's insane. Well, I never saw I'd see liquid cooled SSDs. So the world is changing under our feet, gentlemen. Well, they're talking about liquid-cooled HBM, too.
Starting point is 00:35:47 Oh, God. The way that works is that you put the HBM on its side and then somehow run fluid in between the different dice on the HBM. Yeah, and that HBM is typically, like, you know, in the perspective I've been working with it, like that's attached to the GPU. So it's basically, it's kind of natural to liquid cool it anyway. I guess that makes sense.
Starting point is 00:36:13 And yeah, you should see the size of the things going into the back of those helios racks. It looks like basically four fire hoses like going into the back of the helios rack. To support the cooling. Cold water and then two of them hot water exchange. Just for grins, Jason, how many servers and GPUs in a helios rack? 72 GPUs, 18 sleds. as far as like trays that are running
Starting point is 00:36:43 in Venice CPUs. Yeah. But yeah, and but a rack sucks about a one quarter of a megawatt of power. Jesus. Wow. So it's called
Starting point is 00:36:55 Helios because it's as hot as the sun. Just about. I think it's a Greek thing, but hey, you know what? Yeah, no, Helios is the sun god. Yes. And yes, it's as hot as the sun. Yeah. Oh, God. What else is going on here? Yeah, I'm trying to think of other interesting things that I saw there.
Starting point is 00:37:19 There were a lot of people who were talking about how everything has changed and the memory prices are going to stay up. I'm sure a lot of OEMs were saying, shoot me now. Right. But, you know, it's still a commodity market. And I think a lot of people aren't recognizing that. And there is probably going to be a crash at some point where hyper-scalers decide that, okay, we've built enough. We're going to stop for a while.
Starting point is 00:37:48 Yeah. Yeah. Well, the whole debt picture starts to intervene here and the interest rates and all that stuff and bonds payments and things. Yeah. You know, but we did see that happen with the internet bubble. Right, right. You know, hey, we built enough internet here.
Starting point is 00:38:04 There isn't room for anymore. Yeah, I remember when events. Surf was, he was the CTO over at MCI, and then he's like, the internet's grown at a thousand percent a year, like back in the late 90s. And MCI, like, built out all this fiber. And then, you know, by the time the dot-com crash happened, like 98% of it was dark. Yeah. Yeah.
Starting point is 00:38:30 So I think there's, you know, I think there's a real opportunity to, when you're looking at it, to operationalize AI and to get as much. out of as you possibly can, right? Right. It's like not just building it for the sake of building it. There needs to be more operational efficiency put into the way that actually GPUs are run. Right.
Starting point is 00:38:53 And I think a lot of the, you know, the, you know, between the, you know, open AIs and Anthropics and GROC and all those folks, they realize that, you know, as much as they're utilizing the GPUs, I think, There's still a lot of efficiencies that need to be made that can reduce that overall cost and spend for them. You actually brought that up before during this call too, Jason. And I thought about saying something that I used to say in the 90s, which was that people talked about Moore's Law stopping. And I said, if it stops, we'll start using those transistors a whole lot smarter than we have back when we thought that we could get as many as we want for free. It's, I think, I think we're an interesting point where there's a, there's a lot of ways in which we can operationalize efficiency that's been overbuilt at this point.
Starting point is 00:39:49 So, so you think AI is going to be used to do that? For sure. For sure. My whole coding world has changed in the last six months. Yeah. I don't, I don't code anymore, which is a thankful thing I might add. but I do use AI to do all that stuff. I mean, they're using AI to do routing. They're using AI on the chips themselves to, you know, operationalize and, you know, better efficiency on the chips and stuff like that.
Starting point is 00:40:22 Yeah, and AI is really good at those problems with a whole lot of dimensions. Right. With routing, you've got, you know, signal length and the capacitance and, you know, hot spots on the chip that you have to worry. about and also minimizing the diarrhea, the chips so that the costs stay in control. So yeah, it's a good thing. Yeah, now one of the things I've noticed about AI, you know, in the last, you know, nine months that, you know, I've owned a Tesla with full self-drive.
Starting point is 00:40:53 One of the things that I've actually noticed about that is not necessarily the fact that it can process data better than I can. The reality is it can process more inputs at the same time than I can. It's got like eight cameras on the thing and has the ability to basically make decisions based on things that I can't see that's effectively a blind spot to me. Yeah. Right. Yeah. And I think AI, when you think about what it can do from operationalizing, you know, a lot of IT environments, it can observe things that you can't observe because it's got more inputs coming into it.
Starting point is 00:41:30 Well, yeah. Think about how many times you look over your shoulder before changing lanes and do stuff like that. You know, it's, your stupid head only has eyes on one side. I was in a Tesla. My friend had full self-driving enabled. It cut across six lanes of traffic to take a left turn. This is amazing. But yeah.
Starting point is 00:41:50 Yeah. Yeah. You know, it's like the power, I think, I think a lot of things that people miss is the fact the, the number of inputs and the observability that it has is far greater than, you know, its ability to process the data. I mean, if we were looking specifically at that data, yeah, we can process it just fine. But the fact that it can take so many more inputs in,
Starting point is 00:42:13 that's where the operating, making it operational. Operationalizing IT, you know, that becomes a very, very, very powerful component. I would say in the last 12 months,
Starting point is 00:42:31 maybe 24 months, you know, the whole discussion on AI, has gone from training to inference. And now from inference to agenic workloads. And I think one of the keynotes, the guy had an agenic workload generated 4,5 million tokens or something like that. It was obscene. Ray, I think I told you this probably three years ago.
Starting point is 00:42:53 I'm just like training is where you spend money on AI and inferencing is where you're going to make money on AI. And I think we're starting to see that come to fruition. Yeah, but Jason, maybe agents could be a little bit more frugal with the number of tokens they used. Maybe. This is why I run it. So, you know what? I think another paradigm component that we're going to see is actually the utilization of almost, you know, this switching component where you can run stuff local. Yeah.
Starting point is 00:43:28 I've got some of our like, rising, basically our rising AI boxes. you know, at, you know, at my house. And, you know, it's one of those things where if it's a non-critical workload, I can run it on that. But then I also have the capability of basically shifting, you know, shifting gears up into that token. And that token spend is going to be huge. That token spend is huge.
Starting point is 00:43:53 I've talked to several of the larger cloud service providers. One of those loud, those cloud service providers, their token spend went up to $4,000 per week per engineer. Whoa. Yeah. So they're not happy. The CFO is not not happy. Yeah, yeah, yeah.
Starting point is 00:44:22 I like that you called them loud service providers. That reminds me during FMS there was an ACDC concert going on across the street. Yeah. I heard from somebody who attended it that it was a really, really loud concert. All right, Jason, any last questions for Jim before we low? Go? No, I did the great talk as always, Jim. Always good to hear your perspective on this. Jim, is there anything you'd like to say to our listening to the audience before we close? No, you know, I've come out with a new report on CXL and Tom Coughlin and I have a report on new memories, which is why I seem somewhat conversant on M-RAM.
Starting point is 00:45:04 Good for you. Yeah, I got an HBM report coming out. So anybody who wants to throw any money my way is welcome to. Good, good. Well, this has been great. Jim, thanks again for being on our show today. Oh, thank you for having me. It's always really a lot of fun.
Starting point is 00:45:20 Yeah, that's it for now. Bye, Jim. Bye, Jason. Bye, Ray. Until next time. Next time, we will talk to the system storage technology, If any questions you want us to ask, please let us know. And if you enjoy our podcast, tell your friends about it.
Starting point is 00:45:37 Please review us on Apple Podcasts, Google Play, and Spotify, as this will help get the word out.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.