@HPC Podcast Archives - OrionX.net - @HPCpodcast-110: John Gustafson on AI, HPC, Making Every Bit Count – In Depth
Episode Date: September 17, 2026Can AI and high-performance computing get better answers with fewer bits? HPC pioneer John Gustafson, author of Every Bit Counts, joins Doug Black and Shahin Khan to challenge the default reliance on... 64-bit floating point, and explain why numerical precision belongs at the center of the debate over computing performance, memory bandwidth and energy use. [audio mp3="https://orionx.net/wp-content/uploads/2026/09/110@HPCpodcast_ID_John-Gustafson_Numerical-Precision_Posit-Format_20260917.mp3"][/audio] The post @HPCpodcast-110: John Gustafson on AI, HPC, Making Every Bit Count – In Depth appeared first on OrionX.net.
Transcript
Discussion (0)
So if what you're paying for is kilowatt hours, going to half precision is a factor of four at least.
They call it quantization error in AI.
You really didn't need to make up a new word for that.
It's called rounding error.
We're getting all these sparse matrix problems that now solve with 32-bit pauses that used to require 64-bit floating point to do.
If you're doing something like nuclear, that's where I would reach for interval arithmetic.
There's one place where you need absolute confidence.
that would be the best possible benchmark for high-performance computing, because you make FFTs
faster, you're going to make everything faster.
From Orionx.net, this is the At-HPC podcast.
Join Shaheen Khan and Doug Black as they discuss supercomputing, AI, quantum technologies,
and the applications, markets, and policies that shape them.
Thank you for being with us.
Hi, everyone, I'm Doug Black.
This is the At-HPC podcast, and with me is my co-host, Shaheen,
of OrionX.net.
Shaheen and I are very pleased today to have with us as our special guest, John Gustafson,
who is a very well-known computer scientist and businessman and an HPC pioneer.
He's known for, among other things, the invention of Gustafsson's law,
the introduction of the first commercial computer cluster,
and the invention of new number format and computational systems.
John is currently chief scientist at VQ Research and is a visiting scholar at Arizona State
University.
He has held senior positions at AMD Intel Labs and other companies.
John, does that about cover it?
Would you like that?
That covers it perfectly.
Thank you.
Yes, the academic people think I'm a businessman and the business people think I'm an
academic.
Yeah.
He's also the author of some very influential books in the field of computer science that we'll
get into. The very hot topic these days in AI and then AI for science and how that works with
HPC is numerical precision. You've been looking at it for decades. As I say, it's become a very
hot topic. Yes. Could you share with us some of your most recent thinking or activity in this
area? Well, first of all, Moore's Law is not getting us more speed through smaller transistors
anymore. And really, we've been communication bound for decades. So the only
way I see forward is to get more information per bit. If we're going to communicate a bit,
it better mean as much as it possibly can. And we've been very sloppy about how we communicate
real numbers in particular. We're pretty sloppy about integers, too. But the real number crisis
is really hitting the AI community. And now that they're using megawatts for months to try to
train their systems, they're starting to realize that they're spending an awful lot of electricity
just communicating bits back and forth. So anything we can do to get more information,
about real numbers into a small number of bits is going to save millions of dollars, maybe billions.
Shane, do you want to dive in with a follow-up?
Yeah, Shaheen, you mentioned mixed precision, and that phrase is used a lot, but I think I'd rather
you could say right-sizing precision, so don't use any more precision than you need.
But I also am wary of putting a big burden on people to pick the precision.
I think we need to let the computer pick the precision and do as much of it automatically as possible
because programming is hard enough as it is.
But people have a certain accuracy in their inputs,
and they have a certain accuracy they want the outputs.
And our habits have been to just throw 64-bit floating point
at everything in between.
And that's really very, very wasteful.
And I'm finding increasingly that I can certainly get it down to 32 bits
and still get very acceptable answers.
And that's more than a doubling of speed
and a halving of the energy and power required.
Well, so you're taking it.
like to the next level by saying,
normally do I have a better way of representing numbers,
I can do it dynamically so that you will get just as much as you need and not more.
And of course,
the title of your latest book is every bit counts,
and that's very consistent with that.
Yes, every bit counts and every bit pattern counts.
If you take a look at a current day floating point,
you say, what are all the possible bit patterns,
and what is the complete set of numbers that represents,
then you look at what numbers do I actually need my application?
It's about 2% of those bit patterns.
So there's a lot of headroom for improvement in the way that we represent real numbers.
We do not need 600 orders of magnitude in order to do almost all of scientific computing and AI.
From what I can tell, there's been a divide between AMD and their upcoming generation of accelerator.
chips and Nvidia.
And Nvidia seems to have some workarounds for using lower precision.
Meanwhile, what is it she and the MI430X?
That's the next generation.
The one they just announced.
Yeah, it will include 64 bit.
John, do you come down on either side, which company or maybe taking more interesting
approach or the better approach?
Well, both companies are really taking the same approach, which is what's the smallest
possible amount of engineering we can do in order to get by.
Okay.
What is the least amount of thinking and originality that looks like what we've always been doing
for decades?
That seems to be the main goal of the engineering.
They don't sit down and do a clean slate design of what is the set of numbers you need
to represent and what is the most efficient way to represent them, which is what my arithmetic
is designed to, the positive arithmetic in the latest instantiation.
So every time I've come up against a comparison on AI of let's try it with different number formats.
The posit format that I introduced almost 10 years ago always comes out on top.
It does the best job with the smallest number of bits.
But so far, only small companies are adopting it in hardware.
And I think eventually it's going to find its way into Risk 5.
And those risk 5 processors with positive arithmetic are going to be used in AI centers to good effect.
And what will that mean? Just better results, greater efficiencies, which as we know is a huge issue right now or both? Or what will be the result of that?
Right now, I think the large language models are being mostly trained on 16 bit. Mostly Google's B float 16, if you're familiar with that. It looks like a 32-bit float with the bottom bits kind of saw it off like a shotgun. That gives you the enough dynamic range, but it's not really ideal for the accuracy that you need.
It's funny, they call it quantization error in AI.
You really didn't need to make up a new word for that.
It's called rounding error.
It's been around for a while.
Rounding error is the problem.
And if you only have eight bits of significance,
that's really not enough for the final stages.
And it's also usually not for the inputs.
You need higher precision at the beginning of a training,
and you need a higher precision at the end.
I think that's also true for inference,
because you're trying to distinguish between numbers that really are spaced rather closely together
when you're trying to classify images or sounds.
And if you're dealing with an image where you have red, green, blue, eight bits each,
and obviously you need more than 16-bit precision at the beginning.
But then in the middle, it's amazingly tolerant of low precision.
And so people are starting to exploit that, but I can get down to eight bits with no problem.
And Facebook took a look at my number system,
And they managed to squeeze it down to five bits and get accurate inference with five bit posits.
So at least Facebook adopted some of my thinking.
So, John, paint us like the bigger picture of what is the complexity in number formats?
Because certainly the existing situation is that you get a chip and it supports certain fixed formats and that's it.
Right.
So we fixed, I think we got integers right after the 1950s.
In the 1950s, IBM did sign magnitude format.
They said, let's use the first bit for the sign and the rest of it for the number because they wanted to look like a grade school arithmetic.
There's this human, very anthropomorphic tendency to make things look like whatever we learned in school.
And that's not necessarily what you want for a computer.
But pretty soon they realize that what you really want is two's complement, which, you know, it gives you exactly one way to represent the number of zero.
it goes up halfway and then it flips over to the negative and counts back down to zero.
And that's really, really efficient for hardware and very logical.
And the integer arithmetic is kind of perfect as if you don't overflow.
If you don't go outside the range of the integers and you can multiply, divide, and subtract
and add to your heart's content.
But that's not at all the way real number formats were designed.
The real number format stems from scientific notation, which was invented in World War I,
in 1914 in Spain by Torres-Caedo.
And he just came up with it as a kind of a lark.
Let's just use two integers,
one for the significant and one for the exponent.
And the people have been gravitating back to that ever since
and not really thinking it through.
I believe that real number of representation
should be like integers and that you should be able to read it
from left to right.
Each bit that you read gives you more information
about what range, what part of the real number line you're looking at. And that's the way positives
are defined is you can actually read them from left to right. They map perfectly to the two's complement
integers. And that means you don't need a different comparison unit. It means you don't need
it more hardware to do in the negative of a real number. You get a lot of things for free that way.
And it's just, it's much more elegant and much more mathematical. And it takes care of the really
weird set of exception cases that are in floating point arithmetic. That's brilliant, especially
the analogy with two's complement really is very interesting.
But the other part of this is that you're trying to map an infinite number of numbers
to a fixed number of bits.
Yeah, I always laugh when people bring that up because for some reason it didn't bother them about the integers.
No, it does bother me.
There's an infinite number of integers, but it bothers them when they start talking about real numbers.
I see.
But even for integers, I have the question, is that we uniformly cover every integer until we overflow.
Is that what we want?
Should we just skip a few?
Should we just do prime numbers?
What is like how the distribution profile of an infinite set to a finite set should work?
What are the parameters that would impact that?
Well, think about any integer on a computer, let's say any bit integer.
You count up, then it overflows, which means it turns into a negative number,
and then it counts back up to zero.
So it makes a circle.
So you can map that to a circle.
Well, you can also map the real number line onto a circle.
If you put a circle on top of the number zero and imagine a light bulb projecting from the top of the circle, it projects down onto the infinite real number line.
And that's called the projective reels.
So suddenly that infinite set of real numbers looks like a very finite circle.
And I map the two's complement integers into that and then I map the real numbers onto exactly the same way as to where they project on the real number line.
And you get something that's defined as rigorously as Dedican cuts, if you're,
mathematician and you can actually, you can make sense of things. It's not like just let's cut
this 64-bit number up into bit fields and make up things that each bitfield represents. That's
where we were for most of the history of computing is bitfield mentality. And it makes a real
mass. That's where you get things like negative zero. Well, the other part is that because you have
systems that are bite addressable, it's kind of harder to be dynamic about all of this. Otherwise,
you would say for this number, I only need 17 bits and for the other one I need.
Yeah, exactly.
So my view is you need an approximate way to represent real numbers that's rounded.
That's what everybody accepts.
But then there's one place where you don't want to be approximate.
And that's when you're doing sums and dot products.
And people reach for 64 bit thinking that that's going to give them a decent dot product
or a decent summation.
But in fact, what you really want is a fixed point number that covers the dynamic range
and makes no rounding errors.
So you can total up any number of this wide range of real numbers that you're representing.
And that's the approach that Ulrich Kulish came up with in the 1970s.
And he always tried to get it into the I-Triplea standard but failed because of a lot of reasons.
One is that there's so many exceptions in I-Tri-E like Infinities and NAN, not a number, stuff like that.
How do you turn that into a fixed-point number?
You really can't.
It makes a huge mess.
But my number system is designed to actually use those six-point numbers.
register. I call it a choir, Q-U-I-R-E. And the choir accumulates both sums and sums of products
without any rounding error up to like two billion numbers in a row. And that means that all your
linear algebra becomes extremely reliable and solid. And that's why we're getting, we're getting
all these sparse matrix problems that now solve with 32-bit posits that used to require a 64-bit
floating point to do. That's a big savings. Wow, that's huge. You mentioned I-T-E-E. Let's chat a
little bit about how that came to be. Yes. Well, there's a misconception that was invented by William
Khan of Berkeley. He was actually not the inventor. The person who selected the format and came up with the
rules was a guy named John Palmer, who I worked for when he became the CEO of NQ. And I know the story from both
men. I know them both quite well, spent many, many hours talking to them about where things came from.
And Khan did an interview in which he said the only things he added to the ICCI standard was
the nan data type, not a number, and the inexact flag within the processor that a flag would
go up if you had rounded the number. Other than that, it was all the choices of John Palmer,
who was really basing it on format from deck, digital equipment. It was very similar to a digital
equipment format. It was also kind of similar to the one user floating point systems.
You might be called it. Yes.
At floating point systems, we had a 64-bit double that had an 11-bit export.
That's right.
But it used to his complement.
And it was to come.
That's right.
Yeah.
But Palmer reverted to sign magnitude format and these goofy things with a bias in the
exponent, which increases the cost of adding exponents.
And a lot of other choices that were just quite dubious.
And then he hired William Kahn as a consultant.
And Kahn had a lot of very different ideas about how to build a number system.
He was in very stark disagreement with.
Palmer and they really duked it out. But eventually he gave in to every single one of the
design choices that Intel had already made because they said they were too far down the design
of the 8087 co-processor. And I remember John Palmer gleefully tortled to me at NGube that
whatever is in the ICCLEA standard, that's the 8087. And they got, they managed to hijack
the ICCI standardization process and do exactly what you most fear will will happen, which is
that just one company dictates what the rest of the industry has to do, whether it's a good idea or not.
And that's what we were stuck with.
Oh, interesting.
So they were just too far up the road to quit, basically, and they had to absorb it.
They had the co-processor designed with an 80-bit temporary register, and constant it should be 128-fits, but he lost that one.
He also wanted its support for binary, I mean, for decimal expression, like you get on an HP calculator,
because he didn't want there to be any rounding error in going into a co-processor and then coming back out.
converting from binary to decimal.
He lost that one too.
He also wanted perfect reproducibility.
And they said, no, we don't want perfect reproducibility.
We want it to be different on Intel processors
because we've got 80 bits to guard the precision
and we want our numbers to be better.
So if you actually, if you do a 32-bit calculation
or a 64-bit calculation, honestly,
you're not going to get the same as something
that's secretly and covertly using 80 bits.
And I thought that would actually help sell co-processes.
Maybe it did.
But bit reproducibility was not been in floating point arithmetic for forever.
That's right.
No, I remember Rudvanda Pass's examples with very simple Fortran and how I Tripoli just absolutely fails.
And we did have a working group for posits that did finally agree on a standard.
And if you follow that posit standard from 2022, you will get bitwise reproducible perfect results,
even if you change the order of addition and even if you go parallel,
it still gives you exactly the same answer because it restores the associative property
of addition, which is completely absent in floating point arithmetic.
That's true.
That has been a favorite rant of how can basic laws of arithmetic not hold.
Exactly.
But you can get them back with positive arithmetic, and people are discovering that.
John, we've seen the expansion and the various LNPAC benchmark.
used in HPC.
Yeah.
Do you have views about that, especially as it relates to the whole mixed precision issue?
Well, I understand that you sometimes do things in making up the rules of a benchmark that
don't age well, but you have to do it in order to be fair.
So one of the rules of the HPL benchmark based on a Limpact is that it has to use 64-bit
representation of floating point.
It does not have to be a standard floating point, because that benchmark's been around
since the days of things like the Kray won.
And I've actually written a paper about this
that's being refereed right now.
But the original benchmark was to solve
100 by 100 set of equations with 64-bit floats
and see how fast you could do that
using just standard Gaussian elimination with pivoting.
And then that got to be a really silly benchmark
because it was so small that a Fujitsu could do it
in like 150th of a second.
And it was too small to measure.
So Dengera raised it to 300 by 300 and relaxed the rules slightly as to what the compiler could do.
And pretty soon, by Moore's law, that also became too small of a problem.
So he went for 1,000 by 1,000 and fully allowed parallel processing to be used.
And I had a lot to do with convincing him to make matrix, matrix multiply the kernel and let people be flexible with that.
Because remember, we had the 164 max machine.
Yeah.
and that was exactly when I was talking to Dengarra to let him relax the rule so we could use matrix accelerators.
And he agreed with me on that.
But finally, I just said, why do you make it scalable and use weak scaling?
You know, Gosperson's law?
And just record how many operations per second you can get doing Gaussian elimination on any size problem that fits your memory.
And he did that.
And as a result, we've had the same benchmark since 1992.
Once you make things scalable like that and use weak scaling instead of trying to use Amdahl's law,
then you're going to be able to keep the benchmark going forever.
What is that?
A lot of decades here we've been using the HPL benchmark.
60s or plus additions and twice a year.
So that's like 33 years or something.
Yeah.
Now the way the benchmark is set up, you generate random numbers in the matrix.
You add up the rows and set that to the right hand side so that it looks like the right answer
is a vector of all one values.
But that misses the fact that you're rounding
when you sum up all the values.
So in fact, the exact answer is not exactly all one values.
If you fix that and tweak the random numbers
so that they really do add up to a representable right-hand side,
then I can actually solve that with 16-bit pauses
and nail it every time.
Oh, interesting.
Yeah.
Well, that kind of leads me to the question of
error propagation. Yes.
Because I sort of naively understood early in my career that all these, you know, like the
FPS 190L was 36 bits or 38 bits was it? 38, yes.
38 bits, right. The joke in the Salesforce was, why 38? Oh, it's two times 19.
But, but like I asked, like, why is 38? And I was told that for signal image processing,
which is like what it wanted to do.
The real desired numbers was 38 and look at this CDC,
and CDC was also like some other weird number.
It wasn't 64 bits.
Well, let me use this platform to tell the world the real reason why it was a silly number like 38.
It's because they were an accelerator for all these mainframes,
machines from all different vendors that had different choices for the fraction size
and the exponent size.
And they never wanted to be accused of flipping the dynamic range.
So they had to have at least as many exponent bits
as the largest exponent they were going to accelerate.
And they had to have at least as many fraction bits
as the largest fraction they were going to accelerate.
And that meant that they had to go from 32 all the way up to 38 bits
to accommodate everyone with no loss.
So they wanted to be a superset.
It's a super set.
If everybody's existing 32-bit floating point,
that's what 38 came.
from. Huh. Okay, so it had nothing to do with error propagation. Well, it's certainly handy to have,
you know, a nice big fast. Good, of course. And so you never have to apologize for degrading the
quality of the numerical computation. But yeah, it was a... But it's not like somebody did the analysis
and said for this kind of a math, this is how errors propagate. They never do.
There's many bits to contain it. The main motivation for these like large exponents is that
that takes away from the fraction, and that means your multiplier doesn't have to be as big.
It's a convenience for the hardware designer if, let's say you only have a 48-bit fraction in the cray.
That was the cray, the format.
Right.
And then because they really struggled to build a 48 times 48 integer multiplier for the significance.
And in fact, they did a sloppy job of it and left off the least significant bits.
As a result, on a craye one, you can take A times B and get a different answer from B times A.
Interesting, nice. So they were ahead of the game with I-Tripley.
Well, yeah, they were trying to pack as much arithmetic into that machine as possible.
Remember, they only had like four gates per chip.
Right.
Fairchild technology.
Right, right, right.
So then likewise on the software side, everybody started using 64 bits because, well, it was there.
Yes.
And in the early days, it wasn't even slower.
Right.
They had these like partial differential equations that have a mass.
a massive amount of discretization error, just in there's not enough grid points to do it very
accurately.
You might get maybe a thousand grid points in each direction, so you only have an accuracy of
a tenth of a percent.
But because there's 64-bit arithmetic, you say, well, let's at least make sure that the arithmetic
rounding error is not the problem.
And then we can just worry about dealing with the error that we're made by pretending that
a calculator or a computer can do calculus.
which you can't.
You have finite differences instead of real differentiation.
Well, how does your number format deal with that aspect of thing,
but just like the algorithmic error situation?
And I remember our mutual friend, you know,
Bill Walster and Interval Arithmetic,
and other efforts and methods to try to contain errors as computation
proceed?
Well, my first book,
The End of Error,
was about fixing all the difficulties
of interval arithmetic.
But in the marketplace of ideas,
interval arithmetic is a loser.
I have to,
you take a look at it,
how much scientific computation today
is being done with interval arithmetic?
I'd be surprised.
Oh, because it's too slow, right?
Well, it's not just that.
It's that you have to have
even more expertise
than a normal computer programmer
working with real numbers.
You have to know how,
to stop the problem of the interval expanding.
So your answer comes out negative infinity to infinity,
and you say, that wasn't useful.
I'm not going to do that again.
Because if you don't know what you're doing with the interval arithmetic,
the error bounds become way too big to be useful.
And as I said, that's what my first book was about.
But I also understand people want a system that guesses its way
through a computation and rounds.
And they've been doing it forever.
and I thought about calling them guesses, the guess being the name of the number type,
but that was kind of a pejorative.
So I set up on posit instead, which is an assumption that you expect to be true, you hope to be true.
And so it's a little bit more, you know, less insulting to the number system.
If you can posit that this is a number that's close enough and then do everything in your power to make it accurate,
then it's going to do a better job.
Yes, it sounds kind of sciencey, if you will.
I suppose.
I always use five-letter names for my things, like choir, posit,
and there's something called a tack-um that I encourage them to name a guy named Blaswell Hunhold came up with.
I said, look, you can do a text substitution of float for a posit,
and it won't change your pagination.
Nice.
It's meant to be a drop in replacement for floats.
Excellent.
Then the other topic that is current is so-called emulation.
And Ozaki model gets a lot of airtime.
Yes.
Which, of course, I kind of take exception with calling it emulation,
but I understand that there's a difference in different disciplines
what the meaning of the word emulation is.
So I sort of resorted to calling it algorithmic emulation
because it is context-aware.
Yes.
And I think if you're context aware, then you cannot be emulating.
It's like you have some side knowledge that is impacting you.
Whereas the context unaware, and I call that like numerical or arithmetic emulation,
because now you really don't know anything about.
So what is your take on the Ozaki model?
And is it really true that scientific computing substantially is mappable to Ozaki?
And the part that isn't, I'll just kind of do it.
the slow way.
Yeah.
I'm going to come out ahead.
Well, there's a couple things that the Ozaki approach reminds me of.
One is back in the 1940s, and this is partly because of John von Neumann, he really wanted
people to think about the location of the radix point in the number and program it manually.
He was opposed to the idea of using bits in the number to represent an exponent.
He thought that people would be sloppy with it.
They'd make all kinds of mistakes, and that just was not good programming.
And he was right.
People are sloppy with it.
But eventually, Floating Point was introduced by IBM and caught on,
and people started depending upon it.
But before that, they had to do things like block scaling of a set of numbers.
You say all of these numbers are actually scaled by 2 to the negative 6th or something like that.
And the programmer was required to keep track of all that.
And it was very, very tedious.
And what we now see is a kind of a reversion to the bad old days of asking programmers to manually manage the look
of the radix point in a fixed point number and float these numbers back and forth with
a recorded shift in the exponent instead of letting the hardware do it like we we had the luxury
of doing the other thing I see about Ozaki is do you remember I don't know it was 15 years ago
that somebody noticed how cheap Sony Playstations were oh yeah absolutely and it said let's build a
cluster out of Sony Playstations like 300 right on just go down to the target you know and or
or Walmart and get yourself a whole pilot.
I think UCSD did that.
I remember with UCSD or something.
It might even have been in San Diego, but somebody did it.
And you never heard really, how well did that work for you?
Well, it was probably a really fun project.
It was probably fun, but HPC, supercomputing,
has a long history of using junk that's kind of just out there in the market
for a completely different reason,
because they don't have enough clout to get things,
custom design for the real requirements
who have high performance computing
so they're always using stuff made for video games
or something else
and then trying to apply it
and now we have
we're a wash in this really low precision floating point
like FP4
and yes once you hand out
a whole lot of accelerators
that have many terra ops per second of FP4
you say well what can we do with those
and it's not like you would ever
have redesigned a number system that way.
But that's exactly what's happening is that people say,
well, it's there and it's fast, and let's figure out what we can do with it.
I think that's where Ozaki comes from.
It's from the same exact position.
I've looked at FP4, and it's amusing.
The dynamic range goes from one half to six.
So you're betting that a block of numbers doesn't have a dynamic range bigger than 12
for the largest number to the smallest number.
You scale to the largest number,
and I did some tests on AI,
and about 10% of the numbers underflow to zero.
So you're throwing away 10% of the information
when you scale to the largest number at Lach of FP4.
FP6 is a little better.
It gives you a range of a factor of 30,
and that's what AMD is looking at.
But all of these things,
they require you to manage the scale factor
on a block by block basis.
And if you look at, what is the set of numbers you can represent with that system?
It looks nothing like the set of numbers that you need.
It's generally way too large, far too large.
What you want is something that looks like a Gaussian distribution,
a lumped around magnitude zero, that is two to the zero.
Numbers have magnitude around one,
and then they fade out and such that you almost never go outside the range of 10 to the minus 14th
to 10 to the 14th, something like that.
And if you can just do a really good job at representing numbers
with high accuracy in your 1 and less accuracy for the largest and smallest magnitude
numbers, then every bit counts because you're really matching the distribution of
numbers you can record with the distribution numbers that you need,
which almost nobody else does.
So how many bits do you end up with when you do that?
What is, I mean, there's still occasions where you need a lot of bits, right?
the latest thinking, I think, and as I say, there's a working group for posits.
We have something called a bounded posit that never goes outside the range of 10 to the minus 57 to 10 to the 57.
And that turns out to suffice to represent all the physical constants, the biggest one being the mass of the universe,
and the smallest being the cosmological constant, which is like 10 to the minus 50th, something.
So, and if you can represent those accurately, and you can do any cosmological calculation, like NBod
problems involving simulation of galaxies, then you've got everything covered. You don't not need
anything larger than that. If you do change your algorithm, you know, think about it a little bit
because you're probably using the wrong units. Use light years, not millimeters, please.
Right. And I have yet to find something that really demands numbers outside that range. And once
you get that number range fixed. Now, adding accuracy only changes the fraction. You have a way
of representing the exponent. And if you're switching from one precision to another, it's trivial.
Right now, if you try to change from single to double precision, you have to decode the number
into its exponent and fraction. You have to account for the bias. You have to account for exceptions.
It's really, really expensive. If you compare that with like changing from a 16-bit integer to a 32-bit integer,
there's nothing to do. You just put a bunch of zeros in front of it, a 16-0s, and you've got a 32-bit integer.
Oh, that's interesting.
But now with posits, they convert from one precision to another trivially.
You just add bits at the end and it automatically increases both the dynamic range and the accuracy at the same time, which is kind of amazing.
John, is it appropriate to get into quantum at all?
Are you involved in?
The only overlap I have with quantum is that people have said, we managed to get seven cubits in a row to act a little bit like a number, and they want to represent real numbers.
they use positive arithmetic when they do that.
Eric DeBenedictus, I think, is my contact in the quantum computing realm.
And he agrees you don't want to use some kind of a very wasteful format when you're trying to represent real numbers because they're so precious.
Every single cubit is so precious these days.
So to the extent that.
Oh, interesting.
Yeah.
So when they do try to do real number calculations as opposed to integer calculations with quantum computing, they almost certainly start with positive arithmetic because they are,
each bit from left to right has the maximum possible meaning.
It's like playing 20 questions for what number am I thinking of.
Each bit tells you a little bit more about what number range you're in.
So every qubit already counts.
Yes.
They're very expensive and very unstable.
And there's incentive for them to be efficient.
Let's talk a little about memory because obviously if I have fewer bits,
then I have fewer bits to store, fewer bits to move around.
I get sort of a implied improvement in memory bandwidth, which is the name of the game.
Absolutely.
There is massive memory shortage in the market, thanks to AI's voracious appetite.
Yes.
What's your take on that situation?
As I say, the memory wall, as they call it, just gets worse every year.
I remember Dongara recently pointing out that a dot product, the actual multiply ads,
goes 2,800 times faster on an Intel chip than getting the data out of memory.
that is just so absurd.
And people say, well, you can use cash.
So I recently did a 2026 update to a table I've been using for a long time of what is the cost of getting something out of a register, out of level one cache, out of level two, level three, out of high bandwidth memory.
What happens if you have to actually get it out of main DRAM on a server?
What happens if you have to go to a different server to get the memory?
and it all comes out to being 100 times more expensive at least,
maybe more like a thousand times as expensive as actually doing the arithmetic itself.
So if you really want to make an improvement to an algorithm,
find a way to use half the precision you're using now
or do something to reduce the precision,
and you'll get a massive improvement almost for free
just by moving a fewer bits around.
You'll increase your storage.
And if something didn't quite fit into, let's say, level two cache, maybe it does once you cut the size of the operands in half.
Now you could keep things much closer to the processor and then it runs even faster.
But if you run in half the time and if you have the precision, you're going to probably double the speed of an ad and maybe quadruple the speed of a multiply.
But now if you're using power to do that, you're using half the power and you're doing it for half the time,
that's one-fourth of the energy.
So if what you're paying for is kilowatt hours,
going to half precision is a factor of four, at least.
And it's actually a little bit more than that
in terms of the energy savings.
We had a conversation with Thorsten Heffler a couple of years ago.
And if I recall correctly,
he was telling the story of working with the weather forecasting,
Bura,
and how it took a long time for them
to just simply try running it in 32 bits.
Yes.
And they got all the same answers.
And I was like, again, back to my error propagation analysis,
I was surprised that they had not obviously done that analysis.
Or if they'd done it, it was like so far back.
So I believe there are probably quite a few apps in the HPC world
that maybe don't even need 64 bits.
Or at least they think they do.
They think they do.
And if you look at the discretization error of the grid that they're using,
there's way more error coming from that than there is from the floating point.
So it swamps out the rounding error.
There's a group at Oxford.
I'm trying to remember the name of Milan.
What's this last name?
It's got an umlaude in it.
It won't be hard to find.
He's used 16-bit posits to do weather forecasting.
And he's getting the same answers as the 64-bit W-RF models.
And so I point people to that.
That was one of the first applications of posits.
that was a big success.
I mean, I understand that you kind of need to run it for a while to gain confidence that,
okay, it does work and it takes time and all that.
And I can see that if you're doing something like nuclear or heart surgery or something,
maybe you want to be super careful.
And whether forecasting is very much kind of mission critical like that.
But it seems like it is really worth doing.
If you're doing something like nuclear, that's where I would reach for interval arithmetic.
If there's one place where you need absolute confidence,
bound your answer rigorously and be able to trust it.
But I suppose you could argue that let's do 64 bits so that at least that's not contributing.
But if you use a 64-bit posit, it's about 100 times more accurate than using a float.
That's 64-bit.
And Lawrence Livermore tried it out on shock hydrodynamics and came to that conclusion.
They just make better use of the existing bits with the tapered accuracy.
Well, then it's just not a question of the number of bits is how accurately you're representing the numbers on them.
If you're willing to spend 64 bits on a calculation, you're better off not using floats, but instead using the more modern formats, for sure.
Right. Now, you were saying in one of our exchanges that there's been some advances that allow posits to be implemented more efficiently, less energy, all that?
Yes, about a year and a half ago, by bounding the exponent size, the decoding becomes as cheap as a float, or maybe even a little cheaper.
It is definitely cheaper for 64 bit, and it's about the same for 32.
And for these really low precisions, it's kind of hard to measure.
But if you look at just the cost of pulling apart a number into its scale factor and significant end, and then putting it back together, that's the thing to compare, because otherwise the algorithms are very similar for posits and for floats.
as to how you actually do arithmetic with them.
But they are now on par or actually smaller.
And I think I can take up less silicon now with a posit arithmetic unit than you would
with a floating point unit if it's an honest floating point unit that actually processes
all the exception values.
I think you and I were at Sun about 2005 when they decided to take out all the exceptions
in the spark processor.
When they went to the Spark 3, they said, let's just assume every floating point number
is a normal floating point number, not a subnormal, not an infinity, not a nan.
And if an exception comes along, we'll throw that to software. We'll process it in the operating
system. And it will... Oh, on the assumption that that's not going to happen often.
Exactly. And it ran like a million times slower, of course, if it hit an exception.
And the financial community just went nuts. They screen bloody murder because they rely on
not a number to fill in fields in financial trades where they don't have the information.
So they never make a trade if they have incomplete information about a stock.
They have to fill in every single bit field.
And so they were using nan as a fast and convenient way to stop things from happening.
Oh, interesting.
And suddenly with Spark 3, Ultra Spark 3, everything ran a million times slower.
But soon all of the vendors, Intel, AMD, even ARM, were throwing the exception cases of I-Triple
arithmetic off to a handler of a software.
some kind, like they'd do it in microcode, let's say, and it would run for 200 cycles instead of
being able to do a multiply add in 8 clock cycles pipeline. Instead, it got thrown to a, to a handler.
And that meant that you could actually tell what people were computing by how long it was taking,
which gives you a side channel attack. And some people use that to crack open into websites and
hack them. So it's a security risk as well as being just generally a bad idea, at least
deposits, there's only one exception, which is not a real. And that is easy to handle as fast as any other
computation. So there's no reason not to have a fully compliant hardware support chip for positive arithmetic.
Excellent. Just kind of side conversation. I was talking with Doug Garnett, who you might remember
from all times too. I don't know whether he was at FPS when you were there. But we were talking about
how a company changes their product roadmap, thinking that some features are not used
anymore.
Then they lose like 20% of their customers because the customer, they were using the product
in a way that they didn't expect it to be used.
Yes, exactly.
Yeah.
Had they asked the people in the market development at Sun, we would have told them, yes,
there are people who rely on not a number.
Yeah.
They didn't bother it.
They just did it from their armchairs and said, oh, this looks like a good idea.
Amazing.
Yeah.
Amazing.
Let's talk if you have time about future system architectures.
What do you see as what the HPC community should focus on?
There are folks who think, let's just go to cloud and call it good.
Those who think it's a national treasure to be able to build these systems
and not outsource your expertise to those who may really not care about HPC
at some point in the future.
And we're sort of learning the challenges of trusting the supply chain
only to realize when it isn't delivering.
Yes.
And then what do you do?
Do you just keep going with CPUs, GPUs, QPUs, and call it good?
Or, yeah, what's your perspective on it?
Well, it's not very often I see an architecture design that just really lights the light bulb and says,
that's it.
That's a brilliant way to do things.
I certainly had that reaction back in the array processor days when I saw the wide
instruction word approach of FPS.
That puts all the burden on the software.
and then the hardware can run incredibly fast
because it doesn't have to have as much control logic.
All the control is built into the instruction itself.
And I had that same reaction a few years ago
when I saw the design of Next Silicon,
which is an Israeli company that notice that you could put
like a half a million integer processors on a chip
and embed data flow on it.
So it's like laying out a factory floor
where you connect them to do whatever flowchart
you have in mind, and then you get a result every clock cycle out of every one of the little
configured integer connections that you put together. It's a lot like an FPGA, except instead
of just doing logic gates, you're now doing 16-bit integer calculations, all the basics. And,
and oh my God, that was just like, that is the way to go forward, is you create a factory floor
and you don't change instructions. Each, you, you know,
your thing does the exact same thing over and over again triggered by its inputs and producing an
output and finally we can do data flow the way it was always hoped we could do it and uh if you're
not trying to always transport data to memory but just transport it to the next part of the chip
the uh the speed is just ungodly fast uh the way you can run things in parallel and
pipelined and it just it just runs and nick silicon
so that you don't have to think too hard about how to lay out the factory floor.
But I think there may be one or two others that are coming to the same conclusion
that what we need is to be able to put a data flow on a chip
and not use conventional of Monimon motion to memory and back.
That's going to help a lot with the memory bottleneck.
Right.
So, you know, a challenge of FPGA is how long it takes to program them
and the whole place and route and to do the layout of the interconnect essentially.
Yes.
And I understand that NexSylican makes that a lot faster so you can reroute things on the fly
and optimize it as you go.
There was also something.
They also, I think an FPGA lets you hook just about anything to anything else like a breadboard.
And it's much more restrictive on a Nex Silicon.
You can only throw things a certain distance horizontally to a different processor.
And that makes it really fast to do the interconnect.
And you don't pay the price you pay on an FPGA for a full capability interconnect.
Right.
have the coarse-grained reconfigurable arrays that I think Fugaku either has or talked about,
I remember reading it in that context. So it seems like there's a spectrum of just how much
you can place en route and how can you restrict it in a way that still gives you the benefits
but makes it a lot faster. That seems to be what's going on, right? Yes. Well, if the building blocks
become something that's a little easier to deal with, like integer operations, then we won't be
talking about whose floating point format is best because it won't matter. You could have a different
format for every calculation. Right. And just adjust it to exactly what you need. And I think that's the
direction things are going to go with respect to number formats and computing and relief of the
memory wall. More dynamic anyway. Yeah. That makes sense. John, that's been a theme of a lot of what
you're saying to us today, really, isn't it? Yes. The dynamic. But as I say, always conscious of
putting the burden on the computer and not the programmer, wherever possible, or the compiler,
at least. Let the compiler make the decisions. Yeah, right on. Beautiful. Now, John, you mentioned
your recent activity with VQ research. Yes. Would you like to tell us more about?
Well, just very briefly, it uses additive manufacturing, 3D printing to build passive components.
And for example, you can do an integrated component that has multiple capacitors, inductors,
resistors in it, something that has to be fired in a kiln that you can't possibly do on a piece of
silicon because the highest performance passive components are all things that really you make in
high temperatures, ferroelectrics included. I was a double major of applied math and applied physics
at Caltech and I almost never get to use my applied physics, but I've managed to build up a
patent portfolio with this VQ research and we're having some fun with it. But totally far afield from
high-performance computing.
For the moment, anyway.
At least for the moment.
It's handy to know how to solve Laplace's equation.
Let me tell you.
You need electrostatics in order to design capacitors.
But if you have full flexibility and make them any shape you want,
then you can build a much higher performance capacitor with the same materials and the same
weight and volume and have a higher voltage breakdown.
And it's a pretty fundamental component.
Oh, yeah.
If you look at any processor chip, open up an iPhone, you'll find it surrounded.
by ceramic capacitors.
And those are installed frequently as surface mount devices robotically.
And they're like really tiny, like one by two millimeters.
And they're starting to integrate those and use the ideas of VQ research.
So actually integrate the parts into a single component.
It's really looking like 1960s technology mixed with 2020 technology.
If you look at the silicon density and the transistor density,
and then you see all these discrete parts on a printed circuit board surrounding that process.
and it just, it screams for a more integrated solution to me.
Excellent.
You mentioned Laplace transform and for some reason that makes me think of FFTs and the
benchmarking world.
And I remember that you were in favor of having FFTs as a benchmark.
Oh, yes.
That would sort of set a lower bound and it's a really difficult thing to do and it's a very
fundamental algorithm.
Right.
And we now have conjugate gradient, HBCG, that does.
serve as a bit of a lower bound?
Well, I think FFT is one of the like six benchmarks that you can compute, but it is so totally
bandwidth bound that it encourages companies to do everything in their power to improve the
speed of memory references.
And the problem with the LIMPAC benchmark is that it puts all the emphasis on matrix,
matrix multiply, which is the easiest thing you can possibly do with the processor.
And it makes them do silly.
things like put 16 parallel multiplier adders on an Intel chip, even though you would only use it
for the Limpact benchmark, that kind of stuff, when you should be putting the money more into
figuring out a better memory delivery system. You're absolutely right. One of my favorite
sayings is that anything that's measured is gamed. Yes, exactly. You get what you measure.
And benchmarks are very much like that. So you want to build it such that the gaming actually
benefits something. Yes. So I talked to the people behind the NASS parallel benchmarks that
suggested that they take the FFT one and make it perfectly scalable, at least by factors of two.
And they already have different sizes of it, sort of like the early days of Limpact.
There's a small, medium, large, and even larger and even colossal. But if they just make it a scalable
benchmark, then FFT would be, to me, that would be the best possible benchmark for high-performance
computing because you make FFTs faster, you're going to make everything faster.
Right, right.
John, I understand you're a perennial.
Yes.
At the SC conference, meaning you've been to every one of them since 1988.
Exactly.
It's really been a measure of the revolutionary changes going out around with HPC as it
relates to AI to see how that show has changed over the last three, four, five years.
I guess since the lifting of the pandemic, actually.
Well, 1988 was quite a.
year because Seymour Cray gave the keynote at the first one. And it was because Cray finally
consented to actually participating in a trade show that it was possible to put on supercomputing
in the first place. Because without Cray, you don't really have a supercomputing.
Yeah.
Especially back in 1988. But that was also the same year that I wrote the paper,
re-evaluating Matt Amtall's law. And suddenly parallel processing didn't look so stupid anymore.
It didn't look like an academic exercise, but something you could actually.
actually work. So the watershed in going from a vector mainframe to a massively parallel collection
of commodity processors, like we were pitching it at 40-point systems, that happened the very first
year of supercomputing. And I just got to sit back and watch the revolution happen. It was George
Michaels that really put together supercomputing and figured out how to get the critical participants
that would make it really super interesting. And it all fit in one hotel. It was a Hyatt Regency
in Orlando, Florida, they could actually use the facilities within the hotel to have the trade show.
Well, they're expecting over 20,000 attendees this time. So it's come a long way. It's getting the size of
Sigraph. Yeah. Well, I'm glad that everyone has embraced the idea that AI is part of HPC. It has
exactly the same kind of aesthetic workloads, matrix matrix multiply. It's just lower precision.
Otherwise, I would like just to say for the record that I have consistently been seen.
saying AI is a subset of HPC.
Yes.
My kind of soundbite is that AI is an HPC app and quantum computing is an HPC subwoutine call.
Yes.
But yeah, I am delighted that that perspective seems to be gaining.
Yes.
And I just happened to luck out in that I introduced a new number type just at the time when
AI was really taking off and heading for lower and lower precision.
So they realized it's time to part with the IEEE standard and do something that's much more
appropriate. You know, there is definitely economic reasons to adopt it now, not just, you know, accuracy.
So who is all adopting it? What is the uptake for POSIT? Well, of course, I mentioned Facebook.
Huawei put it into their 5G network. Bosch put it into a vision processor. And I think there
are some people using it in AI, but they were keeping it a trade secret that they were doing so.
But I saw a picture of their booth display and it had the Posit accuracy plot right behind the
the CTO. So I thought, oh, ho. So I think there's more going on than I know about. For the original
paper now has over 600 citations. And of course, although those papers also have citations,
Caligo Technology in Bangalore has a built an accelerator that has 32 and 64 bit posit support
and full software stack. Wow. And even converted applications like Romax, for instance,
ready to run. Oh, nice. You think, yeah. So I think they want like 10 grand for the
accelerator and all the software or something like that.
We've got one in Arizona State, and I know there's one at Iowa State, and I think MIT has one.
So they're working on their next generation, but it's a risk five multi-core.
It's got eight cores, and the next one will have 64.
And so people are starting to build processors.
But this is the way it always works.
The small companies have nothing to lose.
The big companies, like Intel and AMD, they're not going to change anything for the next four years.
Yes, they plan out every single transistor that they're going to build.
it takes an awful lot of competition and losing a few benchmark wars before they change what they're doing.
So it's the same story as every other technological revolution.
It's a lot like the one trying to go to parallel processing.
And I watched that one carefully and, you know, I had a lot to do with it.
But this one is different.
This one, people are just really glad to be rid of floating point arithmetic and they're anxious to find something better.
Whereas with parallel processing, they said, oh, do we have to recode everything?
this is going to be very painful. And they were very reluctant. Even now, universities do not
teach parallel computing until at least the third year of a computer science degree. Yeah. And in fact,
they should really teach it to the science curriculum, too, not just computer science. Yeah.
A funny thing has happened with that original paper on Amdahl's Law. Right. It's getting cited
like every two days. I think it's because of all the multi-core out there, people now realize
there is no path forward except to use the multi-core, multiple cores on a CPU, the processor chips.
Oh, interesting.
And so a new generation is rediscovering how to find scalable parallelism.
The run out of things that do not need it.
Exactly.
Everything now needs it.
And so it's amazing to me that...
It's a juggernaut, John.
It's an absolute juggernaut.
Yeah, I mean, for a 1988 paper to still be getting cited that much, it really blows me away,
It's got like 2,500 citations.
Nice.
I think it's a very profound insight in that paper.
I thought it was common sense.
I didn't want to publish it.
Somebody twisted your arm.
Yeah, the director at Sandia said, no, you really need to publish this, John.
So, okay.
So I wrote up just a quick two-pageer.
The other thing I think that made it work is when you have an idea,
you have to come up with a theory of why it will work,
but then you also have to have experimental results
that prove that in practice it actually does work.
Right.
You put those two out at the same time,
and then you can't deny that there's something there.
And that's what I did in that paper,
because we had three examples of thousand-fold speed-up
that were real problems at Sandia,
and that made it not be an academic exercise anymore,
but something that really applaud.
I remember that. I remember that.
Back to Posit, there are a lot of cores,
Like, for example,
Nvidia is set to be one of the largest shippers of Risk 5 cores,
but they're all kind of buried here and there.
So do you see that posit might be used in that capacity
that allows even the big guys to use it without having to say so?
Yeah, actually, for years we've been using Nvidia GPUs
to do very fast positive operations, custom positive operation,
because remember, these GPUs have logic on them.
It's not just floating point.
And so if you have shift, integer add, integer multiply,
then it's not hard to build a parallel engine into a GPU
that does very fast positive arithmetic.
Like for AI, there's an environment called Pi Torch.
Yeah.
And Hamesi De Silva at National University of Singapore built Q-Torch
that uses the Nvidia processors to do very fast positive operations for AI.
makes it pretty fast to convert any neural network you you want to throw at it.
And then you can play with different variations on the posit format to get really,
really good results.
But as I say, it just blows away FP8, things like that.
If you're going to go down to 8-bit arithmetic, just take a look at some of the papers that have been written.
And there's dozens of papers showing that posits do a better job at 8-bit AI and float suit.
Excellent.
Very well.
Doug, anything else we want to cover?
Thank you, John, for your time.
We went way, way over.
Yeah.
No, I'm good, Jehien.
But, John, real pleasure to be with you again.
And many thanks for coming back.
Yes, pleasure to speak with you.
Thanks for the opportunity.
So I looked it up in our previous episode
where we had the conversation was episode 37.
So I highly encourage our listeners to go look that up as well.
Then you've got, like, the full story.
It's been a delight as always, John.
Thank you so much.
Thank you, John.
Great talking to you guys.
All right.
Take care, everybody.
That's it for this episode of the At-HPC podcast.
If you like the show, please rate and review it on Apple Podcasts or wherever you listen.
Every episode is posted on Orionx.net.
Contact us with any questions or proposed topics of discussion.
The At-HPC podcast is a production of OrionX.
Thank you for listening.
Thank you.
