Algorithms + Data Structures = Programs - Episode 302: From PyTorch to GPU MODE with Mark Saroufim
Episode Date: September 4, 2026In this episode, Conor and Bryce chat with Mark Saroufim about his path from AI theory and Microsoft to Graphcore and PyTorch, what he learned from open source, building the GPU MODE community, and hi...s work automating AI systems research at Core Automation.Link to Episode 302 on WebsiteDiscuss this episode, leave a comment, or ask a question (on GitHub)SocialsADSP: The Podcast: TwitterConor Hoekstra: LinkTree / BioBryce Adelstein Lelbach: Twitter | BlueSkyAbout the Guest:Mark Saroufim is a cofounder at Core Automation, PyTorch maintainer and cofounder of GPU MODE.Show NotesDate Recorded: 2026-07-20Date Released: 2026-09-04Core AutomationGPU MODEGraphcoreThe Robot Overlord ManualMachine Learning: The Great StagnationPyTorchIntro Song InfoMiss You by Sarah Jansen https://soundcloud.com/sarahjansenmusicCreative Commons — Attribution 3.0 Unported — CC BY 3.0Free Download / Stream: http://bit.ly/l-miss-youMusic promoted by Audio Library https://youtu.be/iYYxnasvfx8
Transcript
Discussion (0)
You know, at the time, my thinking was like, okay, well, I really struggled to get people to care about my ideas for money.
You know, but like, what if I could at least get them to care about it for, like, for free?
And so that's kind of what got me into open source.
I was like, okay, it's a good prerequisite for am I on the right track.
So I got started to writing like an e-book called the Robot Overlord Manual, which is, it's an interesting e-book in hindsight.
It was like somewhere in between a robotics versus RL book versus math book.
It's a very strange book.
I don't think it's that good in hindsight, but I had a really good time.
writing it. And basically, like, a couple of those chapters made it to the front page of
hacker news. And I think they all sort of had this sort of similar essence of the math was a bit
different. But like the math helped motivate how to write the systems. One canonical example of
this is like how like lie algebras can help you like represent rotations and robotics. And so
those are really fun for me or using Lagrangian physics to sort of formulate physics as an
optimization problem that minimized energy. So a lot of these problems for me were like very,
very intellectually appealing.
Welcome to ADSP the podcast, episode 302 recorded on July 20th,
20206.
My name is Connor, and today with my co-host, Bryce, we interview Mark Serafim.
Mark is co-founder of both GPU mode and Core Auto.
And in today's episode, we chat about his path into computing.
It includes GraphCore, Pytorch, meta, and more.
I'm going to have to take up the B2.
hat because hats are not super compatible with my hair.
Well, we're here with Mark.
I guess Mark, maybe you should intro yourself.
Sure, yeah. Thanks for inviting me, Bryson. Thank you, Connor.
Yeah, my name is Mark Serufim.
I've been working in, like, AI systems for, like, about, like, six years now.
Most of it was spent, like, working on PyTorch, which is a pretty popular, like,
framework for deep learning.
But I guess, like, what I ended up being most known by was this, like, fairly large
ML systems community called GPU mode, where we do things ranging from, like,
GPU education to competitive GPU programming, and we have a whole variety of working groups.
Most recently, like, I'm no longer purely working in open source like I used to, but I
helped co-found a new Neo Lab called Core Automation.
Our goal there is basically to build the world's most automated lab and build the world's
best model.
And, like, my research agenda there is, like, getting AI systems to be really good at writing,
getting, like, AI to be really good at writing systems code.
And I think this will be, like, fairly transformational, like, really for everyone.
And, you know, so yeah, thank you guys.
And I'm excited to be here.
How did so what did you do before you were working on AI systems?
Like what did you do before Pi Torch?
It's a very good question.
So, yeah, I had like an interesting early career, like pretty bumpy.
But the way I got started in AI was like in AI theory.
So I was a grad student at UC San Diego.
I was fortunate to have like two amazing advisors with like Charles Elkin and Sanjord
the Scoopto.
And they're like primarily like I said, I was in AI theory.
Like I taught like seminars on quantum computing.
I was like trying to figure out if like, you know, we could rewrite like axioms of probability theory using like game theory instead.
And I had a really great time.
But basically, you know, at the time, like I didn't really sort of have like a good understanding of like how important AI would be.
It was more like I just thought it was really cool.
And I ended up like my first job out of school was at Microsoft.
So at the time, like I remember getting like the job offer about a year before I graduated.
And so I just like spent the year doing math.
But I joined as a product manager at Microsoft, and I worked on a whole bunch of wacky things.
Like, I, you know, my first project was selling Microsoft's display ads business to AOL, which was like an interesting first project.
Sock 3-E compliance as well for Outlook.
It's like very exciting stuff.
Wait, what year were you selling things to AOL?
So this was 20, like 2014, 2015, roughly.
This is when I first started my career.
But yeah, I mean, it was a weird project because I remember.
at the time being fairly AI-pilled and like thinking like, okay, I don't know how big of a deal
this will be, but I did get a sense of like this is a really important problem.
But at the time, I would say there was this widespread view in the industry that this was sort
of like a luxury on top, like as in like it's sort of a further down the tech tree because
what you needed to do is to get like big data, which was proprietary company data.
And once you accumulated enough of it, the hypothesis was that like you'd be able to build
very smart systems.
But funnily enough, this turned out to be incorrect.
because private data turned out to be less valuable than public data,
which was like very, I think a very strange conclusion.
But it was the reality.
And so like it was really only in my last year and a half at Microsoft that I worked as a research scientist.
So there I was like working mostly on like email intelligence.
So think things like, you know, the AI systems that would figure out that this is a flight.
And then given that it's a flight correctly parsed information out of it,
these systems have to be fairly conservative because if you got someone's flight information wrong,
would be really angry.
And so, but at the time, like, I mean, a lot of these were like fast rules-based systems,
like more into traditional AI.
I myself had worked on like the first like deep learning system to do this like contact
extraction and email.
But at the time it was considered too heavyweight because the model that I built was like
around like 50 megs.
And so that was considered at the time like unrealistic deploying fraud.
And primarily because like there was again this thesis at the time that like you would want
like a small model to put on the client directly, that it would.
wouldn't really make sense to have a central one because maybe people's preferences would be different.
But it was interesting to me, like basically we were sort of correct in formulating the problem,
but because we had certain ideological beliefs, like one about private data and two, about local models,
I would say that prevented us from building, like, really, like a state-of-the-art model that people could use.
So, yeah, it's a longer story, but that was the Microsoft days.
And let me know if you want me to keep going, if you want to interrupt me here.
Because there's like maybe three more years, basically, worth of stuff that I can do.
Yeah, yeah, I mean, keep going. Keep going.
I'll keep going. Okay.
So after Microsoft, I remember at the time feeling like it was around the time when the Dota
AI bots came out.
And so for me, this was a like, this was kind of like in the same way I see a lot of younger
folks feeling the same way about alums today.
Like that was sort of like my field AGI moment.
Primarily because like before I got started in working on software, like I got started
in software because of video games.
Like my sort of like my first experience using a computer to do work was pirating
Warcraft 3 and reselling it to my schoolmates. So that was kind of how, that was my introduction
to computers. I made a decent amount of money like doing this. But, but yeah, like, basically for,
for me, what always felt like the most magical was like, you see like an agent, like doing things
on a screen. And to me, that always seemed like a, like, the real way to like look at the, like,
it felt like tangible in a way that like, I don't know, grammar based models or like parsing
trees, like never, never really did. And so, so my thinking then was that like, this
work came out. And I was like, okay, well, this seems really transformational. Like, what if we could
do this for all of video games? And we could have, like, really sophisticated game AI. And of course,
like, I myself had spent a lot of time playing games like Starcraft and Dota. You know, I was a competitive
chess player in school. And so, like, I just sort of quit my job and I'm like, okay, I'll figure
this out. Which turned out to me not, not like a very wise career choice for like a variety of
reasons. One is, like, I'd no idea how to sell. I had no idea how to market my work. I was
even that good of an engineer back then. And so, like, it's like struggling on all sorts of, like,
skill issues. And most notably, and I think what ended up being like a really critical flaw of
this work was that like, it was like really expensive to actually get anything working. And so like
here I was like basically running on the equivalent of like a 1080TI. And what I would notice is like running
sort of simple experiments within unity, which is like, you know, I would sort of change whether an
agent is standing like this cord in a system or like increment the X coordinate by just a bit and
like nothing would converge. And I was like very confused. Like these techniques felt really brittle to me.
and it wasn't really clear how to get out of it.
And at the time, I remember, you know,
I was like mostly going off of like my savings from Microsoft days.
And then COVID hit.
So very quickly, like my sort of money that was like in stocks was like evaporating towards
zero.
And of course, I wasn't making income either.
So I really, really needed a job.
And then I got a couple of them.
A lot of them got rescinded because of COVID.
So it was like very like, very like a weird time.
And so like, you know, at the time my thinking was like, okay, well,
I really struggled to get people to care about my ideas.
for money, you know, but like, what if I could at least get them to care about it for, like,
for free? And so that's kind of what got me into open source. I was like, okay, it's a good
prerequisite for Am I on the right track. So I got started with writing like an e-book called
the Robot Overlord manual, which is, it's an interesting e-book in hindsight. It was like somewhere
in between a robotics versus RL book versus math book. It's a very strange book. I don't think it's
that good in hindsight, but I had a really good time writing it. And basically, like, a couple of
those chapters, like, made it to the front page of Hacker News, and I think they all sort of
had this sort of similar essence of the math was a bit different, but like the math helped
motivate how to write the systems. One canonical example of this is like how like lie algebras
can help you, like, represent rotations and robotics. And so those are really fun for me or
using Lagrange in physics to sort of formulate physics as an optimization problem that
minimized energy. So a lot of these problems for me were like very, very intellectually appealing.
And yeah, and I guess eventually I started applying and like the only company that got back
to me. It was this company called Graphcore, which was an old, I guess, now defunct competitor.
You were at Graph Corps?
For 11 months, yes. It was a fun story as well.
So GraphCore was great. I had like amazing colleagues, but basically the sort of just like walking into it was we know we would basically go to these customers and we'd ask them, okay, well, we need to be deploying.
Like, you know, do you want to buy our chip? And they'd be like, well, does it make such and such model faster?
And we'd run a whole bunch of experiments and then, you know, it turns out with like a whole bunch of work.
Yeah, often the answer was like, yes.
The problem was that like, there was like two layers of problems.
One was like, Nvidia wasn't a static target.
Like basically we would do all this work.
And then like a year later,
Nvidia would just like double all the numbers we'd need to hit that always like really sucked.
And then the second thing was like the world basically changed.
Like I think GraphCore's like thesis at the time was that.
So it was like a small Sram chip basically with like 300 megs and then later 900 megs.
The idea was like if you can get a model to fit, it's incredible.
But if it doesn't fit, you're sort of going distributed by default.
and unfortunately, at the time, like, the Burt based on large, language models were quite large.
And so it just, like, wasn't, like, a natural fit.
We had to introduce, like, distributed programming, like, much earlier than the community basically was ready to understand these ideas.
And so that was hard.
And, but, funnily enough, around that time, basically, was my first, I remember at the time writing, like, a, like, a very popular post called the Great Stagnation of Machine Learning.
I don't know if you might have read it.
That post kind of put me on the map at the time because it was, like, sort of this, like, the very,
very off-the-cuff rant about how, like, most academic AI is, like, basically morally bankrupt
because it sort of formulates itself as, like, this, like, innovative field that's, like,
risk-taking, but a lot of it were sort of simple ideas, like, on scaling.
And so, like, I think it was well written.
I had a couple of, like, very catchy phrases in it, like, graduate student descent.
And there I was, like, attacking the auto-research community, which is now defunct.
But I was also attacking the optimizer community, which is, turns out was defunct for a long time,
but later had the introductions of like Mouan and Champu.
Basically, they sort of salvage themselves.
And it was like pretty interesting because like right after I wrote this,
I remember sort of getting into trouble at work because like I was like also talking about like
Graphcore about this.
It was like it was like this very volatile post.
But, you know, I remember over like Christmas break basically, there was like a job opening
at Pytorch and I had started using Pitech myself at work.
And my framing there was that like, you know, but like I started with Piano basically.
that was sort of my first deep learning framework,
and I remember none of them felt like a joy to use
because it was just, I had an idea,
and then I would get bogged down with how to implement it.
But at the time with Pytarch, I remember,
I was reading the UNET paper.
The UNET paper is this architecture for vision encoders
and has this very interesting structure
where the encoders just get smaller and smaller,
and then they, up until a point,
then they get bigger and bigger.
And I looked at this and was like at this fairly large architecture diagram,
and I remember just reading the paper,
And then just, like, typing out the Pytarch code without any reference.
And then it just, like, worked.
And I'm like, oh, my God, like, this thing is insane.
Like, like, what is this alien technology?
And I felt like I had to work on Pytarch and, you know, I applied.
And I got very lucky that I passed interviews because I think a good chunk of my interviews,
people just want to talk to me about my blogs, which was, like, pretty funny.
And, yeah, and Pytar's was, like, one of the best places I've ever worked.
That's why I ended up staying there for, like, a whole five years.
And, yeah, I mean, I just think for me sort of, like, it taught me.
me a lot about like building things that people want. I felt like the engineering culture of the team
was quite strong, but like even which was maybe rare is like I also had the product culture was
quite strong. Like people really understood what is it that researchers wanted. And they had a clear
idea of like their their customer base. This would change over time, but at least like in the early
days I felt like this was like one of the best places in the world. Yeah, it's so fascinating to me
the arc of how Pi Torch
sort of won as the dominant, you know,
ML framework. Because, like,
I remember when, you know, TensorFlow was the,
the thing that we used for production, the big
serious thing, you know, it's in C++,
it's fast, it's got this, you know,
you know, this declarative model. And,
you know, I think at the time it seemed like
this kind of like crazy idea of like,
well, just write it all in Python and, you know,
be easy and simple. And people,
people would ever do that at scale, but like, that does, like, how did we get to the point where,
I mean, if I'm, if I understand correctly, like most of the big labs, or at least some of the
big labs, do train in Pi Torch at scale. Like, how did we, how did we get there?
Yeah, it's strange. And it's like something I've thought a lot about for what it's worth.
But, but basically, like, my sense is that like, and this is like something I felt like a lot of the
maybe hardware vendors with maybe the exception of Nvidia never truly internalized, which is
that like if you ask a customer like what they want, they'll often tell you I want performance
because it's like an easy way of like you put two bars on a slide and then you're like,
oh, I pick the bigger bar.
Therefore my decision is wise.
But then like often like you realize like, okay, well, people want performance, but like they
also want print statements and like they want like dynamic control flow and they want like
ragged shapes and they want sparsity and they want all these things.
And this is like very like sort of weird for a performance person to understand generally.
is. And it's obvious in hindsight, but at the time, I think people thought it was generally weird. And so, like, the way I think of it as sort of our job as systems people today. And again, I'm a sort of a newer systems person than you are, Bryce. But, like, basically, I sort of view my job as being at the service of AI researchers. Like, it's not to tell them your, the way you're doing things is incorrect. It's basically, well, they need to express their ideas and explore them. And it's like my job to figure out how to make this, like, class of people, like, productive. And turns out, like, it's not you write. You write.
of training framework in like raw kuda and it's and then you expect people to understand all the
all the details of the kuda manual and for what it's worth like i think the reason why like i would say
this is no longer true today to be clear but i think the reason why this was probably true for a
long time is because like memory bandwidth relative to flops like there wasn't sort of like this
big dichotomy and so you didn't really like fusion compilers were important but they weren't like
absolutely critical as they are today and so like pie torts very
basically hit a sweet spot where they basically figured out that like a certain tradeoff didn't
really matter for a long time. And then they figured out how to convert that into like a product
advantage and basically all the papers are written at PyTorch. And again, in a world pre-LLMs
where there was like a big variety and the kinds of ideas people were exploring and remember,
there was things like Nerfs and GANS and language models and image models and diffusion and
like a whole bunch of like ideas that have been lost the time today. Like these wouldn't be
possible if like every time a researcher needed to do any work, they needed to recruit like a
research engineer to do their job. And I think there's like a deeply unpleasant way of working,
which is like, you know, you go to the research engineer and then they're like, oh, grumble,
grumble, like, you shouldn't do this. And like, this is a bad idea. And don't you spread. And,
you know, that's not what you want to hear actually, yeah. I think that to some degree,
like, if, if the research had stopped, if we had figured out the answer in, you know, 2019 or
2020, something like TensorFlow would have won. But the way that this industry played out is, like,
the research continued. And that meant that the people doing the work were researchers and the
work being done was exploration. And I think to some degree continues to be today. And so it was
necessary to be able to move fast. Like, I don't know, maybe it maybe it's, maybe when people are
in a tech revolution, it always feels like this. But it feels like the industry moves faster now than it
ever has in the past. And maybe people who are around during the internet revolution will tell us
that it was like that back then. But it really does feel like, you know, like every week, every
month, you know, there's some big new development. I mean, obviously there's some things have
solidified, you know, it seems like transformers just take over most of the world. But there's still,
you know, the idea of the researcher remains. It's not the production engineer, but it's the researcher,
and we're going to explore a bunch of different models
and, you know, that the productivity continues to be key.
I mean, yeah, like, for what it's worth,
like I have heard people say things like,
oh, like, since LLMs, this is sort of like,
suck the oxygen out from a performance engineer's, like, perspective.
But actually, like, I think even the label of transformer
is, like, quite a broad one as far as I'm concerned
when people use it, like, basically, you know,
is a transformer with deltonet, like, can me,
like, is this a transformer or is this, like,
Arnan. Have you seen the alignment chart? Somebody tweeted the only...
That's exactly what I was thinking about. Yeah. And so when I look at this alignment chart,
what I'm thinking is like, okay, well, we found like an architecture that's really good.
And then it's only natural that like people want to find some that are even better because
like the amount of flops are putting into models are remarkable. Like yes, you can always get
more GPUs, but there are sort of like economic limits. Like you can't get $7 trillion.
Like a, you can't get a $7 trillion cluster easily. And there are like real bottlenecks.
the system beyond just money. And so, like, I continue to feel like there's, there's a big place
for, like, architecture research. And I think for architecture research, I mean both, like,
the kinds of models you pick or, like, variants of it. So think of things like a, you know,
a transformer with residual connections or, like, with the kind of like the, yeah, so, like, is this
a transformer? Like, yeah, kind of. Like, and so what are the tradeoffs there? And then, you know,
turns out, like, those do have real tradeoffs. Like, for example, you use, like, a lot of VRAM
for activations. You know, turns out that doesn't matter for training because you need to save
for your backwards past anyways,
but it really matters for inference
where now they use up a lot more memory.
So, like, I, like,
and I've realized this, like, especially since joining Core Auto,
that, like, there is sort of, like, a big spectrum of, like, ideas
that people want to explore.
And then every time the research team sort of comes to me with, like,
a specific idea or when we look at, like, papers and open source as to what people
are doing, they have very subtle tradeoffs for what makes them better,
like, worse or better for inference and on what dimensions.
But, yeah, I mean, I think ultimately the goal is capability.
And so, like, that's why, like,
You know, unfortunately, as systems guys, we have to accept that we're like second class
because we want to see the remarkable capabilities.
And then from that, we decide how can we best make this fast?
Because if it was up to us, how do we design architectures?
We would just like, for example, send the same like weight matrix to GPU.
It would be as big as possible.
And then we would just do like a matmole versus itself over and over again.
And there would be no non-linearities.
It would just be like one giant matmole that you do infinitely and we get a GI.
Like if researchers came to me with this, it's amazing.
Unfortunately, such a model would not converge or be useful.
So that's why I think inherently our opinions matter a bit less than we would like.
And when you say you're on the Pytorch team, was that at Meta formerly Facebook at the time?
Or was that like contributing to open source and you were working elsewhere?
Yeah.
So in my case, like I started as a user of Pytorch when I was at Graphcore and mostly spent
a bit of time on the forums and stuff.
I'd never sent a serious PR to PITR to PITR before joining the team.
So when I joined it, it was like for five years and I worked indeed.
I worked on the Pytur's team
at the time the company
was still called Facebook
and at the time as well
like it was sort of a like again
it's all funny for me to think through this
in hindsight but like it was pretty
peak basically because it was like
it was around the same time like
Pytorch was sort of like this like dominant programming
language where everyone loved it was my first time
working on a product that people adored
as opposed to a product that we're trying
to get people to adore which is like very different
psychologically and then Lama 2 was also
like a big revolution because it was like also
So like a llama 1 was a good model, but its license was very restrictive.
Lama 2 had a good license.
And so like now there was this community building on top of it.
Then you had things like local llama and like unsloth.
And local alums were actually like becoming good.
And Pytur's was sort of like the substrate behind that.
And so I think for us for like those like two, three years, like it felt incredible because
you really felt like you're like the good guys are winning in some sense.
Yeah.
And then the work was cutting edge.
And that was like, I think some of the most rewarding time actually to work in Pytritch
was like around this period.
Yeah, did you overlap with Adam them or had he moved on to other things?
Adam almost never worked at Madda as far as I know, because he only worked at, he only worked
as an intern, which is like a funny story.
So, like, I met Adam for the first time last year.
So like, because he lives in Poland or Germany, I forget, but like he doesn't fly to the
U.S. very often.
He's a bit of a hermit, but like one of the smartest people I've ever met.
So I kind of understand.
I kind of understand his perspective.
But yeah, Adam basically was like the main person who wrote.
most of the code for the first prototype of PyTorch.
But then later, I would say, like, the project was mostly scaled up by, like,
Sumit in terms of community and Adyang, in terms of, like, the code, like, basically CI and the
code and all the sort of, like, more hardcore features.
And then, of course, like, the project ended up, like, I would say there's maybe, like,
30 or 40 people that I think were very critical to the project's success.
But, like, you know, what I think of the founders, like, there's basically, yeah, there's,
like, there's, there's, there's, there's, there's, there's, there's, there's, there's, there's,
Sumit, there's Adam, and then there's Ed.
For me, these are the people who I would consider it to be the OGs.
And there were, of course, like, others, like, you might recognize, like, you know,
Natalia, who used to work at Nvidia, who did, like, the most of our, like,
Kuda performance work.
Horace and Jason Ansel, who did, like, the majority of our, like, newer compiler work
with Torch Compile.
So many, like, many great engineers, like, touched a project at various points in time.
But, yeah, I mean, like, the, the, Adam and Sumit did something more remarkable, which
is, like, create something.
from nothing. And I always have a lot more respects for that because I think scaling a project
in some sense is easier. It's hard. But I think creating something from scratch is not something
that, like, I don't know, it's not something I take for granted or that I necessarily know how to
reproduce really well. Right, right. Be sure to check these show notes, either in your podcast app or
at ADSP the podcast.com for links to anything we mentioned in today's episode, as well as a link to
a get-up discussion where you can leave thoughts, comments, and questions. Thanks for listening.
We hope you enjoyed and have a great day.
Low quality, high quantity.
That is the tagline of our podcast.
It's not the tagline.
Our tagline is chaos with sprinkles of information.
All right.
And before we get started, you might have to do a lot of the heavy lifting, Bryce.
My Wi-Fi is fine, but my Deco-Mesh node keeps on, like, turning off, like, failing every 15 minutes.
What is a Deco-Mesh node?
It's a very fancy Wi-Fi.
I have the same one.
Yeah, it's just like, it's not the best, but it's...
up there. And it's just like if you want to set up a mesh network, which means like, you know,
a lot of people, if you're in a condo, you can just have your single rotor that's fine. But if you
have like a bigger space and you need multiple nodes, you can either like pay your Wi-Fi provider
an extra like $5 per little Wi-Fi extender. But I was on principle never going to pay these
companies more money unless if it's for faster internet. And so you can just buy these like,
I don't even know what the correct technical term for it. Like, do you know, Mark?
at like if you live in a new york in a new york apartment you don't have any of these problems you just
they come in they put in one Wi-Fi router and it's always fast except for the last month they're pretty
amazing by the i looked into this tech a while back but they do things like they try to predict for
example like they can like measure which parts of your house like need the most bandwidth and they
will like just route more bandwidth towards it so each of these routers is itself like this like
interesting, like, statistical AI node.
It's pretty, like, remarkable.
I mean, I get, like, 800 megs per second on Wi-Fi speeds because of it.
And before it, I was getting 30.
Wow.
Okay.
Maybe I do need one of these.
It's great.
Like, I mean, if you play, like, especially, like, when I first set it up, I was just, like, download games and just, like, uninstall them for fun.
You know, it was just like, let's just, like, waste capacity.
Anyways, all right.
Now we'll keep that in the edit, maybe, but I'll throw it over to Bryce, you can do introductions here.
But anyways, if I am silent for minutes on end to the listener, now you know it's because my little node is acting up.
Over to Bryce.
