Algorithms + Data Structures = Programs - Episode 299: Do We Need Humans in the Loop?
Episode Date: August 14, 2026In this episode, Conor and Bryce discuss the role of humans in the loop with AI and code review in the age of AI.Link to Episode 299 on WebsiteDiscuss this episode, leave a comment, or ask a question ...(on GitHub)SocialsADSP: The Podcast: TwitterConor Hoekstra: LinkTree / BioBryce Adelstein Lelbach: Twitter | BlueSkyShow NotesDate Recorded: 2026-08-12Date Released: 2026-08-14Red-tailed HawkBroad-winged HawkIntro Song InfoMiss You by Sarah Jansen https://soundcloud.com/sarahjansenmusicCreative Commons — Attribution 3.0 Unported — CC BY 3.0Free Download / Stream: http://bit.ly/l-miss-youMusic promoted by Audio Library https://youtu.be/iYYxnasvfx8
Transcript
Discussion (0)
So we have to turn our sense of design and style.
And we have to quantify what good code is.
And I don't think this is going to necessarily be an easy problem.
But I do think it's doable.
I'll give you an example.
Welcome to ADSP, the podcast, episode 299 recorded on August 12, 2026.
My name is Connor.
And today with my co-host, Bryce, we chat about code review in the age of AI.
and whether humans are still needed in the loop.
Connor.
What's up, buddy?
Did you misspell your own name?
Well, first of all, how's my audio quality?
Your audio quality is very good.
That's surprising.
Wait, I got to get my headphones.
Okay.
By good, I mean relative to my computer.
Okay.
Okay.
All right.
Now, I found one pair of headphones.
I don't know where my other headphones are
and my computer is, this thing's really on its last leg.
It's, you know, it's 10 years old.
I had been hoping to upgrade it three years ago
when the new Nvidia laptop chip came out,
but I've been waiting and waiting and waiting.
Hopefully soon, I'll be able to get the RTX Spark laptop.
I'm very excited about it.
But, yeah, this thing barely works now.
But okay, so you sent me an invitation that's Connor C-O-N-N-O-R.
I did.
Yes, that is not how you spell your name.
At least...
You asked earlier, did I spell my name wrong?
And I did not.
So question, Bryce, what are the top three reasons?
The name would be spelled wrong in my invite.
Well, I suppose it's possible that this entire time I have been spelling your name wrong.
And you have just never corrected me.
No, it's definitely not it.
My name is one end.
It's the Irish spelling because I'm half Irish.
So guess again?
Because you can't type well today?
Because your hands are busy?
I can't type at all right now.
You're very hot.
Indeed.
Because you're running or you're doing some other activity.
I'll take myself off camera for a second.
Are you driving?
I am driving, folks.
It's an ADSP first.
Road trip.
Why are you driving?
Is this unexpected driving?
Yeah, I was supposed to be back by now, but I've been running late all day long,
and I was going to message you to say I'm going to be an hour late,
so maybe we can just chat, and then we can record when I get home.
but according to you, my audio is good.
Will the recording be good?
I don't know.
I mean, now that I have my headphones in, it is less good, but...
I say what we do.
I mean, and man, oh, man is the traffic bad.
So it's going to take me twice as long to get home.
I may or may not be picking up Shima on the way home,
but I'm supposed to be home at five.
The meeting recording is supposed to end at 530.
So we will definitely get one solid 30 minute episode once I get home.
So maybe we should treat that.
And I have a really good topic, which I'm hesitant to start talking about now because I don't know how bad the audio quality is.
And I want it to be good audio.
So, I mean, we can just chat for the next 40 minutes.
And maybe we can keep this as bonus.
What do you call it?
like recording.
It's not footage,
but,
you know,
if it's good enough to use,
we'll release this not as a bonus episode.
If it's not good enough,
maybe when I've got the kid
and I can't record,
this will be a backup.
I was about,
I was about to say like,
you know,
this is good prep for the listener
because,
my first thought was like,
was like,
oh,
the baby has come early
and Connor is like on his way
to the hospital
or waiting at the hospital
or something.
No, no, no.
I mean, the baby could come any day now.
I think Shima is 36 weeks or whatever full term is.
So now we're at the point where...
Wait, why do you have to pick her up?
Isn't she taking time off work before she's due?
No, no.
She's working right up until...
My mom did the same thing.
Yeah, yeah.
It's...
Although, no, yeah, she is working up.
I think right up until, and apparently that actually I learned about something, there is both
maternity leave and pregnancy leave, apparently in Canada for certain roles.
And you have to take, at least for her job, as a public health doctor, you have to take
the pregnancy leave before you take the maternity leave.
But you can take the pregnancy leave as soon as the baby's born.
So it's a bit of a misnomer.
It's like optionally pregnancy leave.
Anyways, the baby is still cooking in there.
the oven. She was healthy. Remind me,
remind me, did you have a name yet?
We have two names,
my favorite and her favorite.
And her favorite is my second favorite.
Wait, it's a boy, right?
It's a boy, yeah.
And my favorite, at one point
was her second favorite, but she's kind of fallen out of
love with it. So I'm pretty sure.
It's not conclusive. Although, I mean, this episode
might come out two days from now, and
Schuma may or may not listen, but we haven't even discussed the finalization of this.
But all indicators point towards we unofficially have a name picked that we actually
haven't like finalized.
But yeah, I really love my name, but, you know, it's a team effort.
You know, all parties must contribute.
Oh, man.
So why were you late?
Why were you late? What were you doing today that made you late?
A lot of stuff.
I was in the morning watching Odyssey in IMAX 70mm, which I bought the tickets for a month ago because it's sold out everywhere.
The tickets in New York are like $500 or something.
You have dynamic pricing on IMAX theater tickets?
Or is it like second?
I think it's like the second market or something.
That's just conjecture.
That's just like something.
heard. It could be wrong.
I ended up going by myself
because there were no
like basically in Canada
Cineplex puts the
theater tickets on sale
a month out.
And for that entire month, basically
IMAX 70 millimeter
of which there are two in the greater
Toronto area are basically
also note except for the
like the 11 a.m.
and the
I think like 3 p.m. shows.
and even those, it's only like the bottom section
with like one-seaters.
And anyways,
we actually considered driving
to Rochester, New York, because they had
tickets available
for the 1045 showing.
Anyways, so that was in the morning,
and then that,
ooh, I ended up
just seeing, oh,
what is that?
It was a hawk.
It was a hawk.
I knew it.
I knew it was a bird, folks.
Yeah.
I called it.
It might have been a red tail.
But it might have been something else.
It might have been a broad-winged hawk.
I mean, you know, you got to keep your eyes on the road.
Or at least that's what you should be doing.
While recording a podcast and having a...
I didn't realize that Ramona was home.
And I just shouted that very loudly.
And she just walked into the room.
And she gave me a look.
And then she closed the door.
Yeah.
Oh, man.
I am looking for an Odyssey ticket.
I didn't want to see it.
won't say anything now because I don't want to spoil it because I imagine many of our listeners
have not seen it.
Although, I guess maybe, because the thing is, it's really the IMAX 70 millimeter that are
the ones that are like sold out because that's what the show or the movie was filmed for.
Anyways, no spoilers, no reviews.
Maybe in like, remind me around like November or Christmas time, at which point you have
no excuse if you haven't seen it by then.
And you'll have seen it by then and maybe we can do a
Odyssey review.
Anyways, then I had to run and then there was a meeting in the middle of that.
And so, yeah, it's just been a crazy day.
And here we are now, recording on the way back home.
What have you been up to?
How's life?
How's work?
How's AI?
There are tickets available.
Maybe, yeah, maybe I'm on.
I don't see it.
Oh, man, I have, I just got back from Greece.
Ramona and I were in Greece.
She was doing this acting residency, and that was very cool.
She got to, like, act in this play under the, like, right next to the Temple of Apollo and Corinth, like, at the Archaeological Society.
It was very cool.
And, like, they have to, the whole play is in Greek.
So these, this theater company, like, they have, like, three weeks to, like, learn how to say everything in Greek.
And, like, all the Greek people kept coming up to her.
and being like, oh, your accent's perfect.
Like, hot, like, I, you know,
it's like talking to her in Greek and whatnot.
And I'm just like, I could not.
I could not do that.
But we just got back.
I was, I was in Greece.
He was working from Greece.
It was very hot.
There was no AC.
I'm very glad to be back in America where there's air conditioning.
But Greece was really lovely.
And, man, I've just been, I've just been burning all the tokens.
Burning all the tokens.
The thing I've been trying to answer recently,
has been what programming languages are most token efficient.
And it's specifically for Cuda kernel writing.
It's a mixed bag.
It's a mixed bag.
And I got to get some more data and do some more analysis before I know for sure.
But yeah, I just been, I feel like I've never been busier.
Just so many projects, so much to do.
It's a little time to do it all.
you know, I've been thinking more and more about how we're going to review all this stuff.
You know, in the past we did this thing where somebody would propose change to code base,
and then a human would go and read the code, and we would all talk about the code,
and we would iterate on the code.
And it just increasingly seems to me like that process is not sustainable going forward.
And so I've started thinking about what is the alternative.
And I have come to the conclusion that all design, aesthetic, style, taste, and sense references that we have are only meaningful if they can be machine checkable.
So we have to turn our sense of design and style.
and we have to quantify what good code is.
And I don't think this is going to necessarily be an easy problem.
But I do think it's doable.
I'll give you an example.
Let's say that you get a PR that add some new feature to your code base,
but it also introduces this utility function.
And it turns out that there's another utility function
already in the existing code base that you should really be using instead.
At first glance, this seems like something that is hard to quantify, right?
Like the lesson here is like, don't reinvent yourself.
Don't duplicate things, right?
Like you don't want to have duplicated functionality within the code.
So like how could we possibly make this something that we have a hard check for?
Well, it turns out this one's actually kind of tractable.
Because what you can do is you could build a tool that looks at the AST of the PR, of all the changes introduced by the PR, and compares all of the entities introduced by the PR against everything existing in the code base.
And so, like, if you added some new function and that function was like 95% similar to an existing function in the code base, you could build a tool.
that could detect that, and that could then leave a review comment saying,
hey, did you consider using this function instead, this existing function?
So for things like abstractions, genericity, don't repeat yourself,
some of that stuff we could maybe build deterministic tools for.
And I've just been wondering, like, how much of style and sense do,
can we actually turn into something that's machine checkable?
Because I think anything that we can't make machine checkable
is not going to matter,
because the only way to review code changes in the future
is going to be to have testing in place that's good enough
that we trust that if the tests pass,
that the code is good to merge,
and that the tests have to not only test functionality,
but also the quality of the code.
What percentage of code do you think actually needs to be, like,
reviewable in the future?
Like, does code quality
from, like, a human point of view matter?
Like, what, or so, the real
question is, because I know,
like, I think that,
actually, I don't know if everyone can agree,
that some percentage
of code, like, doesn't need to be
reviewed. So, like, what,
what code needs to be reviewed?
I think,
like, yeah, like,
examples of code that doesn't need to be reviewed.
If it's, like, something for, like, some, like,
internal info project or some like script or something like that. Yeah, maybe it doesn't need to be
reviewed. But for something that's like going into like, you know, a software application or like a
library in particular or something like that, it's not that the code needs to be reviewed. It's never
been about the review of the code. It's about quality control. It's always been about just quality
control and testing. And like there's nothing sacred about the review process. The, the reason
that we do the reviews today is because we have not been able to automate the checking for quality.
And if you think about it, there was a time when we did not have CI in this industry,
where we did not have continuous testing, where if you, like, the only way to know whether the commit was good or not was like you would go and run manual testing.
And like we think of that now as the dark ages.
And I think in the future, we may think of the days of code review as being the Dark Ages 2.
But I do think that, like, there, again, it's not about the review.
It's about whether we can create systems that we are confident enough that the automated testing process is sufficient to ensure the quality of the product.
And today, I think there are very few large, important, critical production code bases for which the automated validation is sufficient to allow agents to merge PRs without human review.
And that is a problem that the industry needs to solve.
that we do not
every code base
has some amount
of unwritten,
intangible
guidelines and practices
and rules
that exist
largely for good reason
and which if not followed
might make the code
difficult to maintain
or introduce bugs in the future
or make the code
difficult to audit or
for difficult for humans to understand,
which I do think human understandability is important.
And I don't think that any code base today
is in a place where it has sufficiently automated
all of those rules and guidelines
to a point where you could just let agents,
if they pass this set of checks, you can let them commit.
I don't know.
I mean, maybe.
You're like picturing a very still,
I don't know, I feel like you're projecting,
projecting the flaws of like human programming onto the future of like AI programming.
And I have honestly over the last like month become more and more infuriated with the quality.
And I am like I am omitting many, many expletives that my wife has heard about like how bad software is.
Like, like, we seem to think that like we talk about like, oh, AI, blah, blah, blah.
It's not as good.
Like, look at the state of the software industry.
I have, I want to make this YouTube video that I just don't have time to make.
That's called the worst software, like, ever written.
And it is my Toyota phone app.
It never works.
It never works.
Every single time you open the app, whatever you try to do and like the information that it provides you is stale.
and like you have to open the app and you can't just close it right away you have to wait like you have to open it up close it
maybe do that a couple times and like wait a minute or two every single time i click charge now it tells me
failed to do that operation but it's always successful and uh and like so many times at least like
four or five times a month it turns off my manual charging and like it's scheduled to start at seven
PM and so like if that it get that gets turned off like the car doesn't get charged and then we it's like it's literally polluting the world this app is making me cause like use more gasoline like every once in a while the car just starts using gasoline like right now I'm on electric mode and the thing about the the Toyota Prius 2025 model that I have it's like as long as you have a battery you can keep it in EV mode I found though that if you go below minus 13 Celsius it turns off the ability to be electric irritate
but whatever, but then sometimes it just like turns on for no reason.
Anyways, like, we did this. We did this and like AI is so good. Codex GPT 5.6, like, if you ask it to go and diagnose all the bugs in some program, it'll, it goes and finds so much stuff and that's stuff that humans wrote. And anyway, so I just like, I just have started to notice like how bad all software is. There's not like a single piece of software.
or like technology that I use that just constantly works well
and that like I either don't notice because it works so well
or like I'm happy because I'm just like,
wow, this is a really great piece of software.
Like all of it's terrible to some extent.
And anyway, so when I hear you talking about like,
oh, you know, we gotta figure out how to do this and that.
I'm just like, I will give you an example.
I will give you an example.
Right now.
Good software?
Of good software?
Right now, right now the Kuda core compute library is
repo has 300 open PRs and 1.6K open issues. And I think if we merged all of the PRs that were currently
passing tests, I think it would break, it would be chaos. Like if you took all of the PRs that were
currently passing the test and you merged them all, I am sure that bad things would happen.
And that to me means that, like, we're not ready.
Like, we don't have sufficient, we have not sufficiently spec or validated the functionality that this software is supposed to provide.
And likewise, I think if you told an agent, go fix every one of these 1.6K issues and close every, like, close every issue by, like, merging commits to the, merging commits or close them if they're invalid.
or whatever, again, I think it would be chaos.
Would you disagree?
I disagree. I disagree.
I think it's not as easy as just sending 5.6 soul off to fix everything.
But I think that like if you, whether just individually or as a team, like guided it because
some of these things, I guarantee you are going to conflict.
And so like it's something that like a decision needs to be made.
If you gave the AI direction saying, like, listen, this is our overarching goals.
If you end up seeing two things that are in conflict with each other, like make a decision
as like a project manager.
Honestly, I think it might be, it might do just as well as humans.
But like, and I'm not saying that it would get it on one shot, but if you set this thing up
in a loop and like told it what the existing infrastructure is and said, hey, like add whatever
more infrastructure.
I honestly, I think it would succeed.
And I'm of the opinion that these models programmatically are better than like almost all humans in like every domain.
But do you think, do you think if we merged all 300 of these PRs, do you not think that it would introduce instabilities into the code base that over time would make it unmantainable?
Well, I mean, if you ask the AI to merge all 300 of these things, it wouldn't just do it all at once.
I like, you know, you have to give it, you know, a command that says, you know,
priority, prioritize these in terms of like some goal, minimizing the conflicts, blah, blah,
and like, you know, I honestly, I think I think it could happen.
You know, you need to give it the right parameters.
The real question is, do you think that the current tests are sufficient?
Look, it'll add tests if they're not sufficient.
I don't think that while I agree that AI is a great tool.
not think that it is at the point where we could let it formulate the, like,
I don't think it's at the point where we could let it wholesale manage the development of
software like CCCL.
Yeah, I don't know.
I think we would have to codify more, I think we would, I think humans would need to
codify more validation specs and requirements.
before that could happen.
Why does the human need to do it?
Because I do not think that current models have sufficient judgment to do so.
While they may be good at writing code, they don't always, like when I have a model
do a task, its first instinct about how to go about doing it is wrong frequently enough
that I would not want to entirely turn over all responsibility for a large production code to it.
What are you using? Are you using Opus 4.8, 5.8?
Any model, 5.6 soul or any frontier model?
I don't know.
Yours and my experience is, I mean, I'm, yeah.
But the topic that we're going to talk about next has to do with, like, my existential
like what do you call it?
Midlife, not midlife crisis,
but just the, I did a point of like...
I do think our experiences are very different here,
and I do wonder if it's that we're working on different things.
Like, I have been working on getting the agent to merge,
to merge the histocash algorithm into CCCL,
and like, GBT 5.6 sole,
its first PR attempt, it used the completely wrong approach for the tuning policies.
Like instead of adding a parameterized tuning policy that would parameterize, that would allow you to customize the tunings for each different GPU architecture,
which any human engineer working on the project would know that that's what you're supposed to do,
it just hard-coded in many places the constants for the architecture that it was tuned for.
And I had to tell it that, no, you need to follow the existing design patterns.
And I frequently, almost always the first shot of a design that it gives me as flaws that I need to give it feedback on.
Is your experience not been that you frequently have to give it feedback?
or is your experience that it always gets things right on the first try?
Or rather, has your experience been that you are able to set up a loop
and that with a very high likelihood on the first time that you do the loop,
that it gets things exactly the way that you want it?
Well, I'll dodge your question and answer it after I ask you a question,
why are humans able to do that correctly and not the agent?
because my guess is that it's, I don't know if tribal knowledge is politically incorrect,
but it's like it's written, it's knowledge that's not written down.
And so how's the agent's supposed to know about it?
So that's like that's another human flaw.
Like, why isn't that document?
Connor, but that's what I'm, that's why I'm saying.
That's why I said that we need to turn all of our design sense in guidelines and rules
into machine checkable metrics,
because that's even better than writing it down.
Sure, we could just document all of it in the codebase,
but if we instead figure out how to distill the design principles
into something that can be machine checked,
then it's just a part of CI.
Well, I mean, that's the thing is like...
It's something that we can check deterministically.
What you're saying is like machine-reaching.
to me is just like it's just a cop it like that should be documented anyways like having something
that's not written down that needs to be like human to human like verbally communicated that's like
that's a flaw with the process right that's a first step but i think even if you wrote everything down
i think agents would still get some things wrong and that's why you need to have machine checkable
tests for these things yeah i would i would i would say that like honestly
Everything that, like, make the difficult to the agents to, like, do things right are just, like, literally flaws of, like, the human process and that if we were better at, you know, writing down and documenting, et cetera, et cetera, that, like, the agents, the agents are would be totally able to, and I just think actually that is AI.
Hang on, hang on, hang on, hang on, hang on, hang on.
Are you claiming that, that agents do not make a mistake?
No, I'm not claiming that.
I'm claiming, though, that they're better than humans, basically, at every aspect.
If agents, if agents make mistakes, how are we going to check?
How are we going to catch those mistakes?
With, like, basically looping systems that monitor logs and, like, exception catches and stuff.
And so they're just, like, self-healing.
Like, anytime there's some problem, they are able to, like, self-correct because they can monitor that something's
gone wrong.
And so, like, you, you...
How am I going to, how will I know that a PR should be merged to CCCO?
How, sorry, ask that question again?
How should I decide whether a PR should be merged to the Kuta Core compute libraries?
I mean, well, in an ideal world, you have like a bunch of automatic CI and GitHub action checks that, like, when they're all green, it's all good.
But that is what I am saying.
What I am saying is that today, that is not sufficient.
People have to do this code review process where they read every line.
And the reason for that is because we have not sufficiently automated in CI,
all of the things, the design sensibilities, et cetera, that we care about.
And that is why people are still doing code review,
because the CI process is not sufficient.
Yeah, I see what you mean.
Yeah, and I guess my meta point, or is it a meta point,
just a regular point?
Is that like the AI would have set that stuff up from the get-go, you know?
No, because I think that some of those things,
some of those things require substantial research and innovation,
which, yes, we can do with AI,
but they require substantial research and innovation
that has not been done yet in our industry.
My underlying premise is that we must replace the manual code review.
We must take the humans out of the loop.
I mean, I'm totally with you there.
I just, I think we think we're at like different points on the curve.
Yeah, but I don't see, I do not think there are large important projects today that have the human entirely out of the loop.
Let me make it more tractable.
I do not think that there are large software projects that have been around for more than five years that are all.
automatically accepting, like, any AI PR without human review because they have so much confidence in their testing system.
Yeah, but that's because five years ago, like, it was pre-AI.
Like, that's like, you know, I promise you that there are companies out there that are, like, fully AI driven.
We're like AI from the ground up, you know, sure, are they massive software or, like, codebases?
Like, well, probably not because they only have started in the last, you know,
two or three years, but, um, but I think, I think many of them still have humans in the loop
at some point in the review process. I'm sure, I don't know if I agree with many. If you look at
I'm sure like, well, I don't know. We'd have to do some kind of survey, but like, can you, can you,
let's, you as homework, your homework is to find an example of an open source project that has
the human completely out of the loop that has been around for more like more than a,
more than two years at least I'll say.
It has to be something that's been around
for a while
because that's the only
way to understand
like the software life cycle
over time.
I mean, what does that prove though?
Like, you know? I don't know. I mean, what's
open clause contribution
policy? I think even OpenClaught,
they still have a human
that merges the PRs.
I mean, is Clicking?
a button really like a I mean one of my my array box thing like I click a button every
a couple days that says merge but like I don't do anything other than that open claw has
2.1k PR's open that tells me that their process is not entirely automated if the process
was entirely automated there wouldn't be there would be maybe 100 open 2.1k PR's open means that
there's like a backlog I mean I don't know I feel like this is like a
What are we, what's the word?
Like, fallacy or something is like, just because there's a human in the loop
deciding what they want added or not added, it's not really like an actual human in the loop,
like in terms of like coding and code review.
Like you can make decisions at like a project level of the direction and the features and stuff like that.
But like I'm talking about like the coding, the CI, the CD, all that stuff.
I will again, I will challenge you to find me an example.
of a large adopted open source project where they're accepting PRs without doing any code review today.
By a human?
Yeah, without any human code review today.
All right.
I guarantee you that there's a bunch.
So whether they're two years old, whether they have 50,000 GitHub stars, I don't know about that.
I will.
If you're challenging me just to find some open source GitHub projects that have like non-human.
They have to be important projects that are used by a large, like a large base of users.
I don't have a specific metric in mind.
I will waive the specific two-year requirement, but it needs to be something that's used in production.
All right.
My wife's about to get in the car.
so I don't know.
Well, she's not in the car yet.
She's like 10 seconds away,
and she may or may not want to be on the recording.
We'll find out in a sec.
I do have.
Oh, sweetie.
Would you like to be on the podcast?
She shakes her head and says no.
Be sure to check these show notes,
either in your podcast app
or at ADSP the podcast.com for links to anything we mentioned in today's episode,
as well as a link to a get-up discussion
where you can leave thoughts, comments, and questions.
Thanks for listening.
We hope you enjoyed and have a great day.
Low quality, high quantity.
That is the tagline of our podcast.
That's not the tagline.
Our tagline is chaos with sprinkles of information.
Hi, Shima.
How you doing?
Hi, everybody.
I'm pretty good.
I'm pregnant.
Yeah, I hear you're going to be due soon.
Yep, I'm 35 weeks.
So I'm almost any week now.
Congratulations.
Are you excited, scared?
How does it feel being an almost mother?
Oh, I'm really excited.
Yeah, I can't wait. I'm super excited.
The fetus moves a lot. We're very lucky.
You refer to our baby as a fetus?
You use our baby as a baby.
Okay, our baby.
Do not insult my future son by calling him a fetus.
Well, I think he has a fetus.
By definition?
Yeah.
