How I AI - This solo builder runs 24/7 local AI on his own hardware | Alex Finn
Episode Date: July 13, 2026Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across...: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He’s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it.What you’ll learn:How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good forWhy Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleetHow the build loop and review loop in Claude Code workHow to allocate tasks by machine and modelWhy unlimited local inference changes the use-case math in a way a $20 cloud subscription never canWhat OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in—Brought to you by:Runway—The creative AI platform for images, video, and moreJira Product Discovery—Prioritize with insights, build with confidence—In this episode, we cover:(00:00) Intro(02:58) Alex's hardware stack(03:48) What "ambient AI" means(04:15) Alex's red-pill moment with OpenClaw(07:04) Mac Studio vs. DGX Spark vs. RTX 5090(13:24) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes)(17:16) Fleet control dashboard: assigning 24/7 tasks across machines(20:42) Local models as security scanners feeding Claude Code(22:25) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task(24:28) OpenClaw vs. Hermes: the honest comparison(26:55) The software factory: build loop, review loop, rocket emoji(31:55) Lightning round: favorite hardware, favorite model, prompting style(34:46) Where to find Alex—Tools referenced:• Claude Code: https://claude.ai/code• OpenClaw: https://openclaw.ai/• Hermes: https://hermes-agent.nousresearch.com/• Tailscale: https://tailscale.com/• Codex (OpenAI): https://openai.com/codex• GLM 5.2 (z.ai): https://huggingface.co/zai-org/GLM-5.2• Qwen 3.6 (Alibaba): https://huggingface.co/Qwen/Qwen3.6-35B-A3B• Ornith 1.0: https://github.com/deepreinforce-ai/Ornith-1• Gemma 4: https://huggingface.co/collections/google/gemma-4• Playwright (browser testing): https://playwright.dev/• Vercel (preview deploys): https://vercel.com/—Other references:• DGX Spark (Nvidia): https://www.nvidia.com/en-us/products/workstations/dgx-spark/• Mac Studio (Apple): https://www.apple.com/mac-studio/• How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex: https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules—Where to find Alex Finn:LinkedIn: https://www.linkedin.com/in/alex-finn-1848684aYouTube: https://www.youtube.com/@AlexFinnOfficialX: https://x.com/AlexFinn—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Transcript
Discussion (0)
What is stacked around your office right now?
I have three Mac Studio 512 gigabytes.
We got a DGX Spark, as well as a computer.
I just built an RTX 5090.
Basically at all times of the day,
each one of these computers is just burning tokens.
The number one pushback I get on all this
is your computers are so expensive.
Isn't cloud models cheap?
Isn't it $20 for a ChadGBT subscription?
Well, that's not the point.
The point is in pure ROI.
The point is the,
The use cases it unlocks.
You now have, because you have AI models running locally,
the ability to run unlimited intelligence
around the clock 24-7.
If you were to do that with a cloud model
like Chad GBT or Claude,
you would be spending outrageous amounts of money.
What else fun are you doing with AI?
The most fun I've been having lately
is building out my software factory.
I have two loops in Claude Code.
going. I have a build loop and a review loop. First, it has a build loop where it'll take all those
tasks and start building out the tasks it came with over and over and over again. And then it has a
review loop that takes all the tasks that were built and has another agent go in and review it.
Once that's reviewed, it pings me on Slack and I can just leave a rocket emoji. And when I leave
the rocket emoji, it says merged. And my Henry loop goes and merges it. It's been a blast kind of
cracking this nut of how do you build your own software factory.
Welcome back to How IAI.
I'm Claire Vow, product leader and AI obsessive,
here on a mission to help you build better with these new tools.
Today I have Alex Finn and he's going to walk us through his two Mac studios,
DGX Spark and the Nvidia-backed computer he built for himself
and demystify what it means to run local models
and have ambient AI working for you 24 hours a day.
Let's get to it.
This episode is brought to you by Runway, a new kind of creative platform that has everything
you need to generate any image, video, or piece of content you want, all in one place.
With Runway, it's now possible to go from initial idea to a finished deliverable in a matter
of minutes. From turning low fidelity product shots into campaign-ready imagery all the way
through putting together big brand films, Runway can help your team scale your creative ambitions
while keeping your budgets and timelines from doing the same.
Runway brings together the world's most advanced AI models,
which is why enterprises like Microsoft, Robin Hood, Amazon, and Adobe,
along with studios like Lionsgate and Legendary,
all use Runway to ship real work every day.
Try it yourself at runwayml.com slash how I.AI.
promo code HowI.A.I.
Alex, welcome to How IAI Hardware Edition, Local Model Edition.
I am so excited.
So before we jump in, tell us just what is stacked around your office right now.
It feels like Asana in this office right now, to be quite honest with you, because I got a lot going.
I have three Mac Studio 512 gigabytes, which apparently puts me in the wealth class of Elon Musk having these now.
I think they resell for like $30,000 each.
We got a DGX Spark.
As well as a computer, I just built two days ago a basically a big computer built around an RTX 5090.
So we have, I think what's that come out, to five or six different computers for AI.
So a lot of hardware around here.
And you are making really good use of it.
Yeah.
So I use it basically as what I call ambient AI to support.
my entire life and we'll go through all the things is doing.
But basically at all times of the day, each one of these computers is just burning tokens,
doing things, helping my life, right?
Unlimited AI is basically how local AI works.
And I'll have them doing jobs around the clock for me, which you really can't do a cloud,
you know, APIs.
So how do we get here?
How did you become hardware, local model, like, like, why is it so hot in your, in your office?
What brought you to this moment?
My big awakening was back in January when I discovered OpenClaw.
And so I was scrolling X one day on a Friday.
I saw a blog post around OpenClaw.
I opened it.
I'm like, wow, this is really interesting.
Don't know why my gutton things.
Like, I got to buy a MacMini.
This is before anyone was talking about OpenClaw, by the way.
I get a Mac Mini, put it on, start using it.
It was one of the most aha awakening moments of my life.
and something about building this like personal bond
with the open claw
and with this agent was like
I want this to live in the computer
I don't want this to come from the cloud
I don't want this I want this working for me on my computer
and like this is the future
just having this kind of personal assistant
on your computer
and so I started doing research
okay how can I run models locally
and I came in inclusion
I'm going to go now after buying this Mac Mini
and buy a Mac studio with tons and tons of RAM
and run local models and run it locally.
And so that was kind of my, I guess,
red pill moment into local AI was using OpenClown.
It only just advanced the last couple months.
You know, then they started banning frontier models like Fable
and then the hardware prices start exploding.
And it just felt like everything was moving in this direction
of like sovereign own your own intelligence.
and so I just can invest in more and more, more into it over the last couple months.
So, I mean, I had a very similar moment, which is like just like got deeply claw-pilled in January.
Like it just, I remember I told this story, I woke up on, they can turn to my husband because this is kind of things I say to him.
And I was like, I'm having a chat, like truly having a chat GPT moment where I think everything from here on out is going to be totally different.
So every red pill that we took is that little lobster color.
I think. And I, look, I accidentally, by accidentally, I just like waited too long to get a studio and now literally I can't afford one or find one. But I do have little fat stacks of minis back here that are doing plenty of good work, but through cloud API. So I'm a little bit behind you.
I'm a strong believer, though, that those minis will be very useful soon enough.
Like Google is doing a lot of research around like optimizing models so they can run on hardware like minis.
So you won't be out of the game for long.
Well, and I still find use for them.
They're still happily at work.
They're just at expensive, expensive cloud work.
But let's go to, you know, you have all these machines.
You have bought them and you have built them.
kind of what's good for what?
Can you walk me through, like, how you make decisions
and how you actually just get these models set up
and running on these machines?
Yeah, so I've been experimenting the last few months
around the different machines,
the different capabilities, what they're good for,
what they're weak for.
Obviously, I started out with the max video.
I bought three of these 512 gigabytes.
They've been great.
But since then, I've experimented with kind of AI-focused computers
like the DGX Spark.
And then just recently, most recently, I've been buying kind of traditional GPUs from
Nvidia like the 5090, soon the 6,000 pro.
I put together a little chart here, which I'll show you.
So you basically have four different options.
Your Mac studios, your AI computers like the DGX Spark, which is getting very popular now.
Your kind of powerhouse traditional Nvidia chips like the 5090, building computers around
these chips.
and then basically everything else, laptops, Mac minis, whatever you got sitting around your house.
They all have different strengths and weaknesses, different reasons why you'd want to go for one over the other.
First, there's the Mac Studio.
What's very powerful about Macs and the reason why a lot of people are going for them is they have what's called high unified memory.
When you buy like an Nvidia chip, for instance, and you build it into your computer, you have kind of
two separate types of memory. You have your memory that runs all your programs, and then you
have your memory, which is like your V-RAM, which is for your graphics. And when you run AI models
on that, it's all in the V-RAM. The issue is the V-RAM is very small. With the Mac, it's unified,
so everything's the same. So whatever memory you have on your computer can be used for graphical
processing. And so if you buy a Mac with a ton of memory, 512 gigabytes, 256, even 1,000, you
128, you can use all of that for graphical processing, and that graphical processing is what,
you know, runs the AI models. So Mac studios are fantastic for a huge, big models because you can
use unified memory, right? I'm running GLM 5.2, which is Opus's 48 level, you know,
intelligence on one Mac studio right now, which is unbelievable. The downside is you have very
low memory bandwidth with Mac computers, which means it can't process a lot of it all at once,
which means speeds are very slow.
Right.
So the models are very, very slow.
If I send a prompt to GLM 52 right now on my Mac studio, it might take five minutes for me
to get a response back.
So it is very, very slow, but you have frontier intelligence roaming on your hardware, which
is amazing.
Then you have AI computers.
This is like the DGX Spark, very popular.
popular right now. It's like, I think, the price just went up like, I think, $4,600 or a micro center
for $4,000. These are plug-in-play AI workstations. They have actually unified memory with
Nvidia, so you do get a lot of memory, 128 gigabytes, and you also get pretty decent bandwidth,
and you get the Nvidia architecture, which is called Kuta, which basically gives you a lot
of speed as well. So it's kind of this sweet spot where you get decent memory and decent
speed. And so you can run
like kind of these mid-sized good models
like Quinn-36
pretty quickly, which is
nice. It's also very plug-in-play,
these AI computers. You don't even need a monitor.
He's plugged into the wall and you can connect to it.
And then lastly, you have
your traditional Nvidia chips. This is when you
go on Twitter and you see
the people saying buy GPUs and they have these
huge racks of all these chips.
These are all invidia chips.
and these are
probably the most powerful
of all the three
you get kind of lower V-RAM
right the 5090 which is a $4,000
chip only has 32 gigs of
V-RAM but it is lightning
fast it is extremely
high bandwidth and you're getting
like cloud speeds
but locally which is really amazing
so these are really like your
three I and then you have everything else the Mac
minis your old laptop
you can still run local models
on them. There's options out there like Gemma 4, like a few other really small ones.
They're not going to be anything close to frontier, but you can do small things like embedding,
which is basically managing memory for your agents. So you have these options. And so Mac Studio
and just like big beefy models, high intelligence low, kind of these AI computers, kind of like
sweet metal spot, and then these chips, very, very speedy. And then, you know, does feel computers.
maybe these like point solution models that are like small for a specific use case that might be
beneficial to be running in some context but are probably not going to be the thing that you just
rely on over and over again. Though you still, on dusty old computers, what I say is like you can
still put them to work in terms of parallelizing your cloud work across cross computers. So we have
a bunch of old computers and it's like, I cannot do enough on my MacBook pro. I'm kicking off
stuff to other computers just to like run work trees and do all sorts of interesting things.
I'm sure I could do that in the cloud too. But you know, I like to make use of all this,
this hardware sitting around. No, and it's, it's, uh, I like to do that as well. I have two Mac minis and
I do the exact same thing. And these apps are so good now. Like Codex, for instance, is maybe the
best agent harness out there right now. You can just go to Codex and be like, hey, go on the Mac
mini, right, start building an app on there, test it yourself, do playwright so you can go through
the flow, green shot, send me videos. So I actually like running it locally, like on the Mac
Mini as opposed to cloud, because you can have it actually click through and do things. Yep.
And then like, then do screenshots of what it's doing. So it's a lot better for that too.
Completely agree. Okay. And then how do you get these, you know, just high level set up with
these models? What is kind of your typical install on any of these missions?
So the good news is OpenClaw and Hermes has made this process 10 trillion times easier, right?
Before you would have to go find the right model, find the right version of it, make sure it can
fit into memory, download it, running on a server, all these really complex things that a normal
person would never be able to do in like a thousand years.
Well, the good news is you can now take Hermes or OpenClaw, basically make it your IT guy, say, hey, check out my Mac Studio, check out my DGX Spark, whatever you got, see what the hardware is, and then find whatever model you think is most appropriate for that hardware and load it up.
And then so as long as you have an agent, Hermes or OpenClaught, as well as tail scale, which basically allows you to create a private network across all your devices.
Your agents, basically your IT guy, and can go across all of your devices,
install whatever models and needs, set it up, do whatever you need.
I'm going to show you a dashboard soon, which shows all my models working and running.
It's all coordinated and run by my Hermes and my OpenClaw.
So as long as you have Hermes OpenClaw going, you're going to say,
hey, openClaw, check out the new Mac Studio I just bought.
Look at all the hardware, figure out which model we should run.
think about the use cases that are appropriate for me
and then load up a model and get it going
and you don't need any technical knowledge whatsoever
they'll go across your devices on tail scale and set it up for you
there really is like no technical work needed
and so to repeat this for folks just so you want to I'm understanding
you have a machine set up with open claw
you also have all your machines networked on tail scale
which allows you to have this like little virtual private network
and then that one sort of like IT guy, OpenClaw or Hermes agent,
can then use TALScale to just go into these different machines,
configure, and manage them.
Yeah, exactly.
So,
Tailscale, like, is worth getting even if you just have one computer
because it creates this private network where even if you're vibe coding,
like, on your MacBook Pro and you just have a phone,
you can, like, go on the local host on your phone.
And from, that's on your computer and, like, test your local apps out
because it's all in the same network now.
But it's even better if you have multiple computers.
So maybe you have your MacBook Pro, then you buy a Mac Studio for AI or a DGX Spark.
You install, tailscale and all of them.
And then you can just say, hey, go on this other computer and do this.
Load up the model, whatever.
And it will jump between all your devices.
No technical knowledge needed.
And load up and run anything you want.
This episode is brought to you by Jira Product Discovery.
AI has made individual PMs incredibly productive.
But multiplayer mode is where it.
still breaks, getting everyone aligned on what should actually get built. Decisions live in a
markdown file from last week. The roadmap's a spreadsheet no one's looking at. Giro product discovery
is where teams actually decide what to build. Capture ideas, prioritize them as a team, and share a
living roadmap everyone works from. It's powered by Atlassian's teamwork graph so it can pull in
customer feedback, what your team's shipped, plus your goals, and suggest what to build next. And when a
decision is made, you can hand it off straight to Jira, so a developer or even an agent can pick
it up and start building. Teams at Canva, Deliveroo, and Toast already use Jira product discovery.
Join more than 25,000 teams at Atlassian.com slash how IAI. Start building the right things together.
Okay, so we've talked enough about hardware. We've talked enough about how to get set up.
show me how you use it. So what are you using all this sort of intelligence for and how do you keep it burning tokens, you know, effectively?
I've been building a system over the last couple months to coordinate all these computers and all these models.
I do a lot of different things just to kind of set the stage here.
The number one pushback I get on all this is, well, your computers are so expensive.
Isn't cloud models cheap? Isn't it $20 for a ChadGBT subscription? Why the hell would I buy a $10,000 Mac Studio? That's like 11 years of Chad GBT usage. Well, that's not the point. The point is in pure ROI. Not everything in your life is pure ROI dollars and cents, right? The point is the use cases it unlocks, right? You now have, because you have AI models running locally, the ability
to run unlimited intelligence around the clock 24-7.
If you were to do that with a cloud model like Chad GBT or Cloud,
you would be spending outrageous amounts of money.
So you wouldn't be running at 24-7 burning tokens around the clock.
But because it's local, because you have unlimited usage of it,
you can burn tokens around the clock.
And so that brings me to my fleet control I call it.
This is my fleet dashboard, which allows me to see all my computers,
which models are running on them currently,
and monitor everything they're doing and organize their 24-7, 365 tasks.
Throughout the day, the local models I have running
are constantly doing work for me to support my life,
to support my many lines of business.
I'm building a SaaS right now, Henry Intelligent Machines.
Every 30 minutes to an hour, one of these local,
models does a security scan. So actually picks out an API endpoint or some part of my code runs
a security scan on it and makes sure it's secure. Another local model every half hour or so does a code
review, picks out some piece of code and deslopifies it, right? Finds ways to optimize it, finds ways to
speed it up, finds way to make it better. Another local model will every 20 minutes look at Twitter,
Reddit, product hunt, hacker news, and look for signal, right?
As a problem solver, the only way I can solve problems is if I find them, right?
And so I have this ambient AI going online, 24 hours a day, reading all the social media
sites, looking for signal, looking for challenges people are having.
If someone goes, man, I really wish I had a piece of software that allows me to edit my videos
like this or something like that,
my agent will find that signal,
put it in my queue,
and I'll be like, okay,
can I build a SaaS to solve this?
Can I build a program to solve this?
And this is,
there's many screens here I can go through,
but that is from a high level
what these agents are doing
and the strength of local AI
is you can have it running 24-7, 365,
just doing different things online for you.
And so for that,
let's just talk about the coding use cases really quickly.
So for these like automated security scans,
automated like quality checks.
Do you feel like the local models are of sufficient intelligence to get the job done?
And are there specific models that you've applied to, in particular, the coding use case?
When getting into local models, you really want to understand what is the delineation between
what you want to do with local models and what you want to do with frontier intelligence, right?
You don't want your entire security apparatus to be run by local models.
The intelligence just isn't there yet.
But where the advantages and the strengths come in is basically it runs this security scan, looks for these challenges, and then what it does is every day it builds a report.
And you can see some of them here where it'll build a report around what the security issue is, what the code snippet is, what the problem is.
Put it all in this markdown file that describes, like, for instance, today's has 374 findings in it of security.
security issues. Every day I have a loop running in Claude code. So slash loop 24 hours, where it goes and takes
whatever the latest finding is from the local models and reviews it. And then goes through the code,
exactly where it points to and sees if it's a real issue and how to fix it or not. So it's, you know,
if I were to loop Claude code like the local models doing every 20 minutes, look at different,
I'd be spending thousands, thousands, thousands of dollars a month. But the, you know, the, you know,
this is almost like the business development rep, right, that's going qualifying leads and then
Claudecodes like the closer, taking those leads and like, okay, what can we fix? So like, you don't
have the closer doing all the work. How do you federate work across these machines? And so we're like,
do you have some that you're like, this is my coding machine, this is my market research machine.
Are you like round robin? Again, we're like using very SDR terms. I'm like, are we round robining these
leads? Like, how we've both been in the SaaS world.
way too long.
Exactly.
Yeah, how do you do that allocation?
Again, every model computer has different strengths.
GLM 5-2 is opus level, right?
But it's very, very slow.
So I don't need it doing super fast work.
So I can have it doing the security scans, right?
I can have it go and do security scans.
Even if it takes 24 hours, that's fine.
Claudecodes only checking the report once a day.
So it can do the kind of super critical smart stuff.
On the other end, Quinn 36, which is not quite as smart as GELM,
I can have that reading Twitter and Reddit all day, just looking for signal.
That's a very simple thing to find.
Look for someone who has a challenge.
And so it's just pulling in data, crunching it and reading for challenges,
which you don't need the highest intellect.
So you just need to know, like, what's the strength of the computer?
What's the strength of the model?
And what would be tasked that be appropriate for those strengths?
Is it in this daily brief that you have these tasks assigned to you?
Like, is this the dashboard for you, the agents and the machines all to collaborate?
So I'm going to advance it to the point where it's like collaborative with Claude Code and Codex eventually.
Right now the system where Claudecode is like looping every day and going in, that's happening separately.
It's just in its own chat inside Claudecode where loops over and over again.
But no, this is all around my local models.
But I think where the connector is with everything else going on is with my open claw and my Hermes.
They're basically the I-I guy that runs this as well, this dash four, then communicates with Claude Code.
Hey, can you kick this off?
Can you do this?
So it's kind of all glued together by my personal agent.
Okay, I have to ask you a question.
Open claw Hermes.
You run both?
You have a favorite?
Tell me.
I use both.
Because much like clawed code and codex,
it feels like there's some days where one is really, really dumb.
And the other is significantly smarter.
And I feel like uninstalling one and getting rid of the other.
I'll say this, though.
The dependability of OpenClaw turned me off.
There was a run for like a month straight where every update I did broke it.
And then I had to spend half an hour fixing it.
Hermes, I've never had that issue.
It seems to be a much more dependent.
application. I would say this, if both were like 100% dependable, never broke, I can lean on it no matter
what, I'd probably be using OpenClaw because I think I've had the most wow, impressive big bang
moments with OpenClaugh. I just can't afford to be like spending a half hour once a week
fixing it from breaking for some reason. And so that's why I think I've been using Hermes agent
a little bit more lately. That's, I mean, that was my general takeaway as well. I was like,
OpenClaw doesn't doesn't functionally work, but it works in my heart.
And Hermes functionally works, but has not cracked my heart yet.
The trick that I have, I mean, we all have, like, agents, managing agents is now I have
the lifeguard.
He's an open claw.
He runs on his own gateway.
And his only job is to keep the other agents alive.
He does not get upgrades as much as everybody else.
Because truly, I was spending, like you, all my time trying to figure out why my agents
were broken. But I just love my little open claws. I can't get over it. It's just too good.
I have an emotional attachment to my open claw. I've never had an emotional attachment to my Hermes.
But I eventually got to the point where like, I don't care about emotions anymore. I just got to
get work done. But I have a similar setup where I have so much failover. So I have, I think three Hermes
agents I'm running, an Opus one, a Chad GBT one, and one on one on a local,
model and then two open clause, an opus and a chat GPT1. And like, I didn't give him point of those
five agents, like three are always down for one reason or other. But the good news is I have two that
can go and fix the other. So I have that failover as well. Perfect. Well, you should have so much
in terms of like how to pick hardware, how to set up hardware, how to manage all this compute across
all your machines, use cases from engineering perspective, from a market discovery perspective.
What else fun are you doing with AI?
Any other ones where you're like, I didn't get to show a fun workflow,
but just something that's really been a delight for you to use with either local or kind of cloud models?
The most fun I've been having lately is building out my software factory.
So there's been this big trend the last month.
I'll show this to you as well in a second of people talking about loops online.
Everyone's kind of like vague posting about loops, which is, you know, at first it was kind of angry me.
It was pissing me off.
Like, why are these people vague posting about loose?
Why wouldn't they go into it?
So I spent like days locked in trying to figure out how it can make a really productive loop.
And I've just had like a blast the last few weeks trying to make this loop system completely autonomous.
And this is like part of what I showed you ties into it.
like the security reviews, the code reviews ties into this loop.
But basically what I have doing, I'll try to show you enough here.
Let's see what I can show you.
Basically, what I have doing is I have two loops in Claude Code going.
I have a build loop and a review loop.
And basically what I do is first, I'll show you Claude Code in a second.
First, what I do is I go in a Claude in the morning and I go,
So morning build and asks me a bunch of questions what I'm thinking about.
Comes up with a bunch of tasks to build for my SaaS Henry Intelligent Machines.
And then what it does is it goes in here and first it has a build loop where it'll take all those tasks
and start building out the tasks that came with over and over and over again.
And then it has a review loop that takes all the tasks that were built and has another agent go in
and review it and fix any code or anything like that.
And then from there, once that's reviewed,
I can go in to Slack.
It pings me on Slack.
And it shows me everything that was built and reviewed.
And I can just leave a rocket emoji.
And when I leave the rocket emoji, it says merged.
And my Henry Loop goes and merges it.
And so I have this new workflow.
I've had a blast building out,
which I think is the future of like software development where,
you know, before you, you,
you prompt your agent all day,
hey, build this, now build this,
now build this.
And you kind of handhold it as it builds things.
Now it's, I go in,
we talk to each other a little bit,
comes up with a bunch of ideas,
spends all day building it
and reviewing itself and testing it out.
At the end of the day, I come in,
I see what's merge ready,
and I just leave a rocket chip emoji
on every single thing that looks ready,
shows me exactly how to test it,
puts it on its own Versel preview site,
and I can go in, test it out,
out and then give it a rocket chip emoji.
So this has been like the most fun, interesting thing I've been doing recently.
And, you know, I'm figuring out how the local models can tie into it and support this.
But it's been a blast kind of cracking this nut of how do you just like build your own software factory?
I love it.
I love it so much.
It's giving me some inspiration on some things that that I'm going to do with my own loops.
And in case folks missed it, I did do a WTFR loops episode about a week ago.
very clear.
I do three loops that you can copy and paste.
So if you're still trying to figure out,
I too was like,
why are people being vague about loops?
Like,
this is not a scary or mysterious thing.
So we just popped up the screen share and showed a couple of them.
Thesis, by the way,
I'll share the thesis.
My thesis is that why everyone's being vague
and not like sharing how they're doing it.
Like,
neither open AI or clutter,
like they talk about a lot,
but they haven't shared it is like,
I think this is kind of the last moat for a lot of these companies.
is like, you know, anyone could build anything they want,
but the actual infrastructure around it and like how you automate that to put out more code,
like they probably have their own systems that pump out or ridiculous amount of high quality code.
If they were to share how they do it, other people would be able to copy and they'd be able to be equally as productive.
But like this is kind of a motive theirs now.
Like if they can figure out a really good system that builds high quality codes,
I think that's why they're being vague and not sharing it.
And I have a, I have a suspicion about why non kind of model people are being vague, which is one, their use of loops is probably pretty boring.
Yes.
They're not like, guess what?
This is my loop in the morning.
It runs a cron and does this thing.
And two, they probably get more eyes being vague than being specific.
So I am very cynical about the vague posting, which is why we are a screen sharing podcast here.
Well, Alex, this has been super fun.
Let's do quick lightning round, and then I will get you out of here.
Question number one of all the hardware, which one is your favorite?
I'm torn between the Macs Studio and the 5090.
I like the Macs, because I love the integration with all Apple device.
Everything owns Apple, iPhone, iPad.
The integration being able to run models side by side locally.
First of all, I think is the future of computing.
I think Apple will start running local models built in that help you out with an
next 10 years.
But I feel like I kind of lean the 50-90 because I can play
cyberpunk 2027 in all Ultra with all the hyper-realistic mods.
So I think I had to lean the 50-90.
Okay.
And then on the model side of all the models that you have locally installed,
do you have one that you just really love?
So I was on Quinn, first 3-5, then 36, for basically the entire
span of last five months. I've been on local models. But a new model just came out. I know nothing
about the team. I've done zero research. So I hope to God I'm not promoting bad people. But it's called
Ornith 1.0. And I think they did some like reinforcement learning on Quinn and improved it. It made
it even better at coding. And every eval I've run on it has shown that it's better than Quinn. It's
faster and smarter. And so Ornith 1.0, 35B, has been my, my
most used model recently, and you can run on a DGX spark.
So anyone who has that, you can load it up and it works great.
Who thought that you and I,
SaaS people, would just be like Ornith 1.083B, DGX,
like just letter after letter after letter.
But this is our life now.
I love it.
We both had the same background working for SaaS marketing tools,
spamming people all day with email to now the nerdiest technology on planet Earth.
That's exactly right.
And then last question, I ask everybody, when your open claw, your Hermes agent is being real dumb and not listening to you.
What is your prompting strategy?
Are you extremely polite or less so?
Let's just say this.
If my chat logs were to ever leak to the internet, first of all, you would never have me on your podcast again.
Second of all, I think I'd be taken off every single social media site.
So I am a pretty nasty person to my agents.
I am not nice to them.
I'll just say I've threatened them multiple times.
The threats never seem to work, even though I threatened to hurt their agent family.
It doesn't work. They still fail.
But, yeah, I'm pretty mean.
But I find that when just much like Claudecote and Codex, one stupid, I just go to the other until things calm down and they figure out what's going on and it gets fixed.
And then eventually they're smart again.
Amazing. I love it.
Well, Alex, this has been so fun.
Where can we find you and how can we be helpful?
I'm Alex spin on YouTube.
Alex spin on X.
I have a community.
You can join the Vibe Code Academy.
I also have two SaaS I'm working on.
Creator Buddy, which is a basically operating system for Twitter,
as well as Henry Intelligent Machines, which is coming soon.
So if you just subscribe to me on YouTube, you'll hear about all the other things eventually.
Amazing.
Alex, I will get you back to your machines and back to building.
Thanks for joining Howie AI.
Thanks so much for having me, Claire.
Thanks so much for watching.
If you enjoyed this show, please like and subscribe here on YouTube.
YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on
Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review,
which will help others find the show. You can see all our episodes and learn more about the show
at how IAIIPod.com. See you next time.
