How I AI - ChatGPT Codex Voice + browser + Sites: an expert’s AI workflow | Nick Baumann (OpenAI)
Episode Date: August 3, 2026Nick Baumann is on the Developer Experience team at OpenAI, where he spends his days building with, testing, and communicating the capabilities of ChatGPT Codex and ChatGPT Work. In this episode, Nick... walks me through several features that have launched or evolved recently: the new voice interface with its screen-reading orb, the Heartbeats automation system in ChatGPT Work on mobile, the live ChatGPT Sites deployment feature, and his personal use case for AI-assisted UGC video editing.What you’ll learn:How two-person voice chat worksHow Heartbeats workHow to build and deploy a live website with ChatGPT SitesHow to delegate a flight search, hotel booking, and expense report to Codex in a single voice conversation without opening a single app manuallyWhy ChatGPT Work on mobile is the most underutilized AI workflow for people already using the ChatGPT appHow to use a custom UGC Video plugin to feed 50 raw clips into ChatGPT, let it pull transcripts, pick the best takes, and assemble a finished vertical video overnight—Brought to you by:Bolt.new—Turn your idea into a real productHyperagent—Deploy fleets of agents that handle real work—In this episode, we cover:(00:00) Introduction to Nick Baumann(02:56) What’s new in Codex(05:40) ChatGPT Work and Heartbeats(06:40) Live Codex voice demo(13:25) Latency vs. intelligence(14:36) Quick recap(15:04) Voice on mobile and the ChatGPT Sites workflow(21:24) Live UGC video demo(32:30) How I AI website results(34:04) Lightning round and final thoughts—Tools referenced:• ChatGPT Codex: https://chatgpt.com/codex• ChatGPT Sites: https://chatgpt.site—Where to find Nick Baumann:LinkedIn: linkedin.com/in/nick--baumann—Where to find Claire Vo:ChatPRD: https://www.chatprd.ai/Website: https://clairevo.com/LinkedIn: https://www.linkedin.com/in/clairevo/X: https://x.com/clairevo—Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.
Transcript
Discussion (0)
If we go back to 2020 to 2023, around the release of chat chbt, I think a lot of people felt the AI then.
A chat bot that I say things to and it intelligently says things back to me.
I think with the advent of Cody agents and like the Cody agent harness in particular,
now this AI is going out and like reading files and creating files and running commands and this is crazy.
I saw it one thing and then it figures out all the other things on it sound.
I haven't gone deep with voice. I want to see it from the pro.
How should we be using it? What's cool about it?
So I'm just going to trigger voice. I've got a hot-
key for it. I'm listening. Go ahead.
Okay, there should be some Amazon receipt for a mic I just bought recently.
It's for DX uses. I'm hoping you can help me expense it.
Could you spend up another thread to handle taking care of that expense report for me?
The expense task is running. It will find the Amazon receipt and stage the Navon report,
but not submitted until you confirm. I'll let you know when it's ready to review.
If anybody has been lucky enough to have an assistant, this is how it gets done. And
it is just like such a nice delegation experience to be like, hey, can you take?
care of this, hey, can you fix this? Tell me what's going on. Typically, voice is great.
You can like talk back and forth, but it's usually with a less strong model that is less capable.
Whereas this, it's fully able to delegate and manage these fully, you know, five, six sole threads
on its own, which is great because then you can have it do things on your behalf.
Hot take question. What do you think matters more on voice experience, latency or intelligence?
Welcome back to How IAI. I'm Clairevaux, product leader and AI obsessive,
here on a mission to help you build better with these new tools. Today we have Nick Bowman at
OpenAI and he's going to show us some of the advanced use cases of Chat ChachapT Codex and
chat ChachyPT work, including how you can talk to your computer to book your flights, use chat
you pt sites to build websites you can share with anybody or nobody and use my favorite
workflow, which is editing creator content in Codex. Let's get to it. This episode is brought to you
by Bolt.new, the AI app builder for people who have ideas and want to ship them.
Most AI tools spit out code that looks great in a demo and falls apart the second you try
to do anything real with it.
Or they lock you into their own platform with no real way out.
Bolt is different.
You describe what you want to build, a startup MVP, a landing page, an internal tool, a side project.
And Bolt generates production ready code in minutes.
Connect Stripe or any other MCP, hook up your domain and deploy it live.
Founders are using Bolt to build businesses doing real revenue.
Product managers are shipping prototypes their teams actually use.
Designers and marketers are launching campaigns without waiting in line.
Anyone can build. Engineering can ship.
Everyone wins.
You just need an idea and a weekend.
Check it out at bolt.new slash how I AI.
Nick, welcome to How IAI.
Hello, Claire. Thanks for having me.
I am going to make you laugh because I might be the number one Codex fan, fanboy,
fanboy. In the world, I am like constantly telling people,
wait, but have you tried Codex yet?
It is just my daily driver, so I'm so excited to have you on because even though I use it all
the time and for some things that you and I think both think are valuable, I have not touched
most of the new features that have come out in the last week.
It's really hard to keep up.
And so I'm really psyched that you're here to just show us some of the new stuff
and maybe how you use Codex both for work and life.
So tell us what are the things in Codex?
Do you feel like people don't really know about or brand new
that you think are completely going to change how we all work?
Yeah.
I feel like with yesterday's release of Voice in the Chat, QBT, QVT, Codex app,
it's kind of like the amalgamation of all these primitives.
we've been like putting together kind of quietly,
and now we're super useful.
So a lot of people don't know that the chat ChbT app,
you can ask it to create threads,
you could ask it to message your existing threads,
and these threads actually talk to each other too.
And it's like not super kind of forward in the app,
but these are all capabilities.
And now we've added this voice orchestration layer really,
where now you can, you know, with the hot key,
you can trigger this orb that like pops on your screen.
And no matter where you are in your computer,
You can talk to it.
It'll talk back to you.
And it can see your screen.
You can see what you're working on.
And because we have those primitives of see your existing threads, create new ones, talk to them.
It can basically manage your entire chat.
You can be able to out for you.
Yeah.
One of the things that I don't want people to miss before we even get into the fancy voice stuff is I really do feel like most people do not know the meta capabilities of codex and even like we won't say, but like the other one.
In that, like, its ability to not only spin off subagents, but actually start completely new threads, fork threads, look through all data, look through its own memory.
It's a really good mechanism to let the AI manage context and decide when forking off different tasks makes a lot of sense.
And so I've been doing that a lot, but I actually had to discover that capability sort of organically when one day I was sitting there and it was probably 5'6 was like, I'm going to kick off a new thread.
And I was like, you're going to do what?
Like, what now?
And so I do feel one of the challenges with these core kind of platforms like codexes and chat
GPT is people just don't know how to discover the features.
And they're so rich.
And so do you have any, you know, like we'll talk about voice.
But are there any other like sort of edge features of codex or chat GPT that you feel like
people underuse, not just this thread one?
The chat GPU work experience, you know, people are talking about.
about, you know, should we using chat chbt work or codex.
I think chat chapti work on the web, specifically mobile,
and that's where I use it mostly, is so insanely powerful.
It's one, it's essentially like the codex harness,
which has this incredible compaction.
You have all your plug-ins, so if you've already connected to
Gmail and g-cal, like, whatever you use, it's already there.
And then you have it all accessible on your phones.
You can like create these automations,
also it's heartbeats, and it's already there.
It's like, I mean, a few times a day, I'll just get,
like a notification saying, hey, I got this email about this thing you bought. Are you sure you bought that?
Or like, hey, this package is arriving. Make sure you grab that. And I'm not watching my email
or my finances, but that's just kind of being monitored for me. And that's, that's really magical
experience. Okay. I love it. All right. So let's get into chat chitipit voice and what you,
what you use it for. Because I, it came out yesterday. I've been super busy and around humans. We're
going to talk about that a little bit later. So I haven't gone deep with voice. I want to see it
from the pro. How should we be using it? What's cool about it? Yeah, I think it's really magical.
So I'm just going to trigger voice. I've got a hot key for it. I'm listening. Go ahead.
I'm hoping you can help me handle a few logistical things right now. Can you see the offsite?
I've got pull up and slack? I'll take a look. Let me check that Slack view. Yep, it's the DX team
offsite in Paris, August 26 through 28. You're due in Tuesday with the program starting
Wednesday morning and wrapping Friday afternoon. Flights and hotel are supposed to be booked in
Navon this week. All right, so I've muted myself so that Kedex doesn't hear me right now.
But what happened there is it can already see my screen. You can see that I've got Slack opened up.
And what it did is it took an app shot. And so an app shot is something that you can do or
whatever window you have opened, you just hit command command,
and it essentially takes like a smart screenshot loaded with context.
I can use computer use or browse use on it instantly.
But, you know, kind of a meta feature,
Codex Voice already has that built it.
So I'm hoping you can help me book my flights in a hotel for the Parasoft site.
I'm also thinking about sticking around through the weekend
and maybe a couple days after.
Could you maybe check my calendar, see what my availability is like
and see if it's feasible for me to stick around?
I'll check your calendar against the offsite dates and then look at travel options.
Is this using the connectors?
Good question. Yeah. So this is using the plug-ins underneath the hood right now.
So it's not having to actually go.
You're clear after the offsite through Monday, August 31st.
A Tuesday, September 1st return would conflict with a late afternoon meeting.
The Simplice plan is returning Monday, August 31st.
Do you want me to plan around that?
You know what? I'm thinking I might just take PTO on that Monday and Tuesday.
I'm flying out of SFO, so why don't you just go ahead and use Nivon, spin up another task to figure out the travel, pick a hotel that's reasonable and within our company policy, and then when it's ready, let me know, and I'll click book now.
Sure, I'll sort that out. I'll find options for flying out of SFO arriving Tuesday before the Wednesday start and returning after Tuesday, September 1st.
then bring you choices that fit policy for you to look.
So is this chat going to be in?
Oh, sorry.
I'm talking to, I'm talking to his codex.
Invite Claire to the Paris offsite.
Put her up in the ritz.
Get her in business class.
And honestly, we don't really have a Paris offsite scheduled yet,
but so much of our DX team is French that I'm hoping that I can like,
you know, start seeding the ideas here.
That task is running now.
comparing flights from SFO and a practical hotel for August 25th through September 2nd.
I'll share the suggested itinerary when it's ready.
Cool. So we've got this task already running.
And make sure it uses my existing Chrome browser for this one.
I'll check that.
I've updated the travel task to use your existing Chrome session and keep it to preparing options, not looking.
Awesome. So right now it just sets a message back to that task. It's already running.
but I think while we've got that going, I'm going to ask it to do another thing for me.
While I've got you, I've got an email, I think.
There should be some Amazon receipt for a mic I just bought recently.
It's for DX uses.
I'm hoping you can help me expense it.
Could you spin up another thread to handle taking care of that expense report for me?
Sure, I'll take a look.
I've started that expense task, and it'll use your existing Chrome session.
The expense task is running.
It'll find the Amazon receipt and stage the Navon report, but not submitted until you confirm.
I'll let you know when it's ready to review.
Awesome.
So those are both running in the background.
I've muted our voice so we can talk without being interrupted.
But we can already see that in Chrome, this task is using the browser.
It's figuring out our flight options.
It'll figure out the hotel.
Frankly, this is something I just enjoy not having to do.
It's nice to just have this handle in the background for me.
What I like about voice is I do think sometimes people with AI really struggle still with like blank chat window syndrome, which is like they look at and they're like, what am I going to ask?
What am I going to ask this brilliant AI to do?
And so what I love about first just like voice transcript to text, which a lot of people have been using is it is what our beloved multi-time guest Hillary calls the Yapper's API, which is like the best highest bandwidth way to communicate.
with an LLM is just to yap out loud and just context now.
But what I love about this like voice to voice interaction is if anybody has been lucky enough
to have an assistant, like this is how it gets, it gets done.
And it is just like such a nice delegation experience to be like, hey, can you take care
of this? Hey, can you fix this? Tell me what's going on. Blah, blah, blah, blah, blah.
And so I'm guessing the hypothesis here with this voice kind of experience is,
there's going to be like better discovery of use cases and sort of like more sprawling use cases
because people can kind of delegate in a more natural way as opposed to having like sit there
and think through how to instruct kind of in text the models.
Have you found that like there are specific things that you reach for voice with that you don't,
that you don't like type with your human fingers?
Are you a full voice pill?
Like where are we?
Yeah, a few thoughts there.
One, I think it's like this iteration itself is like kind of a new primitive in that
Typically, voice is great.
You can talk back and forth,
but it's usually with a less strong model
that is less capable.
Whereas this,
it's fully able to delegate
and manage these,
you know, fully, you know,
five, six sole threads on its own,
which is great,
because then you can have it do things on your behalf.
Other thoughts there?
I find the, like,
the thinking process and voice
just a lot better.
When it comes to, you know,
I could dictate for, you know,
know, two minutes, send off a long paragraph, get back a couple long paragraphs and read those.
And it's really hard for me to get in flow. I think of like planning in voice to be really,
really effective. Hot, hot take question. What do you think matters more on voice experience,
latency or intelligence? If you could only pick one, I'm going to make you pick one.
I would say when I use chat chbtee, I put it on high intelligence for voice. I think there's like
a middle ground. I think if there are good tools you can delegate to, then I clear bar about latency.
Yeah. So I don't know. I guess I still say latency. Yeah. I think I think it's what we're going to
be talking. And this is the Claire of a prediction. I think it's we're going to talk more and more about
latency in the in the second half of a year because I just think it's like almost the thing to
differentiate on right now is like how real time can these experiences really be. I see a lot of people
abandon great AI workflows because of a spinner or because of a delay. And so like the more you
can close that gap, the more I think people can discover cool things and actually pull through and
get some stuff done. Yeah. I feel like so many product assumptions in design and AI tooling is
built around this like there's going to be latency. There's going to be delay. Yep. And the more
we shrink that, we can kind of drop those assumptions and new products will, you know, will emerge.
Amazing. Okay. So what we've seen here, just to reiterate for people who are maybe not
watching or listening is you can hockey spin-up voice. It can manage these like very intelligent
five, six threads. You can delegate tasks. It can be hook up to plugins. It can use your browser.
It can use app shots. It can use all these things. And it's just like a very high bandwidth,
low friction way to talk to super intelligence that can get nice stuff done. You mentioned that a little
bit earlier that you're also loving, like, mobile and this combination of, like, dictation
and managing through mobile. Can you tell me how, like, being able to do some of this on
your phone has changed your workflow, and maybe, like, give me an example of that?
So I put together the site, and so this is just a aggregation of all these really cool
tweets that, you know, both people in, like, the Codex and chatbGG community and also
people work at OpenAI have shared with how they're using Codex and ChatGBTBTBT.
This was done just with chat ChbT, like the chat ChbT app where I had it, use its own in-app browser, go through Twitter for like a couple hours, and then embed these links and kind of add a little bit of description to them.
And now what I do is I've got a thread on my work phone that is essentially managing this site.
And when I see something I like, I just drop in that link as they add it.
And it does.
we also have this prompts tab, and I do the same thing there.
If I have an idea for how somebody could be, maybe he's like newer to chat chitb-t work,
I'll just like dictate my phone, explain like, hey, there's this prompt for like having to manage your inbox and your finances.
Can you add this to the site as well?
And in terms of actually editing the site, deploying it, that can all happen from chat chatt chb-t work on mobile.
And I think what people are not seeing here, which again,
Then so many new features coming out that I think people aren't picking advantage of is the domain here is chatGBT.t.com.
So you can now deploy these like artifacts, these sites, live.
So you just walk us through a little bit how that works or maybe we can do one.
Just so people can understand the power of these sites and how you might use them or what their limitations are even.
Yeah, I think like for like 10 second description for the type of little audience, it supports a SQL database.
It has S-free storage or files, and it even has environmental variables you can put there.
For the non-taglo audience, it's essentially a website that you can store things on, which is great.
And you can also control who sees it.
You can filter by email, so it can be private, it can be fully public.
That's up to you.
But what we could do if you're interested, if you could try building a site live.
Let's do it.
We love to build live.
Yeah.
All right.
So let's jump into it.
So I've got an idea for a site.
I'm here right now with Claire Vaux.
She hosts the Howie AI podcast.
And we're hoping that we can build a site that takes all the best tips from the videos on
her YouTube channel and makes it so you can like easily filter through them and have quick
links to like different timestamps on our videos.
They can match the How IAA branding.
We use good old AI blurple, purple, um, and bluepil.
black and white. And let's have it, if possible, categorize by tool. So like what AI tool
is being used in the workflow and also function. So if it's for designers, it's if it's for
engineers, it's for personal productivity, any of those things. Otherwise, I'm excited to see what
you come up with. Awesome. That sounds great. And go ahead and deploy it privately. And then we'll
take a look at it when it's ready and that you can deploy it publicly from there. If it,
if it looks good. While this is running, let me make you laugh because this two human chat to chat
GPT, I did recently with my mother and she was like stuck on an administrative task and she was like
getting all frazzled about. She's like, oh, I'm going to have to look up all this stuff and blah,
blah, blah, blah, blah, blah, blah. And I was on the phone with her and I was like, Mom, just hold on one second.
And I opened up chat GPT on my desktop and I turned on voice. And then I was like,
can you just please tell me what you need to do?
And I turn the speaker phone off the phone and I put it up to the computer and had her like babble kind of like we did.
And I asked her a couple questions.
And then I was like, thanks.
Press to enter and I was like I did it for you.
I did it for you.
So I really do feel like people are underusing like two person voice chat.
You can do it on your phone.
You can use speaker phone.
Like it is this really nice again like high bands width way to get requirements into the system.
And it doesn't just have.
have to be you alone. Yeah, I've had times where, you know, I need context, like, something I work
with. Yeah. And I'm like, please just like ramble, dictate into Slack and send me, like,
garbage, I don't care. And they're like, no, let's just like set up a meeting. And I'm like,
no, don't do that. I don't need that. I don't need this to be like pretty or anything. Just
just tell me. That's fine. So are we at this meeting? Could have been a voice note, like,
level? Yeah, I think so. I mean, it's more like, give me information that unstructured,
however it is, that my agent can decipher, and that's good enough for me. This episode is brought to you
by HyperAgent, the platform for deploying always on agents that actually run your business. With
HyperAgent, you build agents in the cloud and deploy them where your work already happens,
like Slack, Telegram, or Email. An agent will scan your inbox and draft replies to vendor
follow-ups, another monitors competitors, and spins up rich ad kits and landing pages.
A third notices a deal going cold in Salesforce and writes the save email with full account
context. These aren't chatbots waiting for a perfect prompt. They're proactive, learning your
preferences, retaining your playbooks, and getting better with every run. One user built four
agents to run an outbound sales pipeline, prospecting outreach, follow-ups, CRM updates, all in a
single afternoon. No local setup, no VPS bills, no fragile permissions on your laptop,
just powerful agents with full control over skills, tools, and guardrails. HyperAgent was built by
the team behind Airtable and How IAI listeners get $1,000 in free inference to start building.
Claim yours at hyperagent.com slash how IAI.AI. Okay, so this is going to go spin up. We're going to let
it kind of like whirl a little bit.
Maybe we'll come back to what we see.
But there's one other use case that both you and I really like in Codex that I thought
you could show in particularly how you prompt it, which is Codex for video editing.
And for those that don't know it, I do a lot of this for the podcast.
I also do a lot of this just for like general work stuff.
I'm cutting a lot of videos.
So tell me like how you came to this.
use case, why you feel it so useful, and then how easy is it to prompt kind of like video editing?
You know, I do a little bit of content as part of the DX team, and, you know, we're trying
to reach people that might not be inside like the Twitterverse, which, you know, there's a lot of us
there, but there's more of us not there. And I think explaining, you know, how you can use these
tools in really relatable ways where you're explaining to a camera in like 45 seconds is great.
And so I started just kind of like going to a park, filming myself, filming my camera, getting a, you know, a smattering of little clips.
And so this is from, like this is from yesterday.
I went to the park.
These are like most of these are takes that we're not going to use.
But what I'll do is I'm literally going to drag these into chat Q2T work.
And then so what I've done, and I'm going to share this, but I've created.
I created a plugin called UGC Video.
Actually, let me let me get the one that I use for.
I've made a version that has opening eye branding,
but you can actually see it here.
And so this basically, after I've gone through
a few runs of this, what I've noticed is that
just having like some guidance around like,
hey, make sure that we're not entering
like the safety zone with Instagram ads
or let's do these formats.
We want nine by 16, also four by five.
That's great.
But what I'll generally do,
and I'll dictate this.
So we've got a bunch of clips here.
I recorded a video for the record and replay feature.
It starts with me just explaining that I'm apartment hunting
and that's kind of a pain.
And then I walk through and show the feature in real life.
I show myself actually apartment hunting,
and I show how Curtis can do it for me.
And then I have a closer about how this skills
that's been made for me can be put on a heartbeat,
and I can have Codex just search for apartments in the background.
Can you go through these clips, first, like, pull the transcripts,
find the best takes, and then kind of piece this together
into a UGC-style video?
So that's running.
A few things to kind of explain how this works.
So I've only submitted, like, 20-some-odd video clips.
And so what Codex is going to do is going to start by,
processing those clips to get just the transcripts.
This will help it understand the story,
what we're talking about,
and even like my light description
will help it understand that.
And then we'll do then, actually it's gonna ask you the questions.
Sure, we'll do organic vertical.
Let's do, complete and clean.
Sure.
So first get the transcript,
and that's gonna go through the takes
and actually process images from them
to understand what I'm showing.
And then it'll choose like the best takes.
And when I'm recording, I'll say things like cut or that's a bad take or that's a good take.
And it's able to like use that as guidance and know like, you know, what's actually useful from these.
You're going to give me so many ideas because I have like deeply abandoned my once great TikTok because
it's like, you just like, short form, short form is hard.
And I just have such a high bar that I'm constantly like recording these clips and being like that was garbage.
and then I get frustrated and I walk away.
But you've given me this idea, which is I could just keep recording
until like there's one that feels good and then just dump it all.
I don't even have to remember which one is good.
Dump it all, give some voice notes about how to edit and then and then kind of assemble it.
And then I love this process of just having the model, transcribe, look at the video,
come up with like good, good cuts and put it together.
I use it a lot for our trailers.
So anytime I do a very long talk, like a 30 to 60 minute talk, or I do a podcast that's like 60 minutes, I dump it in and I'm like, give me a 60 second hype video with the funniest parts of the talk and like give me a little bit of buffer for clipping.
And then it clips that really well.
It actually is like pretty good funny taste, which is nice.
And so I do that.
And then I also do like I teach like a three hour long workshop every quarter on how to like transform your engineering organization to.
be kind of like AI engineers. It's very, very long. It's like three-hour zoom clip. And I also do that
to like clip chapters, clip shorts, do all that sorts. It's just like all this tedium of a video
editing. It's quite quite good at. And then just tell us like what does what does this plugin actually
do? Yeah. So the plugin is really just based on like all the steering I've done really this week
of like making these videos, whether it's, you know, describing the kind of formats I need or,
you know that you don't want to have text in certain areas
because it'll be cut off or natural kind of breakpoints
or how to do captions.
It's, you know, literally I just,
I had this massive thread where I did a few of these videos
and there was so much essentially like bespoke knowledge
in this process built up that was like,
hey, can you make a plug-in out of this?
And there's actually a plugin creator skill
in the chat, TEPT app, so it's pretty easy.
And then I put a few tweaks,
and now it's more,
I'll probably get, I guess, the happy path that I want.
And can you show us an app?
I know you did one of these for yesterday.
Can you show us, like, what the kind of output is that you can get?
Yeah.
So here's my starting prompt.
Again, I just jumped in a bunch of clips.
And one thing I want to know, you talked about, like, the tedious stuff.
Yeah.
Even, like, little things, like, going through and deleting bad takes
or finding, like, the right segments, I feel like I'm so willing to,
to hand that over and just get that off my plate.
I don't want to think about it.
So yeah, I started a pretty similar prompt
that I just kicked off with a little bit before.
At one point, and actually I was in an Uber,
I said this on my phone.
I remember, I was like, oh yeah, I need four by five also.
So I sent that in as like a steer.
And it was like, all right, got it.
Updating the delivery matrix.
And then this is actually where, you know,
I was like, why aren't you asking questions?
So I had to update the plugin at the same time.
And this is what the output looks like.
So this is actually a few different, like, you know, back and forth steers.
Here, we'll actually haven't watched this full one, but we'll see how it looks.
It's insane trying to find an apartment in SF or area city right now.
And it's impossible to get to all the listings before they're going.
So what's funny, I can see that.
It even added blurs to these addresses here.
Real smart.
No, I can see it.
Yeah. Actually, I did like a couple of these for, like, that we're releasing like Instagram next week.
And there's like sensitive like release material information in my Slack that it like bowled up in my chat.
And it did this like insanely granular work of like blurring out, literally lines and then following it while I'm scrolling.
And it does all this like verification where it has these blurs.
It checks its own work.
And then it like does that over and over again until it's good.
And then it gives you, like, you know, a good result.
I probably six months ago coded a redactomatic for the How IAI podcast because I had so many people being like,
whoops, I left like my API key open in that like one piece of code that we were talking about.
And so we actually like built one.
I need to try it again with Codex because it's probably, the one challenge we have is the blurs following kind of like screen chairs.
Yeah.
And actually, that gets back to the, like, I don't, my, I don't even want to deal with the tedious work.
Like, not only, actually, it's a non-starter that I'm going to make blurs automatically, do that manually.
That's just, that's not going to happen.
What's even more knowing, though, is having to think about, all right, well, how do I generate outputs that don't have confidential information?
And how do I still make it seem genuine?
That's also a pain that I just, I don't think about.
I just, I leave it to, leave it to the chat, should be T.
I'm going to give one other benefit for the content creators' out.
there for this like AI editing of content which is it's real embarrassing to edit your own content
it's like really embarrassing to look at yourself and watch all these videos of yourself and be like
oh yeah i was like real cute and funny in that one and that one i am not like i make no sense as a human
so i like to offload like the critique of myself and the critique of my performance off to
another model so i don't have to look at my face so much as well especially
as like a solo creator, solo founder, somebody who's like working on their own work.
So I do like having a third, you know, like a neutral third party do these cuts.
So I don't have to sit there and like listen to my own videos all day.
Yeah.
I know.
It's great.
Back to like the, like it's definitely lowers like the bar, like the barrier to entry for
stuff like this.
If I just need to like go to a park, you know, I recorded like three videos.
I saw it recorded to three videos over the course of like an hour and a half a couple days ago.
And then I uploaded like 50 or 60 clips.
Yep.
And I dictated like what the different narratives of those three videos were.
And because I have a plug-in and a task that already understands what I want,
you know, it took a while.
I went to bed and I came back to it.
But I had what I needed and actually shipped one or Instagram a couple days ago.
Okay.
I have one more question on this.
And then maybe we'll see if the website is up.
Do you need Soul Ultra for this?
because that's what I see below.
Yeah.
I would.
In my experience, there's just so many,
as much as editing just for speed,
there's so many little, like,
microtasks involved in checking all these frames
and checking these transcripts.
And there's a lot of parallel processing that can happen.
And a lot of like parallel judgment, if you will.
And what's so ultra is,
is it's basically this framework for multi-agent,
for five, six, soul. And so I find that it's just more efficient and I get better outputs when he's
Soul Ultra. Love it. Okay, cool. Let's see if the HowIAA website is up and running. Look at that.
Black and purple and white, just like I said. We also have AI identify which thumbnails
are going to do the best. By the way, this episode, just you know, it's all, look, ma' no hands.
We had some problems.
Yeah.
Oh, I see what it is.
It's the use cases.
So each episode is showing up multiple times.
That was a very good episode about Codex browser use, where I let it shop for me for my Hawaiian vacation.
And it did quite a lovely, lovely job.
If we want to get really fancy, you could probably have it, have different thumbnails for each clip.
We could even have like use ImageGen to make like a, you know, a singular thumbnail.
Love it.
Amazing.
there's good old Alex over and over and over again,
and then we're going to, three of me.
It just built it.
Awesome.
And then this can be published.
Just remind me how it gets published,
who I can share it with, how that all works.
Yeah.
So let's say I wanted to share this with you.
I can say, you know, share this.
And so now you can put in, you know,
whoever you really want to.
It can be totally publicly available.
You could add individual email addresses.
Yep.
And it'll share with those.
folks. Amazing. And then how do you log in? It's logging with chat chbtee. Great. Easy peasy.
I do have a chat chbtee account underneath that you know for us. We are good to go.
Just to recap, we did a couple of use cases. So we talked a little bit about voice,
high bandwidth, yappers API, just basically like talking to a computer empowered assistant.
I can just go do a bunch of your work for you with access to.
and other things, including just being able to screenshot whatever is on the screen and have context.
We did a Chajibati site, which I did not know.
It has its own computer and memory and browser and all that stuff.
So that's fun fact for Claire, where you can have it actually go off and make a website,
which you can share publicly or share with specific email addresses, which you just showed.
And then again, my favorite use case, which is take a bunch of clips for UGC, dump them in,
make it the models problem, including doing redaction and blurring, which is for any of my fellow
YouTubers out there, a huge and very annoying issue. I'm very happy personally to have seen this
workflow and get it solved. Nick, a couple of questions, lightning round questions for you.
We will get you back to all your voice chatting. Question number one is on that topic.
In this world, where we're all loving our chatypte voice experience, are, is everybody in the
open AI office just like mumbling to themselves and to their computer like how are we going to
intersect all this voice capability with the realities of like existing around humans that find
all this chitter chatter annoying like how does it actually work i think it's like a almost
genuinely uncomfortable question because when i'm at home i exclusively dictate and when i'm at my desk
around other people i will sometimes whisper
but I'm more likely to use my keyboard.
And, you know, we have all these meeting rooms
and these, like, little pods people that hop into.
And I see people, like, hopping in there for not calls
or often these days,
because they're clearly going in there to, like, use dictation
or to use voice and codex.
And clearly people are adapting to this as, like, a better experience.
But frankly, we, you know,
I don't know if the answer is we all walk around,
like, with those masks around us
that, like, shield our voice.
to anybody with the agent, but it's not ideal right now.
But you can tell that there's a better, there's a good version, and that's using voice,
but it's not ideal for our current, you know, office makeup.
Infrastructure.
Yeah.
I imagine we like press a button and like a tube comes down.
That would be ideal.
Yeah, that's what I'm, that's what I can talk about.
And then we press our button and the tube, the tube goes up.
Okay, we're going to, we're going to figure this out until we get the direct codex to mind
connection, which I'm sure will be a 2027 release. Okay. Second question, if people really wanted to
like feel the AGI, what are what are the things that you really think are just huge step changes
in the last three months? Maybe like voice aside, what is like the one thing that you would tell
people to try or that you're really thinking is is the new interface? Because there's so many
things that you've shown here. In terms of like in terms of feeling the AGI,
I think a lot of people like you and me have already felt the AGI.
I think if we go back to, you know, 2020, 2003,
like release of ChatGBT, BT, I think a lot of people felt the AI then.
You know, what AI can be became understandable.
You understand this conception of a chat bot that I say things to
and it intelligently says things back to me.
That's pretty cool.
And then I think with the advent of Cody agents
and like the Cody agent harness in particular,
I think people in the developer community and like AI tinkerers,
they felt this AGI moment where, wow, now this AI is going out and like reading files
and trading files and running commands and this is crazy.
I tell it to do, I sell it one thing and then it figures out all these other things on its own.
And that was a huge AGI moment for people.
But I do think there was a certain barrier to entry with that where if you're not an engineer
or you're not a tinkerer, there's, there's a,
this is limited to a small-ish group of people.
But now, you know, with something like Chat Chb-T work in the web,
I think we're taking a lot of the primitives,
you know, like the coding agent harness,
like having an always-on agent,
and we're making them not only accessible at the web,
but they're accessible to your phone.
And so for people that are already using the ChatGBTGPT app,
which is a lot, they can just switch over the work tab.
They can ask a question like, hey, can you check my Gmail,
and also my finances, which I know are two, you know,
plugins that you can install really quickly.
And just monitor those and let me know if anything stands out
or anything has gone awry.
And now they suddenly have this agent that's, you know,
maybe twice a day, checking their email, checking their finances.
And I think that like small thing, you know,
when they first get a notification, you know,
letting them know maybe that packages arrived
or there's some weird charge,
I think those are like the small moments
that people might kind of like,
it might be a foot the door for a lot of people.
Okay, amazing.
And then I have one last question I ask every guest,
and because you speak to AI,
I'm really curious here, what you say.
When, I saw a little bit of your prompting strategy,
it's a little bit of what I do,
but when AI is not responding,
when it is not doing what you want,
when it is not using that plugin,
what is your prompting strategy?
Do you yell?
Do you all caps?
Well, it depends.
If I'm sitting on my desk or if I'm in my living room, I guess my treatment of AI is different.
It might be all caps in the office as I like banging my keyboard.
At home, I'll turn on dictation, maybe share a few expletives, try to redirect things.
I'm just like constantly being like, this is trash.
Why?
Why are you, why are you this way?
I used to be so polite and now my bar is so high.
I'm like, you can do better.
I believe in you.
I would say lately, I feel like I have been less upset.
Things are just working better, you know?
I've been more upset.
People are, are you getting more upset?
Yeah, because my bar is higher.
I'm like, I know you can do this.
You're not a dumb, dumb.
Come on.
That's interesting.
Yeah.
All right, yeah.
Treat it as an equal or superior.
You're surprised when it doesn't come through.
Okay. Well, Nick, this has been super fun. Where can we find you? How can we be helpful?
Yeah. And I'll find me on Twitter X. My name, Nick Bowman underscore.
If you have any questions about Codex, chat ShpT work, anything we're releasing Open AI,
feel free to reach out to me. DMs are open. Great to be with you, Claire.
Yeah, thanks for being here.
Thanks so much for watching. If you enjoyed this show, please like and subscribe here on YouTube
or even better, leave us a comment with your thoughts.
You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app.
Please consider leaving us a rating and review, which will help others find the show.
You can see all our episodes and learn more about the show at how IAIIPod.com.
See you next time.
