Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 800: Celebrating our 800th Episode: 8 AI Truths, 10 Smart AI Moves and 10 Questions You Must ask
Episode Date: June 17, 2026For 799 episodes, we’ve cut it to you straight on AI. (Episode 800 will be no different.) To celebrate our 800th episode, we took a ‘State of AI’ type look across the sector and broke down thr...ee big categories: 8 Uncomfortable AI Truths, 10 Moves Smart Teams Are Making, 10 Questions Every AI Leader Must Ask. Join us as we break it all down. Celebrating our 800th Episode: 8 AI Truths, 10 Smart AI Moves and 10 Questions You Must ask — An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Eight Uncomfortable AI Truths for LeadersAI Model Risk: Single-Model DependencyMoats Shifting to AI Operating LayersWorkflow Design vs. Prompting in AIStatic Business Artifacts Becoming AI DebtWorkflow Automation and AI Task ReshapingOpen Weight Models as AI ResilienceAgentic AI: Operational vs. Theoretical RisksTen Smart AI Team Moves ExplainedPortable AI Stack Building StrategiesWorkflow Mapping Before AI Tool AdoptionBuilding Role-Specific AI Skill PacksAutomating Reports and Briefings with AIConverting Static Files to Living ArtifactsAI Agent Deployment: Read-Only Mode TestingRouting Work by Model Value and CostMeasuring AI Output Quality, Not ActivityRed Teaming AI Workflows Beyond ModelsCapturing AI Decision Reasoning for BusinessTen Critical AI Questions for LeadersTimestamps:00:00 Episode 800: 8 AI Truths, Leader Moves05:40 Rising Trend of AI Super Apps07:37 AI series and future capabilities13:01 Challenges of AI Implementation14:33 Impact of AI on Work Processes19:44 Microsoft's open-source AI strategy23:24 Avoiding shiny new AI tools25:44 Using marketing skill packs29:38 Criticism of human-in-the-loop systems33:44 Managing AI and Cost Efficiency35:50 Evaluating AI content effectiveness40:15 Key questions for leaders42:35 Challenges in AI workflow management44:52 Evaluating AI costs and outputs48:20 The rise of AI capabilitiesKeywords: AI truths, AI workflow automation, generative AI, artificial intelligence, AI model risk, single model strategy, business continuity, model export control, government AI policy, operating layer, AI super apps, codex, unified memory, large language models, agentic AI, workflow design, AI prompting, scheduled tasks, AI skills, context engineering, static business artifacts, organizational debt, AI native deliverables, autonomous AI, AI job redesignSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist.
Transcript
Discussion (0)
This is the Everyday AI show, the Everyday Podcast where we simplify AI and bring its power to your fingertips.
Listen daily for practical advice to boost your career, business, and everyday life.
Well, we made it to 800 episodes and guess what?
We're still here and so is artificial intelligence.
Turns out this AI thing really matters.
So much so we're still doing the Everyday AI podcast three and a happy.
years after we started. But here's the other realities. No, there is no AI bubble. Yes,
AI is smarter than all of us. And no, AI has not taken 50% of white color jobs,
regardless of what some CEO that's trying to drum up IPO hype tries to sell you. At everyday
AI, I've spent the last 799 episodes trying to break down complex theories into digestible
daily takeaways with my random rumblings sprinkled in, obviously.
So for today's 800th episode, we're going to continue the no-nonsense trend as we give
you eight uncomfortable AI truths, 10 moves smart leaders are making, and 10 questions
every leader must ask. See, it's 8 times 10 times 10. I'm not going to give you 800 random
AI facts or something like that like I used to do when we were at episode 100.
So and also sometimes these hundred-ish episodes, right, 500, 600, 600, sometimes they go kind of long.
I'm going to try my hardest to make this one a shorter value-packed episode.
So let's get straight into it.
And if you are new here, well, welcome to Everyday AI.
My name's Jordan Wilson.
We do this thing, well, every day.
It's your daily, unedited, unscripted, live stream, podcast and free daily newsletter to help everyday business leaders,
not just keep up with what's happening in the world of AI, but how we can make sense of it
to grow our companies and careers.
So it starts here, but make sure to go to our website at your EverydayAI.com.
We're going to be recapping the highlights from this show, as well as all of the other AI
news.
That's important that you need to know to get ahead.
All right, let's get into it.
And if you did miss our 700 episode, we're taking a kind of similar take.
But if you listen to this episode and you're like, wow, this was really helpful.
I think you'll like our 700th as well, although some things might be a little bit out of date
now, even though it's only 100 episodes ago, that's how quickly AI moves, but you might want to
go listen to it. So we went over seven ways AI is reshaping how we work, 10 AI workflows that
actually deliver ROI and 10 AI skills every professional needs in 2026. But let's start off this
round with our eight uncomfortable AI truths. So number one, single model AI strategies are now
continuity risks. And I mean, the timing of this.
is obvious, right? We saw enthropic roll out a very capable and much hyped model in their
Fable 5, which is of the Mythos 5 family, right? It's Mythos 5 without guardrails. And I saw a lot of
chatter online, you know, companies saying like, oh, you know, we're goodbye, chat GPT, goodbye, Gemini,
goodbye co-pilot and trying to move all of their operations inside this new model, because it is a
step change or I could say it was or kind of is because it's currently no one in the world can use it.
So that's why this is a single model strategy is a risk because we saw the U.S. exports control
action on Fable 5.
So for companies that rushed to try to move all of their operations into Fable 5 from Anthropic,
number one, okay, Scroogey McScruge duck writing blank checks, how can you afford that?
it's a very expensive model.
But number two, I think that shows where everything is headed right now,
because you have to be thinking across multiple models.
And I don't know if this kind of U.S. government export action could be a common thing.
You know, the only thing with the new president's new executive order is a 30-day review period.
So this kind of export control action was not part of that initial, you know, executive order.
But we've seen that anthropic in the U.S. government.
have kind of had some ongoing beef for the last couple of months.
So this isn't shade at Anthropic.
This is just the reality of if you can trust your entire organization's AI strategy
on a single model that the federal government literally at 5 p.m. on a Friday can click,
nope, and then the entire world loses access.
Because if one model breaks your workflow, I mean, procurement misses business continuity planning.
Right.
I did a training for a company that was using clawed models and they were trying to set up all of these complex workflows.
And one thing I talked about them after the session I did is what's your backup plan?
You have to be building these things modularly because if you spend so much time, you know, especially weaving some core part of your business's operations into a single model.
You could be screwed come Friday at 5.01 p.m. when said model is done and gone, whether that's last for,
a couple of days, a couple of weeks, or indefinitely.
All right.
So that is the uncomfortable AI truth number one.
Truth number two, the moat has moved from models to operating layers.
And that kind of goes with uncomfortable truth number one.
And we actually did have a start here series on this one.
I believe it was trying to do my math.
I think it was yesterday, right?
When we talked about AI super apps, kind of being the new wave or the new trend on where work
happens.
But, I mean, Codex, you can't overla.
looked at with now five million weekly users.
So yes, in codex at least, you are still using GPT models, but you are using models, multiple
of them in the same conversation with a unified memory, being able to use a browser,
a file browser, an internet browser, being able to control your computer, being able to use your
browser, all of those things.
That's why I think the models are becoming less and less important as are capabilities
that we get through large language models become more and more agentic, right?
As they start using browsers and using computers and using terminals and using files
in the same way that humans do, I do think that there is maybe less of an overall emphasis
away from a single model.
And that's why I do think when we talk about moats, when we're, you know, which I think
for, you know, business leaders that are working on the front end AI strategy before you, you
know, say, hey, we're for this department, we're going to go kick off a one year program
with this big AI company.
You have to understand what their moat is, right?
Like, are they developing a harness that is going to be able to fully use any of these tools?
Is it maybe better to look at a more open harness?
And then you can plug and play different frontier models at will.
So I think that's a really important thing.
You know, talking about things like plugins that bundle all of these different skills
and apps together from Codex, I think is another important thing.
Also, as we talk about context, connectors, all of the,
those things, all of these things on the outside, the second layer, if you think of the model as
the base layer, all of the things surrounding them, I think are going to become more and more
important. I'll probably do a maybe my final start here series. And if you've missed that,
you know, make sure you go to start here series.com. Go listen to that. It's a, you know, an ongoing
series. I think we're going to cap it at about 30 episodes because that's probably enough. But I think
it's been great for, you know, beginners to AI leaders to tackle all of these ongoing changes.
But I do think that there's maybe going to become a certain capability gap,
not where humans are, you know, not using the full capabilities of models,
but actually where we, well, for most of our day-to-day work, at least as it is today,
our work will not require 95% of what these models are capable of.
You know, we kind of saw this with Microsoft looking at DeepSeek as a potential, you know,
maybe cost alternative to running the GPT and Claude models.
And at first, I'm like, well, that doesn't make sense, at least now.
But it might make sense in six to nine months because I think at least until the work
that knowledge workers are required to do changes drastically, right?
For the most part, I think in six to nine months, you know, some of the more basic,
I'm not saying, you know, deep seek, right?
I'm not saying that model, but I do think some of the more basic models, or even a company's
second or third most powerful model, maybe for a period of a couple of years.
might be sufficient enough to do the majority of all this work,
which is why I think, you know, yes,
when we were talking about models,
capabilities from, you know,
2024 and 2025,
the model really matter.
But I think it matters less and less today than it did.
Truth number three,
prompting is losing to workflow design.
All right.
And if you haven't noticed this,
again,
this goes along with the push toward the super app and the harnessing
and all of those things.
But, you know,
even Google, right?
Google is pushing,
Agents, scheduled tasks, subagents, orchestration, right?
If you are still spending a majority of your time or your team is still spending the majority
of your time, prompting large language models versus reading and acting on their outputs,
you are behind.
Let me just say that very bluntly, right?
If you are still going in and typing things and, you know, your team has a prompt library,
it's not like that anymore, right?
We have skills.
That's a universal language that all large language models speak.
For the most part, they all have scheduled task or automations.
So for the most part, humans should be working more on the front end,
you know, making sure that your context stays up to date,
making sure that, you know, you are making sure whether you're working with
markdown files, skills across your organization,
making sure those stay up to date with your most up to date context.
But then from there, you should really be spending, right?
Humans should be spending more of their time, taste making, right?
If you have a project that you would normally go in and prompt a large language model for,
right and that project happens every Monday afternoon instead of you prompting every single day
at Monday afternoon well you should be looking at the deliverables that five different agents
delivered and saying which one is best and then Tuesday after that deliverable is due you
pour in with new updated context and make sure that hey next next Monday you know those five
you know reports that five different agents are going to take a swing at they actually
have more better context and directions truth number four
static business artifacts are becoming organizational debt.
And I didn't say have become on this one.
Okay, I'm saying they are becoming.
Let me give you the quickest example.
Again, did in a recent episode on this,
actually when we talked about codex sites.
So make sure to go back and listen to that one.
But as the harnesses become more popular, right?
As things like codex cursor, right,
Claude desktop, I think has some catching up to do.
We'll see if Google continues to invest in the anti-gravity 2.0 app or not.
But as these harnesses become, number one, more capable, but more team friendly,
I think the concept of emailing files back and forth or even dropping, you know,
static files in a shared drive, I think it's pretty soon going to become archaic.
And I think our artifacts first need to become AI native.
Yes, we need to figure out what that means for history and versioning and all of those things.
But I think Codex sites is maybe the first iteration.
Maybe it's not the final or the best, right?
But I think that's a good example of the difference between how a static business artifact can
actually become organizational debt.
And I didn't think it at first, but I thought back throughout the course of my career,
How much time that I spent either communicating about different file versions, trying to find old versions, you know, corrupted files, right?
All these things, just tackling, communicating, finding, organizing around different files, sharing them.
You know, you're waiting for approval on all these files because, oh, it turns out they had V4 and they actually needed V4 final, right?
All these things.
I think we have to look at what does an AI native deliverable look like?
Maybe it's a codec site.
Maybe it's a, you know, clawed live artifact.
I don't know.
But that is truth number four is we have to start thinking about AI native deliverables.
Truth number five, AI can create more work if leaders don't remove old work.
All right.
Let me tell you what that means.
AI often just adds more and more, right?
If you're looking at text, walls of text, that doesn't necessarily mean that your organization is winning
back time, right? When we talk about R.O.Y, if all you're doing is continually experimenting with
different AI, with different large language models, or if your employees, if your team members are
just banking their save time, which I would argue the majority of companies that saw, you know,
AI gains in theory, but not on the books in 2024 and 2025. Yeah, that's just your, you know,
employees having, yeah, I do all the work for them doing the same amount of work that they were doing
pre-AI and then saying, well, look, right.
But on the flip side of that, AI, if you don't properly haul it in, can actually create
way more work if you don't redesign workflows to be AI native.
So if your job description hasn't changed, if your SOPs haven't changed, especially since
late 20, late 2024, early 2025, when models became more agentic, you have to do so because
all you're technically doing is you are just doing.
doing a disservice actually now that I think about it a little bit more right if you are still
producing artifacts for clients industry reports whatever the the novel work is that you're
creating business value within your organization if you're still doing it the old way let's just
say the 22 way with 2026 tools that's bad right yes you are still creating more work but one of the
reasons you're creating more work is because you are creating more work is because you are
creating all of these processes that in theory are not going to be transferable once your company,
your organization, or your artifact becomes AI native or with whatever your, whether it's
for a client, an internal team member, et cetera, right?
What about when their demands change?
And I think the consumer demand will begin to change, right?
We're getting overslopified with all this AI slot.
But I think that demands are going to change.
So if you are just trying to use as much AI as possible on old processes, old job descriptions, companies and in organizations that are set up in a pre-AI way, I think ultimately you are just setting yourself up for failure, even if it seems like success in the short term.
So real ROI starts when the AI actually goes in there and removes steps, removes meetings, and reworks what you're actually working on.
Truth six, longer AI task completion will reshape jobs.
All right.
So good example here.
Infropics says autonomous task length has doubled roughly every four months.
And as an example, infropic said that Claude authored over 80% of Anthropics merge code in 2026.
So what does this mean?
I'm not sure yet, right?
This is one of those things that still kind of keeps me up at night.
And I think we, it's, it's going to take even people on the edge of AI.
It's going to take them a while to figure out what does the future job look like, right?
Not like, oh, what happens when we have, you know, embodied AI and human, you know, human
AI robots doing all these things, even before that, right?
What happens when autonomous AI harnesses are doing all of our normal deliverables that we
were doing three years ago?
What happens then?
But I think that managers must start to read.
design work around delegation, supervision, review, and exceptions, right?
We have to start redesigning old work processes because as these models can work,
whether we're talking about in loops and managed and controlled loops or just working
truly autonomously toward a goal, right?
As they start doing things that would normally take many, many hours or many, many
dates, right?
We don't know how jobs are going to look.
And I'm not talking about job disruption.
I'm not talking about job creation.
I'm talking about the jobs that will still be here,
which I think is the majority of jobs,
how are they going to look differently?
But as you have these models now that can run
and do these long tasks,
and they can work for 8, 10, 12, 15, 18 hours
on a single type of project that would normally take a team,
what does that mean for the future jobs?
I don't know yet.
but it will reshape jobs.
Truth seven, open weight models are becoming resilience insurance.
Right.
So I already talked as an example about the recent Microsoft news that said that they're
looking at DeepSeek as a potential rollback or different offering.
Right.
So they're looking at fine-tuning a version of DeepSeek and maybe using it internally,
offering it up to customers within co-pilot.
We'll see right now it's just reporting that's still shaking itself out.
but the reality is open source and open weight models will soon become at least a resiliency insurance, right?
This is something I've even started to experiment with on my own.
I will say this, though, open weight models, they're great.
If you think that you can slap, right, a general open weight model and, you know, have it run, you know, locally on your machine like,
why would I ever, you know, use Fable 5 or why would I ever use GPD 5, 6?
I can have an open weight model running on my computer.
No, that makes you look foolish.
FYI, I'm just saying that.
In that case, open weight models will still, you know, open weight or open source models will
still be multiple years behind in terms of what a consumer can go and run on hardware.
So unless you're spending $30,000 on hardware, it is much more economically feasible
and just responsible for you to be paying $200 a month
and just use the frontier model, right?
But open weight models are becoming a great backup plan
for specific use cases.
So as an example, and this one's kind of recent as well,
GLM 52 Max is now technically the number one model in the world
for front end coding because the only model that's ahead of it,
Fable 5 is not technically available.
So it is the number one available model.
Think of that, an open source model.
All right.
Again, you can't really run this locally unless you have a cluster of supercomputers.
But again, thinking about how this shapes the future of work and how this impacts decisions that your company makes.
I mean, you have to look at what Microsoft is doing there.
If Microsoft is thinking that, hey, we might be able to, whether it's internally or for co-pilot customers, if we put enough human effort into this,
we could fine tune an open source model and run it at production for even half of internal or
external purposes.
That's big.
And I think that business leaders need to start looking at that.
For most companies, it's not there yet.
But you have to understand that open fallbacks matter because vendor changes are going to happen.
Policy prices.
Access, I think through the rest of 2026, we'll see what shakes out between the Anthropic and
White House.
But you have to start thinking of these things.
and you need to start routing workloads by sensitivity, volume, compliance, cost, and quality threshold.
Truth number eight, agents make AI risk operational, not theoretical.
All right.
And I talked about this a little bit more in our 2026 AI prediction and roadmap series,
but as the default, as the default models become more agendic, they're getting read,
right actions.
Let's even take out the fact of running these, you know, desktop super apps that can access
every single file in your computer, they can read, right, use the terminal, all those things.
Even look inside something like chat chpityt, even look inside something like, I don't know,
GROC, perplexity computer.
Even the quote unquote, non-technical AI chatbots are getting more agentic power, right?
Very small example.
ChatGPT can send emails in it.
Right, it doesn't sound like a big thing, but chat GPT has always had the
ability to draft emails and to read your emails and to pull in all this context.
But just as of a week ago, it can actually send an email.
Do you know how much risk there is when you give any AI model the ability to send an email?
Right.
There's great promise, but there's also a great downside if you don't have the right expert
driven loops because every agent needs an owner, limits, approval, logs, and rollback plan.
All right.
So those are our eight uncomfortable truths.
Now let's get into the 10 AI moves that's smart.
AI teams are making.
And yeah,
we're going to pick up the pace here.
All right,
move number one is they build portable AI stacks
before disruption hits.
So I'm talking about primary backup,
cheap private and manual fallback packs.
Pats,
these all have to be documented, right?
What happens when AI works?
You have to be able to answer that first.
And then what happens when that AI that works fails.
All right.
So smart teams are making those moves already.
People always are so quick to say once you find a win and you're like, oh my gosh, this is going to save us 40% of staffing costs.
So we're just going to go ahead and cut 10% of staff.
Call it a win, right?
No.
Use those humans and start building these portable AI stacks because disruption will hit, whether it's internal, external, from a policy standpoint, whatever you need to be kind of creating these redundancies.
Move number two, they map workflows before buying.
more tools.
All right.
I think so many organizations that I talk to, they find one success.
And then they roll that success out across different teams, right?
Which is great.
That's a great way to go and do things.
But then what happens is then they just say, okay, then what's the next good tool that
we could do this same process, right?
Instead of replicating that workflow mapping across the entire organization.
Instead, right?
And I get it.
It's hard because companies always want to be chasing the shiniest,
latest,
greatest AI tool and model.
Let me,
let me be very direct here.
Even if you,
I'd say,
if your organization,
unless you have a team of 100 people,
whose only job it is is AI deployment in your organization,
which is no one,
they have no outside responsibilities or deliverables,
unless you have 100 people that that's their only job,
you and your company will get much more,
much more value.
If you literally stop,
don't try GPD 56,
don't try Fable 51 or whenever it comes out.
Don't try Gemini 35 Pro.
You will find more business value.
If you literally stop trying the next best model.
And instead, look at your next or your last big win
and then use the findings from that and map that out to everything within your organization.
That's very hard to do.
Instead, normally you just go through this concept of we're just going to map or we're just going to try to replicate our big win.
And oh, by the time we, you know, rolled out that one big win across all organizations, there's a new model.
So we're just going to go chase that.
It's hard not to, but you're going to get much more ROI by just mapping entire workflows before you just jump on to the next big model.
Move three. Smart teams are building role-specific AI packs.
Luckily, Anthropica was really ahead of the game with these skills.
And surprisingly, it did take Open AI a little bit longer than I thought it would to kind of jump, I won't say on the bandwagon, but to follow suit.
Skills are extremely powerful, right?
this kind of takes away a lot of the grunt work of prompting, right?
So skills, I mean, you could make the argument.
Skills are just a bunch of text prompts stacked together, right?
And they sometimes do some technical things behind the scenes, right?
But this is the basics of, well, prompt engineering, right?
The skills are essentially this prepackaged markdown files that tell a large language model
in a very specific way, what to do, what not to do, but around different skill sets.
So as an example, there's great skill packs for different type of marketing.
So if you're a marketer, you should probably at least start with a marketing skill,
even before you go in and prompt something or customize a workflow to work for how you need it.
But this is huge.
And, you know, skills are obviously transferable across different systems.
Codex just rolled out a bunch of kind of, you know, plugins, I guess is what they're saying
because their version is more using skills.
plus apps, which is great, right?
But every organization needs to be experimenting with skills.
And you need to find a good skill that's already pre-built for your organization.
There's great open resources for skills because you can plug and play them into anywhere, right?
You need to start with a skill that is highly rated, bet it across your organization, test it in a
sandbox way, make sure it works, customize it, and go on to.
to the next. All right, move number four, smart teams are automating recurring reports and briefings.
My gosh, if you are still having 20 people in this boring three hour meeting every single week
and someone's still typing up to do's because no one has access to the AI note taker,
oh, or can't use it for this meeting. And then you're preparing a deck for the meeting and a deck
after the meeting. My gosh, no, stop. You have to start automating all of those things because,
number one, the majority of the people in those meetings don't want to be in there anyways,
right? The majority of people reading a 30-page slide deck, they only care about one
little slide anyways, right? That's the reality. I think you have to start pulling in all of
these things that are spending, that you are spending too much human capital on in a non-AI native
And you have to start blowing up those old processes.
All right.
Move number five.
Already kind of talked about this, but it's worth talking about again, converting static files into living artifacts.
So, sorry, spreadsheets, PDF, PowerPoints, all these things.
You have to, in a good one too is skill, skill libraries, right?
Your company, your organization needs to have living AI native documents.
old static files are going to slowly die.
Probably not by this year, right?
It'll probably still be another year or two before they actually die.
And until we see the rise, right, even look in the AI community,
how popular things like Markdown files are now, right?
We are going to get, whether it's something like Codex sites,
I don't know what it is,
but you have to start converting your static files right now into living artifacts.
Move number six.
Smart teams are testing agents in read-only mode before giving them control.
My gosh, I can't tell you the amount of times I've seen this, right?
Whether it's, you know, hearing about it on, on, you know, conversations I don't have here on on the show, reading about it in news articles.
the process of sandboxing an agent and putting an agent into production is pivotal, right?
I think a lot of teams think, okay, well, let's sandbox the agent, let's test it against the guardrails,
make sure it's safe, make sure it doesn't drift.
You know, they go through, they check their boxes, right, that makes the compliance team happy.
And they're like, look, we're super responsible, right?
And then they do this.
You know, this is usually those, those companies that I think are a little too AI trigger happy.
And then all of a sudden, they start setting these agents out.
And when they're like, oh, don't worry, Bill from IT, these are human in the loop.
You know what?
I know I talk about this a lot, but the world would be a lot better place if no one ever came up with this human
in the loop thing because it's absolutely garbage.
Because in, I tell you, 90% of the time, the human that's in the loop, number one,
is the wrong person.
Number two, they're not driving or steering.
They just look in there when it's more of,
they're a crash scene investigator
instead of an FAA flight operator.
Most of the time, a human in the loop is only,
quote unquote, involved when something goes wrong.
So you have to test agents in read only mode
before giving them control.
Because as the models drive,
these Asians become more sophisticated. So too does your process of the human or humans you have overseeing
them. AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company
might lag behind while AI native competitors leap ahead. But you don't have 10 hours a day to
understand it all. That's what I do for you. But after 700 plus episodes of everyday AI, the most common
questions I get is, where do I start? That's why we created the Start Here series, an ongoing
podcast series of more than a dozen episodes you can listen to in order. It covers the AI
basics for beginners and sharpens the skills of AI champions pushing their companies forward.
In the ongoing series, we explain complex trends in simple language that you can turn into action.
There's three ways to jump in. Number one, go scroll back to the first one in episode 691.
Number two, tap the link in.
your show notes at any time for the start here series or you can just go to start here series.com
which also gives you free access to our inner circle community where you can connect with
other business leaders doing the same. The start here series will slow down the pace of AI so
you can get ahead. Move seven smart AI guys are making. They route work by value, risk,
and cost. So we had an episode recently in the start here
series talking about going from token maxing to token efficiency.
And this part's huge.
You should probably not, if I'm being honest, you should probably not be using Fable 5 for 99% of your
tasks.
You shouldn't be using GPT 55x high for 99% of your tasks, right?
It's crazy the amount of people that are just using the best model for everything,
a blanket.
That's not the way to do it.
Right.
I think most companies take these.
I won't say their habits, but maybe best practices.
Because for the most part, in 2023 through 2025,
we were all living in this subsidized AI nirvana, right?
Where it's like, oh my gosh, I can just use the best model for anything regardless.
All right.
But times are changing.
Times are changing.
Token maxing has led to token efficiency.
Right.
We saw all these stories, all these huge organizations, you know, hit their AI
spend limit or their token budget within three or four months.
And now they have to change.
And the way you change that is by routing the type of work to the proper model or models that it needs.
That's why I think that model routers, mixture of model setups.
I think there's going to be some great AI startups.
I've already seen a few launch in the last few months that do exactly this.
Right.
So you are not wasting as an example, a billion tokens a day.
It's not crazy for, you know, I very easily go through 500 million to a billion tokens a day sometimes.
Me, small business owner, right?
Luckily, I'm still heavily relying on the subsidized, you know, Open AI $200, chat chitp T pro claim.
But subsidies are all going to eventually fall or diminish.
So you have to start making these calls appropriately.
And similarly, you do have these now very capable models like GLM 52, the Kimi models, right?
We're not saying use them for every task.
But instead of using, you know, one model for 50 different tasks within your organization,
you might be using, you know, five different models for eight different tasks.
or something like that or 15 different models for three different tasks each.
So you have to start routing them by value, risk, and cost now.
Because if not, you are going to be for a rude awakening.
Because as an example, right, if Anthropics like, yeah, you know,
here's enough Fable 5.
We'll see if the June 22nd date gets extended or not.
But, you know, it's very likely.
Anthropic might go down this path of, you know,
releasing these models just to all subscribers.
people are going to be reworking their workflows,
then they get pulled and you have to start paying API costs.
So, you know, $200 becomes $14,000, give or take.
So you can't keep doing that.
You have to work smarter.
Move number eight, I'm seeing smart teams make.
Measure accepted output, not AI activity.
Right.
Lines of code is not return on investment, right?
tokens burned doesn't mean a thing.
Are your deliverables getting better?
Are they moving the needle?
Are your blog posts bringing in more human visitors?
All of these things that I think we rushed to start using AI for.
Oh my gosh.
Look at Claw design is so great now.
I can go get out 5x the number of proposals.
Okay, number one, did you?
And number two, are you bringing in 5x the amount of business?
or are you just forcing AI slop down people's throats who are already being avalanche upon
with just, you know, more and more content?
Just because we can create more and more and more text, more and more photos, more and more and
more videos, more and more audio, doesn't mean that we should be.
I think human taste really matters here.
So you have to look at accepted outputs, right?
Which of those, you know, to pull from my random example, random rambling from earlier,
You know, when my agent came back and gave me five versions of a certain report, which one was number one accepted by the expert driving that agent loop?
And number two, which one ultimately won over my peers, one over the internal stakeholders, one over the external stakeholders as well, which one got us more business.
So you have to start measuring those things, not activity.
Activity is not movement.
Move number nine, they red team the workflow, not just the model.
Okay.
Yes.
Large enterprises need to have red teams.
You absolutely have to.
Right.
We saw this with the, you know, even though we don't have the full story yet, but we saw
this with the, the Fable 5 rollout, right?
Essentially, at least according to the White House and the federal government, you know,
they said anthraxic shipped a model that wasn't safe.
So what happens if you put that into production and for whatever reason, one of your
customers, one of your clients who maybe relies on your services or your outputs that you use
somehow stumbles upon a piece of code that wasn't safe. And you're like, well, well, this model
did it. Well, okay. Yes, we assume that labs like Anthropic and Microsoft and Open AI and Google
have the best red teamers in the world. It's those, those engineers that go in and make sure a new
model is safe. And they try to break it internally before releasing it.
That's kind of like what red team is.
That's red teaming for dummies 101, right?
But you have to have an internal team that does that as well, not just the model,
but the harness.
So in the same way, you know, if you say, hey, we're going to roll codex out to 10,000
people, we're going to roll out Claudecode.
We're going to roll out cursor, right?
I think before when you would roll these things out, it was usually developers using
them the right code.
But as we start producing non-technical work documents, as that becomes the norm,
how are you red teaming that internally?
That's what you have to do.
Agent risk includes the inputs, the tools, permissions, actions, and the integrations.
So your teams must test prompt injection, what happens when you use bad data, tool misuse,
and also runaway costs.
And Move 10, smart teams are capturing decision reasoning in systems they own.
All right, models are rented, but company reasoning history can become durable memory.
I've been talking about this for years.
One of the most important things your companies can do is to start collecting first company reasoning data, right?
FCRD.
I've used the acronym once or twice before.
So the simplest way to think about this, the old school transformer models needed structured data.
Right.
It gobbles up everything on the internet.
Models that reason, models that think.
That's today's models.
Yes, they're still good.
Yes, they still need structured data.
Good data.
Your company's good data.
It also needs how your company thinks and how your company makes decisions.
Because the smartest thinking model in the world is only going to be as helpful as the nuanced information that you give it.
What do I mean by that?
In the same way that you would teach and invest in a human over years to,
understand why you made this decision. Hey, here's why I rejected this proposal and here's why
I crossed off this line item. Those things aren't always found in a spreadsheet. That's not always
structured data, right? These are the things that smart teams are capturing, putting them in documents
and training models to work through. All right. So now let's get to 10 questions every leader must ask.
All right.
Question number one, what breaks if our primary model disappears tomorrow?
Yeah, maybe a little recency bias in putting this all together.
But I think the whole situation with the Anthropic Fable 5 and the U.S.
government really sheds a much needed light on this AI race that has become a frenzy.
What happens if that model is gone?
You have to be able to identify the workflows.
customers, employees and systems that are dependent on that one model.
And what happens if it goes away?
All right?
I'm not going to answer these questions.
These questions are for you and your team.
Play this episode for them.
Share this episode with them.
You need to be talking about these.
Question two.
Are we solving a task workflow or operating model problem?
Because tasks need prompts, workflows need systems,
and operating models need redesign.
So if you just have an AI summarizing meetings,
That's a lot different from updating projects and owners because you need to be not just solving the task, but the bigger picture as well.
Question three, what proprietary context makes RAI better?
And this goes back to that first company reasoning data.
Generic models just produce generic outputs.
And it's the same thing, even though these models are becoming smarter and smarter.
If all you're doing is feeding your data to a world's smartest model in the correct way,
that's not going to help.
I mean, it's going to help a little bit.
But that's not going to be what ultimately separates you because your competitors are going to be doing the exact same thing.
They're going to be putting their best structured data with the best reasoning models or the best harness.
Right.
If competitors can match your outputs instantly, then the model was never your moat.
Question number four, what can AI read, write, change, or send?
You need to know all of those things.
And there are very few.
Yeah, let me say this honestly.
There are very few AI leaders out there that are in charge of AI rollout across their
organization that can accurately answer these questions.
And it's not their fault.
It's because this space moves too quickly.
So unless you have a team of people reading every single change log, right?
Or unless your whole team listens and reads the,
You know, if you listen to our newsletter every day, if you, or sorry, if you listen to the podcast
every day, read the newsletter.
You're probably ahead of most teams.
But still, even at that point, you have to know every single model, every single harness,
every single connector, which one has read, right, access, which one requires approval,
what permissions, right?
Which permissions are needed, you know, when using which AI super app, what can do what?
You have to know those things.
Question five.
who owns every AI workflow?
What one person?
Asian crash is going to happen at your organization, in your industry, your competitors,
etc.
You have to know now.
Number one, you have to map the workflows.
But number two, you have to know who owns each workflow.
And that person has to be able to answer those questions like question four.
What can AI read, write, change, or send?
And you have to know who owns that workflow?
Question number six, what is the cost?
ceiling. The ceiling is important. We saw some stories. We don't know if they're actually true or not.
There's one story that came out that one company, quote unquote, accidentally ran up their monthly
AI bill to $500 million because they didn't put a spend cap on cloth. Was that true? I don't know.
But what is the cost ceiling? You have to have those measures in place. Not just how do we monitor
what the cost is, but what cost should you be paid?
Right. Should you be paying, I don't know, a million dollars a month to summarize emails?
Maybe, I don't know, if your organization, if you have hundreds of lawyers reading very important emails, maybe, right?
But you have to know what the cost ceiling is because agetic AI can spend money invisibly while appearing to be productive, right?
That's why you need to look at things like artificial analysis is cost per intelligence, right?
Because a lot of times people just look at a single benchmark and be like, oh, my gosh, this fable five thing is great.
Well, what if it costs three times as much as something like GBT55?
If you get the same exact output, but it costs three times as much.
What's your cost ceiling?
Question seven, what counts as a useful output?
My gosh, I've seen people just really redefine and twist what a certain end artifact
should be because an AI created it.
Generated content means nothing until the business accepts it, uses it, and you can prove an
ROI on it.
So you need to say, what counts as a useful output?
Where is your AI spend going?
And is it producing an actual useful output?
Question 8.
What happens when the AI is wrong?
Okay.
Not when it goes out, but what happens when it's wrong?
Do you have the right expert-driven loop that can identify?
when an AI output is wrong.
Hallucinations, yeah, they are not really a problem anymore,
but not everyone is using the right model.
So they obviously still exist.
If you do the context engineering thing correctly,
if you have expert driven loops,
if you're using the right model for the right purpose
and you've gone through the guard railing and the scoping and all those,
let me just say it.
Hallucinations are extremely rare if you do all those things,
but most companies do not do those things.
So you have to ask the question,
what happens when the AI is wrong?
Question number nine.
What evidence could we show six months later?
Six months later.
Yeah, leaders, you need records of models, data, prompts, actions, approvals, and changes.
Some, right, at some enterprise level, like if you're using chat GPT enterprise,
Claude Enterprise, right, co-pilot, they do have some useful metrics, I guess, right?
I like Microsoft's approach with their intra ID.
and a lot of the other companies are starting to adopt something similar as well.
But every single piece of value that is created via an augmented relationship between a human and
AI, it needs to be observable, it needs to be traceable, it needs to be auditedable.
You need to be able to go back six months later and say, oh, my gosh, we ended up winning
this huge RFP, right?
And we had this, you know, RFP agent cooking it all up.
let's go back.
How did we do that?
Who set that up?
What models were we using?
What was the workflow design process?
And why did we get wildly different results from all these other RFPs when we were using the same process?
You have to not only be able to go back and say, hey, this worked.
Let's go back and make sure that we apply some of our successes across the board.
But you have to also say, why did it work in some scenarios and not others?
And then the last question, question 10, what should we stop doing if this works?
And this is the big one.
This is the thing about shifting human agency.
And maybe all, this is good to end this episode, episode 800 with a bigger question,
something maybe thought provoking because it's something I'm thinking about all the time.
What happens in six months?
What happens in nine months?
What happens in one year?
when these agents are even more and more capable.
Go back and look at where we were a year ago and look at where we are now.
It's scary.
There is no wall.
It is a ramp going straight up.
If capabilities continue, if the harnesses continue,
if these agents that run on your local machine 24-7 and can produce artifacts in the same way that humans can,
and they are economically as valuable or more valuable than expert,
humans, if all these things continue.
What should we stop doing when this continues to work?
Should we stop spending time as humans on the things that we've spent the majority of
our career on?
Should we stop?
Let's just say, let's just say you love creating PowerPoint presentations, right?
Let's just use that as a small example.
You love getting the team in the room and, you know, brainstorming and, you know, pulling in all
these documents and creating a bangor presentation and you present it to the client.
and they love it.
What's,
that's your job,
right?
If that's your job and you love it,
but it's shown that an agent is better,
what do you stop doing?
I don't have the answer,
but maybe on episode 900,
we will.
But that's a wrap for today,
celebrating our 800th episode,
eight uncomfortable truths,
10 moves smart teams are making,
and 10 questions.
Every AI leader must ask.
I hope this was helpful.
You know what?
I'll say one thing that is helpful for me, hearing from all of you.
So if you didn't make it this far 50 minutes into this episode, reach out.
Let me know what would you like to see in episode 801?
What would you like to see down the road?
I know we normally on Wednesdays do our AI working Wednesday.
So thanks for, you know, letting us throw a little curveball in the rotation.
We'll be back next Wednesday with our normal working Wednesday series.
But thank you, whether you've been around for one episode or all.
800. I hope we can keep going to 900,000 and beyond. So if this was helpful, do me a favor.
Tell someone about it. I can only make it to 900,000 and more by giving you unbiased,
straight AI info when you tell people. So please, number one, subscribe to the podcast.
If you haven't already on Spotify or Apple Podcasts, then make sure you go to your everyday
AI.com. Sign up for the free daily newsletter. Do me a favor. Drop me a line today. Drop me
an email. I would love to hear, you know, whether if something's been helpful along the way,
how long have you been listening? What would you like to see us do differently in the next 800
episodes? I work for you. So thank you for tuning in. Hope to see you back tomorrow and every day for
more Everyday AI. Thanks y'all. And that's a wrap for today's edition of Everyday AI. Thanks for
joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going.
For a little more AI magic, visit Your EverydayAI.com and sign up to our daily newsletter so you don't get left behind.
Go break some barriers and we'll see you next time.
