Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 800: Celebrating our 800th Episode: 8 AI Truths, 10 Smart AI Moves and 10 Questions You Must ask

Episode Date: June 17, 2026

For 799 episodes, we’ve cut it to you straight on AI. (Episode 800 will be no different.) To celebrate our 800th episode, we took a ‘State of AI’ type look across the sector and broke down thr...ee big categories: 8 Uncomfortable AI Truths, 10 Moves Smart Teams Are Making, 10 Questions Every AI Leader Must Ask. Join us as we break it all down. Celebrating our 800th Episode: 8 AI Truths, 10 Smart AI Moves and 10 Questions You Must ask — An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Eight Uncomfortable AI Truths for LeadersAI Model Risk: Single-Model DependencyMoats Shifting to AI Operating LayersWorkflow Design vs. Prompting in AIStatic Business Artifacts Becoming AI DebtWorkflow Automation and AI Task ReshapingOpen Weight Models as AI ResilienceAgentic AI: Operational vs. Theoretical RisksTen Smart AI Team Moves ExplainedPortable AI Stack Building StrategiesWorkflow Mapping Before AI Tool AdoptionBuilding Role-Specific AI Skill PacksAutomating Reports and Briefings with AIConverting Static Files to Living ArtifactsAI Agent Deployment: Read-Only Mode TestingRouting Work by Model Value and CostMeasuring AI Output Quality, Not ActivityRed Teaming AI Workflows Beyond ModelsCapturing AI Decision Reasoning for BusinessTen Critical AI Questions for LeadersTimestamps:00:00 Episode 800: 8 AI Truths, Leader Moves05:40 Rising Trend of AI Super Apps07:37 AI series and future capabilities13:01 Challenges of AI Implementation14:33 Impact of AI on Work Processes19:44 Microsoft's open-source AI strategy23:24 Avoiding shiny new AI tools25:44 Using marketing skill packs29:38 Criticism of human-in-the-loop systems33:44 Managing AI and Cost Efficiency35:50 Evaluating AI content effectiveness40:15 Key questions for leaders42:35 Challenges in AI workflow management44:52 Evaluating AI costs and outputs48:20 The rise of AI capabilitiesKeywords: AI truths, AI workflow automation, generative AI, artificial intelligence, AI model risk, single model strategy, business continuity, model export control, government AI policy, operating layer, AI super apps, codex, unified memory, large language models, agentic AI, workflow design, AI prompting, scheduled tasks, AI skills, context engineering, static business artifacts, organizational debt, AI native deliverables, autonomous AI, AI job redesignSend Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info) Start Here ▶️Not sure where to start when it comes to AI? Start with our Start Here Series. You can listen to the first drop -- Episode 691 -- or get free access to our Inner Cricle community and all episodes: StartHereSeries.com Also, here's a link to the entire series on a Spotify playlist. 

Transcript
Discussion (0)
Starting point is 00:00:00 This is the Everyday AI show, the Everyday Podcast where we simplify AI and bring its power to your fingertips. Listen daily for practical advice to boost your career, business, and everyday life. Well, we made it to 800 episodes and guess what? We're still here and so is artificial intelligence. Turns out this AI thing really matters. So much so we're still doing the Everyday AI podcast three and a happy. years after we started. But here's the other realities. No, there is no AI bubble. Yes, AI is smarter than all of us. And no, AI has not taken 50% of white color jobs,
Starting point is 00:00:46 regardless of what some CEO that's trying to drum up IPO hype tries to sell you. At everyday AI, I've spent the last 799 episodes trying to break down complex theories into digestible daily takeaways with my random rumblings sprinkled in, obviously. So for today's 800th episode, we're going to continue the no-nonsense trend as we give you eight uncomfortable AI truths, 10 moves smart leaders are making, and 10 questions every leader must ask. See, it's 8 times 10 times 10. I'm not going to give you 800 random AI facts or something like that like I used to do when we were at episode 100. So and also sometimes these hundred-ish episodes, right, 500, 600, 600, sometimes they go kind of long.
Starting point is 00:01:37 I'm going to try my hardest to make this one a shorter value-packed episode. So let's get straight into it. And if you are new here, well, welcome to Everyday AI. My name's Jordan Wilson. We do this thing, well, every day. It's your daily, unedited, unscripted, live stream, podcast and free daily newsletter to help everyday business leaders, not just keep up with what's happening in the world of AI, but how we can make sense of it to grow our companies and careers.
Starting point is 00:02:00 So it starts here, but make sure to go to our website at your EverydayAI.com. We're going to be recapping the highlights from this show, as well as all of the other AI news. That's important that you need to know to get ahead. All right, let's get into it. And if you did miss our 700 episode, we're taking a kind of similar take. But if you listen to this episode and you're like, wow, this was really helpful. I think you'll like our 700th as well, although some things might be a little bit out of date
Starting point is 00:02:27 now, even though it's only 100 episodes ago, that's how quickly AI moves, but you might want to go listen to it. So we went over seven ways AI is reshaping how we work, 10 AI workflows that actually deliver ROI and 10 AI skills every professional needs in 2026. But let's start off this round with our eight uncomfortable AI truths. So number one, single model AI strategies are now continuity risks. And I mean, the timing of this. is obvious, right? We saw enthropic roll out a very capable and much hyped model in their Fable 5, which is of the Mythos 5 family, right? It's Mythos 5 without guardrails. And I saw a lot of chatter online, you know, companies saying like, oh, you know, we're goodbye, chat GPT, goodbye, Gemini,
Starting point is 00:03:18 goodbye co-pilot and trying to move all of their operations inside this new model, because it is a step change or I could say it was or kind of is because it's currently no one in the world can use it. So that's why this is a single model strategy is a risk because we saw the U.S. exports control action on Fable 5. So for companies that rushed to try to move all of their operations into Fable 5 from Anthropic, number one, okay, Scroogey McScruge duck writing blank checks, how can you afford that? it's a very expensive model. But number two, I think that shows where everything is headed right now,
Starting point is 00:03:59 because you have to be thinking across multiple models. And I don't know if this kind of U.S. government export action could be a common thing. You know, the only thing with the new president's new executive order is a 30-day review period. So this kind of export control action was not part of that initial, you know, executive order. But we've seen that anthropic in the U.S. government. have kind of had some ongoing beef for the last couple of months. So this isn't shade at Anthropic. This is just the reality of if you can trust your entire organization's AI strategy
Starting point is 00:04:35 on a single model that the federal government literally at 5 p.m. on a Friday can click, nope, and then the entire world loses access. Because if one model breaks your workflow, I mean, procurement misses business continuity planning. Right. I did a training for a company that was using clawed models and they were trying to set up all of these complex workflows. And one thing I talked about them after the session I did is what's your backup plan? You have to be building these things modularly because if you spend so much time, you know, especially weaving some core part of your business's operations into a single model. You could be screwed come Friday at 5.01 p.m. when said model is done and gone, whether that's last for,
Starting point is 00:05:20 a couple of days, a couple of weeks, or indefinitely. All right. So that is the uncomfortable AI truth number one. Truth number two, the moat has moved from models to operating layers. And that kind of goes with uncomfortable truth number one. And we actually did have a start here series on this one. I believe it was trying to do my math. I think it was yesterday, right?
Starting point is 00:05:42 When we talked about AI super apps, kind of being the new wave or the new trend on where work happens. But, I mean, Codex, you can't overla. looked at with now five million weekly users. So yes, in codex at least, you are still using GPT models, but you are using models, multiple of them in the same conversation with a unified memory, being able to use a browser, a file browser, an internet browser, being able to control your computer, being able to use your browser, all of those things.
Starting point is 00:06:13 That's why I think the models are becoming less and less important as are capabilities that we get through large language models become more and more agentic, right? As they start using browsers and using computers and using terminals and using files in the same way that humans do, I do think that there is maybe less of an overall emphasis away from a single model. And that's why I do think when we talk about moats, when we're, you know, which I think for, you know, business leaders that are working on the front end AI strategy before you, you know, say, hey, we're for this department, we're going to go kick off a one year program
Starting point is 00:06:49 with this big AI company. You have to understand what their moat is, right? Like, are they developing a harness that is going to be able to fully use any of these tools? Is it maybe better to look at a more open harness? And then you can plug and play different frontier models at will. So I think that's a really important thing. You know, talking about things like plugins that bundle all of these different skills and apps together from Codex, I think is another important thing.
Starting point is 00:07:16 Also, as we talk about context, connectors, all of the, those things, all of these things on the outside, the second layer, if you think of the model as the base layer, all of the things surrounding them, I think are going to become more and more important. I'll probably do a maybe my final start here series. And if you've missed that, you know, make sure you go to start here series.com. Go listen to that. It's a, you know, an ongoing series. I think we're going to cap it at about 30 episodes because that's probably enough. But I think it's been great for, you know, beginners to AI leaders to tackle all of these ongoing changes. But I do think that there's maybe going to become a certain capability gap,
Starting point is 00:07:53 not where humans are, you know, not using the full capabilities of models, but actually where we, well, for most of our day-to-day work, at least as it is today, our work will not require 95% of what these models are capable of. You know, we kind of saw this with Microsoft looking at DeepSeek as a potential, you know, maybe cost alternative to running the GPT and Claude models. And at first, I'm like, well, that doesn't make sense, at least now. But it might make sense in six to nine months because I think at least until the work that knowledge workers are required to do changes drastically, right?
Starting point is 00:08:32 For the most part, I think in six to nine months, you know, some of the more basic, I'm not saying, you know, deep seek, right? I'm not saying that model, but I do think some of the more basic models, or even a company's second or third most powerful model, maybe for a period of a couple of years. might be sufficient enough to do the majority of all this work, which is why I think, you know, yes, when we were talking about models, capabilities from, you know,
Starting point is 00:08:57 2024 and 2025, the model really matter. But I think it matters less and less today than it did. Truth number three, prompting is losing to workflow design. All right. And if you haven't noticed this, again,
Starting point is 00:09:10 this goes along with the push toward the super app and the harnessing and all of those things. But, you know, even Google, right? Google is pushing, Agents, scheduled tasks, subagents, orchestration, right? If you are still spending a majority of your time or your team is still spending the majority of your time, prompting large language models versus reading and acting on their outputs,
Starting point is 00:09:33 you are behind. Let me just say that very bluntly, right? If you are still going in and typing things and, you know, your team has a prompt library, it's not like that anymore, right? We have skills. That's a universal language that all large language models speak. For the most part, they all have scheduled task or automations. So for the most part, humans should be working more on the front end,
Starting point is 00:09:56 you know, making sure that your context stays up to date, making sure that, you know, you are making sure whether you're working with markdown files, skills across your organization, making sure those stay up to date with your most up to date context. But then from there, you should really be spending, right? Humans should be spending more of their time, taste making, right? If you have a project that you would normally go in and prompt a large language model for, right and that project happens every Monday afternoon instead of you prompting every single day
Starting point is 00:10:24 at Monday afternoon well you should be looking at the deliverables that five different agents delivered and saying which one is best and then Tuesday after that deliverable is due you pour in with new updated context and make sure that hey next next Monday you know those five you know reports that five different agents are going to take a swing at they actually have more better context and directions truth number four static business artifacts are becoming organizational debt. And I didn't say have become on this one. Okay, I'm saying they are becoming.
Starting point is 00:10:58 Let me give you the quickest example. Again, did in a recent episode on this, actually when we talked about codex sites. So make sure to go back and listen to that one. But as the harnesses become more popular, right? As things like codex cursor, right, Claude desktop, I think has some catching up to do. We'll see if Google continues to invest in the anti-gravity 2.0 app or not.
Starting point is 00:11:25 But as these harnesses become, number one, more capable, but more team friendly, I think the concept of emailing files back and forth or even dropping, you know, static files in a shared drive, I think it's pretty soon going to become archaic. And I think our artifacts first need to become AI native. Yes, we need to figure out what that means for history and versioning and all of those things. But I think Codex sites is maybe the first iteration. Maybe it's not the final or the best, right? But I think that's a good example of the difference between how a static business artifact can
Starting point is 00:12:06 actually become organizational debt. And I didn't think it at first, but I thought back throughout the course of my career, How much time that I spent either communicating about different file versions, trying to find old versions, you know, corrupted files, right? All these things, just tackling, communicating, finding, organizing around different files, sharing them. You know, you're waiting for approval on all these files because, oh, it turns out they had V4 and they actually needed V4 final, right? All these things. I think we have to look at what does an AI native deliverable look like? Maybe it's a codec site.
Starting point is 00:12:43 Maybe it's a, you know, clawed live artifact. I don't know. But that is truth number four is we have to start thinking about AI native deliverables. Truth number five, AI can create more work if leaders don't remove old work. All right. Let me tell you what that means. AI often just adds more and more, right? If you're looking at text, walls of text, that doesn't necessarily mean that your organization is winning
Starting point is 00:13:11 back time, right? When we talk about R.O.Y, if all you're doing is continually experimenting with different AI, with different large language models, or if your employees, if your team members are just banking their save time, which I would argue the majority of companies that saw, you know, AI gains in theory, but not on the books in 2024 and 2025. Yeah, that's just your, you know, employees having, yeah, I do all the work for them doing the same amount of work that they were doing pre-AI and then saying, well, look, right. But on the flip side of that, AI, if you don't properly haul it in, can actually create way more work if you don't redesign workflows to be AI native.
Starting point is 00:13:55 So if your job description hasn't changed, if your SOPs haven't changed, especially since late 20, late 2024, early 2025, when models became more agentic, you have to do so because all you're technically doing is you are just doing. doing a disservice actually now that I think about it a little bit more right if you are still producing artifacts for clients industry reports whatever the the novel work is that you're creating business value within your organization if you're still doing it the old way let's just say the 22 way with 2026 tools that's bad right yes you are still creating more work but one of the reasons you're creating more work is because you are creating more work is because you are
Starting point is 00:14:39 creating all of these processes that in theory are not going to be transferable once your company, your organization, or your artifact becomes AI native or with whatever your, whether it's for a client, an internal team member, et cetera, right? What about when their demands change? And I think the consumer demand will begin to change, right? We're getting overslopified with all this AI slot. But I think that demands are going to change. So if you are just trying to use as much AI as possible on old processes, old job descriptions, companies and in organizations that are set up in a pre-AI way, I think ultimately you are just setting yourself up for failure, even if it seems like success in the short term.
Starting point is 00:15:23 So real ROI starts when the AI actually goes in there and removes steps, removes meetings, and reworks what you're actually working on. Truth six, longer AI task completion will reshape jobs. All right. So good example here. Infropics says autonomous task length has doubled roughly every four months. And as an example, infropic said that Claude authored over 80% of Anthropics merge code in 2026. So what does this mean? I'm not sure yet, right?
Starting point is 00:15:59 This is one of those things that still kind of keeps me up at night. And I think we, it's, it's going to take even people on the edge of AI. It's going to take them a while to figure out what does the future job look like, right? Not like, oh, what happens when we have, you know, embodied AI and human, you know, human AI robots doing all these things, even before that, right? What happens when autonomous AI harnesses are doing all of our normal deliverables that we were doing three years ago? What happens then?
Starting point is 00:16:30 But I think that managers must start to read. design work around delegation, supervision, review, and exceptions, right? We have to start redesigning old work processes because as these models can work, whether we're talking about in loops and managed and controlled loops or just working truly autonomously toward a goal, right? As they start doing things that would normally take many, many hours or many, many dates, right? We don't know how jobs are going to look.
Starting point is 00:17:02 And I'm not talking about job disruption. I'm not talking about job creation. I'm talking about the jobs that will still be here, which I think is the majority of jobs, how are they going to look differently? But as you have these models now that can run and do these long tasks, and they can work for 8, 10, 12, 15, 18 hours
Starting point is 00:17:25 on a single type of project that would normally take a team, what does that mean for the future jobs? I don't know yet. but it will reshape jobs. Truth seven, open weight models are becoming resilience insurance. Right. So I already talked as an example about the recent Microsoft news that said that they're looking at DeepSeek as a potential rollback or different offering.
Starting point is 00:17:48 Right. So they're looking at fine-tuning a version of DeepSeek and maybe using it internally, offering it up to customers within co-pilot. We'll see right now it's just reporting that's still shaking itself out. but the reality is open source and open weight models will soon become at least a resiliency insurance, right? This is something I've even started to experiment with on my own. I will say this, though, open weight models, they're great. If you think that you can slap, right, a general open weight model and, you know, have it run, you know, locally on your machine like,
Starting point is 00:18:30 why would I ever, you know, use Fable 5 or why would I ever use GPD 5, 6? I can have an open weight model running on my computer. No, that makes you look foolish. FYI, I'm just saying that. In that case, open weight models will still, you know, open weight or open source models will still be multiple years behind in terms of what a consumer can go and run on hardware. So unless you're spending $30,000 on hardware, it is much more economically feasible and just responsible for you to be paying $200 a month
Starting point is 00:19:02 and just use the frontier model, right? But open weight models are becoming a great backup plan for specific use cases. So as an example, and this one's kind of recent as well, GLM 52 Max is now technically the number one model in the world for front end coding because the only model that's ahead of it, Fable 5 is not technically available. So it is the number one available model.
Starting point is 00:19:29 Think of that, an open source model. All right. Again, you can't really run this locally unless you have a cluster of supercomputers. But again, thinking about how this shapes the future of work and how this impacts decisions that your company makes. I mean, you have to look at what Microsoft is doing there. If Microsoft is thinking that, hey, we might be able to, whether it's internally or for co-pilot customers, if we put enough human effort into this, we could fine tune an open source model and run it at production for even half of internal or external purposes.
Starting point is 00:20:05 That's big. And I think that business leaders need to start looking at that. For most companies, it's not there yet. But you have to understand that open fallbacks matter because vendor changes are going to happen. Policy prices. Access, I think through the rest of 2026, we'll see what shakes out between the Anthropic and White House. But you have to start thinking of these things.
Starting point is 00:20:25 and you need to start routing workloads by sensitivity, volume, compliance, cost, and quality threshold. Truth number eight, agents make AI risk operational, not theoretical. All right. And I talked about this a little bit more in our 2026 AI prediction and roadmap series, but as the default, as the default models become more agendic, they're getting read, right actions. Let's even take out the fact of running these, you know, desktop super apps that can access every single file in your computer, they can read, right, use the terminal, all those things.
Starting point is 00:20:59 Even look inside something like chat chpityt, even look inside something like, I don't know, GROC, perplexity computer. Even the quote unquote, non-technical AI chatbots are getting more agentic power, right? Very small example. ChatGPT can send emails in it. Right, it doesn't sound like a big thing, but chat GPT has always had the ability to draft emails and to read your emails and to pull in all this context. But just as of a week ago, it can actually send an email.
Starting point is 00:21:32 Do you know how much risk there is when you give any AI model the ability to send an email? Right. There's great promise, but there's also a great downside if you don't have the right expert driven loops because every agent needs an owner, limits, approval, logs, and rollback plan. All right. So those are our eight uncomfortable truths. Now let's get into the 10 AI moves that's smart. AI teams are making.
Starting point is 00:21:56 And yeah, we're going to pick up the pace here. All right, move number one is they build portable AI stacks before disruption hits. So I'm talking about primary backup, cheap private and manual fallback packs. Pats,
Starting point is 00:22:09 these all have to be documented, right? What happens when AI works? You have to be able to answer that first. And then what happens when that AI that works fails. All right. So smart teams are making those moves already. People always are so quick to say once you find a win and you're like, oh my gosh, this is going to save us 40% of staffing costs. So we're just going to go ahead and cut 10% of staff.
Starting point is 00:22:36 Call it a win, right? No. Use those humans and start building these portable AI stacks because disruption will hit, whether it's internal, external, from a policy standpoint, whatever you need to be kind of creating these redundancies. Move number two, they map workflows before buying. more tools. All right. I think so many organizations that I talk to, they find one success. And then they roll that success out across different teams, right?
Starting point is 00:23:06 Which is great. That's a great way to go and do things. But then what happens is then they just say, okay, then what's the next good tool that we could do this same process, right? Instead of replicating that workflow mapping across the entire organization. Instead, right? And I get it. It's hard because companies always want to be chasing the shiniest,
Starting point is 00:23:29 latest, greatest AI tool and model. Let me, let me be very direct here. Even if you, I'd say, if your organization, unless you have a team of 100 people,
Starting point is 00:23:41 whose only job it is is AI deployment in your organization, which is no one, they have no outside responsibilities or deliverables, unless you have 100 people that that's their only job, you and your company will get much more, much more value. If you literally stop, don't try GPD 56,
Starting point is 00:24:00 don't try Fable 51 or whenever it comes out. Don't try Gemini 35 Pro. You will find more business value. If you literally stop trying the next best model. And instead, look at your next or your last big win and then use the findings from that and map that out to everything within your organization. That's very hard to do. Instead, normally you just go through this concept of we're just going to map or we're just going to try to replicate our big win.
Starting point is 00:24:30 And oh, by the time we, you know, rolled out that one big win across all organizations, there's a new model. So we're just going to go chase that. It's hard not to, but you're going to get much more ROI by just mapping entire workflows before you just jump on to the next big model. Move three. Smart teams are building role-specific AI packs. Luckily, Anthropica was really ahead of the game with these skills. And surprisingly, it did take Open AI a little bit longer than I thought it would to kind of jump, I won't say on the bandwagon, but to follow suit. Skills are extremely powerful, right? this kind of takes away a lot of the grunt work of prompting, right?
Starting point is 00:25:21 So skills, I mean, you could make the argument. Skills are just a bunch of text prompts stacked together, right? And they sometimes do some technical things behind the scenes, right? But this is the basics of, well, prompt engineering, right? The skills are essentially this prepackaged markdown files that tell a large language model in a very specific way, what to do, what not to do, but around different skill sets. So as an example, there's great skill packs for different type of marketing. So if you're a marketer, you should probably at least start with a marketing skill,
Starting point is 00:25:53 even before you go in and prompt something or customize a workflow to work for how you need it. But this is huge. And, you know, skills are obviously transferable across different systems. Codex just rolled out a bunch of kind of, you know, plugins, I guess is what they're saying because their version is more using skills. plus apps, which is great, right? But every organization needs to be experimenting with skills. And you need to find a good skill that's already pre-built for your organization.
Starting point is 00:26:27 There's great open resources for skills because you can plug and play them into anywhere, right? You need to start with a skill that is highly rated, bet it across your organization, test it in a sandbox way, make sure it works, customize it, and go on to. to the next. All right, move number four, smart teams are automating recurring reports and briefings. My gosh, if you are still having 20 people in this boring three hour meeting every single week and someone's still typing up to do's because no one has access to the AI note taker, oh, or can't use it for this meeting. And then you're preparing a deck for the meeting and a deck after the meeting. My gosh, no, stop. You have to start automating all of those things because,
Starting point is 00:27:15 number one, the majority of the people in those meetings don't want to be in there anyways, right? The majority of people reading a 30-page slide deck, they only care about one little slide anyways, right? That's the reality. I think you have to start pulling in all of these things that are spending, that you are spending too much human capital on in a non-AI native And you have to start blowing up those old processes. All right. Move number five. Already kind of talked about this, but it's worth talking about again, converting static files into living artifacts.
Starting point is 00:27:51 So, sorry, spreadsheets, PDF, PowerPoints, all these things. You have to, in a good one too is skill, skill libraries, right? Your company, your organization needs to have living AI native documents. old static files are going to slowly die. Probably not by this year, right? It'll probably still be another year or two before they actually die. And until we see the rise, right, even look in the AI community, how popular things like Markdown files are now, right?
Starting point is 00:28:25 We are going to get, whether it's something like Codex sites, I don't know what it is, but you have to start converting your static files right now into living artifacts. Move number six. Smart teams are testing agents in read-only mode before giving them control. My gosh, I can't tell you the amount of times I've seen this, right? Whether it's, you know, hearing about it on, on, you know, conversations I don't have here on on the show, reading about it in news articles. the process of sandboxing an agent and putting an agent into production is pivotal, right?
Starting point is 00:29:09 I think a lot of teams think, okay, well, let's sandbox the agent, let's test it against the guardrails, make sure it's safe, make sure it doesn't drift. You know, they go through, they check their boxes, right, that makes the compliance team happy. And they're like, look, we're super responsible, right? And then they do this. You know, this is usually those, those companies that I think are a little too AI trigger happy. And then all of a sudden, they start setting these agents out. And when they're like, oh, don't worry, Bill from IT, these are human in the loop.
Starting point is 00:29:40 You know what? I know I talk about this a lot, but the world would be a lot better place if no one ever came up with this human in the loop thing because it's absolutely garbage. Because in, I tell you, 90% of the time, the human that's in the loop, number one, is the wrong person. Number two, they're not driving or steering. They just look in there when it's more of, they're a crash scene investigator
Starting point is 00:30:07 instead of an FAA flight operator. Most of the time, a human in the loop is only, quote unquote, involved when something goes wrong. So you have to test agents in read only mode before giving them control. Because as the models drive, these Asians become more sophisticated. So too does your process of the human or humans you have overseeing them. AI moves too fast to follow, but you're expected to keep up. Otherwise, your career or company
Starting point is 00:30:45 might lag behind while AI native competitors leap ahead. But you don't have 10 hours a day to understand it all. That's what I do for you. But after 700 plus episodes of everyday AI, the most common questions I get is, where do I start? That's why we created the Start Here series, an ongoing podcast series of more than a dozen episodes you can listen to in order. It covers the AI basics for beginners and sharpens the skills of AI champions pushing their companies forward. In the ongoing series, we explain complex trends in simple language that you can turn into action. There's three ways to jump in. Number one, go scroll back to the first one in episode 691. Number two, tap the link in.
Starting point is 00:31:29 your show notes at any time for the start here series or you can just go to start here series.com which also gives you free access to our inner circle community where you can connect with other business leaders doing the same. The start here series will slow down the pace of AI so you can get ahead. Move seven smart AI guys are making. They route work by value, risk, and cost. So we had an episode recently in the start here series talking about going from token maxing to token efficiency. And this part's huge. You should probably not, if I'm being honest, you should probably not be using Fable 5 for 99% of your
Starting point is 00:32:15 tasks. You shouldn't be using GPT 55x high for 99% of your tasks, right? It's crazy the amount of people that are just using the best model for everything, a blanket. That's not the way to do it. Right. I think most companies take these. I won't say their habits, but maybe best practices.
Starting point is 00:32:36 Because for the most part, in 2023 through 2025, we were all living in this subsidized AI nirvana, right? Where it's like, oh my gosh, I can just use the best model for anything regardless. All right. But times are changing. Times are changing. Token maxing has led to token efficiency. Right.
Starting point is 00:32:59 We saw all these stories, all these huge organizations, you know, hit their AI spend limit or their token budget within three or four months. And now they have to change. And the way you change that is by routing the type of work to the proper model or models that it needs. That's why I think that model routers, mixture of model setups. I think there's going to be some great AI startups. I've already seen a few launch in the last few months that do exactly this. Right.
Starting point is 00:33:30 So you are not wasting as an example, a billion tokens a day. It's not crazy for, you know, I very easily go through 500 million to a billion tokens a day sometimes. Me, small business owner, right? Luckily, I'm still heavily relying on the subsidized, you know, Open AI $200, chat chitp T pro claim. But subsidies are all going to eventually fall or diminish. So you have to start making these calls appropriately. And similarly, you do have these now very capable models like GLM 52, the Kimi models, right? We're not saying use them for every task.
Starting point is 00:34:14 But instead of using, you know, one model for 50 different tasks within your organization, you might be using, you know, five different models for eight different tasks. or something like that or 15 different models for three different tasks each. So you have to start routing them by value, risk, and cost now. Because if not, you are going to be for a rude awakening. Because as an example, right, if Anthropics like, yeah, you know, here's enough Fable 5. We'll see if the June 22nd date gets extended or not.
Starting point is 00:34:50 But, you know, it's very likely. Anthropic might go down this path of, you know, releasing these models just to all subscribers. people are going to be reworking their workflows, then they get pulled and you have to start paying API costs. So, you know, $200 becomes $14,000, give or take. So you can't keep doing that. You have to work smarter.
Starting point is 00:35:12 Move number eight, I'm seeing smart teams make. Measure accepted output, not AI activity. Right. Lines of code is not return on investment, right? tokens burned doesn't mean a thing. Are your deliverables getting better? Are they moving the needle? Are your blog posts bringing in more human visitors?
Starting point is 00:35:39 All of these things that I think we rushed to start using AI for. Oh my gosh. Look at Claw design is so great now. I can go get out 5x the number of proposals. Okay, number one, did you? And number two, are you bringing in 5x the amount of business? or are you just forcing AI slop down people's throats who are already being avalanche upon with just, you know, more and more content?
Starting point is 00:36:03 Just because we can create more and more and more text, more and more photos, more and more and more videos, more and more audio, doesn't mean that we should be. I think human taste really matters here. So you have to look at accepted outputs, right? Which of those, you know, to pull from my random example, random rambling from earlier, You know, when my agent came back and gave me five versions of a certain report, which one was number one accepted by the expert driving that agent loop? And number two, which one ultimately won over my peers, one over the internal stakeholders, one over the external stakeholders as well, which one got us more business. So you have to start measuring those things, not activity.
Starting point is 00:36:47 Activity is not movement. Move number nine, they red team the workflow, not just the model. Okay. Yes. Large enterprises need to have red teams. You absolutely have to. Right. We saw this with the, you know, even though we don't have the full story yet, but we saw
Starting point is 00:37:07 this with the, the Fable 5 rollout, right? Essentially, at least according to the White House and the federal government, you know, they said anthraxic shipped a model that wasn't safe. So what happens if you put that into production and for whatever reason, one of your customers, one of your clients who maybe relies on your services or your outputs that you use somehow stumbles upon a piece of code that wasn't safe. And you're like, well, well, this model did it. Well, okay. Yes, we assume that labs like Anthropic and Microsoft and Open AI and Google have the best red teamers in the world. It's those, those engineers that go in and make sure a new
Starting point is 00:37:50 model is safe. And they try to break it internally before releasing it. That's kind of like what red team is. That's red teaming for dummies 101, right? But you have to have an internal team that does that as well, not just the model, but the harness. So in the same way, you know, if you say, hey, we're going to roll codex out to 10,000 people, we're going to roll out Claudecode. We're going to roll out cursor, right?
Starting point is 00:38:13 I think before when you would roll these things out, it was usually developers using them the right code. But as we start producing non-technical work documents, as that becomes the norm, how are you red teaming that internally? That's what you have to do. Agent risk includes the inputs, the tools, permissions, actions, and the integrations. So your teams must test prompt injection, what happens when you use bad data, tool misuse, and also runaway costs.
Starting point is 00:38:41 And Move 10, smart teams are capturing decision reasoning in systems they own. All right, models are rented, but company reasoning history can become durable memory. I've been talking about this for years. One of the most important things your companies can do is to start collecting first company reasoning data, right? FCRD. I've used the acronym once or twice before. So the simplest way to think about this, the old school transformer models needed structured data. Right.
Starting point is 00:39:15 It gobbles up everything on the internet. Models that reason, models that think. That's today's models. Yes, they're still good. Yes, they still need structured data. Good data. Your company's good data. It also needs how your company thinks and how your company makes decisions.
Starting point is 00:39:34 Because the smartest thinking model in the world is only going to be as helpful as the nuanced information that you give it. What do I mean by that? In the same way that you would teach and invest in a human over years to, understand why you made this decision. Hey, here's why I rejected this proposal and here's why I crossed off this line item. Those things aren't always found in a spreadsheet. That's not always structured data, right? These are the things that smart teams are capturing, putting them in documents and training models to work through. All right. So now let's get to 10 questions every leader must ask. All right.
Starting point is 00:40:23 Question number one, what breaks if our primary model disappears tomorrow? Yeah, maybe a little recency bias in putting this all together. But I think the whole situation with the Anthropic Fable 5 and the U.S. government really sheds a much needed light on this AI race that has become a frenzy. What happens if that model is gone? You have to be able to identify the workflows. customers, employees and systems that are dependent on that one model. And what happens if it goes away?
Starting point is 00:41:00 All right? I'm not going to answer these questions. These questions are for you and your team. Play this episode for them. Share this episode with them. You need to be talking about these. Question two. Are we solving a task workflow or operating model problem?
Starting point is 00:41:13 Because tasks need prompts, workflows need systems, and operating models need redesign. So if you just have an AI summarizing meetings, That's a lot different from updating projects and owners because you need to be not just solving the task, but the bigger picture as well. Question three, what proprietary context makes RAI better? And this goes back to that first company reasoning data. Generic models just produce generic outputs. And it's the same thing, even though these models are becoming smarter and smarter.
Starting point is 00:41:54 If all you're doing is feeding your data to a world's smartest model in the correct way, that's not going to help. I mean, it's going to help a little bit. But that's not going to be what ultimately separates you because your competitors are going to be doing the exact same thing. They're going to be putting their best structured data with the best reasoning models or the best harness. Right. If competitors can match your outputs instantly, then the model was never your moat. Question number four, what can AI read, write, change, or send?
Starting point is 00:42:30 You need to know all of those things. And there are very few. Yeah, let me say this honestly. There are very few AI leaders out there that are in charge of AI rollout across their organization that can accurately answer these questions. And it's not their fault. It's because this space moves too quickly. So unless you have a team of people reading every single change log, right?
Starting point is 00:42:52 Or unless your whole team listens and reads the, You know, if you listen to our newsletter every day, if you, or sorry, if you listen to the podcast every day, read the newsletter. You're probably ahead of most teams. But still, even at that point, you have to know every single model, every single harness, every single connector, which one has read, right, access, which one requires approval, what permissions, right? Which permissions are needed, you know, when using which AI super app, what can do what?
Starting point is 00:43:21 You have to know those things. Question five. who owns every AI workflow? What one person? Asian crash is going to happen at your organization, in your industry, your competitors, etc. You have to know now. Number one, you have to map the workflows.
Starting point is 00:43:39 But number two, you have to know who owns each workflow. And that person has to be able to answer those questions like question four. What can AI read, write, change, or send? And you have to know who owns that workflow? Question number six, what is the cost? ceiling. The ceiling is important. We saw some stories. We don't know if they're actually true or not. There's one story that came out that one company, quote unquote, accidentally ran up their monthly AI bill to $500 million because they didn't put a spend cap on cloth. Was that true? I don't know.
Starting point is 00:44:13 But what is the cost ceiling? You have to have those measures in place. Not just how do we monitor what the cost is, but what cost should you be paid? Right. Should you be paying, I don't know, a million dollars a month to summarize emails? Maybe, I don't know, if your organization, if you have hundreds of lawyers reading very important emails, maybe, right? But you have to know what the cost ceiling is because agetic AI can spend money invisibly while appearing to be productive, right? That's why you need to look at things like artificial analysis is cost per intelligence, right? Because a lot of times people just look at a single benchmark and be like, oh, my gosh, this fable five thing is great. Well, what if it costs three times as much as something like GBT55?
Starting point is 00:45:05 If you get the same exact output, but it costs three times as much. What's your cost ceiling? Question seven, what counts as a useful output? My gosh, I've seen people just really redefine and twist what a certain end artifact should be because an AI created it. Generated content means nothing until the business accepts it, uses it, and you can prove an ROI on it. So you need to say, what counts as a useful output?
Starting point is 00:45:34 Where is your AI spend going? And is it producing an actual useful output? Question 8. What happens when the AI is wrong? Okay. Not when it goes out, but what happens when it's wrong? Do you have the right expert-driven loop that can identify? when an AI output is wrong.
Starting point is 00:45:54 Hallucinations, yeah, they are not really a problem anymore, but not everyone is using the right model. So they obviously still exist. If you do the context engineering thing correctly, if you have expert driven loops, if you're using the right model for the right purpose and you've gone through the guard railing and the scoping and all those, let me just say it.
Starting point is 00:46:12 Hallucinations are extremely rare if you do all those things, but most companies do not do those things. So you have to ask the question, what happens when the AI is wrong? Question number nine. What evidence could we show six months later? Six months later. Yeah, leaders, you need records of models, data, prompts, actions, approvals, and changes.
Starting point is 00:46:35 Some, right, at some enterprise level, like if you're using chat GPT enterprise, Claude Enterprise, right, co-pilot, they do have some useful metrics, I guess, right? I like Microsoft's approach with their intra ID. and a lot of the other companies are starting to adopt something similar as well. But every single piece of value that is created via an augmented relationship between a human and AI, it needs to be observable, it needs to be traceable, it needs to be auditedable. You need to be able to go back six months later and say, oh, my gosh, we ended up winning this huge RFP, right?
Starting point is 00:47:14 And we had this, you know, RFP agent cooking it all up. let's go back. How did we do that? Who set that up? What models were we using? What was the workflow design process? And why did we get wildly different results from all these other RFPs when we were using the same process? You have to not only be able to go back and say, hey, this worked.
Starting point is 00:47:38 Let's go back and make sure that we apply some of our successes across the board. But you have to also say, why did it work in some scenarios and not others? And then the last question, question 10, what should we stop doing if this works? And this is the big one. This is the thing about shifting human agency. And maybe all, this is good to end this episode, episode 800 with a bigger question, something maybe thought provoking because it's something I'm thinking about all the time. What happens in six months?
Starting point is 00:48:11 What happens in nine months? What happens in one year? when these agents are even more and more capable. Go back and look at where we were a year ago and look at where we are now. It's scary. There is no wall. It is a ramp going straight up. If capabilities continue, if the harnesses continue,
Starting point is 00:48:30 if these agents that run on your local machine 24-7 and can produce artifacts in the same way that humans can, and they are economically as valuable or more valuable than expert, humans, if all these things continue. What should we stop doing when this continues to work? Should we stop spending time as humans on the things that we've spent the majority of our career on? Should we stop? Let's just say, let's just say you love creating PowerPoint presentations, right?
Starting point is 00:49:02 Let's just use that as a small example. You love getting the team in the room and, you know, brainstorming and, you know, pulling in all these documents and creating a bangor presentation and you present it to the client. and they love it. What's, that's your job, right? If that's your job and you love it,
Starting point is 00:49:19 but it's shown that an agent is better, what do you stop doing? I don't have the answer, but maybe on episode 900, we will. But that's a wrap for today, celebrating our 800th episode, eight uncomfortable truths,
Starting point is 00:49:35 10 moves smart teams are making, and 10 questions. Every AI leader must ask. I hope this was helpful. You know what? I'll say one thing that is helpful for me, hearing from all of you. So if you didn't make it this far 50 minutes into this episode, reach out. Let me know what would you like to see in episode 801?
Starting point is 00:49:57 What would you like to see down the road? I know we normally on Wednesdays do our AI working Wednesday. So thanks for, you know, letting us throw a little curveball in the rotation. We'll be back next Wednesday with our normal working Wednesday series. But thank you, whether you've been around for one episode or all. 800. I hope we can keep going to 900,000 and beyond. So if this was helpful, do me a favor. Tell someone about it. I can only make it to 900,000 and more by giving you unbiased, straight AI info when you tell people. So please, number one, subscribe to the podcast.
Starting point is 00:50:34 If you haven't already on Spotify or Apple Podcasts, then make sure you go to your everyday AI.com. Sign up for the free daily newsletter. Do me a favor. Drop me a line today. Drop me an email. I would love to hear, you know, whether if something's been helpful along the way, how long have you been listening? What would you like to see us do differently in the next 800 episodes? I work for you. So thank you for tuning in. Hope to see you back tomorrow and every day for more Everyday AI. Thanks y'all. And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit Your EverydayAI.com and sign up to our daily newsletter so you don't get left behind.
Starting point is 00:51:19 Go break some barriers and we'll see you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.