Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 846: AI Hallucinations: What they are, why they happen, and the right way to reduce the risk (Start Here Series Vol 5)
Episode Date: August 21, 2026Let's talk about the AI elephant in the room: hallucinations. 🐘Maybe hallucinations are the reason your company has been hesitant on AI. But here's the thing, y'all. If you know wha...t you're doing, hallucinations are largely manageable. But first, you gotta understand what they are, how they happen, and how to reduce the risk. Let's get started cutting down hallucinations together. AI Hallucinations: What they are, why they happen, and the right way to reduce the risk (Start Here Series Vol 5) -- An Everyday AI Chat with Jordan WilsonNewsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageJoin the discussion on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:AI Hallucinations Definition and CausesLarge Language Models' Hallucination MechanismsHallucination Types: Fabricated Claims & SourcesModel Improvements Reducing Hallucination RateContext Window Impact on AI AccuracyAI Hallucinations in Legal and Enterprise SettingsFour-Layer Method for Minimizing HallucinationsCustom Instructions and Retrieval-Augmented GenerationExpert-Driven Verification and Agent Safety PracticesTimestamps:00:00 "Modulate's Velma: Smarter AI Insights"03:18 "Reducing AI Hallucinations Explained"08:36 "Minimizing AI Hallucinations with Skill"12:30 "Model Retention and Recall Decline"13:23 AI Advances: Improved Accuracy and Recall19:24 "AI Hallucinations and Their Causes"21:07 "Customizing AI Behavior Effectively"24:47 "Connecting Data to Reduce Hallucinations"28:47 "AI Oversight and Expert Input"30:56 "Reducing AI Hallucinations Simplified"Keywords: AI hallucinations, large language models, next token prediction, AI error, human error, fabricated claims, reinforcement learning with human feedback, context window, context engineering,Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info)
Transcript
Discussion (0)
Welcome to the Everyday AI podcast.
My name is Jordan Wilson, and for the past three and a half years, we put out more than 800 episodes.
Yet, one of the most common questions I get I didn't really have an answer for.
Where do I start on the Everyday AI podcast?
And that's why we started the Start Here series.
And with the fall now back in full swing, the Everyday AI podcast is going back to school and playing back the entire Start Here series from front to back.
We've hit pause on our normal Monday to Friday programming to run back our most popular series ever for the next 30 days.
We made the Start Here series for beginners and AI champions alike.
So whether you're just trying to get a grasp on large language models or grappling with the best coding harness for multi-agentic workflows, the Start Here series covers it all.
Plain language, no jargon, and easy to follow along each day.
So make sure to subscribe to the podcast.
podcast and check back each day for new insights day by day.
The series is a culmination of spending more than 10,000 hours covering generative AI
over the past three and a half years.
So you don't want to miss a single episode of the start here series.
Let's get into it.
Let's talk about the elephant in the room when it comes to AI.
That one word, hallucinations, right?
When you rely on an AI chatbots to provide you something in it straight up lies.
or it tells you a very non-truthful version of what you might be looking for.
And if you don't keep up with the advancements of AI, you might still think that AI lies all the time.
And well, if you're using the wrong model, if you're not using best practices,
hallucinations, even in 2026, can still be a huge problem.
But here's what not enough people are talking about with the latest technology in today.
thinking models and the ability to very easily ground responses in your company's data without
much tech know-how essentially get rid of at least the high rate of hallucinations.
So that's why we're going to be diving into this sticky topic on the fifth volume of our
start here series talking about AI hallucinations, what they are, why they happen, and the right
ways to reduce the risk. All right, let's get into it. I'm excited for today's conversation. I hope you are, too.
If you're new here, this is our fifth installment of the Start Here series. This is the essential
podcast series to both learn the AI basics and to double down on your AI knowledge. So whether you
are a beginner trying to understand where do I start with AI, or you're really just trying to deepen
your expertise, our Start Here series is definitely for you. And if you haven't already,
please go to Start Herepseries.com. Here's why. Well, number one, that's going to give you
free access to our inner circle community. All right. And you're going to get straight
into our Start Here series channel there. So you can go back and listen to all of our Start
Here series right there in our free community, but you can also get free access to our prime
prompt and polish, prime prompt polish chat GPT prompt engineering course.
All right.
And if you miss our last episode of Start Here series, that was Volume 4.
We talked about the human AI collaboration and best practices for working alongside AI, which
leads us right into today's show.
Because if you are doing the best practices of human AI,
collaboration what we've tackled in volume four, then that leads you straight into reducing
hallucinations. Because if you are working with AI in the right way versus blindly trusting its
outputs, well, you're going to be able to fight back against the hallucinations. So here's what
we're going to cover on today's show. First, we're going to tell you what hallucinations are
and why they are there. Then we're going to show you how to assess the risks and talk about how
models have improved on hallucinations, but they still run into them. And then last but not least,
I'm going to suggest to you a kind of four-layer method to reduce the errors and show you how you
don't really need to be that scared of hallucinations as long as you are being a smart human.
All right, let's jump into it and start at the beginning. What the heck are hallucinations and why do
they still happen? Well, in general, here's as, as, as, as,
bluntly as I can put it. Large language models, right? And go back and listen to volume one and volume
two, if you need a refresher on how large language models are built. But essentially, they scrape
the entirety of the internet, online, offline data sets, and then humans train these models. So when
you and I ask a question, right, through this process called reinforcement learning with human
feedback, these smart people at, you know, Open AI and Google and Anthropic and all these other labs
have trained models that when someone asks about XYZ, here's what you're going to be. You're
you should respond. However, for the most part, AI models are very super smart next token prediction.
So sometimes that means if something was maybe incorrect on the internet or if a model is
confused about what you're actually asking, it might present a made-up answer.
And it might do so very confidently because at their core, AI models are trained to be
helpful assistance. That is in almost every single system prompt that an AI model uses.
So that's why sometimes they are going to make things up because they want to be helpful more than anything else, right?
Because they predict the next word from patterns.
And that's kind of why they exist.
So it's almost like maybe you've heard this saying that hallucinations are a feature, not a bug.
Because at the same time, kind of that same methodology of how they're built help large language models be extremely creative, strategic.
And even if we're talking about, you know, scientific discovery, drug discovery,
mathematics, that's actually how they're able to solve new problems that humans haven't been able to.
But I think a lot of people, number one, are using the wrong model.
They're not going through the context engineering 101 of how you should be working with a model.
And they're using, right?
They're using, like I said, they're using the wrong one.
That's why hallucinations have been a huge problem.
And they can get you and your company in a lot of trouble if you're not doing the best practices.
But here's what's changing.
Models now they can think and they can reason, much like humans do.
And that has led to a drastic reduction in hallucinations, right?
Obviously, the average person is using way more inference, right?
They're eating up way more tokens to go through basic problems.
But that's why I think hallucinations are really on the decline.
So the combination of better models that can think and reason like a human do
combined with kind of this rag-esque features that front-end large language models give you
are really leading to that reduction.
But this is how they work, right?
Because people are always confused.
Like, hey, if I give a large language model a simple problem,
why does it get it wrong or, you know, a riddle or asking it how many R's there are in
the word strawberry, why does it get it wrong?
Why is there this, you know, this really jagged, you know,
almost polarizing output that a large language model can do.
It can do something that's absolutely genius, but then it can get something absolutely wrong.
Right.
So that is, well, because they are next token or next word predictions.
They are based on patterns trained by humans.
They don't verify the truth against reality necessarily.
And it's that same capabilities that allow AIs to be extremely creative and to generate code
and to write poetry.
Well, at the same times, that can lead them to.
lie to us, right? And by definition, a hallucination is just a confident answer containing either
fabricated claims because the model just generates the text or it can make up sources, right,
which is a big problem as well. So I'm going to be talking about when we go over our kind of
forward layer approach, how you can avoid that type of hallucination. So there's really
different categories of hallucinations, right? One is it just lies. One is it, it's not
really lying, but it's sounding overly confident in just giving you some generic information.
And then other times, it can just make up sources, right? So it might give you true facts,
but it might make up where it got those facts from. Right. In most cases, like I said,
hallucinations, if I'm being honest in 2026, they're more human error than AI error. And I know
a lot of people won't agree with that. That's fine. But if you get me a room full of
that are using AI models the right way, right?
Get a room full of humans that have taken our free, you know, prime, prompt polish courses.
They're going to run into very few hallucinations, right, compared to the average user.
Because if you actually know how models work, how to feed them the right data,
and how to check and verify on the back end being a smart, you know, human user,
you're not going to run into them very much.
But let's talk about the early days because it was bad, right?
So go back to the early models, the GBT3 or the GPD 3.5, you know, the early days of chat GPT
in late 2022.
Studies showed that GPT5 fabricated up to 40% of academic citations, right?
Fast forward to the next version of GPD4, that went down to about 29%.
So hallucinations got a little less, but still rampant.
Fast forward to today, GPT52 reported.
a 6.2 error rate on general queries and OpenAI claims a 30% reduction in errors in GPT52 versus GPT51.
So we're not talking, you know, 30% fewer errors between GPT52 and that version from three years ago.
No, we're talking about in a three month period.
And that jump is huge.
And I'm going to explain why that jump has occurred.
and I think why in a year from now, we may be not even talking or talking very much about hallucinations.
And the main reason is just the ability for these labs, right?
So I'm going to give an example here from Open AI, but I think that Anthropic in Google and Open
AI have made tremendous strides with their models and through a lot of different techniques,
I'm not going to get into the technical side.
They've made AI models that have much more reliable outputs.
And one of the reasons is their ability to handle longer context.
All right.
So for our live stream audience here, I have a screenshot from OpenAI's GPT52 model release.
And I want to kind of talk about this long context.
So there is a test.
It's essentially called the four needle test.
So what this is, they have it pull out and they ask the model questions.
And it's a needle in the haystack test.
And they see over a wide range of a conversation, right?
If you're using a model into the hundreds of thousands of tokens, right, like at 256,000 tokens.
So a very long conversation because what has happened in the past, a lot of these hallucinations come when you are hitting a longer point in the context.
window, right? So think of the 3 p.m. brain fog, right? Let's just say you work 9 to 5. At 3 p.m.
You're probably not as sharp as you were at 9.30 a.m. right? When that second espresso hits and you're
like, let's go. And you're firing all similarers, right? For the most part, that's how large language
models had been, I would say, even late into 2025. But think of that 9 to 5 instead as a context window,
right? Because all large language models have their constraints. They have their
kind of confines that you can't break through. And one of those is the context window.
That's how much information a large language model can retain until it starts to forget.
And potentially when it starts to forget, it will start to hallucinate.
So even great models like GPD-5-1, as I go here on the results of that four needles test,
you saw its ability to properly pull out those facts over a large context window.
When it started out, right, it was about 95%.
It was very good at GPD-51 thinking.
But then toward the end of that context window, it dropped down to like 45%, a 45% ability
to recall that information.
Whereas GPT-5-2 thinking hardly any decline at all.
I believe it was at like a 95 or 96% at the end of the context window.
Okay, so think that very drained human at 3 p.m.
They might not be able to recall facts, right?
But with today's, and when I say today's, I'm saying the latest generation, Gemini
3 Pro, Opus 4.5 from Anthropic, Claude Opus 4.5, and GPD 52 thinking from OpenAI,
their ability to recall information across that context window has greatly improved,
which is one of the main reasons why that hallucination rate has gone down, because now models
can think. They can only think and plan ahead and reason and call tools on their own to
provide you more accurate information, but they're able to do that across the entirety
of their context window, which is one of the main reasons why hallucinations are going down.
Here's why this still matters for your business. Well, let me cut it to you.
straight. One of the biggest gaps right now and one of the most blaring things that are going wrong
with AI implementation across the enterprise is a lack of training and education. Because you have
companies that I've talked to personally. There's plenty of case studies out there that are rolling
out access, whether you're talking about, you know, Microsoft co-pilot 365 licenses, you know,
Chad GipT Enterprise, Claude Enterprise, Gemini Business, Gemini Enterprise, etc.
They're rolling out AI access, a lot of this was in 2025, to thousands, tens of thousands of
employees, but not giving any best practice training or even education, which is why hallucinations
are still rampant. People don't even understand, oh, I need to click this model selector and
choose the best model for the job. Or I should be having this model.
you know, call a certain tool.
It should be running Python.
I should be uploading files.
People know the basics.
Yeah, they're just trying to save as much time as possible with AI.
And that's led to a lot of hallucinations with a lot of high profile and a lot of press around
it, right?
So as an example, there was the AI hallucination cases database at HEC Paris, documented 486
different legal cases worldwide involving fabricated AI content.
Yeah, a lot of these like hallucination stories.
They come at the most embarrassing point, which is in the legal sector, right?
So there's been over 128 lawyers that have been cited for filings with hallucinated cases.
One of the most popular is the, I think it's the Meda versus Aviancha.
I don't know if I got the pronunciation of that, right?
I might have hallucinated the pronunciation.
But that's where you saw attorney sanctioned for six fake citations.
That was one of the more infamous cases of AI in legal early on.
And even Deloitte reportedly refunded.
part of a $300,000 Australian government contract after it was, they reportedly found some
AI generated phantom citations in their report. So this isn't just people who are, you know,
using AI to write blog posts, right? A lot of times some of these hallucinations make their way
to the forefront in very high, high value and just extremely visible places, right? Like consulting
companies like lawyers.
We've seen plenty on the financial side as well.
So that's why this still matters, right?
Because maybe what I laid out for you and like saying, hey, if you use the right
model and look at these, you know, the needle in the haystack, that's improving the context.
Well, most people aren't using models the right way.
And that's why this is still extremely important for your business.
Don't worry.
We're going to lay it out for you here.
But I want to be very clear.
just because I am extremely optimistic about hallucinations decreasing with proper human involvement
and training, it doesn't mean they're going away, right?
The very nature of what a large language model is, how it works, means that hallucinations
will be there for a while, right?
Because the models are always going to optimize for what word or what sets of words could
be next, even if those things are incorrect.
Think of how many times that you've been on the internet researching,
something and you're like, ah, this isn't right. That's not right. In the same way,
well, large language models take all that information from the web. So let's say you are a domain
expert and there's something, you know, that in your field people constantly get wrong or there's
some pieces of information out there that exist that are circulating that aren't exactly true.
Well, in the same way that you might find incorrect data on websites or you might go to a conference
or hear colleagues speak about something that you're an expert in and you're like, nope,
that's not right. A lot of times large language models are reflecting inaccuracies or, you know,
half lies or half truths that exist in the real world. Second, the second reason why hallucinations
probably are going to go away is, well, when they're needed information isn't there by default,
by default, large language models are going to fill in the gap, right? They're going to do everything
they can to be a helpful assistant because that is the default behavior unless you change it and you
should and I'll tell you how.
And then third, the chat interfaces rewards that fast, confident answers, right?
I'm not going to get too much into, you know, reinforcement learning with human feedback and
kind of scaling laws and inference and all that.
But for the most part, chatbots are trained to quickly give you an answer and being token
efficient, right?
So if it thinks, oh, I can spit out an answer rather quickly, it may just do that unless
you told it not to.
Right?
So they'll at times will reward kind of giving you an answer, even if it's not confident,
versus just saying, I don't know.
I will say that's models more of late 20, 25, because the models from 2026, you'll see
by default, again, using the right models in the right context.
They are going to say more, and I'd love to hear if you've seen this more.
I have models that actually say, I don't know or I'm not sure, right?
even without giving, you know, anything extra in your prompt, anything in the special instructions,
today's models are more likely to say, hey, I'm not, I'm not sure, which is a good thing.
So here's how to spot and cut down on hallucinations, a quick four-step plan.
All right.
Number one, you need to change the model's behavior.
Like I said, I do think that, I think that in the future, the AI, the AI labs,
are going to find that sweet spot between instructing models to be helpful assistance,
but not too helpful and not fabricating things.
But in the short run, you can do that, right?
So whether it's in your prompting sequence when you are working with a large language model,
or what I would recommend is setting custom instructions, right?
So there's different kind of places that you can set custom instructions.
But if you are non-technical, if you're not really sure what that means,
that's in, there's settings in most of the big providers that you can essentially put in your own
rules in any chat that you use, any response that you are using the model, it will always go through
and read and usually adhere to the own set of custom instructions that you put in there.
So, you know, even putting something simple, like if you're not certain or the information
isn't provided, say, I don't know rather than guessing. Something simple is that I have a set of
custom instructions that I've done a lot of testing with over the years that are much more robust,
but even putting something simple like that, like, hey, if you're not 100% sure, say you're not sure,
right? Or to require every factual claim to include a source or, you know, labeling everything
afterwards, right, saying, hey, this is what I'm 100% telling the model. This is what I'm 100% sure
on. This is what I'm not sure on. You can even do something like in each response, having it
giving you a confidence score. That's another great thing that you can do. And then have it to
separate facts from inferences, right? Because here's the thing, especially if you're using it as a
strategy, creative partner, brainstorming. Not all those things are black and white, right? Maybe 99%
of what you might use a large language model for is in that gray area, right? It's strategies,
creative, right? Being creative. So have it separate facts from inferences. You can ask for a table with
columns for confirmed fact versus assumption in a structured output.
So number one is doing that either in your prompt or in your custom instructions.
Number two, my gosh, the fact that we have this and you don't have to pay any extra is wild, right?
Make it retrieve the information, right?
This whole context engineering, something I've been teaching since 20, 23 before it was a thing.
right? If you've taken our free prime prompt polish course, this is the refine queue in the priming.
So we call it the fetch in the insights, but essentially using a smaller simplified version of RAG,
Retreatable augmented generation. Now you have a simplified version of RAG available with a couple of clicks
by connecting your company's data, right? So both in obviously Microsoft 365 copilot in Claude,
in Google Gemini, in OpenAIs, chat, GPT,
they have different ways for you to connect your business data in a few clicks.
So you can essentially ground it in a way you can grout it in a way you kind of can't.
But the combination of number one, those kind of custom instructions,
and number two, first putting your company's information,
if you combine those two things and you say, always check, you know,
this ABC document before you respond.
and then if you make sure you check and verify it that it did, I mean, those two things right there,
amazing. But it doesn't matter if you're, you know, a Microsoft organization, you can connect your,
you know, your OneDrive and SharePoint data in Open AI's products, right? If you're using,
you know, Microsoft 365 copilot like online, their online version, you can connect your, if you use
Google, if you use Google Drive, right? So being able to connect your company's data to a large
language model and then instructing it to always look at those things first is huge.
And it makes a big difference.
So there was a 2024 Stanford study that found RAG combined with reinforcement learning
with human feedback and guardrails achieved 96% hallucination reduction versus the baseline.
Just doing those basic things, right?
Like having a version of your company's data and proper instructions are going to cut down
on hallucinations, just those two things alone.
And then steps three and four kind of combined into one is just the verification
workflows and agent safety.
So it's kind of like three and three beats.
All right.
So what is that?
Well, you have to be able to catch errors before they escape.
This is the expert driven loops that I always talk about, not the lazy, passive human in
the loop.
I'm talking about the active, proactive, expert driven loops.
So as these models, they think by default.
You know, I say GPD 52, right?
And if you're listening to this, this episode in July, right, maybe it's GPD 53 or Gemini
3.5, I don't know.
But today's latest models are agented by nature.
They're going to go through, they're going to think, they're going to decide how much to
think.
Should I spend five minutes on this answer?
Should I spend 30 seconds?
Right.
But you have control over that.
And you can always kind of build, and I always recommend doing something like this,
build a second pass review where that model.
model's only job is checking claims.
That's why I also do kind of a mixture of models set up.
But I have system set up where, you know, I maybe have a Google Gemini deep research run.
And then I'll take some of those results and I will verify it.
And I have something set up.
The model that I like doing this for is GPD 52 Pro.
I will then use GPD 52 Pro only to go through and verify every single thing that I get from a different model.
That doesn't mean I don't trust model A versus model B, right?
And sometimes I'll flip-flop it.
It means you should always be doing this, right?
Especially for high value, highly visible projects, right?
If you're just trying to see like, you know, what's the weather next week or, you know, something topical, right?
Like what's the best, you know, what are the three best softwares for tackling this issue?
I don't think you necessarily need to go through all those steps.
But if this, if you are using an AI model as part of a high value workflow,
you should definitely be doing this second pass review.
And then, again, requiring even that second pass to show the sources.
And then more than anything, you need to be tracing and observing how these models are getting to those conclusions.
And what that means is, well, you should be using a thinking model and then reading the summarized or the chain of thought.
Right.
So what that is, most of the models, you know, different models give you different level of visibility on what they're actually doing under the hood.
but just if you were to hand off an important assignment to a brand new employee,
hopefully you wouldn't just hand that off to the client, right?
You would take a step in between and be like, hey, new employee, tell me how you got to this
answer, right?
Let's say they spent an entire week on this big project.
I would hope you would take at least an hour or so to sit down with that new employee and
be like, okay, how do we get to this conclusion?
Walk me through, right?
So the good thing with these agentic models by default, you can see all that.
And you, smart human, need to be able to go through and check and see what it did at each and every step.
Did it follow your instructions of, you know, here's what's factual.
Here's what I inferred.
Did it make sure to look at the right documents at the right time?
You can go through and sequentially check all of those things.
And if it didn't do it, then you can course correct.
But that is layers three and four.
And especially as AI can take action, right?
As we're talking about agentic AI, that's.
when you really need to treat it like a junior employee. But no, they can still make things up.
And like I said, most of the time, if you are doing the proper context engineering one-on-one,
if you're using the right model, and then if you're going through kind of layers three and four right
here, the second pass in the observability and traceability of going through that chain of thought,
hallucinations aren't going to be the biggest problem for you, right? You still need to have
those expert-driven loops because as you're looking at that chain of thought, as you're
you're providing data on the front end, you have to do that at an expert level.
You can't just say, here's a thousand files.
Good luck.
Agent, no, you need to, you know, point them in the right direction.
The same thing, checking responses on the back end.
You have to know what all of that means.
But if you have expert-driven loops and if you go through those four steps that I talked about,
I think hallucinations are no longer going to be the big elephant in the room.
You're going to be able to reduce them.
All right.
But remember, it's not going away.
Hallucinations are a property that you need to manage,
not something that you hope will get fixed one days.
And it's not that the winners right now are picking the best models.
They're just building those four steps of verification into every single episode
or into every single piece of work that they're doing.
All right.
Speaking of episodes, I hope this episode was helpful because the Start Here series
I want you to know all the details, right?
As someone that's done this everyday AI thing now, more than 700 times,
I've got to speak with some of the smartest people in the world,
I've realized if you're educated, if you're trained,
if you understand how these models work, and if you keep up,
that's the big caveat there, right?
If you keep up, you don't have to be as worried about hallucinations.
I'm not saying you can write them off.
You shouldn't.
But if you're going through the right steps and what we went over in today's episode,
you are definitely able to reduce the risk and to get more out of large language models
because that's what we're all about here at everyday AI,
cutting through the fluff, giving you the facts and the right information to grow your company
and your career.
So I hope volume five of the Start Here series was helpful.
I hope you know more about hallucinations now, why they happen,
and how you can reduce the risk.
So if this was helpful, please go to start here series.com.
That's going to give you free access to our inner circle community,
the prompt engineering course,
as well as an easy spot to go listen and catch up
and engage with others who are going through this Start Here series with you.
So thank you for tuning in.
I hope to see you back later for more Everyday AI.
Thanks, y'all.
And that's a wrap for today's edition of Everyday AI.
Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everyday AI.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.
