The AI Daily Brief: Artificial Intelligence News and Analysis - Why a New Class of AI “Judgment Models” Could Have Big Business Implications
Episode Date: September 16, 2026A new AI model called Jev is built to make fast, inexpensive judgments rather than generate text. NLW explores how this approach could reshape business automation, help agents check their work, and co...ordinate decisions across teams. In the headlines: Zuckerberg pushes back on a collective AI slowdown, Bernie Sanders and Steve Bannon find common ground on AI regulation, and Salesforce announces a new model and tools for third-party agents.Multiplayer AI Sprint - https://multiplayerai.ai/Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/SophisticatedHarbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidailyHyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybriefRackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/Robots & Pencils - Cloud-native AI solutions that power results https://robotsandpencils.com/The AI Daily Brief helps you understand the most important news and discussions in AI. Newsletter: https://aidailybrief.beehiiv.com/Interested in sponsoring the show? sponsors@aidailybrief.ai
Transcript
Discussion (0)
It's not every day that we get a new model to play around with, and it's certainly not every day that we get an entirely new approach to model building with some fairly different implications for how we even use it.
Today, though, we are talking about a new class of models which you might refer to as AI judgment models.
Rather than producing long strings of text, these judgment models, like the one we're discussing today, Jev from TypeSafe, produce probabilities around specific questions.
Is this customer angry? Is there a new dependency in this email? Do we need to change the operational?
operational plan because of this. Today we're exploring the idea behind these models, how they're
trained differently, how they can produce these judgments much more quickly and much less
expensively, and most importantly, where they're going to fit in your overall model stack.
The AI Daily Brief is a daily podcast and video about the most important news and discussions
in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's
sponsors, KPMG, Blitzy, Section, and Hyperagent. To get an ad-free version of the show, go to
Patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts.
Subscriptions are just $3 a month for ad-free. And if you want to learn about sponsoring the show,
send us a note at sponsors at AIDilybrief.aI.i. Or just go to AIDilybrief.A.I.
daily brief. where you can learn all about it.
The AI safety discourse continues to trickle out through the tech industry as well as mainstream
society. But for now, unless something absolutely seismic happens, we're going to move it
into the headlines and away from the main episode. With that in mind, after staying quiet over
the weekend, Mark Zuckerberg has made his thoughts known on this idea of an AI slowdown.
On Tuesday, Zuckerberg wrote in a post on X,
Every lab has the responsibility and incentive to move at the pace required to train its model safely
and the ability to take its own actions to ensure that happens.
Basically, his view is that pacing is the responsibility of individual labs rather than a
collective action, and that view hinges on two core ideas.
First, that, quote,
people don't want to use agents that are misaligned with them and that don't do what they ask,
so labs have a strong natural incentive to make their models more aligned,
and two, labs face significant liability if their models cause harm,
so they have a strong incentive to prevent this as well.
Emphasizing the point, Zuckerberg said that Meta had delayed the release of Muse by several months
to work on safety. He continued,
We didn't call for everyone else to do this before we would.
We just did it as part of our day-to-day work because it was clearly the right thing for people and for us.
essentially Zuckerberg is saying that the individual incentives and consequences that are already in
place are enough to force AI labs to work on alignment and to pace the frontier correctly,
rather than needing some exogenous government-in-forced slowdown.
Now, Zuckerberg did support the idea of independent evaluators and advisors as a matter of best practice
rather than regulation. He claimed that meta has already engaged outside evaluators,
not because it's required of them, but because it helps produce better work.
Finally, he concluded,
Committing the significant majority of compute
towards serving people,
rather than racing towards recursive self-improvement,
is one of the best ways to ensure that we develop this technology safely.
Meta has made this commitment, and other labs can do this as well.
I believe the key to building a positive future for everyone
is maintaining the right balance of power.
This is within our power to do.
It's carefully worded, but basically this whole thing says,
come on, guys.
Let's please stop with the theatrics.
And a lot of people, frankly, found this a breath of fresh air.
YouTuber Joseph Carlson wrote,
Hold up a minute, you are telling me companies can slow down,
make sure things are safe without telling all their competitors to slow down?
And I would say that broadly speaking, reactions fell into one of two categories.
The first was, like that one and like this from Matthew Berman,
love this, model safety and alignment is a feature and economically incentivized.
Or Bill Ackman, who simply called this the proper approach to AI development.
But then on the other side were those,
arguing effectively that Zuckerberg just does not have the trust or standing to make this argument,
regardless of the merits of the argument itself.
Hard Forks Kevin Ruse wrote,
whether you're a Democrat or a Republican,
an EA or an EACC,
I think we can all agree that the person
best suited to protect us against the harms
of powerful new technology is Mark Zuckerberg.
Now, meanwhile, over in Strange Bedfellows Daily,
Bernie Sanders and Steve Bannon
have joined forces to call for human-centric AI regulations
in the strangest alliance of this political cycle.
The two highly ideological leaders
spoke from the same stage on Tuesday at the Future of Life Institute's pro-human assembly in Washington.
Sanders told the crowd,
if the lives of every man, woman, and child are going to be fundamentally changed by this technology,
then the people of this country must make decisions about AI and not just a handful of oligarchs.
Bannon had very similar remarks stating,
The American citizens are not going to be supplicants to the oligarchs anymore.
We can't do it.
This is a hinge in history.
We have to handle this correctly.
To handle it correctly, number one, we can never trust when an oligarch says.
Now, I will note that while this seems strange at first, these two highly ideologically opposed people sharing the same stage, the broader movements they connect to have a fair bit of shared context in history.
The 2016 presidential election, in which Bernie was narrowly beaten by Hillary to miss out on becoming the Democrat nominee, and in which Bannon obviously architected the first Trump administration, both had their roots in populist anger at the post-GFC financial landscape and the lack of accountability for the institutions and institutional leaders.
who were involved in creating that particular economic crisis.
Obviously, the left and the right's reaction to that particular context were different,
but it is, I would contend, perhaps, less surprising than you think,
that at some point these two would find common ground.
And to be clear, I am not dismissing the common ground that they find just on the merits
of this particular issue itself, and the power of this particular issue to scramble
existing political alliances.
I'm just making the point that their stories are actually more intertwined than you might
think at first glance.
In any case, the event seemed somewhat less about the existential risks of AI that have filled
the headlines this week and more about the class struggle the technology has come to represent.
TV screens played a parody interview from a fictional AI CEO, described as someone who
loves people but isn't crazy about humans. Throughout the event, it seemed that AI itself wasn't
the risk that needs addressing, but rather the unchecked power of oligarchs.
Now, you might remember that back in August, journalist Jasmine Sun towards the country speaking
to real people who were working in opposition to data centers and found that the most common
complaint was not about electricity, water use, or noise pollution, but a lack of control.
and the sense that the future was being forced upon their communities with little input.
The same was true about this event.
Across multiple speakers, the risk of AI was not framed around cybersecurity, bioterrorism,
or other ex-risk.
It was about a lack of agency in determining the shape of the future.
From this followed their main concern, which was simply the idea of losing control of AI.
As Texas Democrat Greg Kassar, who is sponsoring Bernie Sanders Superintelligence Bill said,
the answer is simple.
We ban AI systems that are too powerful for humans to control.
And as strange as it might seem, despite all the intensity of this rhetoric recently, I still
contend that we've moved into a new phase that's all about negotiating the relationship where
citizens and governments have a stake in this, and where given that that is now pretty much
where everyone is, the next phase is likely to include a lot more specificity and, dare I say,
nuance. Glenn Beck, for example, had a long monologue on his show about how he can both
have signed the pro-human AI declaration alongside them, but also disagree with them on specifics
of data centers, and honestly, as crazy as it sounds, I think that the more that the political
discourse kind of frags your brain for how confusing and all over the place it is, that might
just be a sign that we're actually making progress. The conversation also found its way into
Salesforce's annual Dreamforce event, where the company unveiled a new AI model, but where a lot
of the chatter on social media focused on appearances from Sam Altman, Jensen Huang, and Dario
Amadeh, who each iterated their own safety views. Dario argued that he was just trying to put
forward a set of standards that the industry could organize around, while Altman said that he was
very confident in our company's ability and our industry's ability to do this safely.
Jensen Wong, meanwhile, reiterated the Zuckerbergian view, saying, run as fast as you can, but if you
feel at any given point in time the company's out of control or the product's not going to be safe,
take a pause and make sure you get it right. And as for Salesforce CEO Mark Benioff, he believes
that every company has a responsibility to uphold ethical standards. Still, to give the safety
debate a bit of arrest, Salesforce also had two big practical AI announcement.
First, they're releasing their first in-house model in quite some time.
Called COA, the model is a fine tune of Nvidia's Nemotron and is designed to handle sales
management within the CRM.
The announcement reinforces the role that open source has to play in the enterprise by enabling
this sort of narrow vertical model.
Secondly, Salesforce unveiled a new initiative called AI Force.
This will be the umbrella term for Salesforce connectors that will allow third-party agents
to access Salesforce data.
The release reinforces Salesforce commitment to moving towards headless software in a platform
agnostic way, allowing any agent to become the interface. Now, obviously, these are big conversations
that are important and are going to continue, but for now, that is where we will close the headlines.
Next up, the main episode. A new study from KPMG in the University of Texas at Austin found that
when people work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than
500 early career professionals and found that the best performers consistently amplified
the value of AI by guiding, evaluating, and refining its outputs.
These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI.
Learn more about what separates AI amplifiers from everyone else at KPMG.com slash US slash AI amplifiers.
Here's why most legacy modernization projects fail.
The AI doing the work can't understand codebases at scale.
It sees a small slice of context, examine syntax, and misses years of decisions distributed across the global application ecosystem.
Blitzy solves this the way it solves everything.
Grounded in your code before any migration begins,
Blitzie's agents reverse engineer the entire legacy system
into a persistent knowledge graph,
every dependency, every constraint,
every piece of tribal knowledge that used to live in one engineer's head.
From that understanding,
Blitzy autonomously executes language migrations,
framework upgrades,
and monolith to microservices transformations,
all validated end to end.
One Blitzie customer modernized a $10 million monolithic insurance stack
in 16 weeks,
against a 137-week baseline with coding agents.
That's 9x.
compression. Retire technical debt while accelerating your roadmap. See how at blitzie.com. That's
BLiTZY.com. Here's a harsh truth. Your company is probably spending thousands or millions of
dollars on AI tools that are being massively underutilized. Half of companies have AI tools,
but only 12% use them for business value. Most employees are still using AI to summarize meeting notes.
If you're the one responsible for AI adoption at your company, you need Section.
Section is a platform that helps you manage AI transformation across your entire organization.
It coaches employees on real use cases, tracks who's using AI for business impact,
and shows you exactly where AI is and isn't creating value.
The result, you go from rolling out tools to driving measurable AI value.
Your employees move from meeting summaries to solving actual business problems,
and you can prove the ROI.
Stop guessing if your AI investment is working.
Check out section at sectionaI.com.
That's SECT-I-O-N-AI.com.
This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.
Forget local agents and chat workflows waiting on your laptop to be prompted.
HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses.
Marketing agents turn competitor moves into landing pages.
Sales agents enrich leads, draft emails, and updates the CRM.
Ops agent chases the paperwork and tracks the budget.
Every agent has access to shared context and follows your rules about scope and approvals.
It's time you add agents that feel like teammates.
Hire yours at Hyperagent.
Get $100 in credits at hyperagent.com slash AI Daily Brief.
Welcome back to the AI Daily Brief.
On today's main episode, we are looking at something really different and quite rare,
which is, in short, a totally different approach to AI that is not just another LLM.
Now, in the nearly four years since ChatGBTGBT was released and the three and a half
that the show has been around, the vast majority of the things that we and the rest of the industry
have focused on have been in some ways related to large language models.
And yet, as some have pointed out, although often in a way that didn't get much traction
and was more or less screaming into the void, LLMs are not, in fact, the totality of artificial
intelligence.
Yesterday, Diogo Al-Meda posted on X,
After co-inventing Chat-GabitT, I kept asking myself,
why have superhuman chat models not led to AGI?
I've spent the last two years in stealth building a new way to train models, R-L-CD,
or reinforcement learning for calibrated decisions,
and a new type of frontier AI model that we are releasing today.
Jev.
Jev is 20 to 200 times faster,
40 to 400 times cheaper with output tokens free,
frontier composable intelligence optimized for decisions.
As far as I can tell,
the shortest path to AI-based economic revolution.
Now, one could absolutely be forgiven for seeing numbers like 20 to 200 times faster and 40 to 400
times cheaper and being a bit skeptical of the claims to say the least. And were this just another
LLM that skepticism would be entirely warranted? However, with Jev and this new strategy, we're
dealing with something that is quite different. The company behind Dev is called TypeSafe,
and in the announcement blog post, they talk a little bit more about what makes their approach
different. They write, we built a new stack entirely focused on automation with a new model architecture,
parallel sampler for maximum efficiency, and training method we call reinforcement learning for
calibrated decisions. Whereas existing LLMs optimized for human preference, i.e. write-ups and chat
responses that human raiders prefer, the new system one models, the first of which is
Jev, optimized for calibrated decisions, or answers with epistemically honest probabilities.
So what does that actually mean? Well, let's look at how Mike Taylor from
every describes it. He writes,
Think of it as a smart if-then statement that determines what happens next when you're
automating a workflow. Say you're building software that prioritizes customer service requests,
and you write code that asks the model, does this customer sound angry?
Jev might answer 0.9, which means there's an estimated 90% probability that the answer is yes
based on what the model learned in training. You could also provide categories you define
like annoyed, irritated, offended, furious, and enraged, and learned that the customer was
60% likely to be classified as furious, with only a 10% probability of being enraged.
Going on to explain why this matters and how it differs in LLMs,
Mike continues, with an answer of 0.9, very likely to be angry,
the software might automatically proceed to escalate the customer concern to a manager,
or if it answers 0.1, not likely to be angry, that request might be deprioritized.
However, chatbots are trained to respond with flowery text, like, you're absolutely right,
this customer does sound very angry. Would you like me to compose a draft email response in a
friendly supportive tone? This text response would cause the program you're building to crash
because it was expecting a number between zero and one, not an essay. Teal Fellow Michael Lee says
this allows a class of decision-making that was neither suited to dumb, unintelligent code,
nor to slow, expensive LLMs. As Chubby sums up, Jev is an AI model built for decisions
rather than text generation. And it is important to note that the trade-off here is that this model
does not generate text. It is not, in other words, a replacement for LLMs in general. It's a replacement
for a certain category of work that LLMs do, where they've been very square peg mashed into a round
hole to do it. The idea continues, Chubby, is to embed fast, cheap AI decisions into software.
So what is this model actually built for? Well, think of how much of office work consists of reading
something and deciding what should happen next. Does this message need a response? Which department
should handle it? Does this document answer the question? Is this customer describing a bug or asking
for a feature? Does this draft make a claim its source doesn't support? Is the situation routine
enough to automate or should someone review it? These are judgments about meaning, and they're
often difficult to express as fixed rules. Jev then is designed to take the relevant information
and answer narrowly defined questions with probabilities, categories, or scores. The surrounding
software then uses those answers to route, rank, flag, or proceed. In the typesafe documentation,
explicitly recommend breaking complex decisions into small questions and then combining their results
in code.
So where would this show up in normal business?
One obvious area is customer support, where the small judgment the model could make would
be something like, is the customer frustrated and how if previous replies failed to address it?
With those small judgments in hand, the software could then next route the ticket, raise its
priority, or request human review.
In the sales domain, the model could judge, is this a buying inquiry?
Does the prospect fit the product? Are they requesting a meeting? The software could then take those
judgments to sort inbound leads and assign follow-up. In marketing and editorial, the model might judge
whether the copy meets specific style rules or whether the offer is clear. The software that surrounds
it could then flag passages for revision before publication. And importantly, where a lot of people
went was not just understanding where they would use this instead of LLMs, but how they might use it
alongside LLMs. YC founder Nathan Flurry wrote,
I'd imagine a lot of workflows that look like LLM proposes options,
Jev decides, code executes. To put a clear example on this, imagine that a customer writes,
this is the third time I've contacted you, we still can't export our reports,
and our renewal is next week. A support workflow supported by something like Jev
could ask several questions together. One, is the customer describing a product problem?
two, does the message indicate repeated unsuccessful support?
Three, is a commercially significant deadline approaching.
Four, which team is best equipped to help?
The software that surrounds it could then combine those signals with actual account information,
such as the renewal date, and escalate the ticket.
From there, you would still have a generative model draft the reply,
but Jev, once again, could check the draft against narrow criteria.
Does the response acknowledge the repeated contacts?
Does it address the export problem?
Does it promise something unsupported by the information provided?
And part of what people are excited about opening up with this new approach
is that cheap judgment makes frequent checking more practical.
If a check adds a noticeable delay or expense,
a team may run it only on selected cases or at the end of a task.
If it becomes sufficiently fast and inexpensive,
it could run on every incoming request after each draft revision
across many candidate documents before an agent takes a consequential step.
Back in Mike Taylor's article from Every, he gave Jev the text from all 27 of his articles,
alongside 10 deliberately AI-styled counterpoints, then asked the same 21 questions
concurrently across all articles to check for AI tells.
Basically, the check that he was doing with Jev was, does this essay do specific things
that indicate to people that it is AI composed?
Things like, does the text repeat an idea without adding evidence?
Does it force a symmetrical both sides argument?
Does it over-explain a straightforward point?
In less than 0.7 seconds, Mike said,
Jev quote unquote read all 37 documents and answered all 21 questions for each,
returning 777 judgments for an estimated quarter of a cent.
As he points out, that's fast and cheap enough to AI check everything everyone at your company has ever written
and get the results back in an instant.
A comparison that Mike makes is a code linter for knowledge work.
He writes,
In software development, a code linter is a tool that analyzes your work
and almost instantly flag syntax errors,
catches bugs, spots bad patterns,
and it forces stylistic consistency.
TypeSafe's model is so fast at turning fuzzy tasks
into clear, structured answers,
that it could act as a kind of codenter for knowledge work.
Give codex or clawed access to Jev and a list of questions,
and it can quickly check its own work for problems you've told it to avoid.
Perengrat says,
most software is ultimately a giant tree of,
if this, do that, if this root here,
if this escalate, if this reject, if this, ask a human.
Jeff is basically asking, what if those if statements could understand messy human context?
That's a much more interesting framing than another AI model.
I can see this being very useful for fraud and risk, support routing, moderation,
PR and QA, automation, lead scoring, compliance, workflow orchestration, and agent routing.
Early tech, obviously, he says, but the category itself makes a lot of sense.
Now, interestingly, Matt Stockton points out that in some ways,
companies adopting this amounts to a post-LLM AI technology, making pre-LLM machine learning techniques
a little bit more accessible. As he writes, lots and lots of problems in business are
classification or regression problems. Lots and lots of companies don't know that the types of problems
they have are solvable by classic machine learning methods. They often solve them with people
in process instead of technology. With the emergence and popularity of LLMs, more companies are
thinking, maybe we can use AI for that and are solving classification and regression
problems with LLMs. This is good in some ways because companies are potentially automating
some manual work, but also bad in some ways because it's often the wrong tool for the job,
and possibly not as good as classic ML methods for what they are trying to do. But he points out
the classical techniques require you to label your data, train a model, and host that model
somewhere. They aren't as easy to use compared to calling an LLM API, and it requires you and your
org to be aware of those techniques and capable of investing in them. Without going too far on
the analogy, he basically says one way to look at Jev is as a U.S.
for using LLM-style user interaction patterns for classical ML techniques.
Now, one interesting question that comes up is whether this is for individuals or teams
building systems.
And the short answer is that, while it is absolutely both, it also puts a fine point on
the multiplayer AI themes that we've been talking about recently.
Certainly individuals could use this sort of capability to have personal tools that
sort an inbox against their own priorities, check drafts of their writing against
an editorial rubric, rank saved articles against research interests, or flag commitments in
meeting transcripts, things like that, that are going to personally help you do your work better.
But I think that where this sort of technique is going to really shine is in the domain of
teamwork that happens through small judgments about who needs to know, who should act, and
whose approval is required.
When an agent is serving a single person, it gets pretty far simply by learning that person's
preferences.
An agent operating across a team, however, needs to understand the relationships between people's
work.
That creates a different set of questions.
Who owns this?
Whose work does this effect?
Is someone waiting on this decision?
decision? Does this promise create an obligation for another team? Can the current owner decide or does
this need broader agreement? In some ways, the interesting unit of work becomes the handoff.
Consider a salesperson telling a customer, we should be able to support that integration
before your renewal. For the salesperson's personal agent, the next steps might be straightforward.
Update the account record, draft a follow-up, and create a reminder. But inside the organization,
that sentence implicates several responsibilities. For the sales folks, that sentence means for them a
potential way to secure the renewal, i.e. supporting that integration. For engineering,
that sentence means a possible delivery commitment involving uncertain work. For product, it means
a potential change to roadmap priorities, and for customer success, it's an expectation they may
have to manage. A multiplayer agent would need to recognize the sentence as a possible cross-team
commitment, and a judgment model could assess specific questions. Does this message imply a delivery
promise? Does the promise concern work outside the speaker's authority? Does it conflict with the
supplied roadmap, is there evidence that the responsible team agreed? The larger system could then
create a proposed commitment, identify the necessary owners, and request the missing decisions.
Now again, to reinforce costs and benefits, the big cost of Jev and this type of judgment model in
general, to the extent that this becomes a category, is that it is an incomplete category by
definition. It cannot do all the work that we currently have generative AIs do. Judgment models are
going to have to be part of a more complex model architecture, the type of the type of.
of model stack that we've been discussing for the last several months. The benefit, of course,
though, is that it can do this extraordinarily inexpensively. And because this sort of
judgment intelligence can be applied so cheaply, it means that done well, it can be integrated
incredibly deeply into the automated systems we're all building. This is obviously just the
first day of a very new concept and a concept which is significant enough that this very
competent team has spent two years working on. So obviously we're going to need to see how
it all plays out in practice. But it does have the feel when you dig
begin, of something both important and obvious, the type of thing that once it exists, we will
be surprised in the future that we didn't have it for so long.
Certainly, I'm going to be keeping an eye on this, and I will continue to look out for more
examples of how people are using it, as well as where companies are running into challenges
as they try to build these new types of systems.
For now, though, very cool stuff to go check out, and that's going to do it for today's
AI Daily Brief.
Appreciate you listening or watching, as always.
Until next time, peace.
