The AI Daily Brief: Artificial Intelligence News and Analysis - How Big a Deal is Llama 4's 10M Token Context Window?
Episode Date: April 8, 2025Meta’s new Llama 4 models have a massive 10 million token context window and a fresh architecture using a mixture of experts. Scout and Maverick are out now, and Behemoth is still training. Despite ...strong benchmark scores, many users report underwhelming real-world performance.Get Ad Free AI Daily Brief: https://patreon.com/AIDailyBriefBrought to you by:KPMG – Go to https://kpmg.com/ai to learn more about how KPMG can help you drive value with our AI solutions.Vanta - Simplify compliance - https://vanta.com/nlwThe Agent Readiness Audit from Superintelligent - Go to https://besuper.ai/ to request your company's agent readiness score.The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614Subscribe to the newsletter: https://aidailybrief.beehiiv.com/Join our Discord: https://bit.ly/aibreakdown
Transcript
Discussion (0)
Today on the AI Daily Brief, meta launches Lama 4 with a massive new context window.
Before that in the headlines, Mid Journey launches V7.
The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
To join the conversation, follow the Discord link in our show notes.
Welcome back to the AI Daily Brief Headlines edition, all the daily AI news you need in around five minutes.
Big model release day here on the AI Daily Brief.
Our main episode is all about Metas Lama 4, and Mid Journey also has released their first
new model in almost a year. Called V7, the model is obviously incredibly gorgeous. It's hugely
capable in both photorealism as well as stylized modes. It features things like voice prompting,
image personalization based on your own preferences and multiple speed settings. The new model doesn't
really introduce any novel features. It's just a better version of the mid-jorney experience
that has been so popular so far. Now, of course, the context for this is very different in the wake
of OpenAI's image gen release. Indeed, when that came out,
and was using a different approach than diffusion,
many wondered if the AI approach that underpins MidGernie was going away
to be supplanted by natively multimodal image generation.
Swix even tagged in David holds the founder of MidGernie and said,
I try not to drink hyperbole, but will Mid Journey go the same way?
Now that we have this in Gemini 4-0, I don't see how I ever go back to anything else.
David's one-word answer was Nah.
On the release of V7, he wrote,
this is an entirely new model with unique strengths and probably a few weaknesses.
We want to learn from you what it's good and bad at,
but definitely keep in mind it may require different styles of prompting.
So, play around a bit.
The reaction was honestly a little mixed.
Content creator FreeBatar wrote,
two-year mid-jorney power user.
Gotta say it, kind of disappointed.
Open AI set the bar sky high.
Talk to your image gen like it's your bro.
Mine equals blown.
MJ7 looks more realistic, but did we really need that?
Mid Journey Plus Magnific already nailed it.
might pause my sub to be honest.
Yavi Lopez, the founder of Magnific, added,
yep, I mostly agree.
But Mid Journey still has that super cool, artistic, aesthetic,
and richness of styles.
Though, to be fair, they already had that in version 6.1 in previous.
The problem is V7 doesn't really feel like V7.
It feels more like version 6.2.
Jamie Ortega spotlighted,
they knew it wasn't ready yet, which is why it wasn't released.
Then 4-0 ImageGen took off and they were forced to put out whatever they had to gain momentum.
It's just a retaliatory move and a poor one at that.
Professor Ethan Malik wrote,
I'm a big fan of Mid Journey if you want to do visually interesting AI images and have been using it since
2022. I like their new release, but the problem with the new V7 release today is that V6 was already
really good. Experimenting will be fun, though. Tatiana Siegelleva, however, the creative ambassador
at Perplexity doesn't know what everyone's complaining about, posting, my mind is blown exploring
Mid Journey V7, huge jump in quality, planning to have a lot of fun this weekend. Now, I've been
thinking a lot about Mid Journey ever since the release of the new image gen. On the one hand, the style and
aesthetic and quality of mid-journey continues to just be tops for me relative to the other image
generation models out there. And when I have any sort of deep creative or artistic project,
it's still my go-to. However, even before ImageGen, I had found myself switching almost entirely
to ideogram when it came to day-in-and-day-out type of usage. Now, a big part of that was that
ideagram had better text adherence so I could actually create cover images for podcasts and YouTube
videos and things like that, which is a primary use case for me. But part of it is that it just had much
better adherence to my prompt. It always felt to me like I was fighting with Mid Journey,
like it thought it knew better than I did how to make something cool. And so when I was willing to
let the AI kind of do its thing, it was still great for that. It's getting harder and harder
though now with ImageGen coming out, because not only now do those other tools have better
prompt adherence, you also have the ability for inline editing where you can just talk at it,
and it actually makes the changes you want. That is so transformative and so different across so many
different use cases. It makes the band of things that I want Mid Journey for more and more narrow.
I've generated tens of thousands of images on Mid Journey, absolutely love it, and I'm a person
who has no problem spending money on multiple different subscriptions. But even I myself am wondering
at what point I'm going to pause Mid Journey, because I'm just not using it enough anymore.
Mid Journey is also interesting just as a case study. The company is bootstrapped and highly profitable,
so it's not clear that they actually need to do anything other than maintain being a useful model.
Back in December, David Holtz wrote,
By VC standards, we should either conquer the world or die in a fire, and neither of these are
spiritually compelling to me.
I never wanted a company I just wanted a home.
At this point, we have a large and loyal paid community, we build tons of features for them,
and they're pretty happy.
We have enough revenue to fund tons of crazy R&D, and our models are still the best by
the metrics we care about, which is how the images look and how fun it is to make things.
We have a huge backlog of exciting things to make our models way better.
Zero risk.
We did this all with no investors.
Honestly, it feels like we are successful.
The next metric of success I think about most is a big well-funded R&D lab with cool people free to work on whatever they want.
Can we now build something that would make Baby David proud?
Can we now tell bold stories about a human future that people want to be part of?
I think we can.
And I think it's awesome that we have a company that's really pushing the boundary of doing it this way.
At the same time, I will be interested to see how durable that large and loyal paid community is, as all these things change around them.
Next up, at an event celebrating their 50th anniversary, Microsoft has rolled out the agents.
Co-Pilot is now able to handle internet use with agenic features allowing the model to book tickets,
reserve restaurants, and more.
The feature is configured to work in the background so you can keep working on other tasks
while having an agent run digital errands.
The Agentic Assistant also now has memory so can remember the user preferences between sessions.
Microsoft has also introduced a podcast generation feature similar to Google's audio
overviews, and Co-Pilot Now finally also has a deep research feature.
Now, of course, none of these features are pushing the limits.
Basically, they just represent Microsoft keeping up with their AI rival.
Then again, as we see Apple and Amazon and their struggles, being able to offer feature complete
agentic AI if these things really work, would still be a great deal better than some of the other
big tech firms are doing.
What's more, Microsoft AI CEO Mustafa Sullyman clearly sees the importance of agents commenting,
I think that this will completely change the way we use computers forever.
And speaking of Microsoft, the company is getting back in the game with an AI-generated version
of Quake 2.
They released a tech demo of the classic 90-shooter powered by its muse generative game model.
The demo is very basic, has blurry enemies, low resolution, and only replicates a single level,
but it's still a playable game generated frame-by-frame using AI.
Back in February, when Muse was first unveiled, Microsoft Gaming CEO Phil Spencer said,
You could imagine a world where from gameplay data and video that a model could learn old games
and make them portable to any platform where these models could run.
We've talked about game preservation as an activity for us,
and these models and their ability to learn completely how a game plays
without the necessity of the original engine running on the original hardware opens up a ton of opportunity.
Many claimed that this was not technically possible at the time, but the Quake demo seems to be a viable if limited proof of concept.
Researchers wrote in a blog post, much to our initial delight, we were able to play inside the world that the model was simulating.
We could wander around, move the camera, jump, crouch, shoot, and even blow up barrel similar to the original game.
Additionally, since it features in our data, we can also discover some of the secrets hidden in this level of Quake 2.
At the same time, they were careful not to oversell the technology.
They wrote that they didn't intend to fully replicate the original game and that the demo should be thought of as, quote,
playing the model as opposed to playing the game.
Derek Strickland wrote,
I'm trying the quake thing and Gen A.I. is just so weird.
It's like playing a dream where you turn around and things are different,
never-ending hallways, etc.
I know it's experimental, but it's still freaky.
And indeed, whereas most of the conversation around this
is just how useful it is for game preservation,
I think it's much more interesting as an early example
of what it's going to look like to generate games on the fly.
A big part of my thesis for how the world evolves
is more and more custom experiences,
and part of that is going to be, I think,
live generation of those experiences.
based on the user that's actually experiencing it.
This is a very small step in that direction, but a step nonetheless.
That, however, is going to do it for today's AI Daily Brief Headlines edition.
Next up, the main episode.
A quick note before we get into today's ads, for those of you who are looking for an ad-free
experience, we are now up on Patreon.
You can go to patreon.com slash AI Daily Brief.
Right now, the benefits of this are that you will get the episodes without ads,
and they will also come out a little bit earlier.
We'll be exploring things like community and additional content in the weeks to come,
but for now, I've heard from lots of you that you are looking for an ad-free experience,
and so again, it's patreon.com slash AI Daily Brief. Thanks as always for supporting the show.
Today's episode is brought to you by Vanta. Trust isn't just earned, it's demanded.
Whether you're a startup founder navigating your first audit or a seasoned security professional
scaling your GRC program, proving your commitment to security has never been more critical
or more complex. That's where Vanta comes in.
Businesses use Vanta to establish trust by automating compliance needs across
over 35 frameworks like SOC2 and ISO-2701.
Centralized security workflows, complete questionnaires up to 5X faster, and proactively manage vendor risk.
Vanta can help you start or scale up your security program by connecting you with auditors and
experts to conduct your audit and set up your security program quickly.
Plus, with automation and AI throughout the platform, Vanta gives you time back, so you can
focus on building your company.
Join over 9,000 global companies like Atlassian, Cora, and Factory, who use Vantage to manage risk
improve security in real time.
For a limited time, this audience gets $1,000 off Vanta at vanta.com slash NLW for $1,000 off.
Today's episode is brought to you by Super Intelligent and more specifically Super's Agent
Readiness Audits.
If you've been listening for a while, you have probably heard me talk about this.
But basically the idea of the Agent Readiness Audit is that this is a system that we've
created to help you benchmark and map opportunities in your organizations where agents could
specifically help you solve your problems, create new opportunities in a way that, again,
is completely customized to you. When you do one of these audits, what you're going to do is
a voice-based agent interview where we work with some number of your leadership and employees
to map what's going on inside the organization and to figure out where you are in your agent
journey. That's going to produce an agent readiness score that comes with a
deep set of explanations, strength, weaknesses, key findings, and of course, a set of very specific
recommendations that then we have the ability to help you go find the right partners to actually
fulfill. So if you are looking for a way to jumpstart your agent strategy, send us an email at
agent at besuper.a.i. And let's get you plugged into the agentic era. Welcome back to the AI Daily
Brief. Some exciting new model announcements to close out the end of last week. On Friday,
Meta revealed their new Lama 4 family of models.
As is the case every time meta announces a new set of models, there is a lot to dig into here.
These models feature all-new architecture, including multimodal functionality for the first time.
The models are the first to utilize the mixture of experts architecture that most recently has been seen in Deepseek.
It's an architecture that allows the models to access a subset of parameters within a larger model, making inference more efficient.
The Lama 4 family includes three different models.
Lama 4 Scout is a 17 billion parameter model with 16.
experts, which meta claims is the, quote, best multimodal model in the world in its class
and is more powerful than all previous generation Lama models while fitting in a single
Nvidia H-100 GPU.
Lama 4 Maverick has the same 17 billion active parameters, but includes 128 experts,
basically meaning it's a total of 400 billion parameters.
Meta states that the model is the, quote, best multimodal model in its class,
beating GPT-40 and Gemini 2O Flash across a broad range of widely reported benchmarks,
while achieving comparable results to the new Deepseek v3 on reasoning and coding at less than half
the active parameters.
Lama 4 behemoth is still in training.
It's set to feature 288 billion active parameters with 16 experts for a total of 2 trillion
parameters.
So meta here is taking the same strategy that they did with Lama 3, which is release a couple
of the smaller models early to get people excited and then release the biggest version of
the model a couple months later.
Now when it comes to Lama 4 behemoth, this will be the first time a model has reached into the
trillions of parameters that we know for sure.
and the first mixture of experts model of this size, so we don't really know how model performance
will be affected. Looking at costs, Lama Force seems to be pretty competitive. Inference service
provider Grock has the hosted model available already. Scout costs 11 cents per million input
tokens and 34 cents per million output tokens, while Mavericks prices are 50 cents and 77 cents per million
for input and output respectively. In that, both models undercut deep-seek Gemini 2.0 Flash
and Quen's QWQ32B. When it comes to benchmarks,
The new models look comparable to their peers. Scout outperforms models like Mistral 3.1, Gemini
2.0 Flashlight, and Gemma 3 on some benchmarks, while Maverick beats out GPD40 and Gemini
2O Flash on most multimodal reasoning benchmarks. Notably, neither of those models are a true
reasoning model utilizing chain of thought or test time compute. Now, one thing that's really
important to note is the context into which Lama 4 is entering. A couple months ago, we got this leak from
inside the company, which claimed that the meta-gen-AI organization was in panic mode.
The leaker wrote, it started with Deepseek v3, which rendered the Lama 4 already behind in benchmarks.
Adding insult to injury was the unknown Chinese company with 5.5 million training budget.
Engineers are moving frantically to dissect Deepseek and copying anything and everything we can from it.
I'm not even exaggerating.
Management is worried about justifying the massive cost of the Gen.A.I.org.
How would they face the leadership when every single leader of Gen.A.I.org is making more than what it cost to train Deepseek V3 entirely, and we have dozens of such leaders.
DeepseekR1 made things even scarier. I can't reveal confidential info,
but it'll soon be public anyways.
It should have been an engineering-focused small organization,
but since a bunch of people wanted to join the impact grab
and artificially inflate hiring in the org, everyone loses.
So this was the type of report that we were getting behind the scenes.
And in the wake of these announcements,
there is a lot of discussion about the feeling that maybe this released was rushed,
and that there might even be something more nefarious than that going on.
Min Choy writes,
Yikes, Lama4 benchmarks looked insane, but something feels off.
Reddit leak claims meta cooked it.
In the 24 hours following the announcement, as people started to dig in, they seemed to be
finding a fairly big difference in output between what meta was claiming and what seemed to be
the reality. TechCrunch writes, researchers on X have observed stark differences in the behavior
of the publicly downloadable Maverick compared to the model hosted on Elm Arena.
The Elm Arena version seems to use a lot of emojis and give incredibly long-winded answers.
Even more concerning was a Reddit post from someone who claimed that they were a meta-engineer.
The post they shared said this.
Despite repeated training efforts, the internal model's performance still falls short of open-source
state-of-the-art benchmarks, lagging significantly behind.
Company leadership suggested blending test sets from various benchmarks during the post-training
process, aiming to meet the targets across various metrics and produce a presentable result,
presentable in air quotes.
Failure to achieve this goal by the end of April deadline would lead to dire consequences.
Following yesterday's release of Lama 4, many users on X and Reddit have already reported
extremely poor real-world test results.
As someone currently in academia, I find this approach utterly unacceptable.
Consequently, I have submitted my resignation and explicitly requested that my name be excluded
from the technical report of Lama 4.
Notably, the VP of AI at Meta also resigned for similar reasons.
There have been a lot of people referencing this post without a ton of verification yet.
Bernie Tech wrote,
Lama 4 gamed benchmarks so hard LMAO, completely out of touch with reality and practice.
Andrew Allen summed it up this way.
He wrote,
META just dropped Lama 4 and score number two on LM Arena, beating GPT-4O and GROC,
but users are calling it garbage and vaporware.
Let's unpack the biggest benchmark controversy of 2025 so far.
The numbers look incredible on paper.
10 million token context window, 1417 ELO score on LM Arena, the second highest,
beating many top-closed models.
But something doesn't add up when users actually try it.
He pointed to a tweet from Harsh Varden that writes,
tried out META's Lama 4 for coding-related tasks,
found it super basic and almost useless.
Didi Das from Menlo Ventures writes,
Lama 4 seems to be actually a poor model for coding.
ELO maxing on L-M Arena doesn't create the best models.
Back to Andrew, he continues.
The disconnect is stark on paper second highest on L.M. Arena leaderboard,
in practice super basic and almost useless for coding.
Marketed revolutionary capabilities,
reality struggling with basic construction following.
The most serious allegation,
META may have submitted a different model for benchmarks
than what's publicly available.
This raises major questions about benchmark integrity.
User reports highlight specific failures, freezing when run locally on Macs, poor coding
capabilities compared to Claude and GPT, inability to follow instructions consistently, declining
quality with longer contexts.
Many users are calling the 10 million token context window marketing fluff that doesn't translate
to better performance.
And we'll be coming back to that 10 million token context window in just a minute.
But Andrew also points out there are bright spots.
It's fast 512 tokens per second on GROC, cost effective, improved vision capabilities over
Lama 3, and open source enabling community.
innovation. Ultimately, he writes what this reveals about AI development, benchmark scores do not
equal real-world utility, the gap between lab performance and practical use is widening, and users
increasingly value reliability over raw specs. Obviously, we'll get a lot more information in the days to
come, and even if there hasn't been nefarious behavior here, there's still some pretty big gaps
between the marketing promise and what people are actually finding in practice. Outside of all
that dubiousness. The big point of discussion and the thing that has everyone's mind racing
is that Lama 4 Scout theoretically features a 10 million token context window. Until now, Google's
development of a functional million token context window for their Gemini models was state of
the art. It was five times as large as the same class of models from OpenAI and Anthropic.
Now, ultra-long context windows are a really big deal for a variety of use cases, for example,
for coding assistants. The longer the context window, the more a coding assistant is able to ingest an
entire code base to be understood all at once. For agents, long context allows four much longer
tasks to be completed before losing coherence. Meta demonstrated the performance with a retrieval
needle in a haystack test across 10 million lines of code. Scout didn't have a single failure
across their testing. Now, independent benchmarks weren't anywhere near as impressive. And yet still in this
case, most of the conversation wasn't so much about Meta and Lama 4 specifically, but about
what the implications are as the tech improves. Representing around a million variants on this take,
Marvin Aziz, the community manager at Lindy wrote, Rag is dead. Why bother with a knowledge base when
you can shove 10 million tokens into a context window and call it a day? Rag, of course, refers to
retrieval augmented generation, which is the process of hooking up an LLM to a database or knowledge
source to search up any information it might need. Then again, the opposite take was just as prolific.
With AI evaluations designer Hamil Hussein writing, Rag is dead post our annoying AF. R is
retrieval and AG is the LLM. This means you think retrieval is dead. Seriously, you think retrieval is
keyword search, metadata filtering like dates and users, grep and other filtering are retrieval.
Good luck without retrieval. Charles Fry writes, rag is dead is also the sort of thing only said
by someone who has never run LLM inference themselves, let alone been on the hook for cost and latency.
Enigmatically, Swix writes, unpopular opinion right now, but Lama4's 10 million token window
will finally actually end the long context versus rag debate, but not the way the other guy is
thinking. For those trying to tow a more middle of the line, they basically point out that we just
don't know enough yet to declare the end of rag or really understand how well long context windows
are going to work. Near Cyan wrote, I haven't played with the Lama 4 series, but needle in a haystack
is woefully insufficient to know the strength of a context window. If you want needle in a haystack,
we have grep for that. Grep is a Linux command for searching databases. OpenAI co-founder
Andre Carpathy falls into the category of wanting to believe, adding, my reaction to when
reading all the rag as dead tweets earlier today. Huge amount of optimism that the context window
is also usable in practice for real problem solving and not just in theory. Could very well be true,
I just don't super know. Now, the community with the most enthusiasm about an ultra-long context window
was the vibe coders. Plain game creator Peter Levels wrote, this is insane and makes it
finally possible to vibe code up to giant code sizes. The limit just weeks ago was context window.
AI would get lost once your vibe-coded gamer app became too big. Imagine an AI with memory loss
that starts breaking stuff. With 10 million tokens, there's practically no
limit, really quite big for vibe coding and another big hit for the perpetual naysayers.
AI consultant Sasha Lecti added,
At this point, you can throw the entire documentation of multiple libraries with examples
and your project into the context window, and it will handle tasks in one shot.
In my opinion, the bottleneck now lies more on the agenic side.
These systems need to operate without me babysitting them.
Still, there were many trying to harsh the vibes with practical issues of using a 10 million
token context window.
They assumed that loading that many tokens would be painfully slow,
and questioned whether a Gemini 2.0 flashlight class model would be,
be up to the task of generating functional code. Developer Nick Dobos rebutted,
lazy take, use it to ask questions and plan, use the high-tier models to write the actual code,
not hard. His point being that even if the model isn't really up to writing code or even developing
a plan, simply creating an outline of a large code base for use in another LLM is a new feature
that hasn't previously been accessible. LinkedIn co-founder Reid Hoffman had a less
combative take, posting, spending the day playing with Lama 4. One of the many interesting things
that massive context window is a game changer. I don't think it's a
the end of RAG, but for a surprising number of workflows, the long context alone is enough.
And I think this is an important point. Ultra-long context doesn't have to be perfect or
completely replace RAG to be a really big deal. To the extent that it holds up at all, this feature
could unlock a huge range of functionality that wasn't possible before. Orchestration Platform
Oblix commented that this is just one tool in future workflow design, writing,
long context doesn't replace RAG, but it absolutely shifts the tradeoffs. For structured contained
workflows like contracts, single docs, or chat history, context alone is similar.
simpler, faster, and good enough. Rag still shines when you need external dynamic or filtered
retrieval. The future probably blends both. Long context for memory, rag for knowledge access,
orchestrators for choosing the best tool in real time. Matthew Berman zoomed out even more.
While noting a ton of Lama4 shortcomings, he added, here's the strategic insight that everyone's
missing. Meta's 10 million token context window isn't about today's performance. It's about
signaling tomorrow's direction. They're showing us a future where AI doesn't just retrieve knowledge
but transforms your entire knowledge base into manipulable working memory.
Zuckerberg understands the truth Google accidentally leaked.
Close source AI has no moat.
Foundation models are becoming commodities faster than anyone predicted,
and meta is accelerating this transformation.
Meta strategy becomes clear when you connect the dots.
Commoditize foundation models through open source,
make context the new competitive battleground,
force innovation up the application layer,
leverage their massive social graph advantage,
and ultimately create an open ecosystem
where social and application data become the true moats.
Still, at the end of the day, as much as they are helping shape the conversation,
it's hard not to view this release so far as a disappointment.
Professor Ethan Malik even commented that their flagship model doesn't stack up,
writing,
looks like even Lama Behemoth doesn't come that close to Gemini 2.5,
so no open model parody with the state-of-the-art-art-en-closed models.
We will see what happens when people slap a reason around Lama, though.
It doesn't seem like they're launching with one.
And indeed, this was another common take.
Andrei Borkoff writes,
If today's disappointing release of Lama 4 tells us something,
it's that even 30 trillion training tokens and 2 trillion parameters doesn't make your non-reasoning
model better than small reasoning models. Model and data size scaling are over. And so as we wrap up here,
I'm not yet exactly sure what to make of this. On the one hand, it feels a little rushed. It does seem
like the deep seek pressure is getting to meta. At the same time, given that they are taking an open
strategy, the consequences of releasing earlier a little bit less severe for them than perhaps for other
companies. If it's cost-effective, better than some of the things that people had access to before,
there's still going to be a lot of developers building on it. Indeed, holding aside wanting
every single model to break the mold every single time, ultimately for developers, this just
represents another set of choices, which in a very fast-moving environment is nothing but a good thing.
For now, though, that is going to do it for today's AI Daily Brief. Appreciate you listening or
watching as always, and until next time, peace.
