TBPN Live - Model Mayhem, GPT-6 Astra, Why Nvidia Bought Hugging Face | Diet TBPN
Episode Date: September 3, 2026Diet TBPN delivers the best of today’s TBPN episode in 30 minutes. TBPN is a live tech talk show hosted by John Coogan and Jordi Hays, streaming weekdays 11–2 PT on X and YouTube, with ea...ch episode posted to podcast platforms right after.Described by The New York Times as “Silicon Valley’s newest obsession,” the show has recently featured Mark Zuckerberg, Sam Altman, Mark Cuban, and Satya Nadella.TBPN is made possible by:Ramp - https://ramp.comPublic - https://public.comCisco - https://www.cisco.comConsole - https://www.console.comCrowdStrike - https://www.crowdstrike.comFigma - https://www.figma.comMongoDB - https://www.mongodb.comNYSE - https://www.nyse.comRailway - https://railway.comShopify - https://www.shopify.com/Follow TBPN: https://TBPN.comhttps://x.com/tbpnhttps://open.spotify.com/show/2L6WMqY3GUPCGBD0dX6p00?si=674252d53acf4231https://podcasts.apple.com/us/podcast/technology-brothers/id1772360235https://www.youtube.com/@TBPNLive
Transcript
Discussion (0)
It's model mayhem, folks. It's model mayhem. We got tons and tons of new AI model releases. It's a great week to be into AI. I think all the lab leaders, they got together. They said, you know what? People just love AI. Let's give them more. Even more. Let's all team up to launch new AI for them the same week so that everyone has something just to be happy about. That's right. We got Anthropic, Fable 5.1. We got Mew Spark 1.3. We got Gemini 3.8.4.
Flash. Open AI's T's an Astra is a GPT6. There's a six where the S goes in one of the videos.
People will figure it out. But the model mayhem is continuing. Everyone got back from their long
summers, their vacations. They said, we got a, we got a long something new.
It's time to ship. We got to update this stuff. And Grock, Claude. Yeah. Oh, yeah. They
open AI. We're all down this morning. So much demand. People thought Astra might have escaped.
Kind of a deflock moment for AI maybe.
What do we think is going on?
Possibly.
Do we actually understand why all of the different models went down at the same time?
Because it's easy if it's like AWS went down and it took down a bunch of stuff.
Yeah, but why is Gemini?
I think it was U.S. East 1.
Why is Gemini down then?
No, no, Gemini wasn't down.
Oh, Gemini was never down.
Oh, okay, okay.
Amazon, clearly very critical to the global internet.
Good luck to the folks over at Amazon that are fighting the good fight.
A couple big announcements.
One, which we got.
Chamath's birthday.
It is.
Big 5-0.
Should we sing full happy birthday?
I think full happy birthday.
Happy birthday to you.
Happy birthday to you.
Happy birthday, dear Chamas.
Happy birthday to you.
Fantastic.
And almost all in conference is coming to Los Angeles, I believe, a couple weeks.
Very exciting.
Sign up.
Get for our invites.
So wait for our invites.
I think we, I think we, I think,
invited to pay and go if you want to go. Good to know. Equally important.
TBPN's Road to Christmas. How many days you got? 112 days out. 112 days out.
112 days out. I got to say it's feeling like it's going a little slow. Yeah. I wish there
I wish I was feeling more pace. Okay, think about it this way. We're only 13 days away from double
digits. That's a big moment. That's a moment everyone's going to be talking about on the road to Christmas.
When we get to 99 days till Christmas, that's what you can start a countdown.
That's big.
If you get a really big advent calendar, you can basically start that.
Let me give you the roundup on the model mayhem that's going on.
So, Anthropic launch, Claude Fable 5.1, alongside the restricted Claude Mythos 5.1.
Google released Gemini 3.8 Flash and a cybersecurity focus version.
That's good news.
Meta released Muse Spark 1.3.
So it might not be that much of a surprise that Anthropic seems to have the strongest
model of the three with Fable 5.1 scoring 66 on the artificial intelligence in index. That is the
bar chart that everyone has been posting in this cycle. It feels like we're sort of maybe getting
to the end of the benchmark era. It feels like when these models are released, it's much better
to solve a novel math problem or do something else. The Pelican on the bicycle is still one of my
favorites. Yeah. I love that one. But it changes everything.
It does. But the benchmarks, you know, they've been accusations of bench hacking, odd, hard to interpret.
I just think people are very, very low trust in benchmarks. They do. At this point, everyone has had enough experience using various models.
They have their own sort of internal benchmark. Yeah. And so the demos of like, I built this game, I did this thing with it. And then also just the trusted voices of people who, you know, use a bunch of these models and they kind of give you the breakdown of what they like, what they don't. That has been where people lean a lot more. But the artificial intelligence.
The artificial analysis intelligence index, this bar chart that you see, has been a good way to
kind of compress down a bunch of benchmarks into one meta benchmark.
So Fable 5.1 get the highest result ever on the index.
The score is also ahead of Opus 5, which got 63, and Fable, which got a 62.
Anthropics says Fable 5.1 is also cheaper and more efficient made possible by an improved
caching system.
It should make ordinary workloads 25% cheaper and Long Horizon Agentic Jobs 45% cheaper, the company
says.
That's good news.
And as Ben Thompson pointed out, Anthropics also sort of dropping its no zero data retention policy, which there was a whole news cycle around a few weeks ago.
People were saying, you know, why is Fable not taking off an adoption?
It's a really a great model.
And there were a bunch of different explanations.
One of them was companies demand zero data retention.
They don't want closed source AI labs to be hoovering up their private information.
Yeah, and I think they said they're testing functionality that will allow for.
Data retention, but it's on servers and infrastructure that the company owns.
Yeah.
So you do keep some of the data.
You still are monitored for hostile usage, but it's not going straight into Anthropics databases.
So Alex Karp and Satinaele both warned against this idea that companies should have data sovereignty.
So the policy is going to be replaced.
The no zero data retention, no ZDR is going to be replaced with something called EFS, Enterprise Frontier Safeguards.
and that may have contributed to lower fable adoption among enterprises,
and it sounds like it was a direct response to user feedback.
So good news that the people spoke and the companies listened.
So over in Google World, Gemini 3.8 Flash is the company's third flash release in six weeks.
They are flashing out these flash releases.
And scored 73.7% on Deep Swee just behind Opus 5 and competitive with models that cost several times more.
independent testing gave it a 59 intelligence score, which isn't the absolute frontier,
but it's a great result for a model generating roughly 300 tokens per second.
So very quick, very cheap, and very good at coding, at least on this particular benchmark,
deep sui.
We'll see what adoption looks like and where enterprise spend goes.
Aura Karazian has some very interesting data from the Ramp Economics Lab.
You can go check out.
He also has a new post that's very interesting that we can talk about in a second.
But last model, meta-muse Spark 1.3, did very well on benchmarks, scoring 75.4% higher than Gemini 3.8 flash on DeepSuite, beating both Opus 5 and GPD 5.6 sole. It didn't sweep the board. Opus still beats it on several professional work and computer use evaluations, but it got 62 on the intelligence index, which is only behind the newest Claude models. Tons of stuff to think about and discuss here.
So the interesting post from Ara Karasi, and I don't know if we have it in the timeline, if we can pull it up,
But he was saying that there's a lot of concentration in the enterprise AI revenues right now.
Open AI Anthropic, 80% of their enterprise revenue comes from just 1% of the companies.
And I was like, 1% that seems crazy.
And he notes that this is uncommon for software categories.
Like if you look at CRM, if you look at databases, if you look at all sorts of different software spend,
typically you don't see as much concentration.
You don't see 1% driving 80% of the spend.
And I was wondering about this.
And so I started looking up like, where else do we see this type of inequality,
if you can call it that, this distribution, this power law?
Power laws are everywhere.
But where else does this exist?
And you might go to hiring.
Like, is AI a drop-in replacement for hiring?
Is it going to be proportional to hiring?
And in fact, the top 1% of biggest companies in America,
they do hire a ton of people.
The top 1% of American businesses employ 65% of the total workforce, not 80%.
But interestingly, the top...
Concentration risk there, John.
65% of jobs are tied to just 1% of companies.
There is.
Yeah, there is.
And I mean, yeah, you definitely see that.
Although the top 1% companies tend to be pretty, like, Lindy, you're talking about...
No, no, I know.
I'm joking.
Yeah, the government and whatnot.
But the interesting 1%, 80% core.
correlation comes from sales. So the top 1% of American companies by sales generate 80% of total
revenue. And so there's this weird dynamic where I don't know exactly how correlated it is,
how causal it is, but there is an interesting dynamic there where it feels like if you
look at the total AI spend, it's around 150 billion a year, something like that, and then you
look at total revenue for all U.S. businesses, AI,
is roughly a quarter of a percent of total U.S. business revenue. And it tracks fairly closely to
the revenues of those individual firms. So you see that one percent of the top businesses generate 80
percent of the revenue. They also spend 80 percent. They also generate 80 percent of the AI revenue.
And so there's this interesting dynamic where because Enterprise AI, particularly, you're not going
to be on the $20 plan. You're not going to be on a $200 plan. You're going to be
consumption-based, and you're going to look at it a lot more like a marketing line item
that's proportional to your revenue, potentially. That's at least one interpretation of this.
Another fellow over at Ramp said that this is a roar shock test for how you feel about AI.
Either you look at this and you're like, it's great, or you look at this like, it's over.
But fun, fun, fun, fun chart to dig into. Anything else on this, you guys?
Read anything else? Nah. Let me tell you about codex.
I just was realizing...
Codex is a powerful workspace for getting work done.
with AI agents, whether you're writing code, analyzing data, creating content, or automating
business workflows.
Codex helps you move projects forward from start to finish.
Yeah.
We forgot to cover.
Pablo Torre joining at 1130.
11.
To talk about the clippers and Kwai Leonard.
So these are people that go, they watch live streams and they clip them and they put them out on social
media.
Yeah.
It's a big boom.
I think that's what the team, why they name the team that.
Yeah, kind of an homage.
Because there's a lot of clipping that happens in L.A., TikTok clips, Instagram clips.
So they call them.
But yeah, this whole Balmer-Kauai Leonard thing, you had an interesting pronunciation of Kauai's name earlier, because I don't think you'd ever heard of it.
Well, I was calling him Steve Balmay.
I dropped the R because I thought it was friends.
Oh, nice.
No?
Yeah.
So the other big story in the news, the open source community is stronger than every.
You saw three, four, who knows five, closed source releases.
All the big labs are duking it out.
Meanwhile, Nvidia is going even bigger on open source with the $13 billion acquisition for Hugging Face.
All over the timeline, this was leaked a couple weeks ago, I feel like, rumored.
What are you laughing at?
No, just every single day, I would see a headline about Nvidia hugging face.
Yeah.
Okay, now it's official.
Yeah, no.
Today is the day they can talk about it.
They can explain it.
And there's a lot that makes sense.
There's not too many questions.
It's just a great outcome generally.
But it's an interesting story because it's a true 10-year overnight success.
And they're having fun with the acquisition price.
And Vidi agreed to pay $12.9303 billion for Hugging Face,
which just happens to be the exact decimal code for the hugging face emoji,
which of course is the icon used by the company.
And then also, if you take that number and you turn it into a color code, I think you get a green that sort of hints it at NVIDIA.
So they're having fun both ways.
Like, yeah, this was always in the plan.
Symbolism.
It's just funny to be having fun with a price this big.
I remember the Instagram acquisition and the idea of a billion dollar outcome is being insane during the social media boom.
And now we're seeing like deca corn liquidity events every couple weeks.
I'm, of course, thinking of open router.
It's honestly an incredible time to be investing in AI seven years ago.
Yes, yes.
Three to seven years ago was an amazing time to be investing in AI.
It was.
And so obviously, this shouldn't come as that much of a surprise because Jensen has been probably the loudest voice on open source.
He put out that open letter that everyone signed on to.
And he wants to maintain Nvidia's dominant position in the AI ecosystem.
both selling chips to close source labs, who might wind up making their own chips as well,
and there's a whole tug of war there, but for open source AI development, he wants to be
the place where developers and companies go to pick their models and then hopefully rack them
on Nvidia GPUs. The simple distillation is just Hugging Face is the GitHub of AI. It's a little
more complicated than the Microsoft GitHub deal, but it still makes a lot of sense. So HuggingFace
doesn't own the smartest models. They don't even try to build them. And there's some interesting
financial dynamics there about how capital efficient they were because of that decision. But they
created this nexus for people to upload, discover, test, modify models. And the numbers are good.
They have over 18 million developers, 200,000 companies using the product, 3 million models, and
over half a million datasets. And so they have certainly created this vortex of activity that's
really valuable. How strong is the network effect? It's, you know, it's certainly cooking and it's
certainly driving a lot of value here. So the interesting thing about Hugging Face is that it did not
start as an AI GitHub for AI. It started as a completely different idea. The founder worked at a French
computer vision startup called Moodstocks that was eventually acquired by Google. And in 2016,
he teamed up with two co-founders, one who was a mathematician and the other one who was a scientist who
had worked in patent law, apparently. And they started building an AI that could basically talk about
everything. So this was post-Syri, post-Alexa, but instead of focusing on like, tell me the weather
and be an assistant, you know, set a timer, he wanted just to be able to talk to you.
Still not solved, by the way.
Wait, which one?
Siri.
Yeah, yeah, yeah. So, yes. It's pretty good at setting timers.
But so the goal was to build something like a tomogachi, something very cute, hence the hugging face
icon. A funny, emotional, digital friend targeted teenagers. The app let users name the bot, text it,
send selfies, trade emojis. It was explicitly marketed as an AI best friend for bored teenagers.
And they scaled it. I think this is pretty significant. It was doing a million messages a day.
They had more than 100 million messages in total by 2018. That seems pretty significant. That doesn't seem
like you're languishing in the app store with no downloads. Because how many messages
a day can a bored teenager possibly put up with an AI agent.
Even if it's like a thousand, you still have, I guess, a thousand users, power users?
I don't know, probably the average user's doing 20 messages of the day.
So you're seeing pretty significant adoption.
And so they were able to raise a series of financing rounds.
The big one that grabbed headlines was Kevin Durant.
Durant for three.
2018, this is pre-GPT3, and it didn't become a durable consumer business.
So in 2018, Google released BERT, which was sort of the first language model, very primitive, but people were really excited about it.
But it wasn't delivered just as weights that you could download on the internet.
It was delivered as a paper from Google.
And the paper was implemented in Google's TensorFlow framework.
And people like Pytorch.
So the Hugging Face team converted Burt from TensorFlow to Pytorch and released the conversion for free.
And so developers really like that.
and that became sort of like the initial go-to-market flywheel for developer adoption.
And eventually they added more and more models, eventually thousands.
Now I think they have millions of models, which is sort of crazy.
But when you think about all the forks and fine tunes, it makes sense.
Eventually the team stopped trying to build this one application and focused on building tools,
became the picks and shovels trade.
So instead of trying to pick a winner, you just host every model, they became the Switzerland of AI to some degree.
And so the flywheel started compounding.
more models, more developers, more model creators, more companies.
And it was the same basic network effect as GitHub.
You know, GitHub was the default home for open source software.
Hugging Face very quickly became the default home for AI models.
Over time, Hugging Face grew from a code library to a place where developers could publish models,
version them, attach datasets, discuss changes.
And they even allowed them to build demos.
Huggingface eventually launched a Spaces product where you could demo the different models.
Companies can maintain private repositories, same GitHub strategy.
So 2019, Lux comes in with $15 million.
Then, yeah, Lux got in early, Series A, 15 mil.
They also came back for the Series C in 2022.
That was $100 million at a $2 billion valuation.
Sequoia and Koto were in that round.
There was also a Series B in 2021.
That was 40 mil.
And then the big step up was in August of 2023.
Hugging Face raised $235 million at a $4.5 billion valuation.
and it's a murderer's row of potential acquires.
You've got Salesforce, Google, Amazon,
Nvidia, AMD, Intel, Qualcomm, and IBM.
So you're, you know, it's not like they were doing a roadshow to sell the company,
but it's very much like we want to be the Switzerland of AI.
We want good partnerships with everything.
We're going to be chip agnostic.
So, yes, we have Nvidia on our cap table,
but we also have AMD and Intel and Qualcomm.
So, you know, you can count on Hugging Face as being like an independent place.
we're not purely Nvidia backed, which is maybe one of the things they'll have to deal with now,
but they're not purely Nvidia backed at that time.
So it's very much like, oh, yeah, we'll host a model that runs well on AMD.
We'll host a model that runs well on Nvidia.
We'll host models that are from Google or from Amazon, et cetera.
And so it looked like this like peace treaty moment from the major AI infrastructure companies.
You get everyone around the table.
Everyone's aligned with the mission.
And Hugging Face becomes this neutral territory where it's supported competing clouds.
chips, frameworks, models, no single company could control the platform. And so even though they did a number
of rounds, Huggy Face, I'm going to say only raised under 400 million, which is a lot of money, but not at a $12 billion
outcome. And it's pretty small considering the outcome. And it was very capital efficient because they
weren't actually buying chips or serving models directly. And they had this flywheel that sort of
spurred growth through the network effect naturally. So not a lot of cost in the business. They
became profitable in 2025, still had half the money that they raised. So,
NVIDIA came in to offer $500 million, late 2025, at a $7 billion valuation, but they turned it down.
Whoa.
Turned it down. Because he said, hey, if we're going to go deeper with one particular area,
it's got to be the whole shebang. And so that's what wound up happening. So pretty high revenue
multiple. My question is, I wonder what Jensen's vision for Hugging Faces. They do offer model routing
and, you know, they rank a bunch of inference providers.
Is this something that could we see Huggin' Face
and OpenRouter and Ramp's Router
competing more and more?
Yeah, sure. I think it's two things.
I think one is close source is already its own business line,
sell chips to the labs, but the labs are building A6.
They're doing a lot of stuff,
and there's this whole back and forth tug of war there.
But on the flip side, you have open source,
which is continuing to grow,
and if you can be sort of the front door to that,
and then say, hey, you found your best model,
your framework on Hugging Face and now you're ready to go and buy chips or by inference or by
compute and Nvidia is right there. That is a very logical flow and anything that they can do
to make open source powerful, exciting, a place where you can build a career, have a great
outcome. I think that's beneficial to Nvidia because a lot of people will see this and say,
yeah, like you can go and build a great company and having a fantastic outcome. I mean, the retention
packages are apparently a billion dollars for a pretty small team. And so if you're, if, if, if, if, if, if, if, if,
video is just trying to send a massive signal to the world that you can make it at in the open
source world, that's a really good signal to send. And it feels like it's landing loud and clear,
especially today. What else is on your mind regarding hugging face and invidia?
People are still waiting for an official launch from, uh, of, of astra. Lisan al-Gaibe is
sharing some Astro benchmarks.
Arc AGI3, 98.6.
He had previously said,
are you ready for a nuke to hit Arc AGI3?
So almost fully saturated.
Frontier math, Tier 4V2 gets a 97.6.
Deep Swee, 74.1.
Explate bench also saturated at 100%.
So, it seems pretty good.
Let's go through what else is in the timeline.
What is John Palmer saying these days?
He says, I actually think Snapchat for work
might be a good idea in today's big companies. Work platforms, work communications platforms have always
followed what teenager were doing 10 years ago. I did use IRC when I was a teenager.
Unkst. Tyler has to look it up. He doesn't know what internet relay chat is. Wow.
For what it's worth, I didn't use IRC either. No, did you use AOL? You didn't use AIM, AOL Instant Messenger?
Wow. Youngman. What was your first community?
communication platform on the internet email email email i remember i remember being inundated with emails
i remember there was like a summer there was a summer he's laughing at me because i had a hotmail wow
well i i just i remember spending a summer as like yeah not even a teenager yet just being super
stressed about my inbox because like every kid had just started using email yeah and so they were
just sending these super long emails and i would and i would be thinking i'd just be like playing outside in
the grass and thinking man i got i got to i got to check my email do you think there's
anything actually to this? It sounds like he's being serious. In the age of AI slop, the most efficient
form of communication is just short videos of yourself speaking. What do you think? I would love to use
Snapchat for work. So if we said, hey, as a team, we're going to communicate through short selfie
videos. Yeah, with the filters. With the film? You got to have the dog filter or whatever, the hot dog
filter. He says he's serious. A two-minute demo video. Yeah, I guess in terms of actually just taking a video
of your screen showing people what you're working on.
You know, there's a lot of different things you can do.
It's official hot bot summer is over.
Super Grock is saying goodbye to companions.
I remember we were debating, you know, there's obviously a lot of people that had
ethical concerns about AI romantic companions, but we were more discussing just like,
is there actually a business here once you break the seal of like, okay, we're doing it,
he did it, would it actually be successful?
because replica has seemed to get to scale.
There's been other products that have played in this world
and seemingly reached adoption and scale.
But this is the same thing as the SORA discourse
where everyone was caught up in SORA is either going to be the most powerful thing ever
and we're not going to be able to stop watching it
or it's going to be good and so fun and it's going to be amazing and dominant.
But no one was counting just like, oh, it might just go away in six months.
and is the same thing here.
I don't know.
I think I was saying it was going to go away.
Just because I thought it was going to be a tool.
Yeah.
And I think we benchmarked the market for this.
Like if you look at other romantic stuff,
there's just not a trillion dollars of revenue.
When Elon first said, okay, I'm going all in on adult entertainment,
I did think it was a potential path for GROC to get into the single digit billions.
I probably overestimated.
overestimated the market there.
But it felt like one of the plays that he had to get back in the game at the time.
But again, this was last year.
And back then if you could get into the single digit billions, you were doing pretty well.
And then the game obviously, you know, really shifted to energy.
Yeah, it makes sense to wind it down, focus on enterprise software.
That was what was in the SpaceX S-1 was the, what was it, $13 trillion market he was going after,
or something like that.
It was a shocking, shocking number,
maybe $20 trillion or something.
Absolutely huge, huge numbers, and it makes sense.
There's a lot of value in the enterprise,
much less in this controversial topic.
Well, fish.ad audio ran a bellboard campaign.
We love out of home.
This one, sort of confused people.
Heshi Brody says $100 for anyone
that can explain what this company does
without looking it up.
This is on the New York City subway today.
And we'll read it to you, and you can take a guess.
Fish.orgio says, we put voice AI on a silent sign.
You see the problem.
Is this one of those jokes, though?
I think it's a joke.
I think this is an 11 labs competitor.
But is this a real ad from a real company, or is this a prankster making fun of tech ads?
I think it really makes you think.
It really makes you think. It made me go to the website.
Okay.
They make text to speech, speech to text, audio, separation, voice changers, translation.
And they partner with global innovators, John.
They're working with Hey Gen, retail, a bunch of games companies, it looks like.
Clout Kitchen.
You ever heard of Cloud Kitchen, John?
No.
You ever been in the kitchen cooking clout?
Okay, so isn't it like they put voice AI on a silent sign?
You can't like if you have a billboard, you can't listen to it, but their product is audio.
Oh, so if you can't listen to the billboard.
That's the problem.
Yeah.
Right?
I guess I would expect your explanation is a problem.
GPT6 Astra has landed.
The blog post is up.
You can go check it out on openaI.com.
To the stars.
To the stars.
And the headline, I believe, good performance on a bunch of things.
but the RKGI 3 number is crazy.
The score is 99.9%.
So it feels like they just beat that.
They just beat RKGIB3.
So that's the video games,
the ones that Tyler was briefly.
Globally ranked.
Globally ranked.
You're out of a job, Tyler.
Say goodbye.
Hang up your Arc AGI V3 hat.
Don't worry.
The team over at Arc AGI is working on V4.
We're going to move the goalpost soon.
We're going to move the goalpost soon.
But this is very impressive if you've played around with Arc AGIV3.
It requires some real creative thinking to actually learn how the games work, how to be efficient.
What else is sticking out?
Has anyone been able to monitor the timeline at all?
Very chaotic launch.
So I think it's hard to get a real reaction yet.
Tyler, what are you seeing?
Yeah, I mean, there's still no actual post on opening eyes like X account.
It's just the blog post that's now live.
Okay.
But, I mean, there's a bunch of benchmarks in there that are old.
Terminal Bench Science.
We're going over into the World of Science.
GPT6 Astra scored 64.6% on reasoning effort max costs $26,26.
I'm sure there will be a lot more.
Also, the exploit Jim Honeypot, lower is the better.
GPT6 Astra, 0%.
So good performance there.
The world's best computer use model.
I'm putting that to the test ASAP.
Let's see if it can 1V1 me on Rust
because if it's good at using a computer
should be able to no scope, right?
That's the bar.
That's where my goalposts are.
Agents' last exam, good performance.
So you can go check it all out.
Sam is going to be on Bloomberg in just a few minutes
over with our buddy at Ludlow.
So that would be fun.
So we'll be digging into this
and we have some special guests lined up
to talk more once we've been able to digest,
take it for a spin,
and have a lot of fun with it.
Mike over at Arc is moving the goalposts.
Moving the goalposts.
They're moving.
Thank you.
Let's do it.
Thank you for not waiting to move the goalposts either.
You know.
He said, Astra is the new state of the art on ARC, AGI 3.
It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising.
That said, we lack evidence to call this AGI yet.
While we are still studying the human capability gaps, we believe open-ended invention is unsolved,
and this will form the new basis for ARC AGI4.
So now it's like you've got to invent new physics, new science.
You've got to go to that.
You actually, Astra has to actually go to the stars.
Yeah.
He really said, what have you done for me lately?
I love it.
I love it.
It's AGI when Mike says it's AGI.
It's a good time.
Anyway, very interesting.
You can dig into his post.
He gives a lot more context.
there about RKGIV3, benchmarks, Astra, you can go check it all out.
And of course, we'll be discussing it tomorrow.
We have a bunch of special guests.
So have a great day.
We'll see you tomorrow.
Leave us five stars on Apple Podcasts.
It's a Spotify.
Sign for a newsletter at TBPN.com.
And we will see you tomorrow.
