Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 876: The Most Important AI Model You’ll Probably Never Use That Just Dropped (Replay)
Episode Date: October 7, 2026You've probably never heard of Inkling. It's the newest (and first) model from Thinking Machines Labs, and it could very well be a small snowball that picks up major momentum in today'...s enterprise AI landscape. If you haven’t heard of Thinking Machines, they’re led by Mira Murati, the former CTO at OpenAI. The big bet with Inkling? The future of AI could be using smaller models fine-tuned and optimized for smaller tasks. Will it work? Tune in live as we dive in. The Most Important AI Model You’ll Probably Never Use That Just Dropped -- An Everyday AI Chat With Jordan Wilson (Replay) Newsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Inkling AI Model Launch OverviewThinking Machines Lab Leadership HighlightInkling's Multimodal and Agentic CapabilitiesOpen Source vs. Proprietary AI ModelsEnterprise Procurement with American AI ModelsAI Fine Tuning as a Service (Tinker)Benchmark Scores: Inkling vs. Frontier ModelsCustomization and Model Shopping for EnterprisesAI Token Costs Driving Model EfficiencyBridgewater Case Study: AI Model CustomizationFrontier Models Enabling Efficient Fine-TuningFuture Trends: Specialized Small Language ModelsTimestamps:00:00 Inkling: A new AI model release05:43 Inkling AI model details09:08 China's dominance in open source AI11:48 Launch and model updates discussed15:21 Concerns over using Chinese open-source models19:06 Training smaller AI models20:22 Using GPT for AI Model Training23:54 Predicting Rise of Small Language Models28:38 Choosing the right AI modelKeywords: Inkling, Thinking Machines Lab, Meera Muradi, former OpenAI CTO, open source AI model, American AI model, fine tuning as a service, enterprise AI, multimodal AI, agentic models, customizable AI, Tinker, enterprise distribution, model procurement, Chinese open source models, strategic reset, model overhang, capabilities gap, AI model shopping, model routing, cost-conscious enterprises, artificial intelligence index, 975 billion parameter model, text-image-audio AI, open weights, proprietary AI models, customization accessibility, small language models, AI workflows, context window, Bridgewater use case, model distillation, GPU infrastructure, API costs, token efficiency, fine-tuned models, post training, AI competitive leverage, recurring financial judgment, AI benchmarks, middle tier models, automated model evaluation, privacy and workflow mapping, economical AI models, model rental, model routing automation.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info)
Transcript
Discussion (0)
One of the more important AI models you've probably never heard of and likely won't use just got released.
This may get slept on, but inkling is a huge release from thinking machines lab.
So if you haven't heard of thinking machines, they're led by Mira Murati, the former CTO of OpenAI.
So why is inkling maybe the most important AI model you probably won't use?
because it's now the best open source model from an American company
and thinking machines is betting on the future of AI, fine-tuning as a service.
The model itself, though, it's multimodal, it's agentic, it's customizable,
and it's available through enterprise distribution on day one.
Incling doesn't need to beat every Chinese model or Claude Fable to be relevant.
It only needs to peak the interest of a few cost-conscious enterprises to not only be extremely profitable, but to also help shift the conversation around enterprise AI.
Frontier models are becoming so powerful.
They can actually fine-tune smaller practical models without much iteration or without even much expertise.
So, as a new category reemerged is fine-tuning back and,
what might thinking machines first major product mean for the broader AI competition?
Well, here's the big picture.
Inklink could change how enterprises buy AI.
So this was just released Wednesday, so hours ago.
And it is not the smartest model, but it is strategically important.
I think it gives companies a credible American open weight model alternative to Chinese models.
And Inkling could be the best general purpose open model because it is multi-modal, right?
A lot of the open source, open-weight Chinese models aren't.
Many of them are text-only.
So Inklink, it is multimodal.
It is agentic.
And as workflows in the real world lags so far behind model capabilities,
I do think that there's a real market for bespoke middle of the pack AI.
So on today's.
show, here's what you're going to learn. You're going to learn why an open model matters, even if you're
never going to deploy it. You're going to understand how inkling reopens enterprise AI options beyond just
Chinese open models. You're going to know why frontier models could make fine-tuning practical
beyond research teams. And you're going to know how model shopping is going to change budgets,
vendors, and competitive leverage. All right. Let's get to it. Welcome to Everyday AI. If you're new here,
my name is Jordan Wilson. We do this every day. It's your
your daily live stream podcasts and free daily newsletter helping business leaders like you and me
keep up with the nonstop AI updates because my gosh can't take an hour off. I help you decide what's
important. I tell you how to use all this information to grow your company and your career.
So it starts here with the unedited, unscripted live stream podcast, but make sure you go to our
website at your everyday AI.com. That's your cheat code. We're going to be not just recapping
the highlights from today's show, but go sign up for our free daily newsletter. We're going to be
giving you everything else that you need to know that's happening in the world of AI today.
Because, yeah, it's one of those things.
You got to, like, go and work it out.
It's a muscle.
Use it every single day.
And yeah, all the AI news will be in our newsletter.
All right.
Let's get into it.
Live dream audience.
Good to see you.
Adam joining us from St.
Louis, Jose, from Santiago.
Angie, joining from Montana, Amico, Tokyo.
Brian.
What's up, Brian?
Joining from Minnesota.
So,
let's talk about.
maybe the most important AI model you'll probably never use.
So there's a lot of different factors that have been compounding over the last,
I would say three months.
And to put it in a very short summary,
Frontier models are probably for the most part,
too much for many enterprises, right?
I'll say this,
You know, probably the Fortune 500, they can squeeze as much juice and, you know, make it worth the cost.
But I'd say for many companies, especially those probably in the Fortune 5 or like the Fortune 501 to the 5,000, right?
So this isn't everyone.
You know, I don't think this changes the equation for every single company, every single business out there because so many people are still going to want, you know, their chat GPT enterprise, their, you know,
Gemini, they're Claude, they're co-pilot. They want to make things easy and they're not necessarily,
you know, looking at, okay, how do open models or fine-tuning models change our strategy? But for so
many in our audience, it does. And this is a really big deal. So here's a little bit more about
the model itself. So model, the model is from thinking machines and it's called inkling. It is a
975 billion parameter model and it's multimodal. So text, images,
and audio. What is interesting as well, previously thinking machines did demo kind of a similar
dual purpose or, you know, two-way street audio model as well. So that can kind of hear and listen
as well as talk at the same time like OpenAI's new GPT live. So Inkling is led by former OpenAI CTO
Miramaradi and its quick arrival raises, I think, the next big question, which is how good is it?
Well, here's from the company themselves from their release.
They say our model called inkling is a mixture of experts transformer with 975 billion total
parameters, 41 billion active.
It supports a context window of up to 1 million tokens.
It was pre-trained on 45 trillion tokens of text, images, audio, and video.
It is the first in family of models of different sizes.
Alongside it, we are sharing a preview of inkling small, a lighter weight model with 12,
billion active parameters trained with a smaller recipe that achieved strong performance with even
lower costs and latency inkling reasons natively over text images and audio and balances costs with
performance through efficient and controllable thinking effort we trained it to be a broad
balanced foundation model strong across many domains flexible enough to adapt inkling is not the strongest
overall model available today open or close instead a combination of quality of
make it a good open weight space for customization, multimodal capabilities, efficient thinking,
and availability on Tinker for fine tuning.
Inkling is just the start.
Our first release in a model family we will continue to build on.
We want to make customization accessible for more use cases.
So inkling is available for fine tuning on Tinker today.
Picking the right base model to fine tune is a qualitative judgment that combines measurable
benchmarks with a unique feel of a model that comes from playing with.
it. To enable the latter, we're adding the inkling playground in the Tinker Council, a developer-facing
interface for chatting with Inklinkling. To show what customization means in practice, we ask Inklinkling
to fine-tune itself using Tinker. The model wrote its own fine-tuning job, ran it, and evaluated
the result. So, long story, kind of short, right? If you want to go use Tinkling, Tinkling, right?
And I don't know why.
The combination of thinking machines lab and tinker, it just doesn't roll off the ton, right?
But if you want to go try it out, so they do have like a dev counsel where you can go quote unquote chat,
which is interesting that they just didn't release a chat version.
Anyways, this is, I think, a really big deal.
And here's one of the reasons why.
Yes, benchmarks.
So podcast audience is showing.
the artificial analysis intelligence index.
But this is shaded here by our closed proprietary in our open weights models.
So let me zoom way out for our very non-technical audience or if you're very new to AI.
What's the difference?
What's proprietary open source, right?
proprietary are models that you can't really modify them to the core.
You can add custom instructions to them and, you know, change their behavior that way.
but you can't really change the foundations of these proprietary models, right?
Those are the clods, the GPTs, the Gemini's, et cetera, right?
Then you have these open source or open weight models.
And these are ones that, for the most part, China has been absolutely dominating on.
And they've been doing it through distillation, which is, you know, not exactly something the American labs are happy with.
But the Chinese labs are essentially, you know, stealing or borrowing, whatever you might call it,
the work of the American labs and making their own versions of these models.
And then they serve those, right?
So you can, if you're a big company and if you have the server racks, you can use these models and the open source, open weight models.
If you have the infrastructure, you can download them and run them 24-7 and not really pay any additional costs.
if you have that capacity.
So that is the allure of these open source models.
Or as consumer hardware becomes more capable being able to run some of these locally.
All right.
So some of these models, not necessarily the ones I'm showing here on screen,
but some of these open source models,
if you do have a very expensive, very beefy machine,
you can run a slower version of them.
So that's kind of the premise and the difference between proprietary and
open source models. But here's why I think it's interesting. Because inkling, where it came at and
where it landed on the artificial intelligence index, a 41. So right now, your leaders are
Fable 5 with a 60 and GPD 56 sole with a 59. All right. So 41, you know, seems like,
okay, that's a pretty big drop off, middle of the pack, right? True. But when you put into context
that go back about seven-ish months.
The leaders of the pack at that time,
well, it was GPT 5.2 with a 42.
So the numbers change, all right,
because the benchmarks that go into this,
the artificial analysis intelligence index,
those benchmarks get updated.
But it's essentially about a dozen or so
different benchmarks that are always updated and rotated
that tell you how good is a model
compared to, you know, the most important factors.
And inkling, again, only being about seven months behind the frontier is actually pretty
impressive for the first release from a company that we didn't really know what they were
working on, right?
We really didn't start hearing from a Thinking Machines Lab for their first, like, year that
they launched, right?
So they launched, I think it was quarter one or quarter two of 2025.
We didn't really hear anything for them for a year.
And then we heard they're working on.
you know, Tinker and then we heard that they were making their own model for fine-tuning.
And then we saw this, you know, this bi-directional voice model.
But we didn't actually see the inkling model until, well, less than 24 hours ago.
But if you put it like that, it is about seven months behind the frontier.
But it is an American model, which is actually important.
And it's multimodal.
Those two things alone, I think, have the potential to reshape what's possible.
Because my thought is, right, that model right there,
it's not going to you know if your team is AI native if you have you know agents running if you have
workflow set you're not going to be able to slip like let let me just be honest right you can't just
you know click copy and paste and put inkling in there but for those companies that are still finding
their footing or larger enterprise companies that are looking to you know chunk off a big piece of
their workflows to something maybe more affordable inkling is actually not a bad option right
So the real product, though, is Tinker.
It is they are trying to turn fine-tuning into a service.
So Inkling, yes, it can be downloaded, but even those compressed versions need 600 gigabytes of memories.
So, yeah, you're probably not running this or using this unless you are a large enterprise organization with your own, you know, GPU server infrastructure.
So Tinker, though, manages the training so companies can customize the models without actually owning the GPUs.
And the business model essentially turns that openness into what I think could be the first
strategic reset.
So let's talk about those potential strategic resets.
Reset number one, American open weights could reopen procurement.
So many big enterprises haven't been able to touch some of these open source Chinese models.
Well, because of the current, you know, China-U.S. relationship.
right, especially for those companies that do business with the government, right?
And we've seen the U.S. government gets much more involved here recently between, you know, export controls.
But also a huge thing here is that we've also seen reports in the past week that China may actually, which is, I don't know, funny or interesting, right?
that China may shut down its models to other countries.
So even though they are distilling from U.S. companies, we've seen reports they may not let,
you know, who knows how that will be set up.
But they may not, quote unquote, allow overseas companies to use their open source
models.
So I don't know how open source that actually makes them and especially if they're just
distilling them from U.S. labs anyways, but that's beside the points.
But I think for so many enterprises, they haven't been able to look at the open source category yet because of that reason.
Right.
If you're a company that has big government contracts, if you are using Chinese open source models, that's going to put you in a sticky situation.
Or your RFP might be DOA, right?
You may not even be considered if you are a company that has been using Chinese open source models.
So that's a big unlock.
All right.
And I think that many companies have also just wanted to use open source models, but the whole fact that these are Chinese models.
And we've seen reports on, you know, can you actually trust, right?
Like what's kind of the messaging that may be coming out of this that may not be in line with companies that have stronger, you know, American values or stronger Western values?
I'm not going to get into that.
But, you know, there's obviously many different reasons, not just geopolitical.
reasons that so many companies here in the U.S. haven't really been able to, you know, convince
their board or convince, you know, anyone to go down the open source route just because it is all
Chinese models. So here's potential reset too. The model overhang makes this customization very
timely. So I've talked about this a lot over the past like three months, but I think we're now
at this point where there's a model overhang. It's a little bit.
different than the capabilities gap that I talked about.
There's the great anthropic study from, it seems like it was from so long ago,
but it was only from a couple months ago.
Their labor index report, you know, that essentially showed the capabilities of these
models and then what companies were actually using them for, right?
They anonymized, I think it was 400,000 agenic chats and they mapped it all out.
And essentially what they said is these models are so capable.
But, you know, maybe, you know,
Enterprises are using, you know, on average, about 10 to 20% of the model capabilities.
So essentially, these models right now, the frontier models, are way more powerful than most than the average company actually needs or even has the actual capabilities to take advantage of.
So that kind of adoption gap makes fine-tuning middle of the pack models for stable, repeated work, maybe an actual new and intriguing area of,
of AI. You know, couple that with the fact that we have seen this whiplash, the token maxing to,
you know, value maxing or token efficiency whiplash where, you know, earlier in, you know,
from December 2025 to I would say March 26, you know, enterprise companies were like, yes,
you know, we've been all in on AI. So go use as many tokens as you can. Right. And then it's like,
wait, these token costs are getting higher and higher. And, you know, certain model providers
aren't token efficient, right?
The data says that is anthropic, right?
So all of a sudden, these companies have these huge API bills.
So now we've seen, you know, in the May, June, July, this whiplash of companies being like,
wait, we have to start raining spend in.
And part of it is, well, they're realizing that we don't need, you know, a Fable 5 type model
for, you know, 100 employees rewriting their emails.
And that's what I think these middle of the pack bespoke customized models could actually be a big thing.
And why is that? Well, fine tuning couldn't return because now we actually have frontier models that are smart enough to teach, right?
Which we haven't really had before.
Actually have, you know, a couple exciting examples that I'm going to go over here.
But, you know, this, the concept of fine tuning, right?
So this is the teacher student model.
This is essentially, right?
And we've also got reports.
If you read our newsletter every single day, we've gotten now some sniffs of real like,
okay, we're at this point of recursive self-improvement where, you know, models are actually,
according to research, right?
There's actually a very interesting paper from earlier this week that, yeah, the models are actually improving themselves.
Right.
But that also lends itself to the big, big models.
I think maybe GPD-5-6 soul is the first one I've seen in mass that we,
we've seen actual real examples from that can teach and train smaller, much, much,
much, much smaller versions of themselves to be used for many different purposes, right?
Because these frontier models can generate data evaluations, code, failure analysis.
So I think it is not just, you know, tinker, this fine tuning as a service from thinking
machines and their new model is the combination of that plus this,
new tier of frontier models that make fine-tuning a reality, right? Because it has to become
common language, right? And I think we're going to see that. It's like I could right now,
and I have an example that I think is a really awesome one, but I think I forgot to include the photo
of it in my slide. But you can, if you have a decent enough machine, even a local one,
right, like a Mac studio, you can start building very, very small models yourself.
with hardly any idea of what the heck you're doing for very small purposes.
All right.
So three of these more recent examples of this concept of, you know, well, in this case,
GPD 56 Seoul being a fine-tuned models.
So Open AIs Jason Luai said that GPD 56 Seoul actually helped post-trained open AI's small model,
GVD-56 Luna, completing the work estimated to require two researchers roughly
two weeks. This is the one I forgot to put the screenshot in on my on my slides here.
But Pietro Scherano, hopefully I got the name right. So he's CEO of an AI company,
but he said just a fun little project that he worked on. He used GPT56 to build a small
local model from his I message history. So essentially, I think he had to work overnight or
something like that. And it literally read every single message in his Mac history.
You know, so on your Mac, you can access your I message.
If you're, you know, one of our green bubble friends, you're like, what does that mean?
Right. So on, on my computer, I have my text messages.
So he just let it go.
It read everything.
It understood everything, ran some models, some evaluations.
And all of a sudden, he had a literal, a small language model that GBT 5-6 made for him
trained on that.
So this isn't like a custom GPT, right?
It's like, oh, I'm using the big model to give it in the
instructions. No, it is a separate model itself, right? Absolutely crazy. Then NVIDIA used
Kodax. They just put a blog post out about this, I think on Tuesday. Invidia just used Kodax and
GVD56 to post-train its Cosmos 3 nanomodel, reportedly improving accuracy in it from 54% to 93% in one day
using two prompts. All right. I'll make sure to reshare that in today's new
newsletter. I think we did put it in Tuesdays or yesterdays, but since I'm mentioning it on the show,
I'll make sure to put it in there. But here's the reality. Two prompts. These are off the shelf
creating or improving existing models with today's frontier technology. So it's not just inkling,
right? It is more of a reflection of the combination of these very powerful frontier models that can
literally build, run the evaluations, run the testing, run the QA, can literally build large, small
language models on their own and also train, you know, other variations of themselves.
So one example, though, that we got from thinking machines, they did share recently one use
case about Bridgewater, right? So using their tinker service, right? So this fine-tuning as a service,
Bridgewater customized Quinn, which is a Chinese open source model. So this is, you know,
obviously before inkling was available or at least before they were ready to use it and talk about it.
they did come out with this customer use case of Bridgewater,
customizing the Chinese open source Quinn model through thinking machines tinker
platform for recurring financial judgment tasks.
So it reportedly beat the best tested frontier model while costing 13.8 times less.
So in this case, specialized judgment beat the frontier defaults.
And that kind of explains this potential model.
shopping shift. So let me just put it out here, y'all, not wanting to be like, hey, I told you guys
this a long time ago. Maybe I was just a little too early or a little too weird. But I think two
years ago in my AI prediction series, right, I talked about this very thing. And the, the rise of potentially
seeing thousands of small language models created by the large language models. And all
Although we're still maybe not there because, you know, there is this whole thing called, you know, compute is scarce, right?
Where necessarily you can't have Anthropic and Open AI and Google and Microsoft and meta and GROC.
They can't necessarily, you know, pull gigawatts of compute to create thousands of these bespoke middle tier and small language models.
But that's still the reality.
And I think that this here, the combination of.
GPD 56 soul, being able to create small language models and post train its own versions
and the combination of thinking machines coming up with this as a service.
I think there's finally my, I forgot if it was late 2023 or late 2024 prediction could be coming
true where I think we are eventually and very soon going to see large language models
create hundreds or thousands of versions of small language models because they're cheaper.
And there's finally the appetite to care about it because nine months ago, people didn't
necessarily care because we were still, right?
We were still on this, you know, this richy rich blank check AI, right?
For $20 a month, you couldn't most 90% of employees, right?
If you had a team or a business plan paying $20 to $50 a month per seat,
most employees couldn't get through that.
So companies didn't necessarily care about being efficient with their AI.
The concept of a small language model or a medium-sized model fine-tuned four specific tasks
didn't necessarily matter because the big model could still do it.
And everyone essentially, we were playing with monopoly money until about four months ago.
So the whole concept of small language models or fine-tuned models literally didn't matter
because it was free money.
Now, because of the whiplash, this is more important than ever.
So as we wrap, here's what I want to talk about.
Reset three, right?
Model shopping is now big.
And this is now, I think, part of the executive playbook.
And we saw Microsoft, right?
Microsoft is reportedly going to start moving away from Open AI and Anthropic models.
Not moving away, but they're going to start mixing in their own models as well as
reportedly deep seek models.
So even the biggest enterprise customers are starting to realize that for a majority of day-to-day tasks,
you might not need the number one model in the world to do a big chunk of the work.
And I think now it's more of this concept of maybe renting frontier intelligence for those, you know, for the ambiguous and risky and changing work.
But I think the future, which I've talked about, is using the right model for the right purpose.
at the right time.
And eventually that will become automatic, right?
I've actually built some things like that.
I'm like, I wonder why no one else is doing this
because it's not like, especially since GPD 56 came out.
Like I built skills that essentially model route, right?
I can use Fable inside of codex.
I can use GPD 56 inside of Claude desktop, right?
If you have a little bit of skill and enough patience,
you can do those things, right?
And I've done the same things where not that I ever hit my usage,
but just to practice, right, where I have, I've built in model routing within codex or,
you know, chat GPT work.
So if you are a cost conscious, you know, company, this is the future of working with multiple
models.
And you're not just going to throw, you know, every single request at one big model.
So here's the new AI playbook.
You own the workflow, but you're probably just going to rent the frontier or just
exclusively use the frontier for those specific purposes.
So you have to score every workflow.
by the stakes, the volume, the privacy,
and how much proprietary judgment actually differentiates it.
So you have to start to default to more of this economical model, right?
Like you have to start mapping out that workflow and saying,
okay, maybe we send the 20% to, you know,
especially if you're using it via the API to that frontier model, right?
But I mean, my gosh, talk about middle and lower tier models.
I mean, open AIs, uh, Terra.
in Luna, when it comes to a cost efficiency are legit off the charts, right?
So if I'm advising companies, you know, I'm saying like, you should be using these models
for the most part, right?
The Luna model on like max setting is probably more than enough for almost 90% of the work
that you would do.
So it is kind of shifting away from this one model for all purposes to understanding,
which I know is confusing because it is easy to hit that easy.
button, but we can't just hit that easy button every single time because that button is going to start to cost more and more money as it requires more and more power.
Right. So you have to default to those sometimes more economical models route by difficulty and then fine tune potentially stable, repeated measurable work.
And you have to stop asking which model is the smartest and start asking which model is enough for each job.
All right. That's a wrap. I hope this one was held.
the pretty exciting release, maybe not just for the model itself, but more for what it represents.
So I hope this was helpful.
If so, please let me know about it.
Go sign up for our free daily newsletter.
Drop me in line when you get that automated, you know, email, welcome email, but also, if you could, do me a favor.
Go subscribe to the podcast on Spotify or Apple Podcasts.
Appreciate you tuning in.
We'll see you back tomorrow and every day for more Everyday AI.
Thanks, y'all.
And that's a wrap for today's edition of Everyday AI.
Thanks for joining us.
If you enjoyed this episode, please subscribe and leave us a rating.
It helps keep us going.
For a little more AI magic, visit Your EverydayAI.com
and sign up to our daily newsletter so you don't get left behind.
Go break some barriers and we'll see you next time.
And that's a wrap for today's edition of Everyday AI.
Thanks for joining us.
If you enjoyed this episode, please subscribe and leave us a rating.
It helps keep us going.
For a little more AI magic, visit your everyday AI.com and sign up to our daily newsletter so you don't get left behind.
Go break some barriers and we'll see you next time.
