Everyday AI Podcast – An AI and ChatGPT Podcast - Ep 876: The Most Important AI Model You’ll Probably Never Use That Just Dropped (Replay)

Episode Date: October 7, 2026

You've probably never heard of Inkling. It's the newest (and first) model from Thinking Machines Labs, and it could very well be a small snowball that picks up major momentum in today'...s enterprise AI landscape. If you haven’t heard of Thinking Machines, they’re led by Mira Murati, the former CTO at OpenAI. The big bet with Inkling? The future of AI could be using smaller models fine-tuned and optimized for smaller tasks. Will it work? Tune in live as we dive in. The Most Important AI Model You’ll Probably Never Use That Just Dropped -- An Everyday AI Chat With Jordan Wilson (Replay) Newsletter: Sign up for our free daily newsletterMore on this Episode: Episode PageToday's Episode on LinkedIn: Thoughts on this? Join the convo on LinkedIn and connect with other AI leaders.Upcoming Episodes: Check out the upcoming Everyday AI Livestream lineupWebsite: YourEverydayAI.comEmail The Show: info@youreverydayai.comConnect with Jordan on LinkedInTopics Covered in This Episode:Inkling AI Model Launch OverviewThinking Machines Lab Leadership HighlightInkling's Multimodal and Agentic CapabilitiesOpen Source vs. Proprietary AI ModelsEnterprise Procurement with American AI ModelsAI Fine Tuning as a Service (Tinker)Benchmark Scores: Inkling vs. Frontier ModelsCustomization and Model Shopping for EnterprisesAI Token Costs Driving Model EfficiencyBridgewater Case Study: AI Model CustomizationFrontier Models Enabling Efficient Fine-TuningFuture Trends: Specialized Small Language ModelsTimestamps:00:00 Inkling: A new AI model release05:43 Inkling AI model details09:08 China's dominance in open source AI11:48 Launch and model updates discussed15:21 Concerns over using Chinese open-source models19:06 Training smaller AI models20:22 Using GPT for AI Model Training23:54 Predicting Rise of Small Language Models28:38 Choosing the right AI modelKeywords: Inkling, Thinking Machines Lab, Meera Muradi, former OpenAI CTO, open source AI model, American AI model, fine tuning as a service, enterprise AI, multimodal AI, agentic models, customizable AI, Tinker, enterprise distribution, model procurement, Chinese open source models, strategic reset, model overhang, capabilities gap, AI model shopping, model routing, cost-conscious enterprises, artificial intelligence index, 975 billion parameter model, text-image-audio AI, open weights, proprietary AI models, customization accessibility, small language models, AI workflows, context window, Bridgewater use case, model distillation, GPU infrastructure, API costs, token efficiency, fine-tuned models, post training, AI competitive leverage, recurring financial judgment, AI benchmarks, middle tier models, automated model evaluation, privacy and workflow mapping, economical AI models, model rental, model routing automation.Send Everyday AI and Jordan a text message. (We can't reply back unless you leave contact info)

Transcript
Discussion (0)
Starting point is 00:00:02 One of the more important AI models you've probably never heard of and likely won't use just got released. This may get slept on, but inkling is a huge release from thinking machines lab. So if you haven't heard of thinking machines, they're led by Mira Murati, the former CTO of OpenAI. So why is inkling maybe the most important AI model you probably won't use? because it's now the best open source model from an American company and thinking machines is betting on the future of AI, fine-tuning as a service. The model itself, though, it's multimodal, it's agentic, it's customizable, and it's available through enterprise distribution on day one.
Starting point is 00:00:55 Incling doesn't need to beat every Chinese model or Claude Fable to be relevant. It only needs to peak the interest of a few cost-conscious enterprises to not only be extremely profitable, but to also help shift the conversation around enterprise AI. Frontier models are becoming so powerful. They can actually fine-tune smaller practical models without much iteration or without even much expertise. So, as a new category reemerged is fine-tuning back and, what might thinking machines first major product mean for the broader AI competition? Well, here's the big picture. Inklink could change how enterprises buy AI.
Starting point is 00:01:41 So this was just released Wednesday, so hours ago. And it is not the smartest model, but it is strategically important. I think it gives companies a credible American open weight model alternative to Chinese models. And Inkling could be the best general purpose open model because it is multi-modal, right? A lot of the open source, open-weight Chinese models aren't. Many of them are text-only. So Inklink, it is multimodal. It is agentic.
Starting point is 00:02:15 And as workflows in the real world lags so far behind model capabilities, I do think that there's a real market for bespoke middle of the pack AI. So on today's. show, here's what you're going to learn. You're going to learn why an open model matters, even if you're never going to deploy it. You're going to understand how inkling reopens enterprise AI options beyond just Chinese open models. You're going to know why frontier models could make fine-tuning practical beyond research teams. And you're going to know how model shopping is going to change budgets, vendors, and competitive leverage. All right. Let's get to it. Welcome to Everyday AI. If you're new here,
Starting point is 00:02:56 my name is Jordan Wilson. We do this every day. It's your your daily live stream podcasts and free daily newsletter helping business leaders like you and me keep up with the nonstop AI updates because my gosh can't take an hour off. I help you decide what's important. I tell you how to use all this information to grow your company and your career. So it starts here with the unedited, unscripted live stream podcast, but make sure you go to our website at your everyday AI.com. That's your cheat code. We're going to be not just recapping the highlights from today's show, but go sign up for our free daily newsletter. We're going to be giving you everything else that you need to know that's happening in the world of AI today.
Starting point is 00:03:33 Because, yeah, it's one of those things. You got to, like, go and work it out. It's a muscle. Use it every single day. And yeah, all the AI news will be in our newsletter. All right. Let's get into it. Live dream audience.
Starting point is 00:03:45 Good to see you. Adam joining us from St. Louis, Jose, from Santiago. Angie, joining from Montana, Amico, Tokyo. Brian. What's up, Brian? Joining from Minnesota. So,
Starting point is 00:03:57 let's talk about. maybe the most important AI model you'll probably never use. So there's a lot of different factors that have been compounding over the last, I would say three months. And to put it in a very short summary, Frontier models are probably for the most part, too much for many enterprises, right? I'll say this,
Starting point is 00:04:27 You know, probably the Fortune 500, they can squeeze as much juice and, you know, make it worth the cost. But I'd say for many companies, especially those probably in the Fortune 5 or like the Fortune 501 to the 5,000, right? So this isn't everyone. You know, I don't think this changes the equation for every single company, every single business out there because so many people are still going to want, you know, their chat GPT enterprise, their, you know, Gemini, they're Claude, they're co-pilot. They want to make things easy and they're not necessarily, you know, looking at, okay, how do open models or fine-tuning models change our strategy? But for so many in our audience, it does. And this is a really big deal. So here's a little bit more about the model itself. So model, the model is from thinking machines and it's called inkling. It is a
Starting point is 00:05:21 975 billion parameter model and it's multimodal. So text, images, and audio. What is interesting as well, previously thinking machines did demo kind of a similar dual purpose or, you know, two-way street audio model as well. So that can kind of hear and listen as well as talk at the same time like OpenAI's new GPT live. So Inkling is led by former OpenAI CTO Miramaradi and its quick arrival raises, I think, the next big question, which is how good is it? Well, here's from the company themselves from their release. They say our model called inkling is a mixture of experts transformer with 975 billion total parameters, 41 billion active.
Starting point is 00:06:07 It supports a context window of up to 1 million tokens. It was pre-trained on 45 trillion tokens of text, images, audio, and video. It is the first in family of models of different sizes. Alongside it, we are sharing a preview of inkling small, a lighter weight model with 12, billion active parameters trained with a smaller recipe that achieved strong performance with even lower costs and latency inkling reasons natively over text images and audio and balances costs with performance through efficient and controllable thinking effort we trained it to be a broad balanced foundation model strong across many domains flexible enough to adapt inkling is not the strongest
Starting point is 00:06:50 overall model available today open or close instead a combination of quality of make it a good open weight space for customization, multimodal capabilities, efficient thinking, and availability on Tinker for fine tuning. Inkling is just the start. Our first release in a model family we will continue to build on. We want to make customization accessible for more use cases. So inkling is available for fine tuning on Tinker today. Picking the right base model to fine tune is a qualitative judgment that combines measurable
Starting point is 00:07:23 benchmarks with a unique feel of a model that comes from playing with. it. To enable the latter, we're adding the inkling playground in the Tinker Council, a developer-facing interface for chatting with Inklinkling. To show what customization means in practice, we ask Inklinkling to fine-tune itself using Tinker. The model wrote its own fine-tuning job, ran it, and evaluated the result. So, long story, kind of short, right? If you want to go use Tinkling, Tinkling, right? And I don't know why. The combination of thinking machines lab and tinker, it just doesn't roll off the ton, right? But if you want to go try it out, so they do have like a dev counsel where you can go quote unquote chat,
Starting point is 00:08:04 which is interesting that they just didn't release a chat version. Anyways, this is, I think, a really big deal. And here's one of the reasons why. Yes, benchmarks. So podcast audience is showing. the artificial analysis intelligence index. But this is shaded here by our closed proprietary in our open weights models. So let me zoom way out for our very non-technical audience or if you're very new to AI.
Starting point is 00:08:37 What's the difference? What's proprietary open source, right? proprietary are models that you can't really modify them to the core. You can add custom instructions to them and, you know, change their behavior that way. but you can't really change the foundations of these proprietary models, right? Those are the clods, the GPTs, the Gemini's, et cetera, right? Then you have these open source or open weight models. And these are ones that, for the most part, China has been absolutely dominating on.
Starting point is 00:09:10 And they've been doing it through distillation, which is, you know, not exactly something the American labs are happy with. But the Chinese labs are essentially, you know, stealing or borrowing, whatever you might call it, the work of the American labs and making their own versions of these models. And then they serve those, right? So you can, if you're a big company and if you have the server racks, you can use these models and the open source, open weight models. If you have the infrastructure, you can download them and run them 24-7 and not really pay any additional costs. if you have that capacity. So that is the allure of these open source models.
Starting point is 00:09:55 Or as consumer hardware becomes more capable being able to run some of these locally. All right. So some of these models, not necessarily the ones I'm showing here on screen, but some of these open source models, if you do have a very expensive, very beefy machine, you can run a slower version of them. So that's kind of the premise and the difference between proprietary and open source models. But here's why I think it's interesting. Because inkling, where it came at and
Starting point is 00:10:22 where it landed on the artificial intelligence index, a 41. So right now, your leaders are Fable 5 with a 60 and GPD 56 sole with a 59. All right. So 41, you know, seems like, okay, that's a pretty big drop off, middle of the pack, right? True. But when you put into context that go back about seven-ish months. The leaders of the pack at that time, well, it was GPT 5.2 with a 42. So the numbers change, all right, because the benchmarks that go into this,
Starting point is 00:11:00 the artificial analysis intelligence index, those benchmarks get updated. But it's essentially about a dozen or so different benchmarks that are always updated and rotated that tell you how good is a model compared to, you know, the most important factors. And inkling, again, only being about seven months behind the frontier is actually pretty impressive for the first release from a company that we didn't really know what they were
Starting point is 00:11:26 working on, right? We really didn't start hearing from a Thinking Machines Lab for their first, like, year that they launched, right? So they launched, I think it was quarter one or quarter two of 2025. We didn't really hear anything for them for a year. And then we heard they're working on. you know, Tinker and then we heard that they were making their own model for fine-tuning. And then we saw this, you know, this bi-directional voice model.
Starting point is 00:11:50 But we didn't actually see the inkling model until, well, less than 24 hours ago. But if you put it like that, it is about seven months behind the frontier. But it is an American model, which is actually important. And it's multimodal. Those two things alone, I think, have the potential to reshape what's possible. Because my thought is, right, that model right there, it's not going to you know if your team is AI native if you have you know agents running if you have workflow set you're not going to be able to slip like let let me just be honest right you can't just
Starting point is 00:12:22 you know click copy and paste and put inkling in there but for those companies that are still finding their footing or larger enterprise companies that are looking to you know chunk off a big piece of their workflows to something maybe more affordable inkling is actually not a bad option right So the real product, though, is Tinker. It is they are trying to turn fine-tuning into a service. So Inkling, yes, it can be downloaded, but even those compressed versions need 600 gigabytes of memories. So, yeah, you're probably not running this or using this unless you are a large enterprise organization with your own, you know, GPU server infrastructure. So Tinker, though, manages the training so companies can customize the models without actually owning the GPUs.
Starting point is 00:13:11 And the business model essentially turns that openness into what I think could be the first strategic reset. So let's talk about those potential strategic resets. Reset number one, American open weights could reopen procurement. So many big enterprises haven't been able to touch some of these open source Chinese models. Well, because of the current, you know, China-U.S. relationship. right, especially for those companies that do business with the government, right? And we've seen the U.S. government gets much more involved here recently between, you know, export controls.
Starting point is 00:13:56 But also a huge thing here is that we've also seen reports in the past week that China may actually, which is, I don't know, funny or interesting, right? that China may shut down its models to other countries. So even though they are distilling from U.S. companies, we've seen reports they may not let, you know, who knows how that will be set up. But they may not, quote unquote, allow overseas companies to use their open source models. So I don't know how open source that actually makes them and especially if they're just distilling them from U.S. labs anyways, but that's beside the points.
Starting point is 00:14:36 But I think for so many enterprises, they haven't been able to look at the open source category yet because of that reason. Right. If you're a company that has big government contracts, if you are using Chinese open source models, that's going to put you in a sticky situation. Or your RFP might be DOA, right? You may not even be considered if you are a company that has been using Chinese open source models. So that's a big unlock. All right. And I think that many companies have also just wanted to use open source models, but the whole fact that these are Chinese models.
Starting point is 00:15:15 And we've seen reports on, you know, can you actually trust, right? Like what's kind of the messaging that may be coming out of this that may not be in line with companies that have stronger, you know, American values or stronger Western values? I'm not going to get into that. But, you know, there's obviously many different reasons, not just geopolitical. reasons that so many companies here in the U.S. haven't really been able to, you know, convince their board or convince, you know, anyone to go down the open source route just because it is all Chinese models. So here's potential reset too. The model overhang makes this customization very timely. So I've talked about this a lot over the past like three months, but I think we're now
Starting point is 00:16:02 at this point where there's a model overhang. It's a little bit. different than the capabilities gap that I talked about. There's the great anthropic study from, it seems like it was from so long ago, but it was only from a couple months ago. Their labor index report, you know, that essentially showed the capabilities of these models and then what companies were actually using them for, right? They anonymized, I think it was 400,000 agenic chats and they mapped it all out. And essentially what they said is these models are so capable.
Starting point is 00:16:34 But, you know, maybe, you know, Enterprises are using, you know, on average, about 10 to 20% of the model capabilities. So essentially, these models right now, the frontier models, are way more powerful than most than the average company actually needs or even has the actual capabilities to take advantage of. So that kind of adoption gap makes fine-tuning middle of the pack models for stable, repeated work, maybe an actual new and intriguing area of, of AI. You know, couple that with the fact that we have seen this whiplash, the token maxing to, you know, value maxing or token efficiency whiplash where, you know, earlier in, you know, from December 2025 to I would say March 26, you know, enterprise companies were like, yes, you know, we've been all in on AI. So go use as many tokens as you can. Right. And then it's like,
Starting point is 00:17:30 wait, these token costs are getting higher and higher. And, you know, certain model providers aren't token efficient, right? The data says that is anthropic, right? So all of a sudden, these companies have these huge API bills. So now we've seen, you know, in the May, June, July, this whiplash of companies being like, wait, we have to start raining spend in. And part of it is, well, they're realizing that we don't need, you know, a Fable 5 type model for, you know, 100 employees rewriting their emails.
Starting point is 00:18:04 And that's what I think these middle of the pack bespoke customized models could actually be a big thing. And why is that? Well, fine tuning couldn't return because now we actually have frontier models that are smart enough to teach, right? Which we haven't really had before. Actually have, you know, a couple exciting examples that I'm going to go over here. But, you know, this, the concept of fine tuning, right? So this is the teacher student model. This is essentially, right? And we've also got reports.
Starting point is 00:18:35 If you read our newsletter every single day, we've gotten now some sniffs of real like, okay, we're at this point of recursive self-improvement where, you know, models are actually, according to research, right? There's actually a very interesting paper from earlier this week that, yeah, the models are actually improving themselves. Right. But that also lends itself to the big, big models. I think maybe GPD-5-6 soul is the first one I've seen in mass that we, we've seen actual real examples from that can teach and train smaller, much, much,
Starting point is 00:19:09 much, much smaller versions of themselves to be used for many different purposes, right? Because these frontier models can generate data evaluations, code, failure analysis. So I think it is not just, you know, tinker, this fine tuning as a service from thinking machines and their new model is the combination of that plus this, new tier of frontier models that make fine-tuning a reality, right? Because it has to become common language, right? And I think we're going to see that. It's like I could right now, and I have an example that I think is a really awesome one, but I think I forgot to include the photo of it in my slide. But you can, if you have a decent enough machine, even a local one,
Starting point is 00:19:56 right, like a Mac studio, you can start building very, very small models yourself. with hardly any idea of what the heck you're doing for very small purposes. All right. So three of these more recent examples of this concept of, you know, well, in this case, GPD 56 Seoul being a fine-tuned models. So Open AIs Jason Luai said that GPD 56 Seoul actually helped post-trained open AI's small model, GVD-56 Luna, completing the work estimated to require two researchers roughly two weeks. This is the one I forgot to put the screenshot in on my on my slides here.
Starting point is 00:20:38 But Pietro Scherano, hopefully I got the name right. So he's CEO of an AI company, but he said just a fun little project that he worked on. He used GPT56 to build a small local model from his I message history. So essentially, I think he had to work overnight or something like that. And it literally read every single message in his Mac history. You know, so on your Mac, you can access your I message. If you're, you know, one of our green bubble friends, you're like, what does that mean? Right. So on, on my computer, I have my text messages. So he just let it go.
Starting point is 00:21:13 It read everything. It understood everything, ran some models, some evaluations. And all of a sudden, he had a literal, a small language model that GBT 5-6 made for him trained on that. So this isn't like a custom GPT, right? It's like, oh, I'm using the big model to give it in the instructions. No, it is a separate model itself, right? Absolutely crazy. Then NVIDIA used Kodax. They just put a blog post out about this, I think on Tuesday. Invidia just used Kodax and
Starting point is 00:21:44 GVD56 to post-train its Cosmos 3 nanomodel, reportedly improving accuracy in it from 54% to 93% in one day using two prompts. All right. I'll make sure to reshare that in today's new newsletter. I think we did put it in Tuesdays or yesterdays, but since I'm mentioning it on the show, I'll make sure to put it in there. But here's the reality. Two prompts. These are off the shelf creating or improving existing models with today's frontier technology. So it's not just inkling, right? It is more of a reflection of the combination of these very powerful frontier models that can literally build, run the evaluations, run the testing, run the QA, can literally build large, small language models on their own and also train, you know, other variations of themselves.
Starting point is 00:22:35 So one example, though, that we got from thinking machines, they did share recently one use case about Bridgewater, right? So using their tinker service, right? So this fine-tuning as a service, Bridgewater customized Quinn, which is a Chinese open source model. So this is, you know, obviously before inkling was available or at least before they were ready to use it and talk about it. they did come out with this customer use case of Bridgewater, customizing the Chinese open source Quinn model through thinking machines tinker platform for recurring financial judgment tasks. So it reportedly beat the best tested frontier model while costing 13.8 times less.
Starting point is 00:23:21 So in this case, specialized judgment beat the frontier defaults. And that kind of explains this potential model. shopping shift. So let me just put it out here, y'all, not wanting to be like, hey, I told you guys this a long time ago. Maybe I was just a little too early or a little too weird. But I think two years ago in my AI prediction series, right, I talked about this very thing. And the, the rise of potentially seeing thousands of small language models created by the large language models. And all Although we're still maybe not there because, you know, there is this whole thing called, you know, compute is scarce, right? Where necessarily you can't have Anthropic and Open AI and Google and Microsoft and meta and GROC.
Starting point is 00:24:16 They can't necessarily, you know, pull gigawatts of compute to create thousands of these bespoke middle tier and small language models. But that's still the reality. And I think that this here, the combination of. GPD 56 soul, being able to create small language models and post train its own versions and the combination of thinking machines coming up with this as a service. I think there's finally my, I forgot if it was late 2023 or late 2024 prediction could be coming true where I think we are eventually and very soon going to see large language models create hundreds or thousands of versions of small language models because they're cheaper.
Starting point is 00:25:03 And there's finally the appetite to care about it because nine months ago, people didn't necessarily care because we were still, right? We were still on this, you know, this richy rich blank check AI, right? For $20 a month, you couldn't most 90% of employees, right? If you had a team or a business plan paying $20 to $50 a month per seat, most employees couldn't get through that. So companies didn't necessarily care about being efficient with their AI. The concept of a small language model or a medium-sized model fine-tuned four specific tasks
Starting point is 00:25:41 didn't necessarily matter because the big model could still do it. And everyone essentially, we were playing with monopoly money until about four months ago. So the whole concept of small language models or fine-tuned models literally didn't matter because it was free money. Now, because of the whiplash, this is more important than ever. So as we wrap, here's what I want to talk about. Reset three, right? Model shopping is now big.
Starting point is 00:26:09 And this is now, I think, part of the executive playbook. And we saw Microsoft, right? Microsoft is reportedly going to start moving away from Open AI and Anthropic models. Not moving away, but they're going to start mixing in their own models as well as reportedly deep seek models. So even the biggest enterprise customers are starting to realize that for a majority of day-to-day tasks, you might not need the number one model in the world to do a big chunk of the work. And I think now it's more of this concept of maybe renting frontier intelligence for those, you know, for the ambiguous and risky and changing work.
Starting point is 00:26:49 But I think the future, which I've talked about, is using the right model for the right purpose. at the right time. And eventually that will become automatic, right? I've actually built some things like that. I'm like, I wonder why no one else is doing this because it's not like, especially since GPD 56 came out. Like I built skills that essentially model route, right? I can use Fable inside of codex.
Starting point is 00:27:13 I can use GPD 56 inside of Claude desktop, right? If you have a little bit of skill and enough patience, you can do those things, right? And I've done the same things where not that I ever hit my usage, but just to practice, right, where I have, I've built in model routing within codex or, you know, chat GPT work. So if you are a cost conscious, you know, company, this is the future of working with multiple models.
Starting point is 00:27:39 And you're not just going to throw, you know, every single request at one big model. So here's the new AI playbook. You own the workflow, but you're probably just going to rent the frontier or just exclusively use the frontier for those specific purposes. So you have to score every workflow. by the stakes, the volume, the privacy, and how much proprietary judgment actually differentiates it. So you have to start to default to more of this economical model, right?
Starting point is 00:28:07 Like you have to start mapping out that workflow and saying, okay, maybe we send the 20% to, you know, especially if you're using it via the API to that frontier model, right? But I mean, my gosh, talk about middle and lower tier models. I mean, open AIs, uh, Terra. in Luna, when it comes to a cost efficiency are legit off the charts, right? So if I'm advising companies, you know, I'm saying like, you should be using these models for the most part, right?
Starting point is 00:28:38 The Luna model on like max setting is probably more than enough for almost 90% of the work that you would do. So it is kind of shifting away from this one model for all purposes to understanding, which I know is confusing because it is easy to hit that easy. button, but we can't just hit that easy button every single time because that button is going to start to cost more and more money as it requires more and more power. Right. So you have to default to those sometimes more economical models route by difficulty and then fine tune potentially stable, repeated measurable work. And you have to stop asking which model is the smartest and start asking which model is enough for each job. All right. That's a wrap. I hope this one was held.
Starting point is 00:29:26 the pretty exciting release, maybe not just for the model itself, but more for what it represents. So I hope this was helpful. If so, please let me know about it. Go sign up for our free daily newsletter. Drop me in line when you get that automated, you know, email, welcome email, but also, if you could, do me a favor. Go subscribe to the podcast on Spotify or Apple Podcasts. Appreciate you tuning in. We'll see you back tomorrow and every day for more Everyday AI.
Starting point is 00:29:54 Thanks, y'all. And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit Your EverydayAI.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.
Starting point is 00:30:18 And that's a wrap for today's edition of Everyday AI. Thanks for joining us. If you enjoyed this episode, please subscribe and leave us a rating. It helps keep us going. For a little more AI magic, visit your everyday AI.com and sign up to our daily newsletter so you don't get left behind. Go break some barriers and we'll see you next time.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.