The AI Daily Brief: Artificial Intelligence News and Analysis - The New Enterprise Battle Over Who Owns the Model

Episode Date: July 16, 2026

Thinking Machines Lab’s new open-weight model Inkling may signal a new enterprise battle over who controls the model, the data, and the learning built on top of it. NLW explores its promise—and wh...y fine-tuning may be harder than advocates suggest. In the headlines: Cursor, Apple’s AI chip hunt, and Microsoft’s model push.Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠kpmg.com/us/Sophisticated⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Hyperagent - Hire a fleet of always-on agents. New users get $1,000 in inference. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Retool - Secure your vibecoded apps. New enterprise customers get up to $10,000 in AI credits per year. ⁠⁠⁠⁠⁠⁠⁠⁠⁠retool.com/aidaily ⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Scrunch - The AI customer experience platform - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://scrunch.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://pod.link/1680633614⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Our Newsletter is BACK: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Today on the AI Daily Brief, the month of model continues, and businesses might want to pay attention to this new open-weight model introduced yesterday. Before that in the headlines, Apple hunting for a chip acquisition, Microsoft competing with OpenAI and Anthropic, and cursor also getting deeper into the model game. AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Rackspace, Splitsy,
Starting point is 00:00:30 an air table. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note at sponsors at AIDailybrief.AI. Today's one of those days where the headlines part of the episode actually has a pretty coherent theme with the main episode, because friends, we are in our model era. The information recently reported on an all-hands meeting that cursor held back in May, where CEO Michael Truill laid out the future direction for the company. The meeting he took place a couple of weeks after the SpaceX deal was announced, which at the time was a compute and training partnership with an acquisition option, although even then it seemed pretty
Starting point is 00:01:08 likely that the deal would be finalized after the SpaceX IPO. Truel told staff that Curser would aim to become a top-tier model developer in their own right, not exclusively limited to AI coding. Indeed, in the short term, he aimed to produce a state-of-the-art model by the end of the year, and by 2027, he wanted to accrue a significant compute advantage and actually push the frontier forward. True will also apparently acknowledge some unease among staff about the SpaceX acquisition, he described Curser as being in the midst of rapid change and significant growth and promised to provide more clarity around the deal as it developed. True also explained that SpaceX wanted to leverage Cursor's brand, existing enterprise relationships, and larger
Starting point is 00:01:43 go-to-market team. Now, subsequent to this, we've seen the release of the first SpaceX AI model trained in partnership with Cursor in GROC 4.5. And holding aside recent controversy around data retention, the actual model itself has been well received and seems competitive, especially as we get into more advanced model architectures, where companies are pairing state-of-the-art and frontier, with slightly less performant but slightly more affordable models. Behind the scenes, Cursor is also rumored to be working on a competitor to Claude co-work, which would be their first big expansion beyond coding. Now, it's very clear that Elon and SpaceX AI have big plans for this integration,
Starting point is 00:02:16 and for those who have been paying attention, this is more evolution than revolution, given that Cursor had already decided as early as the end of last year that the model development game was going to have to be a game they played. Still, we have now confirmation that all of the signals we're seeing from outside are what they're talking about inside. And I think as consumers, you've got to think that this just means more quality options for us, so I'm excited to see what they do. Now, speaking of SpaceX on the market side, the company's stock dipped below its IPO price for the first time. On Wednesday, SpaceX traded down to $133.00,000 IPO price for the first time. It did recover very slightly to
Starting point is 00:02:49 close the day at $135.27, but the stock is now down 33% from its all-time high, and Elon Musk has lost his trillionaire status. Musk, of course, doesn't seem all that concerned focusing instead on the next Starship launch. Meanwhile, there are plenty of bears pointing out that insider lockups won't even start for another month. The more significant implications of this than anything happening in the SpaceX world may be actually around what it means for other AI companies looking to IPO. Last we had heard, Open AI were increasingly inclined to wait until next year for their IPO plans. As of late June, advisors were reportedly telling Sam Altman that they thought that they were unlikely to achieve his target of a trillion-dollar market cap. Anthropic, on the other hand,
Starting point is 00:03:28 seemed to still be on track for an IPO in the fall. Bloomberg reports that Anthropic has appointed investment banks to lead the offering and are setting meetings with investors over the coming weeks. Anthropic had been targeting an IPO as early as September or October, which could still be possible. The company has also added a couple billion dollars in a revolving credit facility, sending all the signals that they are, in fact, gearing up for a run at the public markets. Still, it feels like the next few weeks are going to be an extremely important period, as Anthropic will be meeting with public investors to test the waters and see whether an IPO is viable at this time. Moving back to models, Nvidia has released another new open source model, this time attempting
Starting point is 00:04:02 to fill a missing niche. The model called Cosmos 3 Edge is a small physical AI model, weighing in it just 4 billion parameters. It's designed to be run on edge devices that are physically installed in robots, such as Nvidia's Jensen hardware, which also saw a product refresh this week. The model takes a dual approach to robotics, able to function as both a world model, as well as a vision language model similar to standard LLM architecture. Now, last month, Nvidia's release of Nemotron 3 Ultra was pretty timely. That model was a larger and more capable version of their open model family, addressing the demand for powerful and cheap US-based models to deal with spiraling token budgets.
Starting point is 00:04:34 This time, Nvidia's new model for physical AI arrives as another wave of robot designs are unveiled. Last week, 1X previewed the new version of their neo-humanoid with superhuman hand movements. Invidia themselves are striking up a ton of partnerships across the space, including an expanded partnership with Toyota. After working together on self-driving cars in recent years, the partnership will now include robots, smart cities, and automated factories. Now, as a slight detour, one thing that I do keep a bit of an eye on is everything going on with physical AI. Right now, obviously, the main
Starting point is 00:05:01 focus of most of our discussions here is in the generative LLM world with a particular bias towards what matters at work, but you have to think that over the next few years, we're going to have a lot more context to focus on AI in the physical world. One really interesting example of that of a company that just launched is called Chip. Chip is, honestly, one of the first new takes we've seen on a family vehicle for some time. Now, golf cart style smaller vehicles have been getting increasingly popular over the last couple years, but Chip kind of takes things to the next level. The idea is that it combines AI and autonomy to more officially turn your car into a robot that can do things on your behalf. The company just went live in the initial launch video,
Starting point is 00:05:36 shows examples like a father at a soccer game sending Chip to run errands and pick up snacks, to go park itself after the couple decides that they're going to take an Uber home from the bar, and clearly the idea is to rethink what a vehicle can actually do. You can check it out at the just-launched chip motors.com, but as a herald of the type of ways, we're going to see AI integrate itself into the physical world in the coming years. Now, speaking of chip, but a very different type of chip, Apple is apparently shopping for acquisitions to build AI server chips. According to the information, Apple is quite far along in an effort to buy up a chipmaker. They've reportedly spoken with bankers about funding possible deals
Starting point is 00:06:08 and have approached semiconductor startups to gauge interest. Now, Apple has famously shied away from big acquisitions, preferring to develop capacity internally. In fact, their biggest acquisition ever is the Beats deal in 2014, which costs just $3 billion. There is a sense that the AI era might be changing their tune on that, and the information speculates that their most recent push grew out of a realization that Apple doesn't have the technology they need to support AI product. The M-Series chips have been a massive hit when it comes to local inference in the pro-sumer market, i.e. selling out Mac minis, but Apple relies on the same M-2 Ultra chips to run their AI servers, and they might not be up to the task as Apple rolls
Starting point is 00:06:42 out AI Siri. Not only did Apple contract with Google to develop the models of Power Siri, they're also outsourcing server capacity to Google Cloud running on Nvidia chips. Apple did have a solution in development, a server-grade chip codenamed Baltro, which was expected to ship this year, but that project has reportedly been delayed. Now, this could end up being a very large deal if Apple can find a suitable acquisition target, but even if the company is willing to pay up, the selection of proven chipmakers is pretty slim. Invitya just paid 20 billion to take Grok off the board, and Cerebris is currently valued at 40 billion in public markets. Apple has been working with Broadcom on server chips since 2014, but that company's $1.8 trillion dollar valuation.
Starting point is 00:07:16 which would mean a merger rather than an acquisition, which not only would be not particularly appellee, but would also come with a huge amount of antitrust scrutiny. Earlier, there were rumors of 10 Storrent fielding acquisition offers from Broadcom, although that was later denied, and that could be a natural fit given that legendary chip designer and CEO Jim Keller worked at Apple in the late 2000s. At this stage, there's no solid rumors. All we know is that Apple is in the market to buy their way back into the AI race, which will be music to Apple fans' ears.
Starting point is 00:07:41 Now, our last story today is actually kind of a setup to the main episode. Microsoft is preparing their sales teams to compete directly with OpenAI and Anthropic, training them to emphasize the drawbacks of the major model labs. According to Bloomberg reporting, Microsoft leadership laid out the plan in a department-wide meeting on Tuesday. It was framed as a new strategy for the fiscal year that began this month. Sales staff were instructed to push the efficiency and cost-effectiveness of Microsoft's in-house MAI models compared to their rivals, and Microsoft's vertically integrated AI stack was also highlighted as a key selling point.
Starting point is 00:08:11 Selling against Claude appears to be the initial focus, and Microsoft is, increasingly sidestepping overall model quality. Instead, one executive reportedly told sales staff to say that when it comes to performance in Microsoft's office suite specifically, Claude is, quote, slower and less accurate and lack the proper security integrations. Microsoft can point to the fact that customers have already begun using the MAI models in copilot and might not have noticed the difference. Earlier this month, Bloomberg reported that Microsoft had begun switching some copilot functions over to MAI as a cost-cutting measure compared to OpenAI and Anthropic. Now, this new sales pitch aligns directly with Satinadella
Starting point is 00:08:45 recent commentary on the risk to corporate data in the AI era. The not-so-s subtle subtext of this post and one that he published a couple of weeks ago was that companies shouldn't trust OpenAI or Anthropic with their data because both of those companies have an incentive to build competing services. But it also feels to me like this pivot is in part about giving Microsoft a foot in the door with their own models before pushing customer-specific fine-tuning as an upsell. Last month, Microsoft AI CEO Mustafa Sullyman laid out this long-term plan with the introduction of Microsoft Frontier tuning. Now, as we'll see from the main episode, This idea of enterprises tuning their own models is something that a number of labs are going to be
Starting point is 00:09:20 pushing as a major theme for the second half of this year. What sort of enterprise uptake of that there is remains to be seen, but it is certainly going to be in the conversation. And interestingly enough, it was not Microsoft that was the first company to announce this sort of fine tuning, but Thinking Machines lab back at the end of last year with their Tinker API. Well, Tinker just got a big new update today, and that is going to be the subject of our main episode. So with that, we'll close the headlines and head on over to the main. One of the most important AI questions right now isn't who's using AI, it's who's using it well. KPMG in the University of Texas at Austin just analyzed 1.4 million real workplace AI interactions
Starting point is 00:09:59 and found something surprising. The highest impact users aren't better prompt engineers. They treat AI like a reasoning partner. They frame problems, guide thinking, iterate, and push for better answers. And the good news? These behaviors are teachable at scale. If you're trying to move from AI access to real capability, KPMG's research on sophisticated AI collaboration is worth your time.
Starting point is 00:10:21 Learn more at KPMG.com slash US-s sophisticated. That's KPMG.com slash us slash sophisticated. One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations. As AI moves into core workflows, regulated data environments, and agentic systems, enterprises need governed infrastructure and inference that can operate reliably day-to-day with clear operational accountability built in from the start.
Starting point is 00:10:49 As those systems scale, the operating model increasingly becomes part of the AI strategy itself. Rackspace technology is the operator of the full enterprise AI stack, from agents to infrastructure across private cloud, hybrid cloud, and edge environments. Rackspace builds and operates governed AI infrastructure, inference, and production AI systems for organizations where sovereignty, compliance, and uptime are non-negotiable. Their forward-deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments. To learn more about where enterprise AI runs and outcome scale, go to rackspace.com. Want to accelerate enterprise software development velocity by 5X?
Starting point is 00:11:26 You need Blitzy, the only autonomous software development platform built for enterprise codebases. Your engineers define the project, a new feature, refactor, or greenfield build. Blitzy agents first ingest and map your entire codebase, then the platform generates a bespoke agent action plan for your team to review and approve. Once approved, Blitzy gets to work autonomously generating hundreds of thousands of lines of validated end-to-end tested code. More than 80% of the work completed in a single run. Blitzy is not generating code, it's developing software at the speed of compute.
Starting point is 00:11:53 Your engineers review, refine, and ship. This is how Fortune 500 companies are compressing multi-month projects into a single sprint, accelerating engineering velocity by 5X. Experience Blitzy firsthand at blitzie.com. That's BLYTZY.com. This episode of the AI Daily Brief is brought to you by HyperAGent. where you run fleets of agents your team can manage together. New users get $1,000 in inference.
Starting point is 00:12:16 Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales as agent enriches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates.
Starting point is 00:12:40 hire yours at Hyperagent built by the team at Airtable. Claim your $1,000 in inference at hyperagent.com slash AI Daily Brief. Welcome back to the AI Daily Brief. Today, the month of models continues. And for the casual observer, this first one we're going to talk about could at first glance be a little underwhelming. What we'll see, though, throughout this episode is that this is actually part of a much more significant set of trends, and I think particularly for enterprises, is really worth
Starting point is 00:13:09 paying attention to. Thinking Machines Lab has released their first large language model. Called Inklink, the model is built on a mixture of experts' architecture with 975 billion total parameters and 41 billion active parameters. It supports a million token context window and reasons natively across text, images, and audio. Now, TML, which is of course the spin-off lab built by former OpenAI CTO Miramaradi after she left that company, says that this is the first in a family of models they plan to release, including a 12 billion parameter model called Inklink Small, which was also released in preview on Wednesday.
Starting point is 00:13:44 For context, at 975 billion parameters, inkling is quite a bit smaller than the 1.6 billion parameters of Deepseek V4 Pro, but a little larger than the 750 billion parameters of ZAI's GLM 5.2. Importantly, TML pre-trained inkling from scratch. And aside from a small bootstrap distilled from Kimmy K2.5, the rest of the fine-tuning run was based on a set of human-created and synthetic data collected by TML. They wrote, We trained Inkling to be a broad, balanced foundation model, strong across many domains, flexible enough to adapt. Inkling is not the strongest overall model available today, open or closed. Instead, a combination of qualities make it a good open-weight space for customization, multimodal capabilities, efficient thinking, and availability
Starting point is 00:14:28 on Tinker for fine tuning. And indeed, it is that availability, which, as we'll see, is the most important part. Still, to get a sense of where this ranks when it comes to performance, a fair way to put it would be to call it somewhat competitive, but certainly not state-of-the-art across a range of different benchmarks. On Humanity's last exam, which measures knowledge and web search, it scores 29.7% without tools and 46% width. This is ahead of Nemotron 3 Ultra from Nvidia, roughly in line with Kimi K2.5, but around 10 points behind Kimmy K2.6, GLM 5.2 and GPD 56 sole, and around 20 points behind Fable 5. On coding, Inklinkling scored 54.3% on Sweet Bench Pro and 63.8% on Terminal Bench 2.1, again putting it ahead of Nemotron and K2.5, but a little behind K2. K2.
Starting point is 00:15:13 And add a significant shortfall compared to GLM 52, 56 sole, and Fable 5. On GDP Val, which measures real-world tasks and is a strong test of tool call execution, it scored 1238, placing it ahead of all major open models other than GLM 5.2 and Deepseek V4 Pro, but a long way behind 56 and Fable 5. TML also chose to highlight inkling scores on some of the audio and vision benchmarks, which typically aren't published for frontier models. On audio, it seems well ahead of most open models, a match for Quinn 3.5 Omni Plus and pretty close to Gemini 3.1 Pro, which is notable for its multimodal trading, and on vision it's roughly a match for Kimmy 2.5 and only slightly behind
Starting point is 00:15:50 K2.6 and Gemini 3.1 Pro. Artificial analysis gave Inkling a score of 41 on their intelligence index, ranking 19th behind Kimi 2.6, Deepseek V4 Pro and Minimax V2.5 Pro, and ahead of Deepseek v4 Flash in Nemotron 3 Ultra. It's 18 points behind 56 sole and 19 behind Fable 5. A.A. did note very strong token efficiency compared to open weights leaders, averaging around two-thirds as many tokens per task as 5-2, 2.6, and Deepseek V4 Pro. And yet, as regular listeners well know, benchmarks don't usually tell much of a story. Unfortunately, although there aren't all that many first impressions, the ones that we have aren't great. Professor Ethan Mollick wrote, happy to see a new open weights model, but so far,
Starting point is 00:16:33 Inkling is pretty rough in my tests, not close to frontier Chinese open weights models. As one example, here it is failing to pass the Lem test, done by every frontier model since DeepSeekar 1 and Sonnet 3.5. Ethan followed up, are people actually trying inkling? I can't seem to get it to work solidly on any of my tests even on extra high, and the chain of thought goes crazy at even simple requests. Am I missing something? Dreams ASI was even more harsh writing, thinking machines should rename two lobotomized machines. First impression of inkling feels like a GPT5.2 that's less accurate. Corporate guardrails galore, plenty of unnecessary refusals, and treating the user as by default an adversarial prompt injector. Lassan scaling 01 writes,
Starting point is 00:17:12 Benchmarks don't look that great. It's basically another Kimmy K2.6 and loses to all closed models in GLM 5.2. This makes it more feel like they are rushing this release before Kimi K3 and Deepseek V4GA. But this is not the only interpretation. Yun Feng Zhang writes, this is going to be an important model when we look back in a few years. The world needs, open LLMs from as many players as possible. Given TLM's talent, know-how funding, and compute, they will soon reach Opus, if not Fable 5 level. So basically, Yunfan's argument is that this should be treated as the first, not last model that TML is going to put out, and seen as a precursor for better things to come. Certainly a reasonable interpretation, but not even the most
Starting point is 00:17:50 positive thing out there. To your taxes, writes, I like that inkling is mediocre on benchmarks. This suggests that they haven't been cutting corners too much with distillation, so their independent data pipeline will shine through and it's not reducible to scores. the second, third update they become a meaningful player. So this is similar to that last comment in the sense that it's a precursor for better things to come, but also will point out that they're clearly not benchmark maxing in the same way that many Chinese models do. Some people viewed this in fundamentally different terms, though, where the benefits weren't just about the future, but about aspects of what we have right now. Burton Shaw wrote, the sickest thing about Inkling
Starting point is 00:18:24 is the manner of their release. They brought out a model with genuinely interesting architecture and applications, told us what it does and does not do, and gave us tools to serve and post train it, which we can choose to use or not. The open scene is beginning to get diverse voices with their own valid takes on how AI should be built. Jack Morris, though, goes farther, arguing that people are underestimating what a big deal this is. Jack writes, This is the only open weight model that's trained without distilling from open AI or anthropic. Kimmy distills. GL.M. Distills. Qwen distills. Nemotron distills Kimi and Deepseek, which counts. Basically, a fully different textack, the first pure open front front.
Starting point is 00:19:03 to your coding model. Now, a community note did point out that TML itself did say in their notes that while Inklink had been pre-trained from scratch, it did use a small bootstrap on synthetic data from open models including K2.5, as well as pointing out that other models like Lama 3.1 were also trained from scratch without open-AI and anthropic distillation. But the significant point that Jack is getting at is that in the world that we're heading to, with the unspoken piece of this being that enterprises are starting to pay more attention to open weights models, the fact that this is an American model that is not primarily distilled from one of the closed models could become more significant very, very soon. Meanwhile, open source expert Nathan Lambert pointed out
Starting point is 00:19:39 that clearly this is one part of the story alongside TML's Tinker platform. Nathan wrote, Inkling was somewhat inevitable once Tinker took off. Integration of post-training services across TML's entire stack, and it's one of the best open model business stories to date. Jeffrey Emanuel explored this dimension of it as well, writing, this is a pretty smart and differentiated business strategy, which makes sense because it's hard to compete head-to-head with the biggest labs at their own game. So you focus on their weakness, which is the emerging competitive paranoia among big companies about leaking alpha. To do that, you need open-weight models so that the companies can run it on their own infrastructure without leaking anything. But the problem with open weights is obviously how to monetize it because inference becomes a race to the bottom and commoditized in capital-heavy.
Starting point is 00:20:23 TML's solution seems to be to monetize the fine-tuning process of customizing their model for the particular the problems a company wants to solve, using that company's own internal data in a way that keeps the learning and benefits of that data exclusive to that one company. Smart! And you better believe that this is the thrust of TML's marketing for this. The last section of their launch post is called customizing inkling. And it starts, many real-world problems aren't solved well by even the best generalist models, with the gap being closed by fine-tuning that utilizes an organization specialized knowledge. The experience of our Tinker customers points in the same direction. Our post-training and results of reinforcement learning at scale
Starting point is 00:20:58 suggests that Inklinkling is capable of rapidly learning from fine-tuning. And once you view it as the model that was designed, to be the base model for Tinker, it starts to make a lot more sense. In another part of that blog post, they wrote, Inkling is designed to be broad. We train it across agentic, reasoning, coding, instruction following, factuality, vision, and audio tasks, rather than narrowly optimizing for one domain.
Starting point is 00:21:19 That breadth matters for customization and real-world use. Different users need models that can adapt to very different workflows, not just Excel on the benchmarks. Now, if you're watching Microsoft closely, this is clearly the direction they're pushing as well. Satya Nadella's recent blog posts on X have been all about how companies need to own their own learning infrastructure, and one of the big product pushes aligning with the new MAI models that Microsoft has released is their Microsoft Frontier Tuning product. And yet, Microsoft Frontier Tuning using MAI still requires enterprises to trust Microsoft with their data. Now, Microsoft is in a good position to cash in on the trust that they've built
Starting point is 00:21:53 over the past decades, making it less of a leap in a jump to trust them as opposed to an open AI or an anthropic for many companies, but that still does look very different to an open weights version of that, such as the one provided by inkling. Now, different companies are going to have different senses of the tradeoffs. There will be, I guarantee, companies who do experiment with that sort of post-training or fine-tuning approach, but who don't ultimately care that much about it being open versus closed. In other words, they want the cost benefits more than the data security benefits or are comfortable with Microsoft as the partner, regardless of the data security issues, but others you have to think might jump at a model that has even more
Starting point is 00:22:28 independence and sovereignty from a big player in the way that an open weights model like Inkling might represent. And when you start to pile the lawyers in the risk group on top of things, the fact that it's an American model that is not distilled becomes pretty unique relative to the other options out there. Daniel Kaplan thinks that this could be the beginning of an entire new subdivision of a services category. He writes, enter the forward deployed fine tuner. Thinking Machines' two GA products equal the perfect go-to-market for a new foundation model in the U.S. Want a custom on-prem model? Got an AI research team? Inkling plus Tinker. Full self-serve and you're welcome to pay us a moderate amount of money for additional expertise, early access to new models, etc.
Starting point is 00:23:04 Want your own custom model but no AI research team? Tinker seem confusing? No problem. We'll send over a forward-deployed fine tuner for a lot of money. Scales to big companies in AI startups who want to protect their alpha, AI app layer companies in search of cheaper tokens, and new startups founded by runaways from the behemoth labs. I think there is very clearly a logic here, and when both Microsoft and this buzzy lab are telling the same story, enterprises will pay attention. At the same time, the fundamental technology underneath this
Starting point is 00:23:31 is not without controversy. In fact, in some ways it runs counter to what people's broad perceptions have been about how generalist models compare to fine-tuned versions. Referring to the bitter lesson, which is a theory that oversimplistically says that human ways of learning aren't necessarily going to be the best ways of machine learning, and that in general the best strategies for machine learning are going to be giving it as much
Starting point is 00:23:53 data and compute as possible and letting it do its thing, the rather than overly relying on unique human expertise, Simon Smith writes, So thinking machines is basically a bet against the bitter lesson? That's how it feels, and I'm personally not convinced. From my experience, fine-tuning models, it's way more effort than people think, that effort is ongoing to address new data edge cases and model updates, models can lose capabilities or have unexpected issues introduced, and ultimately a big general model with a bit of context, eG skills files, comes along and beats your hard work, and then the cost of models with that capability drops off a cliff. When people talk about tokens and fine-tuning and running
Starting point is 00:24:27 their own models, they often only factor the cost per token of their fine-tune, not the fully-loaded cost of continuously collecting and curating data, developing and maintaining data and training pipelines, deploying and administering server infrastructure, and so forth. Anyway, time will tell, my money's still on the bitter lesson. I think it's a super important point, and one of the really important things to keep in mind right now is that the recognition that there is a problem and a challenge, with both data sovereignty and the cost of tokens, does not mean that the immediate answers that the market is jumping into provide, i.e. Microsoft Frontier tuning or Tinker, is ultimately going to be the right answer. In just a few sentences, Simon Smith pointed out all of those different
Starting point is 00:25:04 challenges that may end up showing that this sort of fine tuning isn't necessarily the right way to solve those problems. The flip side, though, is that if you are TML right now, you have to feel good that your solution, because it is open weights, is addressing both sides of that equation, both the token cost side as well as the data sovereignty side. In other words, it might be that a fine-tuning solution based on a closed model like Microsoft is two in between, and that even if it's a niche market that requires both the data sovereignty and the cost side for the people who actually care about it. We don't know, but it's very clear that this is going to be an emerging trend and one to keep an eye on. Former White House AI advisor and A16Z investor Shreryam Krishnan
Starting point is 00:25:41 wrote, It is clear open source models and harnesses are having a moment. There's a few factors at work. One, it is now obvious that you can catch up to near state of the art performance and do so with a clear training lineage. See, TML Zinkling launched today. Two, there are several well-funded talented teams building open-weight models now in the U.S. and abroad. Along with the explosion of other near state-of-the-art models like Grock and Cursor, Mew Spark. It's clear that we are going to have a diverse ecosystem of models, at least on coding and agendic use. Three, organizations are increasingly looking for control over how their data is used and are willing to trade off some access to frontier level tokens for this control.
Starting point is 00:26:15 Organizations and countries are increasingly nervous about the frontier labs potentially competing with them down the road and don't want their data to enable a future competitor. Four, open source is a slider. You could bring your own open harness, your evals, your business context, and are free to pick and choose your model of choice. Five, companies have now actively shifted from how do we get our people to use tokens, to being uncomfortable with their token costs ballooning without a clear line to revenue. Six, geopolitically, countries will be weighing open-weight models as a way to get frontier-level tokens inside controlled environments that may not be otherwise possible.
Starting point is 00:26:46 All of this leads to more choice for all of us. And I think this is the key point. A few months ago, one would be forgiven for thinking that the way things were headed, pretty much every enterprise was just asking the question of whether it should sign up with OpenAI or Anthropic or both, and that was the extent of their decision. Now, we're talking about so much more choice not only in terms of models, but in terms of the harnesses in which model lives, and complex model architectures which could involve multiple models, and entirely new approaches to models such as these fine-tuning approaches. It feels very much
Starting point is 00:27:15 like a moment where we're going to see a lot of new shots on goal and new explorations emerge. Now, as is always the case, when a whole bunch of new flowers bloom, not all of them will work. It may be that the juggernauts in the space, figure out where there's adoption and uptake over the next six to 12 months, take in the best parts, and we're back to most enterprises simply choosing between a very small handful of pre-described solutions. But it's very hard to imagine that, A, this experimental period we're moving into, won't have a significant and deterministic impact on the shape of those solutions, and B, that there isn't a ton of room for alpha, especially among small teams and more nimble enterprises that can take advantage of this flourishing and
Starting point is 00:27:56 highly competitive period. That doesn't mean that you at enterprises all need to rush out and sign up for Tinker or Microsoft Frontier tuning or anything like that, in the same way that you don't need to immediately switch all your activity to open router or go start hacking on GLM 5.2. But at this point, it's pretty clear that if you are in charge of any sort of enterprise buying decisions and you're not at least playing around with and experimenting with or closely paying attention to all these different solutions, you're missing quite a bit of opportunity.
Starting point is 00:28:23 I will continue to keep track of these things as they develop, but for now, that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always. and until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.