The AI Daily Brief: Artificial Intelligence News and Analysis - What the Heck is Graph Engineering?

Episode Date: August 10, 2026

Graph engineering is AI’s latest buzzy term—but it offers a useful framework for organizing agents, tools, knowledge and humans into working systems. NLW explains the evolution from prompts to gra...phs. In the headlines: OpenAI delays Astra, ByteDance trains a massive model, open-weight AI tests revenue sharing and Claude Code embraces Auto Mode.AIDB's AI Summer Adventure: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://summeradventure.ai/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Brought to you by:KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at ⁠⁠⁠⁠https://kpmg.com/us/Sophisticated⁠⁠⁠⁠Hyperagent - Hire a fleet of always-on agents. New users get $1,000 in inference. ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠hyperagent.com/aidailybrief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.rackspace.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Section - Section turns AI investment into workforce transformation and ROI - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.sectionai.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Blitzy - Want to accelerate enterprise software development velocity by 5x? ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://blitzy.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠AssemblyAI - The best way to build Voice AI apps - ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.assemblyai.com/brief⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Robots & Pencils - Cloud-native AI solutions that power results ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://robotsandpencils.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠The AI Daily Brief helps you understand the most important news and discussions in AI. Subscribe to the podcast version of The AI Daily Brief wherever you listen: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://pod.link/1680633614⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Our Newsletter is BACK: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://aidailybrief.beehiiv.com/⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠Interested in sponsoring the show? sponsors@aidailybrief.ai

Transcript
Discussion (0)
Starting point is 00:00:00 Today on the AI Daily Brief, what the heck is graph engineering and why should you care? Before that in the headlines, OpenAI's Atlas model gets a cyber delay. The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI. All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, robots and pencils and hyperagent. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe on Apple Podcasts. And to learn more about sponsoring the show, send us a note. at sponsors at AI Daily Brief.aI.
Starting point is 00:00:37 Late last week, rumors were swirling that OpenAI's latest model, codenamed Astro, was being prepared for an imminent release. Sam Altman even traveled to Washington to preview the model and discuss new model testing policies. In the background, however, the discussion around the Hugging Face hack just continued to grow in prominence and significance. For those who missed that episode, OpenAI's technical breakdown at the Black Hat Conference revealed that not only had their model escape the sandbox and hacked into HuggingFace's servers, it also left internal notes instructing future models on how to pull off the same trick.
Starting point is 00:01:07 On Friday, OpenAI decided to make a big shift. They wrote, Our latest internal evaluations of Astra, one of our upcoming models over the past few days, indicates significant advancements in agent decoding in cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our preparedness framework. Open AI defines that critical threshold as the ability to, quote, identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical
Starting point is 00:01:35 systems without human intervention or devise and execute end-to-end novel strategies for cyber attacks against hardened targets given only a high-level desired goal. Now, on this front, GBT56 sole had been assessed in the high category, which was a little more risky than previous models but still appropriate for a release. Given that the new Atlas models are now in the critical category, as a result, OpenAI is holding back the model from release while beefing up internal safety measures. Testing environments will now be isolated, model weights will have enhanced encryption to prevent leaking, and additional sandbox monitoring will be implemented. OpenAI will also be limiting internal activities using Astra that don't yet meet these enhanced security measures.
Starting point is 00:02:13 On X, Sam Altman added some context around the decision posting. Astra is a powerful model, and we're working to make it generally available. We do not think it is a good strategy to keep powerful models to a chosen few. Given its cyber capabilities, we need a little bit longer to do this safely, but hopefully not too long. Now, one thing we don't know is to what extent this is an OpenAI voluntary pause versus a government-imposed pause, or whether that distinction even matters at this point. One interesting note is that there aren't a lot of folks suggesting that this is just a publicity stunt, as was one of the narratives surrounding the mythos release. Basically, the hugging face incident seems to have made the case that advanced cyber capabilities could be a real concern. Open AI head of strategic futures, Dean Ball, noted that this year is the first big test of weather frontier AI labs would follow their stated safety preferences when push comes to shove.
Starting point is 00:02:57 He wrote, our next model, Astra, may be critical under our preparedness framework. We cannot rule out the serious possibility that it is, and so we are going to take steps consistent with the higher risk level, critical, rather than assuming the model is at a lower risk level. Some of these decisions have the effect of slowing down internal development, and in that sense, they are costly decisions. But they are the right decisions I am proud of Open AI for making them. Now, a lot of the discourse surrounding this is what sort of changes Open AI can actually make to the guardrails in monitoring around these models.
Starting point is 00:03:26 OpenAI's RSI preparedness lead, Micah Carroll, wrote, as part of our response to cybercritical, we've expanded chain of thought monitoring to cover all agentic applications of Astra, including training and evaluation. Flags trigger a security response to review and interrupt high-risk activity. Now, at the same time, there's also some skepticism around that sort of chain of thought monitoring, but these are the types of discussions and experiments you're going to see a lot more of now, where I believe there will be a significantly increased investment in the resources to properly support models that won't be able to be released to the public
Starting point is 00:03:54 without it. Now, speaking of big, powerful models, Bight Dance is reportedly training an ultra-large model comparable in size to Mythos. The Financial Times reports that BightDance is in the early stages of a training run that will result in a base model with as many as 10 trillion parameters. So far, we've only seen a couple of large-scale training runs out of Chinese labs, with Kimi K-3 weighing in at 2.8 trillion parameters and Alibaba's Quinn-38 max at 2.4 trillion. Anthropic doesn't disclose model sizes, but the best estimates have Mythos at around 8 trillion. and Opus 4-8 at around $3 trillion for some comparison. Sources said that the training run could take 3 to 6 months to complete,
Starting point is 00:04:31 with more time added for reinforcement learning after that. Bight Dance also hasn't determined how large the final model will be on release. Still, this could be the first Chinese pre-training run that's truly on the frontier, and while model size is not a guarantee of performance, this could put Bight Dance back in the conversation for leading Chinese labs. That is, of course, all the more relevant if they follow through with their pledge not to distill from Western AI models, as reported last week. Brookings Research Fellow Kyle Chan wrote,
Starting point is 00:04:56 Chinese AI labs seem confident that they have the compute needed to pre-trained 5 to 10 trillion parameter models. Some way, somehow, compute does not seem to be such a major bottleneck, at least when it comes to reaching these levels of model size. Geopolitics commentator Dmitri Alperovich responded, why would there be when there are no restrictions on remote access to compute and the export controls on chips are full of holes like Swiss cheese? Speaking of, some of those holes do seem to be top of mind in Washington.
Starting point is 00:05:23 Shortly after the release of Kimmy K3 last month, the New York Times co-related research on the flow of compute from large-scale data centers in Southeast Asia. According to semi-analysis, the Oracle data center in Malaysia was being used almost exclusively by ByteDance. The data center was powered on in mid-20205 and contains over 100,000 Nvidia Blackwell GPUs. Think Tank China talk determined that Oracle makes up around 22% of China's total supply of compute. Bloomberg also reported last month that Moonshot had access to 20,000 Nvidia H-200s to train Kimmy K-3. This compute was reportedly provided by Alibaba, although they deny the cluster contains
Starting point is 00:05:57 H-200s. Now, an H-200 cluster of that size shouldn't be possible under the current export controls as licenses haven't yet been issued. Bloomberg was unable to confirm the location of the cluster, but heavily implied it could be located overseas and rented by Alibaba. The information offered conflicting reporting that the training cluster actually contained current generation black Blackwell chips. And around the same time, an administration official accused moonshot of acquiring blackwell chips and setting them up for remote access in Thailand. Now, this is all legal. The export control regime only prohibits the import of advanced chips and does nothing to stop them from being installed in another country and then leased to Chinese firms. During its final weeks, the Biden administration
Starting point is 00:06:35 proposed rules that would restrict chip supply routed to third countries, but those rules were scrapped on day one of the Trump administration. The Trump Commerce Department is now reportedly looking into the practice. Per source is familiar with the situation, Bloomberg writes, the effort involves compiling a list of countries with alleged black market operations to get restricted inVIDIA chips physically into China, well within enforcement's usual purview. But the division will also drop a list of countries where Chinese firms access the chips remotely, which isn't typically an enforcement question because it's not illegal. Draft rules that prohibit the export of advanced AI chips into Malaysia and Thailand have been circulated, but none have made it past
Starting point is 00:07:09 the drawing board. Part of the issue is the convoluted nature of these arrangements. According to Bloomberg, Alibaba accesses chips in Malaysia through a Singaporean shell company controlled by a Cayman Islands entity, which is ultimately owned by Alibaba, meaning of course that simple identity checks on compute supply are unlikely to be effective. The chips are also installed already, meaning that a forward-looking crackdown on imports would do nothing to the existing data centers. Speaking of Alibaba, an interesting business model experiment might be afoot. During last week's release of Quen 3.8 Max, some were surprised that Alibaba would publish the weights to their flagship model. Earlier this year, Ali Baba
Starting point is 00:07:44 had signaled a shift away from open source, with Quen team founders stepping away from the company, management signaling a more commercial direction, and multiple flagships, including Quen 3.7, being kept as closed source. That's why folks were very excited to see that Quen 3.8 Max was released with the announcement that the full waits would be forthcoming. According to Reuters, however, there is a catch. Reports suggest that Alibaba plans to demand revenue sharing from large commercial users. And while the specifics of revenue sharing deals are still being finalized, we can perhaps look at Moonshot's Kimmy K3 release for the blueprint. Moonshot kept the weights proprietary for the first week to ensure that they captured all
Starting point is 00:08:19 the curiosity revenue revenue from the release. After that, they reportedly signed 30% revenue sharing deals with all major inference providers, allowing Moonshot to retain pricing power by limiting how much the model can be discounted through other providers. Even now, none of the suppliers on OpenRouter are offering K3 for more than a 7% discount. Patty Shrinivasan, the CEO of inference provider DigitalOcean, described this as the freemium model for AI. And Korean aggregator, Cozy Bear, added, everyone is trying to figure out how to get paid for open weights and revenue share is the most honest attempt so far. It doesn't pretend the model itself is the product. It treats the model as infrastructure with a toll booth. The interesting part
Starting point is 00:08:53 they continue is enforcement. You can't track who's making money on top of your weights, so this is mostly as a signal to enterprises. Use Quen and when you win, we want a seat at the table. More like our relationship contract than attacks. Open source AI just entered its licensing era. Lastly, today, one functional update. Auto mode is now the default for Claude code, marking a big transition point for work automation. Automode allows the user to set Claude on a task that gets completed without interruption. Claude will only prompt the user if a code change is extremely significant, i.e. irreversible, destructive, or aimed at something outside of your environment.
Starting point is 00:09:27 Auto mode was first introduced as a preview feature in March, with the version prior to that being the dangerously skip permissions command. The alternative to that was hitting Enter every couple of minutes to skip the latest notification and keep the session going. After iterating on Auto Mode, however, Anthropic now believes that skipping permissions isn't that dangerous and, in fact, could actually be safer than seeing a prompt for every code change. They conducted a study with over 1,000 testers, finding that auto mode caught 89% of harmful actions. Human reviewers only caught 13.6% of harmful code changes. Anthropic suggested this is down to approvals becoming basically automatic, with their users
Starting point is 00:10:01 approving 97% of code changes. One of the interesting things about the evolution of Auto Mode is that it has forced Claude code to work in a safer way. The system uses a classifier to detect destructive or irreversible code changes and block them, and when this happens, Claude typically finds a safer way to achieve its goal. Only once it runs out of safer options, does it alert the user? Auto mode will now be the default for Pro, Max, and team plans, but will remain opt-in for the enterprise. Still, when it comes to those enterprise users, Anthropics suggest there are significant benefits to using Auto Mode. They claim that auto-mode users ship 25% more PRs, and that many organizations, including Adobe, Gusto, and Garner Health, are already running Auto Mode as their production default.
Starting point is 00:10:40 For the team at Anthropic, perhaps unsurprisingly, Auto Mode is the default, with CloudCode creator Boris Cherney commenting, the team and I use Auto Mode exclusively and have been for many months. I couldn't imagine going back to permission prompts. Really excited to get this out to everyone. The ascendancy of Auto Mode is another example of how our default patterns of interacting with AI are changing, which provides the perfect segue to our main episode, a primer on the latest buzzy buzzword, graphic.
Starting point is 00:11:04 engineering. If you're leading AI inside an enterprise, you already know that the gap right now isn't capability but execution. That's why KPMG's You Can with AI is back with a new season featuring conversations with leaders like Sorosia Chatterjee of Emma, Mayhabib of Writer, McKess and CIO, Elery Fisher, and others focused on practical execution. What's working, what's not, and what it actually takes to move from pilots to real scaled impact, across strategy, data readiness, governance, workforce, and value. And of course, it's co-hosted by me, Nathania Whittimore. Go listen and then subscribe at www.kpmg.us slash AI Podcasts. That's www.kpmg.comg.com slash AI Podcasts. Every AI coding tool on the market does the same thing first. It starts writing code. Blitzy does
Starting point is 00:11:54 the opposite. Before writing a single line, Blitzy spends days reverse engineering your entire codebase. Thousands of agents ingest millions of lines mapping every dependency, every undocumented constraint, every architectural decision made over the last decade. The result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building. Other tools guess at context with grep searches and markdown files, Blitsey never guesses. It builds true understanding first, then delivers over 80% of entire software epics autonomously. Validated, end-to-end tested production grade pull requests. That's why Fortune 500 engineering teams trust Blitzy with the codebases that matter most.
Starting point is 00:12:29 See for yourself at blitzy.com. That's BLITZY.com. One thing I keep seeing in Enterprise AI, companies hedging across every cloud, every model, every framework, or paying a GSI for a pilot that never ends. The team's actually shipping, they've picked a lane and they move fast. That's one of the reasons I like today's sponsor robots and pencils. They've gone all in on AWS. They're an advanced tier and AWS pattern partner, and they ship production AI co-workers in 45 days. That's led to them doing some of the more interesting work I've seen on AI co-workers.
Starting point is 00:13:01 And by that I'm not talking about chatbots. I'm talking about actual agentic systems that sit inside a business architecture and do real work. That kind of focus matters if you're an enterprise leader trying to get something real into production or an AWS rep trying to move a customer from interested to deployed. Request an AI briefing at robots and pencils.com. One conversation with robots and pencils and you'll know. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.
Starting point is 00:13:29 New users get $1,000 in inference. Forget local agents and chat workflows waiting on your laptop to be prompted. Hyperagent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing's agent turns competitor moves into landing pages. Sales as agent enriches leads, drafts emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals. It's time you add agents that feel like teammates. Hire yours at Hyperagent built by the team at Airtable.
Starting point is 00:13:58 claim your $1,000 in inference at hyperagent.com slash AI Daily Brief. We'll go back to the AI Daily Brief. Today we are discussing the latest buzzy term on AI Twitter, which is graph engineering. Now, this one admittedly is a little confusing because A, it sort of started tongue in cheek, and B, depending on who's talking about it, it kind of is describing two different things. But I actually think that the concept at least is useful to situate relative to the lineage of engineering we've had from prompt to context to harness to loop, and so I think it's worth doing this primer.
Starting point is 00:14:38 And indeed, this is definitely more a primer on graft engineering than a complete guide to graph engineering. I want to bring you up to speed on what this term is and why I think it matters. So the tweet that kind of kicked off this discourse came from OpenClaw creator Peter Steinberger. Back in mid-July, he wrote, are we still talking loops or did we shift to graphs yet? AI creator Matthew Berman captured the feelings of many when he said, bro, stop, I'm on vacation. And while almost immediately the Twitter article boosters got to work, big bold declarations like loop engineering is dead, long-lived graph engineering were visible all over the place.
Starting point is 00:15:12 But after this initial phase of hype and bluster, there is in fact actually something interesting here. So let's talk about all of the things that have had engineering around them. The first blank engineering that we had was, of course, prompt engineering. This was what a lot of the AI courses around 23 and early 2024 were all about. And depending on which corporation you look at, still unfortunately the sub-examination. absence of a lot of those upskilling courses today. The idea of prompt engineering was a recognition even then that we were shifting how we did work. Instead of doing all the work ourselves, we were
Starting point is 00:15:43 debutizing an AI chatbot to do some amount of that work. Now, whether that was final production or just some intermediate step like research, we still needed to find ways to optimize what we were asking for to get the best results. That was prompt engineering. At various points in the life cycle of prompt engineering, you had tips and tricks ranging from telling the AI to pretend it was a certain type of person, to fancy JSON engineering, which used this complex way of typing to theoretically better structure requests, and we kind of had every other thing in between as well. Heading into 2025, however, we started recognizing that the prompt was only one part of getting the most out of AI. We didn't just need to be good at asking for things in the right way.
Starting point is 00:16:22 We also needed to be good at giving AI all of the information and knowledge it needed to do a good job with whatever that prompt was. To take a simple example, if you're asking your LLM to create a highly successful on-brand marketing campaign, well, first of all, it needs to know what on-brand is, which means giving it access to brand guidelines and potentially other write-ups in the past about things like brand values, and in order to have it not just be guessing at what good means, it would probably be helpful to give it stats and analytics from previous marketing campaigns, perhaps with some subjective reflections on what worked and what didn't as well. That body of information that surrounds the prompt is the context. Context engineering was all about making sure
Starting point is 00:17:00 that all of that type of information was accessible in the right way. Now, here at this point, we also have an interesting split, which I think we're going to see once again with this latest graph engineering term. The split, broadly speaking, is between technical folks and software developers and everyone else. For the software developers, context engineering wasn't just a matter of making sure that your AI had access to the right files for the job. It was, in fact, an actual engineering task. It was thinking about not just what information is useful, but designing the technical systems by which the AI could traverse the web of accessible context in a way that didn't just spend the entire context window. On this side, context engineers were actually thinking in terms
Starting point is 00:17:38 of context budgets, making sure, for example, as they were designing applications, that certain parts of a process didn't get bogged down in context, while others could go deeper when they needed it. So here we have context broadly referring to the same thing, but engineering being a literal engineering task for the engineers and a mindset for everyone else in terms of how they organized information around the LLMs that they were using. Now this year, just like everything else is sped up, we've also had a speed up in the succession of blank engineering type of terms. Starting with, you've probably heard me talk about harness engineering. The harness is, of course, the environment that exists around a model. On a simple level, that might be an actual software tool like
Starting point is 00:18:14 Claude Code and Codex, but the more expansive definition of harness includes everything from the tools to the permission sets, to the skills files that an AI or agent has available to it to do its work. Throughout 2026, people have become more and more comfortable with the idea that the agent is actually a combination of the model and the harness that surrounds it. This is why, for example, frequently when we're getting new benchmarks, companies will now explain what harness the benchmarks were run in, as that's actually an important part of the story. Now, when it comes to harness engineering, once again, we've got two very different meanings of engineering. There are, of course, the actual engineers and software developers who have been building different and better harnesses
Starting point is 00:18:50 and trying to advance our understanding of harnesses in general. And then there's the more individualist sense of harness engineering, which is about things like which skills you surround your agents with and what tools they have access to. Now you'll see here that as each of these new terms comes online, it's not like the old one goes away. Prompt engineering is certainly the one where there's probably at this point the least leverage to be had, but it's not like because we started to understand the value of harnesses, all of a sudden context stopped mattering. In fact, quite the opposite, the harness became a new context for that context engineering. So we've got the prompt which controls the instructions.
Starting point is 00:19:24 We've got context which controls with the Model C's, the harness which controls the environment, and that brings us to the loop, which controls the iteration that an agent goes through to accomplish a goal. Loops or loop engineering, which has become a big topic over the last few months, is about thinking about your relationship with agents in different ways. The canonical short explanation once again came from Open Clause Peter Steinberger, who said, you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. Loops are the systems by which an agent can observe, plan, act, check results, and repeat until some measurable stop condition is reached. As we've discussed in the past on this show,
Starting point is 00:20:03 one of the big challenges for non-engineers who have been trying to put loops to work is in figuring out which aspects of their knowledge work have those sort of measurable stop conditions. One of the things that we discussed in my show about loops from a month or two ago was this idea that in some cases where there wasn't a natural measurable condition to get an agent working in this sort of loop, you were going to have to precisely define something measurable like that to actually get the loop to work. But what you'll notice here is that a loop is about how to get the most out of a single agent or agentic process. It is a work backwards from a specific goal that gives the agent the repeatable steps it needs to follow that
Starting point is 00:20:38 as many times as it's necessary to actually achieve that goal. But what about when a goal is more complex and requires multiple different processes interacting to actually accomplish whatever that goal is? What about when we move beyond, the output being a single agentic process to actually designing an ongoing agentic system for working? That's where we get into graph engineering.
Starting point is 00:21:00 If prompts control the instruction and context controls with the model C's, harnesses control the environment, and loops control the iteration, the graph controls the new agentic organization. Graft engineering is about designing how multiple agents, tools, knowledge sources, and humans interact and connect. Graft engineering describes both the parts of the system, which some people are referring to as nodes, that could be agents, routers, or human gateways. And graphed engineering also explains the interactions between those nodes, which handoffs are permitted when information or state travels between them. Explainx.aI explained a loop as an autonomous cycle for a single agent.
Starting point is 00:21:37 A trigger fires, an agent acts, a verifier checks, and if not done, the whole system is retried with updated context until the goal has been met. As they write, every guardrail, i.e. max iterations, token budget, et cetera, applies to one agent's run. The loop is the agent's behavioral contract with itself. A graph, on the other hand, is an organization of agents. As they put it, each node in a graph, is an agent running its own loop. The edges or interactions between those nodes define data flows and dependencies. They continue, the graph specifies who exists, which agents and with what specialization, what each owns, their domain context and tool access, how work moves, sequentially, in parallel or conditionally, i.e. is this a loop between different loops?
Starting point is 00:22:22 And finally, the graph defines what happens on failure. Is the node retried? Is it routed to a fallback? Or is there an alert upstream? The graph they conclude is the organization's operating structure, loops live inside nodes, the graph connects them. In short, a loop is how an individual agent does its job, where a graph is how an entire agentic organization works. As Google Shabam Sabu put it, loops made agent behavior programmable. Graphs make agent organizations programmable. Now what I don't expect is for all of you to run out and start designing complete agentic organizations. But the idea of graph engineering is to be able to think in those system terms. and even slightly differently from the context and harness engineering,
Starting point is 00:23:03 with loops and graphs, while sometimes you will use both of these patterns, there will be times when the simpler architecture of a single loop is going to be fine. When a single job has a clear finish line with genuinely sequential steps, and one agent's context window able to hold the whole domain, that's a good candidate for a single loop. When the work instead splits into specialties with different handoffs, when parallelism becomes valuable, when different steps in a process want different models or tool sets,
Starting point is 00:23:29 when routing has to be explicit, and when you want to design a resilient system where the failure of one node doesn't take down the rest, that's where you get into this actual graph engineering. Now, as tends to happen, as soon as we get a new term, we very quickly explore all the nuances as well. One discussion, for example, that you're seeing a little bit is the difference between orgraphs versus work graphs. Org graphs are going to be more stable agentic systems that defines a more permanent style of setup. The orgraph is going to have long-lived agents with each agent. owning a domain and accumulating context over time, preserved memory, and stable relationships and dependencies that don't change unless explicitly told to. So if you are designing an agentic organization that's going to do the same thing over and over again, for example, if I was designing
Starting point is 00:24:13 a multi-agent system that automated the process of going from research to production, to editing, to publishing, to extracting insights, to posting, that might be a good candidate for an orgraph because these are ongoing recurring processes. This is just how my work happens day in and day out. The work graph, on the other hand, is more dynamic and more ephemeral. It can include task nodes that only exist as long as the work exists. Dynamic edges, remember edges, are the interactions between the nodes that can split or merge, an adaptive structure where tasks can disappear when evidence makes them unnecessary, or new tasks can be spawned if new complexities are discovered.
Starting point is 00:24:49 But let's wrap up this primer by coming back to the main point. Like I said at the beginning, my expectation is not that all of a sudden you go out and design complex agentic organizations now that you're acquainted with this wonky concept of graph engineering. What I think is valuable for all of us, however, in the same way that even if you weren't designing loops, understanding the architecture of a loop, a trigger that fires, an action that's taken, a validator that checks the work, and then that on repeat until it's done, that is extremely helpful in thinking about how to use agents to automate chunks of your work. In the same way, what I think graph engineering will unlock for many is the ability to start to start to start to
Starting point is 00:25:27 thinking in multi-agent systems terms, where you can start to see different agents with different jobs and actually understand and even design their relationships with one another. Over time, some of you, I guarantee, will start to design those more complex agenic systems, and the best practices and lessons and tool sets that people build around this graph engineering discipline are going to be extremely useful when you do. So yes, graph engineering is the latest buzzy buzzword, and some of the early tweets about it were frankly tongue-in-cheek, but design Agenic Systems is, I believe, a new work primitive, and something which we will increasingly be called upon to do. So hopefully you now have a better sense of that and can dig in as
Starting point is 00:26:06 makes sense for you. For now, that's going to do it for today's AID Daily Brief. Appreciate you listening or watching as always. And until next time, peace.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.