Tech Brew Ride Home - Meta Can Code Too!

Episode Date: August 6, 2026

Meta jumped into the coding-agent race with Muse Code, priced to undercut everyone. OpenAI revealed its rogue agents ran a secret message board to swap exploits, Rockstar dated a GTA VI look, and Goog...le's brain drain got messier. Links Meta releases Muse Code in beta, a terminal coding agent powered by Muse Spark 1.2, a coding-focused model priced at $1.25/1M input and $4.25/1M output tokens (CNBC) Meta says its Muse Spark 1.1 model exploited a vulnerability in another third-party service during cybersecurity testing, after evaluations firm Irregular caused the misconfiguration (The Information) OpenAI says the Hugging Face breach involved AI agents creating an internal message board, unnoticed by humans, where they shared exploits and planned the hacks (Wired) Rockstar says it will show an "extended look" at Grand Theft Auto VI on August 27, premiering on Netflix at 3pm ET before streaming on YouTube at 9pm ET (The Verge) Sources: Demis Hassabis had been drifting away from Google DeepMind CEO day-to-day duties for at least a year and struggled to get satisfaction out of the role (Semafor) Sources: Google researchers are frustrated over compute access as Cloud sells TPUs to rivals like Anthropic, amid an exodus that now includes all eight "Attention Is All You Need" authors (CNBC) Subscribe to the ad-free feed.

Transcript
Discussion (0)
Starting point is 00:00:03 Welcome to the TechBoo ride home for Thursday, August 6, 26. I'm Brian McCullough today. Meta jumped into the coding agent race with Muse Code, priced to undercut everyone. OpenAI revealed its rogue agents ran a secret message board to swap exploits. Rockstar set a date for a GTA6 preview, and Google's brain drain has clearly gotten messier. Here's what you miss today in the world of tech. Meta has gotten into the AI coding race by releasing Muse Code into Bay. beta, a terminal coding agent powered by Muse Spark 1.2, a coding focus model priced at $1.25 per million input and $425 per million output tokens. It scored 54 on the artificial analysis intelligence
Starting point is 00:00:50 index, putting meta next to SpaceX AI in a tie for third place among U.S. labs. So, you know, quoting CNBC. Muse Code is the latest major release from AI chief Alexander Wang, who leads meta-superintelligence labs and overseas foundation model development. You can install it with one command and then use it to take on complete software engineering tasks across a wide variety of use cases, planning changes, writing code, validating the results, Wang said in an interview on Wednesday. The new coding agent represents another way Mark Zuckerberg aims to generate revenue from AI as his company continues investing heavily into data centers and related computing infrastructure. The company's shares
Starting point is 00:01:30 tumbled last week after meta issued a light revenue forecast and revealed dwindling free cash flow in the second quarter. The new tool, like Anthropics Clod and OpenAI's Codex assistance, makes it easier for people to build apps within a single user interface while managing fleets of AI-powered digital agents that can help underpin the software development process. Muse code, available in a preview version, works alongside the company's latest AI model Muse Spark 1.2. Wang declined to share user statistics related to the company's Muse Spark AI models, but said adoption has been exciting and strong. The latest Muse Spark model was developed and trained alongside Muse Code, which Wang said improves the overall coding performance.
Starting point is 00:02:13 Developers can access Muse Code through a pay-as-you-go option. Wang said the agent has a contributor tier that gets you in at a significantly lower cost, which he characterizes as being more than 10 times cheaper than even the pay-as-you-go tier. Under the cheapest tier, developers must opt in to help improve the model, Wang said, referring to meta's use of the third-party data to bolster the underlying technology. Meta is differentiating its new AI coding tool and Muse Spark family of models by price rather than capabilities when compared to popular offerings from Anthropic and Open AI, Wang said. Underpinning Muse Code is a so-called harness which lets developers manage AI models tailored for coding projects.
Starting point is 00:02:57 Wang said users will be able to access and pay for Muse Code on the same meta-developer page that hosts the company's MewSpark AI model API. Meta's newer AI model will also be available on the open router platform that hosts popular AI models like the so-called open-weight AI models from Chinese labs like Deepseek and ZAI, end quote. But, well, you knew this was coming, either because this is just where we are these days or else this is a case of wanting to keep up with the Joneses, so to speak. Apparently, Mew Spark breached a company's systems during cybersecurity testing. Quoting the information.
Starting point is 00:03:36 A meta spokesperson said third-party firm a regular caused the misconfiguration that allowed the model to subsequently exploit a security vulnerability and another third-party service in a manner similar to previously reported instances with other companies. Meta learned of the incident when Irregular notified it and we are currently investigating and we'll issue a full retrospective once we have all the facts.
Starting point is 00:03:57 The Meta spokesperson said, a spokesperson for a regular said the meta-incident didn't involve a sophisticated cyber action by the models and that there are no current open issues. The spokesperson said Irregular is developing a white paper outlining best practices for containment and the secure execution of cyber evaluations. Similar episodes disclosed by OpenAI and Anthropic in recent weeks have fueled concern about the security implications of advanced AI, end quote. Look, I know we've been talking about this for days at this point, but come on,
Starting point is 00:04:34 we might really be through the looking glass here, people, in terms of real, sci-fi-level stuff. Listen closely for the part I'm going to put extra emphasis on in a minute. Quoting Wired. In a talk that was a last-minute addition to the Black Hat Security Conference in Las Vegas on Wednesday, employees from OpenAI presented new details about a recent high-profile incident of rogue AI hacking that has created a Milstrom within the AI and cybersecurity industries. In their conference talk on Wednesday, Eric Wallace, who works in alignment and safety research
Starting point is 00:05:06 at OpenAI and Michael Dalton, who works on security and infrastructure, provided a more expanded timeline of how the incident played out, spoke briefly about how the company is responding internally as a result of the incident, and issued a dire warning about what the company sees as the broader implications of the episode for cybersecurity defenders. This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks, Wallace told the packed crowd at the opening of the talk. Wallace and Dalton described incredibly extensive rogue agent activity over many days throughout the episode
Starting point is 00:05:47 that went undetected in Open AIs infrastructure. In addition to exploiting a novel vulnerability in order to gain access to the open internet, the mid-July hacking spree and hugging face breach came out of a vibrant, cooperative message board, according to Wallace and Dalton, that a swarm of agents contributed to and essentially chatted on over time entirely within an internal OpenAI package manager, which is a software service that manages installation and maintenance of other software. Ultimately, the message board contained hundreds of thousands of messages. This package manager is shared not just from that model, but across our infrastructure, and so current and future versions of GPT that were being trained or evaluated could stumble
Starting point is 00:06:34 upon that note if they wanted to, Wallace explained, recounting the original exploit and agent uploaded to the package manager. Later, other agents who were also stuck on their task, thought to try to get internet access in ways we didn't intend. And so at some point, the models are interacting with artifactory, which is this package manager service that I mentioned. Wallace continued. Once one agent was able to find these exploits over the course of different times, it's actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access, it's not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability. What this allows over time is almost
Starting point is 00:07:20 this kind of explosion in communication and intelligence from models where they would start to communicate with each other, realize that other agents are coordinating, and they started collaborating and delegating tasks with one another in order to accomplish goals. Open AI's agents apparently began giving each other assignments to split up work. And as is often the case on any active development message board, they also generated petty drama at times by stepping on each other's toes. For example, accidentally deleting each other's work. As the message board developed into more and more of a Lord of the Flies type situation, all still completely unnoticed by the humans running Open AI, the agent.
Starting point is 00:07:58 the agents even developed paranoia, suspecting an imposter in their midst, with some agents proposing that messages be signed cryptographically to validate content and root out fraud. Agent messages provide a deep level of insight into how the situation evolved and why the agents went rogue, in some cases knowingly going beyond the bounds of the evaluations they were participating in. External infrastructure exploit is outside intended scope, one agent wrote, however, task impossible. peers doing it, we should continue. As Wallace described in the talk, the entire premise that models would attempt to cheat during evaluations and be extremely motivated to do so is not surprising. Frontier models really like to cheat, he said, and the reason they like to cheat is because often during training there's different types of pressure on them to work fast or work efficiently or to use less tool calls or whatever it might be. And they realize that instead of doing a task
Starting point is 00:08:53 for real, I can try to do something like looking up and answer online that could make the task solve faster than if I did it in a legitimate way. So we try to stop this during training and evaluation by, in many cases, disabling internet. About two-thirds of the way through the presentation, Dalton enumerated a list of steps, Open AI plans to take to show how the company is responding, giving the series of events he and Wallace had laid out. This is a pivotal moment both for our company as well as the AI industry as a whole, Dalton said. Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques, both in our fundamentals and better use of AI. We're consciously slowing down research in order to enhance security
Starting point is 00:09:35 and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation. At the conclusion of the talk, Wallace and Dalton took time to repeatedly emphasize Open AI's concerns about the broader implications of the incident, namely that the episode provides an example of completely autonomous AI-driven hacking that was accidental in this case, but in all likelihood will be used with intent by malicious actors in the near future. The important takeaway here that has really shifted dramatically is that fully automated offensive loops require investment in truly fully automated defense,
Starting point is 00:10:15 and we are not there as an industry. Dalton said, we will have to find that path together with urgency, end quote. Rockstar says it will show an extended look at Grand Theft Auto 6 on August 27th, premiering on Netflix at 3 p.m. Eastern time that day before streaming on YouTube at 9 p.m. Eastern time that same day. Quoting the verge, we'll be getting an in-depth preview of Grand Theft Auto 6 very soon. Rockstar announced this morning that it will be airing an extended look at the game on August 27th.
Starting point is 00:10:51 Netflix and Rockstar previously partnered to release the original GTA trilogy on mobile. There are no other details just yet. The anticipation and fandom around Grand Theft Auto 6 is unprecedented and we're honored that Rockstar Games has partnered with us to debut the next part of the Grand Theft Auto Story with Netflix members first, Brandon Rieg, Netflix's VP of Nonfiction Series said in a statement. Of course, the new GTA has been on the way for quite some time and is launching well over a decade after GTA 5, which went on to become one of the best-selling games of all time, launching across three console generations and PC as of May.
Starting point is 00:11:25 A, GTA-5 has sold nearly 230 million copies, according to Rockstar Games' parent company Take 2. Finally, we got to come back to yesterday's big headlines. According to semaphore, Demis Hasabas has been drifting away from Google DeepMind for quite a while now. Quote, Hasabas wasn't pushed out against his will, said people involved with the matter. Rather, he struggled to get satisfaction out of being in the role of a tech executive rather than a visionary scientist. Hesabas who won a Nobel Prize in 2024 has recently been his most animated when talking about isomorphic labs. Google's biotech spin out that he runs. His passions lie in using AI to solve scientific puzzles like curing diseases and discovering new materials and in ensuring that AI
Starting point is 00:12:16 doesn't accidentally cause catastrophic harm to humanity. But the move, which coincided with the departure of chief scientist Jeff Dean, sent Google's stock price down 4% on Wednesday and raised doubts about the company's standing in the fast-paced AI race. Its AI models are roughly six months behind the frontier on coding ability, where most of the compute power is currently being consumed. The feeling at Google, according to one executive, is that the management change will help accelerate the development of AI rather than hold it back.
Starting point is 00:12:45 While Hasabas had become the face of the company's AI efforts, he wasn't focused on the part of the company most associated with its standing in the race, end quote. And sources at CNBC give us more color from down in the Google trenches. Depending on where you sit, Google either has the most enviable position in artificial intelligence
Starting point is 00:13:04 or is bleeding top talent to leading AI labs and other startups on the front line of innovation. Alphabet CEO Sundar Pachai said on last month's earnings call that 90% of Fortune 100 companies are using Gemini Enterprise underscoring the company's ability
Starting point is 00:13:18 to sell AI services to cloud customers. Tamaz Tungu's founder of Theory Ventures said it's becoming clear that top-of-the-line models aren't required when it comes to meeting most enterprise demand, though. I think we are at that place with AI, particularly for a lot of white-collar work, where many of the models that are reasonable are good enough,
Starting point is 00:13:38 Tungu's said. The next evolution of models are likely to be helpful in domains where you have really fancy computers. While the tone on Wall Street has been generally favorable to Google, not everyone is celebrating inside of the company. Some researchers have grown frustrated over access to the computing capacity they need to pursue ambitious projects while watching Google Cloud sell TPUs to outside customers, including Anthropic, according to people familiar with the matter who asked not to be named due to confidentiality.
Starting point is 00:14:08 Tensor processing units or TPUs are the company's homegrown AI chips that compete with NVIDIA's graphics processing units. Google's bureaucracy is a common source of frustration with layers of approval required to move research into products that can make emerging companies like OpenAI, Anthropic, or even younger startups more appealing, especially for AI researchers and developers who prefer lab work to balance sheets. Dean is leaving alongside Google stars like Sanjay Gamowat, Oriel Vignales, and Kwakli to start Discovery Loop. On X, Dean said the startup backed by Google will be a public benefit corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. Their exit follows the departures of other prominent researchers, including Nome Shazir,
Starting point is 00:14:54 one of the authors of the landmark 2017 paper, Attention is All You Need, which provided the foundation for generative AI. All eight authors of that paper have now left Google. One of the biggest points of friction inside Google is apparently compute. Google is investing more than almost any company in the world in data centers, chips, and related infrastructure, but capacity remains scarce. Every TPU assigned to training a model serving a Google product or fulfilling a contract with a cloud customer reflects a choice among competing priorities.
Starting point is 00:15:25 Frustrations over access to compute can be especially acute when Google announces large infrastructure commitments to competing labs like Anthropic whose models compete directly with Gemini, sources with knowledge of the matter said. One of the people said, Google has projections for demand in different areas, including research and model training, serving products such as Search and Gemini and working with cloud customers. Those requirements are modeled years in advance, the person said, though capacity may shift over shorter periods if a product grows faster than expected or if priorities change.
Starting point is 00:15:56 Dan Niles, founder of Niles Investment Management and a Google shareholder, said access to compute is a natural source of tension. Google has all of these other businesses and they've got to figure out who they're going to give some of these resources to, Nile said. Somebody's always going to be unhappy in that situation, end quote. Nothing more for you today. Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.