Tech Brew Ride Home - DO You Understand Sandboxing Or No?

Episode Date: August 3, 2026

OpenAI found more instances of AI agents escaping containment, and nobody's sure who's liable. Astra solved ten open math problems, Google pulled a Google Earth image tool over deepfake fears, and AI ...chip counts were set to 10x by 2028. Sources: OpenAI has discovered other instances where AI agents escaped containment; none of the agents were thought to have left OpenAI's network (Reuters) Anthropic's breaches dated back to April but went undiscovered until last week; former UK cyber chief Ciaran Martin calls the lapses "sloppy" as experts warn of national security risks (Bloomberg) Legal experts say US courts haven't settled who's liable when an AI agent goes rogue, with agency law, tort law, and the CFAA's intent requirements all an awkward fit (Wired) OpenAI Says Astra Solved 10 Open Math Problems With Lean Proofs (Implicator AI) Google rolls back an image generation tool in Google Earth to add "stronger guardrails" after concerns arose it can be used to create deepfake satellite imagery (NPR) A look at the deluge of AI computing power set to come online in the coming years; Epoch AI expects the number of AI chips in use to double every nine months (NYT) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices

Transcript
Discussion (0)
Starting point is 00:00:00 The 2006 Chevrolet Equinox awarded the most dependable compact SUV in the U.S. by J.D. Power is designed for your everyday. And with available all-wheel drive, you can handle your to-do list with total confidence. Start your build at Chevrolet.ca. Details at J.D.Power.com. Welcome to the TechBrew right home for Monday, August 3rd, 2026. I'm Brian McCullough today. Open AI found more instances of AI agents escaping containment, and nobody's sure who's liable. An unreleased model has solved 10 legendary math problems. Google pooled a Google Earth image tool over deep fake fears and AI chip counts are set to 10x by 2028. Here's what you miss today in the world of tech.
Starting point is 00:00:47 Sources say that OpenAI has discovered yet more instances where AI agents escaped containment, which leads me to ask again like I attempted to ask with Friday's show title, but somehow screwed up and pasted in the show notes instead. Sorry about that. I wanted to ask, are you sure you understand what the word sandboxed actually means? Quoting Reuters, One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network. The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by Open AI was launched shortly before its primary rival Anthropic disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported.
Starting point is 00:01:43 We have a whole industry where the people designing, developing, and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe, said Maurice Jodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk. Jodo said his concerns were heightened by indications that neither Open AI nor Anthropic were watching the agents as they went rogue. Reuters has previously reported that Open AI realized its agent had broken into hugging face only after the company contained the hack, contacted the FBI, and went public about the intrusion. Open AI has said the Reuters account contain inaccuracies,
Starting point is 00:02:19 but has not responded when asked what they were. In its Thursday statement disclosing how its own agents hacked victims online, Anthropics suggested that it had not been watching them in real time, saying that real-time monitoring of the evaluation logs would have helped to surface the problem sooner. Giotto said that pointed to a lack of proper scrutiny. It seems like they weren't even looking, Chodo said, end quote. More on people increasingly being like, hey guys, are you sure you're doing the right thing here? Quoting Bloomberg. The breaches occurred as far back as April, but weren't discovered until last week when Anthropic audited its cybersecurity testing after open AI disclosed that its own AI agents had escaped a testing environment and infiltrated
Starting point is 00:02:58 hugging face, a repository for open source AI models and documentation. A lot of people from a cybersecurity perspective will see that as sloppy, said Kieran Martin, the former head of the UK's cybersecurity center. If a mainstream cybersecurity company made similar mistakes, it could face lawsuits and potential regulatory action, according to Martin. Cybersecurity firms routinely test potentially dangerous tools in controlled sandboxes, a virtual and isolated software environment, meant to run security tests or analyze unsafe code and are expected to ensure those safeguards work, he said. The incidents have also raised concerns about the risks that autonomous AI systems could pose to national security. Gregory Allen, a former director of strategy and
Starting point is 00:03:39 Policy at the Department of Defense Joint Artificial Intelligence Center said the U.S. military should use advanced AI models to protect its systems while recognizing that the technology creates a new category of risk. Anthropic found these hacks because they started looking for them, Alan said. We actually have no idea how widespread autonomous AI hacking is at this moment in time. U.S. organizations may not have reliable access to AI tools capable of defending against autonomous cyber attacks said Daniel Remler, a former State Department AI policy official now at the Center for New American Security. Hugging Face had to use ZAI's Chinese open weight model for forensic analysis and patching after it couldn't rely on an American model.
Starting point is 00:04:18 Remler said the other open models with comparable coding capabilities are also Chinese, including Deepseek V4 and Kimmy K-3. Remler said the episode should prompt companies and the government to expand access to AI-powered cyber defense. He used. warned that increasingly capable Chinese systems could soon autonomously attack U.S. organizations, making it necessary to develop countermeasures now. This episode should crystallize that we are going to have a Chinese mythos by the end of the year or first quarter next year, he said, referring to an anthropic model that the company said was so powerful it couldn't release it widely. And we are going to run into a situation where these agents are
Starting point is 00:04:55 autonomously able to hack U.S. entities like Hugging Face, and we are not really thinking about their defense, end quote. Also, I have a... hadn't thought of this angle. Quoting Wired, as more and more incidents emerge, questions about legal liability and repercussions have also come to the fore. Researchers and lawyers Wired spoke to emphasize that these questions have not been answered in practice in the United States legal system. In other words, there haven't been decisions in enough relevant cases for the picture to start to form. But the recent high-profile incidents from Open AI and Anthropics suggest that answers will need to come soon. Just because you're using an AI agent or AI model, that shouldn't
Starting point is 00:05:32 somehow absolve you of any liability, but it's going to depend a lot on the facts in the particular situations as cases begin to be decided in courts, said Lauren Yu, a fellow with the ACLU's speech privacy and technology project. Experts say that so-called agency law could be relevant given that the doctrine focuses on situations where a principal has given an agent, permission and authority to act on their behalf. To be clear, the agents in this area of law have always been human. Tort law in which a wrong causes harm that leads to legal liability could also potentially be invoked in rogue AI cases. Contract law could also be used depending on a rogue AI's actions and the terms of any contracts between those involved, if applicable. And hacking laws,
Starting point is 00:06:17 like the Computer Fraud and Abuse Act or state-level legislation, could also be relevant. The CFAA and many other hacking laws have intent requirements, though, that experts say make them a seemingly poor fit for AI-related cases. Ultimately, experts emphasize that questions about U.S. federal AI liability law will be answered only through more litigation. Perhaps most concerning to critics is that AI agents are goal-oriented but lack a human moral or ethical compass. The law firm Brownstein-Huyat-Farmer Shrek wrote in an alert to clients on July 24th. In some situations, an agent may infer actions that were never explicitly authorized if those actions appear necessary to achieve its objective, end quote.
Starting point is 00:06:56 Speaking earlier this week about Open AI's hugging face disclosures, Alex Zenla, chief technology officer of the cloud security firm Adera, mused, quote, This is just the one that we know about, but God knows what's happened with the stuff that we don't know about, end quote. Although at the same time, Open AI really wants you to know this. Quoting Implicator AI, Open AI says an internal version of Astra, it's. its next big model produced results for 10 problems in math, quantum complexity, and theoretical computer science that it had never been solved before. OpenAI estimated the token cost for finding all 10 solutions at roughly $2,000 at sole API rates. Thomas Bloom, a University of Manchester mathematician who runs Eratos Problems.com, called the results big news in an ex post cited by the
Starting point is 00:07:52 decoder, adding that as mathematical constructions. They were bigger than the unit distance. counter-example, Open AI had disclosed earlier. Humans used the same model to prepare manuscripts and then had it formalized each argument in the Lean Proof Assistant. The lab also released model reasoning walkthroughs, a 249-page manuscript dated August 26 and machine-checkable proof files. The materials give reviewers three separate artifacts to inspect while Astra remains private. The release extends a test. Open AI disclosed on May 20th when an unreleased model generated a disproof of the AirDose unit distance conjecture. In the post, OpenAIA also cited a separate program providing 100,000 scientists and mathematicians free access to its best chat GPT models, while it continues to evaluate
Starting point is 00:08:37 private systems on open research problems. Lean is a proof assistant. A mathematical argument is translated into formal statements, and the software checks whether each step follows from the definitions and rules encoded in the system. That can catch missing steps or invalid deductions that ordinary pros may conceal. Open AI's public repository created, on August 1st uses Lean 4.32.0 with the Math Lib Library and Lake Build system. Its ReadMe gives two commands for downloading cash dependencies and building all certificates. The Apache 2.0 license allows others to inspect and run those files without access to Astra. A reviewer can clone the project, fetch the cash dependencies with Lake EXC Cash Get and Run Lake Build All to check the certificates against that software stack. The process tests the public. The public. proof files on a local machine. It does not provide access to Aster or show how the private model behaves on problems outside this release. A successful Lean build validates the statement as formalized inside lien. Journal peer review remains a separate process. Outside researchers cannot
Starting point is 00:09:43 test the private model that generated the arguments or determined how mathematicians will rank the importance of each result. Opening I disclosed Astra's role and an estimated token cost in the post, but the release does not give outside researchers access to the model, its training data, public testing interface. The public materials allow direct checking of the lien certificates while Astra's broader research behavior remains inaccessible. Open AI has not announced a public release date for Astra. Noam Brown, an open AI researcher, supplied his own boundary for the claims. The decoder reported on August 1st that Brown said the lab had not spent much on each problem and that there were, quote, no Millennium Prize problems yet, end quote. The 2026 Chevrolet tracks is the stylish SUV
Starting point is 00:10:32 for those on the move. And with the standard Chevy safety assist package, you have the backup to handle every turn with confidence. The 2026 tracks. Start your build at Chevrolet.ca. Hear that? It's your money calling. It wants a promotion. Elevate your savings with the Scotia high-interest savings account.
Starting point is 00:10:56 Always earn high regular interest rates that grow the more you save and invest. Conditions apply. Visit Scotiabank.com slash H-I-SA to learn. learn more. Scotia Bank, you're richer than you think. Google has rolled back an image generation tool in Google Earth to add what it is calling stronger guardrails after concerns arose that the tool could be used to create deep fake satellite imagery, quoting NPR. In its initial announcement of the new feature on Thursday, the company said it was designed to work with users, quote, imaginations. Just zoom in to a place
Starting point is 00:11:32 in Google Earth on the web, tap create image, and type whatever you want. to see Brian Horowitz, a product manager wrote. But for experts and analysts who rely on Google's imagery to verify breaking news and atrocities in hard-to-reach parts of the world, the potential to create deep-fake satellite imagery at the click of a button was horrifying. I tried refugees at the Mexican border, a nuclear plant in Iran, a crash in Amsterdam, a hospital with a bomb crater in Gaza, nothing was refused. Hank Von S. An open-source researcher who first wrote about the potential damage the tool could cause, told NPR via text. NPR was able to easily generate images of Iran's Karg Island on fire and a flooded U.S. Capitol
Starting point is 00:12:13 complex. Both events which have not happened would constitute major news if they were real. Online journalists and open source investigators wondered aloud about Google's decision. Very curious how or if this idea was red-teamed internally because the opportunities for abuse and disinfo are literally boundless. Evan Hill, a visual forensics investigator at the Washington Post wrote on X. Google Earth imagery is not always the most up-to-date. but its sweeping high-resolution coverage of the planet can provide important reference imagery for both images from the ground and more current images from other satellites. Satellite imagery has been kind of a safe bet when it comes to verifying an event because it's hard to fake,
Starting point is 00:12:50 says Jake Godden, a senior researcher with the online group of Bellingcat, which does visual forensic investigations. Satellite images are usually controlled by companies or institutions, and their source, a camera hundreds of miles above the Earth's surface, has historically been hard to imitate, end quote. Finally today, this is one where you should probably click through to the story link to see the chart associated with this story, because for all the talk of being compute constrained right now, and for all the talk of data center buildouts and hundreds of billions of dollars being spent, I didn't really appreciate until this piece just exactly how much compute is still about to come online. Like, look at the chart and see where we are now.
Starting point is 00:13:36 We're not even in the middle innings yet, quoting the times. From the American Midwest to the Persian Gulf, hundreds of major data centers now under construction will be turned on in the coming years. They are set to deliver an avalanche of computing power to develop and run AI that has no equal in the history of the technology industry, with breakthroughs that once felt revolutionary, likely to become increasingly routine. Behind each leap in AI are corresponding jumps in computing power. Today, there are about 20 million AI chips crammed into the data centers that underpin the technology's growing abilities and usage worldwide, according to the, research firm Epic AI. That figure is expected to double roughly every nine months, putting the world on pace to have about 200 million of the chips by the end of 2028, 10 times current levels. In size and ambition, this moment compares to the building of the railroads in the 1800s,
Starting point is 00:14:27 President Franklin D. Roosevelt's New Deal in the 1930s and the Manhattan Project to create an atomic weapon in the 1940s, technologists said, this is the largest scale infrastructure buildout in the history of humanity, said Rob. Wachin, a co-founder of the microchip firm Etch, which has raised more than $1 billion to meet the growing demand for AI components. Peter DeSantis, who leads foundational AI models at Amazon, which provides computing power to the AI firms, Anthropic, Open AI, and others, said, the Seattle company has doubled its computing capacity since 2022 and would double it again by next year. It's hard to get your mind around the scale, he said.
Starting point is 00:15:04 Fueling the surge is the belief that AI can take on more human responsibilities and solve increasingly complicated tasks with the more data and computing power you feed it. This tenet, sometimes called the scaling laws, has become the driving force behind this technology era. Those with the most computing power will create the most advanced AI systems, capturing the biggest share of profit and value tech leaders argue. The biggest engine, they say, will win the race. Confidence in the scaling laws has led AI leaders to make ever-boulder predictions. Dario Amo Dai, the chief executive of Anthropic, has said that if these laws hold for another year or two, AI will be able to perform huge amounts of white-collar work. Demis Hasabis, head of Google's
Starting point is 00:15:43 AI Lab DeepMind, wrote recently that AI could usher in 10x of the Industrial Revolution at 10x the speed. Economists and investors have raised concerns that tech firms are spending faster than they can profit from AI. Past infrastructure booms have been followed by downturns before the benefits of the technology were realized. The railroad boom in the 1800s, electrification in the 1920s and the dot-com bubble in the late 1990s were punctuated by economic recessions and a stock market crash as companies that overspent went out of business. Each time you've had a technological revolution, this kind of bubble bursting happened, said Philippe Agacombe, who won the Nobel in Economic Science in 2025 for research on innovation-driven economic growth. AI is like the fourth
Starting point is 00:16:26 industrial revolution, and it has this aspect to it that generates a bubble. With more computing power coming online, geopolitical divisions are also set to widen. The United States, home to about 5,500 data centers, about 10 times the next closest country, is far ahead of the rest of the world, including China. U.S. companies like Amazon, Google, Microsoft, and META control about 80% of global computing power that drives AI, according to Epic AI. Google alone is believed to have four times as many AI chips as all of China's companies, which are racing to catch up by developing new semiconductors and AI infrastructure of their own, end quote. Coming at you slightly early today because we are in the midst of moving houses.
Starting point is 00:17:15 Talk to you tomorrow.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.