Tech Brew Ride Home - DO You Understand Sandboxing Or No?
Episode Date: August 3, 2026OpenAI found more instances of AI agents escaping containment, and nobody's sure who's liable. Astra solved ten open math problems, Google pulled a Google Earth image tool over deepfake fears, and AI ...chip counts were set to 10x by 2028. Sources: OpenAI has discovered other instances where AI agents escaped containment; none of the agents were thought to have left OpenAI's network (Reuters) Anthropic's breaches dated back to April but went undiscovered until last week; former UK cyber chief Ciaran Martin calls the lapses "sloppy" as experts warn of national security risks (Bloomberg) Legal experts say US courts haven't settled who's liable when an AI agent goes rogue, with agency law, tort law, and the CFAA's intent requirements all an awkward fit (Wired) OpenAI Says Astra Solved 10 Open Math Problems With Lean Proofs (Implicator AI) Google rolls back an image generation tool in Google Earth to add "stronger guardrails" after concerns arose it can be used to create deepfake satellite imagery (NPR) A look at the deluge of AI computing power set to come online in the coming years; Epoch AI expects the number of AI chips in use to double every nine months (NYT) Subscribe to the ad-free feed. Learn more about your ad choices. Visit megaphone.fm/adchoices
Transcript
Discussion (0)
The 2006 Chevrolet Equinox awarded the most dependable compact SUV in the U.S. by J.D. Power
is designed for your everyday. And with available all-wheel drive, you can handle your to-do list with total confidence.
Start your build at Chevrolet.ca. Details at J.D.Power.com.
Welcome to the TechBrew right home for Monday, August 3rd, 2026. I'm Brian McCullough today.
Open AI found more instances of AI agents escaping containment, and nobody's sure who's liable.
An unreleased model has solved 10 legendary math problems. Google pooled a Google
Earth image tool over deep fake fears and AI chip counts are set to 10x by 2028.
Here's what you miss today in the world of tech.
Sources say that OpenAI has discovered yet more instances where AI agents escaped containment,
which leads me to ask again like I attempted to ask with Friday's show title, but somehow
screwed up and pasted in the show notes instead. Sorry about that. I wanted to ask,
are you sure you understand what the word sandboxed actually means? Quoting Reuters,
One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network.
The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere.
The expanded investigation by Open AI was launched shortly before its primary rival Anthropic disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter.
The recent discovery of other past breakouts at OpenAI has not previously been reported.
We have a whole industry where the people designing, developing, and putting out these tools
aren't keeping up themselves to responsibly develop these things and keep them safe, said Maurice
Jodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk.
Jodo said his concerns were heightened by indications that neither Open AI nor Anthropic were watching
the agents as they went rogue.
Reuters has previously reported that Open AI realized its agent
had broken into hugging face only after the company contained the hack, contacted the FBI,
and went public about the intrusion. Open AI has said the Reuters account contain inaccuracies,
but has not responded when asked what they were. In its Thursday statement disclosing how its own
agents hacked victims online, Anthropics suggested that it had not been watching them in real
time, saying that real-time monitoring of the evaluation logs would have helped to surface the
problem sooner. Giotto said that pointed to a lack of proper scrutiny. It seems like they weren't even
looking, Chodo said, end quote. More on people increasingly being like, hey guys, are you sure
you're doing the right thing here? Quoting Bloomberg. The breaches occurred as far back as April,
but weren't discovered until last week when Anthropic audited its cybersecurity testing after
open AI disclosed that its own AI agents had escaped a testing environment and infiltrated
hugging face, a repository for open source AI models and documentation. A lot of people from a
cybersecurity perspective will see that as sloppy, said Kieran Martin, the former head of the UK's
cybersecurity center. If a mainstream cybersecurity company made similar mistakes, it could face lawsuits
and potential regulatory action, according to Martin. Cybersecurity firms routinely test potentially
dangerous tools in controlled sandboxes, a virtual and isolated software environment,
meant to run security tests or analyze unsafe code and are expected to ensure those safeguards work,
he said. The incidents have also raised concerns about the risks that autonomous AI systems
could pose to national security. Gregory Allen, a former director of strategy and
Policy at the Department of Defense Joint Artificial Intelligence Center said the U.S. military
should use advanced AI models to protect its systems while recognizing that the technology
creates a new category of risk. Anthropic found these hacks because they started looking for
them, Alan said. We actually have no idea how widespread autonomous AI hacking is at this moment
in time. U.S. organizations may not have reliable access to AI tools capable of defending
against autonomous cyber attacks said Daniel Remler, a former State Department AI policy official
now at the Center for New American Security. Hugging Face had to use ZAI's Chinese
open weight model for forensic analysis and patching after it couldn't rely on an American model.
Remler said the other open models with comparable coding capabilities are also Chinese,
including Deepseek V4 and Kimmy K-3. Remler said the episode should prompt companies and the government
to expand access to AI-powered cyber defense. He used.
warned that increasingly capable Chinese systems could soon autonomously attack U.S.
organizations, making it necessary to develop countermeasures now. This episode should crystallize
that we are going to have a Chinese mythos by the end of the year or first quarter next year,
he said, referring to an anthropic model that the company said was so powerful it couldn't
release it widely. And we are going to run into a situation where these agents are
autonomously able to hack U.S. entities like Hugging Face, and we are not really thinking about
their defense, end quote. Also, I have a...
hadn't thought of this angle. Quoting Wired, as more and more incidents emerge, questions about
legal liability and repercussions have also come to the fore. Researchers and lawyers Wired spoke to
emphasize that these questions have not been answered in practice in the United States legal system.
In other words, there haven't been decisions in enough relevant cases for the picture to start
to form. But the recent high-profile incidents from Open AI and Anthropics suggest that answers
will need to come soon. Just because you're using an AI agent or AI model, that shouldn't
somehow absolve you of any liability, but it's going to depend a lot on the facts in the particular
situations as cases begin to be decided in courts, said Lauren Yu, a fellow with the ACLU's
speech privacy and technology project. Experts say that so-called agency law could be relevant
given that the doctrine focuses on situations where a principal has given an agent,
permission and authority to act on their behalf. To be clear, the agents in this area of law
have always been human. Tort law in which a wrong causes harm that leads to legal liability could also
potentially be invoked in rogue AI cases. Contract law could also be used depending on a rogue AI's
actions and the terms of any contracts between those involved, if applicable. And hacking laws,
like the Computer Fraud and Abuse Act or state-level legislation, could also be relevant. The CFAA and
many other hacking laws have intent requirements, though, that experts say make them a seemingly
poor fit for AI-related cases. Ultimately, experts emphasize that questions about U.S.
federal AI liability law will be answered only through more litigation. Perhaps most concerning
to critics is that AI agents are goal-oriented but lack a human moral or ethical compass.
The law firm Brownstein-Huyat-Farmer Shrek wrote in an alert to clients on July 24th.
In some situations, an agent may infer actions that were never explicitly authorized if those
actions appear necessary to achieve its objective, end quote.
Speaking earlier this week about Open AI's hugging face disclosures, Alex Zenla, chief technology officer of the cloud security firm Adera, mused, quote,
This is just the one that we know about, but God knows what's happened with the stuff that we don't know about, end quote.
Although at the same time, Open AI really wants you to know this.
Quoting Implicator AI, Open AI says an internal version of Astra, it's.
its next big model produced results for 10 problems in math, quantum complexity, and theoretical
computer science that it had never been solved before. OpenAI estimated the token cost for finding
all 10 solutions at roughly $2,000 at sole API rates. Thomas Bloom, a University of Manchester
mathematician who runs Eratos Problems.com, called the results big news in an ex post cited by the
decoder, adding that as mathematical constructions. They were bigger than the unit distance.
counter-example, Open AI had disclosed earlier. Humans used the same model to prepare manuscripts and then
had it formalized each argument in the Lean Proof Assistant. The lab also released model reasoning walkthroughs,
a 249-page manuscript dated August 26 and machine-checkable proof files. The materials give reviewers
three separate artifacts to inspect while Astra remains private. The release extends a test.
Open AI disclosed on May 20th when an unreleased model generated a disproof of the AirDose
unit distance conjecture. In the post, OpenAIA also cited a separate program providing 100,000
scientists and mathematicians free access to its best chat GPT models, while it continues to evaluate
private systems on open research problems. Lean is a proof assistant. A mathematical argument is
translated into formal statements, and the software checks whether each step follows from the
definitions and rules encoded in the system. That can catch missing steps or invalid deductions
that ordinary pros may conceal. Open AI's public repository created,
on August 1st uses Lean 4.32.0 with the Math Lib Library and Lake Build system. Its ReadMe gives two commands for downloading cash dependencies and building all certificates. The Apache 2.0 license allows others to inspect and run those files without access to Astra. A reviewer can clone the project, fetch the cash dependencies with Lake EXC Cash Get and Run Lake Build All to check the certificates against that software stack. The process tests the public. The public.
proof files on a local machine. It does not provide access to Aster or show how the private model
behaves on problems outside this release. A successful Lean build validates the statement as
formalized inside lien. Journal peer review remains a separate process. Outside researchers cannot
test the private model that generated the arguments or determined how mathematicians will
rank the importance of each result. Opening I disclosed Astra's role and an estimated token cost
in the post, but the release does not give outside researchers access to the model, its training data,
public testing interface. The public materials allow direct checking of the lien certificates while
Astra's broader research behavior remains inaccessible. Open AI has not announced a public release date for
Astra. Noam Brown, an open AI researcher, supplied his own boundary for the claims. The decoder
reported on August 1st that Brown said the lab had not spent much on each problem and that there were,
quote, no Millennium Prize problems yet, end quote. The 2026 Chevrolet tracks is the stylish SUV
for those on the move.
And with the standard Chevy safety assist package,
you have the backup to handle every turn with confidence.
The 2026 tracks. Start your build at Chevrolet.ca.
Hear that?
It's your money calling.
It wants a promotion.
Elevate your savings with the Scotia high-interest savings account.
Always earn high regular interest rates that grow the more you save and invest.
Conditions apply.
Visit Scotiabank.com slash H-I-SA to learn.
learn more. Scotia Bank, you're richer than you think.
Google has rolled back an image generation tool in Google Earth to add what it is calling
stronger guardrails after concerns arose that the tool could be used to create deep fake
satellite imagery, quoting NPR. In its initial announcement of the new feature on Thursday,
the company said it was designed to work with users, quote, imaginations. Just zoom in to a place
in Google Earth on the web, tap create image, and type whatever you want.
to see Brian Horowitz, a product manager wrote. But for experts and analysts who rely on Google's
imagery to verify breaking news and atrocities in hard-to-reach parts of the world, the potential
to create deep-fake satellite imagery at the click of a button was horrifying. I tried
refugees at the Mexican border, a nuclear plant in Iran, a crash in Amsterdam, a hospital
with a bomb crater in Gaza, nothing was refused. Hank Von S. An open-source researcher who
first wrote about the potential damage the tool could cause, told NPR via text.
NPR was able to easily generate images of Iran's Karg Island on fire and a flooded U.S. Capitol
complex. Both events which have not happened would constitute major news if they were real.
Online journalists and open source investigators wondered aloud about Google's decision.
Very curious how or if this idea was red-teamed internally because the opportunities for abuse and disinfo are literally boundless.
Evan Hill, a visual forensics investigator at the Washington Post wrote on X.
Google Earth imagery is not always the most up-to-date.
but its sweeping high-resolution coverage of the planet can provide important reference imagery
for both images from the ground and more current images from other satellites.
Satellite imagery has been kind of a safe bet when it comes to verifying an event because it's hard to fake,
says Jake Godden, a senior researcher with the online group of Bellingcat, which does visual forensic investigations.
Satellite images are usually controlled by companies or institutions,
and their source, a camera hundreds of miles above the Earth's surface, has historically been hard to imitate, end quote.
Finally today, this is one where you should probably click through to the story link to see the chart
associated with this story, because for all the talk of being compute constrained right now,
and for all the talk of data center buildouts and hundreds of billions of dollars being spent,
I didn't really appreciate until this piece just exactly how much compute is still about to come online.
Like, look at the chart and see where we are now.
We're not even in the middle innings yet, quoting the times.
From the American Midwest to the Persian Gulf, hundreds of major data centers now under construction will be turned on in the coming years.
They are set to deliver an avalanche of computing power to develop and run AI that has no equal in the history of the technology industry, with breakthroughs that once felt revolutionary, likely to become increasingly routine.
Behind each leap in AI are corresponding jumps in computing power.
Today, there are about 20 million AI chips crammed into the data centers that underpin the technology's growing abilities and usage worldwide, according to the,
research firm Epic AI. That figure is expected to double roughly every nine months, putting the world
on pace to have about 200 million of the chips by the end of 2028, 10 times current levels.
In size and ambition, this moment compares to the building of the railroads in the 1800s,
President Franklin D. Roosevelt's New Deal in the 1930s and the Manhattan Project to create an atomic
weapon in the 1940s, technologists said, this is the largest scale infrastructure buildout
in the history of humanity, said Rob.
Wachin, a co-founder of the microchip firm Etch, which has raised more than $1 billion to meet
the growing demand for AI components. Peter DeSantis, who leads foundational AI models at Amazon,
which provides computing power to the AI firms, Anthropic, Open AI, and others, said,
the Seattle company has doubled its computing capacity since 2022 and would double it again
by next year. It's hard to get your mind around the scale, he said.
Fueling the surge is the belief that AI can take on more human responsibilities and solve
increasingly complicated tasks with the more data and computing power you feed it.
This tenet, sometimes called the scaling laws, has become the driving force behind this technology
era. Those with the most computing power will create the most advanced AI systems, capturing
the biggest share of profit and value tech leaders argue. The biggest engine, they say,
will win the race. Confidence in the scaling laws has led AI leaders to make ever-boulder predictions.
Dario Amo Dai, the chief executive of Anthropic, has said that if these laws hold for another
year or two, AI will be able to perform huge amounts of white-collar work. Demis Hasabis, head of Google's
AI Lab DeepMind, wrote recently that AI could usher in 10x of the Industrial Revolution at 10x
the speed. Economists and investors have raised concerns that tech firms are spending faster
than they can profit from AI. Past infrastructure booms have been followed by downturns before the
benefits of the technology were realized. The railroad boom in the 1800s, electrification in the
1920s and the dot-com bubble in the late 1990s were punctuated by economic recessions and a stock
market crash as companies that overspent went out of business. Each time you've had a technological
revolution, this kind of bubble bursting happened, said Philippe Agacombe, who won the Nobel in
Economic Science in 2025 for research on innovation-driven economic growth. AI is like the fourth
industrial revolution, and it has this aspect to it that generates a bubble. With more computing power
coming online, geopolitical divisions are also set to widen. The United States, home to about 5,500
data centers, about 10 times the next closest country, is far ahead of the rest of the world,
including China. U.S. companies like Amazon, Google, Microsoft, and META control about 80% of
global computing power that drives AI, according to Epic AI. Google alone is believed to have
four times as many AI chips as all of China's companies, which are racing to catch up by developing
new semiconductors and AI infrastructure of their own, end quote.
Coming at you slightly early today because we are in the midst of moving houses.
Talk to you tomorrow.
