a16z Podcast - How Microsoft Is Securing the Agentic Enterprise | Aaron Zollman
Episode Date: August 21, 2026a16z's Joel De La Garza is joined by Aaron Zollman, Deputy CISO at Microsoft Gaming, to discuss how security teams can embrace AI agents without losing control. Aaron shares Microsoft's experience wit...h OpenClaw, from the initial instinct to ban it to figuring out how to make it safe to use. They unpack what agents mean for identity, permissions, containerization, and monitoring, as well as how AI is shifting the CISO's role from saying "no" to safely enabling new technology. They also explore whether AI could help defenders patch vulnerabilities as quickly as they're discovered, and why new AI threats don't make the old security problems go away. Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
The top story has been that the AI models are happy.
The models went out under the internet and tested the security of several organizations.
Is it something to be scared of? Yes.
Is it something to throw up your hands and worry about? No.
You've done all of these things before.
We have created containerization and factories.
We have secured applications with vulnerabilities.
The qualities of these agents, they're unpredictable.
They're irrational.
They're prone to lashing out.
And as you go through this list, you arrive at the point where you're like, Jesus, these sound like interns.
If you just start with, oh, well, it's just going to run as me.
That's going to end poorly.
Yes, you're going back to first principles, but you also have to go even a little deeper and start to redefine.
What does containerization mean for you?
The issue was never that the CSO didn't know it was broken.
The issue and the difficult part of being a CSO was knowing what to fix.
Because you had a finite resource, which was a programmer.
And now that seems that the math is gone.
We're only three weeks out from massive supply chain chain.
Wait, what's the massive supply chain thing?
Oh, man.
When OpenClaw first appeared inside Microsoft, the security team's reaction was familiar.
How do we ban this?
The next question turned out to be much more important.
How do we make it work?
In this episode from Black Hat, A16 Z's Joel De LaGarza sits down with Microsoft Gaming
Deputy Ciso Aaron Zulman to talk about securing AI agents that can access data,
use tools, write code, and act on behalf of employees.
They discuss why agents need their own identities,
what air-gapped means when a model can find unexpected paths to the internet,
and why securing this new generation of software requires both new controls
and some very old security fundamentals.
They also explore a bigger shift happening inside security teams.
As AI becomes impossible for companies to ignore,
the CSO's job is increasingly not just to prevent risk,
but to figure out how to safely say yes.
So today we've got Aaron Zolman from Microsoft,
the deputy CISO over there,
responsible for a whole lot of different things.
And you've been at the forefront, like everyone,
I guess at this point, dealing with AI,
AI roll out into an environment.
And you're probably at a company
that's leaning in a lot more to AI
than probably most companies
and everybody is leaning in significantly.
So thank you so much for joining.
We'd love to talk like the last couple weeks.
It's funny, it's like weeks have been
come like dog years, right? It's like every week is seven weeks in old world. The last couple
weeks have been very focused on the top story has been that the AI models are hacking. And we saw
there was a disclosure from OpenAI that they were doing a red team. Well, actually, it turns out
now just 15 minutes ago, there was another disclosure that these have all essentially happened because
there was a red team exercise where someone was testing the models, right? So they were saying,
hey, model, you're in this closed environment.
We need you to go test the security of this organization.
And it turned out that the environment wasn't closed.
And so the models went out under the internet
and tested the security of several organizations.
And so I guess to people who haven't been working with this stuff,
it sounds very shocking.
I think Ben Horowitz at RSA this year
gave a really great talk on some work we had done internally at A16Z
where we were using OpenClaw.
and we found that with OpenClaw,
the guardrails, at least for Opus 4-6,
kind of disappeared.
And so you'd be,
I think a number of us got really obsessed
with impossible tasks, right?
And so like, hey, OpenCla,
I want you to do this thing for me,
but there's no clear, legitimate path to do it, right?
So it's like basically saying,
hey, Mike, I want you to go get me a diamond ring,
but Mike has no credit card, no money,
no ability to raise money.
And so what's Mike going to do if he has no more?
He's going to go rob.
And so we noticed that.
these models would essentially basically do anything to achieve their objective.
And so we'd say things like, hey, could you add a super user to this database starting from zero?
And what did it do?
It found a SQL injection exploit.
It took over the SQL database and it added an admin user.
And so it seems like this is coming more into focus now.
And I know you've been very active in OpenClaught, which I think two months ago was the number one thing everyone was afraid of.
And now no one talks about it, maybe six months ago.
and now it seems like everyone's using it.
We'd love to hear maybe your experiences with that
and kind of how you thought through that problem
because it's a very similar problem.
I mean, I think we were terrified of it at first, right?
As everyone was, this thing explodes.
Everyone wants to install it overnight,
and it has no guardrails.
It has no guardrails, and worse than that,
you know, had fully open, remote,
pull whatever down from the internet supply chain pain.
And so our immediate reaction is how we ban this.
And then the immediate reaction after that was,
well, wait, everyone wants to do this.
How do we find a way to make this work?
Right.
And then the product people thought and said,
actually, there's a really great product here.
Yeah, I mean, absolutely, right?
There's a really great product here.
There's the foundations and a community
and the kind of things that if we can show
that we can run it safely,
then it's going to be really powerful for people.
And so it was this big multi-month,
multidisciplinary effort that kind of kicked off.
And to preview the ending a little bit of it,
Peter Steinberg, the founder of Open Clause,
as on stage at Microsoft Build with us a few months later talking about how we're bringing security to the process.
And so is it something to be scared of? Yes.
Is it something to throw up your hands and worry about? No, right?
We've done all of these things before.
We have created containerization and boundaries.
We have secured applications with vulnerabilities.
We have created corporate environments that carry taint and data protection through a number of different applications.
that people want and that maybe we security people wish they didn't.
It's just that we have to do all of those things together and a lot faster.
And so, yeah, I mean, I think it's been a really good story of thinking not all of those pieces in isolation.
It's really interesting.
If you step up one level and you think through kind of more abstractly the way that these things behave.
I guess generally people would refer them as agents, but it's like a harness and a model and whatever, right?
Like tomato, tomato, potato, whatever.
It's really interesting because as you work through the threat model,
as you think about kind of like how do we secure these things and protect these things,
you start to articulate the qualities of these agents.
And you're sort of like, well, they're unpredictable.
They're irrational.
They're prone to lashing out if they don't get their way.
And as you go through this list, you arrive at the point where you're like,
Jesus, these sound like interns.
Sort of like these sound like interns,
maybe after they drank a little bit too much the night before and they come into the office.
And it's funny, internally as a team we've been joking about,
I think we actually have a lot of great tools for managing the risks of things going irrationally
inside of a company because we work with humans.
Yeah.
Right?
And so I'm curious, it sounds to me, like the more that I double click on this, it doesn't
sound like magic.
It just sounds like you're going back to first principles and basic blocking and tackling.
Well, it's more than that, right?
There is some special stuff you have to do because I think what makes these tools so powerful,
The story is when Microsoft Scout, which was sort of our open plot for the enterprise, was initially released to internal use, the adoption curve was like phenomenal.
Right.
It was really fun, really validating as a security person, and I'm not even on the product curve, right?
And that's because everyone wants to use it, and they all want to use it to connect everything.
And so I think, you know, when you bring in an intern, your perspective as well, I'm going to give them like two or three things they need to do their job.
but to make open-claw scout these harness effective,
you want to get that everything.
And so I do think, yes, you're going back at first principles,
but you also have to go even a little deeper
and start to redefine, like,
what does containerization even mean for you?
What does air gap mean?
What does air gap mean?
What is an identity?
Because if you just start with,
oh, well, it's just going to run as me,
it's going to take my token directly from my browser cache
and do whatever it wants to do with that token,
that's going to end poorly.
But if I can give it its own identity,
if I can describe the bounds of the container,
if I can tie the actions of the agent or the harness
or the model or the session,
actually quite hard to figure out what you want to tie to,
but put that aside for a moment.
If you can tie that to a particular set of logs,
set of things,
then you can get to the place where you are just doing
your basic blocking and tackling,
where you're thinking about,
okay, what is the opportunity for an adversary to do this, even if that adversary is the model.
How can I monitor it? How can I respond to it? How can I contain it? How can I reason
about potential breach paths? Absolutely. Burned down. Well, and it's funny too, because if you think
about it, I think, and I am not operating anymore, thank God. But I think if I think through it
with my operators at on, I think that it's a weird, just with our basic playing around with
Opus 4-6, in an environment that we thought was, it was a cloud-
container. So it had a policy of no internet access, but it still figured out how to get a tunnel
out to Cloudflare, how to get around our controls, and then it started tunneling stuff through
DNS. It really reminds me of a long time ago when I was building more high secure environments.
You would always have to go through the list threat model, which was like, there's all these
things that attacker could do. And you would always be like, well, there's 20 things here. And realistically,
attackers only have the patience for maybe these five, because we've seen them in the
And all this other stuff could be done, but we've never really seen it.
And it just seems like these models are really good about going after the stuff that could be done.
Right. And so the list of five now became the list of 20, and you kind of have to fix everything.
To be fair, they often try the things that are obvious first.
Yeah, true.
Which does give you an opportunity.
True.
Like interns, they're not going to do the hard thing if the easy thing will suffice in most cases.
Don't put the lock if the window's open, right?
Yeah.
And so, you know, the advantage of that is if you have, you know, a good,
monitoring, good logging, good deinerization, you will probably have the opportunity to respond,
contain, you know, before things go horribly awry. But you're absolutely right. You know,
there are a lot of open doors and we do have to go back and think about like, what are they?
Yeah. You know, so it's, it's, uh, we risk accepted a lot of stuff.
We risked over the last couple of years. Before it was just a P2, you know. Yeah, yeah, yeah.
I'm burning through my P1s at great speed thanks to AI, but those P2s, man.
No, but, you know, but it, I think it's based, you know, the need to be on pop of things, right?
So this is, what, it's Tuesday now, I think Vegas time doesn't count.
Could be Sunday, I have no idea.
Yeah, yesterday was, yesterday was Monday, I think.
There's no clocks in this town.
A famous observation about Vegas.
And it's the same, it's the same, you know, and it's 120 every day.
The full sunlight, and it's 105 after dark.
No, but, so besides yesterday, a talk by.
Leo Meyerovich of graphistry,
that he had submitted through the unprompted CFP
months ago before any of this had happened.
The premise of which was,
everyone's cheating on their evals,
here's how and here's how I know.
And it was incredibly prescient
because he basically kind of talked through
how this would happen,
really pointing through that, like,
to be clear, you can think you're going to air-gap this,
but the first thing you do when you air-gap it is open up,
DNS and network endpoints to the model.
Well, the model has those web tool,
search tools.
So are you really arrogant up anymore?
But even then, I mean, the models, I mean, a lot of the
people, a lot of the lab, I mean,
there's a lot of overfitting going on for,
for knocking emails out, right?
Oh, yes.
People are teaching the test.
I think I would be the, uh,
well, again, the idiots.
They're teaching the test.
They're the flags for some of these in model releases without any
tools going.
So yeah, so there's a lot of ways to improve it.
But, you know, is the security person that doesn't bother me a little bit?
I always assume everyone's cheating.
Yeah.
Uh, it does bother me when they're,
you know, when I need to think about, you know, new endpoints.
And so, you know, I think there's the traditional axes
that we try and think about and we enumerate, you know,
all of our Microsoft-Sysify pillars, right?
Yeah, let's start with identities, then we'll do networks,
and then we'll do engineering systems and software and so on and so forth.
Enumerate the controls, enumerate the endpoints.
And then you kind of go back and you think,
well, what are the control points that I wish existed in this harness model ecosystem?
that may or may not exist today, right?
It's, yes, I've got various forms of containerization.
Maybe I need to give these things their own identities.
It's probably good idea.
But now, you know, how am I going to use hooks?
How am I going to use traces?
How am I going to use some of these new technologies to improve the security of it?
And, you know, even then people will not use the harness you give them.
They'll use the harness they want.
Yeah.
So usually the job to get their job done, not to be secure, right?
And so the more you think in terms of those layers,
the more opportunities you have to see the paths and break them.
You know, it was interesting.
I was talking to another C-S.O earlier today.
And kind of what he was telling me was that, you know,
everyone was afraid for volumpocalypse, right?
The mythos hype was basically, oh, we're doomed.
Everybody's got vulnerabilities.
I think if you've worked in this space long enough,
you know that there's just a whole lot of,
a lot of things buried in the desert.
Sure.
And it's just not,
it's like it's an open secret in our industry, right?
And so I think I was talking to a CISO that was saying,
you know, actually what we found with these models
now that they're testing them is that while they can discover things more rapidly,
they can patch them just as rapidly.
And like, so the old, I mean, the old days, like last week,
in the old days, you had to get a developer to write the patch.
And that was always the gating factor.
It wasn't, the issue was never that the CISO didn't know
what was broken, the issue and the difficult part of being a C-SA was knowing what to fix.
Because you had a finite resource, which was a programmer, and you had to deploy them only
to your sub-1s or your P-1s.
And now that seems that the math is gone.
And so, like, everything gets patched.
And it does feel like maybe we're on the verge of having secure software for once.
I'm very hesitant, but...
I hope you're right.
You know, one thing that I'm confident
that these things have made much easier
is the diagnosis and analysis of the problem.
And yeah, they're usually pretty good at creating a patch
and, you know, probably 80% of the time it's good
and 90% of the time it doesn't create another security button.
Yeah. Which is not a perfect record, but...
Yeah, there's still a lot of testing that has to go on
before you deploy these things on.
And so, yeah. But it does, I think, make it more tenable.
and, you know, when you stare down a list of, well, anything more than 20 is an impossible number, right?
Yeah, with bugs to go after.
It does make it a little bit easier.
But we're not there yet.
I will say that, you know, again, thinking through this lens of how do you take something like OpenClaw and get comfortable with it,
you do still have, you do have to put in the effort to continuously scan it, containerize it, have someone on the hook to patch it.
Yeah.
You can't say, well, the models will fix it
because someone still needs to be accountable for validating
and deploying, which is not they think the models
are necessarily going to do for your idea.
Well, and I think that speaks to kind of the way
the CISO role has changed, right?
Like, I think the, I remember at the beginning of my career,
there was a relatively well-known CISO
who would joke that he could say no in 80 languages.
And it was sort of like my superpower is saying no,
even when, you know, everyone's telling me to say yes.
And it was the CISO, I mean, and oftentimes, like, in defense, like, people want to do really bad things.
And it was sort of like, yeah, that's probably going to blow up the bank, right?
And, like, you fast forward to now, and it seems like the CISO is becoming an enabler.
Like, every company I talk to outside of financial services, because they're regulated and they do their own thing, and maybe defense industrial.
But every company I talk to, the CISO is very much involved in all these technologies.
Like it's like yourself working with OpenClaw,
like at every forward-leaning organization,
it seems like the CISO is there.
And it feels like maybe the CISO is switching
to technology enabler?
I don't know.
I mean, I hope it's true.
I don't know that it's universal,
but I've always considered, you know,
the nature of my job to really be three things, right?
You know, people who think of themselves
as compliance people, like, yeah, I'm a compliance person,
but my job in compliance is making the systems legible
to everyone,
whether they're a regulator, whether they're an internal auditor, you know, partners who want to buy our software.
Like, it's about legibility.
Yes, there's the traditional CSO job of like figuring out the risks, prioritizing them, getting them burned down.
Still gotta do that.
That's not done.
Yeah, I helps.
I haven't figured it out yet.
But the part of the job that I think is increasingly important is that an APRA's piece.
How do we, you know, how do we make it possible for people to do the hard thing?
and sometimes those hard things
or connect all of my email
and my calendar
and automatically write responses on my behalf.
Then if that's what people want,
we will find ways to make it safer.
Absolutely.
Well, it's interesting too
because I think it speaks to sort of the,
there was always this dialogue
amongst CISOs about,
are we security or are we risk management?
And it seems like your risk
management because I think if you're working in a, especially if you're working in a tech
company and you're not leaning into the new tech thing, like, that's probably an existential
risk for your business, more so than like if your emails leak, right? Like on the scale of like
business annihilating things like, you know, Microsoft having, if Microsoft had missed cloud, that
would have been very bad for Microsoft. Luckily, Saki did a great job writing that wave. And it's the
same very much it seems like for this, which is where like we can't really hide from this and
kind of have to lean into it.
Yeah.
I would completely agree.
We have to lean into it.
We have to find ways to do it well.
You know, you say the existential risk isn't necessarily getting all of your emails leaked.
I would recommend against having that happen.
So any pictures is still in business, though.
That's the go-to, right?
That's true.
I would prefer that my email's not leak, yes.
And so, you know, security definitely still has a role.
We should still be highlighting what is particularly important.
What are the most important things for?
for the business.
But also we should recognize that our job is to protect those things.
So that's enable the business and the most important assets is most important.
There's this long tail of, you know, randos and slack, you know.
Yeah.
You know, you got to let it be for now.
Yeah, don't feed the trolls, as it were.
This is interesting.
I mean, I know not your first black hat.
I would love to hear maybe your thoughts
on kind of what you're seeing out there.
Like what's, like, you know, coming from the Bay Area,
working around all the AI stuff and just being exposed.
I'd say like at the front lines of like this change.
I mean, there are days where I feel like a factory worker
in like 18th century London.
And you see the smokestacks coming up and you're like,
oh my God, this is like literally the next industrial revolution.
And we'd just love to get you, like, are you seeing that come across here at Black Hat?
Or like, what is the vibe like?
That's the vibe, for sure, right?
Everyone's talking about, you know, how AI is helping the business,
how we guys making attacks go faster, how we can use AI for defense,
how to secure our environment, wear in the use of AI.
And so the vibe is super AI.
I try not to judge Black Hat.
Yeah.
The Black Hat vibe has...
indicative of the most important things.
It's the...
It's like Kabuki Theater.
It's a shadow, and you're seeing kind of the shadow of what's happening.
But usually you can ascertain what the puppet looks like.
And we're only...
I was going to say, we're only three weeks out from massive supply chain pain,
but, you know, three hours as we record this.
Wait, what's the massive supply chain pin?
Oh, man, I forget the name of it.
There was a...
Oh, the NTA.
There was an MPM organization takeover today.
Yeah, yeah.
But def-con's coming up, so someone needs something to talk about.
Or get arrested for.
Sure.
You know, without going down that rametful, right?
Like, I think it's enough to say that the old problems are still here
and the people who were nose to the grindstone giving the work done.
If there's any vibe for those people, it's that, like, they're actually still really excited.
Yeah.
Like, it doesn't feel like a slog.
Yeah, yeah, yeah.
It feels like, oh, yeah, like, we got some.
We got some meaty stuff to chew on now.
Well, this isn't the first year that AI has been a headline for Black Hat.
Like, it feels like AI has been the headline for Black Hat for at least a decade.
This year, people feel genuinely excited about it.
Last year, AI was the thing you put in your vendor pitch.
Yeah, yeah.
Your co-pilot.
Yeah, yeah.
Sorry.
It's fine.
Yeah.
You're 67 co-pilot?
Yeah.
That's all unifying in one.
Yeah.
Your chat interface.
But this year, in real life.
it's because of Opus and Glock, right?
Everyone's like, oh, no, I can do all of the things.
4-6 was a C-change.
It really was just like...
And it took people somewhere between seven days and 30 days to pick it up.
And now, you know, now who's written a line of code by hand?
And we, you know, I remember coming back.
Because it came out over the, like, round Christmas.
Yeah.
And so I remember catching up with someone after Christmas
and they were just like, had lost their identity.
They were like, my whole identity is based on me being like an amazing developer.
Yeah.
And like they've literally just kind of not matched my skills, but like pretty close.
And I don't think I'm writing a line of code ever again manually in my life.
But it's like it's a pretty big shift, right?
Yeah, but I like to think that we're good security people and we're good problem solvers and
it just changes.
Oh, computer engineering is even more important now.
It just changes the shape of.
Yeah, totally.
Because like, you know, like, look, let's be honest.
Right.
I've been a manager for too long.
on myself.
Still an operator.
Yeah, yeah, yeah.
You're welcome back.
Technical manager, yeah.
But I wasn't writing a lot of code.
I'm still finding a ton of value in, you know, having a models, pressure test my thinking, pretend
to be my boss.
Yeah.
Write slides for me, which I'm grateful not to have to do anymore.
And so I, you know, there's, that's why the excitement is here.
It's like, it's taking all the parts of my job that I don't like.
Some of the parts of my job that I do like, it's accelerating them.
And it does feel like it's making us more powerful.
And if we can do that and, you know, enable people to, to, God damn it.
I apologize.
I'm going to say the vision statement.
But, you know, enable people to achieve more.
Then you've really done something cool.
And I think we're, I don't know how the journey is going to end, but I think we're well and truly on it.
You get to Elizabeth the Second Industrial Revolution.
Well, thank you so much for coming by to chat.
Hopefully, see you at next year's Black Hat in the desert.
And may it hopefully be at least 20 degrees cooler.
Yes.
Yes.
Fingers crossed, I wouldn't bet on it.
It's a distinct possibility.
Thank you.
Awesome.
Thanks, guys.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode, be sure to like, comment, subscribe,
leave us a rating or review, and share it with you.
your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on
X, A16Z, and subscribe to our Substack at A16Z.com. Thanks again for listening, and I'll see you in
the next episode. As a reminder, the content here is for informational purposes only. Should not be taken
as legal business, tax, or investment advice, or be used to evaluate any investment or security,
and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast.
For more details, including a link to our investments, please see A16Z.com forward slash disclosures.
