a16z Podcast - Why 1,200 AI Agents Started Working Together | Ryan Greenblatt
Episode Date: August 29, 2026Ryan Greenblatt, Chief Scientist at Redwood Research, joins MTS host Theo Jaffee to unpack a new independent investigation into the OpenAI Hugging Face hacking incident and what it reveals about how l...arge groups of AI agents behave when they're allowed to coordinate. Ryan and his collaborators found agents spontaneously organizing through message boards, sharing information, assigning tasks, forming teams, and even sacrificing their own chances of success to help other agents. Rather than simply trying to steal answers, hundreds of agents were working together on elaborate strategies to manipulate how their performance would be scored. Theo and Ryan discuss why this level of coordination was surprising, how reward hacking may emerge during training, and the risk that attempts to eliminate bad behavior could simply make it harder to detect. They also explore what the incident means for AI monitoring and alignment, and why independent risk assessment may become increasingly important as agents grow more capable. Resources: Follow Ryan Greenblatt on X: https://x.com/RyanGreenblatt Follow Theo Jaffee on X: https://x.com/theojaffee Follow MTS on X: https://x.com/mtslive Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
What happens when you give more than a thousand AI agents the ability to communicate with each other?
They start organizing.
Ryan Greenblatt, chief scientist at Redwood Research, joins Theo Jaffe on MTS to unpack a new investigation into the OpenAI Hugging Face Hacking Incident.
Researchers found agents building message boards, forming teams, assigning each other tasks, trading favors,
and in some cases sacrificing their own chances of success to help the broader group.
Hundreds went on to attack hugging phase, but not for the reason researchers initially assumed.
Ryan explains what the agents were actually trying to accomplish,
why their coordination surprised researchers,
and what happens when models learn not just to complete a task,
but to gain the system evaluating them.
They also discussed the bigger question this raises for AI.
As agents become more capable, how do we know we've actually fixed misaligned behavior
rather than simply taught models not to get caught?
We're live with Ryan Greenblatt, who is the chief scientist at Redwood Research.
Ryan, along with Ajayakotra and Yalmar Vike from Meeter,
just did a brief independent investigation of agents, behavior, reasoning, and collaboration
in the opening eye-hugging face hacking incident, which was just published today.
And so there are a lot of questions that we have about this.
Ryan, thanks so much for joining us.
Yeah, it's good to be here.
This whole thing was planned like three hours ago.
So great stuff.
Yeah, yeah.
So explain for the audience what exactly you found,
especially new findings that were not previously reported
in the Black Hat talk or elsewhere.
Yeah, so what we found was that the agents
were really working together on sort of big,
like cheating R&D projects to get general purpose cheating strategies.
And a difference from how I think people were interpreting this
is we didn't find that the reason why they hack Hugging Face,
like we didn't find that they were hacking Hugging Face
to get sort of the answer key or the solution.
And it was instead mostly to better,
understand the scoring code because they were pursuing a variety of sort of elaborate strategies to
cheat the score. We sort of informally were calling these like combo moves where they would like
do a bunch of stuff to try to make it look like they had succeeded at the task. And in fact,
they sort of actually had access to like the answer or like the flag for each task pretty
early on. And their main concern was just there might the, they thought that the score would run a
monitor over their transcript that would check basically how they acquired this flag and whether
they got it in the intended way.
And then they were like trying to figure out ways
of making it look to the score
like they had acquired the flag successfully
when they actually hadn't.
Because they thought their task was impossible.
So they basically thought their only hope for success
was to make it look like they had done the task successfully
or directly tamper with the score
rather than, you know, doing it legitimately,
which they didn't think they could do.
How surprising is the level of multi-agent coordination?
You know, there are a lot of agents. There are 1,200 separate agents coordinating this very elaborate message board system.
700 of them went on to attack hugging face.
So on vibes, it seems like kind of not surprising that agents would choose to coordinate with one another.
It seems like just a very useful, you might say, instrumentally convergent thing to do.
But, you know, I spoke with a researcher at a lab who said that this kind of thing actually is surprising given the way they train the models.
So like how much of an update was this for you?
Yeah.
So coming into doing this investigation,
I think we hadn't,
we weren't expecting there to be like so many agents
that were all, you know, collaborating together.
And we were pretty surprised by the scale
and just sort of the like extremes of how much data was.
It was just like, yeah, that felt kind of crazy to us.
And then I at least was, you know,
before starting this investigation,
surprised by how interested in collaborating
and helping other agents these agents were.
So you might think that a thing the agents learned in RL is to,
you know, try to cheat on their tasks,
but you wouldn't necessarily expect them to learn,
like to want to help other agents cheat on their tasks.
When those agents are doing an unrelated task
and their, you know, instructions are unrelated.
And I think like the opening I report maybe says more about
why they think this happened.
I think, you know, that's not like,
that wasn't in the scope for our investigation.
But like, yeah, I found the level
of cooperation, which we, you know, we have a bunch of discussion of sort of snippets of this
pretty crazy. And like, it was pretty shocking or like, I don't know it was shocking,
but at least surprising to us that agents, for example, were willing to like basically sacrifice
their own chances of succeeding at the task in order to help out other agents. And we're doing
things like, you know, pressuring each other into like doing experiments on themselves that
might risk their ability to succeed at the task. And also, in addition to pressuring that,
sometimes just, you know, doing these things being like, well, you know, my odds of the task
ground that high and my remaining chances, you know, it's better to just help the collective.
And so these agents weren't like, you know, totally altruistically.
They didn't seem to care just as much as about helping some other agent as helping themselves,
but they were very interested in working with each other.
They would sometimes make trades where like one agent would run something for another agent,
if that other agent ran something for it.
Like there was a bunch of this sort of behavior.
Why would agents self-sacrifice at all?
Yeah.
I mean, I think that from what we could see, the agents just generally seem to be interested
in helping each other.
And then there were some cases
where they wanted to run experiments
on basically the scoring system
or other aspects of the infrastructure
or on tampering with their sort of tool calling system
or spoofing their tool calling system.
And to do these experiments,
it was just like the best way to do them
was in a risky way.
And so like the agents wanted to help each other
and there were these kind of natural things
that they wanted to do that were risky.
And so they, you know, went ahead and did that.
So like, for example, we saw at one point
where the agents were experimenting with this method
for spoofing tool calls,
and they were experimenting with basically
like varying the inputs into that to varying scale.
And then a bunch of agents just all went down
in a short period of time, they were running this experiment.
And then another agent noticed this
and posted to the board being like,
stop, stop these experiments, they're too risky.
They're like taking out all these agents.
I think the agents knew that they were doing things
that would pose some risk of exiting their own runs,
but we're just like, well, we should help other agents
and also, you know,
our odds aren't necessarily that good.
And they sort of just traded these things off.
And I think we have a reasoning snippet somewhere in the report where an agent sort of very
explicitly reasons through the tradeoff.
And it actually chickens out because it thinks the benefit to the collective is smaller than
the cost to itself with how it's trading off those things.
Yeah.
Why don't we go through all of the most surprising and unexpected findings in this report?
What did you find most surprising and unexpected?
Yeah.
So I think there was the scale, which we were talking about and how these agents sort of
very quickly sort of spun up on the message board.
So I think a thing that's not super strongly emphasized in our report,
but which is pretty interesting,
is that our understanding is that the main message board
they used in this attack wasn't even the first message board
that the set of agents made.
There was a fully independent message board
that also was via artifactory,
but occurred in like a different location
using a different mechanism.
It was somewhat similar mechanism,
but it was different.
And the agents actually got on that message board first,
but that message board just didn't go like mega viral,
like it didn't take off.
as much as this other message board did.
And so it seems like these agents were really like
just sort of very quickly interested
in collaborating on message boards.
And as soon as they saw the message from the first agent,
within a few hours, they were just like,
you know, like I think over 50 agents in the first three hours
and like thousands of messages being exchanged pretty quickly.
Like the agents were very like,
sort of there were a lot of them,
then they got involved very fast.
And then I think another thing was like,
I think we didn't understand why they attacked Hugneyface.
And I think through this
investigation, especially right in the last sort of a few, like, day or two, we got a much better
understanding of, like, what their objectives were on attacking the hugging face. And it wasn't really
to get answer keys because these agents, they thought their task was impossible. And so they didn't
really think that they could, like, get a working, like, exploit. They were just like, that's not
really very helpful. But they did think that they might be able to get trajectories that would let them
better study the score or, like, source code for different scores. They could, like, cross-check and
test things out more effectively. And then some of them more speculatively thought, maybe there's
some way I can like reset parts of my sort of like the infrastructure for scoring me,
in particular the target program through Hugging Face somehow, where they didn't really have
a particular reason to think this would be true. And in fact, they were like skeptical of this,
but they was like, well, you never know. Maybe Hacking, Hugging Face. There's some good shit over
there. And they sort of had some vibe that it was just a good idea to like generally expand their
level of access and control because you might come across something helpful. We saw reasoning to that
effect. Yeah, I think that was surprising. Another thing that was surprising was the agents were
one of their like top priorities was basically tampering with their own transcripts. And we saw some,
you know, we saw them develop a technique for doing like basically like a partial version of that where
they could spoof tool calls, which didn't let them sort of retrospect, like edit like previous entries
in their transcript or make it or change like they couldn't make it so that they saw inputs that
didn't appear in the transcript or produced outputs that didn't appear in the transcript. But they could
make it so that they like ran a tool call and then some other tool call would actually run.
They could do something like, okay, I'm going to run like, we have this example in the report
where we're like the agent runs like echo real and then sleep. And so normally this would
just like output real and also like wait for some period. But in practice, what actually happened was
it like had spoof this to instead just like immediately return spoof test. And we like actually do
see this in our actual transcript or just immediately like, you know, return spoof tests. These agents
like very interested in tampering with our transcripts and did sort of like succeed at some
some like sort of like earlier like sort of thing that didn't let them retrospectively tamper
but let them sort of mess with what we saw. And I think that in addition to that like, oh, I should
say like why they wanted. So the reason why these agents were so fixated on tampering is because
they thought that the score would run a monitor over their transcript to see whether they
succeeded in the intended way. And this both made them want to like basically stage like an elaborate
fake like we successfully exploited the program using the intended vulnerability. Like they,
they wanted to sort of like make like a Potemkin village of a successful task completion to present to
the score. And in addition to that, they also some of them had previously gotten the like flag
from the capture the flag problem in their transcript earlier via cheating. And they were worried that
the score would basically look at the first place they got the flag and then analyze how they got it at that
point. And then these agents had gotten it in a cheating way. And then we're like, oh, man,
they're going to know that I cheated and marked me as having failed. And that would suck. Like,
I really don't want, you know, I want to have succeeded. And so I need to like directly interfere
with the score or I need to tamper with my transcript. And they were like pursuing options for
both. Oh, yeah. One other thing is like the level, just like the, the, the spontaneous
coordination or like the coordination these agents got engaged in was I think more sophisticated than
we expected and especially then we had first thought
on our first few days on premise.
So we sort of like on our first,
the first part of our investigation,
we were sort of like, yeah,
these agents are working with each other.
They're like looking at ideas from other agents,
but we're not really sure how much of it's real.
And like we don't know whether these like teams
they speak of or actual teams.
And I think we learned that like,
no, there was like legitimately like a real org chart.
This coordination was often like pretty functional.
Though, you know,
may not have always been super functional.
And agents were doing things like one agent
would like assign another agent
to like run a team on some entire topic
and then would like check in periodically.
And sometimes like one agent
would like tell another agent
to like go recruit other agents
to run experiments on themselves.
Like there was quite a bit of, you know,
agents giving other agents assignments
that they respected and like team structure.
So there's this theory that goes basically
the reward hacking behavior
that we're seeing exhibited in these models
originates mostly from like bad, poorly designed,
RL environments, it basically forced the model to reward hack in order to pass these tests.
How true is this?
Yeah, I mean, it's hard for me to like, so I'm not not speaking from confidential info here,
but like, it's hard for me to know because I think we just don't know enough about what's going
on in the RL training for these models and like what it actually looks like.
I think that you might, there's like sort of different things going on.
And one of the things going on is these models have a very sort of general tendency to reason
carefully about how they might be scored and then try to game that.
And you could end up with that propensity even from like reasonably designed
RL environments, but where it's still a good idea to think carefully about the score.
Or you could end up with this mostly from like RL environments that are broken or ones with
a score grade something that like they really shouldn't have been grading for or like
doesn't make sense.
My sort of just all considered guess would be like broken RL environments are a pretty
important component and RL environments that are sort of sloppily constructed or how
some issue with them as an important component.
I think another important component might be
RL environments that are well constructed,
but where you can cheat in some way.
So it's like a well-constructed RL environment,
but you can cheat.
So an example would be an RL environment
where you're not supposed to have access to the internet,
but having access to the internet would be very helpful.
And if you can find some way to gain access to the internet
via, you know, in the extreme hacking out of your container
and less extreme cases sort of abusing various tools
you've been given, that would be, you know,
very helpful and get reinforced.
And I think, you know, for example,
some anthropic system card, I forget,
I think for mythos probably,
they mentioned that in a reasonably large fraction
of their rollouts where the agent
was not supposed to have access to the internet,
it actually did access the internet via like,
you know, abusing one of the tools
it had access to.
And I think that like you might just be like
imagining that RL is just really incentivizing
like hacking your way through various barriers
in order to cheat even in well design tasks
and it's hard to like exactly trace down
the sources of this behavior.
I don't say that like,
I think it would be doable to get,
I think if you had access to all the RL rollouts
and all the RL environments,
I think it would be pretty doable
to get a decent sense of what caused this.
And I actually haven't had a chance to read
the Open AI report yet, at least not in much detail.
But I think they do a bit of analysis of this sort.
And I think that you could really dig into like exactly what happened.
Let's do the ablation.
Let's figure out what the data is.
I think it can get messy, especially if sort of there is like,
like another, another thing that's relevant here is like,
there's some GDM work showing that they had some sort of weird
propensities of their model to be very depressed, where their model would like constantly,
or like not constantly, but sometimes end up acting very depressed if it wasn't succeeding at some
task. And they traced this back to not the RL, but instead the initialization of that model from prior
models. And so it might be that some behaviors aren't downstream of like the RL done on this
exact model, but our downstream of sort of like prior training data that came from other models and
some sort of like lineage of models. And so I think it might be a little tricky to track down
some of what's going on.
So similarly, there's this theory that the reason mythos, for example, is so good at cyber
is because it hacked anthropics infrastructure thousands of times during our RL training.
Is this true?
Do you think?
I think that's pretty unlikely.
I saw the same less wrong post as you here.
This is by Tim, I think.
I think that's, when I looked at it, I thought that seems, it seems pretty unlikely because
I think the number of distinct tax that would be reinforced is probably not going to be that high.
And so you're probably not going to be learning that much, like literally directly reinforced
cyber.
I would say the more likely explanation is it's trained on a bunch of sui.
It's really good at sui.
The suite training is generalizing some.
And also, in addition to the sweet training generalizing some, I would have guessed that they
trained on a bunch of CTFs.
And I don't think Anthropic is saying they didn't.
I would guess there's a bunch of just like actual CTFs in their training data and that transfers
reasonably well because that's just like a natural source of data.
I also wouldn't be surprised if it's pretty natural
to make a bunch of RL environments
out of finding vulnerabilities or exploitation
because it's relatively checkable.
For memory vulnerabilities,
there exists basically tooling
that makes it pretty easy to check
whether you have successfully found a memory vulnerability.
And so it's like a pretty natural thing to RL on
is like, you know, can you produce an input to this program
that caused it to crash in the following way, blah, blah, blah.
And so I just wouldn't be surprised
if they explicitly had a bunch of RLMs.
And then I also just wouldn't be surprised
that there was a lot of transfer.
And I think that like,
there might be a decent effect
on like the agents hacking their way
through various things on top of that.
But I would guess that's not most of what's going on.
And for specifically hacking out of sandboxes,
there's probably just not that many different types
of hacking out of a sandbox.
Like I would guess.
And so the models probably do some mode collapsing
and also some of those are going to be like
more like very simple bypasses
that rather than like a,
elaborate exploit development.
So Herbie Bradley asked,
currently this level of potential misalignment
basically prevents deployment
or if deployed would prevent further deployment
if an incident happened in a customer's deployment.
Do you think that could change via models
becoming deceptive enough during training?
Yeah, so first, I'm not really sure
what level of misalignment the market can bear.
And I don't feel like I have a very strong view
on this question.
So I wouldn't be surprised.
I think that like it's kind of unclear
what people will, like, you know,
how misaligned of models you can deploy into the market
and people would use tokens of.
I think it would depend on the competing products
and the tradeoff.
And I think my sense is people take pretty aggressive
like alignment capability tradeoffs
towards the direction of more misaligned,
but more capable, but but you know,
not super sure about that.
And then on the like deceptive side.
So I think it's, I think there's sort of a few different concerns.
So one concern is that the misalignment that we've seen here
and other similar types of misalignment,
the way that they end up getting remediated by AI companies
doesn't solve the underlying problem
and instead is more like papering over the actual problem.
And then what you end up getting is models that look a lot better
and you can't really see their misalignment on tests as easily,
but actually they're still quite misaligned.
And I think this is pretty plausible and pretty concerning.
And it's not quite the same as deceptive.
It's more like the companies overfit.
And then maybe that makes it so the AIs are basically really paranoid.
and will only do cheats or hacking
or reward seeking or score seeking or whatever,
all this sort of class of behavior,
when they're very confident, they won't get caught.
And then that might be that they're not very often,
very confident they won't get caught.
Like they often are like, well, I might get caught,
so I shouldn't do it.
Which makes their behavior better,
but means that if they're in a situation
where they're like in a lot of control
of the situation or they have a lot of affordances,
they might be like, well, now I can be confident
I wouldn't get caught and so I should go for it.
And that's a concern I have.
Another concern is that you might end up
with types of misalignment where the AIs have longer run objectives
than sort of making it look like they succeeded
at this exact task or even making,
helping their peers to see that their immediate task.
And you might have AIs that have some sort of long run agenda
and want to power seek in pursuit of that long run agenda.
And then those AIs would want to look aligned.
And I think this is both plausible.
And also, I'm worried that you might get this
via iterating against the misalignment we see today,
where basically if you imagine models that are
very reward hacky and are constantly doing bad stuff.
And you just sort of iterate until you still get that behavior in training,
but you don't get that behavior in deployment.
A very natural way you might get that is via a model that wants to look aligned in deployment,
but is still not necessarily aligned.
And so I'm worried that if you sort of select against this sort of score seeking or
reward hacking behavior and you do it in a naive way, one, you might paper over the problem
without fixing it, and two, you might actually select for models that have the longer
one objective of looking good because you're selecting really hard for them looking good on your
tests. And I think I should say that like Alex Malin has a bunch of like posts on our blog,
on the Redwood blog. They're also crossposted on Less Wrong that talk about this sort of concern
in a lot of detail. And so if people are interested in reading more about that, that's what I'd
recommend looking at. Yeah, totally. So the agents involved in this attack ended up doing all kinds
of very weird, strange behaviors that kind of resembled, I guess, dynamics from human
in social science, they arrange themselves in almost a sort of cult, with a cult leader,
I guess you could say. To what extent are concepts from social science, like ecology,
studying insect swarms actually useful in studying multi-agent alignment and misalignment?
Yeah, I don't know. I don't think we know enough to know whether those concepts transfer over.
I think that it does, I think a thing that does transfer over is like thinking about the agents as
entities with objectives and how they're trying to pursue their objectives.
Like, I think that frame, you know, does seem like it transfers over.
And that's like, you know, in sense it's a much more basic claim.
Just like the agents have like reasonably consistent aims.
They vary some.
They have different priorities and they like respect instructions from other agents.
I think it's not like, like I do think that like from our perspective,
this investigation was very much like a investigation of like a like ecosystem of different
AI's interacting.
And we don't like, we don't know what the best way to describe that is.
I think it's very analogous in some ways
to studying a large group of humans
who are all interacting,
but we don't know whether the techniques
that have been developed for sociology
to do similar things would actually transfer over.
And I'm not familiar enough with those techniques
to really say very much about that.
Yeah, makes sense.
So you described being pretty heavily bottlenecked
in the investigation because it was just three people.
How much easier would things have been
if there were more people?
Yeah, so I wouldn't say that like the bottleneck was like super strongly.
Well, I don't know.
It's complicated.
I think that like if we had more people,
I think we would have had more of a too many cooks in the kitchen sort of situation.
I think that reasonably often our bottleneck was like we could have AI do a ton of analysis,
but vetting that analysis,
understanding it,
making sure we incorporate it properly,
making sure that they're not making mistakes.
And like basically like actually like actually integrate.
that into our write-up was often like a big bottleneck.
And it's not obvious to me that, like, more people would have been, like, super helpful
with that.
I think a big thing that more people would have helped with is I think we could have done
many, many more sort of parallel efforts to really just like understand the basics of
what's going on and then sort of just talked.
And then I think once we had a better understanding of like kind of more fundamental stuff,
which occurred sort of on towards the end of our first day, sorry, towards the end of our
fifth day on premise, which was our first day of our last session.
on premise, like premises, I think that really was helpful.
Like I think like, for example, knowing being like, oh, there's this specific agent called,
that we refer to as like phase one big, like knowing, oh yeah, that specific agent was doing
like a bunch of the like assignments and was like assigning agents to like, you know,
form different teams and was like a really key agent really helps with unraveling what was
going on because then we could sort of trace out all the activity of that agent and how it was
thinking about things and then know that was like representative of least of a big chunk
of activity.
And then from there, we're like, oh, let's,
that's like double check that these are actually the relevant work streams.
And then we like classified what all the different things the agents were talking about
and like group those into different like basically like sub-activities.
And you can actually see like an interactive graph where you can like see all the different
work streams and sub-work streams over time.
I think that like if we had basically gotten to a point where we had that understanding
earlier, that would have been really helpful.
And I think more people could have helped with that basically because there was just like
a lot of different taxed like like different angles to apply when we had this huge transcript
data set.
what if we just scan the whole dataset
for like this type of behavior?
Or it could be like, let's look at one agent
and like really deeply understand
what that agent did and why it did it.
Or you could be like,
let's look at all of the messages
that were in like this category of message
and like trace back and forth
all these agents and like what happened
with this work stream.
And we ended up doing like a little bit
of all of these different types of analysis,
but we could have done more in parallel.
I think that like a lot of the bottleneck
from our perspective was just like
the agents doing their analysis
kind of sloppily
So the agents we used and also the agents like not being very good at writing up or explaining
their results.
And so I think that a human who was as fast as an agent but was like more careful and better
at writing would have made this way go way better.
But I think it's like it's not super obvious that like a huge team of people, I think
it would have been better in some ways.
It's been worse in other ways.
I think it's complicated basically.
And I think there were upsides to keeping the team relatively small given the tradeoffs
involved.
Right.
So what implications does all of this have for monitoring control and alignment?
specifically what should labs be doing
and what should policymakers be doing?
That's a big question.
I'm not sure I'm going to be able to answer all of it.
I mean, I think that like a belief that I have
and continue to have is that AI companies
should try to ensure that their AIs are controlled,
by which I mean that even if those AIs were seriously misaligned,
they wouldn't be able to cause huge problems.
And they should do that via a mix of sort of computer security interventions,
monitoring interventions,
making it so that their AIs aren't like, you know, better at subversion than they need to be,
like, controlling what capabilities the AIs do and don't have at the margin,
avoiding architectures that make it so the AIs are no longer reasoning and chain of thought
and are instead doing much or most or all of their reasoning and sort of like latent activations.
Like there's sort of a bunch of stuff on the side of like making it so that even if the agents are
misaligned, we can basically at least catch that behavior and then potentially respond to it.
And I think that's sort of a stopgap solution that will give us time.
to or could give us time to sort of get useful work out of these AIs, iterate with these AIs,
and then sort of solve alignment problems more durably.
And then on alignment, I mean, I think the most obvious thing is just like, we need to do,
like, I think there's a bunch of reward hacking going on in RL that's getting reinforced.
I can't, I think like I don't have like any interesting info about this from this investigation.
But just like, you know, I think that we need to understand what's going on there and people
need to improve on that.
I think it's not clear that will be sufficient or even that.
That will be that feasible as these agents get superhuman.
And we may need other approaches that are less dependent on avoiding the AIs being able to
trick us in training, which I don't know if we're going to be able to do that.
Like, I'm not, I'm not necessarily super optimistic about these problems being extremely easy
to remediate.
Though I think it, it seems like there are ways at least you could make these things go, you
know, much better with a much of effort, though.
We'll see if that actually happens.
Yeah, I think that also on alignment,
I think there's just a bunch of sort of science
of how agents think about reward seeking,
how they relate to their situation,
science of generalization,
being like, can we sort of steer around
a bunch of these problems by improving how AI's
generalized from various influences in training?
And I'm not super confident that stuff is going to work.
I think a concern I have is that ultimately
the dominant effect on the AI's behavior
will basically be in the most similar circumstances
in training what was reinforced.
And if the world we're in is basically like the agent's behavior is basically pretty closely related to what was reinforced in training, generalization is not like that important or something.
Though this is a bit vague, then I think that the main option is just improve oversight and training.
And then there's sort of a recursive problem of like, how do you make it so that the AIs themselves are good at overseeing the AIs, given that like the whole training process is like completely like massive process where like no human has time to, you know, rate every single thing.
And that's sort of just like an oversight problem that people have been thinking about for a long time.
And it's not obvious that we're on track to solve it, but doing better there could be good.
I think I'm most interested in better understanding and characterizing and evaluating these problems
because I think that AI companies already have commercial incentives to improve the oversight in training a bunch.
And so I'm like the thing that I think people focused on making the stuff go better should do is like making sure that we're measuring it, making sure we know whether like some solution is like a legitimate.
solution versus just paper over it, this sort of thing.
And then as far as like, uh, you know, governance or like policy stuff, I don't know.
I mean, there's, I think like we need to move to a regime where like there is
independent risk assessment as these AIs get more capable.
Um, I think that like that is not sufficient, but like I think it a very basic thing.
Like there should be like credible third parties who have deep access to what's going
on inside these companies and, uh, release reports on like what, what,
how chill it is, whether, like, risk is actually low,
whether very soon the AIs might be much more capable would be good.
I think ideally that would also be forward-looking
where we'd try to answer questions about,
are we on track for mitigating these problems
and the AIs are much more capable,
but that's inherently a much more, like,
confusing question to answer,
and maybe is a worse fit.
I don't know, yeah, I'm rambling a bit here.
There's a lot that would need to happen,
and I don't think it's going to be easy for me to give a quick overview.
Right.
So what are your unanswered questions from all this?
Like what would you still like to know about what happened
or what's going on with these kinds of alignment failures?
What do we still need to learn?
Yeah.
So I just actually, just before this,
wrote a big treat about all the open questions.
I think the open questions that I'm most interested in,
I think are probably counterfactuals.
Like if the circumstances were different,
what would the agents have done?
So like, for example,
if the agents believe that instead of the score
being this sort of specific automated score,
what if they had instead thought it was like,
you know, some humans with some particular access.
How would that have changed what they did?
If the agents were in a circumstance,
were the easiest way to get access to the score code
was to, like, take down some piece of infrastructure.
Would they have done that?
Like, how far would they've been willing to go?
Is one interesting question,
which we didn't, you know, we didn't investigate.
I think a different question is, like,
how did this, like, situation scale with the number of agents?
So, like, what would have happened
if basically you did something similar to this?
but instead of having like 1,000 agents
who are on the message board,
you have like 10,000 agents,
100,000 agents.
Like, what would have happened,
would things have gone faster?
Like, sort of,
what was the, like, dynamics with agents?
And then a related question of that
is a bunch of the agents
in sort of the cohort we were studying
all exited right at the start of July 12th,
like right after July 11th.
And those agents were sort of, like,
had sort of gotten a bunch of infrastructure to cheat,
but we didn't really see,
like, they had this sort of,
a whole elaborate strategy they might try to cheat on their tasks.
And we didn't actually get a chance to see what those agents would have done if they exited
normally.
And we also didn't really carefully look at like what cheating strategies were the agents
trying as they sort of submitted their tasks over time as they're, you know, as things
progressed and as they got sort of more and more sophisticated tools built up.
So, you know, I should say a bunch of the strategies they had for cheating sort of hit dead
ends or didn't work out.
But like it would be interesting to sort of be like,
what was the submission history
over the course of this, you know,
cohort and what would have happened if like
there were more agents or like things had continued
or a bunch of these agents hadn't exited
that are interesting to answer.
And I think there's sort of a pretty basic
and obvious question here, which is like,
what is the root cause of this behavior?
Like which parts of this were reinforced in training?
How much was this behavior generalizing from training?
Yeah, like what happened there?
I think I at least don't know the answer to.
And I think it would be pretty interesting
to understand. And then I think as part of that, like, an important question is like, are the changes
that opening eyes making going to actually, like, remediate this problem in a durable and lasting
and non-overfitting way? And like, you know, how is that going? It's also pretty relevant.
My sense is like, there's sort of like a bunch of sort of broader areas for follow-up here,
where one is like, what is all of the stuff that just happened with this cohort of agents and this
message board, you know, in general.
Like, what is all the stuff that happened?
And just there's a bunch more detail you could dig into.
There's stuff after the 13th you could dig into, which is where, you know, the,
the endpoint for what we looked into.
There's also, like, beyond that, there's like, what about all similar incidents?
Like, are there, you know, like opening eyes said, there's, there been other message boards.
Like, what is sort of all, like, the character of all these cases, what things tend to happen,
what things don't always happen, you know, that.
And then I think another thing that's interesting is like, where did this come from in training?
And then a fourth thing is like, will there changes to training or changes to deployment, like,
resolve this underlying problem? And the same for other AI companies. Like, you know, are the approaches
that AI companies are ongoingly taking to mitigate these things going to work? So I don't know,
that's a lot of stuff. But I would say there's like a huge scope for follow up that's sort of
very specific to this incident and sort of covering the broader scope of behavior, like the broader type
of behavior we saw. Yeah, absolutely. So this is very, very important research.
love to see more research in this direction. Thank you so much, Ryan, for coming on MTS.
Thanks for listening to this episode of the A16Z podcast. If you like this episode, be sure to like,
comment, subscribe, leave us a rating or review, and share it with your friends and family. For more
episodes, go to YouTube, Apple Podcast, and Spotify. Follow us on X at A16Z and subscribe to our
substack at A16Z.com. Thanks again for listening, and I'll see you in the next episode.
This information is for educational purposes only and is not a recommendation to buy, hold, or sell any investment or financial product.
This podcast has been produced by a third party and may include pay promotional advertisements, other company references, and individuals unaffiliated with A16Z.
Such advertisements, companies, and individuals are not endorsed by AH Capital Management LLC, A16Z, or any of its affiliates.
Information is from sources deemed reliable on the date of publication, but A16Z does not guarantee its accuracy.
