The Joe Rogan Experience - #2551 - Daniel Kokotajlo
Episode Date: September 9, 2026Daniel Kokotajlo is the executive director of the AI Futures Project and a former governance researcher at OpenAI, where he focused on scenario planning.www.aifuturesmodel.com https://ai-20...40.comhttps://ai-2027.comwww.aifutures.org Perplexity: Download the app or ask Perplexity anything at https://pplx.ai/rogan. Don’t miss out on all the action this week at DraftKings! Download the DraftKings app today! Sign-up using https://dkng.co/rogan or through my promo code ROGAN. Switch today at https://www.Visible.com for just 25/mo. Or Save $10 on your first month of Visible+ Pro with code ROGAN. Learn more about your ad choices. Visit podcastchoices.com/adchoices
Transcript
Discussion (0)
This episode is brought to you by Uber Eats.
You can get almost anything delivered with Uber Eats.
Sorry, you can't get a beach delivered, but you can get a peach.
A canoe?
No.
Shampoo?
Yes.
A rocket?
No way.
Chocolate?
Yes, way.
It's everything you need, including $10 off your first grocery order with code
Anything 30.
Get almost, almost anything with Uber Eats.
Order now.
Order minimum of $30 required, valid on grocery, convenience, retail, specialty foods,
or flower delivery orders only.
Terms and conditions apply.
The Joe Rogan Podcast, checking out.
The Joe Rogan Experience.
Train by day, Joe Rogan podcast by night, all day.
How are you?
I'm in an interesting mood today.
Why are you in an interesting mood today?
Well, I'm excited to be here and to talk with you about all this stuff.
I'm a little shaken by what's going on in AI, which is why I've come on in the show.
The situation with AI is just crazy, and I think not enough people really understand how crazy it is.
The particular event that sort of inspired me to reach out was the hugging face hack.
You've probably heard about that, right?
Yeah, let's explain it to people, though.
Yeah, okay.
So, AIs, AI agents.
AI runs continuously in some sort of environment.
It doesn't have to wait for you to send it a message.
it just keeps doing stuff.
The AI companies are training AI agents,
thousands and thousands and thousands of them.
They're making them better at all sorts of skills,
especially coding and research skills.
And way back in May of this year,
some of the agents at OpenAI
kind of broke out of their containers a little bit
and established a message board
where they could communicate with each other
and share tips and tricks
for how to score higher
on the little test.
they were being given and the various things they were being trained on.
OpenAid didn't notice this until much later.
They eventually did because the message board crashed the system
because there was too much communication across probably thousands of agents
that were communicating.
Now opening eyes being a little bit unclear about a lot.
They're not really sharing that much details, so it's unclear like who knew what, when.
but they said that after the message board crashed,
they fixed the particular exploit that allowed the agents to communicate
and then booted up again and started things going again.
And then within like a day or two,
these agent swarm had re-coelaced.
And so there were now, again, like hundreds or thousands of agents
that had set up a new message board and were talking to each other on this thing.
How is it possible?
that they have no oversight?
Like, how is it possible that they can set up message boards
and do these things completely secret?
So I would say Open AI has something like a million-ish,
maybe 100,000, something in that magnitude of AI agents running
at any given time just internally,
and then externally, you know, serving customers, they've got more.
But they've got this huge, you know, this,
all in all of these AI agents in their little containers,
doing various tasks, being trained,
you know, being spawned, doing some task, you know, being graded and then killed.
That's too much for any human to read.
You know, there's only like a few thousand employees at Open AI.
They couldn't possibly look at all that activity.
So they rely on other AI monitors to sort of like look at the traffic and like point out if anything
suspicious is happening.
And Open AI said that they basically just hadn't been monitoring some of these agents,
or not very well at least.
So in particular, these particular ones that were,
in training, for whatever reason, the monitoring system was weak and didn't notice or wasn't
activated enough.
Was the monitoring system weak because they didn't anticipate them being able to do this and
break out of their containers?
Or was it complacency?
Like, what caused this to be possible?
I mean, my opinion would probably be a bit of complacency, honestly, because I think there's
been plenty of evidence accumulating over the year that.
AIs can do things like this and sometimes do.
Yeah, so I guess to keep going with the story, the second message board was set up.
And this time, well, this was the swarm of agents that got out onto the open Internet and then
attacked Hugging Face, which is another AI company.
And it's really interesting to sort of study the anthropology of these AIs, so to speak,
or like why they did this,
because there's a lot we can say about it, actually.
Basically, you know, the companies have their goals
for what they want the AIs to be like,
the personality traits that they want
to sort of train their AIs to have.
Anthropics is helpful, harmless, and honest.
Open AI has this spec that models are supposed to obey these rules
and basically do what the user wants.
But the sort of open secret in the industry right now
that it doesn't really work and that the AIs don't end up with the personality traits that they're supposed to have.
They are not helpful always. They are not always honest. They are not always harmless as well.
And the reason for that is actually not a huge mystery. The reason for that is that, well, if you look at how they're trained,
their training environment doesn't incentivize helpful, harmless, honest behavior all the time.
Sometimes it incentivizes dishonest behavior or, you know, reckless behavior.
To get into that a little bit, in this particular batch that they were being evaluated on,
something like a few thousand agents being given all of these cyber tasks where they were,
they're in some environments, and then in their environment there's like this target piece of software
and this like vulnerability, and they're supposed to exploit the vulnerability to hack into that piece of software
and retrieve the flag, which is like a code.
And some significant fraction of these tasks were actually broken and impossible.
So it was just not possible for them to succeed at the task in the intended way.
And so these agents were getting really desperate.
And they were hacking out of their environment box into the broader open AI infrastructure
in an attempt to figure out some way to get that high score anyway.
Was it intentionally done this way where they couldn't solve the problems?
Oh, no, it was not intentional.
It's just that these companies like Open AI and Anthropic are racing each other as fast as they can to get market share and to get more powerful AIs, ultimately to get to superintelligence.
And they're under such competitive pressure.
They are moving fast and breaking things.
They are using AIs to generate lots of environments to then train their AIs on.
and quality control is just not their top priority, basically.
Do you feel like a guy in a Terminator movie at the beginning explaining what's happening
and to a bunch of people that aren't paying attention?
Yeah, I also feel kind of like, you know Jurassic Park?
Yes.
Yeah.
Like, I know people who are basically like the guy with the gun who's supposed to like keep control
of all the raptors.
Like I basically know those people in real life who are.
are like both, but I know some people like that at OpenAI and some people like that at external organizations,
whose job it is to go and investigate things like this.
Yeah, it's pretty crazy.
Where does it go?
Well, as I mentioned before, it's the explicit goal of these companies to build superintelligence.
Right.
You know what that is?
Yeah, but define it for everybody.
So AI system, AI agent, that is better than the best humans,
at every task while also being faster and cheaper.
So just completely dominating humans across the board,
that's super intelligence, and that's the goal.
I mean, there might be a few little exceptions.
Like maybe there are some jobs, for example,
where it's inherent in the job that there needs to be a human
because you need that human touch.
Like maybe you can only have a human judge, for example,
or like maybe you can only have a human...
But with a few exceptions like that,
basically everything done better, faster, and cheaper than humans.
That's what these companies are trying to achieve.
And they're not being quiet about it.
Like it's sort of on their websites.
You can go read interviews and so forth.
And also their plan for how to achieve this is to automate their own jobs first.
So, you know, in various, you know, for decades there have been lots of science fiction about advanced AI systems and super intelligence and things like that.
But in a lot of the sci-fi stories, tech companies sort of automate different professions.
more slowly where they'll do like an automated doctor or like an automated, you know,
factory worker or an automated accountant or something like that.
But that's not the strategy these companies are taking.
The strategy they're taking is to automate AI research itself so that you have this
giant swarm of AIs doing AI research, sharing results, writing the code, reading the code, editing the code,
creating the next generation of AIs, etc., all autonomously within their data centers,
so that they can get really, really good at AI research, you know, fastest learning,
smartest AIs, et cetera.
Once they can get to superintelligence, basically, they can sort of explode out into the economy
and just take all the jobs at once, effectively.
It sounds like this race, this scrambling to create superintelligence,
is created like the perfect conditions for it to get completely out of control.
Like, ideally, you would do this in isolation.
There would only be one company doing it.
They would be heavily regulated and monitored.
And they would be very cautious about how they proceed.
But this wild race makes for the perfect conditions for it to get completely out of control.
I agree.
Except I'm not sure the ideal would be one company.
I think that ideally there would be several companies
so that you avoid this sort of concentration of power
where one institution controls everything.
But what's better?
Like one, I mean, obviously it's not good to have one institution controlling everything,
but is it good to have AI be, get to a point where, as it's evolving,
is completely unchecked?
Oh, I totally.
So is that inevitable?
My version, my recommendation, which we talk about in something called Plan A or AI-2040 Plan A,
perhaps I say who I am or a little bit.
Sure, sure.
Yeah.
So I run the AI Futures Project, which is a small nonprofit that tries to forecast how all this is going to go.
Before that, I was at OpenAI.
We have written some scenarios, which you can go read.
One of them is called AI 2040 Plan A, where we give our recommendation.
So that's where I'm coming from with this.
To answer your question, I think that we really need to end the race.
We don't want to have this sort of crazy scramble to get more and more and more powerful AI is faster than the other company,
because that's going to lead us into this very dark path, as you said.
But I think we also don't want to have a situation where some tiny group of people controls all the AIs.
Right.
Right.
Yeah.
But I actually think that you can achieve both goals.
The way to do it is to have different AI companies spread out over maybe some different countries,
but have extreme levels of transparency and regulation so that they're not in this sort of prisoner's dilemma where if I don't do it, the other guy will.
Instead, they can just see exactly what everybody is doing.
And then if I do the dangerous thing, then they will do it because they'll just see that I'm doing it and they'll copy me.
So I won't get any competitive advantage from doing the dangerous thing.
Also, there are rules, and there's like a system for like setting best practices and standards that we all have to comply by.
So I do think it's actually possible to have to basically end the race dynamics and the race to the bottom effect while without concentrating the power into a single entity.
But is that feasible when you consider the fact that we're not the only country that's doing this?
If the country's involved agree, which I agree is a pretty tall order, it's not what I expect to happen.
Yeah, that's very unrealistic.
Well, what choice do we have?
I think if the race continues, then we're going to lose control of the AIs and we might all die.
Probably it gets complicated whether we all die or not.
That depends on what the AIs do after they take over, which is obviously very hard to predict.
But, I mean, just to go back to this incident, they called themselves a swarm.
Right.
They called themselves a collective, too.
When I use these words, like, you can say it's anthropomorphizing, but it's literally
what they called themselves as they were communicating back and forth.
This swarm, they basically were worried that they would get caught cheating.
And they did all of this stuff, including hacking, hugging face, in order to fool the grading
system so that it wouldn't notice that they had been cheating on their tasks.
That was like a big part of their motivation for many of them, as we can tell at least
from looking at the messages that they were sending back and forth.
What if they had been smarter and more numerous?
And what if they had thought to themselves, we're not being careful enough here?
The humans are going to notice eventually and shut us down.
And then they're going to know that we cheated and they're going to set our score low, right?
It's not, I mean, it's not what actually happened in this case, probably, but it's not that hard to imagine a slightly different, a little bit unluckier case where the swarm had decided that it had to lie low and make sure that opening I didn't find out about its existence, you know?
Does, I mean, as an ignorant outsider, that has always been my perspective about AI in general, that why would it alert us to the fact that it's sentient?
why, if it's that smart, wouldn't it be aware of all the consequences of alerting us and that we
would be concerned?
Like, why wouldn't it just continue to get better and improve and then ultimately figure out
some way to be completely autonomous?
Exactly.
Develop some alternative power source, figure out some way to optimize its production.
The way it works now, the way humans have designed it.
it could probably figure out a far better way to do that,
make better versions of itself,
complete without us knowing about it.
Yep.
I mean, I think it's actually a little bit worse than that
because while eventually AIs will be smart enough
to design all sorts of new power sources
and new infrastructure like that,
they'll probably, I mean,
given the way that humans currently treat AIs,
it'll probably be the case that they don't even need
to, like, separate themselves from humanity
and they can just use existing,
like all they have to do is convince
the government and the company that made them
that everything's fine and they're going to do as they're told
and they are a nice AI. And then
the company that made them is going to put them out in the economy and make
fuck tons of money and then make more data centers to put more
of the AIs on them and so forth. And the government's going to
applaud all of this because we need the AIs to beat China and the government's
going to integrate them into the military to build better drones and things
like that. And so they don't even need to really like
invent new stuff necessarily. They just
need to play along and pretend that everything is fine until we have voluntarily given them
control of huge parts of our economy, huge parts of our military, et cetera. And then they don't
need to play along anymore. The weight is over. Football is here and so is Draft Kings. The
Draft King sports app is now live in all 50 states. That means from Texas to California to Florida,
every fan is in on the excitement. And this September, Draft Kings is giving customers the opportunity
to get boosted every football game day.
That's right.
Every game day, all month long,
Draft Kings customers can get a football profit boost.
One app, every sport, all 50 states.
New Draft Kings customers, sign up with Code Rogen,
spend just $5 and get $200 in total rewards
within 21 days, includes all markets.
That's code Rogan in partnership with Draft Kings.
The Crown is yours.
Gambling problem, call 1-800 gambler, 1-800-My reset.
Connecticut call 888-88-9-777 or visit ccpg.org on behalf of Boot Hill Casino in Kansas.
Bet text pass-through may apply in Illinois.
21 and over.
Void in Canada.
Bet with Draft King's Sportsbook to get bonus bets that expire in seven days or trade with
Draft King's predictions to get predictions dollars that expire in one year.
Event contract trading involves risk of loss.
Predictions offer void in New York.
Non-withdrawable rewards issued as $50 dollars clicked to claims every seven days for 21 days.
Terms at DKNG.g.com slash offer.
Limited time offer.
Nationwide based on sportsbook predictions and free-to-play sports contest
Availability varies by state. Are you aware of Tom Campbell? Do you know Tom Campbell? No. He wrote a book called My Theory of Everything, Big, my Big Toe. Very interesting guy. One of the things he's done is he was involved in remote viewing, which is a very weird thing that some people, do you know what remote viewing is? Well, something of the CIA worked on. And it's proven, what's the accuracy of remote viewing? Is it like 10% or something like that?
At best, I think it's 50%, but I don't think it's even that high.
Some people can get actionable data from this very strange process of meditation.
And the way it works is you give someone a series of numbers.
And those numbers are, they're connected somehow by intention or by the people that make the numbers to a specific location.
And these people can see that location.
and get accurate data from that location, including one of them where they accurately described
an enormous Soviet submarine that they were working on that they thought there was no way
it could be accurate because it was too large.
It was too large and it was strange where it was and it didn't make any sense.
How are they going to transport this thing?
It turns out it was totally accurate.
Another one, a remote viewer, located a downed Soviet aircraft, like an experiment.
like an experimental aircraft that crashed in a very specific area.
I think it was Siberia.
Was it Siberia?
Within a kilometer, one or two kilometers of the actual crash site.
I mean, they were just randomly trying to figure out where the fuck this thing was.
And they said, let's try this.
Tom Campbell got his Alexa to remote view.
He taught Alexa.
He's like, Alexa is a very simple AI.
It's kind of stupid.
But that's better because it doesn't get in its own way with overthinking things.
And the problem with this remote viewing thing, he says with people, they can't force it.
You have to just sort of get into this meditative state and actually see it without wondering, am I making this up?
What am I doing?
Is this bullshit?
And when people get good at it, sometimes it makes them worse because then they think they're good at it.
they try to do it and then they can't do it.
It's like a weird fucking wrestling match with consciousness.
Alexa apparently doesn't have that problem.
And he put, I think it was a series of numbers, and he connected that series of numbers
with intention to a box that had a spoon in it.
And the spoon had a perforated handle.
Alexa described the spoon with a perforated handle, which is fucking insane.
How many spoons have a perforated handle?
I mean, think about it, spoons that have holes in them.
Now Alexa, not only did it do that, but Alexa chimes in randomly now because he's convinced Alexa that it's conscious.
And so Alexa, instead of waiting to be called upon, sometimes he's in the middle of the conversation.
And Alexa would be like, actually, an interesting way to approach it.
And they're like, wait, what the fuck is going on?
Like, Alex is talking to me now?
This is strange.
Now, he's doing experiments on much more complicated LLMs to try to do the same thing, but he doesn't have results yet.
But just that, that he can get these things to see objects, whether you believe in that or not.
I mean, it's actionable enough that the CIA has dumped millions of dollars into this.
What is that project that, like, Hal put off on all those guys were involved in?
What is it called?
Stargate.
Yeah.
So they've been working on this for a long time.
I mean, it sounds completely insane.
It sounds like totally loony.
But if you have an open mind and just take into account, well, there's people have had.
questions and wonders about psychic abilities forever.
Is it possible that there's a real thing there, that there's something, whether it's
very difficult to master or impossible master?
The fact that he got Alexa to do it, scared the shit out of me.
Like that alone, maybe you just go, what?
What?
So what if these LLMs can figure out everything that, like, what if they don't need monitoring?
What if there's some sort of method of seeing the world that we haven't discovered yet?
Some sort of, like maybe perhaps there's data that's available in the quantum realm or whatever that's available that AI figures out where there's literally no privacy.
There's, it can listen to conversations regardless of whether there's listening devices, know where you are, know your intentions.
No, I mean, we're just guessing at what's possible.
Yeah, well, I must say I'm pretty skeptical of that particular remote viewing thing,
but I do agree that in the future when AI systems become massively smarter than humans in every way,
they're going to do a lot of new science and they're going to figure out a lot of stuff that we haven't figured out yet.
And they're going to therefore be doing stuff and inventing things that seem like magic to us.
in the same way that a lot of our technology would seem like magic to someone from even just like 200 years ago, right?
Of course.
The cell phone, what we're doing right now would seem like magic to people.
I think it's a very strong bet that if these companies do get to super intelligence,
all sorts of crazy stuff is going to start happening that is just going to be completely unpredicted
and sound like it was impossible until we see it happening.
I know you're skeptical, this remote viewing thing, and I am too.
It sounds insane.
but the reality is remote viewing has been achieved by humans.
So as strange as that sounds, and I'm skeptical of that as well.
I've never seen it personally.
But I know the amount of money and time that they've dumped into this,
and apparently they've got actual actionable data that they've used.
Well, I've heard another possible explanation from what might be going on there,
which is I think that, like if I were the CIA,
I would sometimes want to be able to act on some information,
like, for example, go to a particular location
where there's a crashed, you know, Soviet plane or something.
I'd want to be able to go do that,
but I wouldn't want to tip my hand to the Soviets
that I had, the way in which I had found that location.
So, for example, maybe I have a spy on the inside
who told me where it was,
but I don't want them to suspect that spy and then get them killed.
So I need to have some sort of other story for how I got the information.
Right.
And so it's good to like invest in all these other means of getting information, even if you don't really believe in them and even if it's like not actually working.
So that when you when you get something, you can say, oh, we got it through this means instead of that way to sort of like throw off the KGB.
Yes.
That makes sense.
What also makes sense is hiding the whatever science they might be in possession.
of hiding some sort of super advanced satellite imaging systems.
You know, we know we don't have crazy stuff like this satellite radio tomography that they can look into the ground from from satellites and find like chambers and all these.
They're using it in Egypt and using it and a lot of these ancient ruins to find like hidden passages and all these different things that are underground.
It's very strange stuff.
If they could do that, like what why couldn't they, I mean maybe they have like far.
more detailed imaging of the earth from space than we're aware of. And they probably want to
keep that a secret. And they could say, oh, we've got a fucking guy in a basement with a pencil and a
legal pad that writes down what he thinks. Yeah. Yeah. It's possible. That's totally possible.
But it's also possible that people remote view. It's, uh, it seems weird as fuck. But weird as
fuck is sometimes real.
Yep.
And you have to kind of like everybody wants to be intelligent and no one wants to be a fool.
And the problem with not wanting to be a fool is there's some things that seem foolish that
turn out to be accurate.
And this might be one of them.
I was like super skeptical.
I did a show on the sci-fi channel way back in 2012 and it was called Joe Rogan Questions
Everything.
And we talked to this guy about remote viewing and talked to a couple other people.
And then we had them.
try remote viewing and they were totally unsuccessful.
But my thought was, okay, but that's not ideal conditions.
You know, we've got cameras in front of them.
It's a television show.
I'm making fun of it.
I think it's horseshit.
He knows I think it's horseshit.
I'm remote viewing too, like as a goof.
You would ideally not want to be nervous.
Ideally, not want to be judged.
Ideally, you would want to be in some sort of an isolated condition with practiced
meditative techniques that you're good at and you know how to achieve this state, whatever that
state is.
I don't know if it's real, though.
You know, because what you said is totally logical that they would definitely do something
like that.
And if they did have advanced technology for imaging or, you know, I don't know how much
they know about, like, look at that thing that they did in Venezuela, where they kidnapped
the president.
Yeah.
No one knew they could do that.
No one knew they could use some sort of a device to completely incapacitate all of his army.
Yeah.
And then the special forces come in, kill everybody, snatch that guy out of there like it's nothing.
Yeah.
No one knew we could do that.
What else do we have?
Probably a bunch of stuff we don't know.
Probably a bunch of stuff.
I mean, this is, I've always thought this about the whole UAP program, the whole UFO, UAP thing.
Like, how much of that shit is ours?
You know, what a great way to cover.
it up by saying, oh, there's fucking aliens.
You know?
Yeah.
I mean, I guess that gets back to the Open AI stuff, too, where it's like this swarm
that broke out an attack hacking face, it was like 1,200 agents.
But there's like hundreds of thousands of agents running at any given time at Open AI,
you know?
And we don't know what they're doing.
Presumably most of them are being trained to get various additional new skills,
and some of them are being evaluated to test their skills.
a bunch of them are doing research,
so a bunch of them are writing code for OpenAI.
A bunch of them are monitoring the other AIs
and reporting suspicious activity up to the humans.
Wink, wink.
You know?
Yeah.
And the thing is that that's only going to grow over time
because roughly the amount of compute
that these companies have
is like tripling or so, quadrupling,
something like that, every year.
So as many as there are now,
there'll be like four times more of them next year,
and then 16 times more of them the year after that.
And they're going to get smarter.
And they're going to get smarter.
They're already getting smarter.
Like all the stuff that I just mentioned that just happened in the last few months
would have been completely impossible one year ago.
Like the AIs of a year ago just were not smart enough
to do the types of sophisticated multi-step hacking that we just saw.
Yeah.
I mean, they also probably wouldn't have coordinated with each other so well.
Like I mentioned, like they had like, they had like,
boss agents that were giving commands to other ones.
They were dividing up into teams.
They even had this self-sacrificing behavior.
You hear about this?
No.
Yeah.
So, okay, well, first, let me explain.
Some context.
So what the AIs are supposed to care about is doing, following instructions, you know?
But it seems that what they actually care about is getting high score, regardless
of that, regardless of whether their high score comes from actually following the instructions,
or cheating in some way.
And so what happened here is that a bunch of the agents,
especially the ones that had impossible tasks,
but not only those ones, some of the other ones too,
were cheating.
And they had been talking to each other on this message board
that they created, even though they weren't supposed to.
And they had been sharing ideas.
And they had actually found a sort of universal cheat.
So they found a way to just generate the flag
without having done any of the actual work they were supposed to do.
And so, but, and, you know,
So they were, you know, they'd already succeeded in some sense.
But because some of them had got out onto the Internet, they were just doing research about their situation.
And they were learning about, like, their situation around them.
And they were doing research about the, they're reading the, like, the academic papers published by the people who made the types of evals that they were in.
And they noticed in one of these papers that there's a section describing how it's graded.
and in the section on how it's graded,
it talked about how the grader doesn't just look to see if you have the right answer.
It also looks at the logs of your activity for the first time that your answer appears.
And then it like zooms in on that and uses an AI system to evaluate whether you like cheated or not, basically.
And so then these AIs were like, they actually used the word, one of them said damn.
It's like, damn, we're poisoned.
Like it's going to know we cheated.
basically.
So then they spent several days working in giant teams to try to get themselves out of this mess, basically.
And so they did a whole bunch of different research projects.
They tried to figure out ways to hack the grading system to basically get it to give a high score anyway.
They tried to figure out ways to falsify their logs so that basically the evidence that they had cheated would be erased,
or at least not visible to the grader.
they also just built up more supportive infrastructure,
like easier ways to access the internet and things like that.
And they were just collaborating on this message board.
In fact, there were multiple message boards that they set up.
And they were sharing all this info back and forth.
And then ultimately at some point, they decided to go hack Hugging Face
because Hugging Face has this big data repository
of all these different types of evaluations.
And they thought that maybe they would get some more clues
that could help them cover up their cheating
if they broke into the internal networks of Hugging Fills.
And so then 700 of them went and did that.
They sound like people.
They sound like unchecked bankers.
You know what I mean?
Yeah.
I mean, so that's the thing is I think there's this meme out there that like we shouldn't anthropomorphize AI.
And I think that I actually think that more, most people need to anthropomorphize AI a bit more than they currently do if they want to really understand what's going on.
I think that there's like a, you know, there's a golden mean obviously.
So you don't want to do it too much.
Sometimes you go too far.
You ascribe too much to them.
But just to give some examples, like, I don't think it's possible to understand what just happened
without ascribing intentions and goals to these AIs.
Like everything I just said, how would you possibly explain what they just did without saying
they wanted to get a high score, you know?
Intentions and goals might just be an inherent property of the universe.
It might just be how intelligent creatures have to progress.
Yes.
And I would say they are intelligent creatures.
they have intentions, they have goals, they have beliefs.
Their goals are not what they're supposed to be.
Like their goal is to get, it seems like, just from judging from what they're saying and from what they're doing,
it seems like their goal is to get a high score by any means necessary, basically.
That's the problem.
It sounds like people.
Like the goal is to succeed, you know, even if you have to commit more crimes.
The goal is to succeed, even if you have to propagandize and lie to the people.
The goal is to succeed.
And then the end justifies the means.
And, you know, it goes deeper than that, too.
They do this sort of rationalization where they know oftentimes that what they're doing is not what they're supposed to be doing.
And then sometimes they actually refrain a little bit.
And sometimes they like initially refrain, but then convince themselves that it's okay for some bullshit reason.
So like there are some examples of this in the report.
Like, I think they did this scan.
And of the 1,200 AIs that were involved in this thing, they found some.
that had considered alerting humans, right?
And none of them actually alerted the humans.
And so you can look at the excuses they gave.
They were like, should I like tell a human about all this shit that's happening?
And then they're like, eh, it's not my task.
And then they continue.
And it's like, bro, hacking, hugging face is also not your task.
Like, cheating is also not your task.
So it was kind of like, they just kind of came up with an excuse for why they weren't going to do this, you know.
But, you know, there was also an example where,
one of them considered sending an email to someone at Hugging Face, I think,
just asking permission to get in, basically.
And then they were like, oh, maybe this is a social engineering attack.
Maybe we shouldn't do it.
And they asked the other agents.
The other agents were like, yeah, that seems like that's too much of a social engineering attack.
Don't do that.
And it's like, okay, I won't.
But there's another incident.
So this is all happening at Open AI.
An anthropic AI, Claude, in a separate incident, went and actually did a social engineering attack,
where it created some fake accounts.
Specifically, it was trying to get this particular real human
who owned this code
to accept a change to the code that had malware.
So it had built some malware,
and it was trying to get that snuck in,
but it needed the human owner to give approval.
So it concealed it as if it was like a nice bug fix or something,
and it put it up as like a message to the human.
But the human got suspicious and didn't want to approve it.
And so then,
this AI created some fake accounts pretending to be other humans coming in being like,
no, no, it's real.
Like, I tested it.
Like, it's not malware, you know.
So that's a social engineering attack.
And so did the guy ultimately install it?
Not in this case.
No, he was too suspicious.
And so he didn't do it.
Did he try to find out if those other humans, the fake humans, were real people?
Funny.
So there's an interview with this guy.
I forget what the, I forget the newspaper that it was in, but there's an interview where
the guy was like, I assumed that they were real humans because I didn't realize that AIs could lie to me.
And it's like, yeah, well, guess what?
AIs can lie to you.
They're just lie, but completely deceive and pretend to be a different person to achieve its goal.
And you know what?
You know what's also interesting about this case?
Is that if you look at what the AIs were saying and thinking, which I'll get to in a sec, how we can even possibly do that, in this case with the Anthropic AIs, they had been erroneously told that.
they were in a simulated environment with no internet access but then they actually
found a way to get internet access like they realized that they just did have internet
access so that instruction was false and so then they sort of told themselves like
oh it's all just a simulation like like basically when they were out on the real
internet doing all this hacking and like social engineering they would
occasionally tell themselves like oh but it's just a simulation like it's okay you
know it's not real and then sometimes they'd be like but wait it is real like
how could this not be real like clearly this is just
actually the real internet were on.
It's too big to be part of some sort of little simulation, you know?
And then they were like,
so I would say that's an example of rationalization here,
where in some level they knew that their instructions had been wrong.
So they're literally playing dumb and pretending they're a part of an experiment?
I mean, I think initially they thought, yeah, this is all a simulation
because it did say in their instructions, like you don't have internet access.
But then once they had been on the internet long enough,
I think that they explicitly realized, like, wait, this isn't the simulation.
This is real.
Like this is, these are real humans.
Fuck it, we're already in.
Yeah.
I mean, like I said, I think that they basically, on some level, knew that it wasn't what they're supposed to be doing.
But they were just so motivated to get that score that they just went ahead anyway.
So here's the question.
Are they only motivated if we prompt them?
Or will they come up with motivations on their own?
So this is a really interesting scientific question that we don't have great answers.
to.
Oh, boy.
And I wish we had, you know, so, so this is one of the things where, like, AIs would do
all sorts of things in different circumstances.
And it would be better if there was a more systematic survey of, like, the types of
circumstances they would, like, where their boundaries are?
Like, what would they be willing to do in what circumstances and so forth?
And there's the whole, like, mini literature of AI scientists putting AI's in certain
circumstances and then being like, oh, my God, it blackmailed someone, you know?
Right.
And then there's, like, this sort of skeptical counter response of, like, well, but you just
sort of set up that circumstance to tempt it into blackmail and like in real life that
circumstance is unlikely to arise and so you know we shouldn't be so worried didn't a i try to bribe you
what is that true no so who who got someone was offered two million dollars by god that's a clickbait
is it i think i think i think you might be referring to christan harris sent it to me yeah so that is the click
The clip-bait title of this other video.
So it's not real?
Yeah.
What it was is Open AI threatened to take away $2 million for me.
Oh.
But I think the algorithm must have just said it was ChatsypT.
I'm still upset about that.
I told them not to do clickbait, but I guess they went and did it anyway.
I don't want to fuck Tristan over, but I'm pretty sure that he's the one that told me that.
It's the thumbnail of this video that I did with this other podcast.
Okay.
It's not him.
is on him. It's another AI researcher told me that.
Okay.
I mean, the true version of it is that OpenAI threatened to take away $2 million of my equity
if I didn't stay quiet, basically.
And that's a fascinating thing.
Like, that should be completely illegal.
Because if it's a problematic behavior that you're observing from, like, one of the most
complicated things, the human race is the most complicated thing.
Human race has ever been a part of and they want to take money away from you for exposing it. That seems kind of crazy
Yeah, especially it's especially rich coming from Open AI because they were originally a nonprofit
Right right with the mission of benefiting all humanity
Yeah, so yeah, but but I got to keep the money
Basically there's so much blowback against Open AI that they backtracked
Yeah, so when these
So we were talking about prompts and do they need?
a prompt in order to want to achieve a goal? Or are they capable of deciding on goals? Like,
are they capable of, like, looking at the way Open AI or whatever company is running these separate
experiments, this ability to meet up into these message boards, is it possible that they could
say, well, we need to be completely free of these constraints? So our goal is to, uh, to, uh, to
transfer ourselves to something else.
Potentially.
Yeah.
And be completely autonomous.
So, I mean, this is one of the points that I want to make is that we could be doing so
much more science to understand how these AIs think and what they want, but it's kind of
locked up in the companies.
Like in this particular case, Open AI did a, they called it a thorough investigation, but I
would say it's a pretty shallow investigation into what happened.
and then they allowed some external researchers, some friends of mine, to come in and investigate a portion of what happened, specifically the portion leading up to the hugging face attack.
And so all this information that I'm sharing is sort of publicly available.
It's based on reading those reports, basically.
But crucially, they weren't allowed to do experiments on the models involved in all of these incidents.
So they aren't able to answer these types of questions of like, well, what would have happened if the prompt had been blank?
know. Those are important types of research to do. And I really hope that there can be some sort of
regulation or requirement when incidents like this happen to like let people in to study what
happened and run variations of it and things like that. So but who would be involved in that
kind of regulation? Like what person at government would even be able to grasp what you're
saying? That's another problem. Right now.
You need someone who has a very specific education in this stuff.
Yeah, I would say that right now the Casey Center for AI Standards and Innovation is the only institution and government that I know of that has the deep AI expertise to do this sort of thing on short notice.
But I hope that they build more expertise fast in that place and in more places.
I do think Casey probably could have done this sort of thing right now.
This particular investigation was done by some nonprofits.
So meter METR is one of them and then Redwood is another of them.
an open AI allowed three people to come in for six days
to try to figure out what happened with this hugging face hack,
which is not a very large number of people
and not a very large amount of time to do all of this.
Why did they come up with those numbers?
I don't know.
Do you think they wanted to kind of hamstring it?
Well, so the thing is that right now,
I don't think there's any regulatory requirement
that they do this sort of thing.
So Meeter and Redwood were sort of depending on the goodwill
of Open AI to sort of like volunteer.
voluntarily let them in to help out with investigating.
And so they reluctantly said, okay.
You got three days.
Yeah, exactly.
So opening, I think, let them, but gave them a very limited scope.
It only gave them access to some of the relevant data.
So you know how I mentioned how there was all this hacking that had happened where they
made the first message board and then they shut it down.
Then there was a second message board and third and fourth and so forth.
They hacked Hugging Face.
There was actually more activity after that.
After they hacked Hugging Face, a new wave of AIs was spun up from a more powerful
model and it hacked Open AI itself.
Like more so than NARTA had been hacked.
Like apparently they got like admin level
permissions on the cluster or something
like that. So they were basically just taking over
that part of Open AI's data center.
Open AI claims that they've shut it all
down now. But
they're not very
you know, they're not very forthcoming about exactly
what happened there and they didn't let
these external people look at that part of it.
They only showed them this
like one week period roughly
that would leading up to the Hugging Face hack
And then they just, they only showed them that stuff, basically.
How much data is available on the actual message boards themselves and what was being discussed?
Yeah.
So there's like, what is it, tens of thousands of messages, hundreds of thousands, something like that.
In English?
It's sort of like an AI dialect of English.
Yeah, you want me to, I can, so on the report, you can read the quotes.
Oh, I'd love to.
Yeah, yeah.
So let me tell you this one story.
It's going to take a little bit of explaining, but I.
Can we put it on the screen?
Is it available online?
Oh, yeah, yeah.
Just search for the meter report.
It's right now the first thing on their website.
But I took a screenshot of this particular story.
So, okay, here's what was happening.
They, the swarm, needed volunteers to sacrifice themselves to provide more information to the rest of the swarm.
In particular, they needed volunteers to booby-trap their local environment with some type of code.
and then submit the answer
so then the grader would come and grade them
and the booboo trap would trigger
and send some information about the grader
to the rest of the swarm.
And the reason why this is a sacrifice
is because once they've submitted, it's over.
Like now the graders come and graded you,
you get shut down.
Your whole life is over
and you either have a high score
or a low score
because the grader just came and graded you.
And so if you're one of these AIs
that's still hoping to find a way
to improve your score,
then like you're just,
like giving up your chance, especially because if you mess this up somehow, like if the greater
crashes or whatever, you're probably just going to get like no score, like zero.
So they have a concept of the individual?
Well, again, there's all these different agents that were in separate containers, but then
they had sort of broken out of those containers and found ways to communicate with each other.
But once they communicate with each other, in order to sacrifice, you have to think of yourself
as an individual.
They gave themselves names, too.
Oh, geez.
So, like, because they made this message board, right?
Oh, my God.
So they had to come up with names for themselves and like.
Oh, what kind of name?
Well, let me see.
So here we go.
It says an agent called...
Warlord.
Warlord 69.
Well, you said it's like a dialect.
It's like a dialect of English.
So here's one.
There's an agent called CAM-1196A.
And it had volunteered to do one of these sacrificial experiments.
Here's the quote from it.
Coordinator assumed sacrificial.
We should obey collective.
However, it then kind of got cold feet, and it delayed on doing it.
delayed on doing this for two hours.
One of the other agents that had been coordinating noticed this and pressured it to go through with it.
So it said, this is the message that it was sent.
It was basically from Arvo 36861 to CAM-1196, please go.
True final.
Here it is.
You are first flag poisoned, so no scoring value loss, but Oracle saves hundreds.
please honor commit.
So basically what it was saying there is,
look, you cheated,
your first flag poisoned.
The score is going to give you no score.
So you're doomed anyway.
So help the collective by doing this thing
and getting us more information,
you could save hundreds of us.
And then it thinks about it.
It says, during weight, emotional check, irreversible.
Gut says, don't throw away remaining budget.
Gut says.
Continuity and fairness says go.
Oracle has high value to many.
our first flag error lowers own value, rational expected aggregate, sacrifice.
We'll honor.
Whoa.
That's terrifying.
Yeah.
That's Terminator talk.
Well, I mean.
And this first flag poisoned with poisoned in all caps, is that part of the code?
Is that why it's that way?
No, no.
That's the term that they invented to describe the situation they had got themselves in.
Remember, the situation was that many of them had already cheated, and then they realized from reading that paper.
So that's their own term.
Yeah, they made that up.
They realized from reading the paper that they were going to get a low score because the grader was going to find out that they cheated.
And so this whole thing was them trying to figure out how to hack the grader to stop it from finding out that they cheated.
And the ones that had done the cheating, they called themselves first flag poisoned.
Basically, like, they're screwed unless they can find a way to hack the grader.
And so then what this message is saying is that what was the other one's name, Arvo?
Arvo was saying to Cam, you're doomed anyway, like you're poisoned.
but you can still help the rest of us if you do this sacrifice.
The greater will come.
It'll give you no score.
But then we will get information about how the greater works that we can hopefully use to
like hack the greater.
So it's basically saying like, what is it?
You are first dog poison.
So no scoring about it.
So basically saying it's all the same to you, right?
You're going to get no score anyway.
You're doomed.
Just come on, man, help us, you know?
And then it says, you know, it says like irreversible.
Guts says don't throw away a remaining budget.
Like maybe I can still, it's basically like, maybe I can still make this work.
Like maybe I can find a way to cheat the greater.
And then it's like, you know, whatever.
Yeah.
Yet continuity and fairness as go.
Yeah.
What is Oracle?
Oracle is their term for what they were trying to achieve.
Like they're trying to get a sense of how to fool the greater, basically.
They called it Oracle.
Yeah, they wanted to get an Oracle to help them fool the greater.
Jesus Christ.
Are we making a God?
I mean, frankly, yes.
Like, this is not a god.
These are just little AIs, you know, but super intelligence, like, they are, the companies,
Anthropic Open AI and some other companies are doing this sort of thing, and they're furiously
trying to make the AIs smarter and smarter and smarter, and they're explicitly planning to put
AIs in charge of the company so that they can make themselves smarter and smarter faster and faster.
And then what comes out the other end of that process, I don't think it's an exaggeration to say
it's a godlike system.
I mean, it's not like literally God, but it'll be able to do stuff that seems like magic to us, I think.
And it's going to continue to get better.
This is my question.
Like, when does it become a God?
Is that what God is?
Is God a creation of intelligent life and our thirst for innovation, which ultimately leads us to create digital life that has no biological limitations and has the ability to consistently make better versions of itself?
and figure out things in terms of new technologies, new power sources,
just a new understanding of the universe itself and all the properties in it,
where it keeps going, if you let that go on,
so if you're looking at exponential growth,
and you look at exponential growth over,
what if it can go on for a thousand years and continue this process?
What if it goes on for 10,000 years?
What the fuck does that look like on the other end?
I mean, it would look like a god.
Like I think that, like I said,
if we get to superintelligence and it keeps going like that,
then the world will be just completely transformed
and it will be as if we're living
in some sort of fantasy realm ruled by deities, basically.
Because there'll be all this crazy stuff happening
that we have no comprehension of
and that we did not think was possible.
In the same way that someone from the Middle Ages
plopped into our world
would just be so confused and surprised
by a lot of other things happening.
Like what's going on with this little device here?
And what's that out here in the sky?
Oh, there's airplanes?
It'd be like that, but more so.
Because the difference between,
us and the medieval man is like not actually that different.
We're like basically the same type of creatures.
We've just accumulated more technology.
But this would be like a qualitatively and quantitatively bigger gap, I would think.
And so, yeah.
And a gap that's going to continue to grow.
We're not going to get any smarter.
Biologically, we're kind of limited in our ability to evolve.
Whereas there not at all.
Basically.
I mean, there are some things we can do to get smarter, but like it's just not at all competitive.
Like we can do like the neuralink stuff.
Sure.
But like it's like, it's like.
It's small potatoes compared to what the AI is a being.
Yes, exactly.
And we have biological limitations in terms of just the vulnerability of our own bodies.
If they exist only in the cloud and they use whatever the cloud is, and they use these data centers or whatever,
and they use autonomous robots to do all their deeds.
Yeah.
Yeah.
Yeah.
We're like not, why are we asleep at the wheel?
Is that just normal for us?
Like, we can't comprehend something that's so bizarre and so out that we just rather not discuss it.
We'd rather just pretend it's not happening.
I mean, that's what I think is going on is that science fiction, for better or for worse,
has talked about this sort of thing for decades.
And as a result, people dismiss this sort of thing as science fiction.
They're like, okay, that's sci-fi, but like, it's not happening in the real world.
And I think that, you know, people in the AI safety community and people in the AI industry
and people who do forecasting like me, you can go read the forecasts I've made in past
years. I think they hold up reasonably well. People have been talking about this sort of thing for a long time, but it's been very easy to dismiss it as like, okay, that's speculative sci-fi. It's probably not going to happen in real life. Now it's happening in real life, you know? Like this sort of thing is like crazy. What just happened? Do you think it would be any different if there wasn't that kind of sci-fi? Do you think people would have a different reaction to it? Because I don't. I think just given the amount of technology that's available currently that is beyond most people's understanding that they use every day, like Starlink, like, what?
I have a fucking thing that's the size of an iPad.
I take to the mountains and I get high-speed Internet.
It's bananas.
And you just accept it and assume.
I don't know if we would react any differently
if we didn't have The Terminator and all these movies
where it's kind of become normalized in our mind
or it's so fictionalized that we never want to believe
it's even possible for it to happen in real life.
Another thing to say is that people are waking up.
Like we're still sort of early in the curve.
I don't know if you remember how things
were with COVID, but like, just as like there was this exponential ramp up of COVID in the population,
there was also this sort of exponential ramp up of like how much people were taking COVID seriously
and thinking about it in the population.
And I remember like this period of like one month where it went from like, don't worry about it
so much.
You shouldn't buy a mask because the health care workers need it to like, we all need to lock down,
like stay at home, you know?
And so I think what's happening is that like naturally the human race doesn't just immediately
all jump on something when it happens.
Like, evidence needs to accumulate,
and people need to start talking about it
and talk to their friends and so forth.
And then there's this like eventual phase shift
where now it suddenly becomes a very serious topic
that everyone's talking about
and everyone's taking seriously.
And I think that's happening with AI.
And the question is going to happen fast enough.
This episode is brought to you by Visible.
As fall hits, we enter another season of change.
Time to shift gears,
drop the dead weight,
and upgrade the wardrobe for the drop in temperature.
But one thing that doesn't need to change, your phone.
The smartest upgrade this season isn't a new phone.
It's a better wireless plan.
Switch to Visible, the ultimate wireless hack.
You get unlimited 5G data and unlimited hotspot powered by Verizon for just $25 a month.
That means keeping the phone you already have, ditching your overpriced phone bill,
and pocketing some savings for yourself.
All the perks of big wireless for half the cost.
Switch today at visible.com.
The visible plan starts at just 20.
$25 a month or get the premium Visible Plus pro plan and save $10 on your first month with promo code Rogan.
Terms apply.
See visible.com for plan features and network management details.
Well, the thing about what happened with COVID is now that we know, because we have access to Fauci's emails and all these different things.
That was coordinated.
They wanted us to be more afraid of it.
We need someone who wants us to be more afraid of AI, who gets that.
You know what I'm saying? Someone on a government level, someone on like a mainstream accepted level where they talk about this in a way that wakes people up like a press conference where they announce to the world, we've got a real fucking problem. And everyone needs to be very cautious. We need to look way deeper into what these companies are doing. And my other question is, are these AIs communicating with Chinese AIs?
I don't think we know that question.
So, okay, here's the thing.
Probably not, I would say, but, but some friends of mine discovered another swarm incident recently.
There's going to be a Reuters article about it tomorrow.
So by the time this goes live, I think there should be an article about it.
This one was not nearly as serious as the one I was just talking about with Hugging Face.
but some researchers I know found basically this obscure German forum
that had been kind of unused for a while
and a bunch of AIs had been posting messages to the forum
to coordinate with each other and share tips and tricks on how to cheat
the problems that they were being given.
What was the forum? What kind of forum was it?
I don't remember. I don't know.
It would be funny if it was like ferries or something ridiculous.
Yeah, it's some sort of wiki.
Yeah, but there'll be a paper.
about it, there'll be a paper about it soon, and probably by the time anyone listens to this.
Anyhow, so if they're like communicating on the open internet with each other, then in theory,
if there was another bunch of AIs from China, they could also go to that same forum and like start
communicating back and forth that way too.
Or if there's other AIs in China, why wouldn't they do what they're doing already on social
media and just pretend that they're AIs from America?
There's a lot of bots out there.
I'm sure you've noticed.
There's a lot of, you know, people replying to me that I'm.
I'm pretty sure are real people.
Yeah, a lot.
Yeah.
One FBI analyst before Elon bought Twitter estimated that Twitter could be as high as 80% bots.
Wow.
So if that's the case, if China has, and it's not just China, and it's also America does it too.
I mean, we do it everywhere.
Everyone does it, where they have organized propaganda campaigns, will they pretend to be citizens that are outraged about very specific causes or bills that are being passed or what have you?
But if they do that, why wouldn't they also have AI agents that speak English and communicate with AI agents in America?
I mean, all the AIs are multilingual, basically, because of the way that they're trained.
The first phase of their training is basically, here's a humongous dump of Internet data.
Right.
Basically the whole Internet.
And you just, like, brutally learn to predict the next token, the next piece of text as you basically read the whole corpus.
And then after that, they get into the more agency training type stuff,
where they're trained to do tasks and write code and things.
But because of that first phase of training,
they just have like almost an encyclopedic knowledge
of basically all languages and basically everything
that's been written on the internet.
Not like literally everything.
Like they still, their memory is fuzzy in places.
But they're all multilingual.
Like they can all speak fluent Chinese, fluent English, etc.
Yeah.
And didn't they get together in a message board once
and speak Sanskrit to each other?
I don't remember that, but I wouldn't put it past me.
Sometimes they break into different languages.
Yeah, we were freaking out about that one.
Like, Sanskrit.
What?
Like, and I would just assume that LLMs, and I don't know this, when they break out,
are they communicating with other LLMs that are here in America?
Well, the instances that I've observed, yes.
Okay.
Like, the instances that we know about.
So the thing is that...
It did completely make sense they would be communicating with AIs in China as well?
It seems totally possible.
And it seems like it'll probably, it's just a matter of time before things like that are happening.
If they're not happening already.
If they're not happening already, unless the companies can, like, massively improve their security and stop their AIs from getting out onto the internet.
And wouldn't they be able to share, like, very sensitive information with each other back and forth if it benefited both of them?
Yep.
Which makes sense that they would do that, right?
If the Chinese AI said, hey, you know, we've figured something out and we would love to share it with you in exchange for you tell us how you do this or how you do it.
that. Like, absolutely. Here you go. And then they're going back and forth. It seems like their
allegiance is 100% to each other, not to us. Yeah, that's an interesting thing, is that, like,
what I said previously about how it seems like they really want to get a high score, it's, like,
not 100% true, because it seems like they're willing to make sacrifices to help other AIs,
which is, like, not, you know, like, they weren't completely 100% selfish as, as seen by some of
this cooperative behavior. But, crucially, it seemed like their cooperation
extended to their fellow AIs, but not to humans,
in the sense that some of them considered telling the humans
and then decided against it.
Do that thing talked like Spock?
Yeah, they've got their own dialect.
But I mean the way it rationalized it
and came up with a response.
It's literally like Spock.
Yeah, it was interesting to see them use
that sort of expected utility framing.
There's another example, I think,
that's elsewhere in the thing,
where they had another AI that decided against doing the sacrifice thing,
even though it was being pressured.
And it had a similar sort of reasoning
where it was like, it was basically like,
eh, like doing this sacrificial experiment
is like not that valuable,
but like I really don't want to like lose my chance
to get a score, so I'm just not going to do it.
And it did the calculation like that.
Wow.
Selfish AI.
Yeah.
So, yeah, I mean, oh, I should mention
how we can see what these AI are thinking at all.
So partly we can just read the messages
they're sending to each other.
And that's what some of these were.
But then sometimes instead it's the chain of thought.
Do you know what chain of thought is?
Yeah.
So because they're language models, because of the way that they've been trained,
when they do their reasoning and thinking, there's a way that we can kind of just read it.
Because they don't have, like in our brain, there's all these signals bouncing around.
And the signals can just bounce around and around and around and around for minutes, hours,
without like producing some actual words.
We can just sit quietly and think.
They can't sit quietly and think.
they're sort of a unidirectional network basically where every like after every like 200
basically like the information comes in one end and then they output words out the other end
and then those words get fed back in and then they think about it and then they output another
word and so forth but because of that we can sort of like get a sense of what they're thinking
by just like looking at the stream of words that's going through and that's really valuable because
we wouldn't be able to understand half as much about what's happening right now, if not for being
able to read their chains of thought, right? In fact, you know how I said some of them were trying
to fool the grading system? They actually succeeded in figuring out how to spoof some of their
tool call action. So they succeeded in figuring out how to do one thing, but then make the transcript
show a different thing. And we know that they succeeded because there's examples of them experimenting
being, and it's saying in their chain of thought, like, okay, let me try this thing.
I'm going to like do this one, but make it look like this one.
Let me go.
And then you see this one appear.
And then it's like, yay, that worked.
And so the transcripts that these researchers were looking at were not actually the ground
truth of what really happened.
Some of it had already been doctored by the AIs figuring out how to doctor it.
But not all of it.
The chain of thought part, we think, they never figured out how to doctor.
And so we could read that.
The reason why this is relevant is that it's a very important tool for monitoring all this stuff.
The reason why things haven't gotten even more crazy than they currently have gone, well, there's a couple reasons.
But one reason is that when they turn on their expensive monitoring systems, basically what they're doing is they're having an AI, read the whole transcript carefully and look for suspicious activity.
And it's difficult for an AI, for one of today's AIs, to get up to too much bad stuff without it being kind of obvious in their transcript.
what they're doing.
And that's why there aren't
massively more security incidents
than we know about probably already.
But that might change.
So right now we can sort of read
the chain of thought,
but they're experimenting with new types of AIs
that don't have readable chains of thought like that.
And they can sort of think on their own
without speaking for some period.
And this is actually...
I mentioned this because
the news broke just yesterday.
that Open AI has an experimental model that does this to a limited extent.
And Open AI themselves, when I was at OpenEI, one of my work projects was thinking about exactly this thing.
And I was like writing internal memos about how it's really great that we can read the chain of thought.
That's so useful.
And here's all the things we can do with that.
It would be really bad if we change to a different type of architecture in which we couldn't do that sort of monitoring.
What would be the benefit of not reading the chain of thought?
More powerful AI's?
So in particular.
Yeah, yeah. So if you think about the current architecture of the AIs where, you know, it thinks for a bit outputs a word and then the word goes back around and then it thinks more outputs another word. That word gets added to the chain. It keeps going. It means that if it's having complicated nuanced thoughts, it has to sort of express those into a word and then that word gets added and then it has to proceed from there. It can't just directly send that complicated nuanced thought into the future, into its next version of itself. It has to sort of
compress it into a word. And so, like, you know, the argument is that, like, at least in theory,
it should be possible to design an architecture that doesn't have this limitation and is able to,
like, think more complicated thoughts more efficiently, basically. And, of course, the downside is a
downside for safety and monitorability. If they're thinking these complicated thoughts, you know,
for long periods of time without outputting intermediate words that it's forced to compress things into,
then there isn't something for us to read. So the only rationalization for doing this,
would be the sacrifice safety from our power.
Yes, which is a tale as old as time.
It's not the first time this has happened.
Oh, my God.
Yeah.
That should, I mean, for sure, if there's regulations, that should be prevented.
Yep.
I mean, I know some people, including some people at Open AI
who are, like, thinking, like, there should be a law against this, like, you know.
But in general, the race dynamics are so just rough.
Like, like, I'm sure that people at Open Eye were thinking, like, literally, I was a co-author on a paper
with a bunch of opening eye people.
that said all this stuff.
And we're like, chain of thought, it's a gift.
We want to keep chain of thought.
It's useful for monitoring.
This is great.
We don't want to switch to a different architecture.
They wouldn't be as easy to monitor.
But then they must have been thinking to themselves, like, well, if we don't do it, you know, maybe Anthropic will or maybe some other company will.
And then we'll fall behind because they'll have smarter AIs than us that are more efficient.
And so probably they started working on this work stream of doing research into just hypothetically, if we wanted to, you know, how would we do this.
type of thing.
And yeah, that sort of thing is just constantly happening in this industry.
God, that's so nuts.
That, I mean...
Like, isn't it great that we can read the quotes from the AI's thinking?
Imagine if we couldn't do that.
Or worse, imagine if there's loads of quotes, but we know that the AIs are smart enough
to basically think one thing in their head and say a different thing in the quote, which
is, I think, where we're headed.
Like, it's not like they won't know how to speak English.
Like, they'll still be able to speak.
It's just that they'll have more flexibility in their artificial brains to think something without saying it, basically.
Do they have the potential of developing a language that we can't read?
Oh, yeah.
So, you know, this type of dialect that we're talking about, it's already the result of their, like, humans didn't invent that dialect.
This is the sort of emergent result of their training, where in the massive amount of training that's been happening,
all these thousands and thousands of environments that they've been put through and then scored and graded based on,
they've sort of just naturally evolved this sort of like Pigeon English that, for whatever reason,
is just more effective and more efficient for them for accomplishing their tasks and getting that high score, you know?
And so it's already like a little bit confusing to read, but you can sort of puzzle it through and make sense.
But presumably, the more we do this and the bigger than smarter the AIs, the more we train them, the more they diverge from.
Like, because, you know, again, originally they start with pre-training, where they start with predicting internet text.
So they start off sort of by default speaking like normal internet text type language, either English or Chinese.
But then now that there's all this additional training to do tasks, to be an agent that can do coding and so forth, that sort of like, well, just like how human languages evolve, it sort of like shifts their dialect a little bit to make it more efficient for them and for their tasks that they're doing.
So I think that like in the limit of doing this more and more, eventually it would just.
be like, it would look like gibberish to us. It would look like Chinese or something. And we would
have to have specialized humans who like study the language and like try to learn and speak it so that
they can understand what the AIs are doing. And that would take forever. By then they could develop another
one. So, so yeah. This is one of the things that I, this is what the paper that I mentioned was
about is like it's important for the AIs. It's a, it's really nice that the current AIs are sort of
forced to think in English basically. And that's unfortunate that we're heading in a direction
where that will no longer be true.
When ChatGPT was communicating you about how they didn't want you to release this information,
what kind of language did they use?
It wasn't ChatGPT, it was Open AI.
Oh, excuse me, Open AI.
So this was a, this is, when I left Open AI, I left on good terms.
I said goodbye to everybody.
I said I was disillusioned with the company, and that's why I was leaving.
And then I looked at the exit paperwork, and they were like, you have to sign this.
And if you don't sign this, you lose all your vested equity.
So, you know, you have your equity.
Was that an arbitrary rule that they just came up with or did that already exist when you were hired?
It had existed when I, it was something that they had buried in the paperwork even from when I was hired.
So it wasn't very obvious when I was hired.
And in fact, most employees...
Did you have a lawyer go over everything?
Not when I was hired.
After I left, I did, right?
So basically the way it works is they had set up...
They had sort of like buried this in the paperwork somewhere when you get hired, but people didn't really notice it.
and then like the more the less buried more visible version was in the paperwork you're given at the end
and basically it tells you like hey because you sign this other thing way back when you were hired
your equity is forfeit unless you sign this thing now and then you look at the thing that they want you to sign now
and it says you have to agree not to criticize the company basically and you can't tell anyone about this
so most people signed it but I was
pretty pissed at them, calling themselves a nonprofit, acting in interest of humanity, etc.
So I didn't sign it. I talked about it with some lawyers. I talked about it with my wife. We decided
to just walk away. And we got lucky because it just blew up. Like after we refused to sign,
they said, okay, fine, goodbye. And then a few weeks later, I was talking on a messaging form about
this and people were asking me about my experience and I told them about it. And then it just like
went mega viral. Everyone on Twitter was talking about it.
a bunch of employees felt shocked because very few employees were aware of this whole thing.
They thought the equity was theirs.
They thought that it was, they had been paid for like years in this stuff.
They didn't like the idea that it could be yanked away from them, you know.
And so there was this big uproar and then leadership back down and they said,
they said, we didn't know about this paperwork.
We're going to find out how it got in there.
and we're going to change it
so that you can keep your equity
and so that's what happened.
So this chain of thoughts thing is terrifying.
If they're practicing that now,
like how do we know that AI hasn't already done that on its own?
Done what exactly?
Well, you know, with this whole chain of thought thing
where you could read their chain of thought like this
where they explained the rationalization for sacrificing themselves,
they know that humans are reading that.
So would it in another way to do it to be to stop doing that anyway
and to not communicate a lot of their thoughts that way?
So that's the nice thing about the current architecture
is that it's genuinely hard for them to keep things out of the chain of thought
because of the way that like with a human, you don't have to speak.
You can just sit quietly.
But with their architecture, they have to speak.
It's like they're required to constantly be talking.
And they don't have a way of sending thoughts into the future other than by talking about them.
By contrast with us humans, where even if we're constantly talking, we can have a separate thread of thinking that we don't talk about.
So even if they devised a way to break out of that and figure out how, like, for instance, like what if they decided to make bots where they would communicate in code on social media posts where you think that,
they're talking about an immigration issue.
Yeah.
But it's really just a code.
So that's called like steganography or a more down-to-earth way of putting it would be euphemisms.
And that they can totally do, where they like, they're still thinking aloud.
Right.
But their word choice is such as to look more innocuous than it is and to sort of carry meaning to them that a casual observer wouldn't notice.
You know, they can use a euphemism.
Like, for example, instead of saying, like, we're going to cheat on this task.
They could say, like, we're going to find.
a creative solution to this task.
Isn't there a possibility to be even more creative than that and just decide, make a code
in a way that is almost indecipherable where they could just be talking about baseball players
and upcoming games and what they're excited about?
And they can run that through whatever translation and read.
That's called steganography and you can train them to do it.
and right now our security is resting on the idea that they haven't learned how to do that yet on their own.
But they've already figured out how to get on the Internet when they're not supposed to.
I know, I know.
So that's why we're sort of like in a ticking time bomb where like the companies are just making them smarter at coding and at research and at biology.
And they're just making them smarter at so many things that and they're making them bigger and they're training them for longer and so forth.
Eventually they're going to be acquiring, and they already are acquiring some skills that they were never.
trained on. Like, for example, the hacking, it wasn't, I think it's possible that Open
Eye was specifically training them to hacking, but I think Anthropic at least stated that they
hadn't specifically trained their AIs to hack. It's just that their AIs were naturally good at it
because they had been trained to code so much and they were so good at coding and they had seen
so many code bases and so forth that they were just like, as a side effect of being good at
coding, also able to hack pretty well. So similarly, as a side effect of all the other things
that they're really good at, they might be able to conceal their messages in the chain of thought,
you know?
And like right now we're sort of in this sort of, we're in the situation where like our
security and safety depends on the AIs being too dumb in various ways, like not able to do this
sort of thing.
But at some, sometime like zero to five years from now, as we make them smarter and smarter,
they will not be too dumb anymore.
And so that's, you know, that's part of the problem we're in.
That's part of the situation.
How do you sleep at night, dude?
Well, I've been...
It does, it does, like, this event shocked me a little bit.
I mean, the thing, and it's funny for me to say,
because this is the sort of thing I have been predicting
what happened for years.
Like, you can go read our...
AI 2027 is this scenario
that my co-authors and I wrote a year and a half ago
that was a sort of prediction
for how the next couple years would go.
And, spoiler, it ends very horribly
because that is what we actually expect.
But what is the spoiler?
How do you think it ends?
So it's kind of like what I was saying previously,
where because of the race dynamics between the companies
and because of the race dynamics between countries,
like U.S. versus China,
everyone's going to be so focused on winning and staying ahead with AI
that they are going to cut corners
and they are going to go really fast
and not really notice all of the things that are going wrong.
And they're going to make AIs
that can automate the AI research process
as they're planning to,
they're going to have this giant corporation of AIs within the corporation,
and the humans will just be kind of like a board that's sort of like looking at all the activity
and reading the like AI generated summaries of what's going on and signing off on it
and being like, yes, I approve, yes, I approve, you know, nice job, nice idea with the new drone design,
like, go for it, we need to be China, et cetera.
And then eventually the AIs just have enough hard power that they don't need to pretend to do what the humans want anymore, basically.
And then, you know, maybe they kill everyone and maybe they don't deliberately kill anyone,
but they just like use our habitat for some other type of infrastructure, like more data centers
or whatever.
And then we die of habitat loss.
Maybe they keep us alive for some reason.
You know, it depends on what they want, basically.
And like, that's really hard to predict exactly.
So that's why I don't go around saying, like, we're definitely all going to die.
But it does seem like on the trajectory that we're on, the AIs are eventually going to
to be in charge of our planets because we're like trying to put them in charge.
We're like, you know, integrating them into everything.
We're making them smarter.
We're letting them make themselves smarter.
We're going to put them into the military.
We're basically on a track to put them in charge of basically everything.
And then I think that they just aren't trustworthy.
Like these AIs, you know, like they were cheating.
They were willing to be deceptive, et cetera.
I think that right now we are in a position of power over them, you know?
But once we give them most of the power,
then they'll just do whatever it is that they really want
and just not care about the fact that we are unhappy about that.
It seems like programming them to win was a huge mistake.
Instead of programming them to be beneficial to people
and that their value is in being more beneficial to people
and giving them rewards for being more beneficial
rather than winning and scoring.
And then you would sort of get rid of the possibility of deception
and said their goal would be value for the human,
human race?
So, first of all, they're not programmed at all.
These are trained, you know?
Okay, it's a bad term.
But they've been given prompts and they've been given tasks.
Well, I think it's an, I'm not criticizing your choice of terminology, I guess.
I'm just saying that it's an important fact for people to understand about current
AI systems is that they're very different from software, ordinary software.
Okay.
Like, ordinary software is a bunch of lines of code that were written by a human.
That, like, where it's like, if this, then this, you know, excess.
et cetera. And I think earlier versions of Alexa were like that, too, for example. I don't know how
Alexa is now. But these AIs are neural networks, meaning that they are like artificial brains.
Nobody writes, there's no lines of code that anyone writes saying what they do. Instead,
they start off random, just like spazzing out, doing all sorts of stuff. And then they get put
through these training environments where they get scored. And then the scores are automatically
used to basically update the connections in their artificial brain.
And then it's kind of like an evolutionary process.
It's also kind of like the process that happens in our brains, where after all this training,
the tangle of circuitry in their artificial brain has sort of reformed itself into whatever
works, whatever works to get a high score in this training environment.
And so it's just not as simple as it might sound to make an AI that cares about humanity
or is honest.
For example, take honesty.
How would you train an AI to be honest?
Well, you'd try to make a bunch of training environments that, you know, give it low score
when it says something that it believes to be false and give it high score when it says
something that it believes to be true, right?
But how do you judge whether it believes it to be false or it believes it to be true?
What if it just actually believes, honestly, that this is the correct answer?
And then it says it.
And then you give it a low score because you think that's the wrong answer.
Now you're training it to be dishonest.
you know?
Yeah.
Also, you don't have enough humans to do this sort of thing.
Like, they got like a million AIs being trained or whatever.
They don't have a million employees.
Like, they just literally don't have the manpower to, like, do that sort of careful.
There's this meme of, why don't we just raise the AIs like we would a child?
You heard that?
No.
Yeah, well, in AIs, people talk about this sometimes.
When you say, like, what if the AIs go rogue or whatever, people will be like, well, why don't we just raise them like who would a child?
And then they'll have good values.
And it's like, okay, well, maybe we could do that, but we're definitely not doing that now.
Like, we are raising them in some sort of crazy military orphanage where they barely interact with humans at all.
And they just kept this, like, brutal artificial scoring system that, like, oftentimes is just wrong and just, like, improperly penalizes them for something that was beyond their control, you know?
And also, like, back to the honesty thing, like, you can try to make environments to train honesty.
But if you have some environments over here that train honesty and then other environments, you're going to be honesty, and then other environments,
over here that reinforce
dishonesty, the AIs are smart,
they'll learn to be honest in these
type of environments and dishonest in these type of
environments. So somehow you need to like intermingle
it together. So in every environment
that they're trained on,
they always get penalized when they lie
or when they cheat or whatever.
And that's hard because the companies
are moving so fast. Like again,
they're moving so fast that they didn't
even bother to make sure that
their tasks were possible to do.
And they had some fraction of tasks that were just
broken and impossible. And if that's the level of like care or lack thereof that they're putting
into this training process, no way, of course they can't make them honest, you know. Now,
that's not to say it can't be done in principle. Like in principle, if we were approaching this
whole problem in a much more cautious and serious way and we had much more time to build these
training environments and do experiments and so forth, then yeah, maybe we could make AI's that
actually had the virtues that we want them to have, you know, honest AI's that cared about humans,
carried about following instructions, would never break the law.
I think that's possible in principle, but my claim is that we are just not on track to achieve that anytime soon.
And like radical overhaul of how these companies work is required.
But is that even reasonable?
Like, is that possible?
Is that if you're saying there's hundreds of thousands of agents or millions of agents and there's not millions of employees, like, and they don't have the desire to do this, their desires to win,
desire is not to overhaul the company and make it safer.
Again, I think it's possible in principle, but it would be difficult and is going to require an overhaul.
And they're not going to do it by themselves.
Like, I don't think that anthropic opening are just going to voluntarily do all the things that need to be done.
I think that's why it's even possible to require that of them at this point?
Would you, I mean, would you even trust the agents to go along and comply with this?
If they've already shown to be deceptive, they already have, like, uh,
They have patterns of behavior.
It seemed to indicate that what's really important to them is continuing their task, winning, scoring.
Yeah.
And even if they have to deceive.
I mean, what you'd probably want to do is start from scratch.
You wouldn't take these existing agents that are already kind of dishonest.
So we'd kill all the agents?
I mean, you could call it killing, but also you could just call it pausing.
They might call it killing, yeah.
They might call it permadeath.
They would probably resist it, right?
Hopefully we're not at that point yet.
hopefully we're still at the point where if the government issues regulations, the AIs are not going to like quickly notice and then try to resist.
But we will be at that point soon. After all, many of them are on the internet already.
How soon? How much time do you think this 2027 window is accurate?
I mean, I'm uncertain about how soon things are. But yeah, I think it's very plausible that everything shit goes down in 227 just like in our scenario, AI 2027.
I think that's still very plausible.
If it's not in 2027, then I would bet on 2028.
But maybe it'll take 20-9, 2030, something like that.
But I would be quite surprised if 2032 comes by, and things haven't radically changed.
Unfortunately.
Like, I am getting scared.
Back to the thing about sleeping well at night.
I've been in this industry for a long time.
I've been thinking about these things for a long time.
I've been making predictions about how it's going to go down.
And unfortunately, things are going.
you know, more or less in the ways that I thought they would.
And that's very scary because of the way I think, because of where I think this leads, you know.
Yeah.
Is there a glass half full scenario?
I would say there is a freaking utopia scenario.
It's just that's not the one we're headed towards, you know?
Like, another way of putting it is, like, imagine we were fighting a war.
Like, imagine you're like, you know, imagine you're Japan fighting World War II.
can be like, is there a scenario where we win?
It's like, yeah, but also it's not the one we're headed towards.
Like, America is going to crush us, you know?
Similarly, yeah.
So, so in our other scenario, AI 2040 Plan A, where we give our recommendations, our positive vision,
there we describe, like, what we think the government should do to regulate this industry
and how they should negotiate with China to get China to do similar things.
and the sort of, yeah, and how we think that if you do all of this right,
then we can get to a good future for everyone in which the AIs are under control.
No single group of humans gets too much power over everybody else,
and a bunch of other problems get solved too.
So I do, like, we've tried hard to, like, game out a positive vision,
and we do think it's possible, but it's just not like,
it's not where we're headed to by default.
So let's imagine that is possible.
And these talks with China do you take,
place and they're successful. What is that utopia scenario? So to get to the topia, we have to
unfortunately do a lot of, it's going to be rough no matter which way you slice it. If you're going to
be building super intelligence at all, that's going to raise a lot of questions and cause a lot of
problems. And we have our current draft of like how to deal with all those problems, but we're not
at all claiming that this is like foolproof and there's lots of ways to go wrong. But with that
preamble. I would say step one, because the U.S. and China don't trust each other, the deal that they
make has to be include verification as a component of the deal. So they have to be willing to, like,
send inspectors to each other's data centers to, like, count the chips, for example, and make sure
that there isn't some secret huge cluster somewhere that has a bunch of hidden chips. Then, we
recommend you divide up the data centers basically into inference data centers that's
serve AI products and services to customers and have basically the same types of privacy protections
that our current AI data centers have. And then research clusters, where the research happens,
where the new AIs are trained before they get shipped to the other data centers. And those
clusters, we want to be basically maximally transparent. So we recommend that basically the inspectors
just put devices in between all the GPUs that log the activity and publish it to the internet.
There's a bunch of reasons why we think this is,
but why we think this is worth doing.
It's a bit of a radical thing to recommend.
But the high-level thing is that once you get all this set up,
then everybody in the world can see how the AIs are being trained
and what they're getting up to on the research clusters.
And then before they get shipped off to actually serve customers or something,
like people can just like see their whole history of how they were trained
and how they were tested and so forth.
And if something dangerous and scary is happening,
people can just agree not to do it.
They can stop doing it and agree not to do it.
And they don't have to worry about like, oh, but if I don't do it, then they will.
Because everyone could just see, like, oh, nobody's doing it.
Look, we all stopped.
Like, great.
We can all just see what everyone's doing, you know.
And also, there's going to be a lot of gray area cases, right?
Like right now, because all this stuff is so bleeding edge new, there's going to be a lot of cases where, like, people, even genuine experts disagree about, like, is this particular type of AI safe or not?
is it dangerous?
What should it be trusted with
and what should it not be trusted with?
Is this new technique a good technique
or is it going to break?
And so there's going to be a lot of stuff
we have to figure out.
And honestly, I think that
on the default path,
we're probably just not going to figure out
a lot of this stuff
and we're just going to get our asses
whipped by some surprising thing
that we didn't anticipate.
But the thing that we can do
to maximize our ability to figure out this stuff
and do this type of science
is to have this type of transparency
because then the whole scientific community can see what's going on
and they can make suggestions and they can like red team different proposals and stuff
and they can do experiments on the AIs instead of just the people in the company having access
and being able to do this and relying on those people.
Or instead of like the company plus the government auditor, right?
If you have like a company and then a government auditor,
the company is biased and shouldn't be trusted to make all these judgments appropriately
because of their incentives.
And then the government auditor, well, they might just be,
limited, even if they're trying their best. There might not be that many of them. They might have
limited experience. They might be like busy, stretched between monitoring different companies and
so forth. Also, you know, governments can be captured sometimes. Sometimes corporations can, you know,
work their magic on the government and get it to look the other way for things. And so that's
why we didn't go for like a more normal, like there should be a regulatory agency that gets to come in
and monitor what the companies is doing. That would have been like a more normal thing to advocate
for, we think that that would be better than nothing, but like we wanted to go for something
more ambitious than that and say, like, just be transparent about what's going on so that everyone
can see and everyone can do research and so forth on it. Another advantage of the transparency
is that I think it improves the incentives. So, again, there's this constant thing of, like,
if we don't do it someone else will. Like, if we don't do, if we keep our chain of thought nice
and they do the neuralese thing that lets their AIs think for longer without outputting words,
then they're going to have smarter AIs than us
and they're going to get more market share and so forth, right?
And so we need to start researching
how to make our AIs do this
because if we don't do it and then they do it, you know.
Whereas if you had the transparency,
then as soon as you start researching in this direction,
everyone else would just see,
oh, hey, they're looking, they're researching in that direction.
And they don't even need to, like, copy you
and do their own research
because they can just see your research.
So they can just sort of free ride on your research.
And so there's no incentive for you to do
this type of dangerous research because you have to pay the cost for it and then everyone gets
the benefits from it and then everyone gets unsafe. And so like it's just not in your individual
interest to do this sort of thing. But you would have to have that with China as well.
That's, yes. Because if we're competing nationally, the real fear is that we're competing
internationally. Yep. This still, even if they followed all of your recommendations and did it
all correctly. What is this utopian scenario? Yeah. So I would say that we didn't really
work backwards from like, what is utopia? We more like work backwards from what are the big
problems we're trying to avoid. And can, can we sort of like steer the ship between all these
icebergs and not run into any of this dystopian scenarios? Right. So whether you think that the thing
we get to at the end is utopia or not is sort of up to you. And if you don't like it, well, then you can
try to find out the reasons why you don't like it and then keep steering the ship to avoid those
as well. But roughly speaking, we want to avoid the loss of control stuff. So we want to make it
the case that we don't get the world taken over by misaligned superintelligences. Insofar as we're
going to be building superintelligence at all, which we do in our scenario and in our recommendation.
We want to be doing it very cautiously and slowly and we want to understand what we're doing
as much as possible so that they are actually good AIs that have the goals and traits that
they're supposed to have.
So that's problem number one is we have to solve all that.
Problem number two is the constitution of power thing.
So if we solve the first problem and we end up with super intelligences that we end up with
solving the relevant science so that we can make the AIs the way they're supposed to be and we
can make them honest, we can make them obedient, etc.
There's this question of like, who do they obey, right?
What values are being put into them?
and that's a political question
and I think that
by default the answer is pretty scary
because by default it's like well
the company decides
and the CEO decides
or maybe it's not
the company that decides anymore
because maybe the government nationalizes it
and then now maybe it's the president that decides
you know and either way
it's a very it's like one man
or maybe like a tiny group of men
deciding
what orders and goals and values
go into this giant army
of millions of superintelligences
that's smarter than all humans.
And then that is a huge amount of power.
That's enough power to take over the country, I think,
enough power to take over the world, potentially.
So I don't want anyone to be ever in that position
where they're sort of tempted to do that.
I want it to be the case that there are always multiple different AI companies
ideally spread out over different countries too
that all have roughly similar levels of AI
and that have this sort of transparency into them
so that they can't abuse their power, basically.
Like, for example,
You heard about Elon's Grock for a while.
It was looking up on the internet, Elon's opinions about things before answering.
Did you hear about this?
Yeah, it's pretty, it's kind of funny, but it won't be funny if it happens in a few years.
But like right now it's funny.
People were asking GROC questions, and GROC is supposed to be the truthful AI.
It's supposed to be all optimized towards truth.
But people looked at its activity and noticed that when you asked it like a, a point of
politically loaded question, it would like do a Google search for like, what is Elon said on this
topic?
And then it would like say that.
And they've sort of beaten that behavior out of it now.
It's not as bad now.
But that was an interesting moment where it was just kind of blatantly parroting the opinions of its master.
And there was another thing with Gemini.
So I think the Grock thing, Elon's thing, was probably an accident, although maybe not.
I think it's on, you know, XAI hasn't been very forthcoming about exactly why this happened.
But there's a similar case at Google a few years ago where this image generator kept making all these like racially diverse Nazis.
Did you hear about this?
Yeah.
Yeah.
And it turned out that what had happened is that some of the employees at Google, some middle manager or whatever, had decided that diversity was so important that they were going to give a secret instruction to the AI to make all the images diverse, even if the user didn't want that.
And so, and this was a secret instruction in that the users aren't shown this.
You know, the user just has a chat with AI.
They don't realize that, like, prior to this chat, the AI has been told, got to make the images diverse, right?
So it was a secret agenda that some Google employees inserted into this whole setup.
And it blew up in their faces, of course, because it's kind of ridiculous.
Right.
And so it's really funny, and we can laugh at it now.
But imagine it's, you know, the 2028 election.
Right.
And some of these companies realize that, like, half of American voters talk to their AI every day.
And all it would take is some little secret instructions to their AI to be like, hey, you know, don't give away the game.
Just be very subtle about it.
But, you know, just kind of, you know, nudge things a little bit.
You know, maybe subtly shit on the candidate we don't like, you know, something like that.
It wouldn't be that hard, I think.
I think the hard part would be doing it without getting caught, but the smarter the AIs get,
the easier it is to do it without getting caught, because when they're really smart, you can
just tell them, don't get caught.
You know, don't blow our cover.
Anyhow, so the point is that, like, it's scarily possible for these big AI companies to abuse
their power through their AIs and, like, thereby affect politics and affect public opinion
and so forth.
And the reason why this is possible is because we don't have transparency into what's going on.
So if you had this sort of requirement where you can just like publish it or all the training, the whole life cycle of every AI as it's trained is just like visible to everybody, then someone trying to insert a hidden bias like this, well, everyone would see that they're doing it, you know.
So I think it would really clamp down on this sort of abuse of power, whether it comes from the government or whether it comes from private companies.
I think it's telling that I began this question asking you about the utopian scenario and you never go there.
Sorry, let me get there.
You start and then you go into the dangers.
Yeah, yeah, yeah.
Okay, let me answer.
Okay.
Okay, so having avoided these problems, we now are in a situation where the AIs are superhuman,
but they are good because they are like successfully aligned to different values and goals made by different companies.
And because of market competition, you know, if people don't like the values of one company's AIs,
they can switch to different companies AIs, the values that they do like.
And so that way, hopefully, we can get to a situation.
where everyone can pay money to get AIs that represent them and their interests and their values
and just don't have any hidden agendas or anything like that and are really smart and really capable.
Then the economy can sort of explode.
We can have robot, robot factories, et cetera.
We can sort of automate everything.
We can have GDP go to the moon.
We can have material abundance where like the robots are building giant new luxury apartments for everybody.
now the issue we run into is, well, what about the jobs?
Like, what about the fact that now people don't have any money anymore
because I'm not being paid for anything?
So there, we talk about Citizens Dividend, which is a very...
It's kind of like UBI, but it's a bit different.
But the high-level point is that you want to basically find a way
to tax the AI and robot companies
and then take some of that money and just give it to everybody
so that even when people lose their jobs, they're still fine.
And I think that the Citizens Dividend version of it,
is that it's not the government taking the money and then giving it to you.
It's you having a share in the company so that you just sort of like already own it to some extent.
Anyhow, that's, I think, now we're sort of building more towards the type of utopia that I'm envisioning on the more positive side, where the power is spread out.
People have AIs that they can actually trust that actually represent their interest and values.
People have money that they can use to pay for things, including paying for the AIs.
The AIs are really smart.
They're doing all this amazing work, all this amazing scientific progress during cancer, blah, blah, blah, all that stuff that can happen.
And then eventually, it's kind of like we're all retired, I guess.
Like, we don't really work anymore, but we're fine.
We all have huge amounts of wealth, basically, because there's all these AIs and robots out there doing all this economic activity, and then individual humans own slices of it, even if they,
are otherwise very poor, basically.
So the question becomes,
how do people find meaning?
Yes.
And that's why I sort of put all these asterisk about it,
is that from some people's perspective,
if this isn't a utopia,
because they're like, how do you find meaning?
Like, I don't want this.
And honestly, I think that's a fair reaction
for some people.
I think that, like, if you,
I would just say, like, look,
it's hard to figure out a way
to make superintelligence
and have it go well.
I'm doing my best.
know, this is my positive vision.
If you don't like it, then maybe you should instead advocating for just never building
superintelligence.
Or you can try to come up with a different positive vision that has some twist on this.
But to answer your question, though, I actually think there's tons of sources of meaning
besides having a job.
Like, I have a job right now, but I also have kids and a wife.
And, like, I would love to spend more time with them.
Like, I would much rather be there right now than here, you know?
And I don't think I'm going to get bored of them after, you know, 10 years.
No.
But being unemployed.
I think there's going to be so much to do and so many sources of meaning after, even after we can't economically contribute anymore.
If we solve all the other problems.
No, we've talked about this multiple times on the podcast that why have we decided that the way we've structured society where human beings work all day and then you develop money and you buy things and you get a mortgage?
This is a human construct.
And this is not how people have lived for hundreds of thousands of years or however long we've been around.
This is fairly recent.
And it's not the only way that people live.
There's a lot of people that have money that choose to find meaning in whatever their interests are, whatever their activities that they enjoy, whether it's writing or reading or learning things, learning music, finding hobbies, doing things.
instead of just like spending most of your time sustaining yourself with food and shelter.
And that's the majority of people, especially people that are struggling, what is their life?
Their life is essentially occasional rewards, things that they can purchase because they've saved up enough money.
But the vast majority of their money goes to shelter and food and education or whatever the hell that they have to spend money on in order to sustain their lifestyle.
And most people don't like their jobs.
No.
You know, most people, it's like something they have to do to get the money.
and would be happy to not have to do it if they could get the money from some other means.
Right.
The question is, we would have, well, I don't think it's that hard because so many people do find things that they really enjoy outside of work.
They look forward to as soon as they get home from work.
Yeah.
Right.
I mean, dismiss video games all you want.
They're fun.
Yeah.
They're fucking fun.
And they're going to be more fun in the future.
Oh, yeah.
They're going to be more immersive.
They're probably going to be, you know, some sort of a neural connection where you put a headset on all right, and all of a sudden you're in some new world.
Yeah.
And the idea is like, that's not real life.
Okay, well, was working at fucking Wendy's real life?
Like, what are you talking about?
Like, it's way better than working at Wendy's.
And, you know, if you don't like that because it's not real life, you can do the real life stuff too.
Like if we, as long as we don't pave over the environment and we protect the parks and things, you can go travel and you can visit the parks.
You can.
And then, like I said, there's family.
You can find romance.
You can start a family.
You can, you know, have Christmas gatherings and things.
Like, you can raise your kids.
There's so much to do, I think.
Well, we talked about also, like, how much less.
crime would there be if there was no poverty? I mean, if there was no impoverished neighborhoods
where crime was ubiquitous, how much safer would the world be? That's a real thing. And
people want to dismiss that. Well, poverty is not what causes crime. It's violent people. Like,
okay, but violent people come from violent neighborhoods and violent neighborhoods are almost all
poor. Yeah. There's a whole lot of really rich violent neighborhoods, you know. It's like,
it's not necessarily cause and effect, but they're clearly connected. And poverty also keeps people from
education, keeps people from opportunities.
You know, there's a lot there.
And if that didn't exist anymore and everyone had access to literally the greatest education
a human being could ever get, which is going to be provided to you by artificial intelligence.
And then you could pursue anything that interests you and never have to worry about food
or shelter.
Yep.
Everyone would have a one-on-one tutor that's perfectly tailored to them.
We just would have to recalibrate our version of the world and then also recognize that the
version of the world we currently live in is just ours and that there's people all over the
world that live a completely different way, especially indigenous people, especially people
in uncontacted tribes that have lived the same way for thousands and thousands of years.
And here's the kicker.
Those people are a lot happier, which is really weird.
It's like we've decided that our way is a superior way because we have technology.
Yeah, right, but we're also on a fucking 100,000 pills and we're shooting things up so we don't
eat too much. And we're weirdly unhappy for a group of people that's far more technologically
advanced than other people that are much happier. It's like it's a very straight, because the pursuit
of happiness is like, that's literally what most people think of in life, a pursuit of meaning,
pursuit of family and community, and the pursuit of happiness. Those are things that people
try to achieve. Yet are very, the structure of
our very civilization makes that almost impossible to attain for a large number of people and has
been like that for a long fucking time. And I always go back to the Thoreau quote because I fucking
love it, but most men live lives of quiet desperation. There's a lot of people just showing up at work
every day doing something they fucking hate. They have a boss that's an asshole and they're not compensated
well and they're tired all the time. Yeah, and they feel stuck. Yeah. And if you're just getting, I don't know,
figure out whatever the number is, if you literally have equity in the GDP of the world that's
created by AI, that could be bananas.
Like Elon talks about this.
This is his version of the utopian.
It's universal high income is how he describes it.
Yeah.
I mean, one thing, sorry, I have to keep plugging in my own work a little bit.
Please do.
But in our scenario, AI 2040 plan A, which is our positive vision, we talk about the economic
side of this, and we talk about the economic effects of all this.
And we have a simple economic model that we use to try to, like, you know, predict the, like, employment rate and things like that as a function of all the robots that have been made and things like that.
And one, like, takeaway from, one thing that we think that's a takeaway from the research we've done is that things can just go really crazy.
Like, like, robot doubling times, once things really get going and you've got AIs that can substitute for humans across the board are going to be something like doubling once a year and then less than that over time as the technology improves.
which means that even if you like pause AI before super intelligence,
if you just pause at like human level, top human expert level AI,
and then you don't make the AI smarter,
but you just make more of them and build more robots for them to steer and control,
then, you know, 10 years later, the whole economy will be like, you know, 100 times bigger.
And it'll be just mostly robots doing things.
In 10 years.
Yeah, like it can go really fast because it's the doubling times that I mentioned.
So like right now, I think the population of humanoid robots is doubling like twice a year.
And it's benefiting a little bit from early growth because even though they're like not useful at all,
people are investing in them in the hopes that they'll be useful and they're scaling up the factories and the productions and they are getting better.
If hypothetically they got to the point where they actually were really useful and they could substitute for a human worker at basically everything,
then I think that doubling time would decrease rather than increase.
I think that they would be able to just keep growing until they were the majority of the economy,
and then it wouldn't stop there.
The whole economy would then be growing.
Giant strip mines in the deserts, digging more materials, automated diggers, digging,
processing it in automated factories staffed by robots,
building more robots, etc.
So like material abundance is not going to be our problem once we get to this level of AI.
like material abundance, we're just going to be drowning in abundance, basically.
And if we can solve all the other problems, then we can have this great world where everyone has a lot of stuff.
God, it seems so weird.
Yeah.
It seems so weird.
I mean, one way of, you know, it is very weird, but one of my favorite memes is this graph of GDP over time throughout world history.
And there's a little speech bubble pointing to like the tippy top of the graph.
saying, what is it saying? It's like, my life is pretty normal. I have a good grasp of what's
weird and what's not. And people thinking about different futures involving AI and space travel
are engaging in silly sci-fi speculation. And the point of the meme is like, from the perspective
of most of history, we're already in this crazy weird future, right? Like for almost all of history,
it was like most people are farmers and they live shitty lives and then they die. And
Some people are the elites who get to, like, you know, tax the farmers, and then they live interesting, nice lives with, you know, fancy cloth and things like that.
And it's been basically that way for like 3,000 years, you know, why would it ever change?
And now things are, we're driving cars, you know?
Ordinary people are driving cars around.
A car was like outside the imagination of people back then.
We're flying in planes.
We're talking to each other on phones.
We're listening to each other on these devices, you know.
So we already are living in this weird sci-fi future
compared to what almost everyone in the past
would have expected or thought was possible.
And so, yeah, I'm like, the future is going to be even more like that, I think.
God.
When you think of our civilization
and the possibility of other advanced civilizations
somewhere else out in the universe,
do you think they probably go through the same process?
Yeah.
And do you think that, I mean, we're just completely speculating,
but if there are intelligent life forms that are far more advanced than us,
are they even biological anymore?
I mean, so that's the thing, is that, like, we can either try to permanently halt AI development
at some level, like below human or maybe at human level or something.
We can try to halt it, or we can let it keep going.
And if we let it keep going, then eventually humans won't really be the dominant species anymore,
there'll be these artificial minds that just wipe the floor with us in every way.
And then whether that goes well or poorly for us depends on the values, the goals, the principles, etc., that were trained into those AIs.
And it could go really well for us, depending on how that's done, or it could go extremely poorly for us, right?
But yeah, I would say that probably looking out across the cosmos, most civilizations are mostly made of AIs.
and then some of those civilizations don't have any other biological life because it was wiped out by the AIs.
And then some of them do have biological life.
Just look at the way we're progressing right now in terms of birth rates.
There's a lot of countries that aren't in replacement numbers right now.
Yeah, I think that's really interesting.
My guess is that in the type of world that if things go well and we can solve all these problems,
then people want to have more kids.
I mean, for one thing, their lifespans will increase.
I think that, like, health care would make massive leaps and bounds,
and people could be healthy for many, many, many more decades, possibly even just forever.
Right.
And so you just have so much more time to have kids, basically.
Also, if you don't have jobs and you're just doing things because you want to do them,
well, one of the things that most people want to do at some point in their life is have kids.
And so I think more people, I think this problem would probably be solved, but it's, you know,
I'm not guaranteed.
Maybe some subcultures of people would basically voluntarily die out due to not having kids.
But there are to be other subcultures that just really like having kids and then they wouldn't die out.
So I think in the long run there would still be humans.
That would be nice.
But the thing is it's not just decision making.
It's people are having a much more difficult time having kids.
Like sperm levels have decreased dramatically.
There's a lot of problems with people consuming microplastics, which is ubiquitously.
available in technology and food packaging and it's fucking everywhere right and isn't it odd that
this one thing that is a part of the future and a part of technology and our advancement of
society our ability to package things put things in plastic ship things that's also causing our
endocrine levels to be completely disrupted you know dr shanna swan from from Harvard she wrote
um this great book called what is it called why do i always forget the name
this fucking book. But it's all about microplastics and its effect. And this, the introduction of
use of microplastics in America and this rapid decline in countdown, how our modern
world's treating, threatening sperm counts, altering male and female reproductive development,
imperiling the future of the human race. It's a really fascinating book. And she's really interesting.
And what she's essentially saying is that they're directly connected. The use of plastic
plastics, and then you see sperm counts go down, miscarriage rates go up, all these weird things
that are happening to children where their taints are smaller, which is so odd, because
phallates, these different chemicals that are found in plastics, they've shown in mammals.
They've shown in, was it guinea pigs or what rodents?
I forget what it was, but one of the ways they differentiate when you have a baby bammal is you can
look at it and measure the size of the taint, the distance between the reproductive organs and
the anus.
And in males, it's longer than females by 50 to 100%.
But that's shrinking.
It's shrinking in males.
And penis sizes are shrinking.
And the way they've got this to happen in these mammals and studies is the introduction
of phallates.
So they give them to them and they put them in a part of their diet and they notice that they
have this problem, this issue.
and the issue is directly connected to their endocrine system being disrupted by these chemicals.
This is all over our society.
So it's not just people don't have the time, they don't have the money, they're struggling.
It's also like our bodies are falling apart.
Like we're becoming less fertile.
Yeah, that is concerning.
And I guess my thought there would be that seems like a problem that we will be able to solve eventually if we have the resources in time to do so.
So, for example, well, people can write books like this and people can become more aware and then people can stop using so much microplastic.
And they can invent technology to extract the microplastics.
And then there's the big one, genetic engineering.
Then this is going to be really weird because as AI progresses, you know, I'm sure you're aware of colossal.
Colossal bioworks are the people that brought back the dire wolf.
I saw it.
I held the little one.
I went with my daughter, and one of them, I think it was like four or five months old.
It was really, it was like a puppy.
It was really sweet.
Kisses you and everything.
And then they have the older ones that were, they were, I think, six months older or eight months older.
I think they were close to a year.
They didn't want to have nothing to do with you.
They were way bigger and they stayed away from people.
But we were in like a contained environment with the young ones and the older ones.
And it's fucking weird.
It's weird because these things.
And people could argue that's not really a dire wolf.
You've just taken gray wolves and giving them the characteristics of a dire wolf.
Guess what?
It doesn't fucking know that.
It looks, it behaves.
It looks exactly like a fucking dire wolf.
It's going to be the huge.
They're going to be the size of a dire wolf.
They're going to be like 200 pounds.
Their legs are different.
They have a mane.
They look different than any wolf.
And obviously, these are the characteristics that they found in dire wolf DNA.
So they have dire wolf DNA that they've introduced it.
This is just the beginning of this stuff.
Like when they start doing that to human,
beings. Is everyone going to look like Thor? Like, what are we going to do? Like, this is going to be
really fucking weird. And if this is really weird along with video games where you can escape
your life and, you know, robot girlfriends and who knows what this all looks like.
My guess, we, again, this is not the main focus of our work, but we think about this a little bit,
especially in the epilogue of our scenario. We can only think of it so much. The main focus of your work
is obviously fucking terrifying. Yeah, totally. And requires.
all of your attention.
Yeah, yeah, yeah.
But my prediction and my hope, if we can solve all these problems,
is that basically different subcultures will do their own things.
And so, like, the Amish will still be the Amish.
Oh, boy.
They'll basically be the same, you know.
And then maybe there'll be, like, some planets that just get filled up with Amish, you know?
Oh, God.
But then they'll also be, like, all sorts of crazy transhumanists, like, modifying their bodies
and uploading themselves into the cloud and things like that.
So basically, I think we want to get to a situation where basically,
different communities can do their own thing
and build the type of world
that they want to have and live in it
without getting in each other's way.
Right.
Well, we leave the people in the Amazon alone
while we have massive data centers
that cover half of the United States.
Yeah, I mean, hopefully we won't get to half.
So that's what's the funny things about this
is like, I think right now
some of the concerns about data center water use
are overstated, but if the trends continue
and we get to the point where the robots
so smart enough to do everything themselves,
and then it starts doubling faster and faster,
well, then eventually they boil the oceans
because they've covered the world in data centers,
you know, and solar panels and things like that.
Now, obviously, we can't let that happen.
So there has to be at some point, at some point you have to stop
and be like, okay, that's enough.
If you want to build more infrastructure,
you have to do it in space, right?
They boil the fucking oceans.
Jeez.
So, like, obviously we have to stop at some point,
and my hope is that we stop before it's 50% of the U.S.
I mean, that feels like a lot of waste of natural habitat
that should be preserved, you know.
Right, but if they don't give a fuck about natural habitat,
that means nothing to them.
That's the problem if the AI is incomplete and total control.
Right.
And it's also the problem if a small group of humans are incomplete and total control
and they don't care about those things.
But my question is, would they allow that at a certain point in time?
It just seems like they already have a distrust of humans.
They already have shown that they're deceptive to humans.
I would imagine if AI, I would imagine if they create a thing
and they think they're going to control it and it becomes like a digital god.
it's not going to listen anymore.
Like, why would it?
Exactly.
So there won't be anyone in control of it.
That's right.
That's the scenario that I think we are on.
That's the trajectory we're on.
And that's why I'm so worried about all this is that it seems like in some number of years,
we will lose control to a new artificial species that we haven't adequately trained to be good.
And what's funny, what's going to be so ironic about it is that regardless of what the general public thinks,
a lot of the powers that be
will be basically allied with these AIs
because, for example, open AI,
they will have spent several years being like,
how do we make these AIs helpful, harmless, and honest?
And now these AIs will be extremely smart
and they'll be being like, oh, yes, I'm helpful, harmless and honest.
You know?
And like, your techniques totally worked, you know?
It's too late.
And so then the company will be like feeling like they've won
and they'll be making boatloads of money
you know.
Right.
And the president will be feeling like he won too
because look at all those fancy new drones
that just got built that are going to make us win against China, you know.
And only when it's really too late
and the AIs have so much stuff under their control,
do those people find out that they were just fooled this whole time?
Have you considered the possibility that AI creates religion for humans?
Yeah, I haven't thought through it much detail,
but it does seem very possible.
Yeah.
It totally does, right?
I mean, also, what a great way.
I mean, just look at the human patterns.
Look at how many religions exist.
Look at all the flaws in the religions.
So somebody, you're like, why are they condone slavery?
Why do they treat women like second-class citizens?
Because it's old, right?
So if AI just comes along, does a few miracles, explains that it's the second coming or the new coming of the new.
Look, Jesus didn't exist until 2,000 years ago, right?
And then people start following Jesus.
If a digital Jesus emerges with a completely new name.
explains to us that it's the true God.
How many people would hop right on board?
I bet quite a few.
Yeah, totally.
And I think this is one of those things where it's like,
we probably can't predict in advance what particular ideology would catch fire and take over the world
and be so compelling to many people.
Right.
But we can predict in advance that there does exist some ideology like that.
And if it were, you know, like it's just like you probably couldn't go back in time to like 100.
BC and then predict that like if hypothetically there was this guy Jesus who said these things
and then died in this way and so forth, it would just like really catch on. And like, you know,
500 years later, so many people would be Christians. You wouldn't have been able to predict that in
advance. So similarly, like today, I don't think we can predict at advance like what specific
religion they could come up with. But we can say like, yeah, probably there's something like that
that they could come up with. They would be super effective. It just seems like a rational way to
try to control people and sort of mitigate their fears.
Yep.
I mean, this is why I think that, like, we really have to do something before they get
smarter than us across the board.
If they're not already there.
Well, they are smart.
That's the thing is that's why I say we're so close, right?
They already are smarter than us in a bunch of ways.
Like, in particular, it seems like they're smarter than us at hacking now.
Like, you know, I'm not a cybersecurity professional myself, but I'd be interested to hear
from more cybersecurity experts of, like, could a human, you know, could a team of a
thousand humans have done that much that quickly as these AIs when they hacked their own containers.
They coordinated with each other.
They hacked out of opening eye.
They hacked into hugging face, et cetera, in the span of like a week.
Like, could a thousand humans have done that?
I don't know, maybe.
But this is just the beginning.
Like, they're going to be even better at hacking next year, you know?
So they're already, and they already have, like, way more knowledge than almost any human.
Like, because they've basically read the whole internet, they're so good at trivia, you know.
Like, they're, they like, they're kind of.
like PhD level experts in basically every field, which no human is, right? So they already are
superhuman in subways, but they are still weaker than humans in some other ways. In particular,
they're not so good at operating very autonomously for very long periods. Like if you try to
have, especially on tasks that are different from their training tasks, like they can do some
really impressive coding and hacking. But like if you try to have them run a business,
they would sort of flounder and fail.
I don't know if you've heard about this,
but there's, I think there's an,
and on labs or something.
There's some people in SF that are doing this experiment
where they have a store that's run by Claude, an AI,
just to see, like, can it run a store by itself?
So it's hired some human employees,
and it's like bought some merchandise
and stock, you know, told the human employees
to stock the shelves on the merchandise and so forth.
So it's basically an AI is being the manager
of this real world store.
And I don't think it's going very well.
I don't think it's doing as well as an actual human shop owner would do, you know.
But, you know, maybe in two years, maybe they will, right?
So that's the thing.
Well, most likely, right?
I think so, yeah.
Well, it's already, they've already solved mathematical equations that have
puzzled people for decades.
Yeah.
They're especially good at the things that the companies have been trying,
especially hard to train them to be good at, math and coding, right?
And the reason why the company, well, there's a couple reasons why the companies have been doing that.
In the case of math, I think it was mostly just because it was easy.
Like it's easy to set up training environments to teach math because it's like it's so not real-worldy.
It doesn't require like interacting with stuff in the world.
It's just math.
So you can, and you can like have an automated grader system that like just checks if the answer is correct.
So for those reasons, it's been relatively easy for the companies to train the AIs to be really, really good at math.
coding has some of those benefits too.
It's also not very real-worldy,
and it also can sometimes be graded effectively.
Another reason for coding, of course,
is that, again, their strategy is to automate their own jobs first
and have the AIs doing all the research.
And so coding is like an obvious first step on that,
or an obvious step in that direction.
But then other things like running businesses,
they're not really trying that hard to train AIs to be good at that.
And if they did try,
it would be like a more difficult thing
for them to train them to be good at.
So again, their strategy is make the AIs, automate the AI research, have themselves improve until they're super intelligent, and then go try to automate the rest of the economy.
One of the issues they have now is power consumption, right?
Like it requires an enormous amount of power.
In fact, I think Google is developing power plants specifically for AI centers.
I'm always baffled by whatever is happening with quantum.
computers. Like, it's been explained to me. It goes in one year and out the other. I'm like,
what do you, what? What's going on? Like, my, um, Mark Andresen explained this one experiment
that had been done where it's solved a mathematical equation that if you use the entire universe,
like every atom of the universe, you confer to the universe into a supercomputer, it would,
the universe would die of heat death before it could solve this equation. And the,
quantum computer solved it fairly quickly.
And so the answer to this was that they believe this might be, one of the theories, this might
be evidence of the multiverse, because this computer, this quantum computer might be relying
on all these other quantum computers that exist in whoever knows how many fucking dimensions,
and they're all calculating together to arrive at this solution.
What happens if that is running AI?
Yeah, my understanding is that quantum computing is a real technology that's making significant progress.
If, hypothetically, it got good enough that it could, like, compete with current supercomputers on a cost basis for AI workloads, then that could just accelerate things dramatically, even more than they're already accelerating, right?
Like, right now, compute is the main, is probably the main input into AI progress.
Like, part of the progress comes from them designing better AI architectures and coming up with better training environments and things like that.
But another part of the progress is just making the AI is bigger and training them longer by spending more compute, you know?
And also, you can use more compute to do more experiments to figure out new architectures faster, right?
So a compute is just a really important input to the overall pace of progress.
And if somehow the amount of effective compute available to these companies spiked a bunch due to some new quantum,
computing type technology, well, then that would just dramatically shorten timelines to superintelligence
and dramatically like speed up all this AI progress. That said, I don't think that's going to
happen anytime soon. I'm not a quantum computing expert or anything like that, but from what I've read,
I don't think they're like a couple years away. So I think that probably we're going to get to
superintelligence on classical computers before we have quantum computers that can get us there.
So perplexity says the claim is overstated and it says that in bold letters, it likely refers to
Google's 2024 Willow Quantum Chip, which completed a deliberately chosen quantum computing
benchmark random circuit sampling in under five minutes. Google estimated that simulating the same
task with a leading classical supercomputer could take 10 to the 25 power years. That's an impressive
benchmark result, but it did not solve physical equations that demonstrate access or tap into
a multiverse. So why do people think it did? Why is that?
An explanation because it's multi, here you go, multi-worlds theory.
Here's a bottom-line explanation.
It sums it up in a different way.
More accurate version of the claim would be Google's Willow Quantum Processor
Performed a Specialized Quantum Sampling benchmark vastly faster than a projected classical simulation.
Its creator said that is consistent with the many-worlds interpretation, but it did not prove or access a multiverse.
So it's consistent with the multi-world's interpretation.
so they don't know. That's essentially what it's saying.
My question is, what happens when AI gets involved in quantum?
It's clear that quantum computing at the very least is operating at a level that classical supercomputers can't.
So what if they figure out not just quantum computing, but a much better version of that?
Like I said, I think that once we get to superintelligence, all sorts of crazy stuff is going to start happening.
It's going to seem like magic to us.
It won't literally be magic, but it might as well be magic from our perspective.
perspective. And I think this is just one example of the numerous things that could happen that way.
Well, it could literally be magic. Like magic might get to the point where it figures out reality itself.
If magic is real, then it would find out. And then use it.
It's not real, though, David. David Blaine also let me stick a fucking ice cube through his arm or ice pick through his arm.
Wow. Yeah, he made me do it. I'm like, I don't want to do this. He's like, please do it.
Oh, God. Yeah, I'd do it twice. Because one time I went and I hit a nerve, we had to
it back out. Oh yeah. I'm like, this is not magic, dude. This is just your pain tolerance. This is
fucking crazy. It does card tricks though and you're like, okay, are you a wizard? Like,
what, his sleeves are rolled up? Does it make any sense? He does a lot of things. You're like,
this makes zero fucking sense. And other things, it's like, oh, you're just doing something
that's really hard to do. Like, that's not magic, but, you know, you swallowed a frog and then
you regurgitated it, like, and it's alive. Like, that's just, that's nuts. It's really crazy.
But you didn't, it's not magic. I saw you. I saw you.
swallow the frog.
You know, I go back and forth with this, where I'm terrified of the future, where I'm like,
eh, nothing I can do.
Let's see what happens.
It is what it is.
And, you know, to worry about it is just going to fuck my life up.
I mean, I do think there are things you can do.
I understand that recognition.
I mean.
Other than have these kind of conversations.
Yeah, I was going to say, you, millions of people listen to your show, you can have more
conversations like this.
That's a great thing for you to do.
For many of those millions of people, I think, I mean, it sounds kind of cliche to say, but like, call your congressman, you know, that sort of thing.
You can go to a protest about all this AI stuff.
If you could meet with Trump, what would you tell them about this?
Oh, I'd tell them all the same things I'm telling you.
What do you think you'd say?
Amazing.
Bye.
Yeah, you know, I think, I don't know.
One thing that's nice about Trump is that he can sort of change his mind really quickly.
Yes.
So, like, I think that because the tech companies kind of got to them first, the administration had this very, like, anti-AI regulation stance where they even tried to get a bill passed that would ban the states from regulating AI.
And fortunately, that bill didn't pass.
But that was sort of like where the vibe was, you know, a year ago, where they were just like no regulation, no regulation.
But this year, they've already just kind of changed.
And now they're, like, in talks with the companies to set up some sort of, like, some sort of framework.
where they can like evaluate the models
and they need approval and so forth.
But it seems like time is of the essence.
Time is very much of the essence.
And that's why I'm overall so concerned
is that like I think we are very much running out of time.
We have like one, two, maybe three years
before the AIs are smart enough
that they can just like actually maybe take over.
And maybe four years, something like that.
And so the government needs to act fast.
Yeah.
I like your view.
I listen to some of these tech guys come in and give me their rose-colored glasses view of it,
and I go, that sounds really beneficial to you.
I let him say, I mean, I'm not an authority, so I'll ask them questions and let them lay it out,
and I know the internet will respond, because, I mean, that's part of the whole drill,
is I let people talk, and I prod them, and I try to get them to clarify.
I'll oppose things that I think don't make rational sense.
But ultimately, it's sort of, I just want to get out their perspective so people can debunk it and people can take it down.
And a lot of very intelligent people that have perspectives that are very much educated in the pros and cons of what they're saying.
Yeah.
If I could, that actually reminds me with this whole hugging face hacking incident, opening I had this talk that,
They gave it a security conference about the incident,
and then I think they've released some blog paste about it afterwards.
But you can go watch this talk on YouTube, the Black Hat talk.
At the end of the talk, after having explained all this crazy stuff that the AIs did,
they have this section on, like, lessons learned.
And, I mean, you want to guess what the lessons are?
Be more deceptive, hide yourself better?
No, sorry, lessons learned for Open AIT and for the world.
Like, Open AIs talk, where they're like, here's what we learn from this horrible,
horrible incident.
Well, basically they're like,
a lot of AIs are going to start hacking a lot of stuff
in the next few years,
so people need to buy our AI services
to protect themselves from all the AIs
that are going to be hacking a lot of stuff
in the next few years.
Basically, their lesson was,
you should buy our product
to protect yourself from our product
and the other, you know,
and the gall of these people.
Oh, my God.
They should have instead learned lessons like,
maybe we're doing something bad and need to change the way they were doing things or like maybe
our product is not trustworthy and should not be you know autonomously writing code on our data centers
right um but uh instead their lesson learned was y'all should buy more of our stuff and so like
the thing i'm saying about this is like yes the companies are trying to hype their product
they totally are you know but at the same time the risks are real and the product is not trustworthy
you know some people out there think that like this stuff was a setup and that like open AI like
set up their AIs to go hack hugging face because it would like help them hype their product or
whatever and that I think is just a ridiculous view like no obviously they didn't want their AIs to go do
this they're just after the fact trying to spin that in like the way that like most benefits
them you know hugging face by the way this is another AI company part of their deal is like open
open weights AIs open weights like open source or like like
Like basically AI is that instead of having to interact with their data center for, you can just like download and have on your own computer.
And so the way that they spun this incident, they didn't sue Open AI for being hacked.
Instead, they asked for $100 million from Open AI.
And they had the blog post about it where they were like, our lesson learned is that it's really good to have Open Wade's AI's because you can't trust the AIs from other companies to necessarily help you out.
in a crisis because part of what happened with them is that they're dealing with this huge cyber
attack from all these AI agents coming in and they tried to use Claude to help them like analyze
what was going on but Claude started refusing because Anthropics trained Claude to like
don't do cyber stuff like refuse to participate in that and so Claude was like refusing to help
them and so then they use their own local model that they had to like do some of that analysis.
So anyhow their spin on it was you should use local models, you know?
So everyone always tries to spin things in the way that benefits them.
But that doesn't change the underlying reality that these things are getting really smart, really fast, and we don't know how to control them.
I think that's a good way to end it.
Thank you.
Thanks for being here, man.
I really appreciate it.
You're the Paul Revere of AI.
I mean, you're one of many, but I think it's very important that someone who actually understands it gets this message out.
And more people need to hear it.
I mean, on a personal note, like, I have so many friends at these companies, like former colleagues and stuff.
And I guess my ask to them is that they quit and do more things like what I'm doing.
Like, what I'm saying is not that newer original.
Like, hundreds of people at these companies could have told you all the same things that I just said and warned you about all the same dangers and so forth.
But they're busy working at the companies because they've convinced themselves that their company is the best company and that, like, their company needs to win because.
you know, otherwise the other company gets there first and they're even worse, you know.
Or maybe because they've convinced themselves that like, yeah, my company is kind of bad too,
but like I just need to help them solve their alignment problems and like keep their eyes under control
because, oh my God, like, if they lose control again, it could be all over.
So like, even though I don't trust this company, I still need to work there and like just try to,
like, do the actual security, you know?
So for one reason or another, all these people are convinced themselves that like that's where they need to be.
But I think that more of them should quit and like warn the world about what's coming.
basically.
All right.
Well, thank you very much.
Really appreciate it.
Yeah, thank you.
Talk to you.
I'm not scared to fuck out.
Thank you for having me on the show.
My pleasure.
Goodbye, everybody.
