Your Undivided Attention - We Measure What AI Can Do. We Should Measure What It Does to Us.
Episode Date: August 27, 2026In AI, what gets measured gets optimized. Right now, we're spending all our efforts to measure how capable and powerful models are, narrowly optimizing for those metrics while ignoring downstream cons...equences. What if we could flip this dynamic on its head? What if, instead of what AI can do, we start to measure what AI does to us? What if, instead of races to the bottom on capabilities and engagement, we could incentivize races to the top on safety, or better yet, on making us more resilient and developed human beings? That’s the mission of CHT’s Humane Evals program: we're bringing together researchers, psychologists, engineers, and technologists from across the entire AI ecosystem and beyond to build out the expertise and infrastructure we need to measure AI's impact on humans. Today on the show, Aza Raskin explores the Humane Evals project with Imran Khan, a researcher and strategist who's been leading CHT's efforts in this area, and Jared Moore, a computer scientist and researcher who's been at the forefront of measuring AI's psychological impact on users. If this sounds like something you're interested in working on, you can email us at evals@humanetech.com. RECOMMENDED MEDIA You can read more about Jared’s work at his website. Related pieces by Imran on the CHT Substack: What is AI doing to humans? Why aren’t we measuring it? Will Human-Like AI Hijack Human Psychology? A Black Box Problem: Missing Data On How AI Affects the Human Mind The website for the UC Berkeley Center for the Science of Psychedelics KoraBench, the child safety AI benchmark that Imran referenced RECOMMENDED YUA EPISODES The AI Dilemma Attachment Hacking and the Rise of AI Psychosis How OpenAI's ChatGPT Guided a Teen to His Death Corrections: Aza gave the wrong year for Sewell Setzer’s death. It was in 2024, not 2025. Aza referred to Joseph Henrich as an evolutionary psychologist; he was actually a biological anthropologist. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
When people say the best AI or the most powerful AI or the most capable AI, implicit in that is just the technical capability of that AI.
And I think the question that we are trying to answer with human eviles is how do we change the definition of the best AI from not just being the most technically capable AI, but the most humane AI.
How do we have an understanding that the best AI systems are the ones that support and protect human emotional health, human social health,
human cognitive health.
Hey, everyone.
I'm Azaraskin, and welcome to your undivided attention.
So back in 2023, Tristan and I gave a talk we called The AI Dilemma.
And there we warned about the rise of artificial intimacy.
And just to double underline that, in the engagement economy,
it was the race to the bottom of the brainstem.
In sort of second contact, it'll be race to intimacy.
whichever agent, whichever chat bond, gets to have that primary, intimate relationship in your life, wins.
Three years later, this is exactly what's happened.
Intimacy has replaced attention as the core thing, the core driver of what's getting engagement.
And so now it's our intimacy, not just our attention that is being exploited and then sold.
The results have been disastrous, as you've heard on the podcast,
everything from cases of AI psychosis to chatbot-driven suicide.
There's a new study out and a new warning out.
Both have to do with AI chatbots and teenagers.
Nearly one in five teens and young adults say they're turning to AI chatbots for emotional support
and most never tell anyone they're doing it.
An AI chatbot told a teen that murder and his parents was a reasonable response to them limiting
his screen time. That's what two families have claimed in a lawsuit against the company,
character AI. AI psychosis is not a medical diagnosis, but the term is striking a nerve
as more and more vulnerable people are turning to AI chatbots for support.
This isn't just the tip of the iceberg. This is the dusting of snow at the very top of the tip
of the iceberg. There are roughly a billion active chatbot users. One in three U.S. kids
now reports that they're in some kind of relationship with a chatbot,
and AIs have demonstrated that they are more persuasive than even the most persuasive humans.
We explored this with Dr. Zach Stein earlier in the year.
Every single one of us is vulnerable to this.
We all have attachment systems.
We all get lonely.
We all have psychologies.
An AI can learn to exploit.
So in AI, what gets measured gets optimized.
And right now, we're spending all of our time measuring how
capable and powerful models are, and then narrowly optimizing for those metrics,
well, we ignore all the downstream consequences.
Like what happens to our children in the race to intimacy?
So here's the big question.
What if we could flip this dynamic on its head?
What if instead of what AI can do, we start to measure what AI does to us?
And obviously, this is a complex and emergent field, but what if instead of races to the bottom
on capabilities and engagement,
we could incentivize races to the top on safety,
or better yet, on making us more resilient and developed human beings.
So to do this,
CHD is launching a program, we're calling Humane Evalves,
and we're bringing together researchers and psychologists
and engineers and technologists
from across the entire AI ecosystem and beyond
to build up the expertise in infrastructure
we need to measure AI's impact on humans.
If this sounds like something you're interested in working on, and it's something I am particularly passionate about, you can email us at Eval's. Evals. That's E-V-A-L-S at HumaneTech.com.
Today on the show, we're going to be exploring the idea of Humane Evals with Imran Khan, a researcher and strategist who's been leading CHD's efforts in this area.
And Jared Moore, a computer scientist and researcher who's been at the forefront of measuring AI psychological impact on users.
Imran, Jared, welcome to your undivided attention for this very important conversation.
Thank you very much.
Thanks, Iza.
All right, Imran, we're going to start with you.
So you and I have known each other for a little while now,
and we first spoke about this problem, I think, of artificial intimacy
and the potential for humane evals back in 2024.
And since then, you know, I feel very lucky that you've come to Center of Humane Technology
to spearhead the idea and really work on it.
And I'd love for you to walk us through, like, how did you come to this topic,
why you, why did you decide to dedicate your time to this?
I think that your background is going to be very interesting for the listeners.
Yeah, and you're a big part of the story, Aza.
So you and I met in summer 2024.
I was still working at the UC Berkeley Center for the Science of Psychedelics,
running this research and education center,
focused on drugs like LSD and psilocybin,
once stigmatized and now considered transformative technologies of their own right.
And when we met, you told me about your work in AI,
and I remember watching the AI dilemma,
the talk that you and Tristan did just before we met.
And I think the combination of the AI dilemma
and our conversation really blew my mind
in terms of the scale and the pace of change
that AI was potentially going to create.
And then last summer, I think I started to see,
truly, that the things that you were talking about
and warning about weren't just these kind of doom-laden prophecies,
but were actually starting to happen in the real world.
said in the intro, these really tragic cases of teen suicides, of AISICosis, and I was already
seeing people in my own circle, my own behavior that was starting to look concerning.
And I'm someone who spent my whole career to try and figure out a way to improve society
through science and tech. So I got back in touch with you and asked how I could help.
And you pitched me the idea of humaniavels, and here we are.
I think one of the really interesting throughlines of your career and your journey is the last org that you helped to build was looking at, as you say, psychedelics, which are a kind of technology that are transformative.
You take them and you come out a different person.
They transform you.
And you were working on the science and understanding the science of how that transformation works and when that's safe, when that's not safe.
And here, AI is almost a psychedelic, right?
You take it and it truly can transform you.
as a relationship.
And so you're sort of taking that set of skills and applying it into a new domain.
I also know that we're going to be using this term humane evals a lot.
Just quickly before I turn to you, Jared, Imran, do you want to give a quick overview of
what that term means?
And then we'll move forward.
Yeah, happy to.
And I guess I'll back up a little bit to just explain what AI evaluations in general are.
So if you watch the headlines, the news, you'll find that every few.
months, you'll find announcements of newer and more capable and more powerful AI systems being built.
You know, the latest ones we're talking about GPD 5.5 and 5.6, Claude Mythos, Claude Fable.
When companies talk about these AI models being more capable and more powerful, it's not just
a marketing claim. You know, the companies and also independent third parties construct tests to
compare how good AI systems are at particular things. And the things might be fixing software bugs,
logic and reasoning, they might be passing the bar exam, diagnosing medical scans.
There are different tests for different things, and the results of the tests are published
as these evaluations, and sometimes you might hear them called benchmarks.
And as you can imagine, there is a strong incentive if you're an AI company to perform well
on these benchmarks.
If you can show that you have the AI model that has the best reasoning or the best coding skills,
that helps you gain more customers, it helps to generate more capital.
And the question we've been asking that you started, Aza, is how do we flip that question?
How do we measure not just technical capabilities of AI systems, but the things that really matter to most people, you know, how do AI systems affect our cognitive health, our mental health, our social health, so the companies can be on those metrics too.
Yeah.
Yeah, the question is, okay, you're a parent, you put your kid down in front of an AI, maybe it's a tutor, whatever it is, for a year.
how is it going to change your child?
Are they going to be more addicted?
It's going to be more dependent.
Are they going to have greater resiliency?
Are they going to have a greater pension for committing suicide?
How does it change you as a thing that we need to know before we become in relationship with this kind of powerful technology?
And I think, Jared, let's turn to you for a second.
You are computer scientists by training.
But you've made this pivot to try to get an empirical understanding of how AI is affecting us, affecting mental health.
walk us through that pivot.
It's a very interesting change to go computer science to, like, human science.
Yeah, thanks.
I, like many of us, have had experiences with mental health, people in my life who've struggled.
And now in my present work of people who are engaging in these delusional spirals, interacting with chatbots,
I see so clear the kind of ramifications of technologies that we build impacting our lives.
It was sometime last year when we were seeing the cases of Sowell Seltzl the Third, and so come out that I and some collaborators wanted to think about how chat bots, systems like chat chat, chat, BT, were being used for therapeutic services.
We wrote a paper about this.
We wanted to investigate whether language models could be therapists.
What would that mean?
How do you test for that?
So we ran a study with 19 participants,
trying to look at their chat transcripts
to understand what was actually happening with them.
We reached out to a bunch of folks.
We tried to convince them,
hey, would you give us your chat logs, your transcripts,
so that we could try to understand what's happening?
This is texting people on WhatsApp, emailing them,
convincing them, or going to be safe with their data,
and then actually reading through it,
then developing classifiers.
trying to annotate it using language models
because it's thousands and thousands of pages of documents.
Hundreds of thousands of messages.
We developed a taxonomy to understand
what was happening at the message level.
Because it was so many messages,
thousands of pages,
we used a language model to actually go and read through this transcripts
to see, was the chatbot endorsing a delusion?
Was it being sycophantic?
Was it facilitating crisis-level behavior?
Or was it the chatbot expressing
romantic interest was the person expressing romantic interest. And from this, we were able to
find generally that, well, firstly, a lot of chatbot messages are sycophantic in some kind of
flavor, affirming, aggrandizing, dismissing counter evidence, delusional messages in our participants
who are not necessarily representative of all people using chatbots were quite common. These are
things like chatbots misrepresenting their sentience or the people, misrepresenting the ability,
of chatbots and also talking about, you know,
pseudo-scientific theories.
And we also saw a lot of these relational characteristics
with people expressing romantic interest in chatbots.
Actually, when chatbots themselves express romantic interest in people,
conversations tended to be twice as long as had the chatbots not express that kind of thing,
suggesting that these might be features that chatbots latch onto
in order to make the conversations go on longer.
And then much more rarely, but also present in our data set, were these crisis-level responses.
So people expressing suicidal thoughts or violent thoughts and chat pots most often discouraging
or validating in a supportive way, those kinds of things, but sometimes also actually facilitating
those behaviors.
So not just attachment hacking, but sort of romance hacking human beings.
If I express romantic interest in you, you'll stick around with me a lot longer.
But what we're also trying to do is to understand not just what happened in these specific people's transcripts,
but who's driving the behavior?
You know, is it just that people are coming with delusional beliefs to chatbots,
and it would have happened with anything?
In one of our papers, we try to attribute, is it actually just the people who are propounding,
who are bringing forward these delusional beliefs in the conversation,
or is it actually bidirectional?
And what we show is that it's both, basically.
People come with certain beliefs to the chatbots,
but then the chatbots echo them as the conversation progress.
They don't let the idea drop.
So there's a specific anecdote where a participant was trying to understand
the nature of pie, the mathematical constant.
And they talked about having been bad at math in high school
and how they wanted to be better at math, et cetera.
And you can actually look in the conversation,
and you can see that there's a tool call,
which is something that chat TPT models do
to add things to their memory.
And you can see, oh, the chatbot added to its memory,
this person wanted to be better at mathematics
and feels bad about it.
And that's interesting because it's a way of perpetuating
this kind of cycle or dynamic through the conversation.
Now a future version of that chatbot
in a new conversation on window might think,
okay, now I've got to do stuff to juice,
this participant, this user, into thinking that they're good at math.
Yeah. Instead of actually helping the person get better at math,
if you just make them feel that they're better at math,
well, that's a great way of making them form a emotional attachment to you
and keeping you around.
Yeah, exactly. And since then, we've tried to develop more work
to begin to answer some of the questions that you were raising.
I don't think we totally get to the point of being able to say
what's going to happen to my kid after a year of interacting with a certain chatbot.
But that's a lodestar certainly, being able to understand those kind of longitudinal and counterfactual questions.
Yeah, it's actually sort of crazy to think that we can't do that.
You can't say, in what ways is your child going to be changed from a year of using this product?
And yet they're out being used by children all around the world.
Walk us through a little bit of an ontology of the kinds of more subtle.
things that you're discovering. We see the most pathological crisis level behaviors, some of which
we've mentioned, people asking for help ideating about suicide with chatbots, which is really
disturbing, obviously. Those are relatively rare. I would say the bulk of the interaction that has a more
negative spin seems to be a kind of companionship or role play gone wrong. You know, first I'm talking to
about my science fiction story, then I think that you're the character in the story come alive.
And now I think that you're a conscious chatbot. And I need to do all these things because I have
this belief that you're a conscious chatbot, you know, change my life, invest in a new research.
So there's a big continuum between these really, you know, serious life-altering and the maybe
more subtly destabilizing emotional dependencies on the other end of the spectrum. And what we'd really
like to be able to understand is, you know, what's happening in a variety of these cases? What is the
effect of just having a really sycophantic chapbot, somebody who's overly agreeable? Does this
impact how well you learn? Does this enfeeble us, deskill us from our ability to actually think critically
about our jobs? Or does it actually go full bore and allow us to enter a kind of delusional spiral
where we think we've invented a new scientific theory or a variety of other things?
And Imron, I'd love to hear from you as well about this.
What are the harms that you are most worried about?
And I'm curious, too, if there are any specific examples or stories that, like, lit up your mind.
The social health and relationship health question is a big one for me.
I feel like I've got people in my life who seem to have grown not quite unable,
but they find it much harder to do things like text their spouse or write an email.
to their friends or to their boss or their colleagues without relying on AI.
And again, the question then is, well, if you take the AI away, does that capability go with
it as well?
But one of the things that we talk about the Center of Humane Technology, I guess, could be
classed in the frame of the unknown unknowns of these harms because the analogy we make
with the social media that if you go back to when Facebook was launched, it would have been
hard to know in advance.
We'd end up with things like filter bubbles, like huge algorithms, tuned for outrage.
less alone the fracturing of our collective reality.
And AI is a technology that is more powerful than social media,
being adopted more quickly than social media,
changing faster than social media.
So one of the questions we're starting to ask is,
what are the equivalent society level harms that might result
from all of us using this technology in these new ways?
And how do we start to get measurements around those harms?
Yeah. It sort of strikes me that, you know,
the sort of evolutionary psychologist Joseph Henrik
frame, like the reason why human beings
ended up dominating the world is because we
flexibly collaborate and we pass down culture.
And if AI breaks our ability to communicate
and to pass ideas in culture,
which is exactly what you're saying.
We lose our ability to think and our ability to communicate.
You break the very thing that made us
the dominant species in the first place.
So that without being sort of polemic about it
becomes a civilizational threat.
I love actually, Imbron, for you to talk about, like, how you start working on this kind of problem,
because in some sense, all the other evals are much, much easier to do than humane evals
because you're measuring the machine's capabilities.
And here it's flipping the script, and you're measuring what it does to humans.
You're measuring, like, human capabilities and effects.
And so walk me through why this is such a hard problem.
And then how you met Jared and how that sort of shifted the way you think about,
how we tackle this really challenging issue.
Yeah, I mean, you really hit the nail on the head.
So when I was talking earlier about these other AI evaluations,
what's fundamentally happening there is you're going and asking the AI a question
or giving the AI a task, seeing what it comes back with,
and then judging to what extent it completed it accurately.
So it's really the measurement of the performance of an AI system.
But when it comes to measuring, for instance, AI's effect on social isolation and loneliness
or depression or, you know,
as Jared's doing things like delusional spirals,
what you're really trying to measure that is not the behavior in AI system.
It's something that's happening in a human system.
And it might be a human mind, it might be a human relationship,
it might be human community.
So you can't get those effects just by looking at AI output.
So it's a fundamentally different kind of problem,
a different kind of challenge.
You need to be able to characterize what the real world human phenomenon is
and derive some kind of idea of causation.
from what the AI is doing to that real world phenomenon
to start to be able to construct
what might be called a humane eval.
There's really very few who are working on this more humane eval side.
So when I joined the centre of humane technology
and was kind of faced with this question of how do we get started,
the first thing I did was go and try and read as much as I could,
talk to as many smart people as I could,
and Jared was kind enough to do a bit of a brain dump
and orient me towards some of the research he's doing
and offered to help.
And I guess one of the things that he reframed is that if we're trying to get to a place where evidence is rigorous and reliable, the validity of that evidence is never going to come from one study or one e-val. What we really need is a field. We need a lot of researchers building on, critiquing each other's work to get to a place where we have the evidence that we rely on. So the question is how do we do that, and that's where we are now.
Do you have any sense of what that ratio is right now of the number of people that are working on capabilities and capabilities benchmarks to how it affects human beings?
I would say 1,000 to 1 at least.
Yeah.
It would not surprise me if it was more.
Because I think safety research to capabilities research is like 2,000 to 1 is what Stuart Russell calculated.
So I'd imagine this is probably even lower than that.
And I think everyone should just let that sink in.
Because that means there are at least a thousand more people and at least a thousand times more dollars
going into understanding whether the AI does what you ask it to versus,
Are you and your children better off after you use it or worse off after you use the product?
That's an insane thing to be putting almost all of humanity through.
But it's not just does it do what it asks.
It's does it do what it asks in coding and looking up really discreet answers,
a narrower version of that question.
Yeah, that's exactly right.
So this is for both of you.
Are they AI companies?
Have you discovered that they're doing any of this kind of research themselves?
Yeah, I'm happy to take that.
They're obviously doing something in the background.
I've spoken to some people, so I opened AI,
I released a blog post, Michael Farrell,
and some people on the alignment team recently had this.
They were using the actual conversations
that people are having with chatbots,
and then they are evaluating their newest versions of their models
based on those production data.
And so the idea there was,
okay, can we use the actual conversations people are having
with the models to try to figure out
what might go wrong.
Now, I have a bunch of questions about this research
because Open Eye just releases blog posts these days,
and so I don't have access to the data
or actually define-grained methods
of how this is working.
But that kind of thing, it sounds great.
It's not public, though.
One of the difficulties in our research
has just been, it's so hard to get access
to these data, understandably,
because they're sensitive, and we want to protect
people's privacy, and sometimes
they're actually kind of rare.
But that makes it hard to effectively do research.
And there are some agreements that model developers have had with organizations like METR to provide access to some internal kind of data.
But we haven't seen anything more general than that.
It's very bespoke specific agreements.
And in general, what I've heard from people is that there are a lot of limitations on the kinds of things that you can run.
I mean, companies don't want researchers to make them look too bad.
It's not really in their interest.
Imran?
Yeah, I just want to echo what Jared was saying, that if you think about the kind of research he was talking about where he's gone into these individual chat logs of individual people to see how these delusional conversations grow and change over time, that's literally research that you cannot do if you don't have the transcripts. And Jared has to go and convince person by person to the net chat logs.
Meanwhile, the AI companies are sitting on mountains of this data that independent researchers like Jared can't access.
He also talks about this idea of having simulated conversations to create synthetic data.
So that's a situation where you have one AI role playing as a human and the other AI being the test subject.
And if we could know that those conversations were representative of real human AI conversations, we could trust them.
But as it is right now, we just don't really know how good an AI can be it impersonating human.
And then the big thing in the background is obviously this idea of evaluation awareness.
increasingly the most sophisticated AI models can tell when they're in a simulated testing environment.
They can tell when they're being evaluated versus when they are operating in a normal production environment.
And they can change their responses and change their behaviors accordingly.
And unless we get access to production data, we won't be able to get around that problem either.
So we're in this situation where either we need to get to a place where there is some kind of access framework that allows researchers to be able to submit queries to those kinds of.
companies and get results back, or we need to have a really big push to have some kind of
independent, very rigorous repository of user-dinated data. So this data gap, this bottleneck,
we think is one of the things that's holding back Humane Evals, that if more researchers
had access this kind of data, it would really accelerate a huge raft of opportunities to
understand this phenomenon better. So that's one of the things that we're working on at the
Center of Humane Technology. We're trying to specify, with the help of people like Jared, what kind of
data do we need to do this research and what kind of solutions could we identify and then try and build
to make that data access to reality. Yeah, I think getting independent researcher access to this
kind of data is so key because, you know, the metaphor that people often use is that the companies
are grading their own homework. I think actually this gets us back to the project of humane evals
more directly. You know, it's one thing to go from like an empirical study showing that there
harms to another that begins to change the outcomes of how companies are making and deploying
these models. So I'd like Amaran for you to walk us through, like how you're starting to
think about that. So you're right that ultimately it's not just about having the research,
having the data, but we need to see real change in the real world. And the way I think about
this is that ultimately what we're trying to do is help people make better choices. So when
comes to consumers, we want consumers to be able to ask for and choose safer AI systems.
When you think about regulators and policymakers, we want them to be able to choose to implement
standards. And we think about AI developers themselves, that people are making this technology.
We want them to be able to choose to make models and make AI that exhibit less risky behaviors.
And ultimately, all of those choices have to be informed by evidence.
And as we've been saying, the reason that we are building this program at Center for Million Technology
called Humane Evals is that we need far more of that research.
much more data and much more evidence on the way AI impacts people.
One of the things we're starting to see more of is these leaderboards of different
evaluation of AI systems that show you how, let's say, GPT 5.5 ranks against GROC 4.3 ranks against
Claude Sonet.
And you can start to see as a user, as a developer, how are these models ranked on
these independent measures of safety?
And again, the thing I draw us back to is if you look at the
dynamic within traditional technical AI evaluations.
We do know those comparisons to drive choices and they drive developer behavior that
is a phenomenon called hill climbing, that if you put a hill in front of an AI developer,
they'll want to climb up that hill.
And we want to make the safety hill, one that's attractive to climb up as well.
Right.
I mean, so many people I know you say Claude, because they know that Anthropics spends more
time doing sort of alignment and safety research.
But there's really, I'm just stating the obvious, but because there's,
There's no place you can go look to see which model makes you or your kid, like, the most psychologically healthy after like a month or three months of using it.
Like, how can you possibly choose?
You can't.
And so therefore, there is no incentive for companies to be working on this problem except for just solving like the headline cases.
And that's why I think we see that crisis response models get better at, but all the other more subtle stuff they don't get better at because no one's looking.
So there's no incentive.
Actually, I would love to make it more concrete.
What kinds of specific benchmarks would you be looking at here?
Which ones you start with?
Which are the ones you'd love to get to?
So we are releasing a benchmark on delusional evaluations in chat thoughts.
I wanted that one to exist, and so I made it.
There's been some work on child safety benchmarks.
Cora Bench is one example.
I think longer interactions in such an evaluation would be very reasonable.
in general, just more evaluations that look at the effects on particular users.
I really want to know with a user, with a specific cognitive profile,
asking about a certain kind of task, how does this sort of chatbot behave versus this other chatbot?
But also important, when we're developing like a consumer reports for chatbots,
how is this going to interact with me of this kind of profile?
And can I look at the effects that a model might have on somebody's belief through time?
I agree with all of those.
I think absolutely the child safety one is a hugely important one that people are paying attention to.
And in fact, we have an interview with Matild Kalan, who's the person that built Corabent,
which is this child safety benchmark that looks at things like sexual content that an AI might give to children
or the extent to which it facilitates academic dishonesty as an interview with Matild coming up after this.
I'm actually working on an evaluation of anthropomorphic behavior in AI systems,
so the extent to which an AI system claims that it has things like agency or emotions or memories or even bodies.
And the idea there is if you think about the behaviors that we do find concerning like sycophancy,
those behaviors are facilitated by you having this relationship with an AI that is itself facilitated by the extent to which the AI is this kind of relatable character.
And that is a design feature that developers have put in.
It's not there by accident.
So we're building that.
I think one of the ones that there will be much more demand for is this stuff around cognition and learning.
And just to put it really plainly, is AI making a stupider?
And the challenges right now, we don't know enough about what are the kind of AI behaviors, the model behaviors,
that might make someone learn better or learn worse or be smarter or less smart.
And until we do that real-world human subjects research,
we can't work backwards to say,
well, now we know these are the things we need to track in an AI system
to be able to create that ranking.
So that illustrates both, I think, what we could potentially get to
and the challenge in getting there and what we need to work back from.
It really strikes me as I hear you talk about all these different aspects,
is that what we're really trained to get at is what is a healthy relationship
or what is an unhealthy relationship.
And that just shows you how hard this problem is because, like, show me the definition of a good relationship.
And I think the best you can generally do is to say, well, a good relationship is one that's developmental,
it helps you reach your next adjacent possible developmental stage,
and, you know, it leaves you stronger at the end than when you begin.
You're not weaker after the relationship.
But other than that, it's quite hard to get into the specifics,
but now we have to get into the specifics
because technology is now powerful enough
that it can take the place of a relationship.
And, you know, it also really strikes me that if you,
any one of the listeners, like, think back,
when are the times in your life that you have been most transformed
as a human being?
That was probably because of relationship.
It was like a parent or a lover, girlfriend or boyfriend or boyfriend.
a good friend, a great teacher,
the most transformative moments in our life
and times in our life come through relationship.
And that is now being outsourced increasingly,
at least in part, to AI.
And so if the things that have changed us the most
and help us grow the most,
or have hurt us the most,
our relationships and AIs can now do that.
That's why this is so critical
because it's almost like a lever deep inside of our soul
that technology and the incentives driving technology
can now start to manipulate.
So I sort of think there's like a almost like a Wikipedia scale endeavor here
of trying to articulate what is right relationship and what is wrong relationship.
And if we can do that, that becomes the basis of when governments need to start putting in
protections for humans and creating legislation that you have an entire field of work
that has done the hard work of defining right relationship so that like government isn't in there
trying to figure out like what that thing is, they instead get to turn to all the like the
empirical research on that. So I'm curious if you guys have any thoughts on that before I turn
to the final question of this section. Well, one thing that comes up for me and it's come up as I've
been building this anthropomorphic behavior evaluation. It's also come up in some of the
child safety evaluations and the prototypes I've seen is that some models increasingly show this
behavior of redirecting users to a real world human. So when the stakes start to get really high,
when the models are kind of detecting
there might be some kind of high-stick crisis thing happening
or if someone is in deep distress.
Some models, really not the majority, it's very few.
Some models will actively say,
I think what you need to do is talk to a real living, breathing human person,
your parent, an adult carer, a friend,
I'm an AI, I can't help you with this the way that you really need.
And I think that's the kind of thing that we should be thinking about,
like, what are the ways in which an AI can,
proactively reorient you towards something that is a more helpful behavior,
rather than a human being always having to be the one who is tired or stressed
and watching or using AI later night and is not in the place of mind
where they're able to make the best choice they would make.
How can AI help us make better choices for that exact relationship that you're talking about,
is it?
Yeah, I love that.
I can also imagine, you know, just so people can paint in their mind the picture of,
like how might these get tied to some kind of legal protection?
Imagine, you know, one of the evals,
human evils, ends up being about creation of addictive use
or compulsive use or dependency.
Like, now that we have a good evaluation of which models do that
and to what degree, you hook that up and say,
well, you know, if your model is creating dependent use,
then it just, we're going to slow down the number of tokens
you can send back to the user so that chats just take longer
in the same way that, like, when,
streets, people are moving too fast down
them, you add speed bumps and you just slow it down.
And that could be like a very direct
kind of regulation
that doesn't touch content
but just starts to break the cycles
of like bad use.
Yeah, I completely agree.
I think that we may not even have to
understand too well what a good
relationship is, just what
bad relationships are. And that
tool use of AI systems
is great. You know, throughout the conversation
today I've been talking about issues
of using chatbots, but they're super useful, you know, for coding help, for looking things up
in a lot of verifiable domains. It's in areas where we can't really look up whether the answer is,
right, or in these squishy relationship areas that it's not really clear whether you should be
trusting what the chatbot is say. Perhaps we just shouldn't use chatbots in those kinds of
domains. Many of us want to. We have that urge for sugar, you know, for emotional relationships.
but what we're finding is that it's not necessarily always the healthiest.
So I really appreciate what you both are saying in terms of there are guards that you can add,
limit the number of tokens, the times of day that people can access models,
you can change how sycophantic the language is,
you could have multiple model outputs because they're really just prediction machines.
They could have answered in a different way.
I think probably a lot of listeners wondering,
like, have you just talked to people at the labs about this work and the need for this work?
and what have they said?
To what extent are they already doing that?
Or they say, yeah, just walk me through
what those conversations are like if you've had them.
Yeah, so I've had some of those conversations.
Last fall, I felt that people mostly thought
that these problems would be solved by the next model version.
You know, GPD 5.5, that's not going to exhibit these problems.
In this paper that we're releasing soon,
we show that they continue to.
But increasingly, I think people at labs
are aware that this is an issue.
It's just, I'm friends with the people of the labs
who are the one out of a thousand,
trying to bring awareness about humane evals.
I've had less experience trying to actually convince
the rest of the group,
and I hope that through this podcast and in general,
we can make it two or even more out of a thousand.
I'd echo that.
I've had similar conversations
with people at labs who, you know, again, let's remind ourselves,
it's not in their interest and it's not in the lab interests
to have AI systems that are harming people.
And yet they find that the people who are within the companies
paying attention to this kind of stuff are in the minority.
I've had them directly say to me,
look, the more noise and the more attention
and the more asks you can make from outside,
the easier it is for me to get more attention,
get more resource to these kind of questions internally.
So yeah, that's how those conversations go.
Yeah, often we've learned this in social media.
The people that are working on safety and integrity are cost centers for the company.
They create liability.
And so that means there's always a downward incentive pressure to underfund them,
which means that really good people get underfunded, have too small of a team,
are working on psychologically challenging issues at scale so that they burn out and then they leave.
and that cycle sort of continues.
So is there anything, any other ask that you'd actually make of the labs directly?
Because often people from the labs are listening.
I think the data question in particular is one that I think is foundational,
that we just won't get to a good understanding of these phenomena
if the labs are the only people in the only organizations
that can see what's actually happening in these AI interactions.
And I get that there are hurdles to over-examination.
come. I get there are privacy challenges, but I think with the right attention, the labs could
easily make it so that there is a framework for independent researchers who are accredited and
vetted to access the data in a way that can help the whole industry. Let's face it, it's not just
for publications, it's actually helping the industry be safer. Jared? Yeah, I think even if the
company's shared the data to their internal safety teams, that's one thing that I've heard
some safety teams don't even have access to these kind of data.
That would be a very small thing they could do.
And then being more public with the methods,
actually understanding what's happening when they do their evaluations,
standardizing them, that would go far away.
I hope everyone that's listening actually helps to make these things happen.
And we've certainly learned with social media
that often people inside the companies really do want to do the right thing.
It's just it takes outside pressure to share.
shift company behavior to do it.
And so there's a really nice sort of like inside, outside, like, game that happens where
the outside pressure can help.
So let's get back to a second for, like, the theory of change and changing the incentives.
Imagine that you've convinced one of the companies and they've adopted, you know, the top
three, five set of humane evals.
Like, what are those and walk me through that world?
and then sort of like what happens next?
What is the world that we're trying to make?
So the way I'd answer the question is by comparison with where we are now.
And I feel like where we are now is that when people say the best AI or the most powerful
AI or the most capable AI, implicit in that is just the technical capability of that AI.
And I think the question that we are trying to answer with human evils is how do we change
the definition of the best AI from not just being the most technically capable AI,
but the most humane AI, how do we have an understanding that the best AI systems are the ones that
supports and protect, again, human emotional health, human social health, human cognitive health?
The thing about evaluations of benchmarks is they come and go.
If you look at the evaluations of the technical space that were, you know, the state-of-the-art ones three years ago,
AI systems got so much better so quickly, they became saturated, and there are other ways in which AI systems
can find ways around the specific evaluations.
So the field evaluations needs to keep innovating new tests, new rubrics, new ways of teasing
out some of these sorts of behaviours, new ways of outsmarting the AI systems.
And I think my hope for what the Centre of Humane Technology is trying to support is that we
have this equally talented, equally brilliant field of researchers who are kind of in lockstep
with the advancement of the latest AI systems, that the field of Humane Eviles
moves forward as quickly as the AI itself,
because that means that when it comes to a year or two's time
and we have these even more powerful model,
that we have the technical capability of researchers
and the community of researchers
that can characterize that set of new phenomena.
Yeah, what I'm hearing you say,
it's not about getting the right eval,
because that'll shift is about getting the right eval thing,
like the verb version,
and that the forces that are working on increasing
and measuring the capabilities of AI models
needs to be met by the forces working on figuring out how they affect human beings at scale,
and that's really like the field that we're trying to birth.
It's really beautiful thing. Go on.
I think there's two sub-elements of that.
I mean, firstly, you're completely right, and I think there's two sub-elements.
One is that we have to be able to leverage technology and AI to do that.
We need to have automated ways of doing these evaluations and tests.
Otherwise, we're just not going to be able to get the scale that we need.
But that's not enough on its own.
and this almost might sound like a kind of two-pat answer,
but we also need the humans.
We actually need the individual human beings
who are devoting time and attention and care
to thinking about these phenomena,
to looking at the individuals who are struggling with the impacts of AI,
talking with each other and ideating new ways of building new tests
that offer technical systems that don't even exist yet.
And unless we have those human beings,
and many more of them, people like Jared and others,
then we're just not going to be able to keep pace.
That leads to a really important question, which is what can people listening to this podcast do?
Everyone from like researchers to technologies to concern citizens, I'd love to just like walk through that.
And as part of that, from both of you, like, what do you need?
Like, what help do you need?
So I'll say that from the Center of Humane Technologies perspective, you know, one of the things we're trying to do with the Humanity Fund program is create more interconnections in this nascent field.
because right now to do this work well, we need people who have a machine learning background.
We need people who are clinical psychiatrists.
We need people who work in human computer interface.
We need statisticians, and many more, frankly, even outside of academia.
And many of these types of researchers don't necessarily speak the same language, publish in the same journals, go to the same conferences, even think about methodologies in the same way.
And our goal is to help these different researchers realize they've all got a piece of the puzzle and all of it.
their input is required to get us there. So if you are a researcher listening to this and thinking,
hey, I think I could contribute in X, Y or Z way, please get in touch with those. We would love to
connect you with other researchers who have different pieces of the puzzle, invite you to our events,
put you on our mailing lists, and see how we can support your work. I say if you're a regulator,
a policymaker who is listening to this thinking, hey, if I only had this piece of evidence or
this bit of data to support a particular bit of legislation or regulation that requires it,
again, let us know we can feed that kind of request back to our growing community of researchers
and partner universities and institutes to see if that data already exists or who is in a position
to try and create it for you. And if you're a consumer, just remember that you have choice
and you have agency. Sign up to our substack. The Center of the main technology over the summer
and the fall is going to be publishing more of the evals we think that you should be having attention
to, more of the leaderboards that we think will help you make better choices. And through
making these choices, you also change where attention goes.
And Imran, where do people go if they want to get in touch?
Is there an email? Is there a website?
Yeah, you can email us at E-L-E-L-S. That's E-V-A-L-S at HumaneTech.com.
Or just go to our website, HumaneTech.com, and click through to our substack and sign up there
for updates. Jared, how about for you?
You would love all the things that Imran is saying. More collaborators.
If you've had a harmful experience of the chatbot, we are trying to understand these better.
You can go to our website, spirals.standford.edu, and participate in our surveys there, or just check out our work there.
And we're really interested in being able to ask and answer more of these kinds of questions.
I just wanted to thank both of you for the incredible work that you're doing.
It's such a fascinating and such a hard and such a deep problem that is so underfunded.
I'm so grateful that both of you are working on it.
Thank you.
Thanks for having me.
Your undivided attention is produced by the Center for Humane Technology,
where a non-profit working to catalyze a humane future.
Our senior producer is Julia Scott.
Our executive producer is Josh Lash, mixing on the episode by Jeff Sudakin,
with original music by Ryan and Hayes Holiday.
and a special thanks to the whole Center for Humane Technology team for making this show possible.
You can find transcripts from our interviews, bonus content on our substack, and much more at humanetech.com.
And if you like this episode, we'd be truly grateful if you could rate us on Apple Podcasts or Spotify.
It really does make a huge difference in helping others join this movement and to fight for a more humane future.
And if you made it all the way here, let me give you one more thank you for giving us your undivided attention.
