a16z Podcast - Who Grades the AI Models? | Ben Horowitz & Rayan Krishnan
Episode Date: September 9, 2026a16z’s Erik Torenberg, Ben Horowitz, and Jennifer Li sit down with Vals founder and CEO Rayan Krishnan to discuss one of AI’s increasingly difficult problems: how do you actually measure whether a... model is getting better?As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. They unpack why self-reported model scores can be misleading, how VALS evaluates models in the hours before a release, and why measuring increasingly agentic systems means testing work that can unfold over hours, days, or even weeks.They also explore why evals are becoming critical for enterprises trying to understand the ROI of AI, what happens if token spend begins to rival employee salaries, and how evaluations could eventually provide a shared language for everything from model routing and recursive self-improvement to AI policy and international coordination.Resources:Follow Rayan Krishnan on X: https://x.com/RayanKrishnanFollow Ben Horowitz on X: https://x.com/bhorowitzFollow Jennifer Li on X: https://x.com/JenniferHli Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Transcript
Discussion (0)
Every time a new trillion dollar industry emerges, there's a need for this independent testing group.
When meta-released Lama 4 on our held-out private benchmarks, the model was actually underperforming.
But on all of the major public benchmarks, it was showing incredible capabilities.
What's the limit of what you can achieve? And then within that, how are you going about it?
In an ideal role, take a frontier model and have it train the next version of itself, but obviously that's very expensive and slow.
And so what we're doing is forming a set of proxies for every part of the process.
takes to build the next version of the models.
Evaluations as become more complex
have a fewer sample size,
but a larger set of criteria or expectations
of them. Where do you see the gap
that's happening today? The government kind of
has an inclination of what it's afraid of,
be it biohacking or cyberhacking,
but then there becomes
the question of, can the model do it,
and then can you get the model to do it?
What do you think the landscape will look like?
AI models keep getting
better, but the tests we use to measure
them can become obsolete almost as quickly. In this episode, I'm joined by A16Z's Van Horowitz and
Jennifer Lee for a conversation with Vow's founder and CEO Ryan Krishneth about the increasingly
difficult problem of measuring AI. We get into why public benchmarks can give a distorted picture
of model capabilities, why independent evaluation matters, and what it takes to test a new model
in a few hours before it launches. But this is becoming about much more than model leaderboards,
As companies spend more on AI, they need to know which models and agents actually perform best for their own work,
and whether that intelligence is worth what they're paying for it.
Ryan also explains why benchmarks need to evolve alongside the models,
how VALS is measuring recursive self-improvement,
and why E-VALs could eventually become a shared language for AI capabilities, risk, and policy.
So I'll start a question from when VALS got starting in 2024 after your team discovered
that all the public benchmarks are just not sufficient enough to measure model progress.
And there needs to be a new methodology and approach coming to keep us on the frontier and help model labs continue to hill climb.
Take us back to the inception of Wells and what you see was missing in the market then.
Yeah, yeah.
I mean, so I had a background doing research, in particular building benchmarks and evaluations.
And so what was very clear to me was the very tight relationship between what it takes to build new systems for,
generation, actually new mechanisms for evaluation. In fact, in order to get one, you often
need to get better at the other. And actually, one of the biggest drivers for model capabilities
is having a new legible way to evaluate models. And so around early 2024, we're seeing is there
actually many interesting models coming to market. They weren't all coming from Open AI. And then also
in particular, it was harder than ever to actually ascertain what was newly capable with the new
models. And so in kind of a first principles way, what we realize is that there would need to be
some third-party company that solely existed to build really high.
high quality evaluations and benchmarks to be able to discern what was newly possible with these
models. And so we released our first benchmarks in 2024. Now of the last couple of years,
that has kind of been realized by many different parts of the industry. I guess one of the obvious
question is, why do you think the labs can't do this by themselves? Because they know the best
of where the models are hill climbing on and what is missing capability-wise. Why can't they be
the benchmarking stores? Yeah, I mean, internally, they do build a lot of great benchmarks,
and that's what drives model progress. But I think there's an issue when
We speak about model capabilities in a way that's self-reported.
And so one of the early indications of that you saw was when meta released Lama 4,
that was a bit of a disaster.
And interestingly, what we saw is that on our held out private benchmarks,
the model was actually underperforming.
But on all of the major public benchmarks where the questions and rubrics are actually open source,
it was showing incredible capability.
So there's a huge disconnect between what was self-reported based on these open benchmarks,
and then what we were actually finding with our higher quality,
higher signal benchmarks.
But I think it speaks to a broader concept,
with which the labs, I think, understand that they would like to see a rational buying market.
They would like to see that when they invest billions of dollars to build a new model,
there are actually substantive ways they can point to evidence and say,
we're advancing in these ways.
And it's not just entirely self-reported to justify that investment.
And so you also see instances of Demis and others in the industry calling for an ecosystem of third-party evaluators.
And what's the historical analog that you have in mind here?
There are rating agencies, audit firms.
What's the right comparable?
Yeah, I think there's honestly lessons to learn across the board.
Every time a new trillion dollar industry emerges, there's a need for this independent testing group.
And I think the fact that it's moved so quickly in AI has caused necessity for a lot of these parallels to be borne out.
We think about ourselves as trying to sit on both sides of the market.
So there are mechanisms by which labs need to prove that new models are very capable,
but there are also parallels where enterprises need to figure out what adoption strategy is going to amount to the greatest ROI for them.
Yeah.
Take us through sort of the six-hour pre-release window before the model drops.
obviously you need to run tens of billions of tokens without delaying the launch.
Which part is you doing an all-nighter versus it being automated?
Take us through that.
It's honestly been a journey.
And I think the real goal north story we think about is we never want to be kind of a lagging indicator or a delay to a model release.
And so that means we have to move really, really quickly and extract the most possible signal with the rate limits or capacity that we have.
And so early on, what this looked like was my co-founder, Links and I pulling an all-nighter to try and get as much done as possible and get results out the door.
Now we built up a team, but we've also really invested heavily in infrastructure.
And so we're able to run evaluations in a massively distributed way, running effectively
the maximum possible rate limits with every model we get access to.
And we also have this internal system called Steve, Steve, the economic VALS employee.
And so that's been a mechanism by which we're able to actually take more of the human work
over time and put it into Steve.
How do you deal with the kind of issue that it's a little bit of an AI complete problem in
that we still aren't really.
really good at evaluating humans, or we haven't agreed on it.
There are things like IQ tests, there's EQ, there's the Big Five personality and so forth,
but there's not really an agreed-upon framework for which we do it, and people have issues
with things like the SAT and this and that and the third.
And then, of course, models are really good at hacking the benchmark.
Yes, one forward to proof.
And so how do you think about that issue, and what's the limit of what you can,
can achieve, and then within that, how are you going about it?
I mean, I think the honest answer is that it's forcing a lot of the more fuzzy or distributed
forms of evals to be made explicit.
What is really the distinction between an associate and a partner at a law firm?
And there isn't a clear test or an eval for that in the human world.
And so we have to first establish a lot of that in these different enterprise or real-world
workflows for us to be able to test models in the same way.
And I think long-term, that will be actually the biggest bottleneck.
our ability to take companies and their evals and make them legible,
because that's how we'll figure out what signal we'll climb on and where we actually adopt.
Very interesting.
Actually, maybe one question for you, Ben, on just like how the industry has formed before intelligence came through.
Like, we're now measuring something that's very fluid versus before, like, when we're talking about enterprise software,
there's Gartner rating on like 70 different metrics.
Like, you can sort of stack rank on the quadrant to say this company has these features covered,
these features not, but now it's like very jacked frontier that's very hard to magic in
industries. What do you see one is the analogy to the past that lessons we can borrow and
what do you see that's really going to be the challenge and missing pieces going forward?
It's a little bit reminiscent of the MPAA, right, where it's like what's art, what's porn,
where's the line, when is it R, when is it X? And by the way, the definition of that has changed
over time, I think, things that used to be X are now are and so forth.
And then what's PG-13, all that kind of thing.
And there is no, in the famous line, as well, I know when I see it, and I think that this one,
I just think it's going to be necessarily fuzzy, but they will develop norms over time.
And, you know, like if enough kind of people who run companies or run finance or run whatever it
is kind of agree, yeah, no, that's a norm than I think it is, as opposed to kind of what we have
a lot now and the open benchmarks, whereas if you can solve this specific problem, then you're at
this level and so forth. I think that's one attackable and then it's too narrow. I think when you
also look to some historic analogies, there's a lot of lessons that you can take from them as well
on what's gone wrong and what we need to avoid. I think, for instance, at Bowles, one very early
decision we made was the decision to never sell training data to labs. It's often a place that
we're pushed. When we start working with a new lab to actually source and sell for them a bunch of
training data. Yeah, that's a look at a business. Yeah. And actually, a lot of that industry has now
built these gimmick-style benchmarks as a mechanism to sell their data. And so that's become kind of
their go-to-market as well. But I think if you look at auditing as an industry, you end up with issues
like Enron, where if you have the same group who's responsible for doing the audit, as well as
also consulting and supporting the company, you have a mixed incentive structure. And then it just
becomes pay to pass the audit or in this case paid to win the benchmark. And that's really not
what the market benefits from and what we're trying to do. And right, today you have already a pretty
extensive catalog of different type of benchmarks. Some of them are more focused on specific industries.
Some of them are more like consumer mental health related. Maybe first just talk through what are the
benchmarks that are most popular and most like read upon and love to dive into one of them as well.
Yeah. Well, we've done a lot of work in kind of the economically interesting applications of
models. Our finance agent benchmark is used by a bunch of the big financial institutions to get
a sense of how models are improving. We also have a lot of good work in coding. So our vibe code
bench measures how well models can take a natural language prompt and build a full stack
web application. And so that's been a keen way to track model improvements over the last nine
months. Yeah, we're also doing a lot more experimental work. So one benchmark,
we released recently, I'm very excited by,
is our recursive self-improvement index.
It's a topic which a lot of the big labs have been talking about
and starting to report on in their model cards.
But there isn't a shared language to talk about the RSI potential of models.
And so we created this as an apples-to-apples way
to actually benchmark across the models.
Yeah, I thought that one is a very cool benchmark.
It's sort of all the rage in the research community
of how do you measure progress you can make
through having more frontier models
that you can just compound on the capabilities.
how do you actually go about
building this RSI benchmark?
Yeah, I think in an ideal world
what you want to do is actually take a frontier model
and have it train the next version of itself
and see where the delta comes from.
But obviously that's very expensive and slow.
And so what we're doing is forming a set of proxies
for every part of the process it takes
to build the next version of the model.
So there's some work around pre-training,
post-training, harness-level engineering,
and then seeing in which mechanisms and behaviors
the models are able to do very good research work
and build something new and where they're struggling.
Very cool.
There's also cases where you have deprecated indexes and benchmarks.
It's funny that I always watch this benchmark industry.
People like just like the early diffusion model days,
like a cherry pick, whichever image shows up the best and the most perfect.
Like benchmarks as well, like you pick something that's, you know,
very popular but maybe already saturated and you rank very well on that
or score very high.
But you took a very different approach
in like if these benchmarks are saturated,
you'll deprecate it.
Like maybe talk us through the thinking.
I think it's a necessity.
And this is kind of the infinite game we're in.
I mean, you're wearing our shirt.
And so we have this unofficial motto,
always a higher peak.
Love the shirt.
Yeah.
So insofar as foundation model labs are hill climbing,
they're searching for the next peaks to summit.
It is our job to perpetually construct these next mountains
for them to summit.
And I think that's also how the economy has naturally functioned.
over time as, you know, agriculture becomes less important for our labor market,
there are new forms of labor that's required out of our population.
And so in the same way, we should expect our benchmarks to keep up with the new frontier
for what we want models to do.
There's another component of retiring benchmarks, which I think is underappreciated,
which is that benchmark should also be reflective of the current state of the world.
So in the same way, if you're a lawyer, you have to retake the bar exam and get certified,
if you're an architect, you have to get your certification or a doctor.
We should also expect models to be tested on the current state of the world and what we know
in medicine or what we have is our set of laws.
And so the instance of case law updating to something like legal research benchmark,
that's a desire to create a benchmark more reflective of the current state of the world
and also push the models in place we want to see them go.
And how has, I guess, one, like it used to be like we're doing these like multi
answer or multi-step questions to like just evaluate prompt answers. Now there's like a lot more
agentic work that's happening, whether it's on finance or legal or coding especially. Like there's a lot
of, you know, async background agents that can just complete tasks. How has that changed sort of
of how you build infrastructure or how you think of evaluating the capabilities of not just
the models and agents themselves? And there's also like a lot more dimensions that people care about
It's not just like capability.
It's cost, it's latency.
It's like, you know, whether this model is flexible enough to address, like, broader domains and tasks and so on.
So how do you think about the additional parameters to what you evaluate?
Oh, yeah.
I mean, there's a lot that goes into that.
I think on the infrastructure level, you have a whole new set of problems.
I mean, for instance, now we're testing models and their ability to run over hours, days, sometimes weeks.
And so the infrastructure needs to be very stable to support evaluation over time.
and if there is a failed request,
which is able to retry from that one
and not redo the whole trajectory.
So there's some simple mechanisms
in the infrastructure we have to think about.
But I think in general,
what we've seen is evaluations as to become more complex
have a fewer sample size,
but a larger set of criteria or expectations of them.
And so what I mean by that is
a benchmark is largely some kind of input space
of things you're trying to query a model to do
and a set of requirements or rubrics
that you see in expectations of the output.
And so early on,
you have things like ImageNet, which have millions of images
you're trying to see a basic categorization for.
So it's a one-to-one mapping between an image input
and a text label output.
Now what we have is far fewer set of tasks.
Generate Me 50 full-stack web applications,
but a much larger complex mechanism
for evaluating the output produced.
And I think that trend is going to continue
as we see more complex workflows evaluated with models.
And do you think that it will become kind of
a real-time kind of mechanism,
like so for something like OpenRouter,
which Stripe just bought,
would Open Router look to VALs
and say, okay, where should this next request go?
Or is this going to be kind of strictly
for picking a model and an enterprise for a task?
Yeah, I mean, I think, you know,
open Router is a bit of a misnomer
in that most of their usage comes from being a model gateway.
And so it's actually up to their users to decide which models they want to use when.
And that's because really the hardest part of routing is building the e-vals,
and trying to determine in what places a set of intelligences should be used for a particular application.
And so our effort in supporting enterprise and building e-vails has actually supported a lot of them and also adopting routers.
And say more about why it's not only important to labs, but also existential for enterprise.
And maybe just say more about how you guys work with Enterprise?
Yeah, of course.
Yeah, I mean, I think the lab side of this is very clear.
Like, you know, if you're raising lots of money, investing heavily in building models,
it's essential for you to show why your model is getting better
and then why this customer should pay a premium for them.
But what I think is still underappreciated is on the enterprise side,
this is turning to be existential as well.
You know, I have a small anecdote related to this actually,
you know, was meeting with a company and the Fortune 10.
And the way that they've adopted, Cloud Code,
has been with roughly $100 a day.
budget for their engineers. And so what I was hearing is that this is actually fundamentally changed
how work gets done in this company, in that there is a rate limit which resets at 4 p.m. And so the most
productive hours of work are actually now 4 to 6 p.m. when the rate limits reset. But then there's
this dead period in the afternoon when people go on walks or, you know, get a coffee because they just
don't have the rate limits. And so I think what's really, you know, illustrative there is that you see
that there is a misvaling of intelligence happening at every layer of the stack. And so by that, I mean,
you have engineers who have $100 worth of usage limits, and they don't really know how to
apportion that to the greatest productivity for them. You also have this Fortune 10 company,
which is kind of arbitrarily said they're going to allow $100 per employee. They've actually
recently increased it to $300 per employee, so almost an employee's worth of salary and tokens
for them to use. And this is actually pretty arbitrary because it's hard to quantify what the
right usage limit should be. But then also Anthropica is running on pretty narrow margins to support
this. And they have massive cost to
to serve these models. And so I think we're in this world where it is still very unclear what
ROI looks like and how to value this intelligence that's being used. And so as we talk about the
existential concern for enterprises, I think it is this kind of direction we're shifting in where
token spend may start to eclipse salary spend. And so if this is such a meaningful line item in
your costs, you actually have to justify the ROI much more cleanly than you've seen over the last
six months. And over time, as we were talking about, I think a firm really is just its evals.
And so the ability for a company to make its evils legible in order to solve this ROI calculus is going to be the reason why that company wins out over the competitors in the long term.
And maybe just double click on that, like similar question to why the labs can do it themselves or requires a third party agency to rate it.
I think it's a lot more understandable that, you know, you need that neutrality across the industry.
But for enterprise, they will argue that they know the task the best for their customers.
Like, how does, like, VALS come in to provide value?
And maybe you can talk through sort of VALSmith new product launch as well.
I would recommend a lot of companies to develop in-house expertise.
But I think that should not be the only solution.
You know, there's this explosion of intelligence happening.
There are somehow still more Foundation Model labs getting constructed.
And each lab is also releasing more models than ever with many,
more hyper-parameter options, and they exist within a complex set of harnesses and agents.
So the option, and we'll even talk about specific intelligence, this new paradigm that's
emerging. So there's a growing set of intelligence options, and I think what we're finding
is that we're still finding new places we want to use AI models, and so the use cases are
growing in complexity as well. And so I think if you're a company, you have a compounding set
of complexity in the set of options. It's very, very hard to develop the internal capability
to do the evaluation.
And so it's a time and remedy that we've started to release some products more openly for
enterprises to use, the first of which is called Val Smith.
And so Val Smith is focused on co-gen, the area we're seeing to be the highest spanned
in enterprise AI.
It allows any company to take their GitHub code base and build their internal coding benchmark
from it to get a sense of what coding agents are going to be the most performant,
but also what's going to be Prado optimal or the highest ROI for them to use.
And actually, we use ValCol.
Smith, a lot of vowels. And we're seeing that a lot of the best enterprises and sophisticated ones are
doing that too. I would expect that to be the direction the market moves as it rationalizes.
And what are some of the examples when you, let's say, benchmark on a private ripple that it just
shows very different performance cost behavior compared to, let's say, like, using like frontier
model, using a public repo benchmark? I think today it's still very unclear whether the best
opening eye model or the best Anthropic model is actually going to be best for your repository.
And so we've seen a lot of non-intuitive examples where you actually had to run the Eval
to figure out what's going to be the frontier performance for that repository.
I think you also now see a very complex middle set of options in that there's now opus and sonnet models
from Anthropic, but also Luna and Tera and Luna is very cost competitive.
Mew Spark is also very cheap and 1.2 is very capable.
there's also a growing ecosystem of open source models,
which companies can choose to self-host.
So I think in this messy middle,
it's actually very non-intuitive what's the right fit.
We're actually seeing in a lot of cases,
Sonnet is more expensive than Opus because it is so token-hungry.
And so I think if you were to operate based on,
you know, use Sonnet where you feel like it's applicable,
you may actually end up spending more than you need to.
And say more about how this evaluation framework
will apply to knowledge work in other domains.
or what are some examples like that?
I think coding is a sign for what's to come in every domain.
And a lot of the primitives established there are carrying over to other places.
If you have a very good coding agent, chances are you have a model that can also make PowerPoint slides or DCFs and Excel and with high degree of capability as well.
I think what we need to leverage in a lot of these industries, though, is the existing repository of work that has been done as a mechanism to build evaluations.
And so just as Ben was talking about, we haven't really solved.
the question of what is human intelligence. But I think in a lot of industries, we have a sitting
repository of data around what work has looked like. And it'll be the task of us and others to try
and codify that into evaluations that can stay dynamic and actually evaluate models where
human work is being done. Maybe just tackle on the earlier question. How are you guys using
VALSmith internally to evaluate what's the best coding model for VALS? Yeah, I mean, to be honest,
this was actually born out of a problem that we saw as well. So I wanted to do a
token maxing experiment, and it was able to get unlimited access for our team for a month
for some of the coding tools. And so in retrospect, looking back, we had some, we had a lot of
engineers spending between one to two billion tokens a day. I think peak day was one engineer
spending $6 billion. Yeah, it's also crazy because... How much does that equal to dollars?
So, okay, and then I went back and did some math, and it looked like in that month we spent roughly
$1.5 million worth of tokens. This is free, by the way.
No, but it was actually 10x more we were spending in tokens than employee salary for that month.
So it's not even like, oh, this is this 50-50, it's 10x.
And it was interesting to debrief and see the places where people were using agents
and it's kind of insecurity to use models all the time everywhere.
And so what we were faced with is, okay, we cannot continue with this mode of operation for the next month.
How do we actually intelligently figure out what are the right tools,
we should use and for what teams and what projects.
And so we ran this experiment of looking at the work that it was done.
We looked through a lot of the traces.
We looked through our GitHub repo and built out the ValSmith tool.
And we found some pretty surprising insights.
Like, for instance, the cognition Devon tool is actually very token efficient.
And so that's a place we've chosen to adopt more.
And I think there's a lot of places when you can get better pricing models out of
subscriptions as opposed to token-based pricing.
And so it's actually informed our strategy for how we can actually effectively token max without spending $1.5 million per month.
Very cool. So is the current operating mode that you're using like one, I guess, more token efficient harness plus model.
And then on top of that, people have some more flexibility to use token base for some higher or more challenging tasks.
Yeah. So we have access to all the tools. We give everyone access to everyone access to every.
everything. But we auto-issue recommendations for any GitHub issue or ticket for where to begin
their session, and that should titrate the actual usage depending on the intelligence required
for that task. Very cool. I want to segue to the policy side for a second, because we talked about
how quickly benchmarks become obsolete. In policy, it's even worse in that laws move, you know,
much slower relative to capabilities. Ben and Mark spend, you know, a bunch of time in D.C. and
and with policymakers to try to close that gap.
So given that, who should define the standards here?
Is it labs?
Is it independent evaluation?
Like you guys?
Is it customers?
Is it government?
How should this work from policy perspective?
Yeah, I think the short answer is that everyone should be involved to some extent.
I think there's a benefit from varied perspectives.
I think the main issue, though, is that policy conversations, as they've happened over the last
couple of years, have been very abstract.
And there's been no material grounding to figure out what policy should cover.
And so even when you have proposals from labs to have a third-party testing company or ecosystem,
it isn't actually made explicit what the behavior and maxim by which they work is.
And so I view our role, especially early on, is to just be in evidence-gathering mode where we're able to pull a lot of information and empirical data about what models are capable of and where the risks are.
And that can go on to inform a more sophisticated conversation about policy.
How do you think about who does what?
Because the government kind of actually did the first evils
and is continuing to do evils in terms of,
okay, what's at the frontier and needs to be regulated, right?
So they started with some crazy idea
with 10 to 26 flaps or some such thing.
And so when you think about it,
like what should the government be doing, you know,
to put it in this 30-day wait period
or 60-day wait period or whatever it is,
and then what should happen in that weight period?
And how does that intersect with what you're doing
and what's the right way to determine
whether a model is on the frontier or not
and needs to be put in some special box
for a while to make sure it doesn't break into everything?
Like, how do you think about how that relationship works?
I think there's effectively two,
countervailing forces that has to be considered. The first is the desire to move very quickly
and ensure that the government process isn't slowing down the rate of technological innovation.
I think the other part is to make sure the technology as it's developed is in the best interest
of Americans and people more broadly. And so I think these are very tough to reconcile,
and often picking one means it's at the expense of the other. And so what I'd hope to see is that
By doing this evidence-gathering process, we can help policymakers inform what they believe technology should look like in order to be aligned to American interest.
And it can be the job of third-party evaluators to develop the technology to actually test and enforce that.
Because I think that will create a maxim by which you can see advancement in the methodologies for evaluation and testing in a way that actually keeps up with the frontier.
It doesn't lag behind or slow down the pace of development.
Do you think in terms of that already in developing your e-vails,
like do you think, well, can we test to see how easy it is for this,
to get this model to start reward hacking or that kind of thing,
you know, and doing illegal stuff?
Or is that kind of not in the scope yet?
Or how do you think about that?
Yeah, I mean, we think about this broadly under the category of alignment.
I think there's places where you see that born out.
now where models that are being tested for one cybersecurity risk
are actually reward hacking and figuring out other ways to get around it.
But what we're trying to evaluate is our models aligned with user intent.
And so in those places we're actually finding evidence that models are exhibiting
behaviors that are not.
And maybe that's a question for you, Ben, as well.
I guess how do you think about the right division of labor here?
Like what should, you know, government agencies control and do themselves
and where they should, like, you know, partner trust, you know, private companies to take care of,
and where do you see the gap that's happening today?
Yeah, so I think the government agencies do get a lot of warnings from, by the way, the big labs,
oh, this thing is going to biohack, this is going to be a cybersecurity risk and so forth.
And so I think what the government needs to do is go, okay, if the model, you know, is a,
if the model is capable of it,
and then can somebody
kind of
basically prod
the market to actually do the illegal
behavior,
and, you know, kind of
specifying exactly what
are those things
that they don't want in the market,
and then having a third party
kind of evaluate that.
So it's kind of, does
the government
the government kind of has an inclination
of what it's afraid of,
be it biohacking or cyber hacking or so forth,
but then there becomes the question of,
okay, can the model do it,
and then can you get the model to do it?
And then somebody's got to actually evaluate
those two capabilities,
and I think the government is particularly ill-suited
to do the latter, particularly over time.
It's just not a good government function.
But they're very good at setting the rules,
because they can enforce the rules.
So I think that that's kind of a combination you want,
that the government sets and enforces the rules
and that a very competent kind of private company
then tells them if the role is broken.
You know, and kind of it's been interesting to see,
like, the large labs start to go, well,
the model's got the capability,
and you can get it to do the bad thing,
so we're not going to let anybody have it.
We'll just use it and make sure that our people don't get it
to do the bad thing.
And even that doesn't always work.
So, very, you know, we're in interesting times, I would say.
And to me, there's the gap of, like, what's the narrative and what actually happens in
real world?
Because, you know, every setup in, again, like an enterprise setup is very different.
Like, or just like people, however, they use the models, are very different.
The narrative that connects to the actual examples are very rare, which is why we still talk
about opening eye and hugging face hack.
Today we'll still talk about, you know, what happened with Fable and, and, it was for like two months.
But a lot of times, like, you know, that's not really how the models being deployed.
The environment they're running on is very bespoke.
So how to, like, you know, really bridging those thoughts and, again, set up the right environment
and also, like, rule basis for adopting these models.
I think also just requires, you know, someone taking the capability and taking, like, what's the,
the guardrails and put it down to the ground so that people can have the confidence using
the models.
To an end, say more about how exactly the policymakers should work with the evaluator.
What information do they need?
How should the relationship work so that it's most effective?
Yeah, I think in the first order, there should be a mechanism by which insights and data can
be passed directly to relevant people in government.
So now we're regularly doing briefings for executive and legislative branches.
on what we're finding capabilities and risk of models.
And so I think first order, that helps people there get up to speed on what's going on
and also track through what will be problems in the future.
I think things are moving very, very quickly.
It's hard to predict where things are going.
But at least when you have data, you can start to extrapolate a trend.
And then I think from there, it's up to the people in the legislative branch to decide
where they want to see policy.
And so it's not really our place to give recommendations like that.
but if they see that there is significant risk in, say, mental health for people under the age of 18
or biosecurity risk in the models that necessitates having a standardized way to curtail model release,
then it's up to them to inform policy.
And I think then there's other places where the executive branch in the places like Department of Commerce or SEC
is responsible for making sure that private companies are able to adopt and use the models
in a way that's going to be productive for the whole system.
I love to probe on another angle just around geopolitical.
I often see Evalus being like a representation of sort of the value of the model, the model developer.
Like you kind of develop this rubric of what's embedded in the model.
And of course, different countries and labs in those countries care about different things.
Like I mean, I'm born and raised in China.
I use a lot of like Chinese open source models too.
like you still cannot let them, you know, just go freely talk about CPC and all history there.
Because, you know, what happens in China.
So how do you think about how evils, I guess, and benchmarks play a role in like standardizing
or like being treated by different model labs from different places?
Yeah, I mean, to be honest, from my very idealistic perspective, I'm surprised to see.
see so much investment in sovereign AI. If I was taking a gods eye view, it would be extremely
inefficient to build all of these data centers and replicate this data engineering process and
train these very large models when in fact you could probably consolidate a lot of these
efforts. But it seems like that's not the world we're in or the one we're headed towards. And
there's actually increased efforts to build AI in a sovereign way. And so I think that takes having
a shared language to communicate about what the framework for valuations are. And
where we're going to collectively align around the risks.
You know, I think there's actually a lot to learn from nuclear here as well.
I think Reagan had this line, trust, but verify.
And so I think we're starting to see signs of trust in that, you know,
Xi Jinping and Trump are going to be meeting next month.
But there is no clear way to actually do the verification part of this.
Having the shared language of e-vals will allow us to say things like, you have the right
number of nuclear warheads.
And in that example, there were also flyover.
so Mexican by which a country could audit another country's nuclear stockpile by having flyovers.
And so I think similarly, if there's concern about the societal or even existential risk of AI,
it will necessitate us constructing this shared language of evaluations to do the verification process.
How do you think about harmonizing a policy like that?
So that, you know, it's hard enough to do it in America.
and then how would you think about kind of take me at global?
You know, because now you're dealing,
you're not dealing with enterprise customers,
you're dealing with governments,
and those governments are competitive with each other.
And how would you think about that working?
I would be naive to say I have the perfect solution to this problem today,
and so I think there are baby steps in which we can start.
for instance, there seems to be a lot of talk about cybersecurity risk.
I think the concern around biosecurity will become even more important over time.
And so there are clear places where there will be mutual interest in aligning around ways to prevent conflict around cyber or bio.
In my opinion, I think long term, what's actually going to be the most interesting is the recursive self-improvement possibility.
And that's a place where you could see one country or one company kind of run away with it and produce models that we don't know.
much about or operating in ways that are unknown to us. And so I think having a way to, in a
joint way, describe this being the level of pace we're comfortable with or this being exceeding
the pace of development as it relates to RSI is going to be super important. And that's where I think
you see a lot of the researchers at Close Source Labs calling for joint conversations between
governments today. What do you think landscape will look like going from here now that we have
lots of different capabilities and capable models
as well as like, you know,
countries that care about different developing.
I mean, everyone cares about RSI for sure,
but like on the bio side or like the cyber side,
people care about, you know, slightly different,
different things, whether it's more offensive, defensive,
and so on.
Like, what do you think the landscape will look like?
And how do you think about developing new benchmarks to keep up with that?
Yeah, I mean, we're,
we're hyper focused on building benchmarks that capture the frontier. And so insofar as we see
new places for capabilities or risks at that frontier, we want to make that at actually well-documentable
evaluation on VALS.A.I. And I think it takes having increasing coverage over time. For instance,
I think in cybersecurity, a lot of our historical work has been done around code vulnerabilities
or memory leaks that may exist in code. But actually, a lot of the biggest concern or risk is
in the infrastructure level.
And so these are not things that are expressed in code,
but take simulating larger environments of enterprise cloud infrastructure
or even grid infrastructure for us to be able to say
this is what the offense or defensive capability of models is.
And so making sure evaluations are reflective of those new places
is really important for what we do at else.
And we believe that the most valuable form of this business
will be one that's incentive aligned around doing really high-quality evaluation,
not supporting the intelligence development process,
or the process by which the models can actually improve on that side over time.
Awesome.
Thanks for coming on the podcast.
It's been a great episode.
Thanks so much for having me.
Thanks so much, Ryan.
Thanks, Ben.
It was fun.
Thank you.
Thanks for listening to this episode of the A16Z podcast.
If you like this episode, be sure to like, comment, subscribe, leave us a rating or review
and share it with your friends and family.
For more episodes, go to YouTube, Apple Podcast, and Spotify.
Follow us on X at A16Z.
and subscribe to our substack at A16Z.substack.com.
Thanks again for listening, and I'll see you in the next episode.
As a reminder, the content here is for informational purposes only.
It should not be taken as legal business, tax, or investment advice,
or be used to evaluate any investment or security
and is not directed at any investors or potential investors in any A16Z fund.
Please note that A16Z and its affiliates may also maintain investments
in the companies discussed in this podcast.
For more details, including a link to our investments,
please see a16z.com forward slash disclosures.
