Software Huddle - Open Source Self-Improving AI with Vignesh Baskaran
Episode Date: August 18, 2026Today we are talking with Vignesh Baskaran, the CTO and co-founder of Hexo Labs, about teaching AI agents to improve themselves. Vignesh has been training neural networks since 2012, back when he was ...still called a data scientist. Then he became an ML engineer and now an AI engineer, though he says the underlying work has never really changed. It's to figure out how to make a system behave the way you intend it to. He built the litigation search engine that Google itself became a customer of, and now he's chasing something new, agents that rewrite and retrain other agents without a human in the loop. We dig into Sia, the meta-agent at the center of Hexo's research, and why improving an agent means touching both its harness and its actual model weights, not just one or the other. We talk about proxy evals for when you don't have much to ground truth. The Darwin-Gödel machine and why formal verification is too strict a bar for anything commercial. How Hexo's work echoes DeepMind's Alpha lineage from AlphaGo to AlphaEvolve, and the spectrum from clearly verifiable to totally subjective tasks? Why VAE evals are quietly wrecking agent quality across the industry, and a great story about an agent that discovered a customer's own eval file was silently corrupted, something buried in hundreds of thousands of traces that no human would have caught.
Transcript
Discussion (0)
Hello, everyone. Welcome to Software Huddle. I'm Sean Falker, and today we are talking with Vignish
Vascaron, the CTO and co-founder of Hexel Labs, about teaching AI agents to improve themselves.
Vignish has been training neural networks since 2012, back when he was still called a data scientist,
then he became an ML engineer and now an AI engineer, though he says the underlying work has never
really changed. It's figure out how to make a system behave the way you intended to. He built the
litigation search engine that Google itself became a customer of, and now he's chasing something new,
agents that rewrite and retrain other agents without a human in the loop. We dig into SIA, the meta-agent
at the center of Hexo's research, and why improving an agent means touching both this harness
and his actual model weights, not just one or the other. We talk about proxy e-vals for when
you don't have much ground truth. The Darwin-Godell machine and why formal verification is too strict a bar
for anything commercial. How Hexo's work echoes deep minds Alpha lineage from AlphaGo to AlphaEvolve
and the spectrum from clearly verifiable to totally subjective tasks. Why Vibe devals are
quietly wrecking agent quality across the industry and a great story about an agent that
discovered a customer's own Eval file was silently corrupted. Something buried in hundreds of
thousands of traces that no human would have caught. And with all that, let's get into it.
Fignish, welcome to Software Huddle.
Hey, hi, Sean. Good to talk to you.
Yeah, thanks so much for being here.
So I was digging a little bit into your background.
And I know you previously built an NLP team.
You went through an acquisition.
Now you're doing kind of frontier software improvement research.
I guess give me a little bit of background on, you know, what was that journey?
What led you from applied NLP to now self-improving systems?
You know, Sean, like I previously, I, I, uh,
built a search engine, right?
And I built a Google Patents litigation search, right?
I used to work for a small startup, but we punched way above our weights.
And we managed to convince Google to use our litigation search, and Google was one of the largest
customers, right?
And I started, I trained my first deep neural network in 2012, right, when AlexNet came up.
From there on, I have been training neural networks and building machine learning models.
Initially, I used to be called data scientists, then ML engineer, then now A.A. engineer, whatever, right?
Like, the roles and responsibilities have changed.
And similarly, I was like a data scientist, principal scientist, now a CTO, all these things.
But fundamentally, the one thing that I do was figure out a way by which I can replicate
the behavior of a system with what I have in my mind.
So I have a dream about, hey, this is what the system is supposed to do.
And then I figure out ways by which I can make the system do that task.
Back then when it was typically traditional machine learning, I used to be like, okay,
now I have a model.
I would like to train this model so that it behaves in a way that I anticipated to behave.
Then how do I do it?
I play with the training data, I play with the evals, I play with the optimization algorithms,
I play with different kinds of hyperparameters to make it work.
And then now in this agentic AI era, I do the same thing, but it's more for a LLM,
which is wired as an agent, right?
How do I make sure that this agent has this singular responsibility?
And it exactly does what I intend to do at any point in time.
So the same principles of making sure that probably not that training,
data per se unless we are training the agent but the evals the tools that are associated to it the harness
the weights the scaffold how do i optimize these things to make sure that the agent is allying to that
of my interest is what i work on so my work is always the same how do i make sure that a mission does
what i anticipated to do yeah so so the goal has been the same through that journey you talked about
you know training uh your first known networks back in you know since 2012
And then you've also gone through a rebranding exercise of data scientists, ML engineer, the AI engineer.
But if the goal is the same, and probably a lot of the technology has changed in some fashion,
but not drastically, like, what has changed in that time?
You know, now we're almost 15 years later.
Has there been a fundamental shift in the type of work that you're doing in order to try to
accomplish that goal from, you know, back in 2012?
No, Sean.
Nothing has changed much, right?
I don't think, see, there are new things that we modify.
Like back then, when we were doing traditional machine learning, we used to create features.
We used to create handcrafted features so that your model can do what you anticipated to do.
And once that went on and then we started working on raw neural network,
then most of the time we spent on fixing the training data, fixing the Eval set,
and then looking for new optimizers so that we are,
able to hill climb towards that metric.
And right now with LLMs, it's more control plane,
like the harness, the tools, the memory,
and all these different features.
But if it all boils down to this thing,
hey, how do you make sure that an artificial neural network
is able to do what I intended to do?
Back then, so back then the only thing that you could change
was the features.
So you used to do feature engineering.
then you could actually change the way how the model is trained itself.
Then you do model engineering.
Now, you can do harness engineering, tool engineering, memory engineering,
all these different kinds of things.
So the number of layers has only consecutively improved,
but at the end of the day, it's the same work that I do all over again.
How do you know which layer it makes sense to try to adjust
in order to actually improve the output?
outcome of what you're trying to do accomplish with an agent or a model?
I don't know.
That is the truth.
No, no, no, Sean.
That's the truth, right?
I don't know.
No one knows.
Even Darry or Sam or Ilya doesn't know.
No one knows, hey, for this agent or the model to do what we anticipate out of it,
what should we change?
No one knows.
We might have some gut feeling.
We might develop some intuition around how to do it.
But the right way to do it is by doing a.
following a scientific process.
Right.
So it's all very methodological science.
Okay.
It's basically like a think of it this way.
Right.
My wife was a researcher in Stanford.
She works on orthopedic surgery.
Right.
So they try to discover the truth about, hey,
what is the way by which I should create a drug
so that I can,
make sure that a certain type of orthopedic disease does not happen to a patient.
So they have a scientific process like, hey, okay, this is how I'll build my baseline.
These are the control variables that I'm going to do.
I'm going to do randomized control trials.
I'll try these different sets of experiments to do it.
So it's a scientific process that you can follow in order to discover what is the right way
by which you modify your agent to find the ultimate truth.
So the scientific process is the base.
Now, how does the scientific process work?
The scientific process has four steps, roughly four, five steps.
One is you take up a problem statement and you come up with hypothesis, that, hey, in order
to solve this problem, these are the different techniques that I can use to solve it.
And then you go ahead and implement the experiments, run the experiments, analyze the results,
learn the truth, which hypothesis worked, which did not work,
and then generate new hypothesis,
and you continue this loop over and over again.
So it's a very scientific process that you do to discover truth.
Yeah, you know, there's been, I think, people trying to do things like,
you know, self-improving machines for a long time.
Like there's research going back many, many decades into this.
And even things like Godell's Godell machine was, you know,
trying to essentially only rewrite when it was mathematically proven that the rewrite is better.
Do you think, you know, you talk the scientific method, is that kind of as close as we can get
to having like a mathematical proof for whether the change is actually leading to a better outcome?
Essentially, we end up having to form a hypothesis, run an experiment, tested against eval set sets,
and then decide this was an improvement or wasn't an improvement.
but can we actually have a mathematical proof to say that?
Yes.
So there are two ways by which you can prove self-improvement, right?
One is a theoretical proof where you derive, you write theorems
and do a rigorous theoretical analysis on whether that is feasible or not.
The other one is empirical, right?
Hey, okay, sure, even if you're able to, so for example, to prove theoretically.
So the reason why Godilmissions was very limiting was because of it,
any proof that it generates has to be formally verifiable.
And in today's world, like, whether you can formally verify something or not is also a big question.
Like, verification is a spectrum, right?
And what is, so you'll be hearing these words, like, is it easy to verify, hard to verify,
proxy verify, all these different kinds of things, right?
So what Godil machine tries to do is, can I use formal verification to verify whether a system is actually self-improving or not,
which is, is it better than its previous generation?
If you loosen this up, that's when the Darwin-Gaudel machine paper came up and said that you don't have to formally verify,
but you can just empirically verify by running experiments and seeing whether it works or not, right?
That is useful enough already for a lot of commercial applications.
Like, you and I don't want RSI to come and change our lives.
Even if there are small harness engineering loops or something that we can build for ourselves to automate and improve the work that we do,
that is more than sufficient.
That's where empirical proof comes into picture, right?
So the work that we do at Hex-O right now is more on making sure that empirical proof is attained,
but there is a small team of researchers that we have funded in Oxford who work with us
on pursuing towards this direction of identifying a theoretical proof for recursive self-improvement.
But it's a small effort that goes on independently in the company
without the major business being affected.
Right.
Do you think other than the, I don't know, like the goal as a scientist or a mathematician of achieving that, is there value that you would get from having a like provably, be able to prove essentially correct that things are improving versus having the empirical evidence?
Like from like a commercial standpoint of like we're building a product, does it actually add value or is it more this is a like a problem?
that we'd like to solve as a human, but not necessarily something that necessarily delivers
higher value right now.
So, Sean, these are kind of two independent scenarios right now in our viewpoint, right?
That's how the world sees it.
If you're able to theoretical prove something, sure, but without having an empirical proof,
I don't care about it, right?
Commercialization.
So I see it as different stages of human progress.
Okay. As a civilization, you would have heard this example million times, which is basically like, hey, we discovered x-rays even before we, someone had thought that, hey, okay, this can be used in order to do photography of your bones.
Right. So what happens is you make technical breakthroughs, sure, but then it takes humanity quite a lot of time to figure out, hey, what is a commercial application that I can build along with this so that I can see.
solve a real world problem so I can get paid for it.
So one is scientific research, the other one is engineering problem, right, where you
identify what is a product, what is a market, is the, can I tweak the fundamental
learnings that I've had in order of fundamental learnings that I've had by conducting this
research so that I can build a better product to solve a problem in a novel way.
Right?
So these are two, these are two disparate efforts.
Both has to happen.
It's non-optional that only one of these two things happen.
So it's like a cycle, right?
It's a super cycle, KAPX, OPEC cycle.
You have to do a lot of capital investment in doing research
so that you have something to sell later, right?
Yeah.
Yeah.
I mean, I think there's tons and tons of examples of that.
Like, even if you look at something like Boolean Algebra was invented in the, like,
I think the early 1800s or something like that.
And I mean, it's hugely, hugely valuable in the computing world,
but that didn't happen for 100 years plus.
Well, let's fast forward to today.
So you're working on, or let's talk a little bit about SIA.
So can you explain to somebody, like essentially that's listening, what is SIA?
How would you describe it to a friend as maybe not an expert in this space?
Sure.
Should I share screen and explain?
Or can I just?
Yeah, you could.
Yeah.
Okay.
So first let me start with,
okay, let me give a layman explanation first.
Right.
Then we can go into a more deep dive as well, Sean.
So all these agents that we use, let it be chat, GPT,
cursor, cloud code, all these different agents that we use
in order to get some work then, right?
Was actually built by an engineer.
Right.
So they follow similar principles as what we have always followed
for software engineering.
It is essentially like, hey, I will do a test driven development.
This is essentially, this is a software that I would like to build.
These are the scenarios in like I would like to test it to make sure that it works effectively.
Once I list down the scenarios, I'll go ahead and build a software, run the software,
and figure out where does it break, where it does not break,
and I'll continuously keep doing this until my software is able to pass all tests that are required.
So I can ship the software.
So a human is actually doing this process of writing these test cases and then writing the software and then running the software on these test cases, figuring out where the software is breaking, which test does it fail, modifying the software and then rerunning the test cases.
So this is a loop that happens over and over again.
So there is a human who used to do this.
Today we have back in the old days.
Back in the old days, right?
And then what happened was now we have agents which are able to write software.
That, hey, let me go ahead and write.
Okay, this is the software that the user wants to build.
So let me go ahead and first write a set of test cases.
Let me write a piece of software.
I'll run the software.
And I'll keep on doing this process over and over again.
Your cloud code or your cursor or one of these agents is the one,
which is actually doing this job and shipping the software.
right but then if you think about it this agent is also a piece of software the agent which is
writing your software is also a software right so now someone is sitting and taking this agent
writing test cases for it writing one version of your agent agent and then running it and then
figuring out where this agent fails and continuously doing this process it's kind of weird that today we
have figured out a way by which we can automate software development with agents. But given that
agent is also a software, we have not figured out a way by which we can automate this agent development
itself. And it's still extremely manual in nature. So this is where we came into picture. And we are like,
hey, okay, we'll figure out ways by which we can automate this process. Is my screen visible, Sean?
No, I don't see it. We see it now? Yeah, there we go. Okay. Awesome.
So you're showing it looks like a scientific paper that you're
Exactly. This is our research paper that we released.
We were the first in the world to discover a way by which we can improve an agent
autonomously, not just in terms of its harness, but also in terms of its weights.
That is a main view prop of this paper.
Now let us look at this.
So this is what I was explaining to you.
Let's take up a use case, okay?
I want to build an agent to do a customer satisfaction survey.
So I want this agent.
So now Wignesh is selling some product.
Wignish wants to build an agent that will go ahead and autonomously make calls to his customers
and call them and ask him,
hey, you have bought Vignish's product?
Is it good enough or not?
If not, what do you think?
How do you think we can improve?
If yes, then can you go ahead?
and rate it on Play Store.
Simple.
Okay?
This is like a customer outreach agent.
Let's call it as customer outreach agent.
So that is the task specific agent that you would like to build.
Now, what do you do is you go ahead and build this agent and you have to run it in an environment.
The environment could be like the telephony that it has to be attached to.
So it can give a phone call.
The web search in case if it has to do some web search to understand what are the products.
that are available that Vignesh is selling and all this information and to go and mine some
Twitter sentiment around Vignish product, all these things.
So that is what your environment is composed of, all these different tools that are available
to your agent to perform the task that it wants to perform.
Now, the way how it happens is there is a meta agent, which is our proprietary agent,
it has nothing to do with Facebook meta.
Meta is just a nomenclature that people use in order to say that when one agent is operating
on another agent to improve that agent, then you call it.
as meta agent. So we have a meta agent which will autonomously take your task specific agent,
run it on the environment, look at what are the ways by which, what is the way by which your
task specific agent is working and then autonomously improve this task specific agent, which means
it might be like, hey, okay, I see, I took the, I created some benchmark data set. The task
specification and the verifier comes from a benchmark data set, which is like you're simulating
some scenarios in which your agent can have a conversation with the customer.
Like in one case, the audio of the customer might be poor, okay, that your task specific
agent is unable to hear.
Then your task specific agent is supposed to tell the customer that, hey, your voice is
unclear.
Would you move to a place where there is better reception?
Right.
Instead of continuing the call in noisy, instead of continuing the call with half context, right?
it has to say something appropriate.
So that could be an input.
The task specification is, hey, you're talking to a customer
where the environment is quite noisy,
so you don't hear the customer clearly.
Then the verifier is, does the task specific agent actually
nudged them towards moving to a place where the reception is better?
Or asking the customer that, hey, can we schedule a call back later
so that we'll be able to speak when you're free?
Something like that, right?
This is a verifier.
Now, your meta agent is supposed to take call.
these evaluation test case scenarios, run your task specific agent through all these scenarios,
and make sure that the sequence of actions that your task specific agent is performing is
as per what you had anticipated to perform. If not, then go ahead and update your task specific
agent accordingly so that your task specific agent is good enough. So all a team has to do is
investigate, sorry, invest time in figuring out what is the best way by which we can improve
the benchmark of tasks that you test your agent on, and then it's a matter of
helming towards improving your agent on the performance metric that you care about.
Yeah, in what's in the feedback agent, you're relying essentially, since you don't have,
I'm assuming you don't have labeled responses for all these different scenarios, you're using
essentially an NLM sort of judge to validate the feedback in performance of agent.
Exactly. So we have what we have, what, we,
we have done is we have built a proprietary methodology by which even in the absence of let me open that.
What we have done is we have built an appropriate methodology.
It's called proxy evaluation agent.
This paper is not publicly available but I can show you a preview of it right now.
So this paper is accepted at ICML.
ICML is one of the most premier conferences for machine learning.
Right. Where what we have done is we have discovered a way, even in the absence of actual ground truth, can we create proxy evils in order to make the agent do the work that it does.
So what's a proxy event?
A proxy avail is in the case of absence of an actual evaluation data set where your customer hasn't defined, our feedback agent, the one that you saw over here, right?
The one that you see over here can actually go ahead, look at the work that was done by your task specific agent and try to guess what are the evaluation scenarios that I can create based on it.
And then show it to the users like, hey, I see these are the scenarios that you have tested already.
These are the new scenarios that you're not testing right now.
And here is my suggestion for these scenarios that you can test.
So it's a proxy Eval that you suggest to the user.
so the user can quickly add new test case scenarios based on it.
In some ways, this reminds me of a little bit of deep minds, alpha-evolve work,
where they're essentially evolving the code or algorithms by generating a much of candidates
and then keeping the one that performs the best.
Yes, sure, yes.
So, okay, then we'll go through the history of alpha series, right?
Alpha-Go, alpha code, alpha-evolve, all these alpha-series, right?
They all form a, it's a lineage.
It's a series of papers that they put out with a singular objective that, hey, how do we build AGI?
So it's really commendable that deep mind, I think AlphaGo paper came in 2016, but they have been working on AGI even before that, right?
So they had a fundamental hypothesis and I'll help you understand how does all of these different, how do all of these different series of papers change?
let me from alpha go to alpha code or even alpha evolve how does it differ so deep mind wanted to
figure out what is they they view the world as a hill climbing machine okay which is essentially
like hey i will have an i will have an i will have an i will have an ai that is extremely powerful
when given a task it can go ahead and accomplish any performance metric that is associated with
the task irrespective of what is the
nature of the task, I'll have a mission that will take up the task and successfully finish the task.
That is what DeepMind believed in.
So, and how does it do it?
It happens via a process called search, where your AI will go ahead and take up the problem
statement and think of some candidate solutions as what I mentioned earlier, some hypothesis
on how to solve the problem.
and then it will go ahead and write some experiments to actually implement the hypothesis, run those
experiments, learn from the truth and then update it.
So this is a search process that has to happen constantly in order to figure out the truth.
So the subject of the search is to help climb towards a performance metric on how well your AI is
able to solve the problem.
This was their initial thought process.
And in order to do this process, then where do these series?
of work come into picture. First, they took up use cases where it's extremely clear that
you have Alpha Go. Go is a task where two players play and it's ultra clear who won the match.
There is no ambiguity along the lines of, hey, who won the match? It's very clear. Which means the
work that your AI does is clearly verifiable. On the other hand, think about a use case like,
like, hey, ask your LLM to write a poem, ask another LLM to write a poem, which of these two poems are great poems.
Some people might agree with one and other people might think the other poem is better.
So it's not clearly verifiable.
It is not objective enough in nature.
So in the alpha series of paper, what they tried to do was they tried to start with clear verifiable reward signals.
tasks which have clear verifiable reward signal and then build an AI that can autonomously
hill climb towards it and they started loosening this definition of verifiability slowly right so
verifiability is from a spectrum right clearly verifiable to non-verifiable outcome and then in alpha
evolve they what they started doing was hey if they expanded the scope and they said this
I'll do it for algorithmic discovery tasks.
Yes, the output is clear, but then the search is too large for me to go.
In Go, if there are like 300 combinations, in Alpha Evol, they'll have like a million combinations.
And then they also added these features of how do I, they add, introduce a feature of memory,
where your AI does not start from scratch every time it tries to solve a problem.
It stores some of the intermediary responsibility,
uh,
intermediate results for it to go ahead and,
uh, solve a new problem, right?
So they added these sophistications.
But at the end of the day, the, the, the, um,
underlying methodology is, hey, can I use a Monte Carlo graph search or Monte Carlo tree
search in order to help him towards a metric?
That's it.
Mm-hmm.
It's not just about, um, verify it.
Verifiability as well.
it's also speed with which you can verify too.
Because if the speed, if it takes too long, the time horizon, the verify is too long,
then you can't learn fast enough.
Like if you were, let's say, I had an agent or an AI that was, I don't know, trying to create laws,
then the amount, the time horizon before you would find out whether that law was passed or something
would be so big that you wouldn't be able to learn fast enough.
Yes, yes, I completely agree, right?
So it's not just based on how objective is it to verify, but how time-consuming is it to verify as well.
Yes, Sean, that also plays a very important role.
Yeah.
You know, I think with a lot of self-improving work, at least based on, you know, my understanding,
I'm not the expert that you are, but I've kind of seen people or prior research kind of take
one lane where in the world of agents, maybe that's working on sort of trying to rewrite
some portion of the harness, the context that you're going to feed in the model.
alternatively we can try to adjust the weight to the model.
And then it seems like what you're trying to do is combine both of those things into the same loop.
So why, I guess, is it important to do both?
And then I'd love to get into a little bit of detail about how the weight adjustments actually work.
Sure, Sean.
So let's let's.
Okay.
Now, I'll address this as two-part question.
Is it essential to do both or is it sufficient to do just one?
And then how does weight adjustment actually work?
Because there is not something that is widely discussed until now.
Right.
Okay.
Maybe we should let the benchmark speak for itself.
Right.
So these are, here are some academic benchmarks that we have released because these are
usually the benchmarks that people use in order to evaluate any agent that they built.
Right.
Let me go ahead and do this.
Okay.
So this is a.
which, this is a better benchmark, okay.
This is a benchmark on evaluating how good is he a agent
at generating Kuda kernels.
Okay. Generating which, sorry?
Kuda kernels.
Okay, yeah, yeah.
GPU kernels.
Yeah, yeah.
Yeah.
And so one of the most time-consuming things that people do
is whenever there is a new type of neural network,
they go out and write a kernel for it
and make sure that kernel is efficient enough
to run that neural network faster.
and it's a very manual labor-intensive process, right?
So how about we build an agent which can autonomously generate kudakon?
This is the question, right?
So here are some benchmark research that you can see, right?
So this is a baseline.
A baseline is essentially just a GPTOS is 120B model
wrapped around a basic react harness.
Okay, that is your baseline.
Now, when you ask,
codex to go ahead and create a GPU kernel, then it's able to attain some performance because
codex's harness has been like really optimized for certain types of tasks. But when you ask clod
code, then it does way better than what codex does. And but when you use CR and you use
CIR and you tell it that, hey, take up open source models and optimize the harness alone, it's able
to do little better than codex but not way better than clod code. Right? Because it's powered
by GPDOS as 120B, which is much inferior model in comparison to that of your Opus 5 or whatever.
Right there, right?
But when you ask SIA to not just explore the harness improvement, but also to explore the weight improvement, then what happens is, say, is able to get way better performance.
Which is basically what SIA does is SIA takes up.
So then we have to, let's go back to this picture.
Okay.
here your task specific agent is a Koda kernel generation agent.
And the way how you evaluate is it,
how fast is the kernel that is being generated running.
So speed is the metric that you're trying to optimize.
Now, what we see is when your meta agent is just updating the harness of the task specific agent,
then you're able to attain some level of performance.
But when you let it improve both the hardness as well as the weight, right?
it's able to attain way better performance.
So what does it show?
It's because it's when it,
I don't know if you have seen Anthropics terms of service that they changed.
And they said that they are,
they are decelerating the capabilities of cloud code
to go ahead and build specific GPU kernels.
It's because it's not in the best interest of Anthropic
to give you those capabilities.
Why?
because they want you to use their models and not open source models
and because you write kernels for open source models, right?
So what you care about,
so therefore the underlying LLM, which is Opus 5 or Sonnet 4.8,
does not have capabilities to go and create Qaeda kernels.
At these scenarios, it makes super high sense for you to have your own agent
and have access to the underlying model
and train the underlying model and just build.
that particular task well.
So this is an extreme example that I told you where they don't want to give you those
capabilities.
But think of an enterprise.
An enterprise has thousands of tools internally that they have built for themselves and
documentation and everything around it.
Neither Open AI nor Anthropic nor DeepMind have access to these tools or to the data or anything
like that.
But so when you use Clods Hornet or one of these agents, one of these models to power your
agent, then they're able to be able to be able to be.
perform to some extent well.
But when you're able to train your own agent, right, then the performance supersedes
what you can get from any of these frontier proprietary close source models as well is
the main lesson, right?
Just because they don't have access to your data to train on it.
Right, yeah.
And in your, since you're doing like of the low rank fine tuning here.
then you're adjusting the weight.
So let's say I was a company that had 100 different types of agents
to solve different purposes, then am I going to end up
with different model weights that's specific to each of those agents?
So essentially, I'm going to run a purpose built harness
and then I have a purpose built weights
that's built on some open weight model to start with as the baseline
for every task that I'm trying to solve for.
Exactly, Sean.
So the way how we believe the future will evolve into us.
You see in software engineering, right now,
multi-cloud is the way to go forward.
Some services from EWS, some services from GCP,
some services, you will have on-prem deployment going on and all those things.
Yeah, hybrid cloud, multi-glown, yeah.
Exactly.
In the future, in the future, how it is going to happen is you'll have multiple model providers.
You'll work with open-a-a-a-a-a-a-and-sropic deep-mind,
as well as your own open-source, fine-tune models and all those things.
I think the current stats is something like 37% of enterprise have at least five models or more going right now.
So I absolutely think you're going to end up, especially with everything that's going on with what happened with Fable and Methos and some companies, especially if you're not in the United States, being concerned about investing in a model and then having a foreign government taken away, you want to have a diversified portfolio of models to protect yourself.
And then probably for certain types of workloads, you'll want to take advantage of cheaper models
because you don't need to run the most premium frontier model to do whatever that task is.
Exactly, Sean.
So there are three main metrics that customers typically care about.
The first one is performance.
The second one is cost.
The third one is latency.
right there are some applications where you are willing to spend quite a lot of money and attain
superior performance and faster latency so that you are able to deliver something for example take
a voice calling agents you have support agents where these support agents are on call with customers
no matter what happens you you don't care about cost over there the first thing that you're
like, hey, is my agent actually doing the right thing?
And is it faster enough for it to happen?
That's what you optimize for.
But let's take up another use case where you are just cold out reaching to some potential prospects and email research agent.
Like you're like, okay, sure.
See, I have, I'll take a better example.
I recently heard about this, okay?
There's a customer that we work with.
they use her agent optimization tool to do abandoned cart optimization.
You go to Amazon, add something on your cart, you forget, you move on, that's it, right?
Now there are like millions of such, millions of such abandoned carts every day.
So what this customer does is they give a phone call to some of the prospects.
And it's very simple.
It just has to call them and tell them that, hey, I see abandoned card, do you want me to check out?
That's it.
Okay, here, there is not much interaction with the customer at all.
It's basically saying that, hey, should I go ahead and check out or not?
That's all it does.
Now, what they care about here is, they don't require some superior performance that out of the box.
You don't need a GPT5 level model there.
It's basic sentences that it has to speak.
It doesn't have to know how to write great poetry and how to debug a Python code.
But it has to be cheap.
And it has to be quite fast.
So what do you care about, right?
Differs as what kind of agent do you build, right?
So Sonnet might be having extremely good performance,
but it also has very high latency,
which your application cannot afford to have.
At that point in time, you might want a different model
or a different harness or a different memory structure.
Right?
So there are these trade-offs that you constantly make in your mind.
Now what happens is all the A.A. engineers,
they go ahead and manually try to try out different experiments to make these tradeoffs.
What we are promising is we are promising a way by which there is a scientific way
by which you can run experiments to figure out what is the right way to optimize your agent.
Yeah.
And then since you're doing some fine tuning there, like what is the delay?
How long does this whole process take to get to a place where you've reached?
some, you know, I don't know, optimal state of improvement.
I guess at some point, I guess you could always try to improve,
but you're going to probably hit some diminishing returns on this.
Exactly, right, Sean.
So the time, right, it largely depends upon how large of an agent that they want to be trained.
What is the performance metrics that they would like to attain?
Or what are the other Eval metrics?
What is the data set that they have in their hand and all these things, right?
If all these things fall in place and it's very clear they want to train a 7B model for like
three apox on a small set of customer data,
then it's a matter of minutes.
A few minutes, you have already one version out there.
But if it involves running more experiments
in figuring out a larger model, larger data set,
and very stringent metrics,
then it can span on for like a few days as well,
like three or four days as well.
So it largely depends on what the customer wants,
and what are the exact tradeoffs that they care about?
But the one thing is it's way better than what you would do by sitting and manually optimizing it.
Because we can just spin off parallel experiments.
How do you prevent overfitting?
The same principles as all what we had always used in machine learning, right?
Have a clean training set, eval set and the test set and make sure that you do careful validation.
And you don't, yeah, the basic scientific hygiene that you follow always to make sure that overfitting doesn't happen.
and add regularization and all those things.
There's no big rocket signs for optimizing agents,
different from what we learned when we optimized RSVM or a random first.
Yeah, although you talk to enough people
and you find a lot of people skipping that.
They don't like having to put in the hard work to put it together good tests.
So that is a thing, right, Sean.
I don't see, everyone is in a pressure.
to deliver quick results.
Okay?
And what has happened with these agents is that demos have become super easy to make,
that everyone is making a Twitter demo and putting it,
and then there is pressure on people who actually do good quality work
on how to deliver something faster, right?
So it's obvious that because of this time pressure to develop,
people might skip some of their steps and they'll go straight there.
Like, have you heard of this term called vibe eval?
vibe what is it vibe r l vibe vials
oh vibe evals yeah okay yeah right like people do white check
yeah yeah yeah what is it i why yes i have yeah yeah yeah i see a lot of the vibe checking
put you know your finger in the wind and it you know it feels good uh people don't usually
like it when i tell them that they need to build their their evals first but exactly right
but everyone but does anyone disagree that it's the right way to do it no everyone agrees that
you need to have a proper evils, say,
try and it,
but you just don't have time to do all these things.
And that is where SIA comes into picture.
It removes off all these routine, mundane tasks
that you manually sit and do, right?
I mean, when we use Psykit Learn or Pytarch hyperparameter optimist,
we did not think about, hey,
should I use a Bayesian optimizer or a hyper, like some other form of hyperbaran.
It was just available as a tool for you to download and just run it.
you will run batches of experiments to figure out what are the right hyperparameter.
Similarly, our goal with CIA is that, hey, it's just available for you, that you can,
just like your Pytarch grid search or Scycuit learns, Bayesian search, you can quickly do some
hyperparameter search in figuring out what is the right way by which you can hill climb
on your Eval matrix.
If that is there, then people will invest your effort into building good evils.
Yeah.
How plug and play is this if I use the framework?
Like, what do I need to do?
It's simple. You have to define your revals and then allocate a budget saying that, hey, I would like to spend $20-30 on being this optimization process and getting some quick results. That's it. And then you need to have a baseline version of the agent already put in place. So you have an idea about, hey, what is that I anticipate?
Okay.
As the framework is like rewriting instructions, changing model weights, is there a version history
to all this?
Exactly, right?
Yeah.
So everything that happens.
So think of it as a very good experimentation engine.
Every time it makes a change, it does a Git commit, creates a new pull request and then
keeps everything in place.
Yeah.
So it's scientifically very, very sound for you to navigate across different versions of history.
Have you seen anything from the system that, I guess, like, truly surprises you where it's something that you wouldn't have thought of if you were hand crafting this?
Yes.
So one of the things, right, Sean, have you heard of this concept called silent failure detection?
So I'll be in a way.
So most of the times what happens is most of the times, what happens is you,
you go ahead and you run okay once I'll tell you this there was there was this customer who had
an agent to be built okay and they had some evils it curated right there's some eval set curated
and they showed me a and they had tried out different models they tried out different models and then they
they reported results on it that, hey, I've already tried my agent with these different models.
And these are the benchmarks course that I get.
Now, can you help use your agent to further optimize it?
We have been trying a lot.
So I looked at it.
Then they have benchmarked, like, I don't know, like 100 different models or something.
Right already?
And they want to figure out, hey, what are the right harnesses that have to go with these models?
Great, so they shipped this.
Now I put it into SIA and I asked Sia, hey, go ahead and just do harness optimization for each of these different models.
That time, Sia noticed something interesting that, hey, I looked at your scores of all these models.
And I see number for different sizes of models.
For example, Lama 7B versus Lama 70B or one of these things, the score differences are extremely small.
right and I when I look at the traces as well that so there are traces that were collected in order to solve these different tasks right when I look at the traces I noticed that there is an eval file that has been corrupted so the customer's agent has gone ahead and created its own eval file and now the metric dashboard that you see is
is essentially coming from the customer's agent itself.
Because the agent is not able to access the Eval file that the customer produced.
But this is buried deep inside 100,000 traces that the customer generated, right?
It's, is it rocket science? No.
Is it some magic? No.
But you just see, right, like even a process of taking up an agent, running Evales,
reading through hundreds of thousands of traces is something that,
cognitively you and I cannot do.
It's just so tiring, so time-consuming, very painful.
But given that SIA is doing all this work and SIA knows how to perfectly manage memory,
when it reads through other agents work, version control and all these things it knows.
It's able to do this job effectively very, very well.
Then this arose a red flag and then the customer was like, oh shit, my evil file.
is corrupted so my agent is not able to access it.
These are some of the places where they're like, so most of the times what I notice that
we might think that customers want SIA to go and discover a novel way to do RL training.
That's not what they anticipate.
They anticipate basic things like can Sia make sure that there is no human oversight happening?
Can Sia make sure that I, the work that my agent does,
is actually sensible enough or not, right?
It's just that I have to review shitload of things, right?
I have to review a shitload of things.
Then instead can see or do that work for me.
So to go back, right?
I don't think customers want Cia to perform
at the level of Ilya and discover a novel algorithm to train the region.
All they want is to just do the routine mundane shitload of
automatable task to be just done well, Sean.
Yeah, yeah, it makes sense.
And in terms of, you know, a lot of agent performance can be driven by the,
especially for like a purposeful agent in B2B,
a lot of that can be based on proprietary data that the customer has.
So they got to go and retrieve it from somewhere.
So then the retrieval process also impacts performance.
Exactly.
So how does that factor in here?
anything within SIA that helps me, you know, tune essentially the way that I'm bringing in data
from my proprietary data sources.
Exactly, right?
So there are two, okay, when you fetch data from your property data sources, it can be used
in two ways.
One is it can be used to create evils for your agent or the other thing is it can be used
as a knowledge base for your agents, right?
For knowledge base, we are not offering anything.
I think the knowledge base has to be created by the customer.
Anyways, every customer has their own environment, their own knowledge base.
So we don't have any viewpoint on what can be done better there.
Sure.
I guess if the search base mechanism into the knowledge base was dynamic, like a tool call,
then you would have probably some leverage over how you're making that tool call if the parameter,
like the search parameter basically is passed in by the ALAM.
Exactly, right?
So the way, so the problem that you mentioned, right, I can help you understand how do we solve it.
So Sia is Socratic in nature, right?
Which means when you deploy Sia, Sia can, Sia also has a inbuilt, critic built within herself.
Okay?
Which means when you bring some knowledge base and say that he gets, try to get,
get this work done by yourself in improving the performance of the agent.
C.R. also has an internal critic agent which will try to critic whether do I have required
knowledge to do this from the knowledge base, right? And it will give feedback on, hey, this task is
just impossible to be done because I just don't have the right knowledge. I don't have,
I don't know what I don't know. So I can't do it. All right. So that is a way by which we
tell you whether the knowledge base is sufficient or not. Right.
On the other hand, the next iteration, what we are trying to do is,
more than the knowledge itself, evils is the bigger problem for customers right now.
So we are introducing a new endpoint.
It's called slash generate evals,
where SIA goes ahead and autonomously generates evils from their existing traces.
So if you have run an agent and it runs into an issue,
then it creates new evalts from those traces itself.
So that is one problem that we are solving today.
But for the knowledge base,
it's we tell,
we are able to tell the,
C is able to tell the customer or the user that,
hey,
I don't have access to this knowledge base in order to do this work,
yes,
but is it able to go and find the right knowledge base?
Not yet.
But surely we should add it in the pipeline or the roadmap shown.
Okay.
Great.
That's fantastic.
I mean, this was super interesting.
I guess anything else you'd like to share,
and also, if you want to check out,
see what you're doing over at Hexel Labs.
Where can you direct them to find out more?
Sure.
So the first thing is going to our GitHub repo.
So if you just Google for,
let me share my screen.
So our open source repo is here,
so people can go ahead and already give it a shot
and then see how it works.
Yes, this is Hexo, AI,
Yes, yes, yes.
So our open source reposes are so people can already go and try it out.
Or people can always go to hexel.a.
Our website, and you'll see a lot of work that we do other than this as well.
There is another paper which people can get to know, Sean.
Fantastic.
Well, Bignish, thank you so much for joining us.
This was great.
Learned a lot, really fascinating.
I feel like we're only scratching the surface, but it was great.
Thank you so much.
Cheers. Thank you, Sean. It was wonderful talking to you.
