Latent Space: The AI Engineer Podcast - 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery
Episode Date: August 11, 2026This January, four big AI × Pharma tools deals were announced at the huge JPM Pharma conference that takes over San Francisco every year. OpenAI-backed Chai Discovery (now worth $4B) was somehow at t...he heart despite being all of 2 years old. The Science team is proud to bring you the first podcast with cofounder Matt McPartlon and product lead Neil Patil to tell the full story! Editor’s note: not to be confused with Chai AI, which was another top pod of ours.Pharma suddenly doing big AI tools dealsFor the non-pharma people, JPM is JP Morgan’s annual conference for pharma deal-making that takes over San Francisco for a week in January with hundreds of side events, etc. It’s a big thing. Tools deals for pharma are also a big (new) thing: companies that start as AI for Pharma usually end up building their own drug pipelines instead, and the reason is something like this: convincing pharma to use your tool requires proof that your tool works. Proof means good targets, maybe with good clinical validation. If you have that, then it’s easier to raise money (with a known, if long path to commercialization) or sell (e.g payment in biobucks) for a specific target than it is to sell to lots of companies on a promise that it will work across their portfolios. The “we’ll just partner / build our own drug” optionality proved to be the only good path up until January. What changed? In short, the tools got good enough for drug design teams to trust.Good-enough-to-trust unlocks the ability to scale discovery: get more, better candidates into the lab and animal trials faster. More screening for toxicity, better delivery, etc. This means that what you push to the clinic is more likely to succeed.Tools also unlock new capabilities: mechanisms that are very hard or impossible to develop using lab-based discovery. Designing an antibody that precisely triggers a very specific molecular cascade takes many years of trial and error. Designing bi-specific antibodies (that bind to two different proteins) is similarly difficult. Good design tools can unlock this.RJ: The fact that the quality of the model has jumped means you’re enabling things you just plain couldn’t do. So it’s a step change. It’s not an efficiency argument at all, or not so much.Matt: Yeah, exactly. It’s kind of interesting, even for us — it took me a while to believe in the thesis, actually. I talked to Josh for months before Chai started... It’s like, can I beat a mouse, and then can I do what mice can’t do? And then how many levels of interaction can you just keep building on top of that?Everyone playing in the structural / binding space has an angle here, and some will be better than others, but Chai is pointing to a different unlock: getting good molecules right out of the gate (meaning they don’t then need as much lab work) means that the iteration time is faster. This turns science into engineering: you can design your systems to reduce friction and hill climb towards one-shotting molecules all the way to the clinic.This, per-se, is not a new thesis: a16z articulated a version of this in 2020. What has changed is that structural models became binding models (how well doesn’t this molecule bind to this molecule, aka “binding affinity). Binding models unlock design, which has been steadily improving. Chai’s observation is that for engineering problems the best product tends to win, and good technology is a necessary but not sufficient condition. Photoshop for moleculesWith that in mind Chai has invested heavily in partnerships that allow them to learn from their Pharma counterparts. What is kind of cool about working so closely and supporting so many of these partners is we get to really learn about what is the stuff that would be helpful in research. So rather than doing research in a vacuum, based on what would hypothetically be cool, we're able to do informed research based on what our partners have just been organically asking us for help with.— Neil Patil, (Chai product lead)This means better UX, such as a molecule editor that is more like a CAD or graphics design program than a chatbot.Their approach has paid off: since June, Chai has announced three more major deals: Lilly, Novartis, argenx, plus an expansion of their Eli Lily program. This episode is too full of quotable moments for a short blog, so tune in to learn about * Why protein tokens have the highest downstream value of any token * Climbing levels of abstraction as models improve * How Pharma, VC, and research are all just portfolio optimization * How better tech changes the whole portfolio * How relentless focus on simplicity leads to scalePlus much more! This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Transcript
Discussion (0)
It looks a lot less like a, you know, a chat chip-t, and a lot more like Autodesk or SolidWorks or Figma, you know, if you've used those things, where you can kind of load up your molecule.
There's this almost like Photoshop-esque, like, design suite. You have this equivalent of a paint tool to kind of paint your epitope.
You have this equivalent of a content-aware fill tool to kind of get your binders generated from Chai.
And I think to add to that, right, yeah, the notion of target discovery and hit discovery and optimization, where each of these has a gate and takes a few months to a few years.
is this very like waterfall model, right, where the cost of trying things and getting things early is very expensive.
But I think to what Matt's saying, right, if you start to get in a regime where you can have models give you really promising candidates, you can start to make that look a lot more like a loop, right?
It's akin to like becoming more agile and software development.
But now the next problem is like agonists, right?
Like how do you reliably one shot hitting a switch like on a cell, right?
Or by specifics or ADCs, right?
And I think this levels of abstraction that we're going to have to climb with the product as, like, the models get better.
If you have, like, these really good primitives for structure prediction and binding and design, and you can kind of compose them, then you can start to just, like, grow into, like, the outer loop of science.
Welcome to Lidenspace AI for Science.
I'm Brandon.
I build R.A. Therapeutics at Atomic AI.
I'm joined by my co-host, R.J. Honicky, CTO and co-founder at Mirroromics.
It's a pleasure to have with us in the studio today.
Mac McPartland and Neil Patil of Chai Discovery.
Chai is a protein design startup, which is about two and a half years old and has made quite a splash in those few years.
They have several very exciting announcements that I think they'll tell us about today.
But yeah, to get started, could you two give us a bit about your background and what you do at Chai?
Yeah, thank you very much for having us.
We're super excited to talk about Chai today.
I'm Matt McPartland.
I'm one of the co-founders at Chai.
My background is in like AI biology-related stuff during my PhD.
I actually started my PhD in like theoretical computer science and then transition to this later.
Yeah, I've been doing this stuff now for like about eight years and I kind of came into the field at an interesting time where protein structure prediction was like just starting to see signs of life.
So this is like Alpha Fold 1 days and was in the field during AlphaFold 2 and like got to see a lot of the interesting developments at that time.
So, yeah, I'd always been pretty interesting, like applying this stuff in the real world,
and Chai was just a perfect opportunity to do that.
And I'm Neil Patil.
I help lead a platform and product here at Chai.
So a lot of the stuff around infrastructure to train models, serve them,
and then the productization piece, you know, the design suite that lets you use the models.
I kind of have a more meandering path, so I kind of got into programming, like 15 years ago,
making apps in the app store, got really addicted to the dopamine hits you get from that.
And then actually got nerds snipped by robotics.
and worked on that for a bit, self-driving cars in like 2018, 2019, got really jaded and was like,
I don't want to touch hardware for a while.
I ended up switching and joining a SaaS company called Vanta is one of the first employees there
and kind of grew with it, started my own security company afterwards, got a few years into that,
and I was like, you know what, Adams are kind of cool.
Like, I want to work on something a little more meaningful.
And so I joined Chai about a year ago right after Chai 2 was announced to help with a lot of the
platform and commercialization pieces.
Awesome.
It's like the five stages of grief or something.
Yeah, we're at acceptance.
Awesome.
You have these, I think, four now big partnerships and raised a whole bunch of money.
Can you tell us a little bit about those partnerships?
And then what I really want to know is what are you telling investors and customers that is so compelling that they're willing to do these big deals?
Yeah.
So we've been very fortunate to partner first with Eli Lilly and then with Pfizer.
Navaros and our GenX. Yeah, I think it's been like a really interesting ride. And I think our
business model is also very compelling to a lot of people. Like we really like to, we care about
the partners succeeding. Like this chai as a company really depends on how the partners succeed.
I think Neil probably has some interesting takes on like, you know, what we actually offer and
what makes that so compelling. So hand it over to you. Yeah. I mean, as you all know,
drug discovery is a very lengthy process, right? And a lot of these pharma companies are spending
lots of time, you know, years and years and billions of dollars trying to find initial therapeutic
candidates. And so at Chai, you know, we trained models that can help accelerate that process
and kind of find those initial binders and then some. And, you know, we, you know, there's a lot of
bio companies, AI for biocompsies that are like making their own drugs. We really don't see ourselves
that way, right? We see ourselves as almost a neutral software factory for making medicines.
And so that's what, you know, lets us go then work with and support all of these other farmers
in their kind of a drug discovery journey.
And so, yeah, I mean, a lot of this capital is just another proof points
that we can sort of start to really accelerate that software factory, right?
Go after harder modalities, train bigger models,
and ultimately just build what our partners and customers ask us for.
But what is it that why you and not other structural companies,
why are they compelled to buy from you?
The thesis of Chai has always been to be the software and modeling layer,
which was, I think, like, very controversial at the time.
Like, everyone, you know, this play has definitely been tried.
Only two years ago, and it's already, like, a completely different world.
Yeah, it's pretty crazy.
Like, people tried this play for a while.
And I think, like, the models just really weren't there yet.
And even, like, for us, we were taking a risk in the very beginning.
Like, we were kind of banking on the models getting there.
And, like, I had seen early signs of life in my work and our CEO, Josh.
Like, he was on the original ESM papers on that team in Meadow.
And he was seeing, like, pretty early signs of life.
like, you know, there might be scaling laws here.
They, like, I think we'll actually be able to start, like, designing things.
Structure prediction is going to really good.
Like, one, like, crazy thought is, like, we didn't have a Multimer structure prediction
model until, like, 2021.
That was five years ago when we could, like, start with deep learning to, like, actually
predict the shape of two proteins at once.
Like, it was an outfold one was like, an outfold two was like this huge breakthrough,
but then, like, outfold two, Multimer came out, like, a year later.
So, like, you really kind of needed that to unlock design in the first.
place anyway. We weren't even trying to predict multiple proteins at once. And then really like
around that time, inverse folding kind of started working and there's like, oh, protein mp and n. This
actually works in the lab. Like credit to the Baker lab for doing all of this really excellent lab
validation and all their models. But I think like we're starting to see them do interesting things and
like actually work on like real world experiments. And now it's probably the time to start betting on this.
I think like before then maybe you could take like some experimental data from a campaign on like this one
target that you had and you care about. And you might be able to like make some progress on that
and like keep hill climbing in this like one very specific case. General models weren't really a
thing back then. So I think like yeah, we took that bet pretty seriously and like we decided to just
like push as hard as possible and to really like shoot for generality in our approach. And then
when Chai 2 came out, our second paper after Chai Wan, we kind of like show the world like,
this is actually possible and it's possible at scale. We didn't show this for like one or two targets.
kind of works. Like we were like, let's just go all in. I think Josh like to say we set up bold
company-wide challenge to design antibodies to 50 targets. And actually like we saw some signs
of life. We're like, all right, let's like, let's do this with real statistics and see if this
actually works. It's an interesting story of how we chose these targets. So we're like, all right,
what targets we're going to choose? We should choose like some interesting targets, whatever. And at
that point, we're like kind of ramping up with CROs and figuring out like what does our wetland
process look like. And we decided.
after trying some stuff with many proteins, whatever.
We're like, here are the interesting targets.
This is what we should look at.
And, like, half the time the targets just, like,
kind of didn't work.
We were still learning, whatever.
And we're like, all right,
maybe we should just go with, like,
targets that the CROs have actually validated.
So let's get the CRO catalog,
see what they've already worked on,
restrict that to, like, an interesting set.
So from that, we chose 50 targets,
designed antibodies against them,
got hits to half.
And at that point, I think Pharma started to realize,
like, okay, there are actually signs of life here.
and this might actually work in some of our programs.
And so antibodies is maybe a more challenging domain
than other structural prediction problems.
So why tackle antibodies?
So maybe back up, what is an antibody?
Yeah.
And what do you do with it?
And why is it an attractive target?
The analogy that everyone gives like this lock and key kind of problem
where like your target, this protein that you're trying to bind to,
it might be some like disease protein, that's kind of like your,
lock and then you want to design this key that fits into it and like in our case just like sticks there.
The interesting thing with the antibodies is like these like really flexible general proteins.
Like in a lot of ways they're very general and a lot of ways they're actually like pretty uniform.
But at least like how they bind to a target is very general.
So like you have a lot of optionality and how you design this kind of binding interface.
The structure prediction problem for antibodies like predict how this antibody actually binds to the target,
how it how the key fits into the lock.
that's been a notoriously difficult problem.
The nice thing is, like, so we've made a lot of progress in a structure prediction.
Kind of the field as a whole has come a long way along, like, in getting structure prediction
to where it is.
But in the design setting, you can be a lot more selective about the types of designs you want
to make and the types of structures you actually want to focus on.
And in some cases, it might actually be even easier to design a protein binder that is
an antibody than to actually predict how it might bind that target in general. So it's kind of like,
if you have the freedom to choose, you can kind of just pick the easy cases, if that makes sense.
So the antibody is like, there's a whole machinery in the body that works with anybody is. What
does the body do with it naturally and what can you do with them that is sort of not natural,
but is useful for therapeutics? This is coming from a non-biologist here, but I think anybody's like
They're these kind of like Y-shaped proteins, like kind of looks like a P-side with your fingers.
Each of these fingers is kind of like an arm of the antibody.
And like it's really actually only the tips of the, of your fingers, the tips of the antibody that engage in binding.
So this makes these really like nice therapeutic design targets for that particular region.
The nice part is that like the rest, apart from the tips, is like actually relatively constant.
So this is called like the framework region of an antibody.
And the design problem, you're typically just as a.
designing like the very fingertips.
And you can actually choose, for the most part, like these kind of framework regions
that your immune system already recognizes.
So antibodies kind of like these Y-shaped proteins that your immune system like recognizes
and knows really well.
It's kind of like your bodies, it's one of the lines in defense against pathogens and other
types of diseases.
So I guess antibodies can, on the one end, like, connects to proteins on the surface of a cell
typically or other things, but typically on the surface of a cell. And then the other end helps the
immune system identify a pathogen typically. But you can also do things like you mentioned ADCs,
anti-anybody drug conjugate. So that's, that means putting a drug on the other side or something like
that. And that causes the, when you bind to something that it releases the drug into the cell.
Right. They're like this very general framework, right? We're kind of on the ends you have these
CDR loops and you can design them to kind of bind to arbitrary.
things where maybe one end you bind to a cancer cell, the other end you bind to a toxic molecule.
You're now precision delivering that toxic molecule to a cancer cell, right? Or you just have two ends
bind to things and kind of force, like, induced proximity to have some effect in the body. Or,
you know, a lot of drugs historically are really just like about like blocking things, right? Like
anti-aginous behavior, right? But maybe you can have agonist behavior. We actually like really
precisely like press a switch. Like there's a GPCR protein, which is a, which is a agonist behavior.
these doorbell proteins that sit in your cell membrane. You have an antibody very precisely engineered
to poke it in a certain way that causes a downstream chain reaction. And I think, like,
one of the things that's really exciting about where we're getting to with some of these models
is we can start to get that precise, right? We can really target a very specific epitope, right,
meaning like binding spot, right? A very specific set of atoms to have the antibody go after,
which, you know, historically, you're with a lot of drugs, you're just kind of brute forcing, you know,
a lot of antibodies and just trying to come up with a bunch of things and see what sticks.
But maybe that gets you a binder to some spot of your target molecule.
But that doesn't let you precisely engineer where you're poking after.
I know you're not biologists, but do you have any like idea about how they used to design
these before, you know, these models came up?
Like what would you, what would you, what was the grueling process you would do to find?
Or what is actually still?
Yeah, what still is the state of the art in terms of drugs which have made it to the clinic?
Yeah, Josh, our CEO likes to say.
say that our biggest competitor is the mouse. So like our nature in certain ways. So like
traditionally these these types of like drug like molecules were either discovered in like these
immunization campaigns. So like you literally will just like infect a mouse of the disease and see
what antibodies it makes to try to like combat that. Other ways of doing this is like super large
yeast display so on. So you might like start with hey I really like this framework and how am I
going to like figure out the right loops to design to bind this target? I'm just going to try as
much as I possibly can and just like literally search for a needle in a hay sack.
And this would be like on the order of like at least billions of potential molecules that
you're screening against this one target.
And in that case, you might like end up with, you know, one to maybe like a dozen potential
hits to this target.
You actually, you don't know much about this hits.
All you know is that they kind of like stick to the target.
You don't know necessarily where, like if they're even necessarily drug-like.
I think like one big separator of chai and like a thing that definitely our partners like to
see is like you can be really intentional with how you want to do this design process.
You can say, I want to bind this target in this particular area.
You can even go back and look to the designs after.
Like we validated that our designs.
So you can go back and look and say like, is this antibody engaging the target in the way
that I expect?
Do I think this will actually have the therapeutic effect that I'm going after?
One of the cool things about knowing that you have the right binding pose is that you can
now also design selectivity into that.
Does your platform have some technique for doing some?
selectivity. Yeah, there's a, there's a nice mix of ideas that went both into the modeling side and
especially on the product side for dealing with selectivity and cross-reactivity. So in some cases,
you want your molecule to bind one target and avoid another one. So you might have like
healthy variants of protein and like disease variant of protein. You want to avoid this,
disease variant or you might have some other similar protein that's like not actually harmful in
your body that you don't want to just like artificially block. So I think like on the modeling
side, yeah, we've come up with ways of doing that, but I think it's even more interesting on the
product side. So, like, how do you enable customers go through or partners to go through and, like,
actually intentionally designed for these things? Yeah, and maybe to, like, back up and define
cross-reactivity, right? Like, it turns out when you're developing a drug, you're not necessarily
going straight to injecting that into a human, right? Like, you might want to put it in monkeys
first, for example. And the monkey might have a maybe mostly similar, but slightly different
variant of it. And so your drug, you know, not only needs to bind to the human variants,
but also the monkey variant, right? And so, you know, the way we've tried to model the models
and the product is to kind of let you account for those very general cases where you say,
hey, I'm trying to design something that can bind to both of these things so that I can
actually go and develop the drug. Let me actually identify maybe the region that's conserved,
and then target, conserved means, you know, doesn't change much between the two and target that exact
region. And then, you know, similarly with selectivity, right, maybe you might want to, there's a very
similar protein in the human that if you accidentally bind that one, that's very bad. And you
only want to bind the target protein. And, you know, that's why a lot of drugs, right, you know,
fail or toxic or have, you know, really bad side effects, right? And so it's kind of, you're kind
of having this like, combinatorial problem of, like, you know, bind only these things and avoid
only these. And I think what's been really exciting with some of the progress recently has been
been like a lot of the improvements we'd have been able to make on the level of specificity
we can get to with those models. So you're not only designing the bind here, but you're also
making sure that it doesn't bind to another thing. Exactly. So other ways that like Hartees have
tried to tackle this by having some molecular or some sort of signaling pathway that says if I
bind, I only fire if I bind this one binds and this one doesn't bind, but you're saying
you just design an antibody that actually only will bind to the thing that you care about.
We're getting to the point where in some case, I mean, it's nuanced, right?
But in some cases, you can actually try that.
Okay, that's amazing.
Yeah.
So you're saying you essentially call it counterscreen or you have, in part of your platform,
you can know reliably counterscreen against like a large, diverse set of proteins,
which might be issues for downstream.
I would say the framing is more you can be very specific about what you care about binding
versus what you care about avoiding.
But I think, you know, for example, like a lot of the money that we're raising now will let us train bigger models that can maybe be even more general and start to account for even more things at the same time, right?
Maybe we should back up.
Let's talk about, so the history of the Chai, you know, series of models.
Well, why don't you tell the story?
We started Chai around two and a half years ago.
At first couple months, we're like, all right, we're going to work on protein design.
And we're working on this.
We're making some progress.
were like, oh, it's pretty interesting.
Like, we had some ideas of models.
And then kind of like, that was right when Alphold 3 came out.
And we were, we'd like been talking about like, man, we really need like an MSA pipeline.
We need like all of this infrastructure.
MSA is multiple sequence alignment pipeline.
Why is it just, we've covered this before, but what is an MSA like in two sentences and why is it important?
So if you want to predict the structure of a protein, it might be really useful to see a bunch of very similar protein sequences.
And what those protein sequences that are really similar to tell you is like kind of what
positions like which amino acids end up being conserved across many variants of this protein.
And if you see like high levels of conservation or like kind of high levels of mutation,
like correlated mutations, it typically gives you some indication that these amino acids are
close in 3D space. So you kind of have this like 2D view of a protein, which can then be used
to help you predict this 3D structure.
So you're learning from evolution, what was conserved because of the things that weren't
conserved probably broke the protein and something died or didn't make it.
Exactly right. Yeah. Yeah, it's pretty remarkable that this works, honestly.
Yeah. One of my favorite, like, biofacts here. Yeah, so we were like kind of thinking like, oh, man, it'd be nice to have like a lot of infra and whatever. So Outflow 3 came out. We're like, hey, we should, we should like open sources model.
We should just like, you know, bunker down, build all the info that we need. I think like this will pay back like in the long term for sure of just like as a forcing function to like be where we are. And also just like to contribute to the community as a whole.
So it's interesting that you chose, okay, this were actually what we're doing here, we're building a model, but what we're really doing is learning how to build the infrastructure. Is that kind of what you're saying?
Yeah, that's exactly right. And like I had built a lot of similar infrastructure in my PhD, but not at a production level for a company. So like at that point, I think we were five people. So there are five of us to try and we're like, all right, this is our forcing function. We have like a clear goal to work towards. It's like very direct. Let's get this thing going and see how fast we can do it.
You guys were at this time sitting in the open AI offices?
We were sitting in the open AI offices, yeah, in the mission.
Right.
So, like, what's the backstory on that?
It's really interesting.
Two of our other co-founders, Josh and Jack, had a relationship with some of the open AI people.
Actually, Open AI co-led our seed round.
So we were, like, kind of thinking, like, all right, should we get an office?
Well, we're only five people.
And it turned out, like, that office was mostly vacant.
So we got to sit in on the, like, in the open-A offices for a while.
Well, Chai 1 built it, open source, learned about infrastructure.
Yeah.
So then after that, like, we really set the size down on protein design.
And worth pointing out, Chai 1 was a structure prediction model, right?
So you have the sequence, what is the structure that it folds to?
And then that was his every shy to.
Yeah, Chai 1. Chai 1's finished.
One other crazy story there.
Let's see if we can actually share this.
But this is a hilarious one.
So, like, we were like, oh, man, we really want to be the first to put this out.
and we're like, okay, we're one week out.
We're like, the model is like almost done training.
We're like, should we build a web server?
And then we're like, oh, yeah, maybe not.
And then, like, we end up spinning up, like, this whole web server.
So, like, people can use it.
Like, rather than just, like, download the Git repo,
it's kind of annoying, especially for biologists.
And, like, we actually wanted people to use this.
So, like, let's spin up a web server.
Let's get the technical report out, all this stuff.
So we ended up, like, we were up for, like, 40 hours straight.
It's, like, getting the paper over the line,
getting all the last things done on the web server.
And then Josh was interviewing with, like, Bloomberg TV or something that morning.
And we've been out for, like, 48 hours straight.
So Josh, like, runs into a room to do this interview on Bloomberg TV.
And, like, I think it was, like, seven in the morning.
Everyone's in the office.
Like, we didn't, like, want to be seen whatever.
And, like, the interviewer's like, oh, like, interesting company.
Doesn't look like there are any employees here.
But, yeah, it was a really fun time.
I think, like, the early startup days were just.
just super fun. So yeah, after that, we kind of set our sites on design. And really what we were
thinking is like we kind of always had antibodies in mind. We thought of this is like the most
tractable problem. The nice thing with proteins is you have this beautiful sequence representation.
There's already a lot of research been done and like how do you auto-regressively generate
sequences? How do you like this sequence generation problem is well studied? So we were thinking like
what's a nice like area to apply sequence generation to in the biospace and it's pretty
natural to do so like linear sequences of amino acids. So we start working on design. A unique thing
about chai is like we're not like we're designing antibodies and like we're an antibody kind of like
we don't we don't really like pigeonhole ourselves into like one therapeutic area. So we like try to
really tackle this problem very generally. So we were thinking like can we design many proteins?
Can we design antibodies? Can we scaffold regular complexes? So like really just take a holistic view on
like how do you design proteins in general? And that eventually led to.
to the Chai 2 model.
So that was our first flagship design model.
And that's where the Chaito paper and like our bold target
discovery project came in.
So we designed antibodies to 50 targets for that paper.
Got binders to about half of them with,
I think on average, around a 20% hit rate for binding.
And then afterwards, started working on Chai 3.
So that's our latest series of model.
But I'll break there.
Before we talk about Chai 3, can you tell us about,
especially for listeners, it may not
be familiar with structure prediction models what does the model look like how does it work in general
let's take a look at chai one uh chai one has this like roughly a tokenizer a transformer something that looks
like language model and then something that kind of looks like an image diffusion model and they're all just like
stick together the tokenizer is like not your kind of typical like words of exile tokenizer uh this is
like i have a bunch of atoms in a molecule and now i want to like pull those into what i would call tokens for my like llm
looking trunk. And then that conditions this like kind of big diffusion model, which will then
emit the image, which is some 3D structure. So is it atoms or is it amino acids that are the
input? It's an interesting question as well. So we have like all these different input tracks. So like
one thing about biology is the data is inherently multi-modality in a sense. You have these,
this like, you know, kind of token sequence representation. Each of these tokens has like a set
of atoms that kind of dangles off. And then you also have, you know, some, some properties
of the different atoms. Like an atom might have, like, a different charge. It might have a different
element type, so like periodic table of atoms. And then these kind of all get bunched together
into tokens. Once tokenized, you can kind of process this in very standard ways. But then ultimately,
you have to get back to these, like, 3D coordinates. So, like, in order to predict the structure,
this is just some 3D object. And that object goes through, or like, to emit that object,
through what looks like an image diffusion model where you kind of go back from tokens back to
the atom representation. I see. So the tokens go in. The transformer establishes the relationship
between the different tokens, and then the diffusion model turns that latent representation
into a 3D structure. This is exactly right. Yeah. Okay, great. So that's Chi 2.
That was try 1.
Okay, so try one folding model.
Yeah, it's like, and like all this bio stuff, it sounds like kind of scary.
Like, atoms, tokens, amino acids.
Like, at the end of the day, my background personally is, like, theoretical computer science.
That's what I spent, like, all of my earlier years doing, transition to this, like, pretty late in my PhD.
But I think, like, the background that you need is really similar to the background that you need for, like, any other field of machine learning.
There are all these domain-specific things that you learn about.
But like one analogy or like anecdote I like to like to say is people think you can't work on like AI bio unless you're biologist.
But it's kind of like you can't work on like video models unless you're like a director or something.
Like there are all these like super domain specific things like oh yeah to understand like lighting and a video, things like that.
But at the end of the day, these are just machine learning problems.
And like they're all solved the same way.
Okay.
So then Chi II, there's a jumping capability as well as an architectural change, right?
Yeah.
What we've disclosed about Chai 2 is like it is an all-atom diffusion model.
So we're trying to predict like, you know, atoms in 3D space still.
But we're doing it in such a way that like the model actually has the ability to like design atoms, place them, decide which atoms actually are there.
So like one way to represent an amino acid like a protein token is by like which atoms are present.
So in the Chai 2 case, we were just predicting like, all right, show the model, let the model just kind of pick what atoms it wants to keep.
and then map that back to what amino acids there are.
What are you able to do with chai two that you can't do with chai one?
It is just like better or are there new capabilities that brings?
It's design, right?
So chai one lets you say, hey, I know the sequence of amino acids, right, that text ring,
and I know the structure that you would get it from like the genome or.
Right, exactly.
Chai two says, okay, I have a target structure, right, that I want to design a binder to.
how chyti two will then generate, you know, candidate molecules, candidate medicines that bind to that target.
And so this is kind of a design model or a design family of models.
And I think that's where you really cross the threshold of usefulness, right?
Like, I mean, Chai 1 alpha will very useful because you can, you know, you can at least intuit and reason about the structure and see what you're looking at.
But, you know, the ultimate goal here is to design medicines, right, and design new molecules.
And I think Chai 2 really crossed the threshold of performance for doing that with antibodies a year ago.
One analogy here would be like, kind of like back to like the image domain.
So, like, Chai Wan would be like, you know, there is a cat in this image.
Like, thanks Chai Wan.
And Chai Two is like, I'll show you a background maybe.
Like I'll prompt you with some like image information like, hey, put a cat in a field.
And Chaito will actually just like give you back an image of a cat in a field.
And you know, like that, that's a good looking image or it's not.
You might have some other model which kind of ranks the image.
But fundamentally it's the generative problem.
So taking an analogy a step further, it's maybe more like you show it a background,
and then it generates, there is a cat, and then it generates an image of the cat at the same time,
and it makes sense that there is a cat in this field and also that the cat works in the image.
So there's a, it's an interesting problem because you have to generate two things at the same time,
both the sequence and the structure.
Then you, I don't know if you can, but could you talk a bit about like how that works?
Like, how do you do that?
So you code, you co-designed the sequence in a way that.
that the structure also fits and makes sense.
One way to think about it is kind of like the classic way of doing this.
Let's talk about both.
And structure prediction, like, all right, I know the sequence of like,
I can from that roughly figure out the 3D shape.
And then there's kind of like the inverse folding problem,
which is like, given a 3D shape, give me back a sequence that would fold into this.
And now you kind of like need to do both things at the same time.
But I think like similar principles apply.
Like you can kind of have the model like think a little bit about what should this
structure look like.
And you can have some other part of the model thinking about
like now what sequence would maybe support this?
And then like a nice thing with diffusion is like you can do this pretty slowly and pretty iteratively.
So you can give the model a lot of time to think about, all right, if I change the structure like this,
how should the sequence change?
And you can kind of just play this back and forth and back and forth.
And eventually it ends up kind of converging on something that's self-consistent.
It's almost like an EM algorithm.
Yeah, exactly.
So you have this model now Chaito, which is able to predict or to sample a structure
and a sequence which generates that structure.
And just because you can generate a structure
doesn't necessarily mean it's necessarily accurate enough to do something.
So do you have other scaffolding on top of that?
Are there additional problems?
Like, are you one-shodding these things?
Or are you needing to generate thousands of them
and then you have a ranking or scoring?
Or just having a candidate is maybe, let's say, not enough.
So what do you do once you sample a structure?
traditionally what's what's done and like when when co-design and like uh protein structure design like
started to become a thing we're like kind of at a loss for metrics is like how do you know that your protein
like you design some some like sequence and structure like how do i know that this is legit or not like
i can tell you it's like anything it's like totally out of domain now right exactly my definition
yeah yeah and like as a human you can look at this thing and be like i don't know it checks out
like even biologists are like i have no idea if this thing actually folds like maybe some of it
looks right. Even our biologists are surprised, by the way, with some of our designs that, like,
do end up working. What was done at the time is, like, we kind of came up with a bunch of metrics,
and, like, alpha fold, it really is what enabled this. So you'd take the sequence that you
predicted. You'd run that through some, like, totally, like, distinct structure prediction
method. So this is completely independent of your model. And you say, if an independent model thinks of
this sequence folds to a similar structure, then it has a higher likelihood of being correct than, like,
you know, just whatever the prior likelihood would be.
So you can take your sequence now and you can measure like how consistent is this structure
prediction method with the structure that you actually predicted for that sequence.
You can now compare your design to an independent model structure prediction.
And that became like a really good way of gaining conviction that your design model was correct.
And people kind of like game these benchmarks for a while and kept pushing, pushing, pushing.
It turns out like it's easy to get self-consistency.
Consistent design the structures if all of your proteins look.
identical. There are a lot of problems that this creates, but then people started adding more and more
on top of this. Yeah, that is an interesting point that I think some people have acknowledged in the
community. So how did you solve that? Yeah, you can see that if you sort of use your Oracle and
also your sampler at the same time, you eventually will converge. What do you do to stop that or
to convince yourselves that you're doing something valuable? One of the nice things about
structure prediction methods is that usually you have some calibration how kind of how confident the model is
in its prediction. It turns out these models, they can give you a pretty well calibrated confidence
prediction. So rather than just say, this is what I think the structure looks like. I'll say this is what
I think the structure looks like and kind of like, here are the parts that I'm not really certain
about. And you can kind of aggregate this down to like a single scalar. And typically what people do
is they'll look at like, okay, like not only how self-consistent am I, how much does this independent
the model even like the structure that it output. So that was one way of early on, I'd say to
like just gain confidence. And then like another thing that people often do is they'll look at like
the diversity of their generations. Because again, you could have a model that's perfectly
consistent, gives you great confidence predictions back. It might be the same structure every time,
like same sequence every time. So you also want to see like, okay, how diverse are the solutions?
How many of these new problems can I solve in a sense? If I had a lot of, a whole lot of money to
validate, how would you do that? Can I go and, you know, do cryoem or something like that and try to
figure out the structure, you know, sort of get some ground truth on that? It's more that the feedback
loop is really slow. So you can validate a few structures like this, but it might take months. And it's,
it's just not like a very scalable direction. So I think that's like a problem for the field as a whole.
And I think people are spending a lot of time, even like especially at Chai, I think thinking about
how do we validate these these problems at like bigger scale? How do we, you know, basically increase
the throughput of our validation or increase the cycle time. Because if you're waiting months to
figure out, hey, it was my model correct? Like, it's just, it's hard to iterate in a research environment
that way. The good news is that this is getting a lot better, right? Like, there's a whole network
now of wet labs that you can work with that will run, you know, these assays, these experiments
and tell you things about, say, you know, does your protein that you came up with bind to its
target well. And so, you know, thankfully we're not at years, right? We're down to like weeks,
which, you know, not as fast as like LLM land where you can just, you know, scale up to.
an eval with and throw more compute and get results back in hours. But, you know, fast enough
to where you can start to recursively self-improve. And, you know, I think we also spend a lot of
time, like, you know, figuring out what are the metrics that we can compute, you know,
in silico, like on the computer that are predictive, perhaps, of lab success. But, you know,
your question about cryo-e-m, yeah. I mean, also you kind of have to measure the structure.
And as you know, that's, like, so expensive because you have to kind of freeze the protein and
shoot these electron beams at it and see how they bounce off. I remember there's like this really
funny anecdote. We'll see if I can share it. But like the, you know, the paper in Chaitou,
we actually, you know, did that. We took some of the, you know, the proteins that the model
predicted and ran cry OEM. And we got the results back. And we're like, wait, the results look
wrong because we had overlaid the kind of prediction over the point, the electron cloud, the point
cloud. We didn't see any difference. And point being, like, we're getting the point now where
these structure prediction models are within, you know, a few angstroms or less of the actual
atomic positions that you validate.
In this case, it was a 0.33 angstrom error, which is one-third the width of an atom.
And we're like, this can't even be right.
Like, clearly they just sent us back the wrong design.
They just sent us back our design.
Yeah, exactly.
Did you check for data leakage?
Yeah, in this case, like, there were no, so, like, we actually chose these targets,
specifically, like, to have no known antibody binder.
Yeah.
So, like, if we did get a hit, like, it was definitely the first antibody hit to this target.
Yeah.
I think that's one of the things I didn't realize about biology was like just how much of it is literally feeling around in the dark.
And that's not even a metaphor.
You literally can't see like how these things luck, right?
So structure models are so huge because now you can, okay, you can actually predict within an atom, you know, how these things luck.
And that enables you to then do things like Chaito with the design models.
This to me is AI for science is one of the cornerstone problems, right?
is that you don't know, you fundamentally don't even know how to measure your problem in a lot of cases.
So it's very difficult to validate.
Yeah.
So you're getting these sub-A-Angstrom predictions with Chai 2.
Chai 3, why Chai 3, what's better or what?
Yeah, I think like with Chai 3, so like honestly, like there was a Chai 2, there's a chai 2.5, there was a chai 2.7, there's eventually a chai 3.
And like each time we saw better and better performance.
And I think, like, the main thing with Tri-3 is, like, we look at Try-2 and, like, we look at the targets it could solve.
There was, like, a lot of internal discussion after Tri-Ti, like, hey, we made, like, successful molecules, binders to half of these 50 targets.
What about the other 25?
You know, what can we do to make those better?
And then, like, you know, we were split.
We're like, all right, should we, like, study these targets that we miss and, like, figure out exactly, like, are there properties of these that we can look at?
Or should we just bet on the models?
Like, will the models just get there if we put more time into, like, you know, just,
be bitter, less impaled in that sense, and just really bet on the models getting better.
And we definitely took the latter approach.
Like, we bet on the models getting better, and we just pushed as hard as we could on that front.
So you're scaling up the model, the data, whatever, to just build more accurate models.
Yeah.
Is it accuracy?
Is that the main thing?
Is it binding affinity?
What do we?
So I think binding affinity is a big one.
Like, you can't just bind weekly.
In order for this to be, like, a useful tool, especially for a partner.
We need to start producing molecules that are like at or very close to therapeutic grade,
which means like they have to bind really tight.
They also have to be developable.
They have to have like all of these nice therapeutic properties.
And developability, I think that we talked about, he mentioned chai 2.5, right, which we
released like a few months after chai two.
There was a study we did on the developability of the molecule, which, you know, for the
audience, like obviously the molecule has to stick good and stick tightly, but, you know,
there are these other properties you care about and to use the non-biological terms, right?
Is it safe? Is it stable?
Is it easy to manufacture?
Does it, you know, self-aggregate?
And we've been pleasantly surprised at, you know, how much we've been able to climb and push the performance in those areas.
It seems like one of the reasons that you want to do antibodies is because the developability.
Yeah, you get a lot for free there, right, with that antibody framework.
Yeah.
It's interesting.
I mean, to me, there are many structure prediction molecules out there.
I mean, models out.
there I feel like it's these other ancillary factors actually that are going to
probably be the most impactful in the usefulness of a of a product.
Yeah, right.
Yeah, absolutely.
The nice thing about structure prediction is there is a ground truth that you can compare against.
For design, you don't really have that.
You're like, here's some new, like, disease molecule.
Give me a binder for that.
And like, if you want to know if this thing really binds, you have to send it off to the
lab and wait a while.
For structure prediction, you can be like, all right, no model hasn't.
seen this sequence before, it's never seen anything close. Does it actually like fold up into the
correct shape? And we can just kind of hold that out of the data set and check. So I think I've
always thought of structure prediction as this really nice speed run kind of benchmark to like validate
ideas on. Right. Sorry, I didn't mean to say, I meant, you know, sort of structural models in
general. Yeah. But yes, exactly. So maybe we can talk a little bit more about get it, start getting
into the product side of things. Thank you for coming. I actually, I mean, like, like I said,
I really think this goes throughout not only for, you know, sort of structural models like this,
but also virtual cell and whatever.
It's really all the other stuff around the drug development process that is going to have the biggest impact.
So you talk a little bit about that?
Yeah, I think that's actually a good thing to talk about it after Chai 2 because I think
Chai 2 is where it's started to get really fun from a product perspective, right?
I think with Chai 2, we cross the threshold of usefulness where after we, you know,
release that paper, we had a lot of, you know,
pharma and biotex approach us and say,
hey, this model might be able to do some stuff for us.
Like, can we use it?
And they were like, oh, man, like, we should build a product, right?
We should build something to let you use that model.
And that's right around when I joined.
And there was sort of this, you know, mad, mad build out to both, you know,
build the product, which we can talk about the shape of
and also go and secure the compute, actually,
so we can go and serve those models to our partners.
And, you know, I think another third piece there that was really interesting
is around security and IP, right?
I think we want to be a very neutral platform
that anyone can design medicines on.
But as you guys know, like,
pharma is this notoriously IP sensitive industry, right?
And I think when I joined, a lot of people told me,
this can't be done.
Like, they're not going to put their data in a platform
and, like, have all their new medicines be generating out a bit
and having a bit of a background and security helped a bit.
Whereas, like, no, actually, if you, like,
just are really aggressive about how you, like, segment data
and set up, like, single tenancy
where you're like almost deploying a separate version
or a separate account in the product per customer,
you can actually like build a platform
and then go and ship it to them.
And so, you know, through the summer of last year,
we started doing that, right?
And, you know, we'd been working with, you know,
or talking to Eli Lilly.
And, you know, they were, you know,
one of the first partners to really work with us closely on that.
Kind of, you know, it made that V1 of that design suite, right,
that you can use to engineer some of those molecules.
on. And maybe it's worth talking a bit about that design suite, right? I think, you know,
I think we have these really, really powerful models from now, right? That can do like all of these
crazy things if you condition them in the right way. If you kind of give them the right context about,
you know, the structure that you're going after or maybe the constraints around the model, right? Like,
hey, I want to design an antibody that hits this GPCR protein, but, you know, it doesn't collide
with the cell membrane and also targets the specific epitope on that as well. And, you know, we
looked at and we're like, I guess we could put a chat bot around it. That'd be like really
easy to talk to, but like really like you're trying to build something almost very visual, right?
And you can finally build something really visual with some of these structure prediction models.
And so if you kind of look at the CHI product, it looks a lot less like a, you know, a chatcheebt
and a lot more like a autodesk or SolidWorks or Figma, you know, if you've used those things,
where you can kind of load up your molecule. There's this almost like Photoshop-esque like design
suite. You have this equivalent of a paint tool to kind of paint your epitope. You have this
equivalent of a content-aware fill tool to kind of get your binders generated from Chai.
You, of course, have a lot of the scientific analysis and plotting and whatever to understand
the results of the models. But we've just been surprised at how much complexity is actually just
in doing that right so that you kind of don't shoot yourself in the foot when you're then
prompting these models to give you binders. So are you sitting with people who are designing these
antibodies, you know, and then they're complaining to you or whatever.
Yeah.
How is that?
Yeah.
How do you convince med chemists to use your tools?
Because med chemists hate AI tools.
Like, notorious, like, I don't want to touch this thing.
Or like, I don't understand it.
And they will not touch things, which they do not understand.
Well, it helps a lot to have the models working really well.
Yeah.
Right.
So when we, you know, when we had the results of Chaito and Chipe 2.5, I think, you know,
that's enough of an activation energy where, you know,
pharma companies and scientists
when these companies are like, oh, let's try it.
Actually, can you guys just try
running the model against a few of these targets
and let's look at the results and then we do that
and the results are good and they're like, okay,
let me try to get on that product and let me try to use it.
No, I think pharma is like incredibly pragmatic actually.
Like I've been very impressed with everyone that we've worked with so far.
They're very, like I was saying, pragmatic about this
and they're like, they're willing to be proven wrong.
And like I actually don't blame them for not trusting the models.
Like, I have used these models.
They've been, like, rightly so.
Like, I am pretty skeptical when I, like, see any release.
I always have been.
So, like, you really just, like, need to show them the proof.
And, like, they can give you this target that they are interested in,
or maybe it's more of something they've worked on in the past.
They probably don't want to, like, share IP right out of the gate.
But they can be like, hey, you know, I've had trouble with this particular target in the past.
Let's see how you guys can do on this.
And then once you show them the proof, they, like, almost overwhelmingly or willing to accept that.
I come from a cybersecurity background or have worked on security products before.
And those were dark, dark years because you spend a lot of your time actually selling to people
who are surprisingly not that technical.
You think cybersecurity people are very technical.
In many cases, they are not.
And it is this kind of like uphill enterprise slog to this very unsophisticated customer.
I think we've been just pleasantly surprised, or I have, by just how much I enjoy working
with our partners and our customers.
You know, these are scientists who have been spending, you know, five, ten,
20 years of their life working on one target, right, often in some cases. And they've studied
everything about it. You know, they're very sophisticated. They're very smart, right? You know,
getting to collaborate with them is just a gold mine. And we learn a lot about how to make
the product better. You know, this anecdote, we, you know, a few months ago, we were actually
showing some of the results that we, from a target that with a pharma partnership. And
one of the scientists in the room, like, started tearing up and crying. Oh, wow.
And she was like, and we were like, what's wrong?
She's like, no, I've just been, I've literally spent 10 years trying to get an initial binder to this thing.
And you guys were able to help me do it.
Oh, that's awesome.
And, you know, that feels really special.
To answer your question, you know, we, you know, there's, of course, the teams of scientists and computational biologists that we're working with within, you know, each of our partnerships.
There's also the people we have within the building, right?
So we, I think one of the things that I really appreciate about Chai is how cross-disciplinary it is.
Like, you know, we have people who are maybe engineering experts and less bio-experts like myself.
We have great, you know, AI scientists, but or ML scientists.
But we also have a bunch of scientists that we work with and have joined Chai to sort of help us both, you know, test the limits of the models, right?
See what is Chai-2 actually capable of?
What targets can it do?
What can't it?
Inform some of the research direction there.
I want to add to that.
In like the Chai two days, like we kind of started with like a bunch of engineers.
People have like AI bio experience.
We didn't have a hardcore lab scientists.
And like one of our first hires on that realm was Nathan Rollins who I think he started working in the Baker lab at 14.
Graduated from Harvard at like 18 and got his PhD by like 21 or something like this in the Mark's lab.
And he was like super skeptical about Chai at first.
And then, you know, the results started to come in.
like, okay, this is kind of interesting.
Like, this could work.
And then, like, once the Chai T results came back, he was like, I need to bulletproof this.
Like, nobody celebrate yet.
Like, all this.
So I think, like, it's been really nice to have that level of rigor and to just have people
who have really, like, they've spent the time in the lab.
They've designed proteins themselves.
They've literally, in the case of, like, Andy, led several therapeutic programs,
brought drugs to the clinic themselves.
And, like, we have all these people internally at Chai, just, like, using the products
and, like, really battle testing that.
So if you don't have your own platforms, right?
So you don't have your own programs, right?
You're a pure platform, your partnership model, right?
Yeah.
How do you battle test something if you basically aren't,
you don't have a use case where you have to continuously push it forward?
Or if you are just pushing things forward when you just end up with your own candidates if you're successful?
And then what do you do about that?
I mean, we have benchmarks of our own internal cases, right?
You know, there's a set of targets that, you know, have our known therapeutics, right?
That have known therapeutics against them.
there's a set of targets that we pick to sort of push ourselves, right?
And so we're constantly refining that set and adding to it.
And that's what that internal science team that we have helps with, right, is expanding that
and almost running the experiments to try to get initial binders there.
We don't care about going and developing those drugs.
Like we just do that in service of validating and making our models better.
And then, of course, there's a loop with our partners, too.
Would you consider yourself hit discovery or are you, I guess, using some jargon,
hit to lead, lead optimization?
Like, where do you live in this?
And, you know, hit discovery might be like one part of it,
which you can do hit discovery.
But the later, the other parts of this are, I think,
oftentimes much more bespoke and kind of special.
I mean, how do you balance that?
And it seems much more difficult to me to be general than it does to solve
general lead optimization than it does to solve like hit discovery.
I think ideally, like we really want to be able to,
rather than think of this as a bunch of,
stages. I think part of the reason why we think of it that way is because the initial molecules
are usually like not good enough to be drugs. And like really like we're kind of at the inflection
point now. We're really seeing this internally at CHI where the models are getting pretty close to
like producing molecules that could eventually are like are very close to drugs. So we we try not to make
too much for a distinction in between, okay, hit discovery, lead optimization, all of the different
parts of this kind of preclinical pipeline are like, you know, the, the, the, the, the
the North Star is to just really produce drug-like molecules straight out of the models.
Of course, this is going to be hard.
And, like, they're going to be, like, tons of roadblocks.
And, like, you need to be able to, like, actually prompt the model to do this.
You need the whole RL stack to, like, learn different properties, things along those lines.
But I think it's very achievable.
Yeah.
And I think to add to that, right, yeah, the notion of target discovery and hit discovery and
optimization, where each of these has a gate and takes a few months to a few years is this
very, like, waterfall model, right?
where the cost of trying things and getting things early is very expensive.
But I think to what Matt's saying, right, if you start to get in a regime where you can have
models give you really promising candidates, you can start to make that look a lot more like
a loop, right?
It's akin to becoming more agile and software development.
Internally, we kind of have two, you know, North Stars, right?
And at first past, they almost sound like contradictory.
But, you know, the North Star and research is to start to de novo one shot, you know, better
and better and better medicinal candidates that are as close to being ready for, you know,
the next phase as possible. But, you know, also within product, we do want to sort of expand
into whatever these iterative workflows look like, right, where maybe I get a binder,
I get some results from the lab. I'm using that to condition my next run of the model.
And I think, you know, they sound contradictory, but I think they're actually not because I think
what's going to happen, you know, the research is going to get better at identifying a de novo
candidate for like a specific class of drugs, right? And say like anti-agnists,
right, like blocking things, right?
A little bit easier maybe, okay, we can get to a state
or we can one shot pretty good drugs there.
But now the next problem is like agonists, right?
Like how do you reliably one shot hitting a switch like on a cell, right?
Or by specifics or ADCs, right?
And I think, you know, there's kind of this levels of abstraction
that we're going to have to climb with the product
as like the models get better.
One of the things I was, I got very existential like a few months ago
because I was like, man, all this stuff we're building in the product
to visualize molecules and do this.
Like, maybe I'm just going to have to throw it all away when, like, Matt ships, like,
shy four, right?
But, you know, I think that's kind of the reality of, like, building products now, right?
You're actually using them less as an end in and of itself.
Like, maybe you'd have built software that was supposed to last, like, 20 years.
Now it's supposed to last maybe one year.
But it is the bridge to deliver value and kind of enable the research that then gets you
to the next thing.
And so I'd imagine we're probably going to rewrite our products at higher and higher levels
of abstraction, right?
Like maybe like right now we have something a little bit more akin to cursor where you're, you know, inspecting the molecule in the same way you're inspecting the code because you really need to verify like the bonds that are forming and the properties of the things that you're getting.
But then, you know, you get to a point where that stuff is solved enough where now the product is actually just helping you orchestrate these like campaigns of hypotheses, right?
Or maybe you have like one target and you're like orchestrating a bunch of different epitope choices or whatever against that.
And then maybe you're going up one level of obstruction where you're now doing a whole.
whole campaign against all of the targets within a pathway, right?
And I think what's really exciting about that is if you have like these really good primitives
for structure prediction and binding and design and you can kind of compose them, then you
can start to just like grow into like the outer loop of science, right?
And then, you know, maybe the thing runs itself and you start to really get to some really,
really, really cool drugs at the end of it.
I actually want to push on what you just said about epitope prediction because I think
A lot of people in the field would argue this might be the much harder problem than finding
antibodies and binders.
Where do you think that the state of the art is in general and also with regards to chai
in terms of epitote prediction?
And like, is this a problem which has a reasonable, solvable time horizon?
Oh, and also maybe can you define epitope prediction?
I'll think of this at like some different levels.
So the most basic level is, okay, I have some disease that I want to target and what
proteins are actually responsible there.
Like actually figuring out biologically what's going on, like what should I be targeting
in the first place with a drug?
Because once you figure that out, it's kind of like a structural biology problem at that point.
You're like, all right, this like set of proteins is responsible and like what's going on there?
Well, this is interacting with some other protein that it shouldn't be interacting with.
And conventionally you just like want to block that interaction or something with anybody.
But that's kind of where these proteins interact and like the type of interactions that you want
to disrupt, that's typically like the epitope.
It's like the actual site on the protein that you want to block.
This is a ridiculously hard problem.
I'm witty on this.
This is like the harder problem.
Like just the amount of context that you need and like the global understanding that you need to get in order to like actually figure out what's interacting and how.
But maybe let's take a few specific cases.
Let's think about what about SARS-CoV-3 comes around or the new flu or whatever.
What would you do there?
I mean, is that something that you think you could actually reasonably tackle?
In that case, like, yeah, you could just run a structure prediction model maybe and like see where the model thinks this thing will bind if it's highly confident in that.
You might say, okay, here is like the site that we want to block.
I think in general, still very hard.
And even like structure prediction, it's getting really good.
And like a lot of people think Alfold 2 like solve structure prediction.
Not really.
Like Alfold 2 got like, I think 11%.
The multimerary version of this got like 11% of antibody antigen prediction cases correct.
That means 90% of the time it's wrong.
Yeah.
But AlphaFold 2 solved a certain class of monomeric proteins with MSAs.
Absolutely.
Yeah, yeah.
Yeah. So, I mean, that's the MSA, I think, might be the key point here because MSAs are sort of the magic, which makes it all work.
It's like a, it's a template in some sense about what the structure should be.
And antibodies almost evolutionarily can't have a template, right?
Yeah.
Everyone has to have unique antibodies custom to the things that they've experienced over the course of their life.
Yeah.
So.
Right. And just to clarify, I had to understand this myself, so maybe I can help the listeners who aren't familiar.
An antibody, the whole point of an antibody is it can be used by the immune system,
identify new things that the body has encountered before.
So the design of antibodies as opposed to other types of proteins is to the system is designed
so that you can quickly recombine different components of it in order to match proteins
that are from unknown pathogens, more or less.
And so this is why it's not conserved in evolution, the way.
that other parts other proteins are yeah so so like back to the the episode prediction problem um i i think
it's still hard i think like there there are a lot of cases that that are maybe tractable but i think
in general like if you want to discover this for a new target uh still still a really difficult problem
maybe virtual cell would be like the closest thing to the state of the art there but that's still
still a ways out i wanted to begin a little bit on the product because i there's something i don't
understand about the economics of basically all all the structural stuff
that's happening right now. And obviously, a lot of people think it's very, very valuable. So
there's, you know, I'm not grocking something. But when you look at the cost of developing an
antibody, you know, it maybe is a couple million dollars, right? When you go from, you, you've
identified a target somehow. And then you say, okay, I need an antibody to match this. And then I have
to sort of optimize it in various ways. And then maybe I try it in, I mean, with antibodies, you go to
the animal typically faster. If you look at how much does it cost to bring if you like are
pressing it and pick the right target and the right technology to get all the way to drug,
it might be half a billion. Typically that $2.6 billion number is amortized over all the
failures as well. So if you look at just the cost of that one success, depending on the
disease may be less, but you know, half a billion might be a good median number or something.
So you're saving like a couple million dollars in a half billion dollar campaign.
So why is this so attractive?
I would maybe challenge the premise a bit like in a few ways, right?
Like, okay, sure, if you're trying to get an antibody for like a very simple kind of target, like maybe, right?
But I think what we've been most excited by is our partners using antibodies in, you know, more sophisticated ways, right?
Like in, for example, in Chaito, we showed like GPCR agonist activity, right, where you can really hit the switch on a,
you know, on a cell doorbell protein, so to speak, right, in a very precise way.
Very, very, very hard to do that with antibodies if you can't be that precise, right?
So you're unlocking a new capability?
Yes, right?
I would think about it as less like, oh, I'm taking the existing drugs that I can do
and making the faster.
I mean, there is some of that too, right?
But it's like, no, they're just like, hey, how do you go after like better targets, right?
That are, you know, maybe more precise, more effective, right?
I see.
I think, like, also on top of that, too, is, like, there are drug modalities that you just can't
discover with immunization. You're not going to design your crazy, multi-specific, war-headed, super-intense formats.
These are really things where you kind of have to design these from first principles. Even just
with Bios specifics in particular, both arms need to now bind different targets. And you've kind
of like have this multiplicative effect on your binding rate. So like if you have a one and a billion
chance to find a binder in arm one and a one a billion chance in arm two, you're not. This isn't going to
work for the traditional approach. Exactly. I think about is, right, you're not just helping your
partner with maybe one drug, right? There might be a portfolio of targets that, or they're going
after a portfolio of drugs that they're trying to make. And the nice thing about the platform
approach, rather than that we're developing individual drugs, so we can sort of scale with them
as they pursue more targets in addition to more ambitious targets. So it lets you concentrate
your learning subdomain of that. And so that you, everybody,
benefits from that. Exactly. That's the, but okay, so I didn't, so what is, what are some of these
capabilities you mentioned a few? Are there more that are really interesting that you guys are
chasing? Yes. I mean, and we talked about like, you know, cross-reactivity. We talked about selectivity.
We talked about some of these like really interesting additional modalities with by specifics, right?
There's a set of things that, you know, our partners have been asking us for that we've been working on
that I can't get too into because then that starts to reveal some of the, the targets that they're going
after. But I think the point being, you can just, once you get precise, like, you can start to do
some really, really cool drugs. It's a new technology, right? So, like, technology in pharma means, like,
how do you deliver your therapeutic? And so this is maybe kind of thinking about, like, Carty is a
technology, right? And so this is maybe a new technology in the sense that you can have these highly,
highly engineered. Right. And that comes, you know, from the mission of the companies to really turn
you know, drug discovery from a scientific experiment to an engineering discipline, right?
How do you sort of get to the precision engineering phase for biology, where you can start with,
you know, almost declaratively define the thing you're trying to get and have the model fill in the gaps
and get you that?
So what is the biggest blocker from going from science to engineering?
Oh, man.
There's so many things.
Like that's the thing about, you know, engineering.
Yeah.
Yeah.
Like her heterogeneous.
What is that?
Define.
I don't even want to talk about this.
Like the amount of headaches.
Too late.
You already met, probably not.
Yeah.
You're committed out.
Okay.
So, like, just, like, when you're actually parsing, like, first of all, file formats for biologists.
Like, I just, they just don't care.
There's, like, no standardized.
There are standardized file formats.
Are they the best?
I don't really know.
But there's, like, also just, like, a lot of information that you want to pack.
I have this structure.
Here are the people who solved it.
This is the method I used to solve it.
There's, like, a lot of stuff going on.
And then depending.
on the method that you use to actually figure out what this 3D structure is, you might have
multiple copies of that structure. Part of it might not have really been resolved. You're like,
it could be here, it could be there. I'm just going to give you like both options. So like the
actual just parsing problem on the engineering side of like working with this type of data is like really
difficult. This seems like something that LOMs can excel at though. They don't know all the edge
cases often, right? This is more back to just like a simplicity approach. Like LMs are very good.
I will absolutely give you that. Then you're thinking about like, do I really want to like, should
this function have 20 special cases or should be like really principled in how we approach
this? And should we be, I guess, more of a opinionated? Opinionated. Yes. Like how opinionated
should we be and how we do this? We want a strategy that's like easy enough for humans to
understand and like when we're reading through the code base. We really need to know what's going
on here. What are the potential problems? And like sometimes that just comes down looking at examples.
But then I think, okay, once you've kind of figured out all the info work and how you get data into the models,
there's then like scaling the model.
There's then scaling the infrastructure around the model to train bigger and bigger versions of this.
And that's like a lot of work that Neil and the product team actually leaves.
Yeah, I mean, that would have been my answer is the infrastructure part.
I mean, you know, not to beat a dead horse, but compute, right?
Getting the compute and using it in the right way is such a challenge.
Especially for startups.
This has been such a theme.
Yeah, yeah.
I think like anthropic is single holding back science.
Exactly.
No, I mean, and to that point like we, I mean, they're also accelerating science, but it's like this weird.
No, totally.
Like one of the things that I help a lot with at Chai is buying compute for the company.
Worst job, man.
I would not recommend it.
It is very stressful.
But, you know, even September of last year, right?
You're back to you, the hardware job.
Yeah, yeah, I know, exactly.
In the wrong way.
But, you know, September of last year, we started to really notice, like, things are getting, getting tight, right?
We were doing a lot of our inference on, you know, spot and on-demand markets.
And we'd have these days where you'd just, like, get these capacity crunches.
And we're like, okay, we should probably start to get ahead of buying some compute for ourselves.
And, I mean, I think everyone probably says this, but, man, it was, it was hard.
Like, I think I didn't, I didn't realize how much of a power law, you know, this is, right?
where, you know, there's, say, 10,000, you know, B300 units that are shipping everywhere, right?
The hyperscalers and the, you know, the biggest AI labs are buying 95 plus percent of it, right?
And then you kind of have the startups, like, fighting over the scraps.
And I think the other thing that's really interesting, especially if you look at these later
computer versions, right, the Vera Rubens or, you know, the B300.
Like, a lot of this stuff has been built very, like, LLM for it, right?
Like you have these systems with like huge KV caches where you have like 72 GPUs that are all
acquired to talk to each other, right?
And, you know, obviously some performance gains there like help us, right?
But like it's kind of interesting just how much the compute market has kind of gotten
LLM pilled.
I think there's like a whole probably set of, you know, compute stack and inference optimizations
and things that need to be made for this class of models.
And, you know, I think this class of models is going to be like just as big, just as impactful
as LLMs.
But it's almost like the compute market like kind of doesn't realize that yet, both in the capacity sense but also in like the software stack sense.
So we actually spent a lot of our time, you know, even just like doing basic optimizations of compute to like get them to work better for the types of models that we have.
Yeah, I know that some structure models are more recursive than than LLMs, for example.
And so that that which changes sort of like you're the maybe the compute to memory ratio that you need and things like that.
What are some of the, like, sort of cool or interesting optimizations that you've done there?
Depending on that type of model.
So, like, we can go back to, like, a Chai Wan-type model.
In that case, we're following the outfold two, three architecture.
And there, you're, like, rather than doing attention over, like, this, like, normal sequence representation,
you're, in a sense, loosely doing attention over this pair representation.
So you can think of this as, like, a sequence of length L squared rather than, like, typically, length L.
If you're doing attention over that, the way that you actually batch this up and ends up being L cubed,
now you're in like a pretty, pretty heavy computer regime.
So the amount of flops that you're putting into every token stays, it's pretty high,
the amount of memory that, like, the memory bandwidth overhead of just transferring that from, like,
S-RAM to whatever, that's a real bottleneck in these architectures.
So, like, even something as simple as like a layer norm can take a long time, actually.
Like, that can be a significant amount of the compute that you're using.
So I think on our side, we've spent a lot of time just optimizing it and engineering, taking engineering very seriously so that these operations are at least better.
We're always looking at how to new chips perform compared to the older versions.
Sometimes that's even different for training versus inference.
And like, of course, Meal knows this really well.
Well, so there's, you know, what you're doing on the individual GPU.
And then there's like, how do you like orchestrate fleets of GPUs, right?
And, you know, you basically shard your computation, right?
And so, you know, when you're designing a molecule on Chai, it's not necessarily like one call, right?
It's a lot of, a lot of GPUs being thrown at the problem, right, across a lot of compute.
And actually, I would say that one of the hardest things to get right in software engineering is durable execution.
Are you all familiar with that?
Can I go on a little like...
Yeah, yeah, yeah.
Ultimately, like, if you're, like, computing a lot of data, you know, model calls across, like, a very wide set of infrastructure, you always run into these problems where,
like some part of the infrastructure is flaky, right?
Like maybe the bucket you're grabbing your data from like goes down or like your database
has a blip because they're like too many transactions against it or you're like
GPU errors out, right?
I've been at companies before where you like spend so much your time just dealing with
this shit, right?
Like you're basically putting like all of these cues and like all of these retries and
you're like duct taping things together and you have a and it becomes this mess where now
what used to be like a ideally like a pretty simple like computation that's just
distributed, you're ending up spending like 95 plus percent of your time on all of this
queuing and retry stuff, right? We're huge fans of this company called Temporal.
Basically, you know, there's this idea, like, look, if you're just trying to get something,
a really long-running job to run, at the end of the day, what do you need? You need a queue.
You know, you need your flaky thing, like pulling off of the queue. You need some retry logic
to put things back on the queue if they fail, right? And then you need some whole like orchestration
system to just like tie all the cues together and monitor them.
What's really cool about temporal is like this is a tech, a company that's kind of invented
a framework for doing this.
And one of the, I think one of the technical decisions we made early on that was very helpful
was to run as much stuff as we can on temporal, right?
So whether those are, you know, calls out to the database from the app, right, to make sure
the database transaction goes through without failing, okay, let's have side effects like sit
on temporal so that they get retried smartly without us having to like write a
our own Q logic, right? Or things related to model calls or things related to orchestrating really
long data pipelines. Point being, like, you know, one of those primitives, like, just like,
hey, you need to get durable execution right so that you're not stuck in, like, retry hell.
A really deep, like, engineering thing that, like, you wouldn't realize if, unless you,
for like me and Jack, you've been, like, burned by this, like, many, many times before.
And I think, like, we're at this state now, right, where we've, you know, we've raised another
$400 million. I have to go buy another.
their compute cluster, like, you know, like, we're going to have like really, really,
really large runs and inference and training sets.
And so getting those foundations right is what's actually going to let us do more ambitious
things.
And to kind of answer your question, actually, that's a lot of the bottleneck to making,
making biology more like engineering is just like having the right engineering primitives
supporting it.
I have an analogous tangent on the model side.
Actually, one of the things that's kind of nice about those problems is they're like super
visible. So like, at least you know, like, hey, this crash, this failed. For us, we just see,
like, loss curve didn't go down or like we see weird gradient behavior or whatever. I think
a lot of these same principles, like, you know, engineering first, that also applies on the
research team. One thing that I like to say is kind of like complexity and being bitter
lesson-pilled or, like, fundamentally at odds. For example, I think like outfold three, I might get this
number on, but I think it was like 23 sub-modules. And at that point, that's a really difficult system to
optimize and study. You're like, all right, what happens if I change? Like, if I tweak this thing
in submodule 30 or like 21, what happens to the whole system? And you can always think,
hey, we can make this better by like adding module 24, but like should you? Or should you
think about just like removing things and lowering that complexity down? But I think that's like a pretty
fundamental thing at Chai is just like the engineering culture and just being like very
simplicity biased. Have you all seen the picture of like the SpaceX engines? It's like Raptor 1 and it has a
bunch of pipes and like Rapture 3.
We have a picture of that like on our office wall because I mean it's just true, right?
Like how do you delete, delete, delete more things?
Yeah.
But the only way you can accomplish that is, I mean, the reason Alpha Fold 2 and Alpha Fold 3 worked,
they were small models, relatively speaking, they were very compute intensive, but they were
very data efficient.
Yes.
And like the, there was inductive bias after inductive bias brought in by human intuition and
probably like hard fought experience.
It was, they're incredibly efficient.
If you try to knock down those things, you know, they're not like a house of cards.
Like everything is a incremental improvement on top of it.
In order to get beyond that, it seems to me like you really just need new sources of data.
You would need at least treat data fundamentally different in a way that is much more efficient.
I mean, I mean, I actually kind of surprised to hear that you have scale to that degree.
because it suggests that you're doing something very different
from the way the community is thinking about it.
I don't know if you can comment about that.
We're pretty first principle.
The whole research team at Chi, except for me and Kevin, really,
like we're the only people with, quote, bio background.
Even still, like, we're pretty far removed.
So I think, like, we try to look at every problem as a core ML problem.
We try to think of, like, what's the analog in other spaces?
So, like, even for image models, like, CNNs were built to
process images. So, like, images should be looked at in patches. Like, that was the nice
inductive bias there. And then people are like, well, you can just kind of tokenize this
thing, throw it into transform and it's going to work. And, like, it did end up working. Even, like,
on a relatively small dataset. But I think for proteins in particular, it is really hard. There's
not as much structural data. There's a ton of sequence data. And, like, that's one of the
unlocks for, like, ESM working. You can get that to just run on a transformer. If you try to do
the same thing with, like, experimental structure data, good luck.
No, not going to work. You need outful. Yeah, absolutely.
There was that Apple paper where they distilled on the Apple Fold, which it was actually really cool that you could distill on a very large data set and you could get good signal.
But it didn't generalize at all because it wasn't reasoning.
It was really pattern matching.
Like one of the things, these triangle layers you were talking about, for example, they do have a very nice inductive bias.
Maybe it's not the triangle inequality like the paper originally proposed.
But it's a clean inductive bias.
And it unambiguously is like one of the things which made it work.
and it just comes at a huge cost.
Yeah.
Yeah, no, I think that's definitely true.
These layers are pretty costly.
And like, that kind of limits what you can do with the architectures.
They're not, like, not only are they like costly in terms of compute.
They're just like not efficient on modern GPUs either.
You have small hidden dimensions, large sequence dimensions.
Like, it's like exactly the opposite of what GPUs are designed to process.
One takeaway from like triangle layers is you're kind of just trading off parameters for compute in that sense.
Like that's like one mental model for thinking about this.
I might want to, like, throw more compute at the problem and just trade that off for parameters.
Because, like, I won't be able to hold as many.
Like, I can't literally store these, you know, large pair representations and still do normal attention.
So I think there are fundamental things you can abstract from the ideas like Alpha Fold,
but you can kind of just, like, tweak these and start billing off of them in your own way.
It sounds like you have quite a bit of research, like fundamental research going to this direction for, I guess,
audience looking for a nerd snipe and ML engineering for new problems, probably something very,
it's a very different research direction than a lot of the communities going in. Yeah, yeah, I think
what we built at Chai is like, it's very unique in a lot of ways, but also very tied to like what
CoreML is good at, kind of what I was saying before. Like, we try to map every problem into like
a CoreML problem. We think, you know, how would you approach this if it were an LLM or something
like that. But yeah, like at the end of the day, we really, really value simplicity. And we, we really
encourage people who don't have a bio background to not be scared of this stuff. And I think that extends
into the product too, where, you know, there's a balance to be had here, right, between, like,
how general do you make the product? Like, do you build a cross-reactivity workflow and a selectivity
workflow and a bi-specifics workflow? Or do you all say, no, like, let's make the model general enough
to say I'm going to like condition on arbitrarily binding or avoiding something. And then you just
have a very general like screen in your CAD suite where you can say, hey, I just want to avoid
or bind to these parts of these different structures, right? And I think, you know, kind of like the
ML team, like I don't, I don't have, you know, a formal bio background. Most of the product
and platform team doesn't have a formal background either. Now, there's some amount of like maybe
regretting my words that I'm going to have right, because I'm sure there are, you know, a million nuances
and, you know, I don't want to come off as, you know, too, too brash or naive there.
But, you know, I think sometimes it's helpful to not be burdened by like all of the,
oh, this nuance and this nuance and this non and this nuance.
And you get to kind of bet and be maximally general because, you know,
that's kind of what we're seeing in the research.
You can, the models are very general that lets the product be very general.
I'm thinking back to like in my CS theory days, my first advisor was like,
we're working on some problem and we need like a polynomial time algorithm for something.
And he would always tell me, like, never underestimate the power of polynomial time.
Like, it's basically like you're allowed to choose, like, whatever exponent you want.
And my first paper was an end to the 20th time algorithm for this problem.
And I was like, Andy, I did exactly what you said.
He's like, wait a minute. I didn't meet it like that.
Yeah.
But I think like it kind of, like you can really help yourself.
Like you can free yourself up a lot when you're like, all right, I can kind of do whatever
I want and then kind of simplify it later.
And I think that's really like a pretty fundamental way of thinking about things that we,
we leverage a lot at Chai.
The space of binders, of protein design and binders in general is actually
a fairly crowded space.
I'm curious about what your general outlook of the field the industry is.
I mean, I can go back to like some anecdote.
I was maybe NERIF three, four years ago,
the one right after RF diffusion came out.
I was talking to someone in the Baker Lab and they're like,
man, I just one-shoted, I don't think they even used one-shot.
One-Wson wasn't even a term back then.
But they just like, I just got picomolar binders out of R-R-F diffusion
and just like threw in the cryo, great, right?
it didn't seem like that just solve the problem.
Like it's not like, oh man, now every, yeah.
But there are lots of people who I think have seen that you can actually do protein design,
at least in some categories, quite well.
I'd say, like, is it many proteins or mini binders?
Ironically, nanobinders are actually smaller than, or larger than many proteins,
or maybe like a little bit harder.
Antibodies are typically considered even harder.
But there's this like, is this something which can.
and will be commoditized, at least in some part, how do you compete? Like, where does this,
where do you, where does the field go from here? I mean, I think the answer is, it's kind of all
of the above. Like, I think there probably will be some commodity layer for, for certain types of
modalities or drugs, right? I think at the same time, we're going to be able to do even more and
more and more ambitious drugs. And you're going to know, it's just like what's happened in LLM land, right?
Like you have your, your open source models that are maybe general and helpful for some things,
but people are still buying frontier models, right?
And actually, if you look at the amount of value captured,
it's actually the closed source frontier models,
you know, the whole pie is growing,
but it's growing so fast that even as the open source models,
like share expands,
the frontier models are still able to capture the majority of the value, right?
Raise your hand if you're using an open source model on your day-to-day.
Right.
And what are the reasons for that, right?
One, like, if you have, you know, more intelligence,
you're going to go after harder tasks, right?
I think if we have more, you know, intelligent bio models we're going to go after more,
more crazy biotasks, right?
But then also, too, like, I mean, a lot of the reason I don't use the open source model is
because, like, you know, I don't get like clod code, right?
I don't get like clot.
You know, I think there's a product layer to be built that is just as important as the model layer.
We learn a lot from our partners and, you know, the people in the building as well,
just like, what are the really tough things that they get stuck on using the models, right?
And some of them are like, you know, the dumbest things, right?
like, you know, I want to be able to better visualize this piece and, like, focus on that.
And some of them are actually, like, very sophisticated things that we then have to build some,
like, pretty vertical product for.
And look, maybe in the fullness of time, like, AGI, like, one shots everything and doesn't matter.
But I think there's quite a bit of time until we get there, right?
And I think the product makes a huge, huge difference for that.
That'd be my answer.
I mean, you probably have a more model forward answer.
No, like, I think, like, like, biology is slow, which is, like, one kind of nice thing.
And there's, like, not that much labeled data.
So, like, you could take all the publicly available sequence information out there.
That might give you a good base model.
But you still need some measurements on that data.
That's still pretty time-consuming.
And then you need to, like, iterate on that.
So I think there are even just data blockers there.
And unlocking, like, if we really want to do this zero-shot design candidate,
start generating molecules that are almost ready to go into the clinic,
I think that's more to that than just, you know,
AGI might not solve that right away.
I think there are definitely like some technical blockers there.
But even in the space of, you know, specialist companies,
I mean, I'm not going to like to start naming them,
but there's, I think, I don't know, probably 10, 15 protein design startups.
I think the two things, which it sounds like Chi has gone on is like one,
all in on product and two, you are not trying to do your own platform.
If you don't have your own data mode, you know, is that going to like help you
want out in the end or is that going to be, you know, a blocker?
I don't, I'm just, I'm just curious about that.
Yeah, that's a great question.
Yeah, so, try definitely no plans of, like, starting a pipeline.
Like, we take the partnership model pretty seriously.
And we, I just like, from a personal stance, I love the incentive alignment.
And between, like, you know, we make the models better.
The partners succeed more.
And just, like, you know, that iterates on itself.
So, like, I think that's, like, a pretty unique part of Chai is, like, one, just being
able to partner with a lot of people, two, getting, like, the feedback on the product.
So, like, you know, knowing that it's very real, this is in, like, like, legit,
big pharma hands and they're actually running campaigns on this stuff. So I think it's interesting.
We really have to be model forward, model focus. Like we need to keep delivering value. So that puts
a lot of pressure like on the research team. The product team, first of all, to like to serve these
things, the research teams always shoot for like better and better versions. The way I think about
this is like if you're a bitter, less impilled forward kind of like thinker or a company,
then there kind of comes a certain point where there's a lot to do on like,
both the model and data side.
But I don't think either is exhausted.
It would be stupid to say, like, we don't need any more data,
but it would also be stupid to say, like,
the models are stuck.
We only can, like, use data to solve these problems.
So I think there's, like, tons of room to grow on both sides.
We're taking, like, both very seriously.
And I would also maybe push back on the no data moat premise, right?
That'd be kind of like saying, hey, like,
all the enterprises that work with Anthropic,
like, you're not letting, like,
entropic train on their data.
So, like, you can't, like, build models
that are good at enterprise workflows, right?
I think, you know, one, we are investing in this, right?
You know, there are ways to turn compute into data and get more and we're doing those, right?
But then also, too, okay, what is the kind of data that you're trying to get, right?
And I think what is kind of cool about, you know, working so closely and supporting so many of these partners is we get to really learn about, you know, what is like the stuff that that would be helpful in research, right?
And so rather than doing research in a vacuum, you know, based on what would hypothetically be cool, we're able to sort of kind of do informed research based on, like, you know, what our partners have just been very organically.
asking us for help with.
I see.
Do you, I assume that you are allowed to train general models based upon your partner's data.
Do you train specialized models for, like, is there a No, artist model and a Pfizer?
Yeah, I mean, like a lot of these deals, you know, and this is all public, right?
We are working with them to, you know, train or fine-tune a version of our model for them.
And I think there's probably like so much more we can do there over time.
My brother started a company called Applied Comput, a great company.
they're kind of doing this thing for, you know, design, for LLMs, right,
and helping enterprises really understand the value of their language data
and do that for specialized tasks.
I think there's a whole world where we could potentially do that for biological data.
What is the value there?
Like, what is the lift that you get from using their data?
I mean, is it just that it's more data or is it more that they're specialized to a problem?
You know, they have a lot of, like, scientific, you know,
data that they're getting from experiments that can maybe help our models do better
in like particular classes of candidates that are targets that they care about.
Yeah.
I mean, even something as simple as like they might just have some preferred way of doing things
that might not be like native to the Chai model.
And they can like, you know, kind of like ask the product team and in a sense to just be like,
hey, we like, you know, our designs have property X.
Can you make sure that they have those?
So I think like even things as simple as that actually have like a pretty big impact for them.
Yeah.
So, I mean, this goes along with the,
a pet hypothesis that I have that all AI companies and especially bio and scientific ones are
actually consulting companies. Pharma, I think, is particularly the case because you're developing
a new drug, right? It's almost by definition new, right? So like the existing stuff has to be
customized in many cases, right? Unless you're doing something that's just reiterationable with stuff.
But a lot of the big pharma are pushing the boundaries of science.
Yeah, I mean, certainly like we aim to make the models very general.
We aim to make the product very general.
We aim to make it powerful.
But, yeah, I mean, there is integration work, right, with every customer.
To answer your question, you do get some defensibility just by doing that, right?
And I think what is nice about building, you know, trusted relationships with these partners is hopefully, you know, if we execute really well over the next, you know, the first year, then they'll continue working with Chai to do more ambitious and more drugs past that.
I mean, it's going to be hard to switch, right?
I hope so, yeah.
Just getting the security review.
Yeah, yeah.
Like, maybe one other interesting point is, like, if you think of this, like, on a per
token basis, I don't know if there's another domain where, like, the downstream value
of a token is, like, as valuable as it is for pharma.
Yeah.
You know, like, you're thinking about, like, the actual drugs that come out.
Like, these can be, like, multi-billion dollar assets.
So, like, the case of GLP-1s, I think the two GLP-1 drugs combined, like, maybe a trillion-dollar
asset, like revenue stream?
Up until, I think, three months.
ago, right, GLP1's like total revenue was more than all of the AI labs put together.
Yeah.
I don't think people realize that.
Like, I didn't realize that.
It's crazy.
But yet the market count way lower.
It's like crazy how relatively speaking to market cap is.
And, you know, I didn't realize how much of like a VC business, you know, you know,
pharma is in, right?
They're in some sense, like, taking really ambitious bets.
You know, I think one of the things that was really cool, you know, is like if you study
the history of Silicon Valley, right?
Like, obviously people think of Silicon Valley with software, but, you know, in the 80s, one of the biggest venture outcomes.
One of the first ones was Genentech, right?
And because it is such a VC model, right?
You get the string of tokens that can then give you so much value downstream.
Just just general shout out to outpostings, blog series about, like, finance and funding.
Yeah.
Yeah.
Yeah.
Before that, I knew a lot of those points, but I did not realize just how deep.
that rabbit hole win.
Yeah.
Yeah, I mean, it's maybe the single biggest problem in biofarm is actually just the funding
model.
There's also, have you heard of Aram's Law?
Yeah.
Oh, yeah.
Yeah, it's more backwards.
Yeah, more's law backwards.
So it's like, and like compute, you know, it's kind of scales.
So you have like this nice exponential scaling, log layer scaling of compute, and you have the exact opposite
in pharma.
So like the cost of actually making a drug in pharma is kind of like increasing exponentially.
So the amount of money put in per drug is growing at kind of like an exponential rate,
which is it's pretty interesting to see this, yeah.
Which guarantees at some point the marginal return on a new drug development will be negative.
Exactly.
So unless someone, I maybe Chi, figures out how to, you know, fix this.
I think that we might be on the verge of sort of flipping some of these.
Face transition, bending the S curve.
Yeah.
Just a double, maybe belabor the point, but that pharma and VC,
fundamentally both are optimizing a portfolio.
Yeah.
And I think that's the,
that's the connection there.
Yeah,
thinking of pharma as like sophisticated capital allocators, right,
where they have this portfolio of targets and they're allocating between them,
I think that was a big reframe for me.
Yeah.
And I think we will just see more of that in the future, right?
And hopefully they can take, you know, in a sense, the VC taking riskier bets.
Like hopefully pharma can take riskier bets and pursue really, really cool target targets in the future.
That analogy is actually like,
one, the kind of like VC type investor-ish model.
It's like actually how we think a lot about research at Chai as well.
Our research team is relatively small.
I think definitely compared to like a lot of the like the isomorphics, deep mines.
Like our research team is like, you know, in the around 10 people.
So like we're a relatively small team, but we kind of think of it as almost like an investing job.
We're like you're investing ideas towards compute in the same sense.
You're really just capital allocators in that respect.
Yeah, I actually think maybe this is too cute, but I would even make.
the broader point, which I think every, we kind of think of
of everyone at CHI as a bit of a capital allocator.
So I think one of the things that surprises people is we're
pretty small, we're only 30 people.
And that's because everyone we hire onto the research team or the
engineering team, you know, especially now that they're in some ways
like very empowered with AI, a lot of it is just like allocating
their attention into the right ideas and allocating their compute.
This is actually, I think, a characteristic to some extent
of mission learning AI projects and also science.
Right, whereas if you're building like an API for some B2B SaaS company that's not building foundation models, whatever, your limit is mostly people, right?
So the resource you're allocating is almost entirely people.
Whereas if you're building hardware, you're building AI models, you're building something scientific, then your constraint is those, the resources that are, you know, the bottleneck is, you know, the lab, it's the compute, it's other things.
and so that you have to really be in that mentality of I have these limited allocation of I have some shots on goal.
How do I allocate those shots?
Well, I would say yes and no.
So I agree it's a bit more like that, right?
But like let's going back to the example of building an API for, you know, a B2B company, right?
That API has incremental cost.
You have to support it.
It adds complexity to the product.
It's another thing you have to go market and sell.
Maybe you should actually be allocating that into like a different bet, right?
A different thing on your product roadmap that you should be prioritized.
instead of the other thing. I think in a world where, like, building things just gets, like,
really cheap and, you know, increasingly free, the scarce thing is the attention, both that you
can put into it, right, to keep your product simple and grokable, and that your customer can put
into to, like, really understand how to use it. I see it less as, like, a binary thing and more just,
like, we're all kind of as engineers going to be a little bit more, like, allocators of attention.
Yeah, which is what executives are. We're all going to speak coming.
Well, I mean, like, there's a sweet of podcast of Satya, Nadella, right?
He says, you know, Microsoft wants to make everyone a manager of infinite minds, right?
If you, like, really take that to your extreme, like everyone's going to be an executive.
I mean, I certainly feel like an executive when I talk to a quad every day, right?
A little suite of interns who are all going out and eagerly solving problems.
You may or may not have actually wanted, but they're solving the problems.
Yeah.
So we have two typical questions that we asked that we've already kind of asked one, but I'm going to ask it again.
maybe more directly, is if you, and you can both answer this, if you could remove a bottleneck
from your problem space by Fiat, what would that be?
That's an interesting question.
I think one thing that would be really nice.
Like, just I'm like always in research land, very hard to turn off.
For me, it's probably just the validation loop of protein design in general.
So like just being able to say like instantly like, hey, this thing works, this thing doesn't.
there's still a bit of walking around in the dark that you're doing.
Just like, you know, you have some ways.
And like, I think at Chai, we've taken this, like, very seriously.
But it's probably along the lines of just, like, validating hypotheses and, like, you know,
knowing for certain that things work.
Yeah, that's unsolved problem for sure.
Yeah.
Yeah.
And it would be hugely valuable.
Usually valuable.
Yeah.
Yeah.
I'm going to take a much more abstract answer to that, which is actually, like, talent obscurity.
I think, you know, there's a lot of smart people going and working on LLMs.
You know, there's a lot of people that are working.
and becoming software engineers for SaaS, right? But I think just like not that many like smart people
go and work on bio. You know, I didn't work on bio like in high school because I was like,
oh, I could like pick up my computer and program apps. But if I want to work on bio, I have to like
go study and get good grades in school and like maybe get a PhD or whatever, right? And, you know,
maybe that's one reason for it. I think another reason is, you know, a lot of the stuff is really
obscure, right? Like we threw around a lot of big words during this podcast. You can't really
visualize the things. It's one of the things we care a lot about it is like how do we make the
whole thing feel visual on our website and in the product. And, you know, part of the reason we're here
is like, you know, I think, you know, more people should realize, like you don't need to like have
like a super, super, super, super specialist bio background to contribute to this like computationally.
And so, you know, I think a lot about like talent flows and like where talent goes in the economy
and right, you know, in the 90s, everyone was flowing to talent. And, you know, since the 2000s,
people have been flowing to tech, but, you know, big tech, like, ate up a lot of the talent,
you know, until, you know, a few years ago and now maybe like LLMs and the big AI labs
are eating up a lot of the good talent. But it's like, you know, at the metal level, like,
how do you allocate talent better? You know, selfishly, I want more talent going to buy.
I mean, we probably want more talent going to manufacturing and physical world things and
these other problems that the U.S. has. But, yeah, I think communicating that better would be
the thing that if I had a megaphone to talk to everyone, I would try to do that.
Okay.
So then that leads to the second question, which is, and maybe the answers the same, but what
is the takeaway, one takeaway, that you would like to people to have from the episode?
Yeah, I mean, I think, you know, biology has been this somewhat obscure feeling field where
you're stumbling around in the dark.
You don't know what you're looking at.
you're dealing with non-determinism in your experiments.
You're having to do a very long and iterative trial and error loop
across a very, very long amount of time.
And at some point, you're crossing that threshold
of what you can do computationally.
When you can get folding models down to being within, you know,
an angstrom, right, where you can get design models
to give you, you know, hit rates, you know, north of 50 percent,
or now you can put them, you know, in a 96 well plate
and actually have like 48 interesting binders,
you start to get to the point where now you can declaratively precision engineer what you want,
rather than betting on nature or trial and error to get you there.
And I think that, look, we had the same thing happen in software
where you can write code and you can deterministically get an outcome
or an electrical engineering where, you know, instead of your schematic being drawn out,
you can put it in cadence design systems and get it on, you know,
it made in software, right?
or CAD for mechanical engineering
where you can sort of precision engineer your part
and get it printed or manufactured.
You know, the same thing is happening in bio
and it's happening very quickly.
Yeah.
And that really opens the door for a lot of really interesting people
or maybe it wasn't as scrutable or accessible before, right?
Like software engineers like myself, researchers like Matt,
you know, obviously we're still going to want the specialists,
but, you know, the generalists can often really accelerate
the precision.
engineering happening in the domain. Yeah, I think for me, like, the base takeaway is that the field is
actually working. And like, like, not only does it have commercial traction, but like the research is,
like, actually showing signs of life. Like, it's not even just showing signs of life. Like, the signs of
life have been shown. We're actually in a place where, like, the models work. They're delivering
value. And, like, there's still tons of really interesting research problems to solve. So I think
there's a lot more low-hanging fruit in this field than there would be in other fields.
And I think the amount of impact you can have, especially like as a researcher, is just like,
unmatched in this in this field for us we're all very mission driven but even if you're not like
it's a lot of fun puzzles to solve like there there's like this kind of 3d geometry angle there's like if
you like diffusion models there's like a million problems to solve in that regard we have this
l-lm looking trunk in like chai one there's just so much of like core machine learning is touched by
these problems we're still although we've made a ton of progress there's still a lot to be done
and i think it's just like one of the most interesting fields to be working in which like while
also having some of the largest impact on just like humanity.
Cool.
Thank you so much.
Thank you for making a long journey.
Yeah.
22-minute walk.
Yeah.
And, you know, we look forward to tracking choice progress.
Awesome.
Thank you guys.
Thank you very much.
