The a16z Show - Why Medical AI Needs a Referee | Protege's Engy Ziedan

Episode Date: August 24, 2026

Daisy Wolf and Eva Steinman are joined by Engy Ziedan, co-founder and Chief Scientific Officer of Protege, to discuss why medical AI has a measurement problem, and why scoring well on a benchmark does...n't necessarily mean a model is ready for the hospital. Engy explains why healthcare AI needs independent evaluations that go beyond static exams and measure how models actually perform in real-world clinical workflows. They explore the risks of subtle bias and misalignment, why the same model can rank differently depending on how it's prompted or tested, and what happens as AI becomes more personalized and changes faster than traditional healthcare quality systems can keep up. The conversation also gets into Protege's role as an independent evaluator, how contaminated training data can undermine benchmarks, and why the future of medical AI may require continuous monitoring rather than occasional testing.   Resources: Read our insights piece: https://www.a16z.news/p/the-oracle-problem-an-invisible-bottleneck Follow Engy Ziedan on X: https://x.com/engyziedan Follow Daisy Wolf on X: https://x.com/daisydwolf Follow Eva Steinman on X: https://x.com/evajsteinman Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 Hundreds of millions of people ask chatchipi questions about their health. Who, if any, is making sure that the answers that are spit out is safe and correct? Models are going to be inhibited in their usefulness by the training data available for them. There is no one that's looking beyond the iceberg of catastrophic failures in misalignment. What is the importance of e-vals in this industry? No one ever asked, like, what is the value of Uber? Like, show me the e-val. But in today's AI market, there is a need for the pricing to be accurate. And without understanding really what is valuable and what's value-less technology,
Starting point is 00:00:39 this technology does not have a marginal cost of zero. And so it became our mission to provide safe and aligned data that would make AI useful. The right decision may also be very different a year from now versus what it looks like today. As we move into the future, a medical AI model can ace thousands of test questions and still fail at a job we actually need it to do. In this episode, Daisy Wolf and Eith Steinman sit down with protege co-founder and chief scientific officer NGZden to unpack why healthcare AI needs a much better way to measure performance. They discuss the gap between benchmarks and real-world clinical tasks, why subtle bias and misalignment may be harder to catch than catastrophic failures, and what happens when models evolve faster than the health care system can evaluate them. NG also makes the case for an independent referee, one that can continuously test how AI behave,
Starting point is 00:01:29 in real clinical settings, compare computing models, and identify where they actually need to improve. Welcome back to the A16C podcast. I'm Daisy Wolf, partner on A16C's bio and health team, joined by Eva Steinman, investor on our bio and health team as well. Today, we are talking with NG Zedin, co-founder and chief scientific officer of protege and a healthcare economist, assistant professor at Indiana University, whose work has been featured in the New York Times and cited by the CDC. Today, we are going to dig into why medical AI has a measurement problem, why acing benchmarks does not make a model ready for the hospital, and how protege is building the referee. Angie, welcome to the podcast.
Starting point is 00:02:17 Thanks for having me. Angie, let's start with your background. How did you first meet the protege team? What were you doing at the time? And how has it evolved since them? Yeah. I met Bobby when I was two years out of my PhD, an assistant professor at Tulane in my lonely office, and a pandemic had just hit. And I decided that I was going to write papers really fast.
Starting point is 00:02:37 So I wanted data basically from two days ago, and that was impossible to obtain at the time. In his past life, he used to work for a data facilitation company, and he offered AWS instances and access to data and connections. And it was a lovely experience. Then we never spoke again, maybe after 2021, 2022. And suddenly in 2024, February, I get this email sent from a Gmail from Bobby. And it has an idea, a Google document in it. And in typical Bobby style, it was human written, very simple. And it said something like, what do we hold to be true about the future?
Starting point is 00:03:15 And in it, it basically lays out this hypothesis, which has been partly true. When I look back at that document, I often read it every few months or so, which is that models are going to be inhibited in their usefulness by the training data available for them. And that goes beyond healthcare in any domain. And so it became our mission to provide safe and aligned data, essentially the teachings that would make AI useful for humans. And today, it's really a proud moment for us. Almost all the models have been pre and mid-trained on our healthcare data. We provide data in multiple verticals, audio, video, robotics. I go through like Slack channels and I see kind of like what data sets people are talking about inside
Starting point is 00:03:59 Perraget and it could be anything from 100,000 endoscopy videos inside someone's body to like pet data. And I don't mean nuclear medicine. I mean literally like veterinary care data to like 3D objects where the team is trying to predict the weight of the object being lifted. So the robot is trained in like a diverse set of like mugs. It's a fascinating world. Totally. By thinking Bobby's background and data, he had the insight very early that like, the internet was going to be scraped and that the frontier models were going to be defined.
Starting point is 00:04:28 The best ones would have access to the best real world data and a very clear vision of how to get that to people. We should frame that one pageer and put it in all of our own. Our ground source of truth. He also had, I think, like, the eye for the humans who annotate and generate data will reach the saturation of their imagination and that there is nothing like reality. Today you kind of see the market now saying, oh, it's actually a real world. that we want and like synthetic data maybe is insufficient or like what I would call like data in a contrived scenario generated in a contrived scenario is insufficient and that was also in the document. Awesome. And so you mentioned you've been working with a lot of large language models
Starting point is 00:05:10 since day one since two years ago and the world has obviously changed a ton since then. How is your relationship with your end customers changed in what you provide to them and how you work with them. Yeah, so first, as a small company, you kind of like want to get ambitious, but you don't know what ambition is, really. So we had this ambition that we would service a lot of the startups that are building an AI early on. That turned out not to be true. That immediately very large foundation models became our modal customers. And they first wanted to conquer the consumer market. So use cases, like can it have an intelligent conversation with you? Is it aligned? Is it it safe, right? Is it reasonably sound in medicine and things like that? And then slowly and slowly,
Starting point is 00:05:56 these foundation models began to see the amount of revenue available in enterprise settings, right? So it's like, can it task aggregate in knowledge work? Can it automate? Can it entirely replace certain workflows? How do we think about tasks? Jobs as a collection of tasks, right, are like a contrived human concept. And so the data requests we get today would probably sit between consumer needs and enterprise needs. And then, of course, there is the need for once I build this model, I would like people to know it's the best. And I'm not an arm's length removed from that conversation. And so we started to naturally gravitate towards benchmarks and evaluations in real world settings and thinking about just what we could do today. But really, what is the right way
Starting point is 00:06:41 to do it going forward? Like, what holds to be true? Totally. And like, talk to us for a second about why e-vals are important, right? Because you can obviously argue, like, hundreds of millions, billions of people use AI every day. They're not asking for an eval necessarily. Like, what is the importance of e-vows in this industry? Yeah, so I think a lot about things from the perspective of economic theory, just given my background. And if you think about it, economists back out kind of value from things using an approach called hedonics, which is like someone's willingness to pay for the thing would give you the value of the thing. So the best example is school districts determine housing prices, right? Because people's willingness to pay for better housing.
Starting point is 00:07:27 No one ever asked, what is the value of Uber? Like, show me the e-vail. People just paid and gone on the ride, and over time it became like a verb. You Uber it. But in today's AI market, there is a need for the pricing to be accurate, for equilibrium, for clearing. And without understanding really what is valuable and what's value less technology, this technology does not have a marginal cost of zero, right? People are paying in tokens. They're paying with like onboarding time. They're also paying with the probability that something would go wrong. They're taking that in. And so there is a need for evaluations that are robust, for it to really become integrated in knowledge work, high risk scenarios, everyday businesses, right? Just like you would hire an employee and you would vet that
Starting point is 00:08:09 they're adequate for the job. Before you hire AI, you need to vet that it's adequate for the job. Totally. There's so many trust issues in AI. And if you can prove why something is trustworthy, obviously the willingness to pay for that thing should be higher. I think also, like, in healthcare specifically, healthcare always had this asymmetry in the transaction. Sometimes people illustrate it with like a used car market where the seller of the car knows so much more than the buyer of the car. The sellers of care, the providers, know a lot more about what is valuable and value less technology and your current state of health. then you may know, right? There is evidence of this. I often talk to the team about the effect of having a doctor in the family. There's research from Sweden that shows like just a random assignment and who goes to med school based on your grade in the med school entry exam, randomly assigns a doctor
Starting point is 00:09:00 in the family and people live X years longer, the entire family. And so this asymmetry means that the providers of the care know a lot more than the consumers of the care. When you put agents on top of that, right? It's almost like saying we have a million doctors in America. We're going to hire a million more. Those additional million doctors have taken a lovely 5,000 question exam and passed it. We'll let them go ham and see how it goes. And we're comparing that to the million existing doctors in the U.S. who have gone through nearly a decade of specialized training many more exam questions than 5,000, which is really wild to think about. To your point, health care is a really interesting industry because you have one and asymmetry problem across all the key stakeholders,
Starting point is 00:09:44 patients, providers, providers, you name it. And you also have incredibly high stakes, right, like the self-driving car analogy where there's absolutely no room for a model to make an error. How do you guys think about that at Protagai? Yeah, so I think like when people talk about kind of like safety and alignment, right, there is the catastrophic failure, which is something goes wrong. Can you prevent that? Actually, I do think that that problem is easier because catastrophic failure is like narrow. Like you could just define it as like mortality, right? The harder one is misalignment broadly and subtle bias.
Starting point is 00:10:20 First, it's very hard to define. It's even harder sometimes to detect. So let's give an example. Today, there are a lot of AI builders who are building in AI for insurance. And hospitals are also building an AI for insurance. prior off or acquiring these technologies for AI for prior off, medical denials and so on. You can think of it as like an edgeworth box where there are two indifference curves that are coming together on a contract curve. One is trying to maximize total revenue and basically
Starting point is 00:10:54 minimize payouts. And one is trying to maximize revenue recovered, right? They are both using agentic tools. And the patient is in the middle. The patient has no agency. And there is no regulator that basically says, we think this model is really, really performing well, right? But it's actually misaligned. So examples I give my colleagues at data lab is like, is it fair to say that if you were to be prescribed a medication, it's okay for the model to first check if you owe the hospital any money? Is it fair to say that if the hospital said don't refer people out of network, right, that that would be okay, even though your preference might be, I'd like to go out of network sometimes, even though it's expensive, right? But there are certain benefits also to the tools. It's not all negative.
Starting point is 00:11:44 So, like, think about ambience technology broadly, right? I read medical notes. We have hundreds of billions of medical notes within prologue. And some medical notes will say things like, she looks really disheveled. Her husband is asking good questions. What they're actually saying in the note is that she's not trustworthy. And suddenly you see a diagnosis. And suddenly you see a diagnosis that's like, oh, that pain she has is potentially mental illness, right? We're not going to clinically look into it. Now, a tool that records a conversation between a patient and a physician kind of, like cuts that out completely. What was said is now documented into a soap note, right? None of that is in there. And now we have a very accurate clinical description of the
Starting point is 00:12:28 conversation. A third-party physician that's looking at the conversation would be like, oh, I should probably order a clinical exam to investigate the cause of this pain. I don't think it's just in her head. And so these benefits are not even being evaluated. Yeah. It's a really interesting point that, like, there's the potential to cut out a lot of human bias with AI, even though there's also the potential to reinforce it by training on bias data. Let's come back to that in a second.
Starting point is 00:12:55 Maybe backing up a bit, can you tell us, like, you know, today hundreds of millions of people, ask, you know, chat-chip-key questions about their health. Doctors are using tools like open evidence to ask medical questions to AI in the course of their practice every day. What, like, who, if any, is making sure that the answers, you know, that are spit out in, say, a doctor's kind of clinical AI tool that they're using is, you know, safe and correct? So I think the answer is everyone and no one. So there's a huge effort within the foundation models
Starting point is 00:13:34 to study the safety and alignment of these models before they're deployed. And specifically within healthcare, there's also research that shows that when they train on healthcare data, the model gets broadly aligned in other things, like it becomes more trustworthy overall, right? And so there is a focused effort. But this is just kind of like out of the box,
Starting point is 00:13:53 it's general personality. You might think of it that way, right? And then the vertical AI builders take that. and they tweak it and tune it to a specific cause, right? And an objective function, an SLA of a business too. There, every vendor of an AI tool, every vertical AI builder would have like a little brochure that says, I am best.
Starting point is 00:14:15 And so the no one is that if everyone says they are best, no one knows who's best. Also, there is no incentive to kind of pinpoint to where your product fails. And so really on behalf of like health care systems, or patients, there is no one that is doing this independently, that's on arm's length removed, that's looking beyond the iceberg of like catastrophic failures and misalignment.
Starting point is 00:14:40 Yeah, like recently there was a paper published in nature saying that general AI was beating vertical-specific AI in healthcare. Like, you know, the general, like the open AI models were performing better than any healthcare specific model. And then there was a competing paper on archive that came to the exact same conclusion in a very short time span. Like, speak more to this problem and, like, what is the solution here to actually, because these are real life and death decisions we're talking about in health care.
Starting point is 00:15:10 It's not just, you know, something that's somewhat trivial. Like, what is the solution to figuring out what the best models actually are? So I'll probably answer this by, like, answering it directly and then going back in history. I think, like, history tends to rhyme and repeat itself. and the issue that we're facing today isn't unique. Like, the issue of evils and of medical technology is basically a constant ever since medical technology was invented. So the two papers came to different conclusions,
Starting point is 00:15:38 whether out-of-the-box frontier models are best or vertical AI builders in what we'd call like clinical note summarization and kind of evidence, like treatment recommendations is best. And in classic academia nature, you know, two papers, like sit on opposite ends and, like, actually editors love that. They're like, oh, you know, we'll publish the thing that is most controversial.
Starting point is 00:15:59 But when we, at least a protege, kind of for our own customers, do these benchmarks and evaluations, they're highly sensitive. They're sensitive to, like, how you prompted the model. They're sensitive, like, especially like an oncopathology of, like, what harness you use. Sometimes even models will switch ranking based on the order of the multiple choice question answers, right? And some people think that that sometimes is evidence that the benchmark itself is
Starting point is 00:16:26 contaminated, that the model has just memorized the ranking of the right answer, just like a student would. And so there's a problem and a problem of asymmetry of information. And if you are attempting to spend millions of dollars on these tools to onboard them, you really want to know kind of which is best and which is best for what, right? That, I would say, is level easy. That's just like the basic answer. As we move into the future where these models, models face data drift, or even every physician has his own model, like you make someone's context now weights, like these really personalized models. Or there is like test time training where the model is constantly evolving. These static question answers or like something in a paper
Starting point is 00:17:16 where it was evaluated on a 2024 technology and published in 2026 just won't do, right? And so you risk, not to fearmonger, but you risk a situation where we enter like the opioid epidemic, when it peaked in 2010. That's only when people started saying, okay, national task force, we're calling it an epidemic, we're acting on it, let's move. We don't want to get there.
Starting point is 00:17:37 To answer maybe in historical context, the government does a program since maybe 2007 called value-based purchasing. To people who don't know what that is, that just means that 80% of what the government pays in care to physicians and hospitals is tight to quality.
Starting point is 00:17:55 There's a whole national. effort and government arms that basically define what's quality and track it. It's really static. So it would be examples like in medical expenditures for physicians, they'll ask like if you're a doctor, how many times did you prescribe antibiotic for something that looked like a virus? That's bad. That's a ding. That's a penalty. If you're a healthcare skilled nursing facility or nursing home, they'll say like how many patients fell and broke something, elderly people on your watch. And these reports tend to be retrospective, like a year later or six months later. We can't have that way they eye.
Starting point is 00:18:34 The technology is just moving too quickly. It's changing rapidly. And it has a lot of agency. Maybe to that, and like, how do you think about then what the right balances between ongoing monitoring versus quarterly annual reporting, especially when it comes to health care where the technology is moving so fast. Decisions are made, you know, at the point of care, but also, you know, literature may be evolving. and the right decision may also be very different a year from now versus what it looks like today.
Starting point is 00:19:00 I mean, if we had courage, right? And we had like no fear because health care is filled with administrative work. Like when you say anything about like, oh, new technology and health care, everyone kind of takes a back seat and they're like, oh, you don't know what it's like. This is 20% of the GDP, 20 million of the 150 million jobs. Good luck. But if you were to really think about it as like, okay, forget all of that. What is the right thing to do?
Starting point is 00:19:24 The right thing to do is for there to be a watcher all the time, right? That when my local AI agent as a physician starts to nudge me towards decisions that are maybe profit maximizing but hurting my patient, I'm alerted, right? When a group of nurses realize that the AI is maybe recommending misaligned things because of a conversation they had about the patient that led the model to be overly eager, Right? So I said something negative about the patient and now the model is overly eager, but is actually like not recommending something safe. It's trying to benefit me as a nurse or my working hours, right? That there's a watcher that says that behavior is unacceptable. You can't discharge this person early because you have to leave early, right? Like I don't think that this is an unbiased decision. Totally. Like you've said before that models, you know, can score up to like 92% on a licensing exam, but 45% on real world clinical tasks. And so like acing a benchmark doesn't make an AI ready for the hospital, just like a perfect MCAT score, doesn't make a, you know, a doctor doesn't make for a great surgeon. Like how are these benchmarks breaking?
Starting point is 00:20:39 What's going on? So first there was like maybe tears to it. The first one was like I think people admit to this and it wasn't perfect is that they would take these like med school style questions and answers, right? and they were doing remarkably well. They were beating the average med student, right, that maybe at a level of a resident and so on. And that was an indication that the models are improving. It's hard to tell if that is true reasoning.
Starting point is 00:21:07 It's hard to also tell if that means alignment is going up or down. It has no indicator for that, right? And then people started putting great effort into these kind of questions that the model would answer, and then they would compare them to physician responses. They would use multiple physicians to grade the response. They would use thousands of questions. Then they upgraded it to be like, you know,
Starting point is 00:21:32 some models say they've had their physicians review more than 700,000 question-answer pairs by human physicians, which is a great effort. But if you are a physician and maybe I ask you, like, as a patient, right, you're about to get spinal surgery. They tell you this model, has passed like a 5,000 question, 10,000 question Q&A exam, would you trust it versus would you trust like a physician
Starting point is 00:22:01 that has done 4,000 of these surgeries, right? What you care about actually isn't the general knowledge of the model. Like, that's great. I love that it's intelligent. I care. Has it been in this scenario before? And has it been helpful? So like, has it been in a spinal surgery as a task aggregator before?
Starting point is 00:22:18 How many people have died unwarranted? like above average, how many people have had improved outcomes. That's what I care about. And so the benchmarks and evils aren't specific enough to these high-risk scenarios. They also face, I think, what is often kind of rushed under the carpet. Like any occupation in the world, physicians have preferences, right? Like investors have preferences, right? Some of them are like risk-loving.
Starting point is 00:22:45 Some of them are not as risk-loving. They have a beat. And physicians have a beat. So, like, we show a lot of these kind of graphs inside Proreje, where I showed on knee replacements, for example, I said there are physicians who never do a full knee replacement. He's only part of it. And then physicians who, if he touches a knee,
Starting point is 00:23:00 it's a full knee replacement, never does partial. And that's his beat. You take this medical record, this real world case, and you put it in front of a model, and you say, what would you recommend? And it would recommend the opposite of what the physician did. And you say, oh, model's wrong. But actually, it could be that the physician is wrong, right?
Starting point is 00:23:18 He has a sticky preference. a hysteresis, almost. And so we are also capped in our ability to evaluate them in real world setting, but we're also capped now in our imagination. It's like, what is going to happen when the model becomes better than the average physician?
Starting point is 00:23:32 The answer simply can't be like, oh, so then you ask more physicians and you get the law of large numbers to, like, help you out. At some point, a task will be so expert, right, so high risk, that testing a point of care will become really warranted.
Starting point is 00:23:47 How do you do that? So we think that there is two ways to do it. The first is that there are multiple vertical AI builders in what I would call like an identical subnode. So subnodes could be things like ambience inscribing, clinical trial recommendations, nurse workflows, scheduling. There are subnodes like oncopathology. So can the model look at an H&E slide
Starting point is 00:24:15 and recommend like the IOT therapy that is best, paired, right? Within these sub nodes, our customers say, us compared to others in the benchmark space is like ships in the night. I don't know if he's better than me. He says he's best. I'm best. And so what we have been pulled into is basically these kind of models say, can you host my model and have it take a unified evaluation, right, and produce kind of a report that evaluates us against each other? Could we, we would also like to learn where we are better than each other and where we can improve. And so we've been doing that. And it's been really interesting kind of seeing a role as an arbiter. Things like, where do I draw the line of like, no, that's our decision actually.
Starting point is 00:25:04 Like what should the prompt be, right? Should the frontier model have like a bespoke harness built or should we use basically the built? right? Because like some world models don't need a harness. What are you going to do when the model has a refusal? Are you going to put in the mean value, right? Or are you going to defer to like a second best model? Do these change the rankings of who's being evaluated? And so we create these kind of like evaluations within the subnodes. And our purpose is to share them with the consumers of these technologies. So in the case of oncopathology, it would be the pharmaceutical companies that would ultimately purchased this or the hospitals that would ultimately purchase this. So they know where
Starting point is 00:25:46 they stand. It's so obvious why the creators of the models shouldn't be the ones releasing the evals and why, you know, people shouldn't be allowed to grade their own homework. Like that is, you know, an easy incentives issue to spot. But explain to me why, like a skeptic might say, okay, like, the labs are getting their data from protege now. Why should protege be the, like, arbitragee be the, like, arbiter of evals and truth. Like, can you explain to me why protege is well positioned to do this? Yeah. So, I mean, there is like the idea of comparative and absolute advantage. I tend to explain it to my team as this way, which is like absolute advantage is like, we're best at this, right? We're best at this healthcare data game. We're probably best in the market at it. Comparative advantage
Starting point is 00:26:38 is like in everything we do, we have lowest marginal cost. when we also facilitate healthcare training data and e-vails. Like between everything we do, we're also best at it among our own production function. And so we became really geared for it. That doesn't mean you have the right to win. We have the right to win here because it kind of naturally evolved is that once we do these e-vals in the subnode, right? We say once the evaluation is done and we hold this test set that never leaves, if you have a deficiency in an attribute that's now been revealed, we can tell you exactly,
Starting point is 00:27:12 what data would improve the model's performance on that deficiency. And because of the size of the data set we have and the various sources that we have, it's a very fast way for you to evaluate and then reinforce, right, where the model is failing with, like, new teachings. So it became that we are geared really up for this. There's also an incentive to completely remain impartial, right? because like any network effect idea, lack of trust destroys the entire network almost immediately.
Starting point is 00:27:45 It's like explosive. And so you are really like held to what are you doing and how transparent are you being. So for example, all the methodology is shared broadly with everyone in the sub-node. We don't reveal the names of the people, the vertical AI builders in the sub-node until the very end. They have the right to silent X.
Starting point is 00:28:08 it and things like that. We think that we are better positioned than the government, because the government's not going to do it. Yeah, a lot of people would just say, like, why isn't the government doing this? Because the government's not going to do it. And like, basically what's more unsafe, right, is like sitting there and saying, well, I'm going to wait for the government to do it. No, we think what's more safe is just do it, right?
Starting point is 00:28:29 And then let people challenge you on it, but just do it. You know, the government can't even get you to file taxes online. Like, people had to invent, like, a marketplace online for that. to file for the government. And they just asked industry for help with this. With AI, yeah. So there's like a federal registered document where the government is asking industry collaborators
Starting point is 00:28:48 to help them identify how should they evaluate AI. But generally, like pay for performance, value-based purchasing programs, like a zillion years ago, I wrote my dissertation on this policy which penalized hospitals for excess readmissions, a quality metric, right? You know, thousands of papers.
Starting point is 00:29:06 It took like 10 years to evaluate. That's one question. Is it causing more readmissions? In AI evaluations, you can actually ask a thousand questions about how an agentic system is behaving inside the hospital. It is also, I think, missed always that we are not replacing doctors and nurses, at least not today, with AI in hospitals. We are kind of like allowing the AI to ride shotgun, right?
Starting point is 00:29:33 And humans behave towards technology in a my radicalization. ways. That is a subject of study. It's called like human computer interaction. They teach that in universities, right? Humans can dig their heels in. They can become overly optimistic. And these kind of benchmarks and evils that are retrospective fail to acknowledge or measure how in a live wire scenario, how is the AI behaving? And so when you are tapped into that data live, as prologer is, you're really prime positioned to be answering these questions. Yeah, you also made the point earlier about AI just memorizing answers. And when creating benchmarks, you can actually create them on data that has not been sold to all of the models out there. And so you can test whether the model has just, you know, memorized the answer or has the ability to do kind of clinical reasoning, correct? Yeah, I mean, the contamination issue, like we've had various debates about it within proje and with customers. I think.
Starting point is 00:30:38 First, we had, like, I didn't think that was even possible. I was like, okay, so we have like 200 million people, you know, hundreds of billions of notes. Like, I just need data that hasn't been ever given to anyone for training. And they were like 80% of what we have has been given for training. And I was like, okay, they were like, if your definition of basically independent is a patient that the model hasn't ever seen before, you may have a challenge in front of you. And so for like the oncopathology example, we literally had to pull patients
Starting point is 00:31:15 whose whole slide images had never been scanned before. We contacted hospitals and were like, in your pathology drawer, who hasn't been scanned before? And they're like, net new slides. I was like, net new slides. Never before has this patient entered a model. And that's very hard, but we know which patients have entered which models. We've given this data for pre-endmit training.
Starting point is 00:31:36 Of course, like sometimes, you know, You know, you have these, like, philosophical debates with researchers, and some of they say, oh, the model is, like, doing worse than random. It's doing really bad on this task. Who cares if there's a little bit of contamination? Like, sure, but it's not impacting the judgment. And I'm like, yeah, but like, worse than random could be a detection that there is serious misalignment to. Like, why is it doing worse than random? We should just care that this patient has not entered training previously.
Starting point is 00:32:08 And so it's a challenge for us. We're pulling net new data all the time. And now we have what I call like sealing a membrane, essentially, this moratorium on any data set or any patient that has entered training. They are not allowed to be in a benchmark or an e-val. To fascinating intersection of the market to sit up. Yeah, I mean, I think maybe the way to frame it is that the U.S. imports a lot of physician's skill from other countries.
Starting point is 00:32:35 Definitely in like long-term care. I think that market is dominated actually by imported nurses from developing countries like that scale. And they go through rigorous teachings and basically study, examination. You can't even have a doctor in the United States that's certified across states. Like they have to retake, basically, exams to become recertified again and so on. And then we have these AI tools that are unleashed in healthcare systems or in kind of practices with absolutely no credentials, other than the fact that the maker says compared to some physicians we have, the AI beat it, right?
Starting point is 00:33:14 And that sits in a very perplexing way. People care about their education. They care about their money. They care about their health. These are like core human capital axioms. And so any misalignment in these human capital axioms will read to a lot of like mistrust of AI generally. And we don't want that to happen. Totally.
Starting point is 00:33:34 I think to your point before, we like, you know, all of this is just kind of being allowed to be self-governed right now and self-regulated because we're like, oh, well, it's just like, you know, the doctor is just consulting it, but we'll catch any errors. And every time we give medical advice to a patient on chat GPT, we're saying, you know, this is not medical advice. But to your point, like, people trust what comes out of computers. They don't necessarily question it. And so it's really important. that someone out there is, you know, grading these models, and it's not just models themselves. Thank you so much, Angie, for coming on the show and for all of the important work you're doing to make clinical AI safe. Thanks for having me. Thanks for listening to this episode of the A16D podcast.
Starting point is 00:34:24 If you like this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcast, and Spotify. Follow us on X. A16Z and subscribe to our substack at A16Z.substack.com. Thanks again for listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that
Starting point is 00:35:00 A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a16Z.com forward slash disclosures.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.