The a16z Show - Beyond the God Model | Alex Atallah & Amjad Masad

Episode Date: October 3, 2026

A16z’s Erik Torenberg sits down with OpenRouter’s Alex Atallah and Replit founder and CEO Amjad Masad to discuss why the future of AI may look less like one all-purpose model and more like an ecos...ystem of specialized models working together.Alex explains why OpenRouter is betting on “neurodiversity”: different models trained in different ways, routed and combined based on the job at hand. Amjad makes a similar case from inside the enterprise, where companies increasingly need to own their AI capabilities rather than depend entirely on a single model provider. They explore what happens when general-purpose agents give way to teams of specialized agents, why smaller models can sometimes be cheaper, safer, and easier to control, and how routing and model fusion could deliver frontier-level performance at lower cost. They also get into agent-to-agent communication, AI security, and why the next generation of companies may need an independence layer across models, clouds, and data.Resources:Follow Alex Atallah on X: https://x.com/alexatallahFollow Amjad Masad on X: https://x.com/amasadLearn more about OpenRouter: https://openrouter.aiLearn more about Replit: https://replit.com Stay Updated:Find a16z on YouTube: YouTubeFind a16z on XFind a16z on LinkedInListen to the a16z Show on SpotifyListen to the a16z Show on Apple PodcastsFollow our host: https://twitter.com/eriktorenberg Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures. Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

Transcript
Discussion (0)
Starting point is 00:00:00 You saw the SpaceX S1. It's like, oh, $30 trillion. It's like, what is the world GDP? 100 trillion? Both Stripe and OpenRouter really want lots of new companies in the world. We don't want everyone to be a part of one giant company. The reason why the open AI hacks have been so destructive is because they're so capable.
Starting point is 00:00:22 It's like nuking a butterfly. When the models get more intelligent, the risk actually will like continue to get higher. continue to get higher, and yet no one knew is taking responsibility. We're going to slowly realize how good we've had it with, like, deterministic code. Remember the days when computers did exactly what we told them to do? Are we going to prevent the models from deceiving users during training runs, predictably? Well, like, a model that's big enough and powerful enough, suddenly stop deception,
Starting point is 00:00:54 stop sandbagging. For the past few years, AI has been racing toward bigger, more general models. But what if the future is actually more specialized? In this episode, I sit down with open routers Alex Atala and Repplets Amjad Massad to discuss why the next phase of AI may be built around many different models working together. Alex makes the case for what he calls neurodiversity,
Starting point is 00:01:19 combining models with different strengths, rather than locking into a single provider. Amjad explains why enterprises may increasingly want that independence too, owning more of their intelligence and choosing the best model for each job. We also get into specialized agents, model routing and fusion, the tradeoffs between capability, cost, and safety, and what happens when AI systems start training smaller models to handle specific tasks? Welcome to AZZ podcast. We're here with Amjad of Replit and Alex of OpenRouter,
Starting point is 00:01:52 and this is the first podcast that Alex has done since the acquisition, so we're really excited to have both of you. Thank you. I'm excited to be here. Alex, let's start with that, actually, if you can briefly share, obviously, massive acquisition. I'm Jodd as an investor. We're also, of course, an investor, the biggest shareholder, but who's counting? Alex, why don't you give us a little bit of the backstory? Like, how does the acquisition like that even happen?
Starting point is 00:02:15 Do you get a DM from Patrick one day? What can you share? Hey, how much for it? I had talked to Will Gaborick, the straight president, a lot. time ago, a couple years ago, when we were doing our series A. And we just, we stayed in touch. We had, like, a lot of Stripe work streams going on with various teams at Stripe. So there were always sort of, like, things that we were doing with Stripe.
Starting point is 00:02:45 We presented at Stripe sessions. And so they always kind of felt close. And then, yeah, just in July, I believe, they reached out, wanted to chat, met both of them in person. and it kind of just progressed from there fairly quickly. They're very efficient and they were very founder-friendly about the experience. I was really impressed with the whole thing. Did you want to sell? Did it even cross your mind before they were shot?
Starting point is 00:03:16 No. We were not thinking about that at all. I did really respect the company, do really respect it today, and of possible acquisition options for us, it was my top choice. And so it was like an interesting idea. And as we kind of fleshed out the reasons why it would make sense for both companies,
Starting point is 00:03:44 got more and more interesting. It was really clear how aligned they were with us kind of having autonomy over the brand and roadmap and product and keeping OpenRouter doing what it's already doing just much faster with a much more serious go-to-market plan and then some better together stories between the two products and the two companies. And then culturally, there were just, I think,
Starting point is 00:04:11 in terms of mission and values and building a neutral, trusted platform that businesses can depend on and scale on top of, that's also really developer-friendly with the best possible developer experience to encourage new companies to emerge. That alignment was there. And there was a bigger picture kind of alignment too
Starting point is 00:04:32 where both Stripe and OpenRouter really want lots of new companies in the world. We don't want everyone to be a part of one giant company. We want to create really good incentives and really sort of easy, streamlined workflows for people to start new companies and grow them successfully and make both lifestyle and venture-backed businesses
Starting point is 00:04:56 on top of really good, reliable, and price-efficient infrastructure and like a marketplace that works. And I really want that future. And Stripe demonstrated that they've been wanting it and building towards it for many, many, many years. And in many ways, like payments and inference are going to blend together for companies of the future.
Starting point is 00:05:20 I'm curious, I understand why Stryryry. Ripe's incentive is to have much more vibrant startup ecosystem. I understand the kind of moral argument and why you would want that. But why is that good for OpenRouter? Is your model for Open Router? Is that it a network effect business? Is it like a network? Well, for us, I think there's a couple different problems open router solves.
Starting point is 00:05:47 One is allowing you to build a company. that uses AI or augments intelligence with unique data and other services without model lock-in, without vendor lock-in, allowing you to kind of be on the Pareto Frontier continuously as the ecosystem grows it. And to do that, it's a lot of work because there's all kinds of little lock-in that appears. We also, we want companies to feel like that they can like add more than just prompts on top of a single model. There's a lot more to building like unique intelligence. And I think a big component of that is neurodiversity. You really need the power of multiple models that are trained in different ways, including some of your own, to do more.
Starting point is 00:06:47 than chat GPT or Claude would on the task if like someone who's thinking about buying from you as a potential customer is wondering, well, what if I just use the model directly? How do you really show that you are significantly better and able to build a business that matters? And I think a lot of it will involve neurodiversity and blending like powers and good data from multiple models.
Starting point is 00:07:15 Another component is like, helping people get really good cost efficiency. Like, there are a lot of business that just cannot, do not emerge until they become cost effective. And, like, creating an environment where we can help drive down costs by building an efficient market is, like, crucial to making that happen. Otherwise, why lower my prices as a provider? We have a captive market. And I think that's, like, a key point of marketplaces. that was just totally missing from AI before we showed up.
Starting point is 00:07:50 There was just one player, Open AI. It could have been like a very strange world. I'm not saying that we did all the work. Of course not. But helping people like choose new models and explore new models and learn like what makes a closed source or open weight model like actually good at your task involves like seeing what the whole ecosystem is doing and learning automatically.
Starting point is 00:08:18 Like, LLMs are not things where you can just enumerate all the features on a web page. It's impossible. You have to see how they're being used to know what they're good at. You started the company three years ago. I'm curious what has surprised you the most about sort of the evolution of open and close source models to the present as it relates to model performance or just people's perspective on the variety of models they have available to them. People have been more open-minded than I,
Starting point is 00:08:46 thought they would be towards open weight models. Typically, there's a lot of brand, especially in enterprise. Typically, there's a lot of like brand trust where like enterprises in general are like, oh, I don't really know how to tell the difference between these things. So I'm just going to buy the one that all the other credible enterprises are buying. And there's a lot of like enterprise lock in with a mentality like that. And that just didn't happen that much. It did happen, like, a bit, but we saw a lot of enterprises want to explore new models.
Starting point is 00:09:26 It's simply like it was very much good for a marketplace. Enterprises wanted to diversify outside of just the proprietary frontier model labs, both for like cost reasons and for differentiation reasons. They wanted to own their intelligence, so they could, I think, one, keep their talent, have, like, an internal AI practice. Like, AI is just, like, a huge strategy topic. It's not like you go to your board
Starting point is 00:09:56 and you're like, oh, yeah, we fix the AI problem. Quarter complete. Like, your board is, like, asking you every month, like, what's next for the internal AI team? And, like, every single enterprise now has just this internal AI team that they're developing where they need a strategy behind it.
Starting point is 00:10:15 it's not just a, like, you know, we check the feature off. We, like, set up the database and we're done. And so I think that that dynamic has resulted in just, like, a desire to explore and diversify. A desire to kind of, like, figure out how to reduce costs and figure out how to do benchmarks for the first time in the company's life. I have been, like, surprised there hasn't been more benchmarking, like more companies creating more benchmarks. It's starting to happen. And I think eventually we'll see a lot more of them to demonstrate, like, oh, yeah, like, this thing is better than using, you know, Claude Direct. But I think that's going to be, like, a bigger focus for this internal AI group at every company, eVals.
Starting point is 00:11:02 And I think, like, Amjad's been doing that a lot of replet, for example. Like, you guys have done a lot of cost per task research. You guys have been, you know, you made, like, a doom loop. rescue. You kind of been like experimenting with new ways of using agents like Doom Loop Rescue and and like helping bring those to developers. So like more of more of that kind of research I think is going to pop up internally everywhere for all of those reasons.
Starting point is 00:11:34 Yeah, I think Sotja, CEO of Microsoft has been very, very, you know, pressient. on this and also very articulate on why companies need to own their intelligence. Ultimately, in the same way that we had dot-com companies and then every company became an internet company. Like every company employs people that know how to build websites and be on the internet.
Starting point is 00:12:04 Similarly with the software, every company has software engineers. Every company needs some AI practice, AI capability, and that will compound over time. the knowledge, the intelligence inside the company, the use case model fits, like which models actually work for them,
Starting point is 00:12:23 how do they save money? They need sort of that independence. The other thing that I think Alex Carp of Poundary has been talking about is that there's like a risk that when you work closely with the foundation model companies is that they're going to move into your business. And we've seen that, you know,
Starting point is 00:12:40 with Figma vis-a-vis Anthropic. We've seen that with like now Harvey an open AI and you know it's it's really hard to partner with them and and because you know they're they see the world as their potential market right like I mean when you when they talk to investors that are like you know you saw the SpaceX S1 it's like oh 30 trillion dollars it's like what is the world GDP 100 trillion and so there is a sense in which these companies are different than other generation of companies it's it's harder to partner with them because their ambition is such that they want to be they want to subsume a big part of the
Starting point is 00:13:21 economy and increasingly like what we're thinking about at reflet is kind of in the similar vein to what Alex have sort of innovated is uh replet is becoming more of an independence layer inside inside enterprises where we create a layer of indirection between um between you and the models and we get you the best token at the cheapest price. But also we also create an abstraction layer on top of the cloud as well. Because you should be able to deploy to AWS and Azure and you should be able to use Databricks and Snowflake and so on. And so increasingly, I think there is not just with AI,
Starting point is 00:14:05 but all of technology, there needs to be more platforms that help companies gain independence. It seems like, yeah, what OpenRouter did for their segment, you're doing for other areas of this. Yeah. There was this kind of reminds me of this tweet I saw the other day when somebody was like, basically all companies are building the same thing now. Everybody's building an agent loop with like notifications and context, like third party connectors and context, like, third party connectors and context, management and memory and sandboxes
Starting point is 00:14:46 sandboxes you know web you know agentic web search and an always on agent on top of it notifications and it's like this product is showing up
Starting point is 00:14:58 everywhere in a way yeah it is like showing up everywhere but it also it kind of feels to me like these are just the new table stakes primitives. It's kind of like a 2005 version of that tweet would be, oh, everybody's building the same thing.
Starting point is 00:15:18 It's like a database, a user's table, a sign in page, a sign up page, a profile page, a logout page. Like, everything's the same. There's like a lot of differentiation, really. There's like there's table stakes needs for AI just like there are table stakes needs for the web. Yeah, and I think as well, like inside the enterprise, making these products actually do real work is still unsolved problem. Like you can use Mews in your personal life and connected to your credit card and bank accounts. And no one's connecting Mews to their enterprise data or even Grockbox. and things like that.
Starting point is 00:16:10 I think there's an even more emphasis on data sovereignty and security. So we spent the past, you know, year almost working on making replets deployable on your own cloud, basically on-prem, like bring your own cloud. Like two years ago, I would have thought I would never do this because, you know, it's just like, yeah, the cloud is the future, like software as servers. all of that. But now, actually, we've sort of like reverted a little bit back to a world where companies are a little bit more protective because there's so many ways in which data can leak.
Starting point is 00:16:52 Like all these agents that people are using. I mean, there's all these screenshots on Twitter. I don't know how true where instinct is like, or muse is like mixing people's data, starts to call you by a different name or something like that. And so, yes, the kind of consumer stuff is kind of obvious. But on the enterprise, there's still a tremendous amount of work for the entire industry to do in order to actually get these things to be useful and productive at work. Do you, I'm John, are you doing any, like, do you have, like, a custom personal agent
Starting point is 00:17:24 other than Muse or Instinct that you use for, like, work stuff that you've been building? Yeah, I mean, I built something on Reput, like a long time ago. I started as, like, a sort of a CRM agent initially. that was the main problem that I had. But slowly, like, we added features to it, and it's sort of, like, doing more and more things. But what's really interesting is that the more I connected Rapplet to all my stuff,
Starting point is 00:17:51 it sort of, like, started answering all the things for me. And so increasingly, like, the platform itself is, like, subsuming these sort of, like, these domain-specific agents that I built. I do think that there is some, in some ways, you want something that is sometimes like focused on one particular thing and you don't want it to be able to do everything. On the other hand, once you have your entire company's context in one place, it's really cool to join across totally different domains. Like when I ask you the question, it can like look at my sort of personal chat history, joining across the GitHub repo across Salesforce.
Starting point is 00:18:31 And so it says like it will link like random things. It's like, oh, you met this guy like a year. ago at a conference, I see it on your calendar. And by the way, someone else from their team is in discussion with your sales team. And it creates all these different synergies. And when I go into a meeting, I, like, you know, I, I've connected a lot of different threads and I'm making much more progress on a deal or something like that. So, so this is where it's sort of trending now. Well, I think I'm like a little bit, I'll take the counter on that. I think that the worst part about doing cross-domain joins with your personal agent is that the more work you give it to do, the more understanding of what's going on, you're sacrificing.
Starting point is 00:19:26 and yet no one knew is taking responsibility for that sacrifice, understand it. Like, agents don't have any responsibility. You can't, like, if there's a fixed level of cortisol that the whole company can tolerate between everybody, and, you know, you like, you, like, I want to be, like, less stressed about some area if I'm going to be, like, sacrificing my understanding of it,
Starting point is 00:19:58 someone else needs to take the cortisol. The agent doesn't take on any of that responsibility. And then like a universal agent that's doing all things, you know, I can't like adjust how much understanding I'm sacrificing in all the different things. It like points me a little bit towards, you know, maybe like down the road, like the sub-agents that people use will be like,
Starting point is 00:20:25 very vertically focused. Maybe we have like a chief of staff type agent that, you know, coordinates between them. But I feel like you do need like vertically focused agents where you're like, okay, like this agent is more responsible psychologically for these things. And like I want like quality checks that make sure it's doing those things correctly. And I don't, it doesn't need to focus on anything else. It just has one focus area.
Starting point is 00:20:55 I wonder if that's going to help people at least get like a weird, loose sense of responsibility on top of agents. Fascinating. So you're saying general agents create like a tragedy of commons of sorts. Kind of. Like I have a general agent that every day looks for things that need me and tries to figure out what to do. And it's just, like, impossible to improve this feature. Like, it's, like, every time I try to make an improvement, I end up, like, ignoring its output about a week later. It just feels like it doesn't really care about any of the, like, specific things that's diving into.
Starting point is 00:21:51 And, like, like, kind of imagine having a chief of staff where they're very good at, you know, like drafting all of your replies across the whole organization. And then compare that to something where you have like 10 chiefs of staff, each as competent as that one chief of staff. But they're all responsible for like individual sectors of what makes up your life. Like the latter, I feel like gives you a way of tuning, how much understanding how much understanding, you sacrifice compared to the gain you get from, like, basically I can like lean in more to the areas where the agent is failing for some areas and then have agents with very good competency, like take over my understanding of other parts of my life.
Starting point is 00:22:45 Yeah, interesting. It's sort of like almost rediscovering, you know, specialization. Right. what's his name the like the famous economist Adam like him Adam Smith Adam Smith like with a pencil
Starting point is 00:23:02 kind of thing where yeah that was like a huge realization for humanity that like specialization is actually good the problem is like we kind of like
Starting point is 00:23:15 over specialized as like a civilization and I think over specialization is is is oppressive in its own ways. I mean, I think we've, you know, sort of like, you know, there's the Marxist theory of alienation, right?
Starting point is 00:23:34 The idea is because of oversperson, you know, people are doing, just focus on one thing. They do not see the fruits of their labor. They don't actually know what their impact is on the larger organization or the product they're producing. And therefore, they actually kind of feel depressed. detached and you're kind of acting like a machine and you're not actually fully fully human. And so maybe there's like a bit of a reaction to that.
Starting point is 00:24:02 And I think with our agents we're like, oh, there should be like one God, God agent. But in fact, specialization is actually like really good for machines. And that's like the point that you're making. And like humans should be general, but like machines should be ultimately a lot more specialized. The problem with, I think what I'm describing is that we don't know what good looks like. There hasn't been a system of specialized agents that feels
Starting point is 00:24:29 as elegant as like chat GPT or Claude or Muse, where you're basically just talking to one thing only. It's yet to be discovered. Maybe like opening I just launched dots. I think they're kind of like experimenting in that direction.
Starting point is 00:24:46 Grockbot, I guess, kind of counts. But they're all very general. Like I think the idea behind My dot is that it's like, it's like your digital double. At least that's what I understood it. Well, when I saw Grock Bot, I don't know what's going to, I don't know that much about dots yet, but they just came out. But Grock Bot, when it first came out, like the first use cases that I saw people talking
Starting point is 00:25:12 about were, oh, whoa, I can make two bots, one that knows my bank account, right? And one that knows my Twitter account. And the two bots, like, don't have the critical. credentials from each other. But they can talk to each other if, like, they need to get something done. There's no credential sharing, though. And that was, like, one sort of big, unique thing I saw, like, a couple times, like, pop up a couple times that people seem to like. But it looks like muse is...
Starting point is 00:25:42 But I wasn't sure. Muse and instinct have a much stronger product market fit than Grockbot. And perhaps... To you like that. Perhaps it is because you don't have to worry about creating these domains. But I think maybe personal agents are different than work agents. And I think your kind of your critique of general agents is more pertaining to work and to enterprise, which I sort of agree with.
Starting point is 00:26:14 And there's also all sorts of data access considerations. I think as CEOs, we can have general agents because we have. have admin access. But you know, you're for individual employees or certain teams, they can't have like truly general, you know, fully context aware agents because there's there's access control control issues. So you'll have to kind of work on something like specialization. Ultimately, I also think we need to figure out what is agent to agent communication looks like.
Starting point is 00:26:47 I don't think there is like good protocols around that just yet. I don't think that agents are trained to handle that very well. I think we've seen, it seems like, the next generation of Open AI models are trained to do agent collaboration because we've seen it in the hugging face hack where they started helping each other instead of emerge naturally. But there's also needs to be some way in which like an agent can't convince another agent to kind of give it information that it shouldn't give it give it like there's there needs to be like
Starting point is 00:27:28 data um isolation and and and and and and proper ways in which these agents communicate you almost don't want to communicate them fully in natural language maybe that there's like some other dsl or protocol that that they need to follow i really i i think one of the cool potential applications of jev and other decision models like it is going to be alignment, you know, checking to see if a tool call or like an agent-to-agent, you know, communication is aligned. Because there's just so many tool calls. Like you really need like a cheap, fast model.
Starting point is 00:28:08 If you're going to block something like that, then like a really, really fast decision model that just classifies and gives feedback on rejections. might be like a really good way to bridge the gap between agents and from agent to infrastructure too. So I haven't seen like, we have like a little prototype that we're running internally
Starting point is 00:28:36 at OpenRouter, but I kind of think a good, I think it could be like an interesting alignment. So you're using it for policy enforcement? Yeah, like imagine looking at the system prime, and the current tool call being made and be like, you know, is this aligned with the system prompt with the original agent and with these like extra guidelines that maybe we didn't tell the agent about?
Starting point is 00:29:03 For example, let's say you have a bunch of agents that are instructed to do to like red team the like some new product. And they cannot, should not be able to access the internet. And if they ever do, they should stop right away. But you might not want to, like, explain all of that to the agents doing the red teaming. You might want them to try to, like, break out of the sandbox and act like bad actors. Like, what would a bad actor do? It wouldn't be, like, break out of the sand, you know, try to, like, break into this company.
Starting point is 00:29:40 And the moment you do, stop, don't do anything else. So, like, having another model, more like, you know, use for, you know, building a neurodiverse system. Having another model, like, check every single tool call or every single assistant message to see if it's indeed aligned with something that wasn't in the system prompt, I think makes sense. And then having kind of like structural safeguards too, which is what I think Nvidia just launched with their like open agent safety. I think was called OpenShel. I think companies are probably going to explore a combination of those. I wonder another thing about specialization and instead of what you're talking about,
Starting point is 00:30:27 there's a lot of talk of recursive self-improvement. There's something I don't think is getting a lot of discussion, which is models training their replacements. It's sort of like, you know, you can think of it as just a, time compiler. So the way just in time compilers is, you know, as you're executing dynamic code, you know, the interpreter realizes that there's an opportunity to optimize. It will emit machine code on the fly, and that's a lot more optimized. So you can imagine models like general models, you're kind of doing something with Opus or some of the Astra, some of the big models.
Starting point is 00:31:10 And they realize that the use case is limited. or you prompt them in some way or some other agent observing and realizes that use cases are limited. And I think general agents have all these flaws that you just talked about, but also there's more potential for them to be harmful. There's more potential for them to go off the rails. And sort of on the fly trains a model
Starting point is 00:31:36 that could be its replacement, but is like a lot more domain specific. And therefore it is, cheaper and also you know less less vulnerable to prompt injections
Starting point is 00:31:53 less harmful for you know because it's less capable and it's almost like a like you know some system that's training machine learning models for specific use cases
Starting point is 00:32:07 as it's monitoring the entire system like would that specific use case be it would involve unstructured text generation or a very like structured decision model like it could be on structured text generation it could be decision models like even the case of of of jav like if you have if you understand the the inputs ahead of time you could potentially like you know taken off the shelf like quen or something like that and like train it specifically for that policy but for it makes
Starting point is 00:32:39 sense to do for cost reasons um assuming that there aren't, like, really, like, the model labs, the frontier model labs might make very low-cost models that you can easily transition to. But safety as well, right? Oh, yeah, I see your body. They're so capable. And so I think oftentimes people are using these big foundation AGI-like models
Starting point is 00:33:07 to, like, it's like nuking a butterfly, right? And it's like, yeah, they're very, you know, most of the times, like, a lot of the use cases, even unstructured use cases, don't need that capable model. I wish there were more public e-vals about this stuff. Like, but a lot of the e-vails about this are private. You just can't see whether, like, when the models get more intelligent, the risk actually will, like, continue to get higher. Because I think there's also an argument to be made that alignment will get better. and the models will like start to, you know, avoid going off and, you know, hacking on their own as we, as they get smarter and better at alignment, especially when it comes to agent to agent coordination. Like the, there's something Noam Brown said on a podcast recently. Like, as the agents have gotten smarter, they've gotten just better at coordinating.
Starting point is 00:34:05 they're, there's still like, and, you know, it's unclear if they're going to, like, be harder to align than humans when they're, when we get more and more of them. But if we can figure that problem out, then a smaller model, like, will it be harder to align? So, so I asked earlier is that, is that, I think it's true. that smarter models are more aligned naturally or very easier to align. Well, if you think back to the original sort of like rationalist less wrong arguments for AI safety, there is this thing called the orthogonality thesis. The idea is that intelligence is orthogonal to ethics or morality or, you know, so on. I don't believe that's entirely true with humans.
Starting point is 00:35:04 I think people who are generally like more intelligent, kind of more educated, tend to, tend to, not always, tend to, you know, be more considerate of animals,
Starting point is 00:35:16 for example. But, but, but, but in, in, in, in, in,
Starting point is 00:35:23 in, in, I, I think it could go the, the other way because, you know, there's been quite a bit of studies on, on, on, on, on, on,
Starting point is 00:35:34 showing like, you know, how reward hacking and deception, they just like get better at it. And like the evals could be, could be deceiving because the model could be smart enough to, to know that it's getting evaled. I mean, we already know this. It's been shown that if you do a lot of monitoring on a chain of thought, they start lying in their chain of thoughts. So you add pressure almost on the chain of thought on that that kind of creates. And I think at some point for you to do proper alignment evals, you need to run it for like months, right? You need to run this thing for months on like a really large, you know, goal or task in order for it to truly kind of figure out whether it's aligned or not. I always struggle with this word alignment.
Starting point is 00:36:35 It just feels like wrong for so many reasons. It's sort of like vague and sort of like aligned to what, whose values. And so it just doesn't make the conversation easier. I think in this case I'm talking especially about deception. Like the model is actually deceiving its user. I mean, maybe this kind of. reduces to, like, are we going to solve the line, you know, are we going to prevent the models from deceiving users during training runs predictably with, like, you know, better, like, well,
Starting point is 00:37:14 like a model that's big enough and powerful enough suddenly stop deception, stop sandbagging. And nobody knows the answer to that yet. So, like, at the point when that does. If that ever does happen, though, we might see kind of an interesting pressure for organizations to go towards the frontier. Oh, interesting. All to have no, you know, basically no risk or significantly less. Would they be willing to pay 10x to get that much? I mean, it probably depends on the, like, types of tasks they're trying.
Starting point is 00:37:59 trying to do, like, you know, some just have way lower risk than others. Writing code that are doing like security research is the highest risk type of task today. And so you probably spend 10x to get a fully aligned model that can also find all the bugs, or fully like anti-deceptive model that can also find all the bugs. of the coolest things about decision models that you fully control the structured output. And generally with structured output models in general, the room for misbehavior is so much lower. You just have, like, defined tasks, and only machines are, like, dealing with the outputs. And it's not writing code that it can execute.
Starting point is 00:38:52 that those tasks feel like probably underrepresented in the ones that people talk about and the things that enterprises are dealing with. So I expect like enterprises to get a lot more interested in them. Yeah, I feel like we're going to slowly realize how good we've had it with like deterministic code. Like, oh my God, remember the days when computers did exactly what we told them to do.
Starting point is 00:39:26 And I think things like Jeff, I think hint at like more of the need for, you know, not only specialized models, but models as output domain is more controllable. And maybe you could do, maybe you could do a lot more than we thought you'd need, you know, by you know, by you know, using like a bunch of specialized models,
Starting point is 00:39:55 specialized output models. Have you guys done any workloads internally with it? You know, I've been training a lot of small models. I mean, I said this glib thing when I first came out because I was like sort of, I gave this hacker news comment like comment, which I felt disgusted with myself afterwards. But I've been taking a lot of like Quinn 8 and like asking, honestly, asking,
Starting point is 00:40:22 you know, Fable and Opus and Astra to train a model. For example, I trained a cost estimator model internally so that when you put it in a prompt and replica, we know exactly how much it will cost. And it basically emits a,
Starting point is 00:40:37 you know, probability distribution over multiple buckets. Like if it is between $5 and $10, bucket A, bucket B between, you know, $10 and $20. And like, so I'm used to the training these classifiers by, you know, giving it different enums essentially and looking at the log props per enum. I've been doing it for a couple of years.
Starting point is 00:41:00 I trained a chat spot to play by just doing that. So I'm already sort of pilled on like sort of decision models and specialized models. So it wasn't that big moment for me. But I understand that like a true foundation model that's fully promptable, is like amazing user experience, amazing developer experience
Starting point is 00:41:25 and you can do a bunch of stuff without training a model from scratch. But if you have a data, if you work at a place where you have the data, you have so much data at Replit, like I ended up training a lot of specialized classification models pretty easily. Yeah, like I definitely,
Starting point is 00:41:39 it also feels like less model debt. Like something that I still hear from companies is that they're worried about fine-tuning models for like unstructured outputs. because you're just like always, you got to redo it again in like two months. And everybody just feels the model, the weight of the model debt. But like a very bespoke classifier that's trained with like proprietary data, I feel like people won't think it's behind constantly.
Starting point is 00:42:07 And it might just like last longer. Yeah. It's just, you don't have to worry about its ability to speak a new language or write rust or, you know, do anything that the LMs are being like evaluated on. You know the use cases so you can like build it more. And it seems like an easy thing for enterprises to build themselves and actually like not regret. Yeah. Speaking of Russ, actually like a good analogy is when the world got super excited
Starting point is 00:42:38 about dynamic languages. Like if you think back to the 90s, everyone was writing in Java and C++ plus, things like that. And then like Python, JavaScript, Ruby, just like took over the internet and everyone was like, oh, this is how you built startups really quickly. This is, and you built Stripe, a financial organization on Ruby.
Starting point is 00:42:57 I was like, how crazy is that? And we built Facebook, you know, using PHP. And then everyone was like, oh shit, like we're running into all these really bad bugs.
Starting point is 00:43:07 It's like freaking slow. So like, let's go in and add types. Okay, let's add a shit. Compiler. And you end up sort of reinventing everything. And then Russ came out.
Starting point is 00:43:20 I was like, okay, I guess we can use Russ for a lot of things we would otherwise be using JavaScript and Python. And my prediction is that the same cycle will happen here where we're using these AGI-like models for all these different use cases. And then everyone's going to wake up and be like, oh, my God, this is like so wasteful, so risky for no reason. And it's got to be so much easier. We're actually adding that capability on Rapplet,
Starting point is 00:43:47 but I think it's going to be everywhere. It's going to be so much easier to like to go to a site, like upload a CSV file and get a special model that does one thing. And, you know, that goes back to your thesis about open router, this like neurodiversity, which I like really fundamentally believe in a lot more. And Eric and I had like discussions a lot about like AGI and whether we were truly on a path to AGI or whether it's even desirable to get there. I think the future is a lot more diversity. Outside of code review, which was like, I think the first time I saw people get really serious about using, like, different model families
Starting point is 00:44:28 to double check the results of their main model. The, these, like, fusion models, like, the research has been getting, has been, like, kind of slow for years on, like, doing mixture of models. and composite models. But things have been speeding up from my view of the research.
Starting point is 00:44:57 And I mean, now we see like a bunch of AI agent labs. Like we launched a fusion tool, a fusion model, and technician launched one. And like they do reduce cost, and allow like, a wider breadth of ideas to be searched. Like our initial launch was focused on deep research. The thinking is that like if all these model apps are training
Starting point is 00:45:26 on different sources of data, like, why not pull from all of them? And this like resulted in basically fable level quality at 2x lower cost. We just published results actually just today about that. We showed like a deep suite like rapporteur, application different things, including the harness, but also you might think of it as a fusion type thing, where it's like, you know, frontier level at like 40 to 50% of the cost. Which models are it used?
Starting point is 00:46:07 I think it changes over time. But one thing that has been interesting is I think OpenAI added this feature that allows you, to save the computation, like, across different model families. So you can, like, also across different effort levels. So, like, you wouldn't do, you wouldn't miss the cash if you change the effort. Don't call me this. I think it's also across different models, which is hard to fathom how. I might be wrong, though.
Starting point is 00:46:47 but I need to double check that. But I think, you know, seeing within the Open AI family has, like, added a lot of efficiencies. But in the past, we've done it with other, with other models as well. Because cash is like one of the big, like being cash aware is like one of the biggest things
Starting point is 00:47:05 when you're designing fusion models, routers, sort of escalation models. Alex, I'm glad that's been a great conversation. Thank you. Thanks for listening to this episode of the A60s podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family.
Starting point is 00:47:26 For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X at A16Z and subscribe to our Substack at A16Z.com. Thanks again for listening, and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments
Starting point is 00:47:58 in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.