Latent Space: The AI Engineer Podcast - OpenRouter: from Seed to Stripe — with OpenRouter’s Alex Atallah & AMP’s Anjney Midha

Episode Date: September 25, 2026

From the earliest days of open-weight models to becoming the neutral routing layer for more than 10 million developers, OpenRouter is one of the clearest bets that the future of AI will be multi-model.... In this episode, OpenRouter co-founder & CEO Alex Atallah, with AMP’s Anjney Midha returning with swyx to unpack how OpenRouter emerged from the first wave of Llama, Alpaca, Mistral, and Midjourney, why model diversity mattered before it was consensus, and how a company dismissed as “just a wrapper” became critical infrastructure for the AI ecosystem.We go deep on the product and distribution lessons behind OpenRouter: why model labs can spend billions training a checkpoint and still struggle to get it into developers’ hands, how Mistral helped prove the value of a competitive inference marketplace, why OpenRouter chose focus over expanding into fine-tuning, memory, and other adjacent products, and how its rankings became a real-time map of how AI usage was changing. Alex also explains OpenRouter’s early experiments with model fusion, why they deleted the first version and brought it back years later, and how the platform grew to more than 10 trillion tokens per day.Finally, Anjney explains why Stripe and OpenRouter fit together, why token fraud may become one of the defining security problems of the AI economy, and why the next wave of fraud won’t just come from humans but from autonomous agents attacking increasingly valuable token flows.We discuss:* Why OpenRouter bet early that no single AI model would win everything* Alpaca, Llama, and open models becoming impossible to ignore* Why Discord’s early AI deployments exposed the limitations of closed models* Why model labs can spend billions on training and still fail at distribution* How OpenRouter became a neutral distribution layer for model developers* Why VCs dismissed OpenRouter as “just a marketplace” or “just a wrapper”* The Mistral price war and the first real proof of an inference marketplace* How Midjourney scaled through Discord and what it taught the AI ecosystem* Why crypto infrastructure became a dress rehearsal for generative AI* OpenRouter vs. LM Arena and why their missions are fundamentally different* Why focus became one of OpenRouter’s biggest strategic advantages* Anthropic’s early focus on AI pair programming and coding* The OpenRouter products that were prototyped but never launched* MOM, OpenRouter’s early Mixture of Models experiment* Why model fusion failed in 2024 — and why it works much better now* How OpenRouter’s leaderboard became a live map of the AI industry* OpenClaw, auto-routing, and agents reshaping AI usage* How OpenRouter reached 10+ trillion tokens per day* Why inference gateways are increasingly becoming targets for fraud* Why Stripe’s fraud infrastructure is strategically important to OpenRouter* The coming rise of agentic fraud and attacks on the token economy* What changes and what stays the same as OpenRouter joins StripeAlex Atallah* LinkedIn: https://www.linkedin.com/in/alexatallah/* X: https://x.com/alexatallah* Website: https://alexatallah.comAnjney Midha* LinkedIn: https://www.linkedin.com/in/anjney/* X: https://x.com/AnjneyMidha* AMP: https://www.amppublic.com/Timestamps00:00:00 Introduction00:02:12 Alpaca, Llama, and the Multi-Model Bet00:06:04 Discord, Open Models, and OpenRouter’s Origins00:14:28 Why “One Model Wins” Was the Wrong Bet00:17:27 Why Model Labs Struggle With Distribution00:23:04 “Just a Wrapper”: Why VCs Misunderstood OpenRouter00:27:58 Bootstrapping OpenRouter Through Community00:36:16 Crypto, Midjourney, and the Early Generative AI Ecosystem00:43:38 Mistral and the Birth of the Inference Marketplace00:47:10 OpenRouter vs. LM Arena00:52:08 Focus, Anthropic, and Roads Not Taken00:59:34 Mixture of Models and Model Fusion01:02:44 Sonnet, OpenClaw, and OpenRouter’s Explosive Growth01:09:03 Why Stripe Acquired OpenRouter01:12:45 Fraud and the Emerging Token Economy01:17:47 The Coming Wave of Agentic Fraud01:19:07 What’s Next for OpenRouter at StripeTranscriptIntroduction: OpenRouter, Marketplaces, and Pub-Sub as a Product PrincipleSwyx [00:00:00]: Okay, we are here in Anja’s house, which is where all big startups in San Francisco start.Anjney Midha [00:00:08]: Howdy.Swyx [00:00:08]: And, congrats on Cursor, Mistral. I don’- God knows what else. You got so much stuff going on.Anjney Midha [00:00:17]: There’s, there’s a lot going on. Well, OpenRouter is probably the - has been the most, I would say, like, one I’m excited about recently.Swyx [00:00:24]: Yeah. And we have Alex, first time on the pod, but,Anjney Midha [00:00:27]: Thanks for having me.Swyx [00:00:27]: You’ve been in the IE a few times. I appreciate every time you’ve shown up, for the community. Congrats. I just, like, what a journey. When I was looking back at your past posts, one of the earliest principles that I saw you write as a product person is sub as a product principle. And I wanted - you to maybe explain how you think about what should exist in the world.Anjney Midha [00:00:49]: Yeah. The sub piece, which was early 2023, I didn’t think about it until we talked like 10 minutes ago, is about how there is like a way of thinking about products as an intersection between subscribing to data and publishing data. And marketplaces are an easy example of this. You have suppliers that are publishing some product to a SKU. And the SKU is like a sub topic that a consumer is subscribing to and just going to, like, consume whenever they want. And humans consume in a very, like, discreet, ad hoc way. It’s not very scalable. all their attention is on the topic when they’re buying the thing, and their attention is nowhere else when that happens. agents and consumers of inference don’t act like that. They’re consuming continuously, and they’re changing the SKUs that they consume from all the time. So OpenRouter is like a blend between a normal API experience and a marketplace where we create model slug. We have the auto router. We have all kinds of, like, product SKUs that you can subscribe to. And then you can, like, continuously add, like, derive value and make decisions based on those consumers.Alpaca, Llama, and the Multi-Model BetSwyx [00:02:11]: Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and, that people would not use the native SDKs. I guess, for each of you, what was your realization moment that this would be it? I, - You’ve, you’ve given a talk at EIE about Alpaca as,Anjney Midha [00:02:33]: Yeah.Swyx [00:02:33]: One of your inspiring moments.Anjney Midha [00:02:35]: Alpaca, I can, like, rehash the Alpaca moment for a sec. Like, the very beginning, at the end of 2022, OpenAI was the only game in town. There was, like, OpenAI, Cohere,Swyx [00:02:47]: Yes.Anjney Midha [00:02:48]: And then a smattering of, like, early attempts at open weight models.Swyx [00:02:54]: Yeah.Anjney Midha [00:02:54]: When Llama came out in January of 2023, it was like, “Wow, really exciting. This is really big.” It outperforms 3 on, one or two benchmarks. but you can’t chat with it. It wasn’t like - It wasn’t an engaging model, but it seemed like someone just needed to fix a couple things and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, tuned Llama, and made Alpaca, billion parameter model. Or was - Maybe it was thirteen billion parameters. And it was so good. Like, I was just, like, on an airplane using it. I, - in many cases, I, like, you could not discern a ChatGPT versus an Alpaca result. And I figured if it was this easy to make a model, one, we have a whole new way of monetizing data for the first time. you can just, like, take really valuable data and turn it into a service in $600. and that cost will probably go down over time.Swyx [00:04:03]: When you - So sorry. when you say monetizing your data as, what eventually will become an MCP endpoint or as a training data for a model?Anjney Midha [00:04:12]: Yeah, training data for a model.Swyx [00:04:13]: Awesome.Anjney Midha [00:04:13]: Like, an abstract way of saying like, “Hey, I have this data.”Swyx [00:04:15]: Compress it into a model.Anjney Midha [00:04:16]: Like, it makes sense for me in my product, but, like, I could repackage it in the form of a model and sell it. And so it’s just a whole new business model for the economy. It also, of course, provides, like, a way of following what Frontier Labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so - Whenever you have an example of that, like a breakout app that’s doing really well, and then some framework for imitating it with - in your own flavor, you have an immediate ecosystem of, like an immediate ecosystem, like, should arise because there’s just a huge gap between the, like, decisions that the single company is making and all of the variations in those decisions that, like, a wider ecosystem can create themselves. And so then, you need a marketplace to, like, discover all of those, services and all of those products. There wasn’t any place on the internet that, like, was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why.Swyx [00:05:29]: The closest would be Hugging Face.Anjney Midha [00:05:30]: Hugging Face was the closest at the time, yeah.Swyx [00:05:31]: They just started Hugging, like, a few years ago before that.Anjney Midha [00:05:34]: Yeah, and Hugging Face also didn’t have the closed-source models.Swyx [00:05:37]: Yeah.Anjney Midha [00:05:38]: And they didn’- you couldn’t use the models at the time. and there wasn’t data about who was using them. There were, like, a bunch of differences between OpenRouter and Hugging Face, and those differences felt really critical to me, especially when I was just trying to learn about LLMs and, like, why people are choosing, like, Different little ones that are emerging over time.Discord, Open Models, and the Origins of OpenRouterSwyx [00:06:03]: Got it. And then, Ansh, no stranger to wanting more model diversity, at the time, you’re a couple of years into your Anthropic journey, which we covered in the previous podcast as well. What was your introduction to Alex?Alex Atallah [00:06:16]: Well, the introduction was, I think, thirteen years before that.Swyx [00:06:20]: Oh.Alex Atallah [00:06:20]: But the OpenRouter handshake happened right over there, if you remember.Anjney Midha [00:06:23]: Yeah.Alex Atallah [00:06:24]: Which - So Alex and I, met, I believe as sophomores now, if I remember at the Stanford Review,Anjney Midha [00:06:32]: That’s rightAlex Atallah [00:06:32]: Meeting for the first time.Anjney Midha [00:06:33]: I think so, yeah.Alex Atallah [00:06:35]: Yeah.Anjney Midha [00:06:35]: Yeah.Alex Atallah [00:06:35]: So Stanford Review was the libertarian newspaper on campus at Stanford that Peter Thiel started back in the day. And, whatever-- for whatever reason, I, Alex and I both showed up to one of the meetings, and I remember, the editor-chief was a mutual friend of ours. Lisa was really a really great editor-chief, where, part of an editor-chief’s job is to assign responsibilities to people and make sure the work gets done. and I, I may be misremembering the details, but I remember wanting to. It was surprising to me that at the time there was no dedicated technology section in the newspaper.Alex Atallah [00:07:11]: YouSwyx [00:07:13]: Because it’s political, right?Alex Atallah [00:07:14]: It is primarilySwyx [00:07:14]: Like, it’s talkingAlex Atallah [00:07:15]: It originally started as like aAnjney Midha [00:07:16]: Yes.Swyx [00:07:17]: Yeah, states and all those things.Alex Atallah [00:07:17]: Correct.Swyx [00:07:18]: Yeah.Alex Atallah [00:07:18]: But it, - To take us back in time, you may remember this, but, there was this technology, legislation that was being debated called, the Net Neutrality Act. And net neutrality is, like, inherently this political concept, right? It’s, it’s about the regulation of - internet broadband access. And so there was a community of us who were technologists, but also debating the politics of the technology. And I thought the Review would be a great place - to, like, write about that. And I was working on, I think, a net neutrality article, and I remember proposing, “Well, maybe we should start a technology section.” And Alex was one of the only people who said, “Yes, that would be cool.” And said. I forget whether we ended up writing stuff together, but - that’s when we first met,Alex Atallah [00:08:03]: Was 2011 or twelve. I forget which year it was. It was one of those.Anjney Midha [00:08:09]: Yeah.Alex Atallah [00:08:09]: It was at Old Union, if I remember correctly.Alex Atallah [00:08:11]: That’s where we used to meet. But, along the way, Alex and I have had a chance to, To hang out often. And probably the time when we had the most professional overlap was when I was running the platform at Discord, and it had become this explosive platform for cryptoSwyx [00:08:32]: YeahAlex Atallah [00:08:32]: And NFTs in the middle of the pandemic.Swyx [00:08:35]: Which also, by the way, you were in charge of safety and security as well, right?Alex Atallah [00:08:38]: I was the head of platform, which meant all of the crypto - the DAO and NFT launch security debugging fell onSwyx [00:08:45]: And their phishing and.Alex Atallah [00:08:47]: The phishing, the social engineering attacks, the katana DDoS that we were getting hit by. but it’s around the time I first started teaching security at scale at Stanford, CS 153. And Alex was on the, - at OpenSea at the time, and I was trying to figure out how we could defend against all these attacks that we were. Like, and at peak, I forget, if you remember how much NFT volume was running throughSwyx [00:09:10]: DiscordAlex Atallah [00:09:10]: Discord, but it was, like, a meaningful amount of, like, it was, like, several billion dollars in NFT volume of GMV, so to speak, were running through the platform, and it was all coming from OpenSea. It was these, like, buy, sell,Swyx [00:09:20]: TheAlex Atallah [00:09:21]: ServersSwyx [00:09:21]: The D in DAO is Discord.Alex Atallah [00:09:25]: Yes. And so that’s when I think we had hung out professionally. But a year after that, OpenAI gave Discord early access to GPT. Sorry, three. No, it was five. Yeah, five, which is the RL version of three. And that’s around the time we made a Discord bot with, OpenAI for internal deployment, and that’s when I realized we would need. Like, since I was part of the deployment team.Anjney Midha [00:09:50]: What was the use case?Alex Atallah [00:09:51]: There were two that were. And there’s, there’s a post now called “Discord is Your Place for AI with Friends” that somebody sent me recently that I wrote, and published in twenty-three. But There were two use cases. One was Clyde, which was the - like, a party friend inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more. and then there was content moderation. And one of the realizations we had with content moderation was - it would refuse to moderate. Like, it would just refuse our prompts because the The training was. We were very early in the training era, and it would just. Our prompts would trigger it, its, like, guardrails. And we told OpenAI, “Hey, guys, we need access to the weights because if we’re gonna be doing content moderation at scale, we had 250 million monthly active users, we need more reliability that the model will do what we need it to.” And they said, “Well, sorry, guys, that’s not how this works. We’re a closed-source company.” And so that was my first realization that we needed open models, and the enterprises would need more control over capabilities, and then ultimately would need some control plane or management system to orchestrate these open models. But there weren’t no good - there were no good open alternatives until maybeAlex Atallah [00:11:10]: Six months later when Llama came out. And six months after that, I led the series A into Mistral, which was started by Guillaume and the Llama team. And - That, - Around that time is when I remember hearing about Alex launching OpenRouter and going, “These worlds are gonna collide, and I don’t know when it’ll make sense to team up.” But Alex was so early and could see. I think he was totally right about this ecosystem starting with Llama that then needed, like, a, an easy layer to manage for, especially for. I was approaching it from the enterprise perspective because I had been that, like, the. As the VP of platform at Discord, it was my job to ensure that when we deployed models to, like, 250 million users, they did what we wanted them to. And that was very hard, because if you outsourced it to the labs and they controlled the guardrails and their guardrails are their safety policies. Forbid the model from responding to your prompts. That was quite catastrophic.Swyx [00:12:05]: Yeah. But what, a moderation is the thing that they want to support. And obviously, beyond that, they would - OpenAI would work with you, presumably to give you a moderation endpoint, which they offer for free.Alex Atallah [00:12:16]: It was an interesting use case, that - So they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the, - Often, like every, subreddit, Discord servers, public ones have their own rules that the user, the users create.Swyx [00:12:41]: Oh, yeah. We run the LinkedIn Discord in. Yeah.Alex Atallah [00:12:43]: And then humans used to read those norms and then enforce it every day manually, like observing each message in these communities. And these communities have like millions of users. So we had a 5,000+ person team globally in the, on the Discord content moderation team. These are outsourced contractors who had a really tough job. And so the idea was instead, if you could give the norms of that server To the LLM, then the LLM would do custom moderation for that server. It’s almost like a, like context moderation for that server. And many of those servers’ norms just violated OpenAI’s rules. And so - It was like we had our own custom eval. So each server had its own custom eval. But Discord-- at the time, OpenAI’s evals, we were all soAlex Atallah [00:13:28]: Primitive in our thinking about how to deploy these LLMs that often the training prompts were super handed. It said, “Oh, anything about Harry Potter, anything that has trademarked content, don’- refuse.” And if it was a fan - Harry Potter fan community, this is a real use case, that had content moderation, the LLM would just refuse.Swyx [00:13:48]: Yeah.Alex Atallah [00:13:49]: And that was just not precise enough.Anjney Midha [00:13:52]: Another one that we heard was like if someone was trying to write like a detective story, and there’s one chapter with a lot of violence, like maybe someoneAlex Atallah [00:14:01]: RightAnjney Midha [00:14:01]: Like kills someone, the LLMs would just refuse to, like, help with that part of the story.Alex Atallah [00:14:07]: Yeah.Anjney Midha [00:14:07]: And then - like, we used to be like, okay, this is not like structurally inherent to LLMs. There must be, like, some choice out there so that I can, like, switch to another model, when I’m getting, like, a refusal or a bad result from the main one that I have. And that, like, tension also drove me for a marketplace.Why “One Model Wins” Was the Wrong BetSwyx [00:14:28]: Yeah. I think that is well accepted now. What was it like back then when you were raising or, starting this? did people get it? what was the, some of the struggles? I like getting stories out of him about how other VCs don’t get it. So like anything you wanna, talk about, now - Let’s, let’s call it, that the early journey of OpenRouter is done, right? You can obviously talk about some of the early days stuff.Anjney Midha [00:14:54]: Well, I was gonna say that, like, the biggest objection we got is big model win, which is - all of theSwyx [00:15:03]: Scaling laws.Anjney Midha [00:15:04]: Huh?Swyx [00:15:04]: Scaling laws.Anjney Midha [00:15:05]: Yeah, scaling laws, and natural network effects are just gonna accrue to one company, which will be - It’ll be a Google-style monopoly, just like how Google won the search market, by a large margin, and you’ll just be fighting for scraps at the end. That was probably the biggest objection we got. it is interesting that Google won the search engine race with such a huge margin. I think, like, had there been more interesting benchmarks or had, like, search engines been, - had people, like, seen them a little bit more like LLMs where they’re services that you can build companies on top of, that might not have been the case. but LLMs don’t merely have a user interface. They’re also, like, ways of building entirely new businesses. And, a Google-level monopoly would be like the Dutch East India Company times, quadrillion in magnitude because the whole economy ends up, like, depending on the one monopoly as well. So it didn’t seem like would be a really crazy outcome if that happened. And it’s also less likely because the economics of, like, creating good competitors are much, like, much more decentralizable.Alex Atallah [00:16:25]: Everything Alex said is true, And I came at it from a completely different perspective, whichSwyx [00:16:31]: Yes, this is why we’re here.Alex Atallah [00:16:32]: The scaling laws were never - In my mind, were always a feature, not a bug for why OpenRouter would be very valuable. Because, I was one of the first investors in Anthropic, and it was obvious to me that other researchers in our friends - I went to grad school for machine learning, and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, “Oh, fantastic. Now we have at least two proof points that compute scaling works.” It was OpenAI and Anthropic. and by the time I think we decided to team up on OpenRouter, I had already invested in Mistral and Black Forest Labs and Luma. So there was multiple model companies and teams that I was, working with.Why Model Labs Struggle With DistributionSwyx [00:17:14]: But you did other modalities, whereas this is literallyAlex Atallah [00:17:16]: Across different modalities, yesSwyx [00:17:17]: Text.Alex Atallah [00:17:18]: Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created, and that this whole narrative of, like, Only one company will dominate like Google was, well, like maybe true, but one, I don’t believe that. But two, there was so much extraordinary innovation happening across several different research teams. But the shared problem I was noticing across all of them was often, the research teams were fantastic at figuring out how to reason about new capabilities. They think in terms of capabilities, but never - like, are not developer mindset-oriented. Like, what happens after the training is done and the checkpoint comes out? Like, you’d be shocked how, like, similar the early training teams at OpenAI, sorry, Anthropic, BFL, Mistral, were in their, like, default approach to. Taking their research out of the, lab and scaling their impact, which is often, oh, the checkpoint is done, put it out as an API, done, and then there’d be crickets. in the case of Claude, the first Claude checkpoint was done a year before they released it internally. And then ChatGPT came out, and we decided, okay, yes, it’s a good idea to release a Claude version externally.Alex Atallah [00:18:34]: And they had no plan, like no plan for how to get developers to try it out. And so if you go to the Claude one blog post, you’ll notice there are, like, three developer examples for users of the API, and one is a Discord bot, and the second is Vivian, my wife’s startup called Juny Learning, ‘- And then there was, like, Notion, because these were all friends of, like, the Anthropic Because that’s how - like, last minute the planning was around, hey, once the model’s done training, how do you get it out to the world? There was no distribution platform that understood what developers needed, all the key management, provisioning, like, simple, like, endpoint management, versioning control. Like, all these things that the scientists and researchers go, “ that’s plumbing. I don’t really think about it.”Swyx [00:19:15]: Implementation detail.Alex Atallah [00:19:16]: Right. And instead, Alex came at it from that perspective. And so, it was so obvious to me that, like, every single lab I was funding would spend - like, literally sometimes billions of dollars into training, and then a checkpoint would be done, and there’d be crickets, like, during early access because they’re like, “Oh, that’s right.”Alex Atallah [00:19:35]: It’s hard to use a checkpoint to make anything. You need a whole bunch of plumbing around it to make it usable by a developer. And so by the - I think - it was so obvious to me that a distribution platform like OpenRouter was critical to have in the ecosystem if we wanted there to be competition to Google. Like, unless-- ‘cause with Google, DeepMind is done training a new checkpoint, and then they push a button, and it gets blasted out across all their surfaces from Google Docs to,Swyx [00:20:01]: Everywhere, even if I don’t want it.Alex Atallah [00:20:02]: Everywhere. You wanna know about, like, on Android, like, overnight, they can deploy a new checkpoint to, like, a billion devices, right? And that invisible infra advantage, distribution advantage, most people don’t realize, but until OpenRouter showed up, - you had to think about all of that yourself as a model lab. And it was very daunting. at Anthropic, I think it took, well, more than twelve months to get to our first 10 million in revenue. And in contrast with Black Forest Labs, I remember the early days, you guys had a conversation with the BFL team, and, it was so simple for OpenRouter to say, “Oh, no problem. Like, the day you launch, we can send 1 million developers to you.” that was crazy. That was like a step function change in, like, an hour.Swyx [00:20:46]: Is that a real number, a million?Alex Atallah [00:20:47]: I,Swyx [00:20:48]: Okay. All right.Alex Atallah [00:20:48]: I think today it’s, like, 4 million. How many developers are on OpenRouter today?Anjney Midha [00:20:52]: Over ten,Alex Atallah [00:20:54]: Yeah.Anjney Midha [00:20:54]: Over 10 million, but, like, it’s, it’s hard to, youAlex Atallah [00:20:59]: I, yeah, I don’t know how to. Yeah.Anjney Midha [00:21:00]: We do a lot of, like, account duping work, but, noAlex Atallah [00:21:04]: If you could get 1,000 developers, just to put in context If you get 1,000 developers who try the model on day one after you release it and just, like, do inference and give you feedback, that’s a thousandAnjney Midha [00:21:15]: That’s hugeAlex Atallah [00:21:16]: More developers than they knew how to get to on their own.Swyx [00:21:19]: Well, BFL had a reputation, but yes.Alex Atallah [00:21:21]: They had one in Stable Diffusion.Swyx [00:21:22]: Yeah.Alex Atallah [00:21:23]: And with Mistral, I don’t know if you guys remember, but the first checkpoint they released was, like, torrents. It was, like, torrent weights.Swyx [00:21:31]: Yeah, they just put up a magnet link.Alex Atallah [00:21:33]: Yeah, there was no API.Anjney Midha [00:21:34]: Yeah.Alex Atallah [00:21:34]: Because they didn’- they weren’t infra people.Alex Atallah [00:21:37]: ? Like, it’s like, okay, download these weights, and you guys go figure out how to host it.Swyx [00:21:39]: Well, he has a story on his side, yeah.Anjney Midha [00:21:41]: Yeah, in addition to the, like, building a really good developer experience around it, the marketing that we do on, like, for different models is totally different and perceived totally differentlyAlex Atallah [00:21:54]: RightAnjney Midha [00:21:54]: From the marketing that a model lab does for itself.Alex Atallah [00:21:56]: Yes, 1,000%.Anjney Midha [00:21:57]: Right? We are like a, neutral layer looking at this market like it’s a big dark room with all the corners completely obscure to users, and users are walking into the room and, like, feeling aroundAlex Atallah [00:22:09]: YeahAnjney Midha [00:22:09]: And trying to figure out what objects to grab off the tables and, like, build into, their companies. And it’s just an insane way of working. Like, models are not products where you can just enumerate all their features onto a web page. They’re all black boxes, including the open weight ones. So you need to, like, shine lights on all corners of this room, so that people can see what makes this model good, and you need the company shining that light to be a neutral third party, which is what we specialize in. So the, like. It’- In addition to developer experience, there’s also, like, a very important, like, marketing and product packaging componentAlex Atallah [00:22:50]: YeahAnjney Midha [00:22:50]: And a way of, like, routing and discovering models becomes, like, critical to your market as a provider or a model lab or a server tool and more in the future.“Just a Wrapper”: Why VCs Misunderstood OpenRouterAlex Atallah [00:23:03]: And this value, to your earlier point about how many VCs, like, just don’t. One of my biggest frustrations is that venture capitalists, many of them, like, just don’t have any operating experience in the field. so unlike a traditional investor who’s just maybe come up through the ranks as, like, a associate working on financial modeling or maybe hasn’t been a real operator in the field for, like, more than ten years, which is a big part of the industry now, I had just arrived at a16z, like, a year after running the platform. And so I knew what the challenges were of, like, building a real - great developer experience and like, being able to create a working piece of software with a model. And there were a few, I won’t name names, but there were investors who were looking at OpenRouter, and, felt at the time, like, when I would compare notes with people, that it was just, I quote unquote, “just a marketplace.”Swyx [00:23:59]: Yeah, just a thin layer, just aAlex Atallah [00:24:00]: CorrectSwyx [00:24:00]: JustAlex Atallah [00:24:01]: A wrapper or whatever on other people’s APIs. And I was like, “You have no idea how strategic the value that OpenRouter has created by being able to orchestrate even three.” APIs in production. The amount of both engineering work and community design that goes into getting that live and running in production at the scale the OpenRouter team had started just doesn’t happen by default. And that was one of the things that stood out to me about Alex from the earliest days. Like, he just understood, like, - from a systems perspective, like, how do you get these flywheels going? Like, that stood out to me with OpenSea when we were working together on the NFT integration at Discord. Like, Alex had a level of community-- like, systems thinking on how you get these flywheels going that most scientists and machine learning people just don’tAlex Atallah [00:24:48]: Think of. Like, we often think in terms of training.Swyx [00:24:52]: It’s a linear stage.Alex Atallah [00:24:53]: It’s this linear pipeline.Swyx [00:24:53]: There’s no loop yet.Alex Atallah [00:24:54]: Yeah. It wasn’t until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like. Like, mostly we did a lot of ML, like, when I was in grad school on a laptop. So you just, like, download a dataset, ran some ablations, and you looked at the loss curves, and you’re like, “Great, I made AI.” And the idea that you have to, like, deploy those capabilities, collect feedback trajectories, then, like, put those into a continuous loop, like, came much later. And it was very counterintuitive to the - like, the traditional AI mindset. I do remember doing the investment phase for, OpenRouter, I just didn’t try and educate a bunch of other VCs on why it was not just a marketplace. I was like, “ what? I’m just gonna invest.”Anjney Midha [00:25:41]: Yeah.Alex Atallah [00:25:41]: And I’m going to, like, take the opportunity to partner with Alex, and if - no other VCs get it, that’s totally fine. ‘Cause at the time, - it was not obvious, I think, to several of the investors that, like, OpenRouter was not more than just a wrapper around APIs. And - that infuriated me. And I was like, “ what? I don’t have time to debate you. I’m - we’re gonna, we’re gonna invest.” And then I think, like, a month later, Matt Murphy marked it up by 10x. Like, - I think. I forget what the exact money was and so on, but, to his credit, Menlo Ventures realized, “Okay, there’s much more strategic value here as well.” Maybe you didn’t hear all these conversations behind the scenes But that frustrated me a lot. there’s a lot of this, like, opining about wrappers. and if you’re like, “Oh, an app is just a wrapper on a model,” then, like. And, OpenRouter is, like, this wrapper on top of other APIs, and this is the most stupid, reductive framework.Alex Atallah [00:26:31]: And so it’s clearly somebody who has no experience deploying product at scale.Swyx [00:26:34]: It’s the thing you dismiss other things with. Like, you’re a - everyone’s a wrapper on everything, right? Like, and there’s, there’s some Some wrappers have value.Alex Atallah [00:26:40]: Investors are wrappers and LPs, right?Alex Atallah [00:26:42]: Like venture capitalists. So, yeah, it’s all wrappers down, all down to bare metal, I guess, and like energy.Swyx [00:26:46]: Yeah, there - When I started the whole AI engineer, I guess, the coining, in 2023, like, that was, like, the number one pushback is that this is no value. You should just train models.Anjney Midha [00:26:56]: Right.Swyx [00:26:57]: And, yeah, obviously this is, like. you guys are one of the testaments to the fact that you can build very valuable wrappers, but also very valuable model companies.Alex Atallah [00:27:06]: It’s so, hard to be. Like, the day a model launches, the fact that you have an OpenRouter, endpoint for that model frequently at the top of Hacker News on day one, people don’t realize the amount of work that goes into accomplishing that. And OpenRouter used. Like, that would happen over and over again, and I remember going, “People have no idea how hard that is.”Alex Atallah [00:27:30]: That’s not.Swyx [00:27:31]: Yeah, we’ve covered some of the inference engineering that goes behind,Alex Atallah [00:27:34]: YesSwyx [00:27:34]: Some of - with Base Ten and all those. Well, today you have, all those, like, cool code name things that people guess what Oxy Alpha is and all those things. But, like, I guess one of the things that you’re teasing is, how do you get that initial flywheel going, right? Because today you have your scale and your reputation, all these things, so obviously you - you’re driving immense distribution. But when you were early on, when it’s mostlyBootstrapping OpenRouter Through CommunityAlex Atallah [00:27:55]: The bootstrap, yeah.Swyx [00:27:56]: Yeah.Alex Atallah [00:27:56]: What was the bootstrap like?Anjney Midha [00:27:58]: To bring it back to early Discord days, I think we, like, initially connected with. This is an OpenSea story, technically. But, and we initially connected when you were at Discord, and we talked about, like, - the Axie Infinity server.Alex Atallah [00:28:13]: Oh, yes. Yes.Anjney Midha [00:28:14]: This server was, like, the biggest server at theAlex Atallah [00:28:17]: YeahAnjney Midha [00:28:17]: At Discord.Alex Atallah [00:28:18]: That’s right.Anjney Midha [00:28:19]: And you were like, constantly bumping up theAlex Atallah [00:28:22]: The limits on the server. Oh, my GodAnjney Midha [00:28:24]: Of how many people could be in the server.Swyx [00:28:24]: For those who don’t know, like, 10% of Philippines was Axie.Alex Atallah [00:28:29]: Was on that server. That’s a big hit.Swyx [00:28:31]: It was, like, a meaningful contributor to the GDP of the country.Alex Atallah [00:28:33]: It was an NFT, like, crypto game, but itSwyx [00:28:35]: It was like a Pokémon breeding thing.Anjney Midha [00:28:36]: Yeah.Alex Atallah [00:28:36]: Yeah. Similar. Yeah. There was battling, there was breeding, and then there was, like, a marketplace for trading.Swyx [00:28:43]: Earn as well.Alex Atallah [00:28:45]: Yeah, earn. And, like, the graphics were really cute and fun, and you like, you get emotional about your Axie that you make. So to, like, start a community like that, which we had to do many times at OpenSea with every early project, for us to create a marketplace for it, we need to make sure that the, like, the community wants it.Anjney Midha [00:29:09]: Right.Alex Atallah [00:29:09]: And it’s like building something that people want and going and telling them about it. Like, you can do that on a one basis, but there’s way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time, like, building things that the community really wanted. We did the same thing for OpenRouter. And, like, the Axie community was one of, like, a zillion communities we did that with. And Anj, like, saw us doing it and. ‘Cause you could just see people sharing OpenSea links constantly in that Discord. Like, users sharing links is a really clear indicator that, like, something important is going on. So we spent, a lot of time, like, first figuring out what the gap is in the technology that people care about. Like, what was the actual problem that needs to be solved? in early LLM days, it was, OpenAI refusing to finish the prompt or,Anjney Midha [00:30:09]: YeahAlex Atallah [00:30:10]: To, like, complete the task. It was also.Anjney Midha [00:30:13]: Inability to customize models. and so there are communities that, like are just completely blocked on that issue, and those are the communities that are most useful to learn about and dive into and explore.Alex Atallah [00:30:28]: Something that really struck me at that time, - as I was just hearing your talk, I remember noting - you may not remember this, but we - we had these, like working, Zoom calls that we were doing a sprint around for, like this OpenSea integration with Discord. and, we’d, we’d - it was myself, my engineering team. I think you were there. And I remember, Alex, in the middle of one of those calls, just like there was like silence. we were all like, “Oh, yeah, this totally makes sense. Let’s do this.” And then there’s - every, like everybody aligned. And Alex was like, “No, this makes no sense to me.” And everyone’s - I remember going, “What? Like, it works. Like, you click on a link and this, then it bounces you out to, like, OpenSea.” And he was like, “It’s not a good user experience. Yeah, we should not do this.” And I remember going, he was the only one person out of all of us to raise his hand and go, yes, it made sense from a technical implementation perspective. Like, we were bouncing the user out into the, into OpenSea. And so it kinda checked the box of the product manager’s requirements on both sides. But Alex went one step further and was like, “ what would be better, guys? If we just embedded the experience right here inside of Discord so the link opened up as an embedded iframe, and you can just check out right there.”Alex Atallah [00:31:47]: And not one person on the call, and there’s like seven of us who had met, like, week after week.Swyx [00:31:52]: And it’s the guy who doesn’t work for Discord.Alex Atallah [00:31:53]: And it’s the guy who doesn’t work for Discord.Swyx [00:31:55]: Like, technically, you benefit if they bounce.Alex Atallah [00:31:57]: Exactly. And that was, like, adversarial. To keep the user inside of Discord would be adversarial to OpenSea. And yet Alex put that user experience first. And I was like, “That’s special.”Swyx [00:32:08]: Wow.Alex Atallah [00:32:08]: Because it’s very hard to have somebody who’s technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that’s two sides of the flywheel that if you can get spinning, like is often hard to stop. And you just reminded me, like that one was one of those moments where I go, I - I realized I gotta be better at user experience because I should have been the one who came up with that, and I didn’t. And I learned from you. And, I think that went into one of our case studies for the PM training program at Discord.Swyx [00:32:34]: Whoa.Alex Atallah [00:32:36]: I don’t know if it there is Because ofSwyx [00:32:38]: You need an Alex is the conclusion.Alex Atallah [00:32:40]: Yeah. You need an Alex. And this is why I’m not, nobody should be surprised why Stripe decided like they had to buy OpenRouter because it’s a really rare combination of people who understand the machine learning community, the developer experience, and the user experience. And putting all that together has resulted in this extraordinary scale that very few other marketplaces have been able to achieveWindow AI, BYOM, and Finding the Right Form FactorSwyx [00:33:02]: Yeah.Alex Atallah [00:33:02]: Over the last, five years.Swyx [00:33:04]: Yeah. Well, we should talk about the other reasons for acquisitions, whichAlex Atallah [00:33:07]: Yes, we should.Swyx [00:33:07]: You’ve written about. I wanna proceed somewhat chronologically as well. So - there is a point that, one of the questions that, Dave from H of Zero sent in was, when did it - really started to work? And you brought up Mixtral. I don’t know if you wanna bring up that story.Alex Atallah [00:33:22]: Oh, yeah.Swyx [00:33:23]: Which obviously you overlap with, so.Anjney Midha [00:33:26]: Yeah, the MoE was. I don’t know when. there’s no like one moment where I was like, “Oh, this is, officially starting to work.” It wasSwyx [00:33:36]: The moment where you had a Chrome extension, like, really super early on.Anjney Midha [00:33:39]: Oh, yeah. But, well, - yeah. So before OpenRouter, I wanted to, like, explore a bring-your-own-model experiment. And,Swyx [00:33:47]: Which anyone familiar with crypto is like, yeah, Phantom and all these things.Anjney Midha [00:33:50]: Yeah. So it felt like doing a MetaMask analogy for AI would be a fun way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps that were, like, hitting AI - like, hitting an LLM via an API call as there were, like, games just doing it in JavaScript. like there was a, there was a moment in time where it could have been the case that web apps call LLMs through the browser, like through some desktopAlex Atallah [00:34:27]: Yes.Anjney Midha [00:34:27]: Managed app that is controlled by the user. and of course, there are like, I think, many reasons that did not happen. But back when the days were that primordial, I built a Chrome extension called Window AISwyx [00:34:43]: With Plasmo.Anjney Midha [00:34:44]: With Plasmo.Swyx [00:34:45]: I had come across early on, and I was like, “Who’s gonna use this?” You did.Anjney Midha [00:34:49]: Plasmo had a couple, like, I think Phantom was using it. there were some other, like real companies using it.Alex Atallah [00:34:56]: It was like a shim.Swyx [00:34:57]: React for Chrome extension. It compiles to allAnjney Midha [00:35:00]: Yeah.Alex Atallah [00:35:00]: I see.Anjney Midha [00:35:00]: Like Next.js for Chrome extensions.Swyx [00:35:01]: Next.js, Next.js.Alex Atallah [00:35:02]: Okay.Anjney Midha [00:35:03]: And yeah, built Window AI on top of it. The creator of Plasmo, like started contributing code to Window AI, in GitHub, and that turned out to be Louis VicchiAlex Atallah [00:35:15]: Oh, you’Anjney Midha [00:35:15]: Who is the founder of OpenRouter.Alex Atallah [00:35:17]: That’s right. You have told me this is how you met Louis. Yes.Anjney Midha [00:35:19]: Yeah.Alex Atallah [00:35:19]: Okay.Anjney Midha [00:35:20]: So, that allowed users to like configure which model they wanted to use for a web page in their browser, and then, like the app would just call out to that model when it needed to do things. not the right form factor for LLMs, but, it’s like fun experiment. You learn a lot, and like I open sourced it. And the main learning is like, okay, this has to be an API, and it has to look a little bit - like, there has to be more of a developer experience here and more of a discovery experience as well. Like, I don’t know where to use these models, and a little Chrome extension is not gonna help me discover. It’s not enough real estate. I need more space. I need visuals. I need graphs. I need, examples. I need images. I need to, like, I need to be able to, like explore both as a human and as an agent.Crypto, Midjourney, and the Early Generative AI EcosystemAlex Atallah [00:36:10]: Yeah.Anjney Midha [00:36:10]: So that’s how OpenRouter came to be.Alex Atallah [00:36:13]: A meta point that.Alex Atallah [00:36:16]: I think is underappreciated, but Alex is reminding me, is that we were quite lucky that we were so. we were, like, adjacent to the crypto community in those days. Because in hindsight, crypto ended up being like a dress rehearsal for generative models, right? If you think about the Axie experience, Alex is totally right, there were not that many AI apps at the time. And while I was dealing-- my job was to be the head of platform at Discord, which meant to be a general purpose place for communities and friends to create-- for developers to create apps and bots and, other services that could be deployed across Discord. And while 80% of the attention at the time was being spent on crypto, because that’s where all the NFT volume was, there was, like, twenty percent of my time I was spending with a friend, who would get hotbot with me and ask me for. We would play Magic: The Gathering on weekends, and he was working on a little Discord bot that could take a text input and turn it into an image, and it was called Midjourney. YouSwyx [00:37:15]: Is that David?Alex Atallah [00:37:15]: It was David Holz.Alex Atallah [00:37:16]: He was a good friend. And David and I have both been failed ARVR founders, in the before that. And, I remember this. Midjourney was one of the fastest-growing communities we had after Axie Infinity started to peter off. And many of the, like, the abstractions and the infrastructure decisions we made to scale Axie happened just in time because they. Axie did this and then fell off a cliff. And then as Midjourney was taking off, we, like, explicitly decided to help David make the server, the Midjourney server, as the primary place for interaction with the model, because it was very hard for people to understand how to use the model if they couldn’t see other people using it and copy them. And so the single-player Midjourney web app on its own, like midjourney.com, had, like, terrible retention because people would show up, they’d see this empty field. It’s like E 2, and they would type in, like, cat or dog. And it was, like, paralyzing for them to have this blank canvas that they had to fill because they’d never used an AI model before. But instead, in a Discord server, you could see other people using it and riff off of their prompt, and the engagement was off the charts. And so scaling, Midjourney from zero to, like, 10 million monthly actives was a much smoother approach Axie Infinity. And so,Swyx [00:38:29]: Don’t forget the best of four pictures, and you choose one.Alex Atallah [00:38:31]: The best, yeah, and then the other, weSwyx [00:38:32]: Which is the feedback loop.Alex Atallah [00:38:33]: The RLHF feedback loop, which, by the way, separately, like, Tom Brown, David and I used to play Magic: The Gathering on weekends. And so, like, it was one group of friends would hang out, and we’d. Like, these concepts were all being discussed all the time. But, there was.Alex Atallah [00:38:47]: I think there were few of us who bridged both the crypto worlds and the AI worlds. And compared to crypto, where it was - the question was always, what’s the use case, for this technology? There was never any need to ask that for AI because it’s, like, the use case was so visceral. It was like, I can create now anything at - I can imagine. I can write novels, I can code. And the infrastructure that those of us who believed in the distributed systems, like, value of crypto, like the censorship resistance part, found this use case that was explosive. And I think between Midjourney, the, Claude was a Discord bot launch, that we were using internally as an LLM. ElevenLabs had a TTS model that we had on Discord as well. Like, Discord became this petri dish for, like, early apps to innovate. And I don’t think it’s a coincidence that they found a home there before OpenRouter gave the world, like, a public home store or, like, a, storefront. Discord was this, like, almost petri dish storefront that - had, like, piggybacked on the infra we’d built for crypto communities. And then I think Alex was one of the first people to realize, wait a minute, like, these apps need their own home, on the internet. And then OpenRouter, to me, was a continuation of that community’s needs. And of course, there was the crazy distribution that you enabled for a lot of these developers.Why OpenRouter Couldn’t Just Live Inside DiscordSwyx [00:40:07]: So then my question is, how come you were. My perception is OpenRouter is not that Discord-centric, right? You have a Discord.Anjney Midha [00:40:14]: Yeah.Swyx [00:40:14]: And you use it to engage your community, but it’s not like Midjourney where, like, no, that is like the primary way people experience OpenRouter.Anjney Midha [00:40:21]: Yeah, Midjourney, like, it really helps to see visually really quickly how people are using the model and how to prompt it.Swyx [00:40:29]: Yeah.Anjney Midha [00:40:29]: And I think that is partly why the server was so critical. It’s like it is the user experience. It adds a ton.Swyx [00:40:36]: Yes.Anjney Midha [00:40:37]: And you can go the whole mile with just, like, prompting via Midjourney, like, the, via the Midjourney Discord server, getting your images and then sharing them and having fun. For OpenRouter, for LLMs, like, you need a lot of user experience around LLMs to make them, like, really usable.Swyx [00:40:54]: Charge point.Anjney Midha [00:40:55]: And yeah.Anjney Midha [00:40:57]: The, like, seeing the examples of other people is also not as useful because it’s a lot of stuff to read. It takes a long time.Swyx [00:41:03]: Yeah.Anjney Midha [00:41:04]: You need, like, based integration. Not possible to do in a Discord server. You need, Or technic- it’s possible. I shouldn’t say that. It’s just not a great developer experience. you need, like, - you need governance for. At the point where you got based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams. All that stuff needs a lot more than a Discord server can provide. So it’s justSwyx [00:41:30]: YeahAnjney Midha [00:41:30]: It’s not the right.Alex Atallah [00:41:32]: Well, in addition, you’re not wrong, but also there’s the very important distinction that, Midjourney was an end user application.Swyx [00:41:40]: Right.Alex Atallah [00:41:40]: And, that’s why Discord, which has 250 million monthly end consumers, made, it made sense for Discord to be a host for that application experience. What I knew was gonna happen soon after Midjourney found explosive product-market fit, because we. I think when Midjourney launched, from launch to $100 million revenue run rate, it was less than eight months. And shortly thereafter, Stable Diffusion launched. And, all of us used to hang out in the Discord server. There, I think it was the,Swyx [00:42:13]: The Stability Discord?Alex Atallah [00:42:14]: It was theSwyx [00:42:16]: Yeah, LAION.Alex Atallah [00:42:16]: Yeah, the LAION Discord server.Swyx [00:42:17]: The image community that spawned Stable Diffusion.Alex Atallah [00:42:19]: The image community. Yeah. And so when Stable Diffusion came out, I realized- Oh, now other people can build their own Midjourney.Alex Atallah [00:42:27]: Because until then, Midjourney did not have an API, so they were a stack company, right? They were training their own models, and they were deploying them as an application. But if you wanted to build your own Midjourney, there was no API of that quality. and I think E two was still quite primitive. Like, Midjourney had great quality. And then when Stable Diffusion came out, suddenly there was this new person who - there was - this new capability in the world, which is a developer could create their own Midjourney. And that, I think, created the need for something like OpenRouter, because then you need an API to. If you - if you had the creativity of David Holz and you had Stable Diffusion as the model and you wanted to put these things together, how could you do that without having to figure out how to host the weights? And what OpenRouter, - the shape of OpenRouter enabled is that. Right? When you have open model alternatives to closed applications, OpenRouter’s value in the world becomes extraordinary because now any developer can just show up and use theStable Diffusion and the Need for a Model API LayerSwyx [00:43:20]: You just love model diversity.Anjney Midha [00:43:21]: Did you just say the shape of OpenRouter?Alex Atallah [00:43:23]: Oh, no.Anjney Midha [00:43:25]: Were you in cloud? What is this the real Han?Alex Atallah [00:43:26]: I’ve been, I’ve been - I’m, I’m misaligned now. I’ve been overtrained. I’ve been using Cloud way too much, haven’t I?Swyx [00:43:34]: Claude-ish is what people would say.Alex Atallah [00:43:35]: Claude-ish. Oh, God, I gotta untrain myself.Swyx [00:43:38]: Okay. - And I just wanna cap off the Mistral side. my TLDR is there was a Mistral price war, is what they called it, right? Like, round about NeurIPS is twenty-three or twenty-four.Mistral and the Birth of the Inference MarketplaceAnjney Midha [00:43:47]: Yes. DecemberSwyx [00:43:48]: They launched, the Mistral 8x7B, and like the price went down like 80%.Anjney Midha [00:43:54]: Yeah.Swyx [00:43:54]: To me, that’s very positive because it’s like the first, like, real competition to host Mistral. Is there more?Anjney Midha [00:44:01]: Yeah, that was. I’m, like, trying to remember it, all the things that happened. It. Like, we saw that model come out and immediately saw people say that it was the best model in the world.Alex Atallah [00:44:15]: Yes.Anjney Midha [00:44:15]: Like, this was, to my knowledge, the first time an open weights model was called that in real seriousness.Swyx [00:44:22]: It’s hype, right? Is it?Anjney Midha [00:44:25]: It was hype. It was hype. It was also, like, hype from AI influencers at the time. And there were many examples where it was, like, outperforming four. So people really wanted to try it out and see, is this gonna be true for me too? And if so, at what price? And, the, like, inference landscape was really messy.Alex Atallah [00:44:49]: Yes.Anjney Midha [00:44:50]: We cleaned it up. - it allowed, like, providers to compete on price, so we could give you just the best price in one spot. And so it was, I think, the first clear example of, like, a provider marketplace working in a way that adds value to end developers.Alex Atallah [00:45:08]: Sean, you may not remember this, but I think we met for the first time a few days after Mistral came out at NeurIPSAnjney Midha [00:45:15]: Yeah.Alex Atallah [00:45:15]: At a luncheon.Swyx [00:45:16]: Yeah. That’s where I also met BFL as well. Yeah.Alex Atallah [00:45:18]: And Guillaume was there.Swyx [00:45:19]: Yeah.Anjney Midha [00:45:19]: I was at NeurIPS at that time.Alex Atallah [00:45:20]: You were there too. And, we had just announced the Mistral investment, and I remember Guillaume was over there, and I remember turning to Guillaume and asking him, Like, “Is it is all the. Like, how are you feeling after the launch of Mistral and seven B?” And, him in his typical French fashion was like, “ it’s a, it’s an okay model. It’s not that good.” And I was like. It was so, in contrast. But I remember him also saying that part of the reason he felt a lot of people Thought that it was better than four was because of the speed. - it was an MoE model that they had, like, absolutely figured out how to make super efficient. It was on the Pareto frontier. And this is an important thing about LLMs, right? Sometimes when they’re faster, you think they’re smarter, even though, like, if you did, N of, these common, like, evals that are - you do seven tries, and I don’t remember. I think we should go back and figure out what the data says, but I wouldn’t be surprised if it turns out, oh, on an N of seven attempts, four was smarter on evals, but the perception of on, like, or correctness would be smarter or more accurate. But, people, like, from a human preference perspective felt that it was faster because it - or smarter because it’s so fast.Swyx [00:46:36]: Yeah. And most queries do not take that levelAlex Atallah [00:46:39]: Don’t take that. That’s true.Swyx [00:46:40]: Right? So this is the start of humans as routerAlex Atallah [00:46:42]: Yes.Swyx [00:46:42]: Which then eventually becomes OpenRouter as router of like theAlex Atallah [00:46:45]: Oh, that’s interesting way to think about it. Yeah.Swyx [00:46:47]: Like, because humans are the routing mechanism. Like, I will ask the fast model first, and then if, like, oh, not good enough, I’m gonna upgrade manually.Alex Atallah [00:46:52]: Yes.Swyx [00:46:53]: But then he’s gonna auto it.Alex Atallah [00:46:54]: I didn’t, I hadn’t thought of it that way, but that makes sense.Swyx [00:46:57]: Which then there’s, there’s a lot more techniques, like fusion. Fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the early years. one thing that I observe, which you are also an investor in Arena.OpenRouter vs. LM ArenaAlex Atallah [00:47:10]: Right.Swyx [00:47:10]: And we talked about Midjourney having that feedback loop of, A, B, C, D, and choosing that very. being very important. And you understand the flywheel. So how come you didn’t build Arena, and how come Arena didn’t build OpenRouter?Anjney Midha [00:47:23]: Well, Arena started before OpenRouter, right?Swyx [00:47:27]: They had the school projectAnjney Midha [00:47:29]: Yeah, LMSwyx [00:47:29]: And then it became a company.Anjney Midha [00:47:31]: LM Arena, yeah.Swyx [00:47:32]: So, but, and I know you had some Arena experiences, like the up comparison type things.Anjney Midha [00:47:37]: Yeah.Swyx [00:47:37]: But you never really went as hard as Arena did.Swyx [00:47:40]: And,Anjney Midha [00:47:40]: In doing up experiences?Swyx [00:47:42]: Yes. And LM Arena did have a router project based on LM Arena ELOs, which they never commercialized.Anjney Midha [00:47:48]: It’s hard to do a company that does both because one company is taking data and selling it, and the other company really can’t by default. So, I think there is, like, a branding reason that there are two companies here. like, when you set up OpenRouter, there’s no training, there are no prompts, right, aside from what your provider policy set. Like, OpenRou- like, OpenRouter can’t see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it. And so we’re, like, pretty conservative and careful about data policy and security. And privacy. And LM Arena is like, their business model is like oriented around the labs and,Swyx [00:48:34]: Because they give it for free, right? You don’t give it for free to give it for free.Anjney Midha [00:48:37]: Yeah.Anjney Midha [00:48:38]: But we do give some. We like have free endpoints too, but like those free endpoints, we, I think we’re not collecting any prompts. We’re not like monetizing the data unless you, opt into it for some reason.Alex Atallah [00:48:48]: This comparison. you’re not the first person to ask me this, and Alex knows this, but I was the interim, like the founder, like first CEO of Arena for the first five months when, and we were helping Anastasios and Waylin spin out of Berkeley. And, I did invest in that before, OpenRouter, but it was very strange to me the comparisons that outside, folks would make between the two projects because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute because it was there as an eval service. Like the data, so to speak, that they were originally, offering the labs was how do you make the evaluation of models more reliable than like the state of the art at the time, which was like really just finger in the wind.Alex Atallah [00:49:38]: That’s what Anastasios and Waylin’s PhD work was as scientists at Berkeley, was on statistical methodologies for correcting, eval estimates, based on like intrinsic biases and how you collected the data.Swyx [00:49:54]: Yes.Alex Atallah [00:49:54]: AndSwyx [00:49:54]: Style control.Alex Atallah [00:49:55]: Style control and stuff like that. And which is very much like a, hey, how. If you’re a scientist and you’re trying to. the highest expectation customer for Arena was always like a training and, like a researcher at a lab. Whereas the highest expectation customer from my perspective that Alex like really understood and was the mission was to serve was like a developer, right? Who then takes the result of the research and then produces an application that’s deployed to the world. It was a completely different problem and person that these two teams were focused on. And so from the outside in. I don’t know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet, together for OpenRouter, I’d given you a call because we were trying to get a pooled data set together from OpenRouter and from Arena to, create like an open source repository of prompts. these projects were so different in their goals that it was totally normal to me to be like, “Oh, yeah, let’s call Alex and see if he’d want to team up on pooling data,” because they’re so different. We need. We don’t have that data at all. We. Like, we didn’t have API prompts. We didn’t, we didn’t have like what developers want to do with the models, which is very different from what researchers inside a model lab want to do before releasing the model.Swyx [00:51:15]: Yeah.Alex Atallah [00:51:15]: Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000-foot level you could. I guess you could conclude that Arena and OpenRouter are adjacent, but, the roadmaps, the missions and so on at the time at least were like in very different directions.Swyx [00:51:36]: That ideal customer, I get. I totally get that.Alex Atallah [00:51:39]: Yes.Swyx [00:51:39]: As a founder, I want to own everything, right?Alex Atallah [00:51:41]: That’s possible.Swyx [00:51:42]: Like this is clearly an adjacency that I’m like gonna explore that.Anjney Midha [00:51:45]: Own everything meaning like you don’t know what to do yet, so you wanna like make sure you catch PMFocus, Anthropic, and Roads Not TakenAlex Atallah [00:51:51]: No, I think what heAnjney Midha [00:51:52]: As quickly as possible.Alex Atallah [00:51:53]: You want to own the entire infrastructure space, and so you expand to whatever demand you can capture.Swyx [00:51:58]: You want to have a play in each end.Alex Atallah [00:51:59]: Yeah, I think that’s, that’s hard, in reality, because serving multiple customers is difficult.Swyx [00:52:05]: Clearly, this is the one focus, right?Alex Atallah [00:52:08]: Yeah.Anjney Midha [00:52:08]: Yeah. I still think even in the age of AI, like focus is,Alex Atallah [00:52:12]: Is criticalAnjney Midha [00:52:13]: Underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is.Alex Atallah [00:52:22]: One thousand percent.Anjney Midha [00:52:23]: The world can map like, “Oh, I have this issue. Which brand out there is going to help me with that issue? This is the brand that’s known for that focus.”Alex Atallah [00:52:31]: Yes.Anjney Midha [00:52:32]: So like if I want real attention on this issue, like this really matters to me, I should go with the brand that cares the most about it.Alex Atallah [00:52:39]: To underscore Alex’s point about how important focus is, in the early days of Anthropic, it was not easy to. Like people think that the early days of Anthropic were like super easy because they were on their 3 guys who left, but it was very competitive. The company was starting 10 billion dollars behind OpenAI, right? And so to get to the frontier, like the big question was, what do we want to be known for? What’s the mission? And the mission was AGI pair programming. And so to the, exclusion of all kinds of other things that were really shiny at the time, like image models and video models that were getting lots of, momentum, the Anthropic team was like, “We just got to focus on coding.” Like that is the core capability that we’re focused. And today you can see the results, right? It’s a trillion-dollar company within five years. And that focus, I think, like the high. The focus on who your highest expectation customer is and how you exceed their expectations, because exceeding anyone’s expectations is hard, and doing it for multiple like customers is so even more difficult, is part of the reason why OpenRouter succeeded and Anthropic as well.Anjney Midha [00:53:39]: Was the focus on coding that early, though, or did it come later?Alex Atallah [00:53:42]: Literally from day one it was AI pair programming is. Responsibly commercialize an AI pair programmer was the seed memo. That was when I invested, right? We like refined that memo a lot. Well, you got to ask Dario and Tom for permission on that.Alex Atallah [00:53:57]: But it’s an extraordinary piece of writing that they had put together. And AI, commercializing it. Responsibly commercializing an AI pair program was the mission, from day one. And I would say there were maybe like a couple moments in the company’s history where like they did experiments to see if like little detours made sense, like a general chatbot, like Claude.ai when ChatGPT was really taking off. But, at the end of the day, but especially once, they got their like significant training compute online, I think like the. All the main evals at the company, for example, have always Coding evals, long horizon agentic programming. from day one, that was always the plan.Anjney Midha [00:54:34]: Because when, like, Claude Instant came out and Claude 2 cameAlex Atallah [00:54:38]: YesAnjney Midha [00:54:39]: I remember the marketing mostly being focused on pros. Like, thisAlex Atallah [00:54:43]: YeahAnjney Midha [00:54:43]: Could write betterSwyx [00:54:44]: Yeah Long context. It was the first of its kind.Anjney Midha [00:54:47]: Long context,Swyx [00:54:49]: This directly affected me ‘cause I built something on that. Yeah.Alex Atallah [00:54:51]: What did you make?Swyx [00:54:52]: A small developer, which was my Devin before Devin.Alex Atallah [00:54:54]: Oh, yeah. Yes.Anjney Midha [00:54:55]: Yes.Alex Atallah [00:54:55]: Small.Swyx [00:54:56]: Yes. and, so I think, like, there’s, there’s all that really, like, good, like, focus is another thing - That is a question that people do wanna ask. you could have built any other things. Like, and obviously OpenRouter was working. were there other ideas that you wanted to pursue that you turned down? just the paths, roads not taken.Anjney Midha [00:55:16]: We made a couple prototypes for things that we didn’t launch. One was a tuning model as a service.Swyx [00:55:23]: Yeah. Lots of that with OpenPipe and, all those things.Anjney Midha [00:55:25]: But it - It was in a very consumery form factor, where you would give us a YouTube video or two or three. We would then extract all the transcripts from it and try to tune a model to talk like the person in the YouTubeAlex Atallah [00:55:40]: YeahAnjney Midha [00:55:40]: Or the people in the videos that you sent. So, like, a really easy way of creating a tuned model based on, like, some videos that you like.Alex Atallah [00:55:48]: That would be so useful.Anjney Midha [00:55:50]: We,Alex Atallah [00:55:51]: NoAnjney Midha [00:55:51]: We made it too. It wasAlex Atallah [00:55:53]: You don’t think so?Anjney Midha [00:55:54]: It was, itAlex Atallah [00:55:55]: And nobody used it?Anjney Midha [00:55:55]: It - We didn’t like, test it with that many people because the model marketplace was our main focus, and it was, like, growing, and we were building more conviction in it over time.Swyx [00:56:09]: Just, youAlex Atallah [00:56:10]: Yeah. Why,Swyx [00:56:10]: As a creatorAlex Atallah [00:56:11]: Yes. I’m a creator.Swyx [00:56:11]: Have you been pitched many, like, - I have five hundred hours of recorded voice of myself.Alex Atallah [00:56:17]: Right.Swyx [00:56:17]: Make a thing of you, charge access to it. it works for OnlyFans, doesn’t work forAlex Atallah [00:56:23]: I seeSwyx [00:56:23]: As regular people. I think - this is mostly, - It’s just a glorified RAG bot.Alex Atallah [00:56:28]: Right.Swyx [00:56:29]: Whether it’s in the weights or it’s outside the weights, doesn’t really matter. You’re just doing RAG on the videos, and people ultimately always just wanna find the source video, that directly answers it.Alex Atallah [00:56:36]: Oh. my use case was mostly to practice - - with myself ‘cause I often like to see what. Like, the way I practice for a job interview or if I’m hiring a candidate or public speaking or whatever is I wish there was, like, a goodSwyx [00:56:48]: YeahAlex Atallah [00:56:48]: That I could, like, critique ‘cause it’s kinda hard to pull yourself out. I would never get. I would never offer it to other people as a service.Swyx [00:56:54]: I wish there were, like, pick your top five mentors that, then talk to them instead of talking to yourself.Alex Atallah [00:56:57]: That’d be cool too, yeah.Anjney Midha [00:56:58]: That was, that’Swyx [00:56:59]: That’s the creator AI. That’s a replica.Anjney Midha [00:57:01]: And that was the use case we were aiming at.Alex Atallah [00:57:02]: I see.Anjney Midha [00:57:03]: Is like, you wanna create an experienceSwyx [00:57:06]: Like AI Steve Jobs and.Anjney Midha [00:57:07]: And AI Steve Jobs was the initial use case.Alex Atallah [00:57:11]: That’s a,Anjney Midha [00:57:12]: Even though it’s not allowed.Alex Atallah [00:57:14]: That’s a, that’s a common prototype, yeah.Swyx [00:57:15]: Talking about adjacencies, tuning as a service, as part of the router service is something that I would typically think about as well, right? Like, why don’t you do that? ‘Cause if people are running already their inference through you, store everything, log everything, tune to a smaller model that is cheaper, faster, all these things that’s within your control, right? you didn’t do that, but, like, other people would have pitched that in the general state of a infra startup.Anjney Midha [00:57:37]: Yeah. Yeah.Alex Atallah [00:57:37]: I think you were just maybe a little bit early ‘cause today that’s an extraordinarily growing segment. Like, from Mistral, where they do a lot of enterprise deploymentsFine-Tuning as a Service and Infrastructure AdjacenciesAnjney Midha [00:57:44]: RightAlex Atallah [00:57:44]: And stuff and tuning as, custom models for ASML or whatever. And oftenSwyx [00:57:48]: But not as a router. They’re, they’re just like, “I come to you because I like your Mistral models. I want custom Mistral model,” right? It is not, “I want, to run all my OpenAI prompts, - store all my results, and then just move off of OpenAI.” Right? They’re not doing that.Alex Atallah [00:58:01]: As a, as like a way to export off of dependency on a Frontier lab, I have not seen that yet. Yeah.Swyx [00:58:08]: Right.Alex Atallah [00:58:08]: Which was your vision.Swyx [00:58:09]: Is efficient to do.Anjney Midha [00:58:10]: We decided. Really, we, like, leaned into our focus and figured that, like, there aren’t. Like, we just saw the ecosystem develop over time. All these inference providers that do wanna help companies do that, - Like, it makes sense for us to partner with them and to, like, give users lots of choice and to, like, figure out what makes them, what gives them competitive advantages. It’s, it’s a whole new business and there’s, there’s value in being a neutral marketplace that just like, works with those companies.Alex Atallah [00:58:45]: Could you share a little bit, to Sean’s point, like, how you prioritized. What are some ways you prioritize features? ‘Cause you’ve always done it so elegantly that I never. it just happens, and you make all the right decisions that always have product-market fit from the outside looking in. But consistently, you seem to have prioritized, a lot of hit features that worked. And maybe I have a sample set bias or whatever, but Sean’s questionSwyx [00:59:06]: Can you list what you think hit features worked?Alex Atallah [00:59:09]: Oh, the leaderboards.Swyx [00:59:10]: Leaderboard, okay.Alex Atallah [00:59:10]: Yeah. like, from daySwyx [00:59:13]: That’s charting, right? That’s the feedback loop.Alex Atallah [00:59:14]: Charting, BYOK.Swyx [00:59:15]: But, like, he had, like, ins. he had, like, And I think there was a whole thing I wanna get into about, like, completions versusHow OpenRouter Prioritizes ProductAlex Atallah [00:59:22]: Yes.Swyx [00:59:23]: Check completions versus completions. And then also, let’s call it, like, the rise of the reasoning models and how you deal with those, multimodality, all those things, right?Alex Atallah [00:59:31]: Yeah. BYOK.Swyx [00:59:32]: BYOK, yeah.Alex Atallah [00:59:32]: That was a huge one.Anjney Midha [00:59:34]: There’s one I. Like, I think it was in early 2024, very early 2024, we thought it might be interesting to fuse the results of multiple models together, and we launched a prototype called MOM, Mixture of Models, that let you, like, pick a couple models, or we’d pick them for you, and then it would fuse the results together at the end, and it would show you all the intermediate results in this, like, big Kanban looking product.Mixture of Models and Model FusionSwyx [01:00:05]: What does the fusion at the end, another model?Anjney Midha [01:00:07]: Another model. The,Swyx [01:00:08]: The smartest ofAnjney Midha [01:00:09]: The smartestSwyx [01:00:10]: Of the setAnjney Midha [01:00:10]: Of the three, of the set.Swyx [01:00:12]: Okay. So this is like a council idea?Anjney Midha [01:00:13]: Yeah. It was a model. It was like a very early LLM council.Alex Atallah [01:00:16]: This is a agent swarm as, like, they would call it at one of the Frontier Labs, in the early days?Anjney Midha [01:00:23]: Yeah, like some of those ideas are, like, going the right direction, but the devil’s in the details.Swyx [01:00:27]: Yeah.Anjney Midha [01:00:27]: There’s a lot of, like, product refinement needed to make them really work. they take your focus awaySwyx [01:00:34]: RightAnjney Midha [01:00:34]: Whatever else you have going on. And there’s a lot of, like, community building and learning that you need to do. And the technology might be too early. So there are - like, all kinds of reasons they might go wrong. And in our case, the technology was a little too early. In other words, the fused result was a little bitSwyx [01:00:53]: Right. Like a FrankensteinAnjney Midha [01:00:54]: Sometimes the same as the best model that was being used to fuse because the best model was so far ahead of options two and three at the time. over time, the top three or four LLMs have gotten closer together, still neurodivergent, but, like, all capable of inserting, like, pretty interesting ideas. Like, RL has like, expanded the surface area of creativity for machine learning researchers within each lab, and so they can, diversify the reasoning power of different models more effectively. At least that’s my theory forSwyx [01:01:29]: YeahAnjney Midha [01:01:30]: Fusion - it, like, works better than it used to, but early twenty-twenty-four. And, so the technology was a little bit too primitive. The form factor was not right, and so we would have had to go through a couple more iterations. And so we decided to just delete all the code. And, then years later, middle of twenty-twenty-six, or early twenty-twenty-six, we’re like, “Let’s bring it back.” Like, the research is looking kinda promising for fusion. The models now have, like, two, three, four top frontier models that are all really good and, like, I’m, I’m frequently trying to, like, consult multiple models to get the best results. Like, and then I ran a little personal experiment where I was like, “I’m gonna, like, do a, an architecture plan for a code change. I’m gonna give it to all the models. I’m gonna fuse the result, and then I’m gonna ask all the models if the fused result is better than the individual result each model came up with.” And they all said yes, that the fused result was better. And this happened a couple times, and I was like, “Okay, spot check, pretty good. We should, like, benchmark this.” And that’s how we built fusion.Revisiting Fusion as Frontier Models ConvergeSwyx [01:02:40]: Yeah. And it came on your Fable, so you were like, “This is Fable level.”Anjney Midha [01:02:43]: Yeah.Swyx [01:02:44]: Let’s start leading up to this year, which we haven’t gone to this year. can you mark out the main milestones in the journey? I think, it seems like your promise, was, routing. You decided the business model very early.Swyx [01:02:59]: You take a cut. And, like, what are the major milestones that, inflect the growth, right? Like, you’re, you’re growing, like, 9% week on week now? Is this the official number?Anjney Midha [01:03:10]: In terms of token volume, I think that sounds about right, yeah.Swyx [01:03:13]: Yeah. So just, like, can you mark out, like, the brief history of OpenRouter up to, the acquisition? Let’s, let’s call we’re, we’re just, we’re just, talking about, people are, - you have a your birth moment with, the Mistral stuff where people are really competing. You have your state of AI thing where,Anjney Midha [01:03:32]: Yeah.Swyx [01:03:32]: It’s very cute. You have a hundred trillion tokens, ha, ‘cause now you’re doing ten a week, .Anjney Midha [01:03:39]: Yeah. We’re doing ten a day.OpenRouter’s Growth InflectionsSwyx [01:03:41]: Ten a day now?Anjney Midha [01:03:42]: Yeah. More.Swyx [01:03:43]: So yeah, you do this in ten days.Swyx [01:03:45]: Like, what are the major end points there? I just wanna. Like, there’s a smooth curve, but, like, you feel the inflections.Anjney Midha [01:03:51]: A lot of this is oriented around model launches. we had, a huge focus on pros all the way up through May of twenty-twenty-four, because coding was just not there, and no apps were able to build much on top of it. So, a diversity in models, but not a wide diversity and not a wide diversity in use cases. Dream Tavern was one of our top apps at the time. The creator of Dream Tavern now runs product at Cognition, Devon. - Then - In the middle of twenty-twenty-four, we saw Claude 3.5 Sonnet. That came out, incredible leap forward in coding, and we saw the dynamics of, like, apps building on top of us change. we saw a huge surge in volume in, like, users, using OpenRouter. And this is when I think people started to look at the, like, money that they were spending and get a little bit like, “Whoa, what’s going on? I might need to, like, think about, like, more efficient but equivalent models.” And shortly after that, I think it was after Sonnet three five, Mixtral 8x7B came out, and everyone was like, “What? This is the model.” Like, the OpenWeights community delivered. And so it was really good timing from Mistral.Swyx [01:05:17]: All of Anja’s portcos are just helping you out.Alex Atallah [01:05:21]: It takes an ecosystem to grow an OpenRouter?Anjney Midha [01:05:24]: Yeah, that was the. Yeah, it was. It like, it was the, like, this early ecosystem, it was like a swing action where, like, model labs would come up with some frontier innovation. Like, usage would surge. Then users, look at their invoices 30 days later and like, “Whoa, what’s going on here?” And then OpenWeight models would deliver, like, a, like, effective options two, three months later. We saw that happen several times.Swyx [01:05:54]: By the way, oneAnjney Midha [01:05:55]: YeahSwyx [01:05:55]: One thing you also did with the coding agents was that you broke out which are the top coding agents, and they love that. They love that leaderboard. The Klein versus the Rue code versus the what have you.Anjney Midha [01:06:04]: Yeah. Like, Klein was, like, the top of our leaderboard at the time. We, We then, at the end of. And I’ll skip forward a little bit. The end of twenty-twenty-five, there were quite a few coding apps on the leaderboard, but they were all IDs or, terminal-Agents. And at the end of twenty-five, we saw OpenClaw appear. And OpenClaw was, like, particularly interesting because, one, it was like a new form factor that, like, brought in a new type of user, not just a developer, but like a productivity or a, like an internet creator came to AI for the first time. And it also had an interesting architecture where it was, like, calling your chosen model for these heartbeats to see if it was still alive in addition to using the model for real tasks. And the heartbeats are like, they’re kindOpenClaw, Hermes, and the Auto RouterSwyx [01:07:02]: Fréquence.Anjney Midha [01:07:02]: You don’t wanna pay a lot ofSwyx [01:07:03]: Every thirty minutesAnjney Midha [01:07:04]: To do a heartbeat.Swyx [01:07:05]: Yeah.Anjney Midha [01:07:05]: So, the auto router that we provided was really useful to this, like, wide range of users all of a sudden. And so we just saw it rocket exponentially, and then we saw, like OpenClaw just blow up and a couple other, apps lean into that new paradigm and do something similar. Hermes came out and really leaned into things like the auto router and built, like, a really good community and leaned into, like, skill management and making it really easy and effective for people to, like, set their memory in the agentSwyx [01:07:44]: Yeah.Anjney Midha [01:07:44]: And build really good skills.Swyx [01:07:45]: Which another thing you never did, memory skills, sandboxes, all these, like, adjacent things you could have done.Anjney Midha [01:07:52]: Could have, but It’- I think,Swyx [01:07:54]: It’s hard to bet.Anjney Midha [01:07:55]: They’re also - There are things that developer-- that really matter for, like, the developer use cases that were coming out at the time. Like, developers wanted to architect those things.Swyx [01:08:05]: Right.Anjney Midha [01:08:05]: Those were kinda critical to building a good user experience. It’s really-- It was, like, - It’s been hard for companies to find abstractions that work for all developers on the memory layer. It is, it - Yeah, there are some, like Mastra has done a pretty good job, for example. But, like, developers have, like, lots of varied preferences for them. And then we - - the way our leaderboard has changed over time is like a movie of how the AI space has changed over time. If you just like, go to the Wayback Machine and look at the rankings leaderboard and the apps leaderboard over time, it shows you, like, what’s happened in AI over the last couple of years.Swyx [01:08:48]: To me, the coming of age moment was, Andrej Karpathy was like, “I no longer read Local Llama ‘cause, like, I just go to OpenClaw-- OpenRouter’s leaderboard.”Leaderboards as a Map of the AI EcosystemSwyx [01:08:57]: Which I remember that. Yeah. I think he probably, like, said, like, “Sorry, guys, I’m gonna send a bunch of traffic to you.”Swyx [01:09:03]: So I also wanna bring it into the Stripe, thing.Why Stripe Acquired OpenRouterSwyx [01:09:07]: How does that conversation start?Anjney Midha [01:09:09]: We had this longstanding relationship with Stripe, though, from, like, many different projects that we had worked on with them. We invest, a lot of effort in countering abuse,Swyx [01:09:24]: Token fraud.Anjney Midha [01:09:24]: And token fraud.Swyx [01:09:26]: Can you give some numbers just - so people understand?Anjney Midha [01:09:29]: I think I, like, I posted about this. We blocked 10x as much dollar volume last month as the month before. And the types of token fraud are diversifying quite a bit. there are, like, fraudsters going after typical stolen credit cards, but there are also, people trying to resell traffic against the terms of service. There’s, like, hacked accounts. There’s people who just lose - like, their whole company is compromised, and they don’t even realize it, and we help them, like, regain control and detect it. There’- There are accounts that are, like, reselling inference on the side. There’- There are accounts that are dealing with, a, like, an accidental runaway agent, and they don’t realize it. Not a hack, but it’s something that blows up and the company doesn’t want it. And so our trust and safety team, like, works a lot on all of these, like, categories of problems and helps block it and detect it. And so we’ve built these. we have models around them. We - We worked closely with Stripe for a while on this, and I think it’s gonna become a huge problem in the ecosystem. Like, we’re already seeing a lot of companies start to see these fraudsters, like, spread and look for other ways other than OpenRouter to other fraud vectors. And if you’re making a gateway or selling, like, generalized inference, you are a target for fraud. If you’re selling very discreet, like, intelligence products that are, like, doing something pretty specific, but not, like, just reselling inference with some added capability, then you’re way less likely to get these fraudsters. So - I think we’ll see companies also move away from just reselling inference with some like, added capability and move towards like, discreet tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference, like, in a party way.Fraud, Abuse, and the Emerging Token EconomySwyx [01:11:39]: Whoa. Okay. and yeah, obviously you would power that.Anjney Midha [01:11:44]: Right.Swyx [01:11:44]: But you - People pay, for outcomes Or per task?Anjney Midha [01:11:48]: I think people will pay. I think, like, the Datadog pricing page is a good look at, like, the future to come. It’s like companies, like infrastructure companies will, like, charge for different types of events that they’re providing, and there’ll be lots of, like, continuous pricing models that look like that. And of course, there will be, like, if you go down, towards consumer apps, simpler pricing, more subscriptions, fewer events to worry about, and ones that, like, are not. Focus on just adding a markup on top of inference.The Token Economy and Security at ScaleSwyx [01:12:28]: Yeah.Anjney Midha [01:12:28]: Not just because fraud is hard, but also because the pressure from the labs and from - like, good inference providers to, like, do a commit and then bring your inference elsewhere is gonna be very high.Swyx [01:12:44]: Any comments?Alex Atallah [01:12:45]: Two. One, I think Alex has done a very eloquent job of describing something, counterintuitively I knew would be a thing at scale, like four years ago because of Discord. And the particular experience that taught me this was, as we started scaling Midjourney, - one of the primary ways that we used to give away or, like, get people to try Midjourney early on to get to their first ten generations. Because, ten generations - ten images generated was roughly the magic moment activation point we found. Like, once you’d done ten, you were like, “This is extraordinary.” but for that week, so we had a free trial with Midjourney. And one day I woke up, because I was the head of platform and had to monitor, I had all these dashboards, and I had, like, three missed calls from David. And it turns out, like, there had been this flood of new users overnight. And we were like, “This is great.” And he was like, “No, we shut down the free trial.” And I was like, “Why is that?” and he said, “I want you to look at the geolocation IP addresses.” And somebody in China had started to resell Midjourney free, subscriptions with the free trial as a way to, like, you - It was fraud abuse, right?Swyx [01:13:54]: Even for a specialized model like Midjourney.Alex Atallah [01:13:56]: Yeah. And that was an application. So this idea - I think the big picture realization I had back then was, hey, there’s a new type of unit of value that’s being streamed across the internet called a token.Alex Atallah [01:14:11]: And over the next ten years, the entire internet value chain was going to have to deal with the fact that, like, the more valuable tokens got, The more bad actors are gonna go to try to get their hands on those tokens. And anytime you scale something and the payload gets more and more valuable, More bad things, people try to get access to that value. And so it was very obvious to me back then. And so, look, to this day, I don’t think there’s a free turn. Like, I don’t think Midjourney’s ever turned on the free trial since then, because it was really not an easy problem to solve in terms of trust and safety. that’s why I - started teaching the class Security at Scale at Stanford. Like, it was like one of - that and the Anthropic learnings, to me, it was clear that the need for security at scale is gonna be enormous a few years from then. Because if you just do the math, right, think about, like, if we’re. online payments, has started roughly in the eighties and nineties, right, and grew to over a trillion dollars over the next ten years, and we needed to build entirely new payment solutions to deal with online fraud. where we are today is roughly there on tokens, but over the next even five years, we’re expecting the token economy to get to, like, roughly 5 trillion dollars. And over the next ten years, I’d be shocked if we weren’t at 10 trillion dollars of token flow. And so if we were starting to see such aggressive abuse and fraud at subscale, Midjourney, remember Midjourney at this point was, like, less than three $100 million revenue run rate a year.Alex Atallah [01:15:44]: I just realized we were gonna need, like, entirely new, Like, systems to deal with the fraud that was gonna happen for trying to get into the token flow. And so, - I, - I forget the board meeting it was when you brought up that, Stripe wanted to partner up, and it made so much sense to me because Stripe Radar. When I was at Kleiner ten years ago, we invested in Stripe, and the whole pitch that, Patrick and John communicate so eloquently was like, “Hey, unlike traditional payment tools like Braintree that do a day verification, like KYC and AML to get the fraud out of the way, we just bite the fraud cost upfront as customer acquisition cost and - tell a developer, like, just use five lines of code, and we start accepting your payments in five minutes. And what’ll happen is over time, we collect all this data on the developers.”Swyx [01:16:31]: Cloudflare model.Alex Atallah [01:16:32]: Is the Cloudflare model, right? And they did. Five years later, they launched Stripe Radar, and Stripe really today is a security company. That’s the real. People think it’s a payments company. No, the reason. There’s lots of other payments providers today that give you, like, cheaper payments transmission. But the reason Stripe keeps, being the dominant one here and Adyen and Europe is because they have extraordinary fraud detection that they’ve built, - over the years.Swyx [01:16:52]: It’s the same story with Elon and Max LevchinAlex Atallah [01:16:55]: And affirm, yeah.Swyx [01:16:56]: Yeah.Alex Atallah [01:16:57]: So, I think the story shows up over and over again, where every time you have value streamed across the world in large amounts, you need new protection and security infrastructure to fight, to keep the bad guys out and allow the good people to, like, have their transactions happen really fast. And so I think, - this is why - from my perspective, like, the Stripe and OpenRouter story is a security story for the internet ecosystem, for the frontier AI ecosystem. Without a partnership like that, it becomes very hard to defend the quality of experience and the speed and all the good stuff without letting the bad guys get in the way. the second is that, there’s this underappreciated thing about, like, the fact that you need to. Like, - all the bad things that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next ten years.Swyx [01:17:46]: Oof.Alex Atallah [01:17:47]: Right? So think about the, like, recursive scale we’re about to see of bad actors. It’s not just bad human beings, it’s, it’s all the bad agents that are gonna be attacking the token flow. And there’s. It’s very hard if you’re a researcher and at an AI lab to reason about that problem because the only data you have is how agents you’re training are going rogue. But that’s just a fraction of all the bad behavior on the internet that we’re gonna see. And so what you need is defenders, new sheriffs in town, which cowboy hats, that can see all the bad behavior from AI agents across the ecosystem, from different model labs and different trained deployments and different developers, and take all of that data and say, “We’re gonna build a shield for the entire token economy.” Because without that, the amount of fraud we’re gonna see of this 10 trillion dollars in GMV and global GDP growth is, like, a huge percentage of that, I think, is going to be fraud, abuse. And we might never get there if people just don’t trust. Tokens, right? and I don’t think this infrastructure exists. So you have your work cut out for you with, at Stripe, but I don’t think people have realized the scale at which agents, agent, agentic fraud, like bad behavior perpetuated by AI agents is about to hit us like a tsunami.OpenRouter + Stripe: What Changes NextSwyx [01:18:58]: Yeah. there’s a lot to dig into there. I wanna give you the last word. We do have to wrap. what can people expect from OpenRouter and Stripe?Anjney Midha [01:19:07]: I think this is a really good way for us to accelerate market and, to go upmarket more quickly. It’s also, as Ansh eloquently described, this is, there’s a really clear better together story here when it comes to improving trust and safety and making it really easy to, like, accept tokens and let people bring their own inference to your app and to help developers just, like, build on top of inference, going forward. We have a really strong brand with OpenRouter, and we’re keeping the brand. So, like, OpenRouter, like, as a product and the roadmap and the name and the brand, like, is staying the same. And so what, like, you should expect, in the next six months is that most things will be like what we would have done had we been independent, except everything will be moving faster. And that’s like our, term goal. Longer term, hopefully I can comment on it soon, but I can’Closing: Building the Infrastructure for the Token EconomyAnjney Midha [01:20:11]: Now.Swyx [01:20:11]: Okay. Well, we’ll hopefully do a follow-up at some point, but thank you for being so generous with your time, and, congrats on the partnership. this is one of the most beautiful bromances I’ve seen in AI.Alex Atallah [01:20:22]: Just starting out.Swyx [01:20:23]: Starting from StanfordAlex Atallah [01:20:24]: Just starting.Swyx [01:20:24]: To here.Alex Atallah [01:20:24]: Yeah. Lots more to do.Anjney Midha [01:20:26]: Yeah.Alex Atallah [01:20:26]: Lots of sheriff, policing to do of the, ofSwyx [01:20:29]: Yes. The cowboys in town.Alex Atallah [01:20:30]: Of the token economy. We need We need new sheriffs for sure.Swyx [01:20:33]: Yeah. Awesome. Thank you.Anjney Midha [01:20:35]: Thank you. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

Transcript
Discussion (0)
Starting point is 00:00:02 Okay, we are here in Andge's house. It was just where all great startups in San Francisco start. Howdy. And congrats on Cursor, Mistral. God knows what else. You got so much stuff going on. There's a lot going on. Open router is probably the most, has been the most, I would say,
Starting point is 00:00:22 when I'm excited about. Yeah, yeah. And we have Alex, first time on the pod, but you've been in the EIA a few times. I appreciate every time you've shown up for the community. congrats. I just, like, what a journey. When I was looking back at your past post, one of the earliest principles
Starting point is 00:00:37 that I saw you write as a sort of product person is PubSub as a product principle. And I wanted you to maybe explain how you think about what should exist in the world. Yeah, the PubSub piece, which was early 2023. I didn't think about it until we talked like 10 minutes ago is about how there is like a way of thinking about products
Starting point is 00:01:00 as an intersection between subscribing to data and publishing data. And marketplaces are an easy, easy, easy example of this. You have suppliers that are publishing some kind of product to a skew. And the skew is kind of like a pub sub topic that a consumer is subscribing to and just going to like consume whenever they want. And humans consume in a very like discrete ad hoc way. It's not very scalable. you know, all their attention is on the topic when they're buying the thing,
Starting point is 00:01:34 and their attention is nowhere else when that happens. Agents and consumers of inference don't act like that. They're consuming continuously, and they're changing the skews that they consume from all the time. So open router is sort of like a blend between a normal API experience and a marketplace where we create model slug, we have the auto router, we have all kinds of like product skews that you can subscribe to and then you can like continuously add like derive value
Starting point is 00:02:06 and make decisions based on those consumers. Yeah. This is something that was more consensus now, but not consensus when you guys started, which was that there is such a demand for swapping models and changing things out and that people would not use the native SDKs. I guess for each of you, what was your sort of real?
Starting point is 00:02:28 realization moment that this would be it. You've given a talk at EIE about alpaca as like one of your inspiring moments. Alpaca, I can like rehash the alpaca moment for a sec. Like the very beginning at the end of 2022, opening I was the only game in town. There was like opening eye, cohere, and then a smattering of like early attempts at open weight models.
Starting point is 00:02:54 Yeah. When Wama came out in January of 2023, it was like, Well, really exciting. This is really big. It outperforms GPT3 on one or two benchmarks. But you can't chat with it. It wasn't actually an engaging model. But it seemed like someone just needed to fix a couple of things
Starting point is 00:03:11 and do some RLHF on it to get it all the way there. And Alpaca was the first model that I saw that did that. It only took $600 to do. A team at Stanford generated a bunch of synthetic data, fine-tuned Lama, and made alpaca, seven billion parameter model. Maybe it was 13 billion parameters.
Starting point is 00:03:33 And it was so good. Like I was just like on an airplane using it. I, you know, in many cases, like you could not discern a chat GPT versus an alpaca result. And I figured if it was this easy to make a model. One, we have a whole new way of monetizing data for the first time.
Starting point is 00:03:54 You can just like take really valuable data and turn it into a service in $600. And that cost will probably go down over time. When you say, sorry, when you say monetizing your data as what eventually would become an MCPN point or as a training data for a model. Yeah, training data for a model. Like an abstract way of saying like, hey, I have this data. Press it into a model. Like, it makes sense for me in my product.
Starting point is 00:04:19 But like, I could repackage it in the form of a model and sell it. And so it's just a whole new business model. for the economy. It also, of course, provides, like, you know, a way of following what frontier labs are doing, but in a way that, like, a single developer or a small team of developers can roll on their own. And so whenever you have an example of that, like a breakout app that's doing really well, and then some kind of framework for imitating it in your own flavor, you have an
Starting point is 00:04:56 immediate ecosystem of like an immediate ecosystem like should arise because there's just a huge gap between the like decisions that the single company is making and all of the variations in those decisions that like a wider ecosystem can can create themselves and so then you know you need a marketplace to like discover all of those services and all of those products there wasn't any place on the internet that like was like a home base for LLMs in terms of seeing how much they were being used and seeing who was using them and why. The closest would be hugging face. They just started a few years ago before that.
Starting point is 00:05:34 Yeah, and Hugging Face also didn't have the close source models. Yes. And you couldn't use the models at the time. And there wasn't data about who was using that. There are like a bunch of differences between OpenRouter and Hugging Face. And those differences felt really critical to me, especially when I was just trying to learn about LLMs and why people are choosing like these different little ones that are
Starting point is 00:06:01 emerging over time. Got it. And then on no stranger to wanting more model diversity at the time you're a couple of years into your anthropic journey which you've covered in the previous podcast as well. What was your introduction to Alex? Well the introduction was I think 13 years before that. But the open router handshake actually happened right over there
Starting point is 00:06:22 if you remember. which was Alex and I met, I believe it was sophomores now, if I remember, at a Stanford Review. That's what I was eating for the first time. I think so, yeah. Yeah. So Stanford Review was the Libertarian newspaper on campus at Stanford that Peter Thiel started back in the day.
Starting point is 00:06:41 And for whatever reason, you know, Alex and I both showed up to one of the meetings. And I remember the editor-in-chief was a mutual friend of ours. Lisa was really. a really great editor-in-chief. Part of an editor-in-chief's job is to assign responsibilities to people and make sure the work gets done. And I may be misremembering the details,
Starting point is 00:07:03 but I remember wanting to... It was kind of surprising to me that at the time there was no dedicated technology section in the newspaper. Because it's political, right? It's primarily... Originally started as like a... Rights and states and all those things, yeah.
Starting point is 00:07:18 But to take us back in time, you may remember this, but there was this technology kind of legislation that was being debated called the Net Neutrality Act. And net neutrality is like inherently this political concept, right? It's about the regulation of internet broadband access. And so there was a community of us who were kind of technologists, but also debating the politics of the technology.
Starting point is 00:07:44 And I thought the review would be a great place to write about that. And I was working on, I think I had net neutrality article. And I remember proposing, well, maybe sure, started technology kind of section. And Alex was one of the only people who said, yes, that would be cool. And said, I forget whether we ended up writing stuff together, but that's when we first met,
Starting point is 00:08:02 was 2011 or 12. I forget which year it was. Who's one of those? Yeah. Is that old union, if I remember, that's where we used to meet. But, you know, along the way, Alex and I've had a chance to hang out often. And probably the time when we had the most professional,
Starting point is 00:08:21 overlap was when I was running the platform at Discord, and it had become this explosive kind of platform for crypto and NFTs in the middle of the pandemic. Which also, by the way, you were in charge of safety and security as well, right? I was the head of platform, which meant all of the crypto, Dow and NFT launch security debugging fell on me. And they're fishing and the fishing, the social engineering attacks, like a ton of DDoS that we were getting hit by. It's around the time I started teaching security at K.O.
Starting point is 00:08:56 at Stanford, CS-153. And Alex was at OpenC at the time, and I was trying to figure out how we could defend against all these attacks that we were... And at peak, I'd forget, if you remember how much NFT volume was running through... Discord. But it was, like, a meaningful amount of, like,
Starting point is 00:09:12 several billion dollars in NFT volume of GMV, so to speak, were running through the platform, and it was all coming from OpenC. It was these, like, buy, sell, trade... I mean, the servers. The D in D-D-D-N-D Discord. Yes. So that's when I think we had hung out professionally.
Starting point is 00:09:28 But a year after that, OpenAI gave Discord early access to GPT. Sorry, GPT3. No, it was GPD3. Which is the RLed version of GPD3. And that's around the time. We made a Discord bot with OpenAI for internal deployment. And that's when I realized we would need, like, since I was part of the deployment team, What was the use case?
Starting point is 00:09:51 There were two that were, and there's actually a post now called Discord is your place for AI with friends that somebody sent me recently that I wrote and published in 2023, but there were two use cases. One was Clyde, which was like a first-party friend
Starting point is 00:10:08 inside of Discord that could help you set up your Discord server and talk to you about onboarding and get your friends to hang out more. And then there was content moderation. and one of the realizations we had with content moderation was it would refuse to moderate. Like we just refuse our prompts because the RL,
Starting point is 00:10:29 the post training was, we were very early in the post training era. And it would just, our prompts would trigger it, it's like guardrails. And we told Open AI, hey guys, we need access to the weights because if we're going to be doing content moderation
Starting point is 00:10:42 at scale, we have 250 million monthly active users, we need more reliability that the model will do what we needed to. And they said, well, sorry, guys, that's not how this works. We're a close-source company. And so that was my first realization that we needed open models, and the enterprises would need more control over capabilities,
Starting point is 00:10:59 and then ultimately would need some kind of control plane or management system to orchestrate these open models. But there were no good open alternatives until maybe six months later when Lama came out. And six months after that, I led the Susei into Mistral, which was started by Guillaume and the Lama team. And around that time is when I remember hearing what Alex launching OpenRouter and going, these worlds are going to collide. And I don't know when it'll make sense to team up.
Starting point is 00:11:29 But Alex was so early and could see, I think he was totally right about this ecosystem starting with Lama that then needed like an easy layer to manage for it, especially for, I was approaching from the enterprise perspective because I had been that like the, as a VP of platform at Discord, it was my job to ensure that when we, deployed models to like 250 million users, they did what we wanted them to. And that was very hard. Because if you outsourced it to the labs and they controlled the guardrails and their guardrails their safety policies forbid the model from responding to your prompts, that was quite catastrophic. Yeah. But, you know, moderation is the thing that they want to support. And obviously, beyond that, they would work, opening out would work with you, you know, presumably to give you a moderation endpoint, which they offer for free.
Starting point is 00:12:16 It was an interesting use case that they, so they did give us a moderation endpoint. However, as you guys know, every Discord server is like a mini deployment of itself. And so the use case was instead of having human moderators that have to interpret the norms of the community, you just give the, they often like every, you know, subreddit, Discord servers, public ones have their own rules that the user, the users create. Oh, yeah, rerndily in sports. Right. And then humans used to read those norms. and then enforce it every day manually, like observing each message in these communities.
Starting point is 00:12:51 And these communities have like millions of users. So we had a 5,000-plus person team globally on the Discord content moderation team. These are outsourced contractors. They were a really tough job. And so the idea was instead, if you would give the norms of that server to the LLM, then the LLM would do custom moderation for that server.
Starting point is 00:13:09 It's almost like a in-context moderation for that server. and many of those server's norms just violated Open AIS rules. And so that, it was like, we had our own custom e-vales. Each server had its own custom e-valed, but at the time, open AIS e-vales, we were also primitive in our thinking about how to deploy these LLMs.
Starting point is 00:13:30 Often the post-training prompts were super heavy-handed. It said, oh, anything about Harry Potter, anything that has trademarked content, you know, don't refuse. And if it was a fan, Harry Potter fan community, this is a real use case that had a content moderation, the LLM would just refuse. And that was just not precise enough. Another one that we heard was like if someone was trying to write like a detective story
Starting point is 00:13:57 and there's one chapter with a lot of violence, like maybe someone like kills someone, the LLMs were just refused to like help with that part of the story. Yeah. And then they like these would be like, okay, this is not like structurally inherent to LLM's. There must be like some choice out there so that I can like switch to another model when I'm getting like a refusal or a bad result from the main one that I have. And and that like tension also drove me for a marketplace. Yeah. I think that is well accepted now.
Starting point is 00:14:32 What was it like back then when you were raising or, you know, starting this? Did people get it? You know, what was this some of the struggles? Basically, I like I like getting stories out of him about how other VCs don't get it. So like anything you want to talk about now that, let's call it the early journey of open router is done. You can obviously talk about some of the early day stuff. Well, I was going to say that like the biggest objection we got is big model win,
Starting point is 00:15:00 which is every, all the value. Scaling loss. Yeah, scaling laws. And natural network effects are just going to kind of accrue to one company, which will be like, it'll be a Google-style monopoly, just like how Google won the search market by a large, large margin, and you'll just be fighting for scraps at the end, basically.
Starting point is 00:15:24 That was probably the biggest objection we got. You know, it is interesting that Google won the search engine race with such a huge margin. You know, I think, like, had there been more interesting benchmarks or had, like, search engines been, you know, bit, you know, have people, like, seen them a little bit more like LLMs where there are services that you can build companies on top of that might not have been the case. But LLMs don't merely have a user interface. They're also, like, ways of building entirely new businesses.
Starting point is 00:15:56 And, you know, a Google-level monopoly would be, like, the Dutch East India company times, you know, quadrillion in magnitude. Because the whole economy ends up, like, depending on the one monopoly as well. So it didn't seem like, you know, like would be a really crazy outcome if that happened. And it's also less likely because the economics of like creating good competitors are much like much more decentralizedable. Everything Alex said is true. And I came at it from a completely different perspective, which is. Yes, this is why we're here. The scaling laws were never like, in my mind, we're always a feature, not a bug for why open router would be very value.
Starting point is 00:16:38 because I was one of the first investors in Anthropic, and it was obvious to me that other researchers in our friends group, I went to grad school for machine learning, and I just had a lot of friends in the ML community who it was very obvious to us that the bitter lesson holds. And so I was like, oh, like, fantastic, now we have at least two proof points that compute scaling works. It was Open AI and Anthropic.
Starting point is 00:17:01 And by the time, I think we decided to team up on OpenRouter, I had already invested in Mistral and Black Forest Labs, and Luma. So there was multiple companies and teams that I was working with. But you did other modalities, whereas this is literally
Starting point is 00:17:16 different modalities, yes. Exactly. And it was so obvious to me that an ecosystem of different kinds of models were being created and that this whole
Starting point is 00:17:25 narrative of like only one company will dominate like Google was like maybe true but one, I don't believe that but two, there was so much extraordinary innovation
Starting point is 00:17:37 happening across several different research teams, but the shared problem I was noticing across all of them was often, you know, the research teams were fantastic at figuring out how to reason about new capabilities.
Starting point is 00:17:48 They think in terms of capabilities, but never, like, are not developer mindset oriented. Like, what happens after the training is done and the checkpoint comes out? Like, you'd be shocked how, how, like, similar their early pre-training teams
Starting point is 00:18:00 at Open AI, sorry, Anthropic, BFL, Mistral, uh, were in, in, in their, like, default, approach to taking their research out of the lab and kind of scaling their impact, which is often, oh, the checkpoint is done,
Starting point is 00:18:15 put it out as an API, done, and then there'd be crickets. In the case of Claude, the first Cloud checkpoint was actually done a year before they released it internally. And then ChatGPT came out and we decided, okay, yes, it's a good idea to release a Cloud version externally. And they had no plan. Like no plan for how to get developers to actually try,
Starting point is 00:18:37 it out. And so if you go to the Claude 1 blog post, you'll notice they're like three kind of developer examples for users of the API. And one is a Discord bot and the second is Vivian my wife's startup called Juni Learning. And then there was like Notion because these were all friends of like the Anthropic team because that's how like last minute the planning was around, hey, once the model is done training, how do you get it out to the world? There was no distribution platform that understood what developers needed all the key management. provisioning, like simple, like endpoint management, versioning control, like all these things that scientists and researchers go, I mean, that's plumbing. I don't really think about it.
Starting point is 00:19:15 Implementation detail. Right. And instead, Alex came at it from that perspective. And so, you know, it was so obvious to me that like every single lab I was funding would spend like, literally sometimes billions of dollars into training. And then a checkpoint would be done. And there'd be crickets like during early access because they're like, oh, that's right. Like, it's hard to use a checkpoint to make anything. You actually need a whole bunch of plumbing around it to make it usable by a developer. And so by the time, I think we think it was so obvious to me
Starting point is 00:19:44 that a distribution platform like OpenRouter was critical to have in the ecosystem if we wanted there to be competition to Google. Like with Google, DeepMind is done training a new checkpoint and then they push a button and it gets blasted out across all their surfaces from Google Docs to, you know. Everywhere, even if I don't want to.
Starting point is 00:20:02 Everywhere. You want to know about it. Like on Android, like overnight, they can deploy a new checkpoint to like, a billion devices, right? And that invisible infra advantage, distribution advantage, most people don't realize, but until OpenRouter showed up,
Starting point is 00:20:14 you had to think about all of that yourself as a model lab. And it was very daunting. You know, at Anthropic, I think it took more than 12 months to get to our first 10 million in revenue. And in contrast with Black Forest Labs, I remember the early days, you guys had a conversation with the BFL team, and it was so simple for,
Starting point is 00:20:37 Open rudder say, oh, no problem. Like, the day you launch, we can send a million developers to you. You know, that was crazy. That was like a step function change in, like, power. Is that a real number? Million? I think today it's like $4 million. How many developers are on Open Rudder today?
Starting point is 00:20:52 Over 10, but, yeah, over 10 million, but, like, it's hard to, you know. Yeah, I don't know how to, yeah. We do a lot of, like, you know, account de-duping work, but, you know, no one. If you could get a thousand developers, just to put it in context, If you get a thousand developers who actually try the model on day one after you release it and just like do inference and give you feedback, that's a thousand more developers than they knew how to get to on their own. Well, you know, BFL had a reputation, but yes. They had rapid stable diffusion. Yeah.
Starting point is 00:21:23 And with Mistral, I don't know if you guys remember, but the first checkpoint they released was like torrents. It was like torrent weights. Yeah, they just put up a magnet link. There was no API because they didn't, they weren't in for people. you know like saying okay download these weights and you guys go through it he has a story on his side yeah yeah I mean
Starting point is 00:21:42 in addition to the like building a really good developer experience around it the marketing that we do like four different models is totally different and perceived totally differently from the marketing that a model lab does for itself
Starting point is 00:21:56 yes we are like a neutral layer looking at this market like it's a big dark room with all the corners completely obscure to users and users walking into the room and like feeling around and trying to figure out what objects to grab off the tables and like build into their companies. It's an insane way of working. Models are not products where you can just enumerate all their features onto a web page. They're all black boxes, including the open-weight ones. So you need to like shine lights on all
Starting point is 00:22:30 corners of this room so that people can see what makes this model good. And you need the company shining that light to be a neutral third party, which is what we specialized in. So in addition to developer experience, there's also a very important marketing and product packaging component. And a way of like routing and discovering models becomes like critical to your go to market as a provider or a model lab or a server tool and more in the future. And this value to your earlier point about how. many VCs, like, you know, just don't... One of my biggest frustrations is that venture capitalists,
Starting point is 00:23:12 many of them just don't have any operating experience in the field. You know, so unlike a traditional investor who's just maybe come up through the ranks as like of associate working on financial modeling or maybe hasn't been a real operator in the field for like more than 10 years, which is a big part of the industry now, I had just arrived at A16Z, like a year after running the platform.
Starting point is 00:23:33 And so I knew what the challenges were of like building a real great developer experience and actually like being able to create a working piece of software with a model. And there were a few, I won't name names, but they were investors who were looking at OpenRouter and felt at the time when I would compare notes with people
Starting point is 00:23:55 that it was just, I quote unquote, just a marketplace. Yeah, just a thin layer, just a proxy, a wrapper or whatever on other people's APIs. And I was like, you have no idea how strategic the value that OpenRruder is created by being able to orchestrate even three APIs in production. The amount of both engineering work and community design that goes into getting that actually live
Starting point is 00:24:19 and running in production at the scale the OpenRotter team had started just doesn't happen by default. And that was one of the things that stood out to me about Alex from the earliest days. He just understood from a system's perspective like how do you get these fly wheels going? Like that stood out to me with OpenC
Starting point is 00:24:35 when we were working together on the NFD integration at Discord. Like Alex had a level of community, like systems thinking around how you get these flywheels going. Most scientists and machine learning people just don't think, think of it. Like we often think in terms of pre-training, mid-training, post-training.
Starting point is 00:24:52 It's a linear stage. It's a linear pipeline. It's no loop, yeah. Yeah, it wasn't until much later that the modern context feedback loop cycle really got standardized in the industry. But at the time, if you remember, machine learning was like,
Starting point is 00:25:02 like mostly we did a lot of, ML, like when I was in grad school on a laptop. So you just download a dataset, ransom ablations, and you looked at the loss curves, and you're like, great, I made AI. And the idea that you have to, like, deploy those capabilities, collect feedback, trajectories, then, like, put those into a continuous loop.
Starting point is 00:25:21 Like came much, much, much later, and it was very counterintuitive to the science, like the traditionally AI mindset. I do remember doing the investment phase for OpenRouter, I just didn't try and re-educate a bunch of other VCs on why it was not just a marketplace. I was like, you know what? I'm just going to invest.
Starting point is 00:25:41 And I'm going to like take the opportunity to partner with Alex. And if not, no other VCs get it, that's totally fine. Because at the time, it was not obvious, I think, to several of the investors that like OpenRruder was not more than just a wrapper around APS. And that infuriated me. I was like, you know, I don't have time to debate you. I'm just, we're going to invest.
Starting point is 00:25:58 And then I think like a month later, Matt Murphy, he marked it up by 10x. Like our, I think, I forget what the exact post money was and so on. But, you know, to his credit,
Starting point is 00:26:08 Manlo Ventures realized, okay, there's actually much more strategic value here as well. Maybe you didn't hear all these conversations behind the scenes. But that,
Starting point is 00:26:15 that frustrated me a lot. You know, there's a lot of this, like, opining about rappers. And if you're like, oh, an app is just a wrapper on no model, then like,
Starting point is 00:26:24 and, you know, open router is like this wrapper on top of other APIs. This is the most stupid, reductive framework. It's clearly somebody who has no experience deploying products in the field. It's the thing you dismiss other things with. Everyone's a rapper and everything, right?
Starting point is 00:26:38 Like there's some point, some rappers are value. I mean, investors are rappers and LPs, right? Like metric rappers. So, I mean, yeah, it's all rappers down all down to bare metal, I guess. When I started the whole AI engineer, I guess, the coining in 2003, like that was the number one pushback is that this is no value. You should actually just train models. Right.
Starting point is 00:26:57 And, yeah, I mean, obviously this is like, you guys are wanted testaments to the fact that you can actually to build very valuable wrappers, but also very valuable model companies. It's so hard to be, like, the, the day a model launches, the fact that you have an open router endpoint for that model frequently at the top of hacker news on day one,
Starting point is 00:27:19 people don't realize the amount of work that goes into accomplishing that. An open router used, like, that would happen over and over again. And I remember going, people have no idea how hard that is. You know, that's not. Yeah, we've covered. some of the inference engineering that goes behind some of the... Yes.
Starting point is 00:27:35 Base 10 and all those. Well, today you have all those like cool code name things that people guess what OxyAlpha is and all those things. But I guess one of the things that you're teasing is how do you get the initial flywheel going, right? Because today you have your scale and your reputation and all these things. So obviously you drive
Starting point is 00:27:51 immense distribution. But when you're early on, when it's mostly... The bootstrap. Yeah. What is the bootstrap? I mean, to bring it back to early Discord games. I think I think we initially connected with, this is an open C story technically, but we initially connected when you were at Discord
Starting point is 00:28:10 and we talked about the Axi Infinity server. Yes, yes. This server was like the biggest server at the time at Discord. That's right. And you were kind of like constantly bumping up the limits on the server. Oh my God.
Starting point is 00:28:24 For those are normal, like 10% of Philippines was Axi. Was on that server. That's it. It was like a meaningful Contrature to the GDP of the country. It was an NFT, like, crypto game. It was like a Pokemon breeding thing.
Starting point is 00:28:36 Yeah. Similar, yeah. There was battling. There was breeding. And then there was like a marketplace for trading. Play to earn as well. Yeah. Play to earn.
Starting point is 00:28:46 And like the graphics were really cute and fun. And you kind of like, you, you know, you get kind of emotional about your axi that you make. So to like start a community like that, which we had to do many times at OpenC, with basically every early project for us to create a marketplace for it, we need to make sure that the community actually wants it. And it's kind of like building something that people want and going and telling them about it.
Starting point is 00:29:14 Like you can do that on a one-on-one basis. But there's way higher leverage to do that in a community where everyone can talk to you at the same time. So we spent a lot of time like building things that the community really wanted. We did the same thing for OpenRouter. And, you know, like the AXE community was one of, like, a zillion communities we did that with. And Ange, like, saw us doing it. And because you could just see people sharing OpenC links constantly in that Discord. Like, users sharing links is a really clear indicator that, like, something important is going on.
Starting point is 00:29:48 So we spent, you know, a lot of time, like, first figuring out what the gap is in the technology that people, people care about? Like, what was the actual problem that needs to be solved? You know, in early LLM days, it was, you know, Open AI refusing to finish the prompt or like to like complete the task. It was also, you know, inability to customize models. And so there are communities that like are just completely blocked on that issue. And those are the communities that are most useful to sort of learn about and and dive into and explore something that really struck me at that time you as i was just hearing your talk i remember noting how you may not remember this but we were we had these like working uh zoom calls that we were doing a sprint around for like this open c integration
Starting point is 00:30:41 with discord um and you know we'd get it was myself my engineering team i think you were there and i I remember, you know, Alex in the middle of one of those calls, just like, there was like silence. You know, we were all like, oh, yeah, this totally makes sense. Let's do this. And then there's some, everybody aligned. And Alex was like, no, this makes no sense to me. And everyone's like, I remember going, what? Like, it works. Like, you click on a link and this, then it bounces you out to, like, open C. And he was like, it's not a good user experience. Yeah, we should not do this. And I remember going. you know, he was the only one person out of all of us to actually raise his hand and go,
Starting point is 00:31:25 yes, it made sense from a technical implementation perspective, like we were bouncing the user out into the open C. And so it kind of checked the box of the product managers' requirements on both sides. But Alex went one step further. I was like, you know what would be better, guys, if we just embedded the experience right here inside of Discord. So the link opened up as an embedded eye frame
Starting point is 00:31:45 and you can just check out right there. And not one person on the call. like seven of us who had met like, you know, week after week. And it's the guy who doesn't work for Discord. And it's the guy who doesn't work for Discord. Like, technically, you benefit if they bounce. Exactly. And that was like adversarial.
Starting point is 00:32:00 To keep the user inside of Discord would be adversarial to open. And yet Alex put that user experience first. And I was like, that's special. Wow. Because it's very hard to have somebody who's technical like Alex and understands the developer flow, but also understands the best user experience and wants to prioritize that. And that's two sides of the fly wheel that if you can get spinning. is often hard to stop.
Starting point is 00:32:20 And you just reminded me, like that one was one of those moments where I go, I realized I got to be better at user experience because I should have been the one who came up with that and I didn't.
Starting point is 00:32:28 And I learned from you and I think that went into one of our case studies for the PM training program at this room. I don't know if they're going. You need an Alex this is a conclusion.
Starting point is 00:32:40 Yeah, yeah, you need an Alex. And this is why nobody should be surprised why Stripe decided like they had to buy open router because it's a really rare combination of people who understand the machine learning community, the developer experience, and the end user experience,
Starting point is 00:32:55 and putting all that together has a result in this extraordinary scale that very few other marketplaces have been able to achieve over the last five years. Yeah, well, we should talk about the other reasons for acquisitions, which you're afraid about. I want to sort of proceed somewhat chronologically as well.
Starting point is 00:33:11 So there is a point that, you know, one of the questions that Dave from HF0 sent in, was when did you know it really started to work? And you brought up mixed draw. I don't know if you want to bring up that. Oh, yeah. Which obviously you overlap with. Yeah, the MOE was, I don't know when, I mean, there's no like one moment where I was like,
Starting point is 00:33:32 oh, this is, you know, officially starting to work. It was like moments of increasing connection. Like really super early on. Oh, yeah. But, yeah. So before OpenRouter, I wanted to like explore a bring your own model experiment. And... Which anyone familiar with crypto
Starting point is 00:33:48 is like, you know, Phantom and all these things. Yeah, yeah. So it felt like doing a MetaMask analogy for AI would be kind of a fun way of exploring that. And at the time, there were no AI apps. There were probably as many AI apps that were like hitting AI via, like hitting an LLM via an API call as there were like games, just doing it in JavaScript.
Starting point is 00:34:14 basically like there was a there was a moment in time where it could have been the case that web apps call LLMs through the browser like through some kind of desktop managed app that is controlled by the user and of course there are like I think many reasons that that did not happen but back when the when the days were that primordial I built a Chrome extension called Window AI and with Plasma which Plasma. I had come across early on and I was like, who was going to actually use this? You did. Plasma had a couple, like,
Starting point is 00:34:51 I think Phantom was using it. There were some other, like, real companies. There was like a shim basically. React for Chrome extension. It compiles to all these. Yeah, kind of like NextJS for Google engines. Okay.
Starting point is 00:35:03 And, yeah, built Window A on top of it. The creator of Plasma like started contributing code to Window AI in GitHub. And that turned out to be Lewis Vichiex. who is the co-founder of OpenRouter.
Starting point is 00:35:17 You have told me this how you met Lewis. Yes. Okay. So that allowed users to kind of like configure which model they wanted to use for a web page in their browser. And then like the app would just call out to that model wouldn't need to do things. You know, not the right form factor for LLMs. But, you know, it's like fun experiment.
Starting point is 00:35:36 You learn a lot. And like I open sourced it. And, you know, the main learning is like, okay, this has to be an API. and it has to look a little bit, like there has to be more of a developer experience here and more of a discovery experience as well. Like, I don't know where to use these models. And a little Chrome extension is not going to help me discover. It's not enough real estate.
Starting point is 00:35:59 I need more space. I need visuals. I need graphs. I need, you know, examples. I need images. I need to, like, I need to be able to, like, explore both as a human and as an agent. So that's kind of how Open Rider came to be. You know, a meta point that I think is underappreciated, but Alex is reminding me is that we were quite lucky that we were so, we were like adjacent to the crypto community in those days because in hindsight, crypto ended up being kind of like a dress rehearsal for generative models, right?
Starting point is 00:36:31 If you think about the AXE experience, you know, Alex is totally right. They were not that many AI apps at the time. And while I was dealing, you know, my job was to be the head of platform at Discord. which meant to be a general purpose place for communities and friends to create, for developers to create apps and bots and other services that could be deployed across Discord. And while 80% of the attention at the time was being spent on crypto
Starting point is 00:36:57 because that's where all the NFT volume was, there was like 20% of my time I was spending with a friend who would get HotBot with me and asked me, we'd play Magic the Gathering on weekends, and he was working on a little Discord bot that could take a text input and turned it into an image, and it was called Mid Journey. Is that David?
Starting point is 00:37:15 That was David. He was a good friend. And David and I have both then sort of failed AR VR VR founders, you know, in the last, before that. And I remember this, you know, mid-Journey was one of the fastest-growing communities we had after Axi-Infinities started to Peter off. And many of the, like, the abstractions and the infrastructure decisions we made to scale Axi happened just in time because, you know, Axie did this and then fell off a cliff. And then as Mid Journey was taking out,
Starting point is 00:37:43 we explicitly decided to help David make the server, the Mid Journey server is the primary place for interaction with the model because it was very hard for people to understand how to use the model if they couldn't see other people using it and copy them. And so the single player Mid Journey web app on its own, like Midjurney.com, had like terrible retention. Because people would show up,
Starting point is 00:38:05 they'd see this empty field. It's kind of like Dolly too, and they would type in like cat or dog. And it was paralyzing for them to have this blank canvas that they had to fill because they'd never used an AI model before. But instead, in a Discord server, you could see other people using it and riff off of their prompts
Starting point is 00:38:19 and the engagement was off the charts. And so scaling, you know, mid-jury from zero to like 10 million monthly active was a much smoother approach post-AXE infinity. And so... Don't forget the best of four pictures, which is the feedback loop. The RLHF feedback loop,
Starting point is 00:38:35 which, by the way, separately, like Tom Brown, David and I used to play Magic the Gathering on weekends. And so, like, it was one, group of friends would hang out. These concepts were all being discussed all the time. But, you know, there was, I think there were few us who bridged both the crypto worlds and the AI worlds.
Starting point is 00:38:52 And compared to crypto, where the question was always, what's the use case, you know, for this technology? There was never any need to ask that for AI because the use case was so visceral. It was like, I can create now anything I can imagine. I can write novels, I can code. And the infrastructure that those of us who believed in the distributors, systems like value of crypto, like the censorship resistance part, found this use case that was explosive. And I think between Mid Journey, you know, Claude was a Discord bought pre-launch,
Starting point is 00:39:21 you know, that we were using internally as an LLM. 11 Labs had a TTS model that we had on Discord as well. Like, Discord became this Petri dish for like early apps to innovate. And I don't think it's a coincidence that they found a home there before OpenRouter gave the world like a public home store or like a, um, you know, storefront. Discord was this like almost kind of petriish storefront that had kind of like piggybacked on the infra we'd built for crypto communities.
Starting point is 00:39:51 And then I think Alex was one of the first people to realize, wait a minute, like these apps need their own home on the internet. And then OpenRouter, to me, was a continuation of that community's needs. And of course, there was the crazy distribution that you enabled for a lot of these developers. So then my question is, how come you were, My perception is open router is not that discourse-centric, right? You have a Discord.
Starting point is 00:40:13 Yeah. And you use it to engage your community, but it's not like Mid-Journey where, like, no, that is like the primary way people experience Open Router. Yeah, mid-Journey, like, it really helps us see visually, really quickly how people are using the model and how to prompt it. And I think that is partly why the server was so critical. It's like it is the user experience.
Starting point is 00:40:35 It actually adds a ton. Yes. And you can go the whole. with just like prompting via mid-journey, like via the mid-journey Discord server, getting your images and then sharing them and having fun. For OpenRrouter, for LLMs, like, you need a lot of user experience around LLM's,
Starting point is 00:40:53 make them, like, really usable. And, yeah, but, like, seeing the examples of other people is also not as useful, because it's a lot of stuff to read. It takes a long time. You need, like, code-based integration, not possible to do in a Discord server. You need, or technically, it's possible. I shouldn't say that.
Starting point is 00:41:11 It's just not a great developer experience. You need governance for, at the point where you got code-based integration, now you need governance for managing the LLMs that have access to it, the data policies, which teams, all that stuff needs a lot more than a Discord server can provide. So it's just like it's not the right. Well, in addition, you're not wrong, but also there's the very important distinction that, you know, Mid Journey was an end-user application. Right.
Starting point is 00:41:40 And, you know, that's why Discord, which has 250 million monthly end consumers, you know, it made sense for Discord to kind of to be a host for that application experience. What I knew was going to happen soon after Mid-Journey found explosive product market fit because we, I think when Mid-Jurney launched, from launch to 100 million revenue run rate, it was less than eight months. and shortly thereafter stable diffusion launched. And all of us used to hang out
Starting point is 00:42:10 in the Discord server it was the Distability Discord? It was the Lyon The Lyon. Yeah, the Lyon. The image community that spawned stable diffusion
Starting point is 00:42:19 Yeah, and so when stable diffusion came out I realized oh, now other people can build their own Mid Journey because until then MidGernie did not have an API so they were a full stack company right, they were training their own models
Starting point is 00:42:32 and they were deploying them as an application, but if you want to build their own Mid-Journey, there was no API of that quality. And I think Dolly, too, was still quite primitive. Like, Mid-Journey actually had great quality. And then when Stable Defusion came out, suddenly there was this new person who could, there was this new capability in the world, which is a developer, could create their own mid-Journey.
Starting point is 00:42:51 And that, I think, created the need for something like OpenRouter, because then you need an API to, if you had the kind of creativity of David Holes and you had Stable Diffusion as the model, and you wanted to put these things together, how could you do that without having to figure out how to host the weights. And what open router, the shape of open router enabled is that. When you have open models alternatives to closed sort of applications, open router's value in the world becomes extraordinary because now any developer can just show up and use the API. Did you just say the shape of open router? Oh no. Are you in? Oh, that's the real on. I'm misaligned now. I've been overtrained. I've been
Starting point is 00:43:30 over-trained. I've been using cloud way too much, haven't I? Claudish is what people is... Caudish. Oh God, I got untrained myself. Okay, and I just want to cap off the mistrial side. My TLDR is there was a mixed trial price war is what I called it, right? Like roundabout in Europe's 2020 or four.
Starting point is 00:43:47 They launched the mistrial 8 by 7B and like the price went down like 80%. To me that's very positive because it's like the first like real competition to to host Mistral. Is there more? Yeah, that was my choice. trying to remember it, all the things that happened. Like, we saw that
Starting point is 00:44:08 model come out and immediately saw people say that it was the best model in the world. Yes. Like, this was, to my knowledge, the first time an open weight's model was called that in real seriousness.
Starting point is 00:44:22 It's hype, right? Is it, you know? It was hype. It was hype. It was also, like, hype from AI influencers at the time. And there were many examples where it was like outperforming GBT4. So people really wanted to try it out and see if this can be true for me too.
Starting point is 00:44:42 And if so, at what price? And the like inference landscape was really messy. Yes. We cleaned it up. It allowed like providers to compete on price. So we could give you just the best price in one spot. And so it was I think the first clear example of like a provider market place working in a way that adds value to developers.
Starting point is 00:45:08 Sean, you may not remember this, but I think we met for the first time a few days after MixTrol came out at Nureps at a luncheon. Yeah, that's where I also met BFL as well. And Guillaume was there. Yeah, yeah. You were there too. And we had just announced the Mistral investment. And I remember Guillaume was over there.
Starting point is 00:45:26 And I remember turning to Guillaume and asking him, like, is it, is all the, like, how are you feeling after the launch of McStral and 7B? and, you know, him in his typical French fashion was like, I mean, it's an okay model. It's not that good. It was like, it was so, you know, in contrast. But I remember him also saying that part of the reason he felt a lot of people thought that it was better than GPT4 was because of the speed.
Starting point is 00:45:51 You know, it was an MOE model that they had, like, absolutely kind of figured out how to make super efficient. It was on the period of frontier. And this is an important thing about LLMs, right? Sometimes when they're faster, you think they're smarter. even though, like, if you did n of, you know, these common, like, e-vals that are seven, you do seven tries. And I don't actually remember,
Starting point is 00:46:13 I think we should go back and figure out what the data says, but I wouldn't be surprised if it turns out on an end of seven attempts, GPT-4 was smarter on e-viles, but the perception of on, on, like, or correctness would be smarter or more accurate. But, you know, people, like, from a human preference perspective, felt that it was faster because it was smarter because it was so fast.
Starting point is 00:46:36 And actually most queries do not take that level. Don't take that. That's true. So this is the start of humans as router, which then eventually becomes open router as router of like the auto mode. Because humans are the routing mechanism. I will ask the fast model first. And then if like not good enough, I'm going to upgrade manually.
Starting point is 00:46:52 But then he's going to auto it. I didn't thought of it that way. But I mean, that makes sense. Which then there's a lot more techniques. Like fusion is the thing that we should talk about. Before I move on to those things, I just want to close off the sort of early years. One thing that I observe,
Starting point is 00:47:08 which you are also an investor in Arena. Right. And we talked about Mid Journey having that feedback loop of ABCD and choosing that being very important. And you understand the flywheel. So how come you didn't build Arena and how come Arena didn't build OpenRouter? Well, Arena started before OpenRouter, right?
Starting point is 00:47:27 They had the school project. Yeah, Ellum and then they became a company. Yeah. And I know you had some arena experiences, like the heads-up comparison type things. Yeah. But you never really went as hard as Arena did. And doing heads-up experiences? And Elmsis actually did have a router project based on Elamarina Elo's, which they never commercialized.
Starting point is 00:47:48 It's hard to do a company that does both because one company is taking data and selling it. And the other company really can't by default. So, you know, I think there is like a branding. reason that there are two companies here. Like, when you set up OpenRouter, there's no training, there's no prompts, aside from what your provider policies set. Like, OpenRrruder can't see your prompts or completions. If you want to see that as an org, you have to opt into it and enable it.
Starting point is 00:48:22 And so we're, like, pretty conservative and careful about data policy and security and privacy. And El-Marina is like, their business model. is like oriented around the labs and... Because they give for free, right? You don't give it for free, they give it for free. Yeah. I mean, we do give some... We, like, have free endpoints, too.
Starting point is 00:48:41 But, like, those free endpoints, I think we're not collecting any prompts. We're not, like, monetizing the data unless you, you know, opt into it for some reason. This compares, I mean, you're not the first person to ask me this. And Alex knows this, but I, you know, I was the, the first CEO of Arena for the first five months when we were helping Anastasius and Whelan,
Starting point is 00:49:00 kind of spin out of Berkeley. And I did invest in that before OpenRouter, but it was very strange to me the comparisons that outside folks would make between the two projects because the missions were completely different. The founding entity for Arena, we called it the AI Reliability Institute,
Starting point is 00:49:19 because it was actually there as an e-val service. Like the data, so to speak, that they were originally kind of offering the labs was how do you make the... evaluation of models more reliable than kind of like the state of the art at the time, which is like really just finger in the wind. That's kind of what Anastasios and Wayland's PhD work was as scientists at Berkeley was on statistical methodologies for sort of correcting, you know, eval estimates based on like intrinsic biases and how you collected the data.
Starting point is 00:49:54 Style control and stuff like that. Which is very much like a, hey, how, if you're a scientist and trying to kind of, the highest expectation customer for Arena was always like a post-training and like a research at a lab. Whereas the highest expectation customer from my perspective that Alex really understood and the mission was to serve was like a developer, right, who then takes the result of the research and then produces an application that's deployed to the world. It's actually a completely different problem in person that these two teams were focused on.
Starting point is 00:50:31 And so from the outside in, actually, I don't know if you remember this, but I have a distinct memory of a few weeks before we did the term sheet together for OpenRouter. I'd given you a call
Starting point is 00:50:41 because we were trying to get a pooled data set together from OpenRouter and from Arena to create like an open source repository of prompts. I mean, these projects was so kind of different in their goals that it was totally
Starting point is 00:50:56 normal to me to be like, oh yeah, let's call Alex and see if you'd want to team up on pooling data, because they're so different. We actually don't have that kind of data at all. We didn't have API prompts. We didn't have what developers want to do with the models,
Starting point is 00:51:09 which is very different from what researchers inside a model lab want to do before releasing the model. Does that make sense? And so to this day, I think you see that this difference, even though at a 30,000-foot level, I guess you could kind of conclude that arena and open-router are adjacent, but the roadmaps, the missions, and so on at the time, at least,
Starting point is 00:51:32 were like, in very different sort of directions. That ideal customer, I totally get that. As a founder, I want to own everything, right? This is clearly the adjacency, and I'm like going to explore that. Own everything, meaning, like, you don't know what to do yet, so you want to, like, make sure you catch PMF. No, I think what he says is you want to own the entire infrastructure space, and so you kind of expand to whatever demand you can.
Starting point is 00:51:58 Yeah, I think that's hard, you know, in reality, because serving multiple customers is clearly, you know, this is when it's like a focus, right? Yeah, I still think even in the age of AI, like focus is is underrated and critical, not just because you end up with a better product by focusing your humans on it, but also because the world knows what your focus is. The world can map like, oh, I have this issue, which brand out there is going to help me with that issue. This is the brand that's known for that focus. So, like, if I want real attention on this issue, like, this really matters to me, I should go with the brand that cares the most
Starting point is 00:52:38 about it. To underscore Alex's point about how important focus is in the early days of Anthropic, it was not easy to, like, people think that the early days of Anthropic were, like, super easy because they were on the GPT2 guys who left. But it was actually very competitive. The company was starting $10 billion behind Open AI, right? And, so, So to get to the frontier, like the big question was, what do we want to be known for? What's the mission? And the mission was the AGIV repair programming. And so to the exclusion of all kinds of other things that were really shiny at the time,
Starting point is 00:53:10 like image models and video models that were getting lots of, you know, momentum, the Anthropic team was like, we just got to focus on coding. Like that is the core capability that we're focused. And today you can see the results, right? It's a trillion-dollar company within five years. And that focus, I think, like the focus on who your highest expectation cost, and how you exceed their expectation, because exceeding anyone's expectations is hard,
Starting point is 00:53:32 and doing it for multiple customers is so even more difficult is part of the reason why OpenRoders succeeded and Anthropica's as well. Was the focus on coding that early though, or did it come later? Literally from day one, it was AI pair programming, is responsibly commercialized an AI pair programmer was the seed memo. That was when I invested, right? We actually refined that memo a lot. Well, you've got to ask Dario and Tom for permission on that.
Starting point is 00:53:56 that. But it's an extraordinary piece of writing that they had put together. And AI, you know, commercializing an responsibly commercializing AI pair program was the mission, you know, from day one. And I would say there was maybe like a couple moments in the company's history where like they did experiments to kind of see if like little detours made sense like a general chatbot, like Cloud AI when ChatGPT was really taking off. But at the end of the day, but especially once they got their really significant pre-training computer online, I think, like, like, All the main evals at the company, for example, have always been coding e-vials,
Starting point is 00:54:30 long horizon, agentic programming. I mean, from day one, that was always... Because when, like, Claude Instant came out and Cloud 2 came out. Yes. I remember the marketing mostly being focused on pros. Like, this model, right, right, better. Yeah. Long context.
Starting point is 00:54:46 Long context. This directly affected me, because I build something on it. What did you make? A small developer, which is my devon before devian. Oh, yeah. Yes. That's small.
Starting point is 00:54:56 And, you know, so I think there's all that really, like, good, like, focus is another thing. That is a question that people do want to ask. You know, you could have built any other things. Like, and obviously, Open Rado was working, working, working, working. Were there other ideas that you wanted to pursue, that you turned down, you know, just the paths, roads not taken? We made a couple prototypes for things that we didn't launch. One was a fine-tuning model as a service. Lots of that, open pipe and all those things.
Starting point is 00:55:25 But it kind of was in a very consumery form factor where you would give us a YouTube video or two or three, we would then extract all the transcripts from it and try to fine tune a model to talk like the person in the YouTube video or the people in the videos that you sent. So like a really, really easy way of creating a fine-tuned model based on like some kind of videos that you like. That would be so useful.
Starting point is 00:55:51 We made it too. You don't think so? It was... Nobody used it. It was, we didn't actually like test it with that many people because the model marketplace was our main focus. And it was like growing and we were building more conviction in it over time. Just as a creator, I've been pitched many like, I have 500 hours of recorded voice of myself. Make a thing of you charge access to it.
Starting point is 00:56:20 Works for only fans. Doesn't work for us as regular people. I think mostly it's just a glorified rag bot. Whether it's in the weights or outside the weights, it doesn't really matter. You're just doing rag on the videos. And people ultimately always just want to find the source video that directly answers it. My use case was mostly to practice with myself because I often like to see what, like, the way I practice for a job interview or like I'm hiring a candidate or public speaking or whatever is I wish there was like a good mini-me that I could like critique because it's kind of hard to pull yourself out. I would never get, I would never offer it to other people.
Starting point is 00:56:54 Like, pick your top five mentors that didn't talk to them instead of talking to yourself. That would be cool too, yeah. That was the character AI, that's a replica. And that was the use case we were aiming at. I see. It's like you want to create an experience. Like AI Steve Jobs and AI Steve Jobs was the initial use case. That's a, even though it's not allowed.
Starting point is 00:57:13 That's a common prototype. Talk about adjacencies. Fine-tuning as a service as part of the router service is something that I would typically think about as well. Right? Like, why don't you do that? Because if people are running already, the inference through you, store everything, logger thing, fine-tuned to a smaller model that is cheaper, faster,
Starting point is 00:57:29 all these things that's within your control, right? You didn't do that, but other people would have pitched that in the general state of an infrastructure startup. Yeah. I think you were just maybe a little bit early because today that's an extraordinarily fast-growing segment, like, you know, from Mistral, where they do a lot of enterprise deployment.
Starting point is 00:57:44 It's often fine-tuning as custom models for ASML or whatever. But not as a router. They're just like, I come to you because I like your Mistral models. I want custom mistral models. It is not I want to run all my opening eye prompts, store all my results, and then just move off of opening eye. They're not doing that. As like a way to export off of dependency on a frontier lab,
Starting point is 00:58:06 I have not seen that yet. Which was your kind of your... I mean, we decided, really we leaned into our focus and figured that we just saw the ecosystem develop over time, all these inference providers that do want to help companies do that, it makes sense for us to partner with them and to give users lots of choice and to figure out what makes them,
Starting point is 00:58:34 what gives them competitive advantages. It's a whole new business, basically. And there's value in being a neutral marketplace that just kind of works with those companies. Could you share a little bit to Sean's point, like how you prioritized... What are some ways you prioritize features? Because you've always done it so elegant
Starting point is 00:58:53 that never, you know, it just happens and you make all the right decisions that always have product market fit from the outside looking in, but consistently you seem to have prioritized, you know, a lot of hit features that worked and maybe I have samples at bias or whatever, but Sean's question. Can you list what you think hit features worked? Oh, the leaderboard.
Starting point is 00:59:10 Like, leaderboard, okay. Yeah, you know, like from day one, it's charting. That's the feedback. Charting. B.Y. Okay. But like, he had like, plug-ins. You know, he had like, and I think there was a whole thing I want to get into about like completions versus, yes. Check notifications with the completions.
Starting point is 00:59:24 And then also, let's call it, like, the rise of reasoning models and how you deal with those. Multimodality, all those things. D. YOK. That was a huge one. There was one, like, I think it was in early 2024. Very early 20204, we thought it might be interesting to fuse the results of multiple models together. And we launched a prototype called Mom, mixture of models. let you, like, pick a couple models.
Starting point is 00:59:55 We'd pick them for you. And then it would fuse the results together at the end. And it would show you all the intermediate results in this, like, big con bond board-looking product. What does the fusion at the end? Another model? Another model. The smartest of the set. Of the three, of the set.
Starting point is 01:00:11 So this is like a council idea? It was a model. It was like a very early LLM council. This is a multi-agent swarm, as like, they would call it at one of the frontier labs in the early days, you know? Yeah, like some of those ideas are like going the right direction, but the devil's in the details. There's a lot of like product refinement needed to make them really work. They take your focus away from whatever else you have going on. And there's a lot of like community building and learning that you need to do.
Starting point is 01:00:40 And the technology might be too early. So there are like all kinds of reasons they might go wrong. And in our case, the technology was a little too early. In other words, the fused result was a little bit worse. Sometimes the same as the best model that was being used to fuse. Because the best model was so far ahead of options two and three at the time, you know, over time, the top three or four LLMs have gotten closer together, still narrow divergent, but like all capable of inserting like pretty interesting ideas.
Starting point is 01:01:14 Like RL has basically like expanded the surface area of creativity for machine learning researchers within each lab. and so they can diversify the reasoning power of different models more effectively. At least that's my theory for why Fusion works better than it used to early 2024. So the technology was a little bit too primitive. The form factor was not right. And so we would have had to go through a couple more iterations. And so we decided to just delete all the count. And then years later, middle of 2026, or early 20,
Starting point is 01:01:52 26, we're like, let's bring it back. Like, the research is looking kind of promising for fusion. The models now have, like, two, three, four top frontier models that are all really good. And, like, I'm frequently trying to, like, consult multiple models to get the best results. And then I, you know, I ran a little personal experiment where I was like, I'm going to, like, do an architecture plan for a code change. I'm going to give it to all the models. I'm going to fuse the result, and I'm going to ask all the models if the fused result is better
Starting point is 01:02:25 than the individual result each model came up with. And they all said yes, that the fused result was better. And this happened a couple times, and I was like, okay, spot check, pretty good. We should like benchmark this, and that's how we built fusion.
Starting point is 01:02:40 Yeah, and it came on your fable, so you were like, this is fable level. Yeah, yeah. Let's start leading up to this year, should we haven't gone to this year. You know, can you mark out the main milestone, in the journey. I think it seems like your promise was, you know, routing. You decided the business model very early. You take a cut. And like, you know, what are the major milestones that
Starting point is 01:03:03 inflect the growth, right? Like, you're growing like 9% week on week now? Is that the official number? In terms of the token volume? I think that sounds about right. Yeah. Yeah. So just like, can you, can you mark out like the sort of brief history of open router up to, up to, you know, the acquisition. Let's call it. We're just talking about, you know, people are, have, you have sort of your birth moment with the Michelle stuff where people are really competing. You have your state of AI thing.
Starting point is 01:03:32 It's very cute. You have 100 trillion tokens. Ha, ha, ha. Because now you're doing 10 a week. You know. We're doing 10 a day. 10 a day now? Yeah.
Starting point is 01:03:42 More. So, yeah, you do this in 10 days. Like, what are the major points there? You know, I just want to, like, this is a small. smooth curve, but you feel the infections. A lot of this is kind of oriented around model launches. We had a huge focus on pros all the way up through May of 2024, because coding was just not there, and no apps were able to build much on top of it.
Starting point is 01:04:11 So, you know, a diversity in models, but not a wide diversity, and not a wide diversity, and not a wide diversity in use cases. Dream Tavern was one of our top apps at the time. The creator of Dream Tavern now runs product at Cognition, Devon. Then we, in the middle of 2024, we saw Claude Sonnet 3.5. That came out, incredible leap forward encoding. And we saw the dynamics of like apps building on top of us change. We saw a huge surge in Volunt.
Starting point is 01:04:47 in like users using open router. And this is when I think people started to look at the like money that they were spending and get a little bit like, whoa, what's going on? I might need to like think about like more cost efficient but equivalent models. And shortly after that, I think it was after Sondon 3-5, Mix, Mixtrol 8X7B came out. And everyone was like, whoa, this is the model. Like the open weights community delivered. And so it was really good timing from Mistral.
Starting point is 01:05:17 Basically, all of Anjid's portclius are just helping you. It takes an ecosystem to grow an open router, you know? Yeah, that was the, yeah, it was, like, this early ecosystem, it was like a swing action where, like, model labs would come up with some sort of front-tier innovation, like usage would surge, then users, you know, look at their invoices 30 days later, and they're like, whoa, what's going on here? And then open weight models would deliver like a cost effective options to three months later. We saw that happen several times. One thing you also did with the coding agents was that you broke out which are the top coding agents.
Starting point is 01:05:59 And they love that. They love that leaderboard. The client versus the root code versus the what have you. Yeah. Yeah. Like Klein was like the top of our leaderboard at the time. We then at the end of, and all of it was. skip for it a little bit. At the end of 2025,
Starting point is 01:06:16 there were quite a few coding apps on the leaderboard, but they were all IDs or terminal-based agents. And at the end of 2025, we saw OpenClaw appear. And OpenClaw was like particularly interesting
Starting point is 01:06:32 because, one, it was like a new form factor that like brought in a new type of user, not just a developer, but like a productivity or sort of an internet-creative came to AI for the first time. And it also had an interesting architecture
Starting point is 01:06:52 where it was like calling your chosen model for these heartbeats to see if it was still alive in addition to actually using the model for real task. And the heartbeats are like, they're kind of, you don't want to pay a lot of money to a heartbeat. So the auto router that we provided was really, really useful to this wide range of users all of a sudden.
Starting point is 01:07:13 And so we just saw it rocket exponentially. And then we saw, you know, like OpenClaughts blow up and a couple other apps lean into that new paradigm and do something similar. Hermes came out and really leaned into things like the auto router and built like a really good community and leaned into like basically skill management and making it really easy and effective for people like set their memory. in the agent and build really good skills. Which another thing you never did, memory, skills, sandboxes, all these adjacent things. You could have done. Could have, but like, it's, I think like,
Starting point is 01:07:54 it's hard to bet. They're also very, there are things that developer, that really matter for, like, the developer use cases that were coming out at the time. Like, developers wanted to architect those things. Right. Those were kind of critical to building a good user experience. It's really, it was, like, hard,
Starting point is 01:08:09 it's been hard for companies to find abstractions that work for all developers on the memory layer. It is, it is, you know, there are some. Like, Mastra has done a pretty good job, for example. But, like, developers have, like, lots of very preferences for them. And then we saw, you know, the way our leaderboard has changed over time is kind of like a movie of how the AI space has changed over time. If you just sort of, like, go to the way back machine and look at the rankings leaderboard and the apps leaderboard. and the apps leaderboard over time,
Starting point is 01:08:44 it sort of shows you like what's happened in AI over the last couple of years. To me, the coming of age moment was, Andre Carpathie was like, I no longer read local llama because I just go to open router's leaderboard. Which I remember that. I think you probably said like, sorry guys.
Starting point is 01:09:00 I'm going to send a bunch of traffic to you. So I was going to bring it into the Stripe thing. How does that kind of conversation start? We had this longstanding relationship with Stripe, though, from, you know, like many different projects that we had worked on with them. We invest, you know, a lot of effort in countering abuse. token fraud? And token fraud.
Starting point is 01:09:26 Can you give some numbers just to so people understand? I think I, like, I posted about this. We, we block 10x as much dollar volume last month as the month before. and the types of token fraud are diversifying quite a bit. You know, there are like fraudsters going after typical stolen credit cards, but they're also, you know, people trying to resell traffic against the terms of service. There's like hacked accounts. There's people who just lose, you know, like their whole company is compromised and they don't even realize it.
Starting point is 01:10:01 And we help them like regain control and detect it. there are accounts that are like reselling inference on the side. There are accounts that are dealing with, you know, like an accidental runaway agent, and they don't realize it. Not a hack, but it's something that blows up and the company doesn't want it. And so our trust and safety team, like, works a lot on all of these, like, categories of problems and helps block it and detect it. And so we've built these, you know, we have models around them.
Starting point is 01:10:38 We have, we worked closely with Stripe for a while on this. And I think it's going to become a huge problem in the ecosystem. Like, we're already seeing a lot of companies start to see these fraudsters, like spread and look for other ways, other than OpenRouter, to other fraud vectors. and if you're making a gateway or selling like generalized inference, you are a target for fraud. If you're selling very discreet, like intelligence products, intelligence products that are like doing something pretty specific, but not like, you know, just reselling inference with some added capability,
Starting point is 01:11:18 then you're way less likely to get these fraudsters. So I think we'll see companies also move away from just reselling inference with some sort of like added capability and move towards sort of like discrete tasks and charging for those tasks and charging for those enhancements and letting people bring their own inference like in a third party way.
Starting point is 01:11:39 Whoa. Okay. And yeah, obviously you would power that. But people pay for outcomes or per task. I think people will pay, you know, I think like the data dog pricing page is a good look at like the future. to come. It's like companies, like infrastructure companies will like charge for different types of events that they're providing. And there'll be lots of like continuous pricing models
Starting point is 01:12:09 that look like that. And of course, there will be like if you go down towards consumer apps, you know, simpler pricing, more subscriptions, you know, fewer events to worry about. and ones that are not focused on just adding a markup on top of inference. Not just because fraud is hard, but also because the pressure from the labs and from good inference providers to do a commit and then bring your inference elsewhere is going to be very high. Any comments?
Starting point is 01:12:45 Two. One, I think Alex has done a very eloquent job of describing something. you know, counterintuitively I knew would be a thing at scale like four years ago because of discord and the particular experience that taught me this was, you know, as we started scaling mid-jurney, you know, one of the primary ways that we used to give away or like get people to try mid-jurney early on
Starting point is 01:13:08 to get the first 10 generations. Because, you know, 10 generations of, 10 images generated was roughly the magic moment activation point we found. Like once you're done 10, you were like, this is extraordinary. But for that, so we had a free trial with Mid-Journey. And one day I woke up because they had a platform and had to monitor, I had all these dashboards.
Starting point is 01:13:28 You know, I had like three missed calls from David. And it turns out like there had been this flood of new users overnight. And we were like, this is great. And he was like, no, actually we shut down the free trial. And I was like, why is that? And he said, I and took at the geolocation IP addresses. And basically somebody in China had started to resell Mid-Journey, you know, subscriptions with the free trial as a way to like basically, you know, it was fraud abuse, right?
Starting point is 01:13:54 And even for a specialized model like mid-jury. Yeah, and that was actually an application. So this idea that I think the big picture of realization I had back then was, hey, there's a new type of unit of value that's being streamed across the internet called a token. And over the next 10 years, the entire internet value chain was going to have to deal with the fact that like, the more valuable tokens got, the more bad actors are going to try to get their hands on those tokens. And anytime you scale something
Starting point is 01:14:28 and the payload gets more and more valuable, more bad things, people try to get access to that value. And so it was very obvious to me back then. And so, look, to this day, I don't think there's a free turn. Like, I don't think the journey's ever actually turned on the free trial since then
Starting point is 01:14:44 because it was really not an easy problem to solve in terms of trust and safety. And that's why I started teaching the class security at scale at Stanford. Like one of the that and the anthropic learnings, to me, it was clear that the need for security at scale was going to be enormous a few years from then. Because if you just do the math, right, think about if where, you know, online payments, you know, started roughly in the 80s and 90s, right, and grew to over a trillion dollars over the next 10 years.
Starting point is 01:15:13 And we needed to build entirely new payment solutions to deal with online fraud. where we are today is roughly there on tokens, but over the next, even five years, we're expecting the token economy to get to like roughly $5 trillion. And over the next 10 years, I'd be shocked if we want to $10 trillion of token flow. And so if we were starting to see such aggressive abuse and fraud at sub-scale, mid-jurne, remember,
Starting point is 01:15:39 mid-jurne at this point was like less than 300 million revenue, run rate a year, I just realized we were gonna need, like, entirely new, like systems to deal with the fraud that was going to happen for trying to get into the token flow. And so my, I forget the board meeting it was when you brought up that, you know, Stripe wanted to partner up and it made so much sense to me because Stripe radar,
Starting point is 01:16:02 when I was a Kleiner 10 years ago, we invested in Stripe. And the whole pitch that, you know, Patrick and John communicate so eloquently was like, hey, unlike traditional payment tools like Braintree that do a seven-day verification, like K-YC and AML to get the fraud out of the we actually just bite the fraud cost up front as customer acquisition cost and tell a developer, like just use five lines of code and we start accepting your payments in five minutes. And what will happen is over time we'll collect all this data on the developers.
Starting point is 01:16:31 Cloudflare model. Is the Cloudflare model, right? And they did. Five years later, they launched Stripe Radar. And Stripe really today is a security company. That's the real. People think it's a payments company. No, the reason, there's lots of other payments providers today that give you like cheaper payments transmission.
Starting point is 01:16:44 But the reason Stripe keeps, you know, being the dominant. one here in Adian and Europe is because they have extraordinary fraud detection that they've built over the years. This is the same story with Elon and Max Lefchin and... And a firm, yeah. You know, I think the story shows up over and over again where every time you have value streamed across the world in large amounts, you need new protection and security infrastructure to fight to keep the bad guys out and allow the good people to have their transactions happen
Starting point is 01:17:10 really fast. And so I think, you know, this is why, from my perspective, like the stripe and open Router story is a security story for the internet ecosystem, for the frontier AI ecosystem without a partnership like that. It becomes very hard to defend the quality of experience and the speed and all the good stuff without letting the bad guys get in the way. The second is that, you know, there's this underappreciated thing about, like, the fact that you need to, like, these, all the bad things that Alex described as being perpetuated by humans right now is going to be perpetuated by AI agents over the next 10 years.
Starting point is 01:17:46 Right? So think about the recursive scale we're about to see of bad actors. It's not just bad human beings. It's all the bad agents that are going to be attacking the token flow. And it's very hard if you're a researcher and an AI lab to reason about that problem because the only data you have is how agents your training are going rogue. But that's just a fraction of all the bad behavior on the internet that we're going to see. And so what you need is defenders, new sheriff sent down, which in kind of. cowboy hands, that can see all the bad behavior from AI agents across the ecosystem from different model labs and different post-trained deployments and different developers and take all of that data and say we're going to build a shield for the entire token economy. Because without that, you know, the amount of fraud we're going to see of this $10 trillion in GMV and global GDP growth is like a huge percentage of that, I think is going to be fraud, abuse. And we might never get there if people just don't trust tokens, right?
Starting point is 01:18:45 and I don't think this infrastructure exists. So you have your work cut out for you at Stripe, but I don't think people have realized the scale at which agentic fraud, like bad behavior perpetuated by AI agents, is about to hit us like a tsunami. Yeah, I mean, there's a lot to dig into there. I want to give you the last word.
Starting point is 01:19:02 We do have to wrap. What can people expect from OpenRouter and Stripe? I mean, I think this is a really good way for us to accelerate go-to-market and to go-up market more quickly. It's also, as Ange eloquently described, there's a really clear, better together story here when it comes to improving trust and safety
Starting point is 01:19:24 and making it really easy to accept tokens and let people bring their own inference to your app and to help developers just build on top of inference going forward. We have a really strong brand with OpenRouter, and we're keeping the brand. So like open router, like as a product and the roadmap and the name and the brand, like, you know, is staying the same. And so what, like, you should expect, you know, in the next six months is that most things will be like what we would have done had we been independent, except everything will be moving faster. And that's kind of like our, you know, near term goal, longer term.
Starting point is 01:20:06 Hopefully I can comment on it soon, but I can now. Okay. Well, we'll hopefully do a follow-up at some point. But thank you for being so generous for your time. And congrats on the partnership. I mean, this is one of the most beautiful bromances I've seen in AI. Just starting out. Starting from Stanford to here.
Starting point is 01:20:25 Lots more to do. Lots of sheriff policing to do of the token economy. We need new sheriffs for sure. Yeah. Awesome. Thank you. Thank you.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.