Latent Space: The AI Engineer Podcast - State of the Art: Training >70B LLMs on 10,000 H100 clusters
Episode Date: June 25, 2024It’s return guest season here at Latent Space! We last talked to Kanjun in October and Jonathan in May (and December post Databricks acquisition): Imbue and Databricks are back for a rare treat: a d...ouble-header interview talking about DBRX from Databricks and Imbue 70B, a new internal LLM that “outperforms GPT-4o” zero-shot on a range of reasoning and coding-related benchmarks and datasets, while using 7x less data than Llama 3 70B.While Imbue, being an agents company rather than a model provider, are not releasing their models today, they are releasing almost everything else: * Cleaned-up and extended versions of 11 of the most popular NLP reasoning benchmarks* An entirely new code-focused reasoning benchmark* A fine-tuned 70B model, built with Meta Llama 3, to identify ambiguity* A new dataset of 450,000 human judgments about ambiguity* Infrastructure scripts for bringing a cluster from bare metal to robust, high performance training* Our cost-aware hyperparameter optimizer, CARBS, which automatically and systematically fine-tunes all hyperparameters to derive optimum performance for models of any sizeAs well as EXTREMELY detailed posts on the infrastructure needs, hyperparameter search, and clean versions of the sorry state of industry standard benchmarks. This means for the FIRST TIME (perhaps since Meta’s OPT-175B in 2022?) you have this level of educational detail into the hardware and ML nitty gritty of training extremely large LLMs, and if you are in fact training LLMs of this scale you now have evals, optimizers, scripts, and human data/benchmarks you can use to move the industry forward together with Imbue.We are busy running the sold-out AI Engineer World’s Fair today, and so are unable to do our usual quality writeup, however, please enjoy our show notes and the excellent conversation! Thanks also to Kanjun, Ashley, Tom and the rest of team Imbue for setting up this interview behind the scenes.Video podTimestamps* [00:00:00] Introduction and catch up with guests* [00:01:55] Databricks' text to image model release* [00:03:46] Details about the DBRX model* [00:05:26] Imbue's infrastructure, evaluation, and hyperparameter optimizer releases* [00:09:18] Challenges of training foundation models and getting infrastructure to work* [00:12:03] Details of Imbue's cluster setup* [00:18:53] Process of bringing machines online and common failures* [00:22:52] Health checks and monitoring for the cluster* [00:25:06] Typical timelines and team composition for setting up a cluster* [00:27:24] Monitoring GPU utilization and performance* [00:29:39] Open source tools and libraries used* [00:32:33] Reproducibility and portability of cluster setup* [00:35:57] Infrastructure changes needed for different model architectures* [00:40:49] Imbue's focus on text-only models for coding and reasoning* [00:42:26] CARBS hyperparameter tuner and cost-aware optimization* [00:51:01] Emergence and CARBS* [00:53:18] Evaluation datasets and reproducing them with high quality* [00:58:40] Challenges of evaluating on more realistic tasks* [01:06:01] Abstract reasoning benchmarks like ARC* [01:10:13] Long context evaluation and needle-in-a-haystack tasks* [01:13:50] Function calling and tool use evaluation* [01:19:19] Imbue's future plans for coding and reasoning applications* [01:20:14] Databricks' future plans for useful applications and upcoming blog postsTranscriptSWYX [00:00:00]: Welcome to the Latent Space Podcast, another super special edition. Today, we have sort of like a two-header. John Frankel from Mosaic Databricks, or Databricks Mosaic, and Josh Albrecht from MBU. Welcome.JOSH [00:00:12]: Hey, glad to be here.SWYX [00:00:14]: Thank you for having us. Hey, so both of you are kind of past guests. Jonathan, you were actually one of the most popular episodes from last year talking about MPT7B. Remember the days when we trained large models and there was 7B?JONATHAN [00:00:30]: Yeah, back when reproducing LLAMA1-7B was considered a huge accomplishment for the field. Those are the good old days. I miss that.SWYX [00:00:38]: As the things have accelerated a lot. Actually, let's do a quick catch up and Josh, you can chime on in as well. So Databricks got acquired. I talked to you at New York.JONATHAN [00:00:45]: Mosaic got acquired, although sometimes it feels like Mosaic acquired Databricks because, you know, we're having a lot of fun being here. But, you know, yeah.SWYX [00:00:52]: Yeah. I mean, you are chief scientist now of Databricks.JONATHAN [00:00:55]: Chief AI scientist. Careful with the title. As much as I would love to understand how Spark works, I'm going to have to defer that to much smarter people than me.SWYX [00:01:03]: Got it. And I don't know about like what you would highlight so far as a post-acquisition, but the most recent news is that you guys released DBRX. Is that the thing that most people should be aware of?JONATHAN [00:01:13]: Actually, that's no longer the most recent news. Honestly, the most recent news, we announced this, but it was at our Data and AI Summit last week. So it was announced among like 100,000 other things, is that we finally released our text to image model, which has been a year in the making through a collaboration directly with Shutterstock. There was a lot of work put into finding a dataset that we were comfortable with working on and trying to build a model that honestly, I felt like I could trust and that others might be able to trust to put out in the world. So that model was released last week. It's unfortunately just available via API due to the fact that the data is quite sensitive and quite valuable. It's Shutterstock's entire business in a lot of ways, but I'm still really excited that there's now a model that is trained on a dataset where the provenance of every single image is known, and it's a damn good model. So I'm really proud of the team on that.SWYX [00:01:55]: Yeah, amazing. Josh, do you have any thoughts on image model questions?JOSH [00:01:59]: That is not my area of expertise, but I was excited to see the release of it last week as well, and very happy that you guys did a nice job on the data side of everything there. So that was cool to see.SWYX [00:02:09]: I think what's unusual is like, I think Shutterstock's doing multiple deals in multiple labs. So what is the Shutterstock model? Like, I guess, is this the house model for Shutterstock? Is this Databricks' version of the Shutterstock model? Like, what is this?JONATHAN [00:02:22]: The way that I would think about it is that Shutterstock is doing an amazing business in AI across the board. Their dataset is kind of widely known to be the best stock photos dataset in the world, the most comprehensive, the biggest. When you think about like, what dataset am I going to train a multimodal model on? You call Shutterstock. And I, at least I've heard in the news, like OpenAI, Google, Meta, Apple have all called Shutterstock and made those deals. So a lot of models have had Shutterstock data incorporated into them. But this is the only model I know of so far where it was, you know, exclusively and specifically trained just on the vanilla Shutterstock data. There was nothing else mixed in. We didn't go and scrape the web and find other data or combined datasets or anything like that. And so this is, in some sense, the house blend. But the other piece is that it's just a dataset where the provenance of every image is known in public. Where did the data come from? It is the Shutterstock collection. That's it. You know, nothing less, nothing more. And certainly being at Databricks, if I've learned one thing, I've learned about enterprise customers and what they want out of AI. And one of the things they ask for most is just, what can you tell me about the data the model was trained on? And here, especially for text to image models, where images are just tricky subject matter, there's been a lot of kind of legal conversation about images, especially. It's nice to just have something where I can point to it and say, you know, if you want to know where the images came from, these are what they are and this is how they got there.SWYX [00:03:36]: I will talk a little bit about Databricks because it's relevant to the rest of today's episode. So Databricks, sorry, I keep misspeaking. It's DBRX.JONATHAN [00:03:46]: DBRX, actually, there's been a pronunciation update. It is now D-B-Rex. So we have decided to add a dinosaur mascot because what model doesn't like a mascot? So literally, I wish I could pull it up. There is a little plush dinosaur that we had made. It's like the world's cutest dinosaur, but it is the official mascot of D-B-Rex. And there's a little dinosaur logo that, you know, you'll probably see around a little bit more because DBRX is a mouthful, but D-B-Rex, like, you know, it's just kind of...SWYX [00:04:13]: Rolls off the tongue. I love mascots. Like every company should have a mascot. And I think Hugging Face got it right. You need an emoji mascot because that's the minimal viable image.JONATHAN [00:04:21]: I probably shouldn't talk at all about, you know, Velociraptor, but, you know, that's a, maybe that's something we can talk about later in the summer. I'll just leave it at that.SWYX [00:04:28]: Okay. That's a hint to names. I feel like your names leak a lot of alpha. So just to quickly cover the headline details, DBRX, as Make Sure Experts model, that's fairly big, 132 billion total parameters, so 36 billion active on any input, pre-trained on 12 trillion tokens of text and code, and did really well on evals to the point where you had to dye your hair blue. That's my high level conclusion.JONATHAN [00:04:53]: Never make a bet with your team two weeks out from model launch, even when, you know, human eval is looking quite bad. Because if you set some bar, even if it's arbitrary and you think there's no way in hell they're going to hit it, apparently money doesn't motivate people anymore. Humiliating their boss motivates people. So Josh, you should really take a hint from this. You know, you cannot pay someone enough money to make up for you dyeing your hair blue.JOSH [00:05:15]: I'll keep that in mind for our next model.SWYX [00:05:17]: It works. So speaking of Imbue's next model, perhaps Josh, you want to actually just say hi to the general sort of latent space audience and talk about what we're releasing today. Yeah.JOSH [00:05:26]: I'm Josh, CTO of Imbue, and we're not releasing the model. We're not releasing the weights, but we are releasing a bunch of different things that should make it easier for other people to make their own models. So I think right now, training foundation models from scratch is like a very difficult, time-consuming, expensive, kind of risky endeavor, especially for smaller companies. And the things that we're releasing hopefully make that at least a little bit easier. So the things that we're releasing fall into kind of three different buckets. One is infrastructure and scripts for dealing with the kind of hardware and hardware failures and understanding how well is the actually lowest level of thing actually working so that you can actually do your training at all and at a reasonable speed without having to constantly restart, etc. So infrastructure and training scripts. A second set of things is around the evaluation. So after you've trained it, like how well is this actually working and how do you know how well it's working? We're releasing a whole bunch of different data there, a new benchmark about code, reasoning, understanding, as well as our own private versions of 11 different open source benchmarks. So things like pool queue or ANLI, where we've gone through and kind of cleaned up the data as much as possible by looking at all the ones that models get wrong or that are flagged for ambiguity and also our own kind of private reproductions of those where we've done like a kind of clean room black box, like, okay, this is what the data set is supposed to be. Here are some examples. Let's make our own version of this to make sure that there is no data contamination, etc. To make sure that we're actually, you know, not testing on train. And then I think a final thing that we're releasing there is around 450,000 human judgments about ambiguity and question quality, which we used in the process of cleaning these evaluations and we also hope will be helpful for other people training kind of similar models. And then the third thing is CARBS, our hyperparameter, our cost-aware hyperparameter optimizer, which was especially helpful for being able to experiment at much smaller scales and then scale those experiments up to the much larger scale kind of on the first try without having to retry it. You don't want to be training, you know, 10, 20 different 70B models. You really want to get these larger modelsSWYX [00:07:30]: right on the first try.JOSH [00:07:30]: And so the ability to kind of tune things very precisely and learn scaling laws, not just for, you know, the like data and flops, but also for learning rate and all the other hyperparameters and see like how should you scale these things up was extremely valuable to us as we were training the larger models. Yeah, that's a lot of stuff.SWYX [00:07:49]: Yeah, exactly. So there's a bunch of stuffJOSH [00:07:50]: we'll have to go through all of it.JONATHAN [00:07:52]: Yeah, I just want to throw in how excited I am about this. This is the stuff that nobody ever talks about. That is the difference between success and failure in this stuff. Like, can you get your cluster to run? Can you get software on your cluster? Can you figure out what broke? Because fault tolerance is still not really built into any of the fundamental primitives of training models. And so if something breaks, you have to go figure out what broke, your job stops, you have to restart your job. It is a nightmare just to get to the point where anything can train on the cluster. A basic MPI hello world that has the GPUs talk to each other is hard enough, let alone actually training a model, let alone getting good performance out of the GPUs, let alone actually getting a model that converges to anything interesting. There's so many levels of things you have to accomplish. This is the kind of stuff that matters. I think to a point that Josh made earlier, before we got on here, there are plenty of weights out there. Nobody's released this.JOSH [00:08:46]: Yeah, that was part of the motivation actually is that there are lots of other things that are complimentary, but I have not seen nearly as much discussion about some of these other things that we think are pretty important. I mean, in some sense,SWYX [00:08:56]: I'm very excited to have Jonathan on because this is a little bit, you're a bread and butter with Mosaic. And I think you've released some part with Composer. And I think it's just really interesting to see like a different take, basically a full stack take that's kind of open source today.JONATHAN [00:09:18]: Yeah, it's really kind of, it's been an ordeal to figure this out. And every time something changes, whether it's a new GPU or even a new driver update, you get new creative errors and new things go wrong. And, you know, we've dealt with the weirdest things from, you know, our InfiniBand cables getting stolen from the data center twice, like in boxes before they arrived at the data center. Like, you know, Porch Pirate basically had stolen our InfiniBand cables back when those were hard to come by. To like, you know, weird recalls of switches to like the strangest stuff has happened. I have my favorite GPU failures I've seen, like ones where the GPU doesn't fail, it has a correctable memory issue and the memory correction causes the GPU to become a straggler and hold up the whole job. Like weird stuff happens and figuring out how to not just identify all of that, but then eventually productize it, is in some sense, the entire story of Mosaic and now Databricks in terms of our ML offering. Really, the thing we offer is we have gone through this suffering and figured out how to even productize that. It has been a pain in the butt.SWYX [00:10:20]: Yeah, it's a lot of work.JOSH [00:10:20]: I think my favorite failure was GPU is just giving wrong math. Like if they give errors, great, because you can see the errors, but if they just give you the wrong math back, not so fun.SWYX [00:10:30]: When did they give you wrong math?JOSH [00:10:32]: Like literally you could just, you know, add two things. For example, the numbers come back. They're not the numbers that they're supposed to be.JONATHAN [00:10:40]: I think it's important to say at this stage, just because like it, I think it goes without saying for Josh and I, but it's worth saying here, this isn't to say that like anything is wrong with us. It's not like NVIDIA did a bad job or, you know, Mellanox did a bad job or the like the server builder, the data center operator, the cloud provider, like the million other parties that are involved in building this. We are running these insane chips that are huge and complicated and built on tiny transistors at insane frequencies with insane heat in data centers that for the most part, were not built remotely for this kind of power or heat and have been retrofitted for this. Like failures happen on a good day with normal CPUs. And this is not a good day and not a normal CPU for the most part. It's fun to joke about all the weird things we see. This is not to say anybody's done anything wrong. This is just kind of part and parcel of working on a massive cluster running at multiple megawatts of power at a time.SWYX [00:11:32]: It's crazy. Yeah.JONATHAN [00:11:33]: So optical cables, like all sorts, like everything.SWYX [00:11:37]: I'll take the opportunity to start going to the sort of infra piece. There's just like a description of the infra just to give people a sense of what we talk about when we talk about massive clusters. So I'm just going to read off the blog post here. This post is about one cluster that has 4,092 H100 GPUs spread across 511 computers. They use unified fabric manager nodes, which manage the infinite band network. And you talk a little bit about your networking. Is there anything unusual about this setup that you'll call out to people?JOSH [00:12:03]: Yeah, actually this particular cluster is a little bit non-standard. The normal, like vanilla setup for these large clusters as vanilla as it can be is what's normally like a 127 node cluster. So closer to like 1024 GPUs instead of 4,000. Here we have a larger cluster. As you start to get into the larger clusters, the networking becomes a little bit more custom. It's a little bit more, it's a little bit trickier. It's a little bit more difficult to get these things to all be able to talk to each other at the same speed. And so this has, in this particular case, this is a three tier network architecture instead of two tiers, kind of the normal one. So most of the clusters are a little bit smaller. As you get to even larger scales, then this becomes even much more complicated,SWYX [00:12:43]: much more expensive.JOSH [00:12:43]: So we chose this particular scale, kind of knowing our own workloads and kind of what we wanted to do. This was kind of the right size for us. But yeah, I think it's not exactly vanilla already. It's already getting into kind of the custom territory.SWYX [00:12:54]: So my understanding is that there, and is there any part of this that comes with the Voltage Park deal that you guys had? Is that part of the hardware that you got from the deal with them?JOSH [00:13:04]: Yeah, so we worked really closely with Voltage Park to set up all their clusters and infrastructure and everything and kind of decide even like what to order, how should the networking work? Like we were very involved in kind of the construction and bring up of this. And that's what this post is about, is about that process of like bringing up all these, there's like different clusters in different places of different scales. So in this particular post, we're talking about this one 4096 GPU, but there are other clusters that they have as well. And we were very closely involved with figuring out the exact architecture and kind of the trade-offs that go along with picking, you know, those exact components. You really don't want to like place the wrong order because it takes months to get it and it's very expensive. So yeah, we were happy to help out with that.JONATHAN [00:13:43]: And then your bit of good cables get stolen.SWYX [00:13:44]: Yeah, yeah, exactly.JOSH [00:13:47]: We wanted to make sure that we ended up with compute that would work for us and that would also work for their other customers. And so we kind of helped design something so that we would get exactly what we were looking for. We knew that these kinds of details would be super important and that getting down to the level of the hardware and like having these good scripts and everything was going to be a core part of like actually getting this to work. I'm very glad that we did that. I don't think that most companies kind of take that full stack approach, but for us, it certainly paid off.SWYX [00:14:12]: Yeah, it's basically sort of built to spec. It's interesting that relationship because you usually, for the rest of us who don't operate at your scale, we take whatever we can get from cloud providers, but you are basically co-designing from the single machine up. And you described that a little bit. Do you want to take us through the process that you described here?JOSH [00:14:27]: Yeah, so for the actual, like the blog post and kind of bringing these machines online.SWYX [00:14:32]: Yeah.JOSH [00:14:32]: So yeah, I think the process, as we have it broken down in the blog post, there's kind of a few different layers. First is like getting the individual machines to work at all and then getting the machines to actually be able to talk to each other. So getting the InfiniBand networking to work and then getting to a point where, you know, not just the machines are working and they can talk to each other, but everything is actually working correctly. There's a big gap between like it's working at all to it's working perfectly correctly. And then after you have all this stuff working perfectly correctly, nice and healthy, then now you get into kind of the software data, like training issues. And then after that, you're still not done. Like now, even once you're training at full speed, things are going to fail over time. Things are going to change. There's going to be new, you know, firmware updates. Like how do you kind of deal with this change and flux over time without going crazySWYX [00:15:16]: and pulling your hair out,JOSH [00:15:16]: trying to like reproduce things or understand why there were regressions. And so there's a lot of work to kind of automate the infrastructure tooling as well. And kind of the first step, like bringing these things online in the first place, you know, you have hundreds of machines at this point. So you don't necessarily want to be like walking around with like a CD-ROM or a USB drive, like plugging it in with your keyboard, like hitting next, next, next on the OS install. That's not how this works. You do that for one machine. And then you use, we use this thing called Metal as a Service to bring up all the other machines. So it's a kind of server that can kind of install the operating system on these other machines. So most like when you're talking about these machines, like each machine is, you know, on the order of hundreds of thousands of dollars. So they usually come with a kind of out-of-band management interface as well. So they don't, they have their InfiniBand networking. They have their normal 100 gigabit per second Ethernet networking. These are like dual, redundant, et cetera. And then you also have this extra out-of-band management network. So you can log in and you can see like the boot screen or you can see the blue screen of death. You can like get in there and actually see what was wrong, which is pretty fun. And it makes it like possible to automate a lot of this work. So the beginning of that, and the blog post goes into much more detail about like exactly how we set these up and kind of the other errors that we ran into. When you're bringing these online, you'll definitely have failures. Even if they all worked in the factory, they get shipped, some parts come loose, something fails, something goes wrong. So when you're bringing them online, there'll be some that don't quite work for all sorts of reasons. As you start to be working with machines at this scale, like if something happens one in a thousand times, you're like pretty likely to see it. And so you can get pretty rare, weird things, especially since we had fairly early builds and fairly early versions of this hardware. Like these are some of the like first machines that were ever produced, some of the first GPUs. So you've got some extra special things there. We definitely worked with Dell, for example, on making fixes in the firmware level to be like, okay, like this thing is wrong. Like we need to update this at the firmware to like actually fix this particular thing. So we worked pretty closely with Dell and Nvidia. Yeah, that's what I'm saying. Like this stuff gets complicated. And the thing is like, you know, taking a step back, the whole reason we're doing this, right, is that we knew that this was going to be complicated. There would be these kinds of failures. And if we're just using, you know, AWS or some other cloud provider, these errors are still gonna be there and you're gonna have no way to know and no way to debug this and no way to diagnose what's going wrong. And so we would much rather be able to like call up Dell and say, hey, this isn't working. And they're like, yep, okay, cool. Let's debug it together. Oh, I see. Yeah, cool. We'll ship a firmware update and actually fix this for you. That was a much better experience than like, great, just magically fails. I guess we restart and hope that that machine goes away. Like that's not a very good place to be. So yeah, that's kind of the first place is getting to a place where like GPU training is working on your single node machines. You can observe stuff. We have tons of tooling around like, you know, Prometheus and all sorts of other tools for understanding what's going on in these machines because you don't want to be like logging into each one and looking at the temperature or something you really need to have tooling to collect all these metrics, et cetera. Unfortunately, all of the scripts that we have for this are like for this entire cluster and for all this infrastructure are a little bit like special purpose for our particular thing. So it's not that every script that we have, it's not that you can just like take this and plug this in. Even if we did open source all the tooling that we have, you'd still have to do like a lot of work to open source it. What we are releasing is as many of the things that we can that are going to be useful for other people. You're still going to have to have some way of kind of managing these things, making your own like logging aggregators, et cetera, et cetera. So that's kind of bringing them up to the like, you know, the single nodes that are working. From there, it goes into, I'm happy to keep going if you want. Well, I just want to leave the opportunity for JohnSWYX [00:18:53]: to comment if there's anything that's different from how he runs things.JONATHAN [00:18:57]: Oh, I mean, all I'll say is I'll endorse this and say this s**t is hard. Like this is really, really hard. And, you know, I have a special props to, you know, the folks in Vue because they were building this from the ground up. You know, at Databricks and at Mosaic, we typically work with cloud providers because some of this stuff is just, there's too much to handle. It's complicated. There's a lot to deal with. And this doesn't even get into things like physical security, you know, securing power if you're the data center operator. Like this gets infinitely complicated and you have to abstract somewhere. Like, you know, and then you get to the folks who are literally building their own custom chips and like, good God.SWYX [00:19:36]: Like, oh my God, that's, you know,JONATHAN [00:19:38]: if you're one of those folks, you're having, you know, pour one out for the infra people at some of the AI chip startups who are having a really, really interesting time right now. But this stuff is really hard. And I don't think we talk about it much because there's so many other things that are hard. But the other hard things, I think everybody's becoming pretty familiar with at this point. This is something that I don't think there's ever really been a comprehensive discussion of, at least not that I've seen.SWYX [00:20:00]: Yeah, so my impression is that you guys, Mosaic, have your own software for sort of spinning up and down machines, just like Imbue had to build. But Imbue probably, it sounds like Imbue, you guys went fuller stack. I don't know how to describe it. Like Mosaic is not working with Dell on like their firmware.JONATHAN [00:20:21]: No, no, we're typically working with like, you know, pick your cloud provider on their Dell firmware or what have you. Like, it's kind of, I think one of the things, I don't know, Josh, you can correct me on this. It's kind of impossible if you're doing training to not go all the way through the entire stack, regardless of what happens. Like somehow I'm still chatting with cloud providers about power contracts, even though the whole point of dealing with the cloud provider is not to have to think about power contracts. Somehow I'm still asking them about which InfiniBand provider they used this time to see if this is part of the bad batch of cables I encountered on that cloud provider or what have you. Or like, we're still talking about a firmware update from pick your provider. You can't not do this. It's convenient that they have data center staff who are worrying about what to send back to which provider when, and they have people who can go and wait for the InfiniBand cables so they don't get stolen outside. But, you know, it's kind of, it's impossible not to really go full stack if you're thinking about the infrastructure at all. I don't know, Josh, correct me. No, I think that's right.JOSH [00:21:17]: That's what we expected from the beginning as well, is that we would inevitably have to get into the details here. And I'm glad that we kind of just planned for it. I think it made it a lot easier from our perspective to have direct control over this. Instead of having to go to the cloud provider that goes to the data center, that goes to the supplier, we could just go direct to NVIDIA or DellSWYX [00:21:37]: or the data center,JOSH [00:21:37]: whoever was responsible and be like, hey, this thing needs to change. And they're like, oh, okay. Yeah, that is our responsibility. Great, we can fix that. So it was just a lot easier for us to fix these bugs than if we had to go through an extra layer of email.SWYX [00:21:48]: Something we discussed in the pre-show was that you had a rule of thumb for your cluster of reliability. You say here in the post, by and large, you expect around 3% of your machines to break every week. So you're basically going to turn through all your machines in a year.JOSH [00:22:04]: As it says in the post. So that would be true if it was a uniform failure like that. But as it says in the post, it's usually these kind of problematic nodes. And to be clear, that is the number that we've heard from other people is like they're having about 3%. I don't think we're experiencing failure rates that are that high. I think ours is actually quite a bit lower than that, probably because we've taken the time to like dig into a large, maybe larger number than we should have of these failures and get to the root cause of it and be like, oh, okay, like that's exactly what's going wrong.SWYX [00:22:33]: How do we fix this?JOSH [00:22:33]: How do we prevent this from happening? How do we make automated checks for this so that if it does happen, it just goes back to whoever owns that particular part of the process and they can fix it immediately.SWYX [00:22:43]: And that's part of what you're also open sourcing, which is the health checks, right? You got the NIC health checks, GPU health check, this space health check, Docker D message. I don't know what that is.JOSH [00:22:52]: That one is just a lot of stuff.SWYX [00:22:54]: Yeah.JOSH [00:22:55]: That one is one where we realized that actually like when these machines boot, sometimes they wouldn't actually boot cleanly all the way. Or when they rebooted, they had problems that they didn't have when they were working before, which was kind of frustrating. Like usually if you restart your computer,SWYX [00:23:08]: it gets better.JOSH [00:23:08]: Here you restart. It did not get better.SWYX [00:23:10]: It got worse.JOSH [00:23:10]: That was very frustrating. So this health check looks at every particular line we've ever seen from the boot, like in D message, like every single log line that your computer emitsSWYX [00:23:21]: and says like,JOSH [00:23:21]: have we ever seen this before?SWYX [00:23:23]: Is this expected?JOSH [00:23:23]: Is this in the right order? Or is there something out of place? If there's anything out of place, let me say, okay, great. Like now it goes into this, like longer, more triage list of like, all right, great. Like, is this acceptable?SWYX [00:23:33]: Should we flag this?JOSH [00:23:33]: Like, should someone take a look at this? So we're looking down at a very, very granular detail level, what's happening on these computers to make sure that nothing is out of place. And that's critical because without that, if you're running your training, as Jonathan said, and this thing is slow, like what are you supposed to do? Right?SWYX [00:23:49]: Like you really,JOSH [00:23:49]: you really want to be very certain that like all 4,000 of these GPUs are working like they're supposed to.SWYX [00:23:54]: We know that.JOSH [00:23:54]: And so if it's slow, it's because like we messed up the config or something else and not because of this earlier thing that's like really hard to detect in software later.JONATHAN [00:24:01]: Yeah. I think the, I'm just curious to ask,SWYX [00:24:03]: like, you know,JONATHAN [00:24:03]: suppose you were to set up another, let's say another H100 cluster and it were at a different data center. And instead of the vendor being Dell, it was super micro or what have you. How much of this would be repeatable? And how much of this would you have to redo? I, you know, I genuinely don't know.SWYX [00:24:18]: A decent amount.JOSH [00:24:19]: I think it would go a lot faster the second time. I think there's lots of learnings that we had. And also the blog post,SWYX [00:24:24]: you know, yes,JOSH [00:24:24]: we are releasing the health checks, releasing some scripts, but a lot of the valuable stuff is also in the blog post itself, in the details and kind of the, you know, the learnings that we've had and the sort of errors that we run into. We tried to as much as possible surface those to other peopleSWYX [00:24:36]: could learn from thoseJOSH [00:24:36]: and avoid the same mistakes or failures as well. But I think it would go a lot faster.SWYX [00:24:41]: Although, yes,JOSH [00:24:41]: there would certainly be some things that'd be a little bit different. I mean, there'd probably be different CPUsSWYX [00:24:46]: or whatever,JOSH [00:24:46]: but I think a lot of that stuff is less,SWYX [00:24:49]: it's less,JOSH [00:24:49]: that's the like, that's less variable. I think most of it would apply the second time around. Although I'm sure next timeSWYX [00:24:56]: we're building one,JOSH [00:24:56]: it'll probably be, you know, at a scale that's 10x as big with a different chip or something like this.SWYX [00:25:00]: And then who knows?JOSH [00:25:01]: Yeah, with Kinect X8,JONATHAN [00:25:02]: that will have its own fun behavior and all that good stuff. Yeah.SWYX [00:25:06]: Perhaps there's something that people don't discuss about, and you don't even talk about this in the blog, but I always wonder is what is the timeline that's like kind of reasonable for this amount of work, at least the initial stages? And also what does the team composition look like for setting up a cluster, right? Like what are the mix of skills that you typically would require to get all this going?JOSH [00:25:27]: I'm, I can't really speak to typical. One thing I am very proud of is how much we accomplished with such a ridiculously small team. Like our infrastructure team is like, you know, fluctuates from week to week, depending on like how many things are on fire and how much we need to build. But it's like between like three and six people, like it's small. It's not like some huge team of like tons and tons of engineers. But those people are very, very good at what they do. And so that has allowed us to get a lot of mileage out of out of these things. I think it's not that we're building everything, right? It's not that three to six people build this whole thing. I definitely want to like, you know, say thanks very much to Dell and H5 and NVIDIA and the other people that have done a lot of the work, like to bring up this cluster, you know, with 4000 GPUs and three tier networking, networking architecture, you have 12,000 cables. So that's 24,000 things that need to be plugged in. Like that's just a lot of stuff to plug in, right? And you don't want to mess it up. Like each one needs to be done correctly. Like it's a little bit loose. Like it doesn't really work.SWYX [00:26:23]: If you break it,JOSH [00:26:23]: you need to replace it. Like there's a lot of workSWYX [00:26:26]: that goes into this.JOSH [00:26:27]: Yeah.SWYX [00:26:28]: And then, you know,JOSH [00:26:28]: that's just like that's it. That's if you were to do everything right the first time.SWYX [00:26:32]: And if you didn'tJOSH [00:26:32]: have to fix anything. But inevitably, you know, you will have to replace something, which means like taking all the wires out, pulling the thing out, taking all the GPUs out, going and fixing some cable, putting it all back correctly, putting it back in, doing this every time. So there were a lot of people at Dell, NVIDIA and at H5 that all helped a ton with this stuff. I don't know the exact size of the Dell team. It also fluctuated over time.SWYX [00:26:55]: Yeah, excellent. And then, you know, you so you have all the hardware set up and now you're firing it up for a single node. There's a long description that you guys have about just like monitoring the MFU, right? And what each situation might look might be indicative of. One of the most interesting things to me that I saw from here is like, you know, if training immediately starts off at 60 to 80% MFU, something's wrong.SWYX [00:27:24]: But like, you know, like what what are like, you know, some anecdotes or, you know, notable scenarios here that you might you might call out as maybe counterintuitive or super interesting.JOSH [00:27:36]: There's just so many of them. I mean, one of them, which I think is probably pretty common, like common knowledge by this point. But like we did have a sort of likeSWYX [00:27:46]: which one was this exactly?JOSH [00:27:47]: I think for the MFU, like gradually getting worse over time. I think that one, when we saw that the first time we were like, what the heck is going on? Like, why does it get just like a little bit worse? This is so strange. Like, what is it getting lazy or tired or something? Like, is it heat? Like what's going on? And in this particular case, it was memory fragmentation. Because you have hundreds of machines, they're doing garbage collection slightly different times. And then they get slightly further apart and slightly more and more jittered until eventually they're all happening kind of at random times. And just like really messing up each one of your steps. So you just turn off garbage collection and call it a day, basically,SWYX [00:28:20]: to be honest.JOSH [00:28:20]: There's other things you can do if you want to be a little bit more sophisticated about it. But you can also just manuallyJONATHAN [00:28:25]: have it all garbage collect on some interval. Like that's what we've done. We just have a garbage collection callback that just runs. But I've seen the exact same thing.JOSH [00:28:33]: Yeah, yeah, exactly. So I thought that one was kind of funny. And we did trace that one down and look and we did find the actual call. Like, again, this goes to like having good tools. So we had really good tools where we could look at a bunch of like actual traces in C and be like, OK, cool. This is the thing that's taking a lot of time. Or like, you know, this is the thing that doesn't quite line up here. Like, oh, I guess it's garbage collection. OK, cool.SWYX [00:28:52]: Interesting.JOSH [00:28:52]: Yeah, let's just try taking it off.SWYX [00:28:54]: OK, great.JOSH [00:28:54]: That's what it was. Now we can fix it. So for each of them, like basically bugs are not hard if you have good tools. But if you don't have good tools, bugs can be very, very hard. So similarly for like heat, another thing that we saw was like, oh, you know, the CPU is getting throttled. OK, well, it's easy to see if you're monitoring the CPU throttling or monitoring the heat. If you're not monitoring that, it's really hard to know why it's just suddenly one of them is going slower. I noticed also in the pieceSWYX [00:29:17]: that you mentioned FSDP with 0.3. Actually, we met, I went to iClear and Guanhua from the DSP team was there presenting 0++. I was wondering if you want to make any call outs to, you know, particular open source or open library or open whatever implementation teams that were super helpful in your process. I think we ended up actuallyJOSH [00:29:39]: pulling from a whole bunch of different ones to pull things in into our own particular pipeline. So we use things from NVIDIA's, you know, Megatron stuff. We use stuff from probably DeepSpeed. I think we pulled in a bunch of different pieces from a bunch of different places. So it was really nice to see all these working open source like examples. I think I really appreciate all the effort that has gone into actually tuning these things because you can tune them, but it's a lot of work to like tune this stuff and do all this stuff from scratch. It's really nice to have like a working example. I think those are probably the two biggest ones, DeepSpeed and Megatron alone, but there are probably other ones as well.SWYX [00:30:13]: Is there a particular thing in the ecosystem where you would call out as like, you know, there should be something here that is open source, but like it's not really, it's like everyone kind of builds it on their own. I want to say something with the file system because everyone talks about the file system eventually.JOSH [00:30:28]: The file system actually was,SWYX [00:30:30]: I mean, we did somethingJOSH [00:30:31]: kind of dumb there. Like we have our own sort of local mirror so that we can, you know, like a crappy version of S3SWYX [00:30:38]: that's local,JOSH [00:30:38]: but it's just a pretty simple script, right?SWYX [00:30:41]: Like I think we run likeJOSH [00:30:41]: a little web server that just like serves files and then, you know, it can upload themSWYX [00:30:45]: and download them.JOSH [00:30:45]: Okay, great. And part of the reason we did that is that our internet connectionSWYX [00:30:50]: in the beginningJOSH [00:30:50]: was not the like full speedSWYX [00:30:52]: one that we wouldJOSH [00:30:52]: eventually have. And so we are a little bit more kind of bottlenecked in terms of internet bandwidth. And so we had this. I think we looked at a bunch of services out there like Minio and some other ones, but a lot of these like come with a lot of extra overhead and maintenance. And since we already have so much infrastructureSWYX [00:31:09]: to deal with,JOSH [00:31:09]: we kind of didn't want to, you know, bring in a whole other like cloud provider, virtualize something, something.SWYX [00:31:14]: We just wanted something simple.JOSH [00:31:14]: So we went with that, which has been quite helpful. Like our toolsSWYX [00:31:19]: are usually quite simple.JOSH [00:31:19]: It's like Bash and Python and SSH and Docker. Like we'd like to keep things simple so that's easier to debug, like less layers of infrastructure, less layers of abstraction, make it a lot easier to work with. Like we don't use Kubernetes,SWYX [00:31:30]: for example,JOSH [00:31:30]: and we just directly launch these things. And it's just been much easier to debug this way. One tool actually that does come into mind that I will call out is Kraken from Uber. That was great. We love that tool. We were a little bit skeptical. What is it?SWYX [00:31:44]: I'm sorry. Yeah.JOSH [00:31:45]: So Kraken is this, yeah, it's a distributed like Docker registry, basically, that uses BitTorrent to like transfer things between the machines in a sort of nice optimal way. Like in the very beginning, the naive way is like you have this one Docker registry, which was outside of the cluster. So every time we change an image, you know, there's many gigabytes that each of the 500 machines needs to download.SWYX [00:32:07]: So that just takesJOSH [00:32:07]: a really long time. So what this thing does is like just one of them downloads it and then like they all sort of broadcast all the pieces to each other. And it was just like a really nice, fast way of getting these images down. And it was very robust.SWYX [00:32:19]: Like there's a lotJOSH [00:32:19]: going on under the hood, but I think it's a pretty cool tool that we haven't really had any bugs with it at all. Amazing.SWYX [00:32:26]: Yeah. I mean, that's all my questions, I guess, for the info piece. I don't know if, John, you had something that you were sort of burning to ask or.JONATHAN [00:32:33]: No, all I can say is just sameSWYX [00:32:36]: in a lot of places, like, you know, and they're done thatJONATHAN [00:32:38]: seeing this plus one. I think the one big difference, you know, perhaps in philosophies is we've tried to basically standardize on as much commodity stuff as possible, just because, you know, I think the reason I asked about trying to do thisSWYX [00:32:50]: on multiple differentJONATHAN [00:32:50]: pieces of infrastructure is like, I think we're running on like six or seven different clouds right now. And everybody has done something slightly different. And my gosh, the little differences add up as you know, you've seen. And so, you know,SWYX [00:33:04]: our philosophy has been like, whatever the hellJONATHAN [00:33:05]: we can standardize, please let's standardize it. Like vanilla off the shelf FSDB.SWYX [00:33:10]: And like, you know,JONATHAN [00:33:10]: we wrote our own data loader, but we've tried to make that as much of a standard as we can across our infrastructure and in Databricks, because things just start getting really complicatedSWYX [00:33:18]: or like we useJONATHAN [00:33:18]: Kubernetes extensively because it at least gives us a uniform set of APIs. Like that's our hardware abstraction layer to a certain extent for everything else. So it's just, you know, a difference in philosophy there. But otherwise, like, yeah, this stuff is really, really hard. And I feel like we take for granted how much of this, you know, is done for us when you go and you just query chat GPT, for example. Like, oh my God, everything going on underneath that, you know, it's kind of a miracle that the machines boot up, let alone that you can like query a giant language model that's probably doing inference across multiple machines and was trained across thousands of machines. Like, you know, minor miracle.SWYX [00:33:54]: Yeah, it is an awesome amount of power that we invoke with a single API call that we take for granted these days. It's absurd. Yeah, I mean, like Kubernetes, like that point about Kubernetes, I will say as a former AWS employee, like it seems like it would be ideal for imbue to at some point make it more abstracted or agnostic because you're going to want to, you know, replicate your setup. We do have our ownJOSH [00:34:19]: sort of replacement. It's just a much simpler version of Kubernetes. Kubernetes is really designed for running services, not for running experiments. Like that's not its like main architecture. And so for us, like we have everything that's like, cool, you're going to run an experiment. So you want it to run to completion, right?SWYX [00:34:34]: OK, great.JOSH [00:34:34]: Like the primitives are sort of built around a slightly different style. And that makes it a lot easier, like just a lot simpler to fit that the nature of like these machines are going to disappear. They will need to be rebooted for infrastructure upgrades. They will like something will happen to the GPUs. Failure is like baked into this as like a core part of our infrastructure. So it's not that we don't have an abstraction. It's that it's a sort of simpler, more tailored abstraction for the particular work that we're doing.JONATHAN [00:34:58]: Yeah, I think it all depends on what your goals are. And like, I think the challenge in a lot of the deep learning stuff right now is that people are trying to like, people often build things that are more complicated than necessary to get the job done. And the complication is the enemy of everything. You know, don't use a fancier parallelism strategy than you have to. Don't use a fancier set of libraries than you have to.SWYX [00:35:18]: Don't do anythingJONATHAN [00:35:18]: that you don't have to do because it's hard enough as it is. Like, don't overcomplicateSWYX [00:35:23]: your own life.JONATHAN [00:35:23]: Don't try to bring in more tools or more fancy architecture tweaks if you absolutely don't have to.SWYX [00:35:29]: Like getting to the minimumJONATHAN [00:35:30]: necessary to get the job done. And it's really tempting to want to try to use everything. So like, I totally understand that one.SWYX [00:35:37]: I think the last piece I'll maybe call out is that I'm just going to weave this in just because I see the opportunity to do it. Are there any infrastructure shifts that need to be, that need to rise because of changing architecture? So I think, for example,SWYX [00:35:57]: you're announcing a dense model, a 70B dense model, whereas John just worked on DBRX and the image-to-text model, which presumably has different bottlenecks.JONATHAN [00:36:10]: That's correct for us. You know, we train both dense and mixture of expert models. The one we happened to, you know, kind of get permission to open source was a mixture of expert model. And those models are very demanding when it comes to network bandwidth, at least if you're training them in kind of FSTP 03 style, where there's just a lot of parameters getting shuffled back and forth. And your ratio of kind of compute to amount of data that you have to shuffle back and forth becomes a lot worse because you're now, you know, you're only using a fraction of the parameters for every token instead of all the parameters. And so we had to really push the envelope on getting all the stuff to the right places on time. And so actually the networking part of DBRX was the single hardest thing, I think, of the entire process. Just get MOE training, working at scale across a big cluster. We still managed to, I think, do it all with commodity parts, which was very exciting. You know, we were using FSTP and we eventually used HSTP so that we could have HSTP as a version of FSTP where you have multiple smaller replicas and you're doing data parallel within those replicas. And that helped a lot with network latency issues that we were running into just because we were transmitting so much data, you know, for every single part of the process. I think it actually, like, it was instructive for how Google designs their hardware and software together personally. Their training, as far as I understand, using kind of a 03 style of training and have been for a while. They also train mixture of expert models. TPUs have a very different network bandwidth to compute ratio. They have a lot more bandwidth just objectively. And TPUs per chip tend to be a little bit less compute intensive and have a little bit less memory. You know, it's just a different design choice. So the ratio of flops to bandwidth is very different. And that means that it's much easier for Google to be able to pull offSWYX [00:37:54]: some of this stuff.JONATHAN [00:37:54]: They also have interesting, you know, Torus style network architecture or Torus style, like, literal network architectureSWYX [00:38:00]: is not like the model,JONATHAN [00:38:00]: but the network.SWYX [00:38:02]: Is this the sort of block attention? I forgot what you call it. So this is just more or the,JONATHAN [00:38:07]: yeah, this is more, not the ring attention, but these are the ring all reduces. Like you have three different dimensions of rings because they kind of put you in these three dimensional Toruses from what I understand. And so like, you know, Google's infrastructure in some sense is kind of, I wouldn't say built for this, but maybe the way that Google trains models is built for a slightly different bit of infrastructure they have. And it's kind of neat to think about that. You know, as one thing that I think NVIDIA announced for, you know, for, for both the GH200 and the GB200 is this hybrid networking where you'll have blocks of NVLink network chips. I think for the GB200, I think it's like groups of 72 GPUs will all have NVLink to each other. So higher bandwidth, then you'll have normal networking of some kind, InfiniBand or Rocky or what have you between these blocks. And that's kind of a, you know, it's a change due to the fact that, you know, it's hard to build really high bandwidth networks over very large groups, but it is now a blocked networking. And you have to think about how you architect your model and your parallelism differently. You also have to think about fault tolerance differently because it now matters where you lose a GPU, whereas it didn't before. So, you know, it's, it's, it's just all really interesting and really fun speaking personally, but it's going to mean new nightmares when we all move to that generation and have to think about, you know, new versions of these problems.JOSH [00:39:20]: As you go up to larger scales, it gets quite different. Like right now, you know, if you're experiencing, let's say, for example, you experience a GPU failure every day, that's fine.SWYX [00:39:31]: Just restart.JOSH [00:39:31]: If you make your thing 24 times as big, now it's once an hour. Now it stops being quite as easy to just restart, right? So now you have to kind of break, like bake in this sort of redundancy that you didn't have before. So I think as you go up in scale, you end up running into like a lot of really interesting problems that also inform the, the actual like design. Yeah, I mean, as an orchestration guy,SWYX [00:39:52]: this is why I always emphasize like very cheap storage or very fast storage. So you can checkpoint more, but I don't think that's probably not the best solution to for fast, you know, training.JONATHAN [00:40:05]: Which works fine when you're doing language and then you move to vision or video. And then, you know, you have multi petabyte datasetsSWYX [00:40:12]: and getting, you know,JONATHAN [00:40:13]: cheap, fast multi petabyte storage starts to bite. Like I've certainly encountered issues where the literal data center where my GPUs were did not have enough, you know, object store to fit the datasets that people wanted to bring into that data center from whichever users were, were trying to bring them in. And then you get to a wholeSWYX [00:40:31]: different world of hurtJONATHAN [00:40:31]: where you have to keep your data in a different region because the region is just out of storage. So things get fun really fast.SWYX [00:40:39]: Speaking of vision, Josh, actually, you know, Embu is an agents company, but you're only, you're announcing a text-only model. What, where does, where does the vision side come in?JOSH [00:40:49]: I think we've actually done a lot of work in the past and people can see kind of our blog posts about sort of self-supervised learning and some other kind of vision-related stuff in the past as well. So we're very familiar with, with that stuff. But I think our main focus right now is on kind of, as we say, coding and reasoning. And there, there's certainly a visual component to some problems. But, you know, it's not necessarily required for all problems. And actually we found that for most of the kind of like code writing and, and reasoning problems that we care about, the visual part isn't really a huge important part of it. Sometimes if you really need to, you can maybe describeSWYX [00:41:24]: the thing.JOSH [00:41:24]: There are other like, you know, multimodal models that you can use off the shelf to sort of plug in for those particular piecesSWYX [00:41:30]: that you need, right?JOSH [00:41:30]: Like if something is driving a browser or whatever, like you can sometimes get away with not having to have that baked into the original model. So our folk were, you know, in a sense, we kind of do a lot across the stack. We're working on our own infrastructure and pre-training and RL and fine tuning and products and everything. But in another sense, we're very narrowly focused on the application side. So all of the stuff across the stack is kind of going toward a very particular purpose. And so that particular purpose right now doesn't really need vision. So we think that people are going to make all sorts of really cool image modelsSWYX [00:42:00]: like Jonathan, right?JOSH [00:42:00]: And all sorts of interesting multimodal models into the future. We'll let them go do that. That's great. We'll take advantage of that, partner with those people in the future. And right now we're really focused on kind of the core reasoning and coding capabilities and aspects of the model.SWYX [00:42:14]: I wanted to go into carbs since that's kind of the next layer of the stack. We talked about carbs in the first episode with Kanjin because you've actually had a blog post about it like a couple of years ago. Maybe let's introduce it.JONATHAN [00:42:26]: Has that been a couple of years now?JOSH [00:42:28]: No, it must have been at least one year. Hopefully it's not multiple years.SWYX [00:42:32]: Sorry, I'm counting AI time. Yeah, yeah. Yeah, I was going to sayJONATHAN [00:42:35]: you're making me feel really old right now.SWYX [00:42:39]: I count everything before the generally intelligent rename as like, you know, prehistory. Yeah. And now sort of modernity, right? So I actually thought carbs was more about hyperparameter optimization in a sense of like sort of parameters, hyperparameter search. Whereas, you know, when you introduced it, especially in this blog post, it's more about scaling laws and predictability of like, are we sort of in the right ballpark before we scale things up? Maybe sort of recount the history of carbs.JOSH [00:43:10]: Yeah, so it really is a little bit of both. So carbs is, it's maybe a backronym, but it's for cost aware Pareto region Bayesian search. So this is about technically how it works, but carbs is like, you know, we like pastries and stuff.SWYX [00:43:26]: So great, why not? But the point is thatJOSH [00:43:29]: it's a cost aware hyperparameter tuner. So most hyperparameter tuners, you kind of say, OK, here's this objective function. I want you to make this number as big as possible or as small as possible, whichever direction you want to go. So yeah, just go make this number, you know, as small as possible. OK, so it'll try a bunch of differentSWYX [00:43:46]: hyperparameters,JOSH [00:43:46]: a bunch of different configurationsSWYX [00:43:48]: to figure out, like,JOSH [00:43:48]: how do I tweak your network and architecture, et cetera, to get the kind of best performance I possibly can. That's usually saying, like, you know, almost all of these hyperparameter configurations are, let's say they're all going to use the same number of GPUs or the same number of nodes.SWYX [00:44:01]: So it's going to runJOSH [00:44:01]: for the same amount of time.SWYX [00:44:03]: So you can do that.JOSH [00:44:03]: You can get a number out and that's great. But what carbs does is it says,SWYX [00:44:07]: OK, actually,JOSH [00:44:07]: what if we relax that constraint? What if we say each of these different points, we're going to model how expensive it will be to sample this configuration. So if what if we train with just one one hundredth of the data? Like, how well can we do?SWYX [00:44:19]: What if we trainJOSH [00:44:19]: with one tenth of the data? What if we train with all the data? That way you can understand, like, as we get more and more data, as we spend more and more compute,SWYX [00:44:26]: as we make a biggerJOSH [00:44:26]: and bigger network, how does performance change with these things that change? Like how expensive it is to even explore this data point. So by doing that, we can see the scaling laws for not just, you know,SWYX [00:44:36]: the scaling lawsJOSH [00:44:36]: from like the, you know, Chantilla paper, the scaling laws for all parameters. We can see how does how does the number of layers change with this? How does the, you know, the learning rate change? How do the like, you know, various types of regularization change? So you can see these nice scaling laws. And as you're going across costs, like how should this be changing as you're scaling up your model? So that, coupled with the kind of metric that we chose, which is a very precise way of measuring performance, allowed us to really like hone in on parameters that worked really wellSWYX [00:45:05]: and understand, like,JOSH [00:45:05]: how do we want to scale those up, especially as we're changingSWYX [00:45:08]: things about the network?JOSH [00:45:08]: Like one of the things that we did is we used a custom tokenizer. As we change this tokenizer, changes a bunch of other things about the model. So how should we scale up this entirely new tokenizer? Like no one has ever made a model this large with this tokenizer before. And so how do we want toSWYX [00:45:22]: change all these things?JOSH [00:45:22]: Harps kind of shows you, like, look, as you change these parameters, like these other ones are kind of dependent on this.SWYX [00:45:28]: Like this is the, these areJOSH [00:45:28]: the relationships between them. So you can better understand, like, OK, if I'm going to scale this up 10x or 100x, like, where do I want to be? I can only go so far. And so, you know, we did run, like, I think maybe it was like a 14b one or somethingSWYX [00:45:40]: like that to check.JOSH [00:45:41]: But and so we had a bunch of like 1b or 14b and then at 70b. I don't think we had a, I think we just did like one at 14b. So you can, we get to check that like, oh, is this on the curve? Like, is this where we expect? It was like right there. So then great, go on to the next one. Yeah, I mean, that makes a lot of sense.SWYX [00:45:56]: I wonder if, so one of the key questions, and correct me if I'm wrong, but like usually people do search or do their evals just based on loss. But you actually evaluate based on, you know, the sort of end state evals that people might expect, like HellaSwag and Lombata, whatever. What is the norm here? Is there a norm?JOSH [00:46:20]: Yeah, I don't know if there's a hundred percent.SWYX [00:46:21]: I don't know. I only see loss on most people's reports.JOSH [00:46:25]: I think it's easy to, like, loss is very nice because it's very precise. It will tell you, like, very fine grained differences between like really small changes in your hyperparameters or network architecture. Whereas, especially at the smaller scales, if you're looking at like accuracy, it's very noisy. Like it might be zero or a hundred or like, you know, fluctuating by like 10 or 20 percentage points, which makes it really hard to tell, like, did that change actually mean anything? So our loss is sort of a combination of these two. Instead of saying, like, let's just look at perplexity, we say, let's look at perplexity on the tasks that we care about for multiple choice questions effectively.SWYX [00:47:00]: So we're saying like, yes,JOSH [00:47:00]: this is formulated as a multiple choice question, and we're going to look at the, like, you know, the loss of perplexity for this particular answer token. And that ends up being something that's like both targeted to what you actually care about and also very precise. The nice thing about this though is that it's independent of the data that you train on. One thing that's annoying about perplexity or about loss is that as you change your data set, this is really obnoxious because now it fundamentally changes your loss, right? And so you can't tell, like, how do I tweak my data set? But because we have this held out evaluation data set where we're looking at perplexity, we can actually change the data mix. And so CARBs actually control what is the mix of data that we want to see, like how much code, you know, how much internet text, et cetera, in order to figure out what is the best optimal mix of data and we could do that because we have this other metric. So that was one of the things that was really, really helpful.SWYX [00:47:46]: I think there is a trend overall about changing data mix as training goes on. I don't know how, you know, we're deciding not to talk about data sets in this podcast, but what have you observed about the changing data mix question?JOSH [00:48:06]: We did some experimentsSWYX [00:48:08]: and we've actually talkedJOSH [00:48:08]: to a bunch of researchers who are doing work here as wellSWYX [00:48:11]: and looking at kind ofJOSH [00:48:12]: their experiments on this. And we were originally pretty hopeful because it sounds like something that should work and make sense, right? Like, oh, cool. Like maybe you would have your model, like learn the basic featuresSWYX [00:48:22]: and then over time,JOSH [00:48:22]: it could get really good at these complicated math problems or coding or something, right? But it just turns out that like, it's just not the way it works. Like we've done so many experiments and you can get like a tiny, tiny little boost from this, but it just is not like, it's just not the important thing, at least in the experiments that we've seen. So yeah, we've kind of, we're letting other peopleSWYX [00:48:40]: explore that moreJOSH [00:48:40]: if they want, but that just doesn't seem like the most promising direction for us.JONATHAN [00:48:44]: We've had some surprisingly good luck with this. We just released a paper on it. The details matter a lot and it really matters what you're trying to do with the model.SWYX [00:48:53]: Yeah.JONATHAN [00:48:53]: But it's been quite effective for us depending on the setting. And certainly when we're thinking about domain-specific models, this helps a ton. You know, to a certain extent, you can always think of this as like early fine tuning. But yeah, I like, there've been little glimmers of this in the literature for years. Like especially, I think the Gemini 1.5 paper mentions this. And I don't remember whether the Llama 3 paper mentions this,SWYX [00:49:15]: but it's kind of,JONATHAN [00:49:16]: it's one of those, like people have different ways to get to these endpoints.SWYX [00:49:20]: I think, you know,JONATHAN [00:49:20]: there are the architectural tricks that each lab has to mitigate loss spikes or what have you. And everybody's got, you know, their own bag of tricks and it leads to kind of sometimes this contradictory information. It's not contradictory. People are just kind of exploringSWYX [00:49:33]: different parts of the spaceJONATHAN [00:49:33]: in some sense. And there are lots of ways to get a great model. But certainly for us within our config, and it seems like, I guess for the folks at Google, within kind of the part of the world they live in, changing the dataset has helped, but the details matter a lot. And it's really hard to get those details right for the reasons Josh,SWYX [00:49:48]: you know, just mentioned.JONATHAN [00:49:48]: Like there's a lot of search involved and you essentially have to make hard choices aboutSWYX [00:49:52]: what parts of the spaceJONATHAN [00:49:52]: you're going to search and which ones you're going to leave be. And so, you know, some people have done an amazing job. Like I think the, who is it? The Deep Seek folks have done an awesome job looking at like batch size warmup. And that's been really, really fruitful for them. You know, other people are looking really hard at things like data mix, but it just gets tricky to look at everything.JOSH [00:50:09]: Yeah, I think we've found that like we could get some things that looked like gains from datasets. But one of the things that I like about carbs is that when we applied carbs to like properly tune things, then a lot of those kind of evaporated. Whereas like, like if we just tune these other parameters, actually we can get almost the same gains without having to do this more complicated thing. So at least in the experiment and in the settings that we've, like in the particular metricsSWYX [00:50:34]: that we care about,JOSH [00:50:34]: we haven't seen these kind of like pan out or scale up in quite the same way. But not to rule it out. And I think you're right, Jonathan,SWYX [00:50:41]: that there probably areJOSH [00:50:41]: a lot of like details that go into like exactly what is the metric, exactly what is the dataset, exactly which, like what schedule are we using for this. And I certainly wouldn't rule it out working.SWYX [00:50:52]: Quick question about emergence. Doesn't emergence throw a spanner into a theory of carbs? Ah, so there is a paperJOSH [00:51:01]: of which I really liked and I think informedSWYX [00:51:05]: a little bit of howJOSH [00:51:05]: we thought about this, which is are emergent properties of language models a mirage? And I think if you look at that paper, it actually makes a relatively compelling case that in fact, you know, this emergent behavior that you're seeing is not really emergent behavior, but is really a function of the evaluation metrics that we're using. So if you look at accuracy as a metric, what's happening is that accuracy is actually going up continually over training, but it's in log scale. So it starts out at 0.001%, 0.1, 0.1, 10.SWYX [00:51:35]: Only when you're goingJOSH [00:51:35]: between 10 and 90 do you see this happen, right? When you go from one in, you know,SWYX [00:51:40]: a thousand getting rightJOSH [00:51:40]: to one in a thousand getting wrong, like there's many orders of magnitude happening here.SWYX [00:51:44]: So when you're lookingJOSH [00:51:44]: at this in perplexity, then you just see this nice straight line. And so that's actually what carbs is exploiting. Like since we're, since our metric is in this kind of like perplexity log space, like you can see like, oh, it's just like getting better as you make it bigger in this nice, very predictable way. So that, and that is exactly what we saw. Like these things were really, really bad at, you know, predicting the multiple choice answer, just always guess A. OK, it's so terrible at it, but it was like learning to be less confident about that.SWYX [00:52:09]: Yeah. One trick I saw from one of the papers recently was just like, just randomize the order of the multiple choice questions. And if you, if, if, if they, if they over, if that hits the performance a lot, then they're just basically memorizing the test set, which makes a lot of sense.JONATHAN [00:52:28]: Yeah, this is, I, I mean, you know, I, I completely agree with what Josh said.SWYX [00:52:32]: I think the, you know,JONATHAN [00:52:32]: my bigger lesson is that anything can look however you want it to look. If you put it on a log scale to a certain extent and log, we love our log scales and deep learning for various reasons. Everything looks very clean on a log scale until everything looks very flat on a log scale. Um, I don't know. I like log scales always mix me up. That's, that's all I can say.SWYX [00:52:51]: Great. I think the, the last thing I was, I was going to mention on, uh, carbs. Oh, well, I mean, let's, let's just kind of go right into evals because I think that's going to be, uh, the, the sort of crowd favorite. Um, so carbs, we already mentioned, um, you know, leans heavily on, uh, the sort of end evals that we would typically eval LLMs on, except that you had to make your own. Um, there are a lot of documented problems with many of the common evals out there and you fixed all of them. It sounds like, I don't knowJOSH [00:53:18]: about fixed all of them, but, uh, I think in the same way that we like to dig into the infrastructure and hardware and understand, like what actually is goingSWYX [00:53:27]: wrong?JOSH [00:53:27]: Like what is the actual error on this machine with this GPU?SWYX [00:53:31]: And why did that happen?JOSH [00:53:31]: And how do we fix it? We take the same approach to the evaluations. So when we looked at the evaluations and actually looked at the data sets, you know, what we did isSWYX [00:53:39]: like, okay, if we're goingJOSH [00:53:39]: to be, you know, evaluating natural language, understanding and reasoning, like, let's look at all the data sets that are out there. Let's actually look at a bunch of the examples and say, like, is this a good data set that we should use for evaluation? That's kind of how we selected the evaluation data set that we had. Uh, and then when we looked at the actual examples in there, we noticed like a lot of these are very messy. Like some of them messySWYX [00:54:00]: to the point of likeJOSH [00:54:00]: incoherence and some of the ones that we didn't choose. Uh, but even the ones that we chose, like people tried pretty hard onSWYX [00:54:06]: these data sets.JOSH [00:54:06]: They did try and clean them, but there's just a lot of data points in there and it's just easy toSWYX [00:54:10]: make mistakes.JOSH [00:54:10]: Right. And so, you know, it's not that they have aSWYX [00:54:13]: hundred people lookingJOSH [00:54:13]: at every question, like that's just way tooSWYX [00:54:15]: expensive.JOSH [00:54:15]: So you end up with questions that just don't make sense.SWYX [00:54:18]: Somebody didn't reallyJOSH [00:54:18]: see this. Somebody just clicked the wrong box for the answer. Uh, or the question makes sense in your head. When you write it, we've often seen this, it's not even like malice orSWYX [00:54:26]: incompetence.JOSH [00:54:26]: It's really just like, you know, you write this,SWYX [00:54:28]: you're ready.JOSH [00:54:28]: You're like, this makesSWYX [00:54:29]: sense to me.JOSH [00:54:29]: You show it to another person like that makesSWYX [00:54:31]: sense.JOSH [00:54:31]: You show it to a thirdSWYX [00:54:32]: person.JOSH [00:54:32]: They're like, this makes no sense at all.SWYX [00:54:34]: That's because you'reJOSH [00:54:34]: kind of, you know, using a different meaning ofSWYX [00:54:36]: the word.JOSH [00:54:36]: And then when they say that, you're like, Oh,SWYX [00:54:38]: wow, you're right.JOSH [00:54:38]: That is actually really confusing. It's easy for things toSWYX [00:54:41]: kind of make sense inJOSH [00:54:41]: our own head. So what we did for the evaluations is really dug into the details of each of these data sets and tried to ask, like, what makes a goodSWYX [00:54:50]: question?JOSH [00:54:50]: What makes a good answer?SWYX [00:54:52]: Like, what does it meanJOSH [00:54:52]: for it to be ambiguous? We had a whole, like,SWYX [00:54:55]: we looked at lots ofJOSH [00:54:55]: data, broke this down, asked lots of peopleSWYX [00:54:58]: about all theseJOSH [00:54:58]: different questions to build a model of this and help us kind of clean these data sets. That was sort of one big piece of it. A second big piece was making sure that our data that we're training on is not data that we're testing on. So there we kind of took a step back and said, like, OK, well, let's just reproduce, you know, 500 to a thousand examples for every single one of these data sets ourselves. And just make sure that this data is definitely not in the, you know, the training set. So we did that. And then we're able to, like, now be confident about, like, our performance of our model and also performance of other open source and other closed source models. Yeah, there's a lot there.SWYX [00:55:33]: You had 11? I don't know how many data sets. I think so. One, two? Yeah. Any one you want to call out in particular to dive deeper on? Some of these are very famous, like HelloSwag, MitoGrand. Some are less famous, like Race. I don't know if... Race is a great data set.JOSH [00:55:50]: See that one?SWYX [00:55:51]: Yeah. Yeah. Just, you know, anything that's interesting you want on specific data sets? I think there areJOSH [00:55:57]: a few asterisks in there. You know, definitely read the whole paperSWYX [00:56:02]: as you're looking atJOSH [00:56:02]: some of these, like the GSM8K one is a little bit weird. I think one that wasSWYX [00:56:06]: kind of funny,JOSH [00:56:06]: it was, like, low performance on ethics from some of the more recent models. I think that was aSWYX [00:56:11]: little bit funnyJOSH [00:56:11]: because the models, you know,SWYX [00:56:13]: I think there wasJOSH [00:56:13]: a reaction to, like, oh, no, like, you know,SWYX [00:56:16]: the models are sayingJOSH [00:56:16]: bad things.SWYX [00:56:17]: And so they went way,JOSH [00:56:17]: way in the other direction. And now, like, on the ethics data set,SWYX [00:56:20]: it's always like,JOSH [00:56:20]: this is totally unethical, even though it's really fine. So they've just been tuned to, you know, make sure they don't make any PR disasters.SWYX [00:56:28]: I thought that wasJOSH [00:56:28]: a little bit funny. Not to say that it's necessarily like a flaw of the model, but just kind of like, you know, political or tuning opinion. I think the main takeaway, I was just going to saySWYX [00:56:38]: the main takeawayJOSH [00:56:38]: for many of the, like, actual performance is, like, once you fix these ambiguous examples, a lot of these benchmarks are really saturated. Like, I think it'sSWYX [00:56:48]: important to look at,JOSH [00:56:48]: like, you know,SWYX [00:56:50]: like when you'reJOSH [00:56:50]: talking about performance on ANLI or race or pool queue or something, what you're really talking about is, like, performance on questions that make no sense. Like, it's just like, did it guess the answer in this, like, really weird scenario? Like, those are the ones that are left.SWYX [00:57:03]: Like, when you lookJOSH [00:57:03]: at the performance on the ones that actually make sense to everyone, all the models agree.SWYX [00:57:07]: We agree, like,JOSH [00:57:07]: everyone's on the same page, which I think is kind of a really interesting result.SWYX [00:57:11]: The question then becomes, you know, what are the new, like, set of evals that would be like the next frontier that often embeds with it your idea of what reasoning is, because it's obviously you're super interested in reasoning. And yeah, I mean, like, where does this, where does the state of evals go from here?JOSH [00:57:30]: This work and this blog post is talking mostly about the public evaluationsSWYX [00:57:34]: and the thingsJOSH [00:57:34]: that we can release. We do have our own internal evaluations. For example, one of them that we are releasing is the code understanding evaluation, which is about predicting,SWYX [00:57:44]: you know,JOSH [00:57:44]: what will this variable be or asking questions about code, et cetera. And that is one of the early benchmarks that we made that we can release. We can partly release it because we can generate an almost infinite amount of this data because these are programmatically generated. And so, you know, we're not really worried about there being like corruption in the kind of the training or test sets. So that makes it a littleSWYX [00:58:03]: bit easier for us.JOSH [00:58:04]: But I think it's, you know, we have built other data sets as well that we can't release. Some of them, you know,SWYX [00:58:09]: for example,JOSH [00:58:09]: because they maybe use other open source code and so we can't redistribute it necessarily. Other ones, because, you know, that's, I think evaluations and data are like a core, important part of, you know, the business. And I think we take evaluations very seriously and are spending a lot of effort in terms of like, what exactly do we make as part of the evaluation set? How do you evaluate these things? We've done a lot of other stuff, you know, since these evaluations. But I think a lot around like code understanding for us, since that's our main focus. And it's a nice place to explore reasoning as well.SWYX [00:58:40]: It sounds like you talk a little bit about like code understanding as like sort of variable level, like sort of very micro context. Is there a sense of like larger code context as well? I don't know what I mean by that, by the way. It's mostly just like if I told the senior engineer to go look at a code base, they would understand at a broad level, the architecture, but also the design decisions and be able to tell me that. I don't know if that's useful or not, but I mean, that's useful to me as a, as someone who might be working with them. Yeah.JOSH [00:59:06]: This particular dataset is like the more low level code understanding,SWYX [00:59:10]: like just literallyJOSH [00:59:10]: what happens in this code. And this is mostly because, you know,SWYX [00:59:13]: this is part of theJOSH [00:59:13]: carbs tuning metric, etc.SWYX [00:59:15]: Like we care aboutJOSH [00:59:15]: the low scale versionSWYX [00:59:17]: of this as well.JOSH [00:59:17]: We want smaller scale models to be able to do something on this. And so that's kind of the focus for this.SWYX [00:59:22]: And hopefully this is moreJOSH [00:59:22]: useful for other people. But yes,SWYX [00:59:25]: those other questionsJOSH [00:59:25]: are also quite interesting. They get a lot harder to evaluate, like, is this a good architecture or not? Like you and I could probably debate for a while on, you know, different architectures. And so it becomes a lot trickier to do these evaluations as they become more realistic. So I think that's one of the things that we've been playing around with a lot, especially around like code generation.SWYX [00:59:44]: So if you're saying,JOSH [00:59:44]: you know, implement this function, okay, it can be kind of objective, but, you know, even MBPP, we've made our own internal version of this data set, right?SWYX [00:59:52]: Where we've taken likeJOSH [00:59:52]: every single exampleSWYX [00:59:54]: and looked at it and been like,JOSH [00:59:54]: does this actually make sense? Like, what is the type signature? Like, can we remove all ambiguity, et cetera?SWYX [01:00:00]: So you basically like reviewed every single question on, I mean, that's impossible for like HelloSwag, right? Yeah, yeah.JOSH [01:00:05]: We didn't do that for HelloSwag, but this is for MBPP, which is only like a few hundred. So we just sat down and did it. Yeah.JONATHAN [01:00:12]: I'm so excited to get to look at this data set. Like this is such a resource for the community. I absolutely can't wait. We should probably do the,JOSH [01:00:19]: I don't know. I don't know if we were planning on doing the healed MBPP one,SWYX [01:00:23]: but hopefully we can doJOSH [01:00:23]: that one in the future. Did you look at SweetBench?SWYX [01:00:26]: It's the sort of hot new data set of the summer.JOSH [01:00:28]: Yeah, I've taken a quick lookSWYX [01:00:29]: at SweetBench.JOSH [01:00:29]: It's really interesting. I like that it's a much more difficult kind of coding, code related task for bug fixing. I think it gets into some of these problems where it is a lot harder to evaluate these things once they get more realistic. Like we were looking at the AgentBench paper, I think just last week for our paper club and one of the thingsSWYX [01:00:49]: that we noticedJOSH [01:00:49]: is that actually like both of the examples in the appendix that are given as like traces where it got it right. This is actually not the right solution. And it's OK. You know, it's fine. Like it did make it past the test. That's what the metric is.SWYX [01:01:02]: That's what the benchmarkJOSH [01:01:02]: is about, right? But like it just said,SWYX [01:01:05]: you know, like,JOSH [01:01:05]: you know, dot encode ASCII. Like, well, that's not the right way to do this. Like it just dropped all the other edge cases that you actually would have cared about in production for this thing.SWYX [01:01:14]: And there is likeJOSH [01:01:14]: a better way of doing it.SWYX [01:01:16]: And you know,JOSH [01:01:16]: that's what the real golden patch was. But, you know, that's OK. But then how do you test all of that?SWYX [01:01:21]: Like as you start to doJOSH [01:01:21]: more realistic things, the test coverage, like getting test coverage over all possible ways of solving these bugs is really hard. Evaluation is the singleJONATHAN [01:01:28]: hardest part of the whole thing. Like I spend a shocking amount of time just telling our customersSWYX [01:01:34]: we need to find a wayJONATHAN [01:01:34]: to measure what you actually want out of the model before you should ever touch a GPU. And, you know, trying to convince my team and me to follow our own advice a lot of the time on that. And I think everybody like on the one hand,SWYX [01:01:46]: it's easy to laughJONATHAN [01:01:46]: at the state of the evaluations that we have. None of them are good. Like if you go read these eval benchmarks, you'll always come awaySWYX [01:01:52]: disappointed.JONATHAN [01:01:53]: And yet they've given us useful hills to climb. And we do seem to be making progress and measuringSWYX [01:01:58]: progress in the field.JONATHAN [01:01:58]: And I think anecdotally, models are getting better year to year. So I feel like people tend to go and get into one situation or the other, like evals don't matter. I'm just going to look at lossSWYX [01:02:07]: or like, you know,JONATHAN [01:02:08]: the evals matter a lot and they're all broken. So what do I do? And I think like a lot of things in deep learning, we have to make peace with just complete imperfection. Like the most successful scientists I see are the ones who are OK operating in a worldSWYX [01:02:20]: where everything'sJONATHAN [01:02:20]: going to be broken.SWYX [01:02:22]: And yet we can stillJONATHAN [01:02:22]: cobble things together and make somethingSWYX [01:02:24]: interesting happen.JONATHAN [01:02:24]: I mean, we were just discussing that with literal infrastructure. And now we're all the waySWYX [01:02:28]: up to like,JONATHAN [01:02:28]: how do we measure whether a model performed a complex coding task correctly? And everything is broken.SWYX [01:02:34]: And yet we're still ableJONATHAN [01:02:34]: to make huge amounts of forward progress.SWYX [01:02:36]: I think that's right, Jonathan.JOSH [01:02:38]: And that the challengeSWYX [01:02:40]: isn't necessarilyJOSH [01:02:40]: making perfect evaluations. I think our blog post here is about going really into the weeds on these to figure out like, what does that look like? And I think one thing is like, you know,SWYX [01:02:49]: as you said,JOSH [01:02:49]: we have been able to make a lot of progress without making these perfect.SWYX [01:02:52]: That's great.JOSH [01:02:52]: You don't have to have perfect evaluations. And, you know, the more interesting work is the stuff that we can't necessarily publish about, which is the imperfect evaluations that we have for actual coding tasks, for example.SWYX [01:03:04]: Like, what does thisJOSH [01:03:04]: really mean as a person? And there, as you said, it's much messier.SWYX [01:03:08]: So it's a lot harderJOSH [01:03:08]: to put it out and say like, hey, everybody use this because there's so manySWYX [01:03:12]: rough edges.JOSH [01:03:12]: It's so hard to like even say, oh, is this even the right task? Is this even the right way to do it? And there's a lot of judgment.SWYX [01:03:19]: There's a lot of intuitionJOSH [01:03:19]: that it comes down to. But yeah, I think that's where it's critical to doSWYX [01:03:23]: if you actually want toJOSH [01:03:23]: make these systems work.JONATHAN [01:03:24]: Yeah, you have to make peace with with living in that in between.SWYX [01:03:28]: Yeah.JONATHAN [01:03:28]: And I think that in some sense,SWYX [01:03:30]: when I hire researchers,JONATHAN [01:03:30]: that's the number one quality I look for. Like, can they be at peace living in a house that is neither clean nor messy,SWYX [01:03:36]: but it's just kind ofJONATHAN [01:03:36]: somewhere in between? And are they OK with that? Are they OK with a few dishes being out on the table and a few clothesSWYX [01:03:42]: being on the floor?JONATHAN [01:03:43]: Or will that drive them insane? Or will they just end up with all the clothes on the floor and like all the dishes out all the time? Like, it's kind of I'm looking for that perfect balance because, you know, we have to operate in this imperfect world. Like, yeah, go ahead and give me the perfect evaluation for programmersSWYX [01:03:58]: or for an LLMJONATHAN [01:03:58]: that is a program assistant tool. Like there is no perfect evaluation. But clearly we've made progress. And so the most important partSWYX [01:04:06]: is just are weJONATHAN [01:04:06]: climbing the right hills? And so this is why I'm so excited to see the ambiguity aspect of this. We often think we have more room to climb on these benchmarks. It turns out we don't. Or it turns out that actually we're climbing, getting good at the benchmark and not actually getting good at the task we care about underlying the benchmark anymore.SWYX [01:04:21]: Maybe the model,JONATHAN [01:04:21]: like this is the famous example where if you get 100% at MNIST, your model must be broken in some way because there are four examples mislabeled, you know, it's it's that all over again. Welcome to this.SWYX [01:04:33]: Yeah, it's the accidental canary canary in this. I think one thing that'sJOSH [01:04:37]: actually really interesting about this also is that, yes, like the ambiguous examples are sort of, you know, not that great from the perspective of these particular tasks that we're evaluating.SWYX [01:04:46]: But actually, one thingJOSH [01:04:46]: that we're very interested in is ambiguity itself. Like, can we detect whether a task from a user is ambiguous or whether you've, you know, completed a task successfully? Like these are actually hard, messy problems, but are really important from like the user experience of using these models. I would much rather have a coding agent that will give me back a thing. And, you know, it's it's actually the code doesn't work like 10% less of the time than some other model, but it will tell me 100% of the time like when it's not sure. Like that's so much more useful if it can communicate like, I'm not really sure about this or maybe there's some errors here. Then just like, here's some code. I have no idea if it works. And so these kind of like, you know, detecting ambiguity and detecting correctnessSWYX [01:05:25]: or uncertainty,JOSH [01:05:25]: I think are really interesting problemsSWYX [01:05:27]: that we're really likeJOSH [01:05:27]: digging into quite deeply.SWYX [01:05:29]: I want to touch on maybe a couple of hot topics in evals, maybe tangentially related, but we're on the evals train right now. So I'm just going to get on that. So ArcAGI, Francois Chollet's hot new thing, it's sort of my take on it is basically it's trying to measure reasoning through an abstract IQ test. Effectively, I noticed that you don't use it. There's a lot of community debate, pro and con about it. What are your thoughts on just more abstract reasoning and maybe ArcAGI specifically?JOSH [01:06:01]: I think we purposely stayed away from the very, like there's BigBench, for example, that has a lot of, I think, to me, feels sort of similar types of tasks that are like very unrealistic. Like, oh, you know, we have books of different colors and then you're going to shuffle them and like which book is furthest to the left or something like, OK, cool, I guess it's neat. It's neat, I think, for us to explore in terms of like an agent reasoning in a larger loop. And we do care about these types of evaluations there. The types of evaluations we're talking about in the blog post here are for getting at, like, does this model in a base model sense, is this working at all? There's no chain of thought in these evaluations. These are just like, go straight to the answer. Does this make sense?SWYX [01:06:42]: Like, is this a thing thatJOSH [01:06:42]: you can answer very quickly? That's what we were selecting for with these evaluations. This is not to say that these are the only evaluations we have. I think the Arc ones are like a little bit too, probably, visual for us to really be able to integrate with.SWYX [01:06:56]: But I think some of theJOSH [01:06:56]: BigBench ones are... You can tokenize it.SWYX [01:06:59]: Yeah, but, you know,JOSH [01:07:00]: I think it's not really... I think you can spend a lot of time getting really good at these kinds of benchmarks without making, like, kind of more general purpose progress. And so I think we're a little bit leery of going too far in that direction. Similarly, like, coding competitions. Like, we do a lot of code generation, but we don't really do a lot on, like, code competition problems for the very, very hard ones.SWYX [01:07:20]: So I think you can goJOSH [01:07:20]: very far down that routeSWYX [01:07:22]: and make something that's, like,JOSH [01:07:22]: really good at those problems, but not actually that useful as, like, a programmer day to day.SWYX [01:07:26]: Yeah.JONATHAN [01:07:27]: Take a different tactic, which is, like, at the end of the day at Databricks, I have 12,000 customers, or I think that's the latest number, all of whom are trying to do something with, you know, LLMs or AI or machine learning. And those things don't look like these tasks. I don't think I have a single customer that's asking to, you know, have AI solve abstract reasoning problems. Things are pretty, like, they can be ambiguous,SWYX [01:07:53]: they can be challenging,JONATHAN [01:07:53]: they can be really interesting,SWYX [01:07:55]: but none of them look quite like this.JONATHAN [01:07:56]: And so, you know, I think to Josh's point, like, it's really about asking, why are we doing this? Even if you're trying to build AGI, and that's not personally my purpose, and I, you know, Josh has much more interesting things to say about that than I do. I don't even know if this is the kind of intelligence I would get excited about or care about personally, or if I would consider, you know, to Josh's point, this to be the indicia of intelligence.SWYX [01:08:17]: It's neat.JONATHAN [01:08:17]: But, you know, for me, it's, like, more down to earth things, like having a model that can have a conversation with you about dataSWYX [01:08:24]: that on the backendJONATHAN [01:08:24]: is running SQL queries on your literal data. That's a much more interesting task to me. That's something that really matters day to day for my customers and, you know, different perspectives, but, you know, I think Josh and I would probably say the same thing,SWYX [01:08:36]: even though I would,JONATHAN [01:08:36]: I'm guessing, I don't want to put words in your mouth. You would say that you're pursuing more general intelligence in your own way. And I would say that I'm very happy with narrow intelligence. Like, I'm very happy with my little SQL bot and building 12,000 of those because they're moving the needle for a lot of folks every day.JOSH [01:08:51]: Yeah, I think we're, you know, we're not as far away in our position as it might seem. I think we're also excited about, like,SWYX [01:08:58]: how do you actuallyJOSH [01:08:58]: make these things useful? And that does end up being pretty narrow. I think these other tasks can be interesting as, like, ways to explore these more abstract reasoning questions or like, OK, how could an agent actually work through this? But it's important to keep in mind that it's like a toy, not a real problem. It's like it's a scientific tool to tell us something about the models.SWYX [01:09:16]: It's not something we shouldJOSH [01:09:16]: be optimizing for necessarily.SWYX [01:09:18]: The one thing I'll point out is, you know, as a kid, I was graded into a gifted program based on my ability to solve these exact type of problems. And then I entered college based on my ability to solve SATs, which, again, have nothing to do with my college experience, but whatever. So, you know, we have a history in the humanity of doing correlated IQ tests to general capability. OK, so the two more, two more viral evals, and then, you know, I just want to be mindful of your time. Needle in a haystack, long context utilization. Oh, for the love of God. Something, well, OK, like, let's just assume that, you know, on our podcast, we've discussed the, you know, baseline problems with needle in a haystack, but just generally long context, right? It's a useful thing for agents. I assume. And it's something that, you know, it's out there. Like, we don't know, don't really know what the best way to utilize memory is. But like, I assume it's important, right? What I'll say is like, you know,JONATHAN [01:10:13]: I spend a lot of time thinking about RAG these days. And RAG, you know, in one sense, you know, the way that I think about RAG is it's the world's simplest agent. It is an agent that basically, you know, there's at least more than one thing happening in the process of building models, at least a system. If you give the model the ability to decide when it wants to retrieve data from a context or retrieve data from a database, then we're talking about an agent. So RAG kind of, I think, like toes that boundary really nicely. There are a lot of reasons why you do genuinely need a long context. Like, I don't think long contexts are problematic in and of themselves. I know there's some controversy even about that. I love the idea of doing like thousand shot tasks as an alternative to fine tuning. I love the idea of pulling in lots of data into the context. I love the idea of once you get in a multimodal land, you're just going to end upSWYX [01:10:54]: with giant context.JONATHAN [01:10:54]: It's kind of unavoidable. The flip side is I don't know of anyone who like is hiding a secret passphrase in a book and needs the model to find it. Needle in a haystack is, it's interesting. The challenge with long context to my mind, and Josh,SWYX [01:11:08]: I'm curious what you think,JONATHAN [01:11:08]: is simply that annotating long context evals is really hard and really expensive, you know, intrinsically, because you need someone to read 10,000 tokens or 100,000 tokens, or like you need someone to read a 1,000 page book or the equivalent thereof in order to measure those long context benchmarks. I don't know if a human could solve these tasks, let alone that a human could do this in any amount of time where you're willing to pay the money to get the data annotated. And so any long context evalSWYX [01:11:33]: has to, in some sense,JONATHAN [01:11:33]: be correct by construction. And you have to, you know, the, you have to know the answer before you've created the example. And needle in a haystack is kind of the simplest waySWYX [01:11:41]: of doing that.JONATHAN [01:11:41]: I think the problems of needle in a haystack are well known, you know, it doesn't measure anything real. You're not even testing the model's ability to holistically use the context just to identify one part of the context. So you can do some wacky things to your model, like quantize the hell out of the KV cache and still get needle in a haystack to work quite well because it's not trying to holistically take advantageSWYX [01:11:59]: of things.JONATHAN [01:12:00]: You know, I have some thoughts on things that I like more that are also still correct by construction. Like, I really like the idea of doing thousand shot tasks where you can look at the scaling as you go from 10 shot to 100 shot to thousand shot to fine tuning on that data instead. And I like that as a way to, you know, have something that's correct by construction, or at least where you haveSWYX [01:12:19]: a nice baselineJONATHAN [01:12:19]: that you can compare to automatically. So I'm typically looking for like contexts that are situations where long context is one way to solve the task, but not the only waySWYX [01:12:28]: to solve the task.JONATHAN [01:12:28]: And we have some other strong baseline floating around personally. But yeah, needle in a haystack, not my favorite thing in the world, to say the least.JOSH [01:12:35]: Yeah, I mean, I agree with most of what JonathanSWYX [01:12:38]: said, I think.JOSH [01:12:38]: I think one other thing that I will call outSWYX [01:12:40]: is that, you know,JOSH [01:12:40]: from like a coding application perspective, it's useful to have long context because the lazy thing of just like throw the whole repo in the context is like,SWYX [01:12:48]: OK, cool.JOSH [01:12:48]: Like, you know, you can just get started with that. But then in, you know, in real scenarios, you don't necessarily want to put the whole thing in there. You can have code basesSWYX [01:12:56]: that are bigger.JOSH [01:12:56]: You probably want to filter down to the stuff that's relevant anyway to not be confusing. Like you probably even if you did have a lot of context,SWYX [01:13:02]: you might want to sort itJOSH [01:13:02]: in some way to say this is more important than this other stuff. So and, you know, you don't want to wait for you don't want to be wasting all this time and computeSWYX [01:13:09]: on inference and likeJOSH [01:13:09]: doesn't really matter. So, yeah, I don't know that it's the most important thing.SWYX [01:13:15]: I think people will find creative use cases. And like Jon said, I think the multimodality examples will naturally lend themselves to long context. Cool. And then one last one on just general sort of agent related capabilities that we didn't really talk about in the eval section is function calling and tool use. There's a recent trend, I think, basically led again by OpenAI on parallel function calling. There's always there's been a limit on how many tools you can call from four to now, I think, 128. And I think theoretically, Claude and Jem and I support a lot more.JOSH [01:13:49]: So just generally,SWYX [01:13:50]: how do you think about evaling tool use? Is that super important for you guys? We're thinking about itJOSH [01:13:55]: in a slightly different way, which is, yes, you can have this like hard coded list of tools. But if only you could have like this really large open set of like tools, maybe they would be like functions that you could call if only there was like a language or like a programming thing, like being able to write code. I think for us, it's like, well, look, if we can write code, like now you have all these tools accessible at the end of the day,SWYX [01:14:16]: like function callingJOSH [01:14:16]: is just a function invocation, like literally in code. I think our approach to this is likeSWYX [01:14:21]: instead of worrying aboutJOSH [01:14:21]: like weird hard coded agents using tools, like let's just make themSWYX [01:14:25]: able to actuallyJOSH [01:14:25]: write code robustly and make that code work and be able to debug that code, know if that code is safe to run, like get really good at the like code writing and execution part of things, because that will open up the action space like far more than, you know, 128 tools, like just everything is at your fingertips, especially I think over the next few years, like we already have so many really good APIs. As we get better and better at writing code, we'll be able to make APIs to things that don't even have APIs today. That's kind of how we think about it is less as like a special purpose thingSWYX [01:14:52]: and more as likeJOSH [01:14:52]: this is one of the reasons to focus on code.SWYX [01:14:55]: On my end,JONATHAN [01:14:55]: the way that I think about this is, you know, I think a lot about how models interact with data.SWYX [01:15:00]: And so for me,JONATHAN [01:15:00]: tool use is really a question of how do you take modelsSWYX [01:15:04]: that are really builtJONATHAN [01:15:04]: for unstructured dataSWYX [01:15:06]: and have them interactJONATHAN [01:15:06]: with structured data? So, you know, and I get the question a lot from my customers,SWYX [01:15:10]: like what do I doJONATHAN [01:15:10]: with tabular data? Or what do I do with like, you know, JSON? Or what do I do? I mean, you name it, like even what do I doSWYX [01:15:17]: with a PDF?JONATHAN [01:15:17]: Because PDF parsing is still an unsolved problem, even in 2024. And the answer, or even just the basic questionSWYX [01:15:24]: of like, should I botherJONATHAN [01:15:24]: to structure my data anymore? Shouldn't I just toss the table? Shouldn't I flatten itSWYX [01:15:28]: and just throw itJONATHAN [01:15:28]: into the LLM context and like let the modelSWYX [01:15:30]: figure it out?JONATHAN [01:15:30]: Answer is no. We've built all these fun APIs and fun languagesSWYX [01:15:36]: and paradigmsJONATHAN [01:15:36]: for dealing with structured data over the years. Just use them.SWYX [01:15:40]: Have your model use them.JONATHAN [01:15:40]: Train a model that can interactSWYX [01:15:42]: with these thingsJONATHAN [01:15:42]: in a meaningful way. Like text to SQLSWYX [01:15:45]: is still,JONATHAN [01:15:45]: or like having a model be able to make SQL calls in the backend is actually like one of the singleSWYX [01:15:51]: most useful thingsJONATHAN [01:15:51]: for my customers. It sounds really boring. Models are really good at it. And it moves the needle day to day.SWYX [01:15:57]: So tool use for meJONATHAN [01:15:58]: really is that like, how do you just interact with structured data sources and take advantage of the fact that you have someSWYX [01:16:05]: prior knowledgeJONATHAN [01:16:05]: about the structure of your data that an LLM would completely flatten away. In many ways, this is kind of one of the, one of my biggest frustrations with the fact that LLMs work well with code. We have decades and decades and decadesSWYX [01:16:17]: of understandingJONATHAN [01:16:17]: about the structure and interpretation of programs. Like I think that's literally the name of a book on programming, if I remember right. And, you know, we have all this theory. We know everything there is to know about programming languages if they're well-formed languages and have the right properties. And yet when we have an LLMSWYX [01:16:31]: work with them,JONATHAN [01:16:31]: we literally just turn it into a token stream.SWYX [01:16:33]: Despite the fact that we knowJONATHAN [01:16:34]: how to parse it. We know, you know, how to do all sorts of, you know,SWYX [01:16:38]: reference, you know,JONATHAN [01:16:38]: disambiguation and things like that. We're still just flattening it into a model and making the model relearn all of these things from scratch. And it frustratesSWYX [01:16:45]: the hell out of me.JONATHAN [01:16:45]: I don't have a better answer when it comes to code, but I really appreciate that with a lot of data sources that have structure to them. Tool uses and function callingSWYX [01:16:53]: are just,JONATHAN [01:16:53]: in my mind,SWYX [01:16:55]: So I think basically what you're saying is like code is the God tool for Jonathan. Like, you know, SQL is so much the right abstraction for accessing all this data. One thing I do spend a lot of time thinking about is for the stuff that doesn't fit in a SQL table, you know, is knowledge graphs the answer? I think a lot of people are exploring that and I think every now and then people get a bout of knowledge graph religion and then it kind of doesn't work out. So I wonder, I wonder what the end state is. Like, is this an idea where it's a mirage? Or is this the idea where it sometime is going to work? It's about having the right toolsJOSH [01:17:27]: for the problems, right? Like as Jonathan was saying, SQL is sometimes definitely the right tool. Like you've got your, you know, order table or something and you want to know, you know, number of sales last month. Like you should be using SQL sum that column. OK, great. You're all set. Knowledge graphs also,SWYX [01:17:40]: you know,JOSH [01:17:40]: are sometimes the right tool for a particular problem. You have some like weird question about relationships between entitiesSWYX [01:17:46]: that are modeledJOSH [01:17:46]: on some particular ontology that you actually understand and it's like math to the real world. Great. Use a knowledge base. Like use a knowledge graph. This is fine. But I think in the real world, it gets a lot messier than like knowledge graph style of things where it's like, well, is there a relationship between these two nodes? Like, I don't know.SWYX [01:18:04]: Like, is are theseJOSH [01:18:04]: two separate nodes? Like those kind of messy borders, I think, prevent itSWYX [01:18:08]: from being a toolJOSH [01:18:08]: that can like solve everything forever. And so I think it'll always be good for certain problems, just like SQL is goodSWYX [01:18:14]: for certain problems.JOSH [01:18:14]: Like different abstractions are good for different problems. And yeah, I think this is why I'm excited about code. Like code lets youSWYX [01:18:20]: kind of pick the right,JOSH [01:18:20]: like let's use this library for this problem.SWYX [01:18:22]: Let's use this libraryJOSH [01:18:22]: for this other problem.JONATHAN [01:18:24]: I think Josh said it and you said it well, like code is kind of the God tool. It unlocks literally everything. The challenge for me is always like,SWYX [01:18:31]: you know, sometimesJONATHAN [01:18:31]: unlocking too much power can sometimes inconvenient things can happen. And so it's all about balancing thatSWYX [01:18:37]: in some sense,JONATHAN [01:18:37]: language is the God tool.SWYX [01:18:39]: If only, you know,JONATHAN [01:18:39]: we knew how to interpret it all the time. So code is has the really nice propertySWYX [01:18:44]: that at least you canJONATHAN [01:18:44]: always execute it. And sometimes you just literally want your model to be able to do SQL calls and nothing else. And setting those boundaries properly for the problem,SWYX [01:18:52]: I think is going to be, I think at least a lot of my customersJONATHAN [01:18:54]: are going to be thinking very hard about that.SWYX [01:18:56]: Like, should I giveJONATHAN [01:18:56]: the model access to the web?SWYX [01:18:58]: Is that actually helpfulJONATHAN [01:18:58]: for this problem? It sounds great to just like flip yes on all the tools.SWYX [01:19:02]: Is that actually going to meanJONATHAN [01:19:02]: I'm going to get better solutions to my problems?SWYX [01:19:04]: So I want to be mindful of time. I think that's basically our sort of recap of our discussion based on Imbue's releases today. I wanted to leave some time for what's next for both of you guys. Maybe Josh, as a guest of honor, you want to go first as to what happens next.JOSH [01:19:19]: We have these releases. We're happy to put these things out. I think there's a lot of stuffSWYX [01:19:22]: that we haven't released.JOSH [01:19:22]: Like, this is not the only thing we've been working on. Most of our actual focus has been on kind of coding and reasoning. In particular, like the things that we're excited about are can we make these things useful? Like Jonathan is saying, right? Like, it's not about toy problems. It's like, can we use these today in our day-to-day workflow and actually have them accelerate us? And I think we have some kind of internal product prototypes and things that we're excited about. And so we're excited to share more about this in the coming, you know, months to quarters as we get it to a place where like other people could maybe get value out of this as well. But that's kind of our real focus right now is like, how do you take these really cool capabilities that are out there that our models have, et cetera. And like, make sure that they're actually useful today for us, like when we're doing real work and then for other people as well. In particular, focused on generating code, understanding code, testing code, verifying it, like starting with the like robust creation of software. Excellent.SWYX [01:20:13]: Jonathan?JONATHAN [01:20:14]: I never like to talk too much about the future because I think you've heard this from me before. I like for us to speakSWYX [01:20:19]: through our work.JONATHAN [01:20:19]: And so I don't, I don't like to tease too much. Our mission is, to Josh's point, to make this stuff useful to 12,000 customers. And not a lot of that ends up making it into the public eyeSWYX [01:20:30]: and not a lot of thatJONATHAN [01:20:30]: ends up getting released open source. So for this kind of forum where really, you know,SWYX [01:20:34]: where we're talkingJONATHAN [01:20:34]: to the community, I'm asking myself right now, like, you know, what exciting thingsSWYX [01:20:38]: are we going to haveJONATHAN [01:20:38]: to offer the community in the next little while? I think the most exciting part is just we're writing a lot of blog posts right now. We're trying to share more and more of our science because I feel likeSWYX [01:20:47]: we've been doingJONATHAN [01:20:47]: these big pushes to create these really giant models.SWYX [01:20:50]: I think, Josh,JONATHAN [01:20:50]: I'm sure you hadSWYX [01:20:51]: the same experience.JONATHAN [01:20:51]: It's exhausting and all-consuming and you get to the endSWYX [01:20:54]: and you're like,JONATHAN [01:20:54]: oh, I have all this stuffSWYX [01:20:56]: I want to talk about.JONATHAN [01:20:56]: Now I need to find the time to talk about it now that I've survived this huge push. And we're definitely in that mode right now. So there's going to be a lot of that coming in in the next little while. And, you know, we're always cooking up fun new models. I think the real question is, you know, releasing models open source is not our day-to-day bread and butter. It's kind of a fun reward that we get to do sometimes when we have something really cool to share and a little bit of time and spare GPUs in our hands. But for the most part,SWYX [01:21:20]: everything is goingJONATHAN [01:21:20]: toward customers. You know, I think the joke is Databricks has been 18 months away from IPO for five years. So I guess DatabricksSWYX [01:21:26]: is 18 months awayJONATHAN [01:21:26]: from IPO still. But 18 months away from IPO means there's a lot of pressure to deliver for customers. And we're going to keep working on that. But I think you'll see hopefully some cool, interesting thingsSWYX [01:21:36]: get dropped over the courseJONATHAN [01:21:36]: of the summer and into the fall. We'll find out when we get there.SWYX [01:21:39]: I think that's the right wayJONATHAN [01:21:39]: to put it. I know we were talking earlier about kind of Abracadabra and Alakazam. And all I'll say is that, you know, the DBRX small model that we still haven't released yet was called Abra. DBRX was called Kadabra. And there's a third Pokémon in that evolution. And that's all I'll say for now. Cool stuff kind of popping up sometimes on Chatbot Arena. And, you know, keep your eyes out. Yep.SWYX [01:21:59]: I'll leave the links and the hints in the show notes. That was a very fun way to leave some breadcrumbs for people to follow. Cool. I'll leave everything to sort of some calls to action. We're going to be releasing this next week. So I'll be deep in my conference, the AI Engineer World's Fair. So people can just go to AI.Engineer and livestream it. Do you guys have any other calls to action before you wrap?JOSH [01:22:20]: The only one is, you know, we're definitely hiring. So if you're interested in working on coding, reasoning, interested in working on all of this stuff, you know, from the ground up and really deeply understanding not just how does the hardware work, but how do the models work and also designing these, you know, systems to actually be useful for yourself day to day, come say hi.JONATHAN [01:22:36]: The only thing I'll say is, you know, and I like saying it these days, it feels like the field is so crowded and, you know, it requires so many resources to do impactful work. And, you know, on some days it feels like everything's been done or somebody else is doing everything before you can. At least I remember that feeling every single day of my PhD and even more so now. But I hope like what you heard from Josh today tells you there is so much enormously impactful work to do in the field. If only you take a step back and take a fresh look at some of these things and just talk about what you're doing. There's a huge amount left to do here and a huge amount of exciting work happening every day. And for those who are certainly feeling that exhaustion right now, and I count myself among those folks many days, it's refreshing to see these kinds of drops and see that there is so much more even in things that people feel like they understand how to set up a cluster. My God, you know, even in these evals that we think we understand, there is still more to understand and still more work to do. I hope everybody's keeping at it.SWYX [01:23:32]: All right. Keep on keeping on. Well, thanks so much for your time, guys. That was a great discussion and we'll put the links in the show notes for people to read more. Thanks. Thanks a bunch.JOSH [01:23:40]: Thank you so much. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
Transcript
Discussion (0)
Welcome to the Late in Space podcast, another super special edition.
Today we have sort of like a two-header, John Franco from Mosaic Databricks, or Databricks, Mosaic.
And Josh Albrecht from Mb. Welcome.
Hey, glad to be here.
Thank you for having us.
Hey, so both of you are kind of past guests.
Jonathan, you were actually one of the most popular episodes from last year talking about MPT-7B.
Remember the days when we train large models under a 7B.
Yeah, back when reproducing Lama
17B was considered a huge accomplishment for the fields.
Those are the good old days. I missed that.
So things have accelerated a lot.
Actually, let's do a quick catch-up, and Josh, you can chime on in as well.
So Databasex got acquired.
I talked to you at New York City.
Got acquired, although sometimes it feels like Mosaic acquired Databricks
because, you know, we're having a lot of fun being here.
But, you know, yeah.
Yeah, I mean, you're chief scientist now of Databricks.
Chief AI scientist.
Careful with the title.
As much as I would love to understand how Spark
works. I'm going to have to defer that to much smarter people than me.
Got it. And I don't know about what you would highlight so far as post-acquisition,
but the most recent news is that you guys released DBRX? Is that the thing that most people
should be aware of? Actually, that's no longer the most recent news. Honestly, the most recent news,
we announced that it was at our data and AI summit last week. So it was announced among like
100,000 other things, is that we finally released our text image model, which has been a year in the
making through a collaboration directly with shutter stock. There was a lot of work put into
finding a data set that we were comfortable with working on and trying to build a model that
honestly, I felt like I could trust and that others might be able to trust to put out in the
world. So that model was released last week. It's unfortunately just available via API due to the fact
that the data is quite sensitive and quite valuable. It's shutterstocks entire business in a lot
of ways. But I'm still really excited that there's now a model that is trained on a dataset where
the provenance of every single image is known. And it's a damn good model. So I'm really
proud of the team on that. Yeah. Amazing. Josh, you have any thoughts on
image model questions. That is not my
source expertise, but I'm very, I was excited
to see the release of it last week as well.
And I'm very happy that you guys did a nice
job on the data side of everything there. So
that was cool to see. I think what's unusual
is like, I think Shutterstock's doing multiple deals
and multiple labs. So what
is the Shutterstock model? Like, I guess
is this the house model for Shutterstock?
Is this Databricks's version of the Shutterstock
model? Like, what is this?
The way that I would think about it is that
Shutterstock is doing an amazing business
and AI across the board. Their data set is
kind of widely known to be the best stock photos data set in the world, the most comprehensive,
the biggest. When you think about what data set am I going to train a multimodal model on,
you call shutterstock. And I at least I've heard in the news, like opening I, Google,
Meta, Apple have all called shutterstock and made those deals. So a lot of models have had
shutter stock data incorporated into them. But this is the only model I know of so far where it was
exclusively and specifically trained just on the vanilla shutterstock data. There was nothing
else mixed in. We didn't go and scrape the web and find other data or combined data sets or
or anything like that. And so this is in some sense the house blend. But the other piece is that
it's just a data set where the provenance of every image is known in public. Where did the data come from?
It is the shutter stock collection. That's it. You know, nothing less, nothing more. And certainly
being at Databricks, if I've learned one thing, it's, I've learned about enterprise customers and what
they want out of AI. And one of the things they ask for most is just, what can you tell me about
the data the model was trained on? And here, especially for text to image models where
images are just tricky subject matter.
There's been a lot of kind of legal conversation about images especially.
It's nice to just have something where I can point to it and say, you know,
if you want to know where the images came from, these are what they are and this is how they got there.
I will talk a little bit about Databricks because it's relevant to the rest of today's episode.
So Databricks, sorry, I keep miss speaking.
It's DBRX.
DBRX actually, there's been a pronunciation update.
It is now D.D.R.
So we have decided to add a dinosaur mascot because what model doesn't like a
So literally, I wish I could pull it up.
There is a little plush dinosaur that we had made.
It's like the world's cutest dinosaur, but it is the official mascot of DB Rex.
And there's a little dinosaur logo that, you know, you'll probably see around a little bit more.
Because DBRX is a mouthful, but DV Rex, like, you know, it's just kind of...
Rolls off the tongue.
I love mascots.
Like every company should have a mascot.
And I think Hugging Face got it, right?
You need an emoji mascot because that's the minimal viable image.
I probably shouldn't talk at all about, you know, Velociraptor, but, you know, that's
maybe that's something we can talk about later in the summer.
I'll just leave it at that.
Okay, that's a hint to names.
I feel like your names leak a lot of alpha.
So just to quickly cover the headline details, DBREX,
as a major experts model that's fairly big,
132 billion total parameters of 36 billion active on any input,
pre-trained on 12 trillion tokens of texting code,
and did really well on e-vails to the point where you had to dye your hair blue.
That's my high-level conclusion about it.
Never make a bet with your team two weeks out from model launch,
even when, you know, human eval is looking quite bad.
Because if you set some bar, even if it's arbitrary and you think there's no way in hell they're going to hit it,
apparently money doesn't motivate people anymore.
Humiliting their boss motivates people.
So, Josh, you should really take a hint from this.
You know, you cannot pay someone enough money to make up for you dyeing your hair blow.
I'll keep that in mind for our next model.
It works.
So speaking of Embu's next model, perhaps, Josh, you want to actually just say hi to the general sort of
the in space audience and talk about what we're releasing today.
Yeah.
I'm Josh, CTO Evan Bue, and we're not releasing the model.
We're not releasing the weights, but we are releasing a bunch of different things that should make it easier for other people to make their own models.
So I think right now, training foundation models from scratch is like a very difficult, time-consuming, expensive, kind of risky endeavor, especially for smaller companies.
And the things that we're releasing hopefully make that at least a little bit easier.
So the things that we're releasing fall into kind of three different buckets.
one is infrastructure and scripts for dealing with the kind of hardware and hardware failures
and understanding how well is the actually lowest level of thing actually working
so you can actually do your training at all and at a reasonable speed without having to
constantly restart, etc. So infrastructure and training scripts,
a second set of things is around the evaluation. So after you've trained it,
like how well is this actually working and how do you know how it's working or releasing a whole
bunch of different data there, a new benchmark about code reasoning understanding, as well as
our own private versions of 11 different open source benchmarks, so things like pool queue or A&LI,
where we've gone through and kind of cleaned up the data as much as possible by looking at all
the ones that models get wrong or that are flagged for ambiguity, and also our own kind of private
reproductions of those, where we've done like a kind of clean room black box, like, okay, this is what
the data set is supposed to be. Here are some examples. Let's make our own version of this to make
sure that there is no data contamination, etc.
To make sure that we're actually, you know,
not testing on train.
And then I think a final thing that we're releasing there is around 450,000 human
judgments about ambiguity and question quality,
which we used in the process of cleaning these evaluations.
And we also hope we'll be helpful for other people training kind of similar models.
And then the third thing is carbs are hyperparameter,
our cost-aware hyperparameter optimizer,
which was especially helpful for being able to experiment in much smaller scales.
and then scale those experiments up to the much larger scale,
kind of on the first try without having to retry.
You don't want to be training, you know, 10, 20, different 70B models.
You really want to get these larger models right on the first try.
And so the ability to kind of tune things very precisely and learn scaling laws,
not just for, you know, the like data and and flops,
but also for learning rate and all the other hyperparameters and see like,
how should you scale these things up was extremely valuable to us as we were training
the larger models.
Yeah, that's a lot of stuff.
Yeah, exactly.
There's a bunch of stuff we'll have to go through all of it.
Yeah, I just want to throw in how excited I am about this.
This is the stuff that nobody ever talks about.
That is the difference between success and failure in this stuff.
Like, can you get your cluster to run?
Can you get software on your cluster?
Can you figure out what broke?
Because fault tolerance is still not really built in to any of the fundamental primitives
of training models.
And so if something breaks, you have to go figure out what broke.
Your job stops.
You have to restart your job.
It is a nightmare just to get to the point where anything can train on the cluster.
A basic MPI hello world that has the GPUs talk to each other is hard enough, let alone
actually training a model, let alone getting good performance out of the GPUs, let alone actually
getting a model that converges to anything interesting.
Like there's so many levels of things you have to accomplish.
Like, this is the kind of stuff that matters.
You know, I think to a point that Josh made earlier, you know, before we got on here,
there are plenty of weights out there.
Nobody's released this.
Yeah, that was part of the motivation actually is that.
There are lots of other things that are complementary,
but I have not seen nearly as much discussion
about some of these other things that we think are pretty important.
I mean, in some sense, I'm very excited to have Jonathan on
because this is a little bit, you're bread and butter with mosaic.
And I think you've released some part of it with composer.
And I think it's just really, really interesting to see,
like a different take, like basically a full stack take that's kind of open source today.
Yeah, it's really kind of, it's been an ordeal.
to figure this out. And every time something changes, whether it's a new GPU or even a new driver update, you get new creative errors and new things go wrong. And, you know, we've dealt with the weirdest things from, you know, our infiniband cables getting stolen from the data center twice, like in boxes before they arrived at the data center. Like, you know, porch pirate basically had stolen our infiniband cables back when those were hard to come by to like, you know, weird recalls of switches to like the strangest stuff has happened. I have, you know,
my favorite GPU failures I've seen, like ones where the GPU doesn't fail, it has a correctable
memory issue, and the memory correction costs the GPU to become a straggler and hold up the whole
job. Like, weird stuff happens. And figuring out how to not just identify all of that, but then
eventually productize it is in some sense the entire story of Mosaic and now Databricks in terms of our ML
offering. Really, the thing we offer is we have gone through this suffering and figured out how to even
productize that. It has been a pain in the butt. Yeah, it's a lot of work. I think my favorite failure
was GPU is just giving wrong math.
Like, if they give errors, great, because you can see the errors.
But if they just give you the wrong math back, not so fun.
When did they give you wrong math?
Like literally, you could just, you know, add two things, for example, the numbers come back.
They're not the numbers that they're supposed to be.
I think it's important to stay at this stage just because, like, I think it goes without saying
for Josh and I, but it's worth saying here.
This isn't to say that, like, anything is wrong with us.
It's not like Nvidia did a bad job or, you know, Melanox did a bad job or the
like the server builder, the data center operator,
the cloud provider,
like the million other parties that are involved in building this,
we are running these insane chips that are huge and complicated
and built on tiny transistors at insane frequencies with insane heat
in data centers that for the most part were not built remotely for this kind of power
or heat and have been retrofitted for this.
Failures happen on a good day with normal CPUs,
and this is not a good day and not a normal CPU for the most part.
It's fun to joke about all the weird things we see.
This is not to say anybody's done anything wrong.
This is just part and parcel of working on a massive cluster running at multiple
megawatts of power at a time.
It's crazy.
But optical people, like all sort of like everything.
I'll take the opportunity to start going to the infra piece.
There's just like a disarrishment of the infra just to give people a sense of what we talk about
when we talk about massive clusters.
So I'm just going to read off the blog post here.
This post is about one cluster that has 4,092 H-100 GPUs, spread across
511 computers.
They use Unified Fabric Manager nodes,
which manage the Infinity Band Network,
and you talk a little bit about your networking.
Is there anything unusual about this setup
that you'll call out to people?
Yeah, actually, this particular cluster is a little bit non-standard.
The normal, like vanilla setup,
or, you know, these large clusters, as vanilla as it can be,
is what's normally like a 127 node cluster,
so closer to like 1024 GPUs instead of 4,000.
here we have a larger cluster. As you start to get into the larger clusters, the networking becomes a little bit more custom. It's a little bit more trickier. It's a little bit more difficult to get these things to all be able to talk to each other at the same speed. And so this has in this particular case, is the three tier network architecture instead of two tier is kind of the normal one. So most of the clusters are a little bit smaller. As you get to even larger scales, then it becomes, this becomes even much more complicated, much more expensive. So we chose this particular scale kind of knowing our own workflows and kind of what we wanted to do. This was kind of the right size for us. But yeah, I think it's not exactly.
Exactly vanilla already is already getting in kind of the custom territory.
So my understanding is that there, and it's any part of this that comes with the Voltage Park deal that you guys had?
Is that part of the hardware that you got from the deal with them?
Yeah.
So we worked really closely with Voltage Park to set up all their clusters and infrastructure and everything and kind of decide even like what to order.
How should like, how should the networking work?
Like we were very involved in kind of the construction and bring up of this.
And that's what this post is about is about that process of like bringing up all these.
There's like different clusters in different places of different scales.
So in this particular post, we're talking about this one, 4,096 GPU, but there are other clusters that they have as well.
And we were very closely involved with figuring out the exact architecture and kind of the tradeoffs that go along with picking, you know, those exact components.
You really don't want to like place the wrong order because it takes months to get it and it's very expensive.
So yeah, we were happy to help out with that.
And then you're going to get stolen.
Yeah, yeah, exactly.
We wanted to make sure that we ended up with compute that would work for us.
And that would also work for their other customers.
And so we kind of help design something so that we would get exactly what we were looking for.
We knew that these kind of details would be super important and that getting down to the level of the hardware and like having these good scripts and everything was going to be a core part of like actually getting this to work.
I'm very glad that we did that.
I don't think that most companies kind of take that full stack approach.
But for us, it certainly paid off.
Yeah, it's basically sort of built to spec.
It's interesting that relationship because usually for the rest of us who don't operate at your scale, we take whatever you can get from cloud providers.
but you are basically co-designing from the single machine up.
And you describe that a little bit.
You want to take us through the process that you're described here?
Yeah.
So for the actual, like, the blog post and kind of bring these machines online.
Yeah.
So yeah, I think the process, as we haven't broken down in the blog post, there's kind of a few different layers.
First is like getting the individual machines to work at all.
And then getting the machines to actually be able to talk to each other.
So getting the infinite van networking to work.
And then getting to a point where, you know, not just the machines are working and they can talk to each other,
but everything is actually working correctly.
There's a big gap between like it's working at all to it's working perfectly correctly.
And then after you have all this stuff working perfectly correctly, nice and healthy,
then now you get into kind of the software data like training issues.
And then after that, you're still not done.
Like now even once you're training at full speed, things are going to fail over time,
things are going to change.
There's going to be new firmware updates.
Like how do you kind of deal with this change and flux over time without going crazy and
pulling your hair out, trying to like reproduce things or understand why they're
regressions. And so there's a lot of work to kind of automate the infrastructure tooling as well.
And kind of the first step, like, bringing these things online in the first place, you know,
you have hundreds of machines at this point. So you don't necessarily want to be like walking
around with like a CD-ROM or a USB drive, like plugging it in with their, you know,
keyboard, like hitting next, next, next, on the LS install. That's not how this works.
You do that for one machine. And then you use, we use this thing called Metal as a Service to bring up
all the other machines. So it's a kind of server that can kind of insert.
stall the operating system on these other machines.
So most like when you're talking about these machines,
like,
each machine is,
you know,
on the order of hundreds of thousands of dollars.
So they usually come with a kind of out of ban management interface as well.
So they don't,
they have their infinite band networking.
They have their normal 100 gigapet per second Ethernet networking.
They're like dual redundant,
et cetera.
And then you also have this extra out of band management network.
So you can log in and you can see like the boot screen or you can see the blue screen
of death.
You can like get in there and actually see what was,
wrong, which is pretty fun, and it makes it possible to automate a lot of this work.
So the beginning of that, and the blog post goes into much more detail about, like,
exactly how we set these up and kind of the other errors that we ran into.
When you're bringing these online, you'll definitely have failures.
Even if they all worked in the factory, they get shipped, some parts can lose, something
fails, something goes wrong.
So when you're bringing them online, there'll be some that don't quite work for all sorts
of reasons.
As you start to be working it with machines at this scale, like, you know, if something
happens one and a thousand times, you're like pretty likely to see it.
And so you can get pretty rare, weird things, especially since we had fairly early builds and fairly early versions of this hardware.
Like these are some of the first machine that were ever produced, some of the first GPU.
So you got some extra special things there.
We definitely worked with Dell, for example, on making fixes in the firmware level to be like, okay, like this thing is wrong.
Like, we need to update this at the firmware to like actually fix this particular thing.
So we worked pretty closely with Dell and Nvidia.
Yeah, that's what I'm saying.
Like, this stuff gets complicated.
And the thing is, like, you know, taking a step back, the whole reason we're doing this, right, is that we knew that this was going to be complicated.
There would be these kind of failures. And if we're just using, you know, AWS or some other cloud provider, these errors are still going to be there.
And you're going to have no way to know and no way to debug this and no way to diagnose what's going wrong.
And so we would much rather be able to like call up Dell and say, hey, this isn't working.
And they're like, yep, okay, cool, we'll see if I get together. Oh, I see. Yeah, cool. We'll ship a firmware update and actually fix this for you.
That was a much better experience than like, great, just magically fails. I guess we restart and hope this.
that machine goes away, like that's not a very good place to be.
So yeah, that's kind of the first place is getting to a place where like GPU training is
working on your single mode machines. You can observe stuff. We have tons of tooling around like,
you know, Prometheus and all sorts of other tools for understanding what's going on
these machines because you don't want to be like logging into each one and looking at the temperature
or something. You really need to have tooling to collect all these metrics, etc.
Unfortunately, all of the scripts that we have for this are like for this entire cluster and
for all the simple structure are a little bit like special purpose for our particular
thing. So it's not that every script that we have, it's not, you can just like take this and
plug this in. Even if we did open source all the tooling that we have, you'd still have to do
like a lot of work to open source it. What we are releasing is as many of the things that we can that are
going to be useful for other people. You're still going to have to have some way of kind of managing
these things, making your own like logging aggregators, et cetera, et cetera. So that's kind of
bringing them up to the like, you know, the single nodes are working. From there, it goes into, I
happy to keep going if you want. Well, I just want to leave the opportunity for John to comment if
there's anything that's different from how he runs things. Oh, I mean, all I'll say is I'll endorse this
and say, this shit is hard. Like, this is really, really hard. And, you know, I have a special
props to, you know, the folks in view because they were building this from the ground up, you know,
at Databricks and at Mosaic, we typically work with cloud providers because some of this stuff is just,
there's too much to handle. It's complicated. There's a lot to deal with. And
this doesn't even get into things like physical security, you know, securing power if you're the
data center operator. Like, this gets infinitely complicated and you have to abstract somewhere.
Like, you know, and then you get to the folks who are literally building their own custom chips and like,
good God. Oh my God. That's, you know, if you're one of those folks, you're having, you know,
pour one out for the, the infra people at some of the AI chip startups who are having a really,
really interesting time right now. But this stuff is really hard. And I don't think we talk about it much
because there are so many other things that are hard.
But the other hard things, I think everybody's becoming pretty familiar with at this point.
This is something that I don't think there's ever really been a comprehensive discussion of,
at least not that I've seen.
Yeah.
So my impression is that you guys, Mosaic, have your own software for sort of spinning up and down machines
just like In Bue had to build.
But Bue probably, it sounds like in Bue, you guys went fuller stack.
I don't know how to describe it.
Like, Mosaic is not working with Dell on like their firmware.
No, no.
We're typically working with like, you know,
pick your cloud provider on their Dell firm,
where or what have you?
Like, it's kind of, I think one of the things,
I don't know, Josh, you can correct me on this.
It's kind of impossible if you're doing training
to not go all the way through the entire stack,
regardless of what happens.
Like, somehow I'm still chatting with cloud providers about power contracts,
even though the whole point of dealing with the cloud provider
is not tough to think about power contracts.
Somehow I'm still asking them about which infiniban provider they used this time
to see if this is part of the bad batch of cables I encountered on that cloud
provider or what have you or like we're still talking about a firmware update from pick your
provider like you can't not do this it's convenient that they have data center staff we're worrying
about what to send back to which provider when and they have people who can go and wait for the
infinibin cable so they don't get stolen outside but you know it's kind of it's impossible not to
really go full stack if you're thinking about the infrastructure at all i don't know josh correct me
no i think that's right that's what we expected from the beginning as well is that we would
have to get inevitably have to get into the details here, and I'm glad that we kind of just planned for
it. I think it made it a lot easier from our perspective to have direct control over this.
Instead of having to go to the cloud provider that goes to the data center that goes to the supplier,
we could just go direct to Nvidia or Dell or the data center, whoever was responsible and be like,
hey, this thing needs to change. And they're like, oh, okay, yeah, that is our responsibility.
Great, we can fix that. So it was just a lot easier for us to fix these bugs than if we had to go
through an extra layer of email. Something we discussed in the pre-show was that,
you had a rule of thumb for your cluster of reliability.
You say here in the post, by and large, you expect around 3% of your machines to break every week.
So you're basically going to turn through all your machines in the year.
As it says in the post, so that would be true if it was a uniform failure like that.
But as it says in the post, like it's usually these kind of problematic nodes.
And to be clear, that is the number that we've heard from other people is like they're having about 3%.
I don't think we're experiencing failure rates that are that high.
I think ours is actually quite a bit lower than that.
probably because we've taken the time to like dig into a large,
maybe larger number than we should have of these failures and get to the root cause of it and be like,
oh, okay, like that's exactly what's going wrong.
How do we fix this?
How do we prevent this from happening?
How do we make automated checks for this?
So that if it does happen, it just goes back to the whoever owns that particular part of the process and they can fix it immediately.
And that's part of what you're also open sourcing, which is the health checks, right?
You got the Nick health checks, GPU health check, this space health check, Docker, D message.
I don't know what that is.
That one is...
Just a lot of stuff.
Yeah, that one is one where we realize that actually, like, when these machines boot,
sometimes they wouldn't actually boot cleanly all the way.
Or when they rebooted, they had problems that they didn't have when they were working before,
which was kind of frustrating.
Like, usually if you restart your computer, it gets better.
Here you restart, it did not get better.
It got worse.
That was very frustrating.
So this health check looks at every particular line we've ever seen from the boot, like,
in D message, like every single log line that your computer emits and says,
like, have we ever seen this before? Is this expected? Is this in the right order? Or is there something
out of place? If there's anything out of place, and we say, okay, great, like, now it goes into
this, like, longer, more triage list of like, all right, great. Like, is this acceptable? Should we flag this?
like, should someone take a look at this? So we're looking down at a very, very granular detail level
what's happening on these computers to make sure that nothing is out of place. And that's critical
because without that, if you're running your training, as Jonathan said, and you're, this thing is
slow, like, what are you supposed to do? Right? Like, you really, you really want to be very certain
that all 4,000 of these GPUs are working like they're supposed to.
We know that.
And so if it's slow, it's because, like, we messed up the config or something else
and not because of this earlier thing that's, like, really hard to detect in software later.
Yeah, I think they, I'm just curious to ask, like, you know, suppose you were to set up
another, let's say, another H-100 cluster and it were at a different data center,
and instead of the vendor being Dell, it was super micro or what have you.
How much of this would be repeatable and how much of this would you have to redo?
I, you know, I genuinely don't know.
A decent amount.
I think it would go a lot faster the second.
time. I think there's lots of learnings that we had. And also the blog post, you know,
yes, we are releasing the health checks, releasing some scripts, but a lot of the valuable stuff is also
in the blog post itself, in the details and kind of the learnings that we've had and the sort of errors
that we run into. We tried to as much as possible surface those so other people could learn from
those and avoid the same mistakes or failures as well. But I think it would go a lot
faster. Although, yes, there would certainly be some things that'd be a little bit different.
I mean, there'd probably be different CPUs or whatever, but I think a lot of that stuff is less,
it's less, that's the, like, that's less variable.
I think most of it would apply the second time around.
Although I'm sure next time we're building one, it'll probably be, you know,
at a scale as 10x as big with a different chip or something like this and then who knows.
Yeah, with Connect X8 that will have its own fun behavior and all that good stuff.
Yeah.
Perhaps there's something that people don't discuss about, and you don't even talk about this in the blog,
but I always wonder is what is the timeline that's like kind of reasonable for this amount of work,
at least the initial stages?
And also what does the team composition look like?
for setting up a cluster, right?
Like, what are the mix of skills that you typically would require to get all this going?
I can't really speak to typical.
One thing I am very proud of is how much we accomplished with such a ridiculously small team.
Our infrastructure team is like, you know, fluctuates from week to week depending on like how many
things are on fire and how much we need to build.
But it's like between like three and six people.
Like it's small.
It's not like some huge team of like tons and tons of engineers.
But those people are very, very good at what they do.
and so that has allowed us to get a lot of mileage out of these things.
I think it's not that we're building everything, right?
It's not that three to six people build this whole thing.
I definitely want to say thanks very much to Dell and H5 and Nvidia and the other people
that have done a lot of the work.
To bring up this cluster with 4,000 GPUs and a three-tier networking, networking architecture,
you have 12,000 cables.
So that's 24,000 things that need to be plugged in.
That's just a lot of stuff to plug in.
And you don't want to mess it up.
like each one needs to be done correctly.
Like it's a little bit loose.
Like it doesn't really work.
If you break it,
you need to replace it.
Like,
there's a lot of work that goes into this.
Yeah.
And then,
you know,
that's just like,
that's if you were to do everything right the first time.
And if you didn't have to fix anything.
But inevitably,
you know,
you will have to replace something,
which means like taking all the wires out,
pulling the thing out,
taking all the GPUs out,
going and fixing some cable,
putting it all back correctly,
putting it back in,
doing this every time.
Like,
there's a lot of work that goes into it.
So there were a lot of people at Dell,
Nvidia and at H5 that all helped
a ton with this stuff. I don't know the exact
size of the Dell team. It also fluctuated
over time. Yeah, excellent.
And then, you know, you
have all the hardware set up and now
you're firing it up for a single node. There's a long
description that you guys have about
just like monitoring the MFU
and what each
situation might look might be
indicative of. One of the most
interesting things to me that I
saw from here is like, you know, if training immediately
starts off at 60 to 80% MFU,
something's wrong.
But like, you know, like what are like, you know, some anecdotes or, you know, notable
scenarios here that you might call out as maybe counterintuitive or super interesting.
I mean, there's just so many of them.
I mean, one of them, which I think is probably pretty common knowledge by this point.
But like, we did have a sort of like, which one was this exactly?
I think for the MFU, like gradually getting worse over time.
I think that one when we saw that the first time,
we're like, what the heck is going on?
Like, why does it get just like a little bit worse?
This is so strange.
Like, what is it getting lazy or tired or something?
Like, is it heat?
Like, what's going on?
And in this particular case, it was memory, fragmentation.
Because you have hundreds of machines,
they're doing garbage collection at slightly different times.
And then they get slightly further apart and slightly more and more jittered
until eventually they're all happening kind of random times
and just like really messing up each one of your steps.
So you just turn off garbage collection and call it a day, basically, to be honest.
There's other things you can do if you want to be a little bit more sophisticated about it.
but you can also just manually have it all garbage collect on some interval like that's what we've done we just
have a garbage collection callback that just runs but i've seen the exact same thing yeah yeah exactly
so i thought that one was kind of funny and we did trace that one down and look and we did find the actual
call like again this goes to like having good tools so we had really good tools where we could look at
a bunch of like actual traces in sea and be like okay cool this is the thing that's taking a lot of time
or like you know this is the thing that doesn't quite line up here like oh i guess it's garbage collection
Okay, cool, interesting.
Yeah, let's just try taking it off.
Okay, great, that's what it was.
Now we can fix it.
So for each of them, like, basically bugs are not hard if you have good tools.
But if you don't have good tools, bugs can be very, very hard.
So similarly for like heat, another thing that we saw was like, oh, you know, the CPU is getting throttled.
Okay, well, it's easy to see if you're monitoring the CPU throttling or monitoring the heat.
If you're not monitoring that, it's really hard to know why it's just suddenly one of them is going slower.
I noticed also in the piece that you mentioned FSDP with 03.
Actually, we met, I went to Iclear and Guanhua from the DSP team was there presenting 0++.
I was wondering if you want to make any callouts to, you know, particular open source or open library or open whatever implementation teams that were super helpful in your process.
I think we ended up actually pulling from a whole bunch of different ones to pull things in into our own particular pipeline.
So we use things from invidias, you know, Megatron stuff.
We use stuff from probably deep speed.
I think we pulled in a bunch of different pieces from a bunch of different pieces from a bunch of different.
from places. So it was really nice to see all these working open source, like, examples. I think I really
appreciate all the effort that has gone into actually tuning these things because you can tune them,
but it's a lot of work to, to like tune this stuff and do all the stuff from scratch. It's really
nice to have like a working example. I think those are probably the two biggest ones as deep speed
and Megatron alone, but there are probably other ones as well. Is there, is there a particular
thing in the ecosystem where you would call out as like, you know, there should be something here
that is open source, but like it's not really, it's like everyone kind of builds it on their own.
I want to say something with the file system
because everyone talks about the file system eventually.
The file system actually was, I mean,
we did something kind of dumb there.
Like we have our own sort of local mirror
so that we can, you know,
like a crappy version of S3 that's local.
But it's just a pretty simple script, right?
Like I think we run like a little web server
that just like serves files and then, you know,
can upload them and download them.
Okay, great.
And part of the reason we did that is that our internet connection
in the beginning was not the like full speed one
that we eventually have.
and so we are a little bit more kind of bottlenecked in terms of internet bandwidth.
And so we had this.
I think we looked at a bunch of services out there, like Minio and some other ones.
But a lot of these come with a lot of extra overhead and maintenance.
And since we already have so much infrastructure to deal with,
we kind of didn't want to, you know, bring in a whole other like cloud provider,
virtualized something, something.
We just wanted something simple.
So we went with that, which has been quite helpful.
Like our tools are usually quite simple.
It's like Bash and Python and S's S&H and Docker.
Like we'd like to keep things simple so that's easier to debug.
like less layers of abstraction
make it a lot easier to work with.
We don't use Kubernetes, for example.
I would just directly launch these things
and it's just been much easier to debug this way.
One tool actually that does come into mind
that I will call out is Craken
from Uber. That was great. We love that tool.
We were a little bit skeptical.
What is it? I'm sorry. Yeah, so Craken is this.
Yeah, it's a distributed
Docker registry, basically, that uses BitTorrent
to like transfer things between
the machines and a sort of nice optimal way.
In the very beginning, the naive way is like you have this one Docker registry, which was outside of the cluster.
So every time we change an image, you know, there's many gigabytes that each of the 500 machines needs to download.
So that just takes a really long time.
So what this thing does is like just one of them downloads it.
And then like they all sort of broadcast all the pieces to each other.
And it was just like a really nice, fast way of getting these images down.
And it was very robust.
Like there's a lot going on under the hood.
But I think it's a pretty cool tool that we haven't really had any bugs with it at all.
Amazing.
Yeah.
I mean, that's all my questions.
I guess for the info piece.
I don't know if John, you had something that you were sort of burning to ask of or.
I know, all I can say is just same in a lot of things.
Like, you know, they're done that scene as plus one.
I think the one big difference, you know, perhaps in philosophies is we've tried to basically
standardize on as much commodity stuff as possible just because, you know, I think
the reason I asked about trying to do this on multiple different piece of infrastructure
is like, I think we're running on like six or seven different clouds right now.
And everybody has done something slightly different.
And my gosh, the little differences add up as, you know, you've seen.
And so, you know, our philosophy has been like, okay, whatever the hell we can standardize,
please let's standardize it, like vanilla off the shelf FSTP.
And like, you know, we wrote our own data loader, but we've tried to make that as much
of a standard as we can across our infrastructure and in data bricks.
Because things just start getting really complicated.
Or like, we use Kubernetes extensively because it at least gives us a uniform set of APIs.
Like, that's our hardware abstraction layer to a certain extent for everything else.
So it's just, you know, a difference in philosophy there, but otherwise, like,
Yeah, this stuff is really, really hard.
And I feel like we take for granted how much of this, you know, is done for us when you go and you just query chat GPT, for example.
Like, oh my God, everything going on underneath that.
You know, it's kind of a miracle that the machines boot off, let alone that you can, like, query a giant language model that's probably doing inference across multiple machines and was trained across thousands of machines, like, you know, minor miracle.
Yeah, it is an awesome amount of power that we invoke with a single API call that we take for granted these days.
It's absurd.
Yeah, I mean, like Kubernetes, like at that point about Kubernetes, I will say as a former AWS employee, like, it seems like it would be ideal for InBew to at some point make it more abstracted or agnostic because you're going to want to replicate your setup.
We do have our own sort of replacement.
There's just a much simpler version of Kubernetes.
Kubernetes is really designed for running services, not for running experiments.
Like that's not, it's like main architecture.
And so for us, like we have everything that's like, cool, you're going to run a.
experiment. So you wanted it to run to completion, right? Okay, great. Like, the primitives are sort of built
around a slightly different style. And that makes it a lot easier, like, just a lot simpler to
fit. The nature of, like, these machines are going to disappear. They will need to be rebooted for
infrastructure upgrades. They will, like, something will happen to the GPUs. Failure is, like,
baked into this as, like, a core part of our infrastructure. So it's not that we don't have an
abstraction. It's that it's a sort of simpler, more tailored abstraction for the particular work that we're
doing. Yeah, I think it all depends on what your goals are. And, like, I think the challenge in
lot of the deep learning stuff right now is that people are trying to, like, people often build
things that are more complicated than necessary to get the job done. And complication is the enemy of
everything. You know, don't use a fancier parallelism strategy than you have to. Don't use a fancier
set of libraries than you have to. Don't do anything that you don't have to do because it's hard enough
as it is. Like, don't overcomplicate your own life. Don't try to bring in more tools or more fancy
architecture tweaks if you absolutely don't have to. Like getting to the minimum of necessary to get the job
done. And it's really tempting to want to try to use everything. So like I totally understand that
one. I think the last piece I'll maybe call out is that I'm just going to weave this in just because
I see the opportunity to do it. Are there any infrastructure shifts that need to be, that that
need to rise because of changing architecture? So I think, for example, in view, like you're announcing
a dense model, a 70B dense model. Whereas John just worked on DBRX and and the,
is an image to text,
the text image model,
which presumably has different bottlenecks.
That's correct for us.
You know,
we train both dense and mixture of expert models.
The one we happen to,
you know,
kind of get permission to open source
was a mixture of expert model.
And those models are very demanding
when it comes to network bandwidth,
at least if you're training them
in kind of FSTP 03 style,
where there's just a lot of parameters
getting shuffled back and forth.
And your ratio of kind of compute
to amount of data
that you have to shuffle back and forth
becomes a lot worse because you're now, you know,
you're only using a fraction of the parameters for every token instead of all the parameters.
And so we had to really push the envelope on getting all the stuff to the right places on time.
And so actually the networking part of DPRX was the single hardest thing,
I think of the entire process, just get MEO training, working at scale across a big cluster.
We still managed to, I think, do it all with commodity parts, which was very exciting.
You know, the, like, we were using FSDP and we eventually used HSTP so that we could
have HSDP is a version of FSTP where you have multiple smaller replicas, and you're doing
data parallel within those replicas, and that helped a lot with network latency issues that we
were running into, just because we were transmitting so much data, you know, for every single
part of the process. I think it actually, like, it was instructive for how Google designs their
hardware and software together personally. Their training, as far as I understand, using kind of a
zero-three style of training and happened for a while. They also train mixture of expert models.
TPUs have a very different network bandwidth to compute ratio.
They have a lot more bandwidth, just objectively.
And TPUs per chip tend to be a little bit less compute intensive
and have a little bit less memory.
You know, it's just a different design choice.
So the ratio of flops to bandwidth is very different.
And that means that it's much easier for Google to be able to pull off some of this stuff.
They also have interesting, you know, Taurus-style network architecture,
or Tore-style, like, literal network architecture is not like the model, but the network that...
Is this the sort of block attention?
I forget what you call it.
So this is just more, or the, yeah, this is just more,
Or the, yeah, this is more, not the ring attention, but these are the ring all reduces.
Like, you have three different dimensions of rings because they kind of put them these three-dimensional
toruses from what I understand. And so, like, you know, Google's infrastructure in some sense is kind of,
I wouldn't say built for this, but maybe the way that Google Trans models is built for a slightly
different bit of infrastructure they have. And it's kind of neat to think about that.
You know, as one thing that I think Nvidia announced for, you know, for both the GH-200 and the GP-200,
is this hybrid networking where you'll have blocks of NVLink networked chips.
I think for the GB 200, I think it's like groups of 72 GPUs will all have NVLink to each other,
so higher bandwidth.
Then you'll have normal networking of some kind, infiniband, or rocky, or what have you,
between these blocks.
And that's kind of a, you know, it's a change due to the fact that, you know,
it's hard to build really high bandwidth networks over very large groups,
but it is now a blocked networking.
And you have to think about how you architect your model and your parallelism
differently. You also have to think about fault tolerance differently because it now matters where
you lose a GPU, whereas it didn't before. So, you know, it's, it's just all really interesting
and really fun speaking personally. But it's going to be new nightmares when we all move to that
generation and have to think about, you know, new versions of these problems. As you go up to
larger scales, it gets quite different. Like, right now, you know, if you're experiencing,
let's say, for example, you experience a GPU failure every day. That's fine. Just restart. If you
make your thing 24 times as big, now it's once an hour. Now it stops being quite as easy to just
restart, right? So now you have to just kind of break, like, bake in this sort of redundancy that you
didn't have before. So I think as you go up and scale, you end up running into like a lot of
really interesting problems that also inform the actual like design. Yeah, as an orchestration
guy, this is why I always emphasize like very cheap storage or very fast storage so you can
checkpoint more. But I don't think that's probably not the best solution to, to, for fast, you
I was training. Which works fine when you're doing language and then you move to visioner video
and then you have multi-petabyte datasets and getting cheap, fast multi-petabyte storage starts
to bite. I've certainly encountered issues where the literal data center where my GPUs were
did not have enough object store to fit the datasets that people wanted to bring into that data
center from whichever users were trying to bring them in. And then you get to a whole different
world of hurt where you have to keep your data in a different region because
the region is just out of storage. So things get fun really fast. Speaking of vision, Josh,
actually, you know, imbue is an agent's company, but you're only, you're announcing a text-only
model. Where does, where does the vision side come in? I think we've actually done a lot of work in the
past. And people can see kind of our blog posts about sort of self-supervised learning and some other
kind of vision-related stuff in the past as well. So we're very familiar with, with that stuff.
But I think our main focus right now is on kind of, as we say, coding and reasoning. And there,
there's certainly a visual component to some problems,
but it's not necessarily required for all problems.
And actually we found that for most of the kind of code writing
and reasoning problems that we care about,
the visual part isn't really a huge important part of it.
Sometimes if you really need to,
you can maybe describe the thing.
There are other like, you know, multimodal models
that you can use off the shelf to sort of plug in
for those particular pieces that you need, right?
Like if something is driving a browser or whatever,
like you can sometimes get away with not having to have that baked into the
original model. So our folks, we're, you know, in a sense, we kind of do a lot across the stack.
We're working on our own infrastructure and pre-training and RL and fine tuning and products and
everything. But in another sense, we're very narrowly focused on the application side.
So all of that stuff across the stack is kind of going toward like a very particular purpose.
And so that particular purpose right now doesn't really need vision. So we think that people are
going to make all sorts of really cool image models like Jonathan, right? And all sorts of interesting
multimodal models into the future. We'll let them go do that. That's great. We'll take advantage of
that partner with those people in the future. And right now we're really focused in kind of the
core reasoning and coding capabilities and aspects of the model. I wanted to go into carbs,
since that's like kind of the next layer of the stack. We talked about carbs in the first episode
with Kanjin, because you've actually had a blog post about it like a couple years ago. Maybe let's
introduce it. Have it been a couple of years now? No, it must have been at least one year.
Hopefully it was not multiple years. Sorry. I'm counting AI time. Yeah, yeah.
Yeah, I was going to say, you're making me feel really old right now.
I count everything before the generally intelligent rename as like, you know, prehistory.
And now sort of modernity, right?
So I actually thought carbs was more about hyperparameter optimization in a sense of like sort of parameter
hyperparam search.
Whereas, you know, when you introduced it, especially in this blog post, it's more about
scaling laws and predictability of like, are we sort of in the right fall part before we
scale things up.
Just maybe sort of recount the history of carbs.
Yeah, so it really is a little bit of both. So Carbs is, it's maybe a backonym, but it's for cost-aware,
Pareto region, Bayesian search. So this is about technically how it works. But carbs is like, you know,
we like pastries and stuff. So great, why not? But the point is that it's a cost-aware hyperparameter
tuners. So most hyperparameter tuners, you kind of say, okay, here's this objective function.
I want you to make this number as big as possible or as small as possible, whichever direction you want to go.
So yeah, just go make this number, you know, as small as possible.
Okay, so it'll try a bunch of different hyperparameters, a bunch of different configurations to figure out, like, how do I tweak your network and architecture, et cetera, to get the kind of best performance I possibly can.
That's usually saying, like, you know, almost all of these hyperparameter configurations are, let's say they're all going to use the same number of GPUs or the same number of nodes, or they're going to run for the same amount of time.
So you can do that, you can get a number out, and that's great.
But what Karp says is it says, okay, actually, what if we relax that constraint?
What if we say each of these different points,
we're going to model how expensive it will be to sample this configuration.
So what if we train with just one one hundredth of the data?
Like how well can we do?
What if we train with one-tenth of the data?
What if we train with all the data?
That way, you can understand, like, as we get more and more data,
as we spend more and more compute,
as we make a bigger and bigger network,
how does performance change?
But these things have changed, like,
how expensive it is to even explore this data point.
So by doing that, we can see the scaling laws
or not just, you know, the scaling laws from like the, you know,
Chantilla paper, the scaling laws for all parameters. We can see how does the number of layers
change with this? How does the learning rate change? How do the like, you know, various types of
regularization change? So you can see these nice scaling laws and as you're going across cost,
like how should this be changing as you're scaling up your model. So that coupled with the kind
of metric that we chose, which is a very precise way of measuring performance, allowed us to
really like hone in on parameters that worked really well and understand like how do we want to
scale those up, especially as we're changing things about the network. Like, one of the things that we did
is we used a custom tokenizer. As we change this tokenizer, changes a bunch of other things about the model.
So how should we scale up this entirely new tokenizer? Like, no one has ever made a model this large
with this tokenizer before. And so how do we want to change all these things? Harbs kind of shows
you like, look, as you change these parameters, like these other ones are kind of dependent
on this. Like, this is the, these are the relationships between them. So you can better understand,
like, okay, if I'm going to scale this up 10x or 100x, like, where do I want to be? Now, you can only go
so far. And so, you know, we did run like a, I think maybe it was like a 14B one or something
like that to check. But, and so we had a bunch of like 1B or 14B and then at 70B. I don't think
we had a, I think we just did like one at 14B. So you can, we get to check that like, oh, is this
on the curve? Like, is this where we expected? It was like right there. So then great,
go on to the next one. Yeah, I mean, that makes a lot of sense. I wonder if, so one of the key
questions and correct me if I'm wrong, but like usually people do search or do their e-vals just
based on loss, but you actually
evaluate based on the
end state evals that people
might expect, like HeloSwag
and Lombada, whatever.
What is the norm here?
Is there a norm?
Yeah, I don't know if there's 100%.
I don't know.
I only see loss on most people's reports.
I think it's easy to
like loss is very nice because it's
very precise. It will tell you
like very fine-grained differences
between like really small changes in your hyperparameters
of network architecture. Whereas,
especially at the smaller scales, if you're looking at accuracy, it's very noisy.
Like it might be zero or 100 or like, you know, fluctuating by like 10 or 20 percent of points,
which makes it really hard to tell.
Like, did that change actually mean anything?
So our loss is sort of a combination of these two.
Instead of saying, like, let's just look at perplexity, we say, let's look at perplexity
on the tasks that we care about for multiple choice questions effectively.
So we're saying like, yes, this is formulated as a multiple choice question.
And we're going to look at the like, you know, the loss of perplexity for this particular
answer token. And that ends up being something that's like both targeted to what you actually
care about and also very precise. The nice thing about this, though, is that it's independent of
the data that you train on. One thing that's annoying about perplexity or about loss is that as you
change your data set, this is really obnoxious because now it fundamentally changes your
loss, right? So you can't tell like, how do I tweak my data set? But because we have this held out
evaluation data set where we're looking at perplexity, we can actually change the data mix.
And so Carbs actually controlled what is the mix of data that we want to see, like how much
code, you know, how much internet text, et cetera, in order to figure out what is the best
optimal mix of data. And we could do that because we have this other metric. So that was one of
the things that was really, really helpful. I think there is a trend overall about changing
data mix as training goes on. I don't know how, you know, we're deciding not to talk about
datasets in this podcast. But what have you observed about the changing data mix question? We did. We
did some experiments, and we've actually talked to a bunch of researchers who were doing work here as well
and looking at kind of their experiments on this. And we were originally pretty hopeful because it
sounds like something that should work and make sense, right? Like, oh, cool, like maybe you would
have your model, like, learn the basic features. And then over time, it could get really good
at these complicated math problems or coding or something, right? But it just turns out that, like,
it's just not the way it works. Like, oh, we've done so many experiments. And you can get like a
tiny, tiny little boost from this. But it just is not, like, it's just not the important thing,
at least in the experiments that we've seen.
So, yeah, we've kind of, we're letting other people explore that more if they want,
but that just doesn't seem like the most promising direction for us.
We've had some surprisingly good luck with this.
We just released a paper on it.
The details matter a lot, and it really matters what you're trying to do with the model.
Yeah.
But it's been quite effective for us depending on the setting.
And certainly when we're thinking about domain specific models, this helps a ton.
You know, to a certain extent, you can always think of this as like early fine tuning.
But yeah, there have been little glimmers of this in the literature for years, like, especially, I think the Gemini 1.5 paper mentions this. And I don't remember what the Lama 3 paper mentions this, but it's kind of, it's one of those, like, people have different ways to get to these endpoints. I think, you know, there are the architectural tricks that each lab has to mitigate lost bikes or what have you. And everybody's got, you know, their own bag of tricks. And it leads to kind of sometimes this contradictory information. It's not contradictory people who are just kind of exploring different parts of the space in some sense. And,
There are lots of ways to get a great model.
But certainly for us within our config,
and it seems like I guess for the folks at Google
within the part of the world they live in,
changing the data set has helped,
but the details matter a lot.
And it's really hard to get those details right
for the reasons Josh just mentioned.
Like, there's a lot of search involved.
And you essentially have to make hard choices about
what parts of the space you're going to search
and which ones you're going to leave be.
And so, you know, some people have done an amazing job.
Like, I think the, who is it, the deep seek folks,
have done an awesome job looking at like patch size warmer.
And that's been really, really fruitful for them.
Other people are looking really hard at things like data mix,
but it just gets tricky to look at everything.
Yeah, I think we found that we could get some things that looked like gains from
datasets, but one of the things that I like about carbs is that when we applied carbs to
properly tune things, then a lot of those kind of evaporated.
Whereas like, if we just tune these other parameters, actually we can get almost the
same gains without having to do this more complicated thing.
So at least in the experiment, in the settings that we've, like in the particular metrics that we
care about. We haven't seen these kind of like pan out or scale up in quite the same way, but not to
rule it out. And I think you're right, Jonathan, there probably are a lot of like details that go into
exactly what is the metric, exactly what is the data set, exactly which, like what schedule are
using for this. And I certainly wouldn't rule it out working. Quick question about emergence.
Doesn't emergence throw a spanner in the theory of carbs?
So there is a paper of which I really liked and I think informed a little bit of how we thought
about this, which is our emergent properties of language models and mirage. And I think if you look at
that paper, it actually makes a relatively compelling case that in fact, you know, this emergent
behavior that you're seeing is not really emergent behavior, but is really a function of the
evaluation metrics that we're using. So if you look at accuracy as a metric, what's happening is that
accuracy is actually going up continually over training, but it's in log scale. So it starts out at 0.001%.
0.1, 10. Only when you're going between 10 and 90 do you see this happen, right? When you go from
1 in 1,000 getting right to 1 in 1,000 getting wrong, like there's many orders of magnitude happening here.
So when you're looking at this in perplexity, then you just see this nice straight line.
And so that's actually what carbs is exploiting. Like since our metric is in this kind of like perplexity, log space, like you can see like, oh, it's just like getting better as you make it bigger in this nice, very predictable way.
So that, and that is exactly what we saw. Like these things were really, really bad at,
you know, predicting the multiple choice answer
just always gets A. Okay, it's so terrible at it.
But it was like learning to be less confident about that.
Yeah. One trick I saw from one of the papers recently
was just like just randomize the order of the multiple choice questions.
And if you, if they over, if that hits the performance a lot,
then they're just basically memorizing the test set, which makes a lot of sense.
Yeah, this is, I mean, you know, I completely agree with what Josh said.
I think my bigger lesson is that anything can look however you want it to look if you put it in a log scale to a certain extent.
And we love our log scales and deep learning for various reasons.
Everything looks very clean on a log scale until everything looks very flat on a log scale.
I don't know.
Log scales always mix me up.
That's all I can say.
Great.
I think the last thing I was going to mention on carbs.
Oh, well, I mean, let's just kind of go right into e-vals because I think that's going to be the sort of crowd favorite.
So carbs, we already mentioned, you know, leans heavily.
on the sort of end evals that we would typically evalms on,
except that you had to make your own.
There are a lot of documented problems with many of the common evals out there,
and you fixed all of them, it sounds like.
I don't know about fixed all of them,
but I think in the same way that we like to dig into the infrastructure and hardware
and understand what actually is going wrong,
like what is the actual error on this machine with this GPU and why did that happen
and how do we fix it?
We take the same approach to the evaluations.
So when we looked at the evaluations and actually looked at the data sets,
what we did is like, okay, if we're going to be evaluating natural language understanding
and reasoning, like, let's look at all the data sets that are out there.
Let's actually look at a bunch of the examples and say, like, is this a good data set that
we should use for evaluation?
That's kind of how we selected the evaluation data set that we had.
And then when we looked at the actual examples in there, we noticed, like, a lot of these
very messy, like, some of them messy to the point of like incoherence and some of the
ones that we didn't choose.
but even the ones that we chose, like, people tried pretty hard on these data sets.
They did try and clean them, but there's just a lot of data points in there,
and it's just easy to make mistakes, right?
And so, you know, it's not that they have 100 people looking at every question,
like, that's just way too expensive.
So you end up with questions that just don't make sense.
Somebody didn't really see this.
Somebody just clicked the wrong box for the answer.
Or the question makes sense in your head when you write it.
We've often seen this.
It's not even like malice or incompetence.
It's really just like, you know, you write this.
You're like, this makes sense to me.
You show it to another person or like, that makes sense.
You show it to a third person.
And they're like, this makes no sense at all.
It's because you're kind of using a different meaning of the word.
And then when they say that, you're like, oh, wow, you're right.
That is actually really confusing.
It's easy for things to kind of make sense in our own head.
So what we did for the evaluations is really dug into the details of each of these
data sets and tried to ask like, what makes a good question?
What makes a good answer?
Like, what does it mean for it to be ambiguous?
We had a whole, like, we looked at lots of data, broke this down, asked lots of people
about all these different questions to build a model of this and help us kind of clean
these datasets.
That was sort of one big piece of it.
A second big piece was making sure that our data that we're training on is not data that we're testing on.
So there we kind of took a step back and said like, okay, let's just reproduce, you know, 500 to 1,000 examples for every single one of these data sets ourselves.
And just make sure that this data is definitely not in the, you know, the training set.
So we did that and then we're able to like now be confident about like our performance of our model and also performance of other open source and other closed source models.
Yeah, there's a lot there. You had 11. I don't know how many data sets. I think so.
One, two. Yeah. Anyone you want to call out in particular to dive deeper on. Some of these I have, are very famous like Hello Swag, Miro Grand. Some are less famous like race. I don't know if. Race is a great data set one. Yeah. Yeah. Just, you know, anything that's interesting you want to call out on specific datasets. I think there are a few asterisks in there. You know, definitely read the whole paper as you're looking at some of these. Like the GSM8K one is a little bit weird. I think one that one that one that one,
was kind of funny. It was like low performance on ethics from some of the more recent models. I think
that was a little bit funny because the models, you know, I think there was a reaction to like,
oh no, like, you know, the models are saying bad things. So they went way, way in the other direction.
And now like on the ethics data set, it's always like, this is totally unethical, even though it's
really fine. So they just been tuned to, you know, make sure they don't make any of PR disasters.
So I thought that was a little bit funny. Not to say that it's necessarily like a flaw of the model,
but just a kind of like, you know, political or tuning opinion. I think the
I was just going to say the main takeaway
from any of the actual performance is like
once you fix these ambiguous examples,
a lot of these benchmarks are really saturated.
I think it's important to look at like,
you know,
like when you're talking about performance on ANLI or race or pool queue or something,
what you're really talking about is like performance on questions that make no sense.
Like it's just like did it guess the answer in this like really weird scenario?
Like those are the ones that are left.
Like when you look at the performance on the ones that actually make sense to everyone,
all the models agree, we agree, like everyone's on the same page,
which I think is kind of a really interesting result.
The question then becomes, you know,
what are the new like set of evals that would be like the next frontier
that often embats with it your idea of what reasoning is?
Because it's obviously you're super interested in reasoning.
And yeah, I mean, like, where does this, where does the state of evals go from here?
This work and this blog post is talking mostly about the public evaluations
and the things that we can release.
We do have our own internal evaluations.
For example, one of them that we are releasing is the code understanding evaluation,
which is about predicting, you know, what will this variable be or asking questions about code,
et cetera.
And that is one of the early benchmarks that we made that we can release.
We can partly release it because we can generate an almost infinite amount of this data because
these are programmatically generated.
And so, you know, we're not really worried about there being like corruption in the kind
of the training or test sets.
So that makes it a little bit easier for us.
I think it's, you know, we have built other data sets as well that we can't release.
Some of them, you know, for example, because they maybe use other open source code and so we can't redistribute it necessarily.
Other ones because, you know, that's, I think evaluations and data are like a core important part of, you know, the business.
And I think we take evaluations very seriously and are spending a lot of effort in terms of like what exactly do we make as part of the evaluation set.
How do you evaluate these things?
We've done a lot of other stuff, you know, since these evaluations.
But I think a lot around like code understanding for us, since that's our main focus.
and it's a nice place to explore reasoning as well.
It sounds like you talk a little bit about like code understanding as like sort of variable level,
like sort of very micro context.
Is there a sense of like larger code context as well?
I don't know what I mean by that, by the way.
It's mostly just like if I told the senior engineer to go look at a code base,
they would understand at a broad level, the architecture,
but also the design decisions and be able to tell me that.
I don't know if that's useful or not, but I mean, that's useful to me as someone who might be working with them.
Yeah.
this particular data set is like the more low level code understanding like just literally what happens in this code and this is mostly because you know this is part of the carbs tuning metric etc like we care about the low scale version of this as well we want smaller scale models to be able to do something on this and so that's kind of the focus for this and hopefully this is more useful for other people but yes those other questions are also quite interesting they get a lot harder to evaluate like is this a good architecture or not like you and i could probably debate for a while on you know different architectures and so it becomes a lot trickier to do these evaluations
as they become more realistic.
So I think that's one of the things
that we've been playing around with a lot,
especially around like code generation.
So if you're saying, you know, implement this function,
okay, it can be kind of objective,
but even MVP,
we've made our own internal version of this data set,
where we've taken like every single example
and looked at it and then like,
does this actually make sense?
Like what is the type signature?
Like, can remove all ambiguity, etc.
So you basically like reviewed every single question on,
I mean, that's impossible for like hello swag, right?
Yeah, yeah, we didn't do that for hello swag,
but this is for MBPP.
which is only like a few hundred, so we just sat down and did it.
Yeah.
I'm so excited to get to look at this data set.
Like, this is such a resource for the community.
I absolutely can't wait.
We should probably do the, I don't know.
I don't know if we were planning on doing the healed MPP one,
but hopefully we can do that one in the future.
Did you look at SweetBench?
This is the sort of hot new dataset of this summer.
Yeah, I've taken a quick look at SweetBench.
It's really interesting.
I like that it's a much more difficult kind of coding,
code-related task for bug fixing.
I think it gets into some of these problems.
where it is a lot harder to evaluate these things once they get more realistic.
Like we were looking at the agent bench paper, I think, just last week for our paper club.
And one of the things that we noticed is that actually like both of the examples in the appendix
that are given as like traces where it got it right, this is actually not the right solution.
And it's okay.
You know, it's fine.
Like it did make it pass the test.
That's what the metric is.
That's what the benchmark is about, right?
But like it just said, you know, like, you know, dot encode askey.
Like, well, that's not the right way to do this.
It just dropped all the other edge cases that you actually would have cared about in production for this thing.
And there is like a better way of doing it.
And that's what the real golden patch was.
But, you know, that's okay.
But then how do you test all of that?
Like, as you start to do more realistic things, the test coverage, like getting test coverage over all possible ways of solving these bugs is really hard.
Evaluation is the single hardest part of the whole thing.
Like, I spend a shocking amount of time just telling our customers, we need to find a way to measure what you actually want out of the model before you should ever touch a GPU.
you and trying to convince my team and me to follow our own advice a lot of the time on that.
And I think everybody, like, on the one day it's easy to laugh at the state of the evaluations that we have.
None of them are good.
Like, if you go read these evel benchmarks, you'll always come away disappointed.
And yet they've given us useful hills to climb.
And we do seem to be making progress and measuring progress in the field.
And I think anecdotally, models are getting better year to year.
So I feel like people tend to go and get into one situation or the other.
Like, evils don't matter.
I'm just going to look at loss.
or like, you know, the evals matter a lot and they're all broken, so what do I do?
And I think, like a lot of things in deep learning, we have to make peace with just complete
imperfection.
Like the most successful scientists I see are the ones who are okay operating in a world where
everything is going to be broken.
And yet we can still cobble things together and make something interesting happen.
I mean, we were just discussing that with literal infrastructure.
And now we're all the way up to like, how do we measure whether a model performed a complex
coding task correctly?
And everything is broken.
And yet we're still able to make huge amounts of forward progress.
I think that's right. Jonathan, and that the challenge isn't necessarily making perfect evaluations.
I think our blog post here is about going really into the weeds on these to figure out like, what does that look like?
And I think one thing is like, you know, as you said, we have been able to make a lot of progress without making these perfect.
That's great. You don't have to have perfect evaluations. And, you know, the more interesting work is the stuff that we can't necessarily publish about, which is the imperfect evaluations that we have for, you know, actual coding tasks, for example.
like what does this really mean as a person? And there, as you said, it's much messier. So it's a lot
harder to put it out and say like, hey, everybody use this because there's so many rough edges.
It's so hard to like even say, oh, is this even the right task? Is this even the right way to do it?
And there's a lot of judgment. There's a lot of intuition that it comes down to.
But yeah, I think that's where it's critical to do if you actually want to make these systems work.
Yeah, you have to make peace with living and that in between. Yeah. And I think that in some sense
when I hire researchers, that's the number one quality I look for. Like can they be at peace living in a
house that is neither clean nor messy, but it's just kind of somewhere in between. And are they okay with
that? Are they okay with a few dishes being out on the table and a few clothes being on the floor?
Or will that drive them insane? Or will they just end up with all the clothes on the floor and
like all the dishes out all the time? Like it's kind of, I'm looking for that perfect balance because
we have to operate in this imperfect world. Like, yeah, go ahead and give me the perfect
evaluation for programmers or for an LLM that is a program assistant tool. Like there is no
perfect evaluation. Yeah. But clearly we've made progress. And so the most
important part is just are we climbing the right hills? And so this is why I'm so excited to see the
ambiguity aspect of this. We often think we have more room to climb on these benchmarks.
It turns out we don't. Or it turns out that actually we're climbing, getting good at the benchmark
and not actually getting good at the task we care about underlying the benchmark anymore.
Maybe the model, like this is the famous example where if you get 100% an MNIST,
your model must be broken in some way because there are four examples mislabeled.
You know, it's that all over again. Welcome to the same goal. Yeah. It's the accidental
canary canary in this.
I think one thing that's actually really interesting
about this also is that, yes,
the ambiguous examples are sort of
not that great from the perspective
of these particular tasks that we're evaluating,
but actually one thing that we're very interested in is
ambiguity itself. Can we detect
whether a task from a user is ambiguous
or whether you've completed a task
successfully? These are actually hard,
messy problems, but are really important
from the user experience of using these models.
I would much rather have a coding agent
that will give me back a thing
and, you know, it's actually the code doesn't work like 10% less of the time in some other model,
but it will tell me 100% of the time when it's not sure.
Like that's so much more useful if it can communicate.
Like, I'm not really sure about this.
Or maybe there's some errors here than just like, here's some code.
I have no idea if it works.
And so these kind of like, you know, detecting ambiguity and detecting correctness or uncertainty,
I think are really interesting problems that we're really like digging into quite deeply.
I'm going to touch on maybe a couple of hot topics and e-vowels,
maybe tangentially related, but we're on the Evels train right now.
So I'm just going to get on that.
So ARC AGI, Francois Cholet's hot new thing,
it's sort of my take on it is basically it's trying to measure reasoning
through an abstract IQ test effectively.
I notice that you don't use it.
There's a lot of the community debate pro and con about it.
What are your thoughts on just more abstract reasoning and maybe RKGI specifically?
I think we purposely stayed away from the very abs.
Like there's Big Bench, for example,
that has a lot of, I think, kind of, to me, feels sort of similar types of tasks that are like very
unrealistic. Like, oh, you know, we have books of different colors and then you're going to shuffle
them and like, which book is furthest to the left or something? Like, okay, cool. I guess it's neat.
It's neat, I think, for us to explore in terms of like an agent reasoning in a like larger loop.
And we do care about these types of evaluations there. The types of evaluations we're talking about
in a blog post here are for getting at like, does this model, like in a base model sense,
is this working at all? Like, there's no change.
of thought in these evaluations. These are just like, go straight to the answer. Does this make sense?
Like, is this a thing that you can answer very quickly? That's what we are selecting for with these
evaluations. This is not to say that these are the only evaluations we have. I think the arc ones are like a
little bit too probably visual for us to really be able to integrate with. But I think some of the
big bench ones are. You can tokenize it. Yeah, but you know, I think it's not really, I think you can
spend a lot of time getting really good at these kinds of benchmarks without making like kind of more
general purpose progress. And so I think we're a little bit leery of going too far in that direction.
Similarly, like, coding competitions. Like, we do a lot of code generation, but we don't really do a lot
on like code competition problems for the very, very hard ones. So I think you can go very far down
that route. It makes something really good at those problems, but not actually that useful as like
a programmer day to day. Yeah. Take a different tactic, which is like at the end of the day at
Databricks, I have 12,000 customers or I think that's the latest number, all of whom are trying to do
something with, you know, LMs or AI or machine learning.
And those things don't look like these tasks.
I don't think I have a single customer that's asking to, you know, have AI solve abstract
reasoning problems.
Things are pretty, like, they can be ambiguous, they can be challenging, they can be really
interesting.
But none of them look quite like this.
And so, you know, I think to Josh's point, like, it's really about asking, why are we
doing this?
If, even if you're trying to build AI, and that's not personally my purpose.
And I, you know, Josh has much more interesting things.
say about that than I do. I don't even know if this is the kind of intelligence I would get excited
about or care about personally. Or if I would consider, you know, to Josh's point, this to be the
indicia of intelligence. It's neat. But, you know, for me, it's like more down-to-earth things like
having a model that can have a conversation with you about data that on the back end is running
SQL queries on your literal data. That's a much more interesting task to me. That's something that
really matters day to day for my customers and, you know, different perspectives. But, you know,
I think Josh and I would probably say the same thing, even though I would, I'm guessing I don't want to put words
in your mouth, you would say that you're pursuing more general intelligence in your own way.
And I would say that I'm very happy with narrow intelligence.
Like, I'm very happy with my little SQL bot and building 12,000 of those because they're,
they're moving the needle for a lot of folks every day.
Yeah, I think we're, you know, we're not as far away in our position as it might seem.
I think we're also excited about, like, how do you actually make these things useful?
And that does end up being pretty narrow.
I think these other tasks can be interesting as like ways to explore these more abstract
reasoning questions or like, okay, how could an agent actually work through this? But it's
important to keep in mind that it's like a toy, not a real problem. It's like, it's a scientific
tool to tell us something about the models. It's not something we should be optimizing for necessarily.
The one thing I'll point out is, you know, as a kid, I was graded into a gifted program based
on my ability to solve these exact type of problems. And then I entered college based on my
ability to solve SATs, which again, I have nothing to do with my college experience, but whatever.
So we have a history
And in a humanity of doing correlated IQ tests
To general capability
Okay, so there are two more viral evils
And then I just want to be mindful of your time
Needle and a Haystack
Long context utilization
Oh for the love of God
Something well, okay
Let's just assume that
On our podcast we've discussed
The baseline problems with needle and haystack
But just generally long context, right?
It's a useful thing for agents
I assume
And it's something that
You know
It's out there like we
don't know, don't really know what the best way to utilize memory is, but like, I assume it's
important, right?
What I'll say is like, you know, I spend a lot of time thinking about RAG these days.
And RAG, you know, in one sense, you know, the way that I think about RAG is it's the
world's simplest agent.
It is an agent that basically, you know, there's at least more than one thing happening in the
process of building model.
It's at least a system.
If you give the model the ability to decide when it wants to retrieve data from a context or
retrieve data from a database, then we're talking about an agent.
So RAG kind of, I think, like, toes that boundary really nicely.
there are a lot of reasons why you do genuinely need a long context. Like, I don't think long
contexts are problematic in and of themselves. I know there's some controversy even about that.
I love the idea of doing like a thousand shot tasks as an alternative to fine tuning. I love the
idea of pulling in lots of data into the context. I love the idea of once you get into
multimodal land, you're just going to end up with giant context. It's kind of unavoidable.
The flip side is, I don't know of anyone who like is hiding a secret passphrase in a book and
needs the model to find it. Needle and a haystack is, it's interesting. The challenge with long
context to my mind and Josh, I'm curious what you think, is simply that annotating long context
evils is really hard and really expensive, you know, intrinsically because you need someone to read
10,000 tokens or 100,000 tokens or like, you need someone to read a 1,000 page book or the equivalent
thereof in order to measure these long context benchmarks. I don't know if a human could solve
these tasks, let alone that a human could do this in any amount of time where you're willing to pay the
money to get the data annotated. And so any long context eval has to in some sense be correct by
construction. And you have to, you know, the, you have to know the, you have to know the answer
before you've created the example. A needle and haystack is kind of the simplest way of doing that.
I think the problems with needle and haystack are well known, you know, it doesn't measure anything
real. You're not even testing the model's ability to holistically use the context just to identify
one part of the context. You can do some wacky things to your model like quantize the hell out of
the KV cache and still get needle and haystack to work quite all because it's not trying to
holistically take advantage of things. You know, I have some thoughts on things that I like more that
are also still correct by construction. I really like the idea of doing thousand shot tasks where you can
look at the scaling as you go from 10 shot to 100 shot to thousand shot to fine tuning on that data instead.
And I like that as a way to have something that's correct by construction, at least where you have a
nice baseline that you can compare to automatically. So I'm typically looking for like contexts that
or situations where long context is one way to solve the task, but not the only way to solve the task.
And we have some other strong baseline floating around personally. But yeah, needle and a haystack,
not my favorite thing in the world, to say the least.
Yeah, I mean, I agree with most of what Jonathan said, I think.
I think one other thing that I will call it is that, you know, from like a coding application
perspective, it's useful to have long context because the lazy thing of just like,
throw the whole repo in the context is like, oh, okay, cool.
Like, you know, you can just get started with that.
But then in, you know, in real scenarios, you don't necessarily want to put the whole
thing in there.
You can have code bases that are bigger.
You probably want to filter down to the stuff that's relevant anyway to not be
confusing.
Like, you probably, even if you did have a lot of context, you might want to sort it in
some way to say this is more important than this other stuff.
And you don't want to wait for, you don't want to use
wasting all this time and compute on inference and like doesn't really matter.
So yeah, I don't know that it's the most important thing.
I think people find creative use cases.
And like John said, I think the multimodality examples will naturally lend themselves
to long context.
Cool.
And then one last one on just general sort of agent related capabilities that we didn't
really talk about in the eval section is function calling and tool use.
There's a recent trend, I think basically.
led again by OpenEI on parallel function calling.
There's always, there's been a limit on how many tools you can call from four to now,
I think a hundred and twenty-eight.
And I think theoretically Claude and Gemini support a lot more.
So just generally, like, how do you think about e-valling tool use?
Is that super important for you guys?
We're thinking about it in a slightly different way, which is, yes, you can have this,
like, hard-coded list of tools.
But if only you could have, like, this really large open set of, like, tools.
Maybe they would be, like, functions that you could call.
if only there was like a language or like a programming thing, like being able to write code.
I think for us it's like, well, look, if we can write code, like now you have all these tools
accessible. At the end of the day, like function calling is just a function invocation like literally in code.
I think our approach to this is like instead of worrying about like weird hard coded agents using tools,
like let's just make them able to actually write code robustly and make that code work and be able to debug that
code, know if that code is safe to run, like get really good at the like code writing and execution part of things
because that will open up the action space
far more than, you know, 128 tools.
Like, just everything is at your fingertips,
especially, I think, over the next few years,
like we already have so many really good APIs.
As we get better and better at writing code,
we'll be able to make APIs to things that don't even have APIs today.
That's kind of how we think about it.
It's less as like a special purpose thing
and more as like this is one of the reasons to focus on code.
On my end, the way that I think about this is, you know,
I think a lot about how models interact with data.
And so for me, tool use is really a question of how do you take models
that are really built for unstructured data
and have them interact with structured data.
So, you know, and I get the question a lot for my customers,
like, what do I do with tabular data?
Or what do I do with like, you know, JSON or what do I do?
I mean, you name it.
Like even what do I do with a PDF?
Because PDF parsing is still an unsolved problem,
even in 2024.
And the answer or even just the basic question of like,
should I bother to structure my data anymore?
Shouldn't I just toss the table?
Shouldn't I flatten it and just throw it into the LM context
and like let the model figure it out?
The answer is no.
We've built all these.
fun APIs and fun languages and paradigms for dealing with structured data over the years,
just use them. Have your model use them. Train a model that can interact with these things in a
meaningful way. Like, text a SQL is still, or like having a model be able to make SQL calls in the
backend is actually like one of the single most useful things for my customers. It sounds really
boring. Models are really good at it and it moves the needle day to day. So tool use for me really
is that. Like how do you just interact with structured data sources and take advantage of the fact
that you have some prior knowledge about the structure of your data, that an LLM would completely
flatten away. In many ways, this is kind of one of the, one of my biggest frustrations with
the fact that LLMs work well with code. We have decades and decades and decades of understanding
about the structure and interpretation of programs. Like, I think that's literally the name of a book
on programming, if I remember right? And, you know, we have all this theory. We know everything there
is to know about programming languages if they're well-formed languages and have the right
properties. And yet when we have an LLM work with them, we literally just turn it into a token stream.
despite the fact that we know how to parse it,
we know how to do all sorts of, you know,
reference, you know, disambiguation and things like that,
we're still just flattening it into a model
and making the model relearn all of these things from scratch.
And it frustrates the hell out of me.
I don't have a better answer when it comes to code,
but I really appreciate that with a lot of data sources
that have structured to them,
tool uses and function calling are just,
in my mind, the right way to deal with this.
So I think basically what you're saying is like,
code is the God tool for Jonathan,
like, you know, SQL is so much the right abstraction
for accessing all this data.
One thing I do spend a lot of time thinking about
is for the stuff that doesn't fit in a SQL table,
is Knowledge Graphs the answer?
I think a lot of people are exploring that.
And I think every now and then people get about
of Knowledge Graph Religion and then it kind of doesn't work out.
So I wonder what the end state is.
Like, is this an idea where it's a mirage
or is this the idea where sometime is going to work?
It's about having the right tools for the problems, right?
Like, as Jonathan was saying,
SQL is sometimes definitely the right tool.
like you've got your, you know, order table or something and you want to know, you know,
number of sales last month.
Like, you should be using SQL, sum that column.
Okay, great, you're all set.
Knowledge graphs also, you know, are sometimes the right tool for a particular problem.
You have some, like, weird question about relationships between entities that are modeled
on some particular ontology that you actually understand and is like maps to the real world.
Great.
Use a knowledge base.
Like, use a knowledge graph.
This is fine.
But I think in the real world, it gets a lot messier than like knowledge graph style of things where
it's like, well, is there a relationship between these two nodes? Like, ah, I don't know.
Like, are these two separate nodes? Like, those kind of messy borders, I think, prevent it from
being a tool that can like solve everything forever. And so I think it'll always be good for
certain problems, just like SQL is good for certain problems. Like, different abstractions are good
for different problems. And yeah, I think this is why I'm excited about code. Like, code lets you
kind of pick the right, like, let's use this library for this problem. Let's use this library for this
library for this other problem. I think Josh said it and you said it well, like code is kind of
the God tool. It unlocks literally everything. The challenge for me is always,
like, you know, sometimes unlocking too much power can sometimes inconvenient things can happen.
And so it's all about balancing that. In some sense, language is the God tool. If only, you know,
we knew how to interpret it all the time. So code has the really nice property that at least you can
always execute it. And sometimes you just literally want your model to be able to do SQL calls and
nothing else. And setting those boundaries properly for the problem, I think is going to be,
I think at least a lot of my customers are going to be thinking very hard about that. Like,
should I give the model access to the web? Is that actually helpful for this problem? It sounds
great to just flip yes on all the tools.
Is it actually going to mean I'm going to get better solutions to my problem?
So I want to be mindful time.
I think that's basically our sort of recap of our discussion based on imbused releases today.
I want to leave some time for what's next for both of you guys.
Maybe Josh, as a guest of honor, you want to go first as to what happens next.
We have these releases.
We're happy to put these things out.
I think there's a lot of stuff that we haven't released.
This is not the only thing we've been working on.
Most of our actual focus has been on kind of coding and reasoning, in particular,
or like the things that we're excited about are can we make these things useful like
Jonathan is saying right like it's not about toy problems it's like can we use these today in our
day-to-day workflow and actually have them accelerate us and I think we have some kind of internal
product prototypes and things that we're excited about and so we're excited to share more about this
in the coming you know months to quarters as we get it to a place where like other people could
maybe get value out of this as well but that's kind of our real focus right now is like how do you take
these really cool capabilities that are out there that our models have etc and like make sure that
they're actually useful today for us, like when we're doing real work and then for other people as
well, in particular, focused on generating code, understanding code, testing code, verifying it,
like starting with the like robust creation of software.
Excellent.
Jonathan?
I never like to talk too much about the future because I don't know.
I think you've heard this for me before.
I like for us to speak through our work.
And so I don't, I don't like to tease too much.
Our mission is to Josh's point to make this stuff useful to 12,000 customers and not a lot
of that ends up making it into the public eye and not a lot of that ends up getting released open
source. So for this kind of forum where really, you know, where we're talking to the community,
I'm asking myself right now, like, you know, what exciting things are we going to have to offer
the community in the next little while? I think the most exciting part is just we're writing a lot
of blog posts right now. We're trying to share more and more of our science because I feel like we've
been doing these big pushes to create these really giant models. I think, Josh, I'm sure you had the
same experience. It's exhausting and all consuming. And you get to the end and you're like, oh, I have
all this stuff I want to talk about. Now I need to find the time to talk about it.
now that I've survived this huge push, and we're definitely in that mode right now.
So there's going to be a lot of that coming in in the next little while.
And, you know, we're always cooking up fun new models.
I think the real question is, you know, releasing models open source is not our day-to-day bread and butter.
It's kind of a fun reward that we get to do sometimes when we have something really cool to share
and a little bit of time and spare GPUs in our hands.
But for the most part, everything is going toward customers.
You know, I think the joke is Databricks has been 18 months away from IPO for five years.
So I guess Databricks is 18 months away from IPO still, but 18 months away from
from IPO means there's a lot of pressure to deliver for customers, and we're going to keep
working on that. But I think you'll see hopefully some cool interesting things get dropped over
the course of the summer and into the fall. We'll find out when we get there. I think that's
the right way to put it. I know we were talking earlier about kind of Aber-Kadabra and Al-Qazam,
and all I'll say is that, you know, the DVRX small model that we still haven't released yet
was called Abra. DBRX was called Cadabra. And there's a third Pokemon in that evolution.
And that's all I'll say for now. Cool stuff kind of popping up sometimes on chatbot arena.
And, you know, keep your eyes on.
Yeah, I'll leave the links and the hints to the show notes.
That was a very fun way to leave some breakruns for people to follow.
Cool. I'll leave everything to some calls to action.
We're going to be releasing this next week.
So I will be deep in my conference, the AI Engineer World's Fair.
So people can just go to AI.orgian and live stream it.
Do you guys have any other calls to action before you wrap?
The only one is, you know, we're definitely hiring.
So if you're interested in working on coding, reasoning,
interested in working on all this stuff, you know, from the ground up
and really deeply understanding not just how does the hardware work,
but how do the models work,
and also designing these systems to actually be useful for yourself day to day.
Come say hi.
The only thing I'll say is, you know,
and I like saying it these days,
it feels like the field is so crowded
and, you know, it requires so many resources to do impactful work.
And, you know, on some days it feels like everything's been done
or somebody else is doing everything before you can.
At least I remember that feeling every single day of my PhD
and even more so now.
But I hope, like, what you heard from Josh today tells you,
there's so much enormously impactful work to do in the field. If only you take a step back
and take a fresh look at some of these things and just talk about what you're doing,
there's a huge amount left to do here and a huge amount of exciting work happening every day.
And for those who are certainly feeling that exhaustion right now and I count myself
among those folks many days, it's refreshing to see these kinds of drops and see that there is
so much more, even in things that people feel like they understand how to set up a cluster,
my God, you know, even in these e-vails that we think we understand, there's still more to understand
and still more work to do. I hope everybody's keeping at it. All right. Keep on, keep on,
keep on. Well, thanks so much for your time, guys. That was great discussion. And we'll put the
links in the show notes for people to read more. Thanks. Thanks much. Thanks so great.
Thank you so much.
