Latent Space: The AI Engineer Podcast - The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka

Episode Date: July 5, 2024

Livestreams for the AI Engineer World’s Fair (Multimodality ft. the new GPT-4o demo, GPUs and Inference (ft. Cognition/Devin), CodeGen, Open Models tracks) are now live! Subscribe to @aidotEngineer ...to get notifications of the other workshops and tracks!It’s easy to get de-sensitized to new models topping leaderboards every other week — however, the top of the LMsys leaderboard has typically been the exclusive domain of very large, very very well funded model labs like OpenAI, Anthropic, Google, and Meta. OpenAI had about 600 people at the time of GPT-4, and Google Gemini had 950 co-authors. This is why Reka Core made waves in May - not only debuting at #7 on the leaderboard, but doing so with all-new GPU infrastructure and 20 employees with

Transcript
Discussion (0)
Starting point is 00:00:05 Welcome back, friends. It's only been a week since the World's Fair, and it was incredible gathering the community to see the latest and greatest in AI engineering. You can catch up now on the four live-streamed track days on the AI engineer YouTube, and our team is busy editing the remaining workshops and five other tracks, including the surprisingly popular AI leadership track. Thank you all for your support, and stay tuned for news about the next event, the 2024 AI Engineering. Summit. Last week, we did a very special deep dive with Josh and John of Inbue and Databricks Mosaic on training LLMs and setting up massive GPU clusters. And today, we're pleased to follow that up with a very special conversation with Yi-Tay, formerly tech lead of Palm 2 at Google Brain, and now chief scientist of Raker AI. Raker's largest model, Raker Core, was at launch, the fifth best model in the world,
Starting point is 00:01:05 and the only GPT 4-class model not trained by a big lab like OpenAI, Google, Anthropic or Meta. In fact, while Google Gemini has 950 co-authors, Raker only has 20 employees, only five people actually working on pre-training. One year after our RWKV episode, Swix was excited to return to Singapore to delve into Ye Raker and building a new AI model lab outside of Silicon Valley. Stay tuned to the very end for a special bonus clip from Yi's recent appearance at the Tekkenasia meetup
Starting point is 00:01:42 for his spiciest take on why senior management is overrated, and why this is the time to build up senior 10,000 X individual contributors. Watch out and take care. Welcome, Yi Tay to Layton Space. This is a long time coming, but I'm so excited to have you here. Yeah, thanks for inviting and excited to be here. I'm glad about a lot of stuff, yeah. So you are interesting to research and introduce.
Starting point is 00:02:06 You are now chief scientists of RECA, which is a super interesting model lab, but before that you were at Google Brain. You were architecture co-lead on Palm 2. You were inventor of you all two. You were a co-contributed on Flann. You were a member of the Bard Corps team, and you also did some work on generative retrieval.
Starting point is 00:02:21 That's a very, very illustrious three-year career at Google Brain. Yeah, thanks, thanks, yeah. And then since then, Rika, you joined in March, 2023 announced a $58 million series A in June 2023. I don't know if you know the post-money valuation or the pre-money valuation is public. So crunch basis is 200. I did not know that. 50-something million. So you don't even have to leak. It's on the internet. Okay. Raker's stated goals were to work on
Starting point is 00:02:45 universal intelligence, including general purpose, multimodal and multilingual agents, self-improving AI and model efficiency. In February, you release Raker Flash. In April, you released Raker Core and Edge. And then most recently, you released Vibe Bell. Is that a good summary of the last six years? No, it's not four years? Four years, yeah. Oh my God.
Starting point is 00:03:03 Okay. We're talking about AI. Yeah, I was like, I'm wondering like since when did I like step into a time machine or something. Yeah. Okay, so can we just talk about your transition into, you did your PhD and we can talk about your PhD,
Starting point is 00:03:13 transition into brain and research and all that. You know, I saw you do some work on recommend your systems. I saw you do some work on quaternions. What the fuck was that? Okay, let's forget about that. Describe your path into modern elements. LMS, right? Because you didn't start there. Yeah, okay, sure. I think the world also didn't start there, right?
Starting point is 00:03:33 I mean, I think in, so I joined Google in 2019, end of 2019, and the world looked like really different at that time, right? I think there was around the time the first GPD was released by GPD1 or something was released by OpenE Eye. So research, like ML research and NLP research looked very different at that time. So I was mostly, I identified as like a language researcher. I don't like to use the word NLP. Jason will kill me if I used the word NLP, but I was like, okay, a language research, right? But I was more like an architecture kind of researcher. And when I joined Google, I was also, I continued on this, like, as a model architecture
Starting point is 00:04:07 research. I worked a lot on, like, efficient transformers. There was your first viral paper. Yeah, yeah. And I worked like long-range arena. I spent quite a lot of time looking of like could we do without attention. Like, there was a synthesizer paper back in 2020. I think that was like my early days in Google.
Starting point is 00:04:22 There wasn't like a, at that point of time, transformer research was mainly like, like, WMT, like machine translation and like Popaxity and stuff like that, it's not really about, you know, that there wasn't like, I think only feel short, feel short in context learning, came only about when GPD3 came out and beyond, right? And then, so I think that at that time,
Starting point is 00:04:40 the meta, I would say, the meta looked very different. And at the time, a lot of the work will focus on like fine-tuning things like T5 or bird or something like that, right? So I think a lot of the research, not only myself, but like around me or like even the broader community were working on those kind of things. And so I think that was, which I feel that in hindsight today,
Starting point is 00:04:59 it's actually pretty useful to like kind of think about because a lot of people came into like AI and into right after chat GPD came out. So they saw AI as kind of, I think there's a lot of benefits of, you know, understanding how transformers and I've broken this thing apart so many times, it's like these things actually help to improve intuition and it's not totally disconnect. I think a lot of things are still relevant today. and it's just the scale has gotten much larger and also the paradigms shift a little bit
Starting point is 00:05:28 from single task fine tuning to like generally do everything kind of universal foundation models. Foundation models, right? I think it's just a slight change in paradigm. But fundamentally, I don't think the underlying principles of research hasn't really changed that much
Starting point is 00:05:42 except for like compute. So basically algorithms stay put and then compute and data scaled. So I have some thoughts about this, right? So I think back then a lot of the academic research. I think people have talked about this, like, like, Sasha Raj has talked about this, or like, other people have talked about this. It's like, the conferences were always organized by like applications, right?
Starting point is 00:06:03 They were always organized by like, oh, like question answering. It was always by this, right? I think there's like a bit of a transpose going on. Things become universal and then becoming like, okay, there's a data work stream. There's a model architecture work stream. And then people work on improving like a universal model and general purpose algorithms to improve this model rather than finding domain specific tricks. I think for, even when 2019, I think I've already been like focusing on works that are like, you know, you could improve one general architecture.
Starting point is 00:06:32 At that time, it was like, like, maybe LSDMs in 2017 or something. And then you try on like 10 different tasks and the kind of thing, right? But like a lot of the research community have been focused more on like, how do I get that extra 2% on question answering or like and then sediment analysis? I think that there was this phrase of like in 2017, 2018, where this type of work was still very fashionable in academia and conferences, right? And then I think the big thing about the chat GPD moment of like 2022, the thing that changed drastically is like it completely like, it was like a dis-sharp,
Starting point is 00:07:05 made all this work like obsolete. So November 2022, you're saying, exactly chat-GPC launch, because I feel like if you're in the research community, this was coming. Yeah, yeah. So I'm saying that the big labs and stuff, people have already been moving towards general. Even T5 was already general purpose. Yeah. And the thing, right, but there was a bit of a time, basically,
Starting point is 00:07:21 just like Google and meta, OpenAI, we will be working on things three years ahead of everybody else. Then academia would be still working on these past specific things. Got it, got it, got it. And then I think the forcing function was the chat gibbet moment actually really like, it was coming, it was coming. It was just the final, the last straw. And then it's finally like.
Starting point is 00:07:38 Yeah, now it's serious. Yeah. Now it's really the thing completely changed. I don't know how it turned from my background to like talking about the meta. I think that you navigate the matter very well. And part of my goal here is that also. also isolate how you think about the matter for other people to reflect on, because I think obviously you do it very well. Oh, thanks.
Starting point is 00:07:57 So I'm looking at your papers published. Somewhere around 2021, you had a hard cut to 22, YOL2 and Palm. You did YAL2 Palm, emergent abilities, DSI, recitation, augmented generation, all in the same year-ish. So like, there was, did you change teams? Did you, did you, like, have a research focus? Like, when did you become... Oh, so you're saying that, like, my research model guy. My research became emergent, right? It was very obvious. No, I don't think I'm like a person that like, I'm not like super, super great at like forcing a trend like two years like a head and then especially like especially like plan for that.
Starting point is 00:08:34 Yeah. I think I smoothly and as like kind of like as the few moves. I never actually really thought about this. This way I just at every step I just optimise for what I found to be most impactful and most promising. And then that gradually. And also it's also a lot of influence by talking to. people, right? At the time, I started working more with, I had some close collaborations with
Starting point is 00:08:56 Jason and other people. I mean, Google is, you can work with anybody you want, basically. So, you're kind of partly the environment shift. And I think the environment shifts very quickly, but I was always polling in the environment. I was not, I think it's always good to have an open mind and then move along with the field rather than, okay, this is my research area. I'm going to get stuck really two years. I think I just move along to find things that interest me. And And naturally, I think that turned out to be the things that were most impactful at that time. In retrospect, I kind of did well, but I never actually really saw it as intentional. Sure.
Starting point is 00:09:27 I didn't do anything really intentional, except that's doing what I find interesting, actually. Cool. Well, we'll just talk about the main work at Google Brain, and then we'll move to RECA. So out of UL2, Palm, Emergent Abilities, which of these came first? Actually, I can't really actually remember. Okay. We'll make you talk about UL2 then. UL2 and DSI, the differential bus search index.
Starting point is 00:09:47 I was working on it the December of 2021. So at Google, there are projects that are big efforts that a researcher will be part of the effort. And then this will be kind of top down to some extent. And then they will also bottom up research that one could do. I can't speak for the Google now for sure, but at least at that time, right? So UL2 and DSI differential search index works that I kind of tinkered with in the December break where nobody was around. Okay. So there's differentiation because there's Palm 1 and there's Palm 2.
Starting point is 00:10:16 right. So Palm 2 I was actually a co-lead of one of the work streams but Palm 1 I was more of a contributor and Palm 2 I was so now I'm to think back of okay what's the timeline which came first right? In general there were three categories of works. One is broader efforts that org level efforts and then there are some that
Starting point is 00:10:31 do and there's I was my own projects. Projects. I used the compute that I had and then I just played with it. You accidentally left the or two running for a month. Yeah, yeah. Yeah, that was in the paper. It was fun. It was really fun. And then there was also the third category where those were the efforts that
Starting point is 00:10:46 my good friends were driving and I contributed. So Flann was just one of the, I would like to maybe say this publicly. You're very publicly. I talked a lot about Flan. Your Flan show number one. But like, yeah, but like the first author is actually Hyeong-Wan, who is great. And then like another guy, I was a core contributor. But I mean, just because I'm a little more visible, so I kind of accidentally took a little bit more credit for that.
Starting point is 00:11:07 But as in, I was a core contributor, but I was not like. The lead authors are obvious. Yeah. So the third categories were like projects that my friends, like emergence was also like, Imurgent abilities. No, actually, that paper was actually supposed to be only me and Jason on the paper. And I actually became friends with Jason from that paper. And then that led to this streak of, like, I don't know,
Starting point is 00:11:26 10 papers or something together with Jason. And now we're like super good friends. The ultimate romance. But that was like the emergent paper. But emergent people was also like a bottom up kind of like a thing. And fun times. Yeah, it was fun. Maybe I'll pick on Palm 2 because I feel like,
Starting point is 00:11:43 I'll pick on Palm 2 in emergence. is I really want to make sure I tell those stories. Those are important stories. Palm 2, I think it's a career story. Effectively became a co-lead on the second version of a very high-profile company-wide effort. How did that happen? I think people would like to know what's like the career strategy there. To be clear, I was one of the co-leads, but there were a lot of co-leads.
Starting point is 00:12:04 So I don't want to take too much credit for that. But my involvement with Palm 2 came from after UL2 was working well, and then it was gaining some visibility within Google. Was YL2 the largest model that Google had released at the time? Yeah, I think so. That was the largest. And you just, it was a personal project. It was a personal project.
Starting point is 00:12:23 Yeah, yeah. Isn't that unusual? How can it be like one person's decision to like suddenly release something that effectively changed? I think how we work worse? I mean, 20B is not that much larger, but from 11B, the 11B, T5. Actually, at that time, there was 13B empty 5, right? So I think UL2 is an encoder decoder 20B model. I think when we got it approved, it was released as like the big brother of T5, you know,
Starting point is 00:12:49 kind of like, okay, we updated T5 with like a new objective and train this new model and do DBM. You want to, and it uses the same pre-training dataset and everything, right? So like from PRC4. Yeah, that was the easiest because there was precedence, right? Yeah, yeah. But yeah, there was some architecture like the mixture of denizers. Yeah, yeah, yeah. So back to Palm 2.
Starting point is 00:13:08 I think my involvement with Palm 2 came from the work to add UL2 to Palm. to. And then, I mean, it was from the top-down point of view, I mean, the leads were decided in a top-down manner. It's not like, like, what, there was not much, like, fighting or, like, or any major things, right? It was like, it was a mixture, like, bottom-up, top-down-ish, like, half-half situation. And then, like, from the top, it was like, okay, like, these are the people who are the most visible in contributing to this workstream. And then, okay, how about E and this other guy becomes, would be in charge of this, like, modeling workstream and something like that, right? So I think it just happened that way organically.
Starting point is 00:13:47 And yeah, I think that that was how I kind of was co-leading the modeling upstream of Palm 2. I think in retrospect, you understand now that this is very valuable experience. And I think now, today, it will be much more competitive to get the job that you got, whereas you didn't, two years ago, you didn't have to try that hard to get it. Or like, you kind of lucked into it with you all too, and then like it just compounded from the initial good decision. I think it's very hard to counterfactually analyze this type of things. I think it's definitely true that there are more people working on generally AI now. And if you are in a big company, it's way harder to navigate this type of things, right?
Starting point is 00:14:21 I wouldn't say that there were like nobody or so wanting to work on this at that time. In fact, they were actually... Well, you were the obvious choice. There were less people. There were definitely less people. But I think I would say that maybe it's slightly harder now, but it's also not like it was easy at that time. Yeah. Yeah.
Starting point is 00:14:36 I imagine it's sensitive. But also in my mind, this is now the most... valuable on the job training in the world. And so people want to know how to get it. This is why I'm trying to figure out. Like actually individually, we also cannot take like somebody else like experience and then try to replicate it on. Because everybody's circumstances, their initialization point,
Starting point is 00:14:55 their thing is kind of also like indifferent. Yeah. This is not only true for ALMs in general, right? Because a lot of times like, oh, okay, you did this in this position and then because of this is, like, it's very hard to trace all this style to find a causal path of this thing. So I think everything in life, there's some luck involved. Yes. Yeah, there is. Emergent abilities, very influential paper, subsequently contested by the Mirage paper. Oh, yeah, yeah. So before we get to the Mirage, was there a story behind emergent abilities?
Starting point is 00:15:20 Yeah, you know, I'm sure it's Jason's thesis or like, just tell more about like the behind the scenes. Like was there a discussion that led to it that? This one was like, this is the idea, the inception of it was like mostly Jason. Okay. Right. I think I helped out to like, you know, shape up a little bit of the paper, get some, stakeholders involved and stuff. I was discussing quite a bit with Jason, but this, the idea itself was Jason itself. So actually when the Mirage thing and everything came out, I didn't, okay, I was this hot takes for the sake of hot takes. I didn't feel, I believe in emergence. I have to just go on
Starting point is 00:15:52 the record and just say, I believe in emergence. And then I was not feeling very strongly because I think that I can't speak for Jason, but I would as imagine that he would be maybe personally offended because, because I'm thinking, I know, Jason is a person that takes a lot of like feedback like very well. He's a very, like, he's not offended by harsh feedback, and he rebuts well, like, online as well, right? But, like, I would as imagine. He would be the one that is the most, like, affected by criticisms of emergence. I was believing in it, but I have to say that that paper. I mean, that's why he's the first author and I'm second, but that was mostly Jason's thesis. And I have to really say that Jason has really good ideas. And I was more of a support role for that
Starting point is 00:16:32 paper, yeah. Sure. Yeah. Lots more to discuss there, but you believe in emergence, that's enough for me to work with. I also think that the Mirage paper is mostly like, I don't know who, actually I don't even remember, who wrote it? Rylan Schaefer. I covered him on my Nirobs podcast. Okay, okay.
Starting point is 00:16:48 He's a very good speaker, and the paper was well done. It's just that people drew the wrong conclusions from the paper, because he had a very good title. Do you believe in emergence? Of course. Okay, high five. I mean, how can you read any paper, read any, the progress of LLMs and not believe in emergence?
Starting point is 00:17:04 It's so stupid. Like, just because you read, paramaritize some benchmarks and evils and make it linear, doesn't mean the emergence is completely gone. And even in the Mirage paper, they acknowledged that there were some metrics that were true genuine emergence according to that. I think it was something like 25-ish percent in the ballpark. That's not the exact number. Yeah, yeah, yeah, yeah. So I was like, okay, fine, like some benchmarks you disagree with, but on the whole, there is emergence. It's just, now we're just talking about the magnitude. Yeah, yeah, yeah, for sure. I think, I don't think the authors of the paper had really very,
Starting point is 00:17:34 like they didn't, I mean nobody, we should just assume people don't have bad intentions, right? But like, they definitely were just doing this. But like, I think I was, I was more like annoyed by the nearest best people. I mean, okay, best people was, just take a bit of a grain of salt, right? But like, there were people who come on me like, oh, you should care about this because it's the nearest best people. Because it's the nearest best paper. I'm like, does best people awards mean anything? Actually, it doesn't mean anything, right?
Starting point is 00:17:57 Like, I think that was more of my, where my angst was coming from. I don't think I really had. I don't even remember who were the authors of them. That paper, right? I'm sure they're doing well for themselves. We don't have to dwell too much on that. Okay, okay. Okay, so a couple more things on Google, and then we can go to Reka.
Starting point is 00:18:12 Quokler was a manager. Yeah, yeah. What is... I had another manager called Dawn. Like, I had two managers during my time at Google. So I'm just basically going to ask for quick hits from what did you learn from Qok? What did you learn from, you know, one? Oh, okay, very interesting.
Starting point is 00:18:24 Who they represent to you, like, how they advise you and all that. So Quok, as a manager, he was more like a friend, and we were, like, talk a lot about research. I think Quok is a very researchy person. he has a lot of like good, like, he's more of like intuition. I learned a lot about like, from him about like, there was no like concrete, like, it was more like over time and it was very implicit soft kind of feeling. But I think like a lot of research science, we were like brainstorm a lot about like, I quite quite like that when we were, but there was this Yelpam paper that didn't like get
Starting point is 00:18:50 as much attention that I feel it deserves. But like I think that was one of the works that I kind of like discussed with Quark quite a bit and like at that time we were releasing the Flan 2 stuff and everything. And then like I think Quok has a lot of good sense about like, like what makes a work a good hit and like, like, you know, publicly a good hit and like a lot of research sense about like what, what makes like research cool. So I think he has good like intuition as a researcher and I learned quite a little bit about. And I also say that I think Jason also probably learned like quite a bit from Quok and there's also influence his like, like,
Starting point is 00:19:21 it was not only just like me getting influenced, but like there was like Jason getting influenced and then Jason influenced me. So I think overall what I learned from Quarks probably is more like intuition, research tastes. We would like chat about AGI sometimes, singularity and stuff like this. Like it was like, it's nice to talk to as a friend manager, kind of, he's like kind of a friend figure to me. He's very much a researcher more than like a corporate manager kind of thing.
Starting point is 00:19:48 I totally expect that. It was fun. It was fun. Jason way, would you learn from him? What is your distillation? Okay, Jason is very interesting. So I learned in my career, I learned two or three things, major things from Jason, right?
Starting point is 00:19:59 So I think the first thing I learned from him is that So Jason was actually, okay, I'm going to talk about the more casual, more fun stuff first. Jason was the more spicy on Twitter first before me. There was an era where I was Goody Two Shoes. I only had my main account. My only tweets would be like New Paper Alert, right? And then Jason was starting to post like hot takes, right?
Starting point is 00:20:19 And I just thought to myself, oh damn. And there were types that I was like, Jason, you should not post this. You're going to get cancer. Right. And he was fine. He always braved through the storm and everything. Until I looked at him and I'm like, maybe it's not that bad after all. Do they just be, right?
Starting point is 00:20:33 So that was like kind of like, which is very interesting because Jason is much younger than me. And I and the other thing also, we out of accounts, right, we created them around the same time. And the interesting story behind it was that Jason's all the account and my account our own original. It was not like an anime character out that nobody, I know who is it. We have our identity. It's pseudonymous, right? And then I asked Jason, why do you want to have a pseudo? like, why don't you just make, like, right?
Starting point is 00:20:59 And he told me this thing, which was quite true, was that like, okay, you can post a take that it's spicy and it's hot, but if you cannot stand by the opinion, then you should not have the opinion in the first place, right? Wow. Right. So there was something that, oh, okay, I thought that was profound, because so far this, I mean, there are times where, okay,
Starting point is 00:21:12 I post something and it's spicy and then, okay, it gets a little bit bad. And then I, okay, I cannot agree that, okay, this is bad. Then I will retract it. But if I could stand by the opinion, then I would just stand by it, because that's the point of making it. It should be said. Right? It should be said because I can put my name behind it.
Starting point is 00:21:26 So this is part of the first bucket about how kind of influence my online persona, like, a little bit. And then I mean, it turns out that now AGIHIPO is so much more spicy than the cola. Coala is just hibernating somewhere. It's not even around, right? So, I mean, Jason also is more constrained because he works for, he has like an actual employer, right? And he has to be a little bit more. My God, the worst thing about Twitter, you know, anytime anyone from openly I tweets anything, they're like, did you see this researcher from openly?
Starting point is 00:21:54 I said something. And they read tea leaves that are not there. And it makes you very cautious to tweet anything. And so it kills the golden goose, is what I say. There was one tweet. I mean, at the time when somebody was, people were speculating the GPD two chatbots, right? And then Jason just posted something on his main account, something excited about new experiments being run.
Starting point is 00:22:12 Just a random. And then people screenshot that and posted. Yeah. I hate that. So I think, now I think for all the account is mostly like personal, like personal stuff. very personal. I think he would stay away from non-work things.
Starting point is 00:22:24 The golden goose has been killed because people on Twitter cannot control themselves from drawing random conclusions from all these hints and all that. Yeah, yeah, yeah. But going to like the actual, like this is like filler,
Starting point is 00:22:35 this is filler. This is filler. I think the second thing I learned from Jason is more about like, from my, you know, kind of like, from my own career is like the importance of like marketing and PR. So Jason is actually like super good at
Starting point is 00:22:49 I mean, I would just like, he was actually like really the emergence, like how many blog posts he wrote about the emergent abilities and how many talks he's given about about immersion. Like a lot. Probably like the other day I was just at this webcam keynote and he was giving a keynote again about immersion abilities and it's been two years, right?
Starting point is 00:23:05 So I think one big success of him is that like he does the work. He thinks a lot about like marketing the work itself. I did not like in my early parts, my career, early parts in Google, right? I was putting out a lot of work, but I didn't put in a lot of like effort in like thinking about the like how the word is going to be received, I'll just be like, here's a paper, here's a paper, here's a paper, right?
Starting point is 00:23:25 But Jason will be like, I'm going to write his paper and I'm going to like market the shit out of it. So I think I learned a lot about like, so every single first author paper that Jason writes in the last, has like 1,000 citations in one year. Oh my God. Like, no, I mean, not every, but like most of it that he leads. So his head rate is very high.
Starting point is 00:23:41 He's hit rate, like impact density. Like it's very high, right? So it's pretty interesting. But I kind of see him as like a peer and I learn a lot from his basically. Some people are just like talented in, different ways. And I think that, like, I looked at how he markets his own work and markets himself, actually, right? If someone is studying from zero, like no Twitter presence, what is the second best thing to do? You mean as a researcher? For marketing, yeah. I think you were like, the most obvious
Starting point is 00:24:05 thing to do, like if you're like a research, like say hypothetically you're like a researcher in like a place without visibility or without, and then you have no personal visibility. The first goal is always to try to find a mentor or co-author that is like within this circle. And then you start from there, right? Because, and then you get people from, like, who has a visibility and following to retreat. So you would, you would like work with them. The big goal is not about like, I learned this. I mean, this is like probably a career mistake in my early days. It was that, you know, instead of like focusing on like, so-called people, okay, if you do good work, it's more of like, okay, how am I going to? I see this visible researcher from deep mind,
Starting point is 00:24:42 right? Or how can I collaborate with this person and then kind of do something that feel is cool and like I can win their respect and that they were like, you know, they would be willing to co-author for me because the exercise itself also about how to, you're not trying to place reviewers or anything. If you can find one semi-visible, you don't even have to be like a famous person. That's like a semi-few thousands of followers, has a good reputation of research. And then you collaborate with this person. And then when you post the work, you are co-author with this person.
Starting point is 00:25:12 And then you get the person to vouch for you or like this. Over time, this would, it could be from internships. It could be from just DMs. I think people are nicer. Then some people, they seem scary, but if you DM them, they're actually reading to collaborate, actually. I was scared of you, actually. No, no, no. And when I DM you, you turned out a lot nicer than I feared.
Starting point is 00:25:29 So thank you for being nice. That's really great advice for people. I just want to leave that all that for people. For others who follow the work that, the career advice that I give, the title topic of this is pick up what others put down, and specifically pick out what your mentors put down. Like mentors always have more work to do than they have personally time for, the high visibility mentors. and if you can show that you're a good collaborator with them, they will lift you up accordingly.
Starting point is 00:25:54 That's a pretty good formula for career growth. Should I ask about Hyeong-1? I don't know how close you are. Oh, we're still good friends. Hul-one is a great engineer, and he's very systematic in the way he thinks. I think Hyeong-1 is, without going into detail too. I still spend a lot of time talking to Hiong,
Starting point is 00:26:10 even after we both are different places, about very interesting, algorithmic ways to think about life. Very interesting. like perspectives on life rather than research. But he's a great engineer. And the one thing that scares me about, he doesn't have multiple monitors.
Starting point is 00:26:28 He just codes you one small screen. And he does everything with very hyper-optimized. And then back at... This is, I want those you curve where one screen, one screen, and then many screens. Yeah, yeah, yeah. So I think Shawah would feel. Shaw one scares me because it's like,
Starting point is 00:26:41 I think that was at Newark's 2022. Like, we were doing some work at the New Orleans. And then he would be like coding like perfectly fine with like this, you know, 13-inch MacBook with like one terminal, and then he'll be like, he keeps telling us, okay, it's more optimal to using keyboard. Like, keyboard is more optimal than moving your head because if you can switch your screen fast enough, it's faster than your head, like moving to different screens and stuff. I did not like actually distill that because it's too painful to do that. But like, it's very interesting in a way that he belongs to one
Starting point is 00:27:11 of those hardcore people with one monitor. Maybe this is a relevant question to close out the Google side. What do you think is a good programmer for AI research? I mean, set up or you think? No, not set up. Not not set up. Not even lifestyle. It's more about skills. Like what should people have? What do you interview for maybe? What do you see the high performers do differently than less high performers? I mean, okay, like generally, there's like, I think like for, for AI researchers, like being a strong IC is like probably like the thing that I feel like it is like important for AI researchers. Like not, not, I think there's certain level like sacrifice to be like.
Starting point is 00:27:47 like an AI engineers, that AI research, especially if you're training like L. and because you cannot really be detached from, your jobs could die on a Saturday at 4 a.m., right? And then there are people who like would just leave it dead until like Monday morning. And then or,
Starting point is 00:28:04 but there will be people who will crawl out of bed at 4 a.m. to restart the job or to check the tensorboard or something like that, right? I think a lot of being a successful AI researcher, I want to say passion is also the entire thing, but it's more of just the, the kind of personality that if something there's a buck at 3 a.m on Saturday night or something, right? And then you would be like, you couldn't go back to sleep unless you, you, I'm not, this is very unhealthy by the way. Like people should not do this for a long time. You know,
Starting point is 00:28:31 I think this kind of things actually like allows people to make progress faster, but it's unhealthy. So I'm also not even sure like what's like the checking out on like Friday, Saturday, Sunday and like work at 9 to 5 if you want to like make progress or like some people are just so good at detaching like, okay, like 8 p.m, I'm not going to, my job can die, and then the chips can stay either for like the whole night. But I want to watch Netflix, right? You cannot, I think there's a level, like it's like a sport. You cannot win an Olympic goal if you want to have super ultra good work like parents.
Starting point is 00:29:02 Yeah, passion, intensity, dedication. Yeah, intensity, right. So those are really good personal qualities. Just technical qualities-wise, how much of the stack should people know, you know? Okay, so there was the question. No, no, no, but that was important. well. It's just harder to interview for because you really just see it on the job. I think stack is not not not that like should know Kuda kernels. I don't know Kuda
Starting point is 00:29:24 kernels. Exactly right? Okay good. For all you listening out there you don't have to feel like an imposter. No but but you need to be willing to learn if you have to I think. Well you haven't had to so far. Yeah I haven't had to so far right. So if I think pie torch, okay great you know what kind of do what do I know like distributed systems like do I know like what is the stock that you recommend for people that gets like a well-rounded end-to-end researcher. I don't think there's any specific thing. In fact, I will try to...
Starting point is 00:29:51 I don't really say like, okay, you need to learn jacks. You need to learn this. By the time you finish letting there's a new framework out. Anyway, so it's more of like staying, like, constantly, like, trying to... Being able to continuously learn and update. I don't think that's a single stack or like a single, like, workflow. I don't think there's a single one. Yeah.
Starting point is 00:30:11 Well, that leads us to Rika. What's the founding story? So I met some of my other... co-founders while we were collaborating at D-Mind. I was at Brain and they were like a deep mind. I'm not like a startup person. I identify even today as a scientist and a researcher more than like a startup person, right? My co-founder, Danny, started this story, right? And then this record was like in the works from like late 2022. I finally left in 2003. Danny kept asking me he wants to do something. Do I want to go with him and do it? And it took a while for me.
Starting point is 00:30:44 I was like kind of the last co-founder to kind of form. Was the plan always for you to leave at some point and joined him? No, no. He was just like convincing you to do it. It was like six months, more than, in fact, like, I think more than six months period of like, I always had this at the back of my mind for since like August. Actually, I didn't want to do it in the first place, but I think eventually in March, I felt that, okay, it's time for me to experience something new.
Starting point is 00:31:09 Like my leap of faith was more of like, I want to experience something new. I've, okay, I've like wrapped up this palm to work at Google and then like more of like, okay, let me experience this new life and see where we can go with this. And I also, I mean, we don't have a lot of like, okay, the funny thing was that like many, many years ago before my PhD, I wanted to do a startup actually at that point. And then over time, I realized that like I was better off as a researcher and I just forgot about the startup thing. And it's quite funny that today I end up doing a bigger startup, right?
Starting point is 00:31:37 But even until now, I actually identified more as like a researcher and scientist. Well, I mean, it's not, when you left brain, you already had a high profile coming out of brain. You could have gone to any startup out there. They all have wanted you. Yeah, okay, yeah. So why did you choose this one, basically? Like, it was just because of preexisting relationships because it wasn't obvious to me. A lot of the other coworkers went to Open AI.
Starting point is 00:31:59 Others went to, if you're fair, you went to mistrial, that kind of stuff, right? Like, Rico was like not on the, on the map. I think it was, for me, it was a mission between staying at Google and, like, co-founding something. I didn't want to, like, it was more of the experience of being a co-founder that, like, was attracted me, right, and wanting to experience that I wouldn't have left for inflection or something like that. Like, I mean, inflection is gone now, but like, RAP? They're still alive.
Starting point is 00:32:25 They're selling themselves as a model foundry or something. I don't know. They're a services company now. Yeah, I don't. But I also think that, like, like, for example, if you to join, like, another, it would be, like, a very big tech experience again, right? I don't know. I felt like the experience I get is very complex.
Starting point is 00:32:40 inventory to what I have, that's the experience I had at Google, right? But if I were to join something else, right, then I wouldn't have, I would have just stay at Google, to be honest. Because to me, it was very clear, just two dishes that I didn't really, I was talking to a bunch of other startups, but I didn't really actually have the intention to go. I was happy at Google, actually, to be honest. I'm sure. Yeah.
Starting point is 00:32:59 I'm sure they have a lot of things to keep you happy. I was happy at Google, yeah, actually. So you described yourself as GPU poor, but also you had $60 million to play with. you got a whole bunch of GPUs. I think you disclosed somewhere, but I don't remember the exact number, and you had a good training run for Flash and then Core and Age.
Starting point is 00:33:17 How would you tell that sort of the story? Like, people can read the technical report, but also, you know, what was that overall experience like? And I should also point people to the blog post that you wrote. There were a lot of interesting things that happened along the way
Starting point is 00:33:31 that led to our... So I think I left around like early April, like March, end of March, April and everything, right? But most of our compute actually came in December, actually. Yeah. And there were delays. So H100,
Starting point is 00:33:42 they were major delays, right? So we were sitting around, right, bunch with like... It's clear you don't own the compute, you're renting.
Starting point is 00:33:47 Yeah, yeah, yeah. So we were sitting around, like, with, for a long period of time, we had 500 A100, because we made a commitment and they were constantly
Starting point is 00:33:57 being delayed. I think because H100, supply demand, whatever, reasons. And it was also very hard to get a lot of
Starting point is 00:34:03 compute in one place, right? And then we were locked in, and we had to wait for the compute to come. Right. So I think it was very painful because even when the compute came, it was mostly broken most of the time. And it was broken to a very bad extent that before I left Google, I was like, even the early stage, I was very optimistic about, okay, this compute translates to do this amount of flop.
Starting point is 00:34:24 This is a model, right? But I never expected the reliability to be so poor that it just threw off all the calculations. And then we had to work 10 times harder just to make the thing go smoothly. So it was a bearable pain. I think the pain was bearable, but it was. just way more than expected. I think you'd address this in your post, but the temptation would have been just to run everything on TPUs, which is the stack that you already know very well,
Starting point is 00:34:47 that works there. No, no, no. So TPUs outside Google and TPS inside Google are probably very different things. Oh, how come? Okay, firstly, it's like infrastructure. Like, there wasn't like a lot of good code bases like outside Google that was like still, right?
Starting point is 00:35:01 And the code base that I was most familiar with was like T5X, it was a jackspace. It would have been like, by the time we wanted to consider it, it was really like debricated for nine months, right? And then TPUs, I mean, we weren't sure about, I mean, the availability of TPUs were also not great, great.
Starting point is 00:35:19 Oh, my perception is it was a lot better. People have the learning curve. Yeah, but at the point of time, we had our infraset up. We were training already training models and it would be so much cost to switch to TPUs. So I think TPU's, the experience of TPUs inside and outside Google, I have not actually run a single TPU job outside Google, by the way,
Starting point is 00:35:34 but just looking through documentation from what I see outside, and from like how much I think that people inside Google don't care about what people think outside Google I kind of feel like okay we were a bit like I don't think we considered
Starting point is 00:35:46 I mean not like forever not considering this but like just like at that point of time it was like The obvious choice is just stick to pipelines does stick to GPUs and Pythodge and make like I mean it's not as if the chips we ordered were not there they were there they're just not in the best shape
Starting point is 00:36:00 Right yeah so I think it was too much work to kind of migrate suddenly to DPUs yeah For those who haven't read the report, you had a very traumatic description about the chaotic and stable phases of various compute providers. And I was just winceing when I was reading all those things. Yeah, there was like a three-body problem reference to chaotic and stable phases. I mean, I was watching three-body problems at the time. And I thought it was fun to be, it was fun to.
Starting point is 00:36:25 There was a lot of like, I think we had a lot of fun adding a lot of references and mean into the type report. I think it goes to show like how fun the environment is within record, right? We had a lot of fun with this. But so I think chaotic and stable phase, mostly it's like we actually found that like usually when like provider provisions new notes or they would like. Yeah, you don't want to be the first to use it. Yeah, it's usually like bad like dog shit like at the start. And then it gets better as you go through the process of returning notes and draining them, giving it back to them. They will send it back for repairs and everything.
Starting point is 00:37:00 And then over time, because it's more of it's a more of a numbers game, right? if there's one bad note, it kills the entire job, right? So, like, the fact of the game became like, just eliminating bad notes from the thing, right? And then, I mean, just because of, maybe because of the supply issue or something, when the deadline comes to ship this, for example, like,
Starting point is 00:37:18 I just give rough numbers. Like, say you order 1,000 H-100s, right? They will not be able to, usually they don't meet the demand of, like, 1000s,000 h-100s at the date. They'll give you, like, 500 first stuff not to piss you off, and then it will give you like another 100. Like, every over two, three weeks,
Starting point is 00:37:30 they were just like, okay, I added like four notes, added like eight notes, that kind of thing. And then over time, we reach the capacity that you, or you actually, maybe you never actually ever reach the capacity that you ordered for. And then, like, as they add these notes, right, sometimes these notes are bad. And then they just kill entire training runs. And the thing, which I feel that, I mean, for all those people trying to sell GPUs, there are a lot of people trying to sell GPUs now. We sell, sell package, whatever, GPUs, right?
Starting point is 00:37:53 And I think the most important thing that there are obviously, there are SLAs, all this in, in the contract and everything. And obviously, you might be entitled to something, something, if something goes wrong, right? The thing that for large model training runs is that one bad note queues the entire job, right? So should the compute provider be liable to pay for all the node waste stage that... No, it's
Starting point is 00:38:14 unlikely because otherwise... It's unrealistic. No one will take that on. No one to take that on, right? So I think that's also like a tricky thing. Who is taking the risk? Is the LOM startup taking the risk? Or is the compute provider taking the risk? I think that, I mean, this is my sense. I'm not 100% sure, but I think
Starting point is 00:38:30 as there are more providers trying to sell our GPUs, we get all this inbound so much about people trying to sell us GPUs, right? The key differentiator is actually to find a way to balance the risk of no failure with, as long as the provider, I'm not going to say 100%, but if somebody can come and tell me that my notes are so stable that I can share some cost with you if your job dies, this is green flag, green flag, right? The moment they start to, I cannot do any of the big clouds do that as far as I know. Because they have the size to guarantee that. But I think for anybody who was watching. or if you do it like a compute startup or anything,
Starting point is 00:39:05 the biggest green flag would to be to share the cost of note failures with your customers, right? Because the whole run? No, no. Like if the note, it's very hard to go. Because you need software to like, you need software to. So let's say you run it for 12 hours, right? And it dies after 12 hours, right? You get 12 hours of throughput, right?
Starting point is 00:39:22 But then you get like some wastage because of like the, you know, the downtime and everything. Right. You know, I think it would be fair to find some middle ground to kind of split the cost. of the failures, right? And this brings back to my point about like work-life balance because if the notes fail so, fail so badly, right?
Starting point is 00:39:38 Like, it actually, basically, right, your engineers cannot sleep at all. Do you have babies sitting, rosters and everything, but you are living life with like constant anxiety because even in the case,
Starting point is 00:39:47 right, where the note failures are refunded, right? You still lose time. You lose three hours. Sure. You lose everything, right? So I don't know how to go around this, but I think if there are a lot of
Starting point is 00:39:56 like compute providers like fighting over, I think a good, good thing to do. do is to figure out like this pain point otherwise or at least figure out some hot shopping but but so far most of the providers that we tried don't don't have this they will also get confused when you try to ask them so my job is dead can you pay for the food can you refund for or at least they will get confused because like this is a LM specific thing that are the large notes
Starting point is 00:40:20 they don't care about yeah yeah they did they get confused about about this right so current status code is the LM startup pays for everything but maybe you could negotiate some like like refunds but usually they will not be so generous to pay for the say you you run 500 GPUs, right, if you break for four hours, they would in their mind, they would be thinking, I should refund you for one node, but in your mind, you just think that I should, they should refund you for the full job, right? Everyone who is from my background is going to be asking this, how is it so fragile? Like, what's your frequency of checkpointing? Our checkpointing is kind of like, we see how stable the job is, and then we decide, because
Starting point is 00:40:52 checkpointing takes a, without a good file system, checkpointing takes actually quite long. So it could be It's like a few hundred gigs, right? Yeah, I think so. I think so. I don't remember of But sometimes if your file system is slow, right, your file IO is slow, your checkpointing for a 20B model could be like, what, 30 minutes or something? Okay. I don't know this by head. Sure, sure. But it's not hours.
Starting point is 00:41:15 If you go larger, what if it's like a 200B model, right? Okay. So you should have some kind of ideal checkpointing to run ratio that is not catastrophic if you run into a known failure. Yeah, no. So we see of it as like an MFU, like because you can average out your flop utilization and then you can see how many percent hit, like how much slow down, right?
Starting point is 00:41:35 So you probably go for something like if it's like you're taking off 1% of your speed, 2% of your speed. So basically it's actually fine to the checkpoint more, more regularly, right? So I think checkpointing, like you also never fully, you can get like from the clean slate like nothing, right? As you optimize like engineer like the system to automatic restart,
Starting point is 00:41:54 everything, you get some of the time back, but you'll never be perfect, perfect. So you still lose, lose stuff like that. If you checkpoint too often like everyone, what, every 30 minutes, then your file system is going to blow up, right? If you're going to checkpoint every, like, like, for us, we just see as like how much storage is cheap compared to compute. No, when your model is like very, very large, your storage can, can easily blow up.
Starting point is 00:42:14 Going on to the models, I feel like I digress so much about all this fun side things. You like compute, right? You like hardware and compute. I know hardware and compute. And also I'm an orchestration guy. So one part of the question, one of the questions I'm skipping right now is there's, I came from temporal. I'm familiar with Kubernetes, I've used Airflow. These are all the data engine cloud or cloud engineer type tools.
Starting point is 00:42:36 It's surprising to me that you guys don't have your set of organization tools that it's solved. You wrote in your blog post you had like the pain of multi-cluster setups. And like to the rest of us, this is completely solved. Okay. I don't know if you know that. We use Kubernetes for a bunch of stuff. But like I think like for experimentation and like stuff like this is still not fully like we, we, We didn't have the time to actually build something that is.
Starting point is 00:43:01 It should exist in open source. Someone should have done this. Okay, okay. I'm not, it is what it is, but I'm surprised, that's all. Okay, okay, okay. It seems like a valuable problem and someone should do it. Okay, okay, okay, yeah, yeah, yeah, good to know. Good to know.
Starting point is 00:43:14 Okay, so Raker Flash Core Edge, you know, congrats on beating a whole bunch of state-of-the-art models, especially much bigger than each. People can see the papers for all the other stuff. Was this your expectation from the start that you would basically definitely be frontier? Like, how do you, like, from the start of, like, haven't trained anything yet, and you're about to kick off the runs. Like, are you able to call your shots and say, we will beat GP3.5?
Starting point is 00:43:35 Nobody can predict the future. No, how much confidence? Okay, we were confident. Like, we were confident. How? Why? All right. It's a good question.
Starting point is 00:43:43 Because it'll be a shame to do a whole bunch of work and then end up, this is the middle of the pack, which a lot of people end up. We were confident. I think that a lot of it was like Yolo. I mean, I mentioned in the, in the thing. I think we would, like, require a lot less iteration than this is because of our prior experience in like training these models. Like so I was confident in,
Starting point is 00:44:01 in myself about like, our models will turn out to be, to be, to be, to be good. And I, about exactly how I actually don't really like pinpointed to a particular reason of like,
Starting point is 00:44:13 I mean, we de-risk stuff. So a lot of part of it is like, de-risking and like, okay, you run like 4B as ablation and you can see, okay, this is like my spice,
Starting point is 00:44:21 if you run 4B and your loss is like going crazy, you know that, okay, this is going to be a shit model, right? But I think it's like, we train enough like, okay, we don't have a lot of compute to do a lot of violations, but we did enough experiments to know that, okay, our infrastructure and everything is set up to be good, right?
Starting point is 00:44:39 Obviously, the field moves, right? I wouldn't say that everything was like smooth, like the first time round is like smooth and everything. But I think we were confident in our ability to like make the least, like we're not like really, we're more confident about like the ability to like move with as little steps as possible to the goal, more so than my model is going to be this
Starting point is 00:45:00 level at this time you know what I mean it's more of like for example we let's say we run the first round of human evaluations right and then we see our number
Starting point is 00:45:09 as this right and then we're confident that in five more tries we'll get to this you know kind of like get to like like this is more of that kind of confidence
Starting point is 00:45:18 rather than actually it's also a little bit of you see a new leaderboard hypothetically like in like if as a researcher you see a release a new lead later about, right? You approach it like a puzzle. You don't know, like, whether you, at the
Starting point is 00:45:31 start of it, you might not have the answer to the puzzle. But if you're good at solving puzzles, like generally, right, you know, you know that with one hour, I'll be able to solve it. You know, that kind of confidence, like, it's the ability to, to heal climb or the ability to improve over arbitrary things, right? Rather than, I think we were confident more about that rather than everything is different, right? The stack is different. The infrastructure is different. The data is also different from what, I mean, we have a lot of, right? It's just, we have a lot of, yeah, we have a lot of experience from prior, like, our jobs, but like, it's not going to be the, like, we don't have actually, like, exactly the same thing because different companies have different stacks, everything, right? So it's more about derisking, being confident in, like, solving the general problem of, like, improving over things, which is why also I think that the team is valuable in the sense that we are not, like, valued by our model itself, but we are just valued about, like, like, like, how we can see one problem and we can just, like, solve it, like, super quickly, right?
Starting point is 00:46:25 That's what we're confident about, right? Actually, like, the artifact itself. Mentioning your team, you said at the largest, your team was three to five people on the pre-training side. It was that the team that you recruited? Was it all your ex-colleges? How do you find people that would have this kind of solid intuition? So I think that some of the people in our team were like, I worked with them at Google, at ex-colleagues and stuff.
Starting point is 00:46:47 And some of them were like fresh hires. Like they were like fresh PhDs or like everything. Okay. So I do want to comment on no architecture. So if you want to, people have variants of all these. Swigloo, GQA, rope, RMS norm, and then obviously the big one is encoder, decoder versus decoder.
Starting point is 00:47:02 Could you comment on each of those? Like, were you just like, we're confident and known got it right? Or did you actually do a evaluation of each of your architecture choices? Oh, I mean like, okay, architecture-wise is something that I feel like I'm easily able to, like, I've run so many architecture experiments that, like, I look at architecture and I, like,
Starting point is 00:47:21 I don't want to be like overly, I think it's very hard to outperform the OG G G. Why? I mean, on the surface of it, like, we have to have learned something in the last seven years. No, all the changes, all the changes that
Starting point is 00:47:32 like, like, Suiglou was this, like, okay, Suigu is like probably one of my favorite papers all of time just because of the divine benevolence. Like, the gnome actually wrote like, we owe this success to divine binel. Like that was like, it's always a meme thing, right? Okay, so like GQA, MQA was always like, the multi-currier that was always like
Starting point is 00:47:49 a big controversial thing because MQA usually you get a hit because it's MQA and everything. So people kind of know that, like, it was a very... Like, hit or miss. Like, you could get a hit in the performance from MQA, like, MQA alone. MQA was always like, you know, the choice, right? It's always like, okay, should you use MQA?
Starting point is 00:48:06 Should you not use MQA? Right? It's always, should you not use MQA. When GQA came in, right, it became like a no-brainer to use GQA because you don't get the hit anymore. Rish, right? The number two, the 70, GQA, right? But, I mean, the reason why we call it
Starting point is 00:48:26 Nome architecture, because MQA came from Nome and GQA was like a follow-up paper by some of my colleagues at Google, right? So I think GQA became a point where, okay, this is already accepted. Like, it's a no-brainer to use GQA. Sway-Glu was an interesting thing because there was a very long period of time.
Starting point is 00:48:42 It also Swiglou was a single-authored paper by Nome. And very few papers were, like, Swigrew had very few citations, like, at the start. Only Google papers were citing Suiglidu at one time. And a lot of them I was, at one point I was like probably like 30% of Sweet Glue citations. Because every time like, like, Swigroup became popular because of the updated T5, the T5, the T5 1.1 that uses Svigl, right?
Starting point is 00:49:04 And nobody actually really cared about Svigl for a long time. Because I was checking why is this like underrated paper like not getting much citations. And then I think probably now it has like a few hundred citations by now. But I think Svigrew is one of the things that like, you know, I played around with a lot like at Google. So Svroup really works. there was also a paper we wrote about, like do transformer modifications, blah, blah, blah.
Starting point is 00:49:27 Like, it was a paper with Norm and Sharon and Tianwan and stuff like that. And then we ablated like so many transformer variants. Yes, yeah, I saw that. Some of them matter, but most of them don't. Most of them don't. And then the only thing that matter
Starting point is 00:49:41 in that two paper was, in that paper was Suiglil. I forgot which exact sweet glue variant was it, but and sparsity at that time, right? So that was strong enough like to finding to... For the listeners, this is the inductive bias. Scaling loss versus model architectures, how does inductive bias...
Starting point is 00:49:56 No, no, not this one. There was another one, like, to transfer one modifications, something, something, something. Okay. First, portal was run, I think. You should run around. You gave the keywords. Yeah, yeah.
Starting point is 00:50:06 I think the RMS norm-Rope thing... Not controversial. Like, it's not like... Obviously, I think rope is probably, like, it has that extrapolation thing, which is nice. And then it's also like default now. Nobody wants to add production and embedding. anymore, right? And I think, I mean, I like the T5 style relative attention for a bit, but like, I think, okay, rope is, I actually ran that ablation for palm, like the T5 relative attention versus rope. I think rope is similar to other things, but he has this extrapolation thing, which is nice. And like... Which is why your long context version can go to 256. Okay.
Starting point is 00:50:39 This, for all, most of the long context models, they use the rope extrapolation thing, which is nice property, right? So that there was for rope. I think there was also like some things like the layer norm. like petitions and stuff like that that were like, it mattered a little bit maybe not too much and everything. But I think in general, there was not a lot of like, there are not a lot of things that people could do to the transformer to, it's been like four, five years, right? And then the vanilla transformer,
Starting point is 00:51:05 I think if you use it as it is today, would not be like that optimal, but like the transformer that we slowly evolve to now is like the gnome transformer is probably like very, very, very strong baseline that is very hard to like. I think you need a drastic shift to, to beat that, right?
Starting point is 00:51:21 Mistakes-based model type-ons. Or you could find more like, like, like, Swigrew is a small change, right? You could find like some small change that are like a big enough impact widely that don't cost a lot of because a lot of architecture changes, right? The moment they are tedious to implement.
Starting point is 00:51:35 Like nobody, when Su-Go is a simple thing, right? It's a very simple thing to implement. Maybe that's why it's caught on because it has like a additional boost that's for the simplicity of it, right? So there's also a bit of implementation lottery, if you will, right?
Starting point is 00:51:47 A little bit of, if you propose some very complicated thing for like 0.1%. Yeah, nobody will use that, right? The biggest, biggest, I mean, I can't believe we have, we're taking so long to come to this topic, but the biggest gnome architecture decision is encoder-decoder versus decoder only. So encoder-decoder is not like a gnome.
Starting point is 00:52:04 The gnome architecture is more the... Okay, maybe like more old-school transformers. Maybe we want to just talk about the decision on encoder-decoder versus decoder only. So, okay, I wouldn't be able to comment about like exactly our setup, but like, I think encoder-de-corder-es. are kind of very misunderstood from thing, right? So there's encoder decoder, non-causal decoder,
Starting point is 00:52:26 which is a prefix LM, and then there's a decoder only model, right? Technically, a causal decoder and a non-causal decoder are very similar in the sense that it's just a bi-directional mass, right? And then a prefix LRM and an encoder decoder has only, the only difference is that encoder decoder splits the inputs and targets into different non-shad transformer stacks, and then like there's encoder buttonnet in the data. end, right? So technically, people like kind of always associate like encoder decoders with like,
Starting point is 00:52:56 I like bird or like something. You know, people get confused about these things, right? But I think in the UL2 paper we really like kind of explored this and also like maybe some of the big science papers that also talk about this, right, is that prefix IOMM and causal decoders are very similar. That's the most. At the end of the day, they're all auto-aggressive transformers. That's actually like the only big benefit of encoder decoders. It has this thing called like, I mean what I like to call intrinsic sparsity. Okay. So basically an encoder decoder with like N parameters is like basically if it's like
Starting point is 00:53:25 it has the cost of like an over two decoder model. So it's a bit like a sparse model because you actually spend the same amount of flops. It's just that you have two sets of parameters like for encoder and decoder, right? So it's actually flop matched with a decoder model of like half the the parameters. So like a like UL220B is actually about a 10b decoder only model. So you get free sparsity from that. It's something that, okay, the old G-T5 paper talks about this. You can look at that.
Starting point is 00:53:51 There's this complexity chart. I didn't, like, when doing the UL2 paper, I kind of like was mind-blown by while encrode-corder is so much more, not bounded by the causal mass anymore. A lot of the efficient transformers, like a lot of the sparse transformers, like, I mean, the early days, there's like lean formal and like whatever things like this. They cannot maintain the causal mass, and that's why you cannot train proper language model with this, right? if you separate out your very long context into the encoder,
Starting point is 00:54:16 this encoder has no loss, right? You could just do like aggressive pooling. You could do some crazy sparse attention that has like final transformer or something like that. Right. And then you could make that smaller than the decoder. You could make that faster than the decoder. That are just some of the advantages of like
Starting point is 00:54:34 why splitting into encoder decoder. It could be beneficial to like just using like a decoder only model. At the end of the day, the decoder in an encodecoder is a language model. It's still a regular auto-rogressive language model. So there's actually, I mean, it's not that much different from like a retrieval or mentored language model. This is news to me. I don't know if you've ever expressed this, but yeah, this is actually, it makes sense. Okay, okay.
Starting point is 00:55:01 I don't know, unfortunately, I don't know enough to push back on this. But on the surface of it, it seems to make sense. Would you make the same choices if you were not so focused on multimodality? You know, that's one of the ways in which I was thinking like, oh, and, Dakota-decoder makes sense, that it's more natively multimodal. I just have to say that it's relevant. It's relevant. Yeah, it's relevant.
Starting point is 00:55:18 Yeah. Then we can move on to broader trends in LLM, just commentary on just ecosystem stuff, like completely independent from Rika. commented on a few things, like Lama 1 to 3 glowed up a lot. I call this the Lama 1 to 3 glow up. Like they improved into like an actual top tier open source model. Yeah. Phi, one, had a lot of criticism, but it seems like Phi 3 is getting a lot of love.
Starting point is 00:55:38 Do you just generally see like in your open model tier list Like what's going up and down? I think Lama 1 and Lama 2 are like quite mid, right? But Lama 3 actually got good, right? I think Lama 3 is actually strong, right? I don't really follow far and much just that... Their whole thesis is the textbooks is all you need thing, right? Like that we can use way less data than everyone else and still...
Starting point is 00:56:00 But I think you cannot cheat the scaling laws, right? Because like you... I remember thinking like vaguely saying that, oh, they match like Mixtra 8 by 22 or like something like that on like some... Okay, I don't think these academic benchmarks are like that meaningful anymore, right? So, but then like, then when you go, they go on IMCs, they get like, what, 47? And then they get like, maybe it just like seem slightly. Maybe it's like, five two. I don't know about five.
Starting point is 00:56:23 Oh, there's five three. No, I think. It was just released like yesterday. Oh, I don't even. Yeah, but, but I, I don't know. I think there's, there's some, I don't follow five that much, but I don't like, a model that is synthetically, actually, I don't even know this where, like, I didn't read the paper.
Starting point is 00:56:38 But I think that, like, a model that is like based on the premise of like distilling, and stuff, something like that is like not that interesting to me. But I think that like Lama Tree actually shows that kind of like meta got a pretty good stack around training these models. Oh,
Starting point is 00:56:51 and I've even started to feel like, oh, they actually kind of maybe caught up to Google now, right? They kind of feeling. There's also maybe a hot tag guy itself. But yeah, I mean, FI don't really kind of follow you that much.
Starting point is 00:57:02 And I just, yeah, I mean, there's too much, too much things to follow. So I think it's like, I, I think like,
Starting point is 00:57:07 Lama Tree is probably like the most, the first most legit. When you say these kinds of things, like most legit, obviously there's some, there's vibes Eval or whatever, but I feel like a lot of people, the very common feeling is MMLE is kind of saturated. Yeah. So like what do you look at now? Is it just LMSS?
Starting point is 00:57:25 Okay. So I think that LMSS has these problems also. Yeah. So Almsis is not like exactly like, I mean it's probably better than all these regular benchmarks, right? But I think like a serious LIM that's create their own Eval. And a good Eval set is one that you don't release. a good Eval set is the one that you
Starting point is 00:57:42 like, okay, you release some of it but like it's like you don't let it be contaminated by the community. Yeah, I think I don't see this is probably the most legit one. I mean, the things like GSM make human care human eval they're all like contaminated.
Starting point is 00:57:58 They're all like saturated contaminated. No. GSMK whether you're 92, 91 like no one cares right, a kind of thing, right? But we'll still report three decimal places in all of our reports. Yeah, yeah, yeah, yeah. But it's kind of like almost like this like obligatory thing to do. You're a table of numbers of your thing at the bowl.
Starting point is 00:58:16 It's interesting to see how the feel evolves also over time for this type of like benchmarks. But I think evolves are going to be important. And it's on the actually interestingly, it's on probably on the academics to set the correct. I mean, they have like there have been academics have always been, oh, we have no computer. But like, okay, this is your chance to like steer the field in the right direction, right? I think the challenge is getting attention. So now MMRU was reaching its end of its life. Like what is next, right?
Starting point is 00:58:40 There's MMU or there's MMLU Hard, which someone recently released. It's pro or MMU Pro, I think. It's called MMU Pro. Oh, yeah, that's right. That's right. But like that only lasts you like a year. Right. And then you have to find something else.
Starting point is 00:58:53 So I don't really know what is that. Well, so one thing, you had a comment, I think, in your Rika paper about there's two types of evels. This is a vibe evel paper. One is LMS as judge. And then two is arena style. Right. as sort of the two ways forwards for just general evils that cannot be gained. Although there's also human evals that you, like, instead of L.O.R.M.
Starting point is 00:59:13 as a judge, there's also like human evils that you run. Like, that's kind of similar to arena, but kind of different to some extent or so. Different in a sense that. By the way, do you use your own staff to do that? Or do you, like, hire an outsourcing firm? Now, we don't even, we have a, like, we work with third party data companies to, like, there are a bunch of this, like, around, right? But, like, obviously, we don't like, eval them ourselves.
Starting point is 00:59:32 Like, I don't know how much, how many eval how many, you want to do, right? Sometimes the best researchers do their own evils. Yeah, looking at the outputs and stuff is something that, like, researchers should do. Yeah. Well, there is one element of parametric evils, which I'm hoping that more people come up with, where like you kind of, the benchmark is generated from a seed, that's it. And you can withhold the seed or like, you can vary the seed. I can report how your model did on the benchmark, given it, given a certain set of seeds or whatever, and you can maybe average them. But in that way, it becomes harder, much harder to contaminate. I wonder if that
Starting point is 01:00:12 is an example of this. Not specifically, this is just something I'm wondering for myself, but I did, someone did recently put out GSM 1K, which was, oh, the scale thing. I think, is it scale AI? Yeah, yeah, yeah. Which is similar in that respect, like, make it easy to make variations of a WAMLenem benchmark, but like, that is more likely to be worthheld from training data. Yeah, yeah, yeah, but eventually those would, like, so it's always a sign. Like, even if we put out vibe be very obvious. So I quite like upfront with like if the more people
Starting point is 01:00:40 use it, there's a lifetime, it's like a car, right? After you run a certain mouse, it's time to shelf it, right? So I don't think there's like actually like a good solution. In general, I'm also like a bit, I mean, I think this is like important for the community to think about, right?
Starting point is 01:00:55 But like is it like a fundamental limitation that any benchmark that goes out? Like also there's also one thing is that in the past people used to like which whole test set, right? Like squat or so they used to which whole test set. But then like the like, After a while, I think people also realize that, like, when you withdraw, like, MMMMU,
Starting point is 01:01:08 Kaggle. No, like, when you withdraw, it's like so much extra work for, like, the community to, like, on this, that they just don't do that, right? It's either your data set become, your benchmark becomes unpopular. I think it's also incentive things, right? So if, let's say you are, you want to run, like, a contest, right? And then your goal as an academic is to get as much citations as possible on this benchmark paper, right? Like, then you, or like, this, you want to be as famous as possible.
Starting point is 01:01:34 you would not want to withhold the test set because if you withhold a test set and then people have like there was once like I mean like many years ago there were even some benchmarks where you had to like package your model and send it to them to run like this benchmarking never ever like took off like took off just because like
Starting point is 01:01:49 so at the end of the day right it's like it's the root problem like incentives like also the benchmarking problem is also like an incentive problem right so like it's also people on the show their model is the best and then the game masters want to gain as much cloud as possible and I think also AMCs also get into got into some I don't have a take of this,
Starting point is 01:02:05 but there's people who also feel that they're also optimizing for hype, right? Their own cloud, right? So there's all this. I think it's a lot of interesting. Like, I don't know what I feel this will be, but like sociologic. I don't know.
Starting point is 01:02:16 Like, like, I think there's a lot of papers to be written, right? I mean, about how these incentives, like rewards and incentives, like kind of be, it might not be soft. So I don't know. I'll say sweet bench is probably the one that's kind of broken out this year as like now a thing that everyone wants to compete on is if you're a coding agent. I don't know if you have a view on it.
Starting point is 01:02:34 But it's just like, it should be known to be hard. You should be able to make progress on it quickly. That makes you popular and cited a lot. Yeah, yeah, yeah, yeah. Multimodality versus Omni modality. So this is a little bit of commentary on GPD40 and Camelion. I don't know if you saw the Camelian paper from Meta. Briefly saw it.
Starting point is 01:02:54 Yeah, I'm not, I didn't really take a look at that. Basically, the general idea is that most multimodal models, like Lava or Flamingo, which are late fusion, which is you freeze, freeze, and then you join together versus early fusion where you do it properly, where like everything is, all the modalities are present in the early pre-trained stage. And it seems like things are trending from late fusion to early fusion is the general thesis.
Starting point is 01:03:16 With GPC40 being very obviously early fusion, you guys, I will class it as early fusion. I don't know if you have commentary on whether this is obvious to you or this is the way or they will just be, they will coexist. I think whenever possible, like early fusion is better. I think there will still be a lot of works that do late fusion just because of, like, it's a... GPU poor. No, no, not GPU.
Starting point is 01:03:40 Okay, partially, right, I see this as an artifact of the line between language researchers and vision researchers. And more of like, okay, like people who are training language models, they put out like a llama or whatever and then somebody takes it and then do late fusion on top of it. It's more like a... It's always a... It's coming the orchard. Yeah, yeah, yeah, I think so. I don't know what law was it. law. Okay, I didn't know about that. But it's kind of like an artifact of the organization
Starting point is 01:04:08 doing thing. Right. No, it's just because people don't have money to train things from scratch. I don't know. Even in big companies, right? Like, I mean, I don't know how things have evolved in many companies, but like... You're talking about Flamingo? Like language and vision teams don't use to be the same team, right? So I think this is like a artifact of this. But as early fusion models get more traction, I think the teams will start to get more and more. It's a bit of how all the tasks
Starting point is 01:04:36 unify, like, from 2019 to like, now it's like all the tasks are unifying. Now it's like all the modality is unifying. And then I think like
Starting point is 01:04:44 eventually everything moved to us like early fusion. Yeah. The other element of multimodality is I've been calling this screen modality, screen vision versus general vision. In a sense that
Starting point is 01:04:55 ADAPT is like very, very focused on screens, tables, charts. Most vision models focus on things in the real world and embodied sort of images. Do you have a view on the usefulness for this? I don't think there's like a huge, like, I mean, I think at the end of the day,
Starting point is 01:05:13 like maybe screen intelligence is like more useful in general. But like what if you have like a natural image in the screen? Yeah. No, I mean, no, I think in the end of the day, it should be mixed, right? If a model can do natural images well, it should be able to do screen well and everything. I think at the end of the day, like the models would become like, I don't, I don't see that there will be screen agents and like natural image. Humans like you can read what's on the screen.
Starting point is 01:05:36 You can go out and appreciate the scenery, right? You're not like, say, I only can look at screens. Right. So I mean, I think eventually the models would like be this good on everything. I look at it from a point of like capabilities and screen is. Even screen, there's also like mobile phone screen and there's also laptop screen. Like also different type of interfaces and everything like reading emails, whatever, right? But like reading a page from a website or buying something for Amazon or something, like all kinds of things.
Starting point is 01:06:00 right, and then even in the picture of like a shopping website, there could be like a natural, like for example, like picking Airbnb, right? Then there's a natural image in there and then it's like, you have to understand like how nice is the scenery, right? Or like, like, where is it, right? Yeah. So I think the end of the day is probably like the same.
Starting point is 01:06:15 If you want to build a general model. Yeah, yeah. But I think the natural images is like way easier. Like as in this way like the models currently, current models are actually already very pretty good at this natural, natural images. And I think like screen images are just something that people need to enhance the capability
Starting point is 01:06:31 a little more. That's why there's like some focus on. Got it. I'll touch on three more things and then we'll just go to career stuff. Scaling laws. Palm 2 was Chinchilla, which is one-to-one scaling of model parameters and data. Now you are training a 7B model with 5 trillion tokens. What are you thinking
Starting point is 01:06:48 about the trend in scaling laws for data versus clients? Chinchela scaling laws that's like optimal for like with this amount of compute how much it's the thing, right? But like actually the optimal like there's no, I mean this is something that that even before I left, we really knew that chinchilla scaling laws are not the end of it, right? Obviously, there's also an inference optimal scaling law, which is, obviously, you take a
Starting point is 01:07:08 spot model, and then you just blast it with as much compute and data as you can, until, until you saturate on everything that you care about, right? So I think, like, Lamar trees are what, 15T tokens or something, right? So I think... Which is ridiculous. It's ridiculous to be honest. But at a certain point of time, your value per flop is like not great anymore because you just... Your models are, you just... eventually gets like saturated. But then the problem of like, the question of like, where is this saturation? It's also like, you always find like some metric that you still continue to improve a little
Starting point is 01:07:37 bit. And then you're like, okay, maybe oh, 100K more is worth it to continue training. Like just a little bit more, right? But then it's like, where does it end? Right. But I think at the end of the day, like, the thing about Chinchilla's getting lost is that, like, it was a bit misunderstood as though this model you need this compute. And if you train the Chinchilla's going law, like you kind of, I don't know why so many people
Starting point is 01:07:55 had this idea that you won't improve past the Chinchilla's scaling law. And then people make so much big deal about trading past chinchilla scaling law. Like, oh, Lamar do is the first model. Like T5 base, right, was 1 trillion tokens. That was really so much beyond chinchilla scaling a lot, right? Because that was T5 base, right? I think OPT and GPT maybe set that as an industry standard. It's GPT three specifically.
Starting point is 01:08:17 No, sorry. Wait, GPG3 was not chinchilla. No, I think like OPP and Bloom, right, models like this, they train a large model and with a very small number of tokens and the model turned out to be bad. Yeah, yeah. So I'm talking about Kaplan, the pre-Chinchancella one, the Kaplan scaling laws. Oh, okay, okay.
Starting point is 01:08:33 That one was from opening eye. Anyway, death of chinchilla covered, agreed. But Chechnya is still a cool paper. I think Chechnya is still a cool paper. I love any scaling laws paper, to be honest. It's like such a service to the community in general. Hanging face recently did one data blations, which is like a data scaling loss paper, looking at data constraints, which is just kind of nice.
Starting point is 01:08:54 I see. Long context. People are telling million token contexts, two million token contexts. 2 million token from Gemini, Magic is talking about 100 million token. How important is it, do you think? I think we need to solve benchmarks first before solving the long context, right?
Starting point is 01:09:07 We have your benchmark. No, no, not like the benchmarks for long context. Okay, yeah. Because like you, the needle in his stack is basically like the MNIS, like it's always like a unit test for this style of things, right? But I think there's one part about like hitting the context line and the other part about like actually utilizing. Utilizing.
Starting point is 01:09:23 Right. I think Gemini's long contact is surely like amazing, right? But I think for the community to move forward in this, then it comes to a problem of like, how do we evaluate this? I think I see some long context benchmark on, like coding one and stuff like that. Like I think making those are important and for the community to heal climb. But I think long context is important. It's just that you don't have a very good way to like measure them like properly now. And yeah, I mean, I think long context is definitely the future rather than wreck.
Starting point is 01:09:52 But I mean, they could be used in conjunction. Definitely. Okay. Okay. That's not a take. Which part of the... Long context is the feature rather than RAG. Like you would...
Starting point is 01:10:03 They will coexist but you are very positive on long context. I will put myself on the other mirror image which is like long context is good for prototyping, but any production system would just move to REC. There are a lot of application use cases where you want a model to take that time and then come out with the right answer, right? Sure. Because Rack is like... But you'll use those sparingly because they're expensive calls. Yeah, it depends on like the nature of the application, I think, because you know, because, you know, you...
Starting point is 01:10:25 I think because in reg, right, like you, there's a lot of issues like, okay, how you, like the retrieval itself is the issue or you, you might get fragmented. It's like, what if it's like a very complex story, right? Then you like a storybook or like a complex like thing, right? And then like, like, like, like rec is very like, you kind of chunks, chunks and chunks, right? The chunking is like, and you definitely have lots of information, right? So there, I think there are a lot of application use cases where you just won the model. It's like, okay, like, 100 bucks, like take your time, take one whole day.
Starting point is 01:10:55 come back to me with like that answer, right? Rather than like, I pay like, like, like one cent and then like get back a wrong answer. So I think that's like, like, that it's actually very easy to show that wreck is better than long context because there are a lot of tasks that don't need this long context. You like, like fact retrieval, you just like rack and then you do this thing, right? So like long contacts may get a unfairly bad rap sometimes because like it's very easy to show like,
Starting point is 01:11:18 wreck is like 100 times cheaper and it's very easy to show this. Right. But then it's also. like not so easy to emphasize the times where you actually really need, like the long context will really make like very, very, very, very, very good decisions. So yeah, I mean, I think both have pros and cons depending on the use cases. Using them together is also interesting. Like at the end of the day, it's like a H-Pram that you have to wiggle around.
Starting point is 01:11:43 Yeah. There's another wiggle on the H-Pram. There's another fog on H-per-M, which is how much you fine-tune new knowledge into the model. Are you positive on that? Do you have any views? So, for example, instead of doing Rack, on a corpus and then inserting into context, you would just find two in your model on the corpus. So it learns the new knowledge in whatever capacity, right?
Starting point is 01:12:04 This is cumbersome, I guess. This is cumbersome and you don't want like, you don't want so many of, like, the point of in context learning is so that you don't actually have to do. I think this one is depending on like a business use case, right? If you find it is actually like the, you are very clear like you want this knowledge and then you just find you once and then you don't ever have to pay like context, like in the context window cause again, then maybe that makes sense. But if the domain is changing, then you might not like.
Starting point is 01:12:28 Yeah, obviously it doesn't make sense if the domain keeps changing. But I think for the model to maybe update fundamental assumptions or, you know, reweight associations between words for, let's say, illegal context versus financial or medical context, like it might work. This is the arguments that some people are talking about. So I see this as a trio. Like, it's long context, it's rag and it's fine tuning. Like people always have this like whether either of them will kill Rag, basically.
Starting point is 01:12:52 because Rag is kind of the simplest approach. Yeah, yeah, okay. I mean, I could see, like, if you want a, like, a model for medical domain, legal domain, then fine tuning really works. It's always the move, like, the, you know, domain specialized model, universal model, and, you know, the kind of this tension between both of them. I think it definitely, like, makes sense. It also makes sense, like, to fine-tuning can also be, like, an alternative to, to rack, yeah.
Starting point is 01:13:14 Yeah, well, there's some, there's some companies that are set up entirely just to do that for people. So it's interesting that, I mean, I sort of view Rika as, like, like not working in that space, but you could potentially offer that if you wanted to want it to. Okay, I was going to ask about efficiency and scaling. I'll just mention this briefly. And then we can talk about MOUEs because I discovered that you will rewrote your co-author on the sparse upcycling paper, which is. Oh, no, I was just advising on that.
Starting point is 01:13:39 Oh, okay. Yeah, yeah. But you can talk about sparse upcycling. It's a topic that's hot. But more generally, efficiency in my mind, when I go to ICA, I go to New York, I see efficiency paper, 90% of the chance, I'm just going to ignore it. because I don't know if it's going to work. And I think this is related to some of your scaling work and your inducted.
Starting point is 01:13:57 Oh, okay, scaling law was an induct. Which is like, okay, there was this, to your taxes. I don't know who this person is on Twitter. He keeps talking about me. It's fucking amazing. Oh, yeah, he does have some obsessions, but like he's good. I don't know who he is, but he's good. So he says if 2024 papers are you be trusted, you don't need most attention,
Starting point is 01:14:13 you don't need high precision, you don't need most KV cash, you don't need most fee-for network layers. You don't need a reward model. blah blah. Like, it's like a lot of efficiency papers are just like, hey, on this like small example, we cut this thing out, works fine or works great, works better, whatever. And then it doesn't scale, right? Like, or, so it's a very interesting observation where like most efficiency work is just
Starting point is 01:14:36 busy work or like it's work in a small scale that doesn't, that just ignores the fact that like this thing doesn't scale because you haven't scaled it. It's just fine for grad student. But as for someone who's trying to figure out what to pay attention to, it's very difficult to figure out what is a worthwhile direction in efficiency. Yeah, that's a good point. I think there's a couple. I agree with you fundamentally that, like, it's actually quite easy to tell.
Starting point is 01:14:56 Like, when you see a paper, okay, this one doesn't work, this one works, this one doesn't work. I guess the HIPO account will just tell you that. Sometimes it's not entirely about this thing doesn't work, this thing works everything. Right. Sometimes it's not, you can always find a task in the data set where your efficiency method gets neutral results, right? You can always find one thing that has, okay, I have comparable complexity. And you know what's the most, the cutest thing ever? every time some people propose like this,
Starting point is 01:15:20 they run like some zero short score on like some LMEval harness or something like that. And at 1B scale, all the numbers are random basically. Like all your Buckeal, they're all like random chance performance, right? And they'll be like, okay, I get like 50 versus 54, I'm better.
Starting point is 01:15:36 But like dude, that's all random chance, right? Sometimes I see papers that we run experiments at like, and then it's right. That's a good tell. I think it's very, like the sad truth is that like it's very hard to tell until you scale out. And sometimes the benchmarks that we have don't even probe entirely about what,
Starting point is 01:15:53 I mean, especially all the works about, the transformer alternatives, right? You can always find like this alternative that at 7B scale, at 1, 3B scale, you kind of like, okay, I met transformer on this and this, this, this, right? But then what's the implications when you go to like 200B? What's the implications when you go to 100B?
Starting point is 01:16:09 No one knows that, right? So that's one thing, right? And I think developing your own intuition of like what works and what doesn't work is important. For example, if somebody's like, okay, to be honest, all researchers, like sometimes are also like guilty of this sometimes because you cannot test on like everything.
Starting point is 01:16:25 I cannot test on everything, right? So sometimes you also just want to show your method works on this. But it depends on the objective. If the objective is to write a paper to XML, sure, you can find two datasets. Your stuff works, right? But when you get adopted, I am not sure. Yeah, researcher meta game is one thing,
Starting point is 01:16:42 but as a consumer of research, I'm also trying to figure out how do I know what is what is a useful direction that that's the interesting thing so for example MEOEs seem to have worked out I'll go so far to say
Starting point is 01:16:56 it's the first form of sparsity that worked because there's so much sparsity research like we can chop all these parameters and look we still still perform the same but then it never actually works but MOE is really Oh you mean like the pruning line of work pruning line of work sorry
Starting point is 01:17:11 I should have used that word So I don't know if you have any commentary on like Deep Seek, Snowflake, Kwen, all these proliferation of MEOs, MEOE models that seems to all be sparse upcycle because you were advisor on the sparse upcycling paper. So the sparse upcycling paper was
Starting point is 01:17:26 mostly vision focused with a little bit of T5 experiment. So it was early stage of like sparse upfacking, but it was good that Google was really thinking about this long ago. And Nome also had paper on it, right? Yeah. I think only is the way to go. Is it like 100 experts, or 1,000 experts? For some reason
Starting point is 01:17:42 the community settled on 8? You probably get more gains from more than 8, I think. But I think in general, it's like MOUs are just a trade-off with like prime and and flop, right? And then you're able to make like you kind of make that that scaling law increase from that additional. So you can keep a low flop but kind of have more parameters. It's just changing the flop ratio.
Starting point is 01:18:07 Keeping in mind, there's a lot of inefficiency between the experts. Yeah, yeah. I think as an architecture itself, the flop program ratio makes it, like, worth it, right? But I think the thing that's not very well understood is that, like, how does M-O-E, like, for me, as a research question, is that, like, when you, like, how does it, like, relate to capabilities and stuff like that? Like, does this inductive bias actually, for example, when you do, like, massive instruction tune, I think there was this paper, like, flood M-O-E or something. Like, they show that, like, instruction tuning. I'm not, like, fully sure, I don't recall fully, but, like, when you do massive instruction tuning, like, MOE models are. like they behave differently from from dance models and stuff like that.
Starting point is 01:18:44 Like I think, okay, like fundamentally I just think the MOEs are just like the way to go in terms of like flop parameters. They show they bring the benefit from the scaling curve. If you do it right, if you bring the benefit from the scaling curve, right. And then that's the performance per flop argument, activated programs, whatever. That's like, that's a way to slightly cheat the scaling law a little bit, right? By having more parameters, right? I think the more interesting thing is about like what tradeoffs do you make in terms of
Starting point is 01:19:10 capabilities because of this new architecture. I think that's actually like the question that I think I guess all the frontier labs are. They already know this. Nobody's writing papers anymore about this. So like you just have to live with what's outside. But I think I'm bullish about MOUEs. Yeah.
Starting point is 01:19:25 I had to, I made an exercise for myself on rating research directions and what their asymptotic value is. And I put MOUEs pretty low because I think you have a good base model and then you upcycle it. and it bumps you a little bit. And I think that's it.
Starting point is 01:19:43 But, like, I'm always seeking to invalidate my hypothesis, right? But, but, like, from scratch, MOUE is also promising, right? From scratch, MOUEs promising. I think in the I do MOU case, you do MOU from scratch. Yeah. Okay. The last part that makes me uncomfortable about MOUE debate is actually related to another paper that you wrote about the efficiency misnomer.
Starting point is 01:20:00 In a sense that now people are trying to make the debate all about the active parameters rather than total parameters. But it seems like it sounds like that's something that you're comfortable with, like flops at inference is a relevant metric and it's not that. Well, thanks for like actually reading or like reading the papers. I'm trying man. Thanks for. It's very hard to cut.
Starting point is 01:20:16 It's very hard. You have a lot of papers. Well, I actually very impressed that like, oh, you're bringing up these papers. Yeah, I'm using attention. Okay. Okay. Yeah, thanks. Thanks.
Starting point is 01:20:26 And also, I mean, I'm interested in efficiency that works. It's just very hard to find efficiency that works. And so like anything that helps me have high signal on efficiency is helpful. So I think for the inefficiency, misnomer by the way, I love the paper, by the way, it's had a fun time working on it. I think efficiency misnormal was like, we found that a lot of people,
Starting point is 01:20:44 like, they use params, like especially to do kind of like, right, and then MOEs was not very hot like in the community at that time, right? But MOEs were like a thing long ago at Google, right? So I think using active params, I'm comfortable with using active brands
Starting point is 01:20:56 to kind of approximate like cost on the model. But like in the efficiency mismanormal paper, we actually made it quite clear that you should always like look holistically about like, like, because you have serving like additional serving costs. Yeah. Like fitting in the GPUs, like fitting on single node and something like that. Interesting one was speed.
Starting point is 01:21:11 Nobody really talks about speed. But your paper actually tried on. Okay. I have something to say about speed. There are so many methods, right, that are proposed about efficiency, right? They are like theoretically like faster because of like complexity, like something like that. But because there's no way to work around the implementation or like your implementation becomes so hard. It becomes like 10x slower.
Starting point is 01:21:34 Okay. There's so many papers. It's not a hard way or where. Like, it could be hard. It might not be, it could be hard where it could be just the way that, like, you have a convenient way to, like, in its mathematical form, it's actually like, okay, linear complexity, like, whatever. And it's actually theoretically faster. But, like, just because you have to, like, do a scan or something like that. And then it becomes, like, actually, like, 10 times slower in practice, right?
Starting point is 01:21:56 There are a lot of things, like, not a lot, but, like, there are some things that are, like, some methods that are like, like, like this, where you don't take it into account throughput, right? which is also the problem of like sometimes like the incentives of like people who are working efficiency, you can easily sell a paper as like more efficient and then people will not suspect that because the reason why we wrote the paper is that so many people were confused
Starting point is 01:22:17 about like efficiency itself, right? Yes. And then they will be like, okay, like a lot of these unsuspecting reviewers, especially like even academics or they don't have like that that real real feeling they were less like, okay, less parameters, more efficient, right? So you could have a method that's like less parameters
Starting point is 01:22:32 but like three times slower. because a lot of times when you add things to the model, it becomes slow. You add complexity, especially if it's like something that's not hard where optimized, no kerners or like something that is like bad for TPUs or whatever.
Starting point is 01:22:44 Your model just become like slow. That's a temporary issue. People can fix it. But some things are not like so. Some things may not be like so easily fixed or like it just adds a lot of like three costs to, to optimize it and everything. Right.
Starting point is 01:22:58 But then it's always marketed as like because I save brands. So I save. Right. And then the brands will add a different place of the model. like, for example, like, if let's say you, even in the case where you brand match models, right, if I take out like some brands from like FFN, right, and I put it to like embedding layer, right? It's a cheap operation for abetting layer, right?
Starting point is 01:23:21 But my model becomes like lopsided, right? I could say I brand match this, but it's not true put match, right? Yeah. Because it's unbalanced on the side. It's unbalanced on the side, right? So there's also of this type of tricky things that like when mixed commons, all the comparisons, like very, very, very, very difficult. And because you cannot even put like flop, throughput and speed,
Starting point is 01:23:40 flop, parliaments and speed, like extra speed, right, in the same plot, right? And then there's always like one money shot in a, like, there's always a like a parietal kind of compute, like whatever plot, right? Like for marketing in papers or something like that, it's always very easy to like, I mean, not intentionally, but like to subconsciously, like, show one story when it's actually like there's like all these other things, to consider. Yeah.
Starting point is 01:24:04 It's a selection bias, self-biased, whatever. Very cool. Okay. Well, that was mostly of most of the technical side. We have one commentary that will happen today on the future of open source models. It basically founders fund said like the future is closed source. You were agreeing with it. And a lot of the open source fanatics are up in arms over this.
Starting point is 01:24:24 I don't know if you care to comment about just open versus close and close whatever. I mean, I don't really like, when I mean like, if you're referring to the tweet that I wrote, but I wrote something about... But this is huge. Like, so many people are commenting about it because they are personally physically offended that open source cannot catch up. Okay, wait, okay.
Starting point is 01:24:42 So I want to say, it's like, I'm not... Like, I contributed to open source in the past, so I'm not, like, against, like, open source per se. But the interesting thing that I want to talk about here is that, like, there's a difference between... Like, I draw a line with, like, open source as in, like, okay, to me, Lama tree is like...
Starting point is 01:24:57 It's like, matter has an org that is like, okay, hypothetically very similar to to like gem night or something but they just decide to release the weights, right? Yeah, it's open weights. Right, it's open, it's open weights everything, right? I think when most people try to say that, like, open source is catching up everything.
Starting point is 01:25:13 They kind of mean like this grassroots, like, yeah, this, this bottom up people that are like, these indie developers that are like coming together to fight, like, it's romanticized and it's dramatized to some extent to fight against like this. It definitely is. Right. And to be very fair, I think that there isn't really much like,
Starting point is 01:25:31 like so far, if you just look at, the fractions of people, the big labs are just pushing and pushing and pushing. The academics, like Stanford and stuff, they came out with DPO, they came out with things like that. They make some, like, but they're kind of in between the line of like open source community. And then there's also like the developers that are like fine tuning on GPT4 distilled models and everything, right? I don't, I think the open source, the underlying like thing about like collectively improving
Starting point is 01:25:57 something. I'm not like criticizing it for the sake of criticizing it, but like, I'm just saying, saying that like in order to make progress, right? I think the incentives of open source are like, what I observe is that like people like to do things like, they like to take somebody else model, they rename it, they make a quick win from there. And then like you notice that like when people realize that like
Starting point is 01:26:18 this starting on the GPD4 tab and running some DPO, it's not going to give them the reward signal that they want anymore. Right. Then all these variants gone, right? You know, there was this era where there's, wow, there's so many of this like, I lost track of all these model variants, but now they're all gone because people realize that
Starting point is 01:26:35 that you cannot climb LMSIS because you need something more than that's something that is lightweight, right? So I think that was just my overall, like... Honestly, the Hanging Face leaderboard contributed to most of that. It's not LMS. No, no, I think LLSI is probably they realized that they could not, yeah, right? The open LM leaderboard is probably like a big problem, to be honest.
Starting point is 01:26:55 We're talking to Clementine in one of our future episodes. Okay, okay, okay. They dedicate a lot of, I mean, there's so much attention to them. It's a tough problem, but they're providing a public service for sure. Yeah, yeah, yeah. I mean, good intentions are always good. I mean, like, good intentions are always good. Yeah.
Starting point is 01:27:28 Is any, any, take that in any order that you wish. I don't have much of a life, actually. But I'm trying more to have more. I mean, you're a father now. I have a baby now. So like, I'm trying more to have more life and everything like this. I think the productivity hack that I have is just like, I didn't have like a boundary between my life and my work like for a long time.
Starting point is 01:27:47 So I think I just cat a lot about working most of the time. Actually for the last like, through my PhD, at Google and everything, I'll be just like working all the time. It's not like the most healthy thing. ever. But I think that was actually one of the biggest productivity. And I like to spend a lot of time writing code and I just enjoy running experiments, writing code and stuff like that. Right. So you kind of, if you enjoy something, it's not work
Starting point is 01:28:11 right? So like, it's very strange. It's like, I would get distracted by sometimes I have to watch some Netflix series because like my wife asked me to like watch it. Or somebody tells me that I'm back on time on some shows, right? But then I get distracted by my experiments running and I just end up like like writing code instead of like so things like like like this it's not the most healthy thing but I think that's one
Starting point is 01:28:32 looking for like a practice where like okay so Andre recently had a thing where like before when he wakes up he doesn't look at social media he only goes trick to work damn I checked Twitter the moment every I know it's just something I do as well but I'm like that's a smart rule and like I'm looking for like rules like that like do you have a rule
Starting point is 01:28:48 no he doesn't check social media because his phone is exploding all the time all the time yeah I don't have so many likes and followers so like it's fine for me yeah you get there Like rules like that, mantras that you've developed for yourself where you're like, okay, I must do this. So for example, recently for me, I've been trying to run my life on calendar for a long time. And I found that the only way that I work is I write things down on pen and paper and I cross them off individually. And like that physical action really, really helps me get things sorted.
Starting point is 01:29:14 And that's work-wise. Reading-wise, I don't know if you know, but I've been running this like AI newsletter. Like all those summarizes, all Twitter, Reddit this points and all that. So that helps me keep up because I have like a socially graded. and I personally vetted the entire pipeline from beginning to end. So, like, this is my input algorithm. I know how to keep up with news because I now have a information condenser.
Starting point is 01:29:37 So, like, I'm trying to figure out what's your algorithm or what's your rules for keeping up. I got something, I got something for keeping up. So I used to check archive, like, every morning when the gate opens, I just check archive. I will wake up 9.30 a.m. Singapore time the archive gate opens, right? And then I'll be very sad if there's no papers to read. But you usually just pick one paper or two papers that you find interesting.
Starting point is 01:29:58 I don't read them. I just like skim like the thing, right? Yeah. So I used to that. I don't do that anymore. I mean, ever since like I have in the startup, I read, right? Yeah, you have a real job now.
Starting point is 01:30:06 I read, I read. I read, I read less papers, right? But I used to camp at the door of archive quite frequently just to see. Isn't that? That's not a good use of time. I'll come on and say it. It's not a good use of time. No, no.
Starting point is 01:30:16 It's a newness bias. Sorry, go ahead. No, no. It's just because like, I ran out of things. I say, yeah. It's just that like the new stuff comes out, right? Yeah. Like, and then like the new stuff come out, so that's how I keep up to date.
Starting point is 01:30:27 So in the space of three years, you read every... No, no, I didn't read everything. It's just that. It's just that. But these days, I realize I don't have to do that anymore just because if the paper is important enough, Twitter will show you to me. Sure. So I, there isn't really, like...
Starting point is 01:30:40 And one thing I do is that I actually don't read papers like that much anymore. I just like skimed them like almost, right? So that's for keeping up, like, with papers, research and everything. And the other thing more of like this, like just like, a productivity point of view is that I used to always keep like the text. Like I usually start writing the thing while working on that thing itself. Like so even like let's say if you want to launch something like the angle is like a blog post or shipping something and everything, right?
Starting point is 01:31:08 I like not really a launch or say all that's papers. I always like to look at it from like what's the story and the end. And then I just like figure out what I need to do to get to kind of right. So I think as a researcher like this is something like I would have like. like so many drafts of like when I start the project I don't know the experiment yet everything right
Starting point is 01:31:26 but I like to imagine like what the title would be right and then I always vibe check but like I always like so I mean my friends at Google would know that I always have like like the overlay draft of like so many
Starting point is 01:31:36 and then I'll just spend time looking at it like to looking at it like took the title is it better two seconds like I can't about I used to care about a lot of things but this actually helped like product because every time I look at it I'm like okay this is the final product
Starting point is 01:31:46 I'm like working towards it right because I think a lot of researchers they tend to like they sew around in the experiments and they never like ship the final story. It's like the shipping like I mean it started out with ship products but like as a researcher your product management. Yeah.
Starting point is 01:32:00 You're shipping the thing. So I like to I like to hang around a lot in my in my draft and I get motivated from that and that's like one productivity thing that I did as a researcher. Yeah. So I think that other than that I don't really have any things that I do that probably different from others. Probably you don't know it. This is unconscious competence versus.
Starting point is 01:32:19 Okay. Was it like just NTIU PhD? Just the story of like How was it coming out from NTU Which is like a good school But like not not typical target school For like a big lab I did my PhD unknowingly
Starting point is 01:32:32 Like I didn't have very Like when I was a very regular undergrad I had decent grades But not the best grades I was not like super smart in school Or something like that I was I wanted to do a PhD That's because I was like curious
Starting point is 01:32:45 And I mean like and then I wanted to stay in Singapore at a time So I just like naturally just did a PhD there. I didn't even vet my advisor. I didn't even think too much. I just like fell into the PhD program. And then there was when I realized that, oh, actually I can do research.
Starting point is 01:32:59 Like I'm like pretty decent at research. Like I just fell into a PhD like unknowingly. Yeah. And I definitely like, NTU leaves a lot to be desired. Actually, to be honest, I think that I mean, Singapore leaves a lot to be desired in general.
Starting point is 01:33:11 Like the research community here is like like probably not great. So how did you like break out? If I was you, I would have no idea. how to break onto the international scene. I think it was, okay, to be honest, like in retrospect, it's a bit of, like, a bit of a miracle. Or, like, I mean, it's not easy to,
Starting point is 01:33:28 I think I could not, if I had, like, a product, like, someone to mentor, like, I could not, like, tell somebody how to replicate the same thing that I did. It's much easier now, maybe compared to in the past, but I've been mostly self-supervised during my PhD. Like, my advisor was basically, like, like, Gramerly. Like a free paid plan of Gramerly.
Starting point is 01:33:47 You won't watch this, so it's fine. But, like, there's a lot. things that it was like this strange art of my life where I was figuring out research by myself and everything and okay maybe going back to the change your opinion is that like the biggest culture shock I had like when like was moving from Singapore PhD to Google I think my research like taste you went straight to Mountaineville yeah I went to Mountedville I started at Mountain View like my research tastes and everything like like I was it was so different like the research culture is so different in in US and in Asia I had to grow so much like during my time at Google too
Starting point is 01:34:19 actually evolve. And then whenever I come back, right, I still have friends in, like, faculty in here and everything. I either think that I'm a snob or they think that I'm like being a like a very nasty person. Because I think to be honest, the research here is like in Singapore. It's just basically like they just care about publishing papers and stuff like that. And then it's not impact driven. I think at US is mostly focused on impact driven.
Starting point is 01:34:43 And the thing needs to make real impact, right? And what, to be fair, you're also working at an industrial lab versus an academic circle, right? Like, you're comparing apples and oranges here a little bit. I mean, at the end of the thing, I think research is like fundamentally, like, as an industry, RIS, you still write papers. Your goal is to advance science and everything. To be honest, it's all the, the incentives, rewards the system is like different
Starting point is 01:35:07 and maybe like slightly different and everything. But, like, at the end of the day, I still feel that researchers are researchers, scientists or scientists, no matter, like, really, like, where you are. I will get so much dissonance when I come back and I talk to people. I would feel like, oh, why do you think like this? But then I used to think like this. So like the environment shapes the way a researcher thinks. The taste is very important.
Starting point is 01:35:29 Sometimes I try to communicate this to people. And then maybe I come across as a snob to, to like the local community here. Right. But like it's just that there's like maybe there's so much dense information that I want to bring back. But like, there's no like fast way to like transfer like all the like transfer all the things that I've learned. also a big cultural shock because I was in brain in the Singapore office for a while and I'm reporting to... You're the only brain in person.
Starting point is 01:35:54 Yeah, yeah, in Britain in Singapore. And then I had like, I took on an intern from NU.S. actually. And the research like vibes and the thing was so much of a conflict for me that it was almost like my body was rejecting it, you know? But this person so like grew and became, I'm happy with how this person grew from my mentorship. So he's now in a way better situation. But I
Starting point is 01:36:18 say that like a lot of people in the in universities here like not like a bit like like ignorance is bliss right maybe sometimes well no it's exposure i didn't know any better myself until i went to to the u.s for for college and then yeah my world was expanded and it's a little bit of a pandora's box because once you've tasted that you're never happy yeah yeah you know so okay last question would be just just a sort of Singapore question so i i like to be visible visibly non-American covering the AI scene because it's very US-centric. Every non-American I talk to always wants to be like, how can we build Silicon Valley in my city? You know, my country, my city, whatever. That is not Silicon Valley. I feel like you have basically just kind of like me. You kind of operate
Starting point is 01:37:04 in the US circles, but you just don't live there. Do you have any advice for like if Singapore, okay, so I'm wearing a register today. This is the official Singapore government's sort of community group that is that is trying to guide Singapore AI policy. If we want a hundred more you days to come out. What should governments be doing? What should communities, ecosystems should be doing? So I actually think that like sometimes like not doing too as like too much is maybe less is small, maybe. I don't think there's actually much like the government can do to like influence like this kind of thing is like a natural, like an organic, like an organic natural thing. Right. The worst thing to do is probably like to create like create a lot of artificial things that like exchange
Starting point is 01:37:44 programs. Okay. I mean Singapore used to have a lot of exchange programs. Like it's send people to, I mean, just talking about AI specifically, right? I think that, for example, like, sometimes, like, trying to do too much or, like, moving in the right, wrong direction is just better than not moving at all. Especially if you, if you accelerate in the wrong direction, you actually get into a worse than possible, right? So I think it's very dangerous to, like, move in a bad, like, direction. I think respect your talent more, maybe.
Starting point is 01:38:08 The government should just respect the talent more. And I don't know whether this is too much of a... No, no, no, no. But not, but not, maybe not moving in a wrong direction is, To me, it's already a very good thing. Funding for startups, incubation, holding academic conferences, I clear next year it's going to be in Singapore,
Starting point is 01:38:27 so people come here and expose to it. But, like, I don't know. This is just very interesting. Like, everyone wants to build up AI expertise within their own country, and, like, there's a massive range into the US. I'm part of that. Like, I live there.
Starting point is 01:38:40 I feel guilty. I don't see any other way around it. It's such a huge problem. I also do think that there is, like, cultural hegemony. it's called it, like US values basically being asserted on the whole world, right? Because we decide RLHF
Starting point is 01:38:53 on these models and now you shall use all our models. And it's just troubling for, like, national sovereignty should be AI sovereignty, and I don't know how to achieve it for people. It's very scary. Okay, there's a lot to unpack. Yeah, this is not technical, but I was just curious.
Starting point is 01:39:09 We can make this the ending conversation, which is, I think your inspiration to a lot of other people who want to follow your career path. I'm really glad that we got the chance to go through your career a bit. Yeah, I'm sure this is just the start. So hopefully there's more to come. And I want to inspire more of you.
Starting point is 01:39:23 Yeah, yeah, sounds good. So I'm just glad that you shared it with us today. As a special coda to this conversation, we were invited to join the Tekkenasia meetup featuring Yi by managing editor Terence Lee. Terence asked a similar question on how other countries can create conditions for top AI labs to spring up outside of Silicon Valley. So, like, where do you see Singapore playing a role in? in AI. So how would you?
Starting point is 01:39:50 Oh, okay, right. I got a practical one. Okay. I got a practical one that is actually actionable. I feel like one thing that people don't get, the advice that, practical advice, like, that is that like, the era of like people who talk versus people who do, like, the people who talk is like gone, right? So like it's no longer about like, I have a team, I have like 10 interns from Southeast Asia or like the region and then they're going to do this, do this, do this, do this for me, right? So I think one thing that, that, that, that, so, I think one thing that, that, senior people in any government may not get, right, is that the world has shift into this
Starting point is 01:40:25 paradigm where senior ISIS, ISIS as individual contributors, right, are actually making the most impact in AI, right? So in GDM and in Open AI, I mean, Frontier Labs, they're all very driven by individual contributors and not, actually this is not even related to, this is, I'm talking about, like, this is the advice I give, but it's actually general, like, so many general thing. So multi-purpose basically. It's not AI-specific. No, it's also, it's very AI-specific because the level, the difficulty of making impact and making breakthrough has started to become like, it's no longer about like, it's not
Starting point is 01:41:00 like software engineering where, where it's, I think AI is a little bit harder, like and then like it's mostly about like getting very senior people who are hands-on and have a lot of experience rather than like management style people that like try to like think they know what they're doing but they actually don't so I think I mean I I'm not going to like say like name obviously right but like I mean I meet a lot of like people like this or like in general I mean not only in Singapore but like right but AI has shifted quite a lot into this I see driven paradigm where the people making impact are the people who are like
Starting point is 01:41:43 on the ground fighting the war, right? So it's no longer about, I have 10 interns, 20 interns, 100 interns, you do this, you do this, you do this, I just take meetings, right? No, right? The senior person writes code, everybody writes code, nobody should not write code, right? And then everybody, so I think this is, okay, this is a big extreme, but, but, but, this is a bit on the extreme side, but I think from people, like, I just, the advice is just like, maybe like, just take 20% of what I say, and incorporate
Starting point is 01:42:13 So instead of like if you if you if you for example hypothetical hypothetical situation right say you want you want to organize like an AI conference in Singapore right and then you want to make it like you want to show Singapore as like the AI hub in the world right I mean you don't invite like policy people and like you know if invite like policy people to come and talk about AI safety AI safety right you invite people who like actually know the stuff right and then if you organize a conference and then like 100 people like go there and then they feel very productive and everything but like the problem is that like Singapore doesn't have people who really can do it you know right so I mean I've through the great vines I mean I hear about people like fighting for territory here and there I mean this is what I hear I I don't want to hear this but I hear this somehow right and then like and then sometimes I just ask them like who's actually going to do it right who's going to do it right the work is the model is not going to train itself right unless we have AGI right
Starting point is 01:43:15 So yeah, I mean, understand that like times have changed, it's no longer about like, it's no longer about like, oh, I'm very senior, very senior, very senior, okay, okay, okay, can you quote, right, that's the question, right? I think that's like the, yeah, yeah. Well said, spicy, spicy already. Okay, okay. Okay, I think we've hit the number. You're like cocoa in baya, raise the cocoa egypt area to the maximum. Yeah, almost there really. Okay, questions, anyone?
Starting point is 01:43:41 Indeed, questions are very welcome. Head over to the latent space substack to leave a question or tweet at Yitai ML or at AGIHippo directly with your feedback.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.