Latent Space: The AI Engineer Podcast - The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Episode Date: July 5, 2024Livestreams for the AI Engineer World’s Fair (Multimodality ft. the new GPT-4o demo, GPUs and Inference (ft. Cognition/Devin), CodeGen, Open Models tracks) are now live! Subscribe to @aidotEngineer ...to get notifications of the other workshops and tracks!It’s easy to get de-sensitized to new models topping leaderboards every other week — however, the top of the LMsys leaderboard has typically been the exclusive domain of very large, very very well funded model labs like OpenAI, Anthropic, Google, and Meta. OpenAI had about 600 people at the time of GPT-4, and Google Gemini had 950 co-authors. This is why Reka Core made waves in May - not only debuting at #7 on the leaderboard, but doing so with all-new GPU infrastructure and 20 employees with
Transcript
Discussion (0)
Welcome back, friends. It's only been a week since the World's Fair, and it was incredible gathering the community to see the latest and greatest in AI engineering.
You can catch up now on the four live-streamed track days on the AI engineer YouTube, and our team is busy editing the remaining workshops and five other tracks, including the surprisingly popular AI leadership track.
Thank you all for your support, and stay tuned for news about the next event, the 2024 AI Engineering.
Summit. Last week, we did a very special deep dive with Josh and John of Inbue and
Databricks Mosaic on training LLMs and setting up massive GPU clusters. And today, we're pleased
to follow that up with a very special conversation with Yi-Tay, formerly tech lead of Palm 2
at Google Brain, and now chief scientist of Raker AI. Raker's largest model, Raker Core, was at launch,
the fifth best model in the world,
and the only GPT 4-class model not trained by a big lab like OpenAI, Google, Anthropic or Meta.
In fact, while Google Gemini has 950 co-authors,
Raker only has 20 employees,
only five people actually working on pre-training.
One year after our RWKV episode,
Swix was excited to return to Singapore to delve into Ye Raker
and building a new AI model lab outside of Silicon Valley.
Stay tuned to the very end for a special bonus clip from Yi's recent appearance at the Tekkenasia meetup
for his spiciest take on why senior management is overrated,
and why this is the time to build up senior 10,000 X individual contributors.
Watch out and take care.
Welcome, Yi Tay to Layton Space.
This is a long time coming, but I'm so excited to have you here.
Yeah, thanks for inviting and excited to be here.
I'm glad about a lot of stuff, yeah.
So you are interesting to research and introduce.
You are now chief scientists of RECA,
which is a super interesting model lab,
but before that you were at Google Brain.
You were architecture co-lead on Palm 2.
You were inventor of you all two.
You were a co-contributed on Flann.
You were a member of the Bard Corps team,
and you also did some work on generative retrieval.
That's a very, very illustrious three-year career at Google Brain.
Yeah, thanks, thanks, yeah.
And then since then, Rika, you joined in March,
2023 announced a $58 million series A in June
2023. I don't know if you know the post-money valuation
or the pre-money valuation is public. So crunch basis is
200. I did not know that. 50-something million. So you don't even have to
leak. It's on the internet. Okay. Raker's stated goals were to work on
universal intelligence, including general purpose, multimodal and
multilingual agents, self-improving AI and model efficiency. In February,
you release Raker Flash. In April, you released Raker Core and Edge.
And then most recently, you released Vibe Bell.
Is that a good summary of the last six years?
No, it's not four years?
Four years, yeah.
Oh my God.
Okay.
We're talking about AI.
Yeah, I was like,
I'm wondering like since when did I like
step into a time machine or something.
Yeah.
Okay, so can we just talk about your transition into,
you did your PhD and we can talk about your PhD,
transition into brain and research and all that.
You know, I saw you do some work on recommend your systems.
I saw you do some work on quaternions.
What the fuck was that?
Okay, let's forget about that.
Describe your path into modern elements.
LMS, right? Because you didn't start there.
Yeah, okay, sure. I think the world also didn't start there, right?
I mean, I think in, so I joined Google in 2019, end of 2019, and the world looked like really
different at that time, right? I think there was around the time the first GPD was released
by GPD1 or something was released by OpenE Eye. So research, like ML research and NLP
research looked very different at that time. So I was mostly, I identified as like a language
researcher. I don't like to use the word NLP. Jason will kill me if I used the word NLP,
but I was like, okay, a language research, right?
But I was more like an architecture kind of researcher.
And when I joined Google, I was also, I continued on this, like, as a model architecture
research.
I worked a lot on, like, efficient transformers.
There was your first viral paper.
Yeah, yeah.
And I worked like long-range arena.
I spent quite a lot of time looking of like could we do without attention.
Like, there was a synthesizer paper back in 2020.
I think that was like my early days in Google.
There wasn't like a, at that point of time, transformer research was mainly like, like,
WMT, like machine translation and like
Popaxity and stuff like that,
it's not really about, you know, that there wasn't like,
I think only feel short,
feel short in context learning, came only about
when GPD3 came out and beyond, right?
And then, so I think that at that time,
the meta, I would say, the meta looked very different.
And at the time, a lot of the work
will focus on like fine-tuning things like T5 or bird
or something like that, right?
So I think a lot of the research, not only myself,
but like around me or like even the broader community
were working on those kind of things.
And so I think that was, which I feel that in hindsight today,
it's actually pretty useful to like kind of think about
because a lot of people came into like AI and into right after chat GPD came out.
So they saw AI as kind of, I think there's a lot of benefits of, you know,
understanding how transformers and I've broken this thing apart so many times,
it's like these things actually help to improve intuition and it's not totally disconnect.
I think a lot of things are still relevant today.
and it's just the scale has gotten much larger
and also the paradigms shift a little bit
from single task fine tuning
to like generally do everything
kind of universal foundation models.
Foundation models, right?
I think it's just a slight change in paradigm.
But fundamentally, I don't think
the underlying principles of research
hasn't really changed that much
except for like compute.
So basically algorithms stay put
and then compute and data scaled.
So I have some thoughts about this, right?
So I think back then a lot of the
academic research. I think people have talked about this, like,
like, Sasha Raj has talked about this, or like, other people have talked about this.
It's like, the conferences were always organized by like applications, right?
They were always organized by like, oh, like question answering. It was always by this, right?
I think there's like a bit of a transpose going on. Things become universal and then becoming like,
okay, there's a data work stream. There's a model architecture work stream. And then people
work on improving like a universal model and general purpose algorithms to improve this model rather
than finding domain specific tricks.
I think for, even when 2019,
I think I've already been like focusing on works that are like,
you know, you could improve one general architecture.
At that time, it was like, like, maybe LSDMs in 2017 or something.
And then you try on like 10 different tasks and the kind of thing, right?
But like a lot of the research community have been focused more on like,
how do I get that extra 2% on question answering or like and then sediment analysis?
I think that there was this phrase of like in 2017, 2018,
where this type of work was still very fashionable in academia and conferences, right?
And then I think the big thing about the chat GPD moment of like 2022,
the thing that changed drastically is like it completely like, it was like a dis-sharp,
made all this work like obsolete.
So November 2022, you're saying, exactly chat-GPC launch,
because I feel like if you're in the research community, this was coming.
Yeah, yeah.
So I'm saying that the big labs and stuff, people have already been moving towards general.
Even T5 was already general purpose.
Yeah.
And the thing, right, but there was a bit of a time, basically,
just like Google and meta, OpenAI,
we will be working on things three years ahead of everybody else.
Then academia would be still working on these past specific things.
Got it, got it, got it.
And then I think the forcing function was the chat gibbet moment actually really like,
it was coming, it was coming.
It was just the final, the last straw.
And then it's finally like.
Yeah, now it's serious.
Yeah.
Now it's really the thing completely changed.
I don't know how it turned from my background to like talking about the meta.
I think that you navigate the matter very well.
And part of my goal here is that also.
also isolate how you think about the matter for other people to reflect on, because I think obviously you do it very well.
Oh, thanks.
So I'm looking at your papers published. Somewhere around 2021, you had a hard cut to 22, YOL2 and Palm. You did YAL2 Palm, emergent abilities, DSI, recitation,
augmented generation, all in the same year-ish. So like, there was, did you change teams? Did you, did you, like, have a research focus? Like, when did you become...
Oh, so you're saying that, like, my research model guy.
My research became emergent, right?
It was very obvious.
No, I don't think I'm like a person that like,
I'm not like super, super great at like forcing a trend like two years like a head
and then especially like especially like plan for that.
Yeah.
I think I smoothly and as like kind of like as the few moves.
I never actually really thought about this.
This way I just at every step I just optimise for what I found to be most impactful and
most promising.
And then that gradually.
And also it's also a lot of influence by talking to.
people, right? At the time, I started working more with, I had some close collaborations with
Jason and other people. I mean, Google is, you can work with anybody you want, basically. So,
you're kind of partly the environment shift. And I think the environment shifts very quickly,
but I was always polling in the environment. I was not, I think it's always good to have an
open mind and then move along with the field rather than, okay, this is my research area. I'm going
to get stuck really two years. I think I just move along to find things that interest me. And
And naturally, I think that turned out to be the things that were most impactful at that time.
In retrospect, I kind of did well, but I never actually really saw it as intentional.
Sure.
I didn't do anything really intentional, except that's doing what I find interesting, actually.
Cool.
Well, we'll just talk about the main work at Google Brain, and then we'll move to RECA.
So out of UL2, Palm, Emergent Abilities, which of these came first?
Actually, I can't really actually remember.
Okay.
We'll make you talk about UL2 then.
UL2 and DSI, the differential bus search index.
I was working on it the December of 2021.
So at Google, there are projects that are big efforts that a researcher will be part of the effort.
And then this will be kind of top down to some extent.
And then they will also bottom up research that one could do.
I can't speak for the Google now for sure, but at least at that time, right?
So UL2 and DSI differential search index works that I kind of tinkered with in the December break where nobody was around.
Okay.
So there's differentiation because there's Palm 1 and there's Palm 2.
right. So Palm 2 I was actually
a co-lead of one of the work streams but Palm 1
I was more of a contributor and Palm 2
I was so now I'm to think back of
okay what's the timeline which came first right?
In general there were three categories of works.
One is broader efforts that
org level efforts and then there are some that
do and there's I was my own
projects. Projects. I used the
compute that I had and then
I just played with it. You accidentally
left the or two running for a month. Yeah, yeah.
Yeah, that was in the paper. It was fun. It was
really fun. And then there was also
the third category where those were the efforts that
my good friends were driving and I contributed.
So Flann was just one of the, I would like to maybe say this publicly.
You're very publicly.
I talked a lot about Flan.
Your Flan show number one.
But like, yeah, but like the first author is actually Hyeong-Wan, who is great.
And then like another guy, I was a core contributor.
But I mean, just because I'm a little more visible, so I kind of accidentally took a little bit more credit for that.
But as in, I was a core contributor, but I was not like.
The lead authors are obvious.
Yeah.
So the third categories were like projects that my friends, like emergence was also like,
Imurgent abilities.
No, actually, that paper was actually supposed to be only me and Jason on the paper.
And I actually became friends with Jason from that paper.
And then that led to this streak of, like, I don't know,
10 papers or something together with Jason.
And now we're like super good friends.
The ultimate romance.
But that was like the emergent paper.
But emergent people was also like a bottom up kind of like a thing.
And fun times.
Yeah, it was fun.
Maybe I'll pick on Palm 2 because I feel like,
I'll pick on Palm 2 in emergence.
is I really want to make sure I tell those stories.
Those are important stories.
Palm 2, I think it's a career story.
Effectively became a co-lead on the second version of a very high-profile company-wide effort.
How did that happen?
I think people would like to know what's like the career strategy there.
To be clear, I was one of the co-leads, but there were a lot of co-leads.
So I don't want to take too much credit for that.
But my involvement with Palm 2 came from after UL2 was working well,
and then it was gaining some visibility within Google.
Was YL2 the largest model that Google had released at the time?
Yeah, I think so.
That was the largest.
And you just, it was a personal project.
It was a personal project.
Yeah, yeah.
Isn't that unusual?
How can it be like one person's decision to like suddenly release something that effectively changed?
I think how we work worse?
I mean, 20B is not that much larger, but from 11B, the 11B, T5.
Actually, at that time, there was 13B empty 5, right?
So I think UL2 is an encoder decoder 20B model.
I think when we got it approved, it was released as like the big brother of T5, you know,
kind of like, okay, we updated T5 with like a new objective and train this new model and do DBM.
You want to, and it uses the same pre-training dataset and everything, right?
So like from PRC4.
Yeah, that was the easiest because there was precedence, right?
Yeah, yeah.
But yeah, there was some architecture like the mixture of denizers.
Yeah, yeah, yeah.
So back to Palm 2.
I think my involvement with Palm 2 came from the work to add UL2 to Palm.
to. And then, I mean, it was from the top-down point of view, I mean, the leads were decided
in a top-down manner. It's not like, like, what, there was not much, like, fighting or,
like, or any major things, right? It was like, it was a mixture, like, bottom-up, top-down-ish,
like, half-half situation. And then, like, from the top, it was like, okay, like,
these are the people who are the most visible in contributing to this workstream. And then,
okay, how about E and this other guy becomes, would be in charge of this, like, modeling
workstream and something like that, right? So I think it just happened that way organically.
And yeah, I think that that was how I kind of was co-leading the modeling upstream of Palm 2.
I think in retrospect, you understand now that this is very valuable experience. And I think now,
today, it will be much more competitive to get the job that you got, whereas you didn't,
two years ago, you didn't have to try that hard to get it. Or like, you kind of lucked into it
with you all too, and then like it just compounded from the initial good decision.
I think it's very hard to counterfactually analyze this type of things.
I think it's definitely true that there are more people working on generally AI now.
And if you are in a big company, it's way harder to navigate this type of things, right?
I wouldn't say that there were like nobody or so wanting to work on this at that time.
In fact, they were actually...
Well, you were the obvious choice.
There were less people.
There were definitely less people.
But I think I would say that maybe it's slightly harder now, but it's also not like it was easy at that time.
Yeah.
Yeah.
I imagine it's sensitive.
But also in my mind, this is now the most...
valuable on the job training in the world.
And so people want to know how to get it.
This is why I'm trying to figure out.
Like actually individually, we also cannot take like somebody else like experience
and then try to replicate it on.
Because everybody's circumstances, their initialization point,
their thing is kind of also like indifferent.
Yeah.
This is not only true for ALMs in general, right?
Because a lot of times like, oh, okay, you did this in this position and then because
of this is, like, it's very hard to trace all this style to find a causal path of this thing.
So I think everything in life, there's some luck involved.
Yes. Yeah, there is. Emergent abilities, very influential paper, subsequently contested by the Mirage paper.
Oh, yeah, yeah. So before we get to the Mirage, was there a story behind emergent abilities?
Yeah, you know, I'm sure it's Jason's thesis or like, just tell more about like the behind the scenes.
Like was there a discussion that led to it that?
This one was like, this is the idea, the inception of it was like mostly Jason.
Okay.
Right. I think I helped out to like, you know, shape up a little bit of the paper, get some,
stakeholders involved and stuff. I was discussing quite a bit with Jason, but this, the idea itself was
Jason itself. So actually when the Mirage thing and everything came out, I didn't, okay, I was
this hot takes for the sake of hot takes. I didn't feel, I believe in emergence. I have to just go on
the record and just say, I believe in emergence. And then I was not feeling very strongly because I think
that I can't speak for Jason, but I would as imagine that he would be maybe personally offended
because, because I'm thinking, I know, Jason is a person that takes a lot of like feedback like very
well. He's a very, like, he's not offended by harsh feedback, and he rebuts well, like, online
as well, right? But, like, I would as imagine. He would be the one that is the most, like,
affected by criticisms of emergence. I was believing in it, but I have to say that that paper.
I mean, that's why he's the first author and I'm second, but that was mostly Jason's thesis.
And I have to really say that Jason has really good ideas. And I was more of a support role for that
paper, yeah. Sure. Yeah. Lots more to discuss there, but you believe in emergence,
that's enough for me to work with.
I also think that the Mirage paper is mostly like,
I don't know who, actually I don't even remember,
who wrote it?
Rylan Schaefer.
I covered him on my Nirobs podcast.
Okay, okay.
He's a very good speaker, and the paper was well done.
It's just that people drew the wrong conclusions from the paper,
because he had a very good title.
Do you believe in emergence?
Of course.
Okay, high five.
I mean, how can you read any paper,
read any, the progress of LLMs and not believe in emergence?
It's so stupid.
Like, just because you read,
paramaritize some benchmarks and evils and make it linear, doesn't mean the emergence is completely
gone. And even in the Mirage paper, they acknowledged that there were some metrics that were true
genuine emergence according to that. I think it was something like 25-ish percent in the ballpark.
That's not the exact number. Yeah, yeah, yeah, yeah. So I was like, okay, fine, like some benchmarks you
disagree with, but on the whole, there is emergence. It's just, now we're just talking about the
magnitude. Yeah, yeah, yeah, for sure. I think, I don't think the authors of the paper had really very,
like they didn't, I mean nobody, we should just assume people don't have bad intentions, right?
But like, they definitely were just doing this.
But like, I think I was, I was more like annoyed by the nearest best people.
I mean, okay, best people was, just take a bit of a grain of salt, right?
But like, there were people who come on me like, oh, you should care about this because it's the nearest best people.
Because it's the nearest best paper.
I'm like, does best people awards mean anything?
Actually, it doesn't mean anything, right?
Like, I think that was more of my, where my angst was coming from.
I don't think I really had.
I don't even remember who were the authors of them.
That paper, right?
I'm sure they're doing well for themselves.
We don't have to dwell too much on that.
Okay, okay.
Okay, so a couple more things on Google, and then we can go to Reka.
Quokler was a manager.
Yeah, yeah.
What is...
I had another manager called Dawn.
Like, I had two managers during my time at Google.
So I'm just basically going to ask for quick hits from what did you learn from Qok?
What did you learn from, you know, one?
Oh, okay, very interesting.
Who they represent to you, like, how they advise you and all that.
So Quok, as a manager, he was more like a friend, and we were, like, talk a lot about research.
I think Quok is a very researchy person.
he has a lot of like good, like, he's more of like intuition.
I learned a lot about like, from him about like, there was no like concrete, like,
it was more like over time and it was very implicit soft kind of feeling.
But I think like a lot of research science, we were like brainstorm a lot about like,
I quite quite like that when we were, but there was this Yelpam paper that didn't like get
as much attention that I feel it deserves.
But like I think that was one of the works that I kind of like discussed with Quark quite a bit
and like at that time we were releasing the Flan 2 stuff and everything.
And then like I think Quok has a lot of good sense about like,
like what makes a work a good hit and like, like, you know, publicly a good hit and like a lot of
research sense about like what, what makes like research cool. So I think he has good like
intuition as a researcher and I learned quite a little bit about. And I also say that I think
Jason also probably learned like quite a bit from Quok and there's also influence his like, like,
it was not only just like me getting influenced, but like there was like Jason getting influenced
and then Jason influenced me. So I think overall what I learned from Quarks probably is more like
intuition, research tastes.
We would like chat about AGI
sometimes, singularity and stuff like this.
Like it was like, it's nice to talk to as a friend manager,
kind of, he's like kind of a friend figure to me.
He's very much a researcher more than like a corporate manager kind of thing.
I totally expect that.
It was fun.
It was fun.
Jason way, would you learn from him?
What is your distillation?
Okay, Jason is very interesting.
So I learned in my career, I learned two or three things,
major things from Jason, right?
So I think the first thing I learned from him is that
So Jason was actually, okay, I'm going to talk about
the more casual, more fun stuff first.
Jason was the more spicy on Twitter first before me.
There was an era where I was Goody Two Shoes.
I only had my main account.
My only tweets would be like New Paper Alert, right?
And then Jason was starting to post like hot takes, right?
And I just thought to myself, oh damn.
And there were types that I was like, Jason, you should not post this.
You're going to get cancer.
Right.
And he was fine.
He always braved through the storm and everything.
Until I looked at him and I'm like, maybe it's not that bad after all.
Do they just be, right?
So that was like kind of like, which is very interesting because Jason is much younger than me.
And I and the other thing also, we out of accounts, right, we created them around the same time.
And the interesting story behind it was that Jason's all the account and my account our own original.
It was not like an anime character out that nobody, I know who is it.
We have our identity.
It's pseudonymous, right?
And then I asked Jason, why do you want to have a pseudo?
like, why don't you just make, like, right?
And he told me this thing, which was quite true,
was that like, okay, you can post a take that it's spicy
and it's hot, but if you cannot stand by the opinion,
then you should not have the opinion in the first place, right?
Wow.
Right.
So there was something that, oh, okay, I thought that was profound,
because so far this, I mean, there are times where, okay,
I post something and it's spicy and then, okay, it gets a little bit bad.
And then I, okay, I cannot agree that, okay, this is bad.
Then I will retract it.
But if I could stand by the opinion, then I would just stand by it,
because that's the point of making it.
It should be said.
Right?
It should be said because I can put my name behind it.
So this is part of the first bucket about how kind of influence my online persona, like, a little bit.
And then I mean, it turns out that now AGIHIPO is so much more spicy than the cola.
Coala is just hibernating somewhere.
It's not even around, right?
So, I mean, Jason also is more constrained because he works for, he has like an actual employer, right?
And he has to be a little bit more.
My God, the worst thing about Twitter, you know, anytime anyone from openly I tweets anything, they're like,
did you see this researcher from openly?
I said something.
And they read tea leaves that are not there.
And it makes you very cautious to tweet anything.
And so it kills the golden goose, is what I say.
There was one tweet.
I mean, at the time when somebody was, people were speculating the GPD two chatbots, right?
And then Jason just posted something on his main account,
something excited about new experiments being run.
Just a random.
And then people screenshot that and posted.
Yeah.
I hate that.
So I think, now I think for all the account is mostly like personal, like personal stuff.
very personal.
I think he would stay away from
non-work things.
The golden goose has been killed
because people on Twitter
cannot control themselves
from drawing random conclusions
from all these hints and all that.
Yeah, yeah, yeah.
But going to like the actual,
like this is like filler,
this is filler.
This is filler.
I think the second thing I learned from Jason
is more about like,
from my, you know,
kind of like, from my own career
is like the importance of like marketing and PR.
So Jason is actually like super good at
I mean, I would just like,
he was actually like really the emergence,
like how many blog posts he wrote about the emergent abilities
and how many talks he's given about about immersion.
Like a lot.
Probably like the other day I was just at this webcam keynote
and he was giving a keynote again about immersion abilities
and it's been two years, right?
So I think one big success of him is that like he does the work.
He thinks a lot about like marketing the work itself.
I did not like in my early parts, my career,
early parts in Google, right?
I was putting out a lot of work,
but I didn't put in a lot of like effort in like thinking about the
like how the word is going to be received,
I'll just be like, here's a paper, here's a paper, here's a paper, right?
But Jason will be like, I'm going to write his paper
and I'm going to like market the shit out of it.
So I think I learned a lot about like,
so every single first author paper that Jason writes in the last,
has like 1,000 citations in one year.
Oh my God.
Like, no, I mean, not every, but like most of it that he leads.
So his head rate is very high.
He's hit rate, like impact density.
Like it's very high, right?
So it's pretty interesting.
But I kind of see him as like a peer and I learn a lot from his basically.
Some people are just like talented in,
different ways. And I think that, like, I looked at how he markets his own work and markets himself,
actually, right? If someone is studying from zero, like no Twitter presence, what is the second best
thing to do? You mean as a researcher? For marketing, yeah. I think you were like, the most obvious
thing to do, like if you're like a research, like say hypothetically you're like a researcher in like a
place without visibility or without, and then you have no personal visibility. The first goal is
always to try to find a mentor or co-author that is like within this circle. And then you start
from there, right? Because, and then you get people from, like, who has a visibility and following
to retreat. So you would, you would like work with them. The big goal is not about like,
I learned this. I mean, this is like probably a career mistake in my early days. It was that,
you know, instead of like focusing on like, so-called people, okay, if you do good work,
it's more of like, okay, how am I going to? I see this visible researcher from deep mind,
right? Or how can I collaborate with this person and then kind of do something that
feel is cool and like I can win their respect and that they were like, you know,
they would be willing to co-author for me because the exercise itself also about how to,
you're not trying to place reviewers or anything.
If you can find one semi-visible, you don't even have to be like a famous person.
That's like a semi-few thousands of followers, has a good reputation of research.
And then you collaborate with this person.
And then when you post the work, you are co-author with this person.
And then you get the person to vouch for you or like this.
Over time, this would, it could be from internships.
It could be from just DMs.
I think people are nicer.
Then some people, they seem scary, but if you DM them, they're actually reading to collaborate, actually.
I was scared of you, actually.
No, no, no.
And when I DM you, you turned out a lot nicer than I feared.
So thank you for being nice.
That's really great advice for people.
I just want to leave that all that for people.
For others who follow the work that, the career advice that I give, the title topic of this is pick up what others put down,
and specifically pick out what your mentors put down.
Like mentors always have more work to do than they have personally time for, the high visibility mentors.
and if you can show that you're a good collaborator with them,
they will lift you up accordingly.
That's a pretty good formula for career growth.
Should I ask about Hyeong-1?
I don't know how close you are.
Oh, we're still good friends.
Hul-one is a great engineer,
and he's very systematic in the way he thinks.
I think Hyeong-1 is, without going into detail too.
I still spend a lot of time talking to Hiong,
even after we both are different places,
about very interesting,
algorithmic ways to think about life.
Very interesting.
like perspectives on life rather than research.
But he's a great engineer.
And the one thing that scares me about,
he doesn't have multiple monitors.
He just codes you one small screen.
And he does everything with very hyper-optimized.
And then back at...
This is, I want those you curve where one screen, one screen,
and then many screens.
Yeah, yeah, yeah.
So I think Shawah would feel.
Shaw one scares me because it's like,
I think that was at Newark's 2022.
Like, we were doing some work at the New Orleans.
And then he would be like coding like perfectly fine
with like this, you know, 13-inch MacBook with like one terminal, and then he'll be like,
he keeps telling us, okay, it's more optimal to using keyboard. Like, keyboard is more optimal
than moving your head because if you can switch your screen fast enough, it's faster than
your head, like moving to different screens and stuff. I did not like actually distill that
because it's too painful to do that. But like, it's very interesting in a way that he belongs to one
of those hardcore people with one monitor. Maybe this is a relevant question to close out the
Google side. What do you think is a good programmer for AI research?
I mean, set up or you think? No, not set up. Not not set up. Not even lifestyle. It's more
about skills. Like what should people have? What do you interview for maybe? What do you see
the high performers do differently than less high performers? I mean, okay, like generally,
there's like, I think like for, for AI researchers, like being a strong IC is like probably
like the thing that I feel like it is like important for AI researchers. Like not, not, I think
there's certain level like sacrifice to be like.
like an AI engineers,
that AI research,
especially if you're training like L.
and because you cannot really be detached from,
your jobs could die on a Saturday at 4 a.m., right?
And then there are people who like would just leave it dead
until like Monday morning.
And then or,
but there will be people who will crawl out of bed at 4 a.m.
to restart the job or to check the tensorboard or something like that, right?
I think a lot of being a successful AI researcher,
I want to say passion is also the entire thing,
but it's more of just the,
the kind of personality that if something there's a buck at 3 a.m on Saturday night or something,
right? And then you would be like, you couldn't go back to sleep unless you, you, I'm not,
this is very unhealthy by the way. Like people should not do this for a long time. You know,
I think this kind of things actually like allows people to make progress faster, but it's unhealthy.
So I'm also not even sure like what's like the checking out on like Friday, Saturday, Sunday and
like work at 9 to 5 if you want to like make progress or like some people are just so good at detaching
like, okay, like 8 p.m, I'm not going to, my job can die,
and then the chips can stay either for like the whole night.
But I want to watch Netflix, right?
You cannot, I think there's a level, like it's like a sport.
You cannot win an Olympic goal if you want to have super ultra good work like parents.
Yeah, passion, intensity, dedication.
Yeah, intensity, right.
So those are really good personal qualities.
Just technical qualities-wise, how much of the stack should people know, you know?
Okay, so there was the question.
No, no, no, but that was important.
well. It's just harder to interview for because you really just see it on the job.
I think stack is not not not that like should know Kuda kernels. I don't know Kuda
kernels. Exactly right? Okay good. For all you listening out there you don't have to
feel like an imposter. No but but you need to be willing to learn if you have to I
think. Well you haven't had to so far. Yeah I haven't had to so far right. So if I
think pie torch, okay great you know what kind of do what do I know like distributed
systems like do I know like what is the stock that you recommend for people that gets
like a well-rounded end-to-end researcher.
I don't think there's any specific thing.
In fact, I will try to...
I don't really say like, okay, you need to learn jacks.
You need to learn this.
By the time you finish letting there's a new framework out.
Anyway, so it's more of like staying, like, constantly, like, trying to...
Being able to continuously learn and update.
I don't think that's a single stack or like a single, like, workflow.
I don't think there's a single one.
Yeah.
Well, that leads us to Rika.
What's the founding story?
So I met some of my other...
co-founders while we were collaborating at D-Mind. I was at Brain and they were like a
deep mind. I'm not like a startup person. I identify even today as a scientist and a researcher
more than like a startup person, right? My co-founder, Danny, started this story, right? And then
this record was like in the works from like late 2022. I finally left in 2003. Danny kept asking
me he wants to do something. Do I want to go with him and do it? And it took a while for me.
I was like kind of the last co-founder to kind of form.
Was the plan always for you to leave at some point and joined him?
No, no.
He was just like convincing you to do it.
It was like six months, more than, in fact, like, I think more than six months period of like,
I always had this at the back of my mind for since like August.
Actually, I didn't want to do it in the first place, but I think eventually in March,
I felt that, okay, it's time for me to experience something new.
Like my leap of faith was more of like, I want to experience something new.
I've, okay, I've like wrapped up this palm to work at Google and then like more of like,
okay, let me experience this new life and see where we can go with this.
And I also, I mean, we don't have a lot of like, okay, the funny thing was that like many,
many years ago before my PhD, I wanted to do a startup actually at that point.
And then over time, I realized that like I was better off as a researcher and I just forgot
about the startup thing.
And it's quite funny that today I end up doing a bigger startup, right?
But even until now, I actually identified more as like a researcher and scientist.
Well, I mean, it's not, when you left brain, you already had a high profile coming out of brain.
You could have gone to any startup out there.
They all have wanted you.
Yeah, okay, yeah.
So why did you choose this one, basically?
Like, it was just because of preexisting relationships because it wasn't obvious to me.
A lot of the other coworkers went to Open AI.
Others went to, if you're fair, you went to mistrial, that kind of stuff, right?
Like, Rico was like not on the, on the map.
I think it was, for me, it was a mission between staying at Google and, like, co-founding something.
I didn't want to, like, it was more of the experience of being a co-founder that, like, was
attracted me, right, and wanting to experience that I wouldn't have left for inflection or
something like that.
Like, I mean, inflection is gone now, but like, RAP?
They're still alive.
They're selling themselves as a model foundry or something.
I don't know.
They're a services company now.
Yeah, I don't.
But I also think that, like, like, for example, if you to join, like, another, it would be,
like, a very big tech experience again, right?
I don't know.
I felt like the experience I get is very complex.
inventory to what I have, that's the experience I had at Google, right?
But if I were to join something else, right, then I wouldn't have, I would have just
stay at Google, to be honest.
Because to me, it was very clear, just two dishes that I didn't really, I was talking to
a bunch of other startups, but I didn't really actually have the intention to go.
I was happy at Google, actually, to be honest.
I'm sure.
Yeah.
I'm sure they have a lot of things to keep you happy.
I was happy at Google, yeah, actually.
So you described yourself as GPU poor, but also you had $60 million to play with.
you got a whole bunch of GPUs.
I think you disclosed somewhere,
but I don't remember the exact number,
and you had a good training run
for Flash and then Core and Age.
How would you tell that sort of the story?
Like, people can read the technical report,
but also, you know,
what was that overall experience like?
And I should also point people
to the blog post that you wrote.
There were a lot of interesting things
that happened along the way
that led to our...
So I think I left around like early April,
like March, end of March, April and everything, right?
But most of our compute actually came in December,
actually.
Yeah.
And there were delays.
So H100,
they were major delays,
right?
So we were sitting around,
right,
bunch with like...
It's clear you don't own
the compute,
you're renting.
Yeah, yeah, yeah.
So we were sitting around,
like,
with,
for a long period of time,
we had 500 A100,
because we made a commitment
and they were constantly
being delayed.
I think because
H100,
supply demand,
whatever,
reasons.
And it was also very hard
to get a lot of
compute in one place,
right?
And then we were locked in,
and we had to wait
for the compute
to come.
Right. So I think it was very painful because even when the compute came, it was mostly broken most of the time.
And it was broken to a very bad extent that before I left Google, I was like, even the early stage, I was very optimistic about, okay, this compute translates to do this amount of flop.
This is a model, right? But I never expected the reliability to be so poor that it just threw off all the calculations.
And then we had to work 10 times harder just to make the thing go smoothly.
So it was a bearable pain. I think the pain was bearable, but it was.
just way more than expected.
I think you'd address this in your post,
but the temptation would have been
just to run everything on TPUs,
which is the stack that you already know very well,
that works there.
No, no, no.
So TPUs outside Google and TPS inside Google
are probably very different things.
Oh, how come?
Okay, firstly, it's like infrastructure.
Like, there wasn't like a lot of good code bases
like outside Google that was like still, right?
And the code base that I was most familiar with
was like T5X, it was a jackspace.
It would have been like,
by the time we wanted to consider it,
it was really like
debricated for nine months, right?
And then TPUs, I mean, we weren't sure about,
I mean, the availability of TPUs were also not great, great.
Oh, my perception is it was a lot better.
People have the learning curve.
Yeah, but at the point of time,
we had our infraset up.
We were training already training models
and it would be so much cost to switch to TPUs.
So I think TPU's, the experience of TPUs inside and outside Google,
I have not actually run a single TPU job outside Google, by the way,
but just looking through documentation
from what I see outside,
and from like how much
I think that people inside Google
don't care about what people think outside Google
I kind of feel like
okay we were a bit like
I don't think we considered
I mean not like forever
not considering this but like just like
at that point of time it was like
The obvious choice is just stick to pipelines
does stick to GPUs and Pythodge
and make like I mean it's not as if the chips
we ordered were not there
they were there they're just not in the best shape
Right yeah so I think it was too much work
to kind of migrate suddenly to
DPUs yeah
For those who haven't read the report, you had a very traumatic description about the chaotic and stable phases of various compute providers.
And I was just winceing when I was reading all those things.
Yeah, there was like a three-body problem reference to chaotic and stable phases.
I mean, I was watching three-body problems at the time.
And I thought it was fun to be, it was fun to.
There was a lot of like, I think we had a lot of fun adding a lot of references and mean into the type report.
I think it goes to show like how fun the environment is within record, right?
We had a lot of fun with this.
But so I think chaotic and stable phase, mostly it's like we actually found that like usually when like provider provisions new notes or they would like.
Yeah, you don't want to be the first to use it.
Yeah, it's usually like bad like dog shit like at the start.
And then it gets better as you go through the process of returning notes and draining them, giving it back to them.
They will send it back for repairs and everything.
And then over time, because it's more of it's a more of a numbers game, right?
if there's one bad note, it kills the entire job, right?
So, like, the fact of the game became like,
just eliminating bad notes from the thing, right?
And then, I mean, just because of,
maybe because of the supply issue or something,
when the deadline comes to ship this,
for example, like,
I just give rough numbers.
Like, say you order 1,000 H-100s, right?
They will not be able to,
usually they don't meet the demand of, like,
1000s,000 h-100s at the date.
They'll give you, like, 500 first stuff not to piss you off,
and then it will give you like another 100.
Like, every over two, three weeks,
they were just like, okay,
I added like four notes,
added like eight notes,
that kind of thing. And then over time, we reach the capacity that you, or you actually,
maybe you never actually ever reach the capacity that you ordered for. And then, like, as they
add these notes, right, sometimes these notes are bad. And then they just kill entire training
runs. And the thing, which I feel that, I mean, for all those people trying to sell GPUs,
there are a lot of people trying to sell GPUs now. We sell, sell package, whatever, GPUs, right?
And I think the most important thing that there are obviously, there are SLAs, all this in, in the contract
and everything. And obviously, you might be entitled to something, something, if something goes
wrong, right? The thing that for
large model training runs is that
one bad note queues the entire job, right?
So should the compute provider be
liable to pay for all the
node waste stage that... No, it's
unlikely because otherwise...
It's unrealistic. No one will take that on.
No one to take that on, right? So I think that's also like
a tricky thing. Who is taking the risk?
Is the LOM startup taking the risk?
Or is the compute provider taking the risk?
I think that, I mean, this is
my sense. I'm not 100% sure, but I think
as there are more providers trying to sell
our GPUs, we get all this inbound so much about people trying to sell us GPUs, right?
The key differentiator is actually to find a way to balance the risk of no failure with,
as long as the provider, I'm not going to say 100%, but if somebody can come and tell me that
my notes are so stable that I can share some cost with you if your job dies, this is green flag,
green flag, right? The moment they start to, I cannot do any of the big clouds do that as far as I know.
Because they have the size to guarantee that. But I think for anybody who was watching.
or if you do it like a compute startup or anything,
the biggest green flag would to be to share the cost of note failures with your customers, right?
Because the whole run?
No, no.
Like if the note, it's very hard to go.
Because you need software to like, you need software to.
So let's say you run it for 12 hours, right?
And it dies after 12 hours, right?
You get 12 hours of throughput, right?
But then you get like some wastage because of like the, you know, the downtime and everything.
Right.
You know, I think it would be fair to find some middle ground to kind of split the cost.
of the failures, right?
And this brings back to my point
about like work-life balance
because if the notes fail so,
fail so badly, right?
Like, it actually,
basically, right, your engineers
cannot sleep at all.
Do you have babies sitting,
rosters and everything,
but you are living life
with like constant anxiety
because even in the case,
right, where the note failures
are refunded, right?
You still lose time.
You lose three hours.
Sure.
You lose everything, right?
So I don't know how to go around this,
but I think if there are a lot of
like compute providers
like fighting over,
I think a good,
good thing to do.
do is to figure out like this pain point otherwise or at least figure out some hot shopping
but but so far most of the providers that we tried don't don't have this they will also get
confused when you try to ask them so my job is dead can you pay for the food can you refund for
or at least they will get confused because like this is a LM specific thing that are the large notes
they don't care about yeah yeah they did they get confused about about this right so current
status code is the LM startup pays for everything but maybe you could negotiate some like like
refunds but usually they will not be so generous to pay for the say you
you run 500 GPUs, right, if you break for four hours, they would in their mind, they would be
thinking, I should refund you for one node, but in your mind, you just think that I should,
they should refund you for the full job, right? Everyone who is from my background is going
to be asking this, how is it so fragile? Like, what's your frequency of checkpointing?
Our checkpointing is kind of like, we see how stable the job is, and then we decide, because
checkpointing takes a, without a good file system, checkpointing takes actually quite long. So it could be
It's like a few hundred gigs, right? Yeah, I think so. I think so. I don't remember of
But sometimes if your file system is slow, right, your file IO is slow, your checkpointing
for a 20B model could be like, what, 30 minutes or something?
Okay.
I don't know this by head.
Sure, sure.
But it's not hours.
If you go larger, what if it's like a 200B model, right?
Okay.
So you should have some kind of ideal checkpointing to run ratio that is not catastrophic
if you run into a known failure.
Yeah, no.
So we see of it as like an MFU, like because you can average out your flop utilization
and then you can see how many percent hit,
like how much slow down, right?
So you probably go for something like
if it's like you're taking off 1% of your speed,
2% of your speed.
So basically it's actually fine to the checkpoint more,
more regularly, right?
So I think checkpointing, like you also never fully,
you can get like from the clean slate like nothing, right?
As you optimize like engineer like the system to automatic restart,
everything, you get some of the time back,
but you'll never be perfect, perfect.
So you still lose, lose stuff like that.
If you checkpoint too often like everyone,
what, every 30 minutes, then your file system is going to blow up, right?
If you're going to checkpoint every, like, like, for us, we just see as like how much
storage is cheap compared to compute.
No, when your model is like very, very large, your storage can, can easily blow up.
Going on to the models, I feel like I digress so much about all this fun side things.
You like compute, right?
You like hardware and compute.
I know hardware and compute.
And also I'm an orchestration guy.
So one part of the question, one of the questions I'm skipping right now is there's, I came from temporal.
I'm familiar with Kubernetes, I've used Airflow.
These are all the data engine cloud or cloud engineer type tools.
It's surprising to me that you guys don't have your set of organization tools that it's solved.
You wrote in your blog post you had like the pain of multi-cluster setups.
And like to the rest of us, this is completely solved.
Okay.
I don't know if you know that.
We use Kubernetes for a bunch of stuff.
But like I think like for experimentation and like stuff like this is still not fully like we, we,
We didn't have the time to actually build something that is.
It should exist in open source.
Someone should have done this.
Okay, okay.
I'm not, it is what it is, but I'm surprised, that's all.
Okay, okay, okay.
It seems like a valuable problem and someone should do it.
Okay, okay, okay, yeah, yeah, yeah, good to know.
Good to know.
Okay, so Raker Flash Core Edge, you know,
congrats on beating a whole bunch of state-of-the-art models,
especially much bigger than each.
People can see the papers for all the other stuff.
Was this your expectation from the start that you would basically definitely be frontier?
Like, how do you, like, from the start of, like,
haven't trained anything yet, and you're about to kick off the runs.
Like, are you able to call your shots and say, we will beat GP3.5?
Nobody can predict the future.
No, how much confidence?
Okay, we were confident.
Like, we were confident.
How?
Why?
All right.
It's a good question.
Because it'll be a shame to do a whole bunch of work and then end up, this is the middle of the
pack, which a lot of people end up.
We were confident.
I think that a lot of it was like Yolo.
I mean, I mentioned in the, in the thing.
I think we would, like, require a lot less iteration than this is because of our prior
experience in like training these models.
Like so I was confident in,
in myself about like,
our models will turn out to be,
to be,
to be,
to be good.
And I,
about exactly how I actually don't really like pinpointed to a particular
reason of like,
I mean,
we de-risk stuff.
So a lot of part of it is like,
de-risking and like,
okay,
you run like 4B as ablation and you can see,
okay,
this is like my spice,
if you run 4B and your loss is like going crazy,
you know that,
okay, this is going to be a shit model, right?
But I think it's like,
we train enough like,
okay, we don't have a lot of compute to do a lot of violations,
but we did enough experiments to know that,
okay, our infrastructure and everything is set up to be good, right?
Obviously, the field moves, right?
I wouldn't say that everything was like smooth,
like the first time round is like smooth and everything.
But I think we were confident in our ability to like make the least,
like we're not like really, we're more confident about like the ability to like
move with as little steps as possible to the goal,
more so than
my model is going to be this
level at this time
you know what I mean
it's more of like
for example we let's say
we run the first round
of human evaluations
right
and then we see our number
as this
right and then we're confident
that in five more tries
we'll get to this
you know
kind of like get to like like this
is more of
that kind of confidence
rather than actually
it's also a little bit of
you see a new leaderboard
hypothetically
like in like if as a researcher
you see a release
a new lead
later about, right? You approach it like a puzzle. You don't know, like, whether you, at the
start of it, you might not have the answer to the puzzle. But if you're good at solving puzzles,
like generally, right, you know, you know that with one hour, I'll be able to solve it.
You know, that kind of confidence, like, it's the ability to, to heal climb or the ability
to improve over arbitrary things, right? Rather than, I think we were confident more about that
rather than everything is different, right? The stack is different. The infrastructure is different.
The data is also different from what, I mean, we have a lot of, right?
It's just, we have a lot of, yeah, we have a lot of experience from prior, like, our jobs, but like, it's not going to be the, like, we don't have actually, like, exactly the same thing because different companies have different stacks, everything, right?
So it's more about derisking, being confident in, like, solving the general problem of, like, improving over things, which is why also I think that the team is valuable in the sense that we are not, like, valued by our model itself, but we are just valued about, like, like, like, how we can see one problem and we can just, like, solve it, like, super quickly, right?
That's what we're confident about, right?
Actually, like, the artifact itself.
Mentioning your team, you said at the largest, your team was three to five people on the pre-training side.
It was that the team that you recruited?
Was it all your ex-colleges?
How do you find people that would have this kind of solid intuition?
So I think that some of the people in our team were like, I worked with them at Google,
at ex-colleagues and stuff.
And some of them were like fresh hires.
Like they were like fresh PhDs or like everything.
Okay.
So I do want to comment on no architecture.
So if you want to, people have variants of all these.
Swigloo, GQA, rope, RMS norm,
and then obviously the big one is encoder,
decoder versus decoder.
Could you comment on each of those?
Like, were you just like,
we're confident and known got it right?
Or did you actually do a evaluation of each of your architecture choices?
Oh, I mean like, okay,
architecture-wise is something that I feel like I'm easily able to,
like, I've run so many architecture experiments
that, like, I look at architecture and I, like,
I don't want to be like overly,
I think it's very hard to outperform
the OG G G.
Why?
I mean, on the surface of it,
like, we have to have learned something in the last
seven years.
No, all the changes, all the changes that
like, like, Suiglou was this, like,
okay, Suigu is like probably one of my favorite papers
all of time just because of the divine benevolence.
Like, the gnome actually wrote like,
we owe this success to divine binel.
Like that was like, it's always a meme thing, right?
Okay, so like GQA,
MQA was always like, the multi-currier that was always like
a big controversial thing because
MQA usually you get a hit because it's
MQA and everything.
So people kind of know that, like, it was a very...
Like, hit or miss.
Like, you could get a hit in the performance from MQA, like, MQA alone.
MQA was always like, you know, the choice, right?
It's always like, okay, should you use MQA?
Should you not use MQA?
Right?
It's always, should you not use MQA.
When GQA came in, right, it became like a no-brainer to use GQA because you don't get the hit anymore.
Rish, right?
The number two, the 70,
GQA, right?
But, I mean, the reason why we call it
Nome architecture, because MQA came from
Nome and GQA was like a follow-up paper
by some of my colleagues at Google, right?
So I think GQA became a point where,
okay, this is already accepted.
Like, it's a no-brainer to use GQA.
Sway-Glu was an interesting thing
because there was a very long period of time.
It also Swiglou was a single-authored paper by Nome.
And very few papers were, like,
Swigrew had very few citations, like, at the start.
Only Google papers were citing Suiglidu at one time.
And a lot of them
I was, at one point I was like probably like 30% of Sweet Glue citations.
Because every time like, like, Swigroup became popular because of the updated T5,
the T5, the T5 1.1 that uses Svigl, right?
And nobody actually really cared about Svigl for a long time.
Because I was checking why is this like underrated paper like not getting much citations.
And then I think probably now it has like a few hundred citations by now.
But I think Svigrew is one of the things that like, you know, I played around with a lot like at Google.
So Svroup really works.
there was also a paper we wrote about,
like do transformer modifications,
blah, blah, blah.
Like, it was a paper with Norm and
Sharon and Tianwan and stuff like that.
And then we ablated like so many
transformer variants.
Yes, yeah, I saw that.
Some of them matter, but most of them don't.
Most of them don't.
And then the only thing that matter
in that two paper was,
in that paper was Suiglil.
I forgot which exact sweet glue variant was it,
but and sparsity at that time, right?
So that was strong enough
like to finding to...
For the listeners, this is the inductive bias.
Scaling loss versus model architectures, how does inductive bias...
No, no, not this one.
There was another one, like, to transfer one modifications,
something, something, something.
Okay.
First, portal was run, I think.
You should run around.
You gave the keywords.
Yeah, yeah.
I think the RMS norm-Rope thing...
Not controversial.
Like, it's not like...
Obviously, I think rope is probably, like, it has that extrapolation thing, which is nice.
And then it's also like default now.
Nobody wants to add production and embedding.
anymore, right? And I think, I mean, I like the T5 style relative attention for a bit, but like, I think, okay, rope is, I actually ran that ablation for palm, like the T5 relative attention versus rope. I think rope is similar to other things, but he has this extrapolation thing, which is nice. And like...
Which is why your long context version can go to 256. Okay.
This, for all, most of the long context models, they use the rope extrapolation thing, which is nice property, right? So that there was for rope. I think there was also like some things like the layer norm.
like petitions and stuff like that
that were like, it mattered a little bit
maybe not too much and everything.
But I think in general, there was not a lot of
like, there are not a lot of things that people could do
to the transformer to, it's been like four, five years, right?
And then the vanilla transformer,
I think if you use it as it is today,
would not be like that optimal,
but like the transformer that we slowly evolve to now
is like the gnome transformer
is probably like very, very, very strong baseline
that is very hard to like.
I think you need a drastic shift to,
to beat that, right?
Mistakes-based model type-ons.
Or you could find more like, like,
like, Swigrew is a small change, right?
You could find like some small change
that are like a big enough impact widely
that don't cost a lot of
because a lot of architecture changes, right?
The moment they are tedious to implement.
Like nobody,
when Su-Go is a simple thing, right?
It's a very simple thing to implement.
Maybe that's why it's caught on
because it has like a additional boost
that's for the simplicity of it, right?
So there's also a bit of implementation lottery,
if you will, right?
A little bit of, if you propose some very complicated thing
for like 0.1%.
Yeah, nobody will use that, right?
The biggest, biggest, I mean, I can't believe we have,
we're taking so long to come to this topic,
but the biggest gnome architecture decision is encoder-decoder
versus decoder only.
So encoder-decoder is not like a gnome.
The gnome architecture is more the...
Okay, maybe like more old-school transformers.
Maybe we want to just talk about the decision on encoder-decoder
versus decoder only.
So, okay, I wouldn't be able to comment about like exactly our setup,
but like, I think encoder-de-corder-es.
are kind of very misunderstood from thing, right?
So there's encoder decoder, non-causal decoder,
which is a prefix LM, and then there's a decoder only model, right?
Technically, a causal decoder and a non-causal decoder
are very similar in the sense that it's just a bi-directional mass, right?
And then a prefix LRM and an encoder decoder has only,
the only difference is that encoder decoder splits the inputs and targets
into different non-shad transformer stacks,
and then like there's encoder buttonnet in the data.
end, right? So technically, people like kind of always associate like encoder decoders with like,
I like bird or like something. You know, people get confused about these things, right? But I think in the
UL2 paper we really like kind of explored this and also like maybe some of the big science papers
that also talk about this, right, is that prefix IOMM and causal decoders are very similar.
That's the most. At the end of the day, they're all auto-aggressive transformers. That's actually like
the only big benefit of encoder decoders. It has this thing called like, I mean what I like to call
intrinsic sparsity.
Okay.
So basically an encoder decoder with like N parameters is like basically if it's like
it has the cost of like an over two decoder model.
So it's a bit like a sparse model because you actually spend the same amount of flops.
It's just that you have two sets of parameters like for encoder and decoder, right?
So it's actually flop matched with a decoder model of like half the the parameters.
So like a like UL220B is actually about a 10b decoder only model.
So you get free sparsity from that.
It's something that, okay, the old G-T5 paper talks about this.
You can look at that.
There's this complexity chart.
I didn't, like, when doing the UL2 paper, I kind of like was mind-blown by
while encrode-corder is so much more, not bounded by the causal mass anymore.
A lot of the efficient transformers, like a lot of the sparse transformers, like,
I mean, the early days, there's like lean formal and like whatever things like this.
They cannot maintain the causal mass, and that's why you cannot train proper language
model with this, right?
if you separate out your very long context into the encoder,
this encoder has no loss, right?
You could just do like aggressive pooling.
You could do some crazy sparse attention
that has like final transformer or something like that.
Right.
And then you could make that smaller than the decoder.
You could make that faster than the decoder.
That are just some of the advantages of like
why splitting into encoder decoder.
It could be beneficial to like just using like a decoder only model.
At the end of the day, the decoder in an encodecoder is a language model.
It's still a regular auto-rogressive language model.
So there's actually, I mean, it's not that much different from like a retrieval or mentored language model.
This is news to me.
I don't know if you've ever expressed this, but yeah, this is actually, it makes sense.
Okay, okay.
I don't know, unfortunately, I don't know enough to push back on this.
But on the surface of it, it seems to make sense.
Would you make the same choices if you were not so focused on multimodality?
You know, that's one of the ways in which I was thinking like, oh, and,
Dakota-decoder makes sense, that it's more natively multimodal.
I just have to say that it's relevant.
It's relevant.
Yeah, it's relevant.
Yeah.
Then we can move on to broader trends in LLM, just commentary on just ecosystem stuff,
like completely independent from Rika.
commented on a few things, like Lama 1 to 3 glowed up a lot.
I call this the Lama 1 to 3 glow up.
Like they improved into like an actual top tier open source model.
Yeah.
Phi, one, had a lot of criticism, but it seems like Phi 3 is getting a lot of love.
Do you just generally see like in your open model tier list
Like what's going up and down?
I think Lama 1 and Lama 2 are like quite mid, right?
But Lama 3 actually got good, right?
I think Lama 3 is actually strong, right?
I don't really follow far and much just that...
Their whole thesis is the textbooks is all you need thing, right?
Like that we can use way less data than everyone else and still...
But I think you cannot cheat the scaling laws, right?
Because like you...
I remember thinking like vaguely saying that, oh, they match like Mixtra 8 by 22 or like something like that on like some...
Okay, I don't think these academic benchmarks are like that meaningful anymore, right?
So, but then like, then when you go, they go on IMCs, they get like, what, 47?
And then they get like, maybe it just like seem slightly.
Maybe it's like, five two.
I don't know about five.
Oh, there's five three.
No, I think.
It was just released like yesterday.
Oh, I don't even.
Yeah, but, but I, I don't know.
I think there's, there's some, I don't follow five that much, but I don't like,
a model that is synthetically, actually, I don't even know this where, like, I didn't
read the paper.
But I think that, like, a model that is like based on the premise of like distilling,
and stuff,
something like that is like
not that interesting to me.
But I think that like Lama Tree
actually shows that kind of like meta got a pretty
good stack around training these models.
Oh,
and I've even started to feel like,
oh,
they actually kind of maybe caught up to Google now, right?
They kind of feeling.
There's also maybe a hot tag guy itself.
But yeah,
I mean,
FI don't really kind of follow you that much.
And I just,
yeah,
I mean,
there's too much,
too much things to follow.
So I think it's like,
I,
I think like,
Lama Tree is probably like the most,
the first most legit.
When you say these kinds of things, like most legit, obviously there's some, there's vibes
Eval or whatever, but I feel like a lot of people, the very common feeling is MMLE is kind
of saturated.
Yeah.
So like what do you look at now?
Is it just LMSS?
Okay.
So I think that LMSS has these problems also.
Yeah.
So Almsis is not like exactly like, I mean it's probably better than all these regular
benchmarks, right?
But I think like a serious LIM that's create their own Eval.
And a good Eval set is one that you don't release.
a good Eval set is the one that you
like, okay, you release some of it but like
it's like you don't let it be contaminated
by the community. Yeah, I think
I don't see this is probably
the most legit one.
I mean, the things like GSM make human care human eval
they're all like
contaminated.
They're all like saturated contaminated.
No. GSMK whether you're 92, 91
like no one cares right, a kind of thing, right?
But we'll still report three decimal places
in all of our reports.
Yeah, yeah, yeah, yeah.
But it's kind of like almost like this like obligatory thing to do.
You're a table of numbers of your thing at the bowl.
It's interesting to see how the feel evolves also over time for this type of like benchmarks.
But I think evolves are going to be important.
And it's on the actually interestingly, it's on probably on the academics to set the correct.
I mean, they have like there have been academics have always been, oh, we have no computer.
But like, okay, this is your chance to like steer the field in the right direction, right?
I think the challenge is getting attention.
So now MMRU was reaching its end of its life.
Like what is next, right?
There's MMU or there's MMLU Hard, which someone recently released.
It's pro or MMU Pro, I think.
It's called MMU Pro.
Oh, yeah, that's right.
That's right.
But like that only lasts you like a year.
Right.
And then you have to find something else.
So I don't really know what is that.
Well, so one thing, you had a comment, I think, in your Rika paper about there's two types of evels.
This is a vibe evel paper.
One is LMS as judge.
And then two is arena style.
Right.
as sort of the two ways forwards for just general evils that cannot be gained.
Although there's also human evals that you, like, instead of L.O.R.M.
as a judge, there's also like human evils that you run.
Like, that's kind of similar to arena, but kind of different to some extent or so.
Different in a sense that.
By the way, do you use your own staff to do that?
Or do you, like, hire an outsourcing firm?
Now, we don't even, we have a, like, we work with third party data companies to, like,
there are a bunch of this, like, around, right?
But, like, obviously, we don't like, eval them ourselves.
Like, I don't know how much, how many eval how many,
you want to do, right? Sometimes the best researchers do their own evils.
Yeah, looking at the outputs and stuff is something that, like, researchers should do.
Yeah. Well, there is one element of parametric evils, which I'm hoping that more people come up with,
where like you kind of, the benchmark is generated from a seed, that's it. And you can withhold
the seed or like, you can vary the seed. I can report how your model did on the benchmark,
given it, given a certain set of seeds or whatever, and you can maybe
average them. But in that way, it becomes harder, much harder to contaminate. I wonder if that
is an example of this. Not specifically, this is just something I'm wondering for myself, but I did,
someone did recently put out GSM 1K, which was, oh, the scale thing. I think, is it scale AI?
Yeah, yeah, yeah. Which is similar in that respect, like, make it easy to make variations of
a WAMLenem benchmark, but like, that is more likely to be worthheld from training data.
Yeah, yeah, yeah, but eventually those would, like, so it's always a sign. Like, even if we put out
vibe be very obvious.
So I quite like
upfront with like if the more people
use it, there's a lifetime, it's like a car, right?
After you run a certain mouse,
it's time to shelf it, right?
So I don't think there's like
actually like a good solution.
In general, I'm also like a bit,
I mean, I think this is like important
for the community to think about, right?
But like is it like a fundamental limitation
that any benchmark that goes out?
Like also there's also one thing is that
in the past people used to like
which whole test set, right?
Like squat or so they used to which whole test set.
But then like the like,
After a while, I think people also realize that, like, when you withdraw, like, MMMMU,
Kaggle.
No, like, when you withdraw, it's like so much extra work for, like, the community to, like,
on this, that they just don't do that, right?
It's either your data set become, your benchmark becomes unpopular.
I think it's also incentive things, right?
So if, let's say you are, you want to run, like, a contest, right?
And then your goal as an academic is to get as much citations as possible on this benchmark paper, right?
Like, then you, or like, this, you want to be as famous as possible.
you would not want to withhold the test set
because if you withhold a test set
and then people have like there was once like
I mean like many years ago
there were even some benchmarks where you had to like
package your model and send it to them to run
like this benchmarking never ever like took off
like took off just because like
so at the end of the day right it's like
it's the root problem like incentives
like also the benchmarking problem is also like an incentive
problem right so like it's also people on the show
their model is the best and then the game masters
want to gain as much cloud as possible
and I think also AMCs also get into got into some
I don't have a take of this,
but there's people who also feel that
they're also optimizing for hype, right?
Their own cloud, right?
So there's all this.
I think it's a lot of interesting.
Like, I don't know what I feel this will be,
but like sociologic.
I don't know.
Like, like, I think there's a lot of papers to be written, right?
I mean, about how these incentives, like rewards and incentives,
like kind of be, it might not be soft.
So I don't know.
I'll say sweet bench is probably the one that's kind of broken out this year
as like now a thing that everyone wants to compete on
is if you're a coding agent.
I don't know if you have a view on it.
But it's just like, it should be known to be hard.
You should be able to make progress on it quickly.
That makes you popular and cited a lot.
Yeah, yeah, yeah, yeah.
Multimodality versus Omni modality.
So this is a little bit of commentary on GPD40 and Camelion.
I don't know if you saw the Camelian paper from Meta.
Briefly saw it.
Yeah, I'm not, I didn't really take a look at that.
Basically, the general idea is that most multimodal models, like Lava or Flamingo,
which are late fusion, which is you freeze, freeze,
and then you join together versus early fusion
where you do it properly, where like everything is,
all the modalities are present in the early pre-trained stage.
And it seems like things are trending
from late fusion to early fusion is the general thesis.
With GPC40 being very obviously early fusion,
you guys, I will class it as early fusion.
I don't know if you have commentary on whether this is obvious to you
or this is the way or they will just be, they will coexist.
I think whenever possible, like early fusion is better.
I think there will still be a lot of works that do late fusion just because of, like, it's a...
GPU poor.
No, no, not GPU.
Okay, partially, right, I see this as an artifact of the line between language researchers and vision researchers.
And more of like, okay, like people who are training language models, they put out like a llama or whatever and then somebody takes it and then do late fusion on top of it.
It's more like a...
It's always a...
It's coming the orchard.
Yeah, yeah, yeah, I think so.
I don't know what law was it.
law. Okay, I didn't know about that. But it's kind of like an artifact of the organization
doing thing. Right. No, it's just because people don't have money to train things from scratch.
I don't know. Even in big companies, right? Like, I mean, I don't know how things have evolved
in many companies, but like... You're talking about Flamingo?
Like language and vision teams don't use to be the same team, right? So I think this is like
a artifact of this. But as early fusion models get more traction, I think the teams will start to get
more and more.
It's a bit
of how all the tasks
unify,
like,
from 2019 to like,
now it's like all the tasks
are unifying.
Now it's like all the modality
is unifying.
And then I think like
eventually everything moved to
us like early fusion.
Yeah.
The other element of multimodality
is I've been calling this
screen modality,
screen vision versus general vision.
In a sense that
ADAPT is like very,
very focused on screens,
tables, charts.
Most vision models
focus on
things in the real world and embodied sort of images.
Do you have a view on the usefulness for this?
I don't think there's like a huge, like, I mean, I think at the end of the day,
like maybe screen intelligence is like more useful in general.
But like what if you have like a natural image in the screen?
Yeah.
No, I mean, no, I think in the end of the day, it should be mixed, right?
If a model can do natural images well, it should be able to do screen well and everything.
I think at the end of the day, like the models would become like, I don't, I don't see
that there will be screen agents and like natural image.
Humans like you can read what's on the screen.
You can go out and appreciate the scenery, right?
You're not like, say, I only can look at screens.
Right.
So I mean, I think eventually the models would like be this good on everything.
I look at it from a point of like capabilities and screen is.
Even screen, there's also like mobile phone screen and there's also laptop screen.
Like also different type of interfaces and everything like reading emails, whatever, right?
But like reading a page from a website or buying something for Amazon or something, like all kinds of things.
right, and then even in the picture of like a shopping website,
there could be like a natural,
like for example, like picking Airbnb, right?
Then there's a natural image in there and then it's like,
you have to understand like how nice is the scenery, right?
Or like, like, where is it, right?
Yeah.
So I think the end of the day is probably like the same.
If you want to build a general model.
Yeah, yeah.
But I think the natural images is like way easier.
Like as in this way like the models currently,
current models are actually already very pretty good at this natural,
natural images.
And I think like screen images are just something
that people need to enhance the capability
a little more. That's why there's like some focus on.
Got it. I'll touch on
three more things and then we'll just go to career stuff.
Scaling laws. Palm 2 was
Chinchilla, which is one-to-one
scaling of model parameters and data.
Now you are training a 7B model
with 5 trillion tokens. What are you thinking
about the trend in scaling laws
for data versus clients?
Chinchela scaling laws that's like optimal for
like with this amount of compute how much
it's the thing, right? But like actually the optimal
like there's no, I mean this is something that
that even before I left, we really knew that chinchilla scaling laws are not the end of it,
right? Obviously, there's also an inference optimal scaling law, which is, obviously, you take a
spot model, and then you just blast it with as much compute and data as you can, until,
until you saturate on everything that you care about, right? So I think, like, Lamar trees are what,
15T tokens or something, right? So I think...
Which is ridiculous. It's ridiculous to be honest. But at a certain point of time, your value per flop
is like not great anymore because you just... Your models are, you just...
eventually gets like saturated.
But then the problem of like, the question of like, where is this saturation?
It's also like, you always find like some metric that you still continue to improve a little
bit.
And then you're like, okay, maybe oh, 100K more is worth it to continue training.
Like just a little bit more, right?
But then it's like, where does it end?
Right.
But I think at the end of the day, like, the thing about Chinchilla's getting lost is that, like,
it was a bit misunderstood as though this model you need this compute.
And if you train the Chinchilla's going law, like you kind of, I don't know why so many people
had this idea that you won't improve past the Chinchilla's scaling law.
And then people make so much big deal about trading past chinchilla scaling law.
Like, oh, Lamar do is the first model.
Like T5 base, right, was 1 trillion tokens.
That was really so much beyond chinchilla scaling a lot, right?
Because that was T5 base, right?
I think OPT and GPT maybe set that as an industry standard.
It's GPT three specifically.
No, sorry.
Wait, GPG3 was not chinchilla.
No, I think like OPP and Bloom, right, models like this,
they train a large model and with a very small number of tokens
and the model turned out to be bad.
Yeah, yeah.
So I'm talking about Kaplan, the pre-Chinchancella one, the Kaplan scaling laws.
Oh, okay, okay.
That one was from opening eye.
Anyway, death of chinchilla covered, agreed.
But Chechnya is still a cool paper.
I think Chechnya is still a cool paper.
I love any scaling laws paper, to be honest.
It's like such a service to the community in general.
Hanging face recently did one data blations, which is like a data scaling loss paper,
looking at data constraints, which is just kind of nice.
I see.
Long context.
People are telling million token contexts, two million token contexts.
2 million token from Gemini,
Magic is talking about 100 million token.
How important is it, do you think?
I think we need to solve benchmarks first
before solving the long context, right?
We have your benchmark.
No, no, not like the benchmarks for long context.
Okay, yeah.
Because like you, the needle in his stack is basically like
the MNIS, like it's always like a unit test for this style of things, right?
But I think there's one part about like hitting the context line
and the other part about like actually utilizing.
Utilizing.
Right.
I think Gemini's long contact is surely like amazing, right?
But I think for the community to move forward in this, then it comes to a problem of like, how do we evaluate this?
I think I see some long context benchmark on, like coding one and stuff like that.
Like I think making those are important and for the community to heal climb.
But I think long context is important.
It's just that you don't have a very good way to like measure them like properly now.
And yeah, I mean, I think long context is definitely the future rather than wreck.
But I mean, they could be used in conjunction.
Definitely.
Okay.
Okay.
That's not a take.
Which part of the...
Long context is the feature rather than RAG.
Like you would...
They will coexist but you are very positive on long context.
I will put myself on the other mirror image which is like long context is good for prototyping,
but any production system would just move to REC.
There are a lot of application use cases where you want a model to take that time and then come out with the right answer, right?
Sure.
Because Rack is like...
But you'll use those sparingly because they're expensive calls.
Yeah, it depends on like the nature of the application, I think, because you know, because, you know, you...
I think because in reg, right, like you, there's a lot of issues like, okay, how you,
like the retrieval itself is the issue or you, you might get fragmented.
It's like, what if it's like a very complex story, right?
Then you like a storybook or like a complex like thing, right?
And then like, like, like, like rec is very like, you kind of chunks, chunks and chunks, right?
The chunking is like, and you definitely have lots of information, right?
So there, I think there are a lot of application use cases where you just won the model.
It's like, okay, like, 100 bucks, like take your time, take one whole day.
come back to me with like that answer, right?
Rather than like, I pay like, like,
like one cent and then like get back a wrong answer.
So I think that's like, like,
that it's actually very easy to show that wreck is better than long context
because there are a lot of tasks that don't need this long context.
You like, like fact retrieval, you just like rack and then you do this thing, right?
So like long contacts may get a unfairly bad rap sometimes because like it's very easy to show like,
wreck is like 100 times cheaper and it's very easy to show this.
Right.
But then it's also.
like not so easy to emphasize the times where you actually really need,
like the long context will really make like very, very, very, very, very good decisions.
So yeah, I mean, I think both have pros and cons depending on the use cases.
Using them together is also interesting.
Like at the end of the day, it's like a H-Pram that you have to wiggle around.
Yeah.
There's another wiggle on the H-Pram.
There's another fog on H-per-M, which is how much you fine-tune new knowledge into the model.
Are you positive on that?
Do you have any views?
So, for example, instead of doing Rack,
on a corpus and then inserting into context, you would just find two in your model on the corpus.
So it learns the new knowledge in whatever capacity, right?
This is cumbersome, I guess.
This is cumbersome and you don't want like, you don't want so many of, like, the point
of in context learning is so that you don't actually have to do.
I think this one is depending on like a business use case, right?
If you find it is actually like the, you are very clear like you want this knowledge and
then you just find you once and then you don't ever have to pay like context, like in the context
window cause again, then maybe that makes sense.
But if the domain is changing, then you might not like.
Yeah, obviously it doesn't make sense if the domain keeps changing.
But I think for the model to maybe update fundamental assumptions or, you know,
reweight associations between words for, let's say, illegal context versus financial or medical
context, like it might work.
This is the arguments that some people are talking about.
So I see this as a trio.
Like, it's long context, it's rag and it's fine tuning.
Like people always have this like whether either of them will kill Rag, basically.
because Rag is kind of the simplest approach.
Yeah, yeah, okay.
I mean, I could see, like, if you want a, like, a model for medical domain, legal domain,
then fine tuning really works.
It's always the move, like, the, you know, domain specialized model, universal model,
and, you know, the kind of this tension between both of them.
I think it definitely, like, makes sense.
It also makes sense, like, to fine-tuning can also be, like, an alternative to, to rack, yeah.
Yeah, well, there's some, there's some companies that are set up entirely just to do that for people.
So it's interesting that, I mean, I sort of view Rika as, like,
like not working in that space, but you could potentially offer that if you wanted to want it to.
Okay, I was going to ask about efficiency and scaling.
I'll just mention this briefly.
And then we can talk about MOUEs because I discovered that you will rewrote your co-author
on the sparse upcycling paper, which is.
Oh, no, I was just advising on that.
Oh, okay.
Yeah, yeah.
But you can talk about sparse upcycling.
It's a topic that's hot.
But more generally, efficiency in my mind, when I go to ICA, I go to New York, I see
efficiency paper, 90% of the chance, I'm just going to ignore it.
because I don't know if it's going to work.
And I think this is related to some of your scaling work and your inducted.
Oh, okay, scaling law was an induct.
Which is like, okay, there was this, to your taxes.
I don't know who this person is on Twitter.
He keeps talking about me.
It's fucking amazing.
Oh, yeah, he does have some obsessions, but like he's good.
I don't know who he is, but he's good.
So he says if 2024 papers are you be trusted, you don't need most attention,
you don't need high precision, you don't need most KV cash,
you don't need most fee-for network layers.
You don't need a reward model.
blah blah.
Like, it's like a lot of efficiency papers are just like, hey, on this like small
example, we cut this thing out, works fine or works great, works better, whatever.
And then it doesn't scale, right?
Like, or, so it's a very interesting observation where like most efficiency work is just
busy work or like it's work in a small scale that doesn't, that just ignores the fact
that like this thing doesn't scale because you haven't scaled it.
It's just fine for grad student.
But as for someone who's trying to figure out what to pay attention to, it's very difficult
to figure out what is a worthwhile direction in efficiency.
Yeah, that's a good point.
I think there's a couple.
I agree with you fundamentally that, like, it's actually quite easy to tell.
Like, when you see a paper, okay, this one doesn't work, this one works, this one doesn't work.
I guess the HIPO account will just tell you that.
Sometimes it's not entirely about this thing doesn't work, this thing works everything.
Right.
Sometimes it's not, you can always find a task in the data set where your efficiency method gets neutral results, right?
You can always find one thing that has, okay, I have comparable complexity.
And you know what's the most, the cutest thing ever?
every time some people propose like this,
they run like some zero short score
on like some LMEval harness or something
like that. And at 1B scale, all the numbers
are random basically.
Like all your Buckeal,
they're all like random chance performance, right?
And they'll be like, okay, I get like 50 versus 54,
I'm better.
But like dude, that's all random chance, right?
Sometimes I see papers that we run experiments
at like, and then it's right.
That's a good tell.
I think it's very, like the sad truth is that like
it's very hard to tell until you scale out.
And sometimes the benchmarks that we have
don't even probe entirely about what,
I mean, especially all the works about,
the transformer alternatives, right?
You can always find like this alternative
that at 7B scale, at 1, 3B scale,
you kind of like, okay, I met transformer
on this and this, this, this, right?
But then what's the implications when you go to like 200B?
What's the implications when you go to 100B?
No one knows that, right?
So that's one thing, right?
And I think developing your own intuition of like what works
and what doesn't work is important.
For example, if somebody's like,
okay, to be honest, all researchers,
like sometimes are also like guilty of this sometimes
because you cannot test on like everything.
I cannot test on everything, right?
So sometimes you also just want to show your method works on this.
But it depends on the objective.
If the objective is to write a paper to XML,
sure, you can find two datasets.
Your stuff works, right?
But when you get adopted, I am not sure.
Yeah, researcher meta game is one thing,
but as a consumer of research,
I'm also trying to figure out
how do I know what is
what is a useful direction
that that's the interesting thing
so for example
MEOEs seem to have worked out
I'll go so far to say
it's the first form of sparsity that worked
because there's so much sparsity research
like we can chop all these parameters
and look we still still perform the same
but then it never actually works
but MOE is really
Oh you mean like the pruning line of work
pruning line of work sorry
I should have used that word
So I don't know if you have any commentary on like
Deep Seek, Snowflake, Kwen,
all these proliferation of MEOs,
MEOE models that seems to all be sparse
upcycle because you were
advisor on the sparse upcycling paper.
So the sparse upcycling paper was
mostly vision focused with a little bit of
T5 experiment. So it was
early stage of like sparse
upfacking, but it was good that Google was really
thinking about this long ago. And Nome also had
paper on it, right? Yeah. I think only
is the way to go. Is it like
100 experts, or 1,000 experts? For some reason
the community settled on 8?
You probably get more gains from more than 8, I think.
But I think in general, it's like MOUs are just a trade-off with like
prime and and flop, right?
And then you're able to make like you kind of make that that scaling law increase
from that additional.
So you can keep a low flop but kind of have more parameters.
It's just changing the flop ratio.
Keeping in mind, there's a lot of inefficiency between the experts.
Yeah, yeah.
I think as an architecture itself, the flop program ratio makes it, like, worth it, right?
But I think the thing that's not very well understood is that, like, how does M-O-E, like, for me, as a research question, is that, like, when you, like, how does it, like, relate to capabilities and stuff like that?
Like, does this inductive bias actually, for example, when you do, like, massive instruction tune, I think there was this paper, like, flood M-O-E or something.
Like, they show that, like, instruction tuning.
I'm not, like, fully sure, I don't recall fully, but, like, when you do massive instruction tuning, like, MOE models are.
like they behave differently from from dance models and stuff like that.
Like I think, okay, like fundamentally I just think the MOEs are just like
the way to go in terms of like flop parameters.
They show they bring the benefit from the scaling curve.
If you do it right, if you bring the benefit from the scaling curve, right.
And then that's the performance per flop argument, activated programs, whatever.
That's like, that's a way to slightly cheat the scaling law a little bit, right?
By having more parameters, right?
I think the more interesting thing is about like what tradeoffs do you make in terms of
capabilities because of this new architecture.
I think that's actually like the
question that I think I guess all the
frontier labs are. They already know this.
Nobody's writing papers anymore about this.
So like you just have to live with what's outside.
But I think I'm bullish about MOUEs.
Yeah.
I had to, I made an exercise for myself
on rating research directions
and what their asymptotic value is.
And I put MOUEs pretty low
because I think you have a good base model
and then you upcycle it.
and it bumps you a little bit.
And I think that's it.
But, like, I'm always seeking to invalidate my hypothesis, right?
But, but, like, from scratch, MOUE is also promising, right?
From scratch, MOUEs promising.
I think in the I do MOU case, you do MOU from scratch.
Yeah.
Okay.
The last part that makes me uncomfortable about MOUE debate is actually related to another
paper that you wrote about the efficiency misnomer.
In a sense that now people are trying to make the debate all about the active parameters
rather than total parameters.
But it seems like it sounds like that's something that you're comfortable with,
like flops at inference is a relevant metric and it's not that.
Well, thanks for like actually reading or like reading the papers.
I'm trying man.
Thanks for.
It's very hard to cut.
It's very hard.
You have a lot of papers.
Well, I actually very impressed that like, oh, you're bringing up these papers.
Yeah, I'm using attention.
Okay.
Okay.
Yeah, thanks.
Thanks.
And also, I mean, I'm interested in efficiency that works.
It's just very hard to find efficiency that works.
And so like anything that helps me have high signal on efficiency is helpful.
So I think for the inefficiency, misnomer by the way,
I love the paper, by the way,
it's had a fun time working on it.
I think efficiency misnormal was like,
we found that a lot of people,
like, they use params,
like especially to do kind of like,
right,
and then MOEs was not very hot
like in the community at that time, right?
But MOEs were like a thing long ago at Google, right?
So I think using active params,
I'm comfortable with using active brands
to kind of approximate like cost on the model.
But like in the efficiency mismanormal paper,
we actually made it quite clear that you should always like look holistically
about like,
like, because you have serving like additional serving costs.
Yeah.
Like fitting in the GPUs, like fitting on single node and something like that.
Interesting one was speed.
Nobody really talks about speed.
But your paper actually tried on.
Okay.
I have something to say about speed.
There are so many methods, right, that are proposed about efficiency, right?
They are like theoretically like faster because of like complexity, like something like that.
But because there's no way to work around the implementation or like your implementation becomes so hard.
It becomes like 10x slower.
Okay.
There's so many papers.
It's not a hard way or where.
Like, it could be hard.
It might not be, it could be hard where it could be just the way that, like, you have a convenient way to, like, in its mathematical form, it's actually like, okay, linear complexity, like, whatever.
And it's actually theoretically faster.
But, like, just because you have to, like, do a scan or something like that.
And then it becomes, like, actually, like, 10 times slower in practice, right?
There are a lot of things, like, not a lot, but, like, there are some things that are, like, some methods that are like, like, like this, where you don't take it into account throughput, right?
which is also the problem of like sometimes
like the incentives of like
people who are working efficiency,
you can easily sell a paper as like more efficient
and then people will not suspect that
because the reason why we wrote the paper
is that so many people were confused
about like efficiency itself, right?
Yes.
And then they will be like, okay,
like a lot of these unsuspecting reviewers,
especially like even academics or they don't have like that
that real real feeling they were less like,
okay, less parameters, more efficient, right?
So you could have a method that's like less parameters
but like three times slower.
because a lot of times when you add things to the model,
it becomes slow.
You add complexity,
especially if it's like something that's not
hard where optimized,
no kerners or like something that is like
bad for TPUs or whatever.
Your model just become like slow.
That's a temporary issue.
People can fix it.
But some things are not like so.
Some things may not be like so easily fixed or like
it just adds a lot of like three costs to,
to optimize it and everything.
Right.
But then it's always marketed as like because I save brands.
So I save.
Right.
And then the brands will add a different place of the model.
like, for example, like, if let's say you, even in the case where you brand match models, right,
if I take out like some brands from like FFN, right, and I put it to like embedding layer,
right?
It's a cheap operation for abetting layer, right?
But my model becomes like lopsided, right?
I could say I brand match this, but it's not true put match, right?
Yeah.
Because it's unbalanced on the side.
It's unbalanced on the side, right?
So there's also of this type of tricky things that like when mixed commons,
all the comparisons, like very, very, very, very difficult.
And because you cannot even put like flop, throughput and speed,
flop, parliaments and speed, like extra speed, right, in the same plot, right?
And then there's always like one money shot in a, like,
there's always a like a parietal kind of compute, like whatever plot, right?
Like for marketing in papers or something like that,
it's always very easy to like, I mean, not intentionally, but like to subconsciously, like,
show one story when it's actually like there's like all these other things,
to consider.
Yeah.
It's a selection bias, self-biased, whatever.
Very cool.
Okay.
Well, that was mostly of most of the technical side.
We have one commentary that will happen today on the future of open source models.
It basically founders fund said like the future is closed source.
You were agreeing with it.
And a lot of the open source fanatics are up in arms over this.
I don't know if you care to comment about just open versus close and close whatever.
I mean, I don't really like, when I mean like, if you're referring to the tweet that I wrote,
but I wrote something about...
But this is huge.
Like, so many people are commenting about it
because they are personally physically offended
that open source cannot catch up.
Okay, wait, okay.
So I want to say,
it's like, I'm not...
Like, I contributed to open source in the past,
so I'm not, like, against, like, open source per se.
But the interesting thing that I want to talk about here
is that, like, there's a difference between...
Like, I draw a line with, like, open source as in,
like, okay, to me, Lama tree is like...
It's like, matter has an org that is like,
okay, hypothetically very similar to
to like gem night or something
but they just decide to release the weights, right?
Yeah, it's open weights.
Right, it's open, it's open weights everything, right?
I think when most people try to say that, like,
open source is catching up everything.
They kind of mean like this grassroots, like,
yeah, this, this bottom up people that are like,
these indie developers that are like coming together to fight,
like, it's romanticized and it's dramatized to some extent
to fight against like this.
It definitely is.
Right.
And to be very fair, I think that there isn't really much like,
like so far, if you just look at,
the fractions of people, the big labs are just pushing and pushing and pushing.
The academics, like Stanford and stuff, they came out with DPO, they came out with things
like that.
They make some, like, but they're kind of in between the line of like open source community.
And then there's also like the developers that are like fine tuning on GPT4 distilled models
and everything, right?
I don't, I think the open source, the underlying like thing about like collectively improving
something.
I'm not like criticizing it for the sake of criticizing it, but like, I'm just saying,
saying that like in order to make progress, right?
I think the incentives of open source are like,
what I observe is that like people like to do things like,
they like to take somebody else model, they rename it,
they make a quick win from there.
And then like you notice that like when people realize that like
this starting on the GPD4 tab and running some DPO,
it's not going to give them the reward signal that they want anymore.
Right.
Then all these variants gone, right?
You know, there was this era where there's, wow,
there's so many of this like,
I lost track of all these model variants,
but now they're all gone because people realize that
that you cannot climb LMSIS
because you need something more than
that's something that is lightweight, right?
So I think that was just my overall, like...
Honestly, the Hanging Face leaderboard contributed to most of that.
It's not LMS.
No, no, I think LLSI is probably they realized that they could not, yeah, right?
The open LM leaderboard is probably like a big problem, to be honest.
We're talking to Clementine in one of our future episodes.
Okay, okay, okay.
They dedicate a lot of, I mean, there's so much attention to them.
It's a tough problem, but they're providing a public service for sure.
Yeah, yeah, yeah.
I mean, good intentions are always good.
I mean, like, good intentions are always good.
Yeah.
Is any, any, take that in any order that you wish.
I don't have much of a life, actually.
But I'm trying more to have more.
I mean, you're a father now.
I have a baby now.
So like, I'm trying more to have more life and everything like this.
I think the productivity hack that I have is just like, I didn't have like a boundary
between my life and my work like for a long time.
So I think I just cat a lot about working most of the time.
Actually for the last like, through my PhD, at Google and everything, I'll be just like working
all the time.
It's not like the most healthy thing.
ever. But I think that was actually one of the biggest
productivity. And I like to spend a lot of time
writing code and I just enjoy running experiments, writing code
and stuff like that. Right. So you kind of, if you enjoy something, it's not work
right? So like, it's very strange. It's like, I would get distracted by
sometimes I have to watch some Netflix series because like my wife asked me to
like watch it. Or somebody tells me that I'm back on time on some shows, right?
But then I get distracted by my experiments running and I just end up like
like writing code instead of like
so things like like like this
it's not the most healthy thing but
I think that's one
looking for like a practice where like
okay so Andre recently had a thing where like
before when he wakes up he doesn't look at social
media he only goes trick to work
damn I checked Twitter the moment every
I know it's just something I do as well but I'm like
that's a smart rule and like I'm looking for
like rules like that like do you have a rule
no he doesn't check social media because his phone is exploding
all the time all the time yeah I don't have so many
likes and followers so like it's fine for me
yeah you get there
Like rules like that, mantras that you've developed for yourself where you're like, okay, I must do this.
So for example, recently for me, I've been trying to run my life on calendar for a long time.
And I found that the only way that I work is I write things down on pen and paper and I cross them off individually.
And like that physical action really, really helps me get things sorted.
And that's work-wise.
Reading-wise, I don't know if you know, but I've been running this like AI newsletter.
Like all those summarizes, all Twitter, Reddit this points and all that.
So that helps me keep up because I have like a socially graded.
and I personally vetted the entire pipeline from beginning to end.
So, like, this is my input algorithm.
I know how to keep up with news
because I now have a information condenser.
So, like, I'm trying to figure out what's your algorithm
or what's your rules for keeping up.
I got something, I got something for keeping up.
So I used to check archive, like, every morning when the gate opens,
I just check archive.
I will wake up 9.30 a.m. Singapore time the archive gate opens, right?
And then I'll be very sad if there's no papers to read.
But you usually just pick one paper or two papers that you find interesting.
I don't read them.
I just like skim like the thing, right?
Yeah.
So I used to that.
I don't do that anymore.
I mean,
ever since like I have in the startup, I read, right?
Yeah, you have a real job now.
I read, I read.
I read, I read less papers, right?
But I used to camp at the door of archive quite frequently just to see.
Isn't that?
That's not a good use of time.
I'll come on and say it.
It's not a good use of time.
No, no.
It's a newness bias.
Sorry, go ahead.
No, no.
It's just because like, I ran out of things.
I say, yeah.
It's just that like the new stuff comes out, right?
Yeah.
Like, and then like the new stuff come out, so that's how I keep up to date.
So in the space of three years, you read every...
No, no, I didn't read everything.
It's just that.
It's just that.
But these days, I realize I don't have to do that anymore just because if the paper is important enough,
Twitter will show you to me.
Sure.
So I, there isn't really, like...
And one thing I do is that I actually don't read papers like that much anymore.
I just like skimed them like almost, right?
So that's for keeping up, like, with papers, research and everything.
And the other thing more of like this, like just like,
a productivity point of view is that I used to always keep like the text.
Like I usually start writing the thing while working on that thing itself.
Like so even like let's say if you want to launch something like the angle is like a blog post
or shipping something and everything, right?
I like not really a launch or say all that's papers.
I always like to look at it from like what's the story and the end.
And then I just like figure out what I need to do to get to kind of right.
So I think as a researcher like this is something like I would have like.
like so many drafts of like
when I start the project
I don't know the experiment
yet everything right
but I like to imagine like
what the title would be
right and then I always vibe check
but like I always like so I mean
my friends at Google
would know that I always have like
like the overlay draft of like
so many
and then I'll just spend time looking at it
like to looking at it like
took the title is it better two seconds
like I can't about
I used to care about a lot of things
but this actually helped
like product because every time I look at it
I'm like okay this is the final product
I'm like working towards it right
because I think a lot of researchers
they tend to like
they sew around
in the experiments and they never like ship the final story.
It's like the shipping like I mean it started out with ship products but like as a researcher
your product management.
Yeah.
You're shipping the thing.
So I like to I like to hang around a lot in my in my draft and I get motivated from
that and that's like one productivity thing that I did as a researcher.
Yeah.
So I think that other than that I don't really have any things that I do that probably different
from others.
Probably you don't know it.
This is unconscious competence versus.
Okay.
Was it like just NTIU PhD?
Just the story of like
How was it coming out from NTU
Which is like a good school
But like not not typical target school
For like a big lab
I did my PhD unknowingly
Like I didn't have very
Like when I was a very regular undergrad
I had decent grades
But not the best grades
I was not like super smart in school
Or something like that
I was I wanted to do a PhD
That's because I was like curious
And I mean like and then
I wanted to stay in Singapore at a time
So I just like naturally just did a PhD there.
I didn't even vet my advisor.
I didn't even think too much.
I just like fell into the PhD program.
And then there was when I realized that,
oh, actually I can do research.
Like I'm like pretty decent at research.
Like I just fell into a PhD like unknowingly.
Yeah.
And I definitely like,
NTU leaves a lot to be desired.
Actually, to be honest,
I think that I mean,
Singapore leaves a lot to be desired in general.
Like the research community here is like like probably not great.
So how did you like break out?
If I was you,
I would have no idea.
how to break onto the international scene.
I think it was, okay, to be honest,
like in retrospect, it's a bit of, like, a bit of a miracle.
Or, like, I mean, it's not easy to,
I think I could not, if I had, like, a product,
like, someone to mentor, like,
I could not, like, tell somebody how to replicate the same thing that I did.
It's much easier now, maybe compared to in the past,
but I've been mostly self-supervised during my PhD.
Like, my advisor was basically, like, like,
Gramerly.
Like a free paid plan of Gramerly.
You won't watch this, so it's fine.
But, like, there's a lot.
things that it was like this strange art of my life where I was figuring out research by myself
and everything and okay maybe going back to the change your opinion is that like the biggest
culture shock I had like when like was moving from Singapore PhD to Google I think my research
like taste you went straight to Mountaineville yeah I went to Mountedville I started at Mountain View like
my research tastes and everything like like I was it was so different like the research culture
is so different in in US and in Asia I had to grow so much like during my time at Google too
actually evolve.
And then whenever I come back, right,
I still have friends in, like, faculty in here and everything.
I either think that I'm a snob or they think that I'm like being a like a very nasty person.
Because I think to be honest, the research here is like in Singapore.
It's just basically like they just care about publishing papers and stuff like that.
And then it's not impact driven.
I think at US is mostly focused on impact driven.
And the thing needs to make real impact, right?
And what, to be fair, you're also working at an industrial lab versus
an academic circle, right?
Like, you're comparing apples and oranges here a little bit.
I mean, at the end of the thing, I think research is like fundamentally, like,
as an industry, RIS, you still write papers.
Your goal is to advance science and everything.
To be honest, it's all the, the incentives, rewards the system is like different
and maybe like slightly different and everything.
But, like, at the end of the day, I still feel that researchers are researchers,
scientists or scientists, no matter, like, really, like, where you are.
I will get so much dissonance when I come back and I talk to people.
I would feel like, oh, why do you think like this?
But then I used to think like this.
So like the environment shapes the way a researcher thinks.
The taste is very important.
Sometimes I try to communicate this to people.
And then maybe I come across as a snob to, to like the local community here.
Right.
But like it's just that there's like maybe there's so much dense information that I want to bring back.
But like, there's no like fast way to like transfer like all the like transfer all the things that I've learned.
also a big cultural shock because I was in
brain in the Singapore office for a while
and I'm reporting to... You're the only brain in person.
Yeah, yeah, in Britain in Singapore. And then I had like, I took
on an intern from NU.S. actually. And the research
like vibes and the thing was
so much of a conflict for me
that it was almost like my body was rejecting it, you know?
But this person so like grew and became, I'm happy with how this
person grew from my mentorship. So he's now
in a way better situation. But I
say that like a lot of people in the in universities here like not like a bit like like
ignorance is bliss right maybe sometimes well no it's exposure i didn't know any better myself
until i went to to the u.s for for college and then yeah my world was expanded and it's a little
bit of a pandora's box because once you've tasted that you're never happy yeah yeah you know so
okay last question would be just just a sort of Singapore question so i i like to be visible visibly
non-American covering the AI scene because it's very US-centric. Every non-American I talk to always
wants to be like, how can we build Silicon Valley in my city? You know, my country, my city, whatever.
That is not Silicon Valley. I feel like you have basically just kind of like me. You kind of operate
in the US circles, but you just don't live there. Do you have any advice for like if Singapore,
okay, so I'm wearing a register today. This is the official Singapore government's sort of community
group that is that is trying to guide Singapore AI policy. If we want a hundred more you
days to come out. What should governments be doing? What should communities, ecosystems should be doing?
So I actually think that like sometimes like not doing too as like too much is maybe less is
small, maybe. I don't think there's actually much like the government can do to like influence
like this kind of thing is like a natural, like an organic, like an organic natural thing. Right.
The worst thing to do is probably like to create like create a lot of artificial things that like exchange
programs. Okay. I mean Singapore used to have a lot of exchange programs. Like it's
send people to, I mean, just talking about AI specifically, right?
I think that, for example, like, sometimes, like, trying to do too much
or, like, moving in the right, wrong direction is just better than not moving at all.
Especially if you, if you accelerate in the wrong direction, you actually get into a worse
than possible, right?
So I think it's very dangerous to, like, move in a bad, like, direction.
I think respect your talent more, maybe.
The government should just respect the talent more.
And I don't know whether this is too much of a...
No, no, no, no.
But not, but not, maybe not moving in a wrong direction is,
To me, it's already a very good thing.
Funding for startups, incubation,
holding academic conferences,
I clear next year it's going to be in Singapore,
so people come here and expose to it.
But, like, I don't know.
This is just very interesting.
Like, everyone wants to build up AI expertise
within their own country,
and, like, there's a massive range into the US.
I'm part of that.
Like, I live there.
I feel guilty.
I don't see any other way around it.
It's such a huge problem.
I also do think that there is, like,
cultural hegemony.
it's called it, like US values
basically being asserted on the whole world,
right? Because we decide RLHF
on these models and now
you shall use all our models. And it's
just troubling for, like,
national sovereignty should be AI sovereignty, and
I don't know how to achieve it for people.
It's very scary. Okay, there's a lot
to unpack. Yeah, this is
not technical, but I was just curious.
We can make this the ending conversation, which is, I think
your inspiration to a lot of other
people who want to follow your career path.
I'm really glad that we got the chance
to go through your career a bit.
Yeah, I'm sure this is just the start.
So hopefully there's more to come.
And I want to inspire more of you.
Yeah, yeah, sounds good.
So I'm just glad that you shared it with us today.
As a special coda to this conversation,
we were invited to join the Tekkenasia meetup featuring Yi by managing editor Terence Lee.
Terence asked a similar question on how other countries can create conditions for top AI labs
to spring up outside of Silicon Valley.
So, like, where do you see Singapore playing a role in?
in AI. So how would you?
Oh, okay, right. I got a practical one.
Okay. I got a practical one that is actually actionable.
I feel like one thing that people don't get, the advice that, practical advice, like,
that is that like, the era of like people who talk versus people who do, like, the people
who talk is like gone, right? So like it's no longer about like, I have a team, I have like
10 interns from Southeast Asia or like the region and then they're going to do this, do this,
do this, do this for me, right? So I think one thing that, that, that, that, so, I think one thing that, that,
senior people in any government may not get, right, is that the world has shift into this
paradigm where senior ISIS, ISIS as individual contributors, right, are actually making the most
impact in AI, right? So in GDM and in Open AI, I mean, Frontier Labs, they're all very driven
by individual contributors and not, actually this is not even related to, this is, I'm talking about,
like, this is the advice I give, but it's actually general, like, so many general thing.
So multi-purpose basically.
It's not AI-specific.
No, it's also, it's very AI-specific because the level, the difficulty of making impact
and making breakthrough has started to become like, it's no longer about like, it's not
like software engineering where, where it's, I think AI is a little bit harder, like and
then like it's mostly about like getting very senior people who are hands-on and have a lot
of experience rather than like management style people that like try to like
think they know what they're doing but they actually don't so I think I mean I I'm
not going to like say like name obviously right but like I mean I meet a lot of
like people like this or like in general I mean not only in Singapore but like
right but AI has shifted quite a lot into this I see driven paradigm where the
people making impact are the people who are like
on the ground fighting the war, right?
So it's no longer about, I have 10 interns, 20 interns, 100 interns,
you do this, you do this, you do this, I just take meetings, right?
No, right?
The senior person writes code, everybody writes code, nobody should not write code, right?
And then everybody, so I think this is, okay, this is a big extreme, but, but, but,
this is a bit on the extreme side, but I think from people, like, I just,
the advice is just like, maybe like, just take 20% of what I say, and incorporate
So instead of like if you if you if you for example hypothetical hypothetical situation right say you want you want to organize like an AI conference in Singapore right and then you want to make it like you want to show Singapore as like the AI hub in the world right
I mean you don't invite like policy people and like you know if invite like policy people to come and talk about AI safety AI safety right you invite people who like actually know the stuff right and then if you organize a conference and then like 100 people like
go there and then they feel very productive and everything but like the problem is that like
Singapore doesn't have people who really can do it you know right so I mean I've
through the great vines I mean I hear about people like fighting for territory here and there
I mean this is what I hear I I don't want to hear this but I hear this somehow right
and then like and then sometimes I just ask them like who's actually going to do it right
who's going to do it right the work is the model is not going to train itself right unless we have AGI right
So yeah, I mean, understand that like times have changed, it's no longer about like, it's no longer about like, oh, I'm very senior, very senior, very senior, okay, okay, okay, can you quote, right, that's the question, right?
I think that's like the, yeah, yeah.
Well said, spicy, spicy already.
Okay, okay.
Okay, I think we've hit the number.
You're like cocoa in baya, raise the cocoa egypt area to the maximum.
Yeah, almost there really.
Okay, questions, anyone?
Indeed, questions are very welcome.
Head over to the latent space substack to leave a question or tweet at Yitai ML or at AGIHippo directly with your feedback.
