Y Combinator Startup Podcast - #45 - Building Dota Bots That Beat Pros - OpenAI's Greg Brockman, Szymon Sidor, and Sam Altman
Episode Date: November 8, 2017Greg Brockman is the CTO and cofounder of OpenAI.Szymon Sidor is a Research Scientist at OpenAI.Sam Altman is the President of Y Combinator and Co-Chairman of OpenAI.Watch their bot compete at The Int...ernational.
Transcript
Discussion (0)
Hey, this is Craig Cannon and you're listening to Y Combinators podcast.
Today's guests are Greg Brockman, Shimon Sador, and Sam Allman.
Greg is the CTO and co-founder of OpenAI.
Shimon is also at OpenAI.
He's a research scientist there.
And before we get going, if you haven't yet subscribed or reviewed the podcast yet,
that would be awesome if you did.
All right, here we go.
Now, if you look forward to what's going to happen over upcoming years,
is the hardware for these applications for running neural nets really, really quickly,
are going to get fast, faster than people expect.
And I think that what that's going to unlock is you're going to be able to scale up these models
and you're going to see qualitatively different behaviors from what you've seen so far.
At Open AI, we see this sometimes.
For example, we had a paper on this unsupervised learning where you train a language model to,
you train a model to predict the next character in Amazon reviews.
And just by learning to predict the next character in Amazon reviews,
somehow it learned a state-of-the-art sentiment analysis classifier.
And so it's kind of crazy if you think about it, right?
You just were told, hey, predict the next character.
And, you know, if you were told to do this,
well, the first thing you'd do is you'd learn the spelling of words and you learned punctuation.
The next thing you do is you start to learn semantics, right, if you have extra capacity there.
And that this effect goes away if you use a slightly smaller model.
And what happens if you have a slightly larger model?
Well, we don't know because we can't run those models yet.
But in upcoming years, we'll be able to.
what do you guys think are the most promising under explored areas in AI if we're trying to make it come faster
what should people be working on that they're not yeah so there are many areas of AI that we already
develop by quite a bit there's some some basic researching just classification deep learning and in
reinforcement learning and what people do is people kind of try to invent problems and the
such as solving some complicated games of hierarchical structure and they try to add kind of extra
features to their models to combat those problems.
But I think there is very little research happening on actually understanding the existing
methods and their limits.
So for example, it was a long-head belief in deep learning that kind of to parallelize the
to paralyze your computation, you need to cram as small batches as possible on every device.
And in fact, by due this impressive engineering feed where they took recurrent neural networks
and they implemented the kind of GPU assembly code to make sure that you can fit like
bat size 1 at an end on every GPU.
And, you know, despite like all those smart people working on those problems,
like only very recently that Facebook kind of took a cold, quiet look at just like very
basic problem of classification. And they, in their great paper called Image That in one hour,
they showed that if you actually take a code that does image classification and if you fix
all the bugs, you actually can get away with much larger batch size and therefore finish the
classification probably much faster. And, you know, it's not the kind of sexy research that
people want to see where you have like some hierarchy,
how big at an end, but it actually, this kind of research,
I think at this point will advance field the most.
So Greg, you mentioned hardware in your initial answer.
In the near term, what are the actual innovations that you foresee happening?
So the big change is that the kinds of computers that we've been trying to make really
fast are general purpose computers that are built on the bond no way in architecture.
You know, you basically have a processor, you have a big memory, and you have you have some some
bottleneck between the two.
With the applications that we're starting to do now, suddenly you can start making use of massively
parallel compute that the architectures that these models can run on sort of the fastest are going
to look kind of like the brain, where, you know, the brain is basically you have a bunch of
of neurons that all have their own memory right near to them.
and that they all talk to their neighbors,
and maybe there's some kind of longer-range skip connections,
and that just no one's really had incentive to develop hardware like this.
And so what we've seen is that, well, you move your neural networks from running on a CPU to a GPU,
and now suddenly you have 1,000 kuda cores running in parallel,
and that you can get massive performance boosts there.
Now if you move to specialized hardware that is sort of much more brain-like
and that runs a bunch of, you know, they sort of runs in parallel,
with a bunch of tiny little cores,
that you're going to be able to run these models
sort of insanely faster.
Okay.
So I think one of the most common questions
or threads of question
that were asked on Twitter and Facebook
were generally how to get into AI.
Could you guys give us just a primer
of where someone should start
if they're just a CS major in college?
Yeah, absolutely.
So it really depends on the nature of the project
that you would like to do.
I can tell you a bit about our project, which is essentially developing large-scale reinforcement learning for Dota 2.
And their majority of the work is actually engineering.
And, you know, like essentially taking the algorithms that we have already implemented and trying to scale them up is usually the, it's usually the,
the fastest way to get
improvement in
our experiments.
So essentially
becoming a good engineer
for our team is much more valuable
than, for example, people spending
months upon months
implementing exotic models
in TensorFlow. So just to echo this, because
I hear this come up all the time, people say it's like my dream
to work at OpenAI, but I got to go get an AI PhD,
so I'll see you in like five or seven years.
If people are just a really solid engineer but no experience at all with AI, how long does it take someone like that to become productive for the kind of work at Open AI that we're looking for?
So someone like that can actually become productive from day one.
And with different engineers who show up in Open AI is that there's a spectrum of where they end up specializing.
There are some people who focus on building out infrastructure and that actually this infrastructure can range from, well, we have a big Kubernetes deployment.
that we run on top of cloud platform and building tooling and monitoring and sort of managing
this underlying layer.
It actually looks quite a bit like running a startup and that a lot of the people who are most
successful at that have quite a bit of running at large scale in a startup or production environment.
There's kind of a next level of getting closer to the actual machine learning, where if you think
of how machine learning systems look, that they tend to be this like magical
black box of machine learning, this core.
And you actually try to make that core be as small as possible because machine learning is
really hard, eats a lot of compute, it's really hard to tell what's going on there.
And so you want it to be as simple as possible.
But then you surround it by as much engineering as you possibly can.
So what percent of the work on the Dota 2 project would you guys say was what people
would really think of as like machine learning science versus engineering?
Essentially, as far as day-to-day work goes, this kind of work was almost not.
existence, there was like maybe few person weeks spent on that compared to like personal
spent engineering.
And I think maybe placing some good bets was one of it.
Good bets on the machine learning side?
On the machine learning side.
Yeah.
And they're often about more about what not to do rather than what to do.
So at the very beginning of the project, we knew we wanted to solve a game, a hard game.
We didn't know exactly which one we wanted to do because these are great test beds for pushing the limits of our algorithms.
And one of the great things about it, too.
And just to be clear, you guys are two of the key people, the entire team was like 10 people.
About 10 people.
That's right.
And these things, good test beds for our algorithms, see what the limits are to really push the limit of what's possible.
And you know for sure that when you've done it, that you've done it.
It's a very binary testable.
And so actually the way that we selected the game was we went on Twitch.
and just look down the list of most popular games in the world.
And starting, you know, number one is League of Legends.
The thing about League of Legends is it doesn't run on Linux,
and it doesn't have a game API.
And little things like that actually are the biggest barrier
to making AI progress in a lot of ways.
And so looking down the list,
Dota actually was the first one that kind of had all the right properties,
runs on Linux,
that it has a big community around replay parsing,
that there's a built-in Lua API.
It was actually this API was meant
for building mods rather than for building bots.
And we were like, but we could probably use it to build bots.
And the, you know, one of the great things about Valve as a company is that they're very
into having these open hackable games where people can go and do a bunch of custom things.
And so kind of philosophically, it was very much the right kind of company to be working,
to be working with.
So the, we actually did this initial selection back in November.
And we're working in some other projects at the time.
really get started until late in December.
And one of the funny things is that by total coincidence, in mid-December, Valve released a new
bot-focused API, and that they were saying, hey, our bots are famously bad, that maybe
the community can solve this problem, so we'll actually build an API specific for people
to do this.
And that was just one of those coincidences of the universe that just worked out extremely well.
So we were kind of in close contact with the developer of this API.
and kind of all throughout.
So at the very beginning of the project, well, what are you going to do?
Right.
So the first thing was we had to become very familiar with this game API to make sure we understood all the little semantics
and all of the different corner cases and also to make sure that we can run this thing on large scale
and to turn it into a pleasant development environment.
And so at the time it was just two of us.
One person was working with the bot API building a scripted bot.
And so basically this is the learn all the game rules.
rules think really hard about how it works.
The particular person who wrote it, Raffle, has played about three or four games of Dota
in his life, but he's watched over a thousand hours of Dota gameplay and has now written
the best Dota scripted bot in the world.
And so that, you know, sort of a lot of just writing this thing in Lua, getting very intimately
familiar with all those details.
In the meanwhile, what I was working on was trying to figure out how do you turn
this thing into a Docker container.
And so they had this whole build process.
It turns out that Steam can only be an offline mode for two weeks at a time, that they
pushed new patches all the time so you needed to go from this like, you know, sort of
manually download the game and whatever to actually have an automated, repeatable process.
It turns out that it's, the full game files are about 17 gigabytes and that our Docker registry
can only support five gigabyte layers.
And so I had to write a thing to chunk up things into five gigabyte tarballs and put those in
S3 and start them back down.
So a bunch of sort of like things there where it's really just about figuring out what the right workflow is, what the right abstractions are.
And then the next step was, well, we know we want to be writing our bots in TensorFlow and Python.
How do you get that?
So why was that?
Well, because so machine learning, you know, that it's actually quite interesting that a lot of the highest order bit on progress, just like having the game API is a higherder bit.
It's also can you use tools that are familiar and sort of easy to iterate with before the world of kind of modern machine learning very.
marks ever write their code in Matlab.
If you had a new idea, it would take you two months to do it.
Like, good, good luck making progress.
And so it was really all about iteration speed.
And so if you can get into the Python world, well, we have these large code bases that we
built up of high quality algorithms, that there's just so much tooling built around it,
that that's like the optimal experience.
And so the next step was to port the scripted bot into Python.
And so the way I did that was I literally just renamed all of the dotlua files to dot pi
commented out the code and then started uncommenting function by function.
And then, you know, you run the function, you get an exception.
You then go and uncomment whatever code it depends on.
And as mechanically as possible, I tried to be like a human transpiler.
And, you know, that Lua's one index, Python zero index.
So you have to do that.
That Lua has, doesn't distinguish between an array type and a dictionary type.
And so you kind of have to disemiguate those two.
But for the most part, I did something that could have been like sort of totally mechanically done.
And it's great because I didn't have to understand any of the game logic.
I didn't have to understand anything that was going on under the hood.
I could basically just port over and it just kind of came together.
But then you end up with a small set of functions that you do not have implementations of,
which are all of the actual API calls.
And so I ended up with a file with a bunch of dummy calls,
and I knew exactly which calls I needed,
and then implemented on top of GRPC a protobuf-based protocol
where on every tick the game would dump the full game state,
send the diff over the wire,
reassemble that into an in-memory state object, Python.
And then all of these API methods would be implemented in Python.
And so at the end of this,
it sounds like a bit of a Frankenstein process,
but it actually worked really well.
And in the end, we had something that looked just like
a typical open AI gym environment.
And so all you have to do is you say,
gym.
Dotta environment ID,
and suddenly you're playing Dota.
And that your Python code just has to call
into some object that implements the Lua API,
and suddenly these characters are running around the screen
doing what you want.
And so this was like a lot of the kind of thing
that I was working on in the peer engineering side.
And actually, as time went on,
so Shimon and Yakup and Jay and others joined the project,
and most people were building on top of this API
and really didn't have to dig into any of the underlying implementation details.
So personally, my one machine learning contribution to the project, I'll tell you about that.
Because, you know, my background is primarily startup engineering, building large infrastructure,
not sort of machine learning, definitely not a machine learning PhD.
I didn't even finish college.
So I kind of reached a point where I got in the infrastructure into a pretty stable point
that I felt like, all right, like I don't have to be fighting the fires there very constantly.
I have some time to actually focus on digging to some of the machine learning.
One particular piece that we were interested in doing was behavioral cloning.
So we had one of the systems that we had built was to go and download all of the replays that are publish each day.
And so the way this game works is that there's about 1.5 million replays that are available for public download.
Valve clears them out after two weeks.
And so you have to have some discovery process.
You have to stick them in S3 somewhere.
Originally we were downloading all of them every day and realized that was about two terabytes worth of data a day.
That adds up quite quickly.
So we ended up filtering down to the sort of most expert players.
But we wanted to actually take this data, parse it, and use it to clone the behavior for a bot.
And so I spent a lot of time with like sort of, you know, it's basically this, you have to need this whole pipeline to download the replays to parse them to, you know, kind of iterate on that, to then take it, train a model and try to try to predict what the behavior would be.
And, you know, first it's just like, like one thing I find very interesting is the sort of different.
workflow that you end up with when doing machine learning.
Like there are a bunch of things where when software engineers join OpenAI that are just very
surprising.
For example, if you look at a typical research workflow, you'll see a lot of files named like,
you know, experiment, whatever the name of the experiment is one, two, three, four.
And you look at them and they're just like slight forks of the same thing.
And you're like, is this what version controls for it?
Why are you doing that?
And after doing this cloning project, I learned exactly why.
because the thing is if you have a new idea for okay well I've kind of got this thing working
and now I'm going to try something slightly different as you're doing the new thing well machine
learning is to some extent very binary that at the start it just like doesn't work at all and you don't know why
or it kind of works but it has some weird performance and you're not sure exactly is it a bug is it just how
this data set works like you just don't know and so if you've gotten it working at all and then you make a change
you're always going to want to go back and kind of compare to the previous thing you had running
And so you actually do want the new thing running side by side with the old thing.
And if you're constantly stashing and unstashing and checking out and whatever, then you're just going to be sad.
And there are a lot of like kind of workflow issues like that that you just, you got to bang your head against the wall.
And then you see like, I've been enlightened.
So before we progress further on the story, can you just explain the basics of training a bot in a game?
Like, how are you actually giving it the feedback?
Oh, yeah.
So on a high level, we are using reinforcement.
learning with self-play. So what that means, I mean, it's no rocket science, even though
like reinforcement learning sounds so fancy. Essentially what's happening is we have a bot which
observe some state in the environment and perform some actions based on that state. And based on those
actions that it executes, then it continues playing and eventually, you know, either does well
or poorly.
So that's something that we can quantify in a number,
and that's more of an engineering problem than research problem,
how to quantify how good the bot is doing, right?
You need to come up with some metric.
And then, you know, the boat gets feedback or whether it's doing good or not,
and then tries to select the actions that yield to high,
that positive feedback, to high reward.
And to give us a sense for how well that works,
so the bot plays against itself to get better,
once you had everything working,
how good would a bot from day N do
against a bot from day N minus 1?
So I guess we have a story
that kind of illustrates
what to expect from those techniques.
So when you start this project,
our goal wasn't to really do research.
I mean, at some high level it was,
but we are very goal-oriented.
All we wanted to do is we wanted to solve problem.
Right?
we want to solve.f55 and we had our mice love that I won v1.
And the way it started, it was like early days when there was just Greg and Raphao
and Raphael was implementing scripted bot.
So he just like literally right the logic.
I think this is what bots should do when he sees a creep, he needs to attack it, yada yada.
And he spent like three months of his time.
And Raffo is actually a really good engineer.
So we had a really good scripted bot.
So what happened then, you know, kind of like he got it to the point
I wish he didn't think you could improve it much more.
So we tried, okay, let's try some reinforcement learning.
And, you know, I was actually at vacation on the time.
But there was other engineer, Jakub, who throughout my vacation,
which I found super surprising.
So I leave, there is nothing.
I come back.
There is a reinforcement learning bot.
And actually, it's beating our scripted bot after like a week worth of engineering effort.
Possibly it was two weeks.
But it was something very miniature compared to the development of scripted bot.
So actually our bot, which didn't have any assumptions baked about the game,
figured out the underlying game structure well enough to beat anything that we could call it by hand,
which was pretty amazing to see.
And at what point do you decide to compete in the tournament?
Well, so maybe I should finish up my story.
Sorry if it's running a bit long, but it'll get good shortly.
So just finish up my machine learning contribution.
So basically spent about a month really learning the workflow,
got something that like, you know, was able to do some signs of life where it like run to the
middle and you're like, oh, it knows what it's doing. It's so good. And it's very clear,
like, when you're just doing cloning that like these, these algorithms like learn to imitate
what it sees rather than the actual intent. And so it'd get kind of confusing and like kind
of run, you know, try to do some like creep blocking or something, but like the creeps
wouldn't be around. And so it'd just be like zigzagging back and forth. And anyway, I got this
to the point where it was actually creep blocking pretty reliably pretty well. And, and, and then
at that point I turned it over to Jay, who's also working on the project, and he used reinforcement
learning to fine-tune that. And so suddenly it went from only understanding the actions rather
than the intent to suddenly it really knew what it was doing and kind of has the best creep block
that anyone has seen. And so that was my one machine learning contribution in the project.
So time went on. And one of the most important parts of the project was having a scoreboard.
So we had a metric on the wall, which was the true skill of our best bot. So true skills basically
an evil-like rating that measures the win rate of your bot versus versus others.
And you put that on the wall, and each week people just try all the ideas and some of them work.
Some of them improve the performance.
And that we actually ended up with this very smooth, almost linear curve.
So we posted it in a blog post.
And that that really means kind of like exponential increase in the strength of this bot over time.
And that part of that is, you know, sometimes these data points where you just train the same experiment for longer.
typically our experiments
would last for maybe up to two weeks
but also a lot of those were
while we had a new idea
we tried something else
we made this tweak,
we added this feature,
removed this other component
that wasn't necessary
and so we knew that
so we chose the goal of 1B1
I don't recall exactly when
but it must have been you know
in the spring or maybe even early summer
but we really didn't know
are we actually going to be able to make it
and unlike normally
when you're building an engineering system
you think really hard about
all the components.
It's like, well, you decompose it into this subsystem, that subsystem, that subsystem,
and you can measure your progress as what percent of the components are built.
Here, you really have ideas that you need to try out and that it's sort of unpredictable in some sense.
And actually, one of the most important changes to the project and making progress was initially
the way that the project management was happening was that each week, well, so we'd written down
our milestones of let's beat this person by this date, let's beat this other person by this state,
let's be able to do kind of these outcome-based milestones on kind of a weekly or bi-weekly basis
that those things would come and go and you wouldn't have them.
And then what are you supposed to do?
It's completely unactionable, right?
It's not like there was anything else you could have done.
It's just you have more ideas you need to try.
And instead of shifting it to what are all the things we're going to try by next week?
That's a good insight.
And then you do that.
And then, yeah, if you didn't actually do everything you said you were going to do,
then you should feel bad.
and then you should do more of it.
And if you did all those and like it didn't work, then, you know, fair enough.
But you achieve what you want to achieve.
And so even going into the international, so two weeks before the international was kind of our cutoff for at this point, there's not much more we can do that we're going to do our biggest experiment ever, put all of our compute into one basket and see where it goes.
And at that point, like at two weeks out, how good was the bot?
Oh, it was barely sometimes winning with professionals that we had testing.
But not even always.
No, no, it sometimes happened.
So, yeah, so to be specific, I'm just pulling this back in.
So July 8th is when we had a first win against our semi-protester.
And then...
A sequence of losses.
A sequence of losses.
And then we were kind of more consistent with it.
And then he went on vacation, and so he was on some laptop somewhere that was not very good.
And then we were consistently beating him.
But that was not very reliable data.
This was the week before the international.
And so we didn't really know how good we were getting.
We knew that the true skill was going up.
When was the last time that, like, an open AI employee beat the bot?
How far out was that from?
I think, like, a month or two before the eye, although we're not very good at Dota.
Okay.
But so, like, a month or two out, it could beat all.
to open AI people. Two weeks out, it could one time beat a semi-pro?
I'd say, so four weeks was, yeah, the first time that it beat the semi-pro.
Okay.
And then, you know, two weeks out, we don't know.
We still can't really find out.
I mean, I guess we could rerun that bot.
Yeah.
But, you know, we really didn't know how good it was.
At that time, we just knew, hey, we're able to beat our semi-pro occasionally.
And we, going into the international, figured that, hey, there's a 50-50 shot.
And I think we were telling Sam the whole way, like,
with these things,
you never really trust the probabilities.
You just trust the trend of the probabilities.
Even that was just swinging wildly.
You guys would text me every night.
It would be, oh, we're going to, you know,
no chance.
We're not going to.
We're definitely going to win every game.
Yep.
Yep.
And so it was very clear that our own estimates
of what was going to happen were miscalibrated.
And throughout the week of TI, actually,
we still didn't know.
And what was happening is we...
See, you guys all went to Seattle for this week.
Most of the team went there.
Okay.
So you're like hold up in a hotel or a conference center or something in Seattle?
Well, actually the reality of it was that we were hold up like near the stadium
which is happening.
Let me describe how we were hold up.
So we were given a locker room in the basement of Key Arena.
So we all had production badges and so you feel very special.
As you walk in, you're just like, oh yeah, you know.
I just get to kind of skip the line and go to the backstage area.
but it was literally a locker room that they converted into a filming area.
And we all had our laptops in there and that they would also bring in pro players every so often.
We had a whole filming setup and then we'd play against the pros.
And we had a partition that we set up, which was just like a black cloth, basically,
between like the whole team sitting there being like, are we going to be able to beat this pro maybe?
And trying to be as quiet as possible.
And these pros who were playing.
and you know on Monday they brought three or I think two pros and like one very high ranked analyst by
and we had our first game and you know we really didn't know what was going to happen and we beat
this person three three zero and you know this was actually a very exciting thing for everyone
at opening eye where so at the time what I was I was kind of live uh live slacking the updates as the
game is like this person you know just said this and like you know now it's this many last
And yeah, were you winning by a large margin?
So, yeah, do you remember the details of that one?
Which game specifically?
This was Blitz.
Oh, Blitz.
I think we won every game.
Yeah, we did 3-0.
I don't know exactly what the margin was.
We have all the data, but Valve brought in the second pro, this professional
named Pyecat, and he played the bot, and we beat him once, we beat him twice,
and then he beat us.
Oh, okay.
And that looking at the game, we knew exactly what had happened.
Yeah, essentially what happened is he accumulated a bunch of want charges, right?
There's this item that accumulates charges, and he accumulated more charges than our bot has ever seen in game.
Because just our bots don't, it turns out that like there was a small, I think it's safe to save kind of a bug in a setup.
In Dota?
In our setup.
Oh, okay.
So basically it passed some threshold that your bot was not ready for.
Well, I'd say very specifically, the kind of the root cause here was that he had gone for an item, an early wand build.
Okay.
And we had just never done an item early wand build.
And so it's just like our bot had just never seen this particular item build before.
Okay.
And so it had never had a chance to really explore.
What does it mean?
And so it had never learned to save up stick charges and to use them and whatever.
And so that it would do, it's very good at calculating like, who's going to win a fight.
Wild.
Okay.
But because, and I've got recognized that he's like, I wonder what happens if I push on this axis.
And sure enough, it was an axe that the bot hadn't seen.
And so then we played a third match against another pro, went 3-0 on that.
And it's actually very interesting getting the pros of reactions because we also didn't really know, are they going to have fun?
It's going to be cool.
I think you're going to hate it.
And we got a mix of reactions.
Some of the pros are like, this is the coolest thing ever.
I want to learn more about it.
One of the pros is like, this thing's stupid.
I would never use it.
But apparently, after the pros left that night, they spent four hours just like talking about the bot and kind of what it meant.
Yeah, and the players were like highly emotional in their reactions to the bot.
They were never beaten by the computer.
So it's kind of unbelievable.
So for example, one of the players that actually managed to eventually beat the bot,
he was like, okay, this bot is effing useless.
Like I never want to see it.
And then he kind of called down.
And like after like five hours term minutes, he's like, okay, this is actually great.
This is going to improve my practice a lot.
And so after your bot, it lost that first time, did they start talking about counterintuitive
strategies to beat it?
Well, so I think at that point that, you know, I think that, well, actually, I don't know,
maybe you can answer that particular one.
Yeah, so I don't think pro players are interested in that.
The pro players are mostly interested in the aspect where it lets them get better at the game,
which means that.
But there was a point after the event where we set up this big land party where we had like
50 computers running the boat.
we kind of unleash this swarm of humans
to kind of add our boat
and they found all the exploits
and we kind of expected them
to be there because
the boat can only learn
as well as the environment
in which it plays allows it to
right so there are some things that it just never seen
and of course those those
ones will be exploitable
and we are kind of excited
about our you know our next day
is 5V5 because 5B5 is one giant exploit.
Like essentially it's about like exploiting the other team like being where they don't
expect you to be kind of like doing auto-dissevision things.
So naturally we know we will have to solve those problems head on for 5E5.
Right.
So one thing I think was pretty interesting about the training process is that a lot of our job
while we were doing this was seeing what the exploits were and then making a small tweak
that fixes them.
And like the way that that I now think about machine learning systems is that they're really a way to make the leverage of human programmers go way up.
Okay.
Right.
Because again, normally like when you're building a system, you build component one, component two, component three.
And that kind of your marginal return on like building component four is like, you know, similar to your marginal return on component one.
Whereas here, a lot of the early stuff that we did is just like your thing goes from being like crappy to like slightly less crappy.
But once we're at the international and we had this loss to PiCat, we knew.
okay, well, the root cause here is just, it's never seen this item build before.
Well, all we had to do was make it tweak to add that to our list of item builds.
And then it played out this scenario for the next, you know, however long.
Can you walk me through actually how that tweak works on the technical side?
Because my impression is kind of what you guys have been saying.
It's just been in a million games.
So it kind of has learned all this stuff.
And some people talk about, you know, these networks as just very gray.
and they don't actually know how to manipulate what.
How are you guys getting in there and changing things?
Yes.
So it's kind of funny.
In some sense, on a high level, you can compare this person to teaching a human.
Like, you know, like, kind of, you see a kid doing maths,
and it's kind of, like, confusing, like, addition with subtractions, suppose, right?
And you're kind of like, look here at this symbol.
This is what you're not seeing clearly, right?
And the same of those tweaks to our boss.
So clearly our bot has never seen this one build that Greg mentioned.
And all we had to do is we had to say that like when the bot plays games and chooses what items to purchase,
we just need to add some probability of sampling that specific build that it has never seen.
And when it plays a couple of games against opponents that use that build,
when it uses this build a couple of times itself,
then it kind of becomes more comfortable with the idea of what happens,
what are the in-game consequences of that build.
Okay.
And kind of a, I have kind of a couple different levels to that answer,
I think are pretty interesting.
So one is at a very like kind of object level.
So the way that these models work is you basically do have a black box,
which takes in some list of numbers and outputs a list of numbers.
And, you know, it's very smart and how it does that mapping.
But that's what you get.
Yeah.
And then you think of this as this is my primitive.
Now what do I build on top of that so that as little work as possible has to be done inside of the learning here?
And a lot of your job is, well, one thing that we noticed that we'd forgotten as well on Monday was, well, it wasn't that we'd forgotten.
We just hadn't gotten around to it was the passing in data that corresponds to the person was passing in the visibility of a teleport.
So as a human, you can see when someone's teleporting out.
Our bot just did not have that feature.
The list of numbers passed in did not have that feature.
And so, well, one of the things you need to do is you need to add it.
And so that kind of goes from your feature is, you know, your feature vectors however long it was.
And now it's got one more, one more feature on it.
And the bot wasn't recognizing that as an on screen thing.
So it doesn't see the screen.
It's passed data from the bot API.
Oh, okay.
And so it really is given whatever data we give it.
Okay.
And so it's kind of on us to do some of this feature engineering.
And you want to do as much as you can to make it as easy as possible so that it has to do as little work inside as possible.
So you can spend, you know, you think of it as you've got some fixed capacity.
Do you want to spend that on learning the strategy?
Do you want to spend it on learning how to like, you know, map, you know, choose which creep you want to hit?
Like, do you want to spend that on trying to parse pixels?
Like, you know, at the end of the day that I think a lot of our job as the system designers here is to push as much of that model capacity and as much of the learning towards.
the like interesting parts of the problem that you can't script that you can't possibly do any
processing for. And so that's kind of that's kind of one level is that a lot of the work ends
of being identifying which features aren't there or kind of engineering the observation
in action spaces in an appropriate way. Another thing is I think is like another level where you
zoom out is like the way that this actually happened was so you know we're there on Monday and
people got dinner and then Shimon and Jakop and Raffel and I think you know,
maybe one or two others stayed up all night to do surgery on our running experiment.
And so it was very much like you've got your production outage and like everyone's there,
like all hands on deck trying to go and make the improvements.
Yeah.
So specifically like to kind of zoom in and to give you a bit of a feel, what it felt like working on the world.
Like, you know, this is like very tiring week.
Every day we were like the day was just like meeting with.
the pros and kind of watching our boat getting excited and the nights were kind of coding up the
next version of experiment because actually it's a little known fact but from day to day like each
version of the experiment was not good enough to beat the next player next next next day's professional
so so just that morning we would download the new parameters of the network and it would be
good enough to beat it but the day before it wasn't how are you discerning that well this was this was
again something of almost a coincidence I mean yeah there might be something a little bit deeper
But so the, you know, kind of the full story of the week was we did the Monday play.
Okay.
And that, you know, there we'd lost to PaiCat.
And so just to clarify, are you guys in the competition or not in the competition?
So the thing that we did was we did a special event.
Okay.
To play against Dendi, who's, you know, one of the best players of all time.
And while we were there, that we were also like, well, let's test this out against all these other pros.
Got you physically here right now.
Okay.
Let's see how we do.
Got it.
All right.
So Monday happens.
You start training it.
Yep. And so actually, yeah, so this experiment we've kicked off, you know, maybe sometime the prior week.
Yep.
Three weeks before, I think.
Yeah, something like that.
And we've been running this experiment for a while.
And our infrastructure is really meant for you run an experiment from scratch.
You know, you start from complete randomness and then you run it.
And then you, you two weeks later, go and see how it does.
We didn't have two weeks anymore.
And so we had to do this surgery and this very careful, like, you know, read every single character of your commit to make sure.
to make sure that you're not going to have any bugs
because if you mess it up, we're out of time.
There's nothing you can do.
And it's not one of those things like,
if you're just a little bit more clever,
that you can go and do a hot patch
and have everything be good.
It's just literally the case
that you've got to let this thing sit here
and it's got to bake.
Yeah.
And so Monday came and went.
We were running this experiment
that we performed surgery on.
And the next day, we got a little bit of reprieve
where we just played against some kind of lower ranked players
who were kind of commentated.
and popular in the community, but, you know, we're not pushing the limit of our bot.
On Wednesday, at 1 p.m., our contact from Valve came by and said, hey, I'm going to get you Artizi
and Sumil, who are basically, you know, the top players in the world.
And I was like, could we push them off to Thursday maybe?
And he was like, their schedules booked.
You're going to get them when you get them.
And we're going to, we're scheduled to get them at 4 p.m.
So we looked at our bot to see how it was doing.
We've kind of been along the way gauging it.
We tested it against our semi-pro player.
And he said this body is completely broken.
Oh, no.
And, you know, kind of pictures of maybe we had a bug during the surgery, like went through
our head.
And he showed us the issue.
And he said, look, first wave, this bot takes a bunch of damage.
It doesn't have to take.
There's no advantage to that.
I'm going to run and I'm going to go kill it.
I'll show you how easy it is.
He ran into kill it and he lost.
Okay.
And then again.
But don't jump ahead.
Explain what happened?
So he played it five times, and he lost each time until he finally did figure out how to exploit it.
And we realized what was going on was that this bot had learned a strategy of baiting.
You pretend to be a really dumb bot.
You don't know what you're doing.
And then when the person comes in to kill you, you just turn around and you go super bot.
It was legitimately a bad strategy, you know, if you're really, really good.
But I guess it was good against the whole population of bots that it was playing against.
And you had never seen it until that day.
So, yeah, we had not seen that behavior.
And we did not at all expect it.
It was like one of the major examples of the things that we kind of didn't have explicit incentive for and yet the bot actually learned them.
And yeah, essentially, I mean, it was kind of funny because, of course, when the bot played against its other versions, it was just like good baiting strategy that was kind of all the just.
But it got a very interesting psychological effect on humans because optimal strategy was not to fall for the weight,
it was kind of to wait it out a little bit because the bot already is at a disadvantage.
But he's like, okay, look how stupid this bot is.
I'm going to go for a kill.
So it kind of had interesting psychological effect on humans, which I thought was like, it's kind of funny.
It almost knows it's a bot.
Yeah, it knows how it's attacked.
Yeah, it's funny to see a bot which kind of seems like it's playing with emotions of the
Of course it was not what actually happened, but it seemed this way.
So now we were faced with the dilemma.
It's 1 p.m. on Wednesday that these best players are going to be showing up at 4 p.m.
We have a broken bot.
What are we going to do?
And we know that our Monday bot is not going to be good enough.
We know it's not going to cut it.
And so the first thing we do is we're like, well, Monday bot, it is pretty good at the first wave.
This new bot is a super bot thereafter.
Okay.
So can we stitch the two together?
So we wrote, so we already, we had some code for doing something similar.
So we kind of revived that.
And then in the three hours, Jay spent his time doing a very careful stitch where you run the first bot and then you cut over at the right time to the second bot.
And this is literally just like bot one plays the first X amount of time and then Bot two takes over.
And he finished it 20 minutes before the year.
Of course.
Before the thing.
We ran it by our semi-pro.
Semi-pro is like, this is great.
So we at least got that done in the nick of time.
But the other question was, how do we actually fix this bot?
I mean, actually just to finish your story, it was like one aspect because we are also kind of uncertain what happens when you switch over from one boat to the other.
So I was actually standing by the pro who was playing it.
And I was looking at the time at the moment when it was switching.
I was like,
I kind of distract the guy in case something stupid.
Just trying to distract him for once in a moment.
And of course, it was probably completely unnecessary.
But we weren't sure what would happen there.
So I didn't know about that part of the story.
So the question of how do you actually fix it?
So there was a little bit of debate of like maybe we should abandon ship on this, switch back to our old experiment, run that one for longer.
And I forget who suggested it.
But someone was like, I think we just have to let it run for longer because you learn a strategy of baiting.
Well, the counter strategy for that is just don't bait.
Play well the whole time.
And so we got that run for the additional three hours.
And so we first played Artizzi, who showed up on.
on our Switchbot, you know, kind of the Franken bot, and, you know, that beat him three times.
And right, all right, let's try out this other bot and just see what happened with the additional three hours of training.
Because, you know, our semi-protester, it at least validated that, like, it looks like it's fixed.
And so in that three hours of training, how many games is it actually playing simultaneously?
It's a good question, quite a bit.
Yeah, okay.
And so we played this new bot against Ortizzi.
didn't know how he was going to do, and sure enough, it beats him.
And he loved it.
He was having a lot of fun.
He ended up playing 11 games that day, or maybe it was 10, but I think that he was just
like, oh, this is so cool.
We were supposed to have Sumil that day as well.
But due to his scheduling snafu, he had to be at some panel, and so, like, timing didn't
work out.
But Artizzi and his coach, who also coaches Sumil both said, yeah, both said,
Sumail's going to beat this bot.
Like, it's going to happen.
You know, maybe he'll have a little bit of trouble.
to figure it out for the first game, but like after that, you're in trouble.
And so, you're like, all right, we've got one more day to figure out what to do.
And so what did we do?
I don't know.
It's kind of like some nice dinner.
Kind of what did you say?
Kind of went for some nice dinner.
We kind of rested a lot.
We kind of chatted, you know, like, you know, Slack with some people at home.
And then in the morning we download the new parameters of the network and just let it play.
We literally just hung out and just let it go.
Just let it play.
It's the exact opposite of how I'm used to engineering deadlines happening.
Yeah.
Normally it's your work right up until the minute.
So you guys weren't like you guys were getting like full nights of sleep, nice and relaxed.
Oh no, no, absolutely no.
Okay.
So to make this clear, the night before the night, like two nights before the day where we got the restaurant relaxation, the night looked something like the following.
We had like, we had full day of dealing with the pros and kind of like emotional highs and so on.
is absolutely not.
Come midnight, we start working.
Okay, we need to make all those changes.
Like, the one thing that we talked about around midnight,
we start with four people.
And we are also tired that, like, you know,
we look through all the committees
that we are going to add to the experiments.
There were actually two people looking at them
because we didn't trust a single person,
given how tired we all were.
So they were, like, looking at those coins to the 6 AM.
I was doing this, like, updating the model,
which is a lot of nasty, like, off-by-one indexing
thing. So even though it was a short call, it took me like six hours to do. Somewhere around
3 a.m. we had like a phone call with Azure because it turns out with certain number of
machines you start exceeding some limits. So we tried to make them raise the limits. And around
6 a.m., we are okay. We are ready to deploy this. And then there was, deploying is just like one-man job.
So, Jakub was just like, you know, hitting, like, you know, clicking the deploy and kind of fixing all the issues that came up.
I was staying around just exclusively to make sure that Jakub doesn't fall asleep.
And eventually at 11 a.m. experiment was running and we kind of went to sleep, like walk up at 4 p.m. or something.
And then it was like all.
So it had over 24 hours to train.
I think I think it ended up being like one and a half day until the game with Samail.
Sorry.
Yeah, just to repeat the timeline.
So this was Monday is when we played first set of games, had the loss, did the surgery that night.
You know, I played it, I guess, starting on 11 on Tuesday.
Then, yeah, Wednesday, 4 p.m. is when we played Artizzi and then trained for longer.
I don't think we made any changes after that.
Maybe we made some small ones.
But then on Thursday is when we played Sumil.
Okay.
So I think Tuesday to Wednesday was the night where we made last changes.
Yeah.
And there was like quite a bit, like it's, you know, there's quite a bit of different work going on that all kind of came together at once.
Like one thing I think is really important was, so we, we, one of our team members, his handle of Saihou, who's a very well known program competition competitor was spending a lot of time just watching the bot play and seeing why does it, do.
this weird thing in this case?
You know,
what are all the like,
you know,
weird tweaks and really getting intuitions for,
oh,
because we're representing this feature in this way.
And so if we change it to this other thing,
then it's going to work in a different way.
And I think that like this really trying to,
it's almost this very human-like process of watching this like expert playing the game
and try to figure out what are all the little micro decisions
that are going into this macro choice.
And it's kind of interesting,
starting to have this very different relationship to the system you build.
because normally the way that you do it is, well, your goal is to have everything be very observable.
And so, yeah, you want to put metrics on everything.
And like, you know, if something's not understandable, add more logging.
Like, you know, that's how you design the systems.
Whereas here, you have this, you know, you do have that for the surrounding bits, but for this machine learning core,
that there you really do have to understand it at more of a sort of behavioral level.
Was it ever stumping you where you're just like, oh, it's being creative in a way that we didn't expect it to?
and it may be even working,
but you don't know why
or how it decided to make that choice.
Yeah, I think the debating story
that we shared is the main one.
It's the main story like this.
We got a few small ones,
like where, you know,
there was like some early days of the project
where like we have professionals
like playing the next stage of the body.
He's like,
hmm, everybody's really good at crippling.
And we are like, oh, what is crippling?
I'd say that there is also one
one other part of the story that I think is interesting, and then I think probably we can wrap up this part.
But so to see how well, you know, because our semi-protester had played hundreds of games against this bot over the past, you know, a couple months.
And so we wanted to see just how to see benchmark relative to Artizi.
And so we had him play against Artizzi.
And, you know, Artizu was up the whole game.
It was just like beating him to the last hit by like 500 milliseconds every single time.
And so our semi-pro was like, all right, I've got one last-stitch effort to go try this strategy that the bot always does.
to me. And it's like some, some, you know, strategy where you, like, do something complicated and
then you, like, triple wave the, the, the, the, your opponent, you get him under the tower,
you have regent, you go and you, you go in for the kill. And he did it, and it worked.
Oh. And this was the, the, the bot had, like, taught him the strategy that you could use
against a human. And I think that was, like, very interesting and a good example of the kinds of
things that you can get out of these systems, that they can discover these very sort of, you know,
non-obvious strategies that can actually be taught to humans.
And how did it go with Sumil?
So with Sumil, we went undefeated and I think it was 5-0 that day.
One thing that's actually interesting, so we'll probably blog about this in upcoming weeks.
But we've actually been playing against a bunch of pros since then.
So our bot has been very high demand.
And some of these pros have been live streaming it.
And so we've gotten a better sense of kind of watching as humans go from just being
completely unable to beat it to if you play against it for long enough you can actually get pretty
good and so there's actually a very interesting uh you know a set of stats there that you know will be
kind of pulling and analyzing in a bit are there humans that consistently beat the bot today yeah so
i think there's one who has like a 20% win rate or something i think it might be actually 30
and that player played hundreds of games and just finds strategies to exploit uh no actually because
essentially as good as the bot
at what the bot is doing
she finds extremely surprising
but it turns out that
he played hundreds of games with it
so it's actually
and is he a top player
like does he beat most humans
these are all professionals
okay it's not just some random kid
who's good at beating the bot
that's right
the way to think about this is that
like yeah I mean
being a professional video game player
is a pretty high bar
yeah I think everyone wants to be a professional
video game player
you know who play these games
and the number of pros is very small.
And there are some who I have really like, you know,
when you're playing hundreds of games against it,
you're going to get very, very good at the things that it does.
And so talking to Artizi,
I was asking him,
has it changed your play style at all?
And he said,
he thinks that the thing that it's done for him is it's helped him focus more.
Because like, you know,
while you're just there in lane last hitting,
now suddenly like,
that's just so wrote, right?
Because you just have been doing it so much.
You've gotten so good at it.
And I think that one really,
interesting thing to see is going to be how can you improve can you improve human play style can
you change human play style and i think that we're starting to see some positive answers in that
direction so i know we're almost out of time i could do like a little lightning round just
quickly go through some of this question of this guys actually like to the question of like what kind
of skills uh you need to to work at open i could we have like a very small that was going to be the
first lightning round question so so specific list of things uh that that we've
found very useful, at least in the Dota team, is some knowledge of distributed systems,
because we build a lot of those and those are easy to not do properly.
And another thing that we found very important is actually writing back free code.
Essentially, I know it's kind of taken for granted in computer science community that
like everybody makes bugs and so on, but here it's even more important than other projects
that you minimize it because they're very hard to debug.
Specifically, many bugs manifest in kind of lower training performance where to get that number
it takes a day and in like a spree of hundreds of convince.
It's really easy to miss.
And primary way of debugging this is actually reading the code.
So every bug has very high cost associated with it.
So actually writing like this correct bug free code is quite important to us.
And we sometimes actually kind of sacrifice good engineering habits, good kind of code modularity
to make our code shorter and simpler and kind of having less, essentially less lines where you can make bugs.
And I guess lastly, as we mentioned, like primary skill,
is good engineering.
But if somebody really feels like, gosh, I really need to brush up my maths,
I really need to kind of go in there and feel comfortable, like,
not have like somebody asking questions about maths and I understand that.
I think mostly like getting good basics in linear algebra, in linear algebra and in basic
statistics, that's especially when doing experiments, it's easy to make like,
elementary statistics mistakes and linear algebra is just kind of most of what you need to know
like basic optimization as well to follow what's what's happening in those models but but
this is kind of compared to being a good engineer quite easy to pick up at least in project like
the one we are doing yeah so I wanted to talk about some non-technical skills that I think are
are really important. So one is that I think that there's like a real humility that's required.
If you're coming from an engineering background like I am and working in these projects where
you're no longer the technical expert in the way that you're used to, right? And I think that,
you know, if you go and you build, you talk to, you talk to like, you know, let's say you want to
build a product for doctors, right? I think you can talk to 10 doctors. And honestly, whatever thing
you're going to build is probably going to be a pretty valuable addition to their workflow because
doctors can't really build their own software tools. You know, maybe some can, but, you know,
as a general rule, no.
Whereas with machine learning research, you know, everyone that you're working with is very
technical can build their own tools.
But if you inject engineering discipline in the right place, if you build the right tool
at the right time, if you kind of look at the workflow and think about, oh, we could do it
in this other way.
That's where you can really add a bunch of value.
And so I think it's about knowing when to inject the engineering discipline, but also
knowing when not to.
And being, you know, to Shimon's point, you know, sometimes, well, we really just want
the really short code because we're really terrified of bugs.
and so that can yield different choices
than you might expect
for something that's just a pure production system.
Who writes the least bugs at all of Open AI?
That's a contentious question.
What's the question?
Who writes the least bugs per line of code
at all of Open AI?
I'm definitely not going to say me.
Possibly, Jacob.
Yeah.
It's hard to say.
But it's, yeah, it's three cuts.
It could be Greg.
I read a lot of bugs.
I caught Jacob the least amount of times on bugs.
So, okay.
It's more okay to have bugs that are going to cause exceptions.
Right.
And my bugs usually cause exceptions.
So that's fine.
That's fine.
Yeah.
It's what you don't want is the things that cause correctness issues where it gets 10% worse.
Yeah.
So there was another question related to skills.
But this is for non-technical people.
Yep.
So Tim Beko asks, how can non-technical people be helpful to AI startups?
Well, I was going to say, I think, I think one.
important thing is that for AI generally right now, I think there's a lot of noise.
And I think it could be hard to distinguish what is real from what's not.
I think just like simply educating yourself, I think is like a pretty important thing.
Like I think it's very clear that AI is going to have a pretty big impact.
And, you know, that's just look at what's already been created and extrapolate that without any new technology development, any new research.
And it's pretty clear there's going to be baked into lots of different systems.
There are a lot of ethical issues to work through.
and I think that being kind of a voice in those conversations and educating yourself,
I think is like a really important thing.
And then you look to, well, what are we going to be able to develop next?
And I think that that's where the really transformative stuff's going to come.
Okay.
I once saw a post of Greg's Rescue Time Report and was pretty shocked.
Do you have any advice for work for working such long, focused hours?
I think it's not a good goal.
I would not have a goal of trying to maximize the number of hours you sit at your computer.
For me, I do it because I'd do it.
love it. And that the thing that the activity that I love most in the world is when you're in the
zone writing code, producing it for something that's meaningful and worthwhile. And so I think that
as a second order effect, it can be good. But I wouldn't say that like that is the way to have an
impact. I will also say more specifically the only way I've ever seen people be super productive
as if they're doing something they love. There is nothing else that will sustain you over a long enough
period of time. Okay. Is the term AI overused by many startups just to look good in the press?
Yes.
Indeed.
Okay.
What is the last job that will remain as AI starts to do everything else?
The last human job?
What is going to be the hardest thing for AI to do?
It's a hard question to answer in general.
Because I think it's actually not AI researcher.
The AI researcher will kind of go before.
It's actually very interesting when you ask people this question.
I think that everyone tends to say whatever their job is, as a lot of
hardest one. But I actually think that AI research is going to be one that you're going to want to
make these systems very good at doing. Totally. I think the last question, maybe this is obvious,
is can you just connect the dots between how playing video games is relevant to building AGI?
Yeah. It's actually maybe one of the most surprising things to me, the degree to which games end up being
used for AI research. And the real thing that you want, right, is you really want to have
algorithms that are operating in complex environments where they can learn skills and that you want
to increase the complexity of those skills that they learn. And that's either you push the environment,
you push the complexity of the algorithms, you scale these things up. And that that's really
the path that you want to take to building really powerful systems. So games are great because
they are a pre-packaged environment that some other humans have spent time sort of making, first
of all, putting in a lot of complexity, making sure that there's like actual intellectual things
to solve there, or not even just intellectual, but, you know, like interesting mechanical challenges
that you kind of can get human level baselines on them, so you know exactly how hard they are,
that they're very nice, unlike, you know, something like robotics, where you can just run them
entirely virtually, and that means you can scale them up and you can run many copies of them.
And so they're a very convenient test bed.
And I think what you're going to see is that there's a lot of work that's going to be done in games.
But the goal is, of course, bring it out of the game and actually use it to solve problems in the real world and to actually be able to interact with humans and do useful things there.
So I think they're a very good sort of starter and a very good place to, like, I think one thing that I really like about this Dota project and bringing it to all these pros is that we're all going to be interacting with super advanced AI systems in the future.
And right now I think we don't really have good intuitions as to how they operate, where they fail, what it's like to interact with them.
And this is a very low-stakes way of having your first interaction with very advanced AI technology.
Cool.
If someone wants to get involved with Open AI, what should they do?
Well, we have a job posting at our website.
I guess the tips that are giving about how to get a job at Open AI are very geared towards the specific job posting that we have there,
which is a large scale for a learning engineer.
Cool.
And in general, we look for people who are very good at whatever technical access they specialize in, and we can use lots of different specialties.
Great.
Thanks, guys.
Just to echo that, like everyone thinks they have to be an AI PhD.
Not true.
Neither of these guys are.
All right.
Thanks a lot.
Cool.
Thank you.
Thank you.
All right.
Thanks for listening.
So, as always, the video and transcript are at blog.w.wik combinator.com.
And if you have a second, please subscribe and review the show.
All right. See you next week.
