Y Combinator Startup Podcast - #5 - An AI Primer with Wojciech Zaremba

Episode Date: May 15, 2017

Wojciech Zaremba is a cofounder of OpenAI (https://openai.com). OpenAI is a non-profit AI research company, focused on discovering and enacting the path to safe artificial general intelligence. Read... the transcript here (http://blog.ycombinator.com/an-ai-primer-with-wojciech-zaremba).

Transcript
Discussion (0)
Starting point is 00:00:00 Hey, this is Craig Cannon, and you're listening to Y Combinators podcast. Today's guest is Vojchezuremba, who's a co-founder of Open AI. Open AI is a non-profit AI research company. They're focused on discovering and enacting the path to safe, artificial, general intelligence. This episode is a bit of a primer on AI, as we have several AI interviews coming up, and they tend to be a bit more specific than this one. All right, here we go. Hey, today we have Vojchecks Suramba, and we're going to talk about AI.
Starting point is 00:00:26 So, Voichick, could you give us a quick background? I'm a founder at OpenAI. I'm working on robotics. I think that deep learning and AI is a great application for robotics. Prior to that, I spent a year at Google Brain and I spent a year at Facebook AI research. And same time I graduated from, I have finished my PhD at NYU. Can you explain how you pulled that off? That seems pretty rare. So the great thing about both of these organizations is that they are focused on research. So throughout my PhD, I was actually publishing papers over there. I highly recommend both organizations as well as, of course, OpenAI. Yeah, okay, so most people probably don't know what OpenAI is. So could you just give that quick explanation?
Starting point is 00:01:21 So Open AI focuses on building AI for the good of human beings. We are a group of researchers and engineers collaborating together who essentially try to figure out what are the missing pieces of general artificial intelligence and how to build it in a way that would be maximally beneficial to humanity as a whole. Open AI is greatly supported by Elon Musk and Sam Altman. in total we gather an investment of one billion dollar in the group which is quite a lot and so what are I mean I know some but what are the open AI projects so there is several large projects going on simultaneously we have also we are doing also basic research so let me first enumerate large projects these are robotics so In terms of robotics, we are working on manipulation.
Starting point is 00:02:29 We think that manipulation is the complete, it's one of the parts of robotics, which is the most unresolved. Sorry, just to clarify, what does that mean exactly? It means that, so in robotics, there are essentially three major, three major families of tasks. One is locomotion,
Starting point is 00:02:50 which means how to move from, let's say how to walk, how to move from point, to point B. Second is navigation. It's say you are moving in the complicated environment such as, for instance, a flat or a building and you have to figure out actually to which rooms, which rooms have you visited before, which not, and where to go. And the last one is manipulation. So it means you want to grasp an object, let's say, open, an object, place objects in various locations.
Starting point is 00:03:29 And the third one is the one which is currently the most difficult. So it turns out that when it comes to arbitrary objects, current robots are enabled to just grasp an arbitrary object. For any object, it's possible to hand-code a single solution. So say as long as, let's say, in fact, or if you have the same object, like, I don't know, you are producing glasses. and there exists hand-coded solution to it. There is a way by code to write a program saying,
Starting point is 00:04:02 let's place a hand in the middle of the class and then let's close it. But there is no way so far to write a program such that it would be able to grasp an arbitrary object. Okay, got you. And then just very quickly, the other open-AI projects going on. So another one has to do with playing a complicated computer. game and the third one has to do with playing large number of computer games and you might ask why it's interesting and in some sense I would like to see yeah that so so human is has an incredible skill of being able to learn extremely quickly and it has to do
Starting point is 00:04:54 with a prior experience. So let's say even if you haven't played ever a volleyball, if you try it out for the first time within 10 or 15 minutes, you would be able to grasp how to actually how to play. And it has to do with all the prior experience that you have from different games. If you would put the child, like if you would put an infant on the volleyball baller ball ball court and ask him or her to play, it would fail miserably. But I mean, due to the fact that it has experience coming from large number of other games, or let's say other life situations, it's able to actually transfer all the knowledge. So at OpenAI, we're able to pull together a large number of computer games.
Starting point is 00:05:46 And computer games can be, it's quite easy to quantify. how good they are in the computer game. Currently, best AI systems. So, first of all, it's possible for many computer games to write a program that solves it pretty well or plays it well. There are also results from, there are results in terms of reinforcement learning or in terms of so-called deep reinforcement learning, showing that it's possible to learn how to play a computer game.
Starting point is 00:06:20 These are, these are, like the initial results are coming from deep mind. And, but simultaneity, simultaneously, it takes extremely long time,
Starting point is 00:06:36 like in terms of real-time execution to learn to play computer games. So, for instance, Atari games, for instance, in terms of real-time execution, it takes something around three years of play to learn to play simple games.
Starting point is 00:06:59 I mean, it can be hugely paralyzed, therefore it takes a few days to train it on current computers. But it's way shorter for human. In 10 minutes, we can kind of... Teach it how to play and win? Yes. Okay. And is that through you giving it feedback? So the way how it works in case of computer games,
Starting point is 00:07:20 the feedback comes from the score. So it looks at the score in the game and tries to optimize it. And I would say that's kind of reasonable, but I would say simultaneous is not that satisfying to me. So the reason why it's not that satisfying to me, so the assumption underlying reinforcement learning is that there is some environment.
Starting point is 00:07:46 And environment, you are an agent, and you're acting in environment by executing actions and getting rewards from the environment. And the rewards might be taught as, let's say, pleasure or so. And the main issue is that it's actually not that easy to figure out what are the rewards in the real world. Further on, other underlying assumption is in being able to reset environment to kind of get to repetitive the same situation
Starting point is 00:08:14 so the system can try thousands or millions of times to actually finish a game. So there are some small discrepasses. People also believe that it might be possible somehow to hard code into system rewards, but I would say that's actually one of the big issues that it's kind of unresolved. Like when I look how my nephew plays computer game,
Starting point is 00:08:39 he actually doesn't look on score because he cannot read. And still my nephew, yeah, they can play pretty well. So I mean, you can say maybe reward is somewhat different. Maybe reward comes from like a sing a nice, hearing nice voice in the game or so. But I would say that's something what is very unclear how to build a system and what system should optimize. So in some sense, if we have a metric that we want to optimize, it's possible to build a system that could optimize for it. but it turns out that in many cases it's not that easy and I would say that's actually
Starting point is 00:09:22 one of the motivations why I wanted to work on robotics because in case of robotics it's way closer to the system that we care about so what I mean by that for instance let's say you would like your robot to creeper scramble X for you and so the question is So how should I build a reward? And in computer games, actually, the nice thing is they are getting reward extremely frequently. So let's say any time you kill an enemy or let's say won't die, it's quite great. But in case of scrambling eggs, it would mean,
Starting point is 00:10:04 or the way how people write rewards for systems, it would mean distance from hand to a pen. Then let's say somehow you have to quantify what's the, if the egg is, if you were able to crack open an egg, or let's say if you fried it's sufficiently, and how to kind of quantify, it turns out to be extremely difficult. And also there is no way even to reset the system, how to reset the system to the same place. So, these are like fundamental issues. And the reason why I'm personally interested in robotics is thing that actually this challenge.
Starting point is 00:10:46 will tell us how to solve. So let's start by defining a couple things. So what is artificial intelligence, what is machine learning, and then what is deep learning? Okay. These are pretty good questions. Okay. So artificial intelligence is actually extremely broad.
Starting point is 00:11:08 It's an extremely broad domain, and machine learning is subpart of this domain. And in essence, artificial intelligence consists of writing any software that tries to solve some problems through some intelligence. It might be hand-coded solution, rules-based system. Yeah, so pretty much it's actually very hard to say what is not artificial intelligence. You can say that. So initial version, for instance, of Google search. was based on, it was avoiding any machine learning,
Starting point is 00:11:53 and it was, there was like a well-defined algorithm called page rank, and essentially page rank counts how many incoming links are from other websites, and that's artificial intelligence. It's an essentially system that does intelligent things for you. Then over the time, Google Search started to use machine learning because it was it helps to improve results at Simultensity they wanted to avoid it for some time as it's more difficult to interpret the results and it's more difficult to actually understand what system does so what is
Starting point is 00:12:38 machine learning machine learning it's essentially way of building, or let's say that's essentially you have data and you would like to generate based on data program with some behavior. So like the most common example, which is still sub-branch of machine learning, so-called supervised learning. So you have pairs of examples, X, comma Y, which means like I would like to map X to Y. for instance, either if given email is spam or not spam or let's say if an image, what is the category of an image or for instance, to whom should I recommend given product. And based on this date, I would like to generate a program, some sort of the black box or some
Starting point is 00:13:41 function that for new examples would be able to give you similar answers. And that's an example of supervised learning. But the sense, machine learning means that you would like to generate program from data. Okay. And this usual uses statistical machine learning method. So somehow you can't somebody, how many times given events occurred or so. Okay. got you
Starting point is 00:14:11 and then the third being deep learning so deep learning that's that's one paradigm in terms of machine learning and idea behind is
Starting point is 00:14:28 ridiculously simple so so people realized that if you want to, as I said, machine learning means that you get data as an input and program as the output. And deep learning says the computation of the program, what I'm actually doing with this data should involve many steps. Not one step, but many.
Starting point is 00:15:04 And pretty much that's it in terms of meaning of deep learning. So you might ask why it's so popular now and how it's so different from what was there before. So it turns out that if you assume that you do one step of computation, let's say that you take your data and you kind of have single if statement or small number of if statements, then like for instance, say if you have a, I don't know, let's say your data is a regular,
Starting point is 00:15:38 according from a stock market and you're saying you're going to sell or buy depending on value speaker or smaller than something or if let's say depending or who is the new president or so you are making some decisions. So in sense there's not that in case of models that are based on single step, people are able to prove plenty of stuff mathematically. And in terms of models that require multiple steps of computation, mathematical, antical proofs are extremely weak. And for a long time, models that do single step of computation,
Starting point is 00:16:14 they were outperforming models that do many steps of computation. But recently it kind of changed. And it was, for many people, it was obvious for a long time that true intelligence cannot be done in single step, but it would require many steps. But so far, many systems,
Starting point is 00:16:33 actually they worked in the way that they had kind of very, very shallow, they were very shallow, but simultaneously extremely gigantic. So what I mean by that, you could generate, let's say, for the task of interest, let's say the recommendation, you could generate large number of features, let's say thousands of them. These are features saying, for instance, let's say you want to do movie, recommendation. You can say, is movie longer or shorter than two hours? Is it longer or shorter than one hour? There are two features. You can say, is it drama? Is it trailer? Is it something? And you can
Starting point is 00:17:21 generate a million of these. And then, or let's say 100,000, that's actually quite reasonable value. And then your shallow classifier can determine based on the combination of these features, either to recommend it to you or not. In case of deep learning, you would say, let's kind of combine it for multiple steps. And that's essentially that's entire difference. And in case of deep learning, the most successful embodiment of deep learning is in terms of neural networks.
Starting point is 00:18:06 Okay. So let's define that too. So neural networks, it's also extremely simple concept, and that's something that people came out with a long time ago. And it means it follows. You have an input. This might be, say, vector, or it might have some additional structure like a, let's say, image.
Starting point is 00:18:34 So it's kind of a matrix, two-dimensional. And you neural network, it's a sequence of layers. Layers are represented by matrices. And what you do is you multiply your input by a matrix and apply some non-linear operation and multiply it again by a matrix and apply non-linear operation. You might ask, why would I? even need to apply this non-linear operation,
Starting point is 00:19:08 it turns out that if you would multiply by two matrices, it can be reduced to multiplication by single matrix. Like a composition of two linear operators can be written as single linear operator. You could multiply these matrices together and the result of the, and you could condense it into single matrix. Okay. And non-linearity is something like the classical non-denarity. So I'd say there are extremely large number of variants in terms of what I said.
Starting point is 00:19:50 But what I just described is so-called feed-forward neural network. So it essentially takes input, multiplies it by matrix, non-linearity multiplies it by matrix. Examples of non-denarities, there is something that one which is classical, something called Sigmoid. So, sigmoid is a function that it has a shape of S character, S letter. It's kind of close to zero for negative values. It grows to half at zero and then goes up to one when the values are larger. it kind of modulates the input
Starting point is 00:20:34 and that's the most classical version of activation function it turns out that the one which is even simpler empirically works way better which is called RELU rectify linear unit and this one is ridiculously simple
Starting point is 00:20:57 RELU is just maximum of zero comma So when you have negative value, you set zero. You have positive value. You just copy the value. And that's it. So you might ask, so first of all, what are the successes of deep learning? Why we actually believe that it works?
Starting point is 00:21:19 Why? What change? And why it's so much different than it was before? And they're like some few differences. This is a good question. No, it's exactly where I was going to go, but I was going to ask beforehand, yeah, why neural networks are a thing now as opposed to in the past? The main difference is all of a sudden we can train them to solve various problems.
Starting point is 00:21:48 And let's say one family of problems. These are problems in supervised learning. So better than any other method, they can map these examples to labels. And then on the holdout data, on test data, they out. anything else and in many cases they get superhuman results and is that just a function of like computational power that we have access to when it comes to models and neural networks is an example of model there is always a question so how to figure out parameters of a model so there is some training procedure and the most common procedure for neural networks is so-called stochastic gradient descent
Starting point is 00:22:24 It's also ridiculously simple procedure. And it turns out that empirically it works very well. So people came out with vast number of learning algorithms. Stochastic gradient descent is an example of one learning algorithm. Others, let's say there is something called Hebian learning that's motivated by the way how neurons in human brain learning. learn, but this one so far empirically is working the best. Okay. So then let's go to the question you asked yourself, which is why now?
Starting point is 00:23:04 Like, what's happening to make people care about it right now? So since 20 years ago, there were several small differences in terms of how people train neural networks. and there is a large increase in computational power. So I can speak about the major advances. So number one advance, I would say that's even the one advanced. That's actually an old one, but it seems to be extremely critical, something called convolutional neural network. Okay.
Starting point is 00:23:48 And what does that mean? Yeah, so it's actually a very simple concept. So let's say your input is an image. And let's say your image is of a size 200 by 200. It has also, let's say, three colors. So that would, the number of values in total is actually 120,000. So if you would actually squash it into a vector, this vector would be of this size.
Starting point is 00:24:25 And then I can think that if you would like, let's say, to apply neural network to essentially multiply it by a matrix. And let's say if you would like to have output of the multiplication of similar size, let's say 120,000, then all of a sudden the matrix to multiply it would be of a gigantic. size.
Starting point is 00:24:49 And learning, learning consists of estimating parameters of a neural network. It turns out that empirically, that wouldn't essentially work. If you would use algorithm of back propagation, you would get quite poor results. And people realized that in case of images, you might want to multiply by a little bit special matrix that also allows to do way faster computations. So you can think that neural network, as I say, apply some computation to the input. So neural network applies some computation to the input.
Starting point is 00:25:37 You might want to constrain this computation in some sense. So you might think as you will have several layers, maybe initially you would like to do very local computation. and it should be pretty much similar in every location. So you'd like to apply the same computation in the center as in the corners. Maybe later on you need some diversification, but you want to pre-process image the same way. So the idea is that when you take an image or any actually two-dimensional structure,
Starting point is 00:26:14 So the other example is you can take voice And it turns out that you can By applying Fourier transform You turn voice into image And all the It's like a two-dimensional image So like a waveform? Yeah, so you take a waveform
Starting point is 00:26:29 Yeah And you apply Fourier transform Okay And essentially On the X axis you have time As the speech Uh goes on
Starting point is 00:26:43 and on y-axis you have different frequencies and that's an image and speech recognition systems they also they treat sound as it would be an image I didn't realize that that's really cool okay
Starting point is 00:26:57 so that's why I'm saying that the technique that like also like a kind of as a side track the cool thing about neural networks is it used to be the case
Starting point is 00:27:12 that people special in processing text, images, sound. And these days, this is the same group of people. That's really cool. We are using the same method. So coming back to what is convolutional neural network, as I mentioned, you would like to apply the same computation all over the placing image.
Starting point is 00:27:36 And essentially, convolutional neural network says when we take an image let's just connect a neuron with local values on the image
Starting point is 00:27:57 and let's copy the same way it's over and over again so this way you will multiply kind of multiply values in the center in the corners
Starting point is 00:28:12 by the same values in the matrix. So an input to the convolution is an image and output is kind of also an image can think that there is also some specific vocabulary. So in this kind of three-dimensional
Starting point is 00:28:32 image is like you have height and you have also depth. So let's say in case of image that's three dimensions and then you apply convolution you can kind of change number of that dimensions usually people go to let's say i don't know 100 dimensions or so okay gotcha and then you kind of have several of these layers and then there are so-called fully connected layers which are just conventional matrices so i would say that's one of advances that actually happened 20 years ago
Starting point is 00:29:02 already another one which is it might sound kind of funny but for a long time people didn't believe that it's possible to train deep neural networks. And they were thinking quite a lot about what are the proper learning algorithms. And it turns out that. So let's say when you train a neural network, you start off by initializing weights to some random values. And it turns out that, very important to be careful to what magnitude you initialize weights.
Starting point is 00:29:46 And if you set it to right values, and I can even give you some, let's say, intuition what it means, turns out that then simplest algorithm, which is called static gradient descent, actually works pretty well. Okay. So, some sense, as I said, let's say, layers of neural network, they kind of move. they multiply input by matrices. And a property that you would like to retain, you don't want the magnitude of values to blow up
Starting point is 00:30:23 and also you don't want it to shrink down. And if you kind of multiply, if you choose random initialization, it's easy to choose some initialization that will kind of, you know, turn the magnitude to go keep on increasing. And then if you have 10 layers, and let's say in each of them you multiply by,
Starting point is 00:30:41 2, 2, 2, 2, 2. Yeah, yeah. And then the output, all of a sudden is of completely different magnitude. And learning is not happening anymore. And if you kind of just choose them, and it's a matter of choosing variance or like a magnitude of initial weights. Huh. And if you said it, starts at, let's say, output is of the same magnitude as input, and everything
Starting point is 00:31:04 works. So basically just adjusting those magnitudes was what proved that you could do this with a neural network? Yes. Oh, wow. Okay. That's kind of ridiculous that, let's say, people haven't realized it for a long time, but that's what it is. And when and where did that happen?
Starting point is 00:31:20 It happened actually at the University of Toronto. Oh, okay. So at the Jeffreys Hinton Lab. So the crazy thing is people had several schemes in terms of how to train deep neural networks. And one was called generative pre-training. and so let's say there was some scheme what to do in order to get to such a state of neural network that all of a sudden you can use
Starting point is 00:31:49 this trivial algorithm called stochastic gradient descent so there was like an entire involved procedure and at some point Jeffrey asked his student to you know compare it to like the simplest solution which would be adjusting magnetes and like a showing how big difference there is. That's crazy, man. Oh my God.
Starting point is 00:32:15 Okay, so a question that's a little bit broader is just like, then what has happened in the past, say, five years to excite people so much about AI? So I would say the most stunning were so-called image net results. So first of all, I should tell you where was computer vision five years ago. Then I will tell you what is ImageNet. Then I will tell you about the results. So computer vision is a field where essentially you try to make sense of images. Like a computer tries to interpret what is on images.
Starting point is 00:33:02 And it's extremely simple to say, oh, here on an image. image there is a cow, a horse or so. But for computer image is just the collection of numbers. So it's a large matrix of numbers and it's very difficult to say, oh, like how to, it's very difficult to interpret what's the content. And it was the case that people came out with various schemes how to do it. You know, you could imagine, I don't know, maybe let's quantify how much of a brown color. There such that you can say it's a horse. Like a simple stuff. People, of course,
Starting point is 00:33:43 came out with more clever solutions, but systems were quite better. I mean, you could fit the picture of a sky the system and it was telling you that there is a car. It's like...
Starting point is 00:34:00 So not so good. Yeah. Yeah. So then then Fayfayley, Fayley is a professor at Stanford. She, together with her students, she collected the large data set of images and the dataset is called ImageNet.
Starting point is 00:34:22 It consists of 1 million images and 1,000 classes. So that was by the time actually the largest data set of images. In a class just to clarify, being like car might be a class? Yes. So there is the data set, I would say, it's not perfect. It has, for instance, it doesn't contain people. That was one of constraints over there.
Starting point is 00:34:49 It contains large number of breeds of dogs. So that's a queer key thing about it. But same time, I mean, that's the essential data set that made deep learning. happen. Types of dogs? No, the fact that it's so large. So what happened
Starting point is 00:35:10 there was like a plenty of teams actually participating in the management competition and and I'll say even as I'm saying there is 1,000 classes over there so if you have a guess
Starting point is 00:35:25 random guess then probability that your guess is correct is essentially 0.1%. The metric there was slightly different. actually if you make five guesses and if one of them is correct, then you are good. Okay.
Starting point is 00:35:40 Because there might be some other objects and so on. And I remember for the first time when I have seen that someone made, you know, that someone created a system that had 50% error. I was impressed. Okay. I was like, oh man. It's like 1,000 classes and it can say, okay, we 50% percent. percent error what is there.
Starting point is 00:36:06 I was quite impressed. But then during competition, like a pretty much like all the teams got around 25% error rate. There was a difference by one person. They were like a, for instance, a team from University of Amsterdam, Japanese team, like plenty of people around the world. And a team from University of Toronto led by Jeffrey Hinton. on and that's like
Starting point is 00:36:34 the on team was Alex Rischewski and Ilyoszoucaver they actually got to something like 15%. So let's say all other teams they were like a 25%
Starting point is 00:36:45 the difference was 1% and these two guys they got to 15% okay? Yeah and the crazy thing is that we've been so so
Starting point is 00:37:00 So within following three years on this data set, the error dropped dramatically. I remember like next year, the error got to, let's say, 11%, 8%. I was kind of, you remember, by that time, I was wondering what's the limit, how good can you be? And I was thinking, 5%. That's like that's the best. And even there's like a human strength.
Starting point is 00:37:30 to see how far they can get if they spend arbitrary amount of time on, let's say, looking on other images and kind of comparing to be able to figure out what is there. I mean, it's not that simple for human. For instance, if you have plenty of breeds of dogs and like who knows. But let's say, if you can use some external images to kind of compare and so on, that that helps. But in the sense, within several years, people got down, I believe, to 3% error,
Starting point is 00:38:06 and that's essentially superhuman performance. And as I'm saying, it used to be the case that systems in computer vision you take a picture of sky. They were telling you, it's a car. And all of a sudden, you are getting to superhuman performance. And it turns out that these results actually are not just limited to computer vision.
Starting point is 00:38:31 People were able to get amazing other systems, let's say, speed recognition or so. So because that's like the underlying question, right? Because like it's not, I mean, to someone not in the field like me, it's not necessarily intuitive that computer vision, computer image recognition would, you know, seed artificial intelligence. So, I mean, like what came after that?
Starting point is 00:38:54 So in the sense, the, the crazy thing is that the same architectures were worth for various tasks and all the sudden that the fields which seem to be unrelated, they started benefit from each other. So as I mentioned, it turns out that problems in speech recognition can be in very similar way
Starting point is 00:39:25 you can essentially take speech, apply for your transform, and then speech starts to look like an image, and you apply similar object recognition network to kind of recognize what are the sounds over there, and like phonemes. And so phonemes are like kinds of sounds out there. And then you can turn it into text. And so that's where it went, so it went to speech after images
Starting point is 00:39:54 And then, yeah. Then the next big thing was essentially translation. Translation was extremely surprising to people. That's the result by Eliasus' cover. So translation is an example of another field that actually lived there by its own. And one of the crazy things about translation is input is of a variable length and output is of variable length and it was unclear even how to kind of consume it with neural network how to produce variable length input variable length output and and ilia came out with an idea there is
Starting point is 00:40:40 something called recurrent neural network so i mean let's say recurrent neural network and convolutional neural network they shared an idea which is you might want to use the same parameters if you are doing similar stuff. And in case convolutional network, it means let's share the same parameters in space. So let's say let's apply the same transformation to the middle of image as in the corners and so on. And in case of recurrent neural network, this is as we'll be reading text from left to right, I can consume first word, can create some hidden state representation, and then then next time step when I'm consuming next word, I can take it together with this hidden
Starting point is 00:41:35 representation and generate next hidden representation. And you are applying the same function over again, and this function consumes hidden representation and next word, hidden representation and worth, hidden representation and the word. So it's relatively simple. The cool thing is if you are doing it this way, regardless of length of your input, you have the same size of a network. And the way how his model works,
Starting point is 00:42:06 and as described in a paper called Sequence to Sequence, essentially consume word by word sentence that you want to translate and then when you are about to generate translation you essentially start emitting word by word
Starting point is 00:42:29 and at the end when you emit a dot that's end that's so cool and it was quite surprising to people by that time they got to decent performance they were not able to beat
Starting point is 00:42:42 a phrase based systems and now it's like outperform like a long time ago already and yeah the one other issue that people have so with neural network systems like in case of translation the problem with deploying it on the large scale is that it's quite computationally expensive and it requires essential And in deep learning literature, there are various ideas how to make things way, way cheaper computationally after you train it. So it's possible to throw a large number of weights or essentially turn floats, a 32-bit floats into smaller size numerics and so on and so forth. And pretty much that's the reason why things are not largely deployed in production systems out there. but neural network-based solutions are actually outperforming anything what is out there.
Starting point is 00:43:48 There are a couple more things I would like to just define for a general listener. So there are a couple words being thrown around a lot. So narrow AI, general AI, and then super intelligence. Can you just break those apart? Sure. So pretty much all AI that we have out there is narrow AI. No one built so far general AI. No one built super intelligence.
Starting point is 00:44:16 So narrow AI means artificial intelligence. So it's like a piece of software that solves a single predefined problem. General AI means it's a piece of software that can solve huge, vast number of problems, all the problems. So you can say that human is general. generally intelligent because you can give an arbitrary problem and human can solve it. But for instance, battle opener can solve only battle opening. So pretty much when we look at any tools out there, at any software, our software is good in solving single problem.
Starting point is 00:45:04 For instance, our chess playing programs cannot drive. private card. And for any problem, we have to create a separate piece of software. And general artificial intelligence is a software that could solve arbitrary problems. So how we know that it's even doable? Because there is an example of a creature that has such a property. Yeah. And then super intelligence is just, I assume, just the next step, yeah.
Starting point is 00:45:45 Essentially, super intelligence means that it's more intelligent than human. Cool. So given all of that, given that like we're basically at a state of narrow AI across the board at this point, where do you think is like, what's the current status of this stuff? Where do you see it going in the next five or so years? So as I mentioned, they're essentially machine learning. There are also paradigms. So one of them is supervised learning.
Starting point is 00:46:21 There is something called unsupervised learning. There is also something called reinforcement learning. And so far, the supervised learning paradigm is the only one that works so remarkably well that it's ready to be applied in business applications. All other are not really there. And so you ask me where we are. So we can solve these problems, other problems they require further work. It's very difficult to plan with ideas, how long it will take to make them work.
Starting point is 00:47:05 the thing which is very different with contemporary artificial intelligence is that we are using precisely the same techniques across the board. Simultaneity, majority of business problems can be framed as supervised learning and therefore they can be solved with current techniques as long as we have sufficient number of input examples and what we want to predict. And as I mentioned, the purse can be extremely rich. Output might be a sentence. And current systems work pretty well with it. And nonetheless, it requires an expert to train it.
Starting point is 00:47:57 And so then given the, given like pretty substantial hype, we see. What do you think of it all? The field is simultaneously underhyped and overhyped. So from perspective of business application, as long as you have pairs of examples, pairs from like that indicate mapping, what's the input, what's the output?
Starting point is 00:48:28 It's, we can pretty often get to superheaval. human performance. But in all other fields, we are still not there, and it's unclear how long it will take. So give some example, let's say for recommendation systems, you have often companies like Amazon, they have examples of millions of users, and they know what they bought when they were happy or not. And that's an example of a task that is pretty good for neural network to learn what to recommend to new users. Simultaneously, Google knows what is the good search query for you because on the search result page,
Starting point is 00:49:16 we are clicking on the links that you are interested in, and therefore they should be displayed first. and in other fields it's actually quite often more difficult. In case of, let's say, Apple picking robot, it's difficult to provide supervised data telling how to move an arm toward Apple. Therefore, that's way more complicated. Same time, the problem of detecting where Apple is, it's where better defined and can be out there. outsource to human to annotate plenty of images and to give localization of the apple.
Starting point is 00:50:00 And quite often the rest of the problem can be prescripted by engineer. But the problem of how to place fingers on an apple or how to grip it, it's not well scientifically solved. And so we have a couple questions then at this point. If people were to be interested in learning more about AI and maybe working with Open AI or doing something, how would you recommend they get involved and educate themselves? So, let's say, a good place to start is Coursera.
Starting point is 00:50:45 Coursera is pretty good. There is also a lot of TensorFlow tutorials. TensorFlow is an example of framework. to train neural networks. Okay. Also, Andre Carpatti's class at Stanford, it's extremely accessible. You can find it on, I believe, on YouTube. Yeah.
Starting point is 00:51:10 And then, like, in terms of actual exercises? So in case of TensorFlow tutorial, many of the problems so I believe in case of Andre's class there might be homework and in case of TensorFlow exercises it's quite often easy to come out with some random thought after let's say
Starting point is 00:51:36 reading like I mean you can take for instance let's say like the simple task over there is let's classify let's classify digits and let's classify pictures of digits. Let's assign them classes. You can try maybe download some images from some other source, like a flicker.
Starting point is 00:52:03 Let's try to classify it toward tags. Okay, so given that you guys are working on with robots at this point, one of the other things that's thrown in, like, kind of part and part. with AI is automation specifically of like a lot of these low level blue collar jobs. What do you think about the future, maybe the next 10 years of those jobs? So I believe that we'll have to offer to people basic income. I super strongly believe that actually that's the only way. So I don't think that it will be possible for 40. years old taxi driver to reinvent himself every 10 years. I think it might be extremely hard.
Starting point is 00:52:58 Other crazy thing is people define themselves through job, and that might be another big social problem. Simultaneously, they might not even like their jobs. Like if you ask someone, would you like your kid to sell in the supermarket, to be a seller in the supermarket, they would answer no. And maybe it's possible to live in the world that there is abundance of resources and people can just enjoy their life.
Starting point is 00:53:45 I think we're going to have to figure out a way. I mean, maybe people will always find purpose, but I think making it easier to find that purpose will become much more important in the future if automation actually happens to the degree people talk about. And what about just like influences on you that maybe have inspired you to work with robotics
Starting point is 00:54:08 and in AI? Are there any books or films or any media that you really enjoyed? It's pretty good book called Homodeus. actually describes the history of humans and then speaks, then has various predictions about the future or where we are heading. That's one pretty good.
Starting point is 00:54:40 I mean, nowadays, there is like plenty of movies about AI and how it can go wrong. What's the best one? I think hair is pretty good. Okay. Yeah, X-Machina is also pretty good. Cool. All right, do you have any other last things you want to address?
Starting point is 00:54:58 Oh, I think, no, thank you. Okay, cool. Thanks, man. All right, thanks for listening. Please remember to subscribe to the show and leave a review on iTunes. After doing that, you can skip this section forever. And if you'd like to learn more about YC or read the show notes, you can check out blog.
Starting point is 00:55:14 commodator.com. See you next week.

There aren't comments yet for this episode. Click on any sentence in the transcript to leave a comment.