Cognifest NYC 2017: Krishna Bala, Why Deep Learning Works
Recording: Cognifest NYC 2017: Krishna Bala, Why Deep Learning Works
[Music] these learning works but before that let me tell you a bit about Metis Magnus is a data science training home and you get skill you get connection and you get hired that's that describes our bootcamp essentially which is an intense 12-week experience and it's a lot of fun for those who like that and for those who prefer that are also evening classes that cover aspects of exercise so with that why deep learning works the short answer is I don't know so I'll take you through what does deep learning nothing reasonably quickly in and maybe point you some references that might take you further the feet forward step and then how you do back propagation to learn and then we start looking at questions about widely bloody works it works immensely well I think you know building simple models with carrots and so on and they're just so easy to put together and you get really good strong models and that the method seems so simple so it's really puzzling you know what just why does it work so much you could be caught somewhere in something you know complex optimization minimum and instead you get good answers so so we look at the energy landscape which is you know starting to explain why it might work then we look at another approach by some physicists tying it to the renormalization group which helped explain phase transitions we'll see how that goes and then and then there's some more recent work which sort of describes the flow of the learning process by dish being and associates so we'll look at that as well so that's three different perspectives right none is guaranteed to be right they explain different features of the problem and we'll take it from there so what is deep learning we look roughly at a feed-forward Network meaning what all connections are forwarded to the next layer right the layers nodes built into a layer and they all connect forward and we typically look at fully connected layers where you know each node is connected to every node in the previous layer let's simplify some of them the mechanics so so there's a bunch of inputs X I coming in at the top and what's called the input layer and each of these hidden layers has a bunch of weights that operate on the excise it's a selenium or affine transformation and then there are these nonlinear activation functions that come in and that non-linearity turns out to be essential to learn complex boundaries and things like that and basically that non-linearity applies to the linear operation on the X's and the beat just means that there are many hidden layers that's deep learning in one slide the activation functions typically looked like we saw the logistic sigmoid represented in the previous time slide there's another patch which you know goes from minus 1 to plus 1 at plus infinity there are some more recent ones that that are easier to compute with all right and they have roughly the same behavior they the point is that the sigmoid and thank functions because of the nonlinearities are harder to compute and they also tend to saturate so that saturation causes problems where the network server sits and does nothing so people experiment with introducing different kinds of activation and these are some of those that were tried so the the first step of the network is to set the inputs right and then pick typically a small batch of inputs and apply you know random weights throughout the system randomize the weights at all the nodes the W's that we saw and then what they do is they apply the weights that the equations that we saw at each layer and compute forward right so given an input you can compute the result of applying those random weights and the activations all the way forward right and then when you reach the end of the network layer the output layer what you do is then you you compute either a classification or regression error at that point and now our purpose so far as we've you know set an input and computed an error now we won't go back and see how that error can be reduced how should we modify the weights to reduce the error and that's the major learning step which and that's TV so yeah after this slide sort of steps through the lost computation which is essentially when you apply the weights and apply up apply the weights on the input and include the biases and then apply a sigmoid to them you end up here and after that you compute the probabilities using as a softmax function and in this case what happened is you're dealing with a three class problem here and the correct class happens to be a blue so it sort of guessed wrong the highest probability should have been a green so you get up sort of higher loss in this case so once we have the loss we try and figure out how we take that loss and back propagated to to figure out how to adjust each of the weights and the basic process is simply that starting at the end at the output layer you have the lost computation you can see that this error depends on whatever is defined in this layer so if I take all the weights in this layer I can completely represent the change in this error and similarly with any of the layers if I want to determine how this particular weight impacts this layer all I have to do is introduce its derivative with respect to all the layers along the way right so what what you do essentially is compute the gradient of this loss function with respect to each of the weights on each of the layers using the derivative chain rule once you have the derivative you have to figure out you know how you take a step in that direction how big your step should be and there are a number of ways of figuring that out essentially they all stem from gradient descent or some sort where the aim is you compute the gradient and you take a step in the direction opposite to that the gradient this is a slide that shows a number of different optimizers that are used and they perform differently some faster than others and the whole aim is to get quick convergence so then we get into asking you know what what does the loss function landscape look like which is why does gradient descent work why doesn't it get stuck in all these valleys in you know that arise in the in the landscape of the deep learning problem there are so many weights and all of them interacting through nonlinearities we would get some complex surface like that why don't we get stuck why do we get good answers a few people have looked at this one way of looking at it is since since the space is too complex to take in right the idea that in Goodfellow and colleagues hired was to was to look at a linear path from the starting point to the ending point right and look at all the you know non convexity x' and things that happen along the path right and when they did that they found typically that the part was was pretty smooth if you take the starting set of weights say W initials and then you run it through the training through a bunch of epics backpropagation and so on and you arrive at an ending set of weights W final and then you plot a linear path from W initial to W final right they found that it was pretty smooth so so the question then was what what's happening I mean you can see that if you plot the actual gradient path outside the linear part it's it's sort of wobbles but it sort of stays close to that the main linear fold from the beginning to end so then they looked at the training and validation error on this happens to be on the M this data set with with max out and what they found was that the the linear interpolation that they did right gets the energies very close to the full gradient descent except of course that after a certain point it takes off right so the basic idea is here is that when you start training you see some non convexity at the beginning but after that initial period it looks like a normal convex function and you're descending you know easily without hills and valleys or so it seems so that that was a big find and they they've tried this with different models and you know found similar results with some variations so they have some theoretical backing as well that was developed earlier with spin glass models and and and so on and so their experiments sort of confirm the earlier you know on sources that that there are saddle points that at high values of the cost function I mean the cost the cost function is I guess very high dimensional so if you have some kind of an optimum the odds are it's a saddle point right because I mean for a minimum you need all the curvatures to be like positive for a maximum you need all the curvatures to be negative with so many dimensions in so many curvatures involved the odds are you'll get some kind of saddle point right and that turns out to be true at higher energies and sort of less so at low energies because in low energies you have to have the right curvature to to be a low energy so so what they found was that the minimum are tend to lie at low values of J that are close to the global minimum in in a few different model experiments and and these minima turn out to be pretty good and had pretty good test error so the reason they say you don't get caught you know in some high up minimum that far away from the global minimum is because there aren't so many out there especially when your model is large and there's a further discussion by mr. Martin who who suggests that it looks like a spin class model a spin glass funnel used in in protein folding and I haven't really appreciated that analogy so the other argument said that came up is the idea of using the renormalization group and what happened was meta and Schwab found an exact mapping between Oh a Boltzmann machine and a kind of renormalization group calculation the normalization group calculations were used in physics to try you know solve problems involved involving multiple length scales essentially so if you have if you have a small model of interacting particles then you combine them and sort of make bigger globs that interact slightly differently and so on and if the interaction form is the same you can repeat that process and ask you know where does your model end up right and and that turns out to give interesting properties and it turns out that near phase transitions these models are self-similar so you can actually plot the trajectory is you know analyze the transition and it's thought that similar effects might pertain here which is in in some sense you'll get similar self similar models in certain cases in which case you can analyze them so then the the idea is that you know as you do cause grading of particles into globs in in renormalization group the deep the deep network actually does the same from layer to layer you know you have a huge number of features and the next feature course grades them and and picks up something and so on and you go down this just shows examples of their calculation or nice Ising model and you're comparing it to a restricted Boltzmann machine there's no major takeaway I guess except you have coarse-grained features and but I guess then the main finding was the actual mapping right that there is an exact mapping between a restricted Boltzmann machine and a renormalization calculation so then we look look at a third model which was presented by Schwarz even dish B in 2017 what they did was they analyzed the the mutual information between the layers and the input and the layers and the output so looking at the bottom I guess it's the number of bits of mutual information and on the y-axis we have the mutual information of any particular layer with the output and the colors are the layers are the reason you have many points as they did many different rounds to get you know reasonable statistics so then what do you see is I guess maybe the move we might say better so you see these layers sort of move up as their information about the output increases and then they start to drift left which means they're losing information about the input what is considered irrelevant information so that was a fascinating finding and they have some explanation of this process where where initially they think it's it's it's it's a learning phase up to that point where then it starts diffusing left slowly right and so they're able to explain why the diffusion works much better with with you know deep layers with many layers because you have to diffuse less at each step because you have many layers so it's it's faster learning with many layers which is observed sometimes so so it's it's it's a fascinating result so again you can see the sort of move up and to the left as the number of epochs go they also point out that this is just this the mean and the standard deviation of the weights they point out that the weights are initially sort of you know I guess moved and then later they become noisy alright so so there's an introduction of noise starting at about you know five hundred eight books or three hundred epics and and that sort of is part of the forgetting process or forgetting the inputs so it's a very appealing theory and they construct some I guess compression metrics the idea is that you you learn first initially and then you sort of compress your learning to reduce the number of bits so that's fascinating and with that Island statistical mechanics can still be useful would be fun to understand how it all works yeah I don't believe so I think they've tried models with regularization and not and I I don't I don't think it does yeah ever very wide Network where they're supposed to be the question for that recording yeah so the question is whether when I say large networks I mean deep networks or white networks and I think it's both I mean both would I mean if you had it was yes they are if you have a large number of variables I guess we don't run into these problems so the variables can come from the width of the day but my guess is that there's probably some limiting cases so I'll battle and then if you have a single layer so I think they might be that's not a sense I got that's why I use the word large rather than people instead work out to employee conducts a vacation that that's by using different kind of norms next initially Wow I mean you want some complex energy landscape because we're building pretty complex models here's an inside straight for example in seismic methods recently they apply something to basement optimal transport like Worcester spinning distance using the different norm connects okay yeah even solver before I don't I don't see why not on it but stochastic gradient seems to be working so I'm not sure how much better it will do because ultimately the thought is if you find the global minimum you might be over learning oh yeah I'm sure someone will devise another thank you [Applause] [Music]