Devreal

Computer Vision for Coders

Event: AI Vision

ai.bythebay.io: Jeremy Howard, Computer Vision for Coders

Recording: ai.bythebay.io: Jeremy Howard, Computer Vision for Coders

Okay, hi everybody. So I'm going to be talking to you about learning deep learning. And I've got a little picture here which was from our most recent class where we learned about drawing interesting pictures of cats. One of the many useful things you can do with computer vision. Before we talk about cats though, I'm going to tell you about our journey to creating this course. So I've been involved in and around machine learning for about 25 years. And throughout that time I've been very much on the application side. I don't have any technical degree myself, but I spent nearly 10 years in management consulting focused on the more analytical side

And then I've built a number of startups that will heavily use machine learning. More recently I was the president and chief scientist of Kaggle, which was obviously in the news today. Built a medical diagnostics company called analytic. And throughout this there was, there's always been a lot of things that I wanted to see improve. In terms of how machine learning generally is dealt with. And these things have become a lot worse with the advent of deep learning. I've really seen deep learning portrayed as something that's a very kind of exclusive thing. It requires lots of data and lots of GPUs and lots of math and specialist languages and so on and so forth

And I knew from my experience that really most of this was not true. For example, at analytic we built a state of the art lung cancer malignancy classification system on a single GPU with four people, none of whom had a PhD in machine learning in about a month. So I really wanted to find a way to make deep learning, bring deep learning to the folks who don't have the machine learning PhDs but who are making the world run every day. One of the big issues is around education. If you want to use any technical subject you need to learn it. And deep learning has been pretty heavily in the math world. And there's been a long history of terrible math education in the West. Perhaps as best described by this fantastic mathematician Paul Lockhart

Who describes it as a system that seems expressly designed to be as soul crushing as possible. And I certainly sometimes have seen that in deep learning educational materials. Just like other math materials. Specifically the kinds of problems that we saw. And when I say we I'm referring to my co-founder Rachel Thomas who actually does have a math PhD. And I is that what we've seen about successfully applying deep learning is that really it requires a code centric approach. Generally speaking people who apply deep learning successfully do a lot of engineering. And a lot of experiments and write a lot of code

But a lot of the teaching has been math centric. We've seen like in most technical education that it suffers from what the fantastic education researcher David Perkins calls elementitis. refers to this idea that you start by learning probability. And then you learn statistics. And then you learn information theory and blah blah blah blah. And five years later you write your first piece of deep learning code. So we really wanted to turn this upside down. And use what David Perkins refers to as a whole game approach to teaching

The whole game approach to teaching is like the way you learn baseball. You don't learn first of all the physics of a parabola as you throw a ball and techniques for winding balls. And using a lathe to create a bat and five years later you see your first game of baseball. No you introduce kids to a very simplified game of baseball but it is a complete game. And you start playing and at every point in that journey you're playing baseball. So we really wanted to set this up so that from the very first lesson people were solving useful problems with deep learning from the very start. The other thing I've seen again and again with deep learning teaching materials is that they tend to leave you at a point where they say okay there's a way of dealing with object classification and computer vision. But the way they show you is far far far short of the state of the art

It's something which doesn't even really scrape by. So we wanted to teach people from the very start how to create state of the art results and provide the tools to do that. So these are some of the things that we had in our head as our goals. So we looked around a lot to see kind of what educational materials are out there. And we just didn't find much. In fact a lot of the schools required a PhD even to get admitted to you know things like data insight or the deep learning summer school. And so that definitely didn't really count as inclusive to us. So we actually set up and paid for out of our own pockets diversity scholarships

We set up an international fellowship. So we've had folks from Ivory Coast and Nigeria and Bangladesh doing things like creating the first ever corpus of Urdu word vectors through to studying domestic violence against women in Bangladesh. To looking at the use of gendered language who's an English literature PhD who's gone through our program. So looking at all these things we wanted to do we realized the only way to do it was to set something up ourselves. So we set up a research lab called Fast AI and we set up a course. And we were lucky enough that USF agreed to actually put on a course in person. And they were fantastic. They basically let us do whatever we like

So we curated this fantastic group of interesting people. Now many of whom are in the audience today I can see. And we then recorded all of these courses. And then at the end we put it online. Did a lot of, created a lot of materials and so forth. And turned it into a MOOC. And so the whole thing is, is free. There's no advertisements

And so there's now tens of thousands of people have gone through and learnt deep learning. So we called it practical deep learning for coders. Which is not really where we wanted to end up. You know deep learning for coders means there's a prerequisite here which is you have a year of coding experience. That was just the best we could do. We just found with the current state of technology we weren't able to create a practical, deep learning for non-programmers. But that's our goal. is to eventually get to that point

So one of the, so we kind of looked at how are people currently being told when you look on Hacker News or something. And somebody says how do I learn deep learning. And inevitably somebody will say oh look at the Ian Goodfellow book. And so if you look at the Ian Goodfellow book you'll see things like this. Which says to gain some intuition into how backprop through time behaves. Here's an example. So if this gives you an intuition that's great. You're smarter than me

Even my partner who's a math PhD did not find this as intuitive as it might be. Which is not to say anything about the book. The book's actually fantastic for people who want to study the theory deeply. And become an expert theoretician. But what we did was we created something where in class one we showed this. And so we showed six lines of code. And if you run these six lines of code you will get a state of the art result in image classification. And for a very wide variety of image classification problems

So one of the tricks to this was that we spent many months creating the course. And creating the course was not just to say okay what's out there. Let's teach it. But it was to say every time there's more than half a screen full of code to do anything. We would keep working and refactoring and simplifying until we built something that was that simple. So the library that we're showing here is a fairly thin wrapper over Keras. So Keras was absolutely critical to us in being able to deliver this you know really beginner friendly approach. But even Keras didn't go far enough

So we created like these simple little wrappers where with this now with these six lines of code. As long as your images have been put into folders. So you might have a folder called dogs and a folder called cats. And you run these six lines of code. And it will download a pre-trained ImageNet network. It will customize the ImageNet network for your folder structure. It will retrain the appropriate layers. And we'll spit out something which would have gotten in the top ten of the Kaggle dogs versus cats competition

And so that was generally our goal. Was to say with a tiny bit of code this will get you in the top ten of this competition. So we tended to use Kaggle competition data sets a lot. Because one of the things we found again and again was that academic papers like the number of academic papers that claim to have a state of the art result on sci-fi 10. It's not physically possible that a thousand papers have the state of the art result on sci-fi 10. One paper has the state of the art result. But the problem is the details of the engineering of how you train something is so important in deep learning. That those results in papers are nearly meaningless

Whereas you look at the result of a competition and you've got thousands of people competing against each other to do the very best they can to get everything right. So if you can give a student something that can get top ten in a Kaggle competition then you know that you're really giving them something that works truly well. So that was our goal. It was very interesting to hear the feedback that we got from students. So the original set of students that came through USF was about a hundred people. And they were from all walks of life. Maybe about half of them had PhDs but none of them had PhDs in machine learning. They were PhDs in English literature or neuroscience or space engineering or whatever

And about half of them were more like entrepreneurs or engineering VPs from big companies or whatever. And we tended to hear this basic comment again and again when we talked to people later in the course. Which was it was so easy for people to go back to the math textbooks. You know to go back to studying the theory. And what we've heard again and again from people is that the thing that actually made them better deep learning professionals was writing code. And specifically writing code doing lots of experiments. Like trying things and find out what works. So that would be I would say the most common feedback we've heard from students

And even though the whole thing was called practical deep learning for coders. We actually had a rule that was for part one we would never show a mathematical equation. Like nothing was dumbed down. But every time we had math we would show you know working numpy code. So despite all of that many many students found that they just had these habits of kind of going back and studying textbooks rather than writing and running code and learning from experiments. So one of the promises I said in my little abstract this talk is I would give some lessons learned. From this experience of teaching initially a hundred people in person deep learning from scratch. And then since that time maybe 40,000 people online

This would be our biggest learning. Some of the things that I think has made this the most successful is about getting creating a community. And getting that community involved. And one of the most important ways was that from the very start we set up a forum. For those of you that haven't used discourse before it's this fantastic forum software. It's what the PyTorch discussion forums for example use. And if you go to the fast AI forums today you will see you know dozens of posts you know in the last few hours. There's lots and lots of people now talking online about their experiences of learning deep learning

And they've helped us so much. You know this community now has become self-sufficient. We've created moderators from within that community of students. Some of the students who finished the course earlier are now helping the other students who are going through it from the start. And explaining some of the things that they find unclear. So this has been critical for us in trying to make deep learning accessible. Is to try and create a self-sustaining community. So these forums have been an important part of that

Another really important part of that was setting up a wiki. And allowing our students then to create their own learning materials. And so we encourage people as they see questions that pop up on the forum. To then like move that into the wiki. And so we now have a huge you know like hundreds and hundreds of pages of learning resources. That we and the students together have created. So for example every lesson has many pages of prose describing what's happening in that lesson. Links to additional information and so on and so forth

The kinds of things that our community have done for us have been surprising and wonderful. For example one person went through and created a timeline for every single video. And so now that's really great. When anybody is going back and revising. They can go back and click on one of these links and bang. There they are looking at that part of the video. Somebody else transcribed every single video. So we now have captions

And so for people who are studying where English isn't their first language. Most of those people much prefer to look at subtitles. So we now have those subtitles. And of course it also means that we're compliant with accessibility legislation. So the community building that community was a key goal for us. And we're thrilled with how well that's worked. So the next thing I thought I would tell you is. For those of you that are interested in building your deep learning expertise

What are the kind of five key things that we ended up having as the five key takeaways for our students. And I haven't seen these five things really written down anywhere before. But I think they're kind of critical. So here are the five. So the first is this basic insight of if you take non-linear functions that are differentiable and stack them on top of each other, you can solve pretty much any predictive modeling problem. And this is an insight which we repeated again and again during the course. But it took repeated examples of people seeing it to basically see like, OK, today we're going to learn to recognize cats versus dogs pictures. OK, tomorrow we're going to learn to recognize positive versus negative movie reviews

OK, today we're going to do collaborative filtering for product recommendations. And you know, each time we would learn something new. And each time the vast majority of the code would be identical. You know, it's like, OK, we set up this Keras network. We have these layers. We call compile. We call fit. We call predict

There we go. And so for a long time, a lot of the students just, they kept on asking, how is this working? Even though they'd learnt SGD from, you know, stochastic gradient descent from scratch. And they'd learnt about all the different variants of SGD. And they'd learnt how convolutions worked. And they knew all the pieces. The idea that this incredibly simple, single thing could solve all of these problems. It was just hard for people to really intuitively believe that and feel that. And so by the end of week seven, I think they had seen so many different kinds of problems solved with this exact same approach that people were starting to be like, OK, I see

It really is just taking simple stacks of nonlinear functions, chucking them through SGD, and it works. As long as there's enough parameters. And then the second key insight was that we told our students, if you can, always use transfer learning. So transfer learning refers to taking a model that somebody else has built and reusing it with minor changes. And so those minor changes are generally twofold. The first is generally removing the last layer and replacing it with one that solves your problem rather than their problem. And then the second is running a few epochs of SGD again to update the weights so that it solves your problem rather than their problem. To give you a sense of how powerful this is, at my medical diagnostics company, when we looked at lung cancer diagnostics, we took an image net network, so a pre-trained network that had learnt to recognise dogs from cats from jumbo jets in full colour 2D images

We then used that as a starting point for looking at CT scans of lungs. And we transfer learnt from the image net network into something that gave us a state of the art result. In fact, beat the world's best radiologists at recognising the malignancy of lung cancer. So with transfer learning, it's not even that you need somebody else to have solved a very similar problem. It really needs to be just a problem very vaguely in the same domain. A similar approach is very popular in natural language processing now, which is anytime you have a word, replace it with a word vector. And you can download from the internet lots of word vectors, which are basically pre-trained distributed representations from big models. So this is a kind of critical insight

And this insight is also critical for particularly for start-ups, but for any company that's investing in deep learning, is that you can kind of build this ongoing knowledge base of models that are stacked on top of models that are stacked on top of models. And so over time, it's kind of like creating sourdough bread, you know. You've got these sets of weights which have come from so many different areas and they became more and more generalisable and they become more and more powerful. The third thing that we then looked at is, okay, a stack of differentiable non-linear functions can solve pretty much any problem, but some functions solve it better than others. So what are the architectures which allow you to solve particular types of problems more quickly or less quickly? So we spent a lot of time on CNNs and a lot of time on computer vision. And the reason for that really is that our course name is Practical Deep Learning. And so we wanted to show techniques that work right now. And like as we really looked into it, a lot of the stuff that's being hyped doesn't really work in practice

You know, like chatbots, for example. Deep learning for chatbots, very popular area, it's rapidly developing. But, you know, you go and actually look at Facebook that has perhaps the most advanced deep learning lab in the world and you use their chatbot and it's totally crap, right? And that's because that, you know, that area of technology is not practical yet. Computer vision on the other hand, very, very much it's. fantastically successful. So we spent a lot of time looking at computer vision and convolutional neural networks are a fantastic architecture for that. We spent a reasonable amount of time looking at recurrent neural nets. I think interestingly though, the students kept on asking us the same question with RNNs, which was why are you teaching us this when it seems that CNNs tend to do better? And like quite often we didn't have a good answer for that

So even as we were teaching using RNNs for natural language processing, that week Google came out with the WaveNet paper showing that CNN actually is the new state of the art for natural language processing. So I think that's interesting. I think it's, it's some, you know, we're still teaching RNNs, but our focus is on convolutional neural networks because more and more it just seems that they're becoming the state of the art, even in areas that RNNs have traditionally been more successful at. And then of course we learned about different types of activation functions, such as Softmax and ReLU. And you know, so if you know these basic pieces, CNNs, RNNs, when to use what kind of activation functions, you're pretty much know what you need to know about building an architecture that's going to make your model train faster. So it's very, very well to train a model. But then the key thing, next key thing that we needed to teach people is the idea that there's no use training a model if you then try and use it on a different data set and it doesn't work anymore. And so that obviously tends to happen because of overfitting

So we taught a really simple five step process. And the process is basically this. Start off by overfitting as much as you can. Like if you can't build a model that totally overfits and gives 100% accuracy, then you haven't found a functional form that's flexible enough to actually solve this problem. So you start off with no regularization, no data augmentation, you know, and just try and overfit the hell out of it. And once you've done that, you know that you now have an architecture that can solve this problem. And then you can gradually add in these five steps. Right? So you can add in more data

You can then add in data augmentation, which is like faking more data by flipping things left to right or rotating them or changing their colors. You can use more generalizable architectures like batch normalization. And then if you have to, you can add regularization, such as weight decay or dropout. But any kind of regularization is basically adding noise and therefore removing some predictive power from your network. So we really try to say that's one of the last things you should try. And then if you're still overfitting, then finally you can try to reduce the complexity of your architecture, most probably by reducing the number of activations or the number of filters that you have. So another piece of feedback that we got again and again as our students tried to get to the top of more and more Calcul competitions was, oh my God, this approach actually works. So, you know, people tend to dive in by using what they think is going to be the perfect approach, based on their kind of experience and intuition

And then it doesn't work and then they spend a week debugging it and try to figure out why. And eventually in frustration they give up and they actually listen to the teacher and they throw away all their regularization and they try this approach and we heard again and again from students is, okay, this actually really can't go wrong. So this is a great system. One minor issue but super important to recognize is data augmentation is something which a lot of people get wrong. So like what kind of data augmentation and how much? And there's a really simple way to do this, which is if you think about it, the correct amount of data augmentation and the correct types of data augmentation will not change depending on whether you're using a small sample or the whole data set, right? If horizontal flipping is useful, it would be useful for a small sample, just like it would be useful for the whole data set. So what we taught in the class was this idea that you always start with a small sample and you just write a bunch of loops that figure out which types of data augmentation actually work. So forget intuition and guesses and whatever else, you know, how much rotation if any, run experiments to find out. Which kinds of flipping to use, what types of color adjustments to use

By doing it on a small sample, it takes like an hour to run all of those experiments and then you know. And then you can scale up to the full data set. So that was a handy technique that we showed people. And then finally the fifth big takeaway is about the power of embeddings. So embeddings refers to, embeddings are generally, have generally used for things like replacing a word in NLP with a vector. Which is kind of, has all of the information that's in that word. Or in collaborative filtering, the idea of latent factors. So replacing a person with a list of basically an embedding

Or a movie that's being recommended with an embedding. It turns out that much underappreciated is that you can use the same embedding approach for any kind of categorical data. So this is being very rarely written about. But there's at least two examples of Kaggle competitions that have been won. By people looking at time series and structured data problems. And winning them by taking the categorical variables and replacing them with embeddings. And so using embeddings as much as possible was another big takeaway. So that was last year

And so we're now doing something, kind of taking it to the next step. Which is asking those students to now come back and move from practical deep learning. Which is basically here are all the things that deep learning does really well right now. And here's how to do the best practices. To state of the art research. So we're now doing part two of this course. Where the students have come back. And are learning what's basically the edge of research and deep learning

So there's a number of changes that we made. One of the things we did in part one was we intentionally used Python 2. Because so many of the basic tutorials for people that haven't done that much programming before. Tend to be in Python 2. But then for part two we're now moving to Python 3. Because Python 3. You know we're now teaching things like parallel programming. And you know there's a lot of stuff that basically Python 3 makes easier

So that was one change we made. One of the big reasons also was that IPython announced that they're not going to support Python 2. In their next version. So given that all of our stuff runs on top of IPython or Jupyter Notebook. And that was really important as well. Another change we made was the choice of library. We still use Keras. But in the first part we used Keras on top of Theano

And that worked really well. Because Theano just couldn't be easier. As a back end to Keras you really hardly have to notice it's there. And when you do have to dig in and customize things it tends to be super easy. We've now moved to TensorFlow for the second part of the course. Which honestly is a lot more verbose. It's a lot more complicated. A lot of the defaults are kind of ridiculous

But it is also fantastically powerful. Particularly when you want to put stuff in production. You can use TensorFlow serving. And bang. You've now got yourself an API. So that's been a couple of big changes we made. Another big change we've made is to introduce PyTorch. For those of you who haven't used this yet

Don't worry. You're not that far behind. It's only been out for about six weeks. But it's absolutely extraordinary. The PyTorch discussion forums are one of the most active deep learning communities on the internet. Or perhaps the most active already. And it's what's called a dynamic system. Or a defined by run system

Which means you don't have to set up a whole computation graph ahead of time. And then run it. You can. It just looks like normal Python code. So it's much easier to write. And it's much easier to debug. And as a result it tends to be faster. Because you can spend more time on writing a optimized algorithm

So we're using PyTorch increasingly. One of the things that was really important for the students was to make it easy for them to actually use GPUs. Deep learning is a pain in the ass if you have to use CPUs. But getting GPUs up and running traditionally has been difficult. And a lot of tutorials assume a whole lot of background about either Docker or AWS or whatever. So we set up basically a bunch of scripts where you run a script and it does everything. It sets up your subnet. And it sets up your private key

And it sets up your server. And it provisions it. And it logs you in. And so making it really easy for everybody to be able to use GPU servers has been critical. And a lot of people told us that they were kind of terrified of this part. You know, because a lot of these people they've like been doing a .NET line of business apps. Or they've, you know, maybe they've been riding Kobo most of their life or whatever. And the idea of like all this AWS stuff and Bash and whatever else seemed terrifying

But the amount of pride and excitement we heard from people when they ran it and they logged into the AWS instance, it was fantastic. It was, you know. So now the whole class, everybody is, you know, up there using AWS and it's been great. So for part two, we're going further and actually trying to encourage people to build their own box. So there's some great forum threads now as people, you know, going through the same thing. Like, holy shit, I never thought I would build a computer. And we're hearing stories and people are writing medium posts and so forth, describing how they've built their own box. And, you know, they're very proud and very excited as they should be

Also in part two is we actually start writing papers, right? So there becomes a point where if you actually want to move from practical, you know, best practices to cutting edge research, it doesn't live anywhere else other than in papers. So we're kind of trying to now say to people, all right, this is how you can read a paper and understand it even if you don't have a math background. And that actually seems to be going pretty well. A lot of people who have never read a math paper before are now, you know, sharing discussions about what they're finding about new papers that they're looking at. So overall in part two, we've kind of moved from this, you know, best practices, very computer vision focused, now to more of a generative models focus. Because a lot of the current research really, when you look at it, is focused on generative models. So things that the output is itself a picture or the output is itself a sentence. Captioning, question and answer systems, artistic style, colorization, so forth

And so it turns out that, you know, this area of deep learning really is much less mature, but it's pretty exciting. And we do think that people are going to need to understand it increasingly in the coming months and certainly the coming year. So this is one of the focuses of our teaching now. And we're also explaining much more how to use much bigger data sets. So we're explaining how do you process the whole of ImageNet. And it turns out that there's very little, if any, information out there that takes you from step to step through, here's how to create a whole ImageNet model. And there's just a lot of little details, little engineering details you have to get right. And then finally, we're showing people how to move beyond kind of just the classic NLP and computer vision type areas to looking at structured data sets like fraud data sets or credit data sets, and also looking at time series data sets

Things which, you know, very little of the academic community in deep learning have spent time on. But for, you know, real world problems, this is actually where most people spend most of their time. So it's been a really exciting journey. I'm really happy to be able to say that I think the experiment's been successful. You can take people who have nothing more than a year of coding experience and teach them deep learning in seven weeks. We ask them to put in about 70 hours of work in total. And people who have done that have come away with some great skills. So thanks very much for listening

And if you're interested in checking it out yourself, it's course.fast.ai. Thanks very much. It's a great program. So I wonder if you ever have to do very specifically targeted clients, like for a specific company, and they have a very specific area and you tailor this for that. Yes. So we've been asked about that. And in our current course, quite a few big companies you'll be very familiar with have sent teams to the course. Because what I've said to them each time is, okay, if you need a custom course, I'll do a custom course

But first of all, do the non-custom course and see. And because my guess is, and seems to be being bad at borne out, is that almost nobody needs a custom course. You know? Particularly because in the course everybody has a capstone project, which is often something they're bringing in from work. And so they can, you know, work with us and the fellow students on that project. And so they kind of learn about it as they go. So, yeah, so far there doesn't seem to be too much need for custom courses. But if we find there is a need, we'll, we'll do them. Great

Thank you. Once again, for what we have a lot.