Scale By The Bay 2021 : Lukas Biewald, Reproducible Machine Learning at Scale
Recording: Scale By The Bay 2021 : Lukas Biewald, Reproducible Machine Learning at Scale
thumbs up all right all right um i will tell you it is a skill that i am still struggling with to um give these talks out into the ether like this i have no idea what you're thinking or um or seeing so i'm gonna do my best um you know my my goals here is to kind of talk about three things that i think are important topics you know kind of cover what's actually happening out there in the real world in ml today because there's so much hype um and i i think that what's important to me is making ml actually work um you know secondly um the difference between um engineering and um machine learning engineering because i just think about that a lot and then also wanted to talk about at uh suggestion of the the organizers talk about building products um you know for ml teams and how i think about that so you know briefly kind of introduce myself in my career and my biases i think these are relevant um i came from um a i started a company called crowdflower figure eight back in 2008 and you know back then when you would try to raise investor money you would hide the fact that what you were doing um was related to machine learning machine learning sounded like a science project um and so um i ran that company for about a decade and then i started um weights and biases which is what i'm working on now which we think of as a developer first ml ops platform i hope some of our users are out there in the audience today and i hope that um if my talk is boring you go to wnb.com which is our website and and try a product which would make me even happier than you listening to this talk i'll say one more thing um where a lot of this is coming from is you know we've been um recording this podcast with a lot of people working on real world machine learning um projects for the last three or four years and this has been a real lesson for me to get to see in lots of different verticals um what's relevant to people in the um in the machine learning space so um check out gradient descent if you if you have the time i thought it might be fun to kind of talk about my very first um uh machine learning stack that i built because really my first job um was essentially the um the ml engineer and the ml ops person and all that together working for a company um called yahoo which was um you know super exciting back in 2004 2005 when i was there and um the goal was to make the search engine results better and at that moment we were switching from a rule-based system to decide the the search results um to a machine learning system and i was actually you know the person that deployed um the code that that um you know ranked the search results for um a lot of the the yahoo properties and i was kind of reflecting back on what my stack looked like which actually is not that different um than than what you'll see today so i don't know if there's any r users in the audience uh but back then you know we would trade our models in r typically um and we use the gbm library which is actually still around um there's graded boosted machines but then you know there was no way to actually get that into um production automatically now there's tons of companies that do this back then what we would literally do is we would i wrote code that would take that r stuff and generate um c code with like a whole bunch of go to's right because it was just automatically um generated and then we just check it into um to production which this is like before git and these files were enormous and actually brought the um version control code to its knees because it wasn't used to such gigantic files um getting pushed in there and then we actually had an evaluation set which is something that you know people do now did at the time right and we would run on that then we'd do you know an a b test and some of the traffic and then finally we go into um into production the whole process you know took months actually so the the development cycles here um were once we now we see in lots of companies they do on the order of days but at the time it seemed fast so it's amazing how much things have changed in 2005 there really were not a lot of applications when i was graduating and looking for an interesting job in machine learning now we see almost every company um of a certain size is doing something in machine learning that really matters to them i thought it might be fun to do a quick um survey of some of the stuff that we that we see out there um i mean first of all i think everyone feels the um the impact of speech recognition i mean this is like the device that you know is like obvious to my mother right like you know like now there's like alexa and there's siri and we um you know we actually interact them and kind of um enjoy it um and um and there's this amazing graph um that that actually economists put together which i really love because i think it sort of shows what's really happening in a lot of these machine learning applications where you're moving from you're just making sort of steady progress right so like here on this graph lowers better and it's showing kind of different um different benchmarks and it is in log scale right but you see from like you know 1993 to you know 2016 um just folks are making continuous improvement right but but that's not the experience that we have um as a consumer of ml products i mean sort of you know speech recognition goes from incredibly frustrating um like it was in the arts to you know how it is now where it's incredibly um impressive right and and and we're willing to sacrifice our privacy many of us to put it in our um in our houses um you know just some other examples that that we see that we wouldn't have necessarily um you know known uh would exist is this this amazing example this is one of my favorite examples from from blue river which is a subsidiary of john deere and we're actually just talking about them on the podcast in detail but about how this works but um they are in fields in texas with sprayers and they identify weeds with cameras on these gigantic sprayers and and hundreds of cameras and when they identify a weed they spray just the weed with um herbicide so it actually has like two huge benefits right i mean one is that um you can use a lot less um herbicide like a lot lot less which is good for the environment and um you know good for the the farmers it saves money and it makes um a cleaner world to live in which is a huge deal um and it also makes the um herbicide more effective because they can use different types of herbicide for the different um types of of weeds that there are so it's not just like um you know is it a weed or is it a lettuce it's like which um you know which exact type of weed is it and there's an incredible amount of effort that goes into actually deploying this right so you know i've been working with these folks since um you know maybe 2016 and it seems so simple right this seems like such an easy maybe vision task but you can imagine like you know you're taking this vision task and then you're deploying it into computers putting them in the hot you know texas sun right and they're not like you know putting a huge cloud of dust around the the cameras and the the technology and so it actually is an incredible feat you know not just the vision task but the deployment but um the reward is a better world for all of us and so i'm just i'm hoping that we see more and more um stuff get deployed like this i think another place where you know this is just sort of reflecting on my career when i was at um and i was working at figure 8 we had basically no um healthcare customers uh maybe maybe a handful and a weights and biases one of the biggest places that we sell into and it's really exciting because um i mean if you if any of you have family members that have had any kind of like serious illness um you know that like medicine is just like um it becomes like the most important thing that that that you care about and so um you know i think drug discovery has gone from um you know basically a non-ml task to totally a an ml task like everywhere and we haven't seen any drugs come through um fda trials yet but that's there's a little bit of lag here i think where the technology takes off um and then you know you have to like make sure that you know it's working and safe but um you know i think we saw um you know some amazing stuff with google's um alpha folding if i you know could go back to grad school i think this would be the thing that i would study because this just seems like you know kind of going from zero to to to a lot lot of impact in really short order um you know we see a lot of robotics and vision tasks and the the applications are kind of surprising right like you wouldn't think that like you'd want to make an autonomous robot with good vision and then the application would be you know kind of looking at where stuff is in the supermarket but you know commercial applications are funny um and and now you know you can actually go to a supermarket and see a robot um you know really doing something really using ml to navigate and to check on what's in its world we see um you know medical imaging tasks of all types but i mean this is one where it's looking at cancer cells and trying to make um prognoses um and and eventually maybe this will turn into um to diagnoses right to sort of see what type of cancer you have how bad is it and all that i kind of love these medical imaging tests especially because um you know it's something that um the human bar is lower right like we're you know we're designed to navigate through the world um we're probably not designed to look at gigantic images from microscopes and you know figure out exactly what type of cancer um we're dealing with so it seems like there's a real opportunity um for computer vision to make a real difference in um in people's lives and then of course autonomous um vehicles is is is world changing and even things like you know we see credit scoring so we just see such a breadth of applications it's totally different from um when i was graduating and you know the main applications were essentially um you know go work for a hedge fund and try to you know make a little bit more money or or go um you know basically work for um you know google or yahoo and rank search or rank advertisement so um anyway i think it's a really really exciting time to be in the um ml space and it's an important time to kind of ask ourselves this question like how is ml um different than than engineering right is it just sort of a subfield of engineering now say some people do think that ml is just sort of like a subfield of engineering and it's not that different including one of my podcast guests if you want to get the counterpoint um you know to me you can look at this this interview i did with ananta contrera who actually made some really smart um kind of points on the other side but i i really believe that ml actually requires kind of fundamentally different processes and techniques and so i'll sort of make that case i think you know the first thing is that you know sort of like rips up the whole stack like you know back when i was in college we barely studied um you know compilers i remember i would run into like linker errors and i would kind of just give up um and i didn't see linker errors you know from like i'd say you know 2005 my first job until maybe 2015 when i started to kind of come back to deep learning it's like okay they've swapped out that cpu um with a gpu um and so it's kind of gone like deep into the stack and made some changes i mean i don't know how many of you have struggled with just installing um cuda like when we put out articles that weights and devices on these kind of simple um practical topics of sort of like low-level setup they end up always being our kind of most popular content because i think i think it's actually just so hard to get stuff um up and running in a way that it kind of got easier for a few years and now it's kind of harder again and we need to kind of really fix um the the sort of like low-level setup of machine learning but you know maybe more importantly um i think the code that you generate um for ml is actually really different than the code that you generate um in standard um engineering and i guess like you can sort of see it here on the left and the right of like you know left you have like you know code that i wrote right it's like uh it's keras you know to to generate a double model and then right you see the mo model that i create right so like you know people often ask like why don't folks use git for version control for their um ml models and you know there's there's a couple reasons right they're just practical reasons it's different right so the mo models they're like giant files right so like you know the version control um breaks these files don't diff right so like you know code like you could change you know one line and um you know the the code might you know function in a new way with it with an ml model i mean you could train the same model in the same data you know the same time um functionally identical models like you couldn't find any difference but like literally every character in that file is going to be different right so you don't get that same diffing and obviously like you know we really we write code sort of for computers but really we write it um for each other right to kind of understand what we're trying to do and so there's just no inspectability explainability is such a problem um in in machine learning and so i think what this what this creates this kind of fundamental difference between like writing um you know code that runs like a recipe versus writing um you know essentially code that creates a world for a model to learn in kind of means that debugging is a lot more science than engineering and i like love this um this slide from um andrei karpathy actually presented it at a figure eight conference many years ago for the first time and you know talks about you know how the training didn't really set him up well um for the work that he does right like as in his phd spending all of his time looking at models and algorithms very little time looking at data sets right but then when you're doing actual ml in the real world you spend most of your time kind of tweaking data sets um you know to make the model work and trying to see what it's what it's doing and i think it's kind of telling you know some of the explainability stuff right like um you know to try to figure out what a model is doing you often act more like a scientist right so you'll end up um you know this is a really interesting um uh thing from a paper that i see a lot of people doing intuitively right where you have a network um trained on imagenet and then it's looking at um you know this picture on the left of a you know person i guess with a dog mask you know playing guitar and so you know it says electric guitar acoustic guitar and labrador and so you know uh acoustic guitar and labrador is accurate in some sense and then electric guitar is not and you look at which pixels if you removed them um would change the model's confidence the most right so you can actually see which pixels the model is then using um to make the decision of electric guitar or acoustic guitar or um labrador and so you can see the fretboard is actually the thing that makes the model think that it's electric guitar inaccurately right so it kind of makes sense that fretboard and electric guitar looks a lot like an acoustic guitar but in this sense you know we're we have some sense of explainability but the explainability is really coming from you know not like inspecting the network like we might used to do with like a decision tree or even sort of with a random forest they're really just coming from like observing what the um the the model is doing in practice and trying different uh mutations of the data that the model is looking at to get explainability it feels again a lot more like kind of science than um that engineering um i should say you know you know i want to put a bunch of like links here to stuff i think is really good content if you want to if you want to learn more since this is a little more of a survey talk but you know when you think about um debugging um models there's like so many different things that could go into it right but the output is always basically that you get a worse um model right and so you know this this um website fullstackdeeplearning.com has a ton of free content on this topic that i would like really encourage you to kind of check out and they have like actually really really prescriptive ideas and best practices for looking at um you know how to figure out what's going wrong with a particular um particular model and the challenge is that right you know you can it can be over um it can be um uh overdrive right so like you know you can have bad hyper perimeter choices and bad data construction and bad data model fits or if you fix any one of them it's not going to make your model performance um increase and you get these like absolutely maddening um bugs i thought it was funny actually these in these lecture notes they talk about this bug that i've actually personally encountered so here's a little aside if you if you run into this one i could save you weeks of time because this actually cost me over a week to get to so the glob um the python glob function is actually not deterministic in its order right so if your features and your labels you're both getting from running this glob command they're going to be out of order right and so you know when you when you when you're trained you're actually training on noise right because your features and your labels here um are not matching and it's just funny i mean i literally spent like a week just like banging my head like trying like every possible um you know tweak to figure out what was going on it turns out actually the guy um josh who runs full stack deep learning encountered the exact same um error and got stuck in the exact same way i think there's another kind of practice i mean these like to spend a lot of time with like executives that are trying to get ml to work and i think like the root of a lot of the problems here is that we we actually have terrible terrible intuitions about you know what's hard and easy um for ml and so like i love this example from the movie 2001 which came out in the late 70s right where a computer plays the astronaut in chess and um beats him in chess and also kind of makes fun of him right and and it's fun to watch this clip of this this movie it's uh it's funny what people thought were hard and easy back back in that day and it turns out actually that the um the they almost removed the part where the computer beats the astronaut chest because it seems so unrealistic but no one thought it was unrealistic that the computer kind of banters casually with the um the astronaut and i think it's like really telling right because you know actually it was not that long after this the computers um got better than humans at chess right that happened in the 90s but i mean i think the casual banter um is is almost an ai complete task like i've not had a computer make fun of me in a convincing way um still and it's over 20 years after that so i think we should just always acknowledge that look like you know the best scientists the smartest people we don't know what's hard and what's easy and inside of teams trying to accomplish some specific business objective we also don't know um what's hard and what's easy i'll give you another example um this goes back to my figure eight days but you know i did a kaggle competition um if you guys don't know what kaggle's you should you should go check out kaggle it's a amazing site we can kind of crowdsource um models and so you have lots of people kind of try your tasks and do the best performances so i had a task um the the task was actually a search relevance task and the accuracy went from 35 percent to about 60 percent in just an amazing um such a feat right like i was so excited to see that i kind of wondered how good the accuracy would get actually expand this out you know for another um you know month or two and it completely flatlined right and so i think that there's a real lesson here because actually you might think well maybe people got bored they weren't like trying stuff it turns out more and more and more and more teams were trying to um win it by task and this is a typical curve that you see in capital competitions where you get you know diminishing um returns but you know if you imagine you're kind of the manager and you're seeing this graph you're getting really optimistic and then you're seeing this graph and you're getting really frustrated right this is like the best you could possibly because there's like thousands of people around the world all competing to do this task as well as possible and so you know it's this is not just like in your um in your stand-up that you get these projects that run long right i mean you see like you know in self-driving cars even going back to 2015 you get these exponential um looking curves and then i've been presenting this slide you know for years right but you look at um this is it also in 2015. you know elon musk drops his prediction of autonomous driving from three years to just two i remember thinking at the time that seems a little crazy like i wonder um if we'll really see that right but it um you know i don't think that we have autonomous driving um you know still today right but you can imagine kind of looking at this exponential looking curve on the quality of the self-driving you're thinking wow like this is about to get solved and we look at you know other exponential curves as startup ceos like you know like user growth or um things like that you know you can kind of extrapolate them out right so i think you know ml is particularly hard for um executives and startups to to to reason about i say that as one um you know there's also this um this this reproducibility crisis which i think everyone kind of understands now right but you know there's this real challenge you know um to even reproduce any of the academic results and i know this because i've i've tried to reproduce um you know many of them and people and i did a great job of naming this machine learning's reproducibility crisis and we actually just had a conversation on grading descent about this right but you see like all these different places that come from uh whereas the casticia you can can can break in right like i mean you i'm used to like you know you set your random number seed and you're done i actually think like true reproducibility at this point is is almost impossible right like the gpus just just existing and just kind of sending data different orders if you're working even with one gpu it can be very hard to to make your model um completely reproducible that said i think like everything that you do to add a little bit of reproducibility to your model um it it helps you in many different ways you really can't have explainability without reproducibility you can have auditability without reproducible i don't think you can really say that you're doing real high quality engineering without um some sense of reproducibility and so there's actually a really great um checklist um by uh by um uh put out by mcgill that i think um i think as many of these things as you can check um you'll be able to to reproduce it better but you know i think like the problem is getting harder not not worse i mean i think andrewing you know kind of early said this you know in a very clear um way that basically um you know as we get like more data and more and more computation available the scale just increases right so the the the reproducibility problem and the explanability problem just gets harder and harder um over time and you see this in the you know exponentially increasing um resource consumption and you know i've been talking with like a lot of journalists and others now about like hey what do we do about this right like you know um we keep seeing kind of more and more um you know resources getting used for these models like what if we made these models kind of um you know simpler like easier to train will that lower the resource consumption and actually i don't think it will i think this is the underrated um graph from the same um blog post that open ai you know put it out a while back which shows that um you know basically just to get alex net level performance you need less and less compute resources so actually our algorithms are getting a lot better right where you know every year um we need less resources to get to um a certain quality of performance from an algorithm but that i think actually just inspires companies to use even higher levels of compute because the benefit then um is even higher and there's more compute available so this i think is a really i mean an increasing problem for you know environmental concerns but also maybe more importantly just being able to reproduce any of the academic work that's that's out there and you see like an increasingly complicated dependency graph right like i mean you know back in the day you know you would kind of train in one set of data and just just um deploy it but now we see that like most things like a self-driving car for example try to tries to break all the different things that it's doing into different sub components right and so you know what happens is um you know over time you have different sub components with different release cycles um and they can accidentally um mess each other up and another thing which we used to talk about is like something people should really do transfer learning it's kind of gone from a research topic to just table sticks that i think everyone um does today because the performance is really incredible and there's nowhere that it's more impressive than um with essentially word embeddings and and things that you see like um you know originally burnt now gbt3 and and um you know some of the models in in hugging face to the point where you know you'll see people that um are essentially you know calling themselves prompt engineers right so they're they're saying you know look like my goal is not to like train the models to put the right prompts into um a model to make the model um perform the specific tasks that i want which seems to work particularly well um in the nlp domain um and then kind of before moving on to briefly kind of the stuff that i'm working on i should mention that you know i think responsible ai is becoming more and more um critical to to to anyone that's actually deploying stuff into the real world you see you know daily um issues of like bias creeping into models and i just wanted to point you to i think um alyssa um simpson put out a really excellent um book on you know kind of a breadth analysis of all the things that you'd want to do to to deploy ai responsibly um so i think in the last you know three minutes here i want to kind of talk briefly about um you know what i'm working on at weights and biases the thing i was seeing you know i started the company is like man there's all these like great tools for um the developer workflow where you know we kind of know what the workflow is and i think we kind of know what the workflow is now for you know ml developers but you know what a cml developer is doing is basically custom things text files shell scripts um you know even though there's there's starting to be stuff available it just doesn't seem to like past that bar um where developers really want to use it and so you know the goal with weights and biases was to make essentially developer tools that that people could pick and choose from use the best of breed thing and kind of just solve the specific issues that they're running into in the um the production the real world workflows i should say there's a lot of other um you know approaches to mlaps there's um you know kind of like auto ml where it's like look we're just going to like automate the whole thing and you're not going to need to worry about it like maybe we don't even need the ml engineer kind of interesting but i don't think that's like where most companies are headed right now there's also like a lot of pre-built models also really exciting that's definitely not what we do um i think the premium model is often um you know really impressive but hard to get to use for like an individual specific um use case and so you know we haven't been around that long but i will say we work with um thousands of customers that i feel really proud of and i think like you know it's amazing to see how many different industries um are doing real world ml applications where they're they're using um weights and vices and i think they pick us because of like really really good focus on interoperability um you know we work with the most like we work with more open source repos i think than any other um mls product on the planet and i think we've done a really nice job of um making um strong principles that that we stand by right so you know the key here is we want the integration to be super simple because we know people are busy there's a ton of different things they can try and we really work to get that integration down to a tweet um length we really want to make the ml developer ml practitioner animal engineer successful and we want to be the something they really want to use and it makes our life better and all the stuff they do and then i think one thing that i think other mo ops platforms miss is like collaboration reproducible is really critical right as teams grow they need a way to kind of share um the results of what they're doing and it feels really good to help them with that so a lot of reproducibility is not like a computer science problem it's like a it's like a human problem right of making it really easy like you know everyone kind of knows all these steps they should take to to make models reproducible but the pain in the butt you know so you don't do it you forget about it and if you kind of miss even one of these things on the check box then it's not really um reproducible um machine learning um so i'm running a little bit i know i'm running a little bit over time here but i wanted to say you know for folks that want to learn more about um you know ml engineering um there's three things i'd recommend you you check i think fasta has incredible courses um and we love them um kaggle also you know now actually has courses too but like i think gives you like real applications to get your hands on um doing ml stuff and then you know the way somebody says community we're hoping to build a community that's really welcoming people of all um skill levels and and we'd love to see you there at community.wmb.ai ai so yeah i think i'm right at the the 30 minutes to maybe open up to questions from there you