Devreal

Scale By The Bay 2020: Jennifer Prendki, When the Only Way to Scale up, is to Scale Down

Scale By The Bay 2020: Jennifer Prendki, When the Only Way to Scale up, is to Scale Down

Recording: Scale By The Bay 2020: Jennifer Prendki, When the Only Way to Scale up, is to Scale Down

[Music] good morning everyone uh it's my pleasure to be with you today uh but if there's one thing that always makes me nervous when giving a talk it's to be the last speaker to speak just before lunch so i'll try to keep it as engaging as possible and hopefully that should be easy because i feel very passionate about the the problem and the topic i'm gonna present today so if you are attending this conference chances are you do love working with big data and you work with big data on a daily basis but in fact i believe that you probably are more likely to have a love-hate relationship with the data that you work with and this is the topic we're gonna discuss today so this topic of love hate relationship to our data is something that has followed me throughout my entire career and so i remember when i started as a machine learning scientist at first i was very frustrated because usually the companies i was working for did not have enough data and suddenly when we saw this huge boom of large data sets everywhere i started complaining about not knowing what to do with it and later on when i became a manager i constantly had to manage budgets and figure out ways for for my teams to basically manage to do machine learning with big data and i felt so passionate about trying to provide a solution to the community that at first i joined one of the largest labeling companies figure eight as their uh vp of machine learning and i eventually started my own company uh electeo to basically try to provide solutions to the problems of uh the problems associated to doing machine learning with big data so just a few words about electros are still a very young company you probably haven't heard of us uh and we're currently 13 people less than two years old and everything we do is essentially to help machine learning teams and machine learning companies do machine learning in the edge of big data and just like the rest of you guys we believe that machine learning is awesome there is no question that we've done a lot of progress recently but there is also no question that doing machine learning with big data is expensive and time consuming so the way we sort of try to solve those challenges associated to doing machine learning with big data is essentially by trying to provide solutions to help people reduce the amount of data that's actually necessary to train machine learning models and by doing so we hope to help people reduce training time and essentially time to market we help people reduce operational costs whether those costs are associated to data labeling compute or storage and in fact our crazy hope is given that we potentially believe and we've seen cases where training a machine learning model with less data has actually helped companies achieve better performance just because you're solving a garbage in garbage out problem but before we talk about how it's possible to do machine learning with less data and achieve similar results that you would have with large data sets i would like to try to make the case why i believe it's time for the community for to have a paradigm shift in this direction and so the very first thing i'd like to discuss is basically the history of big data and when i discuss the history of big data i usually like to start with the history of deep learning because i believe these two are so closely tied to each other something that very few people even experts in the field do realize is that the mathematical framework and the early version of our modern neural nets have actually been around for a very very long time in fact like 70 years essentially and even the modern version of the math behind back propagation and our neural networks have been around for decades right but as we all know uh is it took decades and it's only in the last 10 years that we essentially started seeing like a true usage and practical applications for deep learning right i mean so why why did it take so long right so uh even though backprop was invented uh in uh 1975 it took 15 years or almost 15 years for jan lucun to be able to practically test bug propagation and then it essentially took another uh like 20 years to reach the point where we started seeing an explosion of uh people using and deep learning on a daily basis right so what happened there right so of course there is no question that this is essentially associated to uh the the the the coming of uh the big data era which is tightly uh coupled and associated to uh an era where we started having better hardware right i mean so it's unclear whether we got better hardware because there was a need for more data or if we got more data because uh suddenly better hardware became possible but these two are tightly uh coupled to each other right so it seems that now there is no limits anymore to what we can try to achieve and we can finally put in application all of this research we've done for such a long time right so uh the problem though is that doing machine learning with big data is is actually very challenging for many different reasons right so when you work with big data essentially like the more data you have the more data labeling you're gonna have to do it it does not necessarily seem to be a big problem specifically if you're not in charge of a labeling budget uh it does not maybe you you don't realize that this is a problem for for your company but a lot of the labeling today uh tends to be done manually and so that tends to like uh incur costs for the company it's difficult to manage and it takes time right even data storage which doesn't seem to be such a big problem because typically data storage isn't the most expensive part of doing machine learning we all know that s3 instances are significantly cheaper than an ec2 instance but still for many companies like given the volume of data that they they collect it might eventually become a problem and so even though we're light years from the age of the floppy disks uh this this can this can still be an issue for many companies uh then there is the problem of compute so that's something that even ics probably can be more familiar with when you train the machine learning model and if this deeper like for example you work with a deep learning model which is extremely deep or complicated basically training this model either on your own servers or on the cloud might be very expensive in fact i wouldn't be surprised if your manager has uh um has been complaining about uh the costs that this has cost to the team and then eventually like even if you forget like the monetary aspect of things working with larger data sets obviously tends to uh make the entire process a lot slower either when when you think about data preparation or when you think about the training process itself so uh i mean at this point i'm pretty sure a lot of people will feel like there is an easy solution to this problem and that solution is to get better hardware right i mean and so in to some extent it does make sense because with better hardware we got to where we are today right so this is essentially what moore's law is promising us moore's law is this uh empirical law that states that uh the performance of the hardware that we create will keep increasing proportionally uh basically at the same speed than uh the data that we collect right and that that actually makes sense because we need better hardware also to collect more data right uh the problem though uh that some people will point out and some more and more experts actually tend to believe in is that to get better faster like more for higher performance hardware you need to get smaller transistors right and the problem at some point is like the size and the scale of the transistors were reaching at this point is putting us in a situation where that scale is reaching uh the skills where uh the basically like a scale is governed by the the law of quantum physics right i mean so in other terms there is a chance that we won't be able to do any more significant pros uh progress with hardware unless we start mastering quantum computing and i don't necessarily think we're there yet right so even though for many years many people tend to see what i call the optimistic view of moore's law if you are being honest with yourself you probably believe that there might be some slow down right now right so now whether or not you agree with this conclusion and you agree that uh there may be like some slowdown in uh moore's law you also have to admit that even if better hardware exists it's not necessarily always available uh to the common machine learning scientists out there right and so the fact that good technology exists doesn't mean that it can practically be used by the various machine learning companies out there now something that i feel unfortunately is not being talked about enough is the environmental aspect of big data right so if you're going to work with more data and better hardware you still have to keep in mind that this hardware functions on electricity right i mean so just to make a point i would like to give you a few numbers about data centers right so uh as of 2020 we now have 40 000 zettabytes of data collected worldwide if you think about this number that's basically 40 000 billion of gigabytes that's obviously a very very large number right and you you need to put this data somewhere so typically you put this data in one of the many hyper hyperscale data centers somewhere in the world and these are actually very large consumer of electricity of power right and uh about uh i mean depending on the source that this this data and those numbers come from like the the average is basically like people estimate that two percent of all power consumption on planet earth actually come from from big data uh so the reason why uh this is happening is that those data centers are typically not very efficient so it's it's basically like one of the reason why this is happening is that you don't only need to make those machines can function but you also have to cool them and so if you have an ac at home you know how expensive uh cool a cooling system might be and unfortunately many of these hyperscale data centers tend to be in places that are not particularly cold and so since we've talked about the impact on the power grid and obviously on our electricity bills as machine learning scientists or the electricity bills for our companies uh i think it's worth talking about carbon emissions as well so right now about two percent of all human generated emissions come from big data and data centers specifically right it does not seem like a lot certainly compared to other industries like the closing industry or agriculture for example but what's sort of like really concerning here is that we see this uptrend like for example we predict that by 2025 just like five years from now uh we'll reach like 5.5 of all uh generated emissions and this is sort of the opposite trend that you see with other industries where one of the very few industries in fact that is not working hard to try to reduce this number and when i first saw this number my hope of course was that we could adopt more solar electricity and be like be more sustainable but unfortunately uh if you look the numbers here this is this is actually not the case so before like basically uh jumping to the next session and discussing how it's possible for anyone to try to control the amount of data that they work with i would like to close this argument and try to finalize the case for smaller data sets by sharing with you a few numbers did you know for instance that 55 of a company's data was typically what you would call dark data so dark data is essentially data that even like basically like uh it exists but the company either doesn't know how to get to it and they don't know how to access it or they might not even know it exists in the first place in fact 85 percent of all data is either dark or rot which is essentially an acronym for redundant obsolete or trivial data just think about this number 85 percent of the data that we have stored somewhere is actually not something that you can use right and 90 of all unstructured data is never even used either for analytics or for machine learning purposes right i mean so uh so now again when i saw this number my hope was that maybe in the future when uh companies get uh you're making possessions of like better servers and so forth and so on maybe they will leverage the data unfortunately uh for us it turns out that 60 of all data we collect loses its value immediately when collected so essentially 60 of this data cannot even be used at all so this should basically like prove you something very important we're all data hoarders the machine learning community and the data community is hurting data that it cannot even be used eventually right so this sounds like a terrible statement to make right but when i look at this i look at this as being an opportunity because this should prove to you that there is a lot of room for improvement so at this point i hope that uh you so you're sort of convinced that scaling down is is the way to go and that is time for all of us uh to shift our opinion about big data being necessary for a good machine learning models so this is sort of like uh weird for us to say because of course we've all been trained and uh to believe that the more data the better our machine learning algorithms would perform right and i've been taught that i'm sure you've been taught that as well the good news is that there is actually a lot of hope and there is a lot of research in many areas that are sort of like either try to control the cost associated to working with big data are to try to control the actual amount of data that we work with altogether for example if you just focus on data labeling cost you not only have like new breeds of a machine learning algorithm that help us like automatically label data for example you might have heard of the human in the loop paradigm that helps you like identify a poorly labeled data or leverage automated processes on top of humans essentially humans just focus on the corner cases you might have heard of the snorkel algorithm which is a using weak supervision to automatically label data but like for me like the biggest promise is going to come from uh other learning paradigms such as transfer learning meta learning and active learning which all attempt to try to reduce the amount of data uh that we train our machine learning models on right so now of course none of these methods and there are many i didn't list here can solve all of our problems at once but there is a lot of promise in a lot of different areas now of course uh i'm sure you won't prove that you can actually train a model with less data so this is basically just like a a couple of screenshots that were taken from one of our blogs uh one of our blog posts and that's basically some research we've done on an open source data set uh called deepweeds and the problem is essentially image classification for uh four weeds right and the plot that you see on the right here is basically a series of violin plots where we try to see uh how models perform or how a model performs uh when you vary the amount of data that it's trained with right so each one of these violins here are essentially uh showing the distribution of the performance of uh the same model being trained on 20 different models or 23 different data sets excuse me uh trained with the same amount of data so the first violin plot that you see here is essentially the distribution of um 20 experiments that were all trained uh examples of like uh like in each one of these cases the training data set was a sample of 20 of the entire data set and what we displayed here is the accuracy that was measured on the set for each one of these experiments right and so we repeated the same experiment for various amounts of data right and there are actually two very important conclusions here right if you first you realize that the very best sample of size 20 actually over performs the worst samples of size 80 right so it should prove to you two very important things first and foremost it matters which part of the data you train with right because like for example even when you pick 80 of the data there is a huge variance in the accuracy of the data that you're going to work with right and this is all just like random data and samples that were uh randomly picked right number two you should see that it's possible to get a better model with less data just because like there are many configurations where uh the best uh performance of uh you know like even for 40 20 60 the best one all over performs a large majority or a large fraction of the the cases where we trained on 80 of the data right so now that you believe that and hopefully you do believe that it's possible to trade models with less data i would like to explain a little bit deeper how active learning works and because active reading is uh at the very root of a lot of things that we do at electio right so you probably heard about active learning but chances are you haven't heard the right definition because i believe that active learning is one of the concepts that is that is the most misunderstood in the industry so the definition the actual definition of active learning is that active learning is the semi-supervised learning technique based on an incremental learning approach that typically people are going to use to reduce their labeling cost so what does it mean it means that uh because it's something supervised we work with only partially labeled data the fact that it's incremental means you're going to retrain with by adding more and more data and we basically process in loops so each loop is the process where you train your model with a subset of the data i'm going to show you how it works later and typically people have been using this mostly for operational reasons up to now there's been very little research in trying to use active learning to try to get better models right and so uh even though it's kind of difficult to put in application and in practice in real life the process of active learning is actually fairly simple it's a repeated set of loops where you select data you label this data you train with this data and you run inference uh with this data so i'll show you how what the pseudocode looks like so that you can try and implement this for yourself i just want to emphasize that active learning has nothing to do with what uh with auto labeling and it has nothing to do with human in the loop uh even though a lot of people tend to believe that all right so what are the pros and cons of active learning well first and foremost because active learning was originally designed to help people save on labeling of course this is like one of the major like her pros of uh the technology and the paradigm right active learning is also doing a lot better than any other semi-supervised learning techniques uh when it comes to like generating biases in fact uh it's very it's much a lot more difficult to generate biases with active learning than it would be with a regular service supervised learning technique which you might have used or heard of elsewhere and finally like from my perspective i see active learning as being one of the very first steps towards like becoming more programming when when trying to do machine learning unfortunately active learning has a lot of downsides as well and you'll see why when i show you how essentially how the algorithm works first and foremost it tends to increase training time it tends to increase increase compute costs because you have to retrain the model over and over again and it's actually pretty difficult to put an application certainly more so with uh data sets that are messy like they are in the industry so even though you have you tend to see a lot of research in academia not a lot of people are using it in the industry so the good news is that uh there is a lot of like uh it's very easy to put or it's fairly easy to understand how it works so i'm going to try to explain how this works now right so when you do active learning you essentially start with the same training data set you would work with if you had uh if you were doing supervised learning right you're gonna work with the same model typically you already have a model that's optimized and one of the most common applications for active learning is to retrain a model that's already been optimized and you're gonna have a similar data set just like the same way you would have for supervised learning right now within your training data set you're gonna split your data into the unlabeled part of this data set and the label part of this data set that's that's essentially one of the only like uh uh difference at this point right so what you're going to do is like as i said earlier the process starts with a select process so at first typically people tend to randomly pick some of the data then this data gets labeled it gets labeled like uh you can label it whichever way you like then you're going to send this to your model essentially train your model and get a model instance uh with this specific subset of the data now because you've used a small part of your data set you you shouldn't have any expectations that this model is going to perform really well in fact probably it won't right and then uh once you're done you basically put the data which you just labeled and just used to train your first version of the model uh into the batch of data that you call the label data set right and then this is where things get interesting at this point you're going to evaluate how good your model is on your test set and simultaneously you're going to run inferences to basically try to probe the model to see where the model is at right and so you're going to basically get a lot of predictions from the unlabeled part of your training data set but you cannot use those predictions per se to make a decision for example you cannot say that because your model seems to be underperforming on cats for example you should pick more cuts within your training data set because if your model is struggling with cuts the data that was predicted as a cat is actually very unlikely to actually be a cat right i mean and so uh the the fact that you are not relying on the prediction itself is the very reason why it's unlikely that you will generate biases um it's more unlikely that you will generate biases compared to other uh semi-supervised learning techniques right and so then you keep going you increment the loop number right i mean so we went through the first loop and now we're gonna do the second loop where we're going to select additional data and this is where you can be smarter about selecting the data because you know more about where the model stands uh and you're going to label the subset of data which means you're going to have to do labeling incrementally and then you're going to retrain your model not only with the data which you've just selected but also also with the data which was already labeled previously and this is actually one of the reasons why active learning in its original format tends to be very compute greedy because you keep retraining on the same data over and over again right so now let's jump into the the pseudocode so the pseudocode is actually fairly easy to understand and uh once you understand how uh the the essentials work you're gonna be able to uh start experimenting with it yourself and i i sure hope that you you try to do that so when you initialize the process you will basically like um initialize your unlabeled data set to the data because the assumption is that you start without any labels in your data set at all and then you're going to jump into the first loop um so for the first loop and usually what you do is like when you select the first subset of the data you tend to do that randomly because you don't know much about which data is actually going to be useful to the model and then you're going to label this data how the data gets labeled does not matter it can be done manually it can be done with a human in the loop process it can be done with machine learning alone right and so the reason why i mentioned earlier that active learning per se is not a human in the loop process is that you could actually use a fully automated process to label this data in fact you could use something like snorkel and in this case there was there would not be a a human in the process or human involved in the process at all uh then uh quite obviously so since you label some of your data you're gonna have less data to unlabel um uh or less data to pick from in the next loop you're going to train uh your model with data you just selected and then you're going to run your inferences uh and basically so this is where it gets interesting because this inference step is essentially the point where it is your opportunity to try to measure where the model is struggling and try to help the model get better so usually what people do at this point is to try to identify or measure the level of uncertainty of the model and a lot of people do that by using a confidence strategy or sorry uh acquiring strategy uh called confidence as we're gonna talk about later so you keep going over but like what gets interesting for the next loop is like now you can use the inference data to try to be smart about which data you should pick and so the selection function here is what people refer to when they speak about acquiring strategy in an active learning process so now the question is like how do you tune an active reading process how do you know how many loops you need to use well it's difficult it's difficult to predict ahead of time and usually people pick a number of loops based on operational constraints such as a specific amount of time or specific amount of money that they have to spend on data labeling how large should the loops be well that's another hyper parameter usually when people do regular active learning they tend to control or set a loop size ahead of time but there's actually no reason why this number should be fixed why this number should be constant and why this number should actually be decided for in advance right i mean so in other terms that's a parameter that can be tuned as you go through the active learning process and so this is one of the space where i believe there is a ton of research that needs to be done right finally maybe like the most important question here is like what do you pick for requiring strategy so the most common thing that people are going to do is to try to measure the uncertainty of the model you measure uncertainty by looking at the prediction like so the confidence level with which a prediction is made so even though you cannot look at the prediction itself you can try to see how confident the model is unfortunately predicting uncertainty is very difficult specifically with deep learning models so traditionally those querying strategies tend to be static functions in other terms you're using the same sampling technique over and over again but you can actually become very creative you can try to have something that would change over time it does not need to be decided ahead of time the only constraint for you is really to identify like the only thing that acquiring strategy is is technically it's simply a sampling method so just to show you a few results like these are like uh like again going back to deep weeds like these are a couple querying strategies that we work with and so one thing you can see right away is that not all querying strategies are going to be equally good right which sort of like reflects what i said earlier that uh the subset of the data you're going to pick sort of matters a lot right here in this case the most standard querying strategy that everybody uses which is um low confidence actually performs very poorly in fact it performs worse than randomly picking the data just think about this right uh and so uh we actually like obviously do a lot of research in this space and uh one of our preparatory uh technology is what we call auto active learning which is a process where you you do some of the things i mentioned earlier you use like a non-static like amounts of data or or basically querying strategies and uh and you can see that without doing any tuning at all we're able to achieve the same performance for 80 of the data that uh the margin requiring strategy did right i mean so uh of course margin is performing pretty well the problem is that if you do not know which querying strategy is going to work you would have to essentially like a bootstrap and brute force a lot of growing strategy to get to that level right i mean so there's obviously a lot of research to do there so i seriously encourage everybody to start looking into active learning because active learning shows a lot of promise not only to try to reduce the amount of data and reduce costs but also in terms of many different things right so what we are trying to do and what we do a lot of research on is basically building a better breed of active learning so what does this mean right first and foremost it means building something that's more robust one of the problems of active learning is that it assumes that all data is correct right i mean so so like active learning is trying to identify the most useful data within a data set but this assumes that no data can cause harm to the model which is usually not the case when you work with very messy data we also believe that it's time for the industry to start adopting uh machine learning driven querying strategy the traditional querying strategies typically are preset pre-determined rules are based on like it's essentially a rules-based sampling method and there's actually no reason not to do this in fact just the same way that uh eventually at some point people decided to use machine learning as opposed to rules-based models it's time for us to adopt machine learning driven querying strategies i also believe it's time for us to realize that active learning is actually very promising in different areas as well for example by trying to analyze which data was prioritized by the system you can try to understand which data was the most useful and you can try to start diagnosing why this data was more useful and maybe collect more of the same data identify the weaknesses of the model and so forth and so on finally you might remember that i mentioned that active learning tends to be very compute uh expensive uh and so one of the things that we're extremely focused on these days is to try to create a version of active learning that doesn't increase the amount of compute being used to train machine learning models right i mean so you can do this by figuring out uh how when you need to stop active learning tends to to lead to this plateaus uh or plateauing learning curves and basically like a there is a point in time where it seems that the model isn't learning anything anymore so the sooner you identify when this happens the better and we're also trying to find ways to avoid retrain on the entire set of selected data and essentially only train on the data that just got selected so my conclusions today is that i think it's high time for the machine learning community for a shift in uh in paradigm right i think people need to realize that less can be more when even when it comes to data i obviously believe that big data is amazing but i believe big data is amazing because we now have the opportunity to become more picky about the data that we we use right and the good news is that it seems to be that awareness generally speaking uh is starting to race there is more and more research in this space in fact like there was never as many papers on the space on active learning and on the many other concepts we talked about earlier than there was in 2019 active learning is obviously very promising there is a lot of promise in using active learning in the industry as well and it can be used beyond the use case of trying to reduce the amount of data the amount of money spent on data labeling the problem though is doing active learning is extremely challenging as you could have seen from my presentation vanilla active learning uh is really hard to use with messy data so that's one of the things we're trying to do a lot of research on and doing hyper parameter tuning is difficult and time consuming regardless of all of this i hope all of you enjoyed the presentation and i hope you'll be interested in trying to to validate or test active learning on your own use cases we'd love to hear about how this goes we'd love to help and i also want to encourage everybody to visit our blog post we publish key studies all the time we talked about we talk a lot about our technology so uh i hope you enjoyed the presentation and i hope you'll be interested in hearing more about active learning in general thank you