ai.bythebay.io: Vitaly Gordon, Machine Learning, The Right Way
Recording: ai.bythebay.io: Vitaly Gordon, Machine Learning, The Right Way
[Music] hi guys so I'm botella Goran I lead data science and engineering for Soulforce Einstein and I think Alexei might kind of a little bit oversold me about the like today I still occasionally when my developers let me write code and I think Alexei kind of took it too far is that Batali I want you to you know to start in conference I want you to keynote a lot coding and think okay no pressure so here I am please expect some tough to break it'll be like jumping between a lot of Windows and so if something doesn't work that's less you know at least you guys know it's real so without further ado this is the talk about artificial intelligence for the 99% and this is about all to what we're trying to do in software science and I will not talk about because again my my talk is not to kind of represent the word that my team is doing but more like what I personally passion about and it's kind of bringing whatever we're going to discuss today to larger map I'll start with her rent and the rent is it's actually came from ma'am how many of you here may be know heard of a Giri Timothy used to be at the Twitter and he was one of the authors of scalding okay not so many but some of you do you know and he at least rants over the weekend about data science machine learning in between it was like a very long Hui term and one of the things that said there is like you know deploring is kind of irrelevant for 99.99% of the problems out there and I didn't pull this weed because again there will be a recording of this and I not necessarily you know want to represent because I not necessarily agree with with that statement her I did fine is funny that yeah if you look if you look at you know today's schedule you can see that some words are fairly common in all the session so if it's hard to see what exactly the word I'm referring to I did a work lab for you so deep learning by the way I'm not here to bash the flooring the flooring is great for cell precise line we have with wired minimize as well even Mary will be here later today I think Richard soldier goes and talks on the first day of the conference we all love D flooring and you know in this conference is great because it's you know try also to bring the applied kind of value of deploring and not just going to be theoretical but still it kind of feels like the clip that I'm about to show you when I made the clip it was like from a show that I personally loved it very much but then I realize that maybe it's actually not reaching it as broad of an audience as I saw so just with the kind of show of hands how many of you know top gear or you know then you kind of okayed more the thought than great so here is you know being number one that I hope the sound will work insight that we don't feature enough affordable cars on the show so we're kicking off tonight with the cheapest Ferrari of them all so if you didn't get a joke it's top gear is usually about cars that people get regular people cannot afford so in a lot of people complain about it and then top gear kind of said okay we're taking off the shell was the cheapest Ferrari of them all and kind of this is what I also going to feel a little bit about kind of these a lot of to talk about deploring is kind of a version of the especially apply one without the cheapest Ferrari of them all that ends my rant so what is the problem like why like why did I decide to talk about it and it's a lot to do with kind of the work that we're doing and we're seeing it's not again it's not there is a problem with the chlorine in but something I experienced myself coming from I was previously at LinkedIn which is you know one of these classic Silicon Valley a kind of pet tech companies and now actually at my word we're so first I get to meet with a lot of surfers has hundreds of thousands of customers across the world and work will follow them and in terms of most of these customers and when you meet with them and talk about their machine learning application it's just a little bit like you know Manhattan in the 19th century so the problem is still talking to these companies about the latest events in the flooring it's kind of like you know having Elon Musk show up and talk about Hyperloop when people you know ride the horse on carriages so it's great it's interesting you know Hyperloop - look is amazing SpaceX is amazing occupying Mars is amazing but there is more to be said but some of you might say well you know Vitaly we are in the age of data there's so much data is generated there is n abide the autumn parantha bytes and gilt bytes and you know IBC project like surely you know we have a lot of data and everyone can can benefit from that and yes and you know this is like why I don't not a big fan of a lot of these kind of Alice they they take a very very simplistic you yes the amount of data is growing whatever percent was created whatever the last year's again and but now access all for a kind of machine learning purposes or you know the sexier name AI how many of them how much of that data is actually labeled and could be useful so let's talk about you know one of the largest labels and data sense that we have today to work it's imaging it and as you can see the total number of images in image net today is roughly 14 million images there's the YouTube eight million data set which funny enough and I take it from the hood the YouTube eight million data set contains seven million videos so I think it's you know they chose the title in order to show some progress or you know so they will have room for growth there they didn't want to commit that it will stay at seven million so it has seven million has like you know three point four average labeled for video you can do the map it's around you know 25 million labels again an amazing data set that you know a lot of ther of the latest advances in kind of an artificial intelligence world videos come from them having that data set here is another one the Amazon reviews data set and again first five million number four views seven billion a lizard products a bunch of users with reviews you get it so I get the point I was trying to make here is even though you know the data is the amount of data grows exponentially as you can see actually this data set are extremely hard to come up with and they all you know in the millions which is again it's way way bigger than what we used to have but still it's not so much the exponential growth and a lot of these data set actually have not and have not been you know growing AB fast as all all the other days if you actually think all of them you know we talk about the age of the data it's actually I looked at the file all of them combined can fit on my iPhone so that kind of the age of big data that's and we're talking about so the second kind of problem with would say you know artificial intelligence on the state of the world today is that we kind of you know we talk about deep luring and this kind of assumes that shallow learning if this is kind of the term I will use for what call it you know classic ml that this is kind of a soul problem that you know everything that we could have souls already with traditional machine learning techniques have already been sold and you know it's time to move on to solve kind of the next the next problem and again from our experience with working with the large myriad of companies this is this is not true at all this is not even close to the honor for us to souls and kind of all the problems but to move to the next stage so here a little short demo that hopefully will explain kind of what I mean by kind of shallow learning is not sold and this will be an extremely trivial a machine learning application that let's see if it's which can we switch to green somehow from here terminal or am i doing something wrong Oh okay and what soon Oh well that well it's very small okay is that big enough to see and okay so so just a ratchet library okay and yeah how can i maybe blue screen all of its call nevermind so first of all what I want to show is the Train function of loop linear so look at how many parameters it has so first of all there is the you know type of solver and as I consider eight different solvers for just religious regression and you know some of them are l2 l2 areas some of them l1 so probably most people are in the room either find a trivial or have no idea what I'm talking about and that's fine because kind of the point is there are a lot of parameters that's what I'm trying to say here and then they could have caught there's the cost parameter and there's the you know - excellent parameters there is a bias parameter there's a weight parameters or some cross-validation and you know these parameters are there's the kind of this you are not so much about the actual algorithm but some more without doesn't stuff like that the point is people assume that machine learning algorithms are kind of like quicksort right it's the deterministic output that you just give it in input and we will find the best solution without realizing that oh actually you know we have a bunch of knobs so that you know you might ask yourself but how important these knobs are so in order to actually show I'll take an extremely simple kind of the data set benefit that was like published in the 90s or something with us and hopefully you can see here so it's basically it's called a those data set to turn to predict whether it's very based on the kind of a demographic data in socio-economic data what are the person's income is above the pKa year or below and as you can see there is like aged work class their education marital status occupation raise so again the point here is not so much about that data set but kind of trust me on this this is a an extreme er presentation of a trivial data set okay so and kind of a think there is like capital gain Kevin loss and again just the way there is the data it's a preprocessor there is only 14 features but it will find like more than 100 teachers because what they did here is on the continuous data set they just broke it to quantiles and and all the categorical were basically one hot encoded which basically means is like you know for every single for example country now be comforted by neri variables that is either 1 or 0 whether that person lives in that country so let's try to kind of train train a classifier to predict that data set so first let's kind of start with the services so just I have some so a but the also kind of took this data center is like nine different organization each of them has slightly different different amounts of them and posit but it's all from the same data set you can see that in this a 1 a which has 1,200 negatives and 400 positives so in the kind of a dataset let's just you know train a classifier and I will use the first sold consoler and I will use the C parameter which is actually goes through all the possible lower case C parameters not all the possible but many of them and tries to find the best for this dataset so for this is the one a and as you can see here it tries different C parameters and you can see here that based on this on that single parameter you can actually find there's a quite a difference between the accuracy of the algorithm apply on this extremely trivial data set fringes anyway from doesn't matter what rate is but from 75 to like 83 point seven so just by tweaking one parameter we can actually you know kind of get fairly poor even though 75 is not so poor but again it's a trivial data set so to something that is bigger and you can see here is that the best B is the kind of one which as you can see the log here corresponds to you know then a zero and this is kind of the bit so same data set same exact data but let's just try a slightly different folder and as you can see here suddenly actually it changes and the best scene is now something else so again the point here is trivial algorithm trivial data set same exact data set to completely kind of different like two different tweaking of the same kind of knob that gives us different different results so and then let's take the same kind of data set but maybe a little bit larger which is the ni name which i think is the if the full data set you can see here you know there is 24,000 negatives about 8,000 positives and we're going to try again the same one against the same data set just more labels as you can see because there is more data it's slightly slower and again now the see still and if we try a different solver will get again now to sees 0.25 so but you might ask the kind of okay but there and this is what again I didn't actually touch the data set I was just trying different configuration of the solver like what are kind of a common technique that people might do is say okay you know whenever we I showed you guys there now the positive negative level term something that data scientists do is for example you know we can up sample or down sample the data and let you know see how it might affect our data so I have this you know kind of a very very very simple script in Python scripts that just does you know hey let's up sample the plus one means positive by three times and we'll get like roughly 50/50 correspondence and then you know let's write it to a nine a dot you sent for example and let's just make sure that we are optimal correctly and as you can see it's roughly fifty PP and then we can you know run the same training on the up sample data as you can see actually what's interesting about is actually the lowest now rate is actually higher than the lowest before but the highest is actually lower and now we have similar see but it's actually it looks like an the cross-validation we actually made it worse by up sampling and again people that are don't really understand what all these knobs are we say okay you know if by up sampling we actually hurt at the performance maybe I need to down sample so you know let's try it let's see if we down sample the negative like by a hundred times and write it here and again now just to show that the distribution is now sorry I did I get something wrong what did they do wrong no I I wanted to actually down sample the okay something is Oh okay something is not oh oh sorry yes my bad yeah that's the long part yeah okay so I just remove 99% of the positives and now if we train again and we have here success 99% cross-validation accuracy as you can see it should always down simple so hopefully you're laughing you just understanding that this is not true but what I'm trying to again if from all that bit mode the two main takeaway is just how unintuitive that kind of machine learning is and this is again a trivial trivial example right it's like I try it and different things I'm saying okay computation is infinite we can you know just the brute force through all of it and you know I just empirically test I don't need any theory and you know this is kind of the results right I try to take I get you know 99% accuracy and say okay work and I'll study the diesel also have a tester that you know we didn't train on and you know just trust me that if we try to model I just train on the test set it will not do very well in other words it will over fit so this is kind of my point if we go back to the PowerPoint that the coal you know shallow learning is being a so problem for the 99% it's also not actually true but the biggest problem the mole is kind of how do we take the kind of cute demo I just did it makes it to actually a real viable product so a lot of the companies that kind of we deal with our real companies with real products and little customers so actually the quality bar is as high as any kind of Silicon Valley company in terms of what they need to get into production however and this is kind of a favorite slide of mine that I show pretty much as any internal or external talks that I give that is from a Google paper that's called the hidden technical dead in the shimmering systems that you can read it in anyone here that ever you know try to take a machine-learning product into production probably recognizes every single box here you probably knows that it doesn't matter what it is kind of you know your parameters that say you count the size of the box and line of code time spans engineering hours whatever you want the black box is actually smaller it doesn't mean that the black box is not extremely important or it's only extremely sophisticated all that doesn't matter there's like in real systems they're just a lot more stuff that needs to happen and this is also where we find our customers probably fail the most because honestly the you know queue demo that I did you can find you know it's kind of bright students from a college as you know get them to your company even from you know companies in this flyover States as we call them and they will be able to do something at a PHP level but actually taking into production pouring kind of millions of for real customers you know that's a problem so the question is like why should we care and I kind of touched upon it and this comes also from the personal experience after working at LinkedIn and doing a great job that got me promoted which is kind of you know improving our job recommendation algorithm by two percent I've worked in it for a year and the point is like most companies are not Amazon Google or Facebook these companies are great they're you know I love the products and honestly when I was the data scientist L in working with it even if you fire every single machine learning engineer data science at this company I'm not saying like just you know a remove their work but keep the word that they did and like never let them improve any of it again these companies would still have great products the problem is what about the other ninety nine for some of the companies think about your healthcare provider about your bank about your insurance company about your interpreters these companies is kind of kind of mind-blowing that they probably even have more data about to you then some of the companies omission or some of these Silicon Valley companies however they actually kind of fail to do much with it like whenever I log into any one of these providers this experience feels like oh yeah they've seen you me for the first time even though I might be a customer for a decade now right so this is kind of the problem that we're trying to solve at solar science and and kind of but what is the problem like why is it so hard to so kind of that problem for these companies hey there's more of them usually it doesn't matter what why you're looking at 99% more than 1% the other problems that we have aspect to that so forth is we care about privacy a lot and what when I say we care of our privacy everyone cares about privacy I get it but let me tell you why you know we care probably more foot press itself was the first cloud company so think about it's like not in the year 99 with focus or and we're now celebrating 18 years so yes it's 9999 is telling like a lot of companies put some of your most sensitive data in the thing that everyone gets at the cloud right but then it's basically the cloud which is you know someone else's computer put your data and when I talk about sensitive data if you're a public company the trends the business transaction and you know how many deals you want or loss or basically can be used in order to know how you enter the quarter it can be easily used to it's not even predicting the stock market it's fairly obvious what is company so for us it's isin actually even though they decide this do not look at their data all the work that we're doing is without actually looking at the customers data and again there's hundreds of thousands of them the second problem is performance matters why performance matters again it's like this obvious statement but let me try to give you the analogy in a consumer insurance company if you improve like if you have an algorithm you improve it for 90% of your customers that's a huge deal if they use success you will get promoted you know if and you will be celebrated with us every single customer pays us for that service for example I cannot go to Acme corporation and tell them hey story but it works for the outer 90% but they still want to collect your check it doesn't work this way it has to work for every single customer or you know we're not getting paid in the / problem which is probably the the most difficult is we have different type of beef we have huge customers and we have tiny customers as well and all these customers and here I'm mostly talking about still let's take volume of the data is not like their market cap you know so company like Fitbit it might be smaller but you know because it's a something I didn't mention we also have an IOT business but for example IOT generates way more data than some other businesses that we can add some other products we but the point is is we have to support all ranges and all use cases of data so because I'm kind of running at a time I will job very quickly and to demo and this is an exciting moment because this is kind of you seeing it before both kind of both of the company have seen it and it will be a very small sliver of the things that we're doing in order to solve our problem but kinda hopefully it will just give you a rough rough idea what does it mean to build a you know machinery product for hundreds of thousands of customers of different sizes without actually looking at their data so hopefully it will work okay so let's jump to it large and large larger is that big enough for should I make larger okay awesome so this is kind of we have our biggest business unit is fill clouds or CRM and products and here let me just three count so okay so I just checked out my my branch which is my code and I'll jump right quickly presentation mode and so this is what the code looks like it's something that we call internally Optimus Prime because it's a bunch of transformers and we gave a lot of talk talks about it and hopefully you can see so what is the absence prime is we call this kind of metadata modeling language so we have here basically this is just definition of like what is the data what is worth lo what Valerie reused and this is kind of the cool part because what I'm about to show you is actually a kind of trance image the tensorflow is the computational graph we you can say any node in that graph you want to compute up to that node and this is kind of what this represents okay yeah how do I jump between tabs and presentational sorry oh okay so first of all and by the way the cool thing about it is we also have like this beta mode for up in front that the code you see can be also kind of auto-generated so you can input you can in a kind of an input the customer and their data and that will generate so here's bunch of the kind of features around particularly one of the self or the object to something to know about so first helpers has a very kind of a well-defined kheema that is you know all these helpers customer uses but that schema is by many but also can be customized so here you can some of the fields that we have most of the color code is around oh you so like ignore in the extract function but some cool things say I can show here is like if you see the phone number sorry for some reason oh it's the old and so but we also have you know and kind of a phone and phone type so I guess this is did not recognize it as the phone but it can be also a phone that we can then do some cool stuff with it and the other part of it is actually you see there is also custom fields which are field we actually don't know and every customer has a different set of these fields and their schema and basically as you can see we actually treat them as a key value map any type of transformation for all of them so now we'll see very quickly what are those the kind of em how does the computation graphs looks like and again because I don't know how to switch oh I'll go this so here's the extra code so because that you will see it's not a DSL it's actually it's a DSL but it's the library you can it's the read on top of part ml but what's great about it because I said we have customers of all sizes that code will be can be executed as a part job as a spark bad job as a streaming job for real time and purposes and also on a single machine for the smaller customer that it kind of doesn't make sense for them to have this code run on you know on a cluster and this is almost like 80 or 90% of the customers out there it do not require a digitalization but that code is extremely defensive also on the performance side and I what am about to show you is also on the modeling side so basically you can see here we do you know a bunch of features we so here the pivot table key is basically one house encoding with the top 20 being their own and then the others everything else pulling the outer pocket description this is kind of cool you can see the valid code here is actually because we recognize it's a phone type then we can now do in code completion say is valid phone like things like that in is valid form with the context of the country so basically there is a rich rich text system and I'm you know we can have a Transformers like Jaccard similarity on top of it and then we put all of it like in a feature vector and story that running through this is running out of time but the point here is also this is kind of where a lot of the magic happens is we automatically remove bad for less features you know features with low entropy and I'll just show you the results and here we also apply a bunch of different models here you as you can see there are actually four models that we try it's legit regression random forest but the logistic regression a show that has two params and the random forest has to the in this case it's actually kind of doing research because there is like a very key prompt but in cases where there are a lot of params and even it's not only the kind of a hyper progress of the algorithm but also the parameters inside the feature transformation it bit smarter than just than just research so basically I'll skip the step will I kind of input this but the way the way we do it sorry okay so the way we actually execute it as you can see here it found a the way we award because we have all these different tenants like I said we we actually here found that we have ten different customers and again this is wrong to my local doctor but in our kind of staging and prod environment it will run on all of them and and as you can see there is the IDS of either smart jobs of whatever jobs it is it will actually the one piece of code will run on all the customers at once and you can collect metrics so what kind of the metric look like in this system will be and last thing I show so this is kind of a Zeppelin that we use and again this is in order to you know no one actually has any data on the real app tops I hope you can see and some of the things here and and the point the main point here as you can see for this this is one customer that a these are all the fields that were dropped as you can see some of them are custom and some of them are actually standard fields that were dropped and this is the interesting part here it actually tries for models as you can see here the logistic regression models actually outperform the random forest random forest model and because here the positive labels were very low compared to negative the data set was both up sampled on deposit and down sampled on the negative and the next customer that here and I'm showing it by the way these are real results from our kind of production environments here you can see a completely different set of fields were dropped and as you can see here actually things we're seeing here is a random forest mall actually way better than the logician regression models and again in production you know tries more things and we're now integrating it with slow it's an algorithm tool so I think this is about all I had for today sorry so I just wanted to say thank you everyone and you know if what I talk to you about is interesting your founders interesting so please reach out to our recruiting team upstairs or just me personally as Vitaly at software comm so thank you [Applause] [Music] [Applause] you