Devreal

Deploying and Scaling Spark ML and Tenso...

Event: Oh Hai Ai

ai.bythebay.io: Chris Fregly, Deploying and Scaling Spark ML and Tensorflow AI Models

Recording: ai.bythebay.io: Chris Fregly, Deploying and Scaling Spark ML and Tensorflow AI Models

you [Music] so like Alexa said today is going to be at least for the next 40 minutes will be rightly mostly about the plane models I'm going to try to sneak in a little spark in there but we'll also talk about tuncer flow as well too as like i said i have this combined spark and tensorflow meet up and actually so yeah today's talk is going to be all Jupiter notebook and it's going to be demos where the demos break i'm just going to kind of fake through it and show you guys right so the results that were generated right before this talk if anything happens yeah so here we are today let's see who am I Chris fregley I I'm pretty famous for this emoji at least i think i'm famous in my head for one of my friends well it got me this emoji that looks just like me am I wearing oh no I'm wearing my bringing yes I'm bringing taxi back that with this hat here so the yes I bought this off a taxi driver down in San Jose like six years ago I was all drunk in gave him 50 bucks and it was like the happiest day of his life and I woke up with this hat and 50 less dollars in my pocket so but yeah the emoji is the Cubs from Chicago and right like we won now where the world champions or the world series champions but it's now I'm scared to not wear the hat because maybe something some other I don't know yeah there's a lot of stuff here okay so research scientist pipeline I oh I haven't even changed my slides I go back and forth between pipeline AI and pipeline i/o depending on who I'm talking to I've got this thing called the pancake stack kind of funny things to remember from this talk well yes I should mention that the last time I gave this talk it was four hours and that was like a couple Tuesday's ago so we're going to do four hours in 40 minutes so there's a lot more here if you go to pipeline I oh you'll see all the code off of github everything's hundred percent open source I don't make money the investors are asking why I don't make money I just don't call them back that's my plan so pancakes yeah there's a lot going on is presto and yeah there were a lot of a's in pancake stack so i started getting a little creative I got in trouble from the Apache folks for not putting the word Apache and right like in front of everything and I had to remind them that it would just be the like paw stop you know stock like kind of thing you know it'd just be a bunch of a's if it was all Apache Apache Apache so don't ya don't photo this and tweet it too far because then they start coming after my ass working on high-performance sensor flow in production that'll be coming out later the end of this year it's going to be or sorry it's going to be end of this month there's some online training yeah you can click through I think the way that the o'reilly people do it now for the online training is that you have to sign up for Safari so they're kind of going like the netflix route you know the like recurring subscription seems to be a big right like thing around here in the Bay Area right yes of BC's like these recurring things versus just a one-off so I think if your Safari member you can get to it we're going to be covering a lot of optimizations going into way more detail than we do here once you've trained your tensorflow model there's things that you can do to the tensorflow model to make it score faster right and right like smaller devices things like that so I'll mentioned some of them here and in the context of deploying yeah so again here's the meetup yeah there's Chester there was Chester Yeah Yeah right there Chester yeah between me like alexion Chester 90 right leg meet up last year for spark and spark related things here's the github repo docker hub all the SlideShare worked for netflix works for data bricks dinner brief stints one of the early members of the IBM spark tech center we're trying to grow that thing out pretty quickly it's pretty fun challenge you know trying to wear shorts in meetings where everyone had a keys and Riley not just khakis but pleated khakis like from the 80s like old navy side khakis so yeah I just bring my car Bo like the cargo pants and yeah so we change that culture a little bit there that was fun yeah actually IBM stark tech center now has I think for committers or yeah maybe even five smart committers in that group which was pretty amazing seeing as we have zero when we started so I'm a spark contributor not a committer all right some of the tools that we're working with today are dr. Cooper Nettie's anyone here use Cooper Nettie's in production anyone use may sauce in production they're used for slightly different things at this point but ok how about yeah I guess yarn all right it's like every one yard no one yarn all right good crowd so this is we've schoo this is a fun little tool on top of coover Nettie's this is one of riley gilben source rightly products by this company we've we've worked so i'll yeah let me zoom in a little bit so i don't forget and i learned this trick the other day so it doesn't take up half of the screen so i'll actually show this live here in a sec but this is a so this is the training cluster yeah oh actually this is both a training cluster and a like prediction cluster so right like here is where you're going to have your tensorflow your high meta story your Redis got kafka things like that so shows all the moving parts so these are all doc arised right this is all within that same github repo these are all pushed out to like docker hub you can pull these down you can essentially with just a couple clicks you can create right like this whole environment using Cooper Nettie's yeah so you can do it locally now there's something called mini cube right the like really changes the game right like similar to how there's doctor for mac and doctor for Windows now you can actually run right like an entire crew burn Eddie's like cluster on like your local laptop it's not always successful just because of the available RAM and things like that but it's fun to try anyway will be as in Jupiter we probably will not leave Jupiter much yeah I like during this because we could do everything from within here who here uses air flow do you think Yeah right we're just talking about that who knows Jenkins yeah anyone heard of Jenkins right where you push you like bills code the kind of hot new thing is that you continuously build code right you have your little source java directory and like you make changes throughout the day every 10 minutes every hour that code right like you want to push it out so when I first and yeah so think of continuous like model deployments right this yes think of it similar to how you would do continuous code deployments I started studying up Jenkins when I first stood out to like tackle this particular use case the continuous deployment of models I realize I had to learn like right like another son like goofy groovy dsl that I guess was the old workflow dsl that they renamed to pipeline like dsl it was just a lot of like heavyweight stuff I'd be Jenkins rather all the time pretty much like the past four or five jobs like we always use Jenkins or Hudson or some form of that I was trying to find something different like something a little bit more Python II I have a Java background but most of the people i talk to have data science backgrounds have Python backgrounds so I was trying to kind of fit like their world and not force people to learn java and the JVM and you know some of these things so yeah these air flow operators they're called you can set up these dads so similar to a tensorflow dag right where you have operations or your spark dag where you have a map and a filter and joins and things like that think of this sort of like a production workflow dag right so here's a case we see i think i could show one here so here's one so picture the ability to trigger these dags right that's the term that like air flow uses and a valid trigger would be committing something into github so like imagine training a model within the notebook which will do here in a sec and then commit that that's going to then do a call back right from its a like post commit call back right like web hook back into air flow and there's an airflow server running out there and it actually triggers this particular thing passes this github hash right the commit hash build a new docker image with that model that you built that tensorflow model or spark model that you built and then write like does this thing pushes it out to whatever like dr. hub type repo that you have for your doctor can take for you write like for your docker images and then in parallel can actually update rightly both amazon if you're running in Google where like you can update that as well I had as you're running for a little bit and then I forgot my hotmail password and like I don't know what to do anymore I'm just kind of locked out still like paying money for something that I can't seem to delete so let's see back to get back to hear ya air flow has become kind of a critical piece now I just went to a meetup the other day at like galvanized i think was last week two weeks ago it was one of the Airbnb people speaking about air flow he's actually on the air flow team right he's an airflow committer sort of talking about the history of it the problems right like that they were having right like the other tools that they had things like Jenkins that kind of stuff and one thing he showed me was so one thing that's been on like my mind recently is not just deploying these models but actually a be testing them right like doing multi-armed bandit testing where you kind of shift traffic towards the winning person so that you're not completely screwing all the other like bad experiments right the people using that had been slotted into those bad experiments for the duration for the three weeks or four weeks of this like right like particular test and he popped up there was this like like HTML you I that was tied to airflow here and it actually like let you set the prior distributions on those different tests right like ABCDEF tests and then give it the duration for the test itself and that would actually generate right like gamal that's used or yeah no sorry this is actually python-based right so each of these is an actual Python I think I can get to it from here yeah these operators are just Python code so if I back up should be able to drill in and go to code so this is actually one so write like this is the example where I've created a new tensorflow model using some new hyper parameters trying out right like a new architecture which is a form of hyper parameter and I then push it and from this creates a canary right so just to kind of show that that visual version that I had showed here here's the actual code version of it so right like the interesting thing will be these two happening in parallel which is the very last step so right now i just kinda because i'm like still getting i'm still trying to learn air flow and some of these life oh yeah so just picks your air flow just like it's a python library right and yeah you can extend these things yes i'm using rightly for here bash operators which is just just kind of a generic where they execute this right like batch command bash command so the very end here like you'll see so once the image has been pushed to dock or hub then right like these two things then trigger off that which is update amazon and an update gcp so like that's where you actually get this parallel ISM you would use air flow in conjunction with celery yes I didn't know any of these terms before cuz like I've always been kind of a Java guy but apparently everyone knows celery there's this tool on top of celery called flower right that actually lets you inspect the tasks so picture celery is kind of your just submitting tasks and right so as there's like task workers available it'll just start gobbling them up then doing them yeah so you can scale up yeah and then flower guys but it's kind of a funny name on top yeah so that's how you would actually build these things sort of end-to-end okay oh and then if there's failures and errors and things like that of course you can notify yourself or have it send an email to pagerduty right like one thing to mention i mentioned the like github dependency right like never ever ever depend on right like github for like anything right that's so like Netflix for example they run their own stash service that's like highly available and write like supports the get api and all that but they have control over the availability this is such a critical part oh yes I didn't finish my thought back here it was yeah so this particular run actually goes through build a canary from that newly committed model Riley builds the canary and so that term canary think of it just like canary code right that goes out that whole like canary in a coal mine where you stick something out right like into the coal mine see if it lives yeah if it fails then we're at like tear it down don't go that direction so yeah this particular flow would would create that one canary that sits alongside the other thousand model servers right so now you have a thousand and one that are all taking traffic right so there's one load balancer coming in that's spraying it out you then have a dashboard where you can actually compare right like not just like system stats like load and Riley keep and like things like that but you can also look at like predictive performance as well so you can actually compare predictive performance so for the same inputs right like the old model gave you this and now this new model gives you this and if they're wildly out of sync if one just Riley keep saying no this is not spam not spam but like the real model or the current model is then I think you would know what to do and now one thing for these deep neural networks one of the hyper parameters of course is how deep like your network is right so one thing to keep in mind is that you may see Layton sees go up right so if your data sciences person sitting down building a new model on your life wait a minute I'm just going to make this thing 50 layers deep right yes that has to be better of course and so you go to deploy that model and now right like the average latency 90th percentile latency goes up to you know right like 50 milliseconds whereas you have kind of a like a 10 millisecond limit and so you would actually see 50 milliseconds vs 10 milliseconds right that's a system metric this is actually a dashboard that's shown to the data scientist through this process right so I think someone yes that was talked earlier today about kind of keeping or like data scientists really don't have to worry about infrastructure and like things like that like yes they sure they certainly shouldn't have to really know the details but yeah from riley r / from right like our standpoint we're like we want to show them these dashboards right like give data scientists data not just about their training like results and all that but like how is this actually doing in production on live data okay so mention the canary lease think of continuous deployment more than just code like but also your train models I'm a big fan of dashboards this is something I I learned and was kind of beaten into me at Netflix was if you are pushing out code that you don't have insight into where like new code that you pushed out new models that you pushed out and you can't really defend what's going on or if something is happening right then you basically already lost you should just take it out and every instrument it and then put it back out yeah so something that was tough under at the time we were under like two-week iterations and of course metrics is always the last thing you know things like error counts and all that and sometimes you would instrument specifically for that use case and then you would just kind of put it to do to take out the instrumentation later and of course that never happened right so like you really have to be disciplined but here's one example this is like one of my favorites so this this particular concept is directly related to serving like models right so this concept of a circuit breaker yes the netflix has thus open source library called hystrix this is kind of how it looks with all these kind of detailed out it's very very visual and show you guys this is across both amazon and google you actually fire this up i think i'm pointing at Google right now yeah so what you see again right like this is metrics this is data for the data scientists to see what's going on out there he doesn't necessarily have to write like respond to things that are going on but if you arguing canary analysis right like you would want to compare and see what's going on but yeah those pages judy alerts right yeah they should like typically go to the ops person I wouldn't guess I wouldn't like necessarily expect that like data scientist person can actually fix the problem but they can definitely fix things that they know right like Layton sees and things like that so this is a super small load test and actually this is being limited by my laptop because i'm running jmeter locally here I've got jmeter in the cloud but if we have time we can load it up and and I could switch over and show amazon as well too so these are like deployed services here so I'm kind of showing you guys the fun metrics part and then we'll back into how we got here so we should see the amazon stuff light up and we can actually compare this you i was a little bit tighter like we might even be able to oh yes i'm only running one at a time but at one point I was getting for the same price per prediction right this is one metric you can pay attention to so the the same size of instance I think it's that n1 or it's that the n1 hime m8 or something on the Google side which is equivalent to like an r32 XL so it's basically 84 is about 50 to 60 gigs of ram at one point I was seeing Google twice as cheap or half the cost for right like / prediction so the same right like load was going in it was able to serve it up faster and you haven't really looked into any of the like respects between those two or which ones I was getting at that time but at like something to like keep an eye on and then think of the abilities to sort of auto-shift right like we know about auto scaling and scaling up throughout the day scaling down throughout the day now right look you can cross clouds and at any given point in time say in the morning right like we just lost all of our spot instances which are cheaper instances within amazon and now maybe Google instances are cheaper so we could slowly shift traffic over right like based on this rightly price / like prediction yeah also to the obvious like high availability right we have it's now multiple clouds that are able to serve the same load so if amazon goes down right yeah so god forbid s3 goes down right like when's the last time that happened we can definitely shift over all right so i had mentioned before in these tests i purposely isolated out the prediction cluster now prediction cluster is so i'm using netflix based microservices right so that has a lot of the circuit breaker stuff built in that has all these metrics built in I'm actually pushing also like along the lines of metrics here I've got Prometheus going for me Theus is very very very Google kuber Nettie's friendly but it's really totally generic here so there's kind of the generic like Prometheus then of course / fauna on top right so all this is is there for you guys to play around I think I have once I killed it might kill the dashboard yeah ah yes I killed it right for coming yes so you can write like build these dials and things like that that urn here these like gauges right that's one of the types of metrics early gage would be like us like your speedometer on your car kind of thing and that that will come in handy in a sec here in your show yeah so I purposely split out so that the prediction cluster isn't competing with the spark jobs of course and you know tensorflow builds and all that kind of stuff so if i show you guys live the cuban Eddie's cluster here on amazon we should see i could filter by host sometimes this thing likes to bounce off the pods all right there's pods five minutes oh my god oh right because of like QA is that this gotcha okay yeah so this is your like prediction cluster here so yeah I haven't gotten to it yet but there's really different ways to deploy things there's just generic kind of key value so that's this guy key value there's tensorflow right which has the tensorflow runtime it's got tools to like optimize the actual graph afterward I've been using p.m. ml with my sparkly male models it's pretty handy that it seems to work for most of the just like vanilla spark ml I've got a generic Python thing so you can actually deploy like code out there you're scikit-learn code that kind of thing and it's going to execute that and this is all wrapped in these netflix services right so if your Python likes like it learned stuff starts to get a little wonky like it'll notify it'll be circuits broken you can compare Python or you can like compare scikit-learn to your like spark and male models to your like tensorflow models all within the same framework here so all right so let me keep going tensorflow serving is pretty much the the heart of serving models with tensorflow they they're like a little bit weird about upgrading the actual version so you pretty much always want to be on master that gets you in trouble also you know for like demos and things like that it's a bit scary they're not as like diligence as the actual tensorflow poor team but so yeah here's some like model deployments I want you guys to think about and write like just like generic key value pairs so just think of where like offline generating recommendations so that so right this something Netflix would you write the for each user ID just have the top right yeah right like the top 100 like movies that that this person when they like go to login that should be served up so it would literally be key of the user ID and then value is this array of early video IDs three minutes all right and so tentacles serving funny enough right like you always think of tensorflow is kind of a neural network system there's actually this we're like generic like concept here hash map and yeah so they realized why I mean yeah we're rightly not just serving up neural net models here but we can also serve out key value pairs so they must have had that use case pop-up and so much that yappy FML it's a four-letter word here in the Bay Area right like we all know this I see it used I've seen it customized I've seen in Riley all different ways a lot of people are using it I was just doing a lot of like performing stuffs with it and actually ended up being pretty handy there's there's a couple other ways I want you to think about deploying code right you can actually take p.m. ml or take some you know like take some sort of source like like spark ml and actually generate right like native Java code native C code the cool thing about generating native Java code is that this can now be part of the just-in-time right like the whole piling thing that I like these JVMs do right it can do the inlining and all that and of course tends to flow model exporting this is kind of a handy tool now this this tool is called well yeah actually freeze graph so like just real quick on tensorflow when you build your model what you end up with is a directory right so it's very rare that you get just like a dot model file or whatever so one part of that directory 1 files in that directory is a static graph so that's the dag right that's all your operations that's your matrix multiplies that's all that stuff there's a second file for the right lip for the trained weights themselves and these are called checkpoints the new version of tensorflow 10 makes us a little bit cleaner changes the API like a tiny bit but it's a little bit more understandable cool so oh and so what you have to do when you actually go to to serve these models right like you have to call something called freeze graph or that's kind of logically what it's called it's physically the name of this Python script here and it's so what would be a variable inside of your graph you know for the matrix multiple the matrix multiply operator the script goes through and slams in the current trained wait and put it rightly from that check phone file into the actual variables and so they become constants and so the reason that these checkpoint files kind of live separate is that you could always pick up from that like current checkpoint file and then continue to train right so those would be the starting values and then just keep training training you could potentially over fit yeah couple things to think mutable not so I mentioned before about building a new doctor image one docker image / model right so it's a nice way to kind of get traceability all the way from the commit hash back and write like you would use that same commit pass version for example on your doctor image so that you always know what's out there a couple other things this is this graph transform transform tool is sort of the replacement for some of these smaller scripts that or write these kind of early scripts to optimize things like see if it's taking a little bits come up but think about like during the training process right you're doing things like drop out right so these are very specific to the actual training process when you go to score right for that for like the inference side you want these operations to not be there and yeah so drop out slightly different in that sensor so typically you on the serving side you would overcompensate for what was dropped out yeah but actually come to field knows this and so right like during training it's doing tricks in there to actually write so that nothing has to be done so that you can cleanly just remove those like drop out nodes but yeah things like summaries you know if you're writing out like summaries of your variables during the training process those are not needed during the inference so yeah check out graph transform tool Pete worden was picked up by Google to do sort of small device and 8-bit quantization things like that you train at full 32-bit like precision and then you actually when you're done training you would just lop off and go back to an 8-bit in and because this is all kind of fuzzy anyway right now on the inference side it's not going to be the exact like predictions that you would have gotten what with the full 32-bit but it's much much smaller and on GPUs right yeah this is actually like quite a bit faster because you can pack them in make them more dense one really cool thing to keep in mind is it a hard stop at 140 is that oh the next speaker starts ok so this concept of like of this like so it yeah so if you picture that top request so I like for whatever reason this model needs to do this matrix times vector for right like what it's trying to classify something it just wants that result so that's one request coming from one device and right like Japan somewhere that next one right like the server happens to get both of these requests at the same time to do like matrix times vector and like the sort of naive way to do it or the sort of low throughput way to do it is to just do these separately right so pass them both in and picture gpo right like on a GPU where it has to then channel through pcie express or like whatever so now what you can actually do is wait you know 5 milliseconds 10 milliseconds so here you're like trading off latency for throughput and so let's say this is the two requests I picked up within some window of you know 5 milliseconds or whatever I'm going to pass both of those into a little padding here because I know the offsets on the way back so this one two three four just to make it clear is just this guy and this guy and these two are going to put rightly edge to edge right there and so this is the exact same results based on these offsets and now you can actually just return to each caller that's specific so right like with one big matrix multiply yeah huge huge improvement for throughput if you have low throughput it's probably not worth it because you do have that five or ten millisecond latency all right so demos have a couple minutes here let me just kind of yeah this is like a little bit janky like mainly because I haven't built any real tooling into a rally Jupiter notebook I just got to sit down and figure out tornado I guess but you know think of this happening behind the scenes were actually this were so we would build a model it would end up in this directory 0 0 27 that's a new version we would commit it we push it that we start the new air flow you could actually watch our flow from within here you can actually do shift enter you should be able to see it here's the Cuban Eddie's cluster and yeah that's pretty much all i got that will fit in yeah this is much less than four hours [Laughter] any questions no time for questions all right that was my favorite question thank you [Applause] [Music]