Devreal

Less Code, More Intelligence with PyTorch on SageMaker

Event: Featuring PyTorch with FB, Autodesk and AWS

SF Scala: Emily Webber, Less Code, More Intelligence with PyTorch on SageMaker

Recording: SF Scala: Emily Webber, Less Code, More Intelligence with PyTorch on SageMaker

[Music] from from now until the expected start time we're just gonna knickers so good evening my name is Emily Webber I am a machine learning specialist Solutions Architect at Amazon Web Services what the heck does that mean I ask myself the same question the the reality is that I spend most my time working with customers I work directly with customers across every stage of their cloud journey we're talking very small companies very large companies you know all all over the spectrum here um I feel like my actual skills kind of fall somewhere in between architect plus data scientist plus research scientist there there are so many interesting ideas in the space but it's it's really fun to stay up-to-date I spend a lot of time working with customers who are new to machine learning particularly folks who are quite technical especially folks who are familiar with AWS but may be quite new to the to the machine learning space and so when I'm working with my customers I always like to actually ask a pretty broad question and the broad question I like to ask is well what the heck is a computer right I think a lot of the times you take this question for granted right particularly in today's world where many of us you know wake up in the morning and we start looking at our phones then we continue looking at our laptops and you leave looking at her phones and so it's it's actually a it's actually an interesting question for my standpoint if you look back a computational history right so the first person who proposed what a computer could be right was Charles Babbage so that was back in the 1800s and that was Babbage's picture right he called it an analytical engine and the goal of the analytical engine was to solve any arbitrary problem and so you formulate your problem on the right hand side and then you shoot it down the center and it will break out all the sub components and figure out the answer to this problem one of the first people who criticized this idea said hey you know that sounds useful but it won't be able to do anything new that was actually Ada Lovelace right so Ada Lovelace is known as being the first programmer on our planet but more than anything then she was a thinker she was an intellectual she was a scientist and so she's reading through Babbage's proposal she says wow that sounds really useful but again it won't have a life of its own won't have any creativity and a sense of originality that became known as the ada lovelace critique when you look at Alan Turing's work what's interesting is he find that he was actually obsessed with artificial intelligence right touring was convinced that it was possible to not just design a computer to perform general intelligence but that that computer could be so intelligent that it would mislead another person and that that person would think the machine was actually a person and not a machinist so that's what the Turing test actually refers to when you speak with folks who are based in in London and the UK about artificial intelligence frequently they actually start with Tory Margaret Hamilton was the first director of software and engineering on the planet because she invented the term software engineering software engineering was actually invented by women right and this this was from Margaret Hamilton if you're not familiar with this picture this is Hamilton standing next to the suite of code that she herself wrote to land Apollo on the moon prior to Admiral Grace Hopper which is this picture down here computers literally just did math right that was the extent of what a computer can do was arithmetic and so everyone told Grace Hopper that that was all a computer could do and so Grace Hopper says to herself in order to productionize this technology what I actually need to do is develop a way for people to write something in a language that makes sense to them and then translate it in a way that computers can read and today we call those compilers Catherine Johnson became especially famous recently in the film hidden figures and so she was back in the early day when a computer was a woman who did the math for the men right that really happened computers were women who performed arithmetic for the scientists who were men and so Katherine Johnson was actually necessary to launch multiple spacecrafts because she was the best and so she kicked some serious butt and then Carlos question has been a key figure in helping actually boots come to life outside of the NLP world a vast majority of customers are still performing classification right by and large classification tends to be the first way to solve many a machine learning problem and so actually boost is still a very common algorithm and so I bring these up for a number of reasons and there we go personally I have two theses if you will the first thesis is that I just think that since computers have been around people have always wanted computers to be intelligent and they've always wanted computers to have their own sense of agency and so what I think we're seeing today is less about this is a new fad that's gonna go down and it's more about this is actually where computing has always gone and personally anyone who knows me knows that I'm a very firm believer in in diversity and inclusion I have a number of projects that are directly related to that and so will be I'll be walking you through those so in terms of AWS for AI ml there are actually just three things that we really want to hit and the first one is called s3 so hands up if you have worked with us three all right beautiful recruit in so s3 ec2 uh and customers right so Amazon starts with customers and we work backwards from there I cannot tell you how many sage maker workshops or tsuge maker workshops primarily I've been involved in where I get through and here's what sage maker does and then people stand on the room say oh but I actually need a model repository or I need batch transform or I need XYZ and then what happens is that actually gets incorporated into our product roadmap 90 to 95 percent of our roadmap comes directly from customer requests so don't be shy don't hold back tell us what you want to see and then and then as I like to say wait until Christmas comes and then we'll see what happens at the end of the year so in terms of machine learning and AWS is really interested in helping make machine learning easier than it has been machine learning is historically very difficult challenging reserved the elite few and what we want to actually do is open up access to this so to literally put machine learning in the hands of every developer and data scientist and that leads us to what we call the machine learning stack well we really want to understand here that there's actually a very clear philosophy guiding how Amazon thinks about its services and particularly around AML and it comes in these three different flavors and so essentially these are stocks in a these are stacks in the technology stack right these are layers in a technology stack at the lowest level if PowerPoint is working with me okay so at the lowest level of the stack we have what I call customer managed services and essentially these are pieces of infrastructure that you as the customer can manage the end-to-end like lifecycle of this this can be an ec2 instance at least can be outpost if you're trying to run on local content this can be wavelength if you're trying to put together a 5g solution this can be a local zone if you're trying to deploy something that's actually near a metro metro area so there's a lot of innovation that's happening at the lower level of the AWS stack I'd encourage you to check out some of the reinvents keynotes that happened in December to take a look at that I spent most of my time one level above that and so I call that AWS managed because AWS is managing the infrastructure we're turning on the ec2 instances you use them and then we're turning them off we're pip installing all the packages we've got lots of software we've got lots of features built out for you so you don't have to write as much code there's literally less work required because we're doing what we call that undifferentiated heavy lifting we're actually building out all of those basic steps that you as customers do on a day to day basis and we're adding that to the service so you can move faster but on the highest level of the stack is what I call fully managed and these are fully trained machine learning models that you can leverage so as a developer you can hit recognition like an API and get a classification response for the content in that image or you can use comprehends to automatically extract sentiment keywords you can actually build a recommendation system using personalized and your own data you can build a forecasting system you can leverage code guru to analyze the code in your commits and make sure you're using your cloud resources cost-effectively huge amount of innovation that's that's going on with the top to the stack and so what we actually need to do is pick the solution that works the best for us and so there's overlap here and that's okay because sometimes you'll start at the top tier of the stack right because maybe you realize that you can actually use personalize to build your own recommender but then you realize that you want to add another component so you'll drop down to Sage maker but if you have really extensive you know a simulation software you want to run you might just spin up another ec2 instance there is a tremendous amount of action in terms of machine learning on AWS the reason why this matters to you is that in the event there is a single customer up here who has a workload that is similar to yours in the slightest degree the odds are they've spoken with us about it and the odds are they've actually given us feedback about what they want to see and we're acting on that feedback and so what that means is that you get access to all of the innovation and feedback that we get with our customers basically for free which is pretty cool if you didn't know this my slides are not working there we go if you didn't know this the NFL uses Amazon sage maker right if folks were watching the the Superbowl that happened here it's the NFL uses Amazon sage maker for next-gen stats on the way this works of football players are outfitted with IOT sensors helmets shoulder pads cleats the ball and as they're running down the field we can capture all that IOT and telemetry data you can combine that with a twenty two cameras that are around the field and then label it with the commentators actually so you can label it with where folks have run on the field and then perform drought predictions so you can actually take a picture of the field and draw a line on the screen for where the receiver is going to run prior to the play even starting and so this is in act this is in use currently so cool stuff so undifferentiated machine learning life cycle here this is this is what happens you build your models we need to train them but then we need to deploy them and so when I work with teams most custom have something in play for build and train it might not be the most efficient use of their time and it might not be giving them the most value they can get but they're able to get it done it's it's a little bit more rare to come across a team who has a really slick solution for for that last step and so Amazon Sage Maker is a tool that was designed to meet each stage of this workflow so let's break this down so step one on em some stage maker is a managed ec2 instance and so Alex did a great job of calling out resizing the notebook instance so that means we can start running your models and writing your analyzing your data on a smaller ec2 instance and then you can just resize it when you actually need something larger there are 17 built-in algorithms in Sage Maker so if you want to just use a built-in image classifier you absolutely can you do not need to and we're gonna cover the other examples here especially because we want to bring our own PI torch models there is what we call a machine learning training service which is to say when you fit your model we're gonna turn on a new ec2 instance we're gonna copy your model on to that instance we're gonna get your data out of s3 and on to that cluster that cluster is gonna be up for the number of seconds your model is training these second your model finishes training the that cluster comes back down and you can go on your merry way we have a managed hyper parameter tuning service this is using Bayesian optimization and so you can leverage Sage Maker to automatically find the best hyper parameters again with our Bayesian optimizer there is a hosting service so after you have your model defined and trains with a single line of code you can automatically create a RESTful API around model we're gonna cover some examples of how to do that but literally you just get your model and then you say model dot deploy we're gonna create that web server for you some folks needed on a batch schedule so if you have a time-based solution you are solving for batch transform is a great option we've got a bunch of SDKs and a lot of content out there and so to hit this home there are actually five ways you can train a model saij maker alright and so number one is the built-in algorithm and when I'm working with teams usually that's the one we start with because the infrastructure is slightly different and at least for folks in Chicago usually it takes a little bit of time for us to really get comfortable and proficient on the new platform so we'll start with the built-in algorithms but I definitely encourage folks to grow their skills and move into you know a place where they can write their own models such as in PI tours and so script mode and docker container are two ways that you can write PI torch models on Sage maker the docker container means you're literally bringing a docker file you're copying your code into that docker file you're building your docker container pushing it to ECR and then dropping it onto the stage maker cluster script mode is a lot easier script mode is where you can literally just bring a script drop it into our pie chart estimator as we learned about last session and then run that on a trainee job employee the AWS marketplace has over 250 models that you can actually just say use last fall I did a demo at strata on a solution we called mansplaining it's it's an Alexia time to laugh and so it's a so it's an Alexis skill where you can be recording your meeting such as with Skype or or chime but you record your conference bridge when you're finished you get your mp3 or mp4 file right you drop that thing in - three bucket drop it through my pipeline I mean use transcribe to figure out when everyone is speaking and then I'm getting at the timestamps and use transcoder to get these tiny tiny bite files I can hit a model from the marketplace that does audio gender classification write back the analytics drop that in a dynamodb table which hits my alexis skill and then it can say alexa analyze my meeting and she'll say men were speaking for 90 percent of the time in this meeting even though they were twenty five percent of the participants so fun stuff are any Phills team uh no sales teams are not using that yeah no but we where I'm at personally I'm looking for customers on that solution so if you're interested let me know no so so this solution starts with a recording so its start no no no the recording comes from you you you record it like in chime and then chime or your web conferencing bridge creates an audio file Alexa is only receiving and then playing back the audio okay yeah yeah that's it that's all it's a one-way though it's not a two-way you can record it anywhere yeah it takes it it takes an mp3 file it takes an audio file yeah question in the back yes yes so the question was do you have model performance monitoring the answer is overwhelmingly yes we're about to get to that so I am talking about sage maker classic right now and I'm going to talk about the classic infrastructure so we can really understand that then we're gonna flip a switch and look at the reinvent launches from 2020 and see that it's actually doubled which is really exciting so the most common place that machine learning is developed is on a laptop there are clear upsides to developing on your laptop obviously this is very flexible right you can do whatever you want to this is very personal and it's very easy to get started my friends there are downsides to developing on your laptop the number one being that obviously you can't scale as soon as you want to it if you want to train your own Bert model you can't do that on a laptop I'm sorry to say or if you want to fine-tune a Bert model on a data set that is larger than will fit in the RAM of your computer it's going to break your computer right we've we've we've battle-tested that we know that that one is true and so obviously you can't scale you can't run an application in production off your laptop as soon as you close your laptop your app is toast right that's that's a sad fact of life and then the last step in our space there is so much innovation right there are these new solutions that just come out obviously PI torch has all these awesome things that we want to play with but in order to get our hands on those things or across some of the other open source packages that are out there the first thing we actually need to do is upgrade pips upgrade virtual environments create a virtual environment pip install all the solutions into that virtual environment and try and make sure that all of my software dependencies in that virtual environment are compatible with anything else that I need so I'm not breaking my machine and if you're a person like me that's actually really challenging and that's gonna take me a solid four to six hours to do on a good day even worse and so what I'm happy to say is that on Sage Maker you literally don't need to do that you just turn on a new machine the other place that machine learning is commonly developed is on servers right and we sometimes feel really really familiar with the servers right for for lots of folks that says feels very familiar but there are actually some clear downsides to to developing on service I'm not gonna go too deep into this but this is a core value proposition for why cloud is helpful which is to say teams are usually stuck in this in this path of over or under utilizing the servers right let's say you're an innovative company and you just had a really successful hiring season and so you have a whole new fleet of data scientists which is great but in your data center you only literally have two servers and so if you have 12 data scientists are you gonna have each six of those people hitting your two servers that's actually gonna break your servers right because they're they're gonna be over utilized those nodes are going to go down and so then you say okay well here's what I'm gonna do I'm gonna order more servers and it turns out it's actually gonna take me about six months to get that order through procurement to get those servers actually shipped to my place rack and staff provision and configure them meanwhile my data science team is sitting there doing nothing right and so that's under utilization and so this is why companies are moving to the cloud because you can elastically blend your compute to your needs and so the key point that I really want you to understand is that sage maker gives you dedicated compute for every stage of your flow so you getting dedicated notebooks dedicated training jobs and dedicated endpoints on top of a lot more these are your notebook instances literally we're talking about an ec2 instance your ec2 instance has an EBS volume already lis attached you can actually drop a GPU on there so you can just attach what's called an elastic inference it's a portion of a GPU attach it to your ec2 instance a whole bunch of things we're gonna cruise here it is very common to clean your environment with a lambda function so you can leverage lambda to basically turn off your your notebook instances and make sure they're not being utilized common flow on a training job you're on your notebook instance and you call model dot fit once you fit your model we're gonna shoot your data out to us three we're gonna turn on new ec2 instances those ec2 instances are dedicated to you and your model no one in the world has SSH access onto those machines right they are managed by Amazon and that is that you have the elastic container registry that is hosting your algorithm if you are putting your PI torch code in a docker container and then registering it it's gonna live there otherwise you're gonna avoid it and use the script mode which we're going to talk about but in any case your image is downloaded to that ec2 instance or to that fleet of ec2 instances we ran a benchmark test last summer on nine terabytes of data in a single extra boost model across 45 machines and it trained in 95 minutes so you can really scale up or scale out rather and then you're gonna write your trained model back to s3 so you get that model artifact after your trainee job is finished again the whole process comes down no to the training job no no that's that's fully managed a really good thing to know about is that you can decrease your costs by up to 90 percent by using sage maker spot so we talked about this sage maker spot is a great way to again just save on your cost of your training on spot instances again just to hit this home automatically create a restful api or on your model we're gonna go really fast here because we're already a little bit over it is very common to actually develop a multi account strategy so under no circumstances do I want to see company run their entire business and a single a tub use account that just turns into a nightmare multiple years down the road what you actually want to do is develop a multi account strategy so over there on the left hand side is where your resources are running in praat it's got a prod ec2 instance if you're responding to millions of requests an hour which apparently isn't that much traffic say my friends in in caching and micro services so if you're responding to millions of requests an hour you can actually cache those in like Redis to to get model responses faster Rattus responds in less than a millisecond sage maker the endpoints are gonna respond in a few milliseconds so single-digit milliseconds but again Redis is going to be even faster than that drop an account next to that and use this kind of buffer account to promote new models so to push new models into production and then demote your logs obfuscate your production data drop that back in your data science account where you can set up a dashboard and have your team's actually analyze the result of your model that's running in production so exciting stuff it is very common to also leverage batch transform so batch transform just starts with cron job you're gonna hit lambda that's gonna hit s3 and then sage maker yada yada yada so that's based on again a time of day all right we're gonna very quickly here walk through all the new features and then we're gonna jump into PI torch and then we're gonna call it a day so let's let's get through this so sage maker studio is now in general preview in Ohio so sage maker studio is a fully integrated IDE we wanted a solution that would feel very very natural to developers and that would give us all the productivity boosts that we heard from the field over the last couple years was was really helpful so sage maker CEO experiments you can use sage maker to manage your experiments when you have projects and you want to see how well your team is performing and what activities were done you can use sage maker experiments processing jobs if you just want to turn on an ec2 instance analyze a bunch of data and turn that off processing is for you we have a debugger so you can run a training job and then actually debug what that training job is doing why your gradients are vanishing and then keep keep cruising autopilot is a solution here auto pilots will look at your data run feature engineering based on just general stats figure out your prediction problem in some cases and then train your model kubernetes operators model monitoring multi model endpoints and augmented intelligence and now we're gonna bring it all home here with PI torch and sage maker so number one your general rule of thumb here if it fits and docker it's gonna fit in Sage Maker so just really let's let's internalize that anything you can do in docker generally speaking you can do in sage maker so that means you can copy your code into this docker container specify your script however you don't need to touch docker if you don't want to and so that is why we have what's called script mode and so script mode generally means that I have a script my script is sitting right here maybe in my local directory and I'm gonna pass that in to what's called this PI torch estimator and so the pie chart estimator is just using the PI torch container so it's using the the AWS managed PI torch container that we have then I'm actually about to show you and then you can just run your script using our container so you can use our container and you can actually extend that container so let's say you want to import sort of our container as your base in a much smaller docker file and then extend it by just running your code in that container you can actually do that number four you do not need to write your own web server you can absolutely train a PI torch model however you want you don't even have to tree know and sage maker if you really don't want to but you can still deploy it so number one you're gonna use this SDK right so using our sage maker SDK to reference a pre-trained model that model is sitting in s3 you've got your sage maker role and the script right and the actual the the pie-chart script then again one line of code predictor dot deploy model dot deploy rather model dot deploy number of ec2 instances and the type so you can default to having two ec2 instances plus an elastic load balancer not San Antonio because it's at the application level so again you're going to use this when you don't want to manage your own API number five yeah question so it varies based on the framework and the the source you're using most of our Python examples are using flask it's using the sage maker service and then it lets you build a micro service architecture number five you do not need to scale your training jobs yourself so here is an example using a reinforcement learning estimator we have we have loads of reinforcement learning examples should that be of interest and you can specify a Python script right here number of training instances and you can really just increase that we were learning earlier about different ways you can distribute your your model and your data and so that that definitely holds true on Sage maker there's a lot of variety there and I'm about to show you some of those and so this is one way of in your your training script setting the number of your hosts and then basically looking at a at a global level I'm sure there's a better way of doing it but that's that's one way of getting it done so some learning resources here for you at risk of being self promotional there's actually a reasonably nice YouTube video series so I have eleven videos that you can watch I think we just hit 33,000 views so I guess there's some market adoption there but 11 videos I go very very deep on all those components you can pick fast forward rewind whatever it is you want to do so that's if you like consuming via content are via videos we have 250 example notebooks that you can just use right off the bat so this github repository is probably the single best place you can use this to evaluate all sorts of capabilities all of our new features have their own set of example notebooks so you can walk through those to gain some familiarity and with that thank you very much definitely feel free to reach out on LinkedIn I love staying in touch with folks and I hope you had a good time [Applause] okay any questions yes in the back yeah us yeah so we have a feature of stage maker called inference pipeline that lets you connect up to five containers and then in your five containers you can just operate those in serial so just one right after the other you can have up to five on top of that again if it fits in docker it fits in stage maker so anything you want to write in that image is up to you yeah yeah so there I mean that there is a there is a small of charge but if you look at the entire lifecycle of the amount of you know engineering that's required to accomplish the same set of features on your own then it's a it's a it's a clear one-way bet there we also asked you specifically for the notebook instances not for the training part but for that yeah so that I mean the pricing varies all the time so it's possible that you know there was an update recently but I know you had a question earlier about you know like can we just do it for free and so there is nothing stopping you from creating AWS accounts that just run on the free tier I do this myself actually I just create a bunch of AWS accounts like literally just keyed off of my email and those those are just so every aw secound do you create has a free tier that free tier is based on the length of that account that's based on the compute resources that using so you want to monitor it closely but again there's nothing stopping you from just creating as many of those as you need do you recommend any tools for monitoring performances models like weights of balances or anything like that yeah so I mentioned earlier let me cruise over there so one of our many reinvent updates it's called model monitoring and so model monitoring is right down here and this is going to monitor an endpoint that is running in production and it's gonna let us monitor the level of the statistical data that is hitting that endpoint and so for all of the features in your data set you can actually set a baseline for what you want the stats on that data set to be and then if the nature of that market actually changes over a period of time you can detect that via model monitoring and we're intending to to improve that over time all the services and what would be there and how does that compare with accurate yeah so I mean is I mean essentially we actually see that it's really helpful when so essentially AWS is very committed to lowering our costs of operating over the past multiple years that AWS has been in operation we've actually lowered our price many many many many times that is because every single team has goals around lowering their costs of operating so that we can pass those costs directly back to the customer I don't want to go into too much more details out there now but essentially we work really hard to lower cost across all of our customers yes so that is another very interesting area I know we absolutely have some folks who are taking a look at that but I also think there's an opportunity around some of the solutions in Pike perch and specifically how they could be connected I'm just hypothesizing here without mental intelligence because what mental intelligence does is it leverages our managed data labeling solution which I didn't go into detail here it's called ground truth and ground truth lets you take a bunch of data and then specify how you need that labeled whether that's image text you know just numerical data augmented intelligence lets you run a model in production when the confidence level off of that model is below a threshold that you can sense then that's actually gonna trigger a manual review so that's gonna trigger a person who's actually going to review that model prediction and then can reclassify that data and so I think it'd be really awesome if we integrated a pie chart solution there and so that after a model is running in production you could potentially see not just the data but also the the weights and/or the gradients to see exactly what was most informative so that the person can then correctly real able sometimes we have a conference coming up Romar's if you don't know about this this is a new conference that Amazon is doing that's gonna be coming up in June here in Las Vegas it's fun because all the lots of leaders from Amazon will come in and talk about how they're using machine learning across their business we don't really have time for this but let me just show you a couple slides while we're still here last year we shared some information publicly for the first time about Amazon go yeah so this is go so essentially what's going on here so this was shared publicly for the first time that remar is last year so over on the left hand side we see the raw video feed of people who are in stores right people who are actually in the NGO stores over here all the way on the right hand side we see what appears to be a binary semantic segmentation or a pixel mapping right so that's a binary prediction problem where we're actually extracting the people from the background then you see the next in what I call like a logical decomposition then we see this this number 3 which is a multi class problem right clearly we're seeing heads shoulders arms hands and then we see step number 4 and this this fourth component here is a stick-figure representation right so it's interesting is that again because there's this this logical decomposition that's happening that that pixel that stick-figure representation is actually much easier to classify the last piece I want to show you again just because we're here and so this is a drone from our primary solution that was also announced at Reimers last year and this drone was produced in part by a machine learning model so I see a lot of interest in using machine learning to help improve product design actually because for many companies particularly when you've been in business for multiple years you have a lot of data about products that you've tried and products that have worked versus not worked and so you can imagine just building a simple binary classifier right to automatically compare a new product or a new design against the database of historical designs that you have and then automatically giving your team's feedback about how well their product their proposed product is predicted to perform relative to your historical ones so definitely exciting stuff there all right I think we are solidly out of time thank you for for sticking through to the end [Applause] you