Devreal

Inside NVIDIA’s AI infrastructure for se...

Event: Scale by the Bay

Scale By The Bay 2018: Clément Farabet, Inside NVIDIA’s AI infrastructure for self-driving cars

Recording: Scale By The Bay 2018: Clément Farabet, Inside NVIDIA’s AI infrastructure for self-driving cars

you and thanks for having me and thanks so so tonight I'm gonna tell you a little bit about what NVIDIA has been up to in terms of building up our internal AI infrastructure to support massive scale projects like self-driving cars also you know busy and involved supporting other internal projects but to me is what we're doing for self-driving cars is really unique and and worth talking about because of the scale involved and the challenges we have in developing the systems my name is chemin Farah Bay I've been with Nvidia for almost three years now and I'm a VP of AI infrastructure out there and so this this product I'm gonna walk you through is really one of the biggest things were focusing on today in my team so before I dive in and start telling you about what we're building it's worth kind of like recapping where Nvidia is at in terms of self-driving concept well and what our strategy is as you guys know and video is historically a gpo company we we've you know it's a company that's been selling graphics cards and then building software to enable those graphics cards to be used seamlessly into video games and then a few years back it started building this great piece of software called CUDA and since then and NVIDIA GPUs and it's been used in all sorts of high-performance computing applications one of them being deep learning and deep learning has been really you know tapping into the GPU power massively over the past six seven years in fact around 2010 eleven folks working in deep learning labs all started tapping into GPUs and that's really how we started getting amazing results so I remember back in the days I was doing my PhD with yellow cone at NYU and the the the GPU was literally the thing that took us from you know scrappy demos to things that actually started to work and so and so could has been the Olin owned what were doing with drive is kind of like one of these big verticals we're taking what the idea is that we're building the full stack from an SOC that's that super high performance it's a family of chips that go into the car and that was selling to to om scale companies and so on we call that drive a GX there's a family of chips the most high-performance one we have on the market today is called Pegasus it's literally packing something like 300 there are ups of compute power in that SOC and so you can imagine the amount of stuff you can actually pack into that car and the reason we are doing that and the reason the entire Terra industry is moving in that direction is that we believe that the future of channels is going to be entirely safe quality find in fact in the previous talk we had a quick example of the the Ford pickup which uses something like 150 million lines of code was it some pretty pretty large number and and believe me a Ford pickup truck is not a self-driving car yet and and so and so where we're going is we want to enable the entire industry to essentially go put that in the car and program this whole thing and build up the entire AI stack all the way to two high-level planning functions mapping functions and so on so why is this challenging so it turns out that you know to drive a car autonomously you actually need to implement you know more like north of 2025 unique functions that are essentially mapping everything that surrounds the car into some kind of internal representation that you can use to safely locate the car and then actuate the car in any environments so these functions include anything from you know safely detecting objects around the car like you know trucks cars and pedestrians and so on to things like estimating the distance of all the objects around you actually mapping them in some kind of like 3d representation to things like building an understanding of path and that's that goes well beyond just recognizing lanes it's really about you know estimating all the possible paths that can be driven and taken back by the ego car and so these functions are incredibly complex but what really makes them complex I think is two things one of them is the diversity of environments and scenarios that you have to actually get them to work in so it's not enough to to you know train well one of these models and get them to kind of like do the right thing on the highway on the other one you actually have to go get them to work it literally any environment you know all types of roads all types of weather conditions night conditions when it's raining when it's foggy on when it's like smoky like you today so it's like all these weird types of light conditions and and and whatever you do there's always gonna be this tale of things that you haven't seen before so it's really a super-tough I think it's one of the most challenging machine learning I've had to work with for the an etic with the main reason really is that you know for many machine learning problems we have this this this this you know for most products it's kind of out here to deploy something that's not perfect anything that kind of beats a human written baseline is enough to ship where as self-driving cars has this kind of like binary step function of you know you have to reach that level of quality before you can actually put it on the road and then and then so and it gets me to my next point which is that it's really all about testing so for most problems that we get to work with in machine learning training data set is really the big problem test data sets are typically smaller for self-driving tells it's kind of the opposite the scale of testing we're talking about is really daunting and the reason for that is that the industry is basically basically wants to be able to prove that all the functions we implement beat human performance and we have statistics and how many accidents and and issues human drivers cause and real-world data and that typically pretty good and so and so this basically leads to these types of neurons we think that we're gonna have to essentially validate every single neural network every single function we implement millions of miles of real data the real data that we're gonna go collect on the road that covers all these types of scenarios and we're gonna have to vary that these networks at that scale and then we think we're gonna have to go all the way to billions of simulated miles so that we can extend all the the combinatorial order combinations of things that that we think we're going to be observing the in the real world so the scale is really daunting and I think that's the main point I want to make you the the red curve is kind of showing the the our roadmap in terms of like failures and how far we want to push these things and the green curve is showing how much data were collecting to support this this this one map and so just to make this very concrete this is where we are today you know you know you know and devil towards building these these self-driving car modules the first color and tells you about the the type of data we are dealing with so we have a fairly rich rig with twelve cameras mounted around all around the car Radames multiple radars laydowns on the car as well and is rig is mounted and roughly 30 cars that we have driving worldwide and we're roughly ingesting something that picks two to a petabyte of road data which and so the four folks working in large-scale web companies this doesn't sound like you're crazy high number but the thing that's really tricky about this is that all of it is actually used that every time by our team so today we have to read it roughly 15 petabytes of training and test data and that's actually our active set that we use to do machine learning on a daily basis and that's what really makes it complicated and and coming back to the the last the last slide that's what we are today and it's really a scary exponential curve we're gonna have to we're gonna have to cut off so much more educate the second column shows you the augmentation we have to do on these data one thing that's that's a little bit tricky about seven cars is that there's a lot of things we can get from the driver you know like some of our networks actually predict to the paths that need to be taken and they predict the roads based solely and imitating the drivers that collect the data but some of these things I showed before can't be done like that they actually need labeling so we employ roughly 1500 levels full-time working two shifts and it's a fairly big you know just maintaining this thing and and and fitting them is the right data and and and making this efficient is actually a big endeavor as well so we're roughly collecting leveling twenty million objects a month and that's covering a few you know like 20 unique models 50 leveling tasks and I'll say more about this you know the act of selecting the right data for training and testing is is is really important and third is the compute and storage so once you have all of this it's critical that you actually feed this into Eccleston 220 on butter's and and and and test them at scale so here we're talking about roughly 4000 GPUs in a class down fact that's 500 petaflops of compute and our goal is to essentially build infrastructure to to support the development and and all these models and and sustain a computer that care star so I think you know at that scale there's a bunch of fundamental challenges one of them is the fact that I think that you know as an industry getting AI and machine learning into products is really better than AG buy a few things one of them is our ability to develop the right data sets whether it's for training or for testing different industries have different challenges for self-driving cars it's incredibly challenging because you just can't have access to the true distribution the real data is basically incredibly diverse and scattered all over the world we don't really have access to the the target distribution of where the girls are going to drive ultimately that alone is a huge challenge once you collect the raw data organizing them into data sets and making that available to the team is already also really challenging because of the scale and then things like you know running the radix experiments instrumenting everything so that your whole methodology is metrics KPI driven really important and ultimately automation and I'll say more about that and our belief is that you know if we sell these problems and we have great infrastructure to organize all this work to develop these models at scale then we move into a world where adding more computing to this cluster is literally turning into more rapid progress which is not really true if you don't have enough infrastructure because you're typically bottlenecked and the number of data scientist in your team so back almost two years ago we essentially decided to start this project that we call maglev can think of it as our Nvidia's internal AI platform to support our industry grade-a eye needs and and this is really being driven by the scale and challenges of organizing our data and workflows and methodology for for AV so let me get into it and start telling you about what this this this platform is doing at a very high level we're trying to serve three big high-level problems the first one is data set lifecycle management from the moment the the data hits the data center tu-tu-tu-tu-tu when it starts being consumed by the team we want to have great a great methodology to essentially index everything and make it searchable and make it queryable for all sorts of downstream tasks and that's an all about that the second pillar is ml pipeline automation we want to have the ability to essentially describe our end-to-end machine learning pipelines and then schedule them of the underclass town and have them reproducible and I'll say more about that as well and and at the end of the whole pipeline were interested in model analytics and insights and and basically helping the team understand what action to take next to keep improving their models so focusing on the first you know pillar so data set lifecycle management is I think you know probably one of our most challenging problems again because of the scale of the data what we're dealing with and and the variety of access patterns that the team wants to be able to run so we've kind of like built two worlds the first one in green is essentially our large scale you know beta by beta bed scale management system and it's a pipeline where we're essentially writing a lot of the data into Parker and then providing access to the team through multiple interfaces so we're using presto spark hive and then exposing multiple front ends to essentially let the team compose datasets and query arbitrary datasets dynamically and on the right side we've basically invested a lot of effort into building Python side clients to essentially enable the folks developing the models to rapidly compose and assemble data says dynamically and let them sample from multiple sources and and and compose datasets at runtime and the challenge which went to self here is that you know companies with with established large-scale infrastructure always great MapReduce infrastructure typically tend to push a lot of the Welkin on the left side essentially any data set creation is going to be a big MapReduce job on outside we've tried to split that into to do less things on the MapReduce set of things and enable the client-side consumers to actually do much more at runtime the second thing I want to show you is the actual how we've kind of like organized this whole thing so the whole platform we've built is is based on kubernetes kubernetes lets us essentially scale and manage all the services we have today we're running roughly 20 micro services doing all sorts of things from managing the workflows we run on the crystal to tracking our experiments the data sets were consuming so and all these things are kind of like seamlessly running on top of kubernetes this lets us do one thing that's really interesting we went for a hybrid approach where a lot of our services and data management is actually entirely hosted on AWS which is really great because we can scale some of these services elastically so a lot of the ETL we have to do on the data we ingest needs to be adjusted that Emily sometimes we're going to need a few thousand nodes sometimes just a few hundred so we can kind of like the elastically scale all of that but then at some point the workflows themselves run exclusively on our on premise cluster we've backed up 4,000 GPUs so far and and our goal is to basically keep these guys busy with with training job training and testing our models and skip that the so I like and so this view is really interesting to me I think I think a lot of what we're trying to achieve with this infrastructure is to essentially start reconciling production engineering DevOps and and machine learning developers these two four typically you know don't interact much the machine-learning developers tend to work in a world where they write a lot of Python code that deal with fairly small scale data sets and they move very fast and are very focused on iterating on new and novel approaches to algorithms and how they're solving the problem whereas in the economic divert and in production engineering set of things the whole goal is to automate everything scale up the the the work that's been done to really large scale problems and so what we've done here is essentially a world in which we can programmatically describe workflows running on kubernetes an oculist on the machine developers in our teams typically describe these workflows and run them in ad hoc ways and they do a lot of science like that the workflows let them do things like hyper optimization schedule multiple jobs at scale compare results of multiple machine learning approaches and techniques and so on but then the beauty is that once the workflows are finalized were basically versioning them and then we have a way to essentially a schedule workflow straight out of get so that the whole end-to-end pipeline is reproducible and can be cron scheduled so ultimately that's probably want to be we want to be in a place where research and an rng on the models is decoupled from production but using the same codebase and so the production engineers can essentially schedule those things and the output of that is trend models version data sets metrics associated to each models gives us a really healthy way to track everything that's produced and deeper in poverty on the side effect of that is is so automation is really the top goal the side effect is also traceability and I want to draw a little parallel here you know for the past 20 years in in traditional software development we've really learned to manage traditional software on the right way we use versioning systems like get to trace and everything we do we rely on deterministic compiler ions so that any code version in git can be turned into an executable in a deterministic way machine learning is not there at all machine learning is still very wide but I believe we need to get to the same level and I'd like to draw an analogy where you know you can think of machine learning where the data we use is basically another use to to to source code and the algorithm itself like deep learning on machine learning is basically a compiler and it turns this data into some kind of predictor that you deploy in production and so what we're trying to do here is get to a point where we have the same level of sanity where every data set is versioned tracked and then our deep learning and machine on extract is entirely deterministic so that we can compile the data into our predictors and so it's a it's a massive waveform because the industry is not standardized on any of this yet and and we think we have a lot of things to contribute to that one more view which is perhaps the the the most concrete I have is showing you the two big clusters we're running on the top side it's our clouds cloud set of things running on AWS today at the bottom it's our on premise closed down with 4000 GPUs and it's really showing you the the pipeline of how we deal with data so we're ingesting a fairly large amount of data then that ends up being stored in that data Lake and then one of the main type of workflow we're running is what we call data selection which is really about applying pre-trained neural networks to go do inference on data and select which data needs to be leveled the most so one of the things you observe when you start building applied systems with machine learning is that at some point your performance starts plateauing and the reason it plateaus is that all the new data you sample and label you're actually classifying it perfectly well so one way to think about that is you know if you have a model that has 99% accuracy and random data it means that 99% of all the data you get is already well classified 99 positive it's zero there's just one part of it that that's really wrong and so if you keep labeling data randomly you're basically wasting your time 99% of the time and so what so and so this is basically active learning active learning is really this body of a research aimed at biasing the dead that you decide to level to keep growing your training sets and so to us that a selection is really really really critical because very rapidly you get to a point where your models work 99.99% of the time it's all going to be about finding the new edge cases in that whole pile of data we keep collecting and then the second big job is is testing so we're continually increasing the size of this test data set so that they cover more and more scenarios and we want to be able to test any mother which when at these ridiculously large scales and then of course training which which people are typically more familiar with one last thing and that the infrastructure we're building we've put a lot of effort into the hardware layer as well and the reason for that is we essentially spend our days training and testing these models and moving data around and so it's really critical that we don't have any bottlenecks across this whole pipeline our entire platform running on top of kubernetes is built against object storage so all our data sets are stored and object and-and-and-and typical AWS s3 and then what we've done is we get a replica of s3 locally and premise which is using the same API so it's all s3 based and then we've kind of like racked up our systems in a way that's hybrid where half the racks are dgx systems DG X's are these big fat nodes with eight voltage GPUs so these are monsters and we've packed none of them pound rack and then twelve CPU nodes and that gives you roughly something like 1 petabyte of distributed storage power rack feeding into 90 G X's and that's 9 times 8 is 72 GPUs power rack so it's a super high density model where data can be locally cached we use objective plies other way as it's kind of like a three-tiered system all the data is safely stored on s3 and that gives us replication and worldwide we can ingest in different continents but then the unparalleled caching and then ultimately next to the computer it's a cool cool cool design and we're starting to recommend that design to our just along so that they can deploy high-performance computer styles all right i'll step therefore for for maglev i want to move on to another project which is mainly related but largely aaanthor granola so this is something years that we've studied in our team a while back and we actually just launched this around a month ago now so it's it's all new we're super excited about this you can check check it out that Rapids that a I so the project is basically an open-source Python first and GPU first data science ecosystem of libraries and and conceptually the the way we've organized and built this whole thing is it's quite simple if you guys are familiar with spark what spark did for MapReduce and Hadoop was essentially starting to move away from writing data in and out of disks and start starting to do everything in CPU memory what we're trying to do here is essentially push this a little bit further and push everything into GPU memory so the idea is that we've standardizing your on Apache arrow for the format we're using for all the data we're dealing with and then our goal is to basically load everything up into GPU memory and then have the data never leave GPU memory and then build and provide all sorts of interconnects into multiple libraries one of them is called qdf you can think of could EF as a pandas clone so panel is basically library to create data frames and manipulate data sets and do joins and group buys and and solve data set and so on this lets you do the same exact thing entirely based on could in GPU memory then KU ml is a collection of machine learning primitives from k-means to XD boost were supporting XG Boostability decidual forests and so on our end goal here is to support the full scikit-learn api + XG boost as well were way advanced and XG boost we should have some work to do around a circuit loan and then KU graph is a whole bunch of primitives to manipulate graphs efficiently in GPU memory again and then the other cool thing we're doing is that to build bridges into deep longing this is all new starting to basically connect this this format to libraries like Python so that once you've loaded up your data sets into memory you can do some feature engineering some quantization and then ultimately pipe your data into a deep learning server so we're super excited about this one thing that's really great is that when it's open source we're trying to stick to two existing Python API is to make the unborn and this really seamless and the latest thing we've started now is we've started integration into a spark as well well the goal is to is to is going to be to make this also seamlessly available through the j'ni in spark so that you can schedule the same types of jobs through spark and one thing that's really great is is looking at the end-to-end benchmarks so it used to be that traditional ml was really hard to speed up on GPUs because compared to deep learning you have a lot less dense you know metrics multiplies and computes and so it's kind of like it's harder to optimize and one of the challenges with with these types of workflows in data science is that you spend a lot of time doing IO and feature engineering and ml is typically just one aspect of the whole pipeline but because we were taking a holistic approach and we're loading up everything into memory doing you know even CSV decoding and partly decoding is gonna be done in once it's in GPU memory you basically don't waste time on Io so you learn a lot of things at once and then you're basically unbounded and the the bandwidth between the GPU memory and the GPU is incredibly higher than the bandwidth between CPU and CPU memory so really cool stuff the the coolest result here is DJ x2 is what is the quest system that nvidia is building its 1616 GPUs packed in one node so an incredible amount of compute power the craziest thing about DJ x2 is that it has half a terabyte of GPU Ram in one node so you can you can start doing a lot of things without having to distribute anything realistically though systems like DDX one which I think 128 gigs of ram more common and what we've done here is is basically spelling out the whole data processing and machine learning through desk to distribute the load and again in terms of development the whole project is open source so we've been working with quite a few partners on that particularly the folks from you know anaconda blazing DB like very very tight collaboration with them in real and scikit-learn data breaks and the sparks I'll and I'm trying to find yeah and our cell labs who is basically leading the the ro2 definition they were working very closely with West McKinney who came up with arrow in the first place and and really trying to align on that so that everything we do is send other on that alright and obviously my team is hiring very intensely so if you guys are interested in any you know type of job at the intersection of high performance compute GPU and Big Data either for internal applications like what were doing for self-driving cars all these types of open source projects were building through Rapids feel free to reach reach out to me right we don't [Applause] yeah so we don't so right now this is our internal infrastructure and so today we really were focusing and enabling all our internal teams to to build these new stacks including self-driving channels but we don't have a good to market plan but as soon as we do with a bit weekend so yeah yeah it's a great question so I mean so tilde approach well so in each DJ's here we have seven terabytes of SSD voltage X and so this typically gets you to you know like like there's a whole bunch of like training that assess the typical fit in that but then you know a lot about test sets are actually significantly larger than that and so then we start doing everything everything is mapped and and and not even cached locally but it's typically cached in in the upper part of the rack which holds one petabytes and so one petabyte is is pretty much enough for anything we need to do at that point but yeah exactly exactly yeah yeah and infrastructure we've built and perplex like one of the service services we have running on kubernetes is basically responsible for the data management and so we have this system to essentially create virtual images so that the active said that's gonna fit into a job is pretty fine so it's not completely open-ended otherwise very hard to manage yeah yeah it's a great question so the so it's actually one petabyte of data and it's an incredible challenge and the only system we've found so far that has enough bandwidth is FedEx it has shitty latency that amazing bandwidth so we we literally use FedEx to to ship everything it's all SSDs and you know back in the days it used to be that you know the folks at Google Maps Street View used to do the same thing actually don't know what to do now but given the volumes of data you're collecting locally in one specific region you never have enough been with anyone to actually ingest fast enough and so we ship data and we have multiple recipients worldwide and then the the is basically disk base old school SSD base you know it push the SSD ins and it's except everything to the cloud yeah thank you so much