data.bythebay.io: Subbu Rama, The Promise of Heterogeneous Computing
Recording: data.bythebay.io: Subbu Rama, The Promise of Heterogeneous Computing
hey nice to meet you it all seems like a very crowded room so I'm going to basically not talk about a little different thing in our lab so I probably know might be actually know about hearing people talking about algorithms and applications so I'm you know as these you know deep learning applications can become really prevalent one of the major challenges is going to have is you know how do you compute them faster and there's actually a huge trend wherein off course you know people have been using CPUs for a long time but then the last few years we've been using gpus so pretty much if you are doing you know let's say machine learning and deep learning especially you use GPUs know beyond gaming cards and the other things beyond GPUs are actually not evolving I mean for example esterday google release the ship the tensile ship which is essentially a chip essentially for doing tensorflow so you're going to see a lot of those things I'm just going to give you like a like an overview of you know all these different things that are happening and you know what are some of the challenges in accessing each of them and you know what's the best way to access them so I know as you said you know I run a company called bit fusion i know we we have offices here and we also have offices in austin tell you what we do is as these new kinds of architectures actually come in a data center like for example your cell phone basically has you know has like so many different architectures it's got a GPU it's got a speech processor its got like a Wi-Fi chip everything the same thing is happening in data center you have like you know a bunch of different pizza boxes with different types of you know components as this happens the biggest challenge is how do you access them so actually what we do is we basically build software that allows you to combine all these different things across multiple servers and create one mega server a virtual server so in some ways we kind of do what you know people who have been doing in software-defined storage that you can pull storage we do the same type of compute but pool different kinds of computer like you know CPUs GPUs you know FPGAs DSP ease you know new architectures come in you know we basically have a platform that allows you to pull them in an interesting way so do you know that's what we do but in order to get into the fundamental problem if you see the fundamental problem in computing the grass kind of looks like this so on one side your software it's actually you know very abstract know as you know as software developers actually built just software there's one thing with algorithms and applications don't want to worry about hardware it's just becoming like use case driven and then on the other side this hardware which is becoming complex and like really fast right so cpus which are basically hitting the limits of Moore's Law GPUs you know they're becoming better and better at gtc you know the envidia essentially announced that they basically putting a in our new chip with like 150 billion transistors you know there's new chips that are coming up like you know like Google stencil float trip there is no there are FPGAs their DSPs there's Nirvana which is actually a company in San Diego which is basically building a chip which basically runs a deep learning framework of neon so there's a big gap you know so the fundamental problem is no as you know software becomes abstract and hardware is becoming complex as a huge gap and at the end of the day you know why do actually you know why the software becoming why is it a problem it's a problem because you know application delivering performance is the only thing that matters to applications so to put put things in context you know the software world is increasingly abstract right as you can see no new github repositories you know there's so many different programming languages are coming up and if you look at like let's say the job trends now if you look at you know machine learning towards a job trends are like growing up you know in get up on the other hand transistor scaling is ending essentially more Intel basically was saying hey you know 18 months is what Moore's Law is I'm going to basically build new chips but now it's 2.5 years and it might actually become even worse because as the process becomes smaller and smaller it's challenging and you can kind of see you know that's happening over the last 23 years and it's going to become even more prevalent and you know as I said it's 2.5 years more than two years you know these days so this is kind of how the background looks like right so doing the era of frequency you had Intel IBM and AMD building ships faster and faster chips no frequency frequency frequency then they said ok multi-core it's a cute build more cores and actually do things more in parallel now that's when MPP and everything came again the same big guys Intel AMD intel then the error mini corps came you know they say okay multi-core is interesting but let's actually put more course and actually make it more distributed can the same big guys then you start building special course right which is where i mean i mean obviously you know in the mobile you know world you already have actually know guys like Qualcomm doing it but then you even had you know other people like for example you have network processors like the tiny rows of the world who basically said I'm going to put network processors for doing you know you know our network ships again the same guys and a few other guys like Qualcomm and then arm came in and said hey and what okay now I can actually do things why do you need just Intel and you know IBM why do you just need power on x86 hey you can actually have this arm thing I can lie everybody can license this arm that's what every cell phone has me know you know most of the cell phones none of the cell phones have Intel today it's all armed and Qualcomm resistance your consumer apart they basically building you know both you know chips for in all compute as well as for know DSPs as well as for MOBA for telecommunications then easy cheap you know they basically building like you know network processors and so on now if you look at the world today that's how the world looks like today you know yeah in media MF course in media is no man I say now the last 10-15 years right so Nvidia there's you know Google which is a new player you know as software companies are basically starting to build ships because they say hey look I have this problem so I'm going to basically build ships to solve my own problem there's media tag and even we might see in a cloud service providers I can you're building your own ships there is no Cylons alterra there are no there are companies all over the world in all companies in Asia no company in Europe actually building you know building ships to do what they want so if I basically have a problem I'm going to just build the hardware to do this thing so soon there's one particular problem that exists the first graph that I showed you where software is becoming abstract and hardware is becoming complex the gap that is actually not full it still exists because now you have all these different things that Google is going to build a chip to just do their thing you know amazon is going to build a chip to do only third thing so what happened to all these developers like in all developers in a building like no startups you know mid-market companies who cannot go and build a chip how do i access all of these things so that is one of the first you know fundamental problem that still exists now there are some I mean like I can have no led to this thing you know a little bit right developers are the guys like you make everything happen right we know again there is no specialized hardware you know lower burnley pricing know as the libraries api's and everything how are you going to access it so the reminder of the talk is primarily going about is going to be about in details about the hardware that's out there and how to develop for them you know so there's all these Hardware how do you develop for them so the reason this is important is the current state of developer experience for anything other than a cpu is challenging right for example if it's a cpu i can run it on my laptop but if it's anything beyond a cpu it's not easy now if it's a even for a GPU now installing to installing drivers is sometimes I know is a pain in the butt you know if you start looking at DSPs on fpgas it's again even more pain in the butt I was actually not talking to a guy who actually boosts the compilers for the Qualcomm and a chip the hexagon shape and starting to a bunch of people they were telling me that hey I asked him have you guys ever developed as a user on your chip does it know so you don't even own that ship no they basically right compilers for it at work but they don't own the chip because it's really complicated developed right so mean this is very known right GPUs you know me we all know this right if anybody has done developing with GPUs you know I mean stack workflow you know all the problems that people put on stack overflow so the first set of you know simple you know things are integrated GPUs you know essentially that's something that sits on everybody's laptop you know Intel laptops every laptop basic Hasse's and essentially you know the you know it's essentially a simdi you know architecture that they have which is basically single instruction you know multiple data kind of architecture which is very shared resource and primarily the workloads that are tied targeted are no words are not really compute intensive but our latency sensitive and cost sensitive for example you know media you know encoding media transcoding and things like that and the programming models that are there is basically opencl opencl you know if you guys actually opencl is basically another you know in a parallel programming paradigm like cuda and support direct compute suppose c++ as well because inta like a bolted they want to make sure things actually work with the c++ level then you have SP ir and I has a HS al all those things the ecosystem maturity for this is very high because you know everybody has this on the laptop now anybody can develop it so that's the state of integrated GPU and the way now the way actually know things actually workers is essentially it's on the same chip so if you basically have for course on our CPU the fourth code is a strictly integrated you know GPU so this this is basically the most common you know thing that's out there the next one is discrete GPUs right which is which is what nvidia cards are the AMD cards that are out there essentially they are essentially pci cards the device when you put in your PC I you know they can are either come in and all essentially like you know gen 3 by 16 and you know things I keep going to become even faster they are not a first class citizen there's still a second class citizen because it still has to go through the CPU to do any kind of work you need a cpu it doesn't run an operating system or anything again same simdi but you can put larger workloads because it's much bigger and it's not small and it's it's much more you know it's not it's not latency sensitive it's very true put sensitive because it has this thing that has to go through the CPU to do certain thing it has to go through pci and pci is basically no 60 gigabytes per second right and you know you have the programming models opencl of course now AMD supports opencl and you know and even intel has actually discrete cards like VCS which are visual computing accelerators and of course the most popular one is invidious and envidia basically built this thing called CUDA which is essentially their programming paradigm and they built it in like you know like like about seven to ten years ago and that's now really becoming popular and one of the reasons I think that became really popular is because people start building a lot of libraries and they said okay there's all these things out there just like the fabric if you want to build a website I essentially our libraries I can build stuff for my website again the maturity is pretty you know pretty high for this and very good for matrix computations then you are mixed this is essentially you know what int'l calls it xeon phi so it's actually a class of processor intel bills the right notes the the version is called knights landing which is not public it's kind of public in universities and labs but it's not public on public to everybody to use the most public version is nice corner which is you know if you basically Google Xeon Phi you can message you buy something for 200 bucks so again a pci card extremely power-hungry large size workloads high throughput but primarily for hpc people know people doing computational fluid dynamics you know all these guys are using it and intel is now actually not targeting this even for machine learning and deep learning as well so if you guys basically know and i'm assuming you know that this presentation will probably be up online if you actually go to the link you can actually find some example and how axing they were using it for machine learning and such again supports opencl support OpenMP doesn't support cooler suppose NPI and you know and it also supports C++ you know in some ways so again ecosystem maturity is actually pretty high in hpc world but it's not very high in the deep learning or the machine learning world yet and I think this might change as Intel releases the Knights landing which is actually going to be the next version that makes its kind of public but it's not widespread when that actually happens I think this is going to be no become very widespread and once it becomes integrated in the in the first into the first class citizen as it become a socket because Knights landing is essentially going to be a socket you don't need to be a pci card anymore it became very very you know very popular it's going to be interesting to watch in the next few years then the interesting boy comes your FPGAs I mean FPGAs are nothing but field programmable gate arrays they essentially it's kind of like putting it's like it's like a big layer it's a big switch with a bunch of you know programming logic and I can wire it to do however I want for one moment it becomes a speech processing chip another moment it can be a no image processing chip another moment it can be I know a visual processing chip so again it's essentially a bunch of luts logical units with with a very tight fabric that it's connected with so essentially it's kind of like a big multi-bit multiplexer in the front and you have a bunch of logic gates you know in between our logic units in the you know before that again it's again uses coprocessors again it's a pci card the problem here is right now its supports opencl just recently in the last 23 years but the problem is you know so far it's been supporting only VHDL and verilog which no software guy understands so what they've been doing is people have been building IPS and libraries on top of it and you know people have been actually sharing this with everybody then you have you know things like you know automata which is again the next you know the next guy in this mix this is actually a very interesting chip this is a chip actually that's actually done by micron what they actually their vision for doing this is they won't hesitantly say what is the best place to do compute you want to get closer to the data so what they actually want to do is they won't basically put compute next to the memory micron is known for making dims law and memory dimms so what they basically are trying to do is then you build this and put it as a part of the dim so now essentially you actually load things up in memory and you can essentially do computation and they are targeting very special purpose things and these are actually no sub trees that came out of like not things like fpgas so us you can actually it's an NFA you know if you guys know it's an external commenter if you guys actually understand you guys and I'm like you know computer signs no Dana script classes right an NFA on a chip essentially it's a programmable fabric you know it's again you know it's like multiple inputs and single data stream and it's all pattern matching essentially it's completely pattern matching they're very good for things like by informatics very good for regular expressions so for example if you're doing like you know a snot and if you're doing let's say you know you know dns detection and things like that this is actually very good and since its regular expression it can actually be applied for machine learning as well because you know if you look at machine learning not the you know the the deep learning set of things but the old-school machine learning where you're just trying to the pattern matching this could actually be a very very interesting thing and this could actually work in a lot of low-power scenarios again the ecosystem maturity is it is infancy right now you can actually go to you know automata micron automata calm and they actually have an sdk and they have a whole simulator and everything can actually play with it in fact this is actually available for axis at university of virginia so you can probably request access and you know you can probably not play with it if you guys know if you guys want to and if you guys anybody wants to play with it let me know you know we we work with these guys and I can probably not helped get you guys access on this as well again it's as I said right it's very very low they basically building API is there in the infancy where envidia was you know probably 10 15 years ago then comes application-specific processors so just like how automata said I'm going to do this for a particular stream application process like hey I'm going to I nude molecule accumulations I'm going to build a chip for that I need deep neural networks I'm going to build a chip for that which is what you know Google built right I need a thing to the speech processing I know like 15 other companies that are basically building a chip specifically do do feast processing right because the speech processing it's basically a bunch of ordnance and you really the network is like a a bunch of small networks but a lot of them disperse so the ecosystem maturity is hardly any right only big companies with big budgets essentially actually do this thing then you know Google's yesterday's release right essentially a tensorflow chip so which is section you'll be power for deep learning they already been applying it for you know the rank brain to apply improve the search relevancy and also in the street view to essentially improve the arena nap accuracy and in fact the the Alpha go game was actually powered by TP use so this is actually the data center so if you actually have been using Google's actually API for the wha speech recognition you'll actually be using this chip you just didn't know it it's become swaps track and then you know the more interesting thing is quantum computers architecture it could be varied right targeted workloads it's basically salts an optimization problem which is which i think is going to probably happen in the next ten years and there are some really interesting startups doing it there's a rickety computing which is actually doing quantum computers of course you know d-wave is has been doing quantum computers IBM just opened up the quantum computer access to everybody so you can probably go and get access to it and programming model nobody knows what the programming model is ecosystem maturity nobody has a clue but this actually it will be very popular in the next few we know the next few years still you know the accessing this is still a problem right so so the point i was actually i actually make here is there is no so as these different types of architectures actually are actually not getting out in this space once you actually make it accessible and make it available to normal people when I say normals like developers magic could happen so our company essentially what we actually do is we tried actually solve this problem so if you basically know essentially how do you actually access like fpg is how do you access DSPs how would you actually GPUs all from your laptop so you don't even have to actually go to the club so that says you what we actually do hope this is useful you know you can follow me on Twitter no sugu rama or you can follow in a bit fusion-io and are happy to answer any questions