sfspark.org: Alexy Khrabrov interviews Alexander Ulanov about Deep Learning on Spark
Recording: sfspark.org: Alexy Khrabrov interviews Alexander Ulanov about Deep Learning on Spark
hello everybody i'm alexi car broke the organizer office of spark and he were on location at night raw where I we have a very interesting meet up tonight about deep learning on spark and this is a very highly anticipated topic we had about 230 signups at the peak and folks are really interested in in this in this area and today we also have an example rollin off who is like a senior researcher at Hill apocalypse hi nice meeting you yeah it was great to have you here so tell us a little bit about Hill bucket labs why do guys you spark and how did you come to implement your own version of the blue mmm yeah so in Phillips I'm in the analytics labs so that's a lab that actually does the research and software in analytics and our current goal is to come up with some new analytics methods or some techniques that will make our you hard or shine okay so in particular currently HP is developing a shared memory machine and she will have a lot of memory okay and we are also planning to develop software that will take full advantage of it okay because right now we don't have a lot of such software and the tonic might be one of the applications for this shared memory machine and since we can use some of the existing platforms for it so we don't need to come up with some media platform for deep learning or something so we can use park ok because park provides other means like for data processing can for machine learning can for streaming and everything so that's that's one of the reasons why we develop the planning on spark so it's it's one of the analytic locations for their shared memory machine okay this is interesting right because you know we used to have in the 60s shared memory machines so you know the cluster is kind of recent development so so are you saying that you will have basically kind of tradition architecture if I have multiple CPUs working against the shared memory bank because this is how is different from a mainframe is the new generation of a mainframe uh yeah you can say so actually one of the main advantages of this actually one of the main differences from the mainframe is that this memory will be in non-volatile mm-hmm so it won't disappear when you switch off the the computer right okay and that means that you don't go to the data every time that you need to speed up like you I use the memory so it's one of the differences that another is that that architecture is quite different so I can't tell but all the details but it's just completely different because it just has much more memory than you can imagine on a super computer or made friends with just orders of magnitude more and it should replace not only the drm it should replace SSD and ATT also so that's that's one of the differences well that's really interesting so I remember so your 10 texts by the bay right so you're basically interested in text as well so is this kind of odd this different applications or is deploying work connected to your text mining interests I actually it's not connected I used to do a lot of research in text analytics for there isn't five or six years but then I switched more to like distributed machine learning mm-hmm but I still try to follow there isn't papers in this area and I kind of wait until someone will come up with good application of deep learning for text analytics I heard such a no I think meta mind has an interesting work right where they can do sentiment detection real well using the blurring so maybe you guys can can run some text analytics on your on your platform yeah that's true but unfortunately it's not as good as deploying for images or for speech recognition so so it's not kind of something extremely which is like a much better than previous approaches right it is better but it's not as as good as I want it to be okay so so deeper into spark it's I think it's in general and distribute machine learning is hard because normally so you have this iterative algorithms you have this gradient computing machines and basically you need to know what the gradient is and so how do you generally salt like do take advantage of this fact of this your new architecture or do you just hold in general keys what is your approach to to a diverse one yeah so since we are using spark and the main the name the main paralyzation spark is that for realism so that we can like compute the gradient on separate notes and then do a global update of the Moodle and there are three key things with that because like the stochastic gradient its screenshot algorithmic it's very hard to provide it then base gradient doesn't work well for large data sets but it's extremely profitable so we are trying to find like the sweet spot between those two approaches and the straight memory architecture just provides a very fast in to connect like in case of cluster computing when you have network bandwidth and the model is very weak then you can spend a lot of time on those some days I can since you may have like millions of iterations and those millions for dates it will just spend enormous time and the certain number architecture it will provide much faster update time and basically what we are interested in is in finding the trade-off between the computation and communication okay and definitely there is a an optimal number of nodes for a particular task after which if you add more nodes in the tub the tub does not execute faster and can even decrease in performance is there a specific number did you come up with we come up with the some heuristic which pretty close to the empirical results so it allows you to estimate how many nodes you might want to use and if you have a specific number of nodes how much slower you would be compared to the optimal case so you can have such estimates and they're pretty useful I think okay so so you preserve this work in Absalom and spark summit you're here right so I think it you know it looked like one of the most kind of interesting talks to me at least on the program so what what kind of feedback did you get in a look like is there a lot of folks using the chlorine in Europe what was the general pressure what kinda questions people ask yeah actually I was very surprised that there are so many data scientists and all of those folks they use tighten our ears so they were mainly coming from that perspective mm-hmm so that how can i easily use your zip planning implementation without a big effort of learning something new right because like that there are a lot of libraries for deployment and you can just use them and spark it tends to somehow also develop in the area of like data scientists when you have a plan to different machine learning methods and that's processing things so those people were asking about this simplicity so how how simple is it to try mm-hmm so that that was a my main surprise and that they are interested in fight in actually and we have quite an interface for the plane okay so so they they can use it so that's a first garba compression and the second one was that people were suggesting something like some interesting workarounds and approaches like they were suggesting to you to write a separate separate implementation for GPUs mm-hmm that will take full advantage of the GPU and also to explore like other ways of communication yeah that's as was the second second interesting outcome so in the book of patent this was recently interesting to me because I noticed right a lot of they designed especially the men using Python mr right and so and so you know spark is written scholar for a reason so as kind of feel skyler compete organizer I'm always interested in how do we bring Python and arc pull into scholar so what is your impression do do is it possible to teach a lot of data scientist to you scholar or is it kind of safer to stay you know with kind of some API is in Python and explain it to two of these folks but you know if they have to actually tinker like when can they contribute conveyed 4qn you know work with your scholar implementation if they use not only spasibo what do you think is a good strategy and my person was that they don't really want to switch to something here mm-hmm and they have very good tools with a lot of options and a lot of libraries mm-hmm like Skype I and other is redder and I think they want just to use the same API but at the same time to be able to use it on the scale mm-hmm like without any visible switch in the code mm-hmm just to scale it mmhmm yeah and i think it's it's also one of the directions that data breaks guys are following great so and so do you provide basically equivalent python api for everything else you doin scholar yeah for 4d planning that there is a python api as well sure to the scalability yes okay yeah we tried to keep it very simple so it wasn't hard to implement it although i didn't do it it was done by the other guy okay and so this is a joint who are great because we talked with chugging yawning so general data bricks and another folks yes yes as a gentle work and it's surprising there that we didn't know each other unless we meet in spark developers Miriam please uh-huh and it enables us to collaborate and it's really interesting because we're like we have different backgrounds were for different companies but we can come up with something and it's beneficial for everyone I think as compared to that everyone will develop its own system or its own machine learning methods it wouldn't be as effective as the collaboration of people right and we have a lot of the blurring libraries right now so there was this funny post on linkedin today from O'Reilly saying that answer flow matters despite the fact that every 47 days there is a new deep learning library right attends our focus on important things about it so so what do what is your take on all of this right basically so have you know theano and have caffeine all these things are kind of not really scalable and then you have tons of flow which turned out to be not as efficient and some other things right now we'll have some other libraries so how do you see your library in this kind of set of other libraries do you bench market do you plan to maintain it and make it the best library at least the best library for spark like how do you see a library yeah yeah yeah so first of all it's not a library so it's inside spark so if you download spark then you have it or you could already yeah oh wow let's also this is part of him a loop is about a distribution yeah it's a part of the distribution okay so you don't need to download something else if you if you have spark since 11.5 then then then you have you go all right this was yeah that's good i think and actually is a lot of people who are asking me like similar questions on this park Simas year event I kind of come up with with some good answers in that actually if you are a deep learning creature you would not use spark because it isn't the right platform for doing research mm-hmm at least for right now integrating itself yeah and dependent itself because you need to have a lot of features you need to have very efficient implementations for everything and like it is not for research okay but there are two basically these cases that are appropriate here so on one of them is that production so sometimes you can train your model somewhere else and then put it into spark and use it in production and it will be scaled like very matronly mm-hmm so this is one option the second option is dead probably if your data resides in Hadoop or in stark or in cluster and it's expensive to move it to some specific machine that does machine learning but does the clinic and then move it back then probably you will sacrifice some performance for the ease of use mm-hmm and I think it's like a general general a decimal advertisement for spark yes it is it it's easier to use them so that's the second one you stay within the same just yet so that's a that's not another option and also what's what's also interesting is that we kind of work on the ways how to scale and deep learning as it's not a very simple question and some other libraries and they just don't interest it in that mm-hmm particular question they are interested in like more features amazing GPU like one hundred percent real and yeah our our work is related to to try to scale it and it is it's different it's not after the no thing but it's very important it's funny to hear for me you know that somebody would not be interested in scaling right because if you hear somebody explaining why deploring became successful in our times you know and the Wrangell explains that you know if you want to build a rocket ship you need a huge engine and a huge back you know tank of fuel so the huge tank of fuel is data and huge engine is scalable engine combination right because you know you can build a deep architecture and so I don't know you know like why would you not want scale yes that's right i'm not talking to 440 libraries but like some of them they just not interest like a Theon or like a fad that there are implementations that use cafe for example and like try to put it on the cluster read and from yahoo right do some some yes and also from birthday and there is a first paper yeah and i think that's that's an interesting line of work yeah but it's kind of at the same time it's kind of coming up with some new and distributed platform and i think that probably it's better to use some common platform you're right maybe i'm not correct here what kind of try to use as much as as much library said as I candies and if I can just use something developing kid and probably is a good way here right so so you develop that library that that features in skul so and we talked about why in all of the signs you sponsor because i used to interact about if you want to motivate them right because you chose scholar to run it people are ready for spark and spark itself is written skull right so i think it's kind of it's reasonable effort to you know teach the two scientists that they can actually develop this stuff is Carla and I wonder like how would you explain to your colleagues would not use call this park yet they may be kind of trying spark with both of them right but they may be curious you know why is it in Scala what is this color thing how do you explain to them the advantages of scholar and why thing you know it may be a good idea for the design is to try its call that's that's a really hard question action yeah it's a really hard question because it's it's kind of sometimes it's really personal preference yes like you you like fight on for some reasons yes like you you like spell of some reasons I ok I can tell why I like yeah why do you look all I for me it seems more natural to write code in Scala because very functional and it enables me to focus on writing the actual algorithm as opposed to writing a boilerplate code for something like for some for some sucking kills mm-hmm and the code looks very condensed and at the same time it's very readable but you can say the same about an answer language actually no java it does not look very concise to me yeah that's the this room but at the same time a lot of people are used to it and like you can't appreciate them off using something else true yeah and I think it's perfectly okay to have liked to API is in fightin and Scala and then probably at some point in time people just the people who didn't have the ground in Python they will just start from using power but you scroll for your own implementation right and you found that it's reasonable right so I think that's kind of to me it's pretty good evidence yeah that's right that's like in a very there is a small difference here because I'm I'm actually as a developer so right and data scientist is not a developer all right so probably there is a difference okay and you probably don't need all these scholar features for doing data science probably yeah you might want like very useful libraries with a lot of parameters pipelines and like things you used to have but you might not you might not want to use all the features right so it's a subset which you can do not you don't have done right they can use kind of you know subset of scholars just roughly smell to python all right and but it kind of already will give you probably better performance and even type safety and things like that so that's kind of fun but that's how I think about it thing ever that designers can use collections right enough spark is basically is Martin dirty centers can be distributed signals conference is the ultimate skull collections yes and also its kind of adjust it yourself mark is just a deal on skull so mo which allows you to work with big data mmhmm yeah yeah so so it's a yes it's really an interesting question and I think if people are if people like use bitin we should provide them with the proper okay absolutely especially if your portal spark rather this part provides people at guy in bothell Javagal know are as well well excellent so thank you very much for sharing this we're looking forward to the evolution of your of your deploy implementation let's see how this develops you know it's kind of see how patent people use across call people use it let's have you back once you make some improvements to it and you're looking for example yeah thank you thank you very much happy here thanks