Devreal

Data Fellas Q&A with Alexy Khrabrov

Data Fellas Q&A with Alexy Khrabrov

Recording: Data Fellas Q&A with Alexy Khrabrov

hello everybody I'm Alexa crab Roe for the organizer biceps column and here are a location at night Ron we will have a great meet up tonight with Bill banners talking about Scylla dest 3.0 and Teddy fellows talking about a gel data science with skull and we have both data fellows here loss let them introduce themselves and tell us what they the fells are all about yeah so I'm Andy Petrella I've co-founded the data fellows feed Xavier initially was the creator of the spine notebook which is one of the most important notebook around when you want to deal with spark but especially with scala and i will let w present himself yes oh I'm se va todo co-founder of data that's with Andy I'm a physicist stuff don't know a lot of work in bioinformatics as well and so we are very found of you know distributed technologies and data science we want to solve problems in this field of data science so and you mentioned Sparkman book and that's people we've been using at night run children joint work on explaining some pieces of machine learning including your networks using smart notebook with retinol so can you tell you know though deals a little bit about what is partnering book why is it good for for this site yeah what is it good for data science is pretty easy i would say because i am a data scientist so i have a mathematical background and three years ago i started with park after one year of scala and and very soon i had to have this kind of tool to do my homework right and it's actually i had created a company at that time already and i had to create a product so i had to you know do a lot of machine learning is a spark and i was contrary fed up about either creating as BTW project or repel using the rebels so i decided very quickly to to create a spy notebook i started with something o mater then I switch to another one which is the current one is based on the officially initially on the disk Alan notebook that I made evolved towards what is now so what is good is because it includes a lot of things that data scientist needs like reactive plotting in order to you know to explore quickly and efficiently the data sets that they have it's not about having data dashboards is more about know getting the sense about the data rather than you know trying to expose final results using it so it has a lot of feature coming from different tools that we've used a data scientist like the list of variables that have been defined notebook and also you know easy configuration so you don't have to bother with connection with the clusters and thing like these so it has a lot of features that are both in the scalar data science plotting machine learning stuff and also connecting with clusters which is more kind of data engineer / develop things and emotional awareness is called Smart Notebook and so it's basically systematically optimized for usage with luis abarca so many inventions per cluster that's because you know can sparkle to be cluster and usually them it was all the work for kind of the most personal data side is probably not fair you know enough time to do that right so to help them with that it's kind of a big deal yeah exactly and each user can tune the way spark will be used because each notebook is actually behind the scene a riddle JVM that is it has its own spark context that means that the data scientist can chew how the process is behaving on the cluster by by tuning how many courses dedicated up for the job and also the memory ticketed for the job and so on so far so you can tune it the way you like it so you know for those folks who are not really familiar with sparking Scott like a sister right and probably most of the scientists tried no vocal you know I Python can you apply and kind of advantages of using spark you know what kind of problems are most suitable for spark and get a spark than just in a plain old ipod right actually so the spy notebook UI is mainly the Jupiter one however we understand everything has been represented to use on the JVM stack so you don't need to have any Python or whatever and stone the machine and so there were difference between both of them is that the UI has been extended a lot in order to provide extra features like you know this panel of variables but also chat window and then you know plotting libraries are very sticky to the spine notebook so that when you return the data frame for instance automatically it will be rendered as a as a table but also if you return a list of doubles for instance is gonna plot a a line chart automatically if you like so there is something like this that are very adapted to the way you will work with span notebook because we only work with Scala we can do a lot of things which is not available in the others so yeah spark obviously is a distributed system which basically can go over later exactly so there are a few features that are very interesting when we deal with this park and and the spiral notebook because there is something really painful that you use park is that you have to go regularly to yarn or message to check the logs for instance in order to see if you had some problems or whatever in the spinal book for instance we have implemented thanks to ikea and web sockets and reactive programming we have implemented a way to capture all the locks in the Indus by cluster and gather them into the browser console so you don't have to go to the cluster manager you just have to go to your browser and console and check all the logs in that so you will have all the possible excellent we may have and the classical ones are of course the receiver errors and and null pointers or errors well that's great I mean I was still playing with spark on one elbow breton it gives you all the links to all the pieces of sparkle I itself rancho this actually browser based environment is really good way to learn a lot about spark itself and see the liberation see what day it is cached right then we'll tasks are executing so i think this is a convenient so let's kind of zoom up a little bit and you know you talk about genocide data science of scholars part in kind of you so you are sitting of this company to do that right and so the first time I met can go to work and severe joint item and so I won't ask you guys a what what kind of data science challenges are you solve for advanced honors what kind of areas right and basically how do you see kind of Sparta notebook given them an advantage how it's kind of a great tool for them to use the soul the challenges ok I want to give it a few so few words before passing that Mike to to my body it's a view so the thing is that we we had discovered a lot of problems while using notebooks in general not only the spine notebook in general right and the the main purpose of the company that we have credit is not to predict eyes the spinal book per se but is to two hello people using notebooks to work efficiently towards the production environment right so because when you have a notebook you're happy because you have done your work but it's not finished because you have to integrate with other parts parties of your company and more importantly you have to provide information to the business because those guys are the one that you have to satisfy of course and when I say business is not only customers but also internal businesses right so this is what you want to do face and now I will leave exactly explaining what we do in these different sectors yeah so what you observed often in the in the project we've been working on these these years is that you know the business comes with a question it can be we work with insurance companies for example and there are there are many questions coming for you know for marketing which actually terms product we should recommend to our customers and at what time so these kind of questions are no addressable with you know the promise of a big data let's say and so they come and they talk to the data scientist and you know he knows how to solve the problem for from a mathematical perspective and so is working on that and he provides a solution and it will come maybe with a report to do to the business they do that and it will work but things in practice should go further than that because the business guy will ask the question several times so you have to productize this this work and in order to do that you have to enable something downstream of the work of the data scientists usually work with tools in may work with Python it may work with our how do you make it a product so that you know the business pushes on a button and you get the answer it asks for before but with new data because data is streaming along the way so you have to be you know really reactive and agile with the is the data flowing in and you cannot ask the data scientist to make the end product for the for the business so you need data engineers you need devops to develop downstream and this is where the problem arises because it's a lot of work to get you know the data scientists work efficiently with what happens downstream because those guys there are very efficient they know their their work but they have they need a way to communicate because it's hard for a data engineer to understand exactly what the data scientist did and so we're working on this kind of issues and the spark notebook is any neighbor in some way because we provide the data scientist with a tool that allows him to do is work but because we're working with scala with types and so on we're able to extract very valuable information on water it with you know types if you write the data into into at a billion database or if he the output is a machine learning model we can extract the information on these and provide these as a scheme at Avro for instance to to the web engineer and we can even you know generate the the basic services as projects that are able to query this data to make the service to to to make a prediction with the machine running model so the work is prepared and we have you know docker image that are being prepared as well and can be deployed on environment like amazes so even the DevOps work is prepared with this solution and so we have you know the communication between the different team players is made with the code that is generated so they know what they're talking about because they have the code that is telling them so it's easier for them to communicate understand that the needs of each other so that's really what what we are doing at the divers shared place decided engineers to look at in a graphical environment understand the data right and c0 know yeah exactly so I remember ND mentioning that you guys work with genomics right and I wonder if you can explain and kind of to generally decide as you know what makes College Park good tool chain for genomics and house part notebook and help and kind of genomic were closed yeah so besides insurance were targeting genomics we have a background so we have some knowledge into this field and we know that you know there's a huge pile of we could say legacy code because a lot of work has been done to analyze this data coming from these bio technologies that are really growing at the amazing pace and ideally as a it's a hard work for IT to to follow so distributed technologies are will be there to help you know process this data when you when you think of genomics actually the data is very simple it's you know having the state of the DNA a different position along the chromosome so it's very simple data okay yeah yeah yeah yeah it's a TCG's but it's just like zero and once that that's true and you know most of the works consist of making kind of descriptive statistics so it's just counting what happens a certain position in on a chromosome for certain population for example people that are susceptible to a certain disease what what is their feature in the generic sense so it's descriptive statistics and spark is really good at counting so it's it's the right tool to do this kind of work and because we're talking about a lot of data so a genome is a 3 3 billion bases when we look at the data set like the thousand genomes which is the complete genomes 4,000 individual we're talking about 30 million variable position on this on these chromosomes so it's quite a lot of data and so spark is really there to help and we've been doing some work on this data set using this part notebook on clusters and you can really interactively make those counting's make those classifications of genomes in in seconds or minutes so yeah it sounds like it's perfect for spark because it just you know has preset unlimited memory if you can hook up enough machines and then it's just a simulation right so I guess you know this set is not trying this before and that's really stellar example how can just a lot of all data set right on its citizen memories it's fast in the control is faster so all those units called the plan first of all so they felt like I really like the design you have these cute animals can you show a little bit like what they signify away like I feel like looking for new animal when we can describe yeah regaining anymore so okay so maybe it's due to the fact that we are Belgian and maybe we are screwed by the basis not actually so initially so we wanted to do a lot of my learning a lot of genomics and a lot of medical center and so for us and when we looked at the spectrum right now of toolkits say not only in genomics man general for data scientist we thought okay actually people enterprises companies tool kits are thinking that the data and gene the data scientist is a monkey right is a monkey because you just have to push buttons and automatically everything will be generated from him you know the models will be tuned everybody everything we go just fine right so they don't open hold the doors for the people to work with the real data behind and creating new models okay so we decide to yeah actually this is monkey because they I met the better sent his nose that is not a monkey right he knows that he can do better than that right but is sticking with the toolkit that they have been provided right so we said okay maybe there's a way to open those doors and then we can make them cool monkeys you know so and of course there is that there is a monkey which is the the logo of the spinal book because this one opens all the doors it yeah and then there is somebody which is very fancy you know it's very smart and very not to say is how to say it has a certain notoriety his dilemma you know neva as you know has these kind of things so this most probably representing savvier monkeys miss me because i have the eye of the cap and there is a rat the rat is the poor animal in there that is constantly you know the target for new medicines and thing like this so the guy is not part of the team because it would suffer quite a lot actually so the yeah there is a bit of history behind that due to the refraction and the the way we were understanding the current landscape oh great that's a great algae no I know better like you know it's a good to be a big monkey it was spot on right yeah that enables a monkey right so yeah I had something to add regarding genomics for instance there is something interesting when you play with atom which is the library that we use mentally when we deal with genomics is that those guys as understood they have the understood that types is very important and organizing the data is very important right so again so we don't have to think that the data is somewhere so we can process it right away thinking about how the data and structure and how the data is distributed is very important when we deal with data submitted this with a computing system so those guys has you know taken the specification of genomics data and they revamped it into another schema and store them in to park a witch then which then allows adam to be very efficient when it comes to computation because you can do it in Perl efficiently and and not only on the other side they were smart enough to create types on that but also on the other side of the spectrum there is a onion named genomic sorry and Global Alliance for genomics and health and those guys are like the w3c of genomics and they are defining types all sewn to the different data sets that we use in genomics and health and also a series of services that we have to implement in order to be compliant with them so that different entities into indo sector can communicate in a efficient manner in a comprehensible matter right so this is this is a smart way of thinking about data pipe and these days and look like that actually general mix sector has you know overstepped you know they jumped like two or three generation of technologies and they go for further than the other sector so this is very exciting ibadah that is why initially we wanted to go in the direction and we still go there but also you know genetics and genetics and the health is also very related to health care in general and insurance so we got a lot of traction this is why we have a lot of customers in this area no I love this kind of interest because on one hand right through there is Adam but on the other hand most of bio informatics percent is still in Perl and Python Oh surprise half of them were using parallel and I did a talk at fam last year so you know and i think it's like on the one hand technology percolate slowly so i think it will actually be great health and great service to spread the word right I'll try this thing so well kind of this is really kind of my last question or a kind of brainstorm right so we we are planning to do a preparing a hackathon global hackathon for Scola and data right because it kind of merges that you know all the people using scholar they kill a pump a lot of data through it and they use it you know forget the pipelines or maybe ice they use it will spark right they uses a different context but can it turns out that whatever reason skull is good for data so you know and we will actually have this hackathon both on the tool inside the pipe landscape that also will see we can put together some interesting puzzles in genomics and working with vampire and you guys right so I just want you to kind of to to think aloud like why skull is good for the right like a lot of people will start using data size with Python the top pilots on the rides are not bindles coming yet with the here'bouts part not many people should know what this point that spark is all written skull so it's kind of good to kind of spread the word and explain why Scala is especially good for data so i just want you guys to kind of think about that so I like I mentioned why and what would make an interesting problems people can try to play the data users call yeah yeah actually in my in my track on working with data I've seen many systems I've started probably with with some see then went to see prosperous then I went to genomic so pearl was the way to go and again see if you want to be fast sometimes and then some are a lot of our and then some Java which I never used for data analytics because it simply doesn't work and then when I went into skaara it was really kind of a perfect language for me to manipulate data going into tional style so really thinking about processing in another way and especially when you are going into into spark you know you're thinking about your data transformation and it goes into the distributed world in in a pretty much transparent way so but that is more specific to you know functional programming and and using it as an interface a parting for distributed computing now the specifics of Skala for for me is that you know writing code writing a function that will do some statistics on data so a bit of transformation selecting some features counting them in a single function it requires applying several transformation that are chained working on a notebook that is interactive with Scala that means types it means that I just implement one step executed and I see the type that is coming back if it's good then I write the next element in the chain so writing a complex function may take just 20 minutes and it's it's right when it's written while with other languages or other other environment it it's much more difficult because you right dolci bank then you write some tests and you have to test everything yeah maybe yeah that's for the specifics of Scylla of course is that you you you manage problem in a much clever way you differ the problems to the end and it doesn't enter into your function definition so it's it's really great for me it's is the way I work now and the way we work with data yeah businesses know this is a great summary right like in types felt you not know and you were smelly and you put together another thing when you put together things which I being each other's types they usually work rightly the workers way I mean maybe you need to add a kind of tweak the data a bit the processing but at least they will compile the war on and that's a huge huge knowledge there are two thing as well so regarding functions is very interesting to think about it I want to reiterate on that because when we deal with look at machines what we do is that we bring data in memory to the function and then we apply it right so this is how our local machine will work but now we have to think differently so that we have to send functions to the to the data in order to be applied right so this is why functional programming is were interesting because you can they are first-class citizen so you can think about the functions and then the system will send them over the data but the data a remote know and that means that helping being helped by the type system is very interesting because you do that on your local system and when you send them you're sure that it's going to come to to to cope with the overall structures that you have defined so far now when we deal with so when we have skala code interesting interesting thing is that the tool chain that you will create towards the data itself to the production environment can remain the same same language I mean because you did the data you consume it you will take it from the outside world you will put it in something like Kafka using a highly available and highly scalable conductors and scada which is easily to deploy easy to deploy and maintain the new position into into Kafka which is by the basic element a shin then this disc after can be confusing special come on Spock streaming and I is going to be pushed in in in Cassandra which is javabe JVM again and then you go through the machine learning that you can write in since gala thanks to spark in the new libraries that are coming out of the wide right now in scanner and then you put you can create the micro services in skala again highly available highly scalable we're very reactive you push it into production and then you can declare some new customers that you can consume it and then if you are smart enough that you declare web services per se they can consume it the way the like it might be in Python might envy not GS or whatever they can do the way they like it because this is not how a responsibility to make them hide available or highest caliber you know so using one single language in one single platform you can do the whole thing from one end to the other end and this is why this is why I pretty much like scatter in general but why not java because kayla for me is a bit easier to think of and to end to deploy new stuff by the time you finish writing java call the problem went away already right talks you don't in the solution anymore chris is doing yeah you know actually this is funny it means me thing right because you're talking about the hardest problems in india to talk about machine learning right genomics and then they talk about microservices I thought maybe it's like it's it's a it's a Miss name maybe we need to call the macro services right because they gonna I guess so big problems so it's a good microphone yes I they compose it composable but they can actually capture the most part to find the service right which can be seen there for working of genetic data right answer so so it doesn't mean the smaller it's just like they are composable right yeah they'd have to be highly mobile and using a resilient system is very important so that so once we have kind of an easy way to have this microservices spin up and combine probably you can compose them in scale with the same way the most functions and maybe like you can use part number two visually yeah and then something interesting to note is that if you deploy services that are disconnecting from the cluster you know the company cluster that means that you are you can you know very scale very at very different levels because if you have a new bunch of user that coming into the platform you don't have to scale your spark on computing frame Ram cluster which is very expensive you just have to scale the services which are targeting another system behind that can scale already so I mean this is also a matter of cost you know scaling the services which are generally do small machines versus scanning the computing framework I mean there is no this is not no brainer I mean you should always scale the micro services instead of the computing clusters so this is the difference for me between tools that are generating dashboard within the same environment without passing by the internment step which is the service dinner no this sounds great actually I think it's kind of very nice if it's the velocity of scholars of composable exact instructions link impossible functions all right and we can do the same great well thank you very much that was really good update right on day two fellows so we're looking forward to great things to come out of this error looking for a dog yeah thank you thank you for saving us