SF Scala: Q&A with Andy Petrella by Alexy Khrabrov
Recording: SF Scala: Q&A with Andy Petrella by Alexy Khrabrov
hello everybody this is SF schola and I'm the organizer Alexi kro we're here on location at Nitro for a great meet up fall in schola days uh at scholar days everybody from around the world comes to San Francisco this year and we have with us here Andy pet the author of spark notebook early adopter of play and other great skull goodies and he just did an amazing talk about spark notebook with a lot of Graphics graph x uh GitHub mining uh EUR exchanges and all of this is available functional TV alongside this Q&A so we're going to first welcome Andy it's great to have you I know you're jet lagged your laptop shows 5 a.m. so we really appreciate that after whole day of spark train training uh you came and give us a great talk and you still have a few minutes to answer some questions so uh um when play became popular I've seen some block posts from you you were trying it out I think you did it on Hoku as well um and now you're are an adopter of of spark and you really do amazing things you bring up some doy instances and you use kaf and AKA and like you really know a lot of aspects of the schol E system how did you become interested in scholar and how did you become so it in all aspects of the system yeah actually uh 3 years ago I started my company and at that time it was it was a start up at that time so we wanted to create a a um a tool for data mining um mainly in the building environment so so for construction and so on so in in order to organize the the work in the construction so at that time I had to pick up some technologies uh to create these tools and so when I did the um I to say the the overview of what was a available at that time I knew already some Scala um I saw display Frameworks uh slowly raising mhm it was not there yet it was not even in alha or in beta it's just something there that has had to compile I picked it up and I believ directly in it because it was something quite amazing it it was just there and it were it ran fine mhm so uh also I picked up some other things like new 4J other stuff like this did you know an next city guys did you find it socially or did you just find it by yourself by myself that time because yeah there there were not that many documentation on um on Scala thing in general at that time right three or four years ago I had to read the epfl papers in order to get some some insight into the the monad and you know I remember having read the the reader monad out of apfl paper that time because I didn't knew that much hcal uh and then you know you have to move to hcal at some point in time when you do Scala apparently the hcal that's right the hcal lator yeah exactly uh anyway so uh yeah I picked a lot of different uh Technologies at that time then the the startup went down and I keep up um uh looking for new technologies and then I discovered haachi spark quite soon around maybe two two years and a half already now mhm and I created also some blogs about it uh just just to to show what we can do with that and yeah that's mainly it now I'm going to create another startup and I don't know which new platform I'm will use that's great so yeah I mean that's like the stuff you you can do is really impressive shows like the whole whole different aspect so uh so why spark why Big Data why Kafka why all this e system yeah so why why is Big Data or mainly why distribute Computing because big data is like a bad word to me true um so why distribute Computing but in in a case in some sense actually I wanted to do some data analyzis right so um because I'm a background in mathematics M and um the project I wanted to have had to be on mathematics mhm and since uh we henter the world where the data is not cannot fit into a single machine I had to choose a model that has to run onto this data and I it of course it was the D Computing and um so why spark because I didn't want it to do Hadoop MH so I looked for something else than Hadoop uh scalding was not enough and then you know the the model that spark um proposed you with this memory caching and this uh lazy computations and so on was very attracting and then I picked it up mhm yeah so uh so do you think that spark is the future of distributed computing is this where you know we're going to all uh end up to me yeah actually the thing is that it you know new technologies has New Kids on the blog gener H has to bring something disruptive in order to be the new one right and Spark comes with this um with this three main features uh first is the reactive um the the inter interactive programming using the Ripple or tools like notebooks the second one is the cache integration and the third is the API of course so with these three items they I actually they they are farer than than Hado and can who cannot they get that that just cannot keep up and and reach reach them back so that's why spark is very interesting and that's why also why I I think that spark will stay there mhm other thing that they integrate with they integrate quite well with different languages is also interesting of course um okay so the I think spark will be there at least for 5 to 10 years has being the the very cool thing in the blog and the the the way to go before it anwers the mainstream area like had is right now so IBM can sell it um so um the maybe the next thing uh what will be the next thing or in my opinion what can be the next thing it's it's most probably be the distribute databases mhm um you know when you look at Cassandra or you look at rayak I mean these guys are there to manage data from the ground they are just there to manage data right and this is basically what the distributed processing um uh framework is doing manipulate data but is doing it a middle level uh when the data has been stored somewhere although right and Cassandra can store the data in a very efficient manner replicate and so and so on and also if you look at Rak it can also process the data in the Lo the nodes so actually you are not that far from something like spark mhm so yeah that's something that might be uh come with a new set of features that maybe spark hasn't although spark pron is also a good fit I think it's a good point so I've seen um a few new things come so there is a something called FB which is a new object uh data store from folks Twitter right so I think uh I've just heard about the recently it looks like uh people are doing something similar because spark has an object model so the difference with a file system that you you have the storage but they also have object model So you you're talking in terms of rdds right so if you if basically if you have persistent rdds which just start the database and you don't need to serialize De serialize them right then you will have something really convenient so hopefully we're get we're getting and I think I think that uh uh datab bricks is going to read uh a park files loaded into memory directly with the data frames right and that's probably will simplify this processing a lot so uh what um um what pain points do you experience right now with spark and c and scull ecos system what is most difficult for you if there is anything which kind of comes to mind um there is nothing really difficult with that because it's I mean when when I wor with Hadoop and storm it was it that there were some pain points for sure and then spark clears cleared them almost all of them uh the thing which is a bit um boring sometime to solve is this calization problem that happens time to time yes and actually it's not really the fault of spark it's it's it's something it's a great feature that Sparks brings to the table and it has Throwbacks because actually at some point you can include some stuff that are coming from somewhere and you didn't knew it and then you have to to hunt um until you find it so you have to enable some weird parameters in jvm that's that's the that's um the thing that dis like Lum the must I mean it's not really a spark problem it's a more General problem like a system yeah so yeah that's that's the the main point the other points I don't I don't know maybe I don't know it's it works quite good I I have no real issue with that cool so uh this is you know a rare occurrence when a lot of folks from Europe for coming to visit us and I know you mostly work with European customers so I'm wondering uh what is the state there you know because we live in a bubble here we think everything starts here and ends here how does European industry differ how what what do you what patterns do you see working with European customers how big is scull adoption how big is spark adoption if you kind of characterizes broadly how you see it um maybe I can talk for the Central Europe Europe or several projects in France Italy and London and Belgium of course M um I I think that Scala has a a single second um uh how to say how to say a second slope you know it's it's raises again the first uh the first time it was um thanks to play framework yes you know Scala was that that much wasn't that much known then play framework um arrived and then a lot of people came into it and um was able to deliver some web applications in a very short time because of this lifetime death lifetime MH now it's starts um to be a bit flat so no more uh people are are are are interested in or new people are coming in because you know now there is this spring thing the spring boo that can do something like this so ja people are uh staying away from from play but you know spark comes with this uh with this interesting um uh API in a word well where distribute Computing is very Buzzy so a lot of old people are coming into the game so at that time it was at the time of Play It Was the web developer and now we have the data scientist are interested in Scala or forc to do Scala yes I I'm don't know if or both or both yes so yeah the Scala is is Raising so in Europe we can see it that now the Meetup are are a bit more um or a bit less empty mhm so before it was like we were like 5 to 10 now we can go up to 30 at some point in time depending on the on the on the subject you know the last time when we I was given the uh the spark um to work with um uh we were like 50 or something it was amazing in Belgium you know 50 in bing it's like the half of the population right yes that's right um so um so it was quite interesting so the big data is 50 are French speaking and the other 50 are uh I would say maybe 80% was Dutch P were Dutch speaking and the other uh where uh um German speaking no no the rest was were simply German and and French in order um yeah uh so yeah this distributed computing or Big Data engine is is attracting a lot of people uh but the market has some hiccup to start I must say I I Leave myself in the bubble because we are not that many in Europe knowing very well spark and so on or this with Computing so I I'm I'm I'm qu different places to to to to explain uh how to do distributed computing or distrib but I don't know how the ERS are doing I think it's it's evolving in a good way that's great to know that's great to know and you know in so in August we have a big data Scala conference with two major themes end to end data pipelines and data science on the jvm and actually it brings me kind of to the second point because you're a mathematician you're you know anti patella spell with a natural number set right which you know NP which you know as a mathematician you know related people you know can appreciate uh so uh I think there is this huge task before the scholar Community to bring data scientists onto the scholar platform right because you know what we can do with spark is definitely Beyond reach of python or R and uh uh we can be more efficient right and we can be as concise and so forth uh the question is we know that we know we can have beautiful dsls can have efficiency uh but most of the data scientists are uh brought up on Python and R they like their tool so I wonder what's your taking is how can we enable them how can we help them to move to Scala and what should we do as Scola uh Community as spark Community to help you know General data scientists to to come to Scola yeah um about that um what misses Scala for instance what what they are missing the most is is is visualization MH right if you are proficient with r or python or even Julia Matlab and things like this you are you have all the the the plot libraries at your ability in order to have to to make very easily uh representation of your data right Scala is still hard to have it okay um if the spr notebook for instance you you have some but still you don't have the the fancy box plotted or or box block plot or the the funnel or all these things that you can have quite easily with DG plot to but quite easily between codes um using BG PL 2 or other um python uh libraries like maybe mbook lip or things like this so SC when they enter the Scala word they say okay man I just I can't print the console really it's not really what I want to do right um um then the the other thing that we might have to do and actually with started in belum is to to pick these guys at the University directly mhm because University they're already doing a lot of python everything related to data science nowadays is python correct so that's why we with my associate we we took some students and we started learning them Scara and Spark already at at the very at the very start and they start learning them because then they are not BS by by I don't know just a python R whatever mhm Alo is still interest still very interesting to know our Julia or or python beside because um there are so many models already implemented in there so at least to to build your expertise in in modeling and machine learning stuffs you have to go there there you have to go there otherwise you won't have uh all the the different models you can imagine Implement to uh Escala right and the other thing regarding that maybe we can show them the programming prog probability sorry probabilistic programming book mhm I'm going say that this book is very good one and shows how you could use Scala and its type system and the the monad even even though it doesn't say about the Monet um how you can use it to constrict a very um elaborated models uh using probabilities mhm and actually this is something that you cannot do with all the others uh languages because they don't have these kind of types and and you know function with strongly typed and this monad structures so it's a very good one it's a very good book that I would recommend to all the data scientist uh nowadays who are the authors H who are the authors of this book U I can't remember um it's it's an U it's a manic book uhhuh um I can't remember exactly it's Pro um uh probalistic programming it's F nemon I can't remember okay we'll we'll figure it out and we'll add a link to this uh uh episode so you guys can can the library is Figaro that's that's how I know okay okay that's great to know that's great to know that's you know that's I think I I agree with you we we need to kind of fill in the gaps right and we can certainly uh learn from R and python but you know there's a lot of uh skull developers so we can just you know identify what are the problems and kind of uh implement the necessary things exactly yeah uh agreed great uh well this this was uh very interesting so thank you very much for you know Vis us and we'll definitely follow your progress and hope to see you here in know our uh at some of our future events and uh thank you very much thank you thank you for