scala.bythebay.io: Moon, Complete big data pipeline with Apache Zeppelin
Recording: scala.bythebay.io: Moon, Complete big data pipeline with Apache Zeppelin
today I'm going to talk about a little bit about participial in the house and how it fits to your data pipeline and please interrupt me and any time if you have any questions I am my name is moon I'm a creator of our party supply and i'm a co-founder onf lives who donated our projects to the party foundation and before I start I want to ask questions so how many people here are heard about up exactly ok how many people uses our past happening ok thank you so I think it I can go through a little bit of a history of the project really quickly so Zeppelin it wasn't on pressors from the beginning so in late year 2012 in my company we have our work we were working on a commercial product called peloton on top of a spark and sharks that was it I think it really all related with sparking shark and it was a protocol solution that protocol solution for the dynamics that has a dashboard and user management and data import an interactive analysis and one over its menu well it's mr. interactive analysis was the most popular of peach among our customers so a year later we decided to open source that feature that became as a clean project and it wasn't look like now from the beginning it was much more ugly in the beginning but we went through multiple iterations and year late another year later we decided to sorry yeah we decided to our donate a peplum project to the Apache Software Foundation so and over year 2014 it became a Apache Stephanie went to went into the Apache Incubator and it's been almost two years since the incubation so the Flynn made became a top-level project in this year and in the beginning they put in had ten contributors or and or maybe less than 100 stars on it we talked depository but now we have more than 160 contributors from worldwide and we have more than 2000 stars on our github depository and please don't forget star after this talk and sublimated 60 religious and I am sure that Leninism are one of the most popular projecting or Apache Software Foundation and this is I think it that some some idea behind exactly what I'm trying to do so I would say this is a general lifecycle over your data at your work so data is being collected and then as it processed or transformed and then you will analyze your data and the results become a report or result become our data products that's I think a basic life cycle of your data and this life cycle you will need to use many different technologies or frameworks libraries or languages for each steps so you probably will combine at least four or five or six different technologies are to build this pipeline and also different type of users are involved in this pipeline so not only are there are engineers but data scientist or even business users are involved in this pipeline so what the pudding trying to do is it Zeppelin wanted to provide a unified environment for analytics for this all these different type of people to use all these different technology in a single place and definitely the planet looks like this and I think it many will you here use discipline so already familiar so it's happening is a interactive web-based node tool that you can use multiple different technology at the same time with a nice visualization and you can share your results directly to the other people so I think it although many people here know that plain but I'm sure here I saw a lot of a lot of people still don't know exactly so I wanted to share I wanted to do that some quick demo how that plane looks like and how it works so this is it can you can you see out of our text this is how the plane looks like and the planes one of the major menu it's a primary menu notebook and that's where you work and the plane have a configuration menu and the one over important menu in the configuration is an interpreter so interpreter is how the plane integrate with your integrate to your back-end systems for example Zeppelin have integration with your party spar have integration with the JDBC the innocents have integration with a lot of different tools or languages so here this menu is where you set up your backend integration for example let's see some more building integration so they are the list of available into integration we call it interpreter in Japanese and I can set up multiple configuration in this menu and then I come to not hook and I can select what kind of interpreter I want to use in this module so once I have this interpreter set up then I can use multiple different back-end in the same node tool so let's see this one let's run this one so this one is a share interpreter and I'm downloading some data the data is also cake data so where and and how strong earthquake has it happened and then I can list it the file I have downloaded with this again sharing computer and oh by the way this Audrey active selects the interpreter so person sign spark select spark interpreter % as they just like to share interpreter and I can load downloaded file using spot it takes some time yeah and then I can print the contents what it what how the file looks like first 100 lines and it's got a lot of comments in the beginning but it looks like some kind of CSV format so I I wrote a case class here and here some spa API to parse the data and register the data is a temp table and if I run this one then spark discotheque generate creates of the refrain in spark and register it as a table and then the plane hi boss sparks current operators so I can all directly query it okay and I can make another another type of query okay once you have once you made us some query then you will see built-in visualizations and some buttons that you can make your output data combust your type of data to nice visualizations and I'm sorry but oh okay can I can I stop the screen recording because I think my my laptop is struggling with i-i-i just stopped a screen-recording so yeah now it's I think it's fine so yeah you can leverage ability in visualizations here and there are some nice keyboard functions that you can play around your data and once you have done then you can change the not talk to the report format and it becomes nice-looking reports you can directly share to your co-workers or business people so that's I think it basic primary uses you're exactly so people who can handle the data work in the new tool to collect the data and transform the data and people who can analyze the data can work in the same new tool to analyze data and visualize the data and then it can be reports that business people can consume directly so you don't need any export import result into the database and input from business bi tool you can everyone can work in the same tool the same time and I can go to go through it some of other examples so one on one of a good picture of Zeppelin is a plane supports multiple interpreter at the same time that means you can work with the Scala SPARC API but at the same time you can do a Python and in the same network and exchange the data between Scala and price so this is a simple example that I put one string here this is it from Scala and I read that string from Tyson and Tyson print a string object from Scala and of course we can put our data from from the Python and read data in Python and read from Scala and this is it as one example in the Python site that we have loads this all-skate data in are from the scholar site and this code is a Python code that uses spark Emily and do a clustering algorithm on the tail on the terror attacks color it so let the mirror on yeah this yeah so there are three clusters created and it's been nicely visualized with the scatter chart so we can see how we can mix different languages in the same tool and one of the convenient feature in the plane is now this is something we call a dynamic form so any language integration can create this dynamically created reformed or not sure why this visualization is broken yes little bit yeah so probably if you are engineer with the aid of scientists you you can definitely understand this simple code and if you wanted to change the number of cluster on your k-means clustering you can simply change the number and run this part on this code again but let's say you're a business user and if your business user if your own business user then you will have no idea what this code does and it's a pretty scary call for but if you have a distal form that you can change the number input parameters then even though you are not understanding all this code you can simply change it a number here and run this machine learning algorithms on your own so it gives some ability to interact with non-technical people in the notebook very easily yeah I'm not sure why this visualization is you keep broken but let's let's continue and you you can see there are couple of ability visualizations here but that doesn't mean salinity limited to those type of visualizations if you have your own library visualization library like map will live in Python for example you can use inside of exactly yeah I'm sorry let me reload all the neutral it doesn't usually it's not happening this problem is not happening what always happening during the talk okay so let's see how the plane can use other visualization library like math live in Python so this is the same scanner chart but in math live in Python not built in visualization so you can see you are first lose our data in Scala and truth made of some transformations or process processing in Scala side and then run some machine learning in Python and visualize it in Python in the signature and this is another example of visualization actually macula busy are rendered in a back-end side but if you have a front and side visualization like Java Script visualization cause then you can put your JavaScript inside of no tool so you can visualize your data so this as I said this is all fake data so this starts shows where the earthquake was a happen and again we can leverage the dynamic form here to change the date range so let's see let's let me put dynamic one here then any business user can now simply change the range of the date let me see from 1996 and we can see a lot of earthquakes and we can you can see how the plate looks like in the earth right so if I change this one this notebook to the report type it becomes a shareable nice clean report that any any people can consume so it I think it it really shows how our simple how how deeply make interaction between engineers and business user simple and easy right so let me go back to my slide ok because Paula and the question is how do I use a matte lip for the data I loaded from this color if you can can you lipstick Asia are you saying that you the way [Music] yeah I think it okay the question is what is the recommended way to map will live if I'm if I'm using our Scala so one way is like it this example this demo lost data from the Scala and do some processing and pass the results to the Python side using Japanese feature and then read that data from Python and draw mackerel the visualization or I think there is a project that provides some Escala API for Maithili I think it's it developed one of the people from Netflix I think yeah yeah so I think it there are a couple of options and I hope that yeah hope that it comes out I don't know you know well integrate with exactly in the future okay let me is there any other question about the demo or let me go back to my slide so yeah interpreter to another is just written to disk in the next one okay the question is how how can data how data can be passed from one interpreter to the other interpreter so one way is it right the airin into the disk can read from the other interpreter the other way is a plane I have something called resource for there is a distributed you can think it as a distributed map among all the interpreters so if your data is are small enough to fit it into your memory and you can put your data into the resource pool and then the other interpreter process can read it okay top custom is it yeah if I didn't see shareable using resourceful right yeah so if because of its interpreters are running as a separate JVM process so on object need to be serializable to you know read read by the other interpreter so how did the I if we I think I don't think it's a serializable but if if any object is realizable and you can read it from the other interpreter an interpreter I wanted to talk a little bit about the interpreters so interpreter is a back-end integration on abstraction layer for the backend in the plane and in the beginning of the project we had only three interpreters which was a spark and a markdown and our shell I think in two years ago but now we have more than 20 interpreters and there are even third-party interpreters on github so there are much more choices over back integration that you can use and like I mentioned the interpreter is basically running as a separate process and communicate with each opening demon using a strict messaging and there are three different modes that you might in interested because of that's it related you to how the plane integrated with the spark and Scala so there are three different modes of interpreter and our first one is a shared mode and shared mode is one interpreter process and there are interpreter group inside of a process and it eats interpreter group have multiple interpreter or instances and a single processing single interpreter group service all the new tool and the second mode is a scope mode in this mode still a single process but multiple interpreter group service each individual neutral and the last mode is isolated one and we see create individual separate process per notebook so now let's see how SPARC interpreter leverages these three different modes so first our shared mode spoke interpreter create spot context and the spot context is being shared by spark and spaghetti pie Spock Spock our interpreter inside so on although you are using Python although you are using Scala you can see the same you can access the same table you have registered in the scholar context and your job submitted by spa or top submitted in a PI spark was spa our spark and scale all they go into a single spark texts and and also you have a single skull on a pole so that and this process service all new to book so that means if you define or let's say you define a variable Twitter in neutral a and include two P you can leave that variable and you can update that variable so sometimes that expect you to behavior but sometimes that's not a not expected behavior so the scope mode Spock interpreter runs a little bit differently so still or there are a single spark context but each module will have now their own skull Aleppo so it's not work will have their own namespace where they are variable they are not conflicting each other they cannot read each other update each other but all the jobs up we needed from all these new to and all these interpreter will go to the single spark context and the table is scheduled by fare schedule on inside of a spark context so that's it scoped mode and isolated Modi to look like this so eating it's no two will have their own spark context and their own color ripple so that is three different modes and you I think you can choose three different modes depends on your workload or how you want to use so that's it how the plane integrated with the spark and color I think so any questions so far yeah okay let me go to the other components so interpreter is one of the component of Zeppelin and that like I mentioned before in the beginning there were made for interpreter three interpreters but as soon as we make it palatable community contributed a lot of interpreters so now it became more tweeny so we expected to that the same thing happened to the other other module inside of Zeppelin so no true stories are not on pluggable components that are Zeppelin hat so basically you're new to is it persisted in your local file system but sometimes you want to put your note to in a shared storage or your own companies the other sturdy system that have ovulated to a person control or backbone so the plane made the pluggable layer for the node to stir it and now we have a git repository git no to Aleppo and s3 and Ezer and they're clean how as well and you can also plug in your only path tree it's the API is really simple so you can create your own integration in a few hours I think and another problem module that we are still working on is visualizations and although sapling you can use your own library like Maithili or any any JavaScript visualization library inside what you're not true if using routine visualization is actually much simpler because you don't need to write any code you just need to click the button right so we want to have our visualizations pluggable now it's a hard the list of visualization built in is they are fixed but we want to have a pluggable layer for it so eventually we want to see our interpreter is probable and not two stories is probable and visually make visualization pluggable as well so it's a working in progress and so our pluggable interface or probability is one big topic in Chaplin and that's it how we expect more contribution to the project and another good another big topic of the cleanest enterprise the pictures we will now see more and more users using user discipline on their production or on their enterprise workloads so I think it is a plane from zero point six I think definitely now have a basic authentications and authorizations I think it dis awesome authentication and authorization feature especially this authorization Prasad here are from Twitter contribute today and I think that's it that shows the where that plane goes in the future so chuckling we wanted to have a enterprising level pictures so the plane can just can be used in the enterprise without custom and usual by some enterprise personal promotion personal destiny so Zeppelin now supports authentications and authorization of notebook so you can create a user you can integrate the pin video LDAP or Active Directory and you can let your user user login and it's it's not too can be accessed controlled and we have with more roadmap on enterprise side actually so except for authentication and authorization there can be more enterprising library features especially multi-user support like for example like impersonation is it we are working in progress and we are also working something called a personalized mode which is its individual Zeppelin not too if you have uses a plane then one of the cool features Appling is what you see is what others see in their browser if you change the type of type of graph then the other brow other in the other browser session immediately see real-time changes sometimes that's cool but sometimes that's not you want you want to provide personalized selection I want to see the data with the pie chart the other people want to see the same data with the other type of chart and even with the different parameters some I want to see from year 2010 to 2015 the other people want to see at the same time in different ranges so we are currently working on it and also a lot of people are asking how can how deeply no two can be included in not they our production pipeline so once user made some new tool in Japanese they wanted to include that code in the production pipeline so we are trying to make some progress on that side so top management easy hopefully we can do some useful feature we can provide some user feature for that use case for example event 2 who can be one of that feature also you can integrate a plan with your production pipeline and another one is a table data processing engine so as you see Zeppelin automatically shows built in visualization like your output is a table data so we we wanted to improve that little bit more so we eventually that's it I think a kind of little bit long term goal but we wanted to improve our this table data support so one we want to make sapling not only for data engineers or scientists but presidency we want to business usually do more job on data with their mouse click not using code so that's it I think what table data processing engine will provide in the future and of course we wanted to have a better support for Python and our I think it that please scholar support is really I think it's not bad really good so but our Python and our support which is also popular among our data scientists which is I think it's a little bit behind so we are really working on it and sapling is written out the front and side of Japanese written in angular based on angular 1 and as you all know angular 1 is a you know end of its life so we wanted to improve our UI as well our UX as well in the future and the one of the the best thing of Zeppelin is a plane is open-source project so this is actually my own agenda and you can use you can add yours and you can participate the community so this is the users and contributors and companies who do business with step are using the plane and some technology integration so it's happening have a lot of users and contributors like Twitter and even a lot of different companies our company I have Zeppelin in their product for example Amazon heavily also putting on their email and Google provides our script for their plot the data Pro and Microsoft Azure have a certainly not to be inside so the community already being adopted in many commercial product and there are a lot of users and technology integrations so please this is open source and you can you can participate so please doing the community and yeah I think it that's it so thank you for this and if you have any more Kisan please is is a question challenges around the iron our interpreter yeah so actually tackling at the moment have two different iron interpreter implementation one is the inside of us Park interpreters history the other is on separate as a completely separate implementation but they are basically they're doing the same there are some history but I mean there are two duplicated implementation for spa are so I want to change this one to look like Python case for example Python also definitely have a two different Python interpreter implementation one is a spark PI spa which using integrity the fully integrated with the spa the other is a pure Python interpreter written works without spa but so I want to in the future I want to see these two our interpreter integration one becomes part R and the other become a pure our interpreter yeah we don't have oh I cannot really say that timeframe here but I I would say if there are more usually men stand more contributor in the community will I spend time on it so if you have a cousin or your requirement or request then you can create authorize you you can write an email in the mailing list and contributor will see and change the directions I think [Applause]