data.bythebay.io: Konstantin Boudnik, Data in the Apache "big data" ecosystem
Recording: data.bythebay.io: Konstantin Boudnik, Data in the Apache "big data" ecosystem
thank you very much um okay only may do this and I'm not a data scientist I'm not an expert in machine learning but this is a data conference I'll start with a joke right so the statistician walks into an average bar and the bartender saying like you know we don't serve your kind here and statisticians like you know it's mean so anyway so I'll going to be talking about what helps to move the data right so I'm not going to talk about the you know data algorithms or data frame works or how you can do the optimization of your data pipelines I'll be talking about something that usually is hidden behind the scenes and helps to achieve your you know ultimate goal of crunching the data right so i'll be talking about in infrastructure about the projects that helped you to do this stuff right so I'm a little bit about myself I'm the CEO and the founder of the map correo we do in essentially one-click cloud delivery for in-memory convenience tax so basically you can do all up royalty p streaming messaging whatever it is with a single click on a few cloud providers right now so you can provision the cluster you can pause it is humid to save the costs and you can do all sorts of data in jest and processing at all anyway so with my essential Big Data Apache guy head on let's talk a little bit about the big data right and in reality this this whole talk should have been should have been called size doesn't matter but you know big is probably better right anyway so how we got into this whole big data mess right so information explosion right so we started all of a sudden we started seeing people wanted to collect all sorts of events and all sorts of the data pieces and beats from everything right so like Comcast would want you collect all the all the data from there you know septic TVs or set set of devices from their little modem and routers and what's not and they would get into like gigabytes of data in take on a daily basis for it and more of these companies will will come to do to the existence and all of them would say you know we're not sure that we can actually analyze the data and make the right decisions based on our analysis because we don't know how good the analysis is right so we cannot use the enterprise features that out there already because they too expensive they cannot scale actually to you know whatever the data sizes we're dealing with so um let's figure out what to do with it right so we can notice yell it properly because again there is no tool set and stuff whatever we'll buy from Nikki so what's not doesn't cut it because you know it's expensive the updates are coming you know and not that this is satisfying piece and and again I'm not picking up on you teaser here but in general the business purdy the solution providers enterprise solution providers they did not react to the market demands as quickly as sometimes in it so and the most important of all is that cookie cutter solutions are not feeding everyone right so they they they actually sort of like the same however customer might need variances of of the data flows of the you know data patterns and stuff like this so and here come the glorious Apache Hadoop and the whole could do package system and it says you know what everything is fine it's all downhill from here so no need to worry about anything you can store as much data is unique you says Apache Hadoop let's know schema bother you because you and process it without any schema whatsoever says whatever Apache Cassandra you can scale it anywhere says MapReduce framework but we already heard that MapReduce framework actually has a few snuggles here and there so and then we've heard that once you have all the data you can actually figure out everything about the the data and your customers and what's not and here comes the machine learning but there is as we know lies damn lies and statistics we probably cannot figure out everything about everything at any rate either suddenly enough it's not that big right so if you look at six months ago soory from the Forester they did online survey of eighteen hundred companies eighteen hundred plus companies they figured out that literally there is a lesser than ten percent of the company is dealing with the petabyte scale data sets I said like three companies here of course it's not three but you know a little bit more than that but still there's very few companies that deal with the huge humongous data set so there is google there's facebook there's yahoo back in the day actually at a yahoo hadoop cluster we had whatever for four and a half thousand nodes at twenty one petabyte or something like that right but again you can probably count these companies with one hand so about seventy-three percent of the company is actually dealing with the datasets under 99 our turbines actually turbines yet durable so it's like what two two boxes the rack to have all your data so what was the point of having the whole cluster doing was that they doing was the data so if you if you can't have two boxes and an average data size or every enterprise data size is actually under 80 terabytes right so it's like my laptop has whatever 250 actually gigabytes okay so it's a little bit more than my left but nonetheless so it's not that huge right so the price of this storage is actually when significantly down the price of the memory when significant role down so you don't need to build these huge solutions usually read huge infrastructure setups to to work with the data you you normally dealing on a daily basis and interestingly enough despite all the hype around like no no sequel and all all the stuff so schemas are still highly relevant right if people are still working with the data structures and structured data so in the same sir way they shown that actually more than thirty seven percent of the data in enterprise data sets are represented by the structured data right so it's still easier to work with the structured data it's still more you know customary to work with the structured data and of course the rest of the data which is literally you know no sequel kinda no not not not structured scheme only kind of stuff you can work with it with pretty much any any tool out there but the problem is as soon as you start calculating these little tools with one another you need to figure out actually how to make these you know connections how to pass the data around around the component boundaries and you know new projects do not speak the DL anymore so you need to figure out actually how to how to deal with the stuff you keep the data in the storage but there is no state on this data right so there is no state on the objects so you don't know if the version was changed from yesterday or not right so I mean like it's good if you have the HDFS filesystem which is sort of like read-only right so you cannot really over read something you can attend the Tran kid or completely remove the file and replace it something else but still the versioning of the data in your storage is is an important characteristic which you know sometimes it's not addressed and of course there is like serialization mechanism like like a ver that that carries around version along with the data stuff and it's not but again it's all it's all kind of loosely coupled and not necessarily solve all your problems and what are those problems essential right what needs to be solved is it too much data is it like to slow process them do we need to do updates do we need to do transactions on a data right so if scheme on readers to too difficult is you know schema on right is distinctive and you you risked enough losing some data so does yours sequel performs well right do you have a visualization tools bi tools be a tools you know that kind of stuff so what you're trying to solve and surprisingly actually you have to solve all of these problems right so if you if you try to build a versatile enough data analytic platform you have to deal with seneschal with all these bits and pieces right and then you start looking around and check what the 22 sets you you can can use so and here's the apache comes right so Apache has this extensive set of the tools that actually are focused on dealing with the data processing right and and of course this slide particularly is not pretending to be complete but just just a few pieces so you wanna steam and then store only you know what you need so you can go and do something like flink ignite kafka for the messaging spark for for micro Beijing red and then my Christina or microfiche screamin if you will and now we have a higher level frameworks and there's the SDKs like Apache be incubating rent so do we need to do something with the real time right so close to real time there is a passionate night there's a patch in each base if you need to you know really externalise your schema and be able to share this kima somehow well there's H catalog there's there's Cassandra in that kind of stuff so do we need sequel like I you hardcore sequel guy you don't want to learn you know God forbid skull Oh something like this so sure there is igniters hog there's Phoenix there's spark again you know sparks equal in this case stuff like this doing it decent visualization tools sure there's a partial settlement again right and there's an interesting story about the Apaches Ethel and wish I was mentoring in opposition incubation wife incubator a year ago so at the apogee big data conference they believe in Texas in Austin a son for that was it no it wasn't Budapest was here anyway so it's irrelevant but it's funny so the discipline guys they were have in the presentation and they were shown you know nice clean shots and you know some demos how they doing this and that and what's not and somebody in in the audience actually stood up and said like excuse me is it like open source version of the data bricks UI and we throw with the joke caught up so quickly that we can my flag at every Zeppelin talk or every spark talk somebody would raise the end and said like by the way could you comment on the fact that the Zeppelin is the open-source version of the day the bricks you guys that's pretty funny but anyway so Zeppelin is actually pretty cool projects i encourage everyone to check it out if you if you haven't here so angry so let's look at something like this um there is you know components on the left and there is you know characteristics of the system at the components on the lab sorry yeah I know so and the characteristics of the system horizontally right so and I'm not pretending to give you like an exact technical advice or how build your data pipeline and all the data analytic system but essentially you can see that if you need transactional sequel support there is not as much component that can help you who is right if you need a real time or course the real-time data processing there is whatever two components in the whole zoo right so if you need machine learning there's a couple of them and so on and so far right so basically having having this this kind of approximate matrix you can figure out actually wear and what you'd like to use in order to gauge their where you need to be right so and again it's all kind of hunky dory it's nice picture and all but the problem is you start facing the exponential growth of complexity red so all these little projects they have different requirements they have completely uncontrollable you know proliferation of diversion so each base goes at one please Spargo at the different pace and you know they're not necessarily Cardini can lose each other in terms of the API is compatibility and binary compatibility so basically you start seeing the impedance mismatch at the integration points right so the civilization needs to be addressed because they don't talk the same languages and they don't use the same civilization formats and as I mentioned the development is actually sometimes crazily fast right so you can see like hundreds of commits amounts in the most active projects right so how you can keep up with that this is unimaginable and the release drains are different so how you how you deal with this complexity how you can actually navigate through this and fortunately Apache has an answer for you so this is Apache big top or this huge ten that covers the whole zoo for you write the whole circus and it actually there are 22 exactly provide you the the the help with all the complexity red so Apache big top is essentially the system the framework and the philosophy to deal with the complex tax right so you can specify your your software stack your components diversions whatever it is and you can actually convert this tech definition into standard Linux packagin which would be accepted by everyone and anyone and while going there it will guarantee that all these components into your stack I actually compatible with each other they guarantee to work with each other they compatible with you know at the API and the binary levels red because we have this extensive integration and system test series and there it also provides the open open deployment interfaces and specifications so basically we can hook it up to any orchestration system in order to get your solution on their residence not it's not enough to help the packages right it's not enough to have the jar jar files and tar balls you have to get it out in order to make it usable cool so and that's exactly what the different dozen and because of that it's very widely accepted by by the industry actually you know the Hadoop distributions are using this all of them in the innocence well actually with the role of it and how you can get it to spin if you want you so it's actually really easy right so the Apache big top gives you essentially the deployment mechanism that works at pretty much any machinery right so you can do this on vm you can do this on docker you can do this on lxc you can do this on cloud anywhere right verb 1 hardware and stuff so you can use third-party tools to do the modeling of your cluster solution with big top so for instance canonical Ubuntu juju lets you to lay out the model for your data processing stack which will effectively hide out all the complexity of the deployment and provisioning knowledge for you from you read so all you need to say is like I want to have hae HDFS name node blast spark plus i don't know Zeppelin coo coo button get yourself a cluster on whatever 20-plus are different cloud providers right so um we are work in actually or contemplation rather the the way of integrate d ck yo engine to be the part of the Apache big top so you would have your own orchestration asian so you don't need to rely on you know commercial tools and what's not to manage this sort of deployment into the cloud right so in fact the the company I'm doing the mem choreo we using this decayed the very dedicated your engine actually two to do the provision so and of course if you need to you can have the blueprints to coordinate with things to to be deployed from operational value whatever reason right so essentially you have this massive stack of the software which is whatever 30 plus components right we should tested on seven plus different linux distributions on three different hardware platforms rights its x86 its arm and open power now right so you can actually choose and pick which in it so and interestingly now speaking of data so a partial big top is actually having the storage variability are built in red so the HDFS camp concept is actually highly supported so Hadoop compatible file system so you can go with HDFS you can go with Gloucester we can go with qfs you can go with this and that we thinking of air and seven s3 and that's pretty much it so let's talk about questions yes there's not that much relation so for us UDP is yet another downstream like whatever cloud there Hortonworks and in fact actually I did the work to produce the the first version of the ODP I reference implementation earlier this year using big top so it took me oh gosh these days actually to do this well the LM I'm consulting them so they their client of mine but yep so the user knows as a framework and deployment mechanism yes how the top Achebe top is different from a pajama party it's very different so a pajama party is essentially a cluster monitoring software or if you go a little bit wider you can say that it's an orchestration system which is not exactly but anyway so Apache big top is deployment development and testing framework for the software stacks so I'm Bari is a tool set that allows you to deploy and monitor one particular stack namely Apache Hadoop stack so with big top you can actually build anything you can build a lamp stack you can build Hadoop stack you can do whatever rest is the universal framework actually for the production of complete complex tax and that's pretty much summarizes and cause the session thank you very much