Devreal

sfspark.org: John Leach talks with Alexy Khrabrov

sfspark.org: John Leach talks with Alexy Khrabrov

Recording: sfspark.org: John Leach talks with Alexy Khrabrov

hello everybody I'm Alyssa cover the organizers offensive spark in France and here are on occasion splice machine we have a very interesting talk tonight which is marrying spar with some real time analytics and here ones who have joined each who is the CTO and co-founder of ice cream thank you thanks gelinas very excited to learn about machine so maybe you can tell us a little bit what is closed mention how did you come to co-found and then once kind of what is the need in the market for for the product availability yeah so a lot of times products come from some great innovation or some great idea that you had and this product came out because when the basically Hadoop came to market about six or seven years ago large companies were really excited about it and they said conceptually they looked at it as a database but in actuality was a file system and it was a file system with no concurrency model and we I struggled personally trying to bring that successfully to market at different mostly fortune 500 type companies just didn't work they wanted it to behave more like a database they wanted to load data all during the day they wanted to do things like deletes and updates they want to do constraints all this data management stuff that you know probably isn't terribly interesting to us but you realize when you don't have it and you have a multi terabyte data set or petabyte data set with no constraints no governance that's not realistically something you can do so based on that failure and that there was a really pertinent question someone posed to me and they came up to me they said look we have this massive amount of data and I've got these applications that are struggling that are under tremendous duress but you know what this Hadoop I need someone to say look I can run my campaign management software on to do I can run my order management software on it I can run my internet of things application on it I can do all of that have all the benefits of the ecosystem and all the great intellectual property that's being generated and cool ideas but I still need it to function do constraints do the things that we expect and when that was posed to me you know Monty and I kind of came together with gene Davis the third co-founder we said it's a great opportunity because Hadoop can't be over here at applications over here and you know Monty had an experience with rocket fuel where they were using HBase and real-time doing a lot of real-time optimization and you really realize that's nice for a company as sophisticated as rocket fuel but I want distributed computing to be available to everybody they always you know we always hear about they face because to know people wasn't a tune it right like oh gosh yeah yeah it will talk a little bit about that too that scene on the side that the new spotters ways and never stopping means that's not an easy beast fitness and takes up the company to manage it properly yeah in what we've done a splice machine is to really think about all the problems we have in the problems that we're solving and we'd really try to isolate them so when we look at H face a great oltp system but it struggles when it tries to do analytical processing more compaction all these things that were problematic and it's splice machine we didn't accept that as a reality we said look capac shins are a problem in HBase well you know what we're not going to run them in HBase we're going to run them in spark why should your transactional system try to read massive amounts of files and merge them together would you do I mean does that seem go here that's not it was actually all processing exactly so for us we look at everything say look here's oltp and here's all lap that's a gross simplification but really it's about is there something that needs to go through a block cash go through you know different structures for optimization for transactional processing or is this a destructive process ie an analytical process sometimes reading a lot of information in performing computations that are expensive if that's happening put it in a system that's appropriate for that and make sure that you can both isolate oltp and OLAP and inside olap you can prioritize all the jobs that are happening compaction you know should a compaction be prioritized over the CEOs query probably not should compaction be prioritized over index creation that's a good that's a good thing for an administrator to determine well which should be prioritized so that's really our strategy is to isolate oltp an OLAP and why were here talking to you guys is we really determined this is over two years ago now that spark would do that and Michael Franklin is on my advisory board he really helped us like conceptualize and you know what it is we're doing here so when we look at it it made sense to us the investment there spark has come a massive I mean huge amount you know the sophistication there is just you can see it's snowballing that's exciting for us and you know we've been like I said we've been at it for two years this is a really that's kind of far ahead of because enough I mean I tried on a tripod you know Sparky cloud 2012 frame that was a very wrong thing oh gosh you know people you know where she'll tell me like I don't know about the smart thing I don't know about this kind of thing your worries right yeah and so and so recently I read that kind of mess of acceptance happened but I was a two years ago it was still not such clear views and michael is a great guy can feel I know I block folks in this I think this is really interesting but I wonder how so so you had this portion by founder experience before I saw and it wasn't a good one to be just be honest it was it was you know I'm one of those guys admits I failed a big data initially like it it didn't meet really yeah it may have been a successful project for some people but if you talk to the business executives they didn't get the value out of it they needed to and that's just not acceptable because this this is what's going on here is going to change businesses innovation but but we have to get it to a place where people can go beyond just data management and start actually doing advanced algorithms smart application which Monte that's his passion building smart applications with machine learning real-time data I mean that's that's inspires me with the way he communicates to the whole team here at splice machine about we are going to build the next generation of smart applications they're going to be built on our platform this is really exciting to me right back okay no we receive this data pipeline Samaritan wearing all from inception from api from we lose my home again on IDI then they update the kind of person through right and then they are irrigated and there is some of the process happening spark and some some kind of action happens in it is to happen facts right so correct sounds like what you guys doing really it's got its couple but I'm kind of I want to kind of follow up on this value match because a lot of a lot of developer the trenches they're not realizing right that like decisions are made by this executive center of this is what this usually it's procured by company which has messaged me to come a nail and so so and it's a religious ability to imagine you know you want a little value to 2 to the businesses right so how do you compute comedians and researchers and researchers i mean that's like that's going to be fascinating and how do we how do we tell them about this value right how do you come to commit I I can I couldn't understand researchers dig in this but how become any evening shoes to you know to executives basically need to make a choice do they do my oracle or I loosen up a face do it yourself you know then the bricks about sure how do kind of community the the the key points of what is different about yeah so our differentiation is if you have an existing like Oracle RAC or Oracle Exadata what we would do is say if you have applications running on that that's a key qualifier for us you're not just kind of throwing data into a file system you have active applications you have users you have a business owner who's trying to get some value or run something that's driving their business and what we do is we see oracle or another relational database out there and people want this data updated more frequently they want to be able to run both transactional and we call it as kind of real-time analytics operational analytics if we see that that's really our sweet spot is to going that initially it's been mostly digital marketing their early adopters they have a ton of stress they have existing packaged applications that they've already invested in whether it's unica redpoint those types of products great products they're all sequel based and they're transactional and you know what that's really our focus because they have a tremendous amounts of data already they want to add Twitter data when I had all kinds of different data sets but that expense has to be commoditized yes it cannot be you know this is going to be a five-million-dollar system they'll say whoa whoa okay you know a couple of thousand i can get it but five million i'll just spend that naughty no what happened hi mark you know so so for us we're really focused on bringing applications and what it was an application me it means transaction ality it means that things are being concurrently updated and the database has to apply those rules hmmm interesting so yeah reminds me of multiple at tech companies I know they kind of M&A Liam they can observe Oregon they would install a gigantic like 80 even the road ace you know terabyte drive machine or get stolen vessel the 70 million dollar type of the pier it would keep logs in there right and go and when I was resolved years ago basically everything was open source except for oracle handling money right because it will entrust the most important transactional system to handle money right so in a manner so everything you know Tasha morning was Oracle so even but even when you have those transactional systems that operational data store that next layer it still has all of the data management components in it and then then you have the spin-off columnar type read-only optimized stores and things of that nature in the analytical systems or machine learning like a spark reading from that or having its own environment the case of SAS or something like that but I think work will still cemented in that operational data store which is a very expensive it's not the actual app in some sense running the business but it's it's still a very expensive proposition and you know that's something that we attack pretty aggressively yes yeah it's not the source guy right there park is in here another source and you know the great thing about that you know fashion scarf is always going to be open source write the correct editor is this correct so so I can easily see you know if I need some lesser amount of stuff to memory like a draw the US open source tools for this so would you say that you can or you can provide another section processing for financial data which is currently you know can be stored on oracle or and a movie rhetorical and kind of be the operational the dementia can you take on the whole you know suite of transactional processing well we run the TPC see benchmark so yes we would love to do that but I'm also realistic right if I come to you and say I'm going to replace your trading platform you're gonna say but if I say look every year you're not gonna have to write that check to Oracle for your operational data store and here are the suite of applications or the larger data sets there and hey can we transition that on displacement and cap your Oracle spent that's interesting to people and by the way when we built our system we built it in a completely open manner and what I mean by an open manner is we're a relational database that allows you to in parallel read in you know read and write data to our system from hive Impala spark all of our sequel will return for the spark community it will actually return all the instructions it won't actually return result sets our result sets are actually rdds they're actually data sets so you have all the instructions you haven't even run your sequel and you say here that other manipulations execute it and the whole package is in a resilient structure of these the instructions need to be run to perform whatever operations you're running from an analytical perspective nice yeah but it's also makes a lot of sense to me because you know oracle is compliant so you can entrust kind of the kind of final you know financial do this but i don't see a reason to basically store logo operational data in it right back to you if you can use more efficient with other systems to do this ah so so i guess you guys are usually sigh you using this colorful this no sir it's good question so we're mostly a Java shop and then we've got obviously some functional with regard to how we integrate with spark with Scala there but mostly jbn based yeah you know our focus has really been like I said isolating the workloads resource management on the analytical side we've spent a lot of time with you barn so spark on your art is something that is just second nature is that's how even developers they run little yarn apps with spark in them when we do it so we're really focused on how can you get the repeatable results because that's really where confidence comes in the database mmm interesting yeah this is this isn't just because you know I all of that same friend survey in terms of the skull come when you then you don't do that everybody surprised that you know the majority of the world was basically even a lot of loan on scholarships Java shops are hanging spark right i mean all the data sizes using photo so it's really they can hold this kind of this common way to have a lot of data in memory yeah and i don't want to pick like for us we support sequel you know and that's our strength obviously the applications written in java and all that but from an external point of view you can use scala you can use the spark client Scala and do that an issue sequel and get those instructions back but the cool thing is when you write data from spark and a splice machine it's still going to do triggers it's still going to do check constraints it's going to do your foreign key value that validation there's no second-class citizen so tools that use that we integrate with they have all the capabilities that we have it splice machine when they write and read data they read data consistently so if there are big updates going on and things like that they'll see a consistent view of the database that's absolutely critical as we get these more complex environments where data is coming in at all the different times it was an advanced dtl environment or an application environment it's absolutely critical so it's if that's what's exciting to me because now I can start to see how can you have really complex applications at scale and yet still have an R dbms where the knock against our DB messes mine right there they were like a data it you only the sequel in there and if you want to get it out it's a result set or ODBC JDBC and it's mine right and it completely isolated all other products from it and we said no no that's that's a fundamental mistake spark should be able to go in parallel and splice machine and access our data natively off of HDFS is kind of a common common piece and that was we're passionate about that piece because there's so much innovation going on not just spark but I mean there's a whole host of different machine learning algorithms that people are developing you name it all right that's absolutely system and you know you basically get the benefit right exactly exotic so this term and this is this is great here I mean this is kind of you know what folks in the limit ops you know a passion poblet basically know how this open-source technologies will have been different different errors I would ask you for European about this because you know what I noticed you know the evolution of spark is actually it seems like it's evolving towards a database so and if you look you know they started the basically roar you know already needs right and then they the frames appear right and then the project hansen wants to push it down and optimize it and so essentially what is happening in you know hundred percent of our users are using sparks equal so basically the sequel Jeff Locke died right like it basically you know it will ring her name you need a new post right and it will show up and basically so so what happened is you know the market basically is driving spark to provide sequel again but now it's different kind of sequel right this is basically necessarily not satisfy constraints but it looks like sickle or I think it was the computers of it you know what do you think about essentially and it seems like since calm percent of our users the business part signal so then they're demanding basically that that it is optimized and the unity frames are going crazy right do you think that it's kind of where where the market is driving spark and that we know the data frame difference will become effectively this kind of in-memory database what was the government's yeah I think when we talk internally about spark we call it a data flow engine it is technically what it does it applies relational algebra it doesn't focus on the storage piece that's very abstract right deliberately because it makes it very flexible there's great characteristics of that but when you talk about database the sticky bits and when you get down to the storage right yes that's the concurrency part that's the you know how do you create an index and still be transactionally correct when people are writing into it the same time you know all these really difficult relational database and transactional sort of questions that's where we focus the splice machine and why we're I think very different cuz we focused on that piece where what sparks doing it and I don't to discount it but they're following the same thing that like hive did right I was just hey we want to be able to use this but i want to write java force let's not do this right it's equal is something everybody on new and there's a lot of legacy code for you know actual sequel apps so i think it's interesting but for me the key is really the storage layer because from and for me it's like well what kind of applications are you enabling if what's absolutely critical for splices i want somebody who can run you know a marketing application some sort of real application that affects the business without having to build it from the ground up with a hundred engineer you know that can just even take just a packaged app off the street or if someone says you know what we have a company that's a healthcare company building smart applications that's amazing to me because then you're starting to see people realize i can build really complex apps the dude are dbms hmm but same time i can take advantage of classification algorithms and spark some of those things and you start to say wow this is going to be cool or the atom libraries in spark where they're trying to do genetic processing you know all of these things i'm like wow we're starting to commoditize data and we're trying to build these applications that we really wanted to build when we started this is when we start whenever I did a Hadoop and started sparked it was it was really about can we build these applications that are interactive that are smart that give us what we want when we want it and why didn't it happen because we got bogged down in the distributed file system we got bogged down into just the data management yes we really did in the differentiation the companies are going to finally achieve is ok we have the data management piece but I want to work on feature development for my algorithms developing the right features and measurements to build a feed to an algorithm when they get there and they're not talking about gosh my data is corrupt doesn't you know they can get out of that world and they can get into the other world of actually applying algorithms and our topics are not people trying to do indexing strategies and stuff like that mmm but really our topics are what algorithms appropriate what features are appropriate that's when we know we've progressed in my mind to building these applications that everybody's starting to expect but they're really expensive right now for people to buy interest and it sounds like really really new strategy so so when we was going to wrap up with some predictions eyes like I think you're not right but it sounds like really available I hear you know we filed it right and I and all the other people when like they're building genomics closed and there are multiple genomic startups in the avocados conference call David by the way coming up in May and then multiple verses and people to have life sciences and kind of you know if five senses folks can take this little problem wheel right so so where do you see yourselves in in the next year we have industries you know what kind of setup sounds best to use a sponsorship makenewcastle perspective companies the customers kind of seen this you know who do suggest and is a good pilot who is kind of the early customer connection yeah so digital marketing they're always early adopters people with large operational data stores that they're probably still in an Oracle RAC and it's probably stressed a little bit but it's sizing Internet of Things type applications with high concurrency and you want to be able to do operational analytics and maybe even drive an application you know we're all starting to wear these smart devices I think that industrial Internet we've started see that whether it's cell phone towers those sorts of things are very latest since yeah that's extremely interesting to me so I think there's a whole host of those types applications and those are the things but you know to be honest we're really just focused on our customers I mean maniacally focus on making our customers successful in that's you know we partner with our with our prospects and customers only focus on making those success for whether it's you know genetic information marketing information we really want to just focus there on customer Scott so that's how we're going to be successful cool so so like you know if we win next year we'll make a baby what kind of civil affairs look how philip chang would make you know you happy in next year it's funny when we started the company montine i sat down and said what would make you excited about splice machine and we talked about things like we'd like to be able to say we save someone's life by providing algorithms at time in that sort of platform for real-time and mati talks about computing the uncap you tible and um computer was having a word and he loves that which is great and his point is and I really do concur with them as we have applications that can start taking real-time data but still give us analytical capabilities and machine learning capabilities I want to see those use cases in splice machine um and I want to see people start pushing it that's what would make me happy yeah it'd be great to replace a lot of Oracle instance is something I'm happy with that just you know from a revenue and acquiring customers but to be honest I really would love to see those smart applications because as a technologist that's really exciting all right so you know hope you know that we see a lot of smart apps and it sounds like a really just flattened to try this so I hope you know folks even you can give the spinner and see what kind us and was looking forward to doc thank you thanks